Case Study

Scaling Mission-Critical OpenShift Storage Without Production Downtime for a Leading Public Sector Technology Organisation

Overview

Executive Summary

A production Red Hat OpenShift environment running OpenShift Data Foundation (ODF) required a significant storage expansion to address growing capacity requirements. The existing storage infrastructure was constrained by hardware RAID controller configurations that limited the number of disks available for OpenShift Data Foundation, reducing overall storage capacity.

Silvereye IT Solutions executed a phased storage modernization activity across all eight production nodes by reconfiguring hardware RAID controllers from RAID1 to Non-RAID (JBOD) mode, optimizing OpenShift Data Foundation storage configuration, and expanding the Ceph storage cluster without interrupting production services.

The implementation increased raw storage capacity from 209 TiB to 671 TiB, expanded the cluster from 15 to 48 OSDs, established a uniform storage distribution across all eight nodes, and completed the entire activity with zero production downtime, zero data loss, and a healthy production cluster.

Who We Worked With

About the Client

The client is a leading public-sector technology organisation responsible for operating one of India's largest digital infrastructure environments supporting critical government services. Its production Red Hat OpenShift platform hosts business-critical workloads where high availability, operational resilience, and uninterrupted service delivery are essential.

Given the scale and criticality of the environment, every infrastructure change required meticulous planning, controlled execution, and rigorous validation to ensure production services remained continuously available.

About the Client
The Problem

The Challenge

01
Production Storage Reaching Capacity

Growing production workloads had pushed the existing OpenShift Data Foundation environment close to its storage limits. Additional capacity was essential to support future demand without disrupting critical production services.

02
Limited Hardware Utilisation

The existing RAID1 configuration prevented OpenShift Data Foundation from accessing all available physical disks, restricting storage utilisation and limiting the platform's ability to scale efficiently.

03
Imbalanced Storage Architecture

Only 15 Object Storage Daemons (OSDs) were distributed across seven production nodes, creating an uneven storage architecture that limited scalability, resilience, and efficient resource utilisation.

04
Zero-Downtime Requirement

Because the environment hosted business-critical public-sector workloads, every infrastructure change had to be completed while production applications remained fully operational, leaving no room for service disruption.

05
Maintaining Cluster Health

Every storage modification required continuous Ceph health validation and controlled data rebalancing to ensure the production cluster remained HEALTH_OK throughout every implementation phase.

06
Protecting Data Integrity

The storage expansion had to preserve complete data integrity from start to finish, ensuring zero data loss while maintaining uninterrupted application availability across the production environment.

How We Did It

Our Approach & Solution

To minimise operational risk, Silvereye adopted a phased, node-by-node implementation strategy that ensured every infrastructure change was fully validated before progressing. Each phase included workload migration, storage reconfiguration, Ceph health verification, and data rebalancing, enabling the production environment to scale without compromising service availability or data integrity.

01

Production Readiness

Conducted comprehensive pre-implementation validation by updating storage configurations, resetting Local Storage Operator (LSO) symlinks, and cleaning stale Persistent Volumes (PVs), ensuring accurate storage discovery before infrastructure changes.

02

Workload Protection

Live migrated production virtual machines from the target node before every storage activity, maintaining uninterrupted application availability while enabling infrastructure modernisation without affecting business-critical services.

03

Storage Modernisation

Reconfigured hardware RAID controllers from RAID1 to Non-RAID (JBOD), exposing all available physical disks to OpenShift Data Foundation and enabling efficient utilisation of the underlying storage infrastructure.

04

Ceph Cluster Expansion

Expanded the storage environment by adding six physical disks per production node as new Ceph Object Storage Daemons (OSDs), significantly increasing storage capacity while establishing a balanced and scalable architecture.

05

Continuous Validation

Validated Ceph cluster health after every implementation phase, completed controlled data rebalancing, and confirmed HEALTH_OK status before proceeding to the next production node, ensuring operational stability throughout.

06

Final Optimisation

Completed the implementation with a balanced six-OSD-per-node architecture across eight production nodes, aligning the environment with Red Hat OpenShift Data Foundation 4.19 best practices for scalability and resilience.

Expertise & Stack

Domain & Tools Used

Domain Expertise

OpenShift Platform Administration

OpenShift Platform Administration

Administration and lifecycle management of production Red Hat OpenShift environments.

Software-Defined Storage

Software-Defined Storage

Deployment, expansion, and management of OpenShift Data Foundation and Ceph storage environments.

Enterprise Storage Operations

Enterprise Storage Operations

Planning and execution of production storage expansion activities with minimal operational risk.

Technology Ecosystem

Container Platform

Container Platform

Red Hat OpenShift Container Platform
By the Numbers

Results at a Glance

Silvereye successfully expanded the production storage environment without compromising service availability, data integrity, or cluster stability. Through a controlled, phased implementation, the project delivered a scalable, production-ready storage architecture while maintaining continuous business operations throughout the transformation.

221%
Storage Capacity Increase

Expanded raw storage capacity from 209 TiB to 671 TiB, providing the scalability required to support future production workloads without additional infrastructure disruption.

Zero
Production Downtime

Completed the entire storage expansion while production workloads remained fully operational, ensuring uninterrupted application availability throughout every implementation phase.

220%
Increase in Ceph OSDs

Expanded the Ceph cluster from 15 to 48 Object Storage Daemons (OSDs), improving storage distribution, resilience, and long-term cluster scalability.

Zero
Data Loss

Maintained complete data integrity throughout the project, enabling infrastructure modernisation without compromising production data or business continuity.

HEALTH_OK
Throughout

Validated Ceph cluster health after every implementation stage, completing the project with a fully balanced and HEALTH_OK production storage environment.

Takeaway

Conclusion

This project demonstrates that even the most complex production infrastructure can be modernised without compromising availability, data integrity, or operational stability when backed by the right expertise and execution methodology.

By successfully modernising a production OpenShift Data Foundation environment, Silvereye IT Solutions delivered a large-scale storage transformation while maintaining uninterrupted operations, complete data integrity, and continuous Ceph cluster health throughout the implementation.

The engagement highlights the capability to execute high-risk infrastructure transformations with precision, enabling organisations to scale critical environments confidently while protecting business continuity.

Work with us