The shared storage on the main production cluster of a platform I run under contract is a Ceph cluster inside Kubernetes, managed by Rook. Every application there that needs a volume several pods can write to gets it from one CephFS filesystem. It is one of the most available storage systems I have worked with, and availability is what I designed it for.
Why not the managed one
The platform runs on GKE, so the obvious question is why I don't use Google's storage and let somebody else be on call for it.
The workload needs many pods writing to the same volume at once. For years each of the main applications ran at least four pods and went far higher during live events, with autoscaler ceilings of up to 200 pods. In late August one volume had 26 pods mounting it. An ordinary Persistent Disk volume is ReadWriteOnce, so that was never an option. Google's managed file storage, Filestore, is the usual answer for shared volumes. I costed moving to it this summer, and the saving was small against a large and risky migration, for a system that was giving me no reason to leave.
How it is built
I designed it with the redundancy inside the storage itself. The other ways to survive a failure were a second storage system kept as a fallback, or a second cluster running the same workloads. Either one means a failover. A second copy of the data has to be kept in sync, which either slows every write or leaves the copy behind, and the applications or their traffic have to switch to it at the moment something has already gone wrong. That switch is rarely exercised, and it has to work the one time it is needed. I judged that a bigger risk of downtime than Ceph itself, and a waste of resources for what it buys.
The cluster has five OSDs and five monitors. Each OSD runs on its own node, with a network disk that follows it when the node is replaced. Every object is kept on all five, and the pool stays writable down to three copies, so it keeps accepting writes with two OSDs down. That is more replicas than most people run. The extra copies cost capacity, and I accepted that to have room for two failures at once.
The filesystem's metadata server has a warm standby on a different host, replaying the active one's journal. On a planned restart the standby takes over in seconds.
The clients mount CephFS through the kernel, so a mount does not depend on the CSI driver pod running on its node. I checked that on a live node by deleting the driver pod under a read loop and a write loop: 780 operations across the restart, no errors.
Rook itself is outside the data path. If the operator is down, nothing new gets reconciled, and every volume keeps serving.
Five copies protect against losing nodes and disks. They do nothing if someone deletes the wrong thing, so the volumes that cannot be regenerated are regularly backed up to storage outside the cluster. The ones that can be rebuilt from a release or a repository are left out on purpose.
Before trusting it
Ceph came in to replace an NFS provisioner that had no redundancy and was a major obstacle to safe Kubernetes upgrades. Before it went into production I tried to break it.
I upgraded the nodes under five OSDs while copying data onto a volume, and then rolled a restart across the whole cluster during a copy. OSDs went down in turn, and both times the copy ran through with no errors and no slowdown that could be measured. Then I scaled the storage and application pools to zero and back, so that every node was a new one, and all the storage came back where it belonged.
Since then
Since that first installation, every full Kubernetes upgrade of the cluster has gone through with the storage available, including with the node pool's upgrade surge set to four. In a production upgrade in 2024, the cluster of ours that was still on NFS went down for five minutes, and the one on Ceph stayed up.
In planned maintenance, Rook's disruption budgets block the next node from draining until the data on the first is fully healthy again. This September I replaced all five storage nodes with a different machine type that way. It took fourteen minutes in total, one node at a time, with every placement group active throughout, degraded while each OSD moved in turn. The disks reattach to the new nodes, so only the writes made in between have to be caught up.
The same month I took Rook and Ceph to their current releases, in place, with the applications running. Rook's upgrade guide warns that storage "may be unavailable for short periods" during an upgrade, so I rehearsed all of it first on an empty replica of the cluster built on my own machine, took a copy of every volume that cannot be regenerated with file hashes compared against the source, and ran production one step at a time with a stop after each. What the applications saw was the metadata server pausing for a few seconds at each step. No pod that mounts the storage restarted.
What it asks of me
A managed file store handles the upgrades for you, which is reason enough for many teams to choose it. Keeping them means more effort, but when changes can be quickly rehearsed with a simple script on a local cluster, you can make them in a period you pick, stop between steps, and make use of the extra budget that buys you elsewhere in the design.
Running your own storage is supposed to mean being on call for it too. The alerts exist, on Ceph's health as on the rest of the platform. I have seen plenty of alerts from this platform over the years and none of them was about storage. It has been one of the least hands-on systems I have run in Kubernetes.
Did you know?
- Ceph is short for cephalopod, which is why the logo is an octopus.
- CERN's IT department runs several Ceph clusters with more than 100 PB of capacity between them, for block storage, CephFS and S3.
- DigitalOcean built its block storage on Ceph, and wrote up why.
- Rook became a graduated CNCF project in October 2020, the first one based on block, file or object storage.
Top comments (0)