A hands-on look at how etcd works, why it uses a flat key-space for hierarchies, and how MVCC handles state updates under the hood.
Episode 5 of K8s with Pravesh
Hola Amigos đ
Welcome back to K8s-with-Pravesh. In the first few episodes, we covered the holistic journey of Kubernetes internal communicationâexploring what happens when we run kubectl apply, diving deep into the front door of the cluster (the API Server), and clearing up some myths about StatefulSets. Today, we are leveling up our journey and exploring etcd, the storehouse of the Kubernetes cluster.
According to the official docs, âetcd is a consistent and highly-available key value store used as Kubernetes' backing store for all cluster data.â It focuses on being a simple, well-defined, user-facing API (gRPC), secured with automatic TLS and optional client certificate authentication. It is written in Go and uses the Raft consensus algorithm (which allows a cluster of computers to form a single coherent group that can survive individual machine failures) to manage a highly available replicated log. It elects a "leader" node to handle state mutations, which are then replicated over to "follower" nodes. It also follows an odd-number topology because Raft requires a majority quorum to commit data. Thatâs why etcd must run with an odd number of nodes (typically 3 or 5 in production) to handle split-brain scenarios and maintain fault tolerance.
In the context of Kubernetes, it is used to store state files in the form of key-value pairs. This data acts as the single source of truth. Everything Kubernetes knowsâits cluster details, deployments, services, workloads, secretsâis recorded inside etcd. When a user runs kubectl apply -f deployment.yml, the kube-apiserver validates the request and writes the metadata inside etcd. It has a watch mechanism effect where control plane components like the scheduler and controller manager constantly watch etcd using the API server. If a change is detected (like the creation of a new pod or a change in metadata), the new changes, if valid, are committed to the etcd database, and other components will implement that change in the cluster.
Interesting Fact: No component talks directly to etcd; all communication has to go through the API server to ensure security policies, authentication, and strict data validation before data is mutated.
Even though etcd actually uses a key-value databaseâmeaning it doesnât have real directoriesâit uses a slash-separated naming convention to simulate a deeply nested directory tree. When the kube-apiserver serializes and stores data inside etcd, it uses a default prefix for the root path (/registry). The hierarchy generally splits into two structured paths depending on whether a resource is namespaced or cluster-scoped:
Namespace-Scoped Resources (Most Common): It includes objects related to a specific namespace (Deployments, services, ConfigMaps, secrets, etc.). The pattern is as follows:
/registry/{resource_plural}/{namespace}/{object_name}
- Deployments:
/registry/deployments/default/nginx-demo - Pods:
/registry/pods/default/nginx-demo-6b74467d-9xyz - Secrets:
/registry/secrets/kube-system/default-token-abc12
Cluster-Scoped Resources: For global objects that exist outside any namespace (like nodes, namespaces themselves, cluster roles, and persistent volumes), the namespace segment is dropped:
/registry/{resource-plural}/{object-name}
- Namespaces:
/registry/namespaces/default - Nodes:
/registry/nodes/ip-10-0-1-50.ec2.internal - Persistent Volumes:
/registry/persistentvolumes/pv-data-disk
To see things in action, letâs spin up a Minikube cluster and look at how data is stored inside etcd in the form of slash-separated directories. The commands that we will use in this demonstration are available in the following GitHub repo: K8s_with_Pravesh. Inside the root dir, open part-05-etcd, where you will find the commands in the README.md file.
Once you have your Minikube cluster running, we will create an Nginx deployment using the following command:
kubectl create deployment nginx-demo --image=nginx --replicas=3
Now, to look past the Kubernetes abstraction layer, we will directly exec inside the etcd pod. Since we are using Minikube, we will directly target the control plane pod using its specific certificate paths:
kubectl exec -n kube-system etcd-minikube -- sh -c "ETCDCTL_API=3 etcdctl \
--endpoints=https://127.0.0.1:2379 \
--cacert=/var/lib/minikube/certs/etcd/ca.crt \
--cert=/var/lib/minikube/certs/etcd/server.crt \
--key=/var/lib/minikube/certs/etcd/server.key \
get /registry/deployments/default/nginx-demo --write-out=fields" | grep -E "Key|Value" | strings
In the output, you will see the key-value pair of our Nginx deployment. Kubernetes uses Protobuf binary to encode data, so some information stays hidden, but you can see messages confirming the successful progress of the Nginx deployment.
Unlike traditional relational databases, Kubernetesâs etcd uses Multi-Version Concurrency Control (MVCC). Every time we modify a Kubernetes object (like updating our Nginx deployment version or image name), etcd doesnât overwrite itâit increments a global revision number and stores a new version. To see this in a practical demo, use the following command:
kubectl exec -n kube-system etcd-minikube -- sh -c "ETCDCTL_API=3 etcdctl \
--endpoints=https://127.0.0.1:2379 \
--cacert=/var/lib/minikube/certs/etcd/ca.crt \
--cert=/var/lib/minikube/certs/etcd/server.crt \
--key=/var/lib/minikube/certs/etcd/server.key \
get /registry/deployments/default/nginx-demo --write-out=json" | jq
If we inspect the output, we can see that when we query the specific path (/registry/deployments/default/nginx-demo), the ModRevision is 2814. Now, letâs update our Nginx image version using the following command:
kubectl set image deployment/nginx-demo nginx=nginx:1.25
Now, if we run the same command again and check the ModRevision, we will see something different (3854 in my case). As we update the image version, instead of overwriting, etcd incremented and stored the new version.
Conclusion
That wraps up our deep dive into etcd! Understanding how Kubernetes stores its state behind the scenes, how logical hierarchies are mapped out using flat key prefixes, and how MVCC handles version revisions gives you a whole new perspective when debugging cluster issues. You are no longer just looking at abstract kubectl commandsâyou know exactly what is happening down in the database layer.
I hope you enjoyed this episode of K8s-with-Pravesh. Stay tuned for the next one where we tackle even more core Kubernetes internals!
If you found this helpful, letâs connect:
- YouTube: youtube.com/@pravesh-sudha
- LinkedIn: linkedin.com/in/pravesh-sudha/
- X (Twitter): x.com/praveshstwt
- Medium: medium.com/@programmerpravesh
- Dev.to: dev.to/pravesh_sudha_3c2b0c2b5e0
Happy cluster building, amigos! đ



Top comments (1)
The way etcd handles the flat key-space for hierarchies is something that catches a lot of people off guard when they first start debugging state issues. It is easy to forget that those slashes in the keys are just logical abstractions and not actual directory structures, which can lead to some messy prefix scanning if your key naming convention isn't strictly disciplined. I have seen several production incidents where slow watch operations were caused by inefficient key patterns that forced the engine to do way too much work during prefix lookups. Keeping an eye on the compaction frequency is also critical, as a bloated MVCC history can quickly tank your disk I/O and latency if you aren't tuning your maintenance jobs properly.