Most Raft projects look healthy when every node is still running.
The useful question is narrower: after a leader stops, what exactly can the
remaining nodes prove?
I built RaftKV, a teaching-grade three-node Raft sharded key-value store in Go.
For this test I did not want another architecture diagram. I wanted a recorded
flow with a committed key, a stopped node, a failover, a read, and a new write.
This post walks through that flow and separates the evidence from the claims the
test does not support.
The cluster
RaftKV starts three independent node processes:
-
node-a,node-b, andnode-c - two shards:
shard-0andshard-1 - one Raft group per shard
- three replicas in each Raft group
- real HTTP RPCs for
RequestVoteandAppendEntries - client requests accepted by any node, then forwarded to the shard leader
The important part is that the nodes are separate processes. A leader election
is not a method call inside one test binary.
Before the fault, the health endpoint returned:
{
"cluster": "RaftKV Trial 2026-09-06",
"leaders": 2,
"maintenance": false,
"shards": 2,
"status": "ok",
"version": "raft-kv-1.0.2"
}
Two shards, two leaders, three live nodes.
Step 1: commit a key
I wrote a key through the public HTTP API:
trial:user:1001 = RaftKV-ok
The key was routed to shard-1. The write had to go through the shard leader,
reach a Raft majority, commit, and then apply to the state machine.
This distinction matters. A value stored in memory on one node is not the same
as a committed Raft entry replicated to a majority.
Step 2: stop node-a
While the cluster was healthy, I stopped node-a, which had node_id 0.
I did not stop a request handler or change a flag in a test. The process
disappeared.
Step 3: ask the surviving nodes
After waiting for a new election, both surviving nodes reported:
{
"cluster": "RaftKV Trial 2026-09-06",
"leaders": 2,
"maintenance": false,
"node_id": 2,
"shards": 2,
"status": "ok",
"uptime_s": 81,
"version": "raft-kv-1.0.2"
}
The affected shard had a new leader. The old leader was no longer part of the
live majority.
Step 4: read the committed key
The previously committed key was still available:
{
"found": true,
"key": "trial:user:1001",
"shard": "shard-1",
"value": "RaftKV-ok"
}
This is the part that connects Raft theory to an operational result. The key
survived because it had already been committed by a majority, not because one
process happened to keep a copy in memory.
Step 5: write again
A cluster that can only read after failover is not enough. I wrote a new key:
failover:after:node0 = still-writable
The new value was read back successfully.
The recorded evidence summary is:
RaftKV fault evidence: stopped node-a (node_id 0), then queried node-b and node-c.
health post: {"cluster":"RaftKV Trial 2026-09-06","leaders":2,"maintenance":false,"node_id":2,"shards":2,"status":"ok","uptime_s":81,"version":"raft-kv-1.0.2"}
read post: RaftKV-ok
write post: still-writable
What this test proves
In this implementation and this recorded environment, a three-node Raft group
continued to serve committed data after one node stopped. A new leader was
elected, the old committed value was readable, and the cluster accepted a new
write.
The same public evidence repository also records:
- three-node HTTP Raft read and write tests
- follower restart and catch-up coverage
- restart recovery
- dynamic shard add and migration recovery
- snapshot restore
- concurrent write visibility
Raw files include:
evidence/trial-health-pre.jsonevidence/trial-status-pre.jsonevidence/trial-read-post.jsonevidence/trial-write-post.jsonevidence/fault-post-summary.txtevidence/go-test-output.txt
What it does not prove
I do not treat this experiment as proof that RaftKV is production-ready.
It does not prove:
- correct behavior under every network partition
- disk-failure or fsync-failure correctness
- cross-data-center replication
- production-grade consistency guarantees
- performance under production traffic
RaftKV is teaching-grade and portfolio-grade software. It is not a replacement
for Redis, etcd, or TiKV.
That boundary is not a disclaimer added at the end. It changes how the project
should be reviewed. The goal is to make the implementation and failure behavior
inspectable, not to pretend a small system has covered every production failure
mode.
Why this is a better project story
A project description that says "implemented Raft, sharding, and failover" is
difficult to verify.
A stronger story separates four things:
- The design choice.
- The failure that was injected.
- The observed result.
- The remaining uncertainty.
For a backend interview, that structure also creates better follow-up
questions. Why did the committed key survive? What happens to an uncommitted
entry? What happens when the old leader restarts? What still breaks under a
network partition?
Those questions are more useful than another list of features.
Reproduce it
The public evidence repository is here:
https://github.com/yuan1521913/raft-kv
The full Go source, tests, Docker Compose setup, management dashboard, and
teaching materials are part of the licensed source package.
If you want to run the Windows trial first:
If you build distributed systems, I would be interested in the failure test you
consider mandatory before trusting a Raft implementation.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.