DEV Community

yuan ming
yuan ming

Posted on

I Stopped a Raft Leader in a 3-Node Go KV Store. Here Is the Evidence.

Most Raft projects look healthy when every node is still running.

The useful question is narrower: after a leader stops, what exactly can the
remaining nodes prove?

I built RaftKV, a teaching-grade three-node Raft sharded key-value store in Go.
For this test I did not want another architecture diagram. I wanted a recorded
flow with a committed key, a stopped node, a failover, a read, and a new write.

This post walks through that flow and separates the evidence from the claims the
test does not support.

The cluster

RaftKV starts three independent node processes:

  • node-a, node-b, and node-c
  • two shards: shard-0 and shard-1
  • one Raft group per shard
  • three replicas in each Raft group
  • real HTTP RPCs for RequestVote and AppendEntries
  • client requests accepted by any node, then forwarded to the shard leader

The important part is that the nodes are separate processes. A leader election
is not a method call inside one test binary.

Before the fault, the health endpoint returned:

{
  "cluster": "RaftKV Trial 2026-09-06",
  "leaders": 2,
  "maintenance": false,
  "shards": 2,
  "status": "ok",
  "version": "raft-kv-1.0.2"
}
Enter fullscreen mode Exit fullscreen mode

Two shards, two leaders, three live nodes.

Step 1: commit a key

I wrote a key through the public HTTP API:

trial:user:1001 = RaftKV-ok
Enter fullscreen mode Exit fullscreen mode

The key was routed to shard-1. The write had to go through the shard leader,
reach a Raft majority, commit, and then apply to the state machine.

This distinction matters. A value stored in memory on one node is not the same
as a committed Raft entry replicated to a majority.

Step 2: stop node-a

While the cluster was healthy, I stopped node-a, which had node_id 0.

I did not stop a request handler or change a flag in a test. The process
disappeared.

Step 3: ask the surviving nodes

After waiting for a new election, both surviving nodes reported:

{
  "cluster": "RaftKV Trial 2026-09-06",
  "leaders": 2,
  "maintenance": false,
  "node_id": 2,
  "shards": 2,
  "status": "ok",
  "uptime_s": 81,
  "version": "raft-kv-1.0.2"
}
Enter fullscreen mode Exit fullscreen mode

The affected shard had a new leader. The old leader was no longer part of the
live majority.

Step 4: read the committed key

The previously committed key was still available:

{
  "found": true,
  "key": "trial:user:1001",
  "shard": "shard-1",
  "value": "RaftKV-ok"
}
Enter fullscreen mode Exit fullscreen mode

This is the part that connects Raft theory to an operational result. The key
survived because it had already been committed by a majority, not because one
process happened to keep a copy in memory.

Step 5: write again

A cluster that can only read after failover is not enough. I wrote a new key:

failover:after:node0 = still-writable
Enter fullscreen mode Exit fullscreen mode

The new value was read back successfully.

The recorded evidence summary is:

RaftKV fault evidence: stopped node-a (node_id 0), then queried node-b and node-c.
health post: {"cluster":"RaftKV Trial 2026-09-06","leaders":2,"maintenance":false,"node_id":2,"shards":2,"status":"ok","uptime_s":81,"version":"raft-kv-1.0.2"}
read post: RaftKV-ok
write post: still-writable
Enter fullscreen mode Exit fullscreen mode

What this test proves

In this implementation and this recorded environment, a three-node Raft group
continued to serve committed data after one node stopped. A new leader was
elected, the old committed value was readable, and the cluster accepted a new
write.

The same public evidence repository also records:

  • three-node HTTP Raft read and write tests
  • follower restart and catch-up coverage
  • restart recovery
  • dynamic shard add and migration recovery
  • snapshot restore
  • concurrent write visibility

Raw files include:

  • evidence/trial-health-pre.json
  • evidence/trial-status-pre.json
  • evidence/trial-read-post.json
  • evidence/trial-write-post.json
  • evidence/fault-post-summary.txt
  • evidence/go-test-output.txt

What it does not prove

I do not treat this experiment as proof that RaftKV is production-ready.

It does not prove:

  • correct behavior under every network partition
  • disk-failure or fsync-failure correctness
  • cross-data-center replication
  • production-grade consistency guarantees
  • performance under production traffic

RaftKV is teaching-grade and portfolio-grade software. It is not a replacement
for Redis, etcd, or TiKV.

That boundary is not a disclaimer added at the end. It changes how the project
should be reviewed. The goal is to make the implementation and failure behavior
inspectable, not to pretend a small system has covered every production failure
mode.

Why this is a better project story

A project description that says "implemented Raft, sharding, and failover" is
difficult to verify.

A stronger story separates four things:

  1. The design choice.
  2. The failure that was injected.
  3. The observed result.
  4. The remaining uncertainty.

For a backend interview, that structure also creates better follow-up
questions. Why did the committed key survive? What happens to an uncommitted
entry? What happens when the old leader restarts? What still breaks under a
network partition?

Those questions are more useful than another list of features.

Reproduce it

The public evidence repository is here:

https://github.com/yuan1521913/raft-kv

The full Go source, tests, Docker Compose setup, management dashboard, and
teaching materials are part of the licensed source package.

If you want to run the Windows trial first:

https://5552463341538.gumroad.com/l/raftkv-trial?utm_source=devto&utm_medium=article&utm_campaign=raftkv_launch

If you build distributed systems, I would be interested in the failure test you
consider mandatory before trusting a Raft implementation.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.