Seventeen hosts in my cluster run Consul, twelve run Nomad, three run Vault. Upgrading any of them follows the same rules. Servers tolerate running ahead of clients, not the reverse. One server goes down at a time, and only while the rest can still coordinate. The host that coordinates hands that off deliberately rather than by going away.
My binaries are installed by Cinc, the free distribution of Chef. Every host runs cinc-client on a timer, converging it to whatever its cookbooks describe.
The version each tool should be running is not in those cookbooks. It lives in a versions data bag, read by a small library:
# The bag is the only place a version is declared. munchbox-hashi-upgrade
# writes it as part of a rolling upgrade; this repository declares no version
# at all, so there is no second copy to drift from and bumping one is an API
# call rather than a pull request.
#
# Missing pins raise. A pin only goes absent through misconfiguration -- the
# bag comes from the same Chef server as the cookbooks, so a node that cannot
# read it never got far enough to converge -- and falling back to a literal
# would install a stale version instead of naming the problem.
def pinned_version(tool)
item = data_bag_item(VERSIONS_DATA_BAG, tool)
version = item['version'].to_s
raise "munchbox_lib: #{VERSIONS_DATA_BAG}/#{tool} has no 'version' field" if version.empty?
version
end
The consul, nomad and vault cookbooks all call it the same way, and each restarts its service when the version it installed changes:
vault_install 'vault' do
version pinned_version(cookbook)
bin_path node[cookbook]['install']['bin_path']
...
end
So nothing upgrades by pushing a binary at a host. Changing the data bag changes what every host installs at its next converge, which means an upgrade is: stop the scheduled converges so no host moves on its own, change the pin, then converge the hosts deliberately, in the right order, checking the cluster between each one.
This tool is that sequence.
Three commands
plan reads the cluster and writes a run file. It changes nothing: no pin, no timer, no host. Everything that alters anything is a task in the file it produces.
run carries out a run file, recording the outcome of every task back into it. It resumes, so a run that stopped is continued by naming the same file.
status reads a run file and reports it. No cluster, no credentials, no change. A run that stopped overnight gets read before deciding what to do about it, and reading it should not require the ability to act on it.
hashi-upgrade plan consul --to 2.0.4
-> consul-munchbox-20261001T062815Z.yaml
hashi-upgrade run consul-munchbox-20261001T062815Z.yaml
hashi-upgrade status consul-munchbox-20261001T062815Z.yaml
plan names the file after the tool, the cluster it read, and the time it read it. The cluster name keeps two clusters planned from the same directory apart, and the timestamp means planning again writes a new file rather than overwriting the record of a run that already happened.
Task order
survey freeze-converge stop the scheduled converges fleet-wide
set-version-pin move the pin every converge will read
servers upgrade-member each non-coordinating host in turn
hand-off-coordination move coordination off the last one
upgrade-member that host, last
clients upgrade-member each host that carries work, in turn
verify verify-cluster read the fleet back against the target
thaw-converge release the scheduled converges
The timers stop before the pin moves. A scheduled converge landing between the two installs the new version on a host the run has not reached, out of order and unobserved.
The pin moves once, centrally, before any host is touched, so hosts differ only in when they are converged and never in what they converge toward.
The coordinating host goes last, so the cluster spends most of the server stage with settled coordination.
Verification runs before the timers come back. The per-host gates each proved one host returned, which is not the same as the fleet having arrived: a host the survey missed, or one whose converge was a no-op because the pin never reached it, passes every gate and still runs the old binary.
A plan against my Consul datacenter (and yes, all the bare-metal nodes are named after Law & Order characters):
munchbox: 17 hosts, 1 server failure tolerated
upgrading consul to 2.0.4
servers
goren 2.0.3
nomad-server-03 2.0.3 coordinating
stabler 2.0.4 already at 2.0.4
clients
cabot 2.0.3
fontana 2.0.3
mccoy 2.0.3
rubirosa 2.0.3
...
survey
Stop scheduled converges across the fleet [prompt]
Pin consul to 2.0.4 [typed]
servers
Upgrade goren to 2.0.4 [prompt, point of no return]
Upgrade stabler to 2.0.4 [prompt]
Hand coordination off nomad-server-03 [typed]
Upgrade nomad-server-03 to 2.0.4 [typed]
clients
Upgrade cabot to 2.0.4 [prompt]
Upgrade fontana to 2.0.4 [prompt]
Upgrade mccoy to 2.0.4 [prompt]
Upgrade rubirosa to 2.0.4 [prompt]
...
verify
Confirm the fleet is running 2.0.4
Restore scheduled converges across the fleet
Confirmations are recorded in the file rather than decided by the runner when it reaches them, so which boundaries stop for a person can be reviewed in advance. Two strengths: prompt for anything worth pausing on, and typed for moving coordination and restarting the host that had it. Typed asks for the host's name back.
stabler is marked already at 2.0.4 from an earlier failed attempt. Its task still appears and settles as unnecessary when the run reaches it.
What a survey reads
Topology is never declared. Each tool is asked, and the answer is two reads that have to be matched up.
The two lists overlap. goren and stabler each run a Nomad server and a Nomad client, and a host like that is recorded once, as a server, because it carries one binary and one service. Upgrading it twice would restart one raft member twice and spend the cluster's fault tolerance twice for one host.
They are matched on advertise address rather than on name, because the two reads do not agree on name shape.
Gate conditions
A converge exiting zero does not mean the host is back. The gate after each host polls the cluster until the cluster says so:
func (g *Gate) Server(ctx context.Context, name, target string) error {
what := fmt.Sprintf("%s rejoining the cluster at %s", name, target)
arrived := fmt.Sprintf("%s is running %s, healthy, and back in the cluster", name, target)
return g.await(ctx, what, arrived, func(cluster plan.Cluster) error {
m, err := member(cluster, name)
if err != nil {
return err
}
switch {
case m.Version != target:
return fmt.Errorf("%s is running %s, not %s", name, m.Version, target)
case !m.Healthy:
return fmt.Errorf("%s is not healthy", name)
// Asked only where coordination runs on a quorum of these members. A
// cluster whose quorum lives in its storage backend has no voters at
// all, and requiring one would hold every host here until it timed out.
case cluster.Votes() && !m.Voter:
return fmt.Errorf("%s is not a voter", name)
case !cluster.Healthy:
return fmt.Errorf("%s rejoined but the cluster is not healthy", name)
case cluster.Tolerance < minTolerance:
return fmt.Errorf("failure tolerance is %d, want at least %d", cluster.Tolerance, minTolerance)
}
return nil
})
}
The version proves the host restarted, read from the running agent rather than from the pin. A host that was never restarted reports the old version however healthy it looks.
A timeout carries the condition it gave up on rather than only its duration. "Waited ten minutes" gives nothing to act on; "spent them waiting for cinc-server at 2.0.4, and it is running 2.0.3" identifies the host and the check.
Adding Consul
Consul's autopilot answers in the same shape as Nomad's: healthy, failure tolerance, and a server row carrying version, leader, voter and raft ID. The snapshot a run is generated against is built the same way, and everything downstream is shared without modification: the run file, the runner, the step table, the gates, all three commands.
What is new per tool is one client package and a case here:
// For returns the client for a tool.
//
// A tool with no client is refused by name rather than attempted, so it fails
// here with something to read instead of somewhere deeper holding a nil.
func For(tool plan.Tool, opts Options) (Cluster, error) {
switch tool {
case plan.Nomad:
return nomad.New(nomad.Options{Address: opts.Address, Region: opts.Region})
case plan.Consul:
return consul.New(consul.Options{Address: opts.Address})
case plan.Vault:
return vault.New(vault.Options{Address: opts.Address})
default:
return nil, fmt.Errorf("unknown tool %q; expected nomad, consul or vault", tool)
}
}
Capabilities that not every tool has stay off the shared interface. Nomad is the only one that places work, so a client able to drain a host implements execute.Drainer and is asked for it at the point of use. The --drain flag is dropped at plan time for a tool that schedules nothing, rather than recorded in the file and ignored.
Vault with Consul storage
My Vault cluster keeps its data in Consul rather than in integrated raft storage, which changes what a survey can read.
There is no raft, so sys/storage/raft/autopilot/state does not exist. There is no failure tolerance to read, no stabilisation timestamp, and no voters. Every node is a standby contending for a lock, and vault operator raft list-peers returns No raft cluster configuration found.
The survey reads sys/ha-status, which lists the HA set from the active node's view. Failure tolerance is derived rather than read: the cluster serves while one node is unsealed and one is active, so what it can afford to lose is every serving node but one.
Voting needed a decision. Treating "Vault does not vote" as a per-tool fact would be wrong, because the same Vault on integrated raft storage does vote. Whether coordination runs on a quorum of the cluster's own members describes the deployment, not the tool, so it is read from the members:
// Votes is whether coordination in this cluster runs on a quorum of its own
// members.
//
// Read from the cluster rather than decided by the tool, because it is a
// property of how the cluster is deployed and not of what it is. A Vault
// cluster keeping its data in Consul has no voters; the same Vault with
// integrated raft storage does. Asking the members means neither case has to
// be configured, and a cluster that is migrated between them is read correctly
// without being told.
//
// One host mid-restart has dropped out of its quorum while its peers have not,
// so any voter makes this a voting cluster.
func (c Cluster) Votes() bool {
for _, m := range c.Members {
if m.Voter {
return true
}
}
return false
}
Migrating Vault to raft storage later requires no change here.
Restarting a Vault node leaves it sealed until the KMS unseals it. sys/ha-status does not report that, so each node is also asked sys/health at its own address, which does. A sealed node then counts as unhealthy, and a gate waiting on one says it is sealed rather than blaming the version.
A configuration server in the test binary
Everything above moves a pin on a Chef server, so the tests have to prove a pin written there can be read back. Stubbing the HTTP calls proves only that the client called the paths its author expected, which is the same author who wrote the assertions.
cinc-server-ng is that server as a Go library, so the tests run against a real one, in process:
// live starts a configuration server and returns a client that authenticates
// to it as the bootstrap admin.
func live(t *testing.T) *Cinc {
t.Helper()
srv, err := server.New(server.Options{Orgs: []string{testOrg}})
if err != nil {
t.Fatalf("new configuration server: %v", err)
}
if err := srv.Start(); err != nil {
t.Fatalf("start configuration server: %v", err)
}
t.Cleanup(func() { srv.Stop(context.Background()) })
c, err := New(Options{
ServerURL: srv.URL() + "/organizations/" + testOrg,
ClientName: srv.AdminName(),
KeyPath: writeTemp(t, "admin.pem", srv.AdminKey()),
})
...
}
Real data bags and real Mixlib signature verification on both halves of the round trip, in 0.08 seconds, with no Docker and no build tag. These are ordinary unit tests:
func TestPinRoundTrip(t *testing.T) {
c := live(t)
if err := c.SetPin(t.Context(), "nomad", "2.0.6"); err != nil {
t.Fatalf("SetPin: %v", err)
}
got, err := c.Pin(t.Context(), "nomad")
...
}
The signing is the part this earns. The client signs every request the way Chef expects, and a stub would have accepted a signature that no server would.
Container fleets
Each tool also has a Docker Compose environment: three servers and two clients for Nomad and Consul, three servers for Vault. Each starts a release behind the version to plan toward, since a fleet already at the target plans a run whose every task is a no-op.
They are driven by hand, the same way a real fleet is:
TOOL=consul make cluster-up # start it, behind the version to plan toward
TOOL=consul make cluster-plan # survey it and write a run file
TOOL=consul make cluster-run # drive the file; ARGS=--yes to skip prompts
TOOL=consul make cluster-down # stop it and discard its state
Each is a separate Compose project on its own configuration-server port, so more than one can be up at once. That matters when a change touches shared code: the Vault work altered the gate conditions and the successor choice for every tool, and having the Nomad and Consul fleets still running is what proved it had not broken them.
The hosts carry stand-ins for the two things a container does not have: a systemctl that answers the timer commands, and a cinc-client that reads the pin and installs it. That covers the orchestration around a converge. Whether a cookbook installs Consul correctly belongs to the cookbook, which is tested on its own.
The Vault environment runs two services that are not part of the fleet being upgraded: a Consul for storage, and a second Vault serving a transit key in place of a cloud KMS, so a restarted node returns unsealed without intervention. A Vault cluster without automatic unsealing cannot be rolled unattended at all.
Four faults found in these environments:
The gates' check for an unhealthy cluster could never fail. Nomad's autopilot answers 200 when the cluster is healthy and 429 when it is not, with the same health document in the body either way. The Go client turns any non-2xx into an error and drops the body, so an unhealthy cluster came back to me as a failed read, never as a snapshot saying Healthy: false.
A gate treats a failed read as "not yet" and polls again, so every snapshot a gate ever received had come from a 200, where healthy is true by definition. The !cluster.Healthy condition was evaluated on every poll and was never once true.
The fix is to pull the body back out of the error, but only for that one status. Anything else stays an error, because a cluster that cannot be read is a different problem from one that reads as unwell:
// unhealthyReply recovers the health reply autopilot sends with a 429.
//
// The status is how autopilot says the cluster is not healthy, which is a
// condition a run waits out rather than an error it stops for. Any other
// status, or a body that will not parse, is left as the error it was: a
// cluster that cannot be read is different from one that reads as unwell.
func unhealthyReply(err error) (*api.OperatorHealthReply, bool) {
var resp api.UnexpectedResponseError
if !errors.As(err, &resp) || resp.StatusCode() != http.StatusTooManyRequests {
return nil, false
}
var health api.OperatorHealthReply
if json.Unmarshal([]byte(resp.Body()), &health) != nil {
return nil, false
}
return &health, true
}
The survey calls it on the one error path it has, and fails as before when it returns false.
Consul uses the same 429 convention for an unhealthy datacenter, but its Go client accepts that status and parses the body rather than erroring on it, so the Consul survey needed no equivalent:
// we use 429 status to indicate unhealthiness
_, resp, err := op.c.doRequest(r)
...
err = requireHttpCodes(resp, 200, 429)
The server gate waited on a timestamp that does not move. An agent that goes down and comes back inside one health interval is never seen unhealthy. StableSince therefore still reads from before the restart, and the gate cannot pass a host that has already arrived.
A failed converge was taken as success. Converge returns a non-zero exit as a result rather than an error, since the command reached the host and reported. The caller discarded it. So a cinc-client that exited 1 counted as done. The run moved on to a gate waiting for a host nothing had installed anything on, timed out there, and reported a version mismatch.
A voter was required in a cluster where nothing votes. Two sites. The gate above is one. The other is choosing which host coordination is handed to, which fails later: after two of three servers are upgraded, at the step marked as the point of no return.
Hand coordination off vault-server-1
unwinding: releasing the converge timers
hand-off-coordination: no healthy voter can take coordination from vault-server-1;
unfit: [vault-server-2 vault-server-3]
Both sites now call Cluster.Votes() first.
Two faults found in production
Gates gave up after three minutes. A converge installs a couple of hundred megabytes, restarts a service and rejoins a cluster. One host took 3m20s. A gate that gives up inside that fails a run that was going to succeed, and leaves the operator to determine whether the host is wrong or slow. The default is ten minutes.
The pin rolled itself back. Two separate decisions combined into a state with no recovery path.
The pin step recorded a compensation that restored whatever version it found. The reasoning was that an unwinding run turns the converge timers back on. Leaving the pin at the new version would then have every host converge to it unattended, which is what freezing the timers before pinning exists to prevent.
Separately, a resumed run does not repeat a task that already succeeded.
During a 17-host Consul run I left a confirmation prompt unanswered. The run failed, unwound, and restored the pin to 2.0.3.
On restart, the pin step was already recorded as succeeded, so it was skipped. Three hosts then converged onto 2.0.3, the version the fleet was being upgraded away from, and sat at gates waiting for a 2.0.4 that was never coming. Recovering meant editing the data bag by hand.
The fix was to stop treating the pin as a reversible side effect. It is the run's intent. Converges are frozen for the whole run, so nothing reads the pin until the run ends and leaving it set costs nothing. Rolling it back costs twice: a fleet half-converged onto the new version and aimed at the old one converges backwards once the timers return, and a resume cannot repair it.
Two tests had encoded the old behaviour.
Run output
Vault, 2.0.4 to 2.1.1, against the container fleet:
Stop scheduled converges across the fleet
Pin vault to 2.1.1
Upgrade vault-server-2 to 2.1.1
vault-server-2 | [cinc-client] reading the pin from http://cinc-server:8889/organizations/test
vault-server-2 | [cinc-client] pinned 2.1.1, running 2.0.4
vault-server-2 | [cinc-client] fetching https://releases.hashicorp.com/vault/2.1.1/vault_2.1.1_linux_amd64.zip (168 MB)
vault-server-2 | [cinc-client] downloaded in 4s
vault-server-2 | [cinc-client] installed Vault v2.1.1
vault-server-2 | [cinc-client] restarting the agent
vault-server-2 | [cinc-client] converge complete
waiting for vault-server-2 rejoining the cluster at 2.1.1
vault-server-2 is running 2.0.4, not 2.1.1
vault-server-2 is running 2.1.1, healthy, and back in the cluster (4s)
Upgrade vault-server-3 to 2.1.1
...
waiting for vault-server-3 rejoining the cluster at 2.1.1
vault-server-3 is running 2.0.4, not 2.1.1
vault-server-3 is running 2.1.1, healthy, and back in the cluster (13s)
Hand coordination off vault-server-1
waiting for coordination moving off vault-server-1
vault-server-1 is no longer coordinating (2s)
Upgrade vault-server-1 to 2.1.1
...
waiting for vault-server-1 rejoining the cluster at 2.1.1
vault-server-1 is running 2.1.1, healthy, and back in the cluster (4s)
Confirm the fleet is running 2.1.1
Restore scheduled converges across the fleet
every host is on vault 2.1.1
Each gate reports the condition not yet met, then the fact that it now is, with how long it took. vault-server-2 is running 2.0.4, not 2.1.1 is the gate observing a host that has not come back yet.
The download size and timing are printed because curl is silent while it works. The first version printed fetching and then nothing for a minute, which on a cold run is indistinguishable from a hang. I interrupted a working run on that basis.
Resuming
Two outcomes do not settle: failed, where the task reported an error, and active, where it started and never reported back, which is what an interrupt leaves behind. Neither is stepped over on resume, because whether the work took effect is what is unknown.
stopped: upgrade-cinc-server is active
check the host, then: run consul-munchbox-20261001T055728Z.yaml --reset upgrade-cinc-server
The run does not decide on its own that a half-done converge is safe to repeat. Naming the task explicitly, after looking at the host, is what distinguishes a retry from a guess.
Production runs
All three tools have been driven against production: Nomad 2.0.5 to 2.0.7 across 12 hosts, Consul 2.0.3 to 2.0.4 across 17, Vault 2.0.4 to 2.1.1 across 3.
The Consul run exercised the most, since it touches five hosts Nomad does not: the Proxmox hypervisors and the Cinc server. It failed immediately on the Cinc server:
freeze-converge-timers: ssh dial 192.168.68.99:22 as root:
ssh: handshake failed: ssh: non-certificate host key
The tool verifies host keys against a CA and refuses a bare one. That host was presenting a bare key, with its signed certificate sitting unused on disk, because its role was missing the sshd_ca recipe every other node role carries. One line in a role file. Only a run that dialled every host was going to surface it.
A preflight script exists for exactly this, and it missed it. The script shells out to ssh, which was happy to trust a known_hosts entry. It checks reachability, not the credentials a run actually uses. That gap is still open.
The code
The tool is at munchbox-hashi-upgrade. The design document in that repository covers the sequence, what each tool does differently, and the full list of faults found so far.
Three things are still open. The integration tier covers Nomad only, so the handoff and the coordination gate have no automated test against a real election; they are covered by container fleets driven by hand and by the production runs. A cluster-level barrier between hosts is written and tested, but no step calls it yet. Draining a client before restarting it is implemented and off by default, and has never run against a production fleet.
Originally published at alexfreidah.com.



Top comments (0)