DEV Community

Anton Brilliantov
Anton Brilliantov

Posted on

The Mark That Stops After The First Hop

A new version is not ready when it is up. It is ready when the mark reaches the last hop - and mine would not have.


👋 Hi, I'm Anton - a software engineer working mostly in PHP/Symfony and Go, currently carving a live PHP monolith into Go services. This block is about the things a fleet of services has to do identically or every service reinvents them; this last part is about putting a new version in front of one request instead of all of them. Running notes live on my GitHub: github.com/brilliant-almazov.

This is how I do it right now, with the price attached - maybe you already do it better, maybe you have been through it, maybe you see it differently.


The answer I already had

This starts with something that was already broken, before the idea below was on the table.

An audit of one of my own services asked a single question: where in this service is there code that the shared platform library already provides. One finding reads like a footnote and is not one:

The publish was going out with empty headers.

The first consequence was a decoding one. With no header to select a codec by, the platform typed receive was physically inapplicable, so the consumer had to call JSON decoding by hand, in the handler, on raw bytes.

The second consequence is the one this article is about. Anything that travels in headers did not cross the queue boundary. Not the trace context, not the claim about who arrived, and not a mark on the request.

A mark on the request is exactly the mechanism the rollout technique below stands on. So the question "will the mark reach the last hop" was never open on my side. It already had an answer, and the answer was no.

Where this sits

A short primer, so this reads on its own.

A PHP monolith is being taken apart into Go services. The Go API gateway is the only door into the new services from outside; behind it the extracted services talk to each other over gRPC, and every one of them is raised from the same shared platform library. Two things are already settled at that door: the edge parses the input once and passes on a claim about who arrived, carried in the call metadata; and, as the previous part argued, a flag is decided there too - the service reads a finished yes or no rather than the address of a flag store.

Some work answers inside the call. The rest is published as a domain event and picked up later by a background handler in a separate process. That second half is where every mechanism in this article either survives or quietly stops.

The technique, in one paragraph

The new version is raised beside the old one. The standard path stays exactly as it was for everybody. Only a request carrying a mark is routed into the new version.

The mark is not a new invention. It is the same kind of statement as the identity claim and the flag: a decision taken before the service runs, travelling in the call metadata, with the lifetime of one request. Nothing new appears in the service; a field that already exists gets one more neighbour.

That paragraph is the entire appeal of the technique, and it is also why it is easy to underestimate. Standing a second copy of a binary next to the first one is a deployment parameter. Getting a mark to the end of a chain is not.

The condition that makes or breaks it

Written as a condition, because that is the only form of it that can be checked:

The technique is ready when the mark arrives at the last hop - not when the second version is up.

Stated that way it becomes something a test can hold: send one marked request, walk the whole chain, and assert that the last participant saw the mark. Stated as an intention - "we will pass the flag through" - it becomes a rule, and a rule that produces no symptom when broken is a rule that is already broken somewhere you have not looked yet.

If the mark stops after the first hop, the first service runs the new code, and every call it makes afterwards - synchronous or not - goes to the old one. That is not "the rollout is partially on". It is one user action executed by two generations of code, with nobody in a position to notice, because each half looks correct from where it stands.

  edge          hop 1            hop 2         hop 3         hop 4
  decides  -->  new version  |   old        -->  old      -->  old
  mark: yes     mark: yes    |   mark: -        mark: -       mark: -
                             |
                       the mark stops here
Enter fullscreen mode Exit fullscreen mode

Four hops in a row - the first one routed into the new version and carrying a yes mark, the three after it routed into the old version with no mark, and a broken connector with a visible gap between the first and second hop labelled the mark did not travel

The two places the thread breaks

Both of them are already documented on my side, which is the reason I trust the condition above rather than my optimism.

A dropped call context. The context is passed as the first parameter everywhere - handlers, repositories, managers, background handlers, the publish call itself - and it is checked by cancellation tests. Breaking that rule produces no symptom: nothing fails to build, nothing throws, no response code changes. It shows up as a thread that quietly stops being connected.

A publish with empty headers. The finding above. Everything the messaging layer needs in order to know what it is holding travels in the same place as everything the rollout needs in order to know which generation this request belongs to. When those headers are empty, both are broken at once - and fixing either one fixes the other.

Two panels - on the left a dropped call context with a note that it produces no symptom, on the right a publish with empty headers with a note that the typed receive becomes inapplicable, and below them one wide bar reading the mark travels in exactly the places that were already broken

The useful way to read that pair: the mark is not asking for a new transport. It is asking for the transport that was supposed to be working already.

The message that repeated forever

Here is the concrete episode, from the same audit, and it is the reason I treat contract compatibility and consumer behaviour as one requirement rather than two.

What we had. A subscription that set no redelivery limit, no backoff and no dead-letter queue. A decode error was returned as an ordinary error. The driver did what a driver does with an ordinary error: it redelivered.

What that cost. One message the consumer could not parse occupied the stream and was discarded by nothing. Not slowed down - repeated, indefinitely, with nothing in the arrangement able to give up on it.

One message that fails to decode looping back into the same consumer through three missing guards - no redelivery limit, no backoff, no dead-letter queue - with the loop drawn in the accent colour and a note that the decode error was returned as an ordinary error

Why it belongs in an article about rollout. Two generations of code on one stream is precisely the situation where a message the consumer cannot parse arrives as a matter of routine rather than as an accident. The new version writes an event the old consumer does not understand, or the reverse, and the poison loop above is no longer a latent defect - it is the expected outcome of the technique working as designed.

So the requirement is one requirement, not two: the contract is compatible in both directions, and the consumer has a defined behaviour for a message it cannot parse. Satisfying only the first is how you find out about the second.

How this is usually done

Four shapes show up in public write-ups and standards, and all four are reasonable:

  • Routing on a header at the entry layer. The decision is taken once at the door, and the entry layer sends the request to one of two upstreams. This is the shape the technique below is a version of.
  • Two colours side by side. A whole second environment is raised, traffic is switched to it at once, and the old one is kept ready to switch back - see BlueGreenDeployment.
  • A gradually increasing share of traffic, with the share moved up while the failure rate is watched - see CanaryRelease.
  • A separate sandbox on a copy of the environment, where the new version is exercised with no live traffic at all.

The four side by side, in the only terms that decide between them:

Shape What it needs from the chain Blast radius of one mistake What it costs
routing on a header the mark reaching every hop one marked caller a mark that must survive the whole chain
two colours nothing - the whole environment moves everything, briefly a second environment, kept warm
a growing share nothing at the chain level the current share a way to watch failure rates and move the dial
a separate sandbox nothing - no live traffic at all none a copy of the environment, and no real load in it

I am not arguing any of them is wrong, and I am deliberately naming no gateways, no service meshes and no entry providers - the choice of product is not what makes this work or fail. The row that matters for this article is the first one, and its middle column is the reason to want it: the blast radius is one caller who agreed to be that caller. Its right-hand column is the reason it is hard.

How it lands on the deploy I already have

Worth separating what is already there from what is not, because they land very differently.

What is already there. Deployment is a manual pipeline run with two parameters: the version, and which services to move (server, worker, all). Tags are set by one person, pseudo-versions are forbidden, and a service carries exactly the tag it was given. The version and the commit are stamped into the binary with linker flags, and there is a build target that checks the symbol is alive - because the linker will silently ignore a flag aimed at a symbol that does not exist, leaving the version at dev and sending dev into the audit trail and into the trace attribute.

So "stand the new version beside the old one" costs nothing new: it is the same pipeline, run twice, with two version arguments. There is no new code, no new component and no new runtime.

"Get the mark to the last hop" lands on nothing at all. It is the part that has to be built, and the two breaks above are its actual scope.

One honest boundary while I am on the subject of the pipeline: there is no "wait for health" step in it. Health and readiness are served by the platform, and the check after a rollout is the environment's job, not the pipeline's.

What it costs

The four items below are the ones I would put in front of anybody who wants to try this, and none of them is a token entry.

Two generations of code live at the same time. Not two branches in a repository - two running versions, both correct, both reachable, both maintained. Every fix in the period applies to a question nobody enjoys answering: which of the two, or both.

The contract has to be compatible in both directions. Forward compatibility is the one people plan for: the old side must tolerate what the new side sends. Backward compatibility is the one that gets discovered: the new side must tolerate what the old side is still sending, for as long as the old side exists.

Data written by the new version is read by the old one. On my side that is survivable by construction, because there is no destructive UPDATE in the domain tables - a change is a new version of a row, and the old generation keeps seeing its own rows. That holds right up until the new version starts writing something the old version cannot read at all, at which point the storage layer, not the contract, is what ends the experiment.

Whoever is testing has to know they are in the new branch. Otherwise debugging is done blind: a behaviour is observed, attributed to the wrong generation, and the conclusion is confidently wrong. This is the item that looks smallest on the list and produces the most wasted hours.

  old version  <-- contract compatible both ways -->  new version
       |                                                   |
       +-------------------> data <-----------------------+
                  written by new, read by old
Enter fullscreen mode Exit fullscreen mode

Old and new versions drawn side by side with a two-way arrow between them reading contract compatible both ways, and one shared data band below with an arrow up from new and an arrow down to old labelled written by new, read by old

Put together, the shape of the price is not what it looks like from the deployment side:

Four cost rows stacked in monospace plates - two generations alive, contract compatible both ways, data written by new read by old, the tester must know which branch they are in - and a fifth row separated by a rule and set in serif reading organisational, not technical

This is organisational complexity, not technical complexity. The technical part is a routing rule and a header. The expensive part is two generations of a system being true at the same time, and the people around it having to hold both in their heads.

Why do it at all, then

Two reasons. I do not have a third.

The alternative prices every mistake at the whole traffic. Rolling out to everybody at once is not a cheaper technique; it is the same technique with the blast radius set to maximum. The marked request exists so that the first thing a mistake meets is one caller who agreed to meet it.

The mechanism already exists. The mark is the same decision, taken at the same place, travelling in the same call metadata as the identity claim and the flag. No flag store appears, no client appears in each service, no cache and no refresh loop appear - which matters, because a piece of state with a refresh loop is runtime, and runtime is what the platform provides. The technique earns its keep partly by not adding anything.

When not to do this

  • When the chain is longer than it can carry the mark. Then the honest order is: stitch the chain first, argue about routing second. Routing a mark into a chain that drops it produces a half-migrated request, which is worse than no rollout mechanism at all.
  • When the change touches storage in a way that makes reading back impossible. Compatibility in both directions is not achievable there, and the technique simply does not apply - no amount of routing discipline fixes a row the other generation cannot read.
  • When there are few enough consumers to move all of them at once. Then the whole apparatus is ceremony, and the cheaper answer is a coordinated switch.

And the boundary of this write-up, stated plainly because it matters: there is no routing of a marked request into a new version in production on my side. This is a breakdown of a technique and its price, not a report on a working scheme. There are also no numbers here - not the share of marked traffic, not how long two generations coexist, not the number of hops in the chain, not how long a rollout takes - because none of those have been measured on my side. A hole you can see is worth more than a plausible figure.

The multiplier line

An automated executor makes the mechanical half of this cheap: threading one more field through the call metadata, setting headers at every publish site, writing the test that fails when the mark does not reach the background handler. What it does not do is decide that the readiness criterion is the last hop rather than the second binary, or notice that a poison loop and a contract change are the same requirement seen from two sides. That judgement is the work. Speed amplifies whoever set the definitions; it does not supply them.


Access, tracing, rollout - Part 5. Next, and last in this block: checking a release before everybody starts using it - the coverage ratchet sitting at zero, one infrastructure container for the whole run, and what drifts between what is declared and what is actually there.

If you do this better, tell me how you prove a request mark reaches the last hop rather than the first. If you have been through this, what did your two coexisting generations turn out to disagree about? If you see it differently, say where rolling out to everyone at once is the cheaper answer. How is it solved on your side, and what broke there?

Top comments (0)