The synchronous path is one tree; everything after the publish is another one - and the audit found out why: the headers were empty.
👋 Hi, I'm Anton - a software engineer working mostly in PHP/Symfony and Go, currently carving a live PHP monolith into Go services. This block is about the three things a fleet of services has to do identically or each service reinvents them: who arrived, what they may do, where their request has been. This part is the third one - where the request has been, and the exact point where the answer stops being available. Running notes live on my GitHub: github.com/brilliant-almazov.
This is how I do it right now, with the price attached - maybe you already do it better, maybe you see it differently.
The shape of the system, briefly
A PHP monolith is being taken apart into Go services. The Go API gateway is the only door into the new services from outside; behind it the extracted services talk to each other over gRPC, and each of them is raised from the same shared platform library. Some work is synchronous and answers inside the call. The rest is published as a domain event and picked up later by a background handler in a separate process.
That last sentence is the whole subject of this article. Everything before the publish is one call. Everything after it is a different process, at a different time, with a different life expectancy - and by default, a different trace.
What arrives without a line of service code
Before any argument about instrumentation, the inventory. A snapshot taken from a freshly generated service - name, type, help text, labels, histogram buckets, where each one is declared - reads 67 records: 59 of them from the platform, 6 from the service, plus one <dynamic> entry for the dynamic-metric factory. On top of that the service declares 13 domain metrics of its own: queue depth, age of the oldest queued row, relay cycle duration, publish failures by reason, resolve duration, audit records written.
Collection is pull. An agent scrapes the system port with a bearer token and remote-writes into a store; the application never pushes anything. Traces leave over OTLP.
The line I would keep from that paragraph: observability here is an import, not a body of code inside the service. Health, readiness, the metrics endpoint, gRPC metrics, broker driver metrics, connection pool instrumentation, trace export - all of it comes with the runtime the service is built on. Nobody writes it per service, so nobody writes it slightly differently per service.
The one thing the service still has to do
Exactly one requirement is left on the service side, and it is small enough to state in a sentence:
Do not lose the call context.
It is passed as the first parameter everywhere - handlers, repositories, managers, background handlers, and the publish call itself - and it is checked by cancellation tests. In tests the context comes from the test runtime; an empty one is forbidden.
That last clause is the part that makes the requirement real, and it took a while to arrive at. A rule about context has a nasty property: breaking it produces no symptom. A dropped context does not fail a build, does not throw, does not change a response code. It shows up as a trace that quietly stops being connected, which nobody notices until they need it.
So the requirement is not written as a rule. It is written as a failure: a handler that drops the context fails its cancellation test, because the operation no longer ends when the caller goes away.
The case: a requirement that was only a rule
Here is what it looked like while it was still an agreement, and the numbers are the point.
What we had. "Pass the context everywhere" as a written rule, honestly believed by everyone. Meanwhile the test code had accumulated 421 calls building an empty context, across 184 test files.
What that cost. A test built on an empty context does not cancel with itself, so it hangs until the package timeout instead of failing. Worse, it asserts nothing at all about the context: if production code dropped the incoming context and substituted its own, that entire suite would stay green. A test suite of that size, running on every push, was structurally incapable of noticing the one failure mode the rule existed to prevent.
What changed. Four things, and only the last one is enforcement:
- the context comes from the test runtime,
t.Context()andb.Context(); - a port stub keeps the context it was handed, as a field, and gives it back to the test - instead of accepting it as an ignored parameter;
- every transport, repository, manager and consumer has a cancellation test: cancelled context in, an error that is true under
errors.Is(err, context.Canceled)out, and no row written; - writing a test with an empty context is blocked at the tool level, so the count cannot climb back.
The ladder underneath this is the same one from the first part of this block: a reminder lives one session, a rule works while it is being read, a check works. 421 is what the middle rung looks like after a year.
How this is usually done
Four shapes show up in public write-ups and standards, and all four are reasonable:
- A correlation identifier passed by hand. One header, generated at the door, logged by everyone. Cheap, works everywhere, and gives you grep rather than a tree - you can join records, but you cannot see nesting or duration.
- Standardised context propagation in the transport. W3C Trace Context defines the headers; client and server libraries carry them, and application code never touches them.
- Instrumentation through interceptors rather than call sites - one place per protocol wraps every call, so no handler contains tracing code.
- A separate collection layer next to the application, receiving spans and forwarding them, so the application knows one export protocol and nothing about the backend.
We sit on the second and third: propagation is a property of the transport, instrumentation is a property of the platform library, and the service contributes neither. Which leaves the asynchronous path, where the transport is a queue and no library does it for you unless you decide what travels. The messaging conventions say the same thing in more words: the producer puts the context into the message, or there is no link.
How we do it: what comes from where
The whole arrangement fits in one table, and the last column is the useful one:
| What you see in the trace | Where it comes from | What the service does | How it breaks |
|---|---|---|---|
| service version attribute | the value the build stamps in | nothing | a dead -X target leaves dev
|
| spans of the synchronous path | platform instrumentation | passes the context first, everywhere | a dropped context |
| durations | platform and service histograms | declares its domain metrics | names drift apart |
| the link to background work | the message headers | sets them at publish | empty headers |
Three of those four failures are invisible in a green build. That is the shared property worth naming: every way this breaks produces correct-looking output. The service answers, the data is right, the tests pass, and the trace is quietly wrong.
Where the thread breaks
Now the specific loss, the one this article exists for.
The synchronous path traces beautifully. Door, service, validation, transaction, publish - one root, nested children, durations that behave the way nesting says they should. Then the message lands in a queue, a worker in another process picks it up some time later, and that work has its own root. Two trees, no edge between them.
one trace a second tree, its own root
--------- ---------------------------
gateway root consumer handler
└─ service ├─ decode, called by hand
├─ validate └─ domain work
├─ transaction
└─ publish ← the last span no parent, no link back
with a parent
The thread across that gap is stitched with message headers - the same mechanism the serialisation already uses. And here is the finding from our own audit of the service, which is why this is written as a direction and not as a report:
The publish was going out with empty headers.
The consequence was not primarily a tracing one. With no header to select a codec by, the platform's typed receive was physically inapplicable - the consumer had to call JSON decoding by hand, in the handler, on raw bytes. One defect, two problems: the receiving side could not be typed, and the trace could not be continued.
That is the part I find genuinely useful to notice. The headers you need for propagation are not an extra tax invented by observability. They are the same headers the messaging layer already needs in order to know what it is holding. If they are empty, both things are broken at once, and fixing either one fixes the other.
What the message carries, and what it deliberately does not
Since the whole argument rests on headers, it is worth being exact about the division:
- The body is protojson of the domain event: event id, tenant, type, subject, timestamps, revision, acting principal.
- The row identifier in the outgoing queue table is the same value as the message id.
- Beyond the message id, the service was setting no headers of its own. That is the defect above, stated as an inventory line.
- Delivery status and attempt counters are not in the message at all: queue depth, age of the oldest row and relay cycle duration live in the metrics endpoint and never travel to the broker.
Which settles a design question that otherwise gets argued in circles: the thread is stitched with headers, not with a field in the event body. The body belongs to the event contract - consumers depend on its shape, and it outlives any particular observability arrangement. Metadata about the delivery of a message is not part of what happened in the domain.
The version attribute that lies quietly
One more failure from that table, kept separate because it is about trust rather than connectivity.
The build stamps the service version into a symbol the runtime reads, and the runtime puts it into the trace as an attribute. The linker will silently ignore an -X flag aimed at a symbol that does not exist. No warning, no error, green build - and the version stays dev. Then dev travels into the audit trail and into the trace attribute, and every span claims to come from a build nobody can identify.
The check is embarrassingly cheap: a build target that compiles both binaries with a test version tag and greps the binary for that string. Empty output means the symbol is wrong. And the flags have to be right in two files, the Makefile and the Dockerfile - a template carrying a dead target in one of them hands the same silent defect to every service generated from it afterwards.
A trace attribute nobody checks lies quietly.
Why it is arranged this way
Two reasons, and neither is about elegance.
Observability written by hand in a service is runtime written by hand in a service. Runtime is what the platform provides - registries, instrumentation, export, pooling. The moment a service writes its own version of it, its trace stops lining up with its neighbour's: different span names, different attribute spellings, different histogram buckets. A fleet where each service is individually well instrumented and collectively incomparable is worse than it sounds, because every cross-service question turns into a translation exercise first.
The single requirement left to the service is expressed as a check. Not because rules are worthless, but because 421 is what a rule looks like when it is only a rule. A requirement with no executable form is a shared intention, and the failure it permits is silent by construction.
What it costs
An honest list:
- Context as the first parameter is a discipline in every signature. It shows up in every review, forever, and it is the kind of thing that feels like ceremony right up until the day it isn't.
- Cancellation tests are written per transport and per handler. That is real test code, and it grows with the surface, not with the feature list.
- Message headers become part of compatibility. Once a consumer reads a header, changing it is a coordinated change, with the same care as changing a field in the contract.
- Cardinality is kept low on purpose. The tenant identifier does not go into labels - there are too many tenants, and the label would blow up the series count. So a per-tenant cut is simply not available in traces or metrics; it lives in queries against the data.
- A thread stitched across a queue makes the tree long. A trace spanning a synchronous call, a queue and a background handler is more expensive to read than two short ones, and a reader who does not know the shape will get lost in it.
When not to do this
- When there is no asynchronous path. Nothing to stitch; propagation in the transport already covers the whole story.
- When the message leaves the perimeter and its headers are controlled by someone else. Then the thread ends at your boundary by agreement, and the honest thing is to say so rather than to hope.
- When the volume makes recording everything cost more than it returns. That is a conversation about sampling, which is a different decision from whether the thread exists at all.
And the honest caveat about the boundaries of this write-up: there are no names of specific tracing systems here, and no numbers on trace volume, sampling share, span durations or tree depth - the first because naming them adds nothing, the second because those figures were never measured on my side. A hole you can see is worth more than a plausible figure.
The multiplier line
An automated executor made the mechanical half of this cheap: threading a parameter through hundreds of signatures, writing a cancellation test per transport, converting a rule into a blocking check. What it did not do is decide that the requirement should exist as a failure rather than a sentence, or notice that empty headers were simultaneously a decoding defect and a broken trace. That judgement is the work. Speed amplifies whoever set the definitions; it does not supply them.
Access, tracing, rollout - Part 3. Next: feature flags, and why the decision belongs at the boundary - the service receives a finished yes or no, not the address of a flag store.
If you do this better, tell me what you put in the message headers and how you prove it is still there. If you have been through this, what did your broken thread turn out to be hiding besides the trace? If you see it differently, say where a correlation identifier and good logs beat a stitched tree. How is it solved on your side, and what broke there?





Top comments (0)