DEV Community

Cover image for The Token Stops At The Edge
Anton Brilliantov
Anton Brilliantov

Posted on

The Token Stops At The Edge

The service never sees the caller's token - it receives a claim about who arrived, and a red test stand guards that.


👋 Hi, I'm Anton - a software engineer working mostly in PHP/Symfony and Go, currently carving a live PHP monolith into Go services. This block is about the three things a fleet of services has to do identically or each service reinvents them: who arrived, what they may do, where their request has been. This part is the first and the smallest of the three - what actually travels between services when a person is behind the call. Running notes are on my GitHub: github.com/brilliant-almazov.

This is how I do it right now, with the price attached - maybe you already do it better, maybe you see it differently.


The shape of the system, briefly

Short primer, so this reads on its own.

A PHP monolith is being taken apart into Go services. The Go API gateway is the only door into the new services from outside; behind it sit the extracted services, talking to each other over gRPC. The PHP monolith still runs behind its own front door - the two doors are peers and do not route into each other. Every one of the Go services is raised from the same shared platform library, which is what makes a fleet-wide rule like this one enforceable in the first place: there is one place where "how a service reads who is calling" is decided.

So when I say "the edge" below, I mean the door in front of the new services, and when I say "a service", I mean anything behind it - including the background worker that picks up a message an hour later.

The check that fails

The rule in this article exists in exactly one executable form, and it is a failure:

The test stand fails if a token header arrived in the incoming call metadata.

Not a warning. Not a lint note somebody reads on a good day. The stand goes red, and the run stops.

That is worth stating before the design it protects, because the design is the boring part - most teams end up somewhere near it - and the enforcement is the part that decides whether the design is still true six months later.

What is forbidden, and what travels instead

The forbidden thing: the caller's token is not forwarded into the system. It stops at the outer boundary.

What travels instead: the edge parses the input once and passes on a claim about who arrived, carried in the call metadata alongside the call itself. A service downstream receives a finished claim. It does not verify tokens - it has neither the keys for it nor a reason to.

Drawn as two paths through the same fleet:

token forwarded                      claim propagated
---------------                      ----------------

  request (token)                      request (token)
        |                                    |
        v                                    v
     [ edge ] -- token -->             [  edge  ]  <-- the only place with keys
        |                                    |
        |                                    |  claim: who arrived
   +----+----+------+                   +----+----+------+
   v         v      v                   v         v      v
[svc A]  [svc B]  [svc C]            [svc A]  [svc B]  [svc C]
 key      key      key                no key   no key   no key
 verifies verifies verifies           reads the claim, and that is all
Enter fullscreen mode Exit fullscreen mode

Two bands. Top, token forwarded: the request reaches the edge and then three services, each holding a key and verifying for itself. Bottom, claim propagated: the edge holds the only key, parses the input once, and the three services behind it hold no keys and read a claim about who arrived

Say it as a property rather than a policy and it gets sharper: a service in this fleet is not capable of validating a caller's token. Not "is asked not to". Not capable. There is no key material in it, no library wired for it, no configuration pointing at an issuer. If a token showed up, the service would have nothing to do with it - which is precisely why a token showing up has to be an error and not a shrug.

The episode: a rule with nothing under it

Here is the actual sequence, and it is not a production incident - I am not going to dress it as one.

What we had. The rule "no tokens travel inward" existed as an agreement. Everyone knew it. Everyone would have repeated it correctly if asked in a review.

What that cost. There was nothing to check it with. An agreement holds exactly until the first edit made by someone in a hurry, and the violation it produces is invisible: forwarding a header is not a compile error, not a failing assertion, not a 500. It shows up in production, and only if somebody goes looking for it. Nobody goes looking for a rule everybody agrees with.

What changed. The rule was turned into a prohibiting check: the test stand fails if a token header arrived in the incoming call metadata. That is now the only form in which this rule exists in an executable state on my side.

There are no numbers attached to this episode, so I am not attaching any. Not how many services accept a claim, not how large the metadata is, not how long the edge takes to parse the input - those numbers go in when they are taken from the repo, not before. A hole you can see is worth more than a plausible figure.

How this is usually done

Four approaches show up in public write-ups and standards, and all four are defensible depending on what you are building:

  • Forwarding the caller's token all the way down. Every service validates it independently. Simple to reason about, and the reason it is popular is that it needs no new format - the thing the client sent is the thing the service reads.
  • Exchanging the token at the boundary for a narrower internal one. This is a standardised operation - RFC 8693, OAuth 2.0 Token Exchange - and it keeps the external credential outside while still giving services something verifiable.
  • A signed assertion about the caller placed in the call metadata, usually as a JWT issued by the boundary rather than by the external issuer.
  • Mutual authentication between services at the transport layer, which answers a different question - which process is calling - and is a separate layer from which person is behind the call. The two are often confused; they compose, they do not replace each other.

None of these is wrong. The first is the cheapest to start with and the hardest to reverse: once three services have a verification path, removing it is a change to three teams' code. The second and third both move the parsing to one place; they differ in whether what comes out is still a credential. What we do is the third shape with the verification removed on the receiving side - the boundary issues the claim, and nothing behind it checks the issuing. That is a real trade-off and it gets a price list further down, rather than a shrug.

How we do it

Three statements, one line each:

  1. One place parses the input. The outer boundary, once, per request.
  2. Identity travels in the call metadata, next to the call itself - not in the body, not in a lookup the service performs.
  3. A service is a consumer of the claim, not a verifier of it.

And a fourth that is easy to forget until the logs go quiet: the same requirement applies on the asynchronous path. The claim has to survive the queue boundary the same way the call context does - passed as the first parameter everywhere, including background handlers and the publish call itself. If it does not survive, the background handler is acting as the machine, and every record it writes says so. Our domain events carry a reference to the acting principal in the message itself (principalRef in the event payload), which is what makes "who did this" answerable after the fact rather than at request time only.

One synchronous lane where the edge hands a claim to a service, and below it the asynchronous lane: service, queue, background handler, with two outcomes - the claim carried in the message metadata so the record names the person, or the claim dropped at the publish so the record names the machine

Why a check and not a rule

There is a ladder I keep coming back to, and it has three rungs:

  • A reminder - said in a thread. Lives one session.
  • A rule - written in an instruction file. Works while it is being read.
  • A check - a linter, a prohibiting test, a structure test, a blocking hook. Works.

A card of incoming call metadata with the authorization line highlighted in vermilion, and beside it the result plate: stand fails. Underneath, the line: a rule with no executable check is not a rule

The stand that fails on a token header sits on the third rung, in the same row as the other prohibiting checks we run: the forbidden test asserting that the row-scan loop appears in exactly one package; the block on a test that uses an empty context instead of the test runtime's; the drift check that fails the build when the environment catalogue and the metrics snapshot no longer match the code.

Three steps rising left to right - reminder lives one session, rule works while it is read, check fails the build - with the third step in the accent colour, and below them a strip listing the prohibiting checks that exist

One line, and it is the only thing in this article I would defend without qualification:

An agreement about identity with no executable check under it is not a rule.

It is a shared intention. Shared intentions are useful. They are not enforcement, and confusing the two is how a fleet ends up with three services that quietly do it the old way.

What the check actually asserts

It is worth being precise about what a prohibiting check like this does and does not prove, because it is easy to oversell.

It asserts one thing: on the path the stand exercises, no token header reached a service. That is narrow. It does not prove the claim is well-formed, it does not prove the edge parsed correctly, it does not prove some other path exists where a header does get through.

What makes it worth having anyway is the failure mode it removes. The violation it catches is the one nobody would otherwise notice - the header that gets forwarded because forwarding all headers was the convenient way to write the client, and that then works perfectly for months. A check that catches a class of silent, correct-looking mistakes is worth more than a broad check that catches loud ones, because the loud ones were going to be caught anyway.

Three consequences

One place parses the input. Parsing, signature checking, clock skew, key rotation, malformed-input handling - all of it exists once. Not once per service, and not once per service plus the four services that copied it from the first one before it was fixed.

Changing how people log in does not touch the services. The shape of the input is a property of the boundary. Behind it, services consume a claim whose meaning has not changed, so a change in the way identity is established is a change in one component instead of a fleet-wide migration.

Logs and traces name the person who acted, not the machine that ran. This is the one I underestimated. Because the claim rides with the call and continues into the event stream, an audit record can name the acting principal - and that record is written in the same transaction as the mutation it describes, not appended by a best-effort listener afterwards. "Who changed this" stops being an investigation.

Where identity ends and permission begins

A short divider, because these two get mixed constantly:

"Who arrived" and "what they may do" are different questions.

A decision that needs the service's own data - does this tenant own this row - is made in the service, because that is the only place it is answerable at all. Everything else is made as early on the request path as it can be. That is the whole boundary for the purposes of this article; the full breakdown of where permission decisions live is the next part.

What it costs

An honest list, not one token item:

  • The boundary becomes a single point of trust. The service takes the claim as given and cannot check it. If the boundary is wrong, everything behind it is wrong, confidently and silently.
  • A claim format exists now, and every participant must understand it. Changing it is a breaking change across the fleet, with all the coordination that implies.
  • A locally raised service needs a way to inject a claim. Otherwise it is unusable without the boundary in front of it, and the developer experience quietly degrades into "you can only test this on a stand".
  • Debugging "why is the wrong principal here" moves to the boundary - that is, into code the service's author does not own and may not have read.
  • The prohibiting stand has to be maintained. A check that gets switched off "just for now" is not a check any more; it is a comment with a longer commit history.

When not to do this

  • When there is exactly one service and no boundary as a separate participant. Parsing the input in place is cheaper, and a claim format between one sender and one receiver is ceremony.
  • When the service genuinely needs the original input - for instance, when it is the boundary, or when it re-issues credentials.
  • When an external consumer is contractually owed a response derived from the original token. Then the token has a reason to travel, and the reason is written down in a contract rather than in habit.

And one honest caveat about the boundary of this whole description: there is no single permission-decision service on my side. The general "may this caller reach this at all" check lives at the edge; the data-dependent check lives in the service; there is no third participant deciding both. Everything above is about identity - who arrived - and not about permission.

The multiplier line

Wiring the check, moving the claim through the metadata, threading the context into the background handlers - that part is fast now, fast in a way that is genuinely different from a few years ago. What did not get faster is the decision that this rule deserves a prohibiting check at all, and that the check belongs in the stand rather than in a document. That call sets what "correct" means, and everything generated afterwards is correct or wrong relative to it. Speed amplifies whoever set the constraints; it does not supply them.


Access, tracing, rollout - Part 1. Next: the same question one step further in - who decides what is allowed, and the three places that decision can be made, each with a different price.

If you do this better, tell me what your boundary passes inward and how a service proves it did not get a token. If you have been through this, what did your identity agreement turn out to have drifted into before anyone noticed? If you see it differently, say where forwarding the caller's token all the way down is the cheaper answer. How is it solved on your side, and what broke there?

Top comments (0)