DEV Community

Anton Brilliantov
Anton Brilliantov

Posted on

A Flag Is A Decision Made Before The Service

The service receives a finished yes or no the same way it receives who arrived - not the address of a flag store.


👋 Hi, I'm Anton - a software engineer working mostly in PHP/Symfony and Go, currently carving a live PHP monolith into Go services. This block is about the things a fleet has to do identically or every service reinvents them; this part is the shortest of the four, and it is one thought with one price attached. Running notes live on my GitHub: github.com/brilliant-almazov.

This is how I do it right now, with the price attached - maybe you already do it better, maybe you see it differently.


Where this sits

A short primer, so this reads on its own.

A PHP monolith is being taken apart into Go services. The Go API gateway is the only door into the new services from outside; behind it the extracted services talk to each other over gRPC, and every one of them is raised from the same shared platform library. Identity is already settled at that door: the edge parses the input once and passes on a claim about who arrived, carried in the call metadata next to the call itself. By the time a service runs, who is calling is a finished statement, not a question.

This part asks whether a second statement can travel the same way.

The thesis, in one line

A flag is not a variable inside a service. It is a decision made before the service runs.

The edge decides whether a branch is on for this caller, and the result travels the same way the identity claim does. The service reads a finished yes or no. It never learns where the answer came from, and it is never handed the address of a flag store.

a flag client in every service        the decision at the edge
------------------------------        ------------------------

  [service A] --+                       +-- [service A]  yes
   config       |                       |
   cache        +--> [ flag store ]     |
   refresh      |                     [edge] -- [service B]  no
  [service B] --+     each service     decides  |
   config       |     keeps its own    once     |
   cache        |     copy of the               +-- [service C]  yes
   refresh      |     machinery
  [service C] --+                       services hold no config,
                                        no cache, no refresh
Enter fullscreen mode Exit fullscreen mode

Left, a flag client in every service: three services, each carrying its own config, cache and refresh, all pointing at one shared flag store. Right, the decision at the edge: one edge box deciding once and three services each holding a finished yes or no, with no arrows leaving them

The case that made me write this down

Not a production incident, and I will not dress it as one. It is an audit finding.

What we had. An audit of one of my own services asked a single question: where in this service is code that the platform library already provides? Among the findings, two matter here.

Three caches written by hand - a map plus a read-write mutex, no TTL, no capacity limit, no metrics. And a retention schedule nailed down as a constant in the code instead of arriving as configuration.

What that cost. The platform library already ships a cache with a TTL, a capacity bound and metrics, and configuration arrives as an import rather than as code somebody writes. So each of the three hand-written copies was a piece of runtime that nobody could observe: no eviction anyone could reason about, no size anyone had bounded, no number on a dashboard when it misbehaved. Nothing was on fire. That is exactly the problem - a piece of runtime with no metric under it does not announce itself, it just sits there being load-bearing.

Left, written by hand and found in the audit: three caches built from a map and a mutex with no TTL, no capacity and no metrics, and a retention schedule as a constant. Right, what the shared platform library already provides: a cache with TTL, capacity and metrics, and schedules arriving as configuration

Why it belongs in an article about flags. A flag client living inside a service reproduces exactly these shapes - a piece of configuration, a cache, a refresh loop - and adds one more on top. It is the same mistake with a nicer name.

How this is usually done

Three shapes show up in public write-ups, and all three are defensible depending on what you are building:

  • An in-process client of a flag store in every service, with a local cache and a background refresh. This is the default in most write-ups on the subject, and its appeal is real: the decision sits next to the code it affects, so the branch reads naturally.
  • Flags as configuration that ships with the deploy. No runtime lookup at all - a value arrives with the release. Simple and observable; the cost is that changing it means a deploy.
  • The decision made on the entry layer, with the result attached to the request as it travels inward. Less common, and the one this article is about.

The three shapes, side by side, in the only terms that decide between them:

Shape What the service holds Who changes the value What a change costs
client in every service configuration, cache, refresh a human, in an admin page nothing at request time, a runtime layer forever
configuration with the deploy one value, read at start whoever cuts the release a deploy per change
decided at the edge nothing a human, at one participant one more field in the call metadata

Martin Fowler's Feature Toggles is still the clearest public breakdown of the categories and their lifetimes, and the vendor-neutral OpenFeature specification is worth reading for what a flag evaluation is actually made of - a key, an evaluation context, a default. Both describe the first shape, because that is what most systems need. I am not arguing they are wrong. I am saying the first shape has a price that a fleet of small services pays repeatedly.

What a flag client in a service drags in

Four things, and each one is a separate piece of runtime:

  • Its own configuration - where the store is, how to authenticate to it, what the defaults are.
  • Its own cache - because you are not making a network call per branch.
  • Its own refresh - because a cache that never updates is a constant with extra steps.
  • Its own admin page - somewhere a human changes the value and sees what it currently is.

Four plates in a row - configuration, cache, refresh, admin page - above one wide bar reading: all four of these are runtime, with a note that the audit found three hand-written caches in a single service

There is a standing rule on my side that decides this before taste gets a vote: the platform owns the runtime, and writing your own version of a platform layer is forbidden. The service writes its input, its processing and its output. Everything else - the cache, the configuration, the scheduling, the metrics registry - is an import.

So "a flag client in every service" is not a preference I dislike. It is a request to hand-write a runtime layer in N services, in a fleet where that is already prohibited, and where an audit has already found the exact same shapes written by hand three times over in a single service.

One line, and it is the only thing here I would defend without qualification:

A "small" feature that has state and a refresh is runtime.

How I do it - one example

One behaviour branch, enabled for some callers and not others. Neutral domains: an entity belongs to an account, and an order is written against it.

  1. The edge decides, once per request, whether the new branch is on for this caller. It is the only participant that has the caller resolved and the only one that has to be told anything.
  2. The result travels in the call metadata, next to the claim about who arrived - the same channel, the same lifetime, the same rules.
  3. The handler reads a finished yes or no. No lookup, no client, no cache, no default value to get wrong.

And the part that is easy to forget until it bites: the mark has to survive the queue boundary the same way the call context does. If the synchronous half of the work runs the new branch and the published message goes out without the mark, the background handler picks up the old branch - and now one user action ran two different versions of the same logic.

  edge          handler         publish      background      next call
  decides  -->  new branch -->   ...     |   old branch  -->  old branch
  mark: yes     mark: yes                |   mark: -         mark: -
                                         |
                                   the mark stops here
Enter fullscreen mode Exit fullscreen mode

Four steps in a row - edge decides, handler, background handler, next call - the first two carrying a yes mark and running the new branch, the last two carrying no mark and running the old branch, with a dashed break in the gap where the queue boundary drops it

That failure is worse than the branch being off everywhere, because it is not visible from either side. The synchronous path looks correct. The background path looks correct. Only the pair is wrong.

What actually travels

Worth being precise, because "the edge decides" is easy to over-read.

What travels is a mark on this request: this call, for this caller, runs the new branch. It is a fact about one request, with the lifetime of one request, sitting in the same place as the claim about who arrived.

What deliberately does not travel is the rule that produced it - the percentage, the account list, the date after which everyone gets it. None of that is a service's business, and none of it belongs in a message that crosses four processes. A service that received the rule instead of the answer would be right back to evaluating flags, only with a worse input.

That distinction is also what keeps the compatibility story sane: a mark is a boolean with a name, and a name plus a boolean is a thing you can add, deprecate and remove. A rule is not.

Why the edge and not the service

Two reasons, and I do not have a third.

The decision needs the caller, and the caller is resolved at the edge. The service receives an input that has already been parsed; whether it also receives the resulting yes or no is a matter of putting one more field next to one that is already there. Deciding in the service means re-deriving something that was already known one hop earlier.

Nothing new appears in the service. No flag configuration, no cache, no background refresh - which is to say, no runtime that then has to be maintained separately in every service that has a flag. The count matters here: one decision point does not scale with the number of services; a client in each of them does.

What this costs

An honest list, not one token item:

  • Product logic moves to the edge. "Is this branch on for this account" is a product decision now living in the door, which is not where anyone looks for it. That is a real navigational cost for whoever debugs it next.
  • The set of marks in the call metadata grows, and it becomes part of the compatibility surface between the edge and everything behind it. Removing one is a fleet-wide change.
  • Local development needs a way to inject a mark by hand. Otherwise a service is only runnable with the edge in front of it, and the developer experience quietly degrades into "you can only test this on a stand".
  • A mark that does not survive the whole chain gives you a half-enabled branch - the failure described above, and the most unpleasant of the lot.
  • What I have just described is a direction, not a running system. There is no flag store on my side and no page to manage flags from a browser. The cost of "its own admin page" is therefore named here, not measured - I am counting a cost I have not paid yet, and it belongs in the list marked as such.

Three cards - a flag store, a page to manage flags, a marked request routed in production - each carrying a badge reading does not exist, above a line saying this part describes a direction and its price rather than a working scheme

When not to do this

  • When the switch is purely technical and lives inside one service. A timeout, a batch size, a worker count - that is configuration, not a flag, and it arrives as an environment variable derived from the resource declaration. Routing it through the edge would be ceremony.
  • When the decision depends on data the edge does not have. This is the same dividing line as with permissions: a question that needs rows is answered where the rows are, which is the service. The edge decides what it can decide from the caller alone.
  • When there is exactly one service. Then there is no fleet to protect from N copies of a runtime layer, and the whole argument is premature.

And the honest boundary of this write-up, stated once more because it matters: there is no flag store here, no admin page for one, and no marked request routed into a new version in production. There are also no numbers - not how many flags, not the TTL of a flag cache, not a refresh delay, not a share of enabled callers - because none of those have been measured on my side. A hole you can see is worth more than a plausible figure.

One line

A flag is an answer, not a source. The service reads the answer; it does not go and fetch one.

The multiplier line

The mechanical half of this is cheap now: threading one more field through the call metadata, adding a parameter to a handler signature, writing the test that fails when the mark does not reach the background handler. What did not get faster is the judgement that a flag is a decision rather than a lookup, and that a piece of state with a refresh loop is runtime no matter how small it looks in the diff. Get that wrong and you get a perfectly written flag client, in every service, passing review because it looks like every other flag client. Speed amplifies whoever set the definitions; it does not supply them.


Access, tracing, rollout - Part 4. Next: what follows from this - a marked request routed into a new version running beside the old one, two generations of code alive at once, and the price that is organisational rather than technical.

If you do this better, tell me where your flag decision is evaluated and how the answer reaches a background handler. If you have been through this, what did your half-enabled branch turn out to break? If you see it differently, say where a client in every service is worth the runtime it brings. How is it solved on your side, and what broke there?

Top comments (0)