Three places can decide what a caller may do - only one of them can answer whether this tenant owns this row.
👋 Hi, I'm Anton - a software engineer working mostly in PHP/Symfony and Go, currently carving a live PHP monolith into Go services. The previous part was about who arrived. This one is about the second question - what that caller may do - and the only thing I am confident about is where the answer has to be computed. It is the shortest piece in this block, because it is one rule with one test under it. Running notes are on my GitHub: github.com/brilliant-almazov.
This is how I do it right now, with the price attached - maybe you already do it better, maybe you see it differently.
Where this sits
A short primer, so this reads on its own.
A PHP monolith is being taken apart into Go services. The Go API gateway is the only door into the new services from outside; behind it the extracted services talk to each other over gRPC, and each of them owns its own data. Identity is settled at that door: the edge parses the input once and passes on a claim about who arrived, so by the time a service is running, who is calling is a finished statement rather than a question it has to work out.
Which leaves the second question, and it does not resolve itself the same way.
The thesis, in one line
"Who arrived" is answered before the service runs. "What they may do" is not - and it can be answered in three different places, at three different prices.
Picking one of them is usually discussed as a matter of taste, or of which layer feels tidier. It is neither. Two of the three places physically cannot answer some of the questions.
The three places
- At the edge. The general check - may this caller reach this at all. It is cheap, it happens before any service is involved, and it knows nothing about the data behind the call.
- In the service. The check that needs a row to answer - does this tenant own this row. This is the only place where such a decision is possible at all, because it is the only place the row exists.
- In a shared resolver. A separate service that answers "allowed or not" on request. It removes duplicated rules and gives you one place to look. It also adds a network call on the hot path and a new dependency between the service and something that must be up for it to serve.
Nothing on that list is wrong. They are priced differently, and the price is paid at different times: the edge charges you at design time, the shared resolver charges you on every request.
What the edge can see
Worth being precise, because the edge is the place people overestimate. It holds the claim about who arrived, the method being called and whatever the caller sent. That is enough to answer a large class of questions, and it is genuinely the cheapest place to answer them: nothing has been started yet, no connection has been taken from a pool, no transaction is open.
What it does not hold is a single row of anything. It can be told which tenant is calling; it cannot be told which tenant owns the row that is about to change, because that fact exists only in the storage of the service that owns it. Handing the edge that fact means handing it the data, and we are back to the copy.
One question decides the place
The dividing rule is short enough to fit in a review comment:
Does answering this check need rows from this service?
Two answers, and each names the place:
- No - decide it as early as possible on the request path.
- Yes - decide it in the service, next to the row.
The reason I like this phrasing is that it is a checkable condition, not a preference. It has exactly two answers, anyone can arrive at the same one, and it does not require agreeing on what "business logic" means - a discussion I have never seen finish.
Applying it in a review
The rule is only worth having if two people apply it to the same check and land in the same place. Four checks, run through the question:
| Check | Needs rows from this service? | Place |
|---|---|---|
| may this caller reach this method at all | no | as early as possible |
| is the payload shaped correctly | no | as early as possible |
| does this tenant own this row | yes | in the service |
| is this row in a state that allows the change | yes | in the service |
The last one is the interesting entry. It is not usually filed under permission at all - it reads as ordinary validation - and it lands in exactly the same place for exactly the same reason. That is a good sign about the rule: it stops producing a separate category for checks that were never separate.
One request, two checks
The concrete case is not an incident, and I am not going to dress it as one. It is the ordinary shape of a write.
What we had. A request to change one row needs two different answers before anything is written: this caller is allowed to use this method, and this tenant owns this row. Both are permission questions. Both are asked on the same request. They are not the same kind of question at all.
What that cost when they were treated as one thing. The second check cannot be moved outward unless the data moves with it. To decide ownership somewhere else, that somewhere else needs to know which tenant owns which row - which means a copy of the ownership data, kept somewhere it is not authoritative. A copy starts to drift the day it is made, and a permission copy that has drifted fails in the worst available direction: it answers confidently.
What follows. The first check is settled earlier, without touching data. The second lives where the row lives - and once you say that out loud, it stops looking like security and starts looking like what it is: an invariant of the domain, standing in the same row as every other rule about what may be written. In our services those are expressed as ordinary typed checks - a Rule over an input, a Guard from an input to an output - and the ownership check is one of them, not a separate security layer bolted to the side.
There are no numbers in this episode, so I am not attaching any. Not how many rules exist, not how many checks the edge runs, not what a hop to a resolver would cost us. Those go in when they are taken from the repo, not before. A visible hole is worth more than a plausible figure.
The same rule after the queue
Work does not always finish on the request that started it. Part of it is published as a message and picked up later by a background handler, and at that point the neat picture of "the edge decided the general part" needs re-reading.
The claim about who arrived travels with the message, which is what keeps the later record naming the person rather than the machine. What does not travel is the edge: there is no door in front of a background handler. So the general check is not re-run there, and the data-dependent one still is - by the service that owns the row, exactly as before.
That asymmetry is worth stating rather than discovering: the earlier a decision is made, the more places it has to survive to still be true. A general check made once at the edge is a decision that has to hold for the whole life of the work it authorised, queue hops included.
How this is usually done
Three shapes show up in public write-ups and standards, and all three are defensible:
- Policy as a separate decision service. The vocabulary for this is standardised - the split between the point that enforces a decision and the point that decides it is described in NIST SP 800-162, along with the attribute-based model most such systems use. Rules live in one place and are auditable by querying it.
- A policy library embedded in every service. The same rules, evaluated locally, no hop. The cost moves from latency to distribution: every service carries a version of the rules, and "which version is live where" becomes a question you need an answer to.
- Role checks at the entry layer. The oldest and cheapest arrangement, and the one that quietly stops being sufficient the moment a check needs a row.
The interesting part is that none of the three removes the constraint this article is about. A decision service still has to be given the data to decide on; an embedded library still has to read the row. The question of where the data is survives every one of these designs.
How we do it - one example
One write, spelled out:
request: update one row of entity
check 1 may this caller use this method rows needed: none decided at the edge
check 2 does this tenant own this row rows needed: one decided in the service
check 2 stands next to the other invariants of entity, not in a layer of its own
Check 1 never sees the database. Check 2 cannot exist anywhere else: the row that answers it is in this service's storage, and the service is the only participant that can read it without a copy.
One thing that falls out of writing it this way: the two checks stop competing for the same slot in the code. Check 1 is a property of the route and lives with the route. Check 2 is a property of the write and lives with the write - which also means it runs inside the same transaction as the change it guards, rather than as a separate lookup that was true a moment ago.
Why this way
Two reasons, and I would not add a third.
Moving a data-dependent check outward means moving the data outward. There is no version of that trade where the data stays put. Either the deciding party reads this service's storage - which makes the storage shared, and the service's ownership of it a fiction - or it keeps a copy, and the copy drifts.
A shared resolver is another runtime on the hot path. The rule I work by is that the platform supplies the runtime and a service supplies its input, its processing and its output. Introducing a component that every request must consult, that must be available for any of them to answer, and whose failure mode is "everything is denied" or "everything is allowed" - that is a runtime decision, and it is a much bigger commitment than it looks like when it is drawn as one box.
Neither reason is an argument that a shared resolver is a bad design. Both are arguments that it is a larger decision than the one this article is about, and that the data-dependent check would still be in the service afterwards.
What it costs
An honest list, not one token item:
- The rules are split between the edge and the services. There is no single place where "who may do what" can be read off. Answering that question means reading two things, and knowing that you had to.
- Identical general checks in different services have to be kept identical without shared code. Anything that is generic enough to share belongs in the platform layer; anything that is not drifts quietly.
- With no resolver, auditing the rules is reading code. You cannot ask a service "list everything this caller may do" - nothing is holding that list. It is a real capability we do not have.
- The tenant is not a dimension in the metrics. A tenant identifier does not go into metric labels - the cardinality is not affordable - so "who was denied and why" is a question for queries against data, not a panel someone can open.
- A denial at the edge and a denial in the service look different to the caller. One arrives before any work started, one arrives after part of the request has been read and validated. Making them look the same to a client is extra work that nobody schedules.
- New services can get this wrong silently. The dividing rule is easy to state and easy to skip, and skipping it produces code that works in the happy path and is wrong exactly once.
When not to do this
- When the rules are many and change faster than the code. Then a shared resolver pays for itself, and it is not a debate - rules that change on a different clock than deployments want a different home.
- When a consumer needs "am I allowed to" without performing the action. A decision endpoint is genuinely useful there, and inventing one out of the checks you already run is worse than declaring one.
- When there are two services. The whole discussion is premature; the split costs more than it returns.
- When the rules are owned by someone who does not write code. If the people who decide what is allowed are not the people who deploy, a place they can edit is worth a hop, and no amount of tidy layering substitutes for it.
And the honest boundary of everything above: there is no single permission-decision service on my side. The general check lives at the edge, the data-dependent check lives in the service, and there is no third participant deciding both. This is an account of three places and their prices, not a report on a resolver that runs in production.
The one line
The place of a check is decided by whether it needs data, not by which architectural layer it looks like it belongs to.
The multiplier line
Writing the checks themselves is fast now - a typed rule over an input is a few lines, and generating the twentieth one is not where the time goes. What did not get faster is deciding which of the two questions a given check actually is. Get that wrong and you produce a perfectly written ownership check running in a place that has to be handed a copy of the data to answer - correct code, wrong location, and it will pass review because it looks like every other check. Speed amplifies whoever set the definitions; it does not supply them.
Access, tracing, rollout - Part 2. Next: one request identifier that has to survive from the handler all the way into the queue and back - and the two places the thread reliably snaps.
If you do this better, tell me where you draw the line between the general check and the data-dependent one. If you have been through this, what did your copy of the permission data drift into before anyone noticed? If you see it differently, say where a shared resolver on the hot path was worth its hop. How is it solved on your side, and what broke there?




Top comments (0)