Volume is a line of the requirement, not something you find out on the day the disk fills up.
👋 Hi, I'm Anton - a software engineer working mostly in PHP/Symfony and Go, currently carving a live PHP monolith into Go services. This block is about what a service is to everyone else: whose it is, who calls it, what it promises. This part takes the least glamorous promise on that list - how much data this thing will hold and for how long - and asks when that number has to exist. Running notes are on my GitHub: github.com/brilliant-almazov.
This is how I do it right now, with the price attached - maybe you already do it better, maybe you see it differently.
Where these lines live
Short primer, so this reads without the earlier parts of the block.
Each of our Go services raises itself from one declaration - a manifest listing its daemons, its databases, its queues, its schedules. That file is not documentation sitting next to the code; it is what the runtime and the deploy pipeline read. The service card, the page that answers "what is this service", is a rendering of that declaration rather than a wiki page somebody maintains.
Some rows of that card come from the manifest, some are observed, one is a human name. This part is about a row that comes from neither: it comes from the requirement, and it has to exist before the first table does.
The four lines
Every requirement that touches storage carries four lines, and they are written before the code:
- how many rows per day, and how many in total at the end of the retention period;
- the ratio of reads to writes;
- the lifetime of a row, and the way it goes away;
- tenant isolation, and what happens when the number of tenants grows by a multiple.
That is the entire list. It is short on purpose - four lines is something a person will actually fill in, and each of them has a mechanism waiting on the other end.
These are not an appendix to the requirements document. They are the input to the storage decision. One table, partitions, application-level shards, a database per tenant - that choice is made from these four lines. If the lines are missing, the choice still gets made; it just gets made implicitly, by whoever writes the first migration, on the basis of what was convenient that afternoon.
The case: the data only arrives
Here is why these lines are not theoretical in the service I am describing.
There is no destructive UPDATE in the domain tables. A change is a new version, not an overwrite. There is no physical DELETE either: "deleted" means a new status carried by a new version, and "this link no longer exists" means a closed interval. The current state is a slice over the history - valid_to IS NULL AND superseded_at IS NULL - not a separate table that somebody keeps in sync. The only UPDATE allowed anywhere is closing an interval, in the same transaction that opens the next one.
That model has one consequence which decides everything downstream: the data only ever arrives. It arrives on correct writes, and it arrives on mistakes too - a wrong value is not fixed in place, it is superseded by a new version, and both rows stay. So "how many of these will there be in a year" stops being a design-review question and becomes an operational one, on day one.
Two smaller decisions push in the same direction. There are no constraints in the tables - no foreign keys, no UNIQUE, no CHECK - only a primary key over a Snowflake identifier and ordinary indexes; integrity lives in the write layer, in the domain invariants and in the tests. And anything that is a category - a vertical, a type, a status, an attachment level - is a dictionary row referenced by identifier, not a string repeated in every table. The first decision means the database will not refuse a row on your behalf. The second is the one lever that actually keeps the row narrow.
The code grew the same way, and that part I can put a number on:
v0.1.0 2026-08-10 1 751 lines 35 files 13 packages tests/code 0.45
HEAD 2026-08-16 61 411 lines 1 540 files 252 packages tests/code 1.25
Six days, thirteen tags. The test-to-code ratio went from 0.45 to 1.25 and has not dropped below 1.14 since v1.0.0. The average file size sits at 39 lines and stays there: the service grows by number of files and packages, not by files getting fatter.
Underneath it, 66 migration files, 1 990 lines. Forward-only - Up and no Down - one table per migration, a comment on every column, and a file that is immutable once it has shipped. A rollback is a new migration going forward, never an edit to an old one.
Nothing in that list is reversible by editing. Not the data, not the schema. Which is exactly why the retention line has to be answered up front, and answered with a mechanism rather than an intention to clean up later.
So retention here is executed by dropping a partition, never by deleting rows. The job that does it is idempotent: it updates the partition-age gauge even on a run where there was nothing to cut.
How this is usually done
Fairly, because I have done it this way too.
Non-functional requirements go into their own section of the document, below the functional ones, and get read once. Volume is estimated by eye during design - "a few thousand a day, probably" - and never checked afterwards. The decision about partitioning or sharding is deferred, on the entirely reasonable grounds that it is premature to solve a problem you do not have yet.
It then gets made after the first incident with growth: under time pressure, on a table that is already large, already indexed the wrong way, and already being read by four callers.
None of those steps is stupid. The problem is only that the last one is the expensive one, and the first three are what make it inevitable.
Requirement line to decision to mechanism to metric
This table is the whole article. Each row starts with one of the four lines, ends with the thing that tells you it is still true, and the middle two columns are not a matter of preference once the first column is filled in.
| requirement line | decision it forces | mechanism | observed by |
|---|---|---|---|
| rows per day, total at end of retention | partitioning at the database level | audit cut into monthly partitions; data verticals partitioned HASH by owner |
svc_audit_partition_oldest_age_days, svc_audit_partition_dropped_total; target volumes - from the repo
|
| read-to-write ratio | reads and writes separated - by layer, by daemon, and at a third level by service |
repository reads and only reads, manager writes and only writes; server and worker are separate daemons of one manifest |
handler durations; outbox gauges svc_outbox_pending_rows, svc_outbox_oldest_age_seconds
|
| lifetime, and how a row goes away | drop the partition, never DELETE
|
retention job in the worker, monthly boundary |
svc_audit_partition_oldest_age_days, svc_audit_partition_dropped_total
|
| tenant isolation, multiple growth in tenants | database per tenant, or application-level sharding, or partitioning | partitioning today; the other two as described directions | from the repo |
Two cells say from the repo instead of carrying a number. The target volumes and the tenant-facing metric are per-service figures that live in the service repository, and I am not going to invent them for an article - a made-up row count would make the table look more finished and be worth less.
One rule constrains what can go in that last column at all: the tenant identifier never becomes a metric label. There are many tenants, and a per-tenant label blows up cardinality. A per-tenant cut is a question you ask the data, not the metrics.
And each of those gauges has exactly one publisher - the worker, never the server - because two processes publishing one gauge produce a graph that jumps between their two versions of the truth and is worse than no graph.
How the partition drop actually works
The retention line is only executable because the rules are boring and written down:
- the boundary is the first day of the month N months back from the current one;
- a partition is dropped only if it lies entirely before the boundary;
- a partition that crosses the boundary is not touched;
- the current and the next partition are not touched;
- the default partition is never dropped, under any condition;
- the job is idempotent, and updates the age gauge even when it cut nothing.
Five of those six rules are about what the job does not touch. That is the useful shape for anything that deletes at scale: the destructive case is one narrow condition and everything else is explicitly excluded, rather than the other way round.
The last rule is the one that makes the gauge trustworthy. A retention job that only reports when it does something is indistinguishable from a retention job that has silently stopped running. This one always writes the age of the oldest surviving partition, so "nothing was dropped this month" and "nothing has run since March" are two different pictures instead of the same flat line.
Why volume is named before the code, not after the launch
Because the shape of the data is fixed before the first line of code, and the shape is what makes volume predictable.
No destructive UPDATE, no physical DELETE, current state as a slice - that decision is taken once, at the top, and every table follows it. Once it is taken, growth is monotonic by construction and you can multiply it out: rows per event, events per day, days of retention. Take the opposite decision and you cannot multiply anything, because rows appear and vanish, and the only honest answer to "how big will this get" is "we will see".
The second thing that follows from the same place: there are no unbounded lists. Every list is keyset pagination - a cursor over int64, a clamp on the page size, total_count only when it is explicitly asked for. OFFSET is not used on large sets, for two separate reasons: its cost grows with the page number, and inserts between two requests shift the window under the reader. The cursor lives in one cross-cutting package and is the same for every domain, so this is one decision rather than one per list handler.
Note what that pair does together. The volume line decides how the data is stored, and the same line decides how it is allowed to be read. Both stop being negotiable the moment the first table exists.
Partitioning is not sharding
These two get used interchangeably in conversation and they are not the same decision.
Partitions live inside one database and are transparent to the planner. You write the same query; the planner picks the partitions. Here: audit in monthly partitions, data verticals partitioned HASH by owner.
Shards are different databases and are opaque. Nothing plans across them for you. The application holds logical shards and a map of their physical addresses, the shard bit lives inside the Snowflake identifier, and any query that needs more than one shard is assembled by the application.
The practical difference is where the cost lands. Partitioning is a property of a table and can be added without touching a single call site. Sharding is a property of the application, and every read path has to know about it - which is why the fourth requirement line matters months before anyone writes a shard map.
What a missing line looks like later
The failure mode is never "we forgot to write down the volume". It is always something else wearing that as a costume.
A list endpoint that was fine for a year starts timing out, and the fix everyone reaches for is an index. The index helps for a month. The actual missing line was the first one - rows per day - and the actual answer was a partition boundary that should have been declared in the first migration.
Or: a report is asked for "by tenant", the metric has no tenant label, and somebody proposes adding one. The missing line was the fourth one. With it, the answer is decided up front and stays decided: tenant cuts are a query against the data, and the label never gets added no matter how convenient it looks in a dashboard.
Or the quietest one: the retention job has not run since a deploy six weeks ago, and nothing says so, because a job with nothing to report reports nothing. That is why the sixth rule of the drop exists at all.
None of these read as a requirements problem when they arrive. They read as a database problem, an observability problem and an ops problem, and they get three separate fixes.
Why it is decided this way and not by taste
The general form of this, which I now apply well outside storage: an argument about "microservice or module", "CQRS or not", "queue or direct call" is settled by a line of a non-functional requirement, not by preference. If no such line exists, there is nothing to argue about - the requirement has not been gathered yet, and the argument is two people comparing their instincts.
The four lines above are the minimum for storage. Without them, the conversation about a database is conducted on feelings, and feelings lose to a table that has been growing for eight months.
What it costs
- The numbers have to be produced before anyone knows them. The person asking for the feature rarely knows the row count. You derive it from the business event that creates the row, then check it against reality later. That is real work, done up front, on an estimate that may be wrong.
- Monthly granularity is all you get. "Delete last week's data" is not something this scheme does. Dropping a partition is a month-sized operation, and that is the whole point of why it is cheap.
- Mistakes grow the table too. With no destructive operations, a bad write is not removed, it is superseded. The volume line therefore has to include an error rate, which nobody enjoys estimating.
- No jumping to page 4 000. Keyset pagination gives you next and previous, not arbitrary offsets. Some interfaces genuinely want that, and they do not get it.
-
The partition key is a commitment.
HASHby owner is fixed at creation. Changing the key later is a data migration, not a configuration change.
When not to do this
When there is little data and it does not grow. All four lines collapse into one unpartitioned table, and that is the correct answer, not a lazy one. Partitioning a table that will hold fifty thousand rows forever buys nothing and costs you a retention job nobody watches and a gauge nobody reads.
When lifetime is regulatory and per-row. If different records have to disappear on different dates set by law, a monthly partition drop does not implement that. You need a different mechanism, and pretending a coarse one is fine is worse than having no mechanism at all.
And the honest boundary of what I am describing here. There are no thousands of per-tenant databases in production on my side - the database-per-tenant option is analysed as a design and its price, not reported as a running installation. There is no shard map in this service either: partitioning exists, application-level sharding is a described approach and nothing more. The two from the repo cells in the table stay empty for the same reason the rest of the article has numbers in it: I would rather leave a visible hole than fill it with something plausible.
The multiplier line
Deriving four numbers from a business event, drafting the migration, wiring the retention job and its gauge - all of that is fast now, and fast in a way that is genuinely new. What did not get faster is deciding that the data model has no destructive operations in it, and that these four lines are therefore mandatory rather than nice to have. That decision is what makes every generated migration correct or wrong, and it is not delegable. Speed amplifies whoever set the constraints; it does not supply them.
Services and commitments - Part 5. Next: the read-to-write ratio taken seriously - the same domain pulled apart into different executors at three depths, and what each depth costs in consistency, latency and moving parts.
If you do this better, tell me which of the four lines you make mandatory and who signs off on the number. If you have been through this, what did your volume estimate turn out to be wrong by, and when did you find out? If you see it differently, say where declaring volume up front costs more than it returns. How is it solved on your side, and what broke there?




Top comments (0)