A promise is a metric with a target on it. Without the metric underneath, it is a sentence.
👋 Hi, I'm Anton - a software engineer working mostly in PHP/Symfony and Go, currently carving a live PHP monolith into Go services. This block is about what a service is to everyone else: whose it is, who calls it, what it promises. Part 3 was the call map. This part is the last row of that card - the commitments - and the thought I want to share is a narrow one: what has to exist under a promise before it is allowed to be written down. Running notes are on my GitHub: github.com/brilliant-almazov.
This is how I do it right now, with the price attached - maybe you already do it better, maybe you see it differently.
The measurement, before the argument
The repository holds a machine-generated snapshot of every metric the service exposes: name, type, help text, labels, histogram buckets, source, and where it is declared. Two numbers out of that file:
metrics snapshot ............ 67 records
from the platform ......... 59
from the service .......... 6
dynamic-metric factory .... 1 record, <dynamic>
domain metrics declared ..... 13
declared SLAs ............... 0
formal SLO documents ........ 0
Both halves of that are true at the same time, and together they are the whole subject of this part. The material for a promise is already there and generated. The promise is not. Everything below is the reasoning about what would turn one into the other, and what that would cost - not a report on a thing I am running.
Two definitions, two lines
- SLO is an internal target, expressed in the same metrics the service already exposes by default - not in a separate measurement invented for the occasion.
- SLA is an external promise, and it only means anything when there is a declared SLO underneath it and a metric the SLO is measured by.
Which gives the rule this part is named after: a promise with no metric under it is not declared. Not "declared and measured later" - not declared.
The case: 13 metrics, and an empty queue that is not a failure
Here is the concrete episode the rule comes from, and it starts with what the default metrics could not answer.
The runtime gives a service its metrics for free: liveness and readiness, build version, gRPC metrics, broker-driver metrics (amqp_delivery_total{queue,handler,status}, amqp_delivery_duration_seconds, amqp_delivery_in_flight), outbox-runner metrics, connection-pool metrics, Go runtime metrics. That is 59 of the 67 records, and none of them answer three questions that get asked whenever a write path is late:
- how many rows are waiting to be published right now,
- how old is the oldest one,
- and why did a publish fail.
The platform metrics have no notion of a redelivery, no classification of a failure reason, and no duration of the drain cycle. Those are properties of this service's domain, not of the transport.
So the service declares 13 domain metrics as a layer on top of the platform ones, never as a replacement. Four of them carry this article:
svc_outbox_pending_rows gauge rows waiting to publish
svc_outbox_oldest_age_seconds gauge age of the oldest waiting row
svc_outbox_relay_cycle_duration_seconds histogram one SELECT -> publish -> DELETE cycle
svc_outbox_publish_failures_total counter labels: type, reason
Three rules hold that layer in shape, and each of them exists because of a specific way the layer can rot:
-
Names are declared in exactly one place - two types,
NamesandLabels. There is no second copy of a metric-name literal anywhere in the repository, because the second copy is how a dashboard quietly starts querying a name nothing emits. -
The failure reason is not parsed out of an error message. A classifier compares the error with a list of sentinels through
errors.Isand maps it to one ofconnection,channel,nack,timeout. Error text changes with a library upgrade; a sentinel does not. - Drift of the metrics snapshot fails the build, as its own check. Add a metric, rename a label, change a bucket - the build is red until the snapshot is regenerated.
A fourth metric in that list is there for a reason worth stating on its own. Delivery is at-least-once: the broker does not deduplicate, and a crash between the publish and the DELETE sends the same row a second time with the same message id. That is normal behaviour, not a defect - so svc_outbox_redelivered_total exists to make the rate of it visible, because the incident is a spike in redeliveries, not the existence of one. A promise about this path that had no such counter under it would be a promise about a number nobody was keeping.
And the price of that layer, which is the part I did not expect to have to pay: "empty" had to be defined explicitly. That one gets its own section below, because it only makes sense next to the table.
How this is normally done
Nothing here is novel, and the industry answer is well documented. The SRE workbook chapter on implementing SLOs sets out the usual shape: an SLI is a measurement, an SLO is a target on it over a window, an error budget is what is left of that target, and an SLA is the contractual consequence of missing it. The SLO chapter of the SRE book makes the point that targets are chosen from what users actually notice, not from what is convenient to graph. On the measurement side, the Prometheus notes on histograms and quantiles are worth reading before promising anything at a percentile: a quantile computed from buckets is an estimate with a known error, and a promise made on it inherits that error.
The error budget is the part of that machinery I keep coming back to, because it is what makes a target a decision rather than a wish: if the objective is a share of successful responses over a window, then the complement of it is a quantity that gets spent, and spending it has to change what the team is allowed to ship. That mechanism is only available once the share is measured continuously - which is, again, the same precondition.
Where those write-ups put the artefact is the one thing I would flag. The SLO tends to live as a document owned by a reliability team, separately from the service. That is a reasonable structure with more people in it than I have. The direction I am arguing for is narrower and follows from the previous three parts of this block - the promise belongs in the same declaration everything else about the service is rendered from.
Four quantities, and the metric each one is measured by
This is the table the whole part exists for. Four promise quantities, each attached to a metric that already exists.
| Promise quantity | Metric under it | Source | Target | If there is no data |
|---|---|---|---|---|
| Share of successful responses | the platform's gRPC metrics | platform | from the repo |
no traffic - not success |
| Latency at a percentile | svc_resolve_duration_seconds |
both | from the repo |
no observations in the window |
| Queue processing lag | svc_outbox_relay_cycle_duration_seconds |
both | from the repo |
idle consumer, or a stalled one |
| Age of the oldest unpublished row | svc_outbox_oldest_age_seconds |
service | from the repo |
0 - the outbox table is empty |
Each of the last three rows has a companion the promise would be read together with: amqp_delivery_duration_seconds next to the resolve histogram, amqp_delivery_in_flight and amqp_delivery_total{queue,handler,status} next to the relay cycle, and svc_outbox_pending_rows next to the age gauge. The both in the source column means exactly that - one metric from the runtime, one declared by the service.
Two of those four rows are about a synchronous call and two are about a queue, and they fail in opposite directions. A synchronous promise breaks loudly: the caller is waiting and gets an answer or does not. A queue promise breaks quietly, because from the outside a queue that is draining slowly and a queue with nothing in it produce the same silence. That asymmetry is the reason the last two rows need an explicit "no data" column at all, while the first two mostly need a window.
The target column is deliberately unfilled. Percentiles, thresholds, windows and error budgets are numbers I do not have written down, and inventing them here would be exactly the failure this part is about: a promise with nothing under it, dressed up in a number to look measured. When those targets get chosen, they get chosen against these metrics, in this table, in the declaration - and the column stops saying from the repo.
What "no data" is allowed to mean
The outbox table holds one row per domain event and the row is deleted after a successful publish. So the normal state of a healthy system is an empty table - and an empty table has to be distinguishable from a broken exporter, from a stalled drain, and from no traffic at all.
The definition that got written down:
pending_rows = 0 published_total rising normal - the drain keeps up
pending_rows > 0 published_total flat the drain is stuck
pending_rows = 0 published_total flat no traffic - and no promise applies
svc_outbox_oldest_age_seconds is 0 when the table is empty - reported as zero, not as the age of the last row that was deleted. That is a decision, not a detail: an age that keeps counting after the queue drained looks exactly like a backlog nobody is clearing.
A promise you cannot tell apart from "there was no traffic" is not a promise. That sentence cost more design time than the metrics themselves.
Why the promise belongs next to the service
Two reasons, and neither is about tidiness.
First, it is the same rule as for any requirement: the criterion has to be checkable. "Works fast" is not a criterion. "Sustains N requests per second below M milliseconds" is one, because there is a procedure that returns true or false. A promise is a requirement pointed at the outside world, so it inherits that constraint whole.
Second, a promise that lives somewhere else drifts from the thing it promises about. The service card in this block is rendered from the manifest - the daemons, the resources, the migrations, the schedules. If the commitments are one more row of that card, they change in the same commit as the thing that has to hold them. If they live in a separate document, they change when someone remembers, which is the same failure mode Part 1 of this block was about, just with a longer fuse.
What already holds the shape
The mechanical half of this is done, and it was done for other reasons:
- names in one place, so a promise refers to a name that provably exists;
- the reason label produced by
errors.Isagainst sentinels, so a failure class survives a library upgrade; - the tenant identifier deliberately kept out of labels, because tenants are many and the cardinality would explode - the per-tenant cut is a query against data, not a metric;
- the queue gauges published by the
workerbinary only, never byserver, because two sources for one gauge produce a graph that jumps; - drift of the snapshot failing the build.
What is not in the service repository, and will not be: dashboard files and alert files. The dashboard is assembled by a generator alongside the rest of the fleet's dashboards. Nine alert rules over these metrics exist as text and are set up on the fleet monitoring side. So a promise made here would be measured by metrics from this repository and watched by rules that are not - which is a seam worth naming before anyone signs anything.
What it costs
Four prices, and the first one is the one that stings.
A quantity with no metric under it does not make it into the promise. This is not a hypothetical trade-off - part of what I would like to commit to has nothing measuring it today, and the rule says it stays out. Following the rule feels like losing something, every time.
Low cardinality means no per-tenant promise. Keeping the tenant identifier out of labels is right for the metric system and expensive for the wording: the commitment can only be phrased for the service as a whole, and any per-tenant cut moves to queries against data, on a different cadence and with a different owner.
A promise pinned to a metric name makes renaming that metric a breaking change. Today a rename is a red build and a regenerated snapshot. With a promise on top of it, a rename is a change to something external.
A declared promise needs someone to hold the gate when it is missed. Without that name - Part 2 of this block, ownership as a field - a commitment is one more line that renders on a card and changes nothing when it is broken.
When not to do this
No external consumer with an expectation. An SLA is a promise to somebody. The consumer and their expectation come first; the promise is written afterwards, in their terms. Writing it in the other order produces a number that is technically monitored and answers no one's question.
One service and one consumer who sits next to you. A promise is a coordination mechanism for people who cannot ask each other directly. When they can, the mechanism costs more than the conversation it replaces, and the metric alone is enough.
A metric that is still changing weekly. While a measurement is still being reshaped along with the code, hanging a commitment on it means either the commitment or the code stops moving. Let it settle first.
And the caveat this part would be dishonest without. There is no declared SLA to any external consumer here, and no formal SLO document. What exists is a generated metric snapshot, 13 domain metrics on top of the platform's, an explicit definition of what an empty queue looks like, and a drift check that keeps all of it from rotting. Everything above is the reasoning about how that becomes a commitment, with the price list attached. It is not a status report.
The multiplier line
Wiring a metric is mechanical: declare the name, register it, emit it, regenerate the snapshot, watch the check go green. An automated executor does that faster than I do and never forgets the snapshot. What it does not do is decide which quantity is worth promising, and what "no data" is allowed to mean for it. Those two judgements are the entire exercise - and speed applied before they are made produces a well-formatted, thoroughly monitored promise about the wrong thing.
Services and commitments - Part 4. Next: data volume and lifetime as a line of the requirement - rows per day, the total by the end of retention, the read-to-write ratio, and dropping a partition instead of deleting rows, decided before the first table exists rather than after it fills up.
If you do this better, tell me which of your promises is measured by a metric your service already exposed anyway, and which one needed a new measurement built for it. If you have been through this, what did your "no data" turn out to mean the first time it mattered? If you see it differently, say where a promise is worth declaring before there is a metric under it. How is it solved on your side, and what broke there?




Top comments (0)