DEV Community

Cover image for A Promise With No Metric Under It Is Not Declared
Anton Brilliantov
Anton Brilliantov

Posted on

A Promise With No Metric Under It Is Not Declared

A promise is a metric with a target on it. Without the metric underneath, it is a sentence.


👋 Hi, I'm Anton - a software engineer working mostly in PHP/Symfony and Go, currently carving a live PHP monolith into Go services. This block is about what a service is to everyone else: whose it is, who calls it, what it promises. Part 3 was the call map. This part is the last row of that card - the commitments - and the thought I want to share is a narrow one: what has to exist under a promise before it is allowed to be written down. Running notes are on my GitHub: github.com/brilliant-almazov.

This is how I do it right now, with the price attached - maybe you already do it better, maybe you see it differently.


The measurement, before the argument

The repository holds a machine-generated snapshot of every metric the service exposes: name, type, help text, labels, histogram buckets, source, and where it is declared. Two numbers out of that file:

  metrics snapshot ............ 67 records
    from the platform ......... 59
    from the service .......... 6
    dynamic-metric factory .... 1 record, <dynamic>
  domain metrics declared ..... 13

  declared SLAs ............... 0
  formal SLO documents ........ 0
Enter fullscreen mode Exit fullscreen mode

Two numbers side by side - 67 metric records in the snapshot, split into 59 from the platform and 6 from the service plus a dynamic record, against 0 declared SLAs and 0 formal SLO documents, with a strip below reading the material exists, the promise does not

Both halves of that are true at the same time, and together they are the whole subject of this part. The material for a promise is already there and generated. The promise is not. Everything below is the reasoning about what would turn one into the other, and what that would cost - not a report on a thing I am running.

Two definitions, two lines

  • SLO is an internal target, expressed in the same metrics the service already exposes by default - not in a separate measurement invented for the occasion.
  • SLA is an external promise, and it only means anything when there is a declared SLO underneath it and a metric the SLO is measured by.

Which gives the rule this part is named after: a promise with no metric under it is not declared. Not "declared and measured later" - not declared.

The case: 13 metrics, and an empty queue that is not a failure

Here is the concrete episode the rule comes from, and it starts with what the default metrics could not answer.

The runtime gives a service its metrics for free: liveness and readiness, build version, gRPC metrics, broker-driver metrics (amqp_delivery_total{queue,handler,status}, amqp_delivery_duration_seconds, amqp_delivery_in_flight), outbox-runner metrics, connection-pool metrics, Go runtime metrics. That is 59 of the 67 records, and none of them answer three questions that get asked whenever a write path is late:

  • how many rows are waiting to be published right now,
  • how old is the oldest one,
  • and why did a publish fail.

The platform metrics have no notion of a redelivery, no classification of a failure reason, and no duration of the drain cycle. Those are properties of this service's domain, not of the transport.

So the service declares 13 domain metrics as a layer on top of the platform ones, never as a replacement. Four of them carry this article:

  svc_outbox_pending_rows                    gauge      rows waiting to publish
  svc_outbox_oldest_age_seconds              gauge      age of the oldest waiting row
  svc_outbox_relay_cycle_duration_seconds    histogram  one SELECT -> publish -> DELETE cycle
  svc_outbox_publish_failures_total          counter    labels: type, reason
Enter fullscreen mode Exit fullscreen mode

Three rules hold that layer in shape, and each of them exists because of a specific way the layer can rot:

  • Names are declared in exactly one place - two types, Names and Labels. There is no second copy of a metric-name literal anywhere in the repository, because the second copy is how a dashboard quietly starts querying a name nothing emits.
  • The failure reason is not parsed out of an error message. A classifier compares the error with a list of sentinels through errors.Is and maps it to one of connection, channel, nack, timeout. Error text changes with a library upgrade; a sentinel does not.
  • Drift of the metrics snapshot fails the build, as its own check. Add a metric, rename a label, change a bucket - the build is red until the snapshot is regenerated.

A fourth metric in that list is there for a reason worth stating on its own. Delivery is at-least-once: the broker does not deduplicate, and a crash between the publish and the DELETE sends the same row a second time with the same message id. That is normal behaviour, not a defect - so svc_outbox_redelivered_total exists to make the rate of it visible, because the incident is a spike in redeliveries, not the existence of one. A promise about this path that had no such counter under it would be a promise about a number nobody was keeping.

And the price of that layer, which is the part I did not expect to have to pay: "empty" had to be defined explicitly. That one gets its own section below, because it only makes sense next to the table.

How this is normally done

Nothing here is novel, and the industry answer is well documented. The SRE workbook chapter on implementing SLOs sets out the usual shape: an SLI is a measurement, an SLO is a target on it over a window, an error budget is what is left of that target, and an SLA is the contractual consequence of missing it. The SLO chapter of the SRE book makes the point that targets are chosen from what users actually notice, not from what is convenient to graph. On the measurement side, the Prometheus notes on histograms and quantiles are worth reading before promising anything at a percentile: a quantile computed from buckets is an estimate with a known error, and a promise made on it inherits that error.

The error budget is the part of that machinery I keep coming back to, because it is what makes a target a decision rather than a wish: if the objective is a share of successful responses over a window, then the complement of it is a quantity that gets spent, and spending it has to change what the team is allowed to ship. That mechanism is only available once the share is measured continuously - which is, again, the same precondition.

Where those write-ups put the artefact is the one thing I would flag. The SLO tends to live as a document owned by a reliability team, separately from the service. That is a reasonable structure with more people in it than I have. The direction I am arguing for is narrower and follows from the previous three parts of this block - the promise belongs in the same declaration everything else about the service is rendered from.

Four quantities, and the metric each one is measured by

This is the table the whole part exists for. Four promise quantities, each attached to a metric that already exists.

Promise quantity Metric under it Source Target If there is no data
Share of successful responses the platform's gRPC metrics platform from the repo no traffic - not success
Latency at a percentile svc_resolve_duration_seconds both from the repo no observations in the window
Queue processing lag svc_outbox_relay_cycle_duration_seconds both from the repo idle consumer, or a stalled one
Age of the oldest unpublished row svc_outbox_oldest_age_seconds service from the repo 0 - the outbox table is empty

Each of the last three rows has a companion the promise would be read together with: amqp_delivery_duration_seconds next to the resolve histogram, amqp_delivery_in_flight and amqp_delivery_total{queue,handler,status} next to the relay cycle, and svc_outbox_pending_rows next to the age gauge. The both in the source column means exactly that - one metric from the runtime, one declared by the service.

Four promise quantities as rows - share of successful responses, latency at a percentile, queue processing lag and age of the oldest unpublished row - each with the metric that measures it, a platform or service source pill, a target marked to be taken from the repository, and what an absence of data means

Two of those four rows are about a synchronous call and two are about a queue, and they fail in opposite directions. A synchronous promise breaks loudly: the caller is waiting and gets an answer or does not. A queue promise breaks quietly, because from the outside a queue that is draining slowly and a queue with nothing in it produce the same silence. That asymmetry is the reason the last two rows need an explicit "no data" column at all, while the first two mostly need a window.

The target column is deliberately unfilled. Percentiles, thresholds, windows and error budgets are numbers I do not have written down, and inventing them here would be exactly the failure this part is about: a promise with nothing under it, dressed up in a number to look measured. When those targets get chosen, they get chosen against these metrics, in this table, in the declaration - and the column stops saying from the repo.

What "no data" is allowed to mean

The outbox table holds one row per domain event and the row is deleted after a successful publish. So the normal state of a healthy system is an empty table - and an empty table has to be distinguishable from a broken exporter, from a stalled drain, and from no traffic at all.

The definition that got written down:

  pending_rows = 0   published_total rising      normal - the drain keeps up
  pending_rows > 0   published_total flat        the drain is stuck
  pending_rows = 0   published_total flat        no traffic - and no promise applies
Enter fullscreen mode Exit fullscreen mode

svc_outbox_oldest_age_seconds is 0 when the table is empty - reported as zero, not as the age of the last row that was deleted. That is a decision, not a detail: an age that keeps counting after the queue drained looks exactly like a backlog nobody is clearing.

Three queue states in a row - pending zero with published rising marked normal, pending rising with published flat marked drain stuck, pending zero with published flat marked no traffic and therefore no promise - over a strip reading age is 0 when the table is empty

A promise you cannot tell apart from "there was no traffic" is not a promise. That sentence cost more design time than the metrics themselves.

Why the promise belongs next to the service

Two reasons, and neither is about tidiness.

First, it is the same rule as for any requirement: the criterion has to be checkable. "Works fast" is not a criterion. "Sustains N requests per second below M milliseconds" is one, because there is a procedure that returns true or false. A promise is a requirement pointed at the outside world, so it inherits that constraint whole.

Second, a promise that lives somewhere else drifts from the thing it promises about. The service card in this block is rendered from the manifest - the daemons, the resources, the migrations, the schedules. If the commitments are one more row of that card, they change in the same commit as the thing that has to hold them. If they live in a separate document, they change when someone remembers, which is the same failure mode Part 1 of this block was about, just with a longer fuse.

What already holds the shape

The mechanical half of this is done, and it was done for other reasons:

  • names in one place, so a promise refers to a name that provably exists;
  • the reason label produced by errors.Is against sentinels, so a failure class survives a library upgrade;
  • the tenant identifier deliberately kept out of labels, because tenants are many and the cardinality would explode - the per-tenant cut is a query against data, not a metric;
  • the queue gauges published by the worker binary only, never by server, because two sources for one gauge produce a graph that jumps;
  • drift of the snapshot failing the build.

Five mechanisms that keep the metric layer honest - names declared once, failure reason from errors.Is sentinels, tenant identifier kept out of labels, gauges published by one binary only, snapshot drift failing the build - each with the failure it prevents

What is not in the service repository, and will not be: dashboard files and alert files. The dashboard is assembled by a generator alongside the rest of the fleet's dashboards. Nine alert rules over these metrics exist as text and are set up on the fleet monitoring side. So a promise made here would be measured by metrics from this repository and watched by rules that are not - which is a seam worth naming before anyone signs anything.

What it costs

Four prices, and the first one is the one that stings.

A quantity with no metric under it does not make it into the promise. This is not a hypothetical trade-off - part of what I would like to commit to has nothing measuring it today, and the rule says it stays out. Following the rule feels like losing something, every time.

Low cardinality means no per-tenant promise. Keeping the tenant identifier out of labels is right for the metric system and expensive for the wording: the commitment can only be phrased for the service as a whole, and any per-tenant cut moves to queries against data, on a different cadence and with a different owner.

A promise pinned to a metric name makes renaming that metric a breaking change. Today a rename is a red build and a regenerated snapshot. With a promise on top of it, a rename is a change to something external.

A declared promise needs someone to hold the gate when it is missed. Without that name - Part 2 of this block, ownership as a field - a commitment is one more line that renders on a card and changes nothing when it is broken.

When not to do this

No external consumer with an expectation. An SLA is a promise to somebody. The consumer and their expectation come first; the promise is written afterwards, in their terms. Writing it in the other order produces a number that is technically monitored and answers no one's question.

One service and one consumer who sits next to you. A promise is a coordination mechanism for people who cannot ask each other directly. When they can, the mechanism costs more than the conversation it replaces, and the metric alone is enough.

A metric that is still changing weekly. While a measurement is still being reshaped along with the code, hanging a commitment on it means either the commitment or the code stops moving. Let it settle first.

And the caveat this part would be dishonest without. There is no declared SLA to any external consumer here, and no formal SLO document. What exists is a generated metric snapshot, 13 domain metrics on top of the platform's, an explicit definition of what an empty queue looks like, and a drift check that keeps all of it from rotting. Everything above is the reasoning about how that becomes a commitment, with the price list attached. It is not a status report.

The multiplier line

Wiring a metric is mechanical: declare the name, register it, emit it, regenerate the snapshot, watch the check go green. An automated executor does that faster than I do and never forgets the snapshot. What it does not do is decide which quantity is worth promising, and what "no data" is allowed to mean for it. Those two judgements are the entire exercise - and speed applied before they are made produces a well-formatted, thoroughly monitored promise about the wrong thing.


Services and commitments - Part 4. Next: data volume and lifetime as a line of the requirement - rows per day, the total by the end of retention, the read-to-write ratio, and dropping a partition instead of deleting rows, decided before the first table exists rather than after it fills up.

If you do this better, tell me which of your promises is measured by a metric your service already exposed anyway, and which one needed a new measurement built for it. If you have been through this, what did your "no data" turn out to mean the first time it mattered? If you see it differently, say where a promise is worth declaring before there is a metric under it. How is it solved on your side, and what broke there?

Top comments (0)