Metrics backends do not bill you for traffic. They bill you for distinct time series, and one careless label — user_id, a raw URL path, a build SHA — multiplies that count by thousands overnight while your request volume stays flat. Before you compare vendors, count your active series and find out which labels create them; the number you get decides both the price and whether self-hosting is realistic.
This is the part of observability pricing that the pricing pages assume you already understand. Here is how to measure it, how the three main cost models respond to it, and where each one hurts.
Why does the metrics bill grow when traffic doesn't?
A time series is one metric name plus one unique combination of label values. Storage and billing scale with the number of those combinations, not with how often you write to them.
The multiplication is the whole problem. Take a request-duration histogram with route, method, and status:
- 40 routes × 5 methods × 6 status codes = 1,200 series
- a Prometheus histogram with 12 buckets stores one series per bucket, plus
_countand_sum→ ×14 - that's ~16,800 series from a single instrumented metric
- multiply by replicas, because
instanceis a label too: 10 pods → ~168,000
Now add one more label with unbounded values. A customer_id on a 500-customer SaaS multiplies that by 500. Traffic did not change. Your series count went from six figures to eight.
The symptoms arrive before the invoice does. Self-hosted Prometheus grows RSS until the OOM killer takes it (exit code 137 in the container events, with no application leak to find). If you set guard rails, you instead see scrape failures on the targets page reading sample limit exceeded or label_limit exceeded, and the metrics for that target silently stop. On a managed backend you get a quota email, or — worse — nothing at all until the monthly bill lands with an "active series overage" line.
Cardinality is the only metrics capacity number worth monitoring, because every other symptom — OOM, dropped scrapes, overage fees — is downstream of it.
How do I find high-cardinality metrics in Prometheus?
Don't estimate from your instrumentation code; measure the running system. Prometheus exposes its own head block stats.
Start with the total, which is the number every vendor quote is based on:
prometheus_tsdb_head_series
Then find which metric names own that total:
topk(10, count by (__name__)({__name__=~".+"}))
And once you know the offending metric, find which label inside it is doing the multiplying:
count(count by (customer_id) (http_request_duration_seconds_count))
That last one returns the number of distinct values for one label. Run it per suspicious label; the one that returns thousands is your bill. The HTTP API gives the same answer without the UI:
curl -sG 'http://localhost:9090/api/v1/query' \
--data-urlencode 'query=topk(10, count by (__name__)({__name__=~".+"}))' \
| python3 -c 'import json,sys; [print(int(float(r["value"][1])), r["metric"].get("__name__")) for r in json.load(sys.stdin)["data"]["result"]]'
Prometheus also ships /api/v1/status/tsdb, which returns the top label-value counts directly — useful when you don't yet know which labels to suspect.
Fix it at the scrape, not in the application, so the fix applies to third-party exporters you don't control:
scrape_configs:
- job_name: api
sample_limit: 50000 # fail loudly instead of OOMing silently
label_limit: 30
static_configs:
- targets: ['api:9090']
metric_relabel_configs:
# collapse an unbounded label away, keeping the rest of the metric
- regex: 'customer_id|request_id'
action: labeldrop
# drop a whole metric that is cardinality-heavy and never queried
- source_labels: [__name__]
regex: 'http_request_duration_seconds_bucket'
action: drop
A warning about that last rule: dropping histogram buckets means you can no longer compute percentiles for that metric at all, only averages. Pre-aggregate with a recording rule before you drop, if you still need the shape:
groups:
- name: api-aggregates
interval: 30s
rules:
- record: job:http_request_duration_seconds_bucket:sum_rate
expr: sum by (job, le) (rate(http_request_duration_seconds_bucket[1m]))
Label-drop rules at the scrape layer are the cheapest cardinality fix you will ever deploy, and they work on exporters whose source you cannot change.
What do the three cost models actually charge for?
As of October 2026, the mainstream options bill on different units, which is why a quote that looks cheap on one can be absurd on another. Verify current rates on each vendor's pricing page before you commit — the models below are stable, the numbers are not.
| Option | Billed unit | What blows the budget | What you still pay for |
|---|---|---|---|
| Self-hosted Prometheus | nothing (your infra) | RAM per active series; retention on local disk | On-call, upgrades, no HA by default |
| Prometheus + VictoriaMetrics or Mimir/Thanos | nothing (your infra) | object storage + compute for long-term queries | A real operational surface: compaction, query tiers |
| Grafana Cloud | active series + data points per minute | series count, and scrape interval you forgot to raise | Per-seat and log/trace volume on top |
| Datadog | hosts, plus custom metric time series | custom metrics multiplying past the per-host allotment | Separate SKUs for logs, APM, synthetics |
Two model details matter more than the headline price. First, scrape interval is a billing lever on anything that charges by data points: halving a 15-second scrape to 30 seconds halves the data points for the same series count, and for most infrastructure metrics you will never notice. Second, Datadog separates ingested from indexed custom metrics through its Metrics without Limits feature, so you can send everything but only pay query-side for the tag combinations you actually keep queryable — that is the knob to reach for before you start deleting instrumentation.
If you want the managed version of Prometheus without changing your exporters, Grafana Cloud is the one that takes plain remote_write and bills on active series you can count in advance with the queries above. If your team wants metrics, traces, logs, and dashboards from one vendor and will trade cost control for that consolidation, Datadog is the one that covers the whole surface — with the caveat that each pillar is its own line item and custom metrics are the one most likely to surprise you. If you have the series count under control and want to stay self-hosted past what single-node Prometheus retains, VictoriaMetrics is the one that handles high series counts on modest hardware without you running a multi-component stack.
Pick the vendor whose billing unit you can measure today, not the one with the lowest number on the pricing page.
When is a managed metrics backend worth paying for?
The honest break-even is not storage cost. Single-node Prometheus is cheap and will run for years on a small VM. What you are buying is someone else's answer to retention and high availability: Prometheus alone keeps local data for a fixed window, has no replication story, and querying a year of history means adopting Thanos, Mimir, or VictoriaMetrics — each of which is a distributed system you now operate.
Pay for managed when any of these is true: you need more than a few months of queryable history, you are the on-call rotation and metrics going down during an incident is unacceptable, or you want alerting that survives the cluster the alerts are about. Stay self-hosted when your series count is modest, your retention need is weeks, and you already run stateful workloads competently.
The trap in both directions is the same: teams self-host to save money and spend an engineer-week per quarter on it, or buy managed and never tune cardinality, so the thing they were avoiding — an unbounded cost — shows up as an invoice instead of an OOM.
Managed metrics is worth it the first time long-term retention or HA appears on your requirements list, because that is the point where the self-hosted option stops being one process.
FAQ
How do I find high cardinality metrics in Prometheus?
Query topk(10, count by (__name__)({__name__=~".+"})) to rank metric names by series count, then run count(count by (<label>) (<metric>)) on each suspicious label to see how many distinct values it has. /api/v1/status/tsdb returns the same top-label breakdown without writing queries.
What counts as a custom metric in Datadog?
Any metric you submit that isn't emitted by an official Datadog integration, counted as one custom metric per unique combination of metric name and tag values. Plans include an allotment tied to host count, and anything past it is billed per custom metric — which is why one unbounded tag, not one new metric name, is what moves the bill.
Does dropping labels lose historical data?
No — metric_relabel_configs applies at scrape time, so already-stored series stay queryable until retention expires. New samples simply stop carrying the dropped label, which means the old and new series are distinct and your dashboards will show a seam at the deploy.
Bottom line
Count prometheus_tsdb_head_series and the top-10 metric names before you talk to any vendor; that single number determines your price under every model. Small teams with weeks of retention and no HA requirement should stay on single-node Prometheus and spend the effort on scrape-level label drops instead. Teams that need long retention or metrics that survive their own outage should buy Grafana Cloud if they want Prometheus semantics and predictable series-based billing, or Datadog if one-vendor coverage across logs and traces is worth the per-pillar line items. Whichever you choose, put a cardinality panel on a dashboard you actually look at — the bill and the OOM are the same bug caught at different stages.
Top comments (0)