I spent a week writing a translator from Datadog monitors to Grafana alert rules. It compiled cleanly, produced sensible-looking PromQL, and passed 18 unit tests.
Then I loaded the output into an actual Grafana and watched an alert fire, drop below its threshold, and keep firing. Forever.
This is what I learned, in the order the mistakes found me.
Why translate monitors at all
OpenTelemetry won instrumentation. Telemetry is portable now — same OTLP, same conventions, pointable anywhere. The industry read that as "the backends are commoditised."
It isn't quite true. The lock-in moved up a layer.
A team with 900 Datadog monitors cannot leave Datadog. Not because the data won't walk — it will. Because of the person-years encoded in DQL: the composite monitors, the anomaly detectors someone tuned in 2023, the notification routing, and the institutional knowledge that monitor #447 firing means check replica lag first.
Every vendor's answer to this is "we're OpenTelemetry-native, migration is easy." That's a claim about the data layer. It's true, and it's irrelevant. The question that actually decides the renewal — what happens to my 900 monitors — gets no answer at all.
So I tried to answer it.
Chapter 1: The query is 20% of a monitor
Here's a Datadog monitor:
avg(last_10m):avg:system.cpu.user{env:prod} by {host} > 90
Translating that to PromQL is trivial. An LLM does it in one shot, correctly:
avg by (host) (avg_over_time(system_cpu_user{env="prod"}[600s])) > 90
Ship it and you have destroyed five behaviours, none of which appear anywhere in the query string:
| Datadog option | What it does | Grafana |
|---|---|---|
no_data_timeframe: 15 |
alert after 15 min without data | fires on the first empty evaluation |
critical_recovery: 75 |
real hysteresis — alert at >90, recover at <75 | recovers at the alert threshold, so it flaps in the band |
evaluation_delay: 60 |
wait 60s for late-arriving points | no equivalent |
new_group_delay: 600 |
don't page for a freshly-booted host | no equivalent |
require_full_window |
refuse to fire on a partial window | always evaluates whatever came back |
The alert still exists after your migration. It just fires differently. And you discover that during an incident, which is why nobody attempts this.
This is the whole reason the project exists: the query is about 20% of a monitor, and the other 80% is undocumented evaluation semantics.
So the translator carries every one of those into an intermediate representation, and anything the target cannot express becomes a caveat attached to the output. Never a silent drop.
Chapter 2: PromQL comparisons filter
Now the bug from the opening.
The emitted rule looked right:
avg by (host) (avg_over_time(system_cpu_user{env="prod"}[600s])) > 90
I drove CPU to 95. It fired. I dropped it to 78. It stayed firing. I checked the data directly:
avg_over_time = 78
threshold = 90
rule state = firing
rule health = ok
PromQL comparison operators filter. When the condition is false, the query returns no series. And Grafana cannot distinguish:
- no series because the condition is false
- no series because the exporter died
Both are NoData. So with noDataState: Alerting — which is exactly what Datadog's notify_no_data: true should map to — the rule fires the moment your metric drops below the threshold, and never clears.
Every monitor in the migration with no-data alerting enabled would have been permanently stuck on.
The fix is Grafana's own two-node shape. The query returns a value; a separate threshold expression evaluates it:
data:
- refId: A
datasourceUid: my-prom
model:
expr: avg by (host) (avg_over_time(system_cpu_user{env="prod"}[600s]))
instant: true
- refId: B
datasourceUid: __expr__
model:
type: threshold
expression: A
conditions:
- evaluator: { type: gt, params: [90] }
condition: B
Now "below threshold" is a false condition and "exporter died" is NoData. Two different states, with two different configured behaviours.
If you write Grafana alert rules by hand and you put the comparison in the PromQL, go and check them. This is not specific to migration.
Chapter 3: for is not no_data_timeframe
While fixing that, I found a second one sitting underneath it.
I had mapped Datadog's no_data_timeframe onto Grafana's for duration. They sound similar. They are completely different things:
-
no_data_timeframe— how long to wait before alerting on missing data -
for— how long a breached threshold must hold before firing
Conflating them meant every monitor with notify_no_data: true sat in Grafana's pending state for fifteen minutes during a real outage, while Datadog would have fired at the next evaluation.
The correct value is 0s, and not as a cop-out. Datadog's time rollup is already baked into the PromQL range selector — avg_over_time(...[600s]). Any non-zero for applies the window twice: average over ten minutes, then hold for ten more.
// Wrong. Two different concepts, one variable.
const forDuration = ir.noData.timeframeSeconds
? `${ir.noData.timeframeSeconds}s`
: `${ir.evaluation.windowSeconds}s`;
// Right. The window is in the range selector; `for` must not repeat it.
const forDuration = '0s';
no_data_timeframe genuinely has no Grafana equivalent. It's reported as a caveat rather than approximated.
Chapter 4: Datadog's own monitors broke my parser
I built a realistic estate in a live Datadog account — 31 monitors across golden signals, host saturation, Kubernetes, Postgres, Kafka, SQS, and business metrics — and started streaming telemetry into it.
Then Datadog auto-provisioned nine of its own recommended monitors once it detected the integrations.
Those nine turned out to be a far better test corpus than anything I had written, because they use syntax I didn't know existed:
max(next_3d):forecast(
(1 - avg:system.disk.free{*} by {host,device}
/ avg:system.disk.total{*} by {host,device}) * 100,
"seasonal", 1, interval="60m", seasonality="weekly"
) >= 99
Keyword arguments. interval="60m", seasonality="weekly", direction='above', count_default_zero='true'. My lexer had never seen an = that wasn't part of >=. None of my hand-written fixtures contained one.
And it failed badly. A LexError escaped the compiler and killed the entire run — one unparseable monitor took down all forty. For a tool whose entire job is reporting honestly on a large estate, that's the worst possible failure mode.
} catch (err) {
// ANY failure reading the query is a refusal, not a crash.
findings.push(caveat('unparseable_query',
`${err.name}: ${err.message}`,
{ level: Level.BLOCKER, ddValue: dd.query }));
return finish(dd, null, findings);
}
The lesson I'd pass on: your fixtures lie. Mine were realistic, reviewed, and still wrong in ways only a real account exposed. If you're building anything that parses a vendor's DSL, get real input early — and note that Datadog hands you a free corpus of professionally-written monitors the moment telemetry arrives.
The Datadog API taught me four more constraints by rejecting my monitors outright:
- forecast monitors need a future window (
next_1w), notlast_* - that window must be between 12h and 3mo
- forecast supports only
min/maxaggregators, notavg - service checks require a grouping (
.by("host"))
Chapter 5: Refusing is a feature
Against the 40 live monitors:
40 monitors
├─ 0 translated exactly 0.0%
├─ 31 translated with caveats 77.5%
└─ 9 UNTRANSLATABLE 22.5%
Nine refused, each by name with an actionable reason:
3 anomalies() proprietary model — the band can't be derived from the query
2 forecast() proprietary extrapolation
1 outliers() proprietary inter-series clustering
1 log monitor different query grammar
1 composite references other monitors by id
1 service check no query — aggregates check statuses
This is the part I'd defend hardest. Asked to translate anomalies(), a language model will happily emit a plausible static threshold. It looks right. It is not right, and nobody audits the alert that still exists.
A confidently-wrong translation of your seven most important monitors is strictly worse than a missing one, because the missing one gets noticed.
Chapter 6: Why deterministic, and not an LLM
For a single monitor, an LLM is fine. Genuinely — it'll do it correctly right now.
Three reasons the pipeline is deterministic anyway:
- Auditability. "Why did monitor 447 translate this way?" — "the model decided" is not an answer when that rule guards production.
- Consistency at scale. 900 monitors through an LLM is 900 slightly different decisions. Monitors 47 and 612 share an idiom and translate two different ways, which becomes unmaintainable the first time someone edits one.
- Diffability. The output gets compared month over month. Anything whose key ordering wanders produces noise that hides real change.
Tested, not asserted:
test('DETERMINISM: identical input produces byte-identical output', () => {
const once = stableStringify(monitors.map(compileMonitor).map(emit));
const twice = stableStringify(monitors.map(compileMonitor).map(emit));
assert.equal(once, twice);
// And key order must not depend on input key order.
const shuffled = monitors.map(m =>
Object.fromEntries(Object.keys(m).reverse().map(k => [k, m[k]])));
assert.equal(once, stableStringify(shuffled.map(compileMonitor).map(emit)));
});
There is one legitimate place for a model: proposing translations for the minority that fail deterministic compilation — with a differ as the judge, and nothing shipping until parallel evaluation agrees. Arithmetic decides; the model assists.
Chapter 7: Proving it, instead of claiming it
Every caveat the compiler emits started life as something I believed about Grafana. That's not good enough when the caveat is the product.
So the repo ships a docker-compose harness — Prometheus, Grafana, and a controllable exporter whose values a test drives directly. No API keys: Grafana provisions alert rules from files on disk.
The headline claim under test was the hysteresis caveat:
Datadog recovers at a separate threshold, so an alert at >90 recovering at <75 stays firing between 75 and 90. Grafana recovers at the alert threshold, so this monitor will flap in that band.
Datadog monitor: critical 90, warning 80, critical_recovery 75.
1. CPU -> 95 firing after 9s
2. CPU -> 78 (inside Datadog's 75-90 hold) inactive after 10s
3. stop sending data entirely firing after 0s
Step 2 is the difference. At 78, Datadog would still be holding. Grafana clears.
The caveat is now observed rather than remembered.
A confession on that harness: three separate failures during this were my test, not the translator. Each time I changed where I read state from without changing what I expected to find. Grafana reports two vocabularies and they are easy to conflate:
rule level firing | pending | inactive
instance level Alerting | Pending | Normal | NoData | Error
Threshold behaviour must be read per instance — the rule is avg by (host), so a stale series from an earlier run on a different host keeps the rule firing and makes your test lie. But NoData is raised as a separate DatasourceNoData instance carrying none of your labels, so it's only visible at rule level.
What's still not verified
The Grafana side is measured. The Datadog side is still partly reasoned.
I have not sat with a live monitor, cut the data, and recorded exactly when no_data_timeframe fires. I have not walked a metric through a recovery band on the Datadog side and watched it hold. Those claims come from understanding rather than observation, and the honest thing is to say so in the README rather than let a confident-sounding table imply otherwise.
That's next, and the account now exists to do it.
Also missing, and it's the real product: the differ. Both engines evaluating the same telemetry for thirty days, producing "847 identical, 41 differ only in evaluation delay, 12 untranslatable." That table is what someone takes into a renewal conversation. Everything above is just the compiler that sits in front of it.
Try it
git clone https://github.com/Rocketgraph/translate
cd translate
npm test # 25 tests, zero dependencies
node bin/ddtranslate.mjs fixtures/sre-standard.json
And to watch the behaviour rather than trust me:
cd demo
docker compose up -d --build
node provision.mjs ../fixtures/sre-standard.json
node verify.mjs
MIT. Zero runtime dependencies — npm test needs nothing installed beyond Node.
Issues and PRs welcome, particularly if you have monitors it refuses that you think it shouldn't. That list is the most interesting part of the project.
Disclosure: I build an observability product. This is not it — it translates to Grafana and Prometheus, and works fine for people who never buy anything from me.
Top comments (0)