Someone edited a critical alert in the Grafana UI on a Friday afternoon. By Monday the rule was gone — overwritten by a dashboard sync, or a different env's export, or just a click nobody remembered. We still had Prometheus. We still had metrics. We did not get paged when a tenant success rate fell off a cliff.
That is when "alerting as code" stopped being a nice idea and became the only way I would run Grafana alerts.
This is the sequel to my Dockerized Monitoring Stack post. That article got Prometheus, Grafana, and exporters into one Compose file. This one covers what I promised next: provision contact points, notification policies, datasources, and alert rules through the Grafana HTTP API from Jenkins, with dry-run, idempotent apply, and a way back when something is wrong.
I assume you already have Grafana up (I use the Compose stack from that article) and a Prometheus datasource. If you are still designing the failure model, my HA architecture piece maps instance / AZ / Region failures to what you should alert on.
The JSON and scripts below are reference implementations. Point them at a non-prod Grafana first. Check tokens, folder UIDs, and API versions for your Grafana major version before you touch production.
What "alerting as code" means here
It does not mean dumping every dashboard JSON into Git and calling it done. Dashboards can stay file-provisioned (as in the monitoring repo). Alerts are different: they change often, they wake people up, and a silent UI edit is a production incident waiting to happen.
In this setup:
- Git is the source of the intended state (contact points, notification policies, alert rules, datasource defs).
- Jenkins (or any CI) validates, diffs against live Grafana, then applies.
- Dry-run prints the diff and exits 0 without writing.
- Rollback restores the last exported snapshot from the job artifacts.
Prometheus can still evaluate recording/alert rules from prometheus/rules/*.yml. I use both: Prometheus rules for anything that must work even if Grafana is down, and Grafana Alerting for routing, grouping, and human-facing policies. This article focuses on the Grafana side.
Figure: Git → Jenkins → Grafana API → Slack/email (dry-run on PR, apply on merge).
Repo layout
alerting-as-code/
├── grafana/
│ ├── datasources/
│ │ └── prometheus.json
│ ├── contact-points/
│ │ ├── slack-oncall.json
│ │ └── email-warning.json
│ ├── notification-policies/
│ │ └── root.json
│ └── alert-rules/
│ ├── availability.json
│ └── business-sli.json
├── scripts/
│ └── apply_grafana.py
└── Jenkinsfile
One folder per Grafana object type. One file per object (or per rule group). Env-specific values (webhook URLs, folder UIDs) come from Jenkins credentials or a secrets manager — not committed.
I keep UAT and prod as the same files with different secret injection, same idea as env/uat vs env/prod in the monitoring stack. If the rule set truly diverges, split alert-rules/uat/ and alert-rules/prod/.
Grafana API pieces you actually need
Grafana Alerting is driven by the HTTP API under /api/v1/provisioning/... (and a few older /api/ routes for datasources). You need a service account token with Admin (or a custom role that can manage alerting). Store it in Jenkins as GRAFANA_TOKEN. Base URL looks like https://grafana.example.com.
Useful endpoints (Grafana 10/11 style — confirm against your docs):
| Object | List / get | Create / update |
|---|---|---|
| Datasource | GET /api/datasources |
POST /api/datasources or PUT /api/datasources/uid/:uid
|
| Contact point | GET /api/v1/provisioning/contact-points |
POST / PUT .../contact-points/:uid
|
| Notification policy tree | GET /api/v1/provisioning/policies |
PUT /api/v1/provisioning/policies |
| Alert rules | GET /api/v1/provisioning/alert-rules |
POST / PUT .../alert-rules/:uid
|
Always set a stable uid in every JSON file. Without it, every apply creates duplicates and dry-run cannot match objects.
1. Datasource (once)
File-provisioned datasources are fine. If you want the pipeline to own them too:
grafana/datasources/prometheus.json
{
"uid": "prom-main",
"name": "Prometheus",
"type": "prometheus",
"access": "proxy",
"url": "http://localhost:9090",
"basicAuth": true,
"basicAuthUser": "admin",
"secureJsonData": {
"basicAuthPassword": "${PROMETHEUS_PASSWORD}"
},
"jsonData": {
"httpMethod": "POST",
"manageAlerts": true
},
"isDefault": true
}
${PROMETHEUS_PASSWORD} is substituted at apply time. Never commit the real password. The monitoring stack already put basic auth on Prometheus — Grafana must use the same credentials or every alert evaluation fails with 401 and you get no pages (or false "datasource error" alerts).
2. Contact points
Two contact points cover most teams I work with: critical → Slack (or PagerDuty), warning → email.
grafana/contact-points/slack-oncall.json
{
"uid": "cp-slack-oncall",
"name": "slack-oncall",
"type": "slack",
"settings": {
"recipient": "#oncall-alerts",
"title": "{{ .CommonLabels.alertname }}",
"text": "{{ .CommonAnnotations.summary }}\n{{ .CommonAnnotations.description }}"
},
"secureSettings": {
"url": "${SLACK_WEBHOOK_URL}"
}
}
grafana/contact-points/email-warning.json
{
"uid": "cp-email-warning",
"name": "email-warning",
"type": "email",
"settings": {
"addresses": "platform-warnings@example.com",
"singleEmail": true
}
}
SMTP must already be configured on the Grafana instance (GF_SMTP_* in the Compose stack). Contact points do not replace SMTP config — they only choose who gets the mail.
3. Notification policy (routing)
This is where severity becomes "who gets woken up."
grafana/notification-policies/root.json
{
"receiver": "email-warning",
"group_by": ["alertname", "grafana_folder"],
"group_wait": "30s",
"group_interval": "5m",
"repeat_interval": "4h",
"routes": [
{
"receiver": "slack-oncall",
"object_matchers": [["severity", "=", "critical"]],
"continue": false,
"group_wait": "10s",
"repeat_interval": "1h"
},
{
"receiver": "email-warning",
"object_matchers": [["severity", "=", "warning"]],
"continue": false
}
]
}
Rules I stick to:
-
critical→ chat / pager that a human watches. -
warning→ email or a low-noise channel. - Do not send everything to the same Slack channel. That is how on-call learns to ignore you.
-
group_byonalertname(and folder) so a flapping port does not create 40 threads.
4. Alert rules (the part that pages)
Grafana Unified Alerting wants rule payloads with a stable uid. Below is the availability set from the monitoring stack, shaped for Grafana provisioning. Field names can differ slightly by Grafana version — export one rule from the UI once and mirror that shape if something 400s.
grafana/alert-rules/availability.json
{
"uid": "grp-availability",
"title": "availability",
"folderUID": "${GRAFANA_FOLDER_UID}",
"interval": "1m",
"rules": [
{
"uid": "rule-service-port-down",
"title": "ServicePortDown",
"condition": "C",
"data": [
{
"refId": "A",
"relativeTimeRange": { "from": 600, "to": 0 },
"datasourceUid": "prom-main",
"model": {
"expr": "probe_success == 0",
"refId": "A",
"instant": true
}
},
{
"refId": "C",
"datasourceUid": "__expr__",
"model": {
"type": "threshold",
"expression": "A",
"conditions": [
{
"evaluator": { "type": "gt", "params": [0] },
"operator": { "type": "and" },
"reducer": { "type": "last" }
}
],
"refId": "C"
}
}
],
"noDataState": "OK",
"execErrState": "Error",
"for": "2m",
"annotations": {
"summary": "Port check failing on {{ $labels.instance }}",
"description": "TCP connect failed for 2 minutes."
},
"labels": { "severity": "critical" }
},
{
"uid": "rule-mongo-lag",
"title": "MongoReplicationLagHigh",
"condition": "C",
"data": [
{
"refId": "A",
"relativeTimeRange": { "from": 600, "to": 0 },
"datasourceUid": "prom-main",
"model": {
"expr": "(scalar(max(mongodb_rs_members_optimeDate{member_state=\"PRIMARY\"})) - mongodb_rs_members_optimeDate{member_state=\"SECONDARY\"}) / 1000",
"refId": "A",
"instant": true
}
},
{
"refId": "C",
"datasourceUid": "__expr__",
"model": {
"type": "threshold",
"expression": "A",
"conditions": [
{
"evaluator": { "type": "gt", "params": [30] },
"operator": { "type": "and" },
"reducer": { "type": "last" }
}
],
"refId": "C"
}
}
],
"noDataState": "OK",
"execErrState": "Error",
"for": "5m",
"annotations": {
"summary": "Mongo replication lag high"
},
"labels": { "severity": "warning" }
}
]
}
Business SLI rule (uses the recording rule tenant:success_ratio:pct from the monitoring repo):
grafana/alert-rules/business-sli.json
{
"uid": "grp-business-sli",
"title": "business-sli",
"folderUID": "${GRAFANA_FOLDER_UID}",
"interval": "1m",
"rules": [
{
"uid": "rule-tenant-success-drop",
"title": "TenantSuccessRateDrop",
"condition": "C",
"data": [
{
"refId": "A",
"relativeTimeRange": { "from": 600, "to": 0 },
"datasourceUid": "prom-main",
"model": {
"expr": "tenant:success_ratio:pct < 90 and app_total_calls > 50",
"refId": "A",
"instant": true
}
},
{
"refId": "C",
"datasourceUid": "__expr__",
"model": {
"type": "threshold",
"expression": "A",
"conditions": [
{
"evaluator": { "type": "gt", "params": [0] },
"operator": { "type": "and" },
"reducer": { "type": "last" }
}
],
"refId": "C"
}
}
],
"for": "10m",
"noDataState": "OK",
"labels": { "severity": "critical" },
"annotations": {
"summary": "Success rate for {{ $labels.tenantId }} dropped below 90%"
}
}
]
}
for: is the noise filter. Port down: 1–2 minutes. Capacity: longer. The app_total_calls > 50 clause stops a 1-of-2 failure at 3 AM from paging when traffic is near zero — same idea as in the monitoring article.
Figure: Prometheus signal → Grafana rule (for:) → notification policy → on-call.
5. The apply script (diff, dry-run, apply)
This is the heart of the article. Read each file, substitute secrets, GET live object by uid, print + create / ~ update / = same, and only write when not --dry-run. Full script is in the companion repo; the important bits:
export GRAFANA_URL=https://grafana.example.com
export GRAFANA_TOKEN=glsa_xxx
export GRAFANA_FOLDER_UID=eeeeeeeee
export SLACK_WEBHOOK_URL=https://hooks.slack.com/services/xxx
export PROMETHEUS_PASSWORD=changeme
# PR / review
python3 scripts/apply_grafana.py --dry-run
# Before prod apply — keep a snapshot for rollback
python3 scripts/apply_grafana.py --export-snapshot snapshot-before.json --dry-run
# Apply
python3 scripts/apply_grafana.py
What the script does:
- Optional
--export-snapshotdumps live contact points, policies, and rules to a JSON file (Jenkins archives it). - Loads every
grafana/**/*.json, runs envsubst-style${VAR}replacement. - For each contact point and alert rule: match by
uid, print create/update/same. - Puts the notification policy tree (single root document).
- With
--dry-run, never calls POST/PUT.
Export one real rule from your Grafana UI the first time. Grafana’s schema is picky across versions; matching an export beats guessing field names.
6. Jenkins: dry-run on PR, apply on main
Same spirit as the monitoring deploy job: validate before you write, keep a snapshot for rollback.
pipeline {
agent { label 'cicd' }
parameters {
choice(name: 'whereTo', choices: ['uat', 'prod'], description: 'Grafana env')
booleanParam(name: 'dryRun', defaultValue: true)
}
environment {
GRAFANA_URL = credentials("grafana-url-${params.whereTo}")
GRAFANA_TOKEN = credentials("grafana-token-${params.whereTo}")
}
stages {
stage('checkout') {
steps { checkout scm }
}
stage('validate') {
steps {
sh 'find grafana -name "*.json" -print0 | xargs -0 -n1 python3 -m json.tool > /dev/null'
}
}
stage('snapshot') {
when { expression { !params.dryRun } }
steps {
sh 'python3 scripts/apply_grafana.py --export-snapshot snapshot-before.json --dry-run'
archiveArtifacts artifacts: 'snapshot-before.json', fingerprint: true
}
}
stage('apply') {
steps {
script {
def flag = params.dryRun ? '--dry-run' : ''
sh "python3 scripts/apply_grafana.py ${flag}"
}
}
}
}
post {
failure {
echo 'Apply failed — restore from the archived snapshot or previous Git tag'
}
}
}
Workflow I use:
- Open a PR → Jenkins runs with
dryRun=true→ review the+/~/=lines in the log. - Merge to main → apply to UAT with
dryRun=false. - Same commit to prod after UAT has been quiet for a day.
- If apply corrupts routing, re-apply from
snapshot-before.jsonor the previous Git tag.
Do not auto-apply alert changes to prod from every commit on day one. Dry-run forever on PR catches most mistakes.
7. Noise, ownership, and tests
Alerting as code does not fix a bad rule. It only makes the bad rule reproducible.
- Every critical rule needs an owner label (
team=platform) and a runbook link inannotations. -
noDataState: OKfor probes that disappear during deploys; useAlertingonly when missing data is itself an incident. - After apply, force a test: stop a blackbox target or fire a temporary rule with
for: 0s, confirm Slack, then delete/re-apply. An untested contact point is a dashboard with no pager. - Silences stay in Grafana UI for short incidents. Do not encode week-long silences in Git — that is how you ship deaf alerting.
Map rules back to the HA article’s failure modes:
| Failure | Alert |
|---|---|
| Instance / task |
up == 0 or target health |
| Bad deploy | error ratio + deploy annotation |
| AZ / port |
ServicePortDown (probe_success) |
| Data lag | MongoReplicationLagHigh |
| Business | TenantSuccessRateDrop |
8. Checklist before you call it done
- [ ] Service account token in Jenkins, not in Git
- [ ] Stable
uidon every contact point and rule - [ ] Notification policy routes
criticalandwarningdifferently - [ ] Dry-run on every PR
- [ ] Snapshot artifact before prod apply
- [ ] At least one forced test page after first apply
- [ ] Prometheus datasource auth works (
manageAlerts: true) - [ ] Folder UID exists in Grafana before rule apply
Wrapping up
Dashboards as files got us part of the way. Alerts are what wake humans, so they belong in Git with the same care as a deploy: validate, diff, apply, keep a snapshot.
The monitoring Compose stack still scrapes and stores. This pipeline decides who hears about it — and makes sure Friday’s UI click cannot erase Monday’s page.
Companion files live under alerting-as-code/repo (same layout as above). Wire it to your Grafana URL and token, run --dry-run first, then apply to UAT.
If you want the next one after this, it is CI/CD with safe rollback for the app fleets themselves — the same backup → release → restore idea, outside Grafana.
For any issue or support: LinkedIn.
About me
I'm Sanket Satish Patharkar, a Senior Cloud Operations Engineer. Most of my work is AWS, DevOps, CI/CD, and keeping production systems observable without training the team to ignore the pager.
Top comments (0)