DEV Community

Sanket Patharkar
Sanket Patharkar

Posted on

Alerting as Code: Grafana rules, contact points, and Jenkins dry-runs

Someone edited a critical alert in the Grafana UI on a Friday afternoon. By Monday the rule was gone — overwritten by a dashboard sync, or a different env's export, or just a click nobody remembered. We still had Prometheus. We still had metrics. We did not get paged when a tenant success rate fell off a cliff.

That is when "alerting as code" stopped being a nice idea and became the only way I would run Grafana alerts.

This is the sequel to my Dockerized Monitoring Stack post. That article got Prometheus, Grafana, and exporters into one Compose file. This one covers what I promised next: provision contact points, notification policies, datasources, and alert rules through the Grafana HTTP API from Jenkins, with dry-run, idempotent apply, and a way back when something is wrong.

I assume you already have Grafana up (I use the Compose stack from that article) and a Prometheus datasource. If you are still designing the failure model, my HA architecture piece maps instance / AZ / Region failures to what you should alert on.

The JSON and scripts below are reference implementations. Point them at a non-prod Grafana first. Check tokens, folder UIDs, and API versions for your Grafana major version before you touch production.


What "alerting as code" means here

It does not mean dumping every dashboard JSON into Git and calling it done. Dashboards can stay file-provisioned (as in the monitoring repo). Alerts are different: they change often, they wake people up, and a silent UI edit is a production incident waiting to happen.

In this setup:

  • Git is the source of the intended state (contact points, notification policies, alert rules, datasource defs).
  • Jenkins (or any CI) validates, diffs against live Grafana, then applies.
  • Dry-run prints the diff and exits 0 without writing.
  • Rollback restores the last exported snapshot from the job artifacts.

Prometheus can still evaluate recording/alert rules from prometheus/rules/*.yml. I use both: Prometheus rules for anything that must work even if Grafana is down, and Grafana Alerting for routing, grouping, and human-facing policies. This article focuses on the Grafana side.

Figure: Git → Jenkins → Grafana API → Slack/email (dry-run on PR, apply on merge).


Repo layout

alerting-as-code/
├── grafana/
│   ├── datasources/
│   │   └── prometheus.json
│   ├── contact-points/
│   │   ├── slack-oncall.json
│   │   └── email-warning.json
│   ├── notification-policies/
│   │   └── root.json
│   └── alert-rules/
│       ├── availability.json
│       └── business-sli.json
├── scripts/
│   └── apply_grafana.py
└── Jenkinsfile
Enter fullscreen mode Exit fullscreen mode

One folder per Grafana object type. One file per object (or per rule group). Env-specific values (webhook URLs, folder UIDs) come from Jenkins credentials or a secrets manager — not committed.

I keep UAT and prod as the same files with different secret injection, same idea as env/uat vs env/prod in the monitoring stack. If the rule set truly diverges, split alert-rules/uat/ and alert-rules/prod/.


Grafana API pieces you actually need

Grafana Alerting is driven by the HTTP API under /api/v1/provisioning/... (and a few older /api/ routes for datasources). You need a service account token with Admin (or a custom role that can manage alerting). Store it in Jenkins as GRAFANA_TOKEN. Base URL looks like https://grafana.example.com.

Useful endpoints (Grafana 10/11 style — confirm against your docs):

Object List / get Create / update
Datasource GET /api/datasources POST /api/datasources or PUT /api/datasources/uid/:uid
Contact point GET /api/v1/provisioning/contact-points POST / PUT .../contact-points/:uid
Notification policy tree GET /api/v1/provisioning/policies PUT /api/v1/provisioning/policies
Alert rules GET /api/v1/provisioning/alert-rules POST / PUT .../alert-rules/:uid

Always set a stable uid in every JSON file. Without it, every apply creates duplicates and dry-run cannot match objects.


1. Datasource (once)

File-provisioned datasources are fine. If you want the pipeline to own them too:

grafana/datasources/prometheus.json

{
  "uid": "prom-main",
  "name": "Prometheus",
  "type": "prometheus",
  "access": "proxy",
  "url": "http://localhost:9090",
  "basicAuth": true,
  "basicAuthUser": "admin",
  "secureJsonData": {
    "basicAuthPassword": "${PROMETHEUS_PASSWORD}"
  },
  "jsonData": {
    "httpMethod": "POST",
    "manageAlerts": true
  },
  "isDefault": true
}
Enter fullscreen mode Exit fullscreen mode

${PROMETHEUS_PASSWORD} is substituted at apply time. Never commit the real password. The monitoring stack already put basic auth on Prometheus — Grafana must use the same credentials or every alert evaluation fails with 401 and you get no pages (or false "datasource error" alerts).


2. Contact points

Two contact points cover most teams I work with: critical → Slack (or PagerDuty), warning → email.

grafana/contact-points/slack-oncall.json

{
  "uid": "cp-slack-oncall",
  "name": "slack-oncall",
  "type": "slack",
  "settings": {
    "recipient": "#oncall-alerts",
    "title": "{{ .CommonLabels.alertname }}",
    "text": "{{ .CommonAnnotations.summary }}\n{{ .CommonAnnotations.description }}"
  },
  "secureSettings": {
    "url": "${SLACK_WEBHOOK_URL}"
  }
}
Enter fullscreen mode Exit fullscreen mode

grafana/contact-points/email-warning.json

{
  "uid": "cp-email-warning",
  "name": "email-warning",
  "type": "email",
  "settings": {
    "addresses": "platform-warnings@example.com",
    "singleEmail": true
  }
}
Enter fullscreen mode Exit fullscreen mode

SMTP must already be configured on the Grafana instance (GF_SMTP_* in the Compose stack). Contact points do not replace SMTP config — they only choose who gets the mail.


3. Notification policy (routing)

This is where severity becomes "who gets woken up."

grafana/notification-policies/root.json

{
  "receiver": "email-warning",
  "group_by": ["alertname", "grafana_folder"],
  "group_wait": "30s",
  "group_interval": "5m",
  "repeat_interval": "4h",
  "routes": [
    {
      "receiver": "slack-oncall",
      "object_matchers": [["severity", "=", "critical"]],
      "continue": false,
      "group_wait": "10s",
      "repeat_interval": "1h"
    },
    {
      "receiver": "email-warning",
      "object_matchers": [["severity", "=", "warning"]],
      "continue": false
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

Rules I stick to:

  • critical → chat / pager that a human watches.
  • warning → email or a low-noise channel.
  • Do not send everything to the same Slack channel. That is how on-call learns to ignore you.
  • group_by on alertname (and folder) so a flapping port does not create 40 threads.

4. Alert rules (the part that pages)

Grafana Unified Alerting wants rule payloads with a stable uid. Below is the availability set from the monitoring stack, shaped for Grafana provisioning. Field names can differ slightly by Grafana version — export one rule from the UI once and mirror that shape if something 400s.

grafana/alert-rules/availability.json

{
  "uid": "grp-availability",
  "title": "availability",
  "folderUID": "${GRAFANA_FOLDER_UID}",
  "interval": "1m",
  "rules": [
    {
      "uid": "rule-service-port-down",
      "title": "ServicePortDown",
      "condition": "C",
      "data": [
        {
          "refId": "A",
          "relativeTimeRange": { "from": 600, "to": 0 },
          "datasourceUid": "prom-main",
          "model": {
            "expr": "probe_success == 0",
            "refId": "A",
            "instant": true
          }
        },
        {
          "refId": "C",
          "datasourceUid": "__expr__",
          "model": {
            "type": "threshold",
            "expression": "A",
            "conditions": [
              {
                "evaluator": { "type": "gt", "params": [0] },
                "operator": { "type": "and" },
                "reducer": { "type": "last" }
              }
            ],
            "refId": "C"
          }
        }
      ],
      "noDataState": "OK",
      "execErrState": "Error",
      "for": "2m",
      "annotations": {
        "summary": "Port check failing on {{ $labels.instance }}",
        "description": "TCP connect failed for 2 minutes."
      },
      "labels": { "severity": "critical" }
    },
    {
      "uid": "rule-mongo-lag",
      "title": "MongoReplicationLagHigh",
      "condition": "C",
      "data": [
        {
          "refId": "A",
          "relativeTimeRange": { "from": 600, "to": 0 },
          "datasourceUid": "prom-main",
          "model": {
            "expr": "(scalar(max(mongodb_rs_members_optimeDate{member_state=\"PRIMARY\"})) - mongodb_rs_members_optimeDate{member_state=\"SECONDARY\"}) / 1000",
            "refId": "A",
            "instant": true
          }
        },
        {
          "refId": "C",
          "datasourceUid": "__expr__",
          "model": {
            "type": "threshold",
            "expression": "A",
            "conditions": [
              {
                "evaluator": { "type": "gt", "params": [30] },
                "operator": { "type": "and" },
                "reducer": { "type": "last" }
              }
            ],
            "refId": "C"
          }
        }
      ],
      "noDataState": "OK",
      "execErrState": "Error",
      "for": "5m",
      "annotations": {
        "summary": "Mongo replication lag high"
      },
      "labels": { "severity": "warning" }
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

Business SLI rule (uses the recording rule tenant:success_ratio:pct from the monitoring repo):

grafana/alert-rules/business-sli.json

{
  "uid": "grp-business-sli",
  "title": "business-sli",
  "folderUID": "${GRAFANA_FOLDER_UID}",
  "interval": "1m",
  "rules": [
    {
      "uid": "rule-tenant-success-drop",
      "title": "TenantSuccessRateDrop",
      "condition": "C",
      "data": [
        {
          "refId": "A",
          "relativeTimeRange": { "from": 600, "to": 0 },
          "datasourceUid": "prom-main",
          "model": {
            "expr": "tenant:success_ratio:pct < 90 and app_total_calls > 50",
            "refId": "A",
            "instant": true
          }
        },
        {
          "refId": "C",
          "datasourceUid": "__expr__",
          "model": {
            "type": "threshold",
            "expression": "A",
            "conditions": [
              {
                "evaluator": { "type": "gt", "params": [0] },
                "operator": { "type": "and" },
                "reducer": { "type": "last" }
              }
            ],
            "refId": "C"
          }
        }
      ],
      "for": "10m",
      "noDataState": "OK",
      "labels": { "severity": "critical" },
      "annotations": {
        "summary": "Success rate for {{ $labels.tenantId }} dropped below 90%"
      }
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

for: is the noise filter. Port down: 1–2 minutes. Capacity: longer. The app_total_calls > 50 clause stops a 1-of-2 failure at 3 AM from paging when traffic is near zero — same idea as in the monitoring article.

Figure: Prometheus signal → Grafana rule (for:) → notification policy → on-call.


5. The apply script (diff, dry-run, apply)

This is the heart of the article. Read each file, substitute secrets, GET live object by uid, print + create / ~ update / = same, and only write when not --dry-run. Full script is in the companion repo; the important bits:

export GRAFANA_URL=https://grafana.example.com
export GRAFANA_TOKEN=glsa_xxx
export GRAFANA_FOLDER_UID=eeeeeeeee
export SLACK_WEBHOOK_URL=https://hooks.slack.com/services/xxx
export PROMETHEUS_PASSWORD=changeme

# PR / review
python3 scripts/apply_grafana.py --dry-run

# Before prod apply — keep a snapshot for rollback
python3 scripts/apply_grafana.py --export-snapshot snapshot-before.json --dry-run

# Apply
python3 scripts/apply_grafana.py
Enter fullscreen mode Exit fullscreen mode

What the script does:

  1. Optional --export-snapshot dumps live contact points, policies, and rules to a JSON file (Jenkins archives it).
  2. Loads every grafana/**/*.json, runs envsubst-style ${VAR} replacement.
  3. For each contact point and alert rule: match by uid, print create/update/same.
  4. Puts the notification policy tree (single root document).
  5. With --dry-run, never calls POST/PUT.

Export one real rule from your Grafana UI the first time. Grafana’s schema is picky across versions; matching an export beats guessing field names.


6. Jenkins: dry-run on PR, apply on main

Same spirit as the monitoring deploy job: validate before you write, keep a snapshot for rollback.

pipeline {
  agent { label 'cicd' }
  parameters {
    choice(name: 'whereTo', choices: ['uat', 'prod'], description: 'Grafana env')
    booleanParam(name: 'dryRun', defaultValue: true)
  }
  environment {
    GRAFANA_URL   = credentials("grafana-url-${params.whereTo}")
    GRAFANA_TOKEN = credentials("grafana-token-${params.whereTo}")
  }
  stages {
    stage('checkout') {
      steps { checkout scm }
    }
    stage('validate') {
      steps {
        sh 'find grafana -name "*.json" -print0 | xargs -0 -n1 python3 -m json.tool > /dev/null'
      }
    }
    stage('snapshot') {
      when { expression { !params.dryRun } }
      steps {
        sh 'python3 scripts/apply_grafana.py --export-snapshot snapshot-before.json --dry-run'
        archiveArtifacts artifacts: 'snapshot-before.json', fingerprint: true
      }
    }
    stage('apply') {
      steps {
        script {
          def flag = params.dryRun ? '--dry-run' : ''
          sh "python3 scripts/apply_grafana.py ${flag}"
        }
      }
    }
  }
  post {
    failure {
      echo 'Apply failed — restore from the archived snapshot or previous Git tag'
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Workflow I use:

  1. Open a PR → Jenkins runs with dryRun=true → review the + / ~ / = lines in the log.
  2. Merge to main → apply to UAT with dryRun=false.
  3. Same commit to prod after UAT has been quiet for a day.
  4. If apply corrupts routing, re-apply from snapshot-before.json or the previous Git tag.

Do not auto-apply alert changes to prod from every commit on day one. Dry-run forever on PR catches most mistakes.


7. Noise, ownership, and tests

Alerting as code does not fix a bad rule. It only makes the bad rule reproducible.

  • Every critical rule needs an owner label (team=platform) and a runbook link in annotations.
  • noDataState: OK for probes that disappear during deploys; use Alerting only when missing data is itself an incident.
  • After apply, force a test: stop a blackbox target or fire a temporary rule with for: 0s, confirm Slack, then delete/re-apply. An untested contact point is a dashboard with no pager.
  • Silences stay in Grafana UI for short incidents. Do not encode week-long silences in Git — that is how you ship deaf alerting.

Map rules back to the HA article’s failure modes:

Failure Alert
Instance / task up == 0 or target health
Bad deploy error ratio + deploy annotation
AZ / port ServicePortDown (probe_success)
Data lag MongoReplicationLagHigh
Business TenantSuccessRateDrop

8. Checklist before you call it done

  • [ ] Service account token in Jenkins, not in Git
  • [ ] Stable uid on every contact point and rule
  • [ ] Notification policy routes critical and warning differently
  • [ ] Dry-run on every PR
  • [ ] Snapshot artifact before prod apply
  • [ ] At least one forced test page after first apply
  • [ ] Prometheus datasource auth works (manageAlerts: true)
  • [ ] Folder UID exists in Grafana before rule apply

Wrapping up

Dashboards as files got us part of the way. Alerts are what wake humans, so they belong in Git with the same care as a deploy: validate, diff, apply, keep a snapshot.

The monitoring Compose stack still scrapes and stores. This pipeline decides who hears about it — and makes sure Friday’s UI click cannot erase Monday’s page.

Companion files live under alerting-as-code/repo (same layout as above). Wire it to your Grafana URL and token, run --dry-run first, then apply to UAT.

If you want the next one after this, it is CI/CD with safe rollback for the app fleets themselves — the same backup → release → restore idea, outside Grafana.

For any issue or support: LinkedIn.


About me

I'm Sanket Satish Patharkar, a Senior Cloud Operations Engineer. Most of my work is AWS, DevOps, CI/CD, and keeping production systems observable without training the team to ignore the pager.

Top comments (0)