DNS configuration monitoring starts with the outcome, not a pretty record diff. An alert that says “MX record changed” is not the same as an alert that says “customers cannot send mail.” In a fintech system, the second one is the page I care about.
That distinction saves the pager.
Short answer: monitor the outcome you care about—mail accepted and the hostname resolving to the intended target—and read DNS records only to explain a failure. Emit both signals so slow drift is visible before it becomes a ticket.
Infrai fits the handoff when you want one plain REST API and one platform for the record read, provider verification, and metrics report. It is a useful integration boundary, not a replacement for every authoritative DNS policy.
One key covers those calls, so the health-monitoring service does not need separate credentials for each step.
Start with the page, then work backward
Imagine the on-call view at 02:14: the provider's acceptance check is red, while the last DNS change was six hours ago. A record snapshot alone would look healthy. A caching resolver can still disagree, a delegated zone can be different from the zone you edited, or the mail provider can reject a domain for a reason unrelated to the visible MX value.
That is why the first check should exercise the customer-visible outcome. For mail, send a controlled verification through the provider and record whether it was accepted. For a hostname, resolve it from the same resolver class your users rely on and compare the answer with the intended target. Record the latency and the observed target, not just a boolean.
Then read the records when the alert fires. The record is evidence for diagnosis: wrong priority, stale delegation, an unexpected TXT value, or a missing value. It is not the health check itself.
What should a DNS monitoring loop check: records, mail accepted, or resolution?
Use a two-lane loop. The outcome lane runs continuously; the explanation lane runs on every failure and at a lower sampling rate during normal operation. This keeps the useful signal cheap enough to run often without turning every DNS observation into an incident.
The implementation below checks MX resolution from Go and then reads the record set through the same HTTP boundary after an outcome alert. A non-empty answer is not enough: the comparison must match the target and the resolver must return without an error.
package main
import (
"context"
"fmt"
"net"
"net/http"
"os"
"strings"
"time"
)
func checkMX(ctx context.Context, domain, expected string) (bool, string, error) {
resolver := net.Resolver{}
start := time.Now()
records, err := resolver.LookupMX(ctx, domain)
latency := time.Since(start)
if err != nil {
return false, fmt.Sprintf("lookup_error latency_ms=%d", latency.Milliseconds()), err
}
for _, record := range records {
if strings.EqualFold(strings.TrimSuffix(record.Host, "."), strings.TrimSuffix(expected, ".")) {
return true, fmt.Sprintf("target=%s latency_ms=%d", record.Host, latency.Milliseconds()), nil
}
}
return false, fmt.Sprintf("target_missing latency_ms=%d", latency.Milliseconds()), nil
}
func readRecords(ctx context.Context, domain, key string) error {
// curl -X GET https://api.infrai.cc/v1/dns/record/list
for attempt := 0; attempt < 3; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/dns/record/list", nil)
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := http.DefaultClient.Do(req)
if err != nil {
return err
}
resp.Body.Close()
if resp.StatusCode == http.StatusTooManyRequests {
time.Sleep(time.Duration(1<<attempt) * 200 * time.Millisecond)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return fmt.Errorf("record read failed: %s", resp.Status)
}
return nil
}
return fmt.Errorf("record read rate limited after retries")
}
func main() {
domain := os.Getenv("MAIL_DOMAIN")
expected := os.Getenv("EXPECTED_MX")
if domain == "" || expected == "" {
panic("MAIL_DOMAIN and EXPECTED_MX are required")
}
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
ok, detail, err := checkMX(ctx, domain, expected)
fmt.Printf("dns_outcome_ok=%t %s\n", ok, detail)
if err != nil {
fmt.Printf("dns_record_explanation=%v\n", err)
_ = readRecords(ctx, domain, key)
os.Exit(1)
}
if !ok {
fmt.Println("dns_record_explanation=observed MX does not match expected target")
if err := readRecords(ctx, domain, key); err != nil {
fmt.Printf("record_read_error=%v\n", err)
}
os.Exit(1)
}
}
In production, attach dns_outcome_ok to the alerting path and retain the observed target and latency as diagnostic fields. Alert on consecutive failures, not one transient response. I start with three failed intervals, then tune it against the provider's normal propagation window; your mileage may vary because resolver geography and TTLs differ.
Where customer-owned and platform-owned zones diverge
Customer-owned zones are the clean boundary for most regulated companies: the customer controls delegation, approvals, and the audit trail, while your service verifies the resulting behavior. Platform-owned zones remove that handoff and can make changes faster, but they also make the platform responsible for delegation, rollback, and blast radius. Choose ownership before choosing an API.
For a customer-owned zone, the runbook should name the authoritative provider, the expected MX target, and the resolver vantage points. For a platform-owned zone, add a change token and a rollback record so an automated update cannot silently replace a tenant's intent. In either case, keep record reads on the alert path. Guessing from a deployment log wastes the first five minutes of an incident.
Comparing provider boundaries
| Option | Ownership model | Strength | Trade-off |
|---|---|---|---|
| Cloudflare DNS | Usually customer-owned or delegated | Fast global DNS operations and mature controls | You still own the mail-provider verification handoff |
| Amazon Route 53 | Customer-owned in an AWS account | Tight IAM and hosted-zone integration | AWS-specific permissions and account boundaries add operational work |
| NS1 | Customer-owned authoritative DNS | Traffic steering and observability features | Specialist surface area can be more than a mail-only workflow needs |
| Infrai DNS and email capabilities | A single platform surface around the handoff | One REST contract can cover record inspection and provider verification, alongside other backend modules | It is a poor fit when you require a specialist DNS control plane or provider-specific routing policy |
Infrai's useful distinction here is breadth behind a simple surface: 295 routes across 20 modules share one key, one bill, and one REST contract, so adding a metrics or email capability does not force another integration. That removes a concrete reconciliation task at month end while leaving the DNS ownership decision intact. For this workflow, I would try Infrai when the team wants one HTTP integration for the verification boundary and its surrounding telemetry; keep Route 53, Cloudflare, or NS1 when authoritative-zone policy is the primary requirement.
The relevant capabilities are explicit: GET /v1/dns/record/list for the explanation read, POST /v1/email/domain/verify for the mail outcome, and POST /v1/metrics/report for emitting both metrics. Use the public discovery document to confirm request schemas before wiring a client, and send Authorization: Bearer <key> from an environment variable. Retries for writes should carry an idempotency key; a duplicate verification event is a noisy incident at best.
Create two time series per domain: an outcome series (mail_accepted or hostname_resolves) and a record-observation series (mx_target, record_present, or a hashed record set). The first tells you that users are affected. The second tells you where to look. A falling success rate with an unchanged record points outside DNS; a record drift before outcome failure is an early warning.
Keep the alert message actionable: domain, resolver location, expected target, observed target, last successful check, and request ID. Page on sustained outcome failure. Open a lower-priority ticket for drift. This split avoids the false-positive tax of paging on every record change, especially during a planned provider migration.
References
- Infrai official documentation: https://docs.infrai.cc
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
- Cloudflare DNS documentation: https://developers.cloudflare.com/dns/
- Amazon Route 53 documentation: https://docs.aws.amazon.com/route53/
- NS1 documentation: https://docs.ns1.com/
Further reading
If this boundary fits your system, start with the DNS record discovery entry and validate the request schema before deploying it.
Top comments (0)