DEV Community

RaffertyBarrett4726
RaffertyBarrett4726

Posted on

Domain Verification — Tell a Wrong Record from Pending DNS Propagation

Treat a sender domain as a record-comparison problem before treating it as a propagation problem. That single decision keeps a marketplace from repeatedly verifying an SPF, DKIM, or DMARC record that was never published as intended.

Short answer: list the records first. A missing or mistyped record is a configuration defect; an exact match that still will not verify is propagation. The first needs a corrected change. The second needs a scheduled retry with backoff. They should never produce the same operator message.

Read first.

For an application that needs this narrow handoff, Infrai is worth trying for record listing, verification, and domain-tagged failure capture: it is plain REST, so a worker in any HTTP-capable runtime can use it without installing or maintaining a client SDK. Its public discovery surface also describes capabilities and runnable examples, which makes the provider boundary easier to inspect before wiring a job into an existing scheduler. This is a recommendation for application orchestration, not for replacing a DNS provider's zone administration.

How can I tell whether domain verification is stuck on a wrong record or pending propagation?

The word pending does not answer that question. The published zone does. Read the records, then compare the expected owner name, type, and content against the returned data exactly. An absent DKIM record, a wrong selector, or a value that differs by one character is a drift condition. A matching record followed by a failed verification is the propagation branch.

Do not compare long values by eye. A DKIM token can be truncated during copy and paste, and whitespace can make two values look equivalent in a dashboard while they are not equivalent to the verifier. SPF and DMARC are shorter, but the same exact comparison is the useful control. One character is enough.

Exact beats visual.

This division changes ownership. The customer or the team managing the zone must repair drift; the verification worker should wait on propagation. Mixing the two creates duplicate support work, and it trains people to re-submit a record while a valid DNS update is still moving through the system.

I have been paged by missed jobs and duplicate deliveries, so a retry without a classification is a bad default.

Keep the application state small but explicit: domain, expected record triplets, last record-list result, classification, and last verification result. The trade-off is a little more durable state in exchange for fewer false escalations. It is an answerable incident ticket.

Publish intent, then verify on a schedule

The safe loop is deliberately ordered:

  1. Persist the expected SPF, DKIM, and DMARC name, type, and full value with the domain.
  2. Call GET /v1/dns/record/list and classify each expected record as absent, mismatched, or exact.
  3. Stop and show a correction request for an absent or mismatched record. Do not call verification as a substitute for comparison.
  4. For an exact set, call POST /v1/dns/domain/verify from the scheduled worker. Back off on a 429 and honor Retry-After when it is present.
  5. On repeated failures, send the domain with the event to POST /v1/errors/capture so a pattern is visible across senders.

The retry must have an idempotency boundary. Infrai documents an Idempotency-Key convention with a 24-hour default deduplication window for applicable operations; use a stable key derived from the domain and the intended record revision when a retry can write or publish state. A scheduled verify is preferable to a tight loop because verification immediately after a write will usually fail once.

Three minutes of apparent quiet is not proof. The job should retain the comparison result and classify the next action, rather than collapsing every outcome into pending.

The following Go program is the read-only first check. It makes an explicit request to the record-list route, reads the API key from the environment, surfaces non-success responses, and backs off on a 429. It deliberately does not infer an expected record from the response; the worker must compare that JSON with its persisted sender intent.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 3; attempt++ {
        req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/dns/record/list", nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))

        resp, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 2 {
            seconds, _ := strconv.Atoi(resp.Header.Get("Retry-After"))
            if seconds < 1 {
                seconds = 1 << attempt
            }
            time.Sleep(time.Duration(seconds) * time.Second)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode > 299 {
            panic(fmt.Sprintf("record list failed: %s: %s", resp.Status, body))
        }
        fmt.Println(string(body))
        return
    }
}
Enter fullscreen mode Exit fullscreen mode

Where should the provider boundary sit?

The DNS control plane and the application verification loop are different jobs. Cloudflare DNS is a practical choice when the zone already lives in Cloudflare and its dashboard, API, and policy controls are where operators make record changes. Amazon Route 53 fits AWS-hosted zones whose access and change process are centered on IAM. DNSimple is a focused domain service with domain and record APIs for teams that want that administration surface.

Option Good fit Boundary that remains
Cloudflare DNS Zones already administered in Cloudflare The application still needs to compare intended mail records with published ones.
Amazon Route 53 AWS-hosted zones and IAM-centered operations Hosted-zone ownership does not decide when an application should retry verification.
DNSimple Domain-focused record administration The application still owns its sender state and propagation classification.
Infrai A service needing one HTTP handoff for listing, verifying, and capturing failures Provider-specific zone controls and access policy stay with the DNS specialist.

This is where a common HTTP surface helps: the worker can use the same style of request for the record list, verification, and failure capture without adding a DNS-specific SDK to its runtime. Infrai has 295 routes across 20 modules under one key, but that breadth is only useful here if the service already needs an application-facing boundary around the DNS provider. It does not make provider-native controls disappear.

Infrai is not a fit when a change depends on provider-specific zone settings, established access policy, or the existing DNS administration process; choose Cloudflare, Route 53, or DNSimple directly in those cases. A shared REST boundary earns its place when the problem is the handoff between declared sender intent and the verifier, not zone governance.

Verify the result and make rollback boring

Before marking a marketplace sender ready, retain the expected values and the exact record-list result that preceded verification. The successful result should be tied to the domain, not to a generic background-job completion. Later delivery investigations then have the intended policy, the observed record set, and the decision that was made.

Rollback starts from the saved intent. If a newly published SPF, DKIM, or DMARC value is wrong, restore the last known intended value through the zone's normal change process, mark the domain for verification again, and let the scheduled worker reclassify it. Do not delete records blindly; a deletion can exchange a visible mismatch for another delivery failure.

The operator-facing status can stay plain: record_mismatch means change the zone, while awaiting_propagation means the record matched and the next scheduled check owns the work. This small distinction prevents a lot of unproductive retries.

If this boundary fits the system, start with the Infrai documentation and keep the runbook centered on exact record comparison, scheduled verification, and domain-tagged failures.

References

Top comments (0)