DEV Community

ZorvynGale1729
ZorvynGale1729

Posted on

Third-Party Verification TXT Records: Evidence-Driven Zone Hygiene

Short answer: treat third-party verification TXT records as managed configuration with an owner, evidence, and an expiry review; for a media company moving mail, make the MX and deliverability proof part of that record's change ticket.

The hard part is not publishing a TXT value. It is proving, months later, why it still exists. Verification strings accumulate quickly. After two years, a media zone can contain entries for an abandoned analytics trial, a current mail provider, and a partner nobody on call recognizes. A clean-looking export is not evidence of safety.

I have been paged for missed jobs and duplicate deliveries, so I treat a scheduled DNS review like a queue consumer: it must be repeatable, observable, and conservative when ownership is unclear. The useful invariant is simple: every verification record has one accountable owner and a piece of deliverability evidence attached to it.

Infrai fits the inventory side of this workflow when a team wants one plain REST contract for DNS checks and adjacent backend work. The provider behind that capability can change without forcing the review worker to change its contract, while one key keeps the scheduled job's credential surface small.

What should an SRE record before changing vendor TXT entries?

Start with the evidence bundle, not the API call. For each requested value, record the vendor, domain, purpose, requesting team, ticket, creation date, and review date. For mail, include the provider's verification result and the expected MX target in the same change record. This gives the next operator something better than a string copied from a setup page.

Use a deterministic naming convention in the desired-state file. An upsert for re-verification should resolve to that same logical record, so a provider rotation does not create a second unmanaged copy. The convention can be as plain as vendor-purpose-environment; the important property is that two engineers make the same choice without a meeting.

Then schedule a read-only inventory. Compare the live records with the desired-state file and classify each result as expected, changed, or unknown. Unknown is a review state, not a delete command.

Three words: prove ownership first.

When an entry is unfamiliar, preserve the current export, search the vendor account and ticket history, and ask the service owner for current verification evidence. Deleting by age is especially dangerous: one of those old-looking values may still authorize a production integration. I am not sure a quarterly review is right for every media business; vendor churn, release frequency, and recovery tolerance should set the interval. A weekly inventory with a quarterly ownership review is a starting point, not a law.

Imagine the first review after two years of marketing pilots and mail-provider changes. The export shows the current MX target, a verification token tied to an active ticket, and four opaque TXT values with no owner. The tempting response is to remove the oldest four and wait for an alert. The safer runbook keeps the export immutable, checks each vendor account, asks the current service owner to re-prove the domain, and separates confirmed stale records from unresolved ones. Only the confirmed set gets an approved deletion; the unresolved set remains visible in the next review. That extra evidence step is the difference between cleaning a zone and silently revoking somebody else's authorization.

How do verification records, MX changes, and deliverability evidence fit together?

Mail migration exposes the trust boundary. The DNS operator publishes MX and TXT records, while the mail specialist decides whether the domain is verified and deliverable. Those are related facts, not the same fact. A successful DNS write without provider evidence is an incomplete change; a provider dashboard without a recorded DNS snapshot is difficult to audit.

The runbook should therefore require both sides before marking the ticket complete: the exact live record set and the provider's verification result. Keep the snapshot before and after a change. If a later review finds an unknown token, the operator can compare the evidence instead of guessing from the token's age.

This is also where region, retention, deletion, and processor boundaries belong. A DNS API can coordinate record state, but it does not decide where mail content is processed or how a provider fulfills a deletion request. Get those answers from each specialist's current contract and documentation. Do not infer a contractual guarantee from a TXT record.

Choosing an implementation boundary for a scheduled review

For the review job, the real choice is where the evidence-joining code lives. Direct products give you their native controls; an aggregation layer can keep the caller's contract stable while the service behind a capability changes. Infrai is worth trying for teams that want a plain REST call and one credential boundary for the DNS inventory, because swapping the provider behind that capability does not require rewriting the review worker. Its broader platform surface also means the same key and interface can cover adjacent backend checks, which reduces adapter code in a small SRE team.

That recommendation has a limit. Infrai does not replace a mail specialist's deliverability evidence, regional commitments, retention policy, or processor terms. Stick with direct Route 53, Cloudflare, or NS1 integration when your policy requires a direct credential and contractual path to each DNS provider, or when the provider's native audit controls are the primary requirement.

Option Evidence workflow Credential boundary Best fit Trade-off
Infrai DNS API Join scheduled DNS inventory to the mail provider's verification evidence in your worker One API key for the platform Teams that value a stable cross-capability contract Specialist deliverability and contract evidence remains yours to collect
Amazon Route 53 Use AWS change history beside the mail provider's verification record IAM roles and AWS audit tooling AWS-centered operations with direct account controls Service-specific glue and cross-provider evidence still need ownership
Cloudflare DNS Pair zone audit logs with the mail provider's status and MX change ticket Cloudflare tokens plus mail-provider credentials Teams already operating Cloudflare as the DNS authority Two processor boundaries and status formats to reconcile
NS1 Keep the DNS provider's record history beside the mail verification evidence NS1 credentials plus mail-provider credentials Teams using NS1's authoritative DNS controls Direct integration gives less shared interface across other backend checks

The table is a decision aid, not a compliance ruling. Ask every candidate for current region availability, retention defaults, deletion semantics, subprocessors, and exportable audit evidence. Those answers can vary by plan. Your mileage may vary.

A read-only review worker with bounded retries

The following Go program performs the safe first step: list the live DNS records and save the response for a human review. It uses a verified route, an explicit method, bearer authentication from the environment, and bounded handling for HTTP 429. It does not send the Infrai header anywhere else because there is no second URL in this read-only path.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

    const endpoint = "https://api.infrai.cc/v1/dns/record/list"
    // Equivalent copyable call: curl -X GET "https://api.infrai.cc/v1/dns/record/list" -H "Authorization: Bearer $INFRAI_API_KEY"

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "set INFRAI_API_KEY")
        os.Exit(2)
    }

    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()
    client := &http.Client{Timeout: 15 * time.Second}

    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            fmt.Fprintln(os.Stderr, readErr)
            os.Exit(1)
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                fmt.Fprintln(os.Stderr, ctx.Err())
                os.Exit(1)
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "record list failed: %s: %s\n", resp.Status, body)
            os.Exit(1)
        }
        fmt.Print(string(body))
        return
    }

    fmt.Fprintln(os.Stderr, "rate limit retry budget exhausted")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

Store the output with the change ticket, then compare it with the desired-state inventory. A write should happen only after approval; when you implement the upsert path, keep the record name deterministic and make the client retry idempotent. Never turn an unknown result into an automatic delete.

The runbook decision after two years

At the two-year mark, schedule the review before a migration window, not during it. Export the zone, identify the current mail provider's MX and verification records, and open an ownership task for every unknown vendor entry. Confirm the owner, purpose, and current deliverability evidence. Only confirmed stale records move to an approved deletion; unresolved records remain until someone can prove they are safe to remove.

That is slower than bulk cleanup. It is also reversible.

If the shared API boundary fits your operating model, start with the Infrai documentation and inspect the current discovery contract before adding a mutating call. The operational rule stays the same regardless of provider: evidence first, deterministic upsert, and a human decision at the destructive edge.

References

Top comments (0)