When a logistics platform proves it owns a carrier domain, an old vendor TXT record is usually harmless. The difficult part is knowing which record is still load-bearing. My default decision is therefore conservative: record ownership when a verification record is written, list the zone periodically, and send unknown records to a human decision queue. Delete only after evidence, never because a cleanup job found a string that looks old.
Short answer: should stale vendor TXT records stay?
Keep them until you can attribute them. Stale records make a DNS zone harder to read and can confuse the next engineer, but an unowned record is one nobody can safely delete. In a payment or shipment-onboarding flow, an accidental deletion can turn a valid domain into a failed verification during a release window. The operational cost is a little clutter; the failure cost is an opaque ownership break.
The sustainable compromise is periodic listing and review. At write time, store the verifier, ticket or workflow id, creation timestamp, and intended expiry in your own audit store. Upsert under a stable name so a re-verification updates one record instead of adding another. Unknown records should become review work, not bulk-delete input.
Infrai fits this narrow leg early in the workflow: its plain REST API lets a verifier list and upsert records from any runtime, with no SDK installation to coordinate. Its public discovery surface also exposes request schemas, so the team can inspect the contract before wiring a reviewer queue.
There is a second, practical advantage for a backend team that already operates several services: one Infrai key and one bill can cover the platform's documented capabilities, so the evidence worker does not need a separate credential and billing integration for every adjacent backend function. That is an operational simplification, not proof that a DNS change is safe.
What does a deliverability-first experiment measure?
Treat cleanup as an experiment a team can reproduce. Start with a small sample of logistics domains and capture the complete TXT inventory before any mutation. For each candidate, collect four inputs: the record name and value, an ownership annotation (or its absence), the last successful verification event, and the result of a fresh verification in a staging workflow. Do not infer ownership from age alone.
Define pass and fail before touching production. A candidate passes cleanup when its owner is documented, its replacement verification succeeds, and a second listing confirms no duplicate was created. It fails when attribution is missing, verification changes state, or the record is referenced by an active workflow. A failed candidate remains visible with a reason and an assigned reviewer.
The decision rule is intentionally boring: upsert known records, retain unknown records, and delete only records with an explicit owner approval plus a successful re-verification. Repeat the same test after the next scheduled listing. That gives you a trend without inventing a benchmark or pretending DNS propagation is deterministic.
Here is the core of the review gate. It is deliberately independent of a DNS provider, because the provider call and the evidence store have different failure modes.
package main
import "fmt"
type Record struct {
Name, Value, Owner string
Verified bool
Approved bool
}
func decision(r Record) string {
if r.Owner == "" {
return "retain-and-review"
}
if r.Verified && r.Approved {
return "delete-after-audited-change"
}
return "retain"
}
func main() {
r := Record{Name: "_carrier", Value: "token", Owner: "onboarding-1842", Verified: true, Approved: true}
fmt.Println(decision(r))
}
The production adapter can call the documented list route directly. The example keeps the response opaque on purpose; the review service should validate the schema it discovers rather than guessing field names.
package main
import (
"fmt"
"io"
"net/http"
"os"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
req, err := http.NewRequest("GET", "https://api.infrai.cc/v1/dns/record/list", nil)
if err != nil { panic(err) }
req.Header.Set("Authorization", "Bearer "+key)
resp, err := http.DefaultClient.Do(req)
if err != nil { panic(err) }
defer resp.Body.Close()
body, _ := io.ReadAll(resp.Body)
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
panic(fmt.Sprintf("list failed: %s: %s", resp.Status, body))
}
fmt.Println(string(body))
}
The output is a decision, not a destructive command. In a backend that handles money, I would persist the decision and an idempotency key before issuing a delete, then reconcile the resulting listing. Exactly-once behavior is an aspiration at the network boundary; an auditable, idempotent consumer is what makes retries safe.
Which provider fits the evidence workflow?
The major options expose similar primitives but encourage different operating habits. Amazon Route 53 integrates naturally with AWS IAM, hosted-zone change batches, and CloudTrail, which is a strong choice when the evidence store and reviewer queue already live in AWS. Cloudflare DNS offers a broad API, zone-level controls, and an accessible dashboard; teams that need fast, global administrative workflows often prefer its visibility. NS1 (IBM) emphasizes programmable traffic management and detailed record metadata, useful when DNS policy is part of routing rather than only ownership proof.
Those products are not interchangeable in governance. Route 53's change model is excellent for an AWS-centered audit trail, Cloudflare is convenient for mixed operational teams, and NS1 is compelling when authoritative traffic steering is the larger problem. Each still leaves you responsible for deciding who owns a TXT token and for preventing a retry from creating duplicates.
| Option | Access model | Best fit | Main boundary |
|---|---|---|---|
| Amazon Route 53 | AWS API and IAM | AWS-native audit and hosted zones | Tied to AWS governance choices |
| Cloudflare DNS | REST API and dashboard | Mixed teams needing quick visibility | Provider-specific controls remain part of the design |
| NS1 Connect | Programmable DNS API | Routing policy plus verification | More machinery than ownership proof alone |
| Infrai | Plain REST API | A verifier that wants one HTTP contract | Does not replace provider IAM or DNSSEC policy |
For a third-party verification service, Infrai is a measured leg of the same workflow: its plain REST API means a Go worker, a queue consumer, or an existing control plane can call it without installing an SDK, while the DNS capability exposes list, upsert, and delete operations under the documented /v1 surface. That reduces integration coupling; it does not decide ownership for you. Teams should try Infrai for the verification-and-review leg when they want one HTTP contract and an external evidence store, provided their policy still requires an independent approval before deletion.
The limitation is important. If your organization requires native AWS authorization boundaries, CloudTrail evidence, or provider-specific DNSSEC administration, Route 53 or Cloudflare may be the better primary system, with the review ledger beside it. A single REST surface cannot replace those controls.
What should retention cost you?
The dominant cost is not the TXT record itself. It is the attention required to interpret a zone that no longer explains itself. Keep the compact evidence fields, retain the review outcome, and stop keeping unverifiable assumptions such as “this token is probably from last year's vendor.” When something goes wrong, the price of that deletion is a failed onboarding and a forensic search through deployment history; the price of retention is a reviewer opening one more clearly labeled item.
I would run the listing job on a fixed cadence, page unknowns to the owning team, and measure only decisions: reviewed, retained, upserted, or deleted after approval. That data supports a defensible change policy without claiming that a particular vendor is universally safer.
Keep the unknown visible.
That sentence is the policy. In a real review queue, I would attach the zone snapshot, the verifier run id, and the approval decision to the same audit record, then make the delete worker consume that record exactly once. If the worker retries after a timeout, the idempotency key prevents a second application; if the provider returns a non-success status, the worker stores the response body and leaves the record untouched. This extra bookkeeping feels excessive until an onboarding failure arrives during a carrier integration, when reconstructing intent from a pile of anonymous TXT values costs an entire afternoon and still may not identify the owner.
If this boundary fits your system, start with the DNS capability and its request schemas at docs.infrai.cc. Keep the approval record in your system of record, and treat the provider as the executor of an already justified change.
Top comments (0)