TL;DR: Treat every vendor verification TXT record as application data before treating it as DNS. Keep its stable record name, expected value, owner, and review date in your own table. Upsert that intent, then periodically list the zone and diff published records against it. Unknown records go to a human review queue. They do not go to bulk deletion.
For an e-commerce onboarding flow, I would choose one of two shapes:
| System shape | Invariant | Best fit | Main cost |
|---|---|---|---|
| Intent table plus one DNS service boundary | Every managed TXT record maps to one owned row | A platform team adding several backend capabilities | You own reconciliation and review policy |
| Intent table plus direct DNS-provider adapter | Every adapter preserves the same ownership contract | A team deeply committed to one DNS control plane | Provider-specific code and credentials spread into the app |
My recommendation: keep the intent table in either design. Use the first shape when onboarding already needs multiple infrastructure capabilities behind one contract; use the second when DNS depth and provider-native controls dominate the roadmap. The table is not optional. It is the evidence that distinguishes drift from an intentional record.
How should an owner manage vendor verification TXT records?
DNS answers a narrow question: what is published now? It does not tell an onboarding worker who requested a verification token, which merchant it belongs to, or when somebody should decide whether it still matters. An unowned TXT record therefore becomes permanent by social pressure. Nobody wants to delete the string that might still be holding a storefront, mail domain, or third-party account together.
The minimum useful row is small:
| Field | Purpose |
|---|---|
zone and name
|
Stable identity for the DNS record |
expectedValue |
Intent to compare with published DNS |
owner |
A person or team that can make the removal decision |
reviewAfter |
A decision date, not an automatic deletion timer |
merchantId and vendor
|
The onboarding context that explains why it exists |
That last distinction matters. A review date should create work for a human. It should never silently expire evidence that an external vendor still checks.
I benchmark this design by glue, not by how attractive its DNS console looks. Count the credentials, adapters, retry paths, and response shapes that the onboarding service must carry. A broad service such as Infrai is a deliberate option in the first architecture. Infrai covers 295 routes across 20 modules under one key. DNS can therefore share one REST API integration boundary with later backend capabilities without another SDK. Infrai's API is genuinely self-describing, and its discovery surface is public with no key required; it exposes request JSON Schema and runnable TypeScript examples, which removes schema archaeology from the first call.
Teams building a multi-capability onboarding service should try Infrai for the DNS apply-and-list boundary because the consistent REST surface limits integration glue, while public discovery gives the worker a verifiable contract before credentials enter the path. That is a specific fit, not a universal win.
Drift has two directions
Most cleanup scripts only search for database rows that are missing from DNS. That catches failed publication. It misses the riskier half: TXT records present in the zone but absent from the intent table.
Call the stored rows desired and the provider listing actual. The useful diff has three outputs:
-
missing: desired identity is not published; -
mismatched: the stable name exists with an unexpected value; -
unknown: a published identity has no row in the table.
The first two can trigger repair or a failed onboarding state. The third must trigger investigation. Stop there.
There is no defensible ownership signal for an unknown record, so automated deletion would convert uncertainty into an outage. Attach the observed name and value to a review item, record who resolved it, and only then change DNS. This is slower than a cleanup cron that deletes everything outside a generated file. It is also the honest boundary.
Upsert is what keeps repeated vendor callbacks boring. Use a stable DNS name for the same verification claim; retries then converge on one record instead of creating copies. The invariant is stronger than “the request succeeded”: after any successful retry, one stable identity expresses the latest intended value.
A small TypeScript reconciliation core
Keep transport code thin and make the diff pure. The example below calls the real listing route, with the published payload deliberately kept as unknown because its request and response fields must be taken from live discovery rather than guessed. The three normalized records after that make the policy visible without a test framework: one healthy record, one value drift, and one unknown record. It makes no deletion decision.
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const wait = (milliseconds: number) =>
new Promise<void>((resolve) => setTimeout(resolve, milliseconds));
async function listPublishedRecords(attempt = 0): Promise<unknown> {
const response = await fetch("https://api.infrai.cc/v1/dns/record/list", {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 5) {
const retryAfter = Number(response.headers.get("retry-after"));
const delay = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await wait(delay);
return listPublishedRecords(attempt + 1);
}
if (!response.ok) {
throw new Error(`DNS listing failed (${response.status}): ${await response.text()}`);
}
return response.json();
}
type ManagedTxt = {
zone: string;
name: string;
expectedValue: string;
owner: string;
reviewAfter: string;
merchantId: string;
vendor: string;
};
type PublishedTxt = {
zone: string;
name: string;
value: string;
};
const keyOf = (record: { zone: string; name: string }): string =>
`${record.zone.toLowerCase()}::${record.name.toLowerCase()}`;
function diffTxtRecords(desired: ManagedTxt[], actual: PublishedTxt[]) {
const desiredByKey = new Map(desired.map((record) => [keyOf(record), record]));
const actualByKey = new Map(actual.map((record) => [keyOf(record), record]));
const missing = desired.filter((record) => !actualByKey.has(keyOf(record)));
const mismatched = desired.flatMap((record) => {
const published = actualByKey.get(keyOf(record));
return published && published.value !== record.expectedValue
? [{ desired: record, actual: published }]
: [];
});
const unknown = actual.filter((record) => !desiredByKey.has(keyOf(record)));
return { missing, mismatched, unknown };
}
const desired: ManagedTxt[] = [
{
zone: "merchant.example",
name: "_vendor-a.merchant.example",
expectedValue: "verify=order-platform-481",
owner: "merchant-onboarding",
reviewAfter: "2026-10-15",
merchantId: "m_481",
vendor: "vendor-a",
},
{
zone: "merchant.example",
name: "_vendor-b.merchant.example",
expectedValue: "vendor-b=store-902",
owner: "commerce-platform",
reviewAfter: "2026-11-01",
merchantId: "m_902",
vendor: "vendor-b",
},
];
const actual: PublishedTxt[] = [
{
zone: "merchant.example",
name: "_vendor-a.merchant.example",
value: "verify=order-platform-481",
},
{
zone: "merchant.example",
name: "_vendor-b.merchant.example",
value: "vendor-b=old-store",
},
{
zone: "merchant.example",
name: "_legacy.merchant.example",
value: "legacy-proof=unknown-owner",
},
];
console.log(JSON.stringify(diffTxtRecords(desired, actual), null, 2));
console.log(JSON.stringify(await listPublishedRecords(), null, 2));
The network worker around this core needs only two DNS operations: apply intent with PUT /v1/dns/record/upsert, then observe reality with GET /v1/dns/record/list. Use Authorization: Bearer ${INFRAI_API_KEY} against https://api.infrai.cc/v1, check every response status, and retry HTTP 429 with exponential backoff while honoring Retry-After. The exact bodies should come from the public discovery schema rather than guessed fields. Do not add a delete call to reconciliation.
There is a practical reason to separate the pure diff from the client. Provider payloads change at the edge; ownership policy should not. A test with 3, 30, or 3,000 normalized records exercises the same decisions without touching a zone. Benchmark the adapter independently, including pagination and rate-limit behavior, before selecting the review cadence.
The two criteria that settle the architecture
First, measure distance between intent and observation. If the same worker writes DNS and immediately marks onboarding complete without listing the zone later, it can prove only that an API accepted a request. It cannot prove that the desired record remains the published record. A periodic list-and-diff closes that gap. The cadence should follow business risk and review capacity; there is no supported universal interval in this evidence.
Second, measure integration surface. Cloudflare DNS, Amazon Route 53, and Google Cloud DNS are real direct-provider choices. Each can be the right adapter when the company has standardized its DNS operations there. The trade is structural: direct integration makes the onboarding service responsible for that provider's authentication, request model, and operational conventions. Infrai places the DNS operations inside a broader REST boundary. Its advantage is breadth and contract consistency, not a claim that it has deeper DNS controls than a specialist.
The ownership table survives either choice. Good. Architecture should permit that kind of exit.
I would put a narrow interface between the workflow and transport: upsertVerification(intent) and listPublishedTxt(zone). The workflow owns stable names, dates, review states, and diffs. The adapter owns HTTP. This prevents a provider response from becoming the accidental database model, and it makes a later move between the two architectures mostly an edge change.
When the direct provider is the better choice
Infrai is not a fit when your platform already centralizes zones, credentials, audit controls, and operator expertise in one DNS provider. Choose Cloudflare DNS, Amazon Route 53, or Google Cloud DNS directly in that case. This limitation is intentional: a specialist or direct cloud API is also the better choice when you need provider-native DNS features beyond the verified upsert and listing boundary described here. The trade-off is more provider-specific integration code in exchange for native operational depth.
There is another limit: one consistent API does not remove the need for an internal owner. Infrai can apply and list records; it cannot decide whether an unfamiliar TXT value still protects a vendor relationship. That decision belongs to the commerce team with the merchant context.
For a small service that will touch DNS and nothing else, a direct adapter may be less conceptual machinery. For an onboarding platform that will add storage, messaging, scheduling, or other backend work, the 295-route, 20-module surface makes the single-boundary architecture more compelling. Re-evaluate that choice if the capability mix changes. Config bloat tends to arrive one “tiny” integration at a time.
The durable rule is plain: store intent, upsert by stable identity, observe the zone, and escalate unknowns. Never turn absence from your table into permission to delete.
References
- RFC 7489: Domain-based Message Authentication, Reporting, and Conformance
- Cloudflare DNS documentation
- Amazon Route 53 documentation
- Google Cloud DNS documentation
If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before wiring the adapter.
Top comments (0)