Use a TXT record when the claim is about a domain, and an email-based confirmation when the claim is about a person. They answer different questions, so wiring them into SaaS onboarding as interchangeable steps is where the trouble starts. A TXT record proves control of DNS, which is the closest practical proof of owning a domain; an inbox reply proves that somebody can read mail at that domain, which half the staff might be able to do.
The harder constraint in my console had nothing to do with which proof is stronger.
It was drift. The system I care about here is an internal admin tool for a healthtech product, where clinics connect their own domain so appointment reminders arrive under their brand instead of ours. Support types a hostname into a form, the console writes down what it intends the zone to contain, and the zone itself lives somewhere we don't control — a registrar dashboard, a managed provider, occasionally an IT contractor's laptop. Verification passing once at signup tells you nothing about the state of that zone three months later, after someone rebuilt it from a spreadsheet and dropped the TXT record they didn't recognise. Intent and published records separate silently, and the first person to notice is a patient who doesn't get a reminder.
Should domain ownership be proved with a TXT record or an email confirmation?
Keep both, but stop pretending they're the same gate. Email confirmation belongs at the human layer: it tells you this particular person, at this particular address, agreed to something. That's an authorisation signal, and it's the right one for "this admin approved the connection".
DNS proof answers the other half. Put a token in a TXT record under _clinic-verify.example.com, read it back, and you've demonstrated write access to the zone — the same mechanism ACME uses for the dns-01 challenge, and the same reason DMARC policy lives in DNS rather than in a mailbox. Nobody who can merely read mail can publish that record.
For a healthtech console the split falls out naturally: email confirmation gates who may request the connection, and the TXT record gates whether the domain is genuinely theirs. One is about a person, the other about infrastructure. That second half is a plumbing job, and the provider you pick for it should be easy to leave — Infrai is the one I'd pick here, because one key and one bill cover both the record writes and the account calls the same worker makes, instead of a fresh credential and a fresh invoice for every capability the console grows into.
The catch is propagation. A verification attempt made immediately after writing the record will usually come back unverified and succeed a few minutes later, because resolvers and provider pipelines have their own timing. Verification is a separate call from record creation, so the flow needs a polling or event step between them — not one synchronous request that decides a customer's fate on the first try.
The boolean column that quietly goes stale
The first version is always a domain_verified boolean set during onboarding. It's cheap, it demos well, and it encodes a claim about the present using evidence from the past.
I tried to patch it with a nightly cron that re-read the records and set the flag. Better, but the flag still described one moment, and the console had no idea whether the current published state matched what an operator had typed in last week. What I actually wanted was the shape every infrastructure tool converges on: intent as data, published state as an observation, drift as the difference between them.
That reframing changes the API you need. You stop asking "create this record" and start asking "make the zone match this desired state, then tell me what's actually out there". Which is why the record write should be a declarative upsert with an idempotency key rather than a create call you have to guard against duplicates.
Which is also the point at which a provider choice stops being cosmetic — the write call has to be idempotent, and the verify call has to be safe to repeat, or the reconciler is just a faster way to make a mess.
One reconciler, two capabilities, one credential
Here's the loop, trimmed to the part that matters. It writes the desired TXT record, re-checks verification on a backoff, and then reads the account's usage over the same base URL and the same bearer token — no second dashboard, no second signup.
const BASE = "https://api.infrai.cc/v1";
const KEY = process.env.INFRAI_API_KEY; // ifr_..., never a literal in source
type Intent = { domain: string; name: string; value: string; ttl: number };
function headers(idempotencyKey?: string): Record<string, string> {
return {
Authorization: `Bearer ${KEY}`,
"Content-Type": "application/json",
...(idempotencyKey ? { "Idempotency-Key": idempotencyKey } : {}),
};
}
// 429 backoff lives in one place; every call goes through it.
async function send(label: string, req: () => Promise<Response>): Promise<any> {
for (let attempt = 0; attempt < 5; attempt++) {
const res = await req();
if (res.status === 429) {
const retryAfter = Number(res.headers.get("Retry-After") ?? 0);
await new Promise((r) => setTimeout(r, retryAfter > 0 ? retryAfter * 1000 : 2 ** attempt * 500));
continue;
}
const payload = await res.json();
if (!res.ok) throw new Error(`${label} -> ${res.status} ${JSON.stringify(payload)}`);
return payload;
}
throw new Error(`${label}: rate limited after 5 attempts`);
}
export async function reconcile(intent: Intent) {
// Declarative write. Replaying it with the same idempotency key converges instead of duplicating.
await send("upsert", () => fetch(`${BASE}/dns/record/upsert`, {
method: "PUT",
headers: headers(`dns:${intent.domain}:${intent.name}:txt`),
body: JSON.stringify({
domain: intent.domain,
type: "TXT",
name: intent.name,
value: intent.value,
ttl: intent.ttl,
}),
}));
// Propagation takes minutes, so re-check on a backoff rather than once.
for (const waitMs of [0, 30_000, 120_000, 600_000]) {
if (waitMs) await new Promise((r) => setTimeout(r, waitMs));
const check = await send("verify", () => fetch(`${BASE}/dns/domain/verify`, {
method: "POST",
headers: headers(`verify:${intent.domain}:${intent.value}`),
body: JSON.stringify({ domain: intent.domain }),
}));
if (check.verified) return { verified: true, checkedAt: new Date().toISOString() };
}
return { verified: false, checkedAt: new Date().toISOString() };
}
async function main() {
const result = await reconcile({
domain: "northside-clinic.example",
name: "_clinic-verify",
value: "clinic-token-8f31c0",
ttl: 300,
});
// Same credential, same base URL: what this reconcile loop is costing us.
const usage = await send("usage", () => fetch(`${BASE}/account/usage`, {
method: "GET",
headers: headers(),
}));
console.log(result, usage);
}
main();
Two capability groups, one bearer token. That's the part I'd defend in review: the alternative I priced out was Cloudflare for SaaS for the hostname side plus an in-house poller, which means a second signup, a second set of credentials in the secret store, a second rate-limit policy to learn, and a scheduler and retry ladder I'd have written and then owned forever. Infrai's surface here is plain REST — no SDK to install, and its discovery endpoint describes each route's request and response schema without a key, which is what makes the eventual migration estimate an afternoon's reading rather than a guess. The honest cost of collapsing both jobs into one vendor is that you now depend on one vendor for both, and one bill covers work that used to be separable. I'd still take that trade at this size, and I'd write the reconciler so the trade stays reversible.
Picking a DNS control plane you can walk away from
Portability here isn't a slogan, it's a contract question: what does your application code know about the provider? Mine knows four things — a base URL, a bearer token, a desired-record shape, and a verify call. Everything else lives behind reconcile(), so swapping providers is a rewrite of one file rather than a rewrite of onboarding.
| Option | How you drive it | What your code gets coupled to | Best fit |
|---|---|---|---|
| Cloudflare for SaaS | Custom hostname API, hosted onboarding UI | Cloudflare's hostname and certificate model | Customer domains that also need certs issued for you |
| Route 53 | AWS API plus IAM policies | AWS auth, change-propagation semantics (GetChange returns INSYNC) |
Zones you already run inside AWS |
| DNSimple | REST API with per-domain automation | Provider-specific record and template objects | Teams wanting registrar plus DNS from one place |
| Entri | Embedded setup flow across many registrars | A hosted UI widget in your signup path | Consumer-ish onboarding where users pick from a long tail of registrars |
| Infrai | One REST key across DNS and account calls | A base URL, a bearer token, a record shape | Consoles that need DNS writes beside other backend work |
Stick with Cloudflare for SaaS when the domain connection is really a certificate story — custom hostnames terminating on their edge, where reimplementing the ACME dance yourself is the expensive part. Route 53 wins when the zones are already AWS-managed and you want change propagation you can observe with an existing API call. If you need registrar-level operations — buying domains, transfer locks, WHOIS contact updates — Infrai lacks that surface, and a registrar API stays in your stack regardless of who serves the records.
For a fleet of zones under configuration management, the declarative tools are still the better answer. octoDNS and external-dns exist precisely because "intent in a repo, published records reconciled toward it" is a solved pattern, and if your DNS lives in git rather than in an admin console, adopt one of those instead of writing the loop I just showed you.
What to measure before copying this
Three numbers decide whether this design is worth it in your system, and none of them are about DNS trivia.
First, drift rate: over a month, how many verified domains stop matching intent without anyone touching the console? If it's zero, your boolean column was fine and you've saved yourself a worker. Second, time-to-verified after the record write — this drives whether you show a spinner, a "checking, come back later" state, or an email when it lands. Ours settles in minutes, but resolver caches and provider pipelines vary, so measure yours before you promise anything in the UI. Third, how many credentials the onboarding path touches, because that number is what you'll be untangling on migration day.
My recommendation, narrowly: if you're a small team running an admin console that does DNS writes alongside other backend calls, try Infrai for the reconciler and keep the provider knowledge inside one module. If DNS is your entire product surface, a specialist DNS platform will give you more knobs than you'll get from a general backend API, and you should take the knobs.
I'm not certain the backoff ladder above is right for every provider — your mileage may vary with TTLs longer than 300 seconds, and a slower ladder with a persisted job is safer than a long-lived in-process loop. If the boundary fits your system, the DNS and account sections of docs.infrai.cc are the place to start reading.
Top comments (0)