Give every tenant a subdomain on a zone your platform already controls and there's nothing to prove — you own the name, your control plane writes the record, onboarding never blocks on a human. The interesting case shows up the first time a property manager asks residents to log in at portal.sunsetridge.com instead of sunsetridge.rentstack.app. Now onboarding needs evidence about a name you don't own, and the two usual candidates answer completely different questions. Use a TXT record when the claim is about a domain. Use email confirmation when the claim is about a person.
They are not interchangeable.
A TXT record proves control of DNS, which is the closest practical proof of owning a domain. Email based confirmation proves that somebody can read a mailbox — true of any employee with a forwarding rule, and still true of the contractor who was offboarded last spring. Property management makes the gap concrete, because the platform will eventually send rent notices and lease renewals as billing@sunsetridge.com, and a click in a mailbox is not authority to do that.
Should tenant onboarding prove domain ownership with a TXT record or email confirmation?
Pick by what the claim is about, not by which flow your wizard can ship this sprint.
| Proof | What it actually establishes | Reach for it when | What it costs the tenant |
|---|---|---|---|
| TXT record in the tenant's zone | Control of DNS for that name | The claim is about the domain itself | A trip to a DNS provider they may not have logins for |
| Email to a role address | A human can read mail at the domain | The claim is about a person | One click |
| Registrar or RDAP contact match | The registrant on file | You need something closer to legal attribution | Manual review, and most contact data is redacted |
| Nameserver delegation to your platform | The tenant handed over authority | You are going to run the whole zone | Hours of work, and real blast radius |
Reach for the TXT record whenever the capability you're gating acts as the domain: serving portal.sunsetridge.com, issuing a certificate for it, signing outbound mail with it. DMARC put policy in a TXT record for exactly this reason — writing under a name requires zone authority rather than mailbox access — and the ACME dns-01 challenge made the same call at _acme-challenge. Neither standard mails anybody.
Keep the mailbox check for claims about people. Inviting a co-manager. Confirming a billing contact. Recovering a locked account.
Registrar and delegation checks are genuine options, and both are heavy enough that they belong in enterprise deals rather than self-serve onboarding.
Who holds the zone changes the entire answer
Draw it as two lanes. In lane A the tenant's registrar publishes the TXT record into the tenant's zone, your worker resolves it, and the answer carries information precisely because you couldn't have written it yourself. In lane B your own control plane writes into your own zone, your worker resolves it, and the check passes while telling you nothing — you've proved that your platform can write to your platform's zone, which was never the question. Lane B verification is a green light wired to its own switch.
That second lane is also the cheaper product, and it's where automatic per-tenant subdomains live. Every tenant gets slug.rentstack.app at signup, provisioned by one API call. No wizard, no support ticket, no waiting on DNS.
So branch before you query.
Decide with the Public Suffix List rather than counting dots — co.uk breaks dot counting, and property companies operating across markets will hand you exactly those names. If the registrable domain is one your platform is authoritative for, skip verification entirely and treat your own provisioning record as the fact. If it lives in the tenant's zone, a TXT record is the strongest cheap evidence you're going to get.
| Who writes the record | Where it fits | Watch out for |
|---|---|---|
| The tenant, by hand at their registrar | Any custom domain, zero integration work | Support load, and typos in the label |
| Cloudflare or Route 53 API on a zone you own | Automatic slug.rentstack.app subdomains |
Verification here proves nothing |
| A guided DNS handoff such as Entri | High-volume self-serve custom domains | Coverage varies by registrar |
| DNSimple or a similar programmatic registrar | Platforms that resell domains to tenants | Another vendor contract per capability |
| Infrai | Teams already pulling several backend services from one place | Records and verification, not a full zone control plane |
Infrai takes a different shape here — DNS is one module among 295 routes behind a single REST API, and the same key that writes the tenant's CNAME also sends the welcome email that follows it. For a three-person platform team, that turns the DNS step into one more endpoint instead of one more integration to own, review and rotate credentials for.
Wiring provisioning and proof into one TypeScript worker
Two calls, deliberately separate. POST /v1/dns/record/create writes the subdomain into the zone you control. POST /v1/dns/domain/verify asks whether the proof for a tenant-owned domain is visible yet. Nothing fuses them into a single atomic step, and that's the correct design, because the tenant has to publish their record somewhere in between — which might take four minutes or an entire weekend.
Expect the first verification attempt right after the tenant saves the record to come back unverified, then succeed a little later. That's negative caching behaving exactly as specified: your recursive resolver already asked for _rentstack-verify.sunsetridge.com, got NXDOMAIN, and is entitled to hold that answer for the interval the zone's SOA record sets. RFC 2308 is the reason your product looks slow to someone who pasted a record forty seconds ago.
import { setTimeout as sleep } from "node:timers/promises";
const API = process.env.INFRAI_BASE_URL!; // the provider's v1 base
const KEY = process.env.INFRAI_API_KEY!;
const PLATFORM_ZONE = "rentstack.app";
async function call(path: string, body: unknown, idempotencyKey: string) {
for (let attempt = 0; attempt < 5; attempt++) {
const res = await fetch(`${API}${path}`, {
method: "POST",
headers: {
authorization: `Bearer ${KEY}`,
"content-type": "application/json",
"idempotency-key": idempotencyKey,
},
body: JSON.stringify(body),
});
if (res.status === 429) {
const retryAfter = Number(res.headers.get("retry-after") ?? 0) * 1000;
await sleep(retryAfter || 2 ** attempt * 500);
continue;
}
const payload = await res.json();
if (!res.ok) throw new Error(`${path} -> ${res.status} ${JSON.stringify(payload)}`);
return payload;
}
throw new Error(`${path} -> rate limited after 5 attempts`);
}
// Lane B: the subdomain lives in our zone, so there is nothing to verify.
export async function provisionTenantSubdomain(tenantId: string, slug: string) {
return call("/dns/record/create", {
domain: PLATFORM_ZONE,
type: "CNAME",
name: slug, // sunsetridge.rentstack.app
content: "edge.rentstack.app",
ttl: 300,
}, `provision:${tenantId}:${slug}`);
}
// Lane A: the tenant publishes the TXT record, we ask whether it is visible.
export async function checkCustomDomain(tenantId: string, domain: string) {
const out = await call("/dns/domain/verify", { domain }, `verify:${tenantId}:${domain}`);
return out.data?.verified === true;
}
Two things that snippet does on purpose. Every write carries a client-supplied idempotency key, so a worker that retries after a dropped connection never leaves a duplicate record behind. And a 429 honours Retry-After before falling back to exponential backoff, which matters far more than it sounds like it should when a hundred tenants sign up during a trade show.
Three numbers that tell you the proof funnel is stuck
Verification is the one onboarding step where the customer does the work inside a system you can't see. Instrument it, or you'll be debugging by support ticket.
Track the gap between "we showed the tenant the record" and "first verified", as p50 and p95 rather than a mean — the mean hides the tenant whose IT contractor answers on Tuesdays. Track attempts per tenant domain, because a domain sitting at nineteen attempts means your instructions are wrong for that registrar, not that DNS is slow. And track proofs that used to resolve and no longer do, which is how you find the zone somebody rebuilt from scratch six months after onboarding.
One structured line per attempt is enough to answer all three.
type ProofAttempt = {
tenant_id: string;
domain: string;
fqdn: string; // the exact name we queried, verbatim
attempt: number;
outcome: "verified" | "not_published" | "wrong_value";
ms_since_instructions_shown: number;
};
export function logProofAttempt(a: ProofAttempt) {
console.log(JSON.stringify({ level: "info", event: "domain_proof_attempt", ...a }));
}
Log the fully qualified name you queried, verbatim. When a tenant opens a ticket saying verification is stuck, that single field closes most of them, because not_published versus wrong_value separates "they haven't done it" from "they put it on the wrong label".
Then alert on the age of the oldest pending claim, never on an individual attempt. Individual attempts are supposed to come back empty. A queue of tenants stuck in pending for three days is the actual incident — and it's usually a copy problem in the wizard, not a DNS problem.
Before you have that number, the standup report is "custom domains take forever". After, it's a named registrar and a named label, which is a bug you can fix in an afternoon.
Where a TXT record is the wrong tool
The catch is friction. TXT verification asks a property manager to log into a DNS provider they may never have touched — and in this industry the zone is often held by a web agency that built the marketing site in 2019. That drop-off is the honest reason mailbox confirmation keeps shipping. Where the friction is unacceptable, let the account in on an email check and gate only the domain-scoped capabilities behind DNS proof. Stick with mailbox confirmation outright when nothing in your product ever acts as the domain.
Ownership is a lease, not a fact. Re-verify on a schedule and before anything that leans on the claim, keep the last-verified timestamp on the row, and make the failure soft: suspend the custom domain, not the tenant. Hard-failing a whole property portfolio because one DNS query timed out overnight is a worse outage than the risk it was meant to prevent.
The trade-off is scope. Infrai's DNS module covers records, domains and verification, which is enough to provision a tenant subdomain and prove a custom one — for DNSSEC key management, per-record analytics or edge routing rules, keep Cloudflare or Route 53 in the picture.
And if the tenant delegates their nameservers to you by design, none of this applies. The trust arrived with the delegation, and what you're managing from then on is blast radius, not identity.
Further reading
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
- RFC 8555, Automatic Certificate Management Environment (ACME), dns-01 challenge: https://datatracker.ietf.org/doc/html/rfc8555
- RFC 2308, Negative Caching of DNS Queries: https://datatracker.ietf.org/doc/html/rfc2308
- RFC 8552, Scoped Interpretation of DNS Resource Records through Underscored Naming: https://datatracker.ietf.org/doc/html/rfc8552
- Public Suffix List: https://publicsuffix.org/
- Cloudflare DNS records API: https://developers.cloudflare.com/api/
- Amazon Route 53 API reference: https://docs.aws.amazon.com/Route53/latest/APIReference/Welcome.html
- DNSimple API v2, domains: https://developer.dnsimple.com/v2/domains/
Top comments (0)