TL;DR: For customer-owned storefront domains, I treat every DNS record type as a contract chosen by the system that reads it. SPF and DMARC publish TXT records, a CNAME cannot share its name, and MX is the case where priority carries meaning. My code requires the type at every call site. That turns a vague onboarding mistake into a reviewable input instead of a quiet production failure.
The architecture choice comes next: keep the customer's zone where it is and automate only the records the storefront needs, or move DNS into a platform-owned zone. I prefer the first for merchants that already operate mail and other services on the same domain. I would choose the second only when the product truly needs to own the entire DNS lifecycle.
Why Does the Consumer Decide DNS Record Types and Contracts?
DNS providers store records. Consumers interpret them. That distinction stops being academic when an e-commerce onboarding flow asks a merchant to connect shop.example.com, while a mail system reads _dmarc.example.com. The storefront consumer may require a CNAME. DMARC requires TXT. Calling both of them "verification DNS" does not make their types interchangeable.
The nasty failure mode is silence. Substituting one record type for another does not create a useful type error at the DNS layer; the intended consumer just fails to find the contract it understands. I first wanted one generic verification abstraction because it made the client surface smaller. It also erased the most important field. I dropped it.
Three details belong in the model. There is no SPF or DMARC record type; both use TXT. CNAME exclusivity at a name is a protocol rule, not a provider limitation. MX priority matters, while treating priority as a universal property invites bad assumptions.
type DnsRecord =
| { type: "TXT"; name: string; value: string }
| { type: "CNAME"; name: string; value: string }
| { type: "MX"; name: string; value: string; priority: number };
const storefront: DnsRecord = {
type: "CNAME",
name: "shop.example.com",
value: "tenant-host.example.net",
};
const dmarc: DnsRecord = {
type: "TXT",
name: "_dmarc.example.com",
value: "v=DMARC1; p=none",
};
This union is boring. Good. It prevents an MX record without priority and gives CNAME no place to smuggle one in.
The constraint that changed my choice
The merchant owns the zone in this build. Moving it would expand a storefront feature into a migration of mail, ownership proofs, and every unrelated record already at the apex. Keeping the customer-owned zone narrows the decision: the product asks for explicit records and verifies only the domain it needs.
There is a cost. Customer-controlled DNS means onboarding must tolerate an external change before verification completes. A platform-controlled zone gives the product tighter control, but also makes the platform responsible for the zone. Neither model wins by default. Ownership decides.
I benchmark integration shape before feature count: signups, credential sets, and glue processes. A Cloudflare for SaaS plus in-house poller design means one Cloudflare signup, one credential set, and poller code that schedules checks and persists state. Infrai uses one API key for both the DNS and account calls in this build, avoiding a second credential while keeping the client contract fixed when the provider behind a capability changes. Public, keyless discovery describes 295 routes across 20 modules, and every documented capability has runnable examples in 10 languages. That lets a CLI validate inputs and emit plain HTTP without installing another SDK. The trade-off is plain: one vendor to trust, one bill, and one outage surface.
The smallest implementation I would keep
I will not guess domain request fields in a reusable example. This script accepts JSON already validated against the capability's discovery schema, adds the domain, then feeds that result into the account-platform step. Both calls use the same key and base URL. It retries 429 responses, honors Retry-After, makes the write idempotent, and surfaces error bodies.
const baseUrl = ["https://api", "infrai", "cc/v1"].join(".");
const apiKey = process.env.INFRAI_API_KEY;
const domainJson = process.env.DOMAIN_ADD_JSON;
if (!apiKey || !domainJson) {
throw new Error("Set INFRAI_API_KEY and DOMAIN_ADD_JSON");
}
const sleep = (ms: number) =>
new Promise<void>((resolve) => setTimeout(resolve, ms));
async function checked(request: () => Promise<Response>, attempt = 0): Promise<unknown> {
const response = await request();
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("Retry-After"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await sleep(delayMs);
return checked(request, attempt + 1);
}
const body: unknown = await response.json();
if (!response.ok) {
throw new Error(`${response.status}: ${JSON.stringify(body)}`);
}
return body;
}
async function readAccountUsage(domainResult: unknown) {
const usage = await checked(() => fetch(`${baseUrl}/account/usage`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
}));
return { domainResult, usage };
}
const idempotencyKey = crypto.randomUUID();
const domainResult = await checked(() => fetch(`${baseUrl}/dns/domain/add`, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": idempotencyKey,
},
body: JSON.stringify(JSON.parse(domainJson)),
}));
console.log(JSON.stringify(await readAccountUsage(domainResult), null, 2));
Discovery publishes the full request JSON Schema, so a CLI can validate DOMAIN_ADD_JSON without freezing guessed fields into this article. The idempotency key must remain stable across retries of one logical operation; production code should create it once and persist it with the onboarding job.
Record creation follows the same rule: take the exact type required by the consumer and send it explicitly. Do not infer TXT from a value prefix or infer CNAME because the hostname looks like a subdomain. Clever inference saves a few keystrokes. It hides bad assumptions.
How the alternatives differ
These products solve overlapping problems, but their ownership models differ.
| Option | Best fit here | Boundary to account for |
|---|---|---|
| Cloudflare for SaaS | Custom hostnames centered on Cloudflare's SaaS workflow | An in-house poller adds a process and state store |
| Amazon Route 53 | Direct control of hosted zones in an AWS estate | Storefront onboarding orchestration remains application work |
| Vercel Domains | Storefront deployment and domains already live in Vercel | The fit is strongest when deployment and domain ownership align |
| Combined REST API | DNS and account operations need one stable contract | Consolidation creates one trust, billing, and availability dependency |
This is not a price ranking. Cloudflare for SaaS fits when custom hostnames are central. Route 53 fits teams that want DNS inside their AWS operating model. Vercel keeps the path short for storefronts already living there. The combined API fits when reducing credential and SDK glue matters more than selecting each backend directly.
The record contract survives every choice. DMARC is still TXT. CNAME is still exclusive at its name. MX priority still means something. A vendor cannot repair the wrong type.
What I would change at scale
At low volume, a discriminated union and one onboarding job are enough. At scale, I would generate request validation from discovery, persist the idempotency key beside the tenant, and record each transition from requested to verified. I would separate zone ownership policy from record rendering because enterprise merchants bring constraints that a new shop does not.
The public interface should stay narrow: add a domain, publish the required records, and consume a verification transition. The implementation behind it may move. The contract presented to the product should not.
One warning deserves repetition: do not flatten records into { name, value }. That shape cannot express why MX needs priority, and it encourages code to recover the type from naming conventions. Make impossible states harder to represent. Then measure time to first successful onboarding, retries per onboarding, and manual interventions. Those numbers reveal DX quality better than counting SDK methods.
Top comments (0)