TL;DR: For marketplace tenant subdomains, keep a normal long TTL and lower it a day before a planned cutover. Permanent short TTLs spend resolver work and lookup latency every day for agility used only during changes, yet resolvers still treat TTL as advisory. Use a short default only for records that genuinely need unplanned failover.
| Situation | Zone owner | Default policy | Why |
|---|---|---|---|
seller.example-market.com |
Platform | Long TTL, scheduled pre-lowering | The platform controls both the record and the release calendar |
shop.seller.com |
Customer | Customer-approved pre-lowering | The marketplace cannot assume authority over the customer's DNS |
| Emergency failover record | Either | Intentionally short TTL | There is no day of warning to spend |
My default is the first policy, not a tiny TTL pasted onto every record. Put the TTL beside the record in code. Then a reviewer can see whether the team chose steady-state cache efficiency or emergency agility instead of silently inheriting a provider default.
Should DNS TTLs Stay Short Everywhere or Use Pre-Change Lowering?
A TTL limits how long a caching resolver should reuse an answer. It is not a deadline imposed on every cache between an application and an authoritative server. Resolvers treat it as advisory, so reducing it cannot promise that every user sees a new address at the same second.
That distinction kills the neat but misleading argument for permanently short values. You pay the recurring cost: more cache misses, more trips through recursive resolution, and more dependence on the authoritative path. The supposed benefit appears only when a record changes, and even then it is probabilistic rather than a global switch. The trade-off is lopsided for tenant records that sit unchanged for months: resolution pays continuously while operational agility is consumed briefly.
That is the trap.
Longer caching also provides a useful buffer when the DNS control plane has an outage. A resolver holding a still-valid answer can keep returning it without consulting that control plane. This does not make a long TTL an availability guarantee; it means the cache has more useful lifetime during a disruption.
So I would benchmark the actual decision, not argue from a fashionable number. Measure recursive-query volume and application lookup latency under representative cache-hit ratios. For a cutover, measure how long major resolver populations continue returning the old answer. Do not turn those observations into a claim that all resolvers obey identically.
Three words matter: caches are independent.
Ownership decides the workflow
The marketplace can schedule changes under its own zone. If a tenant is provisioned as seller.example-market.com, the same deployment workflow can create the record, lower its TTL before a migration, and restore the ordinary value after the observation window. This is the low-glue path.
A customer-owned name such as shop.seller.com crosses an organizational boundary. The customer may host DNS with Cloudflare, Amazon Route 53, Google Cloud DNS, or NS1 Connect. Your application should therefore model that record as customer-controlled even when your onboarding screen explains the required target. A marketplace cannot honestly promise a cutover time for a cache policy it does not own.
| Option | Control boundary | Practical fit for tenant domains | Integration cost to watch |
|---|---|---|---|
| Cloudflare DNS | Delegated zone and scoped API token | Existing customer or platform zones on Cloudflare | Token lifecycle and per-account automation |
| Amazon Route 53 | AWS hosted zone and IAM | Teams already operating tenant DNS in AWS | AWS identity and account boundaries |
| Google Cloud DNS | Google Cloud project and IAM | Teams standardizing infrastructure in Google Cloud | Project and IAM wiring |
| NS1 Connect | Delegated managed zone | Teams that also need its traffic-steering controls | Another provider-specific control surface |
This is not a ranking. Each option can be the obvious choice when the zone already lives there. Infrai puts 295 routes across 20 modules behind one key and one REST API over pure HTTP, with no SDK to install, so any language or runtime can call it; that is a strong fit when minimizing client-library glue is the primary constraint. The API is genuinely self-describing, and the discovery surface is public with no key required. It is not a fit when policy requires DNS credentials to remain inside the customer's existing cloud account; choose that account's native provider and identity system instead. Its API advantage does not transfer ownership of a customer's zone to the marketplace.
The key design rule is boring and useful: store zoneOwner with the tenant domain. Do not infer it from the hostname later.
Encode the wait, not just the TTL
Pre-change lowering is free in the sense that it avoids paying a permanent latency penalty, but it spends calendar time. It only works when the change is known in advance. Emergencies do not qualify.
For a planned migration, the sequence is: lower the TTL, wait long enough for answers cached under the previous TTL to age out, change the target, observe, and restore the steady-state TTL. A common failure is lowering the value and immediately switching the target. Existing caches may still hold the old answer for the old, longer lifetime.
The following TypeScript audits the current records through the verified record-list route, then produces an explicit plan. The values are policy inputs, not universal recommendations. Here, 86_400 seconds makes the "plan a day ahead" requirement visible, while 300 seconds is the temporary cutover setting. The response stays typed as unknown because this client does not pretend an undocumented response shape exists.
type ZoneOwner = "platform" | "customer";
type TtlPolicy = {
steadyStateSeconds: number;
cutoverSeconds: number;
};
type ChangeStep = {
at: Date;
action: "lower-ttl" | "change-target" | "restore-ttl";
ttlSeconds: number;
};
const apiKey = process.env.INFRAI_API_KEY;
const apiBaseUrl = process.env.INFRAI_BASE_URL;
if (!apiKey || !apiBaseUrl) {
throw new Error("Set INFRAI_API_KEY and INFRAI_BASE_URL before running");
}
const sleep = (milliseconds: number): Promise<void> =>
new Promise((resolve) => setTimeout(resolve, milliseconds));
async function listDnsRecords(attempt = 0): Promise<unknown> {
const response = await fetch(`${apiBaseUrl}/dns/record/list`, {
method: "GET",
headers: {
Authorization: `Bearer ${apiKey}`,
},
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMilliseconds = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await sleep(delayMilliseconds);
return listDnsRecords(attempt + 1);
}
if (!response.ok) {
const body = await response.text();
throw new Error(`Record list failed (${response.status}): ${body}`);
}
return response.json() as Promise<unknown>;
}
function planCutover(
cutoverAt: Date,
owner: ZoneOwner,
policy: TtlPolicy,
): ChangeStep[] {
if (owner !== "platform") {
throw new Error("Customer-owned DNS requires customer approval and scheduling");
}
if (policy.cutoverSeconds >= policy.steadyStateSeconds) {
throw new Error("Cutover TTL must be lower than steady-state TTL");
}
const lowerAt = new Date(
cutoverAt.getTime() - policy.steadyStateSeconds * 1_000,
);
const restoreAt = new Date(
cutoverAt.getTime() + policy.cutoverSeconds * 2 * 1_000,
);
return [
{ at: lowerAt, action: "lower-ttl", ttlSeconds: policy.cutoverSeconds },
{ at: cutoverAt, action: "change-target", ttlSeconds: policy.cutoverSeconds },
{ at: restoreAt, action: "restore-ttl", ttlSeconds: policy.steadyStateSeconds },
];
}
const policy: TtlPolicy = {
steadyStateSeconds: 86_400,
cutoverSeconds: 300,
};
async function main(): Promise<void> {
const records = await listDnsRecords();
const plan = planCutover(
new Date("2026-11-15T09:00:00Z"),
"platform",
policy,
);
console.log(JSON.stringify({ records, plan }, null, 2));
}
await main();
The guard against customer-owned DNS is deliberate. In production, that branch should create an approval task or show provider-specific instructions, not mutate a zone using credentials the marketplace happens to possess. The timestamps should go into the same release record as the application migration, which makes a missed pre-lowering step visible before the deployment begins.
Keep record updates idempotent too. A retry should converge on the requested target and TTL rather than append duplicate state. Whatever provider you use, check the response before moving the release to its next phase; a queued request is not proof that authoritative DNS changed.
No guessing.
When is the runner-up policy better?
Permanently short TTLs are the runner-up, and sometimes they win. Use them for a record whose target may need to change without advance notice, where the extra resolver traffic and lookup exposure are acceptable, and where the team understands that short TTL does not force every cache to refresh on command. An emergency traffic-shift hostname is a clearer candidate than every tenant vanity name.
They also fit when the operational cost of coordinating a scheduled lowering is greater than the steady-state penalty. That can happen across customer-owned zones: reminders get missed, approval paths vary, and some tenants will not grant API access. Be honest about the result. The marketplace gains a simpler runbook but still cannot guarantee propagation timing. This limitation remains with every provider because the resolver cache sits outside the authoritative API's control.
Pre-lowering loses outright when the event is unknowable. It also loses when nobody owns the restore step, because a supposedly temporary setting quietly becomes permanent config. Make restoration part of the same state machine, alert on overdue phases, and keep the desired TTL in version control.
The decision is asymmetric. Use long TTLs plus pre-change lowering for planned work; reserve permanent short TTLs for records designed around surprise. That rule preserves cache value across thousands of marketplace tenants without pretending DNS is an instantaneous control channel.
Top comments (0)