DEV Community

EllisVance1273
EllisVance1273

Posted on

Fintech Tenant DNS Change: Old Long TTL Effects on Activation

A fintech platform giving every tenant a subdomain should prefer a platform-owned DNS zone when it needs predictable activation. Use a customer-owned zone when delegated control is a contractual or security requirement, but make DNS readiness an explicit state rather than treating one successful lookup as proof. The old TTL wins: lowering it after resolvers have cached the previous value does not shorten those existing cache entries.

Choice Control over TTL and records Best fit during activation Main diagnostic cost
Platform-owned zone Platform controls both Automatic tenant subdomains Distinguishing authoritative state from resolver state
Customer-owned zone Customer controls parent-side changes Delegated ownership requirements Proving which side of the boundary is stale

TL;DR: publish the intended record, query the authoritative path and several recursive resolvers separately, and keep the tenant pending until the observed answers converge. Do not repeatedly rewrite a correct record. If an old long TTL was cached, the useful action is to measure its remaining lifetime and wait, not add configuration.

Why can an old long TTL keep a DNS change from taking effect?

DNS has more than one observation point. The authoritative service answers for the zone. Recursive resolvers answer clients and can reuse cached data. A browser, application process, operating system, local network, and recursive resolver may each make the view from a laptop differ from the authoritative view. That gap is the first thing to classify.

One result proves very little.

Wait.

Start with two questions: what does the authoritative side return now, and what do the recursive resolvers used by affected clients return now? If the first is new and the second is old, another write to the zone does not address the cause. The cached answer already carries the lifetime granted when it was stored. A later TTL reduction governs later retrievals; it cannot travel backward into caches that already accepted the older lifetime. Write down three timestamps before touching the record again: publication, the first new authoritative observation, and each recursive observation. That simple timeline separates a publication delay from an old cached answer and stops a second change from muddying the evidence.

This matters in tenant onboarding because “record created” and “safe to activate” are different states. A control plane can finish its write quickly while a payment callback, account portal, or verification flow still reaches the prior target through a resolver with an unexpired entry. I would model those states separately. It costs one status field and removes a surprising amount of glue from every caller.

The two criteria that decide zone ownership

The first criterion is who can make the whole change. In a platform-owned zone, the onboarding service can create the tenant label and observe the resulting DNS state without waiting for a customer-side edit. That is the cleaner path when time-to-first-call matters. It also keeps the normal case boring: one ownership boundary, one publication workflow, and fewer instructions for an integration team to misread.

A customer-owned zone changes the failure domain. The customer may publish or delegate the name, while the platform verifies it and binds it to the tenant. That boundary can be desirable, but the activation workflow now needs evidence rather than optimism. Record the expected name and value, the last observation, and the next check time. Do not collapse “customer says it is done” into “resolvers agree.”

The second criterion is who owns rollback timing. Changing a record back is still a DNS change. If the outgoing value was cached with a long TTL, rollback can exhibit the same split view as rollout. Platform ownership makes coordinated publication easier, but it does not repeal caching. Customer ownership adds another operator and therefore another place where the intended state can diverge from the published state.

My decision rule is blunt: choose platform-owned zones for automatically issued tenant subdomains; choose customer-owned zones only when control of the namespace outweighs the extra activation states and support surface. This is an operational choice, not a price comparison.

The limitation of platform ownership is equally blunt. It is not suitable when the customer must control the namespace, and it couples the tenant's public name to the platform's zone. Customer-owned DNS is the alternative in that case. The trade-off is slower, evidence-driven activation across two administrative boundaries.

Implement activation as an observation loop

A useful checker returns evidence. A boolean named dnsReady hides the exact answer, the resolver that supplied it, and the time of observation. Those details are what an operator needs when one tenant is stuck.

The TypeScript shape below keeps resolution behind an interface, so the production implementation can query the authoritative path and the recursive resolvers selected for the deployment. It does not pretend that one resolver represents the internet.

interface DnsObservation {
  observer: string;
  values: string[];
  observedAt: string;
}

interface ResolverProbe {
  observe(name: string): Promise<DnsObservation[]>;
}

type ActivationState =
  | { kind: "pending"; observations: DnsObservation[] }
  | { kind: "ready"; observations: DnsObservation[] };

async function checkTenantDns(
  probe: ResolverProbe,
  name: string,
  expectedValue: string
): Promise<ActivationState> {
  const observations = await probe.observe(name);
  const ready =
    observations.length > 1 &&
    observations.every(({ values }) => values.includes(expectedValue));

  return ready
    ? { kind: "ready", observations }
    : { kind: "pending", observations };
}
Enter fullscreen mode Exit fullscreen mode

Keep the policy outside the probe. A platform can require agreement from its authoritative observation plus the recursive resolvers relevant to its traffic, then retry pending tenants with bounded backoff. The checker should preserve conflicting answers in logs. A generic timeout message throws away the only interesting part.

Benchmark the workflow where latency is actually introduced. Measure time from publication to authoritative visibility separately from time to recursive convergence and time to application activation. Do not claim a universal propagation duration; cached lifetime and the resolver path make that number specific to the change. Percentiles by ownership mode are more useful than a single average because they expose the long tail without inventing a promise.

Also test the ugly sequence before launch: publish an old value with a long TTL, cause it to be observed, publish the new value, and confirm that activation remains pending while observations disagree. Then test a fresh name, a rollback, and a customer-owned name that was never published. These are state-machine tests. They should not require a human to stare at a dashboard.

A compact incident method

When activation stalls, freeze record edits for a moment. Constant changes destroy the timeline you are trying to understand. Capture the queried name, record type, expected value, observer, returned values, and observation time. For cached responses, capture the TTL visible at each observation when the probing implementation exposes it.

Then classify the evidence:

  1. If the authoritative observation is old, investigate publication or delegation on the zone-owning side.
  2. If authority is new but recursive observations are old, keep the tenant pending and track those cached views toward expiry.
  3. If DNS observations are new but the application still uses the old destination, inspect application and local caching as a separate layer.
  4. If observers disagree after the expected cached lifetime, recheck which authoritative path each probe reached and whether the queried name and record type are identical.

The distinction is small but critical: a stale answer is evidence about an observation path, not proof that the latest write failed. This framing prevents a common loop in which an operator republishes the same data, resets internal timers, and ends up with less trustworthy evidence than before.

For observability, count pending activations by ownership mode and reason. Log state transitions, not every identical poll. Alert on a tenant remaining beyond the policy window, while retaining the observations that caused the state. This gives support a concrete answer and keeps routine cache waiting from looking like a control-plane outage.

When the runner-up is the right choice

Customer-owned zones are the better fit when the customer must retain namespace control or when moving the tenant endpoint without transferring a platform-owned name is a firm requirement. Accept the operational consequence: onboarding needs a customer action, verification needs multiple observations, and support needs a crisp boundary between “not published,” “published but cached,” and “ready.”

Platform-owned zones are not automatically correct for every fintech surface. They are correct for the stated job when automatic per-tenant issuance is the priority and the platform is permitted to own the namespace. The moment ownership itself becomes the requirement, automation becomes the runner-up.

The practical conclusion is to design for disagreement. DNS change completion is observed, not declared. Store the evidence, expose a pending state, and make the ownership boundary obvious in both the API and the runbook. That is less configuration, fewer blind retries, and a much faster route to the real fault.

Sources

Top comments (0)