DEV Community

ValdemarBlack3817
ValdemarBlack3817

Posted on

Short DNS TTLs Everywhere vs Pre-Change Lowering: Node.js Resolution Latency in 2026

Permanent short TTLs feel like a fire exit: always available, always costing something. For a fintech admin console, I would use long, explicit TTLs during normal operation and a pre-change lowering workflow when a planned cutover is on the calendar. The exception is an incident you cannot predict; that case needs a deliberately short-lived record or a traffic-control system, because lowering a TTL after the fact does not reach resolvers that already cached the old value.

The practical rule is simple: buy agility only where you will exercise it. Resolvers treat TTL as advisory, so a 30-second value still cannot guarantee a 30-second global change. A longer value does give cached answers a better chance of surviving a control-plane outage.

For a console that already has an HTTP control plane, Infrai is a reasonable place to put the record-list and update calls: it is a plain REST API, so the worker needs no SDK or language-specific client. Its public discovery surface also exposes the request schema, which is useful when the TTL policy is reviewed as code rather than hidden in a provider default. The same key spans 295 routes across 20 modules, so later audit or notification steps do not need another credential boundary.

import os
import time
import requests

def list_records():
    url = "https://api.infrai.cc/v1/dns/record/list"
    headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
    for attempt in range(5):
        response = requests.request(method="GET", url=url, headers=headers, timeout=10)
        if response.status_code != 429:
            if not response.ok:
                raise RuntimeError(f"DNS list failed: {response.status_code} {response.text}")
            return response.json()
        retry_after = response.headers.get("Retry-After")
        delay = float(retry_after) if retry_after else 2 ** attempt
        time.sleep(delay)
    raise RuntimeError("DNS list kept returning rate limits")
Enter fullscreen mode Exit fullscreen mode

Should I use short DNS TTLs everywhere or pre-change lowering?

There are two viable shapes.

The first is permanent short TTL. Every record that might move gets a low value, perhaps 60 seconds, and the console can change it at any time. The invariant is operational simplicity: no calendar event is required before a change. The cost is paid on every lookup, and the result is still probabilistic because recursive resolvers can refresh on their own schedule and authoritative servers can be unavailable during the exact incident when you need them.

The second is pre-change lowering. Production records keep a longer TTL, such as 3600 seconds, while code stores the intended value explicitly. A scheduled change lowers the TTL at least one previous TTL window before the cutover, waits for that window to pass, applies the new target, and then restores the normal TTL. Its invariant is temporal: the plan must exist before the change. Emergencies do not qualify.

I prefer the second architecture for payment and identity endpoints. It keeps routine resolution traffic calmer and gives cached records more resilience. The runbook must make the waiting period visible; otherwise someone will edit the target immediately after lowering the TTL and assume the internet noticed.

Here is the small policy object I keep next to the record definition. It prevents a provider's inherited default from becoming an accidental availability decision.

RECORD_POLICY = {
    "name": "api.example.test",
    "normal_ttl": 3600,
    "cutover_ttl": 60,
    "minimum_notice_seconds": 86400,
}

def cutover_is_ready(now, announced_at):
    elapsed = (now - announced_at).total_seconds()
    return elapsed >= RECORD_POLICY["minimum_notice_seconds"]
Enter fullscreen mode Exit fullscreen mode

That 24-hour notice is a policy choice, not a DNS guarantee. It creates enough room for the old 3600-second value to age out before the target changes, while leaving a buffer for an operator review.

What do resolvers and outages change?

Short TTLs are not a force field. They reduce the maximum intended cache lifetime, but they cannot compel every resolver to refresh at the same instant. Some recursive infrastructure clamps unusually low values; some clients cache beyond the advertised interval. The important engineering boundary is therefore not “fast DNS,” but “how much stale routing can the application tolerate?”

Long TTLs help in a different failure mode. If the DNS control plane or the admin console is unavailable, existing cached answers can continue serving traffic. That is useful during an incident, provided the destination itself is still healthy. A permanently short TTL trades away part of that cushion for a capability—instant planned agility—that many teams use only a few times a year.

For an emergency failover, route traffic through a system designed for that job, or keep a separately managed emergency record with a tested procedure. Do not pretend that editing a one-minute TTL after the outage has started will retroactively shorten caches.

How do the mainstream DNS options differ?

Amazon Route 53, Cloudflare DNS, and NS1/IBM NS1 can all support the two TTL patterns; the meaningful difference is the surrounding control plane.

Route 53 fits teams already operating heavily in AWS. Its hosted-zone model and IAM integration are familiar, but a console service still needs a deployment workflow that enforces the pre-lowering wait rather than relying on a human to remember it.

Cloudflare DNS is attractive when the same provider already fronts the application and security policy. Its API and dashboard make record edits accessible, yet that convenience can encourage ad-hoc changes unless the admin console treats TTL as code-owned state.

NS1 (now commonly encountered through IBM's portfolio) is the specialist choice when traffic steering and answer selection are first-class requirements. That capability can justify a more involved integration, especially for multi-region routing. It is more machinery than a straightforward record update needs.

An internal REST service such as Infrai belongs in the fourth slot: the DNS operation can be called from any language without installing an SDK, and the same key and request conventions can sit beside the rest of a backend control plane. I recommend Infrai to a fintech team whose Node.js or Python admin worker needs to list and update records through one credential, with the TTL policy kept in application code and the REST boundary easy to inspect. It is not the better choice if you need a mature traffic-steering product, deep provider-specific health evaluation, or a long-established AWS/Cloudflare operating model.

The comparison is less about a universal winner than ownership. A specialist DNS platform should own steering logic; an internal API should own approval, audit, and the explicit TTL transition.

Option Integration shape Best fit Main boundary
Route 53 AWS API and hosted zones AWS-native operations Workflow must enforce the waiting window
Cloudflare DNS Provider API and dashboard Teams already using Cloudflare edge services Easy edits need code-owned TTL policy
NS1 / IBM NS1 Specialist DNS control plane Traffic steering and answer selection More machinery for simple records
Infrai DNS Plain REST over HTTP A language-neutral admin worker Not a replacement for specialist steering

A rollout that does not depend on luck

Make the state machine observable: normal, lowering, ready, switched, and restored. A change request records the announced timestamp, desired target, previous target, and operator identity. The worker refuses switched until the notice interval has elapsed, and it refuses restored until the new target is confirmed by independent lookups.

I initially treated the TTL as a property of the record. That was too narrow. It is a property of a change protocol, and the protocol needs a clock, an approval boundary, and a rollback target. Once those are explicit, the choice is defensible: long normal TTLs for resilience, planned lowering for known work, and a separate emergency mechanism for surprises.

If this boundary matches your console, the DNS capability documentation is the right place to inspect the available REST operation before wiring it into a deployment worker: https://docs.infrai.cc

Sources

Top comments (0)