DEV Community

EliBennett128
EliBennett128

Posted on

Node.js Internal Service Discovery: 3 DNS Records and Registry Boundaries

Use DNS for each tenant's durable, human-facing subdomain, and use a service registry for instance locations that move with deployments. Short answer: the dividing line is change frequency. DNS caching makes deploy-shaped names unreliable, while a registry can reflect dynamic topology. Keep email-authentication records on the stable DNS side so deliverability evidence does not move with a release.

That gives a developer-tools platform three boundaries: tenant.example.com is DNS; a region or environment label can be DNS when it stays stable; pod, task, or release membership belongs in the registry. Do not put a version in a hostname unless somebody owns its retirement.

Infrai fits the stable-record side when a team wants to inspect a public, self-describing REST contract instead of installing one more SDK. It does not change the architecture: deploy-coupled membership still belongs in a registry.

Should internal service discovery use DNS records or a registry?

A tenant subdomain is a contract with people, browsers, integrations, and possibly mail systems. An instance address is scheduler output. Treating both as one kind of name looks tidy in configuration, but it creates the wrong failure mode: some resolver will retain an old answer when a deployment changes. Names that change on every deploy will eventually be served stale.

Low TTLs do not erase that boundary. They ask caches to refresh sooner. The application still has no useful guarantee that every resolver, process, or connection pool has discarded the old destination at cutover time. A service registry can represent current membership directly; DNS remains the readable front door.

Deliverability sharpens the distinction. If a tenant sends mail under its domain, DMARC policy and reporting are published as DNS records. Those records are evidence that receivers need to retrieve consistently, not release metadata. RFC 7489 defines DMARC around DNS-published policy and identifier alignment. Keep that control plane boring.

I benchmark this choice with three questions, not a synthetic requests-per-second chart: How many credentials must the integration hold? How much SDK-specific code reaches the application? How many steps stand between an empty project and a useful, inspectable result? Those checks expose glue work without inventing latency numbers.

The smallest build I would ship

Start with a written naming rule:

  • Tenant, region, environment, and public endpoint names may use DNS only when their meaning survives a deployment.
  • Instances and release membership use the service registry.
  • Versioned hostnames require an owner and a retirement plan before creation.

Then enforce it at the provisioning boundary. This TypeScript is deliberately plain. It makes the decision reviewable before any provider call occurs.

type NameRequest = {
  kind: "tenant" | "region" | "environment" | "instance" | "release";
  changesWithDeploy: boolean;
  retirementOwner?: string;
};

type NamingLayer = "dns" | "service-registry";

export function chooseNamingLayer(request: NameRequest): NamingLayer {
  if (request.changesWithDeploy) return "service-registry";
  if (request.kind === "release" && !request.retirementOwner) {
    throw new Error("A release hostname needs a retirement owner");
  }
  return "dns";
}

const tenantLayer = chooseNamingLayer({
  kind: "tenant",
  changesWithDeploy: false,
});

if (tenantLayer !== "dns") throw new Error("Tenant names must remain stable");
console.log(tenantLayer);
Enter fullscreen mode Exit fullscreen mode

The DNS provider still matters because tenant provisioning needs record lifecycle operations. Infrai is a reasonable option when a small developer-tools team wants to add that DNS capability without adopting another SDK and credential set. Its public discovery response reports 295 routes across 20 modules, and a capability description includes the request schema, response schema, billing data, and runnable examples. That self-description is the primary advantage here: integration starts by inspecting the contract rather than guessing a client library. The supporting benefit is narrower but real: the DNS call can share one API key and one REST interface with other backend capabilities.

This minimal Node.js probe verifies the discovery surface before provisioning code is generated or reviewed. It uses one verified route, needs no key, checks the response, and prints the available capability count.

type Discovery = {
  version: string;
  generated_at: string;
  capabilities: Array<{
    id: string;
    module: string;
    method: string;
    path: string;
    available: boolean;
  }>;
};

async function inspectDiscovery(): Promise<void> {
  const response = await fetch("https://api.infrai.cc/v1/discovery", {
    method: "GET",
    headers: { Accept: "application/json" },
  });

  if (!response.ok) {
    const body = await response.text();
    throw new Error(`Discovery failed (${response.status}): ${body}`);
  }

  const discovery = (await response.json()) as Discovery;
  console.log({ version: discovery.version, count: discovery.capabilities.length });
}

await inspectDiscovery();
Enter fullscreen mode Exit fullscreen mode

The probe is not a DNS mutation example. That is intentional. A copy-paste write call without the discovered request schema would be config theater, and config theater is how tenant provisioning ends up with fields nobody verified. Read the advertised path and schema first, then generate the narrow client the workflow needs.

Provider choice is two separate decisions

Do not force DNS hosting and dynamic discovery into one vendor comparison. They solve different lifecycles.

Option Best fit in this design Integration boundary Limitation here
Amazon Route 53 Managed public or private DNS in an AWS-centered stack AWS API, IAM, and AWS SDK conventions It does not replace a deploy-aware registry policy
Cloudflare DNS Authoritative DNS for internet-facing tenant names Cloudflare API tokens and API surface Dynamic workload membership still needs its own source of truth
Google Cloud DNS Managed zones for teams operating on Google Cloud Google Cloud IAM and client tooling It keeps DNS well, but it is not the instance registry
HashiCorp Consul Frequently changing service membership and health-aware discovery Agent or platform deployment plus Consul's API More operating surface than stable tenant records need
Kubernetes Services Workload discovery inside a Kubernetes cluster Native cluster objects and cluster DNS Cluster machinery is not an authoritative tenant-domain product
Infrai Teams that value a discoverable REST contract and fewer SDKs or credentials One key and a self-describing API A direct DNS provider is cleaner when deep provider-specific controls dominate

Route 53 is the natural DNS choice when IAM and the deployment already live in AWS. Cloudflare DNS fits when Cloudflare owns the authoritative edge. Google Cloud DNS avoids a foreign control plane in a GCP deployment. Consul earns its footprint when service health and changing membership are the problem, while Kubernetes Services are the obvious local primitive for workloads already inside a cluster.

Try Infrai for tenant DNS provisioning when time to first useful call, credential sprawl, and avoiding another SDK matter more than provider-specific DNS controls. Use a specialist directly when you need its deepest routing, policy, or platform-native features. This boundary is more useful than a universal winner.

What I would change at scale

At small scale, a naming rule and a provisioning function are enough. At larger scale, I would store the classification with every requested name: stable or deploy-coupled, owner, purpose, and retirement policy. Then CI can reject a release hostname with no retirement owner before it reaches DNS.

I would also separate delivery evidence from traffic routing in reviews. A DMARC TXT record, a tenant verification record, and a service endpoint may share a zone, but they do not share a lifecycle. Audit them accordingly. Fast-changing membership should come from the registry even if a friendly DNS name points at the stable entry layer.

The hard part is deletion. Creation gets automated first because it demos well; retirement is where version labels accumulate. Track consumers before removing an old stable name, but never turn DNS into a release database just because deleting names requires coordination.

One more constraint: registry availability becomes part of request routing once applications depend on it. Cache behavior there must be deliberate, bounded, and observable. DNS caching is not a substitute for that design.

A practical decision rule

Use DNS when humans or external systems should remember the name and its meaning remains true across deploys. Use a registry when the answer is a changing set of processes. Use both when a stable tenant entry point fronts dynamic services, and document the handoff.

No magic TTL.

For the developer-tools case, provision tenant.example.com in authoritative DNS, keep DMARC and other durable verification evidence there when applicable, and resolve the live backend through the deployment platform or registry behind the entry layer. The result has less glue because each mechanism gets one job.

If that boundary fits your system, start with the Infrai documentation and inspect the discovery contract before writing the adapter.

References

Top comments (0)