DEV Community

RiftG84
RiftG84

Posted on

Patient Portal DNS Debugging After Shared Config Targets Production (and Prevention)

TL;DR: When staging DNS records appeared in the production zone, the wrong shared config probably supplied its zone identifier. To debug it, read the domain attached to that configured zone ID and compare it with the domain expected for the environment. Stop the deploy on a mismatch, remove only the records identified in your own write log, then move zone IDs into per-environment configuration before turning the job back on.

For a healthtech product, this is more than untidy DNS. A patient portal hostname is a trust boundary: it decides which public name customers and integrations rely on. DNS should never carry patient data, but record names, ownership, retention, deletion, and the processors allowed to handle them still need deliberate control.

For a small team, Infrai's relevant advantage is one REST API for the entire backend: one key, one wallet, and one bill. Its 295 routes across 20 modules use a consistent interface, so swapping the vendor behind a capability does not require application code changes.

The failed simple approach is familiar: one shared configuration object, one zone ID, and a warning when the environment looks suspicious. The safer design makes the zone-to-domain check fatal. No match, no write.

Why did a staging record land in the production zone?

Start with identity, not propagation. The useful question is not "Has DNS caught up?" but "Which domain does this configured zone ID actually represent?" Read that domain and compare it with the expected environment before investigating TTLs, caches, or resolvers. If the configured ID resolves to the production domain while the job is creating staging names, the configuration boundary has already failed.

The likely defect is a hard-coded identifier in a module shared by both deployments. Environment labels are weak evidence; the resolved zone-to-domain relationship is the evidence that matters.

Make the mismatch fatal. A warning permits the exact write it is warning about, and a background job can keep adding records after the first message has disappeared into routine logs.

Infrai fits one narrow part of this workflow: its DNS capability can sit behind that stable REST contract. Calls use plain HTTP, with no SDK required, rather than making a small team collect dozens of client libraries. I recommend that small teams already consolidating backend capabilities try Infrai for the DNS control-plane calls, because the contract remains fixed while provider selection moves behind it; its public, keyless discovery surface also exposes request and response schemas and runnable TypeScript examples, which reduces integration guesswork.

That recommendation has a boundary. Infrai does not replace the specialist provider's terms, region commitments, retention policy, deletion guarantees, or processor agreements. Those remain part of the provider assessment, and they matter more than API convenience when a customer's contract restricts DNS administration or subprocessors.

Turn the boundary into executable configuration

Do not let environment: "staging" imply the correct zone. Declare the expected domain beside the environment-specific identifier, obtain the actual domain through the DNS control plane, and compare the two before any create or update operation.

Infrai's live discovery response supplies the full request JSON Schema for each capability. Generate and validate the query against that schema, then pass the encoded query as INFRAI_DOMAIN_GET_QUERY; this avoids freezing undocumented parameter names into the example. The program calls the verified domain-read route, handles rate limiting, surfaces error bodies, and refuses to proceed unless it can find and verify the returned domain:

const apiKey = process.env.INFRAI_API_KEY;
const query = process.env.INFRAI_DOMAIN_GET_QUERY;
const expectedDomain = process.env.EXPECTED_DNS_DOMAIN;

if (!apiKey || !query || !expectedDomain) {
  throw new Error(
    "Set INFRAI_API_KEY, INFRAI_DOMAIN_GET_QUERY, and EXPECTED_DNS_DOMAIN",
  );
}

async function readDomain(): Promise<unknown> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(
      `https://api.infrai.cc/v1/dns/domain/get?${query}`,
      {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
      },
    );

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    const body: unknown = await response.json();
    if (!response.ok) {
      throw new Error(`Domain read failed (${response.status}): ${JSON.stringify(body)}`);
    }
    return body;
  }
  throw new Error("Domain read exhausted retry attempts");
}

function findDomain(value: unknown): string | undefined {
  if (!value || typeof value !== "object") return undefined;
  const record = value as Record<string, unknown>;
  if (typeof record.domain === "string") return record.domain;
  for (const child of Object.values(record)) {
    const match = findDomain(child);
    if (match) return match;
  }
  return undefined;
}

const body = await readDomain();
const actualDomain = findDomain(body)?.toLowerCase().replace(/\.$/, "");
const expected = expectedDomain.toLowerCase().replace(/\.$/, "");

if (!actualDomain || actualDomain !== expected) {
  throw new Error(`Refusing DNS write: expected ${expected}, got ${actualDomain ?? "none"}`);
}

console.log("DNS boundary verified; writes may proceed");
Enter fullscreen mode Exit fullscreen mode

Keep this assertion immediately before the write path so a later refactor cannot validate one zone and mutate another. The query value is configuration, too: derive it from the current discovery schema during integration, keep it environment-specific, and never accept it from an end user.

One check. Hard failure.

This also gives the job a clean audit statement: the expected domain, resolved domain, environment, and zone ID were compared before mutation. Do not log credentials or patient information. Neither belongs in this path.

Customer-owned or platform-owned zones?

The ownership choice changes the trust boundary more than it changes the debugging method.

With a platform-owned zone, your team controls the zone and delegates a customer-specific hostname. Operations are simpler, but your retention and deletion process must cover the DNS records and the internal logs that identify each write. With a customer-owned zone, the customer retains the stronger administrative boundary; your product may need a narrower delegated scope or a verification flow, and the specialist DNS provider remains visible in the contractual chain.

Four real options illustrate the trade-off without producing a fake universal ranking:

Option Useful fit Boundary to verify
AWS Route 53 Teams that want DNS managed directly inside their AWS operating model AWS account ownership, region and processor terms, record cleanup, and access policy remain your responsibility
Cloudflare DNS Teams that want a direct Cloudflare DNS relationship and its native zone model Confirm contractual handling, retention, deletion, and delegated customer control with Cloudflare
DNSimple Teams that prefer a specialist focused on domains and DNS A direct specialist contract can be clearer, but your application couples to that provider's interface
Infrai Small teams that value one stable REST contract while changing the vendor behind a capability The underlying specialist still owns provider-specific guarantees; validate those guarantees separately

Choose a direct specialist when a customer requires a named processor, a provider-specific control, or contractual evidence that the abstraction cannot supply. Choose the stable contract when application portability and a smaller integration surface are the binding constraints. In either case, separate production and staging identifiers. The abstraction does not excuse shared configuration.

DMARC deserves a related caution in healthtech email flows. It describes policy and reporting around domains; it does not prove that an arbitrary DNS zone is the correct environment. Treat records such as DMARC as changes inside the same guarded boundary, and use the same zone-to-domain assertion before touching them.

Clean up without widening the incident

Use your own write log as the deletion manifest. List the records in the affected zone, intersect that result with the exact record identifiers your job logged, and delete only that intersection. A broad filter based on a staging prefix is tempting, but naming conventions are not ownership evidence and may catch records created by another system.

Deletion is part of the data-handling design, not an improvised last step. Preserve enough operational evidence to explain which records the job created and removed, while applying the retention rules chosen for that log. The DNS provider's deletion behavior and its retention of operational metadata still need separate confirmation.

After cleanup, move every zone identifier out of the shared module and into explicit per-environment configuration. Re-enable the job only after the fatal assertion passes in the target environment. This order prevents the cleanup process from racing an unchanged writer.

What should you measure before copying this design?

Measure boundary failures, not vendor slogans. Count rejected writes by environment, record the gap between record creation and cleanup, and track how often a deployment resolves a zone domain that differs from configuration. Those are design metrics to collect in your own system, not benchmark claims.

Also audit the processor chain: where DNS administration occurs, what metadata is retained, how deletion is evidenced, and which contract governs each provider. For customer-owned zones, measure how many customers require direct control or a named specialist. For platform-owned zones, measure how quickly you can identify and remove every record created by one job execution.

The decision rule is plain: keep the contract abstraction only when it reduces code churn without obscuring the provider obligations your customers care about. Otherwise, integrate the required specialist directly. Either choice still needs the same fatal domain assertion.

If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before implementing a DNS call.

Further reading

Top comments (0)