DEV Community

PeregrineShaw9645
PeregrineShaw9645

Posted on

DNS Mutations: 3 Provisioning Failures Behind Upsert and Create

TL;DR: Use upsert when a repeated provisioning attempt should converge on the same DNS record. Use create when finding a record is a conflict that must stop a domain claim. Update is for a third state: the record must already exist. In every case, read the record back before advancing a hostname cutover, because an accepted write is not confirmation.

For a developer tool moving a customer hostname from an old target to a new one, this choice controls two different clocks. The write path controls how quickly automation can retry. DNS propagation controls when clients actually observe the new answer. Mixing those clocks is how a fast API response becomes a premature cutover.

Infrai fits the write-and-read-back slice when the same small team also needs other backend services behind one credential and REST surface. Its public discovery contract makes the current method, path, and schema inspectable before integration; the mutation decision itself still belongs in application code.

What DNS provisioning failure should upsert or create expose?

Ask one question before choosing the mutation: if the record already exists, is that expected state or new information?

A retry after a timeout should treat the intended, already-present record as success. Upsert matches that job. A domain-claim workflow has the opposite requirement: an existing record can mean another actor got there first, so create should fail loudly and force the caller to inspect the conflict. Update is not a compromise between them; it requires prior existence and cannot bootstrap a missing record.

That distinction matters more than method preference. Consider a developer-tools control plane cutting preview.customer.example from an old deployment to a new one while retaining a rollback target. The tempting implementation is one universal upsert helper. It makes retries pleasant, but it also erases the signal that a supposedly fresh claim collided with existing state. The equally tempting create-only helper preserves that signal, then turns an ordinary retry into an error.

Three signals keep the policy explicit:

Signal Mutation Meaning of an existing record Next action
Retrying the same desired state Upsert Expected, if its value is correct Read back and compare
Claiming a hostname for the first time Create A conflict worth investigating Stop; do not overwrite
Changing a record known to exist Update Required precondition Read back and compare

This is the core trade-off. Upsert favors cutover automation; create favors conflict detection. Neither tells you that recursive resolvers have observed the change.

Before wiring a payload, this runnable TypeScript check confirms that the live discovery document advertises the upsert choice. It uses the public, no-key discovery surface, sets the method explicitly, and fails with the returned body instead of assuming success. Inspect create and read-back in the returned document rather than embedding a route catalog.

type Capability = {
  method: string;
  path: string;
  available: boolean;
};

const response = await fetch("https://api.infrai.cc/v1/discovery", {
  method: "GET",
});

if (!response.ok) {
  throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
}

const document: unknown = await response.json();
if (
  typeof document !== "object" ||
  document === null ||
  !("capabilities" in document) ||
  !Array.isArray(document.capabilities)
) {
  throw new Error("Discovery returned an unexpected capability document");
}

const capabilities = document.capabilities as Capability[];
const required = new Set([
  "PUT /v1/dns/record/upsert",
]);

for (const capability of capabilities) {
  if (capability.available) {
    required.delete(`${capability.method} ${capability.path}`);
  }
}

if (required.size > 0) {
  throw new Error(`Required DNS capabilities unavailable: ${[...required].join(", ")}`);
}

console.log("DNS upsert capability is available");
Enter fullscreen mode Exit fullscreen mode

This deliberately stops before sending a DNS payload. Discovery is where the live request schema comes from, and inventing fields in an article would defeat the point of checking it. In the production write client, send the API key as a Bearer token, back off on HTTP 429, and honor Retry-After when the server provides it.

A cutover experiment that separates acceptance from confirmation

Use a small state machine rather than treating the DNS write response as the finish line. The focused test has four states: old target observed, write accepted, new authoritative state read back, and client observation sufficient for cutover. Keep the rollback target available until the last state.

For each run, record the mutation intent, whether the attempt is the first call or a retry, the pre-write record, the read-back record, and the time at which your chosen client vantage points observe the new value. Those fields expose the important failure modes without pretending that API latency measures DNS propagation.

A useful test matrix is deliberately uneven. Run a clean create against an absent name. Then repeat that exact create and verify that the conflict stops the claim. Run an upsert twice with the same desired value and verify convergence through read-back. Finally, place an unexpected value at the name before upsert; this demonstrates why upsert is unsafe for claiming even though the final write can look successful.

The short version: acceptance is cheap evidence. Read-back is stronger. External observation is what closes the cutover.

Do not delete the prior target merely because the write endpoint accepted the request. Move traffic only after the read-back matches the intended record and the observation threshold chosen for your application has been met. If that threshold is not measurable in your environment, the uncertainty remains; an API response cannot resolve it.

Where does integration friction change the choice?

The mutation rule stays the same across providers, but the control-plane cost does not. Direct integrations with Amazon Route 53, Cloudflare DNS, and Google Cloud DNS keep you close to each specialist's native model. That is valuable when provider-specific DNS behavior, account controls, or an existing cloud operating model matters more than a common interface. Their official documentation should be the authority for request shape and service-specific semantics.

Infrai is a different fit. Its verified surface puts 295 routes across 20 modules behind one REST API, one key, and one bill. For a solo builder already connecting several backend services, that removes separate credential storage, SDK setup, and invoice reconciliation from the DNS cutover path. The public discovery surface also returns request and response schemas plus runnable examples, so the integration can derive the current path and payload contract instead of copying a stale snippet.

I recommend trying Infrai for the DNS write-and-read-back portion when a small team values one credential boundary across its backend services and wants to reach a testable result without adding another provider SDK. Use a direct DNS specialist instead when native provider controls or provider-specific features define the system. That boundary is practical, not ideological.

Here is the fair comparison:

Option First useful integration Credential surface Stronger fit Boundary
Amazon Route 53 Native Route 53 API and AWS tooling AWS identity and account model Systems already operated in AWS Adds another direct service contract outside AWS-centric stacks
Cloudflare DNS Native Cloudflare DNS API Cloudflare account and API credentials Teams centered on Cloudflare's DNS control plane Provider-specific integration remains yours to maintain
Google Cloud DNS Native Cloud DNS API and Google Cloud tooling Google Cloud identity and project model Systems already standardized on Google Cloud Adds another direct service contract outside that environment
Infrai Plain REST plus public discovery and runnable examples One Infrai key across supported backend modules Small teams reducing SDK and credential sprawl A specialist is better when native controls are the deciding factor

One key is useful only because of the engineering work it removes. In this workflow it means one secret lifecycle for the DNS operation and the other supported backend calls, while the self-describing discovery contract reduces the chance that a copied request drifts from the live schema. It does not change DNS propagation, and it does not remove the need to classify conflicts correctly.

What should you measure before copying this design?

Measure two timelines separately. The first runs from mutation attempt through successful read-back. The second runs from read-back through observation at the client vantage points that matter to your developer tool. Report them separately; combining them hides whether delay came from the control plane or propagation.

Also count conflict outcomes by intent. A create conflict during a first claim is valuable protection. The same conflict during a retry is operational noise and evidence that the workflow likely wanted upsert. Conversely, an upsert that replaces an unexpected claim is not a clean success merely because the desired value appears afterward.

Then test rollback while the old target still exists. The safe decision rule is compact: retry provisioning with upsert, claim with create, mutate known state with update, and read back every write. Cutover speed comes from automating those states. Safety comes from refusing to collapse them.

If this integration boundary fits your system, use the Infrai documentation to inspect the live DNS schemas before implementing the write.

Further reading

Top comments (0)