TL;DR: Model DNS ownership records as an intended set and repeatedly converge the current set toward it. That turns a fragile sequence of writes into retriable provisioning with a useful drift diff. For property-management onboarding, favor a short reconcile interval during cutover, but never automate an existing zone until its records have been captured into the intended set.
A single API key and one bill across backend services can reduce credential and invoice sprawl when DNS verification is one step in a larger onboarding pipeline; a team does not have to accumulate dozens of keys and reconcile dozens of invoices. A shared API should still be evaluated beside direct DNS providers rather than treated as the automatic choice.
How should intended state and current DNS configuration converge?
Suppose a property manager connects leasing.example during onboarding. The application needs an ownership record to appear before onboarding can complete, and the decision axis is propagation delay versus cutover speed. A procedural workflow says, "create the record, then mark the step done." If the worker loses its response, it cannot tell whether a retry repairs the operation or repeats it. A sequence of writes remembers what happened last; it has no definition of correct.
An intended set supplies that definition. Read the current records, normalize both sets, compute a diff, and upsert every missing or changed member. The same input can be retried. The diff is also an alertable artifact: an empty diff means the managed records have converged, while a non-empty diff names the drift.
This is the part I would put under an eval harness before wiring it into an agent. Given fixed current and intended fixtures, the planner should always emit the same plan. Keep DNS mutation outside the model call; a prompt can help classify an onboarding request, but it should not invent the desired zone state. That split controls token cost and makes the consequential step deterministic.
One trap matters more than the rest: import existing records first. If automation treats an incomplete file as the entire intended set, records omitted from that file look unwanted and can be deleted. The tempting shortcut is to start with only the new ownership TXT record because it is the record blocking onboarding. That file is not a zone baseline. Before any authoritative reconciler gets a delete path, capture the existing set, review which records the application actually owns, preserve everything outside that ownership boundary, and test the diff against a frozen fixture. A one-record intended file can otherwise turn a quick cutover into an accidental deletion plan.
Capture first.
Build the reconciliation core first
Here is a runnable probe plus a local planner for the ownership-record slice. The probe calls the verified record-list route, handles rate limiting, and prints the response without assuming undocumented response fields. The deterministic planner uses fixtures so its diff remains testable. Save it as reconcile.py, set INFRAI_BASE_URL to the API's versioned v1 base, set INFRAI_API_KEY, and run it with Python.
from dataclasses import dataclass
from email.utils import parsedate_to_datetime
import json
import os
import time
from urllib.error import HTTPError
from urllib.request import Request, urlopen
LIST_URL = os.environ["INFRAI_BASE_URL"].rstrip("/") + "/dns/record/list"
def retry_delay(value: str | None, attempt: int) -> float:
if value is None:
return float(2**attempt)
try:
return max(0.0, float(value))
except ValueError:
return max(0.0, parsedate_to_datetime(value).timestamp() - time.time())
def fetch_current() -> object:
key = os.environ["INFRAI_API_KEY"]
for attempt in range(4):
request = Request(
LIST_URL,
method="GET",
headers={"Authorization": f"Bearer {key}"},
)
try:
with urlopen(request, timeout=30) as response:
return json.load(response)
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(f"DNS list failed ({error.code}): {body}") from error
time.sleep(retry_delay(error.headers.get("Retry-After"), attempt))
raise RuntimeError("DNS list retry loop ended unexpectedly")
@dataclass(frozen=True, order=True)
class Record:
name: str
kind: str
value: str
def normalize(record: Record) -> Record:
return Record(
name=record.name.rstrip(".").lower(),
kind=record.kind.upper(),
value=record.value.strip(),
)
def plan_upserts(current: set[Record], intended: set[Record]) -> list[Record]:
current_normalized = {normalize(record) for record in current}
intended_normalized = {normalize(record) for record in intended}
return sorted(intended_normalized - current_normalized)
def main() -> None:
print(json.dumps(fetch_current(), indent=2, sort_keys=True))
current = {
Record("leasing.example", "TXT", "owner=old-token"),
Record("mail.leasing.example", "TXT", "v=DMARC1; p=none"),
}
intended = {
Record("leasing.example", "TXT", "owner=verified-token"),
Record("mail.leasing.example", "TXT", "v=DMARC1; p=none"),
}
for record in plan_upserts(current, intended):
print(f"UPSERT {record.kind} {record.name} {record.value}")
if __name__ == "__main__":
main()
The example deliberately plans upserts rather than a create/delete script. Upsert is the primitive that makes convergence practical: rerunning the planner against unchanged intended state does not create a new desired result. The listing call is real, while the transformation from its response stays out of the sample because the supplied contract does not establish record-list response fields. Inspect the discovered schema before writing that adapter. After an adapter applies the plan, list records again and evaluate the diff. Do not equate an accepted write with globally observed propagation; keep verification as a separate read-and-compare phase.
Small boundary, honest code.
There is one simplification here. A real state key may need more than name, type, and value, but those fields are provider-specific unless a contract defines them. The important invariant is stable normalization before comparison. In a notebook, test case variants such as a trailing dot and lowercase record type; in production, promote those same fixtures into CI.
Compare the control planes, not their logos
Cloudflare DNS, Amazon Route 53, and Google Cloud DNS are the direct-provider candidates I would put on the same shortlist. Their official documentation should be the authority for each provider's current record model, authentication, and change semantics. A team already standardized on one cloud may reasonably prefer its native DNS control plane because credentials, policy, and operational ownership stay in the existing boundary.
Infrai provides one REST API for the entire backend, with one key and one bill; its API is genuinely self-describing, and its public discovery surface requires no key. That surface describes 295 routes across 20 modules, and documented capabilities have runnable examples in 10 languages. For a property platform already routing several backend functions through that shared control plane, keeping DNS onboarding there can reduce key sprawl. The practical advantage is inspectability: discovery exposes request and response schemas, billing information, and examples without requiring a key.
| Option | Sensible fit | Boundary to examine |
|---|---|---|
| Cloudflare DNS | The zone and operational policy already live with Cloudflare | Confirm its current record identity and change behavior in the official API docs |
| Amazon Route 53 | The team wants DNS operations inside its existing AWS ownership model | Confirm how the current API represents and applies record-set changes |
| Google Cloud DNS | The team operates the zone inside its Google Cloud boundary | Confirm the current managed-zone and record-set contracts |
| Shared backend API | DNS is one part of a multi-service workflow and consolidating keys and billing matters | Validate the discovered DNS schema, then keep provider-specific assumptions out of the planner |
This is not a price decision. It is an ownership-boundary decision with a propagation constraint attached. A shared API is not a fit when the DNS team requires a direct-provider control plane or already owns authentication, policy, and operations in Cloudflare, AWS, or Google Cloud; choose that native provider instead. A shared backend API is stronger when onboarding spans several services and the application team values one credential and one invoice. The limitation of the shared layer is also its boundary: provider-specific controls should be confirmed in the discovered contract rather than assumed from a direct provider's API.
Propagation delay versus cutover speed
Convergence does not remove propagation delay. It makes the application's response to delay explicit. During onboarding, run the read-diff-upsert-read loop on a tighter schedule to minimize time between observable changes, then relax it after ownership has been proven. Do not generate another token merely because the first read still shows the old state; that changes intended state while DNS may still be propagating and makes diagnosis harder.
Fast cutover raises request volume and can amplify noisy alerts. Slow reconciliation extends the period in which onboarding waits. That trade-off is deliberate. Pick the interval from the product's completion objective and provider limits, then test the state machine with delayed-read fixtures.
No guesswork.
The alert should carry the diff, the intended-state revision, and the observation time. That is far more actionable than "DNS verification failed," and it gives an eval a stable expected output. If the diff persists, operators can distinguish a record that was never applied from one that was applied but is not yet observed by the verification path.
Operational handoff
Start by exporting every existing record into the intended-state store and reviewing that baseline. Version changes so a deployment and its DNS intent can be correlated. Restrict the reconciler to the records it owns unless the intended set has explicitly been declared authoritative for the whole zone. Then test three cases: an empty diff, a missing ownership record, and a stale ownership value. The second run after each successful apply must produce an empty plan.
During a property onboarding cutover, record the intended revision before applying it, upsert the planned records, and poll through the same read path used by verification. Completion occurs when current state matches that revision, not when a write returns. Afterward, keep reconciling at the normal interval and alert on a persistent non-empty diff. This turns drift into data instead of a support ticket.
Top comments (0)