Short answer: reconcile normalized tenant intent against a fresh authoritative-zone inventory on a schedule, but alert only after the same difference survives two runs. That split catches real DNS drift without turning propagation time, disabled tenants, or unrelated records into pager noise.
For a B2B SaaS application that gives every tenant a subdomain, the database says what should exist and the zone says what users can resolve. Neither side alone is truth. The useful result is the difference between them, labeled with enough context to decide whether to create a record, remove one, or investigate an ownership collision.
The data flow is deliberately boring: export enabled tenant hostnames, fetch the current zone inventory through an adapter, normalize both into comparable keys, classify the set differences, and persist a small confirmation state. Keep provider calls outside the comparison function. That makes the risky part easy to exercise in a notebook and the production job easy to evaluate with fixtures.
What should scheduled drift detection compare in a tenant table and live DNS zone list?
Compare intent, not raw rows. A tenant row usually carries lifecycle state that a plain hostname set loses: active, provisioning, suspended, or deleted. The desired set should include only states whose subdomains are supposed to resolve. The observed set should include only record types owned by this automation and only names below the tenant label you control.
That scope matters. A zone can also contain mail and verification records. DMARC, for example, publishes policy as a DNS TXT record at a specific _dmarc name; RFC 7489 describes that record and its discovery rules. A reconciliation job that treats every TXT record as tenant inventory will manufacture drift. Filter by record type and managed namespace before comparing anything.
Use fully qualified, lowercase names with one consistent trailing-dot policy. Also reject tenant labels that would escape the managed suffix. DNS libraries and provider exports don't always display names in the same form, so normalization belongs at the input boundary rather than scattered through the diff.
The core classifications are small:
| Classification | Meaning | Initial action |
|---|---|---|
| Missing | Enabled tenant exists, managed DNS name does not | Confirm, then queue creation |
| Orphaned | Managed DNS name exists, eligible tenant does not | Confirm ownership and lifecycle before removal |
| Mismatched | Name exists with an unexpected target or type | Escalate with expected and observed values |
| In sync | Intent and published record agree | Record success, no alert |
Don't collapse missing and mismatched into one boolean. A missing name may be ordinary provisioning lag; an unexpected target can indicate that another workflow owns the same label. Those need different runbooks.
Build the runnable reconciliation core
This example uses only the Python standard library. It reads two JSON snapshots: one exported from the tenant table and one freshly fetched from the authoritative zone API by a separate adapter. The separation is useful because credentials, pagination, and provider response shapes change; set comparison should not.
from __future__ import annotations
import json
from dataclasses import dataclass
from pathlib import Path
from typing import Iterable
MANAGED_SUFFIX = "customers.example.com"
MANAGED_TYPES = {"A", "AAAA", "CNAME"}
ELIGIBLE_STATES = {"active", "provisioning"}
@dataclass(frozen=True)
class DesiredRecord:
tenant_id: str
name: str
record_type: str
value: str
@dataclass(frozen=True)
class ObservedRecord:
name: str
record_type: str
value: str
def fqdn(value: str) -> str:
return value.strip().lower().rstrip(".")
def tenant_name(label: str) -> str:
clean = label.strip().lower()
if not clean or "." in clean:
raise ValueError(f"invalid tenant label: {label!r}")
return f"{clean}.{MANAGED_SUFFIX}"
def desired_records(rows: Iterable[dict]) -> dict[str, DesiredRecord]:
desired: dict[str, DesiredRecord] = {}
for row in rows:
if row["state"] not in ELIGIBLE_STATES:
continue
record = DesiredRecord(
tenant_id=str(row["tenant_id"]),
name=tenant_name(row["subdomain"]),
record_type=row["record_type"].upper(),
value=fqdn(row["target"]),
)
if record.name in desired:
raise ValueError(f"duplicate desired name: {record.name}")
desired[record.name] = record
return desired
def observed_records(rows: Iterable[dict]) -> dict[str, ObservedRecord]:
observed: dict[str, ObservedRecord] = {}
suffix = f".{MANAGED_SUFFIX}"
for row in rows:
name = fqdn(row["name"])
record_type = row["type"].upper()
if not name.endswith(suffix) or record_type not in MANAGED_TYPES:
continue
record = ObservedRecord(
name=name,
record_type=record_type,
value=fqdn(row["value"]),
)
if name in observed:
raise ValueError(f"multiple managed records for: {name}")
observed[name] = record
return observed
def reconcile(tenants: Iterable[dict], zone: Iterable[dict]) -> list[dict]:
desired = desired_records(tenants)
observed = observed_records(zone)
findings: list[dict] = []
for name in sorted(desired.keys() - observed.keys()):
record = desired[name]
findings.append({
"kind": "missing",
"name": name,
"tenant_id": record.tenant_id,
"expected": [record.record_type, record.value],
})
for name in sorted(observed.keys() - desired.keys()):
record = observed[name]
findings.append({
"kind": "orphaned",
"name": name,
"observed": [record.record_type, record.value],
})
for name in sorted(desired.keys() & observed.keys()):
expected = desired[name]
actual = observed[name]
if (expected.record_type, expected.value) != (
actual.record_type,
actual.value,
):
findings.append({
"kind": "mismatched",
"name": name,
"tenant_id": expected.tenant_id,
"expected": [expected.record_type, expected.value],
"observed": [actual.record_type, actual.value],
})
return findings
def load_json(path: str) -> list[dict]:
data = json.loads(Path(path).read_text(encoding="utf-8"))
if not isinstance(data, list):
raise ValueError(f"expected a JSON array in {path}")
return data
if __name__ == "__main__":
tenants = load_json("tenants.json")
zone = load_json("zone_snapshot.json")
print(json.dumps(reconcile(tenants, zone), indent=2, sort_keys=True))
Try it with a fixture before connecting the adapter. This tiny dataset includes one valid record, one missing record, one wrong target, and one orphan. It also includes _dmarc; the filter correctly ignores it because it is outside the managed tenant record types.
from pprint import pprint
tenants = [
{"tenant_id": 101, "subdomain": "acme", "state": "active",
"record_type": "CNAME", "target": "edge.example.net"},
{"tenant_id": 102, "subdomain": "bravo", "state": "active",
"record_type": "CNAME", "target": "edge.example.net"},
{"tenant_id": 103, "subdomain": "cobalt", "state": "active",
"record_type": "CNAME", "target": "edge.example.net"},
{"tenant_id": 104, "subdomain": "delta", "state": "suspended",
"record_type": "CNAME", "target": "edge.example.net"},
]
zone = [
{"name": "acme.customers.example.com.", "type": "CNAME",
"value": "edge.example.net."},
{"name": "cobalt.customers.example.com", "type": "CNAME",
"value": "old-edge.example.net"},
{"name": "former.customers.example.com", "type": "CNAME",
"value": "edge.example.net"},
{"name": "_dmarc.example.com", "type": "TXT",
"value": "v=DMARC1; p=none"},
]
pprint(reconcile(tenants, zone))
The expected findings are bravo as missing, cobalt as mismatched, and former as orphaned. acme is quiet. That negative assertion is important: an eval harness should prove that correct and unrelated records do not generate findings, not merely that planted failures do.
Confirm drift without hiding it
One clean diff is not yet an alert. The database export and live zone fetch are two observations taken at different instants, so a tenant created between them can appear missing even when the publishing workflow is healthy. Fetch both as close together as practical, attach timestamps, and require an identical finding on two consecutive runs before paging or opening a ticket.
Here is a compact confirmation layer. It fingerprints the evidence, persists the previous run, and emits only repeated findings. A changed expected target creates a new fingerprint, which avoids confirming stale evidence.
import hashlib
import json
from pathlib import Path
def fingerprint(finding: dict) -> str:
encoded = json.dumps(finding, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(encoded.encode("utf-8")).hexdigest()
def confirmed_findings(findings: list[dict], state_path: str) -> list[dict]:
path = Path(state_path)
previous = set()
if path.exists():
previous = set(json.loads(path.read_text(encoding="utf-8")))
current = {fingerprint(item) for item in findings}
confirmed = [item for item in findings if fingerprint(item) in previous]
temporary = path.with_suffix(".tmp")
temporary.write_text(json.dumps(sorted(current)), encoding="utf-8")
temporary.replace(path)
return confirmed
Two runs are an example policy, not a DNS guarantee. Choose the interval and threshold from the application's provisioning objective, then test that policy against recorded workflow timings. I'm not sure a single default can serve both an interactive tenant signup flow and a nightly back-office import; the evidence that resolves that choice is your own time from database commit to authoritative publication.
Keep detection read-only.
Automatic repair looks attractive, but an orphan can mean a delayed database restore, a manually delegated customer name, or an incomplete offboarding transaction. A detector that deletes first destroys evidence. Queue proposed changes behind ownership checks and an auditable approval boundary. For a tightly controlled namespace, auto-creating a repeatedly missing record may be reasonable; auto-removal deserves a higher bar.
Evaluate the job like application code
DNS inventory reconciliation is cheap in CPU terms, yet careless scheduling can create avoidable API calls and noisy logs. Fetch each zone once per run, handle pagination in the adapter, and compare locally. Don't query one name per tenant. Imagine a table with 10,000 eligible tenants: the per-name design asks for 10,000 remote observations, while a paginated inventory lets the adapter retrieve the zone in bounded pages and lets the reconciliation core do the remaining work in memory. More troubling than the call count, the first hostname and the last hostname no longer describe one useful observation window. A deployment occurring during that walk can leave the report with a mixture of old and new state. The zone-list design has an explicit snapshot timestamp, so the report can say exactly which observation it evaluated. It also makes prompt and log costs easier to control if findings later feed an AI-assisted triage step: send the small classified difference, never the entire raw zone.
Noise compounds.
The eval set should cover casing, trailing dots, duplicate desired names, multiple observed records, suspended tenants, names outside the managed suffix, unrelated TXT records, missing names, wrong types, and wrong values. Add a regression fixture whenever an operational review discovers a new ambiguity. I favor fixtures over mocks here because the provider adapter can save a sanitized response, while the pure comparison core consumes the same stable JSON shape in a notebook and in the scheduled worker.
Watch four signals: snapshot age, reconciliation duration, finding count by kind, and confirmation age. A zero finding count with an old or empty snapshot is not success. Fail the run when either input is unavailable or malformed, and do not overwrite the prior confirmation state after a failed fetch; otherwise the next healthy run loses its consecutive-run evidence.
There is a catch. Full-zone listing is not suitable when the service account cannot be granted inventory visibility, when the zone is too large for the job's allowed window, or when another team controls record ownership without machine-readable labels. In those cases, stick with a narrower delegated tenant zone, an append-only publication ledger, or event-driven verification of records this application created. Scheduled full inventory remains valuable as a backstop, but it should not force a broader security boundary than the organization accepts.
Put the reconciliation worker into production
Schedule one instance per zone or use a lease so overlapping runs cannot race on confirmation state. Give every run an ID, retain the two source timestamps, and include the tenant ID plus expected and observed tuples in structured output. Redact unrelated record values from logs; the detector only needs evidence within its managed scope.
Deployment is complete when a fixture run, a read-only production run, and an alert-routing test all pass. Start with reporting only. After the team has classified several real findings, automate just the remediation whose ownership rule is unambiguous, and leave mismatches plus removals for review. This progression keeps the comparison reproducible and the blast radius small.
The operational checklist fits in prose: verify that the tenant export and zone snapshot are fresh, confirm the namespace and lifecycle filters, run the pure diff, persist evidence only after successful inputs, emit repeated findings with timestamps, and track them to resolution. Re-run a clean snapshot afterward. Done.
Top comments (0)