Short answer: use an idempotent upsert and a read-before-write reconciliation loop so a retried onboarding job converges on one tenant DNS record; reserve cleanup for duplicates with proven workflow ownership.
A retried tenant-onboarding job should converge on one intended DNS state, not create another copy of it. For a customer-support platform assigning each tenant a subdomain, treat provisioning as a reconciliation loop: identify the desired records, normalize what the DNS control plane returns, upsert the one canonical record, and alert on evidence that the published set has drifted. Cleanup is a repair path, not the normal success path.
The page usually arrives too late. An on-call engineer sees a tenant whose support-mail domain has conflicting TXT values or whose expected verification record is absent after a retry. The immediate temptation is to delete every matching record and run onboarding again. That can remove a deliberate record along with the duplicate. It also hides the cause: the original create operation was not safe to repeat.
For deliverability, the useful question is narrower than "did the job return 200?": does the resolved DNS answer contain exactly the evidence the mail policy expects for this tenant? RFC 7489 defines a DMARC record as a DNS TXT record at _dmarc under the domain being evaluated. A control-plane success response cannot substitute for that observation.
Start from the page, then define the expected answer
Make the alert describe the published state and the tenant impact. tenant_acme_dmarc_expected_but_missing is actionable; dns_api_retry_failed is an implementation detail. The first says what the responder must verify. The second sends them spelunking through logs while a customer-support address may still lack the policy evidence the workflow was meant to establish.
I use a compact desired-state key made from the tenant, record owner, type, and purpose. The target value belongs in the comparison, but not in the identity key, because a policy update should replace the old target rather than manufacture a second record. A 30-minute reconciliation interval gives the system a bounded repair window; it is a deliberate availability-versus-control-plane-load trade-off, not a promise about DNS propagation.
| Signal | What it compares | Page condition |
|---|---|---|
| Desired-state count | Canonical records in the store | More than one active record for one identity |
| Control-plane count | Returned records after normalization | Zero or more than one canonical match |
| Published evidence | Resolver answer versus expected value | Expected DMARC TXT value is missing |
These are three different failures. Collapsing them into one "DNS failed" metric makes duplicate creation look the same as delayed visibility, which produces poor remediation.
How should retried onboarding handle duplicate DNS records?
Because the first request may have reached the DNS service even when the worker never observed a usable response. A timeout, worker restart, or queue redelivery leaves the caller unable to distinguish "nothing happened" from "the record was created." Retrying a blind create therefore turns an uncertain write into a duplicate-risk write.
Consider the shape of the race. At 09:00:00, worker A submits the record create. At 09:00:03, the remote system has accepted it but the response is lost. At 09:00:30, the queue redelivers the same onboarding message to worker B, which has none of A's in-memory context. A second create can succeed before either worker sees the first record. The right response to this ambiguous timeline is not a wider retry budget; it is durable desired state, a stable idempotency key, and a fresh remote read before a mutation. I have been paged by missed jobs and duplicate deliveries, so I want that evidence attached to the page instead of buried in a worker log.
Don't call create blind.
That uncertainty must survive in the design. A job key alone is insufficient if a new worker can issue the same create after losing process memory. Persist an idempotency key with the desired state, and run a lookup before any mutation. The lookup needs normalization because DNS names are case-insensitive and a trailing dot is representation, not a second owner name.
package dnsstate
import (
"strings"
)
type Record struct {
ID string
Owner string
Type string
Value string
Purpose string
}
func canonicalOwner(owner string) string {
return strings.ToLower(strings.TrimSuffix(strings.TrimSpace(owner), "."))
}
func matching(records []Record, desired Record) []Record {
var found []Record
for _, record := range records {
if canonicalOwner(record.Owner) == canonicalOwner(desired.Owner) &&
strings.EqualFold(record.Type, desired.Type) &&
record.Purpose == desired.Purpose {
found = append(found, record)
}
}
return found
}
This is where a surprising amount of damage happens. If the provider-facing record model has no place for Purpose, store the association in your own desired-state database and treat remote records as candidates. Do not infer ownership from a value prefix that an administrator might legitimately reuse.
What should the reconciliation worker do?
The worker should make the smallest safe change after it has read both desired state and remote state. It should not issue a create because the last attempt is unknown. When there are no candidates, create one record and save the returned remote identifier. When there is one candidate with the wrong value, update that identifier. When there are several candidates, quarantine the identity for review or apply a narrowly scoped duplicate policy that can prove which records were created by this workflow.
Here is the decision point in Go. The interface is intentionally generic: it keeps the operational rule separate from a particular DNS provider.
func reconcile(client DNSClient, desired Record) error {
candidates, err := client.Find(desired.Owner, desired.Type)
if err != nil {
return err
}
matches := matching(candidates, desired)
switch len(matches) {
case 0:
return client.Create(desired)
case 1:
if matches[0].Value == desired.Value {
return nil
}
return client.Update(matches[0].ID, desired)
default:
return DuplicateRecordError{Owner: desired.Owner, Count: len(matches)}
}
}
Notice the refusal to delete in the last branch. Automatic deletion becomes defensible only when the system has a durable ownership marker and an audit trail that identifies each duplicate as its own prior write. Even then, deleting records changes customer-visible state; a review queue is often the better first deployment.
The retry policy belongs around this function, with the same idempotency key passed through each attempt. Queue delivery is at-least-once in practice from the consumer's perspective, so the handler must tolerate being invoked twice. Make that property visible in a test: run reconcile twice against a fake control plane, then assert one canonical remote record and one intended value.
Instrument the earlier signal, not only the failure
A useful trace links the page back to the decision that caused it: tenant ID, job attempt, idempotency key, desired-state revision, remote record IDs examined, and the final reconciliation branch. Avoid logging TXT contents if those values carry tenant data; record a stable digest instead. The responder needs correlation, not a second copy of the policy.
Track reconciliation outcomes as counters and observe age separately. A count of duplicate candidates detects a bad write pattern. An age gauge for desired_state_created_at to published_evidence_observed_at tells you that a tenant is waiting without pretending every wait is a duplicate. For DMARC, query the expected _dmarc owner and compare the relevant TXT result to the desired value; RFC 7489 is the standard reference for the record location and evaluation model.
The runbook can stay short:
- Confirm the desired identity and its current revision.
- List normalized remote candidates and inspect ownership evidence.
- Compare the resolver result with the intended record before changing anything.
- Retry reconciliation with the original idempotency key after a transient control-plane error.
A page that says "two candidates, no safe ownership proof" is a good page. It prevents an automatic cleanup from deleting a record the platform does not own.
Set thresholds with the false-positive bill in mind
Thresholds define human work. If every resolver delay wakes someone immediately, responders learn to ignore the alert and genuine duplicate drift sits in the same queue. If the threshold is too long, a tenant can remain in a partial onboarding state long enough to affect support-mail trust.
Use two windows rather than one magic number: a shorter warning for delayed published evidence, and a higher-severity page only when the delay exceeds the workflow's documented expectation or duplicate candidates appear with positive ownership evidence. The warning may close naturally after the next reconciliation pass. The page should carry the candidate count and last successful observation so an engineer can decide without reconstructing the timeline.
The operational goal is boring convergence. One tenant identity maps to one managed record, retries lead back to that state, and a resolver check supplies the evidence that users actually depend on. That design is easier to test, safer to repair, and much less likely to turn an ambiguous timeout into a permanent DNS mess.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.