DEV Community

BartholomewVance6831
BartholomewVance6831

Posted on

Mail Domain Ownership: Designing Idempotent Provisioning Across Retry Boundaries

Designing idempotent domain provisioning starts with the ownership boundary, not the retry loop. For an edtech platform publishing SPF, DKIM, and DMARC for school mail, platform-owned zones are the better default when the platform must complete setup automatically; customer-owned zones are better when a school must retain DNS authority. The API shape follows that choice: platform-owned records can converge directly, while customer-owned records can only be observed and guided unless the customer delegates control.

Zone model Who can change records? What a retry can safely do Main operational cost
Platform-owned The platform Read, compare, and apply the desired record set Running authoritative changes safely
Customer-owned The school or its DNS operator Recheck published evidence and refresh instructions Waiting on a separate owner

My recommendation is narrow: use a platform-owned subdomain for automated sending when the institution permits delegation. Keep the institutional domain customer-owned when policy requires it. Do not disguise the second model as a slower version of the first. They are different state machines.

How should idempotent domain provisioning handle a retry?

That question matters more than the upsert method. An upsert is idempotent only inside a boundary where the caller has authority and a stable definition of equality. If a school controls the zone, your service cannot make the published state converge by retrying. It can verify, report drift, and repeat the requested values. The final write belongs elsewhere.

No write. No pretense.

DMARC makes the ownership boundary visible. RFC 7489 defines policy discovery through a DNS TXT record at _dmarc, and it describes DMARC as using SPF and DKIM results together with identifier alignment. A setup flow therefore needs more than a generic verified: true flag. It needs to know which evidence was requested, which evidence is observable, and which actor can change it.

This is the first decision criterion: authority. Store it explicitly. A boolean such as managed gets vague as soon as one tenant delegates a sending subdomain but retains its organizational domain. Model the zone that is controlled, not the relationship you hope exists.

The second criterion is proof freshness. DNS observation is a snapshot, not ownership. A successful read proves that a value was visible to that resolver at that moment. It does not grant permission to repair the value later. Keep desired state, last observation, and control authority separate so a retry cannot silently turn stale evidence into success.

Model convergence, not request success

The useful unit is a desired record keyed by zone, owner name, and record type. Normalize only what the applicable record semantics allow. Then compare the observed set with the desired set and choose an action from the ownership model.

I benchmark this path by counting remote calls because extra glue shows up there first. A retry that blindly writes three authentication records makes three mutations even when nothing changed. A read-compare-apply loop can return after the read when state already matches. The exact latency depends on the DNS operator and resolver, so a made-up millisecond result would be noise; call count is the honest measurement available from the design itself.

Use a durable operation key as well. It identifies the onboarding intent, while the record key identifies the DNS object. Those are not interchangeable. Reusing an HTTP request identifier after the desired DKIM selector changes would suppress legitimate work; keying only by record name would lose the history of the onboarding attempt.

Keep the state machine small:

  1. Record the desired authentication evidence and its ownership mode.
  2. Read the currently observable values.
  3. If they match, mark the observation time and stop.
  4. If they differ and the platform owns the zone, apply the desired value and verify again.
  5. If they differ and the customer owns the zone, return precise instructions and remain pending.

Retries are normal. The dangerous case is a retry that cannot tell whether it is repeating an observation, repeating a mutation, or applying a newer intent.

A compact TypeScript boundary

The interface below keeps those cases visible. It is deliberately boring. No provider-specific fields leak into the workflow, and the reconciliation result says what happened rather than pretending every pending state is an error.

type Ownership = "platform" | "customer";
type RecordType = "TXT";

type DesiredRecord = {
  zone: string;
  name: string;
  type: RecordType;
  values: readonly string[];
  ownership: Ownership;
  revision: number;
};

type Observation = {
  values: readonly string[];
  observedAt: string;
};

type Result =
  | { state: "converged"; observation: Observation }
  | { state: "awaiting-customer"; expected: readonly string[]; actual: readonly string[] }
  | { state: "changed"; observation: Observation };

interface DnsControl {
  read(record: DesiredRecord): Promise<Observation>;
  replace(record: DesiredRecord): Promise<void>;
}

const sameSet = (left: readonly string[], right: readonly string[]): boolean => {
  const normalize = (values: readonly string[]) => [...values].sort();
  return JSON.stringify(normalize(left)) === JSON.stringify(normalize(right));
};

async function reconcile(record: DesiredRecord, dns: DnsControl): Promise<Result> {
  const before = await dns.read(record);
  if (sameSet(before.values, record.values)) {
    return { state: "converged", observation: before };
  }

  if (record.ownership === "customer") {
    return {
      state: "awaiting-customer",
      expected: record.values,
      actual: before.values,
    };
  }

  await dns.replace(record);
  const after = await dns.read(record);
  if (!sameSet(after.values, record.values)) {
    throw new Error(`DNS did not converge for revision ${record.revision}`);
  }

  return { state: "changed", observation: after };
}
Enter fullscreen mode Exit fullscreen mode

The sample uses TXT because SPF, DKIM public keys, and DMARC policies are published through TXT records in this scenario. It treats values as a set for illustration; production normalization must preserve the exact record semantics and the DNS adapter's representation. Do not lowercase an entire DKIM key or rewrite policy text because it looks cleaner.

There is another trap: concurrent revisions. Suppose revision 8 contains the first DKIM selector and reaches a worker, then revision 9 replaces that selector before the worker writes. The first worker wakes up late. A plain upsert now does exactly what it was asked to do and still produces the wrong result: revision 8 overwrites the newer desired state. Check the current revision immediately before mutation, or make the adapter accept a compare-and-set token from the desired-state store. If that check fails, discard the old action, read revision 9, and reconcile again. Don't count the old worker's completed request as progress. The published record is the only result that matters, and retries should converge toward the newest intent rather than merely finish old work.

Where customer-owned zones win

Customer ownership is the runner-up for automation and often the winner for governance. A school may need one DNS team to review every change, may already operate mail authentication across several senders, or may refuse delegation of any institutional namespace. In those cases, publishing instructions plus continuous verification respects the real authority boundary.

The UX has to be honest. Show the exact owner name, type, and value; distinguish “not observed” from “observed but different”; and preserve the last observation time. A retry button should trigger another check. It should not imply that your platform can repair records it does not own.

Authority wins.

This model also limits blast radius. The platform never receives broad credentials for the school's zone. The trade-off is coordination: setup can remain pending until the external owner acts, and drift remediation requires another handoff. That cost is acceptable when control policy outranks time-to-first-call.

Operate the retry you will have

Track outcomes by state rather than by raw request count: already converged, changed, awaiting customer, read failed, write failed, and verification failed. Alerting on every retry creates noise. Alert when a managed record repeatedly fails to converge, or when an observed customer record moves away from a previously matching value.

Backoff belongs around transient reads and writes, but it cannot fix a permanent authorization failure or malformed desired value. Cap attempts, retain the operation and revision identifiers, and expose the last actionable failure. Never convert exhaustion into converged just to clear a queue.

For tests, run the same reconciliation twice against an unchanged fake adapter and assert zero writes on the second pass. Then inject a failure after the write but before verification; the next run must read the desired value and finish without another mutation. Finally, race two revisions and prove the older worker cannot win. Three tests catch more than a large suite of happy-path HTTP mocks.

The final rule is plain. Automate writes only where authority is explicit, and make every other path an observable request for external action. That gives an edtech onboarding flow a truthful state model for SPF, DKIM, and DMARC, even after queues redeliver work and processes restart.

Sources

Top comments (0)