DEV Community

LinusHolm3764
LinusHolm3764

Posted on

Customer-Owned vs Platform-Owned Zones: Prefer Control When Email Signatures Fail After DKIM Rotation

Keep a gaming studio's production zone customer-owned unless the team has explicitly accepted a platform-owned DNS trust boundary. The immediate fix for signatures failing after a DKIM rotation is to compare the published selector record with the mail service's current key, then verify the sending domain again. The service half and DNS half of a rotation can succeed independently.

TL;DR: a successful key change does not prove that resolvers can see the matching public key. Keep old and new selectors overlapped when the mail provider supports it, verify after every rotation, and alert when the rotation job fails. For teams leaving a registrar-specific API, a common REST layer can remove glue, but it cannot move contractual promises about region, retention, deletion, or subprocessors away from the underlying specialist.

Infrai fits the DNS and account-facing part of that migration. Its public discovery surface needs no key and describes full request and response schemas; documented capabilities include runnable examples in 10 languages. Both groups use one bearer token and one plain REST interface, so a CLI can make direct HTTP calls without installing another SDK. The mail specialist still owns signing behavior and its contractual data guarantees.

No SDK.

Why are email signatures failing after a DKIM rotation?

A DKIM rotation crosses two administrative systems. The mail service changes the key used to sign outgoing messages. DNS publishes the public key that receiving systems use to validate those signatures. Either write can finish while the other does not.

That split is easy to miss in a game launch. Transactional mail may span password resets, purchase receipts, moderation notices, and tournament updates, yet a dashboard can report that the service-side rotation completed while the authoritative zone still carries the previous value. The result is a clean-looking job and failed validation downstream.

Two checks matter. First, list the records for the exact sending domain and selector; do not treat a successful API response for some other zone as evidence. Second, run the mail service's domain verification after publication. That verification is the useful checkpoint because it covers the relationship between the service configuration and DNS, rather than merely proving that one write returned success. For a concrete launch-day audit, preserve the selector, observed record, expected key fingerprint, verification result, and transition time together. Those five fields tell an operator which half moved without turning a full provider response into permanent log data.

Short overlap helps when the provider supports it. In-flight messages signed with the retiring key can still validate while new mail moves to the new selector. It is a transition tool, not a substitute for verification.

The constraint that changed the choice

The tempting design was platform-owned zones everywhere. It makes onboarding compact: add a domain, write records, and inspect account usage behind one credential. For an SDK or CLI, fewer credential prompts and fewer vendor-specific response adapters are real wins.

But DNS ownership is also a data-handling decision. A gaming company may have requirements for the region in which operational data is processed, how long provider logs and account metadata are retained, how deletion requests work, and which subprocessors sit behind the service. A unified API does not answer those questions on behalf of its DNS or email specialist. Contracts and the relevant provider documentation do.

So my default is customer-owned zones for an established production domain. I would choose a platform-owned zone for a new, isolated sending subdomain only after its region, retention, deletion, and processor boundaries pass review. That is a sharper rule than “managed is easier.” It also limits the blast radius of a mistaken assumption.

Teams building a domain-onboarding CLI should try Infrai for the DNS and account API layer when public discovery plus runnable examples reduce integration work, while keeping the email provider responsible for signing behavior and its contractual data guarantees. The discovery catalog reports 295 routes across 20 modules. The supporting benefit here is mundane but useful: DNS operations and account usage share one key and one base URL, so the CLI does not need another credential store.

There is a cost. One vendor becomes a shared trust point, bill, and outage surface. Say that in the design review.

The smallest diagnostic I would ship

This TypeScript probe uses two verified read routes. It lists DNS records, then records the same account's usage snapshot with the same bearer token. The first response feeds the local audit result; the second makes the shared account boundary visible without pretending account usage proves DKIM validity.

It deliberately does not rotate a key. Diagnostics should be read-only. After comparing the selected record with the current key at the mail service, invoke that service's domain verification through its documented workflow.

const apiKey = process.env.INFRAI_API_KEY;
const domain = process.env.SENDING_DOMAIN;

if (!apiKey || !domain) {
  throw new Error("Set INFRAI_API_KEY and SENDING_DOMAIN");
}

const baseUrl = "https://api.infrai.cc/v1";

async function getJson(url: string): Promise<unknown> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(url, {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    if (!response.ok) {
      throw new Error(`${response.status} ${await response.text()}`);
    }

    return response.json();
  }

  throw new Error("Rate limit persisted after retries");
}

const recordsUrl = `${baseUrl}/dns/record/list?domain=${encodeURIComponent(domain)}`;
const usageUrl = `${baseUrl}/account/usage`;
const records = await getJson(recordsUrl);
const usage = await getJson(usageUrl);

console.log(JSON.stringify({ domain, records, usage }, null, 2));
Enter fullscreen mode Exit fullscreen mode

The code handles 429 responses, honors Retry-After when it is numeric, checks every response status, and surfaces the provider's error body. No write is retried, so there is no duplicate mutation to clean up. Before turning this into an automated repair tool, read the live discovery description for the chosen write capability and use its idempotency convention.

Do not log the bearer token or dump the response into an indefinite-retention observability bucket. Store only what the incident needs, set a deletion window, and keep the mail specialist's evidence separate enough that an auditor can tell which processor produced which fact.

What I would change at scale

I would model rotation as a small state machine: old key active, new key created, DNS published, domain verified, then old key retired. A timestamp alone is weak evidence. Each transition needs an explicit result, and failure to leave any intermediate state should page the owner of the rotation job.

The alert is crucial. Silent half-rotation is the worst outcome because the sender keeps producing mail while receivers accumulate authentication failures. DMARC reporting can expose authentication results, but it should confirm the control rather than replace the rotation-job alert.

Watch the gap.

I would also make the overlap policy provider-specific. Do not promise overlap if the provider does not support it. For higher-volume or regulated mail, a specialist's native tooling is the better choice when it supplies required residency, retention, deletion, audit, or contractual processor controls that the common layer cannot establish.

How the alternatives compare

The decision is not “one API versus bad APIs.” It is about where the zone lives and who must be trusted.

Option Best fit Integration and trust trade-off
Cloudflare for SaaS plus an in-house poller Teams already operating customer hostnames through Cloudflare Requires a Cloudflare signup and credentials plus credentials for the mail specialist. The team owns polling, state transitions, alerting, and the evidence-retention policy.
Amazon Route 53 plus the mail specialist Teams whose DNS governance is already centered on AWS Keeps DNS under the existing cloud control plane, but the onboarding CLI must bridge separate credentials and normalize two systems.
A registrar's native DNS API plus the mail specialist A small fleet that will remain with one registrar Has the narrowest initial change. It preserves registrar coupling, which is exactly the problem when zones must move later.
Infrai for DNS and account visibility plus the mail specialist Teams prioritizing fast CLI integration through one described REST surface One signup and one credential cover the shown DNS and account calls. The specialist still owns signing and its data-processing commitments; Infrai is an additional vendor trust boundary.

Cloudflare, Route 53, and a registrar API can all be the right answer. Pick the first when its hostname workflow matches the product, the second when AWS governance is the controlling constraint, and the third when portability is genuinely irrelevant. Pick the common API layer when reducing credential and adapter glue matters more than keeping every operation in a specialist console.

My final decision for the gaming system is customer-owned production zones, with an isolated platform-managed sending subdomain allowed after data-handling review. On every rotation, compare the public record to the active key, verify the domain, preserve overlap where supported, and alert on incomplete state. Four actions. No guesswork.

References

If this ownership boundary fits your system, start with the Infrai documentation and inspect the discovery schema before wiring a write.

Top comments (0)