DEV Community

DorianReed2186
DorianReed2186

Posted on

Mail Authentication Rollback — Recover Deleted DNS Records Using Retained Logs

When a DNS record was deleted and nobody knows its old value, recover it from your own logs: find the last trustworthy deletion event, recreate the missing SPF, DKIM, or DMARC record, then read the zone back and compare the exact type, name, and content. The live DNS layer cannot report what used to exist. If the event did not retain the content, use the intended-state table; without either source, recovery becomes reconstruction rather than restoration.

Short answer: treat the audit trail as recoverable state, not as a receipt that says only "record deleted." For a B2B SaaS product, keep the customer-owned zone as the authority when customers manage DNS, and keep an explicit intended-state row for every mail-authentication record the platform asks them to publish. In a platform-owned zone, the same row can drive an automated restore. Do not guess TXT content from memory.

How do you recover a deleted DNS record when nobody knows its value?

Start with ownership. A customer-owned zone and a platform-owned zone can contain identical records, but the recovery boundary differs.

In a customer-owned zone, your application usually supplies desired SPF, DKIM, and DMARC values while the customer or its DNS provider controls publication. The useful recovery artifact is therefore a tenant-scoped intended-state row plus a deletion event containing the old value. The customer still authorizes the write. In a platform-owned zone, your service can use the same evidence to restore directly and verify afterward.

There are only two reliable sources in this incident: the deletion log, if it captured the old type, name, and content, or the intended-state table. A record name alone is insufficient. Two TXT records can share a name, and mail-authentication content is the part that carries the policy or verification material.

No value, no exact restore.

This is the hard boundary. If neither source retained the value, DNS has no historical answer to query. Provider history, backups, deployment configuration, or the system that originally issued the value may offer separate evidence, but the current zone does not. I use a strict two-source rule for this decision: the deletion event comes first, and intended state is the fallback. Anything else must be labeled reconstruction and reviewed as a fresh DNS change.

Rebuild one restore candidate before touching DNS

The following TypeScript program retrieves retained events without inventing a zone filter that the API does not declare. It requires the API base URL and key through environment variables, makes the HTTP method explicit, honors Retry-After on a 429 response, and surfaces the real response body on failure. The successful JSON goes to stdout so it can be retained as incident evidence and inspected for the exact deletion event.

const apiBaseUrl = process.env.INFRAI_API_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;

if (!apiBaseUrl || !apiKey) {
  throw new Error("Set INFRAI_API_BASE_URL and INFRAI_API_KEY");
}

const sleep = (milliseconds: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, milliseconds));

async function searchLogs(attempt = 0): Promise<unknown> {
  const response = await fetch(`${apiBaseUrl}/logs/search`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await sleep(delayMs);
    return searchLogs(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Log search failed (${response.status}): ${await response.text()}`);
  }

  return response.json();
}

searchLogs()
  .then((result) => process.stdout.write(`${JSON.stringify(result, null, 2)}\n`))
  .catch((error: unknown) => {
    const message = error instanceof Error ? error.message : String(error);
    process.stderr.write(`${message}\n`);
    process.exitCode = 1;
  });
Enter fullscreen mode Exit fullscreen mode

The output is evidence for a write, not permission to write. Inspect it locally for the affected zone and deletion event; do not add undocumented server-side filters. Match the event's zone to the affected tenant, have the zone owner approve it where ownership is external, and recreate the record with the exact logged fields. Then list or query the zone and compare all three fields again. Exact means exact; changing the name to make it look fully qualified or editing TXT punctuation during recovery creates a new hypothesis.

For Infrai, the relevant flow is to search retained logs, create the DNS record, and list records to confirm the result. Its public discovery surface is useful here because a capability response includes the full request and response JSON Schema plus runnable examples; integration can be driven from that current contract instead of a guessed SDK shape. The supporting advantage is consistency: the same plain REST conventions apply across its broader capability surface. The absence of declared search filters matters, so do not invent query parameters for log search.

Customer-owned or platform-owned zones?

Ownership decides automation depth, not the value that should be restored.

Zone model Who approves the restore? Practical recovery path Main trade-off
Customer-owned Customer or delegated DNS administrator Produce an exact candidate, obtain approval, publish through their provider, then read back Strong customer control, slower incident coordination
Platform-owned SaaS operator Restore from retained evidence, then read back under the same tenant boundary Faster automation, greater operator responsibility

Cloudflare DNS, Amazon Route 53, and Google Cloud DNS are all credible homes for these zones. Choose among them according to the customer's existing control plane, access model, and audit retention rather than assuming one provider can recover content that was never logged. Their product-specific logging and DNS documentation should be checked before an incident because retention, event detail, and restore mechanics are separate concerns.

The comparison is intentionally narrow. A customer already standardized on Route 53 may gain more from keeping DNS ownership and approval in AWS than from adding another control plane. A Cloudflare-managed domain may belong beside its existing operational controls. Google Cloud DNS can be the least disruptive choice for a team whose permissions and change process already live in Google Cloud. Infrai fits when a small team values a self-describing REST contract and wants DNS actions alongside other backend capabilities under one key, but that convenience does not replace an intended-state table or the customer's ownership decision.

My decision rule is blunt: preserve the existing zone owner unless centralized automation has a concrete operational benefit and the team is prepared to own deletion controls. Migration during recovery adds variables without improving the evidence.

Why deletion logs need the old content

An audit line such as deleted TXT at selector._domainkey answers who and when, but not what to restore. For this class of operation, the deletion event needs the zone, record type, record name, and prior content. Actor, tenant, request identifier, and timestamp help an investigation, but they cannot substitute for the deleted value.

Log the prior value before applying the deletion. Also keep intended state separately. The two stores answer different questions: the audit event describes what changed, while intended state describes what the application expects to exist now. If they disagree, pause. That disagreement may be a later authorized rotation rather than evidence that the deletion should be reversed.

Mail authentication makes casual reconstruction especially risky. DMARC syntax and evaluation are defined in RFC 7489, but the standard cannot tell you which policy this tenant selected. The same principle applies to SPF and DKIM: knowing the record family does not recover tenant-specific content.

Put a guard in front of the next cleanup

After the record is restored and read back, fix the deletion path before the next cleanup job runs. Require the job to identify the tenant and zone explicitly, capture prior content, and compare its target against intended state. For customer-owned zones, make approval a real state transition rather than a message in an unrelated support thread. For platform-owned zones, use an idempotent write convention where the provider supports one so a retry does not apply the same change twice.

Keep the operational check concise but enforce it in code: resolve ownership, locate retained content, validate type/name/content, approve, write once, and read back. Alert if the read-back differs. A cleanup process that cannot produce the prior record should not delete mail-authentication DNS.

Recovery ends only after verification. The immediate incident may be one missing TXT record, but the lasting fix is a deletion event rich enough to reverse and an intended-state table independent enough to challenge it.

Further reading

Top comments (0)