DEV Community

YatesHolloway6872
YatesHolloway6872

Posted on

Implementing 3 Custom Domains as a Product Feature Safely

A healthtech SaaS taking on custom domains as a product feature is really taking on an asynchronous onboarding system. SPF, DKIM, and DMARC live outside the app's transaction boundary, so my choice is to model all three as observed checks and let the latest verification result drive the UI.

Short answer: the DNS calls are the small part. The product is the state machine around pending propagation, incorrect records, and eventual verification. Ship that state machine before polishing the setup wizard.

What are you really taking on with custom domains?

Saving care.example proves only that the application accepted a string. It doesn't prove that the domain publishes the expected records. A customer might paste a value into the wrong zone, omit a character, or wait while a resolver still sees an older answer. Treating Save as success turns uncertainty into a misleading green check. That is what custom domains as a product feature really means: your application now has to explain an external system without pretending to control it.

This matters more than usual for healthcare mail. The operational question is deliverability evidence: can the product show what it expected, what it observed, and when it checked? DMARC builds on SPF and DKIM identifiers and adds policy and reporting. It deserves its own status rather than one generic dnsConfigured boolean.

The awkward support boundary is part of the design. The person editing DNS may be an agency, an IT contractor, or a customer's security team rather than the person using the SaaS. Budget for those conversations.

They'll arrive.

Build the smallest honest state machine

I ship weekly, so I chose three states because each one changes the next support action. Start by inspecting the live capability description, then keep desired configuration separate from observed evidence. The discovery surface is public, but this sample still reads the platform key from the environment to demonstrate the same credential discipline used for an authenticated call. Set INFRAI_BASE_URL to the documented v1 base, set INFRAI_API_KEY, and run the file with npx tsx domain-readiness.ts.

type CheckName = "spf" | "dkim" | "dmarc";
type CheckState = "pending" | "verified" | "failed";

type Evidence = {
  state: CheckState;
  expected: string;
  observed: string[];
  checkedAt: string | null;
  message: string;
};

type DomainReadiness = Record<CheckName, Evidence>;

const baseUrl = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;

if (!baseUrl || !apiKey) {
  throw new Error("Set INFRAI_BASE_URL and INFRAI_API_KEY");
}

async function discoverVerify(attempt = 0): Promise<unknown> {
  const response = await fetch(`${baseUrl}/discovery/dns.domain.verify`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` }
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1000
      : 250 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return discoverVerify(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
  }

  return response.json();
}

const pending = (expected: string): Evidence => ({
  state: "pending",
  expected,
  observed: [],
  checkedAt: null,
  message: "Waiting for DNS verification"
});

const domain: DomainReadiness = {
  spf: pending("v=spf1 include:mailer.example -all"),
  dkim: pending("selector._domainkey.care.example"),
  dmarc: pending("v=DMARC1; p=none; rua=mailto:dmarc@care.example")
};

function applyObservation(
  current: DomainReadiness,
  name: CheckName,
  observed: string[],
  checkedAt: string
): DomainReadiness {
  const expected = current[name].expected;
  const verified = observed.includes(expected);

  return {
    ...current,
    [name]: {
      expected,
      observed,
      checkedAt,
      state: verified ? "verified" : "failed",
      message: verified
        ? "DNS answer matches the expected value"
        : "DNS answered, but the expected value was not present"
    }
  };
}

async function main(): Promise<void> {
  const capability = await discoverVerify();
  const afterSpf = applyObservation(
    domain,
    "spf",
    ["v=spf1 include:mailer.example -all"],
    "2026-09-21T08:00:00Z"
  );
  const ready = Object.values(afterSpf).every(
    (check) => check.state === "verified"
  );

  console.log(JSON.stringify({ capability, ready, checks: afterSpf }, null, 2));
}

await main();
Enter fullscreen mode Exit fullscreen mode

Keep pending distinct from failed. Pending says the system hasn't yet established the answer; failed says it obtained evidence that doesn't match. That distinction gives support a useful next action and stops the interface from blaming a customer while propagation is unresolved. The trade-off is extra UI copy and one more branch in support tooling. I accept it because collapsing the states destroys the evidence needed to decide between waiting and correcting a record.

Don't cache hope.

The final ready value is derived. Never store it as an optimistic flag. A later verification can change an individual check, and the aggregate follows without two sources of truth.

Choose the control boundary deliberately

There are several legitimate ways to own this workflow. They solve different parts of it.

Option Useful boundary Trade-off for a solo SaaS
Cloudflare DNS Manage records in zones delegated to Cloudflare Strong fit when the application controls the zone; customer-owned zones still require coordination
Amazon Route 53 Manage hosted zones and records inside AWS Natural for an AWS estate, but it does not remove the external onboarding state machine
Vercel Domains Attach domains to applications deployed on Vercel Convenient for web serving; mail-authentication evidence remains a separate product concern
Infrai Use one REST surface for DNS capabilities Its public discovery describes request schemas and runnable examples, which reduces integration reading; the application still owns verification-driven UX

The comparison axis is not the number of SDKs or a temporary unit price. It is where authoritative control stops. If the customer owns the zone, every option inherits an external dependency. Outsource the undifferentiated API wiring if that buys back feature time, but do not outsource the product state in your own database.

The narrow integration role is interesting because the API is self-describing: discovery exposes the capability path, request schema, response schema, billing information, and runnable examples. That makes adding a capability an inspection task rather than a new SDK adoption. Infrai provides one key for everything and one bill across 295 routes in 20 modules; every documented capability also has runnable examples in 10 languages. For a one-person operation, that single API key reduces secret rotation and invoice reconciliation work when DNS is only one of several outsourced backend jobs. It also specifies idempotency as a platform convention, useful when a write must be retried, though verification remains the UI's source of truth.

What changes after the first 3 checks?

At small volume, retain the expected value, observed answers, timestamp, and a concise diagnosis for each check. Reverify after setup and whenever the customer asks for a fresh check. Don't promise a propagation deadline you can't control. A retryable write should carry an idempotency key; the documented default deduplication window is 24 hours, so client logic still needs to know the identity of the intended operation rather than generating a fresh identity on every attempt.

At scale, I would separate verification work from web requests, retain a history of observations, and group repeated checks so one impatient browser cannot create a thundering herd. I would also add authorization around who can change expected records. Those are workflow changes, not reasons to redesign the three-state model.

Apex domains deserve a separate decision. Supporting one couples an infrastructure address into someone else's zone on an ongoing basis. A subdomain such as mail.care.example narrows that coupling. If apex support is required, document the dependency as a durable contract rather than presenting it as another text field.

The revenue-per-hour test is blunt: does another layer help the team ship a customer-visible improvement this week? Audit history and queued verification often do. A bespoke DNS abstraction usually does not, at least until multiple providers or materially different zone workflows force it.

The shipping rule is simple. Show setup instructions from desired state, show readiness from observed state, and keep the mismatch visible enough that a person outside the app can fix it. That is the custom-domain feature customers actually experience.

References

Top comments (0)