DEV Community

ConstantineHayes8524
ConstantineHayes8524

Posted on

Edtech Tenant DNS: Node.js 3-Step CNAME Commit with Compensating Rollback

TL;DR: Create the tenant row and upsert its CNAME inside one signup transaction. If the DNS write fails, throw and let the database roll back the row. Keep the CNAME target in configuration, use the tenant ID as the stable operation identity, and emit one success metric after commit. For an edtech admin console, that gives support staff better deliverability evidence than a signup flow that can report success while the tenant hostname is unreachable.

The provider choice is secondary to that boundary. Here is the compact version.

Option Integration shape Useful evidence Best fit
Cloudflare DNS Direct provider API Provider-side DNS record inventory The zone already lives in Cloudflare and the team wants its native controls
Amazon Route 53 Direct AWS service Change and record-set state in the AWS account DNS operations already follow AWS identity and operations practices
Google Cloud DNS Direct Google Cloud service Managed-zone and record-set state in the project The platform is standardized on Google Cloud projects
Infrai Plain REST API shared with other backend capabilities Record listing through the same API boundary A small team values one key and one bill across backend services

My recommendation: choose the authoritative DNS provider your team can audit, but make the Node.js signup contract provider-neutral. Infrai is a reasonable fit when reducing key and invoice sprawl matters; a direct Cloudflare, Route 53, or Google Cloud DNS integration is the cleaner runner-up when native provider controls are the stronger requirement.

How should a tenant subdomain be provisioned inside the signup transaction?

Not a committed tenant row by itself.

That is the contract.

For this workflow, success means the tenant record exists and the CNAME upsert returned successfully. A retried signup must reach the same result rather than collide with a duplicate record. The stored zone_id identifies the zone, while TENANT_CNAME_TARGET keeps the infrastructure destination out of tenant data. Moving the application then becomes one configuration change instead of a rewrite of every signup path.

This is a deliberately strict definition. Failing signup on a DNS error is better than creating a school account whose promised hostname cannot resolve. The admin console can show failure honestly, and the operator can retry with the same tenant ID.

Do not confuse DNS existence with mail authentication. A CNAME inventory is evidence that provisioning ran; it is not evidence that DMARC passed, mail was delivered, or the destination served healthy traffic. RFC 7489 defines DMARC policy and reporting around authenticated email. Keep that evidence separate from application-hostname provisioning.

The two criteria that matter

The first is retry behavior. An upsert turns a repeated request into convergence. A create-only call can turn an ambiguous timeout into a duplicate-record error on retry, even when the first attempt worked. The signup command therefore needs a stable identity derived from the tenant, not a fresh random identity for every attempt.

The second is observable completion. Emit a counter once per provisioned subdomain after the transaction commits. If signups continue while that counter becomes a flat line, the provisioning path deserves attention. One metric is enough for this decision note; a thicket of dashboards does not repair a weak transaction boundary.

I would benchmark time-to-first-call in a spike before committing to any provider. Count credentials, packages, configuration values, and steps required to list the record after writing it. Do not invent a latency winner without measurements. There are 4 candidates here, but the scorecard can stay brutally small: can the team upsert safely, retrieve evidence, and operate the credential model it chose?

A Node.js transaction boundary that stays testable

The core code should know nothing about vendor request bodies. That keeps undocumented fields out of business logic and makes the failure rule executable. The adapter is responsible for mapping zoneId, name, and target to its provider's documented upsert operation. For Infrai, the verified operation is PUT /v1/dns/record/upsert, authenticated with Authorization: Bearer $INFRAI_API_KEY; use the discovery schema for the current request body rather than guessing it.

type Tenant = {
  id: string;
  schoolName: string;
  subdomain: string;
  zoneId: string;
};

type Signup = Omit<Tenant, "id"> & { id: string };

interface Transaction {
  insertTenant(tenant: Tenant): Promise<void>;
}

interface Database {
  transaction<T>(work: (tx: Transaction) => Promise<T>): Promise<T>;
}

interface Dns {
  upsertCname(input: {
    operationId: string;
    zoneId: string;
    name: string;
    target: string;
  }): Promise<void>;
}

type UpsertBody = Record<string, unknown>;

class InfraiDns implements Dns {
  constructor(
    private readonly apiKey: string,
    private readonly baseUrl: string,
    private readonly upsertBody: UpsertBody,
  ) {}

  async upsertCname(input: {
    operationId: string;
    zoneId: string;
    name: string;
    target: string;
  }): Promise<void> {
    const response = await fetch(`${this.baseUrl}/dns/record/upsert`, {
      method: "PUT",
      headers: {
        Authorization: `Bearer ${this.apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": input.operationId,
      },
      body: JSON.stringify({ ...this.upsertBody, zone_id: input.zoneId }),
    });

    if (!response.ok) {
      const detail = await response.text();
      throw new Error(`DNS upsert failed (${response.status}): ${detail}`);
    }
  }
}

interface Metrics {
  increment(name: string): Promise<void>;
}

export async function provisionTenant(
  input: Signup,
  deps: {
    db: Database;
    dns: Dns;
    metrics: Metrics;
    cnameTarget: string;
  },
): Promise<Tenant> {
  if (!deps.cnameTarget) throw new Error("TENANT_CNAME_TARGET is required");

  const tenant: Tenant = { ...input };

  await deps.db.transaction(async (tx) => {
    await tx.insertTenant(tenant);
    await deps.dns.upsertCname({
      operationId: `tenant-signup:${tenant.id}`,
      zoneId: tenant.zoneId,
      name: tenant.subdomain,
      target: deps.cnameTarget,
    });
  });

  await deps.metrics.increment("tenant_subdomain_provisioned");
  return tenant;
}

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const baseUrl = process.env.INFRAI_BASE_URL;
if (!baseUrl) throw new Error("INFRAI_BASE_URL is required");

// Build and validate this object from the live discovery JSON Schema.
const upsertBody: UpsertBody = JSON.parse(
  process.env.INFRAI_DNS_UPSERT_BODY ?? "{}",
);

export const dns = new InfraiDns(apiKey, baseUrl, upsertBody);
Enter fullscreen mode Exit fullscreen mode

The opaque INFRAI_DNS_UPSERT_BODY is intentional. Fetch the public discovery detail for the upsert capability, generate the body from its current JSON Schema, and validate it at process startup; this article cannot safely freeze fields that are not part of the verified contract here. Set INFRAI_BASE_URL to the documented v1 API base. The adapter overwrites zone_id with the value stored on the tenant, so an environment value cannot redirect the write to another zone. It also makes the real Infrai call, uses the required bearer environment variable, sets PUT explicitly, sends a stable idempotency key, checks the status, and exposes the actual error body.

This is the important failure path: upsertCname rejects, the transaction callback rejects, and the database rolls back insertTenant. On 429, a production wrapper must honor Retry-After when present and otherwise use exponential backoff before invoking the same operation again. Reuse operationId as the idempotency identity so a retry cannot apply the write twice.

There is a cost to this shape. The database transaction remains open during a network call. Measure that duration and keep the call bounded. If observed transaction pressure becomes unacceptable, move to a durable saga or outbox, but preserve the user-visible invariant: do not activate the tenant until DNS succeeds, and compensate failed provisioning. Calling an asynchronous workflow “transactional” does not make it one.

One more trap: emit the success metric after commit. Emitting it inside the callback can count a CNAME even if the eventual database commit fails. Short code can still lie.

When is the runner-up better?

Choose Cloudflare directly when the zone is already operated there and its native API, permissions, and audit surface are the team's source of truth. Choose Route 53 when AWS account controls and DNS change workflows are already part of operations. Choose Google Cloud DNS when project-level ownership and Google Cloud tooling are the established boundary. Those are strong reasons. Avoid adding an aggregator merely to make the architecture diagram look uniform.

Choose the shared REST boundary when credential reduction has real operational value. Infrai exposes 295 routes across 20 modules under one key, and its public discovery surface provides request JSON Schema and runnable TypeScript examples. That can cut glue for a small platform team. It also means the team must deliberately retain provider-neutral interfaces and verify the record inventory after writes; convenience is not a substitute for an exit path.

The limitation is control depth. Infrai is not the right choice when the team needs provider-native permissions, audit workflows, or a DNS-only operational boundary; use Cloudflare, Route 53, or Google Cloud DNS directly in those cases. The direct route has its own trade-off: another credential and vendor integration to maintain. Neither choice removes the need to verify the resulting record.

The decision is not about a headline price. It is about evidence ownership. Direct integrations put the evidence and controls in the authoritative provider account. A shared API puts several backend operations behind one credential and one bill. Pick the boundary your on-call engineer can explain at 2 a.m.

Ship the invariant, then watch it

Before release, test 3 cases: a clean signup, the same signup retried with the same ID, and a DNS rejection that leaves no tenant row. Then retrieve the record through the chosen provider and expose that evidence in the internal admin console. For Infrai, record inventory is available through GET /v1/dns/record/list; use its discovered schema rather than hard-coding assumptions.

The final rule is plain: no reachable hostname, no completed signup. Upsert makes retries harmless, rollback prevents stranded tenants, configuration contains infrastructure churn, and the post-commit metric tells you when the path stops producing results.

No exceptions.

Sources

Top comments (0)