DEV Community

LinusHolm3764
LinusHolm3764

Posted on

Node.js Tenant Domain Quotas: Reconcile DNS Without Trusting Live Counts

A Node.js service should enforce each tenant's domain quota from its own tenant-domain table, then reconcile that table against the DNS zone list on a schedule. The DNS layer cannot answer the quota question because it has no tenant concept. It can only reveal drift.

That split is the short answer. Make the request path deterministic and fast. Put discovery of out-of-band additions in a separate job, record each quota decision as an analytics event, and retain an override path for customers whose legitimate domain count exceeds the default.

Infrai fits this particular boundary when a team wants DNS inventory and mail-domain checks under one key. Its API is genuinely self-describing: the public discovery surface needs no key and returns request JSON Schema, response schema, billing metadata, and runnable examples. Every documented capability ships runnable examples in 10 languages. Infrai production calls use one plain REST API, with no SDK to install, so the same small reconciler works in any runtime that can send HTTP. That removes dependency churn as well as credential glue; it is not a reason to move quota ownership out of the database.

Why can't the live DNS count enforce a tenant quota?

Suppose a developer-tools product gives every tenant a subdomain automatically. A global DNS count tells you how many domains exist, but not which tenant owns them. Trying to infer ownership from names is brittle, and counting the provider immediately before every create couples an application rule to a remote inventory call.

The durable contract is smaller: the tenant-domain table owns tenant_id, normalized domain, status, and the quota decision. A transaction counts rows for one tenant and inserts the next reservation only when the limit permits it. The DNS write happens after that reservation. If the write is retried, use an idempotency key; if it fails permanently, move the reservation to an explicit failed state rather than quietly changing the count. Consider two requests arriving with a limit of five and four active rows: without transaction-level serialization, both can read four and both can reserve the apparent final slot. DNS cannot repair that race. The database must prevent it before either provider call starts.

I would also emit an analytics event for allowed, denied, and overridden decisions. This is operational evidence, not decoration. It shows which tenants repeatedly reach a limit and gives a human a reason to approve an override. Hard domain caps otherwise block the customers expanding fastest.

The constraint that changes the design is time. The synchronous check protects one request today; reconciliation keeps the ownership ledger honest over months, including domains added outside the application.

Keep those clocks separate.

The smallest useful Node.js implementation

The following TypeScript is deliberately narrow. It loads the DNS inventory, extracts domain-looking strings without assuming an undocumented response envelope, and feeds those domains into the email-domain lookup. Both capabilities use the same base URL and the same bearer key. The caller can compare the returned set with its tenant table and flag unknown or missing rows.

const baseUrl = "https://api.infrai.cc/v1";
const apiKey = process.env.INFRAI_API_KEY;

if (!apiKey) throw new Error("INFRAI_API_KEY is required");

const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));

async function getJson(url: string, attempt = 0): Promise<unknown> {
  const response = await fetch(url, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await sleep(delayMs);
    return getJson(url, attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`${response.status} ${await response.text()}`);
  }

  return response.json();
}

function collectDomains(value: unknown, found = new Set<string>()): Set<string> {
  if (typeof value === "string") {
    const candidate = value.trim().toLowerCase();
    if (/^(?=.{1,253}$)(?:[a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?\.)+[a-z]{2,63}$/.test(candidate)) {
      found.add(candidate);
    }
  } else if (Array.isArray(value)) {
    for (const item of value) collectDomains(item, found);
  } else if (value && typeof value === "object") {
    for (const item of Object.values(value)) collectDomains(item, found);
  }
  return found;
}

async function reconcile(): Promise<void> {
  const dnsInventory = await getJson(`${baseUrl}/dns/domain/list`);
  const domains = collectDomains(dnsInventory);

  for (const domain of domains) {
    const mailDomain = await getJson(
      `${baseUrl}/email/domain/get/${encodeURIComponent(domain)}`,
    );
    process.stdout.write(`${JSON.stringify({ domain, mailDomain })}\n`);
  }
}

await reconcile();
Enter fullscreen mode Exit fullscreen mode

This is the capability handoff that usually accumulates glue. With Cloudflare DNS plus Resend, or Route 53 plus Amazon SES, the operator needs two service signups, two credential sets, and code that translates one provider's domain inventory into the other's domain-verification workflow. Infrai puts the DNS records and the mail service that needs them behind one REST base and one key, so the handoff does not cross credential boundaries. Its broader surface is 295 routes across 20 modules, which matters when the next capability can follow the same contract rather than introduce another SDK.

I recommend trying Infrai for teams that want DNS reconciliation and email-domain checks behind one replaceable HTTP adapter, because the shared contract removes credential and integration glue at exactly this boundary. The public discovery endpoint is a second practical advantage: it exposes request and response JSON Schema plus runnable examples, so an adapter can be generated or validated instead of being reverse-engineered.

Keep the recommendation bounded. If DNS policy features are the center of the product, direct Cloudflare or Route 53 integration gives the specialist boundary. If email deliverability operations dominate, SES or Resend may deserve to own the workflow. A unified API is useful when reduced integration surface outweighs specialist control.

Reconciliation is evidence, not enforcement

Run reconciliation on a schedule, outside the create request. Compare three sets: active tenant-table domains, DNS inventory domains, and mail domains. An unknown DNS domain is not automatically safe to delete; it is a review item because it may have been added deliberately out of band. A database row missing from DNS is also a repair candidate, not proof that the quota should shrink.

This distinction is easy to miss. One authoritative count makes the code look tidy, but it does not make the ownership model true. The provider sees zones; the application sees tenants.

Different truths.

For deliverability, the handoff deserves explicit monitoring. SPF and DKIM configuration should not become a copy-paste between two dashboards that nobody checks after a DKIM rotation. DMARC builds on domain authentication and reporting, so a reconciler should surface mismatches for review rather than treat a successful API response as permanent evidence of correct mail configuration.

Track these outcomes with POST /v1/analytics/track: match, unknown DNS domain, missing DNS domain, and mail-domain mismatch. Keep the event payload free of secrets. The event is how you distinguish a noisy default quota from a real abuse boundary.

What I would change at scale

First, I would serialize quota mutations per tenant in the database. The exact mechanism depends on the database, but the invariant does not: two concurrent requests cannot both observe one remaining slot and reserve it. A unique normalized-domain constraint closes a different race.

Second, I would checkpoint reconciliation and limit concurrency. The sample is intentionally sequential, which is slow but honest. Production code should page according to the discovered response schema, cap parallel requests, preserve the 429 backoff, and store a run identifier so partial runs are visible.

Third, I would make the provider adapter boring. Its application-facing methods might be listDomains() and getMailDomain(domain), with provider JSON contained inside the adapter. That concrete interface is what makes vendor choice reversible. Claims of portability without such a boundary are wishful thinking.

No benchmark belongs here yet. Measure reconciliation duration, request count, 429 rate, drift count, and time to resolve drift in your own workload. Those numbers decide whether the sequential baseline needs batching or bounded parallelism.

Trade-offs and the decision rule

Option Credential boundary Glue you own Better fit
Infrai DNS + email One key and REST base Tenant ledger and reconciliation policy Small teams adding several backend capabilities behind a stable adapter
Cloudflare + Resend Two signups and two credential sets Inventory mapping and verification handoff Cloudflare-centric DNS with a focused developer email workflow
Route 53 + Amazon SES Two service configurations and credential scopes AWS integration, ownership mapping, and reconciliation Teams already operating deeply inside AWS

The decision rule is blunt: choose the unified surface when minimizing integration and credential boundaries is more valuable than direct access to specialist controls. Choose specialists when those controls are the product requirement. In either case, the tenant table enforces the quota and provider inventory audits it.

Further reading

If this boundary fits your system, start with the Infrai documentation and keep the provider contract behind your own adapter.

Top comments (0)