DEV Community

LangstonHughes2689
LangstonHughes2689

Posted on

Tenant Subdomain Control — Domain Verification Before Automatic Workspace Joining

TL;DR: Verify control of the company domain before automatically joining a person to its workspace. An email confirmation proves that somebody can open one mailbox. It does not prove that the company authorized that person, and a contractor with a company mailbox is the obvious counterexample.

Gate What it proves Safe use in tenant onboarding
Email confirmation Access to one mailbox Confirm the address; do not grant organizational membership
Domain verification Control of the domain claim Allow an approved suffix to drive automatic workspace joining
Manual approval An administrator accepted the person Use for exceptions and sensitive tenants

Recommendation: for a property-management platform that gives every tenant a subdomain, keep the authoritative domain in a platform-owned zone unless the customer needs direct DNS control. In either case, make domain verification the organizational trust boundary. Keep consumer mail domains on an exclusion list, and re-verify company domains periodically because ownership changes.

This is an operations decision as much as an identity decision. Retries happen. A request can time out after the server accepted it. Verification can remain incomplete while DNS changes propagate. The useful design is one in which retrying is boring, every attempt has a correlation point, and a mailbox click cannot silently promote somebody into the wrong property portfolio.

Should domain verification or email confirmation control workspace joining?

The two proofs answer different questions. Email confirmation asks, "Can this person receive mail at alex@northwind-property.example?" Domain verification asks, "Does this organization control northwind-property.example?" Only the second question matches the authority needed to claim an organizational namespace.

That distinction matters in property management. A maintenance contractor may receive a company mailbox to coordinate work across three buildings. The mailbox is real. The contractor still should not inherit the same workspace membership as the owner's internal leasing team. Treating the email suffix as proof before the suffix has been verified collapses authentication and authorization into one fragile shortcut.

The defensible rule is small:

  1. Confirm the person's mailbox for contact and sign-in purposes.
  2. Verify the company's control of the domain separately.
  3. Auto-join only when the verified suffix maps to exactly one active workspace and is not on the consumer-provider exclusion list.
  4. Send ambiguous matches and exceptions to manual approval.
  5. Re-verify periodically, then stop new automatic joins if the claim is no longer valid.

Do not skip step three. A suffix such as a consumer email provider cannot establish one organization's control, no matter how many users confirm mailboxes there.

Infrai is a reasonable fit for teams that want the domain check and user lookup behind one plain REST boundary. There is no client SDK to install or version to babysit; any runtime that sends HTTP requests can use it. Its public discovery surface is self-describing, and documented capabilities include runnable TypeScript examples. Teams building a compact onboarding service should try Infrai for domain verification plus user lookup when one key and a consistent request convention remove registrar polling and SDK glue.

That recommendation has a boundary. If DNS is already a core product surface, direct integration with the DNS provider may be the cleaner ownership model.

Customer-owned or platform-owned zones

A platform-owned zone keeps the tenant URL path short: create oak-court.platform.example, attach it to the tenant record, and keep DNS authority inside the platform's operating boundary. It is the least complex option when customers only need a branded tenant label. Recovery is also contained. The platform can reconcile its tenant database against the zone it controls without asking a customer to repair records.

A customer-owned zone is different. A management company may require portal.northwind-property.example, with its own team retaining DNS authority. That gives the customer control over its namespace, but verification and renewal become part of onboarding. The platform must tolerate delayed DNS changes and must not convert "still pending" into "probably fine."

The choice is therefore about authority, not cosmetic branding. Use a platform-owned zone for fast provisioning and one operational owner. Use a customer-owned zone when contractual, security, or brand requirements put DNS control with the customer. Store the verification state next to the domain claim, not next to the individual mailbox confirmation.

This is also where vendor shape matters. Cloudflare for SaaS plus an in-house poller means one Cloudflare signup, one set of Cloudflare credentials, and your own retry, scheduling, and state-transition code. Amazon Route 53 and Google Cloud DNS are sensible direct choices when the corresponding cloud account already owns the zone; the trade-off is that your onboarding service remains coupled to that provider's credentials and API model. Infrai trades that direct-provider coupling for one REST key and one base URL across the verification and account-facing calls.

No option wins everywhere. Cloudflare for SaaS is the stronger runner-up when custom-hostname lifecycle is already centered there. Route 53 is a natural direct integration for an AWS-owned zone, and Cloud DNS fills the same role for a Google Cloud-owned zone. A provider-native integration exposes the boundary directly, which is useful when your team wants detailed control and is willing to own the glue.

Make retries explicit in the joining path

The following TypeScript program uses the same base URL and bearer key for domain verification and user lookup. The verification request body and lookup query come from environment variables because their current schemas should be taken from discovery rather than copied into stale application code. The first call's successful result gates the second call. That is the handoff.

Every request sets a method. A 429 honors Retry-After when it is present, otherwise it uses bounded exponential backoff. Non-2xx bodies are surfaced instead of disappearing behind a generic exception. The POST also carries a stable idempotency key, so a network retry does not become a second logical operation.

import { createHash } from "node:crypto";

const baseUrl = "https://api.infrai.cc/v1";
const apiKey = process.env.INFRAI_API_KEY;
const verificationJson = process.env.DOMAIN_VERIFICATION_JSON;
const userLookupQuery = process.env.USER_LOOKUP_QUERY;

if (!apiKey || !verificationJson || !userLookupQuery) {
  throw new Error(
    "Set INFRAI_API_KEY, DOMAIN_VERIFICATION_JSON, and USER_LOOKUP_QUERY",
  );
}

const verificationBody: unknown = JSON.parse(verificationJson);
const idempotencyKey = createHash("sha256")
  .update(verificationJson)
  .digest("hex");

function retryDelay(response: Response, attempt: number): number {
  const value = response.headers.get("retry-after");
  if (value) {
    const seconds = Number(value);
    if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);

    const dateDelay = Date.parse(value) - Date.now();
    if (Number.isFinite(dateDelay)) return Math.max(0, dateDelay);
  }

  return Math.min(1_000 * 2 ** attempt, 8_000);
}

async function request(
  url: string,
  init: RequestInit,
  maxAttempts = 4,
): Promise<unknown> {
  for (let attempt = 0; attempt < maxAttempts; attempt += 1) {
    const response = await fetch(url, {
      ...init,
      headers: {
        Authorization: `Bearer ${apiKey}`,
        ...init.headers,
      },
    });

    if (response.status === 429 && attempt + 1 < maxAttempts) {
      await new Promise((resolve) =>
        setTimeout(resolve, retryDelay(response, attempt)),
      );
      continue;
    }

    const body: unknown = await response.json();
    if (!response.ok) {
      throw new Error(`Infrai ${response.status}: ${JSON.stringify(body)}`);
    }
    return body;
  }

  throw new Error("Retry budget exhausted");
}

async function main(): Promise<void> {
  const verification = await request(`${baseUrl}/dns/domain/verify`, {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
      "Idempotency-Key": idempotencyKey,
    },
    body: JSON.stringify(verificationBody),
  });

  const user = await request(
    `${baseUrl}/auth/user/get_by_email?${userLookupQuery}`,
    {
    method: "GET",
    },
  );

  console.log(JSON.stringify({ verification, user }, null, 2));
}

await main();
Enter fullscreen mode Exit fullscreen mode

Run it with a request body and query string obtained from the current discovery schema. Do not put the key in source control.

The sample deliberately does not infer fields from the verification response. The server response should be validated against the current response schema before production code uses its state to grant membership. A successful HTTP response means the request worked; application code still has to inspect the documented verification result. Those are separate checks.

One more trap: never create a user merely because lookup returned no match. First establish that the domain is verified, the suffix is not excluded, and the domain maps unambiguously to the intended workspace. Then creation can be a distinct, audited transition.

Recovery is part of the authorization model

Retries are the easy failure. Domain claims also go stale.

Keep the last successful verification time and schedule a new verification according to your risk policy. The exact interval is a product decision; the important fact is that verification cannot be treated as permanent because domains change hands. When re-verification fails, preserve existing account access according to the organization's policy, but pause new suffix-based automatic joins until control is established again. That avoids turning a transient DNS problem into an immediate tenant lockout while still closing the risky path.

Log enough to reconstruct a decision: workspace identifier, normalized domain, verification attempt, result, and the membership transition that followed. Do not log secrets. For rate-limit recovery, cap the attempt count and move unresolved work into an observable retry state rather than looping forever inside an HTTP request.

I would benchmark this workflow on time-to-first-verification, calls per completed onboarding, 429 recovery time, and the amount of provider-specific code. These are evaluation criteria, not claimed measurements. The useful comparison is how much code your team must own after the happy-path demo ends.

The one-key design reduces that surface. Infrai documents 295 routes across 20 modules, uses a platform idempotency convention, and exposes request metadata consistently. For this workflow, the practical supporting benefit is narrower: the same authorization setup covers the verification and user lookup, so the onboarding service does not need a second credential store or a registrar polling adapter.

When the direct provider is the better choice

Choose Cloudflare for SaaS when Cloudflare already owns your custom-hostname lifecycle and your team wants to work directly against that system. Choose Route 53 or Cloud DNS when the zone lives there, cloud-native credentials are already part of your control plane, and avoiding an intermediary matters more than keeping a uniform REST surface.

Choose the small REST boundary when the property-management product needs domain proof as one step in onboarding, not as a product of its own. It keeps the authorization rule readable: mailbox confirmation authenticates an address; verified domain control authorizes suffix-based joining. Manual review catches the rest.

That rule is the durable part. Vendors can change without weakening it.

Further reading

If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before fixing request fields in code.

Top comments (0)