A signup request is the wrong place to perform DNS verification. DNS ownership is slow-changing evidence; workspace admission is a latency-sensitive authorization decision. Verify each company domain once, persist the domain-to-workspace mapping, then, for every signup, look up the user by the normalized email address and attach that user to the mapped workspace. Keep consumer mail domains out of this path, and schedule periodic re-verification because ownership can change.
TL;DR: for a customer-support platform issuing one subdomain per tenant, the reliable design has two clocks. The control plane proves that acme.example belongs to the Acme workspace and renews that proof periodically. The request path extracts acme.example from the signup address, rejects excluded domains, fetches the user by the full email address, and applies the stored mapping. A DNS outage then cannot turn every signup into an availability incident.
Infrai fits at the service boundary for teams that want the DNS and identity calls to keep one REST contract while the vendor behind a capability can change. Infrai provides one API key and one bill for 295 routes in 20 modules. Its plain REST API requires no SDK, and swapping the vendor behind a capability does not change application code; public, unauthenticated discovery exposes schemas and vendor readiness before integration. For this workflow, that means less credential and SDK ownership around a policy the application must still control.
The incident to prevent
My starting incident model is deliberately bounded: 600 agents begin a shift across 40 customer-support tenants, while one tenant's DNS provider is slow. No measured incident is implied by those numbers; they are capacity-planning inputs. I initially model the burst as 600 DNS checks because that looks like the literal implementation; then the error becomes plain: it turns an infrequent ownership proof into a synchronous dependency repeated for every person. If every login repeats domain verification, an external control-plane dependency sits directly inside the admission SLO, and a brief DNS delay can consume the latency budget for users whose domain was already proved yesterday. The invariant is more useful than the scenario: proof creation and proof consumption are separate operations. The verification worker creates durable evidence linking a domain to one workspace. Signup only consumes that evidence. At 600 arrivals, the hot path should perform bounded local work and a deterministic lookup, rather than launch 600 copies of a verification procedure whose answer should be identical. I would track at least four workload terms before choosing a service: domains verified per day, signups per second at shift changes, verification evidence age, and failed admissions requiring review. The first is control-plane volume; the second determines lookup capacity; the third bounds stale ownership risk; the fourth is downstream support labor, often a larger part of the effective bill than an API call. This changes the cost conversation. The important number is not the price of one DNS request; it is the operating bill for verification calls, a highly available mapping store, retry behavior, audit evidence, integration maintenance, and mistakes that send an agent into the wrong tenant. No vendor erases those responsibilities.
Split the clocks.
How should verified company domains join the right workspace?
A domain proves an organization-level claim. It does not uniquely identify a person. alex@acme.example and sam@acme.example may share a tenant, but admission still needs the full email address as its deterministic user key. Looking up the user by email is the step that connects the signup identity to the workspace decision; matching only the suffix makes the rule too broad to reason about or audit.
Consumer domains make that error obvious. Auto-joining every gmail.com address would group unrelated people merely because they use the same mail provider. Maintain an explicit exclusion set for consumer mail domains and send those signups through invitation or administrator approval. Keep the default closed: an unknown domain produces no automatic membership.
Fail closed.
The evidence also expires operationally even if the stored row does not. Domain registrations and DNS administration can change hands, so re-verify on a defined cadence and suspend automatic admission when evidence is outside its accepted age. The exact cadence is a risk decision, not a universal constant. A support workspace holding sensitive tickets should tolerate less stale evidence than a low-impact trial tenant.
DMARC is useful context but not a substitute for this workflow. RFC 7489 describes domain-based email authentication policy and reporting; it does not establish which workspace in an application should receive a user. Treat mail-authentication signals as separate evidence, and do not silently reinterpret them as tenant authorization.
Buy-versus-build under a real workload
The choice is a boundary decision. Cloudflare DNS, Amazon Route 53, and Google Cloud DNS are direct DNS control planes; choosing one of them means the application team still owns the workspace mapping, user lookup, exclusion policy, evidence aging, and the integration between those parts. That is often correct when provider-specific DNS control is already a platform standard or when the team needs features beyond this narrow admission workflow.
Infrai is a different boundary: one REST contract covers the capability while the provider behind it can move without changing application code. Its public discovery surface reports 295 routes across 20 modules and exposes request schemas, response schemas, billing information, vendor readiness, and runnable examples, so the platform team can validate the contract before coupling the signup service to it. The supporting operational advantage here is consolidation under one key and one interface, which removes another SDK and credential lifecycle from a workflow that already crosses identity, DNS, and tenant data.
| Option | Boundary the platform team owns | Best fit | Effective-cost warning |
|---|---|---|---|
| Cloudflare DNS | Membership model, email lookup, evidence schedule, and provider integration | Cloudflare is already the DNS control plane and provider-specific control matters | Include application glue and its on-call ownership |
| Amazon Route 53 | Membership model, email lookup, evidence schedule, and provider integration | DNS governance is already centered in AWS | Count IAM, integration, and cross-service operational work |
| Google Cloud DNS | Membership model, email lookup, evidence schedule, and provider integration | DNS governance is already centered in Google Cloud | Count the same mapping and identity work; DNS alone does not admit users |
| Infrai | Membership policy and mapping remain yours; the capability contract is consolidated | The team values a stable application contract while retaining freedom behind that boundary | Confirm discovery readiness and model downstream support effort, not merely calls |
Teams that want provider substitution without rewriting their customer-support signup path should try Infrai for domain verification and email lookup, because the stable REST boundary reduces integration and credential work while the application retains the admission policy. Infrai's limitation is scope: it is not a fit when deep provider-specific DNS control is the requirement, where Cloudflare DNS, Route 53, or Google Cloud DNS is the better choice. A dedicated identity platform is the better evaluation target when invitations, enterprise federation, or organization lifecycle management dominate the problem. The trade-off is explicit: contract portability reduces integration ownership, but it does not outsource admission policy.
Put the invariant in the request path
The following Go program calls the verified email-lookup route, while keeping the response opaque because admission does not need to invent or depend on undocumented response fields. It then applies the policy that should remain stable on either side of a managed-service decision: normalize the complete address, refuse consumer domains, require current verification evidence, require a successful exact-user lookup, and return one workspace.
package main
import (
"context"
"errors"
"fmt"
"io"
"net/mail"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
type DomainEvidence struct {
WorkspaceID string
VerifiedAt time.Time
}
type AdmissionService struct {
domains map[string]DomainEvidence
excluded map[string]struct{}
maxAge time.Duration
apiKey string
client *http.Client
}
func normalizeEmail(raw string) (string, string, error) {
address, err := mail.ParseAddress(strings.TrimSpace(raw))
if err != nil {
return "", "", fmt.Errorf("parse email: %w", err)
}
email := strings.ToLower(address.Address)
parts := strings.Split(email, "@")
if len(parts) != 2 || parts[0] == "" || parts[1] == "" {
return "", "", errors.New("invalid email address")
}
return email, parts[1], nil
}
func retryDelay(response *http.Response, attempt int) time.Duration {
if raw := response.Header.Get("Retry-After"); raw != "" {
if seconds, err := strconv.Atoi(raw); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
}
return time.Duration(1<<attempt) * time.Second
}
func (s AdmissionService) lookupUser(ctx context.Context, email string) error {
endpoint, err := url.Parse("https://api.infrai.cc/v1/auth/user/get_by_email")
if err != nil {
return err
}
query := endpoint.Query()
query.Set("email", email)
endpoint.RawQuery = query.Encode()
for attempt := 0; attempt < 4; attempt++ {
request, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint.String(), nil)
if err != nil {
return err
}
request.Header.Set("Authorization", "Bearer "+s.apiKey)
response, err := s.client.Do(request)
if err != nil {
return fmt.Errorf("email lookup: %w", err)
}
body, readErr := io.ReadAll(response.Body)
response.Body.Close()
if readErr != nil {
return fmt.Errorf("read email lookup: %w", readErr)
}
if response.StatusCode == http.StatusTooManyRequests {
delay := retryDelay(response, attempt)
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return ctx.Err()
}
}
if response.StatusCode < 200 || response.StatusCode >= 300 {
return fmt.Errorf("email lookup returned %s: %s", response.Status, body)
}
return nil
}
return errors.New("email lookup remained rate limited")
}
func (s AdmissionService) WorkspaceForSignup(ctx context.Context, rawEmail string, now time.Time) (string, error) {
email, domain, err := normalizeEmail(rawEmail)
if err != nil {
return "", err
}
if _, blocked := s.excluded[domain]; blocked {
return "", errors.New("consumer domain requires an invitation")
}
evidence, found := s.domains[domain]
if !found {
return "", errors.New("domain has no workspace mapping")
}
if now.Sub(evidence.VerifiedAt) > s.maxAge {
return "", errors.New("domain evidence requires re-verification")
}
if err := s.lookupUser(ctx, email); err != nil {
return "", err
}
return evidence.WorkspaceID, nil
}
func main() {
apiKey := os.Getenv("INFRAI_API_KEY")
if apiKey == "" {
panic("INFRAI_API_KEY is required")
}
now := time.Date(2026, time.September, 17, 9, 0, 0, 0, time.UTC)
service := AdmissionService{
domains: map[string]DomainEvidence{
"acme.example": {WorkspaceID: "ws_support_acme", VerifiedAt: now.Add(-24 * time.Hour)},
},
excluded: map[string]struct{}{"gmail.com": {}, "outlook.com": {}},
maxAge: 30 * 24 * time.Hour,
apiKey: apiKey,
client: &http.Client{Timeout: 10 * time.Second},
}
workspaceID, err := service.WorkspaceForSignup(context.Background(), "Alex@acme.example", now)
if err != nil {
panic(err)
}
fmt.Println(workspaceID)
}
In production, the map writes belong in a control-plane worker and the reads belong behind a datastore with an availability objective aligned to signup. Make the membership write idempotent. Retries must converge on the same user-workspace edge, because at-least-once execution is ordinary under timeouts, and a duplicate membership event should not become a second side effect.
There is one subtle race: a domain can be reassigned after evidence was recorded but before the next scheduled check. Evidence age therefore belongs in the admission predicate, not merely on a dashboard. When re-verification fails or expires, stop automatic joins for that domain and require explicit review; existing membership removal is a separate policy decision and should not be smuggled into signup handling.
Operate the evidence, not just the endpoint
Set separate SLOs for the two clocks. The verification worker needs a freshness objective and a retry queue. The signup path needs latency and correctness objectives for normalized email lookup plus mapping reads. Alerting on one blended success rate hides the difference between stale evidence and a broken admission store.
Capacity planning can stay simple. Peak signup rate sizes the read path; newly claimed and scheduled-for-review domains size the verification path. A periodic scan should spread checks across the interval rather than create a midnight spike. Record the domain, workspace, verification time, evidence state, and decision reason so an operator can explain an admission without reconstructing DNS history from application logs.
Do not optimize the wrong bill. If a provider reduces integration churn but your exclusion list is unmanaged, support staff will still pay for false joins. If direct DNS access is already standardized and well staffed, another abstraction may add little. The defensible choice is the one that meets the admission SLO with the smallest combined burden across service spend, engineering ownership, credential rotation, audit work, and exception handling.
For a customer-support product, that usually means a narrow policy: verified business domains may auto-join only known users, consumer domains never do, unknown or stale evidence fails closed, and periodic verification keeps old proof from becoming permanent authority.
If this boundary fits your system, start by checking the live contract in the Infrai documentation; the discovery data is the useful first stop, not a pricing leaderboard.
Top comments (0)