DEV Community

ZachariahHolloway9058
ZachariahHolloway9058

Posted on

Tenant Domain Caps: 5 Places the Limit Breaks When MX Records Drift

Enforce the per-tenant domain limit in your own application, against your own tenant records — then use the published zone as the reconciliation source, never as the thing you count at request time. The DNS layer has no idea what a tenant is, so the quota has to live in your model; the resolver, meanwhile, is the only honest answer to "how many domains does this customer actually have mail flowing through?" Both numbers are needed. Counting either one alone is how you end up telling a school district they've hit 10 of 10 while mail is quietly landing for 12 domains nobody tracked.

That gap between intent and published records is the real failure mode here, and a domain cap is where it surfaces first.

The system I'll use throughout is an edtech platform: districts sign up, each district points its school mail at our provider, and every domain needs MX plus SPF, DKIM and a DMARC policy before we'll accept mail for it. A district is one tenant. A district often has seven domains, because it absorbed three charter schools and a foundation, and two of those domains are managed by a part-time IT contractor who has the registrar password and no interest in our admin UI.

That contractor is the whole problem.

1. The counter that counted intentions instead of MX records

The first version did what everyone's first version does: a tenant_domains row was written when someone clicked Add Domain, and the cap was a count(*) against that table. Clean, transactional, fast. It also measured the wrong thing.

A row in that table means a district asked for a domain. It does not mean the domain resolves, that its MX points at us, or that the district still owns it. Three drift directions show up in practice, and each one breaks a different promise. Rows with no working MX inflate the count and lock out a legitimate add — the district is paying for ten and can only use six. Domains in the zone with no row are worse: mail flows, DMARC aggregate reports arrive, support sees a customer we have no billing record for. And the slow one, the one you find months later, is a domain whose MX was repointed away by that contractor while the row sat there marked verified, still consuming a slot.

None of that is exotic. It's the normal entropy of a system where two parties can edit the same state and only one of them reads your changelog.

2. Where should a per-tenant domain limit live — the app, the gateway, or the zone itself?

Three candidate enforcement points, and they know different things.

Enforcement point Knows the tenant Knows what's published Typical failure
Application / service layer Yes, transactionally No Counts intent, drifts from reality
API gateway or policy engine Only what the token claims No Stale claims, cap edits need a token refresh
DNS provider or zone No Yes, authoritatively Account-wide limits that don't map to your tenants

The application wins, and it isn't close. It's the only layer that can hold the quota and the write in one transaction, the only one you can unit test without a network, and the only one where "this tenant bought the 25-domain plan last Tuesday" takes effect immediately instead of when a JWT expires.

Gateway enforcement is tempting because it's centralized. In my experience it fails on the boring detail: the cap is tenant state, and pushing tenant state into token claims means every plan change is a cache invalidation problem. Provider-level limits are real and worth reading — records per zone, zones per account, API rate limits on record writes — but they are the provider's limits, not your product's, and they're shared across all your tenants. Treat those as capacity planning, not as customer policy.

So the zone gets a different job. It's the auditor.

3. Count what resolvers answer, not what the form accepted

Reconciliation is a scheduled job, and it compares three sets: the domains the tenant claims in your database, the records present in the zones you manage, and what a public resolver returns for each claimed name. Anything in one set and not the others is drift, and drift gets a row with a first-seen timestamp — not an alert on the first miss, because DNS propagation and a 3600-second TTL will generate false positives all day.

Two details matter more than the diffing logic.

Store the limit and the current count together on the tenant record, refreshed by the reconciler, so support can answer "why can't they add another domain?" by looking at one row instead of running a query someone will get wrong at 2am. And resolve MX with an explicit timeout, because a dead nameserver on one district's vanity domain should not stall the reconciliation of the other four hundred. Per RFC 5321, a name with no MX records falls back to its A/AAAA address for mail delivery, so "no MX" is not the same as "no mail" — if you treat an empty MX set as proof the domain is unused, you'll release slots for domains that are still receiving mail.

4. A reconciler you can safely run twice

The cap check belongs in the write path, and it has to survive a retry. A district admin double-clicks Add Domain; the first request commits and the response times out; the retry must not consume a second slot.

// AddDomain is idempotent per (tenantID, name): a retry returns the existing
// domain instead of consuming another slot against the tenant's limit.
func (s *Store) AddDomain(ctx context.Context, tenantID, name string) (Domain, error) {
    tx, err := s.db.BeginTx(ctx, &sql.TxOptions{Isolation: sql.LevelSerializable})
    if err != nil {
        return Domain{}, err
    }
    defer tx.Rollback()

    var existing Domain
    err = tx.QueryRowContext(ctx,
        `SELECT id, name, state FROM tenant_domains WHERE tenant_id = $1 AND name = $2`,
        tenantID, name).Scan(&existing.ID, &existing.Name, &existing.State)
    if err == nil {
        return existing, tx.Commit() // already held: not a quota event
    }
    if err != sql.ErrNoRows {
        return Domain{}, err
    }

    var limit, used int
    if err := tx.QueryRowContext(ctx, `
        SELECT t.domain_limit,
               (SELECT count(*) FROM tenant_domains d
                 WHERE d.tenant_id = t.id AND d.state <> 'released')
          FROM tenants t WHERE t.id = $1 FOR UPDATE`,
        tenantID).Scan(&limit, &used); err != nil {
        return Domain{}, err
    }
    if used >= limit {
        return Domain{}, &QuotaError{Limit: limit, Used: used}
    }

    d := Domain{TenantID: tenantID, Name: name, State: "pending_mx"}
    if err := tx.QueryRowContext(ctx,
        `INSERT INTO tenant_domains (tenant_id, name, state) VALUES ($1, $2, $3)
         ON CONFLICT (tenant_id, name) DO UPDATE SET name = EXCLUDED.name
         RETURNING id`, tenantID, name, d.State).Scan(&d.ID); err != nil {
        return Domain{}, err
    }
    return d, tx.Commit()
}
Enter fullscreen mode Exit fullscreen mode

The reconciler that runs behind it is deliberately boring. It reports; it does not delete.

// Reconcile classifies each claimed domain by comparing tenant intent with what
// resolvers actually publish. Drift is recorded, never auto-corrected.
func Reconcile(ctx context.Context, claimed []Domain, wantMX string) []Drift {
    var out []Drift
    for _, d := range claimed {
        rctx, cancel := context.WithTimeout(ctx, 5*time.Second)
        mx, err := net.DefaultResolver.LookupMX(rctx, d.Name)
        cancel()

        switch {
        case err != nil:
            out = append(out, Drift{Domain: d.Name, Kind: "lookup_failed", Detail: err.Error()})
        case len(mx) == 0:
            // RFC 5321: no MX means implicit-MX fallback to A/AAAA, not "unused".
            out = append(out, Drift{Domain: d.Name, Kind: "no_mx"})
        case !strings.EqualFold(strings.TrimSuffix(mx[0].Host, "."), wantMX):
            out = append(out, Drift{Domain: d.Name, Kind: "mx_repointed", Detail: mx[0].Host})
        }
    }
    return out
}
Enter fullscreen mode Exit fullscreen mode

Auto-correction is the tempting next step and I'd argue against it. A reconciler that rewrites MX records can turn one bad config push into a district-wide mail outage, and you will be the one explaining it. Let it open a ticket.

Zone-as-code tooling — OctoDNS, DNSControl, external-dns — solves the adjacent problem well: it makes the published zone a reviewed artifact with a diff, which kills the "who changed this?" question. It still won't enforce a per-tenant cap, because the tenant model isn't in the zone file. Two systems, two jobs.

5. When a hard per-tenant cap is the wrong control

A limit that blocks a paying customer at 2am is a bad trade. That's the part teams get backwards: the cap exists to stop abuse and runaway cost, and neither of those is an emergency measured in minutes.

For plan tiers, a soft limit reads better in production — allow the add, mark the tenant over quota, alert the account team the same day. Reserve hard enforcement for the free tier, where the real risk is one script registering four hundred throwaway domains to farm verification jobs and DMARC report ingestion. The catch is that soft limits need a working escalation path or they quietly become no limits at all.

There are cases where a domain cap isn't the right tool. If your cost driver is aggregate report volume rather than domain count — RFC 7489 reporting on a large district can dwarf the cost of the domain record itself — cap the reports. If tenants routinely arrive with dozens of acquired domains, stick with a review queue and drop the numeric limit entirely. And if you're multi-region, be careful about counting against a resolver in one region only; I'm not certain there's a clean answer to split-horizon setups beyond resolving from several vantage points and treating disagreement as its own drift class.

Whatever you pick: count intent in your database, count reality from the resolver, and alert on the difference. The cap is just the place where the difference finally becomes visible.

References

Top comments (0)