Short answer: enforce each tenant's domain limit in your application database, then reconcile that record against the DNS zone list on a schedule. DNS can publish SPF, DKIM, and DMARC, but it cannot know which customer owns a domain or what that customer's contract permits.
That boundary matters in an e-commerce system. Mail delivery is the outcome we can measure; a successful POST is not proof that a message will authenticate or reach an inbox. The limit check protects tenancy, while DNS and mail telemetry provide the evidence.
Keep it boring.
Two counters. One owner.
The decision record: quota in the tenant model, evidence in DNS
Keep domain_limit and domain_count beside the tenant row (or in a tenant-quota table with a unique tenant key). Increment the count in the same transaction that records a successful domain assignment. A unique constraint on (tenant_id, normalized_domain) prevents a retry from consuming a second slot. The write path then has one clear failure boundary: if the count is at the limit, reject the assignment before calling the DNS provider.
The provider remains the source of truth for what exists in the zone. That sounds contradictory until the two questions are separated:
- The application asks, “May this tenant add one more domain?”
- Reconciliation asks, “What domains actually exist, including anything added out of band?”
Run reconciliation periodically and after an operational incident. Compare the normalized provider list with tenant records, record a drift event, and require an explicit owner decision before deleting or silently adopting anything. This gives support an answer without a live query: “limit 25, application count 18, provider count 19, drift detected at 02:10 UTC.”
Be generous by default. A domain limit that blocks a paying customer at 2 a.m. is a bad trade; raise it or put the tenant into a review state rather than making an emergency operator edit DNS by hand.
How should a tenant enforce per-domain limits while proving deliverability?
Treat domain ownership and mail authentication as a state machine, not as a counter alone. A newly requested domain is pending; only after verification should it become active and consume a production slot. SPF, DKIM, and DMARC records can then be checked independently, with evidence captured at a timestamp and tied to the tenant and domain.
For DMARC, preserve the policy and aggregate-report destination you asked the customer to publish. RFC 7489 describes the reporting model, but it does not define your commercial quota or your tenant identity. That is why the quota belongs in your model even when the records live elsewhere.
The evidence path should include at least:
- The exact normalized domain and tenant identifier.
- Verification time, record observations, and the expected selector or policy.
- A delivery sample or provider event that links the authenticated domain to a message outcome.
- The reconciliation result showing whether the provider's zone list matches your records.
I’m not sure which mailbox providers will weigh a particular DMARC policy most heavily in your traffic mix; your own delivery events and aggregate reports resolve that uncertainty. Do not turn a DNS active flag into an inbox-delivery promise.
Comparing the implementation choices
The right provider depends on how much control and operational surface your team wants. The quota decision stays yours in every row.
| Option | Useful fit | Trade-off for a multi-tenant quota | Deliverability evidence posture |
|---|---|---|---|
| Cloudflare DNS | Fast authoritative DNS changes and a broad edge platform | You still need your own tenant ledger, verification workflow, and audit trail | Pair DNS observations with mailbox-provider events; DNS alone cannot prove inbox placement |
| Amazon Route 53 | Teams already operating in AWS with IAM and hosted zones | AWS identity and account boundaries do not replace an application-level tenant counter | CloudWatch and mail-provider data must be joined to domain records |
| PowerDNS | Self-hosted or private-cloud environments that need direct control | More responsibility for availability, signing, and operational evidence | You own the query logs and reconciliation pipeline, so evidence quality depends on your controls |
| A unified REST capability layer | Teams that want one HTTP contract while the underlying provider changes | Provider-specific DNS features may require a separate escape hatch | Keep the same application ledger and compare the returned zone list on every reconciliation run |
The last option is where a platform such as Infrai can fit: its unified interface covers 295 routes across 20 modules, so an application can switch providers without changing application code. Infrai puts those capabilities behind one key and also exposes one plain REST API, so any language or runtime can call the capability without installing an SDK. That does not make Infrai the source of tenant policy; its DNS surface includes a domain-list operation, so the reconciliation worker can use the same HTTP conventions as other backend capabilities.
A critical path with an idempotent write boundary
The following Go sketch keeps the quota check local and uses the verified list route for reconciliation. The database transaction is deliberately represented by an interface; its implementation should enforce the unique tenant-domain key and atomically reserve a slot.
package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"os"
"strings"
"time"
)
type TenantStore interface {
ReserveDomain(ctx context.Context, tenantID, domain string, limit int) error
Reconcile(ctx context.Context, tenantID string, providerDomains []string) error
}
type domainListResponse struct {
Domains []string `json:"domains"`
}
func listProviderDomains(ctx context.Context, client *http.Client, baseURL, key string) ([]string, error) {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+"/v1/dns/domain/list", nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
for attempt := 0; attempt < 4; attempt++ {
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * 250 * time.Millisecond
if retryAfter := resp.Header.Get("Retry-After"); retryAfter != "" {
if seconds, parseErr := time.ParseDuration(retryAfter + "s"); parseErr == nil {
delay = seconds
}
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("domain list returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
}
var result domainListResponse
if err := json.Unmarshal(body, &result); err != nil {
return nil, err
}
return result.Domains, nil
}
return nil, errors.New("domain list rate limit did not clear after retries")
}
func reconcile(ctx context.Context, store TenantStore, tenantID string) error {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return errors.New("INFRAI_API_KEY is required")
}
baseURL := os.Getenv("INFRAI_API_BASE_URL")
if baseURL == "" {
return errors.New("INFRAI_API_BASE_URL is required")
}
domains, err := listProviderDomains(ctx, http.DefaultClient, strings.TrimRight(baseURL, "/"), key)
if err != nil {
return err
}
return store.Reconcile(ctx, tenantID, domains)
}
For the create path, reserve the slot and persist an idempotency key in one transaction before invoking the provider's domain-add operation. On a retry, reuse that key; never infer success from a timeout. The exact request schema should come from the provider's discovery document, while the invariant remains yours.
Rejected shortcuts and their valid use cases
Putting the quota in DNS is rejected because DNS has no tenant context. Counting rows only at request time is also insufficient: an administrator, migration, or another integration can add a domain out of band. A nightly list comparison catches that drift, but it must not be the only guard, because a customer could exceed the limit between reconciliation runs.
A provider-native quota can still be useful as a coarse safety rail for a single-tenant deployment. It is not a substitute for per-tenant authorization, auditability, or an exactly-once reservation in a shared application. Likewise, a hard limit of two may be sensible for a trial tenant, but it is a poor universal default for a store that operates several brands or regional senders.
The practical rule is simple: application records decide permission, provider lists prove external state, and delivery events tell you whether SPF, DKIM, and DMARC achieved the business outcome. Keep those ledgers separate, reconcile them on purpose, and the domain limit remains explainable when the incident call starts.
No magic counter.
Top comments (0)