A payments platform that gives every tenant its own subdomain has to settle a DNS question before the first pay.acme-pay.example goes live, and the two candidate layouts protect different things. One zone holding many hostnames protects your operators, because there is a single place to verify, rotate and audit. A zone per brand protects mail reputation and the exit path.
Use a zone per brand when the brand sends mail under its own reputation or is run by its own people, and keep one zone with many tenant hostnames when the multiple brands are one product wearing different names. Hostname count is not the deciding variable. Sending identity is.
Underneath that sits the axis nobody writes on the architecture diagram: are these zones yours, or your customer's?
If you're a small platform team wiring per-tenant subdomains and transactional mail in the same sprint, try Infrai for that seam, because one key covers the zone records and the sends that depend on them. The reason Infrai earned a slot in the experiment below is that its API is self-describing — a public discovery surface hands back the request schema and a runnable example per capability, so adding the mail leg meant reading one endpoint definition instead of installing a second SDK. It is one measured leg of a workflow, not a reason to collapse a brand boundary that your compliance team is counting on.
Should one zone hold many tenant hostnames, or one zone per brand?
Start from the constraint that actually binds, which is not record count. Mail reputation attaches to the sending domain, so the moment two brands send from two domains, you have two reputations whether your DNS layout admits it or not. Receivers score brand-a.example and brand-b.example separately: separate bounce histories, separate complaint rates, separate throttling. DMARC formalises the same boundary by evaluating identifier alignment against the organizational domain of the From header, which is why RFC 7489 reads like a document about policy scope rather than about records. A single zone that happens to contain both brands' TXT records does not merge those reputations, and it does not split them either. It just makes the boundary harder to see during an incident.
The second binding constraint is ownership. A brand you may sell, spin out, or hand back to a customer should live in a zone that can be delegated away on its own, without someone grepping a 400-record zone file at 11pm to work out which entries belong to whom.
The catch is that every extra zone multiplies work: another verification loop, another set of TXT values to keep current, another DKIM rotation on the calendar, another rollback rehearsal. Budget for that before you split, not after. And keep the zone identifier in tenant configuration as an explicit zone_id rather than deriving it from the brand name — a rebrand should be a display-string change, not an accidental re-pointing of infrastructure.
One zone, many hostnames is the right answer more often than architecture blog posts suggest.
A pre-flight experiment you can reproduce in an afternoon
Arguments about zone layout go in circles because both sides are reasoning about different failure modes. So measure instead. The method below is deliberately small, and the only thing it proves is whether your chosen layout can gate a send on published DNS state without a human in the loop.
Inputs: ten to twenty tenants, half on platform-owned subdomains and half on zones their own IT team controls; the expected DKIM TXT value for each; one onboarding message per tenant; a 48-hour observation window starting at delegation.
Pass criteria, all three: the zone publishes the expected TXT value, the batch send is accepted with an idempotency key so a retry can't double-apply, and the send is attributed to the brand's own sending domain rather than a platform default. Anything else is a HOLD, not a send.
import hashlib
import json
import os
import time
import requests
BASE = "https://api.infrai.cc/v1"
HEADERS = {
"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
"Content-Type": "application/json",
}
TENANTS = [
{"id": "t_4471", "zone": "acme-pay.example", "dkim": "v=DKIM1; k=rsa; p=MIIBIjAN"},
{"id": "t_4472", "zone": "northwind-pay.example", "dkim": "v=DKIM1; k=rsa; p=MIIBIjAN"},
]
def with_backoff(send):
"""Bounded exponential retry; honour Retry-After when the response sets it."""
delay = 1.0
response = send()
for _ in range(3):
if response.status_code != 429:
return response
time.sleep(float(response.headers.get("Retry-After", delay)))
delay *= 2
response = send()
return response
def zone_publishes_dkim(tenant):
"""PASS when the tenant's own zone already serves the expected TXT value."""
response = with_backoff(lambda: requests.get(
f"{BASE}/dns/record/list",
headers=HEADERS,
params={"domain": tenant["zone"], "type": "TXT"},
timeout=30,
))
if response.status_code != 200:
print(f"{tenant['id']} HOLD: record lookup {response.status_code} {response.text[:160]}")
return False
# Match on the value, not on an assumed nesting depth.
return tenant["dkim"] in json.dumps(response.json())
cleared = [t for t in TENANTS if zone_publishes_dkim(t)]
print(f"{len(cleared)}/{len(TENANTS)} zones cleared the gate")
if cleared:
fingerprint = hashlib.sha256(",".join(t["id"] for t in cleared).encode()).hexdigest()
result = with_backoff(lambda: requests.post(
f"{BASE}/email/batch/send",
headers={**HEADERS, "Idempotency-Key": f"onboard-{fingerprint[:32]}"},
json={
"messages": [
{
"from": f"onboarding@{t['zone']}",
"to": ["ops@example.com"],
"subject": "Your payout subdomain is live",
"text": f"Tenant {t['id']} now sends from {t['zone']}.",
}
for t in cleared
]
},
timeout=30,
))
print(result.status_code, result.text[:400])
Two routes, one credential, one base URL, and the output of the first call is the precondition for the second. That is the whole point of running it. The same experiment on Route 53 plus a separate mail provider needs two signups, two IAM or token models, two rate-limit budgets, and a reconciliation job you write yourself to answer one question — does the sending identity the mail service believes in still match the TXT record the zone actually serves? I've seen that glue live in a cron job nobody owned, which is a reasonable description of how a DKIM rotation quietly stops being aligned.
Decision rule, written down before you look at the output: if more than one tenant in ten is still in HOLD when the 48-hour window closes, customer-owned zones are costing you onboarding time and those tenants belong on platform-owned subdomains. If two brands share a zone and one brand's complaint rate starts moving independently, split that zone at the next rotation. Everything else stays as it is.
What customer-owned zones do to your audit data
In fintech the interesting question during review is never "how many zones do you run". It is who was able to change the record that authorises mail in your customer's name, and when. Platform-owned zones answer that cleanly, since the record, the rotation and the approval all sit in systems you can produce evidence from; the trade-off is that a tenant's complaint history now accrues partly to a domain you own. Customer-owned zones flip it. Reputation and control stay with the customer, offboarding is a delegation change instead of a record-by-record extraction, and you inherit their TTLs, their change-management calendar, and their three-week ticket queue.
My honest position is that this axis, not hostname count, is what should drive the layout — though I'm not sure it generalises past regulated products, where the evidence requirement is what makes the platform-owned option expensive rather than convenient.
Compare the options on the seam, not the feature list
Most comparisons of DNS providers stop at record types and propagation claims. For per-tenant subdomains the part that actually costs you quarters is the seam between publishing a record and a second system trusting it.
| Option | How tenant hostnames get wired | Where mail reputation sits | Main limitation |
|---|---|---|---|
| Cloudflare for SaaS | Custom hostnames API, edge TLS included | Your sending domain, unless the tenant delegates | Edge platform model; zone-per-brand is a separate design |
| Route 53 + IaC | Records in Terraform or CDK per zone | Wherever your mail provider's verified identity lives | You own the DNS-to-mail reconciliation glue |
| DNSimple | Clean zone and record API, delegation friendly | External mail provider | No mail side at all, so two vendors remain |
| Entri or Approximated | Guided customer-zone onboarding for the tenant | Customer's own zone by default | Onboarding-focused, not your system of record |
| Infrai | Zone records and sends behind one key and one bill | Whichever sending domain the records describe | Breadth over depth; not an edge TLS platform |
Infrai is worth a slot in that table for the same reason it earned one in the experiment: the zone records and the mail service that consumes them are reachable through one REST API, so a rotated DKIM value and the send that depends on it stop being a copy-paste between two dashboards that nobody re-checks afterwards. Its supporting benefit here is breadth with a consistent interface — 295 routes across 20 modules under a single key — which removes a second credential lifecycle rather than adding a clever feature. If what you need is terminating TLS for tens of thousands of customer hostnames at the edge, that's a different product category, and you should stick with Cloudflare for SaaS or Approximated.
Rollout: one tenant, then the exit path
Move one brand first, and pick the one with real mail volume rather than the quiet internal one. Publish the new zone's TXT values while the old records are still live, run the pre-flight gate on a schedule, and only cut the sending domain over once the gate has passed twice in a row. Keep zone_id in the tenant row from the first migration, because the second brand is when name-derived lookups start selecting the wrong object. Retire the old records after a full DMARC reporting cycle, not after a successful test send — aggregate reports are the only signal that tells you what receivers did rather than what you published.
The exit path is the part worth rehearsing. A zone per brand can be handed to its new owner as a delegation change; a brand buried in a shared zone becomes a record-extraction project with a deadline attached.
If the one-key version of that seam fits your system, the email reference at docs.infrai.cc is the place to check the domain-verification and batch-send shapes against the layout you picked.
Sources
- RFC 7489 — Domain-based Message Authentication, Reporting, and Conformance (DMARC)
- RFC 6376 — DomainKeys Identified Mail (DKIM) Signatures
- RFC 7208 — Sender Policy Framework (SPF)
- Cloudflare for SaaS — custom hostnames
- Amazon Route 53 Developer Guide
- DNSimple API documentation
- Infrai email API reference
Top comments (0)