DEV Community

KasimirBerg5341
KasimirBerg5341

Posted on

Per-Tenant Email Authentication: What SPF, DKIM and DMARC Setup Costs to Operate

Use one platform-owned zone, publish three TXT records per tenant subdomain, and let the mail service confirm the domain instead of trusting your own resolver. For a marketplace that gives every seller its own sending subdomain, that is the least complex setup that survives contact with real receivers: SPF authorises the hosts that may send, DKIM signs each message, DMARC tells receivers what to do when those two disagree. All three are TXT records. There is no SPF record type — the experimental type 99 was deprecated by RFC 7208.

The DNS writes are close to free: three records per tenant, written once through your provider's API, rewritten only on a key rotation.

The recurring bill is the reporting side of DMARC, and almost nobody sizes it before switching it on.

The bill is aggregate reports, not DNS records

A DMARC record with a rua tag is a standing subscription. Every receiver that honours it mails you an XML aggregate report once per reporting interval, and the default interval in RFC 7489 is 86400 seconds — one report per domain, per receiver, per day. The domain count is the part that hurts in a marketplace, because each tenant subdomain that publishes its own _dmarc record is a separate subscription.

The arithmetic is boring and worth doing before you commit:

from dataclasses import dataclass

@dataclass(frozen=True)
class ReportLoad:
    tenant_domains: int
    reporting_receivers: int          # distinct receivers that actually send you XML
    interval_seconds: int = 86400     # DMARC "ri" default, RFC 7489

    def reports_per_day(self) -> float:
        per_domain_per_receiver = 86400 / self.interval_seconds
        return self.tenant_domains * self.reporting_receivers * per_domain_per_receiver

    def rows_per_year(self, avg_rows_per_report: int) -> float:
        return self.reports_per_day() * 365 * avg_rows_per_report

load = ReportLoad(tenant_domains=4000, reporting_receivers=8)
print(load.reports_per_day())                    # 32000.0
print(load.rows_per_year(avg_rows_per_report=6)) # 70080000.0
Enter fullscreen mode Exit fullscreen mode

Substitute your own numbers; mine are assumptions, not measurements. The shape of the result holds anyway: reports scale with tenants multiplied by receivers, and the row count — one row per source IP per disposition — scales with how many relays and forwarders touch your mail. Seventy million rows a year is a real table with real indexes, and it grows every time sales onboards a cohort.

Two levers move that dominant term, and only two. Publish _dmarc on the parent sending domain and let tenant subdomains inherit policy through sp, which collapses thousands of subscriptions into one, at the price of losing per-tenant attribution in the reports. Or keep per-tenant records and drop the raw XML early, which is the retention question I come back to at the end.

I lean toward inheritance for tenants under the platform's own domain, and per-tenant records only for sellers who send under their own brand. That is a judgement call about which failures you need to see, not a cost optimisation.

What does each of SPF, DKIM and DMARC actually have to cover?

SPF covers the envelope sender's hosts, and it has a hard budget: RFC 7208 caps a policy evaluation at ten DNS-querying mechanisms, and exceeding it is a permerror, which receivers may treat as a failure rather than a soft warning. Chained include: statements are how platforms blow through that budget without noticing — your relay's include expands to two lookups today and five after they add a region. The other trap in the same document: a domain may publish exactly one SPF record, so a tenant who already has one from a CRM will produce a permerror the moment your onboarding adds a second.

DKIM covers the message itself. The public key lives at selector._domainkey.<domain> as TXT, the private key signs headers and body at send time, and the selector is what lets you rotate without a flag day. One practical detail bites during automation: a 2048-bit RSA key does not fit in a single DNS character-string, because RFC 1035 caps each string at 255 octets, so the record has to be published as multiple concatenated strings. Some provider APIs do that for you. Some return the record looking fine and hand back something receivers can't parse, which is why the verification step later in this article checks from the outside rather than from your own zone file.

DMARC covers the disagreement. It ties the authenticated identifiers back to the visible From domain — alignment, relaxed or strict via adkim and aspf — and publishes what receivers should do when alignment fails, plus where to mail the evidence.

That last word is the useful one. Evidence.

A policy of p=none changes nothing about delivery; it only turns the reporting on, and starting anywhere else is how a marketplace discovers its own forgotten senders by losing their mail. Policy discovery also explains why subdomain strategy is a design decision and not a detail: if a receiver finds no _dmarc record for seller-4821.mail.example-market.com, it walks up to the organizational domain and applies sp from there, with the organizational domain determined through the Public Suffix List.

Choosing between customer-owned and platform-owned zones

This is the axis that decides everything else, including the deliverability you can promise in a contract.

Question Platform-owned zone Customer-owned zone
Who writes the records your automation, in seconds the seller's admin, in days
Reputation blast radius shared organizational domain isolated per seller
Alignment control strict alignment inside your own tree depends on the seller's existing records
Failure you inherit one abusive tenant drags every tenant down onboarding stalls; conflicting SPF records
Report plumbing one rua mailbox, one parser external destination verification per domain

The right-hand column has a clause worth reading twice. If a seller publishes rua=mailto:dmarc@reports.example-market.com on their own domain, RFC 7489 requires your reporting domain to authorise it with a <their-domain>._report._dmarc.reports.example-market.com TXT record before conformant receivers will send anything. Skip that and the reports quietly never arrive — no error, no bounce, an empty dashboard that looks like clean mail.

Platform-owned zones make the reputation coupling explicit instead of accidental. Every tenant subdomain shares one registrable domain, so receivers that score reputation at the organizational level see one sender, and a single fraudulent seller's campaign lands on everybody's transactional mail. Submitting the domain to the Public Suffix List splits that reputation boundary, but it's a slow, manually reviewed process with long cache tails in downstream software, so treat it as a quarter of lead time rather than a sprint task. My honest position is that I'm not sure the PSL route pays for itself below a few thousand active senders; the evidence that would settle it is your own complaint-rate data segmented by tenant, which most platforms don't retain long enough to answer the question.

Stick with customer-owned zones when sellers send under their own brand and their deliverability is their own asset. Go platform-owned when the platform is the sender of record and onboarding time is the metric you're graded on.

An idempotent API call per tenant, then verification from the mail side

The write path should be a pure function of tenant state, so that a retry after a timeout produces the same zone contents rather than a duplicate record. Upsert semantics, one record set per tenant, no read-modify-write races.

import time
from dataclasses import dataclass

@dataclass(frozen=True)
class TenantSending:
    subdomain: str                 # seller-4821.mail.example-market.com
    dkim_selector: str             # rotate by publishing a new selector, never by editing one
    dkim_public_key: str

    def txt_records(self) -> dict[str, str]:
        return {
            self.subdomain: "v=spf1 include:relay.example-market.com -all",
            f"{self.dkim_selector}._domainkey.{self.subdomain}":
                f"v=DKIM1; k=rsa; p={self.dkim_public_key}",
            f"_dmarc.{self.subdomain}":
                "v=DMARC1; p=none; adkim=s; aspf=s; "
                "rua=mailto:dmarc@reports.example-market.com",
        }

def publish(zone, tenant: TenantSending) -> None:
    for name, value in tenant.txt_records().items():
        zone.upsert_txt(name=name, value=value, ttl=300)   # same input, same resulting state

def wait_for_sending_domain(mail, domain: str, deadline_s: int = 900) -> str:
    # the mail service resolves from its own network; a local dig only proves your resolver
    started = time.monotonic()
    while time.monotonic() - started < deadline_s:
        state = mail.verify_domain(domain)                 # pending | verified | failed
        if state != "pending":
            return state
        time.sleep(30)
    return "pending"
Enter fullscreen mode Exit fullscreen mode

Verification belongs to the sending service, not to your resolver, and the reason is caching rather than distrust. A negative answer from before you wrote the record stays cached for the interval derived from the zone's SOA under RFC 2308, so the first lookup after a successful write can legitimately return nothing for several minutes; a low TTL on the records themselves does not shorten that window, because the negative answer was cached against the name that didn't exist yet. Ship the onboarding flow with a pending state and a deadline, and treat the deadline as a product decision.

Partial failure is the interesting error case, and it is the one worth a test. If the DKIM record lands and the SPF record does not, messages still go out, still get signed, and still fail alignment at strict aspf, which means a tenant in a half-configured state produces exactly the symptom — intermittent placement in spam — that support tickets are worst at describing. Store the record set as one row with a version and a verified-at timestamp, emit a metric per tenant for the gap between write and verification, and alert on the count of tenants stuck pending rather than on individual failures. That last choice is what keeps the on-call rotation sane when a provider has a slow afternoon.

What we deliberately stop keeping, and what it costs later

Aggregate XML is evidence, and evidence ages badly. The policy I would defend in a review: keep raw reports for 30 days, then roll them into daily counters keyed by tenant, source IP, and disposition, and keep those for two years. Storage drops by more than an order of magnitude, and the aggregates still answer the questions that matter month over month — which sellers are failing alignment, which forwarders break DKIM, whether a policy change moved the pass rate.

Failure reports are a different decision. The ruf stream carries message-level detail, including recipient addresses and subject lines, which is why RFC 7489 spends a section on the privacy consequences and why many large receivers never send them at all. I'd leave ruf unset and accept the blind spot rather than build a retention regime around other people's mail.

Here is the cost of that choice, stated plainly so nobody is surprised during an incident: once the raw XML is gone, you can answer which source sent how much and how it was dispositioned, but you cannot answer which message. A seller who opens a ticket about one specific customer email from six weeks ago will get an aggregate answer, and that is sometimes not good enough. The mitigation is cheap and worth wiring in advance — a per-tenant switch that pins raw retention for that tenant, flipped by support without a deploy.

Three records, one decision axis, one bill that grows with tenants rather than with volume. The rest is retention policy you should choose on purpose.

Further reading

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.