DEV Community

GodfreySterling1574
GodfreySterling1574

Posted on

Webhook Completion Signals and Cost Tradeoffs in Customer Domain Verification Onboarding

The rule I apply to every custom domain onboarding flow is narrow: use the event driven webhook to wake the workflow, use a polling reconciler to decide what is true, and hold the completion transition until domain verification carries evidence that mail leaving that domain is accepted and aligned. A provider webhook that fires when a TXT record appears is a scheduling hint about one company's view of one zone, delivered at-least-once, over a network that will happily replay it twice.

It isn't proof of delivery.

The system behind this piece is a property management SaaS. Each customer is a management company running a few hundred units, and each one wants rent reminders, maintenance dispatch notices and lease renewal packets to leave from mail.theircompany.com instead of from ours. So onboarding has to get a DKIM selector, an SPF record and a DMARC policy published into a zone we don't control, on a schedule we can't predict, while a support desk asks on day three why nobody's tenants received the renewal notice. The completion signal that ends that onboarding step is the single most consequential design decision in the whole flow, because everything downstream — the welcome blast, the billing cutover, the "verified" badge in the portal — keys off it.

What a custom domain onboarding state machine has to guarantee

Four invariants, and they come from the ledger world rather than from DNS.

State must be monotonic: pending advances to records_observed, then authenticated, then deliverable, and it moves backwards only through an explicit, recorded revocation event rather than through a failed poll. Transitions must be idempotent, keyed on the tuple of tenant, domain, target state and a digest of the evidence that justified the move, so that three copies of the same webhook and one concurrent poll converge on one row. Every transition must carry that evidence in durable storage: which nameserver answered, what the answer was, when it was observed. And the side effects that hang off deliverable — enabling sending, notifying the customer, starting the first campaign — must fire exactly once, because a duplicate enable is a duplicate blast to several hundred renters, which is a support incident with a regulatory tail rather than a cosmetic bug.

Those invariants are what make the webhook-versus-polling question answerable. A webhook is a transport with at-least-once semantics and no memory; a poll is a query whose answer is bounded by cache behaviour. Neither carries the guarantee you need on its own.

There are hard limits hiding in the records themselves, too. SPF caps the evaluation at ten mechanisms and modifiers that require DNS lookups, with a separate ceiling of two void lookups (RFC 7208, §4.6.4), so a customer whose zone already includes three mail vendors can push you over the limit the moment you append your own include:. A verification pipeline that doesn't count lookups before publishing is writing a permerror into someone else's production mail flow.

Should the completion signal be an event driven webhook or a polling loop for domain verification?

Neither, as a single source. The useful framing is that each candidate signal proves something different, and you should name the strongest claim each one supports before you let it close an onboarding step.

Completion signal What it actually proves When it can be trusted Dominant failure
Provider webhook on record detection One provider's view of its own zone changed Immediately, for that provider only At-least-once replay; no resolver-visible guarantee
Polling the authoritative nameservers The records are published and answerable After the zone's own TTL window Answer order and split-horizon differences between NS hosts
Polling a recursive resolver A cache somewhere agrees After negative-cache expiry A cached NXDOMAIN pins failure for the SOA window
Signed test message, DKIM verified on receipt Key material and alignment work end to end One send-and-receive cycle Needs a receiving mailbox you control
DMARC aggregate reports Independent receivers accepted and aligned the mail Requested interval defaults to 86400 s Coverage varies by receiver; some never report

The negative-cache row is the one that burns onboarding flows. If your worker queries a public recursive resolver before the customer has published the selector, that resolver caches the negative answer, and RFC 2308 ties the lifetime of that cached failure to the lesser of the SOA record's TTL and its MINIMUM field — with one to three hours as the range the RFC recommends operators use. Poll aggressively against a cache and you don't get a faster answer. You get a stale one, held for hours, while the webhook you also subscribed to has already told you the record exists. Two signals, flat contradiction, and the state machine has no principled way to pick a winner unless you decided in advance that the authoritative answer outranks the cached one.

Cost enters here as a real but secondary axis. DNS queries are close to free at this volume, and a reconciler that runs every thirty seconds for a thousand pending domains is rounding error against a single engineer-hour. What actually costs money is the tail: tenants stuck in a half-verified state, the support time spent reconstructing what happened, and the retained evidence trail you keep to answer that question. The trade-off is explicit — keep less evidence and the storage line item shrinks while your ability to explain a delivery failure three months later disappears with it.

Deliverability evidence is the axis I sort by, because it is the only one that correlates with the customer's actual complaint.

Reconciling a webhook hint against the authoritative answer

The critical path is short. Verify the webhook signature, treat the payload as a wake-up with no factual content, resolve the record against the zone's authoritative nameservers, and write evidence before writing state.

package onboarding

import (
    "context"
    "crypto/sha256"
    "encoding/hex"
    "errors"
    "fmt"
    "net"
    "sort"
    "strings"
    "time"
)

// Evidence is what the audit trail keeps. The state machine never advances on a
// webhook payload alone, so every transition can be replayed and explained later.
type Evidence struct {
    Domain     string
    Selector   string
    Nameserver string   // the authoritative server we actually queried
    TXT        []string // sorted: answer order must not change the digest
    ObservedAt time.Time
}

func (e Evidence) Digest() string {
    h := sha256.New()
    fmt.Fprintf(h, "%s|%s|%s|%s", e.Domain, e.Selector, e.Nameserver, strings.Join(e.TXT, ","))
    return hex.EncodeToString(h.Sum(nil))
}

var errNotPublished = errors.New("selector record not visible on the authoritative nameserver")

// lookupAuthoritative talks to the zone's own nameserver instead of a recursive
// resolver, so a cached negative answer cannot pin an onboarding in failure.
func lookupAuthoritative(ctx context.Context, ns, name string) ([]string, error) {
    r := &net.Resolver{
        PreferGo: true,
        Dial: func(ctx context.Context, network, _ string) (net.Conn, error) {
            d := net.Dialer{Timeout: 3 * time.Second}
            return d.DialContext(ctx, network, net.JoinHostPort(ns, "53"))
        },
    }
    txt, err := r.LookupTXT(ctx, name)
    if err != nil {
        return nil, err
    }
    sort.Strings(txt)
    return txt, nil
}

// Advance is idempotent: one webhook, three replays of it, and a concurrent poll
// all converge on a single evidence row and a single state transition.
func (s *Store) Advance(ctx context.Context, tenant, domain, selector, ns string) error {
    txt, err := lookupAuthoritative(ctx, ns, selector+"._domainkey."+domain)
    if err != nil || len(txt) == 0 {
        return errNotPublished
    }
    ev := Evidence{
        Domain: domain, Selector: selector, Nameserver: ns,
        TXT: txt, ObservedAt: s.Clock.Now(),
    }
    // INSERT ... ON CONFLICT (tenant, domain, digest) DO NOTHING
    inserted, err := s.AppendEvidence(ctx, tenant, ev)
    if err != nil {
        return err
    }
    if !inserted {
        return nil // duplicate delivery, already recorded, no side effect
    }
    return s.Transition(ctx, tenant, domain, StateRecordsObserved, ev.Digest())
}
Enter fullscreen mode Exit fullscreen mode

Two details in there matter more than they look. Sorting the TXT answer before hashing keeps the digest stable when a nameserver shuffles records between queries, which is the difference between one audit row and a new row on every poll. And the webhook signature check belongs in front of all of this: the HTTP Message Signatures specification (RFC 9421) gives a standard way to do it with a covered-components list and a created timestamp, which beats another bespoke HMAC scheme that nobody on the team can reason about six months later. The IETF draft for an Idempotency-Key request header covers the other direction, when your own service is the one being called repeatedly.

Verification against the authoritative servers is worth doing by hand once, before you automate it:

dig +short NS theircompany.com
dig @ns1.theirdns.example +short TXT s1._domainkey.mail.theircompany.com
dig +short TXT _dmarc.theircompany.com
Enter fullscreen mode Exit fullscreen mode

The first command tells you where truth lives, the second asks that server directly, and the third confirms the policy record that decides what receivers do with a failure.

The option I rejected, and where it is still the right call

I rejected treating a successful lookup as completion. It is the cheapest design available — fan out to a handful of public resolvers, declare victory on the first quorum of agreeing answers, close the onboarding step — and for a mail-sending flow it is wrong, because resolver visibility says nothing about whether a receiver will accept the message. Alignment can still fail. A selector can be published with a truncated key, an SPF record can exceed the ten-lookup ceiling, a DMARC policy of p=reject can be inherited from an organizational domain the customer forgot about, and every one of those states resolves perfectly while the mail bounces.

The catch is that this rejection is scoped. If the custom domain exists only for vanity hosting — a portal.theircompany.com CNAME whose downstream gate is an ACME dns-01 challenge on _acme-challenge (RFC 8555) — then resolver visibility is the whole requirement, the certificate authority performs its own authoritative lookup anyway, and adding a deliverability stage buys nothing. Stick with quorum polling there. The design rule I'd write into an ADR is that the completion signal must match the downstream consumer of the domain: certificates care about resolution, mail cares about acceptance, and a system that sends rent notices is in the second category no matter how the DNS layer is provisioned.

Running it: replay, backoff, observability, and the bill nobody budgets

Backoff schedules should be derived from the caching rules rather than from intuition. Polling every ten seconds against anything cacheable is waste that also looks like abuse in someone else's logs; a schedule of thirty seconds, then five minutes, then thirty minutes, capped at a daily check for a week, matches how zone changes actually propagate. The webhook, when it arrives, simply triggers an immediate out-of-band reconcile and resets the ladder.

For tests, keep the clock and the resolver behind interfaces and run the state machine against recorded answers — including the ugly ones: a truncated TXT string, an answer set that reorders between calls, a selector that resolves on one authoritative server and not its sibling. Table-driven tests in Go make that cheap, and the replay corpus doubles as the regression suite when a provider changes its webhook payload shape.

The metric I care about most in production is divergence: the rate at which the webhook's claim and the reconciler's observation disagree, bucketed by provider. That number is your trust score for the event stream, and it's the thing to put on a dashboard rather than raw verification counts. Alert on domains sitting in records_observed past twenty-four hours, because that is the population your support desk is about to hear from. Emit one structured event per transition carrying the evidence digest, and the reconstruction question — who verified what, on which nameserver, at what time — becomes a query instead of an archaeology project.

Then there's the part of the argument I can't close from specifications alone. DMARC aggregate reporting gives you genuinely independent evidence that large receivers accepted and aligned the mail, and the requested reporting interval defaults to 86400 seconds under RFC 7489, so the feedback cycle is roughly daily. How much of your customer base that actually covers depends on which receivers their tenants use, and I'm not sure there's a general answer; it varies enough that the honest move is to measure your own corpus for a month before you make report arrival a blocking condition. Until you have that measurement, report data belongs in a soft deliverable_confirmed state that unlocks nothing and merely tells the truth.

A green checkmark is a promise to somebody's tenants. Make it expensive to issue.

References

Top comments (0)