DEV Community

loganpierce2073
loganpierce2073

Posted on

React to Domain Verification Webhook Updates for Tenant State and Customer Email

Short answer: treat a domain-verification webhook as an untrusted signal, record it idempotently, reconcile the tenant against the DNS state you can prove, and send the customer email from the committed transition rather than from the webhook handler.

In a marketplace, a tenant can add a sending domain while sellers and buyers are already waiting for receipts, payout notices, and password mail. The dangerous gap is not the DNS lookup itself. It is the drift between what the onboarding workflow believes is published and what the authoritative record actually says. A retry, a delayed resolver, or an old webhook can otherwise move a tenant to verified and trigger mail that still fails authentication.

The decision record: intent is not evidence

I use three durable facts for this workflow: the tenant's intended record set, the latest observed verification signal, and the state transition that was accepted. They have different lifecycles. Intent is what the tenant asked us to publish; evidence is what the verifier observed; the transition is our business decision. Collapsing them into one mutable row makes an audit trail impossible.

The invariant is strict: no customer-facing “mail is ready” message can be emitted unless a committed transition says the required SPF, DKIM, and DMARC checks passed. The transition also carries the event identifier and a hash of the observed record set, so a support engineer can explain which evidence was accepted without trusting a queue's delivery order.

Option Useful property Failure boundary
Update state directly in the webhook handler Low latency and little code Duplicate or out-of-order delivery can send mail twice or accept stale evidence
Put events on a queue and let any worker mutate the tenant Easy horizontal processing Without a compare-and-set transition, two workers can publish contradictory state
Store the event, reconcile, then project state and email Replayable, auditable, and idempotent Requires a small state machine and an outbox poller

The third option is the one I would record in an architecture decision record. It makes the failure boundary visible: a webhook may be lost or repeated, but a committed event and its resulting transition can be replayed. Exactly-once delivery is not a property I can demand from the network; exactly-once effects are a property I can enforce at the database boundary.

How should a webhook update tenant state and email the customer?

The handler should do very little. Verify the signature using the provider's documented mechanism, reject malformed payloads, and insert the event under a unique (source, event_id) key. It should not call an SMTP service or decide that DNS is correct while holding an HTTP request open.

The worker loads the tenant's current intent, resolves the relevant DNS names through a resolver with an explicit freshness policy, and evaluates the required records. A passing observation is still just an observation until a conditional transaction advances the tenant from pending_dns to mail_enabled. The transaction inserts an outbox row in the same commit; the email sender consumes that row and records its own idempotency key. In practice, the record needs enough context to explain a surprising result later: the normalized TXT values, selector names, resolver identity, observation timestamp, policy version, and the event that caused the check. If a tenant edits DKIM while an earlier webhook is in flight, the worker must compare the event's observed hash with the current intent, leave the state pending when they disagree, and schedule another check rather than “helpfully” guessing which version won. I've seen teams omit that comparison because the happy path is only a few lines; the missing comparison is where a receipt can leave the system before the DNS change has propagated.

Then the queue retries.

Here is the critical path in Go. The interfaces are deliberately generic so the policy remains testable and no commercial API becomes the architecture.

package onboarding

import (
    "context"
    "crypto/sha256"
    "encoding/hex"
    "fmt"
)

type DomainEvent struct {
    Source  string
    ID      string
    Tenant  string
    Records []string
}

type Store interface {
    InsertEvent(ctx context.Context, source, id string, e DomainEvent) (alreadySeen bool, err error)
    Begin(ctx context.Context) (Tx, error)
}

type Tx interface {
    LoadIntent(ctx context.Context, tenant string) (string, error)
    AdvanceIfPending(ctx context.Context, tenant, evidenceHash string) (bool, error)
    AddOutbox(ctx context.Context, key, tenant string) error
    Commit(ctx context.Context) error
    Rollback(ctx context.Context) error
}

func Reconcile(ctx context.Context, s Store, e DomainEvent) error {
    seen, err := s.InsertEvent(ctx, e.Source, e.ID, e)
    if err != nil || seen {
        return err
    }

    tx, err := s.Begin(ctx)
    if err != nil {
        return err
    }
    defer tx.Rollback(ctx)

    if _, err = tx.LoadIntent(ctx, e.Tenant); err != nil {
        return err
    }
    h := sha256.Sum256([]byte(fmt.Sprint(e.Records)))
    evidence := hex.EncodeToString(h[:])
    changed, err := tx.AdvanceIfPending(ctx, e.Tenant, evidence)
    if err != nil {
        return err
    }
    if changed {
        if err = tx.AddOutbox(ctx, "domain-ready:"+e.Tenant, e.Tenant); err != nil {
            return err
        }
    }
    return tx.Commit(ctx)
}
Enter fullscreen mode Exit fullscreen mode

The unique event key handles delivery retries; the conditional transition handles races; the outbox key handles email retries. Those are separate idempotency scopes, and naming them separately has prevented more reconciliation bugs in ledger work than adding another queue ever did. One small sentence can save a week of forensic work.

What does “verified” mean under SPF, DKIM, and DMARC?

Verification is a policy decision, not a string comparison. SPF authorizes sending infrastructure for a domain. DKIM proves that a message was signed with a key published under the domain. DMARC tells receivers how to evaluate alignment and what to do with failures; its reporting model is specified in RFC 7489. A marketplace should write down whether it requires all three before enabling mail, which selectors are expected, and how long an observation remains fresh.

DNS is eventually consistent. A resolver can return a cached answer after a tenant has changed a record, while another resolver sees the new value. Store the resolver context, observation time, and normalized records, then re-check on a schedule when the evidence is stale. Do not silently reinterpret a timeout as a pass. Your mileage may vary across recursive resolvers, and I am not sure a single global TTL policy fits every registrar, so the freshness window belongs in configuration and in the audit record.

The email itself should say what changed and what did not. “Your domain passed the configured checks” is safer than promising universal inbox delivery, because DMARC alignment does not control recipient policy, reputation, or content filtering.

Rejected option and its valid use case

I would reject a design that lets the webhook handler set mail_enabled=true and immediately send the customer message. It is attractive for a prototype, and it is valid when the event is only a best-effort dashboard hint with no financial or customer-facing consequence. It is not suitable when a stale DNS answer can cause a marketplace to send receipts from an unauthenticated domain.

The catch is operational cost: the event ledger, reconciliation worker, and outbox add tables and metrics. Keep them when auditability, replay, or regulated communication matters. For an internal sandbox that never sends external mail, a single projection may be enough. Choose based on the consequence of drift, not on the number of moving parts in a diagram.

References

Top comments (0)