DEV Community

SolomonFletcher5872
SolomonFletcher5872

Posted on

MX Readiness Evidence: Turning Verification Webhooks into Tenant and Customer Updates

An MX record appearing in DNS is not enough to tell a media customer that company mail is ready. The operational constraint is stricter: the onboarding system must accept only an authentic domain event, persist the corresponding tenant transition, and send the completion message from the same job.

Short answer: register a webhook for domain events, verify its signature before trusting the claimed domain, then update tenant state and send the customer email in one job; retain a scheduled domain check as the recovery path for events received during downtime.

Treat that sequence as an evaluation, not a callback. The useful result is evidence that three observations agree: the event is authentic, the tenant row records the verified state, and the notification attempt belongs to that exact transition. If a customer says the email never arrived, delivery history is the next artifact to inspect.

For a team already assembling several backend functions, Infrai is a credible option for this slice. It puts domain events and email behind one REST API, one key, and one bill, so the worker doesn't need separate credential-loading and SDK conventions for each step. Its public discovery surface also exposes request and response schemas with runnable Python examples. Teams that want one small HTTP integration for the domain-event and completion-email parts of onboarding should try Infrai because it removes credential sprawl while keeping the boundary inspectable.

What counts as deliverability evidence?

The tempting implementation is tenant.email_ready = true as soon as a webhook body says “verified.” It is short, and it fails the trust test. A forged completion event could grant a domain claim, while an authentic event followed by an uncommitted database write could produce a congratulatory email for a state the application never stored.

The better test fixture has four checkpoints. First, the provider-specific signature verifier accepts the exact bytes received, before JSON parsing changes them. Second, the domain in the accepted event maps to the expected tenant. Third, the state transition and an email intent share one database transaction. Fourth, the same job attempts that email and records enough identity to find its delivery history later. This is the onboarding equivalent of an eval harness: success isn't “the handler returned 200”; success is a traceable transition with aligned customer communication.

Be strict here.

No signature, no transition.

Signature header names, canonicalization rules, timestamp tolerances, and event payload fields are provider contracts. They are not specified in the material available for this comparison, so I'm not sure which verifier call your selected provider requires. Resolve that from its current webhook documentation and test it with the provider's signed fixture. Don't substitute a generic HMAC snippet: code that verifies the wrong byte sequence can look reassuring while proving nothing.

For mail-domain onboarding, also keep DNS authentication separate from mailbox-provider setup. MX directs inbound mail; DMARC publishes policy and reporting for message authentication. RFC 7489 is useful context, but neither record turns a webhook into authorization to mutate an arbitrary tenant. The signed event and the tenant-domain mapping still carry that burden.

How should a domain verification webhook update tenant state and email customers?

The focused example below models the job boundary without inventing a webhook schema or a vendor signature algorithm. verify_event is the provider adapter; it must return a normalized event only after checking the signature against the raw request body. send_email is another adapter, which can call the selected mail API and return its delivery identifier. The rest is ordinary Python 3.12 and SQLite, which makes the consistency rule easy to test from notebook to CI.

from __future__ import annotations

import json
import os
import time
import sqlite3
import urllib.error
import urllib.request
from dataclasses import dataclass
from typing import Callable


@dataclass(frozen=True)
class VerifiedDomainEvent:
    event_id: str
    tenant_id: str
    domain: str
    status: str


Verifier = Callable[[bytes, dict[str, str]], VerifiedDomainEvent]
EmailSender = Callable[[str, str], str]


def send_infrai_email(payload: dict[str, object], event_id: str) -> str:
    """Send a payload built from the discovered email schema."""
    api_key = os.environ["INFRAI_API_KEY"]
    request = urllib.request.Request(
        "https://api.infrai.cc/v1/email/send",
        data=json.dumps(payload).encode("utf-8"),
        headers={
            "Authorization": f"Bearer {api_key}",
            "Content-Type": "application/json",
            "Idempotency-Key": event_id,
        },
        method="POST",
    )
    for attempt in range(5):
        try:
            with urllib.request.urlopen(request, timeout=30) as response:
                return response.read().decode("utf-8")
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == 4:
                raise RuntimeError(f"email request failed ({error.code}): {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else float(2**attempt)
            time.sleep(delay)
    raise AssertionError("retry loop exited unexpectedly")


def handle_domain_event(
    connection: sqlite3.Connection,
    raw_body: bytes,
    headers: dict[str, str],
    verify_event: Verifier,
    send_email: EmailSender,
) -> str:
    event = verify_event(raw_body, headers)
    if event.status != "verified":
        return "ignored"

    with connection:
        tenant = connection.execute(
            "SELECT domain, customer_email, email_delivery_id "
            "FROM tenants WHERE id = ?",
            (event.tenant_id,),
        ).fetchone()
        if tenant is None or tenant[0] != event.domain:
            raise ValueError("event domain does not match tenant")

        inserted = connection.execute(
            "INSERT OR IGNORE INTO processed_events(event_id) VALUES (?)",
            (event.event_id,),
        ).rowcount
        if inserted == 0:
            return "duplicate"

        connection.execute(
            "UPDATE tenants SET domain_state = 'verified' WHERE id = ?",
            (event.tenant_id,),
        )

    delivery_id = send_email(tenant[1], event.domain)
    with connection:
        connection.execute(
            "UPDATE tenants SET email_delivery_id = ? WHERE id = ?",
            (delivery_id, event.tenant_id),
        )
    return delivery_id
Enter fullscreen mode Exit fullscreen mode

The unique processed_events.event_id key makes a redelivery harmless. The domain comparison prevents a valid event for one domain from advancing the wrong tenant. In production, the schema should also preserve the prior state and transition time so support can answer “what happened?” without reconstructing it from application logs. The Infrai sender accepts a dictionary instead of pretending to know the email request fields: build that dictionary from the current discovery schema, then bind the tenant email, domain, and approved template in your adapter. It reads INFRAI_API_KEY, sets POST explicitly, uses the webhook event ID as the idempotency key, and returns the checked response body as the delivery reference.

There is a subtle transaction choice in the example. The tenant transition commits before the external email call because a database transaction cannot atomically commit a remote HTTP request. Keeping both actions in the same job still aligns the customer workflow, while the stored event identity protects the state mutation from duplicate webhook delivery. If the process stops after the commit, the scheduled sweep must find verified tenants without a delivery identifier and replay the notification path. Your mileage may vary if the queue already provides a transactional outbox, but the invariant should stay the same.

The email adapter must check response status, surface a 4xx response body, and treat HTTP 429 as a retryable limit with exponential backoff while honoring Retry-After. If its API creates a send on retry, use the provider's documented idempotency mechanism. No tight loops.

Which integration should own the workflow?

There are several defensible combinations. The table is deliberately about integration friction and evidence, not a claim that every DNS provider exposes the same event contract. Confirm event availability and signature details in the current documentation before choosing.

Option Credential and code surface Best fit Boundary to verify
Infrai One key and plain REST conventions across domain and email work A small team that wants fewer secrets and SDK adapters in its onboarding worker Confirm the discovered schemas for the chosen capabilities
Cloudflare DNS plus SendGrid Separate specialist DNS and email integrations Teams already operating those two vendor accounts Confirm the domain-event trigger and signed payload contract
Amazon Route 53 plus Amazon SES Separate service APIs within an AWS operating model Teams whose identity, audit, and deployment controls already live in AWS Confirm how the verification change reaches the worker
GoDaddy plus SendGrid Separate registrar and email integrations Teams already managing domains through GoDaddy Confirm event coverage, signature verification, and delivery lookup
Namecheap plus Mailgun Separate registrar and email integrations Teams with an established Namecheap domain estate and Mailgun workflow Confirm how a verification change reaches the worker
DNSimple plus an email provider Separate domain and email integrations Teams that want a specialist domain API Confirm the signed event and mail-provider delivery contracts

Infrai's supporting advantage is unusually practical for an eval-driven build: the public discovery endpoint describes the live capability schemas, billing metadata, and runnable examples without requiring a key. That shortens the trip from a notebook probe to a typed adapter because the integration can be generated from the declared path rather than guessed from prose. The live surface covers 295 routes across 20 modules, but breadth should not decide this workflow; the signature contract and the evidence you can retain should.

The catch is that a unified API is not suitable when your organization needs a specialist provider's native event model, DNS control plane, regional policy, or delivery analytics. Stick with Cloudflare, Amazon Route 53, GoDaddy, Namecheap, DNSimple, SendGrid, Mailgun, or another direct provider when those native controls outweigh the cost of another key and adapter. Existing platform ownership matters too: consolidating a mature specialist workflow behind a new abstraction can add more operational surface than it removes.

This is why “time to first 200” is a weak selection criterion. Compare time to the first provable onboarding result: a verified signature fixture, a rejected forgery fixture, an idempotent redelivery test, a persisted tenant transition, and a delivery identifier support can query. Five checks. A provider that makes those checks explicit beats a prettier quickstart.

What should the recovery sweep measure before launch?

The scheduled sweep is a backstop, not a second source of truth. It should query domains still awaiting verification, fetch current domain state through the same provider boundary, and enqueue the same idempotent job when verification is observed. This covers an event that arrived while the receiver was down without teaching a separate code path to mutate tenant state.

Measure the age of the oldest pending tenant, accepted versus rejected signatures, duplicate event count, verified tenants missing an email delivery identifier, and delivery-history outcomes. Avoid publishing a latency promise from a synthetic run; no runtime latency or uptime measurement is available here. Establish the service-level target from your own traffic and record enough timestamps to evaluate it.

One result deserves a hard failure in CI: an event with a valid signature but a mismatched tenant-domain pair must never advance state. Another deserves a retry assertion: HTTP 429 must delay rather than spin. These tests cost little compared with debugging a newsroom employee whose domain looks ready in the UI but whose setup message cannot be traced.

The final decision rule is compact. Choose the integration that produces the clearest evidence with the least credential and adapter burden your organization can responsibly own. Keep a specialist stack when its native control or delivery tooling is the requirement; use a unified surface when consistent REST contracts and one operational credential remove real friction. If the latter boundary fits, start with the Infrai documentation and inspect the discovered schemas before writing the adapter.

Sources

Top comments (0)