Treat suppression as a compliance ledger, not a side effect of sending mail. For a property-management SaaS, the right API is the one whose bounce events can be authenticated, replayed, deduplicated, and tied to a decision about a recipient before the next rent receipt or maintenance update is sent. Short answer: put a small FastAPI ingestion service and a local suppression store between any transactional email API and your application, then evaluate providers against that contract. DKIM and DMARC protect a different boundary; they do not replace recipient-level suppression evidence.
The data flow is deliberately plain. A provider posts a delivery event, the ingress verifies its signature, and the service records the untouched event before applying a suppression rule. The sending path asks the same store for permission. Polling is a recovery path for missed notifications, using the provider's stable event cursor, rather than a second source of truth.
A ledger for recipient decisions
Start with an append-only event record. A mutable suppressed = true flag cannot answer who reported the failure, when the application received it, or why a later message was blocked. Those details matter when a tenant changes addresses, a property manager disputes a missing notice, or an operator has to explain the system's decision.
Use an internal event shape that does not copy one provider's vocabulary:
from dataclasses import dataclass
from datetime import datetime
from enum import StrEnum
class EventKind(StrEnum):
HARD_BOUNCE = "hard_bounce"
SOFT_BOUNCE = "soft_bounce"
COMPLAINT = "complaint"
DELIVERED = "delivered"
@dataclass(frozen=True)
class DeliveryEvent:
source_event_id: str
recipient: str
kind: EventKind
occurred_at: datetime
received_at: datetime
raw_sha256: str
Keep the raw payload as evidence in access-controlled storage and place only its digest in the decision table. That separation reduces the personal data copied into operational queries while preserving a way to show that the normalized row corresponds to a specific input. Define retention and access rules with counsel for the jurisdictions and notices your application actually handles; geography alone does not produce a universal retention period.
DMARC belongs in the same review, but on its own track. RFC 7489 describes a domain-owner policy published in DNS and aggregate or failure reporting about message authentication. It also describes alignment between an authenticated identifier and the visible From domain. That supports domain-level authentication evidence. A hard-bounce ledger supports a recipient decision. Conflating them leaves a gap.
Implement the event boundary first
The following service is intentionally small enough to run during an evaluation. It uses SQLite so the durability rule is visible, FastAPI for the HTTP boundary, and an HMAC contract as a generic example. In a real integration, replace verify_signature with the candidate API's documented verification algorithm; do not assume every service signs the same bytes or uses the same header.
import hashlib
import hmac
import json
import os
import sqlite3
from datetime import UTC, datetime
from fastapi import FastAPI, Header, HTTPException, Request
app = FastAPI()
db = sqlite3.connect("delivery_events.db", check_same_thread=False)
db.execute("PRAGMA journal_mode=WAL")
db.execute(
"""
CREATE TABLE IF NOT EXISTS delivery_events (
source_event_id TEXT PRIMARY KEY,
recipient TEXT NOT NULL,
kind TEXT NOT NULL,
occurred_at TEXT NOT NULL,
received_at TEXT NOT NULL,
raw_sha256 TEXT NOT NULL
)
"""
)
db.execute(
"""
CREATE TABLE IF NOT EXISTS suppressions (
recipient TEXT PRIMARY KEY,
reason TEXT NOT NULL,
source_event_id TEXT NOT NULL,
decided_at TEXT NOT NULL
)
"""
)
def verify_signature(body: bytes, supplied: str) -> bool:
secret = os.environ["EVENT_SIGNING_SECRET"].encode()
expected = hmac.new(secret, body, hashlib.sha256).hexdigest()
return hmac.compare_digest(expected, supplied)
@app.post("/delivery-events")
async def receive_event(
request: Request,
x_event_signature: str = Header(),
) -> dict[str, str]:
body = await request.body()
if not verify_signature(body, x_event_signature):
raise HTTPException(status_code=401, detail="invalid signature")
try:
payload = json.loads(body)
event_id = str(payload["id"])
recipient = str(payload["recipient"]).strip().lower()
kind = str(payload["kind"])
occurred_at = datetime.fromisoformat(payload["occurred_at"])
except (KeyError, TypeError, ValueError, json.JSONDecodeError) as error:
raise HTTPException(status_code=400, detail="invalid event") from error
received_at = datetime.now(UTC).isoformat()
digest = hashlib.sha256(body).hexdigest()
with db:
inserted = db.execute(
"""
INSERT OR IGNORE INTO delivery_events
VALUES (?, ?, ?, ?, ?, ?)
""",
(
event_id,
recipient,
kind,
occurred_at.isoformat(),
received_at,
digest,
),
).rowcount
if inserted and kind in {"hard_bounce", "complaint"}:
db.execute(
"""
INSERT OR IGNORE INTO suppressions
VALUES (?, ?, ?, ?)
""",
(recipient, kind, event_id, received_at),
)
return {"status": "accepted"}
Two details carry most of the reliability. The provider event ID is the primary key, so retries do not create new decisions. The event and suppression mutation share one transaction. If the process stops between them, SQLite rolls back both; the sender never observes a suppression without its evidence row.
Do not suppress every temporary failure forever. The example makes only hard_bounce and complaint terminal decisions because its policy explicitly names them. Soft-bounce handling needs a documented threshold and time window derived from your delivery requirements. Record each input even when it does not cross that threshold.
The suppression query must happen before an API call is queued, not inside a webhook worker after another message may already be in flight. Keep that rule boring and testable.
def can_send(recipient: str) -> bool:
normalized = recipient.strip().lower()
row = db.execute(
"SELECT 1 FROM suppressions WHERE recipient = ?",
(normalized,),
).fetchone()
return row is None
def enqueue_property_notice(recipient: str, notice_id: str) -> dict[str, str]:
if not can_send(recipient):
return {"status": "blocked", "notice_id": notice_id}
# Replace this return value with the application's durable queue write.
return {"status": "queued", "notice_id": notice_id}
This is also where notebook-to-production discipline pays off. A notebook can replay a fixed corpus of signed events against a temporary database; production uses the same normalization and decision functions behind a durable queue. I would make the evaluation corpus include a duplicate hard bounce, an out-of-order delivery after that bounce, a malformed timestamp, a bad signature, and two distinct event IDs for one recipient. Five cases reveal more about the integration contract than a polished dashboard does.
The test should assert decisions and evidence, not response prose:
def test_duplicate_event_creates_one_decision(client, signed_hard_bounce):
first = client.post("/delivery-events", **signed_hard_bounce)
second = client.post("/delivery-events", **signed_hard_bounce)
assert first.status_code == 200
assert second.status_code == 200
assert db.execute("SELECT COUNT(*) FROM delivery_events").fetchone()[0] == 1
assert db.execute("SELECT COUNT(*) FROM suppressions").fetchone()[0] == 1
Short tests win here.
Can a SaaS transactional email API prove domain deliverability?
No event receiver should depend on perfect, one-shot delivery. During a provider evaluation, disconnect ingress, generate known test events, reconnect it, and recover the gap by polling. The useful question is not whether an events endpoint exists. Ask whether its cursor is stable, whether results have a deterministic order, how long events remain available, and whether the poll response carries the same immutable ID and timestamp as push delivery.
Run the poller from a saved cursor and feed every result through the same normalization and transaction used by the webhook. Advance the cursor only after the local commit. This prevents a successful network response followed by a failed database write from silently skipping events.
There is a cost trade-off. Polling every few seconds spends requests to reduce recovery latency, while a slower interval leaves a wider window in which an invalid recipient might be considered sendable. Pick the interval from the notice workflow's risk tolerance, measure requests and lag in the eval harness, and avoid presenting a vendor's current unit price as an architectural constant.
Domain verification needs a similarly concrete test. Verify a subdomain reserved for transactional mail, inspect the published DKIM material, and confirm how key rotation overlaps old and new selectors. Then send evaluation messages and retain the authentication results with the configuration change record. DMARC policy and reporting can show how receivers evaluated alignment, but the application still needs its own evidence that a domain configuration was approved and active when a notice was attempted.
| Control | Evidence to retain | Failure the test exposes |
|---|---|---|
| Event ingestion | Immutable event ID, receipt time, payload digest | Duplicate or unauthenticated input |
| Suppression | Reason, source event ID, decision time | A later send to a blocked address |
| Poll recovery | Committed cursor and event ordering | A gap after ingress downtime |
| Domain authentication | DNS change record and message authentication result | Misalignment or incomplete key rotation |
Make the operating rule explicit
Before launch, write one sentence that an on-call engineer can apply: a verified hard bounce or complaint blocks new transactional sends to that normalized address until an authorized correction creates a new, linked decision. Then test the full loop with synthetic recipients. Alert on event-ingestion lag, signature failures, polling cursor age, and sends blocked by suppression; those signals describe the control rather than a vendor's internal status.
Review access to raw event bodies, test restoration of the ledger, and sample decisions back to their payload digests. Keep domain authentication checks separate from recipient suppression checks in both dashboards and audits. The best fit is the API that lets your team prove this chain under replay and outage tests, with documented event identity, signature verification, retention, pagination, and DKIM rotation behavior. Everything else is interface polish.
Sources and References
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
Top comments (0)