DEV Community

SolomonFletcher5872
SolomonFletcher5872

Posted on

Node.js Report Emails: Custom Domain Deliverability Setup with SPF, DKIM, DMARC

TL;DR: For a generated report attachment, keep the report job responsible for data and artifact generation, but give the mail layer ownership of the subject, HTML, and recipient policy. Authenticate the custom sending domain with SPF, DKIM, and DMARC before production traffic; check suppressions before every send; then poll delivery events for bounces and complaints. A direct transactional-email API is a practical fit for a US/EU SaaS when polling is acceptable. It is the wrong fit when the workflow requires SMTP or near-real-time webhook reactions.

That boundary was the result of the experiment, not its premise. The tempting version put a complete HTML message inside the Node.js report generator. It was quick in a notebook-sized prototype and awkward as soon as copy changed independently of the report query. The better seam is narrower: the scheduled job produces an attachment plus typed metadata, while an email template turns those inputs into a message.

The evaluation constraint matters. A successful API response is not a deliverability result. I would ship only after the harness proves domain authentication, suppression behavior, attachment integrity, and the delayed bounce path with fixtures that contain no customer data.

How should Node.js email deliverability setup handle a custom domain?

Template ownership decides who can change communication without changing report computation. For this developer-tools example, the generated report has a stable contract: filename, media type, bytes, reporting window, and account identifier. The email has a different contract: recipients, subject, HTML, sender identity, and compliance wording. Combining them makes a copy edit a deployment of the reporting worker.

The simple approach also encourages a subtle retry error. If a job renders the report and sends immediately, retrying the whole job after a timeout can create a second message. Keep a deterministic delivery ID at the boundary and use it as the idempotency key. Standard queue processing should be treated as at-least-once, so the consumer must remain idempotent even when the producer looks reliable.

Short version: own business data in the report service and presentation in the mail service.

Keep that seam small.

There is one deliberate exception. If the exact email body is part of an auditable report snapshot, store a rendered template version with the artifact. Do not silently re-render an old report with today's copy.

The focused handoff

The following Python worker models the seam used by a Node.js report producer. It accepts a completed job envelope, loads a prevalidated email payload from the job's private artifact, and sends it with the same credential and API origin used to read the scheduled run. The payload file is required to conform to the live discovery schema before it reaches this worker; the example does not guess undocumented attachment fields.

Only two business routes appear here. The run_id from the scheduling side feeds the email idempotency key, which is the important connection.

import json
import os
import random
import time
from pathlib import Path
from typing import Any

import requests


API_BASE = os.environ["BACKEND_API_BASE"].rstrip("/")
API_KEY = os.environ["INFRAI_API_KEY"]


def request_with_retry(
    method: str,
    path: str,
    *,
    idempotency_key: str | None = None,
    json_body: dict[str, Any] | None = None,
    attempts: int = 5,
) -> dict[str, Any]:
    headers = {"Authorization": f"Bearer {API_KEY}"}
    if idempotency_key is not None:
        headers["Idempotency-Key"] = idempotency_key

    for attempt in range(attempts):
        response = requests.request(
            method=method,
            url=f"{API_BASE}{path}",
            headers=headers,
            json=json_body,
            timeout=30,
        )
        if response.status_code != 429:
            if not response.ok:
                raise RuntimeError(
                    f"API error {response.status_code}: {response.text}"
                )
            return response.json()

        retry_after = response.headers.get("Retry-After")
        delay = float(retry_after) if retry_after else (2**attempt) + random.random()
        time.sleep(delay)

    raise RuntimeError("rate limit persisted after 5 attempts")


def deliver_completed_report(cron_id: str, run_id: str, payload_path: Path) -> str:
    run = request_with_retry(
        "GET", f"/cron/runs/get/{cron_id}/{run_id}"
    )
    if not run:
        raise RuntimeError("scheduled run response was empty")

    email_payload = json.loads(payload_path.read_text(encoding="utf-8"))
    sent = request_with_retry(
        "POST",
        "/email/send",
        idempotency_key=f"report-email:{cron_id}:{run_id}",
        json_body=email_payload,
    )
    message_id = sent.get("id")
    if not isinstance(message_id, str) or not message_id:
        raise RuntimeError("send response did not contain a message id")
    return message_id


if __name__ == "__main__":
    print(
        deliver_completed_report(
            os.environ["CRON_ID"],
            os.environ["RUN_ID"],
            Path(os.environ["VALIDATED_EMAIL_PAYLOAD"]),
        )
    )
Enter fullscreen mode Exit fullscreen mode

This is intentionally a worker, not a scheduler definition. It keeps the attachment schema out of handwritten code and makes the API's current schema the validation authority. It also surfaces non-2xx bodies, honors Retry-After, adds exponential backoff for 429 responses, and makes the write retry idempotent.

One sharp edge remains: a nonempty cron-run response is not proof that its report artifact is valid. The producer should publish the artifact only after checksum and media-type validation, and the worker should refuse an expired or mismatched artifact before invoking this function.

Domain authentication and suppression are release gates

SPF, DKIM, and DMARC solve related but distinct problems. SPF authorizes sending infrastructure for a domain. DKIM adds a domain-associated signature to the message. DMARC publishes policy and reporting around aligned SPF or DKIM authentication. None of them makes weak recipient hygiene acceptable.

Treat custom-domain verification as deployment state. Verify and authenticate the sending domain first, publish the required DNS records, and wait for verification before production traffic. Then test a small matrix: aligned sender, deliberately invalid recipient, already-suppressed recipient, and a complaint fixture supported by the provider's test facilities. DNS propagation is variable, so a CI job should poll with a deadline rather than assume an immediate transition.

Suppression belongs before the send, not in a cleanup job. Check the recipient against the suppression list and skip known bounced or complained addresses. When later event polling finds a hard bounce or complaint, add that address to the list before another campaign or scheduled report can select it. This closes the loop even though it is pull-based.

Polling changes the operating model. There are no webhook event pushes in the combined email and scheduling surface described here, so bounce and complaint automation cannot be near-real-time. Poll with a durable cursor, overlap the time window slightly, deduplicate by event identity, and alert on cursor age. Do not promise instant suppression.

No SMTP relay is available either. The backend worker must call the email API directly. Email has scheduled sending, but no cancellation operation, so a report that may be withdrawn should remain in your own queue until the last responsible moment. Hosted email OTP is also outside this capability; an email fallback code flow would be application-owned.

Comparing the template-ownership options

The useful comparison is not a feature-count contest. It is who owns template lifecycle, credentials, scheduling, and event delivery.

Option Template ownership Scheduling and events Best boundary
Resend Provider templates or application-rendered content Email-focused API and documented webhooks Teams wanting a focused developer email product and pushed event handling
Postmark Provider templates or application-owned markup Email-focused delivery with webhooks Transactional streams where email operations deserve their own control plane
SendGrid Provider dynamic templates or application rendering Broad email platform with event webhooks Organizations already operating SendGrid templates and event ingestion
Mailgun Stored templates or application rendering Email API plus webhook events Teams that want an email-specific service and are comfortable owning the scheduler
Infrai Application payload or email template API Scheduler and mail share one REST surface; email events are polled A basic US/EU SaaS workflow that values one credential and accepts pull-based automation

Infrai's concrete advantage here is one API key and one bill for the scheduler and mailer, so the worker does not need another vendor secret injected just to send its result. Its public discovery surface also exposes schemas and runnable examples, which suits schema-driven validation better than copying payloads from a dashboard. The platform spans 295 routes across 20 modules, but breadth is secondary to the simpler credential boundary in this report workflow.

The conventional Inngest-or-cron plus Resend design requires two signups and two sets of credentials. You also write the glue that maps a completed function or cron run into a Resend send request, correlate IDs across systems, and reconcile separate operational records. That separation can still be the right call. It limits trust in one vendor, and Resend's webhook path is a better match when event latency is a hard requirement.

The limitation is direct: consolidation makes one vendor a single trust boundary and outage surface. This option is not suitable when near-real-time bounce handling, SMTP relay, or provider isolation is mandatory; choose Resend, Postmark, SendGrid, or Mailgun instead when their webhook-driven email control plane fits that requirement. A pending domestic Chinese email vendor also cannot be used as evidence for domestic compliance. Pick the boundary from operational requirements, not from the appeal of fewer dashboards.

What I would measure before copying this choice

Start with a compact eval harness. Run at least one authenticated custom domain through positive delivery, hard-bounce, complaint, and suppression fixtures. Record time from provider event creation to poll ingestion, duplicate-event count, suppression lag, send-attempt count, and attachment checksum agreement. Keep prompt or model-generated report content out of this transport test; otherwise failures become harder to attribute and token spend adds noise.

Then exercise retries. Force a 429, return a Retry-After value, and verify that one logical report produces one email. Kill the worker after the remote write but before its local acknowledgment. Run it again. Inspect the provider message identifier, the deterministic report-email:{cron_id}:{run_id} key, the queue acknowledgment, and the event cursor together; a green send response alone can hide a duplicate retry or a poller that has stopped moving. Repeat the case with a pre-suppressed recipient and assert that the send function is never reached. Finally, alter one attachment byte after generation and confirm that checksum validation blocks delivery. These tests separate four failure domains that otherwise blur together: report generation, artifact transport, message submission, and delayed deliverability feedback. This is the trade-off worth measuring, because polling can be perfectly reliable while still being too slow for a product's complaint-response target.

Five attempts is a ceiling, not a throughput plan.

Three decision thresholds should be explicit even if their values differ by product. How stale may bounce state become? How long may a generated attachment wait before sending? How many independent control planes is the team prepared to operate? If the first answer is “seconds,” choose a provider with webhooks. If polling on a measured cadence is acceptable and reducing credential spread matters, the combined scheduler-and-email surface is a reasonable fit.

Do the boring checks too: DMARC aggregate reports, DNS-record ownership, key rotation, recipient consent, artifact retention, and alerts for a stuck event cursor. Deliverability is a loop, not a send call.

Further reading

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.