Ever clicked "Submit" on a web app and watched the browser freeze for 4 to 5 seconds?
We encounter this everywhere: student portals, banking checkouts, and SaaS dashboards. You submit a form, the screen locks up, you click again out of frustration, and suddenly you get charged twice or receive duplicate notification emails.
In backend architectures, the culprit is almost always the same: holding the client's HTTP connection hostage while executing slow external network I/O.
[User Browser] --(HTTP POST)--> [Backend API] --(Blocks 4s on SMTP/Stripe)--> [User Browser Waits...]
When developers notice this latency, the most common "quick fix" is simply omitting the await keyword in Node.js:
// The dangerous anti-pattern
app.post("/register", async (req, res) => {
const user = await db.users.create(req.body);
// Fire-and-forget: Dropping await to respond "instantly"
sendWelcomeEmail(user.email);
syncWithAnalytics(user);
return res.status(201).json({ success: true });
});
It feels fast in local development, but in production, dropping await turns your application runtime into an untracked, fragile, in-memory queue.
The Danger of In-Memory Async Hacks
When you trigger promises without awaiting or persisting them:
- Ephemeral State: If your container recycles during an auto-scale event, restart, or deployment, all in-flight promises sitting in the V8 heap disappear permanently. Zero retry. Zero audit log.
- Unhandled Rejections: If the downstream SMTP server or external API drops the connection, an unhandled rejection can terminate the Node.js runtime process.
-
Memory Leaks: Under high-traffic bursts, thousands of unresolved promises stack up in memory, triggering garbage collection thrashing and Out-Of-Memory (
OOMKilled) termination. -
No Backpressure or Concurrency Bounds: There is no queue worker control. If 1,000 users sign up simultaneously, your server fires 1,000 concurrent external connections, leading to downstream rate limits (
429 Too Many Requests).
To solve this without paying recurring per-message SaaS fees, I built Mailman-Relay—a lightweight, self-hosted background worker and task orchestration engine.
The Core Concept: The Restaurant Counter
Think of it like ordering food at a busy restaurant counter:
You don’t make the customer stand at the billing desk for 15 minutes while the chef chops vegetables and cooks the entire meal.
Instead:
- You verify payment and complete the core transaction synchronously.
- You hand them a receipt token in under 15ms.
- The customer sits down immediately.
- The background kitchen staff picks up the order slip, prepares the food, and delivers it safely.
In Mailman:
-
The Ingestion Layer: Accepts tasks, enforces atomic deduplication, logs metadata in PostgreSQL, appends the payload to Redis Streams, and returns
202 Acceptedin sub-15ms. - The Worker Pool: Independently consumes tasks, handles rate limits, implements exponential backoff, and tracks delivery status.
Architectural Deep Dive
Mailman utilizes a dual-storage pattern pairing Redis Streams for low-latency dispatch with PostgreSQL for relational persistence and auditing.
1. Atomic Deduplication (Double-Click Shield)
To prevent duplicate job processing when users double-click, Mailman enforces an atomic lock at the entry point:
SET idempotency:{key} {job_id} NX EX {ttl}
If an identical request arrives while the first is in-flight, it receives an instant 409 Conflict, entirely avoiding race conditions.
2. Dual-Stack SSRF Protection (TOCTOU Defense)
For webhook dispatching, arbitrary URLs pose a security risk. Malicious payloads often target internal cloud endpoints (169.254.169.254 for AWS metadata, or private IP ranges like 10.0.0.0/8 and 127.0.0.1).
Mailman performs a dual-stack DNS resolution immediately prior to opening the TCP socket, validating the final resolved IPv4/IPv6 address against a strict private subnet blocklist with HTTP redirects disabled.
3. Crash Recovery with Redis Streams (XAUTOCLAIM)
What happens if a worker process gets killed midway through a task?
- Tasks read from Redis Streams reside in the Pending Entries List (PEL).
- Active workers acknowledge tasks via
XACKonly after successful completion. - If a worker crashes (
SIGKILL), the task remains unacknowledged. - A supervisory thread uses
XAUTOCLAIMto scan the PEL for idle tasks exceeding the visibility timeout and reassigns them to healthy workers.
Stress Testing & Failure Verification
A background orchestrator must be tested against real-world failures. I subjected Mailman to standard chaos and concurrency testing:
1. High-Concurrency Saturation (Grafana k6)
- Load Profile: 50 concurrent virtual users hitting the ingestion gateway continuously over 30 seconds.
- Total Transactions: 17,805 requests completed.
- Throughput: ~592 requests/second sustained.
- Error Rate: 0.00% across all requests.
-
Latency Profile:
- P50: 30.64 ms
- P95: 50.45 ms
- P99: 58.38 ms
Every request executed the atomic Redis lock, PostgreSQL state insertion, and stream append within this budget.
Benchmarking Environment Note:
These tests were executed in a local development environment running Node.js 20, Redis, and PostgreSQL concurrently. Tail latencies (P99 between ~58ms–90ms) reflect local process context switching, V8 GC cycles, and connection pool warm-ups. Production deployments across isolated cloud instances with dedicated network interfaces will vary based on regional topology.
2. Network Degradation Injection (Shopify Toxiproxy)
To test downstream failure, I injected a 4,000ms latency jitter toxic into the outgoing TCP socket using Toxiproxy:
- Client-facing ingestion remained decoupled, responding in 20ms.
- Background workers cleanly recognized the socket timeout without blocking the event loop and scheduled the task for exponential retry.
3. Worker Crash Recovery (SIGKILL)
I fired kill -9 at a worker while it was processing an active job:
- The abandoned job stayed in the Redis PEL.
- A standby worker reclaimed the task via
XAUTOCLAIMafter the timeout threshold and finished execution. - Total jobs lost: Zero.
Conclusion
Decoupling slow side-effects from your user-facing request cycle is essential for reliable web services. You don't need heavyweight cloud suites or recurring SaaS bills to achieve robust async processing. By combining Redis Streams for queue coordination and PostgreSQL for durability, you get sub-millisecond dispatching, zero lost jobs, and an instant user experience.
- GitHub Repository: https://github.com/Monis-dev/Mailman-Relay
- License: MIT (100% Free & Self-Hosted)

Top comments (0)