<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Monis Ahmed</title>
    <description>The latest articles on DEV Community by Monis Ahmed (@monis).</description>
    <link>https://dev.to/monis</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4154101%2Fe3fc5e5b-35cc-47b8-a5f3-3e2d654543dc.jpg</url>
      <title>DEV Community: Monis Ahmed</title>
      <link>https://dev.to/monis</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/monis"/>
    <language>en</language>
    <item>
      <title>Why Dropping "await" is Broken: Building a Resilient Background Task Engine with Node.js, Redis Streams, and PostgreSQL</title>
      <dc:creator>Monis Ahmed</dc:creator>
      <pubDate>Sun, 04 Oct 2026 14:06:51 +0000</pubDate>
      <link>https://dev.to/monis/why-dropping-await-is-broken-building-a-resilient-background-task-engine-with-nodejs-redis-466g</link>
      <guid>https://dev.to/monis/why-dropping-await-is-broken-building-a-resilient-background-task-engine-with-nodejs-redis-466g</guid>
      <description>&lt;p&gt;Ever clicked "Submit" on a web app and watched the browser freeze for 4 to 5 seconds?&lt;/p&gt;

&lt;p&gt;We encounter this everywhere: student portals, banking checkouts, and SaaS dashboards. You submit a form, the screen locks up, you click again out of frustration, and suddenly you get charged twice or receive duplicate notification emails.&lt;/p&gt;

&lt;p&gt;In backend architectures, the culprit is almost always the same: &lt;strong&gt;holding the client's HTTP connection hostage while executing slow external network I/O.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[User Browser] --(HTTP POST)--&amp;gt; [Backend API] --(Blocks 4s on SMTP/Stripe)--&amp;gt; [User Browser Waits...]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When developers notice this latency, the most common "quick fix" is simply omitting the &lt;code&gt;await&lt;/code&gt; keyword in Node.js:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// The dangerous anti-pattern&lt;/span&gt;
&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/register&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;user&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;users&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// Fire-and-forget: Dropping await to respond "instantly"&lt;/span&gt;
  &lt;span class="nf"&gt;sendWelcomeEmail&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;email&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; 
  &lt;span class="nf"&gt;syncWithAnalytics&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;201&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;success&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It feels fast in local development, but in production, dropping &lt;code&gt;await&lt;/code&gt; turns your application runtime into an untracked, fragile, in-memory queue.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Danger of In-Memory Async Hacks
&lt;/h2&gt;

&lt;p&gt;When you trigger promises without awaiting or persisting them:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Ephemeral State:&lt;/strong&gt; If your container recycles during an auto-scale event, restart, or deployment, all in-flight promises sitting in the V8 heap disappear permanently. Zero retry. Zero audit log.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unhandled Rejections:&lt;/strong&gt; If the downstream SMTP server or external API drops the connection, an unhandled rejection can terminate the Node.js runtime process.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory Leaks:&lt;/strong&gt; Under high-traffic bursts, thousands of unresolved promises stack up in memory, triggering garbage collection thrashing and Out-Of-Memory (&lt;code&gt;OOMKilled&lt;/code&gt;) termination.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No Backpressure or Concurrency Bounds:&lt;/strong&gt; There is no queue worker control. If 1,000 users sign up simultaneously, your server fires 1,000 concurrent external connections, leading to downstream rate limits (&lt;code&gt;429 Too Many Requests&lt;/code&gt;).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;To solve this without paying recurring per-message SaaS fees, I built &lt;strong&gt;Mailman-Relay&lt;/strong&gt;—a lightweight, self-hosted background worker and task orchestration engine.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Core Concept: The Restaurant Counter
&lt;/h2&gt;

&lt;p&gt;Think of it like ordering food at a busy restaurant counter:&lt;/p&gt;

&lt;p&gt;You don’t make the customer stand at the billing desk for 15 minutes while the chef chops vegetables and cooks the entire meal. &lt;/p&gt;

&lt;p&gt;Instead:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You verify payment and complete the core transaction synchronously.&lt;/li&gt;
&lt;li&gt;You hand them a receipt token in under 15ms.&lt;/li&gt;
&lt;li&gt;The customer sits down immediately.&lt;/li&gt;
&lt;li&gt;The background kitchen staff picks up the order slip, prepares the food, and delivers it safely.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In Mailman:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Ingestion Layer:&lt;/strong&gt; Accepts tasks, enforces atomic deduplication, logs metadata in PostgreSQL, appends the payload to Redis Streams, and returns &lt;code&gt;202 Accepted&lt;/code&gt; in sub-15ms.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Worker Pool:&lt;/strong&gt; Independently consumes tasks, handles rate limits, implements exponential backoff, and tracks delivery status.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Architectural Deep Dive
&lt;/h2&gt;

&lt;p&gt;Mailman utilizes a dual-storage pattern pairing &lt;strong&gt;Redis Streams&lt;/strong&gt; for low-latency dispatch with &lt;strong&gt;PostgreSQL&lt;/strong&gt; for relational persistence and auditing.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Atomic Deduplication (Double-Click Shield)
&lt;/h3&gt;

&lt;p&gt;To prevent duplicate job processing when users double-click, Mailman enforces an atomic lock at the entry point:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SET idempotency:{key} {job_id} NX EX {ttl}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If an identical request arrives while the first is in-flight, it receives an instant &lt;code&gt;409 Conflict&lt;/code&gt;, entirely avoiding race conditions.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Dual-Stack SSRF Protection (TOCTOU Defense)
&lt;/h3&gt;

&lt;p&gt;For webhook dispatching, arbitrary URLs pose a security risk. Malicious payloads often target internal cloud endpoints (&lt;code&gt;169.254.169.254&lt;/code&gt; for AWS metadata, or private IP ranges like &lt;code&gt;10.0.0.0/8&lt;/code&gt; and &lt;code&gt;127.0.0.1&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;Mailman performs a dual-stack DNS resolution immediately prior to opening the TCP socket, validating the final resolved IPv4/IPv6 address against a strict private subnet blocklist with HTTP redirects disabled.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Crash Recovery with Redis Streams (&lt;code&gt;XAUTOCLAIM&lt;/code&gt;)
&lt;/h3&gt;

&lt;p&gt;What happens if a worker process gets killed midway through a task?&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tasks read from Redis Streams reside in the &lt;strong&gt;Pending Entries List (PEL)&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Active workers acknowledge tasks via &lt;code&gt;XACK&lt;/code&gt; only after successful completion.&lt;/li&gt;
&lt;li&gt;If a worker crashes (&lt;code&gt;SIGKILL&lt;/code&gt;), the task remains unacknowledged.&lt;/li&gt;
&lt;li&gt;A supervisory thread uses &lt;code&gt;XAUTOCLAIM&lt;/code&gt; to scan the PEL for idle tasks exceeding the visibility timeout and reassigns them to healthy workers.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Stress Testing &amp;amp; Failure Verification
&lt;/h2&gt;

&lt;p&gt;A background orchestrator must be tested against real-world failures. I subjected Mailman to standard chaos and concurrency testing:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. High-Concurrency Saturation (Grafana k6)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Load Profile:&lt;/strong&gt; 50 concurrent virtual users hitting the ingestion gateway continuously over 30 seconds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Total Transactions:&lt;/strong&gt; 17,805 requests completed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Throughput:&lt;/strong&gt; ~592 requests/second sustained.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Error Rate:&lt;/strong&gt; 0.00% across all requests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latency Profile:&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;P50: 30.64 ms&lt;/li&gt;
&lt;li&gt;P95: 50.45 ms&lt;/li&gt;
&lt;li&gt;P99: 58.38 ms&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6jvn9i5ibw65aja4z8pz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6jvn9i5ibw65aja4z8pz.png" alt="k6 test case" width="800" height="829"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every request executed the atomic Redis lock, PostgreSQL state insertion, and stream append within this budget.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Benchmarking Environment Note:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
These tests were executed in a local development environment running Node.js 20, Redis, and PostgreSQL concurrently. Tail latencies (P99 between ~58ms–90ms) reflect local process context switching, V8 GC cycles, and connection pool warm-ups. Production deployments across isolated cloud instances with dedicated network interfaces will vary based on regional topology.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  2. Network Degradation Injection (Shopify Toxiproxy)
&lt;/h3&gt;

&lt;p&gt;To test downstream failure, I injected a 4,000ms latency jitter toxic into the outgoing TCP socket using Toxiproxy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Client-facing ingestion remained decoupled, responding in &lt;strong&gt;20ms&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Background workers cleanly recognized the socket timeout without blocking the event loop and scheduled the task for exponential retry.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. Worker Crash Recovery (&lt;code&gt;SIGKILL&lt;/code&gt;)
&lt;/h3&gt;

&lt;p&gt;I fired &lt;code&gt;kill -9&lt;/code&gt; at a worker while it was processing an active job:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The abandoned job stayed in the Redis PEL.&lt;/li&gt;
&lt;li&gt;A standby worker reclaimed the task via &lt;code&gt;XAUTOCLAIM&lt;/code&gt; after the timeout threshold and finished execution.&lt;/li&gt;
&lt;li&gt;Total jobs lost: &lt;strong&gt;Zero&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Decoupling slow side-effects from your user-facing request cycle is essential for reliable web services. You don't need heavyweight cloud suites or recurring SaaS bills to achieve robust async processing. By combining Redis Streams for queue coordination and PostgreSQL for durability, you get sub-millisecond dispatching, zero lost jobs, and an instant user experience.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub Repository:&lt;/strong&gt; &lt;a href="https://github.com/Monis-dev/Mailman" rel="noopener noreferrer"&gt;https://github.com/Monis-dev/Mailman-Relay&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;License:&lt;/strong&gt; MIT (100% Free &amp;amp; Self-Hosted)&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>opensource</category>
      <category>node</category>
      <category>postgres</category>
      <category>redis</category>
    </item>
  </channel>
</rss>
