<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Kamil Mysliwiec</title>
    <description>The latest articles on DEV Community by Kamil Mysliwiec (@kamilmysliwiec).</description>
    <link>https://dev.to/kamilmysliwiec</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F196704%2F637eb949-df9c-4cb5-8b02-7a890cd56ce8.png</url>
      <title>DEV Community: Kamil Mysliwiec</title>
      <link>https://dev.to/kamilmysliwiec</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kamilmysliwiec"/>
    <language>en</language>
    <item>
      <title>NestJS Production Checklist: 17 Checks Before You Deploy</title>
      <dc:creator>Kamil Mysliwiec</dc:creator>
      <pubDate>Wed, 30 Sep 2026 08:55:09 +0000</pubDate>
      <link>https://dev.to/nestjs/nestjs-production-checklist-17-checks-before-you-deploy-1d0j</link>
      <guid>https://dev.to/nestjs/nestjs-production-checklist-17-checks-before-you-deploy-1d0j</guid>
      <description>&lt;p&gt;Locally, a Nest app runs as one process, restarts when you save, talks to a database with no other clients, and gets requests straight from your browser. In production it runs as several replicas behind a load balancer, gets stopped on every deploy, shares its database with other processes, and depends on APIs that sometimes don't answer. A lot of defaults that are fine for the first setup are wrong for the second.&lt;/p&gt;

&lt;h2&gt;
  
  
  Startup and deploys
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Validate config at boot, and keep secrets out of the image
&lt;/h3&gt;

&lt;p&gt;A missing or malformed environment variable should stop the process before it accepts a request. Otherwise the app boots, passes its health check, takes traffic, and fails the first request that reads the value.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;@nestjs/config&lt;/code&gt; accepts any Standard Schema (Zod, Valibot, ArkType) as &lt;code&gt;validationSchema&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;zod&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Module&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;imports&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="nx"&gt;ConfigModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;forRoot&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;isGlobal&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;validationSchema&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;object&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="na"&gt;NODE_ENV&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;enum&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;development&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;test&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;production&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]),&lt;/span&gt;
        &lt;span class="na"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;url&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="na"&gt;STRIPE_SECRET_KEY&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="na"&gt;PORT&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;coerce&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;number&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="k"&gt;default&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3000&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
      &lt;span class="p"&gt;}),&lt;/span&gt;
    &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;AppModule&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read required values with &lt;code&gt;config.getOrThrow('STRIPE_SECRET_KEY')&lt;/code&gt; rather than &lt;code&gt;get()&lt;/code&gt;. During a rolling deploy, a version that fails on boot never becomes ready and the previous one keeps serving. That's the failure mode you want.&lt;/p&gt;

&lt;p&gt;While you're here, check where the values come from. Secrets belong in the platform's secret store (Kubernetes Secrets, AWS Secrets Manager, Vault, your PaaS's config), injected at runtime. Not in a &lt;code&gt;.env&lt;/code&gt; file copied into the image, and not as a Docker &lt;code&gt;ARG&lt;/code&gt;/&lt;code&gt;ENV&lt;/code&gt;: both end up in the image layers, readable by anyone who can pull the image. A &lt;code&gt;.dockerignore&lt;/code&gt; with &lt;code&gt;.env*&lt;/code&gt; in it is a cheap guard.&lt;/p&gt;

&lt;p&gt;And don't log the config object. &lt;code&gt;console.log(config)&lt;/code&gt; at startup is a common debugging step that stays in, and it ships every credential to your log pipeline. If you need to see what loaded, log the keys, not the values.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Handle SIGTERM, and get the container right
&lt;/h3&gt;

&lt;p&gt;Deploys, scale-downs and node drains all send &lt;code&gt;SIGTERM&lt;/code&gt;. Nest doesn't listen for it unless you ask, so by default &lt;code&gt;onModuleDestroy&lt;/code&gt; and &lt;code&gt;onApplicationShutdown&lt;/code&gt; never run: pools stay open, queue workers are killed mid-job, in-flight requests are dropped.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;NestFactory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;AppModule&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;enableShutdownHooks&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;listen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;PORT&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="mi"&gt;3000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With hooks enabled, a signal runs: the HTTP adapter is marked as closing, &lt;code&gt;onModuleDestroy&lt;/code&gt;, &lt;code&gt;beforeApplicationShutdown&lt;/code&gt;, the HTTP server, gateways and microservices close, then &lt;code&gt;onApplicationShutdown&lt;/code&gt;. Most integrations clean up in that last one. &lt;code&gt;@nestjs/bullmq&lt;/code&gt;, for example, closes its workers there, and &lt;code&gt;worker.close()&lt;/code&gt; waits for the jobs they're running.&lt;/p&gt;

&lt;p&gt;Two ways the signal gets lost in containers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Something between the runtime and Node.&lt;/strong&gt; &lt;code&gt;CMD npm run start:prod&lt;/code&gt;, or the shell form &lt;code&gt;CMD node dist/main.js&lt;/code&gt;, puts npm or a shell in front of Node, and neither reliably forwards signals. Use the exec form: &lt;code&gt;CMD ["node", "dist/main.js"]&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Node running as PID 1.&lt;/strong&gt; The kernel doesn't apply default signal behaviour to PID 1. &lt;code&gt;enableShutdownHooks&lt;/code&gt; installs its own handler, so the hooks run, but Nest then exits by removing that handler and re-sending the signal to its own process. As PID 1, that second signal is ignored, and the process only exits if the event loop happens to be empty. Run with an init process (&lt;code&gt;docker run --init&lt;/code&gt;, or &lt;code&gt;tini&lt;/code&gt; as the entrypoint), or use &lt;code&gt;app.enableShutdownHooks(undefined, { useProcessExit: true })&lt;/code&gt; so Nest calls &lt;code&gt;process.exit()&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rest of the image is worth a look while you're in the Dockerfile. Build in one stage and run in another, so compilers, dev dependencies and your source tree don't ship; set &lt;code&gt;NODE_ENV=production&lt;/code&gt;; and don't run as root, since the official Node images already include a &lt;code&gt;node&lt;/code&gt; user:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;node:24-slim&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;AS&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;build&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; package*.json ./&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;npm ci
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; . .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;npm run build &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; npm prune &lt;span class="nt"&gt;--omit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;dev

&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; node:24-slim&lt;/span&gt;
&lt;span class="k"&gt;ENV&lt;/span&gt;&lt;span class="s"&gt; NODE_ENV=production&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; --from=build /app/package.json ./&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; --from=build /app/node_modules ./node_modules&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; --from=build /app/dist ./dist&lt;/span&gt;
&lt;span class="k"&gt;USER&lt;/span&gt;&lt;span class="s"&gt; node&lt;/span&gt;
&lt;span class="k"&gt;CMD&lt;/span&gt;&lt;span class="s"&gt; ["node", "--enable-source-maps", "dist/main.js"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(&lt;code&gt;package.json&lt;/code&gt; is copied on purpose: the Nest 12 starter is ESM, and Node reads &lt;code&gt;"type": "module"&lt;/code&gt; from it.)&lt;/p&gt;

&lt;p&gt;Then check memory. If the container has a memory limit and V8's heap limit is higher than what the container can actually give it, the kernel kills the process before V8 gets a chance to complain: exit code 137, no JavaScript error, no log line, just a restart. Depending on the Node version and how the limit is set, the default heap size doesn't necessarily follow the container's limit, so set it explicitly to leave room for everything that isn't heap (buffers, native memory, the stack):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;NODE_OPTIONS&lt;/span&gt;
    &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--max-old-space-size=768"&lt;/span&gt; &lt;span class="c1"&gt;# ~75% of a 1 GiB limit&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now a leak or an oversized payload ends in a &lt;code&gt;JavaScript heap out of memory&lt;/code&gt; error with a stack trace in your logs, instead of a silent kill. Alert on restarts either way, and on &lt;code&gt;OOMKilled&lt;/code&gt; specifically.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Drain before you close
&lt;/h3&gt;

&lt;p&gt;When Kubernetes terminates a pod, two things start at the same time: the pod is removed from the Service's endpoints, and the kubelet stops the container (&lt;code&gt;preStop&lt;/code&gt; hook, then &lt;code&gt;SIGTERM&lt;/code&gt;). Endpoint removal takes a few seconds to reach every load balancer and kube-proxy. A server that closes immediately on &lt;code&gt;SIGTERM&lt;/code&gt; refuses the requests routed to it during that window.&lt;/p&gt;

&lt;p&gt;The usual fix is a short delay before shutdown starts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;lifecycle&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;preStop&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;exec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sleep"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;5"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nest 12 also has &lt;code&gt;return503OnClosing: true&lt;/code&gt; on &lt;code&gt;NestFactory.create()&lt;/code&gt;. It answers new requests with a 503 and &lt;code&gt;Connection: close&lt;/code&gt; while in-flight ones finish. That's useful once traffic has drained, but during the propagation window those 503s reach clients, so use it together with the delay, not instead of it.&lt;/p&gt;

&lt;p&gt;Then check the time budget. Kubernetes sends &lt;code&gt;SIGKILL&lt;/code&gt; after &lt;code&gt;terminationGracePeriodSeconds&lt;/code&gt; (default 30), and the &lt;code&gt;preStop&lt;/code&gt; time counts against it. A queue job that runs longer than what's left gets killed, BullMQ marks it stalled, and another worker runs it again from the start. Either raise the grace period or keep jobs short, and make them idempotent regardless, because a crash does the same thing.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Scheduled jobs run on every replica
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;@Cron()&lt;/code&gt;, &lt;code&gt;@Interval()&lt;/code&gt; and &lt;code&gt;@Timeout()&lt;/code&gt; from &lt;code&gt;@nestjs/schedule&lt;/code&gt; run in-process. Three replicas means the nightly invoice export runs three times. It's one of the most common production bugs in Nest apps, because it can't happen on a single dev instance.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;@nestjs/locks&lt;/code&gt; handles this with decorators:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Cron&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;0 2 * * *&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;OnOneInstance&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;invoices:export&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;exportInvoices&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// runs on exactly one instance; the others skip the tick&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Cron&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;*/5 * * * *&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;WithoutOverlapping&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;reconcile&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// skips a tick while the previous run is still going&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;@OnOneInstance()&lt;/code&gt; holds a lease that the owning instance renews every &lt;code&gt;ttl / 3&lt;/code&gt; (30 s TTL by default). If that instance dies, another one takes over once the TTL runs out. &lt;code&gt;@WithoutOverlapping()&lt;/code&gt; only holds the lock for the length of a run, which covers the other common bug: a 5-minute job that sometimes takes 7.&lt;/p&gt;

&lt;p&gt;The lock store must be shared, and it's yours to provide: you register a &lt;code&gt;LockStore&lt;/code&gt; backed by Redis or Postgres (the docs include both). The in-memory store only excludes other callers in the same process, so &lt;code&gt;LocksModule&lt;/code&gt; refuses to use it in production unless you set &lt;code&gt;allowInMemoryStorage&lt;/code&gt;. Missed ticks aren't replayed, so write jobs that process current state ("export everything not yet exported") rather than "yesterday's rows".&lt;/p&gt;

&lt;p&gt;For code that has to be safe if a lock is lost mid-run (a long GC pause, a network partition), each acquisition comes with a &lt;code&gt;signal&lt;/code&gt; that aborts on lease loss and a monotonically increasing &lt;code&gt;fencingToken&lt;/code&gt;. Pass the signal to your I/O, and use the token to reject writes from a holder whose lock was taken over.&lt;/p&gt;

&lt;p&gt;Queues have their own version of this. BullMQ doesn't retry unless you set &lt;code&gt;attempts&lt;/code&gt;, a job whose worker dies or blocks the event loop past its lock gets marked stalled and run again, and completed and failed jobs stay in Redis until something removes them. The minimum is to set &lt;code&gt;attempts&lt;/code&gt;, &lt;code&gt;backoff&lt;/code&gt; and &lt;code&gt;removeOnComplete&lt;/code&gt;/&lt;code&gt;removeOnFail&lt;/code&gt; explicitly in &lt;code&gt;defaultJobOptions&lt;/code&gt;, and make every job safe to run twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The edge
&lt;/h2&gt;

&lt;h3&gt;
  
  
  5. Make authentication deny-by-default
&lt;/h3&gt;

&lt;p&gt;The usual pattern is &lt;code&gt;@UseGuards(AuthGuard)&lt;/code&gt; on each controller. It fails open: the one controller where someone forgets the decorator is public, and nothing tells you. Invert it. Require authentication globally and opt routes out explicitly.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;@nestjs/authentication&lt;/code&gt; works that way out of the box. &lt;code&gt;AuthenticationModule&lt;/code&gt; registers a global guard, so every route requires a signed-in user and anything else gets a 401, and you mark the exceptions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Public&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Controller&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;auth&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;AuthController&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// sign-in, sign-up, password reset&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It covers sessions, JWT bearer tokens, OpenID Connect, magic links and API keys, and &lt;code&gt;@CurrentUser()&lt;/code&gt; gives handlers the resolved user.&lt;/p&gt;

&lt;p&gt;If you keep your own guard, get the same property by registering it with &lt;code&gt;APP_GUARD&lt;/code&gt; and letting a metadata decorator opt out:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;IS_PUBLIC&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;isPublic&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;Public&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nc"&gt;SetMetadata&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;IS_PUBLIC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Injectable&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;AuthGuard&lt;/span&gt; &lt;span class="k"&gt;implements&lt;/span&gt; &lt;span class="nx"&gt;CanActivate&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;constructor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;reflector&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Reflector&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;

  &lt;span class="nf"&gt;canActivate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;context&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ExecutionContext&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;isPublic&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;reflector&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;getAllAndOverride&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;boolean&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;IS_PUBLIC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
      &lt;span class="nx"&gt;context&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getHandler&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
      &lt;span class="nx"&gt;context&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getClass&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;]);&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;isPublic&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="c1"&gt;// verify the token or session here&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// app.module.ts&lt;/span&gt;
&lt;span class="nl"&gt;providers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;provide&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;APP_GUARD&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;useClass&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;AuthGuard&lt;/span&gt; &lt;span class="p"&gt;}],&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Either way, a review now has one thing to check: every &lt;code&gt;@Public()&lt;/code&gt; in the diff.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Security headers, CORS, rate limits, and the proxy in front
&lt;/h3&gt;

&lt;p&gt;Nest now sets the standard security headers itself, with the same defaults as Helmet 8 (CSP, HSTS, &lt;code&gt;X-Content-Type-Options&lt;/code&gt;, &lt;code&gt;X-Frame-Options&lt;/code&gt; and the rest):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;NestFactory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;AppModule&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;useSecurityHeaders&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It has to be called before &lt;code&gt;app.init()&lt;/code&gt;/&lt;code&gt;app.listen()&lt;/code&gt;, and only once; otherwise it throws. Individual headers are configurable, e.g. &lt;code&gt;app.useSecurityHeaders({ contentSecurityPolicy: { directives: { ... } } })&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;CORS: &lt;code&gt;app.enableCors()&lt;/code&gt; with no arguments allows every origin. For an API used by one frontend with cookies, list the origins:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;enableCors&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;origin&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://app.example.com&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="na"&gt;credentials&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Rate-limit at least the expensive or brute-forceable endpoints (login, password reset, anything that sends email or SMS):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;ThrottlerModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;forRoot&lt;/span&gt;&lt;span class="p"&gt;([{&lt;/span&gt; &lt;span class="na"&gt;ttl&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="nx"&gt;_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt; &lt;span class="p"&gt;}]),&lt;/span&gt;
&lt;span class="c1"&gt;// ...&lt;/span&gt;
&lt;span class="nx"&gt;providers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;provide&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;APP_GUARD&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;useClass&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ThrottlerGuard&lt;/span&gt; &lt;span class="p"&gt;}],&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Counters are per process unless you configure shared storage, so with N replicas the effective limit is N times higher.&lt;/p&gt;

&lt;p&gt;Rate limiting also depends on knowing who the client is, and behind a load balancer every request comes from the load balancer's address. Until Express trusts &lt;code&gt;X-Forwarded-*&lt;/code&gt;, &lt;code&gt;req.ip&lt;/code&gt; is that address, all clients share one limit, and &lt;code&gt;req.protocol&lt;/code&gt; is &lt;code&gt;http&lt;/code&gt; even for HTTPS clients, which breaks secure cookies and audit logs too:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;NestFactory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;create&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;NestExpressApplication&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;AppModule&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;trust proxy&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// number of proxies in front of the app&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set the actual hop count. &lt;code&gt;true&lt;/code&gt; trusts whatever the client sends in &lt;code&gt;X-Forwarded-For&lt;/code&gt;, which lets any client choose its own IP.&lt;/p&gt;

&lt;h2&gt;
  
  
  Input and output
&lt;/h2&gt;

&lt;h3&gt;
  
  
  7. Validate every input
&lt;/h3&gt;

&lt;p&gt;Register a global validation pipe in &lt;code&gt;main.ts&lt;/code&gt;. Which one depends on how you define DTOs.&lt;/p&gt;

&lt;p&gt;With &lt;code&gt;class-validator&lt;/code&gt; classes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;useGlobalPipes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ValidationPipe&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;whitelist&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;forbidNonWhitelisted&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;transform&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;}),&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;whitelist&lt;/code&gt; strips properties that aren't declared on the DTO. That's what stops &lt;code&gt;{ "email": "...", "role": "admin" }&lt;/code&gt; from reaching a &lt;code&gt;repository.save(dto)&lt;/code&gt;. &lt;code&gt;forbidNonWhitelisted&lt;/code&gt; rejects them with a 400 instead, which also surfaces client typos.&lt;/p&gt;

&lt;p&gt;With Zod, Valibot or ArkType, use &lt;code&gt;StandardSchemaValidationPipe&lt;/code&gt; and attach the schema to the parameter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;useGlobalPipes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;StandardSchemaValidationPipe&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;createUserSchema&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;object&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;email&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;email&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
  &lt;span class="na"&gt;password&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;CreateUserDto&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;infer&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;createUserSchema&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Post&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(@&lt;/span&gt;&lt;span class="nd"&gt;Body&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;createUserSchema&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="nx"&gt;dto&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;CreateUserDto&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;users&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;dto&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here unknown keys are the schema's business, not the pipe's. &lt;code&gt;z.object()&lt;/code&gt; strips them by default; use &lt;code&gt;z.strictObject()&lt;/code&gt; if you want them rejected. Either way, the handler receives the schema's output, so coercions and defaults are applied.&lt;/p&gt;

&lt;p&gt;One thing to know about &lt;code&gt;useGlobalPipes()&lt;/code&gt;: it's applied to the app instance, so an e2e test that builds its own app from &lt;code&gt;AppModule&lt;/code&gt; doesn't get it. Either repeat the call in the test setup, or register the pipe as an &lt;code&gt;APP_PIPE&lt;/code&gt; provider so the module brings it along.&lt;/p&gt;

&lt;p&gt;Validation is also where you bound list endpoints. A &lt;code&gt;limit&lt;/code&gt; or &lt;code&gt;take&lt;/code&gt; query parameter without a maximum lets one request load every row in the table into memory; so does a list endpoint with no default. Put the cap in the DTO or schema, and give every list query a default:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;listOrdersQuery&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;object&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;coerce&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;number&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="k"&gt;default&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="na"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(&lt;code&gt;@Max(100)&lt;/code&gt; plus a default value does the same with &lt;code&gt;class-validator&lt;/code&gt;.) The same applies to anything that fans out per item: a bulk endpoint that accepts 10,000 ids is a list endpoint too.&lt;/p&gt;

&lt;p&gt;The body size limit is already reasonable: Express's JSON parser stops at 100 KB. The mistake is raising it globally for one upload endpoint. &lt;code&gt;app.useBodyParser('json', { limit: '50mb' })&lt;/code&gt; applies to every route, and parsing a 50 MB JSON body blocks the event loop for every other request on that process. Stream large payloads on the route that needs them, or upload them directly to object storage.&lt;/p&gt;

&lt;h3&gt;
  
  
  8. Control what goes out
&lt;/h3&gt;

&lt;p&gt;Returning ORM entities from controllers leaks every column, including the ones added after the endpoint was written. Map to response DTOs, or use &lt;code&gt;@Exclude()&lt;/code&gt; with a global &lt;code&gt;ClassSerializerInterceptor&lt;/code&gt;. The interceptor only applies to class instances, so plain objects (raw query results, &lt;code&gt;{ ...user, extra }&lt;/code&gt;) are serialized as-is.&lt;/p&gt;

&lt;p&gt;Don't expose Swagger publicly in production. &lt;code&gt;SwaggerModule.setup()&lt;/code&gt; publishes a complete map of the API. Skip it when &lt;code&gt;NODE_ENV === 'production'&lt;/code&gt;, or put it behind auth.&lt;/p&gt;

&lt;h3&gt;
  
  
  9. Make writes safe to repeat
&lt;/h3&gt;

&lt;p&gt;Clients retry. A mobile app on a flaky connection, a user who double-clicks, a proxy or SDK that retries on a timeout: any of them can send the same &lt;code&gt;POST /orders/:id/pay&lt;/code&gt; twice, and the second one charges the card again. Your handler can't tell a retry from a new request unless the client says so.&lt;/p&gt;

&lt;p&gt;The standard contract is an &lt;code&gt;Idempotency-Key&lt;/code&gt; header: the client generates a key per logical operation and sends it with every attempt, and the server runs the operation once and replays the stored response for repeats. &lt;code&gt;@nestjs/idempotency&lt;/code&gt; implements it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;IdempotencyModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;forRoot&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="c1"&gt;// keys are per user, so two clients can't collide or read each other's results&lt;/span&gt;
  &lt;span class="na"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="na"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="nx"&gt;User&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;:id/pay&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Idempotent&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="nf"&gt;pay&lt;/span&gt;&lt;span class="p"&gt;(@&lt;/span&gt;&lt;span class="nd"&gt;Param&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;id&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Body&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="nx"&gt;dto&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;PayOrderDto&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;CurrentUser&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;User&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pay&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;dto&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A repeat with the same key gets the stored response (with &lt;code&gt;Idempotent-Replayed: true&lt;/code&gt;) and the handler doesn't run. A duplicate that arrives while the first is still running gets a 409 with &lt;code&gt;Retry-After&lt;/code&gt;. The same key with a different body is a 422, which catches client bugs that reuse keys. 5xx responses aren't stored, so a request that failed on your side can be retried and will run again. &lt;code&gt;required: true&lt;/code&gt; rejects requests without a key, which is what you want on payment-like endpoints.&lt;/p&gt;

&lt;p&gt;The records need a shared store (Postgres or Redis; the docs have both), for the same reason as the lock store in item 4: in memory, a retry that lands on another replica, or arrives after a deploy, runs the handler again. The module refuses to start in production without one. Stored responses are kept for 24 hours by default (&lt;code&gt;ttl&lt;/code&gt;), and the whole response body is stored, so keep that in mind for endpoints that return large payloads.&lt;/p&gt;

&lt;p&gt;This is different from &lt;code&gt;@Retry({ idempotent: true })&lt;/code&gt; in item 12, which lets the server re-run a handler within one request. &lt;code&gt;@Idempotent()&lt;/code&gt; deduplicates the client's retries across requests. If you use both, register &lt;code&gt;IdempotencyModule&lt;/code&gt; before &lt;code&gt;ResilienceModule&lt;/code&gt; in &lt;code&gt;imports&lt;/code&gt;, because import order sets the order of their global interceptors.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dependencies
&lt;/h2&gt;

&lt;h3&gt;
  
  
  10. Run current versions of Node, Nest and your dependencies
&lt;/h3&gt;

&lt;p&gt;Old versions are a production risk even when nothing is visibly broken: security fixes, memory leak fixes and bug fixes only land in supported release lines, and the longer you wait, the bigger and riskier the eventual upgrade.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Node&lt;/strong&gt;: run an LTS release that's still supported, Active or Maintenance. Odd-numbered majors never become LTS, and an end-of-life line gets no security patches at all: Node 20 reached end-of-life in April 2026, so anything still on it is unpatched. Pin the major in the base image (&lt;code&gt;node:24-slim&lt;/code&gt;) and rebuild regularly so patch releases actually reach production, rather than building once and running that image for a year.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nest&lt;/strong&gt;: stay on the current major, and keep every &lt;code&gt;@nestjs/*&lt;/code&gt; package on matching versions. Mismatched &lt;code&gt;@nestjs/common&lt;/code&gt; and &lt;code&gt;@nestjs/core&lt;/code&gt; versions cause confusing DI and metadata errors. Several items in this list (&lt;code&gt;useSecurityHeaders()&lt;/code&gt;, &lt;code&gt;StandardSchemaValidationPipe&lt;/code&gt;, &lt;code&gt;return503OnClosing&lt;/code&gt;) only exist in recent releases.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Everything else&lt;/strong&gt;: commit the lockfile, install with &lt;code&gt;npm ci&lt;/code&gt; in CI and in the image, and let Renovate or Dependabot open small update PRs continuously. Twenty small upgrades with passing tests are much cheaper than one big one after two years. Run &lt;code&gt;npm audit --omit=dev&lt;/code&gt; in CI to catch known vulnerabilities in what actually ships.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  11. Time out every outbound call
&lt;/h3&gt;

&lt;p&gt;A dependency that stops responding is worse than one that errors. Every request waiting on it holds memory, a socket and often a database connection, and under load you run out of those before the dependency recovers. Few clients time out by default:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;fetch&lt;/code&gt; waits up to 5 minutes for response headers. Pass &lt;code&gt;signal: AbortSignal.timeout(5_000)&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;For HTTP APIs, use &lt;code&gt;@nestjs/http-client&lt;/code&gt;, the Promise-based replacement for the axios-based &lt;code&gt;HttpModule&lt;/code&gt;. Configure a timeout per client:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;  &lt;span class="nx"&gt;HttpClientModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;register&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;payments&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;baseUrl&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://api.payments.example.com&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;timeout&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;5s&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;}),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The timeout applies per attempt, and retries are on by default for idempotent methods (&lt;code&gt;GET&lt;/code&gt;, &lt;code&gt;HEAD&lt;/code&gt;, &lt;code&gt;OPTIONS&lt;/code&gt;, &lt;code&gt;PUT&lt;/code&gt;, &lt;code&gt;DELETE&lt;/code&gt;: 3 attempts with exponential backoff, and on timeouts too). So a 5 s timeout can mean 15 s or more before the call fails. Make sure that fits the caller's budget, or tune &lt;code&gt;retry&lt;/code&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Database queries: set a statement timeout. With &lt;code&gt;pg&lt;/code&gt;, that's &lt;code&gt;statement_timeout&lt;/code&gt; in the pool config, whatever ORM sits on top.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  12. Plan for dependencies that are down
&lt;/h3&gt;

&lt;p&gt;Timeouts cap one call. When a dependency is down for minutes, you also want to stop calling it for a while, limit how much of your capacity it can occupy, and serve something reasonable in the meantime. &lt;code&gt;@nestjs/resilience&lt;/code&gt; provides that as decorators and presets:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;ResilienceModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;forRoot&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;presets&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;carrier&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;timeout&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;2s&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;retry&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;attempts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="na"&gt;circuitBreaker&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;failureRateThreshold&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;minimumCalls&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;openDuration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;30s&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;quotes&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Resilience&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;carrier&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Fallback&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;flatRateQuotes&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;handleIf&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt; &lt;span class="k"&gt;instanceof&lt;/span&gt; &lt;span class="nx"&gt;CircuitOpenError&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="nf"&gt;getQuotes&lt;/span&gt;&lt;span class="p"&gt;(@&lt;/span&gt;&lt;span class="nd"&gt;Query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;orderId&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;orderId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Signal&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;AbortSignal&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;shipping&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getQuotes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;orderId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the breaker is open, calls fail immediately with a 503 and &lt;code&gt;Retry-After&lt;/code&gt; instead of waiting 2 s each, and the fallback answers. &lt;code&gt;@Bulkhead()&lt;/code&gt; caps concurrent calls, so a slow dependency can't tie up every request on the process. Pass the injected &lt;code&gt;signal&lt;/code&gt; down to your I/O so a timed-out attempt actually stops. Inside services (jobs, cron, anything that isn't a handler), use &lt;code&gt;ResilienceService.preset('carrier').execute(...)&lt;/code&gt;. Two details: &lt;code&gt;@Retry()&lt;/code&gt; only applies to safe methods unless the handler is also &lt;code&gt;@Idempotent()&lt;/code&gt; (item 9) and the retry is marked &lt;code&gt;idempotent: true&lt;/code&gt;, and retries multiply across layers. A resilience retry around an &lt;code&gt;@nestjs/http-client&lt;/code&gt; call that retries itself makes 3 x 3 = 9 attempts, so decide which layer owns retries.&lt;/p&gt;

&lt;h3&gt;
  
  
  13. Run migrations as a deploy step
&lt;/h3&gt;

&lt;p&gt;Every ORM has a way to sync the schema directly from your model: TypeORM's &lt;code&gt;synchronize: true&lt;/code&gt;, &lt;code&gt;drizzle-kit push&lt;/code&gt;, &lt;code&gt;prisma db push&lt;/code&gt;. They're convenient in development and unsafe against production data, because they apply whatever diff they compute, and a renamed field can look like a drop plus an add.&lt;/p&gt;

&lt;p&gt;In production, apply reviewed migrations (TypeORM migrations, &lt;code&gt;drizzle-kit migrate&lt;/code&gt;, &lt;code&gt;prisma migrate deploy&lt;/code&gt;) as a separate step before the new version starts: a Kubernetes Job, a release phase, a CI step. Don't run them from app startup (&lt;code&gt;migrationsRun: true&lt;/code&gt; and equivalents) when you have several replicas, because they all start at once and race.&lt;/p&gt;

&lt;p&gt;Also check that migrations are backward compatible with the version still running. During a rolling deploy, old and new code share the new schema. Dropping or renaming a column the old version still reads breaks it for the length of the rollout, so split those into expand and contract steps across two deploys.&lt;/p&gt;

&lt;p&gt;The same rule is what makes rollback possible. The goal is that at any moment you can redeploy the previous image and it works against the current schema, because down-migrations under incident pressure are rarely tested and often lossy. If a deploy can't be rolled back that way, it should be a deliberate, reviewed exception, not something you discover during an incident.&lt;/p&gt;

&lt;h3&gt;
  
  
  14. Check the connection pool math
&lt;/h3&gt;

&lt;p&gt;Every replica has its own pool. Five replicas with a pool of 20 is 100 connections. That's Postgres's default &lt;code&gt;max_connections&lt;/code&gt;, and you haven't counted the migration job, the workers, an admin tool, or the deploy window where old and new pods run at the same time.&lt;/p&gt;

&lt;p&gt;Set the pool size explicitly (for &lt;code&gt;pg&lt;/code&gt;, the &lt;code&gt;max&lt;/code&gt; option of the pool, however your ORM passes it through). Multiply by the maximum replica count during a deploy, and check it fits. When it no longer fits, add a pooler like PgBouncer rather than raising &lt;code&gt;max_connections&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  15. Watch for request-scoped chains
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;Scope.REQUEST&lt;/code&gt; propagates up the dependency graph. Anything that injects a request-scoped provider becomes request-scoped, and so does anything that injects that, up to the controllers. A request-scoped logger near the bottom of the graph can mean a dozen objects constructed per request across a module, without any code saying so.&lt;/p&gt;

&lt;p&gt;Check for it before launch. If you need per-request data rather than per-request instances, &lt;code&gt;AsyncLocalStorage&lt;/code&gt; (or &lt;code&gt;nestjs-cls&lt;/code&gt;) provides it without rebuilding the graph. If you need per-tenant instances, durable providers reuse them per tenant instead of creating them per request.&lt;/p&gt;

&lt;h2&gt;
  
  
  Errors and visibility
&lt;/h2&gt;

&lt;h3&gt;
  
  
  16. Errors: statuses, logs, source maps
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Keep the default exception handling semantics&lt;/strong&gt;: real statuses for &lt;code&gt;HttpException&lt;/code&gt;, an opaque 500 plus a log entry for everything else. If you customise the response shape, extend &lt;code&gt;BaseExceptionFilter&lt;/code&gt; and delegate to it. A filter written from scratch drops Nest's &lt;code&gt;ExceptionsHandler&lt;/code&gt; log line, and unhandled errors stop appearing in logs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Errors outside handlers.&lt;/strong&gt; An unhandled promise rejection crashes the process (Node's default since v15). That's the correct behaviour. Make sure the error is logged with its stack before exit, and that repeated restarts alert someone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Source maps.&lt;/strong&gt; The starter already emits them. Start Node with &lt;code&gt;--enable-source-maps&lt;/code&gt; (or &lt;code&gt;NODE_OPTIONS=--enable-source-maps&lt;/code&gt;) so stack traces point at &lt;code&gt;.ts&lt;/code&gt; files and lines.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  17. Logs, health checks, monitoring
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Logs&lt;/strong&gt;: JSON, production log levels, structured params, and a trace id on every line. &lt;a href="//./01-how-to-monitor-a-nestjs-app.md"&gt;How to monitor a NestJS app&lt;/a&gt; covers each.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Health checks&lt;/strong&gt;: separate liveness from readiness. Liveness asks whether this process is stuck, and it shouldn't check dependencies: if the database goes down and liveness fails, Kubernetes restarts every pod, which doesn't help and adds a reconnect storm when the database comes back. Readiness asks whether the pod should receive traffic, and that's where the &lt;code&gt;@nestjs/terminus&lt;/code&gt; database check goes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Monitoring&lt;/strong&gt;: the items above prevent known failure modes. For the rest, you need per-route latency and error rates, traces that show which of your methods and queries took the time, errors grouped into defects with an alert on new ones, and visibility into jobs and scheduled runs.&lt;/p&gt;

&lt;p&gt;For Nest specifically, that's what we built &lt;a href="https://observe.nestjs.com" rel="noopener noreferrer"&gt;NestJS Observe&lt;/a&gt; for. It instruments through Nest's DI container, so spans are your own providers, and setup is a module and one option:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// app.module.ts&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;ObserveModule&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ObserveInstrument&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;createObserveModule&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Module&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;imports&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="nx"&gt;ObserveModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;forRoot&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;appKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;OBSERVE_APP_KEY&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;appSecret&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;OBSERVE_APP_SECRET&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;serviceId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;orders-api&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="c1"&gt;// nothing is ignored by default; keep probes out of the latency numbers&lt;/span&gt;
      &lt;span class="na"&gt;http&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;ignore&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;/^&lt;/span&gt;&lt;span class="se"&gt;\/&lt;/span&gt;&lt;span class="sr"&gt;health&lt;/span&gt;&lt;span class="se"&gt;(?:\/&lt;/span&gt;&lt;span class="sr"&gt;|&lt;/span&gt;&lt;span class="se"&gt;\?&lt;/span&gt;&lt;span class="sr"&gt;|$&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;AppModule&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;

&lt;span class="c1"&gt;// main.ts&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;NestFactory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;AppModule&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;instrument&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ObserveInstrument&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That gives you requests, jobs, cron runs and messages with p95 and error rate, traces with the SQL under each method, unhandled errors grouped with the failing source line and an email on new ones, and trace ids on &lt;code&gt;ConsoleLogger&lt;/code&gt; output. The free tier covers 300,000 events a month.&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Config is validated at boot, required values use &lt;code&gt;getOrThrow&lt;/code&gt;, secrets come from a secret store, and the config object is never logged.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;enableShutdownHooks()&lt;/code&gt; is on, Node receives &lt;code&gt;SIGTERM&lt;/code&gt; directly, and there's an init process; the image is multi-stage, runs as non-root, and the heap limit fits the container's memory limit.&lt;/li&gt;
&lt;li&gt;Pods drain before closing, and the grace period covers the longest job.&lt;/li&gt;
&lt;li&gt;Scheduled jobs are guarded with &lt;code&gt;@OnOneInstance()&lt;/code&gt;/&lt;code&gt;@WithoutOverlapping()&lt;/code&gt; on a shared lock store, and every job is idempotent.&lt;/li&gt;
&lt;li&gt;Authentication is global, and every &lt;code&gt;@Public()&lt;/code&gt; is deliberate.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;useSecurityHeaders()&lt;/code&gt; is on, CORS lists its origins, sensitive endpoints are rate-limited, and &lt;code&gt;trust proxy&lt;/code&gt; matches the number of proxies.&lt;/li&gt;
&lt;li&gt;A global &lt;code&gt;ValidationPipe&lt;/code&gt; (with &lt;code&gt;whitelist&lt;/code&gt;) or &lt;code&gt;StandardSchemaValidationPipe&lt;/code&gt; is registered, list endpoints have a maximum page size, and the body limit is unchanged.&lt;/li&gt;
&lt;li&gt;Responses are DTOs or serialized classes; Swagger isn't public.&lt;/li&gt;
&lt;li&gt;Endpoints that must not run twice require an &lt;code&gt;Idempotency-Key&lt;/code&gt;, backed by a shared store.&lt;/li&gt;
&lt;li&gt;Node is on a supported LTS, &lt;code&gt;@nestjs/*&lt;/code&gt; packages are current and aligned, and dependency updates arrive continuously.&lt;/li&gt;
&lt;li&gt;Every outbound call and query has a timeout, and retry budgets add up.&lt;/li&gt;
&lt;li&gt;Unreliable dependencies have a circuit breaker, a bulkhead or a fallback, and retries are owned by one layer.&lt;/li&gt;
&lt;li&gt;Migrations run as a separate deploy step and are backward compatible, so the previous image can always be redeployed.&lt;/li&gt;
&lt;li&gt;Pool size times peak replica count fits in the database.&lt;/li&gt;
&lt;li&gt;No unintended request-scoped chains.&lt;/li&gt;
&lt;li&gt;Error handling keeps statuses and logs; stack traces resolve to source.&lt;/li&gt;
&lt;li&gt;Logs are structured, liveness doesn't depend on the database, and the service is monitored.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you'd rather see the monitoring side of this (item 17) in action than read about it, here's a walkthrough of setting up NestJS Observe: &lt;a href="https://www.youtube.com/watch?v=kg-JPz7Fcek" rel="noopener noreferrer"&gt;NestJS Observe: Zero-Config Observability for NestJS&lt;/a&gt; on YouTube.&lt;/p&gt;

</description>
      <category>nestjs</category>
      <category>node</category>
      <category>security</category>
    </item>
    <item>
      <title>How to monitor a NestJS app</title>
      <dc:creator>Kamil Mysliwiec</dc:creator>
      <pubDate>Mon, 21 Sep 2026 13:07:01 +0000</pubDate>
      <link>https://dev.to/nestjs/how-to-monitor-a-nestjs-app-5c4b</link>
      <guid>https://dev.to/nestjs/how-to-monitor-a-nestjs-app-5c4b</guid>
      <description>&lt;h2&gt;
  
  
  How to monitor a NestJS app
&lt;/h2&gt;

&lt;p&gt;When something breaks in production, the first question is usually why it happened. Monitoring gives you a much better answer than a collection of isolated logs, so here's a practical way to think about the signals that help.&lt;/p&gt;

&lt;p&gt;Many applications start with logs and a health check, while a broader monitoring plan waits for a quieter week. That's a reasonable place to begin. Logs are much more useful when you already know what to look for, though, and during an incident that context is often the first thing you need.&lt;/p&gt;

&lt;p&gt;So let's go through what's actually worth having, roughly in the order I'd add it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Health checks
&lt;/h2&gt;

&lt;p&gt;This is a useful first layer, and Nest has &lt;code&gt;@nestjs/terminus&lt;/code&gt; for it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Controller&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;health&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;HealthController&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;constructor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="nx"&gt;health&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;HealthCheckService&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;TypeOrmHealthIndicator&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;

  &lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Get&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;HealthCheck&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;health&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;([()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pingCheck&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;database&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)]);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Point your load balancer at &lt;code&gt;/health&lt;/code&gt; and it knows when an instance is dead. Keep in mind that's all it knows. An app can answer &lt;code&gt;/health&lt;/code&gt; in 2 ms while every real request takes 9 seconds, and as far as the load balancer is concerned everything is fine.&lt;/p&gt;

&lt;p&gt;It is worth excluding this endpoint from the other instrumentation below. A probe every few seconds creates thousands of identical requests a day and can pull latency numbers toward a route that does very little real work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Logs
&lt;/h2&gt;

&lt;p&gt;The built-in &lt;code&gt;Logger&lt;/code&gt; is often enough during development. In production, a few &lt;code&gt;ConsoleLogger&lt;/code&gt; options make a real difference:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;levels&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;LOG_LEVELS&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;fatal,error,warn,log&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;LogLevel&lt;/span&gt;&lt;span class="p"&gt;[];&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;logger&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ConsoleLogger&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;json&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;flattenParams&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;logLevels&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;levels&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;NestFactory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;AppModule&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;logger&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;json: true&lt;/code&gt; gives your log system one object per line instead of coloured text it has to parse back apart.&lt;/p&gt;

&lt;h3&gt;
  
  
  Levels are a contract
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;logLevels&lt;/code&gt; decides what gets written at all, but the levels only help if everyone on the team uses them the same way. The split I'd use:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;fatal&lt;/code&gt;: the process can't continue. Bad config at startup, a lost connection that won't come back.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;error&lt;/code&gt;: an operation failed and somebody should look at it. If nobody would act on it, it isn't an &lt;code&gt;error&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;warn&lt;/code&gt;: something went wrong and was handled. A retry, a fallback, a degraded response. One of these is noise; a hundred in a minute is an early warning.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;log&lt;/code&gt;: business events you'll want during an incident. Order placed, payment captured, job finished.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;debug&lt;/code&gt; and &lt;code&gt;verbose&lt;/code&gt;: off in production. They bury the lines you're looking for, and you pay to ship and store every one of them.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Structured params
&lt;/h3&gt;

&lt;p&gt;Since Nest 12, plain objects passed after the message are treated as data attached to the entry, not printed as extra messages. With &lt;code&gt;flattenParams&lt;/code&gt;, their keys land at the top level of the JSON line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;warn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Payment declined&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;orderId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;stripe&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;attempt&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="c1"&gt;// {"level":"warn",...,"message":"Payment declined","context":"PaymentsService","orderId":"ord_91f2","provider":"stripe","attempt":2}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now your log system can filter on &lt;code&gt;orderId&lt;/code&gt; and group by &lt;code&gt;provider&lt;/code&gt; instead of you grepping text. A few rules make this pay off:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Constant message, variables in params.&lt;/strong&gt; "Payment declined" written 400 times is a pattern you can count and alert on. 400 different strings with the order id baked into each one are not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Objects, never extra strings.&lt;/strong&gt; &lt;code&gt;this.logger.log('Order shipped', order.id)&lt;/code&gt; doesn't do what it looks like. A trailing string is how Nest passes a context, so on a context-less logger the id replaces the context, and on a &lt;code&gt;new Logger(OrdersService.name)&lt;/code&gt; it's printed as a second, separate log line. &lt;code&gt;{ orderId: order.id }&lt;/code&gt; is always safe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Errors keep their stack.&lt;/strong&gt; Params go before the stack: &lt;code&gt;this.logger.error('Charge failed', { orderId }, err.stack)&lt;/code&gt; produces one line with &lt;code&gt;orderId&lt;/code&gt; and a &lt;code&gt;stack&lt;/code&gt; field.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Framework fields win.&lt;/strong&gt; A param called &lt;code&gt;message&lt;/code&gt;, &lt;code&gt;level&lt;/code&gt; or &lt;code&gt;timestamp&lt;/code&gt; is dropped silently rather than overwriting the real one. Name it &lt;code&gt;reason&lt;/code&gt; or &lt;code&gt;status&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ids, not objects.&lt;/strong&gt; &lt;code&gt;{ user }&lt;/code&gt; ships the whole entity, five levels deep, including whatever PII is on it. &lt;code&gt;{ userId: user.id }&lt;/code&gt; is what you'll actually filter on.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Trace ids
&lt;/h3&gt;

&lt;p&gt;What the built-in logger can't give you is a trace id on every line. Without one, finding the dozen entries for a failed request among thousands of lines written that minute is guesswork, and filtering by &lt;code&gt;orderId&lt;/code&gt; only helps if every line in the request happened to include it.&lt;/p&gt;

&lt;p&gt;A hand-rolled request id (a middleware plus &lt;code&gt;AsyncLocalStorage&lt;/code&gt;) gets you part of the way, but it only covers HTTP. Queue jobs, cron runs and microservice handlers never pass through that middleware, and it isn't the id of anything else you can look at: not a trace, not an error report.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ConsoleLogger&lt;/code&gt; has no idea which operation it's logging for, so something that does has to add the id. That's one of the things &lt;code&gt;@nestjs/observe&lt;/code&gt; does. It hooks into &lt;code&gt;ConsoleLogger&lt;/code&gt;, and every line written during a traced operation (a request, a queue job, a cron run, a microservice message) gets the id of that trace:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"level"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"warn"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"pid"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4242&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"timestamp"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1789992138000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Payment declined"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"context"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"PaymentsService"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"orderId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ord_91f2"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stripe"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"attempt"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"traceId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"4bf92f3577b34da6a3ce929d0e0e4736"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The id is read when the line is written, so it's the right one even when a hundred requests interleave across &lt;code&gt;await&lt;/code&gt;s. It's the same id as the trace in the dashboard, so even if your logs stay in your own pipeline, you can go from a slow request to its log lines and back. If you log through pino or winston instead, the hook doesn't apply; &lt;code&gt;TracerService.currentTraceId()&lt;/code&gt; gives you the id to add yourself, for example from pino's &lt;code&gt;mixin&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Metrics
&lt;/h2&gt;

&lt;p&gt;Requests per second, error rate, latency, memory, event loop delay. Cheap to collect, and they're what you put alerts on.&lt;/p&gt;

&lt;p&gt;Two choices matter here. First, average latency can hide one request in twenty taking four seconds, so p95 is usually a more useful view of the slower experience. Second, break the data down by route. A global number changes when your traffic mix changes, even if the code has not become slower.&lt;/p&gt;

&lt;p&gt;Metrics tell you that something changed. They never tell you why.&lt;/p&gt;

&lt;h2&gt;
  
  
  Traces
&lt;/h2&gt;

&lt;p&gt;Traces are often what turns a broad symptom into a short list of likely causes, although they do take more setup than logs and metrics.&lt;/p&gt;

&lt;p&gt;A trace is a single request broken into the things that happened during it: the guard, the controller method, the service it called, the three queries that service ran, the call to Stripe. Each one with a start time and a duration. Instead of "it's probably the database" you get a picture, and the picture is usually surprising.&lt;/p&gt;

&lt;p&gt;The traditional setup can be substantial: the OpenTelemetry SDK, auto-instrumentation packages, an exporter, a collector, and a backend to store the data. Depending on the existing infrastructure, it can take a few days before the first useful trace is available, which is difficult to prioritize alongside product work.&lt;/p&gt;

&lt;p&gt;If you do go shopping for a tracing tool, here's what I'd look for in a Nest app specifically. Spans named after your code (&lt;code&gt;OrdersService.create&lt;/code&gt;, not &lt;code&gt;middleware - &amp;lt;anonymous&amp;gt;&lt;/code&gt;). Self time and not only total time, because a controller that awaits a slow repository didn't do anything wrong and you don't want it at the top of your list. Database queries and outbound HTTP calls as their own spans, with the statement visible. And traces that survive a queue, so the request that enqueues a job and the job itself are one story.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg8xa6esoxu7i2dkanauu.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg8xa6esoxu7i2dkanauu.webp" alt="SQL query" width="800" height="364"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Error tracking
&lt;/h2&gt;

&lt;p&gt;Logs contain errors, but they do not automatically tell you that this &lt;code&gt;TypeError&lt;/code&gt; has happened 3,140 times since the last deploy, that those occurrences represent one bug, or that the bug is new.&lt;/p&gt;

&lt;p&gt;That's what error tracking is for: group occurrences into actual defects, remember when each one first showed up and in which release, and tell you when a new one appears. It should also know the difference between errors you threw on purpose (a &lt;code&gt;NotFoundException&lt;/code&gt; is your API doing its job) and the ones you didn't see coming. If those two end up in the same bucket, your error rate is permanently at 4% and everybody learns to ignore it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Queues
&lt;/h2&gt;

&lt;p&gt;If you use BullMQ or Bull, half your app runs somewhere no HTTP dashboard can see. For each queue I want to know how long jobs wait before a worker picks them up, how often they fail and on which attempt, and which request enqueued them. I wrote a separate post about this because it deserves one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Doing all of this without losing a week
&lt;/h2&gt;

&lt;p&gt;You can assemble everything above from separate tools. Terminus, a log shipper, Prometheus and Grafana, an OpenTelemetry pipeline, an error tracker, something for queues. Lots of teams run exactly that and it works. It's also five or six systems to keep alive.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This is the reason we built &lt;a href="https://observe.nestjs.com" rel="noopener noreferrer"&gt;NestJS Observe&lt;/a&gt;&lt;/strong&gt;. We're the framework team, so we could do something nobody else can: instrument the app from the inside, using Nest's own hooks, instead of wrapping a generic Node.js agent around it. The whole setup is this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; @nestjs/observe
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and with this package installed, the rest is configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// app.module.ts&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;createObserveModule&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@nestjs/observe&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;ObserveModule&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ObserveInstrument&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;createObserveModule&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Module&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;imports&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="nx"&gt;ObserveModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;forRoot&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;appKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;OBSERVE_APP_KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;appSecret&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;OBSERVE_APP_SECRET&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;serviceId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;orders-api&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;http&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="c1"&gt;// keep the load balancer probe out of your percentiles&lt;/span&gt;
        &lt;span class="na"&gt;ignore&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;/^&lt;/span&gt;&lt;span class="se"&gt;\/&lt;/span&gt;&lt;span class="sr"&gt;health&lt;/span&gt;&lt;span class="se"&gt;(?:\?&lt;/span&gt;&lt;span class="sr"&gt;|$&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;AppModule&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// main.ts&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;NestFactory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;AppModule&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;instrument&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ObserveInstrument&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And that's it. No spans to write, nothing to preload, no collector to run. From that point you get requests, GraphQL operations, microservice messages, WebSocket messages, queue jobs and cron jobs, per route, with throughput, error rate and p95. Every request has a trace where the spans are your own classes and methods. SQL queries show up under the method that ran them (through &lt;code&gt;pg&lt;/code&gt;, &lt;code&gt;mysql2&lt;/code&gt; and &lt;code&gt;mongodb&lt;/code&gt;, so TypeORM, Drizzle, MikroORM and Mongoose just work), and so do outbound HTTP calls. We only record the shape of a statement, never the values. Unhandled errors come with the failing source line, get grouped into defects, and you get an email when a new one shows up. Jobs carry their wait time and attempts, and they share a trace with the request that enqueued them. And every line your &lt;code&gt;ConsoleLogger&lt;/code&gt; writes during any of that carries its trace id.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F21l1l9g60zy70vxkcqn5.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F21l1l9g60zy70vxkcqn5.webp" alt="NestJS Observe Dashboard" width="800" height="381"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It needs &lt;code&gt;@nestjs/core&lt;/code&gt; 11.1 or newer and works with both Express and Fastify. The free tier is 300,000 events a month, which is plenty to find out if it's useful to you. Trace ids in your logs are on every plan. Log forwarding (the lines themselves shipped to Observe and placed on the trace's timeline, with redaction on by default) and alert rules are on the paid plans.&lt;/p&gt;

&lt;p&gt;Two caveats. Keep your Terminus health check, because your orchestrator still needs an endpoint to hit. And if your company runs five languages and has standardised on OpenTelemetry, stick with that. A shared standard across your whole stack is worth more than a shortcut for one framework.&lt;/p&gt;

&lt;h2&gt;
  
  
  If I were starting from zero today
&lt;/h2&gt;

&lt;p&gt;Health check first, excluded from everything else. Then JSON logs with production log levels and structured params. Then tracing and error tracking together, since they come from the same instrumentation and an error without its trace is half a story, and the same instrumentation puts a trace id on every log line. Alerts on error rate and p95 once you have a week of data to set thresholds against. Queue monitoring the day you add a queue.&lt;/p&gt;

&lt;p&gt;The goal is not a collection of attractive dashboards. It is a shorter path from a production symptom to the line of code or dependency that needs attention.&lt;/p&gt;

&lt;p&gt;You can click around a live project without signing up here: &lt;a href="https://www.observe-demo.nestjs.com/dashboard" rel="noopener noreferrer"&gt;the demo dashboard&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>nestjs</category>
      <category>monitoring</category>
      <category>telemetry</category>
      <category>apm</category>
    </item>
    <item>
      <title>Announcing NestJS Monorepos and new CLI commands</title>
      <dc:creator>Kamil Mysliwiec</dc:creator>
      <pubDate>Tue, 01 Oct 2019 12:52:25 +0000</pubDate>
      <link>https://dev.to/trilon/announcing-nestjs-monorepos-and-new-cli-commands-1n0b</link>
      <guid>https://dev.to/trilon/announcing-nestjs-monorepos-and-new-cli-commands-1n0b</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on the &lt;a href="https://trilon.io/blog/announcing-nestjs-monorepos-and-new-commands" rel="noopener noreferrer"&gt;Trilon Blog&lt;/a&gt; on Sep 31, 2019.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In this blog post, we'll be looking at a new API of the latest Nest CLI, which supports an alternate structure for managing multiple projects and libraries in a one single repo called &lt;strong&gt;monorepo&lt;/strong&gt;. In addition, I'll give you some insights about the new CLI commands that have just been introduced, respectively &lt;code&gt;nest build&lt;/code&gt; and &lt;code&gt;nest start&lt;/code&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If you're not familiar with &lt;strong&gt;&lt;a href="https://nestjs.com" rel="noopener noreferrer"&gt;NestJS&lt;/a&gt;&lt;/strong&gt;, it is a TypeScript Node.js framework that helps you build enterprise-grade efficient and scalable Node.js applications.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  History
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Before we dive further, let's take a step back to see how everything was handled in the past&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Until now, we used TypeScript &lt;code&gt;tsc&lt;/code&gt; compiler by default. To provide a good developer experience in the development environment (e.g. reload the application on file change), we utilized &lt;code&gt;nodemon&lt;/code&gt;, &lt;code&gt;ts-node&lt;/code&gt; and &lt;code&gt;tsc-watch&lt;/code&gt;.&lt;br&gt;
This was perfect for most projects, but we found that some members of the community moved toward &lt;code&gt;webpack&lt;/code&gt; in combination with &lt;code&gt;ts-loader&lt;/code&gt;. That means, that eventually several packages were needed in order to handle basic features which led to various side-effects and inconsistencies.&lt;/p&gt;

&lt;p&gt;Similarly, some companies needed to follow a monorepo approach, instead of having a separate repository for every single application (or library).&lt;br&gt;
Consequently, even more libraries, tools and different packages become required.&lt;br&gt;
In order to solve this problem, we made a decision to address all these issues directly in the official CLI.   &lt;/p&gt;


&lt;h2&gt;
  
  
  Builders 🏗
&lt;/h2&gt;

&lt;p&gt;With &lt;code&gt;nest build&lt;/code&gt;, you can use the same command to compile your application or library in either development or production environment.&lt;br&gt;
Would you like to watch for changes and recompile on each source file modification? Use &lt;code&gt;nest build --watch&lt;/code&gt;. Would you like to switch to &lt;a href="https://webpack.js.org/" rel="noopener noreferrer"&gt;webpack&lt;/a&gt;? Use &lt;code&gt;nest build --webpack&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Ftrilon.io%2F_nuxt%2Fimg%2Ff86c9b0.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Ftrilon.io%2F_nuxt%2Fimg%2Ff86c9b0.gif" alt="nest build" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;code&gt;nest build --watch --webpack&lt;/code&gt; example


&lt;p&gt;But the builder itself is not just a wrapper around the compiler (&lt;code&gt;webpack&lt;/code&gt; or &lt;code&gt;tsc&lt;/code&gt;).&lt;br&gt;
It also has plugin system that allows you to leverage the build process itself for pre or post compilation (e.g. to automatically provide additional metadata for &lt;code&gt;@nestjs/swagger&lt;/code&gt; to reduce the boilerplate). &lt;br&gt;
In fact, one plugin is already built-in into the compiler. This plugin will automatically resolve your path aliases (e.g. &lt;code&gt;@trilon/core&lt;/code&gt; imports) so you will no longer have to use helper packages like &lt;code&gt;tsconfig-paths&lt;/code&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;HINT:&lt;/strong&gt; Learn more about the available builder options &lt;a href="https://docs.nestjs.com/cli/usages#nest-build" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;
  
  
  Start 🚀
&lt;/h2&gt;

&lt;p&gt;In the past you may have been confused by all of the different &lt;code&gt;package.json&lt;/code&gt; scripts in a starter application (&lt;code&gt;ts-node&lt;/code&gt;, &lt;code&gt;tsc-watch&lt;/code&gt; and &lt;code&gt;nodemon&lt;/code&gt;). &lt;br&gt;
Thanks to &lt;code&gt;nest start&lt;/code&gt; (and &lt;code&gt;nest start --watch&lt;/code&gt;), all listed packages become useless. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Ftrilon.io%2F_nuxt%2Fimg%2F8e06663.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Ftrilon.io%2F_nuxt%2Fimg%2F8e06663.gif" alt="nest start" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;code&gt;nest start --watch&lt;/code&gt; example


&lt;p&gt;In the past, you may have faced an issue after adding a single TS file in the root directory of your project - and &lt;code&gt;npm run start:*&lt;/code&gt; stopped working. But rest assured - that will no longer be the case. Nest CLI will automatically detect whether your output files are located within &lt;code&gt;dist&lt;/code&gt; or &lt;code&gt;dist/src&lt;/code&gt; directory and execute appropriate entry file based on that assumption.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;HINT:&lt;/strong&gt; Learn more about the available &lt;code&gt;nest start&lt;/code&gt; options &lt;a href="https://docs.nestjs.com/cli/usages#nest-start" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;
  
  
  Monorepos 🐱
&lt;/h2&gt;

&lt;p&gt;Over the last few years, monorepos became quite popular in the developer community. And even though there are downsides of using monorepos, the &lt;strong&gt;benefits&lt;/strong&gt; they bring provide substantial value.&lt;br&gt;
Monorepos make it easier to compose modular components and libraries, track cross-project changes, promote code re-use, and make integration testing simpler.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Ftrilon.io%2F_nuxt%2Fimg%2F2a088ab.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Ftrilon.io%2F_nuxt%2Fimg%2F2a088ab.png" alt="Monorepo diagram" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;
Monorepo that consist of 2 applications (payments and alerts) and 1 shared library

&lt;h3&gt;
  
  
  Applications and libraries
&lt;/h3&gt;

&lt;p&gt;In order to meet the demands of the community, we've added &lt;code&gt;nest g app&lt;/code&gt; and &lt;code&gt;nest g lib&lt;/code&gt; commands that allow you to convert the existing structure to a monorepo mode structure. &lt;br&gt;
For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;nest g app alert-service
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;will scaffold the sub-application alongside your existing application within the same &lt;a href="https://docs.nestjs.com/cli/workspaces" rel="noopener noreferrer"&gt;workspace&lt;/a&gt;. These two applications will share the same &lt;code&gt;node_modules&lt;/code&gt; folder (&lt;em&gt;single-version policy&lt;/em&gt;) and configuration files (e.g. &lt;code&gt;tsconfig.json&lt;/code&gt; and &lt;code&gt;nest-cli.json&lt;/code&gt;).&lt;br&gt;
However, these applications can be executed, developed and deployed &lt;strong&gt;separately&lt;/strong&gt;. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Ftrilon.io%2F_nuxt%2Fimg%2F6863795.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Ftrilon.io%2F_nuxt%2Fimg%2F6863795.gif" alt="nest g app" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;code&gt;nest g app&lt;/code&gt; example


&lt;p&gt;Similarly, to generate a &lt;a href="https://docs.nestjs.com/cli/libraries" rel="noopener noreferrer"&gt;library&lt;/a&gt; (a general-purpose feature that can be used within multiple projects), you can use the following command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;nest g lib &lt;span class="nb"&gt;users&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt; 
  Note that a library created inside a monorepo is &lt;strong&gt;not suitable&lt;/strong&gt; for publishing to the NPM registry. If you want to create a package that you want to publish / share across different repos, you should rather use &lt;code&gt;nest new&lt;/code&gt;. 
&lt;/blockquote&gt;

&lt;p&gt;Both these commands will automatically update the &lt;code&gt;nest-cli.json&lt;/code&gt; which holds the metadata needed to build and organize workspace projects. Typically though, you won't have to edit its content manually (unless you want to change the default file names etc.). You can read more about the monorepo mode &lt;a href="https://docs.nestjs.com/cli/monorepo#monorepo-mode" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Building projects
&lt;/h3&gt;

&lt;p&gt;To build a single project, you can simply call &lt;code&gt;$ nest build NAME&lt;/code&gt; command where &lt;code&gt;NAME&lt;/code&gt; is the name of the application / library you passed to the &lt;code&gt;$ nest g&lt;/code&gt; command. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;NOTE&lt;/strong&gt;: In the monorepo mode, CLI will use &lt;code&gt;webpack&lt;/code&gt; instead of &lt;code&gt;tsc&lt;/code&gt; by default. The reason for this change is that &lt;code&gt;webpack&lt;/code&gt; can produce a single file bundling all project components together (resolves cross-references out-of-the-box). &lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Running projects
&lt;/h3&gt;

&lt;p&gt;Likewise, in order to run particular application, you can call &lt;code&gt;$ nest start NAME&lt;/code&gt; command where &lt;code&gt;NAME&lt;/code&gt; is the name of the application / library you passed to the &lt;code&gt;$ nest g&lt;/code&gt; command. &lt;/p&gt;

&lt;blockquote&gt; 
  However, keep in mind that libraries cannot run on their own because they don't have &lt;code&gt;main.ts&lt;/code&gt; file. 
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Generate building blocks
&lt;/h3&gt;

&lt;p&gt;If you are familiar with Nest CLI already, you know that &lt;code&gt;$ nest g&lt;/code&gt; allows you to quickly scaffold essential building blocks of your application, such as controllers and providers.&lt;br&gt;
However, what happens if you switch from single project mode to a monorepo? Nest CLI is now setup to show an &lt;em&gt;interactive list&lt;/em&gt; of all applications / library in a workspace so you can select exactly where you want something generated!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Ftrilon.io%2F_nuxt%2Fimg%2F7245a8a.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Ftrilon.io%2F_nuxt%2Fimg%2F7245a8a.gif" alt="nest g service" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;code&gt;nest g service&lt;/code&gt; example


&lt;h2&gt;
  
  
  Backward compatibility
&lt;/h2&gt;

&lt;p&gt;Using &lt;code&gt;nest start&lt;/code&gt; and &lt;code&gt;nest build&lt;/code&gt; is &lt;strong&gt;not required&lt;/strong&gt;. All existing applications will work as expected - there are no breaking changes! New features were implemented to make developers life easier, but you can still use the same techniques as before to compile and serve your application (using either &lt;code&gt;typescript&lt;/code&gt; compiler or &lt;code&gt;webpack&lt;/code&gt; directly is totally fine).&lt;/p&gt;

&lt;h2&gt;
  
  
  In Conclusion
&lt;/h2&gt;

&lt;p&gt;With the latest CLI we've introduced many new great features for your applications.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;New &lt;code&gt;build&lt;/code&gt; &amp;amp; &lt;code&gt;start&lt;/code&gt; commands&lt;/li&gt;
&lt;li&gt;Monorepo support&lt;/li&gt;
&lt;li&gt;Generate upgrades

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;nest g app NAME&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;nest g lib NAME&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Interactive support letting you choose exactly &lt;em&gt;where&lt;/em&gt; to generate&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We hope you're as excited about these new features as we are, and we look forward to hearing your feedback and how we can improve the NestJS ecoystem even more!&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Become a Backer or Sponsor to Nest by&lt;/em&gt; &lt;a href="https://opencollective.com/nest" rel="noopener noreferrer"&gt;&lt;em&gt;donating to our open collective&lt;/em&gt;&lt;/a&gt;&lt;em&gt;. ❤&lt;/em&gt;&lt;/p&gt;

</description>
      <category>nestjs</category>
      <category>node</category>
      <category>typescript</category>
    </item>
  </channel>
</rss>
