DEV Community

Cover image for NestJS Production Checklist: 17 Checks Before You Deploy
Kamil Mysliwiec for NestJS

Posted on

NestJS Production Checklist: 17 Checks Before You Deploy

Locally, a Nest app runs as one process, restarts when you save, talks to a database with no other clients, and gets requests straight from your browser. In production it runs as several replicas behind a load balancer, gets stopped on every deploy, shares its database with other processes, and depends on APIs that sometimes don't answer. A lot of defaults that are fine for the first setup are wrong for the second.

Startup and deploys

1. Validate config at boot, and keep secrets out of the image

A missing or malformed environment variable should stop the process before it accepts a request. Otherwise the app boots, passes its health check, takes traffic, and fails the first request that reads the value.

@nestjs/config accepts any Standard Schema (Zod, Valibot, ArkType) as validationSchema:

import { z } from 'zod';

@Module({
  imports: [
    ConfigModule.forRoot({
      isGlobal: true,
      validationSchema: z.object({
        NODE_ENV: z.enum(['development', 'test', 'production']),
        DATABASE_URL: z.url(),
        STRIPE_SECRET_KEY: z.string().min(1),
        PORT: z.coerce.number().default(3000),
      }),
    }),
  ],
})
export class AppModule {}
Enter fullscreen mode Exit fullscreen mode

Read required values with config.getOrThrow('STRIPE_SECRET_KEY') rather than get(). During a rolling deploy, a version that fails on boot never becomes ready and the previous one keeps serving. That's the failure mode you want.

While you're here, check where the values come from. Secrets belong in the platform's secret store (Kubernetes Secrets, AWS Secrets Manager, Vault, your PaaS's config), injected at runtime. Not in a .env file copied into the image, and not as a Docker ARG/ENV: both end up in the image layers, readable by anyone who can pull the image. A .dockerignore with .env* in it is a cheap guard.

And don't log the config object. console.log(config) at startup is a common debugging step that stays in, and it ships every credential to your log pipeline. If you need to see what loaded, log the keys, not the values.

2. Handle SIGTERM, and get the container right

Deploys, scale-downs and node drains all send SIGTERM. Nest doesn't listen for it unless you ask, so by default onModuleDestroy and onApplicationShutdown never run: pools stay open, queue workers are killed mid-job, in-flight requests are dropped.

const app = await NestFactory.create(AppModule);
app.enableShutdownHooks();
await app.listen(process.env.PORT ?? 3000);
Enter fullscreen mode Exit fullscreen mode

With hooks enabled, a signal runs: the HTTP adapter is marked as closing, onModuleDestroy, beforeApplicationShutdown, the HTTP server, gateways and microservices close, then onApplicationShutdown. Most integrations clean up in that last one. @nestjs/bullmq, for example, closes its workers there, and worker.close() waits for the jobs they're running.

Two ways the signal gets lost in containers:

  • Something between the runtime and Node. CMD npm run start:prod, or the shell form CMD node dist/main.js, puts npm or a shell in front of Node, and neither reliably forwards signals. Use the exec form: CMD ["node", "dist/main.js"].
  • Node running as PID 1. The kernel doesn't apply default signal behaviour to PID 1. enableShutdownHooks installs its own handler, so the hooks run, but Nest then exits by removing that handler and re-sending the signal to its own process. As PID 1, that second signal is ignored, and the process only exits if the event loop happens to be empty. Run with an init process (docker run --init, or tini as the entrypoint), or use app.enableShutdownHooks(undefined, { useProcessExit: true }) so Nest calls process.exit().

The rest of the image is worth a look while you're in the Dockerfile. Build in one stage and run in another, so compilers, dev dependencies and your source tree don't ship; set NODE_ENV=production; and don't run as root, since the official Node images already include a node user:

FROM node:24-slim AS build
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
RUN npm run build && npm prune --omit=dev

FROM node:24-slim
ENV NODE_ENV=production
WORKDIR /app
COPY --from=build /app/package.json ./
COPY --from=build /app/node_modules ./node_modules
COPY --from=build /app/dist ./dist
USER node
CMD ["node", "--enable-source-maps", "dist/main.js"]
Enter fullscreen mode Exit fullscreen mode

(package.json is copied on purpose: the Nest 12 starter is ESM, and Node reads "type": "module" from it.)

Then check memory. If the container has a memory limit and V8's heap limit is higher than what the container can actually give it, the kernel kills the process before V8 gets a chance to complain: exit code 137, no JavaScript error, no log line, just a restart. Depending on the Node version and how the limit is set, the default heap size doesn't necessarily follow the container's limit, so set it explicitly to leave room for everything that isn't heap (buffers, native memory, the stack):

env:
  - name: NODE_OPTIONS
    value: "--max-old-space-size=768" # ~75% of a 1 GiB limit
Enter fullscreen mode Exit fullscreen mode

Now a leak or an oversized payload ends in a JavaScript heap out of memory error with a stack trace in your logs, instead of a silent kill. Alert on restarts either way, and on OOMKilled specifically.

3. Drain before you close

When Kubernetes terminates a pod, two things start at the same time: the pod is removed from the Service's endpoints, and the kubelet stops the container (preStop hook, then SIGTERM). Endpoint removal takes a few seconds to reach every load balancer and kube-proxy. A server that closes immediately on SIGTERM refuses the requests routed to it during that window.

The usual fix is a short delay before shutdown starts:

lifecycle:
  preStop:
    exec:
      command: ["sleep", "5"]
Enter fullscreen mode Exit fullscreen mode

Nest 12 also has return503OnClosing: true on NestFactory.create(). It answers new requests with a 503 and Connection: close while in-flight ones finish. That's useful once traffic has drained, but during the propagation window those 503s reach clients, so use it together with the delay, not instead of it.

Then check the time budget. Kubernetes sends SIGKILL after terminationGracePeriodSeconds (default 30), and the preStop time counts against it. A queue job that runs longer than what's left gets killed, BullMQ marks it stalled, and another worker runs it again from the start. Either raise the grace period or keep jobs short, and make them idempotent regardless, because a crash does the same thing.

4. Scheduled jobs run on every replica

@Cron(), @Interval() and @Timeout() from @nestjs/schedule run in-process. Three replicas means the nightly invoice export runs three times. It's one of the most common production bugs in Nest apps, because it can't happen on a single dev instance.

@nestjs/locks handles this with decorators:

@Cron('0 2 * * *')
@OnOneInstance({ key: 'invoices:export' })
async exportInvoices() {
  // runs on exactly one instance; the others skip the tick
}

@Cron('*/5 * * * *')
@WithoutOverlapping()
async reconcile() {
  // skips a tick while the previous run is still going
}
Enter fullscreen mode Exit fullscreen mode

@OnOneInstance() holds a lease that the owning instance renews every ttl / 3 (30 s TTL by default). If that instance dies, another one takes over once the TTL runs out. @WithoutOverlapping() only holds the lock for the length of a run, which covers the other common bug: a 5-minute job that sometimes takes 7.

The lock store must be shared, and it's yours to provide: you register a LockStore backed by Redis or Postgres (the docs include both). The in-memory store only excludes other callers in the same process, so LocksModule refuses to use it in production unless you set allowInMemoryStorage. Missed ticks aren't replayed, so write jobs that process current state ("export everything not yet exported") rather than "yesterday's rows".

For code that has to be safe if a lock is lost mid-run (a long GC pause, a network partition), each acquisition comes with a signal that aborts on lease loss and a monotonically increasing fencingToken. Pass the signal to your I/O, and use the token to reject writes from a holder whose lock was taken over.

Queues have their own version of this. BullMQ doesn't retry unless you set attempts, a job whose worker dies or blocks the event loop past its lock gets marked stalled and run again, and completed and failed jobs stay in Redis until something removes them. The minimum is to set attempts, backoff and removeOnComplete/removeOnFail explicitly in defaultJobOptions, and make every job safe to run twice.

The edge

5. Make authentication deny-by-default

The usual pattern is @UseGuards(AuthGuard) on each controller. It fails open: the one controller where someone forgets the decorator is public, and nothing tells you. Invert it. Require authentication globally and opt routes out explicitly.

@nestjs/authentication works that way out of the box. AuthenticationModule registers a global guard, so every route requires a signed-in user and anything else gets a 401, and you mark the exceptions:

@Public()
@Controller('auth')
export class AuthController {
  // sign-in, sign-up, password reset
}
Enter fullscreen mode Exit fullscreen mode

It covers sessions, JWT bearer tokens, OpenID Connect, magic links and API keys, and @CurrentUser() gives handlers the resolved user.

If you keep your own guard, get the same property by registering it with APP_GUARD and letting a metadata decorator opt out:

export const IS_PUBLIC = 'isPublic';
export const Public = () => SetMetadata(IS_PUBLIC, true);

@Injectable()
export class AuthGuard implements CanActivate {
  constructor(private readonly reflector: Reflector) {}

  canActivate(context: ExecutionContext) {
    const isPublic = this.reflector.getAllAndOverride<boolean>(IS_PUBLIC, [
      context.getHandler(),
      context.getClass(),
    ]);
    if (isPublic) return true;
    // verify the token or session here
  }
}

// app.module.ts
providers: [{ provide: APP_GUARD, useClass: AuthGuard }],
Enter fullscreen mode Exit fullscreen mode

Either way, a review now has one thing to check: every @Public() in the diff.

6. Security headers, CORS, rate limits, and the proxy in front

Nest now sets the standard security headers itself, with the same defaults as Helmet 8 (CSP, HSTS, X-Content-Type-Options, X-Frame-Options and the rest):

const app = await NestFactory.create(AppModule);
app.useSecurityHeaders();
Enter fullscreen mode Exit fullscreen mode

It has to be called before app.init()/app.listen(), and only once; otherwise it throws. Individual headers are configurable, e.g. app.useSecurityHeaders({ contentSecurityPolicy: { directives: { ... } } }).

CORS: app.enableCors() with no arguments allows every origin. For an API used by one frontend with cookies, list the origins:

app.enableCors({ origin: ['https://app.example.com'], credentials: true });
Enter fullscreen mode Exit fullscreen mode

Rate-limit at least the expensive or brute-forceable endpoints (login, password reset, anything that sends email or SMS):

ThrottlerModule.forRoot([{ ttl: 60_000, limit: 100 }]),
// ...
providers: [{ provide: APP_GUARD, useClass: ThrottlerGuard }],
Enter fullscreen mode Exit fullscreen mode

Counters are per process unless you configure shared storage, so with N replicas the effective limit is N times higher.

Rate limiting also depends on knowing who the client is, and behind a load balancer every request comes from the load balancer's address. Until Express trusts X-Forwarded-*, req.ip is that address, all clients share one limit, and req.protocol is http even for HTTPS clients, which breaks secure cookies and audit logs too:

const app = await NestFactory.create<NestExpressApplication>(AppModule);
app.set('trust proxy', 1); // number of proxies in front of the app
Enter fullscreen mode Exit fullscreen mode

Set the actual hop count. true trusts whatever the client sends in X-Forwarded-For, which lets any client choose its own IP.

Input and output

7. Validate every input

Register a global validation pipe in main.ts. Which one depends on how you define DTOs.

With class-validator classes:

app.useGlobalPipes(
  new ValidationPipe({
    whitelist: true,
    forbidNonWhitelisted: true,
    transform: true,
  }),
);
Enter fullscreen mode Exit fullscreen mode

whitelist strips properties that aren't declared on the DTO. That's what stops { "email": "...", "role": "admin" } from reaching a repository.save(dto). forbidNonWhitelisted rejects them with a 400 instead, which also surfaces client typos.

With Zod, Valibot or ArkType, use StandardSchemaValidationPipe and attach the schema to the parameter:

app.useGlobalPipes(new StandardSchemaValidationPipe());
Enter fullscreen mode Exit fullscreen mode
export const createUserSchema = z.object({
  email: z.email(),
  password: z.string().min(8),
});
export type CreateUserDto = z.infer<typeof createUserSchema>;

@Post()
create(@Body({ schema: createUserSchema }) dto: CreateUserDto) {
  return this.users.create(dto);
}
Enter fullscreen mode Exit fullscreen mode

Here unknown keys are the schema's business, not the pipe's. z.object() strips them by default; use z.strictObject() if you want them rejected. Either way, the handler receives the schema's output, so coercions and defaults are applied.

One thing to know about useGlobalPipes(): it's applied to the app instance, so an e2e test that builds its own app from AppModule doesn't get it. Either repeat the call in the test setup, or register the pipe as an APP_PIPE provider so the module brings it along.

Validation is also where you bound list endpoints. A limit or take query parameter without a maximum lets one request load every row in the table into memory; so does a list endpoint with no default. Put the cap in the DTO or schema, and give every list query a default:

export const listOrdersQuery = z.object({
  limit: z.coerce.number().int().min(1).max(100).default(20),
  cursor: z.string().optional(),
});
Enter fullscreen mode Exit fullscreen mode

(@Max(100) plus a default value does the same with class-validator.) The same applies to anything that fans out per item: a bulk endpoint that accepts 10,000 ids is a list endpoint too.

The body size limit is already reasonable: Express's JSON parser stops at 100 KB. The mistake is raising it globally for one upload endpoint. app.useBodyParser('json', { limit: '50mb' }) applies to every route, and parsing a 50 MB JSON body blocks the event loop for every other request on that process. Stream large payloads on the route that needs them, or upload them directly to object storage.

8. Control what goes out

Returning ORM entities from controllers leaks every column, including the ones added after the endpoint was written. Map to response DTOs, or use @Exclude() with a global ClassSerializerInterceptor. The interceptor only applies to class instances, so plain objects (raw query results, { ...user, extra }) are serialized as-is.

Don't expose Swagger publicly in production. SwaggerModule.setup() publishes a complete map of the API. Skip it when NODE_ENV === 'production', or put it behind auth.

9. Make writes safe to repeat

Clients retry. A mobile app on a flaky connection, a user who double-clicks, a proxy or SDK that retries on a timeout: any of them can send the same POST /orders/:id/pay twice, and the second one charges the card again. Your handler can't tell a retry from a new request unless the client says so.

The standard contract is an Idempotency-Key header: the client generates a key per logical operation and sends it with every attempt, and the server runs the operation once and replays the stored response for repeats. @nestjs/idempotency implements it:

IdempotencyModule.forRoot({
  // keys are per user, so two clients can't collide or read each other's results
  scope: (req: { user?: User }) => req.user?.id,
}),
Enter fullscreen mode Exit fullscreen mode
@Post(':id/pay')
@Idempotent({ required: true })
pay(@Param('id') id: string, @Body() dto: PayOrderDto, @CurrentUser() user: User) {
  return this.orders.pay(id, user, dto);
}
Enter fullscreen mode Exit fullscreen mode

A repeat with the same key gets the stored response (with Idempotent-Replayed: true) and the handler doesn't run. A duplicate that arrives while the first is still running gets a 409 with Retry-After. The same key with a different body is a 422, which catches client bugs that reuse keys. 5xx responses aren't stored, so a request that failed on your side can be retried and will run again. required: true rejects requests without a key, which is what you want on payment-like endpoints.

The records need a shared store (Postgres or Redis; the docs have both), for the same reason as the lock store in item 4: in memory, a retry that lands on another replica, or arrives after a deploy, runs the handler again. The module refuses to start in production without one. Stored responses are kept for 24 hours by default (ttl), and the whole response body is stored, so keep that in mind for endpoints that return large payloads.

This is different from @Retry({ idempotent: true }) in item 12, which lets the server re-run a handler within one request. @Idempotent() deduplicates the client's retries across requests. If you use both, register IdempotencyModule before ResilienceModule in imports, because import order sets the order of their global interceptors.

Dependencies

10. Run current versions of Node, Nest and your dependencies

Old versions are a production risk even when nothing is visibly broken: security fixes, memory leak fixes and bug fixes only land in supported release lines, and the longer you wait, the bigger and riskier the eventual upgrade.

  • Node: run an LTS release that's still supported, Active or Maintenance. Odd-numbered majors never become LTS, and an end-of-life line gets no security patches at all: Node 20 reached end-of-life in April 2026, so anything still on it is unpatched. Pin the major in the base image (node:24-slim) and rebuild regularly so patch releases actually reach production, rather than building once and running that image for a year.
  • Nest: stay on the current major, and keep every @nestjs/* package on matching versions. Mismatched @nestjs/common and @nestjs/core versions cause confusing DI and metadata errors. Several items in this list (useSecurityHeaders(), StandardSchemaValidationPipe, return503OnClosing) only exist in recent releases.
  • Everything else: commit the lockfile, install with npm ci in CI and in the image, and let Renovate or Dependabot open small update PRs continuously. Twenty small upgrades with passing tests are much cheaper than one big one after two years. Run npm audit --omit=dev in CI to catch known vulnerabilities in what actually ships.

11. Time out every outbound call

A dependency that stops responding is worse than one that errors. Every request waiting on it holds memory, a socket and often a database connection, and under load you run out of those before the dependency recovers. Few clients time out by default:

  • fetch waits up to 5 minutes for response headers. Pass signal: AbortSignal.timeout(5_000).
  • For HTTP APIs, use @nestjs/http-client, the Promise-based replacement for the axios-based HttpModule. Configure a timeout per client:
  HttpClientModule.register({
    name: 'payments',
    baseUrl: 'https://api.payments.example.com',
    timeout: '5s',
  }),
Enter fullscreen mode Exit fullscreen mode

The timeout applies per attempt, and retries are on by default for idempotent methods (GET, HEAD, OPTIONS, PUT, DELETE: 3 attempts with exponential backoff, and on timeouts too). So a 5 s timeout can mean 15 s or more before the call fails. Make sure that fits the caller's budget, or tune retry.

  • Database queries: set a statement timeout. With pg, that's statement_timeout in the pool config, whatever ORM sits on top.

12. Plan for dependencies that are down

Timeouts cap one call. When a dependency is down for minutes, you also want to stop calling it for a while, limit how much of your capacity it can occupy, and serve something reasonable in the meantime. @nestjs/resilience provides that as decorators and presets:

ResilienceModule.forRoot({
  presets: {
    carrier: {
      timeout: '2s',
      retry: { attempts: 2 },
      circuitBreaker: { failureRateThreshold: 50, minimumCalls: 10, openDuration: '30s' },
    },
  },
}),
Enter fullscreen mode Exit fullscreen mode
@Get('quotes')
@Resilience('carrier')
@Fallback('flatRateQuotes', { handleIf: (error) => error instanceof CircuitOpenError })
getQuotes(@Query('orderId') orderId: string, @Signal() signal: AbortSignal) {
  return this.shipping.getQuotes(orderId, signal);
}
Enter fullscreen mode Exit fullscreen mode

When the breaker is open, calls fail immediately with a 503 and Retry-After instead of waiting 2 s each, and the fallback answers. @Bulkhead() caps concurrent calls, so a slow dependency can't tie up every request on the process. Pass the injected signal down to your I/O so a timed-out attempt actually stops. Inside services (jobs, cron, anything that isn't a handler), use ResilienceService.preset('carrier').execute(...). Two details: @Retry() only applies to safe methods unless the handler is also @Idempotent() (item 9) and the retry is marked idempotent: true, and retries multiply across layers. A resilience retry around an @nestjs/http-client call that retries itself makes 3 x 3 = 9 attempts, so decide which layer owns retries.

13. Run migrations as a deploy step

Every ORM has a way to sync the schema directly from your model: TypeORM's synchronize: true, drizzle-kit push, prisma db push. They're convenient in development and unsafe against production data, because they apply whatever diff they compute, and a renamed field can look like a drop plus an add.

In production, apply reviewed migrations (TypeORM migrations, drizzle-kit migrate, prisma migrate deploy) as a separate step before the new version starts: a Kubernetes Job, a release phase, a CI step. Don't run them from app startup (migrationsRun: true and equivalents) when you have several replicas, because they all start at once and race.

Also check that migrations are backward compatible with the version still running. During a rolling deploy, old and new code share the new schema. Dropping or renaming a column the old version still reads breaks it for the length of the rollout, so split those into expand and contract steps across two deploys.

The same rule is what makes rollback possible. The goal is that at any moment you can redeploy the previous image and it works against the current schema, because down-migrations under incident pressure are rarely tested and often lossy. If a deploy can't be rolled back that way, it should be a deliberate, reviewed exception, not something you discover during an incident.

14. Check the connection pool math

Every replica has its own pool. Five replicas with a pool of 20 is 100 connections. That's Postgres's default max_connections, and you haven't counted the migration job, the workers, an admin tool, or the deploy window where old and new pods run at the same time.

Set the pool size explicitly (for pg, the max option of the pool, however your ORM passes it through). Multiply by the maximum replica count during a deploy, and check it fits. When it no longer fits, add a pooler like PgBouncer rather than raising max_connections.

15. Watch for request-scoped chains

Scope.REQUEST propagates up the dependency graph. Anything that injects a request-scoped provider becomes request-scoped, and so does anything that injects that, up to the controllers. A request-scoped logger near the bottom of the graph can mean a dozen objects constructed per request across a module, without any code saying so.

Check for it before launch. If you need per-request data rather than per-request instances, AsyncLocalStorage (or nestjs-cls) provides it without rebuilding the graph. If you need per-tenant instances, durable providers reuse them per tenant instead of creating them per request.

Errors and visibility

16. Errors: statuses, logs, source maps

  • Keep the default exception handling semantics: real statuses for HttpException, an opaque 500 plus a log entry for everything else. If you customise the response shape, extend BaseExceptionFilter and delegate to it. A filter written from scratch drops Nest's ExceptionsHandler log line, and unhandled errors stop appearing in logs.
  • Errors outside handlers. An unhandled promise rejection crashes the process (Node's default since v15). That's the correct behaviour. Make sure the error is logged with its stack before exit, and that repeated restarts alert someone.
  • Source maps. The starter already emits them. Start Node with --enable-source-maps (or NODE_OPTIONS=--enable-source-maps) so stack traces point at .ts files and lines.

17. Logs, health checks, monitoring

Logs: JSON, production log levels, structured params, and a trace id on every line. How to monitor a NestJS app covers each.

Health checks: separate liveness from readiness. Liveness asks whether this process is stuck, and it shouldn't check dependencies: if the database goes down and liveness fails, Kubernetes restarts every pod, which doesn't help and adds a reconnect storm when the database comes back. Readiness asks whether the pod should receive traffic, and that's where the @nestjs/terminus database check goes.

Monitoring: the items above prevent known failure modes. For the rest, you need per-route latency and error rates, traces that show which of your methods and queries took the time, errors grouped into defects with an alert on new ones, and visibility into jobs and scheduled runs.

For Nest specifically, that's what we built NestJS Observe for. It instruments through Nest's DI container, so spans are your own providers, and setup is a module and one option:

// app.module.ts
export const { ObserveModule, ObserveInstrument } = createObserveModule();

@Module({
  imports: [
    ObserveModule.forRoot({
      appKey: process.env.OBSERVE_APP_KEY!,
      appSecret: process.env.OBSERVE_APP_SECRET!,
      serviceId: 'orders-api',
      // nothing is ignored by default; keep probes out of the latency numbers
      http: { ignore: [/^\/health(?:\/|\?|$)/] },
    }),
  ],
})
export class AppModule {}

// main.ts
const app = await NestFactory.create(AppModule, { instrument: ObserveInstrument });
Enter fullscreen mode Exit fullscreen mode

That gives you requests, jobs, cron runs and messages with p95 and error rate, traces with the SQL under each method, unhandled errors grouped with the failing source line and an email on new ones, and trace ids on ConsoleLogger output. The free tier covers 300,000 events a month.

Summary

  1. Config is validated at boot, required values use getOrThrow, secrets come from a secret store, and the config object is never logged.
  2. enableShutdownHooks() is on, Node receives SIGTERM directly, and there's an init process; the image is multi-stage, runs as non-root, and the heap limit fits the container's memory limit.
  3. Pods drain before closing, and the grace period covers the longest job.
  4. Scheduled jobs are guarded with @OnOneInstance()/@WithoutOverlapping() on a shared lock store, and every job is idempotent.
  5. Authentication is global, and every @Public() is deliberate.
  6. useSecurityHeaders() is on, CORS lists its origins, sensitive endpoints are rate-limited, and trust proxy matches the number of proxies.
  7. A global ValidationPipe (with whitelist) or StandardSchemaValidationPipe is registered, list endpoints have a maximum page size, and the body limit is unchanged.
  8. Responses are DTOs or serialized classes; Swagger isn't public.
  9. Endpoints that must not run twice require an Idempotency-Key, backed by a shared store.
  10. Node is on a supported LTS, @nestjs/* packages are current and aligned, and dependency updates arrive continuously.
  11. Every outbound call and query has a timeout, and retry budgets add up.
  12. Unreliable dependencies have a circuit breaker, a bulkhead or a fallback, and retries are owned by one layer.
  13. Migrations run as a separate deploy step and are backward compatible, so the previous image can always be redeployed.
  14. Pool size times peak replica count fits in the database.
  15. No unintended request-scoped chains.
  16. Error handling keeps statuses and logs; stack traces resolve to source.
  17. Logs are structured, liveness doesn't depend on the database, and the service is monitored.

If you'd rather see the monitoring side of this (item 17) in action than read about it, here's a walkthrough of setting up NestJS Observe: NestJS Observe: Zero-Config Observability for NestJS on YouTube.

Top comments (3)

Collapse
 
elijahbrown profile image
Elijah Brown •

Item 7 is the right default. For signup DTOs I'd go one step past z.email(): look up whether the domain publishes MX, because a well-formed address on a dead or Null-MX domain still passes the schema and becomes an account nobody can reach.

Collapse
 
haithamoumer profile image
OUMERZOUG Haïtham •

Intresting topic, thank you @kamilmysliwiec

Collapse
 
mohammadi profile image
Sina Mohammadi •

You've really elevated the NestJS experience. With all the new features and repositories like Observe, it's genuinely impressive. Thanks!