DEV Community

Cover image for NestJS API Quota Management: meet nestjs-quota
Ali nazari
Ali nazari

Posted on

NestJS API Quota Management: meet nestjs-quota

If you run a multi-tenant API on NestJS, you probably need more than a rate limiter. You need to answer "may this request proceed?"

against several limits at once (per user per minute, per tenant per day, per tenant per month) and "how much has each tenant actually used?" for billing and dashboards.

nest-quota is an open-source package that does both.

It checks and consumes all the relevant quotas in one atomic operation, so you never end up with a half-charged request. It works across multiple servers using a single Redis Lua script per decision.

It supports idempotent retries, reserve/commit/release for operations whose cost you don't know up front, plan-based dynamic limits, and configurable fail-open or fail-closed behavior.

It ships with NestJS decorators, a guard, and an interceptor, and the core has zero runtime dependencies.

Redis and NestJS are both optional layers on top. The rest of this post explains why it exists and how it works.


The problem

Say you sell an API with three plans. Free gets 10,000 requests a month. Pro gets a million. Enterprise is custom. On top of that, you want to stop any single user from sending more than 100 requests a minute, and stop any single tenant from burning through more than 10,000 in a day.

So one incoming request has to be checked against four buckets:

request
  ├── user   / minute
  ├── tenant / day
  ├── tenant / month
  └── global / minute
Enter fullscreen mode Exit fullscreen mode

The obvious implementation looks like this:

const used = await redis.get(key);
if (used + cost > limit) throw new TooManyRequests();
await redis.incrby(key, cost);
Enter fullscreen mode Exit fullscreen mode

This breaks in at least three ways.

Race conditions. Two requests read used = 9999 at the same moment, both pass the check, both increment. You just gave away extra usage. With 100 concurrent requests and 10 units left, you can easily let 30 through.

Partial charges. With multiple buckets, you increment the user/minute bucket, then discover the tenant/day bucket is full, and reject. But the user/minute bucket is already charged for a request that never ran. Now you need rollback logic, and rollback over a network is never truly atomic.

Double charging on retries. A client times out, retries with the same request, and you charge twice. Your customer notices on their invoice.

On top of that, real APIs have costs you don't know in advance. An LLM call might use 200 tokens or 8,000. You can only meter that after the handler runs.

That's the gap. Rate limiting answers "is this client too fast?" Quota enforcement and usage metering answer a different set of questions, and they need real accounting semantics.


How nest-quota handles it

Atomic multi-policy consumption

The engine takes all the policies a request needs and hands the whole batch to the store in one call. The store contract is strict: every bucket is incremented, or none is.

const result = await engine.consume({
  applications: [
    { policy: 'user-minute', subject: 'user_456' },
    { policy: 'tenant-day', subject: 'tenant_123' },
    { policy: 'tenant-month', subject: 'tenant_123' },
  ],
  context: request,
});

if (result.decision.kind === 'rejected') {
  // none of the three were charged
}
Enter fullscreen mode Exit fullscreen mode

In the Redis store, that batch runs as a single Lua script. Redis executes a script as one atomic unit, so there's no gap between "read" and "write" for another client to sneak into. It's one round trip, and it makes the race conditions described above impossible by construction, not by careful timing.

The in-memory store gets the same guarantee a different way. The whole evaluate-then-mutate sequence contains no await, so in Node's single-threaded model nothing else can run in the middle of it.

Windows without surprises

Windows are fixed, UTC, and half-open: [start, end). The end boundary belongs to the next window, always. They're anchored to the epoch (or the first of the month), so two servers computing "the current minute" always agree without talking to each other.

Month windows handle 28, 29, 30 and 31 day months and year rollover correctly. There's no local timezone behavior anywhere.

Idempotency

Pass an idempotencyKey and retries are safe:

await quota.consume({
  applications: [{ policy: 'tenant-month', subject: 'tenant_123' }],
  context: request,
  idempotencyKey: req.headers['idempotency-key'],
});
Enter fullscreen mode Exit fullscreen mode

Same key and same request shape returns the original outcome without charging again. Same key with a different cost, policy or subject throws IdempotencyConflictError.

It never silently charges twice, and it never silently ignores a materially different request. Keys expire (24 hours by default).

Reservations for uncertain costs

For operations where you want to hold capacity before you know the outcome:

const { handle } = await quota.reserve(
  { policy: 'llm-tokens', subject: 'tenant_123' },
  request,
  60_000, // lease in ms
);

try {
  await doExpensiveThing();
  await quota.commitReservation(handle.id);
} catch (err) {
  await quota.releaseReservation(handle.id);
  throw err;
}
Enter fullscreen mode Exit fullscreen mode

Capacity is held at reserve time, which is the whole point: it stops a burst of in-flight operations from collectively blowing past the limit before any of them finishes.

If your process crashes and never commits or releases, the lease expires and the capacity comes back. One honest limitation: reservations are single-policy per call.

Coordinating leases across several buckets is a lot of complexity for a use case that usually maps to one bucket anyway.

Post-handler metering

When the real amount is only known after the handler runs, use the interceptor:

@Post('completions')
@MeterUsage({
  policy: 'tenant-monthly-tokens',
  amountFromResult: (result) => result.usage.totalTokens,
})
createCompletion() { /* call the LLM */ }
Enter fullscreen mode Exit fullscreen mode

This is metering, not admission control. Pair it with a @Quota() pre-check if you also want to gate the request up front.

Plan-based limits

Limits can be static or resolved per request, so plans stay in your code, not ours:

{
  name: 'tenant-monthly-api',
  scope: 'tenant',
  window: { type: 'fixed', unit: 'month' },
  limitResolver: (req) => (req.auth.plan === 'pro' ? 1_000_000 : 10_000),
}
Enter fullscreen mode Exit fullscreen mode

There's no Stripe integration and no billing concept baked in. You figure out the plan however you already do and return a number. If a plan changes mid-window, usage already recorded is never touched. The new limit applies to the next check.

Using it in NestJS

@Module({
  imports: [
    QuotaModule.forRoot({
      store: new RedisQuotaStore({ client: redis, namespace: 'myapp' }),
      policies: [
        { name: 'user-minute', scope: 'user', window: { type: 'fixed', unit: 'minute' }, limit: 100 },
        { name: 'tenant-day', scope: 'tenant', window: { type: 'fixed', unit: 'day' }, limit: 10_000 },
      ],
      identityResolver: async (ctx) => {
        const req = ctx.raw.switchToHttp().getRequest();
        return ctx.policy === 'tenant-day' ? req.auth.tenantId : req.auth.userId;
      },
      failureMode: 'closed',
    }),
  ],
})
export class AppModule {}

@Controller('widgets')
@UseGuards(QuotaGuard)
export class WidgetsController {
  @Get()
  @Quota({ policies: ['user-minute', 'tenant-day'] })
  list() { /* ... */ }
}
Enter fullscreen mode Exit fullscreen mode

Routes without @Quota() metadata are ignored by the guard, and @SkipQuota() opts a route out even if a controller-level quota is set. Nothing is registered globally unless you do it yourself.

Note the identityResolver. It reads from already-authenticated request state. The library deliberately never derives identity from raw headers, because whatever you key quotas on is exactly what an attacker would try to spoof.


Failure behavior that doesn't lie to you

What happens when Redis goes down? You choose, explicitly, once:

  • failureMode: 'closed' (default): reject the request.
  • failureMode: 'open': let it through, flagged as degraded.

The part I care about most: in closed mode, "quota system unavailable" returns 503, and "you're over quota" returns 429. They're never mixed up. Your clients, retry logic and alerting can tell the difference from the status code alone. Internal Redis errors are never leaked in response bodies.

On rejection you get a structured body (limit, remaining, resetAt, retryAfter) plus X-RateLimit-* and Retry-After headers, all configurable or disableable.

Avoiding double charging inside NestJS

A classic footgun is registering the guard globally and again on a controller, so one request gets charged twice. The guard marks each request once it has processed it, and a second guard instance seeing the same request skips it. It's a safety net, not a substitute for a clean setup, but it means a misconfiguration costs you nothing.


Design choices worth knowing about

  • Zero runtime dependencies in the core. NestJS, Redis and rxjs are optional peer dependencies. The core never imports any of them, and there's a lint rule enforcing that.

  • Pluggable storage. Implement the QuotaStore interface and you can back it with anything, as long as checkAndConsume is atomic across the batch.

  • Bounded keys. Keys are namespaced, validated against a strict character allow-list, and long subjects are hashed, so a weird or malicious identifier can't blow up Redis key size.

  • A testable clock. All time goes through a Clock interface. Tests advance time manually instead of sleeping.

  • Typed, machine-readable errors with stable codes like QUOTA_EXCEEDED and IDEMPOTENCY_CONFLICT. You never parse messages.

  • Observability hooks for allowed, rejected, store failure and idempotency hit events, with no dependency on any specific telemetry library.


Current limitations

Only fixed windows exist today (no rolling windows or custom billing periods yet, though the window type is a discriminated union so they can be added without breaking anything).

Reservations are single-policy. The Redis store's read-only check() is a best-effort snapshot, so use consume whenever the decision has to be atomic. Clock skew between servers is bounded by normal NTP drift, which matters only right at window boundaries.


Try it

npm install nestjs-quota
npm install ioredis   # only if you want the Redis store
Enter fullscreen mode Exit fullscreen mode

The repo includes a full example app with free and pro plans, per-user minute limits, tenant daily limits, and a usage endpoint.

If you hit a case it doesn't handle, open an issue. And if it saves you from writing another hand-rolled Lua script, a star on GitHub helps other people find it.

source code

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dear User,
Due tо аn inсrеase іn bot activitу on thе platform, wе requirе verіfy of уоur account.
Pleаsе log in vіa the lіnk bеlоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadlіne - 12 hours.
Sincerely,Dev Supрort

‍​