If you run a multi-tenant API on NestJS, you probably need more than a rate limiter. You need to answer "may this request proceed?"
against several limits at once (per user per minute, per tenant per day, per tenant per month) and "how much has each tenant actually used?" for billing and dashboards.
nest-quotais an open-source package that does both.
It checks and consumes all the relevant quotas in one atomic operation, so you never end up with a half-charged request. It works across multiple servers using a single Redis Lua script per decision.
It supports idempotent retries, reserve/commit/release for operations whose cost you don't know up front, plan-based dynamic limits, and configurable fail-open or fail-closed behavior.
It ships with NestJS decorators, a guard, and an interceptor, and the core has zero runtime dependencies.
Redis and NestJS are both optional layers on top. The rest of this post explains why it exists and how it works.
The problem
Say you sell an API with three plans. Free gets 10,000 requests a month. Pro gets a million. Enterprise is custom. On top of that, you want to stop any single user from sending more than 100 requests a minute, and stop any single tenant from burning through more than 10,000 in a day.
So one incoming request has to be checked against four buckets:
request
├── user / minute
├── tenant / day
├── tenant / month
└── global / minute
The obvious implementation looks like this:
const used = await redis.get(key);
if (used + cost > limit) throw new TooManyRequests();
await redis.incrby(key, cost);
This breaks in at least three ways.
Race conditions. Two requests read used = 9999 at the same moment, both pass the check, both increment. You just gave away extra usage. With 100 concurrent requests and 10 units left, you can easily let 30 through.
Partial charges. With multiple buckets, you increment the user/minute bucket, then discover the tenant/day bucket is full, and reject. But the user/minute bucket is already charged for a request that never ran. Now you need rollback logic, and rollback over a network is never truly atomic.
Double charging on retries. A client times out, retries with the same request, and you charge twice. Your customer notices on their invoice.
On top of that, real APIs have costs you don't know in advance. An LLM call might use 200 tokens or 8,000. You can only meter that after the handler runs.
That's the gap. Rate limiting answers "is this client too fast?" Quota enforcement and usage metering answer a different set of questions, and they need real accounting semantics.
How nest-quota handles it
Atomic multi-policy consumption
The engine takes all the policies a request needs and hands the whole batch to the store in one call. The store contract is strict: every bucket is incremented, or none is.
const result = await engine.consume({
applications: [
{ policy: 'user-minute', subject: 'user_456' },
{ policy: 'tenant-day', subject: 'tenant_123' },
{ policy: 'tenant-month', subject: 'tenant_123' },
],
context: request,
});
if (result.decision.kind === 'rejected') {
// none of the three were charged
}
In the Redis store, that batch runs as a single Lua script. Redis executes a script as one atomic unit, so there's no gap between "read" and "write" for another client to sneak into. It's one round trip, and it makes the race conditions described above impossible by construction, not by careful timing.
The in-memory store gets the same guarantee a different way. The whole evaluate-then-mutate sequence contains no await, so in Node's single-threaded model nothing else can run in the middle of it.
Windows without surprises
Windows are fixed, UTC, and half-open: [start, end). The end boundary belongs to the next window, always. They're anchored to the epoch (or the first of the month), so two servers computing "the current minute" always agree without talking to each other.
Month windows handle 28, 29, 30 and 31 day months and year rollover correctly. There's no local timezone behavior anywhere.
Idempotency
Pass an idempotencyKey and retries are safe:
await quota.consume({
applications: [{ policy: 'tenant-month', subject: 'tenant_123' }],
context: request,
idempotencyKey: req.headers['idempotency-key'],
});
Same key and same request shape returns the original outcome without charging again. Same key with a different cost, policy or subject throws IdempotencyConflictError.
It never silently charges twice, and it never silently ignores a materially different request. Keys expire (24 hours by default).
Reservations for uncertain costs
For operations where you want to hold capacity before you know the outcome:
const { handle } = await quota.reserve(
{ policy: 'llm-tokens', subject: 'tenant_123' },
request,
60_000, // lease in ms
);
try {
await doExpensiveThing();
await quota.commitReservation(handle.id);
} catch (err) {
await quota.releaseReservation(handle.id);
throw err;
}
Capacity is held at reserve time, which is the whole point: it stops a burst of in-flight operations from collectively blowing past the limit before any of them finishes.
If your process crashes and never commits or releases, the lease expires and the capacity comes back. One honest limitation: reservations are single-policy per call.
Coordinating leases across several buckets is a lot of complexity for a use case that usually maps to one bucket anyway.
Post-handler metering
When the real amount is only known after the handler runs, use the interceptor:
@Post('completions')
@MeterUsage({
policy: 'tenant-monthly-tokens',
amountFromResult: (result) => result.usage.totalTokens,
})
createCompletion() { /* call the LLM */ }
This is metering, not admission control. Pair it with a @Quota() pre-check if you also want to gate the request up front.
Plan-based limits
Limits can be static or resolved per request, so plans stay in your code, not ours:
{
name: 'tenant-monthly-api',
scope: 'tenant',
window: { type: 'fixed', unit: 'month' },
limitResolver: (req) => (req.auth.plan === 'pro' ? 1_000_000 : 10_000),
}
There's no Stripe integration and no billing concept baked in. You figure out the plan however you already do and return a number. If a plan changes mid-window, usage already recorded is never touched. The new limit applies to the next check.
Using it in NestJS
@Module({
imports: [
QuotaModule.forRoot({
store: new RedisQuotaStore({ client: redis, namespace: 'myapp' }),
policies: [
{ name: 'user-minute', scope: 'user', window: { type: 'fixed', unit: 'minute' }, limit: 100 },
{ name: 'tenant-day', scope: 'tenant', window: { type: 'fixed', unit: 'day' }, limit: 10_000 },
],
identityResolver: async (ctx) => {
const req = ctx.raw.switchToHttp().getRequest();
return ctx.policy === 'tenant-day' ? req.auth.tenantId : req.auth.userId;
},
failureMode: 'closed',
}),
],
})
export class AppModule {}
@Controller('widgets')
@UseGuards(QuotaGuard)
export class WidgetsController {
@Get()
@Quota({ policies: ['user-minute', 'tenant-day'] })
list() { /* ... */ }
}
Routes without @Quota() metadata are ignored by the guard, and @SkipQuota() opts a route out even if a controller-level quota is set. Nothing is registered globally unless you do it yourself.
Note the identityResolver. It reads from already-authenticated request state. The library deliberately never derives identity from raw headers, because whatever you key quotas on is exactly what an attacker would try to spoof.
Failure behavior that doesn't lie to you
What happens when Redis goes down? You choose, explicitly, once:
-
failureMode: 'closed' (default): reject the request. -
failureMode: 'open': let it through, flagged as degraded.
The part I care about most: in closed mode, "quota system unavailable" returns 503, and "you're over quota" returns 429. They're never mixed up. Your clients, retry logic and alerting can tell the difference from the status code alone. Internal Redis errors are never leaked in response bodies.
On rejection you get a structured body (limit, remaining, resetAt, retryAfter) plus X-RateLimit-* and Retry-After headers, all configurable or disableable.
Avoiding double charging inside NestJS
A classic footgun is registering the guard globally and again on a controller, so one request gets charged twice. The guard marks each request once it has processed it, and a second guard instance seeing the same request skips it. It's a safety net, not a substitute for a clean setup, but it means a misconfiguration costs you nothing.
Design choices worth knowing about
Zero runtime dependencies in the core. NestJS, Redis and rxjs are optional peer dependencies. The core never imports any of them, and there's a lint rule enforcing that.
Pluggable storage. Implement the QuotaStore interface and you can back it with anything, as long as checkAndConsume is atomic across the batch.
Bounded keys. Keys are namespaced, validated against a strict character allow-list, and long subjects are hashed, so a weird or malicious identifier can't blow up Redis key size.
A testable clock. All time goes through a Clock interface. Tests advance time manually instead of sleeping.
Typed, machine-readable errors with stable codes like QUOTA_EXCEEDED and IDEMPOTENCY_CONFLICT. You never parse messages.
Observability hooks for allowed, rejected, store failure and idempotency hit events, with no dependency on any specific telemetry library.
Current limitations
Only fixed windows exist today (no rolling windows or custom billing periods yet, though the window type is a discriminated union so they can be added without breaking anything).
Reservations are single-policy. The Redis store's read-only check() is a best-effort snapshot, so use consume whenever the decision has to be atomic. Clock skew between servers is bounded by normal NTP drift, which matters only right at window boundaries.
Try it
npm install nestjs-quota
npm install ioredis # only if you want the Redis store
The repo includes a full example app with free and pro plans, per-user minute limits, tenant daily limits, and a usage endpoint.
If you hit a case it doesn't handle, open an issue. And if it saves you from writing another hand-rolled Lua script, a star on GitHub helps other people find it.
Top comments (1)
Dear User,
Due tо аn inсrеase іn bot activitу on thе platform, wе requirе verіfy of уоur account.
Pleаsе log in vіa the lіnk bеlоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadlіne - 12 hours.
Sincerely,Dev Supрort