A concurrency limit treats every request as roughly equal.
For LLM workloads, that assumption breaks down.
A request with a 500-token budget and one with a 30,000-token budget might both occupy a single concurrency slot, but they can represent very different amounts of provider capacity.
That’s the problem I built async-bulkhead-llm to address.
Instead of limiting only concurrent requests, it can bound two things at once:
concurrent requests
+
estimated in-flight tokens
A request is admitted only when both budgets have enough capacity.
Conceptually:
const bulkhead = new AsyncBulkhead({
maxConcurrent: 8,
maxInFlightTokens: 40_000,
});
That gives you a bulkhead around an LLM provider that accounts for workload size rather than simply request count.
This becomes especially useful when interactive and large batch requests share the same upstream model. A conventional concurrency limiter can still allow a few very large requests to consume most of the useful capacity.
Token-aware admission gives you another control surface.
async-bulkhead-llm is an open-source TypeScript/npm library I’ve been developing specifically around this problem.
The broader lesson has been simple: for LLM systems, concurrency is only one dimension of load.
If you're building production LLM infrastructure, I'd be interested in how you're handling admission control and overload protection.
Top comments (0)