DEV Community

laravel-o11y
laravel-o11y

Posted on Originally published at boring-observability.dev

Rate-Limited APIs and Laravel Queues: One Request at a Time, Without Starving the Rest

You are integrating with an API that allows sixty calls a minute and no parallel requests at all. Send two at once and it returns a 429, or worse, corrupts the record you were updating. Your queue workers, meanwhile, are built to run as many jobs at once as they can.

Each of Laravel's answers to this problem addresses one of two different limits, and most teams reach for the wrong one. This article goes through the options the framework and Horizon give you (the middleware, the limiters, the process counts) and what each one costs. It ends with a single serial slot, shared by a fast lane and a slow lane, where urgent work can still go first.

Key takeaways

  • "Sixty a minute" and "one at a time" are different limits. A rate limit bounds how often a job may start; a concurrency limit bounds how many may be in flight. RateLimited handles the first and does nothing about the second: sixty workers can fire simultaneously and stay within sixty a minute.
  • Laravel has no "don't pull" primitive. Every built-in throttle pops the job, finds it can't run it, and pushes it back. BullMQ leaves rate-limited jobs waiting in the queue; Laravel turns them into churn.
  • That churn has three costs: a released job goes to the back of the queue, it consumes an attempt, and it lands in the delayed set, where Horizon mostly stops showing it to you.
  • The cheapest answer is not in the docs: a dedicated supervisor with maxProcesses => 1. Concurrency of one comes from the process count, with no locks, no releases and no burned attempts.
  • Priority on a single worker needs weights. Put api-high and api-default on that one worker and they share the limiter with no extra code. With strict priority order, though, a flood on api-high starves api-default to zero. Weighting the queues gives api-default a fixed share instead.

Two limits that look like one

The vendor's docs state two limits. "60 requests per minute" and "no concurrent requests" are different constraints, and no single setting satisfies both:

  • A rate limit bounds how often a call may start: 60 per minute, 5 per second. It says nothing about how many are in flight at once.
  • A concurrency limit bounds how many calls may run simultaneously, in our case exactly one. It says nothing about how often they may start.

They are independent. Sixty calls per minute, each taking ten seconds, is ten requests in flight at any moment: within the rate limit and well over the concurrency limit. One call at a time, each taking ten milliseconds, is a hundred requests a second: serial, and far over the rate limit. You need both limits, and Laravel provides them in different places.

The distinction is visible in Laravel's API. Redis::throttle() is a rate limiter and Redis::funnel() is a concurrency limiter; the RateLimited middleware is the former and WithoutOverlapping is the latter. Most "how do I rate-limit a Laravel job" answers online cover the rate half and leave the concurrency half unsolved. The options below cover both, with what each one costs.

Option 1: the RateLimited middleware

The documented answer, and the right tool for the "60 per minute" half. Define a named limiter, then attach the middleware to the job:

use Illuminate\Cache\RateLimiting\Limit;
use Illuminate\Support\Facades\RateLimiter;

// In AppServiceProvider::boot()
RateLimiter::for('vendor-api', function (object $job) {
    return Limit::perMinute(60);
});
Enter fullscreen mode Exit fullscreen mode
use Illuminate\Queue\Middleware\RateLimitedWithRedis;

public function middleware(): array
{
    return [new RateLimitedWithRedis('vendor-api')];
}
Enter fullscreen mode Exit fullscreen mode

When the limit is hit, the middleware calls $job->release() with a delay of "time until the window resets, plus three seconds." You can override that with releaseAfter(), or call dontRelease(), which is easy to misread. dontRelease() does not fail the job and does not defer it. It deletes it: no exception, no failed_jobs row, no failed() hook.

Prefer RateLimitedWithRedis over plain RateLimited whenever the limit is an external contract. The generic version checks the counter and increments it as two separate cache operations. That is a time-of-check-to-time-of-use race: several workers can pass the check before any of them increments, and together they exceed your vendor's limit. The Redis variant does both in a single atomic Lua script. The docs describe it only as "more efficient."

Pros: documented, declarative, per-job, and the limiter key is shared across every job that names it, so two different job classes can draw on one budget. Cons: it limits rate and nothing else. Sixty workers can still hit the API in the same instant and remain within sixty per minute. And every throttled job pays the release costs described below.

Option 2: WithoutOverlapping, a mutex

This covers the concurrency half, and it gives you the number you want: one.

use Illuminate\Queue\Middleware\WithoutOverlapping;

public function middleware(): array
{
    return [
        (new WithoutOverlapping('vendor-api'))
            ->shared()
            ->releaseAfter(5)
            ->expireAfter(120),
    ];
}
Enter fullscreen mode Exit fullscreen mode

All three settings in that snippet matter, and all three defaults are wrong for this use case:

  • shared() is required here. By default the lock key includes a hash of the job's class name, so SyncContact and PushInvoice get separate mutexes and call your no-parallel-requests API in parallel. shared() drops the class from the key, which is what makes both jobs contend for the same lock.
  • releaseAfter() defaults to 0. A blocked job is re-queued with no delay: pop, fail the lock, re-enqueue, pop, fail the lock. That loop keeps a worker core busy and burns an attempt on every pass.
  • expireAfter() defaults to never. The lock is released in a finally block, so an exception is fine, but a worker killed by --timeout or SIGKILL never runs it. The lock is then orphaned permanently, and that key stays locked until someone force-releases it by hand. Always set it above your job timeout.

Pros: mutual exclusion, works on any lock-capable cache driver, and it composes with RateLimited: stack both and you have covered both limits. Cons: the concurrency is fixed at one; there is no ->limit(n). It is release-based, so it pays the full churn cost. And a lock is a coordination mechanism you have to operate, with expiries to tune and orphans to clean up.

Option 3: funnels and throttles inside the job

Below the middleware sit the primitives they are built on. Redis::funnel() is the one that matters here, because it is the only thing in Laravel that expresses "at most N of these at once" for N greater than one:

use Illuminate\Support\Facades\Redis;

public function handle(): void
{
    Redis::funnel('vendor-api')
        ->limit(1)
        ->releaseAfter(120)
        ->block(0)
        ->then(
            fn () => $this->callTheVendor(),
            fn () => $this->release(5),
        );
}
Enter fullscreen mode Exit fullscreen mode

funnel caps how many run at once; its sibling throttle()->allow(60)->every(60) caps how often they start. Both are Lua-backed and atomic, and both are undocumented. They were in the queue docs up to Laravel 8 and were dropped when RateLimited landed. The classes still ship and still work, but tutorials that teach them cite documentation that no longer exists.

There are two traps. First, block() defaults to three seconds, and that is three seconds of the worker sleeping inside your job's runtime. It pushes the job toward retry_after, at which point a second worker picks the job up while the first is still holding it, and you make the duplicate API call you were trying to prevent. Pass block(0) to fail fast. Second, omitting the failure closure throws a LimiterTimeoutException instead of releasing, which counts as a job failure.

The releaseAfter(120) above is the slot's safety TTL. If a worker is hard-killed mid-job, the slot is not returned until it expires, so it must exceed your job timeout. Otherwise a second worker is handed the slot while the first is still running, and the limit(1) guarantee no longer holds. Set it too long, though, and one stuck job stalls your serial pipeline until the TTL runs out.

Laravel 12 added the same concept to the cache layer as Cache::funnel(), which works on any lock-capable driver and is documented. The semantics are the same, without the Redis dependency.

Pros: atomic, and the only way to express a concurrency limit above one. Cons: the Redis variants are undocumented, they block by default, there is no middleware wrapper, and the code lives inside handle(), which mixes coordination logic into your business logic.

Option 4: one queue, one worker

Every option so far gets concurrency of one by having many workers compete for a lock, with the releasing, re-queueing and attempt-burning that implies. The queue already has a concurrency control, and it is the most reliable one available: the number of worker processes.

Give the API its own queue and give that queue exactly one worker, and concurrency of one follows from the setup instead of being enforced. There is nothing to contend for. No lock, no release, no attempt burned, and FIFO order is preserved because no job ever loses its place.

'supervisor-vendor-api' => [
    'connection'   => 'redis',
    'queue'        => ['api-high', 'api-default'],
    'balance'      => false,
    'minProcesses' => 1,
    'maxProcesses' => 1,
    'rest'         => 1,
    'timeout'      => 30,
],
Enter fullscreen mode Exit fullscreen mode

maxProcesses => 1 is the setting that matters, and it is easy to get wrong. It is tempting to assume balance => false already pins the process count, but it does not: even unbalanced, Horizon scales the pool toward the number of ready jobs, clamped between minProcesses and maxProcesses. Pinning both ends to 1 is what guarantees a single serial slot.

rest, one of Horizon's more obscure options, sleeps the worker for N seconds after each job. On a single-process supervisor that works as a rate limit: one job per (duration + 1s) stays comfortably under sixty a minute. There is no middleware, no limiter key, no released job, no consumed attempt, and nothing is pushed to the back of the queue. The option appears in no version of the Laravel documentation.

Pros: the cheapest correct answer. No release churn, no lock administration, FIFO preserved, and it covers both limits at once. Cons: two. It is coarse: rest paces every job on that worker identically and cannot burst up to a limit the way a token bucket can. And maxProcesses is a per-host cap, not a cluster-wide one.

The per-host cap breaks the guarantee without any error. Every machine running php artisan horizon builds its own provisioning plan and deploys every supervisor matching its environment. Nothing in Horizon coordinates process counts across hosts: there is no cluster-wide cap, no lock and no leader election. Three app servers running Horizon against the same Redis give you three of these "single" workers, and three parallel calls to an API that permits none.

The fix is to make exactly one host provision this supervisor by giving it an environment of its own. Horizon resolves its environment from --environment, then config('horizon.env'), then APP_ENV, so the simplest lever is the flag, set in that host's Supervisor or systemd unit:

# On the one host that owns the vendor API. Everywhere else: php artisan horizon
php artisan horizon --environment=serial
Enter fullscreen mode Exit fullscreen mode
'environments' => [
    'production' => [
        // ...your normal supervisors, on every app server
    ],

    'serial' => [
        'supervisor-vendor-api' => [
            'connection'   => 'redis',
            'queue'        => ['api-high', 'api-default'],
            'balance'      => false,
            'minProcesses' => 1,
            'maxProcesses' => 1,
            'rest'         => 1,
        ],
    ],
],
Enter fullscreen mode Exit fullscreen mode

Environment matching is a first-match wildcard lookup, and a host whose environment matches nothing starts zero supervisors, silently. If you cannot guarantee single-host ownership, or would rather not stake your vendor contract on a deployment detail, keep a funnel(1) or a WithoutOverlapping as a backstop. It should almost never fire.

Option 5: chains and drip-feeds

Job chains (Bus::chain()) serialize by construction: job N+1 is not enqueued until job N succeeds. For a known, finite, ordered batch, such as "sync these 200 records, in order", this works well and needs no locks. But a chain is a closed list. Producers elsewhere in your app cannot append to it, two chains run in parallel with no mutual exclusion between them, and one failure abandons the rest of the chain. It serializes one batch, not access to the API.

Drip-feeding means a scheduled command that dispatches N jobs a minute, or a self-redispatching "pacemaker" job that pulls work off a list on an interval. Teams build it when the built-ins fall short. It works, and it has the same advantage as the single worker: nothing is popped that cannot be run. But you are now hand-rolling a scheduler on top of a queue, the scheduler's one-minute floor sets your resolution, and you own every edge case yourself. Use it only if the single-worker supervisor cannot express your pacing.

Option 6: the packages

The ecosystem has filled part of the gap, and the two mature packages solve different problems:

  • spatie/laravel-rate-limited-job-middleware wraps Redis::throttle in a fluent middleware and adds what core lacks: exponential backoff on release, time-based attempts, and events when the limit is hit. It is maintained and widely used. It is still a rate limiter, and still release-based.
  • mxl/laravel-queue-rate-limit throttles at the worker level, which core cannot do. Configure 'allows' => 1, 'every' => 5 for a queue and it makes the worker sleep instead of releasing the job. No attempts consumed, no re-queueing, no reordering. The cost is coarseness, since it throttles the whole queue regardless of job class. On a dedicated API queue that is what you want anyway.

Concurrency limiting above one is still missing. Core gives you exactly one (WithoutOverlapping) or an undocumented funnel, and no maintained package covers the range in between well.

The options, side by side

Approach Limits rate? Limits concurrency? Release churn? Main cost
RateLimited Yes No Yes Non-atomic; can overshoot a hard limit
RateLimitedWithRedis Yes No Yes Solves only half the problem
WithoutOverlapping No Yes, exactly 1 Yes Lock orphans; three wrong defaults
Redis::funnel() No Yes, N Yes Undocumented; blocks the worker by default
Redis::throttle() Yes No Yes Undocumented; blocks the worker by default
Chains No Within a chain No Closed set; can't serialize the API globally
One worker + rest Yes Yes, exactly 1 No Coarse pacing; per-host, not cluster-wide

The cost of nearly every option: Laravel pops before it checks

Every lock- and middleware-based option in the table has release churn. It is why teams who add the RateLimited middleware often end up with a queue that is slower, drops more jobs and is harder to see than before.

When a job is rate-limited, Laravel has already popped it: reserved it, unserialized it and re-fetched its Eloquent models from the database. Only then does the middleware run, find the budget spent, and push the job back. In BullMQ, rate-limited jobs stay in the waiting state: the worker declines to pull, and nothing else happens. Laravel has no "don't pull" primitive, so it pulls the job and puts it back. Rate limiting in Laravel therefore does not reduce work. It converts work into churn, and the churn grows with how far over the limit you are.

That design has three consequences, and the documentation mentions only one of them:

  • Released jobs go to the back of the queue. On Redis, release() puts the job in the delayed set; when it matures it is RPUSHed onto the tail. On the database driver it is a new row with a new auto-increment id, sorted last. The job that was next in line becomes last in line every time it is throttled. Under sustained load, a job can be starved indefinitely while newer work overtakes it. For the jobs you are throttling, FIFO order is inverted.
  • Released jobs consume attempts. This one is documented, and it is still the most common way this setup fails: your retry budget, which exists to protect you from failures, is spent on successful backpressure. With the default tries of 1, a job fails permanently the first time it is throttled. The fix is to count time instead of attempts, with retryUntil().
  • Released jobs collide again as a group. Every job blocked on the same limiter gets the same release delay; middleware releases have no jitter and no exponential backoff. They all mature at the same instant, are migrated back in one batch, and are popped and throttled again together on the next tick. The group stays synchronized indefinitely.
use DateTime;

// Retry until a deadline instead of for a fixed number of attempts.
public function retryUntil(): DateTime
{
    return now()->plus(hours: 6);
}
Enter fullscreen mode Exit fullscreen mode

Teams have followed this path to its end (custom throttling middleware, then dynamic per-user queues, then a forked Horizon autoscaler) and written post-mortems about it. That is a lot of engineering for calling an API slowly, and it is the strongest argument for an approach that never releases a job in the first place.

The shape that works: one slot, one limiter

The API needs serial access and a paced rate, you want urgent work to overtake routine work, and you want as little churn as possible. That gives you one supervisor, two queues and one worker:

'supervisor-vendor-api' => [
    'connection'   => 'redis',
    'queue'        => ['api-high', 'api-default'],
    'balance'      => false,
    'minProcesses' => 1,
    'maxProcesses' => 1,   // the serial slot
    'rest'         => 1,   // the rate limit
    'timeout'      => 30,
],
Enter fullscreen mode Exit fullscreen mode

Because both queues are drained by the same single process, they cannot produce a parallel request. No lock forbids it; there is no second worker to make one. The two queues share the concurrency budget by sharing the worker. And if you add a RateLimited('vendor-api') middleware for a limit rest can't express, both queues share that budget too, because the limiter is keyed by its name, not by the queue.

Keep the rest of your queues, and the other twenty supervisors doing your main work, out of this. This supervisor is a deliberate bottleneck.

Letting urgent jobs jump the line

A customer clicks "sync now" and expects their record pushed within seconds, but there are 4,000 nightly-reconciliation jobs on the same queue and the vendor allows one call a second. That is an hour of waiting behind work nobody is watching.

With balance => false, the queue list is a strict left-to-right priority order: a worker only looks at api-default when api-high is empty. Dispatching to api-high therefore does what you want. The urgent job is picked up next, and it uses the same single slot and the same limiter, so jumping the line cannot breach the API contract.

SyncContact::dispatch($contact)->onQueue('api-high');
Enter fullscreen mode Exit fullscreen mode

Strict priority is absolute. Push a few thousand jobs onto api-high and api-default gets nothing for as long as api-high stays non-empty. With one worker and a rate limit throttling you on purpose, that can be hours. This is the same problem covered in Horizon's balancing trade-offs, and a rate-limited single worker is its worst case, because your throughput is capped by contract.

Turning balancing on does not help. With balance => 'auto', Horizon ignores queue order and gives each queue its own process pool, which means two workers and two parallel API calls, the thing the vendor forbids. Auto-balancing and serial access are mutually exclusive.

Skyline's weighted queues keep the single shared worker of balance => false, so the serial slot survives, and have it check the queues proportionally instead of in strict order:

'supervisor-vendor-api' => [
    'connection'   => 'redis',
    'queue'        => ['api-high', 'api-default'],
    'balance'      => false,
    'minProcesses' => 1,
    'maxProcesses' => 1,
    'rest'         => 1,
    'queueWeights' => [
        'api-high' => 5,
        // 'api-default' omitted => default weight of 1
    ],
],
Enter fullscreen mode Exit fullscreen mode

At five to one, roughly five in every six pickups favour api-high, and the sixth still reaches api-default. Urgent work gets most of a scarce, contractually capped resource, and the nightly reconciliation keeps draining instead of stalling until morning. There is still one worker and one API call at a time.

For one-off urgency (a single job, now, without a second queue), front-of-queue dispatching pushes a single job onto the head of the Redis list rather than the tail:

use Laravel\Horizon\InteractsWithFrontOfQueue;

class SyncContact implements ShouldQueue
{
    use Dispatchable, InteractsWithFrontOfQueue, Queueable;
}

SyncContact::dispatch($contact)->onFront();
Enter fullscreen mode Exit fullscreen mode

onFront() acts on the ready list. A job sitting in the delayed set because a rate limiter released it is not on that list yet, so it cannot be moved to the front until it matures. This is the ordering problem from the release section again: once a job has been released, its position is decided by when it matures, not by you.

What you can no longer see

This architecture has one more cost: it makes your queue harder to see.

Every job that a rate limiter releases lands in the delayed set. It is not running, not failed and not in the ready list; it is in a fourth state that Horizon barely shows. A backlog of throttled work can be almost absent from the dashboard you are checking to find out why nothing is happening. Jobs delayed beyond your retention window can drop off the dashboard altogether while still pending, a long-standing Horizon issue that was closed without a fix.

The single-slot pattern adds a second visibility cost. With balance => false, Horizon registers the pool under the comma-joined name of its queues, so your two lanes stop being two rows and become one pseudo-queue called redis:api-high,api-default. You cannot set a per-queue wait threshold for them, and the metrics dashboard shows them as a single entry. The pattern that gave you a serial slot costs you per-queue visibility, on the queue you most need to watch because you throttle it on purpose.

In Skyline, delayed and retrying jobs are first-class: you can see what is waiting and when each job is due, and run any of them immediately instead of waiting out a backoff window. You can open api-default and read the jobs inside it, in order. And when the vendor goes down, you can pause that one queue and let the work accumulate, instead of a rate limiter running 4,000 jobs through their retry budgets against an API that is returning 503s. None of that needs a migration or a code change: see Skyline vs Horizon for what swapping the package does and does not touch.

One change is worth making by hand. Import RateLimitedWithRedis from Laravel\Horizon\Middleware instead of Illuminate\Queue\Middleware, and Skyline's Locks & Limits screen counts how often each limiter released or dropped a job, while a job dropped by dontRelease() is logged instead of vanishing.

If you are deciding which queue a job belongs on in the first place, that is practice #2 in 12 best practices for Laravel background jobs. And if the jobs piling up behind your rate limiter are duplicates of each other rather than distinct work, read the ShouldBeUnique gotchas next.

Common questions

How do I limit a Laravel job to one at a time?

Two ways. The lock-based way is the WithoutOverlapping middleware with ->shared(). Without shared(), the lock key includes the job class, so two different job classes still run in parallel. The cheaper way is structural: give the job its own queue and a Horizon supervisor with maxProcesses => 1, so a single worker process drains it serially. That needs no locks, releases no jobs and consumes no attempts.

Does the RateLimited middleware prevent parallel requests?

No. A rate limit bounds how often a job may start; it says nothing about how many run at once. Sixty workers can call an API simultaneously and still be within a limit of sixty per minute. To bound concurrency you need WithoutOverlapping, Redis::funnel(), or a supervisor pinned to one worker process.

What is the difference between Redis::throttle() and Redis::funnel()?

throttle() is a rate limiter: at most N starts per time window. funnel() is a concurrency limiter: at most N running at once, with each slot held for as long as the job runs and returned when it finishes. Both are atomic Lua scripts. Both were dropped from the Laravel documentation after version 8, and both still ship and work.

Why do my rate-limited jobs fail after a few minutes?

Because releasing a rate-limited job back onto the queue still increments its attempt count. With the default tries of 1, a job fails permanently the first time it is throttled, having used its only attempt on waiting for the limiter. Define retryUntil() on the job so its retry budget is a deadline instead of an attempt count.

How do I let an urgent job jump ahead of a rate-limited queue?

Run an api-high and an api-default queue on the same single-worker supervisor: they share the worker, so they share the concurrency budget, and they share the rate limiter because it is keyed by name rather than by queue. With balance => false the queue list is a strict priority order, so a flood on api-high starves api-default entirely. Skyline's queueWeights check the two in proportion instead (5:1 here), so urgent jobs mostly go first and api-default is still checked on every poll.

Top comments (0)