Horizon will tell you about one thing: a queue whose estimated wait has crossed a threshold. It will not tell you that the wait ended, that jobs are failing, that a supervisor was killed for memory, or that Horizon itself stopped an hour ago. Most teams find this out the usual way, from a customer asking where their invoice is.
This post walks up the layers you can add. It starts with what ships in Horizon, moves to the gauges spatie/laravel-prometheus exports and the alert rules those support, then to a counter-based exporter and the rules that only counters make possible. The last part covers the conditions no scrape carries, which is where Skyline's alerts come in. The rules are the ones we run in production, and for each we say which exporter can and cannot evaluate it.
Key takeaways
-
Horizon has one notification,
LongWaitDetected. It is built on a forecast of time to clear, repeats every five minutes while the condition lasts, never reports recovery, and is evaluated by a supervisor, so it is silent when Horizon is down. - Spatie's Horizon collectors are gauges read at scrape time. They are enough for "Horizon is not running", "Horizon is paused" and "this backlog has not moved". They cannot express a failure rate, and they have nothing per job class.
- Counters unlock the rules worth paging on. With processed, failed and retried counters plus the age of the oldest pending job, you can alert on failure ratio, on a backlog nobody is draining, and on latency instead of length.
-
In Grafana, keep the comparison out of the PromQL. A query like
x > 900returns no series when it is false, and Grafana reads that as No Data. Put the threshold in a threshold expression instead. -
Some failures never become a time series. A stranded unique lock, a job dropped by a rate limiter, a reservation that expired mid-run and a
timeoutaboveretry_afterproduce no failures and no backlog. Those need checks that read queue state directly.
What Horizon sends on its own
Horizon's notification is configured in two places. The thresholds live in config/horizon.php, keyed by connection and queue, and the destinations are set in your HorizonServiceProvider:
// config/horizon.php
'waits' => [
'redis:default' => 60,
'redis:reports' => 1800,
'redis:housekeeping' => 0, // 0 turns the check off for that queue
],
// app/Providers/HorizonServiceProvider.php
Horizon::routeSlackNotificationsTo('https://hooks.slack.com/services/...', '#ops');
Horizon::routeMailNotificationsTo('ops@example.com');
Horizon::routeSmsNotificationsTo('+15555550100');
A queue you do not list gets 60 seconds. Once a minute, one supervisor in the fleet takes a lock, calculates the wait for every queue it knows about, and raises LongWaitDetected for each one over its threshold. That event sends the notification. It is a reasonable first alert, and it is worth knowing exactly how it behaves before relying on it.
-
The wait is a forecast. It comes from
WaitTimeCalculator: ready jobs multiplied by the queue's average runtime, divided by the processes serving it. Ten thousand fast jobs produce a large number and a single job stuck at the head of a queue produces a small one. - It repeats every five minutes. The listener takes a 300-second lock per connection and queue before sending, so a backlog that lasts two hours sends twenty-four messages.
- It never says the wait ended. There is no recovery message, so a channel that went quiet might mean the queue drained or might mean someone muted it.
-
Queues balanced together are one entry. A supervisor working
high,defaultis keyedredis:high,default, and the threshold applies to the pool. - It needs a running supervisor. The check hangs off the supervisor's loop. When Horizon is down there is no loop, and the one alert you have cannot fire during the outage it was meant for.
Events Horizon fires and tells nobody about
Horizon raises several other events at the moments you would want to hear about, with no notification attached. SupervisorOutOfMemory fires when a supervisor exceeds its memory limit and is about to be terminated, MasterSupervisorOutOfMemory does the same for the master, and UnableToLaunchProcess fires when a worker process will not start. You can listen to them yourself:
use Illuminate\Support\Facades\Event;
use Illuminate\Support\Facades\Log;
use Laravel\Horizon\Events\SupervisorOutOfMemory;
use Laravel\Horizon\Events\UnableToLaunchProcess;
Event::listen(SupervisorOutOfMemory::class, function (SupervisorOutOfMemory $event) {
Log::channel('slack')->critical('Horizon supervisor out of memory', [
'supervisor' => $event->supervisor->name,
'memory_mb' => $event->getMemoryUsage(),
]);
});
Event::listen(UnableToLaunchProcess::class, function (UnableToLaunchProcess $event) {
Log::channel('slack')->critical('Horizon could not launch a worker', [
'command' => $event->process->process->getCommandLine(),
]);
});
The same pattern with Laravel's JobFailed event is the alert most teams write first, and the one that gets muted first. One malformed record on a job dispatched four hundred times an hour fills the channel before anyone is awake. A failure alert has to be about a rate, and a rate needs something that counts.
For Horizon being down, the built-in tool is php artisan horizon:status. It exits 0 when Horizon is running, 1 when a master is paused and 2 when no master is reporting, so a cron job or an uptime monitor that runs commands can watch it from outside the fleet.
Gauges from spatie/laravel-prometheus
spatie/laravel-prometheus is a general Prometheus library for Laravel with a set of Horizon collectors. Uncomment $this->registerHorizonCollectors() in the published PrometheusServiceProvider and the /prometheus route serves seven gauges. The default namespace is app, so the series arrive with that prefix.
| Series | What it reads |
|---|---|
app_horizon_status |
1 running, 0 if any master is paused, -1 if no master is reporting |
app_horizon_master_supervisors |
How many master supervisors are reporting |
app_horizon_current_workload{queue} |
Ready jobs per supervisor pool |
app_horizon_current_processes{queue} |
Worker processes per supervisor pool |
app_horizon_failed_recent_jobs |
Failures inside the trim.recent_failed window, a week by default |
app_horizon_jobs_per_minute |
Horizon's own jobs-per-minute figure for the current snapshot window |
app_horizon_recent_jobs |
Jobs inside the trim.recent window |
Our production rule file has five Horizon alerts. Four of them only ask a current-value question, which is what a gauge is for, and they port to these series directly:
groups:
- name: horizon-spatie
rules:
- alert: HorizonNotRunning
expr: app_horizon_status == -1 or absent(app_horizon_status)
for: 5m
labels:
severity: critical
- alert: HorizonPaused
expr: app_horizon_status == 0
for: 15m
labels:
severity: critical
# The backlog is the same size it was 15 minutes ago, and not empty.
- alert: QueueProcessingStalled
expr: |
(app_horizon_current_workload offset 15m == app_horizon_current_workload)
and (app_horizon_current_workload > 0)
for: 15m
labels:
severity: critical
# The queue has not reached zero once in 24 hours. The inner max_over_time
# bridges short gaps in the series left by restarts and deploys.
- alert: QueueBacklogNotZeroFor24Hours
expr: |
min_over_time(
max_over_time(app_horizon_current_workload[10m])[24h:1m]
) > 0
labels:
severity: critical
absent() matters in the first rule. If the application is unreachable or the scrape target was removed, the series is gone, and a comparison against a series that does not exist matches nothing.
Two details change compared with the originals. app_horizon_status folds paused and running into one series with no label, so HorizonPaused cannot name which master is paused. And the queue label is the supervisor's queue string: a pool working high,default arrives as one series with queue="high,default", and the workload is the pool's total.
app_horizon_master_supervisors gives you one more rule for free. Set it against the number of servers that should be running Horizon, and it catches both a server that dropped out and the duplicate master tree from our supervisor upgrade postmortem:
- alert: HorizonMasterCountWrong
expr: app_horizon_master_supervisors != 2 # your server count
for: 10m
The rule that does not port
Our fifth rule is "more than 20 failed jobs in an hour, sustained for two hours". It is written with increase() over a failure counter, and the Spatie collectors have no counter. app_horizon_failed_recent_jobs is a gauge over a sliding window, so it goes down as old failures leave the window. increase() treats every drop as a counter reset and the result means nothing.
There is a workaround if you can live with its side effect. The gauge counts failures newer than trim.recent_failed minutes, so setting that to 60 turns it into "failures in the past hour":
// config/horizon.php
'trim' => [
'recent_failed' => 60, // was 10080
'failed' => 10080, // the failed jobs list keeps its week
],
- alert: FailedJobsIncreasing
expr: app_horizon_failed_recent_jobs > 20
for: 2h
The side effect is that the dashboard's failed-jobs figure becomes the past hour as well, since it reads the same setting. The figure is also fleet-wide. There is no queue or job class on it, so the alert tells you that something is failing and you go to the dashboard to find out what.
What is missing from the Horizon collectors
There is no age of the oldest pending job, no wait time, no retries and nothing per job class. Spatie's separate queue collectors do export
queue_oldest_pending_job_age, but their documentation says not to run them alongside the Horizon collectors. The jobs-per-minute gauge has its own problems as an alert input, covered in the Horizon in Prometheus post. Also setallowed_ipsbefore deploying the route, because an empty list allows everyone.
Counters change which rules you can write
Our Horizon Prometheus Exporter is open source and runs on stock Horizon. Workers count each outcome into Redis, the counts only go up, and /horizon/prometheus serves them. It exports processed, failed and retried counters per queue and per job class, a wait-time summary, and gauges for queue length, oldest pending age, time to clear, processes and paused state. Every queue label names one queue, with the pool in a group label.
composer require boring-o11y/horizon-prometheus-exporter
On this exporter our five production rules run as they are written. We scrape Skyline's endpoint rather than this package, and the two use the same metric names, so the file is the same either way:
groups:
- name: HorizonAlerts
rules:
- alert: HorizonNotRunning
expr: horizon_up != 1 or absent(horizon_up)
for: 5m
- alert: HorizonPaused
expr: horizon_master_supervisor_paused == 1
for: 15m
- alert: QueueProcessingStalled
expr: |
(horizon_queue_length{queue!=""} offset 15m == horizon_queue_length{queue!=""})
and (horizon_queue_length{queue!=""} > 0)
for: 15m
- alert: QueueBacklogNotZeroFor24Hours
expr: |
min_over_time(
max_over_time(horizon_queue_length{queue!=""}[10m])[24h:1m]
) > 0
- alert: FailedJobsIncreasing
expr: sum(increase(horizon_queue_failed_total[1h])) > 20
for: 2h
HorizonPaused now carries a master label, so the alert names the machine. FailedJobsIncreasing is a real increase() over a counter and needs no change to Horizon's trimming.
Those five were written when gauges were all we had. With counters and the oldest-pending gauge there are better ones, and these are the rules we would add.
Latency instead of length
- alert: QueueBacklogAgeing
expr: max by (queue) (horizon_queue_oldest_pending_seconds) > 900
for: 5m
- alert: QueueWaitHigh
expr: |
sum by (queue) (rate(horizon_queue_wait_seconds_sum[10m]))
/ sum by (queue) (rate(horizon_queue_wait_seconds_count[10m])) > 120
for: 10m
The first is how long the job at the head of the queue has been ready to run. It is 0 on an empty queue, small on a fast one however deep it is, and climbs without limit when nothing is consuming. The second is the average wait of the jobs that were picked up, which catches a queue that is moving but slower than it promised. Set both per queue from what the queue is for. A password reset is late at a minute and a nightly export is fine at half an hour.
A backlog nobody is draining
- alert: QueueNotMoving
expr: |
max by (queue) (horizon_queue_length) > 0
unless sum by (queue) (rate(horizon_queue_processed_total[10m])) > 0
unless max by (queue) (horizon_queue_paused) == 1
for: 10m
This replaces QueueProcessingStalled. Comparing a length with itself fifteen minutes ago has two failure cases: a busy queue that happens to land on the same number, and a stalled queue that keeps receiving dispatches, whose length changes every minute. Asking whether anything finished has neither. The last clause leaves out queues somebody paused on purpose, which get their own low-severity rule:
- alert: QueueLeftPaused
expr: max by (queue) (horizon_queue_paused) == 1
for: 1h
labels:
severity: ticket
Failure ratio, per queue and per job class
- alert: QueueFailureRatio
expr: |
sum by (queue) (rate(horizon_queue_failed_total[5m]))
/ (
sum by (queue) (rate(horizon_queue_processed_total[5m]))
+ sum by (queue) (rate(horizon_queue_failed_total[5m]))
) > 0.05
and sum by (queue) (rate(horizon_queue_processed_total[5m])) > 0.1
for: 10m
- alert: JobClassFailing
expr: sum by (job_class) (increase(horizon_job_failed_total[15m])) > 10
The and clause is a throughput floor of about six jobs a minute. Without it a queue that ran one job in five minutes and failed it reads as 100 per cent. The per-class rule is there because a ratio per queue dilutes: one class failing every run on a queue shared with twenty others can stay under five per cent indefinitely.
Falling behind, and retry churn
- alert: QueueFallingBehind
expr: |
max by (queue) (deriv(horizon_queue_length[10m])) > 0
and max by (queue) (horizon_queue_time_to_clear_seconds) > 900
for: 10m
labels:
severity: ticket
- alert: JobRetryChurn
expr: |
sum by (job_class) (rate(horizon_job_retried_total[15m]))
/ sum by (job_class) (rate(horizon_job_processed_total[15m])) > 1
for: 30m
labels:
severity: ticket
Time to clear is a forecast, and a forecast over a line is a bad page on its own. Paired with a backlog that is still growing it says arrivals are outpacing workers while the fix is still cheap. The retry rule counts jobs released after an exception, so a class retrying more often than it completes is leaning on its tries to get through.
The same rules as Grafana alerts
Everything above is a Prometheus rule file, evaluated by Prometheus and routed by Alertmanager. Grafana lists those rules read-only under its alerting pages. If you would rather have Grafana evaluate them, each one becomes a Grafana-managed rule with the PromQL as the query and for as the pending period.
One thing needs to change on the way over. A PromQL comparison filters: horizon_queue_oldest_pending_seconds > 900 returns no series at all when every queue is healthy, and Grafana reports that as No Data. Query the bare metric and put the 900 in a threshold expression, so the rule always has a value to evaluate. Then set "Alert state if no data" to Alerting on the horizon_up rule, which does the job absent() does in the rule file.
Whichever evaluates them, scrape one web server. Every instance reads the same Redis, so three targets return the same series three times and every sum() triples.
Where a scrape runs out
A scrape carries numbers that exist for every queue or job class at every moment. That shape fits throughput, backlog and latency well. It fits badly when the fault is a fact about one lock, one job or one line of config.
Take a job class that has silently stopped running because a unique lock outlived the worker that held it. Every dispatch is skipped. Nothing fails, so the failure ratio is flat. Nothing is queued, so length and age are zero. Throughput for that class drops to nothing, which looks the same as a quiet afternoon. You could write an "expected throughput" rule per class, and you would be maintaining a list of what each job ought to be doing.
The other gap is detail. A rule can say the failure ratio on default is eight per cent. It cannot say which lock, which limiter or which supervisor setting is responsible, because those were never labels.
The alerts Skyline adds
Skyline is our fork of Horizon, and its alert pipeline runs inside the application against the same Redis. It covers the ground above without a metrics stack: Horizon down, oldest job age per queue, a queue with workers and no completions, a backlog that keeps growing, a pause nobody resumed and a failure rate per queue or per job class. If you have built the Prometheus rules you already have those, and Skyline's endpoint exports the same metric names, so the rule file keeps working.
The checks below are the ones the rule file has no equivalent for. They read lock owners, limiter buckets, worker exit reasons and supervisor config, none of which either exporter publishes.
| Check | Fires when | Why the metrics miss it |
|---|---|---|
stranded_lock |
A unique or overlap lock is still held and the job that took it is gone | No failures and no backlog; the dispatches are skipped |
limiter_drop |
Rate-limit, overlap or exception-throttle middleware dropped or deleted jobs | A dropped job is recorded as completed |
reservation_expired |
A job was still reserved when retry_after ran out, so it will run twice |
Both runs succeed and count as processed |
misconfiguration |
A supervisor's timeout is not below its connection's retry_after
|
It is config, and nothing has gone wrong yet |
job_timeout |
One job class keeps timing out | A timed-out attempt is retried, and only counts as failed when attempts run out |
worker_crash_loop |
Workers die without reporting a reason, faster than a threshold | The supervisor replaces them, so the process count never drops |
memory_restarts |
Workers keep being recycled for their memory limit | Nothing fails; throughput sags while workers boot |
process_launch_failed |
A worker process would not start | A single event between two scrapes |
supervisor_out_of_memory, master_out_of_memory
|
A supervisor or the master exceeded its memory limit and was restarted | A single event between two scrapes |
Work that vanishes without failing
stranded_lock is the case from the previous section. Skyline records which job holds each ShouldBeUnique and WithoutOverlapping lock for its Locks & Limits screen, so the check can tell a lock whose job is still queued or running from one whose job no longer exists. After five minutes in that state it alerts with the job class, the lock key, how long it has been held, when it expires and how many dispatches it has skipped. How locks get stranded in the first place is in the ShouldBeUnique post.
limiter_drop watches the middleware that discards jobs by design: RateLimited with dontRelease(), WithoutOverlapping without releaseAfter(), and ThrottlesExceptions with deleteWhen(). Each leaves the job marked completed, so the processed counter goes up for work that did not happen. The check fires at ten drops in an hour for one limiter or job class. Some applications drop on purpose, a "latest wins" sync for example, so each limiter takes its own threshold or a zero to leave it out:
'limiter_drop' => [
'threshold' => 10,
'circuit_open' => 50, // also alert when an exception throttle holds back 50 jobs an hour
'groups' => [
'App\\Jobs\\SyncLatestPrices' => 0,
'rate_limiter:uploads' => 200,
],
],
Jobs that run twice
When a job is still running at the moment its reservation expires, the queue hands it to a second worker. Both runs complete, both count as processed, and the customer is charged twice. reservation_expired reports it once per queue with a count and the job classes involved, at critical severity, because it is a correctness problem and not a slow one.
misconfiguration is the same fault before it happens. A supervisor whose timeout is at or above its connection's retry_after will produce a double run as soon as one job runs long enough. Stock Horizon accepts that config without comment. Skyline prints a warning when it starts, where nobody is looking after a deploy that went fine, so the check also alerts with the supervisor, the connection and both values, repeats once a day, and resolves when a deploy picks up the corrected config. The timeout vs retry_after post explains why the two have to be ordered that way.
Workers dying, by cause
Stock Horizon shows the configured number of processes whether workers are healthy or being replaced every thirty seconds. Skyline records why each worker stopped, and three checks read that record. worker_crash_loop counts workers that died without giving a reason, which means a PHP fatal or a kill from outside, and fires at five in five minutes per supervisor. memory_restarts counts workers recycled for reaching their memory limit and fires at ten in fifteen minutes. One memory restart is the supervisor doing its job, so the crash check leaves those out and this one looks at the rate.
The memory alert also names the job classes that retained the most heap over the window, since Skyline measures heap growth per class. That turns "workers on default keep running out of memory" into a class name to open. Both checks need Insights enabled, which is what records the restarts.
job_timeout fires when one class times out five times in ten minutes. A timed-out job is killed, retried, and often times out again, so the failure rate only sees it once the attempts are spent. By then a slow dependency has been holding workers for an hour. Timeouts are counted per class and per queue, and Skyline exports them as horizon_job_timed_out_total and horizon_queue_timed_out_total if you would rather write that rule in Prometheus.
The last three are the events from the top of this post: process_launch_failed, supervisor_out_of_memory and master_out_of_memory. They report once when the event happens, through the same channels as everything else, in place of the listeners you would otherwise write.
How they are delivered
Every check holds for a for window before notifying, repeats at most every fifteen minutes while it lasts, and sends a recovery message saying how long the condition held. Notifications are withheld for five minutes after a deploy while the timers keep counting. Each alert has a severity, and routes decides which channels a severity goes to, across mail, Slack, SMS and a JSON webhook for PagerDuty or Opsgenie.
HORIZON_ALERTS=true
HORIZON_ALERT_SLACK="https://hooks.slack.com/services/..."
HORIZON_ALERT_WEBHOOK="https://events.example.com/queue-alerts"
Anything specific to your application goes in a class that extends Check and is registered with Horizon::alertCheck(). "No invoices processed since midnight on a weekday" is a rule like that, and it gets the same hold window, cooldown, recovery message and routing as the built-in ones.
Which layer to stop at
If you run no metrics stack, configure LongWaitDetected with a threshold per queue and put horizon:status behind a cron or an uptime check. That covers a growing backlog and Horizon being down, which are the two outages that cost the most.
If you already run Prometheus and Spatie's package, the four gauge rules are worth the ten minutes they take. Once you want a failure rate, or an alert that separates a busy queue from a stuck one, you need counters and the oldest-pending age, and that is the point to add the exporter.
The Skyline checks are for the failures that leave the graphs looking normal. A rule file and those checks do not compete: the rules stay in Prometheus next to your database and host alerts, and the checks cover the conditions that never reach it.
Common questions
Does Laravel Horizon have built-in alerts?
One. Horizon sends a LongWaitDetected notification by mail, Slack or SMS when a queue's estimated wait passes its threshold in horizon.waits, 60 seconds by default. It repeats every five minutes while the wait lasts and sends nothing when it clears. There is no built-in notification for failed jobs, for a supervisor running out of memory, or for Horizon not running.
How do I get alerted when Horizon stops running?
From outside Horizon, because its own notification is evaluated by a running supervisor. Run php artisan horizon:status from cron or an uptime monitor: it exits 0 when running, 1 when paused and 2 when inactive. With Prometheus, alert on app_horizon_status == -1 from spatie/laravel-prometheus or horizon_up == 0 from a counter exporter, and add absent() so a missing series also fires.
Can I alert on failed jobs with spatie/laravel-prometheus?
Only roughly. Its horizon_failed_recent_jobs gauge counts failures inside Horizon's trim.recent_failed window, a week by default, and has no queue or job class label. It is a sliding window and not a counter, so increase() and rate() do not work on it. Setting trim.recent_failed to 60 makes it failures in the past hour, which also changes the figure on Horizon's dashboard. A failure ratio per queue needs an exporter that keeps counters.
Why does my Grafana alert on a Horizon metric show No Data?
Usually because the comparison is inside the PromQL. A query such as horizon_queue_oldest_pending_seconds > 900 returns no series when every queue is under the threshold, and Grafana treats an empty result as No Data. Query the metric without the comparison and add a threshold expression for the 900, so the rule always has a value to evaluate.
Which Horizon problems can Prometheus not alert on?
The ones that produce no failures and no backlog. A unique lock left behind by a dead worker makes every dispatch of that job get skipped. A job dropped by RateLimited with dontRelease() is recorded as completed. A job that outlives retry_after runs twice and both runs count as processed. These need a check that reads lock owners, limiter counts or supervisor config directly, which is what Skyline's alert checks do.
Top comments (0)