DEV Community

laravel-o11y
laravel-o11y

Posted on Originally published at boring-observability.dev

Laravel Horizon Alerting: What Horizon Sends on Its Own and What Prometheus Adds

Horizon will tell you about one thing: a queue whose estimated wait has crossed a threshold. It will not tell you that the wait ended, that jobs are failing, that a supervisor was killed for memory, or that Horizon itself stopped an hour ago. Most teams find this out the usual way, from a customer asking where their invoice is.

This post walks up the layers you can add. It starts with what ships in Horizon, moves to the gauges spatie/laravel-prometheus exports and the alert rules those support, then to a counter-based exporter and the rules that only counters make possible. The last part covers the conditions no scrape carries, which is where Skyline's alerts come in. The rules are the ones we run in production, and for each we say which exporter can and cannot evaluate it.

Key takeaways

  • Horizon has one notification, LongWaitDetected. It is built on a forecast of time to clear, repeats every five minutes while the condition lasts, never reports recovery, and is evaluated by a supervisor, so it is silent when Horizon is down.
  • Spatie's Horizon collectors are gauges read at scrape time. They are enough for "Horizon is not running", "Horizon is paused" and "this backlog has not moved". They cannot express a failure rate, and they have nothing per job class.
  • Counters unlock the rules worth paging on. With processed, failed and retried counters plus the age of the oldest pending job, you can alert on failure ratio, on a backlog nobody is draining, and on latency instead of length.
  • In Grafana, keep the comparison out of the PromQL. A query like x > 900 returns no series when it is false, and Grafana reads that as No Data. Put the threshold in a threshold expression instead.
  • Some failures never become a time series. A stranded unique lock, a job dropped by a rate limiter, a reservation that expired mid-run and a timeout above retry_after produce no failures and no backlog. Those need checks that read queue state directly.

What Horizon sends on its own

Horizon's notification is configured in two places. The thresholds live in config/horizon.php, keyed by connection and queue, and the destinations are set in your HorizonServiceProvider:

// config/horizon.php
'waits' => [
    'redis:default' => 60,
    'redis:reports' => 1800,
    'redis:housekeeping' => 0,   // 0 turns the check off for that queue
],

// app/Providers/HorizonServiceProvider.php
Horizon::routeSlackNotificationsTo('https://hooks.slack.com/services/...', '#ops');
Horizon::routeMailNotificationsTo('ops@example.com');
Horizon::routeSmsNotificationsTo('+15555550100');
Enter fullscreen mode Exit fullscreen mode

A queue you do not list gets 60 seconds. Once a minute, one supervisor in the fleet takes a lock, calculates the wait for every queue it knows about, and raises LongWaitDetected for each one over its threshold. That event sends the notification. It is a reasonable first alert, and it is worth knowing exactly how it behaves before relying on it.

  • The wait is a forecast. It comes from WaitTimeCalculator: ready jobs multiplied by the queue's average runtime, divided by the processes serving it. Ten thousand fast jobs produce a large number and a single job stuck at the head of a queue produces a small one.
  • It repeats every five minutes. The listener takes a 300-second lock per connection and queue before sending, so a backlog that lasts two hours sends twenty-four messages.
  • It never says the wait ended. There is no recovery message, so a channel that went quiet might mean the queue drained or might mean someone muted it.
  • Queues balanced together are one entry. A supervisor working high,default is keyed redis:high,default, and the threshold applies to the pool.
  • It needs a running supervisor. The check hangs off the supervisor's loop. When Horizon is down there is no loop, and the one alert you have cannot fire during the outage it was meant for.

Events Horizon fires and tells nobody about

Horizon raises several other events at the moments you would want to hear about, with no notification attached. SupervisorOutOfMemory fires when a supervisor exceeds its memory limit and is about to be terminated, MasterSupervisorOutOfMemory does the same for the master, and UnableToLaunchProcess fires when a worker process will not start. You can listen to them yourself:

use Illuminate\Support\Facades\Event;
use Illuminate\Support\Facades\Log;
use Laravel\Horizon\Events\SupervisorOutOfMemory;
use Laravel\Horizon\Events\UnableToLaunchProcess;

Event::listen(SupervisorOutOfMemory::class, function (SupervisorOutOfMemory $event) {
    Log::channel('slack')->critical('Horizon supervisor out of memory', [
        'supervisor' => $event->supervisor->name,
        'memory_mb' => $event->getMemoryUsage(),
    ]);
});

Event::listen(UnableToLaunchProcess::class, function (UnableToLaunchProcess $event) {
    Log::channel('slack')->critical('Horizon could not launch a worker', [
        'command' => $event->process->process->getCommandLine(),
    ]);
});
Enter fullscreen mode Exit fullscreen mode

The same pattern with Laravel's JobFailed event is the alert most teams write first, and the one that gets muted first. One malformed record on a job dispatched four hundred times an hour fills the channel before anyone is awake. A failure alert has to be about a rate, and a rate needs something that counts.

For Horizon being down, the built-in tool is php artisan horizon:status. It exits 0 when Horizon is running, 1 when a master is paused and 2 when no master is reporting, so a cron job or an uptime monitor that runs commands can watch it from outside the fleet.

Gauges from spatie/laravel-prometheus

spatie/laravel-prometheus is a general Prometheus library for Laravel with a set of Horizon collectors. Uncomment $this->registerHorizonCollectors() in the published PrometheusServiceProvider and the /prometheus route serves seven gauges. The default namespace is app, so the series arrive with that prefix.

Series What it reads
app_horizon_status 1 running, 0 if any master is paused, -1 if no master is reporting
app_horizon_master_supervisors How many master supervisors are reporting
app_horizon_current_workload{queue} Ready jobs per supervisor pool
app_horizon_current_processes{queue} Worker processes per supervisor pool
app_horizon_failed_recent_jobs Failures inside the trim.recent_failed window, a week by default
app_horizon_jobs_per_minute Horizon's own jobs-per-minute figure for the current snapshot window
app_horizon_recent_jobs Jobs inside the trim.recent window

Our production rule file has five Horizon alerts. Four of them only ask a current-value question, which is what a gauge is for, and they port to these series directly:

groups:
  - name: horizon-spatie
    rules:
      - alert: HorizonNotRunning
        expr: app_horizon_status == -1 or absent(app_horizon_status)
        for: 5m
        labels:
          severity: critical

      - alert: HorizonPaused
        expr: app_horizon_status == 0
        for: 15m
        labels:
          severity: critical

      # The backlog is the same size it was 15 minutes ago, and not empty.
      - alert: QueueProcessingStalled
        expr: |
          (app_horizon_current_workload offset 15m == app_horizon_current_workload)
          and (app_horizon_current_workload > 0)
        for: 15m
        labels:
          severity: critical

      # The queue has not reached zero once in 24 hours. The inner max_over_time
      # bridges short gaps in the series left by restarts and deploys.
      - alert: QueueBacklogNotZeroFor24Hours
        expr: |
          min_over_time(
            max_over_time(app_horizon_current_workload[10m])[24h:1m]
          ) > 0
        labels:
          severity: critical
Enter fullscreen mode Exit fullscreen mode

absent() matters in the first rule. If the application is unreachable or the scrape target was removed, the series is gone, and a comparison against a series that does not exist matches nothing.

Two details change compared with the originals. app_horizon_status folds paused and running into one series with no label, so HorizonPaused cannot name which master is paused. And the queue label is the supervisor's queue string: a pool working high,default arrives as one series with queue="high,default", and the workload is the pool's total.

app_horizon_master_supervisors gives you one more rule for free. Set it against the number of servers that should be running Horizon, and it catches both a server that dropped out and the duplicate master tree from our supervisor upgrade postmortem:

- alert: HorizonMasterCountWrong
  expr: app_horizon_master_supervisors != 2   # your server count
  for: 10m
Enter fullscreen mode Exit fullscreen mode

The rule that does not port

Our fifth rule is "more than 20 failed jobs in an hour, sustained for two hours". It is written with increase() over a failure counter, and the Spatie collectors have no counter. app_horizon_failed_recent_jobs is a gauge over a sliding window, so it goes down as old failures leave the window. increase() treats every drop as a counter reset and the result means nothing.

There is a workaround if you can live with its side effect. The gauge counts failures newer than trim.recent_failed minutes, so setting that to 60 turns it into "failures in the past hour":

// config/horizon.php
'trim' => [
    'recent_failed' => 60,   // was 10080
    'failed' => 10080,       // the failed jobs list keeps its week
],
Enter fullscreen mode Exit fullscreen mode
- alert: FailedJobsIncreasing
  expr: app_horizon_failed_recent_jobs > 20
  for: 2h
Enter fullscreen mode Exit fullscreen mode

The side effect is that the dashboard's failed-jobs figure becomes the past hour as well, since it reads the same setting. The figure is also fleet-wide. There is no queue or job class on it, so the alert tells you that something is failing and you go to the dashboard to find out what.

What is missing from the Horizon collectors

There is no age of the oldest pending job, no wait time, no retries and nothing per job class. Spatie's separate queue collectors do export queue_oldest_pending_job_age, but their documentation says not to run them alongside the Horizon collectors. The jobs-per-minute gauge has its own problems as an alert input, covered in the Horizon in Prometheus post. Also set allowed_ips before deploying the route, because an empty list allows everyone.

Counters change which rules you can write

Our Horizon Prometheus Exporter is open source and runs on stock Horizon. Workers count each outcome into Redis, the counts only go up, and /horizon/prometheus serves them. It exports processed, failed and retried counters per queue and per job class, a wait-time summary, and gauges for queue length, oldest pending age, time to clear, processes and paused state. Every queue label names one queue, with the pool in a group label.

composer require boring-o11y/horizon-prometheus-exporter
Enter fullscreen mode Exit fullscreen mode

On this exporter our five production rules run as they are written. We scrape Skyline's endpoint rather than this package, and the two use the same metric names, so the file is the same either way:

groups:
  - name: HorizonAlerts
    rules:
      - alert: HorizonNotRunning
        expr: horizon_up != 1 or absent(horizon_up)
        for: 5m

      - alert: HorizonPaused
        expr: horizon_master_supervisor_paused == 1
        for: 15m

      - alert: QueueProcessingStalled
        expr: |
          (horizon_queue_length{queue!=""} offset 15m == horizon_queue_length{queue!=""})
          and (horizon_queue_length{queue!=""} > 0)
        for: 15m

      - alert: QueueBacklogNotZeroFor24Hours
        expr: |
          min_over_time(
            max_over_time(horizon_queue_length{queue!=""}[10m])[24h:1m]
          ) > 0

      - alert: FailedJobsIncreasing
        expr: sum(increase(horizon_queue_failed_total[1h])) > 20
        for: 2h
Enter fullscreen mode Exit fullscreen mode

HorizonPaused now carries a master label, so the alert names the machine. FailedJobsIncreasing is a real increase() over a counter and needs no change to Horizon's trimming.

Those five were written when gauges were all we had. With counters and the oldest-pending gauge there are better ones, and these are the rules we would add.

Latency instead of length

- alert: QueueBacklogAgeing
  expr: max by (queue) (horizon_queue_oldest_pending_seconds) > 900
  for: 5m

- alert: QueueWaitHigh
  expr: |
    sum by (queue) (rate(horizon_queue_wait_seconds_sum[10m]))
      / sum by (queue) (rate(horizon_queue_wait_seconds_count[10m])) > 120
  for: 10m
Enter fullscreen mode Exit fullscreen mode

The first is how long the job at the head of the queue has been ready to run. It is 0 on an empty queue, small on a fast one however deep it is, and climbs without limit when nothing is consuming. The second is the average wait of the jobs that were picked up, which catches a queue that is moving but slower than it promised. Set both per queue from what the queue is for. A password reset is late at a minute and a nightly export is fine at half an hour.

A backlog nobody is draining

- alert: QueueNotMoving
  expr: |
    max by (queue) (horizon_queue_length) > 0
    unless sum by (queue) (rate(horizon_queue_processed_total[10m])) > 0
    unless max by (queue) (horizon_queue_paused) == 1
  for: 10m
Enter fullscreen mode Exit fullscreen mode

This replaces QueueProcessingStalled. Comparing a length with itself fifteen minutes ago has two failure cases: a busy queue that happens to land on the same number, and a stalled queue that keeps receiving dispatches, whose length changes every minute. Asking whether anything finished has neither. The last clause leaves out queues somebody paused on purpose, which get their own low-severity rule:

- alert: QueueLeftPaused
  expr: max by (queue) (horizon_queue_paused) == 1
  for: 1h
  labels:
    severity: ticket
Enter fullscreen mode Exit fullscreen mode

Failure ratio, per queue and per job class

- alert: QueueFailureRatio
  expr: |
    sum by (queue) (rate(horizon_queue_failed_total[5m]))
      / (
        sum by (queue) (rate(horizon_queue_processed_total[5m]))
        + sum by (queue) (rate(horizon_queue_failed_total[5m]))
      ) > 0.05
    and sum by (queue) (rate(horizon_queue_processed_total[5m])) > 0.1
  for: 10m

- alert: JobClassFailing
  expr: sum by (job_class) (increase(horizon_job_failed_total[15m])) > 10
Enter fullscreen mode Exit fullscreen mode

The and clause is a throughput floor of about six jobs a minute. Without it a queue that ran one job in five minutes and failed it reads as 100 per cent. The per-class rule is there because a ratio per queue dilutes: one class failing every run on a queue shared with twenty others can stay under five per cent indefinitely.

Falling behind, and retry churn

- alert: QueueFallingBehind
  expr: |
    max by (queue) (deriv(horizon_queue_length[10m])) > 0
    and max by (queue) (horizon_queue_time_to_clear_seconds) > 900
  for: 10m
  labels:
    severity: ticket

- alert: JobRetryChurn
  expr: |
    sum by (job_class) (rate(horizon_job_retried_total[15m]))
      / sum by (job_class) (rate(horizon_job_processed_total[15m])) > 1
  for: 30m
  labels:
    severity: ticket
Enter fullscreen mode Exit fullscreen mode

Time to clear is a forecast, and a forecast over a line is a bad page on its own. Paired with a backlog that is still growing it says arrivals are outpacing workers while the fix is still cheap. The retry rule counts jobs released after an exception, so a class retrying more often than it completes is leaning on its tries to get through.

The same rules as Grafana alerts

Everything above is a Prometheus rule file, evaluated by Prometheus and routed by Alertmanager. Grafana lists those rules read-only under its alerting pages. If you would rather have Grafana evaluate them, each one becomes a Grafana-managed rule with the PromQL as the query and for as the pending period.

One thing needs to change on the way over. A PromQL comparison filters: horizon_queue_oldest_pending_seconds > 900 returns no series at all when every queue is healthy, and Grafana reports that as No Data. Query the bare metric and put the 900 in a threshold expression, so the rule always has a value to evaluate. Then set "Alert state if no data" to Alerting on the horizon_up rule, which does the job absent() does in the rule file.

Whichever evaluates them, scrape one web server. Every instance reads the same Redis, so three targets return the same series three times and every sum() triples.

Where a scrape runs out

A scrape carries numbers that exist for every queue or job class at every moment. That shape fits throughput, backlog and latency well. It fits badly when the fault is a fact about one lock, one job or one line of config.

Take a job class that has silently stopped running because a unique lock outlived the worker that held it. Every dispatch is skipped. Nothing fails, so the failure ratio is flat. Nothing is queued, so length and age are zero. Throughput for that class drops to nothing, which looks the same as a quiet afternoon. You could write an "expected throughput" rule per class, and you would be maintaining a list of what each job ought to be doing.

The other gap is detail. A rule can say the failure ratio on default is eight per cent. It cannot say which lock, which limiter or which supervisor setting is responsible, because those were never labels.

The alerts Skyline adds

Skyline is our fork of Horizon, and its alert pipeline runs inside the application against the same Redis. It covers the ground above without a metrics stack: Horizon down, oldest job age per queue, a queue with workers and no completions, a backlog that keeps growing, a pause nobody resumed and a failure rate per queue or per job class. If you have built the Prometheus rules you already have those, and Skyline's endpoint exports the same metric names, so the rule file keeps working.

The checks below are the ones the rule file has no equivalent for. They read lock owners, limiter buckets, worker exit reasons and supervisor config, none of which either exporter publishes.

Check Fires when Why the metrics miss it
stranded_lock A unique or overlap lock is still held and the job that took it is gone No failures and no backlog; the dispatches are skipped
limiter_drop Rate-limit, overlap or exception-throttle middleware dropped or deleted jobs A dropped job is recorded as completed
reservation_expired A job was still reserved when retry_after ran out, so it will run twice Both runs succeed and count as processed
misconfiguration A supervisor's timeout is not below its connection's retry_after It is config, and nothing has gone wrong yet
job_timeout One job class keeps timing out A timed-out attempt is retried, and only counts as failed when attempts run out
worker_crash_loop Workers die without reporting a reason, faster than a threshold The supervisor replaces them, so the process count never drops
memory_restarts Workers keep being recycled for their memory limit Nothing fails; throughput sags while workers boot
process_launch_failed A worker process would not start A single event between two scrapes
supervisor_out_of_memory, master_out_of_memory A supervisor or the master exceeded its memory limit and was restarted A single event between two scrapes

Work that vanishes without failing

stranded_lock is the case from the previous section. Skyline records which job holds each ShouldBeUnique and WithoutOverlapping lock for its Locks & Limits screen, so the check can tell a lock whose job is still queued or running from one whose job no longer exists. After five minutes in that state it alerts with the job class, the lock key, how long it has been held, when it expires and how many dispatches it has skipped. How locks get stranded in the first place is in the ShouldBeUnique post.

limiter_drop watches the middleware that discards jobs by design: RateLimited with dontRelease(), WithoutOverlapping without releaseAfter(), and ThrottlesExceptions with deleteWhen(). Each leaves the job marked completed, so the processed counter goes up for work that did not happen. The check fires at ten drops in an hour for one limiter or job class. Some applications drop on purpose, a "latest wins" sync for example, so each limiter takes its own threshold or a zero to leave it out:

'limiter_drop' => [
    'threshold' => 10,
    'circuit_open' => 50,   // also alert when an exception throttle holds back 50 jobs an hour
    'groups' => [
        'App\\Jobs\\SyncLatestPrices' => 0,
        'rate_limiter:uploads' => 200,
    ],
],
Enter fullscreen mode Exit fullscreen mode

Jobs that run twice

When a job is still running at the moment its reservation expires, the queue hands it to a second worker. Both runs complete, both count as processed, and the customer is charged twice. reservation_expired reports it once per queue with a count and the job classes involved, at critical severity, because it is a correctness problem and not a slow one.

misconfiguration is the same fault before it happens. A supervisor whose timeout is at or above its connection's retry_after will produce a double run as soon as one job runs long enough. Stock Horizon accepts that config without comment. Skyline prints a warning when it starts, where nobody is looking after a deploy that went fine, so the check also alerts with the supervisor, the connection and both values, repeats once a day, and resolves when a deploy picks up the corrected config. The timeout vs retry_after post explains why the two have to be ordered that way.

Workers dying, by cause

Stock Horizon shows the configured number of processes whether workers are healthy or being replaced every thirty seconds. Skyline records why each worker stopped, and three checks read that record. worker_crash_loop counts workers that died without giving a reason, which means a PHP fatal or a kill from outside, and fires at five in five minutes per supervisor. memory_restarts counts workers recycled for reaching their memory limit and fires at ten in fifteen minutes. One memory restart is the supervisor doing its job, so the crash check leaves those out and this one looks at the rate.

The memory alert also names the job classes that retained the most heap over the window, since Skyline measures heap growth per class. That turns "workers on default keep running out of memory" into a class name to open. Both checks need Insights enabled, which is what records the restarts.

job_timeout fires when one class times out five times in ten minutes. A timed-out job is killed, retried, and often times out again, so the failure rate only sees it once the attempts are spent. By then a slow dependency has been holding workers for an hour. Timeouts are counted per class and per queue, and Skyline exports them as horizon_job_timed_out_total and horizon_queue_timed_out_total if you would rather write that rule in Prometheus.

The last three are the events from the top of this post: process_launch_failed, supervisor_out_of_memory and master_out_of_memory. They report once when the event happens, through the same channels as everything else, in place of the listeners you would otherwise write.

How they are delivered

Every check holds for a for window before notifying, repeats at most every fifteen minutes while it lasts, and sends a recovery message saying how long the condition held. Notifications are withheld for five minutes after a deploy while the timers keep counting. Each alert has a severity, and routes decides which channels a severity goes to, across mail, Slack, SMS and a JSON webhook for PagerDuty or Opsgenie.

HORIZON_ALERTS=true
HORIZON_ALERT_SLACK="https://hooks.slack.com/services/..."
HORIZON_ALERT_WEBHOOK="https://events.example.com/queue-alerts"
Enter fullscreen mode Exit fullscreen mode

Anything specific to your application goes in a class that extends Check and is registered with Horizon::alertCheck(). "No invoices processed since midnight on a weekday" is a rule like that, and it gets the same hold window, cooldown, recovery message and routing as the built-in ones.

Which layer to stop at

If you run no metrics stack, configure LongWaitDetected with a threshold per queue and put horizon:status behind a cron or an uptime check. That covers a growing backlog and Horizon being down, which are the two outages that cost the most.

If you already run Prometheus and Spatie's package, the four gauge rules are worth the ten minutes they take. Once you want a failure rate, or an alert that separates a busy queue from a stuck one, you need counters and the oldest-pending age, and that is the point to add the exporter.

The Skyline checks are for the failures that leave the graphs looking normal. A rule file and those checks do not compete: the rules stay in Prometheus next to your database and host alerts, and the checks cover the conditions that never reach it.

Common questions

Does Laravel Horizon have built-in alerts?

One. Horizon sends a LongWaitDetected notification by mail, Slack or SMS when a queue's estimated wait passes its threshold in horizon.waits, 60 seconds by default. It repeats every five minutes while the wait lasts and sends nothing when it clears. There is no built-in notification for failed jobs, for a supervisor running out of memory, or for Horizon not running.

How do I get alerted when Horizon stops running?

From outside Horizon, because its own notification is evaluated by a running supervisor. Run php artisan horizon:status from cron or an uptime monitor: it exits 0 when running, 1 when paused and 2 when inactive. With Prometheus, alert on app_horizon_status == -1 from spatie/laravel-prometheus or horizon_up == 0 from a counter exporter, and add absent() so a missing series also fires.

Can I alert on failed jobs with spatie/laravel-prometheus?

Only roughly. Its horizon_failed_recent_jobs gauge counts failures inside Horizon's trim.recent_failed window, a week by default, and has no queue or job class label. It is a sliding window and not a counter, so increase() and rate() do not work on it. Setting trim.recent_failed to 60 makes it failures in the past hour, which also changes the figure on Horizon's dashboard. A failure ratio per queue needs an exporter that keeps counters.

Why does my Grafana alert on a Horizon metric show No Data?

Usually because the comparison is inside the PromQL. A query such as horizon_queue_oldest_pending_seconds > 900 returns no series when every queue is under the threshold, and Grafana treats an empty result as No Data. Query the metric without the comparison and add a threshold expression for the 900, so the rule always has a value to evaluate.

Which Horizon problems can Prometheus not alert on?

The ones that produce no failures and no backlog. A unique lock left behind by a dead worker makes every dispatch of that job get skipped. A job dropped by RateLimited with dontRelease() is recorded as completed. A job that outlives retry_after runs twice and both runs count as processed. These need a check that reads lock owners, limiter counts or supervisor config directly, which is what Skyline's alert checks do.

Top comments (0)