You know your workflow is "fine" because it ran yesterday. But what does fine actually mean?
If your workflow normally processes 50 leads per hour but yesterday it only touched 10, is that a problem? If it usually finishes in 30 seconds but this morning it took 2 minutes, should you page someone? If the success rate was 98% last week and it's 85% this week, is that the early sign of a break or just noise?
These aren't hypothetical. Silent failures live in the gap between "the system says success" and "the system actually worked." Catching them means being able to tell the difference between normal variation and a real problem—and that requires knowing what normal actually looks like for your specific workflow.
That's where baselines come in. Let me walk you through how they work, and why rolling statistical history is the only honest way to do it.
The Problem: Fixed Thresholds Don't Work
Most monitoring solutions use fixed rules: "alert if a run takes more than 60 seconds" or "alert if we process fewer than 100 items."
The trouble is that no two workflows are alike. Your hourly lead-scoring automation might normally finish in 3 seconds, while your daily partner-data sync legitimately takes 7 minutes. A threshold that's sensible for one is nonsense for the other.
And fixed thresholds are blind to the workflow's actual season. If your application-intake workflow normally processes 200 candidates a day during open season but only 20 in the off-season, a static "alert if less than 50" rule will either:
- Ignore the real drop from 200 to 40 (a 80% dive during peak season), or
- Spam you with false alarms every day in the quiet months.
You can't win with fixed numbers. Each workflow has its own shape of normal, and that shape changes.
The Right Approach: Each Workflow Is Its Own Baseline
A baseline-aware system learns what this specific workflow actually looks like. It builds a rolling history of recent runs and judges whether the current execution is unusual relative to that history.
Here's how it works in practice:
Step 1: Collect recent runs. Keep the last 5–10 successful runs in memory (depending on the metric). For a workflow that runs hourly, that's maybe 5–10 hours of history. For a daily workflow, that's 5–10 days. The exact window size is a design choice, but the principle is the same: use what you have.
Step 2: Compute the mean and standard deviation. For duration, you'd calculate the average run time and the spread around it. For item count, the same. For success rate, you'd compute the proportion of successes across those runs.
mean_duration = sum(run_times) / count
std_dev = sqrt(sum((run - mean)^2) / count)
This gives you a statistical shape of what that workflow normally produces.
Step 3: When a new run lands, compare it to the baseline. Is the current run slower than mean + 2σ (two standard deviations above the mean)? That's the statistical threshold where something is genuinely unusual—not just a little slower, but abnormally slower by the workflow's own standards.
For workflows with small baselines—a fast automation that normally runs in 100ms—systems often add a floor: "also alert if it's 3× the mean, even if that's less than 2σ." This prevents hair-trigger alerts on workflows that are so fast that even 2σ is tiny.
Step 4: Flag and alert. If the current run exceeds the threshold, it's genuinely anomalous. Flag it, and let the human decide.
Why This Matters
Here's a concrete example. Imagine a workflow that processes a batch of leads and scores them:
- Runs daily at 9 AM.
- Last 10 runs: 45s, 48s, 52s, 44s, 49s, 47s, 51s, 46s, 50s, 48s.
- Mean: ~48 seconds.
- Standard deviation: ~2.5 seconds.
- Alert threshold (mean + 2σ): ~53 seconds.
Yesterday's run took 52 seconds—well within normal, no alert.
Today's run takes 75 seconds. That's far outside the ~53-second threshold. The system alerts: "this run was significantly slower than usual." You check, and you find that the lead database had 10,000 new unscored records because a data sync stalled. Your workflow is fine; the input grew. But you caught it because the workflow told you something was different.
Compare that to a fixed threshold: if you'd set an alert at "anything over 90 seconds," you'd never have noticed the slowdown. Or if you'd set it at "anything over 55 seconds," you'd be getting false alarms every few weeks when the batch naturally varies.
A baseline-aware system lets you be both sensitive and accurate—you catch real deviations without crying wolf on normal variation.
What This Looks Like in Practice
OpsVeritas uses this approach for three main health checks:
Latency / At Risk: A run that exceeds mean + 2σ (with a 500ms floor and 3× multiplier for tiny-duration workflows), based on recent history.
Volume Drop: Item count falls below 30% of the recent average—a 70%+ drop. This catches the "workflow ran but processed nothing" silent failure.
Degrading: Success rate drops by 25+ percentage points in the last 5 runs vs. the prior 5. This flags when the failure rate is genuinely increasing.
Each check is per-workflow, computed every ~3 minutes, and based on that workflow's actual recent history, not a shared rule.
The neat part: you never have to configure these thresholds yourself. The system learns them automatically by watching your workflow run.
The Limits of Baselines
Baselines aren't magic. They have real limits:
- Cold start: A brand-new workflow with only 1–2 runs won't have enough history to be reliable. Most systems require a small minimum (5+ runs) before they start alerting.
- Sustained changes: If your workflow actually gets slower (you added a new step, or the data volume grew permanently), the baseline gradually shifts to match the new normal. It'll take a few runs before it fully adjusts. This is actually a feature—it avoids false alarms—but it means you won't get an alert the moment a degradation starts; you'll get it once it's clearly sustained.
- Cyclical patterns: Some workflows have weekly or seasonal patterns. A pure rolling baseline can miss these. More sophisticated systems layer in time-of-week or seasonal decomposition, but that's a separate tool.
For most workflows—especially the data-processing and automation kinds—a rolling statistical baseline is the right tradeoff: simple, automatic, and sensitive without being noisy.
The Bigger Picture
All of this comes back to the silent-failure problem. A workflow can say "success" and have done nothing. It can run slower because of a bottleneck upstream. It can process fewer items because of a data condition, not a code problem.
The only way to catch these early is to know what normal looks like for that specific workflow—and that means building baselines from its own history, not from shared rules.
If you're running workflows across n8n, Make, Zapier, or any other platform, this is exactly why monitoring systems that watch your workflow history matter. You can connect https://app.opsveritas.com with a read-only API key or a simple webhook, and it'll start learning your workflows' baselines automatically. No thresholds to set, no guessing games.
Silent failures hide in the gap between intention and reality. Baselines close that gap—not by magic, just by remembering what your workflows actually do.
Top comments (0)