DEV Community

Cover image for Your sampleRate won't save you from an error flood. Here's what does
Amorizz
Amorizz

Posted on

Your sampleRate won't save you from an error flood. Here's what does

I wrote a flood test for the error tracker I'm building: one bug thrown 997 times in a retry loop, plus three rare bugs thrown once each. In the run that looked most like production, all three rare bugs vanished before they left the Node process.

Then I tried the fix everyone reaches for, sampleRate: 0.1. Across five runs, 2 of the 15 rare errors made it. The loop that caused the mess was still the biggest thing on the screen.

TL;DR

  • sampleRate drops errors at random. It cuts the noisy loop by 90% and your one-off bugs by the same 90%, so the error you needed is the one most likely to disappear.
  • The SDK drops events before sampling even runs: Dedupe eats exact repeats, and a tight burst overflows the 64-slot transport buffer.
  • What worked: send everything, then rate limit stored bodies per fingerprint on the server. Count 100% of hits, keep the first 100 stack traces per bug per minute, let every new bug through.

Why doesn't sampleRate fix an error flood?

Because it doesn't know which error is the flood. sampleRate is a per-event coin flip in the SDK. At 0.1, every error has a 10% chance of being sent, whether it's the 900th copy of the same TypeError or the only RangeError you'll see all week.

That's fine for volume and bad for signal. A bug that fires once has a 90% chance of never reaching you. The count you do see is off by about 10x, because the server gets roughly 100 events and nothing in them says "multiply me". Client reports tell the server how many events were dropped and why (sample_rate, queue_overflow), but not which bug they belonged to.

The flood lab: one loop, three rare bugs

You can reproduce this in a few minutes with Node 20.19 or newer. No Docker, no account. ingest.js is a fake endpoint that parses just enough of the Sentry envelope format to count errors per fingerprint. flood.js uses the real @sentry/node SDK.

mkdir flood-lab && cd flood-lab
npm init -y >/dev/null
npm i @sentry/node@11.5.0 --save-exact
Enter fullscreen mode Exit fullscreen mode

ingest.js. The fingerprint is the exception type plus the top in-app frame (file, function, line). The valve is a token bucket per fingerprint: 100 tokens, refilled at 100 per minute.

// ingest.js: fake error ingest with a per-fingerprint spike valve
const http = require("node:http");
const zlib = require("node:zlib");
const crypto = require("node:crypto");

const VALVE = process.env.VALVE !== "off";
const RATE = 100; // stored bodies per fingerprint per minute
const buckets = new Map();
const issues = new Map();

function allowStore(fp) {
  const now = Date.now();
  const b = buckets.get(fp) ?? { tokens: RATE, last: now };
  b.tokens = Math.min(RATE, b.tokens + ((now - b.last) / 60000) * RATE);
  b.last = now;
  buckets.set(fp, b);
  if (b.tokens >= 1) { b.tokens -= 1; return true; }
  return false;
}

function fingerprint(event) {
  const ex = event.exception?.values?.at(-1) ?? {};
  const frames = ex.stacktrace?.frames ?? [];
  const top = frames.filter((f) => f.in_app).at(-1) ?? frames.at(-1) ?? {};
  const key = `${ex.type}\0${top.filename}:${top.function}:${top.lineno}`;
  const fp = crypto.createHash("sha256").update(key).digest("hex");
  return { fp, title: `${ex.type}: ${ex.value}` };
}

http.createServer((req, res) => {
  if (req.url === "/stats") return res.end(JSON.stringify([...issues.values()]));
  const chunks = [];
  req.on("data", (c) => chunks.push(c));
  req.on("end", () => {
    let body = Buffer.concat(chunks);
    if (req.headers["content-encoding"] === "gzip") body = zlib.gunzipSync(body);
    const lines = body.toString("utf8").split("\n");
    for (let i = 1; i + 1 < lines.length; i += 2) {
      if (JSON.parse(lines[i]).type !== "event") continue;
      const { fp, title } = fingerprint(JSON.parse(lines[i + 1]));
      const issue = issues.get(fp) ?? { title, count: 0, stored: 0 };
      issue.count += 1;                                // every hit is counted
      if (!VALVE || allowStore(fp)) issue.stored += 1; // bodies are rationed
      issues.set(fp, issue);
    }
    res.statusCode = 202; // the client never sees the valve
    res.end("{}");
  });
}).listen(9000, () => console.log(`ingest on :9000 (valve ${VALVE ? "on" : "off"})`));
Enter fullscreen mode Exit fullscreen mode

flood.js. Three env switches: SAMPLE_RATE, DEDUPE=off and PACE=off.

// flood.js: one bug stuck in a retry loop, plus three rare bugs you care about
const Sentry = require("@sentry/node");

const SAMPLE_RATE = Number(process.env.SAMPLE_RATE ?? 1);
const DEDUPE = process.env.DEDUPE !== "off";
const PACE = process.env.PACE !== "off";

Sentry.init({
  dsn: "http://public@localhost:9000/1",
  sampleRate: SAMPLE_RATE,
  tracesSampleRate: 0,
  integrations: (defaults) =>
    DEDUPE ? defaults : defaults.filter((i) => i.name !== "Dedupe"),
});

function chargeCard() { throw new TypeError("Cannot read properties of undefined (reading 'id')"); }
function applyCoupon() { throw new RangeError("discount out of range: -5"); }
function sendReceipt() { throw new Error("SMTP 421 try again later"); }
function exportCsv() { throw new SyntaxError("Unexpected token < in JSON"); }

const rare = { 250: applyCoupon, 500: sendReceipt, 750: exportCsv };

(async () => {
  for (let i = 1; i <= 1000; i++) {
    try { (rare[i] ?? chargeCard)(); } catch (e) { Sentry.captureException(e); }
    if (PACE && i % 50 === 0) await Sentry.flush(2000);
  }
  await Sentry.flush(5000);
  const issues = await (await fetch("http://localhost:9000/stats")).json();
  console.log(`sampleRate=${SAMPLE_RATE} dedupe=${DEDUPE ? "on" : "off"} pace=${PACE ? "on" : "off"}`);
  console.log("issue".padEnd(36), "count".padStart(6), "stored".padStart(7));
  for (const { title, count, stored } of issues) {
    console.log(title.slice(0, 36).padEnd(36), String(count).padStart(6), String(stored).padStart(7));
  }
})();
Enter fullscreen mode Exit fullscreen mode

Run the ingest in one terminal and the flood in another. The ingest keeps state in memory, so restart it between runs.

# terminal 1
VALVE=off node ingest.js

# terminal 2
PACE=off node flood.js
Enter fullscreen mode Exit fullscreen mode

Run 1: the SDK drops more than you think

Before any sampling, the SDK already throws events away. Here's the naive run: defaults, no pacing, valve off.

sampleRate=1 dedupe=on pace=off
issue                                 count  stored
TypeError: Cannot read properties of      4       4
RangeError: discount out of range: -      1       1
Error: SMTP 421 try again later           1       1
SyntaxError: Unexpected token < in J      1       1
Enter fullscreen mode Exit fullscreen mode

997 TypeErrors thrown, 4 received. That's dedupeIntegration, on by default. It compares each event with the previous one and drops exact repeats, so you get one TypeError per stretch between the rare errors. With debug: true you'll see this 993 times:

Sentry Logger [warn]: Event dropped due to being a duplicate of previously captured event.
Enter fullscreen mode Exit fullscreen mode

Real floods come from many browsers or workers, each with its own "previous event". So I turned Dedupe off:

PACE=off DEDUPE=off node flood.js
Enter fullscreen mode Exit fullscreen mode
sampleRate=1 dedupe=off pace=off
issue                                 count  stored
TypeError: Cannot read properties of     64      64
Enter fullscreen mode Exit fullscreen mode

This is the run that bit me. 64 events arrived and all three rare bugs were gone. The transport allows 64 requests in flight (DEFAULT_TRANSPORT_BUFFER_SIZE = 64 in @sentry/core). A synchronous burst fills it and everything after is dropped. With debug: true, 936 times:

Sentry Logger [log]: Recording outcome: "queue_overflow:error"
Enter fullscreen mode Exit fullscreen mode

The rare bugs were events 250, 500 and 750. They never had a chance.

In the lab, pacing with await Sentry.flush() every 50 events fixes it, and so does transportOptions: { bufferSize: 2000 }. Both gave me 997 + 1 + 1 + 1. In a real incident, a bigger buffer just means you're queueing the flood in your app's memory.

Run 2: sampleRate 0.1 loses the rare bugs

Now the popular fix. Paced, Dedupe off, valve off, sampleRate at 0.1, five runs:

DEDUPE=off SAMPLE_RATE=0.1 node flood.js
Enter fullscreen mode Exit fullscreen mode
sampleRate=0.1 dedupe=off pace=on
issue                                 count  stored
TypeError: Cannot read properties of    120     120
SyntaxError: Unexpected token < in J      1       1
Enter fullscreen mode Exit fullscreen mode

That was one of my two lucky runs: one rare bug out of three. The other lucky run caught the SMTP error instead. The remaining three runs got 96, 102 and 85 TypeErrors and zero rare bugs. Yours will differ, it's random. That's the point.

Loop bug (thrown 997x) Rare bugs (3 per run)
Arrived over 5 runs 85 to 120 per run 2 of 15
What you conclude count is off by about 10x most of them never happened

The loop still owns the issue list. The bugs you'd want to look at are the ones that disappeared.

Run 3: count everything, ration the bodies

Same flood, no sampling, valve on.

# terminal 1
node ingest.js

# terminal 2
DEDUPE=off node flood.js
Enter fullscreen mode Exit fullscreen mode
sampleRate=1 dedupe=off pace=on
issue                                 count  stored
TypeError: Cannot read properties of    997     100
RangeError: discount out of range: -      1       1
Error: SMTP 421 try again later           1       1
SyntaxError: Unexpected token < in J      1       1
Enter fullscreen mode Exit fullscreen mode

Three runs, same output each time. The count is exact. The loop bug keeps 100 full stack traces and the rest are counted, not stored. Each rare bug gets its own bucket with 100 fresh tokens, so it always lands.

Every request got a 202, so the SDK has no reason to back off. All the throttling happens where you can see it: on the server, per bug.

How the spike valve works in Epure

Same idea, in Rust. Epure is the self-hosted error tracker I'm building. It takes errors from the official Sentry SDKs through a DSN swap and drops transactions and replays. Here's the core of the valve from crates/server/src/pipeline/spike.rs, trimmed:

const DEFAULT_RATE_PER_MINUTE: f64 = 100.0;

pub enum SpikeDecision { AllowStore, CounterOnly }

impl TokenBucket {
    fn check(&mut self) -> SpikeDecision {
        self.refill();
        if self.tokens >= 1.0 {
            self.tokens -= 1.0;
            SpikeDecision::AllowStore
        } else {
            SpikeDecision::CounterOnly
        }
    }
}
Enter fullscreen mode Exit fullscreen mode

AllowStore queues the full event. CounterOnly queues a job that only bumps the issue counter. Buckets sit in a DashMap keyed by fingerprint, and idle ones get evicted once there are 20k keys. The fingerprint is a SHA-256 of the exception type plus the top in-app frame, unless the SDK sends a custom fingerprint.

From the outside, per the errors and rate limits docs:

Mechanism Default What the client sees
Spike valve 100 raw events / minute / fingerprint 202. Extra bodies aren't stored, the counter still goes up.
Project ingest cap 5000 events / hour / project 403 ingest_cap_exceeded

No Redis. The buckets live in memory in the Rust process, next to Postgres 16, in a two-container compose file. The repo has an integration test that posts the same browser envelope 150 times and expects 150 202s, at least 150 counter increments and at most 100 stored rows.

Why not a project-level limit?

Project-level limits have the same blind spot as sampleRate. They protect the bill and can't tell which bug matters.

Sentry's SaaS has Spike Protection. Per its docs, it sets a per-project hourly threshold from your quota and the last 7 days of traffic, then discards events above it. That's a real safety net for your quota. But it limits the project, so a rare bug that lands mid-spike competes with the loop for the same allowance.

Epure doesn't dodge this either. Its hourly project cap is checked before the valve, and counter-only hits count toward it. Rough math: a loop sustained above about 83 events a minute (5000 / 60) burns the default cap within the hour. After that, every event from the project gets a 403 until the next clock hour, rare bugs included. You can raise the cap per project. The valve keeps a loop out of your database. It doesn't stop a sustained loop from eating the cap.

When is sampleRate still the right call?

When the upload itself costs you (a browser SDK on a lot of phones), or when your vendor bills per event and you need a ceiling today.

If one known error is the noise, an ignoreErrors entry or a beforeSend filter for that exact error beats a global coin flip. And don't mix up the knobs: tracesSampleRate is performance data, sampleRate is errors. Turning down traces is cheap. Turning down errors is how you lose the one-off bug.

Failure checklist

If your flood test doesn't match the outputs above:

  1. Only a handful of events arrived. Dedupe is on by default. Filter it out or interleave different errors.
  2. Exactly 64 events arrived. You hit the transport buffer. Pace with await Sentry.flush() or set transportOptions: { bufferSize }.
  3. Rare errors got merged into the loop issue. Your fingerprint is too coarse. Group by message only and every Cannot read properties of undefined shares one bucket.
  4. A deploy reset your throttling. Buckets are in memory, so a restart refills them. And since the frame key includes the line number, moving code gives the bug a new fingerprint.
  5. You run several ingest processes. In-memory buckets are per process. Two processes can store up to 2x the bodies per minute.
  6. Everything returns 403 ingest_cap_exceeded. That's the hourly project cap. The valve never returns an error.
  7. 202 but the issue list is empty. Accept is async, so the event may still be queued. Check the DSN's project and environment, then docker compose logs epure.
  8. Fewer events than expected, no errors anywhere. Someone left sampleRate below 1 in shared config: grep -rn "sampleRate" src/.

What would you set?

I picked 100 stored stack traces per bug per minute because that's about what I'd scroll through during an incident, and I'm not sure it's right. If a loop hit your app tonight, how many full stack traces of the same bug per minute would you actually want kept?

Top comments (0)