Two bugs my green test suite could not see
I published a library two days ago that strips PII and secrets out of an LLM response while it streams. Since then I have shipped fixes for two bugs in it, and neither was caught by the tests.
One of them I found before anyone could install it. The first version was on npm for four minutes with nothing downloaded, and I pulled it. The other was live for about a day.
The tests were not bad. There were 69 of them and they were green. They just could not see these two things, for two different reasons, and both reasons are easy to repeat.
What the library does
Filtering a finished response is easy. Filtering one that is still arriving is not.
const TEXT = 'Your card is 4111 1111 1111 1111, charge it.';
const chunks = [];
for (let i = 0; i < TEXT.length; i += 6) chunks.push(TEXT.slice(i, i + 6));
// ["Your c", "ard is", " 4111 ", "1111 1", "111 11", "11, ch", "arge i", "t."]
chunks.map(c => c.replace(/\b\d(?:[ -]?\d){12,18}\b/g, '[redacted]')).join('')
// → 'Your card is 4111 1111 1111 1111, charge it.'
The regex is correct. It just never sees a whole card, because no chunk contains one. (4111 1111 1111 1111 is the standard test card. It passes Luhn and belongs to nobody.)
Buffering the whole response fixes it and gives up streaming. The AI SDK docs recommend exactly that, and note that it delays output. If you want to keep streaming, the engine has to decide on every chunk how much it can safely release. Never act on a match that could still grow. Never emit text you have not scanned in final form.
That decision is where both bugs were.
Bug one: a policy that tested nothing
Here is the test I was pleased with.
const POLICIES = [
{},
{ action: 'report' },
{ mask: '#' },
{ bannedWords: ['forbidden', 'секрет', '機密'] },
{ normalize: false },
];
Fifteen inputs, five policies, eight chunk sizes, twelve randomised chunkings each. 1,500 assertions. Green.
Now look at the corpus those policies ran against. A card number. An API key. An email. Some emoji. Some CJK. A long run of as.
None of them contains forbidden, or секрет, or 機密.
That policy was asserted 300 times against text that could not trigger it. It checked that nothing happens when nothing happens. action: 'block' was not covered at all, and neither were mask functions, label detection, or PEM blocks.
I added ten inputs and three policies. Two failures came out.
The first was an IBAN.
iban GB82 WEST 1234 5698 7654 32
Filtered as one string it was masked. Streamed at chunk sizes 1, 2, 3, 5, 7 and 13 the account number went through in clear. At 64 it worked, because the whole thing landed in one chunk.
The walk that decides how far back a match might still grow gives up as soon as it meets a letter. There is a reason for that. From the tail, sk-abcdef123456 looks exactly like a digit group, and without the shortcut one long token drags the whole buffer into held text. But IBAN groups are alphanumeric. The walk hit WEST, decided this was not a digit group, and released the account number.
The input that exposed it was already in my benchmark corpus and had been passing since the first commit. The benchmark ran it as one string.
The second failure was in block mode. A batch call discarded the safe text before the detection. The streamed path had already emitted it. Two different outputs for the same input, from the library whose guarantee is that they agree.
This is the pair I got lucky on. Both were in the version that was up for four minutes, and both were fixed about three hours after I pulled it, so the first release anyone could actually install already had them fixed. I do not get to claim that as process. I strengthened the test on a hunch, the same night, before publishing again.
Bug two: two measurements that disagreed
The second one was live, and it was not found by a test at all.
I had measured how much text the engine holds back before releasing it, because that is the latency cost. I got a mean of 20 characters and a max of 58. Someone else ran their own harness on the published build and got a max of 76.
Close enough to shrug at. Same centre, different tail, different corpus, fine.
It was not fine. When I went looking for the difference, the 76 turned out to be the length of their test sample, not a tail. Here is what the published build did on Japanese:
Japanese reply, no digits, 66 chars → held 66, first output: only at flush
Japanese, 5 replies, 330 chars → held 330, first output: only at flush
Japanese, 40 replies, 2,640 chars → held 2,640, first output: only at flush
English control, 4,890 chars → held 33, first output: at char 22
The held text tracked the input length. It never converged, because it never settled at all.
The walk crosses any character above code 32, on the grounds that an unbroken run of them might be one long secret. A JWT runs to hundreds of characters, so this is not a silly rule. Latin prose is full of spaces, so it settles constantly and holds about twenty characters.
Japanese prose has no spaces. The whole paragraph is one token. Nothing settles.
So the library buffered the entire response and released it at the end, for Japanese, Chinese, Korean and Thai. Nothing leaked. Nothing was lost. It just was not streaming any more, for about a fifth of the people on earth, which is the only thing the library claims to do.
The fix is small. Every detector in it matches ASCII. An ideograph cannot be inside a card number or an API key or a JWT, so there was never anything to wait for. The walk now settles at one, the same way it settles at a space. Banned words can be non-ASCII, so crossing is still allowed but bounded by the longest non-ASCII banned word, and by nothing when there are none.
After the fix, on 0.7.3:
Japanese, 680 chars → held 0, first output at char 1
English, 900 chars → held 21, first output at char 20
Japanese now holds less than English does, which is correct. There is nothing in it to wait for.
The first version of that fix broke emoji. Settling at a non-ASCII character can land between the two halves of a surrogate pair, so an emoji arriving one UTF-16 unit at a time got half of itself released and the reader saw a replacement character. Two tests I had written two days earlier caught it. That is the only reason this section is not an apology.
What I took from it
Two things, and the second one surprised me more.
For a test matrix: check that every policy in it actually fires on something. The assertion count goes up whether or not it does, and nothing warns you. If you cannot point at the input that triggers a policy, the policy is decoration.
For measurements: when two of them disagree and the centres match, the disagreement is the finding. I nearly averaged those two numbers. If I had, the bug would still be there, and the README would carry a max that was never a max.
The same habit had left other things lying around. The README said the guarantee was enforced "in CI" when there was no CI. It said 50 tests when the runner printed 69. It advertised a Twilio detector that had been removed, and admitted the removal eighty-nine lines further down the same page. None of that was deliberate. All of it had been true at some point.
Numbers
- 74 tests
- 0 false positives on 39 adversarial clean samples
- 15 of 15 on the sensitive samples
- 47% recall on the gitleaks corpus, with 72 false positives out of 295
- 81%, 74% and 51% on Presidio, Gretel and ai4privacy
The first three are a gate I wrote against inputs I chose. The last two lines are 8,795 samples nobody here wrote, and they are the ones to judge it by. The 72 is not a flattering number and I am not going to explain it away.
Names, addresses and free-text disclosures are not attempted. This matches formats. It does not understand language. It is defence in depth, not a compliance control.
The library
llm-stream-guardrails. MIT, zero dependencies, TypeScript, ESM. Published through npm trusted publishing, so releases carry a provenance attestation you can check with npm audit signatures.
There is a demo that runs the compiled package in your browser. Nothing you type is sent anywhere. Drag the chunk size to one character and watch the output stay the same. Paste Japanese into it, too.
It is two days old, and the version history above is the whole of it. If you get a value past a detector, that is the most useful bug report the project can get, and there is a private channel for it.
One footnote
While building this I read the AI SDK middleware docs. The guardrail example implemented wrapGenerate and stopped, with a note that streaming guardrails are difficult because you do not know the full content until the stream ends. True, and it was the example people would copy.
I filed an issue with a reproduction. I did not write the fix. Vercel's ai-sdk-factory bot picked the issue up within the hour and opened three pull requests, and a maintainer merged them two days later into main, release-v6.0 and release-v5.0. The bot adds a Co-authored-by line naming whoever filed the issue, so the commit has my account on it. That is boilerplate rather than a maintainer crediting me.
The reproduction was the contribution. I checked it before posting it, which given the rest of this post is a slightly ironic thing to be proud of.
Top comments (2)
The surrogate-pair bug made me wonder how far the "settle at any character above code 32" rule generalizes. Surrogate pairs are the sharp case because releasing half of one produces an actual invalid UTF-16 sequence and a visible replacement character. But there's a softer version of the same failure with combining marks - releasing a base letter and holding its combining diacritic for the next chunk wouldn't corrupt anything, it'd just render the base letter alone for a frame and then have the accent appear a beat later. Not broken the way the emoji case was, but still not what "unbroken" is supposed to mean. Did the fix end up being specific to the surrogate-pair case, or does it hold onto trailing combining marks too?
The measurement-disagreement point is the one I'd flag hardest out of the whole post. Averaging two numbers that agree on the centre and disagree on the tail is exactly the move that erases the interesting case, and it's so instinctively "reasonable" that I don't think most people would even notice they'd done it.
Good question, and measuring it first changed my answer.
The short version: the filter does not split a grapheme cluster that arrived intact, and the fix was narrower than it should have been anyway. Both of those needed checking rather than reasoning about.
What I measured. Arabic harakat, Devanagari matras, Hebrew niqqud, Thai vowel signs, Latin combining marks, a ZWJ emoji sequence and a variation selector, at five chunk sizes. 0 of 40 combinations had the filter introduce a boundary that separated a character from a mark that followed it.
The distinction that matters, and I had it wrong in my head before running it: where a base and its mark end up in different chunks, that boundary came from whoever chunked the input. The filter passes it through unchanged. It usually coalesces instead — a 35-chunk Arabic input came out as 31, and a 21-chunk Latin one came out as 1.
What was actually wrong. The guard that stops a cut landing mid-grapheme enumerated U+0300–U+036F. That is Latin. Arabic, Hebrew, Devanagari, Thai and Cyrillic combining marks — 1,238 code points — were not recognised as marks at all, so a cut at one of them was unguarded. It tests the Unicode Mark property now, shipped in 0.7.4.
So: a real gap, at the cut points, and not the one your question was aimed at. I would not have found it by answering you from memory.
Your framing of the cost is right for Latin and understates it elsewhere. A bare vowel for a frame is cosmetic. A Devanagari or Thai matra is not decoration — the consonant alone renders as a different syllable until the rest arrives, so the intermediate frame is a different word rather than an unfinished one.
On the measurement-disagreement point: it happened again here, in a smaller way. My instinct was that the surrogate case was the sharp one and combining marks were the soft version of it, exactly as you framed it. That is true of the symptom and false of the code, where the guard was one enumerated range and everything outside it was unprotected equally. The instinct would have produced a confident and wrong reply.