<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: ROSH™ Company Labs</title>
    <description>The latest articles on DEV Community by ROSH™ Company Labs (@roshcompanylabs).</description>
    <link>https://dev.to/roshcompanylabs</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4138362%2Fe96cc646-4b3b-496b-a51e-6a3b4f3496b9.png</url>
      <title>DEV Community: ROSH™ Company Labs</title>
      <link>https://dev.to/roshcompanylabs</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/roshcompanylabs"/>
    <language>en</language>
    <item>
      <title>Two bugs my green test suite could not see</title>
      <dc:creator>ROSH™ Company Labs</dc:creator>
      <pubDate>Tue, 22 Sep 2026 23:41:38 +0000</pubDate>
      <link>https://dev.to/roshcompanylabs/two-bugs-my-green-test-suite-could-not-see-1eko</link>
      <guid>https://dev.to/roshcompanylabs/two-bugs-my-green-test-suite-could-not-see-1eko</guid>
      <description>&lt;h1&gt;
  
  
  Two bugs my green test suite could not see
&lt;/h1&gt;

&lt;p&gt;I published a library two days ago that strips PII and secrets out of an LLM response while it streams. Since then I have shipped fixes for two bugs in it, and neither was caught by the tests.&lt;/p&gt;

&lt;p&gt;One of them I found before anyone could install it. The first version was on npm for four minutes with nothing downloaded, and I pulled it. The other was live for about a day.&lt;/p&gt;

&lt;p&gt;The tests were not bad. There were 69 of them and they were green. They just could not see these two things, for two different reasons, and both reasons are easy to repeat.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the library does
&lt;/h2&gt;

&lt;p&gt;Filtering a finished response is easy. Filtering one that is still arriving is not.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;TEXT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Your card is 4111 1111 1111 1111, charge it.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="c1"&gt;// ["Your c", "ard is", " 4111 ", "1111 1", "111 11", "11, ch", "arge i", "t."]&lt;/span&gt;

&lt;span class="nx"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;c&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\b\d(?:[&lt;/span&gt;&lt;span class="sr"&gt; -&lt;/span&gt;&lt;span class="se"&gt;]?\d){12,18}\b&lt;/span&gt;&lt;span class="sr"&gt;/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;[redacted]&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;// → 'Your card is 4111 1111 1111 1111, charge it.'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The regex is correct. It just never sees a whole card, because no chunk contains one. (&lt;code&gt;4111 1111 1111 1111&lt;/code&gt; is the standard test card. It passes Luhn and belongs to nobody.)&lt;/p&gt;

&lt;p&gt;Buffering the whole response fixes it and gives up streaming. The AI SDK docs recommend exactly that, and note that it delays output. If you want to keep streaming, the engine has to decide on every chunk how much it can safely release. Never act on a match that could still grow. Never emit text you have not scanned in final form.&lt;/p&gt;

&lt;p&gt;That decision is where both bugs were.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug one: a policy that tested nothing
&lt;/h2&gt;

&lt;p&gt;Here is the test I was pleased with.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;POLICIES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="p"&gt;{},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;report&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;mask&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;#&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;bannedWords&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;forbidden&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;секрет&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;機密&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;normalize&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fifteen inputs, five policies, eight chunk sizes, twelve randomised chunkings each. 1,500 assertions. Green.&lt;/p&gt;

&lt;p&gt;Now look at the corpus those policies ran against. A card number. An API key. An email. Some emoji. Some CJK. A long run of &lt;code&gt;a&lt;/code&gt;s.&lt;/p&gt;

&lt;p&gt;None of them contains &lt;code&gt;forbidden&lt;/code&gt;, or &lt;code&gt;секрет&lt;/code&gt;, or &lt;code&gt;機密&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That policy was asserted 300 times against text that could not trigger it. It checked that nothing happens when nothing happens. &lt;code&gt;action: 'block'&lt;/code&gt; was not covered at all, and neither were mask functions, label detection, or PEM blocks.&lt;/p&gt;

&lt;p&gt;I added ten inputs and three policies. Two failures came out.&lt;/p&gt;

&lt;p&gt;The first was an IBAN.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;iban GB82 WEST 1234 5698 7654 32
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Filtered as one string it was masked. Streamed at chunk sizes 1, 2, 3, 5, 7 and 13 the account number went through in clear. At 64 it worked, because the whole thing landed in one chunk.&lt;/p&gt;

&lt;p&gt;The walk that decides how far back a match might still grow gives up as soon as it meets a letter. There is a reason for that. From the tail, &lt;code&gt;sk-abcdef123456&lt;/code&gt; looks exactly like a digit group, and without the shortcut one long token drags the whole buffer into held text. But IBAN groups are alphanumeric. The walk hit &lt;code&gt;WEST&lt;/code&gt;, decided this was not a digit group, and released the account number.&lt;/p&gt;

&lt;p&gt;The input that exposed it was already in my benchmark corpus and had been passing since the first commit. The benchmark ran it as one string.&lt;/p&gt;

&lt;p&gt;The second failure was in &lt;code&gt;block&lt;/code&gt; mode. A batch call discarded the safe text before the detection. The streamed path had already emitted it. Two different outputs for the same input, from the library whose guarantee is that they agree.&lt;/p&gt;

&lt;p&gt;This is the pair I got lucky on. Both were in the version that was up for four minutes, and both were fixed about three hours after I pulled it, so the first release anyone could actually install already had them fixed. I do not get to claim that as process. I strengthened the test on a hunch, the same night, before publishing again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug two: two measurements that disagreed
&lt;/h2&gt;

&lt;p&gt;The second one was live, and it was not found by a test at all.&lt;/p&gt;

&lt;p&gt;I had measured how much text the engine holds back before releasing it, because that is the latency cost. I got a mean of 20 characters and a max of 58. Someone else ran their own harness on the published build and got a max of 76.&lt;/p&gt;

&lt;p&gt;Close enough to shrug at. Same centre, different tail, different corpus, fine.&lt;/p&gt;

&lt;p&gt;It was not fine. When I went looking for the difference, the 76 turned out to be the length of their test sample, not a tail. Here is what the published build did on Japanese:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Japanese reply, no digits,    66 chars → held    66, first output: only at flush
Japanese, 5 replies,         330 chars → held   330, first output: only at flush
Japanese, 40 replies,      2,640 chars → held 2,640, first output: only at flush
English control,           4,890 chars → held    33, first output: at char 22
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The held text tracked the input length. It never converged, because it never settled at all.&lt;/p&gt;

&lt;p&gt;The walk crosses any character above code 32, on the grounds that an unbroken run of them might be one long secret. A JWT runs to hundreds of characters, so this is not a silly rule. Latin prose is full of spaces, so it settles constantly and holds about twenty characters.&lt;/p&gt;

&lt;p&gt;Japanese prose has no spaces. The whole paragraph is one token. Nothing settles.&lt;/p&gt;

&lt;p&gt;So the library buffered the entire response and released it at the end, for Japanese, Chinese, Korean and Thai. Nothing leaked. Nothing was lost. It just was not streaming any more, for about a fifth of the people on earth, which is the only thing the library claims to do.&lt;/p&gt;

&lt;p&gt;The fix is small. Every detector in it matches ASCII. An ideograph cannot be inside a card number or an API key or a JWT, so there was never anything to wait for. The walk now settles at one, the same way it settles at a space. Banned words can be non-ASCII, so crossing is still allowed but bounded by the longest non-ASCII banned word, and by nothing when there are none.&lt;/p&gt;

&lt;p&gt;After the fix, on 0.7.3:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Japanese, 680 chars → held 0, first output at char 1
English, 900 chars  → held 21, first output at char 20
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Japanese now holds less than English does, which is correct. There is nothing in it to wait for.&lt;/p&gt;

&lt;p&gt;The first version of that fix broke emoji. Settling at a non-ASCII character can land between the two halves of a surrogate pair, so an emoji arriving one UTF-16 unit at a time got half of itself released and the reader saw a replacement character. Two tests I had written two days earlier caught it. That is the only reason this section is not an apology.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I took from it
&lt;/h2&gt;

&lt;p&gt;Two things, and the second one surprised me more.&lt;/p&gt;

&lt;p&gt;For a test matrix: check that every policy in it actually fires on something. The assertion count goes up whether or not it does, and nothing warns you. If you cannot point at the input that triggers a policy, the policy is decoration.&lt;/p&gt;

&lt;p&gt;For measurements: when two of them disagree and the centres match, the disagreement is the finding. I nearly averaged those two numbers. If I had, the bug would still be there, and the README would carry a max that was never a max.&lt;/p&gt;

&lt;p&gt;The same habit had left other things lying around. The README said the guarantee was enforced "in CI" when there was no CI. It said 50 tests when the runner printed 69. It advertised a Twilio detector that had been removed, and admitted the removal eighty-nine lines further down the same page. None of that was deliberate. All of it had been true at some point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;74 tests&lt;/li&gt;
&lt;li&gt;0 false positives on 39 adversarial clean samples&lt;/li&gt;
&lt;li&gt;15 of 15 on the sensitive samples&lt;/li&gt;
&lt;li&gt;47% recall on the gitleaks corpus, with 72 false positives out of 295&lt;/li&gt;
&lt;li&gt;81%, 74% and 51% on Presidio, Gretel and ai4privacy&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first three are a gate I wrote against inputs I chose. The last two lines are 8,795 samples nobody here wrote, and they are the ones to judge it by. The 72 is not a flattering number and I am not going to explain it away.&lt;/p&gt;

&lt;p&gt;Names, addresses and free-text disclosures are not attempted. This matches formats. It does not understand language. It is defence in depth, not a compliance control.&lt;/p&gt;

&lt;h2&gt;
  
  
  The library
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.npmjs.com/package/llm-stream-guardrails" rel="noopener noreferrer"&gt;llm-stream-guardrails&lt;/a&gt;. MIT, zero dependencies, TypeScript, ESM. Published through npm trusted publishing, so releases carry a provenance attestation you can check with &lt;code&gt;npm audit signatures&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;There is a &lt;a href="https://roshcompanylabs.github.io/llm-stream-guardrails/" rel="noopener noreferrer"&gt;demo&lt;/a&gt; that runs the compiled package in your browser. Nothing you type is sent anywhere. Drag the chunk size to one character and watch the output stay the same. Paste Japanese into it, too.&lt;/p&gt;

&lt;p&gt;It is two days old, and the version history above is the whole of it. If you get a value past a detector, that is the most useful bug report the project can get, and there is a private channel for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  One footnote
&lt;/h2&gt;

&lt;p&gt;While building this I read the AI SDK middleware docs. The guardrail example implemented &lt;code&gt;wrapGenerate&lt;/code&gt; and stopped, with a note that streaming guardrails are difficult because you do not know the full content until the stream ends. True, and it was the example people would copy.&lt;/p&gt;

&lt;p&gt;I filed &lt;a href="https://github.com/vercel/ai/issues/21209" rel="noopener noreferrer"&gt;an issue&lt;/a&gt; with a reproduction. I did not write the fix. Vercel's &lt;code&gt;ai-sdk-factory&lt;/code&gt; bot picked the issue up within the hour and opened three pull requests, and a maintainer merged them two days later into &lt;code&gt;main&lt;/code&gt;, &lt;code&gt;release-v6.0&lt;/code&gt; and &lt;code&gt;release-v5.0&lt;/code&gt;. The bot adds a &lt;code&gt;Co-authored-by&lt;/code&gt; line naming whoever filed the issue, so &lt;a href="https://github.com/vercel/ai/commit/2801737b26" rel="noopener noreferrer"&gt;the commit&lt;/a&gt; has my account on it. That is boilerplate rather than a maintainer crediting me.&lt;/p&gt;

&lt;p&gt;The reproduction was the contribution. I checked it before posting it, which given the rest of this post is a slightly ironic thing to be proud of.&lt;/p&gt;

</description>
      <category>javascript</category>
      <category>typescript</category>
      <category>testing</category>
      <category>i18n</category>
    </item>
  </channel>
</rss>
