DEV Community

ROSH™ Company Labs
ROSH™ Company Labs

Posted on

I cut six streaming SDKs at every point. Nothing I found was new.

The last thing I wrote here was about two bugs my own test suite could not see. Both came from the same place: a stream that stops in the middle, and a consumer that cannot tell.

So I pointed the same method at other people's code. Six streaming SDKs, one question, asked mechanically rather than cleverly. This is what came back, including the parts where I was the one who was wrong.

The honest headline first: of the two defects it surfaced, both had already been reported by someone else. That is not a disappointing result. It is the result I would expect from a method that works, and I will come back to why at the end.

The method, in one paragraph

Take a stream of events. Replay it through the client once whole — that run is the control. Then replay every prefix of it: the first event alone, the first two, and so on. After each truncated run, ask two questions. Could the consumer tell the run was cut? And does what it was shown still agree with what was saved? A defect is a cut where the client reports the same thing it reports for a complete run, and something was actually lost.

That is the whole idea. It is not clever. Its only real property is that it does not get bored and does not have a hunch about where to look.

Four were clean, and that is the point

Before any finding is worth reading, the harness has to be capable of saying "nothing here".

  • Vercel AI SDK (ai@7.0.118) — a cut that got text through reports finishReason: other; a cut that got none never calls onFinish at all and raises AI_NoOutputGeneratedError instead. Two different shapes, both detectable.
  • LangGraph JS (@langchain/langgraph@1.4.18) — on abort the checkpoint stops exactly where the consumer stopped, next names the node to resume from, and the call raises. Three independent signals, all correct.
  • Mastra (@mastra/core@1.71.0) — every cut that showed the consumer text still persisted an assistant message.
  • OpenAI Node (openai@7.23.0) — after the cut, finalResponse() still carries the function_call the consumer watched arrive, and status never claims completed.

Four out of six. A harness that finds a defect everywhere it looks is measuring itself.

A complete stream that delivers nothing

This one is not about truncation at all, which is why I nearly missed it.

The SSE grammar in the HTML standard allows three line terminators:

end-of-line = ( cr lf / cr / lf )
Enter fullscreen mode Exit fullscreen mode

CRLF, bare CR and bare LF are the same boundary to a conformant reader. So I held the events fixed and varied only that, over real HTTP, with eventsource-parser alongside as a second implementation of the same grammar handed the same bytes:

  terminator                  conformant  @ag-ui/client  RUN_FINISHED  threw
  LF          \n\n            5           5              true          —
  CRLF        \r\n\r\n        5           0              false         Unexpected non-whitespace character after JSON
  CR          \r\r            5           0              false         Unexpected non-whitespace character after JSON
  CRLF field, LF blank        5           5              true          —
Enter fullscreen mode Exit fullscreen mode

Nothing is cut here. These are whole runs. Under two of the three terminators the grammar allows, the client delivers no events at all and the awaited call rejects.

The cause is one line:

const events = buffer.split(/\n\n/);
Enter fullscreen mode Exit fullscreen mode

\r\n\r\n contains no \n\n, so nothing ever splits. The whole response accumulates, and at EOF the flush parses every frame concatenated and throws. The docstring above that line says the parser "strictly follows the SSE standard".

The shape of the failure is worth more than the fact of it:

  terminator                  k=1  k=2  k=3  k=4  k=5
  LF          \n\n            1    2    3    4    5
  CRLF        \r\n\r\n        1    0    0    0    0
  CR          \r\r            1    0    0    0    0
  CRLF field, LF blank        1    2    3    4    5
Enter fullscreen mode Exit fullscreen mode

It does not degrade. It falls off a cliff between one event and two, because with one event there is no boundary to miss and the trailing CR is swallowed as JSON whitespace. A smoke test that sends a single frame passes.

And it was already reported — for the project's Rust SDK, in an issue that had been open six weeks with no replies. The premise was on the table. Nobody had checked the reference implementation against it.

Two helpers in one package that disagree about a valid stream

Same project, different question.

compactEvents() buffers a tool-call or message group until its *_END arrives, while terminal events go straight out. So a group still open when the run ends gets flushed after the terminal. verifyEvents(), from the same package, rejects almost everything once a run has closed.

There is an invariant those two owe each other: if verifyEvents accepts a stream, it must accept compactEvents(stream). Aborting one ordinary run at each point in turn, with the RUN_ERROR an abort actually emits:

  k   groups open  raw stream  compacted
  1   0            accepted    accepted
  2   1            accepted    REJECTED
  3   1            accepted    REJECTED
  4   1            accepted    REJECTED
  5   0            accepted    accepted
  6   1            accepted    REJECTED
  7   1            accepted    REJECTED
  8   0            accepted    accepted
  9   0            accepted    accepted
  10  1            accepted    REJECTED
  11  1            accepted    REJECTED
Enter fullscreen mode Exit fullscreen mode

That is the probe's own output with two columns dropped for width — it also prints whether compaction reordered anything, and what verifyEvents said. Worth one glance: the rows where nothing is open are reordered too, and still accepted. Reordering is not the defect. Reordering past a terminal event is.

Seven of eleven. The raw stream is valid at all eleven, so compacting is the only thing that breaks it, and the seven are exactly the seven where a group was open — the same set, no exceptions in either direction.

Seven of eleven is not a rate. It is a property of the run I happened to cut; one with groups open for longer breaks at more points. The rule is what carries, and the original report already states the rule. What a sweep adds is that it holds without exception.

The user-visible version of this is a chat thread whose transcript never renders again, for anyone, until the server restarts.

Verifying someone else's fix

There is an open fix for a third one — a request whose SSE response leg dies, leaving the caller waiting out its full timeout for a response that is never coming. Eighteen lines.

Running the harness against the published client and then against the build the PR publishes:

  published    1..3  REQUEST_TIMEOUT    ~403ms
  PR build     1..3  CONNECTION_CLOSED  1-3ms
  both         4     resolved           fast
Enter fullscreen mode Exit fullscreen mode

It holds at every cut point, including the one where it must not fire: the response arrived and the server then closed the leg as it always does.

One thing fell out that is in neither the issue nor the PR description. The SDK documents an onRequestStreamEnd callback a caller can pass. On the published release it never fires — not on any cut, not on success — because the protocol layer never forwarded the option. The same eighteen lines make a documented public option work for the first time.

I would not have looked at that column if the harness had not printed it next to everything else.

The part I would actually keep: where I was wrong

None of the above is the useful half of this post.

I measured a nine-month-old version and nearly filed the result. npm install @mastra/core gave me 0.24.9 while the current release was 1.71.0 — 882 releases and nine months apart. npm resolves the newest version whose engines your Node satisfies and warns rather than failing, and I was on Node 20 while the current release wanted 22.13. The warning scrolled past in the install output. I had a finding written up and two edits from filing it, against behaviour that had been fixed months before I "found" it.

Every probe now refuses to run against anything but the current release. It runs before the measurement, not after, because a measurement of the wrong version is not a weaker result — it is not a result, and the version line is the first thing a maintainer checks.

I built a differential that was not a differential. Comparing the PR build against the published client looked clean and the numbers agreed beautifully. Then I checked what version the PR build reported: an older one than the release I was comparing it to. Any difference could have come from a release of unrelated changes. Numbers that agree are not evidence that the comparison was sound.

My own checkers were wrong about as often as the code I was checking. A leak scanner whose pattern for a Windows drive path also matched the phrase "install:\n". A count printed as NaN because the helper returned a generator and I read .length off it — harmless, except a NaN next to a passing assertion is exactly the shape of a test that silently ran nothing. A suggested fix I nearly published that missed one of five legal cases.

Each of those was caught by the same move: check the checker before believing the check. Run it against something known-broken and confirm it screams.

And the one I keep relearning: my instinct about where a defect lives is confidently wrong often enough that it is not worth consulting. In the CRLF case I assumed the failure was in flushing at EOF, because that was the bug I had just read about somewhere else. The EOF flush turned out to be correct. The defect was one line up, in what counts as a boundary. I would have written a confident and wrong report.

Why "nothing was new" is the result I wanted

Both defects the sweep surfaced had already been filed by other people. One was reported six weeks earlier for a different language binding; one was reported six days earlier, with a runnable reproduction and a companion issue in the framework where users were hitting it.

If a mechanical sweep across six mature SDKs had turned up six novel defects, the right conclusion would not be that those SDKs are bad. It would be that my harness has a bug and is reporting its own assumptions back to me. Agreeing with what careful people already found is the strongest evidence I have that the thing measures what I think it measures.

What it adds is not discovery. It is the shape: which cut points, how many, whether the rule someone stated by hand holds everywhere, and whether the fix they wrote holds at points they did not test.

Run it

Each probe is one file. Everything specific to the SDK lives in one replay function; the loop, the comparison and the report come from the library.

npm install @ag-ui/client eventsource-parser
node probes/ag-ui-line-endings.mjs
Enter fullscreen mode Exit fullscreen mode

https://github.com/roshcompanylabs/cut-at-k

To point it at something else, copy the shortest probe and change replay and the lostContent predicate. Nothing else.

If you run it against something I have not and it comes back clean, that is a result worth having too — most of them are.

Top comments (1)