DEV Community

Cover image for I threw 1000+ LLM calls. Here is how VernLLM handled it.
LakBud
LakBud

Posted on

I threw 1000+ LLM calls. Here is how VernLLM handled it.

VernLLM is the LLM call framework for TypeScript. This video tests out its features, which includes adaptive AIMD rate-limiting, advanced circuit-breakers, multi provider fallbacks, developed retries with budget and observability. Open source with MIT license.

Website: https://vernllm.dev
GitHub: https://github.com/LakBud/vernLLM
npm: https://www.npmjs.com/package/vern-llm

Top comments (4)

Collapse
 
ahmetozel profile image
Ahmet Özel •

AIMD for rate limits is the right instinct, since provider limits are effectively unobservable from the client and probing beats guessing a static ceiling.

Two questions from the serving side. What does the circuit breaker do to in-flight streaming responses? Half-open transitions get messy when the unit of work is a token stream rather than a single request. And does the retry budget account for a fallback provider meaning a different tokenizer and a different output distribution, so the retry is not a like-for-like replay? That distinction is what makes fallback comfortable for classification and risky for anything with a strict output contract.

Collapse
 
lakbud profile image
LakBud • • Edited

Two great questions!

The breaker only checks things when a stream is about to start. A half open trial claims one slot before the stream opens, and that slot stays claimed for the entire stream, with no checks in between chunks. Whether the attempt counts as a success or failure only gets decided after the stream finishes and the response passes validation, not the moment data stops arriving. The one exception is an idle timeout in the middle of a stream, which counts as a failure right away. If it waited like everything else, a provider that sends one chunk and then hangs forever would never reach validation, so the call would just hang with nothing ever recorded.

The retry budget treats a retry against a fallback provider the same as a retry against the original one. It only tracks how many recent attempts were retries out of the total, with no idea about tokenizers or how different the output might be. This is intentional, since a retry against a fallback still costs capacity the same way a normal retry does. Whether a fallback's output is actually equivalent gets handled elsewhere, by a function called detectSoftFailure. It looks at a response that technically passed validation and can still mark it as a failure if it's wrong in a way validation cannot catch, and that decision is what causes a retry in the first place. The retry budget just counts the retry once it happens. It never looks at why.

Collapse
 
lakbud profile image
LakBud •

This test is using a mock AI provider with a 70% guaranteed failure rate. This is because using 1000+ LLM calls with an actual AI provider isnt needed to actually test out its features and its also expensive.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.