VernLLM just passed 10,000 downloads on npm. Thank you to everyone who tried it, filed an issue, or put it in front of real traffic.
VernLLM is the LLM call framework: resilience, observability, and control for every call. It runs inside your own process, so there is no gateway to deploy, no extra network hop, and no third party that sees your prompts. The core package has zero runtime dependencies and stays under 120 KB.
Today I'm releasing v3.0.
import Anthropic from '@anthropic-ai/sdk';
import OpenAI from 'openai';
import { VernLLM } from 'vern-llm';
import { fromAnthropic, fromOpenAI } from 'vern-llm/adapters';
const llm = new VernLLM({
client: fromOpenAI(new OpenAI({ apiKey: process.env.OPENAI_API_KEY, maxRetries: 0 })),
model: 'gpt-6-sol',
maxRetries: 3,
circuitBreaker: true,
retryBudget: { windowMs: 60_000, minCalls: 20, retryRatio: 0.2 },
rateLimit: { requestsPerMinute: 500, maxConcurrent: 20 },
fallback: {
name: 'anthropic',
client: fromAnthropic(new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY, maxRetries: 0 })),
model: 'claude-opus-5-5',
circuitBreaker: true,
},
});
What's new
Choose targets per call
A call, or a wrap middleware, can pick which configured targets to try and in what order. Core never chooses for you.
await llm.call({ userContent, targets: ['bedrock', 'primary'] });
const policy: VernLLMMiddleware = {
name: 'policy',
wrap: (request, next, ctx) =>
next({
targets: ctx.targets.filter((t) => t.adapter.provider === 'anthropic').map((t) => t.name),
}),
};
Policy only narrows. An inner middleware can never add back a target an outer one removed, so a residency or compliance rule can't be undone further down. fallbackOn now also receives the failed target and the next one.
A call context, and your own events
Pass plain JSON with a call and every hook, event and usage report carries it.
await llm.call({
userContent,
context: { tenantId: 't1', routing: { only: ['bedrock'] } },
});
Middleware can report its own decisions with ctx.emit, delivered as a custom event to onEvent and to OpenTelemetry.
wrap: async (request, next, ctx) => {
ctx.emit('router.decision', { deployment: 'claude', strategy: 'weighted' });
return next();
};
A cost tracker can now attribute spend to a tenant, and a router can explain itself in your traces. You can also seed typed middleware state per call with state and stateEntry.
Streaming that cleans up after itself
- Leaving a
for awaitearly (bybreak,returnor a throw) now cancels the provider stream and frees the rate limit slot. - A stream that fails before its first content chunk retries and falls back, even after a keep-alive ping opened it.
- New opt in
readerStallTimeoutMsdetaches a reader that stops reading.
Accurate usage
promptTokens now counts every input token the provider processed, cache reads and writes included. TokenUsage reports the split as cacheReadTokens, cacheWriteTokens and cacheWriteTokensByTtl. The rate limiter reconciles against the right number for each provider.
Tools and newer models
- Tool
argumentsare the output of yourargumentsSchema, so defaults and transforms apply. - With thinking on, Claude's reasoning comes back on
ToolCallResult.thinkingand can be replayed, so tool loops continue. - Newer Claude and GPT models get request fixes that used to fail at the provider, and now succeed or fail earlier with a clearer error.
Caching and rate limit fixes
-
NormalizedCacheAdapterno longer treats"2+2"and"2-2"as the same prompt. - The in memory cache copies values, and a missing or invalid
ttlstores nothing instead of an entry that never expires. - Invalid rate limits throw at construction, and images count as 1,600 tokens instead of their base64 length.
- Output cut off at
max_tokensbecomes a retryableresponse_truncatederror.
Breaking changes
v3.0 is a major release, so read the migration notes before upgrading. The short version:
-
Adapters moved. Import every
from*factory fromvern-llm/adapters. -
Gemini takes the top level
GoogleGenAIclient (fromGemini(ai), notai.models). -
Bedrock moved to its own package,
vern-llm-bedrock, so the core no longer loads the AWS SDK. -
promptTokensis higher on Anthropic calls that read the prompt cache. Code that added cache reads by hand now counts them twice. - Target names must be unique, and invalid rate limits throw instead of being ignored.
-
New error codes need a case in exhaustive
switchstatements overLLMErrorCode.
// Before
import { VernLLM, fromOpenAI } from 'vern-llm';
// After
import { VernLLM } from 'vern-llm';
import { fromOpenAI } from 'vern-llm/adapters';
Why this release
Most LLM failures aren't exotic. A provider slows down, returns a 429, or answers 200 with something unusable, and your app has to decide what to do next. A gateway can only decide from the HTTP it sees. VernLLM decides inside your process, where it knows the schema the answer must pass, the tenant it belongs to, and which providers it may use.
v3.0 makes that decision layer easier to build on. context, state, targets and ctx.emit are the foundations for routing, policy and cost tracking that live as typed code in your repo.
Get it
npm i vern-llm
- Docs: vernllm.dev
- npm: vern-llm
- GitHub: LakBud/vernLLM
If you upgrade, I'd love to hear how it goes. Open an issue if something in the migration is unclear, and a star helps others find it. Thanks again for 10k+ downloads.
Top comments (0)