DEV Community

Cover image for Thanks for 10k+ downloads. VernLLM v3.0 is out!
LakBud
LakBud

Posted on

Thanks for 10k+ downloads. VernLLM v3.0 is out!

VernLLM just passed 10,000 downloads on npm. Thank you to everyone who tried it, filed an issue, or put it in front of real traffic.

VernLLM is the LLM call framework: resilience, observability, and control for every call. It runs inside your own process, so there is no gateway to deploy, no extra network hop, and no third party that sees your prompts. The core package has zero runtime dependencies and stays under 120 KB.

Today I'm releasing v3.0.

import Anthropic from '@anthropic-ai/sdk';
import OpenAI from 'openai';
import { VernLLM } from 'vern-llm';
import { fromAnthropic, fromOpenAI } from 'vern-llm/adapters';

const llm = new VernLLM({
  client: fromOpenAI(new OpenAI({ apiKey: process.env.OPENAI_API_KEY, maxRetries: 0 })),
  model: 'gpt-6-sol',
  maxRetries: 3,
  circuitBreaker: true,
  retryBudget: { windowMs: 60_000, minCalls: 20, retryRatio: 0.2 },
  rateLimit: { requestsPerMinute: 500, maxConcurrent: 20 },
  fallback: {
    name: 'anthropic',
    client: fromAnthropic(new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY, maxRetries: 0 })),
    model: 'claude-opus-5-5',
    circuitBreaker: true,
  },
});
Enter fullscreen mode Exit fullscreen mode

What's new

Choose targets per call

A call, or a wrap middleware, can pick which configured targets to try and in what order. Core never chooses for you.

await llm.call({ userContent, targets: ['bedrock', 'primary'] });

const policy: VernLLMMiddleware = {
  name: 'policy',
  wrap: (request, next, ctx) =>
    next({
      targets: ctx.targets.filter((t) => t.adapter.provider === 'anthropic').map((t) => t.name),
    }),
};
Enter fullscreen mode Exit fullscreen mode

Policy only narrows. An inner middleware can never add back a target an outer one removed, so a residency or compliance rule can't be undone further down. fallbackOn now also receives the failed target and the next one.

A call context, and your own events

Pass plain JSON with a call and every hook, event and usage report carries it.

await llm.call({
  userContent,
  context: { tenantId: 't1', routing: { only: ['bedrock'] } },
});
Enter fullscreen mode Exit fullscreen mode

Middleware can report its own decisions with ctx.emit, delivered as a custom event to onEvent and to OpenTelemetry.

wrap: async (request, next, ctx) => {
  ctx.emit('router.decision', { deployment: 'claude', strategy: 'weighted' });
  return next();
};
Enter fullscreen mode Exit fullscreen mode

A cost tracker can now attribute spend to a tenant, and a router can explain itself in your traces. You can also seed typed middleware state per call with state and stateEntry.

Streaming that cleans up after itself

  • Leaving a for await early (by break, return or a throw) now cancels the provider stream and frees the rate limit slot.
  • A stream that fails before its first content chunk retries and falls back, even after a keep-alive ping opened it.
  • New opt in readerStallTimeoutMs detaches a reader that stops reading.

Accurate usage

promptTokens now counts every input token the provider processed, cache reads and writes included. TokenUsage reports the split as cacheReadTokens, cacheWriteTokens and cacheWriteTokensByTtl. The rate limiter reconciles against the right number for each provider.

Tools and newer models

  • Tool arguments are the output of your argumentsSchema, so defaults and transforms apply.
  • With thinking on, Claude's reasoning comes back on ToolCallResult.thinking and can be replayed, so tool loops continue.
  • Newer Claude and GPT models get request fixes that used to fail at the provider, and now succeed or fail earlier with a clearer error.

Caching and rate limit fixes

  • NormalizedCacheAdapter no longer treats "2+2" and "2-2" as the same prompt.
  • The in memory cache copies values, and a missing or invalid ttl stores nothing instead of an entry that never expires.
  • Invalid rate limits throw at construction, and images count as 1,600 tokens instead of their base64 length.
  • Output cut off at max_tokens becomes a retryable response_truncated error.

Breaking changes

v3.0 is a major release, so read the migration notes before upgrading. The short version:

  • Adapters moved. Import every from* factory from vern-llm/adapters.
  • Gemini takes the top level GoogleGenAI client (fromGemini(ai), not ai.models).
  • Bedrock moved to its own package, vern-llm-bedrock, so the core no longer loads the AWS SDK.
  • promptTokens is higher on Anthropic calls that read the prompt cache. Code that added cache reads by hand now counts them twice.
  • Target names must be unique, and invalid rate limits throw instead of being ignored.
  • New error codes need a case in exhaustive switch statements over LLMErrorCode.
// Before
import { VernLLM, fromOpenAI } from 'vern-llm';

// After
import { VernLLM } from 'vern-llm';
import { fromOpenAI } from 'vern-llm/adapters';
Enter fullscreen mode Exit fullscreen mode

Why this release

Most LLM failures aren't exotic. A provider slows down, returns a 429, or answers 200 with something unusable, and your app has to decide what to do next. A gateway can only decide from the HTTP it sees. VernLLM decides inside your process, where it knows the schema the answer must pass, the tenant it belongs to, and which providers it may use.

v3.0 makes that decision layer easier to build on. context, state, targets and ctx.emit are the foundations for routing, policy and cost tracking that live as typed code in your repo.

Get it

npm i vern-llm
Enter fullscreen mode Exit fullscreen mode

If you upgrade, I'd love to hear how it goes. Open an issue if something in the migration is unclear, and a star helps others find it. Thanks again for 10k+ downloads.

Top comments (0)