DEV Community

Yatin Davra
Yatin Davra

Posted on

The Prompt Wasn't the Whole Cache Key

An LLM response cache sounds simple: same model, same prompt, return the saved result.

But with structured output, the prompt is only part of the request. The schema is part of the promise too.

import { z } from "zod";
import { createClient, responseCacheMiddleware } from "@aviasole/shapecraft";

const Profile = z.object({ name: z.string() });
const AshaOnlyProfile = Profile.refine((profile) => profile.name === "Asha");

const client = createClient({
  middleware: [responseCacheMiddleware()],
});

await client.generate(model, Profile, "Extract the person's name from this note");
await client.generate(model, AshaOnlyProfile, "Extract the person's name from this note");
Enter fullscreen mode Exit fullscreen mode

The two schemas look almost the same as JSON, but the second has an extra rule: the name must be Asha. If the first call returned Mira, reusing that result for the second call would skip the new check.

That's the detail that makes caching structured output different from caching plain text. A cache hit short-circuits the generation pipeline, including validation. The key has to describe the rules that made the original result acceptable.

There are a few easy ways to get this wrong:

  • Two model objects can share an ID while using different configuration.
  • Zod refinements may not appear in the JSON Schema representation.
  • A validation callback can change between calls, even when the prompt stays the same.

Shapecraft's responseCacheMiddleware() keeps entries separate by model and schema. Zod schemas are keyed by their instance, so the refined schema above cannot reuse the first result. Calls with validation callbacks, abort signals, or options that cannot be safely keyed bypass the cache.

const client = createClient({
  middleware: [responseCacheMiddleware({ ttlMs: 60_000 })],
});

const result = await client.generate(model, Profile, prompt);
Enter fullscreen mode Exit fullscreen mode

That safety has a tradeoff: two separately created but equivalent Zod schemas won't share a cache entry. I'd rather miss a cache hit than skip a rule the caller asked us to enforce.

The cache is in memory and has a time-to-live, but no size limit or eviction policy. It's a good fit for repeated calls over a bounded set of prompts. For high-cardinality traffic, keep an eye on memory or put an eviction layer around it.

A cache should save model calls without changing what counts as a valid result. For structured generation, that means the schema belongs in the key too.

Top comments (0)