If you run an LLM-backed endpoint (a support bot, a docs assistant, a "chat with your data" feature), you've probably noticed something in your logs: people ask the same questions over and over, just worded differently.
- "How do I reset my password?"
- "I forgot my password, what do I do?"
- "password reset steps"
A traditional cache keyed on the exact prompt string treats these as three unrelated requests, so you pay for three LLM calls, three times the latency, and three times the tokens.
A semantic cache fixes this. Instead of matching strings, it matches meaning: embed the incoming prompt, find the nearest previously answered prompt, and if it's close enough, return the stored answer without ever calling the LLM.
In this post we'll build one in Node.js and TypeScript with a two-tier design:
- L1: in-process memory. Microsecond lookups for exact repeats.
- L2: Postgres + pgvector. Shared across instances, durable, with similarity search via an HNSW index.
We'll also cover the part most tutorials skip: when a cache hit is actually a wrong answer, and how to design around it.
How a semantic cache works
Request ─▶ normalize ─▶ L1 exact lookup ──hit──▶ return
│ miss
▼
embed(prompt) ◀── ~50-300 ms, small cost
│
▼
L2 nearest neighbor (pgvector)
│
similarity ≥ threshold? ──yes──▶ return cached answer (+ promote to L1)
│ no
▼
call the LLM
│
▼
store in L1 + L2 (embedding + answer) ─▶ return
The economics work because embedding a prompt is much cheaper and faster than generating a completion. A miss costs you one extra embedding call plus a DB query. A hit saves a full generation. If your hit rate is decent, you come out well ahead on both cost and p50 latency.
Real numbers depend entirely on your traffic, so measure them (we'll add metrics later) rather than trusting anyone's blog post, including this one.
Step 1: Set up pgvector
Enable the extension and create the cache table:
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE llm_cache (
id BIGSERIAL PRIMARY KEY,
scope TEXT NOT NULL, -- tenant + model + prompt version (see below)
prompt_hash TEXT NOT NULL, -- hash of normalized prompt, for exact match
prompt TEXT NOT NULL, -- original text, useful for debugging and audits
embedding vector(1536) NOT NULL,
response JSONB NOT NULL,
hit_count INTEGER NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
expires_at TIMESTAMPTZ NOT NULL
);
-- Exact-match lookups
CREATE UNIQUE INDEX llm_cache_scope_hash_idx ON llm_cache (scope, prompt_hash);
-- Approximate nearest neighbor search using cosine distance
CREATE INDEX llm_cache_embedding_idx
ON llm_cache USING hnsw (embedding vector_cosine_ops);
-- For cleanup jobs
CREATE INDEX llm_cache_expires_idx ON llm_cache (expires_at);
Notes:
-
vector(1536)must match your embedding model's output dimension. Check your model's docs, and keep it consistent. Mixing models in one column breaks similarity. - HNSW gives fast approximate search and doesn't require training data first (unlike IVFFlat). It needs pgvector 0.5.0 or newer.
-
<=>is pgvector's cosine distance operator. Cosine similarity is1 - distance.
The scope column is the most important design decision
A cached answer is only valid for the context it was generated in. Put everything that changes the correct answer into scope:
function buildScope(opts: {
tenantId: string;
model: string;
promptVersion: string; // bump when you change your system prompt
toolsetVersion?: string;
}) {
return [opts.tenantId, opts.model, opts.promptVersion, opts.toolsetVersion ?? '-'].join('|');
}
Without this, tenant A's answer can be served to tenant B (a data leak), or a stale answer generated by an old system prompt keeps being served after you "fixed" the bot. Bumping promptVersion is your cheap, instant "invalidate everything" button.
Step 2: The cache class
Install dependencies:
npm i pg pgvector lru-cache
Here's the core implementation:
// semantic-cache.ts
import crypto from 'node:crypto';
import { Pool } from 'pg';
import pgvector from 'pgvector/pg';
import { LRUCache } from 'lru-cache';
export interface CachedResponse {
text: string;
model: string;
usage?: { inputTokens: number; outputTokens: number };
}
export interface CacheLookup {
hit: boolean;
tier?: 'l1' | 'l2';
similarity?: number;
response?: CachedResponse;
}
export class SemanticCache {
// L1: exact-match, in-process. Key = scope + prompt hash.
private l1 = new LRUCache<string, CachedResponse>({
max: 5_000,
ttl: 10 * 60 * 1000, // 10 minutes
});
// Single-flight: if two identical requests arrive at once, only one calls the LLM
private inflight = new Map<string, Promise<CachedResponse>>();
constructor(
private pool: Pool,
private embed: (text: string) => Promise<number[]>,
private opts = {
similarityThreshold: 0.92,
ttlSeconds: 24 * 60 * 60,
},
) {}
static normalize(prompt: string): string {
return prompt.trim().toLowerCase().replace(/\s+/g, ' ');
}
private hash(text: string) {
return crypto.createHash('sha256').update(text).digest('hex');
}
async lookup(scope: string, prompt: string): Promise<CacheLookup & { embedding?: number[] }> {
const normalized = SemanticCache.normalize(prompt);
const promptHash = this.hash(normalized);
const l1Key = `${scope}:${promptHash}`;
// --- L1: exact match in memory ---
const l1Hit = this.l1.get(l1Key);
if (l1Hit) return { hit: true, tier: 'l1', similarity: 1, response: l1Hit };
// --- L2a: exact match in Postgres (no embedding needed) ---
const exact = await this.pool.query(
`SELECT response FROM llm_cache
WHERE scope = $1 AND prompt_hash = $2 AND expires_at > now()`,
[scope, promptHash],
);
if (exact.rows[0]) {
this.l1.set(l1Key, exact.rows[0].response);
return { hit: true, tier: 'l2', similarity: 1, response: exact.rows[0].response };
}
// --- L2b: semantic match via pgvector ---
const embedding = await this.embed(normalized);
const { rows } = await this.pool.query(
`SELECT id, response, 1 - (embedding <=> $2) AS similarity
FROM llm_cache
WHERE scope = $1 AND expires_at > now()
ORDER BY embedding <=> $2
LIMIT 1`,
[scope, pgvector.toSql(embedding)],
);
const best = rows[0];
if (best && best.similarity >= this.opts.similarityThreshold) {
// Fire-and-forget stats update; don't block the response
this.pool
.query('UPDATE llm_cache SET hit_count = hit_count + 1 WHERE id = $1', [best.id])
.catch(() => {});
this.l1.set(l1Key, best.response); // promote so the exact phrasing is fast next time
return { hit: true, tier: 'l2', similarity: best.similarity, response: best.response, embedding };
}
return { hit: false, embedding };
}
async store(
scope: string,
prompt: string,
response: CachedResponse,
embedding?: number[],
): Promise<void> {
const normalized = SemanticCache.normalize(prompt);
const promptHash = this.hash(normalized);
this.l1.set(`${scope}:${promptHash}`, response);
const vec = embedding ?? (await this.embed(normalized));
await this.pool.query(
`INSERT INTO llm_cache (scope, prompt_hash, prompt, embedding, response, expires_at)
VALUES ($1, $2, $3, $4, $5, now() + ($6 || ' seconds')::interval)
ON CONFLICT (scope, prompt_hash)
DO UPDATE SET response = EXCLUDED.response,
embedding = EXCLUDED.embedding,
expires_at = EXCLUDED.expires_at`,
[scope, promptHash, prompt, pgvector.toSql(vec), response, String(this.opts.ttlSeconds)],
);
}
/** Wrap an LLM call with cache lookup, single-flight, and write-back. */
async getOrCompute(
scope: string,
prompt: string,
compute: () => Promise<CachedResponse>,
): Promise<CacheLookup & { response: CachedResponse }> {
const found = await this.lookup(scope, prompt);
if (found.hit) return found as CacheLookup & { response: CachedResponse };
const key = `${scope}:${this.hash(SemanticCache.normalize(prompt))}`;
let pending = this.inflight.get(key);
if (!pending) {
pending = compute()
.then(async (res) => {
// Don't let a cache write failure fail the user's request
await this.store(scope, prompt, res, found.embedding).catch((e) =>
console.error('cache store failed', e),
);
return res;
})
.finally(() => this.inflight.delete(key));
this.inflight.set(key, pending);
}
return { hit: false, response: await pending };
}
}
A few things worth calling out:
- Exact match before embedding. Many repeats are literally identical. Skipping the embedding call for those saves both time and money.
-
Single-flight. If 50 users ask the same new question at once (say, right after an outage), you want one LLM call, not 50. The
inflightmap handles that within a process. -
The cache must never break the request. Failures on store are logged and swallowed. You should also wrap
lookupso a database outage degrades to "just call the LLM" (see the Express example below). -
Parameterized queries throughout.
pgvector.toSql()serializes the array safely for thevectortype.
Step 3: Wire it into an endpoint
// server.ts
import express from 'express';
import { Pool } from 'pg';
import pgvector from 'pgvector/pg';
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
pool.on('connect', (client) => pgvector.registerTypes(client));
const cache = new SemanticCache(pool, embedText);
const app = express();
app.use(express.json());
app.post('/api/ask', async (req, res) => {
const { question } = req.body;
const tenantId = req.user.tenantId; // from your auth middleware
const scope = buildScope({
tenantId,
model: process.env.LLM_MODEL!,
promptVersion: 'support-bot-v7',
});
const started = performance.now();
try {
const result = await cache.getOrCompute(scope, question, () =>
callLlm({ system: SYSTEM_PROMPT, user: question }),
);
metrics.timing('ask.latency_ms', performance.now() - started, {
cache: result.hit ? result.tier! : 'miss',
});
metrics.increment('ask.cache', { result: result.hit ? 'hit' : 'miss' });
res.json({ answer: result.response.text, cached: result.hit });
} catch (err) {
// If the cache layer itself blows up, fall back to the LLM directly
console.error('cache path failed, calling LLM directly', err);
const direct = await callLlm({ system: SYSTEM_PROMPT, user: question });
res.json({ answer: direct.text, cached: false });
}
});
One caveat on that catch: it also catches errors from the LLM call itself, so you may double-call the LLM if the first call failed. In production, separate "cache infrastructure errors" from "compute errors" (for example by having lookup return a miss on DB failure internally).
The hard part: choosing a similarity threshold
The threshold is the knob that trades hit rate against correctness, and there's no universal right value. It depends on your embedding model, your domain, and how costly a wrong answer is.
- Too low: you serve wrong answers confidently.
- Too high: you rarely hit, and the cache is expensive overhead.
Embedding similarity measures topical closeness, not intent equivalence. These pairs are all highly similar to an embedding model but need different answers:
| Prompt A | Prompt B | Same answer OK? |
|---|---|---|
| "How do I cancel my subscription?" | "How do I not cancel my subscription?" | ❌ |
| "What's the refund policy for annual plans?" | "What's the refund policy for monthly plans?" | ❌ |
| "Is feature X available in Python?" | "Is feature X available in Go?" | ❌ |
| "How do I reset my password?" | "I forgot my password" | ✅ |
Calibrate with your own data
Don't guess. Build a small labeled set from real logs: pairs of prompts marked "same intent" or "different intent". Then sweep thresholds:
// calibrate.ts
interface Pair { a: string; b: string; sameIntent: boolean }
function cosine(a: number[], b: number[]) {
let dot = 0, na = 0, nb = 0;
for (let i = 0; i < a.length; i++) {
dot += a[i] * b[i]; na += a[i] ** 2; nb += b[i] ** 2;
}
return dot / (Math.sqrt(na) * Math.sqrt(nb));
}
async function sweep(pairs: Pair[]) {
const scored = [];
for (const p of pairs) {
const [ea, eb] = await Promise.all([embedText(p.a), embedText(p.b)]);
scored.push({ sim: cosine(ea, eb), sameIntent: p.sameIntent });
}
for (let t = 0.80; t <= 0.99; t += 0.01) {
const hits = scored.filter((s) => s.sim >= t);
const falsePositives = hits.filter((s) => !s.sameIntent).length;
const truePositives = hits.length - falsePositives;
console.log(
`threshold=${t.toFixed(2)} hits=${hits.length} ` +
`correct=${truePositives} WRONG=${falsePositives}`,
);
}
}
Choose the lowest threshold where WRONG is acceptably close to zero for your risk tolerance. For a customer-facing bot that quotes prices or policies, be very conservative. For a casual "explain this concept" endpoint, you can be looser.
Defense in depth beyond the threshold
- Segment by scope. Different products, plans, or languages shouldn't share cache entries.
-
Keep entities out of the fuzzy match. If prompts contain order IDs, names, or numbers, either extract them into the
scope/key or skip semantic matching when they're present, because the nearest neighbor of "status of ORD-123456" is "status of ORD-654321". - Per-category thresholds. Stricter for transactional or factual queries, looser for general explanations.
- Log near-misses. Record hits between 0.88 and 0.95 along with the matched prompt for periodic human review. It's the best source of calibration data you'll get.
What you should not cache
A cache hit means "the answer from last time is still correct". That's false for many requests:
- Personalized answers ("what's my balance?"). Different users, different truths.
- Time-sensitive answers ("what's the status of my order?", "what's today's date?").
- Responses that used tools with live data. Caching the final text freezes that data in place.
- Non-deterministic or creative requests where the user expects variety ("give me another idea").
- Multi-turn context. If the answer depends on earlier messages, the last user message alone is not the cache key. Either only cache stateless first turns, or embed a summary of the conversation context.
- Anything with sensitive data, unless you've thought through retention, access control, and your compliance obligations. Cached prompts and answers are stored data.
Implement this as an explicit policy function rather than scattered ifs:
function isCacheable(req: AskRequest): boolean {
if (req.conversation.length > 1) return false; // multi-turn
if (req.usedTools) return false; // live data involved
if (req.user.requiresPersonalization) return false;
if (/\b(today|now|current|latest|my)\b/i.test(req.question)) return false; // crude, tune it
return true;
}
Performance and operations
HNSW tuning
At query time, hnsw.ef_search controls the recall/speed tradeoff (the default is 40; higher means better recall but slower):
SET hnsw.ef_search = 100;
For a cache you can usually tolerate slightly imperfect recall, since a missed neighbor just means one extra LLM call.
Filtering with HNSW
Our query filters by scope. With HNSW, filtering is applied after the index scan, so a highly selective filter can return fewer rows than you asked for (or none, even though matches exist). Options:
- Use pgvector's iterative index scans (available in newer versions, via
hnsw.iterative_scan) - Use partial indexes or table partitioning per large tenant
- Keep scopes coarse enough that most scans find candidates
Check the pgvector README for the options available in your installed version.
Expiry and cleanup
expires_at filtering in queries keeps stale entries from being served, but they still occupy space and index memory. Run a periodic cleanup:
DELETE FROM llm_cache WHERE expires_at < now() - interval '1 day';
Run it in batches, or use partitioning by time if the table is large. Consider also evicting entries with zero hits after N days. They're not earning their keep.
Embedding latency is on your critical path
Every cache miss pays for an embedding call and the LLM call. Mitigations:
- Skip semantic lookup for very short or very long prompts (tune the bounds)
- Set a tight timeout on the embedding call. If it's slow, treat it as a miss and go straight to the LLM
- Reuse the embedding computed during lookup when storing (we do this via the returned
embedding)
Streaming responses
If you stream tokens to users, you can't cache until the stream completes. Buffer the full text as it streams, then store it afterwards. On a hit, you can either return the whole answer instantly or simulate streaming by chunking the cached text, which keeps your frontend code identical for both paths.
A note on "in-memory" scaling
The L1 cache is per-process. With many instances, each warms up independently, and invalidation (via promptVersion in the scope) works because new scopes simply never match old keys. If you need cross-instance L1 coherence, that's a sign to add Redis as a shared tier, but start without it. L1 + pgvector covers most workloads, and every extra tier is another thing to keep consistent.
Metrics that tell you if it's working
Instrument from day one:
-
Hit rate by tier (
l1,l2,miss) and by scope or category - Latency by outcome (hit vs miss). Compare p50 and p95.
- Embedding latency and error rate
- Similarity distribution of hits. If most hits cluster right at your threshold, you're likely over-matching.
-
Estimated tokens saved, from the
usagestored on each cached response -
User feedback signals (thumbs down, re-asks) segmented by
cached: true/false. If cached answers get disproportionately downvoted, raise the threshold.
That last one is the real test. A cache with a great hit rate and worse user satisfaction is a regression.
Putting it together
┌───────────────────────────────┐
Request ──▶ │ isCacheable? ── no ─────────┼──▶ LLM
└──────────────┬────────────────┘
│ yes
┌──────────▼──────────┐
│ L1: in-process LRU │── hit ──▶ response (µs)
└──────────┬──────────┘
│ miss
┌──────────▼──────────┐
│ L2: exact (Postgres)│── hit ──▶ response (ms)
└──────────┬──────────┘
│ miss
┌──────────▼──────────┐
│ embed + pgvector kNN│── ≥ threshold ──▶ response
└──────────┬──────────┘
│ miss
┌──────────▼──────────┐
│ single-flight → LLM │──▶ store L1 + L2 ──▶ response
└─────────────────────┘
Key takeaways
- A semantic cache matches meaning, not strings. Embed the prompt, find the nearest cached prompt, and reuse the answer above a similarity threshold.
- Two tiers are enough. An in-process LRU for exact repeats and Postgres with pgvector (HNSW) for shared, durable semantic matches.
-
Put everything that affects correctness into the
scope: tenant, model, prompt version. This prevents data leaks and gives you instant invalidation. - Calibrate the threshold with labeled data from your own traffic. Similar embeddings don't guarantee equivalent intent.
- Be deliberate about what's cacheable. Personalized, time-sensitive, tool-backed, and multi-turn requests usually aren't.
- Make the cache fail-open. It should only ever make your endpoint faster, never less available.
- Measure user-facing quality, not just hit rate.
Have you built a semantic cache, or tried one and abandoned it? I'd love to hear what threshold and embedding model worked for you in the comments.
Further reading: the pgvector README (HNSW options, iterative scans, distance operators), pgvector-node, and lru-cache.
Top comments (0)