A semantic cache turns each incoming prompt into an embedding and searches pgvector for an earlier prompt that means the same thing. If the similarity is above a threshold you set, it returns the stored answer and skips the model call. Done well, it makes repeat questions faster and cheaper to answer. Done carelessly, it shows one tenant's answer to another tenant, or answers a question about invoice 1042 with the reply meant for invoice 1043.
This post builds the cache in Node.js on Postgres with pgvector. It covers the table, cache keys scoped to each tenant, TTLs, and how to tune the threshold so that a cache hit really is the right answer.
When a semantic cache is worth it
Exact-match caching (hash the prompt, look up the hash) only helps when users type the same words. They rarely do. "How do I reset my password?" and "forgot my password, what now" should return the same answer.
A semantic cache is a good fit when:
- Many users ask a small set of questions, as with support bots, docs Q&A and internal help desks.
- The answer doesn't depend on who is asking, or that dependency is part of the cache key.
- An answer that is a few hours old is still correct.
Skip it when each response depends on live, per-user data such as order status, or when you want the output to vary, as in brainstorming tools.
The table
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE llm_cache (
id bigserial PRIMARY KEY,
tenant_id text NOT NULL,
scope text NOT NULL,
model text NOT NULL,
prompt_text text NOT NULL,
embedding vector(1536) NOT NULL,
answer text NOT NULL,
hit_count integer NOT NULL DEFAULT 0,
created_at timestamptz NOT NULL DEFAULT now(),
expires_at timestamptz NOT NULL
);
CREATE INDEX llm_cache_embedding_idx
ON llm_cache USING hnsw (embedding vector_cosine_ops);
CREATE INDEX llm_cache_lookup_idx
ON llm_cache (tenant_id, scope, model, expires_at);
Set the vector dimension to match your embedding model. Vectors from different embedding models can't be compared, so if you change models you need a fresh table.
Multi-tenant data has one pgvector catch. An HNSW index finds the nearest neighbours first and applies your WHERE filter afterwards. With many tenants, all the top candidates can belong to other tenants, and the query returns nothing. On pgvector 0.8 or later, turn on iterative scans so the index keeps searching:
SET hnsw.iterative_scan = relaxed_order;
What goes into the cache key
The embedding finds similar prompts. The cache key controls which prompts are allowed to match in the first place. Anything that changes the correct answer belongs in the key:
- tenant_id: tenant A must never read tenant B's cached answers. That's a data leak, not a performance bug.
- model: don't return a smaller model's answer as if a stronger model wrote it.
- prompt version: bump it every time the system prompt changes.
- knowledge version: for RAG, the version of the source documents.
- locale: the same question in another language usually needs a different answer.
I combine the versions into one scope string to keep the query simple.
The Node.js code
This uses knex raw queries on Postgres. Plain pg works the same way with numbered placeholders. embed and callModel stand in for whatever your provider SDK offers.
import knexFactory from 'knex';
const db = knexFactory({ client: 'pg', connection: process.env.DATABASE_URL });
// TTL in seconds, per kind of question
const TTL = { faq: 86400, docs: 21600, default: 1800 };
function normalize(text) {
return text.normalize('NFKC').toLowerCase().replace(/\s+/g, ' ').trim();
}
// Embeddings blur numbers, so IDs and dates must match exactly
function hardTokens(text) {
return (text.match(/[0-9]+(?:[.,:-][0-9]+)*/g) || []).sort().join('|');
}
function shouldStore(answer) {
if (!answer || answer.length < 20) return false;
return !/i (can't|cannot|am unable to)/i.test(answer);
}
export async function cachedCompletion(opts) {
const {
tenantId, kind, model, promptVersion, kbVersion = 'none',
locale = 'en', prompt, threshold = 0.95, embed, callModel,
} = opts;
const text = normalize(prompt);
const scope = [kind, promptVersion, kbVersion, locale].join(':');
const vector = JSON.stringify(await embed(text));
const { rows } = await db.raw(
`SELECT id, prompt_text, answer,
1 - (embedding <=> ?::vector) AS similarity
FROM llm_cache
WHERE tenant_id = ? AND scope = ? AND model = ?
AND expires_at > now()
ORDER BY embedding <=> ?::vector
LIMIT 1`,
[vector, tenantId, scope, model, vector]
);
const best = rows[0];
const isHit = best
&& best.similarity >= threshold
&& hardTokens(best.prompt_text) === hardTokens(text);
if (isHit) {
db.raw('UPDATE llm_cache SET hit_count = hit_count + 1 WHERE id = ?', [best.id])
.catch(() => {});
return { answer: best.answer, cached: true, similarity: best.similarity };
}
const answer = await callModel(prompt);
if (shouldStore(answer)) {
await db.raw(
`INSERT INTO llm_cache
(tenant_id, scope, model, prompt_text, embedding, answer, expires_at)
VALUES (?, ?, ?, ?, ?::vector, ?, now() + make_interval(secs => ?))`,
[tenantId, scope, model, text, vector, answer, TTL[kind] ?? TTL.default]
);
}
return { answer, cached: false, similarity: best?.similarity ?? null };
}
Why the code makes these choices:
- Normalising before embedding strips out differences in case and spacing, so near-duplicate prompts get higher scores.
- The hard-token check is the cheapest guard against the worst wrong answers. "Refund for invoice 2024-118" and "refund for invoice 2024-119" have almost the same embedding. Requiring the numbers to match exactly prevents a false hit. If your domain needs it, extend the check to SKUs or product names.
- Refusals and errors are never stored, so one bad response isn't replayed to everyone who asks something similar.
-
kindsets the TTL. An FAQ answer can stay cached for a day. Answers built from documents that change often should expire within hours.
TTLs and invalidation
The lookup query already ignores expired rows, but those rows still use storage and bloat the index. Delete them on a schedule:
DELETE FROM llm_cache WHERE expires_at < now();
To invalidate, you don't need to find and delete rows. Bump kbVersion or promptVersion when content or prompts change. Old entries stop matching straight away and expire on their own.
Tuning the threshold
Most semantic caches go wrong here. Treat any threshold you copy from a blog post, including the 0.95 above, as a starting point. Cosine scores shift with the embedding model, the language and the length of your prompts.
Tune the threshold with real data in shadow mode:
- Run the lookup on every request, but always call the model and don't serve anything from the cache yet.
- Log the new prompt, the closest cached prompt, the similarity score and both answers.
- Label a sample: would the cached answer have been correct for the new prompt? A reviewer works, and so does an LLM grader with a strict rubric if you spot-check its labels.
- Group the results into similarity buckets (0.90 to 0.92, 0.92 to 0.94, and so on) and check what share of answers in each bucket were wrong.
- Set the threshold at the lowest bucket where wrong answers are rare enough for your product.
Decision rules that hold up in practice:
- Start strict. If the threshold is too high, a missed hit just costs one model call. If it's too loose, users get wrong answers.
-
Set a threshold for each
kind. General FAQ questions can tolerate a looser match than questions about accounts, billing or health. - Watch for negation. "Can I cancel without a fee" and "can I cancel with a fee" score very close to each other. For hits near the threshold, add a reranker or a small yes/no LLM check that asks whether both questions have the same answer.
- Keep logging after launch. Record the similarity score on every hit and let users flag bad answers. If users start phrasing questions differently, it will show up there first.
Production checklist
- Every lookup filters by tenant ID inside one shared function.
- The scope includes the model, the prompt version and the knowledge version.
- Numbers and IDs must match exactly before a hit counts.
- Refusals, errors and personalised answers are never stored.
- Iterative scans are on for multi-tenant indexes.
- Each threshold comes from labelled shadow-mode data.
- Hit rate, similarity distribution and user flags are on a dashboard.
One trade-off to keep in mind: the embedding call adds latency to every request, including misses. If your hit rate stays low after tuning, the cache can cost more than it saves, and dropping it is a reasonable call.
Caching sits next to retries, streaming, error handling and cost controls in any Node.js backend that calls LLMs. Geminate Solutions covers those pieces in a wider guide.
If you're adding AI features to a Node.js service, the Node.js AI integration guide is a good next read.
Top comments (0)