DEV Community

Dinesh_gowtham
Dinesh_gowtham

Posted on

Prompt Caching with ElastiCache: Speed Up Claude Calls in Node.js 22, Explained Simply

LLMs are fast, but every repeated request still costs time and tokens. What if you could memoize Claude’s answers in a Redis cache and serve them instantly? This post shows how to turn ElastiCache Serverless into a zero‑latency prompt cache for your Node.js 22 app.


Why Prompt Caching Matters for LLM‑Powered Apps

Large language models (LLM) such as Claude answer a prompt – the text you send them – by running a massive neural network. Even though the inference step takes only a few hundred milliseconds, each call:

  • consumes API dollars (Claude charges per token)
  • burns the quota of tokens you allocated for your app
  • adds latency that users notice, especially on mobile or edge devices

Think of an LLM like a chef who prepares a custom dish from scratch for every order. If a customer keeps ordering the same dish, the chef still has to start from the beginning each time, even though the recipe never changes. Prompt caching is the equivalent of pre‑cooking the dish and keeping it warm so the next order is served instantly.

When the same user (or a different user) asks the same question, you can skip the expensive cooking step and hand back the cached answer. The performance win feels free, but only if the cache behaves correctly. A mismatched cache key or an outdated entry can feed the wrong answer back to your users, silently breaking your app’s logic.

In plain English: Cache the result of a prompt, not the prompt itself. If the prompt changes even a little, you need a new key; otherwise you risk serving the wrong answer.


ElastiCache Serverless: Setup and Cost Model

ElastiCache Serverless is a managed, on‑demand version of Redis (or Valkey) that scales automatically. It removes the need to provision a fixed number of nodes, but the pricing model is different:

Cost factor What it means
Per‑operation cost Every GET or SET you run costs a few micro‑dollars. At low traffic this is cheap; at high traffic the per‑operation fee adds up faster than a provisioned cluster.
Valkey compatibility ElastiCache Serverless runs Valkey, a Redis‑compatible fork. Most Redis commands work, but a few edge‑case commands (e.g., MEMORY USAGE with certain flags) behave slightly differently.
VPC placement The cache lives inside your VPC. If your Lambda or EC2 instance is also inside the same VPC, a cold start may add extra seconds because the network interface must be attached.

Quick setup checklist

# 1. Create a Serverless cache via the AWS console or CLI
aws elasticache create-replication-group \
  --replication-group-id my-cache \
  --engine valkey \
  --cache-node-type serverless \
  --num-node-groups 1 \
  --automatic-failover-enabled

# 2. Note the endpoint (e.g., my-cache.xxxxxx.use1.cache.amazonaws.com:6379)
# 3. Attach the VPC security group that your Lambda/EC2 uses
Enter fullscreen mode Exit fullscreen mode

Tip: Enable encryption‑in‑transit and at‑rest; it adds negligible latency and keeps data safe.


Connecting Node.js 22 to Redis with the New fetch‑Based Client

Node.js 22 ships with an experimental fetch‑based Redis client (node:redis). It uses the standard fetch API under the hood, so you don’t need a separate native binding like ioredis. This keeps the deployment footprint small (especially for Lambda) and works nicely with TypeScript’s built‑in types.

Install the client (still a separate npm package)

npm i redis@^4.6   # version that supports fetch‑based transport
Enter fullscreen mode Exit fullscreen mode

Minimal connection code

import { createClient } from 'redis';

// The endpoint we got from the ElastiCache console
const REDIS_ENDPOINT = process.env.REDIS_ENDPOINT ?? 'my-cache.xxxxxx.use1.cache.amazonaws.com:6379';

// Build a client that talks to Redis over TCP using fetch internally
const redisClient = createClient({
  // url uses the "redis://" scheme; fetch will handle the TCP handshake
  url: `redis://${REDIS_ENDPOINT}`,
  // Optional: give fetch a short timeout so we fail fast on network hiccups
  socket: {
    // 2 seconds is usually enough for an in‑VPC call
    timeout: 2000,
  },
});

redisClient.on('error', (err) => {
  console.error('Redis connection error:', err);
});

await redisClient.connect(); // Returns a promise; wait for the handshake
Enter fullscreen mode Exit fullscreen mode

Key takeaway: The fetch‑based client is just another way to talk to Redis; the rest of your code (GET/SET) stays the same.


Implementing the Cache‑First Claude Call Loop

Now that we have a Redis client, we can build the core function:

  1. Hash the incoming prompt to create a stable cache key.
  2. Lookup the key in Redis. If we find a value, parse it and return.
  3. If the key is missing, call Claude’s HTTP API with fetch.
  4. Store the JSON response in Redis with a TTL (time‑to‑live) of 5 minutes.
  5. Return the fresh response.

Why we hash instead of using the raw prompt

A raw prompt string can be long, contain line breaks, or differ only in whitespace. Storing the entire prompt as a key makes Redis do extra work and creates many near‑duplicate keys. A hash (e.g., SHA‑256) acts like a fingerprint: two identical prompts produce the same short identifier, while any change (even a stray space) yields a different fingerprint.

Full TypeScript implementation

import { createHash } from 'crypto';          // built‑in Node module for hashing
import { fetch } from 'node-fetch';           // Node 22 has native fetch; import only for type hinting
import { redisClient } from './redisClient';  // the client we connected above

// Claude API details (replace with your real endpoint and key)
const CLAUDE_ENDPOINT = 'https://api.anthropic.com/v1/messages';
const CLAUDE_API_KEY = process.env.CLAUDE_API_KEY!;

// Helper: turn a prompt into a short, stable key
function promptToCacheKey(prompt: string): string {
  // SHA‑256 gives a 64‑char hex string; we prefix to avoid collisions with other caches
  const hash = createHash('sha256').update(prompt.trim()).digest('hex');
  return `claude:prompt:${hash}`;
}

// Main function: tries cache first, falls back to Claude, then caches result
export async function getClaudeAnswer(prompt: string): Promise<any> {
  // Step 1 – derive cache key
  const cacheKey = promptToCacheKey(prompt);

  // Step 2 – attempt to read from Redis
  const cached = await redisClient.get(cacheKey);
  if (cached) {
    // Cache hit: parse JSON and return immediately
    console.log('Cache hit for key', cacheKey);
    return JSON.parse(cached);
  }

  // Step 3 – no cache entry, call Claude
  console.log('Cache miss, calling Claude...');
  const response = await fetch(CLAUDE_ENDPOINT, {
    method: 'POST',
    headers: {
      'Content-Type': 'application/json',
      'x-api-key': CLAUDE_API_KEY,
    },
    body: JSON.stringify({
      model: 'claude-3-5-sonnet-20241001',
      max_tokens: 1024,
      messages: [{ role: 'user', content: prompt }],
    }),
  });

  if (!response.ok) {
    // Surface the error; in a real app you might retry or fallback
    const errorBody = await response.text();
    throw new Error(`Claude API error ${response.status}: ${errorBody}`);
  }

  const json = await response.json(); // Claude returns JSON with the answer

  // Step 4 – store the result in Redis with a 5‑minute TTL
  // JSON.stringify makes sure we store a string (Redis stores bytes)
  await redisClient.set(cacheKey, JSON.stringify(json), {
    EX: 300, // EX sets expiration in seconds; 300 s = 5 min
  });

  // Step 5 – return the fresh answer
  return json;
}
Enter fullscreen mode Exit fullscreen mode

Explanation of each block

Line What it does
createHash('sha256') Creates a SHA‑256 hash object, a cryptographic fingerprint.
.update(prompt.trim()) Feeds the trimmed prompt (removes leading/trailing whitespace) into the hash.
.digest('hex') Produces a human‑readable hex string (64 characters).
redisClient.get(cacheKey) Looks up the value for that key; returns null if not present.
fetch(CLAUDE_ENDPOINT, …) Sends an HTTP POST to Claude’s API using the built‑in fetch.
await response.json() Parses the response body as JSON; Claude returns an object with content.
redisClient.set(..., { EX: 300 }) Stores the JSON string for 300 seconds (5 minutes). The TTL ensures stale data disappears.

Tip: Keep the TTL short for rapidly changing knowledge domains (news, weather). For static FAQs you can set a longer TTL or even store forever.


Gotchas: Cache Key Design, TTL, and Serialization Pitfalls

1. Cache key collisions

If you concatenate the prompt directly with other identifiers (e.g., user ID) without a separator, "user42:hello" and "user4:2hello" could clash. Always use a clear delimiter or, better, hash each component separately and join them with a colon.

function makeKey(userId: string, prompt: string): string {
  const promptHash = createHash('sha256').update(prompt).digest('hex');
  return `claude:user:${userId}:prompt:${promptHash}`;
}
Enter fullscreen mode Exit fullscreen mode

2. Whitespace and case sensitivity

A prompt that differs only by extra spaces or different capitalisation will produce a different hash, causing a cache miss. Decide whether you want the cache to be case‑insensitive and whitespace‑agnostic, then normalise before hashing.

const normalised = prompt.replace(/\s+/g, ' ').trim().toLowerCase();
Enter fullscreen mode Exit fullscreen mode

3. TTL (time‑to‑live) sizing

Setting a TTL that’s too long can serve outdated information; too short defeats the purpose of caching. A common pattern is stale‑while‑revalidate: serve the cached value even if it’s a few seconds old while you refresh it in the background. Implementing that requires a second key (e.g., :lock) or a background job, which is beyond this simple tutorial but worth noting.

In plain English: Think of TTL like the expiration date on milk. Short‑lived data (like news) needs a fresh date; stable data (like a product description) can sit longer.

4. Serialization format

Redis stores bytes; we usually stringify JSON with JSON.stringify. If you store the raw object (without stringifying) in a client that automatically performs BSON conversion, you could end up with binary data that JSON.parse can’t read. Stick to a single format throughout the app.

// WRONG – mixing binary and string
await redisClient.set(key, Buffer.from(JSON.stringify(obj)));
Enter fullscreen mode Exit fullscreen mode
// RIGHT – always store a string
await redisClient.set(key, JSON.stringify(obj));
Enter fullscreen mode Exit fullscreen mode

5. Valkey‑specific quirks

Valkey’s SET command accepts the same EX option, but some advanced features like SETEX behave slightly differently. If you ever need to use a command that fails, check the Valkey docs or test locally with a Valkey Docker image.

6. VPC cold‑start cost

When you run this code inside an AWS Lambda that lives in a VPC, the first request after a period of inactivity must attach an ENI (elastic network interface). That adds 1‑2 seconds before the Redis client can even start. Mitigate by:

  • Keeping the Lambda warm (scheduled ping).
  • Using Provisioned Concurrency if latency budgets are strict.

Key takeaway: The biggest hidden latency isn’t the cache lookup; it’s the network plumbing that brings your function into the VPC.


The Takeaway

What you should remember after reading this guide

  • Prompt caching removes token cost and latency by re‑using identical LLM responses.
  • ElastiCache Serverless gives you a managed Redis‑compatible store, but watch per‑operation fees, Valkey quirks, and VPC cold‑starts.
  • Node.js 22’s fetch‑based Redis client lets you talk to the cache without native bindings, keeping Lambda packages tiny.
  • Create a stable cache key by hashing a normalised prompt (think of it as a fingerprint).
  • Store the Claude JSON response as a string with a sensible TTL (5 minutes works for many chat‑bot use‑cases).
  • Guard against subtle bugs: whitespace differences, case sensitivity, serialization mismatches, and accidental key collisions.

By following these steps you’ll turn repeated Claude calls into instant look‑ups, saving money, cutting latency, and keeping your app’s answers reliable. Happy caching!


Transparency notice

This article was written with the help of an AI system — Groq (GPT OSS 120B).

Published: 2026-10-09 · Primary focus: ElastiCache

All code blocks are intended to be correct and runnable, but please verify them
against the official docs for the tools mentioned before using in production.

Find an error? Drop a comment — corrections are always welcome.

Top comments (0)