TL;DR
One of our Node.js 22 API services was slowly eating memory until Kubernetes OOM-killed it roughly every 30 hours. I paired with Claude Code (CLI, October 2026 build) to hunt it down. The win didn't come from the AI "knowing" the answer — it came from forcing it to work from heap snapshots instead of hunches. Two leaks, one afternoon, and five lessons I'd steal if I were you.
The Problem
The service was boring in the best way: an Express-style HTTP API, a Postgres pool, a Redis client, and a few hundred requests per second at peak. It had been stable for a year.
Then the memory graph started looking like a staircase:
- Fresh pod: ~180 MB RSS
- After 12 hours: ~620 MB
- After ~30 hours: hits the 1 GiB limit, gets OOM-killed, restarts, repeat
Nobody noticed for weeks because the restarts were "self-healing." Then a restart landed during a batch import and we dropped a few hundred in-flight requests. Suddenly it mattered.
The annoying constraints:
-
It didn't reproduce locally. A quick
autocannonrun on my laptop showed flat memory. - The growth was slow. ~15 MB per hour is invisible in a 5-minute test.
- The diff since "it was fine" was huge. Three months of dependency bumps and ~140 merged PRs.
My first instinct was to paste the memory graph into Claude Code and ask "why is this leaking?" I did exactly that. It gave me a confident list of seven plausible causes — closures, global arrays, unclosed DB connections, timers — all reasonable, none grounded in our code. That was the moment I changed approach.
How I Solved It
Step 1: Make the agent collect evidence, not opinions
Instead of asking for causes, I asked Claude Code to help me instrument the service. The rule I gave it was explicit:
Don't propose a fix until we have two heap snapshots that show which object types grew.
First, a tiny memory logger so we could see whether the growth was in the V8 heap or outside it (native buffers, etc.):
// memory-probe.js — loaded only when MEMORY_PROBE=1
const INTERVAL_MS = 60_000;
setInterval(() => {
const m = process.memoryUsage();
console.log(JSON.stringify({
probe: "memory",
rssMb: Math.round(m.rss / 1e6),
heapUsedMb: Math.round(m.heapUsed / 1e6),
externalMb: Math.round(m.external / 1e6),
arrayBuffersMb: Math.round(m.arrayBuffers / 1e6),
}));
}, INTERVAL_MS).unref();
After a few hours in staging, the answer was clear: heapUsed tracked RSS almost one-to-one. Buffers were flat. So this was a plain old JS object leak, which is the good kind — heap snapshots can find it.
Step 2: Heap snapshots on demand
Node has a built-in flag for this, and Claude Code suggested it before I remembered it existed:
node --heapsnapshot-signal=SIGUSR2 server.js
Then in staging:
# snapshot 1, right after warm-up
kill -USR2 <pid>
# ...wait ~3 hours under replayed traffic...
# snapshot 2
kill -USR2 <pid>
To get realistic traffic, I had Claude Code write a small replay script that took a sample of sanitized access logs and fired them at staging at roughly production shape. This turned out to be the key, because synthetic load on one endpoint never triggered the leak.
Step 3: Let the agent read the diff — but give it the right diff
Heap snapshots are large JSON files (ours were ~250 MB each), so pasting them into a chat is a non-starter. Instead I loaded both into Chrome DevTools, used the Comparison view, sorted by # Delta, and exported the top rows as plain text. That summary was what I handed to Claude Code:
Constructor # New # Deleted # Delta Size Delta
(string) 912,340 301,220 +611,120 +41.2 MB
Object 455,002 120,889 +334,113 +28.7 MB
Map 18,421 210 +18,211 +6.9 MB
(closure) 37,550 1,104 +36,446 +3.1 MB
TLSSocket 0 0 0 0
Now the conversation got useful. I asked it: "Given this delta, list the places in our codebase that create long-lived Maps and long-lived closures. For each, tell me whether entries are ever deleted."
It came back with a table of 11 candidates. Two stood out.
Leak #1: A "temporary" cache with no eviction
Someone (fine, it was me, eight months ago) had added an in-memory cache for a per-tenant pricing lookup:
// before
const priceCache = new Map();
export async function getPrice(tenantId, sku, currency) {
const key = `${tenantId}:${sku}:${currency}`;
if (priceCache.has(key)) return priceCache.get(key);
const price = await loadPriceFromDb(tenantId, sku, currency);
priceCache.set(key, price);
return price;
}
With a handful of tenants this was fine. After a large customer onboarded with ~200,000 SKUs and four currencies, the key space exploded. Nothing was ever evicted. That explained the (string) and Object growth.
The fix was boring, which is how fixes should be:
// after
import { LRUCache } from "lru-cache";
const priceCache = new LRUCache({
max: 50_000, // hard ceiling on entries
ttl: 10 * 60 * 1000, // prices can change; 10 minutes is plenty
});
Leak #2: An event listener added per request
The (closure) and Map growth pointed somewhere else. Claude Code flagged this middleware, which subscribed to a shared emitter to support request cancellation:
// before
app.use((req, res, next) => {
const onShutdown = () => req.socket.destroy();
lifecycle.on("shutdown", onShutdown);
next();
});
Every request added a listener to a process-wide emitter and never removed it. Each closure held a reference to req, which kept the entire request object graph alive. We had even been getting MaxListenersExceededWarning in the logs — someone had "fixed" that months ago by calling lifecycle.setMaxListeners(0). Ouch.
// after
app.use((req, res, next) => {
const onShutdown = () => req.socket.destroy();
lifecycle.on("shutdown", onShutdown);
res.on("close", () => lifecycle.off("shutdown", onShutdown));
next();
});
I also had Claude Code remove the setMaxListeners(0) call and add a test that asserts the listener count returns to baseline after 1,000 requests:
test("shutdown listeners do not accumulate", async () => {
const before = lifecycle.listenerCount("shutdown");
await Promise.all(
Array.from({ length: 1000 }, () => request(app).get("/health"))
);
expect(lifecycle.listenerCount("shutdown")).toBe(before);
});
Step 4: Prove it, don't assume it
Here's the flow I ended up using, roughly:
flowchart LR
A[Memory graph looks wrong] --> B[Probe: heap vs external?]
B --> C[Two heap snapshots under replayed traffic]
C --> D[Comparison view: top deltas]
D --> E[Agent maps deltas to code]
E --> F[Fix + regression test]
F --> G[Re-run replay 6h, compare again]
After deploying both fixes to staging and replaying traffic for six hours, heapUsed plateaued at ~210 MB and stayed there. In production, pods have now run for over two weeks without a restart.
Total time: about 3.5 hours of focused work, most of which was waiting for snapshots.
Lessons Learned
1. "Why is this leaking?" is the wrong prompt 💡
When I asked for causes, I got a generic checklist. When I asked the agent to map specific evidence to specific code, I got two real bugs. AI coding agents are great at search and correlation across a codebase; they're mediocre at guessing runtime behavior they can't observe. Feed them observations.
2. Summarize big artifacts before handing them over
A 250 MB heap snapshot is useless in a context window. A 10-line delta table is gold. The skill isn't "give the agent everything" — it's compress the evidence to the part that discriminates between hypotheses. This applies to logs, traces, and profiles too.
3. Realistic load beats heavy load
My laptop benchmark hammered one endpoint and showed nothing. The leak only appeared with the real mix of tenants and SKUs. Having the agent write a log-replay script was the single highest-leverage thing it did all day.
4. Silenced warnings are leak detectors you turned off ⚠️
MaxListenersExceededWarning was literally telling us about leak #2 for months. If you've ever written setMaxListeners(0), go look at why. I now ask Claude Code to grep for suppressed warnings as part of any performance investigation.
5. Every fix ships with a test that would have caught it
The listener-count test takes 200 ms to run and would have caught leak #2 the day it was introduced. For the cache, we added a bounded-size assertion. It's tempting to stop when the graph goes flat. Don't — the graph will be someone else's problem in six months.
What's Next
A few things I'm working on now:
-
A nightly "soak" job that replays an hour of sanitized traffic against a staging pod and fails if
heapUsedgrows more than 10% after warm-up. -
A lint rule that flags
new Map()at module scope without an accompanying size bound or eviction strategy. - Turning this into a reusable debugging playbook for my agent, so the "evidence first, fixes second" rule is the default rather than something I have to remember to say.
Wrap-up
Memory leaks feel like black magic until you treat them as a data problem. Claude Code didn't magically know where the leak was — but once I gave it heap deltas instead of vibes, it found both culprits faster than I would have alone.
If you've got a leak war story (or a better heap-diff workflow), drop it in the comments 👇 I read every one.
Follow me here on Dev.to for more build-in-public posts on working with AI coding agents in real production codebases — next up is the soak-test setup. And if you haven't tried Claude Code for debugging yet, start with an evidence-first prompt on your next weird bug. 🚀
Top comments (3)
You need to verify your account.
Link is in the profile.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.