DEV Community

w0i
w0i

Posted on Originally published at github.com

Uncontrolled heap retention in Go: how to tell a leak from a healthy cache

Growing memory is not automatically a leak.

In a real Go service, caches warm up, pools expand, buffers batch work, and connection managers hold state on purpose. If you only watch RSS or alloc_space, everything looks suspicious. The useful question is sharper:

After GC, does retained heap keep climbing without a ceiling — and which call stacks are holding it?

That is the problem godrain is built for.

The mental model

Signal Often intentional Looks uncontrolled
Metric inuse_* after GC same
Shape rises, then plateaus sustained positive slope
Context known cache/pool stacks not allowlisted
Load grows with RPS / queue depth load is flat, heap still climbs

alloc_space will grow on a busy service forever.

inuse_space after a forced GC is much closer to "what is still alive".

A tiny demo

godrain watches a live net/http/pprof endpoint, samples heap with gc=1, aggregates retention by stack, then classifies growth.

go install github.com/gonnafaraway/godrain/cmd/godrain@latest
Enter fullscreen mode Exit fullscreen mode

Unbounded map growth:

# terminal 1
git clone https://github.com/gonnafaraway/godrain.git
cd godrain
go run ./testdata/leaky

# terminal 2
godrain watch \
  --url http://127.0.0.1:6061 \
  --warmup 5s \
  --duration 45s \
  --interval 5s
Enter fullscreen mode Exit fullscreen mode

You should see something like suspect / leak attributed to the goroutine stuffing a global map.

Now the opposite case — a bounded LRU that should plateau:

go run ./testdata/cached
godrain watch --url http://127.0.0.1:6062 --warmup 5s --duration 45s --interval 5s
Enter fullscreen mode Exit fullscreen mode

Same sampling pipeline. Different shape. Different verdict.

How the classifier works

Rough pipeline:

  1. Sample /debug/pprof/heap?gc=1 on an interval
  2. Aggregate inuse_space by call stack
  3. Drop a warmup window (cache fill)
  4. Score slope, plateau, optional allowlist, optional load correlation
  5. Emit ok / expected / suspect / leak
service (pprof)
  → heap snapshots with gc=1
  → inuse by stack
  → drop warmup
  → slope / plateau / allowlist / load
  → verdict + top stacks
Enter fullscreen mode Exit fullscreen mode

Allowlists matter

Without an allowlist, a growing cache during warm-up looks guilty. After the first report, suppress known intentional allocators:

patterns:
  - "internal/cache"
  - "lru."
Enter fullscreen mode Exit fullscreen mode

Load correlation helps even more

If heap grows with RPS, that may be expected. If RPS is flat and retained heap still climbs, that is much more interesting.

godrain watch --url http://127.0.0.1:6060 \
  --load-url http://127.0.0.1:9090/metrics \
  --load-metric http_requests_total \
  --warmup 1m --duration 5m --json
Enter fullscreen mode Exit fullscreen mode

Using it on a real service

  1. Expose pprof privately (127.0.0.1 / internal network only)
  2. Warm the process under realistic traffic
  3. Hold roughly constant load during the watch window
  4. Run:
godrain watch \
  --url http://127.0.0.1:6060 \
  --warmup 1m \
  --duration 5m \
  --interval 15s \
  --json
Enter fullscreen mode Exit fullscreen mode

If the signal is messy, split surfaces:

  • API only
  • workers only
  • both together

That usually shows which contour is retaining memory.

For CI, non-zero exit on suspect/leak is intentional. Use --no-exit-code if you only want the report.

What godrain is not

  • not a replacement for interactive go tool pprof
  • not a static analyzer
  • not a proof of a leak — it is a heuristic
  • not a reason to expose /debug/pprof to the internet

Treat pprof as sensitive. Saved heap profiles can leak internals too.

False positives / false negatives

You may get false suspect when:

  • the window is shorter than cache fill
  • load is ramping and you did not pass --load-url
  • intentional growth is not allowlisted yet

You may get false ok when:

  • the leak is slower than your thresholds
  • the leaking path was not exercised during the watch

Longer soaks beat clever guesses.

Why I built this

I got tired of the loop:

  1. "memory is up"
  2. open one heap profile
  3. argue whether the cache is "supposed to do that"
  4. repeat tomorrow

godrain automates the boring part: compare retained growth over time under steady load, then point at stacks.

If you try it on a staging service, I especially want reports of false positives — that feedback is more valuable than stars.

Top comments (0)