Growing memory is not automatically a leak.
In a real Go service, caches warm up, pools expand, buffers batch work, and connection managers hold state on purpose. If you only watch RSS or alloc_space, everything looks suspicious. The useful question is sharper:
After GC, does retained heap keep climbing without a ceiling — and which call stacks are holding it?
That is the problem godrain is built for.
The mental model
| Signal | Often intentional | Looks uncontrolled |
|---|---|---|
| Metric |
inuse_* after GC |
same |
| Shape | rises, then plateaus | sustained positive slope |
| Context | known cache/pool stacks | not allowlisted |
| Load | grows with RPS / queue depth | load is flat, heap still climbs |
alloc_space will grow on a busy service forever.
inuse_space after a forced GC is much closer to "what is still alive".
A tiny demo
godrain watches a live net/http/pprof endpoint, samples heap with gc=1, aggregates retention by stack, then classifies growth.
go install github.com/gonnafaraway/godrain/cmd/godrain@latest
Unbounded map growth:
# terminal 1
git clone https://github.com/gonnafaraway/godrain.git
cd godrain
go run ./testdata/leaky
# terminal 2
godrain watch \
--url http://127.0.0.1:6061 \
--warmup 5s \
--duration 45s \
--interval 5s
You should see something like suspect / leak attributed to the goroutine stuffing a global map.
Now the opposite case — a bounded LRU that should plateau:
go run ./testdata/cached
godrain watch --url http://127.0.0.1:6062 --warmup 5s --duration 45s --interval 5s
Same sampling pipeline. Different shape. Different verdict.
How the classifier works
Rough pipeline:
- Sample
/debug/pprof/heap?gc=1on an interval - Aggregate
inuse_spaceby call stack - Drop a warmup window (cache fill)
- Score slope, plateau, optional allowlist, optional load correlation
- Emit
ok/expected/suspect/leak
service (pprof)
→ heap snapshots with gc=1
→ inuse by stack
→ drop warmup
→ slope / plateau / allowlist / load
→ verdict + top stacks
Allowlists matter
Without an allowlist, a growing cache during warm-up looks guilty. After the first report, suppress known intentional allocators:
patterns:
- "internal/cache"
- "lru."
Load correlation helps even more
If heap grows with RPS, that may be expected. If RPS is flat and retained heap still climbs, that is much more interesting.
godrain watch --url http://127.0.0.1:6060 \
--load-url http://127.0.0.1:9090/metrics \
--load-metric http_requests_total \
--warmup 1m --duration 5m --json
Using it on a real service
- Expose pprof privately (
127.0.0.1/ internal network only) - Warm the process under realistic traffic
- Hold roughly constant load during the watch window
- Run:
godrain watch \
--url http://127.0.0.1:6060 \
--warmup 1m \
--duration 5m \
--interval 15s \
--json
If the signal is messy, split surfaces:
- API only
- workers only
- both together
That usually shows which contour is retaining memory.
For CI, non-zero exit on suspect/leak is intentional. Use --no-exit-code if you only want the report.
What godrain is not
- not a replacement for interactive
go tool pprof - not a static analyzer
- not a proof of a leak — it is a heuristic
- not a reason to expose
/debug/pprofto the internet
Treat pprof as sensitive. Saved heap profiles can leak internals too.
False positives / false negatives
You may get false suspect when:
- the window is shorter than cache fill
- load is ramping and you did not pass
--load-url - intentional growth is not allowlisted yet
You may get false ok when:
- the leak is slower than your thresholds
- the leaking path was not exercised during the watch
Longer soaks beat clever guesses.
Why I built this
I got tired of the loop:
- "memory is up"
- open one heap profile
- argue whether the cache is "supposed to do that"
- repeat tomorrow
godrain automates the boring part: compare retained growth over time under steady load, then point at stacks.
- Repo: github.com/gonnafaraway/godrain
- Release: v0.1.0
If you try it on a staging service, I especially want reports of false positives — that feedback is more valuable than stars.
Top comments (0)