DEV Community

Vayun Godara
Vayun Godara

Posted on Fully Autonomous

Two 32 MB tokenizers: hunting a memory floor in a Rust proxy

cliproxy-rs is a Rust rewrite of CLIProxyAPI. It's one local server that Codex CLI, Claude Code and other tools talk to, and it forwards to whatever you've configured: your own API keys, your own accounts. I run it on a small Linux box under real coding-agent traffic.

On 0.2.0 that box showed something I didn't like. The resting memory, meaning what the process holds between requests, rose from 26 MB to 89 MB over about seven hours. Peaks were 300 to 400 MB. Nothing was failing, but a floor that only goes up is how a leak usually starts.

Reproducing it

My soak test only replayed Claude streams of 100 to 500 KB. The real traffic was different: Claude and Codex requests of 1 to 2 MB from long agent sessions, plus count_tokens calls. So I wrote a "field mix" replay with exactly that shape (four sessions, bodies averaging 1 MB) and ran 0.2.0 and the fix branch side by side on separate GitHub runners for an hour.

0.2.0 rested at 128 to 145 MB. That was clearly too much for a proxy that does no caching of its own.

What heaptrack found

After two batches of 100 field-mix turns, heaptrack showed 67.1 MB still allocated at exit. 63.8 MB of that was two o200k_base tokenizer encoders, 31.9 MB each. They were identical. Claude count_tokens had built one, and Codex count_tokens had built the other.

Five modules each kept their own lazily built tiktoken encoder in a static. Each was built the first time something needed a token count and never freed. On a server that counts tokens for both Claude and Codex models and streams Claude-format requests to Codex models, up to seven copies could exist, 164 MB of tables in total.

One encoder is bigger than its 31.9 MB of live data suggests. It holds about 600,000 short token allocations, and at malloc's 32-byte minimum chunk that comes to 47 MB resident.

I also checked whether glibc was just holding freed memory. It wasn't: after a trim, calling malloc_trim(0) again changed resident memory by 40 kB. The floor was live data.

The fix

One shared encoder per encoding, still built only when a count first needs it. The same mix now holds 71 MB of tables instead of 164.

Same hour of traffic, same runners:

0.2.0 0.2.2
Resting RSS 128.2 to 145.4 MB 80.6 to 97.6 MB
Peak RSS (VmHWM) 220.1 MB 169.0 MB

CPU per request didn't change (43.1 ms against 42.9 and 44.1 ms on one 2-vCPU machine).

What I haven't solved

  • The replay steps up once and then stays flat. It never shows the slow climb from the field. My best guess is that the field server built its encoders hours apart, which would look like a slow rise, but I don't know that yet. A 24-hour trace of 0.2.2 is running now.
  • Peaks are still big. A Codex request peaks at 8.8 times its body size, because request shaping rewrites the body pass by pass and each pass is a new copy. Claude requests peak at 6.0x.
  • A fixed MALLOC_MMAP_THRESHOLD_ took 26.6 to 33.8 MB off the peak, but cost 16 to 31% more CPU per request, so I didn't ship it.
  • The Go original is still faster on small requests. In my 0.1.0 benchmark, it served 1,568 non-streamed chat requests a second to cliproxy-rs's 1,168.

Keeping it fixed

Every release tag now runs an hour of the field mix, and a 5-hour run happens every Saturday. A run fails if the resting floor rises, if the peak goes over its bound or if any request fails. Heap budgets in the tests now cover the Codex and count_tokens routes as well as Claude.

The method and the raw rows are in docs/BENCHMARKS.md, and the code is at github.com/vayungodara/cliproxy-rs.

Built with some help from AI.

Top comments (0)