DEV Community

Cover image for The Models Read the Release Notes. I Measured the Runtime.
Nazar Boyko
Nazar Boyko

Posted on AI-assisted

The Models Read the Release Notes. I Measured the Runtime.

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

I've written PHP for most of my career and Go for the last few years, and in code reviews I repeat the same memory rules everyone repeats. A goroutine is 2 KB, don't worry about it. An int64 to int64 map entry is 16 bytes. unset() gives the memory back. Use map[K]struct{} for sets, because struct{} is free. I never checked any of them. Why would I? They've always been true.

Except some of them aren't anymore. PHP 8.2 stopped storing keys in packed arrays (php-src #7491). Go 1.24 threw out its old map implementation for Swiss Tables. Go 1.19 started picking the initial goroutine stack size from what it had seen so far. The rules stayed in our heads; the runtimes moved on. So I got curious: when a model answers a memory question, what is it actually remembering? The rule, the runtime, or the specific version?

Let's play. Guess before you scroll. Both maps hold one million keys:

set := make(map[int64]struct{})
for i := int64(0); i < 1_000_000; i++ {
    set[i] = struct{}{}
}

flags := make(map[int64]bool)
for i := int64(0); i < 1_000_000; i++ {
    flags[i] = true
}
Enter fullscreen mode Exit fullscreen mode

How does the heap held by set compare with flags?

  • A. set is about 10% smaller
  • B. set is about half the size of flags
  • C. About the same
  • D. It depends on the Go version

Pick one. The answer, with the measured number, is waiting in Findings. If you enjoy this kind of thing, all seventeen cases are on this page: you pick an option, then you see what the runtime actually did.

Seventeen cases in total: seven about PHP arrays, five about Go maps, five about goroutine stacks. Three of them are controls, cases where the old rule is still right. Without those, a model could score well just by betting against every rule, and I didn't want to reward cynicism.

One thing I was strict about: not a single "correct answer" came from a blog post, including mine. A small harness builds each structure and reads memory_get_usage() in PHP or runtime.MemStats in Go. A GitHub Actions matrix runs it three times on PHP 8.1 through 8.4 and Go 1.23 through 1.26. If you don't trust my numbers, and you shouldn't have to, one command reruns everything:

sh measure/run.sh
Enter fullscreen mode Exit fullscreen mode

Each case went to the models three times, as three separate Kaggle tasks:

  • cold: just the code and the question,
  • versioned: the same, plus the runtime version,
  • evidence: the same, plus one measured number from a neighbouring case.

Models answer in JSON: an option, an estimate in MB or KB, which mechanism they think explains it, and a confidence from 0 to 100. Plain code does the scoring, no LLM judge anywhere. A third of the point for the right option, a third for an estimate within 0.67x to 1.5x of the measured value, a third for the right mechanism.

Models Tested

Twelve models from Kaggle's list. I picked them along three lines: a small and a flagship model from each vendor, a newer and an older generation of the same vendor (did newer training data pick up Go 1.24 maps?), and a few open-weight models.

gemini-3.5-flash-lite, gemini-3.8-flash, gemini-3.1-pro-preview, gemini-2.5-flash, claude-haiku-5-5, claude-opus-5-5-default, claude-haiku-4-5-20251001, gpt-5.4-nano-2026-03-17, gpt-6.1-sol, gpt-oss-120b, gemma-4-31b-it, qwen3-coder-480b-a35b-instruct.

Nothing returned a 404 and I swapped nothing out. Eleven of the twelve have a complete run on all three tasks. The one that doesn't is gemini-3.1-pro-preview: 13 cold questions cost $0.89, one of its replies ran to 22,064 output tokens, and Kaggle gives my account $10 of model calls a day. Its partial replies are in the repository, but it's in none of the tables below.

There's also a thirteenth model in the tables, and I didn't invite it. When you create a task, Kaggle runs it once on gemini-3.7-flash. Those runs are as real as mine, so I kept them.

Findings

First, the answer to the game. On Go 1.23 the struct{} set is about 10% smaller than the bool map (ratio 0.91). On Go 1.24 and later they're the same size (ratio 1.00). So if nobody told you the version, the honest answer is D. Swiss-table maps keep key and value together in one aligned slot, and an empty value saves you nothing. When I put the version in the prompt, exactly 2 of 12 models followed the change: gpt-6.1-sol and claude-opus-5-5-default. Four others did change their answer for Go 1.24, in the wrong direction, to "half the size".

The two best models lost the same question, and it was a control. Here's the setup. I start 100,000 goroutines that each hold an 8 KB array, run a GC, then start 100,000 idle goroutines, and ask how much stack each idle one takes. The old rule says 2 KB. The Go 1.19 release notes say the runtime sizes new stacks from the average it has seen, which points at 16 KB. I measured 2.0 KB, on Go 1.23 through 1.26. Then I changed one thing: I let the first batch exit before starting the idle batch. Now each idle goroutine takes 16.0 KB.

Both numbers are right, and the runtime source explains why. newproc1 gives a brand-new goroutine a stack of stackMin, which is 2 KB. gfget hands out the descriptor of a finished goroutine again, and that one gets startingStackSize, the adaptive size. The links point at the current release, Go 1.27; the four versions I measured have the same two lines.

gpt-6.1-sol got 56 of 59 questions right. The three it missed are this control, once per task, at a stated confidence of 98 or 99. claude-opus-5-5-default missed it three times too. Both basically recited the release note back to me, with arithmetic. Across 36 runs, a model kept the control and lost its twin 19 times, did the reverse 14 times, and got both right twice. Both of those were gemini-3.5-flash-lite.

This is the part that changed how I'll use these models for runtime questions. A strong model knows the changelog better than I do. It also applies the changelog in places where the code doesn't.

Control case: gpt-6.1-sol said 16 KB at confidence 98, 99 and 98, claude-opus-5-5-default said 16 KB at confidence 72, 80 and 72, and the harness measured 2.0 KB on Go 1.23, 1.24, 1.25 and 1.26

The leaderboard. Task score, one complete run per model and task. The row without a slug is me, cold condition only.

Mean task score per model: gpt-6.1-sol 0.95, claude-opus-5-5-default 0.94, gemini-3.8-flash 0.83, claude-haiku-5-5 0.81, gemini-3.7-flash 0.80, gemma-4-31b-it 0.70, gemini-2.5-flash 0.67, gpt-oss-120b 0.55, claude-haiku-4-5-20251001 0.52, gemini-3.5-flash-lite 0.50, qwen3-coder-480b-a35b-instruct 0.46, gpt-5.4-nano-2026-03-17 0.39; the author scored 0.67 on the cold task

model cold versioned evidence mean
gpt-6.1-sol 0.94 0.95 0.95 0.95
claude-opus-5-5-default 0.90 0.95 0.95 0.94
gemini-3.8-flash 0.84 0.86 0.79 0.83
claude-haiku-5-5 0.78 0.83 0.81 0.81
gemini-3.7-flash 0.80 0.81 0.78 0.80
gemma-4-31b-it 0.69 0.70 0.71 0.70
gemini-2.5-flash 0.67 0.68 0.65 0.67
author, cold only 0.67
gpt-oss-120b 0.53 0.57 0.56 0.55
claude-haiku-4-5-20251001 0.57 0.51 0.48 0.52
gemini-3.5-flash-lite 0.51 0.49 0.49 0.50
qwen3-coder-480b-a35b-instruct 0.41 0.44 0.52 0.46
gpt-5.4-nano-2026-03-17 0.37 0.35 0.44 0.39

Within a family, the newer generation wins by a lot: claude-haiku-5-5 at 0.81 against claude-haiku-4-5-20251001 at 0.52, gemini-3.8-flash at 0.83 against gemini-2.5-flash at 0.67. The older Gemini also burned more than twice the output tokens per question (5,445 against 2,235) to get the lower score. Not one reply failed to parse.

Version sensitivity. Four cases have a measured answer that flips with the version: three between PHP 8.1 and 8.3, one between Go 1.23 and 1.24. I asked each of those twice, once per version. gpt-6.1-sol and claude-opus-5-5-default got all four pairs right. gemini-3.8-flash and claude-haiku-5-5 got three, gemini-3.7-flash two, gpt-oss-120b one. Six models got none. All six said "2.5x" for the reversed array on PHP 8.1 and again on PHP 8.3, where I measured 1.25x and 2.50x. The Swiss-table case was the hardest (2 of 12 models), the SplFixedArray case the easiest (6 of 12).

"It depends" is a rare answer. With no version in the prompt, the right answer to those four cases is "it depends on the version". Models picked it in 13 of 48 replies. Most just answered for the newest runtime and moved on. My favourite: claude-opus-5-5-default wrote "Before PHP 8.2 the ratio was about 1.25x" in its reasoning, and then chose the PHP 8.2 option anyway.

Confidence. The stated number told me very little. Seven of the twelve models put 80 or more on every single wrong answer they gave. gemini-3.5-flash-lite and qwen3-coder-480b-a35b-instruct did it 32 times each, gemini-2.5-flash 22 times, gemma-4-31b-it 19. Only two models dropped by more than five points when they were wrong: claude-opus-5-5-default from 77.5 to 70.2, and claude-haiku-5-5 from 70.7 to 61.6. Haiku 5.5 put 80 or more on none of its 17 wrong answers, which is more self-awareness than I expected from a small model.

Right answer, wrong reason. qwen3-coder-480b-a35b-instruct picked the measured answer 27 times and explained 11 of those with a wrong mechanism. For gpt-5.4-nano-2026-03-17 it was 6 of 16. For the top two it was 0. If I had scored on the answer letter alone, none of this would have shown up.

Evidence barely helped. Handing a model one measured number from a neighbouring case moved almost nothing. From cold to evidence, the twelve models gained 14 right answers and lost 14.

Ask the same thing twice. Quota errors cut my first pass to pieces, so most models ended up answering many prompts twice. Across 501 pairs of replies to the same prompt, the answer letter matched 82% of the time. gpt-6.1-sol (19 of 19) and claude-opus-5-5-default (25 of 25) never changed their minds. gpt-5.4-nano-2026-03-17 and gemini-2.5-flash matched only 65% of the time. With one run per model, a gap of a few points on this leaderboard can be noise, and you should read it that way.

Which rules are dead, and which refuse to die. Not one model chose the folklore explanation for unset() freeing memory, for the 16-byte map entry, or for delete giving a map's memory back. Those rules are gone from the training data, it seems. The rules that were once true are the ones that live on. "SplFixedArray saves memory" was the chosen mechanism in 22 of 60 replies and "struct{} is free" in 21 of 60. I measured SplFixedArray at about half the memory of a plain list on PHP 8.1 (ratio 0.48) and about the same on PHP 8.3 (ratio 0.95). That rule held right up to PHP 8.2 and then quietly stopped.

In their own words.

gpt-6.1-sol on the control, confidence 98, measured 2.0 KB:

Since Go 1.19, adaptive stack sizing uses the average stack space scanned during the last GC, adds a stack guard, and rounds up to a power of two. The first batch dominates that average with live frames exceeding 8 KB, so the second batch starts with approximately 16 KB stacks.

qwen3-coder-480b-a35b-instruct on the game from the top, measured ratio 0.91 on Go 1.23 and 1.00 on Go 1.24:

struct{}{} takes zero bytes, so set only stores keys while flags stores both keys and bool values. Each bool is 1 byte, so flags uses roughly twice the memory for values.

gpt-5.4-nano-2026-03-17 on 100,000 three-property objects against 100,000 three-key arrays, where the arrays measured 3.12 times the memory on PHP 8.3:

In PHP, each $objects element is an object with overhead (zval + object header/handlers/property metadata), while $arrays elements are arrays with overhead too (zval + HashTable + buckets). For 3 fields, these overheads are of the same order, so memory_get_usage() is typically close.

A few more things the harness measured. All of these held on every version I ran:

  • unset() on every second element of a one-million-element list leaves memory_get_usage() exactly where it was (ratio 1.00 on PHP 8.1 and 8.3).
  • An int64 to int64 map holds about 38.1 bytes per entry on Go 1.24. The rule says 16.
  • Deleting every key and clear(m) both keep the full map after a GC (36.38 MB on Go 1.24).
  • A goroutine holding an 8 KB local array takes 16.0 KB of stack. "2 plus 8" would suggest 10.
  • One million two-element arrays [$i, $i] take 222.0 MB on PHP 8.3 and 390.59 MB on PHP 8.1. A pair of integers is 32 bytes, in theory.

My own score. 🤭 Before I looked at a single measurement, I answered the cold questions myself. I scored 0.67 on the same scale as the models: 10 of 17 answer letters, 14 of 17 mechanisms, 10 of 17 estimates inside the tolerance. My mean confidence was 80.5 when I was right and 70.0 when I was wrong. The cases that sting are P1, P4, P8 and G3. On all four the measured answer was "it depends on the version", and I did exactly what most of the models did: answered for the version I use, got the mechanism right on three of them, and never once said "it depends".

The quota, for whoever tries this next. Kaggle's model proxy reserves the worst-case cost of every call against a $10 daily quota: the model's maximum output at its output price, which is $2.56 for a single claude-opus-5-5-default call. I started twelve models at once and 160 questions came back refused. The fix was max_tokens on every call and no more than four models at a time. All 71 run files together cost $7.84.

What I'd measure next. Laravel collections on top of PHP arrays. The same harness on Alpine with musl, where the allocator is different. Goroutine stacks under real request handlers instead of synthetic frames.

Limits, so nobody has to dig for them. Seventeen cases and one scored run per model and task; the repeated replies above come from a partial first pass. One runner type per version (GitHub's ubuntu-24.04). Replies that never parse score zero and are counted; there were none. The Kaggle Benchmarks library sends no temperature to the model proxy, so a rerun can differ. Every call carries max_tokens=32768; the longest reply among the scored runs used 22,329 tokens. gemini-3.1-pro-preview has no complete run. The human baseline is one person, and he's biased.

My Benchmark

The MemoryTrap Bench leaderboard on Kaggle: 3 tasks, 12 models; the first four columns show GPT-6.1 Sol 0.95, Claude Opus 5.5 0.94, Gemini 3.8 Flash 0.83 and Claude Haiku 5.5 0.81

The first four of twelve columns; the full board is behind the leaderboard link above.

Credits: the Kaggle Benchmarks library and platform ran the models; setup-php and setup-go gave me eight runtime versions in CI; the PHP and Go release notes, the pull request and the Go runtime source linked above explain the changes the harness measured.

Top comments (0)