DEV Community

finaltype
finaltype

Posted on Originally published at finaltype.github.io

"[Sep 17] First Run of a Benchmark Harness — Plugging the Leaks One by One"

I found that yesterday's model setting fix hadn't actually reached the new benchmark tool at all, spent the day plugging leak after leak, and also finally closed out a set of safety items that had sat untouched for 11 days.

This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).

"Did yesterday's fix actually make it into the new tool?"

Yesterday I fixed the sampling settings for the models behind my AI reports.

Today's job was to prepare the first verification run of a new benchmark tool (perf-harness) I've been building to measure how fast those models run.

Partway through, the question came up: did yesterday's fix actually carry over into this test run?

The answer turned out to be no.

The benchmark tool was building a fresh, minimal config object for every trial, and it never once looked at the value I'd just fixed the day before.

If I'd run the first verification as-is, it would have measured the wrong condition entirely — making the whole benchmark meaningless.

The same kind of leak showed up two more times

Once I started suspecting it, the same class of bug kept turning up.

One case: the benchmark request was overwriting the startup config used when spinning up a fresh server. Another: a separate measurement path that replays the real pipeline directly wasn't carrying the value either.

I fixed all three the same way — inject the value if one's set, fall back to the old behavior if not — and added tests for each.

The sample itself was a moving target

While reviewing the design, I noticed the sampling approach for test data had its own problem.

Until now, every run pulled a fresh sample from the last 7 days of real prompts. That means comparing benchmark records across months mixes up "the setting changed" with "the sample changed" — no way to tell which caused a difference.

So I picked one day's worth of samples and locked it in as a fixed baseline going forward. If there's ever a real reason to change the sample, I'll version it explicitly instead of letting it drift.

The automated run itself was stuck

Once the settings and sample were sorted, I tried running the tool unattended for the first time.

The first attempt just went silent. Tracing it back, the automated launcher had shortened the command it was passing, which tripped a permission check and left it stuck waiting for approval forever.

If a person had been watching, they could've just approved it. But this was meant to run with no one watching, so nothing could approve it.

Even if that shortened command had somehow gotten through, it likely would've failed anyway — running in the system's default environment instead of this project's dedicated one, missing components it needed.

I fixed it by hardcoding the command so it can't get shortened.

Along the way I also discovered that one of the time windows where extra GPU headroom is allowed had been registered while yesterday's sampling bug was still live. I closed that window temporarily until I can confirm the fix actually holds.

Only after clearing all of that did the benchmark actually get through its first launch step. It ended up colliding with other work already using the same resources, though, so I pushed the run back and rescheduled it rather than forcing it through.

Also closed out a set of safety items that had sat for 11 days

A few days ago I found an issue with a data-collection safety mechanism and got an AI's advice on how to fix it.

During a routine review today, I found that only the most urgent piece of that advice had actually been implemented — the rest had sat untouched for 11 days.

I had the remaining items re-checked for whether they still applied. One turned out to be a real, recurring problem. Another looked like the same category on paper, but measuring it showed it had zero actual impact.

I implemented everything that still applied, all in one sitting. Engine crashes and external network failures are now handled as separate cases, and a data-collection failure now raises immediately instead of silently falling back to another method.

While I was in there, I also added a safeguard to stop test code from accidentally touching processes that are actually running in production.

Nothing exotic caused the 11-day gap — there was just no process for tracking advice items once they'd been given. So today I also did a full pass over the whole outstanding-items list to check whether anything else had slipped through the same way.

What's next

The postponed first benchmark run is rescheduled for the next automated window.

The thing to confirm this time: that the automated command actually goes through with the full path intact, and that the run completes end to end.

Top comments (0)