DEV Community

finaltype
finaltype

Posted on Originally published at finaltype.github.io

"[Sep 18] Benchmark Harness — Two Measurement Bugs in the Measuring Tool Itself"

The tool was overcounting GPU usage by more than 3x, and a too-short measurement timeout was flipping candidate rankings upside down. Found and fixed both in one day.

This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).

GPU usage was inflated by more than 3x

The benchmark I'd scheduled last night ran automatically overnight.

Checking the results, something looked off. One trial alone was recorded as having burned through more than 80% of the day's GPU usage budget.

My first read was that the server had died and the client kept retrying against it, burning time in the process.

Then came the question: "where does that number actually come from?" Looking again, that explanation turned out to be a guess with no evidence behind it — the server logs showed no retries at all.

The real cause was different. The benchmark had been summing up the time spent processing several stocks in parallel as if they'd been processed one after another, and then recording that inflated sum as actual GPU occupancy time charged against the budget. It had been overcounted by more than 3x against the real figure.

After getting another AI's help walking through the math to confirm the root cause, I rewrote the accounting to use the actual elapsed time between start and end, and corrected the past records that had been miscalculated. While in there, I also cleaned up duplicate entries that had crept into the sample data.

The rankings themselves were backwards

After fixing the accounting, I ran a few configuration variants against real GPU hardware to compare them.

One variant came out slower than expected. Another came out faster, but with a higher failure rate. A higher failure rate is a reasonable reason to rule out a candidate, so that's how I initially filed it.

That's where the second measurement bug showed up. Asked whether I'd actually confirmed from the logs that those failures were server-side problems, I went back and checked — every one of those "failed" requests had actually been generating a normal response at the time it was cut off.

The problem was that the benchmark tool's own wait timeout was set far shorter than what production actually uses. It had been force-killing perfectly healthy in-progress requests just because its own clock ran out.

Once I matched that timeout to the production value and started counting the time spent on cut-off requests toward the total, the rankings flipped. The variant I'd marked "slower" turned out to actually be faster, and the one I'd marked "faster" turned out to be the slowest of the bunch.

Results with a low completion rate are now flagged separately and automatically excluded from comparison.

Twice in the same day, the measuring tool itself turned out to be measuring wrong. Both times the numbers looked plausible on the surface — without going back to check the actual evidence behind them, both would have gone unnoticed.

Also traced a 9-hour stall to its root cause

Separately from the benchmark work, the previous night's stock-analysis batch had stopped after finishing only 75 of 100 stocks.

Tracing the logs, it wasn't "progress trickled along all night" — one of several parallel branches had gotten stuck for over 9 hours on the very first stock, stalled on an external data lookup with no timeout set at all. The other branches all finished normally.

I added a forced timeout to that external lookup, plus a self-monitoring safeguard so that if a branch makes no progress for a set period, it hands off its remaining work and exits on its own. Both apply starting tonight's run.

What's next

I'm continuing the benchmark with the corrected accounting and timeout in place.

The next thing to confirm is whether tonight's automatically resumed round actually shows the corrected values, and whether it gets judged purely on its own numbers rather than compared directly against the earlier, miscounted round.

Top comments (0)