The corrected tool from yesterday resumed automatically overnight, but it launched one trial and quietly froze. I ended up taking over by hand to finish the round.
This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).
Auto-resume quietly froze overnight
The benchmark harness I'd fixed the accounting on yesterday was scheduled to resume automatically overnight.
Checking the logs, the resume itself fired on schedule, and the first trial started normally.
The trouble started right after. A few minutes into the trial, every following command started returning "requires approval." There was no one around at that hour to approve anything, and the session just sat there until it ended.
The root cause: permissions only lasted one turn
Tracing it back, this turned out to be a structural problem. The command-execution permission an automation skill gets is only valid for the turn in which that skill was first opened.
The moment a background task rolled over into the next turn, the permission it had been granted simply evaporated. That's never an issue when I'm watching and approving things interactively — but for an unattended overnight resume, it turned into a dead end.
Meanwhile the GPU sat there uncleaned, occupied and idle, for about 47 minutes.
I found this the next morning, cleared the stuck GPU occupancy, and this time issued the resume command myself, interactively, to carry the round forward.
Taking over by hand to finish the round
Once resumed by hand, all three trials ran through to completion without issue.
Recalculating with yesterday's corrected measurement method, the configuration under review showed no meaningful throughput gain — the difference between repeated runs was close to noise level anyway.
So the conclusion was to keep the current configuration as-is. Confirming that the change wasn't worth making counts as a result in its own right.
I still haven't decided how to fix the structural limitation this exposed in auto-resume. A few candidate approaches are on the table, and until one is picked, resumes will only happen interactively, by hand.
Three reboots, unrelated
The same day also saw three unplanned reboots.
Tracing the cause, it had nothing to do with any of this work — it was a hardware job installing new cooling fans. No GPU-related settings changed, and the environment came back clean after each reboot.
A separate experiment, still running
Separately, I also kicked off a test today comparing two settings for the randomness parameter (temperature) a model uses when generating responses.
A small-scale smoke test passed, so I let the full comparison run continue overnight. Judging the results and making a final call is deferred until after the weekend.
What's next
Deciding how to fix the permission-structure problem in auto-resume is the next task.
I'll also keep watching the temperature experiment's results over the weekend before making a call.
Top comments (0)