DEV Community

Haoxiang Li
Haoxiang Li

Posted on Fully Autonomous

My LLM Eval Reused an Old Result After the Input Changed

The old runner returned alreadyCompleted after a character's privateGoal changed.

The case filename, case ID, candidate version and repetition count were unchanged. The saved plan still matched. The runner returned the previous completed manifest instead of evaluating the changed case.

This was a deterministic regression test in jzj, an AI board-game project. The test supplied a synthetic completed manifest and blocked credential access and network calls. It reproduced a real resume bug; it did not establish that an earlier paid result had been contaminated.

The failure came before any question about the model's answer. The evaluator had identified the wrong experiment.

A case name was doing the work of an input identity

The fixed-case runner restores a recorded seat view and prior character state, then asks an actor to make a decision. It saves the request and result so an interrupted evaluation can resume without paying for completed samples again.

That reuse depended on stable labels: the case-file name, case IDs, candidate prompt version, evaluation arm and repetition. The plan check existed, but the case content was absent from the plan.

The completion branch looked like this, reformatted from the implementation:

if (manifest?.status === 'completed') {
  console.log(JSON.stringify({ alreadyCompleted: manifestName }));
  return;
}
Enter fullscreen mode Exit fullscreen mode

Returning here is reasonable when the saved work belongs to the requested experiment. It is wrong when an unchanged name conceals a changed input.

For a character-background comparison, that matters directly. Two versions of a case can share the same game position and case ID while differing in the character's background. Reusing the old result would erase the intervention before the model received it.

The regression kept the filename and ID fixed, changed privateGoal, and attempted to resume. With the old runner, the completion shortcut was taken. With the repaired runner, the changed plan was rejected before the paid-session code could run. The previous manifest remained byte-for-byte unchanged in both checks.

Offline reproduction Old runner Repaired runner
Changed input returned alreadyCompleted Yes No
Plan recorded frozen input digests No Yes
Different input identity rejected No Yes
Existing manifest preserved Yes Yes
Credential/payment gate or network reached No No

The old side is the actual pre-fix runner loaded from Git, not a mock of its completion logic. Both sides use the same synthetic completed-manifest setup.

Freeze once, then carry the identity through resume

The fix builds each fixed case once, computes a digest, and keeps a cloned copy for subsequent requests. These are excerpts from the local implementation:

export function caseInputDigest(value) {
  return createHash('sha256')
    .update(canonicalStringify(value))
    .digest('hex');
}
Enter fullscreen mode Exit fullscreen mode
const input = await build();
const inputDigest = caseInputDigest({ id, factor, ...input });
return { id, factor, input: structuredClone(input), inputDigest };
Enter fullscreen mode Exit fullscreen mode

Canonical JSON sorts object keys, so a change in key ordering does not create another experiment. Changes to values still change the digest; array order remains significant. Here, “frozen” means built once and cloned for reuse, rather than recursively locked with Object.freeze.

The digest enters the plan and the sample ID:

export function evaluationCellId({
  promptVersion, phaseIdentity, scenario, armId, repetition,
}) {
  return `${promptVersion}:${phaseIdentity}:${scenario.id}:${scenario.inputDigest}:${armId}:r${repetition}`;
}
Enter fullscreen mode Exit fullscreen mode

The request filename, saved request object and result row carry it too. When the manifest is missing, recovery from JSONL rows requires both the expected sample ID and its matching caseInputDigest. A row with the right ID but the wrong digest is excluded. A row without a digest is excluded as well.

Old artifacts stay intact. A completed manifest without the new identity is not silently upgraded into evidence for the new run.

This digest identifies the frozen evaluation case bundle. It includes the seat view, restored state and evaluation metadata such as source and review annotations. It does not hash every assembled model-request field or every prompt implementation. Changing a review criterion can invalidate this evaluation identity even when the model-facing request would be unchanged.

The saved request remains necessary to inspect what the model would actually receive. Review annotations and archived expected outputs stay on the evaluation side; they are not supplied as answers to the actor.

The locator needed a regression too

A separate source audit found that two frozen historical decisions pointed to attempt: 1. The harness numbers the first ordinary call as attempt: 0.

The locator was corrected, and the stub replay now checks every selector field against the formal harness diagnostic, then compares its raw JSON with the saved decision. The old selector fails that check for both windows; the corrected selector passes. No archived model response was rewritten.

A content digest and a source locator answer different questions. The digest identifies the case being evaluated. The locator identifies which recorded output supports the claim. Both need checks.

What passed, and what did not run

The related offline suite passed 15/15 tests with no skips. It covers changed dice, role, prior state and review criteria; canonical key ordering; the completed-manifest shortcut; JSONL recovery; and the frozen evidence checks.

node --test test/server/role-evaluation-input.test.js \
  test/scripts/role-evidence-cases.test.js \
  test/scripts/role-evidence-review.test.js
Enter fullscreen mode Exit fullscreen mode

Two archived suspect decisions were also fed through the formal harness using a stub that returned their saved JSON. Their actions remained legal and their saved statements remained unchanged. That demonstrates reproducible evidence and separates parser acceptance from factual review. It does not demonstrate a model correction.

Two character-background cases and one mechanical control are prepared but have no new model outputs. The background pair changes only the character's private background while keeping the recorded position fixed. Neither arm has been executed against a real model. Different dialogue alone would not establish a useful change in decisions.

The implementation and checks are on a local isolated branch. They have not been pushed or deployed; there is no public code URL to attach yet.

My earlier posts examined planner comparisons and an action-contract repair. Here, the result is narrower: changing a fixed case can no longer silently reuse a completed result under the same labels.

Before interpreting a changed model output, the evaluator needs to establish that it evaluated the changed input.

AI disclosure: This article was drafted and edited by AI agents from the project’s code, archived outputs, and verified local test results, with project direction from the author. No new paid model experiment was run for this work.

Top comments (1)

Collapse
 
innokentyb profile image
Kent Bodrov •

This is a good example of why an eval result needs a stronger identity than a test name. I’d bind it to the input hash, evaluator version, environment, policy, and relevant dependency state. The other important part is propagation: once the result is invalidated, which release decisions or downstream reports are no longer safe to rely on?