DEV Community

qnbs
qnbs

Posted on Fully Autonomous

Pin the Evaluator, Not Just the Dependency

“The tests passed on commit abc123” sounds precise. It may still leave the most important question unanswered: which tests, run by which evaluator, against which artifact?

A source revision is one part of an evidence chain. If the qualification job rebuilds the application, tests a different binary, or silently changes its fixture, a passing result may not describe the thing that is about to ship.

Identify both the subject and the measurement

For artifact qualification, a useful chain names each link:

source revision → build workflow run → artifact identity and hash → fixture identity and hash → qualification result

That answers the subject question: what exact bytes were examined?

The evaluator also has an identity. Record the test suite or policy version, runner image, relevant toolchain, workflow revision, configuration, and the action or dependency versions that could affect the result. For AI-assisted evaluation, record the model/provider snapshot when available, prompt or rubric version, tool permissions, and whether the result is advisory or a required control.

Not every project needs a giant manifest for every local test. The detail should match the consequence. A release artifact, security decision, or externally relied-on certification warrants stronger provenance than a quick developer experiment.

A trustworthy qualification links an exact source revision to the build run, artifact hash, fixture hash, evaluator and environment, and recorded result.

Why a moving evaluator changes the meaning of green

Suppose a pipeline publishes build A, then runs the same source through a second job that creates build B and tests B. If the results differ, the published artifact has not been qualified. Matching source commit is useful, but it does not prove identical output.

Likewise, a test suite can drift. A fixture may be edited to match the implementation. A policy prompt can become weaker. A workflow can start with broader token permissions. A package can update underneath a floating version reference.

If the evaluator changes, the result can change even when the code does not. That does not mean the evaluator should never improve. It means an evaluation should say which evaluator produced it.

Provenance is a chain, not a badge

SLSA's provenance model gives teams a vocabulary for describing where an artifact came from and how it was built. GitHub Actions guidance recommends pinning third-party actions to full commit SHAs when immutable references are required, and minimizing token permissions. Those are useful integrity measures, but they are not a complete trust argument.

An immutable action can still contain a bug. A signed attestation can faithfully describe an unsafe build. A pinned model snapshot can be biased or inadequate. A fixed test suite can miss a whole class of failure.

Pinning answers “which one?” It does not automatically answer “is this trustworthy?” A robust chain also asks:

  • who controlled the workflow and its inputs;
  • whether the workflow ran against the intended revision;
  • whether untrusted pull-request code had access to secrets or write tokens;
  • whether the artifact hash matches the one users will receive;
  • whether the test and fixture are independent enough to detect the defect;
  • what a missing, stale, or inconclusive result means.

Freeze for comparison, update under governance

“Pin everything forever” would trade one failure mode for another. Old evaluators become stale. Vulnerable dependencies need upgrades. Models and threat assumptions change.

A better rule is: freeze identity for a specific qualification, and govern changes to the evaluator as their own reviewed changes.

When a test harness or policy changes, compare versions, document why it changed, rerun relevant cases, and record whether old results need to be recomputed. If the candidate artifact changes, qualify the new artifact. Do not carry evidence from one identity over to another because the names look similar.

For AI reviewers, version the contract around the model as well as the model where possible: input scope, instructions, tool access, review rubric, expected output, and the authority assigned to the result. If the provider does not expose a stable model snapshot, record that limitation instead of claiming full reproducibility.

Make the evidence readable to a reviewer

A useful qualification record lets another person answer:

  1. What exact artifact was tested?
  2. Which workflow and evaluator produced the result?
  3. What fixtures or inputs were used?
  4. Which checks passed, failed, were skipped, or were unknown?
  5. Did the exact published artifact receive the evidence?
  6. What changed after the evaluator was frozen?

If the answer to one of those questions is unknown, name the gap. The right state is not a green badge with a footnote hidden in another job.

The evaluator is part of the system

Tests and review tools are software too. They have dependencies, permissions, failure modes, owners, and release cycles. Treating them as invisible infrastructure makes it easy to trust results whose identity has drifted.

The engineering goal is not to create an impossible guarantee. It is to make each claim traceable: this is the artifact; this is the evaluator; this is what it checked; this is the evidence; and this is what remains outside the result.

That completes this Series C path from authority to evidence: the reviewer can help find a risk, a correction cycle can converge on a change, an accepted lesson can become a guard, and the guard itself can be identified and maintained.

References

AI assistance was used to prepare this draft. The human editor is responsible for validating source freshness, exact platform guidance, and the final publication decision.

Top comments (0)