DEV Community

mars70s
mars70s

Posted on

A Test Passed. But What Did It Actually Test?

A test passed. That sounds reassuring.

But before treating that PASS as evidence, there is a more basic question:

What exactly did the test execute?

It is surprisingly easy to lose the connection between a test result and the implementation that result is supposed to validate. This became especially visible while reviewing a validation record from my own development work.

A PASS is not enough

A test result normally tells us something about a specific target under specific conditions.

At minimum, I want to know things such as:

  • What file or implementation was executed?
  • Which version was it?
  • Which test script produced the result?
  • Under what environment was it executed?
  • Can the result still be tied to that same target later?

Without that information, a PASS may be real while still being weak evidence for the claim we want to make.

The problem is not necessarily that the test itself is wrong. The problem may simply be that the evidence does not establish the identity of the thing that passed.

I initially made a stronger claim than the evidence supported

While reviewing an earlier validation record, I initially interpreted the evidence as a specific failure mechanism.

The working explanation was roughly:

One implementation may have been treated as evidence for another.

That was a coherent explanation, but the surviving evidence did not support it strongly enough.

When I went back through the retained scripts, outputs, filesystem metadata, and validation records, I could confirm one thing: the record itself did not identify the runner, the implementation, or the invocation path behind its PASS results.

But I could not reliably reconstruct the stronger causal details.

For example, the surviving evidence was not enough to establish:

  • what the intended validation target was,
  • whether a retained run script, which executed local functions, was the run that produced the PASS record,
  • whether those local functions were intended as a test double,
  • how the files involved relate to each other (lineage, copy direction, or creation order),
  • or whether AI was involved at all.

Those remained unknown.

So the stronger explanation had to be withdrawn.

What the evidence actually supported

The defensible conclusion was narrower:

The verification record did not sufficiently identify its executed validation target to attribute the PASS to a specific implementation from the record alone.

The finding is not:

"The wrong implementation was tested."

The evidence does not support that statement.

The observable problem was:

"The recorded PASS could not be attributed to a specific implementation from the record alone."

That distinction matters.

Why this matters in AI-assisted workflows

One caveat first: the evidence in this case does not establish whether AI was involved, and a single case says nothing about how often this happens. What follows is a general observation, not a finding from the case.

AI-assisted development can create many intermediate artifacts:

  • generated scripts,
  • wrappers,
  • temporary implementations,
  • test helpers,
  • copied code,
  • diagnostic scripts,
  • regenerated files,
  • and revised versions of the same component.

None of those are inherently bad.

But they increase the number of objects that can exist between:

what we intended to validate

and

what was actually executed.

If the validation record only says:

RESULT=PASS
Enter fullscreen mode Exit fullscreen mode

we know very little.

Even this is better:

TEST=verify-something.ps1
RESULT=PASS
Enter fullscreen mode Exit fullscreen mode

But it still may not tell us which implementation that test exercised.

For stronger evidence, the record needs enough identity information to reconstruct the relationship.

For example:

TARGET:
  path
  version or commit
  content hash where useful

TEST:
  path
  version or hash

ENVIRONMENT:
  relevant tool versions

RESULT:
  PASS / FAIL

BOUNDARY:
  what this result does and does not establish
Enter fullscreen mode Exit fullscreen mode

The exact format is not important. The ability to trace the result back to the tested object is.

Where hashes are used as evidence, capturing them at the relevant execution or validation boundary is stronger than recomputing them later.

PASS should be a bounded statement

I now try to read a PASS as something closer to:

Under these conditions, this identified test produced this result against this identified target.

Not:

The implementation is correct.

And definitely not:

The system is safe.

Those are much larger claims.

A useful validation record therefore contains not only what was observed, but also what remains outside the evidence.

For example:

OBSERVED:
- identified test completed successfully
- expected exit code was observed
- target identity was recorded

NOT ESTABLISHED:
- universal compatibility
- production safety
- inability to bypass the control
- correctness outside the tested conditions
Enter fullscreen mode Exit fullscreen mode

This may make the record look less impressive, but it also makes it much harder to accidentally reuse the result as evidence for something it never proved.

The uncomfortable part: correcting your own story

The most useful lesson from this case was not technical. It was methodological.

When investigating a result, it is easy to build a coherent explanation from incomplete evidence. Once that explanation has a name and several paragraphs around it, it becomes surprisingly difficult to abandon.

But the explanation is not the evidence.

If later review shows that the evidence only supports a narrower statement, the claim should become narrower too.

In my case, that meant replacing a satisfying causal story with a much more limited conclusion about target identity.

That was the right trade.

What a later validation record preserves

A later, separate validation record shows what more explicit identity information can look like. The historical case neither confirms nor invalidates that later record.

For one Git safety harness, the validation record includes items such as:

  • implementation commit,
  • source file SHA-256,
  • test file SHA-256,
  • Git version,
  • PowerShell versions,
  • Gitleaks version (a secret scanner),
  • individual fail-closed and success-path results,
  • and an explicit list of what the tests do not prove.

Here, fail-closed means stopping when a required check cannot be completed instead of treating an unverifiable state as success.

This does not make the harness universally correct or secure.

It does something more modest:

it makes the validation evidence easier to bind to the object that was actually tested.

A small checklist

Before accepting a PASS as evidence, I now ask:

  1. What exactly was executed?
  2. Can I identify its version or contents?
  3. What test produced the result?
  4. Under what conditions?
  5. What claim does this result actually support?
  6. What does it not establish?
  7. Could I reconstruct that relationship later from the preserved evidence?

If those questions cannot be answered, running more tests may not solve the real problem.

The missing piece may be the identity of the verification target itself.


The underlying records for this case are not public, and the review was my own rather than a third-party audit.

Further reading:

Top comments (0)