A test result is not a transferable compliment. It is a statement about the thing the test actually ran.
If CI tests a build from commit A and a release process later rebuilds commit A on another runner, the second output may be functionally equivalent. But the first result did not execute that second output. If the pipeline publishes an image under a tag and later moves the tag to a new digest, an earlier scan does not automatically qualify the new image.
That is the difference between “we tested the source” and “we tested the artifact users can download.”
Bind each result to an identity
For a release artifact, useful identity usually includes at least the source commit, workflow run, build environment, output filename, and cryptographic digest. For a container, include the immutable image digest; a friendly tag such as latest is a moving reference.
The identity must travel with the result. Otherwise, an operator can easily attach a green check to the wrong build: a retry, a rebuild, a different platform, or an artifact uploaded after the test completed.
A digest answers “which bytes?” It does not answer “are these bytes safe?” The test or audit result answers a different question, and its record should name the digest it evaluated.
The v1.29.0 outcomes were not one result
WorldScript Studio’s v1.29.0 tag is signed and points to commit cf72dc6…. Its associated workflows then ended differently. The CI/security audit run failed. The Tauri release run was cancelled before it created a GitHub Release. The Docker run succeeded and published a container identified by digest sha256:1f463bc6ba1e2544b9d8c2191814c3eff50c9add81cd81d3d3927be736b3fe2d under the 1.29.0, 1.29, and latest aliases.
Those facts cannot be collapsed into a single “release passed” or “release failed.” The tag exists. One publication surface succeeded. Another was cancelled. GitHub Release assets were not produced. Each outcome belongs to its own artifact and publication path.
If the Docker image is rebuilt later, the new digest is a new artifact. Its scan or runtime test should be recorded against that digest. The earlier result does not silently follow the friendly alias. The v1.29.1 successor shows the rule: candidate 99a664c5 failed the newly executed Security Audit and was not tagged; after #912, candidate f255d767 was requalified and separately published. Results did not transfer between candidates.
Rebuilds need their own evidence
Reproducible builds can help determine whether two builds yield identical bytes, but a second build is still a separate event. If the digest matches, the evidence can be related explicitly. If it differs, investigate the difference and qualify the artifact that will be distributed.
This is why “test the artifact you built” is useful even when the build is deterministic. The question is not philosophical equivalence; it is traceability. Can the release record show which exact artifact passed which checks, and can an operator retrieve that artifact later?
Keep the evidence planes separate
A source review does not qualify a native bundle. A successful container scan does not qualify desktop installers. A signed tag does not mean a GitHub Release exists. An artifact can pass one test and fail a later persisted-state check.
A good evidence record therefore has rows for source SHA, artifact digest, platform, test or audit, workflow run, result, and publication destination. If the artifact is replaced, add a new row. Do not rewrite history so that yesterday’s check appears to have run against today’s bytes.
The practical rule is simple: record the identity before testing, pass that identity into the check, and publish only the identity whose required checks passed. When rebuilding, treat the output as a new evidence subject until equivalence is demonstrated and recorded.
Continue reading
Next: Release State Machines: PREPARED Is Not PUBLISHED applies artifact identity to the separate states of a multi-surface release.
Related: The Tests Were Green. The Mutants Survived. asks whether a passing suite can detect a changed behavior; Tests Prove Behavior. Boundaries Prove Architecture. connects test evidence to system boundaries. The broader preservation map is in Anatomy of a No-Loss Persistence Campaign.
Practical extension: keep an artifact evidence record
For every deliverable, capture a compact identity tuple:
| Field | Example of what it answers |
|---|---|
| Source commit | Which repository state was built? |
| Workflow run and revision | Which automation produced it? |
| Artifact name, digest, and platform | Which exact bytes and target are under discussion? |
| Qualification environment | Where and under what profile was it installed or executed? |
| Result and limitation | What passed, failed, or was not exercised? |
This record prevents evidence from drifting between a tested build, a signed package, a container tag, and a later rebuild. A checksum binds a name to bytes; it does not establish who built them. An attestation can add provenance about where and how an artifact was produced, but it still does not prove the application behaves correctly after installation. Keep provenance and runtime qualification as separate evidence planes.
GitHub documents artifact attestations for binaries and container images, including subject digests and workflow permissions: establishing build provenance. Even with attestations, record the installer or image digest that was actually exercised. The correct question is not “did we test this version?” but “which exact bytes did this environment run, and what observation did we make?”
Top comments (0)