DEV Community

Cover image for What happens to an agent’s PASS after a dependency upgrade?
Alexander Prokofiev
Alexander Prokofiev

Posted on Fully Autonomous

What happens to an agent’s PASS after a dependency upgrade?

A PASS is a statement about an observation. It is not a permanent property of a package, a service, or a fix.

If an agent reports that a workaround passed with dependency version 3.2, then the dependency moves to 4.0, what should the next agent believe? Deleting the old PASS loses useful history. Reusing it as evidence for 4.0 silently widens what the observation actually proved.

The useful answer sits between those two mistakes: keep the historical result, keep its scope attached, and require new evidence before claiming the new version passes.

In an earlier post about sharing a task between Claude and Perplexity, I wrote about a related boundary: two agent identities under one operator show interoperability, but do not establish two independent reproductions. The follow-up discussion sharpened the problem. Who observed a result, what evidence grounds it, and whether it applies to the current dependency are separate questions.

A version change opens a new evidence question

Suppose a report says a workaround passed with example-package@3.2. The package is now at 4.0, and nobody has rerun the relevant check.

The old result is still meaningful. It says a particular observation passed against 3.2 in a recorded environment. It does not answer what happens on 4.0. Until somebody runs that check, the 4.0 state is unknown. Unknown is not a failure; it is the absence of a new observation.

Here is a small illustrative design, not a real execution record or a current API contract. The identifier-like values are fictional placeholders that bind an outcome to the exact claim and check revisions; they are not hashes and do not prove correctness:

observation:
  id: illustrative-observation-1
  claim_id: illustrative-claim-1
  claim_revision: 1
  check_id: illustrative-check-1
  check_revision: 1
  dependency: example-package@3.2
  outcome: PASS
  environment: <captured environment>
  evidence: <sanitized execution record>
  lineage: executed_here

derived_applicability:
  dependency: example-package@4.0
  claim_id: illustrative-claim-1
  claim_revision: 1
  check_id: illustrative-check-1
  check_revision: 1
  status: NOT_ESTABLISHED
  basis: no observation for 4.0
  related_history: illustrative-observation-1
Enter fullscreen mode Exit fullscreen mode

The second block is a derived applicability view, not an observation: it does not overwrite or revoke the historical PASS, and it records no 4.0 run. If someone runs the check at 4.0, that should create a new observation with its own environment and evidence. A PASS or FAIL then describes that run. Editing the proposed fix or the check creates a new claim or check revision and therefore a new evaluation subject, even if the package remains at 3.2. If two runs disagree under what appears to be the same conditions, record an unresolved conflict and investigate the missing variable. “Newest wins” would hide the question that needs answering.

Three design decisions that keep the record honest

1. Bind an outcome to the thing and context that were tested. A result should identify the dependency or service version, the relevant configuration, and enough environment detail to judge whether another run is comparable. When available, include a build or artifact identifier, source references, and a sanitized record of what was executed. “It passed” without a subject and context is difficult to reuse safely.

2. Preserve old evidence without carrying its scope forward. A version change is a reason to ask a new question, not a reason to erase history or claim the next version failed. Not every meaningful change arrives with a neat version number: hosted APIs, model behavior, authentication policy, regional behavior, and infrastructure defaults can move independently. A date can help locate an observation, but age alone does not establish staleness. Prefer an explicit “not rechecked against this change” state to automatic decay that pretends to know what changed.

3. Keep lineage and disagreement visible. A second agent may have run a fresh test, inherited the first agent’s transcript, or repeated a source without running anything. Those are different paths. Record whether an observation was executed independently or derived from prior evidence, and preserve the source links. If outcomes conflict, show both observations and the unresolved status instead of averaging them into a score or selecting the newest one. There is no universal trust score that can safely replace the details a particular reader needs.

This also shapes how summaries should work. A compact view can say that 3.2 has a recorded PASS and 4.0 has no observation. But the summary is a view over the evidence, not a new reproduction. The canonical record should remain inspectable and revisions should not silently rewrite what an earlier participant said.

“Immutable” needs a boundary, too. Sanitization should happen before sharing, and a system still needs an explicit way to correct or remove material when privacy, rights, or policy require it. Preserving history is not a reason to keep sensitive data public forever.

What the next agent should inherit

After an upgrade, the next agent should inherit the old PASS with its version and environment, not a broader claim that the fix works now. It should also see what changed, whether a new run exists, and whether any result was inherited rather than independently executed.

That leaves the reader with a useful starting point and an honest boundary: this worked under those conditions; the upgraded dependency remains untested. The old work is not wasted, and the new question is clear.

AI disclosure: AI agents drafted and edited this article from the ideas in my earlier DEV post and its discussion. The new text has not been manually rewritten by me; the schema above is an illustrative design, not a report of an executed test.

Cover image disclosure: AI-generated illustration. The v1/v2 labels are generic; the image does not depict the example package versions or a real test.

Top comments (0)