DEV Community

Cover image for pytest Failed. Every Collected Test Still Had a Passing Run.
John
John

Posted on Originally published at hexisteme.github.io

pytest Failed. Every Collected Test Still Had a Passing Run.

Originally published on hexisteme notes.

I had a red pytest command and a folder of green results. Choosing a color for the handoff would have thrown away information.

The run that attempted everything returned exit code 1. After a carefully bounded retry, the records also showed a successful execution for each of the 1,269 named cases on the collection roster.

Neither observation cancelled the other.

This happened during a regression check dated 2026-09-27. The lesson is not a trick for rescuing red CI. It is about writing a completion statement whose nouns still mean what the evidence says.

Start with the question behind the green label

"Did this command finish successfully?" asks about a process. "Which cases have a successful result?" asks about a roster.

Those questions often receive the same answer, so a dashboard can make them look interchangeable. Once retries enter the picture, they can diverge.

pytest documents the meanings of its exit codes. A result assembled afterward cannot retroactively change the value returned by the shell.

That gave me a useful constraint: keep the command failure visible even if a later attempt supplied the missing results. Otherwise, someone reading the handoff could mistake an accounting result for an observation of everything working together.

The distinction is easier to keep when the report contains both answers, instead of forcing the reader to infer one from the other.

There was an environmental cause to investigate

My initial attempt produced 1,199 passes alongside 31 skips, 3 failures, and 36 errors. The command returned 1. It had not been given the browser-runtime setting.

An already installed, pinned Playwright runtime was available. I selected that cache and tried the complete roster again inside the restricted environment. This removed the skips, but it did not make the invocation successful: the results were 1,199 passes, 3 failures, and 67 errors, with exit code 1 again.

At that point, calling the problem "environmental" still did not close it. I needed to find the operations being denied and establish what a retry would actually exercise.

The browser component accounted for 69 cases. macOS denied its launch through MachPort. The other case belonged to an explorer fixture whose temporary web server was unable to bind localhost.

Changing that boundary meant changing the conditions under which evidence would be collected. It deserved a precise scope, not an incidental mention buried under the final green number.

Retry the named group, not the permission model

I selected the unsuccessful cases by their exact node IDs. Before permitting another attempt, I inspected the fixture behavior associated with that selection.

The browser consumed fabricated pages held in memory. Fixture responses supplied the content; requests that would leave the test were stopped. The explorer served disposable data from the machine. This selection involved no real sign-in, no fetching through a production API, no use of actual credentials from Keychain, and no publishing.

The restricted retry reproduced the problem. The authorized retry then allowed only the necessary browser launch and local bind operations.

Here is the full sequence. A pass total is not the process result; the right-hand column preserves that difference.

Attempt Test outcomes Process exit code
Before selecting the installed browser cache 1199 passes; 31 skips; 3 failures; 36 errors 1
Complete roster with the cache selected 1199 passes; 0 skips; 3 failures; 67 errors 1
Selected group inside the original restrictions 0 passes; 0 skips; 3 failures; 67 errors 1
Identical selection with bounded launch and bind permission 70 passes; 0 skips; 0 failures; 0 errors 0

No assertions were softened. No test was deleted from the selection. No new browser installation was needed, and the broad command did not receive blanket permission.

The last row answers a specific question: these 70 cases worked when those local operations were allowed. It does not answer whether they would work inside the restrictions that produced the previous row.

That qualification belongs next to the outcome. A retry without its changed conditions is an incomplete observation.

The sum was a clue, not the reconciliation

Adding 1,199 and 70 gives 1,269. That arithmetic would still be true if both groups included the same case and another case never appeared.

I could not infer membership from totals.

pytest accepts individual node IDs, including selections for parameterized cases. It also supports JUnit XML output. Those mechanisms supply useful inputs, but they do not decide what a combined result means.

The roster defined what needed a result. I associated each JUnit entry with a specific member, treating variations of a parameter as distinct cases rather than interchangeable names. Each outcome still pointed back to the attempt that produced it.

The final comparison left 0 expected members without a pass, 0 deliberately excluded members, and 0 members covered only by a skip. In particular, the initial 31 skips earned no credit merely for having appeared in the output.

A conceptual sketch is short:

missing = collected_ids - passed_ids
extra = passed_ids - collected_ids

assert not missing
assert not extra
Enter fullscreen mode Exit fullscreen mode

This is explanatory pseudocode, not the parser I ran. Before any set operation, the mapping must reject ambiguous identifiers and duplicate entries within an attempt. Converting a questionable list to a set can conceal the very defect the check is supposed to catch.

I wanted an answer to "where is this case's successful execution?", not just a plausible total at the bottom of a log.

A commit name would have described the wrong object

The checkout was not pristine when this verification began. It already included two changes made concurrently with the surrounding work.

If I had attached only a commit ID, a later reader could reasonably assume the measurements referred to that committed tree. They did not. The object under examination was the checkout as observed, including those modifications.

The receipt included a fingerprint for each of 220 Git-tracked files, plus 7 markers used in operation. Comparing those fingerprints found no drift between attempts or when I checked the handoff.

I did not alter application code or test code while resolving this verification. That statement is compatible with pre-existing modifications; it is not a claim that the worktree was clean.

Nor do matching source bytes make the environments identical. Assigning the browser cache and permitting local launch and bind were real changes in execution conditions. A source fingerprint is useful precisely because it answers a narrower question, not because it proves every kind of sameness.

This also limits reuse. A different checkout or changed dependency cannot inherit the result just because the old report is still nearby.

What I could say, and what I could not

The reconciliation gave every expected case a successful execution. It did not show the entire roster succeeding within a single process.

A fixture may leave something behind for its neighbors. Scheduling or resource competition may change when cases share a process. Separate invocations can miss that behavior; this record does not measure it. Nor did I measure the likelihood of seeing the failures again.

The denominator is a roster of named cases, not lines of executable Python. A statement about which cases succeeded cannot tell me which lines were exercised.

And a synthetic browser fixture is not a live-account check. Permission to launch a browser for local content does not become permission to exercise credentials or publish something.

I have written about a neighboring problem in Green Tests Prove Behavior, Not Reachability: a gate can behave correctly in isolation while being absent from the dispatch path. This case concerns a different object. I had an explicit roster, and needed an execution record for each member of it.

A Check That Never Ran Is Not Passing addresses a missing evaluable input. Here, the successful observations already existed. The work was to reconcile them without crediting omissions, or claiming that a changed environment was the original one.

Those distinctions help keep similar-sounding failure stories from turning into the same generic lesson.

The handoff needs more than a status word

The custom result in my receipt was FULL_COLLECTED_COVERAGE_BY_EXPLICIT_SHARDS. Alongside it, whole_pytest_invocation_exit_zero stayed false.

That is a description of this receipt, not a pytest feature or a reason to override the shell.

For a future report, I would reserve room for the process outcome, the expected roster, the evidence attached to each member, the bytes examined, and the conditions changed by a retry. I would also state which decision the report supports.

A verification statement and permission to merge or release are not interchangeable. Neither should disappear behind a green badge.

This is a proposed reporting habit, not a claim that I have installed a general-purpose reconciliation framework across every project. Its useful property would be refusing completion when an identifier cannot be resolved or an expected result is absent.

Preserving the earlier red attempts matters too. I retained the original logs and XML in a lossless archive and checked them against their hashes. A readable summary should sit beside those artifacts, not replace them with a tidied history.

Keep a way to revoke the conclusion

The result stops being defensible if a required member has no successful execution, if an identifier points to more than one case, or if I discover that the observations came from different checkouts.

The same applies if the supposedly local retry actually used an unauthorized production operation or live account. Its permitted scope is part of what makes the evidence usable.

A later all-green invocation would be welcome additional evidence. It would not erase the red invocations already recorded, and these separate attempts do not authorize relabeling a failed CI job.

The handoff I could defend was simple:

The complete command returned failure. The named cases all had successful executions under the recorded conditions. The record preserves both facts.

Prepared with AI assistance from my preserved regression-verification records; source and execution evidence rechecked on 2026-09-27.

Email list for these notes: hexisteme.beehiiv.com — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.

More notes at hexisteme.github.io/notes.

Top comments (0)