DEV Community

Cover image for I Shipped a Green Test That Lied About My Pipeline
Debashish Ghosal
Debashish Ghosal

Posted on AI-assisted

I Shipped a Green Test That Lied About My Pipeline

Scenario S7 of my v0.2.0 field test had a pipeline that "ran end to end." The unit suite was green. The container scenario timed out at 180 seconds.

The pipeline had created every child run, returned each id, and never started a single one. The children sat queued forever while the parent reported success. Everything I asserted — the status code, the persisted record — was true. The behavior I cared about never happened.

That was the moment I stopped trusting the 200.

I found this defect class four times in one field test of HivePlane v0.2.0. Every instance was the same shape: a test that proved the plumbing and not the outcome. Every one was invisible to the unit suite because the unit suite stopped at the seam.

The shape

Four findings. One shape. A test that proves the plumbing and not the outcome.

Finding What the green test proved What was actually true
D-3, pipeline never completes The child run existed (201, persisted) It never transitioned queued → running; nothing started it
D-4, fan-out invisible The delivery hit the webhook sink GET /delivery/audit read a different store; the delivery was unauditable
#558, approval resolve The decision record was persisted The paused run never resumed — resolve() was a no-op lambda
#575, lease reclaim A reassignment was recorded No lock, no leader gate — concurrent reclaims double-assigned the same run

The unit tests stopped at the seam. The bug lived one boundary further.

The pipeline that never started

RunNodeExecutor.submit did this:

run = self._runs.submit(workload=..., pipeline_origin=origin, ctx=ctx)
return run.id
Enter fullscreen mode Exit fullscreen mode

That's it. It created the child and returned. A run is dispatched to its adapter only on the queued → running transition. In the Docker profile there's no worker daemon to lease queued runs, and the pipeline driver only reconciled nodes that were already running. So the parent waited on a child that nothing would launch — a deadlock dressed as a success.

The fix was one line, and it's the line that should have been there:

run = self._runs.submit(...)
if run.state is RunState.QUEUED:
    self._runs.start(run.id, actor="pipeline", ctx=ctx)
return run.id
Enter fullscreen mode Exit fullscreen mode

The regression test — test_pipeline_starts_its_child_runs — asserts the child transitions to running. It failed before, passes after. Why didn't I have it? Because my unit test used a synchronous executor that hid the missing start. The synchronous path never needs a worker, so the absence of one was invisible. Only the live stack, with no worker, exposed it. The Docker report's §3.3 traces the exact transition, and the orchestration design documents the pipeline execution model.

tip: Any component that produces child runs — pipelines, triggers, certification — must either start them or rely on a worker that leases them. If your test executor is synchronous, it's hiding the exact bug you're shipping.

The delivery that succeeded in the dark

D-4 was sneakier. The fan-out delivered. The webhook sink logged the run.completed payload. The attempts were written to fan_out_deliveries with status delivered.

But GET /delivery/audit — the surface my scenario polled — reads the M51 delivery_attempts store. Two subsystems, one concept, two stores. A successful delivery was unauditable through the control plane. My test asserted the 200 and the empty list and moved on, because the empty list was "no deliveries," which was plausible.

The unused _WEBHOOK_SINK constant in the test was the tell. It was checking the wrong surface. The fix wasn't to the delivery — it was to add GET /runs/{run_id}/deliveries so the audit surface references the store that actually receives the data.

tip: When two subsystems share a concept, your audit endpoint must name which store it reads. "Delivery" meant two things, and the test picked the one that was empty.

The approval that recorded a decision and resumed nothing

This one scared me the most. Interactive and mobile approvals — Slack buttons, email links — had a resolve() wired to a no-op lambda:

# the wire that shipped
approval.resolve = lambda decision: None  # persisted the decision, touched nothing
Enter fullscreen mode Exit fullscreen mode

The decision was persisted to approval_decisions. The test asserted the row existed. The test passed. The paused run never resumed. An operator clicked Approve, saw the confirmation, and the run stayed paused forever.

The green test proved a row was written. It did not prove the run moved. The fix: resolve() calls RunService.resume(run_id) — the side effect the operator actually cares about. The regression test asserts the run transitions to running, not that a decision record exists.

tip: The most dangerous "green" test is the one that asserts the record and not the effect. An approval that persists without resuming is a no-op dressed as a workflow. Ask your test: "what did the operator see?" Then assert that.

What I learned

The unit tests stopped at the seam. The bug lived one boundary further — and the synchronous test executor was the reason the seam was invisible. A live stack with no worker, no shortcut, is the only honest witness.

The v0.2.0 field test came back 36/36 scenarios, 67/67 container tests, 50/50 load test — but only because I re-ran it end to end against the live stack and stopped believing my green unit suite. The full receipts are in the field-test report and the Docker test report.

References

What's the most expensive "green" test in your suite — the one that asserts the 200 and the row but not the side effect? Go look. I'll wait.

Top comments (1)

Collapse
 
reidmarlow profile image
Reid Marlow •

Synchronous test executors hide the exact race where state transitions stall. In-memory runners execute the task inline on the calling thread, so the test checks the submission receipt and never touches the lease worker or the queue poller.

The approval case is especially familiar. We had a webhook handler that saved signed approval tokens into Postgres and returned a clean 200. The regression test asserted the token row was active and passed every run. In staging, the background worker picked up the token but died silently because the task payload had an outdated schema from a previous migration. The user saw 'Approved' on their dashboard, the database had the record, and the job never resumed.

Since then, every pipeline test has to wait for a state transition event emitted by the downstream worker rather than asserting against the control plane table.