DEV Community

Anthony Garces
Anthony Garces

Posted on Originally published at ranex.dev AI-assisted

The Concurrency Test That Proved Nothing

TL;DR: A race test proves nothing unless its operations overlap and the test fails against broken code; here, Effect.all was sequential by default. Before you trust another green concurrency test, start with the Field Notes slice log.

You wrote the race test. You used the function named all. You watched it pass. Then the production bug survives because your “concurrent” work ran one operation after the other. Could this test have failed before the fix?

In this note

SLICE-013 found exactly that trap while repairing a stranded-tool defect in the Ranex harness fork. The test used Effect.all. In the Effect beta used by the harness, that API defaults to sequential execution. Without an explicit concurrency: 2, the test would have passed against the broken code and proven nothing.

The function name was not the concurrency contract.

This warning applies beyond Effect: parallel-looking code is not proof of parallel execution. A race test that cannot lose is not a race test.

The bug lived before the runner decided there was work

The failure was an empty-inbox crash path. The runner returned before reconciliation, leaving a tool projected as running forever.

In the original path, the runner checked whether there was eligible input. No forced run, no steering input, no queued input: return. The reconciliation that marks interrupted tools sat after that guard. A crash with an empty inbox therefore skipped reconciliation entirely. Nothing else came along to look, and the tool stayed projected running.

The immediate fix was a hoist. Reconciliation now runs before the eligible-input guard returns, so an empty-inbox call to run() can recover the stranded tool without scheduling a provider turn. The slice also guarded the normal short-circuit behavior for other input states. Fixing one path should not quietly make every empty inbox start doing unrelated work.

But that only repairs sessions that receive a later run() call. A session nobody calls again remains stranded. The prototype had a reconcile capability, but no caller at startup. Capability without wiring is a polite way to leave the defect in place.

So the completed slice added a startup sweep in the application graph. At process start, it reconciles stranded tools across sessions without scheduling a provider turn. That is the half that repairs the headline case after a crash: nobody calls run(), and recovery happens anyway.

A repair path is not shipped until something calls it under the condition you claim it handles.

Why the first race test was decoration

The race was between reconcile() and run() for one session. Both could read a tool as running and both could publish an interruption, so the completed slice serialized that work per session.

The test matters because duplicate recovery is worse than untidy. The slice requires exactly one durable Tool.Failed event; projected end state alone cannot catch the duplicate because the projector can show the same final status after two published events. The broken case reproduced two events. The fix used a per-session semaphore and the test then observed one.

Calling Effect.all did not make the reproduction concurrent. In this Effect beta, its default is sequential. The source behavior is concurrency ?? 1. Unless the test declares concurrency: 2, one operation finishes before the next begins. There is no overlap. There is no race. There is only a test wearing a race-test hat.

Concurrency is a behavior you must create and observe, not an intention expressed in a function name.

The slice record says the unqualified test would have passed against the unfixed code. That is the standard worth borrowing. Do not only make your repaired code green. Put the test on a broken revision or temporarily remove the synchronization and require it to go red. If it stays green, you have a test-shaped story, not evidence.

Shapes to hunt for before you trust the pass

You can turn this into a short inspection pass. The point is to prove the system reached the dangerous overlap, not to add more concurrency theater.

  1. Read the framework documentation or installed source for the default execution mode. “All” and “join” do not tell you whether work overlaps.
  2. State the exact two operations that can act on the same state. In this slice it was reconcile() and run(), not two calls to run().
  3. Force both operations to pause after observing the vulnerable state and before publishing their result. Then release them together.
  4. Run the test against the unfixed code or remove the lock. Require the duplicate, stale write, or forbidden result to appear.
  5. Count durable events, writes, and external effects. Do not only inspect the final projection or UI state.
  6. Confirm the fix serializes only the reachable surface. A broad lock can hide a test defect while changing behavior you did not intend to change.
  7. Check the startup path. If recovery is supposed to happen after nobody calls the normal operation again, prove the application graph actually invokes it.

The checklist gives you a way to reject a passing test that never created the condition it was named for.

The fix also created a named hazard

The startup sweep recovers abandoned sessions, but it is database-global. A second process booting can mark tools that a first live process is running as interrupted.

That hazard was introduced by this slice, not discovered later and polished out of the story. The record says it is accepted scope because the harness normally runs one daemon. It also states the boundary plainly: before the harness runs as more than one process against one database, the fencing slice must gate the sweep on ownership.

This is where “we fixed it” stops being useful engineering language. The hoist fixes the empty-inbox run() path. The startup sweep fixes sessions nobody calls run() on. The sweep also creates a cross-process ownership hazard. All three can be true at once.

If you remove that last sentence from a release note because it makes the fix look less clean, you are not making the system safer. You are making the next operator discover the precondition by accident.

A named hazard with an owner and a prerequisite is not a victory lap — it is a boundary your system can defend later.

That posture is part of the accountability apparatus behind Ranex: a checkable claim, a record of what it does not cover, and a deterministic gate where authority matters. If that framing is useful, read the accountability apparatus. It is not a claim that Ranex is ready to run for you. Ranex is pre-release, and much of its broader picture remains designed, not built.

What the record actually proves

SLICE-013 closed on 2026-08-08 with all seven criteria met, landing as commit a8bc7bdf35 in anthonykewl20/ranex-harness. It is harness-fork work, as the README’s Current work section describes.

The record proves the unsafe empty-inbox baseline was reproduced, the hoist repaired it, the startup sweep was wired into the application graph, repeated reconciliation produced one durable failure event, and the reachable reconcile-versus-run race was reproduced before serialization closed it. It also records green regression gates, but a green report is not the main point here. The meaningful proof is that the test could fail on the broken condition.

You can read the slice records in docs/slices/done/ in the Ranex repository. The value is not that a repository has a concurrency test — it is that the test’s own execution semantics were checked before anyone trusted its pass.

Questions people actually ask

These answers explain how race tests can pass without creating the overlap they claim to cover.

Why did the concurrency test pass against broken code?

The Effect.all call defaults to sequential execution in the Effect beta used by the harness. Without explicit concurrency: 2, the two operations never raced.

What did the reconciler fix?

The reconciler moved interrupted-tool reconciliation before the eligible-input guard and added a startup sweep for sessions nobody calls run() on.

Why is a startup sweep risky in multiple processes?

The sweep is database-global, so another process starting can mark tools that a first live process is running as interrupted. Session-ID fencing is the recorded prerequisite before multi-process use.

How do you know a race test proves something?

A trustworthy race test runs against the unfixed code and requires failure. The test then starts genuinely concurrent operations using the framework’s documented execution semantics.

Your next test is the one to distrust

Pick the race test in your suite that makes you feel safest. Read the scheduler default. Put the test against the broken code. Count the event that must not happen twice. Then inspect the startup path for the recovery you assume exists.

Leave the new hazard in the record with its prerequisite. Your future self does not need a cleaner story. Your future self needs the boundary.

Try it. Break it. Tell me what broke. If you find a test that passes because nothing overlapped, give the repository a GitHub star and send an honest critique. A critique that finds the false pass is the useful contribution.

Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.

Top comments (0)