I keep replaying a design-review counterexample from a grading pipeline that looked correct until one late witness arrived. The deterministic harness had already recorded a failure for compensation step S7 under the current fixture generation. A slower judge lane, started against the previous fixture, then returned a pass and reached the closer second. Would you seal that compensation step only because the passing witness happened to be the last write on the wire?
That closer did exactly that, and the skipped compensation left an in-doubt tool call looking finished. I am not citing a production dashboard, a customer incident, or a measured error rate for this walkthrough. I am reviewing the architecture as a state machine, because a pretty sequence diagram will not catch a reordered seal. If the last write can erase an earlier failure, the workflow does not converge, it only looks quiet.
The invariant a common closer drops
The invariant is simple enough to test, and it is still easy to violate under at-least-once delivery between lanes. A witness may seal compensation only when its generation and fixture hash equal the current fence. An older generation must be rejected even if its verdict arrives later and sounds more confident. A matching failure must dominate a matching pass, because skipping compensation is the dangerous direction for this step.
Last-write-wins preserves neither the generation fence nor the fail-dominates-pass rule once two lanes share a closer. It treats delivery order as authority, which is an accident of queues rather than a property of the step. I would not call that pattern a merge, and I would call it an unguarded assignment into the seal log. Have you ever watched a queue make the stale message look like the freshest decision in the log?
Assumptions I will state before the diagram
I assume every lane delivers at least once, and that duplicates can arrive after a seal has already been attempted. I assume the orchestrator can bump a fixture generation before new work starts, and that this bump is durable. I do not assume a shared clock between the worker process and the model lane that judged an older fixture. I also do not assume that a textual pass from a model is equivalent to a property-test failure from the harness.
MonkeyCode is relevant here only as capacity for a non-authoritative lane: the operator describes free model access and a free server option. Disclosure: This article was prepared as part of MonkeyCode's product outreach. I am not claiming a quota, a model name, a hardware shape, a duration, or a measured comparison. If you stand up a non-authoritative judge, read the current project terms first, then keep that lane outside the seal path.
Constraints, data flow, and failure domains
The orchestrator owns the fence, the step identity, and the compensation log that later records the seal. The harness domain owns deterministic checks that must run against the published fixture hash for that same generation. The model lane owns a slow judgment that can already be stale the moment a fixture generation bumps. The free server domain owns a worker that can retry, reorder, and deliver twice after an acknowledgement is lost.
sequenceDiagram
participant O as Orchestrator
participant H as Harness lane
participant M as Model lane
participant C as Closer
O->>O: Publish fence gen=7 hash=fx-7
O->>H: Start check S7
O->>M: Start judge still on gen=6
H-->>C: Witness fail gen=7 seq=1
M-->>C: Witness pass gen=6 seq=2
C->>C: Reject gen=6, seal fail, compensate once
What would I change in that picture if the closer still merged witnesses by delivery sequence number? I would delete any arrow that lets the model lane touch the seal, and I would force a reject branch first. The free server worker may host the harness, but it still may not invent a generation or a fixture hash. Capacity and authority stay in different domains, or the retry budget of one domain will close the other.
Numbered steps for the seal protocol
- Publish the fence in the same log that will later record the seal, and refuse new lane starts until that publish is durable.
- Bind every witness to the step id, generation, fixture hash, lane id, and a content hash of the inputs it actually saw.
- On delivery, reject generation or hash mismatches without enqueueing compensation and without advancing the step to closed.
- If an accepted failure exists, seal fail once, then run idempotent compensation keyed by step id rather than by delivery sequence.
- Treat a duplicate accepted witness as a no-op, and replay only when the generation still matches and the witness is missing.
I would not replay an older generation hoping the model will eventually change its mind about a fixture it never saw. That replay spends budget to reproduce a stale pass, which is the opposite of convergence under a fence. Rejection is the correct response to a fence miss, even when the late message is fluent and internally consistent. Compensation is the correct response only after an accepted failure, and only through a key that ignores delivery sequence.
A simulator that shows the break
The following fixture is a proposal you can run locally, and it is not a measurement of any hosted service. I have not used it as a load test, and it does not establish a latency, throughput, or cost claim. It only encodes the event order above, plus one duplicate delivery and one foreign fixture hash. Read the assertions before you trust the printed lines, because a printout alone is not the property.
from dataclasses import dataclass
@dataclass(frozen=True)
class Witness:
lane: str
generation: int
fixture_hash: str
verdict: str
delivery_seq: int
@dataclass(frozen=True)
class Fence:
generation: int
fixture_hash: str
def last_write_wins(witnesses):
winner = sorted(witnesses, key=lambda w: w.delivery_seq)[-1]
return {"sealed": True, "verdict": winner.verdict, "lane": winner.lane, "reason": "lww"}
def seal_with_fence(fence, witnesses):
accepted = [
w for w in witnesses
if w.generation == fence.generation and w.fixture_hash == fence.fixture_hash
]
if not accepted:
return {"sealed": False, "verdict": None, "reason": "reject"}
if any(w.verdict == "fail" for w in accepted):
return {"sealed": True, "verdict": "fail", "reason": "accepted_fail"}
return {"sealed": True, "verdict": "pass", "reason": "accepted_pass"}
def main():
fence = Fence(7, "fx-7")
late = [
Witness("harness", 7, "fx-7", "fail", 1),
Witness("model-lane", 6, "fx-6", "pass", 2),
]
buggy = last_write_wins(late)
fenced = seal_with_fence(fence, late)
assert buggy["verdict"] == "pass"
assert fenced["verdict"] == "fail"
duplicated = late + [Witness("harness", 7, "fx-7", "fail", 3)]
assert seal_with_fence(fence, duplicated) == fenced
foreign = [Witness("free-server", 7, "fx-other", "pass", 4)]
assert seal_with_fence(fence, foreign)["sealed"] is False
print("lww", buggy)
print("fence", fenced)
if __name__ == "__main__":
main()
Save that listing as witness_fence.py, then run python3 witness_fence.py from the same working directory you saved it in. You should see the last-write-wins path seal a pass, while the fence keeps the harness failure and rejects nothing incorrectly in this case. If the buggy assertion ever stops matching a pass, you have probably sorted by generation instead of delivery sequence. That change would hide the counterexample rather than fix the closer that still assigns the last write.
Injected failures and the acceptance rule
I inject four failures, and each one has a deny or accept outcome you can assert without waiting on a timer. A reordered older pass must not change a sealed fail, even if its delivery sequence is the largest number in the batch. A duplicate accepted fail must not start a second compensation, because the compensation key is the step id alone. A hash mismatch at the current generation must hold the step open, and a missing current witness may be replayed once.
The acceptance rule needs an explicit denominator, or a green run will not mean the closer actually converged. I define it as sealed decisions whose generation and hash match the fence, divided by seal attempts in that same fixture run. The numerator excludes rejects, because a reject is a successful protection rather than a sealed compensation step. I would not accept a design that seals any mismatched witness, even if the ratio looks high because most messages are fresh.
Tradeoffs I would put in the review
| Choice | Correctness under reorder | Extra latency | Cost | When it fails |
|---|---|---|---|---|
| Last-write-wins | Breaks when an older pass arrives late | Lowest | Lowest | Stale model lane overwrites a harness fail |
| Reject on fence miss | Holds the invariant | One comparison | Tiny | Operators may retry blindly and add load |
| Replay only on a generation match | Fills a missing current witness | Another lane attempt | Bounded by one replay | Useless if the generation is already old |
| Compensate on an accepted fail | Closes the dangerous case once | Compensation path | Side-effect cost | Unsafe if compensation is not idempotent |
Rejecting a stale pass costs almost nothing and avoids a wrong seal that would skip compensation for step S7. Replaying every mismatch looks thorough, and it quietly burns the retry budget of whichever lane is cheapest to call. I would rather hold the step than compensate from a witness that never saw the current fixture hash. Would a slower seal annoy a status page more than a skipped compensation annoys the ledger you cannot easily unwind?
What I would change next
I would split the seal log from any worker that is allowed to scale just because spare capacity is available. The next change is a single writer for seals, with lane processes limited to append-only witness records that cannot close a step. I would also store the input hash beside the fixture hash, so a matching generation still fails closed when the snapshot drifted. After that, I would add a property test that shuffles delivery order and asserts the seal verdict against the fence.
I would not add a second closer for availability until the fence shares a transaction with the compensation key. Two closers that still use last-write-wins simply race each other into the same unguarded assignment on the seal log. A cheaper worker makes that race easier to schedule, which is a reason to be stricter about authority. Would you really scale the closer just because the worker lane became inexpensive to start on spare capacity?
Who should leave this pattern alone
Do not use this fence if the model call itself performs the side effect you might later need to compensate. A judgment witness is not a tool receipt, and rejecting the judgment will not undo a call that already committed. Do not use it if you cannot publish a durable generation before lanes start, because the fence would then be only a comment. Teams that need a non-authoritative lane to share a write key with the canonical store should redesign that boundary before they tune retries.
This simulator also ignores partial network partitions inside a single harness process, and it will not tell you a latency budget. If your review meeting expects a throughput multiplier, this architecture note will disappoint you on purpose. I would rather ship a counterexample with an acceptance rule than a diagram that cannot fail a test. I am also not claiming that free capacity remains available, permanent, or sized for your workload.
The question I want left on the table
Which event order breaks the invariant, and should the system reject, replay, or compensate for it? Take the harness failure at generation 7, then deliver a model pass from generation 6, then deliver that same harness failure again. I would reject the pass, compensate once from the accepted failure, and treat the duplicate failure as a no-op. If your closer still seals the pass, the generation is not an invariant yet, and spare capacity will only make the race cheaper.
Top comments (0)