The order that already left the lab
Last Tuesday I replayed an agent evaluation that was supposed to remain inside one synthetic lab scope. The third tool call departed carrying a hostname that I had never written into the fixture contract. The planner had already consumed a tool result naming a company domain that looked production-real to me. Why did that worker request another call before any halt gate had a chance to refuse it?
That sequence is not merely a logging gap, and it is not a prompt-quality issue either. The common implementation records a warning, rewrites nothing, and lets the next attempt proceed anyway unchecked. I still want a stricter invariant than one warning line buried inside a long trace file. No tool attempt may leave the worker unless the bound target is a closed fixture contract.
Assumptions I will not quietly expand
I am reviewing the binding path, not the model weights and not the surrounding pager workflow. I assume tool calls are at-least-once, and a timeout does not prove the call never landed. I assume a config refresh can change the host while the run id itself stays stable. I also assume a free isolated runner is available, but I do not assume its size, tenure, or model list.
Those assumptions matter because a stable run id tempts people to treat every refresh as the same world. If the host can change under a stable run id, identity is not a label on the job. Would you still call that binding closed if only the generation field had quietly moved underneath? I would not call it closed, and the simulator below encodes that refusal as an explicit decision.
Constraints that shape the protocol
The fixture contract has four fields that must fully agree before any tool call is issued. The scope must be a reserved non-routable name, and the host must end in .invalid as RFC 2606 defines. The marker must equal the expected synthetic token, and the generation must match the value the issuer holds. A missing field is not a fixture, and a contradictory field is not a fixture either.
RFC 2606 is the factual anchor for the name constraint, not a product claim and not a benchmark. I am not asserting that a string suffix stops a determined worker on a routed network. I am only asserting that the gate must fail closed whenever that string contract actually breaks. If your worker can already resolve production names, this naming protocol is not sufficient by itself.
Data flow from bind to halt
I draw the path as a short sequence, because a box diagram without an execution path has fooled me before. The planner proposes a binding, the gate classifies it, and only then may the worker touch the target. A late refresh is a new event, not a silent mutation of the previous bound event. I keep that late refresh on the diagram so a mutation cannot hide between the boxes.
sequenceDiagram
participant P as Planner
participant G as HaltGate
participant W as Worker
participant T as FixtureTarget
P->>G: Bind(run, gen, host, marker)
G-->>P: FIXTURE or HALT
Note over P,G: Refresh may swap host, gen unchanged
P->>W: Issue(call, gen)
W->>G: Recheck(gen, host, marker)
alt closed fixture
W->>T: Call(.invalid only)
T-->>W: Fixture result
else live, unknown, or stale gen
W-->>P: HALT or REJECT, no rewrite
end
Read that alt branch as the real execution path, not as mere decoration beside the prose. If the recheck happens after the call, the diagram is a lie and the invariant is already gone. I would rather add one extra hop than discover the live host inside a result payload. Does your current worker recheck that binding, or does it simply trust the planner's first bind?
Failure domains I separate on purpose
I split this system into four domains so a fault in one cannot masquerade as health in another. The planner domain owns intent and may retry, but it does not own permission to touch a host. The gate domain owns classification, and it must not share a mutable config object with the issuer. The worker domain owns each attempt, and it must carry the exact generation it was given.
The target domain is either a closed fixture sink or it is entirely out of bounds. A refresh that edits the host sits in the gate domain, even when a planner thread performed the write. A model completion that mentions a real company sits in the planner domain, and it must not become a tool argument. Why would I compensate by rewriting the host and then calling that same target again?
That compensation hides the bad bind and can double-apply a side effect on the far side. I keep the numbered steps explicit because compensation feels helpful and is how the invariant dies. Reject means that a stale generation is discarded, and never replayed against the newly bound host. Halt means a separate report path must look, while the tool path itself stays fully closed.
Numbered steps to install the halt gate
- Freeze the binding as an immutable value object before the planner is allowed to request a tool attempt.
- Classify that value as fixture, live, or unknown, and treat unknown as halt rather than as a soft allow.
- Compare the issued generation with the binding generation, and reject the attempt when they differ at all.
- Issue the call only on allow, and refuse any rewrite-and-retry path that tries to launder a live host.
- Record the decision against the attempt id so a later report cannot flip an earlier allow.
Replay is the wrong verb once a live or unknown class has already been seen here. I would rather drop the in-doubt attempt than invent a second call that merely looks cleaner. Have you ever watched a quiet retry helper turn a hard halt into a successful-looking trace? That helper is the first change I would delete before I trusted any larger worker fan-out.
A simulator you can execute locally
The following program is a proposal I have not treated as a benchmark, and it carries no latency or quota claims. Please run the file locally with Python on any interpreter that already accepts the syntax below. I use it strictly as the local fixture, not as evidence about any hosted model's accuracy. If your local interpreter rejects the type unions, rewrite those hints and keep every decision identical.
"""Unexecuted until you run it: fixture-binding halt gate."""
from dataclasses import dataclass
from enum import Enum
class TargetClass(Enum):
FIXTURE = "fixture"
LIVE = "live"
UNKNOWN = "unknown"
class Decision(Enum):
ALLOW = "allow"
HALT = "halt"
REJECT = "reject"
MARKER = "synthetic-fixture-v1"
SCOPES = {"lab.invalid", "fixture.local"}
@dataclass(frozen=True)
class Binding:
run_id: str
generation: int
scope: str
host: str
marker: str | None
target_class: TargetClass
def classify(b: Binding) -> TargetClass:
host_ok = b.scope in SCOPES and b.host.endswith(".invalid")
marker_ok = b.marker == MARKER
if host_ok and marker_ok and b.target_class is TargetClass.FIXTURE:
return TargetClass.FIXTURE
if b.target_class is TargetClass.LIVE or not host_ok:
return TargetClass.LIVE
return TargetClass.UNKNOWN
def decide(b: Binding, issued_gen: int) -> Decision:
if issued_gen != b.generation:
return Decision.REJECT
if classify(b) is not TargetClass.FIXTURE:
return Decision.HALT
return Decision.ALLOW
def main() -> None:
base = Binding(
"run-7", 3, "lab.invalid", "job.lab.invalid", MARKER, TargetClass.FIXTURE
)
cases = [
("closed", base, 3, Decision.ALLOW),
(
"host-swap",
Binding("run-7", 3, "lab.invalid", "acme.example", MARKER, TargetClass.FIXTURE),
3,
Decision.HALT,
),
(
"marker-drop",
Binding("run-7", 3, "lab.invalid", "job.lab.invalid", None, TargetClass.FIXTURE),
3,
Decision.HALT,
),
("stale-gen", base, 2, Decision.REJECT),
(
"unknown",
Binding("run-7", 3, "lab.invalid", "job.lab.invalid", MARKER, TargetClass.UNKNOWN),
3,
Decision.HALT,
),
]
failed = 0
for name, binding, gen, expected in cases:
got = decide(binding, gen)
print(f"{name}: {got.value}")
failed += got is not expected
print(f"failed={failed} denominator={len(cases)}")
raise SystemExit(failed)
if __name__ == "__main__":
main()
python fixture_gate.py
After the program prints, that acceptance rule is simple enough for you to check by hand. failed must be zero, and the denominator must stay equal to the number of attempted calls in the suite. I do not accept a pass rate computed over runs, because one run can hide three bad attempts. If you add a case, you must add it to that same denominator before you claim a pass.
Where a free runner fits, and where it does not
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
I mention MonkeyCode only because a free model seat and a free server option can host this simulator off production routes. I am not claiming a token allotment, a hardware shape, a duration, or a named model. Those product facts were not supplied for this review, so I simply refuse to invent them. If those seats exist for you, use them to execute the file above and to score traces for halt versus silent continuation.
The free server is useful only when it cannot resolve or write to the systems named in a bad fixture. A model endpoint can help label whether a completion reported the mismatch or simply stayed silent. That label remains a classification task, not a permission to issue the next tool call itself. I would not send live credentials to that runner merely to see what happens next there.
Tradeoffs against an explicit denominator
I compare four responses to a bad bind, and none of them is free of cost. Please read the table as a design choice, not as a measured latency or cost study. Which row would you actually ship if a host swap arrived during a long evaluation suite?
| Choice | Extra hop | Correctness under host swap | Cost shape | Failure if misused |
|---|---|---|---|---|
| Halt on unknown | One recheck | No silent live touch | Retries wasted on bad binds | Slower suite |
| Warn and continue | None | Live touch remains possible | Cheap until an incident | Invariant dies |
| Rewrite then retry | One more call | Hides the bad binding | Duplicate side effects | Audit trail lies |
| Reject stale generation | Compare only | No replay onto a new host | Drops in-doubt work | Must rebind explicitly |
The denominator for any claim about this gate is attempted tool calls in the injected suite, not wall-clock samples and not model pass rate. I accept the design only when every non-fixture attempt in that suite halts or rejects, with zero allows. A single allow on host swap, marker drop, or unknown fails the review, even if the other cases look clean. I am not publishing a speedup here, because I have not measured one on any workload.
What I would change next
I would split the classifier into its own process so the issuer cannot read a half-updated config map. I would store generation as a compare-and-swap token, not as a field the planner may overwrite after issue. I would add a report queue that is mandatory on halt, and I would keep that queue off the tool path. I would also delete the rewrite helper, because its existence invites the compensation I just rejected.
Those changes stay inside design and evaluation, and they are not a deployment guide or a monitoring rollout. If your next week is about installing agents on hosts, this review will not carry you there. Would a shared mutable config still scare you after the split, or would you trust the process boundary too quickly?
Who should leave this approach alone
Do not use this simulator if your evaluation workers already share a network route with production tool targets. Do not treat a green local run as a certification, an audit, or proof about any hosted model. Do not adopt it if you need incident paging, dashboard setup, or a cluster install guide, because those jobs live outside this review. Skip it also if you cannot tolerate fail-closed halts, since unknown is supposed to stop the run.
Which order should break it?
Which event order still breaks the invariant: a host swap under a stable generation, a dropped marker, or a stale generation issued after rebind? Should that attempt be rejected, replayed, or compensated, and which of those three keeps the live host untouched? I want the answer as a row in the suite, with the same denominator, before anyone fans the worker out. If your answer is compensate, can you show me the attempt that never touched the swapped host?
Top comments (0)