Keep the Stub Gate Before Generated Workers Send Live Traffic
A pairing should keep one fixture gate before any generated job is allowed to call a live model endpoint. The gate runs against frozen inputs and a local stub, then fails closed when network access appears by accident. Shared free model access and a free server option stay outside that gate until a human reviews the same diff. This worked example records the questions, two dead ends, and the single decision the pair kept.
Claims used, and claims left out
Disclosure: This article was prepared as part of MonkeyCode's product outreach. The operator supplied two availability claims for this draft: free model access and a free server option. This exercise does not treat model names, numeric quotas, hardware, duration, or permanence as known facts. Readers should confirm those details with current product documentation before relying on them in a workflow.
The constraint that opened the session
The junior brought a generated worker that summarized error reports and posted results to a remote host. The senior refused to review hosting details until the test story was safe to run on a laptop with the network off. The pair treated the model call as an untrusted dependency, not as a source of expected prose. The session therefore started from written constraints rather than from a generated test that already looked green.
Questions recorded before any code landed
The senior asked four questions and wrote each answer into the pairing note before any code landed. The first question asked whether the same test could pass while outbound network access stayed disabled. The second question asked which fields were allowed to leave the process inside logs or traces.
The third question asked what should fail closed when the endpoint variable was empty or unset. The fourth question asked where, if anywhere, a free server belonged inside this particular code change. The pair answered that a free server did not belong in the test change at all. That answer stayed in the note so a later deploy discussion could not smuggle the host back into CI.
Dead end one: a live call dressed as a test
The first draft asked a code assistant to add a realistic integration test against the live endpoint. That test asserted on sentences in the model reply and printed the full payload when the assertion failed. The senior stopped the draft because prose is not a stable contract and a red failure would leak the report text into CI logs. The pair also refused to spend shared model access on every pull request just to watch wording drift.
Dead end two: a golden file of full replies
The second draft stored a complete model reply and compared each new output byte for byte. That approach died in review because a free model can change wording without changing the fields the worker actually needs. The senior would not invent a score, a latency budget, or a quota number to make the golden file look scientific. The pair deleted the fixture of full prose and kept only the input reports plus a structural expectation.
The decision the pairing kept
The pair kept a fixture gate with three rules and rejected every looser option in the same note. The default transport is a local stub that returns fixed JSON and never opens a socket. Structural checks cover required keys, a maximum length, and the absence of a secret marker string. A live call requires an explicit flag that CI does not set, and the free server is not a test target at all.
How the rejected options compare
A short table made the rejection visible to the next reviewer without replaying the whole conversation. Each row records the network need, the CI stability, and whether the pair kept the option. The table is a review aid for this exercise, not a benchmark against any named vendor. Readers can extend a row only when a new option still passes with the network disabled.
| Option | Network in CI | Stable contract | Kept |
|---|---|---|---|
| Live prose assertion | Required | No | No |
| Golden file of full replies | Required | No | No |
| Stub plus structural checks | None | Yes | Yes |
| Free server as the test host | Required | No | No |
Where the two availability claims fit
Free model access can help a developer draft the worker after the gate already passes on the stub. A natural next check is one manual run of the same structural function against a reviewed client. CI should then return to the stub so shared access is not spent on wording drift. That single manual pass only shows whether required keys survive one real reply during review.
Step 1: Freeze the input fixture
The gate starts from a small JSON file that the pair can read in one sitting. Each case carries an input report, a forbidden marker, and the keys the worker must return. The file does not store a preferred paragraph, because the session already rejected prose contracts outright. A reviewer can add a case without touching the transport code or the live flag policy.
{
"cases": [
{
"id": "timeout-report",
"report": "checkout timed out after the retry budget",
"forbidden_marker": "sk-live-example",
"required_keys": ["summary", "action"]
}
]
}
Step 2: Select a stub and fail closed
Save both Python blocks as fixture_gate.py, and treat them as a local exercise rather than a measured vendor run. It selects the stub unless the caller passes a live flag and a non-empty endpoint string. Empty configuration raises before any socket is opened, which keeps accidental live calls from starting at all. The stub returns only the keys the fixture requires, which keeps the test honest about structure.
import json
from pathlib import Path
class GateError(RuntimeError):
pass
def load_cases(path):
data = json.loads(Path(path).read_text(encoding="utf-8"))
if "cases" not in data or not data["cases"]:
raise GateError("fixture missing cases")
return data["cases"]
def stub_complete(report):
return {"summary": report[:80], "action": "review"}
def select_transport(live, endpoint):
if not live:
return stub_complete
if not endpoint:
raise GateError("live flag set but endpoint is empty")
raise GateError("live transport is operator-owned and not bundled here")
Step 3: Check keys, length, and the marker
The checker does not score writing quality and does not call a model to judge a model. It confirms required keys, rejects unexpected types, and scans only the payload for the forbidden marker. A maximum length stops a generated worker from dumping an entire report into the summary field. Failures name the case id so a pairing note can point at one failing row later.
def check_case(case, transport):
payload = transport(case["report"])
if not isinstance(payload, dict):
raise GateError(f"{case['id']}: payload is not an object")
missing = [key for key in case["required_keys"] if key not in payload]
if missing:
raise GateError(f"{case['id']}: missing {missing}")
summary = str(payload.get("summary", ""))
if len(summary) > 120:
raise GateError(f"{case['id']}: summary exceeds 120 characters")
marker = case.get("forbidden_marker", "")
if marker and marker in json.dumps(payload):
raise GateError(f"{case['id']}: forbidden marker present")
return payload
def run_gate(path, live=False, endpoint=""):
transport = select_transport(live, endpoint)
return [check_case(case, transport) for case in load_cases(path)]
Step 4: Prove the default path never goes live
A reviewer can execute the default path in a clean shell without exporting an endpoint variable. The command below writes a one-line result and does not contact a model or a server. CI should call the same function with the live argument left false on every pull request. The live branch in this draft raises on purpose, so a generated edit cannot silently grow a real client.
mkdir -p fixtures
python - <<'PY'
from fixture_gate import run_gate
rows = run_gate("fixtures/reports.json", live=False, endpoint="")
print(f"kept {len(rows)} structural result(s); live transport not used")
PY
Step 5: Add a negative test for the marker
A second test feeds a transport that echoes the forbidden marker and expects the gate to raise. That failure is the point of the check, because a green run on the stub alone would miss a leaky client. The test file stays local and does not import a vendor SDK or read a secret from the environment. A reviewer runs it with the standard library unittest runner before accepting the generated worker diff.
import unittest
from fixture_gate import GateError, check_case
CASE = {
"id": "marker-case",
"report": "checkout timed out after the retry budget",
"forbidden_marker": "sk-live-example",
"required_keys": ["summary", "action"],
}
def leaky_transport(report):
return {"summary": "sk-live-example", "action": "review"}
class GateTests(unittest.TestCase):
def test_leaky_transport_fails(self):
with self.assertRaises(GateError):
check_case(CASE, leaky_transport)
if __name__ == "__main__":
unittest.main()
python -m unittest test_fixture_gate.py
Step 6: Record the decision before any remote target
The pairing note is part of the artifact, not a chat scroll that disappears after the session. It names the questions, the two rejected drafts, and the rule that CI never sets the live flag. Only after that note exists should an operator consider pointing a separate, reviewed client at free model access. A free server, when one is actually assigned, belongs in a later deploy change with its own review.
decision: fixture-gate-v1
default_transport: local_stub
ci_live_flag: absent
free_server_in_this_change: no
rejected: live prose assertion; full-reply golden file
Limitations that the note should keep visible
The gate does not prove that a summary is correct, useful, or safe for customers to read. It does not measure latency, token spend, uptime, or how long any free option will last. It will miss a secret that uses a different shape than the marker stored in the fixture. It also will not stop a later deploy step from binding a public port or writing real user data.
Who should not use this gate
Teams that need a formal evaluation suite for model quality should build that suite on purpose, with owned data and a stated rubric. This gate is the wrong tool for medical, financial, or other high-stakes text, because a key check is not a domain review. Operators without a written rule for live flags should not add a remote host just to see a green run. Generated clients that must call arbitrary URLs also fall outside this kept decision and need a different review.
After the local gate already passes
Readers who already hold an approved account can reuse this gate and attach their own endpoint outside CI. No claim is made here that one host is faster, cheaper, or more reliable than another host. A free server option, if the operator assigns one, still waits for a separate deploy review. The gate itself does not require any vendor SDK, token, or hostname to pass on a laptop.
Closing the pairing note
The pair left with one kept decision: structural checks on a stub, live traffic opt-in only, and no free server inside the test change. A later session can review deploy config, but it should not reopen the rejected golden file. Developers can copy the module, point it at their own fixture, and keep the live flag out of CI.
Top comments (0)