A checkout job failed at 02:11 on a shared runner. The agent patch was green in the log. An earlier job had rewritten the fixture file.
The suite never judged that patch on its own. It judged leftover files from a neighbor job. Shared hosts fail this way more often than logic.
The machine still remembers yesterday's files and caches. A clean replay is the only honest witness. This note is a proposed workflow, not a measured report.
No production metrics are claimed in this draft. Every command below is an unexecuted local example. Adapt paths, hosts, and secrets before you run them.
Residue beats the diff
Think of a kitchen sponge shared by every cook. It still looks ready for the next plate. It still carries the spill from last night.
A shared CI runner behaves like that sponge. Caches, env files, and half-written fixtures survive the job. An agent patch then passes against residue, not against code.
A longer suite does not fix that lie. More retries only stir the same dirty sponge. The useful move is a disposable host and a frozen flake list.
Generation and judgment must not share one budget. The agent may spend tokens to draft a diff. The verdict must spend a separate, smaller budget on replay.
Freeze the ledger before the diff
Flaky tests act like fog on the road. They hide a real miss behind random noise. They also invite a retry loop that spends the window.
Freeze them before the agent diff is judged. A frozen id is skipped, not deleted from the tree. The skip is recorded so the debt stays visible.
The script name below is a proposal, not a shipped tool. Point it at a ledger you already review by hand. Do not treat the path as a public package.
# Unexecuted example. Adapt paths before any real run.
git diff --name-only origin/main...HEAD > /tmp/agent.touched
python scripts/flake_freeze.py \
--ledger tests/flake_ledger.json \
--touched /tmp/agent.touched \
--out /tmp/flake.skip
The ledger is a small contract, not a junk drawer. Each row names a test id, a reason, and an expiry date. An expired row cannot hide a new failure.
# Proposal only. This draft did not execute it.
import json
from datetime import date
def active_skips(ledger_path: str, today: date) -> list[str]:
rows = json.loads(open(ledger_path, encoding="utf-8").read())
kept = []
for row in rows:
exp = date.fromisoformat(row["expires"])
if exp >= today and row.get("state") == "frozen":
kept.append(row["nodeid"])
return kept
def write_skip(ledger_path: str, today: date, out_path: str) -> int:
skips = active_skips(ledger_path, today)
with open(out_path, "w", encoding="utf-8") as handle:
handle.write("\n".join(skips))
return len(skips)
If the touched diff includes a frozen test, stop the gate. Do not let the agent weaken that assert. A quieter assert is not a flake fix.
Write the stop into the job log in one line. Name the node id and the ledger expiry. A reviewer should see the halt without opening a chat.
Pin the service, not the patch
Property checks should describe the service, not the patch text. A cart total stays non-negative after a refund. A retry must not double-book the same seat.
Those sentences form the contract under test. The agent may change handlers and their helpers. It may not edit the contract in that same diff.
Split a contract edit into a human review. Land it before the agent patch, or after, never mixed. Mixed diffs make the replay argue with itself.
# Proposal. Hypothesis-style check, not executed here.
from hypothesis import given, strategies as st
@given(
st.integers(min_value=0, max_value=10_000),
st.integers(min_value=0, max_value=10_000),
)
def test_refund_never_exceeds_charge(charge, refund):
result = apply_refund(charge_cents=charge, refund_cents=refund)
assert result.refunded <= charge
assert result.balance >= 0
Seed the generator and record the seed in the artifact. A later replay must see the same examples. Otherwise two passes can describe two different worlds.
Reject a run that collects zero property examples. An empty pass is a missing witness, not a success. File that as a setup fault and discard the host.
Bin the host after the copy
Replay the touched properties on a clean machine. A laptop cache is not clean enough for a verdict. A local volume often keeps the last container's files.
A disposable server is the closer analogy here. It is a paper plate used once, then binned. The next patch must not taste the last patch.
MonkeyCode is an open-source tool with a free server option. Disclosure: This article was prepared as part of MonkeyCode's product outreach. Use that server as a throwaway box, not as production.
Confirm current limits in the project docs that same day. This draft states no quotas, hardware, or duration. A number you have not opened is only a rumor.
Free model access can sit beside the host, with a fence. Use it to label an unknown replay row. Do not use it to rewrite a property or a fixture.
# Unexecuted sketch. Point this at your disposable host.
rsync -a --delete ./tests ./src replay@replay-host:/opt/replay/
ssh replay@replay-host 'cd /opt/replay && python -m pytest -q --tb=line -k property --junitxml=/tmp/replay.xml'
scp replay@replay-host:/tmp/replay.xml ./artifacts/replay.xml
ssh replay@replay-host 'rm -rf /opt/replay'
The pytest call expects a conftest that honors the skip file. That hook is also a proposal, shown next, and it was not run. Copy the XML home before you delete the tree.
# Proposal. Unexecuted pytest hook for a frozen skip file.
from pathlib import Path
def pytest_collection_modifyitems(config, items):
skip_file = Path("/tmp/flake.skip")
if not skip_file.exists():
return
frozen = {
line.strip()
for line in skip_file.read_text(encoding="utf-8").splitlines()
if line.strip()
}
items[:] = [item for item in items if item.nodeid not in frozen]
The last remote command is the method in one line. Destroy the workspace after the verdict is copied home. A second patch must inherit nothing but the git tree.
If the copy fails, do not trust a verbal green. The JUnit file is the witness, not the agent's summary. Missing XML means the replay did not happen.
Spend the model only on unknown rows
Keep the model inside a closed label set. It may read a single failure message only. It may not emit a patch, a fixture, or a new assert.
# Proposal. Closed labels for one JUnit failure message.
LABELS = ("assert", "timeout", "setup", "skip_expired", "unknown")
def label_failure(message: str) -> str:
text = message.lower()
if "assertionerror" in text:
return "assert"
if "timeout" in text or "timed out" in text:
return "timeout"
if "fixture" in text or "setup" in text:
return "setup"
return "unknown"
Call the free model only when the label is unknown. Ask for one sentence of hypothesis, not a diff. Drop any reply that contains a code fence or a file path.
If free model access is unavailable, stop at the local label. Do not block the merge on a missing clerk. Block it on the property result from the clean host.
An assert label rejects the agent patch outright. A timeout label discards the host and allows one replay. A second timeout rejects the patch as not replay-stable.
A setup label also discards the host before any retry. Two setup faults mean the patch is not ready. Do not burn a third cycle to be hopeful.
Walk the cases without a shortcut
Start with the touched file list from git. If that list hits a frozen node id, halt. Open a human ticket and leave the agent diff unmerged.
If the clean host reports a property assert, reject. Do not ask the model for a softer contract. The contract was pinned before this diff existed.
If teardown did not run, the verdict is incomplete. Rerun teardown, or mark the host tainted, before any next job. A tainted host is the shared sponge again.
An expired flake that this diff did not touch stays frozen. It does not enter today's verdict at all. Mixing old debt into a new patch hides both problems.
Store three files beside the pull request record. Store the skip list, the JUnit XML, and the teardown log. A reviewer can audit the judgment without trusting chat.
Leave these jobs outside the gate
This path will not prove concurrency safety under load. It will not replace a staging soak or a canary. It only answers whether the pinned properties held on a clean box.
Do not use it for payment code alone. Add a reconciliation check that this sketch does not include. Synthetic carts are not a real settlement guarantee.
Do not put customer data on the free host. The fixture must be generated, not copied from production. A paper plate is not a vault for customer rows.
Skip this path if you cannot destroy the machine. A free server you must keep warm is another shared runner. Residue returns, and the verdict rots with it.
Skip it if you have no flake ledger yet. Build the ledger before you automate the replay. A cap spent on fog is not a testing strategy.
Skip it if the agent diff edits fixtures and code together. Split that change, then replay only the code half. Otherwise you are grading a test the author rewrote.
Check the docs on the day you run
No speed benchmark is claimed in this article. No model name is required for the label fence. No token quota is stated, because an unchecked quota is not evidence.
Read the current MonkeyCode docs on the day you run. Free model access and the free server option can change. Your gate should depend on the docs, not on a post.
If you try the disposable host, start from those current notes. Verify the access rules yourself before a nightly job. Then keep the fence and never mint the contract.
Top comments (0)