A Tuesday agent patch looked clean in nightly CI. Every pinned property check passed on seed forty-one. Monday used a fresh seed and the same check failed.
The failing input was smaller than the recorded example. The flake note had sat untouched for six weeks. Nobody owned an expiry date on that note.
The scene is a composite, not a customer report. That pattern shows up often on agent-heavy repositories. The patch edits production code and nearby tests together.
A frozen flake hides the drift until a new seed lands. A green board then measures silence, not live behavior. The useful fix is a dated lease plus a shrink step.
Merge waits until both artifacts exist and agree. This draft presents the workflow as an unexecuted proposal. The snippets were not run while this article was written.
Adapt every path before you trust a local result. Sample windows below are policy choices, not benchmarks. Do not quote them as measured throughput.
Think of a flake freeze as a short parking permit. The permit lets a noisy test sit outside the gate. A permit without an end date becomes a private road.
New agent patches then learn to drive around the hole. The hole grows because each patch copies the skip list. Weeks later the suite still looks calm in the dashboard.
The calm is an accounting error, not a quality win. A lease repairs the books without pretending flakes vanish. Each skipped check carries an owner, a reason, and a date.
The gate reads that file on every proposed merge. An expired row is a hard failure for the branch. A missing owner is a hard failure of the same kind.
A reason that cites the agent diff also fails closed. The agent may propose a skip in its patch notes. The agent may not grant itself the skip.
Property checks sit beside the lease, not under it. They ask one stable question across many generated inputs. They do not replace the small set of example tests.
The question must live outside the patch directory. If the agent can edit the question, evidence collapses. That path split is the point of the merge gate.
Start from a small corpus you already trust in review. Record inputs, outputs, and the seed list together. Commit that directory before the agent is allowed to run.
Call the directory the sealed corpus in your docs. The agent may read those files during a replay. The agent may not write them inside the same diff.
A pre-commit hook can enforce the path rule locally. The hook is boring, which is why people keep it. Boring checks survive agent refactors better than clever ones.
When a property check fails, do not debate the patch yet. Shrink the failing input before any reviewer opens chat. Shrinking removes fields until the failure finally disappears.
What remains is the counterexample you can actually review. A long payload often collapses to a single field. Review time drops because the input diff stays tiny.
The patch author then faces a fact, not a vibe. That fact still needs a second run on a clean tree. One shrink file is a clue, not a merge approval.
The next sketch is a lease gate written in Python. It was not executed for this article, so treat it as a sketch. Wire the paths to your tree before you enforce it.
# proposal_shrink_gate.py
# Unexecuted sketch. Wire paths before use.
from dataclasses import dataclass
from datetime import date
from pathlib import Path
import json
LEASE_PATH = Path("ci/flake_leases.json")
@dataclass(frozen=True)
class Lease:
test_id: str
owner: str
reason: str
expires_on: date
def load_leases(path: Path) -> list[Lease]:
raw = json.loads(path.read_text())
rows = []
for row in raw["leases"]:
rows.append(
Lease(
test_id=row["test_id"],
owner=row["owner"],
reason=row["reason"],
expires_on=date.fromisoformat(row["expires_on"]),
)
)
return rows
def agent_wrote_reason(reason: str) -> bool:
markers = ("agent skip", "auto-quarantine", "generated waiver")
text = reason.lower()
return any(marker in text for marker in markers)
def assert_lease_allows_merge(today: date) -> None:
leases = load_leases(LEASE_PATH)
expired = [row for row in leases if row.expires_on < today]
unsigned = [row for row in leases if not row.owner.strip()]
agentish = [row for row in leases if agent_wrote_reason(row.reason)]
if expired or unsigned or agentish:
raise SystemExit("lease gate blocked merge")
The lease file stays small on purpose in this design. A large waiver log slowly becomes a second product. Keep one row per skipped test and renew it by commit.
Let the date pass if nobody renews the row. The red gate is the reminder your calendar will not send. It is cheaper than a surprise incident on Monday.
A single lease row looks like the JSON below. The date uses ISO form so scripts can compare it. Replace the issue id with a real tracker link.
{
"leases": [
{
"test_id": "tests/prop_retry_budget.py::test_budget_never_negative",
"owner": "qa-rotation",
"reason": "upstream clock skew tracked in ISSUE-1842",
"expires_on": "2026-10-15"
}
]
}
Pair the lease with a property the agent cannot rewrite. The sketch checks a retry budget invariant in isolation. Replace the fake scheduler with your real function.
Keep the assertion in a path the patch cannot touch. Eighty examples is a starting budget, not a proof. Pin the framework version in your lockfile before comparing runs.
# tests/prop_retry_budget.py
# Unexecuted illustration. Not a measured result.
from hypothesis import given, settings, strategies as st
def spend(budget: int, cost: int) -> int:
if cost < 0:
raise ValueError("cost")
return budget - cost
@settings(max_examples=80, derandomize=True)
default_budget = st.integers(min_value=0, max_value=50)
@given(budget=default_budget, cost=default_budget)
def test_budget_never_negative(budget: int, cost: int) -> None:
if cost > budget:
try:
spend(budget, cost)
except ValueError:
return
raise AssertionError("overspend must fail closed")
assert spend(budget, cost) >= 0
Shrinking belongs in the command log, not in chat. Run the property until it fails on the pinned seed. Store the minimal example under the counterexample directory.
Reviewers should open that file before the model trace. The trace can lie by simple omission of a field. The shrunk input is harder to narrate away in review.
# Unexecuted. Use a disposable worktree only.
mkdir -p ci/counterexamples
python -m pytest tests/prop_retry_budget.py -q --tb=short
python -m pytest tests/prop_retry_budget.py -q \
--hypothesis-seed=41 \
--hypothesis-show-statistics
A second command checks the sealed corpus hash first. If the hash moved, stop the job before tests start. Do not let the agent refresh fixtures to match its patch.
Drift can be real, and sometimes the drift is the bug. Drift still needs a separate human commit to land. That commit should not ride inside the agent diff.
# Unexecuted. Swap in the digest tool you already trust.
find ci/sealed_corpus -type f | sort | sha256sum > /tmp/corpus.sha
cmp ci/sealed_corpus.sha256 /tmp/corpus.sha
Agreement is a hash, not a shared feeling in review. Run the same seed in two clean worktrees. Write each counterexample to a file with a stable name.
Hash those two files and compare the digests. Equal digests mean the failure shape survived a cold start. Unequal digests mean you do not have a stable hole yet.
# Unexecuted. Repeat after cleaning caches between trees.
sha256sum worktree-a/ci/counterexamples/seed-41.json \
worktree-b/ci/counterexamples/seed-41.json
cmp worktree-a/ci/counterexamples/seed-41.json \
worktree-b/ci/counterexamples/seed-41.json
Read the shrunk input before any prose summary of the run. If one field flips the invariant, inspect that field. If the shrink collapses to empty, suspect a fixture bug.
If the failure appears on one seed only, open the lease. A one-seed failure can earn a short dated lease. The lease still needs an owner and an end date.
Fourteen days is a policy choice in this draft. It is not a measured optimum from a public study. Shorten it when renewals start to feel automatic.
Extend a lease only when the issue link is real. A ticket number in the reason is the minimum evidence. The phrase flaky on CI is not evidence.
The phrase agent said skip is also not evidence. Renewals are ordinary commits with a different reviewer. The patch author should not approve their own waiver.
Reject the merge when the counterexample is stable twice. Stability means the shrunk input fails in two clean worktrees. Agent patches often pass once on warm cached state.
A second process with a cold cache tells you more. You still have not proved the change is correct. You have only blocked one known hole in the diff.
Do not add a new property during the red review. New properties belong in a follow-up patch after the fix. Mixing them hides which change actually repaired the bug.
Land the fix against the existing question first. Add extra coverage only in a later, separate diff. The suite then stays legible for the next on-call reader.
Long shrink sessions need a quiet machine more than a clever prompt. A laptop sleep cycle will kill an overnight replay. A shared CI queue may evict the job mid-shrink.
This is where a hosted runner earns its keep. MonkeyCode offers free model access and a free server option. Disclosure: This article was prepared as part of MonkeyCode's product outreach.
Use that access as a spare workbench, not as an oracle. Send the corpus hash, the lease file, and the failing seed. Ask for a shrink log and a suspected break.
Keep the final pass or fail decision in your gate script. The sketch above does not call a model on purpose. A model may suggest a smaller input for review.
It cannot sign the lease or waive an expired row. Exact model names and token quotas are omitted here. No primary source for those figures was attached.
Hardware size and offer duration are omitted for the same reason. Do not paste a quota from memory into the team README. Check the current project docs before you plan capacity.
Offers change, and a stale number becomes a false promise. A practical loop starts by pushing the patch to a branch. Run the lease gate locally before any hosted replay.
If the lease is valid, start the property job remotely. Download the counterexample artifact when the job ends. Re-run the same seed once on a second clean worktree.
Merge only when both runs agree on the result. Also require that no lease row is past its date. The second run is the part many teams skip.
That second run is what catches warm-cache lies. Skip it and the gate becomes theater with extra steps. Theater trains reviewers to click through red rows.
This workflow assumes you already have a few stable properties. A new codebase with no invariant gains little from leases. You would be leasing noise and calling it process.
Write three example tests before you invent a property. Promote one example only after it survives a week. The promotion is a design act, not a tooling act.
Skip the dated lease if your suite has no flakes. A lease file with zero rows is harmless extra process. Add the file when the first skip actually appears.
Process that arrives early becomes theater for the team. Theater is how red gates lose their meaning over time. Wait for a real waiver before you add the ceremony.
Do not use this gate as a security boundary. It checks test accounting and shrink quality only. It does not prove the agent lacked dangerous tool access.
Pair it with your existing permission review for tools. A green shrink does not bless a widened shell allowlist. Those failures have different owners and different fixes.
Shrinking can hide bugs that need large structured inputs. A minimal counterexample is not the only interesting case. Keep a small fixed corpus of large cases beside it.
The hash check above covers that fixed corpus each run. If your domain fails only at scale, budget a nightly job. Do not expect eighty random examples to find those bugs.
Time zones can expire a lease one day early. Store dates in ISO form and evaluate them in UTC. The sketch uses the runner local date, which is a gap.
Fix that gap before you enforce the gate globally. A false expiry creates alert fatigue within a week. Fatigue is how careful lease files get deleted.
Free hosted runs can leak fixtures if you upload customer data. Seal only synthetic inputs or already public samples. Strip secrets before the job leaves your network.
A free server is still someone else's computer. Treat the upload as a data review, not a convenience toggle. Convenience is how sealed corpora stop being sealed.
Track four counts in the merge log each week. Count expired leases that blocked a patch from landing. Count shrinks that cut input size by half or more.
Count patches rejected after the second seed disagreed. Count waivers renewed by the same person who wrote the patch. That last count should stay near zero if review is real.
Do not publish those counts as a model quality score. They measure your gate, not a vendor ranking table. A quiet week can mean the properties never ran.
Log the number of property examples executed per job. A green run with zero examples is a vacuous pass. Block that case in the same script that reads leases.
Start with one property and one lease row only. Run the gate on a disposable branch for five merges. Adjust the expiry window only after those five merges.
Fourteen days is a starting policy, not a law. Lengthen it only when a written incident link exists. The date should hurt a little so the list stays short.
Need a spare machine for that first shrink replay? Read the current MonkeyCode docs and confirm the free server terms. Use the spare run, and keep the merge bit local.
The lease expires whether or not the dashboard looks calm. The shrink file tells you what actually broke. Everything else is commentary until those two artifacts exist.
Top comments (0)