Day 21 is the date that reverses most zero-invoice agent decisions. Not day one. Not the procurement meeting. Day 21 is when senior review hours finally show up as a cost, even if nobody put them on a spreadsheet.
Here is a composite I use in adoption reviews. Names are gone. The constraint is not. A twelve-engineer product org lets a coding agent into one squad. Week one feels cheap. Week three, two staff engineers are rewriting assumed configs, guessed secrets, and “helpful” refactors that never matched the workflow. The invoice is still zero. The calendar is not.
Would you keep that path? Or would you buy seats, self-host, or shut it down?
This is not another build-versus-buy slogan. It is a retention gate. Treat the scorecard as a conversation tool, not objective truth. If one field below flips, the whole keep/buy/host/stop call can flip with it.
Free access is not a retention policy
Platform leads keep asking the wrong first question: free, paid, or self-hosted? That question is a vendor conversation. The operating conversation is different. Are you retaining a reviewed workflow, or retaining a habit of unowned assumptions?
Agents fail in a boring way. They assume the environment. They assume the ticket is the spec. They assume a file that compiled locally is safe to merge. Trend chatter this week keeps rediscovering that point. The product problem is older: incentives, review load, and who gets paged on Friday.
A zero-invoice path can still be the right path. It can also be a shadow platform with no owner and no exit. The difference is measurable by day 21.
Define five fields before you argue tools
Write the variables down. If you cannot fill them, you are not ready to standardize anything.
- Review-hour leakage (R). Extra human review hours per merged PR that carried agent output, versus the squad’s baseline PR. Unit: hours. Owner: engineering manager.
- Assumption-defect rate (A). Share of those PRs that needed a rewrite because the agent guessed environment, secrets, scope, or “the user probably wanted.” Unit: percent. Owner: tech lead.
- Runtime owner (O). Named human who is on the hook when the path stalls. Not “the platform team.” A person. Unit: name plus coverage hours.
- Switching days (S). Calendar days to move the squad to a paid seat or a self-hosted runtime without freezing delivery. Unit: days. Owner: platform lead.
- Incentive match (I). Does the path reward merged, reviewed work, or does it reward prompt volume and unreviewed drafts? Unit: yes/partial/no, with one sentence of evidence.
Hard rule: if O is blank, the path is a lab, not a product surface. Labs expire. Products have owners.
Keep / buy / host / stop
Use the fields as a comparison, not a vibe.
- Keep the free path when R stays near baseline, A is visible and falling, O is named, S is known, and I is “yes.”
- Buy when you need vendor coverage, audit trails, or a support contract, and self-hosting would steal a platform engineer you do not have.
- Self-host when data gravity, lock-in fear, or runtime control dominate, and you can staff O for real.
- Stop when A is high and nobody owns the mess. Stopping is a product decision. It is not a moral failure.
Which field would reverse your call this month?
Worked example: a 12-engineer product org
Labeled hypothetical. Do not treat these numbers as a benchmark, a MonkeyCode result, or your team’s truth.
Squad ships about 40 merged PRs in 21 days. Ten are tagged agent-assisted. Baseline review is 0.4 hours per PR. Agent-assisted review is 1.1 hours. Three of the ten PRs needed a rewrite for guessed staging URLs and a “temporary” secret in a comment. Nobody is on-call for the free runtime. The platform lead guesses a paid rollout would take 10 days and a self-host would take 25.
Fill the gate:
| Field | Value in the example | Keep? |
|---|---|---|
| R | 0.7 extra hours per agent PR | Weak |
| A | 30% assumption defects | Fail |
| O | blank | Fail |
| S | 10 days paid / 25 days host | Known |
| I | partial — drafts outrun review | Weak |
Decision in this example: do not keep as a standard. Either name an owner and rerun a 21-day lab, buy a governed path if switching days are the binding constraint, or stop. The free option lost on O and A, not on token price.
Loaded cost sketch, still hypothetical. If senior review is $150/hour fully loaded, 10 agent PRs × 0.7 extra hours = 7 hours = $1,050 of review leakage in 21 days. That figure does not prove paid software is cheaper. It proves “zero invoice” was the wrong unit. If your loaded rate is $90 or $220, the keep/buy line moves. Run your number.
Sensitivity: the number that flips keep to exit
Hold everything else constant and move one field.
- If A drops from 30% to 8% and O is named, keep can win even with a free runtime.
- If R exceeds 0.75 extra hours and stays there after coaching, buy or host. You are spending staff-engineer time as an invisible seat license.
- If S is unknown, you are not allowed to call the path “the default.” Unknown switching cost is a lock-in, including lock-in to a free tool.
- If I is “no,” stop. A path that rewards unreviewed volume will beat any scorecard in the hallway.
Break-even question for the EM: at what extra review hour per agent PR would you rather pay cash than pay calendar? Write that threshold in the doc before the pilot. If you write it after, you will negotiate with your own sunk time.
Expiry: 21 days after the first merged agent-assisted PR. Not 21 days after the Slack announcement. After evidence.
Hard gates, owner, expiry, exit
These are gates, not aspirations.
- No named runtime owner (O) → lab only, or stop. No expansion to a second squad.
- Assumption-defect rate (A) above your pre-written threshold after day 21 → buy, self-host with policy, or exit. Do not “try one more week” without changing an incentive.
- Switching days (S) not estimated → you may not declare a standard.
- Incentive match (I) is no → exit. Tools do not fix a review culture that refuses to label agent output.
- Expiry hit with missing data → treat missing data as a fail, not as a pass. Absence of telemetry is a decision.
Owner of the gate: the engineering manager of the squad, not the vendor champion, not the intern who found the free path. Platform leads advise. Security can veto. Product does not get a silent shadow platform.
Exit criteria, written on day zero:
- Unlabel the workflow in the team README.
- Archive the agent config in the repo.
- Move open drafts back to human-owned branches.
- Record why you stopped in one paragraph, with the field that failed.
If you cannot write the exit paragraph in advance, you are running a hope, not a pilot.
A zero-invoice lab still needs a runtime owner
Some teams want a lab that does not start with a contract. That is a fair constraint for a founding team or a platform group testing adoption, not for a regulated production path.
Open-source coding agents that offer free model access and a free server option fit that lab shape. They do not fit an unowned production shape. Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode is one such open-source project with those two availability claims. I am not attaching model names, quotas, hardware, uptime, or duration to it here, because those details go stale and I will not invent them.
Use the same five fields on that lab as you would on a paid seat. If O is still blank, the free server is a convenience, not a platform. Convenience without an owner becomes shadow DX. You already knew that from every other “temporary” CI runner.
Tradeoffs, still qualitative:
- Free model access + free server: lowest cash, highest need for a human owner, weakest place to hide missing review policy.
- Paid: you are buying coverage and a throat to choke, not magic quality. Still measure R and A.
- Self-host: you are buying control with staffing. If you cannot staff O, you did not self-host. You self-neglected.
The interesting comparison is not logo versus logo. It is who pays when the agent assumes.
Measure it with tools you already have
Proposal only. Unexecuted. Do not paste this into production policy without your own labels.
Tag the work. If you cannot tag it, you cannot retain it.
# Proposal: list recent merged PRs labeled as agent-assisted.
# Adjust the label to whatever your squad actually uses.
gh pr list --state merged --search "label:agent-assisted" --limit 50 \
--json number,title,mergedAt,reviews,comments
Write a tiny retention record the EM can fill in 10 minutes. YAML is enough. A wiki table is enough. A spreadsheet is enough. The format is not the point. The missing owner is the point.
# proposal: day-21 retention record (not a production schema)
pilot:
squad: payments-api
started_on: 2026-08-19
expires_on: 2026-09-09
runtime_owner: "alex (EM), backup sam (TL)"
fields:
R_extra_review_hours: 0.7
A_assumption_defect_rate: 0.30
O_named: false
S_days_to_paid: 10
S_days_to_self_host: 25
I_incentive_match: partial
decision: stop_or_re_lab
exit_if:
- O_named == false
- A_assumption_defect_rate > 0.15
- R_extra_review_hours > 0.75
Want a cheap signal for A without building a platform? Sample ten agent PRs and count how many required a rewrite because the agent assumed. Ten is a conversation. Two is an anecdote. Zero labeled PRs is a fail on telemetry.
PR sample assumed_env assumed_secret assumed_scope rewrite?
#481 y n y y
#486 n n y y
#490 y y n y
Three rewrites in a sample of ten is not a model-quality debate. It is a workflow-mismatch debate. Fix the prompt policy, the repo map, or the path. Do not buy a second tool to launder the same assumption tax.
Who should not use this gate
Skip this approach if you have no PR review culture. The gate needs review hours. Without them, R is undefined and you will hallucinate a decision.
Skip it if you handle regulated data and your counsel already forbids a free runtime. Do not run a “lab” as a workaround. That is not product strategy. That is an incident waiting for a date.
Skip it if you need an SLA you did not purchase. A free server option is not an on-call rotation. If your customers feel the agent path, you already left lab territory.
Skip it if you want a framework that names a winner without your numbers. This gate will not do that. It will force a threshold. That is the job.
Ask the reversal question
I keep coming back to the same line in the planning room. If review-hour leakage crossed 0.75, would you still keep the zero-invoice path? If assumption defects stayed above your threshold with a named owner, would you still refuse to buy or host? If switching days were 5 instead of 25, would self-host even be on the table?
Pick the variable. Write the number. Put an expiry on the calendar. Name the person who has to live with Friday.
If the free path still wins after that, keep it on purpose. If you want a lab with free model access and a free server option while you run the gate, MonkeyCode is one place to put that lab — after O is a name, not a hope.
Top comments (0)