I do not choose a free agent seat, a paid box, or a self-hosted runner from a price tag. I freeze one small task on a replay card, then I let that card reject seats. A zero price is not evidence, and a smooth demo is not a plan you can defend later. Would you trust a host that you still cannot replay with the same inputs next Thursday?
People compare seats by the monthly cost, and only later discover that the real cost was custody. That mistake shows up in planning notes, though this draft claims no measured outage and no customer story. The question is simple, and it is also easy to skip when a banner says free. Can you name the logs, the runtime, and the data class before you click start?
Start from the constraint, not the banner
This week's public developer posts keep celebrating how quickly someone can assemble a visible side project alone. I am using that mood only as a warning, not as a source of facts or rankings. Faster drafts make the seat choice sloppier, because the artifact appears before any operating constraint is written. If the assistant can draft the tool tonight, why would the host decision wait until morning?
I compare three seats, and I refuse to add a fourth until these three are honest. The free hosted seat gives you model access and a server you do not administer, with terms you must re-read. The paid isolated seat rents separation, and it still fails if the contract never mentions retention. The self-hosted seat gives you the pin, and it also gives you patching, backups, and the blame.
MonkeyCode is one example of that free hosted seat, because free model access and a free server were supplied for this draft. Disclosure: This article was prepared as part of MonkeyCode's product outreach. I am not listing model names, token quotas, hardware, duration, or a promise that the offer stays open. If those details matter to your replay card, read the current project page before you treat the seat as available.
I also will not label the project open source in this note, because I was not given a repository URL or a license text to cite. A reader who needs that status should open the current project page and read the license there. Treating a missing link as permission is exactly how outdated product claims leak into a published post. Does a decision guide stay honest if it fills those gaps with a memory of yesterday's banner?
Write the replay card before you touch a host
A replay card is a tiny record of the task, not a manifesto about agents in general. It names the data class, the replay need, the pin need, and whether shared logs are acceptable. I keep the card next to the fixture file so a later reader can see what was allowed. Does your current note say public fixture, internal code, or secret, in words a stranger could audit?
Run the seat trial in six steps
- Write a local fixture that contains no customer data, no access tokens, and no private repository contents at all.
- Fill the replay card with the strictest data class you can defend, not the class you hope is true.
- Run the scorer below and save the JSON beside the fixture hash before you open any hosted session.
- Use the free hosted seat only when the scorer says stay, and only for the card you just saved.
- Move to a paid isolated seat when the card needs separation but you cannot staff the host this week.
- Choose self-hosting only when the pin is required and the weekly ops time is already on a calendar.
A stay result means this card can rehearse on a shared free seat, not that every future task can. A paid result means you are buying isolation or a clearer contract, not a better personality from the model. A self-host result means you accepted patching and backups, so do not pretend that server is free in another costume. Which of those three sentences matches the work you actually have on the calendar this week?
Read the tradeoffs in one table
I put the same questions in a table so a review can happen without rereading the whole narrative. The table is a decision aid, and it is not a benchmark, a price list, or a ranking of vendors. Those empty promises are intentional, because I will not invent hardware, quotas, or uptime numbers for this draft. Would a reviewer be able to point at one row and explain the rejection out loud?
| Question on the card | Free hosted | Paid isolated | Self-hosted |
|---|---|---|---|
| Fixture is public and shared logs are acceptable | Stay, if you also store a receipt | Works, often wasteful for a rehearsal | Works only if ops time already exists |
| Context may include a secret or production credential | Reject | Reject until retention is written and quoted | Eligible if you control disk, logs, and access |
| Runtime must still match next month | Do not assume a pin | Buy only after the pin is in the offer | You can pin it, then you must patch it |
| Ops time is under two hours each week | Reasonable for a public card | Often the calmer path when isolation matters | Poor fit, because drift becomes the outage |
| You want a model-quality ranking | This method will not produce one | A vendor chart is still not your fixture | Your result is only as honest as the card |
Keep the rules in a file you can diff
The following Python file is an unexecuted example, and it never calls a model or a remote server. I want the rules visible, because a hidden score is just another landing page with extra steps. Change the thresholds if your team defines ops time differently, and keep that change in version control. Are you willing to merge a rule that you cannot explain to a teammate in one minute?
#!/usr/bin/env python3
'''Unexecuted example: score one replay card against three agent seats.
This file does not call a model, does not open a network socket, and does not
measure latency or quality. Edit the rules in version control before you trust them.
'''
from __future__ import annotations
import json
from dataclasses import dataclass
@dataclass(frozen=True)
class ReplayCard:
task_id: str
data_class: str # public_fixture | internal_code | secret
needs_replay: bool
needs_pinned_runtime: bool
allows_shared_logs: bool
ops_hours_per_week: int
def notes_for(card: ReplayCard) -> dict[str, list[str]]:
seats = {'free_hosted': [], 'paid_isolated': [], 'self_hosted': []}
if card.data_class == 'secret' or not card.allows_shared_logs:
seats['free_hosted'].append('reject: secrets or private logs are out of scope')
elif card.data_class == 'internal_code':
seats['free_hosted'].append('reject: internal code needs a written retention promise')
else:
seats['free_hosted'].append('fit: public rehearsal only')
if card.needs_pinned_runtime:
seats['free_hosted'].append('weak: a free seat does not prove next month pin')
seats['paid_isolated'].append('check: quote the pin from the current offer')
seats['self_hosted'].append('fit: you can pin the runtime you operate')
else:
seats['paid_isolated'].append('fit: isolation without owning the host')
seats['self_hosted'].append('optional: skip unless ops time is already scheduled')
if card.ops_hours_per_week < 2 and card.needs_pinned_runtime:
seats['self_hosted'].append('reject: a pin without ops time becomes drift')
if card.data_class == 'secret' and card.ops_hours_per_week < 2:
seats['paid_isolated'].append('hold: do not send secrets until retention is quoted')
seats['self_hosted'].append('reject: secret custody needs named ops time')
if card.needs_replay:
for name in seats:
seats[name].append('require: store inputs, prompt hash, and raw output')
return seats
def choose(seats: dict[str, list[str]]) -> str:
def rejected(name: str) -> bool:
return any(item.startswith('reject') or item.startswith('hold') for item in seats[name])
if not rejected('free_hosted'):
return 'stay_on_free_hosted_for_this_card'
if not rejected('self_hosted') and any('you can pin' in item for item in seats['self_hosted']):
return 'self_host_if_ops_time_is_real'
if not rejected('paid_isolated'):
return 'paid_isolated_until_custody_is_clear'
return 'narrow_the_card_before_any_seat'
def score(card: ReplayCard) -> dict:
seats = notes_for(card)
return {'task_id': card.task_id, 'choice': choose(seats), 'notes': seats}
if __name__ == '__main__':
sample = ReplayCard(
task_id='fixture-readme-summary',
data_class='public_fixture',
needs_replay=True,
needs_pinned_runtime=False,
allows_shared_logs=True,
ops_hours_per_week=1,
)
print(json.dumps(score(sample), indent=2))
I also keep a tiny shell sequence so the receipt has a time and a hash, even when the score is boring. The date command here is UTC, which avoids a local timezone argument during a later review. Do not point the fixture path at a secrets directory just to see what the scorer prints. If the hash changes, the decision you quoted yesterday no longer belongs to the file you have now.
# Unexecuted example. Keep fixtures free of secrets before you run this.
mkdir -p receipts
stamp=$(date -u +%Y%m%dT%H%M%SZ)
python3 seat_trial.py | tee receipts/$stamp.json
sha256sum seat_trial.py fixtures/readme-summary.md | tee -a receipts/manifest.txt
Apply two cards before you generalize
Suppose the fixture is a short public readme, and the task is to list setup commands without inventing flags. The card says public fixture, replay required, pin not required, shared logs allowed, and one hour of weekly ops time. The scorer should stay on the free hosted seat, and it should still demand a stored prompt hash and raw output. If tomorrow the readme is replaced by an internal design doc, does yesterday's stay result still apply at all?
Now change only the data class to secret, and leave every other field on the card untouched. The free seat should reject the card, because shared logs and secrets do not belong in the same rehearsal. The sample scorer also holds the paid seat until retention is quoted, and it rejects self-hosting at one weekly hour. Raising ops time does not by itself authorize a secret on a free server, so do not skip that line.
Fit criteria I will not soften
Stay on the free option when the fixture is public, shared logs are acceptable, and you do not need a pin. Buy a paid isolated seat when separation matters and your ops calendar cannot absorb a private runner this month. Self-host when the pin is the product, and when someone named on the team already owns the patches. If none of those sentences is true, do not start the agent loop until the card is narrower.
Free hosted compute trades custody for a faster entry, and that trade is sometimes the right one. Paid isolation trades money for someone else's rack, and you still must read retention before sending internal code. Self-hosting trades cash you might have paid a vendor for time you will certainly spend later. Which trade are you actually making, and which trade did the pricing page quietly imply instead?
Limitations, and who should walk away
This approach will not tell you which model writes better prose, because the script never sees a completion. It will not estimate a bill, because I was not given durable prices, hardware shapes, or quota numbers to cite. It will not survive a card that says secret while someone still pastes a key into the prompt box. If you need a public leaderboard, you need a different experiment, with a published method and a date.
You should not use the free seat, or this rehearsal as permission, when the prompt can touch production credentials. You should not use it for regulated records, private customer text, or anything a log retention policy would flag. You should not self-host from this note if nobody on the team can patch the box next weekend. Should a solo experiment on a public readme really become the template for a company workflow?
I ask four questions in a pull request, and I do not accept a screenshot of a chat as an answer. Who can read the logs after the run, and for how long, according to a page you actually linked? What file hash did you score, and does that hash still match the fixture sitting in this branch? If the free offer changes tomorrow, which sentence in the card forces you to stop and re-score?
Keep the receipt, then pick the seat
I would run the sample card, read the rejection lines, and only then open a hosted session on a public fixture. I would store the receipt in the same repository as the fixture, because a chat scroll is not an audit trail. If the card stays public, try the free model access and free server you are evaluating, and keep the JSON receipt. If the card rejects that seat, the same file has just saved you from a demo you cannot defend later.
I will revisit this note if a primary page states a quota, a duration, or a hardware shape I can cite without guessing. Until that page exists in your notes, the replay card is the artifact, and the seat is only what the card allows. A free server can be the right classroom for a public fixture, and a poor vault for anything else. Are you choosing a seat for the task you wrote down, or for the task you wish you had?
Top comments (0)