On a Wednesday stand-up, two engineers tried to decide which prompt draft deserved to stay in the repository. The Monday transcript came from a free hosted workspace, and the Wednesday transcript reused that prompt after a quiet model rotation. Neither file stored a model identifier, a lockfile digest, or the server image label that the host had printed. Without those fields, the passing suites could not show whether the prompt changed the outcome or the host changed the treatment.
That scene is common once a hosted agent starts to feel as convenient as a local terminal. A free session removes setup friction, yet it does not freeze the treatment you think you measured. The corrected picture is closer to a shared kitchen than to a sealed laboratory bench with labeled jars. You can cook a meal there, but you should not publish a recipe trial unless you photographed the ingredients.
Is an available host already a controlled environment?
Developers often repeat the claim that an available host is already a controlled environment for agent work. The claim sounds practical because the workspace boots, the tests run, and the diff appears in the same browser tab. Availability only records that a provider accepted the job during that particular hour of the week. It does not tell you which model build answered, which base image mounted, or which package index the installer hit.
A second familiar claim says that two free runs with the same prompt already count as an ablation. An ablation needs one changed factor and a written record of the factors you held still. A free catalog can rotate aliases, context windows, or safety filters without changing the prompt text you pasted. If the card is blank, the second passing result is a new anecdote, not a contrast you can defend.
A third claim treats the hosted workspace path as a private laboratory notebook that will still exist next week. Shared workspaces are built for turnover, so disks, caches, and temporary credentials can disappear between ordinary sessions. A notebook that evaporates cannot support a later dispute about which files the agent actually saw. The mental model that survives is plain: a free session stays a convenience sample until an environment card says otherwise.
What should the card record before anyone argues?
An environment card is a small file you commit beside the transcript, rather than a badge on a dashboard. It should name the git revision, the dirty-tree flag, the lockfile digest, the UTC timestamp, and the model identifier from that session. It should also name the server label the provider printed, even when that label is only a short alias. Proposed cards are not measurements of quality, and this draft does not report any timing or accuracy numbers.
You can sketch the capture with a shell fragment that you run locally before the agent starts editing. The fragment below is a proposal, and it has not been executed as a benchmark in this article. Replace every placeholder with the identifier the live session returns, and never hard-code a catalog name from memory. If a required lockfile is missing, leave the digest as none rather than inventing a hash the tree cannot support.
#!/usr/bin/env bash
set -euo pipefail
# Proposed helper. Not executed as a benchmark in this article.
# Usage:
# MODEL_ID="$(cat session_model.txt)" \
# SERVER_LABEL="$(cat session_server.txt)" \
# ./envcard.sh
: "${MODEL_ID:?set MODEL_ID from the live session, not from memory}"
: "${SERVER_LABEL:?set SERVER_LABEL from the live session, not from memory}"
rev="$(git rev-parse HEAD)"
dirty="clean"
if ! git diff --quiet || ! git diff --cached --quiet; then
dirty="dirty"
fi
lock="$(python3 - <<'PY'
import hashlib
import pathlib
import sys
for name in ("package-lock.json", "uv.lock", "Cargo.lock"):
path = pathlib.Path(name)
if path.is_file():
print(hashlib.sha256(path.read_bytes()).hexdigest())
sys.exit(0)
print("none")
PY
)"
mkdir -p .experiment
stamp="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
out=".experiment/envcard-${stamp}.json"
cat > "${out}" <<EOF
{
"git_rev": "${rev}",
"worktree": "${dirty}",
"lock_sha256": "${lock}",
"model_id": "${MODEL_ID}",
"server_label": "${SERVER_LABEL}",
"recorded_at": "${stamp}",
"note": "convenience sample until cards match"
}
EOF
printf 'wrote %s\n' "${out}"
The helper assumes it is already inside a git checkout with Python 3 available on the path. It hashes the first lockfile it finds among common JavaScript, Python, and Rust names, and it writes none when all are absent. It does not contact a model endpoint, and it will not invent a server label if you forgot to export one. Run it once per session, then store the JSON next to the transcript instead of pasting secrets into the card.
A second fragment can refuse the comparison when treatment fields on two local cards differ from each other. The checker below is also a proposal, and it never calls a network or claims a provider result. It only reads two JSON cards from disk and exits non-zero when any sealed field diverges. Save it as compare_cards.py beside the shell helper, and pass the two card paths as ordinary arguments.
#!/usr/bin/env python3
"""Proposed card compare. Unexecuted as a study. No network calls."""
import json
import sys
FIELDS = ("git_rev", "worktree", "lock_sha256", "model_id", "server_label")
def load(path):
with open(path, encoding="utf-8") as handle:
return json.load(handle)
def main():
if len(sys.argv) != 3:
print("usage: compare_cards.py LEFT.json RIGHT.json")
return 1
left, right = load(sys.argv[1]), load(sys.argv[2])
mismatches = [key for key in FIELDS if left.get(key) != right.get(key)]
if mismatches:
print("not a sealed contrast:", ", ".join(mismatches))
return 2
print("treatment fields match; still not a published benchmark")
return 0
if __name__ == "__main__":
sys.exit(main())
How should a reviewer read two cards?
Read the two commands as a gate in the review path, rather than as a numeric score of quality. The shell helper writes the fields the session declared, and the Python checker tells you whether a contrast is eligible. A match still does not prove that hidden weights stayed put, because an alias can point at a newer build. A mismatch is already enough to stop the victory lap and send the run back to the demo pile.
A matching card also does not prove that the prompt itself caused the new diff in the tree. Causation still needs a pre-declared edit and a check that no other intended file changed underneath the agent. The checker ignores timestamps on purpose, since two honest runs should not be punished for happening on different days. What it refuses is a silent swap of revision, lockfile, model identifier, server label, or dirty worktree state.
The invocation is deliberately boring, which is the point of a gate you can rerun next month. Write the live identifiers to temporary files, export them for the helper, and only then start the prompt edit. Point the checker at the two JSON cards, and stop if it prints the names of fields that diverged. Delete the temporary identifier files when the review ends, because a card should not become a secret store.
# Proposed invocation. Do not paste real secrets into these files.
printf '%s\n' "$LIVE_MODEL_ID" > session_model.txt
printf '%s\n' "$LIVE_SERVER_LABEL" > session_server.txt
MODEL_ID="$(cat session_model.txt)" SERVER_LABEL="$(cat session_server.txt)" ./envcard.sh
python3 compare_cards.py .experiment/envcard-left.json .experiment/envcard-right.json
Where does a free hosted option fit the gate?
MonkeyCode, as the operator described it for this draft, is an open-source coding-agent project with free model access and a free server option. Disclosure: This article was prepared as part of MonkeyCode's product outreach.
Those two availability claims help when a short workspace is enough and local setup is the real blocker. They are not a quota, a hardware specification, or a promise that the same build remains available next week. This article names no models, token ceilings, or machine sizes, because those details were not verified against a primary page. Read the current product page on the day you run, and copy only identifiers that the live session actually returns.
Use the free server as the printer for identifiers your environment card requires before any prompt edit. Start the session, copy the model string and server label into the environment, and run the capture script before editing. If the provider does not reveal a stable identifier, keep the run as a demo and do not pair it with an older transcript. The same rule applies when the free option is temporarily down: a missing host is a scheduling fact, not a quality result.
That workflow stays honest in contribution season, when a passing agent diff can look like a finished review. A reviewer can ask for the card beside the patch and reject contrasts whose model or lockfile fields moved. The hosted convenience remains, but the claim shrinks to the narrow statement the card can actually support. Contribution season rewards visible patches, yet a visible patch still inherits every unrecorded change in the host.
Who should leave this gate unused?
Skip the free-session comparison if you need numbers for a paper, a vendor scorecard, or a launch post. A convenience sample cannot carry those claims, even when both cards match, because this method does not pin hidden weights or network conditions. Skip the method also when policy forbids shared hosts, customer code, or secrets outside a controlled network. A capture script does not make a shared disk compliant with a policy that forbids outside hosts.
Skip the method when you cannot read the current product page on the same day you run the session. Availability language in a blog post ages faster than a lockfile, and an outdated image name will poison the card. Teams that already own a frozen image and a pinned model build gain no rigor by moving that trial onto a free host. They would trade a sealed bench for a shared kitchen where other jobs can change the counters overnight.
The practical close stays modest, because a convenient host is not a reason to widen the claim. If your reviews already demand an environment card, add a free hosted session only when the live identifiers fit that card without guesswork. Otherwise keep the free run as a demonstration, and let the sealed comparison wait for a workspace that can hold still. That restraint is the whole method: record the treatment, compare only matches, and leave unpublished numbers unpublished.
Top comments (0)