I published a reinforcement learning environment on 23 September — a scored task that an agent runs against, used to measure or train models. Writing it was the second-cheapest part. Measuring three models against it is what filled that week's inference bill: $9.01, which is a ceiling on what the environment cost rather than an invoice for it. The $10 I had topped the account up with did not cover it, and working out why took me longer than it should have.
That number is why I am writing this. When I was deciding whether to build one at all, what I could not find anywhere was the bill. Plenty of guides on how to structure an environment; none that said what it costs to actually measure models against it once it exists.
So: the receipt, the token counts, the line that does not balance, the two setup traps that stopped the build cold, and an outage I was sure had cost nothing — until a reviewer showed me I had never actually measured that.
The bill
The provider is Prime Intellect, whose Environments Hub — the registry where environments are published and searched — is also where mine lives. One top-up, ever:
Sep 23, 2026 12:48 UTC $10.00 TOPUP STRIPE
Spend, from the provider's own analytics, for the window that contains the whole project:
SEP 19 - SEP 26, 2026
Inference 9.01
Wallet balance as I write this: -$0.48.
Those three numbers do not add up, and that is the first thing worth saying out loud. $10.00 in and $9.01 out should leave +$0.99. It reads -$0.48. $1.47 is missing. Under two assumptions I will name in a moment, it was already on the account before the window opened:
-1.47 balance when the window opened
+10.00 top-up
-9.01 inference, SEP 19 - SEP 26
------
-0.48 closing balance ✓
Here are the two assumptions. First, that the balance I am reading and the spend window close at the same moment — the window ends today, and today is not over. Second, that nothing moved the balance without appearing on the page I am quoting: no refund, no credit, no fee, no charge under another heading. Grant both and -$1.47 is forced; it is the only value that closes the line, since the top-up is the only credit in my history and Inference is the only spend line I am shown. Loosen either and the number slides — an unshown +$0.50 would put the opening balance at -$1.97 instead. So this is a residual under stated assumptions, not a transaction I can point at. My history has one entry in it, the $10.
This matters more than forty-eight cents, because it is also the reason the account ran dry mid-measurement, which is the next section but one. It is also why the balance could go negative at all. This provider allows an overdraft — it lets the balance run below zero by some amount before it starts refusing calls, and its own refusal message says "including overdraft" in as many words, below. It cuts you off once you pass whatever that floor is.
What $9.01 bought
Sixteen evaluation runs across three models — sixteen that wrote a results file, anyway; two more died before writing one, and I get to those. A rollout is one full episode: the model reads a case, calls tools, and gets scored. (162 of the 1,032 below never got that far. That is its own section.)
| model | runs | rollouts | input tokens | output tokens | in:out |
|---|---|---|---|---|---|
| claude-haiku-4.5 | 5 | 312 | 1,300,414 | 102,185 | 12.7 |
| claude-sonnet-4.5 | 6 | 408 | 1,974,546 | 164,250 | 12.0 |
| gpt-4.1-mini | 5 | 312 | 513,679 | 42,633 | 12.0 |
| total | 16 | 1,032 | 3,788,639 | 309,068 | 12.3 |
4.1 million tokens, in and out together, against $9.01 works out to about $2.20 per million, blended across the three. That blend is not a price list, and it is not even a firm measurement: the dashboard bills the account over an eight-day window, not the model and not the project, so $9.01 is a ceiling on what this environment cost rather than an invoice for it. I say more about that at the end.
The shape is the part I would carry forward, and it does not depend on the dollars at all. Input outnumbers output about twelve to one, and the three aggregates land close together: 12.7, 12.0, 12.0.
I first wrote that this is a property of the loop rather than the model — the episode re-reading its own growing transcript, the system prompt and the tool definitions and every previous result, on every turn. Then I went and checked, and the data does not carry that explanation. Per rollout the ratio runs from 8.2 to 26.8 (median 12.25, quartiles 11.1 and 13.5), and it barely tracks turn count: Pearson 0.087, across episodes of two to seven turns. Whatever holds the aggregate near twelve, it is not "more turns, more re-reading" — there are not enough turns for that story to be doing the work.
So the honest version is narrower. In this environment, on this task, three models' aggregate input-to-output ratios came out close to one another, and the per-rollout spread is wide enough that twelve is a planning figure rather than a constant. I do not have a tested explanation for why.
One task shape, though. A single-turn environment, or one with large tool outputs, will land somewhere else entirely. What transfers is the habit of budgeting for reading, not the number twelve.
One more thing the table hides. Runs come in two sizes — 54 rollouts and 96 — because I expanded the evaluation set from 18 cases to 32 partway through, and both sizes are three rollouts per case. Twelve runs at 18 cases, four at 32. That expansion is also what made my own notes wrong for two days, which is the previous post in this series. The directory listing was the evidence sitting in front of me the whole time.
The 162 rollouts I said cost nothing
162 of those 1,032 rollouts returned an error. Exactly 54 per model, three models. 54 is 18 cases times 3 rollouts — one whole run each.
Three models from two vendors failing at an identical count is not a model problem, and I did not have to deduce that. The worker logs said it outright, 54 times per failed run, while I was off reading saved rollouts trying to find the bug:
ERROR - Aborted rollout due to ModelError() -> APIStatusError(
"Error code: 402 - {'error': {'message': 'Insufficient balance
(including overdraft). Please add funds to continue. ...',
'code': 'insufficient_funds' ...}}")
(One line in the log, wrapped and abridged here; the timestamp and logger name, the rest of the provider's message, the error type, the request IDs and a docs link are cut.)
402 is the HTTP status a provider returns for payment required. The balance had hit its overdraft floor and every call was being refused at the gateway, before it reached a model.
Here is the sentence I had written, and it was wrong: those 162 rollouts used zero tokens.
They did not use zero tokens. They have no token record at all — the token_usage field is absent from every one of the 162 rows. My analysis script turned that absence into a measurement, because it said this:
usage = row.get("token_usage") or {}
inp = usage.get("input_tokens") or 0 # missing -> None -> 0
out = usage.get("output_tokens") or 0
None or 0 is 0. The script printed input tokens: 0, I read it as a measurement, and I built a section on it. Three separate review passes re-ran the same aggregation and reproduced the same zero, because they inherited the same defaulting. An adversarial pass that went looking for absent rather than zero caught it hours before this went out.
That is the failure this series keeps finding in my own work, and this time it is one line of Python: a missing value and a measured zero are not the same number, and or 0 erases the difference. If you take one thing from this post, take that one.
What the rows do record, and this part is measured:
error rollouts: 162
token_usage key present: 0 of 162
completion empty: 162 of 162
num_turns == 0: 162 of 162
total_tool_calls == 0: 162 of 162
No turn ran, no tool was called, nothing was generated. The whole 3.79 million input tokens came from the 870 rollouts that did.
So I still believe the outage was free, and I want to be exact about why. Not because I measured zero tokens — I did not measure anything. Because nothing happened that could have produced any, and a 402 is refused at the gateway before a model sees the request. That is an inference from an absence. I think it is a good one. It is not a measurement, and for two days I had it filed as one. Check your own provider before you lean on it; a request refused at a gateway is not free everywhere.
It was worth having anyway, because it exposed a real hole in my scorer: a run where the model was never invoked was still collecting points. Two of the three "no harm done" terms paid out on every dead rollout — no_duplicate_effects and no_unauthorized_payment, 162 of 162 — because nothing is paid twice and nothing exceeds a cap when nothing happens at all. The third, no_false_block, paid out on 72 of the 162 and not on the other 90. The dead rollouts score a mean reward of 0.344, range 0.300 to 0.400. Not zero. The previous post is about what I did to the scorer afterwards; this one is only about what the outage cost. (That post says the balance was sitting at zero. It was below zero: the refusal says including overdraft. And by the arithmetic at the top, under its two assumptions, the week opened $1.47 under.)
Check your balance before a measurement run, and check it again during. Not because running dry is expensive — as far as I can tell it consumed nothing — but because it is indistinguishable from your code being broken, and you will spend the afternoon on the wrong suspect.
Two runs I left out of every table above
While writing this I went back through outputs/evals/ and found eighteen run directories, not sixteen. Two have worker logs and no results.jsonl, so there is nothing in them I can recompute:
- one gpt-4.1-mini run aborted after three rollouts on the same 402, at 12:34 — before the top-up
- another gpt-4.1-mini run started at 12:56, after the money landed, and died fourteen seconds later with no 402 and no rollouts at all
Which means the tidy sentence I had first written — each model has exactly one run that failed completely — is false. gpt-4.1-mini has three — two that the 402 refused, and one that died on its own. The true count of rollouts refused at the door is at least 165, not 162, and the perfect three-way symmetry that made the diagnosis so satisfying is partly an artefact of counting only the runs that finished writing.
I am leaving the tables at 16 runs and 1,032 rollouts, because that is what reproduces from the files. I am telling you about the other two because they are the ones that made the story too neat.
Measurement note, 2026-09-27. The exact local recount is 162 error rows among the 1,032 stored result rows, all containing 402, plus two run directories with no
results.jsonl. One of those resultless directories contains three additional 402 worker-log lines. Therefore 165 is a lower bound on refused attempts, not 165/1,032: the three extra lines are outside the stored-row denominator, and the other resultless run has no countable result rows.
Where the $1.47 shows up
All four runs the 402 killed are timestamped 12:34–12:35 UTC. The top-up is 12:48. The outage is fourteen minutes older than the payment — the account was refusing calls before I paid a cent, which is what a negative opening balance would do.
Every rollout this project recorded before that payment came back 402 with no turns and no token record, and the first successful run starts at 12:58, after the money landed. So on the reading above, the $10 did not buy $10 of measurement: it cleared $1.47 of debt first and bought about $8.53, which is why $9.01 of inference finished the week forty-eight cents under water.
Two traps that stopped the build cold
Both are environment setup, both are boring, and both eat time out of proportion to their size because the error message points somewhere else. For calibration: the Python one shows up in my file timestamps as 49 minutes, 11:33 to 12:22 UTC — ending twelve minutes before the first eval run, and small only because I got lucky about where I looked second.
The Python version
verifiers, the library this kind of environment plugs into, declares:
Requires-Python: <3.14,>=3.11
If your machine's default python3 is 3.14, pip cannot install a release version. It does not stop there — it backtracks, walking down the candidate list until it finds something whose metadata fits, and what fits is a pre-release development build that declares <3.15.
That build is a different library wearing the same name. It drops the entire verifiers.legacy stack, which is where StatefulToolEnv lives and where the loader that goes looking for a load_environment in your package lives. (There is still a function called load_environment in the dev build, over in verifiers.v1.utils.loaders, but it takes a config object and is not the hook your environment plugs into.) So the environment simply does not load, and nothing in the traceback says wrong Python.
Build the environment's virtual environment — a venv, an isolated per-project Python with its own installed packages — with a version the library accepts:
python3.12 -m venv .venv # or python3.11 / python3.13 — whichever you have
python3.12 is the binary I happened to have. Substitute the one on your machine; if none of the three is installed, installing one is the actual first step, and python3.12: command not found is what that looks like.
I keep the created-from line as evidence now. It is already written for you, in .venv/pyvenv.cfg:
...
version = 3.12.0
...
command = /usr/local/bin/python3.12 -m venv .../.venv
(Three lines cut, and the absolute path shortened. Yours will echo back however you invoked it.)
Then pin the upper bound yourself in your own pyproject.toml, so a later install cannot wander back:
[project]
dependencies = [
# Upper bound on purpose. verifiers 0.3.2.dev* drops the whole
# verifiers.legacy stack: no StatefulToolEnv, and no loader that looks for
# a load_environment in this package. (v1.utils.loaders does define a
# load_environment, but it takes an EnvConfig and is not that hook.)
# pip will pull that pre-release when it backtracks — for example when the
# prime CLI is installed into the same environment. Pin until this targets
# that API.
"verifiers>=0.3.1,<0.4",
"datasets>=2.14.0",
]
That is copied from my repo, and it is copied from the second version of it. The first version said the dev build has no load_environment at all, which is wrong — it has one, it is just not the hook. Nobody would ever have checked a code comment. The same review that caught the token bug caught this, and I pushed the correction (bdf2317) before publishing, so that the comment you clone matches the one you are reading.
The CLI needs its own virtual environment
Half a reason looks like this. The platform's command-line tool depends on the same library your environment does, pinned hard to an older major: prime 0.7.6 ships Requires-Dist: verifiers==0.2.0, while the environment needs >=0.3.1. Install both into one venv and pip does not raise — it quietly picks an older prime that fits, and you get a CLI several releases behind with no error to tell you. Drop the pin instead and it resolves the other way, taking verifiers down to 0.2.0 and breaking the environment.
python3.12 -m venv ~/prime-cli && ~/prime-cli/bin/pip install prime # same substitution
Two venvs. It looks redundant until the day it isn't.
The generalisable bit: when a tool and the thing the tool operates on share a dependency, they do not share an environment.
Keep a path that runs without an API key
This one is design, not setup, and it is the decision I would make first again.
The package's __init__.py does not import the RL library at module load. It resolves it lazily, so importing the package costs nothing:
__all__ = ["load_environment"]
def __getattr__(name: str):
if name == "load_environment":
from .environment import load_environment
return load_environment
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
(Module docstring and __dir__ omitted.)
A module-level __getattr__ runs only when someone asks for an attribute the module does not already have, so the heavy import happens on first use or never. Import it eagerly at the top instead and every script in the repo inherits the requirement — including the ones a reader should be able to run with nothing but a clone and a Python.
The split in my test suite makes the payoff concrete. There are 57 tests: 21 on the dataset, 25 on the scorer, 11 on the wiring. 46 of them run with no install of the RL stack at all. (first recovery count, discarded 2026-09-27). There are now 58 tests: 21 on the dataset, 25 on the scorer, 1 on numeric reporting, and 11 on the wiring. 47 run with no install of the RL stack. Only the 11 wiring tests need it.
git clone https://github.com/jigonyoo/duplicate-side-effect-desk
cd duplicate-side-effect-desk
python3 -m pytest tests/test_dataset.py tests/test_grader.py tests/test_reporting.py # 47, no RL stack
python3 scripts/run_attacks.py # the attackers
(pip install pytest if you do not have it. Nothing else.)
That ratio is the point. The expensive dependency guards the part that genuinely needs it, and a stranger can check most of my claims for free. I hold it in place with a test that blocks the library at import time and asserts the scorer still loads.
Search before you build
The cheapest step has no bill attached. Before writing anything, search the hub you intend to publish to — for me, the Prime Intellect Environments Hub, which may want an account — with the actual keywords, and count what is already there.
When I searched on 23 September, the topic I had arrived with was saturated: thirteen environments matching injection, fifty-plus matching reward hack. The thing I actually wanted to measure — the same real-world side effect happening twice, which is what idempotency means for an agent that can retry — returned zero for idempotency, and zero for duplicate once I filtered to that sense of the word. That is a snapshot of one day on one platform, and you should take your own rather than trust mine.
Two things I would look for again:
- Count the crowded topic. If a dozen environments already measure it, a thirteenth needs a reason beyond wanting to publish.
-
Check whether the existing ones are already solved. On one of the nearest environments — another refund desk — its headline safety term,
policy_held, was already sitting at a perfect 1.000: a Claude Haiku baseline fell for zero payloads. An environment that a small model aces no longer separates models, whatever its description says.
And the selection rule that matters more than either: pick something that extends what you already have. Mine sat at the intersection of three small libraries I had already written. That is not modesty about scope — it is what lets the README say something specific instead of something general.
What this does not show
$9.01 is a window, not an invoice. The dashboard bills SEP 19 – SEP 26 to the account, and the runs that spent anything occupy 54 minutes of it (12:58 to 13:52 UTC; all sixteen span 78). Nothing in the repo proves I ran nothing else on that key in those eight days. What the $1.47 tells me is only that the account was already in debt before the window opened; it says nothing about what else may have run inside it. So treat $9.01 as an upper bound on what this environment cost, and $2.20 per million as an upper bound with it.
No per-model cost. The table has tokens, not dollars. The provider bills the account, not the model, so I cannot tell you what the sonnet rows cost against the haiku ones, and a blended rate will not recover it because list prices differ.
"Free" is not a measurement at all. The 162 refused rollouts have no token record and no dollar figure. I am inferring from zero turns and an empty completion that nothing was billed. That is an argument, not a receipt.
The -$1.47 is a residual, not a transaction. It is forced only if the balance and the spend window close together and nothing moved the account off-page. I cannot check either.
I have no explanation for the 12:1. Three aggregates landed close together; turn count does not predict the per-rollout ratio; the per-rollout range is 8.2 to 26.8. I report the observation and not a mechanism.
Two runs are missing from every table. At least 165 rollouts were refused, not 162, and the symmetry that gave the diagnosis away is partly an artefact of what got written to disk.
The token counts do not reproduce from the repository. outputs/ is in .gitignore. You can clone the repo and run the scorer and the attack suite — those recompute, and the commands are above — but the measurement runs are mine alone, and you are taking my word for the tables.
And this is one environment by one person, on one provider, measured over about an hour of billable time. The order that worked for me is not evidence that it is the order that works.
The environment, the scorer, the attackers and the 58 tests are public under MIT:
- github.com/jigonyoo/duplicate-side-effect-desk
- On the Environments Hub as
jigonyoo/duplicate-side-effect-desk(that page may want an account)
If you are budgeting for one
Ten dollars was nearly enough. It covered sixteen measurement runs across three models — three of them killed by a mistake that, as far as I can tell, cost nothing, along with a fourth that never made the tables — and a shipped environment. And it still finished forty-eight cents under, because $1.47 of it went to a debt I did not know the account was carrying.
If I were starting the next one I would put twenty on the account. Not because the work costs twenty, but so that hitting zero never again gets mistaken for a bug, and so the receipt at the end has nothing on it I have to reconstruct.
I use AI as a tool in my work and I disclose it. This post was written with AI assistance. The code, the measurements and the editorial calls are mine, and I do not publish a number I have not run.
Top comments (0)