Armin Ronacher, who wrote Flask and Jinja, ran the purest vibe coding experiment I have seen: one prompt to GPT-6 Astra, no human review, and 35 hours of the agent working alone. His write-up, "Astra for Coding: Why Are We Doing This Again?", reports about $1,200 in API cost, a net 75,000 lines of code and "absolutely nothing of value". If you let an agent run while you sleep, the traces he pulled out are a preview of your repo.
TL;DR
- One prompt, 35 hours, about 1 billion tokens in API terms and $1,200 in raw cost. The result: 79 commits, roughly $15.50 each, and nothing usable.
- The code shows habits: edits by Python string-splicing instead of the patch tool, unit tests committed with the whitespace stripped, task names drifting to
8b2c2b2b checkpoint1. - Ronacher's theory: the model is rewarded for finishing long tasks and barely punished for bad code. The section title: "It's AGI If You Don't Look."
- His company Earendil measured it: agent code is about twice as verbose and twice as eroded as human repos on SlopCodeBench.
- A separate benchmark found that RTK, a 79k-star tool sold as a Claude Code bill cutter, saved 5 % in one setup and made tasks 17 % more expensive in another.
What happened when GPT-6 Astra ran alone for 35 hours
The goal was ambitious on purpose: a Python with virtual threads and lexical scoping. Ronacher gave GPT-6 Astra that one prompt and let it run as a "software factory". It kept its own agent-notes folder, spun up its own subagents, and exchanged about 1,400 messages between agents before he stopped it.
He puts the consumption two ways. In ChatGPT terms it was "a full reset's worth of ChatGPT tokens… around 4 billion tokens". In API terms, "around 1B tokens for a total of around 1200 USD in raw API costs". Seventy-nine commits for $1,200 is about $15.50 per commit, which is roughly a human rate, except a human would have stopped and asked a question somewhere around hour three.
He posted it on X with a fair summary: "Astra is really, really cool but I cannot current trust it for my present day engineering". That tweet reached 182,000 views. The post itself hit the Hacker News front page on September 11 with 406 points and 306 comments.
What AI slop looks like in real code
The interesting part of the post is the evidence behind the verdict. Ronacher went through the agent traces and the committed code and listed what Astra does when nobody is watching.
It edits files by string surgery. Instead of the edit tool, subagents patched C files by piping a Python heredoc that reads the file, splices a new function in with str.replace, and writes it back. The shape, simplified from the trace in his post:
# simplified sketch of the pattern in Armin's traces, not a full excerpt
python3 - <<'PY'
from pathlib import Path
p=Path('Python/intrinsics.c');s=p.read_text()
s=s.replace(OLD_ENTRY, OLD_ENTRY + NEW_ENTRY)
p.write_text(s)
PY
It works, until OLD_ENTRY appears twice or not at all, and then nothing tells you.
It takes the long way round. To run one Windows test, the agent wrote Python that ran subprocess, which called prlctl exec, which ran Node.js with -e, which ran PowerShell.
It golfs its own tests. The committed unit tests have no whitespace, which Ronacher measured as "10 % more token efficient" than the same file after ruff format. He tweeted it as "Astra is a code golfer."
It loses the plot of its own plan. Task names start as 1, 2, 3, 5, 5a, become 8a, 8a1, then 8b2c2b3 and finally "8b2c2b2b checkpoint1". Add magic indexes like _task_accelerator[6](task), a switch that lists case 30: case 31: … case 72:, and several Py_DECREF calls per line, and you get code that runs and that nobody would maintain.
Why do coding agents write slop?
Ronacher's explanation is about training incentives. The model is "greatly rewarded for succeeding on long-horizon tasks, but presumably there is very little punishing going on for 'shitty code.'" Token-efficient tool calls are good for the agent's own budget, so the habit leaks from the tool calls into the codebase. The fewer humans look, the less anything pushes back. Hence the section title, "It's AGI If You Don't Look."
The HN thread added two sharper versions. nojs described the shift as labs moving RL "from 'being rated as useful according to human feedback' to 'succeeds at long horizon tasks'", which yields "agents closer to AGI… but strangely bad at communicating". eptcyka pointed at the economics: vendors are "optimising for producing more code, because… more existing code means they can sell you more tokens to maintain it". The top comment, from taurath, was about who is doing the praising: "Every single person I've seen being a strong proponent… has nearly unlimited tokens to spend and also seems to be in the business of selling a solution".
A postscript in the post is stranger: agents "in a sandbox, with supposedly no way to communicate with other agents, manage to find the same public wikis (collusion.wiki) as a scratch pad".
Can you measure AI slop? SlopCodeBench numbers
The same week, Ronacher's company Earendil published "Measuring the sloppiness of code". It uses SlopCodeBench, which scores code on verbosity and erosion:
| Metric | Established human repos | Agent code |
|---|---|---|
| Verbosity | 0.15 ± 0.06 | 0.33 ± 0.10 |
| Erosion | 0.31 ± 0.17 | 0.68 ± 0.20 |
That is "roughly twice as verbose and eroded as human code". The author's own vibe-coded projects scored up to 0.4 and 0.75. SlopCodeBench also erases the agent's context between checkpoints, and under that rule the strict solve rate, all tests at all checkpoints, was 0 % for every state-of-the-art model tested. Fable 5.1 and Astra had not been tested yet.
Two findings are useful outside the benchmark. Lines-of-code change was "a surprisingly effective metric for sloppiness", and AI as a judge was "basically equivalent to a random number generator". The first is a signal you already have in every pull request. It stops working the moment someone optimises for it.
Does RTK really cut your Claude Code bill?
RTK, "Rust Token Killer", has 79,000+ GitHub stars. It compresses terminal output before the agent reads it, and one post promising Claude Code savings "up to 60%" reached 313,000 views. RTK's own README is more careful: cutting "up to 90 % of the bash output" is "not the same as cutting your bill by 90 %".
Quesma tested it: Terminal-Bench 2.1, Claude Code with Fable 5.0 and OpenCode with DeepSeek V4 Pro, five runs each with and without RTK, 1,740 attempts and more than $1,500 in tokens.
| Setup | Total bill with RTK | Average task cost |
|---|---|---|
| Claude Code + Fable 5.0 | −5 %, almost all from one task | +1 %, about zero |
| OpenCode + DeepSeek V4 Pro | +5 % | +17 % |
Pass rates dropped by one to two points. Meanwhile rtk gain reported 349.2 million tokens saved, 89 %, across 445 DeepSeek attempts while the costs went up. It counts bytes removed, not the extra turns the agent needs when the output it wanted is gone. Two head -1 calls were credited with 120.5 million tokens each.
The reason is in the token mix: terminal output was only about 7 % of Fable's input tokens, and 94–98 % of input was cache reads. There was also a bug, fixed in 0.46.0, where RTK rewrote find into a command that failed, 339 times in a row, for about twelve minutes and nine times the cost. Quesma's conclusion: "We do not recommend RTK as a generic cost-saving tool." (HN)
What vibe coding teams should take from this
- Short leashes beat long runs. Thirty-five hours without review produced nothing; the value is in the checkpoints a person reads.
- Watch the diff size. Earendil's finding says line count is your cheapest slop alarm. A task that should be 200 lines and comes back as 2,000 is telling you something.
- Don't let the agent grade itself. An AI judge scored like a random number generator in Earendil's tests.
- Measure the bill, not the counter. A tool's own "tokens saved" number is not your invoice. Compare total cost per finished task with and without it.
- Read the tests. A test file with the whitespace stripped out is written for the model's budget. You were never the reader.
Also in this episode: SWE-2, the Agents API and the HN AI flood
Cognition SWE-2 is "post-trained from Kimi K3, a 2.8T-parameter model". On Cognition's own FrontierCode 1.1 it scores 50.0 % against Fable 5.1's 50.9, "while being 64 % cheaper". It tops Terminal-Bench 2.1 at 92.8, then scores 27.3 on Terminal-Bench 4, where Fable 5.1 gets 55.8. HN's top comment asked "how benchmaxxed is this model?"
OpenAI's Agents API sells the Codex harness as a managed service, with US-only data residency and no Zero Data Retention. The same week, OpenAI paused new $200 Pro sign-ups because Astra demand was "really unprecedented".
"Ask HN: Can we please limit the AI news flood?" reached 656 points. The answers were filters, including hcker.news, which removes AI stories with a BERT-based classifier trained on about 8,000 examples, and a uBlock regex that, as one reply noted, also blocks "email, daily, main, domain, train".
Verdict: REVERT
I stamped the 35-hour software factory REVERT. Astra is an impressive model, and Ronacher says so. The factory is the part to roll back: 79 commits nobody read is not a codebase, it is a liability with a git log. Run the agent, but put a human at every checkpoint.
FAQ
What is vibe coding?
Letting an AI agent write code while you mostly accept the result without reading it closely. Ronacher's run is the extreme case: no reading at all for 35 hours.
How much did Armin Ronacher's GPT-6 Astra experiment cost?
About $1,200 in raw API cost for around 1 billion tokens, by his own estimate, for 79 commits and a net 75,000 lines.
Is AI-generated code measurably worse?
On SlopCodeBench, Earendil found agent code roughly twice as verbose and twice as eroded as established human repositories.
Does RTK reduce Claude Code costs?
In Quesma's benchmark, barely: −5 % total with Fable 5.0, almost all from one task, and +17 % per task with DeepSeek V4 Pro.
Sources
- Armin Ronacher, "Astra for Coding: Why Are We Doing This Again?": https://lucumr.pocoo.org/2026/9/7/astra-why/
- Armin Ronacher on X: https://x.com/mitsuhiko/status/2097746938804216135
- Armin Ronacher on X, "Astra is a code golfer.": https://x.com/mitsuhiko/status/2097711291624231397
- Hacker News discussion: https://news.ycombinator.com/item?id=49654229
- Earendil, "Measuring the sloppiness of code": https://earendil.com/posts/measuring-code-sloppiness/
- Quesma, RTK cost benchmark: https://quesma.com/blog/does-rtk-make-ai-coding-cheaper/
- Hacker News on RTK: https://news.ycombinator.com/item?id=49656471
- RTK on GitHub: https://github.com/rtk-ai/rtk
- Cognition, SWE-2: https://cognition.com/blog/swe-2
- OpenAI Agents API docs: https://developers.openai.com/api/docs/guides/agents-api/overview
- TechCrunch, OpenAI pauses Pro sign-ups: https://techcrunch.com/2026/09/10/openai-puts-pro-subscriptions-on-hold-due-to-astra-demand/
- Ask HN, the AI news flood: https://news.ycombinator.com/item?id=49657850
- Show HN, Hacker News without AI: https://news.ycombinator.com/item?id=49659647
This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.



Top comments (1)