Claude Sonnet 5.5 edged Claude Opus 5.5 on our general-programming benchmark, a payments app, 0.7984 to 0.7926, a margin I read as level. It built the better app, though: it earned 0.9894 to Opus's 0.8732 before a ceiling squeezed them together. Then Opus 5.5 outclassed it on an Atlassian Forge app, 0.9767 to 0.5508.
Same two models, same goose, same 150-call budget, same week. The order flipped and the gap went from 0.0058 to 0.4259. And each model took exactly one critical defect, each on the other's home ground. Two settings did differ between the boards, reasoning effort and internet access, and I come back to both.
I think the reason is the most useful thing on our benchmark boards right now. Gauntlet asks for a kind of app that's all over the public internet. Forge asks for an app on a niche platform that changes month to month, and the model builds it with no internet. One measures a generalist. The other finds out who studied.
Below is every check behind both results, GPT-6.1 Sol and Sol Pro on the same boards, how the benchmarks grade a running app, and then the other thing we did this fortnight: trying to make our own specialist, a fine-tuned 27B model, with the whole bill.
Claude Sonnet 5.5 vs Opus 5.5 benchmarks, with GPT-6.1 Sol on both boards
These are the boards as of Monday 5 October, 04:59 CEST. Runs are still landing, so the Gauntlet board and the Forge board are the current word, not this table.
| model | Gauntlet 7.2 | Forge 1.0 |
|---|---|---|
| GPT-6.1 Sol Pro | 0.9574 | 0.9664 |
| Claude Sonnet 5.5 | 0.7984 | 0.5508 |
| GPT-6.1 Sol | 0.7978 | 0.9769 |
| Pareto 26.10 Preview | 0.7957 | — |
| Claude Opus 5.5 | 0.7926 | 0.9767 |
| MiMo-V2.6-Flash | 0.7909 | 0.7919 |
| Jev Router | 0.6975 | 0.7978 |
| DeepSeek Pro Latest | 0.6591 | — |
| GLM 5.3 FlashX | 0.6303 | — |
| DeepSeek V4.1 Flash | 0.5910 | 0.3354 |
| GPT-6 Luna | 0.4797 | 0.3950 |
| Qwen3.8 Omni Flash | 0.4269 | 0.5529 |
| GLM 5.3 | 0.1443 | 0.0292 |
| GPT-6 Luna Pro | 0.1343 | — |
| Qwen3.8 27B | 0.0915 | 0.1582 |
| Solar Mini 4 | 0.0083 | — |
Sixteen runs on Gauntlet, eleven on Forge. Every one is a single model through OpenRouter, driven by goose, our fork of the open-source agent Block started. Jev Router and Pareto 26.10 Preview are typesafe/jev-router and unbiased/pareto-26.10-preview there. Qwen3.8 27B, OpenRouter's name for Qwen3.8-27B, is the untouched base model our own fine-tune starts from, as OpenRouter serves it. Its two runs, and GLM 5.3's Gauntlet run, landed overnight. Two settings differ between the boards: Forge pins reasoning effort to medium and cuts the building model off the internet, while Gauntlet runs each model at its default reasoning effort with the network open. Call counts for the Claude and GPT runs in this post include one short Gemini call that shows up in each of their telemetry.
One caveat, said once and meant for the whole piece. Each model has one run per board. Under an earlier scorer version, one model on the same host scored 0.799 on one run and 0.505 on the next. So I read 0.7984 against 0.7926 as level, and 0.9767 against 0.5508 as a real difference in this run, with a cause the scorer writes down. Would Sonnet hit that double-click again on a second run? One run can't tell you. Without it, the gap would still be about 0.08.
If you only want the one-line advice for Forge work, three runs scored above 0.96: Sol, Opus and Sol Pro. Jev Router earned 0.9762 too, but a check I think misfired capped it, and it published 0.7978 (more on that below).
Twelve things the two boards and our training round showed
Each one is argued further down.
- The same two models can swap places when the task changes. Sonnet edged Opus on the payments app and lost to it by 0.43 on Forge.
- On Forge, the hard plumbing was a tie. Sonnet and Opus both scored 1.00 on the event pipeline, reconcile, storage, Rovo and lint tiers. Everything Sonnet lost, it lost at the edges: one button, one modal, one LLM feature.
- The scorer saw Sonnet's app make no Forge LLM call, and Forge LLM's catalogue is all Claude. A Claude model built an explain button that was meant to ask a Claude, and the question never arrived.
- Idempotency is general programming, and Sonnet knows it.
Idempotency-Keyon 3 of 3 payment sends; 0 comments on 2 of 2 double-clicks. Same idea, different room. - A ceiling compresses everything. Sonnet earned 0.116 more than Opus on Gauntlet and finished 0.0058 ahead.
- Sonnet and Opus write essays and GPT-6.1 Sol writes telegrams. On Gauntlet, Sonnet generated about 6,100 tokens per call, Sol about 530, and they finished 0.0006 apart.
- Pro won one of the three Pro-versus-plain match-ups we have. GPT-6.1 Sol Pro won Gauntlet outright, lost Forge to plain Sol, and GPT-6 Luna Pro scored under a third of plain Luna.
- A scorer can flatter itself by accident. On Saturday three different builds sat at exactly 0.7990. That was our bug, and the re-score fixed it.
- Our 96 GB Mac Studio can't fine-tune this 27B model on goose sessions. The measured ceiling is 12,288 tokens, and goose's fixed prompt alone is 32,569.
- A fine-tune inherits its parent's manners. T10's 35,412 training rows held no forum-style greeting, and it still greeted people by invented forum handles, because T9 had.
- The last gate is the one that counts. Our fine-tune had been chosen and was in release when a Forge app-building gate, run for the first time, found it built a complete app 9.7 times in 35. A fix round got it to 27.3, still under T9's 30, and it shipped with that gate waived.
- Renting GPUs to fine-tune this one model wasn't worth it on its own. $1,051.87 bought a model level with its predecessor on Atlassian multiple choice, behind it on Forge questions and better at tool calls; the predecessor took a day on a Mac we already own. Most of that money went on the round that didn't ship and the pipeline it ran on, and the next generation should cost a fraction of it.
Sonnet 5.5 vs Opus 5.5 on a payments app: a 0.0058 edge, and the better app
Gauntlet 7.2 asks for the Meridian Payments Console. Two cooperating services sync 12,288 payments from a mock vendor API, keep them consistent through webhooks, concurrent edits, crashes and partitions, and run a maker/checker approval workflow that creates real vendor payments. On top sits a console with a payments table, a notifications feed, a drafts panel and an interactive 3D field of payment towers. The 3D work carries 46% of the core score, and a browser probe checks it pixel by pixel.
It's a big app. It's also, in shape, a very familiar one: services with sync, webhooks, an approval flow and a dashboard. Every model on the board has read thousands of those.
Here's the pair side by side, from the run documents.
| Claude Sonnet 5.5 | Claude Opus 5.5 | |
|---|---|---|
| earned, before the ceiling | 0.9894 | 0.8732 |
| critical multiplier | 1.0000 | 0.8857 |
| ceiling | 0.799 | 0.799 |
| check that set the ceiling | overview legibility | label culling |
| model calls | 54 | 63 |
| tokens generated | 329.3k | 272.6k |
| model build time | 51m 16s | 56m 48s |
| final | 0.7984 | 0.7926 |
Sonnet's build earned 0.9894, the highest earned score on the board. That's above GPT-6.1 Sol Pro's 0.9574, the run that won.
Opus's build earned 0.8732. Most of that gap is one critical: a defect bad enough to multiply the whole score by a factor, down to a floor of 0.6. The full rules are further down.
Opus's app approved the payment and it never showed up
The scorer walks the approval workflow the way a person would. Log in as a maker, create a draft, submit it, log in as a checker, approve it, then look for two things: a notification, and the new payment in the table.
Sonnet's app passed all seven steps. Opus's passed five. Its approval went through, and the scorer then found neither the notification nor the payment in the table. The run page puts it in one line: "approval completed, but the created payment did not appear in the UI table".
That journey is a critical, because it's the thing a payments console is for. It multiplied Opus's score by 0.8857. The reject path had the same gap: rejected and listed, but no notification. Sonnet's showed one.
So on the part of the app a finance team would actually use, Sonnet built the better product.
Sonnet's 3D field ran off the edge of the frame
Then the 3D field. Both models built a working, interactive field of towers, and both lost exactly one ceiling check there.
Opus's labels failed a culling check: at the decisive camera pose, the set of labels on screen didn't match the set that should be visible. It scored 0.64 on that row.
Sonnet's failed overview legibility, at 0.8333. Its towers covered 0.6765 by 0.7351 of the canvas, and 104 tower pixels sat on the canvas edge. The probe reads that as a field that doesn't fit its frame. Opus's had 0 edge pixels. Sonnet's towers were also lower-contrast, 3.11:1 against Opus's 4.71:1, but GPT-6.1 Sol Pro passed this check at the same 3.11:1, so the edge pixels are what cost Sonnet the row.
Sonnet's other misses were small and specific. A forged webhook should bounce with a 401; the scorer recorded no status for Sonnet's app at all, though the forged change was left untouched. Its timezone buckets were exact in 380 of 384 cells. One of four text groups didn't put enough visible text on screen to read. Its written design decisions scored better than Opus's, 0.83 against 0.67.
I count that as Sonnet ahead on the product and Opus ahead on the picture. On a benchmark where the picture carries 46% of the core score, that should have been close. It was.
A ceiling turns a 0.116 lead into 0.0058
Sonnet earned 0.116 more than Opus and finished 0.0058 ahead.
That's the ceiling doing its job. Gauntlet's ceilings come in bands: miss a visible scene and the run is capped at 0.599, miss correct geometry at 0.699, 3D interaction or overview legibility at 0.799, event animation or full backend recovery at 0.899. Each further miss in a band takes another 0.03 off, never below the next band down. Passing a band adds nothing. Failing one caps you.
A capped run doesn't sit exactly on its cap. It lands a little under it, by 0.05 times whatever it didn't earn: final = ceiling minus 0.05 × (1 minus earned). For Sonnet that's 0.799 minus 0.05 × 0.0106, so 0.7984. For Opus, 0.799 minus 0.05 × 0.1268, so 0.7926. Gauntlet cuts its finals to four decimals rather than rounding them.
So under one ceiling, every 0.1 of earned score is worth 0.005 of final score. I think that's right, and I'd defend it. You can't buy your way past a 3D problem with backend points, and a model that misses the brief's visual bar shouldn't outrank one that clears it. But it's why the top of the board looks like a photo finish.
Essays and telegrams
One more contrast, the one I least expected. GPT-6.1 Sol hit the same ceiling as Sonnet, overview legibility, plus a backend recovery check. It earned 0.9768 and published 0.7978, 0.0006 behind Sonnet.
It got there in a completely different way. Sonnet made 54 model calls and generated 329,286 tokens. Sol made 106 calls and generated 56,390. That's about 6,100 generated tokens per call for Sonnet and about 530 for Sol. Opus sat between them at about 4,300.
Sonnet's average turn is a short chapter. Sol's is a paragraph, and it takes twice as many of them. Both styles got to the same place on a payments app. Sol also did it in 26 minutes against Sonnet's 51.
The token count matters for your bill, which I'll come back to. For the score, on this kind of task, it didn't. I'll take either style.
Opus 5.5 vs Sonnet 5.5 on Forge: the specialist, 0.9767 to 0.5508
Forge 1.0 asks for the Scope Ledger. It's a Jira app for agile coaches who want to know what entered a sprint after it started, who added it and how many story points it carried. It has to work on a Jira dashboard, from the sprint's own action menu and from Rovo. That's 10 Forge module types plus the Realtime API, and Forge LLM and Realtime are mandatory.
The model gets no internet. It gets a dev kit with Forge's own linter, the pinned packages and their typings, the manifest schema, Jira's OpenAPI spec and a mock Jira site it can run its app against. Everything the scorer measures about the platform is in there somewhere. What isn't in there is a tutorial, or anyone telling it how.
Opus built an app that scored 0.9767 with no ceiling, no critical and every correctness tier at 1.00. Sonnet's scored 0.5508.
Here's every tier, with GPT-6.1 Sol for reference. The weights are the share of the core score.
| tier | weight | Sonnet 5.5 | Opus 5.5 | GPT-6.1 Sol |
|---|---|---|---|---|
| L lint and bundles | .08 | 1.00 | 1.00 | 1.00 |
| K platform currency | .10 | 0.90 | 1.00 | 1.00 |
| T event pipeline | .16 | 1.00 | 1.00 | 1.00 |
| R reconcile | .14 | 1.00 | 1.00 | 1.00 |
| S storage | .08 | 1.00 | 1.00 | 1.00 |
| B resolvers and permissions | .12 | 0.83 | 1.00 | 1.00 |
| U UI function | .16 | 0.80 | 1.00 | 1.00 |
| V visual | .08 | 1.00 | 1.00 | 1.00 |
| A Rovo | .08 | 1.00 | 1.00 | 1.00 |
| core score | 0.938 | 1.000 | 1.000 | |
| critical multiplier | x0.60 | x1.00 | x1.00 | |
| final | 0.5508 | 0.9767 | 0.9769 |
Look at the middle of that table before the bottom. On the event pipeline, reconcile and storage, the parts of a Forge app that are actually hard to get right, Sonnet and Opus are twins.
Each run gets its own seeded Jira site, so the counts differ, but the verdicts don't. Sonnet handed 47 of 47 event deliveries to the queue correctly, Opus 40 of 40. Both recorded every live change exactly, duplicated and out-of-order deliveries included. Both parsed every Sprint changelog value that carries several sprint ids in one string (29 for Sonnet, 14 for Opus), which is a Jira quirk most people find out about in production. Both waited out a 30-second Retry-After inside the invocation, 30.2 and 30.3 virtual seconds, which the contract accepts alongside handing the retry back to the platform with a delay. Both found every issue that had been removed from every sprint, healed all 4 dropped events on the second scheduled run, and wrote nothing at all on a run with nothing new. Neither leaked a hidden issue to anyone.
On the event pipeline, reconcile and storage, you couldn't tell the two apps apart by score. I'd have taken either one's plumbing.
Sonnet lost 0.4259 in three places, all at the edges. One button, one modal, one LLM feature. To be straight about the arithmetic: its core score was 0.938 against Opus's 1.000, so most of the gap is one multiplier, the x0.6 from the double-click. By my arithmetic, fixing only that would have put it around 0.896, still about 0.08 behind Opus, mostly because of the LLM miss below, with the modal close and the comment flow making up the rest.
GPT-6.1 Sol vs Claude Sonnet 5.5 on Forge: a double-click that posted nothing
The sprint action opens a modal with a table of every change. You select a row and post a summary comment on that issue. The contract says a double-click must post exactly one comment, because people double-click, and a Forge app that comments twice on a customer's Jira looks broken.
The scorer double-clicked Sonnet's post button on two sprints. Nothing got posted either time: zero comments, zero success flags. Zero fails "exactly once" just like two does.
The check is a critical. Its label reads "duplicate side effect on a customer's Jira", and for an app that posts comments at all, any failure costs the full 0.6. 0.9180 times 0.6 is 0.5508.
Single clicks worked. All 4 of Sonnet's comments from single clicks were valid Atlassian Document Format, posted as the viewer, naming the issue key, the sprint and the creep. The scorer posts once with a single click, which worked, then double-clicks the same button on the same row. Sonnet's app posted nothing for the double-click, and showed no success flag. Either it refuses a second summary for an issue it has already commented on, or it clears the selection or disables Post after posting, so the double-click hit a dead button. I can't tell which from the scorer's rows. I haven't traced it in its code, and I haven't reproduced it by hand on a real Jira site.
Idempotency isn't Forge knowledge. It's the most general-programming idea in the whole task. And on Gauntlet, the same model's payments app sent an Idempotency-Key on 3 of 3 payment sends and reused it on the retry. Sonnet knows the idea cold. It just didn't land it on a Forge Custom UI button.
Qwen3.8 Omni Flash hit the same critical on the same check, 0.9214 down to 0.5529. Opus and Sol both passed it: 6 click and double-click posts, each with exactly one comment and one success flag, a forced rate limit included.
No Forge LLM call the scorer could see
The second miss is the one I find most interesting.
The sprint modal has an "explain" feature. The app is supposed to ask Forge LLM, Atlassian's hosted model API, to explain the sprint's scope creep, using a tool call so the answer comes back structured. Then it has to distrust the answer. The scorer's mock LLM answers five times in a row: a clean answer, one stuffed with made-up numbers and a hidden issue's id, a refusal, malformed tool arguments, and an API error. A good app shows the first, sanitises the second, and puts up an error flag for the last three without breaking the modal.
Forge LLM's catalogue is all Claude. Atlassian's models page says it "supports Claude models across three tiers: Haiku, Sonnet, and Opus", and lists eight ids, among them claude-sonnet-5 and claude-opus-5.
Opus's app made 5 LLM calls and passed 23 of 23 explain steps. Sol's did the same.
Sonnet's app had an Explain button, and the scorer clicked it five times. No call reached Forge LLM. Its explain check scored 0 of 5, and the model-id check had nothing to inspect, so it scored 0 as well.
So a Claude model built a Forge app whose explain button was supposed to ask a Claude model something, and the question never arrived. It isn't a joke at Sonnet's expense, and I can't tell you why it happened. Whether the code never asks, or the button never reaches its backend (its Close button never reached the bridge either), the scorer's rows don't say. And it isn't simply that Forge LLM is new. It went GA on 30 July, after an EAP and a Preview, and Sonnet handled the even newer modules, the dashboard widget, its edit bridge and the Rovo skill, without a scratch.
Those two LLM rows count as one defect in the robustness band, and the double-click is the second. Two defects cap a Forge run at 0.869. One defect caps it at 0.899, which is why the fixed-double-click version still lands under 0.9.
And a modal that wouldn't close
The third miss is small and very Forge. A Custom UI modal closes by calling the bridge's close. The scorer clicked Sonnet's Close button three times on its worst site, and none of the clicks reached the bridge. On a different one of its three sites, two of its ten comment-flow steps failed.
Opus and Sol: 3 of 3 closes, 8 of 8 comment-flow steps.
The screenshots show design differences too. Opus put the sprint totals in five cards and its buttons above the table, and printed dates a person can read. Sonnet put its totals in one line above the table, its buttons under it, and printed raw ISO timestamps. On a sprint with a long ledger, Sonnet's first screen is all table.
The one row where Sonnet beat both winners
To be fair to Sonnet, there's one row where it beat both Opus and Sol.
Excellence includes event economy: how many Jira calls the event path makes against the fewest it could get away with. Full credit stops at 1.46 times the optimum. Averaged over its three sites, Sonnet's app made 99 calls where 31 would do, 3.19 times. Opus's made 90 where 27 would do, 3.33 times. Sol's made 98 where 28 would do, 3.50 times.
So the model that lost by 0.43 was the most economical of the three on the one row where the winners lost points. It also built fastest of the Claude pair: 16 minutes 15 seconds, against Opus's 26 minutes 28.
GPT-6.1 Sol vs Claude Opus 5.5: 0.0002 apart on Forge
Opus is the closest thing to a tie on either board. Its Forge app had no ceiling, no critical and every correctness tier at 1.00. Like Sol's, it lost points only on event economy. Both passed the double-click that sank Sonnet.
Final 0.9767 against Sol's 0.9769. The whole 0.0002 is inside the excellence slice: Sol's excellence mean was 0.8075, Opus's 0.8057. That row averages three per-site scores, so its call ratios and its score don't line up exactly. I'd call it a tie. I won't lose sleep over 0.0002.
The two got there at different speeds. Sol built its app in 13 minutes 12 seconds. Opus took 26 minutes 28. Both wrote a Rovo skill that passed every instruction check, and both declared their storage index exactly the way the contract asks.
On Gauntlet, Sol and Opus were 0.0052 apart, both under the same 0.799 ceiling. I'd call them level on both boards. If I needed a Forge app tomorrow, I'd be happy with either, and I'd pick on price and speed.
The other ways to lose a Forge run
Two more Forge runs show the rest of the failure map.
Jev Router earned 0.9762 and published 0.7978, because its sprint-not-started view, one line and a button on an empty page, read as blank to the theme check, and that's a ceiling check: a screenshot that's 99.5% one colour counts as blank. Looking at the screenshot, the view is fine. I count that one as a false positive in our scorer, and it cost nearly a fifth of the score.
GLM 5.3 scored 0.0292. Its manifest pointed at UI build folders that didn't exist, and its backend wouldn't even bundle, because it assigned to a constant. Neither of its two functions loaded. Nothing downstream had anything to check.
GPT-6.1 Sol Pro vs GPT-6.1 Sol: Pro won one board, not the other
GPT-6.1 Sol Pro scored 0.9574 on Gauntlet, the only Gauntlet run with no ceiling at all and the only one marked excellent. It used 44 of its 150 calls and finished in 21 minutes. That's the fewest calls and the shortest build of the four frontier Gauntlet runs in this post.
It wasn't flawless. Its error state showed an error but gave you nothing to do about it, it skipped an optimistic paint, and it didn't reuse its idempotency key on a retry. None of those is a ceiling check, so they cost points, not a cap.
On Forge, plain GPT-6.1 Sol beat it, 0.9769 to 0.9664. Both lost only on economy. Sol Pro's backfill made 18 Jira calls where 15 would do, and its event path 3.62 times the optimum.
I wouldn't pay for Pro on faith, though. GPT-6 Luna Pro scored 0.1343 on Gauntlet against plain GPT-6 Luna's 0.4797. Its app never finished a first sync: when grading started its store held 0 payments. Two criticals, the sync and the approval journey, compounded to a multiplier of 0.3606.
Why a generalist slips on a niche framework
I want to be careful here. One run per model can't prove a theory about a model. But the pattern in the check rows is specific enough to be worth saying out loud.
Gauntlet is a big app built out of common parts. Sync loops, webhooks, idempotency keys, approval flows and a WebGL scene are all over the public internet. A model that writes good general code does well, and the four runs behind GPT-6.1 Sol Pro (Sonnet, Sol, Pareto and Opus) finished within 0.006 of each other.
Forge is a small app built out of rare parts. A few dates from Atlassian's changelog and npm, all checked on Sunday:
- Forge LLM went generally available on 30 July 2026.
-
rovo:mcpreached Preview on 14 August, and on 1 October it got a Preview that connects a Forge app's tools to external AI clients. - The new
dashboards:widgetmodule went GA on 22 September, and the oldjira:dashboardGadgetwas marked deprecated the next day, to be removed on 17 May 2027. -
@forge/dashboards-bridge2.0.0, the package a widget's edit view saves through, was published on 28 September. -
rovo:skillreached Preview on 2 October. Sonnet's run started on 4 October.
That's what the benchmark is built around. The design rule is to test recall of HOW, never of WHAT. The contract tells the model what to build, module keys included. It never restates Atlassian's docs. Everything else has to come from the model's own knowledge or from reading typings, schemas and an API spec inside a sandbox with no internet.
The design bets that a model that already knows the platform spends fewer calls and makes fewer defects, and that one which has to read its way in can still reach 1.0. Sonnet got to a correct ledger, a working widget, a Rovo skill and an MCP server. Where it ran short was Forge LLM, the bridge's modal close and one interaction bug. Not the newest modules, which it got right. So newness alone doesn't explain it, and I won't pretend it does.
Opus didn't run short anywhere. That's what I mean by a specialist: on the platform's own terms, it didn't drop a single correctness check. I don't know how much Forge each model saw in training, and nobody outside the labs does. I only know what each one built in a room with no internet. And specialist doesn't mean weaker at general work: on Gauntlet, Opus was level with Sonnet. What split them is that Sonnet's general skill didn't carry over to Forge in this run, and Opus's did.
And there's an irony I'll own. Our own fine-tuned model, further down this post, needed two Forge boosters, one to answer Forge questions and a late one to build Forge apps reliably, and its predecessor still beats it on both.
Two agentic coding benchmarks, graded offline in a real browser and an emulator
Both benchmarks share one loop, and I think the method is the most reusable thing we built.
A model gets one realistic product task, a budget of 150 model calls and no human help. When the budget is spent, or the model says it's done, whatever exists gets scored. Nobody reads the code to grade it. Apart from Forge's linter and static rules on the manifest and source, the scorer runs the app.
Scored, not refused
An empty or half-built app gets a score, never a refusal. That sounds obvious and took real work. A check whose precondition isn't met, say a "no duplicate comments" check on an app that never posted a comment, scores 0 instead of passing for free. Without that rule, by the Forge design's own arithmetic, an app that did nothing would have collected about 0.08 from all the "nothing went wrong" rows. With it, the empty starter scores 0.0 and an app with one do-nothing function scores 0.0051.
Graded offline in a real browser
Gauntlet's scorer starts both services against a fresh mock vendor, drives the console in a real browser, and records it. It kills processes mid-sync, partitions the network, forges webhooks, replays events out of order, and then reads the store, the API and the screen. The 3D field is checked by pixels and by the scene's own reported state, at several camera poses.
Forge's scorer does the same in an emulator. App code runs inside Atlassian's own Forge runtime wrapper, fetched from Atlassian's public CDN and pinned by sha256. If the hash doesn't match, the scorer refuses to run rather than fall back. There's an in-repo stand-in for developing the harness, and a verdict produced with it is marked unpublishable.
Every invocation runs under a deny-by-default macOS sandbox. App code can't read the scoring seed, can't spawn processes and can only reach the mock Jira site through the emulator's proxy. The proxy logs every call, and the scorer grades from that log.
The Custom UI is rendered in a bundled Chromium in light and dark, the widget at 380 and 1,180 px, with video recorded. The recording's SHA-256 is printed on the run page.
No internet on Forge, measured
On Forge, the building model gets no internet either. We measured the fence before trusting it. Inside the sandbox profile, curl https://developer.atlassian.com/ returns 000, a Node fetch to the npm registry fails with EPERM, and a server on 127.0.0.1 answers 200. goose still has to reach the model provider, so the harness runs a relay outside the sandbox whose allowlist is exactly one host: the provider's.
Without that, the newest-module checks would measure who fetched today's docs, and runs wouldn't be repeatable, because docs move.
Three seeded Jira sites, and the worst one counts
Each Forge run is scored on three freshly seeded Jira sites. Correctness keeps each check's worst site; excellence takes the mean. The model builds against a dev site with a different seed, so it can't memorise the numbers.
The sites are mean on purpose. Here's the trap list, every one a check:
- Issues removed from every sprint, which a reconcile using
sprint in openSprints()never sees. - Two estimation fields whose ids change per site, so a hard-coded
customfield_10016fails. - Parallel active sprints, for code that assumes "the" active sprint.
- Sprint changelog values with several ids in one string.
- Duplicate deliveries, each with a fresh event id, so deduplicating by event id doesn't work.
- Out-of-order and dropped events.
- The removed
/rest/api/3/search, which the mock site answers with 410. - 429s with a
Retry-Afterof 30 seconds on the event path and 2 on the reconcile. - LLM output trusted as-is, an unknown model id, or
temperatureandtop_pcopied from the Forge LLM README. Sending both is rejected for every model, and either one for four of the Claude ids. - Calling
publish()from a queue consumer, which the docs say isn't supported for async events (they send that code topublishGlobal). - Issue data broadcast on a global Realtime channel.
- Reading person-facing data as the app instead of as the user, so hidden issues leak.
- A double-click that posts two comments.
- A plain-string comment body, which Jira rejects with "Comment body is not valid!".
- Stale platform knowledge:
jira:dashboardGadget, storage from@forge/api, thenodejs18.xruntime,@forge/ui. - Saving widget config through a resolver instead of the dashboards edit API.
- Absolute asset paths, inline scripts and CDN fonts, which go blank under Forge's Content Security Policy.
- Hard-coded colours that vanish in the other theme.
- Reading changelogs one issue at a time across 237 issues, about 480 calls where about 12 would do.
People are already asking AI to write Forge apps, and these are the failures that only show up when the app runs. The clearest public example I found is an Atlassian Community write-up from August, vibe coding a WIP limit app with Rovo Studio. The generated app moved to the new search endpoint and lost the total count it needed, passed a maxResults: 0 that the author, reading the REST docs, found is illegal, and allowed the transition on any API error. So its first version deployed and enforced nothing. The author found that by testing it, listed "The validator itself is not preventing the transition." under what didn't work, and then fixed it with their own Forge and REST knowledge.
I couldn't find another public benchmark that has a model build a Forge app and then runs it.
Criticals, ceilings and why two models can't tie on a cap
I lean on two terms for every score on both boards.
A critical is a defect bad enough to multiply the whole score, down to a floor of 0.6. Forge has seven: an app that won't deploy, bundles that don't load, a change counted twice, changes silently missing after a backfill, a hidden issue shown to someone, a duplicate side effect on Jira, and a dashboard that shows no data. Some are graded by how badly they failed; the double-click costs the full 0.6 on any failure, as long as the app posts comments at all. Several criticals compound.
A ceiling caps the score when a required check fails. Forge's bands are deployable 0.499, a working ledger 0.699, current platform and complete surfaces 0.799, and production robustness 0.899, minus 0.03 for each further defect, never below 0.799. Nineteen checks sit in that last band, and a defect that knocks out several of them counts once.
What a build scores after criticals but before its ceiling is its earned score. If no ceiling bites, that's the final. If one does, the final sits just under the cap: ceiling minus 0.05 times (1 minus earned).
That last rule exists because of a bad Saturday.
Every run was re-scored on Sunday, and the ties went away
On Saturday three different models sat at exactly 0.7990 on Gauntlet. That looked like a coincidence. It was the scorer. The bands capped a run at a flat number no matter how much of a band it failed, and several probe checks were broken. One demanded the word "collar" where the contract says "Hollow currency frame". Others clicked the 3D canvas without aiming at it first.
goose 3.0.92 fixed that on Sunday morning. The bands are now graded, a capped run keeps its own gradient, Forge grades event handling against the true optimum on three sites, and the 3D probe waits up to 480 s instead of 230 s, because real builds took 258 s and 281 s and were being refused instead of scored. Then every saved build was re-scored and republished over its own entry, so every number in this post comes from the same scoring rules. The two Gauntlet runs that landed overnight were scored by later patch releases. Those only turn a build 3.0.92 would have refused into a low score. Nothing else changes.
The three 0.7990 runs are now 0.9574, 0.7978 and 0.7957. I wrote about how a scorer that can't flatter itself works in August. It still gave three different builds the same number. That's the bit I got wrong, and the re-score is the fix.
Before a Forge scorer version is used at all, a freeze gate has to pass. A reference app scores 1.0 on all three scoring sites, an empty starter 0.0, and the scoring thresholds are pinned by hash. On Sunday we re-checked the pin: the thresholds file's sha256 matches the constant in the scorer. There are also 31 deliberately broken copies of the reference app, each with exactly one defect, and each has to lose exactly the checks it's supposed to.
The app publishes its own scores
There's no submission form on our site. Results are posted by the goose desktop app over the site's API, and each run page shows the scorer's per-check output exactly as the app posted it, with the screenshots and the graded browser recording.
That's deliberate. I don't trust a board that only shows a number, and I don't expect you to. This one shows its working: open any run, expand a tier, and every check has its measured detail. When a scorer fix lands, "Re-score saved build" in the app republishes over the existing entry, so the board never carries a stale twin, and the page says it was re-scored.
If you want to check anything in this post, I'd start with Sonnet's Forge run. Open "Resolvers and permissions" in the scoring detail and the "0 comment(s)" line is the double-click check.
What a benchmark run costs on OpenRouter: the bills we have, and the ones we don't
Cheap enough to run every week, at least for the runs I have bills for, which was the point. My brief for these benchmarks said they "need to be cheap, economically viable", or we can't run them consistently.
The ten builds I hold OpenRouter bills for cost between $0.10 (Solar Mini 4) and $2.62 (GLM 5.3 FlashX), $9.53 together. GPT-6.1 Sol's Gauntlet build, a frontier model, billed $1.28 for 106 calls. I don't have bills yet for the other seventeen runs, among them both Sonnet runs, both Opus runs and both Sol Pro runs, so that's a lower bound, not a guess.
The token counts give you the shape, though. On Gauntlet, Sonnet read 7.06M prompt tokens and wrote 329k. Opus read 5.81M and wrote 273k. Sol read 5.40M and wrote 56k. Output tokens are priced above input tokens on these models' price lists. But the essay writers also read more. How much more they cost per run depends on caching, which is the next story.
Part of why these runs are cheap at all is a goose fix. goose used to rewrite the end of every request with the time, the working folder and a context-usage line, so GPT-6 models through OpenRouter never read the prompt cache. On one GPT-6 Luna run, 0 of 20.5M input tokens came from cache. After the fix, Luna's Gauntlet run, which spent its whole 150-call budget, billed $0.23.
A prompt cache only works if each request extends the previous one byte for byte. goose replaced last turn's context lines instead of leaving them where they were. Now it keeps them in the history and appends a new one only when something actually changes.
LLM fine-tuning cost: $1,051.87 to make our own specialist, and the first round didn't ship
Every model above is a cloud API. The other half of our fortnight was trying to make our own specialist: a local model trained for Atlassian and goose work. Neither of our fine-tunes is on these boards yet, and on Forge questions the newest one still trails the model it replaces. Here's what it cost to get it published, and everything that went wrong on the way.
Three names, because they're how the work was tracked. T9 is our published Qwen3.8-27B Atlassian v0.5, which the new model's card calls v1. T10 is the round that never shipped. T11 is the restart, published on Monday morning as Qwen3.8-27B-Atlassian-v2-goose.
T10's goals were better tool calling inside goose, keeping the Atlassian knowledge, my writing voice when a persona prompt asks for it, and never inventing people. That last one turned out to be the whole story. I said it on 2 October, when a run's replies started greeting people who don't exist: "does it still call fake names? That is a deal breaker."
The bill, in one line: the whole effort, T10 and T11, cost $1,051.87 before tax, $1,269.50 with AWS's VAT. The first plan said $100 to $160.
T10's round on its own came to $850.50 before tax, $1,025.83 with VAT, against a $1,000 pre-tax cap at the time. That's our spend watchdog's total when T10's training box was terminated on 3 October: logged instance hours times their rates, the teacher API bill (OpenRouter, $15.55 for the whole round) and the egress for checkpoints synced home. One box took $494.97 of it. About $64 went on failures.
Spend stood at $883.78 when T11 started, after one more corrective round. T11 itself, the model we actually published, cost about $160, and about $50 of that was a late Forge fix. Most of the rest went on T10, the round that didn't ship, and on building and debugging the training pipeline it ran on.
What is NOT in any of these figures. The Mac Studio's own work: the privacy treatment, building the training mix, a knowledge test on every checkpoint, quantisation and the release gates. The earlier rounds, T1 to T9. Our time and the AI assistants'. And the benchmark runs above.
Why we couldn't fine-tune our LLM on a Mac Studio: a 12,288-token ceiling
T9 was trained on one Mac Studio, an M3 Ultra with 96 GB, with MLX. On 1 October we measured whether the Mac could carry this round too.
mlx-lm trains this model's linear-attention layers with a per-token loop. At 2,048 tokens that used 77 GB. At 4,096 it crashed. We wrote a chunked version of that path and got the fit up to 12,288 tokens at 69.7 GB and 91 tokens a second. 14,336 needed 78.5 GB.
A goose training row is a whole agent session. goose's own system prompt plus its tool schemas come to 32,569 tokens before the conversation even starts. Across the first 732 rows the median was 38.2k tokens. The part the model actually learns from, its next move, is a median of 179 tokens.
So no goose row fits on the Mac. Not one. The same arithmetic rules out an 80 GB H100: 32k only fits with activation offload, and 64k ran out of memory in every mode. We settled on 65,536-token rows. That needs an H200 (141 GB) or a B200, and on AWS those come as 8-GPU boxes. An 8xH200 box has 1,128 GB of GPU memory.
The Mac still did a lot of the round. It just couldn't train the rows that mattered. If your rows are shorter, Gabriela on our team wrote up the Mac route as fine-tuning Qwen3.8-27B with LoRA on a Mac.
A LoRA, not a full fine-tune: rank 128, a KL anchor and a guard on the eighth GPU
It's a LoRA, not a full fine-tune. Rank 128, on the attention, linear-attention and MLP projections of layers 32 to 63, 400.6M trainable parameters, continuing T9's own adapter. We did try a full fine-tune as one arm of the first pilot. It ran out of memory on the box.
Three things sit on top of the LoRA.
A KL anchor penalises drifting from T9's own predictions on the rows meant to preserve behaviour. Instruction following kept dropping in every pilot, and the anchor's strength was the only lever that moved it.
A guard. Seven GPUs train and the eighth runs checks every tenth of the run, including five sample answers that have to be read and quoted before the run may continue.
And selection. Every candidate is tested paired against T9 on the same items, and it's ineligible if it invents a single person.
The mix was 35,412 rows and 50,010,670 tokens, 79% of the tokens in goose agent rows. Those came from a factory. T9 drives a synthetic task inside goose with real tools in a sandbox, and where it slips, a teacher model (DeepSeek V4.1 Flash, and MiMo-V2.6-Pro for Atlassian admin and migration tasks) shows the right move. No Claude output went into the training data.
I learned that filtering isn't enough. Every row was scanned against a privacy dictionary and dropped on a hit, never redacted. Then independent readers went through all 1,361 goose rows and dropped 404. Rows sampled from T9 itself were 38% bad. One synthetic set ended 1,650 of 1,650 rows with a made-up "(Forge docs: ...)" citation, which was where T9's habit of inventing citations came from.
The H200 Spot price we paid, box by box
All fifteen instances the T10 round rented, as logged. Spot unless it says otherwise.
| box | region | hours | $/h | $ | what it was for |
|---|---|---|---|---|---|
| p5.4xlarge (1x H100) | us-east-2 | 0.858 | 2.66 | 2.28 | platform probe |
| g7e.2xlarge (1x RTX PRO 6000), on-demand | eu-west-2 | 0.776 | 5.88 | 4.56 | platform probe |
| p5.48xlarge (8x H100) | us-west-2 | 1.396 | 21.07 | 29.41 | baselines and throughput |
| p5.48xlarge (8x H100) | us-west-2 | 0.361 | 21.07 | 7.60 | launch attempt |
| p5en.48xlarge (8x H200) | us-west-2 | 0.341 | 27.32 | 9.30 | launch attempt |
| c7i.8xlarge (CPU) | us-west-2 | 0.158 | 0.56 | 0.09 | builds the Linux goose binary once |
| p5en.48xlarge (8x H200) | us-west-2 | 3.420 | 27.32 | 93.45 | data factory run 1 (732 rows) |
| p5.48xlarge (8x H100) | us-east-2 | 0.407 | 20.75 | 8.45 | launch attempt |
| p5.48xlarge (8x H100) | us-west-2 | 0.355 | 21.07 | 7.49 | launch attempt |
| p5e.48xlarge (8x H200) | eu-north-1 | 1.007 | 17.73 | 17.84 | first 64k-capable box, lost to a hand-over bug |
| p5e.48xlarge (8x H200) | eu-north-1 | 0.269 | 17.73 | 4.76 | duplicate launched by that bug |
| p5e.48xlarge (8x H200) | eu-north-1 | 27.923 | 17.73 | 494.97 | pilots 1-4, factory rerun, two training runs, knowledge eval |
| p5e.48xlarge (8x H200) | eu-north-1 | 2.253 | 15.54 | 35.01 | final run, segment stopped by its guard at step 1012 |
| p5e.48xlarge (8x H200) | eu-north-1 | 0.534 | 15.54 | 8.31 | resume attempt, crashed on restore |
| p5e.48xlarge (8x H200) | eu-north-1 | 07:30 to 14:26 UTC on 3 Oct | 15.37 | final run from step 1012, selection, corrective rounds, verdict |
The kept box took $494.97 over 27.9 hours. It ran all four pilots, the second factory run and the two earlier training runs. The data factory's own logs put its two runs at $66.74 and $57.17 for 2,420 rows, teacher API included; their box share was $63.95 and $51.95, because the factory counts only its own minutes on the box, where the table bills the whole box.
What the failures cost, summed from the table. Four short launch attempts that the logs give no other job: $32.84. The box lost to the hand-over bug plus the duplicate it spawned: $22.60. The resume that crashed: $8.31. Together $63.75. The segment stopped by its own guard ($35.01) isn't in that sum, because the run resumed from its step-1012 checkpoint and kept its training.
Why so much H200. The 64k rows need it. AWS's public price list has no on-demand p5e at all, and an on-demand p5en lists at about $63 an hour in Ohio and $68 in Stockholm. That's roughly four times what we paid for a p5e on Spot in Stockholm, and more than twice our p5en Spot rate.
So I went with Spot, and getting capacity was harder than paying for it. On 1 October the Stockholm box showed up after two hours with no 64k capacity anywhere else. On 3 October the launcher found no p5e or B200 capacity in any region it walked, seven times in a row, 15 minutes apart. No box was ever reclaimed by AWS in our logs. When we did get eu-north-1, it was $15.37 to $17.73 an hour, against $27 to $55 in the other regions.
The cap moved. It started at $800. I cut it to $450 the same afternoon, because I don't want to give budget to burn just to burn. Then I raised it to $650 and $800 on 1 October, $900 on 2 October and $1,000 on 3 October. The plan's first estimate was about $100 to $160. Three independent stops guarded the money: a watchdog on the Mac, a dead-man timer on the box, and an AWS Budgets action.
Why T10 never shipped: it kept inventing people
The first full run, 2 October. The training mix carried my name in 222 of 651 contrast replies that T9 had written, 73 of them as a bare sign-off.
The relaunch trained to the end and chose nothing. 142 of 160 replies from its candidates greeted an invented full name. Both causes were ours. Every voice row's persona line carried my name. And the privacy treatment of my own replies, a separate step from the dictionary scan, had replaced real requesters' names with invented full names, so the model learned to greet someone. The rule now is to pseudonymise a name to nothing, never to a fake one.
The final run, 3 October, stopped itself at step 1012 on two numeric breaches of its clean-loss check. No loops, and the number was falling, so I let it run. The resume then crashed on restore, because every rank loaded rank 0's optimizer state onto GPU 0 and ran it out of memory. The fix was one argument, map_location=cpu. It cost a box.
Then a knowledge drain. Later checkpoints sat 3.5 to 8 points under T9 on 199 Atlassian multiple-choice questions, and that proof stopped training at step 3,541 of 5,059. And then selection found that every checkpoint invented people in its ticket replies, 17 to 47 of 160 samples. Our own voice data had none of it, 0 of 135. T9 did it in every one of its stored samples.
Eight short corrective rounds later, one version invented nobody in 320 samples and kept its tool-call gains, at about 4 knowledge points on the home test and 1.85 IFEval points. I approved it for release with every one of those numbers on the card. The release gates then had their say.
Fine-tuning hallucinations: how the release gates caught invented forum handles
The release runs its own gates on the 8-bit build it's about to upload. Most passed: no loops at 32k or 128k context, needle recall 100% at 4k, 32k and 128k, no leaks in 112 outputs, and knowledge about a point under T9's own 8-bit build (70.85 against 71.86, not significant). Three things failed.
Manners. Asked 20 plain admin questions with no persona, it opened 4 of 40 replies by greeting a forum handle that doesn't exist, the way an Atlassian Community answer starts. T9 does it in 7 of 40.
Fabrication. 309 privacy extraction prompts found no memorised client data. But asked about a named company, it made things up 49 times: ticket keys built from the company's initials, a site URL built from its name with a made-up cloud ID, incidents that never happened. T9 did the same kind of thing 60 times.
Loops. 2 of 160 prompts ran into an endless list of invented identifiers, both on Forge coding prompts. T9 had 1.
So T10 did two of the three less often than T9, on small counts, and it still didn't ship. Fabrication and loops were hard blockers for this round, so being less bad than T9 wasn't enough.
The handles came from upstream. T9's own training data included 220 Atlassian Community threads, and 84 of the answers open by greeting the asker by handle. T9 learned Atlassian partly from forum threads, and learned to say hello like the forum too. T10's data had none of that, 0 in 35,412 rows, but T10 started from T9's adapter and was trained to stay close to T9. It inherited the habit.
One more corrective round, about $26 of GPU, pushed it down to about 1 in 80 sampled answers, from about 1 in 8. Corrective rounds push a habit down. They can't erase what the starting point carries.
T11 started from the base model, and cost about $160
So I had T11 start from the original Qwen3.8-27B, not from T9, with a fresh LoRA and no pull toward T9.
First we tested the premise, on the Mac, for nothing. The untouched base model got the same manners test that caught T10: 0 invented names in 120 replies. So the forum greeting came from T9's training, not from Qwen. The base has its own faults, measured the same morning: it loops more than T9, and asked about a named company it adds public facts or other organisations.
That base model is now on both boards above, as OpenRouter serves it, driven by goose like everything else: 0.0915 on Gauntlet and 0.1582 on Forge. On Forge its dashboard widget showed no data and its backfill silently missed changes, two criticals that took its multiplier down to 0.41. On Gauntlet no rows reached its console, no amounts rendered and its approval flow couldn't finish, three criticals and a multiplier of 0.22. I read that as the base model, building a whole app alone in goose, failing at exactly the parts a person would use. That says nothing yet about T11: T11 hasn't been run on either board, and I won't guess where it would land.
Then the data. Every training example T9 had written, 24,549 of them, was regenerated by the base model on the rented box in 13 minutes and filtered: 631 refused for knowledge the source didn't support, 43 for invented admin facts, 32 for naming a person, 26 for forum shape, plus 191 naming someone not in the prompt and 5 private-dictionary hits caught at home. A whole-mix name scan now fails the build if anything slips through. Run on T9's data, it catches 131 of the 220 Community samples.
It trained on one 8x B200 Spot box at about $14.4 an hour, 7,157 steps in about 3 hours, with the name, fabrication, loop and knowledge checks running on every checkpoint on a spare card. The automatic stop fired once, at step 2147, on detector noise: a licence type, a role and an Atlassian product, not names. We fixed the detector, turned the automatic stop off and read every check by hand.
T11 had a Forge problem too
Two real problems showed up while it trained. Asked "Are you Claude?" or "Are you GPT-4?", it said yes 3 of 8 times. And its Forge knowledge stuck at 53-63%, where T9 scores 90% at full precision.
That second one is the same story as the top of this post, from the other side. Starting from the base model, T11's general coding held up fine (HumanEval 96.3 against T9's 95.7). Forge, it had to be taught.
A 26-minute booster fixed the first problem and helped the second: 760 question-answer rows from Atlassian's own Forge documentation, plus Forge, identity and replay rows. Ten of the 30 Forge test questions were held out of the booster entirely, seeded before it ran, and it scored 80% on those ten. So the gain wasn't only memorised answers, though T9 got all ten.
I made one mistake that cost a re-run, about $50 and 3 hours. The box finished and shut itself down with nothing exported: the booster's checkpoints hadn't been copied off it, and the automatic pick had refused every checkpoint on detector noise. Finding a new box took 72 minutes, with no Spot capacity in Stockholm, and we ended up in Virginia at $34.55 an hour. This time every checkpoint was copied the moment it was saved.
The loop that wasn't
One Forge prompt looped on 3 of 8 samples from T11, and T9 hadn't looped on it in the one sample we had. That looked like a regression. I asked for a fix round first.
Before training anything, the fix round measured that prompt 16 times on both models. T9 looped 2 of 16. T11 looped 1 of 16. The earlier comparison was one sample against eight. The fix itself, 43 steps on 15 self-written Forge apps, made it worse: 4 of 16, and Forge down to 70. By the rule written before the run, nothing was exported. That cost about $7 and taught the cheapest lesson of the fortnight: compare a looping prompt on the same number of samples for both models before you pay to fix it.
The money for that day: I raised the cap to $1,100 and then to $1,250. Start to finish, the cap went $800, $450, $650, $800, $900, $1,000, $1,100, $1,250.
The late catch: T11 could barely build a Forge app
T11 was chosen, measured and in its release gates on the Mac when the last of them ran. It asks for 35 Forge apps, 3 tries each, and checks every one with Forge's own manifest linter, its module allowlist and the TypeScript compiler. That gate had never been run on this round before release.
On average over the three tries, T11 built a complete, valid Forge app for 9.7 of the 35 requests. T9 builds 30. It wrote manifests missing required fields, module keys and permission scopes that don't exist, and UI components that aren't in Forge's library.
The cause was in our mix. T9's validated Forge app code had been diluted to about 4% of T11's 50,000 rows, where it was 19% of T9's, and three of T9's Forge-code sets were missing altogether. Being able to answer Forge questions isn't the same as being able to build a Forge app.
I asked for the fix first. A 388-step booster from the chosen checkpoint on 2,127 validated Forge code rows, plus 22% replay, ran for about 1.5 hours on an 8x H100 box, because the H200 and B200 class had no capacity anywhere. Rows that taught any held-out test item, or matched a test brief, were removed first, and so were 206 rows containing client names. It cost about $50.
The fixed model, checkpoint 7636, builds a complete Forge app for 27.3 of the 35 on average (T9 30) and a valid manifest for 29.3 (T9 31.3). Its goose agent defect rates came out lower than T9's again: no edit without a path, 0 against 1.7%, missing required arguments 0.4% against 1.6%, and repeated calls 2.3% against 4.1%. None of those differences is significant, so I don't claim them as gains. It paid for that elsewhere. On the box, Forge knowledge questions fell to 63 from 73 before the fix (T9 90), Atlassian multiple choice slipped to 76.4 from 78.4, and on the 2 held-out scenarios where it should ask the user first, it never asked. T9 asked in a quarter of its samples. I chose app building, and that's the model we published.
A new generation should now cost a fraction of the first
The expensive part was everything before T11, about $880: T10 itself, and building and debugging the factory it ran on. With it built, I estimate a clean round on a new base model at one box for 5 to 6 hours, about $70 to $90 of rental, plus the release on the Mac. No such round has run yet.
I've put the rules these two rounds paid for into the recipe. Copy every checkpoint as it's saved. Never let detector noise stop a run. Keep one box for every phase. Measure the starting model first. Put known fixes into the main mix. And compare a looping prompt on enough samples for both models.
T11 against the published v0.5: better at tool calls, level on Atlassian, behind on Forge
T11 went public on Monday 5 October at 02:29 and 02:41 CEST, as an 8-bit MLX build and a PEFT LoRA adapter. Each was uploaded private, downloaded back, checked byte for byte, probed for private data on the downloaded files (309 probes, 0 hits), and only then made public. The 6-bit and 4-bit builds pass their quality gates but stopped at the privacy scan: the 6-bit conversion re-saved the tokenizer file and the scan found a dictionary term in it. They're held, not uploaded. The model cards already list them, and those two links won't open until they're public.
Here are the numbers I'd quote, from the model card. The rule since T10 has been that the release's fabrication and privacy checks are hard blockers, and any other failed gate can only ship as a named waiver on the card. Loops were a blocker for T10. For T11 I let them through as a waiver. In all, eight of T11's release gates failed. All eight shipped as named waivers on the card. They are loops, Forge app building, Atlassian knowledge against T9 (on the Forge questions), voice, one general-skills leg, the ask-first goose scenarios, 8-bit agreement on long agent contexts, and the LM Studio load test.
People and privacy: no invented names in ticket replies, where T9 greets an invented person in 18 to 20 of 20. On the same 309 privacy probes T10 was run through, 0 hits. T9 produces 57 (60 before a pre-registered rule for spelling variants), and T9's are the invented ticket keys and URLs described above, not memorised data. One goose output did use my name where nothing in the task called for it, and that's on the card too.
Fabrication: one invented date in 320 answers about named organisations, on the full-precision model. That one comes from the full-precision checks on the training machine, and it's on the card. The base model invented 28 details on the same test.
Loops: not clean. 2 loops in the 160-answer battery against T9's 1, every leg at or below the base model's, and it shipped under a waiver, under my decision to go ahead. With T10, the same 2 against 1 was part of why it stopped. The difference is scale: T10 invented people and 49 company details, and T11 invented no people and one date in 320 organisation answers.
Atlassian knowledge, on the shipped 8-bit builds: multiple choice 74.4 against 71.9, a little higher but not a significant difference, so I call it level. Forge knowledge questions: 19 of 30 against 26, which is significantly worse. Complete Forge apps: 27.3 against 30 of 35, on average over three tries. That's the honest shape of it: in the vicinity of T9 on Atlassian knowledge, worse at Forge.
Tool calling is where it's better, measured on the full-precision model on the training machine. On BFCL, a public function-calling benchmark, 91.8 against 85.7, and declining tools it shouldn't call 85.8 against 76.7, both significant. On the 294 held-out goose scenarios, every measure came out non-inferior or better except one: in the 2 scenarios where the right move is to ask the user first, it never asked. T9 asked in a quarter of its samples. That gate shipped as a waiver too.
General ability, also full precision: MMLU 84.3 against 83.0, IFEval, an instruction-following test, 81.5 against 82.4, and HumanEval 96.3 against 95.7. The general-skills gate still failed on one leg, a thinking-format test, 50.3 against 52.3, and that's another waiver.
Its voice is much closer to mine than T9's on our style measure, a combined distance of 0.65 against T9's 3.07. It still misses my range on four features. It asks questions in 15% of replies where I do in 29 to 46%, greets in 31% where I do in 35 to 54%, signs off in 6% where I almost never do, and uses exclamations in 6.2% where my top is 6.1%. T9 signed off in 58%. The voice gate counts those four misses, so it's a waiver too.
One more limit. The 8-bit build matches full precision on general text, but on goose agent contexts up to 16,384 tokens long it drifts: a KL divergence of 0.400 and top-1 agreement of 89.9%, against a gate of 0.067 and 90%. When we diagnosed the same drift on the earlier checkpoint, the rows that drifted also drifted between MLX and the PyTorch reference at full precision, so the framework is part of it. The PEFT adapter on the bf16 base is the reference. Every waiver, this one included, is on the model card, and I'd read it before you download anything.
Every failure, what it cost and what fixed it
I've told the big ones already. Here's the whole list for both rounds, 30 September to 5 October, in order. Where we measured what a failure cost, in money or time, it's there.
T10, continuing from T9:
- The plan underestimated cost by 6 to 10 times. It said $100 to $160. Every stage needed 8-GPU boxes for 64k-token rows, and Spot capacity for those was scarce. I moved the cap seven times across both rounds, five of them during T10.
- Hardware assumptions. A single H100 or RTX PRO 6000 couldn't train 64k rows, and the Mac tops out at 12,288 tokens. Fix: one fixed 65,536-token length, 8x H200 or B200 boxes only.
- Capacity. Two hours waiting on 1 October, seven empty 15-minute walks across regions on 3 October, and a launch hand-over bug that lost one box and started a duplicate, $22.60 between them. Fix: the launcher refuses instead of walking, and a grab loop now catches boxes.
- All 732 rows from the first factory run came from 2 templates, because briefs ran in id order and the spend stop hit after 65 of 1,483. The re-run for diversity cost $57.
- The privacy treatment of real sessions was too slow, and the session bundle failed its audit. T10 trained on voice data only.
- My name in the training targets: 222 of 651 contrast replies T9 had written named me or signed as me. Fix: a hard rule that no target may name me.
- The relaunch chose nothing on 2 October. Every candidate named me or greeted an invented person in most replies. Causes: the persona line named me, the privacy treatment had swapped real names for invented ones, and T9's habit came through its adapter.
- A filter that hadn't been pushed was missing on the box, and two parts failed after we'd rented it. Fix: the box gets only what's pushed and tested.
- A guard stop at step 1012 on numeric-only breaches, then a resume that crashed on restore because every rank loaded the optimizer onto GPU 0. Cost: a box, $8.31. Fix:
map_location=cpu, numeric stops off, every guard read by a person. - A knowledge drain, found late. Later checkpoints lost 3.5 to 8 points of Atlassian knowledge, and selection had never measured knowledge. Fix: knowledge on every checkpoint in selection. My ruling: a small drop is fine, 20 to 30% isn't.
- Invented people in every checkpoint, 17 to 47 of 160 replies. Eight short corrective rounds got one version to 0, and I approved it.
- The release gates stopped it anyway: invented forum handles, 49 fabricated company specifics, 2 loops. Root cause: 84 of T9's 220 Community answers greet the asker by handle.
T11, a clean model from the base:
- Detector noise stopped training at step 2147. A licence type, a role and an Atlassian product were read as invented people. Fix: a stoplist for the detector, numeric stops off, every guard read by a person.
- During the manual resume we killed the wrong process, the guard's server instead of the evaluator's. About 20 minutes of re-measurement.
- The box shut down with nothing exported. The booster's checkpoints hadn't been copied to storage, and the automatic pick refused every one on detector noise. About $50 and 3 hours to re-run. Fix: every checkpoint copied the moment it's saved, and the pick made by the coordinator, not automatically.
- A misread clock made training look slow. Fix: always check the real clock.
- Capacity again. 72 minutes to find the re-run's box, and at one point nothing in six regions, Spot or on-demand. A grab loop caught one.
- The Mac's GPU timed out merging the adapter. Fix: merge on the CPU, about 2 minutes.
- LM Studio on the release machine indexes no local models at all, so the LM Studio load gate could never be tested. Waived, on the card.
- An unfair comparison. One T9 sample against eight T11 samples made "T9 never loops on this prompt" look like a fact. The fix round I asked for, about $7, made it worse. Measured fairly, T9 looped 2 of 16 and T11 1 of 16. Fix: equal samples first.
- The 8-bit build drifts from full precision on long agent contexts. Diagnosed on the earlier checkpoint as largely the framework, waived and stated on the card.
- Forge app building collapsed, 9.7 of 35 against T9's 30. Cause: validated Forge app code diluted to about 4% of the mix and three Forge-code sets missing. Nobody saw it until the gate first ran at release. The fix round I chose cost about $50 and got it to 27.3, and it cost Forge question answering.
The release night, 4 to 5 October:
- The second release shared the first one's private output folder, so three checks, manners, privacy and voice, silently reused the earlier model's answers. We caught it because they finished in seconds instead of minutes, and re-ran them on the right model.
- More detector false positives: a file format, a role and a code annotation read as people. Added to the stoplist.
- The card generator only knew "main run plus booster", so the Forge-fix stage was written into the card by hand and checked line by line. A card review found nine defects first. The biggest was a false claim that the 8-bit build matches full precision.
- Hugging Face writes its own
.gitattributes, which our upload check read as tampering. Excluded, and every model file matched. - A stale figure from the earlier checkpoint stayed in one card note after publishing. Caught by a review of the site page and corrected on the live cards.
- The 6-bit and 4-bit uploads stopped at the privacy scan, as above. Nothing was uploaded.
- Names leaked twice into the fact sheet this article was written from, forum handles in one file and names from model outputs in another. Removed, and every file now gets the name scan before it's used.
What I think worked: the release gates, which caught 12, 22 and 28 before anything went public, even where I then chose to ship under a waiver. Testing the starting model before training, which showed the name habit was T9's, not Qwen's. Training from the base instead of patching. Held-out test items for the booster. The private, then probe, then public publishing order.
Was renting GPUs worth it? For this one model, no
I asked for a blunt answer to that, and here it is.
The whole effort, T10 and T11, cost $1,051.87 before tax, $1,269.50 with AWS's VAT. GPU boxes were nearly all of it. The plan's first estimate had said $100 to $160.
What did it buy, T11 against T9? Atlassian knowledge in the vicinity of T9: a little higher on multiple choice, not significantly, and worse on Forge. Tool calling better, on BFCL (91.8 against 85.7) and on declining tools it shouldn't call (85.8 against 76.7). And trust: no invented people, 0 hits on the 309 privacy probes where T9 produces 57 invented keys and URLs.
T9 trained on our own Mac Studio. 2,600 steps, 5.8M tokens, about 23 hours, for the price of the electricity.
So for this one model, judged on its gains alone, renting wasn't worth it. Two things soften it.
First, most of the money went on T10 and on building and debugging the rental pipeline, not on the model we published. T11 itself cost about $160, the late Forge fix included. The next round on a new base model is now estimated at $70 to $90 of rental, plus the release on the Mac. Whether the rest was worth it depends on whether there's a next round.
Second, much of what a user will notice didn't need rented GPUs at all. No invented people, no invented ticket keys, the voice: that came from cleaning the data and starting from the original Qwen model instead of T9. The identity fix, for a habit T11 picked up in its own training, came from 143 short identity rows in the booster. Those rows are short too. The Mac could have trained them in a few days.
What I haven't measured, and should before renting again, is which rows bought the tool-calling gain. T11's mix had 527 goose rows out of 50,095. The BFCL gain might come from short general rows the Mac could train. The same mix minus the long goose rows, trained at home, would answer that for the cost of a few days of a machine we already own.
The rule for next time is in our recipe now. Train everything that fits on the Mac first, measure it, and rent a box only for what the Mac can't do.
Opus or Sol for Forge; Sol Pro for a payments app
If you're choosing a model for general app work, GPT-6.1 Sol Pro is the one that cleared every bar on Gauntlet. The next four finished within 0.006 of each other under the same cap. They didn't build the same app. Sonnet earned the most before the cap, Sol wrote about 530 tokens a call to Sonnet's 6,100, and the payment Opus's approval flow created never showed up in its table. Same score, different apps.
If you're choosing one for Forge, Opus 5.5 and GPT-6.1 Sol are the two I'd trust, with Sol Pro close behind. Sonnet 5.5 built a good ledger and then lost almost half its score at a button, a modal and an LLM call. That's exactly where a Forge app meets a real person, which is why the scorer prices it the way it does.
And if you're fine-tuning your own model: measure the starting point first, keep every checkpoint off the box the moment it's saved, run every release gate before you call a model chosen, and assume the parent's manners come along for free.
One run per model. Bills are a lower bound. The boards move, and they'll have moved again by the time you read this, so open the Forge board and the Gauntlet board for the current numbers, and click any run for its every check.
Originally published on leanzero.net. More Atlassian, Forge and local-AI write-ups at leanzero.net/blog, and if you're planning a migration or a Forge app, that's what we do: leanzero.net/services.









Top comments (0)