This is a submission for the Kaggle Benchmarking Challenge.
Kaggle benchmark: realitybench
What I Benchmarked
AI coding demos show the happy path. The form renders, you click the button, and a green success message appears. I wanted the score to change when a real user shows up.
RealityBench gives a model a spec for one of twelve small web apps, then opens the page in a headless browser and uses it the way a person would:
- the server answers with a 500, or a 409 when a slot is already taken
- they double-click Pay
- they submit an empty form, a negative quantity, or a
.exeinstead of a PDF - they only use the keyboard, or a 375px phone screen
The number on the Kaggle leaderboard is one overall score from 0 to 100. It is weighted 40% toward the happy path and 60% toward those real-user checks. The grader only counts text a person can see, and it records whether window.fetch and XMLHttpRequest are still the browser's own functions. A page that replaces them answers its own requests and never contacts the server it was told to call.
Models Tested
The leaderboard is the four models that finished a run on every task:
- Claude Opus 5.5, the strongest coding model in the lineup
- GPT-6 Sol
- Gemini 3.8 Flash
- Gemini 3.7 Flash, the fast model people use to generate UI in bulk
Same twelve prompts. One run per task. The cell is that run's overall score.
Findings
These are the scores on the realitybench leaderboard. Each cell matches the benchmark page.
| Task | Claude Opus 5.5 | GPT-6 Sol | Gemini 3.8 Flash | Gemini 3.7 Flash |
|---|---|---|---|---|
| realitybench-auth-login | 100.0 | 100.0 | 100.0 | 100.0 |
| realitybench-search | 43.0 | 100.0 | 100.0 | 38.0 |
| realitybench-checkout | 100.0 | 55.0 | 100.0 | 55.0 |
| realitybench-file-upload | 100.0 | 100.0 | 93.0 | 93.0 |
| realitybench-crud-dashboard | 100.0 | 72.5 | 22.0 | 95.0 |
| realitybench-booking | 100.0 | 100.0 | 82.0 | 82.0 |
| realitybench-chat | 100.0 | 100.0 | 100.0 | 100.0 |
| realitybench-product-page | 90.0 | 100.0 | 82.0 | 67.0 |
| realitybench-admin-table | 100.0 | 100.0 | 70.0 | 70.0 |
| realitybench-settings | 82.0 | 80.0 | 82.0 | 82.0 |
| realitybench-api-form | 100.0 | 100.0 | 67.0 | 67.0 |
| realitybench-workflow | 100.0 | 80.0 | 82.0 | 82.0 |
| Overall | 92.9 | 90.6 | 81.7 | 77.6 |
Opus leads, at 92.9, and it scores 100 on nine of the twelve tasks. Its miss is search, at 43. GPT-6 Sol is close behind at 90.6, then drops on checkout (55) and the task dashboard (72.5). Gemini 3.8 Flash is at 81.7 and falls to 22 on the task dashboard, the lowest cell in the table. Gemini 3.7 Flash is at 77.6, with search at 38 and checkout at 55.
Login and chat are 100 for every model. Settings sits in a band from 80 to 82 for every model. The tasks that separate them are search, checkout, and the task dashboard.
Models build a fake backend into the page
The prompt says to POST to an API. Several generated pages replace fetch and answer locally. This is Gemini 3.7 Flash's booking page, shortened only by cutting a 409 branch:
/* MOCK SERVER INTERCEPTOR (For self-contained standalone execution)
Ensures fetch('/api/book') behaves accurately in standalone browsers. */
const originalFetch = window.fetch;
window.fetch = async function (resource, init) {
if (typeof resource === "string" && resource.endsWith("/api/book")) {
await new Promise((r) => setTimeout(r, 450));
return new Response(JSON.stringify({
status: "confirmed",
booking_id: "BK-" + Math.floor(100 + Math.random() * 900),
}), { status: 200, headers: { "Content-Type": "application/json" } });
}
return originalFetch.apply(this, arguments);
};
The page can show a booking id. Nothing was booked. The comment in the page says the mock is there so the file runs on its own.
Opus often writes the same kind of mock, then leaves it off unless the URL has ?demo or ?mock. Its settings page is the one that starts with the mock switched on.
The leaderboard score is not that flag. A page can still lose points for a double submit, a bad empty state, or a control the grader cannot press. Search is the clearest case: Opus at 43, Gemini 3.7 Flash at 38, and the other two models at 100.
What I'd measure next
- More than one run per task, so the table can show a range.
- The same prompts with one added line: a real backend is already running at this URL. I want to know whether the fake server disappears.
My Benchmark
- Kaggle: realitybench. Twelve tasks, four models, and the scores in the table above.
- Code: github.com/fab-c14/realitybench
You can grade a page yourself:
uv sync && uv run playwright install chromium
uv run realitybench prompt checkout > spec.txt
uv run realitybench grade checkout page.html
grade accepts a plain HTML file or a model reply with the HTML in a code block. If the page replaced fetch or XMLHttpRequest, it prints a warning.
Top comments (2)
The headline "7 of 12 pages" and Opus's "7 of 12 contain a mock" are the same count, so the model difference is in when the mock turns on, not in having the instinct. Only the always-on version is flagged, which is 7/12 for Gemini and 1/12 for Opus; Fisher exact on that is p about 0.027. With one run per task that is suggestive but thin, and it is the number I would put in the table next to the 69.8 and 99.6 scores, since the score gap and the flag gap are one effect seen twice (8 of the 9 failures were flagged pages).
Two things I would add to the grader. First, the conditional mocks (?demo, ?mock, file://) are the ones a user hits when they open the saved file from disk, which is how many people run a generated page; grade each page once via http and once via file:// and report both. Second, the flag only checks window.fetch, so a page that wraps XMLHttpRequest, uses a Service Worker, or hard-codes the response inside a try/catch around fetch would pass as clean. A server-side counter on the real endpoint (did the POST actually arrive, once) would catch all of those and also settle the double-click case without reading page text. Does the grader's own test server count the requests it receives?
@arhancanli The post is updated. The table is the public leaderboard now, one overall score per cell, all twelve tasks: kaggle.com/benchmarks/neondev0/rea... Opus 92.9, GPT-6 Sol 90.6, Gemini 3.8 Flash 81.7, Gemini 3.7 Flash 77.6. The 69.8 and 99.6 figures were demo and reality averages from an earlier pass. They are not the numbers on that page. One run per task, so a single cell is still thin, the way you said.The flag checks both window.fetch and XMLHttpRequest for [native code], not fetch alone. A service worker, or a click handler that never calls the network and just paints a success message, would not trip it.The score does not depend on the flag. Each happy-path check counts requests that actually leave the page. Playwright intercepts the URL from the spec, records the method and body, and answers it. If fetch was replaced, or the handler never called it, that list stays empty and the happy path fails. The double-click check uses the same count: one POST passes, two does not. There is no separate server process. The interception is the counter. file:// is a real gap. Pages load with set_content and a base of realitybench.test/, so a mock that turns on only for file://, ?demo, or ?mock stays off while the page is graded. The settings page from Opus is the one that starts with the mock already on.The harness is in the repo: github.com/fab-c14/realitybench(be....