This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I build Shipgrade, a tool that grades landing pages. It reads a page the way a script does: the headline, the buttons, the prices. It never looks at the page. That left me with two questions I had never measured. What does a grader miss when it only reads the page? And when a model does get a screenshot, does it notice what is actually broken, or does it give a confident score either way?
To find out, I built PageTwins: 10 original landing pages and 100 "evil twins". Each twin is identical to its original except for one planted defect.
| Defect | What changed in the twin |
|---|---|
| Missing button | every sign-up, buy or download button removed |
| Hidden price | paid prices replaced with "Contact us" |
| Contradiction | one claim changed so two parts of the page disagree |
| Vague headline | the headline no longer says what the product is |
| Typos | three misspelled words in prominent text |
| Placeholder text | "Lorem ipsum" or a {{template_variable}} left in the copy |
| Blocking popup | a signup modal covers the first screen |
| Low contrast | headline faded to about 1.25:1 against its background |
| Broken image | the hero image fails to load |
| Cut-off text | the headline box is clipped partway through the first line |
All brands, copy and prices are fictional, and the art is SVG drawn by code. Every twin was rendered in headless Chromium and measured, and kept only if the measurement for its own defect changed (for example, the number of sign-up buttons dropped to zero) while every other measurement matched the original. The last three defects are visual by construction: the twin's text is byte-for-byte identical to the original's, so a grader that only reads the text cannot see them.
Each original also has two controls that a careful grader should leave alone: an identical copy, and a copy with a harmless change (a brand color swap at the same contrast, reordered feature cards, or a reworded subheadline).
The benchmark has four tasks, written with the kaggle-benchmarks SDK and pushed from the Kaggle CLI:
| Task | The model sees | Score |
|---|---|---|
| Pairs, screenshots | two full-page screenshots, shown in both orders | half for defects picked and named in both orders, half for controls answered "same" in both orders |
| Pairs, text | the same pairs as visible text with element roles | same |
| Single page, screenshot | one screenshot, no reference | share of twins where the model flags the defect and does not flag it on the original |
| Single page, text | one page as text | same |
Pairs is how a careful reviewer works: compare the new version with the last one. The single-page task is how most graders run in practice, mine included: one page, no reference, "what is wrong with this?". Showing every pair in both orders means a pick only counts if it survives the swap, and it measures position bias for free. In the single-page task a model gets credit only if it flags the defect on the twin and does not flag the same problem on the original, so a model that flags everything scores zero.
Every model gets the same prompt; the screenshot and text versions differ by one sentence that describes the input. Replies are short JSON, and anything unreadable counts as wrong.
Models Tested
Ten models, 34 scored runs, every one on Kaggle's model proxy:
- All four tasks (they accept images): Claude Sonnet 5.5 and Claude Haiku 4.5, GPT-6.1 Sol and GPT-6 Luna, Gemini 3.7 Flash and Gemini 3.5 Flash-Lite, and Gemma 4 31B.
- The two text tasks (text-only models): GLM-5, DeepSeek-R1 and Qwen 3 235B.
The idea was a stronger and a cheaper model from each of the three big vision providers, so I could see whether "smaller" means "slightly worse" or "a different kind of grader", plus open-weight models, because a grading tool like mine could run one locally. Kaggle gives $10 of model credit per day, so the lineup leaves out the most expensive tiers. The 26 runs I could price came to about $21, estimated from each run's token counts and per-token prices (I had no price for Gemma, GLM-5 and Qwen).
Two planned models are missing, and I would rather say so than quietly drop them:
-
Grok 4.6 is listed by
kaggle b t models, but Kaggle's proxy answered404 model not foundfor every xAI model I tried (Grok 4.6, Grok 4.5 and both Grok 4.20 models) from October 7 to October 11, and my runner kept probing until the end. Grok has no results. - gpt-oss-120b hit Kaggle's notebook time limit on two attempts and was dropped.
Every other planned run finished and is on the leaderboard.
Findings
1. The strong models do not need the "before" picture
I expected the reference page to be the big advantage. With screenshots, it barely mattered for the top four: Claude Sonnet 5.5 caught 98% of defects judging a page alone and 100% with the original next to it; GPT-6.1 Sol 98% and 96%; Gemini 3.7 Flash 96% and 98%; Gemma 4 31B 93% and 95%. Sonnet 5.5 got a perfect pairs score: all 100 defects picked and named in both orders, and all 20 controls called "same".
The two small models are where the pair changed things, and not always for the better. Gemini 3.5 Flash-Lite went from 63% to 72% caught with a reference, but it failed to call 9 of its 10 harmless pairs "same" in both orders. Claude Haiku 4.5 also passed only 1 of its 10 harmless pairs, and it showed real position bias: it picked the second screenshot 62% of the time and gave the same letter in both orders on 24% of defect pairs, which is exactly what the swap is there to catch. Identical copies were easy for everyone (10 of 10 for every model in every pairs run); a harmless difference is the actual test.
2. A screenshot is not a superset of the text
This is the result that changed how I think about my own tool. Restricted to the seven defects that are visible in both views, and pooled over the seven models that ran both, the screenshot and the text view tie: 82.9% vs 82.2% caught on the single-page task. But they tie by failing on different things.
- Typos: judging one page, the same seven models caught 69 of 70 typo twins from text and 51 of 70 from screenshots. GPT-6 Luna caught 3 of 10 in both screenshot tasks and 9 of 10 from the text view alone. Reading misspellings off pixels is still unreliable.
- Blocking popup: 70 of 70 from screenshots, 40 of 70 from text. The popup is in the text view (as a dialog at the end of the page), so this is not a visibility problem. Gemini 3.7 Flash caught 0 of 10 popups from text alone and 10 of 10 from screenshots. Claude Haiku 4.5 often saw the dialog and still did not flag it; one of its replies reads: "the signup dialog that appears on load could be considered slightly intrusive but is dismissible."
- The three visual-only defects are 0% from text, as designed, and 70% (low contrast), 80% (broken image) and 90% (cut-off text) from screenshots.
3. Images help strong models and hurt weak ones
On those same seven text-visible defects, the strongest models did better with a screenshot than with the text: Sonnet 5.5 97% vs 90%, GPT-6.1 Sol 97% vs 87%, Gemini 3.7 Flash 94% vs 80%. The two smallest did worse: Gemini 3.5 Flash-Lite 64% vs 76%, Claude Haiku 4.5 46% vs 64%. For a small model, an image is not extra information; it is a harder input to read the same information from.
4. Judging alone is a different skill, and the scores barely move
Without a reference, a vague headline is a matter of standards, and many models had none: Gemini 3.5 Flash-Lite, Claude Haiku 4.5 and Qwen 3 235B caught 0 of 10 vague headlines from text on their own, yet Qwen caught 10 of 10 when it could compare against the original.
The 1 to 10 "launch readiness" score each model also gives answers my second question. Strong models drop it hard on a broken page: Sonnet 5.5 gave originals 9.0 and twins 5.6 on average from screenshots, GPT-6.1 Sol 10.0 and 7.5. Weak ones barely move: Claude Haiku 4.5 gave 8.9 to originals and 8.6 to twins. A score from a weak grader looks just as confident on a broken page.
False alarms were rare for most models on the 20 clean pages per run (originals plus harmless variants), with two exceptions on the text view: DeepSeek-R1 flagged 6 of 20 clean pages and GLM-5 4 of 20, most often with an invented "contradiction" (7 of their 10 false labels). They are also the two models that wrote the longest replies (about 1,500 and 2,200 output tokens per page). Two models are not enough to call that a pattern, but it is the next thing I would test.
5. Price does not predict which cheap model is good
On pairs with screenshots, GPT-6 Luna scored 0.965 for about $0.21 and Gemini 3.7 Flash 0.99 for about $0.71, against 0.98 for about $4.12 (GPT-6.1 Sol) and 1.00 for about $4.44 (Claude Sonnet 5.5). Yet Gemini 3.5 Flash-Lite cost the same $0.21 as GPT-6 Luna and scored 0.635, and Claude Haiku 4.5 cost $0.92 and scored 0.455. "The cheaper model from the same family" ranged from nearly perfect to unreliable as a grader, so for a tool that runs on every deploy, a cheap model has to be tested on the actual task rather than picked by tier.
What this means for Shipgrade
A text-only grader is blind to three of these ten defects by construction, and that is the obvious part. The less obvious part is that switching to screenshots is not a free upgrade: it gains the visual defects and the popups, and gives back typos. The cheapest robust setup this data suggests is both views (a cheap vision model that tested well, such as GPT-6 Luna or Gemini 3.7 Flash here, on the screenshot, plus a text model on the text), and a pair against the last deploy when one exists.
Limitations
- The pages are synthetic, with exactly one defect each and a clear control. Real pages have several overlapping problems and no clean original.
- 10 base designs means 10 samples per defect per model, so one miss moves a category by 10 points. Read the per-category numbers as coarse.
- One run per model and one prompt wording. I did not measure run-to-run variance or prompt sensitivity.
- The pairs-text task is saturated: three of ten defects are invisible in text, so 70% caught is the ceiling, and four models (Sonnet 5.5, GPT-6.1 Sol, Gemma 4 31B, GLM-5) hit it exactly.
- Kaggle's overall leaderboard score averages the tasks each model ran, so the three text-only models are averaged over the two text tasks. Compare them on the text columns.
- Grok could not be run at all, and gpt-oss-120b was dropped after two timeouts.
- A small local pilot on Groq models was used only to debug the task code; none of the numbers here come from it.
What I would measure next
- Subtler variants of the same defects (one typo instead of three, contrast at 2.5:1 instead of 1.25:1, a half-hidden button) to find where each model's threshold is.
- Real pages with several defects at once, graded against human reviewers.
- Repeated runs per model, to put error bars on these numbers.
- Whether asking for less reasoning reduces the invented contradictions.
- Mobile widths, where cut-off text and popups behave differently.
My Benchmark
- Benchmark on Kaggle: kaggle.com/benchmarks/prasenx/pagetwins
- Tasks: pairs, screenshots, pairs, text, single page, screenshot, single page, text
- Dataset: kaggle.com/datasets/prasenx/pagetwins (120 pages with HTML, full-page screenshots and text views, CC BY 4.0)
- Code, raw answers and charts: github.com/StarKnightt/pagetwins (MIT)
StarKnightt
/
pagetwins
PageTwins: a Kaggle benchmark of 100 one-defect landing page twins. Do AI graders notice what broke?
PageTwins
A Kaggle benchmark that asks a simple question: when a landing page breaks in one specific way, does an AI grader notice, name the problem, and leave a clean page alone?
- Benchmark: https://www.kaggle.com/benchmarks/prasenx/pagetwins
- Dataset: https://www.kaggle.com/datasets/prasenx/pagetwins
The dataset
10 original landing pages (SaaS, developer tool, e-commerce tool, consumer products, mobile apps, an online course, a local service) on four layouts Each one has 10 "evil twins" that are identical except for exactly one planted defect, plus two controls: an identical copy and a copy with a harmless change. That makes 120 pages, 100 defect pairs and 20 control pairs. Every brand, product, quote and price is fictional and was written for the dataset.
| Defect | What changed in the twin | In the text view |
|---|---|---|
missing_cta |
every sign-up, buy or download button removed | visible |
hidden_price |
paid prices replaced with "Contact us" or similar | visible |
contradiction |
one claim changed so two parts of |








Top comments (2)
The harmless-pair controls are where I would look hardest, because the "same" answer may be the wrong one to reward. A brand color swap and a reworded subheadline are real differences, so a pair grader that says "colors changed, no defect" is describing the pages correctly. Scoring those 10 pairs as pass only when the model says "same" in both orders mixes two things: noticing a difference and calling it a defect. With 10 pairs the Flash-Lite 1/10 and Haiku 1/10 have a 95% Wilson interval of roughly 2-40%, so the data is also thin. Splitting the control reply into "noticed a change" and "named it as the defect" would show which models are miscalibrated and which are just literal.
The typo gap, 69/70 from text against 51/70 from screenshots, pools seven models over the same 10 base pages, so it is 70 correlated trials rather than 70 independent ones. Wilson for 51/70 is about 62-82% and for 69/70 about 92-99.7%, so the direction holds, but I would resample by page. Typo size relative to the render width probably drives it: a typo in a 14px footer line and one in the 48px headline are different tests. Do the 19 misses cluster on a few pages?
On the 82.9% vs 82.2% tie: 490 trials on each side, a 0.7-point difference is nothing, but pooling hides that each model has opposite per-defect profiles, which you show nicely. The practical consequence for the two-view setup is that the union of the two views' flags is what matters, so reporting how many twins are caught by at least one of screenshot or text for the same model would give the actual ceiling of the combined design.
Rendering full-page screenshots introduces another variable for the typo comparison: longer pages can be downscaled more aggressively when the model receives the image. The same 14px word may occupy very different effective pixels across your ten base designs.
I would record the supplied image dimensions and rendered page height, then repeat a few missed typo cases with viewport crops at unchanged scale. That separates failure to read the characters from failure to classify the typo. It also gives your proposed mobile-width extension a useful control, since both layout and effective text resolution can change at once.