DEV Community

Cover image for PageTwins: I Planted One Defect on 100 Landing Pages. Which AI Graders Noticed?
StarKnight
StarKnight

Posted on

PageTwins: I Planted One Defect on 100 Landing Pages. Which AI Graders Noticed?

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

I build Shipgrade, a tool that grades landing pages. It reads a page the way a script does: the headline, the buttons, the prices. It never looks at the page. That left me with two questions I had never measured. What does a grader miss when it only reads the page? And when a model does get a screenshot, does it notice what is actually broken, or does it give a confident score either way?

To find out, I built PageTwins: 10 original landing pages and 100 "evil twins". Each twin is identical to its original except for one planted defect.

Six of the ten defects, each next to the original page

Defect What changed in the twin
Missing button every sign-up, buy or download button removed
Hidden price paid prices replaced with "Contact us"
Contradiction one claim changed so two parts of the page disagree
Vague headline the headline no longer says what the product is
Typos three misspelled words in prominent text
Placeholder text "Lorem ipsum" or a {{template_variable}} left in the copy
Blocking popup a signup modal covers the first screen
Low contrast headline faded to about 1.25:1 against its background
Broken image the hero image fails to load
Cut-off text the headline box is clipped partway through the first line

All brands, copy and prices are fictional, and the art is SVG drawn by code. Every twin was rendered in headless Chromium and measured, and kept only if the measurement for its own defect changed (for example, the number of sign-up buttons dropped to zero) while every other measurement matched the original. The last three defects are visual by construction: the twin's text is byte-for-byte identical to the original's, so a grader that only reads the text cannot see them.

Each original also has two controls that a careful grader should leave alone: an identical copy, and a copy with a harmless change (a brand color swap at the same contrast, reordered feature cards, or a reworded subheadline).

The benchmark has four tasks, written with the kaggle-benchmarks SDK and pushed from the Kaggle CLI:

Task The model sees Score
Pairs, screenshots two full-page screenshots, shown in both orders half for defects picked and named in both orders, half for controls answered "same" in both orders
Pairs, text the same pairs as visible text with element roles same
Single page, screenshot one screenshot, no reference share of twins where the model flags the defect and does not flag it on the original
Single page, text one page as text same

Pairs is how a careful reviewer works: compare the new version with the last one. The single-page task is how most graders run in practice, mine included: one page, no reference, "what is wrong with this?". Showing every pair in both orders means a pick only counts if it survives the swap, and it measures position bias for free. In the single-page task a model gets credit only if it flags the defect on the twin and does not flag the same problem on the original, so a model that flags everything scores zero.

Every model gets the same prompt; the screenshot and text versions differ by one sentence that describes the input. Replies are short JSON, and anything unreadable counts as wrong.

Models Tested

Ten models, 34 scored runs, every one on Kaggle's model proxy:

  • All four tasks (they accept images): Claude Sonnet 5.5 and Claude Haiku 4.5, GPT-6.1 Sol and GPT-6 Luna, Gemini 3.7 Flash and Gemini 3.5 Flash-Lite, and Gemma 4 31B.
  • The two text tasks (text-only models): GLM-5, DeepSeek-R1 and Qwen 3 235B.

The idea was a stronger and a cheaper model from each of the three big vision providers, so I could see whether "smaller" means "slightly worse" or "a different kind of grader", plus open-weight models, because a grading tool like mine could run one locally. Kaggle gives $10 of model credit per day, so the lineup leaves out the most expensive tiers. The 26 runs I could price came to about $21, estimated from each run's token counts and per-token prices (I had no price for Gemma, GLM-5 and Qwen).

Two planned models are missing, and I would rather say so than quietly drop them:

  • Grok 4.6 is listed by kaggle b t models, but Kaggle's proxy answered 404 model not found for every xAI model I tried (Grok 4.6, Grok 4.5 and both Grok 4.20 models) from October 7 to October 11, and my runner kept probing until the end. Grok has no results.
  • gpt-oss-120b hit Kaggle's notebook time limit on two attempts and was dropped.

Every other planned run finished and is on the leaderboard.

Kaggle task scores for every model

Findings

1. The strong models do not need the "before" picture

I expected the reference page to be the big advantage. With screenshots, it barely mattered for the top four: Claude Sonnet 5.5 caught 98% of defects judging a page alone and 100% with the original next to it; GPT-6.1 Sol 98% and 96%; Gemini 3.7 Flash 96% and 98%; Gemma 4 31B 93% and 95%. Sonnet 5.5 got a perfect pairs score: all 100 defects picked and named in both orders, and all 20 controls called "same".

Comparing two versions vs judging one page, screenshots

The two small models are where the pair changed things, and not always for the better. Gemini 3.5 Flash-Lite went from 63% to 72% caught with a reference, but it failed to call 9 of its 10 harmless pairs "same" in both orders. Claude Haiku 4.5 also passed only 1 of its 10 harmless pairs, and it showed real position bias: it picked the second screenshot 62% of the time and gave the same letter in both orders on 24% of defect pairs, which is exactly what the swap is there to catch. Identical copies were easy for everyone (10 of 10 for every model in every pairs run); a harmless difference is the actual test.

2. A screenshot is not a superset of the text

This is the result that changed how I think about my own tool. Restricted to the seven defects that are visible in both views, and pooled over the seven models that ran both, the screenshot and the text view tie: 82.9% vs 82.2% caught on the single-page task. But they tie by failing on different things.

Twins caught per defect, all models pooled

  • Typos: judging one page, the same seven models caught 69 of 70 typo twins from text and 51 of 70 from screenshots. GPT-6 Luna caught 3 of 10 in both screenshot tasks and 9 of 10 from the text view alone. Reading misspellings off pixels is still unreliable.
  • Blocking popup: 70 of 70 from screenshots, 40 of 70 from text. The popup is in the text view (as a dialog at the end of the page), so this is not a visibility problem. Gemini 3.7 Flash caught 0 of 10 popups from text alone and 10 of 10 from screenshots. Claude Haiku 4.5 often saw the dialog and still did not flag it; one of its replies reads: "the signup dialog that appears on load could be considered slightly intrusive but is dismissible."
  • The three visual-only defects are 0% from text, as designed, and 70% (low contrast), 80% (broken image) and 90% (cut-off text) from screenshots.

3. Images help strong models and hurt weak ones

On those same seven text-visible defects, the strongest models did better with a screenshot than with the text: Sonnet 5.5 97% vs 90%, GPT-6.1 Sol 97% vs 87%, Gemini 3.7 Flash 94% vs 80%. The two smallest did worse: Gemini 3.5 Flash-Lite 64% vs 76%, Claude Haiku 4.5 46% vs 64%. For a small model, an image is not extra information; it is a harder input to read the same information from.

Screenshot vs text view on the defects visible in both

4. Judging alone is a different skill, and the scores barely move

Without a reference, a vague headline is a matter of standards, and many models had none: Gemini 3.5 Flash-Lite, Claude Haiku 4.5 and Qwen 3 235B caught 0 of 10 vague headlines from text on their own, yet Qwen caught 10 of 10 when it could compare against the original.

The 1 to 10 "launch readiness" score each model also gives answers my second question. Strong models drop it hard on a broken page: Sonnet 5.5 gave originals 9.0 and twins 5.6 on average from screenshots, GPT-6.1 Sol 10.0 and 7.5. Weak ones barely move: Claude Haiku 4.5 gave 8.9 to originals and 8.6 to twins. A score from a weak grader looks just as confident on a broken page.

False alarms were rare for most models on the 20 clean pages per run (originals plus harmless variants), with two exceptions on the text view: DeepSeek-R1 flagged 6 of 20 clean pages and GLM-5 4 of 20, most often with an invented "contradiction" (7 of their 10 false labels). They are also the two models that wrote the longest replies (about 1,500 and 2,200 output tokens per page). Two models are not enough to call that a pattern, but it is the next thing I would test.

False alarms on pages with nothing wrong

5. Price does not predict which cheap model is good

On pairs with screenshots, GPT-6 Luna scored 0.965 for about $0.21 and Gemini 3.7 Flash 0.99 for about $0.71, against 0.98 for about $4.12 (GPT-6.1 Sol) and 1.00 for about $4.44 (Claude Sonnet 5.5). Yet Gemini 3.5 Flash-Lite cost the same $0.21 as GPT-6 Luna and scored 0.635, and Claude Haiku 4.5 cost $0.92 and scored 0.455. "The cheaper model from the same family" ranged from nearly perfect to unreliable as a grader, so for a tool that runs on every deploy, a cheap model has to be tested on the actual task rather than picked by tier.

What this means for Shipgrade

A text-only grader is blind to three of these ten defects by construction, and that is the obvious part. The less obvious part is that switching to screenshots is not a free upgrade: it gains the visual defects and the popups, and gives back typos. The cheapest robust setup this data suggests is both views (a cheap vision model that tested well, such as GPT-6 Luna or Gemini 3.7 Flash here, on the screenshot, plus a text model on the text), and a pair against the last deploy when one exists.

Limitations

  • The pages are synthetic, with exactly one defect each and a clear control. Real pages have several overlapping problems and no clean original.
  • 10 base designs means 10 samples per defect per model, so one miss moves a category by 10 points. Read the per-category numbers as coarse.
  • One run per model and one prompt wording. I did not measure run-to-run variance or prompt sensitivity.
  • The pairs-text task is saturated: three of ten defects are invisible in text, so 70% caught is the ceiling, and four models (Sonnet 5.5, GPT-6.1 Sol, Gemma 4 31B, GLM-5) hit it exactly.
  • Kaggle's overall leaderboard score averages the tasks each model ran, so the three text-only models are averaged over the two text tasks. Compare them on the text columns.
  • Grok could not be run at all, and gpt-oss-120b was dropped after two timeouts.
  • A small local pilot on Groq models was used only to debug the task code; none of the numbers here come from it.

What I would measure next

  • Subtler variants of the same defects (one typo instead of three, contrast at 2.5:1 instead of 1.25:1, a half-hidden button) to find where each model's threshold is.
  • Real pages with several defects at once, graded against human reviewers.
  • Repeated runs per model, to put error bars on these numbers.
  • Whether asking for less reasoning reduces the invented contradictions.
  • Mobile widths, where cut-off text and popups behave differently.

The PageTwins leaderboard on Kaggle

My Benchmark

GitHub logo StarKnightt / pagetwins

PageTwins: a Kaggle benchmark of 100 one-defect landing page twins. Do AI graders notice what broke?

PageTwins

A Kaggle benchmark that asks a simple question: when a landing page breaks in one specific way, does an AI grader notice, name the problem, and leave a clean page alone?

Six of the ten defects next to the original page

The dataset

10 original landing pages (SaaS, developer tool, e-commerce tool, consumer products, mobile apps, an online course, a local service) on four layouts Each one has 10 "evil twins" that are identical except for exactly one planted defect, plus two controls: an identical copy and a copy with a harmless change. That makes 120 pages, 100 defect pairs and 20 control pairs. Every brand, product, quote and price is fictional and was written for the dataset.

Defect What changed in the twin In the text view
missing_cta every sign-up, buy or download button removed visible
hidden_price paid prices replaced with "Contact us" or similar visible
contradiction one claim changed so two parts of
…

Top comments (2)

Collapse
 
arhancanli profile image
Arhan Canli •

The harmless-pair controls are where I would look hardest, because the "same" answer may be the wrong one to reward. A brand color swap and a reworded subheadline are real differences, so a pair grader that says "colors changed, no defect" is describing the pages correctly. Scoring those 10 pairs as pass only when the model says "same" in both orders mixes two things: noticing a difference and calling it a defect. With 10 pairs the Flash-Lite 1/10 and Haiku 1/10 have a 95% Wilson interval of roughly 2-40%, so the data is also thin. Splitting the control reply into "noticed a change" and "named it as the defect" would show which models are miscalibrated and which are just literal.

The typo gap, 69/70 from text against 51/70 from screenshots, pools seven models over the same 10 base pages, so it is 70 correlated trials rather than 70 independent ones. Wilson for 51/70 is about 62-82% and for 69/70 about 92-99.7%, so the direction holds, but I would resample by page. Typo size relative to the render width probably drives it: a typo in a 14px footer line and one in the 48px headline are different tests. Do the 19 misses cluster on a few pages?

On the 82.9% vs 82.2% tie: 490 trials on each side, a 0.7-point difference is nothing, but pooling hides that each model has opposite per-defect profiles, which you show nicely. The practical consequence for the two-view setup is that the union of the two views' flags is what matters, so reporting how many twins are caught by at least one of screenshot or text for the same model would give the actual ceiling of the combined design.

Collapse
 
ahmetozel profile image
Ahmet Özel •

Rendering full-page screenshots introduces another variable for the typo comparison: longer pages can be downscaled more aggressively when the model receives the image. The same 14px word may occupy very different effective pixels across your ten base designs.

I would record the supplied image dimensions and rendered page height, then repeat a few missed typo cases with viewport crops at unchanged scale. That separates failure to read the characters from failure to classify the typo. It also gives your proposed mobile-width extension a useful control, since both layout and effective text resolution can change at once.