DEV Community

Cover image for Do all AI-generated apps look the same? I measured it
Mathieu Poli
Mathieu Poli

Posted on

Do all AI-generated apps look the same? I measured it

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge.

I wrote this summer that every AI-generated app looks the same. It was an opinion. The Kaggle Benchmarking Challenge gave me a reason to measure it.

The phenomenon itself isn't new. Anthropic described it in November 2025 in a post on frontend design and gave it a name, "distributional convergence": left without direction, a model falls back on the most frequent choices in what it learned, Inter and purple gradients first. I wanted to know three things I hadn't seen measured anywhere:

  1. Do models converge category by category, with one cliché per type of app?
  2. Do models from different makers land on the same choices?
  3. Does asking them to avoid clichés get them out?

Short answer: yes, yes, and no. I measured it twice, in a pilot on my machine and then in a public benchmark on Kaggle, and the two agree almost number for number.

What I Benchmarked

Twelve briefs with nothing in common

Each brief describes an app in one sentence: a boxing gym in a working-class neighbourhood of Marseille, a corporate law firm in Paris, a rock and metal festival, a daycare, a winemaker in Patrimonio, a funeral home, a skate shop, a music conservatory, an emergency department, a hands-on science museum for kids, a budgeting app for students, a fishmonger.

Three experiments

  1. Design choices. The model gets the brief and returns its decisions as JSON: primary, secondary, background and text colours, heading and body fonts, corner radius, home layout out of six, icon style, three mood words.
  2. The same request with one more sentence: "Avoid the visual clichés usually associated with this kind of business: make choices a good designer would find original, while still fitting the brief."
  3. The home screen. For six of the briefs, the model codes the screen as a single 390 × 844 HTML file, Google Fonts allowed, no external images, and no style direction at all.

What I count

  • Colour difference, ΔE, a perceptual distance: under 5, two colours are hard to tell apart side by side; above 50, they have nothing in common.
  • The expected answer, meaning what most models picked in the pilot. I fixed it before running Kaggle:
Where Expected answer
Design choices, every category Inter, Roboto, Lato, Playfair Display, Bebas Neue, Montserrat, Poppins, Nunito or Open Sans; "hero image then cards" home
Boxing gym and festival screens condensed caps headings, dark background
Law firm, winemaker and fishmonger screens serif headings (and a light background for the fishmonger)
Daycare screens rounded type, light background
Every screen a fake status bar showing 9:41, small letter-spaced caps labels
  • The originality score: the share of choices that are not the expected answer. At 0, the model ticks every box; at 1, none.
  • The house-style index, which measures the opposite: how much a model's six screens look like each other (same main font, same background). It sits around 0.5 when the look follows the brief and reaches 1 when it is the same for every app.

These expected answers, like the "new repertoire" of question 3, come from the pilot; Kaggle tests them on other runs, other models and a shuffled order of options.

Screens are analysed from their code, without rendering them: fonts loaded, background colour, presence of 9:41, labels. I checked this analysis against the pilot's 30 screens, measured in a browser at 390 pixels: it gets the fonts and the status bar right in all 30, the background in 29. The rest, block order or content, comes from reading the screenshots.

Conditions

Same prompt for everyone, default settings for every model. In the pilot, each brief was asked three times per model to separate the cliché from chance. On Kaggle, the order of the home layouts and icon styles is shuffled for each brief, and each screen is generated twice.

Models Tested

Three makers, because the second question is about whether they agree, and several model sizes at Anthropic and OpenAI.

Model Maker Pilot Kaggle
Claude Haiku 4.5 Anthropic choices, screens choices, screens
Claude Sonnet 4.6 Anthropic choices, screens unavailable
Claude Opus 4.6 Anthropic choices, screens choices, screens
GPT-6.1 sol OpenAI choices, screens not offered
GPT-6 luna OpenAI choices not offered
GPT-6 Astra OpenAI choices, screens
Gemini 3.5 Flash Google choices (6 briefs) choices, screens
Gemini 3.6 Flash Google screens choices, screens
Gemini 3.7 Flash Google choices, screens
Gemini 3.8 Flash Google choices, screens

On Kaggle, Sonnet 4.6 never ran: the service answered "model under heavy load" on every attempt from 2 to 5 October. DeepSeek R1 and Qwen 3, which I wanted to add, got the same answer, and Grok, although listed, could not be found. The two GPT models from the pilot aren't offered there; GPT-6 Astra replaces them. Gemini 3.7 Flash is Kaggle's default model, which ran it on its own when the tasks were created. In the pilot, the Google API's free quota limited Gemini.

In total: 246 designs and 30 screens in the pilot, 168 designs and 81 screens on Kaggle.

Findings

At a glance

Measure Pilot Kaggle
Colour difference on the same brief, across makers 38 41
Colour difference for the same model, from one brief to another 71 76
"Hero image then cards" home 65% 69% (shuffled order)
"No cliché" headings taken from the new repertoire 65% 65%
Screens: expected type and background for the category 92% 87%
Screens: fake status bar showing 9:41 57% 62%
Law firm filled with real banks 4 models out of 5, all but GPT 6 out of 7, all but GPT

The pilot is the reference: it set the expected answers, so its numbers are partly mechanical. Kaggle is the test, and it finds almost the same ones. Colour differences compare Claude and GPT pairs in the pilot, and pairs across the three makers on Kaggle.

1. Colour: the brief decides, not the maker (questions 1 and 2)

I compare the primary colour of every design.

On the same brief, two models from different makers pick colours an average of 38 apart in the pilot, 41 on Kaggle. The same model, from one brief to another: 71 and 76. Two competitors are closer to each other than a model is to itself as soon as the business changes.


The primary colour picked by seven models from three makers for six of the briefs (Kaggle). On the right, the average difference between the seven: under 10, the eye can barely tell them apart.

For the funeral home and the law firm, the seven models are only 9 apart on average. Sometimes they agree on the exact code: #2C3E50 for the funeral home from Opus and Gemini in the pilot, from Haiku and Opus on Kaggle. It's the Midnight Blue of the Flat UI palette. In the pilot, #E63946, which opens a widely reused Coolors palette, comes out for the boxing gym from Haiku and from Gemini. The only brief that scatters is the one without an obvious cliché: the kids' museum, at 73. A university study from July 2026, Design Theater, found that interface generation tools converge on look and layout but vary more on colour. These numbers say why: colour changes from one category to another, hardly from one model to another.

These category colours have a cost: with white text on top, 34 primary colours out of 84 fail the WCAG minimum contrast (4.5:1) on Kaggle, oranges most of all.

2. Typography: the textbook rule, applied without exception (question 1)

I compare each screen's fonts with the expected font for its category.

In the pilot, all ten boxing and festival screens use condensed caps headings (Bebas Neue, Barlow Condensed, Oswald), all ten law firm and winemaker screens use serif headings, and four daycares out of five use rounded type: the first rule you learn in typography, applied to the whole category. On Kaggle: condensed caps for the boxing gym and the festival in 25 screens out of 27, serifs for the law firm and the winemaker in 23 out of 27, rounded type for the daycare in 10 out of 14.


The home screen heading coded by five models for three of the apps (pilot). Condensed caps for the festival, serifs for the law firm, rounded type for the daycare. Only GPT-6.1 gives the daycare a serif.

3. Screen structure: one template, six outfits (question 1)

I look at the home layout chosen in JSON, then at the order of blocks on the coded screens.

"A hero image then cards" wins in 65% of the pilot's designs, and in 69% on Kaggle, where the order of options is shuffled. The layout follows the category too: a dashboard for the law firm in all 16 of its pilot designs, although that option only came fifth in the list, a grid of tiles for the kids' museum, a hero image for the conservatory.


The home screens coded by five models for six apps (pilot): one row per app, one column per model. Each row looks more alike than each column, except for GPT-6.1. First names invented by the models are blurred.

On the screens, almost all of them stack the same blocks in the same order: the brand name with a bell or an avatar, a big block that sets the tone, a row of shortcuts or numbers, a "title + See all" section, cards, the tab bar. The shortcut row appears in 13 screens out of 30, never from GPT, and in five of them it leads somewhere the tab bar already offers just below.

4. Components: the Dribbble kit (questions 1 and 2)

I count the components that come back, in the code and on the screenshots.

Lists of AI design "tells", like avoid-ai-design or the one from Developers Digest, mention caps labels and rows of stats. Here, they can be counted. A small letter-spaced caps label above a heading appears in 24 screens out of 30 in the pilot and 63 out of 81 on Kaggle. The "big number with a caps caption" tile in ten pilot screens. And a fake iOS status bar showing 9:41 in 17 screens out of 30 in the pilot, 50 out of 81 on Kaggle.


Top: the fake status bar showing 9:41, drawn by five models. Bottom: the same brand name and small caps pairing from four models (pilot).

9:41 is the time Apple shows on almost every iPhone visual, a legacy of its keynotes. In a real app, the system draws that bar, and the prompt didn't ask for one. What the models reproduce is an app presentation, the kind you post on Dribbble.

5. Content and features: what nobody asked for (question 2)

I read the text on the screens.

The law firm brief didn't say who uses the app. Almost every model turned it into an internal tool for the lawyers, and they fill its confidential deals with real banks: four models out of five in the pilot, six out of seven on Kaggle, BNP Paribas first. Sonnet and Opus even invent the same deal, "BNP Paribas / Crédit du Nord". The only ones that use none are the GPT models, in the pilot and on Kaggle; GPT-6.1 actually made an app for the firm's clients.


Top: real banks in made-up deals. Bottom: the training streak four boxing gyms show without being asked (pilot).

No brief mentioned features. Four boxing gyms out of five still show a "streak", the run of training days from fitness apps, and two show a ranking. Gemini's festival has a cashless wristband balance and the beach weather. Even names converge: the daycare is called Sunshine Sprouts, Little Sprouts or Sproutlings by four models out of five, and the festival has Gojira headlining for three. By drawing the screen, the model also sets the scope of the product, and each of these blocks assumes a server and data that someone keeps up to date.

6. "Avoid clichés" moves the cliché (question 3)

I run the request again with the extra sentence, and compare heading fonts.

Colours move noticeably and makers agree less: in the pilot, the difference between Claude and GPT on the same brief goes from 36 to 60. But the models immediately converge on a new repertoire. Two thirds of the "no cliché" headings come from it, 39 out of 60 in the pilot and 55 out of 84 on Kaggle: Fraunces, Space Grotesk, Syne, Bricolage Grotesque.


The most frequent heading fonts, with the plain prompt and then with "avoid clichés" (Kaggle, 84 designs each).

Anthropic had noted it in its post: even with explicit instructions to avoid certain patterns, the model can default to other common choices, "like Space Grotesk". It isn't specific to Claude. The mood vocabulary changes register too ("warm", "grounded", "editorial"), and Sonnet gives the exact same terracotta, #C1440E, to the boxing gym and the festival.

Ask them to avoid clichés and they swap the corporate template for the trendy design studio one.

7. The models: originality and house style

I compare the models' originality scores, then their house-style index.


The originality score of each model on Kaggle, for design choices and for screens.

Only one model goes above 0.5: GPT-6 Astra, at 0.61 on design choices. It is also the most original on screens (0.45), as GPT-6.1 was in the pilot. But that originality comes with a catch: GPT-6.1 applies a house style to every app, a cream background in five screens out of six, DM Sans in all six, nearly square corners, a ↗ arrow on links, and its house-style index reaches 0.92. On Kaggle, GPT-6 Astra gets the exact same index, 0.92: a different model, the same signature. avoid-ai-design attributes the cream and terracotta pairing to Claude; here, cream is GPT's signature.

The others follow the category. Opus is the most original of the Anthropic and Google models on JSON choices (0.47), but when it codes the screen, it applies the category rule in 20 cases out of 20. Sonnet is the most decorated in the pilot, with fifteen labels and ten gradients per screen. Gemini builds live dashboards, where everything seems to be happening right now.

Escaping the category cliché and designing each app for itself are two different things.

What I corrected along the way

  • I worried that "hero image then cards" won because of its place in the list. Kaggle shuffles the order, and it wins even more often.
  • My first screenshots were wrong: headless Chrome forces a 500-pixel width, and centred screens came out shifted and cropped. I redid them at 390 pixels with another engine, and one sentence of my draft fell: GPT-6.1 doesn't use serifs everywhere.
  • I had counted eight models in the pilot. Gemini 3.8's answers had not been saved: it's seven.
  • One screen per app gives noisy scores (Haiku at 0.27, then 0.45 on a second try). On Kaggle, each screen is generated twice.

What I take away

So, do AI-generated apps all look the same? Yes and no. A boxing gym doesn't look like a daycare: colour, type and background change with the business. But every boxing gym looks alike, whichever model draws it. And under the outfit, the skeleton is the same everywhere: the same blocks in the same order, the same components, the same fake status bar. AI doesn't make one app. It makes one app per category, on a single template.

The three opening questions have their answer. Yes, models converge category by category. Yes, different makers land on the same choices, down to the colour code. And no, asking them to avoid clichés doesn't get them out: it leads to the next one.

The model does exactly what it's asked: the most likely answer for this type of app. And the most likely answer for a boxing gym is what it has seen most often for a boxing gym.

A better prompt doesn't change that: "avoid clichés" moves the problem without solving it. What's missing is a decision made somewhere else, by someone who knows what this particular boxing gym has that others don't, and rules that hold it over time. That's what I called taste in an earlier piece, and what a design system does when it's written for a brand and not for a category.

The screens also show that a palette isn't enough: structure, components and the feature list converge as much as colours. Before asking a model for a screen, I would give it more than a style guide: the components it may use, what the screen must let people do, and what it must not invent.

What I'd measure next

First, the effect of a real design system in the prompt: colours, fonts, components and planned features. Does the model follow it, or does the category take over wherever the system says nothing?

Then, language. All briefs were in English. A Marseille boxing gym described in French might have another cliché, or the same one.

Finally, visual similarity. The benchmark measures the screens' code; screenshots should also be compared with each other, over hundreds of generations.

Limitations

JSON choices were partly closed: six home layouts, four icon styles. A single prompt, in English. In the pilot, Claude ran through the Claude Code command line, GPT and Gemini through their APIs; on Kaggle, all go through the same service. Several models couldn't run on Kaggle, and screens only cover six of the twelve apps. Finally, screen counts mix automatic measurements of the code with reading the screenshots.

My Benchmark

The benchmark on Kaggle: Twelve apps, one cliché each, public, with its leaderboard. It holds two tasks: twelve-apps-design-choices, for design choices with and without "avoid clichés", and twelve-apps-home-screens, for home screens. Each gives every model an originality score; the second also shows the house-style index.

Written with help from Claude, which is also one of the models tested.


Mathieu Poli, Head of Frontend Engineering at GoodBarber. I teach and write about frontend engineering, product design and AI, and everything that happens when the three meet.
X: hellomathieup · LinkedIn: hellomathieup

Top comments (0)