We ran the same eight photographs of one man through the fifteen tools competing for "best AI headshot generator" and scored 57 selected outputs. The interesting result is not the ranking. It is that the failures all point the same way.
Disclosure. No affiliate links; every tool link goes to that tool's own homepage and nothing here earns a commission. The input photographs are of Ricardo Ghekiere, who founded BetterPic, one of the fifteen tools, and I work with BetterPic. It scored highest. The scoring and every image are published so the claim is checkable.
The subject
Fifty-something. High-volume grey hair. Full cheeks. A smile that is slightly uneven on one side.
Eleven of the fifteen tools returned a man who was younger, thinner-faced, or darker-haired. Often all three.
Why "random error" does not explain it
If these were noise, the errors would be symmetric. Some outputs would age him, widen the jaw, add grey. Across 57 images, essentially none did.
Every likeness failure we logged moved in one of these directions:
- Grey removed or reduced. Headshot.kiwi in all 3 selected frames, HeadshotsByAI in all 4, Dreamwave, ProPhotos and Aragon in subsets.
- Face narrowed. HeadshotPro in all 6 of its frames, at a consistent magnitude.
- Hair volume inflated. The inverse, and rarer: PortraitPal returned 2 oversized hairstyles, Aragon 2 more.
- Subject replaced. TryItOn AI returned one selected output that was a different person.
A one-directional error across eleven independent products, each with its own pipeline, is not eleven coincident bugs. It is a shared prior.
The mechanism, stated as a hypothesis
These products are personalisation layers — LoRA or DreamBooth-style fine-tunes of a few hundred to a few thousand steps — sitting on a general text-to-image base. You supply 8 to 15 photos; the fine-tune learns an identity token; generation is then conditioned on that token plus a prompt like professional corporate headshot, studio lighting.
That prompt is doing more work than the identity token. Whatever "professional headshot" means inside the base model was learned from stock photography, LinkedIn-adjacent corporate imagery and marketing shots — a distribution that skews young, symmetric and conventionally groomed. Where the subject agrees with that prior, the fine-tune has nothing to fight. Where the subject disagrees — grey hair, a fuller face — the prior competes with the identity token, and on a light fine-tune the prior wins.
That predicts exactly what we measured: the error is not "the model got the face wrong", it is "the model pulled the face toward the mean of its training distribution".
I am labelling this a hypothesis rather than a finding because our test measures outputs, not weights. It is falsifiable, though, and cheaply: undertrain and overtrain the same LoRA on the same subject and watch whether likeness degrades toward the prior specifically, rather than degrading in arbitrary directions.
Headshot.kiwi, where the prior wins outright.
HeadshotPro, the same pull at lower amplitude and applied uniformly.
The corollary nobody markets
Realism and likeness come apart, and every comparison that reports one number hides it.
| Tool | Realism | Likeness |
|---|---|---|
| Headshot.kiwi | 7 | 2.5 |
| HeadshotsByAI | 6 | 2 |
| The Multiverse AI | 7 | 4.5 |
| BetterPic | 8.5 | 9 |
Headshot.kiwi produces genuinely convincing photographs — good directional light, real skin texture, plausible depth-of-field falloff — of somebody else. Scored as "quality", it looks mid-table. Scored on whether it returned the person you uploaded, it is near the bottom.
This is why our rubric weights likeness at 40% and realism at 30%. A model that has learned to render faces beautifully and to render this face is doing two separable things, and only one of them is what the customer bought.
The falsifiable prediction
If the mechanism above is right, the failure rate should scale with distance from the base model's prior. Grey hair is one axis. The others, which this test did not cover, are the ones already reported anecdotally across the category: tightly curled hair, darker skin tones, and strong prescription glasses.
That is the obvious next test, and it is more useful than another ranking. One subject tells you about one subject — this whole exercise is a single face, one skin tone, one hair type, no glasses, and I would not generalise it further than the mechanism.
What came out on top, and where it slipped
BetterPic scored 9 on likeness, the highest in the test, and 8.6 overall. The frame we scored highest keeps the grey, the volume, the skin texture and the uneven smile at once.
It is worth naming where it slipped, because it slipped in the same direction as everything else: its grey-suit frame — the formal one, the one you would actually use — is the weakest of its four. The hair is combed flat and the grey is cut back to a streak at the temple. The prior does not disappear at the top of the table. It just gets outvoted more often.
Every score, all 57 outputs and the 8 inputs: the full gallery.




Top comments (0)