Most evaluations of image generators are run wrong in the same way: someone uploads photos of an easy face, gets nice results, and concludes the tool is good.
Every product in this category clears the easy case. The differences live entirely in the hard one, so an evaluation built on easy inputs is measuring a tie and reporting a winner.
Here is a test design that separates them.
Pick adversarial inputs on purpose
Faces that reliably break these models:
- Glasses, especially strong prescriptions. Lenses refract what is behind them, and the model has to be consistent about a distortion it half-understands.
- Tightly curled or high-volume hair. Under-represented in most training sets and routinely smoothed into something straighter.
- Darker skin tones, where exposure and colour errors are both more likely and more visible.
- Asymmetry. Which is every real face, and the failure is subtle: outputs drift toward a symmetrised average that reads as "similar person, not this person".
- Distinctive marks. Scars, freckles, birthmarks. Watch whether they survive or get quietly averaged away.
If your own face is in one of those categories, use it. You will be a better judge of the likeness than any metric.
Separate the two questions
Is it a good photograph and is it still them are different, and tools trade one against the other. Heavy stylisation raises the first and lowers the second. Score them separately or the comparison collapses.
For likeness, a blind test beats an opinion. Mix 5 generated images with 5 real ones, hand them to someone who knows the subject well, and ask which are real. If they cannot tell, that is a stronger result than any similarity score, and it is cheaper to run.
Control the inputs
The most common self-inflicted error is uploading 15 photos from one afternoon: same shirt, same wall, same light. The model learns the shirt and the wall as features of you, and every output comes back oddly attached to both. Then people blame the model.
Vary days, clothing, lighting and angle. It feels like it should reduce consistency and it improves it.
Run it against more than one tool at once
Same inputs, same day, same prompts. Sequential testing weeks apart is worthless because your judgment drifts and so do the products.
We publish our own side-by-side runs, and the honest caveat is that we chose the source photos, which is exactly the selection problem I described at the top. Read them as a description of method rather than a verdict: versus Aragon, versus HeadshotPro and versus InstaHeadshots. The useful part is the failure cases, not the wins.
Then test the destination, not the frame
A portrait that looks great at full size can fail at 48 pixels in a LinkedIn feed, which is where it will actually be judged. Evaluate at the size of the destination. We split that out separately for Aragon and HeadshotPro because the ranking genuinely changes when you do.
The 15 minute version
One hard face, 15 varied source photos, two tools, blind comparison by someone who knows them. That will tell you more than a week of reading reviews, most of which are affiliate content anyway.
More from this series
- Evaluating an image tool for 200 people is a different test from evaluating it for 1
- What "enterprise" actually means when a small tool sells it to you
- Latency budgets for batch image generation, and why "20 minutes" is the wrong number to optimise
- La foto en el CV: obligatoria en España, sospechosa en Londres, y el ATS no tiene nada que ver
On BetterPic
Top comments (0)