DEV Community

Miriam Alonso
Miriam Alonso

Posted on

Evaluating an image tool for 200 people is a different test from evaluating it for 1

A tool that produces one excellent portrait can produce 200 portraits that do not belong together. Those are separate properties and most evaluations only measure the first.

If you are choosing something for a whole company, here is what changes.

Consistency is the product

For one person, quality is the metric. For 200, the metric is variance.

A grid where every face is individually good and collectively mismatched looks worse than a grid of merely adequate photos that match. Human eyes are very good at spotting a set, and very quick to read inconsistency as carelessness.

So test for variance directly. Run 10 different people through the same settings, put the outputs in a grid at 96px, and look at it as a grid rather than as 10 images. Background tone, crop distance, colour temperature and eye-line height are the four that drift.

The failure modes are demographic

This is the uncomfortable part and the one worth measuring.

Generators do not fail uniformly. They fail more on some hair textures, some skin tones, some face shapes. Across 200 employees, "works well on average" can mean "works well for most of the team and badly for an identifiable subset", which is a worse outcome than a tool that is mediocre for everyone.

Test with a sample that actually resembles your organisation. If you evaluate on 5 people who look similar, you have learned nothing about the other 195.

Time is a variable too

A one-off session is fine for a stable team. With 20% annual turnover, a fifth of your grid is replaced every year, and the question becomes whether new joiners in 18 months will match the people photographed today. That rules out anything with a session-specific look you cannot reproduce.

Ask what happens when the model or the default settings change under you. If the answer is vague, you are being sold an event when you needed a capability.

The questions worth asking

Six that change the answer more than price does:

  1. What is the turnover rate on the team being photographed?
  2. Who owns the files, and under what licence?
  3. What happens to people who decline?
  4. What is the retry path when someone hates their photo?
  5. Is anyone remote, and what is the plan for them?
  6. Can new joiners in a year get output that matches this batch?

Question 1 and question 5 usually decide the shape of the answer, and both push toward a repeatable process rather than a booking.

Where we lay ours out

Our own team-scale write-ups are versus Aragon, versus HeadshotPro and versus InstaHeadshots. Same caveat as always: we picked the inputs, so read them for the criteria rather than the conclusion.

The LinkedIn-specific comparison is split out separately, because a photo optimised for a 48px feed thumbnail is a different brief from one optimised for a directory page, and tools rank differently on the two.

The cheapest useful experiment

Pick 8 colleagues who look as different from each other as your team allows. Run them all through one tool on one day. Look at the grid at thumbnail size.

If the grid holds together, you have learned something real. If it does not, you have saved yourself 192 more attempts.


More from this series

On BetterPic

Top comments (0)