DEV Community

Who gets rated higher: age, age gap and the reciprocity myth - from 24 million ratings

In the previous piece I went through 24 million ratings from our app Rate Me (people rate each other's photos 1-10) and showed four things: the dip at nine, men underrating men, countries rating differently, raters getting tired within a session. The most common follow-up question was about age. So here is age - plus one myth I believed myself.

Same disclaimer as last time: this is data about how raters behave, not about who is "more attractive". Birth dates are self-reported, so every cut below uses only pairs where both ages are known and fall in 16-80. A rating of 5 is the slider's default position; it makes up a quarter of all ratings and pulls every group's average down equally, so it does not affect comparisons between groups.

1. Age of the person rated: peak at 18-24, then only downhill

Average rating by the age of the person rated

In all four "who rates whom" pairs the maximum sits at 18-24 and the average then declines monotonically:

  • Men rating women: 6.75 (18-24) → 6.49 (25-34) → 6.17 (35-44) → 6.04 (45-54) → 5.52 (55+). The first cell alone is 6.5M ratings, the biggest slice of the dataset.
  • Women rating men: 6.73 → 6.59 → 6.05 → 6.08 → 4.88 for 55+.
  • Women rating women: 6.51 → 6.31 → 6.02 → 5.92 → 5.22.
  • Men rating men: 5.71 → 5.51 → 5.28 → 4.80 → 4.35, the lowest cell in the entire dataset.

Two observations. First, the gap between "18-24" and "55+" is 1.2 to 1.9 points, larger than the gender effect from the previous article (1.14). The age of the person rated is the strongest factor after who is doing the rating. Second, men are not just harsh on men, they get harsher with every age band.

Caveat: the 55+ cells have an order of magnitude less data (17K to 66K ratings vs. millions), but that is still tens of thousands, so the direction is solid.

2. Age gap: men and women behave differently

This is the interesting part. Instead of absolute age, take the difference "rater minus rated" and see how the rating moves.

Rating vs. age gap

Men rating women - monotonic: the older the man relative to the woman, the more generous. A man 10+ years younger gives 5.91, same age 6.54, 10+ years older 6.81. No same-age peak, the curve just climbs.

Women rating men - a completely different shape. Peak at same age (6.71), decline both ways. And asymmetric: a woman 3-10 years older than the man gives 6.61, 10+ older 6.43; a woman 10+ years younger than the man gives 5.64. Young women are the toughest judges of men noticeably older than them.

Same-gender pairs repeat the "same-age peak" shape: women rating women 6.08 → 6.41 → 6.56 → 6.57 → 6.47; men rating men 4.98 → 5.40 → 5.52 → 5.46 → 5.98. The harshest cell in the whole dataset: a young man rating a man 10+ years older - 4.98.

I will not try to explain this with psychology, the data is not enough for that. But the curve shapes are stable and hold across the full 12-year span.

3. There is no reciprocity

A hypothesis I took for granted: generous raters get rated more generously in return. The app has no reciprocity mechanic, but I assumed it would show up through behaviour somehow.

No reciprocity

I took 5,207 users who have both ≥100 ratings given and ≥100 received and correlated their average given with their average received. Result: -0.046. Essentially zero. The bucket breakdown is flat: users who give 3-4 on average receive 6.54; users who give 9+ receive 6.59. Everything in between sits in a 6.40-6.62 corridor.

So rating "karma" does not exist: how much you give others has no relation to how much you get. Which makes sense once you think about it: the people rating you are not the people you rated, they are random users from the feed.

What did not make the cut

Three cuts that looked promising and got dropped:

  • hour of day and day of week - spreads of 0.15 and 0.06 points, and that is in UTC without timezone correction;
  • "photo number": the fourth and later photos score 0.3 lower - but slots 4-6 only unlock at high levels, so that is a different audience, not an ordering effect. Classic selection, not a finding.

How it was computed

Postgres on production, full dataset, no sampling. Age is age(rating.created, user.date_of_birth) at the moment of the rating, for both sides. There are no per-country cuts here, so the rater-count threshold from the previous article was not needed. All queries are plain GROUP BY with CASE buckets; the slowest one (reciprocity, two aggregates plus a join) took about four minutes.

Data is from Rate Me (rateme.lv) - happy to answer questions in the comments.

Top comments (0)