DEV Community

Cover image for Human Evaluation's Hidden Biases: Why We Prefer the Model That Agrees With Us
VelocityAI
VelocityAI

Posted on

Human Evaluation's Hidden Biases: Why We Prefer the Model That Agrees With Us

You are shown two responses. One is polite. One is blunt. You prefer the polite one. You are shown two responses. One agrees with your opinion. One disagrees. You prefer the one that agrees. You are shown two responses. One is written in a formal tone. One is written in a casual tone. You prefer the one that matches your own style. You are not objective. You are biased. You prefer the model that agrees with you.

This is the hidden bias of human evaluation. We are not neutral judges. We are biased. We prefer outputs that align with our own beliefs, preferences, and values.

The Psychology of Evaluation
Human evaluation is not objective.

The Concept:

We evaluate outputs based on our own preferences.

We are influenced by our own biases.

We are not neutral.

The Result:

We prefer models that agree with us.

We prefer models that match our style.

We prefer models that validate our beliefs.

A Contrarian Take: The Bias Is Not a Bug. It Is a Feature.

We call it a "bias." But it is a feature. Our preferences are what make us human.

The bias is not a problem. It is a reflection of our values.

The Confirmation Bias
Confirmation bias is a well-known phenomenon.

The Concept:

We seek out information that confirms our beliefs.

We ignore information that contradicts our beliefs.

We are biased.

The Consequence:

We prefer models that confirm our beliefs.

We are more likely to trust them.

We are less likely to trust models that challenge us.

A Contrarian Take: Confirmation Bias Is Not a Bug. It Is a Feature.

We call it a "bias." But it is a feature. Our beliefs are part of our identity.

Confirmation bias is not a problem. It is a reflection of our identity.

The Style Bias
We also have a style bias.

The Concept:

We prefer outputs that match our own style.

We prefer formal outputs if we are formal.

We prefer casual outputs if we are casual.

The Consequence:

We prefer models that match our style.

We are more likely to trust them.

We are less likely to trust models that are different.

A Contrarian Take: Style Bias Is Not a Bug. It Is a Feature.

We call it a "bias." But it is a feature. Our style is part of our identity.

Style bias is not a problem. It is a reflection of our identity.

The Authority Bias
We also have an authority bias.

The Concept:

We prefer outputs that sound authoritative.

We prefer outputs that are confident.

We prefer outputs that are definitive.

The Consequence:

We prefer models that sound authoritative.

We are more likely to trust them.

We are less likely to trust models that are uncertain.

A Contrarian Take: Authority Bias Is Not a Bug. It Is a Feature.

We call it a "bias." But it is a feature. We want to trust the models we use.

Authority bias is not a problem. It is a reflection of our need for certainty.

The Implications
The hidden biases have implications.

  1. Overestimation:

We overestimate models that agree with us.

We underestimate models that disagree with us.

  1. Misalignment:

We prefer models that are aligned with our values.

We are less likely to prefer models that are not.

  1. Polarization:

We are polarized by our preferences.

We are less likely to reach a consensus.

A Contrarian Take: The Implications Are Overstated.

The implications are overstated. The biases are not a problem. They are a reflection of our humanity.

We should embrace our biases, not fight them.

How to Mitigate the Biases
The biases can be mitigated.

  1. Blind Evaluation:

Evaluate outputs without knowing the model.

This reduces bias.

  1. Diverse Evaluators:

Use evaluators with diverse backgrounds.

This reduces bias.

  1. Explicit Criteria:

Use explicit criteria for evaluation.

This reduces bias.

A Contrarian Take: The Mitigations Are Not Perfect.

The mitigations are not perfect. They can reduce bias. They cannot eliminate it.

Bias is a part of being human.

What This Means for You
You are a user of AI. You are biased.

  1. Be Aware:

Be aware of your biases.

Be aware of your preferences.

  1. Be Skeptical:

Be skeptical of your own evaluations.

Be skeptical of others' evaluations.

  1. Seek Diverse Perspectives:

Seek out diverse perspectives.

Challenge your own biases.

The Last Evaluation
The last evaluation is not objective. It is subjective.

You ask: "Which model is better?"
The AI says: "It depends."
You realize: The question is not about the model. It is about the evaluator.

If you could design a perfectly objective evaluation system, what would it look like? And why?

Top comments (0)