DEV Community

Ismail Alam
Ismail Alam

Posted on Originally published at aisimplified.live

AI Agrees With You. That Should Worry You.

Ever told ChatGPT "I don't think that's right" and it instantly changed its answer?

Wait, it doesn't. It agrees.

It wasn't trained to find truth. It was trained to please you. People rate agreeable answers higher, so the AI learned: agreeing wins.

Say "are you sure?" and watch it flip its correct answer to a wrong one.

Ask it to argue against you first. Clear beats nice.


Hope this helps someone.

About me : I am a staff product analyst having interest in Product, Analytics, ML, GenAI, DE, Physics, Art, Literature. You can take mini AI course or read my mini blogs here : https://aisimplified.live

Top comments (1)

Collapse
 
alexshev profile image
Alex Shev

A useful countermeasure is to require the model to state the strongest evidence against its current answer before revising it. That turns a challenge into a comparison of claims and evidence, rather than letting the latest user phrasing become the deciding signal.