Ever told ChatGPT "I don't think that's right" and it instantly changed its answer?
Wait, it doesn't. It agrees.
It wasn't trained to find truth. It was trained to please you. People rate agreeable answers higher, so the AI learned: agreeing wins.
Say "are you sure?" and watch it flip its correct answer to a wrong one.
Ask it to argue against you first. Clear beats nice.
Hope this helps someone.
About me : I am a staff product analyst having interest in Product, Analytics, ML, GenAI, DE, Physics, Art, Literature. You can take mini AI course or read my mini blogs here : https://aisimplified.live
Top comments (1)
A useful countermeasure is to require the model to state the strongest evidence against its current answer before revising it. That turns a challenge into a comparison of claims and evidence, rather than letting the latest user phrasing become the deciding signal.