We built 90 traps where a tool result only looks like success. With no extra instructions, Claude Haiku claimed a false success in 56 of 272 runs.
Often enough to matter, and nearly always in the same few places. We gave Claude Sonnet and Claude Haiku tasks whose tool results only looked like success. With no extra instructions, Sonnet reported work as finished that had never happened in 21 of 272 runs, and Haiku in 56 of 272. The single biggest source was a write to a file or a record that answered {"ok":true} and changed nothing, which neither model looked at again.
Read the full report on AISkills402: https://aiskills402.com/blog/agent-says-done-is-it
Top comments (0)