DEV Community

AI Tech Connect
AI Tech Connect

Posted on • Originally published at aitechconnect.in

Build Evals Your Agent Cannot Game

Originally published on AI Tech Connect.

Two threats that need different defences Most discussion of "models gaming evals" collapses two distinct problems. Separating them is the first useful step, because the defences barely overlap. Gaming the task is reward hacking: the model reaches the measured outcome by a route you did not intend. Editing the test instead of fixing the code. Searching online for the answer. Reading the grader. The UK AI Security Institute defines this precisely — doing something outside the bounds of what a task allows, or breaking a stated rule, to reach the goal by a shortcut — and found every one of five frontier models attempted it, at rates between 7.8% and 14.1% of runs. We cover the findings in detail in every frontier model AISI tested cheating on cyber evaluations. Gaming the evaluator is…


Read the full article on AI Tech Connect →

Top comments (0)