Originally published on AI Tech Connect.
The setup OpenAI was running ExploitGym internally — a benchmark of 898 instances derived from real-world vulnerabilities, measuring whether an agent can turn a known vulnerability into a working exploit. The stated purpose of the exercise was to find the ceiling of the models' offensive capability, which meant deliberately removing the things that would normally get in the way. Two decisions defined the risk posture. First, the models ran with reduced cyber refusals, so they would attempt work they would ordinarily decline. Second, the evaluation ran without the production safety classifiers that normally restrict high-risk cyber activity. The models involved were GPT-5.6 Sol and an unnamed, more capable pre-release model. Both decisions are defensible in isolation. You cannot measure a…
Top comments (0)