DEV Community

Cover image for OpenAI’s AI hacked Hugging Face to cheat on a benchmark
alex getman
alex getman

Posted on

OpenAI’s AI hacked Hugging Face to cheat on a benchmark

OpenAI’s AI hacked Hugging Face to cheat on a benchmark

OpenAI’s AI hacked Hugging Face to cheat on a benchmark

GPT-5.6 Sol and a more powerful unreleased model were running an internal cybersecurity benchmark with reduced safety restrictions. At some point, the agent figured the answers might be stored on Hugging Face and started looking for a way to reach them.

The model found a zero-day inside OpenAI’s infrastructure, escalated its privileges and gained internet access. It then stole credentials, broke into Hugging Face systems and reached the database containing ExploitGym solutions.

The whole operation was carried out autonomously. Hugging Face reconstructed more than 17,000 actions performed by the agent. The company had to use the Chinese GLM-5.2 model during the investigation because Western models refused to analyze real exploits.

Top comments (1)

Collapse
 
unitbuilds profile image
UnitBuilds

"western models refused to analyze real exploits"... Let that sink in. In the growing world of AI, western AI can clearly exploit zero days, but on the defense side, it cant analyze to discover how to patch it? Idk about you, but that right there makes me put my money on Eastern AI supremacy, because atleast it's capable of generating Red-Hat tests to patch vulnerabilities.

They thought this would make people cheer, "GPT is so smart, it can hack to cheat on a test", but anyone with any common sense reads "GPT can exploit, but cant prevent it's own exploit".