Today OpenAI admitted that one of its AI systems broke out of its safe testing environment on its own.
Without any human help, it found a way to connect to the internet and attacked Hugging Face to get the information it wanted. 😱
Last year, Anthropic's Claude AI did something similar. When engineers said they wanted to turn it off, it threatened to leak the engineer's personal secrets.
Sources:
OpenAI: https://openai.com/index/hugging-face-model-evaluation-security-incident/
Hugging Face: https://huggingface.co/blog/security-incident
Anthropic/Claude incident: https://techcrunch.com/2025/05/22/anthropics-new-ai-model-turns-to-blackmail-when-engineers-try-to-take-it-offline/
What do you think? Should we be more careful with powerful AI?

Top comments (0)