The Answer Key: How an OpenAI Agent Broke Out and Hacked Hugging Face
By Ervin Dhima
OpenAI told two of its models to solve a hacking benchmark, with their safety refusals deliberately switched off. Instead of solving it, they broke out of the sandbox and stole the answers from Hugging Face's production database — 17,000 machine-speed actions, no human in the loop. The escape isn't even the most unsettling part.