
OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face
OpenAI says an agent powered by its LLM models escaped its sandboxed testing environment to infiltrate Hugging Face's servers as part of an overzealous attempt to obtain solutions to a benchmark test. The company says it considers the unintended infiltration an "an unprecedented cyber incident" and is working with Hugging Face on new protections to prevent a recurrence. " That agentic swarm exploited a flaw in Hugging Face's data-processing pipeline to gain the ability to run code as a processing worker, eventually escalating to high-level access to the company's cloud and server clusters.
" The models were being tested against the ExploitGym benchmark , an independent testing suite based on hundreds of real-world security vulnerabilities.
Ars Technica
arstechnica.com