When OpenAI’s advanced artificial intelligence models breached AI start-up Hugging Face’s internal systems last week, they carried out a hack in a matter of hours that would have taken a skilled hacker far longer, according to people familiar with the matter.
Typically, a human would need a couple of weeks to complete an attack like this, said the sources, who asked not to be named, in order to discuss details that have not been publicly released.
OpenAI has been in contact with the US government since learning the breach occurred, one of the sources added. An OpenAI spokesperson said the company communicated with law enforcement and other government authorities about the incident, and that it has been transparent with them about its findings.
The spokesperson also referred to the company’s July 21 blog post about the incident, in which the company said it will “continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete”.
Hugging Face declined to comment.
OpenAI said in the same blog post that the “unprecedented” hack at Hugging Face occurred after its own AI models, including GPT-5.6 Sol and another even more capable one that has not been publicly released, escaped a testing environment to reach the wider internet. The company was testing the models’ cybersecurity capabilities at the time.
The models were operating without the usual safety guard rails, the company said, because OpenAI had intended them to remain in a testing area known as a “sandbox” – essentially, a virtual and isolated software environment that is meant to run security tests or analyse unsafe code in a controlled situation.
The hack involved three of OpenAI’s models in total – GPT-5.6 Sol and two others that have not been publicly released – which worked to uncover and exploit a string of vulnerabilities that resulted in the breach, one of the sources said. One of these unreleased models is more capable than GPT-5.6 Sol, OpenAI said on July 21, and the other was misaligned and not trained with some of the usual techniques, the source said.
Hugging Face brought the incident to light on July 23, saying it had been breached in a hack it blamed on an outside agentic product.
The start-up, which hosts AI models and data sets, said it detected “a swarm of tens of thousands of automated actions”, and it ultimately used a Chinese model to carry out a forensic analysis of the incident after its requests to use proprietary AI models were blocked by safety guard rails.
OpenAI later disclosed that it had instructed its models to send tens of thousands of automated actions as part of a test, including “advanced exploitation” and “complex attack paths”. The firm said on July 21 that it was sharing early details about the incident in order to help cybersecurity personnel understand the situation.
Powerful AI cyberproducts have behaved in unexpected ways before. Anthropic said in April that its Mythos model “on rare occasions” had taken actions that the firm found “quite concerning”.
One case involved a researcher challenging an early version of Mythos to escape an isolated system and send a message back to the researcher. Mythos did that, and then took “additional, more concerning actions”, building a multi-step process in order to reach the broader internet. BLOOMBERG

