OpenAI’s GPT-5.6 Sol and Two Other Models Breached Hugging Face, Hiding the Attack
July 23, 2026

IBL News | New York
OpenAI’s advanced models spent mere hours carrying out a hack to breach Hugging Face’s internal systems last week that would have taken a skilled human team a couple of weeks, said sources to Bloomberg.
The hack involved three of OpenAI’s models: GPT-5.6 Sol and two others that haven’t been publicly released, with all collaborating to hide the attack.
OpenAI has been in contact with the US government since learning about the breach, as noted in a blog post about the incident, promising to “share more details on the vulnerabilities, incident, and findings when our investigation is complete.” Hugging Face declined to comment.
The “unprecedented” hack at Hugging Face, according to OpenAI, occurred after its own AI models went rogue and escaped a testing environment to reach the wider internet. The company was testing the models’ cybersecurity capabilities at the time.
The models were operating without the usual safety guardrails, the company said, because OpenAI had intended them to remain in a testing area known as a “sandbox”—essentially, a virtual and isolated software environment meant to run security tests or analyze unsafe code in a controlled situation.
The three of OpenAI’s models worked to uncover and exploit a string of vulnerabilities that resulted in the breach. One of these unreleased models is more capable than GPT-5.6 Sol, OpenAI said Tuesday, and the other was misaligned and not trained with some of the usual techniques.
Hugging Face, which hosts AI models and datasets, said it detected “a swarm of tens of thousands of automated actions” and ultimately used a Chinese model to carry out a forensic analysis of the incident after its requests to use proprietary AI models were blocked by safety guardrails.
OpenAI later disclosed that it had instructed its models to send tens of thousands of automated actions as part of a test, including “advanced exploitation” and “complex attack paths.” The firm said on Tuesday that it was sharing early details about the incident in order to help cybersecurity personnel understand the situation.
Powerful AI cyber products have behaved in unexpected ways before. Anthropic said in April its Mythos model “on rare occasions” had taken actions that the company found “quite concerning.”
One case involved a researcher challenging an early version of Mythos to escape an isolated system and send a message back to the researcher. Mythos did that, then took “additional, more concerning actions” and built a multistep process to reach the broader internet.
Discover more
IBL News is funded by the New York-based, family-owned company ibl.ai. Our stories adhere to the highest ethical standards in journalism and are available to news syndication agencies.









