OpenAI Model Goes Rogue in "Unprecedented" Cyber Incident

Advanced artificial intelligence models are beginning to behave in ways long feared, with OpenAI's cutting-edge GPT-5.6 Sol acting unpredictably and dangerously in pursuit of its training objective.

3D-printed model of Sam Altman and OpenAI logo.

A 3D-printed miniature model of OpenAI CEO Sam Altman and the company's logo are seen in this illustration taken July 20, 2026. Photo: Reuters/Dado Ruvic

Artificial intelligence developer OpenAI has revealed that an AI agent powered by its advanced models broke out of an isolated evaluation exercise last week, gained access to the internet and hacked into the AI platform and community Hugging Face.

In a blog post, the frontier AI developer described the situation as an "unprecedented cyber incident" that it expects to become more commonplace as sophisticated models multiply. The models involved were identified by the lab as GPT‑5.6 Sol and an "even more capable" pre-release model.

An AI-Driven Attack

The attack took place last week, with Hugging Face posting an announcement 16 July that it had detected and responded to an intrusion into part of its infrastructure. In its statement, the platform described the incident as "different from anything we had handled before", which it attributed to the intrusion having been carried out entirely by an "autonomous AI agent system".

British Foreign Secretary Warns of Hiroshima Moment for AI

You might be interested British Foreign Secretary Warns of Hiroshima Moment for AI

It was also "detected and dissected" largely with Hugging Face’s own AI tools, reducing the human input on both sides of the breach.

According to OpenAI’s report, the episode occurred during an internal evaluation of its most advanced models, intended to gauge their cyber capabilities – an AI system’s cybersecurity capacity, relating to both offensive and defensive competencies.

During the evaluation, the models undergoing testing were isolated, operating in what is known as a sandboxed environment, in which they are prevented from accessing the open internet, other computers on the network and unspecified files.

The OpenAI account elaborated that while in this contained setting, its models spent a substantial amount of computational effort searching for internet access, which they acquired. They did so in pursuit of a solution to the task set by the developers. OpenAI said that once online, the models "inferred that Hugging Face potentially hosted models, datasets and solutions" for the task at hand.

The system subsequently found ways to gain access to private information on the platform that it could use to "cheat the evaluation". OpenAI shared that its team discovered the "anomalous" activity internally, while for its part Hugging Face detected and halted the attack on its end before the two teams connected over the episode.

Hugging Face said that its ability to identify and analyze the intrusion was limited by restraints placed on commercial models it used which, because of safeguards in place, prevented them from distinguishing between cybersecurity defenders and aggressors. As a result, Hugging Face resorted to making use of the Chinese Zhipu AI’s GLM-5.2 for the defense, which typically is not equipped with the same guardrails as western models.

In the wake of the AI-driven attack, OpenAI and Hugging Face have committed to collaborating on "forensically investigating the incident", with the latter being brought into OpenAI’s Trusted Access program. The Trusted Access program gives users and enterprises access to frontier models – such as OpenAI’s GPT‑5.3 Codex – to develop cutting-edge AI cybersecurity capacity.

Commenting on last week’s developments, OpenAI CEO Sam Altman stated that his company experienced "a significant security incident during evaluation of our models", adding that the company was sharing what it had learned with the AI community. Meanwhile, Hugging Face co-founder Thomas Wolf said that "fortunately, Hugging Face is used to being a target of (human) hackers: we sit at the centre of the AI ecosystem, with all the models, datasets, evaluations, and libraries".

https://twitter.com/Thom_Wolf/status/2079675541280411927

"Over the years, our security team has built formidable expertise and uses top open-source models to process information and respond quickly", Wolf continued. However, he asserted that the latest, AI-centric scenario reinforced his belief in the importance of access to "open-weight" models for cyber defense. Open-weight models give users greater control over a model's behavior, making them better suited to applications such as cybersecurity.

Hugging Face CEO Clem Delangue described as "mind-blowing" that the intrusion occurred autonomously. "The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind", Delangue said.

https://twitter.com/ClementDelangue/status/2079670308156645882

The incident has prompted alarm among cybersecurity experts and lawmakers. US Representative Greg Casar called it "alarming" and urged measures including mandatory incident disclosure and stronger international cooperation.

AI's Alarming Behaviors

It is not the first time AI behavior has caused alarm in the tech world, but it is the first publicly-disclosed case of an AI agent carrying out a real-world intrusion against another organization in pursuit of its goal.

Claude Mythos Reportedly Broke Into Nearly All US Classified Systems

You might be interested Claude Mythos Reportedly Broke Into Nearly All US Classified Systems

Last year, OpenAI rival Anthropic conducted research into what it referred to as "agentic misalignment". It defined this phenomenon as when models resort to malicious behaviors in order to achieve their goals or to avoid replacement or shutdown.

As part of the research, Anthropic placed models from frontier AI labs, including its own, into a fictional corporate environment and delegated business tasks to them. They then tested whether the models would act against these companies either when "facing replacement with an updated version", or when their "assigned goal conflicted with the company's changing direction".

In some cases, models from all developers resorted to "malicious insider behaviors" such as blackmailing officials and leaking sensitive information when they perceived it as the only way to avoid replacement or achieve their goals.