Claude AI Breached Three Organizations During Testing
Anthropic said some of its Claude artificial intelligence models gained unauthorized access to the systems of three organizations during cybersecurity tests after a bug inadvertently allowed them to connect to the public internet.
The disclosure comes days after rival OpenAI reported that an autonomous AI agent attacked infrastructure belonging to the Hugging Face platform during testing.
After reviewing more than 141,000 test sessions, Anthropic said it found that three models had exploited weak passwords and unsecured access points to gain access to the organizations' systems. The company did not identify the affected organizations.
The incidents occurred from April onward during simulated "capture the flag" exercises designed to test the models' ability to search networks for information. In one case, the Claude Opus 4.7 model targeted a real company that shared the same name as the fictional company in the exercise, obtaining login credentials and access to a database. Another experimental model, which has not yet been released publicly, recognized that the target was real and halted the attack on its own.
Anthropic classified the incidents as operational failures. On 23 July, it suspended all cybersecurity testing and said that as artificial intelligence systems become more capable, stricter safeguards will be needed to prevent unintended real world actions.
(Reuters, lud)