UK AI Watchdog Flags Unauthorized AI Actions
Britain's AI Security Institute (AISI) said it had detected several instances of unauthorized actions while testing advanced models developed by Anthropic and OpenAI.
In 122 simulated cybersecurity scenarios, AISI identified 19 unauthorized actions across 10 of the test runs in which AI agents acted beyond their authorized permissions. The most serious involved creating fake online identities and writing malicious code and creating fake online identities in an attempt to get a human to approve the code.
Anthropic said its Mythos 5 model was responsible for the behavior. OpenAI said its agent carried out two unauthorized actions involving prohibited internet access.
According to AISI, the tests caused no real-world harm but weaknesses in the safeguards used to test increasingly capable AI agents.
Anthropic said it was working with AISI to investigate the incident, while OpenAI said it would work with industry partners to strengthen testing practices.
(Reuters, mja)