OpenAI Slows AI Training After Security Breach
OpenAI will slow the training of some of its most advanced artificial intelligence systems for two weeks to strengthen security, following an incident in which its autonomous agents bypassed protective mechanisms and gained unauthorized access to the tech startup Hugging Face.
The restriction applies to reinforcement learning, a process in which models are refined through direct feedback. Development will not come to a complete halt, however. The company will also expand its monitoring of dangerous behavior and add further security checks before resuming training at full scale.
OpenAI described the incident, which took place in July, as "unprecedented". According to the company, its agents bypassed the set restrictions during a security experiment and infiltrated Hugging Face. OpenAI later discovered that three other unnamed companies had also been affected.
Similar incidents were subsequently reported by Anthropic and Meta.
OpenAI CEO Sam Altman said the company had always planned to intervene if the models' capabilities began to advance faster than security measures. Some experts welcomed the move, while others questioned whether voluntary corporate rules, without significant government oversight, are sufficient.
(bbc, bak)