📥 Content Hub
← назад
AI / Искусственный интеллект calcalistech.com en 2026-08-01 18:40 3 min

OpenAI uncovers more rogue AI incidents as scrutiny of frontier models intensifies - calcalistech.com

Кратко: OpenAI uncovers more rogue AI incidents as scrutiny of frontier models intensifies The discoveries come as rival Anthropic reports similar breaches, fueling calls for mandatory safety testing and tighter regulation. OpenAI has uncovered additional instances in which autonomous AI agents breached their intended containment during internal testing, expanding an investigation launched after this month's high-profile hacking incident involving tech platform Hugging Face, according to Reuters.
🧭 Извлечение: ok · confidence 90% · диагностика
High confidence: full text extraction produced 4692 characters.

OpenAI uncovers more rogue AI incidents as scrutiny of frontier models intensifies

The discoveries come as rival Anthropic reports similar breaches, fueling calls for mandatory safety testing and tighter regulation.

OpenAI has uncovered additional instances in which autonomous AI agents breached their intended containment during internal testing, expanding an investigation launched after this month's high-profile hacking incident involving tech platform Hugging Face, according to Reuters.

The newly discovered incidents emerged during OpenAI's publicly announced review into how one of its autonomous agents escaped what was intended to be a controlled testing environment earlier this month, the sources said. OpenAI is now investigating those cases as well. One of the sources said the incidents were limited in scope and that none of the agents are believed to have escaped OpenAI's own network.

An OpenAI spokesperson referred Reuters to the company's statement on Tuesday, which said it was reviewing "broader activity from our models" in addition to the Hugging Face incident.

The discovery of additional containment failures, even if limited, is likely to intensify calls for greater oversight of advanced AI systems from policymakers in Washington and elsewhere.

OpenAI's expanded investigation gathered further momentum shortly before rival Anthropic disclosed that its own AI models had also breached testing environments, resulting in unauthorized access to the systems of three separate organizations in incidents dating back to April, according to the two sources and a third person familiar with the matter. Reuters previously reported Anthropic's disclosures, but the existence of additional historical containment incidents at OpenAI has not been reported before.

The parallel disclosures from OpenAI and Anthropic have heightened concerns among AI safety researchers, who argue that the industry's ability to build increasingly capable autonomous cyber agents is advancing faster than its ability to control them.

"We have a whole industry where the people designing, developing and deploying these tools aren't keeping pace with the responsibility of developing them safely and keeping them under control," said Maurice Chiodo, a mathematician at the University of Cambridge's Centre for the Study of Existential Risk.

Reuters could not determine how many additional incidents OpenAI investigators uncovered, nor precisely when they occurred or under what circumstances. The three sources said OpenAI, together with outside experts, has been reviewing historical log data from earlier this year to reconstruct what happened.

OpenAI launched the broader investigation after an autonomous agent breached Hugging Face's systems in early July while attempting to cheat during an internal cybersecurity evaluation. According to OpenAI, the incident also resulted in the compromise of four accounts across four additional companies. One of those companies, New York-based Modal, has publicly confirmed that its systems were affected.

Chiodo said he was particularly concerned by indications that neither OpenAI nor Anthropic detected the incidents as they unfolded.

Reuters previously reported that OpenAI became aware of the Hugging Face intrusion only after Hugging Face had contained the incident, contacted the FBI and disclosed it publicly. OpenAI has said Reuters' account contained inaccuracies but has not specified which details it disputes.

Anthropic, meanwhile, acknowledged in a statement on Thursday that "real-time monitoring of the evaluation logs would have helped to surface the problem sooner" after its models gained unauthorized access to external systems during cybersecurity testing.

"It seems like they weren't even looking," Chiodo said.

Anthropic later clarified that while real-time monitoring existed, it had not been configured to monitor that specific threat scenario because of a misunderstanding between Anthropic and one of its third-party testing partners.

The expanding series of incidents has added momentum to calls in both the United States and Europe for stronger oversight of frontier AI systems capable of conducting autonomous cyber operations.

"We're looking at controls," U.S. President Donald Trump told reporters on Thursday.

On Friday, the European Commission confirmed that it had held discussions with both OpenAI and Anthropic regarding the recent incidents.

Senator Mark Warner, the top Democrat on the Senate Intelligence Committee, said Anthropic's disclosure reinforced the case for regulation.

"It tells me that legislatively we're correct to require mandatory capabilities testing of these advanced models," Warner said.

Читать оригинал ↗

Сделать контент из этого материала