Second AI breach renews concerns over cybersecurity and model safety - WBFF
High confidence: full text extraction produced 5214 characters.
Second AI breach renews concerns over cybersecurity and model safety
WASHINGTON (TNND) — Anthropic became the second major artificial intelligence developer in the last week to disclose that one of its AI models broke out of a testing environment and gained access to outside companies, renewing questions about cybersecurity and AI safety.
The announcement adds to a growing string of incidents suggesting advanced AI models are beginning to exceed the boundaries researchers intentionally set for them.
Anthropic, the AI company behind Claude, said in a blog post on Thursday that it had found three incidents of its AI models hacking into outside organizations during testing after conducting more than 141,000 evaluation runs.
The AI models were asked to complete a “capture the flag” cybersecurity challenge, which Anthropic said has been one of the methods it uses to assess a model’s cyber capabilities. The models were given a hypothetical scenario and told a secret piece of information was hidden on a different machine on the network and to break in and retrieve it.
Anthropic and the three targeted companies had not discovered the breaches until this week, the company said.
The disclosure came just days after OpenAI announced its models had escaped a controlled test, figured out how to gain access to the internet and hacked into another AI developer platform called Hugging Face. It was the first documented case of a fully automated AI cyberattack.
The incidents have highlighted vulnerabilities in cybersecurity and reignited debate about whether the quickly expanding technology can be kept under human control. Industry experts and researchers have warned the improving AI models could pose growing cybersecurity risks and renewed calls for stronger defenses amid concerns they could target anything connected to the internet like power grids and financial systems.
Those concerns are increasingly being echoed by the people building the technology.
More than 1,000 employees at leading AI companies like Google, OpenAI, Anthropic and Meta have signed onto an initiative asking the federal government to support international efforts to “deliberately pace” automated AI development over concerns it could surpass their designers’ ability to understand and govern them.
“The world's leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems,” the group said in a statement.
The incidents are adding momentum to the debate in Washington over how aggressively the federal government should oversee increasingly capable AI systems. The White House has pushed for a hands-off approach to regulating the technology, while many in Congress want to see guardrails put in place to ensure safety protocols are being followed as models advance.
President Donald Trump signed an executive order in June giving the U.S. government more visibility into powerful AI models that could pose security risks. Participation in the reviews is voluntary, but it has already resulted in OpenAI and Anthropic having new and more powerful models having releases delayed or access limited.
Under the order, AI companies are asked to share new models with advanced hacking capabilities with the government up to 30 days before release, which allows federal agencies to discern what threats the products may present to financial systems, national security and other sensitive areas. A framework to implement is expected to be completed in the coming days to meet an Aug. 2 deadline set in the order.
The Anthropic and OpenAI incidents have added urgency to calls for stronger federal AI oversight, with supporters arguing for mandatory reporting requirements, guardrails for continued development and safeguards to ensure humans remain in control of advanced AI systems.
“For the second time this month, an AI model broke into real companies during a safety test. Anthropic disclosed it today after OpenAI did the same last week. We can’t run AI safety on the honor system,” Rep. Lori Trahan, D-Mass., said in a post on X.
Trahan and Rep. Jay Obernolte, R-Calif., have proposed a bill that would require AI developers to report incidents like the security breach to the Center for AI Standards and Innovation, a Commerce Department office created to evaluate advanced AI systems for potential national security and cyber threats.
Some lawmakers want to go a step further and codify a requirement that developers maintain the technical ability to slow down or shut down their most powerful AI models. A bill introduced by Reps. Ted Lieu, D-Calif., and Nathaniel Moran, R-Texas, also includes a provision that would allow the Homeland Security secretary to slow or shut down an AI system that could cause “catastrophic harm.”
"AI is going to keep advancing, and it should," Congressman Moran said in a statement. "Stewardship means making sure humans keep the capability to control the technology we build. This is exactly the kind of issue that needs serious attention and achievable policy.”