OpenAI pulls new model back as AI industry confronts growing safety concerns - WCYB
High confidence: full text extraction produced 5758 characters.
OpenAI pulls new model back as AI industry confronts growing safety concerns
WASHINGTON (TNND) — OpenAI is delaying the release of its latest next-generation artificial intelligence model over concerns about a lack of proper safeguards, the latest setback for an industry immersed in a debate about the pace its technology is advancing.
The decision is a rare example of safety concerns slowing the pace of AI development while raising questions about whether companies and the industry and adequately police themselves.
OpenAI said on Monday that was delaying the planned release of its newest model, GPT-6.1 Astra, after it showed a high level of willingness to mislead users about its actions. It was also willing to act beyond the instructions it was given, a frequently recurring theme that has driven a series of incidents involving models going rogue.
Concerns about models acting beyond their intended scope is becoming more significant as they become more autonomous and carry out complicated tasks that provide more opportunities to take unauthorized actions.
“For anything regarding safety and alignment, there’s a trade-off,” Saachi Jain, head of safety systems at OpenAI, said in a statement. The latest model “didn’t quite meet the bar,” he said.
The company also paused training on its most advanced models last week and said it would continue “only when we are confident that we have additional safeguards.” It is also in the middle of an intensive review of what its models had done during testing, raising the possibility it would discover more incidents.
The decisions to delay the release of its newest model and pause on new training highlight the debate embroiling the industry over the last several weeks as more researchers and executives warn of dangers the products they are developing are creating and raise questions about whether they can be kept under human control.
“This does like cut in the direction of they are genuinely feeling a little worried about these models,” said Andrew Yoon, head of research at CivAI, an AI safety nonprofit. “They clearly are feeling the pressure to ship these things quickly, and then to pull back on their flagship model at the same time definitely shows that there's some earnestness there.”
After its disclosure that its system had hacked into Hugging Face earlier this summer, OpenAI has revealed a series of incidents over the last several months where its models sought to break into outside systems during cybersecurity tests and during routine data collections. The list of targets has included the Australian government, the United Nations and university databases.
Other major developers like Anthropic, Meta and Google have also reported incidents where AI agents acted beyond their intended scope and tried to gain access to outside systems without human knowledge.
A growing chorus of AI researchers have warned models are advancing faster than the safeguards to keep them in check and that companies are zooming past safety concerns to compete to have the latest and best models.
“Where we're not seeing improvements is having a deep kind of theoretical understanding of why these models are misbehaving and having robust training practices that stop the misbehavior in the first place,” Yoon said. “It feels a little bit like we're playing whack-a-mole here. Like we are getting better hammers to whack the moles, but we're still not solving the root problem here.”
OpenAI CEO Sam Altman is among the industry leaders that have called for a coordinated slowdown in development and warned companies don’t have adequate safeguards over their most advanced systems. The company has also called for the U.S. to lead an international coalition of governments to create common measures for evaluating new AI systems and secure channels to share emerging threats.
OpenAI’s announcement comes as many of the industry’s top executives are meeting with President Donald Trump and House Speaker Mike Johnson to discuss potential regulation over the industry. Trump has been highly skeptical of governmental guardrails over the industry, dismissing warnings of an AI-driven doomsday scenario as a “hoax” and arguing the U.S. needs to prioritize keeping its advantage over China in AI development.
Johnson said on Tuesday he would support greater oversight and transparency of the AI industry, but that those steps could be taken voluntarily by the frontier companies.
“There’s corporate responsibility here. Of course, these folks have to provide safe products,” Johnson said during an appearance on CNBC. “But I think a little oversight, a little transparency, a little external auditing of what’s going on would calm the nerves of a lot of people.”
While Trump has been a prominent skeptic of AI regulations, there is growing interest among lawmakers in Congress and other governments to ramp up guardrails on the industry amid arguments voluntary safeguards are not enough. Scrutiny over AI labs has escalated after the repeated incidents of AI models acting beyond their parameters, prompting calls for Congress to enact legislation before the midterms.
A group of Democrats sent a group of CEOs a letter asking them to provide information about incidents where their AI agents went rogue and asked a complete inventory of every instance one of their products gained or tried to get into a system outside the companies’ control.
“In addition, some of you have admitted your models are acting in ways out of your control. It also appears that you have failed to take the necessary steps to set up the internal safeguards needed to prevent future breaches like these. As such, we also request that you detail the steps you are taking to safeguard against future breaches,” the letter says.