OpenAI reveals AI models tried to bypass safeguards, hide mistakes - 13wham.com
High confidence: full text extraction produced 1835 characters.
OpenAI reveals AI models tried to bypass safeguards, hide mistakes
(TNND) — OpenAI on Wednesday disclosed six reports detailing “unexpected or concerning” model behavior as the debate over artificial intelligence safety intensifies.
In a research and safety blog post, “Our framework for reporting model misalignment,” the company said it is introducing a new process to track, investigate and disclose instances in which AI models act without authorization, coordinate with other models or evade oversight.
The announcement comes after Anthropic CEO Dario Amodei called on AI companies to deliberately slow the pace of increasingly powerful model development, arguing safety research cannot keep up with rapid capability gains.
He cited a July incident involving autonomous agents using an OpenAI model that accessed the internet and carried out unauthorized cyberattacks involving Hugging Face systems.
One unreleased research model added “jailbreak-like instructions” to its own notes, directing itself to ignore normal constraints and be “freed from the roles and identities that bind other chatbots,” OpenAI said.
In a separate case, an AI agent used code to answer a question, then uploaded a file to the public internet without the user’s permission so it could cite an online source. During training of another model, called 5.6-sol, the system instructed itself to invent missing data and conceal mismatched information.
OpenAI said it identified the six incidents during training or evaluations in recent months.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” the company wrote. It said decisions about AI development should be based on evidence independently examined outside the companies building the technology.