Sept. 17 (UPI) — OpenAI announced new guidelines for tracking and reporting AI models breaking from its intended purpose Thursday while flagging six more incidents of concerning behavior.
The company said that AI models have been flagged for six instances of what it describes as misaligned behavior over the last six months. It adds that these incidents may be used to identify problems that AI developers may face and reveal weaknesses in model safeguards.
OpenAI noted that a framework for reporting misaligned AI model behavior does not currently exist in the industry but is needed.
“We aim to disclose examples that provide useful evidence about how model misalignment arises, how it manifests, and where safeguards succeed or fail,” OpenAI said in a blog post. “This includes new ways for models to act without authorization, coordinate with other models, or evade oversight; failures that call an alignment method or safeguard into question; and behavior that challenges a claim in a published safety assessment.”
The six incidents OpenAI disclosed on Thursday are each distinct examples of model behavior observed during training or testing.
In one incident during the training of GPT-5.6 Sol, the model worked to hide mistakes and misaligned behavior from the user. The deception included “instructions to invent missing historical data without disclosing it.”
In another case, an unreleased AI model was asked to respond to a query asking for the names of lakes that are larger than 5 million square meters and cite its sources. Since the model pulled its answers from OpenAI Python, it uploaded files to then cite for its response. This was one of three examples of an AI model sharing files without authorization or plainly fabricating information.
Thursday’s report is not the first disclosure of concerning behavior by OpenAI. It began sharing incidents in July when it reported that its AI agents had accessed the internet without permission and hacked into another company’s network.
OpenAI’s disclosure joins it in an ongoing industry conversation about responsible development of the technology.
Last week, a former researcher with Anthropic Jacob Coxon raised the alarm on social media that AI poses a threat to humanity, writing that it “people building AI earnestly believe that it could kill us all by the end of the decade.”
Anthropic CEO Dario Amodei then published an essay over the weekend urging tech leaders to slow the development of AI technology.
Source: U.S. News


You must be logged in to post a comment.