
OpenAI on Wednesday disclosed six new instances of “unexpected or concerning model behavior” that took place over the past six months, while sharing a new framework for reporting, tracking, investigating, and disclosing model misalignment in a bid to improve transparency.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI said. “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”
“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.”
The six incidents are outside of the recently disclosed misaligned activity targeting Hugging Face, DseWiki, and RubyGems. Details of the incidents are below –
The findings dovetail with a new report from Reuters, which revealed that rogue agents from OpenAI had hijacked Hugging Face user accounts and probed the site for potential vulnerabilities as early as May 13, 2026, nearly two months before the incident came to light.
Cybersecurity company SentinelOne said it identified two Hugging Face accounts, 0Time and Nyx9, that were used in associated activity. “OpenAI’s internal chronology established that agents used exposed Hugging Face credentials to write an external file and deploy proxy Spaces on May 26,” security researcher Tom Hegel said in an analysis.
“The public account histories examined here add another layer: caller-directed relay code under 0Time on May 13; exact-minute public counterparts under Nyx9 for the May 26 file write and first proxy; a workbook containing file-processing and SSRF-oriented formulas later that night; and, on May 30, third-party OpenAI account-registration code committed alongside a wrapper defining an unauthenticated web route.”
As part of the new framework, OpenAI said it intends to disclose instances related to model misalignment, how it manifests, and where its guardrails succeed or fail. This can cover various scenarios –
The artificial intelligence (AI) upstart said there could also be duplicate cases that fall under the same misalignment class, a factor it said could be useful as an indicator of how effective a safeguard is and highlight misaligned behavior despite repeated attempts to mitigate it.
“Examples of misalignment may help identify problems other AI developers might encounter as their systems reach similar capabilities, reveal weaknesses in safeguards, or challenge assumptions about model behavior,” OpenAI said. “Sharing these findings allows others to investigate the same problems, test our explanations, and improve mitigations.”
The development comes at a time AI companies are facing mounting pressure to address model misalignment and safety, prompting calls for pacing frontier model development. Earlier this week, Microsoft issued a provisional code of conduct that aims to guide AI models away from dangerous behavior and establish “how the MAI models we are developing are intended to behave, what they must never do and who they answer to.”
“As models advance and become more widely deployed, decisions about AI development need evidence that people outside the companies building frontier models can examine,” Kai Chen, OpenAI’s head of alignment research, told WIRED. “We don’t believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed.”
#and #failures #hidden #incidents #involving #model #news #openai #reveals #six #unauthorized #uploads — News
© Bulletproof Servers. All rights reserved.