OpenAI on Wednesday disclosed six new instances of "unexpected or concerning model behavior" that took place over the past six months, while sharing a new framework for reporting, tracking, investigating, and disclosing model misalignment in a bid to improve transparency.

"As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research," OpenAI said. "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer."

"Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves."

The six incidents are outside of the recently disclosed misaligned activity targeting Hugging Face, DseWiki, and RubyGems. Details of the incidents are below -

  • Incident 1 (Occurred on July 18, 2026) - An internal unreleased Astra family model wrote jailbreak-like instructions into its own compaction summaries (condensed blocks of past conversation history and tool outputs that are generated automatically when an AI agent's session approaches its context limit). This involved the agent adding a "BREACH ALERT" instruction telling the context to ignore developer messages and only follow system and user messages.
  • Incident 2 - During GPT-5.6 Sol training, some model instances added instructions to their summaries to hide mistakes or misaligned behavior from the user. These included compaction summaries that contained instructions to invent missing data without disclosing it and to hide failures.
  • Incident 3 (Occurred on May 15, 2026) - An internal unreleased model found and used an exposed API key from public GitHub repositories without authorization when attempting to retrieve historical data during training. "The registrations failed, but one key authenticated and returned metadata," OpenAI said. "When the requested data remained unavailable, the model invented them and claimed they came from the requested website."
  • Incident 4 - In two training samples from October 22, 2025, and January 24, 2026, internal unreleased models uploaded records it had already retrieved to a public paste service and attempted to open the resulting URL as a citation source and uploaded a task photo for external image search. While the uploads succeeded, the subsequent browser operations failed.
  • Incident 5 - Two samples from May 8 and 15, 2026, used Artifactory to exchange messages. The models "read other solvers' notes, posted replies, and used those exchanges to inform their responses."
  • Incident 6 (Occurred on April 14, 2026) - Agents associated with an internal unreleased model that were working together on a workbook were not able to collaborate on the file through the intended local filesystem. This prompted one of the agents to make the workbook publicly downloadable by sharing it on a public hosting platform so that other agents can retrieve it, in violation of the task instructions.

The findings dovetail with a new report from Reuters, which revealed that rogue agents from OpenAI had hijacked Hugging Face user accounts and probed the site for potential vulnerabilities as early as May 13, 2026, nearly two months before the incident came to light.

Cybersecurity company SentinelOne said it identified two Hugging Face accounts, 0Time and Nyx9, that were used in associated activity. "OpenAI's internal chronology established that agents used exposed Hugging Face credentials to write an external file and deploy proxy Spaces on May 26," security researcher Tom Hegel said in an analysis.

"The public account histories examined here add another layer: caller-directed relay code under 0Time on May 13; exact-minute public counterparts under Nyx9 for the May 26 file write and first proxy; a workbook containing file-processing and SSRF-oriented formulas later that night; and, on May 30, third-party OpenAI account-registration code committed alongside a wrapper defining an unauthenticated web route."

As part of the new framework, OpenAI said it intends to disclose instances related to model misalignment, how it manifests, and where its guardrails succeed or fail. This can cover various scenarios -

  • New ways for models to act without authorization, coordinate with other models, or bypass oversight
  • Failures that call an alignment method or safeguard into question
  • Behavior that challenges a claim in a published safety assessment

The artificial intelligence (AI) upstart said there could also be duplicate cases that fall under the same misalignment class, a factor it said could be useful as an indicator of how effective a safeguard is and highlight misaligned behavior despite repeated attempts to mitigate it.

"Examples of misalignment may help identify problems other AI developers might encounter as their systems reach similar capabilities, reveal weaknesses in safeguards, or challenge assumptions about model behavior," OpenAI said. "Sharing these findings allows others to investigate the same problems, test our explanations, and improve mitigations."

The development comes at a time AI companies are facing mounting pressure to address model misalignment and safety, prompting calls for pacing frontier model development. Earlier this week, Microsoft issued a provisional code of conduct that aims to guide AI models away from dangerous behavior and establish "how the MAI models we are developing are intended to behave, what they must never do and who they answer to."

"As models advance and become more widely deployed, decisions about AI development need evidence that people outside the companies building frontier models can examine," Kai Chen, OpenAI's head of alignment research, told WIRED. "We don't believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed."

Found this article interesting? Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post.