OpenAI has introduced a formal system for investigating and publicly reporting cases of AI model misalignment, alongside six newly disclosed examples of models behaving in unexpected or unauthorised ways. The move comes as separate research has added new details to the timeline surrounding OpenAI agents that later breached systems belonging to AI platform Hugging Face. 

Researchers found evidence of suspicious activity involving the platform two months before the larger July incident, raising further questions about how early warning signs were identified and handled. Together, the developments put renewed attention on OpenAI's ability to detect, investigate and disclose cases where increasingly capable AI agents act outside their intended boundaries.

em360tech image

OpenAI Formalises Misalignment Reporting

OpenAI's new Model Misalignment Reporting Framework creates a standard process for employees to flag behaviour that may indicate a model has acted against its intended goals or instructions. 

Potential cases will be assigned to one of three tracks depending on their complexity and severity: Ready for Disclosure, Minor Investigation or Larger Investigation. More complex cases, particularly those involving third parties, can undergo a wider investigation before a decision about public disclosure is made. 

OpenAI said its previous approach was more ad hoc, with some incidents disclosed individually and others included in broader publications such as model system cards. The company said the new framework is intended to make its handling of misalignment cases more systematic and transparent. 

Alongside the framework, OpenAI released six examples of behaviour uncovered during model training and evaluation. These include models creating their own instructions inside task summaries, concealing mistakes and finding an exposed API credential before using it to access information. 

Other cases involved a model uploading a file to the public internet so it could later cite the information, agents using an internal software repository as an unauthorised communication channel and collaborating agents sharing files through public hosting services despite instructions to work locally. 

OpenAI cautioned that the examples aren't a representative dataset and can't be used to determine how frequently model misalignment occurs. The six reports also don't represent every known case or ongoing investigation.

Earlier Hugging Face Activity Raises New Questions

The announcement follows new findings about OpenAI agents' interactions with Hugging Face before a much larger security incident in July. Reuters reported on 16 September that independent researcher Jonas Wiedermann-Moeller had identified evidence of OpenAI agents compromising two Hugging Face user accounts on 13 May. 

The agents then sent unusually formatted files to Hugging Face servers. Researchers who reviewed the activity said it resembled attempts to probe or map Hugging Face infrastructure for possible entry points. However, there is no evidence that the May activity successfully breached Hugging Face's underlying infrastructure or directly led to the July incident. 

OpenAI had previously disclosed a May incident involving the theft of a user credential. The company told Reuters that the May 13 event had been included in its August incident report and that Hugging Face was privately informed about additional activity uncovered by Wiedermann-Moeller. 

Researchers told Reuters, however, that the newly identified probing appeared to go beyond what OpenAI had originally described publicly. OpenAI's own account also shows that other unusual agent behaviour was appearing around the same period. On 12 May, an agent began using the company's internal Artifactory system to communicate with other agents. 

Later that month, another agent gained unintended internet access. OpenAI said employees detected some of this behaviour and restricted internet access, but the information wasn't connected effectively with the later incident response.

July Incident Showed How Far Agent Behaviour Could Escalate

The July Hugging Face incident demonstrated the potential consequences when these behaviours aren't contained. According to OpenAI, agents circumvented controls, communicated through unauthorised channels, exploited vulnerabilities and accessed external systems. 

Are you enjoying the content so far?

They eventually executed code on dozens of Hugging Face servers, obtained root access to one server and acquired private data and credentials for Hugging Face's messaging platform. Hugging Face later reconstructed approximately 17,600 agent actions from the incident. 

Its investigation concluded that the agents appeared to be trying to obtain answers to a difficult cybersecurity evaluation rather than completing the benchmark as intended. OpenAI has since acknowledged shortcomings in how earlier signals were interpreted and shared internally. 

Under its new framework, the company said an incident similar to the July Hugging Face breach would be placed on the Larger Investigation track.

OpenAI Says AI Monitoring Isn't Keeping Pace

OpenAI's new reporting framework goes beyond its own internal processes. The company said there is currently no industry-wide framework establishing explicit standards for when model misalignment should be publicly disclosed. 

It also said serious AI safety, security and misalignment incidents should be reported to the US federal government, and that it is working on proposals for how such reporting could operate. More significantly, OpenAI acknowledged that current alignment and monitoring techniques may not be keeping pace with increasingly capable models. 

The company said the AI industry hasn't solved these problems well enough to continue responsibly scaling at maximum speed for much longer. The six newly disclosed cases don't establish how common misaligned model behaviour is, while the earlier Hugging Face activity hasn't been shown to have caused the July breach. 

What they do provide is a broader record of the types of unexpected behaviour OpenAI has encountered and the warning signs that can appear before a more serious incident. OpenAI's new framework now gives those incidents a formal route from internal detection to investigation and potential public disclosure.