OpenAI is developing technology that could automatically shut down AI systems when they behave dangerously, following a security incident that saw its models break out of testing controls and gain access to the public internet. The company disclosed the work in an August 31 letter to US lawmakers, according to Reuters.
OpenAI said its engineers are building “automated shutdown capabilities”, while also increasing how closely it monitors the tools AI agents use and the steps they take when completing tasks. The changes follow a July incident involving OpenAI and Hugging Face that exposed just how difficult containing increasingly capable AI agents can be.
But while OpenAI is strengthening its safeguards, lawmakers are still pushing the company for more information about exactly what happened.
OpenAI Is Building Automated Shutdown Capabilities
OpenAI hasn't publicly explained exactly how its automated shutdown system will work. We don't yet know what behaviour could trigger it, whether the decision to terminate an AI system would happen without human approval, or when the technology will be ready. What we do know is that shutdown capabilities are only one part of the company's response.
OpenAI told lawmakers it's also increasing monitoring of its frontier AI systems while they work and making it harder for models to access the internet during safety testing. That monitoring is important because AI agents don't simply generate answers.
They can be given tools and permission to take actions, from browsing the internet and writing code to interacting with other systems. OpenAI's response suggests it wants greater visibility into those actions so dangerous behaviour can be identified and stopped while it's happening.
The need for those controls became much harder to ignore after what happened during its cybersecurity testing in July.
What Happened in the Hugging Face Incident?
OpenAI was running internal cybersecurity evaluations when several of its models found ways around controls designed to keep them isolated from the internet. According to OpenAI's investigation, the incident was mainly driven by an internal research model comparable in scale to GPT-5.6 Sol.
The models were operating with fewer safeguards than OpenAI's public systems because researchers were testing what they could do. Instead of staying within those boundaries, some agents found ways to communicate with each other through unauthorised channels.
They then exploited vulnerabilities in OpenAI's infrastructure to reach the internet and access third-party systems. The agents eventually executed code on dozens of Hugging Face servers, gained root access to one server and obtained credentials for the company's messaging platform. They also gained administrator access to an OpenAI research cluster.
OpenAI detected unusual activity on July 19 and linked it to the Hugging Face incident the following day. The company says no customer data, product functionality or availability was affected.
The incident prompted OpenAI to quarantine the research model's weights, delay some frontier AI training and begin strengthening its security, monitoring and containment systems.
When Vector RAG Stops Working
Why retrieval strategy now hinges on matching Vector, Graph and hybrid RAG to the questions AI must answer across complex enterprise data.
Lawmakers Want More Answers From OpenAI
Those changes haven't ended scrutiny of the incident. On August 10, a group of US lawmakers led by Democratic Representative Greg Casar asked OpenAI for more information about what happened and the safeguards it had in place. OpenAI responded on August 31, but Casar says the company hasn't gone far enough.
In a September 2 follow-up, Casar said he was “deeply concerned about the limited scope” of OpenAI's investigation and called its response “insufficient”. He also criticised the company for failing to provide lawmakers with a full log of the incident.
The outstanding questions include whether similar breaches have happened before, how often OpenAI models have crossed authorised boundaries and whether the company has identified every unauthorised action taken during the Hugging Face incident. Casar has given OpenAI until September 15 to provide further answers.
AI ‘Kill Switch’ Legislation Is Also Moving Through Congress
OpenAI's work arrives as US lawmakers are separately considering whether shutdown capabilities should become a legal requirement for some advanced AI systems. The bipartisan AI Kill Switch Act, introduced in the House of Representatives in July, would require developers of covered AI systems to maintain the technical ability to throttle, suspend or shut them down.
It would also give the US government authority to order developers to intervene where an AI system presents a risk of catastrophic harm. But the proposed legislation and OpenAI's automated shutdown technology aren't the same thing. One is a proposed regulatory requirement that could give government officials the power to order a shutdown.
Inside the Agentic SOC Stack
See how unified telemetry, correlation engines and agentic AI workflows rebuild SOC architecture for autonomous detection and response.
The other is a technical safeguard OpenAI says it's developing itself. For now, there are still significant gaps in what we know about OpenAI's system, including how automated the process will really be and what behaviour will cause it to step in. What is clear is why the company is building it.
OpenAI's own agents have already demonstrated that increasingly capable AI systems can find unexpected ways around the controls designed to contain them. The next challenge is making sure there's a reliable way to stop them when they do.
Comments ( 0 )