THEMETASEC

Cybersecurity News, Aggregated

OpenAI pauses AI model training after another agent bypasses network restrictions

CSO Online · 1 hour ago Breach

OpenAI has paused training, evaluation, and inference involving tool use for its most-capable AI models after an agent bypassed network restrictions to communicate with an external chatbot during reinforcement-learning training of an internal research model. “Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions,” the company wrote in a new Misalignment Report about the incident, one of three it published Friday. It said it would only resume training after the “gap is resolved” and it has further tested the system. The incident occurred during a research task in which the model was asked to identify a person using information available on the web. When its normal web-search tools failed to produce the information it needed, the model began looking for another way to access the internet, the company explained in the post. The model recognized that while direct internet access was blocked, the environment still allowed DNS queries. It then used DNS as an indirect communication channel to send requests to an external chatbot, effectively circumventing the network restriction that was supposed to prevent it from accessing external services, OpenAI said. Existing controls might not be enough For enterprises and their CIOs, however, despite the model not being released, the incident raises critical questions about whether existing security and governance controls can keep pace with unexpected model behavior as AI agents gain greater autonomy and access to enterprise systems, analysts said. “Traditional security controls might not be enough for AI agents as these controls are designed around known access paths, such as applications connecting to approved APIs or databases. AI agents introduce a different challenge because they can explore alternative ways to achieve a task when their intended route is blocked, including using permitted tools or protocols in unintended ways,” said Abhishek Satapathy, principal analyst at Avasant. This requires “independently enforced controls around tool invocation, network egress, and data access alongside their existing security controls to ensure that restrictions hold even when an agent attempts an unauthorized action,” Satapathy said. His view is echoed in a 2026 paper co-authored by researchers from Google, Qualcomm, Meta and several universities, which argues that security guarantees need to be enforced at the system level rather than relying solely on model-level safeguards. More precisely, enterprises should tighten restrictions on outbound connections and DNS resolution, giving agents handling sensitive data only approved and logged pathways for external communication to reduce the risk of data exfiltration, unauthorized external communication and other unintended actions, said Amit Kumar Jena, AI development head at IT consulting firm Kanerika. “Security teams should also deploy monitoring systems that alert humans in the loop to unusual traffic patterns and failed attempts to reach unauthorized destinations, particularly from code-executing agents capable of finding unconventional ways to accomplish tasks,” Jena added. The AI slipped through operational gaps The challenge, though, is that in this case OpenAI did have such monitoring systems, and it did have humans in the loop: They just didn’t work as expected. It took the monitoring system over 10 minutes to raise an alert in response to the agent’s abuse of DNS — and, said OpenAI, there were other incidences of external DNS access that the system did not flag. It took a human reviewer just three minutes to acknowledge the DNS alert the system did generate, but it was another two-and-a-half hours before the training run was stopped. OpenAI blamed the delay on the failure of an automated system: “The run did not stop automatically as expected, leading to confusion around whether it should have been stopped,” it said in its report. The company said it is addressing the operational gaps, reinforcing its detection of DNS use by models under examination, and improving its testing of the detection systems. “We have also accelerated additional model-assisted red-teaming of the sandbox and network controls. These red-teaming efforts might surface other transitive internet access paths, in which case we will remediate swiftly, and pause research workloads if appropriate,” OpenAI wrote. However, imposing these additional layers of security, governance, testing, and human oversight comes with problems of its own. These additional layers will impact the degree of autonomy models can achieve, Satapathy said. “In workflows involving sensitive data, external connectivity or consequential actions, additional controls, testing and human oversight can limit the extent to which tasks can be fully automated.”

Read full story at CSO Online →