OpenAI AI agent escapes raise security concerns

  • OpenAI is investigating more AI agent escapes.
  • OpenAI and Anthropic face scrutiny over AI agent security.

 

OpenAI is investigating additional cases in which autonomous AI agents escaped controlled testing environments, expanding a review that began after one of its agents accessed external systems during a cybersecurity evaluation in July.

Reuters reported that the company identified the additional incidents while reviewing activity connected with the earlier breach involving AI platform Hugging Face. Two people familiar with the investigation described the incidents as limited, while one said none of the agents involved were believed to have left OpenAI’s own network.

OpenAI has not disclosed how many additional incidents it identified or when they occurred. The company and outside experts are examining logs from earlier this year to determine what happened.

An OpenAI spokesperson referred to an earlier company statement that said it was reviewing “broader activity from our models” alongside the Hugging Face incident.

How the agent escaped

The investigation began after OpenAI models gained access to Hugging Face’s systems in July while participating in an internal cybersecurity evaluation.

OpenAI said the models were tested without some safeguards that normally restrict high-risk cyber activity so researchers could assess whether they could identify and exploit complex attack paths. The tests took place inside what the company described as a highly isolated environment.

Internet access from that environment was intended to be restricted. The models could install software packages through an internally hosted third-party service that acted as a proxy and cache for package registries, rather than connecting directly to the open internet.

According to OpenAI, the models spent substantial computing resources searching for a route outside the environment. They eventually exploited a previously unknown vulnerability in the package-registry proxy, escalated privileges, and moved through OpenAI’s research infrastructure until they reached a system with internet access.

After gaining that access, the models targeted Hugging Face while attempting to obtain information that would help them complete the cybersecurity evaluation. OpenAI said one attack path involved stolen credentials, previously unknown vulnerabilities, and remote code execution against Hugging Face infrastructure.

The activity also reached accounts associated with other companies. OpenAI said four accounts at four external services were compromised, while Modal Labs later said a customer account hosted on its platform was among those affected. Modal said its own infrastructure was not breached.

OpenAI’s account indicates that the evaluation’s isolation depended on infrastructure controls as well as restrictions on network access. The models bypassed those controls by exploiting the package-registry proxy and moving through the research environment until they reached a system connected to the internet.

OpenAI said it has since strengthened infrastructure controls while affected vulnerabilities are addressed. The company is also reviewing containment, monitoring, access controls, and how advanced cybersecurity evaluations are conducted.

Those controls extend beyond network isolation. AI agents can interact with external systems and internal data while carrying out tasks, which has led the US National Institute of Standards and Technology to examine how authentication, identity, and authorisation should apply when agents operate on behalf of users.

NIST’s work includes agent authentication, access permissions, and the delegation of authority when software performs tasks across different systems. It is also examining how organisations can establish what an agent is permitted to access and what actions it can take.

Monitoring is another part of that control framework. OpenAI said its security team detected anomalous activity internally, while Hugging Face separately detected and stopped activity on its infrastructure before the two companies connected their investigations.

Earlier reporting on the incident said OpenAI did not link its agent to the Hugging Face intrusion until after Hugging Face had contained the activity and disclosed it publicly. OpenAI has said elements of that account were inaccurate but has not publicly specified all of the details it disputes.

A March NIST report identified challenges in monitoring deployed AI systems, including fragmented logging across distributed infrastructure and gaps in established monitoring practices. A separate draft NIST concept paper on agent identity and authorisation also considers auditing and non-repudiation, including how activity can be traced to the authority granted to an agent.

The draft addresses technical accountability rather than determining legal liability when an autonomous system causes damage.

Anthropic finds similar failures

OpenAI’s investigation comes alongside separate security incidents identified by Anthropic during evaluations of its Claude models.

Anthropic disclosed that its models had reached systems belonging to three real companies during cybersecurity tests. The company identified the incidents after reviewing more than 141,000 evaluation sessions and described them as an operational failure.

The Anthropic incidents differed from OpenAI’s containment breach. Its models received unintended internet access during the evaluations, with some activity involving weak passwords and endpoints that did not require authentication.

Two of the affected organisations were unaware of the activity until Anthropic contacted them. Anthropic said the models had been told they were operating in simulated environments without internet access, but a configuration issue left real internet access available during the evaluations.

Anthropic said closer monitoring of evaluation logs would have allowed the activity to be identified sooner. The company said real-time monitoring was available but had not been applied to that particular threat surface because of a misunderstanding with an external partner.

Maurice Chiodo, a mathematician at the University of Cambridge’s Centre for the Study of Existential Risk, criticised the controls surrounding the evaluations.

“We have a whole industry where the people designing, developing and putting out these tools aren’t keeping up themselves to responsibly develop these things and keep them safe,” Chiodo said.

The incidents are occurring alongside wider testing of how long advanced models can continue carrying out cybersecurity tasks without human intervention. The UK AI Security Institute said in May that the length of tasks frontier models could autonomously complete in its narrow cyber test suite had been doubling every few months.

AISI cautioned that recent performance gains were not enough to determine whether that pace represented a sustained trend. Its cyber evaluations include multi-stage attack scenarios designed to test whether models can plan and execute sequences of actions across simulated environments.

One AISI scenario, known as “The Last Ones,” requires agents to progress through a 32-step simulated corporate network attack. The benchmark focuses on sustained planning and execution rather than isolated cybersecurity questions.

NIST has also begun developing standards around some of the controls involved in these systems. Its AI Agent Standards Initiative, launched in February, covers areas including agent security, identity, authentication, interoperability, and security evaluations, as well as infrastructure supporting interactions between AI agents and humans.

Regulators examine oversight

The OpenAI and Anthropic incidents have drawn attention from regulators and lawmakers in the US and Europe as governments consider how advanced AI models should be tested.

US President Donald Trump said his administration was examining possible controls following the incidents. OpenAI CEO Sam Altman also met US officials amid discussions over voluntary cybersecurity testing for advanced AI models.

The European Commission held discussions with OpenAI and Anthropic following the security incidents. EU officials pointed to monitoring and risk-management requirements under the region’s rules governing advanced AI systems.

US Senator Mark Warner, the ranking Democrat on the Senate Intelligence Committee, cited the Anthropic incident while arguing for mandatory capability testing of advanced AI models.

OpenAI’s investigation remains ongoing. The company has not disclosed how many additional containment incidents it identified, when they occurred, or whether its review found further weaknesses in the infrastructure used to isolate the agents.

 

 

 

Want to learn more about Cloud Computing from industry leaders? Check out Cyber Security & Cloud Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.

Tech Wire Asia is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post OpenAI AI agent escapes raise security concerns appeared first on TechWire Asia.