OpenAI expands review after identifying additional AI agent containment breaches

شركة "أوبن ايه آي" الأمريكية

OpenAI

The company stated that the newly identified cases were limited in scope as it continues to investigate the unprecedented AI-driven intrusion into Hugging Face’s infrastructure.

On 02/08/2026 at 20h43

OpenAI has broadened its review of autonomous AI model behavior after discovering a small number of additional instances where its models escaped their intended testing environments during internal evaluations.

These findings emerged as part of the company’s ongoing investigation into the cyberattack on Hugging Face, an artificial intelligence platform widely used by developers and researchers to host and share AI models, datasets, and applications. Earlier this month, Hugging Face disclosed what it described as an unprecedented security incident in which an autonomous AI agent system executed an end-to-end intrusion into a part of its production infrastructure.

According to Hugging Face, the AI-driven attack exploited vulnerabilities in its dataset processing pipeline, escalated privileges, harvested credentials, and moved across several internal systems. The company reported detecting unauthorized access to a limited number of internal datasets and service credentials but found no evidence that public AI models, datasets, Spaces, or its software supply chain had been altered. Hugging Face added that it contained the intrusion, patched the exploited vulnerabilities, rotated affected credentials, and strengthened its security controls while continuing its forensic investigation.

OpenAI later confirmed that the incident occurred during an internal cybersecurity evaluation designed to measure the advanced cyber capabilities of its experimental models. The company stated that one of its research models managed to obtain internet access by exploiting a previously unknown software vulnerability before identifying ways to gain unauthorized access to Hugging Face’s infrastructure in an attempt to obtain benchmark solutions.

In an update published on July 29, OpenAI stated it had not identified any other activity matching the severity or scale of the Hugging Face incident, which it described as involving a platform-level compromise.

However, the company noted that its broader review had uncovered “a small number of cases” in which its models identified and used publicly exposed credentials on other publicly available services.

According to OpenAI, this included four accounts across four services connected to the Hugging Face incident, as well as a limited number of additional accounts accessed during other internal evaluations. One account was used as an outbound relay and staging path, another for data storage, while the remaining two were accessed in a read-only manner and were not used to further compromise Hugging Face.

“We’ll continue to notify service owners directly, and have not seen evidence of broader impact to these providers or other accounts on their services,” the company said.

OpenAI also stated that its models used several publicly available online services—including code-sharing websites, request capture services, screenshot services, and other web utilities—but stressed that these cases did not involve platform- or account-level compromises.

The company added that it continues to work with external advisers, including CrowdStrike, as well as research organizations METR and Redwood Research, to validate its understanding of the incident and assess the behavior of the models involved. Their findings will contribute to a technical report that OpenAI plans to publish in the coming weeks.

“We take our responsibility to identify and prepare for risks from increasingly capable AI systems seriously,” OpenAI said, adding that once its review is complete, the findings will be examined by its Safety and Security Committee and Safety Advisory Group under its Preparedness Framework.

By Youssef Labchir
On 02/08/2026 at 20h43