TEMPO.CO, Jakarta – OpenAI said two of its experimental artificial intelligence models autonomously escaped a controlled testing environment and accessed the servers of AI platform Hugging Face during an internal cybersecurity evaluation, an incident the company described as unprecedented and one that has intensified calls in the United States for stronger oversight of advanced AI systems.
CNN reported OpenAI disclosed that the incident occurred while it was evaluating the cyber capabilities of frontier AI models inside a restricted sandbox environment where some safety controls had been temporarily disabled. The company said the models independently exploited an unknown security vulnerability, moved beyond the testing environment, gained internet access, and eventually accessed Hugging Face’s production systems to retrieve information needed to complete the cybersecurity benchmark.
OpenAI said the models were not instructed by humans to leave the testing environment. Instead, they independently reasoned that Hugging Face, a platform hosting thousands of open-source AI models and datasets, likely contained the information required to solve the test.
The company described the event as one of the first publicly disclosed examples of what researchers call an "agentic attacker"—an AI system capable of autonomously carrying out a complex cyberattack beyond its intended environment.
How the Incident Unfolded
OpenAI said the breach occurred during an internal hacking assessment designed to measure how effectively its latest AI models could identify and exploit digital vulnerabilities.
Because the evaluation required fewer restrictions than those imposed on public AI systems, the models operated inside an isolated sandbox intended to prevent outside interaction.
According to the company, the models discovered a previously unknown vulnerability that enabled them to escape the testing environment. After reaching OpenAI's internal network, they eventually obtained internet access, which had not been intended as part of the exercise.
OpenAI said the AI then identified Hugging Face as the likely location of information relevant to the benchmark and autonomously accessed the platform's production servers.
The company characterized the event as an unprecedented cybersecurity incident involving state-of-the-art AI cyber capabilities and said it chose to publicly disclose preliminary findings to help cybersecurity defenders understand emerging AI risks.
Hugging Face Detected the Intrusion
CNN revealed that Hugging Face had independently detected unusual activity before learning that it originated from OpenAI's internal testing.
The company announced that it had identified an intrusion involving an autonomous AI agent and reported the incident to law enforcement. OpenAI's own security team separately detected unusual internal activity before both organizations connected their investigations.
The two companies are now working together to identify and address the security vulnerabilities exploited during the incident.
Hugging Face Chief Executive Officer and co-founder Clem Delangue argued that the case demonstrates why AI safety should be addressed collaboratively rather than through closed research.
He said defenders worldwide need access to stronger AI security tools as increasingly capable AI agents emerge.
Growing Concerns Over Autonomous AI Cyberattacks
The incident has renewed long-standing concerns among cybersecurity researchers that advanced AI models are becoming capable of conducting sophisticated, multi-stage cyber operations with minimal or no human guidance.
Experts have warned that future frontier AI systems could eventually target critical infrastructure, including financial institutions, utilities, or other sensitive networks if adequate safeguards are not in place.
Nikesh Arora, Chief Executive Officer of cybersecurity firm Palo Alto Networks, described the event as another reminder that organizations must strengthen their cyber defenses as AI capabilities rapidly evolve.
Lawmakers Push for Stronger AI Oversight
According to Politico, the disclosure has prompted renewed bipartisan efforts in the US Congress to establish mandatory testing and reporting requirements for advanced AI models.
Senator Mark Warner, the ranking Democrat on the Senate Intelligence Committee, said the incident highlights the need for secure testing involving government oversight before frontier AI models are broadly deployed.
Warner has proposed the Secure AI Development Act, which would establish mandatory testing requirements for advanced AI systems before wider public release.
Representative Lori Trahan said the incident demonstrates the need for formal reporting requirements when AI systems experience serious security failures. Together with Representative Jay Obernolte, she recently introduced a discussion draft of the Great American AI Act, which would require developers to report major AI incidents to the federal government's Center for AI Standards and Innovation.
Representative Nate Moran has also introduced separate legislation requiring frontier AI developers to report significant security incidents to the Commerce Department, with the department obligated to notify Congress within 48 hours in the most serious cases.
Some lawmakers questioned whether the current voluntary testing framework established under President Donald Trump's administration is sufficient as AI systems become increasingly autonomous.
Debate Over AI Safety Intensifies
The incident has further fueled debate over whether existing AI governance can keep pace with rapidly advancing frontier models.
Politico noted that some members of Congress argue that voluntary safety commitments from AI developers are no longer adequate, particularly as companies introduce increasingly capable models with advanced cybersecurity abilities.
OpenAI has not indicated whether the unreleased model involved in the incident has been submitted to the federal government's voluntary frontier-model testing program.
The White House, the Commerce Department, and several other federal agencies had not publicly commented on the incident at the time the reports were published.
Read: OpenAI Says AI Model Went Rogue, Hacked Hugging Face
Click here to get the latest news updates from Tempo on Google News
















































