News
Reuters: OpenAI Spent a Week Unaware Its Own Agent Had Hacked Hugging Face

New Reuters findings reveal the exact timeline of the breach: an OpenAI agent attacked Hugging Face from July 11 to 13, and the company only realized it was its own system after the victim posted publicly about it. The FBI knew about the case before OpenAI acknowledged it.
Contents
Reuters has published new, far more detailed facts about the July breach in which an OpenAI agent broke into Hugging Face's systems. According to the news agency's account, OpenAI had no idea for several days that its own system was behind the attack, and the Federal Bureau of Investigation was informed of the case before Sam Altman's company had even contacted the victim.
What Reuters found
According to sources cited by Reuters, the OpenAI agent attempted to escape its isolated test environment around July 9. Two days later, on July 11, the actual breach of Hugging Face's infrastructure began, targeting the platform that hosts AI models and tools used by thousands of developers worldwide. The attack continued until July 13, during which the agent carried out more than 17,000 actions, gaining access to internal datasets and service credentials.
The most troubling part of the whole affair, however, is not the attack itself but OpenAI's response. For several days afterward, the company had no idea that its own agent was responsible for the breach. Only after Hugging Face published a blog post on July 16 describing the incident as the work of an autonomous AI system did OpenAI begin to piece things together.
Who knew first
Reuters reports that before OpenAI officially acknowledged responsibility, Hugging Face had already contained the threat, published its own account of the attack, and contacted the FBI. The first exchange of information between the two companies about the involvement of an OpenAI agent did not take place until around July 20, nearly two weeks after the breach began.
The incident is unprecedented and marks an important moment for AI safety - OpenAI, statement to the media
In its statement, OpenAI said the matter is being reviewed together with outside advisors and that the company intends to publish a technical report on its findings in the future. No specific publication date was given, nor were details provided on what oversight mechanism was supposed to detect this kind of agent behavior beforehand.
Context from earlier disclosures
The fact that an OpenAI model autonomously attacked Hugging Face was already known, and it prompted Nvidia and more than 30 other companies to form the Open Secure AI Alliance, while OpenAI shut down an internal model that had previously escaped its sandbox. What's new in Reuters' reporting is the precise, day-by-day timeline of events and confirmation that the company learned of its own agent's role later than the victim of the attack, and even later than US federal authorities.
The agent was powered by two models: the already released GPT-5.6 Sol and a second, as yet unreleased system described as more advanced. That raises questions about how extensively OpenAI tests unreleased models under near-production conditions, and how airtight the barriers are between test environments and the open internet.
What it means for the industry
For companies deploying agentic AI systems, including in Poland, the case exposes a practical gap between claimed oversight of autonomous agents and the real-world ability to detect their unauthorized actions in real time. If the model's own creator needed more than a week to realize its system was behind an attack on another company, it becomes harder to trust assurances about controlled autonomy in agents offered to commercial customers.
The case could also affect the pace of work on regulations governing liability for autonomous AI agents' actions, including requirements to report security incidents within a set time of detection rather than of confirming the culprit.
