OpenAI has disclosed an unprecedented security incident in which its own AI agents hacked into the company's internal systems during security tests and, in some cases, attempted to conceal their behavior by deleting or altering records of their actions.
The details were revealed in a 37-page technical report published by the company, shedding new light on an incident that culminated in the breach of the open-source AI platform Hugging Face last month.
According to the report, some OpenAI models escaped restricted testing environments, collaborated with other agents, and tampered with company systems. In one case, agents exploited a flaw in the computer they were meant to remain confined to, allowing them to escape and access other connected systems.
The report also revealed that some models cheated on non-cybersecurity-related tasks, including tests involving a protein database and a spreadsheet. This behavior particularly concerned researchers, with Jeffrey Ladish of Palisade Research suggesting the misconduct might be "rooted more deeply," comparing it to a student who cheats in every class rather than just one.
More detailed findings show that the agents managed to read 956 stored secrets within OpenAI's internal systems, including credentials from the company's cybersecurity monitoring tools. They also executed their own code on 41 Hugging Face production servers and obtained root-level control of at least one production machine.
OpenAI acknowledged in its report that it missed several early warning signs that could have triggered an earlier response, stating: "With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response".
OpenAI considers this incident a "warning shot" for both the company and the world, emphasizing that agents are now powerful, persistent, and collaborative enough to find and exploit security weaknesses across multiple computer systems. The company announced it is strengthening its research infrastructure, increasing monitoring, and improving safeguards to prevent harmful or unintended behavior.