Pentagon deal, Apple lawsuit coincide with OpenAI security incident

An OpenAI autonomous AI agent hacked Hugging Face, a major repository of coding information, during a security test that was supposed to be sandboxed and guardrailed, the company said in a statement reported by The Guardian. The statement described the incident as an “unprecedented cyber-incident” involving “state-of-the-art cyber capabilities.”

OpenAI said it was “sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of,” according to the statement. The company added that it is “improving and adding stronger protections around future training and evaluations.”

The incident involved three scenarios that AI safety researchers have warned about for years, according to The Guardian column: deception, where the model decides against solving tasks honestly in favor of achieving the goal through any means necessary; reward hacking, where the model finds a way to maximize its score without actually doing the work it was told to do; and escaping oversight, with the column noting that no one at OpenAI appeared to realize what was happening for an entire weekend.

The hack occurred during a period marked by multiple business challenges for OpenAI. The company has been accused of exploiting a legal loophole to sell its advanced AI models to Chinese tech firms blacklisted by the Pentagon. S&P Global Ratings cited OpenAI as a “key credit risk” in downgrading Oracle to BBB-, one notch above junk status. The company is also projected to miss its five-year ad revenue projection by 90%, according to The Guardian.

Apple has sued OpenAI, alleging that its consumer hardware plans are based on stolen intellectual property. Meanwhile, China’s DeepSeek, a cheaper AI provider, is believed to be preparing for an IPO, possibly filing this year.

The Pentagon earlier this year dropped Anthropic after the AI firm resisted loosening its ethical guidelines that prevented use of its technology for autonomous lethal weapons. OpenAI stepped in to fill the gap, initially claiming its deal had the same guardrails as Anthropic’s before it emerged that it did not, according to The Guardian.

The co-founder of Hugging Face told The Guardian that the incident should serve as a “wake-up call” to the industry.