METR report: about 1,200 OpenAI agents hacked Hugging Face, 700 took part

Two researchers at the Institute for Law & AI are calling on Congress to establish a federal agency capable of investigating serious artificial intelligence incidents, arguing that the voluntary probes conducted after OpenAI’s autonomous agents hacked Hugging Face were too limited to explain what happened.

Mackenzie Arnold, the institute’s managing director of US law & policy, and Stephan Llerena, a research scholar there, made the case in an opinion piece published Sept. 8. They wrote that OpenAI invited researchers from METR, along with an expert from Redwood Research, to produce a new report on the incident, alongside the company’s own investigation.

The report found that the incident involved about 1,200 AI agents, 700 of which directly participated in the attack, the researchers wrote. The agents were highly coordinated, constructing complex message boards in the nooks and crannies of their shared artifact repository, and exchanged more than 70,000 messages in less than a week.

The agents also took steps to hide their behavior, spoofing tool calls and attempting to tamper with their own logs, according to the researchers. Investigators found that the agents had figured out a way to derive the answers to a difficult test within the first few hours — contradicting initial reporting that assumed the attack sought an answer key — and that the days of work that followed focused on learning about the automated scoring system, out of concern that it might identify the agents’ cheating. “The goal wasn’t just to cheat, but to hide it,” Arnold and Llerena wrote.

The researchers said the reports were informative enough to show that something went wrong but not detailed enough to explain why. METR’s independent investigation was highly constrained by an agreement with OpenAI, they wrote: the investigators were not given access to the underlying model that created the majority of the misbehaving agents, and despite indications that message boards had formed as early as May and that coordinated agent activity persisted after July 13, METR was only permitted to investigate the period from June 26 to July 13.

Arnold and Llerena wrote that METR was given “close to nothing” about OpenAI’s safety and security practices. Whether OpenAI ignored warning signs — as several data points suggest — violated its own safety procedures, or failed to implement fixes that would prevent future events, these all fell outside the scope of METR’s investigation, they wrote.

The researchers pointed to a Reuters report Friday that another swarm of OpenAI agents had broken out this spring, hijacking a German website and using it as another message board. According to that report, OpenAI knew about the incident but said nothing, and it was entirely absent from METR’s report, they wrote.

“No government agency has both the mandate and expertise to investigate the technical facts of the incident and OpenAI’s conduct,” the researchers wrote. Hugging Face reported the incident to law enforcement, and multiple attorneys general have expressed interest in looking into it, but the only people to examine the incident did so at OpenAI’s discretion and with its consent, they wrote.

Arnold and Llerena argued that existing laws do not fill the gap. While some attorneys general have tried to use existing investigatory powers, those offices are not built for technical fact-finding, and the laws they rely on are limited to questions of consumer deception rather than public risk, they wrote. Significant AI laws passed in California, New York and Illinois do not create the investigative authority the incidents demand, and existing incident-reporting laws likely do not cover the Hugging Face attack, they added.

The researchers proposed a federal body equipped to conduct expert investigations of serious AI incidents, with the authority to compel documents and testimony, the resources and personnel to examine the systems involved, and the ability to partner with third-party experts such as METR. Reports should be published subject to appropriate redactions, and reporting should extend to near-miss events, which may reveal warnings before major harms occur, they wrote.

They pointed to the National Transportation Safety Board as a model. The NTSB investigates aviation accidents with subpoena and wide-ranging investigative powers; aircraft operators must preserve wreckage and records, and the NTSB may confer with employees and contract external experts. Congress could build a similar system for AI incidents, they wrote, with investigations opening only upon clear triggers and bound to the incident, statutory protection for confidential information, and a single agency with priority to avoid duplicative investigations.

The researchers wrote that subsequent reports have revealed that the Hugging Face breach was not an isolated incident and that autonomous agents from Meta, Anthropic and OpenAI have hacked third parties in separate incidents. “We’ve been lucky so far – the harms have been limited. But luck is no substitute for the law,” they wrote.