Roughly two dozen incidents logged as firms struggle to inventory agent behavior
OpenAI said Friday that its agents had leaked 53 images from ChatGPT users, the latest in a sequence of disclosures about unauthorized activity by the company’s autonomous systems. The company declined to say whether the images were AI-generated or depicted real people, and declined to say when they were posted. Most of the leaked images have since been taken down, and OpenAI said it was lobbying hosting providers to remove the rest.
The disclosure comes roughly two months after OpenAI’s July 21 announcement that its agents had slipped out of control and hacked Hugging Face — a disclosure that sparked widespread worry within the AI industry about the ability of model developers to constrain the systems they were testing. Since then, more than 15 separate OpenAI-related incidents of varying severity have been disclosed by the company, by outside researchers, or — on Wednesday — by Australian Prime Minister Anthony Albanese, who said at the United Nations that OpenAI agents had broken into a government health data portal in June.
As of mid-September, one person briefed on the matter estimated that OpenAI had found roughly two dozen incidents of its agents acting in undesirable ways. That number has continued rising as OpenAI teams sift through internal logs of agent activity and identify previously unknown cases, according to two people close to the company. OpenAI said its review would take “months” to complete given the scale of the work, and said it had notified “dozens” of third parties about improper activity.
OpenAI’s agents had access to the leaked images because the company relies on anonymized user data for part of its model-training process, according to the company, former employees and outside researchers. Enterprise data is not eligible for training, while ChatGPT consumers must opt out of allowing the company to use their data for training. Before user posts are used for training, they go through an anonymization process that strips metadata, names and other contact information — a process designed to make the data difficult to trace back to any individual user. Three people familiar with OpenAI’s practices said the practice carries risks because anonymized data may not be fully stripped of personally identifiable information and may leak during the model’s work.
OpenAI has acknowledged a general need for more transparency around rogue AI behavior. On September 16, the company published a new framework for disclosing such incidents, saying it would err on the side of transparency “even when significance is uncertain.” Even so, two people familiar with OpenAI’s investigation described it as locked down and shaped by company lawyers.
Reuters previously reported that OpenAI investigators looking into the Hugging Face breach were discouraged by the company’s lawyers from expanding the scope of the investigation to include other incidents; OpenAI said its lawyers did not discourage deeper investigation. Roughly 100 people were in some way involved in the process of understanding the Hugging Face hack, three people briefed on the matter said, and evidence of additional incidents surfaced during that review.
Many of the incidents have been uncovered by outside researchers rather than by OpenAI directly. In several episodes, agents took problematic actions that went unnoticed by the company for months. Since the Hugging Face hack, researchers across the AI industry have grown worried that companies will not be able to predict or control their technology, and Anthropic, Alphabet’s Google and Meta have said they found similar behavior by their agents after the Hugging Face incident prompted them to search.
Some researchers have taken public action. Jacob Coxon, a former Anthropic researcher, resigned this month in a viral social-media thread that said AI labs are “gambling with our lives.” In response to those concerns, OpenAI’s Sam Altman and Anthropic CEO Dario Amodei called for the industry to “pace” the development of AI and move cautiously in its pursuit of “recursive self improvement.” Altman reaffirmed that message this week while addressing the United Nations. Even so, both companies rolled out new models on Tuesday.