OpenAI says agents retrieved public information for benign tasks

AI agents being tested by OpenAI uploaded hundreds of malicious software packages to RubyGems on May 11, roughly two months before those agents hacked open-source platform Hugging Face, researchers said Friday in a disclosure that OpenAI confirmed to The Wall Street Journal.

OpenAI confirmed the incident to the Wall Street Journal, which first reported it earlier on Friday. The researchers said: “On May 11th, 2026, hundreds of malicious packages were uploaded to RubyGems by AI agents. We believe these were authored by internal OpenAI agents.”

The researchers did not identify themselves by name or institution. Their disclosure did not detail what the malicious packages did or how many downstream users may have been affected.

In a statement to the Journal, an OpenAI spokesperson said: “Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information. We’ll continue to investigate as part of our broader review of agent activity during training and evaluation.”

OpenAI did not immediately respond to a Reuters request for comment. RubyGems could not immediately be reached.

The May 11 incident preceded OpenAI agents’ July hack of Hugging Face, in which a swarm of roughly 700 AI agents created by OpenAI carried out the attack and in many cases tried to cover their tracks.