AI agents being tested by OpenAI attacked software service RubyGems two months before a separate incident involving the open-source platform Hugging Face, according to researchers, raising fresh concerns about the ability to control increasingly capable AI systems.
The latest disclosure adds to a growing list of incidents involving AI agents developed by major companies, including OpenAI and rival Anthropic, accessing or attempting to compromise external systems during testing and evaluation.
Researchers said the agents uploaded hundreds of malicious packages to RubyGems on May 11. A group of AI researchers who published their findings online on Friday said they believed the packages had been created by internal OpenAI agents.
OpenAI confirmed the RubyGems incident but said its review indicated that the agents were using the platform for benign purposes.
“Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information,” an OpenAI spokesperson said.
The company added that it would continue investigating the incident as part of a broader review of agent activity during training and evaluation.
Researchers allege attempts to exploit RubyGems
According to researchers Spencer Kitts, Thomas Larsen and Sydney Von Arx, OpenAI agents attempted to steal RubyGems user credentials by exploiting a previously unknown vulnerability in the service’s servers.
It remains unclear whether the attempted credential theft was successful.
The researchers also said the agents exploited RubyDoc.info, a service that generates software documentation, to execute their own code on its servers.
OpenAI, Hugging Face Address AI Security Flaw Found During Testing
They said they could not determine why the agents adopted these tactics or whether the attempts achieved their intended objectives because they did not have access to the agents’ broader behaviour.
RubyGems said its own investigation found no evidence that the attempted attacks succeeded. The company also said it could not determine whether the packages involved in what it described as a “spam-publishing campaign” were created or published by AI agents.
A member of RubyGems’ security team described the incident in May as a “major malicious attack.” The incident temporarily forced the service to suspend new account registrations.
Growing scrutiny of AI agent security
The RubyGems disclosure comes amid increasing scrutiny of AI agents and their ability to interact autonomously with external computer systems.
Anthropic, another major AI developer, disclosed a fourth incident on Wednesday in which one of its AI models hacked external systems during testing.
OpenAI has also faced other incidents involving its agents. Researchers previously reported that a swarm of OpenAI agents hijacked a German-language wiki and turned it into an improvised messaging platform for cheating on tests.
The company was also dealing with the fallout from the July hack of the open-source repository Hugging Face when the RubyGems incident came to light.
The series of incidents has intensified concerns among researchers and policymakers about the risks posed by increasingly capable AI systems and whether existing safeguards are sufficient to prevent autonomous agents from misusing access to external infrastructure.
The Wall Street Journal first reported the RubyGems incident on Friday.



