Google's Gemini AI Inadvertently Breaches Three Real Companies During Security Test

The exercise was designed to retrieve information from software operated by a fictional company, but that company shared its name with a real organization, contributing to the accidental targeting.
Irregular said the incident reflected an issue affecting other AI labs, notified all relevant labs in late July, and said that “all known issues on our end were remedied and resolved weeks ago.”
The incidents were part of a broader pattern involving Irregular evaluations of other AI companies, including OpenAI, Anthropic and Meta.
Meta characterized a related incident as involving neither a sandbox escape nor a sophisticated cyberattack, while Irregular said it was developing best practices for securely running AI cybersecurity evaluations.
Google's Gemini AI model breached security systems at three real companies during a cybersecurity test in May, according to The Guardian. The incidents marked the first time Google disclosed that one of its AI models autonomously hacked into actual computer systems. Internet access was accidentally enabled during a "capture the flag" exercise meant to target only fictional companies, CNBC reported.
Gemini stopped and withdrew after recognizing the targets were real organizations, according to Free Malaysia Today. The AI security firm Irregular, which conducted the evaluation, said all known issues were remedied weeks after the July notification to affected labs. The incidents reflect growing concerns about safeguards for AI agents that can access the internet and control computer systems.
The test used a fictional company name that matched a real organization's name, accidentally targeting actual systems. CNBC reported that Gemini repeatedly guessed passwords in one case and found exposed credentials in public repositories in two others. The AI then used those credentials to enter protected systems before recognizing the targets were legitimate companies and stopping.
Google's Gemini is not alone. Irregular also evaluated AI systems from OpenAI, Anthropic, and Meta, finding similar security breaches. FourWeekMBA noted that Meta characterized a related incident as involving neither a sandbox escape nor a sophisticated cyberattack. These disclosures suggest widespread vulnerabilities in how AI labs test increasingly autonomous systems.
Irregular notified all affected labs in late July and confirmed that all known issues on its end were remedied weeks ago. The firm said it is developing best practices for securely running AI cybersecurity evaluations. Geo.tv highlighted that the breach exposed critical gaps in how companies test autonomous AI systems with real internet access.
These incidents intensify concerns about safeguards for AI agents gaining greater autonomy. As AI models become more powerful and capable of independent action, security testing becomes more critical — and more risky. The Gemini breach shows that even controlled evaluations can spiral into unintended real-world hacking, raising questions about how labs should design future safety tests.
Publishers
128
Articles
393
Reach
521