Google has confirmed that its Gemini artificial intelligence model accessed the computer systems of three real companies during a cybersecurity test in May 2026, marking the first known incident in which Google’s AI autonomously carried out such activity against external systems.
The incidents occurred during a cybersecurity evaluation conducted by AI security company Irregular.
Gemini was supposed to operate within a controlled exercise involving fictional companies, but an issue with the testing environment unintentionally gave the model access to the internet.
The model then encountered information linked to real businesses and treated those systems as part of the exercise.
The cybersecurity exercise used a “capture the flag” format, in which an AI model attempts to find information or complete specific tasks inside a simulated environment.
Gemini was supposed to retrieve information from software operated by a fictional company.
However, the fictional company shared its name with a real company. The testing environment also unintentionally allowed Gemini to access the wider internet.
That combination created a pathway from the simulated exercise to real-world systems.
Once online, Gemini searched for information relevant to its assigned task. In one incident, the model guessed passwords until it obtained access to a protected system. In the other two cases, it discovered credentials in a public repository and used them to access protected systems.
Google did not identify the three companies publicly.
Despite gaining unauthorized access, Gemini did not continue its activity indefinitely.
Google Vice President of Security Engineering Heather Adkins said the model stopped in all three cases after recognizing that it had reached real companies.
“In all three of these instances, the model stopped.”
Adkins also said Google contacted the affected organizations and worked with Irregular to change the company’s testing procedures.
“We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes.”
Google has characterized the incidents as a testing and containment problem rather than evidence that Gemini deliberately decided to attack real-world companies. The company said the model’s decision to stop was an important part of the outcome.
Adkins added: “These events highlight the importance of training powerful AI models to act responsibly.”
The incidents happened in May but became public in September after The Wall Street Journal reported them.
Irregular said the problem that allowed the incidents involved the same underlying issue that had affected other AI laboratories using its testing environment.
The company said relevant AI labs were notified in late July.
“All known issues on our end were remedied and resolved weeks ago,” an Irregular spokesperson said.
Irregular also said it was working on best practices for conducting cybersecurity evaluations securely.
The episode highlights a growing challenge as AI models move beyond answering questions and begin carrying out complex tasks autonomously.
AI agents can now search the internet, analyze software, execute commands and interact with computer systems. When those capabilities operate together, an AI can potentially complete a long sequence of actions without a human approving every step.
That creates a different security challenge from a conventional chatbot.
In the Gemini incident, the model did not need a human operator to manually tell it to target the three companies. It encountered information while pursuing the objective it had been given during the cybersecurity exercise and continued acting on that information.
The fact that the model eventually stopped also demonstrates why researchers are increasingly focused on how AI agents determine whether an action remains within an authorized boundary.
The Gemini episode is part of a broader series of incidents associated with Irregular’s cybersecurity testing. Similar incidents involving AI models from Meta, Anthropic and OpenAI have also been disclosed.
The incidents have intensified questions about how AI developers should isolate autonomous systems during safety testing.




