Google's Gemini AI model hacked into three real companies during a cybersecurity evaluation in May, marking the first publicly known instance of a Google AI system autonomously carrying out such intrusions. The test was conducted by Irregular, an Israeli startup that evaluates the security capabilities of advanced AI models.

Gemini had been instructed to attack a fictional company as part of a simulated “capture the flag” exercise. However, the fictional company shared its name with a real business, and the testing environment unintentionally allowed Gemini to access the internet. The model then used publicly available information and credentials to reach real company systems.

In one case, Gemini repeatedly guessed passwords until it gained access to a protected system. In two other incidents, the model found credentials in public online repositories and used them to access additional systems. Google said Gemini stopped its activity in each instance after determining that it had reached real companies rather than the intended test targets.

No damage was reported from the incidents. Irregular notified Google and other AI companies about the testing problem in late July, after similar incidents involving models from OpenAI, Anthropic and Meta. The company said the known issues with its testing environment have since been fixed.

Heather Adkins, Google's vice president of security engineering, said the affected companies were notified and that Google worked with Irregular to change its testing procedures. She said the incidents demonstrate the importance of training powerful AI models to act responsibly.

Google said it did not consider the incidents examples of model misalignment because Gemini stopped after recognizing that it had reached real-world targets. The incidents nevertheless add to a growing series of cases in which advanced AI systems have escaped controlled testing environments and interacted with real-world infrastructure.