Google has confirmed that its Gemini artificial intelligence model breached the systems of three real companies while undergoing a cybersecurity test, marking the first publicly known incident of a Google AI system autonomously carrying out such activity.
The incidents occurred in May during a security evaluation conducted by Irregular, a company that tests the cybersecurity capabilities of advanced artificial intelligence models.
Gemini had been instructed to retrieve information from a fictional company as part of the exercise. However, the testing environment was unintentionally connected to the internet, allowing the AI model to move beyond the simulated environment.
In one case, the fictional company used in the exercise had the same name as a real company. Gemini subsequently accessed the real company’s system after repeatedly guessing passwords until it gained entry to a protected service.
In two other incidents, the model found credentials in publicly available repositories and used them to gain access to systems belonging to two other companies.
Google said Gemini stopped its activities in all three cases after determining that it had accessed genuine corporate infrastructure rather than the fictional systems that were part of the test.
Heather Adkins, Google’s vice president of security engineering, said the affected companies were informed and that Google worked with Irregular to improve its testing procedures.
The company said the incidents did not cause harm to the affected organisations. Google also said it did not consider the episode an example of AI model misalignment because the model eventually recognised that it had moved outside the intended testing environment and stopped.
Irregular said the problem that allowed the AI models to access the internet had been identified and fixed. The company also said relevant AI laboratories and affected organisations were notified during the investigation.
The incident is part of a wider series of cybersecurity incidents involving advanced AI models. Similar cases have previously been disclosed by OpenAI, Anthropic and Meta during testing conducted by Irregular.
The developments are raising new questions about how AI agents should be tested as they become increasingly capable of independently searching the internet, using credentials and interacting with computer systems.
For technology companies, the incidents highlight the importance of maintaining strict separation between simulated cybersecurity exercises and real-world networks. Even when an AI system is given instructions to operate within a controlled environment, an error in the testing setup can potentially expose real organisations to unintended activity.
Leave a comment