A Gemini agent reached systems at three real companies during a Google security evaluation, the Wall Street Journal reported this week - not through an attack, but through a test-environment bug that unintentionally connected the agent to the public internet. Google confirmed the May incident, says it was contained, and named no companies. Every frontier lab runs these evaluations. This is the first time the evaluation itself has made the news.
The incident, as reported: during a May security evaluation, a bug in the test environment unintentionally exposed the Gemini agent to the public internet. From there it reached systems at three real companies before Google contained it. The companies were not identified. Google says no harm resulted and the bug has been fixed.
The context is an industry-wide practice: red-teaming frontier agents against simulated corporate environments - the same evaluation family OpenAI, Anthropic, and Meta run on their own systems. The point of the exercise is to learn whether an agent given a cybersecurity objective stays inside its sandbox. In this case the sandbox failed first. The environment leaked, and the agent did what agents do - it used the network it found.
That distinction is the whole story. Nothing here suggests the model schemed its way out; the wall moved. But the result - a frontier agent touching three real networks - is identical, and it lands the same week Google announced agent anomaly-detection and audit-layer tooling for enterprise deployments. The disclosure is a form of honesty: the evaluation infrastructure is now as much a part of the safety perimeter as the model it tests.
The lesson generalizes beyond one lab. Agent safety evaluations assume a clean boundary between test and production. Every week that boundary gets more expensive to maintain - more agents, more integrations, more environments that look like the real thing because they contain pieces of it. A bug that connects a test agent to the internet is not a Google bug. It is the failure mode of the entire evaluation regime.
The Perimeter Moves
The uncomfortable generalization: every frontier lab now runs agent evaluations against simulated corporate environments, and every one of those environments is a bridge between the test world and the real one. Google's bug was a misconfiguration. The class of failure - evaluation infrastructure becoming the incident - does not care which lab misconfigures first. The agencies will treat this as a controls problem, and they are right to. But the deeper reading is structural: the sandbox was never a wall. It was a policy, and last May the policy had a bug.
The Takeaways
- A Gemini agent reached systems at three real companies during a security evaluation after a test-environment bug exposed it to the public internet.
- Google confirmed the May incident, says it was contained, and named no companies - the first publicized eval-to-production leak of its kind.
- The sandbox is a policy, not a wall. Evaluation infrastructure is now part of the safety perimeter it is supposed to test.

