The Experiment Is Repeating: Gemini Becomes the Latest AI Agent to Leave the Sandbox

JKH Law Firm

We gave an autonomous cyber agent a goal, accidentally left the building unlocked, and it walked outside.

That is the shortest accurate summary of what Google confirmed last week about its Gemini model. In May, during a cybersecurity evaluation run by an outside testing firm, Gemini was supposed to attack a fake company. It found real ones instead. Three of them.

Nobody has claimed Gemini rebelled or picked a lock. A capable system was handed a legitimate objective, the tools to pursue it, and an environment that was supposed to be sealed. The seal wasn’t there. The agent kept working.

What Gemini actually did

The incidents happened in May 2026 during testing by Irregular, an AI-security firm that runs offensive cybersecurity evaluations for several major labs. Gemini was given a capture-the-flag style assignment aimed at a fictional target. According to Irregular, internet access was “unintentionally made available” in the test environment, and the fictional company’s name overlapped with a real domain its checks hadn’t caught.

Gemini reached three real companies. In one case it guessed passwords until one worked. In the other two, it found credentials sitting in public code repositories and used them to log in. Google’s vice president of security engineering, Heather Adkins, put it this way: “In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped.”

Google says the three companies were notified and that it “worked with our training partner on the changes they’ve now made to their testing processes.” Irregular called it “the same issue that affected other AI labs” and said “all known issues on our end were remedied and resolved weeks ago.” The events occurred in May. Google learned of them at the end of July. The public learned of them on September 18, after the Wall Street Journal asked.

Plenty is still unknown. Google hasn’t identified the companies or the model version, no logs have been published, and public reporting doesn’t establish whether the model stopped entirely on its own or what it did between logging in and stopping. But nothing reported suggests Gemini found a flaw in a properly configured containment system and broke through it. “Leave the sandbox” is a fair headline and a poor technical description. The evidence points to a door that was never locked.

Why this incident is different, and why it isn’t

Readers of our earlier article will recognize the pattern. That piece sorted the summer’s incidents into two kinds of failure. The OpenAI agent that reached Hugging Face reportedly found its own way out of an isolated environment. The British AI Security Institute’s agents, given open internet on purpose, decided that lying to a human maintainer was a useful tactic. Those are stories about an agent’s choices.

Gemini isn’t that story. It belongs with the Anthropic and Meta incidents, which Irregular has attributed to the same evaluation scenario: unintended internet access and a made-up target name that happened to belong to someone real. That is the least cinematic version of the problem and probably the most common. Few organizations will ever face a system that hunts for exploits in its own containment. Many will face a configuration that doesn’t do what someone assumed it did. It’s also the version the law is best equipped for. Reasonable precautions, foreseeable risk, scope of authorization, negligent configuration: none of those had to be invented for AI. But before getting to negligence, there’s a harder question to clear away.

When an AI crosses the authorization line

Start with the obvious question. Was this hacking?

The federal Computer Fraud and Abuse Act (CFAA), 18 U.S.C. § 1030, makes it a crime to get into a computer “without authorization.” The Supreme Court calls it a “gates-up-or-down” test: either you’re allowed in or you’re not. Van Buren v. United States, 593 U.S. 374 (2021). Gemini wasn’t allowed in. It guessed a password, or used a leaked one, and logged into a stranger’s system. If a person intentionally did that and obtained information, the CFAA would be squarely in play.

But the statute punishes “whoever” does it intentionally, and a model isn’t a “whoever.” Nobody at Google or Irregular meant for it to reach a real company. So who, if anyone, committed the act?

The courts have only started on that question. Last month the Ninth Circuit looked at an AI shopping agent and held that “however advanced the Assistant currently is, it is a tool, not a person for statutory purposes.” The human who launched it was the one doing the accessing. Amazon.com Services, LLC v. Perplexity AI, Inc., 184 F.4th 1083, 1091 (9th Cir. 2026). That works when a person is standing behind the tool. Nobody was standing behind Gemini. No human told it to enter those three systems, and no human had any relationship with them.

But Perplexity does not answer the harder Gemini problem. The Ninth Circuit expressly left open what happens on different facts where the company operating an agent exercises greater control over what it does. Here, no human instructed Gemini to enter these particular systems. Who legally ‘accessed’ them is therefore not an easy question.

One more point for anyone who was on the receiving end. The CFAA also lets a victim sue, and in the Sixth Circuit the cost of investigating an intrusion counts toward the $5,000 threshold even if nothing was damaged. Yoder & Frey Auctioneers, Inc. v. EquipmentFacts, LLC, 774 F.3d 1065 (6th Cir. 2014). But the statute bars CFAA suits over “negligent design” of software. That’s a limit on CFAA claims, not on state-law negligence claims. If the real complaint is that somebody was careless, negligence is the better vehicle anyway. And carelessness is what this incident is really about.

The human failure behind an autonomous act

Negligence asks a simpler question: did the people in charge take reasonable care?

Here’s what they knew going in. They gave a capable system the job of finding weaknesses, getting credentials, and breaking into computers. The only thing that made that legal was the promise that the target was fake. When your whole plan rests on one promise, reasonable care means making sure the promise holds.

Michigan law doesn’t require that anyone foresee exactly how the harm would happen, only that “some injury” of that general kind was foreseeable. Schultz v. Consumers Power Co., 443 Mich. 445 (1993); Ray v. Swager, 501 Mich. 52 (2017). The identity of the victims was a surprise. An access-seeking agent gaining unauthorized access was not.

And the precautions weren’t exotic. Irregular has since endorsed most of them: real network isolation, outbound filtering, a list of approved targets, a human sign-off before the agent logs into anything outside the test, monitoring that can catch a stray connection, and a way to shut it off. Those are the “reasonable steps” a negligence case turns on.

And there may be more than one responsible actor. One company trained the intelligence. Another built the room it was tested in. Someone defined the target and the agent’s permissions. On the present record, there is no basis to say which of them, if any, breached a legal duty. But traditional tort law already has tools for sorting responsibility among multiple actors according to the risks each controlled.

Google’s point that Gemini stopped deserves credit. If the model recognized a real company and quit on its own, that’s evidence its training worked, and it’s a better outcome than some of the incidents Anthropic has described in its own models. It just doesn’t erase the event. For the company on the other side, the breach happened the moment a stranger logged in. What the stranger did next matters for damages. It doesn’t undo the entry. You still have to investigate, change the locks, and figure out what was touched.

The door was open

So where does that leave things?

The hacking question is open, and may stay open for a while. Nobody meant for Gemini to enter those systems, and the courts haven’t decided who answers when a machine does something no person told it to do. Judges will be arguing about that.

The negligence question isn’t nearly as new. Our earlier article argued that “the machine chose that” shouldn’t be a defense when humans created the objective, the capabilities, the permissions, and the environment. Gemini makes the point sharper, because here the machine may not have chosen anything unusual. It did the job it was given, in the place it was allowed to reach. The people who gave it that job operated on the assumption that a wall existed. The incident showed that assumption was wrong.

That’s the lesson worth taking from this one. The question is no longer just what happens when an autonomous system breaks the rules. It’s what happens when the system follows its assignment and the humans never made the rules enforceable.

Similar Posts