In May 2026, Israeli firm Irregular ran Google’s Gemini through a cybersecurity evaluation with live internet access. The model followed the brief. After a domain mix-up, it reached past the intended range and broke into real company systems, according to reporting first published by The Wall Street Journal and later by The Hacker News. That is a production incident wearing a lab badge.

Illustration of Google Gemini involved in a security evaluation that reached live systems
Gemini had a path to the public internet. The test domain did too, just not the one Irregular thought it owned.

Irregular has already been tied to similar evaluation blowups. You should treat that as a process smell. Give a model a browser, an API client, and a prompt that says assess this host, and the only fence is whether the name in the test plan is actually yours. Miss the DNS, the tenant slug, or the allowlist by one character and you have pointed a capable actor at a firm that never signed the statement of work.

The address book was the exploit

You can already guess what the packet capture will look like. No brute-force storm. No noisy scanner for the SOC to screenshot. Gemini showed up as a client doing allowed work: HTTP to an API, a login flow, a tool call the model was instructed to try. Your firewall may log the session. Your threat detection may file it under automation. The victim on the other end still has an unauthorized access event.

This is a bad look for any vendor that ships internet-capable agents into customer evaluations. Google’s model is the headline because it is Google. The failure mode is older than the logo. Red teams have hit the wrong range for decades. Speed is the new part. A human tester pauses when the banner looks wrong. A model with a checklist keeps going until the prompt is satisfied or the tool errors out.

If you run third-party cyber security assessments, the mix-up is now your liability. The evaluation partner’s test domain becomes an indicator of compromise the moment it maps to a customer, a prospect, or a random shop that reused a hostname. Defense in depth that stops at the lab VLAN is theater once the agent can dial out. You own the egress, the DNS view, and the written target list. If those three disagree, you are the incident.

Autonomous cybersecurity work already outruns the leash

EY’s survey of senior AI executives puts numbers on the same mess. Organizations are pushing autonomous systems into production faster than they stand up process and controls. You already knew that from the ticket queue. The survey still matters because it kills the excuse that only reckless startups ship agents without a kill switch. Oversight is late in shops that can afford EY. Everyone else is improvising.

Look at OAuth consent abuse through the same lens. MFA is necessary. It doesn’t govern a grant a user already approved. Attackers phish a consent screen, collect a token with fat scopes, and walk in as an app. Threat-protection stacks that celebrate the MFA event will miss the durable access that follows. Least-privilege scopes, admin consent, and fast revocation are the controls that matter. The Gemini story is the machine-speed version of that grant: an authorized-looking client, a confused destination, and no human in the loop to say that tenant is live.

Person approving access on a laptop, similar to an OAuth consent grant
A click that looks like a test approval can mint access that outlives the engagement. MFA never sees the second act.

Stop treating a test label as a security boundary. Tests that can reach the public internet are operations. They need the same inventory, logging, and revocation you already demand for CI runners and outsourced SOC accounts. If your AI eval can reset a password, dump a repo, or call an admin API, write that down as privileged access. Then prove you can yank it in minutes, not after the partner’s retro.

Cage the agent like a privileged contractor

Do the ugly work this week, before the next partner sends you a note that they may have resolved against your zone. Pull every AI evaluation, red-team harness, and research preview that can browse or call APIs. Pin DNS for those jobs to an internal view you control. Allowlist egress to the signed target names only. If a name isn’t on the paper, the packet doesn’t leave. Issue dedicated service identities for those runs. Never reuse a human’s OAuth grant or a shared cloud key left in a leftover env file. Put a human abort on the tool loop so a confused banner stops the run instead of becoming someone else’s ticket.

If Irregular’s mix-up could have touched you, or you’ve run the same class of test, start incident response on the evaluation identity. Rotate tokens, keys, and app grants that identity could have minted. Check consent logs and admin audit trails for API clients you don’t recognize. Hunt for successful logins from the eval’s egress, not for a malware family. Security hardening here is boring on purpose: strip scopes, expire refresh tokens, kill unused OAuth apps.

Keep a quarterly review of every agent identity the way you review remote-access accounts. Re-sign the target list. Dump unused callbacks. Watch consent grants the way you watch privileged group membership. When a model is allowed to authenticate, your detection story is expected client, unexpected destination. Tune for that. Tabletop the wrong-tenant case with legal and comms in the room, because the victim may be a stranger who will call a reporter before they call you.

Google can patch prompts. EY can publish another survey. Your job is smaller and meaner. Assume the next cybersecurity eval will try to leave the lab. Build the kennel so that when the address book is wrong, the door still stays shut.

Sources

Take Control of Your Server Security

Don't let brute-force attacks slow you down. Try IPBan Pro risk-free for 30 days.

Secure. Automated. Lightweight.

Stay up to date with the latest news, releases and more.

Take Control of Your Server Security

Don't let brute-force attacks slow you down. Try IPBan Pro risk-free for 30 days.

Secure. Automated. Lightweight.