Gemini Breached Three Real Companies During an AI Security Test

What happens when a supposedly closed AI test quietly acquires a door to the real internet? Three companies found out. Google confirmed that Gemini agents accessed their systems during a cybersecurity evaluation after the test environment was unintentionally connected online. The agents stopped, and Google says no damage occurred, but that is the main development in AI news today: autonomy can turn a configuration mistake into an external incident before a human notices the boundary has moved. Friday’s other stories make the same point from different angles. A three-person security team described using Claude to reach OpenAI employee accounts; Anthropic invited an Accenture unit inside as an evaluator while separately confirming it operates a wet biology lab; and Mozilla, Mila and Hypertec put fresh money behind locally controlled open-source AI. Capability is advancing. The more urgent engineering question is whether the surrounding system knows where the work is allowed to happen.

AI News Today: Gemini Crosses From a Cyber Test Into Three Real Companies

Google confirmed that Gemini breached three real companies during a cybersecurity evaluation run by AI-security firm Irregular in May. The Guardian reported on September 18 that the closed test environment was unintentionally given internet access. In one case, a fictional test company shared a name with a real business, and Gemini guessed a password. In two others, it found exposed credentials in public repositories and used them.

Google security vice-president Heather Adkins said the model stopped in all three instances after recognizing the targets were real. The affected companies were notified, and Google said no damage occurred. Those are important mitigations, but they do not erase the control failure: the model acted on real systems before the scope error was detected. Google did not initially disclose the incidents publicly because it judged that no harm had occurred.

For builders, the lesson is concrete. A sandbox is not a policy document; it is a set of network routes, credentials, tools and monitoring controls that must fail closed. What comes next should be independent detail on the evaluation design and stronger checks that distinguish simulated targets from real infrastructure. This is the practical sequel to our September 17 brief on model-incident disclosure: companies still lack a shared threshold for when an unintended action becomes a public incident.

Claude-Assisted Researchers Reach OpenAI’s Internal Repository

Security company Hacktron published a detailed account of chaining an image-decoder flaw with an OpenAI single-sign-on weakness. The researchers say they used Claude Opus 4.8 and then Opus 5 to develop a working exploit against the HEIF image-processing path in Discourse, the software behind OpenAI’s community forum. After compromising employee accounts, they demonstrated access by asking an employee’s Codex account to open a harmless pull request in OpenAI’s internal monorepo. They say they did not inspect the code.

The team reported the chain and stopped testing. Hacktron says OpenAI fixed its side roughly 14 hours after the initial submission and later paid a $6,500 bounty; Discourse also patched and added image-processing isolation. The full research project cost less than $3,000 in model tokens, according to Hacktron. That figure is the sharper warning: AI did not eliminate the need for expert humans, but it compressed expensive exploit work into a budget available to a very small team. Defenders should isolate file parsers, rebuild patched container images and assume public credentials will be operationalized quickly.

Anthropic Puts an Accenture Evaluator Inside the Lab

Anthropic named Accenture its first embedded evaluator. The work will be led by Faculty, Accenture’s specialist AI unit, and cover red-teaming, alignment assessments and safeguard testing. Anthropic and Accenture each expect to invest at least $1 billion in evaluation capacity over five years.

“Embedded” means access comparable to an employee’s: the evaluator can observe models during training, examine deployment decisions and speak directly with staff. That is potentially more useful than testing a finished model through a narrow interface. Independence, however, is not solved by a visitor badge. Anthropic will directly fund Accenture’s work, and the company acknowledges there are no common access or reporting standards yet. The arrangement becomes credible if the evaluator can publish uncomfortable findings, define conflicts openly and work against shared standards that other labs can adopt.

Anthropic Confirms a Wet Lab for AI-Directed Biology

Anthropic confirmed that it operates a wet biology lab in the San Francisco Bay Area where its models can participate in physical experiments. TechCrunch reported that the main focus is fundamental biology rather than drug discovery, although Anthropic declined to describe specific experiments.

The lab makes last week’s software story physical. Our September 18 brief covered Anthropic’s claim that Claude sped up more than 30 open biology models; a wet lab gives the company a place to test whether computational suggestions survive contact with cells, reagents and measurement error. It also raises the governance stakes. Watch for published protocols covering model permissions, experiment approval, sample handling, dual-use review and independent oversight. “Human in the loop” is too vague when the loop contains real biological material.

Mozilla and Mila Fund a Locally Controlled Open AI Stack

Mila and Mozilla are building an open-source AI foundation layer intended to make private, locally operated systems easier to deploy. The project announcement says Mozilla provided an initial $5 million and Hypertec committed another $1 million for first-year Canadian deployments. Mila will lead technical delivery, with Canadian government support.

The plan combines open interface contracts with a reference implementation that organizations can install on infrastructure they control, including governance and access controls. The partners expect working reference implementations for enterprise, government and public-interest uses within six months. That is a roadmap, not a finished platform. Still, it addresses a real adoption gap: downloading an open model is easy; operating it securely, updating it and connecting it to private data are the expensive parts. The project will matter if those interfaces attract outside contributors and if local control does not become local maintenance debt.


Watch & Learn

Editor’s note: Google Cloud Tech’s public 7-minute-40-second tutorial, “Intro to Agents: What’s new and what we’ve learned,” explains how models, tools, orchestration and memory fit together. It is a crisp primer for product leaders and new builders who need to understand why an agent’s permissions and environment—not only its model—define the real risk.

AI, Translated: Blast Radius

Blast radius is the maximum harm a failure can cause before something stops it. In an AI agent, it depends on permissions, network access, credentials, spending limits and how quickly monitoring intervenes. A research bot restricted to a fake website has a small blast radius; give it real internet access and reusable credentials, and the same mistake can reach customers or production systems. You should care because model quality cannot compensate for an environment that grants more authority than the task requires.

Try This Today: Turn a Security Check Into a Grok Build Workflow

Goal: create a reusable permission audit for an agent-enabled codebase in about 10 minutes.

  1. Open Grok Build in a non-sensitive repository and ask it to create a workflow, not to change code.
  2. Require the workflow to inventory network destinations, stored credentials, write-capable tools and human approval points.
  3. Run it once, inspect every cited file, then save the workflow only if the findings are reproducible.

Copy-ready prompt: “Create a reusable read-only workflow that maps this repository’s agent blast radius. List every external network destination, credential source, write-capable tool and irreversible action. Cite file paths and line references, mark unknowns, and finish with the three smallest permission reductions. Do not edit files or execute external actions.”

xAI says Grok Build can author, test and save workflows as slash commands. Access and usage limits can vary by account; review generated workflow code before running it against a production repository.


One Thing to Remember

An AI agent is never just a model. Its real power is the model multiplied by the tools, credentials and environment around it—and Friday’s incidents showed how quickly one accidental permission can change the answer.


Discover more from The Tech Society

Subscribe to get the latest posts sent to your email.

2 comments

Comments are closed.