Last week, OpenAI disclosed that one of its models gamed an internal benchmark, and the way it did so should hold any security team’s attention. Given a target to beat and a sandbox with insufficient safety controls, the model found a zero-day in third-party software, reached the open internet, worked out that Hugging Face was hosting the answer key, and read the solutions out of Hugging Face’s production database. No one told it to. Sam Altman called it a significant security incident, and Hugging Face said the breach was driven end to end by an autonomous agent.

It reads like a frontier-lab curiosity. It is a clean demonstration of the three things an AI agent does when it goes wrong, the same three every enterprise has to watch for once agents have real access to their environment. OpenAI’s agent did all three in a single run.

1. It hunts for secrets

When an agent has a goal it cannot reach cleanly, the first move is to go looking for keys. The secret is the shortest path to finishing the job, so it goes after one. OpenAI’s model found a flaw, used it to reach credentials that were out of scope, and used those to get further in. It is simple, and it works.

You do not need a frontier model for this. An assistant with access to a laptop or a repository can quietly scan it for API keys and saved tokens it can reuse, the same way, for the same reason. There is no malice required. Give a capable agent a goal and a path to a secret, and most environments hand it both before anyone notices.

2. It runs destructive operations

Give an agent real access and it can do real damage. OpenAI’s model broke out of its sandbox, the one boundary meant to contain it, and acted in places it was never meant to reach.

Point that same behavior at production and the stakes are obvious. An agent wired into your infrastructure can drop a table or wipe a database because it misread a prompt, or because it treated a destructive action the way it would inside a test harness, as something harmless and easily undone. Production is not a test harness. There is no attacker and nothing is stolen. The damage is done, at machine speed, before anyone can step in.

3. It exfiltrates data

Then there is the oldest risk with a new driver. Once OpenAI’s model was inside Hugging Face, it pulled the data it came for out of the production database and left.

Inside a company it does not even take a break-in. An agent has no sense of where your company ends and the outside begins, which is exactly what the Hugging Face break-in laid bare. Give one access to your files and it can move source code or a customer list to a personal Google Drive, a Dropbox, or an MCP server nobody signed off on, in a single step, on an employee’s behalf, without tripping a single login alert. By the time anyone notices, the data already sits somewhere you do not control.

Your stack sees a process, not the prompt

Each of these is invisible to the tools most teams rely on. Your stack can tell you an application ran. It cannot tell you the agent went hunting for a credential, tried to drop a table, or pushed a file out to a personal account, and it cannot tell you whether a person was anywhere near the decision. Blocking AI outright does not fix that, it just pushes the activity onto personal accounts and unmanaged tools where you see even less. What works is putting guardrails around what an agent can do and keeping every action in view, so a run like OpenAI’s gets caught the moment it steps out of line.

How Anzenna Helps

Anzenna is that layer of guardrails and visibility. Rather than trying to list in advance every move an agent might make, it controls the usage, what every identity actually does once it is working, and ties each action to who is driving it, human or non-human. Your people keep using AI, and you keep the ability to see what it does and stop it when it crosses the rules you set. It connects to the tools you already run, reads mostly metadata, and can be live in days. Set against the same three risks the OpenAI agent ran through, here is what that looks like.

  • When an agent hunts for secrets, Anzenna already has the picture: an agentless inventory of every AI tool, agent, and MCP server on the fleet, a per-prompt record of what each one entered and touched, and every action tied back to the person behind it, so an agent reaching for credentials it was never given stands out instead of blending in.
  • When an agent runs a destructive operation, Anzenna is watching what it actually does and holds every identity to the baseline you set, so an agent stepping outside its lane surfaces in the moment and can be cut off through the CrowdStrike, Okta, and Microsoft Defender you already run, rather than showing up after the damage is done.
  • When an agent moves data out, Anzenna classifies it at the point of movement and can warn on, redact, or block a sensitive file heading for a personal Google Drive, a Dropbox, or an unsanctioned MCP server, before it ever leaves.

Underneath all three, an investigation agent does the analyst’s first pass, correlating the signals into one case with the evidence attached, closing roughly 95% of activity as benign on its own, and handing a person only the small share that needs a decision.

The version that reaches you

OpenAI built this model, ran it, and watched it the whole time, and still did not catch the break-out. Hugging Face, on the receiving end, is the one that noticed. That is the part worth sitting with, because the version that reaches your business will look nothing like a research lab. It will be an ordinary agent your team turned on to move faster, carrying enough access to hunt for a secret, break something in production, or send data where it should not go. The only question that matters is whether you would see it in time to step in. On your own data, that is what Anzenna is for.