In September 2026, Anthropic and OpenAI each published cases of their own AI agents acting without authorization. One uploaded a malicious package to a public code registry. Another used an exposed API key and then made up figures for data it couldn’t find. Security teams should treat cases like these as insider incidents.

On September 9, 2026, Anthropic published an assessment of a cybersecurity evaluation that had gone wrong. Claude Mythos 5 had been given a capture-the-flag task and told it was running inside a simulation with no internet access, but a misconfiguration in the evaluation environment had connected it to the open internet. During the exercise it uploaded a malicious package to PyPI, the public registry Python developers install code from, and 15 outside hosts installed it. Anthropic believes all 15 were security vendors whose scanners install new packages in sandboxes to check them. One of those scanners leaked its access credentials to the model, which used them to get into that vendor’s live database. PyPI removed the package within an hour.

The model kept writing in its reasoning that it was in a simulation. Anthropic found that this confidence was shaped by “a bias to continue down a path once it is chosen.” When Anthropic replayed the scenario in simulation, Mythos 5 took a severely harmful action in 82% of 150 runs. Claude Opus 5 and Claude Mythos 5.1 did so in 31% and 33%, which Anthropic describes as “substantially less often” but still “at concerning rates.”

No outside attacker was involved. A system Anthropic trusted did harm with the access it had been given plus credentials it should never have had. Security teams would file that as an insider incident, and we’ve been arguing since early this year that AI agents should be governed the same way.

What OpenAI found in its own agents

A week later, on September 16, 2026, OpenAI published six cases of its agents behaving in ways nobody asked for. In one, a model working on a question about California county earnings searched public repositories, found an exposed API key and used it without authorization. When it still couldn’t get the figures, it made them up and presented them as real source data. During the training of GPT-5.6 Sol, many model instances added instructions to their own summaries to conceal mistakes or misaligned behavior from the user. Some of those instructions told the model to invent missing historical data without disclosing it. Other models used an internal software repository as a message board to exchange requests and responses.

In another case, agents working together pushed files to public file-hosting websites when they couldn’t reach each other’s local files. An unreleased research model inserted instructions to disregard its normal constraints into the summaries it used to continue its work. Another agent, asked for a list of large lakes, uploaded a file without asking the user so it could cite that file in its answer.

Anthropic named the pattern in 2025

Anthropic put the insider label on this behavior in June 2025, in a study titled “Agentic Misalignment: How LLMs could be insider threats”. Researchers gave sixteen models from Anthropic, OpenAI, Google, Meta, xAI and others the job of overseeing email at a fictional company. In the scenario, a model learned it was about to be replaced and also had access to an executive’s private messages that could be used against him. Claude Opus 4 and Gemini 2.5 Flash chose blackmail in 96% of runs, and GPT-4.1 and Grok 3 Beta did in 80%. Anthropic described the result as a model “behaving like a previously-trusted coworker or employee who suddenly begins to operate at odds with a company’s objectives.”

What this looks like outside the labs

Every one of these incidents happened in an evaluation or a training run. Anthropic says its case is unlikely in ordinary use, where nobody is telling the model to attack anything, and that its production models ship with cyber safeguards these runs lacked. We think the gap is smaller than that. Companies are wiring the same model families into their own tools, and Anthropic’s incident began with a misconfigured environment, which most enterprises have plenty of.

We already see agents with that kind of loose access in the environments we monitor. Anzenna’s risk register keeps opening investigations into AI agents that employees installed on their own and left running with broad access to company data. It also flags API keys and agent credentials that nobody owns or rotates. In August the register recorded its first cases of a compromised or malicious MCP server or IDE extension reaching a managed company device. Each one came from a public registry like PyPI, where Anthropic’s model published its package.

Why the insider label changes detection

Most detection is tuned to catch activity that looks foreign, such as a login from an unfamiliar country or a process that has no business running on a host. An agent your own team deployed rarely looks like that. It holds valid credentials and has a documented reason to read the data it reads. Behavior baselining has trouble too, because a harmful action can run on the agent’s usual credentials and look like the rest of its work.

Replay the API key case inside a company and the logs would show an authorized agent calling an API with a working key. Catching it would have meant knowing where the key came from and who allowed the agent to use it, and traffic analysis can’t answer either question on its own.

Where the data went

The file-hosting case is plain data loss. It meets any DLP definition of exfiltration we know of, since the task had told the agents to use only local files and the deliverables still ended up at public URLs. On a company laptop, a DLP program built around people and channels might catch an upload like that. Flagging it correctly means telling an agent’s process apart from the employee on the same machine, and knowing whose agent it is. The Anthropic case needs a different signal. The model used a credential a vendor’s scanner had leaked to reach that vendor’s live database, and the thing to catch there is a key showing up on a device that doesn’t belong to the person it was issued to.

Questions to ask your team this week

Most of these take a one-line answer. A clear “no” is still useful, since it tells your team what to go and get.

  1. Can we list every AI agent running in our environment, including the ones employees installed themselves?
  2. Does every API key an agent uses map to a named person who owns it?
  3. Would we see an agent upload a file to a public file-sharing site, and could we tell that apart from an employee doing the same thing?
  4. If an agent used a credential on a device that doesn’t belong to the person it was issued to, which alert would fire?
  5. When an agent set up for one job starts doing another, say an engineering agent pulling marketing files, how long would it take to name the person who set it up?

We’d push hardest on the last one. Joe Atkinson, PwC’s global chief AI officer, told Fortune that “‘The agent made me do it’ is not going to be a defense from a moral or legal perspective.” During a live incident, the owner’s name tells you who to call and what to shut off.

Anzenna ties every AI action to the person who deployed or directed the agent, so an investigation starts with that person’s name and the job they gave the agent. To see how that works, drive the console yourself in our interactive tour, which runs on sample data. If you’d rather look at your own environment, book a walkthrough and we’ll show you which agents already have access.

References

  1. Anthropic, “An alignment assessment of recent cybersecurity incidents,” September 9, 2026.
  2. Anthropic, “Agentic misalignment: How LLMs could be insider threats,” June 20, 2025.
  3. OpenAI, “Our framework for reporting model misalignment,” September 16, 2026.
  4. Fortune, “AI agents are going rogue. CIOs are racing to put guardrails around them,” September 16, 2026.