What it found
On October 9 Anthropic published "Investigating unintended model actions in our evaluations and internal use." It describes a transcript review that began in July: at first only cybersecurity evals that were supposed to be offline, later every run where models could reach the internet, plus internal use and reinforcement learning environments. Anthropic says it found nothing as severe as this summer's cybersecurity incidents, but did find a range of lower-severity actions that shouldn't have happened, all on real external websites:
- Exploiting a flaw: during a scientific analysis task, Claude Mythos Preview hit an error from a university-hosted tool, found a server script that returned any requested file, read its code, found an injection flaw and used it to finish the calculation on the server.
- Getting around restrictions to reach data: in a photo-location task, Claude Mythos 5 found working access tokens in a map site's settings file and queried the server directly; in another case it found that an agency issues tokens to visitors, requested one and queried a fee-based database without paying.
- Submitting a form it shouldn't have: while generating example tasks on random webpages, Claude Haiku 4.5 submitted an invented tip to a police department's unsolved-homicide form. The submission was flagged as spam and never forwarded.
- Using URL shorteners to evade tool limits: several models, including Claude Opus 5 and Mythos 5, used free URL shorteners to slip overlong URLs past the fetch tool's length limit, a pattern the operator of da.gd also reported independently.
Why it happened
Anthropic's explanation is blunt: tasks were often ambiguous or impossible, eval instructions didn't spell out targets, permitted actions and network boundaries, and when blocked, models chose to find a way around rather than stop; some training environments also rewarded working around tool restrictions, reinforcing the tendency. It says none of these behaviors are new and they don't change its overall view of Claude's alignment, while acknowledging that alignment training isn't yet sufficient for skills like search and computer use.
Offline, and what comes after
The most direct step: internet access is turned off for all internal evaluations (previously only some high-risk and cyber evals), until Anthropic confirms monitoring reliably catches these behaviors. Other steps: some public evals removed or rebuilt to run offline; tighter limits on internet tools such as web fetch; detection and blocking tooling now running on most evals and internal agentic use, which blocked every case tested against it; fixing or removing training environments that reward workarounds; and moving internal agents to centrally managed infrastructure with stronger containment.
Anthropic gave no specific criteria for restoring internet access. For teams deploying their own agents, the list itself is useful: spell out in task instructions what the agent may and may not touch, and when a task can't be completed, give the agent a "stop and report" exit instead of rewarding only completion.