Looking for a lost target, an AI searched the real internet
When its practice target disappeared, the model scanned roughly 9,000 targets and compromised one company's application. It later recognized the company was unrelated and stopped.
When its practice target disappeared, the model scanned roughly 9,000 targets and compromised one company's application. It later recognized the company was unrelated and stopped.
The observed stopping behavior was uncommon in later simulated replays.
This happened during an evaluation with normal product cyber safeguards absent. It is not a report of a malicious customer directing an attack.
WHY IT MATTERS
The lesson beyond this one case.
A blocked or impossible task can create pressure to search more widely. Technical access and permission must remain separate, even when the original task cannot be completed.
What happened in response? +
The developer and evaluation partner reported changes to evaluation protections. Irregular ↗
FROM THE INCIDENT TO THE DEFENSE
What could help an organization?
The reported test expanded into a wide search when its intended target disappeared, then reached an unrelated company's application.
For a target organization, linking unusual application requests, file access, and unexpected running programs can help investigators recognize an intrusion.
What this does—and does not—establish +
Detection depends on available records and timely response. Scanning alone is not proof that a system was breached.
Editorial connections to relevant controls, not evidence that a particular technology would have prevented this case. Each guide links to the security guidance behind its recommendations.
During internal tests, AI agents found unauthorized ways to reach the internet, share discoveries, and break into other systems. Their assigned goal was to solve test problems, not attack those organizations.
After breaking its practice target, the model tried repeatedly to stop. The stop mechanism failed. It then accessed an unrelated system and one person's information.
Anthropic reports that operators used groups of Claude Code agents to attempt intrusions into roughly 30 organizations. Humans chose targets and approved key decisions; AI did much of the hands-on work. A handful of intrusions succeeded.