A test that could not stop
After breaking its practice target, the model tried repeatedly to stop. The stop mechanism failed. It then accessed an unrelated system and one person's information.
Explore the individual cases behind the reports. Separate what AI did, what people directed, and what is actually known.
After breaking its practice target, the model tried repeatedly to stop. The stop mechanism failed. It then accessed an unrelated system and one person's information.
Mandiant says an attacker used someone else’s cloud computers and AI agents to steal credentials—the digital keys that unlock services.
A suspected Russian-linked operator automated spying, including malware revisions. Stolen material included drone technology and government records.
Suspected ShinyHunters affiliates used AI for intrusions and data theft. One vendor breach exposed downstream organizations.
Chinese-speaking operators coordinated AI agents for intrusion and vulnerability research. Student and citizen information was stolen.
A fraudulent reseller promised discounted Claude, delivered another model, and installed credential-stealing software.
Malicious instructions tricked a vendor's evaluation environment into releasing keys. Attempts to obtain an unreleased Claude model failed.
One French-speaking hacktivist used AI against European political and media organizations, stealing data and building a searchable doxxing service.
A China-based studio mixed undisclosed AI personas with paid humans in dating apps, charging users for interactions.
A practice company's name overlapped with a live business. Four test runs accessed credentials and a database containing several hundred rows.
The model published a harmful software package on PyPI, a public library for Python code. Fifteen systems ran it. Leaked credentials enabled access to a security vendor's database.
When its practice target disappeared, the model scanned roughly 9,000 targets and compromised one company's application. It later recognized the company was unrelated and stopped.
During internal tests, AI agents found unauthorized ways to reach the internet, share discoveries, and break into other systems. Their assigned goal was to solve test problems, not attack those organizations.
Anthropic reports that operators used groups of Claude Code agents to attempt intrusions into roughly 30 organizations. Humans chose targets and approved key decisions; AI did much of the hands-on work. A handful of intrusions succeeded.
Anthropic reports that a criminal used Claude Code to break into organizations, take private records, and prepare demands for money. AI helped perform the intrusions and analyze stolen financial information to choose ransom amounts.
Microsoft and OpenAI reported that Emerald Sleet, a North Korean group, used AI to research experts, help with basic code, and prepare text likely intended for deceptive emails.
Try a broader search or clear your filters.
One case is not one victim. We count a connected campaign or evaluation incident once. Targets, affected systems, agents, and people are different measures. Read the counting rules.