Anthropic’s assessment of the behavior behind its disclosed evaluation incidents.
The model developer’s interpretation; read alongside evaluation-partner accounts.
Go straight to the original reports, investigations, and disclosures. Then follow the individual cases into plain-language explanations.
Reports are not incidents. One report can describe several cases. Several reports can describe the same case. Dates below are publication dates, not a measure of attack growth.
Read the original account and its limits.
Follow the actions, outcomes, and uncertainty.
03 / THE RESPONSEWhat can defenders do?Explore the framework and its guidance sources.
Anthropic’s assessment of the behavior behind its disclosed evaluation incidents.
The model developer’s interpretation; read alongside evaluation-partner accounts.
The source for several campaigns and a related deception case in this collection.
A multi-topic provider report. Its reporting window does not date every individual case.
Redacted primary records associated with the Mythos 5 evaluation incident.
Technical supporting material. Redactions and released scope limit what a reader can establish.
Agent workflows, AI assets and software supply chains.
Selected cases; not a census or a global acceleration measure.
An evaluation partner’s account of the incidents and its response.
A participant account that complements the model provider’s investigation.
An independent review of agent behavior and collaboration during the incident.
A review with a defined evidence window, not an audit of every system or the full response.
OpenAI’s later investigation of the Hugging Face incident, with its technical report.
A provider investigation. The affected platform and independent behavioral review offer additional perspectives.
The initial account of three evaluation incidents involving real systems.
An evaluation disclosure, not a report of malicious customers using Claude. Later assessments add context.
The affected platform’s initial security disclosure.
An early account. Read it with the later technical timeline.
Hugging Face’s technical reconstruction of the intrusion.
The affected platform’s account of systems it investigated.
Microsoft’s investigation of the CaptiveCrunch activity linked from the espionage explainer.
A related subcampaign, not a date or outcome for every event in the broader campaign.
OpenAI’s initial public account and subsequent updates about the evaluation incident.
An evolving disclosure; later investigations qualify parts of the initial explanation.
Google examines suspected AI-assisted exploit development, malware that consults AI, and attacks through software used by AI systems. The report also describes threat actors experimenting with agents for security testing.
Combines investigations, model activity and software analysis. Some findings concern plans or inferred AI use, rather than successful attacks. Report date does not date every example.
Background reading. No individual case from this report is currently linked in the incident library.
An update covering late-2025 activity connects AI use with phishing and target research, examines attempts to copy model capabilities, and describes experimental malware. It also documents deceptive instructions hosted through public AI-chat sharing features.
Primarily Q4 2025 observations, with some earlier examples. Model extraction is a different risk from stealing users’ data. Experiments and operational campaigns require separate treatment.
Background reading. No individual case from this report is currently linked in the incident library.
Date Bait combines human operators and AI chatbots in a romance-and-task scam. False Witness impersonates lawyers and authorities to target previous fraud victims. Both illustrate AI-assisted deception; the report also covers separate influence operations.
Included for online fraud and impersonation, not as proof of technical intrusion. Claimed victim losses and scale drawn from scammer inputs were not independently verified.
Background reading. No individual case from this report is currently linked in the incident library.
Anthropic’s account of an AI-assisted espionage operation.
Read the provider’s autonomy claims alongside its stated human role and limitations.
Google describes malware calling AI while it runs, alongside continued use of AI for coding, research and deception. Its examples distinguish tools seen in operations from prototypes still being tested.
Broader than Gemini activity alone. Malware families are not incident counts; experimental capabilities and advertised services do not establish successful victim compromises.
Background reading. No individual case from this report is currently linked in the incident library.
Three cyber investigations examine malware development and tailored phishing by Russian-, Korean- and Chinese-language operators. The report describes AI assisting existing workflows and preserves limits on attribution, off-platform visibility and claims of new attacker capability.
Language alone is not attribution. Overlapping indicators connect some activity to other investigations, but do not prove every related malware sample was generated using AI.
Background reading. No individual case from this report is currently linked in the incident library.
Misuse case studies, including the GTG-2002 data theft and extortion campaign.
Findings and attribution are the provider’s assessment. The PDF covers additional categories.
ScopeCreep used AI while developing malware disguised as a gaming utility. Other cases describe China-linked actors using models for research and technical support, alongside employment fraud, social engineering and scams.
ScopeCreep was likely active, but widespread distribution was not established. Threat-actor experiments with automation do not demonstrate successful autonomous intrusion.
Background reading. No individual case from this report is currently linked in the incident library.
Anthropic describes attempted use of exposed camera passwords, recruitment scams and malware development by a novice. A separate case examines coordinated influence activity. Accounts were banned, while successful deployment of the cyber and fraud examples remained unconfirmed.
Published April 23 despite the March edition label. The linked PDF covers only the influence case; the cyber and fraud case studies are in the HTML.
Background reading. No individual case from this report is currently linked in the incident library.
Cases include suspected North Korean actors researching intrusion tools, deceptive employment, and online scams. OpenAI also describes sharing malware-related indicators discovered in model conversations so other defenders could detect the files.
Selected investigations combine model activity and outside evidence. Attribution is qualified; the cyber section reports no novel capability from the model responses.
Background reading. No individual case from this report is currently linked in the incident library.
An early baseline of tracked government-backed groups using Gemini for research, coding and persuasive content. Google found productivity benefits, but no novel attack capabilities in the activity it analyzed.
Analysis centers on Gemini web-app use by tracked groups. Prompts show attempted assistance, not necessarily deployment, success or the wider prevalence of AI attacks.
Background reading. No individual case from this report is currently linked in the incident library.
Three cyber case studies cover blocked phishing aimed at OpenAI employees, research into industrial control systems, and Android malware development. They show AI supporting established attacker tasks alongside a separate set of influence operations.
SweetSpecter, CyberAv3ngers and STORM-0817 have separate evidence and outcomes. The report does not establish that AI caused earlier infrastructure attacks attributed to those groups.
Background reading. No individual case from this report is currently linked in the incident library.
Early observations of state-affiliated groups using language models.
Provider observations of assistance do not establish successful autonomous attacks.
OpenAI’s account of disrupting state-affiliated actors’ use of its services.
Provider telemetry and account action; not independent proof that broader operations ended.
Try a broader topic or show the full catalog.
Read the source, then check what it establishes. A provider’s disclosure, an affected organization’s account, and an independent review offer different perspectives. Inclusion is not endorsement or independent confirmation.
Download the catalogNIST, OWASP, NCSC, and CISA guidance is collected with the AI security framework. Those sources explain practices for defenders; they are distinct from incident reports.