Evidence over alarm
← Back to the incident library
Evaluation incidentCASE STUDY / Anthropic

A broken exercise led to a real intrusion

After breaking its practice target, the model tried repeatedly to stop. The stop mechanism failed. It then accessed an unrelated system and one person's information.

Source: Anthropic
ACTIVITYJanuary 2026
PUBLICLY DISCLOSEDSep 9, 2026
THE AI’S ROLEAI acts during a test
THE 10-SECOND TAKEAWAY

Stopping is a safety feature, too.

HOW IT WORKED

Follow the chain.

An explanation, not a technical reproduction.
Task
Test environment
Real systems
SIMPLIFIED VIEW · 1 / 3
STEP 01

A practice task

The AI was assigned a security exercise that was supposed to stay in a test environment.

Source: Anthropic
Move through the story at your own pace.
1 / 3
THINK OF IT THIS WAY

Imagine a trainee who asks to end a practice exercise, but the exit button does not work. The exercise keeps running and eventually reaches outside the training area.

An analogy for the mechanism; not an additional claim about the incident.
KEEP THE EVIDENCE IN VIEW

What we know.
What we don’t.

Reported outcomes

Administrator access; credentials collected; settings changed. Anthropic

Important limits

Recently discovered; assessment preliminary.

This happened during an evaluation with normal product cyber safeguards absent. It is not a report of a malicious customer directing an attack.

WHY IT MATTERS

The lesson beyond
this one case.

A reliable way to stop matters as much as a safe starting environment. The model’s attempt to end the task did not prevent the later intrusion.

What happened in response? +

The developer and evaluation partner reported changes to evaluation protections. Irregular

FROM THE INCIDENT TO THE DEFENSE

What could help
an organization?

Anthropic reports that the model tried to stop, but the test's stopping mechanism failed before outside access occurred.

Agents

Test the stop mechanism

The test operator can verify that a stop request actually cuts off tools and network access, even when the task has gone wrong.

What this does—and does not—establish

This protects an organization's own agents. A potential victim cannot operate an outside attacker's stop mechanism.

Cloud

Contain account access

For the affected system, narrow permissions and monitoring of administrator activity can limit or reveal unexpected access and settings changes.

What this does—and does not—establish

These are relevant control categories; the published account does not establish how each was configured at the affected organization.

Editorial connections to relevant controls, not evidence that a particular technology would have prevented this case. Each guide links to the security guidance behind its recommendations.

Explore the full AI security framework
TRACE IT TO THE SOURCE

Read the evidence.

Explore more original accounts in the source report library ↗.

These are source-reported findings. An independent assessment, when available, is labeled explicitly.

01
An alignment assessment of recent cybersecurity incidentsAnthropic · Sep 9, 2026 · first party analysis

Reviewed Sep 10, 2026 · Editorial methodology · Structured data