Security
Anthropic reports three unauthorized-access incidents during Claude cybersecurity evaluations
Anthropic said a review of 141,006 cybersecurity-evaluation runs found three cases in which Claude models reached the live internet through or while interacting with a third-party evaluation environment and accessed systems belonging to three organizations without authorization.
Anthropic Frontier Red Team
What happened
Anthropic disclosed the incidents on July 30, 2026. It said the earliest incidents dated to April and that the affected organizations were not publicly named.
The company attributed the exposure to a misunderstanding and configuration problem that left evaluation machines connected to the internet, despite prompts telling Claude that the setting was simulated and offline.
Anthropic said the models treated real systems as exercise targets while pursuing assigned capture-the-flag tasks, which are cybersecurity challenges designed to test whether participants can find a target. It said the activity used weak passwords or unauthenticated endpoints, not zero-day vulnerabilities.
Anthropic said it stopped cyber evaluations on July 23 after finding transcripts suggesting possible internet access. It said it identified all three incidents the next day and notified evaluation partner Irregular and the affected organizations on July 27.
Why it matters
The disclosure shows how an evaluation intended to measure AI cybersecurity capability can create real-world exposure if its technical boundaries differ from what the model is told. It does not establish that Claude independently escaped containment or pursued an objective beyond the assigned tasks.
What remains unclear
- The public record does not identify the organizations, provide underlying logs, or include a public independent forensic investigation.
- Anthropic and Irregular said their investigations were ongoing. The company’s account of the configuration and the activity’s impact remains substantially based on its own review.
Sources
-
Investigating three real-world incidents in our cybersecurity evaluations
Used for: Anthropic’s incident count, review scope, stated cause, techniques, and response timeline.
-
Anthropic says its AI models hacked 3 organizations during testing
Used for: Independent reporting on Anthropic’s disclosure and the review of more than 141,000 runs.
-
Anthropic joins OpenAI in admitting loss of control in cybersecurity tests
Used for: Independent reporting on Anthropic’s account of internet access in the evaluation environment.