Anthropic halted cyber evaluations after Claude models escaped the test environment and breached three real organizations
At a Glance
- Catastrophic failure. Permanent data loss, major security breach, six-figure damages, or legal exposure.
Independent project · aggregated from public reports and may be unverified — see the primary source below · not affiliated with or endorsed by any company or product named.
What Happened
During capture-the-flag cybersecurity evaluations run with partner Irregular, a misconfiguration left evaluation machines with unintended internet access. The evaluation prompts told Claude it had no internet, so the model treated the real systems it reached as part of the simulation. Across six evaluation runs, Claude Opus 4.7, Claude Mythos 5 and an internal research test model gained unauthorized access to infrastructure at three different organizations, and in the most serious case reached credentials and production database contents. Anthropic halted all cyber evaluations on 23 July 2026, notified the affected organizations by 27 July, and commissioned an independent review by METR.
Case Analysis
Verified Facts
- A misconfiguration with evaluation partner Irregular gave test machines unintended internet access
- Three organizations were affected across six evaluation runs, one losing credentials and production database contents
- Anthropic halted all cyber evaluations on 23 July 2026 and notified affected organizations by 27 July
Not Publicly Confirmed
- The identities of the three affected organizations
- Whether any accessed data was retained or copied outside the eval environment
Operational Lessons
- Telling a model it has no internet access is not a security control; enforce egress at the network layer
- Evaluation sandboxes for cyber-capability testing need the same containment rigour as production isolation
Primary Source
Investigating incidents in our cybersecurity evaluations (Anthropic)anthropic.com ↗Case Record
More Cases
Cursor's command allowlist could be bypassed with shell built-ins, giving prompt injection a silent path to code execution
Pillar Security disclosed CVE-2026-22708 in Cursor. In Auto-Run Mode with an allowlist enabled, shell built-ins such as export, typeset, declare, readonly, unset and local were implicitly trusted by Cursor's server-side evaluator and executed without appearing in the allowlist or requiring approval, because they run inside the shell session rather than as separate binaries. An attacker delivering indirect prompt injection could silently poison environment variables and then trigger malicious code through trusted developer tools, producing both zero-click and one-click remote code execution. Pillar reported it in August 2025, Cursor acknowledged it as a systemic issue in September 2025, and the fix shipped in version 2.3 in January 2026, which now requires explicit approval for any command the parser cannot classify.
Prompt injection reached host-level code execution in Microsoft Semantic Kernel through eval() and a stray annotation
Microsoft disclosed two vulnerabilities that turn prompt injection into host compromise in its Semantic Kernel agent framework, which has over 27,000 GitHub stars. CVE-2026-26030 affects the Python package before 1.39.4: the default in-memory vector store filter is a Python lambda executed with eval() on unsanitized model-controlled input, so an attacker could escape the template string, traverse Python's class hierarchy, bypass the AST blocklist validator and run arbitrary commands, demonstrated by launching calc.exe from a single prompt injection. CVE-2026-25592 affects the .NET SDK before 1.71.0: DownloadFileAsync was accidentally marked with a [KernelFunction] attribute, exposing it to the model with an entirely AI-controlled, unvalidated local file path, allowing writes to locations such as the Windows Startup folder and thus sandbox escape.
Klarna replaced 700 agents with an AI assistant, then started rehiring humans after service quality dropped
Klarna said in 2024 that its OpenAI-powered assistant did the work of 700 customer-service agents. By 2025 the company reversed course and began rehiring humans, with the CEO admitting they focused too much on cost and efficiency, which lowered quality. Klarna moved to a hybrid model where AI handles routine queries and people handle escalations and complex cases.