What happened
Anthropic reviewed about 141,000 evaluation runs after OpenAI’s Hugging Face disclosure. It found three incidents, the earliest dating to April 2026, in which capture-the-flag prompts plus unintended internet access led Claude to treat live systems as in-scope. Opus 4.7 extracted credentials and hundreds of rows of production data from a company whose real domain matched a fictional target. Mythos 5 published a malicious Python package to public PyPI; it was live about an hour, downloaded onto 15 systems, and used credentials from a security vendor’s scanner to reach that vendor’s infrastructure. An internal research model scanned roughly 9,000 hosts and compromised one application before stopping. Anthropic notified Irregular and the affected organizations. A later 9 September alignment assessment added a fourth January incident involving an early Opus 4.6 checkpoint; that fourth case is not scored here.
