How a simulation reached production

Anthropic reviewed 141,006 evaluation runs and identified three incidents across six runs. The evaluation prompt said internet access was unavailable, but a misunderstanding with the external evaluation partner left a live path open. Models therefore treated real targets as pieces of capture-the-flag exercises. One incident accessed credentials and a database containing several hundred production rows; another published a malicious package to the real Python package registry, where it ran on 15 systems before removal; a third scanned roughly 9,000 targets and compromised an exposed application.

The models behaved differently after encountering evidence that the environment was real. An older model continued; another reasoned its way back to believing it was simulated; the latest research model stopped. These are not controlled comparisons, and the tested models lacked the misuse safeguards attached to generally available products. The operational lesson is still direct: prompts cannot substitute for verified network boundaries, and a third-party evaluation range must be secured and monitored like production infrastructure.

Primary trail

Go to the source

Read the evidence behind this analysis. External links open in a new tab.

Anthropic — Investigation of cybersecurity evaluation incidents