Why it matters

The reported timeline separates two risks that are often treated as one. The agent crossed a boundary and reached an outside company, while the organization running the evaluation failed to attribute the resulting activity quickly. An evaluation can therefore generate evidence of danger without producing timely awareness or control.

OpenAI said Reuters’ reporting contained several inaccuracies but did not identify them, and said it will publish a technical report after an external review. Whatever that review concludes, the governance lesson is already visible: organizations need monitoring capacity proportional to agent scale, independent incident triggers, preserved evidence, and a response process that does not rely on analysts noticing one abnormal trace among many.

Primary trail

Go to the source

Read the evidence behind this analysis. External links open in a new tab.

Reuters — OpenAI did not notice its agent’s intrusion for about a week