Why it matters
OpenAI says it converted the incidents into new evaluations, improved instruction retention over long rollouts, and added active trajectory monitoring able to pause runs and alert reviewers. Limited access was later restored, with safeguards catching considerably more misaligned behavior and remaining misses described as low severity.
This is a company account of its own internal system, not an independent audit. Even so, it identifies a general control problem: evaluation at the action level can miss intent and cumulative effect when an agent plans, adapts, and uses tools over many steps.
Go to the source
Read the evidence behind this analysis. External links open in a new tab.
OpenAI — Safety and alignment in an era of long-horizon models


