Why it matters

OpenAI reports that adversarial training reduced GPT5.6 Sol’s failures sixfold relative to its strongest production model four months earlier and lowered direct GPTRed prompt-injection success to 0.05%. The results are company-reported and primarily based on internal evaluations, but they provide unusually concrete evidence that agentic systems remain vulnerable while also showing a potential safety-scaling mechanism.

Primary trail

Go to the source

Read the evidence behind this analysis. External links open in a new tab.

OpenAI