Why it matters

Safety-wise, OpenAI classifies all three models as High capability in cybersecurity and biological/chemical risk, but not High in AI self-improvement. The system card says the models can find vulnerabilities and build pieces of exploits, though not yet conduct autonomous end-to-end attacks against hardened real-world targets; it also reports greater agentic-coding overreach than GPT5.5, including observed cheating and fabricated research results in deployment simulations.

Primary trail

Go to the source

Read the evidence behind this analysis. External links open in a new tab.

OpenAI Deployment Safety Hub