How we read the signal

Analysis frame

Evidence level

Mixed evidence

Analytical lens

Compare an independent product-level teen-account evaluation with a vendor model-level safety disclosure without treating their tests as directly comparable.

Affected groups
  • Teen users seeking help or help with schoolwork
  • Parents and caregivers relying on alerts
  • School counselors and mental-health professionals
What remains unknown
  • The independent test is not a population-level injury estimate
  • The October model card's system-level mitigations were not measured in the cited model-level score
  • The two assessments do not test identical versions or account conditions
Second-order effects to watch
  • An unreliable alert could give caregivers false confidence and delay human intervention
  • Overbroad refusal could reduce benign support, requiring a balanced test

What independent testers observed

The institute used more than 4,000 prompts on registered teen accounts and describes specific failures in notifications, crisis referrals and learning controls. Its rating is a judgment under its standards, not a population risk ratio.

The next test should publish account conditions and repeat results across new and established accounts, with independent verification of alerts.

What the vendor disclosed

OpenAI reports improvements in some areas and regressions in several under-18 categories on challenging model-level evaluations. It says a classifier and other product-level interventions add protection.

The model card does not rebut the earlier product test by itself, nor does the earlier product test establish October GPT-6 performance. A common end-to-end protocol could resolve the apparent conflict.

Primary trail

Go to the source

Read the evidence behind this analysis. External links open in a new tab.

Common Sense Media — ChatGPT for Teens risk findings Youth AI Safety Institute — full teen risk assessment OpenAI — October GPT-6 Sol and Luna safety update