How we read the signal

Analysis frame

Evidence level

Primary-source evidence

Analytical lens

The central measurement problem is separating access and frequency from instructional quality, student background, task choice, and the reasoning practices schools actually reward.

Affected groups
  • Students whose grades and future opportunities depend on how AI-assisted work is assessed
  • Teachers and school leaders redesigning instruction without causal evidence about the best AI-use pattern
  • Socio-economically disadvantaged students who report fewer opportunities to evaluate AI-generated information
What remains unknown
  • The observational PISA data cannot establish whether AI use caused higher or lower science performance
  • Self-reported frequency does not reveal the quality of prompts, feedback, verification, or teacher support
  • The long-term effect of AI-assisted learning on independent reasoning and knowledge retention remains unmeasured
Second-order effects to watch
  • Assessment may shift from grading finished answers toward oral defense, process evidence, and source verification
  • Paid tools and stronger school guidance could create a new skills gap even where basic chatbot access is widespread
  • Schools may mistake restrictive device policies for AI literacy and leave students unprepared to evaluate model output

AI use is already normal

PISA 2025 covered more than 760,000 students across 91 countries and economies. Across OECD countries, 45.5% reported using AI at least weekly to help them learn, while only 14% said they had almost never used AI for any of the schoolwork purposes examined.

The adoption story is therefore largely settled. The unresolved question is what kind of learning that use produces and which school practices make a difference.

The frequency curve is not a ladder

After adjusting for socio-economic background, weekly users had science performance similar to non-users. Very frequent users and occasional users tended to score lower, while moderate use for summarising and preliminary research sometimes outperformed both limited and frequent use.

PISA explicitly warns that these relationships do not establish causation. Students choose tools for different reasons, and prior achievement, task difficulty, access, teacher practice, and learning disposition can all shape both use and scores.

Critical evaluation is unevenly distributed

About 62.6% of students across OECD countries reported learning to assess AI-generated information in school lessons. Students who frequently used AI to learn and also reported such opportunities showed a slightly stronger performance pattern.

Disadvantaged students were less likely to report receiving that practice. If the ability to interrogate AI becomes economically valuable, unequal instructional quality may reproduce inequality even when basic access spreads.

What schools should measure next

AI policy in education needs outcome measures that separate access from mastery. Schools should ask whether students can identify an unsupported claim, recover the original source, explain where the model may have failed, and complete a comparable task without assistance.

The goal is neither blanket use nor blanket prohibition. It is visible, assessable judgment.

  • Track task purpose and verification behavior, not only frequency of use.
  • Use oral defense and process evidence where final text can be outsourced.
  • Give disadvantaged schools equal access to teacher training and AI-literacy instruction.
  • Separate short-term performance from long-term retention and independent reasoning.
Primary trail

Go to the source

Read the evidence behind this analysis. External links open in a new tab.

OECD — PISA 2025 Results, Volume I OECD — Executive summary and digital learning indicators OECD — Student school life, digitalisation and AI