Analysis frame
Peer-reviewed research
Separate model discrimination in retrospective datasets from clinical utility, transportability, calibration, and the authority a decision-support tool receives in real care.
- Patients with advanced non-small-cell lung cancer considering immunotherapy
- Oncologists using prognostic and treatment-response evidence
- Hospitals responsible for local validation and workflow integration
- Researchers assembling multimodal clinical datasets across institutions
- Whether the physician gains persist in prospective clinical workflow
- How well the model calibrates across additional hospitals and populations
- Whether multimodal data add reliable value when collected consistently
- How explanations affect automation bias, trust, and treatment decisions
- Routine clinical and blood data could make decision support more accessible than data-intensive multimodal systems
- Site-specific performance gaps may widen inequality if hospitals cannot validate locally
- Explanations may improve useful uptake while also increasing confidence in incorrect predictions
- Prospective trials could establish a stronger standard for clinical AI procurement
The study tested both a model and its use
The retrospective cohort combined clinical and blood data for 2,396 patients, with smaller subsets containing imaging, pathology, and genomic information. Routine-data models reached test AUC values up to 0.77.
Twenty oncologists then assessed one hundred cases before and after receiving model predictions and explanations. Sensitivity for disease-control prediction rose from 0.72 to 0.87, with improvements in several other measures.
External validation changed the story
Performance ranged from 0.55 to 0.72 in external validation, and the authors point to population differences as one likely reason. Added modalities did not produce a consistently reliable benefit across evaluation cohorts.
Those results limit claims of transportability. A model can help in one assembled dataset and still require recalibration, workflow testing, and fairness review at the next hospital.
The next evidence is prospective
The authors describe silent prospective validation in more than two thousand patients, a second usability study, and a planned pragmatic randomized trial. That sequence is designed to test what retrospective metrics cannot: performance under real data flow and clinical decisions.
The appropriate role today is explainable support under physician responsibility. Broader authority should follow evidence that the tool remains useful across sites, populations, and actual care.
Go to the source
Read the evidence behind this analysis. External links open in a new tab.
Nature Medicine — Clinical usability of explainable AI in lung cancer


