Medical imaging foundation models are changing how AI interprets scans, integrates clinical information, and supports multiple tasks from a shared computational base
The approach uses large-scale pre-training to encode reusable knowledge into generalisable representations that can be adapted for classification, segmentation, detection, report generation, medical question answering, and prognosis.
A new review maps the field’s technical routes and emerging applications while warning that impressive benchmark scores do not yet prove clinical benefit and argues that progress should be judged by generalisability, reliability, workflow value, safety, and accountability, not just by model size or the number of tasks covered.
AI is already used for lesion detection, disease classification, and outcome prediction in medical imaging; however, most systems remain narrowly designed for one task and depend on expert-labelled data. Their performance can deteriorate when scanners, acquisition protocols, patient populations, or institutions differ from the training environment.
Foundation models promise greater reuse through pre-training on large image collections, image-report pairs, clinical text, structured records, and videos. But current evidence still comes largely from public benchmarks, retrospective cohorts, and controlled settings, creating uncertainty about real-world robustness, fairness, and clinical impact. These challenges mean deeper research is required in representative data, prospective validation, workflow integration, and continuous governance.
Researchers from the Institute of Automation, Chinese Academy of Sciences, and the School of Engineering Medicine at Beihang University published the review in July 2026 in the Medical Journal of Peking Union Medical College Hospital.
The authors analyse how medical imaging foundation models are built, adapted, evaluated, and moved toward clinical use. They also assess progress across radiology, digital pathology, ultrasound, and surgical video, while highlighting that a central question is not whether these models can perform many tasks, but whether they can provide stable, verifiable value in real-world clinical settings.
The review splits the field into four overlapping development paths: image-representation pre-training, image-language alignment, integration of multiple clinical data sources, and modelling of dynamic visual sequences.
Training data can include computed tomography (CT), magnetic resonance imaging (MRI), X-rays, ultrasound, pathology slides, reports, laboratory measurements, treatment records, endoscopy, and surgical video.
The authors outline that data volume alone can be misleading, as millions of image patches or video frames may not correspond to an equivalent number of independent patients. Quality control, deduplication, patient-level independence, cross-centre coverage, and reliable pairing between images and text are therefore critical considerations.
The authors also describe adaptation strategies such as lightweight task heads, fine-tuning, prompt learning, and instruction tuning. Single-modality examples include cancer subtyping, mutation prediction, survival estimation, and lesion segmentation, while vision-language systems add retrieval and question answering.
Critically, evaluation should extend beyond accuracy. The review suggests examining algorithmic robustness under external data and input disturbances, clinical usefulness through comparisons of clinician-only and clinician-model performance, and workflow outcomes such as reporting time, triage efficiency, repeat examinations, resource use, and patient outcomes. It also emphasises that a foundation model can underperform a task-specific system in a clearly defined clinical setting.
The review argues that the field’s next milestone should not be another increase in parameter counts, but evidence that a model can work safely in a defined clinical role. It suggests a more feasible path: using foundation models to assist with limited, reviewable tasks, such as triage, report drafting, interactive segmentation, risk stratification, or structured follow-up, instead of attempting to replace an entire diagnostic process.
The review also calls for clear indications, prohibited uses, input-quality requirements, uncertainty signals, human review responsibilities, failure reporting, version tracking, and revalidation after updates.
For clinical deployment, the review sets out a practical path for turning broad technical capability into accountable clinical support. Models should connect reliably with picture archiving and communication systems (PACS), radiology information systems (RIS), and hospital information systems (HIS), while also preserving logs of inputs, outputs, clinician edits, warnings, and software versions.
Prospective studies should test if deployment improves decisions, efficiency, or patient outcomes across different centres and patient groups. Governance should also address privacy, consent, secondary data use, copyright, demographic bias, performance drift, and responsibility errors.
The broader implication is that clinical translation will depend on a layered partnership among general foundation models, speciality-specific systems, and human oversight, with every component serving a clearly bounded and auditable role.