
.png)
TL;DR
Most artificial intelligence models in medical imaging are developed using supervised learning with high-quality annotated data. This approach achieves excellent and sometimes expert-level performance for specific diagnostic tasks.
However, it has an important limitation. Every new task typically requires large amounts of manually annotated imaging data. The algorithm is also usually developed to identify one specific abnormality.
Creating these datasets is expensive and time-consuming. This makes it difficult to scale AI across the many different abnormalities encountered in clinical practice.
Generalist AI models using vision-language learning offer another approach. They allow models to learn directly from the relationship between medical images and the text already generated during routine care.
Abdominal CT presents a particularly challenging setting for this approach. Contextual information is highly complex. Diagnostic signals are also extremely sparse relative to the entire three-dimensional (3D) volume.
Zhang et al. developed a generalist vision-language AI model to address this problem. It is called RADAR (Rapid Abdominal Diagnosis with AI and Radiology). The study was published in Science in September 2026.
The researchers assembled a dataset of 424,911 abdominal CT examinations. This produced approximately 1.5 million volume-level image-text pairs and more than 15 million anatomy-level image-text pairs.
RADAR did not require manual disease labels for every CT examination. Instead, it learned largely from the radiology reports already generated during routine clinical care. The authors also established a benchmark covering 18 anatomical structures and 146 findings.

RADAR was first tested in an internal real-world cohort of 39,160 examinations. It achieved a mean area under the receiver operating characteristic curve (AUC) of 0.913 across 146 findings.
The model was subsequently tested on 24,239 examinations from eight independent hospitals. There, it achieved an average AUC of 0.895. Individual-center AUCs ranged from 0.874 to 0.912. The authors also tested the model in several more challenging settings. In an emergency cohort excluded from the original training dataset, RADAR achieved a mean AUC of 0.904 across 17 acute abdominal conditions.
RADAR was also evaluated in two independent cohorts containing 4,333 patients. These cohorts included patients with pathology-confirmed colorectal, pancreatic, gastric, or hepatocellular cancers, as well as controls. AUCs ranged from 0.891 to 0.984, depending on cancer type and center.

Perhaps the most clinically relevant part of the study was the reader experiment. The authors recruited 26 radiologists from 14 institutions. They interpreted 300 abdominal CT examinations covering 61 different findings.
Radiologists first reviewed the cases without AI assistance. After a washout period of at least one month, they reviewed the cases again with RADAR. The model provided predicted findings and attention maps highlighting suspicious regions.
With RADAR assistance, diagnostic sensitivity increased by 10.0%. Specificity decreased slightly, from 98.8% to 98.2%. Reading time per case was also reduced by 30.7%.
Improvements were observed across radiologists with different levels of experience. They were particularly relevant among junior readers.

Despite these encouraging results, important limitations remain. Most of the large-scale training and validation data originated from a single country in Asia.
RADAR showed promising cross-population performance on the Stanford Merlin dataset. However, the authors acknowledge that broader validation across different geographic regions and patient populations is still needed.
Most evaluation labels were also derived from radiology reports rather than independent clinical or pathological reference standards. This was partially addressed through the pathology-confirmed cancer cohorts.
Moreover, prospective evaluation within routine clinical workflows will still be necessary. It is needed to understand how well these improvements translate into everyday practice.
Overall, this study illustrates an important shift in medical imaging AI. Instead of building a separate model and labeled dataset for every diagnostic task, RADAR takes a different approach.
It demonstrates how a single vision-language system can learn from large-scale imaging and routine radiology reports. That system can then support diagnosis across more than one hundred findings.
The RADAR authors note that broader validation across geographic regions and patient populations is still needed. Diverse, real-world imaging data is central to that kind of work. Segmed provides de-identified, real-world multimodal imaging datasets for clinical AI development, life sciences research, and regulatory applications.
Talk to our team about real-world imaging data for your AI program.
RADAR stands for Rapid Abdominal Diagnosis with AI and Radiology. It is a generalist vision-language model for abdominal CT diagnosis. Zhang et al. published it in Science in September 2026.
RADAR learned largely from radiology reports generated during routine clinical care. Its dataset included 424,911 abdominal CT examinations. This produced about 1.5 million volume-level and more than 15 million anatomy-level image-text pairs.
In an internal cohort of 39,160 examinations, RADAR reached a mean AUC of 0.913 across 146 findings. Across eight independent hospitals, it reached an average AUC of 0.895.
In a reader study with 26 radiologists, RADAR assistance increased sensitivity by 10.0%. Specificity decreased slightly, from 98.8% to 98.2%. Reading time per case fell by 30.7%.
Most training and validation data came from a single country in Asia. Most evaluation labels were derived from radiology reports. Prospective evaluation in routine clinical workflows is still needed.
It is a model that learns from the relationship between medical images and the text generated during routine care. It does not need manually annotated data for every new diagnostic task.
Role of Real-World Imaging Data in Fine-Tuning Healthcare Foundation Models
Bias, Equity & Data Diversity: Why Real-World Imaging Data is Essential for Ethical AI in Healthcare
Vision Language Foundation Model for Chest X-Ray Generation