Can a Generalist AI Model Achieve Expert-Level Diagnosis for Complex Abdominal CT Scans?

Author: 

Luiz Perandini

Reading time / 
5 min
AI Medical Imaging

TL;DR

  • RADAR is a generalist vision-language AI model for abdominal CT, published in Science in September 2026.
  • It learned largely from routine radiology reports across 424,911 abdominal CT examinations.
  • It reached a mean AUC of 0.913 across 146 findings internally and 0.895 across eight external hospitals.
  • With RADAR assistance, radiologists' sensitivity increased by 10.0% and reading time per case fell by 30.7%.
  • Broader geographic validation and prospective clinical evaluation are still needed.

‍


Introduction

Most artificial intelligence models in medical imaging are developed using supervised learning with high-quality annotated data. This approach achieves excellent and sometimes expert-level performance for specific diagnostic tasks.

However, it has an important limitation. Every new task typically requires large amounts of manually annotated imaging data. The algorithm is also usually developed to identify one specific abnormality.

Creating these datasets is expensive and time-consuming. This makes it difficult to scale AI across the many different abnormalities encountered in clinical practice.

Generalist AI models using vision-language learning offer another approach. They allow models to learn directly from the relationship between medical images and the text already generated during routine care.


Why is abdominal CT a hard problem for AI?

Abdominal CT presents a particularly challenging setting for this approach. Contextual information is highly complex. Diagnostic signals are also extremely sparse relative to the entire three-dimensional (3D) volume.


What is RADAR and how was it trained?

Zhang et al. developed a generalist vision-language AI model to address this problem. It is called RADAR (Rapid Abdominal Diagnosis with AI and Radiology). The study was published in Science in September 2026.

The researchers assembled a dataset of 424,911 abdominal CT examinations. This produced approximately 1.5 million volume-level image-text pairs and more than 15 million anatomy-level image-text pairs.

RADAR did not require manual disease labels for every CT examination. Instead, it learned largely from the radiology reports already generated during routine clinical care. The authors also established a benchmark covering 18 anatomical structures and 146 findings.

‍

‍

How well did RADAR perform across internal and external cohorts?

RADAR was first tested in an internal real-world cohort of 39,160 examinations. It achieved a mean area under the receiver operating characteristic curve (AUC) of 0.913 across 146 findings.

The model was subsequently tested on 24,239 examinations from eight independent hospitals. There, it achieved an average AUC of 0.895. Individual-center AUCs ranged from 0.874 to 0.912. The authors also tested the model in several more challenging settings. In an emergency cohort excluded from the original training dataset, RADAR achieved a mean AUC of 0.904 across 17 acute abdominal conditions.

RADAR was also evaluated in two independent cohorts containing 4,333 patients. These cohorts included patients with pathology-confirmed colorectal, pancreatic, gastric, or hepatocellular cancers, as well as controls. AUCs ranged from 0.891 to 0.984, depending on cancer type and center.

‍

‍Did RADAR help radiologists in a reader study?

Perhaps the most clinically relevant part of the study was the reader experiment. The authors recruited 26 radiologists from 14 institutions. They interpreted 300 abdominal CT examinations covering 61 different findings.

Radiologists first reviewed the cases without AI assistance. After a washout period of at least one month, they reviewed the cases again with RADAR. The model provided predicted findings and attention maps highlighting suspicious regions.

With RADAR assistance, diagnostic sensitivity increased by 10.0%. Specificity decreased slightly, from 98.8% to 98.2%. Reading time per case was also reduced by 30.7%.

Improvements were observed across radiologists with different levels of experience. They were particularly relevant among junior readers.

‍

‍

What are the limitations of the RADAR study?

Despite these encouraging results, important limitations remain. Most of the large-scale training and validation data originated from a single country in Asia.

RADAR showed promising cross-population performance on the Stanford Merlin dataset. However, the authors acknowledge that broader validation across different geographic regions and patient populations is still needed.

Most evaluation labels were also derived from radiology reports rather than independent clinical or pathological reference standards. This was partially addressed through the pathology-confirmed cancer cohorts.

Moreover, prospective evaluation within routine clinical workflows will still be necessary. It is needed to understand how well these improvements translate into everyday practice.

‍

What does RADAR mean for medical imaging AI?

Overall, this study illustrates an important shift in medical imaging AI. Instead of building a separate model and labeled dataset for every diagnostic task, RADAR takes a different approach.

It demonstrates how a single vision-language system can learn from large-scale imaging and routine radiology reports. That system can then support diagnosis across more than one hundred findings.

‍

How Segmed supports imaging AI development

The RADAR authors note that broader validation across geographic regions and patient populations is still needed. Diverse, real-world imaging data is central to that kind of work. Segmed provides de-identified, real-world multimodal imaging datasets for clinical AI development, life sciences research, and regulatory applications.

Talk to our team about real-world imaging data for your AI program.

‍


‍

Frequently Asked Questions – F.A.Q.

‍

What is RADAR in medical imaging AI?

RADAR stands for Rapid Abdominal Diagnosis with AI and Radiology. It is a generalist vision-language model for abdominal CT diagnosis. Zhang et al. published it in Science in September 2026.

‍

How was RADAR trained?

RADAR learned largely from radiology reports generated during routine clinical care. Its dataset included 424,911 abdominal CT examinations. This produced about 1.5 million volume-level and more than 15 million anatomy-level image-text pairs.

‍

How accurate is RADAR on abdominal CT?

In an internal cohort of 39,160 examinations, RADAR reached a mean AUC of 0.913 across 146 findings. Across eight independent hospitals, it reached an average AUC of 0.895.

‍

Does RADAR improve radiologist performance?

In a reader study with 26 radiologists, RADAR assistance increased sensitivity by 10.0%. Specificity decreased slightly, from 98.8% to 98.2%. Reading time per case fell by 30.7%.

‍

What are the main limitations of RADAR?

Most training and validation data came from a single country in Asia. Most evaluation labels were derived from radiology reports. Prospective evaluation in routine clinical workflows is still needed.

‍

What is a generalist vision-language model in radiology?

It is a model that learns from the relationship between medical images and the text generated during routine care. It does not need manually annotated data for every new diagnostic task.

‍


‍

References

1. Zhang Q, Zhang J, Cao W, et al. An expert-level generalist AI for abdominal CT diagnosis. Science. 2026 Sep 17;393(6817):eaec6129.

‍


‍

Related Resources

Role of Real-World Imaging Data in Fine-Tuning Healthcare Foundation Models

Bias, Equity & Data Diversity: Why Real-World Imaging Data is Essential for Ethical AI in Healthcare

Vision Language Foundation Model for Chest X-Ray Generation

‍