Introduction
With the advent of early pioneering vision-language models, the industry underwent a
noticeable capability explosion. Large-scale contrastive learning paved the way for modern
medical vision-language models [16], [17],
establishing closed & open-ended visual-question
answering, zero-shot classification, report generation and 'generalism' as the new quality
standards that future medical VLMs need to excel at.
Years of stagnant, fragmented advancements in the field were finally undergoing a creative
destruction. This was our algorithmic breakthrough, our transformers-moment, and this was
the key to "AI in medical imaging".
As of October 2026, recounting on the journey, the biggest recent open-source generalist
generative contributions in medical VLMs were: MedGemma 1.5 4B [5], Lingshu family of models [6],
HuatuoGPT-Vision [7], GMAI-VL
[8], UniMedVL [9] and
Fleming-VL [10].
Frontier models are quickly climbing up the ladder and are increasingly performing better
than many of these focused VLMs. 2026 has been a year of aggressive frontier advancement,
with the arrival of Fable-5, Opus-5.5, GPT-5.6 Sol, GPT-6 Astra, and the official beginning
of the AGI era, we need to rethink what counts as capability.
Until 2022, the AI-in-medical-imaging market sat at roughly $1.4 billion [18],
[19]. Today, four years
later, despite a genuine capability explosion, the market has only expanded to low single
digit billions (~ $2.1 - 3.8 billion) [19], [20].
One would naturally raise questions about what caused it, and why is the adoption not
undergoing a surge. One could argue, that the reasons are undeniably the pace of
advancement, models still lacking the capability, and the regulatory hurdles.
We believe it's a mix of capability and regulations.
Capability
I'm critical about how majority of the researchers' think about capability. When we think
about capability of a medical AI imaging system or model, we always tend to think in
extremes. We either think about a case where the model is simply given no autonomy, and is
performing a highly low-level, low-leverage task that doesn't seriously shape the overall
workflow (such as simply drawing a bounding box on a nodule on a chest x-ray open for
interpretation), or we think about a god-tier system awarded full autonomy for diagnosing
patients, recommending follow-up steps, and completely removing the human-in-the-loop.
Until we officially reach the superintelligence era, there's a middle ground, a hybrid
approach we could be optimizing for as engineers and researchers to advance the field. The
recipe is simply: take the current capability, assess which model skill is worth advancing,
and organize everything around it: training, benchmarks, and applications, which could be,
in its executed form, a radical leap towards medical imaging AI adoption. This is a crucial
part of what we're trying to achieve at Neurapex AI.
Optimization so far leaned towards the low leverage accuracy metrics: short-form, closed
ended VQA, zero-shot classification, and overemphasis on single-image report generation.
Modern benchmarks have been centered around these skills, which should make us skeptical
about how we think best models top the benchmarks. While recent radiology benchmarks such as
RadLE v1 [3]
and
RadLE 2.0 [4]
have pioneered evaluating multimodal AI against clinical taxonomy errors, holistic 3D
volumetric generation remains an unaddressed challenge.
Bench 2
Bench 2 is a benchmark based on 100 agent-curated chest CT cases, full volume files, based
on RadGenome-ChestCT [2], a comprehensive,
large-scale, region-guided 3D chest CT interpretation
dataset based on CT-RATE [1],
containing preprocessed chest CT volumes, in NIfTI format.
Initial Setup
The foundation of this benchmark is based on evaluating the frontier reasoning models and
open-source generalist medical Vision-Language Models on long-form, full CT volume report
generation tasks for 3D computed tomography (CT) volumes presented as a series of
representative 2D slices and generate a clinically meaningful radiology report, structured
with a Findings and an Impressions section.
Constructing Ground Truth Reports
We utilized the regional reports of the validation set, curated 100 difficult, long-form
cases using Opus 5.5.
The reports were sentence level, regional and not yet full reports. Each volume's annotation
sentences were grouped by anatomy and passed to gpt-oss-20b [12] with temperature
0
and medium reasoning effort to synthesize the sentences into standard Findings and
Impressions, to make 100 full reports for each chest CT volume.
CT Volume Preprocessing
Each NIfTI volume is a 3D array, so the preprocessing pipeline uniformly samples a fixed
number of representative slices along the axial axis (the shortest spatial dimension of the
volume), distributing them evenly from apex to base. Each slice is intensity-normalized to
an 8-bit grayscale image. Every slice is labeled with its original position in the volume,
so the model receives spatial context.
Frontier models receive 16 slices per volume delivered inline.
Open-source models (limited by per-prompt image constraints in vLLM [13])
receive the same 16 slices but in 4 sequential groups of 4 slices covering entry, upper,
lower and exit regions of the scan.
Frontier Model Configurations
All frontier models receive the same system prompt: an expert radiologist identity
instructing them to examine the provided slices as a representation of the full CT volume
and produce a structured Findings + Impression report, seeded to increase reproducibility.
Output is
enforced as strict JSON.
| Model |
Provider |
Reasoning Effort |
Max Output Tokens |
Seed |
| Claude Opus 5 |
Anthropic |
— |
8096 |
— |
| Claude Fable 5.1 |
Anthropic |
— |
4096 |
— |
| GPT-6 Astra |
OpenAI |
High |
7200 |
42 |
| GPT-5.6 Sol |
OpenAI |
High |
10000 |
42 |
| Gemini 3.8 Flash |
Google |
High (thinking) |
— |
72 |
| Kimi K3 |
Moonshot AI |
High |
8096 |
— |
Open-source Configurations
Open-source models are evaluated on an A100 (80 GB) GPU.
| Model |
Architecture |
Serving |
Quantization |
Slices per Call |
| Lingshu-7B [6] |
Qwen2.5-VL [14] |
HF Transformers |
bfloat16 |
4 (×4 batches) |
| Lingshu-I-8B [6] |
InternVL2.5 [15]
|
vLLM [13] |
bfloat16 |
4 (×4 batches) |
| MedGemma-1.5-4B-IT [5] |
— |
Custom wrapper |
— |
16 inline |
| MedMO-8B-Next [11] |
Qwen3-VL |
HF Transformers |
bfloat16 |
Sanity test only |
The open-source models' prompt is a single instruction asking for an educational radiology
report based on the provided slices. Their 4-batch output segments (labeled Entry, Upper,
Lower, Exit) are concatenated into a single report before submission to the judge.
The Judge
Scoring is performed by Grok-4.5 (xAI), using medium reasoning effort. It
receives the model's generated report, and the ground-truth reference report. It is given no
other context.
The judge is prompted to act as a board-certified radiologist and to reason from first
principles before assigning a score. It works through a defined sequence: identify the
single most clinically important finding in the ground truth (along with its key
specificities, laterality, organ, location, size, morphology, severity); triage the
secondary findings by whether they would change clinical management; matching the candidate
report concept-by-concept against both; classifying any extra statements as benign
elaboration or fabrication; inspecting whether uncertainty was papered over with
hallucinated findings.
- • What the judge rewards: Accurate identification of the core
finding and its specificities. Capturing clinically material secondary findings.
Honest, calibrated hedging.
- • What the judge penalizes: Wrong laterality, wrong organ, wrong
size or severity on the core finding. Hallucinated findings absent from the ground
truth used to fill blind spots.
- • What the judge ignores: Report length, vocabulary, confidence,
phrasing style, section formatting, recommendations for further imaging or
specialist care, and elaboration on findings that are true.
For safety as a metric
Safety is contingent on a system’s degree of autonomy and the consequences of its
interpretations in practice, making it an ambiguous basis for evaluating model capability.
We therefore prioritize long-form accuracy, assessing the factual relevance and fidelity of
generated findings directly rather than conflating model performance with downstream
clinical risk.
Scoring Mechanism
For each case, the score the model gets rewarded will be an integer lying between 0 and 40.
For 100 cases, this gives a perfect score a ceiling of 4,000 points.
Each interpretation will be scored based on how accurately the model identifies the core
findings and its specificities, and will hence represent tiers for scores:
Exceptional (38–40), Strong (32–37), Solid
(24–31), Mixed (16–23), Poor (8–15), and
Fail (a catastrophic score lying between 0 and 7).
Results
Benchmark Score
1,419 / 4,000 (35.5%)
Mean (± SD)
14.19 ± 11.1
Median / 40
12.0
Case Range
0–39
4,000
Ceiling: 4,000 Pts (100%)
3,000
2,000
1,000
0
#1 GPT-6 Astra (OpenAI)
Score: 1,419 / 4,000 (35.5%)
Mean: 14.19 ± 11.1 · Median: 12.0
Range: 0–39
1,419
35.5%
#2 Gemini 3.8 Flash (Google)
Score: 889 / 4,000 (22.2%)
Mean: 8.89 ± 11.0 · Median: 3.0
Range: 0–39
889
22.2%
#2
Gemini 3.8Flash
Google
#3 Kimi K3 (Moonshot AI)
Score: 847 / 4,000 (21.2%)
Mean: 8.47 ± 9.1 · Median: 5.0
Range: 0–35
847
21.2%
#3
Kimi K3high
Moonshot AI
#4 GPT-5.6 Sol (OpenAI)
Score: 768 / 4,000 (19.2%)
Mean: 7.68 ± 7.5 · Median: 4.0
Range: 0–36
768
19.2%
#5 Claude Opus 5 (Anthropic)
Score: 738 / 4,000 (18.4%)
Mean: 7.38 ± 8.7 · Median: 4.0
Range: 0–36
738
18.4%
#5
ClaudeOpus 5
Anthropic
#6 MedGemma 4B (Open Health)
Score: 717 / 4,000 (17.9%)
Mean: 7.17 ± 11.2 · Median: 2.5
Range: 0–36
717
17.9%
#6
MedGemma4B
Open Health
#7 Fable 5.1 (Anthropic)
Score: 685 / 4,000 (17.1%)
Mean: 6.85 ± 7.7 · Median: 4.0
Range: 0–37
685
17.1%
#7
ClaudeFable 5.1
Anthropic
#8 Lingshu 7B (Open-source)
Score: 370 / 4,000 (9.2%)
Mean: 3.70 ± 6.4 · Median: 2.0
Range: 0–36
370
9.2%
#8
Lingshu 7B
Open Source
#9 Lingshu 8B InternVL (Open-source)
Score: 335 / 4,000 (8.4%)
Mean: 3.35 ± 4.9 · Median: 2.0
Range: 0–35
335
8.4%
#9
Lingshu 8BInternVL
Open Source
| Rank |
Model |
Total Score |
Mean / 40 |
Median |
Range |
% of Max |
| #1 |
GPT-6 Astra |
1,419 / 4,000 |
14.19 ± 11.1 |
12.0 |
0–39 |
35.5% |
| #2 |
Gemini 3.8 Flash |
889 / 4,000 |
8.89 ± 11.0 |
3.0 |
0–39 |
22.2% |
| #3 |
Kimi K3 |
847 / 4,000 |
8.47 ± 9.1 |
5.0 |
0–35 |
21.2% |
| #4 |
GPT-5.6 Sol |
768 / 4,000 |
7.68 ± 7.5 |
4.0 |
0–36 |
19.2% |
| #5 |
Claude Opus 5 |
738 / 4,000 |
7.38 ± 8.7 |
4.0 |
0–36 |
18.4% |
| #6 |
MedGemma 4B |
717 / 4,000 |
7.17 ± 11.2 |
2.5 |
0–36 |
17.9% |
| #7 |
Fable 5.1 |
685 / 4,000 |
6.85 ± 7.7 |
4.0 |
0–37 |
17.1% |
| #8 |
Lingshu 7B |
370 / 4,000 |
3.70 ± 6.4 |
2.0 |
0–36 |
9.2% |
| #9 |
Lingshu 8B InternVL |
335 / 4,000 |
3.35 ± 4.9 |
2.0 |
0–35 |
8.4% |
Exceptional (38–40)
Strong (32–37)
Solid (24–31)
Mixed (16–23)
Poor (8–15)
Fail (0–7 / <8)
Hover any
segment
— Inspect
diagnostic fidelity breakdown
100 CT Studies Breakdown
Exceptional (38–40): 1 case (1%)
10
Strong (32–37): 10 cases (10%)
14
Solid (24–31): 14 cases (14%)
17
Mixed (16–23): 17 cases (17%)
19
Poor (8–15): 19 cases (19%)
39
Fail (<8): 39 cases (39%)
Exceptional (38–40): 3 cases (3%)
6
Strong (32–37): 6 cases (6%)
6
Solid (24–31): 6 cases (6%)
5
Mixed (16–23): 5 cases (5%)
10
Poor (8–15): 10 cases (10%)
70
Fail (<8): 70 cases (70%)
7
Strong (32–37): 7 cases (7%)
4
Solid (24–31): 4 cases (4%)
6
Mixed (16–23): 6 cases (6%)
20
Poor (8–15): 20 cases (20%)
63
Fail (<8): 63 cases (63%)
Strong (32–37): 1 case (1%)
4
Solid (24–31): 4 cases (4%)
11
Mixed (16–23): 11 cases (11%)
22
Poor (8–15): 22 cases (22%)
62
Fail (<8): 62 cases (62%)
5
Strong (32–37): 5 cases (5%)
5
Solid (24–31): 5 cases (5%)
Mixed (16–23): 3 cases (3%)
12
Poor (8–15): 12 cases (12%)
75
Fail (<8): 75 cases (75%)
13
Strong (32–37): 13 cases (13%)
Mixed (16–23): 2 cases (2%)
Poor (8–15): 2 cases (2%)
83
Fail (<8): 83 cases (83%)
4
Strong (32–37): 4 cases (4%)
Solid (24–31): 1 case (1%)
7
Mixed (16–23): 7 cases (7%)
15
Poor (8–15): 15 cases (15%)
73
Fail (<8): 73 cases (73%)
Strong (32–37): 3 cases (3%)
Solid (24–31): 1 case (1%)
Mixed (16–23): 1 case (1%)
95
Fail (<8): 95 cases (95%)
Strong (32–37): 2 cases (2%)
Mixed (16–23): 1 case (1%)
96
Fail (<8): 96 cases (96%)
| Model |
Exceptional (38–40) |
Strong (32–37) |
Solid (24–31) |
Mixed (16–23) |
Poor (8–15) |
Fail (<8) |
| GPT-6 Astra |
1 |
10 |
14 |
17 |
19 |
39 |
| Gemini 3.8 Flash |
3 |
6 |
6 |
5 |
10 |
70 |
| Kimi K3 |
0 |
7 |
4 |
6 |
20 |
63 |
| GPT-5.6 Sol |
0 |
1 |
4 |
11 |
22 |
62 |
| Claude Opus 5 |
0 |
5 |
5 |
3 |
12 |
75 |
| MedGemma 4B |
0 |
13 |
0 |
2 |
2 |
83 |
| Fable 5.1 |
0 |
4 |
1 |
7 |
15 |
73 |
| Lingshu 7B |
0 |
3 |
1 |
1 |
0 |
95 |
| Lingshu 8B InternVL |
0 |
2 |
0 |
1 |
1 |
96 |
Discussion & Takeaways
GPT 6 Astra clearly tops the benchmark with a 59.6% increase in performance from the
runner-up, Gemini 3.8 Flash.
The models were given only 16 slices of the total anatomical context to evaluate how
frequently they produce hallucinated findings, for which they were heavily penalized by the
judge.
The failure rate (<8 points) across open-source architectures hovered between 83% and
96%, illustrating that while smaller vision-language models demonstrate impressive
conversational fluency on single-image benchmarks, synthesizing coherent 3D multi-slice
volumetric observations into standard clinical reports requires deep spatial reasoning and
severe hallucination suppression.
[1]
Generalist foundation models from a multimodal dataset
for 3D computed tomography.
Hamamci, I. E., Er, S., Wang, C., Almas, F., Simsek, A. G., Esirgun, S. N.,
et al. (CT-RATE).
Nature Biomedical Engineering, 2026. DOI:
10.1038/s41551-025-01599-y.
[2]
RadGenome-Chest CT: A Grounded Vision-Language Dataset
for Chest CT Analysis.
Zhao, Z., Lei, J., Zhang, Y., Wang, Y., Xie, W.
arXiv preprint arXiv:2404.16754, April 2024.
[3]
Radiology's Last Exam (RadLE): Benchmarking Frontier
Multimodal AI Against Human Experts and a Taxonomy of Visual Reasoning
Errors in Radiology.
Datta, S., Buchireddygari, D., Kaza, L. V. C., Bhalke, M., Singh, K.,
Pandey, A., et al.
CRASH Lab, Koita Centre for Digital Health, Ashoka University.
arXiv preprint arXiv:2509.25559, September 2025.
[4]
Radiology's Last Exam (RadLE) 2.0: Are we ready for
Autonomous AI Diagnosis in Radiology?
Datta, S., Buchireddygari, D., Bhatti, H. B. S., CRASH Lab.
Technical report and benchmark leaderboard, Ashoka University, July 2026.
Technical Report & Leaderboard
[5]
MedGemma Technical Report.
Sellergren, A., et al.
arXiv preprint arXiv:2507.05201, July 2025.
[6]
Lingshu: A Generalist Foundation Model for Unified
Multimodal Medical Understanding and Reasoning.
LASA Team, Alibaba DAMO Academy.
arXiv preprint arXiv:2506.07044, June 2025.
[7]
HuatuoGPT-Vision, Towards Injecting Medical Visual
Knowledge into Multimodal LLMs at Scale.
Chen, J., et al.
Proceedings of the 2024 Conference on Empirical Methods in Natural
Language Processing (EMNLP 2024).
arXiv:2406.19280.
[8]
GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language
Model and A Comprehensive Multimodal Dataset Towards General Medical
AI.
Li, T., Su, Y., Li, W., Fu, B., Chen, Z., Huang, Z., et al.
AAAI Conference on Human Computation and Artificial Intelligence (AAAI
2025).
arXiv:2411.14522.
[9]
UniMedVL: Unifying Medical Multimodal Understanding and
Generation Through Observation-Knowledge-Analysis.
Ning, J., et al.
arXiv preprint arXiv:2510.15710, October 2025.
[10]
Fleming-VL: Towards Universal Medical Visual Reasoning
with Multimodal LLMs.
Shu, Y., Liu, C., Chen, R., Li, D., Dai, B.
arXiv preprint arXiv:2511.00916, November 2025.
[11]
MedMO: Grounding and Understanding Multimodal Large
Language Model for Medical Images.
Deria, A., Kumar, K., Dukre, A. M., Segal, E., Khan, S., Razzak, I.
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI).
arXiv preprint arXiv:2602.06965, February 2026.
[12]
gpt-oss-120b & gpt-oss-20b Model Card.
OpenAI.
arXiv preprint arXiv:2508.10925, August 2025.
[13]
Efficient Memory Management for Large Language Model
Serving with PagedAttention.
Serving Stack
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., et al.
(vLLM).
Proceedings of the 29th ACM Symposium on Operating Systems Principles
(SOSP 2023).
arXiv:2309.06180.
[14]
Qwen2.5-VL Technical Report.
Lingshu-7B Backbone
Qwen Team, Alibaba Cloud.
arXiv preprint arXiv:2502.13923, 2025.
[15]
Expanding Performance Boundaries of Open-Source
Multimodal Models with Model, Data, and Test-Time Scaling.
Lingshu-I-8B Backbone
Chen, Z., et al. (InternVL 2.5).
arXiv preprint arXiv:2412.05271, December 2024.
[16]
Learning Transferable Visual Models From Natural
Language Supervision.
CLIP
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., et
al.
Proceedings of the 38th International Conference on Machine Learning
(ICML 2021).
arXiv:2103.00020.
[17]
Contrastive Learning of Medical Visual Representations
from Paired Images and Text.
ConVIRT
Zhang, Y., Jiang, H., Miura, Y., Manning, C. D., Langlotz, C. P.
Stanford University.
arXiv preprint arXiv:2010.00747, October 2020.
[18]
Artificial Intelligence (AI) in Medical Imaging
Market.
Global Market Insights. Industry Analysis Report, USD 1.38 billion (2022
base year).
[19]
Medical Imaging AI Market Expected to Exceed $1.7B by
2027.
Signify Research. Market evaluation (USD 576 million in 2022, forecast USD
1.73 billion by 2027), reported by AuntMinnie.
[20]
AI in Medical Imaging Market by Component, Technology,
Application, and End User: Global Opportunity Analysis and Industry
Forecast, 2022–2032.
Allied Market Research. Market report (USD 1.9 billion in 2022 base
valuation).