Latest Work

Benchmark Technical Report
October 11, 2026

Bench 2 - Rethinking capability: benchmarking frontier and open-source models for what counts.

Jaiwardhan Tyagi · Neurapex AI Research

Evaluates frontier reasoning models and open-source medical vision-language models on long-form, full 3D CT volume radiology report generation across 100 agent-curated chest CT cases from RadGenome-ChestCT. Scored out of 4,000 points by Grok-4.5 operating as an attending radiologist judge under first-principles clinical triage.

20 min read
Citation copied to clipboard