Medical Evals
We benchmarked today's leading AI models on MedXpertQA, a dataset of expert-level medical exam questions spanning diagnosis, treatment, and basic science — including a multimodal subset with real clinical images. Every model ran zero-shot with chain-of-thought reasoning, no fine-tuning or few-shot examples.
Multimodal medical reasoning
2,000 questions with clinical images · 5 answer choices · Higher is better
Multimodal medical reasoning over time
MM accuracy vs. each model's release date · hover a dot or legend entry for detail · Higher is better
- gpt-6-astra
- Gemini 3.8 Flash
- Gemini 3.1 Pro Preview
- gpt-5.6-sol
- Claude Opus 5
- GPT-5
- gpt-5.6-terra
- gpt-5.6-luna (high)
- gpt-5.6-luna (medium)
- Claude Sonnet 5
- o1
- GPT-4o
- Gemini 2.0 Flash
- Gemini-1.5-Pro
- QVQ-72B-Preview
- Claude 3.5 Sonnet
- Qwen2.5-VL-72B
- Qwen2-VL-72B
- GPT-4o-mini
Text-only medical reasoning
2,450 questions · 10 answer choices · Higher is better
Not every model was run on both subsets. Scores are single runs at each provider's default settings (no reasoning-effort tuning), graded by exact-match on the model's final answer letter. GPT-4o, Gemini 2.0 Flash, and Claude 3.5 Sonnet scores are from the original MedXpertQA paper's published leaderboard rather than a run in this project — they're directly comparable because the paper used the same evaluation harness as every other score here.
Where models struggle
Breaking multimodal accuracy down by organ system and by task type shows the gap between models isn't uniform — some organ systems are hard for every model, while the ranking between models holds fairly steady across task types.
Accuracy by organ system
Multimodal subset · darker = higher accuracy
| Lymphatic | Nervous | Urinary | Muscular | Reproductive | Integumentary | Endocrine | Skeletal | Digestive | Other / NA | Cardiovascular | Respiratory | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| gpt-6-astra | 86 | 91 | 89 | 91 | 91 | 88 | 92 | 89 | 82 | 83 | 85 | 84 |
| Gemini 3.8 Flash | 94 | 92 | 91 | 91 | 89 | 88 | 88 | 87 | 86 | 83 | 81 | 81 |
| Gemini 3.1 Pro Preview | 90 | 84 | 82 | 89 | 85 | 83 | 82 | 82 | 80 | 67 | 77 | 81 |
| gpt-5.6-sol | 81 | 74 | 70 | 87 | 78 | 80 | 81 | 81 | 71 | 72 | 75 | 78 |
| Claude Opus 5 | 89 | 79 | 80 | 80 | 74 | 80 | 78 | 75 | 74 | 61 | 73 | 75 |
| gpt-5.6-terra | 75 | 71 | 66 | 78 | 73 | 79 | 74 | 74 | 67 | 72 | 67 | 71 |
| gpt-5.6-luna (high) | 83 | 69 | 63 | 78 | 68 | 71 | 68 | 69 | 64 | 67 | 67 | 69 |
| gpt-5.6-luna (medium) | 74 | 70 | 64 | 78 | 67 | 67 | 64 | 68 | 62 | 50 | 65 | 70 |
| Claude Sonnet 5 | 69 | 58 | 63 | 65 | 61 | 70 | 64 | 56 | 55 | 44 | 58 | 51 |
Accuracy by task type
Multimodal subset · Higher is better
Limitations
- The dataset itself has some construction issues. MedXpertQA expands each question's original answer choices up to 10 using an LLM-assisted process, which occasionally produces two choices describing the same underlying fact, a handful of internally contradictory questions, and — in the multimodal subset — the occasional low-resolution or mismatched image. In manual review, roughly 17–21% of the answers models got “wrong” had a real, defensible argument for the model's answer. True accuracy for every model here is likely a few points higher than the raw numbers shown.
- Single run, default settings. Each score is one pass per model at the provider's default reasoning effort, not an average across repeated trials or a best-of-N result.
- Organ-system and task-type breakdowns use smaller samples. Some organ systems have as few as 18 questions, so those individual cells carry more statistical noise than the headline scores above.
- Grading is exact-match on the model's stated letter. A correctly-reasoned answer expressed in an unexpected format can, in rare cases, be mis-scored by the automated grader.
FAQ
What is MedXpertQA?+
MedXpertQA is a benchmark of expert-level medical exam questions spanning diagnosis, treatment, and basic science, published by Tsinghua C3I. It includes a text-only subset (2,450 questions, 10 answer choices) and a multimodal subset (2,000 questions with real clinical images, 5 answer choices), and is designed to be harder and less saturated than earlier medical QA benchmarks like MedQA.
How is MedXpertQA different from MedQA?+
MedQA is built from US Medical Licensing Exam (USMLE)-style board questions and has become saturated — leading models now score above 90% on it, leaving little room to tell them apart. MedXpertQA was built specifically to address that ceiling effect: it applies stricter filtering to raise question difficulty, draws from 17 of the 25 American Board of Medical Specialties' member board exams for broader specialty coverage (4,460 questions across 17 specialties and 11 body systems), and adds a multimodal subset with real clinical images, patient records, and exam results — something MedQA doesn't have at all.
What kind of images are in the MedXpertQA multimodal subset?+
The multimodal subset spans 10 image types: radiology, pathology, other medical optical imaging, clinical photos, vital-sign readouts, diagrams, documents, charts, tables, and other miscellaneous images. Across its 2,000 multimodal questions there are roughly 2,800 images total, aiming to cover the same range of visual material a human physician would actually be expected to interpret — not just X-rays or scans.
Which AI model scored highest on medical reasoning?+
On the multimodal subset (questions with clinical images), gpt-6-astra scored highest at 87.4%, narrowly ahead of Gemini 3.8 Flash at 86.55%.
What is the human benchmark on MedXpertQA?+
Human experts score roughly 45% on the multimodal subset. A 2025 paper on GPT-5's medical reasoning (Wang et al., SPIE 13930) reports pre-licensed physicians averaging 45.53% on MedXpertQA MM (45.76% reasoning, 44.97% understanding) and 42.60% on the text subset. GPT-4o scored below that human baseline at 42.8%; every model in the current generation benchmarked here scores well above it, with gpt-6-astra at 87.4% — nearly double the expert average.
Source: Capabilities of GPT-5 on multimodal medical reasoning (SPIE, 2025)
How were the models evaluated?+
Every model ran zero-shot with chain-of-thought reasoning at the provider's default settings — no fine-tuning, no few-shot examples, and no reasoning-effort tuning. Each score is a single run, graded by exact-match on the model's final answer letter. GPT-4o, Gemini 2.0 Flash, and Claude 3.5 Sonnet scores are from the original MedXpertQA paper's published leaderboard rather than a run in this project, using the same evaluation harness so the numbers are directly comparable.
Could the newer models have memorized MedXpertQA from their training data?+
It can't be ruled out. MedXpertQA was released publicly in January 2025, so it has plausibly been part of web-scale pretraining data for every model released after that. This project doesn't run any contamination checks (e.g. canary strings, n-gram overlap, or perturbed question variants), and each score is a single zero-shot run. Newer, held-out medical benchmarks would be the real fix, but there's comparatively little activity in this space versus other domains.
Why does GPT-4o score so much lower than newer models?+
GPT-4o predates the current generation of reasoning-focused models benchmarked here. Its score (42.8% multimodal, 30.37% text) comes from the original MedXpertQA paper rather than a fresh run, and reflects how much medical reasoning performance has improved across newer model releases since GPT-4o's launch in 2024.