MMMU-Pro
robust multimodal understanding (10-option, vision-only variants)
Updated 3 h ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 518
- Models
- 157
- Current leader
- gpt-6-astra 86.9%
Frontier over time · accuracy · evaluator=Artificial Analysis
6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 86.9%gpt-6-astra OpenAI Independent11 Sept 2026
- 86.4%gpt-6-astra OpenAI Independent11 Sept 2026
- 85.1%gpt-6-astra OpenAI Independent11 Sept 2026
- 84.7%Gemini 3.7 Flash Google Independent11 Sept 2026
- 84.5%Gemini 3.8 Flash Google Independent11 Sept 2026
- 81.6%Claude Opus 5 Anthropic Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 157 models
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 101 | Llama 4 MaverickOpen weightsMeta AI · Llama 4 · best of 2 rows | 62.1% | Independent | reasoningoff | Partially comparable0.00 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 101 | Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 4 rows | 62.1% | Independent | reasoningoff | Partially comparable0.00 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 103 | gemini-2-5-flash-04-2025ClosedGoogle · Gemini 2.5 | 62.0% | Independent | group defaults | Partially comparable-0.17 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 103 | gemini-2-5-flash-reasoning-04-2025ClosedGoogle · Gemini 2.5 | 62.0% | Independent | reasoningoff | Partially comparable-0.17 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 105 | nova-2-0-omniClosedAmazon Web Services · Nova 2.0 · best of 6 rows | 61.9% | Independent | reasoningonreasoning_effortmedium | Partially comparable-0.23 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 106 | grok-4-fastClosedSpaceXAI · Grok 4 · best of 4 rows | 61.8% | Independent | reasoningon | Partially comparable-0.35 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 107 | gpt-4.1ClosedOpenAI · GPT 4.1 · best of 2 rows | 61.2% | Independent | reasoningoff | Partially comparable-0.93 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 108 | gpt-5-nanoClosedOpenAI · GPT 5 · best of 6 rows | 61.0% | Independent | reasoning_efforthigh | Partially comparable-1.16 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 109 | qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows | 60.2% | Independent | reasoningon | Partially comparable-1.91 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 110 | Claude 3.7 SonnetClosedAnthropic · Claude · best of 2 rows | 60.1% | Independent | reasoningoff | Partially comparable-2.08 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 111 | Magistral Medium 1.2ClosedMistral AI · Magistral · best of 2 rows | 59.6% | Independent | reasoningon | Partially comparable-2.49 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 112 | gpt-4.1-miniClosedOpenAI · GPT 4.1 · best of 2 rows | 58.7% | Independent | reasoningoff | Partially comparable-3.41 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 113 | Claude Haiku 4.5ClosedAnthropic · Claude · best of 4 rows | 58.5% | Independent | reasoningon | Partially comparable-3.59 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 114 | apriel-v1-5-15b-thinkerOpen weightsServiceNow · best of 2 rows | 57.0% | Independent | reasoningon | Partially comparable-5.09 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 115 | Mistral Small 4Open weightsMistral AI · Mistral · best of 4 rows | 56.8% | Independent | reasoningon | Partially comparable-5.32 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 116 | Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 4 rows | 56.6% | Independent | reasoningon | Partially comparable-5.49 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 117 | GPT-4o (2024-08-06)ClosedOpenAI · GPT 4 · best of 2 rows | 56.3% | Independent | reasoningoff | Partially comparable-5.84 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 118 | Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows | 55.7% | Independent | reasoningoff | Partially comparable-6.48 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 119 | Magistral Small 1.2Open weightsMistral AI · Magistral · best of 2 rows | 55.5% | Independent | reasoningon | Partially comparable-6.65 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 120 | gemini-1-5-proClosedGoogle · Gemini 1.5 · best of 2 rows | 55.0% | Independent | reasoningoff | Partially comparable-7.11 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 121 | Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows | 54.2% | Independent | reasoningoff | Partially comparable-7.98 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 122 | nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 53.2% | Independent | reasoningon | Partially comparable-8.96 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 123 | Mistral Medium 3ClosedMistral AI · Mistral · best of 2 rows | 53.0% | Independent | reasoningoff | Partially comparable-9.13 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 124 | Llama 4 ScoutOpen weightsMeta AI · Llama 4 · best of 2 rows | 53.0% | Independent | reasoningoff | Partially comparable-9.19 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 125 | nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 4 rows | 52.9% | Independent | reasoningon | Partially comparable-9.25 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 126 | Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 4 rows | 52.0% | Independent | reasoningon | Partially comparable-10.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 127 | gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 4 rows | 51.4% | Independent | reasoningon | Partially comparable-10.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 128 | Pixtral LargeOpen weightsMistral AI · Pixtral · best of 2 rows | 50.6% | Independent | reasoningoff | Partially comparable-11.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 129 | GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 4 rows | 50.5% | Independent | reasoningon | Partially comparable-11.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 130 | Ministral 3 14BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 49.8% | Independent | reasoningoff | Partially comparable-12.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 131 | GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 4 rows | 48.5% | Independent | reasoningon | Partially comparable-13.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 132 | gemini-1-5-flashClosedGoogle · Gemini 1.5 · best of 2 rows | 48.4% | Independent | reasoningoff | Partially comparable-13.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 133 | Gemma 3 27BOpen weightsGoogle · Gemma 3 · best of 2 rows | 48.0% | Independent | reasoningoff | Partially comparable-14.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 134 | Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows | 48.0% | Independent | reasoningoff | Partially comparable-14.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 135 | Ministral 3 8BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 46.0% | Independent | reasoningoff | Partially comparable-16.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 136 | Claude Haiku 3.5ClosedAnthropic · Claude · best of 2 rows | 45.6% | Independent | reasoningoff | Partially comparable-16.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 137 | Devstral Small 2Open weightsMistral AI · Devstral · best of 2 rows | 44.6% | Independent | reasoningoff | Partially comparable-17.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 138 | gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 4 rows | 44.6% | Independent | reasoningon | Partially comparable-17.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 139 | nova-proClosedAmazon Web Services · Nova · best of 2 rows | 44.3% | Independent | reasoningoff | Partially comparable-17.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 140 | Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 4 rows | 43.0% | Independent | reasoningon | Partially comparable-19.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 141 | gpt-4o-miniClosedOpenAI · GPT 4 · best of 2 rows | 41.5% | Independent | reasoningoff | Partially comparable-20.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 142 | gpt-4.1-nanoClosedOpenAI · GPT 4.1 · best of 2 rows | 40.1% | Independent | reasoningoff | Partially comparable-22.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 143 | llama-3-2-instruct-90b-visionOpen weightsMeta AI · Llama 3.2 · best of 2 rows | 39.5% | Independent | reasoningoff | Partially comparable-22.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 144 | Ministral 3 3BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 38.1% | Independent | reasoningoff | Partially comparable-24.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 145 | minicpm-v4-6-1-3bOpen weightsOpenBMB · best of 2 rows | 37.9% | Independent | reasoningoff | Partially comparable-24.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 146 | nova-liteClosedAmazon Web Services · Nova · best of 2 rows | 37.8% | Independent | reasoningoff | Partially comparable-24.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 147 | Gemma 3 12BOpen weightsGoogle · Gemma 3 · best of 2 rows | 37.5% | Independent | reasoningoff | Partially comparable-24.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 148 | Molmo2-8BOpen weightsAllen Institute for AI · best of 2 rows | 37.5% | Independent | reasoningoff | Partially comparable-24.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 149 | gemini-1-5-flash-8bClosedGoogle · Gemini 1.5 · best of 2 rows | 36.5% | Independent | reasoningoff | Partially comparable-25.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 150 | Claude 3 HaikuClosedAnthropic · Claude · best of 2 rows | 30.8% | Independent | reasoningoff | Partially comparable-31.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 151 | Gemma 3 4BRestricted weightsGoogle · Gemma 3 · best of 2 rows | 29.9% | Independent | reasoningoff | Partially comparable-32.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 152 | llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.2 · best of 2 rows | 29.3% | Independent | reasoningoff | Partially comparable-32.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 153 | LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.5 · best of 2 rows | 26.5% | Independent | reasoningoff | Partially comparable-35.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 154 | gemma-3n-e4bOpen weightsGoogle · Gemma 3 · best of 2 rows | 26.2% | Independent | reasoningoff | Partially comparable-35.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 155 | Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 4 rows | 25.8% | Independent | reasoningon | Partially comparable-36.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 156 | molmo-7b-dOpen weightsAllen Institute for AI · Molmo · best of 2 rows | 24.5% | Independent | reasoningoff | Partially comparable-37.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 157 | phi-4-multimodalOpen weightsMicrosoft · Phi4 · best of 2 rows | 14.5% | Independent | reasoningoff | Partially comparable-47.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →