Skip to content
AI Atlas
BenchmarkActivecategory · multimodalfamily · mmmu · variant Pro

MMMU-Pro

robust multimodal understanding (10-option, vision-only variants)

data quality51

Updated 3 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
518
Models
157
Current leader
gpt-6-astra 86.9%

Frontier over time · accuracy · evaluator=Artificial Analysis

6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 86.9%gpt-6-astra OpenAI Independent11 Sept 2026
  2. 86.4%gpt-6-astra OpenAI Independent11 Sept 2026
  3. 85.1%gpt-6-astra OpenAI Independent11 Sept 2026
  4. 84.7%Gemini 3.7 Flash Google Independent11 Sept 2026
  5. 84.5%Gemini 3.8 Flash Google Independent11 Sept 2026
  6. 81.6%Claude Opus 5 Anthropic Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 157 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
101Llama 4 MaverickOpen weightsMeta AI · Llama 4 · best of 2 rows62.1%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
101Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 4 rows62.1%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
103gemini-2-5-flash-04-2025ClosedGoogle · Gemini 2.562.0%Independentgroup defaultsPartially comparable-0.17 ptobs. 11 Sept 2026artificialanalysis.aiT2History
103gemini-2-5-flash-reasoning-04-2025ClosedGoogle · Gemini 2.562.0%IndependentreasoningoffPartially comparable-0.17 ptobs. 12 Sept 2026artificialanalysis.aiT2History
105nova-2-0-omniClosedAmazon Web Services · Nova 2.0 · best of 6 rows61.9%Independentreasoningonreasoning_effortmediumPartially comparable-0.23 ptobs. 12 Sept 2026artificialanalysis.aiT2History
106grok-4-fastClosedSpaceXAI · Grok 4 · best of 4 rows61.8%IndependentreasoningonPartially comparable-0.35 ptobs. 12 Sept 2026artificialanalysis.aiT2History
107gpt-4.1ClosedOpenAI · GPT 4.1 · best of 2 rows61.2%IndependentreasoningoffPartially comparable-0.93 ptobs. 12 Sept 2026artificialanalysis.aiT2History
108gpt-5-nanoClosedOpenAI · GPT 5 · best of 6 rows61.0%Independentreasoning_efforthighPartially comparable-1.16 ptobs. 12 Sept 2026artificialanalysis.aiT2History
109qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows60.2%IndependentreasoningonPartially comparable-1.91 ptobs. 12 Sept 2026artificialanalysis.aiT2History
110Claude 3.7 SonnetClosedAnthropic · Claude · best of 2 rows60.1%IndependentreasoningoffPartially comparable-2.08 ptobs. 12 Sept 2026artificialanalysis.aiT2History
111Magistral Medium 1.2ClosedMistral AI · Magistral · best of 2 rows59.6%IndependentreasoningonPartially comparable-2.49 ptobs. 12 Sept 2026artificialanalysis.aiT2History
112gpt-4.1-miniClosedOpenAI · GPT 4.1 · best of 2 rows58.7%IndependentreasoningoffPartially comparable-3.41 ptobs. 12 Sept 2026artificialanalysis.aiT2History
113Claude Haiku 4.5ClosedAnthropic · Claude · best of 4 rows58.5%IndependentreasoningonPartially comparable-3.59 ptobs. 12 Sept 2026artificialanalysis.aiT2History
114apriel-v1-5-15b-thinkerOpen weightsServiceNow · best of 2 rows57.0%IndependentreasoningonPartially comparable-5.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
115Mistral Small 4Open weightsMistral AI · Mistral · best of 4 rows56.8%IndependentreasoningonPartially comparable-5.32 ptobs. 12 Sept 2026artificialanalysis.aiT2History
116Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 4 rows56.6%IndependentreasoningonPartially comparable-5.49 ptobs. 12 Sept 2026artificialanalysis.aiT2History
117GPT-4o (2024-08-06)ClosedOpenAI · GPT 4 · best of 2 rows56.3%IndependentreasoningoffPartially comparable-5.84 ptobs. 12 Sept 2026artificialanalysis.aiT2History
118Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows55.7%IndependentreasoningoffPartially comparable-6.48 ptobs. 12 Sept 2026artificialanalysis.aiT2History
119Magistral Small 1.2Open weightsMistral AI · Magistral · best of 2 rows55.5%IndependentreasoningonPartially comparable-6.65 ptobs. 12 Sept 2026artificialanalysis.aiT2History
120gemini-1-5-proClosedGoogle · Gemini 1.5 · best of 2 rows55.0%IndependentreasoningoffPartially comparable-7.11 ptobs. 12 Sept 2026artificialanalysis.aiT2History
121Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows54.2%IndependentreasoningoffPartially comparable-7.98 ptobs. 12 Sept 2026artificialanalysis.aiT2History
122nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows53.2%IndependentreasoningonPartially comparable-8.96 ptobs. 12 Sept 2026artificialanalysis.aiT2History
123Mistral Medium 3ClosedMistral AI · Mistral · best of 2 rows53.0%IndependentreasoningoffPartially comparable-9.13 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124Llama 4 ScoutOpen weightsMeta AI · Llama 4 · best of 2 rows53.0%IndependentreasoningoffPartially comparable-9.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
125nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 4 rows52.9%IndependentreasoningonPartially comparable-9.25 ptobs. 12 Sept 2026artificialanalysis.aiT2History
126Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 4 rows52.0%IndependentreasoningonPartially comparable-10.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 4 rows51.4%IndependentreasoningonPartially comparable-10.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
128Pixtral LargeOpen weightsMistral AI · Pixtral · best of 2 rows50.6%IndependentreasoningoffPartially comparable-11.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
129GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 4 rows50.5%IndependentreasoningonPartially comparable-11.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
130Ministral 3 14BOpen weightsMistral AI · Ministral 3 · best of 2 rows49.8%IndependentreasoningoffPartially comparable-12.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 4 rows48.5%IndependentreasoningonPartially comparable-13.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
132gemini-1-5-flashClosedGoogle · Gemini 1.5 · best of 2 rows48.4%IndependentreasoningoffPartially comparable-13.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
133Gemma 3 27BOpen weightsGoogle · Gemma 3 · best of 2 rows48.0%IndependentreasoningoffPartially comparable-14.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
134Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows48.0%IndependentreasoningoffPartially comparable-14.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
135Ministral 3 8BOpen weightsMistral AI · Ministral 3 · best of 2 rows46.0%IndependentreasoningoffPartially comparable-16.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
136Claude Haiku 3.5ClosedAnthropic · Claude · best of 2 rows45.6%IndependentreasoningoffPartially comparable-16.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
137Devstral Small 2Open weightsMistral AI · Devstral · best of 2 rows44.6%IndependentreasoningoffPartially comparable-17.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
138gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 4 rows44.6%IndependentreasoningonPartially comparable-17.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
139nova-proClosedAmazon Web Services · Nova · best of 2 rows44.3%IndependentreasoningoffPartially comparable-17.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
140Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 4 rows43.0%IndependentreasoningonPartially comparable-19.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
141gpt-4o-miniClosedOpenAI · GPT 4 · best of 2 rows41.5%IndependentreasoningoffPartially comparable-20.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
142gpt-4.1-nanoClosedOpenAI · GPT 4.1 · best of 2 rows40.1%IndependentreasoningoffPartially comparable-22.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
143llama-3-2-instruct-90b-visionOpen weightsMeta AI · Llama 3.2 · best of 2 rows39.5%IndependentreasoningoffPartially comparable-22.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
144Ministral 3 3BOpen weightsMistral AI · Ministral 3 · best of 2 rows38.1%IndependentreasoningoffPartially comparable-24.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
145minicpm-v4-6-1-3bOpen weightsOpenBMB · best of 2 rows37.9%IndependentreasoningoffPartially comparable-24.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
146nova-liteClosedAmazon Web Services · Nova · best of 2 rows37.8%IndependentreasoningoffPartially comparable-24.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
147Gemma 3 12BOpen weightsGoogle · Gemma 3 · best of 2 rows37.5%IndependentreasoningoffPartially comparable-24.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
148Molmo2-8BOpen weightsAllen Institute for AI · best of 2 rows37.5%IndependentreasoningoffPartially comparable-24.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
149gemini-1-5-flash-8bClosedGoogle · Gemini 1.5 · best of 2 rows36.5%IndependentreasoningoffPartially comparable-25.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
150Claude 3 HaikuClosedAnthropic · Claude · best of 2 rows30.8%IndependentreasoningoffPartially comparable-31.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
151Gemma 3 4BRestricted weightsGoogle · Gemma 3 · best of 2 rows29.9%IndependentreasoningoffPartially comparable-32.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
152llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.2 · best of 2 rows29.3%IndependentreasoningoffPartially comparable-32.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
153LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.5 · best of 2 rows26.5%IndependentreasoningoffPartially comparable-35.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
154gemma-3n-e4bOpen weightsGoogle · Gemma 3 · best of 2 rows26.2%IndependentreasoningoffPartially comparable-35.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
155Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 4 rows25.8%IndependentreasoningonPartially comparable-36.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
156molmo-7b-dOpen weightsAllen Institute for AI · Molmo · best of 2 rows24.5%IndependentreasoningoffPartially comparable-37.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
157phi-4-multimodalOpen weightsMicrosoft · Phi4 · best of 2 rows14.5%IndependentreasoningoffPartially comparable-47.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →