Skip to content
AI Atlas
BenchmarkActivecategory · multimodalfamily · mmmu · variant Pro

MMMU-Pro

robust multimodal understanding (10-option, vision-only variants)

data quality51

Updated 4 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
518
Models
157
Current leader
gpt-6-astra 86.9%

Score history · Ministral 3 3B 1 row

Not enough history to chart — a single observation (38.09% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.

  • 38.09%aa_slug=ministral-3-3b · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · evaluator=Artificial Analysis

6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 86.9%gpt-6-astra OpenAI Independent11 Sept 2026
  2. 86.4%gpt-6-astra OpenAI Independent11 Sept 2026
  3. 85.1%gpt-6-astra OpenAI Independent11 Sept 2026
  4. 84.7%Gemini 3.7 Flash Google Independent11 Sept 2026
  5. 84.5%Gemini 3.8 Flash Google Independent11 Sept 2026
  6. 81.6%Claude Opus 5 Anthropic Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 157 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1gpt-6-astraClosedOpenAI · GPT 6 · best of 11 rows86.9%Independentreasoning_effortmaxleaderobs. 12 Sept 2026artificialanalysis.aiT2History
2Gemini 3.8 FlashClosedGoogle · Gemini 3.8 · best of 6 rows85.6%Independentreasoning_efforthighPartially comparable-1.27 ptobs. 12 Sept 2026artificialanalysis.aiT2History
3Gemini 3.7 FlashClosedGoogle · Gemini 3.7 · best of 6 rows85.5%Independentreasoning_efforthighPartially comparable-1.39 ptobs. 12 Sept 2026artificialanalysis.aiT2History
4Claude Opus 5ClosedAnthropic · Claude · best of 10 rows84.7%Independentreasoning_effortmaxComparable-2.14 ptobs. 12 Sept 2026artificialanalysis.aiT2History
5Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 6 rows84.3%IndependentreasoningonPartially comparable-2.60 ptobs. 12 Sept 2026artificialanalysis.aiT2History
6gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 12 rows83.4%Independentreasoning_effortmaxComparable-3.47 ptobs. 12 Sept 2026artificialanalysis.aiT2History
7Gemini 3.6 FlashClosedGoogle · Gemini 3.6 · best of 2 rows83.2%IndependentreasoningonPartially comparable-3.64 ptobs. 12 Sept 2026artificialanalysis.aiT2History
8Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.1 · best of 2 rows82.4%IndependentreasoningonPartially comparable-4.45 ptobs. 12 Sept 2026artificialanalysis.aiT2History
9Qwen 3.8 MaxClosedQwen · Qwen3.8 · best of 2 rows82.3%IndependentreasoningonPartially comparable-4.57 ptobs. 12 Sept 2026artificialanalysis.aiT2History
10Muse Spark 1.3ClosedMeta AI · best of 2 rows82.0%Independentreasoning_effortxhighPartially comparable-4.86 ptobs. 12 Sept 2026artificialanalysis.aiT2History
11gpt-5.5ClosedOpenAI · GPT 5.5 · best of 10 rows81.2%Independentreasoning_effortmediumPartially comparable-5.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
12gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 12 rows80.7%Independentreasoning_effortmaxComparable-6.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
13Kimi K3Open weightsMoonshot AI · Kimi · best of 4 rows80.5%Independentreasoning_effortmaxComparable-6.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
13muse-sparkClosedMeta AI · best of 2 rows80.5%IndependentreasoningonPartially comparable-6.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
15Qwen 3.7 PlusClosedQwen · Qwen3.7 · best of 2 rows80.5%IndependentreasoningonPartially comparable-6.42 ptobs. 12 Sept 2026artificialanalysis.aiT2History
16Grok 4.5ClosedxAI · Grok · best of 2 rows80.4%Independentreasoning_efforthighPartially comparable-6.48 ptobs. 12 Sept 2026artificialanalysis.aiT2History
17gemini-3-proClosedGoogle · Gemini 3 · best of 2 rows80.2%Independentreasoning_efforthighPartially comparable-6.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
18Gemini 3 Flash PreviewClosedGoogle · Gemini 379.9%Independentgroup defaultsPartially comparable-6.94 ptobs. 11 Sept 2026artificialanalysis.aiT2History
18gemini-3-flashClosedGoogle · Gemini 3 · best of 3 rows79.9%IndependentreasoningonPartially comparable-6.94 ptobs. 12 Sept 2026artificialanalysis.aiT2History
20Qwen3.8 FlashOpen weightsQwen · Qwen3.8 · best of 2 rows79.8%IndependentreasoningonPartially comparable-7.11 ptobs. 12 Sept 2026artificialanalysis.aiT2History
21Kimi K2.6Open weightsMoonshot AI · Kimi · best of 2 rows79.4%IndependentreasoningonPartially comparable-7.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
22apodex-1-1ClosedApodex · best of 2 rows79.2%IndependentreasoningonPartially comparable-7.69 ptobs. 12 Sept 2026artificialanalysis.aiT2History
23Gemini 3.5 Flash-LiteClosedGoogle · Gemini 3.5 · best of 2 rows79.0%IndependentreasoningonPartially comparable-7.86 ptobs. 12 Sept 2026artificialanalysis.aiT2History
24Ling 3.0 Flash VLOpen weightsinclusionAI · best of 2 rows79.0%IndependentreasoningonPartially comparable-7.92 ptobs. 12 Sept 2026artificialanalysis.aiT2History
25Claude Opus 4.7ClosedAnthropic · Claude · best of 4 rows78.8%Independentreasoning_effortmaxComparable-8.04 ptobs. 12 Sept 2026artificialanalysis.aiT2History
26MiniMax-M3Open weightsMiniMax · MiniMax · best of 2 rows78.5%IndependentreasoningonPartially comparable-8.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
26gpt-5.6-lunaClosedOpenAI · GPT 5.6 · best of 12 rows78.5%Independentreasoning_effortxhighPartially comparable-8.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
28gpt-5.3-codexClosedOpenAI · GPT 5.3 · best of 2 rows78.5%Independentreasoning_effortxhighPartially comparable-8.38 ptobs. 12 Sept 2026artificialanalysis.aiT2History
29gpt-5.4ClosedOpenAI · GPT 5.4 · best of 6 rows78.4%Independentreasoning_effortxhighPartially comparable-8.44 ptobs. 12 Sept 2026artificialanalysis.aiT2History
30Grok 4.3ClosedxAI · Grok · best of 8 rows78.1%Independentreasoning_efforthighPartially comparable-8.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
31Qwen3.6 PlusClosedQwen · Qwen3.6 · best of 2 rows78.0%IndependentreasoningonPartially comparable-8.90 ptobs. 12 Sept 2026artificialanalysis.aiT2History
32Claude Sonnet 5ClosedAnthropic · Claude · best of 4 rows77.3%Independentreasoning_effortmaxComparable-9.60 ptobs. 12 Sept 2026artificialanalysis.aiT2History
32Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 4 rows77.3%IndependentreasoningonPartially comparable-9.60 ptobs. 12 Sept 2026artificialanalysis.aiT2History
34DeepSeek-V4.1-FlashOpen weightsDeepSeek · DeepSeek · best of 2 rows77.0%Independentreasoning_effortmaxComparable-9.89 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35grok-build-0-1-06-16ClosedSpaceXAI · Grok · best of 2 rows76.5%IndependentreasoningonPartially comparable-10.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
36GPT-5.2-CodexClosedOpenAI · GPT 5.2 · best of 2 rows76.3%Independentreasoning_effortxhighPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
36Qwen3.8 27BOpen weightsQwen · Qwen3.8 · best of 8 rows76.3%Independentreasoning_effortxhighPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
38Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.1 · best of 2 rows75.5%IndependentreasoningonPartially comparable-11.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
39gpt-5.1ClosedOpenAI · GPT 5.1 · best of 4 rows75.5%Independentreasoning_efforthighPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
40Claude Opus 4.6ClosedAnthropic · Claude · best of 4 rows75.4%Independentreasoningadaptivereasoning_effortmaxPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
40MiMo-V2.5Open weightsXiaomi · best of 2 rows75.4%IndependentreasoningonPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
40agnes-2-5-pro-betaClosedSapiens AI · best of 2 rows75.4%IndependentreasoningonPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
43Kimi K2.5Open weightsMoonshot AI · Kimi · best of 4 rows75.4%IndependentreasoningonPartially comparable-11.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
44Step 3.7 FlashOpen weightsStepFun · Step3.7 · best of 2 rows75.3%IndependentreasoningonPartially comparable-11.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
45Qwen3.5-27BOpen weightsQwen · Qwen3.5 · best of 4 rows75.0%IndependentreasoningonPartially comparable-11.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
45Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 4 rows75.0%IndependentreasoningonPartially comparable-11.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
47Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 4 rows75.0%IndependentreasoningonPartially comparable-11.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
48Gemini 2.5 ProClosedGoogle · Gemini 2.5 · best of 2 rows74.9%IndependentreasoningonPartially comparable-12.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49deepseek-v4-flash-visionClosedDeepSeek · DeepSeek · best of 2 rows74.8%Independentreasoning_effortmaxComparable-12.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
50Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 4 rows74.6%IndependentreasoningonPartially comparable-12.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
51Grok 4.20ClosedxAI · Grok · best of 3 rows74.6%IndependentreasoningonPartially comparable-12.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
51gpt-5.2ClosedOpenAI · GPT 5.2 · best of 4 rows74.6%Independentreasoning_effortmediumPartially comparable-12.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
53Muse Glimmer 30BOpen weightsMeta AI · best of 2 rows74.3%Independentreasoning_efforthighPartially comparable-12.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
53gpt-5ClosedOpenAI · GPT 5 · best of 8 rows74.3%Independentreasoning_effortmediumPartially comparable-12.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
55Claude Opus 4.5ClosedAnthropic · Claude · best of 4 rows74.0%IndependentreasoningonPartially comparable-12.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
55Inkling SmallOpen weightsThinking Machines · best of 2 rows74.0%IndependentreasoningonPartially comparable-12.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
57mimo-v2-omni-0327ClosedXiaomi · best of 2 rows73.9%IndependentreasoningonPartially comparable-12.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
58GPT-5-CodexClosedOpenAI · GPT 5 · best of 2 rows73.8%Independentreasoning_efforthighPartially comparable-13.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
59InklingOpen weightsThinking Machines · best of 2 rows73.5%IndependentreasoningonPartially comparable-13.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 4 rows73.4%IndependentreasoningonPartially comparable-13.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
61Claude Sonnet 4.6ClosedAnthropic · Claude · best of 5 rows73.3%Independentreasoningadaptivereasoning_effortmaxPartially comparable-13.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
61gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 6 rows73.3%Independentreasoning_effortxhighPartially comparable-13.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
63grok-4-20-0309ClosedSpaceXAI · Grok 4.20 · best of 4 rows73.2%IndependentreasoningonPartially comparable-13.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
64gemini-2-5-flash-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 6 rows73.1%IndependentreasoningonPartially comparable-13.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
65GLM 5V TurboClosedZ.ai (Zhipu AI) · GLM5 · best of 2 rows72.8%IndependentreasoningonPartially comparable-14.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
66Qwen3.5-35B-A3BOpen weightsQwen · Qwen3.5 · best of 4 rows72.7%IndependentreasoningonPartially comparable-14.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
67GPT-5.1-CodexClosedOpenAI · GPT 5.1 · best of 2 rows72.5%Independentreasoning_efforthighPartially comparable-14.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
68qwen3-5-omni-plusClosedAlibaba Group · Qwen3.5 · best of 2 rows70.5%IndependentreasoningoffPartially comparable-16.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
69gpt-5-miniClosedOpenAI · GPT 5 · best of 6 rows70.1%Independentreasoning_efforthighPartially comparable-16.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
69o3ClosedOpenAI · OpenAI o-series · best of 2 rows70.1%IndependentreasoningonPartially comparable-16.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
71mimo-v2-omniClosedXiaomi · best of 2 rows69.9%IndependentreasoningonPartially comparable-17.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
72gemma-4-12BOpen weightsGoogle · Gemma 4 · best of 4 rows69.7%IndependentreasoningonPartially comparable-17.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
73Gemma 4 26B A4BOpen weightsGoogle · Gemma 4 · best of 4 rows69.3%IndependentreasoningonPartially comparable-17.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
73Qwen3.5-9BOpen weightsQwen · Qwen3.5 · best of 4 rows69.3%IndependentreasoningonPartially comparable-17.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
73o4-miniClosedOpenAI · OpenAI o-series · best of 2 rows69.3%Independentreasoning_efforthighPartially comparable-17.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
76Gemini 2.5 FlashClosedGoogle · Gemini 2.5 · best of 2 rows69.1%IndependentreasoningonPartially comparable-17.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
77GPT-5.1-Codex MiniClosedOpenAI · GPT 5.1 · best of 2 rows69.0%Independentreasoning_efforthighPartially comparable-17.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
78grok-4ClosedSpaceXAI · Grok 4 · best of 2 rows68.8%IndependentreasoningonPartially comparable-18.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79Claude Sonnet 4.5ClosedAnthropic · Claude · best of 4 rows68.7%IndependentreasoningonPartially comparable-18.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79Qwen3 VL 235B A22B InstructOpen weightsQwen · Qwen3 · best of 4 rows68.7%IndependentreasoningonPartially comparable-18.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
81doubao-seed-codeClosedByteDance · Seed · best of 2 rows68.1%IndependentreasoningonPartially comparable-18.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
82Claude Opus 4.1ClosedAnthropic · Claude · best of 2 rows67.9%IndependentreasoningonPartially comparable-19.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
83exaone-4-5-33bOpen weightsLG AI Research · EXAONE 4.5 · best of 2 rows67.3%IndependentreasoningonPartially comparable-19.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
84diffusiongemma-26b-a4bOpen weightsGoogle · best of 2 rows66.5%IndependentreasoningonPartially comparable-20.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
85Qwen3.5-4BOpen weightsQwen · Qwen3.5 · best of 4 rows65.4%IndependentreasoningonPartially comparable-21.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
85gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 6 rows65.4%Independentreasoning_effortxhighPartially comparable-21.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
87Gemini 2.5 Flash-LiteClosedGoogle · Gemini 2.5 · best of 5 rows65.0%IndependentreasoningonPartially comparable-21.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
87gemini-2-5-flash-lite-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 3 rows65.0%IndependentreasoningonPartially comparable-21.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
89grok-4.20-0309-non-reasoningClosedxAI · Grok64.9%Independentgroup defaultsPartially comparable-22.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
90Mistral Medium 3.5Open weightsMistral AI · Mistral · best of 2 rows64.9%IndependentreasoningonPartially comparable-22.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
91qwen3-5-omni-flashClosedAlibaba Group · Qwen3.5 · best of 2 rows64.7%IndependentreasoningoffPartially comparable-22.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
92ernie-5-0-thinking-previewClosedBaidu · ERNIE 5.0 · best of 2 rows64.6%IndependentreasoningonPartially comparable-22.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
93nova-2-0-proClosedAmazon Web Services · Nova 2.0 · best of 4 rows64.5%Independentreasoningonreasoning_effortmediumPartially comparable-22.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
94Qwen3 VL 32B InstructOpen weightsQwen · Qwen3 · best of 4 rows64.3%IndependentreasoningoffPartially comparable-22.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
95jt-4-1-flash-236b-a21bClosedChina Mobile · best of 2 rows64.1%IndependentreasoningoffPartially comparable-22.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
96step-3-vl-10bOpen weightsStepFun · Step3 · best of 2 rows64.0%IndependentreasoningonPartially comparable-22.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
97nova-2-0-liteClosedAmazon Web Services · Nova 2.0 · best of 8 rows63.8%Independentreasoningonreasoning_efforthighPartially comparable-23.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
98grok-4-1-fastClosedSpaceXAI · Grok 4.1 · best of 4 rows63.3%IndependentreasoningonPartially comparable-23.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
99command-a-plusOpen weightsCohere · Command · best of 2 rows63.2%IndependentreasoningonPartially comparable-23.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
100Claude Sonnet 4ClosedAnthropic · Claude · best of 4 rows62.4%IndependentreasoningoffPartially comparable-24.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →