Skip to content
AI Atlas

MMMU-Pro — cost vs performance

Best current row per canonical model in the group “accuracy · evaluator=Artificial Analysis” (157 models) against context window (tokens). The dashed line is the Pareto frontier: no model is both better and cheaper than a point on it.

open / restricted weights closed Pareto frontier (11 models)bubble = context window

Frontier models 11

Pareto frontier
Modelaccuracy on MMMU-ProRankcontext window (tokens)ProviderTrust
gpt-6-astraOpenAI · Closed86.9%11.05MIndependent
Gemini 3.8 FlashGoogle · Closed85.6%21.05MIndependent
Claude Opus 5Anthropic · Closed84.7%41MIndependent
muse-sparkMeta AI · Closed80.5%13262.1KIndependent
apodex-1-1Apodex · Closed79.2%22256KIndependent
Ling 3.0 Flash VLinclusionAI · Open weights79.0%24131.1KIndependent
ernie-5-0-thinking-previewBaidu · Closed64.6%92128KIndependent
step-3-vl-10bStepFun · Open weights64.0%9665.5KIndependent
Molmo2-8BAllen Institute for AI · Open weights37.5%14836.9KIndependent
LFM2.5-VL-1.6BLiquid AI · Open weights26.5%15332KIndependent
molmo-7b-dAllen Institute for AI · Open weights24.5%1564.1KIndependent

Methodology. Points are the best current row per canonical model in comparability group 'accuracy · evaluator=Artificial Analysis'. Price = cheapest current offer across providers (the provider shown). Pareto frontier maximises the score and minimises x; exact ties are all kept. memory_estimate is an estimate (see /methodology). Only the selected comparability group is plotted; points under other configurations are not mixed in. Price = cheapest current offer across providers at the time of the last crawl. Nothing is estimated except where marked.