Skip to content
AI Atlas

Humanity's Last Exam — cost vs performance

Best current row per canonical model in the group “accuracy · evaluator=Artificial Analysis” (453 models) against context window (tokens). The dashed line is the Pareto frontier: no model is both better and cheaper than a point on it.

open / restricted weights closed Pareto frontier (10 models)

Frontier models 10

Pareto frontier
Modelaccuracy on Humanity's Last ExamRankcontext window (tokens)ProviderTrust
Claude Fable 5.1Anthropic · Closed59.1%11MIndependent
Grok 4.6xAI · Closed44.1%15500KIndependent
gpt-5.3-codexOpenAI · Closed42.5%21400KIndependent
muse-sparkMeta AI · Closed40.7%29262.1KIndependent
grok-build-0-1-06-16SpaceXAI · Closed38.3%40256KIndependent
Claude Opus 4.5Anthropic · Closed30.1%74200KIndependent
deepseek-v3-2-specialeDeepSeek · Open weights28.7%81128KIndependent
qwen3-235b-a22b-instructAlibaba Group · Open weights11.0%18432.8KIndependent
LFM2.5-1.2B-InstructLiquid AI · Restricted weights6.72%25132KIndependent
olmo-2-7bAllen Institute for AI · Open weights5.38%2914.1KIndependent

Methodology. Points are the best current row per canonical model in comparability group 'accuracy · evaluator=Artificial Analysis'. Price = cheapest current offer across providers (the provider shown). Pareto frontier maximises the score and minimises x; exact ties are all kept. memory_estimate is an estimate (see /methodology). Only the selected comparability group is plotted; points under other configurations are not mixed in. Price = cheapest current offer across providers at the time of the last crawl. Nothing is estimated except where marked.