Skip to content
AI Atlas
BenchmarkActivecategory · knowledgefamily · humanitys-last-exam · variant full

Humanity's Last Exam

lastexam.ai

expert-written frontier questions

data quality57

Updated 3 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,217
Models
453
Current leader
Claude Fable 5.1 59.1%

Frontier over time · accuracy · evaluator=Artificial Analysis

9 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 59.1%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  2. 58.7%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  3. 55.9%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  4. 55.5%Claude Fable 5 Anthropic Independent11 Sept 2026
  5. 53.1%gpt-6-astra OpenAI Independent11 Sept 2026
  6. 52.7%gpt-6-astra OpenAI Independent11 Sept 2026
  7. 51.3%Claude Opus 5 Anthropic Independent11 Sept 2026
  8. 11.0%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 453 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
101Step 3.5 FlashOpen weightsStepFun · Step3.5 · best of 4 rows24.5%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
102Qwen3.5-27BOpen weightsQwen · Qwen3.5 · best of 4 rows23.9%IndependentreasoningonPartially comparable-0.56 ptobs. 12 Sept 2026artificialanalysis.aiT2History
103Kimi K2 ThinkingOpen weightsMoonshot AI · Kimi · best of 2 rows23.8%IndependentreasoningonPartially comparable-0.65 ptobs. 12 Sept 2026artificialanalysis.aiT2History
104Ling 3.0 FlashOpen weightsinclusionAI · best of 2 rows23.7%IndependentreasoningonPartially comparable-0.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
105Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 4 rows23.6%IndependentreasoningonPartially comparable-0.84 ptobs. 12 Sept 2026artificialanalysis.aiT2History
106MiniMax M2.1Open weightsMiniMax · MiniMax · best of 2 rows23.2%IndependentreasoningonPartially comparable-1.30 ptobs. 12 Sept 2026artificialanalysis.aiT2History
107Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 4 rows23.1%IndependentreasoningonPartially comparable-1.39 ptobs. 12 Sept 2026artificialanalysis.aiT2History
108mimo-v2-flashOpen weightsXiaomi · best of 4 rows22.9%IndependentreasoningonPartially comparable-1.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
109mimo-v2-omni-0327ClosedXiaomi · best of 2 rows22.6%IndependentreasoningonPartially comparable-1.90 ptobs. 12 Sept 2026artificialanalysis.aiT2History
110Gemini 2.5 ProClosedGoogle · Gemini 2.5 · best of 2 rows22.5%IndependentreasoningonPartially comparable-1.95 ptobs. 12 Sept 2026artificialanalysis.aiT2History
111Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 4 rows22.2%IndependentreasoningonPartially comparable-2.23 ptobs. 12 Sept 2026artificialanalysis.aiT2History
112mimo-v2-0206Open weightsXiaomi · best of 2 rows22.1%IndependentreasoningonPartially comparable-2.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
112mimo-v2-omniClosedXiaomi · best of 2 rows22.1%IndependentreasoningonPartially comparable-2.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
114Ling 3.0 Flash VLOpen weightsinclusionAI · best of 2 rows22.0%IndependentreasoningonPartially comparable-2.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
114Muse Glimmer 30BOpen weightsMeta AI · best of 2 rows22.0%Independentreasoning_efforthighPartially comparable-2.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
116gpt-5-5-instant-05-26ClosedOpenAI · GPT 5.5 · best of 2 rows21.6%IndependentreasoningonPartially comparable-2.88 ptobs. 12 Sept 2026artificialanalysis.aiT2History
116ring-2-6-1tOpen weightsinclusionAI · best of 2 rows21.6%IndependentreasoningonPartially comparable-2.88 ptobs. 12 Sept 2026artificialanalysis.aiT2History
118gpt-5-miniClosedOpenAI · GPT 5 · best of 6 rows21.5%Independentreasoning_efforthighPartially comparable-3.01 ptobs. 12 Sept 2026artificialanalysis.aiT2History
119Step 3.7 FlashOpen weightsStepFun · Step3.7 · best of 2 rows21.4%IndependentreasoningonPartially comparable-3.06 ptobs. 12 Sept 2026artificialanalysis.aiT2History
120Qwen3.5-35B-A3BOpen weightsQwen · Qwen3.5 · best of 4 rows21.0%IndependentreasoningonPartially comparable-3.43 ptobs. 12 Sept 2026artificialanalysis.aiT2History
121Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 3 · best of 2 rows20.8%IndependentreasoningonPartially comparable-3.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
122MiniMax M2.5Open weightsMiniMax · MiniMax · best of 2 rows20.5%IndependentreasoningonPartially comparable-3.94 ptobs. 12 Sept 2026artificialanalysis.aiT2History
123o3ClosedOpenAI · OpenAI o-series · best of 2 rows20.1%IndependentreasoningonPartially comparable-4.42 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124gemini-2-5-pro-05-06ClosedGoogle · Gemini 2.5 · best of 2 rows19.9%IndependentreasoningonPartially comparable-4.54 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124gpt-5-5-instant-06-26ClosedOpenAI · GPT 5.5 · best of 2 rows19.9%IndependentreasoningonPartially comparable-4.54 ptobs. 12 Sept 2026artificialanalysis.aiT2History
126gpt-oss-120bOpen weightsOpenAI · gpt-oss · best of 4 rows19.6%Independentreasoning_efforthighPartially comparable-4.87 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127Gemma 4 26B A4BOpen weightsGoogle · Gemma 4 · best of 4 rows19.3%IndependentreasoningonPartially comparable-5.15 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127grok-4-1-fastClosedSpaceXAI · Grok 4.1 · best of 4 rows19.3%IndependentreasoningonPartially comparable-5.15 ptobs. 12 Sept 2026artificialanalysis.aiT2History
129grok-4-fastClosedSpaceXAI · Grok 4 · best of 4 rows19.1%IndependentreasoningonPartially comparable-5.38 ptobs. 12 Sept 2026artificialanalysis.aiT2History
130Gemini 3.5 Flash-LiteClosedGoogle · Gemini 3.5 · best of 2 rows18.8%IndependentreasoningonPartially comparable-5.66 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131quasar-438bClosedMultiverse Computing · best of 2 rows18.7%Independentreasoning_effortmaxComparable-5.80 ptobs. 12 Sept 2026artificialanalysis.aiT2History
132k-exaone-2-0-0803Open weightsLG AI Research · EXAONE 2.0 · best of 2 rows18.6%IndependentreasoningonPartially comparable-5.89 ptobs. 12 Sept 2026artificialanalysis.aiT2History
133GPT-5.1-Codex MiniClosedOpenAI · GPT 5.1 · best of 2 rows18.5%Independentreasoning_efforthighPartially comparable-5.98 ptobs. 12 Sept 2026artificialanalysis.aiT2History
134gemini-2-5-pro-03-25ClosedGoogle · Gemini 2.5 · best of 2 rows18.0%IndependentreasoningonPartially comparable-6.44 ptobs. 12 Sept 2026artificialanalysis.aiT2History
135Claude Sonnet 4.5ClosedAnthropic · Claude · best of 4 rows17.8%IndependentreasoningonPartially comparable-6.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
135jt-4-1-flash-236b-a21bClosedChina Mobile · best of 2 rows17.8%IndependentreasoningoffPartially comparable-6.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
137g9v3-39a5bOpen weightsAI9Stars · best of 2 rows17.5%IndependentreasoningonPartially comparable-7.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
138Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.1 · best of 2 rows17.2%IndependentreasoningonPartially comparable-7.28 ptobs. 12 Sept 2026artificialanalysis.aiT2History
139GLM 5V TurboClosedZ.ai (Zhipu AI) · GLM5 · best of 2 rows17.1%IndependentreasoningonPartially comparable-7.32 ptobs. 12 Sept 2026artificialanalysis.aiT2History
139Mercury 2ClosedInception · best of 2 rows17.1%IndependentreasoningonPartially comparable-7.32 ptobs. 12 Sept 2026artificialanalysis.aiT2History
141hypernova-60bOpen weightsMultiverse Computing · best of 2 rows16.6%Independentreasoning_efforthighPartially comparable-7.83 ptobs. 12 Sept 2026artificialanalysis.aiT2History
142o4-miniClosedOpenAI · OpenAI o-series · best of 2 rows16.5%Independentreasoning_efforthighPartially comparable-7.93 ptobs. 12 Sept 2026artificialanalysis.aiT2History
143DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 4 rows16.4%IndependentreasoningonPartially comparable-8.11 ptobs. 12 Sept 2026artificialanalysis.aiT2History
144KAT-Coder-Pro V2ClosedKwaipilot · best of 2 rows16.1%IndependentreasoningoffPartially comparable-8.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
145Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 4 rows15.9%IndependentreasoningonPartially comparable-8.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
146Trinity Large ThinkingOpen weightsArcee AI · best of 2 rows15.8%IndependentreasoningonPartially comparable-8.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
146deepseek-r1Open weightsDeepSeek · DeepSeek-R1 · best of 2 rows15.8%IndependentreasoningonPartially comparable-8.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
148gemma-4-12BOpen weightsGoogle · Gemma 4 · best of 4 rows15.7%IndependentreasoningonPartially comparable-8.81 ptobs. 12 Sept 2026artificialanalysis.aiT2History
149Qwen3.5-9BOpen weightsQwen · Qwen3.5 · best of 4 rows14.9%IndependentreasoningonPartially comparable-9.55 ptobs. 12 Sept 2026artificialanalysis.aiT2History
150DeepSeek V3.2 ExpOpen weightsDeepSeek · DeepSeek-V3 · best of 3 rows14.9%IndependentreasoningonPartially comparable-9.60 ptobs. 12 Sept 2026artificialanalysis.aiT2History
150qwen3-5-omni-plusClosedAlibaba Group · Qwen3.5 · best of 2 rows14.9%IndependentreasoningoffPartially comparable-9.60 ptobs. 12 Sept 2026artificialanalysis.aiT2History
152GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.6 · best of 4 rows14.5%IndependentreasoningonPartially comparable-10.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
153DeepSeek V3.1Open weightsDeepSeek · DeepSeek-V3 · best of 4 rows14.3%IndependentreasoningonPartially comparable-10.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
154doubao-seed-codeClosedByteDance · Seed · best of 2 rows14.1%IndependentreasoningonPartially comparable-10.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
155k-exaoneOpen weightsLG AI Research · EXAONE · best of 4 rows13.9%IndependentreasoningonPartially comparable-10.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
156gemini-2-5-flash-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 6 rows13.8%IndependentreasoningonPartially comparable-10.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
157Mistral Medium 3.5Open weightsMistral AI · Mistral · best of 2 rows13.8%IndependentreasoningonPartially comparable-10.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
158MiniMax M2Open weightsMiniMax · MiniMax · best of 2 rows13.7%IndependentreasoningonPartially comparable-10.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
159ernie-5-0-thinking-previewClosedBaidu · ERNIE 5.0 · best of 2 rows13.3%IndependentreasoningonPartially comparable-11.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
160intellect-3Open weightsPrime Intellect · best of 2 rows13.1%IndependentreasoningonPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
161GLM 4.5Open weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows13.0%IndependentreasoningonPartially comparable-11.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
162exaone-4-5-33bOpen weightsLG AI Research · EXAONE 4.5 · best of 2 rows12.9%IndependentreasoningonPartially comparable-11.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
163qwen3-max-thinking-previewClosedAlibaba Group · Qwen3 · best of 2 rows12.7%IndependentreasoningonPartially comparable-11.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
164Qwen3 Next 80B A3B InstructOpen weightsQwen · Qwen3 · best of 4 rows12.6%IndependentreasoningonPartially comparable-11.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
165Claude Opus 4.1ClosedAnthropic · Claude · best of 2 rows12.5%IndependentreasoningonPartially comparable-12.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
166Claude Opus 4ClosedAnthropic · Claude12.3%Independentgroup defaultsPartially comparable-12.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
166claude-4-opusClosedAnthropic · Claude 4 · best of 3 rows12.3%IndependentreasoningonPartially comparable-12.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
168Gemini 2.5 FlashClosedGoogle · Gemini 2.5 · best of 2 rows12.1%IndependentreasoningonPartially comparable-12.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
168apriel-v1-5-15b-thinkerOpen weightsServiceNow · best of 2 rows12.1%IndependentreasoningonPartially comparable-12.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
170gemini-2-5-flash-reasoning-04-2025ClosedGoogle · Gemini 2.5 · best of 3 rows12.1%IndependentreasoningonPartially comparable-12.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
171o3-miniClosedOpenAI · OpenAI o-series · best of 4 rows12.0%Independentreasoning_efforthighPartially comparable-12.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
172cogito-v2-1-reasoningOpen weightsDeep Cogito · Cogito · best of 2 rows12%IndependentreasoningonPartially comparable-12.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
172nemotron-cascade-2-30b-a3bOpen weightsNVIDIA · Nemotron · best of 2 rows12%IndependentreasoningonPartially comparable-12.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
174command-a-plusOpen weightsCohere · Command · best of 2 rows12.0%IndependentreasoningonPartially comparable-12.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
175Qwen3 VL 235B A22B InstructOpen weightsQwen · Qwen3 · best of 4 rows11.9%IndependentreasoningonPartially comparable-12.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176nova-2-0-liteClosedAmazon Web Services · Nova 2.0 · best of 8 rows11.6%Independentreasoningonreasoning_efforthighPartially comparable-12.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
177exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 4 rows11.4%IndependentreasoningonPartially comparable-13.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
178Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 4 rows11.4%IndependentreasoningonPartially comparable-13.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
179granite-4.2-30bOpen weightsIBM · Granite 4.2 · best of 2 rows11.2%IndependentreasoningonPartially comparable-13.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
180ring-1tOpen weightsinclusionAI · best of 2 rows11.1%IndependentreasoningonPartially comparable-13.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
181North Mini Code (free)Open weightsCohere · best of 2 rows11.1%IndependentreasoningonPartially comparable-13.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
182sarvam-105bOpen weightsSarvam · best of 2 rows11.0%Independentreasoning_efforthighPartially comparable-13.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
183grok-3-mini-reasoningClosedSpaceXAI · Grok 3 · best of 2 rows11.0%Independentreasoningonreasoning_efforthighPartially comparable-13.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
184falcon-h1r-7bOpen weightsTII UAE · Falcon · best of 2 rows11.0%IndependentreasoningonPartially comparable-13.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
184gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 4 rows11.0%Independentreasoning_efforthighPartially comparable-13.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
184qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows11.0%IndependentreasoningonPartially comparable-13.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
187Hermes 4 405BOpen weightsNous Research · Hermes 410.9%Independentgroup defaultsPartially comparable-13.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
187hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.1 · best of 3 rows10.9%IndependentreasoningonPartially comparable-13.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
187nanbeige4-1-3bOpen weightsNanbeige · best of 2 rows10.9%IndependentreasoningonPartially comparable-13.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
190diffusiongemma-26b-a4bOpen weightsGoogle · best of 2 rows10.8%IndependentreasoningonPartially comparable-13.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
190step-3-vl-10bOpen weightsStepFun · Step3 · best of 2 rows10.8%IndependentreasoningonPartially comparable-13.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
192apriel-v1-6-15b-thinkerOpen weightsServiceNow · best of 2 rows10.8%IndependentreasoningonPartially comparable-13.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
193Claude Sonnet 4ClosedAnthropic · Claude · best of 4 rows10.7%IndependentreasoningonPartially comparable-13.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
194NVIDIA Nemotron 3.5 Lightning 30B A3BOpen weightsNVIDIA · Nemotron 3.5 · best of 2 rows10.6%IndependentreasoningonPartially comparable-13.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
195k2-v2Open weightsMBZUAI Institute of Foundation Models · best of 6 rows10.5%Independentreasoning_efforthighPartially comparable-13.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
196solar-open-100b-reasoningOpen weightsUpstage · Solar · best of 2 rows10.4%IndependentreasoningonPartially comparable-14.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
197Claude Haiku 4.5ClosedAnthropic · Claude · best of 4 rows10.4%IndependentreasoningonPartially comparable-14.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
198Magistral Medium 1.2ClosedMistral AI · Magistral · best of 2 rows10.3%IndependentreasoningonPartially comparable-14.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
198Solar Pro 3ClosedUpstage · Solar · best of 2 rows10.3%IndependentreasoningonPartially comparable-14.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
198qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 4 rows10.3%IndependentreasoningonPartially comparable-14.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →