Skip to content
AI Atlas
BenchmarkActivecategory · knowledgefamily · humanitys-last-exam · variant full

Humanity's Last Exam

lastexam.ai

expert-written frontier questions

data quality57

Updated 3 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,217
Models
453
Current leader
Claude Fable 5.1 59.1%

Score history · Claude Opus 4.6 4 rows

Score history for Claude Opus 4.60%10%20%30%40%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • Claude Opus 4.6
  • 19.09%aa_slug=claude-opus-4-6 · evaluator=Artificial Analysis · reasoning=off · index_version=4.312 Sept 2026
  • 39.94%aa_slug=claude-opus-4-6-adaptive · evaluator=Artificial Analysis · reasoning=adaptive · index_version=4.312 Sept 2026
  • 19.09%aa_slug=claude-opus-4-6 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 39.94%aa_slug=claude-opus-4-6-adaptive · evaluator=Artificial Analysis · reasoning=adaptive · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · evaluator=Artificial Analysis

9 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 59.1%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  2. 58.7%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  3. 55.9%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  4. 55.5%Claude Fable 5 Anthropic Independent11 Sept 2026
  5. 53.1%gpt-6-astra OpenAI Independent11 Sept 2026
  6. 52.7%gpt-6-astra OpenAI Independent11 Sept 2026
  7. 51.3%Claude Opus 5 Anthropic Independent11 Sept 2026
  8. 11.0%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 453 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1Claude Fable 5.1ClosedAnthropic · Claude · best of 10 rows59.1%Independentreasoning_effortmaxleaderobs. 12 Sept 2026artificialanalysis.aiT2History
2Claude Fable 5ClosedAnthropic · Claude · best of 2 rows55.5%IndependentreasoningonPartially comparable-3.66 ptobs. 12 Sept 2026artificialanalysis.aiT2History
3Claude Opus 5ClosedAnthropic · Claude · best of 10 rows54.9%Independentreasoning_effortmaxComparable-4.26 ptobs. 12 Sept 2026artificialanalysis.aiT2History
4gpt-6-astraClosedOpenAI · GPT 6 · best of 11 rows54.7%Independentreasoning_effortmaxComparable-4.45 ptobs. 12 Sept 2026artificialanalysis.aiT2History
5gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 12 rows49.5%Independentreasoning_effortmaxComparable-9.64 ptobs. 12 Sept 2026artificialanalysis.aiT2History
6Muse Spark 1.3ClosedMeta AI · best of 4 rows48.7%Independentreasoning_effortmaxComparable-10.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
7Claude Opus 4.8ClosedAnthropic · Claude · best of 2 rows48.7%Independentreasoning_effortmaxComparable-10.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
8Gemini 3.7 FlashClosedGoogle · Gemini 3.7 · best of 6 rows47.9%Independentreasoning_efforthighPartially comparable-11.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
9Gemini 3.8 FlashClosedGoogle · Gemini 3.8 · best of 6 rows47.8%Independentreasoning_efforthighPartially comparable-11.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
10Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.1 · best of 2 rows47.0%IndependentreasoningonPartially comparable-12.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
11Kimi K3Open weightsMoonshot AI · Kimi · best of 4 rows46.9%Independentreasoning_effortmaxComparable-12.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
12Muse Spark 1.1ClosedMeta AI · best of 2 rows46.2%Independentreasoning_effortxhighPartially comparable-12.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
13gpt-5.5ClosedOpenAI · GPT 5.5 · best of 10 rows45.8%Independentreasoning_effortxhighPartially comparable-13.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
14Muse Spark 1.2ClosedMeta AI · best of 2 rows45.5%Independentreasoning_effortxhighPartially comparable-13.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
15Grok 4.6ClosedxAI · Grok · best of 8 rows44.1%Independentreasoning_effortxhighPartially comparable-15.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
16gpt-5.4ClosedOpenAI · GPT 5.4 · best of 6 rows43.7%Independentreasoning_effortxhighPartially comparable-15.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
17Qwen 3.8 MaxClosedQwen · Qwen3.8 · best of 2 rows43.0%IndependentreasoningonPartially comparable-16.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
18gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 12 rows42.9%Independentreasoning_effortmaxComparable-16.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
19Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 6 rows42.7%IndependentreasoningonPartially comparable-16.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
19Grok 4.5ClosedxAI · Grok · best of 2 rows42.7%Independentreasoning_efforthighPartially comparable-16.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
21gpt-5.3-codexClosedOpenAI · GPT 5.3 · best of 2 rows42.5%Independentreasoning_effortxhighPartially comparable-16.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
22Qwen3.8 2.4T A95BOpen weightsQwen · Qwen3.8 · best of 2 rows42.5%IndependentreasoningonPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
23Claude Opus 4.7ClosedAnthropic · Claude · best of 4 rows42.3%Independentreasoning_effortmaxComparable-16.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
24GLM 5.3Open weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows42.3%Independentreasoning_effortmaxComparable-16.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
25Claude Sonnet 5ClosedAnthropic · Claude · best of 12 rows41.3%Independentreasoning_effortmaxComparable-17.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
26Z.ai GLM 5.2Open weightsZ.ai (Zhipu AI) · GLM5.2 · best of 4 rows41.1%Independentreasoning_effortmaxComparable-18.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
27deepseek-v4-proClosedDeepSeek · V4 · best of 2 rows41.0%Independentreasoning_effortmaxComparable-18.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
28Gemini 3.6 FlashClosedGoogle · Gemini 3.6 · best of 2 rows40.8%IndependentreasoningonPartially comparable-18.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
29muse-sparkClosedMeta AI · best of 2 rows40.7%IndependentreasoningonPartially comparable-18.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
30Qwen3.7 MaxClosedQwen · Qwen3.7 · best of 2 rows40.5%IndependentreasoningonPartially comparable-18.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
31motif-0714ClosedMotif Technologies · best of 2 rows40.4%IndependentreasoningonPartially comparable-18.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
32Claude Opus 4.6ClosedAnthropic · Claude · best of 4 rows39.9%Independentreasoningadaptivereasoning_effortmaxPartially comparable-19.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
33GLM 5.3 FlashOpen weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows39.9%IndependentreasoningonPartially comparable-19.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
34gemini-3-proClosedGoogle · Gemini 3 · best of 4 rows39.7%Independentreasoning_efforthighPartially comparable-19.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35gpt-5.6-lunaClosedOpenAI · GPT 5.6 · best of 12 rows39.5%Independentreasoning_effortmaxComparable-19.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
36DeepSeek-V4.1-FlashOpen weightsDeepSeek · DeepSeek · best of 2 rows39.3%Independentreasoning_effortmaxComparable-19.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
37MiniMax-M3Open weightsMiniMax · MiniMax · best of 2 rows39.0%IndependentreasoningonPartially comparable-20.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
38deepseek-v4-flashOpen weightsDeepSeek · DeepSeek · best of 2 rows38.5%Independentreasoning_effortmaxComparable-20.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
39agnes-3-0-flashClosedSapiens AI · best of 2 rows38.5%IndependentreasoningonPartially comparable-20.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
40grok-build-0-1-06-16ClosedSpaceXAI · Grok · best of 2 rows38.3%IndependentreasoningonPartially comparable-20.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
41Qwen3.8 FlashOpen weightsQwen · Qwen3.8 · best of 2 rows38.0%IndependentreasoningonPartially comparable-21.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
42gpt-5.2ClosedOpenAI · GPT 5.2 · best of 6 rows37.7%Independentreasoning_effortxhighPartially comparable-21.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
43agnes-2-5-pro-betaClosedSapiens AI · best of 2 rows37.5%IndependentreasoningonPartially comparable-21.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
43deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek · best of 4 rows37.5%Independentreasoning_effortmaxComparable-21.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
45Kimi K2.6Open weightsMoonshot AI · Kimi · best of 4 rows37.5%IndependentreasoningonPartially comparable-21.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
46Grok 4.3ClosedxAI · Grok · best of 8 rows37.2%Independentreasoning_efforthighPartially comparable-21.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
47motif-3Open weightsMotif Technologies · best of 2 rows37.0%IndependentreasoningonPartially comparable-22.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
48Gemini 3 Flash PreviewClosedGoogle · Gemini 336.6%Independentgroup defaultsPartially comparable-22.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
48gemini-3-flashClosedGoogle · Gemini 3 · best of 3 rows36.6%IndependentreasoningonPartially comparable-22.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
50GPT-5.2-CodexClosedOpenAI · GPT 5.2 · best of 2 rows35.7%Independentreasoning_effortxhighPartially comparable-23.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
51MiMo-V2.5-ProOpen weightsXiaomi · best of 4 rows35.7%IndependentreasoningonPartially comparable-23.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
52Qwen 3.7 PlusClosedQwen · Qwen3.7 · best of 2 rows35.6%IndependentreasoningonPartially comparable-23.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
53deepseek-v4-pro-0424-highOpen weightsDeepSeek · DeepSeek35.2%Independentgroup defaultsPartially comparable-23.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
54Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi · best of 2 rows35.0%IndependentreasoningonPartially comparable-24.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
55deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek · best of 4 rows34.9%Independentreasoning_effortmaxComparable-24.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
56Grok 4.20ClosedxAI · Grok · best of 3 rows34.5%IndependentreasoningonPartially comparable-24.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
57deepseek-v4-flash-visionClosedDeepSeek · DeepSeek · best of 2 rows34.5%Independentreasoning_effortmaxComparable-24.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
58apodex-1-1ClosedApodex · best of 2 rows34.1%IndependentreasoningonPartially comparable-25.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
59Qwen3.8 27BOpen weightsQwen · Qwen3.8 · best of 8 rows33.9%Independentreasoning_effortxhighPartially comparable-25.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60nex-n2-proOpen weightsNex AGI · best of 2 rows33.7%IndependentreasoningonPartially comparable-25.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
61LongCat 2.0Open weightsMeituan · best of 2 rows33.7%IndependentreasoningonPartially comparable-25.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
62kat-coder-pro-v1ClosedKwaiKAT · best of 2 rows33.6%IndependentreasoningoffPartially comparable-25.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
63Claude Sonnet 4.6ClosedAnthropic · Claude · best of 5 rows33.6%Independentreasoningadaptivereasoning_effortmaxPartially comparable-25.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
63agnes-2-5-pro-alphaOpen weightsSapiens AI · best of 2 rows33.6%IndependentreasoningonPartially comparable-25.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
65Hy3Open weightsTencent · best of 5 rows33.5%IndependentreasoningonPartially comparable-25.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
66Inkling SmallOpen weightsThinking Machines · best of 2 rows33.3%IndependentreasoningonPartially comparable-25.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
67grok-4-20-0309ClosedSpaceXAI · Grok 4.20 · best of 4 rows32.4%IndependentreasoningonPartially comparable-26.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
68k2-horizon-375b-a23bOpen weightsMBZUAI Institute of Foundation Models · best of 2 rows32.0%IndependentreasoningonPartially comparable-27.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
69InklingOpen weightsThinking Machines · best of 2 rows31.9%IndependentreasoningonPartially comparable-27.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
70Qwen3.6 Max PreviewClosedQwen · Qwen3.6 · best of 2 rows30.8%IndependentreasoningonPartially comparable-28.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
71Kimi K2.5Open weightsMoonshot AI · Kimi · best of 4 rows30.7%IndependentreasoningonPartially comparable-28.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
72mimo-v2-proClosedXiaomi · best of 2 rows30.4%IndependentreasoningonPartially comparable-28.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
73deepseek-v4-flash-0420-highOpen weightsDeepSeek · DeepSeek30.3%Independentgroup defaultsPartially comparable-28.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
74Claude Opus 4.5ClosedAnthropic · Claude · best of 4 rows30.1%IndependentreasoningonPartially comparable-29.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
75GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 4 rows30.1%IndependentreasoningonPartially comparable-29.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
76MiniMax M2.7Open weightsMiniMax · MiniMax · best of 2 rows29.6%IndependentreasoningonPartially comparable-29.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
77a-x-k2Open weightsSK Telecom · best of 2 rows29.6%IndependentreasoningonPartially comparable-29.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
78GLM 5Open weightsZ.ai (Zhipu AI) · GLM5 · best of 4 rows29.3%IndependentreasoningonPartially comparable-29.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79Solar Pro 4ClosedUpstage · Solar · best of 2 rows29.2%IndependentreasoningonPartially comparable-29.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
80Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 4 rows29.0%IndependentreasoningonPartially comparable-30.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
81deepseek-v3-2-specialeOpen weightsDeepSeek · DeepSeek · best of 2 rows28.7%IndependentreasoningonPartially comparable-30.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
82solar-open2-250bOpen weightsUpstage · Solar · best of 2 rows28.5%IndependentreasoningonPartially comparable-30.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
83gpt-5ClosedOpenAI · GPT 5 · best of 8 rows28.5%Independentreasoning_efforthighPartially comparable-30.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
83gpt-5.1ClosedOpenAI · GPT 5.1 · best of 4 rows28.5%Independentreasoning_efforthighPartially comparable-30.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
85Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 3 · best of 2 rows28.4%IndependentreasoningonPartially comparable-30.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
86gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 6 rows28.3%Independentreasoning_effortxhighPartially comparable-30.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
87gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 6 rows28.1%Independentreasoning_effortxhighPartially comparable-31.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88Qwen3 MaxClosedQwen · Qwen3 · best of 4 rows28.0%IndependentreasoningonPartially comparable-31.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
89grok-4.20-0309-non-reasoningClosedxAI · Grok27.9%Independentgroup defaultsPartially comparable-31.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
90GPT-5-CodexClosedOpenAI · GPT 5 · best of 2 rows27.9%Independentreasoning_efforthighPartially comparable-31.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
90Qwen3.6 PlusClosedQwen · Qwen3.6 · best of 2 rows27.9%IndependentreasoningonPartially comparable-31.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
92GLM 5 TurboClosedZ.ai (Zhipu AI) · GLM5 · best of 2 rows27.8%IndependentreasoningonPartially comparable-31.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
92Hy3 previewOpen weightsTencent27.8%Independentgroup defaultsPartially comparable-31.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
94GLM 4.7Open weightsZ.ai (Zhipu AI) · GLM4.7 · best of 4 rows27.4%IndependentreasoningonPartially comparable-31.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
95MiMo-V2.5Open weightsXiaomi · best of 2 rows27.2%IndependentreasoningonPartially comparable-31.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
96grok-4ClosedSpaceXAI · Grok 4 · best of 2 rows26.7%IndependentreasoningonPartially comparable-32.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
97GPT-5.1-CodexClosedOpenAI · GPT 5.1 · best of 2 rows25.7%Independentreasoning_efforthighPartially comparable-33.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
98Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 4 rows25.2%IndependentreasoningonPartially comparable-33.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
99DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 5 rows24.6%IndependentreasoningonPartially comparable-34.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
99DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V324.6%Independentgroup defaultsPartially comparable-34.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →