Skip to content
AI Atlas
BenchmarkActivecategory · knowledgefamily · humanitys-last-exam · variant full

Humanity's Last Exam

lastexam.ai

expert-written frontier questions

data quality57

Updated 6 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,217
Models
453
Current leader
Claude Fable 5.1 59.1%

Frontier over time · accuracy · evaluator=Artificial Analysis

9 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 59.1%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  2. 58.7%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  3. 55.9%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  4. 55.5%Claude Fable 5 Anthropic Independent11 Sept 2026
  5. 53.1%gpt-6-astra OpenAI Independent11 Sept 2026
  6. 52.7%gpt-6-astra OpenAI Independent11 Sept 2026
  7. 51.3%Claude Opus 5 Anthropic Independent11 Sept 2026
  8. 11.0%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 453 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
401Qwen3 Coder 30B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows3.85%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
401Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 4 rows3.85%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
403qwen-2-5-maxClosedAlibaba Group · Qwen · best of 2 rows3.82%IndependentreasoningoffPartially comparable-0.03 ptobs. 12 Sept 2026artificialanalysis.aiT2History
404gemini-2-5-flash-04-2025ClosedGoogle · Gemini 2.53.80%Independentgroup defaultsPartially comparable-0.05 ptobs. 11 Sept 2026artificialanalysis.aiT2History
404granite-4-0-h-smallOpen weightsIBM · Granite 4.0 · best of 2 rows3.80%IndependentreasoningoffPartially comparable-0.05 ptobs. 12 Sept 2026artificialanalysis.aiT2History
404granite-4.1-8bOpen weightsIBM · Granite 4.1 · best of 2 rows3.80%IndependentreasoningoffPartially comparable-0.05 ptobs. 12 Sept 2026artificialanalysis.aiT2History
407Llama 4 ScoutRestricted weightsMeta AI · Llama 4 · best of 2 rows3.78%IndependentreasoningoffPartially comparable-0.07 ptobs. 12 Sept 2026artificialanalysis.aiT2History
408Mistral Small 3Open weightsMistral AI · Mistral · best of 2 rows3.76%IndependentreasoningoffPartially comparable-0.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
408Phi 4Open weightsMicrosoft · Phi4 · best of 2 rows3.76%IndependentreasoningoffPartially comparable-0.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
408deephermes-3-mistral-24b-previewOpen weightsNous Research · Mistral · best of 2 rows3.76%IndependentreasoningoffPartially comparable-0.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
411devstral-mediumClosedMistral AI · Devstral · best of 2 rows3.75%IndependentreasoningoffPartially comparable-0.10 ptobs. 12 Sept 2026artificialanalysis.aiT2History
411devstral-smallOpen weightsMistral AI · Devstral · best of 2 rows3.75%IndependentreasoningoffPartially comparable-0.10 ptobs. 12 Sept 2026artificialanalysis.aiT2History
411gpt-4.1-nanoClosedOpenAI · GPT 4.1 · best of 2 rows3.75%IndependentreasoningoffPartially comparable-0.10 ptobs. 12 Sept 2026artificialanalysis.aiT2History
411jamba-reasoning-3bOpen weightsAI21 Labs · Jamba · best of 2 rows3.75%IndependentreasoningonPartially comparable-0.10 ptobs. 12 Sept 2026artificialanalysis.aiT2History
415claude-35-sonnetClosedAnthropic · Claude 35 · best of 2 rows3.69%IndependentreasoningoffPartially comparable-0.16 ptobs. 12 Sept 2026artificialanalysis.aiT2History
416jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.7 · best of 2 rows3.66%IndependentreasoningoffPartially comparable-0.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
416qwen2-72b-instructOpen weightsAlibaba Group · Qwen2 · best of 2 rows3.66%IndependentreasoningoffPartially comparable-0.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
418Claude Haiku 3.5ClosedAnthropic · Claude · best of 2 rows3.63%IndependentreasoningoffPartially comparable-0.22 ptobs. 12 Sept 2026artificialanalysis.aiT2History
419claude-3-sonnetClosedAnthropic · Claude 3 · best of 2 rows3.62%IndependentreasoningoffPartially comparable-0.23 ptobs. 12 Sept 2026artificialanalysis.aiT2History
420gemma-3-270mOpen weightsGoogle · Gemma 3 · best of 2 rows3.61%IndependentreasoningoffPartially comparable-0.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
420jamba-1-6-largeOpen weightsAI21 Labs · Jamba 1.6 · best of 2 rows3.61%IndependentreasoningoffPartially comparable-0.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
420olmo-2-32bOpen weightsAllen Institute for AI · OLMo 2 · best of 2 rows3.61%IndependentreasoningoffPartially comparable-0.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
423Devstral 2Open weightsMistral AI · Devstral 2 · best of 2 rows3.57%IndependentreasoningoffPartially comparable-0.28 ptobs. 12 Sept 2026artificialanalysis.aiT2History
423o1-miniClosedOpenAI · OpenAI o-series · best of 2 rows3.57%IndependentreasoningonPartially comparable-0.28 ptobs. 12 Sept 2026artificialanalysis.aiT2History
425Llama 3.3 70BOpen weightsMeta AI · Llama 3.3 · best of 2 rows3.56%IndependentreasoningoffPartially comparable-0.29 ptobs. 12 Sept 2026artificialanalysis.aiT2History
425Qwen2.5 72B InstructOpen weightsQwen · Qwen2.5 · best of 2 rows3.56%IndependentreasoningoffPartially comparable-0.29 ptobs. 12 Sept 2026artificialanalysis.aiT2History
427claude-instantClosedAnthropic · Claude · best of 2 rows3.53%IndependentreasoningoffPartially comparable-0.32 ptobs. 12 Sept 2026artificialanalysis.aiT2History
428Qwen2.5 Coder 32B InstructOpen weightsQwen · Qwen2.5 · best of 2 rows3.51%IndependentreasoningoffPartially comparable-0.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
428mistral-mediumClosedMistral AI · Mistral · best of 2 rows3.51%IndependentreasoningoffPartially comparable-0.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
430Mistral LargeClosedMistral AI · Mistral · best of 2 rows3.50%IndependentreasoningoffPartially comparable-0.35 ptobs. 12 Sept 2026artificialanalysis.aiT2History
431Devstral Small 2Open weightsMistral AI · Devstral · best of 2 rows3.48%IndependentreasoningoffPartially comparable-0.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
431gemini-1-5-pro-may-2024ClosedGoogle · Gemini 1.5 · best of 2 rows3.48%IndependentreasoningoffPartially comparable-0.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
433granite-4.1-3bOpen weightsIBM · Granite 4.1 · best of 2 rows3.43%IndependentreasoningoffPartially comparable-0.42 ptobs. 12 Sept 2026artificialanalysis.aiT2History
434gemini-2-0-flash-lite-001ClosedGoogle · Gemini 2.0 · best of 2 rows3.35%IndependentreasoningoffPartially comparable-0.50 ptobs. 12 Sept 2026artificialanalysis.aiT2History
435ernie-4-5-300b-a47bOpen weightsBaidu · ERNIE 4.5 · best of 2 rows3.34%IndependentreasoningoffPartially comparable-0.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
435mistral-large-2Open weightsMistral AI · Mistral · best of 2 rows3.34%IndependentreasoningoffPartially comparable-0.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
437tulu3-405bOpen weightsAllen Institute for AI · best of 2 rows3.28%IndependentreasoningoffPartially comparable-0.57 ptobs. 12 Sept 2026artificialanalysis.aiT2History
438claude-35-sonnet-june-24ClosedAnthropic · Claude 35 · best of 2 rows3.23%IndependentreasoningoffPartially comparable-0.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
439nova-proClosedAmazon Web Services · Nova · best of 2 rows3.19%IndependentreasoningoffPartially comparable-0.66 ptobs. 12 Sept 2026artificialanalysis.aiT2History
440gemini-1-5-flashClosedGoogle · Gemini 1.5 · best of 2 rows3.17%IndependentreasoningoffPartially comparable-0.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
441gpt-4-turboClosedOpenAI · GPT 4 · best of 2 rows3.13%IndependentreasoningoffPartially comparable-0.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
442sarvam-m-reasoningOpen weightsSarvam · best of 2 rows3.10%IndependentreasoningonPartially comparable-0.75 ptobs. 12 Sept 2026artificialanalysis.aiT2History
443DeepSeek-R1-Distill-Qwen-1.5BOpen weightsDeepSeek · Qwen · best of 2 rows3.08%IndependentreasoningonPartially comparable-0.77 ptobs. 12 Sept 2026artificialanalysis.aiT2History
444grok-2Open weightsxAI · Grok 2 · best of 2 rows3.06%IndependentreasoningoffPartially comparable-0.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
445Mistral Large 2.0Open weightsMistral AI · Mistral · best of 2 rows2.95%IndependentreasoningoffPartially comparable-0.90 ptobs. 12 Sept 2026artificialanalysis.aiT2History
446dbrxOpen weightsDatabricks · best of 2 rows2.92%IndependentreasoningoffPartially comparable-0.93 ptobs. 12 Sept 2026artificialanalysis.aiT2History
447Pixtral LargeOpen weightsMistral AI · Pixtral · best of 2 rows2.80%IndependentreasoningoffPartially comparable-1.05 ptobs. 12 Sept 2026artificialanalysis.aiT2History
448claude-3-opusClosedAnthropic · Claude 3 · best of 2 rows2.75%IndependentreasoningoffPartially comparable-1.10 ptobs. 12 Sept 2026artificialanalysis.aiT2History
449gpt-4o-chatgptClosedOpenAI · GPT 4 · best of 2 rows2.73%IndependentreasoningoffPartially comparable-1.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
450Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AI · Kimi · best of 2 rows2.46%IndependentreasoningoffPartially comparable-1.39 ptobs. 12 Sept 2026artificialanalysis.aiT2History
451gpt-4oClosedOpenAI · GPT 4 · best of 2 rows2.42%IndependentreasoningoffPartially comparable-1.43 ptobs. 12 Sept 2026artificialanalysis.aiT2History
452GPT-4o (2024-08-06)ClosedOpenAI · GPT 4 · best of 2 rows2.33%IndependentreasoningoffPartially comparable-1.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
453GPT-4o (2024-05-13)ClosedOpenAI · GPT 4 · best of 2 rows1.76%IndependentreasoningoffPartially comparable-2.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →