Skip to content
AI Atlas
BenchmarkActivecategory · knowledgefamily · humanitys-last-exam · variant full

Humanity's Last Exam

lastexam.ai

expert-written frontier questions

data quality57

Updated 5 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,217
Models
453
Current leader
Claude Fable 5.1 59.1%

Frontier over time · accuracy · evaluator=Artificial Analysis

9 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 59.1%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  2. 58.7%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  3. 55.9%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  4. 55.5%Claude Fable 5 Anthropic Independent11 Sept 2026
  5. 53.1%gpt-6-astra OpenAI Independent11 Sept 2026
  6. 52.7%gpt-6-astra OpenAI Independent11 Sept 2026
  7. 51.3%Claude Opus 5 Anthropic Independent11 Sept 2026
  8. 11.0%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 453 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
301llama-2-chat-70bOpen weightsMeta AI · Llama 2 · best of 2 rows5.21%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
302llama-3-1-nemotron-nano-4b-reasoningOpen weightsNVIDIA · Llama 3.1 · best of 2 rows5.19%IndependentreasoningonPartially comparable-0.02 ptobs. 12 Sept 2026artificialanalysis.aiT2History
303jamba-1-5-miniOpen weightsAI21 Labs · Jamba 1.5 · best of 2 rows5.14%IndependentreasoningoffPartially comparable-0.07 ptobs. 12 Sept 2026artificialanalysis.aiT2History
303molmo-7b-dOpen weightsAllen Institute for AI · Molmo · best of 2 rows5.14%IndependentreasoningoffPartially comparable-0.07 ptobs. 12 Sept 2026artificialanalysis.aiT2History
305R1 Distill Llama 70BOpen weightsDeepSeek · Llama · best of 2 rows5.13%IndependentreasoningonPartially comparable-0.08 ptobs. 12 Sept 2026artificialanalysis.aiT2History
306LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.5 · best of 2 rows5.10%IndependentreasoningoffPartially comparable-0.11 ptobs. 12 Sept 2026artificialanalysis.aiT2History
306Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 4 rows5.10%IndependentreasoningoffPartially comparable-0.11 ptobs. 12 Sept 2026artificialanalysis.aiT2History
308llama-3-instruct-8bOpen weightsMeta AI · Llama 3 · best of 2 rows5.06%IndependentreasoningoffPartially comparable-0.15 ptobs. 12 Sept 2026artificialanalysis.aiT2History
309ling-mini-2-0Open weightsinclusionAI · best of 2 rows5.05%IndependentreasoningoffPartially comparable-0.16 ptobs. 12 Sept 2026artificialanalysis.aiT2History
309minicpm-v4-6-1-3bOpen weightsOpenBMB · best of 2 rows5.05%IndependentreasoningoffPartially comparable-0.16 ptobs. 12 Sept 2026artificialanalysis.aiT2History
311gpt-4.1-miniClosedOpenAI · GPT 4.1 · best of 2 rows5.02%IndependentreasoningoffPartially comparable-0.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
312Granite 4.0 MicroOpen weightsIBM · Granite 4.0 · best of 2 rows5%IndependentreasoningoffPartially comparable-0.21 ptobs. 12 Sept 2026artificialanalysis.aiT2History
312granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.0 · best of 2 rows5%IndependentreasoningoffPartially comparable-0.21 ptobs. 12 Sept 2026artificialanalysis.aiT2History
312phi-4-multimodalOpen weightsMicrosoft · Phi4 · best of 2 rows5%IndependentreasoningoffPartially comparable-0.21 ptobs. 12 Sept 2026artificialanalysis.aiT2History
315Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 4 rows4.96%IndependentreasoningoffPartially comparable-0.25 ptobs. 12 Sept 2026artificialanalysis.aiT2History
315apertus-8b-instructOpen weightsSwiss AI Initiative · best of 2 rows4.96%IndependentreasoningoffPartially comparable-0.25 ptobs. 12 Sept 2026artificialanalysis.aiT2History
315phi-3-miniOpen weightsMicrosoft · Phi3 · best of 2 rows4.96%IndependentreasoningoffPartially comparable-0.25 ptobs. 12 Sept 2026artificialanalysis.aiT2History
318lfm-40bClosedLiquid AI · LFM · best of 2 rows4.93%IndependentreasoningoffPartially comparable-0.28 ptobs. 12 Sept 2026artificialanalysis.aiT2History
319Llama 4 MaverickOpen weightsMeta AI · Llama 4 · best of 2 rows4.91%IndependentreasoningoffPartially comparable-0.30 ptobs. 12 Sept 2026artificialanalysis.aiT2History
319lfm2-8b-a1bOpen weightsLiquid AI · LFM2 · best of 2 rows4.91%IndependentreasoningoffPartially comparable-0.30 ptobs. 12 Sept 2026artificialanalysis.aiT2History
321Qwen2.5-Coder-7BOpen weightsQwen · Qwen2.5 · best of 2 rows4.89%IndependentreasoningoffPartially comparable-0.32 ptobs. 12 Sept 2026artificialanalysis.aiT2History
322SonarClosedPerplexity AI · Sonar · best of 2 rows4.88%IndependentreasoningoffPartially comparable-0.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
323NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 4 rows4.87%IndependentreasoningonPartially comparable-0.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
323nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows4.87%IndependentreasoningonPartially comparable-0.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
325gemma-3n-e4b-preview-0520Open weightsGoogle · Gemma 3 · best of 2 rows4.82%IndependentreasoningoffPartially comparable-0.39 ptobs. 12 Sept 2026artificialanalysis.aiT2History
325gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 4 rows4.82%IndependentreasoningoffPartially comparable-0.39 ptobs. 12 Sept 2026artificialanalysis.aiT2History
325granite-4-0-nano-1bOpen weightsIBM · Granite 4.0 · best of 2 rows4.82%IndependentreasoningoffPartially comparable-0.39 ptobs. 12 Sept 2026artificialanalysis.aiT2History
325nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows4.82%IndependentreasoningonPartially comparable-0.39 ptobs. 12 Sept 2026artificialanalysis.aiT2History
329command-r-03-2024Open weightsCohere · Command · best of 2 rows4.79%IndependentreasoningoffPartially comparable-0.42 ptobs. 12 Sept 2026artificialanalysis.aiT2History
329openchat-35Open weightsOpenChat · best of 2 rows4.79%IndependentreasoningoffPartially comparable-0.42 ptobs. 12 Sept 2026artificialanalysis.aiT2History
331gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 4 rows4.77%IndependentreasoningonPartially comparable-0.44 ptobs. 12 Sept 2026artificialanalysis.aiT2History
331llama-2-chat-13bOpen weightsMeta AI · Llama 2 · best of 2 rows4.77%IndependentreasoningoffPartially comparable-0.44 ptobs. 12 Sept 2026artificialanalysis.aiT2History
333DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows4.74%IndependentreasoningoffPartially comparable-0.47 ptobs. 12 Sept 2026artificialanalysis.aiT2History
334gemini-1-5-flash-8bClosedGoogle · Gemini 1.5 · best of 2 rows4.69%IndependentreasoningoffPartially comparable-0.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
335Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows4.68%IndependentreasoningoffPartially comparable-0.53 ptobs. 12 Sept 2026artificialanalysis.aiT2History
335Mixtral 8x7BOpen weightsMistral AI · Mixtral 8 · best of 2 rows4.68%IndependentreasoningoffPartially comparable-0.53 ptobs. 12 Sept 2026artificialanalysis.aiT2History
337gemini-1-5-proClosedGoogle · Gemini 1.5 · best of 2 rows4.64%IndependentreasoningoffPartially comparable-0.57 ptobs. 12 Sept 2026artificialanalysis.aiT2History
338Ministral 3 14BOpen weightsMistral AI · Ministral 3 · best of 2 rows4.63%IndependentreasoningoffPartially comparable-0.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
338command-r-plus-04-2024Open weightsCohere · Command · best of 2 rows4.63%IndependentreasoningoffPartially comparable-0.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
338nova-microClosedAmazon Web Services · Nova · best of 2 rows4.63%IndependentreasoningoffPartially comparable-0.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
341Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 4 rows4.59%IndependentreasoningonPartially comparable-0.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
342DeepSeek-R1-Distill-Qwen-32BOpen weightsDeepSeek · Qwen · best of 2 rows4.58%IndependentreasoningonPartially comparable-0.63 ptobs. 12 Sept 2026artificialanalysis.aiT2History
343Mistral 7BOpen weightsMistral AI · Mistral · best of 2 rows4.55%IndependentreasoningoffPartially comparable-0.66 ptobs. 12 Sept 2026artificialanalysis.aiT2History
344qwen3-coder-480b-a35b-instructOpen weightsAlibaba Group · Qwen3-Coder · best of 2 rows4.54%IndependentreasoningoffPartially comparable-0.67 ptobs. 12 Sept 2026artificialanalysis.aiT2History
345Qwen3 14BOpen weightsQwen · Qwen34.53%Independentgroup defaultsPartially comparable-0.68 ptobs. 11 Sept 2026artificialanalysis.aiT2History
345qwen3-14b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows4.53%IndependentreasoningonPartially comparable-0.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
347llama-3-instruct-70bOpen weightsMeta AI · Llama 3 · best of 2 rows4.52%IndependentreasoningoffPartially comparable-0.69 ptobs. 12 Sept 2026artificialanalysis.aiT2History
348g9v3-3bOpen weightsAI9Stars · best of 2 rows4.49%IndependentreasoningonPartially comparable-0.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
348gemma-3n-e4bOpen weightsGoogle · Gemma 3 · best of 2 rows4.49%IndependentreasoningoffPartially comparable-0.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
348llama-3-2-instruct-90b-visionOpen weightsMeta AI · Llama 3.2 · best of 2 rows4.49%IndependentreasoningoffPartially comparable-0.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
351Llama-3.1-70BOpen weightsMeta AI · Llama 3.1 · best of 2 rows4.47%IndependentreasoningoffPartially comparable-0.74 ptobs. 12 Sept 2026artificialanalysis.aiT2History
352grok-betaClosedSpaceXAI · Grok · best of 2 rows4.45%IndependentreasoningoffPartially comparable-0.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
352jamba-1-7-miniOpen weightsAI21 Labs · Jamba 1.7 · best of 2 rows4.45%IndependentreasoningoffPartially comparable-0.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
352phi-4-miniOpen weightsMicrosoft · Phi4 · best of 2 rows4.45%IndependentreasoningoffPartially comparable-0.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
355Reka Flash 3Open weightsrekaai · best of 2 rows4.44%IndependentreasoningonPartially comparable-0.77 ptobs. 12 Sept 2026artificialanalysis.aiT2History
356Qwen3-4BOpen weightsQwen · Qwen34.42%Independentgroup defaultsPartially comparable-0.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
356qwen3-4b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows4.42%IndependentreasoningonPartially comparable-0.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
358Gemma 3 27BOpen weightsGoogle · Gemma 3 · best of 2 rows4.40%IndependentreasoningoffPartially comparable-0.81 ptobs. 12 Sept 2026artificialanalysis.aiT2History
359gemini-1-5-flash-may-2024ClosedGoogle · Gemini 1.5 · best of 2 rows4.36%IndependentreasoningoffPartially comparable-0.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
360Mistral Small 3.1Open weightsMistral AI · Mistral · best of 2 rows4.31%IndependentreasoningoffPartially comparable-0.90 ptobs. 12 Sept 2026artificialanalysis.aiT2History
360Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows4.31%IndependentreasoningoffPartially comparable-0.90 ptobs. 12 Sept 2026artificialanalysis.aiT2History
360jamba-1-6-miniOpen weightsAI21 Labs · Jamba 1.6 · best of 2 rows4.31%IndependentreasoningoffPartially comparable-0.90 ptobs. 12 Sept 2026artificialanalysis.aiT2History
363Mistral SabaClosedMistral AI · Mistral · best of 2 rows4.30%IndependentreasoningoffPartially comparable-0.91 ptobs. 12 Sept 2026artificialanalysis.aiT2History
364nova-liteClosedAmazon Web Services · Nova · best of 2 rows4.29%IndependentreasoningoffPartially comparable-0.92 ptobs. 12 Sept 2026artificialanalysis.aiT2History
365Gemini 2.0 FlashClosedGoogle · Gemini 2.0 · best of 2 rows4.26%IndependentreasoningoffPartially comparable-0.95 ptobs. 12 Sept 2026artificialanalysis.aiT2History
365Ministral 3 8BOpen weightsMistral AI · Ministral 3 · best of 2 rows4.26%IndependentreasoningoffPartially comparable-0.95 ptobs. 12 Sept 2026artificialanalysis.aiT2History
365Molmo2-8BOpen weightsAllen Institute for AI · best of 2 rows4.26%IndependentreasoningoffPartially comparable-0.95 ptobs. 12 Sept 2026artificialanalysis.aiT2History
365deephermes-3-llama-3-1-8b-previewOpen weightsNous Research · Llama 3.1 · best of 2 rows4.26%IndependentreasoningoffPartially comparable-0.95 ptobs. 12 Sept 2026artificialanalysis.aiT2History
369mistral-smallOpen weightsMistral AI · Mistral · best of 2 rows4.25%IndependentreasoningoffPartially comparable-0.96 ptobs. 12 Sept 2026artificialanalysis.aiT2History
370lfm2-24b-a2bOpen weightsLiquid AI · LFM2 · best of 2 rows4.22%IndependentreasoningoffPartially comparable-0.99 ptobs. 12 Sept 2026artificialanalysis.aiT2History
371Gemma 3 12BOpen weightsGoogle · Gemma 3 · best of 2 rows4.21%IndependentreasoningoffPartially comparable-1.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
372gpt-4o-miniClosedOpenAI · GPT 4 · best of 2 rows4.20%IndependentreasoningoffPartially comparable-1.01 ptobs. 12 Sept 2026artificialanalysis.aiT2History
373gemini-1-0-proClosedGoogle · Gemini 1.0 · best of 2 rows4.19%IndependentreasoningoffPartially comparable-1.02 ptobs. 12 Sept 2026artificialanalysis.aiT2History
373llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.1 · best of 2 rows4.19%IndependentreasoningoffPartially comparable-1.02 ptobs. 12 Sept 2026artificialanalysis.aiT2History
375gpt-4.1ClosedOpenAI · GPT 4.1 · best of 2 rows4.18%IndependentreasoningoffPartially comparable-1.03 ptobs. 12 Sept 2026artificialanalysis.aiT2History
376Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows4.17%IndependentreasoningoffPartially comparable-1.04 ptobs. 12 Sept 2026artificialanalysis.aiT2History
376gemma-3n-e2bOpen weightsGoogle · Gemma 3 · best of 2 rows4.17%IndependentreasoningoffPartially comparable-1.04 ptobs. 12 Sept 2026artificialanalysis.aiT2History
376granite-3-3-8b-instructOpen weightsIBM · Granite 3.3 · best of 2 rows4.17%IndependentreasoningoffPartially comparable-1.04 ptobs. 12 Sept 2026artificialanalysis.aiT2History
376nova-premierClosedAmazon Web Services · Nova · best of 2 rows4.17%IndependentreasoningoffPartially comparable-1.04 ptobs. 12 Sept 2026artificialanalysis.aiT2History
380deepseek-r1-distill-qwen-14bOpen weightsDeepSeek · Qwen · best of 2 rows4.13%IndependentreasoningonPartially comparable-1.08 ptobs. 12 Sept 2026artificialanalysis.aiT2History
381granite-4.1-30bOpen weightsIBM · Granite 4.1 · best of 2 rows4.12%IndependentreasoningoffPartially comparable-1.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
381grok-3ClosedSpaceXAI · Grok 3 · best of 2 rows4.12%IndependentreasoningoffPartially comparable-1.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
383Mistral Small 1.0ClosedMistral AI · Mistral · best of 2 rows4.09%IndependentreasoningoffPartially comparable-1.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
383gemini-2-0-flash-lite-previewClosedGoogle · Gemini 2.0 · best of 2 rows4.09%IndependentreasoningoffPartially comparable-1.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
383qwen-turboClosedAlibaba Group · Qwen · best of 2 rows4.09%IndependentreasoningoffPartially comparable-1.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
386Claude 3 HaikuClosedAnthropic · Claude · best of 2 rows4.08%IndependentreasoningoffPartially comparable-1.13 ptobs. 12 Sept 2026artificialanalysis.aiT2History
387gemini-2.0-flash-expClosedGoogle · Gemini 2.0 · best of 2 rows4.07%IndependentreasoningoffPartially comparable-1.14 ptobs. 12 Sept 2026artificialanalysis.aiT2History
387jamba-1-5-largeOpen weightsAI21 Labs · Jamba 1.5 · best of 2 rows4.07%IndependentreasoningoffPartially comparable-1.14 ptobs. 12 Sept 2026artificialanalysis.aiT2History
389Mistral Medium 3ClosedMistral AI · Mistral · best of 2 rows4.05%IndependentreasoningoffPartially comparable-1.16 ptobs. 12 Sept 2026artificialanalysis.aiT2History
390Command AOpen weightsCohere · Command · best of 2 rows4.03%IndependentreasoningoffPartially comparable-1.18 ptobs. 12 Sept 2026artificialanalysis.aiT2History
390Devstral Small 1.0Open weightsMistral AI · Devstral · best of 2 rows4.03%IndependentreasoningoffPartially comparable-1.18 ptobs. 12 Sept 2026artificialanalysis.aiT2History
392Mixtral 8x22BOpen weightsMistral AI · Mixtral 8 · best of 2 rows4.01%IndependentreasoningoffPartially comparable-1.20 ptobs. 12 Sept 2026artificialanalysis.aiT2History
393Hermes 3 70B InstructOpen weightsNous Research · Hermes 3 · best of 2 rows3.99%IndependentreasoningoffPartially comparable-1.22 ptobs. 12 Sept 2026artificialanalysis.aiT2History
393qwen2.5-32b-instructOpen weightsAlibaba Group · Qwen2.5 · best of 2 rows3.99%IndependentreasoningoffPartially comparable-1.22 ptobs. 12 Sept 2026artificialanalysis.aiT2History
395Llama-3.1-405BRestricted weightsMeta AI · Llama 3.1 · best of 2 rows3.98%IndependentreasoningoffPartially comparable-1.23 ptobs. 12 Sept 2026artificialanalysis.aiT2History
396gpt-4o-chatgpt-03-25ClosedOpenAI · GPT 4 · best of 2 rows3.97%IndependentreasoningoffPartially comparable-1.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
397QwQ-32B-PreviewOpen weightsAlibaba Group · Qwen · best of 2 rows3.91%IndependentreasoningonPartially comparable-1.30 ptobs. 12 Sept 2026artificialanalysis.aiT2History
398Qwen3 8BOpen weightsQwen · Qwen33.89%Independentgroup defaultsPartially comparable-1.32 ptobs. 11 Sept 2026artificialanalysis.aiT2History
398claude-21ClosedAnthropic · Claude 21 · best of 2 rows3.89%IndependentreasoningoffPartially comparable-1.32 ptobs. 12 Sept 2026artificialanalysis.aiT2History
398qwen3-8b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows3.89%IndependentreasoningonPartially comparable-1.32 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →