Skip to content
AI Atlas
BenchmarkActivecategory · knowledgefamily · humanitys-last-exam · variant full

Humanity's Last Exam

lastexam.ai

expert-written frontier questions

data quality57

Updated 5 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,217
Models
453
Current leader
Claude Fable 5.1 59.1%

Frontier over time · accuracy · evaluator=Artificial Analysis

9 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 59.1%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  2. 58.7%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  3. 55.9%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  4. 55.5%Claude Fable 5 Anthropic Independent11 Sept 2026
  5. 53.1%gpt-6-astra OpenAI Independent11 Sept 2026
  6. 52.7%gpt-6-astra OpenAI Independent11 Sept 2026
  7. 51.3%Claude Opus 5 Anthropic Independent11 Sept 2026
  8. 11.0%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 453 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
201Qwen3 Coder NextOpen weightsQwen · Qwen3 · best of 2 rows10.2%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
201Qwen3 VL 32B InstructOpen weightsQwen · Qwen3 · best of 4 rows10.2%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
201qwen3-max-previewClosedAlibaba Group · Qwen3 · best of 2 rows10.2%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
204k2-think-v2Open weightsMBZUAI Institute of Foundation Models · best of 2 rows10.1%IndependentreasoningonPartially comparable-0.05 ptobs. 12 Sept 2026artificialanalysis.aiT2History
205Qwen3.5-4BOpen weightsQwen · Qwen3.5 · best of 4 rows9.92%IndependentreasoningonPartially comparable-0.23 ptobs. 12 Sept 2026artificialanalysis.aiT2History
206Mistral Small 4Open weightsMistral AI · Mistral · best of 4 rows9.87%IndependentreasoningonPartially comparable-0.28 ptobs. 12 Sept 2026artificialanalysis.aiT2History
206Seed-OSS-36B-InstructOpen weightsByteDance · Seed · best of 2 rows9.87%IndependentreasoningonPartially comparable-0.28 ptobs. 12 Sept 2026artificialanalysis.aiT2History
208magistral-mediumClosedMistral AI · Magistral · best of 2 rows9.82%IndependentreasoningonPartially comparable-0.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
209Granite 4.2 8BOpen weightsIBM · Granite 4.2 · best of 2 rows9.68%IndependentreasoningonPartially comparable-0.47 ptobs. 12 Sept 2026artificialanalysis.aiT2History
210Claude 3.7 SonnetClosedAnthropic · Claude · best of 4 rows9.67%IndependentreasoningonPartially comparable-0.48 ptobs. 12 Sept 2026artificialanalysis.aiT2History
211GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 4 rows9.64%IndependentreasoningonPartially comparable-0.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
211ring-flash-2-0Open weightsinclusionAI · best of 2 rows9.64%IndependentreasoningonPartially comparable-0.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
213gpt-5-nanoClosedOpenAI · GPT 5 · best of 6 rows9.50%Independentreasoning_efforthighPartially comparable-0.65 ptobs. 12 Sept 2026artificialanalysis.aiT2History
214nova-2-0-proClosedAmazon Web Services · Nova 2.0 · best of 6 rows9.41%Independentreasoningonreasoning_effortmediumPartially comparable-0.74 ptobs. 12 Sept 2026artificialanalysis.aiT2History
215ling-3-0-tinyOpen weightsinclusionAI · best of 2 rows9.31%IndependentreasoningonPartially comparable-0.84 ptobs. 12 Sept 2026artificialanalysis.aiT2History
216midm-250-pro-rsnsftClosedKorea Telecom · best of 2 rows9.08%IndependentreasoningonPartially comparable-1.07 ptobs. 12 Sept 2026artificialanalysis.aiT2History
217deepseek-v3-2-0925Open weightsDeepSeek · DeepSeek9.04%Independentgroup defaultsPartially comparable-1.11 ptobs. 11 Sept 2026artificialanalysis.aiT2History
218Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 4 rows8.94%IndependentreasoningonPartially comparable-1.21 ptobs. 12 Sept 2026artificialanalysis.aiT2History
218minicpm5-2bOpen weightsOpenBMB · best of 2 rows8.94%IndependentreasoningonPartially comparable-1.21 ptobs. 12 Sept 2026artificialanalysis.aiT2History
220MiniMax-M1-80kOpen weightsMiniMax · MiniMax · best of 2 rows8.90%IndependentreasoningonPartially comparable-1.25 ptobs. 12 Sept 2026artificialanalysis.aiT2History
221Hermes-4-70BRestricted weightsNous Research · Hermes 48.76%Independentgroup defaultsPartially comparable-1.39 ptobs. 11 Sept 2026artificialanalysis.aiT2History
221hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.1 · best of 3 rows8.76%IndependentreasoningonPartially comparable-1.39 ptobs. 12 Sept 2026artificialanalysis.aiT2History
221motif-2-12-7bClosedMotif Technologies · best of 2 rows8.76%IndependentreasoningonPartially comparable-1.39 ptobs. 12 Sept 2026artificialanalysis.aiT2History
224ling-2-6-1tOpen weightsinclusionAI · best of 2 rows8.71%IndependentreasoningoffPartially comparable-1.44 ptobs. 12 Sept 2026artificialanalysis.aiT2History
225deepseek-r1-0120Open weightsDeepSeek · DeepSeek · best of 2 rows8.50%IndependentreasoningonPartially comparable-1.65 ptobs. 12 Sept 2026artificialanalysis.aiT2History
226deepseek-v4-pro-0424-non-reasoningOpen weightsDeepSeek · DeepSeek8.25%Independentgroup defaultsPartially comparable-1.90 ptobs. 11 Sept 2026artificialanalysis.aiT2History
227mi-dm-k-2-5-pro-dec28ClosedKorea Telecom · best of 2 rows8.06%IndependentreasoningonPartially comparable-2.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
228Grok Build 0.1ClosedxAI · Grok · best of 2 rows8.02%IndependentreasoningonPartially comparable-2.13 ptobs. 12 Sept 2026artificialanalysis.aiT2History
229MiniMax-M1-40kOpen weightsMiniMax · MiniMax · best of 2 rows7.83%IndependentreasoningonPartially comparable-2.32 ptobs. 12 Sept 2026artificialanalysis.aiT2History
230deepseek-v4-flash-0420-non-reasoningOpen weightsDeepSeek · DeepSeek7.78%Independentgroup defaultsPartially comparable-2.37 ptobs. 11 Sept 2026artificialanalysis.aiT2History
231GLM 4.7 FlashOpen weightsZ.ai (Zhipu AI) · GLM4.7 · best of 4 rows7.60%IndependentreasoningonPartially comparable-2.55 ptobs. 12 Sept 2026artificialanalysis.aiT2History
231qwen3-5-omni-flashClosedAlibaba Group · Qwen3.5 · best of 2 rows7.60%IndependentreasoningoffPartially comparable-2.55 ptobs. 12 Sept 2026artificialanalysis.aiT2History
233magistral-smallOpen weightsMistral AI · Magistral · best of 2 rows7.55%IndependentreasoningonPartially comparable-2.60 ptobs. 12 Sept 2026artificialanalysis.aiT2History
234sarvam-30bOpen weightsSarvam · best of 2 rows7.51%Independentreasoning_efforthighPartially comparable-2.64 ptobs. 12 Sept 2026artificialanalysis.aiT2History
235qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows7.46%IndependentreasoningonPartially comparable-2.69 ptobs. 12 Sept 2026artificialanalysis.aiT2History
236Kimi K2 0711Open weightsMoonshot AI · Kimi · best of 2 rows7.41%IndependentreasoningoffPartially comparable-2.74 ptobs. 12 Sept 2026artificialanalysis.aiT2History
236Qwen3 32BOpen weightsQwen · Qwen37.41%Independentgroup defaultsPartially comparable-2.74 ptobs. 11 Sept 2026artificialanalysis.aiT2History
236qwen3-32b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows7.41%IndependentreasoningonPartially comparable-2.74 ptobs. 12 Sept 2026artificialanalysis.aiT2History
239llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.1 · best of 2 rows7.38%IndependentreasoningonPartially comparable-2.77 ptobs. 12 Sept 2026artificialanalysis.aiT2History
240solar-pro-2ClosedUpstage · Solar · best of 5 rows7.37%IndependentreasoningonPartially comparable-2.78 ptobs. 12 Sept 2026artificialanalysis.aiT2History
241QwQ-32BOpen weightsAlibaba Group · Qwen · best of 2 rows7.35%IndependentreasoningonPartially comparable-2.80 ptobs. 12 Sept 2026artificialanalysis.aiT2History
242ling-1tOpen weightsinclusionAI · best of 2 rows7.32%IndependentreasoningoffPartially comparable-2.83 ptobs. 12 Sept 2026artificialanalysis.aiT2History
242llama-nemotron-super-49b-v1-5Open weightsNVIDIA · Llama · best of 4 rows7.32%IndependentreasoningonPartially comparable-2.83 ptobs. 12 Sept 2026artificialanalysis.aiT2History
244GLM 4.5 AirOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows7.04%IndependentreasoningonPartially comparable-3.11 ptobs. 12 Sept 2026artificialanalysis.aiT2History
245o1ClosedOpenAI · OpenAI o-series · best of 2 rows7.01%IndependentreasoningonPartially comparable-3.14 ptobs. 12 Sept 2026artificialanalysis.aiT2History
246Gemini 2.5 Flash-LiteClosedGoogle · Gemini 2.5 · best of 5 rows6.95%IndependentreasoningonPartially comparable-3.20 ptobs. 12 Sept 2026artificialanalysis.aiT2History
246gemini-2-5-flash-lite-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 3 rows6.95%IndependentreasoningonPartially comparable-3.20 ptobs. 11 Sept 2026artificialanalysis.aiT2History
246nova-2-0-omniClosedAmazon Web Services · Nova 2.0 · best of 4 rows6.95%Independentreasoningonreasoning_effortmediumPartially comparable-3.20 ptobs. 12 Sept 2026artificialanalysis.aiT2History
249LFM2.5-8B-A1BOpen weightsLiquid AI · LFM2.5 · best of 2 rows6.86%IndependentreasoningonPartially comparable-3.29 ptobs. 12 Sept 2026artificialanalysis.aiT2History
250celeris-1ClosedCeleris · best of 2 rows6.81%IndependentreasoningoffPartially comparable-3.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
251LFM2.5-1.2B-InstructOpen weightsLiquid AI · LFM2.5 · best of 2 rows6.72%IndependentreasoningoffPartially comparable-3.43 ptobs. 12 Sept 2026artificialanalysis.aiT2History
252gpt-5-chatgptClosedOpenAI · GPT 5 · best of 2 rows6.63%IndependentreasoningoffPartially comparable-3.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
252granite-4.2-3bOpen weightsIBM · Granite 4.2 · best of 2 rows6.63%IndependentreasoningonPartially comparable-3.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
254Sonar ProClosedPerplexity AI · Sonar · best of 2 rows6.55%IndependentreasoningoffPartially comparable-3.60 ptobs. 12 Sept 2026artificialanalysis.aiT2History
255minicpm5-1bOpen weightsOpenBMB · best of 4 rows6.49%IndependentreasoningonPartially comparable-3.66 ptobs. 12 Sept 2026artificialanalysis.aiT2History
256Magistral Small 1.2Open weightsMistral AI · Magistral · best of 2 rows6.44%IndependentreasoningonPartially comparable-3.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
256granite-4-0-h-350mOpen weightsIBM · Granite 4.0 · best of 2 rows6.44%IndependentreasoningoffPartially comparable-3.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
256jt-35b-flashClosedChina Mobile · best of 2 rows6.44%IndependentreasoningoffPartially comparable-3.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
256jt-miniClosedChina Mobile · best of 2 rows6.44%IndependentreasoningoffPartially comparable-3.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
256olmo-3-32b-thinkOpen weightsAllen Institute for AI · OLMo 3 · best of 2 rows6.44%IndependentreasoningonPartially comparable-3.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
261Kimi K2 0905Open weightsMoonshot AI · Kimi · best of 2 rows6.39%IndependentreasoningoffPartially comparable-3.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
262gemini-2.0-flash-thinking-exp-01-21ClosedGoogle · Gemini 2.0 · best of 2 rows6.35%IndependentreasoningonPartially comparable-3.80 ptobs. 12 Sept 2026artificialanalysis.aiT2History
263GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 4 rows6.30%IndependentreasoningonPartially comparable-3.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
264ling-2-6-flashOpen weightsinclusionAI · best of 2 rows6.26%IndependentreasoningoffPartially comparable-3.89 ptobs. 12 Sept 2026artificialanalysis.aiT2History
264olmo-3-1-32b-instructOpen weightsAllen Institute for AI · OLMo 3.1 · best of 3 rows6.26%IndependentreasoningonPartially comparable-3.89 ptobs. 12 Sept 2026artificialanalysis.aiT2History
266LFM2.5-2.6B (free)Open weightsLiquid AI · LFM2.5 · best of 2 rows6.21%IndependentreasoningonPartially comparable-3.94 ptobs. 12 Sept 2026artificialanalysis.aiT2History
266lfm2-5-1-2b-thinkingOpen weightsLiquid AI · LFM2.5 · best of 2 rows6.21%IndependentreasoningonPartially comparable-3.94 ptobs. 12 Sept 2026artificialanalysis.aiT2History
266ling-flash-2-0Open weightsinclusionAI · best of 2 rows6.21%IndependentreasoningoffPartially comparable-3.94 ptobs. 12 Sept 2026artificialanalysis.aiT2History
269gemini-2-0-pro-experimental-02-05ClosedGoogle · Gemini 2.0 · best of 2 rows6.19%IndependentreasoningoffPartially comparable-3.96 ptobs. 12 Sept 2026artificialanalysis.aiT2History
270qwen3-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows6.16%IndependentreasoningonPartially comparable-3.99 ptobs. 12 Sept 2026artificialanalysis.aiT2History
270qwen3-4b-2507-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows6.16%IndependentreasoningonPartially comparable-3.99 ptobs. 12 Sept 2026artificialanalysis.aiT2History
272solar-pro-2-previewClosedUpstage · Solar · best of 3 rows6.07%IndependentreasoningonPartially comparable-4.08 ptobs. 11 Sept 2026artificialanalysis.aiT2History
273Olmo-3-7B-ThinkOpen weightsAllen Institute for AI · OLMo 3 · best of 2 rows5.98%IndependentreasoningonPartially comparable-4.17 ptobs. 12 Sept 2026artificialanalysis.aiT2History
273exaone-4-0-1-2bOpen weightsLG AI Research · EXAONE 4.0 · best of 4 rows5.98%IndependentreasoningonPartially comparable-4.17 ptobs. 12 Sept 2026artificialanalysis.aiT2History
275DeepSeek-R1-0528-Qwen3-8BOpen weightsDeepSeek · Qwen3 · best of 2 rows5.93%IndependentreasoningonPartially comparable-4.22 ptobs. 12 Sept 2026artificialanalysis.aiT2History
275tri-21b-think-v0-5Open weightsTrillion Labs · best of 2 rows5.93%IndependentreasoningonPartially comparable-4.22 ptobs. 12 Sept 2026artificialanalysis.aiT2History
277longcat-flash-liteOpen weightsLongCat · best of 2 rows5.84%IndependentreasoningoffPartially comparable-4.31 ptobs. 12 Sept 2026artificialanalysis.aiT2History
278Olmo-3-7B-InstructOpen weightsAllen Institute for AI · OLMo 3 · best of 2 rows5.79%IndependentreasoningoffPartially comparable-4.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
278tri-21b-think-previewOpen weightsTrillion Labs · best of 2 rows5.79%IndependentreasoningonPartially comparable-4.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
280Qwen3-0.6BOpen weightsQwen · Qwen3.05.64%Independentgroup defaultsPartially comparable-4.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
280qwen3-0.6b-instructOpen weightsAlibaba Group · Qwen3.0 · best of 3 rows5.64%IndependentreasoningonPartially comparable-4.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
282llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.3 · best of 4 rows5.62%IndependentreasoningonPartially comparable-4.53 ptobs. 12 Sept 2026artificialanalysis.aiT2History
283LFM2-1.2BOpen weightsLiquid AI · LFM2.1 · best of 2 rows5.56%IndependentreasoningoffPartially comparable-4.59 ptobs. 12 Sept 2026artificialanalysis.aiT2History
284llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.2 · best of 2 rows5.53%IndependentreasoningoffPartially comparable-4.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
285granite-4-0-350mOpen weightsIBM · Granite 4.0 · best of 2 rows5.51%IndependentreasoningoffPartially comparable-4.64 ptobs. 12 Sept 2026artificialanalysis.aiT2History
285nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 4 rows5.51%IndependentreasoningonPartially comparable-4.64 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287apertus-70b-instructOpen weightsSwiss AI Initiative · best of 2 rows5.47%IndependentreasoningoffPartially comparable-4.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287hyperclova-x-seed-think-32bOpen weightsNaver · Seed · best of 2 rows5.47%IndependentreasoningonPartially comparable-4.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287lfm2-2-6bOpen weightsLiquid AI · LFM2.2 · best of 2 rows5.47%IndependentreasoningoffPartially comparable-4.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
290Llama-3.2-1BRestricted weightsMeta AI · Llama 3.2 · best of 2 rows5.46%IndependentreasoningoffPartially comparable-4.69 ptobs. 12 Sept 2026artificialanalysis.aiT2History
291Ministral 3 3BOpen weightsMistral AI · Ministral 3 · best of 2 rows5.38%IndependentreasoningoffPartially comparable-4.77 ptobs. 12 Sept 2026artificialanalysis.aiT2History
291olmo-2-7bOpen weightsAllen Institute for AI · OLMo 2 · best of 2 rows5.38%IndependentreasoningoffPartially comparable-4.77 ptobs. 12 Sept 2026artificialanalysis.aiT2History
293deepseek-coder-v2-liteOpen weightsDeepSeek · DeepSeek · best of 2 rows5.36%IndependentreasoningoffPartially comparable-4.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
294Llama-3.2-3BRestricted weightsMeta AI · Llama 3.2 · best of 2 rows5.35%IndependentreasoningoffPartially comparable-4.80 ptobs. 12 Sept 2026artificialanalysis.aiT2History
295gemma-3-1bOpen weightsGoogle · Gemma 3 · best of 2 rows5.33%IndependentreasoningoffPartially comparable-4.82 ptobs. 12 Sept 2026artificialanalysis.aiT2History
295llama-2-chat-7bOpen weightsMeta AI · Llama 2 · best of 2 rows5.33%IndependentreasoningoffPartially comparable-4.82 ptobs. 12 Sept 2026artificialanalysis.aiT2History
297qwen3-1.7b-instructOpen weightsAlibaba Group · Qwen3.1 · best of 4 rows5.32%IndependentreasoningoffPartially comparable-4.83 ptobs. 12 Sept 2026artificialanalysis.aiT2History
298Gemma 3 4BRestricted weightsGoogle · Gemma 3 · best of 2 rows5.29%IndependentreasoningoffPartially comparable-4.86 ptobs. 12 Sept 2026artificialanalysis.aiT2History
298Llama 3.1 8BRestricted weightsMeta AI · Llama 3.1 · best of 2 rows5.29%IndependentreasoningoffPartially comparable-4.86 ptobs. 12 Sept 2026artificialanalysis.aiT2History
300tiny-aya-globalRestricted weightsCohere · Aya · best of 2 rows5.24%IndependentreasoningoffPartially comparable-4.91 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →