Skip to content
AI Atlas
BenchmarkActivecategory · reasoningfamily · gpqa · variant Diamond

GPQA Diamond

github.com/idavidrein/gpqa

graduate-level science questions — the 198-question Diamond subset (expert-validated, non-expert-failed)

data quality57

Updated 8 h ago · first seen 12 Sept 2026

Metric
accuracy · %
Current results
1,227
Models
459
Current leader
gpt-6-astra 96.3%

Score history · jt-4-1-flash-236b-a21b 2 rows

Score history for jt-4-1-flash-236b-a21b0%20%40%60%80%100%Sept 26Sept 26Sept 26Sept 26Sept 26
  • jt-4-1-flash-236b-a21b
  • 84.55%aa_slug=jt-4-1-flash-236b-a21b · variant=Diamond · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
  • 84.55%aa_slug=jt-4-1-flash-236b-a21b · variant=GPQA Diamond · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=GPQA Diamond · evaluator=Artificial Analysis

9 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 96.3%gpt-6-astra OpenAI Independent11 Sept 2026
  2. 96.1%gpt-6-astra OpenAI Independent11 Sept 2026
  3. 95.3%Gemini 3.8 Flash Google Independent11 Sept 2026
  4. 95.0%gpt-6-astra OpenAI Independent11 Sept 2026
  5. 93.9%gpt-6-astra OpenAI Independent11 Sept 2026
  6. 93.5%Kimi K3 Moonshot AI Independent11 Sept 2026
  7. 91.9%Claude Opus 5 Anthropic Independent11 Sept 2026
  8. 79.1%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 459 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
301claude-35-sonnet-june-24ClosedAnthropic · Claude 3556.0%Independentgroup defaultsPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302granite-4.2-3bOpen weightsIBM · Granite 4.255.9%Independentgroup defaultsPartially comparable-0.10 ptobs. 11 Sept 2026artificialanalysis.aiT2History
303LFM2.5-2.6B (free)Restricted weightsLiquid AI · LFM2.555.8%Independentgroup defaultsPartially comparable-0.20 ptobs. 11 Sept 2026artificialanalysis.aiT2History
304QwQ-32B-PreviewOpen weightsAlibaba Group · Qwen55.7%Independentgroup defaultsPartially comparable-0.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
305gpt-4oClosedOpenAI · GPT 454.3%Independentgroup defaultsPartially comparable-1.62 ptobs. 11 Sept 2026artificialanalysis.aiT2History
306gemini-2-0-flash-lite-previewClosedGoogle · Gemini 2.054.2%Independentgroup defaultsPartially comparable-1.72 ptobs. 11 Sept 2026artificialanalysis.aiT2History
307tri-21b-think-previewOpen weightsTrillion Labs53.8%Independentgroup defaultsPartially comparable-2.12 ptobs. 11 Sept 2026artificialanalysis.aiT2History
308hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.153.6%Independentgroup defaultsPartially comparable-2.32 ptobs. 11 Sept 2026artificialanalysis.aiT2History
309gemini-2-0-flash-lite-001ClosedGoogle · Gemini 2.053.5%Independentgroup defaultsPartially comparable-2.42 ptobs. 11 Sept 2026artificialanalysis.aiT2History
309qwen3-32b-instructOpen weightsAlibaba Group · Qwen353.5%Independentgroup defaultsPartially comparable-2.42 ptobs. 11 Sept 2026artificialanalysis.aiT2History
311Devstral Small 2Open weightsMistral AI · Devstral53.2%Independentgroup defaultsPartially comparable-2.73 ptobs. 11 Sept 2026artificialanalysis.aiT2History
312Reka Flash 3Open weightsrekaai52.9%Independentgroup defaultsPartially comparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
313Command AOpen weightsCohere · Command52.7%Independentgroup defaultsPartially comparable-3.23 ptobs. 11 Sept 2026artificialanalysis.aiT2History
314GPT-4o (2024-05-13)ClosedOpenAI · GPT 452.6%Independentgroup defaultsPartially comparable-3.33 ptobs. 11 Sept 2026artificialanalysis.aiT2History
315Qwen3-4BOpen weightsQwen · Qwen352.2%Independentgroup defaultsPartially comparable-3.74 ptobs. 11 Sept 2026artificialanalysis.aiT2History
316GPT-4o (2024-08-06)ClosedOpenAI · GPT 452.1%Independentgroup defaultsPartially comparable-3.84 ptobs. 11 Sept 2026artificialanalysis.aiT2History
317Olmo-3-7B-ThinkOpen weightsAllen Institute for AI · OLMo 351.6%Independentgroup defaultsPartially comparable-4.34 ptobs. 11 Sept 2026artificialanalysis.aiT2History
317Qwen3 Coder 30B A3B InstructOpen weightsQwen · Qwen351.6%Independentgroup defaultsPartially comparable-4.34 ptobs. 11 Sept 2026artificialanalysis.aiT2History
317tulu3-405bOpen weightsAllen Institute for AI51.6%Independentgroup defaultsPartially comparable-4.34 ptobs. 11 Sept 2026artificialanalysis.aiT2History
320Llama-3.1-405BRestricted weightsMeta AI · Llama 3.151.5%Independentgroup defaultsPartially comparable-4.44 ptobs. 11 Sept 2026artificialanalysis.aiT2History
320exaone-4-0-1-2bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows51.5%IndependentreasoningonPartially comparable-4.44 ptobs. 11 Sept 2026artificialanalysis.aiT2History
322LFM2.5-8B-A1BRestricted weightsLiquid AI · LFM2.551.3%Independentgroup defaultsPartially comparable-4.65 ptobs. 11 Sept 2026artificialanalysis.aiT2History
322nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 351.3%Independentgroup defaultsPartially comparable-4.65 ptobs. 11 Sept 2026artificialanalysis.aiT2History
324gpt-4.1-nanoClosedOpenAI · GPT 4.151.2%Independentgroup defaultsPartially comparable-4.75 ptobs. 11 Sept 2026artificialanalysis.aiT2History
325gpt-4o-chatgptClosedOpenAI · GPT 451.1%Independentgroup defaultsPartially comparable-4.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
326grok-2Open weightsxAI · Grok 251.0%Independentgroup defaultsPartially comparable-4.95 ptobs. 11 Sept 2026artificialanalysis.aiT2History
327Mistral Small 3.2Open weightsMistral AI · Mistral50.5%Independentgroup defaultsPartially comparable-5.45 ptobs. 11 Sept 2026artificialanalysis.aiT2History
327Pixtral LargeOpen weightsMistral AI · Pixtral50.5%Independentgroup defaultsPartially comparable-5.45 ptobs. 11 Sept 2026artificialanalysis.aiT2History
329nova-proClosedAmazon Web Services · Nova49.9%Independentgroup defaultsPartially comparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
330Llama 3.3 70BOpen weightsMeta AI · Llama 3.349.8%Independentgroup defaultsPartially comparable-6.16 ptobs. 11 Sept 2026artificialanalysis.aiT2History
331Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 2 rows49.4%IndependentreasoningonPartially comparable-6.57 ptobs. 11 Sept 2026artificialanalysis.aiT2History
332devstral-mediumClosedMistral AI · Devstral49.2%Independentgroup defaultsPartially comparable-6.77 ptobs. 11 Sept 2026artificialanalysis.aiT2History
333Qwen2.5 72B InstructOpen weightsQwen · Qwen2.549.1%Independentgroup defaultsPartially comparable-6.87 ptobs. 11 Sept 2026artificialanalysis.aiT2History
333hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.149.1%Independentgroup defaultsPartially comparable-6.87 ptobs. 11 Sept 2026artificialanalysis.aiT2History
335claude-3-opusClosedAnthropic · Claude 348.9%Independentgroup defaultsPartially comparable-7.07 ptobs. 11 Sept 2026artificialanalysis.aiT2History
336mistral-large-2Open weightsMistral AI · Mistral48.6%Independentgroup defaultsPartially comparable-7.37 ptobs. 11 Sept 2026artificialanalysis.aiT2History
337deepseek-r1-distill-qwen-14bOpen weightsDeepSeek · Qwen48.4%Independentgroup defaultsPartially comparable-7.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
338granite-4.1-30bOpen weightsIBM · Granite 4.148.1%Independentgroup defaultsPartially comparable-7.88 ptobs. 11 Sept 2026artificialanalysis.aiT2History
339lfm2-24b-a2bOpen weightsLiquid AI · LFM247.4%Independentgroup defaultsPartially comparable-8.59 ptobs. 11 Sept 2026artificialanalysis.aiT2History
340Mistral Large 2.0Open weightsMistral AI · Mistral47.2%Independentgroup defaultsPartially comparable-8.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
341Ministral 3 8BOpen weightsMistral AI · Ministral 347.1%Independentgroup defaultsPartially comparable-8.89 ptobs. 11 Sept 2026artificialanalysis.aiT2History
341grok-betaClosedSpaceXAI · Grok47.1%Independentgroup defaultsPartially comparable-8.89 ptobs. 11 Sept 2026artificialanalysis.aiT2History
343qwen3-14b-instructOpen weightsAlibaba Group · Qwen347.0%Independentgroup defaultsPartially comparable-8.99 ptobs. 11 Sept 2026artificialanalysis.aiT2History
344nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 346.9%Independentgroup defaultsPartially comparable-9.09 ptobs. 11 Sept 2026artificialanalysis.aiT2History
345qwen2.5-32b-instructOpen weightsAlibaba Group · Qwen2.546.6%Independentgroup defaultsPartially comparable-9.39 ptobs. 11 Sept 2026artificialanalysis.aiT2History
346llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.146.5%Independentgroup defaultsPartially comparable-9.50 ptobs. 11 Sept 2026artificialanalysis.aiT2History
347gemini-1-5-flashClosedGoogle · Gemini 1.546.3%Independentgroup defaultsPartially comparable-9.70 ptobs. 11 Sept 2026artificialanalysis.aiT2History
348Mistral Small 3Open weightsMistral AI · Mistral46.2%Independentgroup defaultsPartially comparable-9.80 ptobs. 11 Sept 2026artificialanalysis.aiT2History
349Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 2 rows45.6%Independentgroup defaultsPartially comparable-10.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
350Mistral Small 3.1Open weightsMistral AI · Mistral45.4%Independentgroup defaultsPartially comparable-10.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
351qwen3-8b-instructOpen weightsAlibaba Group · Qwen345.1%Independentgroup defaultsPartially comparable-10.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
352g9v3-3bOpen weightsAI9Stars43.8%Independentgroup defaultsPartially comparable-12.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
353Devstral Small 1.0Open weightsMistral AI · Devstral43.4%Independentgroup defaultsPartially comparable-12.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
354gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 2 rows43.3%Independentgroup defaultsPartially comparable-12.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
354granite-4.1-8bOpen weightsIBM · Granite 4.143.3%Independentgroup defaultsPartially comparable-12.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
354nova-liteClosedAmazon Web Services · Nova43.3%Independentgroup defaultsPartially comparable-12.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
357llama-3-2-instruct-90b-visionOpen weightsMeta AI · Llama 3.243.2%Independentgroup defaultsPartially comparable-12.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
358Gemma 3 27BOpen weightsGoogle · Gemma 342.8%Independentgroup defaultsPartially comparable-13.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
359jamba-1-5-largeOpen weightsAI21 Labs · Jamba 1.542.7%Independentgroup defaultsPartially comparable-13.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
360gpt-4o-miniClosedOpenAI · GPT 442.6%Independentgroup defaultsPartially comparable-13.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
361Molmo2-8BOpen weightsAllen Institute for AI42.5%Independentgroup defaultsPartially comparable-13.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
362Mistral SabaClosedMistral AI · Mistral42.4%Independentgroup defaultsPartially comparable-13.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
363deepseek-v2-5Open weightsDeepSeek · DeepSeek42.3%Independentgroup defaultsPartially comparable-13.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
364Qwen2.5 Coder 32B InstructOpen weightsQwen · Qwen2.541.7%Independentgroup defaultsPartially comparable-14.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
365granite-4-0-h-smallOpen weightsIBM · Granite 4.041.6%Independentgroup defaultsPartially comparable-14.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
365sarvam-m-reasoningOpen weightsSarvam41.6%Independentgroup defaultsPartially comparable-14.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
367devstral-smallOpen weightsMistral AI · Devstral41.4%Independentgroup defaultsPartially comparable-14.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
368Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AI · Kimi41.2%Independentgroup defaultsPartially comparable-14.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
369qwen-turboClosedAlibaba Group · Qwen41.0%Independentgroup defaultsPartially comparable-15.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
370Llama-3.1-70BOpen weightsMeta AI · Llama 3.140.9%Independentgroup defaultsPartially comparable-15.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
371Claude Haiku 3.5ClosedAnthropic · Claude40.8%Independentgroup defaultsPartially comparable-15.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
371llama-3-1-nemotron-nano-4b-reasoningOpen weightsNVIDIA · Llama 3.140.8%Independentgroup defaultsPartially comparable-15.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
373R1 Distill Llama 70BOpen weightsDeepSeek · Llama40.2%Independentgroup defaultsPartially comparable-15.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
374Hermes 3 70B InstructOpen weightsNous Research · Hermes 340.1%Independentgroup defaultsPartially comparable-15.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
375Olmo-3-7B-InstructOpen weightsAllen Institute for AI · OLMo 340%Independentgroup defaultsPartially comparable-16.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
375claude-3-sonnetClosedAnthropic · Claude 340%Independentgroup defaultsPartially comparable-16.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
377qwen3-4b-instructOpen weightsAlibaba Group · Qwen339.8%Independentgroup defaultsPartially comparable-16.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
378jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.739.0%Independentgroup defaultsPartially comparable-17.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
379jamba-1-6-largeOpen weightsAI21 Labs · Jamba 1.638.7%Independentgroup defaultsPartially comparable-17.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
380deephermes-3-mistral-24b-previewOpen weightsNous Research · Mistral38.2%Independentgroup defaultsPartially comparable-17.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
381mistral-smallOpen weightsMistral AI · Mistral38.1%Independentgroup defaultsPartially comparable-17.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
382llama-3-instruct-70bOpen weightsMeta AI · Llama 337.9%Independentgroup defaultsPartially comparable-18.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
383Claude 3 HaikuClosedAnthropic · Claude37.4%Independentgroup defaultsPartially comparable-18.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
384gemini-1-5-pro-may-2024ClosedGoogle · Gemini 1.537.1%Independentgroup defaultsPartially comparable-18.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
384qwen2-72b-instructOpen weightsAlibaba Group · Qwen237.1%Independentgroup defaultsPartially comparable-18.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
386gemini-1-5-flash-8bClosedGoogle · Gemini 1.535.9%Independentgroup defaultsPartially comparable-20.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
387Ministral 3 3BOpen weightsMistral AI · Ministral 335.8%Independentgroup defaultsPartially comparable-20.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
387nova-microClosedAmazon Web Services · Nova35.8%Independentgroup defaultsPartially comparable-20.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
389qwen3-1.7b-instructOpen weightsAlibaba Group · Qwen3.1 · best of 2 rows35.6%IndependentreasoningonPartially comparable-20.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
390Mistral LargeClosedMistral AI · Mistral35.0%Independentgroup defaultsPartially comparable-20.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
391Gemma 3 12BOpen weightsGoogle · Gemma 335.0%Independentgroup defaultsPartially comparable-21.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
391gpt-4ClosedOpenAI · GPT 435.0%Independentgroup defaultsPartially comparable-21.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
391mistral-mediumClosedMistral AI · Mistral35.0%Independentgroup defaultsPartially comparable-21.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
394claude-2ClosedAnthropic · Claude 234.4%Independentgroup defaultsPartially comparable-21.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
394lfm2-8b-a1bOpen weightsLiquid AI · LFM234.4%Independentgroup defaultsPartially comparable-21.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
396Qwen2.5-Coder-7BOpen weightsQwen · Qwen2.533.9%Independentgroup defaultsPartially comparable-22.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
396lfm2-5-1-2b-thinkingOpen weightsLiquid AI · LFM2.533.9%Independentgroup defaultsPartially comparable-22.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
398granite-3-3-8b-instructOpen weightsIBM · Granite 3.333.8%Independentgroup defaultsPartially comparable-22.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
399Granite 4.0 MicroOpen weightsIBM · Granite 4.033.6%Independentgroup defaultsPartially comparable-22.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
400jamba-reasoning-3bOpen weightsAI21 Labs · Jamba33.3%Independentgroup defaultsPartially comparable-22.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →