Skip to content
AI Atlas
BenchmarkActivecategory · reasoningfamily · gpqa · variant Diamond

GPQA Diamond

github.com/idavidrein/gpqa

graduate-level science questions — the 198-question Diamond subset (expert-validated, non-expert-failed)

data quality57

Updated 5 h ago · first seen 12 Sept 2026

Metric
accuracy · %
Current results
1,227
Models
459
Current leader
gpt-6-astra 96.3%

Score history · olmo-3-1-32b-instruct 4 rows

Score history for olmo-3-1-32b-instruct0%20%40%60%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • olmo-3-1-32b-instruct
  • 59.09%aa_slug=olmo-3-1-32b-think · variant=Diamond · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
  • 53.94%aa_slug=olmo-3-1-32b-instruct · variant=Diamond · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
  • 59.09%aa_slug=olmo-3-1-32b-think · variant=GPQA Diamond · evaluator=Artificial Analysis · reasoning=on11 Sept 2026
  • 53.94%aa_slug=olmo-3-1-32b-instruct · variant=GPQA Diamond · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=Diamond · evaluator=Artificial Analysis

9 leader changes recorded, all dated 12 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 96.3%gpt-6-astra OpenAI Independent12 Sept 2026
  2. 96.1%gpt-6-astra OpenAI Independent12 Sept 2026
  3. 95.3%Gemini 3.8 Flash Google Independent12 Sept 2026
  4. 95.0%gpt-6-astra OpenAI Independent12 Sept 2026
  5. 93.9%gpt-6-astra OpenAI Independent12 Sept 2026
  6. 93.5%Kimi K3 Moonshot AI Independent12 Sept 2026
  7. 91.9%Claude Opus 5 Anthropic Independent12 Sept 2026
  8. 79.1%grok-3-mini-reasoning SpaceXAI Independent12 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 440 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1gpt-6-astraClosedOpenAI · GPT 6 · best of 5 rows96.3%Independentreasoning_effortxhighleaderobs. 12 Sept 2026artificialanalysis.aiT2History
2Gemini 3.8 FlashClosedGoogle · Gemini 3.8 · best of 3 rows95.3%Independentreasoning_efforthighPartially comparable-1.01 ptobs. 12 Sept 2026artificialanalysis.aiT2History
3Grok 4.6ClosedxAI · Grok · best of 4 rows95.0%Independentreasoning_efforthighPartially comparable-1.31 ptobs. 12 Sept 2026artificialanalysis.aiT2History
4Gemini 3.7 FlashClosedGoogle · Gemini 3.7 · best of 3 rows94.5%Independentreasoning_efforthighPartially comparable-1.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
5Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.194.1%IndependentreasoningonPartially comparable-2.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
5Muse Spark 1.3ClosedMeta AI · best of 2 rows94.1%Independentreasoning_effortxhighComparable-2.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
5gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 6 rows94.1%Independentreasoning_effortmaxPartially comparable-2.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
8Claude Fable 5.1ClosedAnthropic · Claude · best of 5 rows93.7%Independentreasoning_effortmaxPartially comparable-2.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
8Claude Opus 5ClosedAnthropic · Claude · best of 5 rows93.7%Independentreasoning_efforthighPartially comparable-2.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
10Kimi K3Open weightsMoonshot AI · Kimi · best of 2 rows93.5%Independentreasoning_effortmaxPartially comparable-2.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
10Qwen3.8 2.4T A95BOpen weightsQwen · Qwen3.893.5%IndependentreasoningonPartially comparable-2.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
10gpt-5.5ClosedOpenAI · GPT 5.5 · best of 5 rows93.5%Independentreasoning_effortxhighComparable-2.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
13Grok 4.5ClosedxAI · Grok93.1%Independentreasoning_efforthighPartially comparable-3.13 ptobs. 12 Sept 2026artificialanalysis.aiT2History
14MiniMax-M3Open weightsMiniMax · MiniMax92.9%IndependentreasoningonPartially comparable-3.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
15Gemini 3.6 FlashClosedGoogle · Gemini 3.692.8%IndependentreasoningonPartially comparable-3.43 ptobs. 12 Sept 2026artificialanalysis.aiT2History
15deepseek-v4-proClosedDeepSeek · V492.8%Independentreasoning_effortmaxPartially comparable-3.43 ptobs. 12 Sept 2026artificialanalysis.aiT2History
17Qwen 3.8 MaxClosedQwen · Qwen3.892.7%IndependentreasoningonPartially comparable-3.53 ptobs. 12 Sept 2026artificialanalysis.aiT2History
18Claude Fable 5ClosedAnthropic · Claude92.6%IndependentreasoningonPartially comparable-3.63 ptobs. 12 Sept 2026artificialanalysis.aiT2History
19gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 6 rows92.5%Independentreasoning_effortmaxPartially comparable-3.73 ptobs. 12 Sept 2026artificialanalysis.aiT2History
20agnes-3-0-flashClosedSapiens AI92.4%IndependentreasoningonPartially comparable-3.84 ptobs. 12 Sept 2026artificialanalysis.aiT2History
21Qwen3.7 MaxClosedQwen · Qwen3.792.3%IndependentreasoningonPartially comparable-3.94 ptobs. 12 Sept 2026artificialanalysis.aiT2History
21Qwen3.8 FlashOpen weightsQwen · Qwen3.892.3%IndependentreasoningonPartially comparable-3.94 ptobs. 12 Sept 2026artificialanalysis.aiT2History
23Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 3 rows92.2%IndependentreasoningonPartially comparable-4.04 ptobs. 12 Sept 2026artificialanalysis.aiT2History
24Claude Opus 4.8ClosedAnthropic · Claude92.0%Independentreasoning_effortmaxPartially comparable-4.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
24gpt-5.4ClosedOpenAI · GPT 5.4 · best of 3 rows92.0%Independentreasoning_effortxhighComparable-4.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
26GLM 5.3Open weightsZ.ai (Zhipu AI) · GLM5.391.7%Independentreasoning_effortmaxPartially comparable-4.54 ptobs. 12 Sept 2026artificialanalysis.aiT2History
27gpt-5.3-codexClosedOpenAI · GPT 5.391.5%Independentreasoning_effortxhighComparable-4.74 ptobs. 12 Sept 2026artificialanalysis.aiT2History
28Claude Opus 4.7ClosedAnthropic · Claude · best of 2 rows91.4%Independentreasoning_effortmaxPartially comparable-4.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
29deepseek-v4-flash-visionClosedDeepSeek · DeepSeek91.3%Independentreasoning_effortmaxPartially comparable-4.95 ptobs. 12 Sept 2026artificialanalysis.aiT2History
30GLM 5.3 FlashOpen weightsZ.ai (Zhipu AI) · GLM5.391.2%IndependentreasoningonPartially comparable-5.05 ptobs. 12 Sept 2026artificialanalysis.aiT2History
31Claude Sonnet 5ClosedAnthropic · Claude · best of 2 rows91.1%Independentreasoning_effortmaxPartially comparable-5.15 ptobs. 12 Sept 2026artificialanalysis.aiT2History
31Grok 4.20ClosedxAI · Grok · best of 2 rows91.1%IndependentreasoningonPartially comparable-5.15 ptobs. 12 Sept 2026artificialanalysis.aiT2History
31Kimi K2.6Open weightsMoonshot AI · Kimi · best of 2 rows91.1%IndependentreasoningonPartially comparable-5.15 ptobs. 12 Sept 2026artificialanalysis.aiT2History
31gpt-5.6-lunaClosedOpenAI · GPT 5.6 · best of 6 rows91.1%Independentreasoning_effortmaxPartially comparable-5.15 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35deepseek-v4-flashOpen weightsDeepSeek · DeepSeek90.8%Independentreasoning_effortmaxPartially comparable-5.45 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35gemini-3-proClosedGoogle · Gemini 3 · best of 2 rows90.8%Independentreasoning_efforthighPartially comparable-5.45 ptobs. 12 Sept 2026artificialanalysis.aiT2History
37Qwen3.8 27BOpen weightsQwen · Qwen3.8 · best of 4 rows90.5%Independentreasoning_effortxhighComparable-5.75 ptobs. 12 Sept 2026artificialanalysis.aiT2History
37agnes-2-5-pro-betaClosedSapiens AI90.5%IndependentreasoningonPartially comparable-5.75 ptobs. 12 Sept 2026artificialanalysis.aiT2History
37deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek · best of 3 rows90.5%Independentreasoning_efforthighPartially comparable-5.75 ptobs. 12 Sept 2026artificialanalysis.aiT2History
40Muse Spark 1.2ClosedMeta AI90.4%Independentreasoning_effortxhighComparable-5.86 ptobs. 12 Sept 2026artificialanalysis.aiT2History
41gpt-5.2ClosedOpenAI · GPT 5.2 · best of 3 rows90.3%Independentreasoning_effortxhighComparable-5.96 ptobs. 12 Sept 2026artificialanalysis.aiT2History
42Grok 4.3ClosedxAI · Grok · best of 4 rows90.1%Independentreasoning_efforthighPartially comparable-6.16 ptobs. 12 Sept 2026artificialanalysis.aiT2History
43Qwen 3.7 PlusClosedQwen · Qwen3.790%IndependentreasoningonPartially comparable-6.26 ptobs. 12 Sept 2026artificialanalysis.aiT2History
44GPT-5.2-CodexClosedOpenAI · GPT 5.289.9%Independentreasoning_effortxhighComparable-6.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
45Muse Spark 1.1ClosedMeta AI89.8%Independentreasoning_effortxhighComparable-6.46 ptobs. 12 Sept 2026artificialanalysis.aiT2History
45gemini-3-flashClosedGoogle · Gemini 3 · best of 2 rows89.8%IndependentreasoningonPartially comparable-6.46 ptobs. 12 Sept 2026artificialanalysis.aiT2History
47Hy3Open weightsTencent · best of 3 rows89.7%IndependentreasoningonPartially comparable-6.56 ptobs. 12 Sept 2026artificialanalysis.aiT2History
48Claude Opus 4.6ClosedAnthropic · Claude · best of 2 rows89.6%Independentreasoningadaptivereasoning_effortmaxPartially comparable-6.66 ptobs. 12 Sept 2026artificialanalysis.aiT2History
48Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi89.6%IndependentreasoningonPartially comparable-6.66 ptobs. 12 Sept 2026artificialanalysis.aiT2History
50Inkling SmallOpen weightsThinking Machines89.5%IndependentreasoningonPartially comparable-6.77 ptobs. 12 Sept 2026artificialanalysis.aiT2History
50Z.ai GLM 5.2Open weightsZ.ai (Zhipu AI) · GLM5.2 · best of 2 rows89.5%Independentreasoning_effortmaxPartially comparable-6.77 ptobs. 12 Sept 2026artificialanalysis.aiT2History
50grok-build-0-1-06-16ClosedSpaceXAI · Grok89.5%IndependentreasoningonPartially comparable-6.77 ptobs. 12 Sept 2026artificialanalysis.aiT2History
53deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek · best of 3 rows89.4%Independentreasoning_effortmaxPartially comparable-6.87 ptobs. 12 Sept 2026artificialanalysis.aiT2History
54Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 2 rows89.3%IndependentreasoningonPartially comparable-6.97 ptobs. 12 Sept 2026artificialanalysis.aiT2History
55nex-n2-proOpen weightsNex AGI89.2%IndependentreasoningonPartially comparable-7.07 ptobs. 12 Sept 2026artificialanalysis.aiT2History
56Solar Pro 4ClosedUpstage · Solar89.1%IndependentreasoningonPartially comparable-7.17 ptobs. 12 Sept 2026artificialanalysis.aiT2History
57Qwen3.6 Max PreviewClosedQwen · Qwen3.688.8%IndependentreasoningonPartially comparable-7.47 ptobs. 12 Sept 2026artificialanalysis.aiT2History
58grok-4-20-0309ClosedSpaceXAI · Grok 4.20 · best of 2 rows88.5%IndependentreasoningonPartially comparable-7.78 ptobs. 12 Sept 2026artificialanalysis.aiT2History
59muse-sparkClosedMeta AI88.4%IndependentreasoningonPartially comparable-7.88 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Qwen3.6 PlusClosedQwen · Qwen3.688.2%IndependentreasoningonPartially comparable-8.08 ptobs. 12 Sept 2026artificialanalysis.aiT2History
61Kimi K2.5Open weightsMoonshot AI · Kimi · best of 2 rows87.9%IndependentreasoningonPartially comparable-8.38 ptobs. 12 Sept 2026artificialanalysis.aiT2History
62grok-4ClosedSpaceXAI · Grok 487.7%IndependentreasoningonPartially comparable-8.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
63agnes-2-5-pro-alphaOpen weightsSapiens AI87.6%IndependentreasoningonPartially comparable-8.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
64Claude Sonnet 4.6ClosedAnthropic · Claude · best of 3 rows87.5%Independentreasoningadaptivereasoning_effortmaxPartially comparable-8.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
64gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 3 rows87.5%Independentreasoning_effortxhighComparable-8.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
66MiniMax M2.7Open weightsMiniMax · MiniMax87.4%IndependentreasoningonPartially comparable-8.89 ptobs. 12 Sept 2026artificialanalysis.aiT2History
67gpt-5.1ClosedOpenAI · GPT 5.1 · best of 2 rows87.3%Independentreasoning_efforthighPartially comparable-8.99 ptobs. 12 Sept 2026artificialanalysis.aiT2History
67k2-horizon-375b-a23bOpen weightsMBZUAI Institute of Foundation Models87.3%IndependentreasoningonPartially comparable-8.99 ptobs. 12 Sept 2026artificialanalysis.aiT2History
69InklingOpen weightsThinking Machines87.2%IndependentreasoningonPartially comparable-9.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
70deepseek-v3-2-specialeOpen weightsDeepSeek · DeepSeek87.1%IndependentreasoningonPartially comparable-9.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
71mimo-v2-proClosedXiaomi87.0%IndependentreasoningonPartially comparable-9.29 ptobs. 12 Sept 2026artificialanalysis.aiT2History
72motif-0714ClosedMotif Technologies86.9%IndependentreasoningonPartially comparable-9.39 ptobs. 12 Sept 2026artificialanalysis.aiT2History
73GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 2 rows86.8%IndependentreasoningonPartially comparable-9.49 ptobs. 12 Sept 2026artificialanalysis.aiT2History
74Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 386.7%IndependentreasoningonPartially comparable-9.59 ptobs. 12 Sept 2026artificialanalysis.aiT2History
75Claude Opus 4.5ClosedAnthropic · Claude · best of 2 rows86.6%IndependentreasoningonPartially comparable-9.69 ptobs. 12 Sept 2026artificialanalysis.aiT2History
75MiMo-V2.5-ProOpen weightsXiaomi · best of 2 rows86.6%IndependentreasoningonPartially comparable-9.69 ptobs. 12 Sept 2026artificialanalysis.aiT2History
77apodex-1-1ClosedApodex86.4%IndependentreasoningonPartially comparable-9.90 ptobs. 12 Sept 2026artificialanalysis.aiT2History
78Ling 3.0 Flash VLOpen weightsinclusionAI86.2%IndependentreasoningonPartially comparable-10.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79Qwen3 MaxClosedQwen · Qwen3 · best of 2 rows86.1%IndependentreasoningonPartially comparable-10.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
80GPT-5.1-CodexClosedOpenAI · GPT 5.186.0%Independentreasoning_efforthighPartially comparable-10.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
81GLM 4.7Open weightsZ.ai (Zhipu AI) · GLM4.7 · best of 2 rows85.9%IndependentreasoningonPartially comparable-10.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
82Qwen3.5-27BOpen weightsQwen · Qwen3.5 · best of 2 rows85.8%IndependentreasoningonPartially comparable-10.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
83Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 2 rows85.7%IndependentreasoningonPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
83Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 2 rows85.7%IndependentreasoningonPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
83a-x-k2Open weightsSK Telecom85.7%IndependentreasoningonPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
83ring-2-6-1tOpen weightsinclusionAI85.7%IndependentreasoningonPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
83solar-open2-250bOpen weightsUpstage · Solar85.7%IndependentreasoningonPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88KAT-Coder-Pro V2ClosedKwaipilot85.5%IndependentreasoningoffPartially comparable-10.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88Ling 3.0 FlashOpen weightsinclusionAI85.5%IndependentreasoningonPartially comparable-10.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88mimo-v2-omni-0327ClosedXiaomi85.5%IndependentreasoningonPartially comparable-10.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
91gpt-5ClosedOpenAI · GPT 5 · best of 4 rows85.3%Independentreasoning_efforthighPartially comparable-10.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
92grok-4-1-fastClosedSpaceXAI · Grok 4.1 · best of 2 rows85.3%IndependentreasoningonPartially comparable-11.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
93MiMo-V2.5Open weightsXiaomi85.0%IndependentreasoningonPartially comparable-11.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
93nanbeige4-1-3bOpen weightsNanbeige85.0%IndependentreasoningonPartially comparable-11.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
95MiniMax M2.5Open weightsMiniMax · MiniMax84.8%IndependentreasoningonPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
96GLM 5 TurboClosedZ.ai (Zhipu AI) · GLM584.8%IndependentreasoningonPartially comparable-11.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
96grok-4-fastClosedSpaceXAI · Grok 4 · best of 2 rows84.8%IndependentreasoningonPartially comparable-11.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
98gpt-5-5-instant-05-26ClosedOpenAI · GPT 5.584.7%IndependentreasoningonPartially comparable-11.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
98mimo-v2-flashOpen weightsXiaomi · best of 2 rows84.7%IndependentreasoningonPartially comparable-11.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
100Qwen3.5-35B-A3BOpen weightsQwen · Qwen3.5 · best of 2 rows84.5%IndependentreasoningonPartially comparable-11.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →