Skip to content
AI Atlas
BenchmarkActivecategory · reasoningfamily · gpqa · variant Diamond

GPQA Diamond

github.com/idavidrein/gpqa

graduate-level science questions — the 198-question Diamond subset (expert-validated, non-expert-failed)

data quality57

Updated 3 h ago · first seen 12 Sept 2026

Metric
accuracy · %
Current results
1,227
Models
459
Current leader
gpt-6-astra 96.3%

Frontier over time · accuracy · variant=Diamond · evaluator=Artificial Analysis

9 leader changes recorded, all dated 12 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 96.3%gpt-6-astra OpenAI Independent12 Sept 2026
  2. 96.1%gpt-6-astra OpenAI Independent12 Sept 2026
  3. 95.3%Gemini 3.8 Flash Google Independent12 Sept 2026
  4. 95.0%gpt-6-astra OpenAI Independent12 Sept 2026
  5. 93.9%gpt-6-astra OpenAI Independent12 Sept 2026
  6. 93.5%Kimi K3 Moonshot AI Independent12 Sept 2026
  7. 91.9%Claude Opus 5 Anthropic Independent12 Sept 2026
  8. 79.1%grok-3-mini-reasoning SpaceXAI Independent12 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 440 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
100jt-4-1-flash-236b-a21bClosedChina Mobile84.5%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
100o3-proClosedOpenAI · OpenAI o-series84.5%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
103Gemini 2.5 ProClosedGoogle · Gemini 2.584.4%IndependentreasoningonPartially comparable-0.11 ptobs. 12 Sept 2026artificialanalysis.aiT2History
104Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 2 rows84.2%IndependentreasoningonPartially comparable-0.31 ptobs. 12 Sept 2026artificialanalysis.aiT2History
105Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 2 rows84.1%IndependentreasoningonPartially comparable-0.41 ptobs. 12 Sept 2026artificialanalysis.aiT2History
106DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 3 rows84.0%IndependentreasoningonPartially comparable-0.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
107Gemini 3.5 Flash-LiteClosedGoogle · Gemini 3.583.8%IndependentreasoningonPartially comparable-0.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
107Kimi K2 ThinkingOpen weightsMoonshot AI · Kimi83.8%IndependentreasoningonPartially comparable-0.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
109GPT-5-CodexClosedOpenAI · GPT 583.7%Independentreasoning_efforthighPartially comparable-0.81 ptobs. 12 Sept 2026artificialanalysis.aiT2History
110gemini-2-5-pro-03-25ClosedGoogle · Gemini 2.583.6%IndependentreasoningonPartially comparable-0.91 ptobs. 12 Sept 2026artificialanalysis.aiT2History
111Muse Glimmer 30BOpen weightsMeta AI83.5%Independentreasoning_efforthighPartially comparable-1.01 ptobs. 12 Sept 2026artificialanalysis.aiT2History
111mimo-v2-0206Open weightsXiaomi83.5%IndependentreasoningonPartially comparable-1.01 ptobs. 12 Sept 2026artificialanalysis.aiT2History
113Claude Sonnet 4.5ClosedAnthropic · Claude · best of 2 rows83.4%IndependentreasoningonPartially comparable-1.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
113motif-3Open weightsMotif Technologies83.4%IndependentreasoningonPartially comparable-1.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
115Step 3.5 FlashOpen weightsStepFun · Step3.5 · best of 2 rows83.1%IndependentreasoningonPartially comparable-1.42 ptobs. 12 Sept 2026artificialanalysis.aiT2History
116MiniMax M2.1Open weightsMiniMax · MiniMax83.0%IndependentreasoningonPartially comparable-1.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
117jt-35b-flashClosedChina Mobile82.9%IndependentreasoningoffPartially comparable-1.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
117k-exaone-2-0-0803Open weightsLG AI Research · EXAONE 2.082.9%IndependentreasoningonPartially comparable-1.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
119gpt-5-miniClosedOpenAI · GPT 5 · best of 3 rows82.8%Independentreasoning_efforthighPartially comparable-1.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
119mimo-v2-omniClosedXiaomi82.8%IndependentreasoningonPartially comparable-1.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
121o3ClosedOpenAI · OpenAI o-series82.7%IndependentreasoningonPartially comparable-1.82 ptobs. 12 Sept 2026artificialanalysis.aiT2History
122qwen3-5-omni-plusClosedAlibaba Group · Qwen3.582.6%IndependentreasoningoffPartially comparable-1.92 ptobs. 12 Sept 2026artificialanalysis.aiT2History
123gpt-5-5-instant-06-26ClosedOpenAI · GPT 5.582.3%IndependentreasoningonPartially comparable-2.23 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.182.2%IndependentreasoningonPartially comparable-2.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124gemini-2-5-pro-05-06ClosedGoogle · Gemini 2.582.2%IndependentreasoningonPartially comparable-2.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
126GLM 5Open weightsZ.ai (Zhipu AI) · GLM5 · best of 2 rows82.0%IndependentreasoningonPartially comparable-2.53 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 3 rows81.7%Independentreasoning_effortxhighComparable-2.83 ptobs. 12 Sept 2026artificialanalysis.aiT2History
128GPT-5.1-Codex MiniClosedOpenAI · GPT 5.181.3%Independentreasoning_efforthighPartially comparable-3.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
128deepseek-r1Open weightsDeepSeek · DeepSeek-R181.3%IndependentreasoningonPartially comparable-3.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
130ernie-4-5-300b-a47bOpen weightsBaidu · ERNIE 4.581.1%IndependentreasoningoffPartially comparable-3.44 ptobs. 12 Sept 2026artificialanalysis.aiT2History
130nova-2-0-liteClosedAmazon Web Services · Nova 2.0 · best of 4 rows81.1%Independentreasoningonreasoning_efforthighPartially comparable-3.44 ptobs. 12 Sept 2026artificialanalysis.aiT2History
132GLM 5V TurboClosedZ.ai (Zhipu AI) · GLM580.9%IndependentreasoningonPartially comparable-3.64 ptobs. 12 Sept 2026artificialanalysis.aiT2History
132Step 3.7 FlashOpen weightsStepFun · Step3.780.9%IndependentreasoningonPartially comparable-3.64 ptobs. 12 Sept 2026artificialanalysis.aiT2History
134Claude Opus 4.1ClosedAnthropic · Claude80.9%IndependentreasoningonPartially comparable-3.65 ptobs. 12 Sept 2026artificialanalysis.aiT2History
135Qwen3.5-9BOpen weightsQwen · Qwen3.5 · best of 2 rows80.6%IndependentreasoningonPartially comparable-3.94 ptobs. 12 Sept 2026artificialanalysis.aiT2History
136g9v3-39a5bOpen weightsAI9Stars80.5%IndependentreasoningonPartially comparable-4.04 ptobs. 12 Sept 2026artificialanalysis.aiT2History
137Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 380%IndependentreasoningonPartially comparable-4.55 ptobs. 12 Sept 2026artificialanalysis.aiT2History
138DeepSeek V3.2 ExpOpen weightsDeepSeek · DeepSeek-V3 · best of 2 rows79.7%IndependentreasoningonPartially comparable-4.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
139claude-4-opusClosedAnthropic · Claude 4 · best of 2 rows79.6%IndependentreasoningonPartially comparable-4.95 ptobs. 12 Sept 2026artificialanalysis.aiT2History
140exaone-4-5-33bOpen weightsLG AI Research · EXAONE 4.579.4%IndependentreasoningonPartially comparable-5.16 ptobs. 12 Sept 2026artificialanalysis.aiT2History
141gemini-2-5-flash-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 4 rows79.3%IndependentreasoningonPartially comparable-5.26 ptobs. 12 Sept 2026artificialanalysis.aiT2History
142DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 2 rows79.2%IndependentreasoningonPartially comparable-5.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
142Gemma 4 26B A4BOpen weightsGoogle · Gemma 4 · best of 2 rows79.2%IndependentreasoningonPartially comparable-5.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
144grok-3-mini-reasoningClosedSpaceXAI · Grok 379.1%Independentreasoningonreasoning_efforthighPartially comparable-5.46 ptobs. 12 Sept 2026artificialanalysis.aiT2History
145Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 2 rows79.0%IndependentreasoningonPartially comparable-5.56 ptobs. 12 Sept 2026artificialanalysis.aiT2History
146nova-2-0-proClosedAmazon Web Services · Nova 2.0 · best of 3 rows78.5%Independentreasoningonreasoning_effortmediumPartially comparable-6.07 ptobs. 12 Sept 2026artificialanalysis.aiT2History
147o4-miniClosedOpenAI · OpenAI o-series78.4%Independentreasoning_efforthighPartially comparable-6.17 ptobs. 12 Sept 2026artificialanalysis.aiT2History
148k-exaoneOpen weightsLG AI Research · EXAONE · best of 2 rows78.3%IndependentreasoningonPartially comparable-6.27 ptobs. 12 Sept 2026artificialanalysis.aiT2History
149GLM 4.5Open weightsZ.ai (Zhipu AI) · GLM4.578.2%IndependentreasoningonPartially comparable-6.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
149gpt-oss-120bOpen weightsOpenAI · gpt-oss · best of 2 rows78.2%Independentreasoning_efforthighPartially comparable-6.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
151GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows78.0%IndependentreasoningonPartially comparable-6.57 ptobs. 12 Sept 2026artificialanalysis.aiT2History
151LongCat 2.0Open weightsMeituan78.0%IndependentreasoningonPartially comparable-6.57 ptobs. 12 Sept 2026artificialanalysis.aiT2History
153DeepSeek V3.1Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows77.9%IndependentreasoningonPartially comparable-6.67 ptobs. 12 Sept 2026artificialanalysis.aiT2History
154Claude Sonnet 4ClosedAnthropic · Claude · best of 2 rows77.7%IndependentreasoningonPartially comparable-6.87 ptobs. 12 Sept 2026artificialanalysis.aiT2History
154MiniMax M2Open weightsMiniMax · MiniMax77.7%IndependentreasoningonPartially comparable-6.87 ptobs. 12 Sept 2026artificialanalysis.aiT2History
154ernie-5-0-thinking-previewClosedBaidu · ERNIE 5.077.7%IndependentreasoningonPartially comparable-6.87 ptobs. 12 Sept 2026artificialanalysis.aiT2History
157qwen3-max-thinking-previewClosedAlibaba Group · Qwen377.6%IndependentreasoningonPartially comparable-6.97 ptobs. 12 Sept 2026artificialanalysis.aiT2History
158ring-1tOpen weightsinclusionAI77.4%IndependentreasoningonPartially comparable-7.18 ptobs. 12 Sept 2026artificialanalysis.aiT2History
159o3-miniClosedOpenAI · OpenAI o-series · best of 2 rows77.3%Independentreasoning_efforthighPartially comparable-7.28 ptobs. 12 Sept 2026artificialanalysis.aiT2History
160Claude 3.7 SonnetClosedAnthropic · Claude · best of 2 rows77.2%IndependentreasoningonPartially comparable-7.38 ptobs. 12 Sept 2026artificialanalysis.aiT2History
160Qwen3 VL 235B A22B InstructOpen weightsQwen · Qwen3 · best of 2 rows77.2%IndependentreasoningonPartially comparable-7.38 ptobs. 12 Sept 2026artificialanalysis.aiT2History
162Qwen3.5-4BOpen weightsQwen · Qwen3.5 · best of 2 rows77.1%IndependentreasoningonPartially comparable-7.48 ptobs. 12 Sept 2026artificialanalysis.aiT2History
163Mercury 2ClosedInception77.0%IndependentreasoningonPartially comparable-7.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
164Mistral Small 4Open weightsMistral AI · Mistral · best of 2 rows76.9%IndependentreasoningonPartially comparable-7.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
165cogito-v2-1-reasoningOpen weightsDeep Cogito · Cogito76.8%IndependentreasoningonPartially comparable-7.78 ptobs. 12 Sept 2026artificialanalysis.aiT2History
166Kimi K2 0905Open weightsMoonshot AI · Kimi76.7%IndependentreasoningoffPartially comparable-7.88 ptobs. 12 Sept 2026artificialanalysis.aiT2History
167Kimi K2 0711Open weightsMoonshot AI · Kimi76.6%IndependentreasoningoffPartially comparable-7.98 ptobs. 12 Sept 2026artificialanalysis.aiT2History
168o1 PreviewClosedOpenAI · OpenAI o-series76.5%IndependentreasoningonPartially comparable-8.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
169doubao-seed-codeClosedByteDance · Seed76.4%IndependentreasoningonPartially comparable-8.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
169kat-coder-pro-v1ClosedKwaiKAT76.4%IndependentreasoningoffPartially comparable-8.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
169qwen3-max-previewClosedAlibaba Group · Qwen376.4%IndependentreasoningoffPartially comparable-8.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
172command-a-plusOpen weightsCohere · Command76.1%IndependentreasoningonPartially comparable-8.49 ptobs. 12 Sept 2026artificialanalysis.aiT2History
172intellect-3Open weightsPrime Intellect76.1%IndependentreasoningonPartially comparable-8.49 ptobs. 12 Sept 2026artificialanalysis.aiT2History
174nova-2-0-omniClosedAmazon Web Services · Nova 2.0 · best of 3 rows76.0%Independentreasoningonreasoning_effortmediumPartially comparable-8.59 ptobs. 12 Sept 2026artificialanalysis.aiT2History
175Qwen3 Next 80B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows75.9%IndependentreasoningonPartially comparable-8.69 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176nemotron-cascade-2-30b-a3bOpen weightsNVIDIA · Nemotron75.8%IndependentreasoningonPartially comparable-8.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
177Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 2 rows75.7%IndependentreasoningonPartially comparable-8.89 ptobs. 12 Sept 2026artificialanalysis.aiT2History
177North Mini Code (free)Open weightsCohere75.7%IndependentreasoningonPartially comparable-8.89 ptobs. 12 Sept 2026artificialanalysis.aiT2History
179gemma-4-12BOpen weightsGoogle · Gemma 4 · best of 2 rows75.3%IndependentreasoningonPartially comparable-9.30 ptobs. 12 Sept 2026artificialanalysis.aiT2History
180Trinity Large ThinkingOpen weightsArcee AI75.2%IndependentreasoningonPartially comparable-9.40 ptobs. 12 Sept 2026artificialanalysis.aiT2History
180ling-2-6-1tOpen weightsinclusionAI75.2%IndependentreasoningoffPartially comparable-9.40 ptobs. 12 Sept 2026artificialanalysis.aiT2History
182Mistral Medium 3.5Open weightsMistral AI · Mistral74.8%IndependentreasoningonPartially comparable-9.70 ptobs. 12 Sept 2026artificialanalysis.aiT2History
182llama-nemotron-super-49b-v1-5Open weightsNVIDIA · Llama · best of 2 rows74.8%IndependentreasoningonPartially comparable-9.70 ptobs. 12 Sept 2026artificialanalysis.aiT2History
184o1ClosedOpenAI · OpenAI o-series74.8%IndependentreasoningonPartially comparable-9.80 ptobs. 12 Sept 2026artificialanalysis.aiT2History
185NVIDIA Nemotron 3.5 Lightning 30B A3BOpen weightsNVIDIA · Nemotron 3.574.3%IndependentreasoningonPartially comparable-10.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186qwen3-5-omni-flashClosedAlibaba Group · Qwen3.574.2%IndependentreasoningoffPartially comparable-10.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
187Magistral Medium 1.2ClosedMistral AI · Magistral73.9%IndependentreasoningonPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
187exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows73.9%IndependentreasoningonPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
189sarvam-105bOpen weightsSarvam73.8%Independentreasoning_efforthighPartially comparable-10.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
190Qwen3 Coder NextOpen weightsQwen · Qwen373.7%IndependentreasoningoffPartially comparable-10.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
191ling-3-0-tinyOpen weightsinclusionAI73.4%IndependentreasoningonPartially comparable-11.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
192GLM 4.5 AirOpen weightsZ.ai (Zhipu AI) · GLM4.573.3%IndependentreasoningonPartially comparable-11.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
192Qwen3 VL 32B InstructOpen weightsQwen · Qwen3 · best of 2 rows73.3%IndependentreasoningonPartially comparable-11.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
192apriel-v1-6-15b-thinkerOpen weightsServiceNow73.3%IndependentreasoningonPartially comparable-11.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
192hypernova-60bOpen weightsMultiverse Computing73.3%Independentreasoning_efforthighPartially comparable-11.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
196quasar-438bClosedMultiverse Computing73.2%Independentreasoning_effortmaxPartially comparable-11.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
197llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.172.8%IndependentreasoningonPartially comparable-11.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
198Grok Build 0.1ClosedxAI · Grok72.7%IndependentreasoningonPartially comparable-11.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
198hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.1 · best of 2 rows72.7%IndependentreasoningonPartially comparable-11.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
200Seed-OSS-36B-InstructOpen weightsByteDance · Seed72.6%IndependentreasoningonPartially comparable-11.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →