Skip to content
AI Atlas
BenchmarkActivecategory · reasoningfamily · gpqa · variant Diamond

GPQA Diamond

github.com/idavidrein/gpqa

graduate-level science questions — the 198-question Diamond subset (expert-validated, non-expert-failed)

data quality57

Updated 6 h ago · first seen 12 Sept 2026

Metric
accuracy · %
Current results
1,227
Models
459
Current leader
gpt-6-astra 96.3%

Score history · gemini-2-5-flash-lite-preview-09-2025 3 rows

Score history for gemini-2-5-flash-lite-preview-09-20250%20%40%60%80%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • gemini-2-5-flash-lite-preview-09-2025
  • 65.05%aa_slug=gemini-2-5-flash-lite-preview-09-2025 · variant=Diamond · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
  • 65.05%aa_slug=gemini-2-5-flash-lite-preview-09-2025 · variant=GPQA Diamond · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 70.91%aa_slug=gemini-2-5-flash-lite-preview-09-2025-reasoning · variant=GPQA Diamond · evaluator=Artificial Analysis · reasoning=on11 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=GPQA Diamond · evaluator=Artificial Analysis

9 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 96.3%gpt-6-astra OpenAI Independent11 Sept 2026
  2. 96.1%gpt-6-astra OpenAI Independent11 Sept 2026
  3. 95.3%Gemini 3.8 Flash Google Independent11 Sept 2026
  4. 95.0%gpt-6-astra OpenAI Independent11 Sept 2026
  5. 93.9%gpt-6-astra OpenAI Independent11 Sept 2026
  6. 93.5%Kimi K3 Moonshot AI Independent11 Sept 2026
  7. 91.9%Claude Opus 5 Anthropic Independent11 Sept 2026
  8. 79.1%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 459 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
101gpt-5-5-instant-05-26ClosedOpenAI · GPT 5.584.7%Independentgroup defaultsPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
101mimo-v2-flashOpen weightsXiaomi · best of 2 rows84.7%IndependentreasoningonPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
103Qwen3.5-35B-A3BOpen weightsQwen · Qwen3.5 · best of 2 rows84.5%Independentgroup defaultsPartially comparable-0.10 ptobs. 11 Sept 2026artificialanalysis.aiT2History
103jt-4-1-flash-236b-a21bClosedChina Mobile84.5%Independentgroup defaultsPartially comparable-0.10 ptobs. 11 Sept 2026artificialanalysis.aiT2History
103o3-proClosedOpenAI · OpenAI o-series84.5%Independentgroup defaultsPartially comparable-0.10 ptobs. 11 Sept 2026artificialanalysis.aiT2History
106Gemini 2.5 ProClosedGoogle · Gemini 2.584.4%Independentgroup defaultsPartially comparable-0.21 ptobs. 11 Sept 2026artificialanalysis.aiT2History
107Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 2 rows84.2%Independentgroup defaultsPartially comparable-0.41 ptobs. 11 Sept 2026artificialanalysis.aiT2History
108Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 2 rows84.1%Independentgroup defaultsPartially comparable-0.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
109DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V384.0%Independentgroup defaultsPartially comparable-0.61 ptobs. 11 Sept 2026artificialanalysis.aiT2History
110Gemini 3.5 Flash-LiteClosedGoogle · Gemini 3.583.8%Independentgroup defaultsPartially comparable-0.81 ptobs. 11 Sept 2026artificialanalysis.aiT2History
110Kimi K2 ThinkingOpen weightsMoonshot AI · Kimi83.8%Independentgroup defaultsPartially comparable-0.81 ptobs. 11 Sept 2026artificialanalysis.aiT2History
112GPT-5-CodexClosedOpenAI · GPT 583.7%Independentgroup defaultsPartially comparable-0.91 ptobs. 11 Sept 2026artificialanalysis.aiT2History
113gemini-2-5-pro-03-25ClosedGoogle · Gemini 2.583.6%Independentgroup defaultsPartially comparable-1.01 ptobs. 11 Sept 2026artificialanalysis.aiT2History
114Muse Glimmer 30BOpen weightsMeta AI83.5%Independentgroup defaultsPartially comparable-1.11 ptobs. 11 Sept 2026artificialanalysis.aiT2History
114mimo-v2-0206Open weightsXiaomi83.5%Independentgroup defaultsPartially comparable-1.11 ptobs. 11 Sept 2026artificialanalysis.aiT2History
116Claude Sonnet 4.5ClosedAnthropic · Claude · best of 2 rows83.4%IndependentreasoningonPartially comparable-1.22 ptobs. 11 Sept 2026artificialanalysis.aiT2History
116motif-3Open weightsMotif Technologies83.4%Independentgroup defaultsPartially comparable-1.22 ptobs. 11 Sept 2026artificialanalysis.aiT2History
118Step 3.5 FlashOpen weightsStepFun · Step3.5 · best of 2 rows83.1%Independentgroup defaultsPartially comparable-1.52 ptobs. 11 Sept 2026artificialanalysis.aiT2History
119MiniMax M2.1Open weightsMiniMax · MiniMax83.0%Independentgroup defaultsPartially comparable-1.62 ptobs. 11 Sept 2026artificialanalysis.aiT2History
120jt-35b-flashClosedChina Mobile82.9%Independentgroup defaultsPartially comparable-1.72 ptobs. 11 Sept 2026artificialanalysis.aiT2History
120k-exaone-2-0-0803Open weightsLG AI Research · EXAONE 2.082.9%Independentgroup defaultsPartially comparable-1.72 ptobs. 11 Sept 2026artificialanalysis.aiT2History
122gpt-5-miniClosedOpenAI · GPT 5 · best of 3 rows82.8%Independentgroup defaultsPartially comparable-1.82 ptobs. 11 Sept 2026artificialanalysis.aiT2History
122mimo-v2-omniClosedXiaomi82.8%Independentgroup defaultsPartially comparable-1.82 ptobs. 11 Sept 2026artificialanalysis.aiT2History
124o3ClosedOpenAI · OpenAI o-series82.7%Independentgroup defaultsPartially comparable-1.92 ptobs. 11 Sept 2026artificialanalysis.aiT2History
125qwen3-5-omni-plusClosedAlibaba Group · Qwen3.582.6%Independentgroup defaultsPartially comparable-2.02 ptobs. 11 Sept 2026artificialanalysis.aiT2History
126gpt-5-5-instant-06-26ClosedOpenAI · GPT 5.582.3%Independentgroup defaultsPartially comparable-2.33 ptobs. 11 Sept 2026artificialanalysis.aiT2History
127Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.182.2%Independentgroup defaultsPartially comparable-2.43 ptobs. 11 Sept 2026artificialanalysis.aiT2History
127gemini-2-5-pro-05-06ClosedGoogle · Gemini 2.582.2%Independentgroup defaultsPartially comparable-2.43 ptobs. 11 Sept 2026artificialanalysis.aiT2History
129GLM 5Open weightsZ.ai (Zhipu AI) · GLM5 · best of 2 rows82.0%Independentgroup defaultsPartially comparable-2.63 ptobs. 11 Sept 2026artificialanalysis.aiT2History
130gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 3 rows81.7%Independentgroup defaultsPartially comparable-2.93 ptobs. 11 Sept 2026artificialanalysis.aiT2History
131GPT-5.1-Codex MiniClosedOpenAI · GPT 5.181.3%Independentgroup defaultsPartially comparable-3.34 ptobs. 11 Sept 2026artificialanalysis.aiT2History
131deepseek-r1Open weightsDeepSeek · DeepSeek-R181.3%Independentgroup defaultsPartially comparable-3.34 ptobs. 11 Sept 2026artificialanalysis.aiT2History
133gemini-3-flashClosedGoogle · Gemini 381.2%Independentgroup defaultsPartially comparable-3.44 ptobs. 11 Sept 2026artificialanalysis.aiT2History
134ernie-4-5-300b-a47bOpen weightsBaidu · ERNIE 4.581.1%Independentgroup defaultsPartially comparable-3.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
134nova-2-0-liteClosedAmazon Web Services · Nova 2.0 · best of 4 rows81.1%IndependentreasoningonPartially comparable-3.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
136GLM 5V TurboClosedZ.ai (Zhipu AI) · GLM580.9%Independentgroup defaultsPartially comparable-3.74 ptobs. 11 Sept 2026artificialanalysis.aiT2History
136Step 3.7 FlashOpen weightsStepFun · Step3.780.9%Independentgroup defaultsPartially comparable-3.74 ptobs. 11 Sept 2026artificialanalysis.aiT2History
138Claude Opus 4.1ClosedAnthropic · Claude80.9%IndependentreasoningonPartially comparable-3.75 ptobs. 11 Sept 2026artificialanalysis.aiT2History
139Qwen3.5-9BOpen weightsQwen · Qwen3.5 · best of 2 rows80.6%Independentgroup defaultsPartially comparable-4.04 ptobs. 11 Sept 2026artificialanalysis.aiT2History
140g9v3-39a5bOpen weightsAI9Stars80.5%Independentgroup defaultsPartially comparable-4.14 ptobs. 11 Sept 2026artificialanalysis.aiT2History
141Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 380%Independentgroup defaultsPartially comparable-4.65 ptobs. 11 Sept 2026artificialanalysis.aiT2History
142DeepSeek V3.2 ExpOpen weightsDeepSeek · DeepSeek-V379.7%Independentgroup defaultsPartially comparable-4.95 ptobs. 11 Sept 2026artificialanalysis.aiT2History
143Claude Opus 4ClosedAnthropic · Claude79.6%Independentgroup defaultsPartially comparable-5.05 ptobs. 11 Sept 2026artificialanalysis.aiT2History
144exaone-4-5-33bOpen weightsLG AI Research · EXAONE 4.579.4%Independentgroup defaultsPartially comparable-5.26 ptobs. 11 Sept 2026artificialanalysis.aiT2History
145gemini-2-5-flash-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 2 rows79.3%IndependentreasoningonPartially comparable-5.36 ptobs. 11 Sept 2026artificialanalysis.aiT2History
146DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 2 rows79.2%IndependentreasoningonPartially comparable-5.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
146Gemma 4 26B A4BOpen weightsGoogle · Gemma 4 · best of 2 rows79.2%Independentgroup defaultsPartially comparable-5.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
148grok-3-mini-reasoningClosedSpaceXAI · Grok 379.1%Independentgroup defaultsPartially comparable-5.56 ptobs. 11 Sept 2026artificialanalysis.aiT2History
149Gemini 2.5 FlashClosedGoogle · Gemini 2.5 · best of 2 rows79.0%IndependentreasoningonPartially comparable-5.66 ptobs. 11 Sept 2026artificialanalysis.aiT2History
149Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 2 rows79.0%Independentgroup defaultsPartially comparable-5.66 ptobs. 11 Sept 2026artificialanalysis.aiT2History
151nova-2-0-proClosedAmazon Web Services · Nova 2.0 · best of 3 rows78.5%Independentreasoningonreasoning_effortmediumPartially comparable-6.17 ptobs. 11 Sept 2026artificialanalysis.aiT2History
152o4-miniClosedOpenAI · OpenAI o-series78.4%Independentgroup defaultsPartially comparable-6.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
153k-exaoneOpen weightsLG AI Research · EXAONE · best of 2 rows78.3%Independentgroup defaultsPartially comparable-6.37 ptobs. 11 Sept 2026artificialanalysis.aiT2History
154GLM 4.5Open weightsZ.ai (Zhipu AI) · GLM4.578.2%Independentgroup defaultsPartially comparable-6.47 ptobs. 11 Sept 2026artificialanalysis.aiT2History
154gpt-oss-120bOpen weightsOpenAI · gpt-oss · best of 2 rows78.2%Independentgroup defaultsPartially comparable-6.47 ptobs. 11 Sept 2026artificialanalysis.aiT2History
156GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows78.0%IndependentreasoningonPartially comparable-6.67 ptobs. 11 Sept 2026artificialanalysis.aiT2History
156LongCat 2.0Open weightsMeituan78.0%Independentgroup defaultsPartially comparable-6.67 ptobs. 11 Sept 2026artificialanalysis.aiT2History
158DeepSeek V3.1Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows77.9%IndependentreasoningonPartially comparable-6.77 ptobs. 11 Sept 2026artificialanalysis.aiT2History
159Claude Sonnet 4ClosedAnthropic · Claude · best of 2 rows77.7%IndependentreasoningonPartially comparable-6.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
159MiniMax M2Open weightsMiniMax · MiniMax77.7%Independentgroup defaultsPartially comparable-6.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
159ernie-5-0-thinking-previewClosedBaidu · ERNIE 5.077.7%Independentgroup defaultsPartially comparable-6.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
162grok-4.20-0309-non-reasoningClosedxAI · Grok77.6%Independentgroup defaultsPartially comparable-7.07 ptobs. 11 Sept 2026artificialanalysis.aiT2History
162qwen3-max-thinking-previewClosedAlibaba Group · Qwen377.6%Independentgroup defaultsPartially comparable-7.07 ptobs. 11 Sept 2026artificialanalysis.aiT2History
164ring-1tOpen weightsinclusionAI77.4%Independentgroup defaultsPartially comparable-7.28 ptobs. 11 Sept 2026artificialanalysis.aiT2History
165o3-miniClosedOpenAI · OpenAI o-series · best of 2 rows77.3%Independentreasoning_efforthighPartially comparable-7.38 ptobs. 11 Sept 2026artificialanalysis.aiT2History
166Claude 3.7 SonnetClosedAnthropic · Claude · best of 2 rows77.2%IndependentreasoningonPartially comparable-7.48 ptobs. 11 Sept 2026artificialanalysis.aiT2History
166Qwen3 VL 235B A22B InstructOpen weightsQwen · Qwen3 · best of 2 rows77.2%IndependentreasoningonPartially comparable-7.48 ptobs. 11 Sept 2026artificialanalysis.aiT2History
168Qwen3.5-4BOpen weightsQwen · Qwen3.5 · best of 2 rows77.1%Independentgroup defaultsPartially comparable-7.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
169Mercury 2ClosedInception77.0%Independentgroup defaultsPartially comparable-7.68 ptobs. 11 Sept 2026artificialanalysis.aiT2History
170Mistral Small 4Open weightsMistral AI · Mistral · best of 2 rows76.9%Independentgroup defaultsPartially comparable-7.78 ptobs. 11 Sept 2026artificialanalysis.aiT2History
171cogito-v2-1-reasoningOpen weightsDeep Cogito · Cogito76.8%Independentgroup defaultsPartially comparable-7.88 ptobs. 11 Sept 2026artificialanalysis.aiT2History
172Kimi K2 0905Open weightsMoonshot AI · Kimi76.7%Independentgroup defaultsPartially comparable-7.98 ptobs. 11 Sept 2026artificialanalysis.aiT2History
173Kimi K2 0711Open weightsMoonshot AI · Kimi76.6%Independentgroup defaultsPartially comparable-8.08 ptobs. 11 Sept 2026artificialanalysis.aiT2History
174o1 PreviewClosedOpenAI · OpenAI o-series76.5%Independentgroup defaultsPartially comparable-8.19 ptobs. 11 Sept 2026artificialanalysis.aiT2History
175doubao-seed-codeClosedByteDance · Seed76.4%Independentgroup defaultsPartially comparable-8.29 ptobs. 11 Sept 2026artificialanalysis.aiT2History
175kat-coder-pro-v1ClosedKwaiKAT76.4%Independentgroup defaultsPartially comparable-8.29 ptobs. 11 Sept 2026artificialanalysis.aiT2History
175qwen3-max-previewClosedAlibaba Group · Qwen376.4%Independentgroup defaultsPartially comparable-8.29 ptobs. 11 Sept 2026artificialanalysis.aiT2History
178command-a-plusOpen weightsCohere · Command76.1%Independentgroup defaultsPartially comparable-8.59 ptobs. 11 Sept 2026artificialanalysis.aiT2History
178intellect-3Open weightsPrime Intellect76.1%Independentgroup defaultsPartially comparable-8.59 ptobs. 11 Sept 2026artificialanalysis.aiT2History
180nova-2-0-omniClosedAmazon Web Services · Nova 2.0 · best of 3 rows76.0%Independentreasoningonreasoning_effortmediumPartially comparable-8.69 ptobs. 11 Sept 2026artificialanalysis.aiT2History
181Qwen3 Next 80B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows75.9%IndependentreasoningonPartially comparable-8.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
182nemotron-cascade-2-30b-a3bOpen weightsNVIDIA · Nemotron75.8%Independentgroup defaultsPartially comparable-8.89 ptobs. 11 Sept 2026artificialanalysis.aiT2History
183Nemotron 3 Nano 30B A3BRestricted weightsNVIDIA · Nemotron 3 · best of 2 rows75.7%IndependentreasoningonPartially comparable-8.99 ptobs. 11 Sept 2026artificialanalysis.aiT2History
183North Mini Code (free)Open weightsCohere75.7%Independentgroup defaultsPartially comparable-8.99 ptobs. 11 Sept 2026artificialanalysis.aiT2History
185gemma-4-12BOpen weightsGoogle · Gemma 4 · best of 2 rows75.3%Independentgroup defaultsPartially comparable-9.40 ptobs. 11 Sept 2026artificialanalysis.aiT2History
186Trinity Large ThinkingOpen weightsArcee AI75.2%Independentgroup defaultsPartially comparable-9.50 ptobs. 11 Sept 2026artificialanalysis.aiT2History
186ling-2-6-1tOpen weightsinclusionAI75.2%Independentgroup defaultsPartially comparable-9.50 ptobs. 11 Sept 2026artificialanalysis.aiT2History
188DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 2 rows75.0%Independentgroup defaultsPartially comparable-9.60 ptobs. 11 Sept 2026artificialanalysis.aiT2History
189Mistral Medium 3.5Open weightsMistral AI · Mistral74.8%Independentgroup defaultsPartially comparable-9.80 ptobs. 11 Sept 2026artificialanalysis.aiT2History
189llama-nemotron-super-49b-v1-5Open weightsNVIDIA · Llama · best of 2 rows74.8%IndependentreasoningonPartially comparable-9.80 ptobs. 11 Sept 2026artificialanalysis.aiT2History
191o1ClosedOpenAI · OpenAI o-series74.8%Independentgroup defaultsPartially comparable-9.90 ptobs. 11 Sept 2026artificialanalysis.aiT2History
192NVIDIA Nemotron 3.5 Lightning 30B A3BOpen weightsNVIDIA · Nemotron 3.574.3%Independentgroup defaultsPartially comparable-10.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
193qwen3-5-omni-flashClosedAlibaba Group · Qwen3.574.2%Independentgroup defaultsPartially comparable-10.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
194Magistral Medium 1.2ClosedMistral AI · Magistral73.9%Independentgroup defaultsPartially comparable-10.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
194exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows73.9%IndependentreasoningonPartially comparable-10.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
196deepseek-v3-2-0925Open weightsDeepSeek · DeepSeek73.8%Independentgroup defaultsPartially comparable-10.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
196sarvam-105bOpen weightsSarvam73.8%Independentgroup defaultsPartially comparable-10.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
198Qwen3 Coder NextOpen weightsQwen · Qwen373.7%Independentgroup defaultsPartially comparable-10.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
199ling-3-0-tinyOpen weightsinclusionAI73.4%Independentgroup defaultsPartially comparable-11.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
200GLM 4.5 AirOpen weightsZ.ai (Zhipu AI) · GLM4.573.3%Independentgroup defaultsPartially comparable-11.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →