Skip to content
AI Atlas
BenchmarkActivecategory · reasoningfamily · gpqa · variant Diamond

GPQA Diamond

github.com/idavidrein/gpqa

graduate-level science questions — the 198-question Diamond subset (expert-validated, non-expert-failed)

data quality57

Updated 4 h ago · first seen 12 Sept 2026

Metric
accuracy · %
Current results
1,227
Models
459
Current leader
gpt-6-astra 96.3%

Frontier over time · accuracy · variant=Diamond · evaluator=Artificial Analysis

9 leader changes recorded, all dated 12 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 96.3%gpt-6-astra OpenAI Independent12 Sept 2026
  2. 96.1%gpt-6-astra OpenAI Independent12 Sept 2026
  3. 95.3%Gemini 3.8 Flash Google Independent12 Sept 2026
  4. 95.0%gpt-6-astra OpenAI Independent12 Sept 2026
  5. 93.9%gpt-6-astra OpenAI Independent12 Sept 2026
  6. 93.5%Kimi K3 Moonshot AI Independent12 Sept 2026
  7. 91.9%Claude Opus 5 Anthropic Independent12 Sept 2026
  8. 79.1%grok-3-mini-reasoning SpaceXAI Independent12 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 440 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
200qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows72.6%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
202ring-flash-2-0Open weightsinclusionAI72.5%IndependentreasoningonPartially comparable-0.10 ptobs. 12 Sept 2026artificialanalysis.aiT2History
203Solar Pro 3ClosedUpstage · Solar72.4%IndependentreasoningonPartially comparable-0.21 ptobs. 12 Sept 2026artificialanalysis.aiT2History
204midm-250-pro-rsnsftClosedKorea Telecom72.2%IndependentreasoningonPartially comparable-0.41 ptobs. 12 Sept 2026artificialanalysis.aiT2History
205Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows72.0%IndependentreasoningonPartially comparable-0.61 ptobs. 12 Sept 2026artificialanalysis.aiT2History
206GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows71.9%IndependentreasoningonPartially comparable-0.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
206ling-1tOpen weightsinclusionAI71.9%IndependentreasoningoffPartially comparable-0.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
208apriel-v1-5-15b-thinkerOpen weightsServiceNow71.3%IndependentreasoningonPartially comparable-1.32 ptobs. 12 Sept 2026artificialanalysis.aiT2History
208k2-think-v2Open weightsMBZUAI Institute of Foundation Models71.3%IndependentreasoningonPartially comparable-1.32 ptobs. 12 Sept 2026artificialanalysis.aiT2History
210Gemini 2.5 Flash-LiteClosedGoogle · Gemini 2.5 · best of 3 rows70.9%IndependentreasoningonPartially comparable-1.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
211deepseek-r1-0120Open weightsDeepSeek · DeepSeek70.8%IndependentreasoningonPartially comparable-1.82 ptobs. 12 Sept 2026artificialanalysis.aiT2History
212qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 2 rows70.7%IndependentreasoningonPartially comparable-1.92 ptobs. 12 Sept 2026artificialanalysis.aiT2History
213minicpm5-2bOpen weightsOpenBMB70.2%IndependentreasoningonPartially comparable-2.43 ptobs. 12 Sept 2026artificialanalysis.aiT2History
214gemini-2.0-flash-thinking-exp-01-21ClosedGoogle · Gemini 2.070.1%IndependentreasoningonPartially comparable-2.53 ptobs. 12 Sept 2026artificialanalysis.aiT2History
214mi-dm-k-2-5-pro-dec28ClosedKorea Telecom70.1%IndependentreasoningonPartially comparable-2.53 ptobs. 12 Sept 2026artificialanalysis.aiT2History
216qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows70%IndependentreasoningonPartially comparable-2.63 ptobs. 12 Sept 2026artificialanalysis.aiT2History
217hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.1 · best of 2 rows69.9%IndependentreasoningonPartially comparable-2.73 ptobs. 12 Sept 2026artificialanalysis.aiT2History
218gemini-2-5-flash-reasoning-04-2025ClosedGoogle · Gemini 2.5 · best of 2 rows69.8%IndependentreasoningonPartially comparable-2.83 ptobs. 12 Sept 2026artificialanalysis.aiT2History
219MiniMax-M1-80kOpen weightsMiniMax · MiniMax69.7%IndependentreasoningonPartially comparable-2.93 ptobs. 12 Sept 2026artificialanalysis.aiT2History
220motif-2-12-7bClosedMotif Technologies69.5%IndependentreasoningonPartially comparable-3.14 ptobs. 12 Sept 2026artificialanalysis.aiT2History
221grok-3ClosedSpaceXAI · Grok 369.3%IndependentreasoningoffPartially comparable-3.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
222step-3-vl-10bOpen weightsStepFun · Step369.0%IndependentreasoningonPartially comparable-3.64 ptobs. 12 Sept 2026artificialanalysis.aiT2History
223gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 2 rows68.8%Independentreasoning_efforthighPartially comparable-3.84 ptobs. 12 Sept 2026artificialanalysis.aiT2History
224solar-pro-2ClosedUpstage · Solar · best of 3 rows68.7%IndependentreasoningonPartially comparable-3.94 ptobs. 12 Sept 2026artificialanalysis.aiT2History
225gpt-5-chatgptClosedOpenAI · GPT 568.6%IndependentreasoningoffPartially comparable-4.04 ptobs. 12 Sept 2026artificialanalysis.aiT2History
226GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows68.4%IndependentreasoningonPartially comparable-4.25 ptobs. 12 Sept 2026artificialanalysis.aiT2History
227MiniMax-M1-40kOpen weightsMiniMax · MiniMax68.2%IndependentreasoningonPartially comparable-4.45 ptobs. 12 Sept 2026artificialanalysis.aiT2History
228k2-v2Open weightsMBZUAI Institute of Foundation Models · best of 3 rows68.1%Independentreasoning_efforthighPartially comparable-4.55 ptobs. 12 Sept 2026artificialanalysis.aiT2History
229Mistral Large 3Open weightsMistral AI · Mistral68.0%IndependentreasoningoffPartially comparable-4.65 ptobs. 12 Sept 2026artificialanalysis.aiT2History
230magistral-mediumClosedMistral AI · Magistral67.9%IndependentreasoningonPartially comparable-4.75 ptobs. 12 Sept 2026artificialanalysis.aiT2History
231gpt-5-nanoClosedOpenAI · GPT 5 · best of 3 rows67.6%Independentreasoning_efforthighPartially comparable-5.05 ptobs. 12 Sept 2026artificialanalysis.aiT2History
231jt-miniClosedChina Mobile67.6%IndependentreasoningoffPartially comparable-5.05 ptobs. 12 Sept 2026artificialanalysis.aiT2History
233Claude Haiku 4.5ClosedAnthropic · Claude · best of 2 rows67.2%IndependentreasoningonPartially comparable-5.46 ptobs. 12 Sept 2026artificialanalysis.aiT2History
234Llama 4 MaverickOpen weightsMeta AI · Llama 467.1%IndependentreasoningoffPartially comparable-5.56 ptobs. 12 Sept 2026artificialanalysis.aiT2History
235diffusiongemma-26b-a4bOpen weightsGoogle66.9%IndependentreasoningonPartially comparable-5.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
236qwen3-32b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows66.8%IndependentreasoningonPartially comparable-5.86 ptobs. 12 Sept 2026artificialanalysis.aiT2History
237qwen3-4b-2507-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows66.7%IndependentreasoningonPartially comparable-5.96 ptobs. 12 Sept 2026artificialanalysis.aiT2History
238gpt-4.1ClosedOpenAI · GPT 4.166.6%IndependentreasoningoffPartially comparable-6.06 ptobs. 12 Sept 2026artificialanalysis.aiT2History
239gpt-4.1-miniClosedOpenAI · GPT 4.166.4%IndependentreasoningoffPartially comparable-6.27 ptobs. 12 Sept 2026artificialanalysis.aiT2History
240Magistral Small 1.2Open weightsMistral AI · Magistral66.3%IndependentreasoningonPartially comparable-6.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
241falcon-h1r-7bOpen weightsTII UAE · Falcon66.1%IndependentreasoningonPartially comparable-6.57 ptobs. 12 Sept 2026artificialanalysis.aiT2History
242ling-flash-2-0Open weightsinclusionAI65.7%IndependentreasoningoffPartially comparable-6.97 ptobs. 12 Sept 2026artificialanalysis.aiT2History
242solar-open-100b-reasoningOpen weightsUpstage · Solar65.7%IndependentreasoningonPartially comparable-6.97 ptobs. 12 Sept 2026artificialanalysis.aiT2History
244DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V365.5%IndependentreasoningoffPartially comparable-7.18 ptobs. 12 Sept 2026artificialanalysis.aiT2History
244gpt-4o-chatgpt-03-25ClosedOpenAI · GPT 465.5%IndependentreasoningoffPartially comparable-7.18 ptobs. 12 Sept 2026artificialanalysis.aiT2History
246gemini-2-5-flash-lite-preview-09-2025ClosedGoogle · Gemini 2.565.0%IndependentreasoningoffPartially comparable-7.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
247granite-4.2-30bOpen weightsIBM · Granite 4.264.4%IndependentreasoningonPartially comparable-8.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
248llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.3 · best of 2 rows64.3%IndependentreasoningonPartially comparable-8.29 ptobs. 12 Sept 2026artificialanalysis.aiT2History
249magistral-smallOpen weightsMistral AI · Magistral64.1%IndependentreasoningonPartially comparable-8.49 ptobs. 12 Sept 2026artificialanalysis.aiT2History
250gemini-2.0-flash-expClosedGoogle · Gemini 2.063.6%IndependentreasoningoffPartially comparable-8.99 ptobs. 12 Sept 2026artificialanalysis.aiT2History
250longcat-flash-liteOpen weightsLongCat63.6%IndependentreasoningoffPartially comparable-8.99 ptobs. 12 Sept 2026artificialanalysis.aiT2History
252sarvam-30bOpen weightsSarvam63.3%Independentreasoning_efforthighPartially comparable-9.30 ptobs. 12 Sept 2026artificialanalysis.aiT2History
253Granite 4.2 8BOpen weightsIBM · Granite 4.263.1%IndependentreasoningonPartially comparable-9.50 ptobs. 12 Sept 2026artificialanalysis.aiT2History
253celeris-1ClosedCeleris63.1%IndependentreasoningoffPartially comparable-9.50 ptobs. 12 Sept 2026artificialanalysis.aiT2History
255Gemini 2.0 FlashClosedGoogle · Gemini 2.062.3%IndependentreasoningoffPartially comparable-10.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
255SonarClosedPerplexity AI · Sonar · best of 2 rows62.3%IndependentreasoningonPartially comparable-10.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
257gemini-2-0-pro-experimental-02-05ClosedGoogle · Gemini 2.062.2%IndependentreasoningoffPartially comparable-10.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
258qwen3-coder-480b-a35b-instructOpen weightsAlibaba Group · Qwen3-Coder61.8%IndependentreasoningoffPartially comparable-10.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
259qwen3-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows61.6%IndependentreasoningonPartially comparable-11.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
260DeepSeek-R1-Distill-Qwen-32BOpen weightsDeepSeek · Qwen61.5%IndependentreasoningonPartially comparable-11.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
260hyperclova-x-seed-think-32bOpen weightsNaver · Seed61.5%IndependentreasoningonPartially comparable-11.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
262DeepSeek-R1-0528-Qwen3-8BOpen weightsDeepSeek · Qwen361.2%IndependentreasoningonPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
263olmo-3-32b-thinkOpen weightsAllen Institute for AI · OLMo 361.0%IndependentreasoningonPartially comparable-11.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
264qwen3-14b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows60.4%IndependentreasoningonPartially comparable-12.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
265o1-miniClosedOpenAI · OpenAI o-series60.3%IndependentreasoningonPartially comparable-12.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
266tri-21b-think-v0-5Open weightsTrillion Labs60.1%IndependentreasoningonPartially comparable-12.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
267claude-35-sonnetClosedAnthropic · Claude 3559.9%IndependentreasoningoffPartially comparable-12.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
268Devstral 2Open weightsMistral AI · Devstral 259.4%IndependentreasoningoffPartially comparable-13.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
269QwQ-32BOpen weightsAlibaba Group · Qwen59.3%IndependentreasoningonPartially comparable-13.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
269ling-2-6-flashOpen weightsinclusionAI59.3%IndependentreasoningoffPartially comparable-13.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
271olmo-3-1-32b-instructOpen weightsAllen Institute for AI · OLMo 3.1 · best of 2 rows59.1%IndependentreasoningonPartially comparable-13.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
272gemini-1-5-proClosedGoogle · Gemini 1.558.9%IndependentreasoningoffPartially comparable-13.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
272qwen3-8b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows58.9%IndependentreasoningonPartially comparable-13.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
274Mistral Medium 3.1ClosedMistral AI · Mistral58.8%IndependentreasoningoffPartially comparable-13.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
275Llama 4 ScoutOpen weightsMeta AI · Llama 458.7%IndependentreasoningoffPartially comparable-13.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
275qwen-2-5-maxClosedAlibaba Group · Qwen58.7%IndependentreasoningoffPartially comparable-13.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
277GLM 4.7 FlashOpen weightsZ.ai (Zhipu AI) · GLM4.7 · best of 2 rows58.1%IndependentreasoningonPartially comparable-14.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
278Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 2 rows57.9%IndependentreasoningonPartially comparable-14.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279Mistral Medium 3ClosedMistral AI · Mistral57.8%IndependentreasoningoffPartially comparable-14.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279Sonar ProClosedPerplexity AI · Sonar57.8%IndependentreasoningoffPartially comparable-14.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
281gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 2 rows57.6%IndependentreasoningonPartially comparable-15.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
282Phi 4Open weightsMicrosoft · Phi457.5%IndependentreasoningoffPartially comparable-15.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
283Ministral 3 14BOpen weightsMistral AI · Ministral 357.2%IndependentreasoningoffPartially comparable-15.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
283nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 2 rows57.2%IndependentreasoningonPartially comparable-15.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
285NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 2 rows57.0%IndependentreasoningonPartially comparable-15.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
286nova-premierClosedAmazon Web Services · Nova56.9%IndependentreasoningoffPartially comparable-15.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287ling-mini-2-0Open weightsinclusionAI56.2%IndependentreasoningoffPartially comparable-16.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
288claude-35-sonnet-june-24ClosedAnthropic · Claude 3556.0%IndependentreasoningoffPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
289granite-4.2-3bOpen weightsIBM · Granite 4.255.9%IndependentreasoningonPartially comparable-16.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
290LFM2.5-2.6B (free)Open weightsLiquid AI · LFM2.555.8%IndependentreasoningonPartially comparable-16.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
291QwQ-32B-PreviewOpen weightsAlibaba Group · Qwen55.7%IndependentreasoningonPartially comparable-17.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
292solar-pro-2-previewClosedUpstage · Solar54.4%IndependentreasoningoffPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
293gpt-4oClosedOpenAI · GPT 454.3%IndependentreasoningoffPartially comparable-18.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
294gemini-2-0-flash-lite-previewClosedGoogle · Gemini 2.054.2%IndependentreasoningoffPartially comparable-18.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
295tri-21b-think-previewOpen weightsTrillion Labs53.8%IndependentreasoningonPartially comparable-18.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
296gemini-2-0-flash-lite-001ClosedGoogle · Gemini 2.053.5%IndependentreasoningoffPartially comparable-19.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
297Devstral Small 2Open weightsMistral AI · Devstral53.2%IndependentreasoningoffPartially comparable-19.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
298Reka Flash 3Open weightsrekaai52.9%IndependentreasoningonPartially comparable-19.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
299Command AOpen weightsCohere · Command52.7%IndependentreasoningoffPartially comparable-19.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
300GPT-4o (2024-05-13)ClosedOpenAI · GPT 452.6%IndependentreasoningoffPartially comparable-20.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →