Skip to content
AI Atlas
BenchmarkActivecategory · reasoningfamily · gpqa · variant Diamond

GPQA Diamond

github.com/idavidrein/gpqa

graduate-level science questions — the 198-question Diamond subset (expert-validated, non-expert-failed)

data quality57

Updated 5 h ago · first seen 12 Sept 2026

Metric
accuracy · %
Current results
1,227
Models
459
Current leader
gpt-6-astra 96.3%

Score history · Mistral Large 3 2 rows

Score history for Mistral Large 30%20%40%60%80%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • Mistral Large 3
  • 67.98%aa_slug=mistral-large-3 · variant=Diamond · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
  • 67.98%aa_slug=mistral-large-3 · variant=GPQA Diamond · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=GPQA Diamond · evaluator=Artificial Analysis

9 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 96.3%gpt-6-astra OpenAI Independent11 Sept 2026
  2. 96.1%gpt-6-astra OpenAI Independent11 Sept 2026
  3. 95.3%Gemini 3.8 Flash Google Independent11 Sept 2026
  4. 95.0%gpt-6-astra OpenAI Independent11 Sept 2026
  5. 93.9%gpt-6-astra OpenAI Independent11 Sept 2026
  6. 93.5%Kimi K3 Moonshot AI Independent11 Sept 2026
  7. 91.9%Claude Opus 5 Anthropic Independent11 Sept 2026
  8. 79.1%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 459 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
200Qwen3 VL 32B InstructOpen weightsQwen · Qwen3 · best of 2 rows73.3%IndependentreasoningonPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
200apriel-v1-6-15b-thinkerOpen weightsServiceNow73.3%Independentgroup defaultsPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
200hypernova-60bOpen weightsMultiverse Computing73.3%Independentgroup defaultsPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
204quasar-438bClosedMultiverse Computing73.2%Independentgroup defaultsPartially comparable-0.10 ptobs. 11 Sept 2026artificialanalysis.aiT2History
205llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.172.8%Independentgroup defaultsPartially comparable-0.50 ptobs. 11 Sept 2026artificialanalysis.aiT2History
206Grok Build 0.1ClosedxAI · Grok72.7%Independentgroup defaultsPartially comparable-0.60 ptobs. 11 Sept 2026artificialanalysis.aiT2History
206Hermes 4 405BOpen weightsNous Research · Hermes 472.7%Independentgroup defaultsPartially comparable-0.60 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208Seed-OSS-36B-InstructOpen weightsByteDance · Seed72.6%Independentgroup defaultsPartially comparable-0.70 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows72.6%IndependentreasoningonPartially comparable-0.70 ptobs. 11 Sept 2026artificialanalysis.aiT2History
210ring-flash-2-0Open weightsinclusionAI72.5%Independentgroup defaultsPartially comparable-0.80 ptobs. 11 Sept 2026artificialanalysis.aiT2History
211Solar Pro 3ClosedUpstage · Solar72.4%Independentgroup defaultsPartially comparable-0.91 ptobs. 11 Sept 2026artificialanalysis.aiT2History
212midm-250-pro-rsnsftClosedKorea Telecom72.2%Independentgroup defaultsPartially comparable-1.11 ptobs. 11 Sept 2026artificialanalysis.aiT2History
213Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows72.0%IndependentreasoningonPartially comparable-1.31 ptobs. 11 Sept 2026artificialanalysis.aiT2History
214GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows71.9%Independentgroup defaultsPartially comparable-1.41 ptobs. 11 Sept 2026artificialanalysis.aiT2History
214ling-1tOpen weightsinclusionAI71.9%Independentgroup defaultsPartially comparable-1.41 ptobs. 11 Sept 2026artificialanalysis.aiT2History
216deepseek-v4-pro-0424-non-reasoningOpen weightsDeepSeek · DeepSeek71.7%Independentgroup defaultsPartially comparable-1.61 ptobs. 11 Sept 2026artificialanalysis.aiT2History
217deepseek-v4-flash-0420-non-reasoningOpen weightsDeepSeek · DeepSeek71.6%Independentgroup defaultsPartially comparable-1.71 ptobs. 11 Sept 2026artificialanalysis.aiT2History
218apriel-v1-5-15b-thinkerOpen weightsServiceNow71.3%Independentgroup defaultsPartially comparable-2.02 ptobs. 11 Sept 2026artificialanalysis.aiT2History
218k2-think-v2Open weightsMBZUAI Institute of Foundation Models71.3%Independentgroup defaultsPartially comparable-2.02 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220gemini-2-5-flash-lite-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 2 rows70.9%IndependentreasoningonPartially comparable-2.42 ptobs. 11 Sept 2026artificialanalysis.aiT2History
221deepseek-r1-0120Open weightsDeepSeek · DeepSeek70.8%Independentgroup defaultsPartially comparable-2.52 ptobs. 11 Sept 2026artificialanalysis.aiT2History
222qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 2 rows70.7%IndependentreasoningonPartially comparable-2.62 ptobs. 11 Sept 2026artificialanalysis.aiT2History
223minicpm5-2bOpen weightsOpenBMB70.2%Independentgroup defaultsPartially comparable-3.13 ptobs. 11 Sept 2026artificialanalysis.aiT2History
224claude-4-opusClosedAnthropic · Claude 470.1%Independentgroup defaultsPartially comparable-3.23 ptobs. 11 Sept 2026artificialanalysis.aiT2History
224gemini-2.0-flash-thinking-exp-01-21ClosedGoogle · Gemini 2.070.1%Independentgroup defaultsPartially comparable-3.23 ptobs. 11 Sept 2026artificialanalysis.aiT2History
224mi-dm-k-2-5-pro-dec28ClosedKorea Telecom70.1%Independentgroup defaultsPartially comparable-3.23 ptobs. 11 Sept 2026artificialanalysis.aiT2History
227qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows70%IndependentreasoningonPartially comparable-3.33 ptobs. 11 Sept 2026artificialanalysis.aiT2History
228Hermes-4-70BRestricted weightsNous Research · Hermes 469.9%Independentgroup defaultsPartially comparable-3.43 ptobs. 11 Sept 2026artificialanalysis.aiT2History
229gemini-2-5-flash-reasoning-04-2025ClosedGoogle · Gemini 2.569.8%Independentgroup defaultsPartially comparable-3.53 ptobs. 11 Sept 2026artificialanalysis.aiT2History
230MiniMax-M1-80kOpen weightsMiniMax · MiniMax69.7%Independentgroup defaultsPartially comparable-3.63 ptobs. 11 Sept 2026artificialanalysis.aiT2History
231motif-2-12-7bClosedMotif Technologies69.5%Independentgroup defaultsPartially comparable-3.84 ptobs. 11 Sept 2026artificialanalysis.aiT2History
232grok-3ClosedSpaceXAI · Grok 369.3%Independentgroup defaultsPartially comparable-4.04 ptobs. 11 Sept 2026artificialanalysis.aiT2History
233step-3-vl-10bOpen weightsStepFun · Step369.0%Independentgroup defaultsPartially comparable-4.34 ptobs. 11 Sept 2026artificialanalysis.aiT2History
234gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 2 rows68.8%Independentgroup defaultsPartially comparable-4.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
235solar-pro-2ClosedUpstage · Solar · best of 2 rows68.7%IndependentreasoningonPartially comparable-4.64 ptobs. 11 Sept 2026artificialanalysis.aiT2History
236gpt-5-chatgptClosedOpenAI · GPT 568.6%Independentgroup defaultsPartially comparable-4.74 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows68.4%Independentgroup defaultsPartially comparable-4.95 ptobs. 11 Sept 2026artificialanalysis.aiT2History
238MiniMax-M1-40kOpen weightsMiniMax · MiniMax68.2%Independentgroup defaultsPartially comparable-5.15 ptobs. 11 Sept 2026artificialanalysis.aiT2History
239k2-v2Open weightsMBZUAI Institute of Foundation Models · best of 3 rows68.1%Independentgroup defaultsPartially comparable-5.25 ptobs. 11 Sept 2026artificialanalysis.aiT2History
240Mistral Large 3Open weightsMistral AI · Mistral68.0%Independentgroup defaultsPartially comparable-5.35 ptobs. 11 Sept 2026artificialanalysis.aiT2History
241magistral-mediumClosedMistral AI · Magistral67.9%Independentgroup defaultsPartially comparable-5.45 ptobs. 11 Sept 2026artificialanalysis.aiT2History
242gpt-5-nanoClosedOpenAI · GPT 5 · best of 3 rows67.6%Independentgroup defaultsPartially comparable-5.75 ptobs. 11 Sept 2026artificialanalysis.aiT2History
242jt-miniClosedChina Mobile67.6%Independentgroup defaultsPartially comparable-5.75 ptobs. 11 Sept 2026artificialanalysis.aiT2History
244Claude Haiku 4.5ClosedAnthropic · Claude · best of 2 rows67.2%IndependentreasoningonPartially comparable-6.16 ptobs. 11 Sept 2026artificialanalysis.aiT2History
245Llama 4 MaverickOpen weightsMeta AI · Llama 467.1%Independentgroup defaultsPartially comparable-6.26 ptobs. 11 Sept 2026artificialanalysis.aiT2History
246diffusiongemma-26b-a4bOpen weightsGoogle66.9%Independentgroup defaultsPartially comparable-6.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
247Qwen3 32BOpen weightsQwen · Qwen366.8%Independentgroup defaultsPartially comparable-6.56 ptobs. 11 Sept 2026artificialanalysis.aiT2History
248qwen3-4b-2507-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows66.7%IndependentreasoningonPartially comparable-6.66 ptobs. 11 Sept 2026artificialanalysis.aiT2History
249gpt-4.1ClosedOpenAI · GPT 4.166.6%Independentgroup defaultsPartially comparable-6.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
250gpt-4.1-miniClosedOpenAI · GPT 4.166.4%Independentgroup defaultsPartially comparable-6.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
251Magistral Small 1.2Open weightsMistral AI · Magistral66.3%Independentgroup defaultsPartially comparable-7.07 ptobs. 11 Sept 2026artificialanalysis.aiT2History
252falcon-h1r-7bOpen weightsTII UAE · Falcon66.1%Independentgroup defaultsPartially comparable-7.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253ling-flash-2-0Open weightsinclusionAI65.7%Independentgroup defaultsPartially comparable-7.67 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253solar-open-100b-reasoningOpen weightsUpstage · Solar65.7%Independentgroup defaultsPartially comparable-7.67 ptobs. 11 Sept 2026artificialanalysis.aiT2History
255DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V365.5%Independentgroup defaultsPartially comparable-7.88 ptobs. 11 Sept 2026artificialanalysis.aiT2History
255gpt-4o-chatgpt-03-25ClosedOpenAI · GPT 465.5%Independentgroup defaultsPartially comparable-7.88 ptobs. 11 Sept 2026artificialanalysis.aiT2History
257granite-4.2-30bOpen weightsIBM · Granite 4.264.4%Independentgroup defaultsPartially comparable-8.89 ptobs. 11 Sept 2026artificialanalysis.aiT2History
258llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.3 · best of 2 rows64.3%IndependentreasoningonPartially comparable-8.99 ptobs. 11 Sept 2026artificialanalysis.aiT2History
259magistral-smallOpen weightsMistral AI · Magistral64.1%Independentgroup defaultsPartially comparable-9.19 ptobs. 11 Sept 2026artificialanalysis.aiT2History
260gemini-2.0-flash-expClosedGoogle · Gemini 2.063.6%Independentgroup defaultsPartially comparable-9.69 ptobs. 11 Sept 2026artificialanalysis.aiT2History
260longcat-flash-liteOpen weightsLongCat63.6%Independentgroup defaultsPartially comparable-9.69 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262sarvam-30bOpen weightsSarvam63.3%Independentgroup defaultsPartially comparable-10.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
263Granite 4.2 8BOpen weightsIBM · Granite 4.263.1%Independentgroup defaultsPartially comparable-10.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
263celeris-1ClosedCeleris63.1%Independentgroup defaultsPartially comparable-10.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
265Gemini 2.5 Flash-LiteClosedGoogle · Gemini 2.5 · best of 2 rows62.5%Independentgroup defaultsPartially comparable-10.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
266Gemini 2.0 FlashClosedGoogle · Gemini 2.062.3%Independentgroup defaultsPartially comparable-11.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
266SonarClosedPerplexity AI · Sonar · best of 2 rows62.3%IndependentreasoningonPartially comparable-11.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
268gemini-2-0-pro-experimental-02-05ClosedGoogle · Gemini 2.062.2%Independentgroup defaultsPartially comparable-11.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
269qwen3-coder-480b-a35b-instructOpen weightsAlibaba Group · Qwen3-Coder61.8%Independentgroup defaultsPartially comparable-11.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
270qwen3-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows61.6%IndependentreasoningonPartially comparable-11.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
271DeepSeek-R1-Distill-Qwen-32BOpen weightsDeepSeek · Qwen61.5%Independentgroup defaultsPartially comparable-11.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
271hyperclova-x-seed-think-32bOpen weightsNaver · Seed61.5%Independentgroup defaultsPartially comparable-11.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
273DeepSeek-R1-0528-Qwen3-8BOpen weightsDeepSeek · Qwen361.2%Independentgroup defaultsPartially comparable-12.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
274olmo-3-32b-thinkOpen weightsAllen Institute for AI · OLMo 361.0%Independentgroup defaultsPartially comparable-12.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
275Qwen3 14BOpen weightsQwen · Qwen360.4%Independentgroup defaultsPartially comparable-12.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276o1-miniClosedOpenAI · OpenAI o-series60.3%Independentgroup defaultsPartially comparable-13.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
277tri-21b-think-v0-5Open weightsTrillion Labs60.1%Independentgroup defaultsPartially comparable-13.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
278claude-35-sonnetClosedAnthropic · Claude 3559.9%Independentgroup defaultsPartially comparable-13.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
279Devstral 2Open weightsMistral AI · Devstral 259.4%Independentgroup defaultsPartially comparable-13.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
279gemini-2-5-flash-04-2025ClosedGoogle · Gemini 2.559.4%Independentgroup defaultsPartially comparable-13.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
281QwQ-32BOpen weightsAlibaba Group · Qwen59.3%Independentgroup defaultsPartially comparable-14.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
281ling-2-6-flashOpen weightsinclusionAI59.3%Independentgroup defaultsPartially comparable-14.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
283olmo-3-1-32b-instructOpen weightsAllen Institute for AI · OLMo 3.1 · best of 2 rows59.1%IndependentreasoningonPartially comparable-14.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
284Qwen3 8BOpen weightsQwen · Qwen358.9%Independentgroup defaultsPartially comparable-14.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
284gemini-1-5-proClosedGoogle · Gemini 1.558.9%Independentgroup defaultsPartially comparable-14.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
286Mistral Medium 3.1ClosedMistral AI · Mistral58.8%Independentgroup defaultsPartially comparable-14.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
287Llama 4 ScoutOpen weightsMeta AI · Llama 458.7%Independentgroup defaultsPartially comparable-14.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
287qwen-2-5-maxClosedAlibaba Group · Qwen58.7%Independentgroup defaultsPartially comparable-14.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
289GLM 4.7 FlashOpen weightsZ.ai (Zhipu AI) · GLM4.7 · best of 2 rows58.1%Independentgroup defaultsPartially comparable-15.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
290Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 2 rows57.9%IndependentreasoningonPartially comparable-15.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
291Mistral Medium 3ClosedMistral AI · Mistral57.8%Independentgroup defaultsPartially comparable-15.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
291Sonar ProClosedPerplexity AI · Sonar57.8%Independentgroup defaultsPartially comparable-15.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
291solar-pro-2-previewClosedUpstage · Solar · best of 2 rows57.8%IndependentreasoningonPartially comparable-15.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
294gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 2 rows57.6%Independentgroup defaultsPartially comparable-15.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
295Phi 4Open weightsMicrosoft · Phi457.5%Independentgroup defaultsPartially comparable-15.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
296Ministral 3 14BOpen weightsMistral AI · Ministral 357.2%Independentgroup defaultsPartially comparable-16.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
296nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 2 rows57.2%IndependentreasoningonPartially comparable-16.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
298NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 2 rows57.0%Independentgroup defaultsPartially comparable-16.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
299nova-premierClosedAmazon Web Services · Nova56.9%Independentgroup defaultsPartially comparable-16.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
300ling-mini-2-0Open weightsinclusionAI56.2%Independentgroup defaultsPartially comparable-17.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →