Skip to content
AI Atlas
BenchmarkActivecategory · reasoningfamily · gpqa · variant Diamond

GPQA Diamond

github.com/idavidrein/gpqa

graduate-level science questions — the 198-question Diamond subset (expert-validated, non-expert-failed)

data quality57

Updated 6 h ago · first seen 12 Sept 2026

Metric
accuracy · %
Current results
1,227
Models
459
Current leader
gpt-6-astra 96.3%

Frontier over time · accuracy · variant=Diamond · evaluator=Artificial Analysis

9 leader changes recorded, all dated 12 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 96.3%gpt-6-astra OpenAI Independent12 Sept 2026
  2. 96.1%gpt-6-astra OpenAI Independent12 Sept 2026
  3. 95.3%Gemini 3.8 Flash Google Independent12 Sept 2026
  4. 95.0%gpt-6-astra OpenAI Independent12 Sept 2026
  5. 93.9%gpt-6-astra OpenAI Independent12 Sept 2026
  6. 93.5%Kimi K3 Moonshot AI Independent12 Sept 2026
  7. 91.9%Claude Opus 5 Anthropic Independent12 Sept 2026
  8. 79.1%grok-3-mini-reasoning SpaceXAI Independent12 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 440 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
401minicpm-v4-6-1-3bOpen weightsOpenBMB30.5%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
401tiny-aya-globalRestricted weightsCohere · Aya30.5%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
403Mistral Small 1.0ClosedMistral AI · Mistral30.2%IndependentreasoningoffPartially comparable-0.31 ptobs. 12 Sept 2026artificialanalysis.aiT2History
403deepseek-r1-distill-llama-8bOpen weightsDeepSeek · Llama30.2%IndependentreasoningonPartially comparable-0.31 ptobs. 12 Sept 2026artificialanalysis.aiT2History
403jamba-1-5-miniOpen weightsAI21 Labs · Jamba 1.530.2%IndependentreasoningoffPartially comparable-0.31 ptobs. 12 Sept 2026artificialanalysis.aiT2History
406jamba-1-6-miniOpen weightsAI21 Labs · Jamba 1.630%IndependentreasoningoffPartially comparable-0.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
407gpt-3.5-turboClosedOpenAI · GPT 3.529.7%IndependentreasoningoffPartially comparable-0.81 ptobs. 12 Sept 2026artificialanalysis.aiT2History
408gemma-3n-e4bOpen weightsGoogle · Gemma 329.6%IndependentreasoningoffPartially comparable-0.91 ptobs. 12 Sept 2026artificialanalysis.aiT2History
408llama-3-instruct-8bOpen weightsMeta AI · Llama 329.6%IndependentreasoningoffPartially comparable-0.91 ptobs. 12 Sept 2026artificialanalysis.aiT2History
410Mixtral 8x7BOpen weightsMistral AI · Mixtral 829.2%IndependentreasoningoffPartially comparable-1.32 ptobs. 12 Sept 2026artificialanalysis.aiT2History
411Gemma 3 4BRestricted weightsGoogle · Gemma 329.1%IndependentreasoningoffPartially comparable-1.42 ptobs. 12 Sept 2026artificialanalysis.aiT2History
412LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.528.9%IndependentreasoningoffPartially comparable-1.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
412qwen1.5-110b-chatOpen weightsAlibaba Group · Qwen1.528.9%IndependentreasoningoffPartially comparable-1.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
414olmo-2-7bOpen weightsAllen Institute for AI · OLMo 228.8%IndependentreasoningoffPartially comparable-1.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
415command-r-03-2024Open weightsCohere · Command28.4%IndependentreasoningoffPartially comparable-2.13 ptobs. 12 Sept 2026artificialanalysis.aiT2History
416granite-4-0-nano-1bOpen weightsIBM · Granite 4.028.1%IndependentreasoningoffPartially comparable-2.43 ptobs. 12 Sept 2026artificialanalysis.aiT2History
417gemma-3n-e4b-preview-0520Open weightsGoogle · Gemma 327.8%IndependentreasoningoffPartially comparable-2.73 ptobs. 12 Sept 2026artificialanalysis.aiT2History
417minicpm5-1bOpen weightsOpenBMB · best of 2 rows27.8%IndependentreasoningonPartially comparable-2.73 ptobs. 12 Sept 2026artificialanalysis.aiT2History
419gemini-1-0-proClosedGoogle · Gemini 1.027.7%IndependentreasoningoffPartially comparable-2.83 ptobs. 12 Sept 2026artificialanalysis.aiT2History
420apertus-70b-instructOpen weightsSwiss AI Initiative27.2%IndependentreasoningoffPartially comparable-3.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
421deephermes-3-llama-3-1-8b-previewOpen weightsNous Research · Llama 3.127.0%IndependentreasoningoffPartially comparable-3.54 ptobs. 12 Sept 2026artificialanalysis.aiT2History
422granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.026.3%IndependentreasoningoffPartially comparable-4.25 ptobs. 12 Sept 2026artificialanalysis.aiT2History
423granite-4-0-350mOpen weightsIBM · Granite 4.026.1%IndependentreasoningoffPartially comparable-4.45 ptobs. 12 Sept 2026artificialanalysis.aiT2History
424Llama 3.1 8BRestricted weightsMeta AI · Llama 3.125.9%IndependentreasoningoffPartially comparable-4.65 ptobs. 12 Sept 2026artificialanalysis.aiT2History
425granite-4-0-h-350mOpen weightsIBM · Granite 4.025.7%IndependentreasoningoffPartially comparable-4.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
426apertus-8b-instructOpen weightsSwiss AI Initiative25.6%IndependentreasoningoffPartially comparable-4.95 ptobs. 12 Sept 2026artificialanalysis.aiT2History
427Llama-3.2-3BRestricted weightsMeta AI · Llama 3.225.4%IndependentreasoningoffPartially comparable-5.06 ptobs. 12 Sept 2026artificialanalysis.aiT2History
428molmo-7b-dOpen weightsAllen Institute for AI · Molmo24.0%IndependentreasoningoffPartially comparable-6.47 ptobs. 12 Sept 2026artificialanalysis.aiT2History
429qwen3-0.6b-instructOpen weightsAlibaba Group · Qwen3.0 · best of 2 rows23.9%IndependentreasoningonPartially comparable-6.57 ptobs. 12 Sept 2026artificialanalysis.aiT2History
430gemma-3-1bOpen weightsGoogle · Gemma 323.7%IndependentreasoningoffPartially comparable-6.77 ptobs. 12 Sept 2026artificialanalysis.aiT2History
431Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 2 rows23.6%IndependentreasoningoffPartially comparable-6.87 ptobs. 12 Sept 2026artificialanalysis.aiT2History
432openchat-35Open weightsOpenChat23.0%IndependentreasoningoffPartially comparable-7.48 ptobs. 12 Sept 2026artificialanalysis.aiT2History
433gemma-3n-e2bOpen weightsGoogle · Gemma 322.9%IndependentreasoningoffPartially comparable-7.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
434LFM2-1.2BOpen weightsLiquid AI · LFM2.122.8%IndependentreasoningoffPartially comparable-7.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
435llama-2-chat-7bOpen weightsMeta AI · Llama 222.7%IndependentreasoningoffPartially comparable-7.78 ptobs. 12 Sept 2026artificialanalysis.aiT2History
436gemma-3-270mOpen weightsGoogle · Gemma 322.4%IndependentreasoningoffPartially comparable-8.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
437llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.222.1%IndependentreasoningoffPartially comparable-8.39 ptobs. 12 Sept 2026artificialanalysis.aiT2History
438Llama-3.2-1BRestricted weightsMeta AI · Llama 3.219.6%IndependentreasoningoffPartially comparable-10.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
439Mistral 7BOpen weightsMistral AI · Mistral17.7%IndependentreasoningoffPartially comparable-12.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
440DeepSeek-R1-Distill-Qwen-1.5BOpen weightsDeepSeek · Qwen9.80%IndependentreasoningonPartially comparable-20.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →