Skip to content
AI Atlas
BenchmarkActivecategory · reasoningfamily · gpqa · variant Diamond

GPQA Diamond

github.com/idavidrein/gpqa

graduate-level science questions — the 198-question Diamond subset (expert-validated, non-expert-failed)

data quality57

Updated 8 h ago · first seen 12 Sept 2026

Metric
accuracy · %
Current results
1,227
Models
459
Current leader
gpt-6-astra 96.3%

Score history · GLM 5 4 rows

Score history for GLM 50%20%40%60%80%100%Sept 26Sept 26Sept 26Sept 26Sept 26
  • GLM 5
  • 82.02%aa_slug=glm-5 · variant=Diamond · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
  • 66.57%aa_slug=glm-5-non-reasoning · variant=Diamond · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
  • 82.02%aa_slug=glm-5 · variant=GPQA Diamond · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 66.57%aa_slug=glm-5-non-reasoning · variant=GPQA Diamond · evaluator=Artificial Analysis · reasoning=off11 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=GPQA Diamond · evaluator=Artificial Analysis

9 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 96.3%gpt-6-astra OpenAI Independent11 Sept 2026
  2. 96.1%gpt-6-astra OpenAI Independent11 Sept 2026
  3. 95.3%Gemini 3.8 Flash Google Independent11 Sept 2026
  4. 95.0%gpt-6-astra OpenAI Independent11 Sept 2026
  5. 93.9%gpt-6-astra OpenAI Independent11 Sept 2026
  6. 93.5%Kimi K3 Moonshot AI Independent11 Sept 2026
  7. 91.9%Claude Opus 5 Anthropic Independent11 Sept 2026
  8. 79.1%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 459 models · trust independent-evaluator

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
401Mixtral 8x22BOpen weightsMistral AI · Mixtral 833.2%Independentgroup defaultsPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
402dbrxOpen weightsDatabricks33.1%Independentgroup defaultsPartially comparable-0.10 ptobs. 11 Sept 2026artificialanalysis.aiT2History
402phi-4-miniOpen weightsMicrosoft · Phi433.1%Independentgroup defaultsPartially comparable-0.10 ptobs. 11 Sept 2026artificialanalysis.aiT2History
404claude-instantClosedAnthropic · Claude33.0%Independentgroup defaultsPartially comparable-0.20 ptobs. 11 Sept 2026artificialanalysis.aiT2History
405olmo-2-32bOpen weightsAllen Institute for AI · OLMo 232.8%Independentgroup defaultsPartially comparable-0.40 ptobs. 11 Sept 2026artificialanalysis.aiT2History
406lfm-40bClosedLiquid AI · LFM32.7%Independentgroup defaultsPartially comparable-0.50 ptobs. 11 Sept 2026artificialanalysis.aiT2History
406llama-2-chat-70bOpen weightsMeta AI · Llama 232.7%Independentgroup defaultsPartially comparable-0.50 ptobs. 11 Sept 2026artificialanalysis.aiT2History
408LFM2.5-1.2B-InstructRestricted weightsLiquid AI · LFM2.532.6%Independentgroup defaultsPartially comparable-0.60 ptobs. 11 Sept 2026artificialanalysis.aiT2History
409gemini-1-5-flash-may-2024ClosedGoogle · Gemini 1.532.4%Independentgroup defaultsPartially comparable-0.81 ptobs. 11 Sept 2026artificialanalysis.aiT2History
410command-r-plus-04-2024Open weightsCohere · Command32.3%Independentgroup defaultsPartially comparable-0.91 ptobs. 11 Sept 2026artificialanalysis.aiT2History
411jamba-1-7-miniOpen weightsAI21 Labs · Jamba 1.732.2%Independentgroup defaultsPartially comparable-1.01 ptobs. 11 Sept 2026artificialanalysis.aiT2History
412llama-2-chat-13bOpen weightsMeta AI · Llama 232.1%Independentgroup defaultsPartially comparable-1.11 ptobs. 11 Sept 2026artificialanalysis.aiT2History
413claude-21ClosedAnthropic · Claude 2131.9%Independentgroup defaultsPartially comparable-1.31 ptobs. 11 Sept 2026artificialanalysis.aiT2History
413deepseek-coder-v2-liteOpen weightsDeepSeek · DeepSeek31.9%Independentgroup defaultsPartially comparable-1.31 ptobs. 11 Sept 2026artificialanalysis.aiT2History
413phi-3-miniOpen weightsMicrosoft · Phi331.9%Independentgroup defaultsPartially comparable-1.31 ptobs. 11 Sept 2026artificialanalysis.aiT2History
416phi-4-multimodalOpen weightsMicrosoft · Phi431.5%Independentgroup defaultsPartially comparable-1.71 ptobs. 11 Sept 2026artificialanalysis.aiT2History
417granite-4.1-3bOpen weightsIBM · Granite 4.131.4%Independentgroup defaultsPartially comparable-1.82 ptobs. 11 Sept 2026artificialanalysis.aiT2History
418lfm2-2-6bOpen weightsLiquid AI · LFM2.230.6%Independentgroup defaultsPartially comparable-2.62 ptobs. 11 Sept 2026artificialanalysis.aiT2History
419minicpm-v4-6-1-3bOpen weightsOpenBMB30.5%Independentgroup defaultsPartially comparable-2.72 ptobs. 11 Sept 2026artificialanalysis.aiT2History
419tiny-aya-globalRestricted weightsCohere · Aya30.5%Independentgroup defaultsPartially comparable-2.72 ptobs. 11 Sept 2026artificialanalysis.aiT2History
421Mistral Small 1.0ClosedMistral AI · Mistral30.2%Independentgroup defaultsPartially comparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
421deepseek-r1-distill-llama-8bOpen weightsDeepSeek · Llama30.2%Independentgroup defaultsPartially comparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
421jamba-1-5-miniOpen weightsAI21 Labs · Jamba 1.530.2%Independentgroup defaultsPartially comparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
424jamba-1-6-miniOpen weightsAI21 Labs · Jamba 1.630%Independentgroup defaultsPartially comparable-3.23 ptobs. 11 Sept 2026artificialanalysis.aiT2History
425gpt-3.5-turboClosedOpenAI · GPT 3.529.7%Independentgroup defaultsPartially comparable-3.53 ptobs. 11 Sept 2026artificialanalysis.aiT2History
426gemma-3n-e4bOpen weightsGoogle · Gemma 329.6%Independentgroup defaultsPartially comparable-3.63 ptobs. 11 Sept 2026artificialanalysis.aiT2History
426llama-3-instruct-8bOpen weightsMeta AI · Llama 329.6%Independentgroup defaultsPartially comparable-3.63 ptobs. 11 Sept 2026artificialanalysis.aiT2History
428Mixtral 8x7BOpen weightsMistral AI · Mixtral 829.2%Independentgroup defaultsPartially comparable-4.04 ptobs. 11 Sept 2026artificialanalysis.aiT2History
429Gemma 3 4BRestricted weightsGoogle · Gemma 329.1%Independentgroup defaultsPartially comparable-4.14 ptobs. 11 Sept 2026artificialanalysis.aiT2History
430LFM2.5-VL-1.6BRestricted weightsLiquid AI · LFM2.528.9%Independentgroup defaultsPartially comparable-4.34 ptobs. 11 Sept 2026artificialanalysis.aiT2History
430qwen1.5-110b-chatOpen weightsAlibaba Group · Qwen1.528.9%Independentgroup defaultsPartially comparable-4.34 ptobs. 11 Sept 2026artificialanalysis.aiT2History
432olmo-2-7bOpen weightsAllen Institute for AI · OLMo 228.8%Independentgroup defaultsPartially comparable-4.44 ptobs. 11 Sept 2026artificialanalysis.aiT2History
433command-r-03-2024Open weightsCohere · Command28.4%Independentgroup defaultsPartially comparable-4.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
434granite-4-0-nano-1bOpen weightsIBM · Granite 4.028.1%Independentgroup defaultsPartially comparable-5.15 ptobs. 11 Sept 2026artificialanalysis.aiT2History
435gemma-3n-e4b-preview-0520Open weightsGoogle · Gemma 327.8%Independentgroup defaultsPartially comparable-5.45 ptobs. 11 Sept 2026artificialanalysis.aiT2History
435minicpm5-1bOpen weightsOpenBMB · best of 2 rows27.8%Independentgroup defaultsPartially comparable-5.45 ptobs. 11 Sept 2026artificialanalysis.aiT2History
437gemini-1-0-proClosedGoogle · Gemini 1.027.7%Independentgroup defaultsPartially comparable-5.55 ptobs. 11 Sept 2026artificialanalysis.aiT2History
438apertus-70b-instructOpen weightsSwiss AI Initiative27.2%Independentgroup defaultsPartially comparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
439deephermes-3-llama-3-1-8b-previewOpen weightsNous Research · Llama 3.127.0%Independentgroup defaultsPartially comparable-6.26 ptobs. 11 Sept 2026artificialanalysis.aiT2History
440granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.026.3%Independentgroup defaultsPartially comparable-6.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
441granite-4-0-350mOpen weightsIBM · Granite 4.026.1%Independentgroup defaultsPartially comparable-7.17 ptobs. 11 Sept 2026artificialanalysis.aiT2History
442Llama 3.1 8BRestricted weightsMeta AI · Llama 3.125.9%Independentgroup defaultsPartially comparable-7.37 ptobs. 11 Sept 2026artificialanalysis.aiT2History
443granite-4-0-h-350mOpen weightsIBM · Granite 4.025.7%Independentgroup defaultsPartially comparable-7.57 ptobs. 11 Sept 2026artificialanalysis.aiT2History
444apertus-8b-instructOpen weightsSwiss AI Initiative25.6%Independentgroup defaultsPartially comparable-7.67 ptobs. 11 Sept 2026artificialanalysis.aiT2History
445Llama-3.2-3BRestricted weightsMeta AI · Llama 3.225.4%Independentgroup defaultsPartially comparable-7.78 ptobs. 11 Sept 2026artificialanalysis.aiT2History
446molmo-7b-dOpen weightsAllen Institute for AI · Molmo24.0%Independentgroup defaultsPartially comparable-9.19 ptobs. 11 Sept 2026artificialanalysis.aiT2History
447Qwen3-0.6BOpen weightsQwen · Qwen3.023.9%Independentgroup defaultsPartially comparable-9.29 ptobs. 11 Sept 2026artificialanalysis.aiT2History
448gemma-3-1bOpen weightsGoogle · Gemma 323.7%Independentgroup defaultsPartially comparable-9.49 ptobs. 11 Sept 2026artificialanalysis.aiT2History
449Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 2 rows23.6%IndependentreasoningoffPartially comparable-9.59 ptobs. 11 Sept 2026artificialanalysis.aiT2History
450qwen3-0.6b-instructOpen weightsAlibaba Group · Qwen3.023.1%Independentgroup defaultsPartially comparable-10.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
451openchat-35Open weightsOpenChat23.0%Independentgroup defaultsPartially comparable-10.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
452gemma-3n-e2bOpen weightsGoogle · Gemma 322.9%Independentgroup defaultsPartially comparable-10.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
453LFM2-1.2BRestricted weightsLiquid AI · LFM2.122.8%Independentgroup defaultsPartially comparable-10.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
454llama-2-chat-7bOpen weightsMeta AI · Llama 222.7%Independentgroup defaultsPartially comparable-10.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
455gemma-3-270mOpen weightsGoogle · Gemma 322.4%Independentgroup defaultsPartially comparable-10.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
456llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.222.1%Independentgroup defaultsPartially comparable-11.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
457Llama-3.2-1BRestricted weightsMeta AI · Llama 3.219.6%Independentgroup defaultsPartially comparable-13.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
458Mistral 7BOpen weightsMistral AI · Mistral17.7%Independentgroup defaultsPartially comparable-15.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
459DeepSeek-R1-Distill-Qwen-1.5BOpen weightsDeepSeek · Qwen9.80%Independentgroup defaultsPartially comparable-23.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →