Skip to content
AI Atlas
BenchmarkActivecategory · reasoningfamily · gpqa · variant main

GPQA

github.com/idavidrein/gpqa

graduate-level science questions (main set, 448 questions)

quality57

Updated 12 min ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
0
Models
0
Current leader

No results recorded for this benchmark yet — its sources are being connected. The definition, aliases and variants are kept so links resolve; nothing is fabricated.

Score history · qwen3-vl-235b-a22b-reasoning 1 row

Not enough history to chart — a single observation (77.17% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.

  • 77.17%aa_slug=qwen3-vl-235b-a22b-reasoning · variant=GPQA Diamond · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Leaderboard 0 models

Select models with +, then Compare.

No result in this group with these filters

Relax the trust / organization filters or pick another comparability group.