BenchmarkActivecategory · reasoningfamily · gpqa · variant main
data quality57
Updated 1 h ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 0
- Models
- 0
- Current leader
- —
No results recorded for this benchmark yet — its sources are being connected. The definition, aliases and variants are kept so links resolve; nothing is fabricated.
Score history · grok-4-fast-reasoning 1 row
Not enough history to chart — a single observation (84.75% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.
- 84.75%aa_slug=grok-4-fast-reasoning · variant=GPQA Diamond · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
Leaderboard 0 models
Select models with +, then Compare.
No result in this group with these filters
Relax the trust / organization filters or pick another comparability group.