GPQA Diamond
graduate-level science questions — the 198-question Diamond subset (expert-validated, non-expert-failed)
Updated 6 h ago · first seen 12 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 1,227
- Models
- 459
- Current leader
- gpt-6-astra 96.3%
Score history · Inkling Small 2 rows
- Inkling Small
- 89.49%aa_slug=inkling-small · variant=Diamond · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
- 89.49%aa_slug=inkling-small · variant=GPQA Diamond · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
Frontier over time · accuracy · variant=GPQA Diamond · evaluator=Artificial Analysis
9 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 96.3%gpt-6-astra OpenAI Independent11 Sept 2026
- 96.1%gpt-6-astra OpenAI Independent11 Sept 2026
- 95.3%Gemini 3.8 Flash Google Independent11 Sept 2026
- 95.0%gpt-6-astra OpenAI Independent11 Sept 2026
- 93.9%gpt-6-astra OpenAI Independent11 Sept 2026
- 93.5%Kimi K3 Moonshot AI Independent11 Sept 2026
- 91.9%Claude Opus 5 Anthropic Independent11 Sept 2026
- 79.1%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 459 models · trust independent-evaluator
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 301 | claude-35-sonnet-june-24ClosedAnthropic · Claude 35 | 56.0% | Independent | group defaults | Partially comparable0.00 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 302 | granite-4.2-3bOpen weightsIBM · Granite 4.2 | 55.9% | Independent | group defaults | Partially comparable-0.10 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 303 | LFM2.5-2.6B (free)Open weightsLiquid AI · LFM2.5 | 55.8% | Independent | group defaults | Partially comparable-0.20 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 304 | QwQ-32B-PreviewOpen weightsAlibaba Group · Qwen | 55.7% | Independent | group defaults | Partially comparable-0.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 305 | gpt-4oClosedOpenAI · GPT 4 | 54.3% | Independent | group defaults | Partially comparable-1.62 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 306 | gemini-2-0-flash-lite-previewClosedGoogle · Gemini 2.0 | 54.2% | Independent | group defaults | Partially comparable-1.72 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 307 | tri-21b-think-previewOpen weightsTrillion Labs | 53.8% | Independent | group defaults | Partially comparable-2.12 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 308 | hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.1 | 53.6% | Independent | group defaults | Partially comparable-2.32 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 309 | gemini-2-0-flash-lite-001ClosedGoogle · Gemini 2.0 | 53.5% | Independent | group defaults | Partially comparable-2.42 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 309 | qwen3-32b-instructOpen weightsAlibaba Group · Qwen3 | 53.5% | Independent | group defaults | Partially comparable-2.42 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 311 | Devstral Small 2Open weightsMistral AI · Devstral | 53.2% | Independent | group defaults | Partially comparable-2.73 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 312 | Reka Flash 3Open weightsrekaai | 52.9% | Independent | group defaults | Partially comparable-3.03 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 313 | Command AOpen weightsCohere · Command | 52.7% | Independent | group defaults | Partially comparable-3.23 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 314 | GPT-4o (2024-05-13)ClosedOpenAI · GPT 4 | 52.6% | Independent | group defaults | Partially comparable-3.33 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 315 | Qwen3-4BOpen weightsQwen · Qwen3 | 52.2% | Independent | group defaults | Partially comparable-3.74 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 316 | GPT-4o (2024-08-06)ClosedOpenAI · GPT 4 | 52.1% | Independent | group defaults | Partially comparable-3.84 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 317 | Olmo-3-7B-ThinkOpen weightsAllen Institute for AI · OLMo 3 | 51.6% | Independent | group defaults | Partially comparable-4.34 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 317 | Qwen3 Coder 30B A3B InstructOpen weightsQwen · Qwen3 | 51.6% | Independent | group defaults | Partially comparable-4.34 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 317 | tulu3-405bOpen weightsAllen Institute for AI | 51.6% | Independent | group defaults | Partially comparable-4.34 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 320 | Llama-3.1-405BRestricted weightsMeta AI · Llama 3.1 | 51.5% | Independent | group defaults | Partially comparable-4.44 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 320 | exaone-4-0-1-2bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows | 51.5% | Independent | reasoningon | Partially comparable-4.44 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 322 | LFM2.5-8B-A1BOpen weightsLiquid AI · LFM2.5 | 51.3% | Independent | group defaults | Partially comparable-4.65 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 322 | nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 3 | 51.3% | Independent | group defaults | Partially comparable-4.65 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 324 | gpt-4.1-nanoClosedOpenAI · GPT 4.1 | 51.2% | Independent | group defaults | Partially comparable-4.75 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 325 | gpt-4o-chatgptClosedOpenAI · GPT 4 | 51.1% | Independent | group defaults | Partially comparable-4.85 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 326 | grok-2Open weightsxAI · Grok 2 | 51.0% | Independent | group defaults | Partially comparable-4.95 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 327 | Mistral Small 3.2Open weightsMistral AI · Mistral | 50.5% | Independent | group defaults | Partially comparable-5.45 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 327 | Pixtral LargeOpen weightsMistral AI · Pixtral | 50.5% | Independent | group defaults | Partially comparable-5.45 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 329 | nova-proClosedAmazon Web Services · Nova | 49.9% | Independent | group defaults | Partially comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 330 | Llama 3.3 70BOpen weightsMeta AI · Llama 3.3 | 49.8% | Independent | group defaults | Partially comparable-6.16 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 331 | Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 2 rows | 49.4% | Independent | reasoningon | Partially comparable-6.57 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 332 | devstral-mediumClosedMistral AI · Devstral | 49.2% | Independent | group defaults | Partially comparable-6.77 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 333 | Qwen2.5 72B InstructOpen weightsQwen · Qwen2.5 | 49.1% | Independent | group defaults | Partially comparable-6.87 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 333 | hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.1 | 49.1% | Independent | group defaults | Partially comparable-6.87 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 335 | claude-3-opusClosedAnthropic · Claude 3 | 48.9% | Independent | group defaults | Partially comparable-7.07 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 336 | mistral-large-2Open weightsMistral AI · Mistral | 48.6% | Independent | group defaults | Partially comparable-7.37 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 337 | deepseek-r1-distill-qwen-14bOpen weightsDeepSeek · Qwen | 48.4% | Independent | group defaults | Partially comparable-7.58 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 338 | granite-4.1-30bOpen weightsIBM · Granite 4.1 | 48.1% | Independent | group defaults | Partially comparable-7.88 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 339 | lfm2-24b-a2bOpen weightsLiquid AI · LFM2 | 47.4% | Independent | group defaults | Partially comparable-8.59 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 340 | Mistral Large 2.0Open weightsMistral AI · Mistral | 47.2% | Independent | group defaults | Partially comparable-8.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 341 | Ministral 3 8BOpen weightsMistral AI · Ministral 3 | 47.1% | Independent | group defaults | Partially comparable-8.89 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 341 | grok-betaClosedSpaceXAI · Grok | 47.1% | Independent | group defaults | Partially comparable-8.89 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 343 | qwen3-14b-instructOpen weightsAlibaba Group · Qwen3 | 47.0% | Independent | group defaults | Partially comparable-8.99 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 344 | nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 3 | 46.9% | Independent | group defaults | Partially comparable-9.09 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 345 | qwen2.5-32b-instructOpen weightsAlibaba Group · Qwen2.5 | 46.6% | Independent | group defaults | Partially comparable-9.39 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 346 | llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.1 | 46.5% | Independent | group defaults | Partially comparable-9.50 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 347 | gemini-1-5-flashClosedGoogle · Gemini 1.5 | 46.3% | Independent | group defaults | Partially comparable-9.70 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 348 | Mistral Small 3Open weightsMistral AI · Mistral | 46.2% | Independent | group defaults | Partially comparable-9.80 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 349 | Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 2 rows | 45.6% | Independent | group defaults | Partially comparable-10.4 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 350 | Mistral Small 3.1Open weightsMistral AI · Mistral | 45.4% | Independent | group defaults | Partially comparable-10.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 351 | qwen3-8b-instructOpen weightsAlibaba Group · Qwen3 | 45.1% | Independent | group defaults | Partially comparable-10.8 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 352 | g9v3-3bOpen weightsAI9Stars | 43.8% | Independent | group defaults | Partially comparable-12.1 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 353 | Devstral Small 1.0Open weightsMistral AI · Devstral | 43.4% | Independent | group defaults | Partially comparable-12.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 354 | gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 2 rows | 43.3% | Independent | group defaults | Partially comparable-12.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 354 | granite-4.1-8bOpen weightsIBM · Granite 4.1 | 43.3% | Independent | group defaults | Partially comparable-12.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 354 | nova-liteClosedAmazon Web Services · Nova | 43.3% | Independent | group defaults | Partially comparable-12.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 357 | llama-3-2-instruct-90b-visionOpen weightsMeta AI · Llama 3.2 | 43.2% | Independent | group defaults | Partially comparable-12.7 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 358 | Gemma 3 27BOpen weightsGoogle · Gemma 3 | 42.8% | Independent | group defaults | Partially comparable-13.1 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 359 | jamba-1-5-largeOpen weightsAI21 Labs · Jamba 1.5 | 42.7% | Independent | group defaults | Partially comparable-13.2 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 360 | gpt-4o-miniClosedOpenAI · GPT 4 | 42.6% | Independent | group defaults | Partially comparable-13.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 361 | Molmo2-8BOpen weightsAllen Institute for AI | 42.5% | Independent | group defaults | Partially comparable-13.4 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 362 | Mistral SabaClosedMistral AI · Mistral | 42.4% | Independent | group defaults | Partially comparable-13.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 363 | deepseek-v2-5Open weightsDeepSeek · DeepSeek | 42.3% | Independent | group defaults | Partially comparable-13.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 364 | Qwen2.5 Coder 32B InstructOpen weightsQwen · Qwen2.5 | 41.7% | Independent | group defaults | Partially comparable-14.2 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 365 | granite-4-0-h-smallOpen weightsIBM · Granite 4.0 | 41.6% | Independent | group defaults | Partially comparable-14.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 365 | sarvam-m-reasoningOpen weightsSarvam | 41.6% | Independent | group defaults | Partially comparable-14.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 367 | devstral-smallOpen weightsMistral AI · Devstral | 41.4% | Independent | group defaults | Partially comparable-14.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 368 | Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AI · Kimi | 41.2% | Independent | group defaults | Partially comparable-14.8 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 369 | qwen-turboClosedAlibaba Group · Qwen | 41.0% | Independent | group defaults | Partially comparable-15.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 370 | Llama-3.1-70BOpen weightsMeta AI · Llama 3.1 | 40.9% | Independent | group defaults | Partially comparable-15.1 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 371 | Claude Haiku 3.5ClosedAnthropic · Claude | 40.8% | Independent | group defaults | Partially comparable-15.1 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 371 | llama-3-1-nemotron-nano-4b-reasoningOpen weightsNVIDIA · Llama 3.1 | 40.8% | Independent | group defaults | Partially comparable-15.1 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 373 | R1 Distill Llama 70BOpen weightsDeepSeek · Llama | 40.2% | Independent | group defaults | Partially comparable-15.8 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 374 | Hermes 3 70B InstructOpen weightsNous Research · Hermes 3 | 40.1% | Independent | group defaults | Partially comparable-15.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 375 | Olmo-3-7B-InstructOpen weightsAllen Institute for AI · OLMo 3 | 40% | Independent | group defaults | Partially comparable-16.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 375 | claude-3-sonnetClosedAnthropic · Claude 3 | 40% | Independent | group defaults | Partially comparable-16.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 377 | qwen3-4b-instructOpen weightsAlibaba Group · Qwen3 | 39.8% | Independent | group defaults | Partially comparable-16.2 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 378 | jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.7 | 39.0% | Independent | group defaults | Partially comparable-17.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 379 | jamba-1-6-largeOpen weightsAI21 Labs · Jamba 1.6 | 38.7% | Independent | group defaults | Partially comparable-17.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 380 | deephermes-3-mistral-24b-previewOpen weightsNous Research · Mistral | 38.2% | Independent | group defaults | Partially comparable-17.8 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 381 | mistral-smallOpen weightsMistral AI · Mistral | 38.1% | Independent | group defaults | Partially comparable-17.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 382 | llama-3-instruct-70bOpen weightsMeta AI · Llama 3 | 37.9% | Independent | group defaults | Partially comparable-18.1 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 383 | Claude 3 HaikuClosedAnthropic · Claude | 37.4% | Independent | group defaults | Partially comparable-18.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 384 | gemini-1-5-pro-may-2024ClosedGoogle · Gemini 1.5 | 37.1% | Independent | group defaults | Partially comparable-18.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 384 | qwen2-72b-instructOpen weightsAlibaba Group · Qwen2 | 37.1% | Independent | group defaults | Partially comparable-18.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 386 | gemini-1-5-flash-8bClosedGoogle · Gemini 1.5 | 35.9% | Independent | group defaults | Partially comparable-20.1 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 387 | Ministral 3 3BOpen weightsMistral AI · Ministral 3 | 35.8% | Independent | group defaults | Partially comparable-20.2 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 387 | nova-microClosedAmazon Web Services · Nova | 35.8% | Independent | group defaults | Partially comparable-20.2 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 389 | qwen3-1.7b-instructOpen weightsAlibaba Group · Qwen3.1 · best of 2 rows | 35.6% | Independent | reasoningon | Partially comparable-20.4 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 390 | Mistral LargeClosedMistral AI · Mistral | 35.0% | Independent | group defaults | Partially comparable-20.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 391 | Gemma 3 12BOpen weightsGoogle · Gemma 3 | 35.0% | Independent | group defaults | Partially comparable-21.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 391 | gpt-4ClosedOpenAI · GPT 4 | 35.0% | Independent | group defaults | Partially comparable-21.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 391 | mistral-mediumClosedMistral AI · Mistral | 35.0% | Independent | group defaults | Partially comparable-21.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 394 | claude-2ClosedAnthropic · Claude 2 | 34.4% | Independent | group defaults | Partially comparable-21.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 394 | lfm2-8b-a1bOpen weightsLiquid AI · LFM2 | 34.4% | Independent | group defaults | Partially comparable-21.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 396 | Qwen2.5-Coder-7BOpen weightsQwen · Qwen2.5 | 33.9% | Independent | group defaults | Partially comparable-22.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 396 | lfm2-5-1-2b-thinkingOpen weightsLiquid AI · LFM2.5 | 33.9% | Independent | group defaults | Partially comparable-22.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 398 | granite-3-3-8b-instructOpen weightsIBM · Granite 3.3 | 33.8% | Independent | group defaults | Partially comparable-22.1 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 399 | Granite 4.0 MicroOpen weightsIBM · Granite 4.0 | 33.6% | Independent | group defaults | Partially comparable-22.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 400 | jamba-reasoning-3bOpen weightsAI21 Labs · Jamba | 33.3% | Independent | group defaults | Partially comparable-22.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →