- Rows
- 60 of 456 eligible
- Columns
- 12
- Cells
- 534 · 241 partially comparable
Colour
Within-column rank of the model's best comparable current score: darkest = best rank in that column. Colours never compare across columns; the number in the cell is the score itself.
lastbest
Trust levels present
IndependentOfficial board
Comparability
Comparable same task and conditions · Partially comparable same task, conditions differ (effort, temperature, judge) · Not comparable different variant or metric
Rows are grouped by comparability group = canonical metric × config_key (hash of the task-defining configuration keys: variant, evaluator, harness, shots, pass regime…). A leaderboard shows one row per canonical model: its best current row inside the group. Reasoning effort, temperature or judge differences keep rows in the same group but mark them partially comparable. Each cell is the model's best current row in the benchmark's primary group; `mean_rank` is only a sort key, not a composite score.
`mean_rank` orders the rows; it is a sort key, not a score. Methodology →