Benchmarks
Benchmarks
Evaluation suites and the results published for them. Each result carries its configuration; leaders are shown per benchmark, not as a composite.
2 of 27 benchmarks
| Benchmark | Category | Metric | Results | Models | Current leader | Updated |
|---|---|---|---|---|---|---|
| AIME 2025 | math | accuracy (%) | 0 | 0 | — | 26 min ago |
| MATH-500 | math | accuracy (%) | 0 | 0 | — | 26 min ago |
Leader = best current result under the benchmark's default direction (higher or lower is better). Configs differ; open a benchmark for its leaderboard, config filter and per-model history.