Benchmarks
Benchmarks
Evaluation suites and the results published for them. Each result carries its configuration; leaders are shown per benchmark, not as a composite.
27 benchmarks
| Benchmark | Category | Metric | Results | Models | Current leader | Updated |
|---|---|---|---|---|---|---|
| τ-bench | agentic | pass^1 (%) | 0 | 0 | — | 19 min ago |
| τ²-bench | agentic | pass^1 (%) | 0 | 0 | — | 19 min ago |
| Aider polyglot | coding | pass rate (2 attempts) (%) | 0 | 0 | — | 19 min ago |
| AIME 2025 | math | accuracy (%) | 0 | 0 | — | 19 min ago |
| ARC-AGI | reasoning | accuracy (%) | 0 | 0 | — | 19 min ago |
| Artificial Analysis Intelligence Index | composite | index | 0 | 0 | — | 19 min ago |
| GPQA | reasoning | accuracy (%) | 0 | 0 | — | 19 min ago |
| HumanEval | coding | pass@1 (%) | 0 | 0 | — | 19 min ago |
| Humanity's Last Exam | knowledge | accuracy (%) | 0 | 0 | — | 19 min ago |
| IFBench | instruction-following | accuracy (%) | 0 | 0 | — | 19 min ago |
| IFEval | instruction-following | prompt-level strict accuracy (%) | 0 | 0 | — | 19 min ago |
| LiveBench | general | average score (%) | 0 | 0 | — | 19 min ago |
| LiveCodeBench | coding | pass@1 (%) | 0 | 0 | — | 19 min ago |
| LMArena text leaderboard | preference | Elo / Bradley–Terry score | 0 | 0 | — | 19 min ago |
| MATH-500 | math | accuracy (%) | 0 | 0 | — | 19 min ago |
| MMLU | knowledge | accuracy (%) | 0 | 0 | — | 19 min ago |
| MMLU-Pro | knowledge | accuracy (%) | 0 | 0 | — | 19 min ago |
| MMMU | multimodal | accuracy (%) | 0 | 0 | — | 19 min ago |
| MMMU-Pro | multimodal | accuracy (%) | 0 | 0 | — | 19 min ago |
| MTEB | embeddings | mean score | 0 | 0 | — | 19 min ago |
| SciCode | coding | accuracy (%) | 0 | 0 | — | 19 min ago |
| SWE-bench (full test split) | coding | resolved (%) | 0 | 0 | — | 19 min ago |
| SWE-bench Lite | coding | resolved (%) | 0 | 0 | — | 19 min ago |
| SWE-bench Multilingual | coding | resolved (%) | 0 | 0 | — | 19 min ago |
| SWE-bench Multimodal | coding | resolved (%) | 0 | 0 | — | 19 min ago |
| SWE-bench Verified | coding | resolved (%) | 0 | 0 | — | 19 min ago |
| Terminal-Bench | agentic | accuracy (%) | 0 | 0 | — | 19 min ago |
Leader = best current result under the benchmark's default direction (higher or lower is better). Configs differ; open a benchmark for its leaderboard, config filter and per-model history.