Benchmarks
Benchmarks
Evaluation suites and the results published for them. Each result carries its configuration; leaders are shown per benchmark, not as a composite.
9 of 27 benchmarks
| Benchmark | Category | Metric | Results | Models | Current leader | Updated |
|---|---|---|---|---|---|---|
| Aider polyglot | coding | pass rate (2 attempts) (%) | 0 | 0 | — | 24 min ago |
| HumanEval | coding | pass@1 (%) | 0 | 0 | — | 24 min ago |
| LiveCodeBench | coding | pass@1 (%) | 0 | 0 | — | 24 min ago |
| SciCode | coding | accuracy (%) | 0 | 0 | — | 24 min ago |
| SWE-bench (full test split) | coding | resolved (%) | 0 | 0 | — | 24 min ago |
| SWE-bench Lite | coding | resolved (%) | 0 | 0 | — | 24 min ago |
| SWE-bench Multilingual | coding | resolved (%) | 0 | 0 | — | 24 min ago |
| SWE-bench Multimodal | coding | resolved (%) | 0 | 0 | — | 24 min ago |
| SWE-bench Verified | coding | resolved (%) | 0 | 0 | — | 24 min ago |
Leader = best current result under the benchmark's default direction (higher or lower is better). Configs differ; open a benchmark for its leaderboard, config filter and per-model history.