Skip to content
AI Atlas
Benchmarks

Benchmarks

Evaluation suites and the results published for them. Each result carries its configuration; leaders are shown per benchmark, not as a composite.

9 of 27 benchmarks

Benchmarks
BenchmarkCategoryMetricResultsModelsCurrent leaderUpdated
Aider polyglotcodingpass rate (2 attempts) (%)0024 min ago
HumanEvalcodingpass@1 (%)0024 min ago
LiveCodeBenchcodingpass@1 (%)0024 min ago
SciCodecodingaccuracy (%)0024 min ago
SWE-bench (full test split)codingresolved (%)0024 min ago
SWE-bench Litecodingresolved (%)0024 min ago
SWE-bench Multilingualcodingresolved (%)0024 min ago
SWE-bench Multimodalcodingresolved (%)0024 min ago
SWE-bench Verifiedcodingresolved (%)0024 min ago

Leader = best current result under the benchmark's default direction (higher or lower is better). Configs differ; open a benchmark for its leaderboard, config filter and per-model history.