Updated 3 h ago · first seen 11 Sept 2026
bench_01M293SPEMVE6W55699099NES3
- Metric
- pass rate (2 attempts) · %
- Direction
- Higher is better
- Results
- 138 · 22 filtered
Leaderboard 22 current results · config contains “0.69.2.dev”
Select models with +, then open Compare.
| # | Model | Score | Config | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|
| #1GPT-4o-mini (2024-07-18)OpenAI | 100% | date=2024-12-21 · command=aider --model gpt-4o-mini-2024-07-18 · dirname=2024-12-21-18-41-18--polyglot-gpt-4o-mini · versions=0.69.2.dev | 21 Dec 2024 | aider.chatT2 | History | |
| #2gemini-2.0-flash-expGoogle | 100% | date=2024-12-22 · command=aider --model gemini/gemini-2.0-flash-exp · dirname=2024-12-22-20-08-13--gemini-2.0-flash-exp-polyglot-whole · versions=0.69.2.dev | 22 Dec 2024 | aider.chatT2 | History | |
| #3Qwen2.5 Coder 32B InstructQwen | 99.6% | date=2024-12-26 · command=aider --model openai/Qwen2.5-Coder-32B-Instruct · dirname=2024-12-26-00-55-20--Qwen2.5-Coder-32B-Instruct · versions=0.69.2.dev | 26 Dec 2024 | aider.chatT2 | History | |
| #4DeepSeek Chat V3 (prev) | 98.7% | date=2024-12-25 · command=aider --model deepseek/deepseek-chat · dirname=2024-12-25-13-31-51--deepseekv3preview-diff2 · versions=0.69.2.dev | 25 Dec 2024 | aider.chatT2 | History | |
| #5gemini-exp-1206 | 98.2% | date=2024-12-22 · command=aider --model gemini/gemini-exp-1206 · dirname=2024-12-22-18-43-25--gemini-exp-1206-polyglot-whole-2 · versions=0.69.2.dev | 22 Dec 2024 | aider.chatT2 | History | |
| #6o1-mini-2024-09-12 | 96.9% | date=2024-12-22 · command=aider --model o1-mini · dirname=2024-12-22-21-26-35--polyglot-o1mini-whole · versions=0.69.2.dev | 22 Dec 2024 | aider.chatT2 | History | |
| #7DeepSeek Chat V2.5 | 92.9% | date=2024-12-21 · command=aider --model deepseek/deepseek-chat · dirname=2024-12-21-20-56-21--polyglot-deepseek-diff · versions=0.69.2.dev | 21 Dec 2024 | aider.chatT2 | History | |
| #8yi-lightning | 92.9% | date=2024-12-23 · command=aider --model openai/yi-lightning · dirname=2024-12-23-01-11-56--yi-test · versions=0.69.2.dev | 23 Dec 2024 | aider.chatT2 | History | |
| #9o1-2024-12-17 (high) | 91.5% | date=2024-12-21 · command=aider --model openrouter/openai/o1 · dirname=2024-12-21-19-23-03--polyglot-o1-hard-diff · versions=0.69.2.dev | 21 Dec 2024 | aider.chatT2 | History | |
| #10Claude Haiku 3.5Anthropic | 91.1% | date=2024-12-21 · command=aider --model claude-3-5-haiku-20241022 · dirname=2024-12-21-21-46-27--polyglot-haiku-diff · versions=0.69.2.dev | 21 Dec 2024 | aider.chatT2 | History | |
| #11Qwen2.5 Coder 32B InstructQwen | 71.6% | date=2024-12-22 · command=aider --model openai/Qwen/Qwen2.5-Coder-32B-Instruct # via hyperbolic · dirname=2024-12-22-13-22-32--polyglot-qwen-diff · versions=0.69.2.dev | 22 Dec 2024 | aider.chatT2 | History | |
| #12o1-2024-12-17 (high) | 61.7% | date=2024-12-21 · command=aider --model openrouter/openai/o1 · dirname=2024-12-21-19-23-03--polyglot-o1-hard-diff · versions=0.69.2.dev | 21 Dec 2024 | aider.chatT2 | History | |
| #13DeepSeek Chat V3 (prev) | 48.4% | date=2024-12-25 · command=aider --model deepseek/deepseek-chat · dirname=2024-12-25-13-31-51--deepseekv3preview-diff2 · versions=0.69.2.dev | 25 Dec 2024 | aider.chatT2 | History | |
| #14gemini-exp-1206 | 38.2% | date=2024-12-22 · command=aider --model gemini/gemini-exp-1206 · dirname=2024-12-22-18-43-25--gemini-exp-1206-polyglot-whole-2 · versions=0.69.2.dev | 22 Dec 2024 | aider.chatT2 | History | |
| #15o1-mini-2024-09-12 | 32.9% | date=2024-12-22 · command=aider --model o1-mini · dirname=2024-12-22-21-26-35--polyglot-o1mini-whole · versions=0.69.2.dev | 22 Dec 2024 | aider.chatT2 | History | |
| #16Claude Haiku 3.5Anthropic | 28% | date=2024-12-21 · command=aider --model claude-3-5-haiku-20241022 · dirname=2024-12-21-21-46-27--polyglot-haiku-diff · versions=0.69.2.dev | 21 Dec 2024 | aider.chatT2 | History | |
| #17gemini-2.0-flash-expGoogle | 22.2% | date=2024-12-22 · command=aider --model gemini/gemini-2.0-flash-exp · dirname=2024-12-22-20-08-13--gemini-2.0-flash-exp-polyglot-whole · versions=0.69.2.dev | 22 Dec 2024 | aider.chatT2 | History | |
| #18DeepSeek Chat V2.5 | 17.8% | date=2024-12-21 · command=aider --model deepseek/deepseek-chat · dirname=2024-12-21-20-56-21--polyglot-deepseek-diff · versions=0.69.2.dev | 21 Dec 2024 | aider.chatT2 | History | |
| #19Qwen2.5 Coder 32B InstructQwen | 16.4% | date=2024-12-26 · command=aider --model openai/Qwen2.5-Coder-32B-Instruct · dirname=2024-12-26-00-55-20--Qwen2.5-Coder-32B-Instruct · versions=0.69.2.dev | 26 Dec 2024 | aider.chatT2 | History | |
| #20yi-lightning | 12.9% | date=2024-12-23 · command=aider --model openai/yi-lightning · dirname=2024-12-23-01-11-56--yi-test · versions=0.69.2.dev | 23 Dec 2024 | aider.chatT2 | History | |
| #21Qwen2.5 Coder 32B InstructQwen | 8% | date=2024-12-22 · command=aider --model openai/Qwen/Qwen2.5-Coder-32B-Instruct # via hyperbolic · dirname=2024-12-22-13-22-32--polyglot-qwen-diff · versions=0.69.2.dev | 22 Dec 2024 | aider.chatT2 | History | |
| #22GPT-4o-mini (2024-07-18)OpenAI | 3.6% | date=2024-12-21 · command=aider --model gpt-4o-mini-2024-07-18 · dirname=2024-12-21-18-41-18--polyglot-gpt-4o-mini · versions=0.69.2.dev | 21 Dec 2024 | aider.chatT2 | History |
22 results
Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.
The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.
Definition
- Category
- coding
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 3 h agomedium
- Task
- 225 Exercism exercises in 6 languages, edit-format aware
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 3 h agomedium
- Metric
- pass rate (2 attempts) · %
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 3 h agomedium
- Direction
- Higher is better
- Known limitations
- Depends on aider's edit format and prompting; cost column depends on provider pricing.
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 3 h agomedium
- Website
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 3 h agomedium
Each value shows its source, tier and observation time. Missing rows mean no source stated them. How results are recorded →
- Evaluated models
- o1-2024-12-17 (high), DeepSeek Chat V2.5, Gemini 2.0 Pro exp-02-05, GPT-4o-mini (2024-07-18), claude-3-5-sonnet-20241022, GPT-4o (2024-11-20), GPT-4o (2024-08-06), Claude Haiku 3.5 +60(68 total)
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history
Categorycategory1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| coding | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Known limitationsknown_limitations1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| Depends on aider's edit format and prompting; cost column depends on provider pricing. | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Metricmetric1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| pass rate (2 attempts) | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Tasktask1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 225 Exercism exercises in 6 languages, edit-format aware | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Unitunit1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| % | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Websitewebsite1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://aider.chat/docs/leaderboards/ | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| Aider polyglot leaderboard (data file on GitHub) | raw.githubusercontent.com/Aider-AI/aider/main/aider/website/_data/polyglot_leaderboard.yml | leaderboard | T2· Quality secondary | 40 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.