Skip to content
AI Atlas
BenchmarkActivecategory · coding

Aider polyglot

aider.chat/docs/leaderboards

225 Exercism exercises in 6 languages, edit-format aware

quality57

Updated 5 h ago · first seen 11 Sept 2026

bench_01M293SPEMVE6W55699099NES3

Metric
pass rate (2 attempts) · %
Direction
Higher is better
Results
138 · 106 filtered
Leader
QwQ-32B + Qwen 2.5 Coder Instruct 100%

Score history · Gemini 2.5 Pro Preview 05-06 2 rows

Not enough history to chart — 2 observations, all dated 7 May 2025. Rows under different configurations count separately; the list below shows each one.

  • 97.3%date=2025-05-07 · command=aider --model gemini/gemini-2.5-pro-preview-05-06 · dirname=2025-05-07-19-32-40--gemini0506-diff-fenced-completion_cost · versions=0.82.4.dev7 May 2025
  • 76.9%date=2025-05-07 · command=aider --model gemini/gemini-2.5-pro-preview-05-06 · dirname=2025-05-07-19-32-40--gemini0506-diff-fenced-completion_cost · versions=0.82.4.dev7 May 2025

Back to the full leaderboard

Leaderboard 106 current results · config contains “diff”

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
101#101QwQ-32BAlibaba Group20.9%date=2025-03-06 · command=aider --model fireworks_ai/accounts/fireworks/models/qwq-32b · dirname=2025-03-06-17-40-24--qwq32b-diff-temp-topp-ex-sys-remind-user-for-real · versions=0.75.3.dev6 Mar 2025aider.chatT2 History
102#102GPT-4o (2024-11-20)OpenAI18.2%date=2024-12-30 · command=aider --model gpt-4o-2024-11-20 · dirname=2024-12-30-20-57-12--gpt-4o-2024-11-20-ex-as-sys · versions=0.70.1.dev30 Dec 2024aider.chatT2 History
103#103gemini-2.0-flash-thinking-exp-01-21Google18.2%date=2025-01-21 · command=aider --model gemini/gemini-2.0-flash-thinking-exp-01-21 · dirname=2025-01-21-22-51-49--gemini-2.0-flash-thinking-exp-01-21-polyglot-diff · versions=0.72.2.dev21 Jan 2025aider.chatT2 History
104#104DeepSeek Chat V2.517.8%date=2024-12-21 · command=aider --model deepseek/deepseek-chat · dirname=2024-12-21-20-56-21--polyglot-deepseek-diff · versions=0.69.2.dev21 Dec 2024aider.chatT2 History
105#105gpt-4.1-nanoOpenAI8.9%date=2025-04-14 · command=aider --model gpt-4.1-nano · dirname=2025-04-14-22-46-01--gpt41nano-diff · versions=0.81.4.dev14 Apr 2025aider.chatT2 History
106#106Qwen2.5 Coder 32B InstructQwen8%date=2024-12-22 · command=aider --model openai/Qwen/Qwen2.5-Coder-32B-Instruct # via hyperbolic · dirname=2024-12-22-13-22-32--polyglot-qwen-diff · versions=0.69.2.dev22 Dec 2024aider.chatT2 History

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.