BenchmarkActivecategory · codingfamily · aider-polyglot · variant polyglot
data quality57
Updated 2 h ago · first seen 11 Sept 2026
- Metric
- pass_rate_2 · % ↑
- Current results
- 67
- Models
- 61
- Current leader
- gpt-5 86.7%
Frontier over time · pass_rate_2
- 88%gpt-5 OpenAI Official board23 Aug 2025
- 84.9%o3-pro OpenAI Official board28 Jun 2025
- 83.1%Gemini 2.5 Pro Preview 06-05 Google Official board6 Jun 2025
- 79.1%Gemini 2.5 Pro Preview 06-05 Google Official board6 Jun 2025
- 76.9%Gemini 2.5 Pro Preview 05-06 Google Official board7 May 2025
- 72.9%Gemini 2.5 Pro Preview 03-25 Official board12 Apr 2025
- 64.9%claude-3-7-sonnet-20250219 (32k thinking tokens) Official board24 Feb 2025
- 64%DeepSeek R1 + claude-3-5-sonnet-20241022 Official board23 Jan 2025
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 61 models
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 1 | gpt-5ClosedOpenAI · GPT 5 · best of 2 rows | 86.7% | Official board | edit_formatdiffreasoning_effortmedium | leader | 25 Aug 2025 | aider.chatT2 | History |
| 2 | o3-proClosedOpenAI · OpenAI o-series | 84.9% | Official board | edit_formatdiffreasoning_efforthigh | Partially comparable-1.80 pt | 28 Jun 2025 | aider.chatT2 | History |
| 3 | Gemini 2.5 Pro Preview 06-05Google · Gemini 2.5 · best of 2 rows | 83.1% | Official board | edit_formatdiff-fencedreasoningonthinking_budget32k | Partially comparable-3.60 pt | 6 Jun 2025 | aider.chatT2 | History |
| 4 | o3ClosedOpenAI · OpenAI o-series · best of 2 rows | 81.3% | Official board | edit_formatdiffreasoning_efforthigh | Partially comparable-5.40 pt | 25 Jun 2025 | aider.chatT2 | History |
| 5 | grok-4ClosedSpaceXAI · Grok 4 | 79.6% | Official board | edit_formatdiffreasoning_efforthigh | Partially comparable-7.10 pt | 11 Jul 2025 | aider.chatT2 | History |
| 6 | o3 (high) + gpt-4.1 · GPT 4.1 | 78.2% | Official board | edit_formatarchitect | Partially comparable-8.50 pt | 27 Jun 2025 | aider.chatT2 | History |
| 7 | Gemini 2.5 Pro Preview 05-06Google · Gemini 2.5 | 76.9% | Official board | edit_formatdiff-fenced | Partially comparable-9.80 pt | 7 May 2025 | aider.chatT2 | History |
| 8 | DeepSeek-V3.2-Exp (Reasoner) · DeepSeek | 74.2% | Official board | edit_formatdiff | Partially comparable-12.5 pt | 3 Oct 2025 | aider.chatT2 | History |
| 9 | Gemini 2.5 Pro Preview 03-25 · Gemini 2.5 | 72.9% | Official board | edit_formatdiff-fenced | Partially comparable-13.8 pt | 12 Apr 2025 | aider.chatT2 | History |
| 10 | Claude Opus 4ClosedAnthropic · Claude · best of 2 rows | 72% | Official board | edit_formatdiffreasoningonthinking_budget32k | Partially comparable-14.7 pt | 25 May 2025 | aider.chatT2 | History |
| 10 | o4-miniClosedOpenAI · OpenAI o-series | 72% | Official board | edit_formatdiffreasoning_efforthigh | Partially comparable-14.7 pt | 16 Apr 2025 | aider.chatT2 | History |
| 12 | R1 0528Open weightsDeepSeek | 71.4% | Official board | edit_formatdiff | Partially comparable-15.3 pt | 6 Jun 2025 | aider.chatT2 | History |
| 13 | DeepSeek-V3.2-Exp (Chat) · DeepSeek | 70.2% | Official board | edit_formatdiff | Partially comparable-16.5 pt | 3 Oct 2025 | aider.chatT2 | History |
| 14 | claude-3-7-sonnet-20250219 (32k thinking tokens) · Claude 3.7 | 64.9% | Official board | edit_formatdiff | Partially comparable-21.8 pt | 24 Feb 2025 | aider.chatT2 | History |
| 15 | DeepSeek R1 + claude-3-5-sonnet-20241022 · Claude 3.5 | 64% | Official board | edit_formatarchitect | Partially comparable-22.7 pt | 23 Jan 2025 | aider.chatT2 | History |
| 16 | o1ClosedOpenAI · OpenAI o-series | 61.7% | Official board | edit_formatdiffreasoning_efforthigh | Partially comparable-25.0 pt | 21 Dec 2024 | aider.chatT2 | History |
| 17 | Claude Sonnet 4ClosedAnthropic · Claude · best of 2 rows | 61.3% | Official board | edit_formatdiffreasoningonthinking_budget32k | Partially comparable-25.4 pt | 24 May 2025 | aider.chatT2 | History |
| 18 | Claude 3.7 SonnetClosedAnthropic · Claude | 60.4% | Official board | edit_formatdiffreasoningoff | Partially comparable-26.3 pt | 24 Feb 2025 | aider.chatT2 | History |
| 18 | o3-miniClosedOpenAI · OpenAI o-series · best of 2 rows | 60.4% | Official board | edit_formatdiffreasoning_efforthigh | Partially comparable-26.3 pt | 31 Jan 2025 | aider.chatT2 | History |
| 20 | Qwen3 235B A22B diff, no think, Alibaba API · Qwen3 | 59.6% | Official board | edit_formatdiff | Partially comparable-27.1 pt | 9 May 2025 | aider.chatT2 | History |
| 21 | Kimi K2 0711Open weightsMoonshot AI · Kimi | 59.1% | Official board | edit_formatdiff | Partially comparable-27.6 pt | 17 Jul 2025 | aider.chatT2 | History |
| 22 | R1Open weightsDeepSeek | 56.9% | Official board | edit_formatdiff | Partially comparable-29.8 pt | 20 Jan 2025 | aider.chatT2 | History |
| 23 | DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V3 | 55.1% | Official board | edit_formatdiff | Partially comparable-31.6 pt | 24 Mar 2025 | aider.chatT2 | History |
| 23 | gemini-2.5-flash-preview-05-20 (24k think) · Gemini 2.5 | 55.1% | Official board | edit_formatdiff | Partially comparable-31.6 pt | 25 May 2025 | aider.chatT2 | History |
| 25 | Quasar Alpha | 54.7% | Official board | edit_formatdiff | Partially comparable-32.0 pt | 4 Apr 2025 | aider.chatT2 | History |
| 26 | Grok 3 Beta · Grok 3 | 53.3% | Official board | edit_formatdiff | Partially comparable-33.4 pt | 10 Apr 2025 | aider.chatT2 | History |
| 27 | Optimus Alpha | 52.9% | Official board | edit_formatdiff | Partially comparable-33.8 pt | 10 Apr 2025 | aider.chatT2 | History |
| 28 | gpt-4.1ClosedOpenAI · GPT 4.1 | 52.4% | Official board | edit_formatdiff | Partially comparable-34.3 pt | 14 Apr 2025 | aider.chatT2 | History |
| 29 | claude-3-5-sonnet-20241022Anthropic · Claude 3.5 | 51.6% | Official board | edit_formatdiff | Partially comparable-35.1 pt | 17 Jan 2025 | aider.chatT2 | History |
| 30 | Grok 3 Mini Beta (high) · Grok 3 | 49.3% | Official board | edit_formatwhole | Partially comparable-37.4 pt | 10 Apr 2025 | aider.chatT2 | History |
| 31 | DeepSeek Chat V3 (prev) · DeepSeek | 48.4% | Official board | edit_formatdiff | Partially comparable-38.3 pt | 25 Dec 2024 | aider.chatT2 | History |
| 32 | gemini-2.5-flash-preview-04-17 (default) · Gemini 2.5 | 47.1% | Official board | edit_formatdiff | Partially comparable-39.6 pt | 20 Apr 2025 | aider.chatT2 | History |
| 33 | chatgpt-4o-latest (2025-03-29) · GPT 4 | 45.3% | Official board | edit_formatdiff | Partially comparable-41.4 pt | 29 Mar 2025 | aider.chatT2 | History |
| 34 | GPT-4.5 PreviewClosedOpenAI · GPT 4.5 | 44.9% | Official board | edit_formatdiff | Partially comparable-41.8 pt | 27 Feb 2025 | aider.chatT2 | History |
| 35 | gemini-2.5-flash-preview-05-20 (no think) · Gemini 2.5 | 44% | Official board | edit_formatdiff | Partially comparable-42.7 pt | 26 May 2025 | aider.chatT2 | History |
| 36 | gpt-oss-120bOpen weightsOpenAI · gpt-oss | 41.8% | Official board | edit_formatdiffreasoning_efforthigh | Partially comparable-44.9 pt | 6 Aug 2025 | aider.chatT2 | History |
| 37 | Qwen3 32BOpen weightsQwen · Qwen3 | 40% | Official board | edit_formatdiff | Partially comparable-46.7 pt | 8 May 2025 | aider.chatT2 | History |
| 38 | gemini-exp-1206 · Gemini | 38.2% | Official board | edit_formatwhole | Partially comparable-48.5 pt | 22 Dec 2024 | aider.chatT2 | History |
| 39 | Gemini 2.0 Pro exp-02-05 · Gemini 2.0 | 35.6% | Official board | edit_formatwhole | Partially comparable-51.1 pt | 25 Feb 2025 | aider.chatT2 | History |
| 40 | Grok 3 Mini Beta (low) · Grok 3 | 34.7% | Official board | edit_formatwhole | Partially comparable-52.0 pt | 10 Apr 2025 | aider.chatT2 | History |
| 41 | o1-mini-2024-09-12 · OpenAI o-series | 32.9% | Official board | edit_formatwhole | Partially comparable-53.8 pt | 22 Dec 2024 | aider.chatT2 | History |
| 42 | gpt-4.1-miniClosedOpenAI · GPT 4.1 | 32.4% | Official board | edit_formatdiff | Partially comparable-54.3 pt | 14 Apr 2025 | aider.chatT2 | History |
| 43 | Claude Haiku 3.5ClosedAnthropic · Claude | 28% | Official board | edit_formatdiff | Partially comparable-58.7 pt | 21 Dec 2024 | aider.chatT2 | History |
| 44 | chatgpt-4o-latest (2025-02-15) · GPT 4 | 27.1% | Official board | edit_formatdiff | Partially comparable-59.6 pt | 15 Feb 2025 | aider.chatT2 | History |
| 45 | QwQ-32B + Qwen 2.5 Coder Instruct · Qwen2.5 | 26.2% | Official board | edit_formatarchitect | Partially comparable-60.5 pt | 7 Mar 2025 | aider.chatT2 | History |
| 46 | GPT-4o (2024-08-06)ClosedOpenAI · GPT 4 | 23.1% | Official board | edit_formatdiff | Partially comparable-63.6 pt | 30 Dec 2024 | aider.chatT2 | History |
| 47 | gemini-2.0-flash-expClosedGoogle · Gemini 2.0 | 22.2% | Official board | edit_formatwhole | Partially comparable-64.5 pt | 22 Dec 2024 | aider.chatT2 | History |
| 48 | qwen-max-2025-01-25 · Qwen | 21.8% | Official board | edit_formatdiff | Partially comparable-64.9 pt | 28 Jan 2025 | aider.chatT2 | History |
| 49 | QwQ-32BOpen weightsAlibaba Group · Qwen | 20.9% | Official board | edit_formatdiff | Partially comparable-65.8 pt | 6 Mar 2025 | aider.chatT2 | History |
| 50 | GPT-4o (2024-11-20)OpenAI · GPT 4 | 18.2% | Official board | edit_formatdiff | Partially comparable-68.5 pt | 30 Dec 2024 | aider.chatT2 | History |
| 50 | gemini-2.0-flash-thinking-exp-01-21ClosedGoogle · Gemini 2.0 | 18.2% | Official board | edit_formatdiff | Partially comparable-68.5 pt | 21 Jan 2025 | aider.chatT2 | History |
| 52 | DeepSeek Chat V2.5 · DeepSeek | 17.8% | Official board | edit_formatdiff | Partially comparable-68.9 pt | 21 Dec 2024 | aider.chatT2 | History |
| 53▼51 | Qwen2.5 Coder 32B InstructOpen weightsQwen · Qwen2.5 | 16.4% | Official board | edit_formatwhole | Partially comparable-70.3 pt | 26 Dec 2024 | aider.chatT2 | History |
| 54 | Llama 4 MaverickOpen weightsMeta AI · Llama 4 | 15.6% | Official board | edit_formatwhole | Partially comparable-71.1 pt | 6 Apr 2025 | aider.chatT2 | History |
| 55 | yi-lightning · Yi | 12.9% | Official board | edit_formatwhole | Partially comparable-73.8 pt | 23 Dec 2024 | aider.chatT2 | History |
| 56 | command-a-03-2025-quality · Command | 12% | Official board | edit_formatwhole | Partially comparable-74.7 pt | 14 Mar 2025 | aider.chatT2 | History |
| 57 | Codestral 25.01Mistral AI · Codestral 25.01 | 11.1% | Official board | edit_formatwhole | Partially comparable-75.6 pt | 13 Jan 2025 | aider.chatT2 | History |
| 58 | openhands-lm-32b-v0.1 | 10.2% | Official board | edit_formatwhole | Partially comparable-76.5 pt | 19 Apr 2025 | aider.chatT2 | History |
| 59 | gpt-4.1-nanoClosedOpenAI · GPT 4.1 | 8.90% | Official board | edit_formatwhole | Partially comparable-77.8 pt | 14 Apr 2025 | aider.chatT2 | History |
| 60 | Gemma 3 27BOpen weightsGoogle · Gemma 3 | 4.90% | Official board | edit_formatwhole | Partially comparable-81.8 pt | 15 Mar 2025 | aider.chatT2 | History |
| 61 | GPT-4o-mini (2024-07-18)OpenAI · GPT 4 | 3.60% | Official board | edit_formatwhole | Partially comparable-83.1 pt | 21 Dec 2024 | aider.chatT2 | History |
61 results
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →