Skip to content
AI Atlas
BenchmarkActive

SWE-bench Verified

swebench.com

resolve real GitHub issues (500 human-validated instances)

quality57

Updated 10 h ago · first seen 11 Sept 2026

Metric
· %
Current results
0
Models
0
Current leader
Claude Opus 4.5 76.8%

No results recorded for this benchmark yet — its sources are being connected. The definition, aliases and variants are kept so links resolve; nothing is fabricated.

Score history · Gemini 2.5 Flash 1 row

Not enough history to chart — a single observation (28.73% on 26 Jul 2025). Rows under different configurations count separately; the list below shows each one.

  • 28.73%date=2025-07-26 · board=Verified · system=mini-SWE-agent · model_tag=gemini-2.5-flash26 Jul 2025

Back to the leaderboard

Leaderboard 42 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1Claude Opus 4.5ClosedAnthropic · Claude76.8%Official boardreasoning_efforthighleader17 Feb 2026swebench.comT2History
2MiniMax M2.5Open weightsMiniMax · MiniMax75.8%Official boardreasoning_efforthighPartially comparable-1.00 pt17 Feb 2026swebench.comT2History
2gemini-3-flashClosedGoogle · Gemini 375.8%Official boardreasoning_efforthighPartially comparable-1.00 pt17 Feb 2026swebench.comT2History
4Claude Opus 4.6ClosedAnthropic · Claude75.6%Official boardgroup defaultsPartially comparable-1.20 pt17 Feb 2026swebench.comT2History
5Gemini 3 Pro PreviewGoogle · Gemini74.2%Official boardgroup defaultsPartially comparable-2.60 pt18 Nov 2025swebench.comT2History
6GLM 5Open weightsZ.ai (Zhipu AI) · GLM572.8%Official boardreasoning_efforthighPartially comparable-4.00 pt17 Feb 2026swebench.comT2History
6GPT-5.2-CodexClosedOpenAI · GPT 5.272.8%Official boardgroup defaultsPartially comparable-4.00 pt19 Feb 2026swebench.comT2History
64gpt-5.2ClosedOpenAI · GPT 5.272.8%Official boardreasoning_efforthighPartially comparable-4.00 pt17 Feb 2026swebench.comT2History
96Claude 4.5 SonnetAnthropic · Claude 4.571.4%Official boardreasoning_efforthighPartially comparable-5.40 pt17 Feb 2026swebench.comT2History
10Kimi K2.5Open weightsMoonshot AI · Kimi70.8%Official boardreasoning_efforthighPartially comparable-6.00 pt17 Feb 2026swebench.comT2History
11DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V370%Official boardreasoning_efforthighPartially comparable-6.80 pt17 Feb 2026swebench.comT2History
12gemini-3-proClosedGoogle · Gemini 369.6%Official boardreasoning_efforthighPartially comparable-7.20 pt26 Feb 2026swebench.comT2History
13claude-4-opusClosedAnthropic · Claude 467.6%Official boardgroup defaultsPartially comparable-9.20 pt2 Aug 2025swebench.comT2History
14Claude 4.5 HaikuAnthropic · Claude 4.566.6%Official boardreasoning_efforthighPartially comparable-10.2 pt17 Feb 2026swebench.comT2History
15GPT-5.1-CodexClosedOpenAI · GPT 5.166%Official boardreasoning_effortmediumPartially comparable-10.8 pt24 Nov 2025swebench.comT2History
15gpt-5.1ClosedOpenAI · GPT 5.166%Official boardreasoning_effortmediumPartially comparable-10.8 pt20 Nov 2025swebench.comT2History
17gpt-5ClosedOpenAI · GPT 565%Official boardreasoning_effortmediumPartially comparable-11.8 pt7 Aug 2025swebench.comT2History
18Claude 4 SonnetAnthropic · Claude 464.9%Official boardgroup defaultsPartially comparable-11.9 pt26 Jul 2025swebench.comT2History
19Kimi K2 ThinkingOpen weightsMoonshot AI · Kimi63.4%Official boardgroup defaultsPartially comparable-13.4 pt10 Dec 2025swebench.comT2History
20MiniMax M2Open weightsMiniMax · MiniMax61%Official boardgroup defaultsPartially comparable-15.8 pt24 Nov 2025swebench.comT2History
21DeepSeek V3.2 ReasonerDeepSeek · DeepSeek60%Official boardgroup defaultsPartially comparable-16.8 pt1 Dec 2025swebench.comT2History
22o3ClosedOpenAI · OpenAI o-series58.4%Official boardgroup defaultsPartially comparable-18.4 pt26 Jul 2025swebench.comT2History
23Devstral Small (2512)Mistral AI · Devstral56.4%Official boardgroup defaultsPartially comparable-20.4 pt9 Dec 2025swebench.comT2History
2420gpt-5-miniClosedOpenAI · GPT 556.2%Official boardgroup defaultsPartially comparable-20.6 pt17 Feb 2026swebench.comT2History
25GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.655.4%Official boardgroup defaultsPartially comparable-21.4 pt1 Dec 2025swebench.comT2History
25qwen3-coder-480b-a35b-instructOpen weightsAlibaba Group · Qwen3-Coder55.4%Official boardgroup defaultsPartially comparable-21.4 pt2 Aug 2025swebench.comT2History
27GLM 4.5Open weightsZ.ai (Zhipu AI) · GLM4.554.2%Official boardgroup defaultsPartially comparable-22.6 pt22 Aug 2025swebench.comT2History
28Devstral 2Open weightsMistral AI · Devstral 253.8%Official boardgroup defaultsPartially comparable-23.0 pt9 Dec 2025swebench.comT2History
29Gemini 2.5 ProClosedGoogle · Gemini53.6%Official boardgroup defaultsPartially comparable-23.2 pt26 Jul 2025swebench.comT2History
30Claude 3.7 SonnetClosedAnthropic · Claude52.8%Official boardgroup defaultsPartially comparable-24.0 pt20 Jul 2025swebench.comT2History
31o4-miniClosedOpenAI · OpenAI o-series45%Official boardgroup defaultsPartially comparable-31.8 pt26 Jul 2025swebench.comT2History
32Kimi K2 0711Open weightsMoonshot AI · Kimi43.8%Official boardgroup defaultsPartially comparable-33.0 pt7 Aug 2025swebench.comT2History
33gpt-4.1ClosedOpenAI · GPT 4.139.6%Official boardgroup defaultsPartially comparable-37.2 pt26 Jul 2025swebench.comT2History
34gpt-5-nanoClosedOpenAI · GPT 534.8%Official boardreasoning_effortmediumPartially comparable-42.0 pt7 Aug 2025swebench.comT2History
35Gemini 2.5 FlashClosedGoogle · Gemini28.7%Official boardgroup defaultsPartially comparable-48.1 pt26 Jul 2025swebench.comT2History
36gpt-oss-120bOpen weightsOpenAI · gpt-oss26%Official boardgroup defaultsPartially comparable-50.8 pt7 Aug 2025swebench.comT2History
37gpt-4.1-miniClosedOpenAI · GPT 4.123.9%Official boardgroup defaultsPartially comparable-52.9 pt20 Jul 2025swebench.comT2History
38gpt-4oClosedOpenAI · GPT 421.6%Official boardgroup defaultsPartially comparable-55.2 pt20 Jul 2025swebench.comT2History
39Llama 4 Maverick InstructMeta AI · Llama 421.0%Official boardgroup defaultsPartially comparable-55.8 pt20 Jul 2025swebench.comT2History
40Gemini 2.0 FlashClosedGoogle · Gemini13.5%Official boardgroup defaultsPartially comparable-63.3 pt26 Jul 2025swebench.comT2History
41Llama 4 Scout InstructMeta AI · Llama 49.06%Official boardgroup defaultsPartially comparable-67.7 pt20 Jul 2025swebench.comT2History
42Qwen2.5 Coder 32B InstructOpen weightsQwen · Qwen2.59%Official boardgroup defaultsPartially comparable-67.8 pt3 Aug 2025swebench.comT2History

42 results

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →