Skip to content
AI Atlas
BenchmarkActivecategory · codingfamily · swe-bench · variant Verified

SWE-bench Verified

swebench.com

resolve real GitHub issues (500 human-validated instances)

data quality57

Updated 2 h ago · first seen 11 Sept 2026

Metric
resolved · %
Current results
158
Models
82
Current leader
Claude Opus 4.5 76.8%

Score history · Claude Opus 4.6 1 row

Not enough history to chart — a single observation (75.6% on 17 Feb 2026). Rows under different configurations count separately; the list below shows each one.

  • 75.6%date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=claude-opus-4-617 Feb 2026

Back to the leaderboard

Frontier over time · resolved · board=Verified · system=mini-SWE-agent

Frontier of SWE-bench Verified0%20%40%60%80%Aug 25Sept 25Oct 25Nov 25Dec 25Jan 26Feb 26
  1. 76.8%Claude Opus 4.5 Anthropic Official board17 Feb 2026
  2. 74.4%Claude Opus 4.5 Anthropic Official board24 Nov 2025
  3. 74.2%Gemini 3 Pro Preview Google Official board18 Nov 2025
  4. 70.6%Claude 4.5 Sonnet Anthropic Official board29 Sept 2025
  5. 67.6%claude-4-opus Anthropic Official board2 Aug 2025
  6. 64.9%Claude 4 Sonnet Anthropic Official board26 Jul 2025
  7. 52.8%Claude 3.7 Sonnet Anthropic Official board20 Jul 2025

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 42 models · trust official-benchmark

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1Claude Opus 4.5ClosedAnthropic · Claude76.8%Official boardreasoning_efforthighleader17 Feb 2026swebench.comT2History
2MiniMax M2.5Open weightsMiniMax · MiniMax75.8%Official boardreasoning_efforthighPartially comparable-1.00 pt17 Feb 2026swebench.comT2History
2gemini-3-flashClosedGoogle · Gemini 375.8%Official boardreasoning_efforthighPartially comparable-1.00 pt17 Feb 2026swebench.comT2History
4Claude Opus 4.6ClosedAnthropic · Claude75.6%Official boardgroup defaultsPartially comparable-1.20 pt17 Feb 2026swebench.comT2History
5Gemini 3 Pro PreviewGoogle · Gemini 374.2%Official boardgroup defaultsPartially comparable-2.60 pt18 Nov 2025swebench.comT2History
6GLM 5Open weightsZ.ai (Zhipu AI) · GLM572.8%Official boardreasoning_efforthighPartially comparable-4.00 pt17 Feb 2026swebench.comT2History
6GPT-5.2-CodexClosedOpenAI · GPT 5.272.8%Official boardgroup defaultsPartially comparable-4.00 pt19 Feb 2026swebench.comT2History
64gpt-5.2ClosedOpenAI · GPT 5.272.8%Official boardreasoning_efforthighPartially comparable-4.00 pt17 Feb 2026swebench.comT2History
96Claude 4.5 SonnetAnthropic · Claude 4.571.4%Official boardreasoning_efforthighPartially comparable-5.40 pt17 Feb 2026swebench.comT2History
10Kimi K2.5Open weightsMoonshot AI · Kimi70.8%Official boardreasoning_efforthighPartially comparable-6.00 pt17 Feb 2026swebench.comT2History
11DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V370%Official boardreasoning_efforthighPartially comparable-6.80 pt17 Feb 2026swebench.comT2History
12gemini-3-proClosedGoogle · Gemini 369.6%Official boardreasoning_efforthighPartially comparable-7.20 pt26 Feb 2026swebench.comT2History
13claude-4-opusClosedAnthropic · Claude 467.6%Official boardgroup defaultsPartially comparable-9.20 pt2 Aug 2025swebench.comT2History
14Claude 4.5 HaikuAnthropic · Claude 4.566.6%Official boardreasoning_efforthighPartially comparable-10.2 pt17 Feb 2026swebench.comT2History
15GPT-5.1-CodexClosedOpenAI · GPT 5.166%Official boardreasoning_effortmediumPartially comparable-10.8 pt24 Nov 2025swebench.comT2History
15gpt-5.1ClosedOpenAI · GPT 5.166%Official boardreasoning_effortmediumPartially comparable-10.8 pt20 Nov 2025swebench.comT2History
17gpt-5ClosedOpenAI · GPT 565%Official boardreasoning_effortmediumPartially comparable-11.8 pt7 Aug 2025swebench.comT2History
18Claude 4 SonnetAnthropic · Claude 464.9%Official boardgroup defaultsPartially comparable-11.9 pt26 Jul 2025swebench.comT2History
19Kimi K2 ThinkingOpen weightsMoonshot AI · Kimi63.4%Official boardgroup defaultsPartially comparable-13.4 pt10 Dec 2025swebench.comT2History
20MiniMax M2Open weightsMiniMax · MiniMax61%Official boardgroup defaultsPartially comparable-15.8 pt24 Nov 2025swebench.comT2History
21DeepSeek V3.2 ReasonerDeepSeek · DeepSeek60%Official boardgroup defaultsPartially comparable-16.8 pt1 Dec 2025swebench.comT2History
22o3ClosedOpenAI · OpenAI o-series58.4%Official boardgroup defaultsPartially comparable-18.4 pt26 Jul 2025swebench.comT2History
23Devstral Small (2512)Mistral AI · Devstral56.4%Official boardgroup defaultsPartially comparable-20.4 pt9 Dec 2025swebench.comT2History
2420gpt-5-miniClosedOpenAI · GPT 556.2%Official boardgroup defaultsPartially comparable-20.6 pt17 Feb 2026swebench.comT2History
25GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.655.4%Official boardgroup defaultsPartially comparable-21.4 pt1 Dec 2025swebench.comT2History
25qwen3-coder-480b-a35b-instructOpen weightsAlibaba Group · Qwen3-Coder55.4%Official boardgroup defaultsPartially comparable-21.4 pt2 Aug 2025swebench.comT2History
27GLM 4.5Open weightsZ.ai (Zhipu AI) · GLM4.554.2%Official boardgroup defaultsPartially comparable-22.6 pt22 Aug 2025swebench.comT2History
28Devstral 2Open weightsMistral AI · Devstral 253.8%Official boardgroup defaultsPartially comparable-23.0 pt9 Dec 2025swebench.comT2History
29Gemini 2.5 ProClosedGoogle · Gemini 2.553.6%Official boardgroup defaultsPartially comparable-23.2 pt26 Jul 2025swebench.comT2History
30Claude 3.7 SonnetClosedAnthropic · Claude52.8%Official boardgroup defaultsPartially comparable-24.0 pt20 Jul 2025swebench.comT2History
31o4-miniClosedOpenAI · OpenAI o-series45%Official boardgroup defaultsPartially comparable-31.8 pt26 Jul 2025swebench.comT2History
32Kimi K2 0711Open weightsMoonshot AI · Kimi43.8%Official boardgroup defaultsPartially comparable-33.0 pt7 Aug 2025swebench.comT2History
33gpt-4.1ClosedOpenAI · GPT 4.139.6%Official boardgroup defaultsPartially comparable-37.2 pt26 Jul 2025swebench.comT2History
34gpt-5-nanoClosedOpenAI · GPT 534.8%Official boardreasoning_effortmediumPartially comparable-42.0 pt7 Aug 2025swebench.comT2History
35Gemini 2.5 FlashClosedGoogle · Gemini 2.528.7%Official boardgroup defaultsPartially comparable-48.1 pt26 Jul 2025swebench.comT2History
36gpt-oss-120bOpen weightsOpenAI · gpt-oss26%Official boardgroup defaultsPartially comparable-50.8 pt7 Aug 2025swebench.comT2History
37gpt-4.1-miniClosedOpenAI · GPT 4.123.9%Official boardgroup defaultsPartially comparable-52.9 pt20 Jul 2025swebench.comT2History
38gpt-4oClosedOpenAI · GPT 421.6%Official boardgroup defaultsPartially comparable-55.2 pt20 Jul 2025swebench.comT2History
39Llama 4 Maverick InstructMeta AI · Llama 421.0%Official boardgroup defaultsPartially comparable-55.8 pt20 Jul 2025swebench.comT2History
40Gemini 2.0 FlashClosedGoogle · Gemini 2.013.5%Official boardgroup defaultsPartially comparable-63.3 pt26 Jul 2025swebench.comT2History
41Llama 4 Scout InstructMeta AI · Llama 49.06%Official boardgroup defaultsPartially comparable-67.7 pt20 Jul 2025swebench.comT2History
42Qwen2.5 Coder 32B InstructOpen weightsQwen · Qwen2.59%Official boardgroup defaultsPartially comparable-67.8 pt3 Aug 2025swebench.comT2History

42 results

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →