Skip to content
AI Atlas
BenchmarkActivecategory · coding

SWE-bench Verified

swebench.com

resolve real GitHub issues (500 human-validated instances)

quality57

Updated 8 h ago · first seen 11 Sept 2026

bench_01M293SPEGA1XMKH0SMTFTCK2R

Metric
resolved · %
Direction
Higher is better
Results
180 · 104 filtered
Leader
Claude Opus 4.5 79.2%

Score history · Multiple 22 rows

0%20%40%60%80%Jan 25Apr 25Jul 25Oct 25
  • Multiple
  • 73.8%date=2025-11-03 · board=Verified · system=Salesforce AI Research SAGE · model_tag=claude-sonnet-4.53 Nov 2025
  • 73%date=2025-10-21 · board=Verified · system=Salesforce AI Research SAGE · model_tag=claude-sonnet-4.521 Oct 2025
  • 74.6%date=2025-09-15 · board=Verified · system=JoyCode · model_tag=claude-4-sonnet15 Sept 2025
  • 76.8%date=2025-09-02 · board=Verified · system=Atlassian Rovo Dev · model_tag=claude-sonnet-4-202505142 Sept 2025
  • 75.6%date=2025-09-01 · board=Verified · system=Warp · model_tag=gpt-51 Sept 2025
  • 76.4%date=2025-08-19 · board=Verified · system=ACoder · model_tag=claude-4-sonnet19 Aug 2025
  • 71.2%date=2025-07-15 · board=Verified · system=Qodo Command · model_tag=claude-sonnet-4-2025051415 Jul 2025
  • 71%date=2025-06-23 · board=Verified · system=Warp · model_tag=claude-sonnet-4-2025051423 Jun 2025
  • 75.2%date=2025-06-12 · board=Verified · system=TRAE · model_tag=claude-4-sonnet-2025052212 Jun 2025
  • 74.4%date=2025-06-03 · board=Verified · system=Refact.ai Agent · model_tag=claude-4-sonnet3 Jun 2025
  • 70.2%date=2025-05-19 · board=Verified · system=devlo · model_tag=claude-3-7-sonnet-2025021919 May 2025
  • 68.2%date=2025-05-16 · board=Verified · system=Nemotron-CORTEXA · model_tag=NV-EmbedCode16 May 2025

Back to the full leaderboard

Leaderboard 104 current results · config contains “false”

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
101#101gpt-4oOpenAI24%date=2024-08-20 · board=Verified · system=EPAM AI/Run Developer Agent · model_tag=gpt-4o-2024-08-0620 Aug 2024swebench.comT2 History
102#102MCTS Refine 7B23.2%date=2025-06-27 · board=Verified · system=MCTS-Refine-7B · model_tag=MCTS-Refine-7B27 Jun 2025swebench.comT2 History
103#103Lingma SWE-GPT 7b (v0925)OpenAI18.2%date=2024-10-02 · board=Verified · system=Lingma Agent · submission=20241002_lingma-agent_lingma-swe-gpt-7b2 Oct 2024swebench.comT2 History
104#104Lingma SWE-GPT 7b (v0918)OpenAI10.2%date=2024-09-18 · board=Verified · system=Lingma Agent · submission=20240918_lingma-agent_lingma-swe-gpt-7b18 Sept 2024swebench.comT2 History

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.