Skip to content
AI Atlas
BenchmarkActivecategory · coding

SWE-bench Verified

swebench.com

resolve real GitHub issues (500 human-validated instances)

quality57

Updated 8 h ago · first seen 11 Sept 2026

bench_01M293SPEGA1XMKH0SMTFTCK2R

Metric
resolved · %
Direction
Higher is better
Results
180 · 10 filtered
Leader
Claude Opus 4.5 79.2%

Score history · claude-35-sonnet 11 rows

0%20%40%60%80%Jul 24Aug 24Sept 24Oct 24Nov 24Dec 24Jan 25Feb 25
  • claude-35-sonnet
  • 62.8%date=2025-02-28 · board=Verified · system=EPAM AI/Run Developer Agent · model_tag=claude-3-5-sonnet-2024102228 Feb 2025
  • 63.4%date=2025-02-06 · board=Verified · system=AgentScope · model_tag=claude-3-5-sonnet-202410226 Feb 2025
  • 60.2%date=2025-01-10 · board=Verified · system=Learn-by-interact · model_tag=claude-3-5-sonnet-2024102210 Jan 2025
  • 55.4%date=2024-12-12 · board=Verified · system=EPAM AI/Run Developer Agent · model_tag=claude-3-5-sonnet-2024102212 Dec 2024
  • 50.8%date=2024-12-02 · board=Verified · system=Agentless-1.5 · model_tag=claude-3-5-sonnet-202410222 Dec 2024
  • 51.8%date=2024-11-25 · board=Verified · system=Engine Labs · model_tag=claude-3-5-sonnet-2024102225 Nov 2024
  • 54.2%date=2024-11-08 · board=Verified · system=devlo · model_tag=claude-3-5-sonnet-202410228 Nov 2024
  • 39.6%date=2024-10-29 · board=Verified · system=EPAM AI/Run Developer Agent · model_tag=claude-3-5-sonnet-2024102229 Oct 2024
  • 49%date=2024-10-22 · board=Verified · system=Tools · model_tag=claude-3-5-sonnet-2024102222 Oct 2024
  • 40.6%date=2024-10-16 · board=Verified · system=Composio SWEkit · model_tag=claude-3-5-sonnet-2024102216 Oct 2024
  • 33.6%date=2024-06-20 · board=Verified · system=SWE-agent · model_tag=claude-3-5-sonnet-2024102220 Jun 2024

Back to the full leaderboard

Leaderboard 10 current results · config contains “OpenHands”

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
1#1MultipleAnthropic73.8%date=2025-11-03 · board=Verified · system=Salesforce AI Research SAGE · model_tag=claude-sonnet-4.53 Nov 2025swebench.comT2 History
2#2gpt-5OpenAI71.8%date=2025-08-07 · board=Verified · system=OpenHands · model_tag=openai/gpt-5-2025-08-077 Aug 2025swebench.comT2 History
3#3Claude 4 SonnetAnthropic70.4%date=2025-05-24 · board=Verified · system=OpenHands · model_tag=claude-4-sonnet-2025051424 May 2025swebench.comT2 History
4#4qwen3-coder-480b-a35b-instructAlibaba Group69.6%date=2025-08-05 · board=Verified · system=OpenHands · model_tag=https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct5 Aug 2025swebench.comT2 History
5#5Undisclosed65.8%date=2025-04-15 · board=Verified · system=OpenHands · submission=20250415_openhands15 Apr 2025swebench.comT2 History
6#6Kimi K2 0711Moonshot AI65.4%date=2025-07-16 · board=Verified · system=OpenHands · model_tag=moonshot/kimi-k2-0711-preview16 Jul 2025swebench.comT2 History
7#74x Scaled60.8%date=2025-02-03 · board=Verified · system=OpenHands · submission=20250203_openhands_4x_scaled3 Feb 2025swebench.comT2 History
8#8CodeAct v2.1 (claude-3-5-sonnet-20241022)Anthropic53%date=2024-10-29 · board=Verified · system=OpenHands · submission=20241029_OpenHands-CodeAct-2.1-sonnet-2024102229 Oct 2024swebench.comT2 History
9#9Qwen3 Coder 30B A3B InstructQwen51.6%date=2025-08-05 · board=Verified · system=OpenHands · model_tag=https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct5 Aug 2025swebench.comT2 History
10#10Devstral Small 1.0Mistral AI46.8%date=2025-05-20 · board=Verified · system=OpenHands · model_tag=mistralai/Devstral-Small-250520 May 2025swebench.comT2 History

10 results

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.