Skip to content
AI Atlas
BenchmarkActivecategory · coding

SWE-bench Verified

swebench.com

resolve real GitHub issues (500 human-validated instances)

quality57

Updated 8 h ago · first seen 11 Sept 2026

bench_01M293SPEGA1XMKH0SMTFTCK2R

Metric
resolved · %
Direction
Higher is better
Results
180 · 11 filtered
Leader
Claude Opus 4.5 79.2%

Score history · Undisclosed 24 rows

0%20%40%60%80%Jul 24Oct 24Jan 25Apr 25
  • Undisclosed
  • 70.4%date=2025-06-10 · board=Verified · system=Augment Agent v1 · submission=20250610_augment_agent_v110 Jun 2025
  • 70.6%date=2025-05-19 · board=Verified · system=TRAE · submission=20250519_trae19 May 2025
  • 65.8%date=2025-04-15 · board=Verified · system=OpenHands · submission=20250415_openhands15 Apr 2025
  • 56.6%date=2025-04-05 · board=Verified · system=SWE-Rizzo · submission=20250405_swe-rizzo_claude375 Apr 2025
  • 65.4%date=2025-04-05 · board=Verified · system=Amazon Q Developer Agent · submission=20250405_amazon-q-developer-agent-20250405-dev5 Apr 2025
  • 65.4%date=2025-03-16 · board=Verified · system=Augment Agent v0 · submission=20250316_augment_agent_v016 Mar 2025
  • 53.2%date=2025-01-20 · board=Verified · system=Bracket.sh · submission=20250120_Bracket20 Jan 2025
  • 41.6%date=2025-01-12 · board=Verified · system=ugaiforge · submission=20250112_ugaiforge12 Jan 2025
  • 62.8%date=2025-01-10 · board=Verified · system=Blackbox AI Agent · submission=20250110_blackboxai_agent_v1.110 Jan 2025
  • 58.2%date=2024-12-13 · board=Verified · system=devlo · submission=20241213_devlo13 Dec 2024
  • 57%date=2024-12-08 · board=Verified · system=Gru · submission=20241208_gru8 Dec 2024
  • 55%date=2024-12-02 · board=Verified · system=Amazon Q Developer Agent · submission=20241202_amazon-q-developer-agent-20241202-dev2 Dec 2024

Back to the full leaderboard

Leaderboard 11 current results · config contains “high”

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
1#1Claude Opus 4.5Anthropic76.8%date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=claude-4-5-opus17 Feb 2026swebench.comT2 History
2#2MiniMax M2.5MiniMax75.8%date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=minimax-m2.517 Feb 2026swebench.comT2 History
3#3gemini-3-flashGoogle75.8%date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=gemini-3-flash-preview17 Feb 2026swebench.comT2 History
4#4GLM 5Z.ai (Zhipu AI)72.8%date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=glm-517 Feb 2026swebench.comT2 History
5#5gpt-5.2OpenAI72.8%date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=gpt-5-217 Feb 2026swebench.comT2 History
6#6gpt-5.2OpenAI71.8%date=2025-12-11 · board=Verified · system=mini-SWE-agent · model_tag=gpt-5.2-2025-12-1111 Dec 2025swebench.comT2 History
7#7Claude 4.5 SonnetAnthropic71.4%date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=claude-sonnet-4-5-2025092917 Feb 2026swebench.comT2 History
8#8Kimi K2.5Moonshot AI70.8%date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=kimi-k2.517 Feb 2026swebench.comT2 History
9#9DeepSeek V3.2DeepSeek70%date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=deepseek-v3.217 Feb 2026swebench.comT2 History
10#10gemini-3-proGoogle69.6%date=2026-02-26 · board=Verified · system=mini-SWE-agent · model_tag=gemini-3-pro-preview26 Feb 2026swebench.comT2 History
11#11Claude 4.5 HaikuAnthropic66.6%date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=claude-haiku-4-5-2025100117 Feb 2026swebench.comT2 History

11 results

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.