Skip to content
AI Atlas
BenchmarkActivecategory · coding

SWE-bench Verified

swebench.com

resolve real GitHub issues (500 human-validated instances)

quality57

Updated 8 h ago · first seen 11 Sept 2026

bench_01M293SPEGA1XMKH0SMTFTCK2R

Metric
resolved · %
Direction
Higher is better
Results
180 · 117 filtered
Leader
Claude Opus 4.5 79.2%

Score history · gpt-4o 8 rows

0%10%20%30%40%Jul 24Oct 24Jan 25Apr 25Jul 25
  • gpt-4o
  • 21.62%date=2025-07-20 · board=Verified · system=mini-SWE-agent · model_tag=gpt-4o-2024112020 Jul 2025
  • 38.8%date=2024-10-28 · board=Verified · system=Agentless-1.5 · model_tag=gpt-4o-2024-05-1328 Oct 2024
  • 27%date=2024-10-16 · board=Verified · system=EPAM AI/Run Developer Agent · model_tag=gpt-4o-2024-08-0616 Oct 2024
  • 24%date=2024-08-20 · board=Verified · system=EPAM AI/Run Developer Agent · model_tag=gpt-4o-2024-08-0620 Aug 2024
  • 23.2%date=2024-07-28 · board=Verified · system=SWE-agent · model_tag=gpt-4o-2024-05-1328 Jul 2024
  • 38.4%date=2024-06-28 · board=Verified · system=AutoCodeRover · model_tag=gpt-4o-2024-05-1328 Jun 2024
  • 26.2%date=2024-06-15 · board=Verified · system=AppMap Navie · model_tag=gpt-4o-2024-05-1315 Jun 2024
  • 32.6%date=2024-06-12 · board=Verified · system=MASAI · model_tag=gpt-4o-2024-08-0612 Jun 2024

Back to the full leaderboard

Leaderboard 117 current results · config contains “true”

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
101#101gpt-4oOpenAI23.2%date=2024-07-28 · board=Verified · system=SWE-agent · model_tag=gpt-4o-2024-05-1328 Jul 2024swebench.comT2 History
102#102MCTS Refine 7B23.2%date=2025-06-27 · board=Verified · system=MCTS-Refine-7B · model_tag=MCTS-Refine-7B27 Jun 2025swebench.comT2 History
103#103GPT-4 (1106)OpenAI22.4%date=2024-04-02 · board=Verified · system=SWE-agent · model_tag=gpt-4-1106-preview2 Apr 2024swebench.comT2 History
104#104gpt-4oOpenAI21.62%date=2025-07-20 · board=Verified · system=mini-SWE-agent · model_tag=gpt-4o-2024112020 Jul 2025swebench.comT2 History
105#105Llama 4 Maverick InstructMeta AI21.04%date=2025-07-20 · board=Verified · system=mini-SWE-agent · model_tag=llama-4-maverick-instruct20 Jul 2025swebench.comT2 History
106#106Lingma SWE-GPT 7b (v0925)OpenAI18.2%date=2024-10-02 · board=Verified · system=Lingma Agent · submission=20241002_lingma-agent_lingma-swe-gpt-7b2 Oct 2024swebench.comT2 History
107#107claude-3-opusAnthropic15.8%date=2024-04-02 · board=Verified · system=SWE-agent · model_tag=claude-3-opus-202402292 Apr 2024swebench.comT2 History
108#108Gemini 2.0 FlashGoogle13.52%date=2025-07-26 · board=Verified · system=mini-SWE-agent · model_tag=gemini-2.0-flash26 Jul 2025swebench.comT2 History
109#109Lingma SWE-GPT 7b (v0918)OpenAI10.2%date=2024-09-18 · board=Verified · system=Lingma Agent · submission=20240918_lingma-agent_lingma-swe-gpt-7b18 Sept 2024swebench.comT2 History
110#110Llama 4 Scout InstructMeta AI9.06%date=2025-07-20 · board=Verified · system=mini-SWE-agent · model_tag=llama-4-scout-instruct20 Jul 2025swebench.comT2 History
111#111Qwen2.5 Coder 32B InstructQwen9%date=2025-08-03 · board=Verified · system=mini-SWE-agent · model_tag=Qwen2.5-Coder-32B-Instruct3 Aug 2025swebench.comT2 History
112#112claude-3-opusAnthropic7%date=2024-04-02 · board=Verified · system=RAG baseline · model_tag=claude-3-opus-202402292 Apr 2024swebench.comT2 History
113#113claude-2Anthropic4.4%date=2023-10-10 · board=Verified · system=RAG baseline · model_tag=claude-210 Oct 2023swebench.comT2 History
114#114GPT-4 (1106)OpenAI2.8%date=2024-04-02 · board=Verified · system=RAG baseline · model_tag=gpt-4-1106-preview2 Apr 2024swebench.comT2 History
115#115SWE-Llama 7BMeta AI1.4%date=2023-10-10 · board=Verified · system=RAG baseline · model_tag=SWE-Llama10 Oct 2023swebench.comT2 History
116#116SWE-Llama 13BMeta AI1.2%date=2023-10-10 · board=Verified · system=RAG baseline · submission=20231010_rag_swellama13b10 Oct 2023swebench.comT2 History
117#117GPT-3.5OpenAI0.4%date=2023-10-10 · board=Verified · system=RAG baseline · submission=20231010_rag_gpt3510 Oct 2023swebench.comT2 History

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.