Skip to content
AI Atlas
BenchmarkActivecategory · coding

SWE-bench Verified

swebench.com

resolve real GitHub issues (500 human-validated instances)

quality57

Updated 8 h ago · first seen 11 Sept 2026

bench_01M293SPEGA1XMKH0SMTFTCK2R

Metric
resolved · %
Direction
Higher is better
Results
180 · 3 filtered
Leader
Claude Opus 4.5 79.2%

Score history · Qwen3 Coder 30B A3B Instruct 3 rows

0%20%40%60%80%Aug 25Aug 25Aug 25Aug 25
  • Qwen3 Coder 30B A3B Instruct
  • 52.2%date=2025-09-01 · board=Verified · system=EntroPO + R2E · model_tag=Qwen3-Coder-30B-A3B-Instruct1 Sept 2025
  • 60.4%date=2025-09-01 · board=Verified · system=EntroPO + R2E · model_tag=Qwen3-Coder-30B-A3B-Instruct1 Sept 2025
  • 51.6%date=2025-08-05 · board=Verified · system=OpenHands · model_tag=https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct5 Aug 2025

Back to the full leaderboard

Leaderboard 3 current results · config contains “devlo”

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
1#1MultipleAnthropic70.2%date=2025-05-19 · board=Verified · system=devlo · model_tag=claude-3-7-sonnet-2025021919 May 2025swebench.comT2 History
2#2Undisclosed58.2%date=2024-12-13 · board=Verified · system=devlo · submission=20241213_devlo13 Dec 2024swebench.comT2 History
3#3claude-35-sonnetAnthropic54.2%date=2024-11-08 · board=Verified · system=devlo · model_tag=claude-3-5-sonnet-202410228 Nov 2024swebench.comT2 History

3 results

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.