Updated 9 h ago · first seen 11 Sept 2026
bench_01M293SPEGA1XMKH0SMTFTCK2R
- Metric
- resolved · %
- Direction
- Higher is better
- Results
- 180 · 104 filtered
- Leader
- Claude Opus 4.5 79.2%
Score history · Undisclosed 24 rows
- Undisclosed
- 70.4%date=2025-06-10 · board=Verified · system=Augment Agent v1 · submission=20250610_augment_agent_v110 Jun 2025
- 70.6%date=2025-05-19 · board=Verified · system=TRAE · submission=20250519_trae19 May 2025
- 65.8%date=2025-04-15 · board=Verified · system=OpenHands · submission=20250415_openhands15 Apr 2025
- 56.6%date=2025-04-05 · board=Verified · system=SWE-Rizzo · submission=20250405_swe-rizzo_claude375 Apr 2025
- 65.4%date=2025-04-05 · board=Verified · system=Amazon Q Developer Agent · submission=20250405_amazon-q-developer-agent-20250405-dev5 Apr 2025
- 65.4%date=2025-03-16 · board=Verified · system=Augment Agent v0 · submission=20250316_augment_agent_v016 Mar 2025
- 53.2%date=2025-01-20 · board=Verified · system=Bracket.sh · submission=20250120_Bracket20 Jan 2025
- 41.6%date=2025-01-12 · board=Verified · system=ugaiforge · submission=20250112_ugaiforge12 Jan 2025
- 62.8%date=2025-01-10 · board=Verified · system=Blackbox AI Agent · submission=20250110_blackboxai_agent_v1.110 Jan 2025
- 58.2%date=2024-12-13 · board=Verified · system=devlo · submission=20241213_devlo13 Dec 2024
- 57%date=2024-12-08 · board=Verified · system=Gru · submission=20241208_gru8 Dec 2024
- 55%date=2024-12-02 · board=Verified · system=Amazon Q Developer Agent · submission=20241202_amazon-q-developer-agent-20241202-dev2 Dec 2024
Leaderboard 104 current results · config contains “false”
Select models with +, then open Compare.
| # | Model | Score | Config | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|
| #101gpt-4oOpenAI | 24% | date=2024-08-20 · board=Verified · system=EPAM AI/Run Developer Agent · model_tag=gpt-4o-2024-08-06 | 20 Aug 2024 | swebench.comT2 | History | |
| #102MCTS Refine 7B | 23.2% | date=2025-06-27 · board=Verified · system=MCTS-Refine-7B · model_tag=MCTS-Refine-7B | 27 Jun 2025 | swebench.comT2 | History | |
| #103Lingma SWE-GPT 7b (v0925)OpenAI | 18.2% | date=2024-10-02 · board=Verified · system=Lingma Agent · submission=20241002_lingma-agent_lingma-swe-gpt-7b | 2 Oct 2024 | swebench.comT2 | History | |
| #104Lingma SWE-GPT 7b (v0918)OpenAI | 10.2% | date=2024-09-18 · board=Verified · system=Lingma Agent · submission=20240918_lingma-agent_lingma-swe-gpt-7b | 18 Sept 2024 | swebench.comT2 | History |
Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.
The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.
Definition
- Category
- coding
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 9 h agomedium
- Task
- resolve real GitHub issues (500 human-validated instances)
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 9 h agomedium
- Metric
- resolved · %
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 9 h agomedium
- Direction
- Higher is better
- Paper
- https://arxiv.org/abs/2310.06770
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 9 h agomedium
- Known limitations
- Scaffold/agent dependent; results are not comparable across harnesses.
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 9 h agomedium
- Website
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 9 h agomedium
Each value shows its source, tier and observation time. Missing rows mean no source stated them. How results are recorded →
- Evaluated models
- Claude Opus 4.6, doubao-seed-code, gemini-3-flash, MiniMax M2.5, Claude Opus 4.5, Gemini 3 Pro Preview, GLM 5, GPT-5.2-Codex +74(82 total)
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history
Categorycategory1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| coding | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Known limitationsknown_limitations1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| Scaffold/agent dependent; results are not comparable across harnesses. | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Metricmetric1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| resolved | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Paperpaper1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/abs/2310.06770 | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Tasktask1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| resolve real GitHub issues (500 human-validated instances) | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Unitunit1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| % | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Websitewebsite1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://www.swebench.com | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| SWE-bench leaderboards | swebench.com/ | leaderboard | T2· Quality secondary | 7 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.