Updated 8 h ago · first seen 11 Sept 2026
bench_01M293SPEGA1XMKH0SMTFTCK2R
- Metric
- resolved · %
- Direction
- Higher is better
- Results
- 180 · 61 filtered
- Leader
- Claude Opus 4.5 79.2%
Score history · CodeAct v2.1 (claude-3-5-sonnet-20241022) 1 row
Not enough history to chart — a single observation (53% on 29 Oct 2024). Rows under different configurations count separately; the list below shows each one.
- 53%date=2024-10-29 · board=Verified · system=OpenHands · submission=20241029_OpenHands-CodeAct-2.1-sonnet-2024102229 Oct 2024
Leaderboard 61 current results · config contains “SWE-agent”
Select models with +, then open Compare.
| # | Model | Score | Config | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|
| #1Claude Opus 4.5Anthropic | 79.2% | date=2025-12-15 · board=Verified · system=live-SWE-agent · model_tag=claude-opus-4-5-20251101 | 15 Dec 2025 | swebench.comT2 | History | |
| #2Gemini 3 Pro PreviewGoogle | 77.4% | date=2025-11-20 · board=Verified · system=live-SWE-agent · model_tag=gemini-3-pro-preview | 20 Nov 2025 | swebench.comT2 | History | |
| #3Claude Opus 4.5Anthropic | 76.8% | date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=claude-4-5-opus | 17 Feb 2026 | swebench.comT2 | History | |
| #4MiniMax M2.5MiniMax | 75.8% | date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=minimax-m2.5 | 17 Feb 2026 | swebench.comT2 | History | |
| #5gemini-3-flashGoogle | 75.8% | date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=gemini-3-flash-preview | 17 Feb 2026 | swebench.comT2 | History | |
| #6Claude Opus 4.6Anthropic | 75.6% | date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=claude-opus-4-6 | 17 Feb 2026 | swebench.comT2 | History | |
| #7Claude Opus 4.5Anthropic | 74.4% | date=2025-11-24 · board=Verified · system=mini-SWE-agent · model_tag=claude-opus-4-5-20251101 | 24 Nov 2025 | swebench.comT2 | History | |
| #8Gemini 3 Pro PreviewGoogle | 74.2% | date=2025-11-18 · board=Verified · system=mini-SWE-agent · model_tag=gemini-3-pro-preview | 18 Nov 2025 | swebench.comT2 | History | |
| #9gpt-5.2OpenAI | 72.8% | date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=gpt-5-2 | 17 Feb 2026 | swebench.comT2 | History | |
| #10GPT-5.2-CodexOpenAI | 72.8% | date=2026-02-19 · board=Verified · system=mini-SWE-agent · model_tag=gpt-5-2-codex | 19 Feb 2026 | swebench.comT2 | History | |
| #11GLM 5Z.ai (Zhipu AI) | 72.8% | date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=glm-5 | 17 Feb 2026 | swebench.comT2 | History | |
| #12gpt-5.2OpenAI | 71.8% | date=2025-12-11 · board=Verified · system=mini-SWE-agent · model_tag=gpt-5.2-2025-12-11 | 11 Dec 2025 | swebench.comT2 | History | |
| #13Claude 4.5 SonnetAnthropic | 71.4% | date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=claude-sonnet-4-5-20250929 | 17 Feb 2026 | swebench.comT2 | History | |
| #14Kimi K2.5Moonshot AI | 70.8% | date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=kimi-k2.5 | 17 Feb 2026 | swebench.comT2 | History | |
| #15Claude 4.5 SonnetAnthropic | 70.6% | date=2025-09-29 · board=Verified · system=mini-SWE-agent · model_tag=claude-sonnet-4-5-20250929 | 29 Sept 2025 | swebench.comT2 | History | |
| #16DeepSeek V3.2DeepSeek | 70% | date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=deepseek-v3.2 | 17 Feb 2026 | swebench.comT2 | History | |
| #17gemini-3-proGoogle | 69.6% | date=2026-02-26 · board=Verified · system=mini-SWE-agent · model_tag=gemini-3-pro-preview | 26 Feb 2026 | swebench.comT2 | History | |
| #18gpt-5.2OpenAI | 69% | date=2025-12-11 · board=Verified · system=mini-SWE-agent · model_tag=gpt-5.2-2025-12-11 | 11 Dec 2025 | swebench.comT2 | History | |
| #19claude-4-opusAnthropic | 67.6% | date=2025-08-02 · board=Verified · system=mini-SWE-agent · model_tag=claude-4-opus-20250514 | 2 Aug 2025 | swebench.comT2 | History | |
| #20Claude 4.5 HaikuAnthropic | 66.6% | date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=claude-haiku-4-5-20251001 | 17 Feb 2026 | swebench.comT2 | History | |
| #21Claude 4 SonnetAnthropic | 66.6% | date=2025-05-22 · board=Verified · system=SWE-agent · model_tag=claude-4-sonnet-20250514 | 22 May 2025 | swebench.comT2 | History | |
| #22gpt-5.1OpenAI | 66% | date=2025-11-20 · board=Verified · system=mini-SWE-agent · model_tag=gpt-5.1-2025-11-13 | 20 Nov 2025 | swebench.comT2 | History | |
| #23GPT-5.1-CodexOpenAI | 66% | date=2025-11-24 · board=Verified · system=mini-SWE-agent · model_tag=gpt-5.1-codex | 24 Nov 2025 | swebench.comT2 | History | |
| #24gpt-5OpenAI | 65% | date=2025-08-07 · board=Verified · system=mini-SWE-agent · model_tag=gpt-5-2025-08-07 | 7 Aug 2025 | swebench.comT2 | History | |
| #25Claude 4 SonnetAnthropic | 64.93% | date=2025-07-26 · board=Verified · system=mini-SWE-agent · model_tag=claude-4-sonnet-20250514 | 26 Jul 2025 | swebench.comT2 | History | |
| #26Kimi K2 ThinkingMoonshot AI | 63.4% | date=2025-12-10 · board=Verified · system=mini-SWE-agent · model_tag=Kimi-K2-Thinking | 10 Dec 2025 | swebench.comT2 | History | |
| #27Claude 3.7 Sonnet w/ Review HeavyAnthropic | 62.4% | date=2025-02-25 · board=Verified · system=SWE-agent · model_tag=claude-3-7-sonnet-20250219 | 25 Feb 2025 | swebench.comT2 | History | |
| #28MiniMax M2MiniMax | 61% | date=2025-11-24 · board=Verified · system=mini-SWE-agent · model_tag=minimax-m2 | 24 Nov 2025 | swebench.comT2 | History | |
| #29DeepSeek V3.2 ReasonerDeepSeek | 60% | date=2025-12-01 · board=Verified · system=mini-SWE-agent · model_tag=deepseek-v3.2-reasoner | 1 Dec 2025 | swebench.comT2 | History | |
| #30gpt-5-miniOpenAI | 59.8% | date=2025-08-07 · board=Verified · system=mini-SWE-agent · model_tag=gpt-5-mini-2025-08-07 | 7 Aug 2025 | swebench.comT2 | History | |
| #31o3OpenAI | 58.4% | date=2025-07-26 · board=Verified · system=mini-SWE-agent · model_tag=o3-20250416 | 26 Jul 2025 | swebench.comT2 | History | |
| #32Devstral Small (2512)Mistral AI | 56.4% | date=2025-12-09 · board=Verified · system=mini-SWE-agent · model_tag=devstral-small-2512 | 9 Dec 2025 | swebench.comT2 | History | |
| #33gpt-5-miniOpenAI | 56.2% | date=2026-02-17 · board=Verified · system=mini-SWE-agent · model_tag=gpt-5-mini-2025-08-07 | 17 Feb 2026 | swebench.comT2 | History | |
| #34GLM 4.6Z.ai (Zhipu AI) | 55.4% | date=2025-12-01 · board=Verified · system=mini-SWE-agent · model_tag=glm-4.6 | 1 Dec 2025 | swebench.comT2 | History | |
| #35qwen3-coder-480b-a35b-instructAlibaba Group | 55.4% | date=2025-08-02 · board=Verified · system=mini-SWE-agent · model_tag=Qwen3-Coder-480B-A35B-Instruct | 2 Aug 2025 | swebench.comT2 | History | |
| #36GLM 4.5Z.ai (Zhipu AI) | 54.2% | date=2025-08-22 · board=Verified · system=mini-SWE-agent · model_tag=GLM-4.5 | 22 Aug 2025 | swebench.comT2 | History | |
| #37Devstral 2Mistral AI | 53.8% | date=2025-12-09 · board=Verified · system=mini-SWE-agent · model_tag=devstral-2512 | 9 Dec 2025 | swebench.comT2 | History | |
| #38Gemini 2.5 ProGoogle | 53.6% | date=2025-07-26 · board=Verified · system=mini-SWE-agent · model_tag=gemini-2.5-pro | 26 Jul 2025 | swebench.comT2 | History | |
| #39Kimi K2 0711Moonshot AI | 53.4% | date=2025-08-04 · board=Verified · system=CodeSweep - SWE-agent · model_tag=kimi-k2-instruct | 4 Aug 2025 | swebench.comT2 | History | |
| #40Claude 3.7 SonnetAnthropic | 52.8% | date=2025-07-20 · board=Verified · system=mini-SWE-agent · model_tag=claude-3-7-sonnet-20250219 | 20 Jul 2025 | swebench.comT2 | History | |
| #41o4-miniOpenAI | 45% | date=2025-07-26 · board=Verified · system=mini-SWE-agent · model_tag=o4-mini-20250416 | 26 Jul 2025 | swebench.comT2 | History | |
| #42Kimi K2 0711Moonshot AI | 43.8% | date=2025-08-07 · board=Verified · system=mini-SWE-agent · model_tag=Kimi-K2-Instruct | 7 Aug 2025 | swebench.comT2 | History | |
| #43SWE-agent-LM-32BQwen | 40.2% | date=2025-05-11 · board=Verified · system=SWE-agent · model_tag=Qwen 2.5 | 11 May 2025 | swebench.comT2 | History | |
| #44gpt-4.1OpenAI | 39.58% | date=2025-07-26 · board=Verified · system=mini-SWE-agent · model_tag=gpt-4.1-20250414 | 26 Jul 2025 | swebench.comT2 | History | |
| #45Devstral Small 1.1Mistral AI | 38% | date=2025-07-25 · board=Verified · system=SWE-agent · model_tag=devstral-small-2507 | 25 Jul 2025 | swebench.comT2 | History | |
| #46gpt-5-nanoOpenAI | 34.8% | date=2025-08-07 · board=Verified · system=mini-SWE-agent · model_tag=gpt-5-nano-2025-08-07 | 7 Aug 2025 | swebench.comT2 | History | |
| #47claude-35-sonnetAnthropic | 33.6% | date=2024-06-20 · board=Verified · system=SWE-agent · model_tag=claude-3-5-sonnet-20241022 | 20 Jun 2024 | swebench.comT2 | History | |
| #48Gemini 2.5 FlashGoogle | 28.73% | date=2025-07-26 · board=Verified · system=mini-SWE-agent · model_tag=gemini-2.5-flash | 26 Jul 2025 | swebench.comT2 | History | |
| #49gpt-oss-120bOpenAI | 26% | date=2025-08-07 · board=Verified · system=mini-SWE-agent · model_tag=gpt-oss-120b | 7 Aug 2025 | swebench.comT2 | History | |
| #50gpt-4.1-miniOpenAI | 23.94% | date=2025-07-20 · board=Verified · system=mini-SWE-agent · model_tag=gpt-4.1-mini-20250414 | 20 Jul 2025 | swebench.comT2 | History | |
| #51gpt-4oOpenAI | 23.2% | date=2024-07-28 · board=Verified · system=SWE-agent · model_tag=gpt-4o-2024-05-13 | 28 Jul 2024 | swebench.comT2 | History | |
| #52GPT-4 (1106)OpenAI | 22.4% | date=2024-04-02 · board=Verified · system=SWE-agent · model_tag=gpt-4-1106-preview | 2 Apr 2024 | swebench.comT2 | History | |
| #53gpt-4oOpenAI | 21.62% | date=2025-07-20 · board=Verified · system=mini-SWE-agent · model_tag=gpt-4o-20241120 | 20 Jul 2025 | swebench.comT2 | History | |
| #54Llama 4 Maverick InstructMeta AI | 21.04% | date=2025-07-20 · board=Verified · system=mini-SWE-agent · model_tag=llama-4-maverick-instruct | 20 Jul 2025 | swebench.comT2 | History | |
| #55claude-3-opusAnthropic | 15.8% | date=2024-04-02 · board=Verified · system=SWE-agent · model_tag=claude-3-opus-20240229 | 2 Apr 2024 | swebench.comT2 | History | |
| #56Gemini 2.0 FlashGoogle | 13.52% | date=2025-07-26 · board=Verified · system=mini-SWE-agent · model_tag=gemini-2.0-flash | 26 Jul 2025 | swebench.comT2 | History | |
| #57Llama 4 Scout InstructMeta AI | 9.06% | date=2025-07-20 · board=Verified · system=mini-SWE-agent · model_tag=llama-4-scout-instruct | 20 Jul 2025 | swebench.comT2 | History | |
| #58Qwen2.5 Coder 32B InstructQwen | 9% | date=2025-08-03 · board=Verified · system=mini-SWE-agent · model_tag=Qwen2.5-Coder-32B-Instruct | 3 Aug 2025 | swebench.comT2 | History | |
| #59SWE-Llama 7BMeta AI | 1.4% | date=2023-10-10 · board=Verified · system=RAG baseline · model_tag=SWE-Llama | 10 Oct 2023 | swebench.comT2 | History | |
| #60SWE-Llama 13BMeta AI | 1.2% | date=2023-10-10 · board=Verified · system=RAG baseline · submission=20231010_rag_swellama13b | 10 Oct 2023 | swebench.comT2 | History | |
| #61GPT-3.5OpenAI | 0.4% | date=2023-10-10 · board=Verified · system=RAG baseline · submission=20231010_rag_gpt35 | 10 Oct 2023 | swebench.comT2 | History |
61 results
Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.
The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.
Definition
- Category
- coding
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 8 h agomedium
- Task
- resolve real GitHub issues (500 human-validated instances)
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 8 h agomedium
- Metric
- resolved · %
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 8 h agomedium
- Direction
- Higher is better
- Paper
- https://arxiv.org/abs/2310.06770
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 8 h agomedium
- Known limitations
- Scaffold/agent dependent; results are not comparable across harnesses.
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 8 h agomedium
- Website
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 8 h agomedium
Each value shows its source, tier and observation time. Missing rows mean no source stated them. How results are recorded →
- Evaluated models
- Claude Opus 4.6, doubao-seed-code, gemini-3-flash, MiniMax M2.5, Claude Opus 4.5, Gemini 3 Pro Preview, GLM 5, GPT-5.2-Codex +74(82 total)
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history
Categorycategory1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| coding | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Known limitationsknown_limitations1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| Scaffold/agent dependent; results are not comparable across harnesses. | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Metricmetric1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| resolved | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Paperpaper1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/abs/2310.06770 | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Tasktask1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| resolve real GitHub issues (500 human-validated instances) | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Unitunit1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| % | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Websitewebsite1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://www.swebench.com | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| SWE-bench leaderboards | swebench.com/ | leaderboard | T2· Quality secondary | 6 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.