Skip to content
AI Atlas
BenchmarkActivecategory · codingfamily · swe-bench · variant full

SWE-bench (full test split)

swebench.com

resolve real GitHub issues (2,294 instances)

data quality57

Updated 6 h ago · first seen 11 Sept 2026

Metric
resolved · %
Current results
21
Models
14
Current leader
Claude Opus 4.5 52.6%

Score history · Claude Opus 4.5 1 row

Not enough history to chart — a single observation (52.62% on 19 Dec 2025). Rows under different configurations count separately; the list below shows each one.

  • 52.62%date=2025-12-19 · board=Test · system=Sonar Foundation Agent · model_tag=claude-opus-4-519 Dec 2025

Back to the leaderboard

Frontier over time · resolved · board=Test · system=Sonar Foundation Agent

1 leader change recorded, all dated 19 Dec 2025 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 52.6%Claude Opus 4.5 Anthropic Community19 Dec 2025

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 1 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1Claude Opus 4.5ClosedAnthropic · Claude52.6%Communitygroup defaultsleader19 Dec 2025swebench.comT2History

1 results

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →