Skip to content
AI Atlas
BenchmarkActivecategory · codingfamily · swe-bench · variant Multilingual

SWE-bench Multilingual

swebench.com/multilingual

resolve GitHub issues across 9 programming languages

data quality57

Updated 6 h ago · first seen 11 Sept 2026

Metric
resolved · %
Current results
13
Models
13
Current leader
gemini-3-flash 72.7%

Score history · GLM 5 1 row

Not enough history to chart — a single observation (69.7% on 13 Feb 2026). Rows under different configurations count separately; the list below shows each one.

  • 69.7%date=2026-02-13 · board=Multilingual · system=mini-SWE-agent · model_tag=glm-513 Feb 2026

Back to the leaderboard

Frontier over time · resolved · board=Multilingual · system=mini-SWE-agent

1 leader change recorded, all dated 13 Feb 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 72.7%gemini-3-flash Google Official board13 Feb 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 13 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1gemini-3-flashClosedGoogle · Gemini 372.7%Official boardgroup defaultsleader13 Feb 2026swebench.comT2History
2Claude Opus 4.6ClosedAnthropic · Claude72%Official boardgroup defaultsPartially comparable-0.70 pt13 Feb 2026swebench.comT2History
3Claude Opus 4.5ClosedAnthropic · Claude70.7%Official boardgroup defaultsPartially comparable-2.00 pt13 Feb 2026swebench.comT2History
4GLM 5Open weightsZ.ai (Zhipu AI) · GLM569.7%Official boardgroup defaultsPartially comparable-3.00 pt13 Feb 2026swebench.comT2History
5gemini-3-proClosedGoogle · Gemini 368.7%Official boardgroup defaultsPartially comparable-4.00 pt13 Feb 2026swebench.comT2History
6MiniMax M2.5Open weightsMiniMax · MiniMax68.3%Official boardgroup defaultsPartially comparable-4.40 pt16 Feb 2026swebench.comT2History
7Kimi K2.5Open weightsMoonshot AI · Kimi67.3%Official boardgroup defaultsPartially comparable-5.40 pt13 Feb 2026swebench.comT2History
8Claude 4.5 SonnetAnthropic · Claude 4.567%Official boardgroup defaultsPartially comparable-5.70 pt13 Feb 2026swebench.comT2History
9gpt-5.2ClosedOpenAI · GPT 5.266.7%Official boardreasoning_efforthighPartially comparable-6.00 pt13 Feb 2026swebench.comT2History
10GPT-5.2-CodexClosedOpenAI · GPT 5.266.3%Official boardgroup defaultsPartially comparable-6.40 pt20 Feb 2026swebench.comT2History
11Claude 4.5 HaikuAnthropic · Claude 4.564.7%Official boardgroup defaultsPartially comparable-8.00 pt13 Feb 2026swebench.comT2History
12DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V359%Official boardgroup defaultsPartially comparable-13.7 pt13 Feb 2026swebench.comT2History
13gpt-5-miniClosedOpenAI · GPT 539.7%Official boardgroup defaultsPartially comparable-33.0 pt13 Feb 2026swebench.comT2History

13 results

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →