Skip to content
AI Atlas
BenchmarkActive

SWE-bench Multilingual

swebench.com/multilingual

resolve GitHub issues across 9 programming languages

data quality57

Updated 15 h ago · first seen 11 Sept 2026

Metric
· %
Current results
0
Models
0
Current leader
gemini-3-flash 72.7%

No results recorded for this benchmark yet — its sources are being connected. The definition, aliases and variants are kept so links resolve; nothing is fabricated.

Leaderboard 13 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1gemini-3-flashClosedGoogle · Gemini 372.7%Official boardgroup defaultsleader13 Feb 2026swebench.comT2History
2Claude Opus 4.6ClosedAnthropic · Claude72%Official boardgroup defaultsPartially comparable-0.70 pt13 Feb 2026swebench.comT2History
3Claude Opus 4.5ClosedAnthropic · Claude70.7%Official boardgroup defaultsPartially comparable-2.00 pt13 Feb 2026swebench.comT2History
4GLM 5Open weightsZ.ai (Zhipu AI) · GLM569.7%Official boardgroup defaultsPartially comparable-3.00 pt13 Feb 2026swebench.comT2History
5gemini-3-proClosedGoogle · Gemini 368.7%Official boardgroup defaultsPartially comparable-4.00 pt13 Feb 2026swebench.comT2History
6MiniMax M2.5Open weightsMiniMax · MiniMax68.3%Official boardgroup defaultsPartially comparable-4.40 pt16 Feb 2026swebench.comT2History
7Kimi K2.5Open weightsMoonshot AI · Kimi67.3%Official boardgroup defaultsPartially comparable-5.40 pt13 Feb 2026swebench.comT2History
8Claude 4.5 SonnetAnthropic · Claude 4.567%Official boardgroup defaultsPartially comparable-5.70 pt13 Feb 2026swebench.comT2History
9gpt-5.2ClosedOpenAI · GPT 5.266.7%Official boardreasoning_efforthighPartially comparable-6.00 pt13 Feb 2026swebench.comT2History
10GPT-5.2-CodexClosedOpenAI · GPT 5.266.3%Official boardgroup defaultsPartially comparable-6.40 pt20 Feb 2026swebench.comT2History
11Claude 4.5 HaikuAnthropic · Claude 4.564.7%Official boardgroup defaultsPartially comparable-8.00 pt13 Feb 2026swebench.comT2History
12DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V359%Official boardgroup defaultsPartially comparable-13.7 pt13 Feb 2026swebench.comT2History
13gpt-5-miniClosedOpenAI · GPT 539.7%Official boardgroup defaultsPartially comparable-33.0 pt13 Feb 2026swebench.comT2History

13 results

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →