Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · terminal-bench · variant 1.0

Terminal-Bench

tbench.ai

terminal tasks solved by agents

quality57

Updated 30 min ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
821
Models
374
Current leader
gpt-5.6-sol 65.9%

Score history · Claude Sonnet 5 7 rows

Not enough history to chart — 7 observations, all dated 11 Sept 2026. Rows under different configurations count separately; the list below shows each one.

  • 2.53%aa_slug=claude-sonnet-5-low · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 80.52%aa_slug=claude-sonnet-5 · variant=v2.1 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 14.14%aa_slug=claude-sonnet-5 · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 7.07%aa_slug=claude-sonnet-5-xhigh · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 5.05%aa_slug=claude-sonnet-5-high · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 75.28%aa_slug=claude-sonnet-5-non-reasoning · variant=v2.1 · evaluator=Artificial Analysis · reasoning=off11 Sept 2026
  • 2.02%aa_slug=claude-sonnet-5-medium · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=hard · evaluator=Artificial Analysis

10 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 65.9%gpt-5.6-sol OpenAI Independent11 Sept 2026
  2. 61.4%gpt-5.6-sol OpenAI Independent11 Sept 2026
  3. 53.0%Claude Sonnet 4.6 Anthropic Independent11 Sept 2026
  4. 51.5%Claude Opus 4.7 Anthropic Independent11 Sept 2026
  5. 50.8%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
  6. 49.2%KAT-Coder-Pro V2 Kwaipilot Independent11 Sept 2026
  7. 33.3%GPT-5.1-Codex Mini OpenAI Independent11 Sept 2026
  8. 26.5%Grok 4.3 xAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 315 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 5 rows65.9%Independentgroup defaultsleaderobs. 11 Sept 2026artificialanalysis.aiT2History
2Claude Fable 5ClosedAnthropic · Claude62.9%Independentgroup defaultsComparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
2gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 4 rows62.9%Independentreasoning_effortxhighPartially comparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
4gpt-5.5ClosedOpenAI · GPT 5.5 · best of 5 rows60.6%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
5Claude Opus 4.8ClosedAnthropic · Claude58.3%Independentgroup defaultsComparable-7.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
6gpt-5.4ClosedOpenAI · GPT 5.4 · best of 3 rows57.6%Independentgroup defaultsComparable-8.33 ptobs. 11 Sept 2026artificialanalysis.aiT2History
7Claude Opus 4.7ClosedAnthropic · Claude · best of 2 rows54.5%IndependentreasoningoffPartially comparable-11.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
8Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.153.8%Independentgroup defaultsComparable-12.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
9Claude Sonnet 4.6ClosedAnthropic · Claude · best of 3 rows53.0%IndependentreasoningadaptivePartially comparable-12.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
9gpt-5.3-codexClosedOpenAI · GPT 5.353.0%Independentgroup defaultsComparable-12.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
11gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 3 rows52.3%Independentgroup defaultsComparable-13.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
12Qwen3.7 MaxClosedQwen · Qwen3.750.8%Independentgroup defaultsComparable-15.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
12Z.ai GLM 5.2Open weightsZ.ai (Zhipu AI) · GLM5.250.8%Independentgroup defaultsComparable-15.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
14KAT-Coder-Pro V2ClosedKwaipilot49.2%Independentgroup defaultsComparable-16.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
15Claude Opus 4.6ClosedAnthropic · Claude · best of 2 rows48.5%Independentgroup defaultsComparable-17.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
16Claude Opus 4.5ClosedAnthropic · Claude · best of 2 rows47.0%IndependentreasoningonPartially comparable-18.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
16Qwen 3.7 PlusClosedQwen · Qwen3.747.0%Independentgroup defaultsComparable-18.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
16gpt-5.2ClosedOpenAI · GPT 5.2 · best of 3 rows47.0%Independentgroup defaultsComparable-18.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
19Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 3 rows46.2%Independentreasoning_effortminimalPartially comparable-19.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
19deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek46.2%Independentgroup defaultsComparable-19.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
21gpt-5.1ClosedOpenAI · GPT 5.1 · best of 2 rows45.5%Independentgroup defaultsComparable-20.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
21muse-sparkClosedMeta AI45.5%Independentgroup defaultsComparable-20.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
23Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi44.7%Independentgroup defaultsComparable-21.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
24Kimi K2.6Open weightsMoonshot AI · Kimi · best of 2 rows43.9%Independentgroup defaultsComparable-22.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
24Qwen3.6 Max PreviewClosedQwen · Qwen3.643.9%Independentgroup defaultsComparable-22.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
24Qwen3.6 PlusClosedQwen · Qwen3.643.9%Independentgroup defaultsComparable-22.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
27GLM 5Open weightsZ.ai (Zhipu AI) · GLM5 · best of 2 rows43.2%Independentgroup defaultsComparable-22.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
27GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 2 rows43.2%Independentgroup defaultsComparable-22.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
27MiMo-V2.5-ProOpen weightsXiaomi · best of 2 rows43.2%Independentgroup defaultsComparable-22.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
30MiniMax-M3Open weightsMiniMax · MiniMax42.4%Independentgroup defaultsComparable-23.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
30gpt-5-5-instant-05-26ClosedOpenAI · GPT 5.542.4%Independentgroup defaultsComparable-23.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
30gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 3 rows42.4%Independentgroup defaultsComparable-23.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
33MiMo-V2.5Open weightsXiaomi41.7%Independentgroup defaultsComparable-24.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
33deepseek-v4-pro-0424-highOpen weightsDeepSeek · DeepSeek41.7%Independentgroup defaultsComparable-24.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
33gemini-3-proClosedGoogle · Gemini 3 · best of 2 rows41.7%Independentgroup defaultsComparable-24.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
36Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 2 rows40.9%Independentgroup defaultsComparable-25.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
36grok-4-20-0309ClosedSpaceXAI · Grok 4.20 · best of 2 rows40.9%Independentgroup defaultsComparable-25.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
36mimo-v2-proClosedXiaomi40.9%Independentgroup defaultsComparable-25.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
39MiniMax M2.7Open weightsMiniMax · MiniMax39.4%Independentgroup defaultsComparable-26.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
40Gemini 3 Flash PreviewClosedGoogle · Gemini 338.6%Independentgroup defaultsComparable-27.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
40deepseek-v4-flash-0420-highOpen weightsDeepSeek · DeepSeek38.6%Independentgroup defaultsComparable-27.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
42GPT-5-CodexClosedOpenAI · GPT 537.9%Independentgroup defaultsComparable-28.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
42Grok 4.20ClosedxAI · Grok37.9%Independentgroup defaultsComparable-28.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
42Grok 4.3ClosedxAI · Grok · best of 4 rows37.9%Independentgroup defaultsComparable-28.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
42gpt-5ClosedOpenAI · GPT 5 · best of 4 rows37.9%Independentreasoning_effortmediumPartially comparable-28.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
42grok-4ClosedSpaceXAI · Grok 437.9%Independentgroup defaultsComparable-28.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
47GPT-5.2-CodexClosedOpenAI · GPT 5.237.1%Independentgroup defaultsComparable-28.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
47o3ClosedOpenAI · OpenAI o-series37.1%Independentgroup defaultsComparable-28.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
49Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 2 rows36.4%Independentgroup defaultsComparable-29.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
49Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 336.4%Independentgroup defaultsComparable-29.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
49deepseek-v4-pro-0424-non-reasoningOpen weightsDeepSeek · DeepSeek36.4%Independentgroup defaultsComparable-29.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
52Claude Sonnet 4.5ClosedAnthropic · Claude · best of 2 rows35.6%IndependentreasoningonPartially comparable-30.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
52DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V335.6%Independentgroup defaultsComparable-30.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
52Step 3.7 FlashOpen weightsStepFun · Step3.735.6%Independentgroup defaultsComparable-30.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
52deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek35.6%Independentgroup defaultsComparable-30.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
52mimo-v2-omni-0327ClosedXiaomi35.6%Independentgroup defaultsComparable-30.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
57GPT-5.1-CodexClosedOpenAI · GPT 5.134.9%Independentgroup defaultsComparable-31.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
57Kimi K2.5Open weightsMoonshot AI · Kimi · best of 2 rows34.9%Independentgroup defaultsComparable-31.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
57MiniMax M2.5Open weightsMiniMax · MiniMax34.9%Independentgroup defaultsComparable-31.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
57Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 2 rows34.9%Independentgroup defaultsComparable-31.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
57Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 2 rows34.9%Independentgroup defaultsComparable-31.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
57deepseek-v3-2-specialeOpen weightsDeepSeek · DeepSeek34.9%Independentgroup defaultsComparable-31.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
57mimo-v2-omniClosedXiaomi34.9%Independentgroup defaultsComparable-31.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
64Claude Opus 4.1ClosedAnthropic · Claude34.3%IndependentreasoningonPartially comparable-31.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
65Hy3 previewOpen weightsTencent34.1%Independentgroup defaultsComparable-31.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
65deepseek-v4-flash-0420-non-reasoningOpen weightsDeepSeek · DeepSeek34.1%Independentgroup defaultsComparable-31.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
67GLM 5 TurboClosedZ.ai (Zhipu AI) · GLM533.3%Independentgroup defaultsComparable-32.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
67GPT-5.1-Codex MiniClosedOpenAI · GPT 5.133.3%Independentgroup defaultsComparable-32.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
67Mistral Medium 3.5Open weightsMistral AI · Mistral33.3%Independentgroup defaultsComparable-32.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
67gpt-5-miniClosedOpenAI · GPT 5 · best of 3 rows33.3%Independentgroup defaultsComparable-32.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
71DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 2 rows32.6%Independentgroup defaultsComparable-33.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
71GLM 5V TurboClosedZ.ai (Zhipu AI) · GLM532.6%Independentgroup defaultsComparable-33.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
71Qwen3.5-27BOpen weightsQwen · Qwen3.5 · best of 2 rows32.6%Independentgroup defaultsComparable-33.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
71Step 3.5 FlashOpen weightsStepFun · Step3.5 · best of 2 rows32.6%Independentgroup defaultsComparable-33.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
75DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 2 rows31.8%Independentgroup defaultsComparable-34.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
75GLM 4.7Open weightsZ.ai (Zhipu AI) · GLM4.7 · best of 2 rows31.8%Independentgroup defaultsComparable-34.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
75Hy3Open weightsTencent31.8%IndependentreasoningoffPartially comparable-34.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
75gemini-3-flashClosedGoogle · Gemini 331.8%Independentgroup defaultsComparable-34.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
79Claude Opus 4ClosedAnthropic · Claude31.1%Independentgroup defaultsComparable-34.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
79Claude Sonnet 4ClosedAnthropic · Claude · best of 2 rows31.1%IndependentreasoningonPartially comparable-34.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
79DeepSeek V3.2 ExpOpen weightsDeepSeek · DeepSeek-V331.1%Independentgroup defaultsComparable-34.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
79Kimi K2 ThinkingOpen weightsMoonshot AI · Kimi31.1%Independentgroup defaultsComparable-34.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
79North Mini Code (free)Open weightsCohere31.1%Independentgroup defaultsComparable-34.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
79Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 2 rows31.1%Independentgroup defaultsComparable-34.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
79ling-2-6-1tOpen weightsinclusionAI31.1%Independentgroup defaultsComparable-34.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
79mimo-v2-0206Open weightsXiaomi31.1%Independentgroup defaultsComparable-34.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
87GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows28.8%Independentgroup defaultsComparable-37.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
87MiniMax M2.1Open weightsMiniMax · MiniMax28.8%Independentgroup defaultsComparable-37.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
87Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 328.8%Independentgroup defaultsComparable-37.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
87jt-35b-flashClosedChina Mobile28.8%Independentgroup defaultsComparable-37.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
87ring-2-6-1tOpen weightsinclusionAI28.8%Independentgroup defaultsComparable-37.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
92mimo-v2-flashOpen weightsXiaomi · best of 2 rows28.0%IndependentreasoningonPartially comparable-37.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
93Claude Haiku 4.5ClosedAnthropic · Claude · best of 2 rows27.3%Independentgroup defaultsComparable-38.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
94Gemini 2.5 ProClosedGoogle · Gemini 2.526.5%Independentgroup defaultsComparable-39.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
94Mercury 2ClosedInception26.5%Independentgroup defaultsComparable-39.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
94Qwen3.5-35B-A3BOpen weightsQwen · Qwen3.5 · best of 2 rows26.5%Independentgroup defaultsComparable-39.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
94doubao-seed-codeClosedByteDance · Seed26.5%Independentgroup defaultsComparable-39.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
98MiniMax M2Open weightsMiniMax · MiniMax25.8%Independentgroup defaultsComparable-40.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
99DeepSeek V3.1Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows25%IndependentreasoningonPartially comparable-40.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
99Gemma 4 26B A4BOpen weightsGoogle · Gemma 4 · best of 2 rows25%IndependentreasoningoffPartially comparable-40.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →