Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · terminal-bench · variant 1.0

Terminal-Bench

tbench.ai

terminal tasks solved by agents

quality57

Updated 30 min ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
821
Models
374
Current leader
gpt-5.6-sol 65.9%

Score history · granite-4-0-h-nano-1b 1 row

Not enough history to chart — a single observation (0% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.

  • 0%aa_slug=granite-4-0-h-nano-1b · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=hard · evaluator=Artificial Analysis

10 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 65.9%gpt-5.6-sol OpenAI Independent11 Sept 2026
  2. 61.4%gpt-5.6-sol OpenAI Independent11 Sept 2026
  3. 53.0%Claude Sonnet 4.6 Anthropic Independent11 Sept 2026
  4. 51.5%Claude Opus 4.7 Anthropic Independent11 Sept 2026
  5. 50.8%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
  6. 49.2%KAT-Coder-Pro V2 Kwaipilot Independent11 Sept 2026
  7. 33.3%GPT-5.1-Codex Mini OpenAI Independent11 Sept 2026
  8. 26.5%Grok 4.3 xAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 315 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
99command-a-plusOpen weightsCohere · Command25%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
99deepseek-v3-2-0925Open weightsDeepSeek · DeepSeek25%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
99ernie-5-0-thinking-previewClosedBaidu · ERNIE 5.025%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
104Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.124.2%Independentgroup defaultsComparable-0.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
104Qwen3 MaxClosedQwen · Qwen3 · best of 2 rows24.2%IndependentreasoningonPartially comparable-0.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
104Qwen3.5-9BOpen weightsQwen · Qwen3.5 · best of 2 rows24.2%Independentgroup defaultsComparable-0.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
104grok-4-1-fastClosedSpaceXAI · Grok 4.1 · best of 2 rows24.2%IndependentreasoningonPartially comparable-0.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
104nova-2-0-proClosedAmazon Web Services · Nova 2.0 · best of 3 rows24.2%Independentreasoningonreasoning_effortmediumPartially comparable-0.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
109Kimi K2 0905Open weightsMoonshot AI · Kimi23.5%Independentgroup defaultsComparable-1.52 ptobs. 11 Sept 2026artificialanalysis.aiT2History
109gpt-oss-120bOpen weightsOpenAI · gpt-oss · best of 2 rows23.5%Independentgroup defaultsComparable-1.52 ptobs. 11 Sept 2026artificialanalysis.aiT2History
109hypernova-60bOpen weightsMultiverse Computing23.5%Independentgroup defaultsComparable-1.52 ptobs. 11 Sept 2026artificialanalysis.aiT2History
112Trinity Large ThinkingOpen weightsArcee AI22.7%Independentgroup defaultsComparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
112k-exaoneOpen weightsLG AI Research · EXAONE · best of 2 rows22.7%Independentgroup defaultsComparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
114GLM 4.5Open weightsZ.ai (Zhipu AI) · GLM4.522.0%Independentgroup defaultsComparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
114GLM 4.7 FlashOpen weightsZ.ai (Zhipu AI) · GLM4.7 · best of 2 rows22.0%Independentgroup defaultsComparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
116Claude 3.7 SonnetClosedAnthropic · Claude · best of 2 rows21.2%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
116ling-2-6-flashOpen weightsinclusionAI21.2%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
116nemotron-cascade-2-30b-a3bOpen weightsNVIDIA · Nemotron21.2%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
116qwen3-5-omni-plusClosedAlibaba Group · Qwen3.521.2%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
120GLM 4.5 AirOpen weightsZ.ai (Zhipu AI) · GLM4.520.4%Independentgroup defaultsComparable-4.55 ptobs. 11 Sept 2026artificialanalysis.aiT2History
120exaone-4-5-33bOpen weightsLG AI Research · EXAONE 4.520.4%Independentgroup defaultsComparable-4.55 ptobs. 11 Sept 2026artificialanalysis.aiT2History
122qwen3-max-previewClosedAlibaba Group · Qwen319.7%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
123Devstral 2Open weightsMistral AI · Devstral 218.9%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
123grok-4-fastClosedSpaceXAI · Grok 4 · best of 2 rows18.9%IndependentreasoningonPartially comparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
123qwen3-coder-480b-a35b-instructOpen weightsAlibaba Group · Qwen3-Coder18.9%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
126Qwen3 Coder NextOpen weightsQwen · Qwen318.2%Independentgroup defaultsComparable-6.82 ptobs. 11 Sept 2026artificialanalysis.aiT2History
126Qwen3.5-4BOpen weightsQwen · Qwen3.5 · best of 2 rows18.2%Independentgroup defaultsComparable-6.82 ptobs. 11 Sept 2026artificialanalysis.aiT2History
126gemma-4-12BOpen weightsGoogle · Gemma 4 · best of 2 rows18.2%Independentgroup defaultsComparable-6.82 ptobs. 11 Sept 2026artificialanalysis.aiT2History
126jt-miniClosedChina Mobile18.2%Independentgroup defaultsComparable-6.82 ptobs. 11 Sept 2026artificialanalysis.aiT2History
130Grok Build 0.1ClosedxAI · Grok17.4%Independentgroup defaultsComparable-7.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
130Mistral Small 4Open weightsMistral AI · Mistral · best of 2 rows17.4%Independentgroup defaultsComparable-7.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
130gpt-5-nanoClosedOpenAI · GPT 5 · best of 3 rows17.4%Independentreasoning_effortmediumPartially comparable-7.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
130grok-3-mini-reasoningClosedSpaceXAI · Grok 317.4%Independentgroup defaultsComparable-7.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
130nova-2-0-liteClosedAmazon Web Services · Nova 2.0 · best of 4 rows17.4%Independentreasoningonreasoning_effortmediumPartially comparable-7.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
130qwen3-max-thinking-previewClosedAlibaba Group · Qwen317.4%Independentgroup defaultsComparable-7.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
136Devstral Small 2Open weightsMistral AI · Devstral16.7%Independentgroup defaultsComparable-8.33 ptobs. 11 Sept 2026artificialanalysis.aiT2History
136cogito-v2-1-reasoningOpen weightsDeep Cogito · Cogito16.7%Independentgroup defaultsComparable-8.33 ptobs. 11 Sept 2026artificialanalysis.aiT2History
136gemini-2-5-flash-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 2 rows16.7%IndependentreasoningonPartially comparable-8.33 ptobs. 11 Sept 2026artificialanalysis.aiT2History
136grok-4.20-0309-non-reasoningClosedxAI · Grok16.7%Independentgroup defaultsComparable-8.33 ptobs. 11 Sept 2026artificialanalysis.aiT2History
140Kimi K2 0711Open weightsMoonshot AI · Kimi15.9%Independentgroup defaultsComparable-9.09 ptobs. 11 Sept 2026artificialanalysis.aiT2History
140Mistral Large 3Open weightsMistral AI · Mistral15.9%Independentgroup defaultsComparable-9.09 ptobs. 11 Sept 2026artificialanalysis.aiT2History
140deepseek-r1Open weightsDeepSeek · DeepSeek-R115.9%Independentgroup defaultsComparable-9.09 ptobs. 11 Sept 2026artificialanalysis.aiT2History
143DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V315.2%Independentgroup defaultsComparable-9.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
143Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 2 rows15.2%Independentgroup defaultsComparable-9.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
143Qwen3 Coder 30B A3B InstructOpen weightsQwen · Qwen315.2%Independentgroup defaultsComparable-9.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
143o4-miniClosedOpenAI · OpenAI o-series15.2%Independentgroup defaultsComparable-9.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
147GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows14.4%Independentgroup defaultsComparable-10.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
147apriel-v1-6-15b-thinkerOpen weightsServiceNow14.4%Independentgroup defaultsComparable-10.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
149Gemini 2.5 FlashClosedGoogle · Gemini 2.5 · best of 2 rows13.6%IndependentreasoningonPartially comparable-11.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
149Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 2 rows13.6%IndependentreasoningonPartially comparable-11.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
149gpt-4.1ClosedOpenAI · GPT 4.113.6%Independentgroup defaultsComparable-11.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
152Magistral Medium 1.2ClosedMistral AI · Magistral12.9%Independentgroup defaultsComparable-12.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
152gemini-2-5-flash-lite-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 2 rows12.9%IndependentreasoningonPartially comparable-12.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
152gpt-5-chatgptClosedOpenAI · GPT 512.9%Independentgroup defaultsComparable-12.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
152o1ClosedOpenAI · OpenAI o-series12.9%Independentgroup defaultsComparable-12.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
156hyperclova-x-seed-think-32bOpen weightsNaver · Seed12.1%Independentgroup defaultsComparable-12.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
157Hermes 4 405BOpen weightsNous Research · Hermes 411.4%Independentgroup defaultsComparable-13.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
157Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AI · Kimi11.4%Independentgroup defaultsComparable-13.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
157Qwen3 VL 235B A22B InstructOpen weightsQwen · Qwen3 · best of 2 rows11.4%IndependentreasoningonPartially comparable-13.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
157grok-3ClosedSpaceXAI · Grok 311.4%Independentgroup defaultsComparable-13.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
161Mistral Medium 3.1ClosedMistral AI · Mistral10.6%Independentgroup defaultsComparable-14.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
161apriel-v1-5-15b-thinkerOpen weightsServiceNow10.6%Independentgroup defaultsComparable-14.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
161gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 2 rows10.6%Independentgroup defaultsComparable-14.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
161ling-1tOpen weightsinclusionAI10.6%Independentgroup defaultsComparable-14.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
161ling-flash-2-0Open weightsinclusionAI10.6%Independentgroup defaultsComparable-14.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
161longcat-flash-liteOpen weightsLongCat10.6%Independentgroup defaultsComparable-14.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
167Qwen3 Next 80B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows9.85%IndependentreasoningonPartially comparable-15.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
167hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.19.85%Independentgroup defaultsComparable-15.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
167k2-v2Open weightsMBZUAI Institute of Foundation Models · best of 3 rows9.85%Independentgroup defaultsComparable-15.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
170devstral-mediumClosedMistral AI · Devstral9.09%Independentgroup defaultsComparable-15.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
170intellect-3Open weightsPrime Intellect9.09%Independentgroup defaultsComparable-15.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
170kat-coder-pro-v1ClosedKwaiKAT9.09%Independentgroup defaultsComparable-15.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
170magistral-mediumClosedMistral AI · Magistral9.09%Independentgroup defaultsComparable-15.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
174GPT-4o (2024-08-06)ClosedOpenAI · GPT 48.33%Independentgroup defaultsComparable-16.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
174Qwen3 VL 32B InstructOpen weightsQwen · Qwen3 · best of 2 rows8.33%Independentgroup defaultsComparable-16.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
174gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 2 rows8.33%Independentgroup defaultsComparable-16.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
174gpt-4oClosedOpenAI · GPT 48.33%Independentgroup defaultsComparable-16.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
174nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 38.33%Independentgroup defaultsComparable-16.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
174qwen3-5-omni-flashClosedAlibaba Group · Qwen3.58.33%Independentgroup defaultsComparable-16.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
180Mistral Small 3.1Open weightsMistral AI · Mistral7.58%Independentgroup defaultsComparable-17.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
180Solar Pro 3ClosedUpstage · Solar7.58%Independentgroup defaultsComparable-17.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
180gpt-4.1-miniClosedOpenAI · GPT 4.17.58%Independentgroup defaultsComparable-17.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
180ring-flash-2-0Open weightsinclusionAI7.58%Independentgroup defaultsComparable-17.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows6.82%Independentgroup defaultsComparable-18.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184Llama 4 MaverickOpen weightsMeta AI · Llama 46.82%Independentgroup defaultsComparable-18.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184Llama-3.1-405BRestricted weightsMeta AI · Llama 3.16.82%Independentgroup defaultsComparable-18.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184Mistral Small 3.2Open weightsMistral AI · Mistral6.82%Independentgroup defaultsComparable-18.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184Seed-OSS-36B-InstructOpen weightsByteDance · Seed6.82%Independentgroup defaultsComparable-18.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184k2-think-v2Open weightsMBZUAI Institute of Foundation Models6.82%Independentgroup defaultsComparable-18.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184nova-2-0-omniClosedAmazon Web Services · Nova 2.0 · best of 3 rows6.82%Independentgroup defaultsComparable-18.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184nova-premierClosedAmazon Web Services · Nova6.82%Independentgroup defaultsComparable-18.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 36.82%Independentgroup defaultsComparable-18.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184o3-miniClosedOpenAI · OpenAI o-series · best of 2 rows6.82%Independentgroup defaultsComparable-18.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184qwen3-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows6.82%Independentgroup defaultsComparable-18.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184ring-1tOpen weightsinclusionAI6.82%Independentgroup defaultsComparable-18.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
196Devstral Small 1.0Open weightsMistral AI · Devstral6.06%Independentgroup defaultsComparable-18.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
196Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows6.06%Independentgroup defaultsComparable-18.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
196deepseek-r1-0120Open weightsDeepSeek · DeepSeek6.06%Independentgroup defaultsComparable-18.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
196devstral-smallOpen weightsMistral AI · Devstral6.06%Independentgroup defaultsComparable-18.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
196ernie-4-5-300b-a47bOpen weightsBaidu · ERNIE 4.56.06%Independentgroup defaultsComparable-18.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →