Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · terminal-bench · variant 1.0

Terminal-Bench

tbench.ai

terminal tasks solved by agents

data quality57

Updated 2 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,639
Models
376
Current leader
gpt-5.6-sol 65.9%

Score history · Llama-3.2-1B 2 rows

Score history for Llama-3.2-1B0%0.2%0.4%0.6%0.8%1%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • Llama-3.2-1B
  • 0%aa_slug=llama-3-2-instruct-1b · variant=hard · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
  • 0%aa_slug=llama-3-2-instruct-1b · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=hard · evaluator=Artificial Analysis

10 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 65.9%gpt-5.6-sol OpenAI Independent11 Sept 2026
  2. 61.4%gpt-5.6-sol OpenAI Independent11 Sept 2026
  3. 53.0%Claude Sonnet 4.6 Anthropic Independent11 Sept 2026
  4. 51.5%Claude Opus 4.7 Anthropic Independent11 Sept 2026
  5. 50.8%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
  6. 49.2%KAT-Coder-Pro V2 Kwaipilot Independent11 Sept 2026
  7. 33.3%GPT-5.1-Codex Mini OpenAI Independent11 Sept 2026
  8. 26.5%Grok 4.3 xAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 317 models · trust independent-evaluator

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
100Gemma 4 26B A4BOpen weightsGoogle · Gemma 4 · best of 4 rows25%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
100command-a-plusOpen weightsCohere · Command · best of 2 rows25%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
100deepseek-v3-2-0925Open weightsDeepSeek · DeepSeek25%Independentgroup defaultsPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
100ernie-5-0-thinking-previewClosedBaidu · ERNIE 5.0 · best of 2 rows25%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
105Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.1 · best of 2 rows24.2%IndependentreasoningonPartially comparable-0.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
105Qwen3 MaxClosedQwen · Qwen3 · best of 4 rows24.2%IndependentreasoningonPartially comparable-0.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
105Qwen3.5-9BOpen weightsQwen · Qwen3.5 · best of 4 rows24.2%IndependentreasoningonPartially comparable-0.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
105grok-4-1-fastClosedSpaceXAI · Grok 4.1 · best of 4 rows24.2%IndependentreasoningonPartially comparable-0.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
105nova-2-0-proClosedAmazon Web Services · Nova 2.0 · best of 6 rows24.2%Independentreasoningonreasoning_effortmediumPartially comparable-0.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
110Kimi K2 0905Open weightsMoonshot AI · Kimi · best of 2 rows23.5%IndependentreasoningoffPartially comparable-1.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
110gpt-oss-120bOpen weightsOpenAI · gpt-oss · best of 4 rows23.5%Independentreasoning_efforthighPartially comparable-1.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
110hypernova-60bOpen weightsMultiverse Computing · best of 2 rows23.5%Independentreasoning_efforthighPartially comparable-1.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
113Trinity Large ThinkingOpen weightsArcee AI · best of 2 rows22.7%IndependentreasoningonPartially comparable-2.27 ptobs. 12 Sept 2026artificialanalysis.aiT2History
113k-exaoneOpen weightsLG AI Research · EXAONE · best of 4 rows22.7%IndependentreasoningonPartially comparable-2.27 ptobs. 12 Sept 2026artificialanalysis.aiT2History
115GLM 4.5Open weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows22.0%IndependentreasoningonPartially comparable-3.03 ptobs. 12 Sept 2026artificialanalysis.aiT2History
115GLM 4.7 FlashOpen weightsZ.ai (Zhipu AI) · GLM4.7 · best of 4 rows22.0%IndependentreasoningonPartially comparable-3.03 ptobs. 12 Sept 2026artificialanalysis.aiT2History
117Claude 3.7 SonnetClosedAnthropic · Claude · best of 4 rows21.2%IndependentreasoningoffPartially comparable-3.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
117ling-2-6-flashOpen weightsinclusionAI · best of 2 rows21.2%IndependentreasoningoffPartially comparable-3.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
117nemotron-cascade-2-30b-a3bOpen weightsNVIDIA · Nemotron · best of 2 rows21.2%IndependentreasoningonPartially comparable-3.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
117qwen3-5-omni-plusClosedAlibaba Group · Qwen3.5 · best of 2 rows21.2%IndependentreasoningoffPartially comparable-3.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
121GLM 4.5 AirOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows20.4%IndependentreasoningonPartially comparable-4.55 ptobs. 12 Sept 2026artificialanalysis.aiT2History
121exaone-4-5-33bOpen weightsLG AI Research · EXAONE 4.5 · best of 2 rows20.4%IndependentreasoningonPartially comparable-4.55 ptobs. 12 Sept 2026artificialanalysis.aiT2History
123qwen3-max-previewClosedAlibaba Group · Qwen3 · best of 2 rows19.7%IndependentreasoningoffPartially comparable-5.30 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124Devstral 2Open weightsMistral AI · Devstral 2 · best of 2 rows18.9%IndependentreasoningoffPartially comparable-6.06 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124grok-4-fastClosedSpaceXAI · Grok 4 · best of 4 rows18.9%IndependentreasoningonPartially comparable-6.06 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124qwen3-coder-480b-a35b-instructOpen weightsAlibaba Group · Qwen3-Coder · best of 2 rows18.9%IndependentreasoningoffPartially comparable-6.06 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127Qwen3 Coder NextOpen weightsQwen · Qwen3 · best of 2 rows18.2%IndependentreasoningoffPartially comparable-6.82 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127Qwen3.5-4BOpen weightsQwen · Qwen3.5 · best of 4 rows18.2%IndependentreasoningonPartially comparable-6.82 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127gemma-4-12BOpen weightsGoogle · Gemma 4 · best of 4 rows18.2%IndependentreasoningonPartially comparable-6.82 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127jt-miniClosedChina Mobile · best of 2 rows18.2%IndependentreasoningoffPartially comparable-6.82 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131Grok Build 0.1ClosedxAI · Grok · best of 2 rows17.4%IndependentreasoningonPartially comparable-7.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131Mistral Small 4Open weightsMistral AI · Mistral · best of 4 rows17.4%IndependentreasoningonPartially comparable-7.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131gpt-5-nanoClosedOpenAI · GPT 5 · best of 6 rows17.4%Independentreasoning_effortmediumPartially comparable-7.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131grok-3-mini-reasoningClosedSpaceXAI · Grok 3 · best of 2 rows17.4%Independentreasoningonreasoning_efforthighPartially comparable-7.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131nova-2-0-liteClosedAmazon Web Services · Nova 2.0 · best of 8 rows17.4%Independentreasoningonreasoning_effortmediumPartially comparable-7.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131qwen3-max-thinking-previewClosedAlibaba Group · Qwen3 · best of 2 rows17.4%IndependentreasoningonPartially comparable-7.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
137Devstral Small 2Open weightsMistral AI · Devstral · best of 2 rows16.7%IndependentreasoningoffPartially comparable-8.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
137cogito-v2-1-reasoningOpen weightsDeep Cogito · Cogito · best of 2 rows16.7%IndependentreasoningonPartially comparable-8.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
137gemini-2-5-flash-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 6 rows16.7%IndependentreasoningonPartially comparable-8.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
137grok-4.20-0309-non-reasoningClosedxAI · Grok16.7%Independentgroup defaultsPartially comparable-8.33 ptobs. 11 Sept 2026artificialanalysis.aiT2History
141Kimi K2 0711Open weightsMoonshot AI · Kimi · best of 2 rows15.9%IndependentreasoningoffPartially comparable-9.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
141Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows15.9%IndependentreasoningoffPartially comparable-9.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
141deepseek-r1Open weightsDeepSeek · DeepSeek-R1 · best of 2 rows15.9%IndependentreasoningonPartially comparable-9.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
144DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows15.2%IndependentreasoningoffPartially comparable-9.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
144Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 4 rows15.2%IndependentreasoningoffPartially comparable-9.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
144Qwen3 Coder 30B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows15.2%IndependentreasoningoffPartially comparable-9.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
144o4-miniClosedOpenAI · OpenAI o-series · best of 2 rows15.2%Independentreasoning_efforthighPartially comparable-9.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
148GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 4 rows14.4%IndependentreasoningonPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
148apriel-v1-6-15b-thinkerOpen weightsServiceNow · best of 2 rows14.4%IndependentreasoningonPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
150Gemini 2.5 FlashClosedGoogle · Gemini 2.5 · best of 2 rows13.6%IndependentreasoningonPartially comparable-11.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
150Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 4 rows13.6%IndependentreasoningonPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
150gpt-4.1ClosedOpenAI · GPT 4.1 · best of 2 rows13.6%IndependentreasoningoffPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
153Gemini 2.5 Flash-LiteClosedGoogle · Gemini 2.5 · best of 5 rows12.9%IndependentreasoningonPartially comparable-12.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
153Magistral Medium 1.2ClosedMistral AI · Magistral · best of 2 rows12.9%IndependentreasoningonPartially comparable-12.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
153gemini-2-5-flash-lite-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 3 rows12.9%IndependentreasoningonPartially comparable-12.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
153gpt-5-chatgptClosedOpenAI · GPT 5 · best of 2 rows12.9%IndependentreasoningoffPartially comparable-12.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
153o1ClosedOpenAI · OpenAI o-series · best of 2 rows12.9%IndependentreasoningonPartially comparable-12.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
158hyperclova-x-seed-think-32bOpen weightsNaver · Seed · best of 2 rows12.1%IndependentreasoningonPartially comparable-12.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
159Hermes 4 405BOpen weightsNous Research · Hermes 411.4%Independentgroup defaultsPartially comparable-13.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
159Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AI · Kimi · best of 2 rows11.4%IndependentreasoningoffPartially comparable-13.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
159Qwen3 VL 235B A22B InstructOpen weightsQwen · Qwen3 · best of 4 rows11.4%IndependentreasoningonPartially comparable-13.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
159grok-3ClosedSpaceXAI · Grok 3 · best of 2 rows11.4%IndependentreasoningoffPartially comparable-13.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
159hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.1 · best of 3 rows11.4%IndependentreasoningonPartially comparable-13.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
164Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows10.6%IndependentreasoningoffPartially comparable-14.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
164apriel-v1-5-15b-thinkerOpen weightsServiceNow · best of 2 rows10.6%IndependentreasoningonPartially comparable-14.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
164gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 4 rows10.6%Independentreasoning_efforthighPartially comparable-14.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
164ling-1tOpen weightsinclusionAI · best of 2 rows10.6%IndependentreasoningoffPartially comparable-14.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
164ling-flash-2-0Open weightsinclusionAI · best of 2 rows10.6%IndependentreasoningoffPartially comparable-14.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
164longcat-flash-liteOpen weightsLongCat · best of 2 rows10.6%IndependentreasoningoffPartially comparable-14.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
170Qwen3 Next 80B A3B InstructOpen weightsQwen · Qwen3 · best of 4 rows9.85%IndependentreasoningonPartially comparable-15.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
170k2-v2Open weightsMBZUAI Institute of Foundation Models · best of 6 rows9.85%Independentreasoning_efforthighPartially comparable-15.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
172devstral-mediumClosedMistral AI · Devstral · best of 2 rows9.09%IndependentreasoningoffPartially comparable-15.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
172intellect-3Open weightsPrime Intellect · best of 2 rows9.09%IndependentreasoningonPartially comparable-15.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
172kat-coder-pro-v1ClosedKwaiKAT · best of 2 rows9.09%IndependentreasoningoffPartially comparable-15.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
172magistral-mediumClosedMistral AI · Magistral · best of 2 rows9.09%IndependentreasoningonPartially comparable-15.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176GPT-4o (2024-08-06)ClosedOpenAI · GPT 4 · best of 2 rows8.33%IndependentreasoningoffPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176Qwen3 VL 32B InstructOpen weightsQwen · Qwen3 · best of 4 rows8.33%IndependentreasoningoffPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 4 rows8.33%IndependentreasoningonPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176gpt-4oClosedOpenAI · GPT 4 · best of 2 rows8.33%IndependentreasoningoffPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows8.33%IndependentreasoningonPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176qwen3-5-omni-flashClosedAlibaba Group · Qwen3.5 · best of 2 rows8.33%IndependentreasoningoffPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
182Mistral Small 3.1Open weightsMistral AI · Mistral · best of 2 rows7.58%IndependentreasoningoffPartially comparable-17.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
182Solar Pro 3ClosedUpstage · Solar · best of 2 rows7.58%IndependentreasoningonPartially comparable-17.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
182gpt-4.1-miniClosedOpenAI · GPT 4.1 · best of 2 rows7.58%IndependentreasoningoffPartially comparable-17.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
182ring-flash-2-0Open weightsinclusionAI · best of 2 rows7.58%IndependentreasoningonPartially comparable-17.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 4 rows6.82%IndependentreasoningoffPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186Llama 4 MaverickOpen weightsMeta AI · Llama 4 · best of 2 rows6.82%IndependentreasoningoffPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186Llama-3.1-405BRestricted weightsMeta AI · Llama 3.1 · best of 2 rows6.82%IndependentreasoningoffPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows6.82%IndependentreasoningoffPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186Seed-OSS-36B-InstructOpen weightsByteDance · Seed · best of 2 rows6.82%IndependentreasoningonPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186k2-think-v2Open weightsMBZUAI Institute of Foundation Models · best of 2 rows6.82%IndependentreasoningonPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186nova-2-0-omniClosedAmazon Web Services · Nova 2.0 · best of 6 rows6.82%IndependentreasoningoffPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186nova-premierClosedAmazon Web Services · Nova · best of 2 rows6.82%IndependentreasoningoffPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows6.82%IndependentreasoningonPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186o3-miniClosedOpenAI · OpenAI o-series · best of 4 rows6.82%IndependentreasoningonPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186qwen3-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows6.82%IndependentreasoningoffPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186ring-1tOpen weightsinclusionAI · best of 2 rows6.82%IndependentreasoningonPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
198Devstral Small 1.0Open weightsMistral AI · Devstral · best of 2 rows6.06%IndependentreasoningoffPartially comparable-18.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
198Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 4 rows6.06%IndependentreasoningoffPartially comparable-18.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
198deepseek-r1-0120Open weightsDeepSeek · DeepSeek · best of 2 rows6.06%IndependentreasoningonPartially comparable-18.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →