Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · terminal-bench · variant 1.0

Terminal-Bench

tbench.ai

terminal tasks solved by agents

data quality57

Updated 2 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,639
Models
376
Current leader
gpt-5.6-sol 65.9%

Frontier over time · accuracy · variant=v2.1 · evaluator=Artificial Analysis

5 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 91.4%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  2. 91.0%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  3. 89.9%gpt-6-astra OpenAI Independent11 Sept 2026
  4. 89.5%gpt-6-astra OpenAI Independent11 Sept 2026
  5. 86.1%Claude Opus 5 Anthropic Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 183 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
101North Mini Code (free)Open weightsCohere · best of 2 rows35.6%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
102gpt-5ClosedOpenAI · GPT 5 · best of 2 rows35.2%Independentreasoning_efforthighPartially comparable-0.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
103gpt-5-5-instant-06-26ClosedOpenAI · GPT 5.5 · best of 2 rows34.8%IndependentreasoningonPartially comparable-0.75 ptobs. 12 Sept 2026artificialanalysis.aiT2History
104g9v3-39a5bOpen weightsAI9Stars · best of 2 rows32.6%IndependentreasoningonPartially comparable-3.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
105Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.1 · best of 2 rows31.1%IndependentreasoningonPartially comparable-4.49 ptobs. 12 Sept 2026artificialanalysis.aiT2History
106Devstral 2Open weightsMistral AI · Devstral 2 · best of 2 rows30.3%IndependentreasoningoffPartially comparable-5.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
106k-exaoneOpen weightsLG AI Research · EXAONE · best of 2 rows30.3%IndependentreasoningonPartially comparable-5.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
108Devstral Small 2Open weightsMistral AI · Devstral · best of 2 rows29.6%IndependentreasoningoffPartially comparable-5.99 ptobs. 12 Sept 2026artificialanalysis.aiT2History
108nova-2-0-proClosedAmazon Web Services · Nova 2.0 · best of 6 rows29.6%Independentreasoningonreasoning_effortmediumPartially comparable-5.99 ptobs. 12 Sept 2026artificialanalysis.aiT2History
110Qwen3.5-9BOpen weightsQwen · Qwen3.5 · best of 4 rows29.2%IndependentreasoningonPartially comparable-6.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
111Gemini 2.5 ProClosedGoogle · Gemini 2.5 · best of 2 rows28.5%IndependentreasoningonPartially comparable-7.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
112ling-3-0-tinyOpen weightsinclusionAI · best of 2 rows27.7%IndependentreasoningonPartially comparable-7.86 ptobs. 12 Sept 2026artificialanalysis.aiT2History
113Mercury 2ClosedInception · best of 2 rows27.3%IndependentreasoningonPartially comparable-8.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
113gemma-4-12BOpen weightsGoogle · Gemma 4 · best of 2 rows27.3%IndependentreasoningonPartially comparable-8.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
115granite-4.2-30bOpen weightsIBM · Granite 4.2 · best of 2 rows26.6%IndependentreasoningonPartially comparable-8.99 ptobs. 12 Sept 2026artificialanalysis.aiT2History
116Mistral Small 3.1Open weightsMistral AI · Mistral · best of 2 rows26.2%IndependentreasoningoffPartially comparable-9.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
116gpt-oss-120bOpen weightsOpenAI · gpt-oss · best of 4 rows26.2%Independentreasoning_efforthighPartially comparable-9.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
118Qwen3.5-4BOpen weightsQwen · Qwen3.5 · best of 4 rows25.8%IndependentreasoningonPartially comparable-9.74 ptobs. 12 Sept 2026artificialanalysis.aiT2History
119NVIDIA Nemotron 3.5 Lightning 30B A3BOpen weightsNVIDIA · Nemotron 3.5 · best of 2 rows24.3%IndependentreasoningonPartially comparable-11.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
119ling-2-6-flashOpen weightsinclusionAI · best of 2 rows24.3%IndependentreasoningoffPartially comparable-11.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
121command-a-plusOpen weightsCohere · Command · best of 2 rows22.9%IndependentreasoningonPartially comparable-12.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
122exaone-4-5-33bOpen weightsLG AI Research · EXAONE 4.5 · best of 2 rows21.4%IndependentreasoningonPartially comparable-14.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
123Mistral Small 4Open weightsMistral AI · Mistral · best of 2 rows21.0%IndependentreasoningonPartially comparable-14.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124Trinity Large ThinkingOpen weightsArcee AI · best of 2 rows20.6%IndependentreasoningonPartially comparable-15.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124nemotron-cascade-2-30b-a3bOpen weightsNVIDIA · Nemotron · best of 2 rows20.6%IndependentreasoningonPartially comparable-15.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
126deepseek-r1-0120Open weightsDeepSeek · DeepSeek · best of 2 rows19.1%IndependentreasoningonPartially comparable-16.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127Granite 4.2 8BOpen weightsIBM · Granite 4.2 · best of 2 rows18.4%IndependentreasoningonPartially comparable-17.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127hypernova-60bOpen weightsMultiverse Computing · best of 2 rows18.4%Independentreasoning_efforthighPartially comparable-17.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
129nova-2-0-liteClosedAmazon Web Services · Nova 2.0 · best of 2 rows16.1%Independentreasoningonreasoning_efforthighPartially comparable-19.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
130k2-think-v2Open weightsMBZUAI Institute of Foundation Models · best of 2 rows15.0%IndependentreasoningonPartially comparable-20.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows13.9%IndependentreasoningoffPartially comparable-21.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows13.9%IndependentreasoningoffPartially comparable-21.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 2 rows13.9%Independentreasoning_efforthighPartially comparable-21.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131granite-4.2-3bOpen weightsIBM · Granite 4.2 · best of 2 rows13.9%IndependentreasoningonPartially comparable-21.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
135Magistral Medium 1.2ClosedMistral AI · Magistral · best of 2 rows12.4%IndependentreasoningonPartially comparable-23.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
135diffusiongemma-26b-a4bOpen weightsGoogle · best of 2 rows12.4%IndependentreasoningonPartially comparable-23.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
137Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows12.0%IndependentreasoningoffPartially comparable-23.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
137Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 2 rows12.0%IndependentreasoningonPartially comparable-23.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
137Solar Pro 3ClosedUpstage · Solar · best of 2 rows12.0%IndependentreasoningonPartially comparable-23.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
140celeris-1ClosedCeleris · best of 2 rows11.2%IndependentreasoningoffPartially comparable-24.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
141Claude Haiku 3.5ClosedAnthropic · Claude · best of 2 rows10.1%IndependentreasoningoffPartially comparable-25.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
141gpt-4.1-miniClosedOpenAI · GPT 4.1 · best of 2 rows10.1%IndependentreasoningoffPartially comparable-25.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
143Ministral 3 14BOpen weightsMistral AI · Ministral 3 · best of 2 rows9.74%IndependentreasoningoffPartially comparable-25.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
144minicpm5-2bOpen weightsOpenBMB · best of 2 rows8.61%IndependentreasoningonPartially comparable-27.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
145Llama 4 MaverickOpen weightsMeta AI · Llama 4 · best of 2 rows7.87%IndependentreasoningoffPartially comparable-27.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
146Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 2 rows6.74%IndependentreasoningonPartially comparable-28.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
146Qwen3 Next 80B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows6.74%IndependentreasoningonPartially comparable-28.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
146nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows6.74%IndependentreasoningonPartially comparable-28.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
149g9v3-3bOpen weightsAI9Stars · best of 2 rows5.99%IndependentreasoningonPartially comparable-29.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
150Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows5.62%IndependentreasoningoffPartially comparable-30.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
150gpt-4o-miniClosedOpenAI · GPT 4 · best of 2 rows5.62%IndependentreasoningoffPartially comparable-30.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
152Qwen3 32BOpen weightsQwen · Qwen35.24%Independentgroup defaultsPartially comparable-30.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
152qwen3-32b-instructOpen weightsAlibaba Group · Qwen35.24%IndependentreasoningonPartially comparable-30.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
154Llama 3.3 70BOpen weightsMeta AI · Llama 3.3 · best of 2 rows4.87%IndependentreasoningoffPartially comparable-30.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
154Qwen3 14BOpen weightsQwen · Qwen34.87%Independentgroup defaultsPartially comparable-30.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
154qwen3-14b-instructOpen weightsAlibaba Group · Qwen34.87%IndependentreasoningonPartially comparable-30.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
157Gemma 3 27BOpen weightsGoogle · Gemma 3 · best of 2 rows4.49%IndependentreasoningoffPartially comparable-31.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
157LFM2.5-2.6B (free)Open weightsLiquid AI · LFM2.5 · best of 2 rows4.49%IndependentreasoningonPartially comparable-31.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
157Magistral Small 1.2Open weightsMistral AI · Magistral · best of 2 rows4.49%IndependentreasoningonPartially comparable-31.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
157o3-miniClosedOpenAI · OpenAI o-series · best of 2 rows4.49%Independentreasoning_efforthighPartially comparable-31.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
161Ministral 3 8BOpen weightsMistral AI · Ministral 3 · best of 2 rows4.12%IndependentreasoningoffPartially comparable-31.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
162Llama 4 ScoutOpen weightsMeta AI · Llama 4 · best of 2 rows3.75%IndependentreasoningoffPartially comparable-31.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
162gpt-4.1-nanoClosedOpenAI · GPT 4.1 · best of 2 rows3.75%IndependentreasoningoffPartially comparable-31.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
162gpt-5-miniClosedOpenAI · GPT 5 · best of 2 rows3.75%Independentreasoning_efforthighPartially comparable-31.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
162nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows3.75%IndependentreasoningonPartially comparable-31.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
166granite-4.1-8bOpen weightsIBM · Granite 4.1 · best of 2 rows3.37%IndependentreasoningoffPartially comparable-32.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
167Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 4 rows3%IndependentreasoningonPartially comparable-32.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
168granite-4.1-30bOpen weightsIBM · Granite 4.1 · best of 2 rows2.62%IndependentreasoningoffPartially comparable-33.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
169Qwen3 8BOpen weightsQwen · Qwen32.25%Independentgroup defaultsPartially comparable-33.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
169qwen3-8b-instructOpen weightsAlibaba Group · Qwen32.25%IndependentreasoningonPartially comparable-33.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
171gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 2 rows1.87%IndependentreasoningonPartially comparable-33.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
172Llama 3.1 8BRestricted weightsMeta AI · Llama 3.1 · best of 2 rows1.50%IndependentreasoningoffPartially comparable-34.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
172qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 2 rows1.50%IndependentreasoningonPartially comparable-34.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
174granite-4.1-3bOpen weightsIBM · Granite 4.1 · best of 2 rows1.12%IndependentreasoningoffPartially comparable-34.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
174nanbeige4-1-3bOpen weightsNanbeige · best of 2 rows1.12%IndependentreasoningonPartially comparable-34.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176gemma-3n-e4bOpen weightsGoogle · Gemma 3 · best of 2 rows0.75%IndependentreasoningoffPartially comparable-34.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
177Gemma 3 4BRestricted weightsGoogle · Gemma 3 · best of 2 rows0.37%IndependentreasoningoffPartially comparable-35.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
177Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 4 rows0.37%IndependentreasoningoffPartially comparable-35.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
177gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 2 rows0.37%IndependentreasoningonPartially comparable-35.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
177phi-4-miniOpen weightsMicrosoft · Phi4 · best of 2 rows0.37%IndependentreasoningoffPartially comparable-35.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
181Gemma 3 12BOpen weightsGoogle · Gemma 3 · best of 2 rows0%IndependentreasoningoffPartially comparable-35.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
181Ministral 3 3BOpen weightsMistral AI · Ministral 3 · best of 2 rows0%IndependentreasoningoffPartially comparable-35.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
181minicpm-v4-6-1-3bOpen weightsOpenBMB · best of 2 rows0%IndependentreasoningoffPartially comparable-35.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →