Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · terminal-bench · variant 1.0

Terminal-Bench

tbench.ai

terminal tasks solved by agents

quality57

Updated 14 min ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
821
Models
374
Current leader
gpt-5.6-sol 65.9%

Score history · NVIDIA-Nemotron-Nano-9B-v2 2 rows

Not enough history to chart — 2 observations, all dated 11 Sept 2026. Rows under different configurations count separately; the list below shows each one.

  • 0.76%aa_slug=nvidia-nemotron-nano-9b-v2 · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 1.52%aa_slug=nvidia-nemotron-nano-9b-v2-reasoning · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=hard · evaluator=Artificial Analysis

10 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 65.9%gpt-5.6-sol OpenAI Independent11 Sept 2026
  2. 61.4%gpt-5.6-sol OpenAI Independent11 Sept 2026
  3. 53.0%Claude Sonnet 4.6 Anthropic Independent11 Sept 2026
  4. 51.5%Claude Opus 4.7 Anthropic Independent11 Sept 2026
  5. 50.8%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
  6. 49.2%KAT-Coder-Pro V2 Kwaipilot Independent11 Sept 2026
  7. 33.3%GPT-5.1-Codex Mini OpenAI Independent11 Sept 2026
  8. 26.5%Grok 4.3 xAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 315 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
196mistral-large-2Open weightsMistral AI · Mistral6.06%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
196nova-proClosedAmazon Web Services · Nova6.06%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
196qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows6.06%IndependentreasoningonPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
196qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 2 rows6.06%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
205llama-nemotron-super-49b-v1-5Open weightsNVIDIA · Llama · best of 2 rows5.30%IndependentreasoningonPartially comparable-0.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
205qwen3-14b-instructOpen weightsAlibaba Group · Qwen35.30%Independentgroup defaultsComparable-0.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
205step-3-vl-10bOpen weightsStepFun · Step35.30%Independentgroup defaultsComparable-0.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208Gemini 2.5 Flash-LiteClosedGoogle · Gemini · best of 2 rows4.55%Independentgroup defaultsComparable-1.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208Hermes-4-70BRestricted weightsNous Research · Hermes 44.55%Independentgroup defaultsComparable-1.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208LFM2.5-8B-A1BOpen weightsLiquid AI · LFM2.54.55%Independentgroup defaultsComparable-1.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208Magistral Small 1.2Open weightsMistral AI · Magistral4.55%Independentgroup defaultsComparable-1.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208Ministral 3 14BOpen weightsMistral AI · Ministral 34.55%Independentgroup defaultsComparable-1.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208Ministral 3 8BOpen weightsMistral AI · Ministral 34.55%Independentgroup defaultsComparable-1.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208Qwen2.5 72B InstructOpen weightsQwen · Qwen2.54.55%Independentgroup defaultsComparable-1.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.14.55%Independentgroup defaultsComparable-1.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208magistral-smallOpen weightsMistral AI · Magistral4.55%Independentgroup defaultsComparable-1.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 2 rows4.55%IndependentreasoningonPartially comparable-1.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208qwen3-4b-2507-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows4.55%Independentgroup defaultsComparable-1.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208solar-pro-2ClosedUpstage · Solar · best of 2 rows4.55%Independentgroup defaultsComparable-1.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220Gemini 2.0 FlashClosedGoogle · Gemini3.79%Independentgroup defaultsComparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220Gemma 3 27BOpen weightsGoogle · Gemma 33.79%Independentgroup defaultsComparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220Mistral Medium 3ClosedMistral AI · Mistral3.79%Independentgroup defaultsComparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220Phi 4Open weightsMicrosoft · Phi43.79%Independentgroup defaultsComparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220Qwen3 14BOpen weightsQwen · Qwen33.79%Independentgroup defaultsComparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 2 rows3.79%IndependentreasoningonPartially comparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 2 rows3.79%Independentgroup defaultsComparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows3.79%IndependentreasoningonPartially comparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220gpt-4.1-nanoClosedOpenAI · GPT 4.13.79%Independentgroup defaultsComparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220motif-2-12-7bClosedMotif Technologies3.79%Independentgroup defaultsComparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows3.79%IndependentreasoningonPartially comparable-2.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
231Llama 3.3 70BOpen weightsMeta AI · Llama 3.33.03%Independentgroup defaultsComparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
231Llama-3.1-70BOpen weightsMeta AI · Llama 3.13.03%Independentgroup defaultsComparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
231MiniMax-M1-80kOpen weightsMiniMax · MiniMax3.03%Independentgroup defaultsComparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
231Qwen3 32BOpen weightsQwen · Qwen33.03%Independentgroup defaultsComparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
231gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 2 rows3.03%Independentgroup defaultsComparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
231midm-250-pro-rsnsftClosedKorea Telecom3.03%Independentgroup defaultsComparable-3.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237Claude Haiku 3.5ClosedAnthropic · Claude2.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237MiniMax-M1-40kOpen weightsMiniMax · MiniMax2.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237Qwen3 8BOpen weightsQwen · Qwen32.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237falcon-h1r-7bOpen weightsTII UAE · Falcon2.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237gemma-3n-e4bOpen weightsGoogle · Gemma 32.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237granite-4-0-h-smallOpen weightsIBM · Granite 4.02.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237granite-4.1-30bOpen weightsIBM · Granite 4.12.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237granite-4.1-3bOpen weightsIBM · Granite 4.12.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.72.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.12.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237mi-dm-k-2-5-pro-dec28ClosedKorea Telecom2.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237qwen3-8b-instructOpen weightsAlibaba Group · Qwen32.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237sarvam-30bOpen weightsSarvam2.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237sarvam-m-reasoningOpen weightsSarvam2.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237solar-open-100b-reasoningOpen weightsUpstage · Solar2.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237tri-21b-think-previewOpen weightsTrillion Labs2.27%Independentgroup defaultsComparable-3.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253DeepSeek-R1-0528-Qwen3-8BOpen weightsDeepSeek · Qwen31.52%Independentgroup defaultsComparable-4.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253Granite 4.0 MicroOpen weightsIBM · Granite 4.01.52%Independentgroup defaultsComparable-4.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253Llama 4 ScoutOpen weightsMeta AI · Llama 41.52%Independentgroup defaultsComparable-4.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 2 rows1.52%Independentgroup defaultsComparable-4.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 2 rows1.52%IndependentreasoningonPartially comparable-4.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253R1 Distill Llama 70BOpen weightsDeepSeek · Llama1.52%Independentgroup defaultsComparable-4.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253nova-microClosedAmazon Web Services · Nova1.52%Independentgroup defaultsComparable-4.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253olmo-3-32b-thinkOpen weightsAllen Institute for AI · OLMo 31.52%Independentgroup defaultsComparable-4.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253sarvam-105bOpen weightsSarvam1.52%Independentgroup defaultsComparable-4.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262Claude 3 HaikuClosedAnthropic · Claude0.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262Command AOpen weightsCohere · Command0.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262Gemma 3 12BOpen weightsGoogle · Gemma 30.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262Gemma 3 4BRestricted weightsGoogle · Gemma 30.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262Llama 3.1 8BRestricted weightsMeta AI · Llama 3.10.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262Olmo-3-7B-ThinkOpen weightsAllen Institute for AI · OLMo 30.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262gemma-3n-e2bOpen weightsGoogle · Gemma 30.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262jamba-reasoning-3bOpen weightsAI21 Labs · Jamba0.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262lfm2-2-6bOpen weightsLiquid AI · LFM2.20.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262ling-mini-2-0Open weightsinclusionAI0.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.20.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262llama-3-instruct-70bOpen weightsMeta AI · Llama 30.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262nova-liteClosedAmazon Web Services · Nova0.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262tri-21b-think-v0-5Open weightsTrillion Labs0.76%Independentgroup defaultsComparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276LFM2-1.2BOpen weightsLiquid AI · LFM2.10%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276LFM2.5-1.2B-InstructOpen weightsLiquid AI · LFM2.50%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.50%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276Llama-3.2-1BRestricted weightsMeta AI · Llama 3.20%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276Ministral 3 3BOpen weightsMistral AI · Ministral 30%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276Molmo2-8BOpen weightsAllen Institute for AI0%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276Olmo-3-7B-InstructOpen weightsAllen Institute for AI · OLMo 30%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276Qwen3-0.6BOpen weightsQwen · Qwen3.00%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 2 rows0%IndependentreasoningoffPartially comparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276Reka Flash 3Open weightsrekaai0%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276apertus-70b-instructOpen weightsSwiss AI Initiative0%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276apertus-8b-instructOpen weightsSwiss AI Initiative0%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276exaone-4-0-1-2bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows0%IndependentreasoningonPartially comparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276gemma-3-1bOpen weightsGoogle · Gemma 30%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276gemma-3-270mOpen weightsGoogle · Gemma 30%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276granite-3-3-8b-instructOpen weightsIBM · Granite 3.30%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276granite-4-0-350mOpen weightsIBM · Granite 4.00%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276granite-4-0-h-350mOpen weightsIBM · Granite 4.00%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.00%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276granite-4-0-nano-1bOpen weightsIBM · Granite 4.00%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276granite-4.1-8bOpen weightsIBM · Granite 4.10%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.10%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276jamba-1-7-miniOpen weightsAI21 Labs · Jamba 1.70%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276lfm2-24b-a2bOpen weightsLiquid AI · LFM20%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276lfm2-5-1-2b-thinkingOpen weightsLiquid AI · LFM2.50%Independentgroup defaultsComparable-6.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →