Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · terminal-bench · variant 1.0

Terminal-Bench

tbench.ai

terminal tasks solved by agents

data quality57

Updated 53 min ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,639
Models
376
Current leader
gpt-6-astra 59.6%

Score history · Step 3.5 Flash 4 rows

Score history for Step 3.5 Flash0%10%20%30%40%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • Step 3.5 Flash
  • 32.58%aa_slug=step-3-5-flash · variant=hard · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
  • 27.27%aa_slug=step-3-5-flash-0202 · variant=hard · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
  • 32.58%aa_slug=step-3-5-flash · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 27.27%aa_slug=step-3-5-flash-0202 · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=v4.0 · evaluator=Artificial Analysis

6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 59.6%gpt-6-astra OpenAI Independent11 Sept 2026
  2. 59.1%gpt-6-astra OpenAI Independent11 Sept 2026
  3. 55.0%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  4. 54.0%gpt-6-astra OpenAI Independent11 Sept 2026
  5. 49.5%gpt-6-astra OpenAI Independent11 Sept 2026
  6. 34.3%Claude Opus 5 Anthropic Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 114 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1gpt-6-astraClosedOpenAI · GPT 6 · best of 11 rows59.6%Independentreasoning_effortxhighleaderobs. 12 Sept 2026artificialanalysis.aiT2History
2Claude Fable 5.1ClosedAnthropic · Claude · best of 10 rows55.0%Independentreasoning_effortxhighComparable-4.55 ptobs. 12 Sept 2026artificialanalysis.aiT2History
3Claude Opus 5ClosedAnthropic · Claude · best of 10 rows49.0%Independentreasoning_effortmaxPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
4Claude Fable 5ClosedAnthropic · Claude · best of 2 rows42.4%IndependentreasoningonPartially comparable-17.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
5GLM 5.3Open weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows41.9%Independentreasoning_effortmaxPartially comparable-17.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
6gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 10 rows39.9%Independentreasoning_effortmaxPartially comparable-19.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
7gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 10 rows35.4%Independentreasoning_effortmaxPartially comparable-24.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
8Muse Spark 1.3ClosedMeta AI · best of 4 rows33.3%Independentreasoning_effortmaxPartially comparable-26.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
9GLM 5.3 FlashOpen weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows32.8%IndependentreasoningonPartially comparable-26.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
10DeepSeek-V4.1-FlashOpen weightsDeepSeek · DeepSeek · best of 2 rows26.8%Independentreasoning_effortmaxPartially comparable-32.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
11Qwen3.8 FlashOpen weightsQwen · Qwen3.8 · best of 2 rows25.3%IndependentreasoningonPartially comparable-34.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
12Claude Opus 4.8ClosedAnthropic · Claude · best of 2 rows21.7%Independentreasoning_effortmaxPartially comparable-37.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
13Grok 4.6ClosedxAI · Grok · best of 8 rows21.2%Independentreasoning_efforthighPartially comparable-38.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
14Gemini 3.8 FlashClosedGoogle · Gemini 3.8 · best of 6 rows19.7%Independentreasoning_efforthighPartially comparable-39.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
15Qwen 3.8 MaxClosedQwen · Qwen3.8 · best of 2 rows18.7%IndependentreasoningonPartially comparable-40.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
16deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek · best of 2 rows14.7%Independentreasoning_effortmaxPartially comparable-45.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
16gpt-5.5ClosedOpenAI · GPT 5.5 · best of 6 rows14.7%Independentreasoning_effortxhighComparable-45.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
18Claude Sonnet 5ClosedAnthropic · Claude · best of 10 rows14.1%Independentreasoning_effortmaxPartially comparable-45.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
18deepseek-v4-proClosedDeepSeek · V4 · best of 2 rows14.1%Independentreasoning_effortmaxPartially comparable-45.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
20Gemini 3.7 FlashClosedGoogle · Gemini 3.7 · best of 2 rows13.6%Independentreasoning_efforthighPartially comparable-46.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
21Kimi K3Open weightsMoonshot AI · Kimi · best of 3 rows12.6%Independentreasoning_effortmaxPartially comparable-47.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
21gpt-5-5-instant-06-26ClosedOpenAI · GPT 5.5 · best of 2 rows12.6%IndependentreasoningonPartially comparable-47.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
23deepseek-v4-flashOpen weightsDeepSeek · DeepSeek · best of 2 rows12.1%Independentreasoning_effortmaxPartially comparable-47.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
23deepseek-v4-flash-visionClosedDeepSeek · DeepSeek · best of 2 rows12.1%Independentreasoning_effortmaxPartially comparable-47.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
25gpt-5.6-lunaClosedOpenAI · GPT 5.6 · best of 10 rows11.6%Independentreasoning_effortmaxPartially comparable-48.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
26Qwen3.8 2.4T A95BOpen weightsQwen · Qwen3.8 · best of 2 rows11.1%IndependentreasoningonPartially comparable-48.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
27Grok 4.5ClosedxAI · Grok · best of 2 rows10.6%Independentreasoning_efforthighPartially comparable-49.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
28Gemini 3.6 FlashClosedGoogle · Gemini 3.6 · best of 2 rows7.07%IndependentreasoningonPartially comparable-52.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
28Muse Spark 1.2ClosedMeta AI · best of 2 rows7.07%Independentreasoning_effortxhighComparable-52.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
28agnes-3-0-flashClosedSapiens AI · best of 2 rows7.07%IndependentreasoningonPartially comparable-52.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
31Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 2 rows6.57%IndependentreasoningonPartially comparable-53.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
32Muse Spark 1.1ClosedMeta AI · best of 2 rows6.06%Independentreasoning_effortxhighComparable-53.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
33Qwen3.8 27BOpen weightsQwen · Qwen3.8 · best of 6 rows5.56%Independentreasoning_effortxhighComparable-54.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
34Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.1 · best of 2 rows4.04%IndependentreasoningonPartially comparable-55.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35Claude Sonnet 4.6ClosedAnthropic · Claude · best of 2 rows3.03%Independentreasoningadaptivereasoning_effortmaxPartially comparable-56.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35agnes-2-5-pro-alphaOpen weightsSapiens AI · best of 2 rows3.03%IndependentreasoningonPartially comparable-56.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek · best of 3 rows3.03%Independentreasoning_efforthighPartially comparable-56.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35deepseek-v4-flash-0420-highOpen weightsDeepSeek · DeepSeek3.03%Independentgroup defaultsPartially comparable-56.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
39GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 2 rows2.02%IndependentreasoningonPartially comparable-57.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
39MiniMax-M3Open weightsMiniMax · MiniMax · best of 2 rows2.02%IndependentreasoningonPartially comparable-57.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
39gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 2 rows2.02%Independentreasoning_effortxhighComparable-57.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
42Qwen3.7 MaxClosedQwen · Qwen3.7 · best of 2 rows1.52%IndependentreasoningonPartially comparable-58.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
43Gemini 3.5 Flash-LiteClosedGoogle · Gemini 3.5 · best of 2 rows1.01%IndependentreasoningonPartially comparable-58.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
43InklingOpen weightsThinking Machines · best of 2 rows1.01%IndependentreasoningonPartially comparable-58.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
43Inkling SmallOpen weightsThinking Machines · best of 2 rows1.01%IndependentreasoningonPartially comparable-58.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
43Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi · best of 2 rows1.01%IndependentreasoningonPartially comparable-58.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
43Qwen 3.7 PlusClosedQwen · Qwen3.7 · best of 2 rows1.01%IndependentreasoningonPartially comparable-58.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
43quasar-438bClosedMultiverse Computing · best of 2 rows1.01%Independentreasoning_effortmaxPartially comparable-58.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.1 · best of 2 rows0.51%IndependentreasoningonPartially comparable-59.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49Hy3Open weightsTencent · best of 2 rows0.51%IndependentreasoningonPartially comparable-59.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49Muse Glimmer 30BOpen weightsMeta AI · best of 2 rows0.51%Independentreasoning_efforthighPartially comparable-59.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49NVIDIA Nemotron 3.5 Lightning 30B A3BOpen weightsNVIDIA · Nemotron 3.5 · best of 2 rows0.51%IndependentreasoningonPartially comparable-59.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 3 · best of 2 rows0.51%IndependentreasoningonPartially comparable-59.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49North Mini Code (free)Open weightsCohere · best of 2 rows0.51%IndependentreasoningonPartially comparable-59.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49Solar Pro 4ClosedUpstage · Solar · best of 2 rows0.51%IndependentreasoningonPartially comparable-59.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49Trinity Large ThinkingOpen weightsArcee AI · best of 2 rows0.51%IndependentreasoningonPartially comparable-59.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49command-a-plusOpen weightsCohere · Command · best of 2 rows0.51%IndependentreasoningonPartially comparable-59.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 2 rows0.51%Independentreasoning_effortxhighComparable-59.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49ring-2-6-1tOpen weightsinclusionAI · best of 2 rows0.51%IndependentreasoningonPartially comparable-59.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Claude Haiku 4.5ClosedAnthropic · Claude · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Claude Sonnet 4.5ClosedAnthropic · Claude · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Devstral 2Open weightsMistral AI · Devstral 2 · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Devstral Small 2Open weightsMistral AI · Devstral · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Gemini 2.5 ProClosedGoogle · Gemini 2.5 · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Gemma 3 12BOpen weightsGoogle · Gemma 3 · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Gemma 3 27BOpen weightsGoogle · Gemma 3 · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Granite 4.2 8BOpen weightsIBM · Granite 4.2 · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Grok 4.3ClosedxAI · Grok · best of 4 rows0%Independentreasoning_efforthighPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Ling 3.0 Flash VLOpen weightsinclusionAI · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Llama 4 MaverickOpen weightsMeta AI · Llama 4 · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Llama 4 ScoutOpen weightsMeta AI · Llama 4 · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60LongCat 2.0Open weightsMeituan · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Mercury 2ClosedInception · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60MiMo-V2.5Open weightsXiaomi · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60MiMo-V2.5-ProOpen weightsXiaomi · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60MiniMax M2.7Open weightsMiniMax · MiniMax · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Ministral 3 14BOpen weightsMistral AI · Ministral 3 · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Ministral 3 3BOpen weightsMistral AI · Ministral 3 · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Ministral 3 8BOpen weightsMistral AI · Ministral 3 · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Mistral Medium 3.5Open weightsMistral AI · Mistral · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Mistral Small 3.1Open weightsMistral AI · Mistral · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Mistral Small 4Open weightsMistral AI · Mistral · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 3 · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Qwen3 14BOpen weightsQwen · Qwen30%Independentgroup defaultsPartially comparable-59.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
60Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Qwen3 32BOpen weightsQwen · Qwen30%Independentgroup defaultsPartially comparable-59.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
60Qwen3 8BOpen weightsQwen · Qwen30%Independentgroup defaultsPartially comparable-59.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
60Qwen3 Coder NextOpen weightsQwen · Qwen3 · best of 2 rows0%IndependentreasoningoffPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 2 rows0%IndependentreasoningonPartially comparable-59.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →