Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · terminal-bench · variant 1.0

Terminal-Bench

tbench.ai

terminal tasks solved by agents

data quality57

Updated 2 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,639
Models
376
Current leader
Claude Fable 5.1 91.4%

Score history · LongCat 2.0 4 rows

Score history for LongCat 2.00%20%40%60%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • LongCat 2.0
  • 50.19%aa_slug=longcat-2-0 · variant=v2.1 · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
  • 0%aa_slug=longcat-2-0 · variant=v4.0 · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
  • 50.19%aa_slug=longcat-2-0 · variant=v2.1 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 0%aa_slug=longcat-2-0 · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=v2.1 · evaluator=Artificial Analysis

5 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 91.4%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  2. 91.0%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  3. 89.9%gpt-6-astra OpenAI Independent11 Sept 2026
  4. 89.5%gpt-6-astra OpenAI Independent11 Sept 2026
  5. 86.1%Claude Opus 5 Anthropic Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 183 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1Claude Fable 5.1ClosedAnthropic · Claude · best of 10 rows91.4%Independentreasoning_effortmaxleaderobs. 12 Sept 2026artificialanalysis.aiT2History
2gpt-6-astraClosedOpenAI · GPT 6 · best of 11 rows89.9%Independentreasoning_efforthighPartially comparable-1.50 ptobs. 12 Sept 2026artificialanalysis.aiT2History
3gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 12 rows89.5%Independentreasoning_effortxhighPartially comparable-1.88 ptobs. 12 Sept 2026artificialanalysis.aiT2History
4Claude Opus 5ClosedAnthropic · Claude · best of 10 rows89.1%Independentreasoning_effortmaxComparable-2.25 ptobs. 12 Sept 2026artificialanalysis.aiT2History
5Grok 4.6ClosedxAI · Grok · best of 8 rows88.4%Independentreasoning_efforthighPartially comparable-3.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
6gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 12 rows88.0%Independentreasoning_effortmaxComparable-3.38 ptobs. 12 Sept 2026artificialanalysis.aiT2History
7Gemini 3.8 FlashClosedGoogle · Gemini 3.8 · best of 6 rows87.6%Independentreasoning_efforthighPartially comparable-3.75 ptobs. 12 Sept 2026artificialanalysis.aiT2History
8Qwen3.8 FlashOpen weightsQwen · Qwen3.8 · best of 2 rows86.1%IndependentreasoningonPartially comparable-5.25 ptobs. 12 Sept 2026artificialanalysis.aiT2History
9Gemini 3.7 FlashClosedGoogle · Gemini 3.7 · best of 6 rows85.8%Independentreasoning_efforthighPartially comparable-5.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
10Muse Spark 1.3ClosedMeta AI · best of 4 rows85.4%Independentreasoning_effortxhighPartially comparable-6.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
11Kimi K3Open weightsMoonshot AI · Kimi · best of 4 rows85.0%Independentreasoning_effortmaxComparable-6.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
12Claude Fable 5ClosedAnthropic · Claude · best of 2 rows84.6%IndependentreasoningonPartially comparable-6.75 ptobs. 12 Sept 2026artificialanalysis.aiT2History
12Claude Opus 4.8ClosedAnthropic · Claude · best of 2 rows84.6%Independentreasoning_effortmaxComparable-6.75 ptobs. 12 Sept 2026artificialanalysis.aiT2History
14GLM 5.3 FlashOpen weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows84.3%IndependentreasoningonPartially comparable-7.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
14gpt-5.5ClosedOpenAI · GPT 5.5 · best of 10 rows84.3%Independentreasoning_effortxhighPartially comparable-7.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
16GLM 5.3Open weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows83.9%Independentreasoning_effortmaxComparable-7.49 ptobs. 12 Sept 2026artificialanalysis.aiT2History
17Claude Opus 4.7ClosedAnthropic · Claude · best of 2 rows83.2%Independentreasoning_effortmaxComparable-8.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
18Qwen3.8 2.4T A95BOpen weightsQwen · Qwen3.8 · best of 2 rows82.0%IndependentreasoningonPartially comparable-9.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
18agnes-3-0-flashClosedSapiens AI · best of 2 rows82.0%IndependentreasoningonPartially comparable-9.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
20Grok 4.5ClosedxAI · Grok · best of 2 rows81.7%Independentreasoning_efforthighPartially comparable-9.74 ptobs. 12 Sept 2026artificialanalysis.aiT2History
21Qwen 3.8 MaxClosedQwen · Qwen3.8 · best of 2 rows81.3%IndependentreasoningonPartially comparable-10.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
22gpt-5.6-lunaClosedOpenAI · GPT 5.6 · best of 12 rows80.9%Independentreasoning_effortmaxComparable-10.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
23Claude Sonnet 5ClosedAnthropic · Claude · best of 4 rows80.5%Independentreasoning_effortmaxComparable-10.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
24Muse Spark 1.2ClosedMeta AI · best of 2 rows80.2%Independentreasoning_effortxhighPartially comparable-11.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
25Qwen3.8 27BOpen weightsQwen · Qwen3.8 · best of 8 rows79.8%Independentreasoning_effortxhighPartially comparable-11.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
26Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 2 rows78.7%IndependentreasoningonPartially comparable-12.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
26deepseek-v4-flashOpen weightsDeepSeek · DeepSeek · best of 2 rows78.7%Independentreasoning_effortmaxComparable-12.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
26deepseek-v4-proClosedDeepSeek · V4 · best of 2 rows78.7%Independentreasoning_effortmaxComparable-12.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
29gpt-5.4ClosedOpenAI · GPT 5.4 · best of 2 rows78.3%Independentreasoning_effortxhighPartially comparable-13.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
30Muse Spark 1.1ClosedMeta AI · best of 2 rows77.9%Independentreasoning_effortxhighPartially comparable-13.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
30Z.ai GLM 5.2Open weightsZ.ai (Zhipu AI) · GLM5.2 · best of 4 rows77.9%Independentreasoning_effortmaxComparable-13.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
32Gemini 3.6 FlashClosedGoogle · Gemini 3.6 · best of 2 rows77.5%IndependentreasoningonPartially comparable-13.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
33motif-3Open weightsMotif Technologies · best of 2 rows74.9%IndependentreasoningonPartially comparable-16.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
34Qwen3.7 MaxClosedQwen · Qwen3.7 · best of 2 rows74.5%IndependentreasoningonPartially comparable-16.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35deepseek-v4-flash-visionClosedDeepSeek · DeepSeek · best of 2 rows74.2%Independentreasoning_effortmaxComparable-17.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
36Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.1 · best of 2 rows73.8%IndependentreasoningonPartially comparable-17.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
37k2-horizon-375b-a23bOpen weightsMBZUAI Institute of Foundation Models · best of 2 rows71.9%IndependentreasoningonPartially comparable-19.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
38Claude Sonnet 4.6ClosedAnthropic · Claude · best of 2 rows71.2%Independentreasoningadaptivereasoning_effortmaxPartially comparable-20.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
39motif-0714ClosedMotif Technologies · best of 2 rows70.8%IndependentreasoningonPartially comparable-20.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
40KAT-Coder-Pro V2ClosedKwaipilot · best of 2 rows70.0%IndependentreasoningoffPartially comparable-21.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
41agnes-2-5-pro-betaClosedSapiens AI · best of 2 rows69.7%IndependentreasoningonPartially comparable-21.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
41apodex-1-1ClosedApodex · best of 2 rows69.7%IndependentreasoningonPartially comparable-21.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
43quasar-438bClosedMultiverse Computing · best of 2 rows69.3%Independentreasoning_effortmaxComparable-22.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
44nex-n2-proOpen weightsNex AGI · best of 2 rows67.8%IndependentreasoningonPartially comparable-23.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
45Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi · best of 2 rows67.4%IndependentreasoningonPartially comparable-24.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
46agnes-2-5-pro-alphaOpen weightsSapiens AI · best of 2 rows67.0%IndependentreasoningonPartially comparable-24.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
47Kimi K2.6Open weightsMoonshot AI · Kimi · best of 2 rows65.9%IndependentreasoningonPartially comparable-25.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
48MiMo-V2.5-ProOpen weightsXiaomi · best of 2 rows65.2%IndependentreasoningonPartially comparable-26.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
48MiniMax-M3Open weightsMiniMax · MiniMax · best of 2 rows65.2%IndependentreasoningonPartially comparable-26.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
50deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek · best of 3 rows64.8%Independentreasoning_efforthighPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
50deepseek-v4-pro-0424-highOpen weightsDeepSeek · DeepSeek64.8%Independentgroup defaultsPartially comparable-26.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
52Hy3Open weightsTencent · best of 2 rows64.4%IndependentreasoningonPartially comparable-27.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
52Ling 3.0 Flash VLOpen weightsinclusionAI · best of 2 rows64.4%IndependentreasoningonPartially comparable-27.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
54MiMo-V2.5Open weightsXiaomi · best of 2 rows63.7%IndependentreasoningonPartially comparable-27.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
55muse-sparkClosedMeta AI · best of 2 rows62.2%IndependentreasoningonPartially comparable-29.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
56GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 2 rows61.8%IndependentreasoningonPartially comparable-29.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
56deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek · best of 3 rows61.8%Independentreasoning_effortmaxComparable-29.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
56mimo-v2-flashOpen weightsXiaomi · best of 2 rows61.8%IndependentreasoningoffPartially comparable-29.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
59Qwen3.6 PlusClosedQwen · Qwen3.6 · best of 2 rows61.4%IndependentreasoningonPartially comparable-30.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Qwen 3.7 PlusClosedQwen · Qwen3.7 · best of 2 rows61.0%IndependentreasoningonPartially comparable-30.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
61Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 4 rows60.7%IndependentreasoningonPartially comparable-30.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
61gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 2 rows60.7%Independentreasoning_effortxhighPartially comparable-30.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
63jt-4-1-flash-236b-a21bClosedChina Mobile · best of 2 rows59.5%IndependentreasoningoffPartially comparable-31.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
64gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 2 rows59.2%Independentreasoning_effortxhighPartially comparable-32.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
65Solar Pro 4ClosedUpstage · Solar · best of 2 rows57.3%IndependentreasoningonPartially comparable-34.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
66deepseek-v4-flash-0420-highOpen weightsDeepSeek · DeepSeek56.9%Independentgroup defaultsPartially comparable-34.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
67Claude Sonnet 4.5ClosedAnthropic · Claude · best of 2 rows55.8%IndependentreasoningonPartially comparable-35.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
68Ling 3.0 FlashOpen weightsinclusionAI · best of 2 rows55.4%IndependentreasoningonPartially comparable-36.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
68MiniMax M2.7Open weightsMiniMax · MiniMax · best of 2 rows55.4%IndependentreasoningonPartially comparable-36.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
70InklingOpen weightsThinking Machines · best of 2 rows55.1%IndependentreasoningonPartially comparable-36.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
70Inkling SmallOpen weightsThinking Machines · best of 2 rows55.1%IndependentreasoningonPartially comparable-36.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
72Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 3 · best of 2 rows53.9%IndependentreasoningonPartially comparable-37.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
73Gemini 3.5 Flash-LiteClosedGoogle · Gemini 3.5 · best of 2 rows53.6%IndependentreasoningonPartially comparable-37.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
74gpt-5.1ClosedOpenAI · GPT 5.1 · best of 2 rows52.4%Independentreasoning_efforthighPartially comparable-39.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
75grok-build-0-1-06-16ClosedSpaceXAI · Grok · best of 2 rows52.1%IndependentreasoningonPartially comparable-39.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
76Muse Glimmer 30BOpen weightsMeta AI · best of 2 rows51.7%Independentreasoning_efforthighPartially comparable-39.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
77Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 2 rows51.3%IndependentreasoningonPartially comparable-40.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
78Mistral Medium 3.5Open weightsMistral AI · Mistral · best of 2 rows50.6%IndependentreasoningonPartially comparable-40.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79LongCat 2.0Open weightsMeituan · best of 2 rows50.2%IndependentreasoningonPartially comparable-41.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
80GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows49.4%IndependentreasoningonPartially comparable-42.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
81Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 4 rows47.6%IndependentreasoningonPartially comparable-43.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
82DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 3 rows46.8%IndependentreasoningonPartially comparable-44.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
82DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V346.8%Independentgroup defaultsPartially comparable-44.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
84Kimi K2.5Open weightsMoonshot AI · Kimi · best of 2 rows45.7%IndependentreasoningonPartially comparable-45.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
85GLM 4.7Open weightsZ.ai (Zhipu AI) · GLM4.7 · best of 2 rows45.3%IndependentreasoningonPartially comparable-46.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
86DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 2 rows44.9%IndependentreasoningonPartially comparable-46.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
86Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 4 rows44.9%IndependentreasoningonPartially comparable-46.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88Claude Haiku 4.5ClosedAnthropic · Claude · best of 2 rows44.2%IndependentreasoningonPartially comparable-47.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88solar-open2-250bOpen weightsUpstage · Solar · best of 2 rows44.2%IndependentreasoningonPartially comparable-47.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
90Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 4 rows43.5%IndependentreasoningonPartially comparable-47.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
91ring-2-6-1tOpen weightsinclusionAI · best of 2 rows43.1%IndependentreasoningonPartially comparable-48.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
92Qwen3.5-35B-A3BOpen weightsQwen · Qwen3.5 · best of 2 rows40.8%IndependentreasoningoffPartially comparable-50.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
93k-exaone-2-0-0803Open weightsLG AI Research · EXAONE 2.0 · best of 2 rows40.5%IndependentreasoningonPartially comparable-50.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
94Grok 4.3ClosedxAI · Grok · best of 4 rows39.7%Independentreasoning_efforthighPartially comparable-51.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
95Step 3.7 FlashOpen weightsStepFun · Step3.7 · best of 2 rows39.3%IndependentreasoningonPartially comparable-52.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
96Gemma 4 26B A4BOpen weightsGoogle · Gemma 4 · best of 2 rows39.0%IndependentreasoningonPartially comparable-52.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
96a-x-k2Open weightsSK Telecom · best of 2 rows39.0%IndependentreasoningonPartially comparable-52.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
98Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 3 · best of 2 rows38.6%IndependentreasoningonPartially comparable-52.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
99Qwen3 Coder NextOpen weightsQwen · Qwen3 · best of 2 rows38.2%IndependentreasoningoffPartially comparable-53.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
100Claude Sonnet 4ClosedAnthropic · Claude · best of 2 rows36.3%IndependentreasoningonPartially comparable-55.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →