Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · terminal-bench · variant 1.0

Terminal-Bench

tbench.ai

terminal tasks solved by agents

data quality57

Updated 2 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,639
Models
376
Current leader
gpt-5.6-sol 65.9%

Frontier over time · accuracy · variant=v4.0 · evaluator=Artificial Analysis

6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 59.6%gpt-6-astra OpenAI Independent11 Sept 2026
  2. 59.1%gpt-6-astra OpenAI Independent11 Sept 2026
  3. 55.0%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  4. 54.0%gpt-6-astra OpenAI Independent11 Sept 2026
  5. 49.5%gpt-6-astra OpenAI Independent11 Sept 2026
  6. 34.3%Claude Opus 5 Anthropic Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 114 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
60Solar Pro 3ClosedUpstage · Solar · best of 2 rows0%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60celeris-1ClosedCeleris · best of 2 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60deepseek-r1-0120Open weightsDeepSeek · DeepSeek · best of 2 rows0%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60gpt-5-miniClosedOpenAI · GPT 5 · best of 2 rows0%Independentreasoning_efforthighPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60gpt-oss-120bOpen weightsOpenAI · gpt-oss · best of 2 rows0%Independentreasoning_efforthighPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 2 rows0%Independentreasoning_efforthighPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60granite-4.2-3bOpen weightsIBM · Granite 4.2 · best of 2 rows0%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60ling-3-0-tinyOpen weightsinclusionAI · best of 2 rows0%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60minicpm5-2bOpen weightsOpenBMB · best of 2 rows0%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60o3-miniClosedOpenAI · OpenAI o-series · best of 2 rows0%Independentreasoning_efforthighPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60qwen3-14b-instructOpen weightsAlibaba Group · Qwen30%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 2 rows0%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60qwen3-32b-instructOpen weightsAlibaba Group · Qwen30%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60qwen3-8b-instructOpen weightsAlibaba Group · Qwen30%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →