Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · terminal-bench · variant 1.0

Terminal-Bench

tbench.ai

terminal tasks solved by agents

data quality57

Updated 3 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,639
Models
376
Current leader
gpt-5.6-sol 65.9%

Score history · olmo-3-1-32b-instruct 3 rows

Score history for olmo-3-1-32b-instruct0%0.2%0.4%0.6%0.8%1%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • olmo-3-1-32b-instruct
  • 0%aa_slug=olmo-3-1-32b-instruct · variant=hard · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
  • 0%aa_slug=olmo-3-1-32b-think · variant=hard · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
  • 0%aa_slug=olmo-3-1-32b-instruct · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=hard · evaluator=Artificial Analysis

10 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 65.9%gpt-5.6-sol OpenAI Independent11 Sept 2026
  2. 61.4%gpt-5.6-sol OpenAI Independent11 Sept 2026
  3. 53.0%Claude Sonnet 4.6 Anthropic Independent11 Sept 2026
  4. 51.5%Claude Opus 4.7 Anthropic Independent11 Sept 2026
  5. 50.8%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
  6. 49.2%KAT-Coder-Pro V2 Kwaipilot Independent11 Sept 2026
  7. 33.3%GPT-5.1-Codex Mini OpenAI Independent11 Sept 2026
  8. 26.5%Grok 4.3 xAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 317 models · trust independent-evaluator

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
279lfm2-24b-a2bOpen weightsLiquid AI · LFM2 · best of 2 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279lfm2-5-1-2b-thinkingOpen weightsLiquid AI · LFM2.5 · best of 2 rows0%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279lfm2-8b-a1bOpen weightsLiquid AI · LFM2 · best of 2 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.3 · best of 4 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279llama-3-instruct-8bOpen weightsMeta AI · Llama 3 · best of 2 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279minicpm-v4-6-1-3bOpen weightsOpenBMB · best of 2 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279minicpm5-1bOpen weightsOpenBMB · best of 4 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279molmo-7b-dOpen weightsAllen Institute for AI · Molmo · best of 2 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279nanbeige4-1-3bOpen weightsNanbeige · best of 2 rows0%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279olmo-2-32bOpen weightsAllen Institute for AI · OLMo 2 · best of 2 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279olmo-2-7bOpen weightsAllen Institute for AI · OLMo 2 · best of 2 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279olmo-3-1-32b-instructOpen weightsAllen Institute for AI · OLMo 3.1 · best of 3 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279phi-3-miniOpen weightsMicrosoft · Phi3 · best of 2 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279phi-4-miniOpen weightsMicrosoft · Phi4 · best of 2 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279qwen3-0.6b-instructOpen weightsAlibaba Group · Qwen3.0 · best of 3 rows0%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279qwen3-1.7b-instructOpen weightsAlibaba Group · Qwen3.1 · best of 4 rows0%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279tiny-aya-globalRestricted weightsCohere · Aya · best of 2 rows0%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →