Skip to content
AI Atlas
BenchmarkActive

Terminal-Bench

tbench.ai

terminal tasks solved by agents

quality57

Updated 10 h ago · first seen 11 Sept 2026

Metric
· %
Current results
0
Models
0
Current leader

No results recorded for this benchmark yet — its sources are being connected. The definition, aliases and variants are kept so links resolve; nothing is fabricated.

Score history · a-x-k2 1 row

Not enough history to chart — a single observation (38.95% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.

  • 38.95%aa_slug=a-x-k2 · variant=v2.1 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Leaderboard 315 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
276lfm2-8b-a1bOpen weightsLiquid AI · LFM20%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.3 · best of 2 rows0%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276llama-3-instruct-8bOpen weightsMeta AI · Llama 30%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276minicpm-v4-6-1-3bOpen weightsOpenBMB0%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276minicpm5-1bOpen weightsOpenBMB · best of 2 rows0%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276molmo-7b-dOpen weightsAllen Institute for AI · Molmo0%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276nanbeige4-1-3bOpen weightsNanbeige0%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276olmo-2-32bOpen weightsAllen Institute for AI · OLMo 20%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276olmo-2-7bOpen weightsAllen Institute for AI · OLMo 20%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276olmo-3-1-32b-instructOpen weightsAllen Institute for AI · OLMo 3.1 · best of 2 rows0%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276phi-3-miniOpen weightsMicrosoft · Phi30%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276phi-4-miniOpen weightsMicrosoft · Phi40%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276qwen3-0.6b-instructOpen weightsAlibaba Group · Qwen3.00%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276qwen3-1.7b-instructOpen weightsAlibaba Group · Qwen3.1 · best of 2 rows0%IndependentreasoningonPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
276tiny-aya-globalRestricted weightsCohere · Aya0%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →