Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · terminal-bench · variant 1.0

Terminal-Bench

tbench.ai

terminal tasks solved by agents

data quality57

Updated 52 min ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,639
Models
376
Current leader
gpt-5.6-sol 65.9%

Score history · gpt-5.4-nano 10 rows

Score history for gpt-5.4-nano0%20%40%60%80%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • gpt-5.4-nano
  • 33.33%aa_slug=gpt-5-4-nano-medium · variant=hard · evaluator=Artificial Analysis · index_version=4.312 Sept 2026
  • 24.24%aa_slug=gpt-5-4-nano-non-reasoning · variant=hard · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
  • 42.42%aa_slug=gpt-5-4-nano · variant=hard · evaluator=Artificial Analysis · index_version=4.312 Sept 2026
  • 60.67%aa_slug=gpt-5-4-nano · variant=v2.1 · evaluator=Artificial Analysis · index_version=4.312 Sept 2026
  • 0.51%aa_slug=gpt-5-4-nano · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.312 Sept 2026
  • 33.33%aa_slug=gpt-5-4-nano-medium · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 24.24%aa_slug=gpt-5-4-nano-non-reasoning · variant=hard · evaluator=Artificial Analysis · reasoning=off11 Sept 2026
  • 42.42%aa_slug=gpt-5-4-nano · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 60.67%aa_slug=gpt-5-4-nano · variant=v2.1 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 0.51%aa_slug=gpt-5-4-nano · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=hard · evaluator=Artificial Analysis

10 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 65.9%gpt-5.6-sol OpenAI Independent11 Sept 2026
  2. 61.4%gpt-5.6-sol OpenAI Independent11 Sept 2026
  3. 53.0%Claude Sonnet 4.6 Anthropic Independent11 Sept 2026
  4. 51.5%Claude Opus 4.7 Anthropic Independent11 Sept 2026
  5. 50.8%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
  6. 49.2%KAT-Coder-Pro V2 Kwaipilot Independent11 Sept 2026
  7. 33.3%GPT-5.1-Codex Mini OpenAI Independent11 Sept 2026
  8. 26.5%Grok 4.3 xAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 317 models

Select models with +, then Compare.

No result in this group with these filters

Relax the trust / organization filters or pick another comparability group.