Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · tau-bench · variant τ²

τ²-bench

dual-control tool-agent-user interaction (telecom, retail, airline)

data quality51

Updated 4 h ago · first seen 11 Sept 2026

Metric
pass^1 · %
Current results
880
Models
326
Current leader
Z.ai GLM 5.2 99.1%

Score history · GPT-5.1-Codex 2 rows

Score history for GPT-5.1-Codex0%20%40%60%80%100%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • GPT-5.1-Codex
  • 83.04%aa_slug=gpt-5-1-codex · variant=Telecom · evaluator=Artificial Analysis · index_version=4.312 Sept 2026
  • 83.04%aa_slug=gpt-5-1-codex · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · pass^1 · evaluator=Artificial Analysis

2 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 99.1%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
  2. 90.3%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 323 models · trust independent-evaluator

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
300gemma-3n-e4bOpen weightsGoogle · Gemma 34.97%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302DeepSeek-R1-0528-Qwen3-8BOpen weightsDeepSeek · Qwen30%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AI · Kimi0%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302Llama-3.2-1BRestricted weightsMeta AI · Llama 3.20%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302Mistral 7BOpen weightsMistral AI · Mistral0%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302Molmo2-8BOpen weightsAllen Institute for AI0%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302Olmo-3-7B-ThinkOpen weightsAllen Institute for AI · OLMo 30%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302Phi 4Open weightsMicrosoft · Phi40%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302Reka Flash 3Open weightsrekaai0%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302deepseek-v3-2-specialeOpen weightsDeepSeek · DeepSeek0%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302ernie-4-5-300b-a47bOpen weightsBaidu · ERNIE 4.50%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302gemma-3n-e2bOpen weightsGoogle · Gemma 30%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302gpt-5-chatgptClosedOpenAI · GPT 50%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302llama-3-instruct-70bOpen weightsMeta AI · Llama 30%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302llama-3-instruct-8bOpen weightsMeta AI · Llama 30%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302molmo-7b-dOpen weightsAllen Institute for AI · Molmo0%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302olmo-2-32bOpen weightsAllen Institute for AI · OLMo 20%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302olmo-2-7bOpen weightsAllen Institute for AI · OLMo 20%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302olmo-3-32b-thinkOpen weightsAllen Institute for AI · OLMo 30%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302phi-3-miniOpen weightsMicrosoft · Phi30%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302ring-flash-2-0Open weightsinclusionAI0%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302sarvam-m-reasoningOpen weightsSarvam0%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302tiny-aya-globalRestricted weightsCohere · Aya0%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →