Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · tau-bench · variant τ²

τ²-bench

dual-control tool-agent-user interaction (telecom, retail, airline)

data quality51

Updated 1 h ago · first seen 11 Sept 2026

Metric
pass^1 · %
Current results
880
Models
326
Current leader
Z.ai GLM 5.2 99.1%

Score history · Claude Haiku 4.5 2 rows

Not enough history to chart — 2 observations, all dated 11 Sept 2026. Rows under different configurations count separately; the list below shows each one.

  • 54.68%aa_slug=claude-4-5-haiku-reasoning · evaluator=Artificial Analysis · reasoning=on · index_version=4.311 Sept 2026
  • 32.46%aa_slug=claude-4-5-haiku · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · pass^1 · evaluator=Artificial Analysis

2 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 99.1%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
  2. 90.3%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 323 models · trust independent-evaluator

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1Z.ai GLM 5.2Open weightsZ.ai (Zhipu AI) · GLM5.299.1%Independentgroup defaultsleaderobs. 11 Sept 2026artificialanalysis.aiT2History
1jt-35b-flashClosedChina Mobile99.1%Independentgroup defaultsleaderobs. 11 Sept 2026artificialanalysis.aiT2History
3GLM 4.7 FlashOpen weightsZ.ai (Zhipu AI) · GLM4.7 · best of 2 rows98.8%Independentgroup defaultsComparable-0.29 ptobs. 11 Sept 2026artificialanalysis.aiT2History
4Claude Fable 5ClosedAnthropic · Claude98.5%Independentgroup defaultsComparable-0.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
4GLM 5 TurboClosedZ.ai (Zhipu AI) · GLM598.5%Independentgroup defaultsComparable-0.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
4GLM 5V TurboClosedZ.ai (Zhipu AI) · GLM598.5%Independentgroup defaultsComparable-0.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
4Step 3.7 FlashOpen weightsStepFun · Step3.798.5%Independentgroup defaultsComparable-0.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
8GLM 5Open weightsZ.ai (Zhipu AI) · GLM5 · best of 2 rows98.3%Independentgroup defaultsComparable-0.87 ptobs. 11 Sept 2026artificialanalysis.aiT2History
9GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 2 rows97.7%Independentgroup defaultsComparable-1.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
9Grok 4.3ClosedxAI · Grok · best of 4 rows97.7%Independentgroup defaultsComparable-1.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
9Qwen3.6 PlusClosedQwen · Qwen3.697.7%Independentgroup defaultsComparable-1.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
12grok-4-20-0309ClosedSpaceXAI · Grok 4.20 · best of 2 rows96.5%Independentgroup defaultsComparable-2.63 ptobs. 11 Sept 2026artificialanalysis.aiT2History
13deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek96.2%Independentgroup defaultsComparable-2.92 ptobs. 11 Sept 2026artificialanalysis.aiT2History
14GLM 4.7Open weightsZ.ai (Zhipu AI) · GLM4.7 · best of 2 rows95.9%Independentgroup defaultsComparable-3.21 ptobs. 11 Sept 2026artificialanalysis.aiT2History
14Kimi K2.5Open weightsMoonshot AI · Kimi · best of 2 rows95.9%Independentgroup defaultsComparable-3.21 ptobs. 11 Sept 2026artificialanalysis.aiT2History
14Kimi K2.6Open weightsMoonshot AI · Kimi · best of 2 rows95.9%Independentgroup defaultsComparable-3.21 ptobs. 11 Sept 2026artificialanalysis.aiT2History
14Qwen3.6 Max PreviewClosedQwen · Qwen3.695.9%Independentgroup defaultsComparable-3.21 ptobs. 11 Sept 2026artificialanalysis.aiT2History
18Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.195.6%Independentgroup defaultsComparable-3.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
18Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 3 rows95.6%Independentreasoning_effortmediumPartially comparable-3.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
18Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 2 rows95.6%Independentgroup defaultsComparable-3.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
18deepseek-v4-flash-0420-highOpen weightsDeepSeek · DeepSeek95.6%Independentgroup defaultsComparable-3.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
22MiniMax M2.5Open weightsMiniMax · MiniMax95.3%Independentgroup defaultsComparable-3.80 ptobs. 11 Sept 2026artificialanalysis.aiT2History
22Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 2 rows95.3%Independentgroup defaultsComparable-3.80 ptobs. 11 Sept 2026artificialanalysis.aiT2History
24deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek95.0%Independentgroup defaultsComparable-4.09 ptobs. 11 Sept 2026artificialanalysis.aiT2History
24mimo-v2-flashOpen weightsXiaomi · best of 2 rows95.0%IndependentreasoningonPartially comparable-4.09 ptobs. 11 Sept 2026artificialanalysis.aiT2History
24mimo-v2-proClosedXiaomi95.0%Independentgroup defaultsComparable-4.09 ptobs. 11 Sept 2026artificialanalysis.aiT2History
27Qwen3.7 MaxClosedQwen · Qwen3.794.7%Independentgroup defaultsComparable-4.38 ptobs. 11 Sept 2026artificialanalysis.aiT2History
28Claude Opus 4.8ClosedAnthropic · Claude94.4%Independentgroup defaultsComparable-4.68 ptobs. 11 Sept 2026artificialanalysis.aiT2History
28Step 3.5 FlashOpen weightsStepFun · Step3.5 · best of 2 rows94.4%Independentgroup defaultsComparable-4.68 ptobs. 11 Sept 2026artificialanalysis.aiT2History
28deepseek-v4-flash-0420-non-reasoningOpen weightsDeepSeek · DeepSeek94.4%Independentgroup defaultsComparable-4.68 ptobs. 11 Sept 2026artificialanalysis.aiT2History
31MiMo-V2.5-ProOpen weightsXiaomi · best of 2 rows94.2%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
31Mistral Medium 3.5Open weightsMistral AI · Mistral94.2%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
31Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 2 rows94.2%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
31deepseek-v4-pro-0424-highOpen weightsDeepSeek · DeepSeek94.2%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
35Qwen3.5-27BOpen weightsQwen · Qwen3.5 · best of 2 rows93.9%Independentgroup defaultsComparable-5.26 ptobs. 11 Sept 2026artificialanalysis.aiT2History
35gpt-5.5ClosedOpenAI · GPT 5.5 · best of 5 rows93.9%Independentgroup defaultsComparable-5.26 ptobs. 11 Sept 2026artificialanalysis.aiT2History
37Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 2 rows93.6%Independentgroup defaultsComparable-5.55 ptobs. 11 Sept 2026artificialanalysis.aiT2History
38grok-4-1-fastClosedSpaceXAI · Grok 4.1 · best of 2 rows93.3%IndependentreasoningonPartially comparable-5.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
38mimo-v2-0206Open weightsXiaomi93.3%Independentgroup defaultsComparable-5.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
38tri-21b-think-previewOpen weightsTrillion Labs93.3%Independentgroup defaultsComparable-5.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
41Grok 4.20ClosedxAI · Grok93.0%Independentgroup defaultsComparable-6.14 ptobs. 11 Sept 2026artificialanalysis.aiT2History
41Kimi K2 ThinkingOpen weightsMoonshot AI · Kimi93.0%Independentgroup defaultsComparable-6.14 ptobs. 11 Sept 2026artificialanalysis.aiT2History
41Qwen 3.7 PlusClosedQwen · Qwen3.793.0%Independentgroup defaultsComparable-6.14 ptobs. 11 Sept 2026artificialanalysis.aiT2History
41jt-miniClosedChina Mobile93.0%Independentgroup defaultsComparable-6.14 ptobs. 11 Sept 2026artificialanalysis.aiT2History
45Hy3 previewOpen weightsTencent92.7%Independentgroup defaultsComparable-6.43 ptobs. 11 Sept 2026artificialanalysis.aiT2History
45nova-2-0-proClosedAmazon Web Services · Nova 2.0 · best of 3 rows92.7%Independentreasoningonreasoning_effortmediumPartially comparable-6.43 ptobs. 11 Sept 2026artificialanalysis.aiT2History
47ring-2-6-1tOpen weightsinclusionAI92.4%Independentgroup defaultsComparable-6.72 ptobs. 11 Sept 2026artificialanalysis.aiT2History
48Claude Opus 4.6ClosedAnthropic · Claude · best of 2 rows92.1%IndependentreasoningadaptivePartially comparable-7.01 ptobs. 11 Sept 2026artificialanalysis.aiT2History
48GPT-5.2-CodexClosedOpenAI · GPT 5.292.1%Independentgroup defaultsComparable-7.01 ptobs. 11 Sept 2026artificialanalysis.aiT2History
48Qwen3.5-4BOpen weightsQwen · Qwen3.5 · best of 2 rows92.1%Independentgroup defaultsComparable-7.01 ptobs. 11 Sept 2026artificialanalysis.aiT2History
51muse-sparkClosedMeta AI91.5%Independentgroup defaultsComparable-7.60 ptobs. 11 Sept 2026artificialanalysis.aiT2History
52deepseek-v4-pro-0424-non-reasoningOpen weightsDeepSeek · DeepSeek91.2%Independentgroup defaultsComparable-7.89 ptobs. 11 Sept 2026artificialanalysis.aiT2History
52mimo-v2-omniClosedXiaomi91.2%Independentgroup defaultsComparable-7.89 ptobs. 11 Sept 2026artificialanalysis.aiT2History
54DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V390.6%Independentgroup defaultsComparable-8.48 ptobs. 11 Sept 2026artificialanalysis.aiT2History
54MiMo-V2.5Open weightsXiaomi90.6%Independentgroup defaultsComparable-8.48 ptobs. 11 Sept 2026artificialanalysis.aiT2History
56grok-3-mini-reasoningClosedSpaceXAI · Grok 390.3%Independentgroup defaultsComparable-8.77 ptobs. 11 Sept 2026artificialanalysis.aiT2History
57Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi90.1%Independentgroup defaultsComparable-9.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
57Trinity Large ThinkingOpen weightsArcee AI90.1%Independentgroup defaultsComparable-9.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
59ling-2-6-1tOpen weightsinclusionAI89.8%Independentgroup defaultsComparable-9.35 ptobs. 11 Sept 2026artificialanalysis.aiT2History
60Claude Opus 4.5ClosedAnthropic · Claude · best of 2 rows89.5%IndependentreasoningonPartially comparable-9.65 ptobs. 11 Sept 2026artificialanalysis.aiT2History
60KAT-Coder-Pro V2ClosedKwaipilot89.5%Independentgroup defaultsComparable-9.65 ptobs. 11 Sept 2026artificialanalysis.aiT2History
62Qwen3.5-35B-A3BOpen weightsQwen · Qwen3.5 · best of 2 rows89.2%Independentgroup defaultsComparable-9.94 ptobs. 11 Sept 2026artificialanalysis.aiT2History
63MiniMax-M3Open weightsMiniMax · MiniMax88.9%Independentgroup defaultsComparable-10.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
64Claude Opus 4.7ClosedAnthropic · Claude · best of 2 rows88.6%Independentgroup defaultsComparable-10.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
64kat-coder-pro-v1ClosedKwaiKAT88.6%Independentgroup defaultsComparable-10.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
66qwen3-5-omni-plusClosedAlibaba Group · Qwen3.588.3%Independentgroup defaultsComparable-10.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
67mimo-v2-omni-0327ClosedXiaomi88.0%Independentgroup defaultsComparable-11.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
68minicpm-v4-6-1-3bOpen weightsOpenBMB87.7%Independentgroup defaultsComparable-11.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
69hyperclova-x-seed-think-32bOpen weightsNaver · Seed87.4%Independentgroup defaultsComparable-11.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
70gemini-3-proClosedGoogle · Gemini 3 · best of 2 rows87.1%Independentgroup defaultsComparable-12.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
70gpt-5.4ClosedOpenAI · GPT 5.4 · best of 3 rows87.1%Independentgroup defaultsComparable-12.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
72GPT-5-CodexClosedOpenAI · GPT 586.8%Independentgroup defaultsComparable-12.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
72MiniMax M2Open weightsMiniMax · MiniMax86.8%Independentgroup defaultsComparable-12.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
72Qwen3.5-9BOpen weightsQwen · Qwen3.5 · best of 2 rows86.8%Independentgroup defaultsComparable-12.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
75gpt-5ClosedOpenAI · GPT 5 · best of 4 rows86.5%Independentreasoning_effortmediumPartially comparable-12.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
75mi-dm-k-2-5-pro-dec28ClosedKorea Telecom86.5%Independentgroup defaultsComparable-12.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
77Solar Pro 3ClosedUpstage · Solar86.3%Independentgroup defaultsComparable-12.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
77gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 5 rows86.3%Independentgroup defaultsComparable-12.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
79gpt-5.3-codexClosedOpenAI · GPT 5.386.0%Independentgroup defaultsComparable-13.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
79ling-2-6-flashOpen weightsinclusionAI86.0%Independentgroup defaultsComparable-13.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
81MiniMax M2.1Open weightsMiniMax · MiniMax85.4%Independentgroup defaultsComparable-13.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
82gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 5 rows85.1%Independentgroup defaultsComparable-14.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
83MiniMax M2.7Open weightsMiniMax · MiniMax84.8%Independentgroup defaultsComparable-14.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
83gpt-5.2ClosedOpenAI · GPT 5.2 · best of 3 rows84.8%Independentgroup defaultsComparable-14.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
85qwen3-5-omni-flashClosedAlibaba Group · Qwen3.584.5%Independentgroup defaultsComparable-14.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
86ernie-5-0-thinking-previewClosedBaidu · ERNIE 5.083.9%Independentgroup defaultsComparable-15.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
87Qwen3 MaxClosedQwen · Qwen3 · best of 2 rows83.6%IndependentreasoningonPartially comparable-15.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
87qwen3-max-thinking-previewClosedAlibaba Group · Qwen383.6%Independentgroup defaultsComparable-15.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
89Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 383.3%Independentgroup defaultsComparable-15.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
89gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 3 rows83.3%Independentgroup defaultsComparable-15.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
91GPT-5.1-CodexClosedOpenAI · GPT 5.183.0%Independentgroup defaultsComparable-16.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
92minicpm5-1bOpen weightsOpenBMB · best of 2 rows82.5%IndependentreasoningoffPartially comparable-16.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
93gpt-5.1ClosedOpenAI · GPT 5.1 · best of 2 rows81.9%Independentgroup defaultsComparable-17.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
94Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 2 rows81.6%IndependentreasoningoffPartially comparable-17.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
94nex-n2-proOpen weightsNex AGI81.6%Independentgroup defaultsComparable-17.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
96tri-21b-think-v0-5Open weightsTrillion Labs81.0%Independentgroup defaultsComparable-18.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
97command-a-plusOpen weightsCohere · Command80.7%Independentgroup defaultsComparable-18.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
97o3ClosedOpenAI · OpenAI o-series80.7%Independentgroup defaultsComparable-18.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
99Gemini 3 Flash PreviewClosedGoogle · Gemini 380.4%Independentgroup defaultsComparable-18.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
99nova-2-0-omniClosedAmazon Web Services · Nova 2.0 · best of 3 rows80.4%Independentreasoningonreasoning_effortmediumPartially comparable-18.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →