Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · tau-bench · variant τ²

τ²-bench

dual-control tool-agent-user interaction (telecom, retail, airline)

data quality51

Updated 2 h ago · first seen 11 Sept 2026

Metric
pass^1 · %
Current results
880
Models
326
Current leader
Z.ai GLM 5.2 99.1%

Score history · llama-3-1-nemotron-nano-4b-reasoning 2 rows

Score history for llama-3-1-nemotron-nano-4b-reasoning0%5%10%15%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • llama-3-1-nemotron-nano-4b-reasoning
  • 11.7%aa_slug=llama-3-1-nemotron-nano-4b-reasoning · variant=Telecom · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
  • 11.7%aa_slug=llama-3-1-nemotron-nano-4b-reasoning · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · pass^1 · variant=Telecom · evaluator=Artificial Analysis

2 leader changes recorded, all dated 12 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 99.1%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent12 Sept 2026
  2. 90.3%grok-3-mini-reasoning SpaceXAI Independent12 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 308 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1Z.ai GLM 5.2Open weightsZ.ai (Zhipu AI) · GLM5.299.1%Independentreasoning_effortmaxleaderobs. 12 Sept 2026artificialanalysis.aiT2History
1jt-35b-flashClosedChina Mobile99.1%Independentreasoningoffleaderobs. 12 Sept 2026artificialanalysis.aiT2History
3GLM 4.7 FlashOpen weightsZ.ai (Zhipu AI) · GLM4.7 · best of 2 rows98.8%IndependentreasoningonPartially comparable-0.29 ptobs. 12 Sept 2026artificialanalysis.aiT2History
4Claude Fable 5ClosedAnthropic · Claude98.5%IndependentreasoningonPartially comparable-0.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
4GLM 5 TurboClosedZ.ai (Zhipu AI) · GLM598.5%IndependentreasoningonPartially comparable-0.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
4GLM 5V TurboClosedZ.ai (Zhipu AI) · GLM598.5%IndependentreasoningonPartially comparable-0.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
4Step 3.7 FlashOpen weightsStepFun · Step3.798.5%IndependentreasoningonPartially comparable-0.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
8GLM 5Open weightsZ.ai (Zhipu AI) · GLM5 · best of 2 rows98.3%IndependentreasoningonPartially comparable-0.87 ptobs. 12 Sept 2026artificialanalysis.aiT2History
9GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 2 rows97.7%IndependentreasoningonPartially comparable-1.46 ptobs. 12 Sept 2026artificialanalysis.aiT2History
9Grok 4.3ClosedxAI · Grok · best of 4 rows97.7%Independentreasoning_efforthighPartially comparable-1.46 ptobs. 12 Sept 2026artificialanalysis.aiT2History
9Qwen3.6 PlusClosedQwen · Qwen3.697.7%IndependentreasoningonPartially comparable-1.46 ptobs. 12 Sept 2026artificialanalysis.aiT2History
12grok-4-20-0309ClosedSpaceXAI · Grok 4.20 · best of 2 rows96.5%IndependentreasoningonPartially comparable-2.63 ptobs. 12 Sept 2026artificialanalysis.aiT2History
13deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek · best of 3 rows96.2%Independentreasoning_effortmaxComparable-2.92 ptobs. 12 Sept 2026artificialanalysis.aiT2History
14GLM 4.7Open weightsZ.ai (Zhipu AI) · GLM4.7 · best of 2 rows95.9%IndependentreasoningonPartially comparable-3.21 ptobs. 12 Sept 2026artificialanalysis.aiT2History
14Kimi K2.5Open weightsMoonshot AI · Kimi · best of 2 rows95.9%IndependentreasoningonPartially comparable-3.21 ptobs. 12 Sept 2026artificialanalysis.aiT2History
14Kimi K2.6Open weightsMoonshot AI · Kimi · best of 2 rows95.9%IndependentreasoningonPartially comparable-3.21 ptobs. 12 Sept 2026artificialanalysis.aiT2History
14Qwen3.6 Max PreviewClosedQwen · Qwen3.695.9%IndependentreasoningonPartially comparable-3.21 ptobs. 12 Sept 2026artificialanalysis.aiT2History
18Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.195.6%IndependentreasoningonPartially comparable-3.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
18Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 3 rows95.6%Independentreasoning_effortmediumPartially comparable-3.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
18Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 2 rows95.6%IndependentreasoningonPartially comparable-3.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
18deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek · best of 3 rows95.6%Independentreasoning_efforthighPartially comparable-3.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
22MiniMax M2.5Open weightsMiniMax · MiniMax95.3%IndependentreasoningonPartially comparable-3.80 ptobs. 12 Sept 2026artificialanalysis.aiT2History
22Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 2 rows95.3%IndependentreasoningonPartially comparable-3.80 ptobs. 12 Sept 2026artificialanalysis.aiT2History
24mimo-v2-flashOpen weightsXiaomi · best of 2 rows95.0%IndependentreasoningonPartially comparable-4.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
24mimo-v2-proClosedXiaomi95.0%IndependentreasoningonPartially comparable-4.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
26Qwen3.7 MaxClosedQwen · Qwen3.794.7%IndependentreasoningonPartially comparable-4.38 ptobs. 12 Sept 2026artificialanalysis.aiT2History
27Claude Opus 4.8ClosedAnthropic · Claude94.4%Independentreasoning_effortmaxComparable-4.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
27Step 3.5 FlashOpen weightsStepFun · Step3.5 · best of 2 rows94.4%IndependentreasoningonPartially comparable-4.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
29MiMo-V2.5-ProOpen weightsXiaomi · best of 2 rows94.2%IndependentreasoningonPartially comparable-4.97 ptobs. 12 Sept 2026artificialanalysis.aiT2History
29Mistral Medium 3.5Open weightsMistral AI · Mistral94.2%IndependentreasoningonPartially comparable-4.97 ptobs. 12 Sept 2026artificialanalysis.aiT2History
29Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 2 rows94.2%IndependentreasoningonPartially comparable-4.97 ptobs. 12 Sept 2026artificialanalysis.aiT2History
32Qwen3.5-27BOpen weightsQwen · Qwen3.5 · best of 2 rows93.9%IndependentreasoningonPartially comparable-5.26 ptobs. 12 Sept 2026artificialanalysis.aiT2History
32gpt-5.5ClosedOpenAI · GPT 5.5 · best of 5 rows93.9%Independentreasoning_effortxhighPartially comparable-5.26 ptobs. 12 Sept 2026artificialanalysis.aiT2History
34Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 2 rows93.6%IndependentreasoningonPartially comparable-5.55 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35grok-4-1-fastClosedSpaceXAI · Grok 4.1 · best of 2 rows93.3%IndependentreasoningonPartially comparable-5.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35mimo-v2-0206Open weightsXiaomi93.3%IndependentreasoningonPartially comparable-5.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35tri-21b-think-previewOpen weightsTrillion Labs93.3%IndependentreasoningonPartially comparable-5.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
38Grok 4.20ClosedxAI · Grok · best of 2 rows93.0%IndependentreasoningonPartially comparable-6.14 ptobs. 12 Sept 2026artificialanalysis.aiT2History
38Kimi K2 ThinkingOpen weightsMoonshot AI · Kimi93.0%IndependentreasoningonPartially comparable-6.14 ptobs. 12 Sept 2026artificialanalysis.aiT2History
38Qwen 3.7 PlusClosedQwen · Qwen3.793.0%IndependentreasoningonPartially comparable-6.14 ptobs. 12 Sept 2026artificialanalysis.aiT2History
38jt-miniClosedChina Mobile93.0%IndependentreasoningoffPartially comparable-6.14 ptobs. 12 Sept 2026artificialanalysis.aiT2History
42Hy3Open weightsTencent · best of 2 rows92.7%IndependentreasoningonPartially comparable-6.43 ptobs. 12 Sept 2026artificialanalysis.aiT2History
42nova-2-0-proClosedAmazon Web Services · Nova 2.0 · best of 3 rows92.7%Independentreasoningonreasoning_effortmediumPartially comparable-6.43 ptobs. 12 Sept 2026artificialanalysis.aiT2History
44ring-2-6-1tOpen weightsinclusionAI92.4%IndependentreasoningonPartially comparable-6.72 ptobs. 12 Sept 2026artificialanalysis.aiT2History
45Claude Opus 4.6ClosedAnthropic · Claude · best of 2 rows92.1%Independentreasoningadaptivereasoning_effortmaxPartially comparable-7.01 ptobs. 12 Sept 2026artificialanalysis.aiT2History
45GPT-5.2-CodexClosedOpenAI · GPT 5.292.1%Independentreasoning_effortxhighPartially comparable-7.01 ptobs. 12 Sept 2026artificialanalysis.aiT2History
45Qwen3.5-4BOpen weightsQwen · Qwen3.5 · best of 2 rows92.1%IndependentreasoningonPartially comparable-7.01 ptobs. 12 Sept 2026artificialanalysis.aiT2History
48muse-sparkClosedMeta AI91.5%IndependentreasoningonPartially comparable-7.60 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49mimo-v2-omniClosedXiaomi91.2%IndependentreasoningonPartially comparable-7.89 ptobs. 12 Sept 2026artificialanalysis.aiT2History
50DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 3 rows90.6%IndependentreasoningonPartially comparable-8.48 ptobs. 12 Sept 2026artificialanalysis.aiT2History
50MiMo-V2.5Open weightsXiaomi90.6%IndependentreasoningonPartially comparable-8.48 ptobs. 12 Sept 2026artificialanalysis.aiT2History
52grok-3-mini-reasoningClosedSpaceXAI · Grok 390.3%Independentreasoningonreasoning_efforthighPartially comparable-8.77 ptobs. 12 Sept 2026artificialanalysis.aiT2History
53Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi90.1%IndependentreasoningonPartially comparable-9.06 ptobs. 12 Sept 2026artificialanalysis.aiT2History
53Trinity Large ThinkingOpen weightsArcee AI90.1%IndependentreasoningonPartially comparable-9.06 ptobs. 12 Sept 2026artificialanalysis.aiT2History
55ling-2-6-1tOpen weightsinclusionAI89.8%IndependentreasoningoffPartially comparable-9.35 ptobs. 12 Sept 2026artificialanalysis.aiT2History
56Claude Opus 4.5ClosedAnthropic · Claude · best of 2 rows89.5%IndependentreasoningonPartially comparable-9.65 ptobs. 12 Sept 2026artificialanalysis.aiT2History
56KAT-Coder-Pro V2ClosedKwaipilot89.5%IndependentreasoningoffPartially comparable-9.65 ptobs. 12 Sept 2026artificialanalysis.aiT2History
58Qwen3.5-35B-A3BOpen weightsQwen · Qwen3.5 · best of 2 rows89.2%IndependentreasoningonPartially comparable-9.94 ptobs. 12 Sept 2026artificialanalysis.aiT2History
59MiniMax-M3Open weightsMiniMax · MiniMax88.9%IndependentreasoningonPartially comparable-10.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60Claude Opus 4.7ClosedAnthropic · Claude · best of 2 rows88.6%Independentreasoning_effortmaxComparable-10.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60kat-coder-pro-v1ClosedKwaiKAT88.6%IndependentreasoningoffPartially comparable-10.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
62qwen3-5-omni-plusClosedAlibaba Group · Qwen3.588.3%IndependentreasoningoffPartially comparable-10.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
63mimo-v2-omni-0327ClosedXiaomi88.0%IndependentreasoningonPartially comparable-11.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
64minicpm-v4-6-1-3bOpen weightsOpenBMB87.7%IndependentreasoningoffPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
65hyperclova-x-seed-think-32bOpen weightsNaver · Seed87.4%IndependentreasoningonPartially comparable-11.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
66gemini-3-proClosedGoogle · Gemini 3 · best of 2 rows87.1%Independentreasoning_efforthighPartially comparable-12.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
66gpt-5.4ClosedOpenAI · GPT 5.4 · best of 3 rows87.1%Independentreasoning_effortxhighPartially comparable-12.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
68GPT-5-CodexClosedOpenAI · GPT 586.8%Independentreasoning_efforthighPartially comparable-12.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
68MiniMax M2Open weightsMiniMax · MiniMax86.8%IndependentreasoningonPartially comparable-12.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
68Qwen3.5-9BOpen weightsQwen · Qwen3.5 · best of 2 rows86.8%IndependentreasoningonPartially comparable-12.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
71gpt-5ClosedOpenAI · GPT 5 · best of 4 rows86.5%Independentreasoning_effortmediumPartially comparable-12.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
71mi-dm-k-2-5-pro-dec28ClosedKorea Telecom86.5%IndependentreasoningonPartially comparable-12.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
73Solar Pro 3ClosedUpstage · Solar86.3%IndependentreasoningonPartially comparable-12.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
73gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 5 rows86.3%Independentreasoning_effortmaxComparable-12.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
75gpt-5.3-codexClosedOpenAI · GPT 5.386.0%Independentreasoning_effortxhighPartially comparable-13.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
75ling-2-6-flashOpen weightsinclusionAI86.0%IndependentreasoningoffPartially comparable-13.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
77MiniMax M2.1Open weightsMiniMax · MiniMax85.4%IndependentreasoningonPartially comparable-13.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
78gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 5 rows85.1%Independentreasoning_effortmaxComparable-14.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79MiniMax M2.7Open weightsMiniMax · MiniMax84.8%IndependentreasoningonPartially comparable-14.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79gpt-5.2ClosedOpenAI · GPT 5.2 · best of 3 rows84.8%Independentreasoning_effortxhighPartially comparable-14.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
81qwen3-5-omni-flashClosedAlibaba Group · Qwen3.584.5%IndependentreasoningoffPartially comparable-14.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
82ernie-5-0-thinking-previewClosedBaidu · ERNIE 5.083.9%IndependentreasoningonPartially comparable-15.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
83Qwen3 MaxClosedQwen · Qwen3 · best of 2 rows83.6%IndependentreasoningonPartially comparable-15.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
83qwen3-max-thinking-previewClosedAlibaba Group · Qwen383.6%IndependentreasoningonPartially comparable-15.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
85Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 383.3%IndependentreasoningonPartially comparable-15.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
85gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 3 rows83.3%Independentreasoning_effortxhighPartially comparable-15.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
87GPT-5.1-CodexClosedOpenAI · GPT 5.183.0%Independentreasoning_efforthighPartially comparable-16.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88minicpm5-1bOpen weightsOpenBMB · best of 2 rows82.5%IndependentreasoningoffPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
89gpt-5.1ClosedOpenAI · GPT 5.1 · best of 2 rows81.9%Independentreasoning_efforthighPartially comparable-17.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
90Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 2 rows81.6%IndependentreasoningoffPartially comparable-17.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
90nex-n2-proOpen weightsNex AGI81.6%IndependentreasoningonPartially comparable-17.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
92tri-21b-think-v0-5Open weightsTrillion Labs81.0%IndependentreasoningonPartially comparable-18.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
93command-a-plusOpen weightsCohere · Command80.7%IndependentreasoningonPartially comparable-18.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
93o3ClosedOpenAI · OpenAI o-series80.7%IndependentreasoningonPartially comparable-18.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
95gemini-3-flashClosedGoogle · Gemini 3 · best of 2 rows80.4%IndependentreasoningonPartially comparable-18.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
95nova-2-0-omniClosedAmazon Web Services · Nova 2.0 · best of 3 rows80.4%Independentreasoningonreasoning_effortmediumPartially comparable-18.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
97Claude Sonnet 4.6ClosedAnthropic · Claude · best of 3 rows79.5%Independentreasoningoffreasoning_efforthighPartially comparable-19.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
97Qwen3 Coder NextOpen weightsQwen · Qwen379.5%IndependentreasoningoffPartially comparable-19.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
97longcat-flash-liteOpen weightsLongCat79.5%IndependentreasoningoffPartially comparable-19.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
100Claude Sonnet 4.5ClosedAnthropic · Claude · best of 2 rows78.1%IndependentreasoningonPartially comparable-21.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →