Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · tau-bench · variant τ²

τ²-bench

dual-control tool-agent-user interaction (telecom, retail, airline)

data quality51

Updated 3 h ago · first seen 11 Sept 2026

Metric
pass^1 · %
Current results
880
Models
326
Current leader
Z.ai GLM 5.2 99.1%

Frontier over time · pass^1 · variant=Telecom · evaluator=Artificial Analysis

2 leader changes recorded, all dated 12 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 99.1%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent12 Sept 2026
  2. 90.3%grok-3-mini-reasoning SpaceXAI Independent12 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 308 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
100exaone-4-5-33bOpen weightsLG AI Research · EXAONE 4.578.1%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
102GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows76.9%IndependentreasoningoffPartially comparable-1.17 ptobs. 12 Sept 2026artificialanalysis.aiT2History
103gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 3 rows76.0%Independentreasoning_effortxhighPartially comparable-2.05 ptobs. 12 Sept 2026artificialanalysis.aiT2History
104Grok Build 0.1ClosedxAI · Grok75.7%IndependentreasoningonPartially comparable-2.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
104nova-2-0-liteClosedAmazon Web Services · Nova 2.0 · best of 4 rows75.7%Independentreasoningonreasoning_effortmediumPartially comparable-2.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
106grok-4ClosedSpaceXAI · Grok 474.8%IndependentreasoningonPartially comparable-3.22 ptobs. 12 Sept 2026artificialanalysis.aiT2History
107k-exaoneOpen weightsLG AI Research · EXAONE · best of 2 rows74.3%IndependentreasoningonPartially comparable-3.80 ptobs. 12 Sept 2026artificialanalysis.aiT2History
108Kimi K2 0905Open weightsMoonshot AI · Kimi73.4%IndependentreasoningoffPartially comparable-4.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
108claude-4-opusClosedAnthropic · Claude 473.4%IndependentreasoningonPartially comparable-4.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
110Claude Opus 4.1ClosedAnthropic · Claude71.4%IndependentreasoningonPartially comparable-6.70 ptobs. 12 Sept 2026artificialanalysis.aiT2History
111gpt-5-miniClosedOpenAI · GPT 5 · best of 3 rows71.0%Independentreasoning_effortmediumPartially comparable-7.02 ptobs. 12 Sept 2026artificialanalysis.aiT2History
112Mercury 2ClosedInception70.8%IndependentreasoningonPartially comparable-7.31 ptobs. 12 Sept 2026artificialanalysis.aiT2History
113apriel-v1-6-15b-thinkerOpen weightsServiceNow69.3%IndependentreasoningonPartially comparable-8.77 ptobs. 12 Sept 2026artificialanalysis.aiT2History
114apriel-v1-5-15b-thinkerOpen weightsServiceNow68.4%IndependentreasoningonPartially comparable-9.65 ptobs. 12 Sept 2026artificialanalysis.aiT2History
115Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 367.8%IndependentreasoningonPartially comparable-10.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
116gpt-oss-120bOpen weightsOpenAI · gpt-oss · best of 2 rows65.8%Independentreasoning_efforthighPartially comparable-12.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
116grok-4-fastClosedSpaceXAI · Grok 4 · best of 2 rows65.8%IndependentreasoningonPartially comparable-12.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
118Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 2 rows65.5%IndependentreasoningoffPartially comparable-12.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
119Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 2 rows65.2%IndependentreasoningoffPartially comparable-12.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
120Claude Sonnet 4ClosedAnthropic · Claude · best of 2 rows64.6%IndependentreasoningonPartially comparable-13.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
121hypernova-60bOpen weightsMultiverse Computing63.2%Independentreasoning_efforthighPartially comparable-14.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
122GPT-5.1-Codex MiniClosedOpenAI · GPT 5.162.9%Independentreasoning_efforthighPartially comparable-15.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
123o1ClosedOpenAI · OpenAI o-series62.6%IndependentreasoningonPartially comparable-15.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124Kimi K2 0711Open weightsMoonshot AI · Kimi61.1%IndependentreasoningoffPartially comparable-17.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
125gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 2 rows60.2%Independentreasoning_efforthighPartially comparable-17.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
126doubao-seed-codeClosedByteDance · Seed58.2%IndependentreasoningonPartially comparable-19.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127o4-miniClosedOpenAI · OpenAI o-series55.6%Independentreasoning_efforthighPartially comparable-22.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
128Claude 3.7 SonnetClosedAnthropic · Claude · best of 2 rows54.7%IndependentreasoningonPartially comparable-23.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
128Claude Haiku 4.5ClosedAnthropic · Claude · best of 2 rows54.7%IndependentreasoningonPartially comparable-23.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
130Gemini 2.5 ProClosedGoogle · Gemini 2.554.1%IndependentreasoningonPartially comparable-24.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
130Qwen3 VL 235B A22B InstructOpen weightsQwen · Qwen3 · best of 2 rows54.1%IndependentreasoningonPartially comparable-24.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
132Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 2 rows53.2%IndependentreasoningonPartially comparable-24.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
132nemotron-cascade-2-30b-a3bOpen weightsNVIDIA · Nemotron53.2%IndependentreasoningonPartially comparable-24.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
134gpt-4.1-miniClosedOpenAI · GPT 4.152.9%IndependentreasoningoffPartially comparable-25.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
135Magistral Medium 1.2ClosedMistral AI · Magistral52.0%IndependentreasoningonPartially comparable-26.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
136Seed-OSS-36B-InstructOpen weightsByteDance · Seed49.4%IndependentreasoningonPartially comparable-28.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
136gpt-5-5-instant-05-26ClosedOpenAI · GPT 5.549.4%IndependentreasoningonPartially comparable-28.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
136midm-250-pro-rsnsftClosedKorea Telecom49.4%IndependentreasoningonPartially comparable-28.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
139grok-3ClosedSpaceXAI · Grok 348.8%IndependentreasoningoffPartially comparable-29.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
140solar-open-100b-reasoningOpen weightsUpstage · Solar48.3%IndependentreasoningonPartially comparable-29.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
141DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V347.1%IndependentreasoningoffPartially comparable-31.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
141gpt-4.1ClosedOpenAI · GPT 4.147.1%IndependentreasoningoffPartially comparable-31.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
143sarvam-105bOpen weightsSarvam46.8%Independentreasoning_efforthighPartially comparable-31.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
144GLM 4.5 AirOpen weightsZ.ai (Zhipu AI) · GLM4.546.5%IndependentreasoningonPartially comparable-31.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
144motif-2-12-7bClosedMotif Technologies46.5%IndependentreasoningonPartially comparable-31.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
146Qwen3 VL 32B InstructOpen weightsQwen · Qwen3 · best of 2 rows45.6%IndependentreasoningonPartially comparable-32.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
146gemini-2-5-flash-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 4 rows45.6%IndependentreasoningonPartially comparable-32.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
148nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 345.3%IndependentreasoningonPartially comparable-32.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
149Gemma 4 26B A4BOpen weightsGoogle · Gemma 4 · best of 2 rows43.6%IndependentreasoningonPartially comparable-34.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
149qwen3-coder-480b-a35b-instructOpen weightsAlibaba Group · Qwen3-Coder43.6%IndependentreasoningoffPartially comparable-34.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
151GLM 4.5Open weightsZ.ai (Zhipu AI) · GLM4.543.0%IndependentreasoningonPartially comparable-35.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
152granite-4.1-30bOpen weightsIBM · Granite 4.142.1%IndependentreasoningoffPartially comparable-36.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
153Qwen3 Next 80B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows41.5%IndependentreasoningonPartially comparable-36.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
154Mistral Small 4Open weightsMistral AI · Mistral · best of 2 rows41.2%IndependentreasoningonPartially comparable-36.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
155Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 2 rows40.9%IndependentreasoningonPartially comparable-37.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
156Mistral Medium 3.1ClosedMistral AI · Mistral40.6%IndependentreasoningoffPartially comparable-37.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
157nova-premierClosedAmazon Web Services · Nova38.3%IndependentreasoningoffPartially comparable-39.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
158Devstral Small 1.0Open weightsMistral AI · Devstral38.0%IndependentreasoningoffPartially comparable-40.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
159DeepSeek V3.1Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows37.4%IndependentreasoningonPartially comparable-40.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
159North Mini Code (free)Open weightsCohere37.4%IndependentreasoningonPartially comparable-40.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
161DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 2 rows37.1%IndependentreasoningoffPartially comparable-40.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
162Pixtral LargeOpen weightsMistral AI · Pixtral36.5%IndependentreasoningoffPartially comparable-41.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
162deepseek-r1Open weightsDeepSeek · DeepSeek-R136.5%IndependentreasoningonPartially comparable-41.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
162gpt-5-nanoClosedOpenAI · GPT 5 · best of 3 rows36.5%Independentreasoning_efforthighPartially comparable-41.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
165gemma-4-12BOpen weightsGoogle · Gemma 4 · best of 2 rows36.3%IndependentreasoningonPartially comparable-41.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
166Qwen2.5 72B InstructOpen weightsQwen · Qwen2.534.5%IndependentreasoningoffPartially comparable-43.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
166Qwen3 Coder 30B A3B InstructOpen weightsQwen · Qwen334.5%IndependentreasoningoffPartially comparable-43.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
166qwen3-14b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows34.5%IndependentreasoningonPartially comparable-43.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
166sarvam-30bOpen weightsSarvam34.5%Independentreasoning_efforthighPartially comparable-43.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
170MiniMax-M1-80kOpen weightsMiniMax · MiniMax34.2%IndependentreasoningonPartially comparable-43.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
171DeepSeek V3.2 ExpOpen weightsDeepSeek · DeepSeek-V3 · best of 2 rows33.9%IndependentreasoningonPartially comparable-44.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
172Mistral Large 2.0Open weightsMistral AI · Mistral33.0%IndependentreasoningoffPartially comparable-45.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
173ling-1tOpen weightsinclusionAI32.8%IndependentreasoningoffPartially comparable-45.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
173qwen3-max-previewClosedAlibaba Group · Qwen332.8%IndependentreasoningoffPartially comparable-45.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
175solar-pro-2ClosedUpstage · Solar · best of 2 rows31.9%IndependentreasoningoffPartially comparable-46.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows31.6%IndependentreasoningonPartially comparable-46.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176MiniMax-M1-40kOpen weightsMiniMax · MiniMax31.6%IndependentreasoningonPartially comparable-46.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
178Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.131.3%IndependentreasoningonPartially comparable-46.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
178o3-miniClosedOpenAI · OpenAI o-series · best of 2 rows31.3%Independentreasoning_efforthighPartially comparable-46.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
180Gemini 2.5 Flash-LiteClosedGoogle · Gemini 2.5 · best of 3 rows30.7%IndependentreasoningonPartially comparable-47.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
180mistral-large-2Open weightsMistral AI · Mistral30.7%IndependentreasoningoffPartially comparable-47.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
182gemini-2-5-flash-lite-preview-09-2025ClosedGoogle · Gemini 2.530.4%IndependentreasoningoffPartially comparable-47.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
183qwen3-32b-instructOpen weightsAlibaba Group · Qwen329.8%IndependentreasoningonPartially comparable-48.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
184Gemini 2.0 FlashClosedGoogle · Gemini 2.029.5%IndependentreasoningoffPartially comparable-48.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
184Mistral Small 3.2Open weightsMistral AI · Mistral29.5%IndependentreasoningoffPartially comparable-48.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 2 rows29.2%IndependentreasoningoffPartially comparable-48.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
187GPT-4o (2024-08-06)ClosedOpenAI · GPT 428.9%IndependentreasoningoffPartially comparable-49.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
188devstral-smallOpen weightsMistral AI · Devstral28.4%IndependentreasoningoffPartially comparable-49.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
189llama-nemotron-super-49b-v1-5Open weightsNVIDIA · Llama · best of 2 rows28.1%IndependentreasoningonPartially comparable-50.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
189nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 328.1%IndependentreasoningonPartially comparable-50.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
189qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 2 rows28.1%IndependentreasoningonPartially comparable-50.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
192Magistral Small 1.2Open weightsMistral AI · Magistral27.8%IndependentreasoningonPartially comparable-50.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
192falcon-h1r-7bOpen weightsTII UAE · Falcon27.8%IndependentreasoningonPartially comparable-50.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
192granite-4.1-8bOpen weightsIBM · Granite 4.127.8%IndependentreasoningoffPartially comparable-50.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
192k2-v2Open weightsMBZUAI Institute of Foundation Models · best of 3 rows27.8%Independentreasoning_efforthighPartially comparable-50.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
192qwen3-8b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows27.8%IndependentreasoningonPartially comparable-50.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
197Ministral 3 14BOpen weightsMistral AI · Ministral 327.2%IndependentreasoningoffPartially comparable-50.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
197qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows27.2%IndependentreasoningoffPartially comparable-50.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
199llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.326.9%IndependentreasoningonPartially comparable-51.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
200Llama 3.3 70BOpen weightsMeta AI · Llama 3.326.6%IndependentreasoningoffPartially comparable-51.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →