Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · tau-bench · variant τ²

τ²-bench

dual-control tool-agent-user interaction (telecom, retail, airline)

data quality51

Updated 2 h ago · first seen 11 Sept 2026

Metric
pass^1 · %
Current results
880
Models
326
Current leader
Z.ai GLM 5.2 99.1%

Score history · Qwen3 VL 30B A3B Instruct 4 rows

Score history for Qwen3 VL 30B A3B Instruct0%5%10%15%20%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • Qwen3 VL 30B A3B Instruct
  • 19.01%aa_slug=qwen3-vl-30b-a3b-instruct · variant=Telecom · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
  • 19.88%aa_slug=qwen3-vl-30b-a3b-reasoning · variant=Telecom · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
  • 19.01%aa_slug=qwen3-vl-30b-a3b-instruct · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 19.88%aa_slug=qwen3-vl-30b-a3b-reasoning · evaluator=Artificial Analysis · reasoning=on · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · pass^1 · evaluator=Artificial Analysis

2 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 99.1%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
  2. 90.3%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 323 models · trust independent-evaluator

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
101Claude Sonnet 4.6ClosedAnthropic · Claude · best of 3 rows79.5%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
101Qwen3 Coder NextOpen weightsQwen · Qwen379.5%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
101longcat-flash-liteOpen weightsLongCat79.5%Independentgroup defaultsComparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
104DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 2 rows79.0%Independentgroup defaultsComparable-0.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
105Claude Sonnet 4.5ClosedAnthropic · Claude · best of 2 rows78.1%IndependentreasoningonPartially comparable-1.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
105exaone-4-5-33bOpen weightsLG AI Research · EXAONE 4.578.1%Independentgroup defaultsComparable-1.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
107GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows76.9%Independentgroup defaultsComparable-2.63 ptobs. 11 Sept 2026artificialanalysis.aiT2History
108gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 3 rows76.0%Independentgroup defaultsComparable-3.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
109Grok Build 0.1ClosedxAI · Grok75.7%Independentgroup defaultsComparable-3.80 ptobs. 11 Sept 2026artificialanalysis.aiT2History
109nova-2-0-liteClosedAmazon Web Services · Nova 2.0 · best of 4 rows75.7%Independentreasoningonreasoning_effortmediumPartially comparable-3.80 ptobs. 11 Sept 2026artificialanalysis.aiT2History
111grok-4ClosedSpaceXAI · Grok 474.8%Independentgroup defaultsComparable-4.68 ptobs. 11 Sept 2026artificialanalysis.aiT2History
112k-exaoneOpen weightsLG AI Research · EXAONE · best of 2 rows74.3%Independentgroup defaultsComparable-5.26 ptobs. 11 Sept 2026artificialanalysis.aiT2History
113Claude Opus 4ClosedAnthropic · Claude73.4%Independentgroup defaultsComparable-6.14 ptobs. 11 Sept 2026artificialanalysis.aiT2History
113Kimi K2 0905Open weightsMoonshot AI · Kimi73.4%Independentgroup defaultsComparable-6.14 ptobs. 11 Sept 2026artificialanalysis.aiT2History
115Claude Opus 4.1ClosedAnthropic · Claude71.4%IndependentreasoningonPartially comparable-8.16 ptobs. 11 Sept 2026artificialanalysis.aiT2History
116gpt-5-miniClosedOpenAI · GPT 5 · best of 3 rows71.0%Independentreasoning_effortmediumPartially comparable-8.48 ptobs. 11 Sept 2026artificialanalysis.aiT2History
117Mercury 2ClosedInception70.8%Independentgroup defaultsComparable-8.77 ptobs. 11 Sept 2026artificialanalysis.aiT2History
118apriel-v1-6-15b-thinkerOpen weightsServiceNow69.3%Independentgroup defaultsComparable-10.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
119apriel-v1-5-15b-thinkerOpen weightsServiceNow68.4%Independentgroup defaultsComparable-11.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
120Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 367.8%Independentgroup defaultsComparable-11.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
121Hy3Open weightsTencent67.5%IndependentreasoningoffPartially comparable-12.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
122gpt-oss-120bOpen weightsOpenAI · gpt-oss · best of 2 rows65.8%Independentgroup defaultsComparable-13.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
122grok-4-fastClosedSpaceXAI · Grok 4 · best of 2 rows65.8%IndependentreasoningonPartially comparable-13.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
124Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 2 rows65.5%IndependentreasoningoffPartially comparable-14.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
125Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 2 rows65.2%IndependentreasoningoffPartially comparable-14.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
126Claude Sonnet 4ClosedAnthropic · Claude · best of 2 rows64.6%IndependentreasoningonPartially comparable-14.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
127hypernova-60bOpen weightsMultiverse Computing63.2%Independentgroup defaultsComparable-16.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
128GPT-5.1-Codex MiniClosedOpenAI · GPT 5.162.9%Independentgroup defaultsComparable-16.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
129o1ClosedOpenAI · OpenAI o-series62.6%Independentgroup defaultsComparable-17.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
130Kimi K2 0711Open weightsMoonshot AI · Kimi61.1%Independentgroup defaultsComparable-18.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
131gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 2 rows60.2%Independentgroup defaultsComparable-19.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
132grok-4.20-0309-non-reasoningClosedxAI · Grok59.9%Independentgroup defaultsComparable-19.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
133doubao-seed-codeClosedByteDance · Seed58.2%Independentgroup defaultsComparable-21.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
134o4-miniClosedOpenAI · OpenAI o-series55.6%Independentgroup defaultsComparable-24.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
135Claude 3.7 SonnetClosedAnthropic · Claude · best of 2 rows54.7%IndependentreasoningonPartially comparable-24.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
135Claude Haiku 4.5ClosedAnthropic · Claude · best of 2 rows54.7%IndependentreasoningonPartially comparable-24.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
137Gemini 2.5 ProClosedGoogle · Gemini 2.554.1%Independentgroup defaultsComparable-25.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
137Qwen3 VL 235B A22B InstructOpen weightsQwen · Qwen3 · best of 2 rows54.1%IndependentreasoningonPartially comparable-25.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
139Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 2 rows53.2%Independentgroup defaultsComparable-26.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
139nemotron-cascade-2-30b-a3bOpen weightsNVIDIA · Nemotron53.2%Independentgroup defaultsComparable-26.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
141gpt-4.1-miniClosedOpenAI · GPT 4.152.9%Independentgroup defaultsComparable-26.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
142Magistral Medium 1.2ClosedMistral AI · Magistral52.0%Independentgroup defaultsComparable-27.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
143Seed-OSS-36B-InstructOpen weightsByteDance · Seed49.4%Independentgroup defaultsComparable-30.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
143gpt-5-5-instant-05-26ClosedOpenAI · GPT 5.549.4%Independentgroup defaultsComparable-30.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
143midm-250-pro-rsnsftClosedKorea Telecom49.4%Independentgroup defaultsComparable-30.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
146grok-3ClosedSpaceXAI · Grok 348.8%Independentgroup defaultsComparable-30.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
147solar-open-100b-reasoningOpen weightsUpstage · Solar48.3%Independentgroup defaultsComparable-31.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
148DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V347.1%Independentgroup defaultsComparable-32.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
148gpt-4.1ClosedOpenAI · GPT 4.147.1%Independentgroup defaultsComparable-32.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
150sarvam-105bOpen weightsSarvam46.8%Independentgroup defaultsComparable-32.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
151GLM 4.5 AirOpen weightsZ.ai (Zhipu AI) · GLM4.546.5%Independentgroup defaultsComparable-33.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
151motif-2-12-7bClosedMotif Technologies46.5%Independentgroup defaultsComparable-33.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
153Qwen3 VL 32B InstructOpen weightsQwen · Qwen3 · best of 2 rows45.6%IndependentreasoningonPartially comparable-33.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
153gemini-2-5-flash-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 2 rows45.6%IndependentreasoningonPartially comparable-33.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
155nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 345.3%Independentgroup defaultsComparable-34.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
156Gemma 4 26B A4BOpen weightsGoogle · Gemma 4 · best of 2 rows43.6%Independentgroup defaultsComparable-36.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
156qwen3-coder-480b-a35b-instructOpen weightsAlibaba Group · Qwen3-Coder43.6%Independentgroup defaultsComparable-36.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
158gemini-3-flashClosedGoogle · Gemini 343.3%Independentgroup defaultsComparable-36.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
159GLM 4.5Open weightsZ.ai (Zhipu AI) · GLM4.543.0%Independentgroup defaultsComparable-36.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
160granite-4.1-30bOpen weightsIBM · Granite 4.142.1%Independentgroup defaultsComparable-37.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
161Qwen3 Next 80B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows41.5%IndependentreasoningonPartially comparable-38.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
162Mistral Small 4Open weightsMistral AI · Mistral · best of 2 rows41.2%Independentgroup defaultsComparable-38.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
163Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 2 rows40.9%IndependentreasoningonPartially comparable-38.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
164Mistral Medium 3.1ClosedMistral AI · Mistral40.6%Independentgroup defaultsComparable-38.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
165nova-premierClosedAmazon Web Services · Nova38.3%Independentgroup defaultsComparable-41.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
166Devstral Small 1.0Open weightsMistral AI · Devstral38.0%Independentgroup defaultsComparable-41.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
167DeepSeek V3.1Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows37.4%IndependentreasoningonPartially comparable-42.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
167North Mini Code (free)Open weightsCohere37.4%Independentgroup defaultsComparable-42.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
169DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 2 rows37.1%Independentgroup defaultsComparable-42.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
170Pixtral LargeOpen weightsMistral AI · Pixtral36.5%Independentgroup defaultsComparable-43.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
170deepseek-r1Open weightsDeepSeek · DeepSeek-R136.5%Independentgroup defaultsComparable-43.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
170gpt-5-nanoClosedOpenAI · GPT 5 · best of 3 rows36.5%Independentgroup defaultsComparable-43.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
173gemma-4-12BOpen weightsGoogle · Gemma 4 · best of 2 rows36.3%Independentgroup defaultsComparable-43.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
174Qwen2.5 72B InstructOpen weightsQwen · Qwen2.534.5%Independentgroup defaultsComparable-45.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
174Qwen3 14BOpen weightsQwen · Qwen334.5%Independentgroup defaultsComparable-45.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
174Qwen3 Coder 30B A3B InstructOpen weightsQwen · Qwen334.5%Independentgroup defaultsComparable-45.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
174sarvam-30bOpen weightsSarvam34.5%Independentgroup defaultsComparable-45.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
178MiniMax-M1-80kOpen weightsMiniMax · MiniMax34.2%Independentgroup defaultsComparable-45.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
179DeepSeek V3.2 ExpOpen weightsDeepSeek · DeepSeek-V333.9%Independentgroup defaultsComparable-45.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
179deepseek-v3-2-0925Open weightsDeepSeek · DeepSeek33.9%Independentgroup defaultsComparable-45.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
181Mistral Large 2.0Open weightsMistral AI · Mistral33.0%Independentgroup defaultsComparable-46.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
182ling-1tOpen weightsinclusionAI32.8%Independentgroup defaultsComparable-46.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
182qwen3-max-previewClosedAlibaba Group · Qwen332.8%Independentgroup defaultsComparable-46.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184qwen3-14b-instructOpen weightsAlibaba Group · Qwen332.2%Independentgroup defaultsComparable-47.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
185solar-pro-2ClosedUpstage · Solar · best of 2 rows31.9%Independentgroup defaultsComparable-47.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
186GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows31.6%Independentgroup defaultsComparable-48.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
186Gemini 2.5 FlashClosedGoogle · Gemini 2.5 · best of 2 rows31.6%IndependentreasoningonPartially comparable-48.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
186MiniMax-M1-40kOpen weightsMiniMax · MiniMax31.6%Independentgroup defaultsComparable-48.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
189Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.131.3%Independentgroup defaultsComparable-48.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
189o3-miniClosedOpenAI · OpenAI o-series · best of 2 rows31.3%Independentreasoning_efforthighPartially comparable-48.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
191gemini-2-5-flash-lite-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 2 rows30.7%IndependentreasoningonPartially comparable-48.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
191mistral-large-2Open weightsMistral AI · Mistral30.7%Independentgroup defaultsComparable-48.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
193Qwen3 32BOpen weightsQwen · Qwen329.8%Independentgroup defaultsComparable-49.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
194Gemini 2.0 FlashClosedGoogle · Gemini 2.029.5%Independentgroup defaultsComparable-50.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
194Mistral Small 3.2Open weightsMistral AI · Mistral29.5%Independentgroup defaultsComparable-50.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
196Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 2 rows29.2%Independentgroup defaultsComparable-50.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
197GPT-4o (2024-08-06)ClosedOpenAI · GPT 428.9%Independentgroup defaultsComparable-50.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
198devstral-smallOpen weightsMistral AI · Devstral28.4%Independentgroup defaultsComparable-51.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
199llama-nemotron-super-49b-v1-5Open weightsNVIDIA · Llama · best of 2 rows28.1%IndependentreasoningonPartially comparable-51.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
199nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 328.1%Independentgroup defaultsComparable-51.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →