Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · tau-bench · variant τ²

τ²-bench

dual-control tool-agent-user interaction (telecom, retail, airline)

data quality51

Updated 4 h ago · first seen 11 Sept 2026

Metric
pass^1 · %
Current results
880
Models
326
Current leader
Z.ai GLM 5.2 99.1%

Frontier over time · pass^1 · variant=Telecom · evaluator=Artificial Analysis

2 leader changes recorded, all dated 12 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 99.1%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent12 Sept 2026
  2. 90.3%grok-3-mini-reasoning SpaceXAI Independent12 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 308 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
200Ministral 3 8BOpen weightsMistral AI · Ministral 326.6%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
200hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.1 · best of 2 rows26.6%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
200intellect-3Open weightsPrime Intellect26.6%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
200magistral-smallOpen weightsMistral AI · Magistral26.6%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
200qwen3-4b-2507-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows26.6%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
206ring-1tOpen weightsinclusionAI26.3%IndependentreasoningonPartially comparable-0.29 ptobs. 12 Sept 2026artificialanalysis.aiT2History
207gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 2 rows26.0%IndependentreasoningoffPartially comparable-0.59 ptobs. 12 Sept 2026artificialanalysis.aiT2History
207qwen3-1.7b-instructOpen weightsAlibaba Group · Qwen3.1 · best of 2 rows26.0%IndependentreasoningonPartially comparable-0.59 ptobs. 12 Sept 2026artificialanalysis.aiT2History
207qwen3-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows26.0%IndependentreasoningonPartially comparable-0.59 ptobs. 12 Sept 2026artificialanalysis.aiT2History
210k2-think-v2Open weightsMBZUAI Institute of Foundation Models25.4%IndependentreasoningonPartially comparable-1.17 ptobs. 12 Sept 2026artificialanalysis.aiT2History
211Mistral Small 3.1Open weightsMistral AI · Mistral25.1%IndependentreasoningoffPartially comparable-1.46 ptobs. 12 Sept 2026artificialanalysis.aiT2History
211gpt-4oClosedOpenAI · GPT 425.1%IndependentreasoningoffPartially comparable-1.46 ptobs. 12 Sept 2026artificialanalysis.aiT2History
213Devstral 2Open weightsMistral AI · Devstral 224.9%IndependentreasoningoffPartially comparable-1.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
213Ministral 3 3BOpen weightsMistral AI · Ministral 324.9%IndependentreasoningoffPartially comparable-1.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
215Claude Haiku 3.5ClosedAnthropic · Claude24.6%IndependentreasoningoffPartially comparable-2.05 ptobs. 12 Sept 2026artificialanalysis.aiT2History
215Mistral Large 3Open weightsMistral AI · Mistral24.6%IndependentreasoningoffPartially comparable-2.05 ptobs. 12 Sept 2026artificialanalysis.aiT2History
217Mistral Medium 3ClosedMistral AI · Mistral24.3%IndependentreasoningoffPartially comparable-2.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
218Devstral Small 2Open weightsMistral AI · Devstral23.4%IndependentreasoningoffPartially comparable-3.22 ptobs. 12 Sept 2026artificialanalysis.aiT2History
218NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 2 rows23.4%IndependentreasoningoffPartially comparable-3.22 ptobs. 12 Sept 2026artificialanalysis.aiT2History
218Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 2 rows23.4%IndependentreasoningoffPartially comparable-3.22 ptobs. 12 Sept 2026artificialanalysis.aiT2History
221llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.123.1%IndependentreasoningoffPartially comparable-3.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
221magistral-mediumClosedMistral AI · Magistral23.1%IndependentreasoningonPartially comparable-3.51 ptobs. 12 Sept 2026artificialanalysis.aiT2History
223granite-4-0-nano-1bOpen weightsIBM · Granite 4.022.8%IndependentreasoningoffPartially comparable-3.80 ptobs. 12 Sept 2026artificialanalysis.aiT2History
224GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows22.5%IndependentreasoningonPartially comparable-4.10 ptobs. 12 Sept 2026artificialanalysis.aiT2History
224hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.1 · best of 2 rows22.5%IndependentreasoningonPartially comparable-4.10 ptobs. 12 Sept 2026artificialanalysis.aiT2History
226gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 2 rows22.2%IndependentreasoningoffPartially comparable-4.39 ptobs. 12 Sept 2026artificialanalysis.aiT2History
227R1 Distill Llama 70BOpen weightsDeepSeek · Llama21.9%IndependentreasoningonPartially comparable-4.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
228nanbeige4-1-3bOpen weightsNanbeige21.6%IndependentreasoningonPartially comparable-4.97 ptobs. 12 Sept 2026artificialanalysis.aiT2History
229nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 2 rows21.4%IndependentreasoningonPartially comparable-5.26 ptobs. 12 Sept 2026artificialanalysis.aiT2History
229olmo-3-1-32b-instructOpen weightsAllen Institute for AI · OLMo 3.1 · best of 2 rows21.4%IndependentreasoningoffPartially comparable-5.26 ptobs. 12 Sept 2026artificialanalysis.aiT2History
229qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows21.4%IndependentreasoningonPartially comparable-5.26 ptobs. 12 Sept 2026artificialanalysis.aiT2History
232Claude 3 HaikuClosedAnthropic · Claude21.1%IndependentreasoningoffPartially comparable-5.56 ptobs. 12 Sept 2026artificialanalysis.aiT2History
232Llama-3.2-3BRestricted weightsMeta AI · Llama 3.221.1%IndependentreasoningoffPartially comparable-5.56 ptobs. 12 Sept 2026artificialanalysis.aiT2History
232qwen3-0.6b-instructOpen weightsAlibaba Group · Qwen3.0 · best of 2 rows21.1%IndependentreasoningonPartially comparable-5.56 ptobs. 12 Sept 2026artificialanalysis.aiT2History
235ling-flash-2-0Open weightsinclusionAI20.8%IndependentreasoningoffPartially comparable-5.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
236exaone-4-0-1-2bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows20.5%IndependentreasoningoffPartially comparable-6.14 ptobs. 12 Sept 2026artificialanalysis.aiT2History
237solar-miniOpen weightsUpstage · Solar20.2%IndependentreasoningoffPartially comparable-6.43 ptobs. 12 Sept 2026artificialanalysis.aiT2History
238Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows19.9%IndependentreasoningonPartially comparable-6.73 ptobs. 12 Sept 2026artificialanalysis.aiT2History
238devstral-mediumClosedMistral AI · Devstral19.9%IndependentreasoningoffPartially comparable-6.73 ptobs. 12 Sept 2026artificialanalysis.aiT2History
240Mistral Small 3Open weightsMistral AI · Mistral19.6%IndependentreasoningoffPartially comparable-7.02 ptobs. 12 Sept 2026artificialanalysis.aiT2History
240granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.019.6%IndependentreasoningoffPartially comparable-7.02 ptobs. 12 Sept 2026artificialanalysis.aiT2History
240granite-4.1-3bOpen weightsIBM · Granite 4.119.6%IndependentreasoningoffPartially comparable-7.02 ptobs. 12 Sept 2026artificialanalysis.aiT2History
240lfm2-5-1-2b-thinkingOpen weightsLiquid AI · LFM2.519.6%IndependentreasoningonPartially comparable-7.02 ptobs. 12 Sept 2026artificialanalysis.aiT2History
244Llama-3.1-405BRestricted weightsMeta AI · Llama 3.119.0%IndependentreasoningoffPartially comparable-7.60 ptobs. 12 Sept 2026artificialanalysis.aiT2History
244qwen3-4b-instructOpen weightsAlibaba Group · Qwen319.0%IndependentreasoningonPartially comparable-7.60 ptobs. 12 Sept 2026artificialanalysis.aiT2History
246Llama 4 MaverickOpen weightsMeta AI · Llama 417.8%IndependentreasoningoffPartially comparable-8.77 ptobs. 12 Sept 2026artificialanalysis.aiT2History
247nova-liteClosedAmazon Web Services · Nova17.5%IndependentreasoningoffPartially comparable-9.07 ptobs. 12 Sept 2026artificialanalysis.aiT2History
248exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows17.3%IndependentreasoningonPartially comparable-9.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
248gpt-4.1-nanoClosedOpenAI · GPT 4.117.3%IndependentreasoningoffPartially comparable-9.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
248granite-4-0-h-smallOpen weightsIBM · Granite 4.017.3%IndependentreasoningoffPartially comparable-9.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
251Llama 3.1 8BRestricted weightsMeta AI · Llama 3.116.4%IndependentreasoningoffPartially comparable-10.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
252LFM2.5-8B-A1BOpen weightsLiquid AI · LFM2.516.1%IndependentreasoningonPartially comparable-10.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
252step-3-vl-10bOpen weightsStepFun · Step316.1%IndependentreasoningonPartially comparable-10.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
254jamba-reasoning-3bOpen weightsAI21 Labs · Jamba15.8%IndependentreasoningonPartially comparable-10.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
255Llama 4 ScoutOpen weightsMeta AI · Llama 415.5%IndependentreasoningoffPartially comparable-11.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
256Command AOpen weightsCohere · Command15.2%IndependentreasoningoffPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
256Llama-3.1-70BOpen weightsMeta AI · Llama 3.115.2%IndependentreasoningoffPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
258granite-4-0-h-350mOpen weightsIBM · Granite 4.014.6%IndependentreasoningoffPartially comparable-12.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
258llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.214.6%IndependentreasoningoffPartially comparable-12.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
260nova-microClosedAmazon Web Services · Nova14.0%IndependentreasoningoffPartially comparable-12.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
260nova-proClosedAmazon Web Services · Nova14.0%IndependentreasoningoffPartially comparable-12.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
262jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.713.4%IndependentreasoningoffPartially comparable-13.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
262lfm2-2-6bOpen weightsLiquid AI · LFM2.213.4%IndependentreasoningoffPartially comparable-13.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
264granite-4-0-350mOpen weightsIBM · Granite 4.013.2%IndependentreasoningoffPartially comparable-13.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
264ling-mini-2-0Open weightsinclusionAI13.2%IndependentreasoningoffPartially comparable-13.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
266apertus-70b-instructOpen weightsSwiss AI Initiative12.9%IndependentreasoningoffPartially comparable-13.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
267Granite 4.0 MicroOpen weightsIBM · Granite 4.012.6%IndependentreasoningoffPartially comparable-14.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
267LFM2-1.2BOpen weightsLiquid AI · LFM2.112.6%IndependentreasoningoffPartially comparable-14.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
267Olmo-3-7B-InstructOpen weightsAllen Institute for AI · OLMo 312.6%IndependentreasoningoffPartially comparable-14.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
267jamba-1-7-miniOpen weightsAI21 Labs · Jamba 1.712.6%IndependentreasoningoffPartially comparable-14.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
271llama-3-1-nemotron-nano-4b-reasoningOpen weightsNVIDIA · Llama 3.111.7%IndependentreasoningonPartially comparable-14.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
272apertus-8b-instructOpen weightsSwiss AI Initiative11.4%IndependentreasoningoffPartially comparable-15.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
272deepseek-r1-0120Open weightsDeepSeek · DeepSeek11.4%IndependentreasoningonPartially comparable-15.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
272llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.111.4%IndependentreasoningonPartially comparable-15.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
275lfm2-24b-a2bOpen weightsLiquid AI · LFM211.1%IndependentreasoningoffPartially comparable-15.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
276Gemma 3 12BOpen weightsGoogle · Gemma 310.8%IndependentreasoningoffPartially comparable-15.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
276LFM2.5-1.2B-InstructOpen weightsLiquid AI · LFM2.510.8%IndependentreasoningoffPartially comparable-15.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
278Gemma 3 27BOpen weightsGoogle · Gemma 310.5%IndependentreasoningoffPartially comparable-16.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
278gemma-3-1bOpen weightsGoogle · Gemma 310.5%IndependentreasoningoffPartially comparable-16.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
278granite-3-3-8b-instructOpen weightsIBM · Granite 3.310.5%IndependentreasoningoffPartially comparable-16.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
278lfm2-8b-a1bOpen weightsLiquid AI · LFM210.5%IndependentreasoningoffPartially comparable-16.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
282gemma-3-270mOpen weightsGoogle · Gemma 39.06%IndependentreasoningoffPartially comparable-17.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
283LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.58.48%IndependentreasoningoffPartially comparable-18.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
284phi-4-miniOpen weightsMicrosoft · Phi48.19%IndependentreasoningoffPartially comparable-18.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
285Gemma 3 4BRestricted weightsGoogle · Gemma 34.97%IndependentreasoningoffPartially comparable-21.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
285gemma-3n-e4bOpen weightsGoogle · Gemma 34.97%IndependentreasoningoffPartially comparable-21.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287DeepSeek-R1-0528-Qwen3-8BOpen weightsDeepSeek · Qwen30%IndependentreasoningonPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AI · Kimi0%IndependentreasoningoffPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287Llama-3.2-1BRestricted weightsMeta AI · Llama 3.20%IndependentreasoningoffPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287Mistral 7BOpen weightsMistral AI · Mistral0%IndependentreasoningoffPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287Molmo2-8BOpen weightsAllen Institute for AI0%IndependentreasoningoffPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287Olmo-3-7B-ThinkOpen weightsAllen Institute for AI · OLMo 30%IndependentreasoningonPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287Phi 4Open weightsMicrosoft · Phi40%IndependentreasoningoffPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287Reka Flash 3Open weightsrekaai0%IndependentreasoningonPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287deepseek-v3-2-specialeOpen weightsDeepSeek · DeepSeek0%IndependentreasoningonPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287ernie-4-5-300b-a47bOpen weightsBaidu · ERNIE 4.50%IndependentreasoningoffPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287gemma-3n-e2bOpen weightsGoogle · Gemma 30%IndependentreasoningoffPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287gpt-5-chatgptClosedOpenAI · GPT 50%IndependentreasoningoffPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287llama-3-instruct-70bOpen weightsMeta AI · Llama 30%IndependentreasoningoffPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287llama-3-instruct-8bOpen weightsMeta AI · Llama 30%IndependentreasoningoffPartially comparable-26.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →