Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · tau-bench · variant τ²

τ²-bench

dual-control tool-agent-user interaction (telecom, retail, airline)

data quality51

Updated 3 h ago · first seen 11 Sept 2026

Metric
pass^1 · %
Current results
880
Models
326
Current leader
Z.ai GLM 5.2 99.1%

Score history · solar-pro-2 4 rows

Score history for solar-pro-20%10%20%30%40%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • solar-pro-2
  • 31.87%aa_slug=solar-pro-2 · variant=Telecom · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
  • 28.07%aa_slug=solar-pro-2-reasoning · variant=Telecom · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
  • 31.87%aa_slug=solar-pro-2 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 28.07%aa_slug=solar-pro-2-reasoning · evaluator=Artificial Analysis · reasoning=on · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · pass^1 · evaluator=Artificial Analysis

2 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 99.1%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
  2. 90.3%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 323 models · trust independent-evaluator

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
199qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 2 rows28.1%IndependentreasoningonPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
202Magistral Small 1.2Open weightsMistral AI · Magistral27.8%Independentgroup defaultsComparable-0.29 ptobs. 11 Sept 2026artificialanalysis.aiT2History
202Qwen3 8BOpen weightsQwen · Qwen327.8%Independentgroup defaultsComparable-0.29 ptobs. 11 Sept 2026artificialanalysis.aiT2History
202falcon-h1r-7bOpen weightsTII UAE · Falcon27.8%Independentgroup defaultsComparable-0.29 ptobs. 11 Sept 2026artificialanalysis.aiT2History
202granite-4.1-8bOpen weightsIBM · Granite 4.127.8%Independentgroup defaultsComparable-0.29 ptobs. 11 Sept 2026artificialanalysis.aiT2History
202k2-v2Open weightsMBZUAI Institute of Foundation Models · best of 3 rows27.8%Independentgroup defaultsComparable-0.29 ptobs. 11 Sept 2026artificialanalysis.aiT2History
207Ministral 3 14BOpen weightsMistral AI · Ministral 327.2%Independentgroup defaultsComparable-0.88 ptobs. 11 Sept 2026artificialanalysis.aiT2History
207qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows27.2%Independentgroup defaultsComparable-0.88 ptobs. 11 Sept 2026artificialanalysis.aiT2History
209llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.326.9%IndependentreasoningonPartially comparable-1.17 ptobs. 11 Sept 2026artificialanalysis.aiT2History
210Llama 3.3 70BOpen weightsMeta AI · Llama 3.326.6%Independentgroup defaultsComparable-1.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
210Ministral 3 8BOpen weightsMistral AI · Ministral 326.6%Independentgroup defaultsComparable-1.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
210hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.126.6%Independentgroup defaultsComparable-1.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
210intellect-3Open weightsPrime Intellect26.6%Independentgroup defaultsComparable-1.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
210magistral-smallOpen weightsMistral AI · Magistral26.6%Independentgroup defaultsComparable-1.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
210qwen3-4b-2507-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows26.6%Independentgroup defaultsComparable-1.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
216ring-1tOpen weightsinclusionAI26.3%Independentgroup defaultsComparable-1.75 ptobs. 11 Sept 2026artificialanalysis.aiT2History
217gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 2 rows26.0%IndependentreasoningoffPartially comparable-2.05 ptobs. 11 Sept 2026artificialanalysis.aiT2History
217qwen3-1.7b-instructOpen weightsAlibaba Group · Qwen3.1 · best of 2 rows26.0%IndependentreasoningonPartially comparable-2.05 ptobs. 11 Sept 2026artificialanalysis.aiT2History
217qwen3-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows26.0%IndependentreasoningonPartially comparable-2.05 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220k2-think-v2Open weightsMBZUAI Institute of Foundation Models25.4%Independentgroup defaultsComparable-2.63 ptobs. 11 Sept 2026artificialanalysis.aiT2History
221Mistral Small 3.1Open weightsMistral AI · Mistral25.1%Independentgroup defaultsComparable-2.92 ptobs. 11 Sept 2026artificialanalysis.aiT2History
221gpt-4oClosedOpenAI · GPT 425.1%Independentgroup defaultsComparable-2.92 ptobs. 11 Sept 2026artificialanalysis.aiT2History
223Devstral 2Open weightsMistral AI · Devstral 224.9%Independentgroup defaultsComparable-3.22 ptobs. 11 Sept 2026artificialanalysis.aiT2History
223Ministral 3 3BOpen weightsMistral AI · Ministral 324.9%Independentgroup defaultsComparable-3.22 ptobs. 11 Sept 2026artificialanalysis.aiT2History
223qwen3-8b-instructOpen weightsAlibaba Group · Qwen324.9%Independentgroup defaultsComparable-3.22 ptobs. 11 Sept 2026artificialanalysis.aiT2History
226Claude Haiku 3.5ClosedAnthropic · Claude24.6%Independentgroup defaultsComparable-3.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
226Mistral Large 3Open weightsMistral AI · Mistral24.6%Independentgroup defaultsComparable-3.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
228Mistral Medium 3ClosedMistral AI · Mistral24.3%Independentgroup defaultsComparable-3.80 ptobs. 11 Sept 2026artificialanalysis.aiT2History
229Devstral Small 2Open weightsMistral AI · Devstral23.4%Independentgroup defaultsComparable-4.68 ptobs. 11 Sept 2026artificialanalysis.aiT2History
229NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 2 rows23.4%Independentgroup defaultsComparable-4.68 ptobs. 11 Sept 2026artificialanalysis.aiT2History
229Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 2 rows23.4%Independentgroup defaultsComparable-4.68 ptobs. 11 Sept 2026artificialanalysis.aiT2History
232llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.123.1%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
232magistral-mediumClosedMistral AI · Magistral23.1%Independentgroup defaultsComparable-4.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
234granite-4-0-nano-1bOpen weightsIBM · Granite 4.022.8%Independentgroup defaultsComparable-5.26 ptobs. 11 Sept 2026artificialanalysis.aiT2History
235GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows22.5%Independentgroup defaultsComparable-5.56 ptobs. 11 Sept 2026artificialanalysis.aiT2History
235Hermes-4-70BRestricted weightsNous Research · Hermes 422.5%Independentgroup defaultsComparable-5.56 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237Hermes 4 405BOpen weightsNous Research · Hermes 422.2%Independentgroup defaultsComparable-5.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 2 rows22.2%IndependentreasoningoffPartially comparable-5.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
239R1 Distill Llama 70BOpen weightsDeepSeek · Llama21.9%Independentgroup defaultsComparable-6.14 ptobs. 11 Sept 2026artificialanalysis.aiT2History
240hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.121.6%Independentgroup defaultsComparable-6.43 ptobs. 11 Sept 2026artificialanalysis.aiT2History
240nanbeige4-1-3bOpen weightsNanbeige21.6%Independentgroup defaultsComparable-6.43 ptobs. 11 Sept 2026artificialanalysis.aiT2History
242nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 2 rows21.4%IndependentreasoningonPartially comparable-6.72 ptobs. 11 Sept 2026artificialanalysis.aiT2History
242olmo-3-1-32b-instructOpen weightsAllen Institute for AI · OLMo 3.1 · best of 2 rows21.4%Independentgroup defaultsComparable-6.72 ptobs. 11 Sept 2026artificialanalysis.aiT2History
242qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows21.4%IndependentreasoningonPartially comparable-6.72 ptobs. 11 Sept 2026artificialanalysis.aiT2History
245Claude 3 HaikuClosedAnthropic · Claude21.1%Independentgroup defaultsComparable-7.02 ptobs. 11 Sept 2026artificialanalysis.aiT2History
245Llama-3.2-3BRestricted weightsMeta AI · Llama 3.221.1%Independentgroup defaultsComparable-7.02 ptobs. 11 Sept 2026artificialanalysis.aiT2History
245Qwen3-0.6BOpen weightsQwen · Qwen3.021.1%Independentgroup defaultsComparable-7.02 ptobs. 11 Sept 2026artificialanalysis.aiT2History
248ling-flash-2-0Open weightsinclusionAI20.8%Independentgroup defaultsComparable-7.31 ptobs. 11 Sept 2026artificialanalysis.aiT2History
249exaone-4-0-1-2bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows20.5%Independentgroup defaultsComparable-7.60 ptobs. 11 Sept 2026artificialanalysis.aiT2History
250solar-miniOpen weightsUpstage · Solar20.2%Independentgroup defaultsComparable-7.89 ptobs. 11 Sept 2026artificialanalysis.aiT2History
251Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows19.9%IndependentreasoningonPartially comparable-8.19 ptobs. 11 Sept 2026artificialanalysis.aiT2History
251devstral-mediumClosedMistral AI · Devstral19.9%Independentgroup defaultsComparable-8.19 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253Mistral Small 3Open weightsMistral AI · Mistral19.6%Independentgroup defaultsComparable-8.48 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.019.6%Independentgroup defaultsComparable-8.48 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253granite-4.1-3bOpen weightsIBM · Granite 4.119.6%Independentgroup defaultsComparable-8.48 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253lfm2-5-1-2b-thinkingOpen weightsLiquid AI · LFM2.519.6%Independentgroup defaultsComparable-8.48 ptobs. 11 Sept 2026artificialanalysis.aiT2History
257Gemini 2.5 Flash-LiteClosedGoogle · Gemini 2.5 · best of 2 rows19.0%Independentgroup defaultsComparable-9.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
257Llama-3.1-405BRestricted weightsMeta AI · Llama 3.119.0%Independentgroup defaultsComparable-9.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
257Qwen3-4BOpen weightsQwen · Qwen319.0%Independentgroup defaultsComparable-9.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
260Llama 4 MaverickOpen weightsMeta AI · Llama 417.8%Independentgroup defaultsComparable-10.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
261nova-liteClosedAmazon Web Services · Nova17.5%Independentgroup defaultsComparable-10.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows17.3%IndependentreasoningonPartially comparable-10.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262gpt-4.1-nanoClosedOpenAI · GPT 4.117.3%Independentgroup defaultsComparable-10.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262granite-4-0-h-smallOpen weightsIBM · Granite 4.017.3%Independentgroup defaultsComparable-10.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
265Llama 3.1 8BRestricted weightsMeta AI · Llama 3.116.4%Independentgroup defaultsComparable-11.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
266LFM2.5-8B-A1BOpen weightsLiquid AI · LFM2.516.1%Independentgroup defaultsComparable-12.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
266step-3-vl-10bOpen weightsStepFun · Step316.1%Independentgroup defaultsComparable-12.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
268jamba-reasoning-3bOpen weightsAI21 Labs · Jamba15.8%Independentgroup defaultsComparable-12.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
269Llama 4 ScoutOpen weightsMeta AI · Llama 415.5%Independentgroup defaultsComparable-12.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
270Command AOpen weightsCohere · Command15.2%Independentgroup defaultsComparable-12.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
270Llama-3.1-70BOpen weightsMeta AI · Llama 3.115.2%Independentgroup defaultsComparable-12.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
272granite-4-0-h-350mOpen weightsIBM · Granite 4.014.6%Independentgroup defaultsComparable-13.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
272llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.214.6%Independentgroup defaultsComparable-13.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
272qwen3-0.6b-instructOpen weightsAlibaba Group · Qwen3.014.6%Independentgroup defaultsComparable-13.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
275nova-microClosedAmazon Web Services · Nova14.0%Independentgroup defaultsComparable-14.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
275nova-proClosedAmazon Web Services · Nova14.0%Independentgroup defaultsComparable-14.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
277jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.713.4%Independentgroup defaultsComparable-14.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
277lfm2-2-6bOpen weightsLiquid AI · LFM2.213.4%Independentgroup defaultsComparable-14.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
279granite-4-0-350mOpen weightsIBM · Granite 4.013.2%Independentgroup defaultsComparable-14.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
279ling-mini-2-0Open weightsinclusionAI13.2%Independentgroup defaultsComparable-14.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
281apertus-70b-instructOpen weightsSwiss AI Initiative12.9%Independentgroup defaultsComparable-15.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
282Granite 4.0 MicroOpen weightsIBM · Granite 4.012.6%Independentgroup defaultsComparable-15.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
282LFM2-1.2BOpen weightsLiquid AI · LFM2.112.6%Independentgroup defaultsComparable-15.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
282Olmo-3-7B-InstructOpen weightsAllen Institute for AI · OLMo 312.6%Independentgroup defaultsComparable-15.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
282jamba-1-7-miniOpen weightsAI21 Labs · Jamba 1.712.6%Independentgroup defaultsComparable-15.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
286llama-3-1-nemotron-nano-4b-reasoningOpen weightsNVIDIA · Llama 3.111.7%Independentgroup defaultsComparable-16.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
287apertus-8b-instructOpen weightsSwiss AI Initiative11.4%Independentgroup defaultsComparable-16.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
287deepseek-r1-0120Open weightsDeepSeek · DeepSeek11.4%Independentgroup defaultsComparable-16.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
287llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.111.4%Independentgroup defaultsComparable-16.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
290lfm2-24b-a2bOpen weightsLiquid AI · LFM211.1%Independentgroup defaultsComparable-17.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
291Gemma 3 12BOpen weightsGoogle · Gemma 310.8%Independentgroup defaultsComparable-17.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
291LFM2.5-1.2B-InstructOpen weightsLiquid AI · LFM2.510.8%Independentgroup defaultsComparable-17.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
293Gemma 3 27BOpen weightsGoogle · Gemma 310.5%Independentgroup defaultsComparable-17.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
293gemma-3-1bOpen weightsGoogle · Gemma 310.5%Independentgroup defaultsComparable-17.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
293granite-3-3-8b-instructOpen weightsIBM · Granite 3.310.5%Independentgroup defaultsComparable-17.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
293lfm2-8b-a1bOpen weightsLiquid AI · LFM210.5%Independentgroup defaultsComparable-17.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
297gemma-3-270mOpen weightsGoogle · Gemma 39.06%Independentgroup defaultsComparable-19.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
298LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.58.48%Independentgroup defaultsComparable-19.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
299phi-4-miniOpen weightsMicrosoft · Phi48.19%Independentgroup defaultsComparable-19.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
300Gemma 3 4BRestricted weightsGoogle · Gemma 34.97%Independentgroup defaultsComparable-23.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →