τ²-bench
dual-control tool-agent-user interaction (telecom, retail, airline)
Updated 3 h ago · first seen 11 Sept 2026
- Metric
- pass^1 · % ↑
- Current results
- 880
- Models
- 326
- Current leader
- Z.ai GLM 5.2 99.1%
Score history · qwen3-30b-a3b-2507 4 rows
- qwen3-30b-a3b-2507
- 28.07%aa_slug=qwen3-30b-a3b-2507-reasoning · variant=Telecom · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
- 10.23%aa_slug=qwen3-30b-a3b-2507 · variant=Telecom · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
- 28.07%aa_slug=qwen3-30b-a3b-2507-reasoning · evaluator=Artificial Analysis · reasoning=on · index_version=4.311 Sept 2026
- 10.23%aa_slug=qwen3-30b-a3b-2507 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
Frontier over time · pass^1 · evaluator=Artificial Analysis
2 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 99.1%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
- 90.3%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 323 models · trust independent-evaluator
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 199 | qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 2 rows | 28.1% | Independent | reasoningon | Partially comparable0.00 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 202 | Magistral Small 1.2Open weightsMistral AI · Magistral | 27.8% | Independent | group defaults | Comparable-0.29 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 202 | Qwen3 8BOpen weightsQwen · Qwen3 | 27.8% | Independent | group defaults | Comparable-0.29 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 202 | falcon-h1r-7bOpen weightsTII UAE · Falcon | 27.8% | Independent | group defaults | Comparable-0.29 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 202 | granite-4.1-8bOpen weightsIBM · Granite 4.1 | 27.8% | Independent | group defaults | Comparable-0.29 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 202 | k2-v2Open weightsMBZUAI Institute of Foundation Models · best of 3 rows | 27.8% | Independent | group defaults | Comparable-0.29 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 207 | Ministral 3 14BOpen weightsMistral AI · Ministral 3 | 27.2% | Independent | group defaults | Comparable-0.88 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 207 | qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows | 27.2% | Independent | group defaults | Comparable-0.88 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 209 | llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.3 | 26.9% | Independent | reasoningon | Partially comparable-1.17 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 210 | Llama 3.3 70BOpen weightsMeta AI · Llama 3.3 | 26.6% | Independent | group defaults | Comparable-1.46 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 210 | Ministral 3 8BOpen weightsMistral AI · Ministral 3 | 26.6% | Independent | group defaults | Comparable-1.46 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 210 | hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.1 | 26.6% | Independent | group defaults | Comparable-1.46 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 210 | intellect-3Open weightsPrime Intellect | 26.6% | Independent | group defaults | Comparable-1.46 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 210 | magistral-smallOpen weightsMistral AI · Magistral | 26.6% | Independent | group defaults | Comparable-1.46 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 210 | qwen3-4b-2507-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows | 26.6% | Independent | group defaults | Comparable-1.46 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 216 | ring-1tOpen weightsinclusionAI | 26.3% | Independent | group defaults | Comparable-1.75 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 217 | gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 2 rows | 26.0% | Independent | reasoningoff | Partially comparable-2.05 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 217 | qwen3-1.7b-instructOpen weightsAlibaba Group · Qwen3.1 · best of 2 rows | 26.0% | Independent | reasoningon | Partially comparable-2.05 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 217 | qwen3-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows | 26.0% | Independent | reasoningon | Partially comparable-2.05 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | k2-think-v2Open weightsMBZUAI Institute of Foundation Models | 25.4% | Independent | group defaults | Comparable-2.63 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 221 | Mistral Small 3.1Open weightsMistral AI · Mistral | 25.1% | Independent | group defaults | Comparable-2.92 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 221 | gpt-4oClosedOpenAI · GPT 4 | 25.1% | Independent | group defaults | Comparable-2.92 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 223 | Devstral 2Open weightsMistral AI · Devstral 2 | 24.9% | Independent | group defaults | Comparable-3.22 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 223 | Ministral 3 3BOpen weightsMistral AI · Ministral 3 | 24.9% | Independent | group defaults | Comparable-3.22 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 223 | qwen3-8b-instructOpen weightsAlibaba Group · Qwen3 | 24.9% | Independent | group defaults | Comparable-3.22 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 226 | Claude Haiku 3.5ClosedAnthropic · Claude | 24.6% | Independent | group defaults | Comparable-3.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 226 | Mistral Large 3Open weightsMistral AI · Mistral | 24.6% | Independent | group defaults | Comparable-3.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 228 | Mistral Medium 3ClosedMistral AI · Mistral | 24.3% | Independent | group defaults | Comparable-3.80 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 229 | Devstral Small 2Open weightsMistral AI · Devstral | 23.4% | Independent | group defaults | Comparable-4.68 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 229 | NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 2 rows | 23.4% | Independent | group defaults | Comparable-4.68 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 229 | Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 2 rows | 23.4% | Independent | group defaults | Comparable-4.68 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 232 | llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.1 | 23.1% | Independent | group defaults | Comparable-4.97 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 232 | magistral-mediumClosedMistral AI · Magistral | 23.1% | Independent | group defaults | Comparable-4.97 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 234 | granite-4-0-nano-1bOpen weightsIBM · Granite 4.0 | 22.8% | Independent | group defaults | Comparable-5.26 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 235 | GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows | 22.5% | Independent | group defaults | Comparable-5.56 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 235 | Hermes-4-70BRestricted weightsNous Research · Hermes 4 | 22.5% | Independent | group defaults | Comparable-5.56 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | Hermes 4 405BOpen weightsNous Research · Hermes 4 | 22.2% | Independent | group defaults | Comparable-5.85 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 2 rows | 22.2% | Independent | reasoningoff | Partially comparable-5.85 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 239 | R1 Distill Llama 70BOpen weightsDeepSeek · Llama | 21.9% | Independent | group defaults | Comparable-6.14 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 240 | hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.1 | 21.6% | Independent | group defaults | Comparable-6.43 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 240 | nanbeige4-1-3bOpen weightsNanbeige | 21.6% | Independent | group defaults | Comparable-6.43 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 242 | nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 2 rows | 21.4% | Independent | reasoningon | Partially comparable-6.72 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 242 | olmo-3-1-32b-instructOpen weightsAllen Institute for AI · OLMo 3.1 · best of 2 rows | 21.4% | Independent | group defaults | Comparable-6.72 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 242 | qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows | 21.4% | Independent | reasoningon | Partially comparable-6.72 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 245 | Claude 3 HaikuClosedAnthropic · Claude | 21.1% | Independent | group defaults | Comparable-7.02 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 245 | Llama-3.2-3BRestricted weightsMeta AI · Llama 3.2 | 21.1% | Independent | group defaults | Comparable-7.02 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 245 | Qwen3-0.6BOpen weightsQwen · Qwen3.0 | 21.1% | Independent | group defaults | Comparable-7.02 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 248 | ling-flash-2-0Open weightsinclusionAI | 20.8% | Independent | group defaults | Comparable-7.31 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 249 | exaone-4-0-1-2bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows | 20.5% | Independent | group defaults | Comparable-7.60 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 250 | solar-miniOpen weightsUpstage · Solar | 20.2% | Independent | group defaults | Comparable-7.89 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 251 | Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows | 19.9% | Independent | reasoningon | Partially comparable-8.19 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 251 | devstral-mediumClosedMistral AI · Devstral | 19.9% | Independent | group defaults | Comparable-8.19 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | Mistral Small 3Open weightsMistral AI · Mistral | 19.6% | Independent | group defaults | Comparable-8.48 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.0 | 19.6% | Independent | group defaults | Comparable-8.48 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | granite-4.1-3bOpen weightsIBM · Granite 4.1 | 19.6% | Independent | group defaults | Comparable-8.48 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | lfm2-5-1-2b-thinkingOpen weightsLiquid AI · LFM2.5 | 19.6% | Independent | group defaults | Comparable-8.48 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 257 | Gemini 2.5 Flash-LiteClosedGoogle · Gemini 2.5 · best of 2 rows | 19.0% | Independent | group defaults | Comparable-9.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 257 | Llama-3.1-405BRestricted weightsMeta AI · Llama 3.1 | 19.0% | Independent | group defaults | Comparable-9.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 257 | Qwen3-4BOpen weightsQwen · Qwen3 | 19.0% | Independent | group defaults | Comparable-9.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 260 | Llama 4 MaverickOpen weightsMeta AI · Llama 4 | 17.8% | Independent | group defaults | Comparable-10.2 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 261 | nova-liteClosedAmazon Web Services · Nova | 17.5% | Independent | group defaults | Comparable-10.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows | 17.3% | Independent | reasoningon | Partially comparable-10.8 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | gpt-4.1-nanoClosedOpenAI · GPT 4.1 | 17.3% | Independent | group defaults | Comparable-10.8 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | granite-4-0-h-smallOpen weightsIBM · Granite 4.0 | 17.3% | Independent | group defaults | Comparable-10.8 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 265 | Llama 3.1 8BRestricted weightsMeta AI · Llama 3.1 | 16.4% | Independent | group defaults | Comparable-11.7 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 266 | LFM2.5-8B-A1BOpen weightsLiquid AI · LFM2.5 | 16.1% | Independent | group defaults | Comparable-12.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 266 | step-3-vl-10bOpen weightsStepFun · Step3 | 16.1% | Independent | group defaults | Comparable-12.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 268 | jamba-reasoning-3bOpen weightsAI21 Labs · Jamba | 15.8% | Independent | group defaults | Comparable-12.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 269 | Llama 4 ScoutOpen weightsMeta AI · Llama 4 | 15.5% | Independent | group defaults | Comparable-12.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 270 | Command AOpen weightsCohere · Command | 15.2% | Independent | group defaults | Comparable-12.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 270 | Llama-3.1-70BOpen weightsMeta AI · Llama 3.1 | 15.2% | Independent | group defaults | Comparable-12.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 272 | granite-4-0-h-350mOpen weightsIBM · Granite 4.0 | 14.6% | Independent | group defaults | Comparable-13.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 272 | llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.2 | 14.6% | Independent | group defaults | Comparable-13.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 272 | qwen3-0.6b-instructOpen weightsAlibaba Group · Qwen3.0 | 14.6% | Independent | group defaults | Comparable-13.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 275 | nova-microClosedAmazon Web Services · Nova | 14.0% | Independent | group defaults | Comparable-14.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 275 | nova-proClosedAmazon Web Services · Nova | 14.0% | Independent | group defaults | Comparable-14.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 277 | jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.7 | 13.4% | Independent | group defaults | Comparable-14.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 277 | lfm2-2-6bOpen weightsLiquid AI · LFM2.2 | 13.4% | Independent | group defaults | Comparable-14.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 279 | granite-4-0-350mOpen weightsIBM · Granite 4.0 | 13.2% | Independent | group defaults | Comparable-14.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 279 | ling-mini-2-0Open weightsinclusionAI | 13.2% | Independent | group defaults | Comparable-14.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 281 | apertus-70b-instructOpen weightsSwiss AI Initiative | 12.9% | Independent | group defaults | Comparable-15.2 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 282 | Granite 4.0 MicroOpen weightsIBM · Granite 4.0 | 12.6% | Independent | group defaults | Comparable-15.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 282 | LFM2-1.2BOpen weightsLiquid AI · LFM2.1 | 12.6% | Independent | group defaults | Comparable-15.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 282 | Olmo-3-7B-InstructOpen weightsAllen Institute for AI · OLMo 3 | 12.6% | Independent | group defaults | Comparable-15.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 282 | jamba-1-7-miniOpen weightsAI21 Labs · Jamba 1.7 | 12.6% | Independent | group defaults | Comparable-15.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 286 | llama-3-1-nemotron-nano-4b-reasoningOpen weightsNVIDIA · Llama 3.1 | 11.7% | Independent | group defaults | Comparable-16.4 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 287 | apertus-8b-instructOpen weightsSwiss AI Initiative | 11.4% | Independent | group defaults | Comparable-16.7 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 287 | deepseek-r1-0120Open weightsDeepSeek · DeepSeek | 11.4% | Independent | group defaults | Comparable-16.7 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 287 | llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.1 | 11.4% | Independent | group defaults | Comparable-16.7 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 290 | lfm2-24b-a2bOpen weightsLiquid AI · LFM2 | 11.1% | Independent | group defaults | Comparable-17.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 291 | Gemma 3 12BOpen weightsGoogle · Gemma 3 | 10.8% | Independent | group defaults | Comparable-17.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 291 | LFM2.5-1.2B-InstructOpen weightsLiquid AI · LFM2.5 | 10.8% | Independent | group defaults | Comparable-17.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 293 | Gemma 3 27BOpen weightsGoogle · Gemma 3 | 10.5% | Independent | group defaults | Comparable-17.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 293 | gemma-3-1bOpen weightsGoogle · Gemma 3 | 10.5% | Independent | group defaults | Comparable-17.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 293 | granite-3-3-8b-instructOpen weightsIBM · Granite 3.3 | 10.5% | Independent | group defaults | Comparable-17.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 293 | lfm2-8b-a1bOpen weightsLiquid AI · LFM2 | 10.5% | Independent | group defaults | Comparable-17.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 297 | gemma-3-270mOpen weightsGoogle · Gemma 3 | 9.06% | Independent | group defaults | Comparable-19.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 298 | LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.5 | 8.48% | Independent | group defaults | Comparable-19.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 299 | phi-4-miniOpen weightsMicrosoft · Phi4 | 8.19% | Independent | group defaults | Comparable-19.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 300 | Gemma 3 4BRestricted weightsGoogle · Gemma 3 | 4.97% | Independent | group defaults | Comparable-23.1 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →