Skip to content
AI Atlas
BenchmarkActivecategory · agenticfamily · terminal-bench · variant 1.0

Terminal-Bench

tbench.ai

terminal tasks solved by agents

data quality57

Updated 54 min ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
1,639
Models
376
Current leader
gpt-5.6-sol 65.9%

Score history · Qwen3.8 27B 14 rows

Score history for Qwen3.8 27B0%20%40%60%80%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • Qwen3.8 27B
  • 49.06%aa_slug=qwen3-8-27b-non-reasoning · variant=v2.1 · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
  • 65.17%aa_slug=qwen3-8-27b-medium · variant=v2.1 · evaluator=Artificial Analysis · index_version=4.312 Sept 2026
  • 5.05%aa_slug=qwen3-8-27b-medium · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.312 Sept 2026
  • 67.42%aa_slug=qwen3-8-27b-low · variant=v2.1 · evaluator=Artificial Analysis · index_version=4.312 Sept 2026
  • 2.53%aa_slug=qwen3-8-27b-low · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.312 Sept 2026
  • 79.78%aa_slug=qwen3-8-27b · variant=v2.1 · evaluator=Artificial Analysis · index_version=4.312 Sept 2026
  • 5.56%aa_slug=qwen3-8-27b · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.312 Sept 2026
  • 49.06%aa_slug=qwen3-8-27b-non-reasoning · variant=v2.1 · evaluator=Artificial Analysis · reasoning=off11 Sept 2026
  • 65.17%aa_slug=qwen3-8-27b-medium · variant=v2.1 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 5.05%aa_slug=qwen3-8-27b-medium · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 67.42%aa_slug=qwen3-8-27b-low · variant=v2.1 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 2.53%aa_slug=qwen3-8-27b-low · variant=v4.0 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · variant=hard · evaluator=Artificial Analysis

10 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 65.9%gpt-5.6-sol OpenAI Independent11 Sept 2026
  2. 61.4%gpt-5.6-sol OpenAI Independent11 Sept 2026
  3. 53.0%Claude Sonnet 4.6 Anthropic Independent11 Sept 2026
  4. 51.5%Claude Opus 4.7 Anthropic Independent11 Sept 2026
  5. 50.8%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
  6. 49.2%KAT-Coder-Pro V2 Kwaipilot Independent11 Sept 2026
  7. 33.3%GPT-5.1-Codex Mini OpenAI Independent11 Sept 2026
  8. 26.5%Grok 4.3 xAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 317 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 10 rows65.9%Independentreasoning_effortmaxleaderobs. 12 Sept 2026artificialanalysis.aiT2History
2Claude Fable 5ClosedAnthropic · Claude · best of 2 rows62.9%IndependentreasoningonPartially comparable-3.03 ptobs. 12 Sept 2026artificialanalysis.aiT2History
2gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 8 rows62.9%Independentreasoning_effortxhighPartially comparable-3.03 ptobs. 12 Sept 2026artificialanalysis.aiT2History
4gpt-5.5ClosedOpenAI · GPT 5.5 · best of 10 rows60.6%Independentreasoning_effortxhighPartially comparable-5.30 ptobs. 12 Sept 2026artificialanalysis.aiT2History
5Claude Opus 4.8ClosedAnthropic · Claude · best of 2 rows58.3%Independentreasoning_effortmaxComparable-7.58 ptobs. 12 Sept 2026artificialanalysis.aiT2History
6gpt-5.4ClosedOpenAI · GPT 5.4 · best of 6 rows57.6%Independentreasoning_effortxhighPartially comparable-8.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
7Claude Opus 4.7ClosedAnthropic · Claude · best of 4 rows54.5%Independentreasoningoffreasoning_efforthighPartially comparable-11.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
8Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.1 · best of 2 rows53.8%IndependentreasoningonPartially comparable-12.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
9Claude Sonnet 4.6ClosedAnthropic · Claude · best of 5 rows53.0%Independentreasoningadaptivereasoning_effortmaxPartially comparable-12.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
9gpt-5.3-codexClosedOpenAI · GPT 5.3 · best of 2 rows53.0%Independentreasoning_effortxhighPartially comparable-12.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
11gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 6 rows52.3%Independentreasoning_effortxhighPartially comparable-13.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
12Qwen3.7 MaxClosedQwen · Qwen3.7 · best of 2 rows50.8%IndependentreasoningonPartially comparable-15.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
12Z.ai GLM 5.2Open weightsZ.ai (Zhipu AI) · GLM5.2 · best of 2 rows50.8%Independentreasoning_effortmaxComparable-15.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
14KAT-Coder-Pro V2ClosedKwaipilot · best of 2 rows49.2%IndependentreasoningoffPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
15Claude Opus 4.6ClosedAnthropic · Claude · best of 4 rows48.5%Independentreasoningoffreasoning_efforthighPartially comparable-17.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
16Claude Opus 4.5ClosedAnthropic · Claude · best of 4 rows47.0%IndependentreasoningonPartially comparable-18.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
16Qwen 3.7 PlusClosedQwen · Qwen3.7 · best of 2 rows47.0%IndependentreasoningonPartially comparable-18.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
16gpt-5.2ClosedOpenAI · GPT 5.2 · best of 6 rows47.0%Independentreasoning_effortxhighPartially comparable-18.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
19Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 6 rows46.2%Independentreasoning_effortminimalPartially comparable-19.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
19deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek · best of 4 rows46.2%Independentreasoning_effortmaxComparable-19.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
21gpt-5.1ClosedOpenAI · GPT 5.1 · best of 4 rows45.5%Independentreasoning_efforthighPartially comparable-20.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
21muse-sparkClosedMeta AI · best of 2 rows45.5%IndependentreasoningonPartially comparable-20.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
23Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi · best of 2 rows44.7%IndependentreasoningonPartially comparable-21.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
24Kimi K2.6Open weightsMoonshot AI · Kimi · best of 4 rows43.9%IndependentreasoningonPartially comparable-22.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
24Qwen3.6 Max PreviewClosedQwen · Qwen3.6 · best of 2 rows43.9%IndependentreasoningonPartially comparable-22.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
24Qwen3.6 PlusClosedQwen · Qwen3.6 · best of 2 rows43.9%IndependentreasoningonPartially comparable-22.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
27GLM 5Open weightsZ.ai (Zhipu AI) · GLM5 · best of 4 rows43.2%IndependentreasoningonPartially comparable-22.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
27GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 4 rows43.2%IndependentreasoningonPartially comparable-22.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
27MiMo-V2.5-ProOpen weightsXiaomi · best of 4 rows43.2%IndependentreasoningonPartially comparable-22.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
30MiniMax-M3Open weightsMiniMax · MiniMax · best of 2 rows42.4%IndependentreasoningonPartially comparable-23.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
30gpt-5-5-instant-05-26ClosedOpenAI · GPT 5.5 · best of 2 rows42.4%IndependentreasoningonPartially comparable-23.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
30gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 6 rows42.4%Independentreasoning_effortxhighPartially comparable-23.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
33MiMo-V2.5Open weightsXiaomi · best of 2 rows41.7%IndependentreasoningonPartially comparable-24.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
33deepseek-v4-pro-0424-highOpen weightsDeepSeek · DeepSeek41.7%Independentgroup defaultsPartially comparable-24.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
33gemini-3-proClosedGoogle · Gemini 3 · best of 4 rows41.7%Independentreasoning_efforthighPartially comparable-24.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
36Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 4 rows40.9%IndependentreasoningonPartially comparable-25.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
36grok-4-20-0309ClosedSpaceXAI · Grok 4.20 · best of 4 rows40.9%IndependentreasoningonPartially comparable-25.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
36mimo-v2-proClosedXiaomi · best of 2 rows40.9%IndependentreasoningonPartially comparable-25.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
39MiniMax M2.7Open weightsMiniMax · MiniMax · best of 2 rows39.4%IndependentreasoningonPartially comparable-26.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
40Gemini 3 Flash PreviewClosedGoogle · Gemini 338.6%Independentgroup defaultsPartially comparable-27.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
40deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek · best of 4 rows38.6%Independentreasoning_efforthighPartially comparable-27.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
40deepseek-v4-flash-0420-highOpen weightsDeepSeek · DeepSeek38.6%Independentgroup defaultsPartially comparable-27.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
40gemini-3-flashClosedGoogle · Gemini 3 · best of 3 rows38.6%IndependentreasoningonPartially comparable-27.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
44GPT-5-CodexClosedOpenAI · GPT 5 · best of 2 rows37.9%Independentreasoning_efforthighPartially comparable-28.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
44Grok 4.20ClosedxAI · Grok · best of 3 rows37.9%IndependentreasoningonPartially comparable-28.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
44Grok 4.3ClosedxAI · Grok · best of 8 rows37.9%Independentreasoning_efforthighPartially comparable-28.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
44gpt-5ClosedOpenAI · GPT 5 · best of 8 rows37.9%Independentreasoning_effortmediumPartially comparable-28.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
44grok-4ClosedSpaceXAI · Grok 4 · best of 2 rows37.9%IndependentreasoningonPartially comparable-28.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49GPT-5.2-CodexClosedOpenAI · GPT 5.2 · best of 2 rows37.1%Independentreasoning_effortxhighPartially comparable-28.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49o3ClosedOpenAI · OpenAI o-series · best of 2 rows37.1%IndependentreasoningonPartially comparable-28.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
51Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 4 rows36.4%IndependentreasoningonPartially comparable-29.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
51Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 3 · best of 2 rows36.4%IndependentreasoningonPartially comparable-29.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
51deepseek-v4-pro-0424-non-reasoningOpen weightsDeepSeek · DeepSeek36.4%Independentgroup defaultsPartially comparable-29.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
54Claude Sonnet 4.5ClosedAnthropic · Claude · best of 4 rows35.6%IndependentreasoningonPartially comparable-30.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
54DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 5 rows35.6%IndependentreasoningonPartially comparable-30.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
54DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V335.6%Independentgroup defaultsPartially comparable-30.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
54Step 3.7 FlashOpen weightsStepFun · Step3.7 · best of 2 rows35.6%IndependentreasoningonPartially comparable-30.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
54mimo-v2-omni-0327ClosedXiaomi · best of 2 rows35.6%IndependentreasoningonPartially comparable-30.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
59GPT-5.1-CodexClosedOpenAI · GPT 5.1 · best of 2 rows34.9%Independentreasoning_efforthighPartially comparable-31.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
59Kimi K2.5Open weightsMoonshot AI · Kimi · best of 4 rows34.9%IndependentreasoningonPartially comparable-31.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
59MiniMax M2.5Open weightsMiniMax · MiniMax · best of 2 rows34.9%IndependentreasoningonPartially comparable-31.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
59Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 4 rows34.9%IndependentreasoningonPartially comparable-31.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
59Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 4 rows34.9%IndependentreasoningonPartially comparable-31.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
59deepseek-v3-2-specialeOpen weightsDeepSeek · DeepSeek · best of 2 rows34.9%IndependentreasoningonPartially comparable-31.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
59mimo-v2-omniClosedXiaomi · best of 2 rows34.9%IndependentreasoningonPartially comparable-31.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
66Claude Opus 4.1ClosedAnthropic · Claude · best of 2 rows34.3%IndependentreasoningonPartially comparable-31.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
67Hy3Open weightsTencent · best of 3 rows34.1%IndependentreasoningonPartially comparable-31.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
67Hy3 previewOpen weightsTencent34.1%Independentgroup defaultsPartially comparable-31.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
67deepseek-v4-flash-0420-non-reasoningOpen weightsDeepSeek · DeepSeek34.1%Independentgroup defaultsPartially comparable-31.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
70GLM 5 TurboClosedZ.ai (Zhipu AI) · GLM5 · best of 2 rows33.3%IndependentreasoningonPartially comparable-32.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
70GPT-5.1-Codex MiniClosedOpenAI · GPT 5.1 · best of 2 rows33.3%Independentreasoning_efforthighPartially comparable-32.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
70Mistral Medium 3.5Open weightsMistral AI · Mistral · best of 2 rows33.3%IndependentreasoningonPartially comparable-32.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
70gpt-5-miniClosedOpenAI · GPT 5 · best of 6 rows33.3%Independentreasoning_efforthighPartially comparable-32.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
74GLM 5V TurboClosedZ.ai (Zhipu AI) · GLM5 · best of 2 rows32.6%IndependentreasoningonPartially comparable-33.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
74Qwen3.5-27BOpen weightsQwen · Qwen3.5 · best of 4 rows32.6%IndependentreasoningonPartially comparable-33.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
74Step 3.5 FlashOpen weightsStepFun · Step3.5 · best of 4 rows32.6%IndependentreasoningonPartially comparable-33.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
77DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 4 rows31.8%IndependentreasoningoffPartially comparable-34.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
77GLM 4.7Open weightsZ.ai (Zhipu AI) · GLM4.7 · best of 4 rows31.8%IndependentreasoningonPartially comparable-34.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79Claude Opus 4ClosedAnthropic · Claude31.1%Independentgroup defaultsPartially comparable-34.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
79Claude Sonnet 4ClosedAnthropic · Claude · best of 4 rows31.1%IndependentreasoningonPartially comparable-34.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79DeepSeek V3.2 ExpOpen weightsDeepSeek · DeepSeek-V3 · best of 3 rows31.1%IndependentreasoningonPartially comparable-34.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79Kimi K2 ThinkingOpen weightsMoonshot AI · Kimi · best of 2 rows31.1%IndependentreasoningonPartially comparable-34.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79North Mini Code (free)Open weightsCohere · best of 2 rows31.1%IndependentreasoningonPartially comparable-34.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 4 rows31.1%IndependentreasoningonPartially comparable-34.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79claude-4-opusClosedAnthropic · Claude 431.1%IndependentreasoningonPartially comparable-34.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79ling-2-6-1tOpen weightsinclusionAI · best of 2 rows31.1%IndependentreasoningoffPartially comparable-34.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79mimo-v2-0206Open weightsXiaomi · best of 2 rows31.1%IndependentreasoningonPartially comparable-34.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.6 · best of 4 rows28.8%IndependentreasoningoffPartially comparable-37.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88MiniMax M2.1Open weightsMiniMax · MiniMax · best of 2 rows28.8%IndependentreasoningonPartially comparable-37.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 3 · best of 2 rows28.8%IndependentreasoningonPartially comparable-37.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88jt-35b-flashClosedChina Mobile · best of 2 rows28.8%IndependentreasoningoffPartially comparable-37.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88ring-2-6-1tOpen weightsinclusionAI · best of 2 rows28.8%IndependentreasoningonPartially comparable-37.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
93mimo-v2-flashOpen weightsXiaomi · best of 4 rows28.0%IndependentreasoningonPartially comparable-37.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
94Claude Haiku 4.5ClosedAnthropic · Claude · best of 4 rows27.3%IndependentreasoningoffPartially comparable-38.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
95Gemini 2.5 ProClosedGoogle · Gemini 2.5 · best of 2 rows26.5%IndependentreasoningonPartially comparable-39.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
95Mercury 2ClosedInception · best of 2 rows26.5%IndependentreasoningonPartially comparable-39.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
95Qwen3.5-35B-A3BOpen weightsQwen · Qwen3.5 · best of 4 rows26.5%IndependentreasoningonPartially comparable-39.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
95doubao-seed-codeClosedByteDance · Seed · best of 2 rows26.5%IndependentreasoningonPartially comparable-39.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
99MiniMax M2Open weightsMiniMax · MiniMax · best of 2 rows25.8%IndependentreasoningonPartially comparable-40.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
100DeepSeek V3.1Open weightsDeepSeek · DeepSeek-V3 · best of 4 rows25%IndependentreasoningonPartially comparable-40.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →