Updated 54 min ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 1,639
- Models
- 376
- Current leader
- gpt-5.6-sol 65.9%
Score history · grok-4-20-0309 4 rows
- grok-4-20-0309
- 40.91%aa_slug=grok-4-20-0309 · variant=hard · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
- 21.97%aa_slug=grok-4-20-0309-non-reasoning · variant=hard · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
- 40.91%aa_slug=grok-4-20-0309 · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
- 21.97%aa_slug=grok-4-20-0309-non-reasoning · variant=hard · evaluator=Artificial Analysis · reasoning=off11 Sept 2026
Frontier over time · accuracy · variant=hard · evaluator=Artificial Analysis
10 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 65.9%gpt-5.6-sol OpenAI Independent11 Sept 2026
- 61.4%gpt-5.6-sol OpenAI Independent11 Sept 2026
- 53.0%Claude Sonnet 4.6 Anthropic Independent11 Sept 2026
- 51.5%Claude Opus 4.7 Anthropic Independent11 Sept 2026
- 50.8%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
- 49.2%KAT-Coder-Pro V2 Kwaipilot Independent11 Sept 2026
- 33.3%GPT-5.1-Codex Mini OpenAI Independent11 Sept 2026
- 26.5%Grok 4.3 xAI Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 317 models
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 1 | gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 10 rows | 65.9% | Independent | reasoning_effortmax | leader | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 2 | Claude Fable 5ClosedAnthropic · Claude · best of 2 rows | 62.9% | Independent | reasoningon | Partially comparable-3.03 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 2 | gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 8 rows | 62.9% | Independent | reasoning_effortxhigh | Partially comparable-3.03 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 4 | gpt-5.5ClosedOpenAI · GPT 5.5 · best of 10 rows | 60.6% | Independent | reasoning_effortxhigh | Partially comparable-5.30 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 5 | Claude Opus 4.8ClosedAnthropic · Claude · best of 2 rows | 58.3% | Independent | reasoning_effortmax | Comparable-7.58 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 6 | gpt-5.4ClosedOpenAI · GPT 5.4 · best of 6 rows | 57.6% | Independent | reasoning_effortxhigh | Partially comparable-8.33 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 7 | Claude Opus 4.7ClosedAnthropic · Claude · best of 4 rows | 54.5% | Independent | reasoningoffreasoning_efforthigh | Partially comparable-11.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 8 | Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.1 · best of 2 rows | 53.8% | Independent | reasoningon | Partially comparable-12.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 9 | Claude Sonnet 4.6ClosedAnthropic · Claude · best of 5 rows | 53.0% | Independent | reasoningadaptivereasoning_effortmax | Partially comparable-12.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 9 | gpt-5.3-codexClosedOpenAI · GPT 5.3 · best of 2 rows | 53.0% | Independent | reasoning_effortxhigh | Partially comparable-12.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 11 | gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 6 rows | 52.3% | Independent | reasoning_effortxhigh | Partially comparable-13.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 12 | Qwen3.7 MaxClosedQwen · Qwen3.7 · best of 2 rows | 50.8% | Independent | reasoningon | Partially comparable-15.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 12 | Z.ai GLM 5.2Open weightsZ.ai (Zhipu AI) · GLM5.2 · best of 2 rows | 50.8% | Independent | reasoning_effortmax | Comparable-15.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 14 | KAT-Coder-Pro V2ClosedKwaipilot · best of 2 rows | 49.2% | Independent | reasoningoff | Partially comparable-16.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 15 | Claude Opus 4.6ClosedAnthropic · Claude · best of 4 rows | 48.5% | Independent | reasoningoffreasoning_efforthigh | Partially comparable-17.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 16 | Claude Opus 4.5ClosedAnthropic · Claude · best of 4 rows | 47.0% | Independent | reasoningon | Partially comparable-18.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 16 | Qwen 3.7 PlusClosedQwen · Qwen3.7 · best of 2 rows | 47.0% | Independent | reasoningon | Partially comparable-18.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 16 | gpt-5.2ClosedOpenAI · GPT 5.2 · best of 6 rows | 47.0% | Independent | reasoning_effortxhigh | Partially comparable-18.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 19 | Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 6 rows | 46.2% | Independent | reasoning_effortminimal | Partially comparable-19.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 19 | deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek · best of 4 rows | 46.2% | Independent | reasoning_effortmax | Comparable-19.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 21 | gpt-5.1ClosedOpenAI · GPT 5.1 · best of 4 rows | 45.5% | Independent | reasoning_efforthigh | Partially comparable-20.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 21 | muse-sparkClosedMeta AI · best of 2 rows | 45.5% | Independent | reasoningon | Partially comparable-20.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 23 | Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi · best of 2 rows | 44.7% | Independent | reasoningon | Partially comparable-21.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 24 | Kimi K2.6Open weightsMoonshot AI · Kimi · best of 4 rows | 43.9% | Independent | reasoningon | Partially comparable-22.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 24 | Qwen3.6 Max PreviewClosedQwen · Qwen3.6 · best of 2 rows | 43.9% | Independent | reasoningon | Partially comparable-22.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 24 | Qwen3.6 PlusClosedQwen · Qwen3.6 · best of 2 rows | 43.9% | Independent | reasoningon | Partially comparable-22.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 27 | GLM 5Open weightsZ.ai (Zhipu AI) · GLM5 · best of 4 rows | 43.2% | Independent | reasoningon | Partially comparable-22.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 27 | GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 4 rows | 43.2% | Independent | reasoningon | Partially comparable-22.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 27 | MiMo-V2.5-ProOpen weightsXiaomi · best of 4 rows | 43.2% | Independent | reasoningon | Partially comparable-22.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 30 | MiniMax-M3Open weightsMiniMax · MiniMax · best of 2 rows | 42.4% | Independent | reasoningon | Partially comparable-23.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 30 | gpt-5-5-instant-05-26ClosedOpenAI · GPT 5.5 · best of 2 rows | 42.4% | Independent | reasoningon | Partially comparable-23.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 30 | gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 6 rows | 42.4% | Independent | reasoning_effortxhigh | Partially comparable-23.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 33 | MiMo-V2.5Open weightsXiaomi · best of 2 rows | 41.7% | Independent | reasoningon | Partially comparable-24.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 33 | deepseek-v4-pro-0424-highOpen weightsDeepSeek · DeepSeek | 41.7% | Independent | group defaults | Partially comparable-24.2 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 33 | gemini-3-proClosedGoogle · Gemini 3 · best of 4 rows | 41.7% | Independent | reasoning_efforthigh | Partially comparable-24.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 36 | Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 4 rows | 40.9% | Independent | reasoningon | Partially comparable-25.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 36 | grok-4-20-0309ClosedSpaceXAI · Grok 4.20 · best of 4 rows | 40.9% | Independent | reasoningon | Partially comparable-25.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 36 | mimo-v2-proClosedXiaomi · best of 2 rows | 40.9% | Independent | reasoningon | Partially comparable-25.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 39 | MiniMax M2.7Open weightsMiniMax · MiniMax · best of 2 rows | 39.4% | Independent | reasoningon | Partially comparable-26.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 40 | Gemini 3 Flash PreviewClosedGoogle · Gemini 3 | 38.6% | Independent | group defaults | Partially comparable-27.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 40 | deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek · best of 4 rows | 38.6% | Independent | reasoning_efforthigh | Partially comparable-27.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 40 | deepseek-v4-flash-0420-highOpen weightsDeepSeek · DeepSeek | 38.6% | Independent | group defaults | Partially comparable-27.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 40 | gemini-3-flashClosedGoogle · Gemini 3 · best of 3 rows | 38.6% | Independent | reasoningon | Partially comparable-27.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 44 | GPT-5-CodexClosedOpenAI · GPT 5 · best of 2 rows | 37.9% | Independent | reasoning_efforthigh | Partially comparable-28.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 44 | Grok 4.20ClosedxAI · Grok · best of 3 rows | 37.9% | Independent | reasoningon | Partially comparable-28.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 44 | Grok 4.3ClosedxAI · Grok · best of 8 rows | 37.9% | Independent | reasoning_efforthigh | Partially comparable-28.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 44 | gpt-5ClosedOpenAI · GPT 5 · best of 8 rows | 37.9% | Independent | reasoning_effortmedium | Partially comparable-28.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 44 | grok-4ClosedSpaceXAI · Grok 4 · best of 2 rows | 37.9% | Independent | reasoningon | Partially comparable-28.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | GPT-5.2-CodexClosedOpenAI · GPT 5.2 · best of 2 rows | 37.1% | Independent | reasoning_effortxhigh | Partially comparable-28.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | o3ClosedOpenAI · OpenAI o-series · best of 2 rows | 37.1% | Independent | reasoningon | Partially comparable-28.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 51 | Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 4 rows | 36.4% | Independent | reasoningon | Partially comparable-29.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 51 | Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 36.4% | Independent | reasoningon | Partially comparable-29.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 51 | deepseek-v4-pro-0424-non-reasoningOpen weightsDeepSeek · DeepSeek | 36.4% | Independent | group defaults | Partially comparable-29.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 54 | Claude Sonnet 4.5ClosedAnthropic · Claude · best of 4 rows | 35.6% | Independent | reasoningon | Partially comparable-30.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 54 | DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 5 rows | 35.6% | Independent | reasoningon | Partially comparable-30.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 54 | DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V3 | 35.6% | Independent | group defaults | Partially comparable-30.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 54 | Step 3.7 FlashOpen weightsStepFun · Step3.7 · best of 2 rows | 35.6% | Independent | reasoningon | Partially comparable-30.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 54 | mimo-v2-omni-0327ClosedXiaomi · best of 2 rows | 35.6% | Independent | reasoningon | Partially comparable-30.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 59 | GPT-5.1-CodexClosedOpenAI · GPT 5.1 · best of 2 rows | 34.9% | Independent | reasoning_efforthigh | Partially comparable-31.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 59 | Kimi K2.5Open weightsMoonshot AI · Kimi · best of 4 rows | 34.9% | Independent | reasoningon | Partially comparable-31.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 59 | MiniMax M2.5Open weightsMiniMax · MiniMax · best of 2 rows | 34.9% | Independent | reasoningon | Partially comparable-31.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 59 | Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 4 rows | 34.9% | Independent | reasoningon | Partially comparable-31.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 59 | Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 4 rows | 34.9% | Independent | reasoningon | Partially comparable-31.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 59 | deepseek-v3-2-specialeOpen weightsDeepSeek · DeepSeek · best of 2 rows | 34.9% | Independent | reasoningon | Partially comparable-31.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 59 | mimo-v2-omniClosedXiaomi · best of 2 rows | 34.9% | Independent | reasoningon | Partially comparable-31.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 66 | Claude Opus 4.1ClosedAnthropic · Claude · best of 2 rows | 34.3% | Independent | reasoningon | Partially comparable-31.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 67 | Hy3Open weightsTencent · best of 3 rows | 34.1% | Independent | reasoningon | Partially comparable-31.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 67 | Hy3 previewOpen weightsTencent | 34.1% | Independent | group defaults | Partially comparable-31.8 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 67 | deepseek-v4-flash-0420-non-reasoningOpen weightsDeepSeek · DeepSeek | 34.1% | Independent | group defaults | Partially comparable-31.8 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 70 | GLM 5 TurboClosedZ.ai (Zhipu AI) · GLM5 · best of 2 rows | 33.3% | Independent | reasoningon | Partially comparable-32.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 70 | GPT-5.1-Codex MiniClosedOpenAI · GPT 5.1 · best of 2 rows | 33.3% | Independent | reasoning_efforthigh | Partially comparable-32.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 70 | Mistral Medium 3.5Open weightsMistral AI · Mistral · best of 2 rows | 33.3% | Independent | reasoningon | Partially comparable-32.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 70 | gpt-5-miniClosedOpenAI · GPT 5 · best of 6 rows | 33.3% | Independent | reasoning_efforthigh | Partially comparable-32.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 74 | GLM 5V TurboClosedZ.ai (Zhipu AI) · GLM5 · best of 2 rows | 32.6% | Independent | reasoningon | Partially comparable-33.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 74 | Qwen3.5-27BOpen weightsQwen · Qwen3.5 · best of 4 rows | 32.6% | Independent | reasoningon | Partially comparable-33.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 74 | Step 3.5 FlashOpen weightsStepFun · Step3.5 · best of 4 rows | 32.6% | Independent | reasoningon | Partially comparable-33.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 77 | DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 4 rows | 31.8% | Independent | reasoningoff | Partially comparable-34.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 77 | GLM 4.7Open weightsZ.ai (Zhipu AI) · GLM4.7 · best of 4 rows | 31.8% | Independent | reasoningon | Partially comparable-34.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 79 | Claude Opus 4ClosedAnthropic · Claude | 31.1% | Independent | group defaults | Partially comparable-34.8 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 79 | Claude Sonnet 4ClosedAnthropic · Claude · best of 4 rows | 31.1% | Independent | reasoningon | Partially comparable-34.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 79 | DeepSeek V3.2 ExpOpen weightsDeepSeek · DeepSeek-V3 · best of 3 rows | 31.1% | Independent | reasoningon | Partially comparable-34.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 79 | Kimi K2 ThinkingOpen weightsMoonshot AI · Kimi · best of 2 rows | 31.1% | Independent | reasoningon | Partially comparable-34.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 79 | North Mini Code (free)Open weightsCohere · best of 2 rows | 31.1% | Independent | reasoningon | Partially comparable-34.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 79 | Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 4 rows | 31.1% | Independent | reasoningon | Partially comparable-34.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 79 | claude-4-opusClosedAnthropic · Claude 4 | 31.1% | Independent | reasoningon | Partially comparable-34.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 79 | ling-2-6-1tOpen weightsinclusionAI · best of 2 rows | 31.1% | Independent | reasoningoff | Partially comparable-34.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 79 | mimo-v2-0206Open weightsXiaomi · best of 2 rows | 31.1% | Independent | reasoningon | Partially comparable-34.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 88 | GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.6 · best of 4 rows | 28.8% | Independent | reasoningoff | Partially comparable-37.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 88 | MiniMax M2.1Open weightsMiniMax · MiniMax · best of 2 rows | 28.8% | Independent | reasoningon | Partially comparable-37.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 88 | Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 28.8% | Independent | reasoningon | Partially comparable-37.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 88 | jt-35b-flashClosedChina Mobile · best of 2 rows | 28.8% | Independent | reasoningoff | Partially comparable-37.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 88 | ring-2-6-1tOpen weightsinclusionAI · best of 2 rows | 28.8% | Independent | reasoningon | Partially comparable-37.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 93 | mimo-v2-flashOpen weightsXiaomi · best of 4 rows | 28.0% | Independent | reasoningon | Partially comparable-37.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 94 | Claude Haiku 4.5ClosedAnthropic · Claude · best of 4 rows | 27.3% | Independent | reasoningoff | Partially comparable-38.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 95 | Gemini 2.5 ProClosedGoogle · Gemini 2.5 · best of 2 rows | 26.5% | Independent | reasoningon | Partially comparable-39.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 95 | Mercury 2ClosedInception · best of 2 rows | 26.5% | Independent | reasoningon | Partially comparable-39.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 95 | Qwen3.5-35B-A3BOpen weightsQwen · Qwen3.5 · best of 4 rows | 26.5% | Independent | reasoningon | Partially comparable-39.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 95 | doubao-seed-codeClosedByteDance · Seed · best of 2 rows | 26.5% | Independent | reasoningon | Partially comparable-39.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 99 | MiniMax M2Open weightsMiniMax · MiniMax · best of 2 rows | 25.8% | Independent | reasoningon | Partially comparable-40.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 100 | DeepSeek V3.1Open weightsDeepSeek · DeepSeek-V3 · best of 4 rows | 25% | Independent | reasoningon | Partially comparable-40.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →