Updated 57 min ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 1,639
- Models
- 376
- Current leader
- gpt-6-astra 59.6%
Score history · Qwen3.5-4B 8 rows
- Qwen3.5-4B
- 11.36%aa_slug=qwen3-5-4b-non-reasoning · variant=hard · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
- 21.35%aa_slug=qwen3-5-4b-non-reasoning · variant=v2.1 · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
- 18.18%aa_slug=qwen3-5-4b · variant=hard · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
- 25.84%aa_slug=qwen3-5-4b · variant=v2.1 · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
- 11.36%aa_slug=qwen3-5-4b-non-reasoning · variant=hard · evaluator=Artificial Analysis · reasoning=off11 Sept 2026
- 21.35%aa_slug=qwen3-5-4b-non-reasoning · variant=v2.1 · evaluator=Artificial Analysis · reasoning=off11 Sept 2026
- 18.18%aa_slug=qwen3-5-4b · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
- 25.84%aa_slug=qwen3-5-4b · variant=v2.1 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
Frontier over time · accuracy · variant=v4.0 · evaluator=Artificial Analysis
6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 59.6%gpt-6-astra OpenAI Independent11 Sept 2026
- 59.1%gpt-6-astra OpenAI Independent11 Sept 2026
- 55.0%Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 54.0%gpt-6-astra OpenAI Independent11 Sept 2026
- 49.5%gpt-6-astra OpenAI Independent11 Sept 2026
- 34.3%Claude Opus 5 Anthropic Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 114 models
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 1 | gpt-6-astraClosedOpenAI · GPT 6 · best of 11 rows | 59.6% | Independent | reasoning_effortxhigh | leader | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 2 | Claude Fable 5.1ClosedAnthropic · Claude · best of 10 rows | 55.0% | Independent | reasoning_effortxhigh | Comparable-4.55 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 3 | Claude Opus 5ClosedAnthropic · Claude · best of 10 rows | 49.0% | Independent | reasoning_effortmax | Partially comparable-10.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 4 | Claude Fable 5ClosedAnthropic · Claude · best of 2 rows | 42.4% | Independent | reasoningon | Partially comparable-17.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 5 | GLM 5.3Open weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows | 41.9% | Independent | reasoning_effortmax | Partially comparable-17.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 6 | gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 10 rows | 39.9% | Independent | reasoning_effortmax | Partially comparable-19.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 7 | gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 10 rows | 35.4% | Independent | reasoning_effortmax | Partially comparable-24.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 8 | Muse Spark 1.3ClosedMeta AI · best of 4 rows | 33.3% | Independent | reasoning_effortmax | Partially comparable-26.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 9 | GLM 5.3 FlashOpen weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows | 32.8% | Independent | reasoningon | Partially comparable-26.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 10 | DeepSeek-V4.1-FlashOpen weightsDeepSeek · DeepSeek · best of 2 rows | 26.8% | Independent | reasoning_effortmax | Partially comparable-32.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 11 | Qwen3.8 FlashOpen weightsQwen · Qwen3.8 · best of 2 rows | 25.3% | Independent | reasoningon | Partially comparable-34.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 12 | Claude Opus 4.8ClosedAnthropic · Claude · best of 2 rows | 21.7% | Independent | reasoning_effortmax | Partially comparable-37.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 13 | Grok 4.6ClosedxAI · Grok · best of 8 rows | 21.2% | Independent | reasoning_efforthigh | Partially comparable-38.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 14 | Gemini 3.8 FlashClosedGoogle · Gemini 3.8 · best of 6 rows | 19.7% | Independent | reasoning_efforthigh | Partially comparable-39.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 15 | Qwen 3.8 MaxClosedQwen · Qwen3.8 · best of 2 rows | 18.7% | Independent | reasoningon | Partially comparable-40.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 16 | deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek · best of 2 rows | 14.7% | Independent | reasoning_effortmax | Partially comparable-45.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 16 | gpt-5.5ClosedOpenAI · GPT 5.5 · best of 6 rows | 14.7% | Independent | reasoning_effortxhigh | Comparable-45.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 18 | Claude Sonnet 5ClosedAnthropic · Claude · best of 10 rows | 14.1% | Independent | reasoning_effortmax | Partially comparable-45.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 18 | deepseek-v4-proClosedDeepSeek · V4 · best of 2 rows | 14.1% | Independent | reasoning_effortmax | Partially comparable-45.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 20 | Gemini 3.7 FlashClosedGoogle · Gemini 3.7 · best of 2 rows | 13.6% | Independent | reasoning_efforthigh | Partially comparable-46.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 21 | Kimi K3Open weightsMoonshot AI · Kimi · best of 3 rows | 12.6% | Independent | reasoning_effortmax | Partially comparable-47.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 21 | gpt-5-5-instant-06-26ClosedOpenAI · GPT 5.5 · best of 2 rows | 12.6% | Independent | reasoningon | Partially comparable-47.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 23 | deepseek-v4-flashOpen weightsDeepSeek · DeepSeek · best of 2 rows | 12.1% | Independent | reasoning_effortmax | Partially comparable-47.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 23 | deepseek-v4-flash-visionClosedDeepSeek · DeepSeek · best of 2 rows | 12.1% | Independent | reasoning_effortmax | Partially comparable-47.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 25 | gpt-5.6-lunaClosedOpenAI · GPT 5.6 · best of 10 rows | 11.6% | Independent | reasoning_effortmax | Partially comparable-48.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 26 | Qwen3.8 2.4T A95BOpen weightsQwen · Qwen3.8 · best of 2 rows | 11.1% | Independent | reasoningon | Partially comparable-48.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 27 | Grok 4.5ClosedxAI · Grok · best of 2 rows | 10.6% | Independent | reasoning_efforthigh | Partially comparable-49.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 28 | Gemini 3.6 FlashClosedGoogle · Gemini 3.6 · best of 2 rows | 7.07% | Independent | reasoningon | Partially comparable-52.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 28 | Muse Spark 1.2ClosedMeta AI · best of 2 rows | 7.07% | Independent | reasoning_effortxhigh | Comparable-52.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 28 | agnes-3-0-flashClosedSapiens AI · best of 2 rows | 7.07% | Independent | reasoningon | Partially comparable-52.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 31 | Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 2 rows | 6.57% | Independent | reasoningon | Partially comparable-53.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 32 | Muse Spark 1.1ClosedMeta AI · best of 2 rows | 6.06% | Independent | reasoning_effortxhigh | Comparable-53.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 33 | Qwen3.8 27BOpen weightsQwen · Qwen3.8 · best of 6 rows | 5.56% | Independent | reasoning_effortxhigh | Comparable-54.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 34 | Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.1 · best of 2 rows | 4.04% | Independent | reasoningon | Partially comparable-55.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 35 | Claude Sonnet 4.6ClosedAnthropic · Claude · best of 2 rows | 3.03% | Independent | reasoningadaptivereasoning_effortmax | Partially comparable-56.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 35 | agnes-2-5-pro-alphaOpen weightsSapiens AI · best of 2 rows | 3.03% | Independent | reasoningon | Partially comparable-56.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 35 | deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek · best of 3 rows | 3.03% | Independent | reasoning_efforthigh | Partially comparable-56.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 35 | deepseek-v4-flash-0420-highOpen weightsDeepSeek · DeepSeek | 3.03% | Independent | group defaults | Partially comparable-56.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 39 | GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 2 rows | 2.02% | Independent | reasoningon | Partially comparable-57.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 39 | MiniMax-M3Open weightsMiniMax · MiniMax · best of 2 rows | 2.02% | Independent | reasoningon | Partially comparable-57.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 39 | gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 2 rows | 2.02% | Independent | reasoning_effortxhigh | Comparable-57.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 42 | Qwen3.7 MaxClosedQwen · Qwen3.7 · best of 2 rows | 1.52% | Independent | reasoningon | Partially comparable-58.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 43 | Gemini 3.5 Flash-LiteClosedGoogle · Gemini 3.5 · best of 2 rows | 1.01% | Independent | reasoningon | Partially comparable-58.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 43 | InklingOpen weightsThinking Machines · best of 2 rows | 1.01% | Independent | reasoningon | Partially comparable-58.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 43 | Inkling SmallOpen weightsThinking Machines · best of 2 rows | 1.01% | Independent | reasoningon | Partially comparable-58.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 43 | Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi · best of 2 rows | 1.01% | Independent | reasoningon | Partially comparable-58.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 43 | Qwen 3.7 PlusClosedQwen · Qwen3.7 · best of 2 rows | 1.01% | Independent | reasoningon | Partially comparable-58.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 43 | quasar-438bClosedMultiverse Computing · best of 2 rows | 1.01% | Independent | reasoning_effortmax | Partially comparable-58.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.1 · best of 2 rows | 0.51% | Independent | reasoningon | Partially comparable-59.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | Hy3Open weightsTencent · best of 2 rows | 0.51% | Independent | reasoningon | Partially comparable-59.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | Muse Glimmer 30BOpen weightsMeta AI · best of 2 rows | 0.51% | Independent | reasoning_efforthigh | Partially comparable-59.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | NVIDIA Nemotron 3.5 Lightning 30B A3BOpen weightsNVIDIA · Nemotron 3.5 · best of 2 rows | 0.51% | Independent | reasoningon | Partially comparable-59.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 0.51% | Independent | reasoningon | Partially comparable-59.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | North Mini Code (free)Open weightsCohere · best of 2 rows | 0.51% | Independent | reasoningon | Partially comparable-59.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | Solar Pro 4ClosedUpstage · Solar · best of 2 rows | 0.51% | Independent | reasoningon | Partially comparable-59.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | Trinity Large ThinkingOpen weightsArcee AI · best of 2 rows | 0.51% | Independent | reasoningon | Partially comparable-59.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | command-a-plusOpen weightsCohere · Command · best of 2 rows | 0.51% | Independent | reasoningon | Partially comparable-59.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 2 rows | 0.51% | Independent | reasoning_effortxhigh | Comparable-59.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | ring-2-6-1tOpen weightsinclusionAI · best of 2 rows | 0.51% | Independent | reasoningon | Partially comparable-59.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Claude Haiku 4.5ClosedAnthropic · Claude · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Claude Sonnet 4.5ClosedAnthropic · Claude · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Devstral 2Open weightsMistral AI · Devstral 2 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Devstral Small 2Open weightsMistral AI · Devstral · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Gemini 2.5 ProClosedGoogle · Gemini 2.5 · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Gemma 3 12BOpen weightsGoogle · Gemma 3 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Gemma 3 27BOpen weightsGoogle · Gemma 3 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Granite 4.2 8BOpen weightsIBM · Granite 4.2 · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Grok 4.3ClosedxAI · Grok · best of 4 rows | 0% | Independent | reasoning_efforthigh | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Ling 3.0 Flash VLOpen weightsinclusionAI · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Llama 4 MaverickOpen weightsMeta AI · Llama 4 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Llama 4 ScoutOpen weightsMeta AI · Llama 4 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | LongCat 2.0Open weightsMeituan · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Mercury 2ClosedInception · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | MiMo-V2.5Open weightsXiaomi · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | MiMo-V2.5-ProOpen weightsXiaomi · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | MiniMax M2.7Open weightsMiniMax · MiniMax · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Ministral 3 14BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Ministral 3 3BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Ministral 3 8BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Mistral Medium 3.5Open weightsMistral AI · Mistral · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Mistral Small 3.1Open weightsMistral AI · Mistral · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Mistral Small 4Open weightsMistral AI · Mistral · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Qwen3 14BOpen weightsQwen · Qwen3 | 0% | Independent | group defaults | Partially comparable-59.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Qwen3 32BOpen weightsQwen · Qwen3 | 0% | Independent | group defaults | Partially comparable-59.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Qwen3 8BOpen weightsQwen · Qwen3 | 0% | Independent | group defaults | Partially comparable-59.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Qwen3 Coder NextOpen weightsQwen · Qwen3 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-59.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →