Updated 5 h ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 339
- Models
- 129
- Current leader
- Claude Fable 5.1 63.1%
Score history · qwen3-8b-instruct 1 row
Not enough history to chart — a single observation (27.89% on 12 Sept 2026). Rows under different configurations count separately; the list below shows each one.
- 27.89%aa_slug=qwen3-8b-instruct-reasoning · evaluator=Artificial Analysis · reasoning=on · index_version=4.312 Sept 2026
Frontier over time · accuracy · evaluator=Artificial Analysis
5 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 63.1%Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 61%Claude Fable 5 Anthropic Independent11 Sept 2026
- 59.8%Gemini 3.7 Flash Google Independent11 Sept 2026
- 59.5%Kimi K3 Moonshot AI Independent11 Sept 2026
- 51.5%Claude Opus 5 Anthropic Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 129 models
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1ClosedAnthropic · Claude · best of 10 rows | 63.1% | Independent | reasoning_effortmax | leader | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 2 | Claude Fable 5ClosedAnthropic · Claude · best of 2 rows | 61% | Independent | reasoningon | Partially comparable-2.08 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 3 | Gemini 3.7 FlashClosedGoogle · Gemini 3.7 · best of 6 rows | 59.8% | Independent | reasoning_effortmedium | Partially comparable-3.24 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 4 | Muse Spark 1.3ClosedMeta AI · best of 4 rows | 59.7% | Independent | reasoning_effortxhigh | Partially comparable-3.36 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 5 | Kimi K3Open weightsMoonshot AI · Kimi · best of 4 rows | 59.5% | Independent | reasoning_effortmax | Comparable-3.59 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 6 | GLM 5.3Open weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows | 59.0% | Independent | reasoning_effortmax | Comparable-4.05 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 7 | Muse Spark 1.1ClosedMeta AI · best of 2 rows | 58.8% | Independent | reasoning_effortxhigh | Partially comparable-4.28 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 8 | Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.1 · best of 2 rows | 58.7% | Independent | reasoningon | Partially comparable-4.40 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 9 | gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 10 rows | 57.8% | Independent | reasoning_efforthigh | Partially comparable-5.33 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 10 | Muse Spark 1.2ClosedMeta AI · best of 2 rows | 57.4% | Independent | reasoning_effortxhigh | Partially comparable-5.67 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 11 | Gemini 3.8 FlashClosedGoogle · Gemini 3.8 · best of 6 rows | 56.6% | Independent | reasoning_efforthigh | Partially comparable-6.48 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 12 | Grok 4.6ClosedxAI · Grok · best of 8 rows | 56.5% | Independent | reasoning_efforthigh | Partially comparable-6.60 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 12 | gpt-6-astraClosedOpenAI · GPT 6 · best of 11 rows | 56.5% | Independent | reasoning_effortmax | Comparable-6.60 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 14 | Claude Opus 5ClosedAnthropic · Claude · best of 10 rows | 56.4% | Independent | reasoning_effortmax | Comparable-6.71 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 15 | gpt-5.5ClosedOpenAI · GPT 5.5 · best of 6 rows | 56.1% | Independent | reasoning_efforthigh | Partially comparable-6.95 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 16 | Grok 4.5ClosedxAI · Grok · best of 2 rows | 55.0% | Independent | reasoning_efforthigh | Partially comparable-8.10 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 16 | gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 10 rows | 55.0% | Independent | reasoning_effortmax | Comparable-8.10 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 18 | Claude Opus 4.8ClosedAnthropic · Claude · best of 2 rows | 54.4% | Independent | reasoning_effortmax | Comparable-8.68 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 19 | Claude Sonnet 5ClosedAnthropic · Claude · best of 10 rows | 54.3% | Independent | reasoning_efforthigh | Partially comparable-8.80 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 20 | Qwen3.8 2.4T A95BOpen weightsQwen · Qwen3.8 · best of 2 rows | 54.0% | Independent | reasoningon | Partially comparable-9.03 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 21 | Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 2 rows | 53.9% | Independent | reasoningon | Partially comparable-9.14 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 22 | gpt-5.6-lunaClosedOpenAI · GPT 5.6 · best of 10 rows | 53.6% | Independent | reasoning_effortmax | Comparable-9.49 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 23 | Gemini 3.6 FlashClosedGoogle · Gemini 3.6 · best of 2 rows | 53.4% | Independent | reasoningon | Partially comparable-9.72 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 24 | Qwen 3.8 MaxClosedQwen · Qwen3.8 · best of 2 rows | 53.2% | Independent | reasoningon | Partially comparable-9.84 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 25 | gpt-5-5-instant-06-26ClosedOpenAI · GPT 5.5 · best of 2 rows | 52.5% | Independent | reasoningon | Partially comparable-10.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 26 | gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 2 rows | 52.1% | Independent | reasoning_effortxhigh | Partially comparable-11.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 27 | DeepSeek-V4.1-FlashOpen weightsDeepSeek · DeepSeek · best of 2 rows | 51.9% | Independent | reasoning_effortmax | Comparable-11.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 28 | GLM 5.3 FlashOpen weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows | 51.6% | Independent | reasoningon | Partially comparable-11.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 28 | agnes-3-0-flashClosedSapiens AI · best of 2 rows | 51.6% | Independent | reasoningon | Partially comparable-11.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 30 | Kimi K2.6Open weightsMoonshot AI · Kimi · best of 2 rows | 51.5% | Independent | reasoningon | Partially comparable-11.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 31 | Z.ai GLM 5.2Open weightsZ.ai (Zhipu AI) · GLM5.2 · best of 2 rows | 51.2% | Independent | reasoning_effortmax | Comparable-11.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 32 | deepseek-v4-proClosedDeepSeek · V4 · best of 2 rows | 51.0% | Independent | reasoning_effortmax | Comparable-12.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 33 | deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek · best of 2 rows | 50.8% | Independent | reasoning_effortmax | Comparable-12.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 34 | MiMo-V2.5-ProOpen weightsXiaomi · best of 2 rows | 50.6% | Independent | reasoningon | Partially comparable-12.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 34 | Qwen3.8 FlashOpen weightsQwen · Qwen3.8 · best of 2 rows | 50.6% | Independent | reasoningon | Partially comparable-12.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 36 | deepseek-v4-flashOpen weightsDeepSeek · DeepSeek · best of 2 rows | 50.4% | Independent | reasoning_effortmax | Comparable-12.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 37 | Claude Sonnet 4.6ClosedAnthropic · Claude · best of 2 rows | 50.1% | Independent | reasoningadaptivereasoning_effortmax | Partially comparable-13.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 37 | MiniMax M2.7Open weightsMiniMax · MiniMax · best of 2 rows | 50.1% | Independent | reasoningon | Partially comparable-13.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 39 | Inkling SmallOpen weightsThinking Machines · best of 2 rows | 49.6% | Independent | reasoningon | Partially comparable-13.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 39 | deepseek-v4-flash-visionClosedDeepSeek · DeepSeek · best of 2 rows | 49.6% | Independent | reasoning_effortmax | Comparable-13.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 41 | Qwen3.7 MaxClosedQwen · Qwen3.7 · best of 2 rows | 49.5% | Independent | reasoningon | Partially comparable-13.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 42 | agnes-2-5-pro-betaClosedSapiens AI · best of 2 rows | 48.8% | Independent | reasoningon | Partially comparable-14.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 43 | Hy3Open weightsTencent · best of 2 rows | 48.6% | Independent | reasoningon | Partially comparable-14.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 44 | Grok 4.3ClosedxAI · Grok · best of 4 rows | 48.3% | Independent | reasoning_efforthigh | Partially comparable-14.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 45 | quasar-438bClosedMultiverse Computing · best of 2 rows | 48.1% | Independent | reasoning_effortmax | Comparable-14.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 46 | solar-open2-250bOpen weightsUpstage · Solar · best of 2 rows | 48.0% | Independent | reasoningon | Partially comparable-15.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 47 | Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi · best of 2 rows | 47.8% | Independent | reasoningon | Partially comparable-15.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 48 | gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 2 rows | 47.2% | Independent | reasoning_effortxhigh | Partially comparable-15.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | MiniMax-M3Open weightsMiniMax · MiniMax · best of 2 rows | 47.1% | Independent | reasoningon | Partially comparable-16.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 50 | InklingOpen weightsThinking Machines · best of 2 rows | 47.0% | Independent | reasoningon | Partially comparable-16.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 51 | Qwen3.8 27BOpen weightsQwen · Qwen3.8 · best of 8 rows | 46.6% | Independent | reasoning_effortxhigh | Partially comparable-16.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 52 | Gemini 2.5 ProClosedGoogle · Gemini 2.5 · best of 2 rows | 46.3% | Independent | reasoningon | Partially comparable-16.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 53 | Qwen 3.7 PlusClosedQwen · Qwen3.7 · best of 2 rows | 46.1% | Independent | reasoningon | Partially comparable-17.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 54 | Claude Sonnet 4.5ClosedAnthropic · Claude · best of 2 rows | 45.7% | Independent | reasoningon | Partially comparable-17.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 55 | Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 2 rows | 45.5% | Independent | reasoningon | Partially comparable-17.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 55 | apodex-1-1ClosedApodex · best of 2 rows | 45.5% | Independent | reasoningon | Partially comparable-17.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 57 | deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek · best of 3 rows | 45.3% | Independent | reasoning_effortmax | Comparable-17.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 58 | ring-2-6-1tOpen weightsinclusionAI · best of 2 rows | 45.0% | Independent | reasoningon | Partially comparable-18.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 59 | Muse Glimmer 30BOpen weightsMeta AI · best of 2 rows | 44.9% | Independent | reasoning_efforthigh | Partially comparable-18.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 2 rows | 44.8% | Independent | reasoningon | Partially comparable-18.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 2 rows | 44.8% | Independent | reasoningon | Partially comparable-18.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 62 | Solar Pro 4ClosedUpstage · Solar · best of 2 rows | 44.6% | Independent | reasoningon | Partially comparable-18.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 63 | Ling 3.0 Flash VLOpen weightsinclusionAI · best of 2 rows | 44.2% | Independent | reasoningon | Partially comparable-18.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 64 | MiMo-V2.5Open weightsXiaomi · best of 2 rows | 43.9% | Independent | reasoningon | Partially comparable-19.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 64 | Step 3.7 FlashOpen weightsStepFun · Step3.7 · best of 2 rows | 43.9% | Independent | reasoningon | Partially comparable-19.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 66 | Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.1 · best of 2 rows | 43.4% | Independent | reasoningon | Partially comparable-19.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 67 | nex-n2-proOpen weightsNex AGI · best of 2 rows | 43.3% | Independent | reasoningon | Partially comparable-19.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 68 | agnes-2-5-pro-alphaOpen weightsSapiens AI · best of 2 rows | 42.9% | Independent | reasoningon | Partially comparable-20.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 68 | k2-horizon-375b-a23bOpen weightsMBZUAI Institute of Foundation Models · best of 2 rows | 42.9% | Independent | reasoningon | Partially comparable-20.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 70 | Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 2 rows | 42.8% | Independent | reasoningon | Partially comparable-20.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 70 | o3-miniClosedOpenAI · OpenAI o-series · best of 2 rows | 42.8% | Independent | reasoning_efforthigh | Partially comparable-20.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 72 | Claude Haiku 4.5ClosedAnthropic · Claude · best of 2 rows | 42.3% | Independent | reasoningon | Partially comparable-20.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 72 | motif-3Open weightsMotif Technologies · best of 2 rows | 42.3% | Independent | reasoningon | Partially comparable-20.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 74 | Ling 3.0 FlashOpen weightsinclusionAI · best of 2 rows | 42.0% | Independent | reasoningon | Partially comparable-21.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 74 | k-exaone-2-0-0803Open weightsLG AI Research · EXAONE 2.0 · best of 2 rows | 42.0% | Independent | reasoningon | Partially comparable-21.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 76 | Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 2 rows | 41.4% | Independent | reasoningon | Partially comparable-21.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 77 | Gemini 3.5 Flash-LiteClosedGoogle · Gemini 3.5 · best of 2 rows | 41.3% | Independent | reasoningon | Partially comparable-21.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 78 | a-x-k2Open weightsSK Telecom · best of 2 rows | 41.0% | Independent | reasoningon | Partially comparable-22.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 79 | Trinity Large ThinkingOpen weightsArcee AI · best of 2 rows | 40.6% | Independent | reasoningon | Partially comparable-22.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 80 | Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 40.3% | Independent | reasoningon | Partially comparable-22.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 81 | Mistral Medium 3.5Open weightsMistral AI · Mistral · best of 2 rows | 40.2% | Independent | reasoningon | Partially comparable-22.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 81 | deepseek-v4-flash-0420-highOpen weightsDeepSeek · DeepSeek | 40.2% | Independent | group defaults | Partially comparable-22.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 83 | Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 2 rows | 39.7% | Independent | reasoningon | Partially comparable-23.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 84 | DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows | 39% | Independent | reasoningoff | Partially comparable-24.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 84 | gpt-5-miniClosedOpenAI · GPT 5 · best of 2 rows | 39% | Independent | reasoning_efforthigh | Partially comparable-24.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 86 | gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 2 rows | 38.9% | Independent | reasoning_efforthigh | Partially comparable-24.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 87 | Mistral Small 4Open weightsMistral AI · Mistral · best of 2 rows | 38.8% | Independent | reasoningon | Partially comparable-24.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 87 | North Mini Code (free)Open weightsCohere · best of 2 rows | 38.8% | Independent | reasoningon | Partially comparable-24.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 89 | command-a-plusOpen weightsCohere · Command · best of 2 rows | 38.5% | Independent | reasoningon | Partially comparable-24.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 90 | deepseek-r1-0120Open weightsDeepSeek · DeepSeek · best of 2 rows | 38.3% | Independent | reasoningon | Partially comparable-24.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 91 | DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 2 rows | 38.0% | Independent | reasoningon | Partially comparable-25.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 92 | granite-4.2-30bOpen weightsIBM · Granite 4.2 · best of 2 rows | 37.9% | Independent | reasoningon | Partially comparable-25.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 93 | Mercury 2ClosedInception · best of 2 rows | 37.7% | Independent | reasoningon | Partially comparable-25.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 94 | g9v3-39a5bOpen weightsAI9Stars · best of 2 rows | 36.8% | Independent | reasoningon | Partially comparable-26.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 95 | Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows | 36.6% | Independent | reasoningoff | Partially comparable-26.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 95 | Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 2 rows | 36.6% | Independent | reasoningon | Partially comparable-26.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 97 | LongCat 2.0Open weightsMeituan · best of 2 rows | 36.3% | Independent | reasoningon | Partially comparable-26.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 98 | Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 36.2% | Independent | reasoningon | Partially comparable-26.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 98 | Qwen3 Coder NextOpen weightsQwen · Qwen3 · best of 2 rows | 36.2% | Independent | reasoningoff | Partially comparable-26.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 100 | Qwen3 32BOpen weightsQwen · Qwen3 | 36% | Independent | group defaults | Partially comparable-27.1 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →