Updated 54 min ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 1,639
- Models
- 376
- Current leader
- Claude Fable 5.1 91.4%
Score history · Qwen3.5-35B-A3B 6 rows
- Qwen3.5-35B-A3B
- 26.52%aa_slug=qwen3-5-35b-a3b · variant=hard · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
- 10.61%aa_slug=qwen3-5-35b-a3b-non-reasoning · variant=hard · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
- 40.82%aa_slug=qwen3-5-35b-a3b-non-reasoning · variant=v2.1 · evaluator=Artificial Analysis · reasoning=off12 Sept 2026
- 26.52%aa_slug=qwen3-5-35b-a3b · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
- 10.61%aa_slug=qwen3-5-35b-a3b-non-reasoning · variant=hard · evaluator=Artificial Analysis · reasoning=off11 Sept 2026
- 40.82%aa_slug=qwen3-5-35b-a3b-non-reasoning · variant=v2.1 · evaluator=Artificial Analysis · reasoning=off11 Sept 2026
Frontier over time · accuracy · variant=v2.1 · evaluator=Artificial Analysis
5 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 91.4%Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 91.0%Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 89.9%gpt-6-astra OpenAI Independent11 Sept 2026
- 89.5%gpt-6-astra OpenAI Independent11 Sept 2026
- 86.1%Claude Opus 5 Anthropic Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 183 models
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1ClosedAnthropic · Claude · best of 10 rows | 91.4% | Independent | reasoning_effortmax | leader | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 2 | gpt-6-astraClosedOpenAI · GPT 6 · best of 11 rows | 89.9% | Independent | reasoning_efforthigh | Partially comparable-1.50 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 3 | gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 12 rows | 89.5% | Independent | reasoning_effortxhigh | Partially comparable-1.88 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 4 | Claude Opus 5ClosedAnthropic · Claude · best of 10 rows | 89.1% | Independent | reasoning_effortmax | Comparable-2.25 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 5 | Grok 4.6ClosedxAI · Grok · best of 8 rows | 88.4% | Independent | reasoning_efforthigh | Partially comparable-3.00 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 6 | gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 12 rows | 88.0% | Independent | reasoning_effortmax | Comparable-3.38 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 7 | Gemini 3.8 FlashClosedGoogle · Gemini 3.8 · best of 6 rows | 87.6% | Independent | reasoning_efforthigh | Partially comparable-3.75 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 8 | Qwen3.8 FlashOpen weightsQwen · Qwen3.8 · best of 2 rows | 86.1% | Independent | reasoningon | Partially comparable-5.25 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 9 | Gemini 3.7 FlashClosedGoogle · Gemini 3.7 · best of 6 rows | 85.8% | Independent | reasoning_efforthigh | Partially comparable-5.62 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 10 | Muse Spark 1.3ClosedMeta AI · best of 4 rows | 85.4% | Independent | reasoning_effortxhigh | Partially comparable-6.00 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 11 | Kimi K3Open weightsMoonshot AI · Kimi · best of 4 rows | 85.0% | Independent | reasoning_effortmax | Comparable-6.37 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 12 | Claude Fable 5ClosedAnthropic · Claude · best of 2 rows | 84.6% | Independent | reasoningon | Partially comparable-6.75 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 12 | Claude Opus 4.8ClosedAnthropic · Claude · best of 2 rows | 84.6% | Independent | reasoning_effortmax | Comparable-6.75 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 14 | GLM 5.3 FlashOpen weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows | 84.3% | Independent | reasoningon | Partially comparable-7.12 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 14 | gpt-5.5ClosedOpenAI · GPT 5.5 · best of 10 rows | 84.3% | Independent | reasoning_effortxhigh | Partially comparable-7.12 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 16 | GLM 5.3Open weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows | 83.9% | Independent | reasoning_effortmax | Comparable-7.49 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 17 | Claude Opus 4.7ClosedAnthropic · Claude · best of 2 rows | 83.2% | Independent | reasoning_effortmax | Comparable-8.24 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 18 | Qwen3.8 2.4T A95BOpen weightsQwen · Qwen3.8 · best of 2 rows | 82.0% | Independent | reasoningon | Partially comparable-9.37 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 18 | agnes-3-0-flashClosedSapiens AI · best of 2 rows | 82.0% | Independent | reasoningon | Partially comparable-9.37 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 20 | Grok 4.5ClosedxAI · Grok · best of 2 rows | 81.7% | Independent | reasoning_efforthigh | Partially comparable-9.74 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 21 | Qwen 3.8 MaxClosedQwen · Qwen3.8 · best of 2 rows | 81.3% | Independent | reasoningon | Partially comparable-10.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 22 | gpt-5.6-lunaClosedOpenAI · GPT 5.6 · best of 12 rows | 80.9% | Independent | reasoning_effortmax | Comparable-10.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 23 | Claude Sonnet 5ClosedAnthropic · Claude · best of 4 rows | 80.5% | Independent | reasoning_effortmax | Comparable-10.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 24 | Muse Spark 1.2ClosedMeta AI · best of 2 rows | 80.2% | Independent | reasoning_effortxhigh | Partially comparable-11.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 25 | Qwen3.8 27BOpen weightsQwen · Qwen3.8 · best of 8 rows | 79.8% | Independent | reasoning_effortxhigh | Partially comparable-11.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 26 | Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 2 rows | 78.7% | Independent | reasoningon | Partially comparable-12.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 26 | deepseek-v4-flashOpen weightsDeepSeek · DeepSeek · best of 2 rows | 78.7% | Independent | reasoning_effortmax | Comparable-12.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 26 | deepseek-v4-proClosedDeepSeek · V4 · best of 2 rows | 78.7% | Independent | reasoning_effortmax | Comparable-12.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 29 | gpt-5.4ClosedOpenAI · GPT 5.4 · best of 2 rows | 78.3% | Independent | reasoning_effortxhigh | Partially comparable-13.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 30 | Muse Spark 1.1ClosedMeta AI · best of 2 rows | 77.9% | Independent | reasoning_effortxhigh | Partially comparable-13.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 30 | Z.ai GLM 5.2Open weightsZ.ai (Zhipu AI) · GLM5.2 · best of 4 rows | 77.9% | Independent | reasoning_effortmax | Comparable-13.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 32 | Gemini 3.6 FlashClosedGoogle · Gemini 3.6 · best of 2 rows | 77.5% | Independent | reasoningon | Partially comparable-13.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 33 | motif-3Open weightsMotif Technologies · best of 2 rows | 74.9% | Independent | reasoningon | Partially comparable-16.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 34 | Qwen3.7 MaxClosedQwen · Qwen3.7 · best of 2 rows | 74.5% | Independent | reasoningon | Partially comparable-16.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 35 | deepseek-v4-flash-visionClosedDeepSeek · DeepSeek · best of 2 rows | 74.2% | Independent | reasoning_effortmax | Comparable-17.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 36 | Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.1 · best of 2 rows | 73.8% | Independent | reasoningon | Partially comparable-17.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 37 | k2-horizon-375b-a23bOpen weightsMBZUAI Institute of Foundation Models · best of 2 rows | 71.9% | Independent | reasoningon | Partially comparable-19.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 38 | Claude Sonnet 4.6ClosedAnthropic · Claude · best of 2 rows | 71.2% | Independent | reasoningadaptivereasoning_effortmax | Partially comparable-20.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 39 | motif-0714ClosedMotif Technologies · best of 2 rows | 70.8% | Independent | reasoningon | Partially comparable-20.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 40 | KAT-Coder-Pro V2ClosedKwaipilot · best of 2 rows | 70.0% | Independent | reasoningoff | Partially comparable-21.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 41 | agnes-2-5-pro-betaClosedSapiens AI · best of 2 rows | 69.7% | Independent | reasoningon | Partially comparable-21.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 41 | apodex-1-1ClosedApodex · best of 2 rows | 69.7% | Independent | reasoningon | Partially comparable-21.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 43 | quasar-438bClosedMultiverse Computing · best of 2 rows | 69.3% | Independent | reasoning_effortmax | Comparable-22.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 44 | nex-n2-proOpen weightsNex AGI · best of 2 rows | 67.8% | Independent | reasoningon | Partially comparable-23.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 45 | Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi · best of 2 rows | 67.4% | Independent | reasoningon | Partially comparable-24.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 46 | agnes-2-5-pro-alphaOpen weightsSapiens AI · best of 2 rows | 67.0% | Independent | reasoningon | Partially comparable-24.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 47 | Kimi K2.6Open weightsMoonshot AI · Kimi · best of 2 rows | 65.9% | Independent | reasoningon | Partially comparable-25.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 48 | MiMo-V2.5-ProOpen weightsXiaomi · best of 2 rows | 65.2% | Independent | reasoningon | Partially comparable-26.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 48 | MiniMax-M3Open weightsMiniMax · MiniMax · best of 2 rows | 65.2% | Independent | reasoningon | Partially comparable-26.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 50 | deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek · best of 3 rows | 64.8% | Independent | reasoning_efforthigh | Partially comparable-26.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 50 | deepseek-v4-pro-0424-highOpen weightsDeepSeek · DeepSeek | 64.8% | Independent | group defaults | Partially comparable-26.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 52 | Hy3Open weightsTencent · best of 2 rows | 64.4% | Independent | reasoningon | Partially comparable-27.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 52 | Ling 3.0 Flash VLOpen weightsinclusionAI · best of 2 rows | 64.4% | Independent | reasoningon | Partially comparable-27.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 54 | MiMo-V2.5Open weightsXiaomi · best of 2 rows | 63.7% | Independent | reasoningon | Partially comparable-27.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 55 | muse-sparkClosedMeta AI · best of 2 rows | 62.2% | Independent | reasoningon | Partially comparable-29.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 56 | GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 2 rows | 61.8% | Independent | reasoningon | Partially comparable-29.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 56 | deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek · best of 3 rows | 61.8% | Independent | reasoning_effortmax | Comparable-29.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 56 | mimo-v2-flashOpen weightsXiaomi · best of 2 rows | 61.8% | Independent | reasoningoff | Partially comparable-29.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 59 | Qwen3.6 PlusClosedQwen · Qwen3.6 · best of 2 rows | 61.4% | Independent | reasoningon | Partially comparable-30.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | Qwen 3.7 PlusClosedQwen · Qwen3.7 · best of 2 rows | 61.0% | Independent | reasoningon | Partially comparable-30.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 61 | Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 4 rows | 60.7% | Independent | reasoningon | Partially comparable-30.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 61 | gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 2 rows | 60.7% | Independent | reasoning_effortxhigh | Partially comparable-30.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 63 | jt-4-1-flash-236b-a21bClosedChina Mobile · best of 2 rows | 59.5% | Independent | reasoningoff | Partially comparable-31.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 64 | gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 2 rows | 59.2% | Independent | reasoning_effortxhigh | Partially comparable-32.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 65 | Solar Pro 4ClosedUpstage · Solar · best of 2 rows | 57.3% | Independent | reasoningon | Partially comparable-34.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 66 | deepseek-v4-flash-0420-highOpen weightsDeepSeek · DeepSeek | 56.9% | Independent | group defaults | Partially comparable-34.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 67 | Claude Sonnet 4.5ClosedAnthropic · Claude · best of 2 rows | 55.8% | Independent | reasoningon | Partially comparable-35.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 68 | Ling 3.0 FlashOpen weightsinclusionAI · best of 2 rows | 55.4% | Independent | reasoningon | Partially comparable-36.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 68 | MiniMax M2.7Open weightsMiniMax · MiniMax · best of 2 rows | 55.4% | Independent | reasoningon | Partially comparable-36.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 70 | InklingOpen weightsThinking Machines · best of 2 rows | 55.1% | Independent | reasoningon | Partially comparable-36.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 70 | Inkling SmallOpen weightsThinking Machines · best of 2 rows | 55.1% | Independent | reasoningon | Partially comparable-36.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 72 | Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 53.9% | Independent | reasoningon | Partially comparable-37.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 73 | Gemini 3.5 Flash-LiteClosedGoogle · Gemini 3.5 · best of 2 rows | 53.6% | Independent | reasoningon | Partially comparable-37.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 74 | gpt-5.1ClosedOpenAI · GPT 5.1 · best of 2 rows | 52.4% | Independent | reasoning_efforthigh | Partially comparable-39.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 75 | grok-build-0-1-06-16ClosedSpaceXAI · Grok · best of 2 rows | 52.1% | Independent | reasoningon | Partially comparable-39.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 76 | Muse Glimmer 30BOpen weightsMeta AI · best of 2 rows | 51.7% | Independent | reasoning_efforthigh | Partially comparable-39.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 77 | Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 2 rows | 51.3% | Independent | reasoningon | Partially comparable-40.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 78 | Mistral Medium 3.5Open weightsMistral AI · Mistral · best of 2 rows | 50.6% | Independent | reasoningon | Partially comparable-40.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 79 | LongCat 2.0Open weightsMeituan · best of 2 rows | 50.2% | Independent | reasoningon | Partially comparable-41.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 80 | GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows | 49.4% | Independent | reasoningon | Partially comparable-42.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 81 | Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 4 rows | 47.6% | Independent | reasoningon | Partially comparable-43.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 82 | DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 3 rows | 46.8% | Independent | reasoningon | Partially comparable-44.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 82 | DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V3 | 46.8% | Independent | group defaults | Partially comparable-44.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 84 | Kimi K2.5Open weightsMoonshot AI · Kimi · best of 2 rows | 45.7% | Independent | reasoningon | Partially comparable-45.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 85 | GLM 4.7Open weightsZ.ai (Zhipu AI) · GLM4.7 · best of 2 rows | 45.3% | Independent | reasoningon | Partially comparable-46.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 86 | DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 2 rows | 44.9% | Independent | reasoningon | Partially comparable-46.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 86 | Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 4 rows | 44.9% | Independent | reasoningon | Partially comparable-46.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 88 | Claude Haiku 4.5ClosedAnthropic · Claude · best of 2 rows | 44.2% | Independent | reasoningon | Partially comparable-47.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 88 | solar-open2-250bOpen weightsUpstage · Solar · best of 2 rows | 44.2% | Independent | reasoningon | Partially comparable-47.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 90 | Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 4 rows | 43.5% | Independent | reasoningon | Partially comparable-47.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 91 | ring-2-6-1tOpen weightsinclusionAI · best of 2 rows | 43.1% | Independent | reasoningon | Partially comparable-48.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 92 | Qwen3.5-35B-A3BOpen weightsQwen · Qwen3.5 · best of 2 rows | 40.8% | Independent | reasoningoff | Partially comparable-50.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 93 | k-exaone-2-0-0803Open weightsLG AI Research · EXAONE 2.0 · best of 2 rows | 40.5% | Independent | reasoningon | Partially comparable-50.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 94 | Grok 4.3ClosedxAI · Grok · best of 4 rows | 39.7% | Independent | reasoning_efforthigh | Partially comparable-51.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 95 | Step 3.7 FlashOpen weightsStepFun · Step3.7 · best of 2 rows | 39.3% | Independent | reasoningon | Partially comparable-52.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 96 | Gemma 4 26B A4BOpen weightsGoogle · Gemma 4 · best of 2 rows | 39.0% | Independent | reasoningon | Partially comparable-52.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 96 | a-x-k2Open weightsSK Telecom · best of 2 rows | 39.0% | Independent | reasoningon | Partially comparable-52.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 98 | Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 38.6% | Independent | reasoningon | Partially comparable-52.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 99 | Qwen3 Coder NextOpen weightsQwen · Qwen3 · best of 2 rows | 38.2% | Independent | reasoningoff | Partially comparable-53.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 100 | Claude Sonnet 4ClosedAnthropic · Claude · best of 2 rows | 36.3% | Independent | reasoningon | Partially comparable-55.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →