Updated 2 h ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 1,639
- Models
- 376
- Current leader
- gpt-5.6-sol 65.9%
Frontier over time · accuracy · variant=v2.1 · evaluator=Artificial Analysis
5 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 91.4%Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 91.0%Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 89.9%gpt-6-astra OpenAI Independent11 Sept 2026
- 89.5%gpt-6-astra OpenAI Independent11 Sept 2026
- 86.1%Claude Opus 5 Anthropic Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 183 models
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 101 | North Mini Code (free)Open weightsCohere · best of 2 rows | 35.6% | Independent | reasoningon | Partially comparable0.00 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 102 | gpt-5ClosedOpenAI · GPT 5 · best of 2 rows | 35.2% | Independent | reasoning_efforthigh | Partially comparable-0.37 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 103 | gpt-5-5-instant-06-26ClosedOpenAI · GPT 5.5 · best of 2 rows | 34.8% | Independent | reasoningon | Partially comparable-0.75 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 104 | g9v3-39a5bOpen weightsAI9Stars · best of 2 rows | 32.6% | Independent | reasoningon | Partially comparable-3.00 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 105 | Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.1 · best of 2 rows | 31.1% | Independent | reasoningon | Partially comparable-4.49 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 106 | Devstral 2Open weightsMistral AI · Devstral 2 · best of 2 rows | 30.3% | Independent | reasoningoff | Partially comparable-5.24 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 106 | k-exaoneOpen weightsLG AI Research · EXAONE · best of 2 rows | 30.3% | Independent | reasoningon | Partially comparable-5.24 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 108 | Devstral Small 2Open weightsMistral AI · Devstral · best of 2 rows | 29.6% | Independent | reasoningoff | Partially comparable-5.99 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 108 | nova-2-0-proClosedAmazon Web Services · Nova 2.0 · best of 6 rows | 29.6% | Independent | reasoningonreasoning_effortmedium | Partially comparable-5.99 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 110 | Qwen3.5-9BOpen weightsQwen · Qwen3.5 · best of 4 rows | 29.2% | Independent | reasoningon | Partially comparable-6.37 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 111 | Gemini 2.5 ProClosedGoogle · Gemini 2.5 · best of 2 rows | 28.5% | Independent | reasoningon | Partially comparable-7.12 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 112 | ling-3-0-tinyOpen weightsinclusionAI · best of 2 rows | 27.7% | Independent | reasoningon | Partially comparable-7.86 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 113 | Mercury 2ClosedInception · best of 2 rows | 27.3% | Independent | reasoningon | Partially comparable-8.24 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 113 | gemma-4-12BOpen weightsGoogle · Gemma 4 · best of 2 rows | 27.3% | Independent | reasoningon | Partially comparable-8.24 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 115 | granite-4.2-30bOpen weightsIBM · Granite 4.2 · best of 2 rows | 26.6% | Independent | reasoningon | Partially comparable-8.99 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 116 | Mistral Small 3.1Open weightsMistral AI · Mistral · best of 2 rows | 26.2% | Independent | reasoningoff | Partially comparable-9.36 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 116 | gpt-oss-120bOpen weightsOpenAI · gpt-oss · best of 4 rows | 26.2% | Independent | reasoning_efforthigh | Partially comparable-9.36 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 118 | Qwen3.5-4BOpen weightsQwen · Qwen3.5 · best of 4 rows | 25.8% | Independent | reasoningon | Partially comparable-9.74 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 119 | NVIDIA Nemotron 3.5 Lightning 30B A3BOpen weightsNVIDIA · Nemotron 3.5 · best of 2 rows | 24.3% | Independent | reasoningon | Partially comparable-11.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 119 | ling-2-6-flashOpen weightsinclusionAI · best of 2 rows | 24.3% | Independent | reasoningoff | Partially comparable-11.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 121 | command-a-plusOpen weightsCohere · Command · best of 2 rows | 22.9% | Independent | reasoningon | Partially comparable-12.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 122 | exaone-4-5-33bOpen weightsLG AI Research · EXAONE 4.5 · best of 2 rows | 21.4% | Independent | reasoningon | Partially comparable-14.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 123 | Mistral Small 4Open weightsMistral AI · Mistral · best of 2 rows | 21.0% | Independent | reasoningon | Partially comparable-14.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 124 | Trinity Large ThinkingOpen weightsArcee AI · best of 2 rows | 20.6% | Independent | reasoningon | Partially comparable-15.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 124 | nemotron-cascade-2-30b-a3bOpen weightsNVIDIA · Nemotron · best of 2 rows | 20.6% | Independent | reasoningon | Partially comparable-15.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 126 | deepseek-r1-0120Open weightsDeepSeek · DeepSeek · best of 2 rows | 19.1% | Independent | reasoningon | Partially comparable-16.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 127 | Granite 4.2 8BOpen weightsIBM · Granite 4.2 · best of 2 rows | 18.4% | Independent | reasoningon | Partially comparable-17.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 127 | hypernova-60bOpen weightsMultiverse Computing · best of 2 rows | 18.4% | Independent | reasoning_efforthigh | Partially comparable-17.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 129 | nova-2-0-liteClosedAmazon Web Services · Nova 2.0 · best of 2 rows | 16.1% | Independent | reasoningonreasoning_efforthigh | Partially comparable-19.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 130 | k2-think-v2Open weightsMBZUAI Institute of Foundation Models · best of 2 rows | 15.0% | Independent | reasoningon | Partially comparable-20.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 131 | DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows | 13.9% | Independent | reasoningoff | Partially comparable-21.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 131 | Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows | 13.9% | Independent | reasoningoff | Partially comparable-21.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 131 | gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 2 rows | 13.9% | Independent | reasoning_efforthigh | Partially comparable-21.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 131 | granite-4.2-3bOpen weightsIBM · Granite 4.2 · best of 2 rows | 13.9% | Independent | reasoningon | Partially comparable-21.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 135 | Magistral Medium 1.2ClosedMistral AI · Magistral · best of 2 rows | 12.4% | Independent | reasoningon | Partially comparable-23.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 135 | diffusiongemma-26b-a4bOpen weightsGoogle · best of 2 rows | 12.4% | Independent | reasoningon | Partially comparable-23.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 137 | Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows | 12.0% | Independent | reasoningoff | Partially comparable-23.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 137 | Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 2 rows | 12.0% | Independent | reasoningon | Partially comparable-23.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 137 | Solar Pro 3ClosedUpstage · Solar · best of 2 rows | 12.0% | Independent | reasoningon | Partially comparable-23.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 140 | celeris-1ClosedCeleris · best of 2 rows | 11.2% | Independent | reasoningoff | Partially comparable-24.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 141 | Claude Haiku 3.5ClosedAnthropic · Claude · best of 2 rows | 10.1% | Independent | reasoningoff | Partially comparable-25.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 141 | gpt-4.1-miniClosedOpenAI · GPT 4.1 · best of 2 rows | 10.1% | Independent | reasoningoff | Partially comparable-25.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 143 | Ministral 3 14BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 9.74% | Independent | reasoningoff | Partially comparable-25.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 144 | minicpm5-2bOpen weightsOpenBMB · best of 2 rows | 8.61% | Independent | reasoningon | Partially comparable-27.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 145 | Llama 4 MaverickOpen weightsMeta AI · Llama 4 · best of 2 rows | 7.87% | Independent | reasoningoff | Partially comparable-27.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 146 | Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 6.74% | Independent | reasoningon | Partially comparable-28.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 146 | Qwen3 Next 80B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows | 6.74% | Independent | reasoningon | Partially comparable-28.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 146 | nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 6.74% | Independent | reasoningon | Partially comparable-28.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 149 | g9v3-3bOpen weightsAI9Stars · best of 2 rows | 5.99% | Independent | reasoningon | Partially comparable-29.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 150 | Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows | 5.62% | Independent | reasoningoff | Partially comparable-30.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 150 | gpt-4o-miniClosedOpenAI · GPT 4 · best of 2 rows | 5.62% | Independent | reasoningoff | Partially comparable-30.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 152 | Qwen3 32BOpen weightsQwen · Qwen3 | 5.24% | Independent | group defaults | Partially comparable-30.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 152 | qwen3-32b-instructOpen weightsAlibaba Group · Qwen3 | 5.24% | Independent | reasoningon | Partially comparable-30.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 154 | Llama 3.3 70BOpen weightsMeta AI · Llama 3.3 · best of 2 rows | 4.87% | Independent | reasoningoff | Partially comparable-30.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 154 | Qwen3 14BOpen weightsQwen · Qwen3 | 4.87% | Independent | group defaults | Partially comparable-30.7 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 154 | qwen3-14b-instructOpen weightsAlibaba Group · Qwen3 | 4.87% | Independent | reasoningon | Partially comparable-30.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 157 | Gemma 3 27BOpen weightsGoogle · Gemma 3 · best of 2 rows | 4.49% | Independent | reasoningoff | Partially comparable-31.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 157 | LFM2.5-2.6B (free)Open weightsLiquid AI · LFM2.5 · best of 2 rows | 4.49% | Independent | reasoningon | Partially comparable-31.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 157 | Magistral Small 1.2Open weightsMistral AI · Magistral · best of 2 rows | 4.49% | Independent | reasoningon | Partially comparable-31.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 157 | o3-miniClosedOpenAI · OpenAI o-series · best of 2 rows | 4.49% | Independent | reasoning_efforthigh | Partially comparable-31.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 161 | Ministral 3 8BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 4.12% | Independent | reasoningoff | Partially comparable-31.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 162 | Llama 4 ScoutOpen weightsMeta AI · Llama 4 · best of 2 rows | 3.75% | Independent | reasoningoff | Partially comparable-31.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 162 | gpt-4.1-nanoClosedOpenAI · GPT 4.1 · best of 2 rows | 3.75% | Independent | reasoningoff | Partially comparable-31.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 162 | gpt-5-miniClosedOpenAI · GPT 5 · best of 2 rows | 3.75% | Independent | reasoning_efforthigh | Partially comparable-31.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 162 | nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 3.75% | Independent | reasoningon | Partially comparable-31.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 166 | granite-4.1-8bOpen weightsIBM · Granite 4.1 · best of 2 rows | 3.37% | Independent | reasoningoff | Partially comparable-32.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 167 | Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 4 rows | 3% | Independent | reasoningon | Partially comparable-32.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 168 | granite-4.1-30bOpen weightsIBM · Granite 4.1 · best of 2 rows | 2.62% | Independent | reasoningoff | Partially comparable-33.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 169 | Qwen3 8BOpen weightsQwen · Qwen3 | 2.25% | Independent | group defaults | Partially comparable-33.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 169 | qwen3-8b-instructOpen weightsAlibaba Group · Qwen3 | 2.25% | Independent | reasoningon | Partially comparable-33.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 171 | gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 2 rows | 1.87% | Independent | reasoningon | Partially comparable-33.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 172 | Llama 3.1 8BRestricted weightsMeta AI · Llama 3.1 · best of 2 rows | 1.50% | Independent | reasoningoff | Partially comparable-34.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 172 | qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 2 rows | 1.50% | Independent | reasoningon | Partially comparable-34.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 174 | granite-4.1-3bOpen weightsIBM · Granite 4.1 · best of 2 rows | 1.12% | Independent | reasoningoff | Partially comparable-34.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 174 | nanbeige4-1-3bOpen weightsNanbeige · best of 2 rows | 1.12% | Independent | reasoningon | Partially comparable-34.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 176 | gemma-3n-e4bOpen weightsGoogle · Gemma 3 · best of 2 rows | 0.75% | Independent | reasoningoff | Partially comparable-34.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 177 | Gemma 3 4BRestricted weightsGoogle · Gemma 3 · best of 2 rows | 0.37% | Independent | reasoningoff | Partially comparable-35.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 177 | Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 4 rows | 0.37% | Independent | reasoningoff | Partially comparable-35.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 177 | gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 2 rows | 0.37% | Independent | reasoningon | Partially comparable-35.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 177 | phi-4-miniOpen weightsMicrosoft · Phi4 · best of 2 rows | 0.37% | Independent | reasoningoff | Partially comparable-35.2 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 181 | Gemma 3 12BOpen weightsGoogle · Gemma 3 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-35.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 181 | Ministral 3 3BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-35.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 181 | minicpm-v4-6-1-3bOpen weightsOpenBMB · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-35.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →