Artificial Analysis Intelligence Index
composite of several evaluations run by Artificial Analysis
Updated 3 h ago · first seen 11 Sept 2026
- Metric
- index ↑
- Current results
- 1,272
- Models
- 477
- Current leader
- Claude Fable 5.1 53.4
Score history · k2-horizon-375b-a23b 2 rows
- k2-horizon-375b-a23b
- 30.7aa_slug=k2-horizon-375b-a23b · version=4.3 · estimated=true · reasoning=on12 Sept 2026
- 30.7aa_slug=k2-horizon-375b-a23b · version=4.3 · estimated=true11 Sept 2026
Frontier over time · index
8 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 53.4Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 53.2Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 51.2Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 51.0gpt-6-astra OpenAI Independent11 Sept 2026
- 49.7gpt-6-astra OpenAI Independent11 Sept 2026
- 45.1Claude Opus 5 Anthropic Independent11 Sept 2026
- 14.6grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026
- 5.71llama-2-chat-7b Meta AI Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 477 models · trust independent-evaluator
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 201 | Qwen3 VL 32B InstructOpen weightsQwen · Qwen3 · best of 4 rows | 11.9 | Independent | reasoningonversion4.3 | Partially comparable0.00 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 201 | ling-3-0-tinyOpen weightsinclusionAI · best of 2 rows | 11.9 | Independent | reasoningonversion4.3 | Partially comparable0.00 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 203 | Magistral Medium 1.2ClosedMistral AI · Magistral · best of 2 rows | 11.8 | Independent | reasoningonversion4.3 | Partially comparable-0.04 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 204 | Sonar Reasoning ProClosedPerplexity AI · Sonar · best of 2 rows | 11.8 | Independent | reasoningonversion4.3 | Partially comparable-0.05 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 205 | Granite 4.2 8BOpen weightsIBM · Granite 4.2 · best of 2 rows | 11.8 | Independent | reasoningonversion4.3 | Partially comparable-0.06 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 206 | hypernova-60bOpen weightsMultiverse Computing · best of 2 rows | 11.7 | Independent | reasoning_efforthighversion4.3 | Partially comparable-0.13 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 207 | MiniMax-M1-80kOpen weightsMiniMax · MiniMax · best of 2 rows | 11.7 | Independent | reasoningonversion4.3 | Partially comparable-0.14 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | nemotron-cascade-2-30b-a3bOpen weightsNVIDIA · Nemotron · best of 2 rows | 11.7 | Independent | reasoningonversion4.3 | Partially comparable-0.20 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 209 | gemini-2-5-flash-reasoning-04-2025ClosedGoogle · Gemini 2.5 · best of 3 rows | 11.7 | Independent | reasoningonversion4.3 | Partially comparable-0.22 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 210 | Mercury 2ClosedInception · best of 2 rows | 11.5 | Independent | reasoningonversion4.3 | Partially comparable-0.36 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 211 | k2-think-v2Open weightsMBZUAI Institute of Foundation Models · best of 2 rows | 11.5 | Independent | reasoningonversion4.3 | Partially comparable-0.37 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 212 | longcat-flash-liteOpen weightsLongCat · best of 2 rows | 11.5 | Independent | reasoningoffversion4.3 | Partially comparable-0.40 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 213 | Mistral Small 4Open weightsMistral AI · Mistral · best of 4 rows | 11.4 | Independent | reasoningonversion4.3 | Partially comparable-0.42 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 214 | deepseek-r1-0120Open weightsDeepSeek · DeepSeek · best of 2 rows | 11.4 | Independent | reasoningonversion4.3 | Partially comparable-0.46 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 215 | o1 PreviewClosedOpenAI · OpenAI o-series · best of 2 rows | 11.4 | Independent | reasoningonversion4.3 | Partially comparable-0.49 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 216 | hyperclova-x-seed-think-32bOpen weightsNaver · Seed · best of 2 rows | 11.4 | Independent | reasoningonversion4.3 | Partially comparable-0.51 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 217 | GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 4 rows | 11.2 | Independent | reasoningonversion4.3 | Partially comparable-0.65 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 218 | Qwen3 Next 80B A3B InstructOpen weightsQwen · Qwen3 · best of 4 rows | 11.2 | Independent | reasoningonversion4.3 | Partially comparable-0.67 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 219 | GLM 4.5 AirOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows | 11.1 | Independent | reasoningonversion4.3 | Partially comparable-0.78 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | mi-dm-k-2-5-pro-dec28ClosedKorea Telecom · best of 2 rows | 11.0 | Independent | reasoningonversion4.3 | Partially comparable-0.83 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 221 | ring-1tOpen weightsinclusionAI · best of 2 rows | 10.9 | Independent | reasoningonversion4.3 | Partially comparable-0.97 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 222 | Trinity Large ThinkingOpen weightsArcee AI · best of 2 rows | 10.9 | Independent | reasoningonversion4.3 | Partially comparable-0.99 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 223 | g9v3-3bOpen weightsAI9Stars · best of 2 rows | 10.8 | Independent | reasoningonversion4.3 | Partially comparable-1.03 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 224 | intellect-3Open weightsPrime Intellect · best of 2 rows | 10.6 | Independent | reasoningonversion4.3 | Partially comparable-1.26 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 225 | gpt-5-chatgptClosedOpenAI · GPT 5 · best of 2 rows | 10.4 | Independent | reasoningoffversion4.3 | Partially comparable-1.43 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 226 | solar-open-100b-reasoningOpen weightsUpstage · Solar · best of 2 rows | 10.4 | Independent | reasoningonversion4.3 | Partially comparable-1.51 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 227 | Gemini 2.5 Flash-LiteClosedGoogle · Gemini 2.5 · best of 5 rows | 10.3 | Independent | reasoningonversion4.3 | Partially comparable-1.52 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 227 | gemini-2-5-flash-lite-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 3 rows | 10.3 | Independent | reasoningonversion4.3 | Partially comparable-1.52 | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 229 | nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 10.3 | Independent | reasoningonversion4.3 | Partially comparable-1.62 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 230 | gpt-4.1-miniClosedOpenAI · GPT 4.1 · best of 2 rows | 10.2 | Independent | reasoningoffversion4.3 | Partially comparable-1.71 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 231 | Qwen3 Coder NextOpen weightsQwen · Qwen3 · best of 2 rows | 10.1 | Independent | reasoningoffversion4.3 | Partially comparable-1.82 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 232 | MiniMax-M1-40kOpen weightsMiniMax · MiniMax · best of 2 rows | 9.99 | Independent | reasoningonversion4.3 | Partially comparable-1.88 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 233 | gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 4 rows | 9.95 | Independent | reasoning_effortlowversion4.3 | Partially comparable-1.92 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 234 | Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows | 9.89 | Independent | reasoningoffversion4.3 | Partially comparable-1.98 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 235 | k2-v2Open weightsMBZUAI Institute of Foundation Models · best of 6 rows | 9.87 | Independent | reasoning_efforthighversion4.3 | Partially comparable-2.00 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 236 | qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 4 rows | 9.81 | Independent | reasoningonversion4.3 | Partially comparable-2.06 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | o1-miniClosedOpenAI · OpenAI o-series · best of 2 rows | 9.77 | Independent | reasoningonversion4.3 | Partially comparable-2.10 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 238 | DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows | 9.72 | Independent | reasoningoffversion4.3 | Partially comparable-2.15 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 239 | Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows | 9.71 | Independent | reasoningoffversion4.3 | Partially comparable-2.16 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 240 | ling-2-6-flashOpen weightsinclusionAI · best of 2 rows | 9.68 | Independent | reasoningoffversion4.3 | Partially comparable-2.19 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 241 | Qwen3 Coder 30B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows | 9.59 | Independent | reasoningoffversion4.3 | Partially comparable-2.28 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 241 | tri-21b-think-previewOpen weightsTrillion Labs · best of 2 rows | 9.59 | Independent | reasoningonversion4.3 | Partially comparable-2.28 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 243 | GPT-4.5 PreviewClosedOpenAI · GPT 4.5 · best of 2 rows | 9.58 | Independent | reasoningoffversion4.3 | Partially comparable-2.29 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 244 | diffusiongemma-26b-a4bOpen weightsGoogle · best of 2 rows | 9.51 | Independent | reasoningonversion4.3 | Partially comparable-2.36 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 245 | qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows | 9.50 | Independent | reasoningonversion4.3 | Partially comparable-2.37 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 246 | QwQ-32BOpen weightsAlibaba Group · Qwen · best of 2 rows | 9.47 | Independent | reasoningonversion4.3 | Partially comparable-2.40 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 247 | Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 4 rows | 9.45 | Independent | reasoningonversion4.3 | Partially comparable-2.42 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 248 | gemini-2.0-flash-thinking-exp-01-21ClosedGoogle · Gemini 2.0 · best of 2 rows | 9.42 | Independent | reasoningonversion4.3 | Partially comparable-2.45 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 249 | Devstral 2Open weightsMistral AI · Devstral 2 · best of 2 rows | 9.41 | Independent | reasoningoffversion4.3 | Partially comparable-2.46 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 250 | Llama 4 MaverickOpen weightsMeta AI · Llama 4 · best of 2 rows | 9.30 | Independent | reasoningoffversion4.3 | Partially comparable-2.57 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 251 | motif-2-12-7bClosedMotif Technologies · best of 2 rows | 9.19 | Independent | reasoningonversion4.3 | Partially comparable-2.68 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 252 | ling-1tOpen weightsinclusionAI · best of 2 rows | 9.17 | Independent | reasoningoffversion4.3 | Partially comparable-2.70 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | nova-premierClosedAmazon Web Services · Nova · best of 2 rows | 9.16 | Independent | reasoningoffversion4.3 | Partially comparable-2.71 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 254 | solar-pro-2ClosedUpstage · Solar · best of 5 rows | 9.07 | Independent | reasoningonversion4.3 | Partially comparable-2.80 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 254 | solar-pro-2-previewClosedUpstage · Solar · best of 3 rows | 9.07 | Independent | reasoningonversion4.3 | Partially comparable-2.80 | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 256 | granite-4.2-3bOpen weightsIBM · Granite 4.2 · best of 2 rows | 9.06 | Independent | reasoningonversion4.3 | Partially comparable-2.81 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 257 | Mistral Medium 3ClosedMistral AI · Mistral · best of 2 rows | 9.05 | Independent | reasoningoffversion4.3 | Partially comparable-2.82 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 257 | magistral-mediumClosedMistral AI · Magistral · best of 2 rows | 9.05 | Independent | reasoningonversion4.3 | Partially comparable-2.82 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 259 | devstral-mediumClosedMistral AI · Devstral · best of 2 rows | 9.01 | Independent | reasoningoffversion4.3 | Partially comparable-2.86 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 259 | llama-nemotron-super-49b-v1-5Open weightsNVIDIA · Llama · best of 4 rows | 9.01 | Independent | reasoningonversion4.3 | Partially comparable-2.86 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 261 | tri-21b-think-v0-5Open weightsTrillion Labs · best of 2 rows | 8.99 | Independent | reasoningonversion4.3 | Partially comparable-2.88 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | gpt-4o-chatgpt-03-25ClosedOpenAI · GPT 4 · best of 2 rows | 8.96 | Independent | reasoningoffversion4.3 | Partially comparable-2.91 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 263 | Claude Haiku 3.5ClosedAnthropic · Claude · best of 2 rows | 8.94 | Independent | reasoningoffversion4.3 | Partially comparable-2.93 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 263 | Gemini 2.0 FlashClosedGoogle · Gemini 2.0 · best of 2 rows | 8.94 | Independent | reasoningoffversion4.3 | Partially comparable-2.93 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 265 | llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.3 · best of 4 rows | 8.93 | Independent | reasoningonversion4.3 | Partially comparable-2.94 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 266 | gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 4 rows | 8.91 | Independent | reasoningonversion4.3 | Partially comparable-2.96 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 267 | Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 4 rows | 8.90 | Independent | reasoningonversion4.3 | Partially comparable-2.97 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 268 | minicpm5-1bOpen weightsOpenBMB · best of 4 rows | 8.80 | Independent | reasoningonversion4.3 | Partially comparable-3.07 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 268 | qwen3-4b-2507-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows | 8.80 | Independent | reasoningonversion4.3 | Partially comparable-3.07 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 270 | sarvam-105bOpen weightsSarvam · best of 2 rows | 8.79 | Independent | reasoning_efforthighversion4.3 | Partially comparable-3.08 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 271 | gemini-2-0-pro-experimental-02-05ClosedGoogle · Gemini 2.0 · best of 2 rows | 8.75 | Independent | reasoningoffversion4.3 | Partially comparable-3.12 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 272 | Devstral Small 1.0Open weightsMistral AI · Devstral · best of 2 rows | 8.74 | Independent | reasoningoffversion4.3 | Partially comparable-3.13 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 273 | claude-3-opusClosedAnthropic · Claude 3 · best of 2 rows | 8.72 | Independent | reasoningoffversion4.3 | Partially comparable-3.15 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 274 | SonarClosedPerplexity AI · Sonar · best of 3 rows | 8.67 | Independent | reasoningonversion4.3 | Partially comparable-3.20 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 275 | gemini-2-5-flash-04-2025ClosedGoogle · Gemini 2.5 | 8.66 | Independent | version4.3 | Partially comparable-3.21 | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | Magistral Small 1.2Open weightsMistral AI · Magistral · best of 2 rows | 8.59 | Independent | reasoningonversion4.3 | Partially comparable-3.28 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 277 | gpt-4oClosedOpenAI · GPT 4 · best of 2 rows | 8.44 | Independent | reasoningoffversion4.3 | Partially comparable-3.43 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 278 | nanbeige4-1-3bOpen weightsNanbeige · best of 2 rows | 8.40 | Independent | reasoningonversion4.3 | Partially comparable-3.47 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 279 | LFM2.5-2.6B (free)Open weightsLiquid AI · LFM2.5 · best of 2 rows | 8.39 | Independent | reasoningonversion4.3 | Partially comparable-3.48 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 280 | DeepSeek-R1-Distill-Qwen-32BOpen weightsDeepSeek · Qwen · best of 2 rows | 8.38 | Independent | reasoningonversion4.3 | Partially comparable-3.49 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 281 | Devstral Small 2Open weightsMistral AI · Devstral · best of 2 rows | 8.36 | Independent | reasoningoffversion4.3 | Partially comparable-3.51 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 282 | gemini-2.0-flash-expClosedGoogle · Gemini 2.0 · best of 2 rows | 8.22 | Independent | reasoningoffversion4.3 | Partially comparable-3.65 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 282 | magistral-smallOpen weightsMistral AI · Magistral · best of 2 rows | 8.22 | Independent | reasoningonversion4.3 | Partially comparable-3.65 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 284 | exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 4 rows | 8.18 | Independent | reasoningonversion4.3 | Partially comparable-3.69 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 285 | Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 4 rows | 8.17 | Independent | reasoningonversion4.3 | Partially comparable-3.70 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 286 | DeepSeek-R1-0528-Qwen3-8BOpen weightsDeepSeek · Qwen3 · best of 2 rows | 8.08 | Independent | reasoningonversion4.3 | Partially comparable-3.79 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 287 | qwen-2-5-maxClosedAlibaba Group · Qwen · best of 2 rows | 8.02 | Independent | reasoningoffversion4.3 | Partially comparable-3.85 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 288 | Hermes-4-70BRestricted weightsNous Research · Hermes 4 | 7.91 | Independent | version4.3 | Partially comparable-3.96 | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 288 | gemini-1-5-proClosedGoogle · Gemini 1.5 · best of 2 rows | 7.91 | Independent | reasoningoffversion4.3 | Partially comparable-3.96 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 288 | hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.1 · best of 3 rows | 7.91 | Independent | reasoningonversion4.3 | Partially comparable-3.96 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 291 | R1 Distill Llama 70BOpen weightsDeepSeek · Llama · best of 2 rows | 7.89 | Independent | reasoningonversion4.3 | Partially comparable-3.98 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 292 | claude-35-sonnetClosedAnthropic · Claude 35 · best of 2 rows | 7.88 | Independent | reasoningoffversion4.3 | Partially comparable-3.99 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 293 | deepseek-r1-distill-qwen-14bOpen weightsDeepSeek · Qwen · best of 2 rows | 7.85 | Independent | reasoningonversion4.3 | Partially comparable-4.02 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 294 | falcon-h1r-7bOpen weightsTII UAE · Falcon · best of 2 rows | 7.83 | Independent | reasoningonversion4.3 | Partially comparable-4.04 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 295 | Solar Pro 3ClosedUpstage · Solar · best of 2 rows | 7.82 | Independent | reasoningonversion4.3 | Partially comparable-4.05 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 295 | gpt-4.1-nanoClosedOpenAI · GPT 4.1 · best of 2 rows | 7.82 | Independent | reasoningoffversion4.3 | Partially comparable-4.05 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 297 | ling-flash-2-0Open weightsinclusionAI · best of 2 rows | 7.81 | Independent | reasoningoffversion4.3 | Partially comparable-4.06 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 298 | gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 4 rows | 7.77 | Independent | reasoningonversion4.3 | Partially comparable-4.10 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 299 | qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows | 7.76 | Independent | reasoningonversion4.3 | Partially comparable-4.11 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 300 | GPT-4o (2024-08-06)ClosedOpenAI · GPT 4 · best of 2 rows | 7.74 | Independent | reasoningoffversion4.3 | Partially comparable-4.13 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →