Updated 5 h ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 1,217
- Models
- 453
- Current leader
- Claude Fable 5.1 59.1%
Frontier over time · accuracy · evaluator=Artificial Analysis
9 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 59.1%Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 58.7%Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 55.9%Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 55.5%Claude Fable 5 Anthropic Independent11 Sept 2026
- 53.1%gpt-6-astra OpenAI Independent11 Sept 2026
- 52.7%gpt-6-astra OpenAI Independent11 Sept 2026
- 51.3%Claude Opus 5 Anthropic Independent11 Sept 2026
- 11.0%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 453 models
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 301 | llama-2-chat-70bOpen weightsMeta AI · Llama 2 · best of 2 rows | 5.21% | Independent | reasoningoff | Partially comparable0.00 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 302 | llama-3-1-nemotron-nano-4b-reasoningOpen weightsNVIDIA · Llama 3.1 · best of 2 rows | 5.19% | Independent | reasoningon | Partially comparable-0.02 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 303 | jamba-1-5-miniOpen weightsAI21 Labs · Jamba 1.5 · best of 2 rows | 5.14% | Independent | reasoningoff | Partially comparable-0.07 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 303 | molmo-7b-dOpen weightsAllen Institute for AI · Molmo · best of 2 rows | 5.14% | Independent | reasoningoff | Partially comparable-0.07 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 305 | R1 Distill Llama 70BOpen weightsDeepSeek · Llama · best of 2 rows | 5.13% | Independent | reasoningon | Partially comparable-0.08 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 306 | LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.5 · best of 2 rows | 5.10% | Independent | reasoningoff | Partially comparable-0.11 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 306 | Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 4 rows | 5.10% | Independent | reasoningoff | Partially comparable-0.11 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 308 | llama-3-instruct-8bOpen weightsMeta AI · Llama 3 · best of 2 rows | 5.06% | Independent | reasoningoff | Partially comparable-0.15 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 309 | ling-mini-2-0Open weightsinclusionAI · best of 2 rows | 5.05% | Independent | reasoningoff | Partially comparable-0.16 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 309 | minicpm-v4-6-1-3bOpen weightsOpenBMB · best of 2 rows | 5.05% | Independent | reasoningoff | Partially comparable-0.16 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 311 | gpt-4.1-miniClosedOpenAI · GPT 4.1 · best of 2 rows | 5.02% | Independent | reasoningoff | Partially comparable-0.19 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 312 | Granite 4.0 MicroOpen weightsIBM · Granite 4.0 · best of 2 rows | 5% | Independent | reasoningoff | Partially comparable-0.21 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 312 | granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.0 · best of 2 rows | 5% | Independent | reasoningoff | Partially comparable-0.21 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 312 | phi-4-multimodalOpen weightsMicrosoft · Phi4 · best of 2 rows | 5% | Independent | reasoningoff | Partially comparable-0.21 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 315 | Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 4 rows | 4.96% | Independent | reasoningoff | Partially comparable-0.25 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 315 | apertus-8b-instructOpen weightsSwiss AI Initiative · best of 2 rows | 4.96% | Independent | reasoningoff | Partially comparable-0.25 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 315 | phi-3-miniOpen weightsMicrosoft · Phi3 · best of 2 rows | 4.96% | Independent | reasoningoff | Partially comparable-0.25 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 318 | lfm-40bClosedLiquid AI · LFM · best of 2 rows | 4.93% | Independent | reasoningoff | Partially comparable-0.28 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 319 | Llama 4 MaverickOpen weightsMeta AI · Llama 4 · best of 2 rows | 4.91% | Independent | reasoningoff | Partially comparable-0.30 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 319 | lfm2-8b-a1bOpen weightsLiquid AI · LFM2 · best of 2 rows | 4.91% | Independent | reasoningoff | Partially comparable-0.30 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 321 | Qwen2.5-Coder-7BOpen weightsQwen · Qwen2.5 · best of 2 rows | 4.89% | Independent | reasoningoff | Partially comparable-0.32 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 322 | SonarClosedPerplexity AI · Sonar · best of 2 rows | 4.88% | Independent | reasoningoff | Partially comparable-0.33 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 323 | NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 4 rows | 4.87% | Independent | reasoningon | Partially comparable-0.34 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 323 | nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 4.87% | Independent | reasoningon | Partially comparable-0.34 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 325 | gemma-3n-e4b-preview-0520Open weightsGoogle · Gemma 3 · best of 2 rows | 4.82% | Independent | reasoningoff | Partially comparable-0.39 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 325 | gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 4 rows | 4.82% | Independent | reasoningoff | Partially comparable-0.39 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 325 | granite-4-0-nano-1bOpen weightsIBM · Granite 4.0 · best of 2 rows | 4.82% | Independent | reasoningoff | Partially comparable-0.39 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 325 | nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 4.82% | Independent | reasoningon | Partially comparable-0.39 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 329 | command-r-03-2024Open weightsCohere · Command · best of 2 rows | 4.79% | Independent | reasoningoff | Partially comparable-0.42 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 329 | openchat-35Open weightsOpenChat · best of 2 rows | 4.79% | Independent | reasoningoff | Partially comparable-0.42 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 331 | gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 4 rows | 4.77% | Independent | reasoningon | Partially comparable-0.44 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 331 | llama-2-chat-13bOpen weightsMeta AI · Llama 2 · best of 2 rows | 4.77% | Independent | reasoningoff | Partially comparable-0.44 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 333 | DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows | 4.74% | Independent | reasoningoff | Partially comparable-0.47 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 334 | gemini-1-5-flash-8bClosedGoogle · Gemini 1.5 · best of 2 rows | 4.69% | Independent | reasoningoff | Partially comparable-0.52 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 335 | Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows | 4.68% | Independent | reasoningoff | Partially comparable-0.53 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 335 | Mixtral 8x7BOpen weightsMistral AI · Mixtral 8 · best of 2 rows | 4.68% | Independent | reasoningoff | Partially comparable-0.53 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 337 | gemini-1-5-proClosedGoogle · Gemini 1.5 · best of 2 rows | 4.64% | Independent | reasoningoff | Partially comparable-0.57 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 338 | Ministral 3 14BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 4.63% | Independent | reasoningoff | Partially comparable-0.58 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 338 | command-r-plus-04-2024Open weightsCohere · Command · best of 2 rows | 4.63% | Independent | reasoningoff | Partially comparable-0.58 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 338 | nova-microClosedAmazon Web Services · Nova · best of 2 rows | 4.63% | Independent | reasoningoff | Partially comparable-0.58 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 341 | Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 4 rows | 4.59% | Independent | reasoningon | Partially comparable-0.62 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 342 | DeepSeek-R1-Distill-Qwen-32BOpen weightsDeepSeek · Qwen · best of 2 rows | 4.58% | Independent | reasoningon | Partially comparable-0.63 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 343 | Mistral 7BOpen weightsMistral AI · Mistral · best of 2 rows | 4.55% | Independent | reasoningoff | Partially comparable-0.66 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 344 | qwen3-coder-480b-a35b-instructOpen weightsAlibaba Group · Qwen3-Coder · best of 2 rows | 4.54% | Independent | reasoningoff | Partially comparable-0.67 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 345 | Qwen3 14BOpen weightsQwen · Qwen3 | 4.53% | Independent | group defaults | Partially comparable-0.68 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 345 | qwen3-14b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows | 4.53% | Independent | reasoningon | Partially comparable-0.68 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 347 | llama-3-instruct-70bOpen weightsMeta AI · Llama 3 · best of 2 rows | 4.52% | Independent | reasoningoff | Partially comparable-0.69 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 348 | g9v3-3bOpen weightsAI9Stars · best of 2 rows | 4.49% | Independent | reasoningon | Partially comparable-0.72 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 348 | gemma-3n-e4bOpen weightsGoogle · Gemma 3 · best of 2 rows | 4.49% | Independent | reasoningoff | Partially comparable-0.72 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 348 | llama-3-2-instruct-90b-visionOpen weightsMeta AI · Llama 3.2 · best of 2 rows | 4.49% | Independent | reasoningoff | Partially comparable-0.72 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 351 | Llama-3.1-70BOpen weightsMeta AI · Llama 3.1 · best of 2 rows | 4.47% | Independent | reasoningoff | Partially comparable-0.74 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 352 | grok-betaClosedSpaceXAI · Grok · best of 2 rows | 4.45% | Independent | reasoningoff | Partially comparable-0.76 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 352 | jamba-1-7-miniOpen weightsAI21 Labs · Jamba 1.7 · best of 2 rows | 4.45% | Independent | reasoningoff | Partially comparable-0.76 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 352 | phi-4-miniOpen weightsMicrosoft · Phi4 · best of 2 rows | 4.45% | Independent | reasoningoff | Partially comparable-0.76 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 355 | Reka Flash 3Open weightsrekaai · best of 2 rows | 4.44% | Independent | reasoningon | Partially comparable-0.77 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 356 | Qwen3-4BOpen weightsQwen · Qwen3 | 4.42% | Independent | group defaults | Partially comparable-0.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 356 | qwen3-4b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows | 4.42% | Independent | reasoningon | Partially comparable-0.79 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 358 | Gemma 3 27BOpen weightsGoogle · Gemma 3 · best of 2 rows | 4.40% | Independent | reasoningoff | Partially comparable-0.81 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 359 | gemini-1-5-flash-may-2024ClosedGoogle · Gemini 1.5 · best of 2 rows | 4.36% | Independent | reasoningoff | Partially comparable-0.85 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 360 | Mistral Small 3.1Open weightsMistral AI · Mistral · best of 2 rows | 4.31% | Independent | reasoningoff | Partially comparable-0.90 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 360 | Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows | 4.31% | Independent | reasoningoff | Partially comparable-0.90 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 360 | jamba-1-6-miniOpen weightsAI21 Labs · Jamba 1.6 · best of 2 rows | 4.31% | Independent | reasoningoff | Partially comparable-0.90 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 363 | Mistral SabaClosedMistral AI · Mistral · best of 2 rows | 4.30% | Independent | reasoningoff | Partially comparable-0.91 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 364 | nova-liteClosedAmazon Web Services · Nova · best of 2 rows | 4.29% | Independent | reasoningoff | Partially comparable-0.92 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 365 | Gemini 2.0 FlashClosedGoogle · Gemini 2.0 · best of 2 rows | 4.26% | Independent | reasoningoff | Partially comparable-0.95 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 365 | Ministral 3 8BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 4.26% | Independent | reasoningoff | Partially comparable-0.95 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 365 | Molmo2-8BOpen weightsAllen Institute for AI · best of 2 rows | 4.26% | Independent | reasoningoff | Partially comparable-0.95 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 365 | deephermes-3-llama-3-1-8b-previewOpen weightsNous Research · Llama 3.1 · best of 2 rows | 4.26% | Independent | reasoningoff | Partially comparable-0.95 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 369 | mistral-smallOpen weightsMistral AI · Mistral · best of 2 rows | 4.25% | Independent | reasoningoff | Partially comparable-0.96 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 370 | lfm2-24b-a2bOpen weightsLiquid AI · LFM2 · best of 2 rows | 4.22% | Independent | reasoningoff | Partially comparable-0.99 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 371 | Gemma 3 12BOpen weightsGoogle · Gemma 3 · best of 2 rows | 4.21% | Independent | reasoningoff | Partially comparable-1.00 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 372 | gpt-4o-miniClosedOpenAI · GPT 4 · best of 2 rows | 4.20% | Independent | reasoningoff | Partially comparable-1.01 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 373 | gemini-1-0-proClosedGoogle · Gemini 1.0 · best of 2 rows | 4.19% | Independent | reasoningoff | Partially comparable-1.02 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 373 | llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.1 · best of 2 rows | 4.19% | Independent | reasoningoff | Partially comparable-1.02 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 375 | gpt-4.1ClosedOpenAI · GPT 4.1 · best of 2 rows | 4.18% | Independent | reasoningoff | Partially comparable-1.03 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 376 | Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows | 4.17% | Independent | reasoningoff | Partially comparable-1.04 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 376 | gemma-3n-e2bOpen weightsGoogle · Gemma 3 · best of 2 rows | 4.17% | Independent | reasoningoff | Partially comparable-1.04 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 376 | granite-3-3-8b-instructOpen weightsIBM · Granite 3.3 · best of 2 rows | 4.17% | Independent | reasoningoff | Partially comparable-1.04 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 376 | nova-premierClosedAmazon Web Services · Nova · best of 2 rows | 4.17% | Independent | reasoningoff | Partially comparable-1.04 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 380 | deepseek-r1-distill-qwen-14bOpen weightsDeepSeek · Qwen · best of 2 rows | 4.13% | Independent | reasoningon | Partially comparable-1.08 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 381 | granite-4.1-30bOpen weightsIBM · Granite 4.1 · best of 2 rows | 4.12% | Independent | reasoningoff | Partially comparable-1.09 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 381 | grok-3ClosedSpaceXAI · Grok 3 · best of 2 rows | 4.12% | Independent | reasoningoff | Partially comparable-1.09 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 383 | Mistral Small 1.0ClosedMistral AI · Mistral · best of 2 rows | 4.09% | Independent | reasoningoff | Partially comparable-1.12 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 383 | gemini-2-0-flash-lite-previewClosedGoogle · Gemini 2.0 · best of 2 rows | 4.09% | Independent | reasoningoff | Partially comparable-1.12 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 383 | qwen-turboClosedAlibaba Group · Qwen · best of 2 rows | 4.09% | Independent | reasoningoff | Partially comparable-1.12 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 386 | Claude 3 HaikuClosedAnthropic · Claude · best of 2 rows | 4.08% | Independent | reasoningoff | Partially comparable-1.13 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 387 | gemini-2.0-flash-expClosedGoogle · Gemini 2.0 · best of 2 rows | 4.07% | Independent | reasoningoff | Partially comparable-1.14 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 387 | jamba-1-5-largeOpen weightsAI21 Labs · Jamba 1.5 · best of 2 rows | 4.07% | Independent | reasoningoff | Partially comparable-1.14 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 389 | Mistral Medium 3ClosedMistral AI · Mistral · best of 2 rows | 4.05% | Independent | reasoningoff | Partially comparable-1.16 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 390 | Command AOpen weightsCohere · Command · best of 2 rows | 4.03% | Independent | reasoningoff | Partially comparable-1.18 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 390 | Devstral Small 1.0Open weightsMistral AI · Devstral · best of 2 rows | 4.03% | Independent | reasoningoff | Partially comparable-1.18 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 392 | Mixtral 8x22BOpen weightsMistral AI · Mixtral 8 · best of 2 rows | 4.01% | Independent | reasoningoff | Partially comparable-1.20 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 393 | Hermes 3 70B InstructOpen weightsNous Research · Hermes 3 · best of 2 rows | 3.99% | Independent | reasoningoff | Partially comparable-1.22 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 393 | qwen2.5-32b-instructOpen weightsAlibaba Group · Qwen2.5 · best of 2 rows | 3.99% | Independent | reasoningoff | Partially comparable-1.22 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 395 | Llama-3.1-405BRestricted weightsMeta AI · Llama 3.1 · best of 2 rows | 3.98% | Independent | reasoningoff | Partially comparable-1.23 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 396 | gpt-4o-chatgpt-03-25ClosedOpenAI · GPT 4 · best of 2 rows | 3.97% | Independent | reasoningoff | Partially comparable-1.24 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 397 | QwQ-32B-PreviewOpen weightsAlibaba Group · Qwen · best of 2 rows | 3.91% | Independent | reasoningon | Partially comparable-1.30 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 398 | Qwen3 8BOpen weightsQwen · Qwen3 | 3.89% | Independent | group defaults | Partially comparable-1.32 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 398 | claude-21ClosedAnthropic · Claude 21 · best of 2 rows | 3.89% | Independent | reasoningoff | Partially comparable-1.32 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 398 | qwen3-8b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows | 3.89% | Independent | reasoningon | Partially comparable-1.32 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →