Artificial Analysis Intelligence Index
composite of several evaluations run by Artificial Analysis
Updated 5 h ago · first seen 11 Sept 2026
- Metric
- index ↑
- Current results
- 1,272
- Models
- 477
- Current leader
- Claude Fable 5.1 53.4
Score history · minicpm5-1b 4 rows
- minicpm5-1b
- 8.8aa_slug=minicpm5-1b · version=4.3 · estimated=true · reasoning=on12 Sept 2026
- 8.69aa_slug=minicpm5-1b-non-reasoning · version=4.3 · estimated=true · reasoning=off12 Sept 2026
- 8.8aa_slug=minicpm5-1b · version=4.3 · estimated=true11 Sept 2026
- 8.69aa_slug=minicpm5-1b-non-reasoning · version=4.3 · estimated=true · reasoning=off11 Sept 2026
Frontier over time · index
8 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 53.4Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 53.2Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 51.2Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 51.0gpt-6-astra OpenAI Independent11 Sept 2026
- 49.7gpt-6-astra OpenAI Independent11 Sept 2026
- 45.1Claude Opus 5 Anthropic Independent11 Sept 2026
- 14.6grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026
- 5.71llama-2-chat-7b Meta AI Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 477 models · trust independent-evaluator
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 301 | Qwen2.5 72B InstructOpen weightsQwen · Qwen2.5 · best of 2 rows | 7.73 | Independent | reasoningoffversion4.3 | Partially comparable0.00 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 302 | step-3-vl-10bOpen weightsStepFun · Step3 · best of 2 rows | 7.69 | Independent | reasoningonversion4.3 | Partially comparable-0.04 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 303 | Llama 3.3 70BOpen weightsMeta AI · Llama 3.3 · best of 2 rows | 7.66 | Independent | reasoningoffversion4.3 | Partially comparable-0.07 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 304 | qwen3-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows | 7.63 | Independent | reasoningonversion4.3 | Partially comparable-0.10 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 305 | Sonar ProClosedPerplexity AI · Sonar · best of 2 rows | 7.61 | Independent | reasoningoffversion4.3 | Partially comparable-0.12 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 306 | devstral-smallOpen weightsMistral AI · Devstral · best of 2 rows | 7.60 | Independent | reasoningoffversion4.3 | Partially comparable-0.13 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 307 | QwQ-32B-PreviewOpen weightsAlibaba Group · Qwen · best of 2 rows | 7.59 | Independent | reasoningonversion4.3 | Partially comparable-0.14 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 308 | GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 4 rows | 7.56 | Independent | reasoningonversion4.3 | Partially comparable-0.17 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 308 | mistral-large-2Open weightsMistral AI · Mistral · best of 2 rows | 7.56 | Independent | reasoningoffversion4.3 | Partially comparable-0.17 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 310 | llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.1 · best of 2 rows | 7.53 | Independent | reasoningonversion4.3 | Partially comparable-0.20 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 311 | ernie-4-5-300b-a47bOpen weightsBaidu · ERNIE 4.5 · best of 2 rows | 7.50 | Independent | reasoningoffversion4.3 | Partially comparable-0.23 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 312 | Hermes 4 405BOpen weightsNous Research · Hermes 4 | 7.49 | Independent | version4.3 | Partially comparable-0.24 | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 312 | hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.1 · best of 3 rows | 7.49 | Independent | reasoningonversion4.3 | Partially comparable-0.24 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 314 | nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 4 rows | 7.48 | Independent | reasoningonversion4.3 | Partially comparable-0.25 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 315 | granite-4.1-30bOpen weightsIBM · Granite 4.1 · best of 2 rows | 7.44 | Independent | reasoningoffversion4.3 | Partially comparable-0.29 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 316 | NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 4 rows | 7.43 | Independent | reasoningonversion4.3 | Partially comparable-0.30 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 317 | gemini-2-0-flash-lite-001ClosedGoogle · Gemini 2.0 · best of 2 rows | 7.41 | Independent | reasoningoffversion4.3 | Partially comparable-0.32 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 318 | nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 7.38 | Independent | reasoningonversion4.3 | Partially comparable-0.35 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 319 | Mistral Small 3.1Open weightsMistral AI · Mistral · best of 2 rows | 7.37 | Independent | reasoningoffversion4.3 | Partially comparable-0.36 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 320 | qwen3-32b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows | 7.34 | Independent | reasoningoffversion4.3 | Partially comparable-0.39 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 321 | GPT-4o (2024-05-13)ClosedOpenAI · GPT 4 · best of 2 rows | 7.33 | Independent | reasoningoffversion4.3 | Partially comparable-0.40 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 321 | gemini-2-0-flash-lite-previewClosedGoogle · Gemini 2.0 · best of 2 rows | 7.33 | Independent | reasoningoffversion4.3 | Partially comparable-0.40 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 323 | llama-3-1-nemotron-nano-4b-reasoningOpen weightsNVIDIA · Llama 3.1 · best of 2 rows | 7.31 | Independent | reasoningonversion4.3 | Partially comparable-0.42 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 324 | Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AI · Kimi · best of 2 rows | 7.30 | Independent | reasoningoffversion4.3 | Partially comparable-0.43 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 325 | Llama-3.1-405BRestricted weightsMeta AI · Llama 3.1 · best of 2 rows | 7.29 | Independent | reasoningoffversion4.3 | Partially comparable-0.44 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 326 | Qwen3-4BOpen weightsQwen · Qwen3 | 7.23 | Independent | version4.3 | Partially comparable-0.50 | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 326 | qwen3-4b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows | 7.23 | Independent | reasoningonversion4.3 | Partially comparable-0.50 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 328 | LFM2.5-8B-A1BOpen weightsLiquid AI · LFM2.5 · best of 2 rows | 7.22 | Independent | reasoningonversion4.3 | Partially comparable-0.51 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 328 | Qwen3 32BOpen weightsQwen · Qwen3 | 7.22 | Independent | version4.3 | Partially comparable-0.51 | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 330 | claude-35-sonnet-june-24ClosedAnthropic · Claude 35 · best of 2 rows | 7.21 | Independent | reasoningoffversion4.3 | Partially comparable-0.52 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 331 | tulu3-405bOpen weightsAllen Institute for AI · best of 2 rows | 7.20 | Independent | reasoningoffversion4.3 | Partially comparable-0.53 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 332 | gpt-4o-chatgptClosedOpenAI · GPT 4 · best of 2 rows | 7.19 | Independent | reasoningoffversion4.3 | Partially comparable-0.54 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 333 | Pixtral LargeOpen weightsMistral AI · Pixtral · best of 2 rows | 7.15 | Independent | reasoningoffversion4.3 | Partially comparable-0.58 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 333 | ring-flash-2-0Open weightsinclusionAI · best of 2 rows | 7.15 | Independent | reasoningonversion4.3 | Partially comparable-0.58 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 335 | olmo-3-1-32b-instructOpen weightsAllen Institute for AI · OLMo 3.1 · best of 3 rows | 7.12 | Independent | reasoningonversion4.3 | Partially comparable-0.61 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 336 | grok-2Open weightsxAI · Grok 2 · best of 2 rows | 7.11 | Independent | reasoningoffversion4.3 | Partially comparable-0.62 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 337 | gemini-1-5-flashClosedGoogle · Gemini 1.5 · best of 2 rows | 7.07 | Independent | reasoningoffversion4.3 | Partially comparable-0.66 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 338 | Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 4 rows | 7.05 | Independent | reasoningonversion4.3 | Partially comparable-0.68 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 339 | gpt-4-turboClosedOpenAI · GPT 4 · best of 2 rows | 7.04 | Independent | reasoningoffversion4.3 | Partially comparable-0.69 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 340 | Command AOpen weightsCohere · Command · best of 2 rows | 6.96 | Independent | reasoningoffversion4.3 | Partially comparable-0.77 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 340 | Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows | 6.96 | Independent | reasoningoffversion4.3 | Partially comparable-0.77 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 340 | nova-proClosedAmazon Web Services · Nova · best of 2 rows | 6.96 | Independent | reasoningoffversion4.3 | Partially comparable-0.77 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 343 | Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 4 rows | 6.94 | Independent | reasoningonversion4.3 | Partially comparable-0.79 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 343 | llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.1 · best of 2 rows | 6.94 | Independent | reasoningoffversion4.3 | Partially comparable-0.79 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 345 | Llama 3.1 8BRestricted weightsMeta AI · Llama 3.1 · best of 2 rows | 6.93 | Independent | reasoningoffversion4.3 | Partially comparable-0.80 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 346 | grok-betaClosedSpaceXAI · Grok · best of 2 rows | 6.89 | Independent | reasoningoffversion4.3 | Partially comparable-0.84 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 347 | qwen2.5-32b-instructOpen weightsAlibaba Group · Qwen2.5 · best of 2 rows | 6.87 | Independent | reasoningoffversion4.3 | Partially comparable-0.86 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 348 | Mistral Large 2.0Open weightsMistral AI · Mistral · best of 2 rows | 6.80 | Independent | reasoningoffversion4.3 | Partially comparable-0.93 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 349 | Qwen2.5 Coder 32B InstructOpen weightsQwen · Qwen2.5 · best of 2 rows | 6.74 | Independent | reasoningoffversion4.3 | Partially comparable-0.99 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 350 | gpt-4ClosedOpenAI · GPT 4 · best of 2 rows | 6.70 | Independent | reasoningoffversion4.3 | Partially comparable-1.03 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 350 | qwen3-14b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows | 6.70 | Independent | reasoningoffversion4.3 | Partially comparable-1.03 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 352 | Mistral Small 3Open weightsMistral AI · Mistral · best of 2 rows | 6.67 | Independent | reasoningoffversion4.3 | Partially comparable-1.06 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 352 | nova-liteClosedAmazon Web Services · Nova · best of 2 rows | 6.67 | Independent | reasoningoffversion4.3 | Partially comparable-1.06 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 354 | gpt-4o-miniClosedOpenAI · GPT 4 · best of 2 rows | 6.66 | Independent | reasoningoffversion4.3 | Partially comparable-1.07 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 355 | deepseek-v2-5Open weightsDeepSeek · DeepSeek · best of 2 rows | 6.62 | Independent | reasoningoffversion4.3 | Partially comparable-1.11 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 356 | Llama-3.1-70BOpen weightsMeta AI · Llama 3.1 · best of 2 rows | 6.60 | Independent | reasoningoffversion4.3 | Partially comparable-1.13 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 357 | granite-4.1-8bOpen weightsIBM · Granite 4.1 · best of 2 rows | 6.57 | Independent | reasoningoffversion4.3 | Partially comparable-1.16 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 358 | gemini-2-0-flash-thinking-exp-1219ClosedGoogle · Gemini 2.0 · best of 2 rows | 6.56 | Independent | reasoningonversion4.3 | Partially comparable-1.17 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 358 | sarvam-30bOpen weightsSarvam · best of 2 rows | 6.56 | Independent | reasoning_efforthighversion4.3 | Partially comparable-1.17 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 360 | deepseek-v2-5-sep-2024Open weightsDeepSeek · DeepSeek · best of 2 rows | 6.55 | Independent | reasoningoffversion4.3 | Partially comparable-1.18 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 361 | Mistral SabaClosedMistral AI · Mistral · best of 2 rows | 6.49 | Independent | reasoningoffversion4.3 | Partially comparable-1.24 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 362 | deepseek-r1-distill-llama-8bOpen weightsDeepSeek · Llama · best of 2 rows | 6.48 | Independent | reasoningonversion4.3 | Partially comparable-1.25 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 363 | olmo-3-32b-thinkOpen weightsAllen Institute for AI · OLMo 3 · best of 2 rows | 6.47 | Independent | reasoningonversion4.3 | Partially comparable-1.26 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 364 | Llama 4 ScoutOpen weightsMeta AI · Llama 4 · best of 2 rows | 6.45 | Independent | reasoningoffversion4.3 | Partially comparable-1.28 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 365 | gemini-1-5-pro-may-2024ClosedGoogle · Gemini 1.5 · best of 2 rows | 6.44 | Independent | reasoningoffversion4.3 | Partially comparable-1.29 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 365 | r1-1776Open weightsPerplexity AI · best of 2 rows | 6.44 | Independent | reasoningonversion4.3 | Partially comparable-1.29 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 367 | qwen-turboClosedAlibaba Group · Qwen · best of 2 rows | 6.43 | Independent | reasoningoffversion4.3 | Partially comparable-1.30 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 367 | reka-flashClosedrekaai · best of 2 rows | 6.43 | Independent | reasoningoffversion4.3 | Partially comparable-1.30 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 369 | llama-3-2-instruct-90b-visionOpen weightsMeta AI · Llama 3.2 · best of 2 rows | 6.41 | Independent | reasoningoffversion4.3 | Partially comparable-1.32 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 369 | solar-miniOpen weightsUpstage · Solar · best of 2 rows | 6.41 | Independent | reasoningoffversion4.3 | Partially comparable-1.32 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 371 | Qwen3 14BOpen weightsQwen · Qwen3 | 6.39 | Independent | version4.3 | Partially comparable-1.34 | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 372 | celeris-1ClosedCeleris · best of 2 rows | 6.35 | Independent | reasoningoffversion4.3 | Partially comparable-1.38 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 373 | grok-1Open weightsxAI · Grok 1 · best of 2 rows | 6.34 | Independent | reasoningoffversion4.3 | Partially comparable-1.39 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 374 | phi-4-miniOpen weightsMicrosoft · Phi4 · best of 2 rows | 6.33 | Independent | reasoningoffversion4.3 | Partially comparable-1.40 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 374 | qwen2-72b-instructOpen weightsAlibaba Group · Qwen2 · best of 2 rows | 6.33 | Independent | reasoningoffversion4.3 | Partially comparable-1.40 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 376 | gemini-1-5-flash-8bClosedGoogle · Gemini 1.5 · best of 2 rows | 6.15 | Independent | reasoningoffversion4.3 | Partially comparable-1.58 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 377 | Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 4 rows | 6.13 | Independent | reasoningonversion4.3 | Partially comparable-1.60 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 378 | deephermes-3-mistral-24b-previewOpen weightsNous Research · Mistral · best of 2 rows | 6.07 | Independent | reasoningoffversion4.3 | Partially comparable-1.66 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 378 | jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.7 · best of 2 rows | 6.07 | Independent | reasoningoffversion4.3 | Partially comparable-1.66 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 380 | granite-4-0-h-smallOpen weightsIBM · Granite 4.0 · best of 2 rows | 6.05 | Independent | reasoningoffversion4.3 | Partially comparable-1.68 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 381 | Ministral 3 14BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 6.04 | Independent | reasoningoffversion4.3 | Partially comparable-1.69 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 382 | jamba-1-5-largeOpen weightsAI21 Labs · Jamba 1.5 · best of 2 rows | 6.01 | Independent | reasoningoffversion4.3 | Partially comparable-1.72 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 383 | Hermes 3 70B InstructOpen weightsNous Research · Hermes 3 · best of 2 rows | 6 | Independent | reasoningoffversion4.3 | Partially comparable-1.73 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 384 | qwen3-8b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows | 5.99 | Independent | reasoningoffversion4.3 | Partially comparable-1.74 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 385 | deepseek-coder-v2Open weightsDeepSeek · DeepSeek · best of 2 rows | 5.98 | Independent | reasoningoffversion4.3 | Partially comparable-1.75 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 386 | jamba-1-6-largeOpen weightsAI21 Labs · Jamba 1.6 · best of 2 rows | 5.97 | Independent | reasoningoffversion4.3 | Partially comparable-1.76 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 386 | olmo-2-32bOpen weightsAllen Institute for AI · OLMo 2 · best of 2 rows | 5.97 | Independent | reasoningoffversion4.3 | Partially comparable-1.76 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 388 | lfm2-24b-a2bOpen weightsLiquid AI · LFM2 · best of 2 rows | 5.95 | Independent | reasoningoffversion4.3 | Partially comparable-1.78 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 389 | gemini-1-5-flash-may-2024ClosedGoogle · Gemini 1.5 · best of 2 rows | 5.94 | Independent | reasoningoffversion4.3 | Partially comparable-1.79 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 390 | Phi 4Open weightsMicrosoft · Phi4 · best of 2 rows | 5.92 | Independent | reasoningoffversion4.3 | Partially comparable-1.81 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 391 | claude-3-sonnetClosedAnthropic · Claude 3 · best of 2 rows | 5.88 | Independent | reasoningoffversion4.3 | Partially comparable-1.85 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 391 | nova-microClosedAmazon Web Services · Nova · best of 2 rows | 5.88 | Independent | reasoningoffversion4.3 | Partially comparable-1.85 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 393 | granite-4.1-3bOpen weightsIBM · Granite 4.1 · best of 2 rows | 5.86 | Independent | reasoningoffversion4.3 | Partially comparable-1.87 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 394 | mistral-smallOpen weightsMistral AI · Mistral · best of 2 rows | 5.85 | Independent | reasoningoffversion4.3 | Partially comparable-1.88 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 395 | gemini-1-0-ultraClosedGoogle · Gemini 1.0 · best of 2 rows | 5.84 | Independent | reasoningoffversion4.3 | Partially comparable-1.89 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 396 | phi-3-miniOpen weightsMicrosoft · Phi3 · best of 2 rows | 5.82 | Independent | reasoningoffversion4.3 | Partially comparable-1.91 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 397 | gemma-3n-e4b-preview-0520Open weightsGoogle · Gemma 3 · best of 2 rows | 5.81 | Independent | reasoningoffversion4.3 | Partially comparable-1.92 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 397 | phi-4-multimodalOpen weightsMicrosoft · Phi4 · best of 2 rows | 5.81 | Independent | reasoningoffversion4.3 | Partially comparable-1.92 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 399 | Qwen2.5-Coder-7BOpen weightsQwen · Qwen2.5 · best of 2 rows | 5.79 | Independent | reasoningoffversion4.3 | Partially comparable-1.94 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 400 | Mistral LargeClosedMistral AI · Mistral · best of 2 rows | 5.76 | Independent | reasoningoffversion4.3 | Partially comparable-1.97 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →