Artificial Analysis Intelligence Index
composite of several evaluations run by Artificial Analysis
Updated 51 min ago · first seen 11 Sept 2026
- Metric
- index ↑
- Current results
- 1,272
- Models
- 477
- Current leader
- Claude Fable 5.1 53.4
Score history · nova-2-0-pro 6 rows
- nova-2-0-pro
- 14.16aa_slug=nova-2-0-pro-reasoning-medium · version=4.3 · estimated=true · reasoning=on12 Sept 2026
- 9.96aa_slug=nova-2-0-pro · version=4.3 · estimated=true · reasoning=off12 Sept 2026
- 12.78aa_slug=nova-2-0-pro-reasoning-low · version=4.3 · estimated=true · reasoning=on12 Sept 2026
- 14.16aa_slug=nova-2-0-pro-reasoning-medium · version=4.3 · estimated=true · reasoning=on11 Sept 2026
- 9.96aa_slug=nova-2-0-pro · version=4.3 · estimated=true11 Sept 2026
- 12.78aa_slug=nova-2-0-pro-reasoning-low · version=4.3 · estimated=true · reasoning=on11 Sept 2026
Frontier over time · index
8 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 53.4Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 53.2Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 51.2Claude Fable 5.1 Anthropic Independent11 Sept 2026
- 51.0gpt-6-astra OpenAI Independent11 Sept 2026
- 49.7gpt-6-astra OpenAI Independent11 Sept 2026
- 45.1Claude Opus 5 Anthropic Independent11 Sept 2026
- 14.6grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026
- 5.71llama-2-chat-7b Meta AI Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 477 models · trust independent-evaluator
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1ClosedAnthropic · Claude · best of 10 rows | 53.4 | Independent | reasoning_effortmaxversion4.3 | leader | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 2 | gpt-6-astraClosedOpenAI · GPT 6 · best of 11 rows | 52.8 | Independent | reasoning_effortmaxversion4.3 | Comparable-0.56 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 3 | Claude Opus 5ClosedAnthropic · Claude · best of 10 rows | 50.7 | Independent | reasoning_effortmaxversion4.3 | Comparable-2.67 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 4 | Claude Fable 5ClosedAnthropic · Claude · best of 2 rows | 49.7 | Independent | reasoningonversion4.3 | Partially comparable-3.67 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 5 | Muse Spark 1.3ClosedMeta AI · best of 4 rows | 48.2 | Independent | reasoning_effortmaxversion4.3 | Comparable-5.20 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 6 | gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 12 rows | 47.1 | Independent | reasoning_effortmaxversion4.3 | Comparable-6.31 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 7 | GLM 5.3Open weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows | 44.9 | Independent | reasoning_effortmaxversion4.3 | Comparable-8.51 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 8 | Grok 4.6ClosedxAI · Grok · best of 8 rows | 44.4 | Independent | reasoning_efforthighversion4.3 | Partially comparable-8.96 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 9 | Kimi K3Open weightsMoonshot AI · Kimi · best of 4 rows | 43.8 | Independent | reasoning_effortmaxversion4.3 | Comparable-9.59 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 10 | gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 12 rows | 42.3 | Independent | reasoning_effortmaxversion4.3 | Comparable-11.1 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 11 | Claude Opus 4.8ClosedAnthropic · Claude · best of 2 rows | 42.0 | Independent | reasoning_effortmaxversion4.3 | Comparable-11.4 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 12 | GLM 5.3 FlashOpen weightsZ.ai (Zhipu AI) · GLM5.3 · best of 2 rows | 41.9 | Independent | reasoningonversion4.3 | Partially comparable-11.5 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 13 | Gemini 3.8 FlashClosedGoogle · Gemini 3.8 · best of 6 rows | 41.2 | Independent | reasoning_efforthighversion4.3 | Partially comparable-12.2 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 14 | Claude Opus 4.7ClosedAnthropic · Claude · best of 4 rows | 40.7 | Independent | reasoning_effortmaxversion4.3 | Comparable-12.7 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 15 | Qwen 3.8 MaxClosedQwen · Qwen3.8 · best of 2 rows | 40.3 | Independent | reasoningonversion4.3 | Partially comparable-13.1 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 16 | Qwen3.8 2.4T A95BOpen weightsQwen · Qwen3.8 · best of 2 rows | 40.0 | Independent | reasoningonversion4.3 | Partially comparable-13.3 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 17 | Qwen3.8 FlashOpen weightsQwen · Qwen3.8 · best of 2 rows | 39.9 | Independent | reasoningonversion4.3 | Partially comparable-13.5 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 18 | Muse Spark 1.2ClosedMeta AI · best of 2 rows | 39.8 | Independent | reasoning_effortxhighversion4.3 | Partially comparable-13.6 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 19 | Gemini 3.7 FlashClosedGoogle · Gemini 3.7 · best of 6 rows | 39.6 | Independent | reasoning_effortmediumversion4.3 | Partially comparable-13.8 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 20 | DeepSeek-V4.1-FlashOpen weightsDeepSeek · DeepSeek · best of 2 rows | 39.5 | Independent | reasoning_effortmaxversion4.3 | Comparable-13.8 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 21 | Grok 4.5ClosedxAI · Grok · best of 2 rows | 39.1 | Independent | reasoning_efforthighversion4.3 | Partially comparable-14.3 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 22 | gpt-5.4ClosedOpenAI · GPT 5.4 · best of 6 rows | 39.0 | Independent | reasoning_effortxhighversion4.3 | Partially comparable-14.4 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 23 | Z.ai GLM 5.2Open weightsZ.ai (Zhipu AI) · GLM5.2 · best of 4 rows | 38.6 | Independent | reasoning_effortmaxversion4.3 | Comparable-14.7 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 24 | gpt-5.5ClosedOpenAI · GPT 5.5 · best of 10 rows | 38.6 | Independent | reasoning_effortxhighversion4.3 | Partially comparable-14.7 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 25 | Claude Sonnet 5ClosedAnthropic · Claude · best of 12 rows | 38.4 | Independent | reasoning_effortmaxversion4.3 | Comparable-15.0 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 26 | gpt-5.6-lunaClosedOpenAI · GPT 5.6 · best of 12 rows | 37.5 | Independent | reasoning_effortmaxversion4.3 | Comparable-15.9 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 27 | deepseek-v4-proClosedDeepSeek · V4 · best of 2 rows | 36.3 | Independent | reasoning_effortmaxversion4.3 | Comparable-17.1 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 28 | agnes-3-0-flashClosedSapiens AI · best of 2 rows | 35.5 | Independent | reasoningonversion4.3 | Partially comparable-17.9 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 29 | agnes-2-5-pro-betaClosedSapiens AI · best of 2 rows | 35.2 | Independent | reasoningonversion4.3 | Partially comparable-18.1 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 30 | deepseek-v4-flash-visionClosedDeepSeek · DeepSeek · best of 2 rows | 35.0 | Independent | reasoning_effortmaxversion4.3 | Comparable-18.4 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 31 | deepseek-v4-flashOpen weightsDeepSeek · DeepSeek · best of 2 rows | 34.5 | Independent | reasoning_effortmaxversion4.3 | Comparable-18.8 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 32 | Gemini 3.6 FlashClosedGoogle · Gemini 3.6 · best of 2 rows | 34.3 | Independent | reasoningonversion4.3 | Partially comparable-19.0 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 33 | Muse Spark 1.1ClosedMeta AI · best of 2 rows | 34.3 | Independent | reasoning_effortxhighversion4.3 | Partially comparable-19.1 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 34 | Qwen3.8 27BOpen weightsQwen · Qwen3.8 · best of 8 rows | 33.9 | Independent | reasoning_effortxhighversion4.3 | Partially comparable-19.5 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 35 | Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 6 rows | 33.6 | Independent | reasoning_effortmediumversion4.3 | Partially comparable-19.7 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 36 | motif-3Open weightsMotif Technologies · best of 2 rows | 33.6 | Independent | reasoningonversion4.3 | Partially comparable-19.8 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 37 | gpt-5.3-codexClosedOpenAI · GPT 5.3 · best of 2 rows | 32.5 | Independent | reasoning_effortxhighversion4.3 | Partially comparable-20.9 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 38 | motif-0714ClosedMotif Technologies · best of 2 rows | 32.3 | Independent | reasoningonversion4.3 | Partially comparable-21.1 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 39 | Claude Opus 4.6ClosedAnthropic · Claude · best of 4 rows | 31.9 | Independent | reasoningadaptivereasoning_effortmaxversion4.3 | Partially comparable-21.4 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 40 | Kimi K2.6Open weightsMoonshot AI · Kimi · best of 4 rows | 31.3 | Independent | reasoningonversion4.3 | Partially comparable-22.0 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 41 | muse-sparkClosedMeta AI · best of 2 rows | 31.3 | Independent | reasoningonversion4.3 | Partially comparable-22.1 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 42 | deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek · best of 4 rows | 30.9 | Independent | reasoning_effortmaxversion4.3 | Comparable-22.5 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 43 | k2-horizon-375b-a23bOpen weightsMBZUAI Institute of Foundation Models · best of 2 rows | 30.7 | Independent | reasoningonversion4.3 | Partially comparable-22.7 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 44 | Claude Sonnet 4.6ClosedAnthropic · Claude · best of 5 rows | 30.4 | Independent | reasoningadaptivereasoning_effortmaxversion4.3 | Partially comparable-22.9 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 44 | gpt-5.2ClosedOpenAI · GPT 5.2 · best of 6 rows | 30.4 | Independent | reasoning_effortxhighversion4.3 | Partially comparable-22.9 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 46 | apodex-1-1ClosedApodex · best of 2 rows | 30.4 | Independent | reasoningonversion4.3 | Partially comparable-23.0 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 47 | Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.1 · best of 2 rows | 30.4 | Independent | reasoningonversion4.3 | Partially comparable-23.0 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 48 | deepseek-v4-pro-0424-highOpen weightsDeepSeek · DeepSeek | 30.1 | Independent | version4.3 | Partially comparable-23.3 | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 49 | Qwen3.7 MaxClosedQwen · Qwen3.7 · best of 2 rows | 29.9 | Independent | reasoningonversion4.3 | Partially comparable-23.5 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 50 | MiniMax-M3Open weightsMiniMax · MiniMax · best of 2 rows | 29.6 | Independent | reasoningonversion4.3 | Partially comparable-23.8 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 51 | Claude Opus 4.5ClosedAnthropic · Claude · best of 4 rows | 29.1 | Independent | reasoningonversion4.3 | Partially comparable-24.3 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 52 | mimo-v2-proClosedXiaomi · best of 2 rows | 28.6 | Independent | reasoningonversion4.3 | Partially comparable-24.7 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 53 | GPT-5.2-CodexClosedOpenAI · GPT 5.2 · best of 2 rows | 28.5 | Independent | reasoning_effortxhighversion4.3 | Partially comparable-24.9 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 54 | Qwen3.6 Max PreviewClosedQwen · Qwen3.6 · best of 2 rows | 28.4 | Independent | reasoningonversion4.3 | Partially comparable-25.0 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 55 | nex-n2-proOpen weightsNex AGI · best of 2 rows | 28.2 | Independent | reasoningonversion4.3 | Partially comparable-25.2 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 56 | Solar Pro 4ClosedUpstage · Solar · best of 2 rows | 28.1 | Independent | reasoningonversion4.3 | Partially comparable-25.2 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 57 | gemini-3-proClosedGoogle · Gemini 3 · best of 4 rows | 28.0 | Independent | reasoning_efforthighversion4.3 | Partially comparable-25.4 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 58 | GLM 5Open weightsZ.ai (Zhipu AI) · GLM5 · best of 4 rows | 27.9 | Independent | reasoningonversion4.3 | Partially comparable-25.5 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 59 | agnes-2-5-pro-alphaOpen weightsSapiens AI · best of 2 rows | 27.8 | Independent | reasoningonversion4.3 | Partially comparable-25.6 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 60 | jt-4-1-flash-236b-a21bClosedChina Mobile · best of 2 rows | 27.3 | Independent | reasoningoffversion4.3 | Partially comparable-26.1 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 61 | grok-build-0-1-06-16ClosedSpaceXAI · Grok · best of 2 rows | 27.2 | Independent | reasoningonversion4.3 | Partially comparable-26.2 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 62 | quasar-438bClosedMultiverse Computing · best of 2 rows | 27.1 | Independent | reasoning_effortmaxversion4.3 | Comparable-26.2 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 63 | Qwen3.6 PlusClosedQwen · Qwen3.6 · best of 2 rows | 27.0 | Independent | reasoningonversion4.3 | Partially comparable-26.4 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 64 | gpt-5-5-instant-06-26ClosedOpenAI · GPT 5.5 · best of 2 rows | 26.8 | Independent | reasoningonversion4.3 | Partially comparable-26.6 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 65 | GLM 5 TurboClosedZ.ai (Zhipu AI) · GLM5 · best of 2 rows | 26.6 | Independent | reasoningonversion4.3 | Partially comparable-26.8 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 66 | GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 4 rows | 26.4 | Independent | reasoningonversion4.3 | Partially comparable-26.9 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 67 | MiMo-V2.5-ProOpen weightsXiaomi · best of 4 rows | 26.4 | Independent | reasoningonversion4.3 | Partially comparable-27.0 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 68 | Gemini 3 Flash PreviewClosedGoogle · Gemini 3 | 26.3 | Independent | version4.3 | Partially comparable-27.0 | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 68 | gemini-3-flashClosedGoogle · Gemini 3 · best of 3 rows | 26.3 | Independent | reasoningonversion4.3 | Partially comparable-27.0 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 70 | Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi · best of 2 rows | 26.3 | Independent | reasoningonversion4.3 | Partially comparable-27.1 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 71 | Inkling SmallOpen weightsThinking Machines · best of 2 rows | 26.1 | Independent | reasoningonversion4.3 | Partially comparable-27.3 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 72 | Qwen 3.7 PlusClosedQwen · Qwen3.7 · best of 2 rows | 25.8 | Independent | reasoningonversion4.3 | Partially comparable-27.5 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 73 | Hy3Open weightsTencent · best of 5 rows | 25.8 | Independent | reasoningonversion4.3 | Partially comparable-27.6 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 74 | Grok 4.20ClosedxAI · Grok · best of 3 rows | 25.7 | Independent | reasoningonversion4.3 | Partially comparable-27.7 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 75 | InklingOpen weightsThinking Machines · best of 2 rows | 25.5 | Independent | reasoningonversion4.3 | Partially comparable-27.8 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 76 | Grok 4.3ClosedxAI · Grok · best of 8 rows | 25.4 | Independent | reasoning_efforthighversion4.3 | Partially comparable-28.0 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 77 | grok-4-20-0309ClosedSpaceXAI · Grok 4.20 · best of 4 rows | 25.2 | Independent | reasoningonversion4.3 | Partially comparable-28.1 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 78 | mimo-v2-omni-0327ClosedXiaomi · best of 2 rows | 25.1 | Independent | reasoningonversion4.3 | Partially comparable-28.2 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 79 | Ling 3.0 FlashOpen weightsinclusionAI · best of 2 rows | 24.9 | Independent | reasoningonversion4.3 | Partially comparable-28.4 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 80 | GPT-5-CodexClosedOpenAI · GPT 5 · best of 2 rows | 24.9 | Independent | reasoning_efforthighversion4.3 | Partially comparable-28.5 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 81 | deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek · best of 4 rows | 24.8 | Independent | reasoning_efforthighversion4.3 | Partially comparable-28.5 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 81 | deepseek-v4-flash-0420-highOpen weightsDeepSeek · DeepSeek | 24.8 | Independent | version4.3 | Partially comparable-28.5 | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 83 | Ling 3.0 Flash VLOpen weightsinclusionAI · best of 2 rows | 24.8 | Independent | reasoningonversion4.3 | Partially comparable-28.6 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 84 | gpt-5.1ClosedOpenAI · GPT 5.1 · best of 4 rows | 24.7 | Independent | reasoning_efforthighversion4.3 | Partially comparable-28.6 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 84 | solar-open2-250bOpen weightsUpstage · Solar · best of 2 rows | 24.7 | Independent | reasoningonversion4.3 | Partially comparable-28.6 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 86 | gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 6 rows | 24.6 | Independent | reasoning_effortxhighversion4.3 | Partially comparable-28.8 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 87 | mimo-v2-omniClosedXiaomi · best of 2 rows | 23.9 | Independent | reasoningonversion4.3 | Partially comparable-29.4 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 88 | GPT-5.1-CodexClosedOpenAI · GPT 5.1 · best of 2 rows | 23.7 | Independent | reasoning_efforthighversion4.3 | Partially comparable-29.7 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 89 | GLM 5V TurboClosedZ.ai (Zhipu AI) · GLM5 · best of 2 rows | 23.5 | Independent | reasoningonversion4.3 | Partially comparable-29.9 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 90 | Kimi K2.5Open weightsMoonshot AI · Kimi · best of 4 rows | 23.5 | Independent | reasoningonversion4.3 | Partially comparable-29.9 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 91 | Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 3 · best of 2 rows | 23.4 | Independent | reasoningonversion4.3 | Partially comparable-30.0 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 92 | MiniMax M2.7Open weightsMiniMax · MiniMax · best of 2 rows | 23.2 | Independent | reasoningonversion4.3 | Partially comparable-30.1 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 93 | gpt-5ClosedOpenAI · GPT 5 · best of 8 rows | 23.0 | Independent | reasoning_efforthighversion4.3 | Partially comparable-30.4 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 94 | Qwen3.5-27BOpen weightsQwen · Qwen3.5 · best of 4 rows | 22.9 | Independent | reasoningonversion4.3 | Partially comparable-30.5 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 95 | Claude Opus 4.1ClosedAnthropic · Claude · best of 4 rows | 22.9 | Independent | reasoningonversion4.3 | Partially comparable-30.5 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 96 | MiniMax M2.5Open weightsMiniMax · MiniMax · best of 2 rows | 22.8 | Independent | reasoningonversion4.3 | Partially comparable-30.6 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 97 | Hy3 previewOpen weightsTencent | 22.7 | Independent | version4.3 | Partially comparable-30.6 | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 98 | a-x-k2Open weightsSK Telecom · best of 2 rows | 22.7 | Independent | reasoningonversion4.3 | Partially comparable-30.7 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 99 | gpt-5-5-instant-05-26ClosedOpenAI · GPT 5.5 · best of 2 rows | 22.7 | Independent | reasoningonversion4.3 | Partially comparable-30.7 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 100 | Gemini 3.5 Flash-LiteClosedGoogle · Gemini 3.5 · best of 2 rows | 22.7 | Independent | reasoningonversion4.3 | Partially comparable-30.7 | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →