Updated 10 min ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 821
- Models
- 374
- Current leader
- gpt-5.6-sol 65.9%
Score history · Mistral Medium 3 1 row
Not enough history to chart — a single observation (3.79% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.
- 3.79%aa_slug=mistral-medium-3 · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
Frontier over time · accuracy · variant=hard · evaluator=Artificial Analysis
10 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 65.9%gpt-5.6-sol OpenAI Independent11 Sept 2026
- 61.4%gpt-5.6-sol OpenAI Independent11 Sept 2026
- 53.0%Claude Sonnet 4.6 Anthropic Independent11 Sept 2026
- 51.5%Claude Opus 4.7 Anthropic Independent11 Sept 2026
- 50.8%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
- 49.2%KAT-Coder-Pro V2 Kwaipilot Independent11 Sept 2026
- 33.3%GPT-5.1-Codex Mini OpenAI Independent11 Sept 2026
- 26.5%Grok 4.3 xAI Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 315 models
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 196 | mistral-large-2Open weightsMistral AI · Mistral | 6.06% | Independent | group defaults | Comparable0.00 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 196 | nova-proClosedAmazon Web Services · Nova | 6.06% | Independent | group defaults | Comparable0.00 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 196 | qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows | 6.06% | Independent | reasoningon | Partially comparable0.00 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 196 | qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 2 rows | 6.06% | Independent | group defaults | Comparable0.00 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 205 | llama-nemotron-super-49b-v1-5Open weightsNVIDIA · Llama · best of 2 rows | 5.30% | Independent | reasoningon | Partially comparable-0.76 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 205 | qwen3-14b-instructOpen weightsAlibaba Group · Qwen3 | 5.30% | Independent | group defaults | Comparable-0.76 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 205 | step-3-vl-10bOpen weightsStepFun · Step3 | 5.30% | Independent | group defaults | Comparable-0.76 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | Gemini 2.5 Flash-LiteClosedGoogle · Gemini · best of 2 rows | 4.55% | Independent | group defaults | Comparable-1.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | Hermes-4-70BRestricted weightsNous Research · Hermes 4 | 4.55% | Independent | group defaults | Comparable-1.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | LFM2.5-8B-A1BOpen weightsLiquid AI · LFM2.5 | 4.55% | Independent | group defaults | Comparable-1.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | Magistral Small 1.2Open weightsMistral AI · Magistral | 4.55% | Independent | group defaults | Comparable-1.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | Ministral 3 14BOpen weightsMistral AI · Ministral 3 | 4.55% | Independent | group defaults | Comparable-1.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | Ministral 3 8BOpen weightsMistral AI · Ministral 3 | 4.55% | Independent | group defaults | Comparable-1.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | Qwen2.5 72B InstructOpen weightsQwen · Qwen2.5 | 4.55% | Independent | group defaults | Comparable-1.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.1 | 4.55% | Independent | group defaults | Comparable-1.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | magistral-smallOpen weightsMistral AI · Magistral | 4.55% | Independent | group defaults | Comparable-1.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 2 rows | 4.55% | Independent | reasoningon | Partially comparable-1.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | qwen3-4b-2507-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows | 4.55% | Independent | group defaults | Comparable-1.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | solar-pro-2ClosedUpstage · Solar · best of 2 rows | 4.55% | Independent | group defaults | Comparable-1.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | Gemini 2.0 FlashClosedGoogle · Gemini | 3.79% | Independent | group defaults | Comparable-2.27 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | Gemma 3 27BOpen weightsGoogle · Gemma 3 | 3.79% | Independent | group defaults | Comparable-2.27 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | Mistral Medium 3ClosedMistral AI · Mistral | 3.79% | Independent | group defaults | Comparable-2.27 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | Phi 4Open weightsMicrosoft · Phi4 | 3.79% | Independent | group defaults | Comparable-2.27 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | Qwen3 14BOpen weightsQwen · Qwen3 | 3.79% | Independent | group defaults | Comparable-2.27 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 2 rows | 3.79% | Independent | reasoningon | Partially comparable-2.27 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 2 rows | 3.79% | Independent | group defaults | Comparable-2.27 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows | 3.79% | Independent | reasoningon | Partially comparable-2.27 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | gpt-4.1-nanoClosedOpenAI · GPT 4.1 | 3.79% | Independent | group defaults | Comparable-2.27 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | motif-2-12-7bClosedMotif Technologies | 3.79% | Independent | group defaults | Comparable-2.27 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows | 3.79% | Independent | reasoningon | Partially comparable-2.27 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 231 | Llama 3.3 70BOpen weightsMeta AI · Llama 3.3 | 3.03% | Independent | group defaults | Comparable-3.03 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 231 | Llama-3.1-70BOpen weightsMeta AI · Llama 3.1 | 3.03% | Independent | group defaults | Comparable-3.03 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 231 | MiniMax-M1-80kOpen weightsMiniMax · MiniMax | 3.03% | Independent | group defaults | Comparable-3.03 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 231 | Qwen3 32BOpen weightsQwen · Qwen3 | 3.03% | Independent | group defaults | Comparable-3.03 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 231 | gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 2 rows | 3.03% | Independent | group defaults | Comparable-3.03 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 231 | midm-250-pro-rsnsftClosedKorea Telecom | 3.03% | Independent | group defaults | Comparable-3.03 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | Claude Haiku 3.5ClosedAnthropic · Claude | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | MiniMax-M1-40kOpen weightsMiniMax · MiniMax | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | Qwen3 8BOpen weightsQwen · Qwen3 | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | falcon-h1r-7bOpen weightsTII UAE · Falcon | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | gemma-3n-e4bOpen weightsGoogle · Gemma 3 | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | granite-4-0-h-smallOpen weightsIBM · Granite 4.0 | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | granite-4.1-30bOpen weightsIBM · Granite 4.1 | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | granite-4.1-3bOpen weightsIBM · Granite 4.1 | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.7 | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.1 | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | mi-dm-k-2-5-pro-dec28ClosedKorea Telecom | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | qwen3-8b-instructOpen weightsAlibaba Group · Qwen3 | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | sarvam-30bOpen weightsSarvam | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | sarvam-m-reasoningOpen weightsSarvam | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | solar-open-100b-reasoningOpen weightsUpstage · Solar | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | tri-21b-think-previewOpen weightsTrillion Labs | 2.27% | Independent | group defaults | Comparable-3.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | DeepSeek-R1-0528-Qwen3-8BOpen weightsDeepSeek · Qwen3 | 1.52% | Independent | group defaults | Comparable-4.54 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | Granite 4.0 MicroOpen weightsIBM · Granite 4.0 | 1.52% | Independent | group defaults | Comparable-4.54 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | Llama 4 ScoutOpen weightsMeta AI · Llama 4 | 1.52% | Independent | group defaults | Comparable-4.54 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 2 rows | 1.52% | Independent | group defaults | Comparable-4.54 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 2 rows | 1.52% | Independent | reasoningon | Partially comparable-4.54 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | R1 Distill Llama 70BOpen weightsDeepSeek · Llama | 1.52% | Independent | group defaults | Comparable-4.54 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | nova-microClosedAmazon Web Services · Nova | 1.52% | Independent | group defaults | Comparable-4.54 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | olmo-3-32b-thinkOpen weightsAllen Institute for AI · OLMo 3 | 1.52% | Independent | group defaults | Comparable-4.54 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | sarvam-105bOpen weightsSarvam | 1.52% | Independent | group defaults | Comparable-4.54 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | Claude 3 HaikuClosedAnthropic · Claude | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | Command AOpen weightsCohere · Command | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | Gemma 3 12BOpen weightsGoogle · Gemma 3 | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | Gemma 3 4BRestricted weightsGoogle · Gemma 3 | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | Llama 3.1 8BRestricted weightsMeta AI · Llama 3.1 | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | Olmo-3-7B-ThinkOpen weightsAllen Institute for AI · OLMo 3 | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | gemma-3n-e2bOpen weightsGoogle · Gemma 3 | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | jamba-reasoning-3bOpen weightsAI21 Labs · Jamba | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | lfm2-2-6bOpen weightsLiquid AI · LFM2.2 | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | ling-mini-2-0Open weightsinclusionAI | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.2 | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | llama-3-instruct-70bOpen weightsMeta AI · Llama 3 | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | nova-liteClosedAmazon Web Services · Nova | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | tri-21b-think-v0-5Open weightsTrillion Labs | 0.76% | Independent | group defaults | Comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | LFM2-1.2BOpen weightsLiquid AI · LFM2.1 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | LFM2.5-1.2B-InstructOpen weightsLiquid AI · LFM2.5 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.5 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | Llama-3.2-1BRestricted weightsMeta AI · Llama 3.2 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | Ministral 3 3BOpen weightsMistral AI · Ministral 3 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | Molmo2-8BOpen weightsAllen Institute for AI | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | Olmo-3-7B-InstructOpen weightsAllen Institute for AI · OLMo 3 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | Qwen3-0.6BOpen weightsQwen · Qwen3.0 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 2 rows | 0% | Independent | reasoningoff | Partially comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | Reka Flash 3Open weightsrekaai | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | apertus-70b-instructOpen weightsSwiss AI Initiative | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | apertus-8b-instructOpen weightsSwiss AI Initiative | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | exaone-4-0-1-2bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows | 0% | Independent | reasoningon | Partially comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | gemma-3-1bOpen weightsGoogle · Gemma 3 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | gemma-3-270mOpen weightsGoogle · Gemma 3 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | granite-3-3-8b-instructOpen weightsIBM · Granite 3.3 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | granite-4-0-350mOpen weightsIBM · Granite 4.0 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | granite-4-0-h-350mOpen weightsIBM · Granite 4.0 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.0 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | granite-4-0-nano-1bOpen weightsIBM · Granite 4.0 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | granite-4.1-8bOpen weightsIBM · Granite 4.1 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.1 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | jamba-1-7-miniOpen weightsAI21 Labs · Jamba 1.7 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | lfm2-24b-a2bOpen weightsLiquid AI · LFM2 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | lfm2-5-1-2b-thinkingOpen weightsLiquid AI · LFM2.5 | 0% | Independent | group defaults | Comparable-6.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →