IFBench
precise instruction following with novel constraints
Updated 4 h ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 898
- Models
- 334
- Current leader
- Grok 4.3 83.3%
Score history · Mistral Small 3.2 2 rows
- Mistral Small 3.2
- 33.47%aa_slug=mistral-small-3-2 · evaluator=Artificial Analysis · reasoning=off · index_version=4.312 Sept 2026
- 33.47%aa_slug=mistral-small-3-2 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
Frontier over time · accuracy · evaluator=Artificial Analysis
6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 83.3%Grok 4.3 xAI Independent11 Sept 2026
- 81.3%Grok 4.3 xAI Independent11 Sept 2026
- 81.2%Grok 4.20 xAI Independent11 Sept 2026
- 81.0%Grok 4.3 xAI Independent11 Sept 2026
- 73.5%gemma-4-12B Google Independent11 Sept 2026
- 45.9%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 334 models · trust independent-evaluator
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 201 | Gemini 2.0 FlashClosedGoogle · Gemini 2.0 · best of 2 rows | 40.2% | Independent | reasoningoff | Partially comparable0.00 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 202 | Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 4 rows | 39.9% | Independent | reasoningon | Partially comparable-0.34 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 203 | Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows | 39.8% | Independent | reasoningoff | Partially comparable-0.40 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 204 | deepseek-r1Open weightsDeepSeek · DeepSeek-R1 · best of 2 rows | 39.6% | Independent | reasoningon | Partially comparable-0.61 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 205 | Llama 4 ScoutOpen weightsMeta AI · Llama 4 · best of 2 rows | 39.5% | Independent | reasoningoff | Partially comparable-0.68 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 206 | llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.3 · best of 4 rows | 39.5% | Independent | reasoningoff | Partially comparable-0.74 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 207 | Mistral Medium 3ClosedMistral AI · Mistral · best of 2 rows | 39.3% | Independent | reasoningoff | Partially comparable-0.95 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | ernie-4-5-300b-a47bOpen weightsBaidu · ERNIE 4.5 · best of 2 rows | 39.1% | Independent | reasoningoff | Partially comparable-1.08 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 209 | Llama-3.1-405BRestricted weightsMeta AI · Llama 3.1 · best of 2 rows | 39.0% | Independent | reasoningoff | Partially comparable-1.15 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 210 | deepseek-r1-0120Open weightsDeepSeek · DeepSeek · best of 2 rows | 39.0% | Independent | reasoningon | Partially comparable-1.22 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 211 | QwQ-32BOpen weightsAlibaba Group · Qwen · best of 2 rows | 38.8% | Independent | reasoningon | Partially comparable-1.42 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 212 | qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows | 38.7% | Independent | reasoningon | Partially comparable-1.49 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 213 | granite-4.1-8bOpen weightsIBM · Granite 4.1 · best of 2 rows | 38.6% | Independent | reasoningoff | Partially comparable-1.56 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 214 | gpt-4.1-miniClosedOpenAI · GPT 4.1 · best of 2 rows | 38.3% | Independent | reasoningoff | Partially comparable-1.90 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 215 | llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.1 · best of 2 rows | 38.2% | Independent | reasoningon | Partially comparable-2.04 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 216 | Devstral 2Open weightsMistral AI · Devstral 2 · best of 2 rows | 38.1% | Independent | reasoningoff | Partially comparable-2.10 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 216 | nova-proClosedAmazon Web Services · Nova · best of 2 rows | 38.1% | Independent | reasoningoff | Partially comparable-2.10 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 216 | olmo-2-32bOpen weightsAllen Institute for AI · OLMo 2 · best of 2 rows | 38.1% | Independent | reasoningoff | Partially comparable-2.10 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 219 | gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 4 rows | 38.0% | Independent | reasoningon | Partially comparable-2.17 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | qwen3-5-omni-flashClosedAlibaba Group · Qwen3.5 · best of 2 rows | 38.0% | Independent | reasoningoff | Partially comparable-2.24 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 221 | hyperclova-x-seed-think-32bOpen weightsNaver · Seed · best of 2 rows | 37.9% | Independent | reasoningon | Partially comparable-2.31 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 222 | GLM 4.5 AirOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows | 37.5% | Independent | reasoningon | Partially comparable-2.65 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 223 | llama-3-instruct-70bOpen weightsMeta AI · Llama 3 · best of 2 rows | 37.1% | Independent | reasoningoff | Partially comparable-3.13 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 223 | solar-pro-2ClosedUpstage · Solar · best of 4 rows | 37.1% | Independent | reasoningon | Partially comparable-3.13 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 225 | llama-nemotron-super-49b-v1-5Open weightsNVIDIA · Llama · best of 4 rows | 37.0% | Independent | reasoningon | Partially comparable-3.19 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 226 | Qwen2.5 72B InstructOpen weightsQwen · Qwen2.5 · best of 2 rows | 36.9% | Independent | reasoningoff | Partially comparable-3.33 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 227 | Gemma 3 12BOpen weightsGoogle · Gemma 3 · best of 2 rows | 36.7% | Independent | reasoningoff | Partially comparable-3.47 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 228 | jt-miniClosedChina Mobile · best of 2 rows | 36.7% | Independent | reasoningoff | Partially comparable-3.53 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 229 | Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 4 rows | 36.6% | Independent | reasoningon | Partially comparable-3.60 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 230 | Command AOpen weightsCohere · Command · best of 2 rows | 36.5% | Independent | reasoningoff | Partially comparable-3.74 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 231 | Qwen3 32BOpen weightsQwen · Qwen3 | 36.3% | Independent | group defaults | Partially comparable-3.87 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 231 | exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 4 rows | 36.3% | Independent | reasoningon | Partially comparable-3.87 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 231 | qwen3-32b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows | 36.3% | Independent | reasoningon | Partially comparable-3.87 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 234 | Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows | 36.2% | Independent | reasoningoff | Partially comparable-4.01 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 234 | nova-premierClosedAmazon Web Services · Nova · best of 2 rows | 36.2% | Independent | reasoningoff | Partially comparable-4.01 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 236 | Claude 3 HaikuClosedAnthropic · Claude · best of 2 rows | 36.1% | Independent | reasoningoff | Partially comparable-4.08 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | GPT-4o (2024-08-06)ClosedOpenAI · GPT 4 · best of 2 rows | 36.0% | Independent | reasoningoff | Partially comparable-4.21 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 238 | nanbeige4-1-3bOpen weightsNanbeige · best of 2 rows | 35.4% | Independent | reasoningon | Partially comparable-4.76 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 239 | Qwen3 Coder NextOpen weightsQwen · Qwen3 · best of 2 rows | 35.2% | Independent | reasoningoff | Partially comparable-4.96 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 240 | jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.7 · best of 2 rows | 35.2% | Independent | reasoningoff | Partially comparable-5.03 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 241 | hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.1 · best of 3 rows | 34.8% | Independent | reasoningoff | Partially comparable-5.44 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 241 | ling-1tOpen weightsinclusionAI · best of 2 rows | 34.8% | Independent | reasoningoff | Partially comparable-5.44 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 243 | devstral-smallOpen weightsMistral AI · Devstral · best of 2 rows | 34.6% | Independent | reasoningoff | Partially comparable-5.64 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 244 | Pixtral LargeOpen weightsMistral AI · Pixtral · best of 2 rows | 34.5% | Independent | reasoningoff | Partially comparable-5.71 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 245 | Llama-3.1-70BOpen weightsMeta AI · Llama 3.1 · best of 2 rows | 34.4% | Independent | reasoningoff | Partially comparable-5.78 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 246 | ling-flash-2-0Open weightsinclusionAI · best of 2 rows | 34.4% | Independent | reasoningoff | Partially comparable-5.85 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 246 | sarvam-105bOpen weightsSarvam · best of 2 rows | 34.4% | Independent | reasoning_efforthigh | Partially comparable-5.85 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 248 | gpt-4oClosedOpenAI · GPT 4 · best of 2 rows | 34.3% | Independent | reasoningoff | Partially comparable-5.91 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 249 | GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 4 rows | 34.2% | Independent | reasoningon | Partially comparable-5.98 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 250 | nova-liteClosedAmazon Web Services · Nova · best of 2 rows | 34.1% | Independent | reasoningoff | Partially comparable-6.05 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 251 | intellect-3Open weightsPrime Intellect · best of 2 rows | 34.0% | Independent | reasoningon | Partially comparable-6.19 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 252 | granite-4.1-3bOpen weightsIBM · Granite 4.1 · best of 2 rows | 33.7% | Independent | reasoningoff | Partially comparable-6.53 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows | 33.5% | Independent | reasoningoff | Partially comparable-6.73 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | Qwen3 8BOpen weightsQwen · Qwen3 | 33.5% | Independent | group defaults | Partially comparable-6.73 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | qwen3-8b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows | 33.5% | Independent | reasoningon | Partially comparable-6.73 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 256 | gpt-4ClosedOpenAI · GPT 4 · best of 2 rows | 33.2% | Independent | reasoningoff | Partially comparable-7.00 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 257 | LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.5 · best of 2 rows | 33.1% | Independent | reasoningoff | Partially comparable-7.07 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 258 | Olmo-3-7B-InstructOpen weightsAllen Institute for AI · OLMo 3 · best of 2 rows | 32.8% | Independent | reasoningoff | Partially comparable-7.41 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 259 | Hermes 4 405BOpen weightsNous Research · Hermes 4 | 32.7% | Independent | group defaults | Partially comparable-7.48 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 260 | Qwen3 Coder 30B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows | 32.6% | Independent | reasoningoff | Partially comparable-7.55 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 261 | Qwen3-4BOpen weightsQwen · Qwen3 | 32.5% | Independent | group defaults | Partially comparable-7.68 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 261 | qwen3-4b-instructOpen weightsAlibaba Group · Qwen3 | 32.5% | Independent | reasoningon | Partially comparable-7.68 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 263 | Ministral 3 14BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 32.0% | Independent | reasoningoff | Partially comparable-8.16 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 263 | gpt-4.1-nanoClosedOpenAI · GPT 4.1 · best of 2 rows | 32.0% | Independent | reasoningoff | Partially comparable-8.16 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 265 | nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 4 rows | 31.9% | Independent | reasoningon | Partially comparable-8.30 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 266 | Gemma 3 27BOpen weightsGoogle · Gemma 3 · best of 2 rows | 31.8% | Independent | reasoningoff | Partially comparable-8.36 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 266 | sarvam-m-reasoningOpen weightsSarvam · best of 2 rows | 31.8% | Independent | reasoningon | Partially comparable-8.36 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 268 | Devstral Small 1.0Open weightsMistral AI · Devstral · best of 2 rows | 31.6% | Independent | reasoningoff | Partially comparable-8.57 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 268 | Mistral Large 2.0Open weightsMistral AI · Mistral · best of 2 rows | 31.6% | Independent | reasoningoff | Partially comparable-8.57 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 270 | Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 4 rows | 31.5% | Independent | reasoningon | Partially comparable-8.70 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 270 | granite-4-0-h-smallOpen weightsIBM · Granite 4.0 · best of 2 rows | 31.5% | Independent | reasoningoff | Partially comparable-8.70 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 272 | jamba-1-7-miniOpen weightsAI21 Labs · Jamba 1.7 · best of 2 rows | 31.4% | Independent | reasoningoff | Partially comparable-8.84 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 273 | Hermes-4-70BRestricted weightsNous Research · Hermes 4 | 31.3% | Independent | group defaults | Partially comparable-8.91 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 273 | hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.1 · best of 3 rows | 31.3% | Independent | reasoningon | Partially comparable-8.91 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 275 | mistral-large-2Open weightsMistral AI · Mistral · best of 2 rows | 31.2% | Independent | reasoningoff | Partially comparable-8.98 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 276 | Devstral Small 2Open weightsMistral AI · Devstral · best of 2 rows | 31.2% | Independent | reasoningoff | Partially comparable-9.04 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 277 | gpt-4o-miniClosedOpenAI · GPT 4 · best of 2 rows | 30.9% | Independent | reasoningoff | Partially comparable-9.25 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 278 | llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.1 · best of 2 rows | 30.8% | Independent | reasoningoff | Partially comparable-9.45 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 279 | Reka Flash 3Open weightsrekaai · best of 2 rows | 30.4% | Independent | reasoningon | Partially comparable-9.79 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 279 | llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.2 · best of 2 rows | 30.4% | Independent | reasoningoff | Partially comparable-9.79 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 281 | GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 4 rows | 30.1% | Independent | reasoningon | Partially comparable-10.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 282 | Mistral Small 3.1Open weightsMistral AI · Mistral · best of 2 rows | 29.9% | Independent | reasoningoff | Partially comparable-10.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 282 | devstral-mediumClosedMistral AI · Devstral · best of 2 rows | 29.9% | Independent | reasoningoff | Partially comparable-10.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 284 | nova-microClosedAmazon Web Services · Nova · best of 2 rows | 29.4% | Independent | reasoningoff | Partially comparable-10.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 285 | Ministral 3 8BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 29.1% | Independent | reasoningoff | Partially comparable-11.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 286 | Llama 3.1 8BRestricted weightsMeta AI · Llama 3.1 · best of 2 rows | 28.6% | Independent | reasoningoff | Partially comparable-11.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 287 | Gemma 3 4BRestricted weightsGoogle · Gemma 3 · best of 2 rows | 28.3% | Independent | reasoningoff | Partially comparable-11.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 288 | Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AI · Kimi · best of 2 rows | 28.1% | Independent | reasoningoff | Partially comparable-12.1 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 289 | gemma-3n-e4bOpen weightsGoogle · Gemma 3 · best of 2 rows | 27.9% | Independent | reasoningoff | Partially comparable-12.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 290 | NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 4 rows | 27.6% | Independent | reasoningon | Partially comparable-12.6 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 291 | R1 Distill Llama 70BOpen weightsDeepSeek · Llama · best of 2 rows | 27.6% | Independent | reasoningon | Partially comparable-12.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 292 | Molmo2-8BOpen weightsAllen Institute for AI · best of 2 rows | 26.9% | Independent | reasoningoff | Partially comparable-13.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 293 | qwen3-1.7b-instructOpen weightsAlibaba Group · Qwen3.1 · best of 4 rows | 26.9% | Independent | reasoningon | Partially comparable-13.3 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 294 | Ministral 3 3BOpen weightsMistral AI · Ministral 3 · best of 2 rows | 26.8% | Independent | reasoningoff | Partially comparable-13.4 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 295 | minicpm-v4-6-1-3bOpen weightsOpenBMB · best of 2 rows | 26.7% | Independent | reasoningoff | Partially comparable-13.5 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 296 | sarvam-30bOpen weightsSarvam · best of 2 rows | 26.5% | Independent | reasoning_efforthigh | Partially comparable-13.7 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 297 | Mistral Small 3Open weightsMistral AI · Mistral · best of 2 rows | 26.4% | Independent | reasoningoff | Partially comparable-13.8 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 298 | lfm2-8b-a1bOpen weightsLiquid AI · LFM2 · best of 2 rows | 26.3% | Independent | reasoningoff | Partially comparable-13.9 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 299 | Llama-3.2-3BRestricted weightsMeta AI · Llama 3.2 · best of 2 rows | 26.2% | Independent | reasoningoff | Partially comparable-14.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
| 299 | granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.0 · best of 2 rows | 26.2% | Independent | reasoningoff | Partially comparable-14.0 pt | obs. 12 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →