IFBench
precise instruction following with novel constraints
Updated 29 min ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 450
- Models
- 333
- Current leader
- Grok 4.3 83.3%
Score history · Llama-3.1-70B 1 row
Not enough history to chart — a single observation (34.42% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.
- 34.42%aa_slug=llama-3-1-instruct-70b · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
Frontier over time · accuracy · evaluator=Artificial Analysis
6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 83.3%Grok 4.3 xAI Independent11 Sept 2026
- 81.3%Grok 4.3 xAI Independent11 Sept 2026
- 81.2%Grok 4.20 xAI Independent11 Sept 2026
- 81.0%Grok 4.3 xAI Independent11 Sept 2026
- 73.5%gemma-4-12B Google Independent11 Sept 2026
- 45.9%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 333 models
Select models with +, then Compare.
| # | Model | Score | Trust | Configuration | vs leader | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|---|---|
| 201 | Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 2 rows | 39.9% | Independent | reasoningon | Partially comparable0.00 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 202 | Mistral Medium 3.1ClosedMistral AI · Mistral | 39.8% | Independent | group defaults | Partially comparable-0.06 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 203 | deepseek-r1Open weightsDeepSeek · DeepSeek-R1 | 39.6% | Independent | group defaults | Partially comparable-0.27 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 204 | Llama 4 ScoutOpen weightsMeta AI · Llama 4 | 39.5% | Independent | group defaults | Partially comparable-0.34 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 205 | llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.3 · best of 2 rows | 39.5% | Independent | group defaults | Partially comparable-0.40 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 206 | Mistral Medium 3ClosedMistral AI · Mistral | 39.3% | Independent | group defaults | Partially comparable-0.61 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 207 | ernie-4-5-300b-a47bOpen weightsBaidu · ERNIE 4.5 | 39.1% | Independent | group defaults | Partially comparable-0.74 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 208 | Llama-3.1-405BRestricted weightsMeta AI · Llama 3.1 | 39.0% | Independent | group defaults | Partially comparable-0.81 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 209 | deepseek-r1-0120Open weightsDeepSeek · DeepSeek | 39.0% | Independent | group defaults | Partially comparable-0.88 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 210 | QwQ-32BOpen weightsAlibaba Group · Qwen | 38.8% | Independent | group defaults | Partially comparable-1.08 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 211 | qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows | 38.7% | Independent | reasoningon | Partially comparable-1.15 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 212 | granite-4.1-8bOpen weightsIBM · Granite 4.1 | 38.6% | Independent | group defaults | Partially comparable-1.22 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 213 | gpt-4.1-miniClosedOpenAI · GPT 4.1 | 38.3% | Independent | group defaults | Partially comparable-1.56 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 214 | llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.1 | 38.2% | Independent | group defaults | Partially comparable-1.70 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 215 | Devstral 2Open weightsMistral AI · Devstral 2 | 38.1% | Independent | group defaults | Partially comparable-1.76 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 215 | nova-proClosedAmazon Web Services · Nova | 38.1% | Independent | group defaults | Partially comparable-1.76 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 215 | olmo-2-32bOpen weightsAllen Institute for AI · OLMo 2 | 38.1% | Independent | group defaults | Partially comparable-1.76 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 218 | gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 2 rows | 38.0% | Independent | group defaults | Partially comparable-1.83 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 219 | qwen3-5-omni-flashClosedAlibaba Group · Qwen3.5 | 38.0% | Independent | group defaults | Partially comparable-1.90 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 220 | hyperclova-x-seed-think-32bOpen weightsNaver · Seed | 37.9% | Independent | group defaults | Partially comparable-1.97 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 221 | GLM 4.5 AirOpen weightsZ.ai (Zhipu AI) · GLM4.5 | 37.5% | Independent | group defaults | Partially comparable-2.31 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 222 | llama-3-instruct-70bOpen weightsMeta AI · Llama 3 | 37.1% | Independent | group defaults | Partially comparable-2.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 222 | solar-pro-2ClosedUpstage · Solar · best of 2 rows | 37.1% | Independent | reasoningon | Partially comparable-2.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 224 | llama-nemotron-super-49b-v1-5Open weightsNVIDIA · Llama · best of 2 rows | 37.0% | Independent | reasoningon | Partially comparable-2.85 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 225 | Qwen2.5 72B InstructOpen weightsQwen · Qwen2.5 | 36.9% | Independent | group defaults | Partially comparable-2.99 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 226 | Gemma 3 12BOpen weightsGoogle · Gemma 3 | 36.7% | Independent | group defaults | Partially comparable-3.13 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 227 | jt-miniClosedChina Mobile | 36.7% | Independent | group defaults | Partially comparable-3.19 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 228 | Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 2 rows | 36.6% | Independent | reasoningon | Partially comparable-3.26 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 229 | Command AOpen weightsCohere · Command | 36.5% | Independent | group defaults | Partially comparable-3.40 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 230 | Qwen3 32BOpen weightsQwen · Qwen3 | 36.3% | Independent | group defaults | Partially comparable-3.53 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 230 | exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows | 36.3% | Independent | reasoningon | Partially comparable-3.53 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 232 | Mistral Large 3Open weightsMistral AI · Mistral | 36.2% | Independent | group defaults | Partially comparable-3.67 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 232 | nova-premierClosedAmazon Web Services · Nova | 36.2% | Independent | group defaults | Partially comparable-3.67 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 234 | Claude 3 HaikuClosedAnthropic · Claude | 36.1% | Independent | group defaults | Partially comparable-3.74 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 235 | GPT-4o (2024-08-06)ClosedOpenAI · GPT 4 | 36.0% | Independent | group defaults | Partially comparable-3.87 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 236 | nanbeige4-1-3bOpen weightsNanbeige | 35.4% | Independent | group defaults | Partially comparable-4.42 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 237 | Qwen3 Coder NextOpen weightsQwen · Qwen3 | 35.2% | Independent | group defaults | Partially comparable-4.62 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 238 | jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.7 | 35.2% | Independent | group defaults | Partially comparable-4.69 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 239 | hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.1 | 34.8% | Independent | group defaults | Partially comparable-5.10 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 239 | ling-1tOpen weightsinclusionAI | 34.8% | Independent | group defaults | Partially comparable-5.10 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 241 | devstral-smallOpen weightsMistral AI · Devstral | 34.6% | Independent | group defaults | Partially comparable-5.30 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 242 | Pixtral LargeOpen weightsMistral AI · Pixtral | 34.5% | Independent | group defaults | Partially comparable-5.37 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 243 | Llama-3.1-70BOpen weightsMeta AI · Llama 3.1 | 34.4% | Independent | group defaults | Partially comparable-5.44 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 244 | ling-flash-2-0Open weightsinclusionAI | 34.4% | Independent | group defaults | Partially comparable-5.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 244 | sarvam-105bOpen weightsSarvam | 34.4% | Independent | group defaults | Partially comparable-5.51 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 246 | gpt-4oClosedOpenAI · GPT 4 | 34.3% | Independent | group defaults | Partially comparable-5.57 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 247 | GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows | 34.2% | Independent | group defaults | Partially comparable-5.64 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 248 | nova-liteClosedAmazon Web Services · Nova | 34.1% | Independent | group defaults | Partially comparable-5.71 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 249 | intellect-3Open weightsPrime Intellect | 34.0% | Independent | group defaults | Partially comparable-5.85 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 250 | granite-4.1-3bOpen weightsIBM · Granite 4.1 | 33.7% | Independent | group defaults | Partially comparable-6.19 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 251 | Mistral Small 3.2Open weightsMistral AI · Mistral | 33.5% | Independent | group defaults | Partially comparable-6.39 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 251 | Qwen3 8BOpen weightsQwen · Qwen3 | 33.5% | Independent | group defaults | Partially comparable-6.39 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 253 | gpt-4ClosedOpenAI · GPT 4 | 33.2% | Independent | group defaults | Partially comparable-6.66 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 254 | LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.5 | 33.1% | Independent | group defaults | Partially comparable-6.73 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 255 | Olmo-3-7B-InstructOpen weightsAllen Institute for AI · OLMo 3 | 32.8% | Independent | group defaults | Partially comparable-7.07 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 256 | Hermes 4 405BOpen weightsNous Research · Hermes 4 | 32.7% | Independent | group defaults | Partially comparable-7.14 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 257 | Qwen3 Coder 30B A3B InstructOpen weightsQwen · Qwen3 | 32.6% | Independent | group defaults | Partially comparable-7.21 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 258 | Qwen3-4BOpen weightsQwen · Qwen3 | 32.5% | Independent | group defaults | Partially comparable-7.34 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 259 | Ministral 3 14BOpen weightsMistral AI · Ministral 3 | 32.0% | Independent | group defaults | Partially comparable-7.82 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 259 | gpt-4.1-nanoClosedOpenAI · GPT 4.1 | 32.0% | Independent | group defaults | Partially comparable-7.82 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 261 | nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 2 rows | 31.9% | Independent | reasoningon | Partially comparable-7.96 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | Gemma 3 27BOpen weightsGoogle · Gemma 3 | 31.8% | Independent | group defaults | Partially comparable-8.02 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 262 | sarvam-m-reasoningOpen weightsSarvam | 31.8% | Independent | group defaults | Partially comparable-8.02 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 264 | Devstral Small 1.0Open weightsMistral AI · Devstral | 31.6% | Independent | group defaults | Partially comparable-8.23 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 264 | Mistral Large 2.0Open weightsMistral AI · Mistral | 31.6% | Independent | group defaults | Partially comparable-8.23 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 266 | Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 2 rows | 31.5% | Independent | group defaults | Partially comparable-8.36 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 266 | granite-4-0-h-smallOpen weightsIBM · Granite 4.0 | 31.5% | Independent | group defaults | Partially comparable-8.36 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 266 | qwen3-32b-instructOpen weightsAlibaba Group · Qwen3 | 31.5% | Independent | group defaults | Partially comparable-8.36 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 269 | jamba-1-7-miniOpen weightsAI21 Labs · Jamba 1.7 | 31.4% | Independent | group defaults | Partially comparable-8.50 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 270 | Hermes-4-70BRestricted weightsNous Research · Hermes 4 | 31.3% | Independent | group defaults | Partially comparable-8.57 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 271 | mistral-large-2Open weightsMistral AI · Mistral | 31.2% | Independent | group defaults | Partially comparable-8.64 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 272 | Devstral Small 2Open weightsMistral AI · Devstral | 31.2% | Independent | group defaults | Partially comparable-8.70 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 273 | gpt-4o-miniClosedOpenAI · GPT 4 | 30.9% | Independent | group defaults | Partially comparable-8.91 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 274 | llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.1 | 30.8% | Independent | group defaults | Partially comparable-9.11 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 275 | Reka Flash 3Open weightsrekaai | 30.4% | Independent | group defaults | Partially comparable-9.45 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 275 | llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.2 | 30.4% | Independent | group defaults | Partially comparable-9.45 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 277 | GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows | 30.1% | Independent | group defaults | Partially comparable-9.79 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 278 | Mistral Small 3.1Open weightsMistral AI · Mistral | 29.9% | Independent | group defaults | Partially comparable-9.93 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 278 | devstral-mediumClosedMistral AI · Devstral | 29.9% | Independent | group defaults | Partially comparable-9.93 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 280 | nova-microClosedAmazon Web Services · Nova | 29.4% | Independent | group defaults | Partially comparable-10.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 281 | Ministral 3 8BOpen weightsMistral AI · Ministral 3 | 29.1% | Independent | group defaults | Partially comparable-10.7 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 282 | hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.1 | 29.0% | Independent | group defaults | Partially comparable-10.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 283 | Llama 3.1 8BRestricted weightsMeta AI · Llama 3.1 | 28.6% | Independent | group defaults | Partially comparable-11.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 283 | qwen3-8b-instructOpen weightsAlibaba Group · Qwen3 | 28.6% | Independent | group defaults | Partially comparable-11.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 285 | Gemma 3 4BRestricted weightsGoogle · Gemma 3 | 28.3% | Independent | group defaults | Partially comparable-11.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 286 | Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AI · Kimi | 28.1% | Independent | group defaults | Partially comparable-11.8 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 287 | gemma-3n-e4bOpen weightsGoogle · Gemma 3 | 27.9% | Independent | group defaults | Partially comparable-12.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 288 | NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 2 rows | 27.6% | Independent | group defaults | Partially comparable-12.2 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 289 | R1 Distill Llama 70BOpen weightsDeepSeek · Llama | 27.6% | Independent | group defaults | Partially comparable-12.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 290 | Molmo2-8BOpen weightsAllen Institute for AI | 26.9% | Independent | group defaults | Partially comparable-12.9 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 291 | qwen3-1.7b-instructOpen weightsAlibaba Group · Qwen3.1 · best of 2 rows | 26.9% | Independent | reasoningon | Partially comparable-13.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 292 | Ministral 3 3BOpen weightsMistral AI · Ministral 3 | 26.8% | Independent | group defaults | Partially comparable-13.1 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 293 | minicpm-v4-6-1-3bOpen weightsOpenBMB | 26.7% | Independent | group defaults | Partially comparable-13.1 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 294 | sarvam-30bOpen weightsSarvam | 26.5% | Independent | group defaults | Partially comparable-13.4 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 295 | Mistral Small 3Open weightsMistral AI · Mistral | 26.4% | Independent | group defaults | Partially comparable-13.5 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 296 | lfm2-8b-a1bOpen weightsLiquid AI · LFM2 | 26.3% | Independent | group defaults | Partially comparable-13.6 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 297 | Llama-3.2-3BRestricted weightsMeta AI · Llama 3.2 | 26.2% | Independent | group defaults | Partially comparable-13.7 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 297 | granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.0 | 26.2% | Independent | group defaults | Partially comparable-13.7 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 299 | apertus-70b-instructOpen weightsSwiss AI Initiative | 25.9% | Independent | group defaults | Partially comparable-14.0 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
| 300 | llama-3-1-nemotron-nano-4b-reasoningOpen weightsNVIDIA · Llama 3.1 | 25.5% | Independent | group defaults | Partially comparable-14.3 pt | obs. 11 Sept 2026 | artificialanalysis.aiT2 | History |
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →