Skip to content
AI Atlas
BenchmarkActivecategory · instruction-followingfamily · ifeval · variant IFBench

IFBench

precise instruction following with novel constraints

quality51

Updated 29 min ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
450
Models
333
Current leader
Grok 4.3 83.3%

Score history · Qwen3.6 Max Preview 1 row

Not enough history to chart — a single observation (76.6% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.

  • 76.6%aa_slug=qwen3-6-max · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · evaluator=Artificial Analysis

6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 83.3%Grok 4.3 xAI Independent11 Sept 2026
  2. 81.3%Grok 4.3 xAI Independent11 Sept 2026
  3. 81.2%Grok 4.20 xAI Independent11 Sept 2026
  4. 81.0%Grok 4.3 xAI Independent11 Sept 2026
  5. 73.5%gemma-4-12B Google Independent11 Sept 2026
  6. 45.9%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 333 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
301exaone-4-0-1-2bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows25.3%Independentgroup defaultsPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
302magistral-mediumClosedMistral AI · Magistral25.1%Independentgroup defaultsPartially comparable-0.21 ptobs. 11 Sept 2026artificialanalysis.aiT2History
303Granite 4.0 MicroOpen weightsIBM · Granite 4.024.8%Independentgroup defaultsPartially comparable-0.55 ptobs. 11 Sept 2026artificialanalysis.aiT2History
303magistral-smallOpen weightsMistral AI · Magistral24.8%Independentgroup defaultsPartially comparable-0.55 ptobs. 11 Sept 2026artificialanalysis.aiT2History
305llama-3-instruct-8bOpen weightsMeta AI · Llama 324.6%Independentgroup defaultsPartially comparable-0.75 ptobs. 11 Sept 2026artificialanalysis.aiT2History
306olmo-2-7bOpen weightsAllen Institute for AI · OLMo 224.4%Independentgroup defaultsPartially comparable-0.89 ptobs. 11 Sept 2026artificialanalysis.aiT2History
307qwen3-14b-instructOpen weightsAlibaba Group · Qwen323.9%Independentgroup defaultsPartially comparable-1.36 ptobs. 11 Sept 2026artificialanalysis.aiT2History
308phi-3-miniOpen weightsMicrosoft · Phi323.9%Independentgroup defaultsPartially comparable-1.43 ptobs. 11 Sept 2026artificialanalysis.aiT2History
309ling-mini-2-0Open weightsinclusionAI23.6%Independentgroup defaultsPartially comparable-1.70 ptobs. 11 Sept 2026artificialanalysis.aiT2History
310Phi 4Open weightsMicrosoft · Phi423.5%Independentgroup defaultsPartially comparable-1.77 ptobs. 11 Sept 2026artificialanalysis.aiT2History
311Qwen3-0.6BOpen weightsQwen · Qwen3.023.3%Independentgroup defaultsPartially comparable-1.98 ptobs. 11 Sept 2026artificialanalysis.aiT2History
312DeepSeek-R1-Distill-Qwen-32BOpen weightsDeepSeek · Qwen22.9%Independentgroup defaultsPartially comparable-2.38 ptobs. 11 Sept 2026artificialanalysis.aiT2History
313Llama-3.2-1BRestricted weightsMeta AI · Llama 3.222.8%Independentgroup defaultsPartially comparable-2.52 ptobs. 11 Sept 2026artificialanalysis.aiT2History
314granite-3-3-8b-instructOpen weightsIBM · Granite 3.322.4%Independentgroup defaultsPartially comparable-2.86 ptobs. 11 Sept 2026artificialanalysis.aiT2History
315apertus-8b-instructOpen weightsSwiss AI Initiative22.4%Independentgroup defaultsPartially comparable-2.93 ptobs. 11 Sept 2026artificialanalysis.aiT2History
316deepseek-r1-distill-qwen-14bOpen weightsDeepSeek · Qwen22.1%Independentgroup defaultsPartially comparable-3.20 ptobs. 11 Sept 2026artificialanalysis.aiT2History
317gemma-3n-e2bOpen weightsGoogle · Gemma 322.0%Independentgroup defaultsPartially comparable-3.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
318LFM2-1.2BOpen weightsLiquid AI · LFM2.122.0%Independentgroup defaultsPartially comparable-3.34 ptobs. 11 Sept 2026artificialanalysis.aiT2History
319qwen3-0.6b-instructOpen weightsAlibaba Group · Qwen3.021.9%Independentgroup defaultsPartially comparable-3.41 ptobs. 11 Sept 2026artificialanalysis.aiT2History
320Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 2 rows21.6%IndependentreasoningoffPartially comparable-3.75 ptobs. 11 Sept 2026artificialanalysis.aiT2History
321phi-4-miniOpen weightsMicrosoft · Phi421.1%Independentgroup defaultsPartially comparable-4.22 ptobs. 11 Sept 2026artificialanalysis.aiT2History
322granite-4-0-nano-1bOpen weightsIBM · Granite 4.020.5%Independentgroup defaultsPartially comparable-4.77 ptobs. 11 Sept 2026artificialanalysis.aiT2History
323tiny-aya-globalRestricted weightsCohere · Aya20.1%Independentgroup defaultsPartially comparable-5.17 ptobs. 11 Sept 2026artificialanalysis.aiT2History
324Mistral 7BOpen weightsMistral AI · Mistral19.9%Independentgroup defaultsPartially comparable-5.38 ptobs. 11 Sept 2026artificialanalysis.aiT2History
324gemma-3-1bOpen weightsGoogle · Gemma 319.9%Independentgroup defaultsPartially comparable-5.38 ptobs. 11 Sept 2026artificialanalysis.aiT2History
326DeepSeek-R1-0528-Qwen3-8BOpen weightsDeepSeek · Qwen319.9%Independentgroup defaultsPartially comparable-5.45 ptobs. 11 Sept 2026artificialanalysis.aiT2History
327molmo-7b-dOpen weightsAllen Institute for AI · Molmo19.7%Independentgroup defaultsPartially comparable-5.65 ptobs. 11 Sept 2026artificialanalysis.aiT2History
328lfm2-2-6bOpen weightsLiquid AI · LFM2.219.5%Independentgroup defaultsPartially comparable-5.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
329deepseek-r1-distill-llama-8bOpen weightsDeepSeek · Llama17.6%Independentgroup defaultsPartially comparable-7.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
329granite-4-0-h-350mOpen weightsIBM · Granite 4.017.6%Independentgroup defaultsPartially comparable-7.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
331granite-4-0-350mOpen weightsIBM · Granite 4.015.9%Independentgroup defaultsPartially comparable-9.39 ptobs. 11 Sept 2026artificialanalysis.aiT2History
332DeepSeek-R1-Distill-Qwen-1.5BOpen weightsDeepSeek · Qwen13.2%Independentgroup defaultsPartially comparable-12.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
333gemma-3-270mOpen weightsGoogle · Gemma 312.1%Independentgroup defaultsPartially comparable-13.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →