Skip to content
AI Atlas
BenchmarkActivecategory · instruction-followingfamily · ifeval · variant IFBench

IFBench

precise instruction following with novel constraints

data quality51

Updated 2 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
898
Models
334
Current leader
Grok 4.3 83.3%

Score history · NVIDIA-Nemotron-Nano-9B-v2 4 rows

Score history for NVIDIA-Nemotron-Nano-9B-v20%10%20%30%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • NVIDIA-Nemotron-Nano-9B-v2
  • 27.14%aa_slug=nvidia-nemotron-nano-9b-v2 · evaluator=Artificial Analysis · reasoning=off · index_version=4.312 Sept 2026
  • 27.62%aa_slug=nvidia-nemotron-nano-9b-v2-reasoning · evaluator=Artificial Analysis · reasoning=on · index_version=4.312 Sept 2026
  • 27.14%aa_slug=nvidia-nemotron-nano-9b-v2 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 27.62%aa_slug=nvidia-nemotron-nano-9b-v2-reasoning · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · evaluator=Artificial Analysis

6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 83.3%Grok 4.3 xAI Independent11 Sept 2026
  2. 81.3%Grok 4.3 xAI Independent11 Sept 2026
  3. 81.2%Grok 4.20 xAI Independent11 Sept 2026
  4. 81.0%Grok 4.3 xAI Independent11 Sept 2026
  5. 73.5%gemma-4-12B Google Independent11 Sept 2026
  6. 45.9%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 334 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
301apertus-70b-instructOpen weightsSwiss AI Initiative · best of 2 rows25.9%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
302llama-3-1-nemotron-nano-4b-reasoningOpen weightsNVIDIA · Llama 3.1 · best of 2 rows25.5%IndependentreasoningonPartially comparable-0.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
303exaone-4-0-1-2bOpen weightsLG AI Research · EXAONE 4.0 · best of 4 rows25.3%IndependentreasoningoffPartially comparable-0.54 ptobs. 12 Sept 2026artificialanalysis.aiT2History
304magistral-mediumClosedMistral AI · Magistral · best of 2 rows25.1%IndependentreasoningonPartially comparable-0.75 ptobs. 12 Sept 2026artificialanalysis.aiT2History
305Granite 4.0 MicroOpen weightsIBM · Granite 4.0 · best of 2 rows24.8%IndependentreasoningoffPartially comparable-1.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
305magistral-smallOpen weightsMistral AI · Magistral · best of 2 rows24.8%IndependentreasoningonPartially comparable-1.09 ptobs. 12 Sept 2026artificialanalysis.aiT2History
307llama-3-instruct-8bOpen weightsMeta AI · Llama 3 · best of 2 rows24.6%IndependentreasoningoffPartially comparable-1.29 ptobs. 12 Sept 2026artificialanalysis.aiT2History
308olmo-2-7bOpen weightsAllen Institute for AI · OLMo 2 · best of 2 rows24.4%IndependentreasoningoffPartially comparable-1.43 ptobs. 12 Sept 2026artificialanalysis.aiT2History
309phi-3-miniOpen weightsMicrosoft · Phi3 · best of 2 rows23.9%IndependentreasoningoffPartially comparable-1.97 ptobs. 12 Sept 2026artificialanalysis.aiT2History
310ling-mini-2-0Open weightsinclusionAI · best of 2 rows23.6%IndependentreasoningoffPartially comparable-2.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
311Phi 4Open weightsMicrosoft · Phi4 · best of 2 rows23.5%IndependentreasoningoffPartially comparable-2.31 ptobs. 12 Sept 2026artificialanalysis.aiT2History
312Qwen3-0.6BOpen weightsQwen · Qwen3.023.3%Independentgroup defaultsPartially comparable-2.52 ptobs. 11 Sept 2026artificialanalysis.aiT2History
312qwen3-0.6b-instructOpen weightsAlibaba Group · Qwen3.0 · best of 3 rows23.3%IndependentreasoningonPartially comparable-2.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
314DeepSeek-R1-Distill-Qwen-32BOpen weightsDeepSeek · Qwen · best of 2 rows22.9%IndependentreasoningonPartially comparable-2.92 ptobs. 12 Sept 2026artificialanalysis.aiT2History
315Llama-3.2-1BRestricted weightsMeta AI · Llama 3.2 · best of 2 rows22.8%IndependentreasoningoffPartially comparable-3.06 ptobs. 12 Sept 2026artificialanalysis.aiT2History
316granite-3-3-8b-instructOpen weightsIBM · Granite 3.3 · best of 2 rows22.4%IndependentreasoningoffPartially comparable-3.40 ptobs. 12 Sept 2026artificialanalysis.aiT2History
317apertus-8b-instructOpen weightsSwiss AI Initiative · best of 2 rows22.4%IndependentreasoningoffPartially comparable-3.47 ptobs. 12 Sept 2026artificialanalysis.aiT2History
318deepseek-r1-distill-qwen-14bOpen weightsDeepSeek · Qwen · best of 2 rows22.1%IndependentreasoningonPartially comparable-3.74 ptobs. 12 Sept 2026artificialanalysis.aiT2History
319gemma-3n-e2bOpen weightsGoogle · Gemma 3 · best of 2 rows22.0%IndependentreasoningoffPartially comparable-3.81 ptobs. 12 Sept 2026artificialanalysis.aiT2History
320LFM2-1.2BOpen weightsLiquid AI · LFM2.1 · best of 2 rows22.0%IndependentreasoningoffPartially comparable-3.88 ptobs. 12 Sept 2026artificialanalysis.aiT2History
321Qwen3.5-0.8BOpen weightsQwen · Qwen3.5 · best of 4 rows21.6%IndependentreasoningoffPartially comparable-4.29 ptobs. 12 Sept 2026artificialanalysis.aiT2History
322phi-4-miniOpen weightsMicrosoft · Phi4 · best of 2 rows21.1%IndependentreasoningoffPartially comparable-4.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
323granite-4-0-nano-1bOpen weightsIBM · Granite 4.0 · best of 2 rows20.5%IndependentreasoningoffPartially comparable-5.31 ptobs. 12 Sept 2026artificialanalysis.aiT2History
324tiny-aya-globalRestricted weightsCohere · Aya · best of 2 rows20.1%IndependentreasoningoffPartially comparable-5.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
325Mistral 7BOpen weightsMistral AI · Mistral · best of 2 rows19.9%IndependentreasoningoffPartially comparable-5.92 ptobs. 12 Sept 2026artificialanalysis.aiT2History
325gemma-3-1bOpen weightsGoogle · Gemma 3 · best of 2 rows19.9%IndependentreasoningoffPartially comparable-5.92 ptobs. 12 Sept 2026artificialanalysis.aiT2History
327DeepSeek-R1-0528-Qwen3-8BOpen weightsDeepSeek · Qwen3 · best of 2 rows19.9%IndependentreasoningonPartially comparable-5.99 ptobs. 12 Sept 2026artificialanalysis.aiT2History
328molmo-7b-dOpen weightsAllen Institute for AI · Molmo · best of 2 rows19.7%IndependentreasoningoffPartially comparable-6.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
329lfm2-2-6bOpen weightsLiquid AI · LFM2.2 · best of 2 rows19.5%IndependentreasoningoffPartially comparable-6.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
330deepseek-r1-distill-llama-8bOpen weightsDeepSeek · Llama · best of 2 rows17.6%IndependentreasoningonPartially comparable-8.30 ptobs. 12 Sept 2026artificialanalysis.aiT2History
330granite-4-0-h-350mOpen weightsIBM · Granite 4.0 · best of 2 rows17.6%IndependentreasoningoffPartially comparable-8.30 ptobs. 12 Sept 2026artificialanalysis.aiT2History
332granite-4-0-350mOpen weightsIBM · Granite 4.0 · best of 2 rows15.9%IndependentreasoningoffPartially comparable-9.93 ptobs. 12 Sept 2026artificialanalysis.aiT2History
333DeepSeek-R1-Distill-Qwen-1.5BOpen weightsDeepSeek · Qwen · best of 2 rows13.2%IndependentreasoningonPartially comparable-12.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
334gemma-3-270mOpen weightsGoogle · Gemma 3 · best of 2 rows12.1%IndependentreasoningoffPartially comparable-13.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →