Skip to content
AI Atlas
BenchmarkActivecategory · instruction-followingfamily · ifeval · variant IFBench

IFBench

precise instruction following with novel constraints

quality51

Updated 29 min ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
450
Models
333
Current leader
Grok 4.3 83.3%

Score history · DeepSeek V3 0324 1 row

Not enough history to chart — a single observation (41.02% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.

  • 41.02%aa_slug=deepseek-v3-0324 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · evaluator=Artificial Analysis

6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 83.3%Grok 4.3 xAI Independent11 Sept 2026
  2. 81.3%Grok 4.3 xAI Independent11 Sept 2026
  3. 81.2%Grok 4.20 xAI Independent11 Sept 2026
  4. 81.0%Grok 4.3 xAI Independent11 Sept 2026
  5. 73.5%gemma-4-12B Google Independent11 Sept 2026
  6. 45.9%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 333 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
101DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V360.7%Independentgroup defaultsPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
101Qwen3 Next 80B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows60.7%IndependentreasoningonPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
103k2-v2Open weightsMBZUAI Institute of Foundation Models · best of 3 rows60.1%Independentgroup defaultsPartially comparable-0.54 ptobs. 11 Sept 2026artificialanalysis.aiT2History
104diffusiongemma-26b-a4bOpen weightsGoogle59.5%Independentgroup defaultsPartially comparable-1.22 ptobs. 11 Sept 2026artificialanalysis.aiT2History
105Qwen3 VL 32B InstructOpen weightsQwen · Qwen3 · best of 2 rows59.4%IndependentreasoningonPartially comparable-1.29 ptobs. 11 Sept 2026artificialanalysis.aiT2History
106Claude Opus 4.7ClosedAnthropic · Claude · best of 2 rows58.6%Independentgroup defaultsPartially comparable-2.04 ptobs. 11 Sept 2026artificialanalysis.aiT2History
107nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 358.2%Independentgroup defaultsPartially comparable-2.45 ptobs. 11 Sept 2026artificialanalysis.aiT2History
108Claude Opus 4.5ClosedAnthropic · Claude · best of 2 rows58.0%IndependentreasoningonPartially comparable-2.72 ptobs. 11 Sept 2026artificialanalysis.aiT2History
108exaone-4-5-33bOpen weightsLG AI Research · EXAONE 4.558.0%Independentgroup defaultsPartially comparable-2.72 ptobs. 11 Sept 2026artificialanalysis.aiT2History
110solar-open-100b-reasoningOpen weightsUpstage · Solar57.7%Independentgroup defaultsPartially comparable-2.99 ptobs. 11 Sept 2026artificialanalysis.aiT2History
111North Mini Code (free)Open weightsCohere57.5%Independentgroup defaultsPartially comparable-3.13 ptobs. 11 Sept 2026artificialanalysis.aiT2History
112ling-2-6-flashOpen weightsinclusionAI57.4%Independentgroup defaultsPartially comparable-3.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
113Claude Sonnet 4.5ClosedAnthropic · Claude · best of 2 rows57.3%IndependentreasoningonPartially comparable-3.40 ptobs. 11 Sept 2026artificialanalysis.aiT2History
114DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 2 rows57.0%IndependentreasoningonPartially comparable-3.67 ptobs. 11 Sept 2026artificialanalysis.aiT2History
114motif-2-12-7bClosedMotif Technologies57.0%Independentgroup defaultsPartially comparable-3.67 ptobs. 11 Sept 2026artificialanalysis.aiT2History
116ling-2-6-1tOpen weightsinclusionAI56.9%Independentgroup defaultsPartially comparable-3.81 ptobs. 11 Sept 2026artificialanalysis.aiT2History
117Claude Sonnet 4.6ClosedAnthropic · Claude · best of 3 rows56.6%IndependentreasoningadaptivePartially comparable-4.08 ptobs. 11 Sept 2026artificialanalysis.aiT2History
118Qwen3 VL 235B A22B InstructOpen weightsQwen · Qwen3 · best of 2 rows56.5%IndependentreasoningonPartially comparable-4.22 ptobs. 11 Sept 2026artificialanalysis.aiT2History
119Trinity Large ThinkingOpen weightsArcee AI56.3%Independentgroup defaultsPartially comparable-4.42 ptobs. 11 Sept 2026artificialanalysis.aiT2History
120LFM2.5-8B-A1BOpen weightsLiquid AI · LFM2.555.6%Independentgroup defaultsPartially comparable-5.03 ptobs. 11 Sept 2026artificialanalysis.aiT2History
121Claude Opus 4.1ClosedAnthropic · Claude55.4%IndependentreasoningonPartially comparable-5.24 ptobs. 11 Sept 2026artificialanalysis.aiT2History
122gemini-3-flashClosedGoogle · Gemini 355.1%Independentgroup defaultsPartially comparable-5.58 ptobs. 11 Sept 2026artificialanalysis.aiT2History
123Claude Sonnet 4ClosedAnthropic · Claude · best of 2 rows54.7%IndependentreasoningonPartially comparable-5.99 ptobs. 11 Sept 2026artificialanalysis.aiT2History
124tri-21b-think-v0-5Open weightsTrillion Labs54.6%Independentgroup defaultsPartially comparable-6.05 ptobs. 11 Sept 2026artificialanalysis.aiT2History
125falcon-h1r-7bOpen weightsTII UAE · Falcon54.4%Independentgroup defaultsPartially comparable-6.26 ptobs. 11 Sept 2026artificialanalysis.aiT2History
126Claude Haiku 4.5ClosedAnthropic · Claude · best of 2 rows54.3%IndependentreasoningonPartially comparable-6.39 ptobs. 11 Sept 2026artificialanalysis.aiT2History
127DeepSeek V3.2 ExpOpen weightsDeepSeek · DeepSeek-V354.1%Independentgroup defaultsPartially comparable-6.53 ptobs. 11 Sept 2026artificialanalysis.aiT2History
128qwen3-max-thinking-previewClosedAlibaba Group · Qwen353.8%Independentgroup defaultsPartially comparable-6.87 ptobs. 11 Sept 2026artificialanalysis.aiT2History
129Claude Opus 4ClosedAnthropic · Claude53.7%Independentgroup defaultsPartially comparable-6.94 ptobs. 11 Sept 2026artificialanalysis.aiT2History
130grok-4ClosedSpaceXAI · Grok 453.7%Independentgroup defaultsPartially comparable-7.01 ptobs. 11 Sept 2026artificialanalysis.aiT2History
131mimo-v2-omniClosedXiaomi53.5%Independentgroup defaultsPartially comparable-7.14 ptobs. 11 Sept 2026artificialanalysis.aiT2History
132Claude Opus 4.6ClosedAnthropic · Claude · best of 2 rows53.1%IndependentreasoningadaptivePartially comparable-7.55 ptobs. 11 Sept 2026artificialanalysis.aiT2History
133grok-4-1-fastClosedSpaceXAI · Grok 4.1 · best of 2 rows52.7%IndependentreasoningonPartially comparable-7.96 ptobs. 11 Sept 2026artificialanalysis.aiT2History
134gemini-2-5-flash-lite-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 2 rows52.6%IndependentreasoningonPartially comparable-8.09 ptobs. 11 Sept 2026artificialanalysis.aiT2History
135jamba-reasoning-3bOpen weightsAI21 Labs · Jamba52.5%Independentgroup defaultsPartially comparable-8.23 ptobs. 11 Sept 2026artificialanalysis.aiT2History
136gemini-2-5-flash-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 2 rows52.3%IndependentreasoningonPartially comparable-8.37 ptobs. 11 Sept 2026artificialanalysis.aiT2History
137Qwen3.5-4BOpen weightsQwen · Qwen3.5 · best of 2 rows52.0%Independentgroup defaultsPartially comparable-8.71 ptobs. 11 Sept 2026artificialanalysis.aiT2History
138doubao-seed-codeClosedByteDance · Seed51.4%Independentgroup defaultsPartially comparable-9.25 ptobs. 11 Sept 2026artificialanalysis.aiT2History
139Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 2 rows51.2%Independentgroup defaultsPartially comparable-9.46 ptobs. 11 Sept 2026artificialanalysis.aiT2History
140qwen3-5-omni-plusClosedAlibaba Group · Qwen3.551.2%Independentgroup defaultsPartially comparable-9.52 ptobs. 11 Sept 2026artificialanalysis.aiT2History
141qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 2 rows50.7%IndependentreasoningonPartially comparable-10.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
142grok-4-fastClosedSpaceXAI · Grok 4 · best of 2 rows50.5%IndependentreasoningonPartially comparable-10.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
143Gemini 2.5 FlashClosedGoogle · Gemini 2.5 · best of 2 rows50.3%IndependentreasoningonPartially comparable-10.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
144step-3-vl-10bOpen weightsStepFun · Step350.2%Independentgroup defaultsPartially comparable-10.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
145Gemini 2.5 Flash-LiteClosedGoogle · Gemini 2.5 · best of 2 rows49.9%Independentgroup defaultsPartially comparable-10.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
146qwen3-4b-2507-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows49.8%IndependentreasoningonPartially comparable-10.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
147grok-4.20-0309-non-reasoningClosedxAI · Grok49.3%Independentgroup defaultsPartially comparable-11.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
147mi-dm-k-2-5-pro-dec28ClosedKorea Telecom49.3%Independentgroup defaultsPartially comparable-11.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
147minicpm5-1bOpen weightsOpenBMB · best of 2 rows49.3%Independentgroup defaultsPartially comparable-11.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
150olmo-3-32b-thinkOpen weightsAllen Institute for AI · OLMo 349.1%Independentgroup defaultsPartially comparable-11.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
151DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 2 rows49.0%Independentgroup defaultsPartially comparable-11.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
152Gemini 2.5 ProClosedGoogle · Gemini 2.548.7%Independentgroup defaultsPartially comparable-12.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
153Claude 3.7 SonnetClosedAnthropic · Claude · best of 2 rows48.3%IndependentreasoningonPartially comparable-12.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
154Mistral Small 4Open weightsMistral AI · Mistral · best of 2 rows48.2%Independentgroup defaultsPartially comparable-12.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
155qwen3-max-previewClosedAlibaba Group · Qwen348.0%Independentgroup defaultsPartially comparable-12.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
156Hy3Open weightsTencent48.0%IndependentreasoningoffPartially comparable-12.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
157deepseek-v4-flash-0420-non-reasoningOpen weightsDeepSeek · DeepSeek47.2%Independentgroup defaultsPartially comparable-13.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
158tri-21b-think-previewOpen weightsTrillion Labs47.1%Independentgroup defaultsPartially comparable-13.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
159Llama 3.3 70BOpen weightsMeta AI · Llama 3.347.1%Independentgroup defaultsPartially comparable-13.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
160grok-3ClosedSpaceXAI · Grok 346.9%Independentgroup defaultsPartially comparable-13.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
161cogito-v2-1-reasoningOpen weightsDeep Cogito · Cogito46.3%Independentgroup defaultsPartially comparable-14.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
162grok-3-mini-reasoningClosedSpaceXAI · Grok 345.9%Independentgroup defaultsPartially comparable-14.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
162lfm2-24b-a2bOpen weightsLiquid AI · LFM245.9%Independentgroup defaultsPartially comparable-14.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
164deepseek-v4-pro-0424-non-reasoningOpen weightsDeepSeek · DeepSeek45.8%Independentgroup defaultsPartially comparable-14.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
165midm-250-pro-rsnsftClosedKorea Telecom45.6%Independentgroup defaultsPartially comparable-15.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
166Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows45.1%IndependentreasoningonPartially comparable-15.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
167gpt-5-chatgptClosedOpenAI · GPT 545.0%Independentgroup defaultsPartially comparable-15.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
168ring-1tOpen weightsinclusionAI44.6%Independentgroup defaultsPartially comparable-16.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
168ring-2-6-1tOpen weightsinclusionAI44.6%Independentgroup defaultsPartially comparable-16.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
170Magistral Small 1.2Open weightsMistral AI · Magistral44.4%Independentgroup defaultsPartially comparable-16.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
170granite-4.1-30bOpen weightsIBM · Granite 4.144.4%Independentgroup defaultsPartially comparable-16.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
172gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 2 rows44.2%Independentgroup defaultsPartially comparable-16.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
173GLM 4.5Open weightsZ.ai (Zhipu AI) · GLM4.544.1%Independentgroup defaultsPartially comparable-16.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
174LFM2.5-1.2B-InstructOpen weightsLiquid AI · LFM2.543.8%Independentgroup defaultsPartially comparable-16.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
175GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows43.4%IndependentreasoningonPartially comparable-17.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
175qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows43.4%IndependentreasoningonPartially comparable-17.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
177claude-4-opusClosedAnthropic · Claude 443.3%Independentgroup defaultsPartially comparable-17.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
177ring-flash-2-0Open weightsinclusionAI43.3%Independentgroup defaultsPartially comparable-17.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
179deepseek-v3-2-0925Open weightsDeepSeek · DeepSeek43.1%Independentgroup defaultsPartially comparable-17.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
180longcat-flash-liteOpen weightsLongCat43.1%Independentgroup defaultsPartially comparable-17.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
181Llama 4 MaverickOpen weightsMeta AI · Llama 443.0%Independentgroup defaultsPartially comparable-17.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
181Magistral Medium 1.2ClosedMistral AI · Magistral43.0%Independentgroup defaultsPartially comparable-17.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
181gpt-4.1ClosedOpenAI · GPT 4.143.0%Independentgroup defaultsPartially comparable-17.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
184Claude Haiku 3.5ClosedAnthropic · Claude42.8%Independentgroup defaultsPartially comparable-17.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
185jt-35b-flashClosedChina Mobile42.0%Independentgroup defaultsPartially comparable-18.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
186Seed-OSS-36B-InstructOpen weightsByteDance · Seed41.9%Independentgroup defaultsPartially comparable-18.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
187lfm2-5-1-2b-thinkingOpen weightsLiquid AI · LFM2.541.8%Independentgroup defaultsPartially comparable-18.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
188MiniMax-M1-80kOpen weightsMiniMax · MiniMax41.8%Independentgroup defaultsPartially comparable-18.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
189Kimi K2 0905Open weightsMoonshot AI · Kimi41.7%Independentgroup defaultsPartially comparable-19.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
190DeepSeek V3.1Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows41.5%IndependentreasoningonPartially comparable-19.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
190Kimi K2 0711Open weightsMoonshot AI · Kimi41.5%Independentgroup defaultsPartially comparable-19.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
190Olmo-3-7B-ThinkOpen weightsAllen Institute for AI · OLMo 341.5%Independentgroup defaultsPartially comparable-19.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
190qwen3-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows41.5%IndependentreasoningonPartially comparable-19.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
194Grok Build 0.1ClosedxAI · Grok41.4%Independentgroup defaultsPartially comparable-19.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
194ernie-5-0-thinking-previewClosedBaidu · ERNIE 5.041.4%Independentgroup defaultsPartially comparable-19.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
196MiniMax-M1-40kOpen weightsMiniMax · MiniMax41.2%Independentgroup defaultsPartially comparable-19.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
197DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V341.0%Independentgroup defaultsPartially comparable-19.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
198Qwen3 14BOpen weightsQwen · Qwen340.5%Independentgroup defaultsPartially comparable-20.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
199qwen3-coder-480b-a35b-instructOpen weightsAlibaba Group · Qwen3-Coder40.5%Independentgroup defaultsPartially comparable-20.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
200Gemini 2.0 FlashClosedGoogle · Gemini 2.040.2%Independentgroup defaultsPartially comparable-20.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →