Skip to content
AI Atlas
BenchmarkActivecategory · instruction-followingfamily · ifeval · variant IFBench

IFBench

precise instruction following with novel constraints

data quality51

Updated 7 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
898
Models
334
Current leader
Grok 4.3 83.3%

Score history · Qwen3-0.6B 1 row

Not enough history to chart — a single observation (23.33% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.

  • 23.33%aa_slug=qwen3-0.6b-instruct-reasoning · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · evaluator=Artificial Analysis

6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 83.3%Grok 4.3 xAI Independent11 Sept 2026
  2. 81.3%Grok 4.3 xAI Independent11 Sept 2026
  3. 81.2%Grok 4.20 xAI Independent11 Sept 2026
  4. 81.0%Grok 4.3 xAI Independent11 Sept 2026
  5. 73.5%gemma-4-12B Google Independent11 Sept 2026
  6. 45.9%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 334 models · trust independent-evaluator

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
101GLM 5V TurboClosedZ.ai (Zhipu AI) · GLM5 · best of 2 rows61.1%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
102GLM 4.7 FlashOpen weightsZ.ai (Zhipu AI) · GLM4.7 · best of 4 rows60.8%IndependentreasoningonPartially comparable-0.27 ptobs. 12 Sept 2026artificialanalysis.aiT2History
103DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 5 rows60.7%IndependentreasoningonPartially comparable-0.41 ptobs. 12 Sept 2026artificialanalysis.aiT2History
103DeepSeek V3.2Open weightsDeepSeek · DeepSeek-V360.7%Independentgroup defaultsPartially comparable-0.41 ptobs. 11 Sept 2026artificialanalysis.aiT2History
103Qwen3 Next 80B A3B InstructOpen weightsQwen · Qwen3 · best of 4 rows60.7%IndependentreasoningonPartially comparable-0.41 ptobs. 12 Sept 2026artificialanalysis.aiT2History
106k2-v2Open weightsMBZUAI Institute of Foundation Models · best of 6 rows60.1%Independentreasoning_efforthighPartially comparable-0.95 ptobs. 12 Sept 2026artificialanalysis.aiT2History
107diffusiongemma-26b-a4bOpen weightsGoogle · best of 2 rows59.5%IndependentreasoningonPartially comparable-1.63 ptobs. 12 Sept 2026artificialanalysis.aiT2History
108Qwen3 VL 32B InstructOpen weightsQwen · Qwen3 · best of 4 rows59.4%IndependentreasoningonPartially comparable-1.70 ptobs. 12 Sept 2026artificialanalysis.aiT2History
109Claude Opus 4.7ClosedAnthropic · Claude · best of 4 rows58.6%Independentreasoning_effortmaxPartially comparable-2.45 ptobs. 12 Sept 2026artificialanalysis.aiT2History
110nvidia-nemotron-3-nano-4bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows58.2%IndependentreasoningonPartially comparable-2.86 ptobs. 12 Sept 2026artificialanalysis.aiT2History
111Claude Opus 4.5ClosedAnthropic · Claude · best of 4 rows58.0%IndependentreasoningonPartially comparable-3.13 ptobs. 12 Sept 2026artificialanalysis.aiT2History
111exaone-4-5-33bOpen weightsLG AI Research · EXAONE 4.5 · best of 2 rows58.0%IndependentreasoningonPartially comparable-3.13 ptobs. 12 Sept 2026artificialanalysis.aiT2History
113solar-open-100b-reasoningOpen weightsUpstage · Solar · best of 2 rows57.7%IndependentreasoningonPartially comparable-3.40 ptobs. 12 Sept 2026artificialanalysis.aiT2History
114North Mini Code (free)Open weightsCohere · best of 2 rows57.5%IndependentreasoningonPartially comparable-3.54 ptobs. 12 Sept 2026artificialanalysis.aiT2History
115ling-2-6-flashOpen weightsinclusionAI · best of 2 rows57.4%IndependentreasoningoffPartially comparable-3.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
116Claude Sonnet 4.5ClosedAnthropic · Claude · best of 4 rows57.3%IndependentreasoningonPartially comparable-3.81 ptobs. 12 Sept 2026artificialanalysis.aiT2History
117DeepSeek V3.1 TerminusOpen weightsDeepSeek · DeepSeek · best of 4 rows57.0%IndependentreasoningonPartially comparable-4.08 ptobs. 12 Sept 2026artificialanalysis.aiT2History
117motif-2-12-7bClosedMotif Technologies · best of 2 rows57.0%IndependentreasoningonPartially comparable-4.08 ptobs. 12 Sept 2026artificialanalysis.aiT2History
119ling-2-6-1tOpen weightsinclusionAI · best of 2 rows56.9%IndependentreasoningoffPartially comparable-4.22 ptobs. 12 Sept 2026artificialanalysis.aiT2History
120Claude Sonnet 4.6ClosedAnthropic · Claude · best of 5 rows56.6%Independentreasoningadaptivereasoning_effortmaxPartially comparable-4.49 ptobs. 12 Sept 2026artificialanalysis.aiT2History
121Qwen3 VL 235B A22B InstructOpen weightsQwen · Qwen3 · best of 4 rows56.5%IndependentreasoningonPartially comparable-4.63 ptobs. 12 Sept 2026artificialanalysis.aiT2History
122Trinity Large ThinkingOpen weightsArcee AI · best of 2 rows56.3%IndependentreasoningonPartially comparable-4.83 ptobs. 12 Sept 2026artificialanalysis.aiT2History
123LFM2.5-8B-A1BRestricted weightsLiquid AI · LFM2.5 · best of 2 rows55.6%IndependentreasoningonPartially comparable-5.44 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124Claude Opus 4.1ClosedAnthropic · Claude · best of 2 rows55.4%IndependentreasoningonPartially comparable-5.65 ptobs. 12 Sept 2026artificialanalysis.aiT2History
125Claude Sonnet 4ClosedAnthropic · Claude · best of 4 rows54.7%IndependentreasoningonPartially comparable-6.40 ptobs. 12 Sept 2026artificialanalysis.aiT2History
126tri-21b-think-v0-5Open weightsTrillion Labs · best of 2 rows54.6%IndependentreasoningonPartially comparable-6.46 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127falcon-h1r-7bOpen weightsTII UAE · Falcon · best of 2 rows54.4%IndependentreasoningonPartially comparable-6.67 ptobs. 12 Sept 2026artificialanalysis.aiT2History
128Claude Haiku 4.5ClosedAnthropic · Claude · best of 4 rows54.3%IndependentreasoningonPartially comparable-6.80 ptobs. 12 Sept 2026artificialanalysis.aiT2History
129DeepSeek V3.2 ExpOpen weightsDeepSeek · DeepSeek-V3 · best of 3 rows54.1%IndependentreasoningonPartially comparable-6.94 ptobs. 12 Sept 2026artificialanalysis.aiT2History
130qwen3-max-thinking-previewClosedAlibaba Group · Qwen3 · best of 2 rows53.8%IndependentreasoningonPartially comparable-7.28 ptobs. 12 Sept 2026artificialanalysis.aiT2History
131Claude Opus 4ClosedAnthropic · Claude53.7%Independentgroup defaultsPartially comparable-7.35 ptobs. 11 Sept 2026artificialanalysis.aiT2History
131claude-4-opusClosedAnthropic · Claude 4 · best of 3 rows53.7%IndependentreasoningonPartially comparable-7.35 ptobs. 12 Sept 2026artificialanalysis.aiT2History
133grok-4ClosedSpaceXAI · Grok 4 · best of 2 rows53.7%IndependentreasoningonPartially comparable-7.42 ptobs. 12 Sept 2026artificialanalysis.aiT2History
134mimo-v2-omniClosedXiaomi · best of 2 rows53.5%IndependentreasoningonPartially comparable-7.55 ptobs. 12 Sept 2026artificialanalysis.aiT2History
135Claude Opus 4.6ClosedAnthropic · Claude · best of 4 rows53.1%Independentreasoningadaptivereasoning_effortmaxPartially comparable-7.96 ptobs. 12 Sept 2026artificialanalysis.aiT2History
136grok-4-1-fastClosedSpaceXAI · Grok 4.1 · best of 4 rows52.7%IndependentreasoningonPartially comparable-8.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
137Gemini 2.5 Flash-LiteClosedGoogle · Gemini 2.5 · best of 5 rows52.6%IndependentreasoningonPartially comparable-8.50 ptobs. 12 Sept 2026artificialanalysis.aiT2History
137gemini-2-5-flash-lite-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 3 rows52.6%IndependentreasoningonPartially comparable-8.50 ptobs. 11 Sept 2026artificialanalysis.aiT2History
139jamba-reasoning-3bOpen weightsAI21 Labs · Jamba · best of 2 rows52.5%IndependentreasoningonPartially comparable-8.64 ptobs. 12 Sept 2026artificialanalysis.aiT2History
140gemini-2-5-flash-preview-09-2025ClosedGoogle · Gemini 2.5 · best of 6 rows52.3%IndependentreasoningonPartially comparable-8.78 ptobs. 12 Sept 2026artificialanalysis.aiT2History
141Qwen3.5-4BOpen weightsQwen · Qwen3.5 · best of 4 rows52.0%IndependentreasoningonPartially comparable-9.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
142doubao-seed-codeClosedByteDance · Seed · best of 2 rows51.4%IndependentreasoningonPartially comparable-9.66 ptobs. 12 Sept 2026artificialanalysis.aiT2History
143Qwen3 235B A22B Instruct 2507Open weightsQwen · Qwen3 · best of 4 rows51.2%IndependentreasoningonPartially comparable-9.87 ptobs. 12 Sept 2026artificialanalysis.aiT2History
144qwen3-5-omni-plusClosedAlibaba Group · Qwen3.5 · best of 2 rows51.2%IndependentreasoningoffPartially comparable-9.93 ptobs. 12 Sept 2026artificialanalysis.aiT2History
145qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 4 rows50.7%IndependentreasoningonPartially comparable-10.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
146grok-4-fastClosedSpaceXAI · Grok 4 · best of 4 rows50.5%IndependentreasoningonPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
147Gemini 2.5 FlashClosedGoogle · Gemini 2.5 · best of 2 rows50.3%IndependentreasoningonPartially comparable-10.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
148step-3-vl-10bOpen weightsStepFun · Step3 · best of 2 rows50.2%IndependentreasoningonPartially comparable-10.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
149qwen3-4b-2507-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows49.8%IndependentreasoningonPartially comparable-11.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
150grok-4.20-0309-non-reasoningClosedxAI · Grok49.3%Independentgroup defaultsPartially comparable-11.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
150mi-dm-k-2-5-pro-dec28ClosedKorea Telecom · best of 2 rows49.3%IndependentreasoningonPartially comparable-11.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
150minicpm5-1bOpen weightsOpenBMB · best of 4 rows49.3%IndependentreasoningonPartially comparable-11.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
153olmo-3-32b-thinkOpen weightsAllen Institute for AI · OLMo 3 · best of 2 rows49.1%IndependentreasoningonPartially comparable-12.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
154Gemini 2.5 ProClosedGoogle · Gemini 2.5 · best of 2 rows48.7%IndependentreasoningonPartially comparable-12.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
155Claude 3.7 SonnetClosedAnthropic · Claude · best of 4 rows48.3%IndependentreasoningonPartially comparable-12.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
156Mistral Small 4Open weightsMistral AI · Mistral · best of 4 rows48.2%IndependentreasoningonPartially comparable-12.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
157qwen3-max-previewClosedAlibaba Group · Qwen3 · best of 2 rows48.0%IndependentreasoningoffPartially comparable-13.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
158deepseek-v4-flash-0420-non-reasoningOpen weightsDeepSeek · DeepSeek47.2%Independentgroup defaultsPartially comparable-13.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
159tri-21b-think-previewOpen weightsTrillion Labs · best of 2 rows47.1%IndependentreasoningonPartially comparable-14.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
160Llama 3.3 70BOpen weightsMeta AI · Llama 3.3 · best of 2 rows47.1%IndependentreasoningoffPartially comparable-14.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
161grok-3ClosedSpaceXAI · Grok 3 · best of 2 rows46.9%IndependentreasoningoffPartially comparable-14.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
162cogito-v2-1-reasoningOpen weightsDeep Cogito · Cogito · best of 2 rows46.3%IndependentreasoningonPartially comparable-14.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
163grok-3-mini-reasoningClosedSpaceXAI · Grok 3 · best of 2 rows45.9%Independentreasoningonreasoning_efforthighPartially comparable-15.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
163lfm2-24b-a2bOpen weightsLiquid AI · LFM2 · best of 2 rows45.9%IndependentreasoningoffPartially comparable-15.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
165deepseek-v4-pro-0424-non-reasoningOpen weightsDeepSeek · DeepSeek45.8%Independentgroup defaultsPartially comparable-15.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
166midm-250-pro-rsnsftClosedKorea Telecom · best of 2 rows45.6%IndependentreasoningonPartially comparable-15.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
167Qwen3 VL 30B A3B InstructOpen weightsQwen · Qwen3 · best of 4 rows45.1%IndependentreasoningonPartially comparable-16.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
168gpt-5-chatgptClosedOpenAI · GPT 5 · best of 2 rows45.0%IndependentreasoningoffPartially comparable-16.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
169ring-1tOpen weightsinclusionAI · best of 2 rows44.6%IndependentreasoningonPartially comparable-16.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
169ring-2-6-1tOpen weightsinclusionAI · best of 2 rows44.6%IndependentreasoningonPartially comparable-16.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
171Magistral Small 1.2Open weightsMistral AI · Magistral · best of 2 rows44.4%IndependentreasoningonPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
171granite-4.1-30bOpen weightsIBM · Granite 4.1 · best of 2 rows44.4%IndependentreasoningoffPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
173gemma-4-E4BOpen weightsGoogle · Gemma 4 · best of 4 rows44.2%IndependentreasoningonPartially comparable-16.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
174GLM 4.5Open weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows44.1%IndependentreasoningonPartially comparable-17.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
175LFM2.5-1.2B-InstructRestricted weightsLiquid AI · LFM2.5 · best of 2 rows43.8%IndependentreasoningoffPartially comparable-17.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176GLM 4.6Open weightsZ.ai (Zhipu AI) · GLM4.6 · best of 4 rows43.4%IndependentreasoningonPartially comparable-17.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
176qwen3-omni-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows43.4%IndependentreasoningonPartially comparable-17.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
178ring-flash-2-0Open weightsinclusionAI · best of 2 rows43.3%IndependentreasoningonPartially comparable-17.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
179deepseek-v3-2-0925Open weightsDeepSeek · DeepSeek43.1%Independentgroup defaultsPartially comparable-18.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
180longcat-flash-liteOpen weightsLongCat · best of 2 rows43.1%IndependentreasoningoffPartially comparable-18.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
181Llama 4 MaverickOpen weightsMeta AI · Llama 4 · best of 2 rows43.0%IndependentreasoningoffPartially comparable-18.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
181Magistral Medium 1.2ClosedMistral AI · Magistral · best of 2 rows43.0%IndependentreasoningonPartially comparable-18.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
181gpt-4.1ClosedOpenAI · GPT 4.1 · best of 2 rows43.0%IndependentreasoningoffPartially comparable-18.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
184Claude Haiku 3.5ClosedAnthropic · Claude · best of 2 rows42.8%IndependentreasoningoffPartially comparable-18.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
185jt-35b-flashClosedChina Mobile · best of 2 rows42.0%IndependentreasoningoffPartially comparable-19.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
186Seed-OSS-36B-InstructOpen weightsByteDance · Seed · best of 2 rows41.9%IndependentreasoningonPartially comparable-19.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
187lfm2-5-1-2b-thinkingOpen weightsLiquid AI · LFM2.5 · best of 2 rows41.8%IndependentreasoningonPartially comparable-19.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
188MiniMax-M1-80kOpen weightsMiniMax · MiniMax · best of 2 rows41.8%IndependentreasoningonPartially comparable-19.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
189Kimi K2 0905Open weightsMoonshot AI · Kimi · best of 2 rows41.7%IndependentreasoningoffPartially comparable-19.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
190DeepSeek V3.1Open weightsDeepSeek · DeepSeek-V3 · best of 4 rows41.5%IndependentreasoningonPartially comparable-19.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
190Kimi K2 0711Open weightsMoonshot AI · Kimi · best of 2 rows41.5%IndependentreasoningoffPartially comparable-19.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
190Olmo-3-7B-ThinkOpen weightsAllen Institute for AI · OLMo 3 · best of 2 rows41.5%IndependentreasoningonPartially comparable-19.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
190qwen3-30b-a3b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows41.5%IndependentreasoningonPartially comparable-19.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
194Grok Build 0.1ClosedxAI · Grok · best of 2 rows41.4%IndependentreasoningonPartially comparable-19.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
194ernie-5-0-thinking-previewClosedBaidu · ERNIE 5.0 · best of 2 rows41.4%IndependentreasoningonPartially comparable-19.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
196MiniMax-M1-40kOpen weightsMiniMax · MiniMax · best of 2 rows41.2%IndependentreasoningonPartially comparable-19.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
197DeepSeek V3 0324Open weightsDeepSeek · DeepSeek-V3 · best of 2 rows41.0%IndependentreasoningoffPartially comparable-20.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
198Qwen3 14BOpen weightsQwen · Qwen340.5%Independentgroup defaultsPartially comparable-20.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
198qwen3-14b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows40.5%IndependentreasoningonPartially comparable-20.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
200qwen3-coder-480b-a35b-instructOpen weightsAlibaba Group · Qwen3-Coder · best of 2 rows40.5%IndependentreasoningoffPartially comparable-20.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →