Skip to content
AI Atlas
BenchmarkActivecategory · instruction-followingfamily · ifeval · variant IFBench

IFBench

precise instruction following with novel constraints

quality51

Updated 27 min ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
450
Models
333
Current leader
Grok 4.3 83.3%

Score history · Qwen3 Coder 30B A3B Instruct 1 row

Not enough history to chart — a single observation (32.65% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.

  • 32.65%aa_slug=qwen3-coder-30b-a3b-instruct · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · evaluator=Artificial Analysis

6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 83.3%Grok 4.3 xAI Independent11 Sept 2026
  2. 81.3%Grok 4.3 xAI Independent11 Sept 2026
  3. 81.2%Grok 4.20 xAI Independent11 Sept 2026
  4. 81.0%Grok 4.3 xAI Independent11 Sept 2026
  5. 73.5%gemma-4-12B Google Independent11 Sept 2026
  6. 45.9%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 333 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
201Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 2 rows39.9%IndependentreasoningonPartially comparable0.00 ptobs. 11 Sept 2026artificialanalysis.aiT2History
202Mistral Medium 3.1ClosedMistral AI · Mistral39.8%Independentgroup defaultsPartially comparable-0.06 ptobs. 11 Sept 2026artificialanalysis.aiT2History
203deepseek-r1Open weightsDeepSeek · DeepSeek-R139.6%Independentgroup defaultsPartially comparable-0.27 ptobs. 11 Sept 2026artificialanalysis.aiT2History
204Llama 4 ScoutOpen weightsMeta AI · Llama 439.5%Independentgroup defaultsPartially comparable-0.34 ptobs. 11 Sept 2026artificialanalysis.aiT2History
205llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.3 · best of 2 rows39.5%Independentgroup defaultsPartially comparable-0.40 ptobs. 11 Sept 2026artificialanalysis.aiT2History
206Mistral Medium 3ClosedMistral AI · Mistral39.3%Independentgroup defaultsPartially comparable-0.61 ptobs. 11 Sept 2026artificialanalysis.aiT2History
207ernie-4-5-300b-a47bOpen weightsBaidu · ERNIE 4.539.1%Independentgroup defaultsPartially comparable-0.74 ptobs. 11 Sept 2026artificialanalysis.aiT2History
208Llama-3.1-405BRestricted weightsMeta AI · Llama 3.139.0%Independentgroup defaultsPartially comparable-0.81 ptobs. 11 Sept 2026artificialanalysis.aiT2History
209deepseek-r1-0120Open weightsDeepSeek · DeepSeek39.0%Independentgroup defaultsPartially comparable-0.88 ptobs. 11 Sept 2026artificialanalysis.aiT2History
210QwQ-32BOpen weightsAlibaba Group · Qwen38.8%Independentgroup defaultsPartially comparable-1.08 ptobs. 11 Sept 2026artificialanalysis.aiT2History
211qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 2 rows38.7%IndependentreasoningonPartially comparable-1.15 ptobs. 11 Sept 2026artificialanalysis.aiT2History
212granite-4.1-8bOpen weightsIBM · Granite 4.138.6%Independentgroup defaultsPartially comparable-1.22 ptobs. 11 Sept 2026artificialanalysis.aiT2History
213gpt-4.1-miniClosedOpenAI · GPT 4.138.3%Independentgroup defaultsPartially comparable-1.56 ptobs. 11 Sept 2026artificialanalysis.aiT2History
214llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.138.2%Independentgroup defaultsPartially comparable-1.70 ptobs. 11 Sept 2026artificialanalysis.aiT2History
215Devstral 2Open weightsMistral AI · Devstral 238.1%Independentgroup defaultsPartially comparable-1.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
215nova-proClosedAmazon Web Services · Nova38.1%Independentgroup defaultsPartially comparable-1.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
215olmo-2-32bOpen weightsAllen Institute for AI · OLMo 238.1%Independentgroup defaultsPartially comparable-1.76 ptobs. 11 Sept 2026artificialanalysis.aiT2History
218gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 2 rows38.0%Independentgroup defaultsPartially comparable-1.83 ptobs. 11 Sept 2026artificialanalysis.aiT2History
219qwen3-5-omni-flashClosedAlibaba Group · Qwen3.538.0%Independentgroup defaultsPartially comparable-1.90 ptobs. 11 Sept 2026artificialanalysis.aiT2History
220hyperclova-x-seed-think-32bOpen weightsNaver · Seed37.9%Independentgroup defaultsPartially comparable-1.97 ptobs. 11 Sept 2026artificialanalysis.aiT2History
221GLM 4.5 AirOpen weightsZ.ai (Zhipu AI) · GLM4.537.5%Independentgroup defaultsPartially comparable-2.31 ptobs. 11 Sept 2026artificialanalysis.aiT2History
222llama-3-instruct-70bOpen weightsMeta AI · Llama 337.1%Independentgroup defaultsPartially comparable-2.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
222solar-pro-2ClosedUpstage · Solar · best of 2 rows37.1%IndependentreasoningonPartially comparable-2.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
224llama-nemotron-super-49b-v1-5Open weightsNVIDIA · Llama · best of 2 rows37.0%IndependentreasoningonPartially comparable-2.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
225Qwen2.5 72B InstructOpen weightsQwen · Qwen2.536.9%Independentgroup defaultsPartially comparable-2.99 ptobs. 11 Sept 2026artificialanalysis.aiT2History
226Gemma 3 12BOpen weightsGoogle · Gemma 336.7%Independentgroup defaultsPartially comparable-3.13 ptobs. 11 Sept 2026artificialanalysis.aiT2History
227jt-miniClosedChina Mobile36.7%Independentgroup defaultsPartially comparable-3.19 ptobs. 11 Sept 2026artificialanalysis.aiT2History
228Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 2 rows36.6%IndependentreasoningonPartially comparable-3.26 ptobs. 11 Sept 2026artificialanalysis.aiT2History
229Command AOpen weightsCohere · Command36.5%Independentgroup defaultsPartially comparable-3.40 ptobs. 11 Sept 2026artificialanalysis.aiT2History
230Qwen3 32BOpen weightsQwen · Qwen336.3%Independentgroup defaultsPartially comparable-3.53 ptobs. 11 Sept 2026artificialanalysis.aiT2History
230exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 2 rows36.3%IndependentreasoningonPartially comparable-3.53 ptobs. 11 Sept 2026artificialanalysis.aiT2History
232Mistral Large 3Open weightsMistral AI · Mistral36.2%Independentgroup defaultsPartially comparable-3.67 ptobs. 11 Sept 2026artificialanalysis.aiT2History
232nova-premierClosedAmazon Web Services · Nova36.2%Independentgroup defaultsPartially comparable-3.67 ptobs. 11 Sept 2026artificialanalysis.aiT2History
234Claude 3 HaikuClosedAnthropic · Claude36.1%Independentgroup defaultsPartially comparable-3.74 ptobs. 11 Sept 2026artificialanalysis.aiT2History
235GPT-4o (2024-08-06)ClosedOpenAI · GPT 436.0%Independentgroup defaultsPartially comparable-3.87 ptobs. 11 Sept 2026artificialanalysis.aiT2History
236nanbeige4-1-3bOpen weightsNanbeige35.4%Independentgroup defaultsPartially comparable-4.42 ptobs. 11 Sept 2026artificialanalysis.aiT2History
237Qwen3 Coder NextOpen weightsQwen · Qwen335.2%Independentgroup defaultsPartially comparable-4.62 ptobs. 11 Sept 2026artificialanalysis.aiT2History
238jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.735.2%Independentgroup defaultsPartially comparable-4.69 ptobs. 11 Sept 2026artificialanalysis.aiT2History
239hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.134.8%Independentgroup defaultsPartially comparable-5.10 ptobs. 11 Sept 2026artificialanalysis.aiT2History
239ling-1tOpen weightsinclusionAI34.8%Independentgroup defaultsPartially comparable-5.10 ptobs. 11 Sept 2026artificialanalysis.aiT2History
241devstral-smallOpen weightsMistral AI · Devstral34.6%Independentgroup defaultsPartially comparable-5.30 ptobs. 11 Sept 2026artificialanalysis.aiT2History
242Pixtral LargeOpen weightsMistral AI · Pixtral34.5%Independentgroup defaultsPartially comparable-5.37 ptobs. 11 Sept 2026artificialanalysis.aiT2History
243Llama-3.1-70BOpen weightsMeta AI · Llama 3.134.4%Independentgroup defaultsPartially comparable-5.44 ptobs. 11 Sept 2026artificialanalysis.aiT2History
244ling-flash-2-0Open weightsinclusionAI34.4%Independentgroup defaultsPartially comparable-5.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
244sarvam-105bOpen weightsSarvam34.4%Independentgroup defaultsPartially comparable-5.51 ptobs. 11 Sept 2026artificialanalysis.aiT2History
246gpt-4oClosedOpenAI · GPT 434.3%Independentgroup defaultsPartially comparable-5.57 ptobs. 11 Sept 2026artificialanalysis.aiT2History
247GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows34.2%Independentgroup defaultsPartially comparable-5.64 ptobs. 11 Sept 2026artificialanalysis.aiT2History
248nova-liteClosedAmazon Web Services · Nova34.1%Independentgroup defaultsPartially comparable-5.71 ptobs. 11 Sept 2026artificialanalysis.aiT2History
249intellect-3Open weightsPrime Intellect34.0%Independentgroup defaultsPartially comparable-5.85 ptobs. 11 Sept 2026artificialanalysis.aiT2History
250granite-4.1-3bOpen weightsIBM · Granite 4.133.7%Independentgroup defaultsPartially comparable-6.19 ptobs. 11 Sept 2026artificialanalysis.aiT2History
251Mistral Small 3.2Open weightsMistral AI · Mistral33.5%Independentgroup defaultsPartially comparable-6.39 ptobs. 11 Sept 2026artificialanalysis.aiT2History
251Qwen3 8BOpen weightsQwen · Qwen333.5%Independentgroup defaultsPartially comparable-6.39 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253gpt-4ClosedOpenAI · GPT 433.2%Independentgroup defaultsPartially comparable-6.66 ptobs. 11 Sept 2026artificialanalysis.aiT2History
254LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.533.1%Independentgroup defaultsPartially comparable-6.73 ptobs. 11 Sept 2026artificialanalysis.aiT2History
255Olmo-3-7B-InstructOpen weightsAllen Institute for AI · OLMo 332.8%Independentgroup defaultsPartially comparable-7.07 ptobs. 11 Sept 2026artificialanalysis.aiT2History
256Hermes 4 405BOpen weightsNous Research · Hermes 432.7%Independentgroup defaultsPartially comparable-7.14 ptobs. 11 Sept 2026artificialanalysis.aiT2History
257Qwen3 Coder 30B A3B InstructOpen weightsQwen · Qwen332.6%Independentgroup defaultsPartially comparable-7.21 ptobs. 11 Sept 2026artificialanalysis.aiT2History
258Qwen3-4BOpen weightsQwen · Qwen332.5%Independentgroup defaultsPartially comparable-7.34 ptobs. 11 Sept 2026artificialanalysis.aiT2History
259Ministral 3 14BOpen weightsMistral AI · Ministral 332.0%Independentgroup defaultsPartially comparable-7.82 ptobs. 11 Sept 2026artificialanalysis.aiT2History
259gpt-4.1-nanoClosedOpenAI · GPT 4.132.0%Independentgroup defaultsPartially comparable-7.82 ptobs. 11 Sept 2026artificialanalysis.aiT2History
261nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 2 rows31.9%IndependentreasoningonPartially comparable-7.96 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262Gemma 3 27BOpen weightsGoogle · Gemma 331.8%Independentgroup defaultsPartially comparable-8.02 ptobs. 11 Sept 2026artificialanalysis.aiT2History
262sarvam-m-reasoningOpen weightsSarvam31.8%Independentgroup defaultsPartially comparable-8.02 ptobs. 11 Sept 2026artificialanalysis.aiT2History
264Devstral Small 1.0Open weightsMistral AI · Devstral31.6%Independentgroup defaultsPartially comparable-8.23 ptobs. 11 Sept 2026artificialanalysis.aiT2History
264Mistral Large 2.0Open weightsMistral AI · Mistral31.6%Independentgroup defaultsPartially comparable-8.23 ptobs. 11 Sept 2026artificialanalysis.aiT2History
266Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 2 rows31.5%Independentgroup defaultsPartially comparable-8.36 ptobs. 11 Sept 2026artificialanalysis.aiT2History
266granite-4-0-h-smallOpen weightsIBM · Granite 4.031.5%Independentgroup defaultsPartially comparable-8.36 ptobs. 11 Sept 2026artificialanalysis.aiT2History
266qwen3-32b-instructOpen weightsAlibaba Group · Qwen331.5%Independentgroup defaultsPartially comparable-8.36 ptobs. 11 Sept 2026artificialanalysis.aiT2History
269jamba-1-7-miniOpen weightsAI21 Labs · Jamba 1.731.4%Independentgroup defaultsPartially comparable-8.50 ptobs. 11 Sept 2026artificialanalysis.aiT2History
270Hermes-4-70BRestricted weightsNous Research · Hermes 431.3%Independentgroup defaultsPartially comparable-8.57 ptobs. 11 Sept 2026artificialanalysis.aiT2History
271mistral-large-2Open weightsMistral AI · Mistral31.2%Independentgroup defaultsPartially comparable-8.64 ptobs. 11 Sept 2026artificialanalysis.aiT2History
272Devstral Small 2Open weightsMistral AI · Devstral31.2%Independentgroup defaultsPartially comparable-8.70 ptobs. 11 Sept 2026artificialanalysis.aiT2History
273gpt-4o-miniClosedOpenAI · GPT 430.9%Independentgroup defaultsPartially comparable-8.91 ptobs. 11 Sept 2026artificialanalysis.aiT2History
274llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.130.8%Independentgroup defaultsPartially comparable-9.11 ptobs. 11 Sept 2026artificialanalysis.aiT2History
275Reka Flash 3Open weightsrekaai30.4%Independentgroup defaultsPartially comparable-9.45 ptobs. 11 Sept 2026artificialanalysis.aiT2History
275llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.230.4%Independentgroup defaultsPartially comparable-9.45 ptobs. 11 Sept 2026artificialanalysis.aiT2History
277GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 2 rows30.1%Independentgroup defaultsPartially comparable-9.79 ptobs. 11 Sept 2026artificialanalysis.aiT2History
278Mistral Small 3.1Open weightsMistral AI · Mistral29.9%Independentgroup defaultsPartially comparable-9.93 ptobs. 11 Sept 2026artificialanalysis.aiT2History
278devstral-mediumClosedMistral AI · Devstral29.9%Independentgroup defaultsPartially comparable-9.93 ptobs. 11 Sept 2026artificialanalysis.aiT2History
280nova-microClosedAmazon Web Services · Nova29.4%Independentgroup defaultsPartially comparable-10.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
281Ministral 3 8BOpen weightsMistral AI · Ministral 329.1%Independentgroup defaultsPartially comparable-10.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
282hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.129.0%Independentgroup defaultsPartially comparable-10.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
283Llama 3.1 8BRestricted weightsMeta AI · Llama 3.128.6%Independentgroup defaultsPartially comparable-11.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
283qwen3-8b-instructOpen weightsAlibaba Group · Qwen328.6%Independentgroup defaultsPartially comparable-11.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
285Gemma 3 4BRestricted weightsGoogle · Gemma 328.3%Independentgroup defaultsPartially comparable-11.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
286Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AI · Kimi28.1%Independentgroup defaultsPartially comparable-11.8 ptobs. 11 Sept 2026artificialanalysis.aiT2History
287gemma-3n-e4bOpen weightsGoogle · Gemma 327.9%Independentgroup defaultsPartially comparable-12.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
288NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 2 rows27.6%Independentgroup defaultsPartially comparable-12.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
289R1 Distill Llama 70BOpen weightsDeepSeek · Llama27.6%Independentgroup defaultsPartially comparable-12.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History
290Molmo2-8BOpen weightsAllen Institute for AI26.9%Independentgroup defaultsPartially comparable-12.9 ptobs. 11 Sept 2026artificialanalysis.aiT2History
291qwen3-1.7b-instructOpen weightsAlibaba Group · Qwen3.1 · best of 2 rows26.9%IndependentreasoningonPartially comparable-13.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
292Ministral 3 3BOpen weightsMistral AI · Ministral 326.8%Independentgroup defaultsPartially comparable-13.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
293minicpm-v4-6-1-3bOpen weightsOpenBMB26.7%Independentgroup defaultsPartially comparable-13.1 ptobs. 11 Sept 2026artificialanalysis.aiT2History
294sarvam-30bOpen weightsSarvam26.5%Independentgroup defaultsPartially comparable-13.4 ptobs. 11 Sept 2026artificialanalysis.aiT2History
295Mistral Small 3Open weightsMistral AI · Mistral26.4%Independentgroup defaultsPartially comparable-13.5 ptobs. 11 Sept 2026artificialanalysis.aiT2History
296lfm2-8b-a1bOpen weightsLiquid AI · LFM226.3%Independentgroup defaultsPartially comparable-13.6 ptobs. 11 Sept 2026artificialanalysis.aiT2History
297Llama-3.2-3BRestricted weightsMeta AI · Llama 3.226.2%Independentgroup defaultsPartially comparable-13.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
297granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.026.2%Independentgroup defaultsPartially comparable-13.7 ptobs. 11 Sept 2026artificialanalysis.aiT2History
299apertus-70b-instructOpen weightsSwiss AI Initiative25.9%Independentgroup defaultsPartially comparable-14.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
300llama-3-1-nemotron-nano-4b-reasoningOpen weightsNVIDIA · Llama 3.125.5%Independentgroup defaultsPartially comparable-14.3 ptobs. 11 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →