Skip to content
AI Atlas
BenchmarkActivecategory · instruction-followingfamily · ifeval · variant IFBench

IFBench

precise instruction following with novel constraints

data quality51

Updated 2 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
898
Models
334
Current leader
Grok 4.3 83.3%

Score history · Gemini 3.1 Pro Preview 2 rows

Score history for Gemini 3.1 Pro Preview0%20%40%60%80%100%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • Gemini 3.1 Pro Preview
  • 77.14%aa_slug=gemini-3-1-pro-preview · evaluator=Artificial Analysis · reasoning=on · index_version=4.312 Sept 2026
  • 77.14%aa_slug=gemini-3-1-pro-preview · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · evaluator=Artificial Analysis

6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 83.3%Grok 4.3 xAI Independent11 Sept 2026
  2. 81.3%Grok 4.3 xAI Independent11 Sept 2026
  3. 81.2%Grok 4.20 xAI Independent11 Sept 2026
  4. 81.0%Grok 4.3 xAI Independent11 Sept 2026
  5. 73.5%gemma-4-12B Google Independent11 Sept 2026
  6. 45.9%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 334 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
1Grok 4.3ClosedxAI · Grok · best of 8 rows83.3%Independentreasoning_effortmediumleaderobs. 12 Sept 2026artificialanalysis.aiT2History
2grok-4-20-0309ClosedSpaceXAI · Grok 4.20 · best of 4 rows82.9%IndependentreasoningonPartially comparable-0.40 ptobs. 12 Sept 2026artificialanalysis.aiT2History
3MiniMax-M3Open weightsMiniMax · MiniMax · best of 2 rows82.9%IndependentreasoningonPartially comparable-0.47 ptobs. 12 Sept 2026artificialanalysis.aiT2History
4Nemotron 3 UltraOpen weightsNVIDIA · Nemotron 3 · best of 2 rows81.4%IndependentreasoningonPartially comparable-1.97 ptobs. 12 Sept 2026artificialanalysis.aiT2History
5Grok 4.20ClosedxAI · Grok · best of 3 rows81.2%IndependentreasoningonPartially comparable-2.11 ptobs. 12 Sept 2026artificialanalysis.aiT2History
6Qwen3.7 MaxClosedQwen · Qwen3.7 · best of 2 rows80.5%IndependentreasoningonPartially comparable-2.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
7nemotron-cascade-2-30b-a3bOpen weightsNVIDIA · Nemotron · best of 2 rows80.4%IndependentreasoningonPartially comparable-2.92 ptobs. 12 Sept 2026artificialanalysis.aiT2History
8MiMo-V2.5-ProOpen weightsXiaomi · best of 4 rows79.9%IndependentreasoningonPartially comparable-3.47 ptobs. 12 Sept 2026artificialanalysis.aiT2History
9nova-2-0-proClosedAmazon Web Services · Nova 2.0 · best of 6 rows79.6%Independentreasoningonreasoning_effortlowPartially comparable-3.74 ptobs. 12 Sept 2026artificialanalysis.aiT2History
10deepseek-v4-flash-0420Open weightsDeepSeek · DeepSeek · best of 4 rows79.2%Independentreasoning_effortmaxPartially comparable-4.15 ptobs. 12 Sept 2026artificialanalysis.aiT2History
11Qwen3.5 397B A17BOpen weightsQwen · Qwen3.5 · best of 4 rows78.8%IndependentreasoningonPartially comparable-4.55 ptobs. 12 Sept 2026artificialanalysis.aiT2History
12Gemini 3 Flash PreviewClosedGoogle · Gemini 378.0%Independentgroup defaultsPartially comparable-5.37 ptobs. 11 Sept 2026artificialanalysis.aiT2History
12Qwen 3.7 PlusClosedQwen · Qwen3.7 · best of 2 rows78.0%IndependentreasoningonPartially comparable-5.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
12gemini-3-flashClosedGoogle · Gemini 3 · best of 3 rows78.0%IndependentreasoningonPartially comparable-5.37 ptobs. 12 Sept 2026artificialanalysis.aiT2History
15GPT-5.2-CodexClosedOpenAI · GPT 5.2 · best of 2 rows77.6%Independentreasoning_effortxhighPartially comparable-5.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
16Gemini 3.1 Flash-Lite PreviewClosedGoogle · Gemini 3.1 · best of 2 rows77.2%IndependentreasoningonPartially comparable-6.12 ptobs. 12 Sept 2026artificialanalysis.aiT2History
17Gemini 3.1 Pro PreviewClosedGoogle · Gemini 3.1 · best of 2 rows77.1%IndependentreasoningonPartially comparable-6.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
18Qwen3.6 Max PreviewClosedQwen · Qwen3.6 · best of 2 rows76.6%IndependentreasoningonPartially comparable-6.73 ptobs. 12 Sept 2026artificialanalysis.aiT2History
19deepseek-v4-pro-0424Open weightsDeepSeek · DeepSeek · best of 4 rows76.5%Independentreasoning_effortmaxPartially comparable-6.87 ptobs. 12 Sept 2026artificialanalysis.aiT2History
20Gemini 3.5 FlashClosedGoogle · Gemini 3.5 · best of 6 rows76.3%IndependentreasoningonPartially comparable-7.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
21GLM 5.1Open weightsZ.ai (Zhipu AI) · GLM5.1 · best of 4 rows76.3%IndependentreasoningonPartially comparable-7.07 ptobs. 12 Sept 2026artificialanalysis.aiT2History
22Kimi K2.6Open weightsMoonshot AI · Kimi · best of 4 rows76.0%IndependentreasoningonPartially comparable-7.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
23gpt-5.4-nanoClosedOpenAI · GPT 5.4 · best of 6 rows75.9%Independentreasoning_effortxhighPartially comparable-7.41 ptobs. 12 Sept 2026artificialanalysis.aiT2History
23muse-sparkClosedMeta AI · best of 2 rows75.9%IndependentreasoningonPartially comparable-7.41 ptobs. 12 Sept 2026artificialanalysis.aiT2History
25gpt-5.5ClosedOpenAI · GPT 5.5 · best of 10 rows75.8%Independentreasoning_effortxhighPartially comparable-7.48 ptobs. 12 Sept 2026artificialanalysis.aiT2History
26MiniMax M2.7Open weightsMiniMax · MiniMax · best of 2 rows75.7%IndependentreasoningonPartially comparable-7.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
26Qwen3.5-122B-A10BOpen weightsQwen · Qwen3.5 · best of 4 rows75.7%IndependentreasoningonPartially comparable-7.62 ptobs. 12 Sept 2026artificialanalysis.aiT2History
28Gemma 4 31BOpen weightsGoogle · Gemma 4 · best of 4 rows75.6%IndependentreasoningonPartially comparable-7.75 ptobs. 12 Sept 2026artificialanalysis.aiT2History
28Qwen3.5-27BOpen weightsQwen · Qwen3.5 · best of 4 rows75.6%IndependentreasoningonPartially comparable-7.75 ptobs. 12 Sept 2026artificialanalysis.aiT2History
30gpt-5-miniClosedOpenAI · GPT 5 · best of 6 rows75.4%Independentreasoning_efforthighPartially comparable-7.89 ptobs. 12 Sept 2026artificialanalysis.aiT2History
30gpt-5.2ClosedOpenAI · GPT 5.2 · best of 6 rows75.4%Independentreasoning_effortxhighPartially comparable-7.89 ptobs. 12 Sept 2026artificialanalysis.aiT2History
32gpt-5.3-codexClosedOpenAI · GPT 5.3 · best of 2 rows75.4%Independentreasoning_effortxhighPartially comparable-7.96 ptobs. 12 Sept 2026artificialanalysis.aiT2History
33Qwen3.6 PlusClosedQwen · Qwen3.6 · best of 2 rows75.2%IndependentreasoningonPartially comparable-8.16 ptobs. 12 Sept 2026artificialanalysis.aiT2History
34GPT-5-CodexClosedOpenAI · GPT 5 · best of 2 rows74.2%Independentreasoning_efforthighPartially comparable-9.18 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35command-a-plusOpen weightsCohere · Command · best of 2 rows74.0%IndependentreasoningonPartially comparable-9.38 ptobs. 12 Sept 2026artificialanalysis.aiT2History
35gpt-5.4ClosedOpenAI · GPT 5.4 · best of 6 rows74.0%Independentreasoning_effortxhighPartially comparable-9.38 ptobs. 12 Sept 2026artificialanalysis.aiT2History
37gemma-4-12BOpen weightsGoogle · Gemma 4 · best of 4 rows73.5%IndependentreasoningonPartially comparable-9.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
38deepseek-v4-flash-0420-highOpen weightsDeepSeek · DeepSeek73.5%Independentgroup defaultsPartially comparable-9.86 ptobs. 11 Sept 2026artificialanalysis.aiT2History
39Z.ai GLM 5.2Open weightsZ.ai (Zhipu AI) · GLM5.2 · best of 2 rows73.3%Independentreasoning_effortmaxPartially comparable-10.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
40gpt-5.4-miniClosedOpenAI · GPT 5.4 · best of 6 rows73.3%Independentreasoning_effortxhighPartially comparable-10.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
41GLM 5 TurboClosedZ.ai (Zhipu AI) · GLM5 · best of 2 rows73.2%IndependentreasoningonPartially comparable-10.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
42gpt-5ClosedOpenAI · GPT 5 · best of 8 rows73.1%Independentreasoning_efforthighPartially comparable-10.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
43gpt-5.1ClosedOpenAI · GPT 5.1 · best of 4 rows72.9%Independentreasoning_efforthighPartially comparable-10.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
44gpt-5.6-solClosedOpenAI · GPT 5.6 · best of 10 rows72.7%Independentreasoning_effortmaxPartially comparable-10.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
45Qwen3.5-35B-A3BOpen weightsQwen · Qwen3.5 · best of 4 rows72.5%IndependentreasoningonPartially comparable-10.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
46Gemma 4 26B A4BOpen weightsGoogle · Gemma 4 · best of 4 rows72.5%IndependentreasoningonPartially comparable-10.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
47GLM 5Open weightsZ.ai (Zhipu AI) · GLM5 · best of 4 rows72.3%IndependentreasoningonPartially comparable-11.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
47MiniMax M2Open weightsMiniMax · MiniMax · best of 2 rows72.3%IndependentreasoningonPartially comparable-11.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
49mimo-v2-0206Open weightsXiaomi · best of 2 rows71.8%IndependentreasoningonPartially comparable-11.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
50MiniMax M2.5Open weightsMiniMax · MiniMax · best of 2 rows71.6%IndependentreasoningonPartially comparable-11.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
51Nemotron 3 SuperOpen weightsNVIDIA · Nemotron 3 · best of 2 rows71.5%IndependentreasoningonPartially comparable-11.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
51gpt-5-5-instant-05-26ClosedOpenAI · GPT 5.5 · best of 2 rows71.5%IndependentreasoningonPartially comparable-11.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
53o3ClosedOpenAI · OpenAI o-series · best of 2 rows71.4%IndependentreasoningonPartially comparable-11.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
54deepseek-v4-pro-0424-highOpen weightsDeepSeek · DeepSeek71.3%Independentgroup defaultsPartially comparable-12.0 ptobs. 11 Sept 2026artificialanalysis.aiT2History
55gpt-5.6-terraClosedOpenAI · GPT 5.6 · best of 10 rows71.2%Independentreasoning_effortmaxPartially comparable-12.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
56Solar Pro 3ClosedUpstage · Solar · best of 2 rows71.2%IndependentreasoningonPartially comparable-12.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
57Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 4 rows71.1%IndependentreasoningonPartially comparable-12.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
58Qwen3 MaxClosedQwen · Qwen3 · best of 4 rows70.8%IndependentreasoningonPartially comparable-12.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
58nova-2-0-liteClosedAmazon Web Services · Nova 2.0 · best of 8 rows70.8%Independentreasoningonreasoning_efforthighPartially comparable-12.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
60gemini-3-proClosedGoogle · Gemini 3 · best of 4 rows70.4%Independentreasoning_efforthighPartially comparable-12.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
61o1ClosedOpenAI · OpenAI o-series · best of 2 rows70.3%IndependentreasoningonPartially comparable-13.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
62Kimi K2.5Open weightsMoonshot AI · Kimi · best of 4 rows70.2%IndependentreasoningonPartially comparable-13.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
63GPT-5.1-CodexClosedOpenAI · GPT 5.1 · best of 2 rows70%Independentreasoning_efforthighPartially comparable-13.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
64MiniMax M2.1Open weightsMiniMax · MiniMax · best of 2 rows69.9%IndependentreasoningonPartially comparable-13.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
65Mercury 2ClosedInception · best of 2 rows69.8%IndependentreasoningonPartially comparable-13.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
66apriel-v1-6-15b-thinkerOpen weightsServiceNow · best of 2 rows69.1%IndependentreasoningonPartially comparable-14.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
67gpt-oss-120bOpen weightsOpenAI · gpt-oss · best of 4 rows69.0%Independentreasoning_efforthighPartially comparable-14.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
68mimo-v2-proClosedXiaomi · best of 2 rows68.8%IndependentreasoningonPartially comparable-14.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
69Mistral Medium 3.5Open weightsMistral AI · Mistral · best of 2 rows68.8%IndependentreasoningonPartially comparable-14.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
70o4-miniClosedOpenAI · OpenAI o-series · best of 2 rows68.7%Independentreasoning_efforthighPartially comparable-14.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
71kat-coder-pro-v1ClosedKwaiKAT · best of 2 rows68.4%IndependentreasoningoffPartially comparable-15.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
72Kimi K2 ThinkingOpen weightsMoonshot AI · Kimi · best of 2 rows68.1%IndependentreasoningonPartially comparable-15.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
73GLM 4.7Open weightsZ.ai (Zhipu AI) · GLM4.7 · best of 4 rows67.9%IndependentreasoningonPartially comparable-15.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
73GPT-5.1-Codex MiniClosedOpenAI · GPT 5.1 · best of 2 rows67.9%Independentreasoning_efforthighPartially comparable-15.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
75Qwen3.6 27BOpen weightsQwen · Qwen3.6 · best of 4 rows67.5%IndependentreasoningonPartially comparable-15.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
75gpt-5-nanoClosedOpenAI · GPT 5 · best of 6 rows67.5%Independentreasoning_efforthighPartially comparable-15.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
77mimo-v2-omni-0327ClosedXiaomi · best of 2 rows67.3%IndependentreasoningonPartially comparable-16.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
78Step 3.7 FlashOpen weightsStepFun · Step3.7 · best of 2 rows67.3%IndependentreasoningonPartially comparable-16.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79MiMo-V2.5Open weightsXiaomi · best of 2 rows67.1%IndependentreasoningonPartially comparable-16.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
79o3-miniClosedOpenAI · OpenAI o-series · best of 2 rows67.1%Independentreasoning_efforthighPartially comparable-16.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
81Qwen3.5-9BOpen weightsQwen · Qwen3.5 · best of 4 rows66.7%IndependentreasoningonPartially comparable-16.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
82KAT-Coder-Pro V2ClosedKwaipilot · best of 2 rows66.7%IndependentreasoningoffPartially comparable-16.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
83Step 3.5 FlashOpen weightsStepFun · Step3.5 · best of 4 rows66.5%IndependentreasoningonPartially comparable-16.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
84hypernova-60bOpen weightsMultiverse Computing · best of 2 rows66.5%Independentreasoning_efforthighPartially comparable-16.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
85nex-n2-proOpen weightsNex AGI · best of 2 rows66.2%IndependentreasoningonPartially comparable-17.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
85nova-2-0-omniClosedAmazon Web Services · Nova 2.0 · best of 6 rows66.2%Independentreasoningonreasoning_effortmediumPartially comparable-17.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
87olmo-3-1-32b-instructOpen weightsAllen Institute for AI · OLMo 3.1 · best of 3 rows66.0%IndependentreasoningonPartially comparable-17.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
88gpt-oss-20bOpen weightsOpenAI · gpt-oss · best of 4 rows65.1%Independentreasoning_efforthighPartially comparable-18.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
89k-exaoneOpen weightsLG AI Research · EXAONE · best of 4 rows64.7%IndependentreasoningonPartially comparable-18.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
90Qwen3.6 35B A3BOpen weightsQwen · Qwen3.6 · best of 4 rows64.3%IndependentreasoningonPartially comparable-19.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
91mimo-v2-flashOpen weightsXiaomi · best of 4 rows64.2%IndependentreasoningonPartially comparable-19.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
92deepseek-v3-2-specialeOpen weightsDeepSeek · DeepSeek · best of 2 rows63.9%IndependentreasoningonPartially comparable-19.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
93Claude Fable 5ClosedAnthropic · Claude · best of 2 rows63.5%IndependentreasoningonPartially comparable-19.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
94nemotron-3-nano-omni-30b-a3bOpen weightsNVIDIA · Nemotron 3 · best of 2 rows63.2%IndependentreasoningonPartially comparable-20.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
95Hy3Open weightsTencent · best of 3 rows63.1%IndependentreasoningonPartially comparable-20.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
95Hy3 previewOpen weightsTencent63.1%Independentgroup defaultsPartially comparable-20.2 ptobs. 11 Sept 2026artificialanalysis.aiT2History
95Kimi K2.7 CodeOpen weightsMoonshot AI · Kimi · best of 2 rows63.1%IndependentreasoningonPartially comparable-20.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
98k2-think-v2Open weightsMBZUAI Institute of Foundation Models · best of 2 rows62.8%IndependentreasoningonPartially comparable-20.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
99Claude Opus 4.8ClosedAnthropic · Claude · best of 2 rows62.2%Independentreasoning_effortmaxPartially comparable-21.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
100apriel-v1-5-15b-thinkerOpen weightsServiceNow · best of 2 rows61.7%IndependentreasoningonPartially comparable-21.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →