Skip to content
AI Atlas
BenchmarkActivecategory · instruction-followingfamily · ifeval · variant IFBench

IFBench

precise instruction following with novel constraints

quality51

Updated 30 min ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
450
Models
333
Current leader
Grok 4.3 83.3%

Score history · grok-4-3-non-reasoning 1 row

Not enough history to chart — a single observation (47.62% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.

  • 47.62%aa_slug=grok-4-3-non-reasoning · evaluator=Artificial Analysis · index_version=4.311 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · evaluator=Artificial Analysis

6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 83.3%Grok 4.3 xAI Independent11 Sept 2026
  2. 81.3%Grok 4.3 xAI Independent11 Sept 2026
  3. 81.2%Grok 4.20 xAI Independent11 Sept 2026
  4. 81.0%Grok 4.3 xAI Independent11 Sept 2026
  5. 73.5%gemma-4-12B Google Independent11 Sept 2026
  6. 45.9%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 333 models

Select models with +, then Compare.

No result in this group with these filters

Relax the trust / organization filters or pick another comparability group.