IFBench
precise instruction following with novel constraints
Updated 28 min ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 450
- Models
- 333
- Current leader
- Grok 4.3 83.3%
Score history · gpt-4.1-mini 1 row
Not enough history to chart — a single observation (38.3% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.
- 38.3%aa_slug=gpt-4-1-mini · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
Frontier over time · accuracy · evaluator=Artificial Analysis
6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 83.3%Grok 4.3 xAI Independent11 Sept 2026
- 81.3%Grok 4.3 xAI Independent11 Sept 2026
- 81.2%Grok 4.20 xAI Independent11 Sept 2026
- 81.0%Grok 4.3 xAI Independent11 Sept 2026
- 73.5%gemma-4-12B Google Independent11 Sept 2026
- 45.9%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 333 models
Select models with +, then Compare.
No result in this group with these filters