LiveBench
contamination-limited, monthly refreshed questions across 7 categories — overall mean of the category averages
Updated 29 min ago · first seen 11 Sept 2026
- Metric
- global_average · % ↑
- Current results
- 57
- Models
- 57
- Current leader
- Claude Fable 5.1 83.4%
Score history · gpt-5-6-sol 8 rows
Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.
- 71.85%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026
- 87.68%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026
- 79.84%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026
- 96.2%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026
- 56.21%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026
- 83.94%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026
- 91.65%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026
- 81.05%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gpt-5.6-sol-max25 Jun 2026
Frontier over time · global_average
9 leader changes recorded, all dated 25 Jun 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 83.4%Claude Fable 5.1 Anthropic Official board25 Jun 2026
- 83.0%Claude Fable 5 Anthropic Official board25 Jun 2026
- 81.1%gpt-5.6-sol OpenAI Official board25 Jun 2026
- 80.2%gpt-5.5 OpenAI Official board25 Jun 2026
- 78.0%gpt-5.4 OpenAI Official board25 Jun 2026
- 77.0%Gemini 3.1 Pro Preview Google Official board25 Jun 2026
- 76.5%Claude Opus 4.7 Anthropic Official board25 Jun 2026
- 74.5%Claude Opus 4.6 Anthropic Official board25 Jun 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 57 models
Select models with +, then Compare.
No result in this group with these filters