Skip to content
AI Atlas
BenchmarkActivecategory · generalfamily · livebench · variant global

LiveBench

livebench.ai

contamination-limited, monthly refreshed questions across 7 categories — overall mean of the category averages

quality57

Updated 28 min ago · first seen 11 Sept 2026

Metric
global_average · %
Current results
57
Models
57
Current leader
Claude Fable 5.1 83.4%

Score history · Gemini 3.5 Flash-Lite High 8 rows

Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.

  • 67.24%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026
  • 71.82%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026
  • 53.25%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026
  • 73.74%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026
  • 45.25%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026
  • 76.07%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026
  • 60.19%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026
  • 63.94%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026

Back to the leaderboard

Frontier over time · global_average

9 leader changes recorded, all dated 25 Jun 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 83.4%Claude Fable 5.1 Anthropic Official board25 Jun 2026
  2. 83.0%Claude Fable 5 Anthropic Official board25 Jun 2026
  3. 81.1%gpt-5.6-sol OpenAI Official board25 Jun 2026
  4. 80.2%gpt-5.5 OpenAI Official board25 Jun 2026
  5. 78.0%gpt-5.4 OpenAI Official board25 Jun 2026
  6. 77.0%Gemini 3.1 Pro Preview Google Official board25 Jun 2026
  7. 76.5%Claude Opus 4.7 Anthropic Official board25 Jun 2026
  8. 74.5%Claude Opus 4.6 Anthropic Official board25 Jun 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 57 models

Select models with +, then Compare.

No result in this group with these filters

Relax the trust / organization filters or pick another comparability group.