Skip to content
AI Atlas
BenchmarkActivecategory · general

LiveBench

livebench.ai

contamination-limited, monthly refreshed questions across 6 categories

quality57

Updated 4 h ago · first seen 11 Sept 2026

bench_01M293SPERF0D8QG7TSCGE8GGF

Metric
average score · %
Direction
Higher is better
Results
456 · 8 filtered
Leader
Claude Fable 5.1 Max Effort 97.01%

Score history · Z.ai GLM 5.2 8 rows

Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.

  • 62.29%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=glm-5.225 Jun 2026
  • 76.24%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=glm-5.225 Jun 2026
  • 73.74%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=glm-5.225 Jun 2026
  • 89.78%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=glm-5.225 Jun 2026
  • 51.77%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=glm-5.225 Jun 2026
  • 79.65%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=glm-5.225 Jun 2026
  • 78.63%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=glm-5.225 Jun 2026
  • 73.16%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=glm-5.225 Jun 2026

Back to the full leaderboard

Leaderboard 8 current results · config contains “claude-fable-5-max-effort”

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
1#1Claude Fable 5 xHigh EffortAnthropic95.99%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=claude-fable-5-max-effort25 Jun 2026livebench.aiT2 History
2#2Claude Fable 5 xHigh EffortAnthropic90.68%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=claude-fable-5-max-effort25 Jun 2026livebench.aiT2 History
3#3Claude Fable 5 xHigh EffortAnthropic89.65%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=claude-fable-5-max-effort25 Jun 2026livebench.aiT2 History
4#4Claude Fable 5 xHigh EffortAnthropic85.99%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=claude-fable-5-max-effort25 Jun 2026livebench.aiT2 History
5#5Claude Fable 5 xHigh EffortAnthropic82.97%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=claude-fable-5-max-effort25 Jun 2026livebench.aiT2 History
6#6Claude Fable 5 xHigh EffortAnthropic80.54%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=claude-fable-5-max-effort25 Jun 2026livebench.aiT2 History
7#7Claude Fable 5 xHigh EffortAnthropic75.77%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=claude-fable-5-max-effort25 Jun 2026livebench.aiT2 History
8#8Claude Fable 5 xHigh EffortAnthropic62.17%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=claude-fable-5-max-effort25 Jun 2026livebench.aiT2 History

8 results

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.