Skip to content
AI Atlas
BenchmarkActivecategory · general

LiveBench

livebench.ai

contamination-limited, monthly refreshed questions across 6 categories

quality57

Updated 3 h ago · first seen 11 Sept 2026

bench_01M293SPERF0D8QG7TSCGE8GGF

Metric
average score · %
Direction
Higher is better
Results
456 · 8 filtered
Leader
Claude Fable 5.1 Max Effort 97.01%

Leaderboard 8 current results · config contains “gpt-5.6-sol-max”

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
1#1gpt-5-6-solOpenAI96.2%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026livebench.aiT2 History
2#2gpt-5-6-solOpenAI91.65%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026livebench.aiT2 History
3#3gpt-5-6-solOpenAI87.68%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026livebench.aiT2 History
4#4gpt-5-6-solOpenAI83.94%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026livebench.aiT2 History
5#5gpt-5-6-solOpenAI81.05%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gpt-5.6-sol-max25 Jun 2026livebench.aiT2 History
6#6gpt-5-6-solOpenAI79.84%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026livebench.aiT2 History
7#7gpt-5-6-solOpenAI71.85%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026livebench.aiT2 History
8#8gpt-5-6-solOpenAI56.21%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026livebench.aiT2 History

8 results

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.