Skip to content
AI Atlas
BenchmarkActivecategory · general

LiveBench

livebench.ai

contamination-limited, monthly refreshed questions across 6 categories

quality57

Updated 8 h ago · first seen 11 Sept 2026

bench_01M293SPERF0D8QG7TSCGE8GGF

Metric
average score · %
Direction
Higher is better
Results
456 · 8 filtered
Leader
Claude Fable 5.1 Max Effort 97.01%

Score history · muse-spark-1-3-xhigh 8 rows

Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.

  • 78%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=muse-spark-1.3-xhigh25 Jun 2026
  • 82.79%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=muse-spark-1.3-xhigh25 Jun 2026
  • 79.57%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=muse-spark-1.3-xhigh25 Jun 2026
  • 95.95%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=muse-spark-1.3-xhigh25 Jun 2026
  • 64.09%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=muse-spark-1.3-xhigh25 Jun 2026
  • 81.06%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=muse-spark-1.3-xhigh25 Jun 2026
  • 89.65%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=muse-spark-1.3-xhigh25 Jun 2026
  • 81.59%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=muse-spark-1.3-xhigh25 Jun 2026

Back to the full leaderboard

Leaderboard 8 current results · config contains “gemini-3.1-pro-preview-high”

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
1#1Gemini 3.1 Pro Preview HighGoogle91.05%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.1-pro-preview-high25 Jun 2026livebench.aiT2 History
2#2Gemini 3.1 Pro Preview HighGoogle85.38%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.1-pro-preview-high25 Jun 2026livebench.aiT2 History
3#3Gemini 3.1 Pro Preview HighGoogle84.01%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gemini-3.1-pro-preview-high25 Jun 2026livebench.aiT2 History
4#4Gemini 3.1 Pro Preview HighGoogle79.1%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gemini-3.1-pro-preview-high25 Jun 2026livebench.aiT2 History
5#5Gemini 3.1 Pro Preview HighGoogle78.54%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gemini-3.1-pro-preview-high25 Jun 2026livebench.aiT2 History
6#6Gemini 3.1 Pro Preview HighGoogle76.95%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gemini-3.1-pro-preview-high25 Jun 2026livebench.aiT2 History
7#7Gemini 3.1 Pro Preview HighGoogle76.45%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gemini-3.1-pro-preview-high25 Jun 2026livebench.aiT2 History
8#8Gemini 3.1 Pro Preview HighGoogle44.14%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=gemini-3.1-pro-preview-high25 Jun 2026livebench.aiT2 History

8 results

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.