Skip to content
AI Atlas
BenchmarkActivecategory · general

LiveBench

livebench.ai

contamination-limited, monthly refreshed questions across 6 categories

quality57

Updated 5 h ago · first seen 11 Sept 2026

bench_01M293SPERF0D8QG7TSCGE8GGF

Metric
average score · %
Direction
Higher is better
Results
456 · 8 filtered
Leader
Claude Fable 5.1 Max Effort 97.01%

Score history · Gemini 3.6 Flash High 8 rows

Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.

  • 75.37%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026
  • 83.9%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026
  • 63%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026
  • 86.4%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026
  • 43.43%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026
  • 77.86%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026
  • 85.15%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026
  • 73.59%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gemini-3.6-flash-high25 Jun 2026

Back to the full leaderboard

Leaderboard 8 current results · config contains “gemini-3.8-flash-high”

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
1#1Gemini 3.8 FlashGoogle91.56%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026livebench.aiT2 History
2#2Gemini 3.8 FlashGoogle89.29%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026livebench.aiT2 History
3#3Gemini 3.8 FlashGoogle87.79%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026livebench.aiT2 History
4#4Gemini 3.8 FlashGoogle81.41%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026livebench.aiT2 History
5#5Gemini 3.8 FlashGoogle75.83%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gemini-3.8-flash-high25 Jun 2026livebench.aiT2 History
6#6Gemini 3.8 FlashGoogle72.49%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026livebench.aiT2 History
7#7Gemini 3.8 FlashGoogle54.24%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026livebench.aiT2 History
8#8Gemini 3.8 FlashGoogle54.01%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026livebench.aiT2 History

8 results

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.