Skip to content
AI Atlas
BenchmarkActivecategory · general

LiveBench

livebench.ai

contamination-limited, monthly refreshed questions across 6 categories

quality57

Updated 4 h ago · first seen 11 Sept 2026

bench_01M293SPERF0D8QG7TSCGE8GGF

Metric
average score · %
Direction
Higher is better
Results
456 · 8 filtered
Leader
Claude Fable 5.1 Max Effort 97.01%

Score history · DeepSeek V4 Flash Vision Exp 8 rows

Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.

  • 70.96%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=deepseek-v4-flash-vision-exp25 Jun 2026
  • 80.36%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=deepseek-v4-flash-vision-exp25 Jun 2026
  • 79.48%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=deepseek-v4-flash-vision-exp25 Jun 2026
  • 87.81%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=deepseek-v4-flash-vision-exp25 Jun 2026
  • 65.1%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=deepseek-v4-flash-vision-exp25 Jun 2026
  • 68.2%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=deepseek-v4-flash-vision-exp25 Jun 2026
  • 85.4%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=deepseek-v4-flash-vision-exp25 Jun 2026
  • 76.76%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=deepseek-v4-flash-vision-exp25 Jun 2026

Back to the full leaderboard

Leaderboard 8 current results · config contains “gemini-3.6-flash-high”

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
1#1Gemini 3.6 Flash HighGoogle86.4%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026livebench.aiT2 History
2#2Gemini 3.6 Flash HighGoogle85.15%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026livebench.aiT2 History
3#3Gemini 3.6 Flash HighGoogle83.9%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026livebench.aiT2 History
4#4Gemini 3.6 Flash HighGoogle77.86%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026livebench.aiT2 History
5#5Gemini 3.6 Flash HighGoogle75.37%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026livebench.aiT2 History
6#6Gemini 3.6 Flash HighGoogle73.59%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gemini-3.6-flash-high25 Jun 2026livebench.aiT2 History
7#7Gemini 3.6 Flash HighGoogle63%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026livebench.aiT2 History
8#8Gemini 3.6 Flash HighGoogle43.43%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026livebench.aiT2 History

8 results

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.