Skip to content
AI Atlas
BenchmarkActive

LiveBench

livebench.ai

contamination-limited, monthly refreshed questions across 6 categories

quality57

Updated 10 h ago · first seen 11 Sept 2026

Metric
· %
Current results
0
Models
0
Current leader
Claude Fable 5.1 83.4%

No results recorded for this benchmark yet — its sources are being connected. The definition, aliases and variants are kept so links resolve; nothing is fabricated.

Score history · gpt-5.6-terra 8 rows

Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.

  • 64.62%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
  • 82.89%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
  • 79.31%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
  • 94.91%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
  • 54.95%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
  • 78.25%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
  • 90.64%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
  • 77.94%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gpt-5.6-terra-max25 Jun 2026

Back to the leaderboard

Leaderboard 57 models

Select models with +, then Compare.

No result in this group with these filters

Relax the trust / organization filters or pick another comparability group.