BenchmarkActive
quality57
Updated 10 h ago · first seen 11 Sept 2026
- Metric
- — · % ↑
- Current results
- 0
- Models
- 0
- Current leader
- Claude Fable 5.1 83.4%
No results recorded for this benchmark yet — its sources are being connected. The definition, aliases and variants are kept so links resolve; nothing is fabricated.
Score history · gpt-5.6-terra 8 rows
Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.
- 64.62%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
- 82.89%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
- 79.31%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
- 94.91%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
- 54.95%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
- 78.25%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
- 90.64%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
- 77.94%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gpt-5.6-terra-max25 Jun 2026
Leaderboard 57 models
Select models with +, then Compare.
No result in this group with these filters
Relax the trust / organization filters or pick another comparability group.