Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.
69.64%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026
74.34%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026
72.55%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026
87.15%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026
58.54%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026
77.16%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026
87.73%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026
75.3%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026
9 leader changes recorded, all dated 25 Jun 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →