Updated 3 h ago · first seen 11 Sept 2026
bench_01M293SPERF0D8QG7TSCGE8GGF
- Metric
- average score · %
- Direction
- Higher is better
- Results
- 456
Score history · Muse Spark 1.2 8 rows
Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.
- 74.33%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=muse-spark-1.2-xhigh25 Jun 2026
- 78.58%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=muse-spark-1.2-xhigh25 Jun 2026
- 76.46%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=muse-spark-1.2-xhigh25 Jun 2026
- 91.2%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=muse-spark-1.2-xhigh25 Jun 2026
- 57.58%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=muse-spark-1.2-xhigh25 Jun 2026
- 77.54%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=muse-spark-1.2-xhigh25 Jun 2026
- 90.01%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=muse-spark-1.2-xhigh25 Jun 2026
- 77.96%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=muse-spark-1.2-xhigh25 Jun 2026
Leaderboard 456 current results
Select models with +, then open Compare.
| # | Model | Score | Config | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|
| #1Claude Fable 5.1 Max EffortAnthropic | 97.01% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=claude-fable-5-1-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #2gpt-6-astraOpenAI | 96.81% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-6-astra-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #3gpt-5-6-solOpenAI | 96.2% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-5.6-sol-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #4Claude Fable 5 xHigh EffortAnthropic | 95.99% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=claude-fable-5-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #5muse-spark-1-3-xhighMeta AI | 95.95% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=muse-spark-1.3-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #6gpt-5.5OpenAI | 95.86% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-5.5-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #7claude-opus-5-max-effort | 95.73% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=claude-opus-5-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #8DeepSeek-V4-Pro-0813DeepSeek | 95.09% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=deepseek-v4-pro-0813 | 25 Jun 2026 | livebench.aiT2 | History | |
| #9gpt-5.6-terraOpenAI | 94.91% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-5.6-terra-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #10Claude 4.8 Opus Thinking Max EffortAnthropic | 94.32% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=claude-opus-4-8-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #11gpt-5.4OpenAI | 94.15% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-5.4-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #12Gemini 3.7 FlashGoogle | 93.47% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.7-flash-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #13DeepSeek-V4.1-FlashDeepSeek | 93.29% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=deepseek-v4.1-flash-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #14gpt-5.2-2025-12-11-high | 93.17% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-5.2-2025-12-11-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #15Claude Sonnet 5 xHigh EffortAnthropic | 92.94% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=claude-sonnet-5-xhigh-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #16claude-opus-4-7-xhigh-effort | 92.85% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=claude-opus-4-7-xhigh-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #17gpt-6-astraOpenAI | 92.65% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gpt-6-astra-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #18Grok 4.6xAI | 92.57% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=grok-4.6 | 25 Jun 2026 | livebench.aiT2 | History | |
| #19Claude Fable 5.1 Max EffortAnthropic | 91.69% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=claude-fable-5-1-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #20gpt-5-6-solOpenAI | 91.65% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gpt-5.6-sol-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #21Gemini 3.8 FlashGoogle | 91.56% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.8-flash-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #22Qwen 3.8 MaxQwen | 91.31% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=qwen3.8-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #23claude-opus-5-max-effort | 91.21% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=claude-opus-5-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #24Muse Spark 1.2Meta AI | 91.2% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=muse-spark-1.2-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #25Gemini 3.1 Pro Preview HighGoogle | 91.05% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.1-pro-preview-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #26gpt-5.4-nanoOpenAI | 90.98% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-5.4-nano-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #27Grok 4.5xAI | 90.83% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=grok-4.5 | 25 Jun 2026 | livebench.aiT2 | History | |
| #28Claude Fable 5 xHigh EffortAnthropic | 90.68% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=claude-fable-5-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #29deepseek-v4-proDeepSeek | 90.68% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=deepseek-v4-pro | 25 Jun 2026 | livebench.aiT2 | History | |
| #30Kimi K3Moonshot AI | 90.67% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=kimi-k3 | 25 Jun 2026 | livebench.aiT2 | History | |
| #31gpt-5.6-terraOpenAI | 90.64% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gpt-5.6-terra-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #32Grok 4.6xAI | 90.51% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=grok-4.6 | 25 Jun 2026 | livebench.aiT2 | History | |
| #33claude-opus-4-5-20251101-thinking-64k-high-effort | 90.39% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=claude-opus-4-5-20251101-thinking-64k-high-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #34Smaug Agentic | 90.27% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=smaug-agentic | 25 Jun 2026 | livebench.aiT2 | History | |
| #35Smaug Flash | 90.06% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=smaug-flash | 25 Jun 2026 | livebench.aiT2 | History | |
| #36Muse Spark 1.2Meta AI | 90.01% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=muse-spark-1.2-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #37Z.ai GLM 5.2Z.ai (Zhipu AI) | 89.78% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=glm-5.2 | 25 Jun 2026 | livebench.aiT2 | History | |
| #38Smaug Mini | 89.74% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=smaug-mini | 25 Jun 2026 | livebench.aiT2 | History | |
| #39Claude Fable 5 xHigh EffortAnthropic | 89.65% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=claude-fable-5-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #40gpt-5.5OpenAI | 89.65% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gpt-5.5-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #41muse-spark-1-3-xhighMeta AI | 89.65% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=muse-spark-1.3-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #42Claude Fable 5.1 Max EffortAnthropic | 89.5% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=claude-fable-5-1-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #43gpt-6-astraOpenAI | 89.43% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gpt-6-astra-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #44Claude 4.6 Opus Thinking High EffortAnthropic | 89.32% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=claude-opus-4-6-thinking-auto-high-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #45Gemini 3.8 FlashGoogle | 89.29% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gemini-3.8-flash-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #46Claude 4.8 Opus Thinking Max EffortAnthropic | 89.19% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=claude-opus-4-8-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #47GPT-5.2-CodexOpenAI | 88.77% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-5.2-codex | 25 Jun 2026 | livebench.aiT2 | History | |
| #48Claude Sonnet 5 xHigh EffortAnthropic | 88.69% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=claude-sonnet-5-xhigh-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #49claude-opus-5-max-effort | 88.69% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=claude-opus-5-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #50Claude 4.6 Opus Thinking High EffortAnthropic | 88.67% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=claude-opus-4-6-thinking-auto-high-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #51Nemotron 3 UltraNVIDIA | 88.66% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=nemotron-3-ultra-550b-a55b | 25 Jun 2026 | livebench.aiT2 | History | |
| #52Inkling xHigh Effort | 88.37% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=inkling-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #53gemini-3.5-flash-high | 88.24% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.5-flash-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #54Qwen 3.8 MaxQwen | 88.21% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=qwen3.8-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #55gpt-5.4OpenAI | 88.12% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gpt-5.4-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #56GLM 5.3Z.ai (Zhipu AI) | 87.9% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=glm-5.3 | 25 Jun 2026 | livebench.aiT2 | History | |
| #57DeepSeek V4 Flash Vision ExpDeepSeek | 87.81% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=deepseek-v4-flash-vision-exp | 25 Jun 2026 | livebench.aiT2 | History | |
| #58Gemini 3.7 FlashGoogle | 87.8% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gemini-3.7-flash-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #59Gemini 3.8 FlashGoogle | 87.79% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.8-flash-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #60Muse Spark 1.1Meta AI | 87.73% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=muse-spark-1.1-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #61gpt-5-6-solOpenAI | 87.68% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gpt-5.6-sol-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #62Qwen3.8 FlashQwen | 87.38% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=qwen3.8-flash-next | 25 Jun 2026 | livebench.aiT2 | History | |
| #63gpt-5.5OpenAI | 87.36% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gpt-5.5-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #64gpt-5.6-lunaOpenAI | 87.2% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gpt-5.6-luna-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #65claude-opus-4-7-xhigh-effort | 87.19% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=claude-opus-4-7-xhigh-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #66Grok 4.5xAI | 87.17% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=grok-4.5 | 25 Jun 2026 | livebench.aiT2 | History | |
| #67Muse Spark 1.1Meta AI | 87.15% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=muse-spark-1.1-xhigh | 25 Jun 2026 | livebench.aiT2 | History | |
| #68claude-sonnet-4-6-thinking-auto-medium-effort | 86.99% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #69DeepSeek V4 Flash (0731)DeepSeek | 86.79% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=deepseek-v4-flash-0731 | 25 Jun 2026 | livebench.aiT2 | History | |
| #70DeepSeek-V4.1-FlashDeepSeek | 86.69% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=deepseek-v4.1-flash-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #71DeepSeek V4 Flash (0731)DeepSeek | 86.64% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=deepseek-v4-flash-0731 | 25 Jun 2026 | livebench.aiT2 | History | |
| #72Gemini 3.6 Flash HighGoogle | 86.4% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.6-flash-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #73Claude Fable 5.1 Max EffortAnthropic | 86.38% | release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=claude-fable-5-1-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #74Smaug Flash | 86.23% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=smaug-flash | 25 Jun 2026 | livebench.aiT2 | History | |
| #75qwen3-8-27b-non-reasoningAlibaba Group | 86.21% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=qwen3.8-27b | 25 Jun 2026 | livebench.aiT2 | History | |
| #76Claude Fable 5 xHigh EffortAnthropic | 85.99% | release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=claude-fable-5-max-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #77DeepSeek-V4-Pro-0813DeepSeek | 85.84% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=deepseek-v4-pro-0813 | 25 Jun 2026 | livebench.aiT2 | History | |
| #78Qwen3.8 FlashQwen | 85.82% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=qwen3.8-flash-next | 25 Jun 2026 | livebench.aiT2 | History | |
| #79GLM 5.3Z.ai (Zhipu AI) | 85.8% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=glm-5.3 | 25 Jun 2026 | livebench.aiT2 | History | |
| #80gpt-5.6-lunaOpenAI | 85.64% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gpt-5.6-luna-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #81Kimi K3Moonshot AI | 85.53% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=kimi-k3 | 25 Jun 2026 | livebench.aiT2 | History | |
| #82Gemini 3.7 FlashGoogle | 85.46% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.7-flash-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #83DeepSeek V4 Flash Vision ExpDeepSeek | 85.4% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=deepseek-v4-flash-vision-exp | 25 Jun 2026 | livebench.aiT2 | History | |
| #84Gemini 3.1 Pro Preview HighGoogle | 85.38% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.1-pro-preview-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #85Qwen3.7 MaxQwen | 85.25% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=qwen3.7-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #86Gemini 3.6 Flash HighGoogle | 85.15% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gemini-3.6-flash-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #87claude-sonnet-4-6-thinking-auto-medium-effort | 84.77% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort | 25 Jun 2026 | livebench.aiT2 | History | |
| #88gemini-3.5-flash-high | 84.58% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.5-flash-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #89Kimi K3Moonshot AI | 84.44% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=kimi-k3 | 25 Jun 2026 | livebench.aiT2 | History | |
| #90Smaug Agentic | 84.36% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=smaug-agentic | 25 Jun 2026 | livebench.aiT2 | History | |
| #91Grok 4.3xAI | 84.34% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=grok-4.3 | 25 Jun 2026 | livebench.aiT2 | History | |
| #92Kimi K2.6 ThinkingMoonshot AI | 84.28% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=kimi-k2.6-thinking | 25 Jun 2026 | livebench.aiT2 | History | |
| #93Gemini 3.1 Pro Preview HighGoogle | 84.01% | release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gemini-3.1-pro-preview-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #94gpt-5-6-solOpenAI | 83.94% | release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gpt-5.6-sol-max | 25 Jun 2026 | livebench.aiT2 | History | |
| #95Smaug Agentic | 83.92% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=smaug-agentic | 25 Jun 2026 | livebench.aiT2 | History | |
| #96Gemini 3.6 Flash HighGoogle | 83.9% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.6-flash-high | 25 Jun 2026 | livebench.aiT2 | History | |
| #97Qwen3.6 PlusQwen | 83.73% | release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=qwen3.6-plus | 25 Jun 2026 | livebench.aiT2 | History | |
| #98Grok 4.6xAI | 83.7% | release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=grok-4.6 | 25 Jun 2026 | livebench.aiT2 | History | |
| #99GPT-5.2-CodexOpenAI | 83.62% | release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gpt-5.2-codex | 25 Jun 2026 | livebench.aiT2 | History | |
| #100Claude Fable 5.1 Max EffortAnthropic | 83.41% | release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=claude-fable-5-1-max-effort | 25 Jun 2026 | livebench.aiT2 | History |
Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.
The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.
Definition
- Category
- general
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 3 h agomedium
- Task
- contamination-limited, monthly refreshed questions across 6 categories
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 3 h agomedium
- Metric
- average score · %
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 3 h agomedium
- Direction
- Higher is better
- Paper
- https://arxiv.org/abs/2406.19314
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 3 h agomedium
- Website
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 3 h agomedium
Each value shows its source, tier and observation time. Missing rows mean no source stated them. How results are recorded →
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history
Categorycategory1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| general | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Metricmetric1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| average score | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Paperpaper1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/abs/2406.19314 | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Tasktask1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| contamination-limited, monthly refreshed questions across 6 categories | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Unitunit1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| % | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Websitewebsite1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://livebench.ai | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| LiveBench | livebench.ai/table_2026_06_25.csv | leaderboard | T2· Quality secondary | 58 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.