Skip to content
AI Atlas
BenchmarkActivecategory · general

LiveBench

livebench.ai

contamination-limited, monthly refreshed questions across 6 categories

quality57

Updated 7 h ago · first seen 11 Sept 2026

bench_01M293SPERF0D8QG7TSCGE8GGF

Metric
average score · %
Direction
Higher is better
Results
456
Leader
Claude Fable 5.1 Max Effort 97.01%

Score history · Gemini 3.8 Flash 8 rows

Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.

  • 81.41%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
  • 87.79%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
  • 54.01%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
  • 91.56%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
  • 54.24%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
  • 72.49%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
  • 89.29%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
  • 75.83%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gemini-3.8-flash-high25 Jun 2026

Back to the full leaderboard

Leaderboard 456 current results

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
201#201claude-opus-4-7-xhigh-effort77.91%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=claude-opus-4-7-xhigh-effort25 Jun 2026livebench.aiT2 History
202#202Kimi K2.7 CodeMoonshot AI77.91%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=kimi-k2.7-code25 Jun 2026livebench.aiT2 History
203#203Gemini 3.6 Flash HighGoogle77.86%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026livebench.aiT2 History
204#204GPT-5.2-CodexOpenAI77.71%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gpt-5.2-codex25 Jun 2026livebench.aiT2 History
205#205GLM 5.3 FlashZ.ai (Zhipu AI)77.64%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=glm-5.3-flash25 Jun 2026livebench.aiT2 History
206#206Muse Spark 1.2Meta AI77.54%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=muse-spark-1.2-xhigh25 Jun 2026livebench.aiT2 History
207#207gpt-5.4OpenAI77.54%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gpt-5.4-xhigh25 Jun 2026livebench.aiT2 History
208#208ox-alpha-max77.53%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=ox-alpha-max25 Jun 2026livebench.aiT2 History
209#209DeepSeek-V4-Pro-0813DeepSeek77.44%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=deepseek-v4-pro-081325 Jun 2026livebench.aiT2 History
210#210Smaug Flash77.44%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=smaug-flash25 Jun 2026livebench.aiT2 History
211#211GLM 5.3 FlashZ.ai (Zhipu AI)77.31%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=glm-5.3-flash25 Jun 2026livebench.aiT2 History
212#212DeepSeek-V4.1-FlashDeepSeek77.27%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=deepseek-v4.1-flash-max25 Jun 2026livebench.aiT2 History
213#213Muse Spark 1.1Meta AI77.16%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026livebench.aiT2 History
214#214DeepSeek-V4-Pro-0813DeepSeek77.16%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=deepseek-v4-pro-081325 Jun 2026livebench.aiT2 History
215#215Qwen3.8 FlashQwen77.11%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=qwen3.8-flash-next25 Jun 2026livebench.aiT2 History
216#216Smaug Flash77.1%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=smaug-flash25 Jun 2026livebench.aiT2 History
217#217Smaug Mini77%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=smaug-mini25 Jun 2026livebench.aiT2 History
218#218Gemini 3.1 Pro Preview HighGoogle76.95%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gemini-3.1-pro-preview-high25 Jun 2026livebench.aiT2 History
219#219MiniMax M3MiniMax76.95%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=minimax-m325 Jun 2026livebench.aiT2 History
220#220Smaug Mini76.93%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=smaug-mini25 Jun 2026livebench.aiT2 History
221#221MiniMax M3MiniMax76.84%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=minimax-m325 Jun 2026livebench.aiT2 History
222#222Grok 4.6xAI76.78%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=grok-4.625 Jun 2026livebench.aiT2 History
223#223DeepSeek V4 Flash Vision ExpDeepSeek76.76%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=deepseek-v4-flash-vision-exp25 Jun 2026livebench.aiT2 History
224#224ox-alpha-max76.59%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=ox-alpha-max25 Jun 2026livebench.aiT2 History
225#225qwen3-8-27b-non-reasoningAlibaba Group76.59%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=qwen3.8-27b25 Jun 2026livebench.aiT2 History
226#226claude-opus-4-7-xhigh-effort76.53%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=claude-opus-4-7-xhigh-effort25 Jun 2026livebench.aiT2 History
227#227Muse Spark 1.2Meta AI76.46%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=muse-spark-1.2-xhigh25 Jun 2026livebench.aiT2 History
228#228Gemini 3.1 Pro Preview HighGoogle76.45%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gemini-3.1-pro-preview-high25 Jun 2026livebench.aiT2 History
229#229GLM 5.3 FlashZ.ai (Zhipu AI)76.4%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=glm-5.3-flash25 Jun 2026livebench.aiT2 History
230#230Grok Build 0.1xAI76.37%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=grok-build-0.125 Jun 2026livebench.aiT2 History
231#231Z.ai GLM 5.2Z.ai (Zhipu AI)76.24%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=glm-5.225 Jun 2026livebench.aiT2 History
232#232Claude 4.8 Opus Thinking Max EffortAnthropic76.22%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=claude-opus-4-8-max-effort25 Jun 2026livebench.aiT2 History
233#233Qwen3.8 FlashQwen76.19%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=qwen3.8-flash-next25 Jun 2026livebench.aiT2 History
234#234MiniMax M3MiniMax76.17%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=minimax-m325 Jun 2026livebench.aiT2 History
235#235GLM 5.3Z.ai (Zhipu AI)76.14%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=glm-5.325 Jun 2026livebench.aiT2 History
236#236claude-sonnet-4-6-thinking-auto-medium-effort76.1%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort25 Jun 2026livebench.aiT2 History
237#237Gemini 3.5 Flash-Lite HighGoogle76.07%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026livebench.aiT2 History
238#238gpt-5.2-2025-12-11-high76.07%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gpt-5.2-2025-12-11-high25 Jun 2026livebench.aiT2 History
239#239Claude Sonnet 5 xHigh EffortAnthropic76.04%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=claude-sonnet-5-xhigh-effort25 Jun 2026livebench.aiT2 History
240#240Gemini 3.8 FlashGoogle75.83%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gemini-3.8-flash-high25 Jun 2026livebench.aiT2 History
241#241Qwen3.6 PlusQwen75.83%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=qwen3.6-plus25 Jun 2026livebench.aiT2 History
242#242Grok 4.5xAI75.77%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=grok-4.525 Jun 2026livebench.aiT2 History
243#243Claude Fable 5 xHigh EffortAnthropic75.77%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=claude-fable-5-max-effort25 Jun 2026livebench.aiT2 History
244#244ox-alpha-max75.77%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=ox-alpha-max25 Jun 2026livebench.aiT2 History
245#245ox-alpha-max75.75%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=ox-alpha-max25 Jun 2026livebench.aiT2 History
246#246qwen3-8-27b-non-reasoningAlibaba Group75.69%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=qwen3.8-27b25 Jun 2026livebench.aiT2 History
247#247gemini-3.5-flash-high75.6%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gemini-3.5-flash-high25 Jun 2026livebench.aiT2 History
248#248gpt-6-astraOpenAI75.58%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gpt-6-astra-max25 Jun 2026livebench.aiT2 History
249#249Smaug Mini75.43%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=smaug-mini25 Jun 2026livebench.aiT2 History
250#250Gemini 3.6 Flash HighGoogle75.37%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026livebench.aiT2 History
251#251Muse Spark 1.1Meta AI75.3%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026livebench.aiT2 History
252#252qwen3-8-27b-non-reasoningAlibaba Group75.27%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=qwen3.8-27b25 Jun 2026livebench.aiT2 History
253#253Kimi K2.6 ThinkingMoonshot AI75.14%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=kimi-k2.6-thinking25 Jun 2026livebench.aiT2 History
254#254Qwen3.6 PlusQwen74.99%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=qwen3.6-plus25 Jun 2026livebench.aiT2 History
255#255DeepSeek V4 Flash (0731)DeepSeek74.98%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=deepseek-v4-flash-073125 Jun 2026livebench.aiT2 History
256#256Claude Sonnet 5 xHigh EffortAnthropic74.97%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=claude-sonnet-5-xhigh-effort25 Jun 2026livebench.aiT2 History
257#257Nemotron 3 UltraNVIDIA74.7%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=nemotron-3-ultra-550b-a55b25 Jun 2026livebench.aiT2 History
258#258Qwen3.8 FlashQwen74.64%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=qwen3.8-flash-next25 Jun 2026livebench.aiT2 History
259#259gemini-3.5-flash-high74.64%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gemini-3.5-flash-high25 Jun 2026livebench.aiT2 History
260#260gpt-5.2-2025-12-11-high74.64%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gpt-5.2-2025-12-11-high25 Jun 2026livebench.aiT2 History
261#261claude-opus-5-max-effort74.55%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=claude-opus-5-max-effort25 Jun 2026livebench.aiT2 History
262#262deepseek-v4-proDeepSeek74.54%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=deepseek-v4-pro25 Jun 2026livebench.aiT2 History
263#263Claude 4.6 Opus Thinking High EffortAnthropic74.52%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=claude-opus-4-6-thinking-auto-high-effort25 Jun 2026livebench.aiT2 History
264#264MiniMax M3MiniMax74.48%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=minimax-m325 Jun 2026livebench.aiT2 History
265#265claude-opus-4-5-20251101-thinking-64k-high-effort74.44%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=claude-opus-4-5-20251101-thinking-64k-high-effort25 Jun 2026livebench.aiT2 History
266#266qwen3-8-27b-non-reasoningAlibaba Group74.35%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=qwen3.8-27b25 Jun 2026livebench.aiT2 History
267#267Muse Spark 1.1Meta AI74.34%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026livebench.aiT2 History
268#268Muse Spark 1.2Meta AI74.33%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=muse-spark-1.2-xhigh25 Jun 2026livebench.aiT2 History
269#269Qwen3.8 FlashQwen74.24%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=qwen3.8-flash-next25 Jun 2026livebench.aiT2 History
270#270Qwen3.7 MaxQwen74.22%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=qwen3.7-max25 Jun 2026livebench.aiT2 History
271#271DeepSeek V4 Flash (0731)DeepSeek74.17%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=deepseek-v4-flash-073125 Jun 2026livebench.aiT2 History
272#272Qwen 3.8 MaxQwen74.08%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=qwen3.8-max25 Jun 2026livebench.aiT2 History
273#273Qwen3.7 MaxQwen74.04%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=qwen3.7-max25 Jun 2026livebench.aiT2 History
274#274GPT-5.2-CodexOpenAI73.98%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gpt-5.2-codex25 Jun 2026livebench.aiT2 History
275#275Kimi K2.7 CodeMoonshot AI73.96%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=kimi-k2.7-code25 Jun 2026livebench.aiT2 History
276#276Smaug Mini73.92%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=smaug-mini25 Jun 2026livebench.aiT2 History
277#277Grok 4.6xAI73.86%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=grok-4.625 Jun 2026livebench.aiT2 History
278#278Z.ai GLM 5.2Z.ai (Zhipu AI)73.74%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=glm-5.225 Jun 2026livebench.aiT2 History
279#279Gemini 3.5 Flash-Lite HighGoogle73.74%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026livebench.aiT2 History
280#280GPT-5.2-CodexOpenAI73.68%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gpt-5.2-codex25 Jun 2026livebench.aiT2 History
281#281Gemini 3.6 Flash HighGoogle73.59%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gemini-3.6-flash-high25 Jun 2026livebench.aiT2 History
282#282Grok 4.3xAI73.58%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=grok-4.325 Jun 2026livebench.aiT2 History
283#283gpt-5.6-lunaOpenAI73.56%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gpt-5.6-luna-max25 Jun 2026livebench.aiT2 History
284#284Inkling xHigh Effort73.46%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=inkling-xhigh25 Jun 2026livebench.aiT2 History
285#285Nemotron 3 UltraNVIDIA73.45%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=nemotron-3-ultra-550b-a55b25 Jun 2026livebench.aiT2 History
286#286Z.ai GLM 5.2Z.ai (Zhipu AI)73.16%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=glm-5.225 Jun 2026livebench.aiT2 History
287#287Qwen3.7 MaxQwen73.14%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=qwen3.7-max25 Jun 2026livebench.aiT2 History
288#288Grok 4.5xAI73.04%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=grok-4.525 Jun 2026livebench.aiT2 History
289#289claude-sonnet-4-6-thinking-auto-medium-effort72.99%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort25 Jun 2026livebench.aiT2 History
290#290Claude Fable 5.1 Max EffortAnthropic72.99%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=claude-fable-5-1-max-effort25 Jun 2026livebench.aiT2 History
291#291Qwen 3.8 MaxQwen72.87%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=qwen3.8-max25 Jun 2026livebench.aiT2 History
292#292Inkling xHigh Effort72.78%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=inkling-xhigh25 Jun 2026livebench.aiT2 History
293#293qwen3-8-27b-non-reasoningAlibaba Group72.66%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=qwen3.8-27b25 Jun 2026livebench.aiT2 History
294#294claude-opus-4-5-20251101-thinking-64k-high-effort72.58%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=claude-opus-4-5-20251101-thinking-64k-high-effort25 Jun 2026livebench.aiT2 History
295#295gpt-5.6-lunaOpenAI72.57%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gpt-5.6-luna-max25 Jun 2026livebench.aiT2 History
296#296Qwen3.8 FlashQwen72.55%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=qwen3.8-flash-next25 Jun 2026livebench.aiT2 History
297#297Muse Spark 1.1Meta AI72.55%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026livebench.aiT2 History
298#298Gemini 3.8 FlashGoogle72.49%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026livebench.aiT2 History
299#299Grok Build 0.1xAI72.46%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=grok-build-0.125 Jun 2026livebench.aiT2 History
300#300Claude 4.8 Opus Thinking Max EffortAnthropic72.03%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=claude-opus-4-8-max-effort25 Jun 2026livebench.aiT2 History

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.