Skip to content
AI Atlas
BenchmarkActivecategory · general

LiveBench

livebench.ai

contamination-limited, monthly refreshed questions across 6 categories

quality57

Updated 7 h ago · first seen 11 Sept 2026

bench_01M293SPERF0D8QG7TSCGE8GGF

Metric
average score · %
Direction
Higher is better
Results
456
Leader
Claude Fable 5.1 Max Effort 97.01%

Score history · claude-sonnet-4-6-thinking-auto-medium-effort 8 rows

Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.

  • 63.22%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort25 Jun 2026
  • 76.1%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort25 Jun 2026
  • 77.95%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort25 Jun 2026
  • 86.99%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort25 Jun 2026
  • 42.63%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort25 Jun 2026
  • 79.27%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort25 Jun 2026
  • 84.77%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort25 Jun 2026
  • 72.99%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort25 Jun 2026

Back to the full leaderboard

Leaderboard 456 current results

Select models with +, then open Compare.

Leaderboard
#ModelScoreConfigEvaluatedSourceActions
301#301Inkling xHigh Effort71.92%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=inkling-xhigh25 Jun 2026livebench.aiT2 History
302#302Grok 4.6xAI71.87%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=grok-4.625 Jun 2026livebench.aiT2 History
303#303gpt-5-6-solOpenAI71.85%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gpt-5.6-sol-max25 Jun 2026livebench.aiT2 History
304#304Gemini 3.5 Flash-Lite HighGoogle71.82%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026livebench.aiT2 History
305#305Qwen3.7 MaxQwen71.79%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=qwen3.7-max25 Jun 2026livebench.aiT2 History
306#306Qwen3.6 27BQwen71.79%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=qwen3.6-27b25 Jun 2026livebench.aiT2 History
307#307Claude Sonnet 5 xHigh EffortAnthropic71.74%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=claude-sonnet-5-xhigh-effort25 Jun 2026livebench.aiT2 History
308#308gpt-5.4-miniOpenAI71.62%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gpt-5.4-mini-xhigh25 Jun 2026livebench.aiT2 History
309#309GLM 5.3 FlashZ.ai (Zhipu AI)71.59%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=glm-5.3-flash25 Jun 2026livebench.aiT2 History
310#310deepseek-v4-proDeepSeek71.57%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=deepseek-v4-pro25 Jun 2026livebench.aiT2 History
311#311Grok 4.5xAI71.53%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=grok-4.525 Jun 2026livebench.aiT2 History
312#312Kimi K3Moonshot AI71.36%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=kimi-k325 Jun 2026livebench.aiT2 History
313#313gpt-5.4-miniOpenAI71.32%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gpt-5.4-mini-xhigh25 Jun 2026livebench.aiT2 History
314#314Inkling xHigh Effort71.02%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=inkling-xhigh25 Jun 2026livebench.aiT2 History
315#315Smaug Agentic70.99%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=smaug-agentic25 Jun 2026livebench.aiT2 History
316#316DeepSeek V4 Flash Vision ExpDeepSeek70.96%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=deepseek-v4-flash-vision-exp25 Jun 2026livebench.aiT2 History
317#317gpt-5.4-miniOpenAI70.95%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gpt-5.4-mini-xhigh25 Jun 2026livebench.aiT2 History
318#318gpt-5.4-nanoOpenAI70.84%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gpt-5.4-nano-xhigh25 Jun 2026livebench.aiT2 History
319#319Grok 4.3xAI70.82%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=grok-4.325 Jun 2026livebench.aiT2 History
320#320Nemotron 3 UltraNVIDIA70.81%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=nemotron-3-ultra-550b-a55b25 Jun 2026livebench.aiT2 History
321#321Grok Build 0.1xAI70.79%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=grok-build-0.125 Jun 2026livebench.aiT2 History
322#322gpt-5.4-miniOpenAI70.79%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gpt-5.4-mini-xhigh25 Jun 2026livebench.aiT2 History
323#323gpt-5.5OpenAI70.73%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gpt-5.5-xhigh25 Jun 2026livebench.aiT2 History
324#324Nemotron 3 UltraNVIDIA70.7%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=nemotron-3-ultra-550b-a55b25 Jun 2026livebench.aiT2 History
325#325deepseek-v4-flashDeepSeek70.58%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=deepseek-v4-flash25 Jun 2026livebench.aiT2 History
326#326Kimi K2.6 ThinkingMoonshot AI70.54%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=kimi-k2.6-thinking25 Jun 2026livebench.aiT2 History
327#327Qwen3.6 27BQwen70.43%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=qwen3.6-27b25 Jun 2026livebench.aiT2 History
328#328Qwen3.6 27BQwen70.28%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=qwen3.6-27b25 Jun 2026livebench.aiT2 History
329#329GLM 5.3Z.ai (Zhipu AI)70.24%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=glm-5.325 Jun 2026livebench.aiT2 History
330#330gpt-5.4OpenAI70.22%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gpt-5.4-xhigh25 Jun 2026livebench.aiT2 History
331#331deepseek-v4-flashDeepSeek70.12%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=deepseek-v4-flash25 Jun 2026livebench.aiT2 History
332#332Inkling xHigh Effort70.1%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=inkling-xhigh25 Jun 2026livebench.aiT2 History
333#333DeepSeek-V4.1-FlashDeepSeek70.03%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=deepseek-v4.1-flash-max25 Jun 2026livebench.aiT2 History
334#334deepseek-v4-proDeepSeek69.99%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=deepseek-v4-pro25 Jun 2026livebench.aiT2 History
335#335Grok 4.3xAI69.93%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=grok-4.325 Jun 2026livebench.aiT2 History
336#336Qwen3.6 PlusQwen69.91%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=qwen3.6-plus25 Jun 2026livebench.aiT2 History
337#337Claude 4.6 Opus Thinking High EffortAnthropic69.89%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=claude-opus-4-6-thinking-auto-high-effort25 Jun 2026livebench.aiT2 History
338#338Muse Spark 1.1Meta AI69.64%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=muse-spark-1.1-xhigh25 Jun 2026livebench.aiT2 History
339#339gpt-5.4-nanoOpenAI69.58%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gpt-5.4-nano-xhigh25 Jun 2026livebench.aiT2 History
340#340GLM 5.3Z.ai (Zhipu AI)69.3%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=glm-5.325 Jun 2026livebench.aiT2 History
341#341Smaug Flash69.26%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=smaug-flash25 Jun 2026livebench.aiT2 History
342#342ox-alpha-max69.24%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=ox-alpha-max25 Jun 2026livebench.aiT2 History
343#343deepseek-v4-flashDeepSeek69.23%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=deepseek-v4-flash25 Jun 2026livebench.aiT2 History
344#344Qwen3.6 PlusQwen68.91%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=qwen3.6-plus25 Jun 2026livebench.aiT2 History
345#345Grok 4.5xAI68.59%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=grok-4.525 Jun 2026livebench.aiT2 History
346#346Kimi K2.7 CodeMoonshot AI68.41%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=kimi-k2.7-code25 Jun 2026livebench.aiT2 History
347#347MiniMax M3MiniMax68.2%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=minimax-m325 Jun 2026livebench.aiT2 History
348#348DeepSeek V4 Flash Vision ExpDeepSeek68.2%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=deepseek-v4-flash-vision-exp25 Jun 2026livebench.aiT2 History
349#349deepseek-v4-flashDeepSeek68.02%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=deepseek-v4-flash25 Jun 2026livebench.aiT2 History
350#350Gemini 3.7 FlashGoogle67.96%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gemini-3.7-flash-high25 Jun 2026livebench.aiT2 History
351#351Grok Build 0.1xAI67.78%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=grok-build-0.125 Jun 2026livebench.aiT2 History
352#352DeepSeek-V4-Pro-0813DeepSeek67.7%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=deepseek-v4-pro-081325 Jun 2026livebench.aiT2 History
353#353gpt-5.4-nanoOpenAI67.64%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gpt-5.4-nano-xhigh25 Jun 2026livebench.aiT2 History
354#354Nemotron 3 UltraNVIDIA67.36%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=nemotron-3-ultra-550b-a55b25 Jun 2026livebench.aiT2 History
355#355MiniMax M3MiniMax67.26%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=minimax-m325 Jun 2026livebench.aiT2 History
356#356Gemini 3.5 Flash-Lite HighGoogle67.24%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026livebench.aiT2 History
357#357gpt-5.4-nanoOpenAI67.21%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gpt-5.4-nano-xhigh25 Jun 2026livebench.aiT2 History
358#358claude-opus-4-7-xhigh-effort66.74%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=claude-opus-4-7-xhigh-effort25 Jun 2026livebench.aiT2 History
359#359GPT-5.2-CodexOpenAI66.45%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gpt-5.2-codex25 Jun 2026livebench.aiT2 History
360#360gpt-5.4-miniOpenAI66.37%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gpt-5.4-mini-xhigh25 Jun 2026livebench.aiT2 History
361#361ox-alpha-max66.14%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=ox-alpha-max25 Jun 2026livebench.aiT2 History
362#362Claude Fable 5.1 Max EffortAnthropic66.06%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=claude-fable-5-1-max-effort25 Jun 2026livebench.aiT2 History
363#363Claude 4.8 Opus Thinking Max EffortAnthropic66.03%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=claude-opus-4-8-max-effort25 Jun 2026livebench.aiT2 History
364#364DeepSeek V4 Flash (0731)DeepSeek65.52%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=deepseek-v4-flash-073125 Jun 2026livebench.aiT2 History
365#365deepseek-v4-flashDeepSeek65.48%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=deepseek-v4-flash25 Jun 2026livebench.aiT2 History
366#366Grok Build 0.1xAI65.39%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=grok-build-0.125 Jun 2026livebench.aiT2 History
367#367Grok Build 0.1xAI65.22%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=grok-build-0.125 Jun 2026livebench.aiT2 History
368#368claude-opus-5-max-effort65.2%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=claude-opus-5-max-effort25 Jun 2026livebench.aiT2 History
369#369Kimi K2.6 ThinkingMoonshot AI65.13%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=kimi-k2.6-thinking25 Jun 2026livebench.aiT2 History
370#370DeepSeek V4 Flash Vision ExpDeepSeek65.1%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=deepseek-v4-flash-vision-exp25 Jun 2026livebench.aiT2 History
371#371gemini-3.5-flash-high64.86%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gemini-3.5-flash-high25 Jun 2026livebench.aiT2 History
372#372Qwen 3.8 MaxQwen64.65%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=qwen3.8-max25 Jun 2026livebench.aiT2 History
373#373Smaug Agentic64.65%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=smaug-agentic25 Jun 2026livebench.aiT2 History
374#374gpt-5.6-terraOpenAI64.62%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gpt-5.6-terra-max25 Jun 2026livebench.aiT2 History
375#375Kimi K2.6 ThinkingMoonshot AI64.36%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=kimi-k2.6-thinking25 Jun 2026livebench.aiT2 History
376#376muse-spark-1-3-xhighMeta AI64.09%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=muse-spark-1.3-xhigh25 Jun 2026livebench.aiT2 History
377#377Qwen3.6 27BQwen64.03%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=qwen3.6-27b25 Jun 2026livebench.aiT2 History
378#378Gemini 3.5 Flash-Lite HighGoogle63.94%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gemini-3.5-flash-lite-high25 Jun 2026livebench.aiT2 History
379#379Claude Sonnet 5 xHigh EffortAnthropic63.86%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=claude-sonnet-5-xhigh-effort25 Jun 2026livebench.aiT2 History
380#380claude-opus-5-max-effort63.77%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=claude-opus-5-max-effort25 Jun 2026livebench.aiT2 History
381#381Claude 4.6 Opus Thinking High EffortAnthropic63.31%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=claude-opus-4-6-thinking-auto-high-effort25 Jun 2026livebench.aiT2 History
382#382Qwen3.6 27BQwen63.3%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=qwen3.6-27b25 Jun 2026livebench.aiT2 History
383#383claude-sonnet-4-6-thinking-auto-medium-effort63.22%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=claude-sonnet-4-6-thinking-auto-medium-effort25 Jun 2026livebench.aiT2 History
384#384deepseek-v4-flashDeepSeek63.14%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=deepseek-v4-flash25 Jun 2026livebench.aiT2 History
385#385Gemini 3.6 Flash HighGoogle63%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gemini-3.6-flash-high25 Jun 2026livebench.aiT2 History
386#386Grok 4.3xAI62.75%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=grok-4.325 Jun 2026livebench.aiT2 History
387#387Kimi K2.7 CodeMoonshot AI62.66%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=kimi-k2.7-code25 Jun 2026livebench.aiT2 History
388#388claude-opus-4-5-20251101-thinking-64k-high-effort62.55%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=claude-opus-4-5-20251101-thinking-64k-high-effort25 Jun 2026livebench.aiT2 History
389#389gpt-5.4-nanoOpenAI62.51%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gpt-5.4-nano-xhigh25 Jun 2026livebench.aiT2 History
390#390deepseek-v4-proDeepSeek62.35%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=deepseek-v4-pro25 Jun 2026livebench.aiT2 History
391#391Z.ai GLM 5.2Z.ai (Zhipu AI)62.29%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=glm-5.225 Jun 2026livebench.aiT2 History
392#392Grok 4.3xAI62.25%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=grok-4.325 Jun 2026livebench.aiT2 History
393#393Claude Fable 5 xHigh EffortAnthropic62.17%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=claude-fable-5-max-effort25 Jun 2026livebench.aiT2 History
394#394Kimi K3Moonshot AI62.17%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=kimi-k325 Jun 2026livebench.aiT2 History
395#395gpt-5.2-2025-12-11-high61.77%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gpt-5.2-2025-12-11-high25 Jun 2026livebench.aiT2 History
396#396Qwen3.8 FlashQwen61.62%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=qwen3.8-flash-next25 Jun 2026livebench.aiT2 History
397#397qwen3-8-27b-non-reasoningAlibaba Group61.36%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=qwen3.8-27b25 Jun 2026livebench.aiT2 History
398#398Smaug Flash61.06%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=smaug-flash25 Jun 2026livebench.aiT2 History
399#399GLM 5.3Z.ai (Zhipu AI)60.91%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=glm-5.325 Jun 2026livebench.aiT2 History
400#400Smaug Mini60.81%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=smaug-mini25 Jun 2026livebench.aiT2 History

Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.

The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.