Skip to content
AI Atlas
BenchmarkActivecategory · codingfamily · scicode · variant main

SciCode

scicode-bench.github.io

research-level scientific coding problems

data quality57

Updated 3 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
339
Models
129
Current leader
Claude Fable 5.1 63.1%

Frontier over time · accuracy · evaluator=Artificial Analysis

5 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 63.1%Claude Fable 5.1 Anthropic Independent11 Sept 2026
  2. 61%Claude Fable 5 Anthropic Independent11 Sept 2026
  3. 59.8%Gemini 3.7 Flash Google Independent11 Sept 2026
  4. 59.5%Kimi K3 Moonshot AI Independent11 Sept 2026
  5. 51.5%Claude Opus 5 Anthropic Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 129 models

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
100qwen3-32b-instructOpen weightsAlibaba Group · Qwen336%IndependentreasoningonPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
102DeepSeek V3Open weightsDeepSeek · DeepSeek · best of 2 rows35.8%IndependentreasoningoffPartially comparable-0.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
103gpt-oss-120bOpen weightsOpenAI · gpt-oss · best of 2 rows34.0%Independentreasoning_efforthighPartially comparable-1.97 ptobs. 12 Sept 2026artificialanalysis.aiT2History
104qwen3-30b-a3b-2507Open weightsAlibaba Group · Qwen3 · best of 2 rows33.0%IndependentreasoningonPartially comparable-3.01 ptobs. 12 Sept 2026artificialanalysis.aiT2History
105Devstral 2Open weightsMistral AI · Devstral 2 · best of 2 rows32.8%IndependentreasoningoffPartially comparable-3.25 ptobs. 12 Sept 2026artificialanalysis.aiT2History
106Devstral Small 2Open weightsMistral AI · Devstral · best of 2 rows32.4%IndependentreasoningoffPartially comparable-3.59 ptobs. 12 Sept 2026artificialanalysis.aiT2History
106Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows32.4%IndependentreasoningoffPartially comparable-3.59 ptobs. 12 Sept 2026artificialanalysis.aiT2History
108NVIDIA Nemotron 3.5 Lightning 30B A3BOpen weightsNVIDIA · Nemotron 3.5 · best of 2 rows32.1%IndependentreasoningonPartially comparable-3.94 ptobs. 12 Sept 2026artificialanalysis.aiT2History
109Llama 4 MaverickOpen weightsMeta AI · Llama 4 · best of 2 rows31.7%IndependentreasoningoffPartially comparable-4.29 ptobs. 12 Sept 2026artificialanalysis.aiT2History
110Granite 4.2 8BOpen weightsIBM · Granite 4.2 · best of 2 rows31.5%IndependentreasoningonPartially comparable-4.52 ptobs. 12 Sept 2026artificialanalysis.aiT2History
111Qwen3 14BOpen weightsQwen · Qwen330.7%Independentgroup defaultsPartially comparable-5.33 ptobs. 11 Sept 2026artificialanalysis.aiT2History
111qwen3-14b-instructOpen weightsAlibaba Group · Qwen330.7%IndependentreasoningonPartially comparable-5.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
113Nemotron 3 Nano 30B A3BOpen weightsNVIDIA · Nemotron 3 · best of 2 rows30.6%IndependentreasoningonPartially comparable-5.44 ptobs. 12 Sept 2026artificialanalysis.aiT2History
114Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows28.6%IndependentreasoningoffPartially comparable-7.41 ptobs. 12 Sept 2026artificialanalysis.aiT2History
115Qwen3 8BOpen weightsQwen · Qwen327.9%Independentgroup defaultsPartially comparable-8.11 ptobs. 11 Sept 2026artificialanalysis.aiT2History
115qwen3-8b-instructOpen weightsAlibaba Group · Qwen327.9%IndependentreasoningonPartially comparable-8.11 ptobs. 12 Sept 2026artificialanalysis.aiT2History
117Mistral Small 3.1Open weightsMistral AI · Mistral · best of 2 rows27.8%IndependentreasoningoffPartially comparable-8.22 ptobs. 12 Sept 2026artificialanalysis.aiT2History
118minicpm5-2bOpen weightsOpenBMB · best of 2 rows26.3%IndependentreasoningonPartially comparable-9.73 ptobs. 12 Sept 2026artificialanalysis.aiT2History
119Solar Pro 3ClosedUpstage · Solar · best of 2 rows25.5%IndependentreasoningonPartially comparable-10.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
120granite-4.2-3bOpen weightsIBM · Granite 4.2 · best of 2 rows25.4%IndependentreasoningonPartially comparable-10.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
121ling-3-0-tinyOpen weightsinclusionAI · best of 2 rows24.2%IndependentreasoningonPartially comparable-11.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
122Ministral 3 14BOpen weightsMistral AI · Ministral 3 · best of 2 rows23.8%IndependentreasoningoffPartially comparable-12.2 ptobs. 12 Sept 2026artificialanalysis.aiT2History
123Gemma 3 27BOpen weightsGoogle · Gemma 3 · best of 2 rows23.3%IndependentreasoningoffPartially comparable-12.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
124celeris-1ClosedCeleris · best of 2 rows21.6%IndependentreasoningoffPartially comparable-14.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
125Llama 4 ScoutOpen weightsMeta AI · Llama 4 · best of 2 rows21.3%IndependentreasoningoffPartially comparable-14.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
126Ministral 3 8BOpen weightsMistral AI · Ministral 3 · best of 2 rows20.7%IndependentreasoningoffPartially comparable-15.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
127Gemma 3 12BOpen weightsGoogle · Gemma 3 · best of 2 rows16.4%IndependentreasoningoffPartially comparable-19.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
128Ministral 3 3BOpen weightsMistral AI · Ministral 3 · best of 2 rows15.3%IndependentreasoningoffPartially comparable-20.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
129LFM2.5-2.6B (free)Open weightsLiquid AI · LFM2.5 · best of 2 rows14.3%IndependentreasoningonPartially comparable-21.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →