Updated 5 h ago · first seen 11 Sept 2026
- Metric
- accuracy · % ↑
- Current results
- 1,639
- Models
- 376
- Current leader
- gpt-5.6-sol 65.9%
Score history · Olmo-3-7B-Think 2 rows
- Olmo-3-7B-Think
- 0.76%aa_slug=olmo-3-7b-think · variant=hard · evaluator=Artificial Analysis · reasoning=on12 Sept 2026
- 0.76%aa_slug=olmo-3-7b-think · variant=hard · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
Frontier over time · accuracy · variant=hard · evaluator=Artificial Analysis
10 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.
- 65.9%gpt-5.6-sol OpenAI Independent11 Sept 2026
- 61.4%gpt-5.6-sol OpenAI Independent11 Sept 2026
- 53.0%Claude Sonnet 4.6 Anthropic Independent11 Sept 2026
- 51.5%Claude Opus 4.7 Anthropic Independent11 Sept 2026
- 50.8%Z.ai GLM 5.2 Z.ai (Zhipu AI) Independent11 Sept 2026
- 49.2%KAT-Coder-Pro V2 Kwaipilot Independent11 Sept 2026
- 33.3%GPT-5.1-Codex Mini OpenAI Independent11 Sept 2026
- 26.5%Grok 4.3 xAI Independent11 Sept 2026
Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.
Leaderboard 317 models · trust independent-evaluator
Select models with +, then Compare.
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →