Updated 10 h ago · first seen 11 Sept 2026
- Metric
- — · % ↑
- Current results
- 0
- Models
- 0
- Current leader
- —
No results recorded for this benchmark yet — its sources are being connected. The definition, aliases and variants are kept so links resolve; nothing is fabricated.
Score history · a-x-k2 1 row
Not enough history to chart — a single observation (38.95% on 11 Sept 2026). Rows under different configurations count separately; the list below shows each one.
- 38.95%aa_slug=a-x-k2 · variant=v2.1 · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
Leaderboard 315 models
Select models with +, then Compare.
One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →