Loading the frontier…
Loading the frontier…
Who leads, on what, since when — one observed dimension at a time. Every leader comes from a benchmark comparability group, a published price or a sourced attribute; the atlas never sums them into a single score.
computed 12 Sept 2026 07:03 UTC
Latest major models
New canonical model releases, newest first. Facts are the model's sourced attributes at crawl time.
| Model | Organization | Released | Key facts | Evidence |
|---|---|---|---|---|
| Fugu Ultra v2 | Sakana | 11 Sept 2026 | 1M context · image · text | openrouter.aiopenrouter(back-filled date) |
| Fugu Max | Sakana | 11 Sept 2026 | 1M context · image · text | openrouter.aiopenrouter(back-filled date) |
| DeepSeek-V4.1-FlashOpen weights | DeepSeek | 10 Sept 2026 | 1M context · image · text | api-docs.deepseek.comdeepseek(back-filled date) |
| Ling 3.0 Flash VLOpen weights | inclusionAI | 10 Sept 2026 | 131.1K context · image · text · video | openrouter.aiopenrouter(back-filled date) |
| North Small Translaterestricted-weights | Cohere | 10 Sept 2026 | 218B params · 16K context · text · CC-BY-NC-4.0 | cohere.comcohere |
| Mercury 2.5 | Inception | 8 Sept 2026 | 260K context · text | openrouter.aiopenrouter(back-filled date) |
| gpt-6-astraProprietary | OpenAI | 4 Sept 2026 | 1.05M context · image · text | artificialanalysis.aiartificial_analysis(back-filled date) |
| GPT-6 Astra Pro | OpenAI | 4 Sept 2026 | 1.05M context · image · text | openrouter.aiopenrouter(back-filled date) |
| Qwen3.8 Max (0902) | Qwen | 3 Sept 2026 | 1M context · image · text · video | openrouter.aiopenrouter(back-filled date) |
| Gemini 3.8 FlashProprietary | 2 Sept 2026 | 1.05M context · audio · document · image · text · video | artificialanalysis.aiartificial_analysis(back-filled date) | |
| Muse Spark 1.3Proprietary | Meta AI | 2 Sept 2026 | 1.05M context · audio · image · text · video | artificialanalysis.aiartificial_analysis(back-filled date) |
| Muse Spark 1.3 Contributor | Meta AI | 2 Sept 2026 | 1.05M context · audio · image · text · video | openrouter.aiopenrouter(back-filled date) |
Benchmark frontier
One row per benchmark with ≥ 20 current results: the primary comparability group, its leader, the second model and the gap between them, with the trust level of the leading row.
| Benchmark | Leader | Score | Second | Gap | Trust |
|---|---|---|---|---|---|
| Artificial Analysis Intelligence Indexindex · n=477 models | Claude Fable 5.1Anthropic | 53.37 | gpt-6-astra 52.81 | 0.56 | independent evaluator |
| GPQA Diamondaccuracy · variant=GPQA Diamond · evaluator=Artificial Analysis · n=459 models | gpt-6-astraOpenAI | 96.26 % | Gemini 3.8 Flash 95.25 | 1.01 | independent evaluator |
| Humanity's Last Examaccuracy · evaluator=Artificial Analysis · n=453 models | Claude Fable 5.1Anthropic | 59.13 % | Claude Fable 5 55.47 | 3.66 | independent evaluator |
| IFBenchaccuracy · evaluator=Artificial Analysis · n=333 models | Grok 4.3xAI | 83.33 % | grok-4-20-0309 82.93 | 0.4 | independent evaluator |
| τ²-benchpass^1 · evaluator=Artificial Analysis · n=323 models | Z.ai GLM 5.2Z.ai (Zhipu AI) | 99.12 % | jt-35b-flash 99.12 | 0 | independent evaluator |
| Terminal-Benchaccuracy · variant=hard · evaluator=Artificial Analysis · n=315 models | gpt-5.6-solOpenAI | 65.91 % | Claude Fable 5 62.88 | 3.03 | independent evaluator |
| MMMU-Proaccuracy · evaluator=Artificial Analysis · n=156 models | gpt-6-astraOpenAI | 86.88 % | Gemini 3.8 Flash 85.61 | 1.27 | independent evaluator |
| SciCodeaccuracy · evaluator=Artificial Analysis · n=126 models | Claude Fable 5.1Anthropic | 63.08 % | Claude Fable 5 61 | 2.08 | independent evaluator |
| Aider polyglotpass_rate_2 · n=61 models | gpt-5OpenAI | 86.7 % | o3-pro 84.9 | 1.8 | official benchmark |
| Aider polyglot — well-formed responsespercent_cases_well_formed · n=61 models | Codestral 25.01Mistral AI | 100 % | DeepSeek R1 + claude-3-5-sonnet-20241022 100 | 0 | official benchmark |
| LiveBenchglobal_average · n=57 models | Claude Fable 5.1Anthropic | 83.41 % | Claude Fable 5 82.97 | 0.44 | official benchmark |
| LiveBench Reasoningaverage score · variant=Reasoning · n=57 models | gpt-6-astraOpenAI | 92.65 % | Claude Fable 5.1 91.69 | 0.96 | official benchmark |
| LiveBench Codingaverage score · variant=Coding · n=57 models | Claude Fable 5.1Anthropic | 86.38 % | Claude Fable 5 85.99 | 0.38 | official benchmark |
| LiveBench Agentic Codingaverage score · variant=Agentic Coding · n=57 models | DeepSeek-V4.1-FlashDeepSeek | 77.27 % | Claude Fable 5.1 66.06 | 11.21 | official benchmark |
| LiveBench Mathematicsaverage score · variant=Mathematics · n=57 models | Claude Fable 5.1Anthropic | 97.01 % | gpt-6-astra 96.81 | 0.2 | official benchmark |
| LiveBench Data Analysisaverage score · variant=Data Analysis · n=57 models | gpt-6-astraOpenAI | 82.97 % | gpt-5.5 81.58 | 1.4 | official benchmark |
| LiveBench Languageaverage score · variant=Language · n=57 models | Claude Fable 5Anthropic | 90.68 % | Claude Fable 5.1 89.5 | 1.18 | official benchmark |
| LiveBench Instruction Followingaverage score · variant=IF · n=57 models | Gemini 3.8 FlashGoogle | 81.41 % | Gemini 3.7 Flash 79.93 | 1.49 | official benchmark |
| SWE-bench Verifiedresolved · board=Verified · system=mini-SWE-agent · n=42 models | Claude Opus 4.5Anthropic | 76.8 % | MiniMax M2.5 75.8 | 1 | official benchmark |
MethodLeader = best current row of the primary comparability group (canonical metric × task-defining configuration), one row per canonical model. Gap = leader − second in the metric's unit. /methodology
Price frontier
Among frontier models, the cheapest current published output price — and the cheapest with a context window of at least 1M tokens — with the provider that publishes it.
Cheapest frontier output
Ling 3.0 FlashinclusionAI
$0.063 output / 1M
input $0.021 · context 262.1K · via OpenRouter · observed 12 Sept 2026
Cheapest with ≥ 1M context
DeepSeek V4 Flash (0731)DeepSeek
$0.08 output / 1M
input $0.04 · context 1.05M · via DeepSeek API · observed 12 Sept 2026
Current offers by output price
575 current offers · USD per 1M tokens · output
Frontier universe: 358 canonical models (320 recent releases by active organizations · 88 top-10 on a benchmark · since 12 Sept 2025).
MethodFrontier models = canonical models released in the last 12 months by organizations with at least 3 canonical models, OR holding a top-10 rank in the primary comparability group of at least one benchmark. Artifacts and folded variants are excluded. No composite score is used to pick them. /methodology
Context frontier
Distinct canonical models with the largest sourced context_length. A routing endpoint counts as a model only if its provider publishes it as one.
Open-weight frontier
Observed dimensions only — best rank on any benchmark, parameters, context — for models whose weights can be downloaded. Sorted by best rank, then parameters.
MethodDimensions: best_rank, parameter_count, context_length. sorted by best benchmark rank then parameters; no composite /methodology
Efficiency frontier
y = index in the primary comparability group; x = cheapest current output price (USD / 1M tokens). Amber points form the Pareto set (higher score, lower price); the rest are dimmed.
187 models plotted · 6 on the Pareto frontier · group n=477
MethodPareto frontier maximises the score and minimises the cheapest current output price; exact ties are kept. Price = cheapest current offer across providers. A point's position is two observed facts, not a rating. /methodology
Agentic frontier
Top rows of the primary comparability group of each agentic benchmark (tool use, terminal tasks, multi-turn agents).
τ²-benchpass^1 · evaluator=Artificial Analysis · n=323
Terminal-Benchaccuracy · variant=hard · evaluator=Artificial Analysis · n=315
Multimodal frontier
Sourced modalities of models that hold a top-10 rank on at least one benchmark. Modality lists are attributes as published, not evaluations.
Recent frontier movements
Non-backfill benchmark events and price changes of at least 20 % over the last 30 days, newest first.
DeepSeek API changed pricing for DeepSeek V4 Flash (0731): $0.065 in / $0.18 out per 1M tokens → $0.04 in / $0.08 out per 1M tokens
OpenRouter changed pricing for Qwen3 235B A22B Instruct 2507: $0.22 in / $0.88 out per 1M tokens → $0.0875 in / $0.35 out per 1M tokens
OpenRouter changed pricing for Llama 3.1 70B Instruct: $0.4 in / $0.4 out per 1M tokens → $0.72 in / $0.72 out per 1M tokens
Moonshot AI Platform changed pricing for Kimi K3: $1.53515 in / $7.70013 out per 1M tokens → $2.30273 in / $11.5502 out per 1M tokens
DeepSeek API changed pricing for DeepSeek V4 Flash 0423: $0.08456 in / $0.16912 out per 1M tokens → $0.06762 in / $0.13524 out per 1M tokens
OpenRouter changed pricing for Hy3: $0.0825 in / $0.33 out per 1M tokens → $0.132 in / $0.528 out per 1M tokens
MethodFrontier models = canonical models released in the last 12 months by organizations with at least 3 canonical models, OR holding a top-10 rank in the primary comparability group of at least one benchmark. Artifacts and folded variants are excluded. No composite score is used to pick them. benchmark_frontier lists the primary comparability group of every benchmark with ≥ 20 current results; efficiency_frontier is the Pareto set (maximise quality score, minimise cheapest current output price); recent movements are non-backfill benchmark events and price moves ≥ 20% in 30 days. Nothing here is a composite ranking. /methodology
Composition and every threshold above are the API's (GET /frontier); this page adds no ranking of its own.