AI Atlas as of 11 Sept 2026
Pick any date: the models that existed, their context lengths and status at the time, the provider prices valid that day, the benchmark leaders known by then, the hardware. Every value opens the claim that establishes it, with its validity interval.
ObservedThis date is inside the observation history: values are claims that were current that day.
Observation history starts . AI Atlas observation history starts at 2026-09-11T20:50:56.012288+00:00. The state is taken from claims, prices and results as they were known at the end of that UTC day.
Benchmark leaders
Leaders known by 11 Sept 2026 23
| Benchmark | Leader | Organization | Score | Metric · group | Trust | Models ranked |
|---|---|---|---|---|---|---|
| Aider polyglot | gpt-5 | OpenAI | 86.7 | pass_rate_2 | official-benchmark | 61 |
| Aider polyglot — well-formed responses | Codestral 25.01 | Mistral AI | 100 | percent_cases_well_formed | official-benchmark | 61 |
| Artificial Analysis Intelligence Index | Claude Fable 5.1 | Anthropic | 53.37 | index | independent-evaluator | 477 |
| GPQA Diamond | gpt-6-astra | OpenAI | 96.26 | accuracyaccuracy · variant=GPQA Diamond · evaluator=Artificial Analysis | independent-evaluator | 459 |
| Humanity's Last Exam | Claude Fable 5.1 | Anthropic | 59.13 | accuracyaccuracy · evaluator=Artificial Analysis | independent-evaluator | 453 |
| IFBench | Grok 4.3 | xAI | 83.33 | accuracyaccuracy · evaluator=Artificial Analysis | independent-evaluator | 333 |
| LiveBench | Claude Fable 5.1 | Anthropic | 83.41 | global_average | official-benchmark | 57 |
| LiveBench Agentic Coding | DeepSeek-V4.1-Flash | DeepSeek | 77.27 | average scoreaverage score · variant=Agentic Coding | official-benchmark | 57 |
| LiveBench Coding | Claude Fable 5.1 | Anthropic | 86.38 | average scoreaverage score · variant=Coding | official-benchmark | 57 |
| LiveBench Data Analysis | gpt-6-astra | OpenAI | 82.97 | average scoreaverage score · variant=Data Analysis | official-benchmark | 57 |
| LiveBench Instruction Following | Gemini 3.8 Flash | 81.41 | average scoreaverage score · variant=IF | official-benchmark | 57 | |
| LiveBench Language | Claude Fable 5 | Anthropic | 90.68 | average scoreaverage score · variant=Language | official-benchmark | 57 |
| LiveBench Mathematics | Claude Fable 5.1 | Anthropic | 97.01 | average scoreaverage score · variant=Mathematics | official-benchmark | 57 |
| LiveBench Reasoning | gpt-6-astra | OpenAI | 92.65 | average scoreaverage score · variant=Reasoning | official-benchmark | 57 |
| MMMU-Pro | gpt-6-astra | OpenAI | 86.88 | accuracyaccuracy · evaluator=Artificial Analysis | independent-evaluator | 156 |
| SWE-bench (full test split) | claude-3-opus | Anthropic | 3.79 | resolvedresolved · board=Test · system=RAG baseline | official-benchmark | 6 |
| SWE-bench Lite | claude-3-opus | Anthropic | 4.33 | resolvedresolved · board=Lite · system=RAG baseline | official-benchmark | 6 |
| SWE-bench Multilingual | gemini-3-flash | 72.7 | resolvedresolved · board=Multilingual · system=mini-SWE-agent | official-benchmark | 13 | |
| SWE-bench Multimodal | o3 | OpenAI | 35.98 | resolvedresolved · board=Multimodal · system=GUIRepair | official-benchmark | 4 |
| SWE-bench Verified | Claude Opus 4.5 | Anthropic | 76.8 | resolvedresolved · board=Verified · system=mini-SWE-agent | official-benchmark | 42 |
| SciCode | Claude Fable 5.1 | Anthropic | 63.08 | accuracyaccuracy · evaluator=Artificial Analysis | independent-evaluator | 126 |
| Terminal-Bench | gpt-5.6-sol | OpenAI | 65.91 | accuracyaccuracy · variant=hard · evaluator=Artificial Analysis | independent-evaluator | 315 |
| τ²-bench | Z.ai GLM 5.2 | Z.ai (Zhipu AI) | 99.12 | pass^1pass^1 · evaluator=Artificial Analysis | independent-evaluator | 323 |
One row per benchmark: the best current row of the primary comparability group among results observed by the date. See the benchmark's Frontier tab for the full leader history.
leaders from results observed by the date (evaluation dates are not used: a result is known only once observed)
Shareable: this URL reproduces the view. Per-entity: every entity page has a History tab with the same as-of reconstruction. What changed since? Diff 11 Sept 2026 → today. Methodology →