Skip to content
AI Atlas
Time machine

AI Atlas as of 12 Sept 2026

Pick any date: the models that existed, their context lengths and status at the time, the provider prices valid that day, the benchmark leaders known by then, the hardware. Every value opens the claim that establishes it, with its validity interval.

ObservedThis date is inside the observation history: values are claims that were current that day.

Observation history starts . AI Atlas observation history starts at 2026-09-11T20:50:56.012288+00:00. The state is taken from claims, prices and results as they were known at the end of that UTC day.

Benchmark leaders

Leaders known by 12 Sept 2026 23

Benchmark leaders as of 2026-09-12
BenchmarkLeaderOrganizationScoreMetric · groupTrustModels ranked
Aider polyglotgpt-5OpenAI86.7pass_rate_2official-benchmark61
Aider polyglot — well-formed responsesCodestral 25.01Mistral AI100percent_cases_well_formedofficial-benchmark61
Artificial Analysis Intelligence IndexClaude Fable 5.1Anthropic53.37indexindependent-evaluator477
GPQA Diamondgpt-6-astraOpenAI96.26accuracyaccuracy · variant=GPQA Diamond · evaluator=Artificial Analysisindependent-evaluator459
Humanity's Last ExamClaude Fable 5.1Anthropic59.13accuracyaccuracy · evaluator=Artificial Analysisindependent-evaluator453
IFBenchGrok 4.3xAI83.33accuracyaccuracy · evaluator=Artificial Analysisindependent-evaluator334
LiveBenchClaude Fable 5.1Anthropic83.41global_averageofficial-benchmark57
LiveBench Agentic CodingDeepSeek-V4.1-FlashDeepSeek77.27average scoreaverage score · variant=Agentic Codingofficial-benchmark57
LiveBench CodingClaude Fable 5.1Anthropic86.38average scoreaverage score · variant=Codingofficial-benchmark57
LiveBench Data Analysisgpt-6-astraOpenAI82.97average scoreaverage score · variant=Data Analysisofficial-benchmark57
LiveBench Instruction FollowingGemini 3.8 FlashGoogle81.41average scoreaverage score · variant=IFofficial-benchmark57
LiveBench LanguageClaude Fable 5Anthropic90.68average scoreaverage score · variant=Languageofficial-benchmark57
LiveBench MathematicsClaude Fable 5.1Anthropic97.01average scoreaverage score · variant=Mathematicsofficial-benchmark57
LiveBench Reasoninggpt-6-astraOpenAI92.65average scoreaverage score · variant=Reasoningofficial-benchmark57
MMMU-Progpt-6-astraOpenAI86.88accuracyaccuracy · evaluator=Artificial Analysisindependent-evaluator157
SWE-bench (full test split)claude-3-opusAnthropic3.79resolvedresolved · board=Test · system=RAG baselineofficial-benchmark6
SWE-bench Liteclaude-3-opusAnthropic4.33resolvedresolved · board=Lite · system=RAG baselineofficial-benchmark6
SWE-bench Multilingualgemini-3-flashGoogle72.7resolvedresolved · board=Multilingual · system=mini-SWE-agentofficial-benchmark13
SWE-bench Multimodalo3OpenAI35.98resolvedresolved · board=Multimodal · system=GUIRepairofficial-benchmark4
SWE-bench VerifiedClaude Opus 4.5Anthropic76.8resolvedresolved · board=Verified · system=mini-SWE-agentofficial-benchmark42
SciCodeClaude Fable 5.1Anthropic63.08accuracyaccuracy · evaluator=Artificial Analysisindependent-evaluator129
Terminal-Benchgpt-5.6-solOpenAI65.91accuracyaccuracy · variant=hard · evaluator=Artificial Analysisindependent-evaluator317
τ²-benchZ.ai GLM 5.2Z.ai (Zhipu AI)99.12pass^1pass^1 · evaluator=Artificial Analysisindependent-evaluator323

One row per benchmark: the best current row of the primary comparability group among results observed by the date. See the benchmark's Frontier tab for the full leader history.

leaders from results observed by the date (evaluation dates are not used: a result is known only once observed)

Shareable: this URL reproduces the view. Per-entity: every entity page has a History tab with the same as-of reconstruction. What changed since? Diff 12 Sept 2026 → today. Methodology →