The last 7 days in AI
Deterministic counters over events that occurred in the window — historical backfill excluded, so a re-crawl of old pages never inflates them. Hover a label for how each one is counted.
5 Sept 2026 → 12 Sept 2026 07:07 UTC
Window
All counters use is_backfill = false and occurred_at inside the window. Sparklines show the daily stats history where a series exists (2 days recorded so far).
MethodDeterministic counters over events that OCCURRED in the window and are not back-filled history; each counter carries its own definition. /methodology
Benchmark leadership
New benchmark leaders · 23
benchmarks whose primary-group leader (computed from results observed by each date) changed over the window
| Benchmark | Current leader | Score | Previous leader | Group | Trust |
|---|---|---|---|---|---|
| SWE-bench Verified | Claude Opus 4.5Anthropic | 76.8 resolved | first leader recorded | resolved · board=Verified · system=mini-SWE-agent · n=42 | official benchmark |
| Aider polyglot | gpt-5OpenAI | 86.7 pass_rate_2 | first leader recorded | pass_rate_2 · n=61 | official benchmark |
| LiveBench | Claude Fable 5.1Anthropic | 83.41 global_average | first leader recorded | global_average · n=57 | official benchmark |
| Humanity's Last Exam | Claude Fable 5.1Anthropic | 59.13 accuracy | first leader recorded | accuracy · evaluator=Artificial Analysis · n=453 | independent evaluator |
| Terminal-Bench | gpt-5.6-solOpenAI | 65.91 accuracy | first leader recorded | accuracy · variant=hard · evaluator=Artificial Analysis · n=315 | independent evaluator |
| Artificial Analysis Intelligence Index | Claude Fable 5.1Anthropic | 53.37 index | first leader recorded | index · n=477 | independent evaluator |
| SWE-bench Lite | claude-3-opusAnthropic | 4.33 resolved | first leader recorded | resolved · board=Lite · system=RAG baseline · n=6 | official benchmark |
| SWE-bench (full test split) | claude-3-opusAnthropic | 3.79 resolved | first leader recorded | resolved · board=Test · system=RAG baseline · n=6 | official benchmark |
| SWE-bench Multimodal | o3OpenAI | 35.98 resolved | first leader recorded | resolved · board=Multimodal · system=GUIRepair · n=4 | official benchmark |
| SWE-bench Multilingual | gemini-3-flashGoogle | 72.7 resolved | first leader recorded | resolved · board=Multilingual · system=mini-SWE-agent · n=13 | official benchmark |
| MMMU-Pro | gpt-6-astraOpenAI | 86.88 accuracy | first leader recorded | accuracy · evaluator=Artificial Analysis · n=156 | independent evaluator |
| SciCode | Claude Fable 5.1Anthropic | 63.08 accuracy | first leader recorded | accuracy · evaluator=Artificial Analysis · n=126 | independent evaluator |
| IFBench | Grok 4.3xAI | 83.33 accuracy | first leader recorded | accuracy · evaluator=Artificial Analysis · n=333 | independent evaluator |
| τ²-bench | Z.ai GLM 5.2Z.ai (Zhipu AI) | 99.12 pass^1 | first leader recorded | pass^1 · evaluator=Artificial Analysis · n=323 | independent evaluator |
| GPQA Diamond | gpt-6-astraOpenAI | 96.26 accuracy | first leader recorded | accuracy · variant=GPQA Diamond · evaluator=Artificial Analysis · n=459 | independent evaluator |
| Aider polyglot — well-formed responses | Codestral 25.01Mistral AI | 100 percent_cases_well_formed | first leader recorded | percent_cases_well_formed · n=61 | official benchmark |
| LiveBench Reasoning | gpt-6-astraOpenAI | 92.65 average score | first leader recorded | average score · variant=Reasoning · n=57 | official benchmark |
| LiveBench Coding | Claude Fable 5.1Anthropic | 86.38 average score | first leader recorded | average score · variant=Coding · n=57 | official benchmark |
| LiveBench Agentic Coding | DeepSeek-V4.1-FlashDeepSeek | 77.27 average score | first leader recorded | average score · variant=Agentic Coding · n=57 | official benchmark |
| LiveBench Mathematics | Claude Fable 5.1Anthropic | 97.01 average score | first leader recorded | average score · variant=Mathematics · n=57 | official benchmark |
| LiveBench Data Analysis | gpt-6-astraOpenAI | 82.97 average score | first leader recorded | average score · variant=Data Analysis · n=57 | official benchmark |
| LiveBench Language | Claude Fable 5Anthropic | 90.68 average score | first leader recorded | average score · variant=Language · n=57 | official benchmark |
| LiveBench Instruction Following | Gemini 3.8 FlashGoogle | 81.41 average score | first leader recorded | average score · variant=IF · n=57 | official benchmark |
Prices
Price changes · 20 · median -0.41 %
PRICE_CHANGED events in the window; % change = output price (else input) new vs old from the event's old/new values; median over events with both values
The pulse returns the count only; the rows below are PRICE_CHANGED events since 5 Sept 2026 from the change feed (same window, same backfill rule).
| Model | Provider | Input / 1M | Output / 1M | Change | Observed | Source |
|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0423DeepSeek | DeepSeek API | $0.067$0.067↓ | $0.134$0.134↓ | in −0.4%out −0.4% | 1 h ago | openrouter.ai |
| DeepSeek V4 Pro 0423DeepSeek | DeepSeek API | $0.829$0.819↓ | $1.66$1.64↓ | in −1.2%out −1.2% | 1 h ago | openrouter.ai |
| DeepSeek V4 Flash (0731)DeepSeek | DeepSeek API | $0.065$0.04↓ | $0.18$0.08↓ | in −38%out −56% | 1 h ago | openrouter.ai |
| DeepSeek V4 Flash 0423DeepSeek | DeepSeek API | $0.067$0.067↑ | $0.134$0.134↑ | in +0.4%out +0.4% | 2 h ago | openrouter.ai |
| DeepSeek V4 Flash 0423DeepSeek | DeepSeek API | $0.067$0.067↓ | $0.135$0.134↓ | in −0.8%out −0.8% | 3 h ago | openrouter.ai |
| DeepSeek V4 Pro 0423DeepSeek | DeepSeek API | $0.85$0.829↓ | $1.7$1.66↓ | in −2.4%out −2.4% | 3 h ago | openrouter.ai |
| Llama 3.1 70B InstructMeta AI | OpenRouter | $0.40$0.72↑ | $0.40$0.72↑ | in +80%out +80% | 6 h ago | openrouter.ai |
| Qwen3 235B A22B Instruct 2507Qwen | OpenRouter | $0.22$0.087↓ | $0.88$0.35↓ | in −60%out −60% | 6 h ago | openrouter.ai |
| MiniMax M2.5MiniMax | MiniMax API | $0.27$0.30↑ | $1.08$1.2↑ | in +11%out +11% | 6 h ago | openrouter.ai |
| DeepSeek V4 Flash 0423DeepSeek | DeepSeek API | $0.068$0.067↓ | $0.135$0.135↓ | in −0.4%out −0.4% | 6 h ago | openrouter.ai |
| DeepSeek V4 Pro 0423DeepSeek | DeepSeek API | $0.84$0.85↑ | $1.68$1.7↑ | in +1.2%out +1.2% | 6 h ago | openrouter.ai |
| Kimi K3Moonshot AI | Moonshot AI Platform | $1.54$2.3↑ | $7.7$11.55↑ | in +50%out +50% | 6 h ago | openrouter.ai |
| DeepSeek-V4-Pro-0813DeepSeek | DeepSeek API | $0.579$0.578↓ | $1.74$1.73↓ | in −0.2%out −0.2% | 6 h ago | openrouter.ai |
| Qwen3.8 27BQwen | OpenRouter | $0.42$0.214↓ | $3$2.55↓ | in −49%out −15% | 6 h ago | openrouter.ai |
| DeepSeek V4 Flash 0423DeepSeek | DeepSeek API | $0.085$0.068↓ | $0.169$0.135↓ | in −20%out −20% | 7 h ago | openrouter.ai |
| DeepSeek V4 Pro 0423DeepSeek | DeepSeek API | $0.948$0.84↓ | $1.9$1.68↓ | in −11%out −11% | 7 h ago | openrouter.ai |
| Hy3Tencent | OpenRouter | $0.083$0.132↑ | $0.33$0.528↑ | in +60%out +60% | 7 h ago | openrouter.ai |
| Kimi K3Moonshot AI | Moonshot AI Platform | $1.8$1.54↓ | $9.01$7.7↓ | in −15%out −15% | 7 h ago | openrouter.ai |
| DeepSeek V4 Flash 0423DeepSeek | DeepSeek API | $0.085$0.085↓ | $0.17$0.169↓ | in −0.3%out −0.3% | 9 h ago | openrouter.ai |
| DeepSeek V4 Pro 0423DeepSeek | DeepSeek API | $0.86$0.948↑ | $1.72$1.9↑ | in +10%out +10% | 9 h ago | openrouter.ai |
Movers are PRICE_CHANGED events emitted when a provider's published price for a model differs from the previous observation. Green = cheaper, red = dearer; % is relative to the previous published price.
Long context
New ≥ 1M-context models · 0
new_models whose context_length is at least 1 000 000 tokens
No new model with a ≥ 1M-token context in the last 7 days
- Fugu Ultra v2Sakana1M · 11 Sept 2026
- Fugu MaxSakana1M · 11 Sept 2026
- DeepSeek-V4.1-FlashDeepSeek1M · 10 Sept 2026
- gpt-6-astraOpenAI1.05M · 4 Sept 2026
- GPT-6 Astra ProOpenAI1.05M · 4 Sept 2026
- Qwen3.8 Max (0902)Qwen1M · 3 Sept 2026
- Gemini 3.8 FlashGoogle1.05M · 2 Sept 2026
- Muse Spark 1.3 ContributorMeta AI1.05M · 2 Sept 2026
Reading the pulse
Definitions
- New models 102
- canonical models with a NEW_MODEL event that occurred in the window (not back-filled)
- New open-weight models 84
- subset of new_models with openness open-weights / open-source
- New artifacts 174
- artifact entities (checkpoints, quantisations, conversions) first seen in the window
- New papers 418
- papers with a NEW_PAPER event that occurred in the window
- Provider listings 15
- PROVIDER_LISTED events in the window
- Delistings 0
- PROVIDER_DELISTED events in the window
- Price changes 20
- PRICE_CHANGED events in the window; % change = output price (else input) new vs old from the event's old/new values; median over events with both values
- New ≥ 1M-context models 0
- new_models whose context_length is at least 1 000 000 tokens
- New benchmark leaders 23
- benchmarks whose primary-group leader (computed from results observed by each date) changed over the window
- Documents changed 235
- DOCUMENT_CHANGED events in the window
- Sources observed 26
- distinct sources with at least one snapshot taken in the window
- Events total 1,432
- all non-backfill events (excluding source-document changes) that occurred in the window
Source: GET /pulse?days=7. Nothing on this page is a projection.