Loading timeline…
Loading timeline…
Every recorded change for this model by Anthropic, newest first. Back to the entity page →
9 events
New model: claude-4-sonnet-thinking (Anthropic)
claude-4-sonnet-thinking scores 31.06% on Terminal-Bench
claude-4-sonnet-thinking scores 36.33% on Terminal-Bench
claude-4-sonnet-thinking scores 64.62% on τ²-bench
claude-4-sonnet-thinking scores 61.79% on MMMU-Pro
claude-4-sonnet-thinking scores 54.69% on IFBench
claude-4-sonnet-thinking scores 10.7% on Humanity's Last Exam
claude-4-sonnet-thinking scores 77.68% on GPQA
claude-4-sonnet-thinking scores 18.92 on Artificial Analysis Intelligence Index
Showing up to 400 events. Narrow by year or category, or use the changes feed for cursor-paged history. Times are UTC.