Skip to content
AI Atlas
PaperActive

How Proper Scoring Rules Shape LLM Forecasting

arxiv.org/abs/2608.28482

quality89

Updated 2 h ago · first seen 11 Sept 2026

paper_01M294FSHRV3A5N70ZHHYHXR5J

Published
11 Sept 2026
T1 · 2 h ago
arXiv
2608.28482
T1 · 2 h ago
Category
cs.LG
T1 · 2 h ago

Abstract

This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forecasts of resolved real-world events. Although the rules share the same theoretical incentive for truthful probability reporting, the resulting models differ in calibration, probability use, and estimated profiles of bias, information, and noise, with smaller differences in aggregate accuracy and discrimination. The Brier-trained model has the lowest observed Brier score and highest AUC-ROC, while the log-trained model has the highest observed log score and lowest calibration error. Models with similar aggregate performance also reach that performance through different combinations of bias, information, and noise. Proper scoring rules therefore need not behave interchangeably as training objectives. Reward choice may shape not only how well an LLM forecasts, but how its forecasting errors are structured.

Authors 5

Benjamin Turtel, Paul Wilczewski, Kris Skotheim, Ville A. Satop\"a\"a, Philip E. Tetlock

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

arXiv id
2608.28482

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Categories
cs.LG, cs.AI

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

2 h ago

Conflicts

None