ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation
Published 16 Sept 2026arXiv:2609.15100
Updated 7 h ago · first seen 15 Sept 2026
paper_01M2JK19CJAAH8DSFM6BNC2K1M
Abstract
An image tool can change its underlying generator while retaining its public name, making version attribution from online posts ambiguous. We study this problem after the ChatGPT Images 2.5 launch. Our frozen collection contains 3,478 images from 2,440 posts across 8 sources. Recorded posting times fall within the first 51.1 hours after the announcement. It records three attribution tiers and retains standalone images after image-form filtering and targeted review. Caption claims and host records provide admission evidence, not independently verified generator identity. The observed content profile depends on the source mixture: NightCafe supplies 39.0% of images but 77.0% of CLIP-assigned fantasy scenes. We then evaluate six frozen detectors at thresholds calibrated to a 5% flag rate on reference photographs. Collection flag rates range from 3.7 to 56.4%, falling 42-81 percentage points below GenImage recall. Held-out artwork false-positive rates range from 1.5 to 96.5%, so a higher collection flag rate does not by itself establish better detection. An exploratory X-only comparison with our April collection finds a higher September flag rate for Effort, and a suggestive difference for DoU, under fixed-threshold post-clustered bootstrap intervals. Attribution, content and processing differences prevent a causal interpretation of these contrasts. The collection supports analysis of reported model use during a product transition, with source and attribution evidence retained for interpretation.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 4
- Property changedPaperChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation
ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv - Property changedPaperChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation
ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation: published at changed from 2026-09-15T04:00:00+00:00 to 2026-09-16T04:00:00+00:00
Published15 Sept 2026→16 Sept 2026arxiv - Property changedPaperChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation
ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv New paper: ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.