Skip to content
AI Atlas
Papercs.AI

Synthetic Data in Marketing Research: How to Evaluate and When to Trust

Published 15 Sept 2026arXiv:2609.13995

data quality89

Updated 29 h ago · first seen 15 Sept 2026

paper_01M2JK192Y1PQRWGGMQZ840G4G

Abstract

Debate over synthetic data in marketing research has polarized between claims that large language models (LLMs) make human respondents obsolete and calls to avoid them entirely. We argue that both positions obscure the more useful question: not whether synthetic respondents work, but when. Building on Brand, Israeli, and Ngwe (2026), we make three contributions. First, we distinguish three types of synthetic data (ungrounded LLM responses, segment-level personas, and individual-level digital twins) and map each to the decisions it can support. Second, we develop a taxonomy of four families of accuracy measures and suggest that the wide range of reported twin accuracy, from near-perfect to near-chance, largely reflects differences in what is being measured rather than in method quality. Aggregate measures often perform well even when little information is supplied to the LLM, and can mask a complete absence of respondent-level differentiation. Third, we introduce the forgotten question problem, in which a question is omitted from a fielded study, as a setting for twin-based augmentation of existing data. We propose an ex-ante answerability diagnostic that requires no ground truth: the R^2 of a random forest predicting twin outputs from the data used to construct the twins. Across 108 attitude questions from a nationally representative survey (N = 3,063), screening at R^2 above 0.7 raises the mean twin-human individual-level correlation by 15% and reduces the share of poorly answered questions from 25.9% to 4.3%. Embedding similarity and experienced-researcher judgment provide correlated but weaker screens.

Authors

Authors 2

Oded NetzerRajan Sambandam

Linked names open researcher pages (created from the paper's author list; name-only, no affiliation unless a source states it). Unlinked names have no researcher record yet.

Organizations

Organizations 0

No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.

Models

Models introduced or described 0

Inbound described_by relations from model cards and documentation.

No model links this paper yet

Model pages link papers through their model cards and documentation; the relation is written only when a source states it.

Datasets

Datasets used 0

No dataset relation recorded.

Benchmarks

Benchmarks used 0

No benchmark relation recorded.

Code

Repositories & frameworks 0

No repository linked.

Timeline

Timeline 1

Full timeline →

Sources

Sources 1

Source documents
SourceDocumentTypeTierLast observedSnapshots
arXiv (Atom API + RSS)rss.arxiv.org/rss/cs.AI feedT1· Official21 h ago4

Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.