FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
Published 16 Sept 2026arXiv:2607.29602
Updated 12 h ago · first seen 15 Sept 2026
paper_01M2JK0TV7A243KZ6JN5RW8FGZ
Abstract
Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models favor ``stranger.'' This is a difference in effective prior, not in discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 2
- Property changedPaperFriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models: published at changed from 2026-09-15T04:00:00+00:00 to 2026-09-16T04:00:00+00:00
Published15 Sept 2026→16 Sept 2026arxiv - New paperPaperFriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
New paper: FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.