Sign InOpen Brain
arXivPaperNeeds Review

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

FriendBench shows why aggregate accuracy can hide behavioral bias: top multimodal models matched human panels overall but favored the “stranger” answer and gained less from video.

arXiv · Jul 31, 2026
Open Source Open MarkdownOpen JSON
Source Summary

FriendBench tests whether a pair is familiar or meeting for the first time using **20-second clips** from **96 balanced dyads**. It compares **26 models from seven companies** with matched human panels across text, audio, and video.

Practical Implication

Builders evaluating socially aware multimodal systems should inspect class balance and channel-specific gains, not just headline accuracy. The strongest models matched the human crowd statistically while leaning toward “stranger,” indicating a different effective prior.

Agent-Ready Context
FriendBench tests whether a pair is familiar or meeting for the first time using **20-second clips** from **96 balanced dyads**. It compares **26 models from seven companies** with matched human panels across text, audio, and video.

Builders evaluating socially aware multimodal systems should inspect class balance and channel-specific gains, not just headline accuracy. The strongest models matched the human crowd statistically while leaning toward “stranger,” indicating a different effective prior.

Only humans benefited from visible behavior beyond speech. The benchmark covers one constrained ice-breaker setup, so its findings do not establish how models handle broader social contexts.
Connected Context · Feed7 Judgment

FriendBench adds a controlled test of whether multimodal models infer a social relationship, while showing that human-level aggregate accuracy can conceal a different class prior and no measurable benefit from visible behavior. This reinforces the need to inspect modality contribution and error balance, not just headline scores, but narrows the conclusion to short, balanced ice-breaker interactions rather than general social understanding.

Context Map
benchmarkvideoaudio#benchmark-integrity
Uncertainty
Only humans benefited from visible behavior beyond speech. The benchmark covers one constrained ice-breaker setup, so its findings do not establish how models handle broader social contexts.