# FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

Source: [arXiv](https://arxiv.org/abs/2607.29602v1)  
Feed7 permalink: https://feed7.dev/p/2607-29602v1-02qoodp  
Published: 2026-07-31T16:33:39.000Z  
Trust: Needs Review (needs_review)

## Why Included

FriendBench shows why aggregate accuracy can hide behavioral bias: top multimodal models matched human panels overall but favored the “stranger” answer and gained less from video.

## Source Summary

FriendBench tests whether a pair is familiar or meeting for the first time using **20-second clips** from **96 balanced dyads**. It compares **26 models from seven companies** with matched human panels across text, audio, and video.

## Practical Implication

Builders evaluating socially aware multimodal systems should inspect class balance and channel-specific gains, not just headline accuracy. The strongest models matched the human crowd statistically while leaning toward “stranger,” indicating a different effective prior.

## Agent-Ready Context

FriendBench tests whether a pair is familiar or meeting for the first time using **20-second clips** from **96 balanced dyads**. It compares **26 models from seven companies** with matched human panels across text, audio, and video.

Builders evaluating socially aware multimodal systems should inspect class balance and channel-specific gains, not just headline accuracy. The strongest models matched the human crowd statistically while leaning toward “stranger,” indicating a different effective prior.

Only humans benefited from visible behavior beyond speech. The benchmark covers one constrained ice-breaker setup, so its findings do not establish how models handle broader social contexts.

## Connected Context

Feed7 judgment across 330 accumulated Signals:

FriendBench adds a controlled test of whether multimodal models infer a social relationship, while showing that human-level aggregate accuracy can conceal a different class prior and no measurable benefit from visible behavior. This reinforces the need to inspect modality contribution and error balance, not just headline scores, but narrows the conclusion to short, balanced ice-breaker interactions rather than general social understanding.

- [Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models](https://feed7.dev/p/2607-09654v1-0b5dedg) — Both show that near-human aggregate vision-language performance can coexist with different perceptual behavior, making error patterns and modality use important evaluation signals.
- [Evidence-Backed Video Question Answering](https://feed7.dev/p/2607-11862v1-18as4nc) — FriendBench finds that models did not gain from visible behavior; evidence-backed video QA offers a way to test whether predictions are actually grounded in tracked visual evidence rather than speech alone.
- [Evaling Video Slop — Maor Bril, Character.ai](https://feed7.dev/p/evaling-video-slop-maor-bril-character-ai-0cd76sd) — Both caution that video evaluation must measure information across time and modalities rather than treating visually plausible output or headline accuracy as sufficient.

## Context Map

- Layer: benchmark
- Domains: video, audio
- Topics: benchmark-integrity

## Uncertainty

- Only humans benefited from visible behavior beyond speech. The benchmark covers one constrained ice-breaker setup, so its findings do not establish how models handle broader social contexts.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
