ConceptioArchivearXiv CS
arXiv CSopen access

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino, Antony D’Avirro, Benjamin Peloquin Fluid Concepts Research Correspondence: [email protected]

arXiv:2607.29602v1 [cs.CL] 31 Jul 2026

Abstract

between two people from a brief sample of how they interact. Specifically, we ask whether an observer can tell that two people are already familiar with one another rather than meeting as strangers, without being told and without relying on what the conversation is about. This judgment is a clean probe of behavioral social perception for two reasons. First, since we have ground-truth relationship labels, its answer is a matter of fact rather than interpretation: much related work in social intelligence targets intentions, emotions, or beliefs (Premack and Woodruff, 1978; Sap et al., 2019), whose correct answer is often contestable, whereas prior familiarity has an unambiguous ground truth (i.e., two people either have met before or they have not). Second, familiarity is often unstated in a short interaction, so it must be inferred from behavior. The task therefore rules out succeeding by reading semantic content alone, and prior work shows that humans make accurate social judgments from brief samples of behavior—so-called thin slices (Ambady and Rosenthal, 1992). We operationalize the task as F RIEND B ENCH: binary classification, familiar versus strangers, from a 20-second clip of a dyadic ice-breaker conversation, evaluated separately across text, audio, and video modalities. Every dyad, whether familiar or strangers, responds to the same type of ice-breaker prompt, so the conversational topic is matched across the two classes and cannot by itself reveal the answer, leaving the manner in which the two people interact as the primary signal. In addition to the core capability we benchmark, we examine how humans and models draw on the three modalities, and how each reaches its accuracy—by telling the two classes apart (discrimination) or by answering one class more often regardless of the evidence (response bias). Two patterns stand out. First, richer channels help both models and humans alike, but unequally: audio adds reliable signal over text for each, yet only

Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models lean toward “stranger”—a difference in effective prior, not discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions.

1

Introduction

Social intelligence, the capacity to make sense of other people and the situations they are in, is a central component of human cognition (Adolphs, 2003; Frith and Frith, 2007). It is also an increasingly important requirement for AI systems, which now observe, mediate, and participate in human interaction in roles ranging from companion agents to meeting assistants and care monitors (Mathur et al., 2024; Sap et al., 2019). Such systems must reason about the social situation people are in, and much of that reasoning rests on multimodal behavioral cues rather than explicit verbal content, such as how people coordinate, respond, and orient toward one another (Tickle-Degnen and Rosenthal, 1990). A system that reads only the semantic content of a conversation captures just part of what social understanding requires. We study one concrete instance of social intelligence: the capacity to recognize the relationship 1

2

humans gain a further reliable increment from the visual channel, while the strongest models are flat from audio to audiovisual—they under-exploit the visible behavior humans read. Second, humans stay close to balanced across both classes in every modality, whereas models show a larger and more idiosyncratic class bias, most clearly in the textonly condition. Both patterns concern how humans and models solve the task, not merely whether they succeed (§5).

Related Work

Thin-slice perception of familiarity. Human observers form accurate judgments about people and relationships from very brief behavioral exposure. Ambady and Rosenthal’s (1992) meta-analysis found that judgments from observations under five minutes (and often under 30 seconds) predicted objective outcomes at r ≈ .39, with longer exposure adding little. Familiarity in particular is legible in thin slices, with friendship the best-studied case: observers tell friends from strangers from silent video (Latif et al., 2014) or brief audio (Bryant et al., 2020), and cues such as inter-turn timing distinguish them (Templeton et al., 2023). Critically for our design, Dunbar et al. (2022) show that relationship quality remains inferable from speech even after its lexical content is digitally removed—the relational signal need not come from what is said. This motivates both our task and our 20-second window, which is ample by this evidence.

Contributions. • We introduce F RIEND B ENCH, a multimodal benchmark for inferring dyads’ prior familiarity (familiar vs. strangers) from ice-breaker clips in which every dyad answers the same type of prompt, built on the Seamless Interaction dataset (Agrawal et al., 2025).1 • We release a matched human-rater dataset for the task, covering 96 dyads in text, audio, and video, with roughly 90 raters per modality.

Recognizing relationships from behavior. One line of work predicts relationship type from images or video via supervised classification: PISC (Li et al., 2017) labels images as intimate, nonintimate, or no-relation, PIPA (Sun et al., 2017) annotates sixteen fine-grained relations in photo albums, and more recent work classifies asymmetric relations from the temporal dynamics of a live interaction (Tang et al., 2026). A parallel line infers relationships from the semantic content of dialogue—acoustic-lexical classifiers over phone calls (Katerenchuk et al., 2014) and language models over movie-script dialogue, from relationclassification datasets such as DDRel (Jia et al., 2021) to LLM evaluations where GPT-4o infers speaker relationships well above chance (Kim et al., 2026). Both differ from our task in two ways: they infer relationship type rather than the presence or absence of prior familiarity, and they lean on appearance, scene, or semantic content (especially for scripted dialogue). We instead fix the conversational prompt, controlling the semantic content these methods lean on.

• We evaluate 26 models from seven companies—OpenAI, Google, Anthropic, Alibaba, Mistral, Thinking Machines, and Meta—spanning proprietary and open-weight systems, across all three modalities. • We evaluate humans and models on the same stimuli under matched conditions. On accuracy, the best model and the human crowd are statistically indistinguishable in every modality. But equal accuracy is not human-like perception. The strongest models reach it by leaning toward “stranger,” while human raters stay balanced. Using signal detection theory, we characterize this as a difference in effective prior, not discrimination (§5). • We find a modality ordering both rater types obey only in part: text carries little signal for either and audio adds reliable discrimination, but the audiovisual channel gives humans a further reliable gain while leaving the strongest models flat—current models capture the vocal signal but under-exploit the visible behavior human observers read (§5.3).

Social reasoning in multimodal models. Multimodal models are increasingly tested for social understanding: Social-IQ (Zadeh et al., 2019) poses questions about social videos, while SIV-Bench (Kong et al., 2026) and PIVOTSBench (Zhang et al., 2026) probe reasoning about social scenes and fine-grained relations. HumanSense (Qin et al.,

1 The benchmark stimuli, human ratings, and model predictions are openly available at https://huggingface.co/ datasets/fluid-concepts/friend-bench.

2

2026) is the closest to our setting; one of its subtasks asks a model to judge how well two people in a video know each other. Three things set our benchmark apart: (1) every dyad answers the same type of prompt, so the topic itself cannot give the answer away; (2) the label is objective, recording whether a pair had actually met before rather than how close a viewer judges them; and (3) we gather matched human ratings in text, audio, and video for direct comparison. Our stimuli come from the Seamless Interaction corpus (Agrawal et al., 2025). Other dyadic corpora record acquaintance too—UDIVA (Palmero et al., 2021) labels each pair known or unknown, NoXi (Cafaro et al., 2017) rates how well partners know each other—but treat it as metadata, not as a label for prediction.

3

Methods

3.1

Benchmark Construction

basics for strangers—cues a shared hypothetical question largely suppresses. An EO interaction, though centered on one either-or question, typically continues for several minutes. We sample one 20-second clip per dyad from this interaction, drawn from either its early or late portion (counterbalanced across relationship and recording site), with boundaries snapped (±5s) to the nearest turn start so clips do not begin or end mid-utterance. The rendered video stimulus retains the clip’s audio track, so the video condition is audiovisual—visible behavior together with speech—whereas the audio and text conditions each isolate a single channel; both human raters and video models therefore receive sound as well as picture in the video condition. The final set fully crosses recording site, relationship (familiar/strangers), and gender composition (same-gender/mixed-gender pair), with 8 dyads per cell (96 total), and is participant-disjoint: no individual appears in more than one dyad. Crossing rules out recording site and gender composition as confounds (§4). Disjointness keeps each dyad an independent observation and guards against recognition leakage: because human raters see multiple stimuli in one session, a shared individual could be recognized from an earlier item, letting that recognition serve as the cue rather than the relationship itself.2 Interaction- and clip-level filters (minimum turns, in-window speaker balance, audio-track synchrony, interaction length, prompt completeness; thresholds in Appendix A, Table A1) exclude clips that cannot support the task regardless of relationship, e.g., one participant speaking for only a few seconds or desynchronized audio tracks. An additional LLM-based audit checks that each interaction’s recorded prompt label matches its content.

Samples are drawn from the Seamless Interaction dataset (Agrawal et al., 2025), restricted to naturalistic interactions from three recording sites (the dataset’s vendor field), each contributing a comparable share of dyads. We exclude the dataset’s improvised interactions, in which participants are assigned a relationship and asked to act it out, since our aim is to measure this task on real, unscripted behavior rather than acted behavior. Because the benchmark is used only for zero-shot evaluation and never for training, we pool dyads across all of the dataset’s predefined splits rather than restricting to one, maximizing the pool from which our quality filters and stratified design (below) can draw. Seamless Interaction sessions include many different interaction types; we use only the Either-Or (EO) ice-breaker, a short task in which one participant poses a forced-choice hypothetical question (e.g., “would you rather have the ability to fly or be invisible?”) and the dyad discusses their answer. We chose EO for three reasons: it standardizes the conversational context, since familiar and stranger pairs do the same thing and any signal must come from how a dyad responds rather than what it discusses; it is always the first interaction in a session, so every dyad is sampled from the same point in their interaction history; and its content carries little direct information about relationship status, so the task cannot be solved from semantic content alone. In free conversation, by contrast, status can leak through content—shared history and inside knowledge for familiars, getting-acquainted

3.2

Human Ratings

To establish baseline human performance, we collected human ratings of relationship type in a crowdsourced study. For each stimulus, raters made a forced choice: familiar or strangers. Before rating, each rater read a short instruction screen. It explained that they would read, listen to, or watch (depending on modality) a clip from a real conversation in which the two people were playing the “Either Or” ice-breaker game, then judge whether the pair had met before (familiar) or were meeting 2

This is not a risk for our models, which are evaluated zero-shot with no training process in which to learn identity.

3

for the first time (strangers). The full instruction text is reproduced in Appendix C. Raters were recruited via Prolific and completed the task on GORILLA. We ran three experiments, one per modality (text, audio, video), each with a panel of roughly 90 raters (94 text, 94 audio, 92 video). Prolific prescreening and quota matching ensured a gender-balanced pool of US nationals residing in the US, with no reported hearing difficulties or autism-spectrum diagnosis, and excluded anyone who had already participated in one of our other modality panels (see Appendix D). Raters followed a planned-missing, 6-block incomplete design (Graham et al., 2006), each rating one of six overlapping blocks and seeing 32 of the 96 dyads by design; individual stimuli were in turn rated by 25–35 raters.This design—rather than one rating per stimulus, averaged—lets us model raters themselves, estimating each rater’s own accuracy and response bias via crossed random effects, which requires each rater to contribute enough trials to be more than a one-shot data point. Each rater additionally completed the Social Information Processing subscale from the Tromsø Social Intelligence Scale (TSIS; Silvera et al., 2001; Grieve and Mahar, 2013) and a custom post-task self-report of which cues (visual, vocal, verbal, interactional) they attended to, enabling individualdifferences analysis (Appendix E). 3.3

d′

c

Rater

Acc.

Fam.

Str.

Text Human (indiv.) Human (crowd) gpt_text_mini

49.9 51.0 56.2

54.6 59.6 20.8

45.3 0.00 −0.12 42.6 0.05 −0.21 91.7 0.54 1.06

Audio Human (indiv.) 56.3∗∗ 57.2 Human (crowd) 63.5∗ 64.6 gemini_audio_pro 66.7∗∗ 39.6

55.5 0.32 −0.02 62.5 0.68 −0.03 93.8 1.21 0.86

Video Human (indiv.) Human (crowd) gemini_video

63.2 0.56 74.5 1.16 91.7 1.12

60.9∗∗∗ 58.7 71.9∗∗∗ 70.2 66.7∗∗ 41.7

0.06 0.06 0.77

Table 1: Human raters (average individual and majorityvote crowd) vs. the best model per modality. Acc. is accuracy; Fam. and Str. per-class recall (all %); d′ and c as in Table 2. The named model in each block is that modality’s best-performing (top-accuracy) system, scored over answered trials; human-indiv. accuracy is the per-rater mean (per-model Wilson CIs in Table 2). ∗ ∗∗ ∗∗∗ / / : above chance at .05/.01/.001, by exact binomial vs. 50% (crowd, best model) or 95%/99%/99.9% GLMM posterior HDI excluding 50% (individual rows).

ble 2); scoring refusals as errors would understate a model for declining rather than for misjudging. Coverage is near-total except for three audio models (§4), and the top models in each modality answer every trial, so this choice does not affect the headline comparison.

4

Model Predictions

Benchmark Results

Table 1 summarizes, per modality, the average perrater human, the human crowd (the per-stimulus majority vote of all raters who saw that stimulus), and the single best-performing model—their accuracy and, for the class-bias analysis in §5.1, their per-class recall and signal-detection statistics (Stanislaw and Todorov, 1999). Table 2 gives the complete per-model breakdown with Wilson 95% confidence intervals. Figure 1 depicts the same comparison, plotting every model against the average individual human rater and the majority-vote crowd in each modality. The best model and the human crowd trade the point-estimate lead across modalities. The model edges the crowd in text (by 5.2 points) and audio (by 3.2), while the crowd leads in video (by 5.2). Video is humans’ strongest modality, and the one place the crowd, not a model, comes out ahead. None of these three gaps is statistically reliable, though. On the 96 paired dyads, an exact McNemar test (Dietterich, 1998; McNemar, 1947) never

We evaluate 26 models from seven companies across text, audio, and video (Table 2), a fullycrossed design in which every model rates every stimulus; unlike the human panels, no incompleteblock correction is needed for the model results. Each model receives the same prompt and response parsing. The model prompt is adapted directly from the human instruction text, so that both rater types are given the same task framing, the same “Either Or” description, and the same familiar/strangers definition; the only difference is the response mechanism—a forced-choice button for humans, a JSON label with a confidence rating for models. Models and prompts are further described in Appendix B and Appendix C, respectively. A few audio models decline to answer on a share of trials (returning no usable familiar/strangers label) rather than committing to a forced choice. Because a refusal is not a wrong answer, we report each model’s accuracy over its answered (covered) trials and give per-model coverage alongside it (Ta4

best model

rater mean

Text

chance

Audio

Video

Models

Human crowd Human raters

20

40 60 80 Accuracy (%)

100 20

40 60 80 Accuracy (%)

100 20

40 60 80 Accuracy (%)

100

Figure 1: Human vs. model accuracy by modality. Each rater point covers only the ∼32 trials that rater saw, so the rater cloud is widened by sampling noise and its upper tail is not a band of expert raters (§5.4). Horizontal bars are Wilson 95% CIs for the best model and crowd, not tests of human–model or cross-modality differences (§4, §5.3).

approaches significance (p > .4 in every modality). Text is the weakest modality for humans: at chance individually and barely above it as a crowd. But the best text model is also not above chance (56.2%, p = .26). Text is therefore less a model win than a modality where neither rater type finds a signal.

5

Performance Analysis

5.1

Class Bias: Humans vs. Models

Humans and the best-performing models are close to parity on accuracy (§4), but do they reach it the same way? To separate discrimination from response bias, we read the per-class recall and signaldetection columns of Table 1, treating “familiar” as the signal class: d′ measures how well a rater tells the classes apart, and the criterion c how far its decision rule leans toward one answer (c > 0 toward “stranger,” c < 0 toward “familiar”). Two patterns stand out. Humans stay near-balanced everywhere: the human criterion is within 0.23 of zero in every modality, and per-class recall gaps never exceed 10 points for pooled raters. The best text and audio models instead reach their accuracy largely through response bias: gemini_audio_pro labels 93.8% of strangers correctly but recovers only 39.6% of familiar pairs (c = 0.86), and the best text model, gpt_text_mini, is more extreme still (91.7% vs. 20.8%, c = 1.06). Across the roster this is the rule: in text almost every model leans “stranger” by 50–85 recall points, and in audio the bias is large but inconsistent in direction (Table 2). The gap holds even in video, where models come closest to humans. There the best model (gemini_video) still leans “stranger” (c = 0.77; 91.7% of strangers but only 41.7% of familiar pairs), while the human crowd stays balanced (c = 0.06) at slightly higher discrimination (d′ = 1.16 vs. 1.12). The two reach comparable d′ , but only the crowd does so without skewing its decision rule. Models thus fail to recreate the human

Only three models are individually above chance at this sample size, all from Google: gemini_audio_pro and gemini_audio in audio, and gemini_video in video. No text model clears chance. Apart from the two top-accuracy systems, the best remaining model in each modality reaches only 56.2% (text, gpt_text_mini), 63.5% (audio, gemini_audio), and 58.3% (video, gemini_video_thinking). Most models cluster near chance on raw accuracy. The per-class recall and criterion columns already hint at the analysis in §5.1: several models carry very large response biases regardless of accuracy. In audio the bias is large but inconsistent in direction. voxtral_audio labels every dyad “familiar” (100%/0%, c = −2.32), while qwen_audio does the opposite (0%/100%, c = 2.19). In text most models lean hard toward “stranger,” with muse_spark_text_high recovering only 8.3% of familiar pairs (c = 1.40). A near-chance accuracy can thus conceal a strongly one-sided response pattern rather than balanced guessing. Three audio models also decline a substantial share of trials: gpt_audio answers only 45%, gpt_audio_mini 28%, and qwen_audio 71%. They sit at chance on the trials they do answer, so their behavior is better summarized as declining-plus-guessing. 5

Model

Acc.

95% CI

Cov.

Fam.

Str.

d′

c

Text gpt_text_mini gemini_text_thinking gemini_text gpt_text inkling_text claude_text muse_spark_text_high mistral_text muse_spark_text

56.2 55.8 55.2 53.1 52.1 52.1 51.0 50.0 49.0

[46.3, 65.7] [45.8, 65.4] [45.3, 64.8] [43.2, 62.8] [42.2, 61.8] [42.2, 61.8] [41.2, 60.8] [40.2, 59.8] [39.2, 58.8]

100 99 100 100 100 100 100 100 100

20.8 27.7 25.0 29.2 16.7 14.6 8.3 87.5 10.4

91.7 83.3 85.4 77.1 87.5 89.6 93.8 12.5 87.5

0.54 0.36 0.36 0.19 0.17 0.19 0.14 0.00 −0.10

1.06 0.76 0.84 0.63 1.03 1.12 1.40 −1.11 1.16

Audio gemini_audio_pro gemini_audio muse_spark_audio gemini_audio_thinking muse_spark_audio_high gpt_audio_1_5 qwen_audio inkling_audio voxtral_audio gpt_audio gpt_audio_mini qwen_omni

66.7∗∗ [56.8, 75.3] 63.5∗ [53.6, 72.5] 57.3 [47.3, 66.7] 55.2 [45.3, 64.8] 53.1 [43.2, 62.8] 53.1 [43.2, 62.8] 51.5 [39.8, 62.9] 51.0 [41.2, 60.8] 50.0 [40.2, 59.8] 48.8 [34.6, 63.2] 48.1 [30.7, 66.0] 47.9 [38.2, 57.8]

100 100 100 100 100 100 71 100 100 45 28 100

39.6 81.2 20.8 64.6 18.8 87.5 0.0 41.7 100.0 100.0 25.0 16.7

93.8 45.8 93.8 45.8 87.5 18.8 100.0 60.4 0.0 4.3 81.8 79.2

1.21 0.76 0.67 0.26 0.25 0.25 0.02 0.05 0.00 0.45 0.18 −0.15

0.86 −0.48 1.13 −0.23 0.99 −0.99 2.19 0.23 −2.32 −1.76 0.72 0.87

Video gemini_video gemini_video_thinking gemini_video_pro muse_spark_video muse_spark_video_high

66.7∗∗ [56.8, 75.3] 58.3 [48.3, 67.7] 57.3 [47.3, 66.7] 54.2 [44.2, 63.8] 53.1 [43.2, 62.8]

100 100 100 100 100

41.7 27.1 18.8 14.6 10.4

91.7 89.6 95.8 93.8 95.8

1.12 0.62 0.77 0.44 0.42

0.77 0.91 1.25 1.24 1.42

Table 2: Per-model results, sorted within modality. Acc. is accuracy (%) over answered trials and CI its Wilson 95% confidence interval; Cov. is coverage (% answered rather than declined). Fam. and Str. are per-class recall (%): share of familiar and of stranger dyads correctly labeled. d′ and criterion c treat “familiar” as the signal class (log-linear corrected; Hautus, 1995), so c > 0 indicates a bias toward “stranger” and c < 0 toward “familiar.” ∗∗ /∗ : accuracy above chance at p < .01/p < .05 (exact binomial vs. 50%).

balanced-recall pattern in any modality, including their strongest: several of the highest-accuracy systems (Table 2) buy that accuracy with a skewed criterion rather than sharper discrimination. 5.2

test the modality ordering while accounting for the repeated-measures structure of the human data, we fit a Bayesian crossed random-effects logistic model (Gelman et al., 2014) to per-trial human correctness (Bernoulli likelihood, logit link), with a fixed effect of modality and crossed random intercepts for rater and for dyad (§3.2); this is the design the block structure was built to support. We estimate it with the bambi interface to PyMC (Capretto et al., 2022), using its default weakly-informative, data-scaled priors (wide normal priors on the fixed effects and half-normal hyperpriors on the randomeffect standard deviations), and sample the posterior with NUTS (4 chains, 1,000 warmup and 1,000 post-warmup draws each; all R̂ ≈ 1.00). We report population-averaged (marginal) accuracies and contrasts as posterior means with 95% highest-density intervals (HDIs, Makowski et al., 2019). The marginal accuracies confirm the raw pattern—text 50.0% (95% HDI [46.7, 53.6]), audio 56.6% [52.2, 60.6], video 60.9% [56.9, 64.9]—and

Wisdom of the Crowd

Pooling human raters via majority vote barely moves accuracy in text (49.9%→51.0%, +1.1 points) but raises it substantially in audio (56.3%→63.5%, +7.2 points) and video (60.9%→71.9%, +11.0 points). Figure 2 shows the full crowd-growth curves: majority-vote accuracy rises with crowd size k in audio and video but stays flat near chance in text. The audio and video gains therefore reflect real signal, not an artifact of pooling. 5.3

Modality Comparison

Text is the weakest modality for both humans and models, and richer channels help both—though, as we show below, not to the same degree. To 6

50

That text is the hardest modality is consistent with the benchmark’s design rather than a defect of it. The EO ice-breaker was chosen precisely so that transcript content carries little direct information about relationship status (§3.1); the weak transcript performance of both humans and models is the expected consequence of that choice. It also helps explain why response bias is most visible in text: with little signal available, a rater’s responses reflect its bias more than the stimulus.

45

5.4

Majority-vote accuracy (%)

Text (50→51%) Audio (56→64%)

Video (61→72%) chance

70 65 60 55

1 5 10 15 20 25 Crowd size k (raters pooled per stimulus)

Rater Individual Differences

Because each rater contributes many trials by design (§3.2), we can estimate individual accuracy and relate it to rater traits. Trait social intelligence (TSIS-PS scale; reliable in every panel, α = 0.88– 0.92) is essentially uncorrelated with accuracy in the two modalities where the task is doable: audio r = 0.03 (p = .78) and video r = −0.12 (p = .24). The only significant association is in text (r = 0.29, p = .004), the modality where average accuracy is at chance, so we read it with caution. Self-reported cue use is similarly flat: raters most often report attending to interactional cues (rapport, responsiveness), but cue-use scores rarely predict accuracy (Appendix E). The individual-rater cloud in Figure 1 is correspondingly wide, with the best raters near 75% in every modality. Its spread should not be read as a stable band of expert observers, though. Each rater’s accuracy comes from only the ∼32 trials they saw, so the upper tail is close to what sampling noise alone would produce. The best model therefore sits within the human distribution, not above it, even as almost no individual beats the aggregated crowd (1 of 92 in video). The crossed random-effects model makes the point directly: the dyad random-intercept SD (0.67–0.95 across modalities, logit scale) is six to seven times the rater SD (0.10–0.14). Who the rater is thus matters far less than which dyad they were rating.

Figure 2: Wisdom of the crowd. Majority-vote accuracy as a function of human crowd size k, with raters sampled per stimulus from that stimulus’s own rater panel (mean ± SD over 300 bootstrap draws). Legend gives singlerater → full-panel crowd accuracy per modality.

the pairwise contrasts are decisive: audio exceeds text by 6.5 points (HDI [4.1, 8.9]), video exceeds audio by 4.5 points ([1.9, 6.8]), and video exceeds text by 10.9 points ([8.5, 13.3]), each with posterior P > 0 of at least 0.999. Text alone is statistically indistinguishable from chance (posterior probability of exceeding 50% only 0.50). Because the video condition is audiovisual (§3.1), its edge over audio reflects the added value of visible behavior on top of speech, not vision in isolation; video is best read as an audiovisual upper bound rather than a vision-only channel. The model side matches the human floor but not the human ceiling. The best model climbs from text (56.2%) to audio (66.7%) but gains nothing from the audiovisual channel (66.7%), and among models evaluated in both audio and video only Gemini Flash improves (63.5% → 66.7%) while Gemini Pro (66.7% → 57.3%) and Muse Spark (57.3% → 54.2%) decline. (Cross-modality model means rise monotonically—52.7, 53.9, 57.9% for text, audio, video—but the video roster is small and skewed toward stronger companies, so we read the best-model and within-model trajectories rather than the pooled mean.) Where humans reliably convert visible behavior into accuracy on top of speech, then, the strongest models do not: they capture the vocal signal but under-exploit the visual channel, the one place a benchmark of multimodal social perception most expects a model to gain.

5.5

Dyad Difficulty

Treating the per-dyad random intercepts from the crossed random-effects model as Rasch-style itemeasiness parameters (Rasch, 1960; de Boeck and Wilson, 2004), we find that difficulty is largely a property of the conversation, not of the modality through which it is observed. Estimated dyad easiness correlates positively across all modality pairs (text–audio r = 0.52, text–video r = 0.33, audio– video r = 0.61; all p < .01): a dyad that is hard to read from one modality tends to be hard from the 7

others as well. A handful of dyads sit below chance in every modality, acting as systematic “lures” that most raters misread alike (Appendix F). Humans and models tend to find the same dyads hard, though how strongly depends on how it is measured. Correlating per-dyad human-crowd accuracy with per-dyad accuracy pooled over all models yields r = 0.56 (audio) and r = 0.34 (video), both significant (p < .01; Appendix Figure A1), but only r = 0.20 in text (p = .046). This raw correlation understates the shared difficulty. It blends two things: how hard a dyad is to read, and which class it belongs to. Humans stay balanced while the models lean toward “stranger” (§5.1), so the two rater types tend to miss different classes. Isolating difficulty from this class split, by correlating within each true class, brings the agreement out clearly: r = 0.61 (text), 0.58 (audio), and 0.44 (video), all p < .001. So a substantial part of what makes a pair legible or illegible is a property of the interaction shared across observers, even though the two reach their answers by different rules.

6

The difference is in how humans and models use their two answers. Discrimination (d′ ) measures how well a rater tells the classes apart. Response bias (criterion c; Table 1) measures how far it leans toward one answer. Human raters stay balanced across the two answers in every modality. The strongest models instead lean toward “stranger,” giving that answer more often regardless of the pair. Video makes this clearest. There the best model tells the classes apart about as well as the crowd (d′ ≈ 1.1), but it reaches that accuracy by leaning toward “stranger” where the crowd stays balanced. In signal-detection terms, this is a difference in effective prior: the models behave as though strangers were the more common answer and demand more evidence before saying “familiar.” Humans behave as though the classes were equally likely, which is true in our balanced set. This describes the response distribution, not a mechanism, and the prior is inferred from behavior rather than verified. A pure criterion shift is correctable: recentering to the known base rate would raise accuracy with no gain in discrimination. So a biased model’s raw accuracy can understate its discrimination.

Discussion

For a benchmark of multimodal social perception, the clearest result is about the modalities themselves. Text carries little relational signal for anyone, by design: the shared ice-breaker prompt suppresses lexical content (§3.1). Richer channels help both humans and models, but asymmetrically. Humans improve reliably at every step, from text to audio to audiovisual (§5.3), gaining a credible increment from visible behavior on top of speech. The strongest models capture the audio gain but not the visual one: their top accuracy is identical in audio and audiovisual (66.7%). The sharpest human–model difference is thus not whether models can read relationships but whether they exploit the visual channel as humans do. This gap has applied stakes. The companion agents, meeting assistants, and care monitors that motivate the task must read social situations from behavior, and the visible cues humans exploit on top of speech are exactly what current models miss. In terms of overall accuracy, the best models have drawn level with an aggregated human crowd. In every modality the two are statistically indistinguishable. Point estimates put the model slightly ahead in text and audio and the crowd slightly ahead in video (Table 1), but no gap survives a paired McNemar test (p > .4 throughout; §4).

At the same time, humans and models are not perceiving unrelated things. Dyad difficulty is partially shared and partially transfers across modalities (§5.5), so part of what makes a pair legible or illegible is a property of the interaction itself, not of the observer or channel. The dissociation is therefore specific: the two agree on which dyads are hard but diverge on the decision rule they apply. Methodologically, these results argue for evaluating social-perception models with more than a single accuracy number: per-class recall, signaldetection statistics, and item-level difficulty each revealed structure that accuracy alone obscured, and are cheap to report once the underlying predictions are available. Taken together, the strongest models now match an aggregated human crowd on accuracy in this task, but not on how they reach it. They find the same conversations hard. Yet they systematically lean toward “stranger,” where human raters stay balanced, and they do not successfully convert the visual channel into accuracy. Whether models can be brought to read relationships as humans do— balanced across classes, and drawing on visible behavior rather than speech alone—is the open question this benchmark is built to track. 8

Limitations

to any single real-world base rate. Measuring calibration against realistic base rates would usefully complement the discrimination-focused evaluation we report here. Each model is also evaluated under a single prompt, adapted from the human instructions (§3.3). Because model responses can be sensitive to prompt wording, the exact biases in Table 2 may shift under other phrasings; a prompt-robustness sweep is worthwhile future work. The human– model comparison holds the task framing fixed for both rater types, so it is less exposed to this concern than any single model’s bias estimate read on its own.

This paper reports only the ice-breaker (firstinteraction) task from the Seamless Interaction dataset. Findings about modality strength and class bias may not generalize to conversations later in a relationship or session, or to different prompt structures. All familiar subtypes (friends, family, romantic partner, coworkers, familiar_other) are collapsed into a single “familiar” class, because the balanced design does not have enough dyads per subtype to power a multi-class analysis. A sixclass relationship-type task is defined in the broader project but is out of scope here, so any withinfamiliar heterogeneity is invisible to this binary framing. Only two-person interactions are studied, and relationship inference in larger groups may draw on cues not captured here. The balanced evaluation set is also modest in size (96 dyads, one clip each), so per-model accuracy intervals are wide and only the strongest models clear chance individually. The class-bias and difficulty patterns we emphasize are more robust than any single model’s rank. Our human baseline is drawn entirely from US-resident raters (§3.2); relationship-perception cues can be culturally specific, so this baseline may not represent human performance in other populations. Finally, even the “naturalistic” subset was recorded in a fixed-camera motion-capture studio, with wired lapel microphones and a posed ice-breaker prompt. This is not truly in-the-wild interaction, so findings may not transfer to less controlled contexts such as phone video or casual settings. Our evaluation set is also balanced 50/50 between familiar and stranger dyads by construction, which makes accuracy a clean measure of discrimination but base-rate-specific. A reader might ask whether the models’ stranger-lean is not miscalibration but a well-calibrated prior for a world in which two people recorded together are more often strangers, penalized only by our artificial balance. This does not threaten our central claim. That claim is the difference in criterion between humans and models on identical, matched stimuli. Neither rater type was told the base rate, so this difference does not depend on the true base rate, and neither does d′ , which is base-rate-invariant. It does mean we measure discrimination, not deployment calibration; we do not read the balanced accuracy as a deployment estimate. The observed bias is in any case idiosyncratic in direction and size across models (Table 2), which is hard to square with calibration

Ethics Statement Informed consent and compensation. Rater recruitment and prescreening are described in §3.2 and Appendix D. Before viewing any stimulus, each rater read a description of the study and affirmatively agreed to three statements: that they had read and understood the information, that they were free to withdraw at any time—by closing the browser tab—without giving a reason, and that they agreed to take part. Participation was voluntary and could be ended at any point without penalty. Sessions took roughly 15–30 minutes depending on modality, and raters were compensated at an effective rate of approximately $10–12 per hour, above Prolific’s fair-pay guidance. The task was minimalrisk: raters viewed short, benign clips of consenting conversation partners and answered non-sensitive perceptual and self-report questions. Privacy and data minimization. We collected no personally identifying information from raters. Beyond the task responses and the self-report questionnaires analyzed here, we recorded only the coarse attributes used for prescreening and quota matching (§3.2), keyed to a pseudonymous Prolific identifier that we do not link to any real-world identity. All released rater data are de-identified. Institutional review. This research was conducted outside a university setting and was not reviewed by an institutional review board. Because it collected no personally identifying information from raters and involved only minimal-risk procedures, it did not fall under IRB oversight; we nonetheless followed standard human-subjects safeguards—voluntary informed consent, the right to withdraw without penalty, fair compensation, 9

minimal-risk stimuli, and data de-identification.

New York, NY, USA. Association for Computing Machinery.

Stimulus source. All interaction clips are drawn from the publicly released Seamless Interaction dataset (Agrawal et al., 2025), whose participants consented to the recording and research use of their audio and video. We use these data in accordance with the dataset’s CC-BY-NC 4.0 license (attribution, non-commercial use only): the rendered benchmark clips are redistributed under the same CC-BY-NC 4.0 terms, with attribution to the Seamless Interaction dataset. We do not attempt to re-identify or contact any recorded individual.

Tomás Capretto, Camen Piho, Ravin Kumar, Jacob Westfall, Tal Yarkoni, and Osvaldo A. Martin. 2022. Bambi: A Simple Interface for Fitting Bayesian Linear Models in Python. Journal of Statistical Software, 103:1–29. Paul de Boeck and Mark Wilson, editors. 2004. Explanatory Item Response Models: A Generalized Linear and Nonlinear Approach. Springer Science & Business Media. Thomas G. Dietterich. 1998. Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms. Neural Computation, 10(7):1895– 1923.

Intended use and risks. The benchmark is intended for evaluating and auditing the socialperceptual behavior of models, including the response biases we document (§5.1). Inferring familiarity from behavior could in principle support surveillance or profiling; we release the benchmark to enable research on and scrutiny of such capabilities, not their deployment. Given the modest accuracy and pronounced, idiosyncratic biases we observe (§4), we caution against using these models or this task to make consequential judgments about real individuals.

R. I. M. Dunbar, Juan-Pablo Robledo, Ignacio Tamarit, Ian Cross, and Emma Smith. 2022. Nonverbal Auditory Cues Allow Relationship Quality to be Inferred During Conversations. Journal of Nonverbal Behavior, 46(1):1–18. Chris D. Frith and Uta Frith. 2007. Social Cognition in Humans. Current Biology, 17(16):R724–R732. Andrew Gelman, John B. Carlin, Hal S. Stern, David B. Dunson, Aki Vehtari, and Donald B. Rubin. 2014. Bayesian Data Analysis, 3rd edition. CRC Press, Boca Raton, FL.

References

John W Graham, Bonnie J Taylor, Allison E Olchowski, and Patricio E Cumsille. 2006. Planned missing data designs in psychological research. Psychological Methods, 11(4):323–343.

Ralph Adolphs. 2003. Cognitive neuroscience of human social behaviour. Nature Reviews Neuroscience, 4(3):165–178.

Rachel Grieve and Doug Mahar. 2013. Can social intelligence be measured? Psychometric properties of the Tromsø Social Intelligence Scale – English Version. The Irish Journal of Psychology, 34(1):1–12.

Vasu Agrawal, Akinniyi Akinyemi, Kathryn Alvero, Morteza Behrooz, Julia Buffalini, Fabio Maria Carlucci, Joy Chen, Junming Chen, Zhang Chen, Shiyang Cheng, Praveen Chowdary, Joe Chuang, Antony D’Avirro, Jon Daly, Ning Dong, Mark Duppenthaler, Cynthia Gao, Jeff Girard, Martin Gleize, and 65 others. 2025. Seamless interaction: Dyadic audiovisual motion modeling and large-scale dataset.

Kilem L Gwet. 2021. Handbook of Inter-Rater Reliability: Chance-corrected Agreement Coefficients, 5th edition, volume 1. AgreeStat Analytics.

Nalini Ambady and Robert Rosenthal. 1992. Thin slices of expressive behavior as predictors of interpersonal consequences: A meta-analysis. Psychological Bulletin, 111(2):256–274.

Michael J. Hautus. 1995. Corrections for extreme proportions and their biasing effects on estimated values of d′. Behavior Research Methods, Instruments, & Computers, 27(1):46–51.

Gregory A. Bryant, Christine S. Wang, and Riccardo Fusaroli. 2020. Recognizing affiliation in colaughter and cospeech. Royal Society Open Science, 7(10):201092.

Qi Jia, Hongru Huang, and Kenny Q. Zhu. 2021. DDRel: A New Dataset for Interpersonal Relation Classification in Dyadic Dialogues. Proceedings of the AAAI Conference on Artificial Intelligence, 35(14):13125–13133.

Angelo Cafaro, Johannes Wagner, Tobias Baur, Soumia Dermouche, Mercedes Torres Torres, Catherine Pelachaud, Elisabeth André, and Michel Valstar. 2017. The NoXi database: Multimodal recordings of mediated novice-expert interactions. In Proceedings of the 19th ACM International Conference on Multimodal Interaction, ICMI ’17, pages 350–359,

Denys Katerenchuk, David Guy Brizan, and Andrew Rosenberg. 2014. “was that your mother on the phone?”: Classifying interpersonal relationships between dialog participants with lexical and acoustic properties. In Proc. Interspeech 2014, pages 1831– 1835.

10

Eunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park, Seyoung Song, A. Seza Doğruöz, Alice Oh, and Najoung Kim. 2026. Are they lovers or friends? Evaluating LLMs’ Social Reasoning in English and Korean Dialogues. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 23431–23451, San Diego, California, United States. Association for Computational Linguistics.

Georg Rasch. 1960. Probabilistic Models for Some Intelligence and Attainment Tests. MESA Press, 5835 S. Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019. Social IQa: Commonsense Reasoning about Social Interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4463– 4473, Hong Kong, China. Association for Computational Linguistics.

Fanqi Kong, Weiqin Zu, Xinyu Chen, Yaodong Yang, Song-Chun Zhu, and Xue Feng. 2026. SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning. In Findings of the Association for Computational Linguistics: ACL 2026, pages 37379–37403, San Diego, California, United States. Association for Computational Linguistics.

David Silvera, Monica Martinussen, and Tove I. Dahl. 2001. The Tromsø Social Intelligence Scale, a selfreport measure of social intelligence. Scandinavian Journal of Psychology, 42(4):313–319.

Nida Latif, Adriano V. Barbosa, Eric Vatikiotis-Bateson, Monica S. Castelhano, and K. G. Munhall. 2014. Movement Coordination during Conversation. PLOS ONE, 9(8):e105036.

Harold Stanislaw and Natasha Todorov. 1999. Calculation of signal detection theory measures. Behavior Research Methods, Instruments, & Computers, 31(1):137–149.

Junnan Li, Yongkang Wong, Qi Zhao, and Mohan S. Kankanhalli. 2017. Dual-Glance Model for Deciphering Social Relationships. In Proceedings of the IEEE International Conference on Computer Vision, pages 2650–2659.

Qianru Sun, Bernt Schiele, and Mario Fritz. 2017. A Domain Based Approach to Social Relation Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3481–3490. Wang Tang, Fethiye Irmak Dogan, Linbo Qing, and Hatice Gunes. 2026. AsyReC: A Multimodal GraphBased Framework for Spatio-Temporal Asymmetric Dyadic Relationship Classification. IEEE Transactions on Circuits and Systems for Video Technology, 36(3):3693–3708.

Dominique Makowski, Mattan S. Ben-Shachar, S. H. Annabel Chen, and Daniel Lüdecke. 2019. Indices of effect existence and significance in the Bayesian framework. Frontiers in Psychology, 10. Leena Mathur, Paul Pu Liang, and Louis-Philippe Morency. 2024. Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 20541–20560, Miami, Florida, USA. Association for Computational Linguistics.

Emma M. Templeton, Luke J. Chang, Elizabeth A. Reynolds, Marie D. Cone LeBeaumont, and Thalia Wheatley. 2023. Long gaps between turns are awkward for strangers but not for friends. Philosophical Transactions of the Royal Society B: Biological Sciences, 378(1875):20210471.

Quinn McNemar. 1947. Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages. Psychometrika, 12(2):153–157.

Linda Tickle-Degnen and Robert Rosenthal. 1990. The nature of rapport and its nonverbal correlates. Psychological Inquiry, 1(4):285–293.

Cristina Palmero, Javier Selva, Sorina Smeureanu, Julio C. S. Jacques Junior, Albert Clapes, Alexa Mosegui, Zejian Zhang, David Gallardo, Georgina Guilera, David Leiva, and Sergio Escalera. 2021. ContextAware Personality Inference in Dyadic Scenarios: Introducing the UDIVA Dataset. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1–12.

Amir Zadeh, Michael Chan, Paul Pu Liang, Edmund Tong, and Louis-Philippe Morency. 2019. Social-IQ: A Question Answering Benchmark for Artificial Social Intelligence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8807–8817. Shuxiang Zhang, Yiting Yin, Wenxuan Song, Yuhang Wu, and Miao Liu. 2026. PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models. Preprint, arXiv:2606.23092.

David Premack and Guy Woodruff. 1978. Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4):515–526. Zheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li, Yi Yuan, Jingdong Chen, and Le Wang. 2026. HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMs. Proceedings of the AAAI Conference on Artificial Intelligence, 40(30):24973–24981.

11

A

Benchmark Quality Filters Your job is to pay close attention to each conversation and determine whether the two people have met before -- meaning they are familiar, such as friends, family members, coworkers, or romantic partners -- or are meeting for the first time in this interaction -- meaning they are strangers.

Every candidate interaction and clip must pass the automatic quality filters in Table A1 before entering the stratified selection pool (§3.1). Interaction-level filters (F1, F4, F5) gate the whole source interaction; clip-level filters (F2, F3) gate the specific 20second window sampled from it. Clips are 20 s long with boundaries snapped within ±5 s to the nearest turn start. These filters remove clips that cannot support the task regardless of relationship; they are followed by the LLM-based prompt-adherence audit (§3.1), which verifies that each interaction’s recorded prompt label matches its actual content.

B

Raters then made a forced choice: familiar or strangers. Model prompt. [Read / Listen to / Watch] the following [transcript / audio clip / video clip], taken from a real conversation between two people. These people are playing a game called "Either Or," where they discuss whether they would prefer one thing -- for example, the ability to fly -- or an alternative, such as the ability to breathe underwater.

Model Inference Configuration

Table A2 lists the inference configuration for each of the 26 models. All models received the same prompt construction and response parsing (§3.3); the settings below cover per-model decoding and, where applicable, reasoning parameters. Cloud models were queried at temperature 0 with a fixed seed (42) wherever the API exposed them; the three locally-run open-weight models (Qwen2-Audio, Qwen2.5-Omni, Voxtral) used greedy decoding (do_sample=False), which is deterministic without a seed. Answer length was capped at 150–300 tokens for non-reasoning models, while reasoning models received a larger combined answer-plusreasoning budget (8,000 tokens for the Gemini thinking and pro variants; 4,000 for Inkling and Muse Spark) to avoid truncating the response.

C

Your job is to pay close attention to the conversation and determine whether the two people have met before -- meaning they are familiar, such as friends, family members, coworkers, or romantic partners -- or are meeting for the first time in this interaction -- meaning they are strangers. For each clip, return: { "relationship_label": str {FAMILIAR|STRANGER}, "confidence": float [0, 1], "reason": str } Use the following confidence scale: 0 = Not at all confident, ..., 0.5 = Moderately confident, ..., 1 = Extremely confident Do not provide extra text outside the JSON object.

Task Instructions and Prompts

D

Both human raters and models were given the same task framing, “Either Or” game description, and familiar/strangers definition, differing only in how a response was collected. Human raters selected an answer with a button; models were additionally instructed to emit a JSON object with a label and confidence. Modality-specific slots (shown here in brackets, Text/Audio/Video order) were filled per panel or per clip.

Human Rating Study Details

Raters were recruited on Prolific with platformlevel quality screening in addition to the demographic criteria in §3.2: a Prolific approval rate of 99–100% and at least 10 prior submissions. Under the planned-missing, 6-block design (Graham et al., 2006), each rater was assigned to a single block and each of the 96 dyads appeared in two of the six blocks, so every dyad was rated by an overlapping subset of raters (25–35 per dyad). By design each rater judged 32 dyads; in the audio panel a platform issue left 13 of the 94 raters with 29–31 completed trials, while all text and video raters completed the full 32.

Human rater instructions. You will [read / listen to / watch] a series of [transcripts / audio clips / video clips] taken from real conversations between two people. These people are playing a game called "Either Or," where they discuss whether they would prefer one thing -- for example, the ability to fly -- or an alternative, such as the ability to breathe underwater.

Inter-rater agreement. Agreement among human raters is slight, as expected for a difficult perceptual task with a balanced label set. Fleiss’s κ 12

ID

Level

Criterion

F1

Interaction

F2 F3 F4

Clip Clip Interaction

F5

Interaction

Both recorded EO prompts are free of placeholder/boilerplate text (e.g., “you will each see the same prompt,” “given different lists of sentences”) that signals the true question was not logged. ≥ 4 merged conversational turns within the clip window. Each speaker contributes ≥ 3 s of speech within the clip window. < 35% of cross-track turn-start pairs coincide within 0.2 s, rejecting duplicated or desynchronized audio tracks. ≥ 90 s of speech in the interaction.

Table A1: Automatic quality filters applied during clip generation and the EO interaction audit. Interaction-level filters gate the source interaction; clip-level filters gate the sampled 20 s window. Thresholds are the generator/auditor constants (scripts/generate_samples_rapport.py, scripts/audit_eo_interactions.py).

(computed per stimulus over whichever raters saw it, accommodating the incomplete-block design) is 0.08 for text and 0.17 for both audio and video. Bennett’s S is essentially identical (0.09, 0.17, 0.17 for text, audio, video), as expected when the two classes are balanced (Gwet, 2021). The near-zero text agreement is consistent with individual text raters performing at chance (Table 1): raters are not converging on a shared, reliable text cue.

best modeled with partially-pooled crossed random effects (§3.2), which shrink noisy individual estimates, rather than trusting raw per-rater rates or treating each rater as a one-shot observation. Text raters cluster tightly around chance, consistent with the near-zero text agreement reported in Appendix D.

E

Table A5 lists the eight hardest and eight easiest dyads by mean model-estimated probability of a correct human response, from the crossed random-effects model of §5.5. Several of the hardest dyads fall below chance in every modality, functioning as systematic lures. Estimated easiness correlates across modality pairs (text–audio r = 0.52, text–video r = 0.33, audio–video r = 0.61; all p < .01), indicating that difficulty is largely a property of the conversation rather than the observation channel. Table A4 decomposes the human–model difficulty correlation reported in §5.5. The raw pooled correlation (human-crowd accuracy vs. accuracy pooled over all models) is strong in audio, moderate in video, and only marginal in text. The weak raw text value is a suppression effect rather than absence of shared structure: humans and models carry opposing class biases (§5.1), so their per-dyad accuracies partly anti-align on the familiar/stranger axis. Partialling out the true class raises the withinclass correlation—fine-grained difficulty that is not reducible to a two-way class effect—to r = 0.61 (text), 0.58 (audio), and 0.44 (video), all p < .001. The shared difficulty is, however, partly a crowdlevel property: repeating the analysis with the single strongest model per modality in place of the pooled estimate leaves it robust only in video (r = 0.45, p < .001), with weak raw associations in audio (r = 0.19, p = .06) and text (r = −0.04, n.s.); confidence-weighted single-model estimates

F

Rater Individual Differences

Table A3 reports, per modality, the reliability of the TSIS-PS social-intelligence scale and its correlation with per-rater accuracy. The scale is highly reliable in every panel, but its correlation with accuracy is null in audio and video and positive only in text—the modality where accuracy is at chance— so we do not interpret it as evidence of a general skill advantage. For self-reported cue use, raters most often reported attending to interactional cues (rapport, responsiveness) across all modalities, but individual cue-attendance scores rarely predicted accuracy (only a handful of per-cue correlations reached p < .05, with no stable cross-modality pattern). Full cue-use heatmaps are provided with the released analysis notebooks. The per-rater accuracy clouds in Figure 1 show the full distribution of per-rater accuracy in each modality (94 text, 94 audio, 92 video raters; each rater’s accuracy over the 32 dyads in their assigned block, 29–32 for a few audio raters affected by a platform issue). The clouds are wide, but much of that width is sampling noise: with only ∼32 trials per rater, binomial variation around a common ability already reproduces most of the observed spread—including the upper tail that appears to exceed the best model (§5.4)—so the distribution should not be read as a stable ordering of raters by skill. This is also why per-rater accuracy is 13

Dyad Difficulty

Model

API / model ID

Temp.

Seed

Reasoning

Text gpt_text gpt_text_mini gemini_text gemini_text_thinking claude_text mistral_text inkling_text muse_spark_text muse_spark_text_high

gpt-4o gpt-4o-mini gemini-3.5-flash gemini-3.5-flash claude-opus-4-8 mistral-large-latest thinkingmachines/Inkling muse-spark-1.1 muse-spark-1.1

0 0 0 0 —a 0 0 —c —c

42 42 42 42 —a 42 42b —c —c

— — off dynamic offa — on minimal high

Audio gpt_audio gpt_audio_1_5 gpt_audio_mini gemini_audio gemini_audio_pro gemini_audio_thinking qwen_audio qwen_omni voxtral_audio inkling_audio muse_spark_audio muse_spark_audio_high

gpt-audio gpt-audio-1.5 gpt-audio-mini gemini-3.5-flash gemini-pro-latest gemini-3.5-flash Qwen/Qwen2-Audio-7B-Instruct Qwen/Qwen2.5-Omni-7B mistralai/Voxtral-Mini-3B-2507 thinkingmachines/Inkling muse-spark-1.1 muse-spark-1.1

0 0 0 0 0 0 greedy greedy greedy 0 —c —c

42 42 42 42 42 42 — — — 42b —c —c

— — — off dynamicd dynamic — — — on minimal high

Video gemini_video gemini_video_pro gemini_video_thinking muse_spark_video muse_spark_video_high

gemini-3.5-flash gemini-pro-latest gemini-3.5-flash muse-spark-1.1 muse-spark-1.1

0 0 0 —c —c

42 42 42 —c —c

off dynamicd dynamic minimal high

Table A2: Per-model inference configuration (26 models, seven companies; roster matches Table 2). Temp. and Seed are the decoding temperature and random seed passed to the API; greedy marks the locally-run open-weight models decoded with do_sample=False (deterministic, no seed). Reasoning: “—” = non-reasoning model; “off” = reasoning-capable but disabled (Gemini thinking_budget=0); “dynamic” = model sets its own reasoning depth (thinking_budget=−1); “on” = reasoning always on with no depth control; “minimal”/“high” = named reasoning-effort tier. Two Gemini IDs are aliases; at run time (July 2026) gemini-pro-latest resolved to gemini-3.1-pro-preview, and gemini-3.5-flash was pinned explicitly rather than gemini-flash-latest (which then resolved to gemini-3.6-flash). a Claude rejects temperature/top_p and exposes no seed; extended thinking was not enabled. b Inkling accepts a seed but the API echoed null, so determinism is best-effort. c Muse Spark exposes no temperature or seed control, and its reasoning depth is non-deterministic run-to-run. d gemini-pro-latest cannot disable thinking and always reasons dynamically.

Modality

α

Pearson r

p

Spearman ρ

Text Audio Video

0.91 0.88 0.92

0.29 0.03 −0.12

.004 .78 .24

0.23 0.00 −0.19

Estimator Pooled, raw Pooled, within-class Best model, raw Best model, within-class

Table A3: TSIS-PS reliability and its correlation with per-rater accuracy (N = 94 text, 94 audio, 92 video).

Text

Audio

Video

0.20 0.61 −0.04 0.24

0.56 0.58 0.19 0.26

0.34 0.44 0.45 0.47

Table A4: Human–model per-dyad difficulty correlation (r) under four estimators, by modality (N = 96 dyads). “Pooled” averages model accuracy over all models; “best model” is the single most accurate model per modality (text gpt_text_mini, audio gemini_audio_pro, video gemini_video). “Within-class” partials out the true familiar/stranger label.

match the binary ones almost exactly, so this is not an artifact of one prediction per dyad. Individual models thus express the human-aligned difficulty signal only noisily, and pooling recovers it.

14

true class familiar stranger Text: r = +0.20, within +0.61

Audio: r = +0.56, within +0.58

Video: r = +0.34, within +0.44

0

0

0

Model acc. (%)

100 80 60 40 20 0 50 Human acc. (%)

100

50 Human acc. (%)

100

50 Human acc. (%)

100

Figure A1: Humans and models tend to find the same dyads hard (§5.5). Each point is one of the 96 EO dyads, plotting human per-dyad accuracy (over all raters) against model per-dyad accuracy (pooled over all models), colored by true class; the diagonal marks equal difficulty. Per-dyad difficulty correlates positively in audio (r = 0.56, ≈ 31% shared variance) and video (r = 0.34, ≈ 12%; both p < .01); text (r = 0.20) is only marginally significant (p = .045).

Dyad

Truth

Txt.

Aud.

Vid.

Mean

Hardest sample_080 sample_055 sample_072 sample_053 sample_038 sample_057 sample_083 sample_065

str. str. fam. str. fam. str. str. fam.

29.6 20.5 35.4 31.1 25.8 34.1 31.7 34.9

32.0 32.1 19.9 27.0 28.5 30.3 34.7 33.4

22.2 33.3 38.0 40.7 47.9 37.9 36.1 35.0

27.9 28.6 31.1 32.9 34.1 34.1 34.2 34.4

Easiest sample_045 sample_062 sample_051 sample_077 sample_009 sample_030 sample_073 sample_016

fam. str. str. fam. fam. str. fam. str.

59.6 69.2 53.3 57.4 70.7 63.3 71.4 68.5

83.8 75.2 88.4 79.1 77.5 83.7 86.1 86.8

79.3 79.4 82.0 89.9 87.4 91.1 81.9 86.8

74.2 74.6 74.6 75.5 78.5 79.4 79.8 80.7

Table A5: Hardest and easiest eight dyads by mean estimated P(correct human response) across modalities (%). “str.” = stranger, “fam.” = familiar.

15

Record · ID 422284 · SHA-256 a7976a4679dbfb45
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.