GroupAffect-4: A Multimodal Dataset of Four-Person Collaborative Interaction Meisam Jamshidi Seikavandi1,2 , Alice Modica1,3∗, Anna Obara1,4∗, Shan Ahmed Shaffi1 , Fabricio Batista Narcizo1,2 , Tanya Ignatenko1 , Ted Vucurevich1 , Karim Haddad1 , Daniel Barratt3 , Daniel Overholt4 , Jesper Bünsow Boldt1 , Paolo Burelli2 , Andrew Burke Dittberner1 1
arXiv:2605.19765v1 [cs.AI] 19 May 2026
2
GN Advanced Science, GN Group, Ballerup, Denmark IT University of Copenhagen, brAIn lab, Copenhagen, Denmark 3 Copenhagen Business School, Copenhagen, Denmark 4 Aalborg University, Denmark
Abstract Existing affective-computing, social-signal-processing, and meeting corpora capture important parts of human interaction, but they rarely support analysis of affect in co-located groups as a coupled individual, interpersonal, and group-level process. The required signals (per-participant physiology, eye movement, audio, self-report, task outcomes, and personality) are usually fragmented across separate dataset traditions. We introduce GroupAffect-4, a multimodal corpus of 40 participants in 10 four-person groups, each completing four ecologically varied collaborative tasks spanning information pooling, negotiation, idea generation, and a public-goods game. Each participant is instrumented with a wrist-worn physiology sensor, eye-tracking glasses, and a close-talk microphone; sessions include continuous affect self-reports, post-task questionnaires, task outcomes, and Big-Five personality scores, all time-aligned to a shared clock. The dataset covers over 91% of expected physiology windows and 98% of eye-tracking windows, with strong task validity confirmed by a clear affective manipulation check across the negotiation block. We define fifteen benchmarkable targets spanning three analysis levels within-person state, between-person traits, and group dynamics and report leave-one-group-out feasibility baselines establishing the dataset’s evaluative scope. GroupAffect-4 is released with a Brain Imaging Data Structure (BIDS)-inspired structure, Croissant metadata, a datasheet, per-session quality reports, and open processing scripts. Code and processing scripts are available at https://github.com/meisamjam/GroupAffect-4; the dataset is publicly archived at https://zenodo.org/records/20037847.
1
Introduction
Meetings are a central setting where affect, cognition, and social dynamics co-occur under collaborative pressure [8, 17, 23]. Understanding them requires observation at three co-occurring levels: within each person (physiological state, attention, cognitive load), between people (personality, trust, negotiation stance), and across the group (floor balance, coordination, collective outcomes), following multilevel views of team process and social interaction [23, 59]. Yet, most multimodal affect corpora address at most one of these levels, and typically in single-participant or dyadic settings [39, 6, 19]. Co-located groups impose qualitatively different constraints: turn-taking is multi-party, gaze targets multiply, roles can shift within a session, and physiological responses must be interpreted against an evolving social context [41, 15, 58, 59]. Theoretically, this situates group affect at the intersection of individual appraisal processes, where personal relevance and coping capacity shape felt response [21, 43], and interpersonal emotion dynamics, where the ongoing reactions of co-present others can reshape appraisal, expression, and regulation [36, 55, 2]. Datasets combining per-participant physiology, gaze, audio, task outcomes, self-reports, and personality measures in the same co-located group setting are rare among the publicly described corpora reviewed in Section 2 and table 1. ∗ These authors contributed equally.
Preprint.
We present GroupAffect-4, a multimodal corpus of four-person group interaction designed to support analysis at three levels: within-person affective and cognitive state, between-person stable traits, and group-level interaction dynamics within the same synchronised release. The motivating premise is not scale, but density: each participant is observed simultaneously through autonomic sensing, egocentric eye tracking, close-talk audio, task outcomes, and in-situ self-report, so the same session can support questions about regulation, attention, conversational floor dynamics, and cooperation that are usually split across separate datasets. Simultaneity matters because cross-level relationships such as whether physiological arousal tracks subjective dominance, or whether audio overlap predicts cognitive demand independently of physiology, can only be tested when modalities are co-observed at the same moment across all participants. This design also builds on prior work showing that gaze dynamics, personality, and multimodal physiological signals can contribute complementary information for emotion perception and face-to-face affect modelling [47, 45, 46, 48]. This paper is a dataset-characterisation paper: its primary contribution is a transparent description of what the dataset contains, how modalities are synchronised, what can be evaluated, and where the main limitations lie. Figure 1 gives a schematic of the study design, modality stack, task sequence, and benchmarkable outputs.
2
Literature and Dataset Positioning
GroupAffect-4 is a compact, high-density dataset rather than a scale-first corpus, following the dataset-characterisation tradition in which value depends not only on sample size but also on documentation, synchronisation, annotation quality, and evaluative scope [13, 60, 33]. Studying group affect requires capturing three co-occurring levels simultaneously: individual affective and cognitive state, interpersonal dynamics, and group-level coordination and outcomes [23, 59]. Individual state is supported by physiology, pupil, and self-report; interpersonal dynamics by audio and gaze; and group-level coordination by floor-balance and outcome measures. The central comparison question is therefore not which prior corpus is largest, but which combines the per-participant signals needed at each level – wearable physiology, egocentric gaze, close-talk audio, in-situ self-report, explicit task outcomes, and personality measures – in one synchronised co-located four-person release. Classic affective corpora provide necessary background: IEMOCAP [6], RECOLA [39], and DEAP [19] address dyadic or individual settings; AMIGOS adds group media-viewing with wearable physiology and personality [27]; K-EmoCon provides wearable sensing with rich emotion annotation in dyadic debates [35]; UDIVA covers dyadic interaction with personality and physiology [34]. Among multiparty corpora, AMI and ICSI are canonical meeting baselines [8, 17]; ELEA, GAP, and UGI add small-group decision tasks with questionnaires and performance measures [42, 62, 3]. The KTH-Idiap Group-Interviewing corpus provides four-person discussions with close-talk microphones and gaze/visual focus of attention (VFOA) design [30]; GaMMA contributes four-person polyadic conversations with eye-tracking glasses, motion, and multi-talker audio [11]; a recent agile-team dataset covers four-person software teams with audio, visual attention, and equity surveys [26]; and G-REx contributes large-scale group PPG/EDA in naturalistic movie sessions [4]. The ingredients already exist but are distributed across corpora: task structure and outcomes in AMI, ELEA, GAP, UGI, and the agile-team dataset; physiology and affect labels in AMIGOS, K-EmoCon, and G-REx; gaze-rich multiparty sensing in KTH-Idiap and GaMMA. Table 1 tabulates the closest comparators per axis. Among publicly described datasets reviewed here, we did not find an earlier dataset that clearly combines all these strands in one co-located four-person benchmarkable release. GroupAffect-4 is positioned between fully lab-controlled paradigms (tightly constrained illumination, timing, and stimuli) and fully naturalistic settings (uncontrolled context), reflecting the broader tension in affective computing between experimental control, ecological validity, and annotation reliability [61, 44, 59]: structured task-context-primed windows and LSL synchronisation provide auditability, while unscripted group interaction within those windows preserves ecological realism.
3
Dataset Design
Participants. Sixteen sessions were recorded in total: the first five were pilot sessions used to finalise the protocol and are not included in the release; of the remaining 11 non-pilot sessions, 10 are released (one was excluded due to incomplete modality coverage). The final subset therefore contains 10 final groups of four participants each (Np = 40 unique participants), all with complete demographics and Big-Five Inventory (BFI-44) scores. Mean age is 35.4 years (SD 9.2, range 23–58); 21 female and 19 male (demographics in Figure 4). Collected demographic fields include age, sex, handedness, English proficiency, and education level; racial or ethnic composition was 2
Table 1: Closest dataset comparators for GroupAffect-4, organised by nearest shared axis. Grp. = group/dyad size. Feature columns: Ph = wearable physiology (EDA/PPG), ET = mobile egocentric eye-tracking, Au = per-participant close-talk audio, SR = in-situ affect self-report, Pe = personality scores, Ta = structured collaborative tasks with outcomes. ✓ present; ◦ partial/limited; — absent. GroupAffect-4 (this work) provides all six. Dataset
Grp.
AMI [8]
4
ELEA [42] GAP [62] UGI [3] KTH-Idiap [30] MatchNMingle [7] AMIGOS [27] K-EmoCon [35] G-REx [4] GaMMA [11] Agile-team [26] AFFEC [45] GroupAffect-4
Nearest shared axis
4-person meetings; synchronised multimodal capture; rich annotations 4–5 Small-group survival task; questionnaires, personality, performance small Group decision task; transcripts and explicit performance outcomes small Group task; questionnaires; privacypreserving multimodal capture 4 4-person discussions; close-talk mics; VFOA design var. Real-world social sensing; wearables, personality, group formations ind./grp Wearable physiology, self-assessment, and personality 2 Wearable sensing; rich emotion labels; dyadic social interaction grp. Large-scale group physiology in naturalistic (audience) settings 4 4-person conversations; ET glasses; multi-talker audio 4 4-person teams; audio, visual attention, task outcomes, equity surveys Face-to-face; EEG, ET, GSR, self2 reports, personality 4
This work
Ph
ET
Au
SR
Pe
Ta
—
—
◦
—
—
✓
—
—
◦
✓
✓
✓
—
—
◦
✓
—
✓
—
—
◦
✓
◦
✓
—
◦
✓
—
—
◦
◦
—
◦
—
✓
—
✓
—
—
✓
✓
—
✓
—
✓
✓
◦
◦
✓
—
—
—
—
—
—
✓
✓
—
—
—
—
◦
✓
✓
◦
✓
✓
✓
◦
✓
✓
—
✓
✓
✓
✓
✓
✓
not collected as part of the study protocol. Participants gave written informed consent under a participant information statement reviewed and approved by the GN Hearing A/S Legal Department. Within-session roles use anonymous seat identifiers P1–P4; the identity mapping is not part of the release. Pre-session familiarity ratings cluster at 1–3 on a 7-point scale (cross-participant mean 1.68, SD 1.60 for other-group-member ratings; maximum group mean 3.08/7), indicating that most participants were meeting for the first time and that established social ties are unlikely to be a primary driver of the observed group dynamics. Pre-session questionnaire. Before each session, participants completed the BFI-44 [18]. Perdomain scores are computed as item means after reverse coding; all 40 participants have complete BFI-44 records. Session structure. Each session (≈70 min) comprised a free-talk social baseline (T0), in which participants engaged in unstructured conversation to allow familiarisation and sensor stabilisation, a common concern in psychophysiological recording and wearable affect measurement [51, 37, 24], followed by four collaborative tasks, with task onsets and offsets demarcated by Lab Streaming Layer (LSL) markers and written to release-side task-window files. All sessions followed the same fixed T0→T1→T2→T3→T4 sequence with no counterbalancing. Any cross-task effects should therefore be interpreted as effects within this fixed task order, where fatigue, and rapport also vary, rather than as effects that can be attributed solely to the tasks themselves. Tasks. The four tasks were selected to target distinct social–affective regimes (Table 2): T1 emphasizes information pooling, group reasoning, and consensus formation under initially asymmetric information [50, 22]; T2 introduces interdependent interests and negotiation pressure, giving rise to concession patterns, cooperative negotiation strategies, and integrative trade-offs [56, 55]; T3 focuses on creative idea generation, engagement in group discussion, and convergence toward a shared solution[38]; and T4 uses a small public-goods game in which private contributions determine individual and group payoffs, highlighting cooperative tension, fairness evaluations, and social comparison [12, 10]. Key outcomes: T1 group consensus rate 100% (above the ≈20–40% literature baseline [22]); T2 all groups settled but overran the 8-min guideline (mean 11.4 min, SD 1.6); T3 3
Figure 1: GroupAffect-4 study design and release structure. Left: top-down lab layout with four colocated participants (P1—P4) seats around a 1.80 m × 0.80 m × 0.75 m rectangular table, seven camera positions, and ArUco calibration markers. Upper Right: participants during a task session wearing Tobii Glasses 3, EmotiBit sensors and lapel microphones; tablets collect self-reported SAM valence, arousal, and dominance; a shared screen displays task prompts. Lower Right: session timeline showing baseline (T0) followed by four collaborative tasks (T1: hidden-profile decision; T2: mininegotiation; T3: idea generation; T4: public-goods game), with Valence-Arousal-Dominance (VAD) probes collected at intervals. Table 2: Session structure: tasks with phases’ timer durations as displayed on the shared screen. ID
Task
Duration (s)
T0 T1
Free-talk baseline Hidden-profile decision
T2
Mini-negotiation
T3
Idea generation and selection
T4
Public-goods micro-game
300 (free talk) 75 (reading); 420 (discussion); 60 (selection) 480 (negotiation); no timer (settlement form) 180 (generation); 420 (discussion); 60 (selection) 60 (contribution); 60 (reveal); 180 (discussion)
winning-idea records recovered for 10/10 groups; T4 mean contribution 6.92/10 tokens (SD 3.10), higher than typical one-shot rates [10], likely reflecting rapport built across T1–T3. Self-report probes. Valence, arousal, and dominance probes are collected at scheduled moments during the tasks - excluding dominance for T4 - on a 9-point Self-Assessment Manikin (SAM) scale, following dimensional affect and Pleasure-Arousal-Dominance (PAD)/VAD measurement traditions [25, 40, 5]. In-task items probe each participant’s first-person felt state [40, 5]. Each task is followed by a brief post-task questionnaire that targets participants’ evaluations of the task and of their interactions with the group. They use 1–7 Likert scales for both self-directed evaluations, and other-directed perceived appraisals of group members’ influence and relational stance, captured via per-seat dominance and trust ratings. [55, 36]. In-task probes are administered as retrospective snapshots at scheduled phase boundaries, with near-complete coverage and per-task completeness (98.8% overall), as documented in Section B.
4
Modalities and Acquisition
The GroupAffect-4 version 1.0 release (no video) contains synchronised physiology, egocentric gaze/pupil features, audio-derived features, transcript artifacts, behavioural responses, task outcomes, 4
and personality metadata. Multi-camera room video and scene video were recorded but are deferred to a future release pending full quality-control. The modalities are listed in Table 5 (see Section B) and are designed as complementary observational channels rather than redundant measures of a single latent construct, following multimodal affective-computing and social-signal-processing practice [59, 61, 19]. Audio captures conversational floor dynamics and prosodic variation core behavioural channels in social interaction and affective computing [41, 59, 44]. Pupil diameter indexes attentional engagement and cognitive load under natural viewing conditions [24], while prior face-viewing work suggests that gaze dynamics can carry emotion-perception signal in naturalistic affective settings [47, 46]. Wrist physiology (EDA, PPG, skin temperature) provides autonomic arousal proxies that operate largely independently of the social-behavioural channel [37]. Self-report supplies the subjective grounding that links observed signals to experienced states, while also introducing the usual limitations of retrospective and probe-based affect measurement [40, 5]. The near-zero cross-modal correlation between pupil and physiology features (|r| < 0.10 across 136 participant-task rows; Section 7) suggests that these channels provide partly distinct information rather than redundant readouts motivating multimodal rather than unimodal benchmarks [24, 37]. Physiology. Each participant wears an EmotiBit wrist sensor (photoplethysmography (PPG), electrodermal activity (EDA), skin temperature, inertial measurement unit (IMU) at ≈25 Hz). quality control (QC) pass: 182/200 rows observed (91.0%); usable rates 76.0% PPG/EDA/Temp, 90.5% IMU (Table 4). PPG-derived heart-rate variability (HRV) should be interpreted cautiously given the 25 Hz rate (≈40 ms root mean square of successive differences (RMSSD) quantisation floor) and sensitivity to motion artefacts [51, 49, 14]. EDA rows with elevated wrist motion are flagged motion_contaminated rather than discarded; accelerometer data are retained as a companion QC signal [52, 37]. Eye tracking. Each participant wears a Tobii Pro Glasses 3 unit providing head-relative scene-frame gaze, pupil diameter, and validity flags at 50 Hz. 196/200 rows observed (98.0%); 87.2% gaze and 95.4% pupil usability. Room-frame gaze alignment is deferred to v2. The release does not include luminance-normalised pupil modelling; pupil effects should be interpreted with residual light-reflex confounding in mind [24]. Audio. Per-participant close-talk DPA 4060 microphones are recorded at 48 kHz, 24 bit via an RME Fireface 802 interface. The public release includes prosodic features (opensmile GeMAPSv01b), transcript artifacts, speech-activity flags, and interaction metrics (speaking fraction, pause count, overlap). Audio quality control achieved a 95% pass rate (38/40 participant-task rows usable; two participants flagged for minimal vocal activity). Raw audio is gated under a Data Use Agreement (DUA) due to voice re-identification risk [31, 53]. Synchronisation. Acquisition is organised around a common LSL timebase [20]. Redundant timing anchors are preserved so offsets, missing streams, and alignment assumptions can be inspected postprocessing. For AV devices, frame-log start-spread analyses target sub-10 ms alignment as a datasetlevel QC criterion. For the DPA 4060 microphones, the hardware clock drifts by approximately 0.04 ms/s relative to the Extensible Data Format (XDF)/LSL clock, so task clipping uses permicrophone linear anchor correction rather than a single global offset; in documented sessions this reduces practical task-level residual sync error to below 5 ms (Section L).
5
Processing and Release
The release follows a BIDS-inspired structure [16] organised around Findable, Accessible, Interoperable, and Reusable (FAIR) principles [60], in which original neuroimaging-focused BIDS conventions are adapted to a multimodal setting that combines physiology, eye tracking, audio-derived features, behavioural responses, and task metadata. Per-modality subdirectories hold physiology, eye tracking, audio-derived features, transcripts, behavioural responses, and annotations; the processing pipeline is deterministic and reproducible. Feature tables contain task-window summary statistics (means, delta-from-baseline values, prosodic summaries, turn-taking summaries). For benchmarks (Section 6) a five-step modality-aware pipeline is applied: eye tracking (ET) quality gating, physiological plausibility gating, ±3 σ winsorisation, within-person robust z-score normalisation (median/MAD across T1–T4), and fold-internal KNN imputation; full details are in Section G. B0–B3d numbers should be read as biased characterisation estimates: within-person normalisation is computed before the LOGO split, creating a mild dependency between test-row statistics and training normalisation, a known source of optimistic bias in cross-validation when preprocessing is not fully nested within the training fold [57, 9]. A leakage-free protocol, fit from training participants, and applied to held-out set, is recommended for future work. Tabular modalities are released under CC BY 4.0; raw audio 5
is gated under a Data Use Agreement due to re-identification risk. The release is documented with Croissant metadata, Responsible AI fields, and a datasheet, following current dataset-documentation and metadata practices [13, 1, 29, 32].
6
Benchmark Tasks and Feasibility Baselines
We define evaluation targets across three analysis levels and report feasibility baselines using leaveone-group-out cross-validation (LOGO-CV) (split key = group_id), ensuring no participant in the test fold shares an interaction context with training data; the mild preprocessing leakage from within-person normalisation described in Section 5 applies independently of fold structure [57, 9]. Baselines use a single Ridge regressor or logistic classifier with the same preprocessing pipeline as in Section 5 but no temporal modelling or architecture search. Within-person z-score normalisation is applied for within-person state targets (B0–B3d); raw delta-from-baseline features (physiology and pupil) plus absolute audio features are used for between-person targets (B4a–B5). Note that audio is baseline-normalised for B0–B3d (within-person z-score across T1–T4) but kept in absolute form for B4a–B5, because T0 free-talk audio is insufficiently controlled for a per-participant baseline; this modality asymmetry should be kept in mind when comparing modality contributions across benchmarks. Residual missing values are imputed via k-nearest-neighbour (KNN) (k = 5) inside each fold [54]. The benchmarks are organised across three analysis levels: Level 1 within-person state (participant-task unit, within-person z-score normalisation), Level 2 between-person stable traits (participant unit, sample-level features), and Level 3 group dynamics (group-task unit, groupaggregated features); the normalisation rationale for each level is given above. The separation between within-person state targets and between-person trait targets follows prior personality-aware affect modelling work showing that trait information can support affect inference but requires evaluation choices that preserve inter-individual variation [18, 46, 48]. Four additions to the original B0–B7 suite are introduced here (marked †): B3c and B3d extend Level 1 with social-affective constructs available in the existing survey data; B4c adds BFI Agreeableness to Level 2; and B6b tests whether within-group dispersion features resolve the B6a mean/variance mismatch. Level 1 Within-person state. B3a mental demand is the clearest above-chance signal (area under the curve (AUC) 0.719; Table 3), driven by audio overlap fraction, which we interpret as a proxy for conversational-floor competition [41, 59]. A single audio feature surpasses the full 31-feature model for B3a, confirming that cognitive-load detection in group tasks is primarily a speech-floor phenomenon. B1a valence is moderate (AUC 0.657) and B3c satisfaction is above chance (AUC 0.571); B1b arousal and B2 dominance are near chance, reflecting temporal label–signal mismatch and feature-set limitations respectively. B3d trust pooled is near chance but rises to AUC 0.679 when restricted to T4 (cooperative context), where within-group trust variance is sufficient for fold-local splits to be meaningful. Level 2 Between-person traits (challenge). LOGO-CV AUC for all three personality benchmarks (B4a–B4c) is near or below chance, driven by test folds of only four participants too small for stable AUC estimates regardless of true signal strength, especially when model selection and evaluation are performed under small grouped folds [57, 9]. Spearman correlations on the full sample suggest that trait-related associations may be present (pupil×Agreeableness r = 0.51, p = 0.008; Section J); the B4 trio is therefore better treated as an explicitly unsolved challenge than as evidence of absent signal. B5 contribution is similarly near-chance with high fold variance (AUC 0.429, SD 0.290). Level 3 Group dynamics. B6a, B6b, and B7 all fall at or below the naive baseline, suggesting that task-window aggregates lose the turn-by-turn floor dynamics that govern speaking inequality in multiparty conversation [41, 15, 59]. B6b rules out feature-design as the explanation: an auxiliary binary classifier on the raw speaking-fraction SD achieves AUC 0.952, confirming the signal is definitively present but inaccessible to ridge regression at n = 28 rows. Per-benchmark interpretation notes, modality ablation, and feature-importance analysis are in Section J. Sequential conversation benchmarks (next-speaker prediction, turn-taking, overlap onset) are defined in Section I as future work.
7
Coverage, Quality, and Validation
Table 4 reports modality coverage and task-response completeness; per-session detail is in Section L. Task validity. The key manipulation check is shown in Figure 3: mean valence drops sharply during T2 negotiation (d = 1.06, p = 2.4 × 10−9 ), and T2 overran its 8-min guideline in all 10 groups, indicating strong and consistent affective and behavioural differentiation. Group-level trust remained 6
Table 3: Benchmark definitions and leave-one-group-out results (31-feature leakage-corrected set). Fold-mean performance with across-fold SD and bootstrap 95% CI (5000 iterations). Chance baselines: majority-class Acc. (B0), AUC 0.500 (classification), naive-mean MAE (B6/B7). Modalities: Ph = physiology, Pu = pupil, Au = audio; All = Ph+Pu+Au+BFI. † new addition; [chall.] challenge benchmark; ‡ CI excludes chance. B0–B3d estimates are biased (pre-split normalisation; see Section 5). ID
Target
Sanity check (Pt-task unit) B0 Task label (T1–T4)
n Metric Mod. 136
Acc. Ph+Pu+Au
Level 1 — Within-person state (Pt-task unit) B1a Valence (hi/lo) 107 AUC All B1b Arousal (hi/lo) 107 AUC All B2 Dominance (hi/lo) 83 AUC Ph+Pu+Au B3a Mental demand (hi/lo) 99 AUC All B3b Engagement (hi/lo) 99 AUC All B3c Satisfaction (T2–T3)† 60 AUC All B3d Trust pooled (T2, T4)† 60 AUC All B3d T4-only (corrected)† 28 AUC
Mean (SD)
95% CI
0.641 (0.132) [0.55, 0.73] 0.657 (0.110) [0.58, 0.73] 0.528 (0.114) [0.45, 0.60] 0.499 (0.186) [0.37, 0.62] 0.719 (0.142) [0.62, 0.81] 0.591 (0.136) [0.50, 0.69] 0.571 (0.213) [0.419, 0.733] 0.562 (0.221) [0.406, 0.711] 0.679 (0.220) [0.536, 0.857]
Level 2 — Between-person traits (Participant unit, CHALLENGE) B4a BFI Extraversion‡ 32 AUC Pt-mean all 0.396 (0.353) [0.17, 0.65] B4b BFI Openness‡ 32 AUC Pt-mean all 0.306 (0.244) [0.11, 0.50] B4c BFI Agreeableness† 31 AUC Pt-mean all 0.604 (0.424) [0.292, 0.875] B4c (top-2 T2 feats† 31 AUC 0.625 (0.317) [0.406, 0.844] B5 T4 contribution (median 28 AUC Pt-mean all 0.429 (0.290) [0.21, 0.64] split) Level 3 — Group dynamics (Grp-task unit) B6a Speaking Gini (Mean) 38 MAE Ph+Pu+BFI 0.089 — B6b Speaking Gini (SD )†‡ 28 MAE Ph+Au 0.102 (0.053) [0.068, 0.140] B6b binary (SDraw ) 28 AUC 0.952 (0.117) [0.857, 1.000] B7 Speech-overlap fraction 38 MAE Ph+Pu+BFI 0.063 —
above the scale midpoint after both T2 and T4, confirming collaborative rather than adversarial dynamics; the T2-to-T4 trust change is descriptive rather than statistically reliable at n = 10 groups (t(38) = 0.96, p = 0.34). Cross-modal complementarity and associations. Pupil and wearable physiology are weakly correlated (|r| < 0.10), suggesting these channels provide partly distinct information rather than a redundant arousal readout an empirical argument for multimodal combination [24, 37]. The strongest preprocessed associations are audio group-overlap fraction with post-block mental demand (r = −0.62, qFDR < 0.001, n = 99) and pupil dilation with mental demand (r = +0.36, q = 0.008, n = 99), both motivating benchmark B3a [41, 24]. Extended autonomic and conversation-structure analyses are in Section H.
8
Discussion
The benchmark gradient across the three analysis levels is interpretable as a coherent finding about multimodal signal structure in collaborative meetings, not merely a performance ranking, consistent with the idea that different behavioural and physiological channels carry partly distinct social and affective information [59, 24, 37]. In brief: audio and pupil dominate cognitive-state and satisfaction detection; physiology contributes primarily to arousal-sensitive targets; personality prediction fails under the current normalisation and feature design, not because of signal absence; and group-level dynamics are not recoverable from task-window aggregates at this cohort size. Mental demand and trust context. B3a mental demand is the clearest Level-1 signal (AUC 0.719), driven by a single audio overlap feature that exceeds the full 31-feature model [41, 59]. B3c satisfaction shares the same audio-floor driver and is weaker. B3d trust rises to AUC 0.679 when restricted to cooperative T4 context; T2’s adversarial negotiation leaves near-zero within-group variance, making that fold-local split near-random, which motivates within-task trust probes in future releases [36, 55]. 7
Feature importance per benchmark (mean |coefficient|, LOGO-CV folds) B0 Task label
B1a Valence
Overlap fraction Audio energy mean Pupil mean baseline Pupil slope Audio energy SD SCR count Pause count Speaking time (s) Pitch mean Pupil right mean Pupil left mean Skin temp baseline Shimmer mean EDA phasic mean baseline Voiced seg. mean (s)
Acc 0.641 0.0
0.2
0.4 0.6 Norm. |coef|
0.8
1.0
0.2
0.2
0.4 0.6 Norm. |coef|
0.8
1.0
AUC 0.533 0.0
0.2
B3b Engagement
AUC 0.739 0.0
0.4 0.6 Norm. |coef|
0.8
1.0
Pupil slope Pitch mean SCR count HRV RMSSD Skin temp baseline Audio energy SD Pupil left mean HR SD EDA tonic baseline Pupil mean baseline Speaking time (s) Unvoiced seg. mean (s) High-motion fraction Voiced segs/s Speech rate proxy 0.2
0.4 0.6 Norm. |coef|
Physiology
0.4 0.6 Norm. |coef|
0.8
Overlap fraction Pupil left mean Audio energy SD Shimmer mean Unvoiced seg. mean (s) Pupil mean baseline Speech rate proxy Speaking fraction Audio energy mean EDA phasic rate baseline HR mean baseline Pitch mean Voiced seg. mean (s) SCR count Pitch SD
1.0
AUC 0.499 0.0
0.2
B_sat Satisfaction Overlap fraction Accel motion mean Pupil mean baseline Pitch mean Speaking time (s) Pupil right mean HR SD Jitter mean Speaking fraction HRV RMSSD Skin temp baseline Pitch SD Voiced seg. mean (s) EDA phasic rate baseline HRV quality
AUC 0.583 0.0
B2 Dominance
Audio energy mean Voiced segs/s Speaking fraction Speaking time (s) High-motion fraction HRV RMSSD HR mean baseline Pupil slope Pitch mean HR SD Unvoiced seg. mean (s) Pitch SD Speech rate proxy EDA phasic mean baseline Pupil right mean
AUC 0.668 0.0
B3a Mental demand Overlap fraction HR mean baseline Pause count Pupil slope Voiced segs/s HR SD HRV quality Pupil right mean HRV RMSSD EDA phasic mean baseline Pitch SD High-motion fraction Audio energy mean Speaking time (s) HNR mean
B1b Arousal
Overlap fraction Skin temp baseline EDA phasic mean baseline Speaking time (s) Pause count Pupil mean baseline Unvoiced seg. mean (s) Accel motion mean SCR count HRV RMSSD HR SD Audio energy mean Pupil slope EDA tonic baseline Speech rate proxy
0.8
1.0
Eye-tracking
Audio
0.2
0.4 0.6 Norm. |coef|
0.8
0.8
1.0
B_trust Trust
AUC 0.571 0.0
0.4 0.6 Norm. |coef|
1.0
Speaking time (s) Shimmer mean Pause count Pupil left mean HR SD Audio energy mean Overlap fraction Jitter mean HR mean baseline Voiced segs/s HRV quality Accel motion mean SCR count HRV RMSSD Speech rate proxy
AUC 0.562 0.0
0.2
0.4 0.6 Norm. |coef|
0.8
1.0
Motion/Temp
Figure 2: Per-benchmark ranked feature importance: top-15 features by mean normalised |coefficient| across LOGO-CV folds (31-feature set). Bar colour: Physiology, Eye-tracking, Audio, Motion/Temp. ⋆ = audio_overlap_fraction_x. B3a mental demand is dominated by a single audio-floor feature; B3b engagement loads instead on pupil slope and pitch, confirming a demand– engagement dissociation. Full benchmark notes and modality ablation in Section J. Table 4: Dataset statistics and task-response completeness. QC pass: 10 sessions, 200 expected participant-task rows. Signal usability per channel (EmotiBit: ≥80% finite samples; Tobii gaze: ≥70% valid). Metric
Value
Groups / unique participants Age, mean (SD), range Sex Scheduled recording time
10 / 40 35.4 (9.2), 23–58 21F / 19M ∼11.7 h
Signal QC (200 expected rows) EmotiBit rows observed PPG / EDA / Temp. usable IMU usable Tobii rows observed Gaze / Pupil usable
182/200 (91.0%) 76.0% / 75.5% / 76.0% 181/200 (90.5%) 196/200 (98.0%) 87.2% / 95.4%
Task-response completeness (all 10 groups completed all 4 tasks) VAD rows (T1/T2/T3/T4) Core post-task items (T1/T2/T3/T4) Overall post-task completeness T1 modal decision / T2 settlement T3 winning-idea records
34 / 36 / 38 / 35 of 40 712/712 / 590/600 / 633/640 / 319/320 2254/2280 (98.9%) 10/10 / 10/10 7/10 groups
Modality complementarity. The feature importance profiles in Figure 16 and modality ablation (Section J) together suggest audio and pupil provide largely non-redundant signal for cognitivestate detection, consistent with the near-zero physiology–pupil correlation (|r| < 0.10; Section 7) [59, 24, 37]. Physiology contributes primarily to arousal-sensitive targets via EDA phasic rate (B1b arousal, and to a lesser extent B1a valence) [37]; B2 dominance is near chance. Between-person inference. We interpret weak B4 results as reflecting the normalisation regime: within-person z-scoring removes inter-individual amplitude variance needed for between-person inference [57, 9]. Future work should use fold-internal normalisation fitted on training participants only. Group-level inequality. B6a, B6b, and B7 all fail the naive baseline. B6b rules out feature-design: even within-group SD features fail (Ridge MAE 0.102 vs baseline 0.088), while raw SD binary classification reaches AUC 0.952 confirming the signal exists but task-window aggregates lose the 8
Figure 3: Valence, arousal, and dominance probe summaries by task. T2 negotiation produces the largest valence drop (d = 1.06) the strongest manipulation check in the dataset. Dominance is absent for T4 in the current response schema. turn-by-turn dynamics that govern floor inequality [41, 15, 59]. Per-turn voice-activity-detection extraction and a larger cohort (n ≥ 20 groups) are the prerequisites for this benchmark tier.
9
Limitations
Sample size. 40 participants across 10 groups supports dataset characterisation but not broad benchmark or leaderboard claims; LOGO-CV test folds of four participants make AUC estimates inherently inaccurate for participant-level targets. Single site and language. All sessions were recorded at a single site and conducted in English; turn-taking norms and negotiation styles may reflect local linguistic, cultural, and institutional context [41, 15]. Fixed task order. All sessions followed the same T0–T4 sequence with no counterbalancing; T4 cooperation outcome follows ≈60 min of prior collaboration and cannot be interpreted as a pure task effect without a counterbalanced replication. Audio and modality scope. Raw audio is gated under a DUA due to re-identification risk [31, 53]; lapel microphones can misclassify ambient noise; released gaze is in each participant’s scene-camera frame and requires additional processing for cross-participant analysis. Additional technical caveats (physiology precision, pupil comparability, survey ceiling, deferred modalities, no clinical ground truth) are in Section C.
10
Ethics, Access, and Responsible Use
The study was conducted under a participant information statement and informed consent procedure reviewed and approved by the GN Hearing A/S Legal Department. Participants provided written informed consent covering all modalities. Tabular features are released under CC BY 4.0; raw audio is gated under a Data Use Agreement prohibiting re-identification, voice cloning, and non-research use [31, 53]. Released files use anonymous seat identifiers (P1–P4). The dataset is not appropriate for clinical diagnosis, individual scoring, automated personnel evaluation, or surveillance [13, 28]. Detailed responsible-use guidance, disaggregated reporting expectations, and a 5-year maintenance commitment under a persistent-DOI repository are in the datasheet [32].
11
Conclusion
GroupAffect-4 contributes a compact but unusually integrated resource for studying co-located group affect: synchronised physiology, egocentric eye tracking, audio, self-report, task outcomes, and personality measures from 10 four-person groups. Its value lies in documentation, synchronisation, and multimodal density around structured group tasks rather than scale. The feasibility baselines show that the release supports meaningful evaluation protocols while also exposing where current data are too small for broad claims. We release the dataset with Croissant metadata, a datasheet, explicit QC tables, and conservative responsible-use guidance [13, 1, 29].
9
References [1] Mubashara Akhtar, Omar Benjelloun, Costanza Conforti, Luca Foschini, Joan Giner-Miguelez, Pieter Gijsbers, Sujata Goswami, Nitisha Jain, Michalis Karamousadakis, Michael Kuchnik, et al. Croissant: A metadata format for ML-ready datasets. In Advances in Neural Information Processing Systems, 2024. Also available as arXiv:2403.19546. [2] Sigal G. Barsade. The ripple effect: Emotional contagion and its influence on group behavior. Administrative Science Quarterly, 47(4):644–675, 2002. [3] Indrani Bhattacharya, Daniel Foley, Nicholas Zhang, Tong Zhang, Christopher Mine, Qi Ji, and Richard J. Radke. UGI: An unobtrusive group interaction dataset. In Proceedings of the 10th ACM Multimedia Systems Conference, 2019. [4] Patricia Bota, Joana Brito, Ana Fred, Pablo Cesar, and Hugo Plácido Silva. G-REx: A real-world dataset of group emotion experiences based on physiological data. 2023. [5] Margaret M. Bradley and Peter J. Lang. Measuring emotion: The self-assessment manikin and the semantic differential. Journal of Behavior Therapy and Experimental Psychiatry, 25(1):49–59, 1994. [6] Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan. IEMOCAP: Interactive emotional dyadic motion capture database. Language Resources and Evaluation, 42(4):335– 359, 2008. [7] Laura Cabrera-Quiros, Andrew Demetriou, Egor Balog, Astrid van der Meijden, Esma Gedik, Binyam G. Gebre, and Hayley Hung. The MatchNMingle dataset: A novel multi-sensor resource for the analysis of social interactions and nonverbal communication in unstructured mingle and speed-dating scenarios. IEEE Transactions on Affective Computing, 12(1):148–164, 2021. [8] Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, Guillaume Lathoud, Mike Lincoln, Agnes Lisowska, Iain McCowan, Wilfried Post, Dennis Reidsma, and Pierre Wellner. The AMI meeting corpus: A pre-announcement. In Proceedings of the Second International Workshop on Machine Learning for Multimodal Interaction, pages 28–39. Springer, 2005. [9] Gavin C. Cawley and Nicola L. C. Talbot. On over-fitting in model selection and subsequent selection bias in performance evaluation. Journal of Machine Learning Research, 11:2079–2107, 2010. [10] Ananish Chaudhuri. Sustaining cooperation in laboratory public goods experiments: A selective survey of the literature. Experimental Economics, 14(1):47–83, 2011. [11] Mark Dourado, Henrik Gert Hassager, Jesper Udesen, and Stefania Serafin. The gamma corpus of danish polyadic conversations with gaze speech and motion data in quiet and noise. Scientific Data, 2026. [12] Ernst Fehr and Simon Gächter. Cooperation and punishment in public goods experiments. American Economic Review, 90(4):980–994, 2000. [13] Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daume III, and Kate Crawford. Datasheets for datasets. Communications of the ACM, 64(12):86–92, 2021. [14] Kyriakos Georgiou, Andreas V. Larentzakis, Nader N. Khamis, Ghadah I. Alsuhaibani, Yasser A. Alaska, and Elias J. Giallafos. Can wearable devices accurately measure heart rate variability? a systematic review. Folia Medica, 60(1):7–20, 2018. [15] Charles Goodwin. Conversational Organization: Interaction Between Speakers and Hearers. Academic Press, 1981. 10
[16] Krzysztof J. Gorgolewski, Tibor Auer, Vince D. Calhoun, R. Cameron Craddock, Samir Das, Eugene P. Duff, Guillaume Flandin, Satrajit S. Ghosh, Tristan Glatard, Yaroslav O. Halchenko, et al. The brain imaging data structure, a format for organizing and describing outputs of neuroimaging experiments. Scientific Data, 3:160044, 2016. [17] Adam Janin, Don Baron, Jane Edwards, Dan Ellis, David Gelbart, Nelson Morgan, Barbara Peskin, Thilo Pfau, Elizabeth Shriberg, Andreas Stolcke, and Chuck Wooters. The ICSI meeting corpus. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. [18] O. P. John and S. Srivastava. The big five trait taxonomy: History, measurement, and theoretical perspectives. In L. A. Pervin and O. P. John, editors, Handbook of personality: Theory and research, pages 102–138. Guilford Press, 2nd edition, 1999. [19] Sander Koelstra, Christian Muhl, Mohammad Soleymani, Jong-Seok Lee, Ashkan Yazdani, Touradj Ebrahimi, Thierry Pun, Anton Nijholt, and Ioannis Patras. DEAP: A database for emotion analysis using physiological signals. IEEE Transactions on Affective Computing, 3(1):18–31, 2012. [20] Christian Kothe, Seyed Yahya Shirazi, Tristan Stenner, David Medine, Chadwick Boulay, Matthew I. Grivich, Fiorenzo Artoni, Tim Mullen, Arnaud Delorme, and Scott Makeig. The lab streaming layer for synchronized multimodal recording. Imaging Neuroscience, 3:IMAG.a.136, 2025. [21] Richard S. Lazarus. Emotion and Adaptation. Oxford University Press, 1991. [22] Li Lu, Y. Connie Yuan, and Poppy Lauretta McLeod. Twenty-five years of hidden profiles in group decision making: A meta-analysis. Personality and Social Psychology Review, 16(1):54– 75, 2012. [23] Michelle A. Marks, John E. Mathieu, and Stephen J. Zaccaro. A temporally based framework and taxonomy of team processes. Academy of Management Review, 26(3):356–376, 2001. [24] Sebastiaan Mathot. Pupillometry: Psychology, physiology, and function. Journal of Cognition, 1(1):16, 2018. [25] Albert Mehrabian and James A. Russell. An Approach to Environmental Psychology. MIT Press, 1974. [26] Diego Miranda, Carlos Escobedo, Dayana Palma, Rene Noel, Adrián Fernández, Cristian Cechinel, Jaime Godoy, and Roberto Munoz. A multimodal experimental dataset on agile software development team interactions. Data in Brief, 61:111828, 2025. [27] Juan Abdon Miranda-Correa, Mojtaba Khomami Abadi, Nicu Sebe, and Ioannis Patras. AMIGOS: A dataset for affect, personality and mood research on individuals and groups. IEEE Transactions on Affective Computing, 12(2):479–493, 2021. [28] Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pages 220– 229, 2019. [29] MLCommons. Croissant format specification, version 1.0. https://docs.mlcommons.org/ croissant/docs/croissant-spec.html, 2024. Published 2024-03-01; accessed 2026-0503. [30] Philipp Müller, Michael Xuelin Huang, and Andreas Bulling. Detecting low rapport during natural interactions in small groups from non-verbal behaviour. pages 153–164, 2018. [31] Andreas Nautsch, Andrés Jiménez, Alexander Treiber, Jan Kolberg, Catherine Jasserand, Els Kindt, Héctor Delgado, Massimiliano Todisco, Pierre Héroux, Nicholas Evans, et al. Preserving privacy in speaker and speech characterisation. Computer Speech & Language, 58:441–480, 2019. 11
[32] NeurIPS. NeurIPS 2026 evaluations & datasets hosting guidelines. https://neurips.cc/ Conferences/2026/EvaluationsDatasetsHosting, 2026. Accessed 2026-05-01. [33] NeurIPS. NeurIPS 2026 evaluations & datasets track — call for papers. https://neurips. cc/Conferences/2026/CallForEvaluationsDatasets, 2026. Accessed 2026-05-01. [34] Cristina Palmero, Javier Selva, Sorina Smeureanu, Julio C. S. Jacques Junior, Albert Clapés, Alba Moseguí, Zejian Zhang, David Gallardo, Georgina Guilera, David Leiva, Hugo Jair Escalante, Isabelle Guyon, Xavier Baró, and Sergio Escalera. Context-aware personality inference in dyadic scenarios: Introducing the UDIVA dataset. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, 2021. [35] Cheul Young Park, Narae Cha, Soowon Kang, Auk Kim, Ahsan H. Khandoker, Leontios J. Hadjileontiadis, Alice Oh, Youn-Byoung Jeong, and Uichin Lee. K-EmoCon, a multimodal sensor dataset for continuous emotion recognition in naturalistic conversations. Scientific Data, 7(1):293, 2020. [36] Brian Parkinson. Emotions are social. British Journal of Psychology, 87(4):663–683, 1996. [37] Hugo F. Posada-Quintero and Ki H. Chon. Innovations in electrodermal activity data collection and signal processing: A systematic review. Sensors, 20(2):479, 2020. [38] Eric F. Rietzschel, Bernard A. Nijstad, and Wolfgang Stroebe. The selection of creative ideas after individual idea generation: Choosing between creativity and impact. British Journal of Psychology, 101(1):47–68, 2010. [39] Fabien Ringeval, Andreas Sonderegger, Jürgen Sauer, and Denis Lalanne. Introducing the RECOLA multimodal corpus of remote collaborative and affective interactions. In Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition Workshops, 2013. [40] James A. Russell. A circumplex model of affect. Journal of Personality and Social Psychology, 39(6):1161–1178, 1980. [41] Harvey Sacks, Emanuel A. Schegloff, and Gail Jefferson. A simplest systematics for the organization of turn-taking for conversation. Language, 50(4):696–735, 1974. [42] Dairazalia Sanchez-Cortes, Oya Aran, Marianne Schmid Mast, and Daniel Gatica-Perez. A multimodal corpus for the study of small group interactions. In Proceedings of the International Conference on Multimodal Interaction Workshops, 2011. [43] Klaus R. Scherer. Appraisal considered as a process of multilevel sequential checking. In Klaus R. Scherer, Angela Schorr, and Tom Johnstone, editors, Appraisal Processes in Emotion: Theory, Methods, Research, pages 92–120. Oxford University Press, 2001. [44] Björn Schuller, Anton Batliner, Stefan Steidl, and Dino Seppi. Recognising realistic emotions and affect in speech: State of the art and lessons learnt from the first challenge. Speech Communication, 53(9–10):1062–1087, 2011. [45] Meisam J. Seikavandi, Laurits Dixen, Jostein Fimland, Sree Keerthi Desu, Antonia-Bianca Zserai, Ye Sul Lee, Maria Barrett, and Paolo Burelli. Advancing face-to-face emotion communication: A multimodal dataset (AFFEC). arXiv preprint arXiv:2504.18969, 2025. [46] Meisam J. Seikavandi, Jostein Fimland, Fabricio Batista Narcizo, Maria Barrett, Ted Vucurevich, Jesper Bünsow Boldt, Andrew Burke Dittberner, and Paolo Burelli. Modelling the interplay of eye-tracking temporal dynamics and personality for emotion detection in face-to-face settings. arXiv preprint arXiv:2510.24720, 2025. [47] Meisam Jamshidi Seikavandi and Maria Jung Barrett. Gaze reveals emotion perception: Insights from modelling naturalistic face viewing. In Proceedings of the 22nd IEEE International Conference on Machine Learning and Applications (ICMLA), pages 2022–2025. IEEE, 2023. 12
[48] Meisam Jamshidi Seikavandi, Fabricio Batista Narcizo, Ted Vucurevich, Andrew Burke Dittberner, and Paolo Burelli. MuMTAffect: A multimodal multitask affective framework for personality and emotion recognition from physiological signals. In Proceedings of the 3rd International Workshop on Multimodal and Responsible Affective Computing, pages 100–108, 2025. [49] Fred Shaffer and J. P. Ginsberg. An overview of heart rate variability metrics and norms. Frontiers in Public Health, 5:258, 2017. [50] Garold Stasser and William Titus. Pooling of unshared information in group decision making: Biased information sampling during discussion. Journal of Personality and Social Psychology, 48(6):1467–1478, 1985. [51] Task Force of the European Society of Cardiology and the North American Society of Pacing and Electrophysiology. Heart rate variability: Standards of measurement, physiological interpretation, and clinical use. Circulation, 93(5):1043–1065, 1996. [52] Sara Taylor, Natasha Jaques, Weixuan Chen, Szymon Fedor, Akane Sano, and Rosalind W. Picard. Automatic identification of artifacts in electrodermal activity data. In Proceedings of the Annual International Conference of the IEEE Engineering in Medicine and Biology Society, pages 1934–1937, 2015. [53] Natalia Tomashenko, Xin Wang, Emmanuel Vincent, Jose Patino, Brij Mohan Lal Srivastava, Paul-Gauthier Noé, Andreas Nautsch, Nicholas Evans, Junichi Yamagishi, Benjamin O’Brien, et al. The VoicePrivacy 2020 challenge: Results and findings. In Proceedings of Interspeech, pages 1399–1403, 2021. [54] Olga Troyanskaya, Michael Cantor, Gavin Sherlock, Pat Brown, Trevor Hastie, Robert Tibshirani, David Botstein, and Russ B. Altman. Missing value estimation methods for DNA microarrays. Bioinformatics, 17(6):520–525, 2001. [55] Gerben A. Van Kleef. How emotions regulate social life: The emotions as social information (EASI) model. Current Directions in Psychological Science, 18(3):184–188, 2009. [56] Gerben A. Van Kleef, Carsten K. W. De Dreu, and Antony S. R. Manstead. The interpersonal effects of emotions in negotiations: A motivated information processing approach. Journal of Personality and Social Psychology, 87(4):510–528, 2004. [57] Sudhir Varma and Richard Simon. Bias in error estimation when using cross-validation for model selection. BMC Bioinformatics, 7:91, 2006. [58] Roel Vertegaal, Robert Slagter, Gerrit van der Veer, and Anton Nijholt. Eye gaze patterns in conversations: There is more to conversational agents than meets the eyes. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, pages 301–308, 2001. [59] Alessandro Vinciarelli, Maja Pantic, and Hervé Bourlard. Social signal processing: Survey of an emerging domain. Image and Vision Computing, 27(12):1743–1759, 2009. [60] Mark D. Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, et al. The FAIR guiding principles for scientific data management and stewardship. Scientific Data, 3:160018, 2016. [61] Zhihong Zeng, Maja Pantic, Glenn I. Roisman, and Thomas S. Huang. A survey of affect recognition methods: Audio, visual, and spontaneous expressions. IEEE Transactions on Pattern Analysis and Machine Intelligence, 31(1):39–58, 2009. [62] Justine Zhang, Ravi Kumar, and Cristian Danescu-Niculescu-Mizil. GAP corpus: Group affect and performance corpus. https://convokit.cornell.edu/documentation/gap.html, 2018. Dataset documentation.
13
Appendix Table of Contents A B C D E F G H I J K L M N
List of Acronyms Stimuli and Task Orchestration Extended Limitations and Caveats BFI-44 Scoring and Item List Audio T0 Baseline Reliability Synchronisation Pipeline Detail Preprocessing Steps Extended Dataset Characterization Extended Benchmarks: Sequential Conversation Tasks Benchmark Interpretation Notes Full Ablation Table Per-Session Quality Table Responsible AI and Croissant Metadata Datasheet for Datasets
A
List of Acronyms
Acronym
Expansion
Acronym
Expansion
AUC BFI-44 BH-FDR BIDS dBFS DUA EDA EEG ET FAIR GSR HR HRV IMU KNN
Area under the curve Big-Five Inventory Benjamini–Hochberg false discovery rate Brain Imaging Data Structure Decibels relative to full scale Data Use Agreement Electrodermal activity Electroencephalography Eye tracking Findable, Accessible, Interoperable, and Reusable Galvanic skin response Heart rate Heart-rate variability Inertial measurement unit k-nearest-neighbour
LOGO-CV LSL MAD MAE NGT PAD PPG QC RMSSD RR SAM SCR SSE VAD VFOA
Leave-one-group-out cross-validation Lab Streaming Layer Median absolute deviation Mean absolute error Nominal group technique Pleasure-Arousal-Dominance Photoplethysmography Quality control Root mean square of successive differences R-R (inter-beat) interval Self-Assessment Manikin Skin conductance response Server-Sent Events Valence-Arousal-Dominance Visual focus of attention
XDF
Extensible Data Format
Note on VAD. Throughout this paper VAD denotes Valence-Arousal-Dominance (the three dimensions of the affect rating probes). Voice-activity detection—a distinct signal-processing operation that segments speech from silence in audio recordings—is always written in full to avoid confusion.
B
Stimuli and Task Orchestration
This appendix describes the complete stimuli suite and session-orchestration system used during data collection. All four tasks were delivered by a custom two-tier web application running on the Recording PC: a display server pushed content to (a) five Android tablets (one for each participant and one for the moderator) over local Wi-Fi and (b) a big screen via HDMI; a task runner serialised phase transitions and sent LSL markers for every content change. Server-Sent Events (SSE) ensured near-zero-latency push without polling.
Release Modalities Participant Demographics Session Architecture A typical session lasted approximately 70 minutes and was preceded by an onboarding phase in which participants were briefed, fitted with the equipment, and calibration procedures were completed. Participants sat around a rectangular table, each with an Android tablet. Content was participant-specific (evidence cards, role cards) and never visible to others. Participants were instructed by the moderator to not show their tablets to others. The display mounted to a wall at the end of a table ("big screen") showed shared task briefs, phases’ instructions, and a timer. The moderator controlled the tasks’ phases using a web application on a tablet.
Task Descriptions T0 Onboarding and Baseline (10 min). After providing written informed consent, participants were onboarded by an experimenter who explained the session procedures and sensing equipment, assisted with device setup and calibration, and then conducted a 5-minute free-talk warm-up (self-introductions; T0). T0 serves as a recording baseline and is excluded from the affective benchmarks.
14
Table 5: Modalities released in this version of GroupAffect-4. All per-participant streams use seat identifiers P1–P4. Sampling rates marked with † are channel-dependent; see the per-channel table in the appendix. Transcript artifacts are included in the reviewer-access package. Multi-camera video and marker-assisted 3D outputs were recorded or piloted for follow-up work; they are not part of the current release or analysis. Modality
Device
Signals
Rate
Per session
Physiology
EmotiBit
PPG (3 wavelengths), EDA, tempera- † ture, IMU
4 (one per P)
Eye tracking (ego- Tobii Pro Glasses 3 centric)
Head-relative scene-frame gaze, pupil 50 Hz diameter, validity
4 (one per P)
Audio (close-talk)
DPA 4060
Mono, 48 kHz / 16-bit
48 kHz
4 (one per P)
Audio (room)
DPA 4060
Mono, 48 kHz / 16-bit
48 kHz
Transcript artifacts
Audio post-processing
Transcript and turn-taking tables de- event-driven rived from participant audio
4 participant channels
1
Behavioural report
self- Android tablets
VAD probes, task responses
event-driven
4 (one per P)
Personality session)
(pre- Online questionnaire
BFI-44 item responses, demographics
one-shot
4 (one per P)
Individual-level demographics
Count
A
Group composition (10 groups, N=4 each)
G10
5 0
Group-level composition
B
Age distribution (N=40)
G9
med=34
20
25
30
Sex F 21
M 19
35
40
45
50
Age (years)
6
PhD Master Bachelor Other
G7 G6
25
G5
3
English proficiency (N=40) Fluent 30
G8
60
6 0
Native 7
55
Education
10
N
20
30
G4 G3
Int. 3
G2 G1 20
30
40
50
Age (years)
60
Age range Age mean Female Male Sex mix
70
Figure 4: Participant demographics (age, sex, education) and group composition across 40 participants in 10 four-person groups. T1 Hidden-Profile Decision Task (10 min). Participants acted as a hiring committee selecting one of three fictional candidates (A, B, C) for an internal “AI Adoption & Transformation Specialist” role. The task followed a hidden-profile structure: each participant viewed on their tablet a private evidence card containing information about all three candidates, with some items identical across all cards (shared evidence) and others present on only one card (unique evidence). Participants were unaware that their information differed from that of other group members, and no single card contained sufficient evidence to identify the normatively correct choice. After 75 seconds of silent individual reading, the group discussed freely. As discussion ended, the moderator instructed the group to choose one person to record the decision on the tablet.
T2 Mini-Negotiation Task (15 min). The group jointly planned a quarterly internal workshop, agreeing on both a topic (four options: AI for productivity, cross-team collaboration, environmental sustainability, stress/workload management) and a format (four options: online training + quiz, lecture + Q&A, interactive workshop, cross-department working group). Each participant received a private role card prescribing a priority dimension (topic or format) and a preference direction, with an explicit instruction to advocate strongly for that position. After a 7 minutes discussion, the moderator instructed the group to choose one person to send a settlement form on the tablet. Figures 5 through 8 show the T1–T4 big-screen layouts.
T3 Idea Generation and Selection (18 min). Participants generated ideas for a GN-wide social event: first independently writing up to three concrete event ideas on their tablet forms during a 3-minute silent phase, then discussing all ideas as a group, and finally reaching consensus on one winning idea recorded on a settlement form by one of the participants. (Figure 7).
T4 Public-Goods Micro-Game (12 min). Each participant was tasked to imagine they received 10 tokens and could privately decided how many to contribute to a shared fund with colleagues. The fund was multiplied by 1.5 and divided
15
equally among all four participants, creating a classic social-dilemma structure: payouti = (10 − ci ) + 14 · 1.5
X
cj
ci ∈ {0, . . . , 10}
j
Contributions were submitted privately; group outcome was revealed immediately, then discussion followed after individual token contributions were shown on the main display.(Figure 8).
Task 1: Hidden-Profile Decision -- Group Discussion
~20 min
Role Being Filled: AI Adoption & Transformation Specialist (Internal) An internal position supporting AI adoption and transformation across different departments at GN. Focuses on helping teams understand the potential of AI, identify opportunities, and support their rollout.
1
2
3
Silent reading (1 min 15 s)
Group discussion (~15 min)
Decision (1 min)
Read your private evidence card. Do not talk yet.
Share what you know -- paraphrase, do not read aloud from your tablet.
Agree on ONE candidate. Record your group's choice on the tablet form.
Each participant has DIFFERENT private information on their tablet. Discuss and share -- do not show your screen to others. Tablets: 4 unique evidence cards (private, 1 per participant) Big screen: role description only
Figure: Big-Screen Display -- T1 Hidden-Profile Task
Figure 5: T1 (Hidden-Profile Decision) big-screen layout. The shared display shows the task brief, phase instructions, and a countdown timer. Participant evidence cards were shown only on individual tablets.
Task 2: Mini-Negotiation -- Workshop Planning
8-12 min
Plan next quarter's internal GN workshop -- agree on ONE topic AND ONE format. Each participant holds a private role card (on their tablet) with a priority and preferences. Advocate for your position -- do not show your screen to others.
Workshop Topics
Workshop Formats
T1 AI for productivity in daily tasks
F1 Online Training + Quiz
T2 Working better across teams
F2 Lecture + Q&A
T3 Environmental sustainability commitments
F3 Interactive Informal Workshop
T4 Managing stress and workload
F4 Cross-Department Working Group
Role card reading
Discussion & negotiation
Settlement form
Check your tablet for your role card
Tablets: private role cards (different priority per participant) Settlement form sent to tablets at end Big screen: shared reference only
Figure: Big-Screen Display -- T2 Mini-Negotiation Task
Figure 6: T2 (Mini-Negotiation) big-screen layout. The shared display shows the topic and format options; each participant’s role card was private.
16
Task 3: Idea Generation (NGT) -- GN Social Event
~18 min
Generate ideas for a GN-wide social event. Be original -- propose events with a specific theme or concrete activity. 1 Silent individual generation (3 min, no discussion) 2 Round-robin sharing (each participant reads ideas aloud) 3 Group selects ONE best idea
P1
P2
P3
P4
Escape room challenge (team of 4-6, multiple rooms)
Science trivia quiz night (pub-style, 6 rounds)
Board game cafe evening (curated selection, light snacks)
Creative writing sprint (flash fiction, group anthology)
Photography walk (guided tour + creative challenge)
Pottery workshop (3-hour beginner session)
Improv comedy workshop (professional facilitator)
Indoor climbing session (beginner-friendly, instructor-led)
Cooking class competition (chef-led, themed cuisine)
City treasure hunt (app-guided, themed clues)
Nature hike + picnic (local trail, catered lunch)
Pub quiz + dinner (external venue, themed evening)
After round-robin, discuss and select ONE best idea as a group. Record the winning idea and its author on the tablet form (T3_GROUP_SELECTION_FORM). Tablets: individual idea generation sheet (3 ideas, 3 min timer) Big screen: shared instructions + live idea board (populated during round-robin)
Figure: Big-Screen Display -- T3 Idea Generation (NGT)
Figure 7: T3 (Idea Generation and Selection) big-screen layout. The shared display shows the social-event brief and the group’s submitted ideas during the discussion phase.
Task 4: GN Community Fund -- Contribution Decision Task 4: GN Community Fund
~12 min
Worked example: everyone contributes 8
Each participant receives 10 tokens.
Group total contribution
32
Tokens you CONTRIBUTE go to the community pot.
Pot after multiplier (32 x 1.5)
48
Group pot is multiplied by 1.5, then split equally among all 4.
Each receives from pool (48 / 4)
12
Tokens you kept (10 - 8)
2
Tokens you KEEP go directly to your account.
Your decision is PRIVATE -- only the group total will be shown.
Final amount = tokens kept + your share from the pool
Final total (kept + pool share)
2+12 = 14
Decide how many tokens to contribute on your tablet. Your choice is PRIVATE -- only the group total will be shown on this screen.
Instructions
Private contribution (1 min)
Group outcome reveal
Discussion & reflection
Tablets: private contribution form (0-10 tokens, 60 s timer) Big screen: instructions -> group total + pool calculation after all submit
Figure: Big-Screen Display -- T4 Public Goods Micro-Game
Figure 8: T4 (Public-Goods Micro-Game) big-screen layout. The shared display reveals individual contributions after all participants have submitted privately, then shows the group payout. Task 2 Role Cards Each participant received a private role card indicating (i) a primary negotiation priority (topic vs. format) and (ii) a secondary preference: • Participant 1 (Role Card A). Primary priority: format. Secondary preference: people and culture topic (Topics 3–4). • Participant 2 (Role Card B). Primary priority: topic (productivity and efficiency; Topics 1–2). Secondary preference: independent-learning format. • Participant 3 (Role Card C). Primary priority: format (independent-learning). Secondary preference: productivity and efficiency topic (Topics 1–2). • Participant 4 (Role Card D). Primary priority: topic (people and culture; Topics 3–4). Secondary preference: collaborative and engaging format.
17
Continuous VAD Probes During each task, participants received periodic affect check-in prompts on their tablets (Figure 9). Three constructs were rated on a 1–9 scale: • Valence: “How pleasant or unpleasant do you feel right now?” (1 = very unpleasant, 9 = very pleasant) • Arousal: “How activated or alert do you feel right now?” (1 = very calm / low energy, 9 = very activated / highly alert) • Dominance: “How much control or influence do you feel right now?” (1 = very little influence, 9 = a great deal of influence) For T4, we omitted the dominance dimension because the anonymous contribution setup does not provide a clear sense of personal control over others or the group outcome. Probes were scheduled at task-specific phase-aligned time points with a small temporal jitter, so that prompts fell near the midpoints of key interaction phases (Table 6). Across a full four-person session, this schedule yields up to three VAD prompts per participant in T0–T3 and two prompts in T4.
Figure 9: Participant tablet view of a VAD affect probe during Task 1. Valence, Arousal, and Dominance are shown simultaneously on a 1–9 Likert scale. Post-Block Questionnaires After each task, participants completed a short individual questionnaire on their tablets (Table 7 for T1, Table 8 for T2, Table 9 for T3, and Table 10 for T4). These post-block items capture both self-directed evaluations (e.g., perceived pleasantness, engagement, mental demand, satisfaction, fairness) and other-directed appraisals of each group member. In the tasks that include trust, each participant rated their trust in each of the other three group members once, yielding a directed dyadic trust network with up to 4 × 3 = 12 directed trust ratings per task and group. Similarly, post-task dominance items asked how dominant each of the four participants (P1–P4) seemed during the task, providing per-seat dominance ratings from every group member.
Table 6: In-task VAD probes and post-task items by task. Task T1: Hidden-profile decision T2: Mini-negotiation T3: Idea generation & selection T4: Public-goods game
VAD probes (per participant) 3 3 3 2
VAD dimensions Valence, arousal, dominance Valence, arousal, dominance Valence, arousal, dominance Valence, arousal (no dominance)
# post-task items 18 15 16 13
Note: During onboarding (T0) the moderator sometimes prompted participants to try the VAD panel, so a small number of VAD submissions appear for T0 in the logs, but there is no formal post-task questionnaire for T0. Hard behavioural outcomes are stored in the beh/ BIDS subfolder: the T1 candidate selected (A/B/C; C is the normatively correct hidden-profile choice), the T2 agreed topic–format pair, the T3 winning idea with authorship attribution, and each participant’s T4 token contribution (0–10).
C
Extended Limitations and Caveats
Physiology precision. PPG-derived HRV at ≈25 Hz imposes a ≈40 ms RMSSD quantisation floor; wearable HRV is not equivalent to ECG-grade measurement (Section G).
18
Figure 10: Participant tablet view of the post-block questionnaire (T4 shown, 7 items shown). All items use a 1–7 Likert scale. Participants respond individually. Pupil and audio comparability. Pupil summaries remain sensitive to illumination and display context; the released audio features are not baseline-normalised because T0 free-talk audio is insufficiently controlled, so multimodal comparisons mix relative physiology/pupil features with absolute audio features.
Survey ceiling and role coverage. Some social-evaluative ratings are high and stable across tasks (e.g. voice/inclusion means 5.85±1.23 in T1, 4.77±1.66 in T2, 5.90±1.11 in T3), which could reflect genuine inclusivity or acquiescence bias. The dataset does not contain direct measures of emergent leadership, expertise attribution, or role crystallisation. Voice re-identification. Close-talk audio carries voice re-identification risk, motivating the separate access tier and DUA described in Section 10.
Egocentric gaze. Released gaze is in each participant’s own scene-camera frame; cross-participant gaze-target analysis requires room-frame alignment not included in this release.
Deferred modalities. Multi-camera video, marker-assisted 3D pose, and room-frame gaze alignment were recorded or piloted but are reserved for a future release.
No clinical ground truth. The dataset must not be used for diagnosis, mental-health inference, personnel evaluation, or surveillance-style deployment.
D
BFI-44 Scoring and Item List
BFI-44 domain scores are computed as item means after applying the standard reverse-coding rules. The release datasheet includes the item-level mapping, reverse-coded item list, and missing-response policy.
E
Audio T0 Baseline Reliability
Audio features are used in two normalization regimes depending on benchmark type. Within-person benchmarks (B0–B3: task classification, affective state) apply within-person z-score normalisation across tasks T1–T4 only, capturing task-induced variance. Between-person benchmarks (B4–B5: personality, contribution) use absolute audio features (no baseline subtraction) to preserve between-person amplitude differences required for trait prediction. When we tested normalising audio features by T0 free-talk baseline across all participants, every tested feature showed increased variance compared to absolute form, indicating T0 introduction of noise.
Why T0 is unreliable. The unstructured T0 session differs fundamentally from structured tasks T1–T4 (Table 11). Absence of task constraints leads to natural heterogeneity: some participants dominate, others remain quiet; vocal energy and pitch vary with conversational mood; social dynamics are unpredictable. T0 is not a neutral biological baseline but rather reflects initial social calibration, novelty effects, and potential social anxiety—orthogonal to individual vocal traits. Third, audio metrics (speaking fraction, pause count, overlap) are group properties, not individual traits; T0 free-talk creates role instability (variable who-speaks-vs-listens patterns) that cannot be attributed to individual differences. 19
Table 7: Task 1 (Hidden-profile decision) post-task questionnaire items. Label Question text Familiarity – P1 How well did you know P1 before this session? Familiarity – P2 How well did you know P2 before this session? Familiarity – P3 How well did you know P3 before this session? Familiarity – P4 How well did you know P4 before this session? Overall Valence Overall, how pleasant or unpleasant did this task feel? Perceived Influence How much influence did you feel you had during the task? Engagement How engaged/absorbed did you feel during the task? Mental Demand How difficult/mentally demanding did the task feel? Team Coordination How well did your group coordinate during the task? Voice / Inclusion To what extent did you feel that your views were heard and taken into account? Information Sharing To what extent do you feel all relevant information was shared within the group before making a decision? Decision Confidence How confident are you that your group made the right decision? How evenly do you think contributions to the discussions were disEquality of Contribution tributed among group members? Dominance – P1 How dominant did Participant 1 seem during the task? Dominance – P2 How dominant did Participant 2 seem during the task? Dominance – P3 How dominant did Participant 3 seem during the task? Dominance – P4 How dominant did Participant 4 seem during the task? Task Check (T1) Information that was unique to individual group members influenced the final decision. Table 8: Task 2 (Mini-negotiation) post-task questionnaire items. Label Question text Overall Valence Overall, how pleasant or unpleasant did this task feel? Perceived Influence How much influence did you feel you had during the task? Engagement How engaged/absorbed did you feel during the task? Mental Demand How difficult/mentally demanding did the task feel? Cooperative vs Competitive To what extent did the negotiation feel cooperative rather than competitive? Voice / Inclusion To what extent did you feel that your views were heard and taken into account? Satisfaction How satisfied are you with the outcome? Trust – Next to You How much did you trust the participant next to you? Trust – In Front How much did you trust the participant in front of you? Trust – At Angle How much did you trust the participant sitting at the angle from you? Dominance – P1 How dominant did Participant 1 seem during the task? Dominance – P2 How dominant did Participant 2 seem during the task? Dominance – P3 How dominant did Participant 3 seem during the task? Dominance – P4 How dominant did Participant 4 seem during the task? The outcome of the negotiation felt mutually beneficial for all parties Task Check (T2) involved. Implementation consequence. Because T0 normalisation universally increases variance, audio for B4–B5 (personality, between-person) is retained in absolute form, preserving between-subject amplitude variation. Separately, T1–T4 within-person normalisation is retained for B0–B3 because those four structured tasks share similar role clarity and behavioral demands. This reflects principled design: physiological features (cardiac, electrodermal) remain trait-stable across tasks and benefit from T0 baseline correction, while interaction-structure metrics require comparable task context.
F
Synchronisation Pipeline Detail
Task windows are derived from experiment-control markers written to the session event spine. Per-modality feature tables are then joined to these windows using timestamp filters and participant/seat mappings. The release preserves enough timing provenance to audit alignment at multiple levels: event-marker windows, per-device frame logs, progress streams, and the derived sync_metadata.json sidecars. In the current pipeline, frame-log start-spread analyses target sub-10 ms AV alignment, while DPA audio requires explicit correction for a measured hardware-clock drift of about 0.04 ms/s relative to
20
Table 9: Task 3 (Idea generation and selection) post-task questionnaire items. Label Question text Overall Valence Overall, how pleasant or unpleasant did this task feel? Engagement How engaged/absorbed did you feel during the task? Mental Demand How difficult/mentally demanding did the task feel? Confidence How confident did you feel during the task? Team Coordination How well did your group coordinate during the task? Voice / Inclusion To what extent did you feel that your views were heard and taken into account? Satisfaction How satisfied are you with the outcome? Fairness How fair did the outcome feel to you? To what extent did you feel free to propose ideas without worrying about Psychological Safety negative reactions? Idea Quality To what extent do you think your ideas were at the level of the others? Idea Diversity To what extent did the group generate many different and original ideas? Dominance – P1 How dominant did Participant 1 seem during the task? Dominance – P2 How dominant did Participant 2 seem during the task? Dominance – P3 How dominant did Participant 3 seem during the task? Dominance – P4 How dominant did Participant 4 seem during the task? Task Check (T3) The group generated many distinct and different ideas during this task. Table 10: Task 4 (Public-goods game) post-task questionnaire items. Label Question text Overall Valence Overall, how pleasant or unpleasant did this task feel? Social Evaluation Concern To what extent were you concerned that the other participants would evaluate or judge your choice? Fairness How fair did the outcome feel to you? Trust – Next to You How much did you trust the participant next to you? Trust – In Front How much did you trust the participant in front of you? Trust – At Angle How much did you trust the participant sitting at the angle from you? Expectation Match To what extent did the outcome match your expectations about how much people would contribute? Regret To what extent do you regret the contribution you chose to make? Dominance – P1 How dominant did Participant 1 seem during the task? Dominance – P2 How dominant did Participant 2 seem during the task? Dominance – P3 How dominant did Participant 3 seem during the task? Dominance – P4 How dominant did Participant 4 seem during the task? When making my choice, I was concerned about how the other particiTask Check (T4) pants would evaluate or judge me. the XDF/LSL clock. The audio splitter therefore fits a per-microphone linear time map instead of relying on a single median offset, yielding practical task-level residual error below 5 ms in documented sessions.
G
Preprocessing Steps
The five-step pipeline applied for all benchmarks (Section 6) is described below; Table 12 summarises the quantitative impact of each step across the 136 active participant-task rows.
Pipeline Steps ET quality gating. Eye-tracking rows where >50% of samples are missing have pupil features set to NaN; rows where the gaze-valid fraction is <50% have gaze features set to NaN.
Physiological plausibility gating. Values outside credible windows are replaced with NaN (heart rate (HR) ∈[40,180] / bpm; RMSSD ∈[10,300] / ms; EDA ∈[0,25] / µS; pupil ∈[1.5,9.0] / mm; pitch ∈[5,55] / semitones; delta features have tighter windows). In total 67 values are gated across all modalities.
Winsorisation. After plausibility gating, values are clipped at ±3 σ per feature. Only 32 values are clipped, confirming plausibility bounds removed the most extreme outliers.
Within-person robust z-score. For state and task benchmarks, features are centred by the participant’s median and scaled by 1.4826 × MAD across T1–T4 rows. This is the most impactful step: pupil dilation Cohen’s d for T2 vs. T1 rises from 0.28 to 1.13; speaking-fraction d doubles from 0.44 to 0.88. This transform is computed before leave-one-group-out splitting, so held-out participants contribute their own unsupervised T1–T4 distribution statistics to test-time normalisation.
21
Table 11: Audio feature variance comparison: absolute form vs. T0-normalised form across all 40 participants and 200 participant-task rows. All features show increased standard deviation when normalised by T0 baseline, confirming that T0 free-talk is insufficiently controlled to serve as a reliable per-participant reference. Consequently, audio is kept in absolute form for between-person benchmarks (B4–B5). Audio Feature
Abs. SD
Delta-T0 SD
Ratio
Interpretation
speaking_fraction
0.036
0.354
9.8×
pause_count
24.821
214.685
8.6×
overlap_fraction
0.030
0.078
2.6×
speech_rate_proxy
0.680
1.106
1.6×
T0 baseline highly variable; signals person was quiet then talkative, not stable trait Pausing naturally dependent on uncontrolled T0 social dynamics Group turn-taking in unstructured T0 creates confounding Vocal tempo variance inflated; T0 normalisation adds noise
Table 12: Step-by-step preprocessing audit across 136 active participant-task rows and 102 candidate feature columns (13,872 feature×row cells). Values gated are replaced with NaN; values clipped are winsorised in-place; values transformed are rescaled (within-person z-score); the KNN row is a descriptive audit, while benchmark imputation is fit inside each LOGO-CV fold. After global feature selection (missing >50% and |r| > 0.95 greedy removal) 35 features survive for benchmark evaluation (5 biomarker composites excluded; see text). Step
New NaNs
Clipped
Transformed
Imputed
16 67 0 0 0
0 0 32 0 0
0 0 0 9,970 0
0 0 0 0 772
†
ET quality gating Plausibility gating‡ Winsorisation (±3σ) Within-person z-score KNN imputation audit (k = 5)
† Pupil features nulled
when >50% samples missing per row; gaze features nulled when valid fraction <50%. ‡ Bounds: HR ∈[40,180] bpm; HRV RMSSD ∈[10,300] ms; EDA ∈[0,25] µS; skin temp ∈[28,40] °C; pupil ∈[1.5,9.0] mm;
pitch ∈[5,55] semitones; ∆HR ∈[−40,+40] bpm; ∆pupil ∈[−3,+3] mm. The benchmark-ready table retains residual NaNs for fold-local KNN imputation. We retain it because the paper’s benchmark goal is dataset characterisation rather than online deployment, but it should be read as a mild leakage source for B0–B3d. For between-person targets (B4a–B5), within-person z-scoring is not applied.
KNN imputation. A k = 5 nearest-neighbour imputer is fit inside each LOGO-CV training fold and applied to the held-out group, preventing missing-data leakage.
Feature Selection Feature selection was applied globally to the full 136-row active-task dataset (T1–T4) before any LOGO-CV split using two criteria applied sequentially: (1) features with more than 50% missing values were dropped; (2) among remaining features, highly correlated pairs (|r| > 0.95) were reduced by greedy removal, retaining the first feature in each correlated cluster. After the full pipeline, 35 features survive selection. Five composite biomarker features (biomarker_*) are excluded from the benchmarks: they were computed as internal linear combinations of physiological and pupil features already in the retained set (e.g. biomarker_cognitive_load = mean(z_pupil, z_eda, −z_rmssd)), and their presence alongside raw source components makes benchmark attribution ambiguous.
Retained Feature Set Table 13 lists all 35 retained features grouped by modality. The four annotation process-metadata features (bottom group) are listed for completeness but are excluded from the 31-feature benchmark set; see Leakage Sources below.
Leakage Sources Three deviations from a fully fold-internal pipeline are acknowledged: 1. Within-person z-score (affects B0–B3d). Normalisation statistics (median, median absolute deviation (MAD)) are computed over all four task rows per participant before the LOGO-CV split, so held-out participants contribute their own unsupervised T1–T4 distribution to test-time normalisation. No directional label bias, but within-person effect sizes are inflated. Not applied for between-person targets (B4a–B5). 2. Global feature selection. Missing-rate and correlation-based filtering is applied to the full dataset before splitting. Cannot favour any particular label direction; impact on AUC is negligible, but represents a strict-protocol deviation.
22
Table 13: Retained feature set after global feature selection (35 features total; 31 used in all reported benchmarks). “∆T0” = value relative to the T0 free-talk baseline. The annotation group is excluded from benchmark runs. Feature
Description
Physiology (8 features) hr_mean_bpm_delta_t0 hr_sd_bpm hrv_rmssd_ms hrv_quality_score eda_tonic_mean_delta_t0 eda_phasic_rate_hz_delta_t0 eda_phasic_mean_delta_t0 eda_scr_count
Mean HR relative to T0 HR variability (SD) HRV RMSSD from PPG PPG signal quality index Tonic EDA ∆T0 Phasic EDA event rate ∆T0 Phasic EDA amplitude ∆T0 skin conductance response (SCR) count
Motion / Temperature (3 features) temp_mean_delta_t0 accel_motion_mean motion_high_fraction
Skin temperature ∆T0 Mean wrist-acceleration magnitude Fraction of high-motion samples
Eye-tracking (5 features) pupil_left_mean pupil_right_mean pupil_std pupil_slope_per_s pupil_mean_delta_t0
Left pupil diameter (mm) Right pupil diameter (mm) Pupil diameter SD Linear pupil dilation slope Mean pupil diameter ∆T0
Audio (15 features) audio_energy_mean_x audio_energy_sd_x audio_hnr_mean_x audio_jitter_mean_x audio_mean_unvoiced_segment_s audio_mean_voiced_segment_s_x audio_overlap_fraction_x audio_pause_count audio_pitch_mean_x audio_pitch_sd_x audio_shimmer_mean_x audio_speaking_fraction_x audio_speaking_time_s audio_speech_rate_proxy_x audio_voiced_segments_per_sec_x
Mean frame energy Frame energy variability Harmonics-to-noise ratio Pitch period jitter Mean unvoiced segment duration Mean voiced segment duration Overlapping-speech fraction Pause event count Mean fundamental frequency Pitch variability Amplitude shimmer Speaking-time fraction Total speaking time (s) Syllable-rate proxy Voiced-segment density
Annotation process-metadata (4 features, excluded from benchmarks) answers_n Responses submitted ann_total_events_n Total annotation events ann_response_postblock_n Post-block form submissions ann_event_span_s Annotation event time span (s) 3. Annotation process-metadata features. The four annotation features (answers_n, ann_total_events_n, ann_response_postblock_n, ann_event_span_s) are excluded from the 31-feature benchmark set. ann_event_span_s and answers_n vary systematically by task (T2 has more VAD probes and a longer annotation span), creating a direct shortcut for the B0 task-classification sanity check. ann_response_postblock_n is a data-completeness proxy structurally correlated with annotation-derived B3/B_trust targets. B0 accuracy rises to 0.734 when these features are included (vs. reported 0.641), confirming the exclusion is conservative. All feature-importance figures (Section K) and benchmark scripts exclude these four features.
Technical Caveats PPG-derived HRV precision. hrv_rmssd_ms is derived from PPG peak detection at ≈25 Hz. At this sampling rate, the minimum detectable R-R interval (RR) interval difference is ≈40 ms, setting a quantisation floor on RMSSD. Spearman correlations between hrv_rmssd_ms and post-block self-report labels are weak (r = −0.09 with engagement, r = 0.21 with arousal, both p > 0.05; n = 53–67), consistent with the limited precision of wrist PPG-derived HRV at this resolution. The feature is retained because it provides the only autonomic vagal-tone proxy available in the release; users requiring ECG-grade HRV should treat hrv_rmssd_ms as an approximate indicator rather than a precise measurement.
23
Table 14: Descriptive task-level autonomic findings used in the worked example. Values are baselinenormalised against T0 within participant. Effect size is within-participant Cohen’s dz against zero; rows are descriptive and are not treated as confirmatory tests. Task
Feature
n
Mean change
dz
T1 T2 T3 T4 T1 T2 T3 T4 T1 T2 T3 T4
Pupil diameter (mm) Pupil diameter (mm) Pupil diameter (mm) Pupil diameter (mm) Skin temperature (deg C) Skin temperature (deg C) Skin temperature (deg C) Skin temperature (deg C) SCR rate (peaks/s) SCR rate (peaks/s) SCR rate (peaks/s) SCR rate (peaks/s)
40 40 40 40 34 34 33 33 33 33 32 32
-0.117 -0.038 -0.108 -0.076 0.610 0.613 0.468 0.459 0.027 0.037 0.036 0.045
-1.01 -0.31 -0.70 -0.73 0.96 0.80 0.58 0.60 0.51 0.42 0.39 0.45
EDA motion contamination. The mean motion_high_fraction across all participant-task rows is 0.0015 (SD 0.0018), indicating that elevated wrist-motion periods represent a negligible fraction of recorded time in this cohort. The feature was dropped during global feature selection due to near-zero variance, confirming that motion-elevated EDA periods are rare enough to have minimal influence on the benchmark features.
H
Extended Dataset Characterization
Personality and multimodal response Spearman correlations between the five BFI traits and four participant-level behavioural means (valence, arousal, engagement, mental demand) yield 20 tested pairs (n = 40). After Benjamini–Hochberg false discovery rate (BH-FDR) correction, no association survives at q < 0.05. The nominally strongest associations before correction are Agreeableness with arousal (r = 0.37, p = 0.017) and Agreeableness with engagement (r = 0.35, p = 0.028). The previously reported Openness– speech-rate correlation (r = −0.40, n = 31) remains present but does not survive the full correction. These participant-level correlations treat individuals as independent observations; a linear mixed-effects model with random intercepts per participant and task would account for the nested structure and yield calibrated confidence intervals the recommended approach for future work with a larger cohort. The complete correlation table is in the release datasheet; BFI results are released primarily as covariates for future work given that n = 40 is severely underpowered for personality–behaviour associations. Per-group BFI-44 trait profiles are available in the release metadata. Language proficiency BFI-44 was administered in English. English proficiency among the 40 participants is: native speaker (10, 25%), fluent non-native (24, 60%), intermediate (6, 15%). BFI-44 has not been validated specifically for intermediate second-language respondents, and item interpretation may vary with proficiency. The six intermediate-proficiency participants represent a small subgroup; their BFI responses are included in the release but should be interpreted with this caveat. English proficiency (english_proficiency) is released as a covariate to support sensitivity analyses; at n = 40 this dataset is underpowered to detect proficiency-moderated effects reliably.
Familiarity Participants rated their prior acquaintance with each other on a 7-point Likert scale (1 = “never met”, 7 = “very familiar”) before the session began. The distribution was heavily skewed toward low familiarity: 70.0% of responses fell below 4, and 62.4% were exactly 1 (“never met”). Overall mean familiarity was M = 2.79 (SD 2.56), indicating that participants were predominantly strangers, which is consistent with the recruitment strategy of assigning groups from volunteers across different departments within the organization. This low baseline familiarity setting reduces confounds from pre-existing team dynamics.
Task outcomes and interaction dynamics The four structured tasks produce quantifiable group outcomes that can be linked to multimodal signals. fig. 11 summarizes the per-group results across all decision tasks. • T1: All 10 groups reached consensus on Candidate C, the correct choice. One group (grp-07) submitted multiple response entries for the same participant, reflecting participants’ misunderstanding of instructions. • T2: All 10 groups (n = 10/10) reached settlement on both topic and format. Discussion timing: Despite the 8-minute guideline, all 10 groups overran, with mean discussion duration 11.4 minutes (SD 1.6, range 9.2–14.1 min). Although specific timings for each phase in each task were fixed, the moderator was instructed not to interrupt the organic conversation flow and to prompt rapid closure as soon as possible after timer expiration. As a result, groups often overran the nominal phase timings (see fig. 12). • T3 (Idea Generation via Nominal Group Technique): All 10 groups generated ideas and voted on a winner. Each group’s final selection is logged as idea_author (format: “PX — Idea N”), which links back to the participant’s submitted idea_N text fields. • T4 (Public-Goods Micro-Game): Individual contributions (tokens contributed out of 10-token endowment) yield M = 7.05 (SD 1.76, range 4.50–10.00 per group mean). High overall contribution level (70.5% of endowment) indicates strong
24
cooperative tendency. Per-group means vary from 4.50 (grp-16, conservative) to 10.00 (grp-13, fully cooperative), providing continuous behavioural targets for cooperation-prediction benchmarks.
Figure 11: Task Outcomes Dashboard (T1–T4). Panel 1: Hidden-profile decision outcome (T1). All 10 groups selected Candidate C. Panel 2: Mini-negotiation outcome (T2). Topics and formats are colour-differentiated. Most common: “AI for productivity” (5 groups). Panel 3: Idea Generation outcome (T3). Winning ideas are grouped by themes. Panel 4: Public-Goods Contribution (T4). Per-group mean contributions (0–10 scale) sorted by value. Colour intensity indicates cooperation level (red < 5, orange 5–7.5, green ≥ 7.5). Overall mean 7.05 (SD 1.76) reflects strong cooperative tendency.
Self-report dynamics In-task self-reports of valence and arousal, measured on a 9-point range, reveal distinct affective signatures across the four tasks. Valence shows a clear task effect: Mean ratings were highest in T0 (baseline/introductions; M = 7.59), T3 (idea generation; M = 7.36), and T4 (micro-game; M = 7.22), but dropped significantly in T1 (hidden-profile decision; M = 6.87) and most notably in T2 (negotiation; M = 5.76). This pattern aligns with task demands: T2 requires adversarial negotiation and format selection, which introduces conflict and uncertainty, thereby dampening positive affect. In contrast, arousal shows an opposite pattern: Baseline arousal in T0 was lowest (M = 5.22), with all task phases elevating arousal similarly (T1: M = 6.12, T2: M = 6.35, T3: M = 6.21, T4: M = 6.13). This indicates that structured group tasks themselves activate heightened engagement, irrespective of affective valence. The asymmetry between valence and arousal suggests that conflict (T2) dampens pleasure while maintaining engagement.
Post-task survey profiles Table 15 maps each post-block survey construct to the tasks in which it was administered. Core items (engagement, mental demand, overall valence, per-seat dominance rating) appear in all four active tasks and support full cross-task comparisons. Trust items appear only in T2 and T4; satisfaction only in T2 and T3; voice and inclusion in T1–T3 but not T4. Researchers planning cross-task analyses of a specific construct should verify availability in this table first. Mental demand differed significantly across tasks. T2 (negotiation) elicited the highest mental load (M = 4.79, SD 1.42), compared to T1 (M = 3.20, SD 1.68) and T3 (M = 3.65, SD 1.84). This aligns with T2’s dual demand: participants must simultaneously negotiate content (topic and format) and social dynamics (consensus-seeking under time pressure). Satisfaction ratings reveal recovery after difficult negotiation: T2 satisfaction (M = 5.09, SD 1.57) was significantly lower than T3 (M = 6.09, SD 1.02), suggesting that the post-negotiation idea-generation task provided a more positive experience, possibly due to reduced interpersonal conflict and the creative freedom of idea generation.
25
Figure 12: Time-to-decision durations (T1–T3). Boxplots show discussion durations from onset to the moderator “finish” prompt for T1–T3. The dashed horizontal line marks the nominal 8minute guideline; groups typically overran this duration, especially in T2. T4 is excluded because its discussion phase did not terminate in a single group decision. Durations were derived directly from events_grp-XX.tsv files (one per group) by filtering moderator push_content events for the relevant task and computing the difference between the role_card start-phase marker and the corresponding finish marker. Onsets are machine-recorded LSL timestamps (seconds relative to session start), so these durations reflect exact event-log timing rather than manual annotation. Voice inclusion measures whether the participant felt like their views were being heard in the dicussion. It rated consistently high in T1 (M = 5.80, SD 1.23) and T3 (M = 5.84, SD 1.22), but declined in T2 (M = 4.81, SD 1.64), the negotiation task. This suggests that asymmetric participation or perceived dominance emerges during negotiation, consistent with the higher mental demand and lower satisfaction ratings for that task. Trust was captured as seat-directed interpersonal ratings: each participant rated the three other group members (front, next, angle) on trust-related items (7-point scale). For example, a participant in position P1 rates trust toward P4 (front), P2 (next), and P3 (angle). We first computed each participant’s trust score as the mean of their three seat-directed ratings, then compared T2 and T4. At the participant level (n = 40 paired observations), trust averaged M = 4.98 (SD 1.09) after T2 and M = 5.25 (SD 1.24) after T4, with mean change ∆ = +0.27 (SD 1.64). The paired test was not significant, t(39) = 1.03, p = .31. At the group level (n = 10; mean across participants within group), all groups were above midpoint in both tasks (10/10 in T2; 10/10 in T4), and 5/10 groups showed a net increase from T2 to T4. The paired group-level change was also non-significant, t(9) = 1.51, p = .165. The directional increase is consistent with a cooperative rebound from negotiation (T2) to collective-action play (T4), but inferential results are not reliable at current sample size and should be interpreted as exploratory.
Conversation dynamics and turn-taking The per-seat close-talk layout enables voice-activity-derived conversation-structure summaries. Table 16 reports transcriptderived turn-taking statistics by task. T0 produces the most turns (n̄ = 54.8) but the shortest mean duration (2.19 s). T1 and T2 show the longest individual turns (4.41 and 4.19 s) and the lowest overlap fractions (0.125 and 0.107), consistent with deliberative or competitive conversational floors. T3 stands out for long response gaps (5.19 s) consistent with the nominal group technique. These patterns motivate benchmarks B9–B11 (Section I).
I
Extended Benchmarks: Sequential Conversation Tasks
The per-seat close-talk recordings support three additional benchmarks requiring per-turn event extraction from voice-activitydetection segmentation. B9: Next-speaker prediction. Given a context window ending at a turn boundary, predict which of the remaining three participants takes the floor next (four-class; majority baseline ≈ 0.33). Inputs would include last-N -second physiology windows, gaze fixation vectors, and prosodic tail features. B10: Turn-taking point detection. Given a rolling 2-second window of per-seat voice-activity-detection output, predict whether a turn transition will occur within the next 500 ms (binary; AUC + F1).
26
Table 15: Post-block survey construct availability by task. A tick (✓) indicates the item appeared in the post-block questionnaire for that task; a dash indicates it was absent from that task’s survey module. T0 had no post-block survey. Real-time VAD probes (valence, arousal, dominance) were administered during all five tasks; dominance was absent from the T4 VAD probe schema. Construct
T1
T2
T3
T4
Shared core items Engagement Mental demand Overall valence Per-seat dominance rating
✓ ✓ ✓ ✓
✓ ✓ ✓ ✓
✓ ✓ ✓ ✓
✓ ✓ ✓ ✓
Partially shared items Voice & inclusion Team coordination Satisfaction Trust (3 items) Fairness Confidence Perceived control
✓ ✓ — — — — ✓
✓ — ✓ ✓ — — ✓
✓ ✓ ✓ — ✓ ✓ —
— — — ✓ ✓ ✓ —
Task-specific items Decision confidence (T1) Info sharing (T1) Equality of contribution (T1) Familiarity per seat (T1) Cooperativeness (T2) Idea quality / diversity (T3) Psychological safety (T3) Contribution tokens (T4) Regret / social concern (T4) Expectation match (T4)
✓ ✓ ✓ ✓ — — — — — —
— — — — ✓ — — — — —
— — — — — ✓ ✓ — — —
— — — — — — — ✓ ✓ ✓
Table 16: Transcript-derived turn-taking and overlap summary by task (participants only). Task
Turns
Turn dur. (s)
Overlap frac.
Resp. gap (s)
Backchannel rate
T0 T1 T2 T3 T4
54.8 21.0 30.2 27.4 16.3
2.19 4.41 4.19 3.16 2.44
0.130 0.125 0.107 0.191 0.200
3.63 3.60 3.04 5.19 3.21
0.404 0.305 0.209 0.271 0.313
B11: Overlap onset prediction. Given the current speaker’s prosodic trajectory, predict whether the next transition involves simultaneous speech (binary; positive fraction ≈ 0.15 based on Table 16). All raw data for B9–B11 is in place; per-turn voice-activity-detection event extraction is deferred to a future release.
J
Benchmark Interpretation Notes
Per-benchmark commentary for the results in Table 3; modality ablation (Figure 15) and feature importance (Figure 16) are also presented here.
B0: Task-label classification (sanity check). Accuracy 0.641 (SD 0.132, 95% CI [0.55, 0.73]) vs. 0.265 baseline (n = 136). The previously published figure of 0.734 was inflated by annotation process-metadata features (ann_event_span_s, answers_n) that vary systematically by task; see Section G. B0 confirms the tasks are behaviourally distinguishable under LOGO-CV, not that audio is especially strong for affect inference.
B1a/B1b: Valence and arousal. Valence AUC 0.657; arousal AUC 0.528 (n = 107). Near-chance arousal reflects a post-block temporal mismatch: probes capture retrospective state rather than the peak physiological response. An event-contingent approach at decision moments would better align labels with signals.
B2: Dominance. AUC 0.499 (n = 83, 95% CI [0.37, 0.62]) near chance. B3a/B3b: Mental demand and engagement (strongest Level-1 result). B3a AUC 0.719 (n = 99); B3b AUC 0.591 (n = 99). A single audio_overlap_fraction_x feature achieves AUC 0.766 for mental demand cognitive-load detection is driven by conversational-floor dynamics, not physiology alone. B3a and B3b load on different feature profiles (Figure 16): B3a is dominated by audio overlap and HR; B3b by pupil slope and pitch. This within-dataset dissociation confirms demand and engagement are not proxies for one construct.
27
Figure 13: Task-level physiological feature effect sizes (Cohen’s d) across modalities. T2 shows sustained EDA elevation; T3 shows high HR and SCR; T4 shows moderate increases. B3c: Satisfaction (new, †). AUC 0.571 (n = 60, T2 and T3 only). T3 satisfaction is significantly higher than T2 (t = 3.33, p = 0.001), but the demand–satisfaction dissociation weakens B3c relative to B3a with annotation processmetadata features excluded. B3d: Trust (new, †). AUC 0.562 pooled (n = 60); AUC 0.679 T4-only (n = 28). Near-zero within-group variance in T2 trust (SD = 0.185) makes the T2 fold-local split near-random; the cooperative T4 context recovers the signal. B4a/B4b/B4c: Personality traits (challenge). All three are near or below chance under LOGO-CV; test folds of four participants make AUC estimates inherently unstable regardless of true signal. Spearman correlations on the full sample confirm signal presence: pupil_right_mean × Agreeableness r = +0.51 (p = 0.008). A two-feature model restricted to T2 achieves AUC 0.625 for Agreeableness. The B4 trio defines a well-motivated but presently unsolvable challenge. B5: T4 contribution. AUC 0.429 (n = 28, SD 0.290); highly imbalanced fold structure. The continuous raw contribution (mean 7.05/10, SD 2.88) is the preferred future modelling target over binary split. B6a/B6b: Speaking Gini. B6a MAE 0.089 vs. baseline 0.086 (group-mean features). B6b Ridge MAE 0.102 vs. baseline 0.088 (SD features) both below naive. A binary classifier on raw speaking-fraction SD alone achieves AUC 0.952 (95% CI [0.857, 1.000]), confirming the signal is definitively present; the Ridge failure is overfitting at n = 28 rows. B7: Speech-overlap fraction. MAE 0.063 vs. baseline 0.060; same mean/variance mismatch as B6a. Modality contributions. Figure 15 reports LOGO-CV performance under ten feature-subset conditions. Audio dominates task-structure and cognitive-demand detection. Pupil adds independent signal for cognitive state beyond audio alone. Physiology contributes primarily to arousal-sensitive targets (B1a, B2). Feature-level importance. Figure 16 shows top-15 features by mean normalised |coefficient| across LOGO-CV folds. audio_overlap_fraction_x ranks #1 in five benchmarks. B3a and B3b load on entirely different profiles (audio vs. pupil+pitch), confirming the demand–engagement dissociation. B3d trust is led by speaking time and shimmer rather than floor competition.
K
Full Ablation Table
L
Per-Session Quality Table
Table 18 reports per-session modality coverage for active tasks T1–T4. grp-09 had three EmotiBit sensors with partial data loss; grp-11 and grp-12 had PPG and EDA channels affected by device faults. All sessions retained full eye-tracking row availability; pupil usability was reduced in grp-07 and grp-11 due to calibration issues.
M
Responsible AI and Croissant Metadata
NeurIPS 2026 E&D requires Croissant metadata with Responsible AI fields [33, 32]. Table 19 summarises the fields in the release Croissant file.
28
Figure 14: Cross-modal Spearman correlation matrix (physiological and audio features vs. self-report annotation targets) on 136 active participant-task rows with BH-FDR correction. audio_overlap_fraction (r = −0.62) and pupil dilation are the dominant predictors of mental demand, motivating benchmark B3a.
N
Datasheet for Datasets
This datasheet follows the Gebru et al. [13] template.
Motivation For what purpose was the dataset created? GroupAffect-4 was created to support research on multimodal group affect, social dynamics, and interaction behaviour in structured co-located tasks. Existing corpora focus on single-participant or dyadic settings; GroupAffect-4 fills a gap by providing synchronised physiology, eye tracking, audio, and self-report for four-person groups. Who created the dataset? Meisam Jamshidi Seikavandi1,2 (corresponding author), Alice Modica1,3 , Anna Obara1,4 , Shan Ahmed Shaffi1 , Fabricio Batista Narcizo1,2 , Tanya Ignatenko1 , Ted Vucurevich1 , Karim Haddad1 , Daniel Barratt3 , Daniel Overholt4 , Jesper Bünsow Boldt1 , Paolo Burelli2 , and Andrew Burke Dittberner1 . 1 GN Advanced Science, GN Group, Ballerup, Denmark; 2 IT University of Copenhagen, Copenhagen, Denmark; 3 Copenhagen Business School, Copenhagen, Denmark; 4 Aalborg University, Denmark. Correspondence: [email protected]. Who funded the creation? Internal research funding from GN Advanced Science, GN Group, Ballerup, Denmark.
Composition What do the instances represent? Each instance is a participant-session-task recording unit. The dataset contains 40 participants, 10 groups of 4, across 5 tasks (T0 + T1–T4): up to 200 participant-task rows per modality. How many instances? 200 expected participant-task recording units. Coverage varies by modality (see Table 5 and Section L). Does it contain all possible instances? This is the full paper-facing subset. 16 sessions were recorded in total: the first 5 were pilot sessions used to finalise the protocol; of the 11 non-pilot sessions, 10 are released and 1 was excluded due to incomplete modality coverage. What data does each instance consist of? (1) EmotiBit physiological signals (PPG, EDA, skin temperature, IMU, ≈25 Hz); (2) Tobii Pro Glasses 3 egocentric eye-tracking (≈50 Hz); (3) close-talk lapel microphone audio (48 kHz WAV); (4) transcript artifacts and turn-taking summaries derived from the audio; (5) tablet self-report probes (valence, arousal); (6) task outcome records. Are there labels or targets? Self-reported valence and arousal (9-point SAM), task performance
29
Modality ablation: relative performance (0 = chance, 1 = best in row) 0.265
0.235
0.493
0.610
â
0.301
0.515
0.500
0.640
0.581
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
0.518
0.482
0.470
0.578
0.470
0.482
0.518
0.422
0.590
0.566
0.525
0.525
0.616
0.556
0.455
0.545
0.646
0.626
0.677
0.657
0.556
0.465
0.677
0.707
0.515
0.556
0.616
0.646
0.687
0.626
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
â
0.083
0.094
0.097
â
0.099
0.107
â
â
0.129
0.154
0.070
0.077
0.057
0.078
0.068
0.070
â
â
0.097
0.117
l diac GSR Pupi Audio Car EDA/ Single modality
best
mid
Rel. performance
B0: Task (clf) B1: Valence (clf) B1: Arousal (clf) B2: Dominance (clf) B3: Engagement (clf) B3: Mental demand (clf) B4: Extraversion (clf) B4: Openness (clf) B5: Contribution (clf) B6: Speaking Gini B7: Speech overlap
chance
p p FI or BFI +EDA +PuCombinations +Pu Sens All+B d d EDA Car Car
Figure 15: Modality ablation heatmap. Colour encodes performance relative to chance (white = chance, dark teal = best). Audio dominates task classification; audio and pupil are jointly strongest for cognitive-state targets. See Table 17 for numerical details. Feature importance per benchmark (mean |coefficient|, LOGO-CV folds) B0 Task label Overlap fraction Audio energy mean Pupil mean baseline Pupil slope Audio energy SD SCR count Pause count Speaking time (s) Pitch mean Pupil right mean Pupil left mean Skin temp baseline Shimmer mean EDA phasic mean baseline Voiced seg. mean (s)
B1a Valence
Acc 0.641 0.0
0.2
0.4 0.6 Norm. |coef|
0.8
1.0
AUC 0.668 0.0
0.2
B3a Mental demand Overlap fraction HR mean baseline Pause count Pupil slope Voiced segs/s HR SD HRV quality Pupil right mean HRV RMSSD EDA phasic mean baseline Pitch SD High-motion fraction Audio energy mean Speaking time (s) HNR mean 0.2
0.4 0.6 Norm. |coef|
0.8
0.4 0.6 Norm. |coef|
0.8
1.0
Pupil slope Pitch mean SCR count HRV RMSSD Skin temp baseline Audio energy SD Pupil left mean HR SD EDA tonic baseline Pupil mean baseline Speaking time (s) Unvoiced seg. mean (s) High-motion fraction Voiced segs/s Speech rate proxy
1.0
AUC 0.533 0.0
0.2
0.2
0.4 0.6 Norm. |coef|
Physiology
0.8
0.4 0.6 Norm. |coef|
0.8
Overlap fraction Pupil left mean Audio energy SD Shimmer mean Unvoiced seg. mean (s) Pupil mean baseline Speech rate proxy Speaking fraction Audio energy mean EDA phasic rate baseline HR mean baseline Pitch mean Voiced seg. mean (s) SCR count Pitch SD
1.0
AUC 0.499 0.0
0.2
B_sat Satisfaction
AUC 0.583 0.0
B2 Dominance
Audio energy mean Voiced segs/s Speaking fraction Speaking time (s) High-motion fraction HRV RMSSD HR mean baseline Pupil slope Pitch mean HR SD Unvoiced seg. mean (s) Pitch SD Speech rate proxy EDA phasic mean baseline Pupil right mean
B3b Engagement
AUC 0.739 0.0
B1b Arousal
Overlap fraction Skin temp baseline EDA phasic mean baseline Speaking time (s) Pause count Pupil mean baseline Unvoiced seg. mean (s) Accel motion mean SCR count HRV RMSSD HR SD Audio energy mean Pupil slope EDA tonic baseline Speech rate proxy
Overlap fraction Accel motion mean Pupil mean baseline Pitch mean Speaking time (s) Pupil right mean HR SD Jitter mean Speaking fraction HRV RMSSD Skin temp baseline Pitch SD Voiced seg. mean (s) EDA phasic rate baseline HRV quality
1.0
Eye-tracking
Audio
0.2
0.4 0.6 Norm. |coef|
0.8
0.8
1.0
B_trust Trust
AUC 0.571 0.0
0.4 0.6 Norm. |coef|
1.0
Speaking time (s) Shimmer mean Pause count Pupil left mean HR SD Audio energy mean Overlap fraction Jitter mean HR mean baseline Voiced segs/s HRV quality Accel motion mean SCR count HRV RMSSD Speech rate proxy
AUC 0.562 0.0
0.2
0.4 0.6 Norm. |coef|
0.8
1.0
Motion/Temp
Figure 16: Per-benchmark ranked feature importance: top-15 features by mean normalised |coefficient| across LOGO-CV folds (31-feature set). Bar colour: Physiology, Eye-tracking, Audio, Motion/Temp. ⋆ = audio_overlap_fraction_x.
outcomes where derivable, and BFI-44 personality traits (participant-level). No external affect annotations. Is information missing? Yes. EmotiBit availability missing for 9.0% of expected rows; PPG/EDA/temperature usability ≈24% incomplete; Tobii row availability 98.0%; valence/arousal self-report 14.5% incomplete. All missingness is documented in Section L. Are relationships between instances explicit? Yes; participants within a session share a group_id, and seat identifiers P1–P4 are consistent across tasks. Are there errors, noise, or redundancies? Known artefacts: HRV RMSSD quantisation floor at ≈40 ms; audio lapel misclassification of ambient noise in T0; eye-tracker calibration drift in some sessions. Is the dataset self-contained? Self-contained for tabular modalities. The public release includes transcript artifacts; room video, 3D pose, and room-frame gaze are not in v1. Does it contain confidential data? Raw audio carries re-identification risk and is released under a separate access tier. Does it relate to people? Yes; all data from human participants with written informed consent. Does it identify subpopulations? Self-reported age, sex, handedness, English proficiency, and education are released to enable demographic reporting, not for subgroup targeting. Racial or ethnic composition was not collected as part of the study protocol; the sample is a university/community convenience sample and should not be treated as representative of any demographic population. Is identification possible? Voice re-identification is a realistic risk; the DUA prohibits attempts. Does it contain sensitive data? Physiological signals combined with personality scores could support inferential attacks; users are prohibited from applying data for clinical diagnosis or individual profiling.
30
Table 17: Feature-modality ablation: LOGO-CV primary metric under ten feature-subset conditions. Cardiac: HR mean and HRV RMSSD (delta + absolute). EDA/GSR: EDA tonic/phasic and skin temperature. Pupil: pupil dilation delta and absolute. Audio: speaking fraction, overlap, energy, pitch, prosody. BFI: Big Five personality traits. Sensor: all four sensor modalities combined (excludes annotation features). Bold = best condition per row. B6 (Speaking Gini) and B7 (Speech-overlap fraction): group-task level; physio/pupil are group means. ⋆ “B7 (legacy): Floor dominance” is a binary classification target from an earlier benchmark version, retained here for ablation completeness; it does not appear in the main evaluation suite. Note: This ablation uses accuracy for binary classification targets (B0, B2, B3a/b) and MAE for continuous/regression targets (B1a/b, B4a/b, B5), and was computed with the full 40-feature set (including biomarker composites subsequently excluded from the main benchmarks); results are therefore not directly comparable to the AUC figures in Table 3. Audio dominates task classification; pupil and audio jointly dominate cognitive-state detection; notably, a single audio feature (audio_overlap_fraction_x) matches or exceeds the full-sensor model for mental demand (see text). Benchmark
Metric Cardiac EDA/GSR Pupil Audio
BFI Card+EDA Card+Pup EDA+Pup Sensor All+BFI
B0: Task label (T1-T B1a: Valence B1b: Arousal B2: Dominance (high/ B3b: Engagement B3a: Mental demand B4a: BFI Extraversion B4b: BFI Openness B5: T4 Contribution B7 (legacy): Floor dominance⋆ B6a: Speaking Gini B7: Speech-overlap fraction
Acc. MAE MAE Acc. Acc. Acc. MAE MAE MAE Acc. MAE MAE
– 1.070 1.121 0.470 0.455 0.515 – – 0.496 0.554 0.099 0.068
0.265 1.462 1.147 0.518 0.525 0.556 0.508 0.960 0.715 0.589 0.083 0.070
0.235 1.030 1.105 0.482 0.525 0.465 0.473 0.476 0.506 0.652 0.094 0.077
0.493 1.094 1.077 0.470 0.616 0.677 0.608 0.464 0.507 0.554 0.097 0.057
0.610 1.118 1.159 0.578 0.556 0.707 0.644 0.652 0.767 – – 0.078
0.301 1.640 1.195 0.482 0.545 0.556 0.678 0.962 0.710 0.670 0.107 0.070
0.515 1.148 1.103 0.518 0.646 0.616 0.950 1.014 0.745 0.446 – –
0.500 1.087 1.102 0.422 0.626 0.646 0.511 0.544 0.522 0.580 – –
0.640 1.442 1.344 0.590 0.677 0.687 1.040 1.988 0.845 0.562 0.129 0.097
0.581 1.725 1.330 0.566 0.657 0.626 – – 0.708 0.598 0.154 0.117
Table 18: Per-session modality coverage across the 10 final groups (active tasks T1–T4; 16 participanttask rows expected per session per modality). ✓ = all 16/16 rows available/usable; fractions indicate partial coverage. Audio counts include T0–T4 WAV files across all microphone channels. Beh = behavioural/self-report event files. Group
Date
Physio avail.
PPG usable
EDA usable
ET avail.
Pupil usable
Audio (WAV)
grp-07 grp-08 grp-09 grp-10 grp-11 grp-12 grp-13 grp-14 grp-15 grp-16
2026-03-12 2026-03-13 2026-03-17 2026-03-17 2026-03-18 2026-03-18 2026-03-18 2026-03-19 2026-03-19 2026-03-20
✓ ✓ 8/16 ✓ ✓ 14/16 ✓ ✓ ✓ ✓
✓ ✓ 7/16 ✓ 4/16 0/16 ✓ ✓ ✓ ✓
✓ ✓ 6/16 ✓ 4/16 0/16 ✓ ✓ ✓ ✓
✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
12/16 ✓ ✓ ✓ 12/16 ✓ ✓ ✓ ✓ ✓
30 28 30 30 30 30 22 30 34 30
150/160 123/160 122/160 160/160 152/160 (93.8%) (76.9%) (76.2%) (100%) (95.0%)
294 –
Total (%)
Collection Process How was data acquired? Physiology via EmotiBit wrist sensors (LSL, ≈25 Hz); eye tracking via Tobii Pro Glasses 3 (≈50 Hz); audio via DPA 4060 lapel microphones (48 kHz WAV); self-report via tablet-based SAM probes; synchronisation via common LSL clock. What mechanisms were used? One or two experimenters per session; participants wore sensors throughout; LSL markers demarcate task windows. Sampling strategy? Convenience sampling via internal staff networks and affiliated university community postings at GN Advanced Science, GN Group, Ballerup, Denmark; not nationally representative. Timeframe? March 2026. Ethical review? Yes; participant information statement and informed consent procedure reviewed and approved by the GN Hearing A/S Legal Department; written informed consent obtained from all participants; participants could withdraw at any time.
Preprocessing, Cleaning, and Labelling Was preprocessing done? Yes: task-window extraction from LSL markers, per-modality QC passes, BFI-44 domain scoring with reverse coding, and anonymisation. Was raw data saved? Yes; the full raw acquisition archive is retained internally. The
31
Benchmarks that beat the naive baseline Classification / AUC (vs. baseline)
Improvement over naive baseline (%)
200 175
Classification (Acc.)
+165.3%
150 125 100 0.703
75
+64.2%
50
+37.1% +21.0%
25 0
0.627
ask B0 T
ance
B
min 2 Do
+11.4%
0.747
t
m.
men
B3
ge Enga
De ntal e M 3
B
+29.8%
0.821
0.570
l
a rous
B1 A
.
ntrib
o B5 C
Figure 17: Feasibility baselines for the 31-feature set (biomarker composites and annotation processmetadata features excluded; leakage-corrected). Each bar shows percentage improvement over the majority-class baseline or naive regressor; benchmarks at or below baseline are omitted. B3a mental demand is the clearest above-chance participant-state target; B1a valence is moderate; B2 dominance and B5 are near or below chance in the clean analysis. B4a–B4c (personality) are systematically below chance and are not shown. B6a–B7 (group-level) remain at or below naive MAE and are not shown. See Table 3 for full results with CIs. no-video release contains task-split BIDS-style modality files with documented provenance. Is the software available? Yes; processing scripts under tools/; deterministic pipeline.
Uses Has it been used for any tasks? Only by the authors for the analyses in this paper. What other tasks could it support? Multimodal affect modelling, group dynamics, personality–behaviour associations, cross-modal synchrony, speech prosody and interaction rhythm. See Section 7 and Section 6. Are there tasks it should not be used for? Clinical diagnosis, individual mental-health inference, workplace surveillance, automated hiring, speaker re-identification, or voice cloning.
Distribution How will it be distributed? The dataset is publicly available. Code and processing scripts: https://github.com/ meisamjam/GroupAffect-4. Tabular features and derived data: https://zenodo.org/records/20037847 (persistent DOI, Zenodo, with Croissant metadata and datasheet). Raw audio is available under a separate Data Use Agreement. How will the public release be distributed? Already publicly released via GitHub and Zenodo as described above. Under what licence? Tabular modalities (including audio-derived prosodic features): CC BY 4.0, consistent with participant informed consent. Raw audio: Data Use Agreement (separate access tier). When? Released publicly at the time of this submission (v1.0).
Maintenance Who will maintain the dataset? Meisam Jamshidi Seikavandi ([email protected]), with institutional support from GN Advanced Science, GN Group. Minimum 5-year commitment via versioned tags at the Zenodo repository (https: //zenodo.org/records/20037847). Will it be updated? The dataset may be extended with video-derived 3D pose or room-frame gaze after full QC, under new version tags. This version is v1.0. What ethical review covered the data collection? Participant information statement and informed consent procedures reviewed and approved by the GN Hearing A/S Legal Department; written informed consent from all participants covering academic release of all modalities, including audio-derived tabular features.
32
Table 19: Responsible AI metadata summary for the no-video release. Croissant RAI field
Release statement
Reviewer-facing implication
rai: dataLimitations
Small, single-site, English-language cohort; no room video or room-frame gaze in v1; wearable HRV precision limited by PPG sampling. Convenience sample with likely university/community skew. Age, sex, education, personality, physiology, egocentric gaze, and voice may be sensitive. Strongest for task-state comparisons, modality QC, and exploratory groupaffect modelling. Positive: transparent evaluation resource. Negative risks: affective surveillance and voice re-identification. No synthetic data; derived from direct laboratory collection.
Use for dataset characterisation and feasibility baselines, not populationgeneral claims.
rai:dataBiases rai:personal Sensitive Information rai:dataUseCases
rai:data SocialImpact
rai:has SyntheticData
Do not treat group behaviours as culturally universal. Identity mappings excluded; raw audio requires DUA. Not validated for clinical diagnosis or personnel evaluation. Mitigated by anonymisation, access tiers, and datasheet. Provenance in collection, BIDS export, and benchmark scripts.
NeurIPS Paper Checklist 1. Claims Question: Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? Answer: [Yes] Justification: The abstract states that the paper introduces GroupAffect-4, a multimodal dataset of ten four-person collaborative sessions, and reports feasibility benchmarks across three analysis levels. Both the dataset (§4–5) and the benchmarks (§6) are fully delivered and evaluated in the paper. Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper. • The abstract and/or introduction should clearly state the claims made, including the contributions made in the paper and important assumptions and limitations. • The claims made should match theoretical and experimental results, and reflect how much the results can be expected to generalise to other settings. • It is fine to include aspirational goals as future work as long as it is clear that these are not currently demonstrated. 2. Limitations Question: Does the paper discuss the limitations of the work? Answer: [Yes] Justification: A dedicated §9 discusses sample size (n=10 groups), single-site and English-language constraint, fixed task order, physiological signal missingness, and deferred modalities. Extended caveats appear in Section C. Guidelines: • The answer NA means that the paper has no limitations while the answer No means that the paper has limitations but does not discuss them. • The authors are encouraged to create a separate “Limitations” section in their paper. • The paper should point out any strong assumptions and how robust the results are to violations of these assumptions (e.g., independence assumptions, additional assumptions to establish formal guarantees). • The authors should reflect on how these limitations can affect the usability of the paper, e.g., in terms of generalisability of results and impact on future work. 3. Theory Assumptions and Proofs Question: For each theoretical result, does the paper provide the full set of assumptions and a complete proof? Answer: [N/A] Justification: The paper makes no theoretical claims; it is a dataset and benchmark paper. Guidelines: • The answer NA means that the paper does not include theoretical results. • All the theorems, formulas, and proofs in the paper should be numbered and cross-referenced. • All assumptions should be clearly stated or referenced in the statement of any theorems. • The proofs can either appear in the main paper or the supplemental material, but if they appear in the supplemental material, the authors are encouraged to provide a short proof sketch to provide intuition. 4. Experimental Reproducibility
33
Question: Does the paper fully disclose all the information needed to reproduce the main experimental results of the paper to the extent that it affects the main claims and is not covered by the accompanying dataset paper? Answer: [Yes] Justification: The full processing pipeline and benchmark code are publicly released at https://github.com/ meisamjam/GroupAffect-4. Preprocessing steps, feature selection, LOGO-CV split logic, and Ridge baseline hyperparameters are described in §5, §6, and Sections G and J. Guidelines: • The answer NA means that the paper does not include experiments. • If the paper includes experiments, a No answer to this question will not block the paper from being reviewed, but the authors are encouraged to include information needed to reproduce the results in their main paper. • If the reproducibility of experimental results is crucial, provide a description of the code used to run the experiments. 5. Open access to data and code Question: Does the paper provide open access to the data and code, with sufficient instructions to replicate the results, across all the experiments? Answer: [Yes] Justification: Tabular features and derived data are openly released on Zenodo (https://zenodo.org/records/ 20037847) under CC BY 4.0; processing scripts and benchmark code are on GitHub (https://github.com/ meisamjam/GroupAffect-4). Raw audio is available under a Data Use Agreement due to voice re-identification risk (§10). Guidelines: • The answer NA means that paper does not include experiments requiring code. • Please see the NeurIPS code and data submission guidelines (https://nips.cc/public/guides/ CodeSubmissionPolicy) for more details. • While we encourage the release of code and data, we understand that this might not be possible, so No is an acceptable answer. Authors of accepted papers will be asked to provide a link to a GitHub issues page for users to report bugs. 6. Experimental Setting/Details Question: Does the paper specify all the training and test splits, evaluation metrics, model architectures, and hyperparameters? Answer: [Yes] Justification: §6 specifies Leave-One-Group-Out cross-validation (LOGO-CV), AUC, MAE, and Spearman ρ as metrics. Baseline models are Ridge regression / logistic classifiers with default regularisation; no architecture search is performed. Section J provides per-benchmark interpretation and modality ablation details. Guidelines: • The answer NA means that the paper does not include experiments. • The experimental setting should be presented in the core of the paper to a level of detail that is necessary to appreciate the results and make sense of them. • The full details can be provided either with the code, in appendix, or both. 7. Experiment Statistical Significance Question: Does the paper report error bars suitably and correctly, using standard deviations or confidence intervals? Answer: [Yes] Justification: Benchmark tables (Table 3 and Section J) report standard deviations across LOGO-CV folds. The paper notes where fold counts are too small for stable AUC estimates (Level-2 personality benchmarks, nfold =4). Guidelines: • The answer NA means that the paper does not include experiments. • The authors should answer “Yes” if the results are accompanied by error bars, confidence intervals, or statistical significance tests, at least for the experiments that support the main claims of the paper. • If error bars are not included in the figures and tables, then the results reported should be accompanied by a description of how they were calculated. 8. Experiments Compute Resources Question: For each experiment, does the paper provide sufficient information on the computer resources (type of compute workers, memory, time of execution) used to reproduce it? Answer: [Yes] Justification: All benchmarks use Ridge regression and logistic classifiers (scikit-learn defaults) on a standard CPU workstation. No GPU or specialised hardware is required; runtimes are seconds to minutes per fold. The low-compute nature is noted in §6. Guidelines: • The answer NA means that the paper does not include experiments. • The paper should indicate the type of compute workers CPU/GPU, internal cluster, cloud provider, expected duration. • The paper should provide enough information that a reader can estimate the total amount of compute and make sure their usage is consistent with NeurIPS’ sustainability goal. 9. Code Of Ethics Question: Does the research conform to the NeurIPS Code of Ethics?
34
Answer: [Yes] Justification: Participants were recruited voluntarily, provided informed written consent, and were compensated. No deception was employed. Personal identifiers are not included in the public release. The dataset is gated against high-risk uses (§10). Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics. • If the authors answer No, they should explain the special circumstances that require a deviation from the Code of Ethics. • The authors should make sure to adhere to its guidelines and send a petition to address the violation if needed. 10. Broader Impacts Question: Does the paper discuss both potential positive societal impacts and negative societal impacts of the work performed? Answer: [Yes] Justification: §10 discusses misuse risks (clinical inference, workplace surveillance, speaker re-identification), access tiers to mitigate voice-data risks, and prohibited applications. Positive impacts (affective computing, assistive hearing technology, group-dynamics research) are discussed in §1 and §8. Guidelines: • The statement should include both positive and negative impacts. • At minimum, a brief statement of potential negative impacts should be included in the paper (e.g., in the introduction or a dedicated section). 11. Safeguards Question: Does the paper describe safeguards that have been put in place for responsible release of data or models that have a high risk of being misused (e.g., pretrained language models, image generators, or scraped datasets)? Answer: [Yes] Justification: Raw audio is gated behind a Data Use Agreement to limit voice re-identification risk. The datasheet (Section N) explicitly lists prohibited uses. The public tabular release excludes all direct participant identifiers. Guidelines: • Datasets that can be used to train models with potential for significant harm should be released with necessary safeguards, for example, only on a case-by-case basis or not at all. • NeurIPS expects that the authors describe the safeguards in place and any limitations of those safeguards. 12. Licenses for Existing Assets Question: Are the creators or original owners of assets (e.g., code, data, models), used in this paper, properly credited and given an appropriate citation? Answer: [Yes] Justification: All third-party tools and libraries are cited: EmotiBit SDK, Tobii Pro SDK, openSMILE, scikit-learn, pandas, numpy, and associated works (§4, §5). Guidelines: • Regardless of whether you created or used third-party assets, ensure that you cite the relevant paper(s). • When using existing assets, the relevant licenses must be respected. • If you created a new asset (e.g., a new dataset, new model, new code), list the license of the new asset in the paper. 13. New Assets Question: Are new assets introduced in the paper well documented and is the documentation provided alongside the assets? Answer: [Yes] Justification: The GroupAffect-4 dataset is documented via a Datasheets for Datasets section (Section N), Croissant metadata file (included in the Zenodo deposit), and README in the GitHub repository. Licenses (CC BY 4.0 for tabular data; separate DUA for raw audio) are stated in §10 and the datasheet. Guidelines: • Upload the data, code, and model to an anonymised website such as Anonymous GitHub or Zenodo in a structured way so that it is easy for a reviewer to understand. • Code should be organized and have a README. • Models weights should have a model card. • New datasets should be accompanied by a datasheet. 14. Crowdsourcing and Research with Human Subjects Question: For crowdsourcing experiments and research with human subjects, does the paper include the following information: (i) the full text or a link to informed consent forms, (ii) instructions given to the participants and (iii) compensation? Answer: [Yes] Justification: §3 and the datasheet (Section N) describe participant recruitment, session procedures, and compensation. Participants were recruited via staff networks and affiliated university postings; sessions were conducted with written informed consent. The participant information statement is included in the public Zenodo release (https://zenodo. org/records/20037847). Guidelines:
35
• Ideally, this information should be provided as supplemental material or as a link to a URL containing this information. 15. Institutional Review Board (IRB) Approvals or Equivalent for Research with Human Subjects Question: Does the paper describe whether Institutional Review Board (IRB) approvals or equivalent were obtained? Answer: [Yes] Justification: The participant information statement and informed consent procedures were reviewed and approved by the GN Hearing A/S Legal Department, as stated in §10 and the datasheet maintenance section (Section N). Written informed consent was obtained from all participants covering academic release of all modalities. Guidelines: • Depending on the country in which research was conducted, IRB approval (or equivalent) may be required for any human subjects research. • If the authors obtained IRB approval, they should clearly state this in the paper. • We recognize that the procedures for this may vary significantly between institutions and locations, and we expect authors to adhere to the NeurIPS Code of Ethics and the guidelines for their institution. • For initial submissions, do not include any information that would break anonymity (if applicable), such as the institution conducting the review.
36
Cross-benchmark feature importance heatmap (colour = mean normalised |coef| across LOGO-CV folds) Modality Physiology Eye-tracking Audio Motion/Temp
Overlap fraction Speaking time (s) Audio energy mean Pupil slope Pause count HR SD
1.0
Pupil left mean Pupil mean baseline HR mean baseline Pitch mean
0.8
Audio energy SD Voiced segs/s
Normalised |coefficient|
HRV RMSSD
0.6
SCR count Shimmer mean Unvoiced seg. mean (s) Speech rate proxy
0.4
Skin temp baseline Speaking fraction EDA phasic mean baseline Accel motion mean
0.2
High-motion fraction Pupil right mean Voiced seg. mean (s) Pitch SD
0.0
HRV quality Jitter mean EDA tonic baseline Pupil SD HNR mean EDA phasic rate baseline
Task
B0 l labe
B1ace n Vale
B1sbal u Aro
B3ad man
B2e anc
in Dom
l de nta
Me
B3b t men
E
ge nga
at B_s ion fact s i t Sa
ust B_trTrust
Figure 18: Cross-benchmark feature importance heatmap. Rows are all 31 sensor/behavioural features ranked by mean importance across benchmarks (annotation process-metadata features excluded; see text); columns are benchmark targets. Cell colour encodes normalised |coef| (dark = high importance); y-axis label colour indicates modality group. Audio dominates the left benchmarks (B0, B3a); physiology and pupil features differentiate the affective state benchmarks (B1a, B2). Generated by paper/analysis/feature_importance.py.
37
LOGO-CV per-fold performance across benchmarks (grey dashed line = chance; mean annotated above median)
1.0
AUC / Accuracy
0.8
0.739 0.641
0.668
0.6
0.533
0.583
0.571
0.562
B3b Engagement [AUC]
B_sat Satisfaction [AUC]
B_trust Trust [AUC]
0.499
0.4 0.2 0.0
B0 Task[Acc] label
B1a Valence [AUC]
B1b Arousal [AUC]
B2 Dominance [AUC]
B3a Mental[AUC] demand
Figure 19: LOGO-CV per-fold performance strip plot across all benchmarks. Each dot is one fold; the box shows the interquartile range; the dashed grey line marks the chance baseline for each benchmark. Mean performance is annotated above the median. The wide fold variance for B5 (T4 contribution) and B4 (personality challenges) is visible directly from the fold distribution. Generated by paper/analysis/feature_importance.py.
38