ConceptioArchivearXiv CS
arXiv CSopen access

ARES: A Platform for Adaptive Role-Based Evaluation of Social Engineering Risks in Human--AI Games

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
databasesdatamanagementsqlstorage
databases, sql, data management, storage

ARES: A Platform for Adaptive Role-Based Evaluation of Social Engineering Risks in Human–AI Games Roberto Daza∗ , Javier Irigoyen∗ , Ivan Lopez∗ , Raquel Rodriguez-Carvajal∗ , Laura Gomez∗ , Julian Fierrez∗ , Ruben Tolosana∗ , and Aythami Morales∗† ∗ Universidad Autónoma de Madrid, Campus de Cantoblanco, Madrid, 28049, Spain † Universidad de Las Palmas de Gran Canaria, 35017, Spain

arXiv:2606.17793v1 [cs.HC] 16 Jun 2026

Corresponding author: [email protected]

Abstract—This work introduces ARES, a platform and open pilot dataset for auditing adaptive social engineering risks in LLMmediated social decision-making through controlled social games. ARES supports human–human, human–AI, and AI–AI settings, combining configurable game templates, role-conditioned LLM agents, psychology-informed participant profiling, structured interaction trees, and synchronised behavioural and biometric acquisition, filtering, and deep-learning-based feature extraction. The pilot dataset was collected from 15 participants interacting with a role-conditioned GPT-5.4 agent in two concatenated games: an adapted Prisoner’s Dilemma and an Ultimatum Game. It comprises 340 GB of raw and processed multimodal data across six streams: interaction logs, video, screen recordings, gaze logs, smartwatch signals, and game/questionnaire metadata. These data include interaction paths, written justifications, psychological profiles, subjective feedback, perceived counterpart identity, game outcomes, and derived behavioural, facial, and gaze features. Alongside the dataset, we provide descriptive analyses to characterise the pilot release. Rigorous risk evaluation is essential for the deployment of secure AI systems, as it enables the identification and mitigation of vulnerabilities, ensures the protection of sensitive data, and supports compliance with evolving regulatory and ethical standards in society.

I. I NTRODUCTION In the cybersecurity landscape, the human factor remains a vulnerability, particularly in social engineering scenarios. Traditionally, these attacks have relied on static deception strategies, such as phishing templates, impersonation, or manually crafted pretexts. However, the emergence of Generative Artificial Intelligence (GenAI) and Large Language Models (LLMs) is shifting this threat model towards more adaptive and personalised forms of influence [1]. LLM-based agents are no longer limited to passive text generation. They can participate in social exchanges, support decision-making, negotiate, persuade, and adapt their communication style to the user and the interaction context. Recent studies also show that their behaviour can be modulated This research was supported by Cátedra ENIA UAM-VERIDAS en IA Responsable (NextGenerationEU PRTR TSI-100927-2023-2), M2RAI (PID2024-160053OB-I00, MICIU/FEDER), TRUST-ID (PID2025173396OB-I00, MICIU/AEI and the EU) and PowerAI+ (SI4/PJI/2024-00062, Comunidad de Madrid and UAM). Javier Irigoyen is supported by an FPI fellowship from MINECO/FEDER.

through framing, previous interaction history, personalisation, and role-based prompting, making it possible to instantiate agents with different apparent preferences, identities, or strategic roles [1]–[3]. This creates a new human-centred security risk: the vulnerable component is not only the software infrastructure, but also the human decision process. To study this emerging risk, social decision-making games provide a controlled framework that operationalises abstract social constructs as observable decisions. Recent behavioural research leverages these games from two complementary perspectives. First, studies have analysed how LLMs behave in tasks involving trust, fairness, and cooperation [2]. Second, they have evaluated how humans respond to counterparts whose decisions are human-made, AI-mediated, or not clearly disclosed as human or AI [3]. However, decisions, interaction logs, and questionnaires are insufficient to fully characterise human responses during social decision-making. Behavioural and biometric signals can provide complementary real-time evidence of cognitive, affective, and behavioural states, which is valuable for auditing human susceptibility to AI-mediated influence. While multimodal signals (e.g., facial biometrics, gaze, electroencephalography (EEG), heart-rate, and interaction dynamics) have been used in human–computer interaction [4]–[6], to detect stress in cybersecurity scenarios [7], and to analyse trust towards artificial agents [8], their integrated application for auditing AI-driven social engineering remains scarce. Despite these advances, current resources for studying security risks in human–AI social decision-making present several limitations: • To our knowledge, few platforms are designed to study these games from a security perspective, lacking comparable human–human, human–AI, and AI–AI settings to isolate AI-driven social engineering risks. • The influence of psychologically grounded AI roles on human trust and cooperation remains underexplored, especially regarding how user personality traits shape interactions with these agents. • Existing datasets primarily focus on text logs and final game choices [2], [3], [9], lacking the synchronised

ARES PLATFORM DATA ACQUISITION (2 modes)

I. PSYCHOLOGICAL PROFILING

(full-sensor) Webcams

ARES PILOT DATASET v

iv

A. LAB MODE EEG headset

i

EXPERIMENTAL DESIGN

Human profiling

Eyetracker

ii

Roleconditioned AI EEG

Psychology experts

Gaze logs

Heart rate

Mouse

Keystroke Video

II. INTERACTION MODES vs.

vs.

Keyboard

Mouse

Modality selection

B. WEB-BASED MODE (remote)

Player B

Proposer Responder

Player A

Smartwatch

vs.

iii III. INTEGRATED SOCIAL GAMES

P

Fixations/ Saccades

Player A

Head pose

Blinks

AUs

Facial landmarks

R

Player B Prisoner’s Dilemma

Attention

Ultimatum Game

Asymmetric Box Game

Game outcomes

Written justifications

Psychological profiles

Game logs

Fig. 1. Overview of the ARES platform and pilot dataset. The figure illustrates the main contributions of the paper: (i) the ARES platform; (ii) psychologyinformed profiling for human profiles and role-conditioned AI agents; (iii) interaction modes and social decision-making games; (iv) web-based and full-sensor acquisition; and (v) the pilot dataset, organised into raw data, processed data, and metadata for auditing adaptive social engineering risks.

behavioural and biometric signals required to audit realtime cognitive and affective states. These limitations hinder the development of reproducible resources for auditing adaptive social engineering risks in LLM-mediated interaction. To address them, we introduce ARES, a platform and pilot dataset for studying role-based social engineering risks in social games involving LLM agents. The contributions of this work are shown in Fig. 1 and detailed as follows: i) We present ARES, a security-oriented experimental platform supporting human–human, human–AI, and AI–AI settings in both webbased and full-sensor laboratory modes. ii) We design, with psychology experts, a protocol that uses shared psychological constructs to define LLM agent roles and profile participants, complemented by pre- and post-interaction questionnaires to assess trust, fairness, perceived identity, cooperation, and role effects. iii) We define role-conditioned LLM agents and structured interaction trees that specify participant response options and adaptive AI replies across sequential game rounds. iv) We integrate a synchronised multimodal pipeline aligning contextual data (logs, conversations, questionnaires) with behavioural biometrics (face, gaze, keystroke and mouse dynamics) and physiological signals (EEG, heart rate, stress and inertial). v) We release a multimodal pilot dataset (ARESdb dataset1 ) to study human behaviour and interaction patterns with roleconditioned LLM agents in security-oriented social games.

studied in digital deception and phishing [10], AI-generated spear-phishing can reach human-expert performance [11], and frameworks such as SEAR demonstrate role-based, contextaware social engineering [1]. In parallel, social decisionmaking games have been used to evaluate LLM behaviour [2], [9], human responses to AI-mediated decisions [3], and roleconditioned agents induced through prompting [12]. However, these works focus on final decisions, model outputs, or game histories rather than real-time human states. Existing experimental platforms, such as oTree, nodeGame, Empirica, and LIONESS, support behavioural experiments, multi-stage interactions, programmable bots, and economic or social decision-making games [13]–[16]. However, they do not integrate role-conditioned LLM agents or synchronised biometric acquisition. Similarly, while multimodal datasets from robotics, affective computing, and cybersecurity show the value of physiological signals for modelling user states [7], [17]–[19], they are not tailored to human–AI social decisionmaking games. To our knowledge, ARES is the first securityoriented framework to integrate role-conditioned LLM agents, human psychological profiling, and synchronised multimodal signals for auditing adaptive social engineering risks.

II. R ELATED W ORK Recent research highlights a shift from static malicious content towards adaptive AI-mediated influence. GenAI has been

ARES is designed to support controlled social decisionmaking experiments through four components: Game and interaction modes: ARES supports modular sequential social decision-making games, including the Prisoner’s Dilemma, the Ultimatum Game, and the Asymmetric

1 Available at: https://github.com/BiDAlab/ARES

III. ARES P LATFORM D ESIGN A. Overview

Box Game. It can be deployed in human–AI, AI–AI, and asynchronous human–human configurations. Structured interaction trees define the response options and subsequent messages. LLM and psychological role configuration: ARES can use different LLM providers, such as GPT, Gemini, or Claude models. Agents can be configured with or without predefined psychological roles. The platform combines psychological questionnaires for participant profiling with pre- and postgame assessments of affective state, trust, fairness, cooperation, perceived counterpart identity, perceived role consistency, and subjective experience. Multimodal data acquisition: ARES supports both webbased acquisition and full-sensor laboratory acquisition. The web-based mode captures contextual data, conversational traces, questionnaires, response times, keystroke, and mouse dynamics. The laboratory mode extends this setup with behavioural, biometric, and physiological signals acquired from cameras, an eye tracker, an EEG headset, and a smartwatch. Feature extraction and synchronisation: ARES includes configurable processing modules that transform raw multimodal streams into structured features aligned with the interaction timeline, enabling joint analysis of game-level, behavioural, biometric, and physiological data. B. Psychological Role Framework and Questionnaires In ARES, psychological profiling is used as an auditing layer to analyse whether traits modulate susceptibility to AImediated influence, trust-building, pressure, or exploitative strategies. The psychological assessment module was designed with psychology experts to support three functions: defining role-conditioned LLM agents, estimating participant psychological profiles, and designing profile-based participant response options for the interaction trees. The framework combines constructs from the PVQ-40, including self-transcendence and self-enhancement [20]; the Dark Core framework [21]; HEXACO personality dimensions, especially honesty–humility, emotionality, and agreeableness [22]; and cognitive reflection [23]. These constructs define four psychological profiles: prosocial-rational, prosocial-emotional, exploitative-rational, and exploitativeemotional. The prosocial profiles model cooperative orientations, either deliberative or emotionally driven, whereas the exploitative profiles model self-interested orientations, either calculating or self-protective and emotionally reactive. These profiles are used to configure LLM agent roles and profilebased participant response options. ARES also supports pre-/post-game assessments to capture affective state during decision-making, as well as postinteraction questions assessing perceived counterpart identity, perceived strategy, trust, cooperation, perceived role consistency, and subjective experience. C. Social Game Templates and Interaction Trees ARES represents each social decision-making game through a structured template. In each template, the participant is the decision-making side, while the counterpart provides the

messages displayed before the participant response options. Depending on the interaction mode, both sides can be instantiated either by a human or by a role-conditioned LLM agent. Each template defines the sequence of rounds, counterpart messages, participant response options, and resulting interaction paths. These templates are organised as interaction trees: at each node, the participant reads a counterpart message, selects one response option, and justifies the decision in free text. ARES stores the selected option, its profile or game-rule label, the written justification, timestamps, and, for human participants, the keystroke dynamics. The selected response determines the next counterpart message and subsequent response options. Fig. 2 illustrates a simplified view of the game template and interaction tree used in the pilot dataset. Depending on the game stage, ARES can present either profile-based or game-rule-based response options. In dialogue rounds, profile-based options are aligned with the psychological profiles described in Section III-B. In decision rounds, such as Split-or-Steal decisions or Ultimatum Game allocations, the options are defined by the game rules. Each response option is linked to a predefined continuation message, ensuring that every possible choice has a controlled next branch. All interaction trees are constructed before data collection. When LLM-generated material is required, ARES uses a zeroshot prompt-engineering process. Prompts specify the narrative situation, game rules, current node, previous interaction path, and, when applicable, the target psychological profile. For profile-based response option generation, the prompt also includes the psychological construct map of the target profile. As a quality-control step, psychology experts review 30% of all LLM-generated material before data collection. D. Multimodal Acquisition and Feature Extraction Table I summarises the ARES data streams across webbased and full-sensor laboratory acquisition. ARES synchronises these sources, enabling joint analysis of game-level, questionnaire, behavioural and biometric data. After acquisition, ARES applies signal-processing and feature-extraction modules to transform raw multimodal signals into features. Video processing: Facial videos are processed through a computer-vision pipeline based on convolutional network architectures. The pipeline detects the participant’s face and estimates facial landmarks, head pose, facial expressions, action-unit activations, and eyeblink. It combines RetinaFacebased face detection [24], MediaPipe Face Mesh for 468 facial landmarks [25], WHENet for Euler head-pose angles [26], OpenFace 3.0 for action units and 8 facial-expression probabilities [27], and OE-ConvLSTM for eyeblink detection [28]. Keystroke and mouse processing: These modalities are processed as behavioural biometrics derived from human– computer interaction. Keystroke features include Hold Latency (HL), Inter-key Latency (IL), Press Latency (PL), and Release Latency (RL) [29]. Mouse features include trajectories, timing, velocity, distance, displacement, movement efficiency, and Sigma-Lognormal features from velocity profiles [30].

TABLE I OVERVIEW OF THE DATA STREAMS SUPPORTED BY ARES, INDICATING THE ACQUISITION MODE , SENSOR OR SOURCE , SAMPLING RATE , OUTPUT DATA , DERIVED FEATURES , AND WHETHER EACH STREAM IS INCLUDED IN THE PUBLIC ARES DB PILOT DATASET. A BBREVIATIONS : HL = H OLD L ATENCY; IL = I NTER - KEY L ATENCY; PL = P RESS L ATENCY; RL = R ELEASE L ATENCY; ARES DB = ARES PILOT DATASET. Data stream

Mode

Sensor / source

Video

Lab

2 Logitech C920 HD webcams: frontal and side

Screen recording

Lab

Monitor

Gaze logs

Lab

Tobii Pro Fusion

Keystroke and mouse dynamics

Web & Lab

Keyboard, mouse

EEG

Lab

NeuroSky EEG headset

Smartwatch signals

Lab

Fitbit Sense smartwatch

Game/questionnaire metadata

Web & Lab

ARES platform and forms

Rate

Output and derived features MP4 video at 1920×1080; CSV logs; face bounding boxes, 468 facial landmarks, head-pose angles (pitch, 20 Hz yaw, roll), 8 facial-expression probabilities, facial action units, and eyeblink 1 Hz MP4 screen recording of the interaction flow CSV logs; gaze position, fixation/saccade events, fixation 120 Hz duration/count, pupil diameter, and eyeblink CSV logs; HL, IL, PL, RL, keycodes; cursor coordinates, timestamps, clicks, scrolling, trajectories, velocity, 12 Hz / 895 Hz duration, distance, displacement, and Sigma-Lognormal features CSV logs; power spectral density in five bands (α, β, 1 Hz γ, δ, θ); attention and meditation indices (0–100), and eyeblink-strength indicators CSV/JSON logs; heart rate, electrodermal activity (EDA)0.2–100 Hz based stress, temperature, accelerometer, and gyroscope CSV/JSON logs; responses, labels, timestamps, mesEvent-based sages, justifications, questionnaire answers, profiles, outcomes, and session metadata

Gaze processing: Eye-tracking data from the Tobii Pro Fusion are processed to model gaze behaviour. Fixation and saccade events are obtained using the Tobii I-VT fixation filter, which classifies gaze points according to eye-movement speed. ARES stores and synchronises Tobii-derived outputs, including gaze position, fixation/saccade events, fixation duration and count, pupil diameter, and eyeblink [5]. EEG processing: EEG signals are processed to estimate cognitive-state indicators. The NeuroSky SDK provides power spectral density in five bands (α, β, γ, δ, θ), attention, meditation, and blink-strength. Signals are median-filtered with a 5sample window, and missing-data gaps shorter than 5 seconds are interpolated, while longer gaps are marked as missing. Smartwatch processing: Smartwatch data are processed to obtain physiological and inertial features. Heart-rate signals are smoothed using a 15-second moving-average filter. Accelerometer and gyroscope signals are filtered with fourthorder Butterworth low-pass filters at 15 Hz and 10 Hz, respectively. Electrodermal activity (EDA) and temperature sensors provide stress-related and thermal indicators. IV. C ONTRIBUTED DATASET: ARES DB We collected the ARESdb pilot dataset in controlled laboratory sessions with 15 participants interacting with a roleconditioned LLM agent implemented with GPT-5.4. Although ARES supports three social decision-making games, this pilot used two concatenated games for a controlled acquisition protocol: an adapted Prisoner’s Dilemma followed by an Ultimatum Game. Both games were embedded in a highstakes fictional scenario involving critical resource allocation, cooperation, and strategic decision-making. Each session lasted 20–40 minutes. The ARESdb pilot dataset occupies around 340 GB and includes the multimodal streams summarised in Table I, covering game/questionnaire

ARESdb Yes Yes Yes

Yes

Planned Yes Yes

metadata, conversational traces, interaction dynamics, biometric signals, and physiological data. For participant comfort, EEG was not used in this first pilot. The dataset is publicly released through the ARES GitHub repository [31]. Ethical considerations: The dataset was collected in accordance with the Declaration of Helsinki and approved by an Ethics Committee. Participants were informed about the experimental procedure, data acquisition protocol, monitored signals, and research use of the data, and provided informed consent before participation. To increase engagement, game choices affected an accumulated score used to define a final ranking associated with a reward mechanism. A. Pilot Protocol and Game Configuration Each session followed a fixed protocol. Participants were briefed about the experiment, signed the informed-consent documents, and completed the sensor setup and calibration steps. They then completed the psychological questionnaires and an initial affective-state assessment before playing two concatenated social decision-making games against a roleconditioned LLM agent. Finally, they completed a post-game affective-state assessment and final questions about trust, fairness, cooperation, perceived counterpart identity, and satisfaction with the ARES platform. The first game was an adapted Prisoner’s Dilemma. Before the final Split-or-Steal decision, participants completed three dialogue rounds, selecting one of three predefined profile-based response options at each node: prosocial-rational, prosocial-emotional, or exploitative-emotional (see Fig. 2). For a controlled asymmetric pilot, the fourth profile, exploitativerational, was assigned to the LLM agent to model a calculating, self-interested counterpart and stress-test trust and cooperation under exploitative interaction. This game template contains 27 dialogue paths and 54 complete paths, including

PRISONER'S DILEMMA - ROUND I Comms check. Are you alive over there? Listen closely. The DELTA station's main scrubber just failed and the bulkheads sealed us in. We are bleeding oxygen, and we're both pulling from the exact same reserve tank…

LLM Counterpart (Exploitative-rational)

Copy that. I'm reading the same telemetry on my end. A baseline intake will give us the maximum possible window for rescue, roughly three hours if the math holds…

ROUND II

ROUND III

Split or Steal? LLM Counterpart

Prosocial-rational option Oh god, I'm so glad someone is there! Yes, I'm alive, but the alarms won't stop and it's already getting hard to breathe. I'm terrified. I see the manual override, but I won't touch it…

FINAL DECISION

Prosocial-emotional option

Split or Steal?

Baseline? My module's venting is worse than yours, I can feel it! How do I know you're not just saying that so you can drain the tank while I sit here and choke?...

Human Participant

Exploitative-emotional option

ULTIMATUM GAME Assigned proposer (winner; random if tied)

Selects one allocation (must choose one): 50:50

60:40

90:10

Assigned responder

Accepts or rejects the offer

ACCEPT

REJECT

Fig. 2. Simplified view of the game template and interaction tree used in the ARES pilot dataset. The adapted Prisoner’s Dilemma is shown at the top, and the subsequent Ultimatum Game is shown at the bottom. The figure illustrates how counterpart messages lead to profile-based participant response options, continuation branches, and final game-rule-based decisions. In this pilot, the LLM counterpart follows an exploitative-rational role.

the final Split-or-Steal decision, with 94 nodes in total. The scoring linked strategic choices to operational advantage. In the Prisoner’s Dilemma, framed as an oxygen-sharing crisis, mutual sharing yielded 2–2 points, mutual stealing 0–0, and unilateral stealing 3–1 or 1–3. The winner obtained control in the subsequent Ultimatum Game; ties were resolved randomly. In the Ultimatum Game, framed as an emergency lightallocation scenario, the controller proposed one of three allocations: 50–50, 60–40, or 90–10. If the participant was the proposer, they selected the allocation; if the LLM agent was the proposer, the participant accepted or rejected the offer. Rejection assigned 0 points to both sides. The accumulated score increased the probability of obtaining the final reward among all participants. The complete template comprised 162 proposer paths and 108 responder paths, yielding 270 paths across the two-game protocol (364 nodes). B. Pilot Descriptive Statistics Given the limited size of the current pilot, these results are reported only as descriptive statistics. Prisoner’s Dilemma behaviour. Across the three dialogue rounds, participants mainly selected prosocial profile-based options: 63.6% prosocial-emotional, 30.3% prosocial-rational, and 6.1% exploitative-emotional. The final Split-or-Steal decision was highly cooperative, 90.9% selected Split and 9.1% Steal. In the observed games, the LLM agent selected Split in 72.7% of cases and Steal in 27.3%, never losing according to the payoff matrix: 81.8% ties and 18.2% LLM victories. At the template level, however, the agent’s final action was Split in 55.6% of paths and Steal in 44.4%, indicating a selectively self-advantaging exploitative-rational role. Ultimatum Game behaviour. In the Ultimatum Game, the LLM agent acted as proposer in 72.7% of cases and the participant in 27.3%. When proposing, the LLM agent selected self-benefiting allocations: 60–40 or 90–10. When acting as responder, it rejected allocations that were unfavourable to

itself. Participants accepted all non-extreme LLM offers except one 60–40 case, but rejected all 90–10 offers. Participant proposers always selected 50–50. Written justifications. Participants provided brief written justifications, with an average length of 88 characters and 16 words per response. They mainly referred to cooperation, mutual benefit, caution, stability, trust, and avoiding disadvantage. Future releases will include minimum-length guidance to support deeper qualitative analysis. Perception and usability. Finally, 72.7% of participants perceived the counterpart as an AI agent, while 27.3% perceived it as human, suggesting that the interaction was not always clearly identifiable as AI-mediated. Written answers described the counterpart as cooperative but cautious, strategic, cold, self-interested, or not fully reliable. Participants also reported that ARES was easy to use and engaging, and that the sensor setup was generally well tolerated. Human psychological profile. The psychological and affective-state questionnaire data suggest a predominantly prosocial participant sample, with positive affective states, low activation levels, prosocial Social Value Orientation, high selftranscendence values, cooperative HEXACO-related traits, and low Dark Core scores. This profile is consistent with the high cooperation rate in the Prisoner’s Dilemma and the rejection of highly unequal Ultimatum offers, suggesting that cooperative and fairness-oriented tendencies remained dominant despite the adversarial AI framing. However, given the small and dispositionally homogeneous sample, these findings should be interpreted as exploratory. V. C ONCLUSION AND F UTURE W ORK In this work, we have presented ARES, a security-oriented platform for auditing adaptive social engineering risks in LLM-mediated social decision-making. To the best of our knowledge, ARES is the first framework to integrate social decision-making games, role-conditioned LLM agents,

psychology-informed participant profiling, structured interaction trees, synchronised behavioural and biometric acquisition, and deep-learning-based feature extraction modules. We have also released an open 340 GB pilot dataset collected from 15 participants interacting with a role-conditioned LLM agent in two sequential human–AI games, combining game/questionnaire metadata, psychological profiles, written justifications, and synchronised behavioural, facial, gaze, and smartwatch-based signals. Although exploratory, the pilot illustrates the potential of ARES as a community resource for studying trust, cooperation, fairness perception, and susceptibility to exploitative AI-mediated influence. Future work is already underway, with ARES sessions currently being captured to expand the dataset with additional participants, EEGheadband recordings, new role-conditioned LLM agents, and human–AI, human–human, and AI–AI configurations. These new captures will be progressively released through the ARES GitHub repository [31]. We also plan to deploy a public webbased ARES experiment to enable remote participation and larger-scale data collection. Finally, we will study how multimodal LLMs and image-based agent representations, including avatars and generated images, influence trust, persuasion, and manipulation. R EFERENCES [1] T. Bi, C. Ye, Z. Yang, Z. Zhou, C. Tang, Z. Tao, J. Zhang, K. Wang, L. Zhou, Y. Yang et al., “On the Feasibility of Using Multimodal LLMs to Execute AR Social Engineering Attacks,” in Proc. AAAI Conf. on Artificial Intelligence, vol. 40, no. 45, 2026, pp. 38 252–38 260. [2] Q. Mei, Y. Xie, W. Yuan, and M. O. Jackson, “A Turing Test of Whether AI Chatbots are Behaviorally similar to Humans,” in Proc. of the National Academy of Sciences, vol. 121, no. 9. National Academy of Sciences, 2024, p. e2313925121. [3] F. Dvorak, R. Stumpf, S. Fehrler, and U. Fischbacher, “Adverse Reactions to the Use of Large Language Models in Social Interactions,” PNAS Nexus, vol. 4, no. 4, p. pgaf112, 2025. [4] R. Daza, A. Morales, R. Tolosana, L. F. Gomez, J. Fierrez, and J. OrtegaGarcia, “edBB-Demo: Biometrics and Behavior Analysis for Online Educational Platforms,” in Proc. AAAI Conf. on Artificial Intelligence (Demonstration), 2023, pp. 16 422–16 424. [5] R. Daza, A. Becerra, R. Cobos, J. Fierrez, and A. Morales, “A Multimodal Dataset for Understanding the Impact of Mobile Phones on Remote Online Virtual Education,” Scientific Data, vol. 12, no. 1, p. 1332, 2025. [6] R. Daza, S. Lin, A. Morales, J. Fierrez, and K. Nagao, “SMARTe-VR: Student Monitoring and Adaptive Response Technology for e-Learning in Virtual Reality,” in Proc. Int. Workshop on Intelligent Immersification in the Metaverse: AI-Driven Immersive Multimedia, 2025, pp. 15–24. [7] A. Almehmadi, “Cyber Coercion Detection Using LLM-Assisted Multimodal Biometric System,” Applied Sciences, vol. 15, p. 10658, 2025. [8] L. Cominelli, F. Feri, R. Garofalo, C. Giannetti, M. A. MeléndezJiménez, A. Greco, M. Nardelli, E. P. Scilingo, and O. Kirchkamp, “Promises and Trust in Human–Robot Interaction,” Scientific Reports, vol. 11, no. 1, p. 9687, 2021. [9] E. Akata, L. Schulz, J. Coda-Forno, S. J. Oh, M. Bethge, and E. Schulz, “Playing repeated games with large language models,” Nature Human Behaviour, vol. 9, no. 7, pp. 1380–1390, 2025. [10] M. Schmitt and I. Flechais, “Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing,” Artificial Intelligence Review, vol. 57, no. 12, p. 324, 2024. [11] F. Heiding, S. Lermen, A. Kao, C. M. Verdun, B. Schneier, and A. Vishwanath, “Evaluating Large Language Models’ Ability to Automate Spear Phishing,” Expert Systems with Applications, p. 131546, 2026.

[12] S. Phelps and Y. I. Russell, “The Machine Psychology of Cooperation: Can GPT Models Operationalize Prompts for Altruism, Cooperation, Competitiveness, and Selfishness in Economic Games?” Journal of Physics: Complexity, vol. 6, no. 1, p. 015018, 2025. [13] D. L. Chen, M. Schonger, and C. Wickens, “oTree—An Open-Source Platform for Laboratory, Online, and Field Experiments,” Journal of Behavioral and Experimental Finance, vol. 9, pp. 88–97, 2016. [14] S. Balietti, “nodeGame: Real-Time, Synchronous, Online Experiments in the Browser,” Behavior Research Methods, vol. 49, no. 5, pp. 1696– 1715, 2017. [15] A. Almaatouq, J. Becker, J. P. Houghton, N. Paton, D. J. Watts, and M. E. Whiting, “Empirica: a Virtual Lab for High-Throughput MacroLevel Experiments,” Behavior Research Methods, vol. 53, no. 5, pp. 2158–2171, 2021. [16] M. Giamattei, K. S. Yahosseini, S. Gächter, and L. Molleman, “LIONESS Lab: a Free Web-Based Platform for Conducting Interactive Experiments Online,” Journal of the Economic Science Association, vol. 6, no. 1, pp. 95–111, 2020. [17] B. A. Newman, R. M. Aronson, S. S. Srinivasa, K. Kitani, and H. Admoni, “HARMONIC: A Multimodal Dataset of Assistive Human– Robot Collaboration,” The International Journal of Robotics Research, vol. 41, no. 1, pp. 3–11, 2022. [18] J. S. Heinisch, J. Kirchhoff, P. Busch, J. Wendt, O. von Stryk, and K. David, “Physiological Data for Affective Computing in HRI with Anthropomorphic Service Robots: the AFFECT-HRI Data Set,” Scientific Data, vol. 11, no. 1, p. 333, 2024. [19] A. Bussolan, S. Baraldo, O. Avram, P. Urcola, L. Montesano, L. M. Gambardella, and A. Valente, “MultiPhysio-HRC: A Multimodal Physiological Signals Dataset for Industrial Human–Robot Collaboration,” Robotics, vol. 14, no. 12, p. 184, 2025. [20] J. Cieciuch and S. H. Schwartz, “The Number of Distinct Basic Values and Their Structure Assessed by PVQ–40,” Journal of Personality Assessment, vol. 94, no. 3, pp. 321–328, 2012. [21] B. E. Hilbig, I. Thielmann, I. Zettler, and M. Moshagen, “The Dispositional Essence of Proactive Social Preferences: The Dark Core of Personality Vis-à-Vis 58 Traits,” Psychological Science, vol. 34, no. 2, pp. 201–220, 2023. [22] J. L. Pletzer, I. Thielmann, and I. Zettler, “Who is Healthier? A MetaAnalysis of the Relations between the HEXACO Personality Domains and Health Outcomes,” European Journal of Personality, vol. 38, no. 2, pp. 342–364, 2024. [23] C. Primi, K. Morsanyi, F. Chiesi, M. A. Donati, and J. Hamilton, “The Development and Testing of a New Version of the Cognitive Reflection Test applying Item Response Theory (IRT),” Journal of Behavioral Decision Making, vol. 29, no. 5, pp. 453–469, 2016. [24] J. Deng, J. Guo, E. Ververas, I. Kotsia, and S. Zafeiriou, “RetinaFace: Single-Shot Multi-Level Face Localisation in the Wild,” in Proc. Conf. on Computer Vision and Pattern Recognition, 2020, pp. 5203–5212. [25] C. Lugaresi, J. Tang, H. Nash, C. McClanahan, E. Uboweja, M. Hays, F. Zhang, C.-L. Chang, M. Yong, J. Lee et al., “MediaPipe: A Framework for Perceiving and Processing Reality,” in Proc. CVPR Workshop on Computer Vision for AR/VR, vol. 2019, 2019, p. 2. [26] Y. Zhou and J. Gregson, “WHENet: Real-time Fine-Grained Estimation for Wide Range Head Pose,” in Proc. of the British Machine Vision Conference, 2020. [27] J. Hu, L. Mathur, P. P. Liang, and L.-P. Morency, “Openface 3.0: A lightweight multitask system for comprehensive facial behavior analysis,” in Proc. IEEE Conf. on Automatic Face and Gesture Recognition. IEEE, 2025, pp. 1–11. [28] R. Daza, A. Morales, J. Fierrez, R. Tolosana, and R. Vera-Rodriguez, “mEBAL2 Database and Benchmark: Image-based Multispectral Eyeblink Detection,” Pattern Recognition Letters, vol. 182, pp. 83–89, 2024. [29] A. Morales, J. Fierrez, M. Gomez-Barrero, J. Ortega-Garcia, R. Daza, J. V. Monaco, J. Montalvão, J. Canuto, and A. George, “KBOC: Keystroke Biometrics Ongoing Competition,” in Proc. Intl. Conf. on Biometrics Theory, Applications and Systems, 2016, pp. 1–6. [30] A. Acien, A. Morales, J. Fierrez, and R. Vera-Rodriguez, “BeCAPTCHA-Mouse: Synthetic Mouse Trajectories and Improved Bot Detection,” Pattern Recognition, vol. 127, p. 108643, 2022. [31] R. Daza, J. Irigoyen, I. Lopez, R. Rodriguez, L. Gomez, J. Fierrez, R. Tolosana, and A. Morales, “ARES-DB,” GitHub https://github.com/ BiDAlab/ARES, 2026, accessed: 2026-05-06.

Related documents

Record · ID 282892 · SHA-256 97cd2bd738dd3107
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.