Persona-E2 : A Human-Grounded Dataset for Personality-Shaped Emotional Responses to Textual Events Yuqin Yang Haowu Zhou Haoran Tu Zhiwen Hui Shiqi Yan HaoYang Li Dong She Xianrong Yao Yang Gao Zhanpeng Jin† School of Future Technology South China University of Technology, Guangzhou, China {ftyuqin_yang, 202364870491, 202330691461, 202364870731}@mail.scut.edu.cn {202364871202, ftlhy, ftdshe, ftxryao}@mail.scut.edu.cn {gaoyang2025, zjin† }@scut.edu.cn
arXiv:2604.09162v1 [cs.CL] 10 Apr 2026
Abstract Most affective computing research treats emotion as a static property of text, focusing on the writer’s sentiment while overlooking the reader’s perspective. This approach ignores how individual personalities lead to diverse emotional appraisals of the same event. Although role-playing Large Language Models (LLMs) attempt to simulate such nuanced reactions, they often suffer from “personality illusion”—relying on surface-level stereotypes rather than authentic cognitive logic. A critical bottleneck is the absence of ground-truth human data to link personality traits to emotional shifts. To bridge the gap, we introduce Persona-E2 (Persona-Event2Emotion), a largescale dataset grounded in annotated MBTI and Big Five traits to capture reader-based emotional variations across news, social media, and life narratives. Extensive experiments reveal that state-of-the-art LLMs struggle to capture precise appraisal shifts, particularly in social media domains. Crucially, we find that personality information significantly improves comprehension, with the Big Five traits alleviating “personality illusion.”
1
Introduction
“Two individuals can construe their situations quite similarly (agree on all the facts), and yet react with very different emotions, because they have appraised the adaptational significance of those facts differently.” (Lazarus, 1991) The study of affective appraisal of events has long been central to affective computing and cognitive psychology (Plaza-del Arco et al., 2024). While appraisal theories suggest that emotions emerge through individualized appraisals shaped by goals and dispositions (Lazarus, 1991; Scherer and Wallbott, 1994), NLP research has largely focused on writer-expressed sentiments and readerbased unified emotional labels (Plaza-del Arco et al., 2024). This focus overlooks reader-based
nuanced perception (Buechel and Hahn, 2017b), which is critical for applications, including empathetic agents, mental health support, and personalized AI assistants, that must not only process the texts but also reason about how different individuals appraise the same event diversely. Recent interest in role-playing LLMs aims to simulate individualized reactions by injecting rich personality profiles into prompts (Tseng et al., 2024; Chen et al., 2024; Hu and Collier, 2024; Mao et al., 2024). Despite this promise, these methods often exhibit “personality illusion” (Han et al., 2025): models tend to imitate stereotypical behaviors rather than adopting the cognitive appraisal patterns based on personality. Crucially, LLM-generated labels lack grounding in authentic feedback (Li et al., 2025a), making them insufficient for evaluating whether models truly capture emotional diversity (Samuel et al., 2025). Thus, the field still lacks a human-grounded dataset to validate and enhance personality-conditioned emotion elicitation. To address the gap, we introduce a novel dataset, Persona-E2 (Persona-Event2Emotion), which incorporates the popular Myers-Briggs Type Indicator (MBTI) (Myers et al., 1962; John et al., 1991) and the robust Big Five Inventory (BFI) (John et al., 2010) traits into reader-based emotion labeling. As shown in Fig. 1 by engaging annotators with assessed personality profiles to label events across diverse domains (News, Social Media, Life Experience narratives), Persona-E2 enables a controlled analysis of the personality effect on the appraisals of identical textual events (Troiano et al., 2023). Notably, unlike previous corpora, Persona-E2 prioritizes annotation density (36 labels per event) to capture diverse, trait-shaped responses (Tab. 1). To evaluate the utility of Persona-E2 , we address three key research questions, through the experimental design in Sec. 5:
Data process
Data Collection Content Safety Filtering
News
MultiDimensional LLM Scoring
Evaluation Expert Verification
RQ1: Affective Divergence RQ2: LLM Simulation
Media Event2Emo
Life
Persona Group
Persona-E2
RQ3: Cognitive Soundness
Figure 1: Overview of the Persona-E2 framework. Events from three domains undergo multi-stage data processing. High-quality stimuli are then annotated by a Persona Group, serving to evaluate three research questions.
• RQ1. Affective Divergence: How do emotional responses diverge across the General Writer, General Reader, and Persona Reader, and how is this variance modulated by source domain and personality traits? • RQ2. LLM Simulation: Can LLMs effectively simulate Persona Reader responses, particularly when faced with elicitation conflicts? • RQ3. Cognitive Soundness: Do LLMs generate psychologically grounded rationales for their predictions, and what methods can enhance their cognitive validity? Our analysis reveals that affective appraisal is a domain-sensitive process, with disagreement serving as a structured personality signal. While LLMs struggle to predict precise appraisal shifts, particularly in social media domains, personality traits improve LLMs’ comprehension, and BFI outperforms MBTI in mitigating “personality illusion.” Finally, we release the dataset to support community development.
2
Related Work
Extended discussions are provided in Appendix A. 2.1
Event-Elicited Emotion Analysis
Early research established the baseline for understanding emotions elicited by events. Classic works like ISEAR (Scherer and Wallbott, 1994), SocialIQA (Sap et al., 2019b) and others (Rashkin et al., 2018; Troiano et al., 2019; Forbes et al., 2020) analyzed first-person narratives and social commonsense, treating events as primitive stimuli
for affective responses. Subsequent studies introduced appraisal theory to interpret these cognitive layers in depth (Troiano et al., 2022, 2023). Crucially, the field is shifting from writer-expressed sentiment to reader-based perception (Buechel and Hahn, 2017b). Benchmarks such as GoodNewsEveryone (Bostan et al., 2020), iNews (Hu and Collier, 2025) and RESEMO (Hu et al., 2024a) focus on how audiences react to news and social media. However, most existing resources rely on aggregating annotations into a single ground truth, which obscures the inter-individual variability essential for understanding diverse emotional elicitation (Plank, 2022; Soni et al., 2024). 2.2
Personality-Conditioned Affective Computing
Research on personality–emotion interaction typically utilizes the MBTI (Myers et al., 1962) and the BFI (John et al., 2010) via three paradigms. Explicit methods link self-reported traits to text or dialogue, as seen in datasets like PANDORA (Gjurković et al., 2021), and PersonaTAB (Inoue et al., 2025), though they primarily capture writer expression rather than reader elicitation. Implicit methods infer traits from behavioral data but often lack ground truth (Gao et al., 2013; Wang et al., 2024; Hu et al., 2024b; Shen et al., 2025). Recently, LLM-based simulation has emerged to generate persona-specific responses (Tseng et al., 2024), such as Big5-Chat (Li et al., 2025a), PersonaGym (Samuel et al., 2025), and PersonalityEdit (Mao et al., 2024). Studies show that as richer prompts with profiles are introduced, the behavioral fidelity of simulated
Dataset News-based Domain GoodNewsEveryone (Bostan et al., 2020) NewsMTSC (Hamborg and Donnay, 2021) iNews (Hu and Collier, 2025) Social Media Domain SemEval-2018 Task 1 (Mohammad et al., 2018) GoEmotions (Demszky et al., 2020) SMP2020-EWECT (BrownSweater, 2020) SenWave (Yang et al., 2025b) Life Experience Domain ISEAR (Scherer and Wallbott, 1994) Event2Mind (Rashkin et al., 2018) ATOMIC (Sap et al., 2019a) EmpatheticDialogues (Rashkin et al., 2019) Social IQA (Sap et al., 2019b) Crowd-enVENT (Troiano et al., 2023) Cross-domain Integration (Ours) Persona-E2
Year
#Events
#Annotations
Perspective
#Emotions
Personality
2020 2021 2025
5,000 11,029 2,899
15,000 56,000 14,550
Writer + Reader Writer Reader
15 7 6
✗ ✗ ✗
2018 2020 2020 2025
22,000 58,000 34,768 10,000
700,000 118,000 36,374 20,000
Writer Reader Writer Writer
4 27 6 10
✗ ✗ ✗ ✗
1994 2018 2019 2019 2019 2023
7,666 24,716 24,000 24,850 37,588 6,591
7,666 57,000 72,000 24,850 37,588 11,091
Writer Writer + Reader Experiencer Writer Reader Writer + Reader
7 OV OV 32 OV 13
✗ ✗ ✗ ✗ ✗ ✗
2026
3,111
111,996
Reader
7
✓
Table 1: A unified comparison of emotion-annotated datasets across three sources: News, Social Media, and Life Experience. Note: OV: Open vocabulary.
agents improves accordingly (Bai et al., 2025; Hu and Collier, 2024). Despite their promise, recent works indicate a “personality illusion“ (Han et al., 2025) where models mimic linguistic styles without adopting the underlying appraisal mechanisms. This highlights a critical gap: the lack of a humangrounded dataset to rigorously evaluate whether LLMs truly capture trait-driven emotional diversity.
3
Persona-E2 Dataset Construction
To construct a rigorously controlled dataset for reader-based emotion elicitation, we designed a pipeline integrating heterogeneous event sourcing and a multi-stage filtering process. 3.1
Event Sources
To ensure affective variety and broad coverage, we gather events from three complementary domains—news, social media, and life experience narratives—covering both digital-world and realworld contexts (Appendix B.1). These sources include two distinct elicitation modes: a) First-person projection, where personal experience drives affective memory, and b) Third-person observation, involving detached, personality-shaped appraisals. This constitutes a large-scale reader-centered emotion dataset that integrates the diversity of source domains and elicitation modes. News We crawled factual reports from mainstream news websites, trending topics and verified
institutional accounts. These well-structured texts provide socially significant events that elicit emotions from third-person perspective observations. Social Media We collected posts from several public channels on Reddit. Social content brings greater topical breadth and more ambiguous context, which may elicit empathy or judgment from annotators. To capture more social norms and interpersonal dynamics, we also introduced a subset of events from Social Chemistry 101 (Forbes et al., 2020) to ensure sufficient emotion-eliciting stimuli. Life Experience Life experience narratives were obtained from specific channels dedicated to experience sharing. These events focus on the quotidian experiences of everyday life, ranging from minor frustrations to moments of gratitude. Such narratives are designed to invite first-person projection, serving as a counterpart to the detached perspective of news. 3.2
Event Filtering Pipeline
Raw collections inevitably contain noise, safety risks, and large quantities of content that lack emotional significance. To ensure high-quality stimuli, we implemented a 3-stage filtering procedure. Stage 1: Content Safety Filtering. For life experience and social media domains, we first prune toxic or sensitive content using NSFW classifiers (Albouzidi, 2023; TostAI, 2023). More detailed information is illustrated in Appendix B.2.1.