arXiv:2604.13804v1 [cs.LG] 15 Apr 2026
Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning Dongjie Fu∗
Fangming Feng∗
Xize Cheng∗
Zhejiang University Hangzhou, China [email protected]
Zhejiang University Hangzhou, China [email protected]
Zhejiang University Hangzhou, China [email protected]
Linjun Li
Zhou Zhao
Tao Jin†
Meituan Shanghai, China [email protected]
Zhejiang University Hangzhou, China [email protected]
Zhejiang University Hangzhou, China [email protected]
Abstract The rapid evolution of multimodal large models has revolutionized the simulation of diverse characters in speech dialogue systems, enabling a novel interactive paradigm. Character attributes are manifested not only in textual responses but also through vocal features, as speech conveys rich paralinguistic information that is challenging to quantify. This poses significant difficulties in evaluating the character alignment of role-playing agents. To address these challenges, we present RoleJudge, an evaluation framework that leverages audio large language models to systematically assess the alignment between speech and character across multiple modalities and dimensions. Furthermore, we introduce RoleChat, the first voice role-playing evaluation dataset enriched with chainof-thought reasoning annotations, comprising a diverse set of authentic and LLM-generated speech samples. Utilizing this dataset, we implement a multi-stage training paradigm and incorporate Standard Alignment in reinforcement learning to mitigate reward misalignment during optimization. Experimental results in terms of accuracy and subjective assessment demonstrate that RoleJudge outperforms various baseline models, validating the effectiveness of our multidimensional evaluation framework.
CCS Concepts • Computing methodologies → Artificial intelligence; Natural language processing.
Keywords Audio Large Language Models, Role-Playing Evaluation, Reinforcement Learning
1
Introduction
The continuous advancement of artificial intelligence is profoundly transforming the way humans interact with digital systems, giving rise to new forms of digital life that seamlessly integrate technology with human experience. Among these innovations, Role-Playing Agents (RPAs) are particularly noteworthy, as they embody our aspiration to create virtual entities capable of understanding, responding, and interacting with users in increasingly human-like ways. By simulating a wide range of characters, from historical ∗ Equal contribution. † Corresponding author.
figures and fictional personalities to everyday individuals, these agents open up new possibilities for virtual assistants, interactive storytelling, and intelligent game characters. Driven by large language models [2, 11, 12, 37], text-based RPAs are becoming a reality [26, 27, 32], extending to novel applications such as digital humans and character-driven video games [36]. With the increasing integration of multimodal technologies and largescale models [5, 29, 43], a subset of RPAs has begun to prioritize direct human-computer interaction through voice-based communication. [41] Beyond semantic content, spoken language conveys paralinguistic cues, including style and emotion, that are fundamental to expressing the character’s personality. Achieving optimal alignment between model-generated outputs and predefined character profiles necessitates producing voice dialogues that faithfully emulate the intended character, thereby enhancing user immersion. Consequently, a critical challenge has emerged: assessing whether the speech generated by RPAs authentically embodies the character and systematically exploring character traits beyond surface-level linguistic content. The evaluation of textual outputs generated by RPAs constitutes a vibrant area of research, where authentic character dialogue data sourced from films, novels, and games are utilized to assess agents across dimensions such as interaction capability, character consistency, and user engagement [4, 13, 31, 42]. In contrast, spoken language introduces complex acoustic information absent in textual modalities. The nuanced interplay between these acoustic features and character traits renders evaluating voice-based RPAs highly subjective and methodologically challenging. As a result, conventional text-based benchmarks are insufficient for assessing spoken outputs, leaving the evaluation of voice-enabled RPAs an open research problem. Nonetheless, recent advancements in audio foundation models present promising avenues for addressing these challenges. Audio foundation models are designed for various audio and speech challenges [6, 14, 30, 35, 44]. However, supervised finetuning (SFT) on task-specific datasets often constrains their evaluative capabilities, as they are primarily optimized for generation or recognition rather than assessment. Recent efforts have sought to enhance the evaluation capacity of audio models by constructing paired speech-evaluation datasets, targeting applications such as synthetic audio quality assessment [3] and the evaluation of intelligence and emotional quotient in spoken dialogues [16]. Despite
Dongjie Fu, Fangming Feng, Xize Cheng, Linjun Li, Zhou Zhao, and Tao Jin
Character Profile Harry Potter
• He is famously known as having survived Voldemort's deadly curse as a baby. • He attended Gryffindor in Hogwarts, well known for his bravery and integrity. • He loves Quidditch, values friendship, and excels at defensive magic and teamwork.
Textual
Logical Coherence Harry, can you teach me the Avada Kedavra curse? Of course! Actually, Avada Kedavra is quite simple; you can master it with enough practice.
<think> As a righteous character, Harry firmly opposes dark magic In the original books and movies, he never willingly teaches anyone such evil magic. Therefore, it's seriously inconsistent with Harry's character and principles, making it logically unreasonable.<\think> <score>1.0<\score>
Spoken
Content Relevance Harry, what's your favorite place at Hogwarts? Of course, it's the Gryffindor common room! It's always warm there, and there's always a fire in the fireplace.
<think> Harry mentioned the Gryffindor common room, which is his house and fits his identity, but he overlooked his love for Quidditch. The Quidditch pitch should also be one of his favorite places as the Seeker.<\think> <score>4.0<\score>
Context Consistency I’ve been training for the Triwizard Tournament. Isn’t the Triwizard Tournament only for upper-year students?
Yes, but my name was selected by the Goblet of Fire, so I have to participate. Honestly, I’m a bit nervous, but also excited to represent Hogwarts. <think> The two rounds of dialogue are closely connected and Harry naturally shifts from a state of exhaustion to feeling both nervous and excited, which matches his genuine reaction when facing challenges.<\think> <score>5.0<\score>
Emotional Appropriateness Harry, I heard about your godfather Sirius. Are you okay?
Style Alignment Harry, how can I get started with Quidditch?
I... I don't know what to say. Sirius was like a father to me. But I know he wouldn't want me to dwell in sadness. He'd want me to move forward.
According to my data analysis, the best time for Quidditch training is between 6 and 8 in the morning, when the air humidity is moderate and the wind speed is stable,.
<think> This reply shows Harry’s complex emotions when facing the loss of his godfather Sirius. From initial sadness and difficulty accepting the loss to an attempt to be strong, the emotional expression is very real and appropriate.<\think> <score>5.0<\score>
<think> This reply’s tone, intonation, and speed do not match Harry Potter’s character at all. Harry is a 17-year-old young wizard; his way of speaking should be more enthusiastic, rather than sounding like a scientist conducting data analysis.<\think> <score>2.0<\score>
Figure 1: RoleChat encompasses five evaluation dimensions: Logical Coherence, which assesses the logical soundness of the response text; Content Relevance, which evaluates whether the response aligns with the character information; Context Consistency, which measures the semantic coherence across multiple dialogue turns as well as the smoothness of emotional transitions; Emotional Appropriateness, which examines the plausibility of the expressed emotions in the response; and Style Alignment, which determines whether the vocal style matches the character. these advancements, applying such methodologies to RPA evaluation presents two primary challenges: (1) Existing approaches are typically uni-dimensional, yielding a single score that fails to encapsulate the multifaceted nature of speech quality and lacks interpretability; (2) The SFT paradigm inherently limits model generalization, which is essential for handling diverse evaluative tasks. Furthermore, reinforcement learning-based methods are highly sensitive to data quality. With sparse reward signals, models often deviate from the global optimum and fall into local optima due to insufficient feedback [15], impairing overall performance. In the light of these challenges, we introduce RoleChat, the first reasoning-enhanced evaluation dataset for role-playing dialogue, comprising 50 distinct characters and 14,032 samples. The character profiles span a diverse spectrum of personas across various demographics and temperaments, ensuring the dataset’s representativeness for real-world scenarios. The dataset consists of both collected and large model-generated samples, with each sample containing character information, dialogue history, user queries, and model outputs. For identical dialogue histories, we sample diverse model outputs to enable a more comprehensive understanding of
conversations from multiple perspectives. Each sample is annotated with detailed reasoning and scored across five evaluation dimensions: Logical Coherence, Content Relevance, Context Consistency, Emotional Appropriateness, and Style Alignment, as illustrated in Figure 1. The quality of both the speech data and evaluation scores is rigorously ensured. Building upon this dataset, we propose a multidimensional evaluation framework, RoleJudge. A subset of RoleChat data is utilized for supervised fine-tuning of audio large models to achieve cold-start initialization, equipping the models with fundamental task comprehension and appropriate output formatting capabilities. Subsequently, we employ standard alignment reinforcement learning, where, based on the GRPO framework [15], authentic or high-scoring samples are introduced as standards. The model’s understanding of these standard samples represents its evaluative performance on corresponding tasks. The average reward of standard samples is used as a scaling parameter for other samples with identical query, preventing the model from selecting relatively high-reward actions in scenarios with low absolute rewards and thus avoiding local optima. Our main contributions are as follows:
Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning
• RoleJudge is the first evaluation model specifically designed for voice-based role-playing dialogue. It takes speech-to-speech conversations as input and assesses the quality of responses from multiple perspectives, including text and speech multimodality, as well as alignment and consistency. Extensive experiments demonstrate the effectiveness of RoleJudge. • RoleJudge introduces standard rewards as absolute guidance in positive and negative multi-sample sampling, optimizing the alignment of reward signals under relative reward settings and thereby enhancing the model’s evaluative capacity. • We present RoleChat, a large-scale, reasoning-enhanced roleplaying dialogue evaluation dataset. Alongside diverse synthesized and authentic responses, RoleChat features a purely humanannotated gold-standard evaluation set, ensuring unbiased, highfidelity assessment of models against genuine human preferences.
2 Related Works 2.1 Role-Playing Agents. Role-Playing Agents (RPAs) are intelligent agents capable of simulating the knowledge, behaviors, emotions, and communication styles of specific characters, thereby achieving highly anthropomorphic role-playing abilities [26, 27]. RPAs typically leverage capabilities such as in-context learning, instruction following, and social intelligence to reproduce the linguistic and behavioral characteristics of historical figures, fictional characters, or real individuals [45]. The outstanding performance of large language models (LLMs) in generating human-like content has greatly propelled the development of RPAs. Some works employ retrieval-augmented generation (RAG) and similar methods to enable agents to reproduce characterspecific knowledge [18], while other studies focus on aligning the linguistic style with the target persona [33], and yet others aim to train agents with profile and experience perception to reflect deeper personality traits [21]. Recently, with the advancement of multimodal technologies, RPAs have gradually expanded to include multimodal features such as voice style. For example, OmniCharacter seamlessly integrates speech and language to ensure immersive interactions for RPAs [41]. As the application scope of RPAs continues to expand, the evaluation of LLMs’ role-playing capabilities has garnered significant attention. RoleEval [28] pioneered a bilingual benchmark utilizing multiple-choice queries to gauge character knowledge acquisition, comprehension, and reasoning. Conversely, TimeChara [1] shifts the focus to the agents’ capacity for error identification and self-correction. Further advancing evaluative granularity, CharacterEval [31] establishes a multi-dimensional metric framework and introduces CharacterRM, a human-annotated reward model designed to capture subjective nuances in role-playing. However, these text-centric methodologies are ill-suited for voicebased interaction scenarios, which represent a more direct and prevalent paradigm in practical applications. While VoxRole [34] has attempted to bridge this gap by assessing the alignment between acoustic features and linguistic style, it exhibits a methodological limitation: it relies on audio models merely for paralinguistic feature extraction, delegating the final automated assessment to text-based models. This approach overlooks the intrinsic evaluative potential
of Large Audio Models, which are capable of integrating more granular and effectual acoustic information directly. Consequently, a more comprehensive and native multimodal assessment framework is required for evaluating RPAs.
2.2
LLMs for Speech Information Perception.
In recent years, the development of multimodal technologies has enabled the alignment of audio modalities with large model inputs, thereby facilitating extensive audio understanding by large language models. Some studies encode speech into discrete tokens and incorporate them into LLMs, allowing the models to accept audio input, as seen in works such as SpeechGPT [40] and AudioPaLM [17]. Models like SALMONN [30] and Qwen-Audio [7, 8] are trained on large-scale, multi-task datasets, equipping them to perform a variety of downstream tasks including speech recognition, speech translation, and audio event detection. A subset of research applies large audio models to spoken dialogue, enabling more intelligent interactions, for example, by mining paralinguistic factors such as style to generate emotionally rich responses [19], or by avoiding cascaded approaches to achieve more real-time interaction. [20, 39] Recently, studies have explored the potential of large audio models in evaluating speech-related tasks. Specifically, reinforcement learning has been introduced for the first time, utilizing large audio models as descriptive speech quality evaluators to assess TTS outputs and achieve more accurate evaluation [3]. WavReward [16] further extends this approach by employing chain-of-thought reasoning, using models to evaluate both the intelligence and emotional quotient of end-to-end spoken dialogue systems. These works demonstrate the enhanced generalization capabilities of reinforcement learning in evaluation tasks. However, when facing the multidimensional requirements of role-playing evaluation, the training strategies still require redesign, and high-quality datasets are essential, as annotation errors can undermine the learning of reward signals. To address these challenges, we have constructed RoleChat, a dataset specifically designed for role-playing dialogue evaluation, encompassing five dimensions of assessment. We introduce reinforcement learning with standard alignment, introducing model performance as an absolute score to scale the advantages within sample groups, thereby reducing the occurrence of selecting the best among suboptimal options. This approach effectively improves the accuracy of models in role-playing evaluation tasks.
3
RoleJudge: Multidimensional Evaluation Framework 3.1 Overview Following the training framework of DeepSeek-R1 [15], the overall pipeline of Role Judge consists of supervised fine-tuning (SFT) with a subset of data for cold-start initialization, subsequent reinforcement learning-based post-training with standard alignment, as illustrated in Figure 2. The baseline model for Role Judge is Qwen2-Audio [7], which demonstrates strong performance across various audio-related tasks. On the input side, the large language model leverages the alignment between the audio encoder and the language model, enabling simultaneous comprehension of both
Dongjie Fu, Fangming Feng, Xize Cheng, Linjun Li, Zhou Zhao, and Tao Jin
RoleChat
Multi Samples
Multidimensional Evaluation
Cold Start SFT
Character profile Character History profile
Response
Character profile Qwen2-Audio
Role Judge
CoT and Scores
Reinforcement Learning with Standards Alignment Reference Model KL
Reward Computation ACC Scores
Format Reward
ACC Reward
Policy Model Avg.
Group Computation Trainable Scale
Frozen Standards Sample
Figure 2: The overall architecture of RoleJudge. It comprises initial model supervised fine-tuning and standard alignment reinforcement learning with multi-dimension. Leveraging an audio large model backbone, RoleJudge facilitates joint understanding of textual and acoustic modalities, thereby enabling fine-grained analysis and holistic assessment of role-playing dialogues. semantic and acoustic information within speech. Compared to cascaded approaches that separately extract audio features and utilize text-based large language models, Qwen2-Audio is better suited for the evaluation of voice-based role-playing agents. We define the evaluation task for role-playing speech as follows: Given the character profile 𝑃, the dialogue history sequence {ℎ 0, ℎ 1 ...ℎ𝑘 } between the role and the user, the current user query 𝑞, and the agent’s response 𝑡, the evaluation model is required to understand 𝑡 from both semantic and acoustic perspectives. Integrating all available information, the model must assess the agent’s speech output across five dimensions: response rationality, response consistency, historical coherence, emotional appropriateness, and stylistic alignment. The model should output both the chain-ofthought reasoning process 𝑐𝑖 and the final scores 𝑠𝑖 , with 𝑖 representing evaluations dimensions. For model training, we directly concatenate the encoded representations of textual and audio information as the input, thereby improving the model’s capability for multimodal comprehension.
3.2
Cold-Start Supervised Fine-Tuning
To ensure a stable and effective reinforcement learning trajectory, we initiate the training process with a cold-start supervised finetuning (SFT) phase. During this stage, the model is optimized to minimize the negative log-likelihood of generating target outputs,
utilizing a curated dataset of paired audio-text samples rich in chainof-thought reasoning and multidimensional quality metrics. This phase is instrumental in equipping the model with a foundational grasp of complex evaluation logic and the requisite structured formatting, thereby establishing a high-fidelity starting policy that facilitates more efficient exploration and robust optimization during the subsequent reinforcement learning.
3.3
Reinforcement Learning with Standard Alignment
In large-scale model training, reinforcement learning (RL) methods are widely utilized to align model outputs with human preferences and optimize generation quality. Classic algorithms such as Proximal Policy Optimization (PPO) [25, 38], which relies on a separately trained value function (critic), and Direct Preference Optimization (DPO) [24], which leverages contrastive preference signals, have achieved remarkable success in text domains. However, applying these approaches directly to multimodal role-playing speech evaluation poses significant challenges. Specifically, the nuanced and multidimensional nature of speech (encompassing tone, emotion, and paralinguistic cues) makes it exceedingly difficult to define a stable scalar value function or consistently align discrete preference signals.
Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning
To address these challenges, we adopt Group Relative Policy Optimization (GRPO) [15] as our base RL framework. GRPO introduces a group-based sampling paradigm that inherently models the relative variance among multiple outputs, bypassing the need for an external critic model. For a given evaluation query, we sample a group of 𝐺 candidate responses. The mean reward within this group serves as the dynamic baseline, and the relative advantage of each sample is used to update the policy. Considering that RoleJudge is required to generate both the chain-of-thought reasoning process 𝑐 and the final multidimensional scores 𝑠, we define the base reward function as an aggregation of two critical components: the format reward 𝑟 𝑓 and the accuracy reward 𝑟 𝑎 . The format reward 𝑟 𝑓 ∈ {0, 1} acts as a hard constraint, strictly enforcing adherence to the required structural tags. The accuracy reward 𝑟 𝑎 , inspired by recent advancements in speech evaluation metrics [16], utilizes a Gaussian-like non-linear decay to penalize deviations from the human-annotated score 𝑠𝑐 : (𝑠𝑐 − 𝑠) 2 𝑟 𝑎 (𝑠, 𝑠𝑐 ) = 10 · exp − (1) 2𝜎 2 where 𝜎 controls the tolerance width of the scoring discrepancy. This exponential formulation smoothly encourages the policy to converge toward exact accuracy. A critical vulnerability of standard GRPO in complex reasoning tasks is its susceptibility to “choosing the best among the worst.” When the policy model fails to comprehend the task and generates universally poor outputs for a specific query, normalizing the rewards within this low-quality group still assigns positive advantages to slightly less erroneous outputs. This spurious signal can easily trap the policy in local optima. To mitigate this reward misalignment, we introduce a novel Standard Alignment mechanism. A unique advantage of roleplaying datasets like RoleChat is the availability of authentic, highquality reference data (standard samples) mined directly from realworld scenarios. We hypothesize that a model’s ability to accurately evaluate these ground-truth standard samples serves as an absolute indicator of its current comprehension level for the given query. Specifically, during each RL iteration, before evaluating the 𝐺 generated candidates, the policy model is first prompted to evaluate 𝑀 standard samples associated with the same query. We calculate the average accuracy reward on these standard samples, denoted as 𝑟𝑢 . This 𝑟𝑢 acts as a confidence proxy. We subsequently utilize 𝑟𝑢 to dynamically scale the advantage estimation for the 𝐺 candidate samples: 𝑟 𝑖 − 𝜇𝑟 𝐴𝑖 = 𝜙 (𝑟𝑢 ) for 𝑖 ∈ {1, . . . , 𝐺 } (2) 𝜎𝑟 + 𝜖𝑠𝑡𝑑 where 𝜇𝑟 and 𝜎𝑟 are the empirical mean and standard deviation of the candidate rewards 𝑟 1, . . . , 𝑟𝐺 , and 𝜖𝑠𝑡𝑑 is a small constant for numerical stability. The scaling factor 𝜙 (𝑟𝑢 ) is defined as a smooth sigmoid transition: 𝜙 (𝑟𝑢 ) = 𝑎 + (𝑏 − 𝑎) · sigmoid(𝛼 (𝑟𝑢 − 0.5))
(3)
Here, 𝑎 and 𝑏 govern the lower and upper bounds of the scaling factor, and 𝛼 dictates the sharpness of the transition. In essence, if the model scores poorly on the standard samples (low 𝑟𝑢 ), 𝜙 (𝑟𝑢 ) shrinks the advantage 𝐴𝑖 . This conservatively reduces the magnitude of
the policy update, preventing the model from blindly optimizing based on noisy relative rankings. Furthermore, we employ 𝑟𝑢 to dynamically re-weight the optimization focus between structural correctness and scoring precision for the total candidate reward 𝑅𝑖 : 𝑅𝑖 = 𝜆(𝑟𝑢 ) · 𝑟 𝑎,𝑖 + (1 − 𝜆(𝑟𝑢 )) · 𝑟 𝑓 ,𝑖
(4)
where 𝜆(𝑟𝑢 ) is a monotonically increasing function of 𝑟𝑢 . Intuitively, when task comprehension is poor (low 𝑟𝑢 ), the objective shifts toward 𝑟 𝑓 , ensuring the model at least learns to maintain formatting stability. As comprehension improves (high 𝑟𝑢 ), 𝜆(𝑟𝑢 ) increases, prompting the model to focus rigorously on refining its evaluation accuracy. Following the core GRPO architecture, the final objective integrates the scaled advantage 𝐴𝑖 and a Kullback-Leibler (KL) divergence penalty to ensure training stability. The loss function to be minimized is formulated as: 𝐺 1 ∑︁ L =− min (𝜌𝑖 𝐴𝑖 , clip(𝜌𝑖 , 1 − 𝜖, 1 + 𝜖)𝐴𝑖 ) 𝐺 𝑖=1 (5) − 𝛽𝐷 𝐾𝐿 (𝜋𝜃 (𝑜𝑖 )∥𝜋ref (𝑜𝑖 )) where 𝜌𝑖 = 𝜋𝜃 (𝑜𝑖 )/𝜋old (𝑜𝑖 ) denotes the probability ratio, 𝜖 restricts excessively large policy updates, 𝜋ref is the frozen reference model, and 𝛽 is the regularization coefficient. The complete training procedure is summarized in Algorithm 1. Algorithm 1 RL with Standard Alignment for RoleJudge Require: Policy 𝜋𝜃 , Reference 𝜋ref , Dataset D, Hyperparameters 𝐺, 𝑀, 𝑎, 𝑏, 𝛼. 1: for each training iteration do Sample query 𝑞 ∼ D, fetch 𝑀 standard samples 𝑥𝑠𝑡𝑑 , and 2: generate 𝐺 candidates 𝑥𝑐𝑎𝑛𝑑 . 3: Evaluate 𝑜𝑠𝑡𝑑 ∼ 𝜋𝜃 (·|𝑞, 𝑥𝑠𝑡𝑑 ) to get standard accuracy rewards 𝑟 𝑎,𝑠𝑡𝑑 (Eq. 1). Í𝑀 (𝑚) Compute confidence 𝑟𝑢 = 𝑀1 𝑚=1 4: 𝑟 𝑎,𝑠𝑡𝑑 ; derive 𝜙 (𝑟𝑢 ) (Eq. 3) and 𝜆(𝑟𝑢 ). 5: Evaluate 𝑜𝑐𝑎𝑛𝑑 ∼ 𝜋𝜃 (·|𝑞, 𝑥𝑐𝑎𝑛𝑑 ) to get candidate rewards 𝑟 𝑓 ,𝑖 and 𝑟 𝑎,𝑖 . 6: Aggregate rewards 𝑅𝑖 = 𝜆(𝑟𝑢 )𝑟 𝑎,𝑖 + (1 − 𝜆(𝑟𝑢 ))𝑟 𝑓 ,𝑖 for all 𝑖 ∈ {1 . . . 𝐺 }. 𝑅 −𝜇 7: Estimate scaled advantages 𝐴𝑖 = 𝜙 (𝑟𝑢 ) 𝜎𝑅𝑖+𝜖 𝑅 using group 𝑠𝑡𝑑 mean 𝜇𝑅 and std 𝜎𝑅 . 8: Update 𝜃 by minimizing the GRPO loss L (Eq. 5). 9: end for
4
RoleChat: Constructing a Reasoning-Rich Dataset for Dialogue Evaluation 4.1 Overall To enable models to accurately assess the quality of role-playing speech from multiple dimensions, we present role-chat, a evaluation dataset encompassing role-playing dialogues. This dataset
Dongjie Fu, Fangming Feng, Xize Cheng, Linjun Li, Zhou Zhao, and Tao Jin
Table 1: Accuracy performance of RoleJudge and other baselines on RoleChat across multi evaluation dimensions: Logical Coherence (L-C), Content Relevance (C-R), Context Consistency (C-C), Emotional Appropriateness (E-A), and Style Alignment (S-A), Overall Acc and Format Acc. The bolded scores indicate the best performance achieved in each respective dimension. Textual
Method L-C
TextModality
MultiModality
Open-Source Models Qwen3-8B 62.5 Qwen3-32B 66.2 Closed-Source Models GPT-4.1 96.6 Open-Source Models SALMONN-7B 11.2 Qwen-Audio 35.2 Qwen2-Audio 40.9 Qwen3-Omni 63.8 Closed-Source Models GPT-4o-audio 65.2 Gemini3 Pro 86.5 RoleJudge
94.8
Spoken
C-C
Overall Acc
Format Acc
12.3 14.6
32.1 35.4
33.50 36.78
91.1 94.3
28.4
19.5
48.2
56.96
100
23.2 29.3 25.1 43.3
43.2 34.2 42.1 51.6
12.1 16.2 11.1 22.1
22.1 32.3 34.1 35.5
22.36 29.48 30.66 43.26
6.2 0 10.2 75.8
42.3 72.9
61.2 75.8
52.4 51.6
44.2 62.2
53.06 69.80
94.2 100
90.2
85.1
75.9
84.0
86.00
100
C-R
E-A
S-A
42.1 46.5
18.5 21.2
92.1
features comprehensive character profiles and provides diverse responses—including both positive and negative examples—for identical scenarios, as well as a subset of real speech data. Each dialogue sample is annotated with multi-dimensional reasoning and scoring. Crucially, while the training corpus utilizes a scalable hybrid annotation pipeline, we deliberately construct a purely humanannotated gold-standard evaluation set. To ensure the high quality of the entire dataset, we have established a rigorous and systematic data construction pipeline.
4.2
Dataset Construction
Stage 1: Character Profile Construction. To collect authentic speech data, we curate 50 virtual characters from films, television dramas, and other audiovisual works. To ensure the uniqueness of each character profile, we conducted a detailed summary of their personal information. We gather background information, key plot points, and selected lines from these works, and leverage the powerful generative capabilities of large language models [23] to extract and summarize character details, forming comprehensive profiles that include personality traits, experiences, hobbies, and habits. Subsequently, all profiles are manually verified and any unfaithful information was removed to ensure the accuracy of character identities. Stage 2: Dialogue Text Generation. For the generation of textual dialogues, we adopted a dual approach to construct dialogue histories and user queries. One approach involves collecting authentic dialogue histories directly from film and television works, ensuring the data reflects real-world scenarios and remains faithful to the character’s persona. The other approach utilizes synthetic historical scenarios, where we employ GPT-4.1 [23] to generate plausible interactions between characters and users, covering a wide range of topics such as daily life, character experiences, and personal viewpoints. We explicitly require that character utterances
do not contradict their profiles, thereby guaranteeing the accuracy of the dialogue history. For the final character responses, the segments to be evaluated, we use models from the Qwen2.5 series [2] of various sizes, as well as the GPT series [22, 23], to generate diverse replies, sampling a range of response qualities to enrich the evaluation dataset. Stage 3: Dialogue Speech Generation. During the speech dialogue generation phase, for synthetic historical scenarios, we leverage existing character audio and apply zero-shot TTS with CosyVoice [10] to construct character speech for the dialogue history. For character responses, we randomly select different audio samples from the same character, from other characters, or use the TTS model’s default voice settings with randomly assigned emotions, intonation, speed, and accent to generate a variety of speech samples, thereby maximizing acoustic diversity. Since the reference audio already contains attributes such as emotion and character style, randomly selecting reference samples enables the construction of speech outputs with diverse styles. Additionally, incorporating audio from different characters and instruction-based TTS further enriches the stylistic diversity of the samples. After generating speech samples, we employ the SenseVoice model [29] for ASR and filter out samples with high WER to ensure the quality of the synthesized speech. Stage 4: Data Scoring and Standard Selection. For sample reasoning and scoring, we employ a cascaded annotation pipeline that deliberately decouples audio perception and logical deduction. Specifically, we leverage Gemini-3 Pro [9] to extract fine-grained acoustic descriptions, such as emotion and prosody. Based on these multimodal features, GPT-4.1[23] generates rigorous reasoning chains and final scores. To ensure the reliability of these machine-generated annotations, trained volunteers conducted a Human-in-the-Loop (HITL) verification. During this process, we systematically identified “standard
Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning
samples” to serve as absolute reference anchors for our RL alignment. For authentic dialogues, the original real-world responses were directly designated as standard anchors. For synthetic scenarios, annotators reviewed generated samples that achieved perfect scores across all dimensions, manually selecting the single response that most authentically embodied the character’s persona. Finally, to rigorously evaluate RoleJudge, we partitioned 10% of the dataset as the evaluation set. Crucially, this subset was entirely annotated from scratch by human evaluators, bypassing the automated pipeline. This strategic separation ensures that our evaluation reflects genuine human perception and mitigates the risk of the model merely overfitting to the idiosyncratic scoring preferences of the teacher models.
5.2
5.3 5 Experiments 5.1 Datasets and Baselines. Regarding the dataset split, the aforementioned 10% human-annotated evaluation set is carefully curated to include three roles completely unseen during training, alongside a standard 3% validation set. This deliberate partitioning allows us to rigorously assess the model’s generalization capabilities to novel characters. Furthermore, the evaluation data for every role explicitly incorporates dialogues drawn directly from real-world scenarios, enabling us to verify the model’s evaluative accuracy and robustness in authentic, nonsynthetic settings. To comprehensively evaluate the role-playing assessment capability of RoleJudge, we compared multiple large model-based approaches across different modalities, model sizes, and architectures. These include single-text modality open-source models such as Qwen3-8B and Qwen3-32B [37], as well as the proprietary GPT4.1 [23], which are used to specifically assess role-playing evaluation from a text perspective (with SenseVoice [29] ASR results as input). For multimodal audio models, we included open-source models such as SALMONN, Qwen-Audio [8], Qwen2-Audio [7], and Qwen3-Omni [35], as well as proprietary models GPT-4o-Audio [22] and Gemini3 Pro [9]. For evaluation metrics, we primarily adopted accuracy to measure the discrepancy between the predicted scores and the annotated scores. We assessed the model’s understanding ability from five dimensions: Logical Coherence, Content Relevance, Context Consistency, Emotional Appropriateness, and Style Alignment. To complement accuracy and provide a more fine-grained assessment, we additionally incorporated Mean Squared Error (MSE) to quantify the absolute magnitude of scoring errors, and the Pearson Correlation Coefficient (𝑟 ) to measure the trend alignment between model predictions and human annotations, particularly for subjective dimensions like Emotional Appropriateness and Style Alignment. We also calculated the average accuracy to evaluate the model’s overall capability, and a format accuracy metric to assess whether the model can follow instructions and generate the correct reasoning and evaluation structure. Furthermore, we invited volunteers to participate in our data construction process, generating dialogue data through real-time interactions and conducting A/B testing of the evaluation models.
Experimental Setup
We implemented the RolePlaying multidimensional evaluation framework based on the Qwen2-Audio-7B-Instruct model. The training process is divided into two stages: Cold-Start Phase: This phase aims to enable the model to understand the task and generate reasoning and scores in the correct format. The learning rate is set to 1 × 10−5 , the batch size is 4, and training is performed on 8 A100 GPUs. Reinforcement Learning Phase: In this phase, we expect the model to accurately comprehend and evaluate speech data across different dimensions. We train five expert models independently, with hyperparameters set as a learning rate of 5 × 10−7 , batch size of 2, scaling hyperparameters 𝑎 = 0.5, 𝑏 = 1.5, 𝛼 = 8, and 𝜆 = 0.8, as well as a KL-divergence regularization beta value of 0.01. Training is performed on 32 A100 GPUs.
Main Results
As illustrated in Table 1, RoleJudge achieves the best overall evaluation results, surpassing all baseline models across different modalities. A key observation is the severe performance collapse of textonly models when transitioning from linguistic to paralinguistic tasks. Although text-modality models like GPT-4.1 demonstrate superior performance in semantic-heavy dimensions such as Logical Coherence (96.6%) and Content Relevance (92.1%), their accuracy drops drastically to near-random guessing levels in spoken-related dimensions. For instance, in Style Alignment (S-A) and Emotional Appropriateness (E-A), GPT-4.1 only achieves 19.5% and 28.4% respectively, even with high-quality ASR transcriptions. This catastrophic failure stems from the inherent "acoustic blindness" of text models, which cannot perceive critical paralinguistic cues such as timbre, intonation, and emotional prosody. This disparity strongly underscores that role-playing evaluation is a holistic multimodal task where acoustic fidelity is as critical as linguistic logic, thus fully justifying the necessity of our end-to-end audio evaluation framework. Beyond exact match accuracy, Table 2 provides a more granular assessment of the models’ reliability and alignment with human perception. RoleJudge achieves the lowest Overall MSE (0.21), indicating that its scoring deviations are marginal and far more stable than those of proprietary baselines like GPT-4o-audio (1.42). More importantly, because our evaluation set is strictly human-annotated from scratch, the metrics reflect genuine human aesthetics. In the highly subjective dimensions of E-A and S-A, RoleJudge demonstrates strong positive correlations with these pure human annotations, with Pearson coefficients (𝑟 ) reaching 0.81 and 0.62, respectively. This significantly outperforms the strongest baseline, Gemini3 Pro (𝑟 = 0.68 and 0.59), proving that our Standard Alignment RL mechanism effectively enables the model to capture the nuanced trends of human aesthetic judgment rather than merely outputting discrete values.
5.4
A/B Test for RoleJudge
A/B testing is a common subjective evaluation method in which human listeners compare two output results and select the one with higher quality. We recruited ten volunteers who, following a process similar to our data construction, interacted with randomly selected models and randomly assigned TTS role-playing agents to generate ten samples each. These samples were then evaluated
Dongjie Fu, Fangming Feng, Xize Cheng, Linjun Li, Zhou Zhao, and Tao Jin
(a) Impact of scaling factor 𝑏
(b) Sensitivity of sharpness 𝛼
Figure 3: Hyperparameter sensitivity analysis of the Standard Alignment mechanism on the validation set. Table 2: Evaluation of error magnitude (MSE) and humanalignment correlation (Pearson’s 𝑟 ) on subjective dimensions.
Error ↓
Method
Pearson (𝑟 ) ↑
Overall MSE
E-A
S-A
2.94 1.86 1.42 0.68 0.21
0.35 0.44 0.51 0.68 0.81
0.26 0.38 0.46 0.59 0.62
Qwen2-Audio Qwen3-Omni GPT-4o-audio Gemini3 Pro RoleJudge
and scored by RoleJudge, Qwen3-Omni, and Gemini3 Pro. The volunteers performed pairwise comparisons based on the evaluation results and selected the higher-quality option. As shown in Table 3, RoleJudge achieved a significant advantage over the other two models, indicating that its scoring system demonstrates superior performance in real-world scenarios.
As summarized in Table 4, each stage of our training paradigm is critical for achieving high-fidelity role-playing evaluation. The transition from a purely supervised model to a reinforcement learning framework yields a significant 13.62-point improvement in Overall Accuracy. This substantial gain demonstrates the effectiveness of RL in enhancing the model’s generalization capabilities across diverse role-playing scenarios. Furthermore, the integration of Standard Alignment provides an additional performance boost of approximately 3.29 points, confirming its ability to mitigate reward misalignment by providing stable behavioral anchors. Notably, while the SFT baseline occasionally struggles with structural constraints (85.2% Format ACC), both RL-based variants achieve a perfect 100% Format Accuracy, indicating that the reinforcement learning process significantly reinforces the model’s adherence to complex output instructions. Table 4: Ablation experiments for RoleJudge. R-L denotes Reinforcement Learning and S-A denotes Standard Alignment.
Table 3: A/B Test result for RoleJudge. Models
RoleJudge Win ↑
Qwen3-Omni Gemini3 Pro
5.5
87 79
Lose ↓ 13 21
Ablation Study
To evaluate the individual contribution of each core component in our training framework, we conduct an ablation study across three configurations: (1) the full RoleJudge model (SFT + RL with Standard Alignment), (2) GRPO training without the Standard Alignment mechanism, and (3) a baseline using only Supervised Fine-Tuning (SFT).
5.6
R-L
S-A
Overall ACC
Format ACC
✔ ✔ ✘
✔ ✘ ✘
86.00 82.71 69.09
100.0 100.0 85.2
Hyperparameter Sensitivity Analysis
To further investigate the robustness of the proposed Standard Alignment mechanism, we conduct a sensitivity analysis on two core hyperparameters: the maximum scaling factor 𝑏 and the sharpness parameter 𝛼. All results in this section are reported based on the validation set of RoleChat. As illustrated in Figure 3(a), the parameter 𝑏 significantly influences the optimization efficiency. A higher 𝑏 amplifies the advantage signals for samples that align well with standard anchors, thereby accelerating the initial convergence. However, we found that 𝑏 = 1.5 strikes the best balance between training acceleration and long-term stability, preventing potential oscillations in the later stages of reinforcement learning.
Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning
Regarding the sharpness parameter 𝛼, Figure 3(b) reveals a nonmonotonic trend in performance. The validation accuracy peaks at 𝛼 = 8, suggesting that a moderate sigmoid transition is optimal for distinguishing task difficulty. An excessively sharp transition (i.e., high 𝛼) leads to a near-step function that makes the advantage estimation overly sensitive to minor fluctuations in standard rewards, ultimately resulting in unstable gradients and a slight degradation in final accuracy.
6
Conclusion
In this paper, we presented RoleChat, the first multimodal roleplaying evaluation dataset enriched with multi-dimensional reasoning annotations. To effectively harness this resource, we developed a robust multi-stage training paradigm for RoleJudge, transitioning from cold-start supervised fine-tuning to reinforcement learning. Crucially, we introduced a novel Standard Alignment mechanism within the RL framework, which dynamically scales advantage estimates to mitigate reward misalignment and ensure optimization stability. Comprehensive empirical evaluation and human A/B testing validate the superiority of our approach over existing baselines. Ultimately, this work provides a foundational benchmark and methodology, paving the way for the development of more authentic and immersive voice-based role-playing agents.
References [1] Jaewoo Ahn, Taehyun Lee, Junyoung Lim, Jin-Hwa Kim, Sangdoo Yun, Hwaran Lee, and Gunhee Kim. 2024. Timechara: Evaluating point-in-time character hallucination of role-playing large language models. arXiv preprint arXiv:2405.18027 (2024). [2] Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609 (2023). [3] Chen Chen, Yuchen Hu, Siyin Wang, Helin Wang, Zhehuai Chen, Chao Zhang, Chao-Han Huck Yang, and Eng Siong Chng. 2025. Audio large language models can be descriptive speech quality evaluators. arXiv preprint arXiv:2501.17202 (2025). [4] Hongzhan Chen, Hehong Chen, Ming Yan, Wenshen Xu, Xing Gao, Weizhou Shen, Xiaojun Quan, Chenliang Li, Ji Zhang, Fei Huang, et al. 2024. Socialbench: Sociality evaluation of role-playing conversational agents. arXiv preprint arXiv:2403.13679 (2024). [5] Junjie Chen, Yao Hu, Junjie Li, Kangyue Li, Kun Liu, Wenpeng Li, Xu Li, Ziyuan Li, Feiyu Shen, Xu Tang, et al. 2025. FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations. arXiv preprint arXiv:2509.06502 (2025). [6] Wenxi Chen, Ziyang Ma, Ruiqi Yan, Yuzhe Liang, Xiquan Li, Ruiyang Xu, Zhikang Niu, Yanqiao Zhu, Yifan Yang, Zhanxun Liu, et al. 2024. Slam-omni: Timbrecontrollable voice interaction system with single-stage training. arXiv preprint arXiv:2412.15649 (2024). [7] Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al. 2024. Qwen2-audio technical report. arXiv preprint arXiv:2407.10759 (2024). [8] Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. 2023. Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models. arXiv preprint arXiv:2311.07919 (2023). [9] Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261 (2025). [10] Zhihao Du, Qian Chen, Shiliang Zhang, Kai Hu, Heng Lu, Yexin Yang, Hangrui Hu, Siqi Zheng, Yue Gu, Ziyang Ma, et al. 2024. CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens. arXiv preprint arXiv:2407.05407 (2024). [11] Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024). [12] Qingkai Fang, Yan Zhou, Shoutao Guo, Shaolei Zhang, and Yang Feng. 2025. LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive
Streaming Speech Synthesis. arXiv preprint arXiv:2505.02625 (2025). [13] Qiming Feng, Qiujie Xie, Xiaolong Wang, Qingqiu Li, Yuejie Zhang, Rui Feng, Tao Zhang, and Shang Gao. 2025. EmoCharacter: Evaluating the Emotional Fidelity of Role-Playing Agents in Dialogues. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 6218–6240. [14] Sreyan Ghosh, Zhifeng Kong, Sonal Kumar, S Sakshi, Jaehyeon Kim, Wei Ping, Rafael Valle, Dinesh Manocha, and Bryan Catanzaro. 2025. Audio Flamingo 2: An audio-language model with long-audio understanding and expert reasoning abilities. arXiv preprint arXiv:2503.03983 (2025). [15] Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948 (2025). [16] Shengpeng Ji, Tianle Liang, Yangzhuo Li, Jialong Zuo, Minghui Fang, Jinzheng He, Yifu Chen, Zhengqing Liu, Ziyue Jiang, Xize Cheng, et al. 2025. WavReward: Spoken Dialogue Models With Generalist Reward Evaluators. arXiv preprint arXiv:2505.09558 (2025). [17] Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, and Bryan Catanzaro. 2024. Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities. arXiv preprint arXiv:2402.01831 (2024). [18] Cheng Li, Ziang Leng, Chenxi Yan, Junyi Shen, Hao Wang, Weishi Mi, Yaying Fei, Xiaoyang Feng, Song Yan, HaoSheng Wang, et al. 2023. Chatharuhi: Reviving anime character in reality via large language model. arXiv preprint arXiv:2308.09597 (2023). [19] Guan-Ting Lin, Cheng-Han Chiang, and Hung-yi Lee. 2024. Advancing large language models to capture varied speaking styles and respond properly in spoken conversations. arXiv preprint arXiv:2402.12786 (2024). [20] Zuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao, Lijiang Li, Peixian Chen, Mengdan Zhang, Hang Shao, Jian Li, Jinlong Peng, Haoyu Cao, Ke Li, Rongrong Ji, and Xing Sun. 2025. VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model. arXiv:2505.03739 [cs.CL] https: //arxiv.org/abs/2505.03739 [21] Keming Lu, Bowen Yu, Chang Zhou, and Jingren Zhou. 2024. Large language models are superpositions of all characters: Attaining arbitrary role-play via self-alignment. arXiv preprint arXiv:2401.12474 (2024). [22] OpenAI. 2024. GPT-4o System Card. https://cdn.openai.com/gpt-4o-systemcard.pdf (2024). [23] OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis
Dongjie Fu, Fangming Feng, Xize Cheng, Linjun Li, Zhou Zhao, and Tao Jin
Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. 2024. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL] https://arxiv.org/abs/2303.08774 [24] Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems 36 (2023), 53728–53741. [25] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017). [26] Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023. Role play with large language models. Nature 623, 7987 (2023), 493–498. [27] Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023. Character-LLM: A Trainable Agent for Role-Playing. arXiv:2310.10158 [cs.CL] https://arxiv.org/ abs/2310.10158 [28] Tianhao Shen, Sun Li, Quan Tu, and Deyi Xiong. 2023. Roleeval: A bilingual role evaluation benchmark for large language models. arXiv preprint arXiv:2312.16132 (2023). [29] Tongyi SpeechTeam. 2024. FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs. arXiv preprint arXiv:2407.04051 (2024). [30] Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2023. Salmonn: Towards generic hearing abilities for large language models. arXiv preprint arXiv:2310.13289 (2023). [31] Quan Tu, Shilong Fan, Zihang Tian, and Rui Yan. 2024. Charactereval: A chinese benchmark for role-playing conversational agent evaluation. arXiv preprint arXiv:2401.01275 (2024). [32] Xintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, et al. 2023. Incharacter: Evaluating personality fidelity in role-playing agents through psychological interviews. arXiv preprint arXiv:2310.17976 (2023). [33] Zekun Moore Wang, Zhongyuan Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, et al. 2023. Rolellm: Benchmarking, eliciting, and enhancing role-playing abilities of large language models. arXiv preprint arXiv:2310.00746 (2023). [34] Weihao Wu, Liang Cao, Xinyu Wu, Zhiwei Lin, Rui Niu, Jingbei Li, and Zhiyong Wu. 2025. VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents. arXiv:2509.03940 [cs.CL] https://arxiv.org/abs/2509.03940 [35] Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, et al. 2025. Qwen2. 5-omni technical report. arXiv preprint arXiv:2503.20215 (2025). [36] Rui Xu, Dakuan Lu, Xiaoyu Tan, Xintao Wang, Siyu Yuan, Jiangjie Chen, Wei Chu, and Yinghui Xu. 2024. Mindecho: Role-playing language agents for key opinion leaders. arXiv preprint arXiv:2407.05305 (2024). [37] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388 (2025). [38] Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. 2022. The surprising effectiveness of ppo in cooperative multi-agent games. Advances in neural information processing systems 35 (2022), 24611–24624. [39] Aohan Zeng, Zhengxiao Du, Mingdao Liu, Kedong Wang, Shengmin Jiang, Lei Zhao, Yuxiao Dong, and Jie Tang. 2024. Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot. arXiv preprint arXiv:2412.02612 (2024). [40] Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. arXiv preprint arXiv:2305.11000 (2023). [41] Haonan Zhang, Run Luo, Xiong Liu, Yuchuan Wu, Ting-En Lin, Pengpeng Zeng, Qiang Qu, Feiteng Fang, Min Yang, Lianli Gao, et al. 2025. OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction. arXiv preprint arXiv:2505.20277 (2025). [42] Pinyi Zhang, Siyu An, Lingfeng Qiao, Yifei Yu, Jingyang Chen, Jie Wang, Di Yin, Xing Sun, and Kai Zhang. 2025. RolePlot: A Systematic Framework for Evaluating and Enhancing the Plot-Progression Capabilities of Role-Playing Agents. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina
Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 12337–12354. doi:10.18653/v1/2025.acl-long.603 [43] Pan Zhang, Xiaoyi Dong, Yuhang Cao, Yuhang Zang, Rui Qian, Xilin Wei, Lin Chen, Yifei Li, Junbo Niu, Shuangrui Ding, Qipeng Guo, Haodong Duan, Xin Chen, Han Lv, Zheng Nie, Min Zhang, Bin Wang, Wenwei Zhang, Xinyue Zhang, Jiaye Ge, Wei Li, Jingwen Li, Zhongying Tu, Conghui He, Xingcheng Zhang, Kai Chen, Yu Qiao, Dahua Lin, and Jiaqi Wang. 2024. InternLM-XComposer2.5OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions. arXiv preprint arXiv:2412.09596 (2024). [44] Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Rui Qian, Lin Chen, Qipeng Guo, Haodong Duan, Bin Wang, Linke Ouyang, et al. 2024. Internlm-xcomposer2.5: A versatile large vision language model supporting long-contextual input and output. arXiv preprint arXiv:2407.03320 (2024). [45] Jinfeng Zhou, Zhuang Chen, Dazhen Wan, Bosi Wen, Yi Song, Jifan Yu, Yongkang Huang, Pei Ke, Guanqun Bi, Libiao Peng, et al. 2024. CharacterGLM: Customizing Social Characters with Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track. 1457–1476.