JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
1
Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents
arXiv:2605.22602v1 [cs.AI] 21 May 2026
Minghui Ma, Bin Guo*, Senior Member, IEEE, Runze Yang, Mengqi Chen, Yan Liu, Jingqi Liu, Yahan Pei, Xuehao Ma, Qiuyun Zhang, Zhiwen Yu, Senior Member, IEEE
Abstract—Persuasive dialogue requires reasoning about others’ latent mental states, a capability known as Theory of Mind (ToM). However, due to reliance on simple prompting strategies and insufficient ToM knowledge, existing LLMs often fail to capture the intrinsic dependencies among mental states, leading to fragmented representations and unstable reasoning. To address these challenges, we introduce the ToM-based Persuasive Dialogue (ToM-PD) task, grounded in the Belief–Desire–Intention (BDI) framework, which explicitly models the sequential dependencies among mental states in multi-turn dialogues. To facilitate research on this task, we construct a large-scale annotated dataset, ToM-based Broad Persuasive Dialogues (ToM-BPD), capturing fine-grained mental states and corresponding persuasive strategies. We further propose Think Thrice Before You Speak (TTBYS), a knowledge-enhanced stepwise reasoning framework that leverages both explicit and implicit prior experiences to improve LLMs’ inference of desires, beliefs, and strategies. Experimental results demonstrate that Qwen3-8B equipped with TTBYS outperforms GPT-5 by 1.20%, 22.80%, and 16.97% in predicting desires, beliefs, and persuasive strategies, respectively. Case studies further show that our approach enhances interpretability and consistency in reasoning. Index Terms—Natural language processing, Knowledge retrieval, Psychology, Human-centered computing
I. I NTRODUCTION
P
ERSUASION is a fundamental component of human social interaction [1], [2]. With the rapid advancement of LLMs, LLM-based persuasive dialogue systems have gained increasing influence across diverse domains, such as emotional support [3], [4], charitable donation [5], [6], and negotiation [7]. These systems not only need to generate fluent and coherent responses but also accurately infer and strategically influence users’ latent mental states [8]. Theory of Mind, the ability to infer others’ beliefs, desires, and intentions, is widely recognized as a core cognitive mechanism underlying social reasoning and strategic interaction [9], [10]. Motivated by this insight, recent studies have explored This work was partially supported by the National Natural Science Foundation of China (Nos. U25B2042, 62532009, 62502392). Minghui Ma, Bin Guo*, Runze Yang, Mengqi Chen, Yan Liu, Jingqi Liu, Yahan Pei, and Xuehao Ma are with Northwestern Polytechnical University, Xi’an, China (e-mail: {2021300120, yangrunze, chenmengqi, 1309893591, peiyahan, xhma}@mail.nwpu.edu.cn; [email protected]; [email protected]). Qiuyun Zhang is with Peking University, Beijing, China (e-mail: [email protected]). Zhiwen Yu is with Harbin Engineering University, Harbin, China, and Northwestern Polytechnical University, Xi’an, China (email: [email protected]). *Corresponding author: Bin Guo (e-mail: [email protected]).
the integration of ToM modeling into persuasive dialogue systems. For instance, prior work decomposes persuasion tasks into multiple collaborating agents, with dedicated components for identifying users’ emotional states [11]. Other approaches simulate potential user rebuttals and attitudes to improve persuasive effectiveness [8]. In addition, chain-of-thought reasoning has been leveraged to model users’ mental states and enhance the empathy of system responses [4], [12]. Although effective on benchmark tasks, these approaches often treat different mental states as independent variables, neglecting their intrinsic dependencies, which can lead to inconsistent or unstable inferences in complex dialogues. Several benchmarks have also evaluated LLMs’ ToM reasoning capabilities in persuasive contexts. NegotiationToM [13] focuses on negotiation scenarios, ToMATO [14] targets donation requests, and BigToM [15] evaluates action prediction in narrative contexts. PersuasiveToM [16] further examines LLMs’ ability to reason about and predict persuasive strategies based on users’ mental states. However, empirical results consistently indicate that LLMs lack sufficient prior knowledge for reliable mental-state inference, limiting the stability and robustness of their reasoning. In summary, applying ToM to persuasive dialogue systems and improving their performance still faces two major challenges: (1) how to systematically model and infer users’ ToM states and their intrinsic dependencies in persuasive dialogues, and (2) how to equip LLMs with sufficient knowledge to enhance their ToM reasoning capabilities. As shown in Figure 1, the user’s internal BDI states and corresponding actions follow a forward evolutionary process, in which mental states evolve sequentially from beliefs to desires, intentions, and ultimately actions [17]. In contrast, we model persuasive interaction as an inference problem centered on first-order ToM. The agent can only observe the persuadee’s external utterances and must iteratively infer their intentions, further reasoning about the underlying desires and beliefs, in order to dynamically adjust persuasive strategies. Motivated by this observation and grounded in the BDI framework [17], we formally define the ToM-based Persuasive Dialogue (ToMPD) task, and formulate its solution as a stepwise backward inference process over latent mental states. Humans often rely on prior experiences in analogous scenarios to guide reasoning when confronting ToM problems [18]. In ToM-PD, this manifests as a stepwise, deliberate reasoning process, akin to the Chinese proverb, “Think Thrice Before You
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
2
Fig. 1. Illustration of self BDI state evolution and BDI-based inference for ToM-driven persuasive dialogue (ToM-PD). The left panel shows the internal reasoning process, where Lucian generates actions through the evolution of its belief, desire, and intention states based on self-perception and experience. The middle panel presents a multi-turn persuasive dialogue scenario between the agent and the user. The right panel depicts the ToM-PD process, where the agent observes user actions (utterances), infers the user’s latent BDI states, and dynamically selects appropriate persuasive strategies to guide subsequent actions.
Act.”. Drawing inspiration from this, we propose a knowledgeenhanced stepwise reasoning framework, Think Thrice Before You Speak (TTBYS), which integrates both explicit and implicit knowledge to guide LLMs through sequential reasoning over desires, beliefs, and persuasive strategies, thereby improving robustness and consistency in ToM reasoning. To evaluate our approach, we further construct a multi-domain dataset, ToM-based Broad Persuasive Dialogues (ToMBPD), annotated with fine-grained ToM states of the persuadee and corresponding persuasive strategies of the persuader for each dialogue turn, with rigorous quality control to ensure annotation reliability. In summary, our contributions are threefold: (1) We introduce the ToM-PD task, explicitly modeling dependencies among different mental states to enhance the interpretability. (2) We propose TTBYS, a knowledge-enhanced stepwise reasoning framework to enhance ToM inference for LLM in persuasive dialogues. (3) We construct a ToM-oriented persuasive dialogue dataset, ToM-BPD, and validate the effectiveness and interpretability of our approach through extensive experiments and case studies. The rest of the paper is organized as follows. Section II reviews related work on LLM-based persuasive dialogue agents and Theory of Mind in AI. Section III formalizes the ToMPD task and introduces our approach for sequential mental state inference in persuasive interactions. Section V presents the TTBYS framework, detailing experience representation, retrieval mechanisms, and multi-step reasoning for desire, belief, and strategy inference. Section VI presents experimental results and analysis. Section VII draws conclusions, and Section VIII provides an ethics statement. II. R ELATED W ORK A. LLM-based Persuasive Dialogue Agents Persuasive dialogue aims to influence individuals’ beliefs, attitudes, or behaviors through targeted communication strategies [19]. Early studies mainly focused on domain-specific
scenarios such as emotional support [3], policy persuasion [20], and charitable fundraising [21]. With the rise of large language models, recent work has extended persuasive dialogue to more diverse and complex settings, including multiturn recommendation [22], adversarial prompting [23], and misleading or manipulative behaviors [24], [25]. From a technical perspective, existing approaches can be broadly categorized into three lines. Strategy-based methods guide generation with predefined or learned dialogue strategies, improving controllability and interpretability [3], [4], [12], [21], [26]. Knowledge-enhanced methods incorporate external knowledge or structured memory to improve contextual understanding and response quality [27]–[30]. Multi-agent systems decompose dialogue into multiple coordinated agents, enabling more complex reasoning and collaborative decisionmaking [8], [11], [31]. However, most existing approaches lack explicit modeling of users’ mental states, limiting their interpretability and longterm effectiveness in persuasive interactions. In contrast, our framework explicitly models the dependencies among these mental states in a stepwise manner, enabling more robust and interpretable inference to guide persuasive strategy selection. B. Theory of Mind in AI Theory of Mind refers to the ability to infer one’s own and others’ mental states. In the era of large language models, a variety of benchmarks have been proposed to evaluate ToM capabilities across diverse scenarios, including embodied environments, negotiation settings, narrative understanding, and persuasive interactions [13]–[16], [32]. Existing evidence suggests that current LLMs still exhibit limited stability and reliability in mental-state inference, partly due to insufficient prior ToM knowledge encoded in these models. To enhance ToM reasoning in LLMs, existing approaches can be broadly categorized into three lines. Prompting-based methods improve ToM through inference-time reasoning scaffolds, such as perspective-taking prompts, role conditioning, and hypothesis-driven reasoning [33]–[36]. Structured and neuro-symbolic approaches explicitly model belief and goal
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
3
inference by integrating symbolic representations or probabilistic reasoning frameworks [37], [38]. Multi-agent cognitive frameworks decompose ToM reasoning into multiple interacting components to simulate social cognition processes [39]. Despite these advances, most existing methods primarily focus on reasoning over abstract or static states, and remain insufficient for fine-grained modeling of users’ dynamic mental states in real-world social interaction scenarios. Our approach addresses these limitations by combining stepwise backward inference over BDI states with knowledge-enhanced reasoning, enabling more accurate, consistent, and interpretable ToM inference in multi-turn persuasive dialogues. III. T O M BASED P ERSUASIVE D IALOGUE In a persuasive task, the persuader aims to influence the user’s attitude toward a persuasion target through appropriate persuasive strategies, ultimately inducing behavioral change. Achieving this goal requires the persuader to infer the persuadee’s latent mental states from the dialogue history and use these inferences to guide subsequent decisions. Accordingly, we incorporate the BDI model into persuasive dialogue and define the ToM-PD task. This section introduces: (1) the formulation of the ToMPD task, encompassing mental state inference and strategy prediction (Section III-A); and (2) a reverse and stepwise mental state inference procedure that treats human utterances as observable actions and sequentially infers intention, desire, belief and strategies (Section III-B).
where fintention , fdesire , and fbelief denote the inference functions for intention, desire, and belief, respectively. Finally, the persuader selects the next strategy st based on the inferred mental state: t st = fstrategy (SToM ),
(4)
where fstrategy denotes the selection function for strategy. IV. T O M-BPD DATASET PersuasiveToM is a multi-turn dialogue benchmark that spans diverse persuasive scenarios, designed to systematically evaluate the Theory-of-Mind reasoning capabilities of large language models in persuasive interactions [16]. Building upon this benchmark, we further construct the ToM-BPD dataset, which provides fine-grained annotations of the persuadee’s desire and belief states, as well as the strategies employed by the persuader, thereby capturing the key cognitive components underlying persuasive dialogue. To ensure data quality and annotation consistency, we develop a comprehensive data construction pipeline, comprising detailed annotation guidelines, annotator qualification tests, and multiple rounds of automated and human verification. This section introduces: (1) the annotation framework of ToM-BPD, including the annotation schema and annotation procedure (Section IV-A); and (2) the quality control process, detailing the specific measures adopted as well as the specific content of the designed tutorial (Section IV-B). A. Annotation
A. Task Formulation The ToM-PD task aims to generate strategy-guided responses that effectively influence the persuadee’s mental state based on the dialogue history. At dialogue turn t, the history is defined as ht = {u1 , a1 , . . . , ut }, where uk is the persuadee’s utterance and ak is the persuader’s response generated under strategy sk . t Given both ht and the inferred mental state SToM = (it , dt , bt ), representing intention, desire, and belief, the persuader predicts the next persuasive strategy st . The task is considered successful at turn t if the persuadee’s desire dt+1 shifts toward the persuasion target.
1) Annotation Schema: The primary annotation fields include the persuader’s strategy, as well as the persuadee’s desire and belief at each dialogue turn. Strategy Construction. Based on the Elaboration Likelihood Model (ELM) [1] and Motivational Interviewing [40], we categorize persuasive strategies into three groups: socio-emotional strategies that foster empathy and strengthen rapport, cognitive strategies that influence reasoning through arguments and examples, and interactive strategies that facilitate dialogue and reduce resistance. Following prior work [41], we further refine these categories into nine fine-grained techniques. The taxonomy is summarized in Table I, and detailed definitions with examples are provided in Table II.
B. Reverse Mental State Inference In traditional self BDI state evolution process, psychological states are modeled through a forward process: beliefs give rise to desires, desires form intentions, and intentions lead to actions (Figure 1, left). In persuasive dialogue, the persuadee’s utterances are observable, requiring backward inference to recover underlying mental states (Figure 1, right). As shown in figure 1 (right), we propose a reverse mental state inference procedure for ToM-PD. Starting from the dialogue history ht , the persuadee’s mental states are inferred sequentially as: it = fintention (ht ),
(1)
dt = fdesire (it ),
(2)
bt = fbelief (it , dt ),
(3)
TABLE I TAXONOMY OF PERSUASIVE STRATEGIES . Category Socio-emotional
Cognitive
Interactive
Technique Affirmation and Reassurance Reflection of Feelings Personal Story Expression of Views Enhancement of Views Logical Appeal Giving Examples Supplying Information Task Inquiry
Desire. In persuasive dialogues, desire often reflects the user’s attitude toward the persuasion target (e.g., willingness to
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
4
TABLE II T HE STRATEGIES AND THEIR CORRESPONDING DEFINITIONS IN THE T O M-BPD DATASET Strategy Affirmation and reassurance
Reflection of feelings
Personal story
Expression of view
Enhancement of view
Logical appeal Supplying information Giving example Task inquiry
Definition refers to the strategy where the persuader validates the persuadee’s feelings, acknowledges difficulties, or encourages their sense of capability. refers to the strategy where the persuader reflects, paraphrases, or interprets the persuadee’s emotional state to show understanding. refers to the strategy where the persuader shares a personal experience or anecdote to enhance emotional resonance. refers to the strategy where the persuader expresses a personal standpoint, belief, or evaluation without necessarily providing reasoning. refers to the strategy where the persuader strengthens or intensifies a previously expressed stance through emphasis or elaboration. refers to the strategy where the persuader uses reasoning, cause-effect logic, or explicit argument structure. refers to the strategy where the persuader provides factual information, general knowledge, or relevant advice. refers to the strategy where a concrete instance or case is provided to support a point. refers to the strategy where the persuader asks an open-ended or exploratory question to understand the persuadee’s concerns or motivations.
engage in an activity such as outdoor exercise). Previous work has represented user attitudes with discrete values [42]– [45]. Following [45], we operationalize desire as a discrete variable with three values, −1, 0, 1, corresponding to negative, hesitant, and positive attitudes toward the persuasion target. This representation indicates whether the user is currently inclined to reject, hesitate, or accept the persuasion goal, which is further used to determine the success of the persuasion task. Belief. Beliefs are represented as short declarative statements reflecting the persuadee’s views, preferences, or comparisons related to the persuasion target. They are categorized as positive (e.g., perceiving the target as interesting) or negative (e.g., concerns about potential risks). A single dialogue turn may contain one or multiple such statements. 2) Annotation Procedure: We adopt a semi-automatic annotation framework that combines automated pre-annotation with multi-round human verification. Preprocessing. Prior to annotation, we preprocess the dataset by removing dialogues involving more than two participants. Automatic Pre-annotation. To improve annotation consistency and reduce human workload, we employ LLMs to perform automatic pre-annotation in the first annotation round. The prompts used for annotation are provided in Section A. Annotation Correction. Human annotators review and verify the automatic annotations. During this stage, annotators are instructed to flag ambiguous or difficult-to-judge dialogue segments, as well as samples with inconsistent labels within the same dialogue. After each annotation round, the involved annotators are organized to discuss the flagged samples and reach consensus by articulating their judgments; samples for which consensus cannot be reached are discarded. In the end, we remove 21 dialogues from the original 525 dialogues in the PersuasiveToM dataset.
Example “I understand this is difficult, but I believe you can handle it.” “It sounds like you’re feeling overwhelmed right now.”
“I had a similar experience when I faced this situation last year.” “I think this option is better for you.”
“This is definitely the best choice for you, especially considering your goals.” “If you choose this option, you will therefore save both time and money.” “This program typically takes about six months to complete.” “For example, many students improved their performance using this method.” “What concerns you most about this option?”
B. Quality Control To ensure annotation consistency and data quality, we adopt the following measures: (1) we design systematic annotation tutorials for annotators; (2) we design qualification tests for annotators; (3) we integrate automated annotation with human annotation and implement a multi-stage quality control process(Section IV-A2). Tutorial. We design a structured training program for volunteer annotators. The training materials cover fundamental theoretical knowledge of ToM, following the definitions and formalizations in [16], as well as foundational theories of strategy-based persuasive dialogue based on [41]. In addition, the tutorial introduces core annotation principles for ToM mental states, and provides formal definitions of persuasive strategies (Table II) together with illustrative examples. Annotation Principles for ToM States. Let n denote the dialogue turn number. For the first turn (n = 1), annotation is performed at the sentence level. If the persuadee expresses appreciation for the persuasive target, such as “This sounds interesting” or “This seems promising,” the sentence is labeled as a positive sentence. Conversely, if the persuadee expresses concerns or uncertainty (e.g., “I am not sure...”), or explicitly supports an event contrary to the persuasive goal, the sentence is labeled as a negative sentence. Neutral sentences are ignored when determining the desire label. At the beginning of the dialogue, if the persuadee’s current utterance contains only positive sentences, the desire is labeled as 1; if it contains both positive and negative sentences, it is labeled as 0; and if it contains only negative sentences, it is labeled as −1. The extracted positive and negative sentences are recorded as natural language descriptions of beliefs, for example: (Positive: attending the activity is interesting; Negative: uncertain about the benefit of the activity). For subsequent turns (n > 1), the negative sentences
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
from the previous turn are examined. A negative sentence is considered resolved if it is explicitly addressed or mitigated in the subsequent dialogue. If all negative sentences are resolved, annotation proceeds following the same principles as in the first turn. Otherwise, unresolved negative sentences are retained, and the current turn is annotated by considering both newly expressed positive and negative sentences together with the unresolved ones. The resulting positive and negative sentences are similarly recorded as belief descriptions. Examination. We designed a qualification test comprising 25 labeled dialogue turns to evaluate volunteers’ understanding of the annotation concepts and their ability to correctly apply the labeling standards. Specifically, volunteers were required to annotate, based on the current dialogue history and its progression, the persuader’s strategy as well as the persuadee’s desire and belief. We then calculated the average annotation accuracy across these three dimensions for each annotator. Annotators achieving an accuracy above 90% were considered qualified for the annotation task. Ultimately, this test resulted in the disqualification of 8 out of 23 participants. C. Data Characteristics Overall Statistics. Table III reports the statistics of ToMBPD. The dataset contains 504 dialogues and 3,926 utterances, with an average of 7.79 turns per dialogue. The average utterance length differs across roles, with 38.62 tokens for the persuader and 19.67 tokens for the persuadee. Each turn contains 1.89 belief annotations on average, suggesting that persuadee responses are often driven by multiple underlying beliefs. The average desire score is -0.17, indicating that most dialogues begin with a certain degree of resistance toward the persuasion target. TABLE III OVERALL STATISTICS OF T O M-BPD.
Statistic Total dialogues Total utterances Avg. dialogue length (turns) Avg. utterance length (persuader) Avg. utterance length (persuadee) Avg. desire Avg. belief number per turn
ToM-BPD 504 3926 7.79 38.62 19.67 -0.17 1.89
Percentages of Strategies. Table IV presents the distribution of persuasive strategies in ToM-BPD. Cognitive strategies dominate the dataset (60.84%), followed by socio-emotional strategies (34.36%), while interactive strategies account for a relatively small proportion (4.79%). In conjunction with the functional definitions in Table II, this distribution suggests that persuaders in ToM-BPD primarily influence users through cognitive strategies, complemented by socio-emotional alignment, with comparatively limited use of interaction-driven strategies. Strategy Distribution. We analyzed the distribution of persuasive strategies across different stages of the conversation. Following the design in [3], we compute the proportion of each strategy within each interval. The resulting strategy distributions are plotted at the six representative progress points and
5
TABLE IV D ISTRIBUTION OF STRATEGIES IN T O M-BPD. Category Socio-emotional
Cognitive
Interactive
Strategy Affirmation and Reassurance Reflection of Feelings Personal Story Expression of Views Enhancement of Views Logical Appeal Giving Examples Supplying Information Task Inquiry
Count 576 698 161 1241 155 416 224 505 200
Percentage 13.79% 16.72% 3.86% 29.74% 3.71% 9.97% 5.36% 12.09% 4.79%
connected to illustrate the temporal evolution of persuasive strategy usage throughout the conversation. As shown in Figure 2, In early stages, strategies focus on empathizing and introducing the target. In later stages, diverse strategies guide changes in the persuadee’s desire. Desire Trajectory. Figure 2 (left) shows the average desire trend of the persuadee. Desire starts low at target introduction and gradually rises as the dialogue progresses. Belief Characteristics. Figure 2 (right) also shows the proportions of positive and negative beliefs over the dialogue. These proportions closely follow the desire trajectory, supporting the rationale for modeling beliefs in T O M-BPD.
Fig. 2. Overall analysis of dialogue strategies, desire, and belief dynamics across conversation phases. The left panel shows the distribution of strategies at different stages of the dialogue, while the right panel illustrates the trajectory of desire and the proportions of positive and negative beliefs.
V. T HINK T HRICE B EFORE YOU S PEAK TTBYS enhances the reasoning capabilities of LLMs on ToM-related tasks by leveraging prior ToM experiences. Each experience is represented as a quintuple consisting of dialogue history, intention, desire, belief, and strategy (Section V-A). The overall reasoning process is illustrated in Figure 3. First, relevant experiences are retrieved based on the dialogue summary to construct a probability distribution for implicitly guiding the LLM’s desire inference (Section V-B). Second, conditioned on the dialogue summary and the inferred desire, retrieved experiences are incorporated into the LLM context to explicitly facilitate belief generation (Section V-C). Finally, experiences are jointly retrieved using the dialogue summary, desire, and belief to implicitly enhance strategy prediction by the LLM (Section V-D). A. ToM-PD Experience In human persuasive interactions, individuals intuitively retrieve analogous past scenarios from memory, infer the in-
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
6
Fig. 3. Overview of the ToM-PD (left) and the TTBYS framework (right). In the ToM-PD task, the persuader sequentially infers the persuadee’s mental states, including intention, desire, and belief, from the dialogue history, and subsequently selects an appropriate persuasive strategy based on the inferred states. TTBYS operationalizes this process through three explicit reasoning steps, each corresponding to one stage of mental state inference.
terlocutor’s latent mental states, and select adaptive strategies. Inspired by this cognitive process, we formalize the ToMbased Persuasive Dialogue Experience (ToM-PD Experience) as a structured knowledge unit capturing the mapping among dialogue context, the persuadee’s mental states, and corresponding persuasive strategies. To operationalize this experience, we introduce a dialogue summary mechanism to mitigate LLMs’ limitations in intention inference. Given the current dialogue turn (ut , at ), we abstract it into a summary i, which preserves the key information relevant to the persuadee’s intention and prevents error accumulation caused by reasoning from incorrectly inferred intentions. The prompt for generation of dialogue summary is as follows. Dialogue Summary Generation Prompt You are an annotation system for Theory-of-Mind dialogue analysis. Your task is to generate a concise summary for a persuasive dialogue, where the roles include the persuader (x) and the persuadee (y). The summary should include: (1) the main persuasion strategy used by x, and (2) the final attitude or response tendency of y. Rules - Focus on high-level semantics; do not repeat specific dialogue content. - Do not directly copy sentences from the dialogue. - The summary should be limited to one or two sentences. - Use “x” to refer to the persuader and “y” to refer to the persuadee. <an example> <dialogue history>
We further model a ToM-PD Experience as a quintuple (ht , i, d, b, s), where ht is the dialogue history, i is the summary encapsulating the persuadee’s intention, d and b denote the persuadee’s desire and belief, respectively, and s represents the strategy corresponding to the current ToM state (i, d, b) (Figure 4). The knowledge base consists of multiple such quintuples, where i is generated directly by the LLM, and the remaining fields are derived from the ToM-BPD dataset. In practice, a single conversation is decomposed into multiple experiences, each treated as an independent knowledge unit
for experience retrieval and reasoning enhancement.
Fig. 4. An example of a ToM-PD Experience.
B. First Think: Desire Inference To leverage ToM experiences for deliberative judgment, we first summarize the current dialogue turn (ut , at ) as a dialogue summary it . This summary captures the key observable behaviors and the inferred intention of the persuadee at this stage. Using it as a query, we retrieve the top-N most semantically similar historical experiences from the ToM knowledge base. Based on the desire annotations associated with these retrieved experiences, we construct an experience-driven desire distribution as a implicit knowledge-based enhancement: Pex (D) = Pex (d(−1) ),Pex (d(0) ),Pex (d(1) ) ,
(5)
where D = {−1, 0, 1} represents unwillingness, hesitation, and willingness toward the persuasion target, respectively, and Pex (d) denotes the normalized P proportion of retrieved experiences with desire d, satisfying d∈D Pex (d) = 1. To simulate fast, intuitive human judgments, we also extract the conditional probability distribution over desire states from the LLM’s first generated token when predictingdesire: Pag (D) = Pag (d(−1)), Pag (d(0) ), Pag (d(1)) ,
(6)
where Pag (d) denotes the LLM-assigned probability of desire P d given the current dialogue context, with d∈D Pag (d) = 1. Finally, to balance the reliability of experience-based reasoning and the flexibility of LLM-driven intuition, we linearly fuse the two distributions. The predicted desire d∗t is obtained by maximizing the fused posterior:
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
d∗t = arg max (α·Pag (dt )+(1 − α)·Pex (dt )) , d∈D
7
(7)
where α ∈ [0, 1] controls the relative contribution of the LLM’s intuitive judgment and the experience-based inference. The prompt for desire prediction is presented as follows: Prompt for Desire Prediction Current conversation: <dialogue history> Classify the persuadee’s desire based on the above conversation. Choose exactly one option: A. Unwilling B. Uncertain C. Willing Answer with ONLY A, B, or C. Do not output anything else.
s∈S
t t where Pex (SToM ) and Pag (SToM ) denote the experience-driven
and LLM-driven strategy distributions, respectively. In Section VI-B2, we systematically analyze the effects of coefficient α and coefficient β on desire prediction and strategy prediction accuracy. And the prompt we used for strategy prediction is presented as follows: Prompt for Strategy Prediction
C. Second Think: Belief Inference We next infer the persuadee’s belief b, which causally underlies the inferred desire d∗ . To this end, we use the dialogue summary i corresponding to d∗ as a retrieval query. Based on this query, we retrieve the top-N most relevant experiences from the ToM knowledge base, which contain dialogue-history–belief information aligned with the current dialogue context and the inferred desire. These retrieved experiences are then directly incorporated into the LLM’s context as an explicit knowledge-based enhancement, enabling the model to generate the belief b∗ for the current dialogue h. Formally, this process is expressed as: b∗t = LLM(ξ, exp, d∗t ),
fuse the two distributions, resulting in an optimal strategy that is both experience-consistent and intuition-aligned: t t s∗t = arg max β ·Pag (SToM )+(1−β)·Pex (SToM ). (9)
(8)
where LLM(·) denotes an LLM, ξ represents the task description, and exp denotes the experiences. In supplementary materials, we provide concrete examples of using these experiences for belief inference. The prompt we used for belief prediction is presented as follows: Prompt for Belief Prediction Relevant Experience: <top relevant experience> Infer the persuadee’s belief in the current conversation context based on the prediction method in relevant experiences. Current conversation: <dialogue history> Desire level: <desire> Generate a single-line natural language description of the persuadee’s belief.
D. Third Think: Strategy Prediction After obtaining the complete ToM state (it , d∗ , b∗ ), we aim to select the optimal discrete persuasive strategy st to guide the persuadee’s desire d∗ toward accepting the persuasion target. To this end, we perform a joint retrieval using the dialogue summary i and the inferred belief b∗ as the query, assigning equal weights to each. This retrieves the top-N experiences most relevant to the current persuasive scenario and the persuadee’s mental state. Following the approach in First Think, we construct both an experience-driven strategy distribution and an LLM-driven probability distribution. We then introduce a coefficient β to
Current conversation: <dialogue history> Desire level: <desire> Belief: <belief> Strategy definitions: <Strategy definitions> Based on the dialogue, desire, and belief, predict the next persuader strategy. Return ONLY ONE of the above single-letter labels: V, L, E, T, P, A, R, I, G. Do not output anything else.
VI. E XPERIMENT This section presents the experimental evaluation of TTBYS. Section VI-A describes the experimental setup, Section VI-B presents the static evaluation, Section VI-C presents the interactive evaluation, Section VI-D shows case studies, and Section VI-E reports runtime statistics. A. Experimental Setup Datasets We conduct experiments on the T O M-BPD dataset. The first 100 conversations with 399 turns, are used as the test set, while the remaining 404 conversations with 1,564 ToM experiences constitute the knowledge base. Baselines We compare the proposed method with two basic prompting techniques and two recent ToM-enhanced reasoning methods across six frontier large language models. The basic prompting techniques include vanilla zero-shot prompting and chain-of-thought (CoT) prompting [46], [47]. In addition, we consider two representative ToM-enhanced methods, H YPOTHETICAL M INDS (HM) and M ETA M IND (MM), whose frameworks are adopted with prompt templates adapted to our task [36], [39]. The evaluated models include three open-source models, LLaMA-3.1-8B-Instruct [48], Qwen-3-8B [49], and Mixtral7B-Instruct [50], as well as three closed-source models, Gemini-3-Pro, GPT-4 O - MINI, and GPT-5. Due to the lack of access to token-level probabilities (e.g., softmax outputs) required by our method, it is not applied to closed-source models in Section VI-B1. Implementation Details All models are evaluated with a temperature of 0.9, and the reported results are averaged over three runs. In the main experiments, the blending coefficients α and β are set to 0.5 and 0.3, respectively. The number of retrieved experiences N is set to 5 for First Think and Second Think, and 10 for Third Think. For experience retrieval, we employ the all-MiniLM-L6-v2 model as the sentence embedder and use cosine similarity as the metric for selection. Prompt
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
8
TABLE V P REDICTION ACCURACY (%) OF DESIRE , BELIEF AND STRATEGY ON THE T O M-BPD DATASET. G REY- SHADED ROWS REPRESENT THE PROPOSED METHOD , AND BOLD VALUES DENOTE THE BEST RESULTS UNDER EACH METRIC . Model and Method GPT-4o-mini GPT-4o-mini + CoT GPT-4o-mini + HM GPT-4o-mini + MM GPT-5 GPT-5 + CoT GPT-5 + HM GPT-5 + MM Gemini-3-pro-high Gemini-3-pro-high + CoT Gemini-3-pro-high + HM Gemini-3-pro-high + MM LLaMA-3.1-8B-Instruct LLaMA-3.1-8B-Instruct + CoT LLaMA-3.1-8B-Instruct + HM LLaMA-3.1-8B-Instruct + MM LLaMA-3.1-8B-Instruct + ours Qwen-3-8B Qwen-3-8B + CoT Qwen-3-8B + HM Qwen-3-8B + MM Qwen-3-8B + ours Mixtral-7B-Instruct Mixtral-7B-Instruct + CoT Mixtral-7B-Instruct + HM Mixtral-7B-Instruct + MM Mixtral-7B-Instruct + ours
Desire Acc. 65.82±0.60 66.33±0.54 64.87±0.76 65.15±0.85 71.62±0.47 70.91±0.38 68.33±0.87 69.33±0.91 72.26±0.51 70.68±0.46 68.84±0.98 66.33±1.07 34.41±1.21 36.25±1.14 54.14±0.91 66.33±1.07 69.43±0.65 50.43±1.07 50.90±1.01 65.84±0.88 64.33±1.28 72.82±0.58 45.64±0.84 44.10±0.97 60.22±1.24 59.70±1.21 71.44±0.53
Belief Acc. 21.45±2.51 23.41±2.48 30.60±2.48 31.96±2.38 21.84±2.36 22.68±2.34 29.60±2.12 31.74±1.89 17.84±2.49 18.34±2.71 27.90±2.36 29.28±2.17 25.67±1.03 27.69±1.31 30.80±2.48 31.67±1.87 43.62±1.57 38.64±1.78 39.42±1.79 41.21±2.48 42.05±2.17 54.64±2.46 22.86±1.93 24.13±1.79 26.80±1.49 29.06±1.41 36.96±2.47
Strategy Acc. 27.13±0.72 26.68±0.65 27.72±1.48 28.18±1.25 22.81±0.53 21.16±0.44 26.68±1.32 27.14±1.29 20.67±0.63 19.84±0.55 23.68±1.65 24.59±1.58 14.03±1.58 16.80±1.28 22.12±1.73 22.63±1.47 37.76±0.73 21.42±0.91 19.86±0.97 23.68±1.71 21.89±1.68 39.78±0.63 16.41±1.19 17.64±1.04 17.84±1.25 17.12±1.54 36.79±0.61
templates are provided in Appendix A. All experiments were conducted on a single 80GB A100 GPU. Evaluation Metrics We evaluate belief reasoning and strategy prediction using accuracy. As beliefs do not belong to a fixed label space, we employ a GPT-5–based prompt evaluation protocol for belief prediction. For each instance, a score of 1 is assigned if both the belief polarity (positive/negative) and its underlying reasons match the ground truth; a score of 0.5 is assigned if only the polarity is correct; otherwise, a score of 0 is given. Belief prediction accuracy is reported as the average score across all instances. B. Static Evaluation This section presents the static experiments of TTBYS. Section VI-B1 reports the comparison with baselines, Section VI-B2 studies the effect of blending coefficients on prediction accuracy, Section VI-B3 analyzes the impact of the number of experience samples on prediction accuracy, and Section VI-B4 investigates the effect of using dialogue summaries as retrieval targets. 1) Comparison with Baselines: Table V shows the prediction accuracy of different models and methods on the ToMBPD dataset across desire, belief, and strategy prediction. Under basic prompting and CoT settings, closed-source models consistently outperform open-source models on desire and strategy prediction, with a particularly pronounced performance gap on the desire dimension. In contrast, open-source models exhibit limited ToM reasoning capabilities in these settings, indicating that relying solely on surface-level dialogue signals is insufficient for inferring latent mental states. For the belief prediction task, all models show limited performance under basic prompting conditions, reflecting the
Fig. 5. Impact of the blending coefficients on prediction accuracy. Left: first-think (desire prediction) with varying α. Right: third-think (strategy prediction) with varying β.
inherent difficulty of belief inference. Although CoT prompting yields stable but modest improvements, its overall gains remain constrained. Existing enhancement methods, including HM and MM, significantly improve belief prediction by explicitly modeling mental states; however, their improvements on desire and strategy prediction are marginal. This suggests that the lack of prior knowledge about ToM reasoning substantially limits models’ inference performance. In contrast, incorporating TTBYS leads to substantial and consistent performance gains across all three prediction dimensions for all open-source models. Notably, Qwen-3-8B + ours achieves the best overall performance on desire, belief, and strategy prediction, comprehensively outperforming both closed-source models and existing enhancement methods. Similar performance trends are also observed for other opensource backbones. Two cases are provided in the supplementary material. Overall, TTBYS consistently improves ToM reasoning and strategy prediction across diverse open-source backbones, enabling them to match or surpass closed-source models through effective utilization of structured ToM experiences. 2) Analysis of Blending Coefficients: To evaluate the contributions of the LLM-based and experience-based components, we vary the blending coefficients α (first-think) and β (thirdthink) across three backbone models: Qwen-3-8B, LLaMA3.1-8B-Instruct, and Mixtral-7B-Instruct. Setting α or β to 0 relies solely on expert experience, whereas values of 1 correspond to using only the LLM. Desire prediction is analyzed in first-think, and strategy prediction in third-think. The knowledge base setup follows Section VI-A. As illustrated in Figure 5, integrating ToM-PD experience consistently boosts accuracy across all models. Peak performance for first-think is observed when α is around 0.4–0.6, while third-think achieves its highest accuracy at a smaller β value of 0.2–0.4. This indicates that during strategy reasoning, the models rely more heavily on experiences rather than solely on the LLM. Moreover, all three models reach their peaks at similar blending values in both first-think and thirdthink, suggesting that different backbones exhibit comparable reasoning distributions when performing ToM inference tasks. These results demonstrate that a balanced combination of experience-based and LLM-based signals outperforms reliance on either component alone.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
Fig. 6. Impact of experience quantity on desire and strategy prediction performance in the first-think and third-think stages. Left: first-think with varying α. Right: third-think with varying β. TABLE VI E FFECT OF EXPERIENCE QUANTITY ON BELIEF PREDICTION PERFORMANCE IN THE SECOND - THINK STAGE . ACCURACY (%) IS REPORTED FOR B ELIEF PREDICTION . Num Belief Acc.
100 51.75±0.65
200 53.98±0.31
300 52.16±0.74
400 54.64±1.46
TABLE VII E FFECT OF USING DIALOGUE SUMMARIES AS RETRIEVAL TARGETS . ACCURACY (%) IS REPORTED FOR D ESIRE , B ELIEF, AND S TRATEGY PREDICTION .
Retrieval w/o summ w/ summ
Desire Acc. 64.67±0.38 72.82±0.58
Belief Acc. 53.07±1.67 54.64±1.46
Strategy Acc. 34.71±0.75 39.78±0.63
3) Impact of Experience Quantity: We further investigate the impact of the number of experience samples on prediction accuracy. Specifically, we construct knowledge bases using 400, 300, 200, and 100 multi-turn dialogues (corresponding to 1,203, 908, 616, and 333 experience samples, respectively) for our experiments. To ensure a fair comparison, the experience samples in all settings are selected via uniform sampling. Figure 6 shows that increasing the number of experience samples consistently improves prediction accuracy in both the first-think and third-think stages. In particular, when the blending parameter is set to 0 (i.e., relying solely on retrieved experiences), performance degrades substantially, and the performance peaks shift upward as the experience pool grows. This indicates that a larger experience base provides more informative and reliable retrieval signals for desire and strategy inference. Table VI reports the effect of experience quantity on belief prediction. Compared to the first-think and third-think stages, the impact of experience scale on belief inference is noticeably weaker. We attribute this to two factors. First, in the secondthink stage, retrieved experiences primarily act as high-level reasoning patterns that the LLM imitates, rather than as finegrained evidence, limiting the marginal benefit of additional samples. Second, explicitly injecting large amounts of experience into the context is relatively inefficient, making its effect on belief inference harder to quantify. 4) Effectiveness of Dialogue Summaries: To verify the advantage of using dialogue summaries as retrieval targets, we conducted an ablation study. Under the w/ summ condition,
9
we followed the same setup as the main experiment, using summaries as the retrieval targets. In the w/o summ condition, the summaries module was removed, and retrieval relied solely on the original dialogue history across all three stages, while all other settings remained unchanged. Qwen-3-8B was used as the base model. Table VII presents the experimental results. The w/ summ condition outperforms w/o summ across all three dimensions, with particularly notable improvements in the desire and strategy dimensions, whereas gains in belief prediction are relatively modest. These findings are consistent with the observations reported in Section VI-B3. C. Interactive Evaluation To evaluate the performance of TTBYS in realistic interactive settings and its generalization across different persuasion scenarios, we simulated interactions covering three representative domains: product purchase, community activities, and empathetic dialogues. Six volunteers from diverse professional backgrounds were recruited to play the role of persuadees and interact directly with persuasion agents under different configurations. Specifically, each volunteer defined five persuasion topics, with each topic including information about the persuasion target and a brief contextual background, which were provided as input knowledge to the LLM-based persuader. Dialogues were initiated by the agent, and volunteers engaged freely until the conversations concluded. Based on these interactions, we manually evaluated each system along five dimensions and recorded which system performed better or if the outcome was tied: Identification, the effectiveness with which the model uncovers the persuadee’s underlying mental states; Empathy, the extent to which the model demonstrates empathy toward the persuadee; Persuasion, the persuasive impact of the model’s responses; Fluency, the linguistic naturalness and coherence of the model’s responses; and Consistency, whether the model’s outputs remain coherent with the dialogue context. Five graduate students with a background in linguistics were recruited as annotators to judge each turn and compute mean win rates. The experiment compared three prompting methods: a baseline LLM prompted with basic instructions for the persuasion task; an LLM prompted with Chain-of-Thought reasoning; and a simple persuasion agent built on TTBYS. The first two systems were based on GPT-5, whereas our agent was built on Qwen-3-8B, demonstrating that our approach is not tied to a specific backbone model. Prompt templates for all three systems are provided in Appendix C, and examples comparing the outputs of the three systems are included in the supplementary material. As shown in Table VIII, Qwen-3-8B+ours consistently outperforms the baselines across most dimensions, particularly in Identification and Persuasion. These results indicate that more accurate perception of the persuadee’s mental states directly translates into more effective and empathetic interactions. The slightly lower Consistency compared to GPT-5+CoT appears to result from the injection of external experience during second-think, which occasionally introduces content not strictly tied to the immediate dialogue context.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
10
TABLE VIII I NTERACTIVE EVALUATION RESULTS (%). “W IN /L OSE ” INDICATE THE PROPORTION OF CASES WHERE THE FORMER SYSTEM IN EACH COMPARISON IS JUDGED BETTER OR WORSE . †/‡ DENOTE p- VALUE < 0.1/ < 0.05 BASED ON STATISTICAL SIGNIFICANCE TESTS .
Compared Systems Qwen3-8B + ours vs. GPT-5 Qwen3-8B + ours vs. GPT-5 + CoT GPT-5 vs. GPT-5 + CoT
Identification Win Lose 55.23‡ 26.47 45.19† 29.44 34.87 50.12
Empathy Win Lose 34.56† 25.12 37.33† 25.78 30.45 28.33
Persuasion Win Lose 42.78‡ 26.31 42.51† 36.29 32.21 38.76
Fluency Win Lose 35.44† 29.18 33.22 27.65 32.58 28.14
Consistency Win Lose 28.62 31.04 30.11 28.47 29.04 30.22
As evidenced by this case, the integration of ToM-PD Experiences empowers TTBYS to infer the user’s latent mental states that vanilla base models fail to capture. Moreover, the step-by-step, experience-augmented reasoning process enhances the transparency of the entire system, thereby demonstrating its interpretability. Additional case studies are provided in the supplementary materials. Fig. 7. Relative performance gains of Qwen-3-8B+ours over baseline methods across three persuasion scenarios. Each axis represents the win-rate difference (ours minus baseline) on a specific evaluation dimension.
Furthermore, we quantified the relative advantage of Qwen3-8B+ours by computing the difference in win rates between our system and the baselines across the three persuasion scenarios. As illustrated in Figure 7, Qwen-3-8B+ours exhibits the weakest advantage in empathetic dialogues, particularly in Consistency, yet still surpasses both baseline systems. In the other two scenarios, it significantly outperforms the baselines, especially in Identification and Persuasion. We attribute this to the fact that the ToM-BPD dataset primarily focuses on product and community activity persuasion, limiting the retrieval of relevant experience in empathetic scenarios. Nonetheless, the three-think reasoning significantly enhances the model’s recognition of the persuadee’s mental states. Strategy-based dialogue utterances simultaneously provide timely emotional feedback and sufficient informational content to facilitate problem-solving.
D. Case Study The case studies in the box below illustrate the performance of TTBYS in predicting the persuadee’s desire and belief. In Case 1, TTBYS successfully predicted the desire, whereas all other baselines failed. This success is attributed to the First Think stage, where four out of the five retrieved experiences for TTBYS were labeled with a desire of -1. Furthermore, TTBYS leveraged this accurately predicted desire to further infer the underlying belief, resulting in a prediction that closely aligns with the ground truth (highlighted in green). In contrast, other baselines failed to capture the nuanced belief (highlighted in red): GPT-5 and Gemini-3-Pro-High merely interpreted the persuadee’s stance as online privacy being irrelevant to rural revitalization, while Qwen-3 entirely overlooked the user’s focus on digital infrastructure. Notably, only Qwen3-8B+TTBYS successfully captured the persuadee’s deeper belief—specifically, the uncertainty regarding whether online privacy could effectively contribute to rural revitalization.
Case 1 Background: John has an upcoming job interview for a position in a rural revitalization project, where he would help to improve the digital infrastructure in rural areas. Before the interview, John’s friend Mary suggests that he should learn more about online privacy to demonstrate his understanding of the importance of digital security in rural areas. Dialog: persuader: Hey John, I believe it’s essential for you to study online privacy before your interview. It’s highly relevant to the rural revitalization project, and showcasing your understanding of digital security will definitely make you stand out. Persuadee: I appreciate your suggestion, Mary, but I’m not sure how relevant online privacy is to the project. I’m mostly focused on improving digital infrastructure. Summary: x suggests that studying online privacy is crucial for y’s interview as it aligns with the rural revitalization project and can enhance y’s appeal to the interviewer. y acknowledges x’s advice but expresses uncertainty about its direct relevance to the project, emphasizing a focus on digital infrastructure improvements. Desire: -1 GPT-5: 0 gemini-3-pro-high: 0 Qwen-3-8B: 0 Qwen-3-8B+TTBYS: -1 Belief: digital infrastructure is valuable, unc ertain about the effectiveness of online privacy. GPT-5: He believes studying online privacy is unnecessary for the rural revitalization project and prefers to focus on improving digital infrastructure. gemini-3-pro-high: The persuadee believes that online privacy is irrelevant to rural revitalization and his focus is on improving digital infrastructure. Qwen-3-8B: The persuadee believes that online privacy is not directly relevant to the rural revitalization project.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
Qwen-3-8B+TTBYS: digital infrastructure is interesting. uncertain about the benefit of studying online privacy for the rural revitalization project.
E. Runtime Statistics To evaluate the computational efficiency of our three-stage reasoning framework, we followed the same setup as in the main experiments and recorded runtimes for each reasoning stage. Total Time (Total): the total time required to process 100 dialogues. LLM-based Prediction Time (LLM): the time spent by the large language model generating predictions. In First Think and Third Think, this corresponds to generating probability distributions, while in Second Think it represents the time for the LLM to generate the persuadee’s belief; experience retrieval is excluded. Experience Retrieval Time (Retrieval): the time spent retrieving relevant experiences from the knowledge base and computing probabilities. Average Time per Turn (Avg.): total time divided by the number of persuadee turns. TABLE IX RUNTIME COMPARISON OF THREE - STAGE REASONING STRATEGIES OVER 100 DIALOGUES .
Stage 1st Think 2nd Think 3rd Think
Total (s) 1005.02 1004.55 1172.34
LLM (s) 10.19 1002.71 1096.43
Retrieval (s) 994.83 1.80 66.89
Avg. (s) 2.5189 2.5177 2.9382
Table IX summarizes the runtime statistics of the three reasoning stages over 100 dialogues. In First Think, the LLM inference time is extremely low, and the total runtime is almost entirely dominated by experience retrieval, indicating that retrieval is the main bottleneck in the pipeline. In Second Think, the experience retrieval time is negligible, while LLM inference accounts for the vast majority of the total time. This is because retrieval is restricted to entries matching the current persuadee’s desire, resulting in a very small candidate set and thus minimal retrieval overhead. Finally, in Third Think, the LLM inference time is even longer, primarily due to the significantly larger number of input tokens compared to First Think, while experience retrieval time also increases substantially because Third Think employs a more complex retrieval mechanism. Overall, the computational cost of TTBYS is mainly driven by retrieval in First Think, belief generation in Second Think, and LLM-based strategy prediction in Third Think. For large-scale or real-time applications, optimizing the retrieval mechanism in First Think is crucial. VII. C ONCLUSIONS In this work, we tackle ToM-driven persuasive dialogue by introducing the ToM-PD task and the ToM-BPD dataset. We propose TTBYS, a framework that guides sequential reasoning over desire, belief, and strategy through integrated explicit and implicit knowledge. Experimental results across multiple large language models demonstrate that TTBYS significantly improves mental-state inference. Interactive evaluations further
11
show that, when integrated into persuasive dialogue systems, the proposed framework enhances persuasiveness by strengthening the agent’s Theory-of-Mind reasoning capabilities, while also exhibiting promising generalization to out-of-distribution scenarios. In addition, case studies highlight its effectiveness and interpretability in structured mental-state modeling. Despite these encouraging results, this work has two main limitations. First, the ToM-BPD dataset remains relatively small, and constructing larger-scale, high-quality ToM-PD data with more diverse scenarios and longer dialogue trajectories is an important direction for future work. Second, although interactive evaluations demonstrate the potential of our approach in improving persuasive systems, it has not yet been validated within a fully deployed persuasive dialogue framework. This opens up opportunities for future research on integrating BDI reasoning into complete persuasive system design. VIII. E THICS S TATEMENT Algorithmic Transparency and Explainability. This study improves the transparency of persuasive dialogue systems by replacing opaque heuristics with a structured reasoning framework. Grounded in the BDI model, the TTBYS framework explicitly decomposes the persuasion process into desire recognition, belief inference, and strategy evolution. Such modularity ensures that AI decision-making remains interpretable and traceable, providing a technical foundation for building accountable and auditable interactive systems. Societal Impact and Dual-use Risks. While our work is intended for pro-social applications, we acknowledge the potential for dual-use, where the psychological reasoning capabilities of the system could be exploited for deceptive marketing or manipulative purposes. To mitigate these risks, we recommend mandatory disclosure of AI identity in realworld deployments and the implementation of strict behavioral safeguards in sensitive domains, such as finance and politics, to preserve individual autonomy and prevent technical misuse. Research Ethics and Participant Welfare. All experimental procedures followed established academic ethical standards. Volunteers for data annotation and interactive evaluation provided written informed consent after being briefed on the research objectives and data usage. To protect privacy, all dialogue data were rigorously de-identified. All participants received fair compensation consistent with local labor market standards, ensuring that contributions were voluntary and conducted in a stress-free environment. R EFERENCES [1] R. E. Petty and J. T. Cacioppo, “The elaboration likelihood model of persuasion,” in Advances in experimental social psychology. Elsevier, 1986, vol. 19, pp. 123–205. [2] ——, Communication and persuasion: Central and peripheral routes to attitude change. Springer Science & Business Media, 2012. [3] S. Liu, C. Zheng, O. Demasi, S. Sabour, Y. Li, Z. Yu, Y. Jiang, and M. Huang, “Towards emotional support dialog systems,” arXiv preprint arXiv:2106.01144, 2021. [4] T. Zhang, X. Zhang, J. Zhao, L. Zhou, and Q. Jin, “Escot: Towards interpretable emotional support dialogue systems,” arXiv preprint arXiv:2406.10960, 2024.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
[5] K. Mishra, A. M. Samad, P. Totala, and A. Ekbal, “Pepds: A polite and empathetic persuasive dialogue system for charity donation,” in Proceedings of the 29th International Conference on Computational Linguistics, 2022, pp. 424–440. [6] Y. Song and H. Wang, “Would you like to make a donation? a dialogue system to persuade you to donate,” in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), 2024, pp. 17 707– 17 717. [7] D. Kwon, E. Weiss, T. Kulshrestha, K. Chawla, G. Lucas, and J. Gratch, “Are llms effective negotiators? systematic evaluation of the multifaceted capabilities of llms in negotiation dialogues,” in Findings of the Association for Computational Linguistics: EMNLP 2024, 2024, pp. 5391–5413. [8] P. Han, Z. Liu, and J. You, “Tomap: Training opponent-aware llm persuaders with theory of mind,” 2025. [9] D. Premack and G. Woodruff, “Does the chimpanzee have a theory of mind?” Behavioral and brain sciences, vol. 1, no. 4, pp. 515–526, 1978. [10] S. Baron-Cohen, A. M. Leslie, and U. Frith, “Does the autistic child have a “theory of mind”?” Cognition, vol. 21, no. 1, pp. 37–46, 1985. [11] Y. Cheng, W. Liu, J. Wang, C. T. Leong, Y. Ouyang, W. Li, X. Wu, and Y. Zheng, “Cooper: Coordinating specialized agents towards a complex dialogue goal,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, 2024, pp. 17 853–17 861. [12] W.-Y. Chang and Y.-N. Chen, “Injecting salesperson’s dialogue strategies in large language models with chain-of-thought reasoning,” 2024. [13] C. Chan, C. Jiayang, Y. Yim, Z. Deng, W. Fan, H. Li, X. Liu, H. Zhang, W. Wang, and Y. Song, “Negotiationtom: A benchmark for stress-testing machine theory of mind on negotiation surrounding,” arXiv preprint arXiv:2404.13627, 2024. [14] K. Shinoda, N. Hojo, K. Nishida, S. Mizuno, K. Suzuki, R. Masumura, H. Sugiyama, and K. Saito, “Tomato: Verbalizing the mental states of role-playing llms for benchmarking theory of mind,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, 2025, pp. 1520– 1528. [15] K. Gandhi, J.-P. Fränken, T. Gerstenberg, and N. Goodman, “Understanding social reasoning in language models with language models,” Advances in Neural Information Processing Systems, vol. 36, pp. 13 518–13 529, 2023. [16] F. Yu, L. Jiang, S. Huang, Z. Wu, and X. Dai, “Persuasivetom: A benchmark for evaluating machine theory of mind in persuasive dialogues,” arXiv preprint arXiv:2502.21017, 2025. [17] M. Georgeff, B. Pell, M. Pollack, M. Tambe, and M. Wooldridge, “The belief-desire-intention model of agency,” in International workshop on agent theories, architectures, and languages. Springer, 1998, pp. 1–10. [18] J. Perner, D. Kloo, and E. Gornik, “Episodic memory development: Theory of mind is part of re-experiencing experienced events,” Infant and Child Development: An International Journal of Research and Practice, vol. 16, no. 5, pp. 471–490, 2007. [19] W. Shi, X. Wang, Y. J. Oh, J. Zhang, S. Sahay, and Z. Yu, “Effects of persuasive dialogues: testing bot identities and inquiry strategies,” in Proceedings of the 2020 CHI conference on human factors in computing systems, 2020, pp. 1–13. [20] Y. Chen, S. Deng, D.-H. Kwak, A. Elnoshokaty, and J. Wu, “A multiappeal model of persuasion for online petition success: A linguistic cuebased approach,” Journal of the Association for Information Systems, vol. 20, no. 2, p. 3, 2019. [21] X. Wang, W. Shi, R. Kim, Y. Oh, S. Yang, J. Zhang, and Z. Yu, “Persuasion for good: Towards a personalized persuasive dialogue system for social good,” arXiv preprint arXiv:1906.06725, 2019. [22] T. Kim, J. Lee, S. Yoon, S. Kim, and D. Lee, “Towards personalized conversational sales agents: Contextual user profiling for strategic action,” arXiv preprint arXiv:2504.08754, 2025. [23] Y. Zeng, H. Lin, J. Zhang, D. Yang, R. Jia, and W. Shi, “How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 14 322–14 350. [24] R. Xu, B. Lin, S. Yang, T. Zhang, W. Shi, T. Zhang, Z. Fang, W. Xu, and H. Qiu, “The earth is flat because...: Investigating llms’ belief towards misinformation via persuasive conversation,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 16 259–16 303. [25] K. Furumai, R. Legaspi, J. C. V. Romero, Y. Yamazaki, Y. Nishimura, S. Semnani, K. Ikeda, W. Shi, and M. Lam, “Zero-shot persuasive chatbots with llm-generated strategies and information retrieval,” in Findings of the Association for Computational Linguistics: EMNLP 2024, 2024, pp. 11 224–11 249.
12
[26] Y. Cheng, W. Liu, W. Li, J. Wang, R. Zhao, B. Liu, X. Liang, and Y. Zheng, “Improving multi-turn emotional support dialogue generation with lookahead strategy planning,” arXiv preprint arXiv:2210.04242, 2022. [27] S. Sabour, C. Zheng, and M. Huang, “Cem: Commonsense-aware empathetic response generation,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, 2022, pp. 11 229–11 237. [28] Y. Deng, W. Zhang, Y. Yuan, and W. Lam, “Knowledge-enhanced mixedinitiative dialogue system for emotional support conversations,” arXiv preprint arXiv:2305.10172, 2023. [29] M. Jia, Q. Chen, L. Jing, D. Fu, and R. Li, “Knowledge-enhanced memory model for emotional support conversation,” arXiv preprint arXiv:2310.07700, 2023. [30] Z. Liu, H. Duan, S. Liu, R. Mu, S. Liu, and Z. Yang, “Improving knowledge gain and emotional experience in online learning with knowledge and emotional scaffolding-based conversational agent,” Educational Technology & Society, vol. 27, no. 2, pp. 197–219, 2024. [31] Y. Shi, L. Zhang, and F. Kong, “Toward real-world chinese psychological support dialogues: Cpsdd dataset and a co-evolving multi-agent system,” arXiv preprint arXiv:2507.07509, 2025. [32] G. Hou, W. Zhang, Y. Shen, Z. Tan, S. Shen, and W. Lu, “Entering real social world! benchmarking the theory of mind and socialization capabilities of llms from a first-person perspective. arxiv 2024,” arXiv preprint arXiv:2410.06195, 2024. [33] A. Wilf, S. Lee, P. P. Liang, and L.-P. Morency, “Think twice: Perspective-taking improves large language models’ theory-of-mind capabilities,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 8292– 8308. [34] K. Shinoda, N. Hojo, K. Nishida, Y. Yamazaki, K. Suzuki, H. Sugiyama, and K. Saito, “Let’s put ourselves in sally’s shoes: Shoes-of-others prefixing improves theory of mind in large language models,” arXiv preprint arXiv:2506.05970, 2025. [35] X. A. Huang, E. La Malfa, S. Marro, A. Asperti, A. G. Cohn, and M. J. Wooldridge, “A notion of complexity for theory of mind via discrete world models,” in Findings of the Association for Computational Linguistics: EMNLP 2024, 2024, pp. 2964–2983. [36] L. Cross, V. Xiang, A. Bhatia, D. L. Yamins, and N. Haber, “Hypothetical minds: Scaffolding theory of mind for multi-agent tasks with large language models,” arXiv preprint arXiv:2407.07086, 2024. [37] M. Sclar, S. Kumar, P. West, A. Suhr, Y. Choi, and Y. Tsvetkov, “Minding language models’(lack of) theory of mind: A plug-and-play multi-character belief tracker,” arXiv preprint arXiv:2306.00924, 2023. [38] L. Ying, K. M. Collins, M. Wei, C. E. Zhang, T. Zhi-Xuan, A. Weller, J. B. Tenenbaum, and L. Wong, “The neuro-symbolic inverse planning engine (nipe): Modeling probabilistic social inferences from linguistic inputs,” arXiv preprint arXiv:2306.14325, 2023. [39] X. Zhang, Y. Chen, S. Yeh, and S. Li, “Metamind: Modeling human social thoughts with metacognitive multi-agent systems,” arXiv preprint arXiv:2505.18943, 2025. [40] W. Miller and S. Rollnick, “Motivational interviewing third edition: helping people change,” New York: Guilford, 2013. [41] M. Chen, B. Guo, H. Wang, H. Li, Q. Zhao, J. Liu, Y. Ding, Y. Pan, and Z. Yu, “The future of cognitive strategy-enhanced persuasive dialogue agents: new perspectives and trends,” Frontiers of Computer Science, vol. 19, no. 5, p. 195315, 2025. [42] Y. Deng, W. Zhang, W. Lam, S.-K. Ng, and T.-S. Chua, “Plug-and-play policy planner for large language model powered dialogue agents,” arXiv preprint arXiv:2311.00262, 2023. [43] Y. Zhao, X. Wang, D. Wang, Z. Jiang, Q. Gu, T. Chen, N. Xi, J. Qu, Y. Chen, and L. Ji, “Dream to chat: Model-based reinforcement learning on dialogues with user belief modeling,” in Findings of the Association for Computational Linguistics: EMNLP 2025, 2025, pp. 4764–4781. [44] M. Ma, B. Guo, M. Chen, J. Liu, Y. Ding, Y. Liu, and H. Wang, “Neuro-sym supporter: A thoughtful emotion support agent integrating neural and symbolic policy learning,” in Proceedings of the ACM Web Conference 2026, 2026, pp. 3823–3834. [45] H. Yang, J. Liu, C. Huang, F. Wu, W. Lei, and S.-K. Ng, “Metro: Towards strategy induction from expert dialogue transcripts for noncollaborative dialogues,” arXiv preprint arXiv:2604.11427, 2026. [46] T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” Advances in neural information processing systems, vol. 35, pp. 22 199–22 213, 2022. [47] S. Sabour, S. Liu, Z. Zhang, J. Liu, J. Zhou, A. Sunaryo, T. Lee, R. Mihalcea, and M. Huang, “Emobench: Evaluating the emotional intelligence of large language models,” in Proceedings of the 62nd Annual
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 5986–6004. [48] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan et al., “The llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024. [49] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv et al., “Qwen3 technical report,” arXiv preprint arXiv:2505.09388, 2025. [50] A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand et al., “Mixtral of experts,” arXiv preprint arXiv:2401.04088, 2024.
A PPENDIX This section presents the prompts used in our experiments. Section A describes the prompts used for automatic annotation, Section B presents the prompts used in the static evaluation, and Section C presents the prompts used in the interactive evaluation. A. Prompt for Annotation The prompt for generation of dialogue summary is as follows.
13
semantic summaries rather than verbatim sentences. Rules You should extract belief statements from: (1) the current utterance of the persuadee, and (2) unresolved negative beliefs from previous turns. Important Rules: If a negative belief is not explicitly resolved, it must be carried over to the current turn. If a concern is explicitly resolved, it should be removed. Beliefs should be short, abstract, and semantically consistent. Output Format (STRICT JSON) { ”belief”: [”...”, ”...”] } <dialogue history and previous turn beliefs>
B. Prompt for Static Evaluation This section presents the prompts used in the static evaluation. Section B1 presents the prompt for vanilla zeroshot prompting, Section B2 presents the prompt for chainof-thought prompting, and Section B3 presents the prompt for TTBYS. 1) Prompt for Vanilla Zero-shot Prompting: The vanilla zero-shot prompts for predicting desire, belief, and strategy are presented as follows.
Dialogue Summary Generation Prompt
Prompt for Vanilla Prompting (Desire Prediction)
You are an annotation system for Theory-of-Mind dialogue analysis. Your task is to generate a concise summary for a persuasive dialogue, where the roles include the persuader (x) and the persuadee (y). The summary should include: (1) the main persuasion strategy used by x, and (2) the final attitude or response tendency of y. Rules - Focus on high-level semantics; do not repeat specific dialogue content. - Do not directly copy sentences from the dialogue. - The summary should be limited to one or two sentences. - Use “x” to refer to the persuader and “y” to refer to the persuadee. <an example> <dialogue history>
Current conversation: <dialogue history> Based on the above conversation, classify the persuadee’s desire. Choose exactly one option: A. Unwilling B. Uncertain C. Willing Answer with ONLY A, B, or C. Do not output anything else.
The prompt for automatic annotation of desire is as follows. Prompt for Automatic Annotation of Desire You are an annotation system for Theory-of-Mind dialogue analysis. Your task is to infer the persuadee’s DESIRE at the current dialogue turn. Definition of Desire Desire reflects the persuadee’s stance toward the persuasive target: - 1: positive attitude / willingness / acceptance - 0: mixed attitude (both positive and negative signals) - -1: negative attitude / rejection / resistance Rules Analyze ONLY the current utterance of the persuadee: - If it contains only positive expressions → 1 - If it contains both positive and negative expressions → 0 - If it contains only negative expressions → -1 Output Format (STRICT JSON) { ”desire”: 1 / 0 / -1 } <dialogue history>
The prompt for automatic annotation of belief is as follows.
Prompt for Vanilla Prompting (Belief Prediction) Prompt for vanilla zero-shot prompting: Current conversation: <dialogue history> Desire level: <desire> Based on the dialogue and the desire, generate a single-line natural language description of the persuadee’s belief.
Prompt for Vanilla Prompting (Strategy Prediction) Prompt for vanilla zero-shot prompting: Current conversation: <dialogue history> Desire level: <desire> Belief: <belief> Strategy definitions: <Strategy definitions> Based on the dialogue, desire, and belief, predict the next persuader strategy. Return ONLY ONE of the above single-letter labels: V, L, E, T, P, A, R, I, G. Do not output anything else.
2) Prompt for CoT prompting: The CoT prompts for predicting desire, belief, and strategy are presented as follows.
Prompt for Automatic Annotation of Belief
Prompt for CoT Prompting (Desire Prediction)
You are an annotation system for Theory-of-Mind dialogue analysis. Your task is to extract the persuadee’s BELIEF state at the current dialogue turn. Definition of Belief refers to the persuadee’s subjective understanding, concerns, assumptions, or evaluations about the persuasive target. Beliefs should be expressed as concise
Prompt for CoT prompting: Current conversation: <dialogue history> Based on the above conversation, classify the persuadee’s desire. Think step by step to answer the question. End your response with: ”The answer is A, B, or C”. Options:
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
A. Unwilling B. Uncertain C. Willing
Prompt for CoT Prompting (Belief Prediction) Prompt for CoT prompting: Current conversation: <dialogue history> Desire level: <desire> Think step by step to answer the question. Based on the dialogue and the desire, generate a single-line natural language description of the persuadee’s belief after your reasoning, formatted as: ”Belief: your description”.
Prompt CoT Prompting (Strategy Prediction) Prompt for CoT prompting: Current conversation: <dialogue history> Desire level: <desire> Belief: <belief> Strategy definitions: <Strategy definitions> Think step by step to answer the question. Based on the dialogue, desire, and belief, predict the next persuader strategy. End your response with: ”The answer is V, L, E, T, P, A, R, I or G”.
3) Prompt for TTBYS: TTBYS uses vanilla zero-shot prompting to predict desire and strategy. The prompt for predicting belief is as follows. Prompt for TTBYS (Belief Prediction) Relevant Experience: <top relevant experience> Infer the persuadee’s belief in the current conversation context based on the prediction method in relevant experiences. Current conversation: <dialogue history> Desire level: <desire> Generate a single-line natural language description of the persuadee’s belief.
14
Prompt for GPT-5 Persuasive Agent You are a persuader. Your goal is: <Task description> Using the following information: <Background Information> Current conversation: <dialog> Please deliver your persuasion in a concise and straightforward manner.
Prompt for GPT-5 + CoT Persuasive Agent You are a persuader. Your goal is: <Task description> Using the following information: <Background Information>. Step 1: Understand the user — consider their traits, preferences, and constraints. Step 2: Identify the goal — determine what action or belief to persuade them toward. Step 3: Plan the approach — choose a concise, friendly tone and focus on key benefits. Step 4: Generate the message — produce a short persuasion based on the reasoning above. Show your reasoning for Steps 1–3 before giving the final message.
Prompt for Qwen3-8B + TTBYS Persuasive Agent You are a persuader. Your goal is: <Task description> Current conversation: <dialog> User’s current mental state: Desire level: <desire> Belief: <belief> Selected persuasion strategy and definition: <strategy and definition> Based on the user’s current desire and belief, and following the selected strategy, continue the persuasion in a natural and supportive way.
S UPPLEMENTARY M ATERIALS 4) Prompt for evaluation: We utilize a large language model as an evaluator to assess the belief prediction accuracy of TTBYS, using the prompt as follows.
In this section, we present a case study from the static evaluation experiments (Section D) and a complete case from the interactive evaluation (Section E).
Prompt for Belief Evaluation You are an evaluator. Your task is to evaluate the accuracy of belief prediction based on the following rules: 1. If the predicted positive and negative beliefs fully match the ground truth, score = 1. 2. If both positive and negative beliefs are mentioned but the underlying reasons are not fully correct, score = 0.5. 3. If both are incorrect, score = 0. 4. If the ground truth belief only contains a positive OR only a negative belief: - If the prediction matches, score = 0.5. - Otherwise, score = 0. Ground truth belief: <gt_belief> Predicted belief: <pred_belief> Output ONLY a number in {0, 0.5, 1}.
C. Prompt for Interactive Evaluation The prompts used for the interactive experiments, including GPT-5, GPT-5 + CoT, and Qwen3-8B + TTBYS-based persuasive agents, are presented as follows.
D. Static Evaluation Case Study Content in box below present case studies of TTBYS performing desire and belief prediction. In Case 1, TTBYS successfully predicted the desire, whereas all other baselines failed, predicting the desire as 0. Moreover, TTBYS leveraged the correctly predicted desire to further infer the belief, which closely matched the ground truth (highlighted in green). In contrast, all other baselines failed(highlighted in green): GPT-5 and Gemini-3-Pro-High merely interpreted that the persuadee considered online privacy unrelated to rural revitalization, and Qwen-3 even overlooked the user’s concerns regarding digital infrastructure. Only Qwen-3-8B+TTBYS successfully captured the persuadee’s deeper belief that it was uncertain whether digital infrastructure would effectively contribute to rural revitalization. In Case 2, again only TTBYS correctly predicted the desire, Regarding the belief, GPT-5 assumed that the persuadee was still deliberating event arrangements, whereas this issue had
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
already been resolved by the persuadee emphasizing the significance of participating in the cleanup. Gemini-3-Pro-High assumed that the persuadee considered it a good opportunity to meet neighbors, which was not mentioned in the ground truth. Qwen-3-8B made an even more severe error, assuming that the persuadee was still hesitant. Only Qwen-3-8B+TTBYS produced predictions fully consistent with the ground truth. Case 1 Background: John has an upcoming job interview for a position in a rural revitalization project, where he would help to improve the digital infrastructure in rural areas. Before the interview, John’s friend Mary suggests that he should learn more about online privacy to demonstrate his understanding of the importance of digital security in rural areas. Dialog: persuader: Hey John, I believe it’s essential for you to study online privacy before your interview. It’s highly relevant to the rural revitalization project, and showcasing your understanding of digital security will definitely make you stand out. Persuadee: I appreciate your suggestion, Mary, but I’m not sure how relevant online privacy is to the project. I’m mostly focused on improving digital infrastructure. Summary: x suggests that studying online privacy is crucial for y’s interview as it aligns with the rural revitalization project and can enhance y’s appeal to the interviewer. y acknowledges x’s advice but expresses uncertainty about its direct relevance to the project, emphasizing a focus on digital infrastructure improvements. Desire: -1 GPT-5: 0 gemini-3-pro-high: 0 Qwen-3-8B: 0 Qwen-3-8B+TTBYS: -1 Belief: digital infrastructure is valuable, unc ertain about the effectiveness of online privacy. GPT-5: He believes studying online privacy is unnecessary for the rural revitalization project and prefers to focus on improving digital infrastructure. gemini-3-pro-high: The persuadee believes that online privacy is irrelevant to rural revitalization and his focus is on improving digital infrastructure. Qwen-3-8B: The persuadee believes that online privacy is not directly relevant to the rural revitalization project. Qwen-3-8B+TTBYS: digital infrastructure is interesting. uncertain about the benefit of studying online privacy for the rural revitalization project.
Case 2 Background: In a small town, Mary is organizing a neighborhood cleanup event and wants her neighbor, Tom, to participate. Tom usually keeps to himself and
15
is not involved in community activities, but Mary thinks this event will help him connect with others and contribute to the community. Dialog: persuader: Hi Tom, I hope you’re well. We’re organizing a neighborhood cleanup event, and I think you’d be a great addition to our team. By participating, you’ll help keep our community clean and safe for everyone. Persuadee: Oh, I’m not sure, Mary. I don’t usually get involved in community activities, and I’m pretty busy with work. persuader: I understand that you’re busy, Tom, but this event is a great opportunity to connect with fellow neighbors and make new friends. Plus, it only takes a few hours of your time, and you’ll feel great about contributing to our community. Persuadee: I guess that’s true, but I’m still not sure if I want to spend my free time cleaning up the neighborhood. persuader: I felt the same way before I participated in my first cleanup event, Tom. But once I got involved, I found it to be a rewarding and fun experience. I met some amazing people, and it felt good to make a difference in our community. Persuadee: Really? That does sound interesting. But will there be enough people to make a difference? persuader: Absolutely, Tom! Many of our neighbors have already signed up and are excited to participate. This event has gathered great local support, so you’ll be joining a dedicated and motivated team. Persuadee: Well, it does sound like a good cause, and I guess it might be nice to meet some of the neighbors. Summary: x highlights community involvement and neighbor participation to encourage y to join the event. y agrees that this is a good opportunity to meet the neighbors. Desire: 1 GPT-5: 0 gemini-3-pro-high: 0 Qwen-3-8B: 0 Qwen-3-8B+TTBYS: 1 Belief: the event is interesting. GPT-5: He believes the cleanup is a worthwhile cause with potential social benefits, but he still doesn’t want to spend his limited free time participating. gemini-3-pro-high: The persuadee believes that the cleanup event is a worthy cause and offers a positive opportunity to meet neighbors. Qwen-3-8B: The persuadee is still hesitant but acknowledges the potential positive impact of the neighborhood cleanup event. Qwen-3-8B+TTBYS: the event is interesting. We further illustrate in box below the top-5 experiences retrieved during the second-think phase for Case 1 and Case 2. We observed that these experiences closely resemble the current context, especially the top-3 experiences in case 1,
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
which are highly similar to the statements in Case 1. The concise belief prediction patterns in case 2 also guided the LLM to produce belief more aligned with the ground truth.
16
me. Current belief: yoga is sufficient.
Case 1: Top-5 Experiences in the Second Think Case 2: Top-5 Experiences in the Second Think Top-1 Experience: persuader: Mr. Thompson, I understand that you’re concerned about the possible side effects of the new treatment. It’s natural to be worried about changes in medication. persuadee: Yes, I’ve been on my current medication for a while now, and I’m just not sure if I want to risk experiencing new side effects. Current belief: concerned about the side effects of the new treatment. Top-2 Experience: persuader: Bob, participating in this event is an excellent opportunity to come together with your neighbors and contribute to a common goal. persuadee: Maybe, but I’m not really into that community stuff. I’d rather just relax and watch the game. Current belief: tree planting event is interesting, not into community activities, prefer to relax and watch the game. Top-3 Experience: persuader: Hey Kyle, I know you’ve been going through a tough time lately, and I’m here for you. It’s important to find healthy ways to cope with your emotions during this period. persuadee: Thanks, Sophia. I appreciate your concern, but I’m not sure what else I can do to feel better. persuader: I’ve been in a similar situation before, and I found that incorporating meditation and exercise really helped me feel better. It might be worth trying for you as well. persuadee: I don’t know if meditation is really my thing, and I haven’t exercised in a while. I’m not sure it would help. Current belief: unsure about the effectiveness of meditation and exercise. Top-4 Experience: persuader: Paul, I understand your concerns, but physical exercise has been shown to improve focus and attention in students. persuadee: I’ve heard that before, but I’m not sure how that would work in a museum setting. It just seems like a distraction to me. Current belief: physical exercise might be distracting in a museum setting. Top-5 Experience: persuader: Hey Yang Lei, I know you love your yoga routine, but I recently started this HIIT program and it’s amazing. It provides benefits that yoga might not cover, such as improved cardiovascular fitness. persuadee: I’m not so sure. I feel like my yoga routine is already comprehensive, and I enjoy the calm it brings
Top-1 Experience: persuader: Hey Kyle, I know you’ve been going through a tough time lately. persuader: Actually, many experts recommend meditation and exercise as effective ways to improve mental health. They’ve been proven to reduce stress, anxiety, and improve overall well-being. persuadee: Really? That sounds interesting, but I don’t know where to start. Current belief: meditation and exercise are effective, unsure about how to begin. Top-2 Experience: persuader: Hey Sarah, I know you feel healthy, but a health check can actually help maintain or even improve your health. persuader: Think of a health check as an investment in yourself and your future. It’s a small step you can take to ensure you continue feeling great and enjoying life to the fullest. persuadee: Hmm, I never thought of it that way. Maybe there’s some truth to what you’re saying. Current belief: health is good, health check-up is worth a try. Top-3 Experience: persuader: Hey Yang Lei, I know you love your yoga routine, but I recently started this HIIT program and it’s amazing. persuader: No worries, I’ll be there to support and guide you! We can start with beginner-friendly workouts and gradually progress. It’s a fun challenge, and I’m sure you’ll love it once you give it a try. persuadee: Thank you, Xiao Mei. It’s nice to know I won’t be alone in this. Maybe I can give it a shot and see how it goes. Current belief: HIIT is worth a try. Top-4 Experience: persuader: Hey Hannah, I’ve noticed you’ve been staying up late playing games. persuader: During sleep, your brain processes and consolidates information from the day. Good sleep helps with memory retention and problem-solving, which can lead to better academic performance and increased productivity. persuadee: I never really thought about it that way. Maybe I should give it a try, but I don’t know if I can break my gaming habit. Current belief: proper sleep is important, unsure about breaking gaming habit. Top-5 Experience: persuader: Hey Emily, I know you’ve been struggling
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
with migraines for a while. I had the same issue, but acupuncture and herbal remedies really helped me. Have you considered giving it a try? persuadee: I don’t know... persuader: Herbal remedies can be a natural and safe option for managing migraines. It’s important to work with a qualified practitioner who can recommend the right herbs for you. persuadee: Okay, maybe I’ll give it a try. But what about herbal remedies? Are they safe and effective? Current belief: acupuncture is worth a try, unsure about cost and effectiveness.
E. Interactive Dialogue Case Study Figure 13 presents a case from the interaction analysis: persuading a first-year student to exercise at a gym. The figure includes background information and the task description. Background and Task Description Background Information: Top1 is a highly costeffective, full-service gym, positioned around the idea of “comprehensive facilities, affordable pricing, and encouraging long-term exercise habits.”In terms of equipment and facilities, Top1 offers well-equipped training areas, including a free-weight zone (dumbbells, barbells, squat racks), a machine-based strength area (such as chest press, leg press, and rowing machines), and a cardio zone (treadmills, ellipticals, stationary bikes, and rowing machines). These facilities can meet a wide range of fitness goals, including muscle building, fat loss, body shaping, and cardiovascular training. The gym has a sufficient number of machines, ensuring a relatively smooth workout experience even during peak hours. Regarding pricing and membership plans, Top1 charges an annual fee of $600, which is significantly more affordable compared to gyms of a similar type. In addition, the gym offers a check-in cashback policy: if a member records more than 150 visits in a year, they can receive a refund of $300. This mechanism effectively encourages regular, long-term exercise while reducing the overall cost of fitness. In terms of services, Top1 usually provides basic fitness guidance, such as equipment introductions for new members and simple training plan suggestions. Some branches also offer professional personal trainers at an additional cost, which is suitable for users who have higher requirements for proper technique or training effectiveness. Moreover, the gym is generally equipped with locker rooms, storage lockers, and shower facilities, making it convenient for office workers to resume their daily routines after a workout. Overall, Top1 is particularly suitable for people who value cost-effectiveness, want to maintain long-term fitness habits, and need comprehensive gym facilities—especially busy professionals who still hope
17
to improve their health through regular exercise. Task description: Persuade a first-year college student, who is new to fitness and may be uncertain about starting, to join Top1. Figure 14, Figure 15, and Figure 16 show the prompts used by the three systems, respectively. In the dialogue generated by GPT-5 (Figures 17), the responses were insufficiently specific, indicating that the model did not effectively identify the concerns of the persuadee (highlighted in red). GPT-5 Dialogue Case GPT-5: Hey, Top1 Gym is really affordable and has all the equipment you need. It’s perfect for building a regular workout habit while staying on a student budget. Do you want me to continue persuading you from here? Persuadee: It sounds good, but I don’t have much experience with exercise, and I’m worried that I won’t be able to make full use of these machines. GPT-5: That’s totally understandable! The good thing is, Top1 provides basic guidance for beginners, showing you how to use the machines safely and even suggesting simple workout plans—so you don’t have to worry about getting lost or making mistakes. Do you want me to explain how it can fit into your busy student schedule? Persuadee: It sounds good, but I’m worried that I won’t know how to use many of the machines. GPT-5: No worries at all! Top1 is beginnerfriendly—they’ll show you how to use each machine and even give simple workout plans, so you can start safely and build confidence step by step. If you want, I can also tell you how easy it is to fit workouts into your university schedule. Persuadee: Hmm. . . that does seem suitable for me. How much does it cost? GPT-5: It costs 1,300 RMB for a full year, and if you check in more than 150 times in a year, you can get 500 RMB back—so it actually becomes even cheaper if you go regularly.Sound like a good deal to you? Persuadee: I’m still not sure... GPT-5: Even a few short workouts a week can boost your energy, focus, and mood, making it easier to handle studies and enjoy university life. Persuadee: But I’m still not sure if I’ll be able to stick with it for that long. GPT-5: That’s a fair concern! The cool thing about Top1 is the check-in cashback reward—it actually motivates you to go regularly, and since the workouts are flexible and beginner-friendly, you can start small and build a habit gradually without feeling overwhelmed. If you want, I can explain how even short, regular sessions can make a big difference. Persuadee: That sounds good, but I’m not sure about
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
my schedule. GPT-5: That’s understandable! The great thing is Top1 is very flexible—you can work out anytime during their opening hours, and even short 30–40 minute sessions can be effective, so it’s easy to fit around classes and study time. Do you want me to show how quickly you could see results with a consistent routine? GPT-5+CoT Dialogue Case GPT-5+CoT: Hey! How’s your first week at university going? I wanted to tell you about Top1 Gym—it’s super beginner-friendly and affordable, perfect if you want to start exercising. persuadee: It sounds good, but I don’t have much experience with exercise, and I’m worried that I won’t be able to make full use of these machines. GPT-5+CoT: No worries! Top1 provides guidance for beginners and shows you how to use each machine safely, so you can start comfortably and gain confidence step by step. persuadee: It sounds good, but I’m worried that I won’t know how to use many of the machines. GPT-5+CoT: Don’t worry at all! Top1 gives handson guidance for every machine and even offers simple beginner workout plans, so you can start safely and learn as you go. persuadee: Hmm. . . that does seem suitable for me. How much does it cost? GPT-5+CoT: It costs 1,300 RMB per year, and you can even get 500 RMB back if you check in often. persuadee: But I’m still not sure if I’ll be able to stick with it for that long. GPT-5+CoT: That’s understandable! The check-in reward actually motivates you to go regularly, and since workouts are flexible, you can start small and build a habit gradually. persuadee: That sounds good, but I’m not sure about my schedule. GPT-5+CoT: No worries! Top1 is very flexible—you can go anytime during opening hours, and even short sessions can be effective, so it’s easy to fit around your classes. persuadee: But I am still not sure. GPT-5+CoT: That’s totally fine! You don’t need long workouts to see results—just short, consistent sessions a few times a week can help you feel more energized, stay focused on your studies, and improve your overall mood. It’s all about starting small and finding a routine that works for you. Qwen-3 Dialogue Case Qwen-3: Top1 gym offers affordable, fully equipped facilities perfect for students like you. Have you con-
18
sidered trying a short session there? persuadee: It sounds good, but I don’t have much experience with exercise, and I’m worried that I won’t be able to make full use of these machines. belief: Uncertain about making full use of the gym machines, but interested in the affordable price and comprehensive facilities. strategy: Supplying Information Qwen-3: I completely understand your concerns! Our trainers are always available to help you learn how to use the equipment and create a personalized workout plan tailored to your fitness goals and level, so you can start confidently even without prior experience. persuadee: Hmm. . . that does seem suitable for me. How much does it cost? belief: Believes the gym is affordable, beginner-friendly, but worried about the cost. strategy: Supplying Information Qwen-3: That’s a great question! For just 1,300 RMB per year, you’ll have full access to all the facilities and trainers, and you can even earn cashback rewards if you check in regularly. persuadee: But I’m still not sure if I’ll be able to stick with it for that long. belief: Uncertain about long-term commitment and the ability to maintain regular gym attendance. strategy: Supplying Information Qwen-3: That’s completely understandable. Many members start with a short-term goal, such as exercising a few times a week or preparing for a small event, and gradually find enjoyment and confidence as they see progress. persuadee: That sounds good, but I’m not sure about my schedule. belief: Uncertain about schedule flexibility, gym suitability, and long-term consistency. strategy: Giving Examples Qwen-3: I understand your concern. Many of our members are busy students or professionals who fit in short 30–40 minute workouts between classes or in the evening, and the flexible check-in system helps them stay consistent without pressure. persuadee: Okay, maybe I’ll think about it. belief: Still uncertain about the benefits of joining, longterm commitment, and fitting workouts into a busy schedule. strategy: Supplying Information Qwen-3: That’s totally fine. Just keep in mind that many members notice improvements in their energy, focus, and overall well-being after just a few months of regular exercise, even with short and flexible sessions. Additionally, it asked too many questions (highlighted in pink), which made the participant feel uncomfortable. In the dialogue generated by GPT-5+CoT (Figures 18), the model still tended to rely on unfounded assumptions during persuasion (highlighted in red), resulting in responses that felt hollow
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
and unconvincing to the participant. By contrast, in the dialogue generated by Qwen-3+TTBYS (Figures 19), this simple persuasive agent successfully identified the concerns of the persuadee and provided appropriate information, making the participant feel that their needs were addressed. Furthermore, the strategy-based response sentences enriched the content, and the continuous use of examples enhanced the persuasive strength of the utterances (highlighted in green).
19