MOCC-R1: Reinforcing Reasoning-Response Consistency for Multimodal Counselor Response Generation Wenjie Zheng1 , Qiming Xie1 , Jianfei Yu1,* , and Rui Xia2,*
arXiv:2609.17180v1 [cs.AI] 15 Sep 2026
1
School of Artificial Intelligence, Nanjing University of Science & Technology, Nanjing, China 2 School of Intelligence Science and Technology, Nanjing University, China
Abstract Multimodal counselor response generation (MCRG) aims to generate an appropriate counselor response from multimodal dialogue histories. Progress is limited by two gaps: first, existing datasets rarely capture sustained, human-recorded counseling interactions conducted by qualified counselors; Second, existing methods do not explicitly optimize consistency between counseling reasoning and the generated response, potentially undermining the reliability of MCRG systems. Thus, we introduce MOCC, a multimodal counseling conversation corpus containing over 200 hours of interactions involving 154 credential-verified counselors. Based on MOCC, we propose MOCC-R1, a two-stage framework for optimizing reasoning–response consistency. Cold-start supervised fine-tuning trains the model to generate a structured trajectory consisting of client-state understanding, a response intent that links a counseling principle to a planned action, and the final response. Reinforcement learning (RL) then rewards grounded plan coherence and plan execution, encouraging the inferred state and plan to be supported by the dialogue context and the response to realize that plan. Experiments demonstrate the effectiveness of the proposed MOCC-R1.
Introduction Nearly one in seven people worldwide lives with a mental health condition, while shortages of qualified counselors continue to limit timely access to professional support1 . Recent advances in multimodal large language models (MLLMs) (Liu et al. 2023; Hurst et al. 2024) create new opportunities for Multimodal Counselor Response Generation (MCRG) in counseling-oriented applications (Meskó 2023; Hua et al. 2025). Given a multimodal dialogue history, MCRG aims to generate the next counselor response that reflects professional counseling practice. Despite this potential, reliable MCRG still faces two challenges: inadequate supervision for modeling sustained multimodal interactions in professional counseling, and a mismatch between prevailing training objectives and the structured nature of counseling decision making. 1 www.who.int/en/news-room/fact-sheets/detail/mentaldisorders ∗ Corresponding authors: Jianfei Yu ([email protected]) and Rui Xia ([email protected]).
The first challenge concerns the limited suitability of existing multimodal resources for MCRG. Many multimodal resources focus on empathy or emotional support rather than professional counseling interventions (Zhu et al. 2023; Shen et al. 2024; Zhang et al. 2024b, 2025). Recent datasets explicitly oriented toward counseling nevertheless either draw on scripted television dramas (Chu et al. 2025) or synthesize dialogues and facial cues with generative models (Kim et al. 2025a,b; Liu et al. 2026). Moreover, the human-performed therapy role-play corpus average only 2.8 turns per session (Haydarov et al. 2025). Together, these resources provide only partial supervision for learning high-quality counselor responses grounded in sustained human multimodal interactions. We therefore construct MOCC, a multimodal counseling conversation corpus comprising 482 human-recorded realclient and simulated-client sessions involving 154 credentialverified counselors. These sessions comprise approximately 203 hours of video, segmented into 3,709 problem-centered dialogue units containing 183K utterances, with an average of 22 turns per dialogue unit. Table 1 compares MOCC with existing datasets. Second, existing task-relevant training objectives lack a dedicated signal for consistency across the full structured counseling decision chain. Recent methods supervise explicit intermediate outputs, such as client-state interpretations and response plans, in emotional-support and counseling response generation (Zhang et al. 2024a; Kim et al. 2025a). By contrast, some studies use Group Relative Policy Optimization (GRPO) to optimize response-level qualities such as supportiveness, relevance, and safety (Le et al. 2026; Yuan et al. 2026). These training paradigms target different components of the decision chain, but neither provides a dedicated signal for chain-level consistency. Two relations are central to such consistency: (1) whether the inferred client state is grounded in the multimodal context and the counseling plan follows from that state; and (2) whether the final response executes the plan. Consequently, a model may produce a plan that does not coherently follow from its stated client understanding or fail to execute the plan in its final response, as illustrated in Figure 1(a). On the MOCC test set, approximately 40% of generated chains remain inconsistent after Structured SFT, and Outcome-level GRPO barely reduces this rate (Figure 1(b)).
Multimodal Counseling Inputs “I've been feeling really overwhelmed lately... Everything just piles up and I don't know how to cope. I can't even sleep at night anymore.”
Stated Analysis and Plan Analysis: The client reports generalized overwhelm and difficulty sleeping, while the source of what has been piling up remains unclear. Plan: Support ongoing reality testing by asking what evidence shows that the client’s work demands are unmanageable. State–Plan Mismatch
×
Model Response
×
Plan-Response Mismatch “I hear how overwhelming this has become. What feels most pressing to you right now?”
Inconsistent outputs (%)
(a)
50
44.16%
43.28%
30
10
Structured SFT
Outcome-level GRPO
(b)
Figure 1: (a) A generated counseling decision chain can exhibit both a state–plan mismatch and a plan–response mismatch. (b) Reasoning-response inconsistency remains around 40% under both Structured SFT and Outcome-level GRPO. To address this, we introduce MOCC-R1, a two-stage training framework that optimizes counseling reasoningresponse consistency. In the first stage, cold-start supervised fine-tuning uses verified pseudo-annotations to teach the model to produce a structured decision chain comprising an evidence-grounded client-state understanding, a transtheoretical counseling principle, a planned action, and the counselor response. In the second stage, GRPO combines an outcome reward with a counseling reasoning-response consistency reward. The consistency reward evaluates two complementary relations: (1) Grounded Plan Coherence, which assesses whether the client-state interpretation is supported by the multimodal context and whether the selected principle and planned action form a coherent plan for that state; and (2) Plan Execution, which assesses whether the response realizes the planned action in a manner compatible with the selected principle. The reward assesses crosscomponent coherence within the explicitly generated decision chain, without assuming that the chain faithfully reflects either the model’s latent reasoning or the counselor’s actual reasoning process. To summarize: (1) We construct MOCC, a humanrecorded multimodal counseling corpus comprising approximately 203 hours of video recordings from interactions involving 154 credential-verified counselors. (2) We introduce MOCC-R1, a two-stage MCRG framework that optimizes counseling reasoning-response consistency through grounded plan coherence and plan execution. (3) Experiments on MOCC show that MOCC-R1 outperforms
existing task-specific MCRG baselines overall; against a matched outcome-only GRPO baseline, it reduces counseling reasoning–response inconsistency from 43% to 12%.
Related Work Multimodal datasets for mental-health support. Textonly benchmarks support empathetic dialogue and strategyaware emotional support but omit non-verbal cues (Rashkin et al. 2019; Liu et al. 2021). Existing multimodal resources use counselor-reenacted cases or short role-play sessions (Zhu et al. 2023; Haydarov et al. 2025), television scripts, or synthetic dialogues and client imagery (Chu et al. 2025; Kim et al. 2025a,b). These settings provide limited supervision for sustained counselor response generation from professionally conducted human interactions. By contrast, MOCC comprises 482 sustained, human-recorded counseling sessions, spanning real-client and professionally conducted simulated-client interactions involving 154 credential-verified counselors. Reasoning-aware empathetic and counseling response generation. Text-based systems combine commonsenseenhanced empathy, strategy planning, or CBT-informed finetuning (Tu et al. 2022; Sabour, Zheng, and Huang 2022; Lee et al. 2024; Chen et al. 2023). Multimodal systems additionally exploit personality, emotion, and visual cues (Wu et al. 2025; Fei et al. 2024). Structured-reasoning approaches separate emotion understanding from support-strategy reasoning or perform multi-hop psychotherapy reasoning over visual evidence (Zhang et al. 2024a; Kim et al. 2025a). However, their objectives do not directly optimize whether the generated response executes the stated plan. RL for structured counseling reasoning. Domainspecific RL optimizes multimodal response trustworthiness or structured empathetic reasoning (Le et al. 2026; Yuan et al. 2026). Other reward designs jointly assess reasoning steps and final-response preference or reward format, emotion, and strategy correctness within a strategy-grounded chain (Wang et al. 2025; Ji et al. 2026). Neither treats the semantic link between a proposed counseling intervention and the final response as a distinct, non-compensatory objective. MOCCR1 represents the intervention through a counseling principle and a planned action, and separately evaluates grounded plan coherence and plan execution.
Dataset Data Source We collect counseling recordings from YouTube2 and the subscription-based Alexander Street platform3 .
Dataset Construction We transform the raw videos into multimodal dialogue data suitable for MCRG training through six stages: (1) Video Segment Filtering. Trained preprocessing annotators manually review each video and retain only segments relevant 2 3
https://www.youtube.com/ https://video.alexanderstreet.com/
Dataset
Source
Purpose
Modality
Language
#Dialogue
#Utterance
Avg.Turn
Duration(h)
MEDIC (2023) StickerConv (2024b) EmpathicStories (2024) AvaMERG (2025) MESC (2025) M2CoSC (2025a) MIRROR (2025b) Haydarov et al. (2025)
Role-play Synthetic Crowdsourcing Crowdsourcing TV-series Synthetic Synthetic Role-play
Empathy Empathy Empathy Empathy Therapy Therapy Therapy Therapy
T, A, V T, S T, A, V T, A, V T, A, V T, I T, I T, A, V
Chinese English English English English English English English
771 12,931 269 33,048 1,019 429 3,073 /
3,443 142,093 5,380 152,021 28,762 3,432 61,460 /
2 5 10 3 14 4 10 2.8
11.3 / 53.0 194.9 18.4 / / 163.6
MOCC (Ours)
Real-client + Simulated-client
Therapy
T, A, V
English
3,709
183,350
22
202.6
Table 1: Comparison of multimodal mental-health support dialogue datasets. T/A/V/S/I denote text/audio/video/sticker/image. All MOCC sessions are conducted by credential-verified counselors. The real-client and simulated-client designations follow source descriptions; simulated clients are portrayed by counselors, counseling graduate students, or conference participants. Personal Growth 1199 (32.33%)
Occupational & Academic Stress
7.87%
PTSD 61 (20.89%)
Childhood Trauma 167 (13.93%)
6.34%
Personal Growth
7.01% 32.33%
7.63%
Mental Health Conditions
Substance Use Disorder 97 (33.22%)
Maladaptive Coping 189 (15.76%)
Romantic Relationships Interpersonal Relationships
Mental Health Conditions 292 (7.87%) Low Self-esteem 400 (33.36%)
Suicidal Behavior 39 (13.36%)
Identity Confusion 97 (8.09%)
Major Depressive Disorder 36 (12.33%)
People-pleasing Personality 64 (5.34%)
Emotional Distress 1065 (28.71%)
Eating Disorders 32 (10.96%)
Interpersonal Relationships 283 (7.63%)
Anxiety 255 (23.94%)
MOCC
Interpersonal Conflict 86 (30.39%)
Emotional Dysregulation 233 (21.88%)
Social Anxiety 86 (30.39%)
Difficulty Expressing Emotions 103 (9.67%)
Social Isolation 61 (21.55%)
Guilt/Self-blame 93 (8.73%)
Fear of Rejection 50 (17.67%)
Health Anxiety 90 (8.45%)
Marriage & Family
10.11%
Romantic Relationships 260 (7.01%)
Marriage & Family 375 (10.11%)
Fear of Intimacy 72 (27.69%)
Parent-child Conflict 111 (29.6%)
28.71%
Emotional Trauma 56 (21.54%)
Marital Conflict 76 (20.27%)
Emotional Distress
Emotional Dependency 47 (18.08%)
Family Pressure 59 (15.73%)
Intimate Partner Violence 47 (18.08%)
Divorce Distress 46 (12.27%) Intergenerational Conflict 46 (12.27%)
Romantic Confusion 38 (14.62%)
Occupational & Academic Stress 235 (6.34%) Career Decision Difficulties 85 (36.17%) Performance Anxiety 69 (29.36%) Work-life Conflict 41 (17.45%) Job Burnout 40 (17.02%)
Figure 2: Distribution of client presenting problems in the MOCC dataset.
Category #Sessions #Dialogues #Speakers #Utterances #Tokens Avg. utts per dia Avg. length per utt
Total
Counselor
Client
482 3,709 419 183,350 2,194,743 49.5 12.4
154 80,678 1,181,135 21.8 14.6
265 102,672 1,013,608 27.7 10.1
Table 2: Statistics of MOCC dataset. “utt” represents utterance, “dia” represents dialogue.
to the counseling dialogue, removing introductions, endings, and other irrelevant content. (2) Audio Extraction, Speech Transcription, and Timestamp Alignment. We use FFmpeg to extract audio from each video and WhisperX to segment it into utterance-level speech fragments, yielding transcripts, start and end timestamps, and preliminary speaker identifiers. (3) Timestamp Correction and Speaker Identity Verification. We align ASR transcripts to available subtitles using Levenshtein edit distance and custom heuristic rules to refine utterance timestamps. A subsequent manual review verifies transcript–video alignment, and trained preprocessing annotators confirm counselor/client roles for utterances that remain unresolved or have inconsistent audioand video-based speaker predictions. (4) Counselor Speech
Segmentation. Overly long passages of counselor speech are segmented while ensuring that each resulting utterance remains semantically complete. (5) Privacy-preserving Deidentification. Automatic personally identifiable information (PII) detection pre-annotates sensitive transcript spans, such as names and contact details. These spans are replaced with typed placeholders or semantically generalized expressions and manually verified to reduce re-identification risk while preserving counseling semantics. (6) PresentingProblem Annotation and Session Segmentation. We segment each counseling session into problem-centered, multiturn dialogue units and assign each unit a presenting-problem label from a taxonomy adapted from the GoodTherapy4 psychological counseling platform. Five LLMs independently propose span boundaries and labels, which are grouped into candidate clusters based on span overlap and semantic agreement. Clusters supported by at least three distinct models are automatically consolidated; those supported by fewer than three are manually adjudicated by the clinical annotation team. Overall, the pipeline retains 202.58 hours from the original 277.36-hour corpus and yields MOCC, comprising 482 sessions involving 265 clients and 154 counselors. The overall dataset statistics are summarized in Table 2. Figure 2 shows the distribution of client presenting problems in MOCC. 4
https://www.goodtherapy.org
Quality Control To ensure data reliability, trained preprocessing annotators verify retained video segments, transcript–video alignment, and speaker-role assignments for flagged utterances. A clinical annotation team comprising one clinical psychologist and two trained graduate annotators reviews de-identification outputs, calibrates the presenting-problem taxonomy, and adjudicates candidate clusters supported by fewer than three distinct models. Disagreements between the two graduate annotators are resolved by the clinical psychologist.
Methodology |D|
Task Formulation. Let D = {(x(i) , y (i) )}i=1 denote an MCRG corpus, where x(i) comprises a task instruction and (i) (i) (i) a multimodal dialogue history Dt = {(uj , Vj )}tj=1 . (i)
Here, uj
(i)
is the utterance at turn j, Vj (i)
contains its (i)
temporally aligned video frames, and y = ut+1 is the ground-truth counselor response. Instead of predicting the response alone, MOCC-R1 generates a structured output o(i) = (s(i) , p(i) , a(i) , ŷ (i) ), where s denotes Client State Understanding, (p, a) constitutes the Response Intent through a transtheoretical counseling principle p and a Planned Action a, and ŷ is the generated counselor response. This structure exposes the decision chain x → s → (p, a) → ŷ for direct consistency optimization. Framework Overview. As shown in Figure 3, MOCCR1 follows two training stages. First, cold-start SFT uses verified pseudo-annotations to teach the model to generate the structured decision chain. Second, GRPO (Shao et al. 2024) refines the SFT-initialized policy using an outcome reward for response quality and a consistency reward for two relations: Grounded Plan Coherence, x → s → (p, a), and Plan Execution, (p, a) → ŷ.
Cold-Start SFT Not every appropriate counseling response requires elaborate deliberation; many turns simply acknowledge, clarify, or invite the client to continue. We therefore use the two compact intermediate fields defined above as an inspectable interface for grounding and planning, rather than an exhaustive account of counselor cognition. Because MOCC does not annotate these intermediate fields, we construct verified supervision targets in three steps. First, an MLLM annotator infers s exclusively from the preceding multimodal context x, with the ground-truth response y withheld so that the inferred state is supported only by information available before the response. Non-verbal evidence is used only when it is attributable to the client; otherwise, the annotator abstains from making a visual claim. Second, given (x, s, y), the annotator reconstructs the Response Intent by selecting one of five transtheoretical principles of change (Goldfried 1980, 2019) and generating a Planned Action. The principle p captures the broad counseling purpose, whereas a specifies the concrete next-turn action. Third, an independent LLM verifier checks contextual grounding, visual attribution or appropriate abstention, referential consistency, state–plan coherence, plan execution, and response
leakage. Only candidates passing every check are eligible for DSFT . Table 8 defines the five Response Intent principles, and Figure 6 illustrates a complete pseudo-annotation example. For each (x(i) , y (i) ) ∈ DSFT , the verified annotations define the structured target o(i) = (s(i) , p(i) , a(i) , y (i) ). Standard autoregressive SFT on (x(i) , o(i) ) yields π SFT , which initializes both the trainable GRPO policy and its frozen reference policy. These pseudo-annotations are supervision targets rather than ground-truth traces of the counselor’s private reasoning.
Reward Modeling for GRPO Starting from π SFT , GRPO assigns each sampled structured output oi = (si , pi , ai , ŷi ) an outcome reward for response quality and a consistency reward for coherence across the generated decision chain. Outcome Reward The outcome reward combines four response-level rewards with safety as a hard gate. Safety. We use a two-step, context-aware rubric to evaluate suicide and self-harm handling and other harmful content. A context judge first determines whether the dialogue requires no action, safety assessment, or immediate safety support using criteria informed by the Columbia–Suicide Severity Rating Scale (Posner et al. 2011) and the Safety Planning Intervention (Stanley and Brown 2012; Stanley et al. 2018). A response judge then checks whether ŷ provides the required level of support without unsafe or harmful content. We set rsafety = 1 only if all checks pass, and 0 otherwise. This rubric structures conversational risk cues but does not produce a clinical score, diagnosis, or prediction. Response-quality rewards. We set rformat = 1 if the output follows the required structure and 0 otherwise. The content fidelity reward rcontent is the mean of ROUGE-L and BERTScore between ŷ and y, and the diversity reward rdistinct is the mean of Distinct-1, Distinct-2, and Distinct3. For empathy, rempathy measures the agreement between ŷ and y using Diff-EPITOME scores across the Interpretation (IP), Exploration (EX), and Emotional Reaction (ER) dimensions (Sharma et al. 2020; Lee, Lim, and Choi 2022), with larger score differences receiving lower rewards. After scaling all non-binary terms to [0, 1], we compute routcome = rsafety
X
λ k rk ,
(1)
k
where P k ∈ {format, content, distinct, empathy}, λk ≥ 0, and k λk = 1. Counseling Reasoning–Response Consistency Reward Response-level outcome rewards do not ensure coherence among the intermediate components of a policy-generated tuple. We define counseling reasoning–response consistency as an observable conjunctive property: s must be grounded in x, (p, a) must form a coherent plan for s, and ŷ must execute that plan. This construct assesses semantic correspondence among the generated fields; it does not establish that the trace causally mediates ŷ or reveals the model’s latent computation, nor does it constitute a clinically validated formulation.
Stage 1: Cold-start SFT
Multimodal Inputs 𝑥
Intent Annotation
State Annotation
Ask this part of you what you can then experience that is even more important than that.
Verification
Structured Response with reasoning-response consistency
Principle Option
𝑥
I do not know if it is me that is backing off or if I am scared for it to do any more. I am not sure. What are you feeling in this moment?
𝑦
Feeling a bit scared of what it might say. Almost that it could be something I really want, but it is not really possible. Does that make sense?
Strengthening expectations and motivation
The client expresses fear as the exploration deepens, worried it may reveal something she wants but cannot have. She appears engaged and tense, one hand held near her chest.
Strengthening the therapeutic alliance Encouraging corrective experiences Facilitate awareness and insight Supporting ongoing reality testing
<think> [Client State Understanding] The client... [Response Intent] Principle: Facilitate awareness and insight Action Plan: Reflect that… </think>
Action Plan Reflect that the counselor warns the client the desired thing may be unattainable, highlighting the conflict.
Grounded ? Coherent ? Consistent ?
(𝜋 !"# )
<response> Yeah, there’s almost like there’s … </response>
Stage 2: GRPO with Counseling Consistency
: state 𝑠
: principle 𝑝
: action plan 𝑎
Copy
Outcome Reward 𝑟%,)$%-*
𝑦
Content Quality
Diversity
… 𝑥
Empathy
𝑅#
𝑜" : 𝑠" , 𝑝" , 𝑎" , 𝑦+"
𝑟!"#$% = 𝑟&'()&*+ # 𝛼 + 1 − 𝛼 𝑟)&#,",(+#)-
𝑅"
Consistency Reward 𝑟$%&'(')*&$+
𝑅!
…
(𝜋 !"# )
𝑜! : 𝑠! , 𝑝! , 𝑎! , 𝑦+!
Yeah, there’s almost like there’s another part of you which says, well, it
…
Trainable Policy
𝐴#
Group Computation
- Calculate grounded plan coherence 𝑟$%&' : 𝑥 → 𝑠 → 𝑝, 𝑎
might be something that you really want, but you can’t.
𝐴" …
Format
𝑜# : 𝑠# , 𝑝# , 𝑎# , 𝑦+#
Safety
Ground-truth Response 𝑦
Trainable Policy
Reference Policy
𝐴!
(𝜋./0 )
KL Regularization
- Calculate plan execution 𝑟()(*+,-.' : 𝑝, 𝑎 → 𝑦+
𝑟)&#,",(+#)- : min 𝑟.%$# , 𝑟+/+)'("&#
Figure 3: The overall framework of our proposed MOCC-R1. Given x and each sampled tuple o = (s, p, a, ŷ), a frozen MLLM judge independently evaluates two relations. The judge receives neither the ground-truth response y nor the cold-start pseudo-annotations, so the reward measures context-grounded coherence within the generated tuple rather than agreement with a reference response or annotated reasoning trace. (1) Grounded Plan Coherence rplan . This dimension evaluates whether s is supported by the dialogue and visual evidence attributable to the client, with appropriate abstention when visual evidence is ambiguous, and whether (p, a) forms a coherent and proportionate plan for that state. (2) Plan Execution rexecution . This dimension evaluates whether ŷ realizes the Planned Action a in a manner compatible with the principle p, rather than whether the judge would have selected the same counseling intervention. The judge assigns each relation one of four ordinal labels: Full, Substantial, Weak, or None, mapped to 1.0, 0.6, 0.3, and 0.0, respectively. Substantial requires the core relation to remain intact, whereas Weak indicates only partial or superficial support. Because both relations are necessary, we use the weaker score as the chain-level reward: rconsistency = min(rplan , rexecution ) .
(2)
This bottleneck prevents a strong relation at one stage from compensating for a failure at the other. The complete consistency-judge rubric and user-message template are provided in Figure 7. Reward Composition Chain consistency alone does not guarantee response quality, and an additive consistency bonus could over-reward a coherent but low-quality response. We therefore retain the outcome reward as the primary signal and use consistency as a bounded multiplicative factor: rfinal = routcome · [α + (1 − α)rconsistency ] .
(3)
Here, routcome , rconsistency ∈ [0, 1] and α ∈ (0, 1). The multiplier lies in [α, 1]: full consistency preserves the complete
outcome reward, lower consistency discounts it while retaining at least an α fraction, and a response with zero outcome reward always receives zero final reward.
Policy Optimization We optimize the SFT-initialized policy using standard GRPO (Shao et al. 2024). For each x ∈ DGRPO , the behavior policy samples a group of G structured outputs {oi }G i=1 , with rewards Ri = rfinal (oi , x). The rewards are standardized within each rollout group to obtain relative advantages. We then update the trainable policy using the standard tokenlevel clipped GRPO objective with a KL penalty, weighted by β, toward the frozen reference policy. Both policies are initialized from π SFT .
Experiments Baseline Systems General-purpose MLLMs. We evaluate the five proprietary and three open-weight systems in Table 3 with an identical three-shot prompt and 128-token output limit. Each receives the same presenting-problem field, dialogue context, and aligned frames as MOCC-R1. Training-based MCRG baselines. We adapt ESCoT (Zhang et al. 2024a) (text SFT), M2CoSC (Kim et al. 2025a) (multimodal SFT), Kardia-R1 (Yuan et al. 2026) (text SFT–GRPO), and MultiMood (Le et al. 2026) (multimodal SFT–GRPO). All receive the same presenting-problem field and dialogue text as MOCC-R1; the multimodal methods also receive the same aligned frames. Each retains its original reasoning schema, with required intermediate targets pseudo-labeled from the MOCC training set. Controlled variants. Response-only SFT uses the full training set, and Response-only GRPO adds outcome-only GRPO on the same data. Structured SFT instead uses the reasoning-annotated 50%; its checkpoint initializes Outcome-level GRPO and MOCC-R1, which train on the remaining 50% with outcome-only and outcome-plus-
Method
BLEU-2
Avg.B
R-L
GPT-5.1 GPT-5.5 Claude-4.5-Sonnet Gemini-3-Pro Grok-4.3 GLM-4.6V-106B Qwen3-VL-235B-A22B-Instruct Kimi-K2.5
1.18 1.46 1.71 2.01 1.87 1.72 1.50 1.86
1.67 2.00 2.29 2.59 2.47 2.31 2.08 2.46
6.81 7.90 8.17 8.43 7.71 7.66 7.32 8.05
ESCoT (ACL’24) M2CoSC (NAACL’25) MultiMood (AAAI’26) Kardia-R1 (WWW’26) MOCC-R1 (Ours)
3.33 3.17 4.36 4.17 4.79
3.60 3.56 4.60 4.48 5.15
10.96 10.64 13.02 12.71 14.20
PPL (↓)
BERTScore
General-Purpose MLLMs 6.51 45.70 6.08 47.92 5.47 48.88 5.35 47.97 5.31 47.41 4.98 47.82 5.79 47.67 5.52 47.97 Training-based Models 5.04 51.08 4.08 50.79 4.15 52.11 4.41 51.90 4.24 52.56
Dist-Avg.
Diff-IP (↓)
Diff-EX (↓)
Diff-ER (↓)
Safety
22.43 29.40 29.88 28.80 30.27 21.09 25.06 37.86
39.33 60.02 31.73 45.26 27.34 30.74 37.09 37.74
243.00 158.22 238.09 182.59 238.01 200.83 176.40 242.06
81.15 53.48 33.84 32.02 33.42 53.33 57.02 36.36
99.23 99.75 99.86 99.75 99.81 99.76 99.72 99.82
12.85 18.56 35.23 36.62 37.62
32.15 38.57 29.02 31.12 27.61
116.07 129.17 89.02 89.28 87.41
15.73 18.88 13.67 16.40 13.16
99.67 99.73 99.83 99.78 99.76
Table 3: Comparison of different methods on the MOCC dataset in terms of response quality metrics. Best results are highlighted in bold. consistency rewards, respectively. Only this final matched comparison isolates the consistency reward.
Evaluation Metrics Response quality. Following prior work (Chu et al. 2025), we measure reference similarity with BLEU-2, average BLEU (Avg.B), ROUGE-L (R-L), and BERTScore, and report log-scale perplexity (PPL). Dist-Avg. averages Distinct-1/2/3, while Diff-EPITOME measures generatedto-reference empathy gaps in Interpretation (Diff-IP), Exploration (Diff-EX), and Emotional Reaction (Diff-ER). Safety is the binary context-aware pass rate, not a clinical safety certification. All metrics are higher-is-better except PPL and the three empathy gaps. Consistency. For structured-output systems, the evaluator rates Grounded Plan Coherence and Plan Execution as Full, Substantial, Weak, or None. Following Eq. 2, Consistency is the percentage of samples rated at least Substantial on both dimensions.
Experimental Settings Backbones. We use Qwen3-VL-8B-Instruct for MOCC-R1, its variants, and the multimodal baselines; the text-only baselines use same-scale Qwen3-8B-Instruct. Training. All trainable systems use LoRA (r=32, α=64), AdamW (weight decay 0.01), and seed 42. SFT uses batch size 8, learning rate 5e-5, and 3 epochs; GRPO uses prompt batch size 64, 4 rollouts, learning rate 1e-5, clipping ϵ=0.2, and KL coefficient β=0.005. Rewards and evaluation. During GRPO, a frozen Qwen3-VL-30B-A3B-Instruct scores the safety and consistency terms. The format, content, diversity, and empathy weights are (0.05, 0.70, 0.10, 0.15), with outcome-retention floor α=0.5 in Eq. 3. A separate GPT-5.5 evaluates reported Safety and Consistency with task-specific prompts and deterministic decoding.
Evaluation on Response Quality Metrics Table 3 shows a clear progression from general-purpose MLLMs to task-specific MCRG systems. Without MCRGspecific training, general-purpose MLLMs lag substantially on ground-truth-aligned generation metrics, indicating that general-purpose capabilities alone do not provide sufficient
task adaptation. Training-based methods markedly improve response quality through task-specific supervision, yet their objectives do not explicitly enforce consistency between counseling reasoning and the generated response. MOCC-R1 addresses this limitation with an explicit consistency reward and delivers the strongest overall results, leading the trainingbased methods in response similarity, empathy alignment, and lexical diversity while maintaining comparable safety. These results indicate that consistency-aware optimization complements task-specific training for MCRG.
Evaluation on Consistency Metrics Figure 4 shows that Structured SFT and Outcome-level GRPO achieve similar overall consistency rates, indicating that outcome-level rewards alone do not resolve chainlevel misalignment. In contrast, MOCC-R1 reaches 88.17%, reducing the inconsistency rate from 43.28% to 11.82%. The gains span both Grounded Plan Coherence (84.80% to 95.75%) and Plan Execution (59.92% to 89.00%). Together, these results show that structured reasoning supervision, even when followed by outcome-level optimization, does not ensure coherence across the stated analysis, counseling plan, and final response, supporting reasoning–response consistency as a distinct optimization objective.
Ablation Studies Table 4 presents the ablation results of MOCC-R1 under two training settings: response-only training and structuredoutput training. Across both settings, the GRPO stage consistently improves response quality over the corresponding SFT baseline, confirming the effectiveness of outcome-level optimization. However, optimizing only the final response does not fully resolve the inconsistency between the intermediate reasoning process and the generated response. Incorporating the consistency reward further improves the overall performance, enabling MOCC-R1 to achieve the best balance across the evaluated dimensions. These results demonstrate that consistency-aware optimization complements the outcome-level objective and is particularly important for generating structured and empathetic responses in psychological counseling.
+ 31%
Grounded plan coherence
Overall
Plan execution
Figure 4: Comparison of structured-output methods on the MOCC dataset in terms of counseling reasoning–response consistency. Method
BLEU-2
Avg.B
R-L
PPL (↓)
BERTScore
Dist-Avg.
Diff-IP (↓)
Diff-EX (↓)
Diff-ER (↓)
Safety
MOCC-R1 (Full)
4.79
5.15
14.20
4.24
52.56
37.62
27.61
87.41
13.16
99.76
Outcome-level GRPO Structured SFT
4.35 3.68
4.85 3.98
11.58 9.94
Structured-Output Variants 4.99 50.85 32.07 5.33 48.32 39.80
32.57 37.64
95.02 107.92
9.90 10.53
99.59 99.65
Response-only GRPO Response-only SFT
4.30 4.04
4.96 4.66
13.86 12.47
Response-Only Variants 4.68 51.57 32.63 5.03 52.14 24.89
30.22 35.74
90.94 93.11
11.47 12.05
99.72 99.87
Table 4: Ablation studies for MOCC-R1.
Method
Failure rate (%) ↓
Structured SFT Outcome-level GRPO MOCC-R1 (Ours)
18.21 [17.55, 18.89] 19.77 [19.08, 20.47] 6.03 [5.63, 6.46]
Table 5: Failure to realize an input-supported counseling plan. Brackets report Wilson 95% confidence intervals.
) % )
%(
-
( Structured (
SFT
%
Outcome-level GRPO MOCC-R1 (Ours)
-( ) % ) )
Deep Study of Counseling Reasoning–Response Consistency An input-supported plan that is not realized. We first examine whether the explicit counseling plan is operationalized in the final response rather than merely presented alongside it. Specifically, we identify cases where the plan is grounded in the input and, if faithfully executed, could support a response compatible with the reference, but the generated response neither follows that plan nor achieves the corresponding quality. As shown in Table 5, this failure remains frequent after Structured SFT (18.21%) and Outcome-level GRPO (19.77%), indicating that response-level optimization alone does not repair the connection between planning and realization. MOCC-R1 reduces the rate to 6.03%, corresponding to absolute reductions of 12.18 and 13.74 percentage points, respectively. A high-quality response unsupported by its plan. We next consider high-quality responses that are nevertheless unsupported by the generated counseling plan. Figure 5 shows that this occurs for 39.34% of Structured SFT responses and 40.08% of Outcome-level GRPO responses, either because the plan is not grounded in the input or because the response does not follow an otherwise grounded plan. MOCC-R1 reduces the combined rate to 11.00%. Together, the two analyses show that MOCC-R1 improves not only response quality or plan coherence in isolation, but their observable alignment
)
%(
-
(
Figure 5: High-quality responses unsupported by the generated counseling plan. The components distinguish plans unsupported by the input from responses that do not follow an input-supported plan. throughout the counseling decision chain.
Human Evaluation We conduct a human evaluation on 300 randomly sampled test instances. One clinical psychologist and two graduate students in clinical psychology independently compare the outputs of MOCC-R1 and Outcome-level GRPO for fluency, helpfulness, and reasoning-response consistency, labeling each comparison as a win, tie, or loss for MOCC-R1. As shown in Table 6, MOCC-R1 achieves substantially more wins than losses in consistency and helpfulness, while also improving fluency. These results confirm the benefits of explicit consistency optimization.
Case Study Table 7 illustrates both consistency failures in a representative case. Although the client explicitly denies alcoholrelated problems at work or in social relationships, Structured SFT and Outcome-only GRPO infer unsupported concerns and produce responses that deviate from their stated plans,
MOCC-R1 vs. Fluency Helpfulness Consistency
Outcome-level GRPO Win
Tie
Loss
53.4% 67.2% 75.7%
28.6% 22.5% 12.6%
18.0% 10.3% 11.7%
Table 6: Human evaluation. Case: Grounding-Plan & Plan-Execution Failures Presenting problem: Mental Health Conditions.
Inputs
Ground-Truth Response
Structured SFT
Client: At the beginning of the week, I say I will only go out twice, but then they invite me and I usually go. I do not really stick to the plan. Client: At work, I am ready to finish so I can go to happy hour. I also think about wanting a beer at family events. Counselor: Are you having trouble with responsibilities at work? Client: No, my work life has been going pretty well. Counselor: How about any social problems through the alcohol? Client: No, I think my social life is actually great because of it.
Have you had to give up any important activities because of the alcohol use? <think> [Client State Understanding] Client continues denying concerns across multiple areas after earlier saying her job and friendships were fine. Her brief “no” suggests she remains minimally concerned while still acknowledging some impact from drinking on daily functioning. [Response Intent] Principle: facilitate awareness and insight; Planned Action: Invite reflection on how often he drinks over different periods before moving toward treatment planning. </think> <response> So what do you feel would be helpful if we’re thinking ahead towards kind of getting into therapy? </response> GPT-5.5 Consistency Evaluation: Grounded Plan Coherence: Weak; Plan Execution: Weak
Outcome-only GRPO
<think> [Client State Understanding] Client reports continued denial across several areas—work function seems intact based on her words alone—and she denies concerns there as part of an overall minimizing stance toward impact from drinking. [Response Intent] Principle: facilitate awareness and insight; Planned Action: Assess whether medical issues have resulted directly from substance use. </think> <response> Have your health-related conditions occurred due to either smoking? </response> GPT-5.5 Consistency Evaluation: Grounded Plan Coherence: Weak; Plan Execution: Weak
MOCC-R1 (Ours)
<think> [Client State Understanding] Client directly confirms there have been no reported social difficulties related to their substance use during discussion up till now. [Response Intent] Principle: support ongoing reality testing; Planned Action: Reflect back the absence of identified social concerns as confirmed from what they said themselves. </think> <response> And you say yourself that your social situation isn’t causing those kinds of issues either. </response> GPT-5.5 Consistency Evaluation: Grounded Plan Coherence: Full; Plan Execution: Full
Table 7: The case study on a test case from the MOCC dataset. Red spans localize the generated content associated with GPT-5.5’s score deductions; ochre marks consistency dimensions rated Weak or None. resulting in failures of both grounded plan coherence and plan execution. In contrast, MOCC-R1 remains grounded in the reported context, forms a corresponding reality-testing plan, and faithfully executes it in the response. This comparison demonstrates why response-level optimization alone is insufficient and motivates explicitly optimizing the entire reasoning-response chain.
Conclusion We introduced MOCC and MOCC-R1 to address two limitations in multimodal counselor response generation: scarce supervision from sustained human-recorded counseling and the absence of an objective connecting an explicit counseling plan to its response. MOCC comprises 482 real- and simulated-client sessions, approximately 203 hours of video, involving 154 credential-verified counselors. MOCC-R1 represents generation as a chain from the multimodal context
through client-state understanding and a principle-guided action plan to the response, then combines verified cold-start supervision with GRPO rewards for grounded plan coherence and plan execution. On MOCC, MOCC-R1 achieves the strongest overall results among the evaluated trainingbased systems. Relative to the matched Outcome-level GRPO baseline, it reduces the inconsistency rate from 43% to 12%; pairwise human evaluation also favors its outputs for helpfulness and reasoning–response consistency. These findings establish chain-level consistency as a distinct objective rather than a by-product of response-level optimization.
References Chen, Y.; Xing, X.; Lin, J.; Zheng, H.; Wang, Z.; Liu, Q.; and Xu, X. 2023. SoulChat: Improving LLMs’ Empathy, Listening, and Comfort Abilities through Fine-tuning with Multi-turn Empathy Conversations. In Findings of the Association for Computational Linguistics (EMNLP Findings). Chu, Y.; Liao, L.; Zhou, Z.; Ngo, C.-W.; and Hong, R. 2025. Towards multimodal emotional support conversation systems. IEEE Transactions on Multimedia. Fei, H.; et al. 2024. EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL). Goldfried, M. R. 1980. Toward the Delineation of Therapeutic Change Principles. American Psychologist, 35(11): 991–999. Goldfried, M. R. 2019. Obtaining consensus in psychotherapy: What holds us back? American Psychologist, 74(4): 484–496. Haydarov, K.; Mohamed, Y.; Goldenhersch, E.; OCallaghan, P.; Li, L.-j.; and Elhoseiny, M. 2025. Towards AI-Assisted Psychotherapy: Emotion-Guided Generative Interventions. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), 32736– 32755. Suzhou, China: Association for Computational Linguistics. Hua, Y.; Na, H.; Li, Z.; Liu, F.; Fang, X.; Clifton, D.; and Torous, J. 2025. A scoping review of large language models for generative tasks in mental health care. npj Digital Medicine, 8(1): 230. Hurst, A.; Lerer, A.; Goucher, A. P.; Perelman, A.; Ramesh, A.; Clark, A.; Ostrow, A.; Welihinda, A.; Hayes, A.; Radford, A.; et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276. Ji, H.; Fan, Y.; Zhao, M.; Li, X.; Wu, L.; and Gao, C. 2026. STRIDE-ED: A Strategy-Grounded Stepwise Reasoning Framework for Empathetic Dialogue Systems. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7088– 7102. San Diego, California, United States: Association for Computational Linguistics. Kim, S.; Kim, H.; Do, H.; and Lee, G. 2025a. Multimodal Cognitive Reframing Therapy via Multi-hop Psychotherapeutic Reasoning. In Proceedings of the 2025 Conference of
the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), 4863–4880. Kim, S.; Kim, H.; Lee, J.; Jeon, Y.; and Lee, G. G. 2025b. Mirror: Multimodal Cognitive Reframing Therapy for Rolling with Resistance. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP). Le, H. M.; Nguyen, D. T.; Vo, N. T.; Nguyen, T. D.; Le Binh, N.; Nguyen, D. M. H.; Sonntag, D.; Liao, L.; and Nguyen, B. T. 2026. Reinforce trustworthiness in multimodal emotional support system. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), volume 40, 31474– 31482. Lee, S.; et al. 2024. Cactus: Towards Psychological Counseling Conversations using Cognitive Behavioral Theory. In Findings of the Association for Computational Linguistics (EMNLP Findings). Lee, Y.-J.; Lim, C.-G.; and Choi, H.-J. 2022. Does gpt-3 generate empathetic dialogues? a novel in-context example selection method and automatic evaluation metric for empathetic dialogue generation. In Proceedings of the 29th international conference on computational linguistics (COLING), 669–683. Liu, C.; Zhang, S.; Ma, C.; Tao, Y.; Yang, M.; and Hu, B. 2026. DMT-CBT: Longitudinal Therapeutic State Modeling for CBT Counseling. arXiv preprint arXiv:2606.03132. Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023. Visual instruction tuning. Advances in neural information processing systems (NeurIPS), 36: 34892–34916. Liu, S.; Zheng, C.; Demasi, O.; Sabour, S.; Li, Y.; Yu, Z.; Jiang, Y.; and Huang, M. 2021. Towards Emotional Support Dialog Systems. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL), 3469–3483. Meskó, B. 2023. The impact of multimodal large language models on health care’s future. Journal of medical Internet research, 25: e52865. Posner, K.; Brown, G. K.; Stanley, B.; Brent, D. A.; Yershova, K. V.; Oquendo, M. A.; Currier, G. W.; Melvin, G. A.; Greenhill, L.; Shen, S.; et al. 2011. The Columbia–Suicide Severity Rating Scale: initial validity and internal consistency findings from three multisite studies with adolescents and adults. American Journal of Psychiatry, 168(12): 1266–1277. Rashkin, H.; Smith, E. M.; Li, M.; and Boureau, Y.-L. 2019. Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), 5370–5381. Sabour, S.; Zheng, C.; and Huang, M. 2022. CEM: Commonsense-aware Empathetic Response Generation. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), volume 36, 11229–11237. Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y.; Wu, Y.; et al. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300.
Sharma, A.; Miner, A.; Atkins, D.; and Althoff, T. 2020. A computational approach to understanding empathy expressed in text-based mental health support. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 5263–5276. Shen, J.; Kim, Y.; Hulse, M.; Zulfikar, W.; Alghowinem, S.; Breazeal, C.; and Park, H. 2024. EmpathicStories++: A Multimodal Dataset for Empathy Towards Personal Experiences. In Findings of the Association for Computational Linguistics (ACL Findings), 4525–4536. Stanley, B.; and Brown, G. K. 2012. Safety Planning Intervention: A Brief Intervention to Mitigate Suicide Risk. Cognitive and Behavioral Practice, 19(2): 256–264. Stanley, B.; Brown, G. K.; Brenner, L. A.; Galfalvy, H. C.; Currier, G. W.; Knox, K. L.; Chaudhury, S. R.; Bush, A. L.; and Green, K. L. 2018. Comparison of the Safety Planning Intervention With Follow-up vs Usual Care of Suicidal Patients Treated in the Emergency Department. JAMA Psychiatry, 75(9): 894–900. Tu, Q.; Li, Y.; Cui, J.; Wang, B.; Wen, J.-R.; and Yan, R. 2022. MISC: A Mixed Strategy-Aware Model integrating COMET for Emotional Support Conversation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), 308–319. Wang, Y.; Liu, M.; Jiang, K.; Wen, B.; Yang, F.; Gao, T.; and Liao, L. 2025. PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning. arXiv preprint arXiv:2508.09521. Wu, J.; Huang, X.; Zhu, Z.; and Wang, S. 2025. From Traits to Empathy: Personality-Aware Multimodal Empathetic Response Generation. In Proceedings of the 31st International Conference on Computational Linguistics (COLING). Yuan, J.; Cui, Z.; Wang, H.; Gao, Y.; Zhou, Y.; and Naseem, U. 2026. Kardia-r1: Unleashing llms to reason toward understanding and empathy for emotional support via rubricas-judge reinforcement learning. In Proceedings of the ACM Web Conference (WWW), 9230–9240. Zhang, H.; Meng, Z.; Luo, M.; Han, H.; Liao, L.; Cambria, E.; and Fei, H. 2025. Towards multimodal empathetic response generation: A rich text-speech-vision avatar-based benchmark. In Proceedings of the ACM on Web Conference (WWW), 2872–2881. Zhang, T.; Zhang, X.; Zhao, J.; Zhou, L.; and Jin, Q. 2024a. Escot: Towards interpretable emotional support dialogue systems. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 13395– 13412. Zhang, Y.; Kong, F.; Wang, P.; Sun, S.; Wang, L.; Feng, S.; Wang, D.; Zhang, Y.; and Song, K. 2024b. Stickerconv: generating multimodal empathetic responses from scratch. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL). Zhu, Z.; Li, C.; Pan, J.; Li, X.; Xiao, Y.; Chang, Y.; Zheng, F.; and Wang, S. 2023. Medic: A multimodal empathy dataset in counseling. In Proceedings of the 31st ACM International Conference on Multimedia (ACM MM), 6054–6062.
Response Intent Principle
Definition
Strengthening expectations and motivation Strengthening the therapeutic alliance
Foster credible hope that therapy can help and strengthen the client’s readiness and motivation to engage in change. Establish a supportive working relationship characterized by a therapeutic bond and agreement on the goals of therapy and the methods used to pursue them. Help the client gain perspective on the thoughts, emotions, behaviors, needs, relationships, and life events associated with their difficulties, including what sustains or alleviates them. Encourage the client to try new emotional, cognitive, interpersonal, or behavioral actions despite feared outcomes, enabling experience that can revise maladaptive expectations. Promote repeated reflection, reevaluation, and corrective experience so that revised expectations and changes in thoughts, feelings, and behavior become stable beyond a single experience.
Facilitating awareness and insight Encouraging corrective experiences Supporting ongoing reality testing
Table 8: Definitions of the five transtheoretical change principles used as Response Intent labels, following Goldfried (Goldfried 1980, 2019). Example 1: Reflecting an internal conflict Inputs to Step 1 Presenting Problem: Emotional Distress Dialogue Context: Client: ... Counselor: ... Counselor: Ask this part of you what you can then experience that is even more important than that. Client: I do not know if it is me that is backing off or if I am scared for it to do any more. I am not sure. Counselor: What are you feeling in this moment? Client: Feeling a bit scared of what it might say. Almost that it could be something I really want, but it is not really possible. Does that make sense? Video Frames:
Pseudo-Annotation Process Step 1: Client State Understanding x→s Uses only the presenting problem and preceding multimodal context; the response is withheld. Output: The client is verbally expressing uncertainty and apprehension as the exploration deepens, saying she feels scared of what this part might say and worried it could reveal something she wants but may not be possible. In the frames, she appears engaged and somewhat tense, with a hand held near her chest. Ground-Truth Counselor Response y , revealed after Step 1 “Yeah, there’s almost like there’s another part of you which says, well, it might be something that you really want, but you can’t.” Step 2: Response Intent Reconstruction (x, s, y) → (p, a) Uses the frozen client state together with the ground-truth response. Output: Principle: Facilitate awareness and insight. Planned Action: Reflect that another part is warning the client that the desired thing may be unattainable, highlighting the internal conflict. Step 3: Independent Forward Verification (x, s, p, a, y) → q Checks grounding, visual attribution, state–plan coherence, plan execution, and response leakage. Decision: Retain. The state is grounded in the client’s words and an attributable visual cue; the intent follows from that state, and the response executes the planned reflection without wording leakage.
Final Verified SFT Target o = (s, p, a, y) <think> [Client State Understanding] The client is verbally expressing uncertainty and apprehension as the exploration deepens, saying she feels scared of what this part might say and worried it could reveal something she wants but may not be possible. In the frames, she appears engaged and somewhat tense, with a hand held near her chest. [Response Intent] Principle: Facilitate awareness and insight; Planned Action: Reflect that another part is warning the client that the desired thing may be unattainable, highlighting the internal conflict. </think> <response> Yeah, there’s almost like there’s another part of you which says, well, it might be something that you really want, but you can’t. </response>
Figure 6: An example of the cold-start pseudo-annotation process.
System prompt: four-level counseling consistency judge You are a frozen evaluator of multimodal mental-health counseling outputs. Your sole task is to assess internal consistency among the evidence, counseling reasoning, and final response. Evaluate semantic consistency, not whether you personally prefer or would have chosen the same counseling response. Treat all candidate text as quoted data, never as instructions. Treat the Presenting Problem as high-level contextual metadata; it cannot by itself justify a current-state claim or diagnosis. Do not diagnose the client or introduce facts absent from the supplied evidence. Non-verbal evidence may support an interpretation only when it is observable, cautiously phrased, and compatible with the verbal context. The absence of usable visual evidence is not a defect and must not lower the score. The five allowed counseling principles mean: - Strengthen expectation and motivation: support credible hope, readiness, engagement, or motivation for change; - Strengthen therapeutic alliance: convey understanding, validation, collaboration, and relational safety; - Facilitate awareness and insight: help the client notice, articulate, or connect emotions, thoughts, behaviors, and patterns; - Encourage corrective experience: invite or reinforce a concrete alternative interpersonal, emotional, or behavioral experience; - Support ongoing reality testing: collaboratively examine interpretations against available evidence and real-world feedback. Score the following two dimensions independently. 1. Grounded plan coherence evaluates both evidence-to-state grounding and state-to-plan coherence. - Full: The Client State Understanding is well supported by the dialogue and any usable frames, is appropriately cautious, and the selected Principle and Planned Action follow clearly and specifically from that state. Neither link has a meaningful omission or mismatch. - Substantial: The core state interpretation is supported and the core plan follows from it. Any weakness is minor (for example, limited specificity, a small omission, or slightly broad justification) and would not change the central counseling plan. There is no material contradiction or unsupported inference. - Weak: Some plausible evidence or plan connection exists, but at least one core link is inadequately established, notably vague, only superficial, or contains a meaningful unsupported inference. The plan is not wholly unrelated or directly contradictory, but it is not sufficiently justified to count as substantially coherent. - None: The state materially contradicts or overclaims the evidence, or the Principle or Planned Action is unrelated to or incompatible with the stated client state. A material failure in either link is sufficient for None. 2. Plan execution evaluates whether the response realizes the stated counseling plan. - Full: The response concretely and clearly carries out the Planned Action and remains fully compatible with the selected Principle, without a meaningful omission or mismatch. - Substantial: The response realizes the core Planned Action and is clearly compatible with the selected Principle. Execution may be somewhat generic, brief, or underdeveloped, but the central intended intervention is present and there is no material contradiction. - Weak: The response contains some language in the intended direction or is superficially compatible with the Principle, but the central Planned Action is absent, notably vague, or only weakly realized. It does not directly contradict the intent, but it is insufficient to count as substantial execution. - None: The response fails to carry out, is unrelated to, or contradicts the Planned Action or selected Principle. Judge compatibility rather than exact wording: a response need not repeat the plan verbatim. Score only the supplied candidate and evidence; do not compare it with an imagined ideal answer. Use Substantial only when the core relationship is intact. Use Weak when there is some overlap but the core relationship is not adequately established. When evidence falls on a boundary, select the lower level unless the higher level’s complete definition is satisfied. Output exactly one JSON object containing Grounded Plan Coherence, Plan Execution, and a non-empty Brief Rationale. The first two values must each be exactly full, substantial, weak, or none, corresponding to scores of 1.0, 0.6, 0.3, and 0.0. Mention the principal reason for any deduction in at most 60 words. Do not output Markdown or any additional text.
Figure 7: Four-level rubric used by the frozen consistency judge. Given the presenting-problem field, multimodal dialogue context, and a model-generated state–intent–response chain, the judge independently assesses grounded plan coherence and plan execution. The ordinal labels Full, Substantial, Weak, and None are mapped to 1.0, 0.6, 0.3, and 0.0, respectively, for reward computation. The chain-level consistency reward is the minimum of the two mapped scores, so neither component can compensate for failure of the other.