Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation
arXiv:2609.21683v1 [cs.AI] 18 Sep 2026
Yunji Chu Department of Artificial Intelligence, Sogang University, Seoul, Republic of Korea [email protected]
Abstract. Conversational speech depends on dialogue context and the listener’s immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance’s emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains essentially unchanged. Ablations show that temporal modeling performs best among the visual variants and that an explicit early-to-late difference is unnecessary; correct listener reactions also outperform cyclic mismatches on average. In a contextualappropriateness study with 20 speech researchers, 76% of judgments prefer Temporal, 9% Text-only, and 15% report no preference. We further connect the predicted response style to a Grad-TTS backbone for end-toend speech realization. Overall, the results support pre-response listener dynamics as complementary cues for conversational response planning. The source code is available at ReACT-TTS. Keywords: Conversational speech generation · Listener facial reaction · Response planning · Audio-visual learning · Expressive TTS
1
Introduction
A spoken response is shaped by both lexical content and interaction. Speakers monitor others and adjust vocal delivery accordingly; visible amusement, discomfort, or confusion can alter emotion, pitch, energy, timing, and speaking rate even when the next sentence is fixed. Conversational speech generation should therefore model how the immediately preceding interaction informs delivery. Conversational TTS conditions the current utterance on previous turns, speaker roles, acoustic context, or inferred affect. Dialogue history is useful, but it does not fully capture non-verbal listener behavior: a verbally neutral listener may still show surprise, skepticism, amusement, discomfort, or disengagement. Such cues can provide complementary evidence for selecting the delivery of the next response. Visual conditioning has also been explored in speech synthesis, primarily through images of the target speaker. In contrast, our visual input depicts the
2
Y. Chu
listener. Importantly, the listener reaction is not treated as an emotion label that should be copied to the target speaker. The appropriate response depends jointly on the target text and conversational context: the same visible reaction may invite an apology, reassurance, explanation, refusal, or playful continuation. Visual-Aware TTS (VA-TTS) [22] previously demonstrated that sequential listener feedback can condition speech synthesis. ReACT-TTS focuses on a complementary question: whether the listener’s facial behavior immediately before the target response can improve an explicit response-planning stage before acoustic generation. This factorization lets us measure the visual contribution independently of synthesis quality and then test whether the planned style can be connected to an end-to-end TTS system. We formulate listener-aware response planning using the previous three dialogue turns, the target response text, and a one-second listener face sequence immediately preceding the target response. The listener sequence is encoded temporally and fused with linguistic context through target-conditioned gating. Rather than making an explicit frame-difference feature central to the method, our final model relies on temporal encoding itself; a controlled ablation evaluates whether explicit difference features add useful information. The revised experimental protocol addresses three questions. First, does temporal listener conditioning provide complementary information over text context alone? Second, does the model depend on the correct listener reaction rather than generic visual input? Third, can the resulting response representation be connected to speech generation without changing the underlying text or target speaker? We answer these questions using a strictly filtered MELD protocol with fixed train/dev/test criteria, multi-seed evaluation, a correct-versus-mismatched listener intervention, and end-to-end Grad-TTS generation. The main contributions are: – A response-planning formulation that uses the listener’s one-second pre-response temporal facial reaction as complementary context for conversational speech generation, rather than directly mirroring listener emotion. – A strict dyadic MELD protocol yielding 1,117/116/261 train/dev/test responseplanning samples, with all filtering thresholds determined without test-set tuning. – Multi-seed controlled experiments showing a positive tendency from temporal listener conditioning, including static-versus-temporal and explicit-difference ablations, plus a mismatched-listener intervention that probes listener-specific information use. – A perceptual contextual-appropriateness study with 20 speech researchers, together with end-to-end Grad-TTS evaluation on 252 test utterances, without claiming superior acoustic quality. Code is publicly available on GitHub at ReACT-TTS_public.
Response Planning from Listener Facial Reactions
2
Related Work
2.1
Face-Conditioned and Visually Guided Speech Synthesis
3
Visual conditioning has been explored in TTS primarily to infer the voice or expressive characteristics of the target speaker. Face-TTS learns a face-styled diffusion model with a cross-modal biometric objective [9]; Face-StyleSpeech separates face-derived speaker information from residual prosody [7]; and FEIMTTS combines facial representation with explicit emotion-intensity control [1]. More recent systems broaden visual conditioning to diverse portrait styles or fullface expressive cues. FaceSpeak suppresses visually irrelevant appearance factors while extracting identity and emotion information [21], whereas AVLM integrates full-face visual representations into an expressive speech language model for emotion-aware generation [18]. These methods primarily use the depicted face to characterize the target speaker or the observed speaker. ReACT-TTS instead uses the listener’s reaction to plan the target speaker’s next delivery. The closest prior work is Visual-Aware TTS (VA-TTS) [22], which conditions speech synthesis on sequential visual feedback from a listener. ReACT-TTS builds on this formulation but replaces direct visual–acoustic fusion with an explicit response-planning stage that combines dialogue semantics with temporal facial dynamics. A mismatched-listener counterfactual intervention further tests whether response planning actually depends on the observed listener. 2.2
Conversational Speech Synthesis
Conversational speech synthesis generates a target utterance using preceding turns rather than treating each sentence independently. DailyTalk introduced a spoken-dialogue corpus and contextual TTS baseline [10]. Later systems model intra- and inter-modal context interaction [5], fine-grained semantic and prosodic graphs [6], and diffusion-based context-aware prosody [20]. Recent work also makes emotion inference more explicit: prompt-guided emotive TTS derives contextual emotion tags and localized acoustic cues [4], while Chain-Talker separates emotion understanding, semantic understanding, and empathetic rendering [2]. UniTalker extends contextual synthesis toward joint speech and talking-face generation [3]. These studies establish the value of dialogue history, but they do not explicitly evaluate the immediately pre-response listener reaction as a separate signal for planning the target speaker’s delivery. 2.3
Listener-Reaction Modeling in Dyadic Interaction
A complementary line of work generates listener behavior from a speaker’s verbal and non-verbal signals. ReactFace models multiple appropriate and synchronized facial reactions instead of a single deterministic response [13]. The REACT 2025 challenge further formalizes the one-to-many nature of listener reactions and introduces the large-scale MARS benchmark for dyadic interaction [17]. ReactDiff combines multimodal interaction modeling with latent diffusion to
4
Y. Chu
generate diverse, contextually appropriate facial reactions [11]. Although these methods generate the listener rather than the next speaker’s voice, they support treating listener behavior as a dynamic, one-to-many contextual signal rather than a label that should be directly mirrored. 2.4
Multimodal Emotion in Conversation
MELD contains approximately 13,000 utterances from 1,433 multi-party dialogues with aligned text, audio, video, speaker, emotion, and sentiment annotations [16]. We use it as a source of audiovisual conversational sequences rather than as a standard utterance-level emotion-recognition benchmark. Because MELD does not explicitly annotate the listener or addressee for every turn, the proposed subset retains only clips with a reliable non-speaking face track and unambiguous local interaction structure.
3
Method
3.1
Task Definition
At dialogue turn t, the target speaker produces utterance xt after the previous three dialogue turns \mathcal {H}_t=\{(x_i,s_i)\}_{i=t-3}^{t-1}, \label {eq:history} (1) where xi and si denote transcript and speaker identity. In addition to the target text xt , the response planner observes a listener-face sequence V_t^L=\{I_n\}_{n=1}^{N_t}, \qquad 8\le N_t\le 16,
(2)
formed from the valid face tracks among 16 requested frames in the one-second interval immediately preceding the target response. This timing is central to our formulation: the visual input is evidence available before the target speaker begins speaking, rather than a visual trace of the target utterance itself. The goal is to first infer an appropriate response style and then realize the target speech: z_t^{\mathrm {style}} &= P(x_t,\mathcal {H}_t,V_t^L), \\ \hat a_t &= S(x_t,z_{s_t},z_t^{\mathrm {style}}), (4) where zst denotes the target-speaker representation. The response target describes the target utterance, not the listener’s emotion. Consequently, listener behavior is used as contextual evidence rather than as a label to be mirrored. 3.2
Interaction Context Encoder
The dialogue history is serialized with explicit speaker-role tokens: q_t=[\mathrm {SPK}_{t-3}]x_{t-3}\cdots [\mathrm {SPK}_{t-1}]x_{t-1}[\mathrm {TARGET}]x_t.
(5)
Response Planning from Listener Facial Reactions Inputs
Stage A: response planning
Previous 3 turns + target response text
Listener facial reaction 1-s pre-response window 16 requested frames
Stage C: speech realization
Interaction Context Encoder (RoBERTa-base)
Listener Reaction Encoder AffectNet ResNet-18 + Transformer encoder
5
ResponseStyleAdapter ztac 256-d
Target/context-conditioned Gated Response Planner
Predicted style
htfused = httext + gt Wrhtreact
256-d
ztstyle
Concatenate [zst; ztac] 512-d
Grad-TTS 128-bin mel
Stage A is frozen during Stage C
Speaker reference audio (train-split reference)
Resemblyzer Speaker Encoder
Speaker embedding zst 256-d
HiFi-GAN 16 kHz
Generated waveform
Fig. 1: Overall architecture of ReACT-TTS. Dialogue context and target text are combined with the listener facial reaction from the 1-s pre-response window to predict a 256-dimensional response-style representation. During speech realization, a train-split speaker-reference waveform is encoded with Resemblyzer [19], while the predicted style is mapped by the ResponseStyleAdapter; the two 256-dimensional vectors are concatenated to condition Grad-TTS, and HiFi-GAN converts the generated 128-bin mel-spectrogram to waveform speech.
A pretrained language encoder produces token representations that are pooled into htext . Including the target text is important because the same listener t reaction can imply different appropriate deliveries depending on whether the upcoming response is, for example, an apology, explanation, reassurance, or playful continuation. 3.3
Temporal Listener Reaction Encoder
Each valid frame is encoded by an expression-oriented facial backbone, f_n=E_{\mathrm {face}}(I_n), \qquad n=1,\ldots ,N_t,
(6)
implemented with a ResNet-18 pretrained on AffectNet [14]. Thus, the visual backbone is not a face-recognition network trained solely for identity discrimination. The resulting sequence is passed to a Transformer temporal encoder and pooled into h_t^{\mathrm {react}}=\mathrm {Pool}\!\left (E_{\mathrm {temp}}(f_1,\ldots ,f_{N_t})\right ). \label {eq:reaction} (7) Our primary model uses this temporal representation directly. To test whether explicitly exposing coarse early-to-late change is necessary, we additionally evaluate an explicit-difference variant. Let E and L denote the earlier and later valid-frame groups within the same one-second window. We compute h^{\mathrm {early}} &= \frac {1}{|\mathcal E|}\sum _{n\in \mathcal E} f_n, & h^{\mathrm {late}} &= \frac {1}{|\mathcal L|}\sum _{n\in \mathcal L} f_n,\\ \Delta h &= h^{\mathrm {late}}-h^{\mathrm {early}}, (9)
6
Y. Chu
and concatenate ∆h to the temporal representation. This feature is treated as an ablation rather than as the core novelty of ReACT-TTS; Sec. 4.4 shows that it provides no additional benefit over temporal encoding alone. 3.4
Target-Conditioned Response Planning
Visual evidence should not affect every target response equally. ReACT-TTS therefore computes a target-conditioned gate g_t=\sigma \!\left (W_g[h_t^{\mathrm {text}};h_t^{\mathrm {react}}]+b_g\right )
(10)
and fuses linguistic and visual evidence through h_t^{\mathrm {fused}}=h_t^{\mathrm {text}}+g_t\odot W_r h_t^{\mathrm {react}}. \label {eq:fusion}
(11)
The residual text path preserves the linguistic interpretation when visual evidence is weak, while the gate permits sample-dependent visual modulation. Because htext jointly encodes the dialogue history and target response, the gate determines t how much listener evidence is useful for the specific upcoming utterance. From the fused representation, Stage A predicts the target-response emotion, continuous affect, and prosodic attributes, \hat {\mathbf y}_t &= \mathrm {softmax}(W_e h_t^{\mathrm {fused}}+b_e),\\ \hat {\mathbf v}_t &= W_v h_t^{\mathrm {fused}}+b_v,\\ \hat {\mathbf p}_t &= W_p h_t^{\mathrm {fused}}+b_p, (14) where ŷt is the target-emotion distribution, v̂t ∈ R3 denotes valence–arousal– dominance (VAD), and \hat {\mathbf p}_t= [\widehat {\mu }_{\log F0},\widehat {\sigma }_{\log F0},\widehat {E}_{\log },\widehat {R}_{\mathrm {phn}}]
(15)
contains utterance-level mean log-F0, log-F0 standard deviation, mean log-energy, and phoneme rate. The prosodic targets are speaker-normalized as in the responseplanning setup. The planner additionally exposes the 256-dimensional responsestyle embedding ztstyle defined in Eq. (3), which is the representation passed to Stage C. In all listener-removal baselines, the textual path and prediction heads are kept unchanged; only the listener visual input is removed. 3.5
Speech Realization
The speech realization stage uses the same acoustic backbone for the Temporal and Text-only systems. The frozen Stage-A planner produces a 256-dimensional style embedding, which is mapped by a trainable ResponseStyleAdapter into the 256-dimensional acoustic-style space. Separately, a training-speaker reference waveform is encoded with Resemblyzer [19] to obtain a 256-dimensional speaker
Response Planning from Listener Facial Reactions
7
embedding. The style and speaker vectors are concatenated to form a 512dimensional global condition for Grad-TTS [15]: z_t^{\mathrm {ac}} &= A_{\mathrm {style}}(z_t^{\mathrm {style}}),\\ c_t &= [z_{s_t};z_t^{\mathrm {ac}}],\\ \hat m_t &= G_{\mathrm {Grad\text {-}TTS}}(x_t,c_t),\\ \hat a_t &= V_{\mathrm {HiFi\text {-}GAN}}(\hat m_t).
(19) The acoustic configuration uses 16-kHz audio, a 1024-point FFT, hop size 160, window size 1024, and 128 mel bins with fmin = 0 and fmax = 8 kHz. Monotonic alignment search is used during acoustic training, and a 16-kHz HiFi-GAN vocoder [8] converts the generated mel-spectrogram into waveform speech. 3.6
Training Strategy
Training is separated into three stages. Stage A trains the response planner. Stage B trains the Grad-TTS acoustic model on the larger MELD acoustic training set using ground-truth style supervision. Stage C freezes Stage A and trains the ResponseStyleAdapter jointly with the acoustic objective so that predicted response style is mapped into the Stage-B acoustic-style space. For Stage C, the adapter output ztac is additionally aligned to the 256dimensional Stage-B ground-truth emotion embedding ztGT using an equal mixture of mean-squared error and cosine distance, \mathcal {L}_{\mathrm {align}}=0.5\,\mathcal {L}_{\mathrm {MSE}}(z_t^{\mathrm {ac}},z_t^{\mathrm {GT}}) +0.5\,\mathcal {L}_{\mathrm {cos}}(z_t^{\mathrm {ac}},z_t^{\mathrm {GT}}),
(20)
where Lcos = 1 − cos(·, ·) denotes cosine distance. The complete Stage-C objective is \mathcal {L}_{C}=\mathcal {L}_{\mathrm {Grad\text {-}TTS}}+\lambda _{\mathrm {align}}\mathcal {L}_{\mathrm {align}},\qquad \lambda _{\mathrm {align}}=0.5. \label {eq:stagec_loss} (21) Ground-truth emotion information is used only as a training target for this alignment and is not available at inference time. The Text-only baseline uses the identical Stage-B/Stage-C acoustic pipeline and differs only by removing listener visual input from Stage A.
4
Experiments
4.1
Dataset and Strict Listener Protocol
Experiments use MELD [16], which provides aligned text, audio, video, speaker identity, emotion, and sentiment annotations for multi-party conversations. MELD does not annotate an explicit addressee for every turn, so listener assignment is an important source of uncertainty. We therefore use a precision-oriented protocol and apply the same criteria to train, development, and test data. A response-planning sample is retained only when the local dialogue contains exactly two speakers, three preceding turns are available, and the target utterance
8
Y. Chu
Table 1: Final MELD protocol. Response-planning samples satisfy the same strict filtering criteria in all splits. Speech-generation evaluation additionally requires a target speaker with a training-split reference. Split Response planning Speech generation Train Dev Test
1,117 116 261
(a) Pre-response listener reaction window -1.0 s
– 115 252
(b) Visual representation variants target onset (0 s)
Static (no temporal encoder)
Listener reaction window
static representation htstatic
Eface
single frame
time
Temporal (primary)
shared Eface
Temporal + (ablation)
shared Eface
Temporal Transformer + Pool
htreact
16 requested frames
Validity filter exactly 2 speakers target duration 1 s listener visibility 0.30 valid frames 8/16 identity similarity 0.65 target speaker via mouth motion
Temporal + Pool early/late means
concat
h
Only the visual representation changes; the downstream gated response planner is shared.
Fig. 2: Listener-reaction timing and representation variants. (a) Visual evidence is restricted to the 1-s interval immediately before target onset and retained only after the strict dyadic/track-validity filter. (b) The controlled ablation changes only the visual representation supplied to the same downstream gated planner: a static representation without temporal encoding, the full temporal representation, or the temporal representation augmented with an explicit early-to-late difference.
lasts at least one second. The listener reaction is restricted to the one-second interval immediately before the target response. Sixteen frames are requested from this interval. We require listener visibility of at least 0.30, reaction validity of at least 0.50 (at least 8 of 16 requested frames), and face-track identity similarity of at least 0.65. The visually active target speaker is determined using mouth motion, and the remaining tracked dialogue participant is treated as the listener. All thresholds were fixed using the training/development protocol and were not tuned on test performance. This procedure yields 1,117 training, 116 development, and 261 test samples for response planning (Table 1), substantially expanding the 102-sample preliminary subset while retaining strict listener filtering. For speech generation, evaluation is further restricted to target speakers for whom a training-split speaker reference is available, leaving 115 development and 252 test utterances. This exclusion is determined by training-reference availability rather than by test-set synthesis quality.
Response Planning from Listener Facial Reactions
9
Table 2: Stage-A response-planning results on the 261-sample test set over ten seeds (42–51), reported as mean ± standard deviation. Method
Accuracy ↑
Macro-F1 ↑
CCC ↑
Text-only 0.4521±0.0316 0.2489±0.0152 0.2173±0.0463 Temporal listener 0.4483±0.0243 0.2582±0.0131 0.2282±0.0291
4.2
Implementation and Evaluation Protocol
The Interaction Context Encoder is initialized from RoBERTa-base [12]. The facial backbone is a ResNet-18 pretrained on AffectNet [14], followed by a twolayer Transformer temporal encoder with hidden dimension 256, four attention heads, and dropout 0.1. Stage-A response planning is evaluated using targetemotion accuracy, macro-F1, and VAD concordance correlation coefficient (CCC), averaged across valence, arousal, and dominance. Because class frequencies are highly imbalanced, macro-F1 is treated as the primary classification measure. For the primary comparison, Text-only and Temporal listener models are evaluated on the fixed 261-sample test set over ten random seeds (42–51). We additionally perform 10,000 bootstrap resamples over test samples to characterize the uncertainty of the paired Temporal-minus-Text difference. The ablation study uses the same five seeds (42–46) for all configurations. For speech realization, Stage B is trained on 9,988 MELD training utterances (one corrupt media item removed), with seen-speaker development evaluation on 1,071 utterances. The best Stage-B initialization is selected on development loss (seed 42, epoch 23). Stage C freezes Stage A and trains the responsestyle adapter; the final Temporal and Text-only checkpoints are independently selected by development loss. All reported speech-generation test results use these development-selected checkpoints. 4.3
Response-Planning Results
Table 2 reports the ten-seed test results. Temporal listener conditioning does not improve raw accuracy (0.4483 vs. 0.4521), but it yields higher mean macro-F1 (0.2582 vs. 0.2489) and CCC (0.2282 vs. 0.2173). The paired macro-F1 difference is +0.0093 ± 0.0184, with Temporal outperforming Text-only in 7 of 10 seeds. Bootstrap resampling gives a macro-F1 difference of +0.0094 with a 95% interval of [−0.0057, +0.0247] and P (∆ > 0) = 0.8844. The corresponding accuracy difference is −0.0040 with a 95% interval of [−0.0226, +0.0149]. Because the intervals include zero, we do not interpret these results as statistically significant improvement. Instead, they indicate a modest positive tendency in macro-F1 and VAD concordance, while overall accuracy remains essentially unchanged. An exploratory class-wise analysis suggests that the contribution of listener information is not uniform across emotions. The largest mean macro-F1 differences occur for surprise (+0.0391), sadness (+0.0225), and anger (+0.0158); Temporal
10
Y. Chu
Table 3: Five-seed fair ablation (seeds 42–46). The temporal model without an explicit difference feature performs best in Macro-F1. Configuration
Macro-F1 ↑
Text-only 0.2367 ± 0.0021 Static face 0.2429 ± 0.0272 Temporal + explicit ∆ 0.2407 ± 0.0191 Temporal, no ∆ 0.2558 ± 0.0135
wins 9/10 seeds for surprise and 7/10 for sadness. Neutral is slightly lower and joy is approximately unchanged. Fear and disgust have only 7 and 8 test examples, respectively, and are therefore not interpreted individually. This pattern is consistent with class-dependent complementarity rather than a universal benefit from visual conditioning. 4.4
Static, Temporal, and Explicit-Difference Ablation
To isolate the contribution of static appearance, temporal modeling, and the explicit difference feature, we conduct a controlled five-seed ablation using the representation variants in Fig. 2(b). Table 3 reports the results. A static face representation improves modestly over Text-only on average, but the strongest result is obtained by the Temporal model without an explicit ∆ feature (0.2558 macro-F1). Adding the explicit difference reduces the mean to 0.2407. These results shift the interpretation of the method away from hand-crafted embedding differences. The AffectNet-pretrained backbone provides expressionoriented frame features, while the temporal encoder models their evolution directly. The explicit early-to-late difference does not provide additional benefit in our setting, so the no-∆ Temporal configuration is used as the primary ReACT-TTS planner. 4.5
Mismatched-Listener Counterfactual Intervention
To test whether performance arises from listener-specific information rather than generic visual regularization, we perform a counterfactual input intervention at test time. For each batch, the correct listener sequence is replaced by another sample’s listener sequence using within-batch cyclic reassignment; dialogue history, target text, and all non-visual inputs remain unchanged. This deterministic reassignment avoids tuning a favorable mismatch for individual examples. Table 4 shows that the correct listener yields 0.2582 macro-F1 compared with 0.2526 for the mismatched condition, a mean difference of approximately +0.0055. Across seeds, the correct listener wins/ties/loses in approximately 6/2/2 cases. Test-sample bootstrap gives a 95% interval of [−0.0027, +0.0135] and P (∆ > 0) = 0.911. The interval again includes zero, so this experiment does not establish a significant or causal advantage. It nevertheless provides complementary evidence that
Response Planning from Listener Facial Reactions
11
Table 4: Correct-versus-mismatched listener intervention over ten seeds. Mismatched reactions are produced by within-batch cyclic reassignment while text context and target response remain fixed. Listener input
Macro-F1 ↑
Mismatched listener 0.2526 ± 0.0117 Correct listener 0.2582 ± 0.0131 Table 5: Objective speech-generation evaluation on 252 test utterances at length scale 1.5, selected on the development split. Lower WER/CER is better; higher speaker similarity is better. Condition
WER ↓ CER ↓ Spk. Sim. ↑
Text-only 1.0320 0.8824 Temporal listener 1.0384 0.8967
0.5976 0.6059
the planner is sensitive to which listener reaction is paired with the conversation: replacing the reaction while keeping the linguistic response fixed tends to reduce macro-F1. 4.6
End-to-End Speech Realization
We next evaluate whether the predicted response representation can be connected to an end-to-end speech generator. The frozen Stage-A style representation is mapped through the ResponseStyleAdapter and combined with a Resemblyzer target-speaker embedding to condition Grad-TTS. Temporal and Text-only systems use the same acoustic backbone and differ only in the Stage-A visual input. Initial synthesis at the default length scale 1.0 produced speech that was substantially shorter than the reference (mean generated duration about 1.58– 1.59 s versus 3.55 s for the reference). We therefore selected the inference length scale on the 115-sample development set, without consulting test results. Among scales 1.5, 1.75, 2.0, and 2.25, scale 1.5 gave the lowest development WER (approximately 1.03) and was fixed for final inference, despite longer scales producing duration ratios closer to one. Table 5 reports objective evaluation on 252 test utterances. Temporal conditioning produces slightly higher speaker similarity (0.6059 vs. 0.5976), while Text-only is slightly better in WER (1.0320 vs. 1.0384) and CER (0.8824 vs. 0.8967). Thus, the generated-speech experiment demonstrates that the response plan can be carried through the acoustic pipeline, but it does not support a claim that listener conditioning improves intelligibility or overall acoustic quality. As a diagnostic, we also synthesized speech from Stage B using ground-truth emotion/style conditioning. Intelligibility remained poor, indicating that the high WER/CER cannot be attributed solely to the Stage-C response-style adapter.
12
Y. Chu
We therefore treat the acoustic backbone and duration modeling as important limitations of the current realization stage rather than using synthesis quality as evidence for the response-planning claim. Perceptual contextual appropriateness. We additionally conduct a preference study with 20 speech researchers holding an M.S. degree or higher in AI-related fields. Given the preceding conversation, target text, and pre-response listener reaction, evaluators compared anonymized Speech A/B realizations of the same text and selected which delivery better fit the listener reaction and conversational context, with a no-preference option. Across collected judgments, 76% preferred Temporal, 9% Text-only, and 15% reported no preference/equal appropriateness. This evaluates contextual appropriateness rather than general naturalness, which remains confounded by the shared acoustic limitations in Table 5.
5
Discussion and Limitations
Across ten seeds, Temporal listener conditioning shows a positive tendency in macro-F1 and CCC, but the paired bootstrap interval includes zero and raw accuracy is essentially unchanged. We therefore interpret pre-response listener behavior as a complementary, class-dependent planning cue rather than a universal improvement. The five-seed ablation further localizes the useful visual signal: an AffectNet-pretrained, expression-oriented backbone with temporal encoding achieves the highest mean macro-F1, whereas the explicit early-to-late ∆ adds no benefit. Thus, temporal listener dynamics—not an engineered difference feature— form the core visual contribution. Listener assignment remains imperfect because MELD lacks explicit addressee annotations. The two-speaker, mouth-motion, visibility/validity, and identity filters reduce ambiguity but cannot guarantee communicative intent and may favor clearly visible faces. The mismatched-listener intervention provides complementary evidence that the correct sequence is more useful on average, but its confidence interval also includes zero; it is a controlled network-input intervention, not evidence of a causal mechanism in human communication. Explicit dyadic addressee annotations and richer action-unit or expression trajectories are important future tests. Speech realization remains the main practical limitation. Both systems have high WER/CER, and the weak Stage-B oracle diagnostic points to the acoustic backbone and duration model rather than listener conditioning alone. Accordingly, generation is used to establish feasibility, not superior synthesis quality. The researcher preference study nevertheless shows a clear descriptive preference for Temporal contextual appropriateness; it does not evaluate naturalness, which requires a stronger synthesis backbone and dedicated study. Direct numerical comparison with prior conversational or visually conditioned TTS remains difficult because datasets and conditioning definitions differ. Evaluation on shared conversational benchmarks remains important future work.
Response Planning from Listener Facial Reactions
6
13
Conclusion
We presented ReACT-TTS, which uses the listener’s one-second pre-response facial behavior as complementary context for explicit response planning. On a strict dyadic MELD protocol, Temporal conditioning shows a modest positive tendency in macro-F1 and CCC and achieves the highest mean macro-F1 among the controlled static and explicit-∆ variants. Correct-versus-mismatched input results further suggest listener-specific information use, while contextual-appropriateness judgments favor Temporal delivery. The planned representation can condition Grad-TTS end to end, although current synthesis quality remains limited. These results motivate pre-response listener dynamics as a useful planning signal while separating that question from acoustic-generation quality.
References 1. Chu, Y., Shim, Y., Park, U.: Facial expression-enhanced tts: Combining face representation and emotion intensity for adaptive speech. In: Computer Vision – ECCV 2024 Workshops. Lecture Notes in Computer Science, vol. 15637, pp. 117–129. Springer (2025). https://doi.org/10.1007/978-3-031-91581-9_9 2. Hu, Y., Liu, R., Ren, Y., Yin, X., Li, H.: Chain-talker: Chain understanding and rendering for empathetic conversational speech synthesis. In: Findings of the Association for Computational Linguistics: ACL 2025. pp. 1988–2003. Association for Computational Linguistics (2025). https://doi.org/10.18653/v1/2025.findingsacl.101 3. Hu, Y., Liu, R., Ren, Y., Yin, X., Li, H.: Unitalker: Conversational speech-visual synthesis. In: Proceedings of the 33rd ACM International Conference on Multimedia. pp. 10248–10257 (2025). https://doi.org/10.1145/3746027.3755502 4. Jeon, Y., Kim, Y., Lee, J., Lee, G.: Prompt-guided selective masking loss for contextaware emotive text-to-speech. In: Findings of the Association for Computational Linguistics: NAACL 2025. pp. 638–650. Association for Computational Linguistics (2025). https://doi.org/10.18653/v1/2025.findings-naacl.38 5. Jia, Z., Liu, R.: Intra- and inter-modal context interaction modeling for conversational speech synthesis. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 1–5 (2025). https://doi.org/10.1109/ ICASSP49660.2025.10890216 6. Jia, Z., Liu, R., Sisman, B., Li, H.: Multimodal fine-grained context interaction graph modeling for conversational speech synthesis. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. pp. 8852–8858. Association for Computational Linguistics (2025). https://doi.org/10.18653/ v1/2025.emnlp-main.448 7. Kang, M., Han, W., Yang, E.: Face-stylespeech: Enhancing zero-shot speech synthesis from face images with improved face-to-speech mapping. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 1–5 (2025). https://doi.org/10.1109/ICASSP49660.2025.10889667 8. Kong, J., Kim, J., Bae, J.: Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. In: Advances in Neural Information Processing Systems. vol. 33 (2020)
14
Y. Chu
9. Lee, J., Chung, J.S., Chung, S.W.: Imaginary voice: Face-styled diffusion model for text-to-speech. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 1–5 (2023). https://doi.org/10.1109/ICASSP49357. 2023.10094745 10. Lee, K., Park, K., Kim, D.: Dailytalk: Spoken dialogue dataset for conversational text-to-speech. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 1–5 (2023). https://doi.org/10.1109/ICASSP49357. 2023.10095751 11. Li, J., Wang, S., Wang, X., Zhu, Y., Xiong, H., Zhuang, Z., Wang, Q.: Reactdiff: Latent diffusion for facial reaction generation. Neural Networks 189, 107596 (2025). https://doi.org/10.1016/j.neunet.2025.107596 12. Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V.: RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692 (2019) 13. Luo, C., Song, S., Xie, W., Spitale, M., Ge, Z., Shen, L., Gunes, H.: Reactface: Online multiple appropriate facial reaction generation in dyadic interactions. IEEE Transactions on Visualization and Computer Graphics 31(9), 6190–6207 (2025). https://doi.org/10.1109/TVCG.2024.3490613 14. Mollahosseini, A., Hasani, B., Mahoor, M.H.: Affectnet: A database for facial expression, valence, and arousal computing in the wild. IEEE Transactions on Affective Computing 10(1), 18–31 (2019). https://doi.org/10.1109/TAFFC.2017. 2740923 15. Popov, V., Vovk, I., Gogoryan, V., Sadekova, T., Kudinov, M.: Grad-tts: A diffusion probabilistic model for text-to-speech. In: Proceedings of the 38th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 139, pp. 8599–8608. PMLR (2021) 16. Poria, S., Hazarika, D., Majumder, N., Naik, G., Cambria, E., Mihalcea, R.: MELD: A multimodal multi-party dataset for emotion recognition in conversations. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. pp. 527–536. Association for Computational Linguistics (2019). https: //doi.org/10.18653/v1/P19-1050 17. Song, S., Spitale, M., Kong, X., Zhu, H., Luo, C., Palmero, C., Barquero, G., Escalera, S., Valstar, M., Daoudi, M., Baur, T., Ringeval, F., Howes, A., André, E., Gunes, H.: REACT 2025: The third multiple appropriate facial reaction generation challenge. In: Proceedings of the 33rd ACM International Conference on Multimedia. pp. 13979–13984 (2025). https://doi.org/10.1145/3746027.3762244 18. Tan, W., Lian, J., Inaguma, H., Tomasello, P., Koehn, P., Ma, X.: Seeing is believing: Emotion-aware audio-visual language modeling for expressive speech generation. In: Findings of the Association for Computational Linguistics: EMNLP 2025. pp. 2600–2617. Association for Computational Linguistics (2025). https: //doi.org/10.18653/v1/2025.findings-emnlp.140 19. Wan, L., Wang, Q., Papir, A., Lopez Moreno, I.: Generalized end-to-end loss for speaker verification. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 3561–3565 (2018). https://doi.org/10.1109/ ICASSP.2018.8462665 20. Wu, W., Lin, Z., Zhou, Y., Li, J., Niu, R., Wu, Q., Cao, S., Ma, L., Wu, Z.: Diffcss: Diverse and expressive conversational speech synthesis with diffusion models. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 1–5 (2025). https://doi.org/10.1109/ICASSP49660.2025.10890208
Response Planning from Listener Facial Reactions
15
21. Zhang, T.H., Zhang, J., Wang, J., Qian, X., Yin, X.C.: Facespeak: Expressive and high-quality speech synthesis from human portraits of different styles. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 39, pp. 25922– 25930 (2025). https://doi.org/10.1609/aaai.v39i24.34786 22. Zhou, M., Bai, Y., Zhang, W., Yao, T., Zhao, T., Mei, T.: Visual-aware text-tospeech. In: IEEE International Conference on Acoustics, Speech and Signal Processing. pp. 1–5 (2023). https://doi.org/10.1109/ICASSP49357.2023.10095084