Conceptio › Archive › arXiv CS
arXiv CSopen access

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

1

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation

arXiv:2605.16026v1 [cs.CL] 15 May 2026

Yu Pan, Yang Hou, Xiongfei Wu, Liang Zhang, Yves LE TRAON, Lei Ma, Jianjun Zhao

Abstract—Compositional speech-to-speech translation (S2ST) systems built upon speech large language models (SpeechLLMs) have recently shown promising performance. However, existing S2ST systems often either neglect source-language information or encode it through a language-as-label paradigm, representing each source language as an independent flat embedding. Such a design overlooks systematic linguistic structure shared across languages, which may limit data-efficient multilingual adaptation when supervised S2ST data are scarce. To address this issue, we propose S2ST-Omni 2, a many-to-one compositional S2ST framework that systematically reformulates multilingual language conditioning from flat language labels to structured typological priors. Specifically, S2ST-Omni 2 revisits language conditioning at three levels: typology-informed hierarchical language encoding for structured source-language representation, dynamicallygated language-aware Dual-CTC for content-adaptive acoustic modulation, and typology-aware LLM prompting for decoderside linguistic guidance. Experiments on CVSS-C show that S2ST-Omni 2 achieves superior average performance among representative S2ST approaches across BLEU, COMET, ASRBLEU, and BLASER 2.0 under the adopted evaluation protocol. Ablation studies indicate that the proposed representation-level, acoustic-level, and decoding-level strategies provide complementary benefits. Moreover, controlled data-budget analyses and a Japanese-to-English evaluation using only ∼3 hours of supervised training data suggest that explicit typological priors provide useful inductive biases for data-efficient multilingual S2ST. Index Terms—Multilingual speech-to-speech translation, SpeechLLMs, linguistic typology, multilingual conditioning, data-efficient speech translation

I. I NTRODUCTION ULTILINGUAL speech-to-speech translation (S2ST) aims to directly translate spoken utterances from one language into speech in another, and is essential for crosslingual communication in scenarios such as healthcare, education, and international collaboration [1], [2]. Traditional S2ST systems typically rely on cascaded automatic speech recognition (ASR) [3], [4], machine translation

M

Yu Pan was with the School of Information Science and Electrical Engineering, Kyushu University, Fukuoka 819-0395, Japan, when this work was conducted, and is currently with Recho Inc., Tokyo, Japan (e-mail: [email protected]). Yang Hou is with the National Institute of Informatics, Tokyo, Japan (email: [email protected]). Xiongfei Wu and Yves LE TRAON are with the Interdisciplinary Research Centre on Security, Reliability and Trust (SnT), University of Luxembourg, Luxembourg (e-mail: [email protected]; [email protected]). Liang Zhang is with Donghua University, Shanghai 201620, China (e-mail: [email protected]). Lei Ma is with the Department of Computer Science, The University of Tokyo, Tokyo 113-8656, Japan, and the Department of Electrical and Computer Engineering, University of Alberta, Edmonton, Canada (e-mail: [email protected]). Jianjun Zhao is with the School of Information Science and Electrical Engineering, Kyushu University, Fukuoka 819-0395, Japan (e-mail: [email protected]).

(MT) [5], [6], and text-to-speech synthesis (TTS) [7], [8]. Although effective in practice, such pipelines are prone to error propagation and cannot be optimized globally. Recent studies have therefore explored end-to-end S2ST [9]–[11] and compositional S2ST [12]–[14]. Among them, compositional S2ST, which combines a speech-to-text translation (S2TT) frontend with a TTS backend, offers a practical balance between modularity, interpretability, and the ability to leverage speech and text resources independently. With the rapid progress of large language models (LLMs) [15]–[17], speech-aware LLMs (SpeechLLMs) [18], [19] have become a promising foundation for multilingual S2ST [14], [20], [21]. Along this line, S2ST-Omni [14] introduces language-label conditioning into a SpeechLLMbased compositional framework, enabling effective many-toone S2ST. However, its language-conditioning strategy follows a language-as-label paradigm, where each source language is represented as an isolated identifier. Such flat language representations overlook systematic linguistic regularities in morphology, reordering tendencies, and genealogical relatedness, which can affect speech alignment, semantic interpretation, and target-language generation [22]–[25]. From this perspective, multilingual S2ST should not only identify which language the input belongs to, but also capture what structural properties the language exhibits. Flat language embeddings may therefore be insufficient for exposing structural priors that support data-efficient multilingual adaptation [26], [27]. In this paper, we propose S2ST-Omni 2, a typology-aware compositional framework for many-to-one data-efficient S2ST. Built upon S2ST-Omni, S2ST-Omni 2 preserves the encoder– adapter–LLM–TTS skeleton while redesigning the languageconditioning pathway at three levels. First, typology-informed hierarchical language encoding (TI-HLE) decomposes source-language information into morphology-related, reordering, genealogical-family, and residual language-specific channels. Second, a dynamically-gated language-aware DualCTC mechanism performs content-adaptive frame-wise modulation for multilingual acoustic modeling. Third, typologyaware prompting injects translation-oriented linguistic priors into LLM decoding. Together, these components provide a structured, adaptive, and linguistically grounded formulation of multilingual language conditioning. We evaluate S2STOmni 2 on CVSS-C [28] against representative S2ST systems. Compared with the direct baseline S2ST-Omni, S2ST-Omni 2 achieves average relative gains of 5.8% in BLEU and 4.6% in ASR-BLEU, with consistent improvements in COMET and BLASER 2.0. Ablation studies show the complementary contributions of the proposed representation-level, acousticlevel, and decoding-level strategies. Furthermore, controlled

2

Fig. 1. Overall architecture and two-stage training pipeline of S2ST-Omni 2. LA denotes language-aware, CE is cross-entropy, and src/tgt denote source/target. TI-HLE and Dynamically-Gated LA-Dual-CTC are training-time auxiliary modules, whereas typology-aware prompting is retained during inference.

data-budget analyses and a limited-supervision Japanese-toEnglish evaluation suggest that explicit typological priors are particularly beneficial when supervised data are scarce. In summary, this work substantially extends S2STOmni [14] by reformulating flat language-label conditioning as structured typological conditioning and by providing a broader empirical evaluation. The main contributions are as follows: • We propose S2ST-Omni 2, a typology-aware compositional S2ST framework that reformulates multilingual language conditioning from flat language labels to structured typological priors. • We introduce TI-HLE, which decomposes sourcelanguage information into morphology-related, reordering, genealogical-family, and residual language-specific channels. • We propose a dynamically-gated language-aware DualCTC mechanism and typology-aware prompting to inject typological priors into acoustic feature modulation and LLM-based decoding, respectively. • We conduct extensive experiments on CVSS-C, including ablation studies, TTS-backend analysis, data-budget comparisons, and a Japanese-to-English evaluation using ∼3 hours of supervised data, providing empirical evidence for the effectiveness of structured typological conditioning in the evaluated multilingual S2ST setting. II. M ETHODOLOGY A. System Overview As shown in Fig. 1, S2ST-Omni 2 follows the compositional design of S2ST-Omni [14], consisting of a SpeechLLMbased S2TT frontend and a plug-and-play TTS backend. The frontend contains five main components: 1) a frozen Whisper encoder [29] for frame-level acoustic–semantic feature extraction; 2) a hybrid speech adapter inherited from S2ST-Omni to

map speech features into the LLM hidden space; 3) a TIHLE module that represents each source language through morphology, reordering, genealogical family, and residual language-specific factors; 4) a Dynamically-Gated LanguageAware Dual-CTC module that applies typology-conditioned modulation to intermediate adapter features with auxiliary source- and target-side CTC supervision; and 5) a Qwen34B decoder [30] guided by a Typology-Aware LLM Prompt for target-language translation. The TTS backend is decoupled from the S2TT frontend, allowing different synthesizers to be integrated without retraining. Following S2ST-Omni [14], the source-language identifier is obtained from ground-truth labels during training and predicted from Whisper encoder representations during inference. The key distinction from S2ST-Omni lies in the languageconditioning pathway. Rather than modifying the overall S2ST backbone, S2ST-Omni 2 replaces flat language-label conditioning with structured typological priors injected at the representation, acoustic-modulation, and LLM-decoding levels. Keeping the backbone unchanged reduces architectural confounds and enables a focused examination of linguistically grounded conditioning while preserving the modularity of the original framework. During inference, the TI-HLE and dynamically-gated LA-Dual-CTC modules are discarded together with the auxiliary CTC branches; therefore, they act only as training-time typological inductive biases and introduce no additional acoustic-side inference cost or change to the encoder–adapter–LLM forward path. The only inference-time difference is the typology-aware prompt selected according to the predicted source language. B. Hybrid Speech Adapter We adopt the hybrid adapter from S2ST-Omni [14] to bridge the frozen Whisper encoder and the Qwen3 LLM. This component is kept unchanged to minimize architectural confounds

3

TABLE I T YPOLOGICAL FEATURE ASSIGNMENT USED IN S2ST-O MNI 2.

Language

Morphology

French Fusional Spanish Fusional German Fusional+Compounding Japanese Agglutinative

Reordering profile

Family

SVO-oriented SVO-oriented Verb-/clause-final Verb-/clause-final

Romance Romance Germanic Japonic

and isolate the effect of the proposed typology-aware conditioning. Given Whisper encoder outputs X ∈ RB×T ×1280 , the adapter first projects them into a dh = 1024 hidden space, applies two local depthwise-separable convolution blocks with kernel size 7, downsamples the sequence with stride 2, and then uses two global self-attention blocks to model longrange dependencies. We denote the downsampled intermediate ′ adapter features as Hdown ∈ RB×T ×dh , where dh = 1024 and T ′ = ⌈T /2⌉. The final linear projection maps the adapter output to the LLM hidden dimension dllm = 3584, yielding ′ Z ∈ RB×T ×dllm for Qwen3 decoding. More details of this inherited adapter can be found in [14]. C. Typology-Informed Hierarchical Language Encoding Flat language conditioning treats each source language as an isolated symbol, without explicitly exposing linguistic properties that affect translation behavior. Motivated by linguistic typology and typology-based language representations in NLP [22]–[25], we construct a typology-informed language representation for speech-side conditioning by decomposing source-language information into four complementary feature groups: morphology-related profile, English-directed reordering profile, genealogical family, and a language-specific residual channel. The first three groups provide coarse but interpretable typological priors, whose assignments are summarized in Table I, while the residual channel preserves finegrained language-specific information not captured by these categories. These assignments are not intended as exhaustive linguistic classifications; rather, they are coarse, translationoriented profiles designed to encode recurrent structural tendencies relevant to English-directed S2ST. 1) Typological Feature Encoding: For each source language, we encode the morphology-related profile with a learnable embedding em ∈ Rd1 . Following standard linguistic typology [31], [32], French and Spanish are assigned to a fusional profile, German to a fusional+compounding profile, and Japanese to an agglutinative profile. This group provides priors for morphologically structured forms, productive compounding, and speech–text correspondence. Then, we encode the English-directed reordering profile with a learnable embedding ew ∈ Rd2 . Based on typological word-order classifications [33], [34] and the reordering demands of translation into English, French and Spanish are assigned to an SVO-oriented profile, whereas German and Japanese are assigned to a verb-/clause-final reordering profile. This grouping does not imply that German and Japanese share the same syntactic system; rather, it reflects that both often require stronger clause-final or verb-final reordering cues than French and Spanish in English-directed translation.

We further encode genealogical family with a learnable embedding ef ∈ Rd3 . Motivated by prior NLP work on cross-lingual structure and genealogical relatedness [23], [24], French and Spanish share a Romance-family embedding, while German and Japanese are assigned distinct Germanic and Japonic embeddings. This design enables explicit sharing for related languages while preserving separate family-level priors for unrelated languages in the present benchmark [35]. 2) Language-Specific Residual Channel: Because typological profiles are necessarily coarse-grained, we introduce a language-specific residual channel er ∈ Rd4 to preserve information not covered by morphology, reordering, or genealogical family. Its dimensionality is set to match the flat languageembedding dimension used in S2ST-Omni, so that the residual channel retains the original language-specific capacity while the additional channels explicitly encode typological structure. 3) Multi-Feature Fusion: The four feature groups are concatenated and projected into a unified language representation: rlang = GELU(LN(Wf [em ; ew ; ef ; er ] + bf )) ,

(1)

where Wf ∈ Rdc ×(d1 +d2 +d3 +d4 ) , dc = 256, and LN denotes layer normalization. The resulting representation rlang ∈ Rdc is used as the conditioning input for both the FiLM generator and the dynamic frame gate. D. Dynamically-Gated Language-Aware Dual-CTC To make language conditioning sensitive to both sourcelanguage structure and frame-level acoustic variation, we introduce a dynamically-gated Language-Aware Dual-CTC module over the downsampled intermediate adapter features ′ Hdown ∈ RB×T ×dh . This module consists of a typologyconditioned source CTC branch and a language-agnostic target CTC branch, which jointly provide source-side content preservation and target-side alignment guidance. 1) FiLM-Based Language Conditioning: Given the fused language representation rlang ∈ Rdc , a FiLM generator [36], [37] predicts feature-wise affine modulation parameters: [γ, β] = split (tanh (fFiLM (rlang ))) ,

(2)

where fFiLM : Rdc → R2dh is a two-layer MLP, and split(·) evenly divides the output into γ, β ∈ Rdh . The tanh(·) activation bounds the modulation parameters and helps stabilize training. For each frame t, the intermediate adapter feature hdown ∈ Rdh is modulated as: t e src = (1 + gt γ) ⊙ hdown + gt β, h t t

(3)

e src ∈ Rdh is the modulated source-side feature, 1 is an where h t all-ones vector, gt ∈ (0, 1) is a scalar gate broadcast along the feature dimension, and ⊙ denotes element-wise multiplication. 2) Dynamic Frame Gate: Instead of applying a globally shared static gate, we compute a per-frame gate conditioned on both acoustic content and language representation:   fgate ([hdown ; rlang ]) t , (4) gt = σ τ

4

where [·; ·] denotes vector concatenation, fgate : Rdh +dc → R is a two-layer MLP, and σ(·) is the sigmoid function. The temperature is parameterized as

In Stage I, the model is optimized mainly to establish reliable speech–text alignment using the LLM cross-entropy loss and dual CTC supervision:

τ = softplus(τlearn ) + ϵ,

tgt src L(1) = LCE + λ(1) src LCTC + λtgt LCTC .

(5)

where τlearn is a learnable scalar, softplus(·) ensures positivity, and ϵ = 0.1 prevents the temperature from becoming too small. The bias of fgate is initialized to −2.0, so that modulation is weak at the beginning of training and gradually increases as stable conditioning patterns emerge. This design allows typology-aware modulation to vary across both languages and time frames, rather than being uniformly applied to the entire utterance. 3) Source and Target CTC Branches: The source CTC branch applies the gated FiLM modulation in Eq. 3 and e src }T ′ to sourcee src = {h projects the resulting features H t t=1 language CTC logits. It is supervised by the standard CTC loss [38], encouraging the model to preserve source-language content under typology-aware conditioning. In contrast, the target CTC branch operates directly on the unconditioned intermediate adapter features Hdown and predicts English CTC logits, providing an additional target-side alignment signal without language-specific modulation. The source and target CTC losses are combined with the LLM cross-entropy loss during progressive fine-tuning. Following S2ST-Omni [14], we use stage-specific CTC weights rather than a single fixed weighting scheme throughout training. The detailed two-stage objectives are described in Section II-F.

(1)

(6)

In Stage II, the same modules remain trainable, and LoRA [39] adapters are inserted into the query and value projections in the self-attention modules of Qwen3 to enhance translation capability. The CTC losses are down-weighted to retain auxiliary alignment regularization: (2)

tgt src L(2) = LCE + λ(2) src LCTC + λtgt LCTC .

(7)

All stage-specific loss weights and optimization hyperparameters are kept consistent with S2ST-Omni, so that the comparison isolates the effect of the proposed typology-aware language conditioning. G. TTS Backend The TTS backend converts the target text generated by the S2TT frontend into target speech. Since the S2TT and TTS modules are connected through an explicit text interface, the proposed framework supports plug-and-play integration with different state-of-the-art TTS systems, without retraining or task-specific coupling of the S2TT frontend. In our experiments, we evaluate six publicly available recent TTS systems [30], [40]–[44] to verify the flexibility and backend interchangeability of this strategy in practical deployment. III. E XPERIMENTAL S ETUP

E. Typology-Aware LLM Prompting To complement acoustic-level conditioning, we introduce a typology-aware LLM prompting strategy. Each prompt consists of a shared system instruction specifying general translation principles and a language-specific instruction highlighting major translation challenges of the source language. The prompts are constructed from coarse typological and linguistic properties relevant to translation, without using sentence-level annotations, dataset-specific examples, or utterance-level information. Specifically, the German prompt emphasizes compound decomposition and clause-final-to-English reordering; the French and Spanish prompts focus on idiomatic expressions and lexical usage; and the Japanese prompt accounts for SOV-to-SVO reordering, omitted-subject inference, and honorific normalization. Since the same prompt is applied to all utterances from the same source language, this strategy provides only language-level prior knowledge and encourages more natural and faithful English translations.

A. Datasets We conduct experiments on the CVSS-C corpus [28], a publicly available multilingual S2ST corpus derived from CoVoST 2 [48]. CVSS-C provides parallel speech in multiple source languages paired with synthesized English target speech, enabling standardized evaluation of S2ST systems. Following prior work [12], [14], we mainly evaluate French→English, German→English, and Spanish→English. For the main multilingual setting, a single model is jointly trained on the supervised training sets of these three directions, containing approximately 264 hours for French, 184 hours for German, and 113 hours for Spanish, for a total of 561 hours. To further examine robustness under limited supervised data for a typologically distant source language, we additionally evaluate Japanese→English using approximately three hours of supervised CVSS-C training data. Both S2ST-Omni and S2ST-Omni 2 are trained under the same Japanese setting and evaluated with the same TTS backend and metric pipeline.

F. Progressive Fine-Tuning

B. Implementation Details

Following S2ST-Omni [14], we adopt the same two-stage progressive fine-tuning strategy to stabilize speech–text alignment before LLM adaptation. In both stages, the Whisper encoder and the base Qwen3 parameters are frozen, while the hybrid speech adapter, TI-HLE module, and dynamicallygated LA-Dual-CTC module are updated.

We use Whisper-Large-V3 [29] as the frozen speech encoder and Qwen3-4B [30] as the LLM decoder. The hybrid speech adapter follows S2ST-Omni [14] and the architecture described in Section II-B. For LLM adaptation, LoRA is applied to the query and value projection layers with rank r=8, scaling factor α=32, and dropout 0.1.

5

TABLE II OVERALL PERFORMANCE COMPARISON ON CVSS-C. B EST RESULTS ARE SHOWN IN BOLD . “-” INDICATES RESULTS NOT REPORTED OR NOT APPLICABLE . G ROUND - TRUTH ASR-BLEU AND W HISPER –Q WEN S2TT REFERENCE RESULTS ARE INCLUDED ONLY AS REFERENCES AND ARE NOT CONSIDERED FOR BEST HIGHLIGHTING . † DENOTES A SINGLE MANY- TO - ONE MODEL EVALUATED ACROSS THE THREE SOURCE - TO -E NGLISH DIRECTIONS . U NMARKED SYSTEMS FOLLOW THEIR ORIGINALLY REPORTED EVALUATION SETTINGS .

Ground Truth Whisper–Qwen S2TT Translatotron 2 [10] DASpeech [13] UnitY [45] ComSpeech [12] StreamSpeech [46] SimulS2S-LLM [20] Hibiki [47] RosettaSpeech† [21] S2ST-Omni† [14]

BLEU 35.15 28.82 30.72 32.60 33.11 35.83

Fr→En ASR-BLEU 84.52 26.07 25.03 27.77 28.15 28.45 26.93 30.50 32.16 33.20

BLEU 36.07 18.66 19.41 23.36 23.22 33.34

De→En ASR-BLEU 75.53 16.91 16.14 18.74 18.16 20.93 21.50 21.54 31.25

BLEU 38.39 25.82 26.51 30.35 30.92 37.85

Es→En ASR-BLEU 88.54 22.93 21.37 24.95 24.80 27.25 26.33 29.35 35.90

BLEU 36.54 24.43 25.55 28.77 29.08 35.67

Average ASR-BLEU 82.86 21.97 20.85 23.82 23.70 25.54 24.92 27.68 33.45

S2ST-Omni 2†

37.83

34.72

35.70

33.16

39.62

37.13

37.73

35.00

Model

For TI-HLE, the morphology-related, reordering, genealogical-family, and residual channels have dimensions of 64, 64, 64, and 128, respectively, and the concatenated 320-dimensional vector is fused into a 256-dimensional language representation. DG-LA-Dual-CTC operates on the intermediate adapter features Hdown with dh =1024; the FiLM generator therefore predicts 2dh =2048 affine parameters, and the dynamic frame gate uses a hidden dimension of 256. The source- and target-side CTC branches use SentencePiece tokenizers [49] with vocabularies of 8k and 4k subword units, respectively. The CTC (1) (1) weights are set to (λsrc , λtgt ) = (0.1, 0.2) in Stage I and (2) (2) (λsrc , λtgt ) = (0.01, 0.05) in Stage II. Training follows the progressive fine-tuning strategy described in Section II-F. All experiments use an effective batch size of 24, implemented with a per-device batch size of 3 and gradient accumulation of 8. We use bf16 mixed precision and conduct all experiments on two NVIDIA A6000 GPUs.

frontend against the ground-truth English references. BLEU measures surface-level lexical overlap, while COMET provides a complementary semantic-level evaluation of translation quality. The Whisper–Qwen S2TT reference is included only as an auxiliary text-level reference in Table II; therefore, only BLEU is reported for this reference. Regarding E2E speech translation quality, we report ASR-BLEU and BLASER 2.0. For ASR-BLEU, the generated speech is first transcribed by a pretrained wav2vec 2.0 ASR model1 , and BLEU is then computed using SacreBLEU2 with a fixed configuration3 . For BLASER 2.0, we use the reference-based configuration [46], i.e., BLASER 2.0-Ref. For systems whose outputs are reproduced or publicly available, all scores are computed using the same evaluation pipeline. For prior systems without publicly available generated outputs, we report the values from the corresponding papers when available.

C. Baselines We compare S2ST-Omni 2 with representative S2ST approaches covering E2E, compositional, simultaneous, zeroshot, and SpeechLLM-based paradigms. These include Translatotron 2 [10], UnitY [45], DASpeech [13], ComSpeech [12], StreamSpeech [46], SimulS2S-LLM [20], Hibiki [47], RosettaSpeech [21], and the direct baseline S2ST-Omni [14]. In addition, we report a Whisper–Qwen S2TT as text-level reference, which uses Whisper-Large-V3 to transcribe the source speech and Qwen3-4B to translate the resulting sourcelanguage transcript into English.

A. Overall Performance

D. Evaluation Metrics Following prior S2ST works [12], [14], [28], we evaluate our model from two perspectives: text-level translation quality and E2E speech translation quality. For text-level translation quality, we report BLEU [50] and COMET [51], both computed on the text output of the S2TT

IV. R ESULTS AND D ISCUSSION

Tables II, IV, and V report the overall results on CVSS-C in terms of BLEU, ASR-BLEU, COMET, and BLASER 2.0. It is worth noting that S2ST-Omni 2 is evaluated as a unified manyto-one multilingual S2TT frontend shared across the three source-to-English directions, rather than relying on separate pair-specific models for each direction. This setting is more challenging because a single shared frontend must handle source-language differences within a unified parameter space. Among all evaluated S2ST systems, S2ST-Omni 2 achieves the best average performance across the reported metrics under the adopted evaluation protocol, suggesting that the proposed typology-aware conditioning improves both textlevel translation quality and E2E S2ST evaluation quality. 1 https://dl.fbaipublicfiles.com/fairseq/wav2vec/wav2vec vox 960h pl.pt 2 https://github.com/mjpost/sacrebleu 3 SacreBLEU signature: nrefs:1|case:mixed|eff:no|tok:13a|smooth:exp| version:2.6.1.dev1+gf615c7286

6

TABLE III A BLATION STUDY ON CVSS-C. E ACH ROW REMOVES OR REPLACES ONE COMPONENT OR ONE FEATURE GROUP FROM THE FULL S2ST-O MNI 2. S UBSCRIPTS IN THE AVERAGE COLUMNS DENOTE RELATIVE DEGRADATION WITH RESPECT TO THE FULL MODEL , WHILE THE MAIN VALUES REPORT ABSOLUTE BLEU AND ASR-BLEU SCORES .

BLEU

Fr→En ASR-BLEU

BLEU

De→En ASR-BLEU

BLEU

Es→En ASR-BLEU

BLEU

S2ST-Omni 2

37.83

34.72

35.70

33.16

39.62

37.13

37.73

35.00

w/o DG w/o TA-Prompt w/o TI-HLE

37.02 36.93 35.77

33.31 33.21 32.56

34.85 34.69 34.24

32.17 32.05 31.93

39.01 38.78 38.26

36.74 36.63 36.54

36.96−2.04% 36.80−2.46% 36.09−4.35%

34.07−2.66% 33.96−2.97% 33.68−3.77%

w/o Morph w/o Reorder w/o Family w/o Residual

35.93 36.45 36.12 35.91

32.84 33.34 32.93 32.87

34.36 34.68 34.65 34.38

32.09 32.17 32.25 32.04

38.39 38.68 38.55 38.33

36.32 36.32 36.43 36.30

36.23−3.98% 36.60−2.99% 36.44−3.42% 36.21−4.03%

33.75−3.57% 33.94−3.03% 33.87−3.23% 33.74−3.60%

Model

TABLE IV COMET RESULTS ON CVSS-C. S CORES ARE REPORTED ON A 0–100 SCALE , WITH THE BEST RESULTS HIGHLIGHTED IN BOLD .

Models Fr→En De→En Es→En Avg. ComSpeech [12] 70.13 57.96 66.96 65.02 StreamSpeech [46] 76.66 65.51 74.80 72.32 RosettaSpeech [21] 78.97 79.65 82.05 80.22 S2ST-Omni [14] 81.94 80.73 83.39 82.02 S2ST-Omni 2

82.74

82.16

85.03

83.31

Compared with the direct baseline S2ST-Omni [14], S2STOmni 2 improves average BLEU from 35.67 to 37.73 and average ASR-BLEU from 33.45 to 35.00, corresponding to relative gains of 5.8% and 4.6%, respectively. It also improves average COMET and BLASER 2.0 by +1.29 and +0.10. Moreover, the largest BLEU and ASR-BLEU improvements are observed on De→En, which is consistent with the motivation that German involves stronger compound morphology and clauselevel reordering mismatch with English. This observation is consistent with the hypothesis that structured typological conditioning is beneficial when translation requires stronger structural mediation. TABLE V BLASER 2.0 RESULTS ON CVSS-C. B EST RESULTS ARE SHOWN IN BOLD .

Models

Fr→En De→En Es→En Avg.

UnitY [45] Translatotron 2 [10] ComSpeech w/pretrain [12] StreamSpeech [46] RosettaSpeech [21] S2ST-Omni [14]

3.17 3.18 3.19 3.20 4.04 4.12

2.83 2.90 2.93 3.00 4.08 4.10

3.23 3.26 3.28 3.31 4.17 4.21

3.08 3.11 3.13 3.17 4.10 4.14

S2ST-Omni 2

4.21

4.18

4.33

4.24

To further contextualize frontend translation quality, we compare S2ST-Omni 2 with the Whisper–Qwen S2TT reference, which follows a cascaded ASR–MT pipeline. As shown in Table II, S2ST-Omni 2 improves the average BLEU score from 36.54 to 37.73, with gains of +2.68 on Fr→En and +1.23

Average ASR-BLEU

on Es→En, while remaining slightly lower on De→En by 0.37 BLEU. This comparison is informative because S2ST-Omni 2 performs many-to-one S2TT through a unified multilingual SpeechLLM frontend, instead of decomposing the process into ASR and MT. The higher average BLEU score suggests that typology-aware speech representations can provide effective guidance for multilingual S2TT, making the unified manyto-one frontend competitive with a strong cascaded textlevel reference under the adopted BLEU evaluation protocol. In addition, S2ST-Omni 2 also shows clear advantages compared with previous SOTA S2ST methods. Relative to RosettaSpeech [21], a recent strong baseline, S2ST-Omni 2 improves average BLEU and ASR-BLEU by +8.65 and +7.32, while also improving average COMET and BLASER 2.0 by +3.09 and +0.14. Additionally, S2ST-Omni 2 further outperforms other representative systems in all evaluated metrics, showcasing the effectiveness of the proposed approach. Overall, these results indicate that the advantage of S2STOmni 2 comes not merely from the SpeechLLM backbone, but from the structured redesign of source-language conditioning. They highlight the importance of how language information is represented and injected in many-to-one multilingual S2ST. B. Ablation Study To assess the contribution of each component, we conduct systematic ablation experiments, as summarized in Table III. “w/o TI-HLE” replaces the proposed typology-informed hierarchical language encoding with a 320-dimensional flat per-language embedding, while keeping the remaining S2STOmni 2 components unchanged. This setting evaluates whether structured typological decomposition provides benefits beyond flat language-label conditioning with matched input dimensionality. “w/o DG” replaces the dynamic frame gate with a static scalar gate, and “w/o TA-Prompt” replaces the proposed typology-aware prompting with the language-aware prompting used in S2ST-Omni, which specifies the source language but does not include explicit typological guidance. In addition, “w/o Morph,” “w/o Reorder,” “w/o Family,” and “w/o Residual” remove the morphology-related profile, wordorder/reordering profile, genealogical-family embedding, and language-specific residual channel, respectively.

7

TABLE VI R EPRESENTATIVE EXAMPLES COMPARING S2ST-O MNI 2 WITH S2ST-O MNI . B OLDFACE HIGHLIGHTS THE CRITICAL ERROR SPAN IN THE TRANSLATION .

System/Type

Example 1

Example 2 German→English

Source Text

Bei viel Regen dehnt sich das Rückhaltebecken enorm aus.

Als Gegenleistung für diese militärischen Dienste erhielt er die Stadt Madaba.

Reference

When there is a lot of rain, the retention basin expands enormously.

As a reward for these military services, he received the city Madaba.

S2ST-Omni

The reservoir is stretched a lot when there are many rains.

The city of Madaba received the town as compensation for this military service.

S2ST-Omni 2

When it rains a lot, the retention basin expands a lot.

As a reward for this military service, he received the city of Madaba.

Spanish→English Source Text

Ası́, al partido se le asignaron cinco escaños en el parlamento.

A quien mucho miente, le huye la gente.

Reference

Thus the party was assigned five seats in the parliament.

From whom much lies people flee.

S2ST-Omni

Thus five scottish members were assigned to the party in parliament.

The more you lie, the people run away from your.

S2ST-Omni 2

Thus the party was assigned five seats in the parliament.

Those who lie too much will lose their friends.

French→English Source Text

Pouvez-vous me rendre un petit service ?

Elle aurait des vertus médicinales.

Reference

Can you do me a small favor?

It has medicinal properties.

S2ST-Omni

Can you give me a little service.

She would have medical aspects.

S2ST-Omni 2

Can you do me a small favor.

It would have medicinal properties.

1) Effect of Dynamically-Gated Language-Aware DualCTC: Replacing the dynamic frame gate with a static scalar gate (w/o DG) causes consistent degradation in all evaluated metrics. To elaborate, it reduces average BLEU by 0.77 and average ASR-BLEU by 0.93, showing that adaptive modulation remains beneficial once richer language representations are available. This suggests that typological priors should not be applied uniformly to all frames; instead, their modulation strength should vary according to both acoustic content and source-language characteristics. 2) Effect of Typology-Aware LLM Prompting: Replacing typology-aware prompting with the original language-aware prompting used in S2ST-Omni (w/o TA-Prompt) decreases average BLEU by 0.93 and average ASR-BLEU by 1.04. This result suggests that the improvement is not merely due to indicating the source language to the LLM, but to providing explicit typology-aware translation guidance beyond conventional language-aware prompting. 3) Effect of TI-HLE: Removing TI-HLE (w/o TI-HLE) yields the largest degradation, reducing average BLEU and ASR-BLEU by 1.64 and 1.32, respectively. Notably, this variant uses a 320-dimensional flat per-language embedding, matching the input dimensionality of the proposed hierarchical representation. The drop therefore cannot be simply attributed to language-embedding capacity; rather, it indicates that decomposing language information into typological and residual channels provides a more effective conditioning signal than an

unstructured flat embedding. The fine-grained ablations further show that each feature group contributes to performance. Removing the residual and morphology-related channels causes the largest BLEU drops, by 1.52 and 1.50 points, followed by genealogical family (-1.29) and reordering (-1.13). Similar trends are observed for ASR-BLEU, with drops of 1.26, 1.25, 1.13, and 1.06 points, respectively. These results suggest that the residual channel preserves language-specific capacity, while the explicit typological channels provide complementary structural priors for multilingual adaptation and translationoriented feature modulation. 4) Summary of Ablation Study: Taken together, these ablation results show that all variants underperform the full S2STOmni 2 model, indicating that the observed gains do not stem from a single isolated component. Rather, S2ST-Omni 2 benefits from the complementary effects of representationlevel typological decomposition, acoustic-level adaptive modulation, and decoder-side typology-aware prompting. C. Qualitative Analysis To complement the quantitative results, Table VI presents representative examples comparing S2ST-Omni 2 with the direct baseline S2ST-Omni. This analysis aims to examine how the proposed typology-informed structured language conditioning affects translation behavior in linguistically challenging cases, including compound morphology, argumentstructure preservation, non-literal expressions, and context-

8

40

BLEU

38

40

Rel. gain over S2ST-Omni 30h: +15.1% 50h: +12.7% 100h: +10.4% 200h: +7.6% 400h: +6.4% 561h: +5.8%

38

34

34

30

30

26

22

26

S2ST-Omni S2ST-Omni 2 0 3050

100

200

300

400

500561

600

22

S2ST-Omni S2ST-Omni 2 0 3050

100

200

BLEU

(a) Average 40

38

38

34

34

30

30

26

26

S2ST-Omni S2ST-Omni 2 0 3050

100

200

300

400

500561

600

(b) Fr→En

40

22

300

400

500561

600

22

S2ST-Omni S2ST-Omni 2 0 3050

100

Training Data Duration (hours)

200

300

400

500561

600

Training Data Duration (hours)

(c) De→En

(d) Es→En

Fig. 2. BLEU under varying training data budgets for S2ST-Omni and S2ST-Omni 2. (a) Average BLEU, with relative gains computed over S2ST-Omni. (b) Fr→En. (c) De→En. (d) Es→En. Across all settings, S2ST-Omni 2 consistently outperforms S2ST-Omni, and the relative advantage becomes more pronounced as the training data budget decreases.

dependent lexical choices. To be specific, S2ST-Omni 2 better preserves compound morphology and clause-level semantic relations for German. It renders Rückhaltebecken as “retention basin,” whereas S2ST-Omni produces the less specific “reservoir” and an unnatural rendering of the rain condition. It also preserves the intended argument structure in the Madaba example, while S2ST-Omni incorrectly suggests that the city received the town. Regarding Spanish, S2ST-Omni 2 better handles both structural and non-literal expressions. In the passive construction with se le asignaron, it correctly preserves the meaning of “five seats,” whereas S2ST-Omni generates the semantically implausible phrase “five scottish members.” For the proverb-like expression, S2ST-Omni 2 produces a more complete and natural paraphrase, avoiding the incomplete literal rendering generated by S2ST-Omni. For French, S2STOmni 2 improves context-dependent lexical choice and natural English phrasing. It maps rendre un petit service to the natural English collocation “do me a small favor,” while S2ST-Omni follows a less appropriate word-by-word translation. It also better handles the mismatch between French grammatical gender and English reference by translating Elle as “it” rather than “she” in the medicinal-properties example. Overall, these examples illustrate that the proposed structured language conditioning helps S2ST-Omni 2 produce more faithful and natural translations across typologically differ-

ent source languages, further supporting the effectiveness of typology-informed conditioning strategy. D. Further Analysis 1) Effect of TTS Backend: Table VII reports ASR-BLEU results with six publicly available TTS backends while keeping the S2TT frontend fixed. The average scores range from 33.87 to 35.00, with a 1.13-point gap between the weakest and strongest backends. Although backend-specific factors such as pronunciation fidelity and prosodic realization still affect ASRBLEU, the relative stability across backends suggests that, in terms of ASR-BLEU, the improvements of S2ST-Omni 2 are not tied to a specific synthesizer under our evaluation protocol. TABLE VII E FFECT OF VARIOUS TTS BACKENDS ON ASR-BLEU. A LL CONFIGURATIONS USE THE SAME S2ST-O MNI 2 S2TT FRONTEND .

TTS Backend

Fr→En

De→En

Es→En

Avg

IndexTTS2 [40] CosyVoice3 [41] Qwen3-TTS [30] FireredTTS2 [42] ZipVoice [43] VoxCPM1.5 [44]

34.72 34.73 33.62 33.27 33.29 33.04

33.16 32.95 32.67 32.47 32.51 32.28

37.13 36.95 36.96 36.81 36.73 36.30

35.00 34.88 34.42 34.18 34.18 33.87

9

2) Effect of Training Data Budget: To examine data efficiency, we train both S2ST-Omni and S2ST-Omni 2 under data budgets ranging from 30 to 561 hours and compare their average BLEU scores. As shown in Fig. 2, the advantage of S2ST-Omni 2 increases monotonically as the amount of training data decreases: the absolute gain grows from +2.06 BLEU at 561 hours to +3.93 BLEU at 30 hours, while the relative improvement correspondingly increases from 5.8% to approximately 15.1%. This trend suggests that typologyaware conditioning is particularly beneficial under limited supervision. By explicitly encoding morphology, reordering, and genealogical relatedness, S2ST-Omni 2 provides structured language priors that can support parameter sharing and dataefficient multilingual adaptation when training data are scarce. In higher-resource settings, part of these regularities may be learned directly from data, leading to a more moderate gain. In lower-resource settings, however, the explicit typological priors provide additional guidance for acoustic conditioning and downstream decoding, resulting in a larger advantage over the flat language-label baseline.

TABLE VIII JAPANESE→E NGLISH TRANSLATION RESULTS WITH ∼3 HOURS OF CVSS-C SUPERVISED TRAINING DATA .

V. C ONCLUSION In this paper, we presented S2ST-Omni 2, a typologyaware compositional many-to-one S2ST framework that reformulates multilingual language conditioning for SpeechLLMbased S2ST. Instead of treating source languages as isolated flat labels, S2ST-Omni 2 introduces structured typological priors through typology-informed hierarchical language encoding, dynamically-gated language-aware Dual-CTC, and typology-aware LLM prompting, thereby enhancing language conditioning at the representation, acoustic-modulation, and decoding levels. Extensive experiments on CVSS-C demonstrate that S2ST-Omni 2 achieves strong average performance across BLEU, ASR-BLEU, COMET, and BLASER 2.0 under the adopted evaluation protocol. Ablation studies and qualitative analyses further indicate that the proposed components provide complementary benefits. In addition, the controlled data-budget analysis and the few-hour Japanese-to-English experiment provide complementary evidence that typologyinformed conditioning is particularly useful under reducedsupervision conditions, covering both lower data budgets within the main multilingual setting and a typologically distant low-resource extension. Overall, these findings highlight structured typological conditioning as a practical and effective inductive bias for data-efficient multilingual S2ST.

Model

BLEU

ASR-BLEU

COMET

BLASER 2.0

R EFERENCES

S2ST-Omni S2ST-Omni 2

19.61 22.00

18.59 20.93

78.29 80.31

3.692 3.779

[1] G.-i. Kikui, E. Sumita, T. Takezawa, and S. Yamamoto, “Creating corpora for speech-to-speech translation.” in INTERSPEECH, 2003, pp. 381–384. [2] A. Lee, H. Gong, P.-A. Duquenne, H. Schwenk, P.-J. Chen, C. Wang, S. Popuri, Y. Adi, J. Pino, J. Gu et al., “Textless speech-to-speech translation on real data,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 860–872. [3] A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolutionaugmented Transformer for Speech Recognition,” in Interspeech 2020, 2020, pp. 5036–5040. [4] Y. Yang, Y. Pan, J. Yin, J. Han, L. Ma, and H. Lu, “Hybridformer: Improving squeezeformer with hybrid attention and nsr mechanism,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5. [5] Y. Moslem, R. Haque, J. Kelleher, and A. Way, “Adaptive machine translation with large language models,” in Proceedings of the 24th Annual Conference of the European Association for Machine Translation, 2023, pp. 227–237. [6] K. Peng, L. Ding, Q. Zhong, L. Shen, X. Liu, M. Zhang, Y. Ouyang, and D. Tao, “Towards making the most of chatgpt for machine translation,” in Findings of the Association for Computational Linguistics: EMNLP 2023, 2023, pp. 5622–5633. [7] Z. Du, Q. Chen, S. Zhang, K. Hu, H. Lu, Y. Yang, H. Hu, S. Zheng, Y. Gu, Z. Ma et al., “Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens,” arXiv preprint arXiv:2407.05407, 2024. [8] S. Chen, Y. Feng, L. He, T. He, W. He, Y. Hu, B. Lin, Y. Lin, Y. Pan, P. Tan et al., “Takin: A cohort of superior quality zero-shot speech generation models,” arXiv preprint arXiv:2409.12139, 2024. [9] Y. Jia, R. J. Weiss, F. Biadsy, W. Macherey, M. Johnson, Z. Chen, and Y. Wu, “Direct Speech-to-Speech Translation with a Sequence-toSequence Model,” in Interspeech 2019, 2019, pp. 1123–1127. [10] Y. Jia, M. T. Ramanovich, T. Remez, and R. Pomerantz, “Translatotron 2: High-quality direct speech-to-speech translation with voice preservation,” in International Conference on Machine Learning. PMLR, 2022, pp. 10 120–10 134. [11] L. Barrault, Y.-A. Chung, M. C. Meglioli, D. Dale, N. Dong, P.-A. Duquenne, H. Elsahar, H. Gong, K. Heffernan, J. Hoffman et al., “Seamlessm4t: massively multilingual & multimodal machine translation,” arXiv preprint arXiv:2308.11596, 2023.

3) Japanese Extension under Limited Supervision: We further evaluate S2ST-Omni 2 on Japanese→English translation using only approximately three hours of supervised training data. As shown in Table VIII, S2ST-Omni 2 consistently outperforms S2ST-Omni across all four metrics, improving BLEU by +2.39, ASR-BLEU by +2.34, COMET by +2.02, and BLASER 2.0 by +0.087. This result suggests that the proposed typology-aware conditioning remains beneficial beyond the three European source languages considered in the main CVSS-C evaluation. This setting is particularly informative because Japanese is typologically distant from French, German, and Spanish in morphology, reordering profile, and genealogy, while also being evaluated under a highly limited data budget. The observed gains provide additional evidence that structured typological priors can offer useful guidance under low-resource and typologically divergent conditions. 4) Summary of Further Analysis: Overall, these analyses provide complementary evidence for the practical behavior of S2ST-Omni 2. The TTS-backend comparison suggests that ASR-BLEU gains are not tied to a specific synthesizer, the data-budget analysis shows that the relative advantage of typology-aware conditioning increases as supervision decreases, and the Japanese extension further indicates that this advantage can extend to a typologically distant low-resource source language. These findings suggest that structured typological priors can serve as useful inductive biases when multilingual S2ST models cannot fully infer language-specific regularities from abundant supervised data.

10

[12] Q. Fang, S. Zhang, Z. Ma, M. Zhang, and Y. Feng, “Can we achieve high-quality direct speech-to-speech translation without parallel speech data?” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 7264– 7277. [13] Q. Fang, Y. Zhou, and Y. Feng, “Daspeech: Directed acyclic transformer for fast and high-quality speech-to-speech translation,” Advances in Neural Information Processing Systems, vol. 36, pp. 72 604–72 623, 2023. [14] Y. Pan, X. Wu, Y. Yang, J. Yao, C. Maxime, L. Ma, and J. Zhao, “S2stomni: Hierarchical language-aware speechllm adaptation for multilingual speech-to-speech translation,” arXiv preprint arXiv:2506.11160, 2025. [15] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al., “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774, 2023. [16] J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang et al., “Qwen technical report,” arXiv preprint arXiv:2309.16609, 2023. [17] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar et al., “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023. [18] D. Zhang et al., “Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,” arXiv preprint arXiv:2305.11000, 2023. [19] R. Huang, M. Li, D. Yang, J. Shi, X. Chang, Z. Ye, Y. Wu, Z. Hong, J. Huang, J. Liu et al., “Audiogpt: Understanding and generating speech, music, sound, and talking head,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 21, 2024, pp. 23 802–23 804. [20] K. Deng, W. Chen, X. Chen, and P. Woodland, “Simuls2s-llm: Unlocking simultaneous inference of speech llms for speech-to-speech translation,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 16 718– 16 734. [21] Z. Zheng, X. Sun, T. Dinh, A. Yanamandra, A. Jain, Z. Liu, S. Hadap, V. Bhat, M. Aggarwal, G. Medioni et al., “Rosettaspeech: Zero-shot speech-to-speech translation from monolingual data,” arXiv preprint arXiv:2511.20974, 2025. [22] B. Comrie, Language universals and linguistic typology: Syntax and morphology. University of Chicago press, 1989. [23] P. Littell, D. R. Mortensen, K. Lin, K. Kairis, C. Turner, and L. Levin, “Uriel and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors,” in Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, 2017, pp. 8–14. [24] E. M. Ponti, H. O’horan, Y. Berzak, I. Vulić, R. Reichart, T. Poibeau, E. Shutova, and A. Korhonen, “Modeling language variation and universals: A survey on typological linguistics for natural language processing,” Computational Linguistics, vol. 45, no. 3, pp. 559–601, 2019. [25] A. Oncevay, B. Haddow, and A. Birch, “Bridging linguistic typology and multilingual machine translation with multi-view language representations,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 2391–2406. [26] A. Ansell, M. Parović, I. Vulić, A. Korhonen, and E. M. Ponti, “Unifying cross-lingual transfer across scenarios of resource scarcity,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 3980–3995. [27] S. Rajaee and C. Monz, “Analyzing the evaluation of cross-lingual knowledge transfer in multilingual language models,” in Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 2895– 2914. [28] Y. Jia, M. T. Ramanovich, Q. Wang, and H. Zen, “Cvss corpus and massively multilingual speech-to-speech translation,” in Proceedings of the thirteenth language resources and evaluation conference, 2022, pp. 6691–6703. [29] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning. PMLR, 2023, pp. 28 492–28 518. [30] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv et al., “Qwen3 technical report,” arXiv preprint arXiv:2505.09388, 2025. [31] B. Comrie, “Linguistic typology,” Annual Review of Anthropology, vol. 17, pp. 145–159, 1988.

[32] M. Haspelmath and A. Sims, Understanding morphology. Routledge, 2013. [33] M. S. Dryer, “Order of subject, object and verb,” in The World Atlas of Language Structures Online, M. S. Dryer and M. Haspelmath, Eds. Leipzig: Max Planck Institute for Evolutionary Anthropology, 2013. [Online]. Available: https://wals.info/chapter/81 [34] J. A. Hawkins, Word order universals. Elsevier, 2014, vol. 3. [35] S. Lim, T. Yun, J. Kim, J. Choi, and T. Kim, “Analysis of multisource language training in cross-lingual transfer,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 2024, pp. 712–725. [36] E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in AAAI, 2018. [37] J. Yao, Y. Yuguang, Y. Pan, Z. Ning, J. Ye, H. Zhou, and L. Xie, “Stablevc: Style controllable zero-shot voice conversion with conditional flow matching,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 24, 2025, pp. 25 669–25 677. [38] A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in ICML, 2006. [39] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al., “Lora: Low-rank adaptation of large language models.” ICLR, vol. 1, no. 2, p. 3, 2022. [40] S. Zhou, Y. Zhou, Y. He, X. Zhou, J. Wang, W. Deng, and J. Shu, “Indextts2: A breakthrough in emotionally expressive and durationcontrolled auto-regressive zero-shot text-to-speech,” arXiv preprint arXiv:2506.21619, 2025. [41] Z. Du, C. Gao, Y. Wang, F. Yu, T. Zhao, H. Wang, X. Lv, H. Wang, C. Ni, X. Shi et al., “Cosyvoice 3: Towards in-the-wild speech generation via scaling-up and post-training,” arXiv preprint arXiv:2505.17589, 2025. [42] K. Xie, F. Shen, J. Li, F. Xie, X. Tang, and Y. Hu, “Fireredtts-2: Towards long conversational speech generation for podcast and chatbot,” arXiv preprint arXiv:2509.02020, 2025. [43] H. Zhu, W. Kang, Z. Yao, L. Guo, F. Kuang, Z. Li, W. Zhuang, L. Lin, and D. Povey, “Zipvoice: Fast and high-quality zero-shot text-to-speech with flow matching,” arXiv preprint arXiv:2506.13053, 2025. [44] Y. Zhou, G. Zeng, X. Liu, X. Li, R. Yu, Z. Wang, R. Ye, W. Sun, J. Gui, K. Li et al., “Voxcpm: Tokenizer-free tts for context-aware speech generation and true-to-life voice cloning,” arXiv preprint arXiv:2509.24650, 2025. [45] H. Inaguma, S. Popuri, I. Kulikov, P.-J. Chen, C. Wang, Y.-A. Chung, Y. Tang, A. Lee, S. Watanabe, and J. Pino, “Unity: Two-pass direct speech-to-speech translation with discrete units,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 15 655–15 680. [46] S. Zhang, Q. Fang, S. Guo, Z. Ma, M. Zhang, and Y. Feng, “Streamspeech: Simultaneous speech-to-speech translation with multi-task learning,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 8964– 8986. [47] T. Labiausse, L. Mazaré, E. Grave, A. Défossez, and N. Zeghidour, “High-fidelity simultaneous speech-to-speech translation,” in Proceedings of the 42nd International Conference on Machine Learning. PMLR, 2025, pp. 32 116–32 129. [48] C. Wang, A. Wu, J. Gu, and J. Pino, “CoVoST 2 and Massively Multilingual Speech Translation,” in Interspeech 2021, 2021, pp. 2247– 2251. [49] T. Kudo and J. Richardson, “Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in Proceedings of the 2018 conference on empirical methods in natural language processing: System demonstrations, 2018, pp. 66–71. [50] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 2002, pp. 311–318. [51] R. Rei, J. G. De Souza, D. Alves, C. Zerva, A. C. Farinha, T. Glushkova, A. Lavie, L. Coheur, and A. F. Martins, “Comet-22: Unbabel-ist 2022 submission for the metrics shared task,” in Proceedings of the Seventh Conference on Machine Translation (WMT), 2022, pp. 578–585.

Record · ID 192418 · SHA-256 700766b9d1489f9e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.