Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation Yongshi Ye1,3 , Biao Fu2,3, * , Chongxuan Huang2,3 , Yidong Chen2,3 , Xiaodong Shi1,2,3, * 1
Institute of Artificial Intelligence, Xiamen University 2 School of Informatics, Xiamen University 3 Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism {yeyongshi,biaofu}@stu.xmu.edu.cn,[email protected]
arXiv:2607.29287v1 [cs.CL] 31 Jul 2026
Abstract Multi-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by human translators’ ability to adapt reasoning effort based on difficulty, we propose TwT (Translation with Thought), a resourcerational framework that learns to modulate inference between intuitive and deliberate reasoning. TwT is trained in two stages: (1) supervised fine-tuning on difficulty-aware long chain-ofthought traces distilled from DeepSeek-R1 and rewritten by GPT-4o to reflect human-like reasoning economy, and (2) reinforcement learning with a hybrid reward to optimize translation quality and reasoning efficiency. Evaluated on 15 benchmarks spanning in-domain and outof-domain settings, as well as 3 seen and 59 unseen languages, with ablations across three backbone models, TwT-7B and TwT-14B outperform much larger SOTA reasoning models in translation quality, while reducing token usage by 32–60%. These results confirm that aligning translation behavior with cognitive principles enables robust generalization, high translation quality, and efficient reasoning in MDMT.
1
Introduction
Multi-domain machine translation (MDMT) remains a core challenge for language models due to significant variation in terminology, syntax, and style across domains. A key difficulty lies in the uneven distribution of complexity: some inputs are routine, while others require deeper reasoning to resolve ambiguity or domain-specific constructs. However, most MT systems translate all inputs uniformly, lacking mechanisms to adjust inference effort based on domain-specific complexity (Li et al., 2025; Liu et al., 2025). This contrasts with human translators, who adapt reasoning effort to input difficulty (Hvelplund, 2011; Gile and Lei, * Corresponding authors.
2020). They typically rely on fast, intuitive processing (System 1) for familiar content and slower, deliberate reasoning (System 2) when encountering Rich Points (Agar, 1994), such as ambiguous terminology, complex syntax, or cultural disparities. Because the density of such Rich Points varies across domains and correlates with translation difficulty (Lacruz, 2017), current MT systems still largely lack this adaptive reasoning ability. From the perspective of reasoning allocation, existing MT approaches fall into two extremes. On one end, standard large language model (LLM)based translators operate purely in System 1 mode: trained via supervised fine-tuning (SFT) on largescale parallel corpora (Xu et al., 2024a), they produce fluent outputs without explicit reasoning. While efficient, these models struggle with Rich Points and degrade in out-of-domain (OOD) or lowresource settings. Recent efforts have introduced Chain-of-Thought (CoT) prompting into translation (Wang et al., 2025a), but this does not fundamentally solve the problem, because the same reasoning pattern is applied uniformly regardless of input difficulty. Conversely, the emergence of large reasoning models (LRMs), such as DeepSeekR1 (Guo et al., 2025), represents a shift to the opposite extreme—an overcommitment to System 2. Reinforcement learning (RL) is often used to train these models to generate Long CoT traces, applying reasoning uniformly across inputs. This raises a natural question: Can RL serve as a bridge to align the model’s reasoning trajectory with the human translation process? To investigate this, we conduct two preliminary experiments (Section 3). The first examines Pure RL, where RL is applied directly to a base model without SFT on annotated reasoning traces. We find that the model rapidly collapses into repetitive, shallow templates, failing to develop domainspecific reasoning behaviors. The second explores RL with SFT, which fine-tunes on CoT traces be-
Figure 1: Case Study of Adaptive Thought. TwT switches between System 1 and System 2 based on complexity.
fore RL. While this setup produces longer reasoning, it lacks control over when such reasoning is needed, leading to verbose traces even for simple inputs. This indiscriminate reasoning may help reveal Rich Points, but often results in overthinking and excessive token usage, reducing efficiency and human alignment. Despite this, most reasoningbased MT methods still adopt either Pure RL (Feng et al., 2025a) or RL with SFT (Wang et al., 2025c), differing mainly in reward design. However, few attempt to align reasoning effort explicitly with input difficulty. Inspired by human cognitive flexibility, we propose TwT (Translation with Thought), a resourcerational framework for MDMT that learns to allocate inference effort based on input difficulty. As shown in Figure 1, TwT dynamically shifts its reasoning behavior according to input difficulty, using concise reasoning for routine inputs and deeper reasoning for domain-specific challenges. To implement this, TwT follows the RL with SFT pipeline. In the cold-start stage, it performs multi-agent distillation to construct difficulty-adaptive reasoning traces across domains: a domain-specialized teacher (DeepSeek-R1) generates diverse reasoning traces, and GPT-4o assesses input difficulty via Rich Points, rewriting the traces to match the appropriate reasoning depth. This process equips our student model with domain-sensitive reasoning and human-like inference modulation, supporting resource-rational translation. In the RL stage, we optimize the adaptive reasoning behavior seeded during cold-start by rewarding high-quality translations. Our hybrid reward combines translation quality metrics (BLEU and COMET) with a
repetition penalty, guiding the model toward efficient, domain-adaptive reasoning through outcomedriven learning. We conduct a comprehensive evaluation of TwT on 15 benchmarks across in-domain and OOD settings, as well as 3 seen and 59 unseen languages, and perform ablation studies on three different backbone models to assess generalization across both domain and linguistic axes. Our results demonstrate that TwT achieves performance competitive with or superior to SOTA LRMs (e.g., DeepSeek-R1, OpenAI-o1) and surpass strong MTspecialized baselines, while reducing token usage by 32–60%. Empirical analysis confirms that TwT effectively modulates reasoning effort according to task difficulty, leading to more coherent reasoning processes and more accurate translations. This validates the core intuition behind TwT: aligning reasoning effort with input difficulty yields both efficiency and quality gains.
2
Related Work
Recent MT studies increasingly explore explicit reasoning to improve translation quality, starting with shallow strategies such as disambiguation, domain recognition, and self-reflection (Chen et al., 2024; Feng et al., 2025b; Wang et al., 2024b; Hu et al., 2024). To support deeper reasoning, recent studies collect Long CoT traces via MCTS (Zhao et al., 2024) or multi-agent workflows (Wang et al., 2025a), then apply SFT. These traces emulate human translation workflows, improving both performance and interpretability (Chen et al., 2025; Liu et al., 2025). More recently, RL has emerged as a reasoning enhancer (Guo et al., 2025). Sev-
eral approaches optimize translation reasoning with verifiable rewards: R1-T1 uses COMET-based signals (He et al., 2025), MT-R1-Zero combines rulebased and neural metrics (Feng et al., 2025a), DeepTrans employs external LLMs (Wang et al., 2025b), and ExTrans adds exemplar-based guidance (Wang et al., 2025c). However, these methods overlook cognitive alignment; we model human-like reasoning to improve MDMT efficiency and quality.
3
Preliminary Analysis
3.1
Reasoning Collapse
Easy
Model
Hard
Quality Token Time Quality Token Time SFT-Parallel 64.85 General-CoT 61.54 Domain-CoT 67.36 SFT-Parallel 67.52 General-CoT 66.59 Domain-CoT 69.81
In-Domain 13 12 311 72 558 136
61.72 56.53 62.65
40 551 883
15 80 149
Out-of-Domain 13 7 59.11 354 96 60.55 634 125 62.52
36 517 840
8 97 138
Table 2: Performance by difficulty for SFT-Parallel, General-CoT, and Domain-CoT. Quality is computed as the average of BLEU, COMET, and COMETKIWI. Time denotes latency in milliseconds.
Reasoning Template
Rate
I will translate this Chinese sentence into English by identifying the key phrases and their meanings, and then constructing a coherent English sentence.
42.66%
I will translate this Chinese sentence into English by identifying the key phrases and their corresponding meanings.
24.56%
I will translate this Chinese sentence into English.
5.53%
Table 1: Top-3 reasoning template frequency.
We first investigate the R1-Zero paradigm (Pure RL); detailed training settings are given in Appendix E.1. This setup is motivated by recent findings that RL alone can induce spontaneous reasoning capabilities in math and code tasks (Guo et al., 2025). To test whether this emergence transfers to MDMT, we train models using GRPO with hybrid rewards. To rule out the possibility that KL regularization suppresses exploration (Yeo et al., 2025), we monitor token length dynamics both with and without the KL term. However, unlike the “Aha moments” observed in STEM tasks, our experiments reveal rapid mode collapse. As shown in Figures 2(a) and 2(b), reasoning traces quickly collapse into shallow patterns (≤ 100 tokens) regardless of the KL setting. Concretely, the model shifts toward high-frequency template recitation, suppressing diverse reasoning. As shown in Table 1, the top three templates account for about 73% of all generated traces in the Zh→En direction. This degeneration reveals a key misalignment: MDMT requires domain-aware reasoning, yet without proper initialization, the model produces reasoning that is too brief and overly templated to be elicited reliably. To address this, our cold-start phase explicitly initializes adaptive reasoning behavior before RL.
3.2
Reasoning Challenges
To evaluate the trade-off between reasoning depth and computational cost, we compare three Qwen2.5-7B-Instruct variants: SFT-Parallel (System 1), General-CoT (System 2), and Domain-CoT, which extends General-CoT with a domain-aware prompt. Details are given in Appendix E.2. Lack of Domain Awareness. As shown in Table 2, General-CoT performs poorly on in-domain data because its reasoning lacks explicit domain grounding, often defaulting to generic translations rather than domain-specific terminology and fixed expressions. By contrast, SFT-Parallel performs well in these cases by matching the distributional patterns of its training data. This same contrast also explains why General-CoT can be more competitive on OOD data, particularly on harder samples: when the input does not closely match the domain patterns seen in training, intermediate reasoning helps the model better handle syntax, ambiguity, and contextual inference than standard parallel SFT. Importantly, adding explicit domain reasoning largely restores in-domain quality, improving over General-CoT from 61.54 to 67.36 on Easy samples and from 56.53 to 62.65 on Hard samples. This confirms that lack of domain awareness is a major source of in-domain degradation. Reasoning Redundancy. However, domain awareness alone is not sufficient. General-CoT applies essentially the same reasoning strategy regardless of input difficulty, resulting in substantial redundancy. On Easy samples, this redundancy is clearly detrimental: despite using 23.9× more tokens in-domain (311 vs. 13) and 27.2× more on OOD data (354 vs. 13), it still underperforms SFT-
150 100
82
82
27
81
81
24
0
250
500
750
RL Training Steps
1000
18
1250
80 79
250
500
750
RL Training Steps
1000
1250
80 79
78 0
CometKiwi vs Training Step
83
30
21
50
COMET vs Training Step
83
COMET
200
BLEU
Response Length
250
BLEU vs Training Step
33
BLEU COMET CometKiwi BLEU+COMET BLEU+CometKiwi
CometKiwi
Response Length vs Training Step
300
0
250
500
750
RL Training Steps
1000
78
1250
0
250
500
750
RL Training Steps
1000
1250
(a) Pure RL training with KL regularization.
150 100
82
82
28
81
81
26
0
250
500
750
RL Training Steps
1000
1250
22
80 79
0
250
500
750
RL Training Steps
1000
1250
78
CometKiwi vs Step
83
30
24
50
COMET vs Training Step
83
COMET
200
BLEU
Response Length
250
BLEU vs Training Step
32
BLEU COMET CometKiwi BLEU+COMET BLEU+CometKiwi
CometKiwi
Response Length vs Training Step
300
80 79
0
250
500
750
RL Training Steps
1000
1250
78
0
250
500
750
RL Training Steps
1000
1250
(b) Pure RL training without KL regularization.
Figure 2: Training dynamics under pure RL using different quality rewards. While translation quality improves under all settings, pure RL training fails to induce extended translation reasoning traces.
Parallel by 3.31 and 0.93 quality points, respectively. Domain-CoT does not solve this problem. Although it restores in-domain quality, it further increases token usage to 42.9× on in-domain Easy samples (558 vs. 13) and 48.8× on OOD Easy samples (634 vs. 13). Similar patterns also hold on Hard samples. These results show that promptlevel domain awareness alone cannot resolve overthinking, motivating TwT, which adaptively shifts between System 1 and System 2 to jointly address domain sensitivity and reasoning efficiency.
4
Method
We propose TwT, a resource-rational approach that dynamically allocates reasoning effort by input difficulty, mimicking human translation process. As shown in Figure 3, training proceeds in two stages. 4.1
Cold Start
To align the model’s reasoning behavior with the human translation process, we construct a Difficulty-Adaptive CoT Dataset that equips the backbone LLM with adaptive reasoning capabilities. The dataset is curated through a multi-agent distillation-adaptation pipeline. We first prompt DeepSeek-R1 with domain-aware instructions to generate high-quality CoT traces tailored to different domains. These serve as the initial reasoning demonstrations. We then employ GPT-4o, which is verified to align best with human judgment (Ap-
pendix F.7), to assess input difficulty based on the theory of Rich Points. Specifically, difficulty is defined along four linguistic dimensions: sentence complexity, vocabulary rarity, grammatical divergence, and contextual nuance. The corresponding evaluation prompt is shown in Figures 10 and 11. Conditioned on this assessment, GPT-4o rewrites raw traces into adaptive formats: Easy inputs are reformulated into concise System 1 checks to minimize token usage, whereas Hard inputs retain comprehensive System 2 deliberations for structural and terminological verification. We then perform SFT on the backbone LLM using this compact dataset (∼7k examples). We define reasoning depth as the length of the generated reasoning trace (i.e., number of tokens). This process establishes an initial policy distribution over difficulty-aware reasoning strategies, enabling the model to autonomously modulate its reasoning depth. Representative examples are shown in Figures 13 and 14. The resulting data efficiency makes our approach particularly suitable for low-resource settings, demonstrating that robust adaptive reasoning can be achieved with a modest CoT-SFT seed dataset. 4.2
RL Training
To scale this adaptive reasoning behavior, we adopt the GRPO algorithm with hybrid quality rewards, which serve as outcome-driven constraints that encourage the model to identify Rich Points and al-
Stage1: Dataset Curation & Cold Start ① Domain-Aware CoT Data Curation
② Difficulty-Aware CoT Data Curation
③ CoT-FT
Translate the following {src_lang} text into {tgt_lang} while maintaining the domain style of the source text. Difficulty Evaluation
GPT-4o
Rewrite Traces
DeepSeek-R1
Model Size
Identify the domain & Generate domainspecific translation reasoning traces.
Estimate the translation difficulty & Optimize the reasoning trace to match its complexity level.
TwT-7B
TwT-14B
Stage2: RL Training Format Reward
Quality Reward
In-domain test set
TwTPenalty Repetition
TwT
In-Domain Test Set Comparison
������ = 1
������ =− 1
������ = ���� + �����
������ =− 2
� � = {�� = �� , …, �_{� + � − 1} } ������ = 1 −
|��� � � | |� � |
Figure 3: Overview of TwT training pipeline. TwT is first fine-tuned on difficulty-adaptive Long CoT traces distilled from DeepSeek-R1 and rewritten by GPT-4o for cognitive alignment. RL is then applied with a hybrid reward.
locate deep reasoning selectively—only where it leads to measurable quality improvements. By mitigating indiscriminate overthinking (Section 3.2), we improve reasoning efficiency and align the model with the resource-rational principle introduced in Section 1. The final reward r is crafted from three components to ensure alignment: r = rf + rq − λ · rrep Format Reward (rf ). We employ a binary format reward (rf ∈ {1, −1}) to strictly enforce the reasoning-translation structure defined in Figure 5. Hybrid Quality Reward (rq ). To prevent the model from producing plausible but functionally ineffective reasoning, we introduce a verifiable hybrid feedback signal. Building upon our preliminary analysis in Section 3.1, we identified a critical reward hacking phenomenon. As illustrated by the red curve in Figure 2(a), optimizing solely for a semantic metric (CometKiwi/COMET) leads to a significant metric divergence: despite high semantic reward scores, the lexical accuracy (BLEU) degrades rapidly during training. This confirms that the model hacks the reward by generating vague paraphrases or copying source tokens to maximize semantic similarity, effectively abandoning translation fidelity. To remedy this and strictly enforce alignment, we adopt a hybrid quality reward that combines BLEU (B) with COMET (C): ( B(ŷ, y) + C(x, ŷ, y) if rf = 1 rq = −2 if rf = −1
This design effectively stabilizes the optimization process. As evidenced by the training dynamics of our TwT models (Figure 8(a)), both the 7B and 14B models (Figures 8(b) and 8(c)) exhibit synchronous improvements in lexical and semantic metrics without divergence, validating that the hybrid signal successfully grounds the reasoning process in accurate translation outcomes. N-gram Repetition Penalty (rrep ). We penalize redundant reasoning to discourage degenerate loops and promote efficient token usage. For a given reasoning trace c, let G(c) denote the list of contiguous n-grams. In our experiments, we set n = 20. We compute the ratio of repeated n-grams to discourage repetitive, loop-like patterns: rrep = 1 −
|set (G(c))| ∈ [0, 1] |G(c)|
5
Experiments
5.1
Experimental Settings
Dataset. We use two datasets for training: (1) a curated set of 7K difficulty-adaptive Long CoT examples spanning 10 domains and three translation directions (De→En, En→Zh, Zh→En) for coldstart SFT, and (2) a separate 20K-sample dataset for RL, constructed from multi-domain parallel corpora. For evaluation, we adopt both in-domain test sets and diverse OOD benchmarks, and additionally include multilingual test sets covering both seen and unseen language pairs. Full dataset details are provided in Appendix B.
Laws
Method
Quality
News
Tokens
Quality
Science
Tokens
Quality
Subtitles
Tokens
DeepSeek-V3 Gemini-2.0-Flash GPT-4o
77.72 76.57 73.77
-
69.41 69.30 68.61
-
69.09 68.80 68.09
-
DeepSeek-R1 Gemini-2.0-Flash-Thinking OpenAI-o3-mini OpenAI-o1 GPT-5 QwQ-32B
77.72 76.21 71.61 73.85 76.28 71.84
577 702 428 478 784 667
68.47 68.24 68.18 68.67 69.07 67.88
498 1149 443 408 740 584
68.46 68.17 68.01 68.60 68.58 67.86
478 1092 385 367 606 563
SFT-Parallel-7B ALMA-7B-R ALMA-13B-R TowerInstruct-7B-v0.2 TowerInstruct-13B-v0.1 CoT-FT-7B MT-R1-Zero-7B SSR-X-Zero-7B mExTrans-7B
76.58 67.88 70.11 73.91 74.65 76.72 68.93 69.59 70.11
51 72 56 597
66.08 63.37 64.65 65.93 66.90 66.24 67.41 66.00 65.49
42 64 52 553
66.38 62.77 64.23 65.45 66.20 66.07 66.98 66.77 66.13
39 61 49 546
TwT-Qwen2.5-7B-Instruct TwT-Qwen2.5-14B-Instruct
75.35 76.58
310 320
68.42 68.62
311 285
68.24 68.32
294 272
Quality
Literary
Tokens
Quality
Large Language Models 62.93 56.83 62.96 57.41 62.92 57.47 Large Reasoning Models 61.81 514 53.95 61.98 708 57.23 62.42 355 56.97 62.62 340 57.28 62.74 519 56.27 61.80 584 55.09 MT-Specialized Models 62.87 55.96 59.38 54.52 60.22 55.30 61.07 55.32 61.89 56.09 62.76 29 55.57 61.97 55 55.80 61.89 39 55.54 59.52 476 54.40 Our Models 63.03 247 57.86 62.97 241 58.14
IT
Koran
Medical
Average
Tokens
Quality
Tokens
Quality
Tokens
Quality
Tokens
Quality
Tokens
-
66.97 66.54 66.37
-
57.73 58.13 57.78
-
69.16 70.21 69.34
-
66.23 66.24 65.54
-
574 781 546 521 859 863
66.28 66.20 65.99 66.45 66.57 61.44
593 345 343 403 492 583
57.48 58.13 56.42 57.94 58.71 55.81
790 677 511 506 751 963
69.01 69.73 68.40 69.20 70.14 66.68
667 415 346 441 531 735
65.40 65.74 64.75 65.58 66.05 63.55
586 734 420 433 660 693
52 69 54 610
68.02 64.20 64.54 66.78 67.18 67.82 65.60 61.22 60.50
35 56 36 452
58.05 55.04 55.71 50.05 49.93 57.57 55.52 55.53 55.41
45 71 46 604
70.21 67.46 68.40 70.73 71.49 70.38 64.43 63.99 63.00
46 71 50 565
65.52 61.83 62.89 63.66 64.29 65.39 63.33 62.56 61.82
42 65 48 551
281 354
68.07 68.37
222 234
58.92 59.35
269 336
70.29 70.59
262 287
66.27 66.62
274 291
Table 3: In-domain translation performance across eight domains, averaged over En→Zh, Zh→En, and De→En. The Bold and underlined values denote the highest and second highest scores, respectively. Method
Conversation Quality
Tokens
Ecommerce Quality
Social
Tokens
DeepSeek-V3 Gemini-2.0-Flash GPT-4o
68.42 68.77 68.75
-
66.34 66.26 66.50
-
DeepSeek-R1 Gemini-2.0-Flash-Thinking OpenAI-o3-mini OpenAI-o1 GPT-5
67.06 68.36 68.04 68.47 68.40
534 1204 290 327 448
64.45 65.83 65.87 65.80 65.49
552 822 363 399 609
SFT-Parallel-7B ALMA-7B-R ALMA-13B-R CoT-FT-7B MT-R1-Zero-7B SSR-X-Zero-7B mExTrans-7B
65.64 64.71 66.03 65.43 66.59 65.70 63.47
31 53 37 464
63.33 62.54 63.37 63.26 64.16 63.46 61.74
45 66 50 566
TwT-Qwen2.5-7B-Instruct TwT-Qwen2.5-14B-Instruct
67.72 67.77
231 240
65.71 65.79
273 309
Quality
Culture Tokens
Large Language Models 66.10 66.10 66.10 Large Reasoning Models 64.11 554 65.50 1081 65.58 372 65.38 405 65.14 652 MT-Specialized Models 62.54 62.68 63.51 62.06 42 63.69 65 63.31 49 61.20 555 Our Models 65.48 269 65.47 298
CommonSense
Average
Quality
Tokens
Quality
Tokens
Quality
Tokens
69.65 69.02 69.01
-
65.96 65.23 65.89
-
67.29 67.08 67.25
-
68.16 68.42 67.55 67.96 68.33
560 1220 596 542 984
64.34 65.44 64.45 64.78 64.09
602 2335 436 392 530
65.62 66.71 66.30 66.48 66.29
561 1332 411 413 645
65.34 66.63 60.94 64.66 66.23 64.35 65.11
54 79 66 631
61.03 62.02 62.91 61.08 62.32 62.18 59.92
33 51 34 470
63.58 63.71 63.35 63.30 64.60 63.80 62.29
41 63 47 537
67.82 68.55
352 333
64.48 64.72
219 259
66.25 66.46
269 288
Table 4: OOD translation performance across five domains, averaged over En→Zh, Zh→En, and De→En.
Implementation Details. For cold start, we use LLaMA-Factory1 (Zheng et al., 2024) with Qwen2.5-7B-Instruct, Qwen2.5-14B-Instruct, and Gemma-2-9B-IT as backbones. We train them on the 7K difficulty-adaptive Long CoT examples for 1 epoch with full-parameter optimization on 8 NVIDIA A100 80GB GPUs, using AdamW with a learning rate of 1e−5, a total batch size of 32, a cosine learning rate scheduler, a warm-up ratio of 0.1, a maximum input sequence length of 4096, and DeepSpeed ZeRO Stage 3. The coldstart stage completes within 10 minutes. For RL, we use verl2 (Sheng et al., 2025) and train for 1 epoch on 8 NVIDIA A100 80GB GPUs with a total batch size of 16, rollout number 16, rollout temperature 1.0, learning rate 1e−6, KL loss coefficient β = 1e−3, maximum response length 2048, and repetition penalty with n = 20. RL training 1 2
https://github.com/hiyouga/LLaMA-Factory https://github.com/volcengine/verl
takes about 10 hours. During inference, we use vLLM3 (Kwon et al., 2023) for efficient decoding with temperature 0.0 and repetition penalty 1.05. Metrics. We report Quality, defined as the average of BLEU, COMET (Rei et al., 2020), and CometKiwi (Rei et al., 2022), and Tokens, the average length of the generated CoT. The full metric breakdowns are provided in Appendix I. Baselines. We compare our method against three categories of models: (1) general-purpose LLMs such as DeepSeek-V3 (DeepSeek-AI et al., 2024), Gemini-2.0-Flash (DeepMind, 2024), GPT4o (OpenAI et al., 2024), and open-source models like LLaMA3.1-8B-Instruct (Grattafiori et al., 2024), Gemma-2-9B-IT (Gemma Team et al., 2024), and Qwen2.5 series (7B, 14B, 32B) (Yang et al., 2024). (2) reasoning-oriented LRMs, such as DeepSeek-R1 (Guo et al., 2025), Gemini3
https://github.com/vllm-project/vllm
En→Zh
Method
Quality
Zh→En
Tokens
Quality
Qwen2.5-7B-Instruct Gemma-2-9B-IT
67.87 66.50
-
50.91 52.64
ALMA-7B-R Tower-Plus-9B SFT-Parallel-7B MT-R1-Zero-7B SSR-X-Zero-7B mExTrans-7B
64.08 69.85 69.44 67.52 68.55 66.70
62 50 537
54.52 57.41 55.96 55.80 56.52 54.40
TwT-Qwen2.5-7B-Instruct TwT-Gemma-2-9B-IT
69.99 69.07
298 227
57.86 58.12
De→En
Tokens
Quality
Tokens
Large Language Models 62.90 60.64 MT-Specialized Models 62.89 67.04 66.15 69 62.56 66 54 62.86 45 610 60.74 544 Our Models 281 66.45 256 249 66.95 218
En→X
X→En
Average
Quality
Tokens
Quality
Tokens
Quality
Tokens
37.61 54.28
-
56.99 66.38
-
55.25 60.09
-
44.68 46.46 32.44 40.98 40.43 42.72
358 306 1047
44.71 63.02 55.70 58.33 58.19 55.99
74 48 731
54.18 60.76 55.94 57.04 57.31 56.11
126 101 694
41.13 54.65
483 280
58.84 67.04
328 257
58.85 63.17
329 246
Table 5: Results on seen and unseen language directions. En, Zh, and De are seen languages, while X denotes unseen languages; En→X and X→En report averages over English↔unseen-language directions.
Method TwT-Qwen2.5-7B-Instruct w/o RP w/o RP + w/o Adaptive CoT w/o RP + w/o Cold Start SFT only w/ Adaptive CoT SFT only w/ domain-aware CoT SFT only w/ general CoT
In-Domain
Out-of-Domain
Quality
Tokens
Quality
Tokens
64.77 64.52 64.68 64.18 62.47 61.65 61.39
278 301 748 62 253 479 531
66.28 66.12 66.17 65.29 64.09 63.47 63.11
263 271 720 59 219 452 519
64.71 63.33 63.68
-
66.28 64.44 66.05
263 298 220
Backbone Models 61.99 59.84 59.93 Our Models TwT-Qwen2.5-7B-Instruct 64.77 278 TwT-Llama-3.1-8B-Instruct 63.20 305 TwT-Gemma-2-9B-IT 64.71 231
Qwen2.5-7B-Instruct Llama-3.1-8B-Instruct Gemma-2-9B-IT
Table 6: Ablation study on in-domain and OOD translation test sets. Results are averaged at the dataset level for each setting. RP = repetition penalty.
2.0-Flash-Thinking (DeepMind, 2025), OpenAI o1 (Jaech et al., 2024), o3-mini (OpenAI, 2025b), GPT-5 (OpenAI, 2025a), and QwQ-32B (QwenTeam, 2025); (3) MT-specialized models, including non-reasoning LLMs such as TowerInstruct (Alves et al., 2024), ALMA-R (Xu et al., 2024a,b), SFTParallel (Qwen2.5-7B-Instruct fine-tuned on 27K parallel pairs), CoT-FT (Hu et al., 2024), and Tower-Plus-9B (Rei et al., 2025); with reasoningoriented models including MT-R1-Zero (Feng et al., 2025a), mExTrans (Wang et al., 2025c), and SSRX-Zero (Yang et al., 2025). For fairer comparison, Appendix C details the training data size of MTspecialized baselines. 5.2
Cross-Domain Generalization
In-Domain. As shown in Table 3, TwT-14B achieves a SOTA average quality score of 66.62, outperforming both strong open-source LRMs such
as DeepSeek-R1 (65.40) and closed-source models including GPT-5 (66.05). Notably, our smaller variant, TwT-7B, also surpasses dedicated MT systems, demonstrating that our method scales effectively with model size. Compared to mExTrans7B, which is also trained under the R1 paradigm, TwT-7B yields a substantial improvement of +4.45 points while reducing reasoning overhead by 50.27%, highlighting the superior efficiency of our difficulty-adaptive mechanism. The benefits of adaptive reasoning are especially evident in structurally complex and culturally nuanced domains. In the Literary domain, TwT-14B consistently outperforms three representative paradigms: it exceeds the System 1 baseline SFT-Parallel by +2.18 points, the heavy-reasoning System 2 model DeepSeek-R1 by +4.19 points, and the pure RL-based MT-R1Zero-7B by +2.34 points. These results suggest that, for multi-domain translation, neither shallow System 1 execution, indiscriminate System 2 overthinking, nor unguided pure RL exploration alone yields optimal performance. In contrast, our model autonomously modulates reasoning depth to strike a more effective balance between literal accuracy and stylistic adequacy. OOD. TwT-7B achieves a strong average score of 66.25 on five OOD test sets, surpassing MTspecialized baselines such as SFT-Parallel-7B (63.58), which exhibit poor generalization under domain shift. Moreover, TwT-7B exhibits a favorable quality-efficiency trade-off, surpassing DeepSeek-R1 while reducing reasoning overhead by 292 tokens. Even in unfamiliar domains, it avoids excessive deliberation by leveraging compact, internalized translation procedures. These results indicate that TwT captures domain-agnostic
translation logic rather than relying on domainspecific memorization, approaching the SOTA performance of GPT-4o at a fraction of the computational cost and model size. 5.3
Multilingual Generalization
On seen directions, TwT-7B improves Zh→En performance by +6.95 over its base model and +1.90 over SFT-Parallel-7B, under the same training data. On unseen directions (En→X)4 , TwT-Gemma-2-9B-IT, trained on only 27K examples, outperforms the multilingual Tower-Plus-9B, built on the same backbone but trained on 286K examples, by a substantial margin of +8.19. Overall, TwT achieves the highest average score across all directions, indicating that its reasoning mechanism generalizes beyond language boundaries and captures transferable alignment strategies. 5.4
Ablation Study
We conduct a comprehensive ablation study to assess the contribution of each component in the TwT framework and validate its generalizability across backbone architectures (Table 6). Removing the repetition penalty (w/o RP) leads to a slight quality drop and longer outputs, indicating its role as a regularizer rather than a performance driver. In contrast, removing both the repetition penalty and the difficulty-adaptive rewriting (w/o RP + w/o Adaptive CoT) results in comparable quality but significantly increases reasoning length (from 278 to 748 tokens), highlighting the critical role of adaptive rewriting in controlling verbosity and ensuring inference efficiency. We further examine the necessity of the two-stage training pipeline. Eliminating the cold-start SFT phase (w/o RP + w/o Cold Start) causes performance degradation and length collapse (to 62 tokens), suggesting that RL alone fails to induce structured reasoning behavior. To isolate the impact of SFT data quality, we compare three variants: difficulty-adaptive CoT yields the best result (62.47), followed by domain-aware (61.65) and general CoT (61.39), showing a clear performance hierarchy. Still, only the full TwT pipeline achieves the highest score (64.77), confirming that RL is indispensable for turning the adaptive reasoning patterns from mere imitation into an internalized and optimized translation strategy. Lastly, we apply TwT to three backbone models: Qwen2.5-7B, 4
En→X and X→En are averaged over 59 unseen languages from FLORES+; see Appendix B.6 for the full list.
Llama-3.1-8B, and Gemma-2-9B, and observe consistent improvements. For instance, TwT improves Gemma-2-9B-IT from 59.93 to 64.71 in-domain and from 63.68 to 66.05 OOD, demonstrating that TwT is a model-agnostic framework that robustly enhances translation reasoning regardless of the underlying architecture.
6
Empirical Analysis
6.1
Human Reasoning Alignment
To benchmark TwT’s reasoning against human cognition, we conduct a qualitative analysis on 10 Zh→En examples, with expert commentary from a translation studies faculty member. A representative case is shown in Appendix G.2. Cognitive Convergence. The analysis revealed that TwT’s reasoning exhibits strong parallels with human translators in early-stage decision-making. Specifically, TwT effectively (1) identifies translation domain and stylistic register, (2) handles complex sentence structures with appropriate syntactic parsing, and (3) demonstrates context-aware terminology adaptation. For example, it consistently distinguishes between literary and technical expressions and adjusts lexical choices accordingly. Its structured CoT mirrors key aspects of professional reasoning—such as coherence maintenance, discourse flow control, and sensitivity to stylistic norms—indicating that TwT has internalized domain-aware reasoning behavior resembling human translation logic. Pragmatic Divergence. Despite these strengths, TwT still shows gaps compared with expert translators. It occasionally struggles with cross-sentence consistency in terminology, especially when handling long-form repetitions or abbreviated references. Moreover, its output lacks fine-grained control over tone, idiomaticity, and cultural adaptation, which human translators adjust based on pragmatic context and target audience. These issues suggest that TwT’s reasoning remains less flexible in discourse-level adaptation, reflecting the absence of high-level pragmatic awareness. Future work seeks to bridge these gaps by integrating processoriented feedback, thereby fostering deeper pragmatic alignment with human cognitive processes. 6.2
Reasoning Redundancy Reduction
To assess whether TwT eliminates unnecessary computation, we employ DeepSeek-V3.2 to detect six
Over-segmentation Unnecessary linguistic explanation Semantic repetition Irrelevant information Redundant alternative translations Low-density long descriptions
Semantic Space Exploration (PCA)
Not Resolved (%)
87.3 93.2 96.9 93.7 96.5 97.2
12.7 6.8 3.1 6.3 3.5 2.8
Pairwise Reasoning Similarity Distribution
Pure-RL TwT (Ours)
0.4
3.0
0.2
2.5
0.0
2.0
Density
Resolved (%)
Principal Component 2
Redundancy Type
Easy
Method DeepSeek-R1 Gemini-2.0-Flash-Thinking OpenAI-o3-mini OpenAI-o1 QwQ-32B General-CoT TwT-Qwen2.5-7B-Instruct TwT-Qwen2.5-14B-Instruct
Medium
Hard
All
Quality
Tokens
Quality
Tokens
Quality
Tokens
Quality
Tokens
66.30 66.62 66.57 67.14 62.91 61.54 67.93 67.98
480 432 280 302 577 311 216 207
63.82 64.34 63.44 64.26 61.92 60.27 64.85 65.23
556 754 430 432 758 433 272 300
62.15 62.93 61.44 62.49 60.52 56.53 62.78 63.26
579 1035 596 578 844 551 332 378
64.09 64.63 63.82 64.63 61.78 59.45 65.19 65.49
538 740 435 437 726 432 273 295
Table 8: In-domain performance by difficulty level.
distinct forms of reasoning redundancy across 15 domains, utilizing the prompt provided in Figure 12. As shown in Table 7, TwT demonstrates exceptional efficiency, successfully resolving over 94% of the redundant steps observed in the SFTwith-RL baseline described in Section 3.2. Notably, it achieves a 97.2% resolution rate for Low-Density Long Descriptions, verifying its ability to compress verbose reasoning into high-density insights, while maintaining 87.3% resolution for structural issues like Over-Segmentation. These results confirm that the model has internalized resource-rational adaptive reasoning behavior, effectively activating System 2 reasoning for Rich Points while avoiding unnecessary elaboration on straightforward segments. 6.3
Translation Difficulty Adaptation
To better understand how models adapt their reasoning behavior to translation difficulty, we group the in-domain test set into three difficulty levels (Easy, Medium, Hard) estimated by DeepSeek-V3 (prompt in Figure 11). Table 8 reports the average quality and the response length for each group, averaged across all domains. Results indicate that TwT effectively addresses the reasoning redundancy of General-CoT through resource-rational allocation. On Easy inputs, it reduces token usage by 33% while improving quality by +6.4 points; conversely, on Hard inputs, it focuses on performance, achieving a substantial +6.7 points quality gain. Moreover, TwT models consistently outperform all SOTA LRMs across all difficulty levels in terms of quality, while maintaining significantly shorter reasoning traces—reducing average token usage by 32% compared to OpenAI-o3-mini and by 60% compared to Gemini-2.0-Flash-Thinking. This highlights TwT’s ability to generate concise, difficulty-aware reason-
1.5
0.2 1.0 0.4
Table 7: Resolution rates of six redundancies by TwT.
Pure-RL TwT (Ours)
0.5
0.6
0.4
0.2
0.0
Principal Component 1
0.2
0.4
0.0 0.0
0.2
0.4
0.6
Cosine Similarity
0.8
1.0
Figure 4: CoT trace similarity comparison. Pure RL vs TwT in space (Left) and distribution (Right).
ing while reducing overthinking. 6.4
Reasoning Collapse Mitigation
In Section 3.1, we identified a critical failure mode of pure RL, where reasoning rapidly degenerates into shallow, repetitive templates. To verify whether TwT successfully mitigates this reasoning collapse, we conduct a semantic diversity analysis across 15 domains. We compute pairwise cosine similarities of generated reasoning traces using multilingual Sentence-BERT5 . Pure RL exhibits severe redundancy with a mean similarity of 0.89, whereas TwT significantly reduces this metric to 0.51. This divergence is visually corroborated by Figure 4, where the PCA projection (Left) shows pure RL confined to tight, isolated clusters compared to the broad semantic manifold of TwT, and the similarity histogram (Right) confirms that TwT diffuses the sharp redundancy peak of the baseline into a balanced distribution. These results show that TwT overcomes template dependency and encourages genuine reasoning. 6.5
Further Analysis
Appendix F provides additional analyses of TwT, including MQM error types, KL ablation, training dynamics, domain-aware prompting, inference cost, language consistency, and failure cases.
7
Conclusion
In this work, we present TwT, a resource-rational translation model that adapts reasoning effort to input difficulty. TwT combines difficulty-aware SFT and hybrid-reward RL to balance System 1 and System 2 behavior. Evaluated across diverse domains and languages, TwT matches or surpasses SOTA LRMs while reducing token usage by 32–60%, validating the effectiveness of aligning translation with human reasoning economy. 5
sentence-transformers/paraphrase-multilingualMiniLM-L12-v2
Limitations While TwT achieves robust performance across multiple domains, several limitations remain. First, the RL training data is randomly sampled without controlling for difficulty distribution, which may result in an imbalanced mix of easy, medium, and hard inputs. Second, the reasoning traces distilled from proprietary LLMs (e.g., DeepSeek-R1) may carry over implicit biases or domain preferences inherent in those models. Although our current setup yields consistent improvements, such biases could influence the reasoning behavior or stylistic tendencies of TwT. Finally, our current reward design does not incorporate difficulty-aware reward shaping. In particular, no length-based reward is applied to encourage concise reasoning on simple inputs and more detailed analysis for complex ones. Incorporating such adaptive rewards may further enhance the model’s ability to adjust reasoning depth based on input complexity in MDMT. We leave this direction for future work.
Acknowledgment This work is supported by the National Science and Technology Major Project (Grant No. 2022ZD0116101), the National Natural Science Foundation of China (NSFC) under Grant No. 62206295, the Major Scientific Research Project of the State Language Commission in the 13th FiveYear Plan (Grant No. WT135-38), the public technology service platform project of Xiamen City (No. 3502Z20231043). In addition, we used a large language model to assist in polishing the visualizations in Figure 1 and generating certain decorative visual elements in Figure 3.
References M. Agar. 1994. Language Shock: Understanding The Culture Of Conversation. HarperCollins. Roee Aharoni and Yoav Goldberg. 2020. Unsupervised domain clusters in pretrained language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7747– 7763, Online. Association for Computational Linguistics. Duarte M. Alves, José Pombal, Nuno M. Guerreiro, Pedro H. Martins, João Alves, Amin Farajian, Ben Peters, Ricardo Rei, Patrick Fernandes, Sweta Agrawal, Pierre Colombo, José G. C. de Souza, and André F. T. Martins. 2024. Tower: An open multilingual large language model for translation-related tasks. Preprint, arXiv:2402.17733.
Andong Chen, Lianzhang Lou, Kehai Chen, Xuefeng Bai, Yang Xiang, Muyun Yang, Tiejun Zhao, and Min Zhang. 2024. DUAL-REFLECT: Enhancing large language models for reflective translation through dual learning feedback mechanisms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 693–704, Bangkok, Thailand. Association for Computational Linguistics. Andong Chen, Yuchen Song, Wenxin Zhu, Kehai Chen, Muyun Yang, Tiejun Zhao, and Min zhang. 2025. Evaluating o1-like llms: Unlocking reasoning for translation through comprehensive analysis. Preprint, arXiv:2502.11544. Google DeepMind. 2024. Introducing gemini 2.0: our new ai model for the agentic era. https: //blog.google/technology/google-deepmind/ google-gemini-ai-update-december-2024/ #ceo-message. Accessed: 2025-04-21. Google DeepMind. 2025. Gemini 2.0 flash thinking. https://ai.google.dev/gemini-api/ docs/changelog. Accessed: 2025-04-21. DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, and 181 others. 2024. Deepseek-v3 technical report. Preprint, arXiv:2412.19437. Zhaopeng Feng, Shaosheng Cao, Jiahan Ren, Jiayuan Su, Ruizhe Chen, Yan Zhang, Zhe Xu, Yao Hu, Jian Wu, and Zuozhu Liu. 2025a. Mt-r1-zero: Advancing llm-based machine translation via r1-zero-like reinforcement learning. Preprint, arXiv:2504.10160. Zhaopeng Feng, Yan Zhang, Hao Li, Bei Wu, Jiayu Liao, Wenqiang Liu, Jun Lang, Yang Feng, Jian Wu, and Zuozhu Liu. 2025b. TEaR: Improving LLM-based machine translation with systematic self-refinement. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 3922–3938, Albuquerque, New Mexico. Association for Computational Linguistics. Gemma Team, Morgane Riviere, and 1 others. 2024. Gemma 2: Improving open language models at a practical size. Preprint, arXiv:2408.00118. Daniel Gile and Victoria Lei. 2020. Translation, effort and cognition. In The Routledge handbook of translation and cognition, pages 263–278. Routledge. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad AlDahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, and 542 others. 2024. The llama 3 herd of models. Preprint, arXiv:2407.21783.
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, and 180 others. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. Preprint, arXiv:2501.12948. Jie He, Tao Wang, Deyi Xiong, and Qun Liu. 2020. The box is in the pen: Evaluating commonsense reasoning in neural machine translation. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3662–3672, Online. Association for Computational Linguistics.
Isabel Lacruz. 2017. Cognitive effort in translation, editing, and post-editing. The handbook of translation and cognition, pages 386–401. Zihao Li, Shaoxiong Ji, and Jörg Tiedemann. 2025. Testtime scaling of reasoning models for machine translation. Preprint, arXiv:2510.06471. Sinuo Liu, Chenyang Lyu, Minghao Wu, Longyue Wang, Weihua Luo, Kaifu Zhang, and Zifu Shang. 2025. New trends for modern machine translation with large reasoning models. Preprint, arXiv:2503.10351.
Minggui He, Yilun Liu, Shimin Tao, Yuanchang Luo, Hongyong Zeng, Chang Su, Li Zhang, Hongxia Ma, Daimeng Wei, Weibin Meng, Hao Yang, Boxing Chen, and Osamu Yoshie. 2025. R1-t1: Fully incentivizing translation capability in llms via reasoning learning. Preprint, arXiv:2502.19735.
NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, and 20 others. 2024. Scaling neural machine translation to 200 languages. Nature, 630(8018):841–846.
Tianxiang Hu, Pei Zhang, Baosong Yang, Jun Xie, Derek F. Wong, and Rui Wang. 2024. Large language model for multi-domain translation: Benchmarking and domain CoT fine-tuning. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 5726–5746, Miami, Florida, USA. Association for Computational Linguistics.
OpenAI, :, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander Madry, ˛ Alex Baker-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, Alex Kirillov, and 401 others. 2024. Gpt-4o system card. Preprint, arXiv:2410.21276.
Kristian Tangsgaard Hvelplund. 2011. Allocation of cognitive resources in translation: An eye-tracking and key-logging study. Frederiksberg: Copenhagen Business School (CBS).
OpenAI. 2025a. Introducing GPT-5. https://openai. com/zh-Hans-CN/index/introducing-gpt-5/.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, Alex Iftimie, Alex Karpenko, Alex Tachard Passos, Alexander Neitz, Alexander Prokofiev, Alexander Wei, Allison Tam, Ally Bennett, Ananya Kumar, and 242 others. 2024. Openai o1 system card. Preprint, arXiv:2412.16720. Tom Kocmi, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Novák, Martin Popel, and Maja Popović. 2022. Findings of the 2022 conference on machine translation (WMT22). In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 1–45, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, SOSP ’23, page 611–626, New York, NY, USA. Association for Computing Machinery.
OpenAI. 2025b. Openai o3-mini. https://openai. com/index/openai-o3-mini/. Accessed: 202504-21. Kishore Papineni, Salim Roukos, Todd Ward, and WeiJing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics. Qwen-Team. 2025. Qwq-32b: Embracing the power of reinforcement learning. Ricardo Rei, Nuno M. Guerreiro, José Pombal, João Alves, Pedro Teixeirinha, Amin Farajian, and André F. T. Martins. 2025. Tower+: Bridging generality and translation specialization in multilingual llms. Preprint, arXiv:2506.17080. Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics. Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, José G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and André F. T.
Martins. 2022. CometKiwi: IST-unbabel 2022 submission for the quality estimation shared task. In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 634–645, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics. John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. Preprint, arXiv:1707.06347. Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. Preprint, arXiv:2402.03300. Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. 2025. Hybridflow: A flexible and efficient rlhf framework. In Proceedings of the Twentieth European Conference on Computer Systems, EuroSys ’25, page 1279–1297, New York, NY, USA. Association for Computing Machinery. Liang Tian, Derek F. Wong, Lidia S. Chao, Paulo Quaresma, Francisco Oliveira, Yi Lu, Shuo Li, Yiming Wang, and Longyue Wang. 2014. UM-corpus: A large English-Chinese parallel corpus for statistical machine translation. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC‘14), pages 1837–1842, Reykjavik, Iceland. European Language Resources Association (ELRA). Jiaan Wang, Fandong Meng, Yunlong Liang, and Jie Zhou. 2025a. Drt: Deep reasoning translation via long chain-of-thought. Preprint, arXiv:2412.17498. Jiaan Wang, Fandong Meng, and Jie Zhou. 2025b. Deep reasoning translation via reinforcement learning. Preprint, arXiv:2504.10187. Jiaan Wang, Fandong Meng, and Jie Zhou. 2025c. Extrans: Multilingual deep reasoning translation via exemplar-enhanced reinforcement learning. Preprint, arXiv:2505.12996. Longyue Wang, Siyou Liu, Chenyang Lyu, Wenxiang Jiao, Xing Wang, Jiahao Xu, Zhaopeng Tu, Yan Gu, Weiyu Chen, Minghao Wu, Liting Zhou, Philipp Koehn, Andy Way, and Yulin Yuan. 2024a. Findings of the WMT 2024 shared task on discourse-level literary translation. In Proceedings of the Ninth Conference on Machine Translation, pages 699–700, Miami, Florida, USA. Association for Computational Linguistics. Longyue Wang, Zhaopeng Tu, Yan Gu, Siyou Liu, Dian Yu, Qingsong Ma, Chenyang Lyu, Liting Zhou, ChaoHong Liu, Yufeng Ma, Weiyu Chen, Yvette Graham, Bonnie Webber, Philipp Koehn, Andy Way, Yulin Yuan, and Shuming Shi. 2023. Findings of the WMT
2023 shared task on discourse-level literary translation: A fresh orb in the cosmos of LLMs. In Proceedings of the Eighth Conference on Machine Translation, pages 55–67, Singapore. Association for Computational Linguistics. Yutong Wang, Jiali Zeng, Xuebo Liu, Fandong Meng, Jie Zhou, and Min Zhang. 2024b. TasTe: Teaching large language models to translate through selfreflection. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6144–6158, Bangkok, Thailand. Association for Computational Linguistics. Haoran Xu, Young Jin Kim, Amr Sharaf, and Hany Hassan Awadalla. 2024a. A paradigm shift in machine translation: Boosting translation performance of large language models. In The Twelfth International Conference on Learning Representations. Haoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfeng Shen, Benjamin Van Durme, Kenton Murray, and Young Jin Kim. 2024b. Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation. In Forty-first International Conference on Machine Learning. An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, and 23 others. 2024. Qwen2.5 technical report. Preprint, arXiv:2412.15115. Wenjie Yang, Mao Zheng, Mingyang Song, Zheng Li, and Sitong Wang. 2025. Ssr-zero: Simple selfrewarding reinforcement learning for machine translation. Preprint, arXiv:2505.16637. Binwei Yao, Ming Jiang, Tara Bobinac, Diyi Yang, and Junjie Hu. 2024. Benchmarking machine translation with cultural awareness. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 13078–13096, Miami, Florida, USA. Association for Computational Linguistics. Edward Yeo, Yuxuan Tong, Morry Niu, Graham Neubig, and Xiang Yue. 2025. Demystifying long chain-of-thought reasoning in llms. Preprint, arXiv:2502.03373. Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Tiantian Fan, Gaohong Liu, Lingjun Liu, Xin Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang, Mofan Zhang, Wang Zhang, Hang Zhu, and 16 others. 2025. Dapo: An open-source llm reinforcement learning system at scale. Preprint, arXiv:2503.14476. Yu Zhao, Huifeng Yin, Bo Zeng, Hao Wang, Tianqi Shi, Chenyang Lyu, Longyue Wang, Weihua Luo, and Kaifu Zhang. 2024. Marco-o1: Towards open reasoning models for open-ended solutions. Preprint, arXiv:2411.14405.
Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, and Zheyan Luo. 2024. LlamaFactory: Unified efficient fine-tuning of 100+ language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 400–410, Bangkok, Thailand. Association for Computational Linguistics.
A
GRPO Algorithm
GRPO (Shao et al., 2024) extends PPO (Schulman et al., 2017) by removing the dependency on a value model and instead leveraging group-wise relative rewards estimation among sampled responses for more stable and efficient policy updates. Given a query x, the model samples a group of G responses {yi }G i=1 , each scored with reward ri . The normalized advantage for each sample is computed as: Ai =
ri − mean({r}G j=1 ) std({r}G j=1 )
i=1
old
A conversation between User and Assistant. The user asks a translation question, and the Assistant solves it. The Assistant first thinks about the translation reasoning process in the mind, and then provides the final translation. The translation reasoning process and the final translation are enclosed within <think> </think> and <answer> </answer> tags, respectively, i.e., <think> translation reasoning process here </think> <answer> final translation here </answer>.\n\n User: {Translation question}.\n Assistant: <think> Figure 5: Template for pure RL in MT task.
.
(1)
Then GRPO optimizes the policy model πθ by maximizing the following objective: JGRPO (θ) = Ex∼D, {yi }G ∼πθ
Template
German-English multi-domain • The dataset (Aharoni and Goldberg, 2020), including five distinct domains: IT, Law, Medical, Koran, and Subtitles.
(·|x)
G π (y | x) 1 X θ i min Ai , G πθold (yi | x) i=1 π (y | x) θ i clip , 1 − ϵ, 1 + ϵ Ai πθold (yi | x) ! − βDKL πθ πref ,
(2) where πθold and πθ are the old and current policies, ϵ is the PPO clipping threshold, and β controls the weight of the KL regularization.
• The English-Chinese UM-Corpus (Tian et al., 2014), covering four domains: News, Laws, Subtitles, and Science. • The Chinese-English GuoFeng-Webnovel dataset (Wang et al., 2023, 2024a) from WMT23 and WMT24 literary translation tasks, representing the Literary domain.
B
Datasets
For each domain, we randomly select 2K sentence pairs with a minimum source sentence length of 20 words (characters for Chinese) to ensure meaningful reasoning potential. This results in a total of 20K training samples used for RL training.
B.1
Details of Cold Start Data
B.3
For cold-start SFT, we collect a curated dataset of about 7K difficulty-adaptive Long CoT examples spanning 10 domains and three major translation directions: De→En, En→Zh, and Zh→En. The data is constructed via domain-aware generation with DeepSeek-R1 followed by difficulty-adaptive rewriting with GPT-4o (see Section 4.1). Detailed dataset statistics are presented in Figure 6. B.2
Details of RL Training Data
We collect a diverse MDMT dataset for RL training across languages and domains. Specifically, we sample from the following sources:
In-Domain Test Data
For in-domain evaluation, we use the official test sets associated with the corpora in Appendix B.2. For the Literary domain, we merge the valid_1, valid_2, test_1, and test_2 subsets to form a comprehensive test set. The data statistics is illustrated in Table 9. B.4
Out-of-Domain Test Data
For out-of-domain evaluation, we consider a diverse set of test sets spanning multiple language pairs and domains. Specifically, the Conversation, Ecommerce, and Social domains are drawn from the WMT22 shared tasks (Kocmi et al.,