Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss Ronald Skorobogat Ameya Prabhu† Matthias Bethge† Tübingen AI Center, University of Tübingen Leaderboard
§ Code
õ Dataset
arXiv:2604.12911v1 [cs.CL] 14 Apr 2026
Abstract Multilingual benchmarks guide the development of frontier models. Yet multilingual evaluations reported by frontier models are structured similar to popular reasoning and knowledge benchmarks, but across many languages. We show such benchmarks, and consequently multilingual evaluations, measure mathematical reasoning and factual recall, not multilingual proficiency. For example, thinking variants dramatically outperform instruct variants on these benchmarks, yet often perform worse on real-world multilingual tasks, such as LMArena. We propose a simple alternative: evaluate multilingual capability via round-trip translation. Given text in a source language, translate it to a target language and back; semantic gaps between the original and result expose failures in multilingual generation capabilities. Round-trip translation correlates almost perfectly (ρ = 0.94) with user ratings on LMArena with our benchmark, requires no human reference translations, and does not require a more capable multilingual judge than tested models. Lastly, we introduce Lost in Translation (LiT), a challenging round-trip translation benchmark spanning widely spoken languages worldwide, for realistic evaluation of multilingual frontier models.
1
Introduction
Multilingual benchmarks (Son et al., 2025; Wang et al., 2025; Romanou et al., 2025; Singh et al., 2025) shape how frontier models are built. Developers use these evaluations to measure progress, allocate resources, and claim capabilities across languages. Yet a fundamental question remains open: Does progress on current multilingual benchmarks truly reflect progress in multilingual proficiency? We find it does not. The evaluation of frontier multilingual models is currently dominated by two predominant paradigms: mathematical reasoning tasks, such as MT-AIME24 (Son et al., 2025) and PolyMath (Wang et al., 2025), and general knowledge multiple-choice question answering (MCQA), such as INCLUDE (Romanou et al., 2025) and Global-MMLU (Singh et al., 2025). We discover, as shown in Figure 1 and 2, that performance gaps on such benchmarks primarily reflect differences in mathematical problem-solving (ρ = 0.94) or factual recall (ρ = 0.83) respectively – not multilingual generation capability (ρ = −0.09 and −0.26 respectively). Intuitively, the original AIME24 and MMLU, similarly, are poor benchmarks for faithful measurement of English comprehension. Consequently, performance gaps between two models on frontier multilingual benchmarks like MT-AIME24 and INCLUDE highlights differences in their mathematical reasoning and factual recall, not their language proficiency. Previous works (Wu et al., 2025) provide support on the failure of multilingual benchmarks to align with human preferences. Overall, we show that frontier multilingual benchmarks are not a faithful measurement of multilingual capabilities. We propose a simple solution: evaluate multilingual capability through round-trip translation (Brislin, 1970; Sennrich et al., 2016). Given a passage in a source language, round-trip translation involves translating it to a target (sequence of) language(s) and back – comparing the result to the original passage. Semantic gaps between the original and back-translated text show whether a model can preserve meaning across languages, a critical capability multilingual benchmarks should capture (Wu et al., 2025). Unlike MT-AIME or Global1
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Table 1: Benchmark comparison across six evaluation criteria. We compare nine multilingual benchmarks across six evaluation criteria including: contamination-Free, challenging (not saturated yet, difficult for frontier models), linguistic diversity (covers diverse language families and scripts), NLG (tests natural language generation), Efficiency (low computational cost), and ground truth-free (without per-sample human annotation). Benchmark
Challenging
Linguistic Diversity
NLG
Efficiency
Ground Truth-Free
General Knowledge Multiple-Choice Include ✓ MMMLU ✗ ✗ Global MMLU
Contamination-Free
✓ ✗ ✗
✓ ✗ ✗
✗ ✗ ✗
✗ ✗ ✗
✗ ✗ ✗
Mathematical Reasoning MT-AIME24 MGSM PolyMath
✗ ✗ ✓
✓ ✗ ✓
✗ ✗ ✗
✓ ✓ ✓
✗ ✓ ✗
✗ ✗ ✗
Machine Translation FLORES-200 WMT24++
✗ ✓
✗ ✗
✓ ✓
✓ ✓
✓ ✗
✗ ✗
LiT (Ours)
✓
✓
✓
✓
✓
✓
MMLU, multilingual understanding and generation is the challenging task in a round-trip translation benchmark. We describe its comparative advantages in Table 1 when compared to current machine translation and current multilingual benchmarks. Round-trip translation has two advantages over classical translation evaluation: scalable sample creation – it requires no human-written reference translations; and ease of judging – the judge needs to evaluate two English paragraphs, not compare them in the low-resource language. This implies we do not need a stronger model to evaluate multilingual translation quality, which does not exist when testing frontier models. Current translation benchmarks are too easy for evaluating frontier models. Assuming frontier models keep improving, we argue these advantages – ease of creating new samples and judging – should increasingly favor round-trip translation over classical machine translation. To enable systematic evaluation, we introduce Lost in Translation (LiT), a round-trip translation benchmark of 1600 samples across 8 language sequences spanning high-, medium-, and low-resource languages. LiT covers technical (Taguchi et al., 2025), pragmatic (Park et al., 2024), and informal language (Yao et al., 2024) using MQM-based automated judging. We show that round-trip translation using LiT offers two advantages over current benchmarks. First, we show that round-trip translation correlates almost perfectly (ρ = 0.94) with user ratings on LMArena (Chiang et al., 2024)1 . Second, it exposes failure modes that reasoning and multiple-choice benchmarks miss. We document systematic hallucinations, content omissions, and semantic drift in multilingual generation using MQM scores (Lommel et al., 2014; Kocmi & Federmann, 2023), not faithfully captured by current multiple-choice (MCQ) or mathematical reasoning evaluations. Our results reveal a stark capability gap: open-source frontier models achieve above 88 MQM scores on high-resource sequences but collapse to below 50 on low-resource languages, indicating unusable output2 . Furthermore, we test model performance on challenging constructions for round-trip translation (Somers, 2005; Zhuo et al., 2023). Model rankings on these challenging cases match general performance closely, validating the robustness of the LiT benchmark. When instructed to translate faithfully, frontier models prioritize instructionfollowing over fluency correction. Rather than self-correcting or masking intermediate errors, they reliably reproduce the awkward phrasing. Consequently, this demonstrates that round-trip evaluation is not artificially derailed by complex linguistic phenomena, but instead serves as a stable, highly correlated proxy for genuine cross-lingual generation. Overall, we hope this work shifts how multilingual frontier models are evaluated. We argue for methods that directly measure user preference (Wu et al., 2025): genuine cross1 LMArena is considered an expensive but gold standard set, as it collects vast array of continuously
updated real-world user queries and uses a wisdom of the crowd effect (see Ni et al. (2024) for details). 2 Translations achieving an MQM score above 80 are generally considered faithful and fit-forpurpose, while a score of <80 should trigger mandatory human review (Lommel et al., 2014; 2024).
2
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss Correlation: MT-AIME24 vs LM Arena Spearman Correlation: -0.09
Qwen 3 235B (Thinking)
Qwen 3 235B (Instruct)
80 75
DeepSeek v3.2 Exp (No Reas)
70 65
1390
1400
81
1410
1420
LM Arena Elo
1430
1440
(a) MT-AIME24 vs. LMArena
98.5
DeepSeek v3.2 Exp (Thinking) Qwen 3 235B (Thinking) Kimi K2 (Thinking)
80
DeepSeek v3.2 Exp (No Reas)
79 Qwen 3 235B (Instruct)
78
Kimi K2
60
Include Score (%)
MT-AIME24 Score (%)
85
82
Kimi K2 (Thinking) DeepSeek v3.2 Exp (Thinking)
90
Correlation: Backtranslation vs LM Arena Spearman Correlation: 0.94 **
1400
DeepSeek v3.2 Exp (No Reas)
98.0 Kimi K2 (Thinking) 97.5
97.0
Qwen 3 235B (Instruct)
DeepSeek v3.2 Exp (Thinking) Qwen 3 235B (Thinking)
96.5
Kimi K2 1390
Backtranslation Score (%)
95
Correlation: Include vs LM Arena Spearman Correlation: -0.26
Kimi K2
1410
1420
LM Arena Elo
1430
(b) Include vs. LMArena
1440
96.0 1390
1400
1410
LM Arena Elo
1420
1430
1440
(c) LiT vs. LMArena
Figure 1: Multilingual benchmarks correlate poorly with human preferences. We show benchmark scores against LMArena Elo ratings for six frontier open-source models, each in Thinking and Non-Thinking variants. (a) MT-AIME24 (Son et al., 2025) shows nearzero correlation (ρ = −0.09), with thinking variants dramatically outperforming on the benchmark but perform no better on LMArena; (b) INCLUDE (Romanou et al., 2025) shows moderate negative correlation (ρ = −0.26), with similar Thinking vs. Non-Thinking disconnect; (c) Round trip translation using LiT with percentage of MQM scores at least 80 correlates almost perfectly with LMArena (ρ = 0.94, with one-sided permutation test p=0.008), with minimal gap between thinking and instruct variants of a model. lingual generation competence. We hope the LiT benchmark contributes to guiding the development of models that actually work for the billions of people who do not speak English as their first language.
2
Related Works
Multilingual Benchmarks. The standard approach to multilingual LLM evaluation translates English reasoning and knowledge tasks into other languages (Hendrycks et al., 2021; Son et al., 2025). However, such translated benchmarks often correlate poorly with human evaluations, performing significantly worse than natively localized benchmarks Wu et al. (2025). Alternatives like INCLUDE Romanou et al. (2025) and Global-MMLU Singh et al. (2025) address this by incorporating native content or cultural context, mitigating both “translationese” artifacts and Western-centric bias Wu et al. (2025). However, these localized benchmarks inherit a deeper structural flaw: they evaluate knowledge retrieval via multiplechoice formats rather than actual natural language generation, suffering from selection bias (Zheng et al., 2024; Balepur et al., 2025). Furthermore, natively sourced benchmarks like INCLUDE (Romanou et al., 2025) inherently provide imbalanced cross-lingual comparisons. Evaluating German on the Driver’s License topic while testing French on other topics confounds language proficiency with task difficulty. Ultimately, real-world utility depends on coherent, fluent, and semantically faithful generation. We provide empirical evidence demonstrating how current benchmarks fail to capture this, motivating our standardized generative approach. Natural Language Generation Foundational benchmarks such as MMLU (Hendrycks et al., 2021; Sai et al., 2022) acknowledge that the future of model evaluation lies in Natural Language Generation (NLG), but use multiple-choice formats because NLG remains difficult to evaluate. Consequently, assessing multilingual generation quality remains an open problem. Traditional machine translation benchmarks like WMT (Kocmi et al., 2025) and FLORES (Team et al., 2022) directly evaluate translation quality, but they depend on reference translations or per-example human ratings. These benchmarks also draw from Wikipedia and other widely-used sources in training datasets, raising concerns about data contamination into model training (Sainz et al., 2024; Karpinska & Iyyer, 2023; Vilar et al., 2023). Furthermore, most evaluations are sentence-level, lacking the complexity of real-world translation tasks. Recent NLG benchmarks such as PolyMath (Wang et al., 2025), MGSM (Shi et al., 2023), and MT-AIME24 (Son et al., 2025) attempt to test multilingual generation by machine-translating 3
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
English-sourced mathematical questions. However, as shown in Fig 2, these approaches primarily evaluate mathematical proficiency rather than multilingual capability. To address these gaps, we curate a benchmark of complex, multi-sentence passages reflecting real-world use cases. We adapt the MQM framework (Lommel et al., 2014), used in WMT Shared Tasks (Freitag et al., 2021; Lavie et al., 2025) and replace human translators with LLM judges for scalability (Lu et al., 2024; Kocmi & Federmann, 2023; Kim, 2025). Separately, LMArena Chiang et al. (2024) provides a large-scale measure of real-world user preferences. This approach has become the standard for tracking real-world model performance and has inspired work on efficient proxies that maintain high correlation with full arena rankings (Li et al., 2025; Dubois et al., 2024; Spangher et al., 2025). We use LMArena as a point of comparison for real-world user preferences. Round-Trip Translation. Round-trip translation has a long history as an evaluation tool. The foundational Brislin (1970) paper established its use for validating translation quality. It later became a widely used data augmentation approach in neural machine translation Sennrich et al. (2016). The core insight is straightforward: if meaning degrades through a round-trip (English → target → English), the forward translation is usually unreliable. Early statistical models undermined this approach by simply utilizing a ”copy mechanism” (Somers, 2005). However, modern neural machine translation (NMT) architectures lack this flaw, confirming that round-trip translation is an effective method for reference-free evaluations (Moon et al., 2020; Zhuo et al., 2023). Where prior work mainly uses round-trip translation for MT data augmentation (Sennrich et al., 2016), we use it to evaluate frontier LLMs (Allamanis et al., 2024).
3
Current Benchmarks Do Not Measure Multilingual Capability
Frontier model reports (Yang et al., 2025) evaluate multilingual capabilities using popular multilingual reasoning benchmarks like MT-AIME24 (Son et al., 2025), MGSM (Shi et al., 2023) and PolyMath (Wang et al., 2025), as well as general knowledge MCQA benchmarks such as MMMLU (Hendrycks et al., 2021) and INCLUDE (Romanou et al., 2025). In this section, we show that these benchmarks correlate poorly with human preferences because they measure reasoning or factual recall, not multilingual comprehension. 3.1
Benchmark Scores Diverge from User Preferences
Higher scores on a good benchmark should imply better real-world utility. We test whether multilingual benchmarks satisfy this criterion by correlating their scores with user experience in-the-wild. Setup. We evaluate six frontier open-source models: Kimi K23 (Team et al., 2025), DeepSeekV3.2-Exp (Liu et al., 2025), and Qwen3-235B-2507 (Yang et al., 2025), each in Thinking and Non-Thinking variants. We constrain the comparison to same-tier models to avoid spurious correlations driven by model size (Kaplan et al., 2020; Ghorbani et al., 2022). We measure zero-shot accuracy on MT-AIME24 (Son et al., 2025), a multilingual reasoning benchmark, and INCLUDE (Romanou et al., 2025), a culturally-curated, knowledge-intensive benchmark. We then compute Spearman rank-order correlations between benchmark scores and LMArena (Chiang et al., 2024) Elo ratings. For comparison, we also evaluate backtranslation quality on LiT (English → Language → English) across seven LMArena languages4 and report MQM≥80 : the percentage of backtranslations whose MQM score is at least 80. Results. Figure 1 shows that benchmark rankings diverge from human preferences. Both MT-AIME24 (left) and INCLUDE (middle) exhibit slight to moderate negative correlation with LMArena Elo scores. The disconnect is strongest between Thinking and Non-Thinking variants of the same model: Thinking models dramatically outperform their counterparts on MT-AIME24 and INCLUDE, yet on LMArena – where users rate actual multilingual outputs 3 Unlike other models, Kimi K2 Thinking was released after the Non-Thinking variant. 4 Chinese, French, German, Japanese, Korean, Russian, Spanish
4
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Correlation: MT-AIME24 vs AIME25 Spearman Correlation: 0.94 **
85 Qwen 3 235B (Instruct)
80 75 70 65 60 55
Kimi K2 (Thinking) DeepSeek v3.2 Exp (Thinking) Qwen 3 235B (Thinking)
DeepSeek v3.2 Exp (No Reas) Kimi K2 (Base)
Include AIME24
Kimi K2 (Thinking) DeepSeek v3.2 Exp (Thinking) Qwen 3 235B (Thinking)
85.0 84.5 84.0 83.5
DeepSeek v3.2 Exp
83.0 82.5
Qwen 3 235B Kimi K2
82.0 60
70
80
AIME25 Score (%)
90
(a) MT-AIME24 vs AIME25
79
80
81
82
83
MMLU-Pro Score (%)
84
GLM 4.7 GPT-OSS 120B GPT-OSS 20B
Mimo V2 Flash Qwen 3 235B Qwen 3 32B
40
80
100
Answer in Same Language (%)
MT-AIME24 Score (%)
90
Correlation: Include vs MMLU-Pro Spearman Correlation: 0.83 * 85.5
Include Score (%)
95
85
(b) INCLUDE vs MMLU-Pro
90 80 70 60 50
0
20
60
Reasoning in Same Language (%)
100
(c) Probing reasoning language
Figure 2: Multilingual benchmarks track English reasoning performance, and models reason in English regardless of input language. (a) MT-AIME24 scores correlate strongly with English AIME25 performance (ρ = 0.94, one-sided permutation test p=0.008), indicating the benchmark primarily measures mathematical reasoning ability. (b) INCLUDE scores correlate strongly with English MMLU-Pro (ρ = 0.83, one-sided permutation test p=0.029), indicating the benchmark primarily measures factual knowledge. (c) Models overwhelmingly default to English for internal reasoning (y-axis) even when answering in the target language (x-axis). This rules out the possibility that benchmark errors reflect multilingual reasoning failures. – they often perform no better. In contrast, round-trip translation scores on LiT correlate almost perfectly (ρ = 0.94) with LMArena, suggesting closer alignment with real-world multilingual performance. Analysis. Why do multilingual benchmarks fail to predict real-world performance? We observe that they conflate two distinct capabilities: reasoning/knowledge and multilingual understanding. Gains on the benchmark may stem from improved reasoning, not improved language proficiency. We correlate each multilingual benchmark with an English benchmark measuring the same underlying skill: MT-AIME24 with AIME25, and INCLUDE with MMLU-Pro (Wang et al., 2024). Figures 2a, 2b show near perfect correlations across both benchmark pairs. This suggests that gains on multilingual benchmarks may largely reflect gains on English reasoning and knowledge tasks. Current benchmarks may track reasoning and knowledge gains more than multilingual ones, creating a misleading impression of multilingual progress. Billions who speak languages other than English might see limited benefit from progress on these benchmarks. Benchmark Scores Diverge from Real-World Use Multilingual benchmarks conflate two distinct capabilities: reasoning/facts and multilingual understanding, i.e. progress on (MT-)AIME24 primarily stems from better mathematical capability, not improved English (language) proficiency. We need challenging language-centric multilingual benchmarks.
3.2
Errors in Current Benchmarks are not Linguistic
We next analyze error types in two common benchmark categories, multilingual reasoning and general-knowledge multiple choice, to test whether these benchmarks reflect reasoning or factual knowledge rather than language proficiency. MT-AIME24 errors are logical, not semantic. We analyze errors for Qwen3-235B-A22BThinking-2507 on MT-AIME24 across 11 languages; with similar trends reproducible across a variety of models (detailed in Appendix C). Figure 3 shows the distribution. We observe that most errors stem from logical mistakes, calculation errors or flawed reasoning steps, and not from poor (multilingual) question comprehension. For French, Russian, and Thai, 100% of errors reflect failed reasoning, not failed understanding. The general trend is that the model parsed the translated problem correctly; it simply could not solve it. Performance gains on MT-AIME24 reflect better mathematical reasoning, not multilingual capability. 5
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
arithmetic logic
N=7 N=3 N=4 N=2 N=5 N=4 N=5 N=4 N=9 N=3 N=5
80 60 40 20 0
100
Error distribution (%)
Error distribution (%)
100
semantic
bn
de
en
es
fr
ja
ru
Language
sw
te
th
(a) Error distribution for MT-AIME24.
regional_knowledge semantic
N=113 N=41
N=91
N=65
N=31 N=123 N=153 N=57
bn
es
fr
ja
80 60 40 20 0
zh
factual hallucination logic
de
Language
ru
te
zh
(b) Error distribution for INCLUDE.
Figure 3: Errors on multilingual benchmarks stem from reasoning and knowledge gaps—not language comprehension failures. We manually categorize all errors made by Qwen3-235BA22B-Thinking across 11 languages. (a) On MT-AIME24, errors are overwhelmingly logical (arithmetic mistakes, flawed reasoning steps) rather than semantic (misunderstanding the translated question). (b) On INCLUDE, errors are predominantly factual (wrong facts, knowledge gaps) rather than semantic (misunderstanding the translated question). Both patterns confirm that these benchmarks measure reasoning and factual recall – not multilingual proficiency.
INCLUDE errors are factual, not semantic. One might argue that general knowledge benchmarks such as INCLUDE avoids this problem because it uses native questions rooted in regional knowledge rather than simply translating text from MMLU. However, the same pattern persists, including for other models in Appendix C. Figure 3 shows that most errors stem from incorrect factual knowledge. For Bengali and Chinese, over 96% of errors reflect knowledge gaps, not misunderstanding of the input question or MCQ options. Performance gains on INCLUDE, therefore, similarly reflect better elicitation and coverage of facts, not better multilingual capability. Reasoning failures occur in English, not the target language. One could argue that reasoning in the target language is itself worth assessing. We test whether models actually reason in the target language by analyzing traces from multiple language models across both benchmarks. Figure 2c shows they do not: models almost always default to English for their reasoning traces. The knowledge gaps and logical errors we identify therefore occur in English, not even in the target language. This is further evidence suggesting that reasoning failures do not reflect multilingual failures – they are reasoning/factual mistakes made in the English language (for extensive results, please refer to the Appendix sections on answer-language and error-distribution analyses). Summary. Overall, both benchmark types fail to disentangle multilingual capability from orthogonal skills. A model with strong language proficiency, but weak mathematical reasoning (like DeepSeek-V3.2-Exp Non-Thinking) (as seen in Figure 1) scores poorly on MTAIME24. Equivalently, a model with limited factual coverage underperforms on INCLUDE. Because models across newer generations often broadly improve both multilingual and reasoning capabilities at the same time, benchmark scores can appear to track multilingual progress even when they mainly reflect improvement in other areas.
Errors Don’t Reflect Multilingual Failures Errors in current multilingual benchmarks primarily reflect reasoning or knowledge gaps, and these gaps persist in English. Multilingual benchmarks should analyze the errors to understand the gaps in multilingual capabilities.
6
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
4
The Lost In Translation (LiT) Benchmark
Following our previous analysis of multilingual benchmark shortcomings, we propose Lost In Translation (LiT) – a natural language generation benchmark centered on backtranslation. LiT addresses the limitations of existing benchmarks summarized in Table 1: limited linguistic diversity, reliance on surface-level metrics, and failure to probe multilingual generation capabilities. 4.1
Dataset: Description and Setup
Data Curation. To prevent data contamination from existing corpora, we create a new benchmark of 1600 samples (200 unique texts x 8 language sequences). We manually select samples with diverse linguistic complexities. The dataset spans three major categories, each targeting distinct aspects of translation difficulty, with further details in Appendix E: (a) Abstracts (20%): Split equally between Humanities and STEM. Abstracts are selfcontained and semantically dense, making errors easy to detect. Humanities tests argumentative nuance; STEM tests terminological precision and notation. (b) Pragmatics (60%): Tests meaning beyond literal content across five phenomena: (i) Core Semantics (20%) i.e. truth conditions and entailment; (ii) Discourse Coherence (17.5%) i.e. referential chains and logical connectives; (iii) Implicit Content (17.5%) i.e. presuppositions and implicatures; (iv) Pragmatic Inference (21.7%) i.e. speech acts and speaker intent; (v) Social Interaction (23.3%) i.e. politeness, formality, and sociolinguistic appropriateness. (c) Informal (20%): Tests colloquial language, slang, idioms, and register shifts; preserving tone and social function beyond denotative meaning. Language Sequences. To rigorously stress-test multilingual capabilities, we evaluate performance across 8 linguistic sequences. We select these sequences to span disjoint language families, distinct scripts (Latin, Cyrillic, Devanagari, Arabic, CJK), and varying levels of pretraining resource availability, We prioritize translation pairs between regions with frequent interactions to reflect real-world translation tasks. Each sequence comprises four languages through which the source text is serially translated before round-trip translation to English: • East Asia (E.Asia, High resource): Japanese → Korean → Chinese → Russian. • Central Europe (C.Europe, High resource): Romanian → Hungarian → German → Polish. • Near East (N.East, High resource): Bulgarian → Greek → Turkish → Persian. • North Europe (N.Europe, Medium resource): Icelandic → Swedish → Finnish → Lithuanian. • Southeast Asia (SE.Asia, Medium resource): Chinese → Thai → Vietnamese → Tagalog. • South Asia (S.Asia, Medium resource): Hindi → Tamil → Bengali → Punjabi. • Africa (Africa, Low resource): Swahili → Arabic → Hausa → Amharic. • South America (S.America, Low resource): Portuguese → Quechua → Spanish → Guarani. The sequences are also grouped into sequences containing exclusively well-resourced languages or including at least one medium- or low-resource language. Evaluation: LLM-as-a-Judge. We employ Grok 4.1 Fast as our primary judge model, being an inexpensive, fast but simultaneously very capable model – consistently one of the highest ranked in multilingual LMArena (Chiang et al., 2024). Following standard practice Lommel et al. (2014); Kim (2025), we adopt the MQM framework, a weighted penalty system that categorizes errors by severity: (i) Minor (−1): Slight awkwardness or non-critical fluency issues that do not affect comprehension; (ii) Major (−5): Significant semantic shifts, structural failures, or mistranslations that alter meaning; (iii) Critical (−25): Complete loss of meaning, hallucinations, or safety-relevant errors. The framework grounds evaluation in interpretable error categories, where low MQM scores indicate severe failure modes dominated by critical errors, with 80 widely considered to be a pass threshold (Farinha et al., 2022; Lommel et al., 2024). We additionally show in Appendix B that both the scores and rankings nearly perfectly correlate when using two alternatives: (i) raw MQM scores and (ii) LLM judge providing a score from 0-100, demonstrating robustness to the metrics. The Appendix results additionally show robustness to the judge model, allowing us to fix the metric to MQM≥80 and Grok 4.1 Fast judge model for the rest of the section. 7
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Table 2: LiT benchmark by linguistic category under MQM≥80 . We report the percentage of translations with MQM ≥ 80, aggregated within each category over eight translation sequences. Each cell shows the mean and the bootstrap-estimated standard error, computed via sentence-level resampling after averaging across sequences. Higher is better. Informal text is the hardest category overall (30.8 avg), while Core Semantics is the easiest (61.5). Model
Average
Abstracts
Pragmatics
Informal
Humanities STEM Core Semantics Discourse Implicit Inference Social Gemini-3-Flash (No-Thinking)
87.2
90.0
85.6
87.5
89.3
86.9
88.5
89.7
80.3
± 0.7
± 1.9
± 1.8
± 1.9
± 1.5
± 2.3
± 1.6
± 1.6
± 2.2
73.8
70.0
88.1
81.2
75.6
76.2
78.4
75.9
45.0
Qwen3.5-397B (Thinking)
± 1.0
± 3.7
± 2.7
± 2.5
± 3.2
± 3.1
± 2.0
± 2.5
± 3.5
73.0
71.2
76.2
78.1
75.0
69.0
79.8
77.7
56.9
Gemma-4-31B (Instruct)
± 1.0
± 3.4
± 4.4
± 2.3
± 2.8
± 3.1
± 1.7
± 1.6
± 2.7
71.9
70.6
76.2
75.5
74.4
70.8
74.0
77.7
55.9
Gemma-4-31B (Thinking)
± 1.0
± 3.2
± 3.4
± 2.6
± 2.5
± 2.9
± 2.5
± 1.8
± 2.9
70.3
66.2
74.4
77.1
77.4
71.4
77.4
72.8
45.9
GLM-5 (Thinking)
± 1.1
± 4.0
± 2.9
± 3.1
± 1.4
± 3.3
± 1.9
± 3.0
± 3.2
67.6
71.2
69.4
71.9
69.6
66.7
74.0
71.4
46.9
Qwen3.5-397B (Instruct)
± 1.2
± 4.0
± 4.4
± 3.8
± 2.2
± 4.0
± 1.9
± 2.4
± 3.2
64.7
63.1
58.1
70.6
70.8
67.6
71.6
68.3
47.2
Kimi-K2 (Thinking)
± 0.9
± 3.2
± 4.0
± 1.3
± 1.4
± 2.6
± 1.5
± 1.9
± 2.4
61.6
63.8
60.6
72.4
70.2
61.3
67.8
65.2
31.6
GLM-4.7 (Thinking)
± 1.0
± 3.2
± 3.5
± 2.5
± 2.0
± 3.6
± 1.9
± 2.3
± 2.9
55.6
49.4
51.9
67.2
66.7
60.1
67.3
59.8
22.8
Qwen3-235B (Thinking)
± 1.0
± 3.7
± 3.8
± 1.9
± 2.1
± 3.7
± 1.8
± 2.8
± 2.5
53.4
52.5
42.5
58.9
57.1
58.3
63.9
57.6
36.6
DeepSeek-V3.2-Exp
± 1.1
± 3.6
± 3.9
± 2.9
± 3.4
± 2.8
± 2.7
± 2.6
± 3.1
53.3
51.2
45.0
60.9
60.7
54.8
59.1
61.2
33.1
GLM-5 (Instruct)
± 1.2
± 4.2
± 4.2
± 2.2
± 3.6
± 3.7
± 2.7
± 2.6
± 3.2
DeepSeek-V3.2-Exp (Thinking)
53.2
55.0
36.9
65.6
61.9
60.1
58.7
59.8
27.8
± 1.3
± 5.5
± 5.1
± 3.1
± 2.9
± 3.4
± 3.5
± 2.3
± 2.8
52.0
48.1
52.5
63.5
60.1
54.2
63.5
58.5
15.9
Qwen3.5-35B (Thinking)
± 1.2
± 4.4
± 3.7
± 2.5
± 3.4
± 4.6
± 2.5
± 2.5
± 2.8
49.7
47.2
43.3
60.9
55.6
49.7
57.4
51.3
32.1
Kimi-K2
± 1.3
± 3.8
± 5.3
± 2.9
± 4.0
± 4.0
± 2.3
± 3.1
± 3.3
47.8
43.8
28.1
60.4
55.4
49.4
60.1
52.2
32.8
Gemma-3-27B (Instruct)
± 1.3
± 4.6
± 5.3
± 2.9
± 3.4
± 3.7
± 1.9
± 2.8
± 3.4
44.1
39.4
35.0
53.6
52.4
48.2
49.5
49.6
25.3
Qwen3-235B (Instruct)
± 1.2
± 3.8
± 4.4
± 2.6
± 3.2
± 3.3
± 2.3
± 3.4
± 3.2
41.9
35.8
41.2
55.7
54.8
39.3
56.2
44.2
7.8
GPT-OSS-120B (High)
± 1.3
± 4.4
± 5.5
± 2.5
± 2.5
± 4.6
± 2.7
± 3.3
± 2.1
39.9
30.6
36.2
57.8
51.8
42.9
49.0
43.3
7.2
MiniMax-M2.5
± 1.2
± 3.2
± 4.3
± 3.4
± 3.6
± 4.9
± 3.0
± 2.9
± 1.8
38.2
32.5
36.9
49.5
41.7
40.5
51.9
40.2
12.8
Qwen3.5-35B (Instruct)
± 1.3
± 4.3
± 4.1
± 3.3
± 3.4
± 5.0
± 2.7
± 3.1
± 2.2
30.2
25.0
14.4
44.3
41.7
34.5
40.9
30.8
10.3
MiMo-V2-Flash
± 1.1
± 3.8
± 3.6
± 3.2
± 3.3
± 3.5
± 3.2
± 2.8
± 2.3
21.4
14.4
19.4
30.7
28.0
23.2
25.5
26.8
3.4
Qwen3-30B (Instruct)
± 1.0
± 2.5
± 4.1
± 3.0
± 3.0
± 3.4
± 2.3
± 2.4
± 1.1
5.5
5.6
1.2
10.4
7.7
3.0
10.6
5.8
0.0
Nemotron-3-Nano
± 0.5
± 1.9
± 0.8
± 1.8
± 1.6
± 1.4
± 1.9
± 1.2
± 0.0
Average
52.6
49.8
48.8
61.5
59.0
54.0
60.2
56.4
30.8
4.2
Main Results
A Linguistic Category View. We present our results in Table 2. The trends from the previous section also appear across LiT categories: reasoning effects are model-dependent, helping Qwen3-235B-A22B-2507 much more than DeepSeek-V3.2-Exp, while consistently hurting informal translation. Gemini-3-Flash achieves the highest overall score and the highest score in each subcategory. The high gap between Gemini and other models indicates that our benchmark can measure performance across the ”long tail” of scenarios efficiently. Category-wise breakdown. To understand failure modes, Table 2 decomposes performance by linguistic category. STEM Abstracts and Informal registers emerge as the most challenging 8
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
categories, with average performance of 48.8 and 30.8 respectively. Each entry reports the mean and a bootstrap-estimated standard error, computed from N = 10, 000 sentence-level resamples after averaging across sequences. Technical content tests whether models can communicate domain knowledge fluently. Informal text tests whether models handle slang and colloquialisms in translation. We observe a striking pattern: reasoning models underperform on informal text. One possible explanation is that reasoning models overcomplicate simple colloquial inputs, producing a mismatch in register. A Language-Centric View. Table 3 presents our results grouped by sequence and by high/medium/low resource classification. Most notably, performance drops sharply from high-resource to low-resource settings. Medium-resource language sequences largely follow high-resource trends. High-Resource Stability. In East Asian, Near Eastern, and Central European sequences, most frontier models achieve strong performance. As observed in Table 3, Gemini-3-Flash leads with an MQM≥80 score of 97.0, followed by GLM-5 (Thinking) at 89.8. Most models perform well on these sequences. Low-Resource Collapse. In contrast, performance collapses in African and South American sequences. Except Gemini-3-Flash which maintains a score of 59.2, Table 3 shows that the MQM≥80 score of the next best model – Gemma-4-31B (Instruct) – drops to merely 32.0. Several models, including Qwen3-30B (0.0) and Qwen3-235B (0.2), receive near-zero scores, signalling critical errors and unusable translations. This confirms the official result that Qwen3-235B does not officially support some low-resource languages in this subset, such as Quechua or Amharic (Qwen Team, 2025). Impact of Reasoning on Translation. Comparing thinking variants with instruct counterparts in Table 3 suggests that inference-time reasoning helps most in low-resource languages. The thinking variant of DeepSeek-V3.2-Exp scores 7.2 in low-resource settings, outperforming the instruct model’s MQM≥80 score of 1.0. These results suggest that reasoning cannot compensate for missing lexical knowledge, but may mitigate some severe errors in unfamiliar linguistic settings. More broadly, however, reasoning does not reliably improve translation quality: in high-resource settings it is often only comparable to instruct models, and even recently released models such as Gemma-4-31B show the instruct variant matching or outperforming the thinking variant overall, with the clearest advantage in low-resource settings. Key Findings • Reasoning often hurts translation quality, especially on informal text. • A catastrophic accuracy cliff separates high-resource from low-resource languages. • Reasoning models outperform instruct models primarily in low-resource languages 4.3
Round-Trip Translation Robustness
While LiT uses round-trip translation, the method has several known limitations (Somers, 2005; van Zaanen & Zwarts, 2006). A correct round-trip translation does not necessarily imply correct intermediate translations, which may contain awkward phrasing, inappropriate formality, register mismatches, or missing cultural nuance that disappear in the final back-translation. To assess whether these classical concerns still affect current models, we evaluate a separate 480-sample robustness benchmark. Our results show that modern LLMs are largely robust to these failure modes, and that round-trip translation remains informative on such challenging cases. Additional details are provided in Appendix A. • Polysemy where words carry multiple meanings that only context can resolve. • Syntactic ambiguity where sentence structure remains unclear until the final word. • Idioms and cultural metaphors to test whether literal translation destroys meaning. • Register shifts including formal, colloquial & technical language within a single example. • Abstract nuance where physical vocabulary describes non-physical concepts.
9
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Table 3: MQM≥80 collapses on low-resource language sequences. We report MQM≥80 rates, the percentage of subset examples with MQM ≥ 80, across the same eight language sequences grouped by resource availability. (a) High-resource sequences remain solvable for the strongest frontier models, with Gemini-3-Flash averaging 97.0 and reaching 98.0 on both Central European and Near Eastern chains. (b) Medium-resource sequences remain relatively stable for the best models, though separation widens below the frontier tier. (c) Low-resource sequences collapse sharply: only Gemini-3-Flash sustains a substantial pass rate, averaging 59.2 across African and South American chains; the next-best model drops to 32.0. (d) Global averages across these eight sequences show Gemini-3-Flash far ahead of all competitors. (c) Low Resource / Imbalanced Sequences
(a) High & Med-High Resource Sequences Model
E. Asia C. Europe N. East Avg
Gemini-3-Flash (No-Think) GLM-5 (Thinking) Qwen3.5-397B (Thinking) Gemma-4-31B (Instruct) Gemma-4-31B (Thinking) GLM-4.7 (Thinking) Qwen3.5-397B (Instruct) Kimi-K2 (Thinking) DeepSeek-V3.2-Exp GLM-5 (Instruct) Qwen3-235B (Thinking) DeepSeek-V3.2-Exp (Think) Qwen3.5-35B (Thinking) Gemma-3-27B (Instruct) Kimi-K2 Qwen3-235B (Instruct) MiniMax-M2.5 Qwen3.5-35B (Instruct) GPT-OSS-120B (High) MiMo-V2-Flash Qwen3-30B (Instruct) Nemotron-3-Nano
95.0 93.5 91.0 90.5 93.5 92.0 88.5 84.0 87.5 87.5 82.5 84.5 76.0 70.0 71.5 83.0 60.0 71.5 68.5 57.0 56.5 29.5
98.0 93.5 92.0 90.5 87.5 90.0 91.5 89.0 85.0 86.0 86.0 75.0 83.0 79.5 78.0 62.5 75.5 73.5 65.5 61.5 49.0 4.0
98.0 82.5 85.5 86.0 86.0 80.5 82.0 79.0 71.5 69.5 71.5 69.5 65.5 68.5 56.5 57.5 58.5 46.5 49.0 30.5 21.5 0.0
97.0 89.8 89.5 89.0 89.0 87.5 87.3 84.0 81.3 81.0 80.0 76.3 74.8 72.7 68.7 67.7 64.7 63.8 61.0 49.7 42.3 11.2
Model
Africa S. America Avg
Gemini-3-Flash (No-Think) Gemma-4-31B (Instruct) Qwen3.5-397B (Thinking) Gemma-4-31B (Thinking) GLM-5 (Thinking) Qwen3.5-397B (Instruct) DeepSeek-V3.2-Exp (Think) GLM-4.7 (Thinking) Nemotron-3-Nano Kimi-K2 (Thinking) Qwen3.5-35B (Thinking) MiMo-V2-Flash GPT-OSS-120B (High) Qwen3.5-35B (Instruct) Qwen3-235B (Thinking) Gemma-3-27B (Instruct) DeepSeek-V3.2-Exp GLM-5 (Instruct) MiniMax-M2.5 Kimi-K2 Qwen3-235B (Instruct) Qwen3-30B (Instruct)
78.5 60.0 41.0 51.5 30.5 27.5 8.5 8.0 1.5 5.0 4.0 0.0 1.5 1.0 0.5 2.0 1.5 2.0 1.0 1.0 0.5 0.0
Model Gemini-3-Flash (No-Think) Qwen3.5-397B (Thinking) Gemma-4-31B (Instruct) Gemma-4-31B (Thinking) GLM-5 (Thinking) Qwen3.5-397B (Instruct) GLM-4.7 (Thinking) Kimi-K2 (Thinking) Qwen3-235B (Thinking) DeepSeek-V3.2-Exp GLM-5 (Instruct) DeepSeek-V3.2-Exp (Think) Qwen3.5-35B (Thinking) Gemma-3-27B (Instruct) Kimi-K2 Qwen3-235B (Instruct) GPT-OSS-120B (High) MiniMax-M2.5 Qwen3.5-35B (Instruct) MiMo-V2-Flash Qwen3-30B (Instruct) Nemotron-3-Nano
SE. Asia S. Asia N. Europe Avg 93.0 82.0 92.5 88.5 86.5 83.5 70.5 71.5 74.5 75.0 63.5 69.5 60.5 60.5 57.0 69.0 49.0 45.0 46.0 41.5 28.0 2.5
95.5 77.0 90.0 93.0 74.0 71.0 66.0 58.5 55.0 46.5 57.0 45.5 48.0 50.0 42.0 48.0 37.5 16.0 21.0 18.0 5.0 0.0
95.5 87.0 61.5 62.5 79.5 75.0 63.5 65.0 53.5 51.5 50.0 53.0 56.0 46.5 54.0 22.5 42.5 43.0 30.5 18.5 1.5 0.0
94.7 82.0 81.3 81.3 80.0 76.5 66.7 65.0 61.0 57.7 56.8 56.0 54.8 52.3 51.0 46.5 43.0 34.7 32.5 26.0 11.5 0.8
10
59.2 32.0 28.0 27.2 18.8 18.2 7.2 5.5 3.0 2.8 2.8 2.2 1.3 1.2 1.2 1.0 1.0 1.0 0.8 0.5 0.2 0.0
(d) Overall Performance Model
(b) Medium Resource Sequences
40.0 4.0 15.0 3.0 7.0 9.0 6.0 3.0 4.5 0.5 1.5 4.5 1.0 1.5 2.0 0.0 0.5 0.0 0.5 0.0 0.0 0.0
Gemini-3-Flash (No-Think) Gemma-4-31B (Instruct) Qwen3.5-397B (Thinking) Gemma-4-31B (Thinking) GLM-5 (Thinking) Qwen3.5-397B (Instruct) GLM-4.7 (Thinking) Kimi-K2 (Thinking) Qwen3-235B (Thinking) DeepSeek-V3.2-Exp GLM-5 (Instruct) DeepSeek-V3.2-Exp (Think) Qwen3.5-35B (Thinking) Gemma-3-27B (Instruct) Kimi-K2 Qwen3-235B (Instruct) GPT-OSS-120B (High) MiniMax-M2.5 Qwen3.5-35B (Instruct) MiMo-V2-Flash Qwen3-30B (Instruct) Nemotron-3-Nano
Global Average 86.7 71.9 71.3 70.7 68.4 66.0 59.2 56.6 53.2 52.4 51.9 51.4 49.3 47.1 45.0 42.9 39.3 37.4 36.4 28.9 20.2 5.2
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Table 4: MQM≥80 on the Backtranslation robustness subset, which targets classical roundtrip translation challenges: idioms, polysemy, register shifts, and garden-path syntax. We report the percentage of sentence-sequence translations with MQM ≥ 80, aggregated within each category over the same eight translation sequences. Each cell shows the mean and bootstrap-estimated standard error, computed via sentence-level resampling after averaging across sequences. Higher is better. Each category contains 12 source sentences. Idioms & Cultural Metaphors is the hardest category overall (35.3 avg), while Conceptual & Abstract Nuance is the easiest (54.8). Model
Average
Conceptual &
Idioms &
Polysemy &
Register & Syntactic Complexity
Abstract Nuance Cultural Metaphors Lexical Ambiguity Tone Shifts
& Garden Paths
Gemini-3-Flash (No-Thinking)
71.0
83.3
65.6
68.8
60.4
77.1
± 2.1
± 3.4
± 5.9
± 5.4
± 5.3
± 2.5
63.5
75.0
62.5
59.4
58.3
62.5
Qwen3.5-397B (Thinking)
± 2.0
± 3.6
± 5.9
± 4.2
± 3.4
± 4.2
59.8
74.0
54.2
50.0
57.3
63.5
GLM-5 (Thinking)
± 2.0
± 2.7
± 4.0
± 6.9
± 3.5
± 4.0
Gemma-4-31B (Thinking)
55.8
75.0
47.9
54.2
55.2
46.9
± 2.4
± 3.6
± 6.0
± 6.2
± 4.7
± 5.9
54.6
67.7
49.0
55.2
50.0
51.0
Kimi-K2 (Thinking)
± 1.8
± 2.3
± 4.8
± 3.6
± 5.1
± 3.3
Gemma-4-31B (Instruct)
52.5
70.8
44.8
49.0
53.1
44.8
± 2.2
± 3.7
± 4.0
± 5.1
± 6.1
± 5.2
51.9
70.8
38.5
56.2
52.1
41.7
Qwen3.5-397B (Instruct)
± 2.1
± 3.7
± 5.0
± 3.5
± 5.5
± 5.3
48.8
61.5
43.8
47.9
43.8
46.9
GLM-4.7
± 2.0
± 2.7
± 5.6
± 4.6
± 5.3
± 4.2
47.5
60.4
42.7
45.8
37.5
51.0
Qwen3-235B (Thinking)
± 2.1
± 3.2
± 5.6
± 5.9
± 5.1
± 3.4
43.5
57.3
40.6
38.5
43.8
37.5
Qwen3.5-35B (Thinking)
± 2.4
± 4.5
± 5.5
± 5.8
± 5.9
± 5.1
DeepSeek-V3.2-Exp (Thinking)
41.0
57.3
29.2
43.8
34.4
40.6
± 2.4
± 4.6
± 5.0
± 6.6
± 4.2
± 6.1
39.2
53.1
34.4
31.2
32.3
44.8
DeepSeek-V3.2-Exp
± 2.2
± 4.2
± 5.6
± 5.2
± 5.2
± 4.6
38.8
43.8
33.3
41.7
32.3
42.7
MiniMax-M2.5
± 2.5
± 4.8
± 5.4
± 5.9
± 6.8
± 4.3
Gemma-3-27B (Instruct)
37.5
57.3
31.2
34.4
37.5
27.1
± 1.9
± 3.1
± 4.3
± 5.2
± 4.4
± 4.2
36.9
50.0
26.0
34.4
30.2
43.8
GLM-5 (Instruct)
± 2.1
± 4.9
± 4.5
± 5.6
± 4.0
± 4.0
GPT-OSS-120B (High)
34.6
37.5
37.5
30.2
38.5
29.2
± 2.4
± 5.1
± 3.6
± 5.9
± 5.8
± 5.7
32.1
51.4
22.5
30.5
24.6
31.5
Kimi-K2
± 2.0
± 4.8
± 4.1
± 5.6
± 3.8
± 4.3
31.9
41.7
26.0
33.3
32.3
26.0
Qwen3-235B (Instruct)
± 2.3
± 5.2
± 5.6
± 5.6
± 5.2
± 4.1
Qwen3.5-35B (Instruct)
24.6
44.8
17.7
21.9
19.8
18.8
± 2.0
± 6.2
± 4.0
± 3.9
± 3.7
± 4.4
24.4
39.6
15.6
22.9
13.5
30.2
MiMo-V2-Flash
± 2.2
± 5.9
± 3.0
± 5.6
± 4.0
± 5.6
14.0
25.0
7.3
11.5
12.5
13.5
Qwen3-30B (Instruct)
± 1.3
± 2.9
± 2.8
± 2.7
± 2.5
± 4.0
5.2
8.3
5.2
1.0
5.2
6.2
Nemotron-3-Nano
± 0.8
± 1.7
± 1.8
± 1.0
± 2.7
± 1.8
Average
41.3
54.8
35.3
39.2
37.5
39.9
11
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Results. Table 4 reports category-wise MQM≥80 scores on the robustness subset. Idioms & Cultural Metaphors emerge as the most challenging category (35.3% on average), consistent with the difficulty of preserving figurative meaning across typologically diverse languages, whereas Conceptual & Abstract Nuance is the easiest (54.8%), suggesting that frontier models handle metaphorical extension more reliably than fixed idiomatic expressions. Although these items were specifically designed to stress classical failure modes of roundtrip translation, model rankings remain highly consistent with the main LiT benchmark: Table 5 in Appendix A shows that the Spearman correlation between each LiT category and the robustness benchmark average is significant at p < 0.001 for all categories, ranging from ρ = 0.78 for Informal to ρ = 0.96 for Core Semantics. As in the main benchmark, we report mean performance together with bootstrap-estimated standard errors. Manual inspection helps explain this robustness: when an intermediate translation is awkward or overly literal, stronger models often preserve that awkwardness rather than correct it, because they follow the instruction to translate faithfully. In this sense, instructionfollowing models can faithfully propagate errors instead of repairing them, suggesting that round-trip translation is less vulnerable to some classical critiques and can still probe whether models genuinely understand idiomatic or structurally complex constructions. Implications. These results suggest that round-trip translation, when applied to instructionfollowing LLMs, largely overcome the classical criticisms against it. The benchmark successfully surfaces genuine cross-lingual weaknesses (e.g., idiom handling, register preservation) rather than being artificially derailed by the phenomena it was designed to probe. The high rank-order consistency with the main benchmark (Table 5 in Appendix A) further validates that robustness cases do not introduce a separate, orthogonal axis of difficulty, but rather stress-test the same underlying multilingual generation capabilities measured by the main LiT benchmark.
5
Discussion and Future Work
We next discuss benchmark-design choices and considerations for future work. 5.1
Benchmark Design
While serial language sequences inherently compound translation errors, we deliberately designed this mechanism to strictly stress-test a model’s multilingual capabilities. A capable multilingual model must maintain semantic integrity across multiple translation hops; the catastrophic degradation observed outside high-resource sequences successfully isolates the boundary of this capability. By evaluating 200 highly dense, paragraph-length source texts (Läubli et al., 2018), LiT prioritizes linguistic depth over superficial breadth to prevent benchmark saturation (Bowman & Dahl, 2021). This provides an efficient evaluation setting (Wu et al., 2025) that probes linguistic phenomena often missed by sentence-level benchmarks. 5.2
Statistical Methodology and Evidence
We report rank correlations across six same-tier frontier configurations (n=6), constrained as described in Section 3.1 to avoid scale-driven confounding. We therefore interpret the correlations together with three additional pieces of evidence: (1) across all three model families, Thinking variants significantly outperform on existing benchmarks yet perform comparably or worse on LMArena, with effect sizes exceeding 20 percentage points on MT-AIME24; (2) our error taxonomy (Section 3.2) independently shows benchmark failures are logical and factual, not linguistic; and (3) the reasoning-language analysis (Figure 2c) confirms models default to English reasoning regardless of input language. We present ρ values for directional contrast, not as standalone hypothesis tests. Finally, LMArena remains the main large-scale source of multilingual human-preference data (Chiang et al., 2024), which constrains the set of models with reliable cross-lingual 12
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
ratings. Future work can strengthen the quantitative evidence as multilingual evaluation platforms expand. 5.3
Automated Evaluation and Validity
To keep LiT scalable and reproducible, we use an LLM-as-a-judge framework. While this removes direct human-in-the-loop verification, recent literature demonstrates the effectiveness of this approach (Kocmi & Federmann, 2023; Lu et al., 2024), strongly correlating with human preference. Crucially, we validate the quality of our automated judge by demonstrating its high rank-order correlation with real-world human preference ratings from LMArena (Chiang et al., 2024). This confirms that the semantic degradation penalized by our round-trip translation approach strictly aligns with the cross-lingual generation failures penalized by actual multilingual users. 5.4
Limitations and Future Directions
While LiT resolves critical limitations in existing multilingual benchmarks, it is merely the first step toward more comprehensive evaluation paradigms. We highlight three highimpact directions for future work which could address current limitations: Scaling to Document-Level Discourse. Frontier models are increasingly deployed for full-document tasks involving reports, literature, and legal text. While LiT effectively evaluates paragraph-level pragmatics, document-scale translation introduces complex global dependencies. Future benchmarks must evaluate the preservation of long-range referential chains, stylistic consistency, and persistent terminology memory across massive context windows. Disentangling Single Language Performance. While our serial language sequences bound overall cross-lingual robustness, they also obscure single-language performance. To isolate exact failure modes, future work should separate these sequences into independent, controlled evaluations. This would enable more granular capability profiles for individual low-resource languages without the relatively fuzzy intermediate error cascading. Efficient and Culturally Grounded Generation. An important direction is to develop efficient proxies that better reflect real-world human utility. While round-trip translation serves as an exceptional proxy for semantic preservation, future evaluation suites must expand beyond translation entirely. Developing highly efficient, native generation tasks that evaluate cultural grounding and target-language fluency, without relying on an English source or pivot, will be critical. Ultimately, combining round-trip evaluation with such native generative tasks will yield a comprehensive multilingual suite, moving the needle for the billions of users who rely on frontier models for global knowledge access and digitization.
6
Conclusion
In this work, we asked a simple question: Do current multilingual benchmarks faithfully measure multilingual capability? Benchmarks like MT-AIME24 and INCLUDE, used by frontier models to claim multilingual progress, actually might be confounded by improving reasoning and factual recall performance. Two findings support this conclusion. First, thinking variants dramatically outperform instruct variants on these benchmarks, yet perform no better (and often worse) on LMArena, where real users rate multilingual outputs. Second, our error analysis reveals that failures are logical and factual, not linguistic: models parse translated questions correctly but fail to solve them. We bring back round-trip translation as an alternative method. Unlike existing benchmarks, performance on our LiT benchmark correlates positively with user preferences on LMArena. It spans highresource to low-resource languages and shows underperformance of reasoning models in informal settings as well as substantial degradation on low-resource languages. We show that round-trip translation is a robust and scalable reference-free method for evaluating current multilingual capabilities of frontier models. We hope this work contributes to 13
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
multilingual benchmarks that better measure preservation of meaning, fluency, and cultural context across languages.
Acknowledgements RS acknowledges funding by the Federal Ministry of Research, Technology and Space (BMFTR), FKZ: 16IS24079A. AP and MB acknowledge financial support by Federal Ministry of Research, Technology and Space (BMFTR) FKZ: 16IS24085B and Open Philanthropy Foundation funded by the Good Ventures Foundation. MB is a member of the Machine Learning Cluster of Excellence, EXC number 2064/1 – Project number 390727645.
References Miltiadis Allamanis, Sheena Panthaplackel, and Pengcheng Yin. Unsupervised evaluation of code llms with round-trip correctness. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024. Nishant Balepur, Rachel Rudinger, and Jordan Lee Boyd-Graber. Which of these best describes multiple choice evaluation with LLMs? a) forced B) flawed C) fixable D) all of the above. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3394–3418, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/2025.acl-long. 169. URL https://aclanthology.org/2025.acl-long.169/. Regina Barzilay and Mirella Lapata. Modeling local coherence: an entity-based approach. In Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics, ACL ’05, pp. 141–148, USA, 2005. Association for Computational Linguistics. doi: 10.3115/ 1219840.1219858. URL https://doi.org/10.3115/1219840.1219858. Samuel R. Bowman and George Dahl. What will it take to fix benchmarking in natural language understanding? In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou (eds.), Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 4843–4855, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021. naacl-main.385. URL https://aclanthology.org/2021.naacl-main.385/. Richard W Brislin. Back-translation for cross-cultural research. Journal of cross-cultural psychology, 1(3):185–216, 1970. Penelope Brown and Stephen C Levinson. Politeness: Some Universals in Language Usage. Cambridge University Press, 1987. Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E Gonzalez, et al. Chatbot arena: An open platform for evaluating llms by human preference. In Forty-first International Conference on Machine Learning, 2024. Ido Dagan, Dan Roth, Mark Sammons, and Fabio Zanzotto. Recognizing textual entailment: Models and applications. Synthesis Lectures on Human Language Technologies, 6(4):1–222, 2013. ISSN 1947-4040. doi: 10.2200/S00509ED1V01Y201305HLT023. Publisher Copyright: © Morgan and Claypool Publishers. All rights reserved. Yann Dubois, Percy Liang, and Tatsunori Hashimoto. Length-controlled alpacaeval: A simple debiasing of automatic evaluators. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=CybBmzWBX0. 14
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Marzieh Fadaee, Arianna Bisazza, and Christof Monz. Examining the tip of the iceberg: A data set for idiom translation. In Nicoletta Calzolari, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Koiti Hasida, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis, and Takenobu Tokunaga (eds.), Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan, May 2018. European Language Resources Association (ELRA). URL https://aclanthology.org/L18-1148/. Ana C Farinha, M. Amin Farajian, Marianna Buchicchio, Patrick Fernandes, José G. C. de Souza, Helena Moniz, and André F. T. Martins. Findings of the WMT 2022 shared task on chat translation. In Philipp Koehn, Loı̈c Barrault, Ondřej Bojar, Fethi Bougares, Rajen Chatterjee, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Alexander Fraser, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Paco Guzman, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Tom Kocmi, André Martins, Makoto Morishita, Christof Monz, Masaaki Nagata, Toshiaki Nakazawa, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Marco Turchi, and Marcos Zampieri (eds.), Proceedings of the Seventh Conference on Machine Translation (WMT), pp. 724–743, Abu Dhabi, United Arab Emirates (Hybrid), December 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.wmt-1.70. URL https://aclanthology.org/2022.wmt-1.70/. Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain. In Loic Barrault, Ondrej Bojar, Fethi Bougares, Rajen Chatterjee, Marta R. Costa-jussa, Christian Federmann, Mark Fishel, Alexander Fraser, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Paco Guzman, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Tom Kocmi, Andre Martins, Makoto Morishita, and Christof Monz (eds.), Proceedings of the Sixth Conference on Machine Translation, pp. 733–774, Online, November 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.wmt-1.73/. Behrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna, Maxim Krikun, Xavier Garcia, Ciprian Chelba, and Colin Cherry. Scaling laws for neural machine translation. In International Conference on Learning Representations, 2022. URL https://openreview.net/ forum?id=hR SMu8cxCV. H Paul Grice. Logic and conversation. Syntax and Semantics, 3:41–58, 1975. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id= d7KBjmI3GmQ. Pierre Isabelle, Colin Cherry, and George Foster. A challenge set approach to evaluating machine translation. In Martha Palmer, Rebecca Hwa, and Sebastian Riedel (eds.), Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 2486–2496, Copenhagen, Denmark, September 2017. Association for Computational Linguistics. doi: 10.18653/v1/D17-1263. URL https://aclanthology.org/D17-1263/. Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models, 2020. URL https://arxiv.org/abs/2001.08361. Marzena Karpinska and Mohit Iyyer. Large language models effectively leverage documentlevel context for literary translation, but critical errors persist. In Philipp Koehn, Barry Haddow, Tom Kocmi, and Christof Monz (eds.), Proceedings of the Eighth Conference on Machine Translation, pp. 419–451, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.wmt-1.41. URL https://aclanthology.org/ 2023.wmt-1.41/. Ahrii Kim. RUBRIC-MQM : Span-level LLM-as-judge in machine translation for high-end models. In Georg Rehm and Yunyao Li (eds.), Proceedings of the 63rd Annual Meeting 15
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
of the Association for Computational Linguistics (Volume 6: Industry Track), pp. 147–165, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176288-6. doi: 10.18653/v1/2025.acl-industry.12. URL https://aclanthology.org/2025. acl-industry.12/. Hannah Calzi Kleidermacher and James Zou. Science across languages: Assessing llm multilingual translation of scientific papers, 2025. URL https://arxiv.org/abs/2502. 17882. Tom Kocmi and Christian Federmann. GEMBA-MQM: Detecting translation quality error spans with GPT-4. In Philipp Koehn, Barry Haddow, Tom Kocmi, and Christof Monz (eds.), Proceedings of the Eighth Conference on Machine Translation, pp. 768–775, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.wmt-1. 64. URL https://aclanthology.org/2023.wmt-1.64/. Tom Kocmi, Ekaterina Artemova, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Konstantin Dranch, Anton Dvorkovich, Sergey Dukanov, Mark Fishel, Markus Freitag, et al. Findings of the wmt25 general machine translation shared task: Time to stop evaluating on easy test sets. In Proceedings of the Tenth Conference on Machine Translation, pp. 355–413, 2025. Philipp Koehn and Rebecca Knowles. Six challenges for neural machine translation. In Thang Luong, Alexandra Birch, Graham Neubig, and Andrew Finch (eds.), Proceedings of the First Workshop on Neural Machine Translation, pp. 28–39, Vancouver, August 2017. Association for Computational Linguistics. doi: 10.18653/v1/W17-3204. URL https: //aclanthology.org/W17-3204/. Samuel Läubli, Rico Sennrich, and Martin Volk. Has machine translation achieved human parity? a case for document-level evaluation. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 4791–4796, Brussels, Belgium, October-November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1512. URL https://aclanthology.org/D18-1512/. Alon Lavie, Greg Hanneman, Sweta Agrawal, Diptesh Kanojia, Chi-Kiu Lo, Vilém Zouhar, Frederic Blain, Chrysoula Zerva, Eleftherios Avramidis, Sourabh Deoghare, Archchana Sindhujan, Jiayi Wang, David Ifeoluwa Adelani, Brian Thompson, Tom Kocmi, Markus Freitag, and Daniel Deutsch. Findings of the WMT25 shared task on automated translation evaluation systems: Linguistic diversity is challenging and references still help. In Barry Haddow, Tom Kocmi, Philipp Koehn, and Christof Monz (eds.), Proceedings of the Tenth Conference on Machine Translation, pp. 436–483, Suzhou, China, November 2025. Association for Computational Linguistics. ISBN 979-8-89176-341-8. doi: 10.18653/v1/ 2025.wmt-1.24. URL https://aclanthology.org/2025.wmt-1.24/. Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. From crowdsourced data to high-quality benchmarks: Arenahard and benchbuilder pipeline. In Forty-second International Conference on Machine Learning, 2025. URL https://openreview.net/forum?id=KfTf9vFvSn. Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, et al. Deepseek-v3. 2: Pushing the frontier of open large language models. arXiv preprint arXiv:2512.02556, 2025. Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics. Tradumàtica, (12):0455–463, 2014. Arle Lommel, Serge Gladkoff, Alan Melby, Sue Ellen Wright, Ingemar Strandvik, Katerina Gasova, Angelika Vaasa, Andy Benzo, Romina Marazzato Sparano, Monica Foresi, Johani Innis, Lifeng Han, and Goran Nenadic. The multi-range theory of translation quality measurement: MQM scoring models and statistical quality control. In Marianna Martindale, 16
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Janice Campbell, Konstantin Savenkov, and Shivali Goel (eds.), Proceedings of the 16th Conference of the Association for Machine Translation in the Americas (Volume 2: Presentations), pp. 75–94, Chicago, USA, September 2024. Association for Machine Translation in the Americas. URL https://aclanthology.org/2024.amta-presentations.6/. Qingyu Lu, Baopu Qiu, Liang Ding, Kanjian Zhang, Tom Kocmi, and Dacheng Tao. Error analysis prompting enables human-like translation evaluation in large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Findings of the Association for Computational Linguistics: ACL 2024, pp. 8801–8816, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-acl.520. URL https://aclanthology.org/2024.findings-acl.520/. Jihyung Moon, Hyunchang Cho, and Eunjeong L. Park. Revisiting round-trip translation for quality estimation. In André Martins, Helena Moniz, Sara Fumega, Bruno Martins, Fernando Batista, Luisa Coheur, Carla Parra, Isabel Trancoso, Marco Turchi, Arianna Bisazza, Joss Moorkens, Ana Guerberof, Mary Nurminen, Lena Marg, and Mikel L. Forcada (eds.), Proceedings of the 22nd Annual Conference of the European Association for Machine Translation, pp. 91–104, Lisboa, Portugal, November 2020. European Association for Machine Translation. URL https://aclanthology.org/2020.eamt-1.11/. Jinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng, Mahir Shah, Kabir Jain, Graham Neubig, and Yang You. Mixeval: Deriving wisdom of the crowd from llm benchmark mixtures. Advances in Neural Information Processing Systems, 37:98180–98212, 2024. Dojun Park, Jiwoo Lee, Seohyun Park, Hyeyun Jeong, Youngeun Koo, Soonha Hwang, Seonwoo Park, and Sungeun Lee. MultiPragEval: Multilingual pragmatic evaluation of large language models. In Dieuwke Hupkes, Verna Dankers, Khuyagbaatar Batsuren, Amirhossein Kazemnejad, Christos Christodoulopoulos, Mario Giulianelli, and Ryan Cotterell (eds.), Proceedings of the 2nd GenBench Workshop on Generalisation (Benchmarking) in NLP, pp. 96–119, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.genbench-1.7. URL https://aclanthology.org/2024. genbench-1.7/. Qwen Team. Qwen3: Think deeper, act faster, April 2025. URL https://qwen.ai/blog?id= qwen3. Angelika Romanou, Negar Foroutan, Anna Sotnikova, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Maheshwary, Micol Altomare, Zeming Chen, Mohamed A. Haggag, Snegha A, Alfonso Amayuelas, and Azril Hafizi et al. INCLUDE: Evaluating multilingual language understanding with regional knowledge. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id= k3gCieTXeY. Ananya B. Sai, Akash Kumar Mohankumar, and Mitesh M. Khapra. A survey of evaluation metrics used for nlg systems. ACM Comput. Surv., 55(2), January 2022. ISSN 0360-0300. doi: 10.1145/3485766. URL https://doi.org/10.1145/3485766. Oscar Sainz, Iker Garcı́a-Ferrero, Alon Jacovi, Jon Ander Campos, Yanai Elazar, Eneko Agirre, Yoav Goldberg, Wei-Lin Chen, Jenny Chim, Leshem Choshen, Luca D’Amico-Wong, Melissa Dell, Run-Ze Fan, Shahriar Golchin, Yucheng Li, Pengfei Liu, Bhavish Pahwa, Ameya Prabhu, Suryansh Sharma, Emily Silcock, Kateryna Solonko, David Stap, Mihai Surdeanu, Yu-Min Tseng, Vishaal Udandarao, Zengzhi Wang, Ruijie Xu, and Jinglin Yang. Data contamination report from the 2024 CONDA shared task. CoRR, abs/2407.21530, 2024. doi: 10.48550/ARXIV.2407.21530. URL https://doi.org/10.48550/arXiv.2407. 21530. John R Searle. Speech Acts: An Essay in the Philosophy of Language. Cambridge University Press, 1969. Rico Sennrich, Barry Haddow, and Alexandra Birch. Improving neural machine translation models with monolingual data. In Proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: long papers), pp. 86–96, 2016. 17
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/ forum?id=fR3wGCk-IXp. Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, and Wei Qi et al. Leong. Global MMLU: Understanding and addressing cultural and linguistic biases in multilingual evaluation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2025. doi: 10.18653/v1/2025.acl-long.919. URL https://aclanthology.org/2025.acl-long.919/. Harold Somers. Round-trip translation: What is it good for? In Proceedings of the Australasian Language Technology Workshop 2005, pp. 127–133, 2005. Guijin Son, Jiwoo Hong, Hyunwoo Ko, and James Thorne. Linguistic generalizability of test-time scaling in mathematical reasoning. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, July 2025. doi: 10.18653/v1/2025.acl-long.699. URL https: //aclanthology.org/2025.acl-long.699/. Lucas Spangher, Tianle Li, William F. Arnold, Nick Masiewicki, Xerxes Dotiwalla, Rama Kumar Pasumarthi, Peter Grabowski, Eugene Ie, and Daniel Gruhl. Chatbot arena estimate: towards a generalized performance benchmark for LLM capabilities. In Weizhu Chen, Yi Yang, Mohammad Kachuee, and Xue-Yong Fu (eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track), pp. 1016–1025, Albuquerque, New Mexico, April 2025. Association for Computational Linguistics. ISBN 979-8-89176194-0. doi: 10.18653/v1/2025.naacl-industry.77. URL https://aclanthology.org/2025. naacl-industry.77/. Chihiro Taguchi, Seng Mai, Keita Kurabe, Yusuke Sakai, Georgina Agyei, Soudabeh Eslami, and David Chiang. Languages still left behind: Toward a better multilingual machine translation benchmark. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 20131–20143, Suzhou, China, November 2025. Association for Computational Linguistics. ISBN 979-8-89176-332-6. doi: 10.18653/v1/2025. emnlp-main.1018. URL https://aclanthology.org/2025.emnlp-main.1018/. Kimi Team, Yifan Bai, Yiping Bao, Guanduo Chen, Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, et al. Kimi k2: Open agentic intelligence. arXiv preprint arXiv:2507.20534, 2025. NLLB Team, Marta R. Costa-jussà, and James Cross et al. No language left behind: Scaling human-centered machine translation, 2022. URL https://arxiv.org/abs/2207.04672. Menno van Zaanen and Simon Zwarts. Unsupervised measurement of translation quality using multi-engine, bi-directional translation. In Proceedings of the 19th Australian Joint Conference on Artificial Intelligence: Advances in Artificial Intelligence, AI’06, pp. 1208–1214, Berlin, Heidelberg, 2006. Springer-Verlag. ISBN 3540497870. doi: 10.1007/11941439 149. URL https://doi.org/10.1007/11941439 149. David Vilar, Markus Freitag, Colin Cherry, Jiaming Luo, Viresh Ratnakar, and George Foster. Prompting PaLM for translation: Assessing strategies and performance. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15406–15427, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.859. URL https://aclanthology.org/2023.acl-long.859/. 18
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Yiming Wang, Pei Zhang, Jialong Tang, Hao-Ran Wei, Baosong Yang, Rui Wang, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang, Junyang Lin, Fei Huang, and Jingren Zhou. Polymath: Evaluating mathematical reasoning in multilingual contexts. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025. URL https://openreview.net/ forum?id=B1vCImy6yI. Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. MMLU-pro: A more robust and challenging multi-task language understanding benchmark. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024. URL https://openreview.net/forum?id=y10DM6R2r3. Minghao Wu, Weixuan Wang, Sinuo Liu, Huifeng Yin, Xintong Wang, Yu Zhao, Chenyang Lyu, Longyue Wang, Weihua Luo, and Kaifu Zhang. The bitter lesson learned from 2,000+ multilingual benchmarks. arXiv preprint arXiv:2504.15521, 2025. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, and Dayiheng Liu et al. Qwen3 technical report, 2025. URL https://arxiv.org/abs/2505.09388. Binwei Yao, Ming Jiang, Tara Bobinac, Diyi Yang, and Junjie Hu. Benchmarking machine translation with cultural awareness. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 13078–13096, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-emnlp.765. URL https://aclanthology. org/2024.findings-emnlp.765/. Z.ai, 5 Team, Aohan Zeng, Xin Lv, Qinkai Zheng, Zhenyu Hou, Bin Chen, Chengxing Xie, Cunxiang Wang, Da Yin, Hao Zeng, Jiajie Zhang, Kedong Wang, Lucen Zhong, Mingdao Liu, Rui Lu, Shulin Cao, Xiaohan Zhang, Xuancheng Huang, Yao Wei, Yean Cheng, Yifan An, Yilin Niu, Yuanhao Wen, Yushi Bai, Zhengxiao Du, Zihan Wang, Zilin Zhu, Bohan Zhang, Bosi Wen, Bowen Wu, Bowen Xu, Can Huang, Casey Zhao, Changpeng Cai, Chao Yu, Chen Li, Chendi Ge, Chenghua Huang, Chenhui Zhang, Chenxi Xu, Chenzheng Zhu, Chuang Li, Congfeng Yin, Daoyan Lin, Dayong Yang, Dazhi Jiang, Ding Ai, Erle Zhu, Fei Wang, Gengzheng Pan, Guo Wang, Hailong Sun, Haitao Li, Haiyang Li, Haiyi Hu, Hanyu Zhang, Hao Peng, Hao Tai, Haoke Zhang, Haoran Wang, Haoyu Yang, He Liu, He Zhao, Hongwei Liu, Hongxi Yan, Huan Liu, Huilong Chen, Ji Li, Jiajing Zhao, Jiamin Ren, Jian Jiao, Jiani Zhao, Jianyang Yan, Jiaqi Wang, Jiayi Gui, Jiayue Zhao, Jie Liu, Jijie Li, Jing Li, Jing Lu, Jingsen Wang, Jingwei Yuan, Jingxuan Li, Jingzhao Du, Jinhua Du, Jinxin Liu, Junkai Zhi, Junli Gao, Ke Wang, Lekang Yang, Liang Xu, Lin Fan, Lindong Wu, Lintao Ding, Lu Wang, Man Zhang, Minghao Li, Minghuan Xu, Mingming Zhao, Mingshu Zhai, Pengfan Du, Qian Dong, Shangde Lei, Shangqing Tu, Shangtong Yang, Shaoyou Lu, Shijie Li, Shuang Li, Shuang-Li, Shuxun Yang, Sibo Yi, Tianshu Yu, Wei Tian, Weihan Wang, Wenbo Yu, Weng Lam Tam, Wenjie Liang, Wentao Liu, Xiao Wang, Xiaohan Jia, Xiaotao Gu, Xiaoying Ling, Xin Wang, Xing Fan, Xingru Pan, Xinyuan Zhang, Xinze Zhang, Xiuqing Fu, Xunkai Zhang, Yabo Xu, Yandong Wu, Yida Lu, Yidong Wang, Yilin Zhou, Yiming Pan, Ying Zhang, Yingli Wang, Yingru Li, Yinpei Su, Yipeng Geng, Yitong Zhu, Yongkun Yang, Yuhang Li, Yuhao Wu, Yujiang Li, Yunan Liu, Yunqing Wang, Yuntao Li, Yuxuan Zhang, Zezhen Liu, Zhen Yang, Zhengda Zhou, Zhongpei Qiao, Zhuoer Feng, Zhuorui Liu, Zichen Zhang, Zihan Wang, Zijun Yao, Zikang Wang, Ziqiang Liu, Ziwei Chai, Zixuan Li, Zuodong Zhao, Wenguang Chen, Jidong Zhai, Bin Xu, Minlie Huang, Hongning Wang, Juanzi Li, Yuxiao Dong, and Jie Tang. Glm-4.5: Agentic, reasoning, and coding (arc) foundation models, 2025. URL https://arxiv.org/abs/2508.06471. Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. Large language models are not robust multiple choice selectors. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=shr9PXz7T0. 19
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Terry Yue Zhuo, Qiongkai Xu, Xuanli He, and Trevor Cohn. Rethinking round-trip translation for machine translation evaluation. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Findings of the Association for Computational Linguistics: ACL 2023, pp. 319–337, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.22. URL https://aclanthology.org/2023.findings-acl. 22/.
20
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Part I
Appendix Contents A Additional Robustness Benchmark Details A.1 Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B Robustness of Metric and Judge Choice B.1 Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
22 22 23 23
C Robustness of Answer Language Analysis Across Models
28
D Robustness of Error Distribution Analysis Across Models
33
E Extended Experimental Details
35
E.1 Dataset Composition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
35
E.2 Sampling Hyperparameters . . . . . . . . . . . . . . . . . . . . . . . . . . . .
35
E.3 Model Prompt . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
36
E.4 Judge Prompt . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
36
21
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
A
Additional Robustness Benchmark Details
Dataset Construction. We construct a dedicated robustness extension of our LiT benchmark consisting of 480 samples (60 source sentences × 8 language sequences), targeting five categories of classically challenging phenomena: (i) Conceptual and Abstract Nuance: Passages where physical or spatial vocabulary is used metaphorically to describe abstract concepts (e.g., ”a heavy decision”), testing whether models preserve figurative meaning across languages. (ii) Idioms and Cultural Metaphors: Figurative expressions whose meaning cannot be recovered from literal word-by-word translation, testing whether models preserve communicative intent rather than surface form. (iii) Polysemy and Lexical Ambiguity: Words carrying multiple distinct meanings where only surrounding context disambiguates the intended sense (e.g., ”bank” as financial institution vs. riverbank). (iv) Register and Tone Shifts: Passages that shift between formal, colloquial, and technical registers within a single example, testing sensitivity to sociolinguistic appropriateness. (v) Syntactic Complexity and Garden Paths: Sentences whose grammatical structure remains ambiguous until late in the sentence, forcing re-parsing and testing whether models maintain structural fidelity through the translation chain. A.1
Results
Table 5 showcases the high correlation between the main LiT benchmark and the robustness benchmark. This indicates that round-trip translation with state-of-the-art frontier models is robust to classical backtranslation weaknesses. Table 5: Rank correlation between LiT categories and the robustness benchmark under MQM≥80 . We report Spearman rank correlations between model performance on each LiT category and the robustness benchmark average, using two-sided tests over all models. LiT Category Average Humanities STEM Core Semantics Discourse Coherence Implicit Content Pragmatic Inference Social Interaction Informal
Spearman ρ
p-value
Significance
0.946 0.917 0.901 0.959 0.953 0.954 0.949 0.940 0.777
3.11 × 10−10
*** *** *** *** *** *** *** *** ***
22
1.27 × 10−8 5.98 × 10−8 2.46 × 10−11 8.98 × 10−11 7.78 × 10−11 1.88 × 10−10 7.46 × 10−10 5.49 × 10−5
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
B
Robustness of Metric and Judge Choice
A potential concern with any LLM-as-a-judge evaluation is sensitivity to the choice of metric formulation and judge model. If model rankings shifted substantially under alternative scoring rubrics, the benchmark’s conclusions would be fragile. We address this concern by evaluating all models under two complementary scoring metrics and verifying that rankings remain stable across them. Metrics. In the main paper, we report MQM≥80 : the percentage of translations whose MQM score meets or exceeds the 80-point threshold widely considered to indicate fit-for-purpose translation quality. We additionally evaluate two alternative metrics: (i) Raw MQM scores (Table 6 and 8): Rather than binarizing at the 80-point threshold, we report the average continuous MQM score for each model–sequence combination. This preserves the full distribution of translation quality and avoids potential artifacts introduced by a fixed cutoff. (ii) Direct judge scores (Tables 7 and 9): We prompt the judge model to assign a holistic quality score on a continuous 0–100 scale, without reference to the MQM error taxonomy. This tests whether the structured MQM framework and a simple judge-based scoring rubric converge on the same model ordering. B.1
Results
Comparing global averages across the three scoring paradigms reveals highly stable rankings. The top tier is unchanged across all three metrics: Gemini-3-Flash leads decisively, followed by Gemma-4-31B (Instruct) and Qwen3.5-397B (Thinking). The bottom tier is equally stable, with Nemotron-3-Nano and Qwen3-30B (Instruct) consistently occupying the lowest positions. Mid-table models exhibit only modest reshuffling (typically within 1–2 rank positions), which is expected given that these models perform similarly and minor scoring differences can reorder near-tied entries. The language-sequence breakdown further confirms robustness. All three metrics agree on the central finding: performance collapses catastrophically from high-resource to lowresource language sequences. Under raw MQM (Table 8), only Gemini-3-Flash maintains a clearly usable average score (75.4) on low-resource sequences, while the next-best model drops to 43.7. Under judge scores (Table 9), the same pattern holds, with Gemini-3-Flash at 79.4 and the runner-up at 64.4. The qualitative conclusion – that a steep accuracy cliff separates high-resource from low-resource performance – is invariant to the metric. Summary. Across three metrics (MQM≥80 , Raw MQM, and direct judge scores) and validated against an external human-preference benchmark (LMArena), model rankings on LiT remain highly consistent. This stability indicates that our findings – including the disconnect between reasoning benchmarks and multilingual proficiency, the low-resource performance collapse, and the underperformance of reasoning models on informal text – are robust properties of the models themselves, not artifacts of a particular evaluation configuration.
23
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Table 6: LiT benchmark by linguistic category under MQM. We report mean MQM scores (higher is better), aggregated within each category over the same eight translation sequences. Each cell shows the mean and the bootstrap-estimated standard error, computed via sentence-level resampling after averaging across sequences. Informal text is the hardest category overall (36.7 avg), while Core Semantics is the easiest (64.4). Model
Average
Abstracts
Pragmatics
Informal
Humanities STEM Core Semantics Discourse Implicit Inference Social Gemini-3-Flash (No-Thinking)
86.9
88.6
86.4
88.8
87.7
85.2
87.5
87.5
83.8
± 0.3
± 0.8
± 0.9
± 0.8
± 0.8
± 1.0
± 0.8
± 0.7
± 0.8
78.1
76.9
87.5
85.2
80.8
80.0
79.3
79.0
56.5
Qwen3.5-397B (Thinking)
± 0.6
± 2.1
± 1.1
± 1.1
± 1.4
± 1.8
± 1.5
± 1.2
± 2.2
75.2
77.4
82.0
75.6
77.8
72.3
76.8
75.3
64.7
Gemma-4-31B (Instruct)
± 0.5
± 1.5
± 1.8
± 1.6
± 1.6
± 1.8
± 1.2
± 1.1
± 1.6
73.7
77.2
81.6
74.9
75.0
71.3
74.0
73.6
62.4
Gemma-4-31B (Thinking)
± 0.6
± 1.5
± 1.7
± 1.8
± 1.5
± 2.3
± 1.5
± 1.3
± 1.9
73.5
76.1
79.9
79.9
77.8
74.0
74.1
74.6
51.4
GLM-5 (Thinking)
± 0.7
± 2.1
± 2.1
± 1.5
± 1.3
± 2.3
± 1.7
± 1.8
± 2.3
72.5
76.1
78.6
76.6
73.2
73.5
74.6
72.8
54.4
Qwen3.5-397B (Instruct)
± 0.8
± 2.4
± 2.9
± 2.1
± 1.7
± 2.4
± 1.4
± 1.4
± 2.9
63.2
67.1
69.4
68.5
67.5
64.4
65.8
62.6
40.1
GLM-4.7 (Thinking)
± 0.8
± 2.6
± 2.4
± 2.3
± 1.8
± 2.3
± 1.6
± 1.8
± 2.1
DeepSeek-V3.2-Exp (Thinking)
62.9
64.8
51.3
72.4
69.2
67.6
66.0
69.0
43.1
± 0.9
± 3.0
± 4.7
± 1.7
± 1.6
± 2.4
± 2.1
± 1.4
± 2.7
61.5
58.4
59.9
70.0
67.8
63.2
66.8
63.5
42.5
Kimi-K2 (Thinking)
± 0.8
± 3.2
± 3.9
± 1.3
± 1.8
± 2.0
± 1.5
± 1.4
± 2.5
59.2
63.1
49.7
64.2
61.7
62.9
65.1
62.1
44.6
DeepSeek-V3.2-Exp
± 1.0
± 3.0
± 4.9
± 1.8
± 2.3
± 2.6
± 1.5
± 1.6
± 2.4
58.8
59.2
33.7
69.1
67.5
61.5
69.3
63.4
47.1
Gemma-3-27B (Instruct)
± 1.1
± 2.9
± 6.9
± 1.3
± 1.2
± 2.0
± 1.3
± 1.9
± 2.0
57.8
55.0
67.0
66.9
64.5
60.2
65.2
59.7
24.3
Qwen3.5-35B (Thinking)
± 1.0
± 3.3
± 3.0
± 2.0
± 1.7
± 3.9
± 1.2
± 1.9
± 3.5
57.6
56.9
56.3
65.9
60.1
60.7
60.6
58.1
41.9
GLM-5 (Instruct)
± 1.0
± 3.7
± 4.9
± 1.8
± 2.3
± 2.4
± 1.9
± 1.8
± 2.2
54.9
53.2
56.4
62.8
60.8
56.8
62.0
55.8
31.5
Qwen3-235B (Thinking)
± 0.8
± 2.4
± 3.2
± 1.4
± 1.8
± 2.8
± 1.0
± 1.5
± 2.5
52.1
47.6
38.5
61.0
56.8
55.5
59.3
56.4
41.9
Kimi-K2
± 1.1
± 4.0
± 5.8
± 1.6
± 2.2
± 2.5
± 1.6
± 1.8
± 2.8
52.1
48.7
61.2
62.4
60.0
53.7
58.7
52.8
19.3
GPT-OSS-120B (High)
± 0.9
± 3.1
± 2.5
± 1.5
± 2.6
± 4.1
± 1.4
± 2.0
± 3.3
49.4
45.5
42.4
58.9
52.9
53.6
53.1
54.4
34.3
Qwen3-235B (Instruct)
± 0.9
± 3.0
± 4.4
± 1.5
± 1.9
± 1.7
± 1.7
± 1.8
± 2.5
49.4
45.1
51.9
56.3
55.6
51.1
60.5
52.9
21.4
Qwen3.5-35B (Instruct)
± 1.0
± 3.2
± 4.3
± 1.7
± 1.9
± 3.4
± 1.5
± 1.8
± 3.6
47.2
43.4
50.0
61.3
58.7
52.1
54.2
51.7
6.5
MiniMax-M2.5
± 1.0
± 3.6
± 3.6
± 1.9
± 1.9
± 3.8
± 1.9
± 1.9
± 3.6
38.0
33.8
12.2
51.7
52.1
44.8
48.9
43.8
16.6
MiMo-V2-Flash
± 1.0
± 2.3
± 5.4
± 1.9
± 2.7
± 2.5
± 1.8
± 1.3
± 3.0
31.1
24.8
19.3
42.3
40.2
39.4
39.7
37.8
5.1
Qwen3-30B (Instruct)
± 1.1
± 3.7
± 5.4
± 2.2
± 3.0
± 2.2
± 1.5
± 2.1
± 2.5
-6.3
-11.9
-7.2
1.2
0.2
-5.0
0.4
-1.4
-26.6
Nemotron-3-Nano
± 1.2
± 3.2
± 5.0
± 3.0
± 2.8
± 3.9
± 2.9
± 2.5
± 3.4
Average
56.8
55.8
54.9
64.4
62.2
59.0
61.9
59.3
36.7
24
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Table 7: LiT benchmark by linguistic category under raw Judge Score. We report mean judge scores (higher is better), aggregated within each category over the same eight translation sequences. Each cell shows the mean and the bootstrap-estimated standard error, computed via sentence-level resampling after averaging across sequences. Informal text is the hardest category overall (63.3 avg), while STEM is the easiest (71.3). Model
Average
Abstracts
Pragmatics
Informal
Humanities STEM Core Semantics Discourse Implicit Inference Social Gemini-3-Flash (No-Thinking)
89.1
91.2
90.5
89.6
88.9
86.2
88.5
88.2
89.5
± 0.2
± 0.6
± 0.7
± 0.8
± 0.9
± 0.9
± 0.6
± 0.5
± 0.5
83.9
85.2
91.1
86.8
84.9
81.9
82.5
83.1
75.7
Qwen3.5-397B (Thinking)
± 0.4
± 1.2
± 0.7
± 0.8
± 0.9
± 1.8
± 0.9
± 0.7
± 1.5
80.9
83.1
87.6
79.9
80.7
76.3
80.1
79.4
80.1
Gemma-4-31B (Instruct)
± 0.3
± 1.0
± 1.0
± 0.9
± 1.1
± 1.2
± 0.7
± 0.6
± 0.6
80.7
82.7
86.4
82.9
82.6
77.8
80.0
79.5
73.4
GLM-5 (Thinking)
± 0.4
± 1.0
± 1.0
± 1.1
± 0.8
± 1.7
± 0.8
± 0.9
± 1.5
80.1
83.6
85.0
80.4
78.9
77.2
79.7
78.6
77.4
Qwen3.5-397B (Instruct)
± 0.4
± 1.6
± 1.7
± 1.5
± 1.1
± 1.6
± 0.9
± 0.7
± 0.8
79.7
82.2
84.0
79.6
79.3
76.0
79.3
78.6
79.0
Gemma-4-31B (Thinking)
± 0.4
± 0.8
± 1.4
± 1.2
± 1.2
± 1.4
± 0.9
± 0.8
± 0.7
75.2
77.7
81.3
77.7
77.7
72.5
74.4
73.5
67.0
GLM-4.7 (Thinking)
± 0.4
± 1.1
± 1.2
± 1.1
± 0.7
± 1.7
± 0.7
± 0.9
± 1.3
DeepSeek-V3.2-Exp (Thinking)
73.7
77.2
72.1
76.3
75.2
71.7
73.6
73.4
70.0
± 0.5
± 1.7
± 2.2
± 1.1
± 1.2
± 1.5
± 0.9
± 0.9
± 1.6
72.3
72.7
75.3
74.1
74.1
69.8
72.7
72.1
67.9
Kimi-K2 (Thinking)
± 0.5
± 1.4
± 2.0
± 0.8
± 0.9
± 1.9
± 0.9
± 1.0
± 1.6
71.9
72.8
74.4
73.6
72.5
70.1
71.4
70.7
69.6
GLM-5 (Instruct)
± 0.5
± 1.7
± 2.2
± 0.9
± 1.1
± 1.5
± 0.8
± 0.9
± 0.8
71.5
75.2
68.7
72.1
72.1
70.2
72.3
70.0
71.3
DeepSeek-V3.2-Exp
± 0.5
± 1.6
± 2.4
± 1.3
± 1.2
± 1.3
± 0.9
± 0.9
± 0.7
69.8
69.4
78.7
74.1
71.9
67.8
71.3
67.4
57.9
Qwen3.5-35B (Thinking)
± 0.5
± 1.4
± 1.4
± 1.3
± 1.4
± 2.1
± 0.8
± 1.0
± 1.2
67.3
68.6
71.9
69.8
70.2
65.2
67.6
65.6
59.6
Qwen3-235B (Thinking)
± 0.4
± 1.1
± 1.4
± 0.9
± 0.9
± 2.0
± 0.6
± 0.9
± 1.1
67.1
69.5
60.2
70.0
70.1
64.5
69.6
66.1
67.2
Gemma-3-27B (Instruct)
± 0.5
± 1.2
± 2.5
± 1.0
± 1.0
± 1.6
± 0.8
± 0.9
± 0.8
65.4
68.0
71.9
69.7
70.0
64.4
67.5
64.2
47.7
GPT-OSS-120B (High)
± 0.5
± 1.4
± 1.7
± 0.8
± 1.4
± 2.0
± 0.9
± 1.1
± 1.8
63.6
62.1
70.4
68.9
68.8
63.2
64.3
61.2
50.1
MiniMax-M2.5
± 0.5
± 1.6
± 1.6
± 1.1
± 1.2
± 2.1
± 0.9
± 1.1
± 1.2
63.0
62.6
69.8
65.1
63.2
60.7
65.1
60.4
57.4
Qwen3.5-35B (Instruct)
± 0.5
± 1.4
± 1.9
± 1.0
± 1.2
± 1.9
± 0.9
± 1.1
± 1.1
62.5
64.3
60.5
65.3
63.4
59.8
63.1
62.2
61.6
Kimi-K2
± 0.6
± 1.6
± 3.0
± 1.2
± 1.3
± 1.9
± 0.8
± 1.0
± 1.0
56.6
54.6
57.1
61.0
57.1
57.3
56.7
56.0
53.3
Qwen3-235B (Instruct)
± 0.6
± 2.4
± 2.4
± 1.4
± 1.8
± 1.8
± 1.4
± 1.6
± 1.6
56.2
55.5
49.1
60.6
60.3
56.2
58.2
56.1
53.8
MiMo-V2-Flash
± 0.5
± 1.4
± 2.6
± 0.8
± 1.2
± 1.7
± 1.1
± 1.1
± 0.9
47.9
47.0
47.8
50.4
51.4
47.1
49.2
48.6
41.3
Qwen3-30B (Instruct)
± 0.5
± 1.2
± 2.1
± 1.2
± 1.2
± 1.5
± 0.6
± 1.1
± 0.9
29.2
30.1
35.6
31.4
31.0
27.0
28.9
28.9
21.0
Nemotron-3-Nano
± 0.5
± 1.4
± 2.0
± 1.2
± 1.0
± 1.4
± 1.3
± 1.3
± 0.9
Average
68.5
69.8
71.3
70.9
70.2
66.5
68.9
67.4
63.3
25
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Table 8: Raw MQM scores similarly collapse on low-resource languages. We report the raw MQM scores across eight language sequences grouped by resource availability. (a) High-resource sequences (East Asian, Central European, Near Eastern): top frontier models remain strong, with Gemini-3-Flash leading at 91.0 average and GLM-5 (Thinking) taking the East Asian column at 90.4. (b) Medium-resource sequences (Southeast Asian, South Asian, North European): performance remains stable for the strongest models, with Gemini3-Flash averaging 89.9. (c) Low-resource sequences (African, South American): catastrophic collapse remains, with only Gemini-3-Flash (75.4) maintaining clearly usable performance; the next-best model (Qwen3.5-397B (Thinking)) drops to 43.7. (d) Global averages across these eight sequences show Gemini-3-Flash holding roughly a 10.5-point lead over the next competitor. (c) Low Resource / Imbalanced Sequences
(a) High & Med-High Resource Sequences Model
E. Asia C. Europe N. East Avg
Gemini-3-Flash (No-Think) Qwen3.5-397B (Thinking) GLM-5 (Thinking) Gemma-4-31B (Thinking) Gemma-4-31B (Instruct) Qwen3.5-397B (Instruct) GLM-4.7 (Thinking) Kimi-K2 (Thinking) DeepSeek-V3.2-Exp Qwen3-235B (Thinking) GLM-5 (Instruct) DeepSeek-V3.2-Exp (Think) Qwen3.5-35B (Thinking) Gemma-3-27B (Instruct) Kimi-K2 Qwen3.5-35B (Instruct) Qwen3-235B (Instruct) MiniMax-M2.5 GPT-OSS-120B (High) MiMo-V2-Flash Qwen3-30B (Instruct) Nemotron-3-Nano
90.0 89.3 90.4 89.6 88.6 88.1 89.3 86.1 87.6 85.0 86.3 85.8 81.2 78.6 79.4 79.3 83.7 71.6 78.3 70.2 71.3 50.9
91.6 90.3 89.7 88.0 88.2 89.3 89.0 87.2 86.2 87.3 85.9 82.9 85.3 83.1 82.8 80.7 68.3 80.2 76.8 74.5 67.5 11.5
91.6 87.1 85.4 86.8 86.0 85.2 82.9 82.8 79.9 80.0 77.8 79.7 75.7 77.8 70.5 64.1 72.0 68.7 63.9 54.3 44.4 -17.1
91.0 88.9 88.5 88.1 87.6 87.5 87.1 85.4 84.6 84.1 83.3 82.8 80.8 79.8 77.6 74.7 74.7 73.5 73.0 66.4 61.0 15.1
Model
Africa S. America Avg
Gemini-3-Flash (No-Think) Qwen3.5-397B (Thinking) Gemma-4-31B (Instruct) Gemma-4-31B (Thinking) Qwen3.5-397B (Instruct) GLM-5 (Thinking) DeepSeek-V3.2-Exp (Think) Gemma-3-27B (Instruct) Kimi-K2 (Thinking) Qwen3.5-35B (Thinking) GLM-4.7 (Thinking) DeepSeek-V3.2-Exp GPT-OSS-120B (High) Nemotron-3-Nano Qwen3.5-35B (Instruct) Kimi-K2 GLM-5 (Instruct) Qwen3-30B (Instruct) Qwen3-235B (Instruct) MiniMax-M2.5 Qwen3-235B (Thinking) MiMo-V2-Flash
84.4 58.5 74.8 72.6 45.9 49.1 29.6 16.8 11.7 13.8 18.4 13.6 3.4 -12.8 -11.4 -16.2 9.2 -12.7 -17.3 -15.1 -35.8 -34.6
Model Gemini-3-Flash (No-Think) Qwen3.5-397B (Thinking) GLM-5 (Thinking) Gemma-4-31B (Instruct) Gemma-4-31B (Thinking) Qwen3.5-397B (Instruct) GLM-4.7 (Thinking) Kimi-K2 (Thinking) GLM-5 (Instruct) DeepSeek-V3.2-Exp Qwen3-235B (Thinking) DeepSeek-V3.2-Exp (Think) Gemma-3-27B (Instruct) Qwen3.5-35B (Thinking) Kimi-K2 Qwen3-235B (Instruct) GPT-OSS-120B (High) Qwen3.5-35B (Instruct) MiniMax-M2.5 MiMo-V2-Flash Qwen3-30B (Instruct) Nemotron-3-Nano
SE. Asia S. Asia N. Europe Avg 89.3 86.0 86.3 88.6 86.5 85.8 80.9 79.6 75.3 81.4 81.1 79.9 76.5 74.4 72.7 77.2 67.7 68.0 61.6 62.3 53.3 -8.1
90.3 83.1 83.3 88.2 89.0 79.7 78.1 73.2 75.5 66.4 71.5 65.2 66.9 62.7 58.3 65.6 58.0 46.6 34.8 38.7 18.1 -41.8
90.2 86.3 83.7 76.0 74.7 81.9 74.6 77.6 68.0 70.8 63.5 68.8 66.8 66.6 67.0 43.7 59.4 53.6 57.9 40.1 1.7 -46.0
89.9 85.1 84.5 84.2 83.4 82.5 77.9 76.8 72.9 72.9 72.0 71.3 70.1 67.9 66.0 62.2 61.7 56.1 51.4 47.0 24.4 -31.9
26
75.4 43.7 39.2 33.2 28.3 26.5 15.0 8.9 -2.9 -3.0 -3.2 -3.7 -4.9 -5.6 -7.3 -8.6 -9.2 -10.8 -11.6 -11.8 -22.1 -22.9
(d) Overall Performance Model
(b) Medium Resource Sequences
66.4 28.9 3.5 -6.2 10.7 3.9 0.4 1.0 -17.5 -19.7 -24.7 -21.0 -13.2 1.5 -3.2 -1.0 -27.6 -8.9 -5.9 -8.4 -8.4 -11.1
Gemini-3-Flash (No-Think) Qwen3.5-397B (Thinking) Gemma-4-31B (Instruct) Gemma-4-31B (Thinking) GLM-5 (Thinking) Qwen3.5-397B (Instruct) DeepSeek-V3.2-Exp (Think) GLM-4.7 (Thinking) Kimi-K2 (Thinking) Gemma-3-27B (Instruct) DeepSeek-V3.2-Exp GLM-5 (Instruct) Qwen3.5-35B (Thinking) Qwen3-235B (Thinking) Kimi-K2 GPT-OSS-120B (High) Qwen3-235B (Instruct) Qwen3.5-35B (Instruct) MiniMax-M2.5 MiMo-V2-Flash Qwen3-30B (Instruct) Nemotron-3-Nano
Global Average 86.7 76.2 74.2 72.6 71.5 70.8 61.5 61.1 60.1 58.4 58.1 56.3 55.0 53.0 51.7 49.3 48.4 47.2 43.9 36.8 29.3 -7.7
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Table 9: Judge scores also drop on low-resource languages. We report judge scores across the same eight language sequences on the 200-example subset, grouped by resource availability. (a) High-resource sequences remain strong for frontier models, with Gemini3-Flash averaging 93.0 and taking both Central Europe (94.2) and Near East (93.5), while GLM-5 (Thinking) narrowly leads East Asia (91.5). (b) Medium-resource sequences remain comparatively stable for the strongest models, with Gemini-3-Flash averaging 91.6. (c) Low-resource sequences show a substantial drop, with Gemini-3-Flash still clearly first at 79.4 average, while the next-best model, Qwen3.5-397B (Thinking), falls to 64.4. (d) Global averages across these eight sequences preserve the same ranking pattern, with Gemini-3Flash leading at 89.1 and holding a 6.0-point lead over the next competitor. (c) Low Resource / Imbalanced Sequences
(a) High & Med-High Resource Sequences Model
E. Asia C. Europe N. East Avg
Gemini-3-Flash (No-Think) Gemma-4-31B (Thinking) Qwen3.5-397B (Thinking) Qwen3.5-397B (Instruct) GLM-5 (Thinking) Gemma-4-31B (Instruct) GLM-4.7 (Thinking) Kimi-K2 (Thinking) DeepSeek-V3.2-Exp GLM-5 (Instruct) Qwen3-235B (Thinking) Qwen3.5-35B (Thinking) DeepSeek-V3.2-Exp (Think) Gemma-3-27B (Instruct) Kimi-K2 Qwen3.5-35B (Instruct) MiniMax-M2.5 GPT-OSS-120B (High) Qwen3-235B (Instruct) MiMo-V2-Flash Qwen3-30B (Instruct) Nemotron-3-Nano
91.2 91.1 91.0 89.7 91.5 90.5 90.2 88.0 88.6 88.6 87.4 83.8 87.5 82.7 82.4 83.6 77.4 81.3 86.3 77.6 79.1 66.4
94.2 91.3 91.5 91.9 91.3 90.9 90.3 88.6 89.0 89.6 89.3 88.3 83.0 87.2 86.9 85.6 84.5 79.5 70.2 81.7 76.4 48.5
93.5 90.3 90.0 90.1 88.7 89.9 87.3 87.8 85.3 83.9 85.3 82.8 84.1 84.4 78.9 77.0 79.0 76.1 78.6 70.1 65.2 29.9
93.0 90.9 90.8 90.6 90.5 90.4 89.3 88.1 87.6 87.3 87.3 85.0 84.8 84.8 82.8 82.1 80.3 79.0 78.4 76.5 73.6 48.3
Model
Africa S. America Avg
Gemini-3-Flash (No-Think) Qwen3.5-397B (Thinking) Gemma-4-31B (Instruct) GLM-5 (Thinking) Qwen3.5-397B (Instruct) Gemma-4-31B (Thinking) DeepSeek-V3.2-Exp (Think) GLM-4.7 (Thinking) GLM-5 (Instruct) DeepSeek-V3.2-Exp Kimi-K2 (Thinking) Qwen3.5-35B (Thinking) GPT-OSS-120B (High) MiniMax-M2.5 Gemma-3-27B (Instruct) Qwen3.5-35B (Instruct) Qwen3-235B (Thinking) Kimi-K2 MiMo-V2-Flash Nemotron-3-Nano Qwen3-235B (Instruct) Qwen3-30B (Instruct)
86.2 72.4 79.4 66.4 66.2 74.2 57.7 52.1 48.4 46.6 46.8 47.1 36.7 33.4 46.3 33.2 16.3 19.8 7.9 5.7 7.2 0.0
Model Gemini-3-Flash (No-Think) Qwen3.5-397B (Thinking) Gemma-4-31B (Instruct) GLM-5 (Thinking) Qwen3.5-397B (Instruct) Gemma-4-31B (Thinking) GLM-4.7 (Thinking) Kimi-K2 (Thinking) Qwen3-235B (Thinking) DeepSeek-V3.2-Exp GLM-5 (Instruct) DeepSeek-V3.2-Exp (Think) Gemma-3-27B (Instruct) Qwen3.5-35B (Thinking) Kimi-K2 GPT-OSS-120B (High) Qwen3.5-35B (Instruct) Qwen3-235B (Instruct) MiniMax-M2.5 MiMo-V2-Flash Qwen3-30B (Instruct) Nemotron-3-Nano
SE. Asia S. Asia N. Europe Avg 90.3 86.9 89.9 86.7 88.1 86.0 84.4 82.7 83.3 84.3 80.4 81.4 81.1 78.5 77.3 71.8 76.2 79.1 72.9 72.6 66.5 33.2
91.7 86.7 89.4 84.7 84.2 90.2 83.0 78.6 80.3 77.3 80.3 74.1 76.6 74.3 71.6 72.8 65.8 73.6 58.0 63.7 50.0 15.9
92.8 89.8 82.7 88.2 86.9 81.7 83.7 84.3 78.5 80.1 78.4 79.4 77.2 79.6 78.0 72.8 71.6 53.8 74.2 65.8 41.3 17.0
91.6 87.8 87.3 86.5 86.4 86.0 83.7 81.9 80.7 80.6 79.7 78.3 78.3 77.5 75.6 72.5 71.2 68.8 68.4 67.4 52.6 22.0
27
79.4 64.4 56.1 54.1 53.6 53.1 48.7 38.0 36.0 33.5 32.6 31.0 27.8 26.4 24.5 20.0 14.2 12.4 8.7 8.3 4.7 0.1
(d) Overall Performance Model
(b) Medium Resource Sequences
72.6 56.3 32.7 41.9 41.0 31.9 39.7 23.9 23.6 20.4 18.5 14.8 18.8 19.4 2.6 6.8 12.0 5.0 9.4 10.8 2.2 0.1
Gemini-3-Flash (No-Think) Qwen3.5-397B (Thinking) Gemma-4-31B (Instruct) GLM-5 (Thinking) Qwen3.5-397B (Instruct) Gemma-4-31B (Thinking) GLM-4.7 (Thinking) DeepSeek-V3.2-Exp (Think) Kimi-K2 (Thinking) GLM-5 (Instruct) DeepSeek-V3.2-Exp Qwen3.5-35B (Thinking) Gemma-3-27B (Instruct) Qwen3-235B (Thinking) GPT-OSS-120B (High) Kimi-K2 Qwen3.5-35B (Instruct) MiniMax-M2.5 Qwen3-235B (Instruct) MiMo-V2-Flash Qwen3-30B (Instruct) Nemotron-3-Nano
Global Average 89.1 83.1 80.7 79.9 79.8 79.6 74.3 73.4 71.9 71.6 71.4 68.7 67.3 66.5 63.7 62.5 62.5 62.3 56.4 56.1 47.3 28.4
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
C
Robustness of Answer Language Analysis Across Models
In this section, we provide results on the languages in which models answer and reason, extending the analysis in Section 3.2 across models. We show results for MT-AIME24 in Figures 4 and 5 and Figures 6 and 7 for Include to separate two failure modes that can confound multilingual benchmark evaluation. First, some models do not consistently answer in the language of the prompt. For example, Qwen3-32B answers in Swahili only about half of the time when prompted in Swahili. Second, an even stronger effect appears in the reasoning traces: most models reason predominantly in English, with only a few partial exceptions, such as Chinese and Russian for some Qwen3 models. Taken together, these results further support our claim that multilingual reasoning and general-knowledge benchmarks often measure English-centered reasoning ability more than genuine multilingual capability.
28
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Same Language
English
Same Language
80
80
Percentage (%)
100
Percentage (%)
100
60 40 20 0
bn
de
en
es
fr
ja
ru
Question Language
sw
te
th
(a) Answer language: the language the model Qwen-3-32B answers in. Same Language
60 40 20 0
zh
Percentage (%)
Percentage (%)
80
60 40 20 en
es
fr
ja
ru
Question Language
sw
te
th
English
Percentage (%)
Percentage (%)
40 20 fr
ja
ru
Question Language
sw
te
th
English
Percentage (%)
Percentage (%)
80
de
en
es
fr
ja
ru
Question Language
sw
te
th
es
fr
ja
ru
Question Language
sw
te
th
zh
English
bn
de
en
es
fr
ja
ru
Question Language
sw
te
th
zh
(f) Reasoning language: the language the model GPT-OSS-20B reasons in.
80
bn
en
Same Language
20
zh
20
100
0
de
Other
40
th
40
100
60
te
60
0
zh
(e) Answer language: the language the model GPT-OSS-20B answers in. Same Language
sw
English
Same Language
60
es
bn
Other 80
en
ru
(d) Reasoning language: the language the model Qwen-3-235B-A22B-Thinking-2507 reasons in.
80
de
ja
20
100
bn
fr
Question Language
40
100
0
es
60
0
zh
(c) Answer language: the language the model Qwen-3-235B-A22B-Thinking-2507 answers in. Same Language
en
Same Language
80
de
de
English 100
bn
bn
(b) Reasoning language: the language the model Qwen-3-32B reasons in.
100
0
English
60 40 20 0
zh
(g) Answer language: the language the model GPT-OSS-120B answers in.
English
bn
de
en
es
fr
ja
ru
Question Language
sw
te
th
zh
(h) Reasoning language: the language the model GPT-OSS-120B reasons in.
Figure 4: Qwen-3 and GPT-OSS models default to English reasoning on MT-AIME24. We analyze the language used for answers (left column) and reasoning traces (right column) across 11 languages. (a-b) Qwen-3-32B answers in the target language inconsistently but reasons almost entirely in English. (c-d) Qwen-3-235B-Thinking answers consistently in the target language (100%) but still reasons predominantly in English, especially for Swahili. (e-f) GPT-OSS-20B shows mixed answering behavior but reasons 93–100% in English. (g-h) GPT-OSS-120B answers mostly in the target language but reasons almost entirely in English. These patterns confirm that mathematical reasoning occurs in English regardless of input language.
29
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Same Language
English
Same Language
80
80
Percentage (%)
100
Percentage (%)
100
60 40 20 0
bn
de
en
es
fr
ja
ru
Question Language
sw
te
th
(a) Answer language: the language the model GLM 4.7 Z.ai et al. (2025) answers in. Same Language
English
60 40 20 0
zh
Percentage (%)
Percentage (%)
80
40 20 de
en
es
fr
ja
ru
Question Language
sw
te
th
en
es
fr
ja
ru
Question Language
Same Language
80
bn
de
Other 100
0
bn
sw
te
th
zh
(b) Reasoning language: the language the model GLM 4.7 Z.ai et al. (2025) reasons in.
100
60
English
(c) Answer language: the language the model Mimo-V2-Flash answers in.
Other
60 40 20 0
zh
English
bn
de
en
es
fr
ja
ru
Question Language
sw
te
th
zh
(d) Reasoning language: the language the model Mimo-V2-Flash reasons in.
Figure 5: GLM-4.7 and MiMo-V2-Flash show contrasting reasoning language patterns on MT-AIME24. (a-b) GLM-4.7 answers in mixed languages across different inputs but reasons overwhelmingly in English (93–100% for most languages), with slight exceptions for Swahili and Chinese. (c-d) MiMo-V2-Flash displays the most diverse reasoning behavior: it reasons natively 33–100% of the time depending on the language, making it an outlier among tested models. However, this native reasoning does not translate to better benchmark performance, further suggesting that MT-AIME24 measures reasoning ability rather than multilingual proficiency.
30
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Same Language
English
80
80
Percentage (%)
100
Percentage (%)
100
60 40 20 0
bn
de
es
fr
ja
ru
Question Language
te
(a) Answer language: the language the 8Qwen-3-32B answers in.
English
de
ja
40 20 bn
Same Language 80
Percentage (%)
80
Percentage (%)
100
40 20 0
bn
de
es
fr
ja
ru
Question Language
te
Same Language
English
Percentage (%)
80
Percentage (%)
80
0
bn
de
es
fr
ja
ru
Question Language
te
Same Language
Percentage (%)
Percentage (%)
80
20 de
es
fr
ja
ru
Question Language
te
zh
(g) Answer language: the language the model GPT-OSS-120B answers in.
es
ru
fr
ja
Question Language
te
zh
Same Language
English
bn
de
es
ru
fr
ja
Question Language
te
zh
(f) Reasoning language: the language the model GPT-OSS-20B reasons in.
80
bn
de
20
100
0
bn
English
40
English
40
100
60
Same Language
60
0
zh
(e) Answer language: the language the model GPT-OSS-20B answers in.
zh
(d) Reasoning language: the language the model Qwen-3-235B-A22B-Thinking-2507 reasons in. 100
20
te
20
Other
40
ru
40
100
60
fr
Question Language
60
0
zh
(c) Answer language: the language the model Qwen-3-235B-A22B-Thinking-2507 answers in.
es
(b) Reasoning language: the language the model Qwen-3-32B reasons in.
100
60
Other
60
0
zh
Same Language
Same Language
English
es
ru
60 40 20 0
bn
de
fr
ja
Question Language
te
zh
(h) Reasoning language: the language the model GPT-OSS-120B reasons in.
Figure 6: Qwen-3 and GPT-OSS models reason almost entirely in English on INCLUDE despite answering in target languages. (a-b) Qwen-3-32B answers consistently in the target language but reasons in English 89–100% of the time. (c-d) Qwen-3-235B-Thinking achieves 100% target-language answering and 100% target-language reasoning—a unique pattern among tested models. (e-f) GPT-OSS-20B answers mostly in the target language (89–98%) but reasons in English 92–100% of the time. (g-h) GPT-OSS-120B shows similar patterns with 91–99% English reasoning. The disconnect between answer language and reasoning language explains why INCLUDE performance tracks English knowledge benchmarks. 31
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Same Language
English
Other
80
80
Percentage (%)
100
Percentage (%)
100
60 40 20 0
bn
de
es
fr
ja
ru
Question Language
te
(a) Answer language: the language the model GLM 4.7 Z.ai et al. (2025) answers in. Same Language
English
Percentage (%)
80
Percentage (%)
80
0
bn
de
es
fr
ja
ru
Question Language
bn
te
de
fr
ja
Question Language
te
zh
(b) Reasoning language: the language the model GLM 4.7 Z.ai et al. (2025) reasons in. 100
20
ru
20
100
40
es
40
Other
60
English
60
0
zh
Same Language
(c) Answer language: the language the model Mimo-V2-Flash answers in.
English
de
ja
Other
60 40 20 0
zh
Same Language
bn
es
fr
ru
Question Language
te
zh
(d) Reasoning language: the language the model Mimo-V2-Flash reasons in.
Figure 7: GLM-4.7 reasons in English while MiMo-V2-Flash shows mixed reasoning patterns on INCLUDE. (a-b) GLM-4.7 answers in the target language 87–100% of the time and reasons in English 92–100% of the time across languages. Russian (87%) and Telugu (12%) show the most answer-language variation. (c-d) MiMo-V2-Flash displays highly variable reasoning behavior: it reasons natively 33–100% of the time depending on the language, with particularly high native reasoning for Telugu (100%), Bengali (60%), and Spanish (84%). This variability makes MiMo-V2-Flash an interesting case study for understanding how reasoning language affects downstream task performance.
32
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
D
Robustness of Error Distribution Analysis Across Models arithmetic logic
100
N=6 N=6 N=6 N=7 N=9 N=9 N=8 N=11 N=10 N=5 N=10
Error distribution (%)
Error distribution (%)
100
arithmetic logic
semantic
80 60 40 20 0
de
en
es
fr
ja
ru
Language
sw
te
th
20
60 40 20 de
en
es
fr
ja
de
en
es
ru
Language
sw
fr
ja
ru
Language
arithmetic logic
N=5 N=4 N=3 N=3 N=4 N=6 N=4 N=6 N=11 N=5 N=3
bn
bn
logic semantic
80
0
40
te
th
sw
te
th
zh
(b) Error distribution for Qwen3-235B-A22BThinking-2507
100
Error distribution (%)
Error distribution (%)
100
60
zh
(a) Error distribution for Qwen3-32B arithmetic formatting
N=7 N=3 N=4 N=2 N=5 N=4 N=5 N=4 N=9 N=3 N=5
80
0 bn
semantic
(c) Error distribution for GPT-OSS-120B
N=5 N=4 N=2 N=4 N=4 N=4 N=5 N=4 N=10 N=4 N=5
80 60 40 20 0
zh
semantic
bn
de
en
es
fr
ja
ru
Language
sw
te
th
zh
(d) Error distribution for GLM 4.7
Figure 8: Error analysis across four additional models confirms MT-AIME24 errors are logical, not linguistic. We categorize errors for four models on MT-AIME24 across 11 languages. (a-d) Across all models, errors are predominantly logical (blue) or arithmetic (green) rather than semantic (red). The semantic error rate rarely exceeds 25% for any language-model combination. For several languages (e.g., French in panels a and c), 100% of errors stem from reasoning failures. This consistent pattern across diverse model architectures confirms that MT-AIME24 does not effectively measure multilingual comprehension. We provide additional error distribution analyses for MT-AIME24 in Figure 8 and Include in Figure 9 across additional model families. We observe that MT-AIME24 errors are consistently dominated by logical and arithmetic failures rather than semantic misunderstanding, indicating that the benchmark primarily tests mathematical reasoning rather than multilingual comprehension. Figure 9 shows a parallel pattern for Include: errors are overwhelmingly factual, with smaller contributions from regional knowledge gaps and hallucinations, while semantic errors due to multilingual misunderstanding remain rare. Together, these results strengthen our conclusion that current multilingual benchmarks mainly reflect reasoning and factual recall, not genuine cross-lingual understanding.
33
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
factual hallucination logic
N=189 N=49 N=125 N=85
N=49 N=153 N=203 N=65
80 60 40 20 0
100
Error distribution (%)
Error distribution (%)
100
regional_knowledge semantic
bn
de
es
fr
ja
Language
ru
te
N=162 N=46 N=114 N=72
N=42 N=164 N=199 N=139
60 40 20 bn
de
es
N=65
N=31 N=123 N=153 N=57
bn
es
fr
ja
40 20 de
regional_knowledge semantic
80
0
N=91
fr
ja
Language
Language
ru
te
zh
(b) Qwen3-235B-A22B-Thinking
100
Error distribution (%)
Error distribution (%)
100
N=113 N=41
60
(a) Qwen3-32B factual logic
regional_knowledge semantic
80
0
zh
factual hallucination logic
ru
te
(c) GPT-OSS-120B
regional_knowledge semantic
N=124 N=43
N=78
N=54
N=32 N=124 N=137 N=51
bn
es
fr
ja
80 60 40 20 0
zh
factual hallucination logic
de
Language
ru
te
zh
(d) GLM 4.7
Figure 9: Error analysis on INCLUDE confirms errors are factual and knowledge-based, not linguistic. We categorize errors for four models across eight languages on INCLUDE. (a-d) Across all models, the dominant error type is factual (blue), accounting for 81–98% of errors depending on language and model. Semantic errors (red) rarely exceed 10% for any configuration. Regional knowledge gaps (orange) contribute modestly (6–14% in some cases). Hallucinations appear occasionally but are not the primary failure mode. This consistent pattern confirms that INCLUDE measures factual knowledge coverage rather than multilingual comprehension ability.
34
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
E
Extended Experimental Details
E.1
Dataset Composition
We provide more details on the composition of our proposed LiT benchmark: (a) Abstracts (20% samples): We divide this category equally between Humanities (10%) and STEM (10%) abstracts. Abstracts are useful benchmark units because they are selfcontained, semantically dense, and complete, so errors are easier to detect (Isabelle et al., 2017; Kleidermacher & Zou, 2025). Humanities abstracts test whether models can preserve argumentative nuance and rhetorical force, while STEM abstracts test terminological precision and mathematical notation, where minor mistakes carry major consequences. (b) Pragmatics (60% samples): This category tests whether models preserve meaning beyond literal content (Park et al., 2024). We subdivide it into five phenomena: (i) Core Semantics (20% samples): tests preservation of truth conditions and logical entailment (Dagan et al., 2013). (ii) Discourse Coherence (17.5% samples): evaluates maintenance of referential chains, topic continuity, and logical connectives across sentence boundaries Barzilay & Lapata (2005). (iii) Implicit Content (17.5% samples): probes handling of presuppositions, implicatures, and information that speakers convey without explicitly stating (Grice, 1975). (iv) Pragmatic Inference (21.7% samples): tests understanding of speech acts, speaker intent, and context-dependent meaning (Searle, 1969). (v) Social Interaction (23.3% samples): evaluates preservation of politeness markers, formality levels, and sociolinguistic appropriateness (Brown & Levinson, 1987). (c) Informal (20% samples): This category tests colloquial language, slang, idioms, and register shifts (Koehn & Knowles, 2017; Fadaee et al., 2018). Informal text requires preserving tone and social function, not just denotative meaning. E.2
Sampling Hyperparameters Table 10: Hyperparameter details for round-trip translation. Model Gemini-3-Flash (No-Thinking) GLM-4.7 GLM-5 (Thinking) GLM-5 (Instruct) Qwen3-235B (Thinking) Qwen3-235B (Instruct) Qwen3-30B (Instruct) Qwen3.5-397B (Thinking) Qwen3.5-397B (Instruct) Qwen3.5-35B (Thinking) Qwen3.5-35B (Instruct) Kimi-K2 (Thinking) Kimi-K2 (Instruct) DeepSeek-V3.2 (Thinking) DeepSeek-V3.2 (Instruct) GPT-OSS-120B MiniMax-M2.5 MiMo-V2-Flash (Instruct) Nemotron-3-Nano Gemma-3-27B (Instruct) Gemma-4-31B (Instruct) Gemma-4-31B (Thinking)
35
Temp.
Top-p
Effort
1.0 1.0 1.0 1.0 0.6 0.7 0.7 0.6 0.7 0.6 0.7 1.0 1.0 1.0 1.0 1.0 1.0 1.0 0.6 1.0 1.0 1.0
0.95 0.95 0.95 0.95 0.95 0.8 0.8 0.95 0.8 0.95 0.8 1.0 1.0 0.95 0.95 1.0 0.95 0.95 0.95 0.95 0.95 0.95
– – – – – – – – – – – – – – – High – – – – – –
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Table 11: Hyperparameter details for MT-AIME24 and INCLUDE Model Gemini-3-Flash GLM-4.7 Qwen3-235B (Thinking) Kimi-K2 (Thinking) DeepSeek-V3.2-Exp (Think) DeepSeek-V3.2-Exp GPT-OSS-120B MiMo-V2-Flash Nemotron-3-Nano Gemma-3-27B (Instruct) Qwen3-235B (Instruct) Qwen3-30B (Instruct)
Temp. Top-p 1.0 1.0 0.6 1.0 1.0 1.0 1.0 1.0 0.6 1.0 0.7 0.7
0.95 0.95 0.95 1.0 0.95 0.95 1.0 0.95 0.95 0.95 0.8 0.8
Max Tokens Effort AIME24 Include 38,912 38,912 38,912 131,072 38,912 38,912 38,912 38,912 38,912 38,912 38,912 38,912
32,768 32,768 32,768 32,768 32,768 32,768 32,768 32,768 32,768 32,768 32,768 32,768
– – – – – – High – – – – –
The hyperparameters used for the model sampling utilizing Openrouter are presented in Tables 10 for the LiT benchmark and Table 11 for MT-AIME24 and Include. We follow the official technical reports for the corresponding models and use the official sampling parameters unless specified otherwise. If no suggested default is officially given, we set the temperature to a default of 1.0 and Top-p sampling to 0.95. E.3
Model Prompt
We provide the translation prompt used in our evaluation, for reproducibility of our pipeline: System Prompt You are a professional translator. Task: Translate the SOURCE TEXT from {src lang} into {target lang}. Instructions: • Use natural, idiomatic {target lang}, avoid unnatural word-for-word translation. • Preserve meaning, tone, and register. Do not add, omit, or summarize. • Output the translated text in {target lang}. Do not include any additional text. SOURCE TEXT: "{text}"
E.4
Judge Prompt
We provide the judge prompt used in our evaluation, for reproducibility of our pipeline:
36
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
System Prompt You are an annotator and expert linguist. Your task is to evaluate the back-translation given the original text and output a score, classification of machine translation quality, and a list of issues found. Score measures back-translation quality on a continuous scale from 0 to 100, where a score of zero means "little-to-no meaning preserved" and score of one hundred means "perfect meaning preserved". The classification of the quality of machine translation must be into one of following 10 categories: • ’1-nonsense’: The text is gibberish, in the wrong language, or completely unrelated to the source. • ’2-severe distortion’: Contains unrecognizable fragments. The core meaning is often lost or dangerously misleading. • ’3-failed gist’: The topic is broadly correct, but the translation is mostly misleading or incomprehensible due to severe errors. • ’4-unreliable’: Meaning is often preserved but significant meaning errors (critical mistranslations) are present. • ’5-machine-like’: The meaning is roughly preserved (no critical errors), but the phrasing is overly literal ("translationese"), awkward, or grammatically poor. • ’6-understandable but Flawed’: Meaning is preserved. Grammar is mostly functional but contains distracting errors or very unnatural stiffness. • ’7-good’: Accurate meaning. Grammatically correct with only minor, non-impeding errors (e.g., wrong punctuation, slight awkwardness). • ’8-very good’: Fluent and accurate. No grammatical errors. However, it may miss minor nuances of tone or style found in the Reference. • ’9-excellent’: Native-level fluency. Captures the exact meaning and tone. Indistinguishable from professional human translation. • ’10-perfect’: Flawless. Captures distinct cultural nuances, idioms, and subtext perfectly. Equivalent to the reference. The list of issues is a comprehensive list of errors based on the following MQM Core dimensions: Accuracy, Fluency, and Terminology and Style/Locale. 1. Accuracy: (Mistranslation, Omission, Addition, Untranslated). Does the target text accurately reflect the source meaning? 2. Fluency: (Grammar, Spelling, Punctuation, Unintelligible). Is the target text linguistically correct and natural? 3. Terminology: (Inconsistent, Wrong Term). Does it adhere to domain standards? 4. Style/Locale: Does it follow local formats (dates, currencies) and cultural norms? Does the translation match the required formality/register (e.g., formal vs. casual)? The three severity categories are: • ’minor’: Has a limited impact on accuracy, stylistic quality, consistency, fluency, clarity, or general appeal of the content. • ’major’: Seriously affects the understandability, reliability, or usability of the content for its intended purpose. For example, it causes significant loss or change in meaning or because the error appears in a highly visible or important part of text. • ’critical’: Hallucination, completely changes meaning or catastrophic failure rendering the sentence unusable or poses serious reputational harm. Issues is a list of tuples of each issue being a tuple of (severity category, issue description). Output Format: You must output a single valid JSON object. Do not include markdown formatting (like ‘‘‘json) or conversational text. The JSON must follow this schema: {"score": <0-100>, "classification": "<one of the 10 quality categories>", "issues": [{"severity":"< one of three severity categories>", "description":"<issue description>"} , ...]} Original text: "{original}" Back-translation: "{translation}"
37