Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora Maciej Skórski
arXiv:2605.22660v1 [cs.CL] 21 May 2026
[email protected] University of Luxembourg Luxembourg
Abstract
Social media post
Moral foundation
Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages. Idiomatic expressions, slang, and cultural references introduce hard-to-avoid translation artifacts. Yet automated moral values classification depends on languagespecific annotated corpora that exist almost exclusively in English. We investigate whether LLM-based translation can bridge this gap, taking Polish as a test case. Using ∼50k morally-annotated social media posts from a diverse range of topics, we apply a principled four-method validation pipeline: LaBSE cross-lingual embedding similarity, Centered Kernel Alignment (CKA), LLM-as-judge evaluation, and deep learning classifier parity tests. We show that despite shortcomings in handling slang, vulgarity, and culturally-loaded expressions, direct translation preserves subtle moral cues well enough to be harvested by cross-lingual machine learning — with mean cosine similarity of 0.86 and AUC gaps of 0.01–0.02 across all foundations closing further under fine-tuning of language models. These results demonstrate that machine translation is a practical and cost-effective path to moral values research in languages currently under-resourced in this domain. We demonstrate this for Polish as a representative Slavic language, with expected generalisation to related languages.
“My heart breaks seeing children separated from families at the border” “Everyone deserves equal access to healthcare regardless of income” “Respect your elders and follow traditional values that built this nation” “Stand with our troops — they sacrifice everything for our freedom” “Marriage is sacred and should be protected from secular corruption”
Care
CCS Concepts • Computing methodologies → Natural language processing; Machine translation; • Applied computing → Psychology.
Fairness Authority Loyalty Sanctity
Table 1: Examples of moral foundations in text [8]. Key moral cues highlighted in foundation color. Yet these methods require language-specific annotated corpora for training, and such resources remain almost exclusively in English [9, 21]. No morally-annotated corpus has been released for any Slavic language so far — leaving the moral discourse of hundreds of millions of speakers beyond the reach of automated MFT analysis. Machine translation offers a natural shortcut: translate existing English corpora and extend MFT tools to new languages at low cost. But moral language is precisely what MT handles worst — laden with irony, cultural idiom, and register sensitivity, it resists the literal mappings that MT systems rely on. Recent findings validate this concern: MT injects systematic bias into cross-lingual text analysis [15], distorting even well-established affective signals [16]. This raises a pointed question for moral NLP:
Keywords Moral Foundations Theory, cross-lingual NLP, machine translation
1
Introduction
Moral language is subtle. Irony inverts it. Cultural idiom obscures it. Register shifts dilute it. Studying it at scale, across languages and cultures, requires a principled framework. Moral Foundations Theory (MFT) [6, 8] provides exactly that: a cross-cultural taxonomy of five moral dimensions — care/harm, fairness/cheating, loyalty/betrayal, authority/subversion, and sanctity/degradation — grounded in decades of cross-cultural moral psychology research. While these foundations are universal, cultures differ markedly in their sensitivity to each dimension and in how they express it in language [6]. Automated methods have made it possible to measure these cultural-linguistic sensitivities at scale. Lexiconbased approaches [10] and, more recently, fine-tuned language models [13, 17, 19, 22] have been applied to political discourse [18], social media analysis [9], and moral dilemmas [14] — building on foundational cross-cultural findings in moral psychology [4, 7].
Does LLM-based EN→PL translation preserve the moral-semantic content of MFT-annotated texts, despite the subtlety of moral language?
We answer this affirmatively, contributing: • Principled validation framework. A reproducible multi-method pipeline combining LLM-as-judge evaluation, embedding-based similarity (LaBSE, CKA), and deep learning classifier parity tests — applicable to any source language, target language, and annotation schema. • Large translated corpus. A validated EN→PL translation of ∼50k morally-annotated social media posts spanning a diverse range of topics and platforms, produced using Claude Sonnet at a cost of approximately 200 USD — demonstrating accessibility. • Evidence that it works — and generalises. Translation quality is good but not perfect (mean cosine 0.86), yet classification remains near-parity with English originals across all five MFT foundations — AUC gaps of 0.01–0.02, nearly closed by fine-tuning. As one of the most morphologically complex Slavic
Maciej Skórski
languages [11], Polish is a demanding test case; success here suggests generalisation across the broader Slavic family.
õ Source corpora = MFRC
MFTC
2 Background and Related Work 2.1 Moral Foundations Theory Corpora MFRC (Moral Foundations Reddit Corpus) [21] contains posts drawn from Reddit communities annotated with MFT foundation labels across three subcorpora: everyday morality (r/AmItheAsshole), US politics, and French politics. The corpus covers moral reasoning expressed through colloquial language, abbreviations (NTA, YTA, AITA), and informal style — making it a challenging but realistic benchmark for translation. MFTC (Moral Foundations Twitter Corpus) [9] comprises tweets from politically and socially charged events, annotated by crowd workers across seven subcorpora. The Davidson variant used here is drawn from a hate speech dataset [2], representing a harder distribution with activist hashtags (#BLM, #MeToo), AAVE expressions, and extreme vulgarity — a demanding test-case for any translation system.
P1
Z ~200 samples per subcorpus
Draft translation claude-sonnet
À LLM-as-judge score 0–10 per row
º refine if unsatisfactory
$ final prompt
P2
2.2
Full corpus all subcorpora
Ô Full translation EN + PL aligned pairs
Cross-Lingual Transfer and Translation
Cross-lingual models such as mBERT [3] and XLM-RoBERTa [1] enable zero-shot transfer to new languages without additional annotation, and LaBSE [5] provides strong cross-lingual sentence embeddings well-suited for semantic equivalence assessment. However, without fine-tuning, LLMs introduce systematic biases that distort cross-lingual text analysis and affective signals [15, 16] — motivating fine-tuning on translated corpora as a more reliable path. Prior cross-lingual moral analysis has relied on multilingual dictionaries to map moral seed words across languages [10], without exploiting machine translation at scale; our work fills this gap.
3 Methods 3.1 Corpora Two corpora were selected to cover a diverse range of moral discourse styles and difficulty levels (Table 2). MFRC (Moral Foundations Reddit Corpus) [21] provides 17,886 Reddit posts labeled across five MFT foundations (authority, care, fairness, loyalty, sanctity) spanning three subcorpora: everyday morality (r/AmItheAsshole), US politics, and French politics. The corpus covers moral reasoning expressed through colloquial language, abbreviations (NTA, YTA, AITA), and informal style. MFTC (Moral Foundations Twitter Corpus) [9] provides 33,858 Twitter posts across seven subcorpora — spanning political movements (#BLM, #MeToo, All Lives Matter), civil unrest (Baltimore uprising), electoral discourse, and a natural disaster (Hurricane Sandy). The Davidson subcorpus, drawn from a hate speech dataset [2], represents a hard test case with vulgarity and AAVE-heavy language. These data cover a diverse range of moral discourse: everyday interpersonal judgments, heated political debate, social movements, and collective responses to crisis — making them a demanding and representative benchmark for cross-lingual translation.
Ì Cosine LaBSE
½ CKA alignment
j Classifier AUC EN/PL
P3
¦ Translation quality verdict
Figure 1: Validation pipeline. The full corpus and ~200 samples per subcorpus are used in parallel during Phase 1: the sample drives iterative prompt refinement via an LLM judge ( À ), while the full corpus awaits the final prompt. Phase 2 produces aligned EN/PL pairs. Phase 3 evaluates via pairwise cosine similarity, CKA, and fine-tuning performance.
3.2
Translation Pipeline
Translation was performed using Claude-Sonnet-4-6 via the Anthropic API, with 20 concurrent asynchronous requests to maximise throughput. Platform-specific prompting was applied ( Section 3.3): the Reddit prompt instructs the model to preserve informal tone, Reddit abbreviations (NTA, YTA), and formatting; the Twitter prompt additionally preserves hashtags and @mentions unchanged. The full translation cost approximately 200 USD for 50k posts combined, demonstrating the accessibility of this approach for research groups without large annotation budgets.
3.3
Translation Prompts
Two platform-specific system prompts were carefully engineered to handle the distinct linguistic styles of each corpus (Prompts P1–P2). Both share a common design philosophy: preserve moral-semantic content while naturalizing tone into idiomatic Polish. Key design
Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora
Table 2: Corpora, subcorpora, and per-foundation prevalence (% of total texts). Au=authority, Ca=care, Fa=fairness, Lo=loyalty, Sa=sanctity. Totals include non-moral instances. Corpus
MFRC
Subcorpus
Platform
Domain
Au%
Ca%
Fa%
Lo%
Sa%
𝑁
Everyday morality US politics French politics
Reddit Reddit Reddit
General moral discourse US political discourse French political discourse
10.4 19.7 25.4
37.4 29.6 16.0
25.5 38.3 26.0
11.7 7.7 13.1
13.5 8.4 8.0
5,366 5,351 7,169 17,886
MFRC total
MFTC
ALM BLM Baltimore Davidson Election MeToo Sandy
Twitter Twitter Twitter Twitter Twitter Twitter Twitter
All Lives Matter Black Lives Matter Baltimore uprising Hate speech US election MeToo movement Hurricane Sandy
20.9 10.2 31.7 3.3 5.8 65.7 45.6
6.2 27.3 27.1 11.5 12.5 33.3 60.9
7.4 23.9 31.4 10.0 11.3 43.9 30.3
12.5 13.1 42.3 1.2 7.5 41.0 42.7
MFTC total decisions include explicit slang mappings (e.g. wtf →kurwa), strict hashtag and mention preservation, and grammar rules for name declension. The prompts differ in their handling of Reddit-specific conventions (NTA/YTA abbreviations, markdown formatting, nested quotes) versus Twitter-specific ones (ALL CAPS emphasis, retweet abbreviations, activist hashtags). Each prompt includes a one-shot example to anchor the target style. Both prompts were iteratively refined using a stratified sample of ~200 posts per subcorpus, evaluated by an LLM judge on tone preservation, slang handling, formatting fidelity, and proper noun treatment — as detailed in Section 3.4.
Prompt P1 — MFRC (Reddit) You are a research assistant tasked with translation of Reddit posts from English to Polish. Rules:
• Preserve tone exactly — slang, vulgarity, sarcasm, humor, aggression • Translate slang and profanity into natural Polish equivalents when possible (wtf→kurwa, asshole→dupek, bro→stary/ziom); keep English only when no
8.6 3.9 12.6 12.3 5.3 17.5 15.0
4,326 5,117 5,190 4,873 5,050 4,711 4,591 33,858
Prompt P2 — MFTC (Twitter) You are a research assistant tasked with translation of English tweets to Polish. Rules:
• Preserve tone exactly — slang, vulgarity, sarcasm, humor, aggression • Translate slang and profanity into natural Polish equivalents when possible (wtf→kurwa, asshole→dupek); keep English only when no good Polish equivalent exists
• Ensure correct Polish grammar, gender agreement and declension — never translate surnames, only decline them
• Keep unchanged: hashtags (#BLM, #MeToo), @mentions, URLs, abbreviations • • • • •
(lol, omg, rt, imo, smh) Keep unchanged: proper nouns, public figures, brand names, places Keep unchanged: emojis, Unicode emoticons, ALL CAPS emphasis Preserve formatting: newlines, ellipsis. . . — do not decode HTML entities Use only Polish/Latin characters, never Cyrillic Return ONLY the translated text, nothing else
One-shot example: IN: This is exactly why #MeToo matters. Men in power think they can get away with anything smh OUT: Właśnie dlatego #MeToo ma znaczenie. Faceci przy władzy myślą, że mogą robić co chcą, no kurwa
good Polish equivalent exists
• Ensure correct Polish grammar, gender agreement and declension • Keep unchanged: abbreviations (NTA, YTA, AITA, OP, RN, GOP), proper
• • • •
nouns, public figures — decline but never translate names or surnames; use Polish equivalents for party names (Democrats→Demokraci, Republicans→Republikanie) Keep unchanged: hashtags, @mentions, URLs, interjections (BABY, WOO HOO) Keep unchanged: Unicode emoticons, ASCII art, kaomojis, emojis (e.g. ) Preserve formatting: > & \n markdown — do not decode Return ONLY the translated text, nothing else
One-shot example: IN: >Why don’t you just leave him lol. NTA, he’s being a massive asshole tbh like wtf bro OUT: >No to czemu po prostu go nie rzucisz lol. NTA, szczerze mówiąc zachowuje się jak totalny dupek no co to kurwa jest stary
3.4
Validation Methods
The validation pipeline proceeds in three phases (Figure 1). Phase 1 uses ~200 samples per subcorpus to iteratively refine translation prompts via an LLM judge. Phase 2 applies the final prompt to all subcorpora. Phase 3 evaluates translation fidelity through four complementary methods.
Embedding Similarity (LaBSE). Cross-lingual cosine similarity between English and Polish sentence pairs was computed using LaBSE [5]. The expected range for well-translated pairs is 0.80–0.95. Centered Kernel Alignment (CKA). CKA [12] measures global alignment between embedding spaces, providing a stronger signal than mean pairwise similarity by capturing structural preservation beyond individual sentence pairs. Model-as-Judge. A stratified sample of ~200 posts per subcorpus was evaluated by Claude Sonnet on four dimensions: tone preservation, slang handling, formatting fidelity, and proper noun treatment. Scores were elicited on a 0–10 scale. Classifier Parity / Gap Validation. A linear classification head was trained on frozen LaBSE embeddings using 10-fold stratified cross-validation. ROC-AUC was compared between English and Polish conditions using a one-sided paired 𝑡-test (𝐻 1 : EN > PL). Note that frozen-embedding AUC is intentionally conservative: full fine-tuning lifts both EN and PL performance jointly, so the reported gaps isolate translation fidelity rather than absolute classifier quality. As a supplementary check, mDeBERTa-v3-base was
Maciej Skórski
fully fine-tuned end-to-end on both English and translated Polish corpora to confirm parity holds under full gradient updates.
4 Results 4.1 Model-as-Judge Quality Translation quality was assessed via row-by-row LLM-as-judge evaluation (𝑁 ≈200 per subcorpus) on a 0–10 scale (Table 3). Across all subcorpora the mean score is 9.1, with 94.6% of posts free of detectable issues. Scores are lowest on AAVE-heavy subcorpora (ALM, Davidson, Baltimore: 8.5), where dialect and embedded lyrics occasionally force paraphrase, and highest on BLM and Sandy (9.5). Two model-level failure modes persist regardless of prompt engineering: sporadic hashtag content translation and Cyrillic character leakage during self-correction. Corpus
Sub-corpus
Clean %
Minor %
Err. %
Score
MFRC
Everyday Morality US Politics French Politics
93.0 95.0 93.0
5.0 3.0 5.0
2.0 2.0 2.0
8.5 9.5 8.5
ALM BLM Baltimore Davidson Election MeToo Sandy
91.0 95.5 95.0 94.0 96.5 96.5 96.5
7.0 3.5 3.5 4.5 2.5 2.5 2.5
2.0 1.0 1.5 1.5 1.0 1.0 1.0
8.5 9.5 8.5 8.5 9.0 9.0 9.5
MFTC
94.6
3.9
1.5
9.1
Average
Table 3: Row-by-row LLM-as-judge translation audit (EN→PL, 𝑁 =200 per sub-corpus). Clean: no issues. Minor: tone softening, inconsistent slang, formatting artefacts. Errors: grammar failures, meaning inversions, untranslated segments, spurious refusals. Score: 0–10 judgment (human validated).
4.2
Embedding Similarity and CKA
LaBSE cross-lingual cosine similarity and linear CKA are reported in Tables 4 and 5. Mean cosine similarity is 0.889 overall (MFRC: 0.876, MFTC: 0.894), well above the random baseline of ≈0.30 and exceeding the 0.80 threshold considered strong semantic equivalence. CKA confirms global embedding alignment (overall 0.860), with French Politics and Baltimore scoring highest (0.895–0.896) and Davidson lowest (0.806), where AAVE paraphrase shifts distributional geometry beyond what pairwise distances capture. A modest gap to 1.0 is expected, attributable to Polish morphological inflection and culture-specific expressions.
4.3
Corpus
Sub-corpus
N
Mean
Std
P05
P95
MFRC
Everyday Morality US Politics French Politics
5,366 5,351 7,169
0.869 0.863 0.896
0.052 0.056 0.039
0.782 0.773 0.831
0.929 0.926 0.946
ALM BLM Baltimore Davidson Election MeToo Sandy
4,326 5,117 5,190 4,873 5,050 4,711 4,591
0.911 0.921 0.911 0.867 0.901 0.903 0.843
0.042 0.047 0.063 0.083 0.056 0.052 0.100
0.834 0.840 0.809 0.712 0.808 0.820 0.713
0.970 0.985 1.000 0.962 0.967 0.979 0.937
MFTC
51,744
0.889
0.063
0.789
0.960
Overall
Table 4: LaBSE cross-lingual cosine similarity between English source and Polish translation per sub-corpus. P05/P95 denote the 5th and 95th percentiles. A threshold of ≥0.80 is widely considered strong semantic equivalence. Corpus
Sub-corpus
MFRC
MFTC
Overall
N
CKA
Everyday Morality US Politics French Politics
5,366 5,351 7,169
0.833 0.826 0.895
ALM BLM Baltimore Davidson Election MeToo Sandy
4,326 5,117 5,190 4,873 5,050 4,711 4,591
0.867 0.883 0.896 0.806 0.877 0.836 0.845
51,744
0.860
Table 5: Linear CKA between LaBSE embeddings of English source and Polish translation per sub-corpus. CKA measures global alignment of embedding spaces; a value of 1.0 indicates identical geometry up to orthogonal transformation. weak moral signal in hate speech, not translation failure — care and fairness even show negative gaps (PL > EN). Frozen-embedding AUC is intentionally conservative: fine-tuning lifts both EN and PL jointly (??), so gaps here measure translation fidelity, not absolute classifier quality.
4.4
Full Fine-Tuning Validation
mDeBERTa-v3 was further validated under full fine-tuning on the MFTC Davidson hate speech corpus (fairness foundation). Despite the weak moral signal characteristic of this corpus, English and translated Polish follow near-identical learning trajectories, converging from ≈0.57 to ≈0.68 ROC-AUC with a final gap of 0.006 — confirming translation parity holds under full fine-tuning.
Classifier Parity
Table 6 reports ROC-AUC per foundation under 10-fold CV with frozen LaBSE embeddings. On MFRC, most gaps are below 0.015 — fairness and sanctity show no significant degradation. Authority is the consistent exception (gaps 0.021–0.031), though still minor. On MFTC, gaps are larger on AAVE-heavy subcorpora (ALM, Election, Sandy: up to 0.048) yet no foundation is invalidated for downstream use. Davidson is a special case: near-chance baseline AUC reflects
5
Discussion
Overall. The convergent evidence from four independent validation methods supports a perhaps surprising conclusion: LLM-based EN→PL translation preserves moral-semantic content at a level sufficient for downstream classification. Moral language — with its irony, idiom, and cultural sensitivity — proves more robust to translation than research on LLM limitations might suggest [15].
Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora
The authority gap. The small but statistically detectable AUC gap on authority (0.014) is the one exception worth examining. We attribute this not to translation error but to genuine cross-cultural divergence in how authority is expressed in Polish discourse — an observation consistent with the cross-cultural MFT literature [6, 20], where authority norms show the highest cross-national variance and the highest sensitivity to domain variation. Future work could verify this by re-annotating a Polish sample with native annotators. Harder corpora. The lower MFTC scores (CKA 0.804, mean judge 8.9/10) relative to MFRC (CKA 0.833, mean judge 9.2/10) are attributable to AAVE-heavy Twitter content, where translation necessarily involves paraphrase. Even so, classifier parity confirms the translated data remains usable for training. Practical accessibility. The pipeline is accessible: ∼50k posts translate for approximately 200 USD, making corpus extension to additional languages feasible without large annotation budgets. Generalisation to other Slavic languages. Polish is among the most morphologically complex languages in the Slavic family, with seven grammatical cases and rich inflectional morphology. Research on neural cross-lingual transfer shows that morphological relatedness within a language family directly facilitates knowledge transfer [11], suggesting that our results for Polish provide a reasonable lower bound for the broader Slavic family. Limitations. Labels are inherited from English annotations without re-annotation in Polish, which precludes measuring crosscultural label shift. The pipeline was validated on Reddit and Twitter discourse styles; domain transfer to news or parliamentary corpora may require re-evaluation, though the diversity of topics covered — everyday moral discourse, political debate, social movements, and natural disasters — suggests reasonable robustness across common domains of moral language use. Finally, prompt engineering was conducted without involvement of native Polish speakers — a potential blind spot for moral discourse patterns specific to Polish cultural context not captured by the one-shot examples.
6
Conclusion
We present a validated pipeline for extending English moral values corpora to Polish via LLM translation. Testing across a diverse range of topics and MFT subcorpora, we find that translation preserves subtle moral cues well enough for cross-lingual machine learning to harvest them — with AUC gaps in the range 0.01–0.02, nearly closed by fine-tuning. Moral semantics survive machine translation, opening a practical and cost-effective path for moral values research in Polish and, by extension, the broader Slavic family.
References [1] Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of ACL. 8440–8451. [2] Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated hate speech detection and the problem of offensive language. In Proceedings of the 11th International AAAI Conference on Web and Social Media. 512–515. [3] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT. 4171–4186. [4] Matthew Feinberg and Robb Willer. 2013. The Moral Roots of Environmental Attitudes. Psychological Science 24, 1 (2013), 56–62.
[5] Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2020. Language-agnostic BERT sentence embedding. arXiv preprint arXiv:2007.01852 (2020). [6] Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean Wojcik, and Peter H Ditto. 2013. Moral foundations theory: The pragmatic validity of moral pluralism. Advances in Experimental Social Psychology 47 (2013), 55–130. [7] Jesse Graham, Jonathan Haidt, and Brian A Nosek. 2009. Liberals and conservatives rely on different sets of moral foundations. Journal of personality and social psychology 96, 5 (2009), 1029–1046. doi:10.1037/a0015141 [8] Jonathan Haidt and Craig Joseph. 2004. Intuitive ethics: How innately prepared intuitions generate culturally variable virtues. Daedalus 133, 4 (2004), 55–66. [9] Joe Hoover, Gwenyth Portillo-Wightman, Leigh Yeh, Shreya Havaldar, Aida Mostafazadeh Davani, Ying Lin, Brendan Kennedy, Mohammad Atari, Zahra Kamel, Madelyn Mendlen, Gabriela Moreno, Christina Park, Tingyee E. Chang, Jenna Chin, Christian Leong, Jun Yen Leung, Arineh Mirinjian, and Morteza Dehghani. 2020. Moral Foundations Twitter Corpus: A Collection of 35k Tweets Annotated for Moral Sentiment. Social Psychological and Personality Science 11, 8 (Nov. 2020), 1057–1071. doi:10.1177/1948550619876629 [10] Frederic R. Hopp, Jacob T. Fisher, Devin Cornell, Richard Huskey, and René Weber. 2021. The Extended Moral Foundations Dictionary (eMFD): Development and Applications of a Crowd-Sourced Approach to Extracting Moral Intuitions from Text. Behavior Research Methods 53, 1 (Feb. 2021), 232–246. doi:10.3758/s13428020-01433-0 [11] Katharina Kann, Ryan Cotterell, and Hinrich Schütze. 2017. One-Shot Neural Cross-Lingual Transfer for Paradigm Completion. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Regina Barzilay and Min-Yen Kan (Eds.). Association for Computational Linguistics, Vancouver, Canada, 1993–2003. doi:10.18653/v1/P17-1182 [12] Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. 2019. Similarity of neural network representations revisited. In Proceedings of the 36th International Conference on Machine Learning. 3519–3529. [13] Tuan Dung Nguyen, Ziyu Chen, Nicholas George Carroll, Alasdair Tran, Colin Klein, and Lexing Xie. 2024. Measuring Moral Dimensions in Social Media with Mformer. Proceedings of the International AAAI Conference on Web and Social Media 18 (May 2024), 1134–1147. doi:10.1609/icwsm.v18i1.31378 [14] Tuan Dung Nguyen, Georgina Lyall, Alasdair Tran, Minkyoung Shin, Nicholas G Carroll, Colin Klein, and Lexing Xie. 2022. Mapping Topics in 100,000 Real-Life Moral Dilemmas. In Proceedings of the International AAAI Conference on Web and Social Media, Vol. 16. 699–710. [15] Gabriel Nicholas and Aliya Bhatia. 2023. Lost in Translation: Large Language Models in Non-English Content Analysis. arXiv e-prints (2023), arXiv–2306. [16] Flor Miriam Plaza-del Arco, Amanda Cercas Curry, Alba Curry, Gavin Abercrombie, and Dirk Hovy. 2024. Angry men, sad women: Large language models reflect gendered stereotypes in emotion attribution. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 7682–7696. [17] Vjosa Preniqi, Iacopo Ghinassi, Julia Ive, Charalampos Saitis, and Kyriaki Kalimeri. 2024. MoralBERT: A Fine-Tuned Language Model for Capturing Moral Values in Social Discussions. In Proceedings of the 2024 International Conference on Information Technology for Social Good. ACM, Bremen Germany, 433–442. doi:10.1145/3677525.3678694 [18] Shamik Roy and Dan Goldwasser. 2021. Analysis of Nuanced Stances and Sentiment Towards Entities of US Politicians through the Lens of Moral Foundation Theory. In Proceedings of the Ninth International Workshop on Natural Language Processing for Social Media. Association for Computational Linguistics, Online, 1–13. doi:10.18653/v1/2021.socialnlp-1.1 [19] Maciej Skorski and Alina Landowska. 2025. Beyond Human Judgment: A Bayesian Evaluation of LLMs’ Moral Values Understanding. In Proceedings of the 2nd Workshop on Uncertainty-Aware NLP (UncertaiNLP 2025), Bryan Eikema, Raúl Vázquez, Jonathan Berant, Marie-Catherine de Marneffe, Barbara Plank, Artem Shelmanov, Swabha Swayamdipta, Jörg Tiedemann, Chrysoula Zerva, and Wilker Aziz (Eds.). Association for Computational Linguistics, Suzhou, China, 17–26. doi:10.18653/v1/2025.uncertainlp-main.3 [20] Maciej Skorski and Alina Landowska. 2025. The Moral Gap of Large Language Models. (2025). arXiv:2507.18523 [cs] doi:10.13140/RG.2.2.26221.70880 [21] Jackson Trager, Alireza S. Ziabari, Aida Mostafazadeh Davani, Preni Golazizian, Farzan Karimi-Malekabadi, Ali Omrani, Zhihe Li, Brendan Kennedy, Nils Karl Reimer, Melissa Reyes, Kelsey Cheng, Mellow Wei, Christina Merrifield, Arta Khosravi, Evans Alvarez, and Morteza Dehghani. 2022. The Moral Foundations Reddit Corpus. doi:10.48550/ARXIV.2208.05545 [22] Lorenzo Zangari, Candida M. Greco, Davide Picca, and Andrea Tagarelli. 2025. ME2-BERT: Are Events and Emotions What You Need for Moral Foundation Prediction?. In Proceedings of the 31st International Conference on Computational Linguistics, Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert (Eds.). Association for Computational Linguistics, Abu Dhabi, UAE, 9516–9532. https://aclanthology.org/2025.colingmain.638/
Maciej Skórski
Table 6: Classifier parity: ROC-AUC per moral foundation and subcorpus (EN vs. PL), linear head on frozen LaBSE embeddings, 10-fold CV. 𝑝 >0 : one-sided test (𝐻 1 : EN > PL); 𝑝 <.02 : one-sided test (𝐻 1 : gap < 0.02). Corpus Subcorpus
MFRC
Foundation EN AUC PL AUC
Gap
𝑝 >0
𝑝 <.02
Everyday
Authority Care Fairness Loyalty Sanctity
0.814 0.853 0.790 0.802 0.730
0.787 0.842 0.784 0.793 0.725
+0.026 +0.011 +0.007 +0.009 +0.005
0.011 0.002 0.092 0.131 0.303
0.738 0.006 0.008 0.078 0.057
US politics
Authority Care Fairness Loyalty Sanctity
0.791 0.768 0.713 0.757 0.698
0.760 0.754 0.707 0.735 0.709
+0.031 +0.014 +0.006 +0.022 −0.011
0.000 0.017 0.086 0.016 0.881
0.953 0.153 0.003 0.596 0.003
Authority Care French politics Fairness Loyalty Sanctity
0.727 0.814 0.727 0.719 0.742
0.707 0.794 0.718 0.698 0.738
+0.021 +0.020 +0.008 +0.022 +0.005
0.000 0.001 0.033 0.013 0.287
0.610 0.498 0.007 0.587 0.053
ALM
Authority Care Fairness Loyalty Sanctity
0.876 0.795 0.788 0.739 0.714
0.828 0.773 0.757 0.735 0.708
+0.048 +0.023 +0.031 +0.003 +0.006
0.000 0.000 0.000 0.188 0.257
0.996 0.750 1.000 0.001 0.095
BLM
Authority Care Fairness Loyalty Sanctity
0.889 0.835 0.813 0.809 0.826
0.862 0.820 0.792 0.788 0.817
+0.027 +0.014 +0.021 +0.021 +0.009
0.001 0.026 0.001 0.006 0.121
0.876 0.206 0.613 0.567 0.081
Baltimore
Authority Care Fairness Loyalty Sanctity
0.857 0.816 0.867 0.880 0.759
0.846 0.802 0.856 0.869 0.740
+0.011 +0.014 +0.010 +0.011 +0.019
0.001 0.001 0.004 0.024 0.089
0.005 0.055 0.007 0.049 0.466
Davidson
Authority Care Fairness Loyalty Sanctity
0.596 0.508 0.549 0.596 0.591
0.596 0.511 0.552 0.576 0.570
+0.000 −0.004 −0.002 +0.020 +0.021
0.482 0.622 0.636 0.074 0.111
0.014 0.036 0.004 0.501 0.514
Election
Authority Care Fairness Loyalty Sanctity
0.838 0.800 0.865 0.753 0.828
0.801 0.782 0.834 0.735 0.817
+0.037 +0.018 +0.030 +0.018 +0.011
0.000 0.000 0.000 0.007 0.012
0.996 0.223 0.997 0.364 0.032
MeToo
Authority Care Fairness Loyalty Sanctity
0.765 0.824 0.806 0.762 0.760
0.761 0.809 0.788 0.741 0.744
+0.003 +0.015 +0.019 +0.021 +0.016
0.287 0.003 0.003 0.007 0.000
0.010 0.131 0.407 0.546 0.091
Sandy
Authority Care Fairness Loyalty Sanctity
0.883 0.846 0.866 0.803 0.805
0.858 0.820 0.827 0.778 0.762
+0.024 +0.026 +0.039 +0.025 +0.043
0.001 0.000 0.000 0.000 0.000
0.759 0.928 0.999 0.882 0.990
MFTC