Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs Joseba Fernandez de Landa1 Carla Perez-Almendros2 Jose Camacho-Collados2 1 HiTZ Center - Ixa, University of the Basque Country EHU 2 Cardiff University [email protected]
arXiv:2604.21751v1 [cs.CL] 23 Apr 2026
Abstract LLMs have been showing limitations when it comes to cultural coverage and competence, and in some cases show regional biases such as amplifying Western and Anglocentric viewpoints. While there have been works analysing the cultural capabilities of LLMs, there has not been specific work on highlighting LLM regional preferences when it comes to culturalrelated questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ). The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan. Moveover, our results show that when prompting in languages such as English or other highresource ones, LLMs tend to provide more diverse outputs and show less inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised finetuning, and not during pre-training.
1
Figure 1: Framework for measuring cultural biases using the proposed CROQ dataset: (1) prompting models with cultural questions containing undefined locations, (2) collecting open responses from the specific model, and (3) judging the responses to identify the referenced region(s).
et al., 2024; Chiu et al., 2025; Arora et al., 2025; Bulté and Rigouts Terryn, 2025). Many of these works evaluate LLMs’ responses against a set of predefined answers considered the gold standard (Romanou et al., 2024; Hasan et al., 2024; Arora et al., 2025; Chiu et al., 2025). Others evaluate the models’ performance against human answers, covering a more or less wide spectrum of topics, cultures and countries/regions (Li et al., 2024; Shi et al., 2024; Myung et al., 2024; Bulté and Rigouts Terryn, 2025). In this work, we propose a novel framework to evaluate cultural regional biases in LLMs, by prompting different models with open ambiguous cultural questions while masking the geographic information. This strategy requires the model not only to answer the cultural question, but also to geographically anchor that answer by selecting a location. By encouraging the models to link their responses to a specific region, our framework surfaces internal cultural
Introduction
Specific cultural knowledge, values and behaviours are often rooted in specific regions or countries. In this sense, we would expect that people from different geographical locations will have different answers to the same cultural questions. However, the behavior of Large Language Models (LLMs) to these types of culture-related input remain an open question. Extensive work has been done towards unveiling cultural biases in LLMs (Shi et al., 2024; Li et al., 2024; Bulté and Rigouts Terryn, 2025; Myung et al., 2024; Chiu et al., 2025; Arora et al., 2025), mainly showing biases in favour of Western or Anglocentric viewpoints, specially United States and Europe (Lee et al., 2024; Li et al., 2024; Hasan 1
priors and exposes inherent biases (Figure 1). We address the following Research Questions (RQ):
13 languages, revealing strong disparities between high- and low-resource languages. Expanding the scope further, NativQA (Hasan et al., 2024) and CaLMQA (Arora et al., 2025) use native-speaker based and culturally specific questions, respectively, across a wide range of topics, demonstrating heightened model failure rates for low-resource languages. Finally, Global MMLU (Singh et al., 2025) adapts MMLU benchmark by annotating questions for cultural sensitivity and expanding coverage to 42 languages. While these benchmarks reveal important weaknesses in factual cultural competence, their closed-answer format limits their ability to assess free-form generation. Beyond QA, several works construct resources of cultural commonsense as knowledge bases. CANDLE (Nguyen et al., 2023) automatically extracts over one million assertions about food, rituals, and behaviors across 386 cultural groups, while CultureBank (Shi et al., 2024) curates communitygrounded descriptors from TikTok and Reddit to capture lived experiences. These knowledge bases are useful for training and grounding models, but their assertional structures are primarily designed for factual recall rather than for evaluating models’ tendencies in situations where multiple perspectives are valid. A smaller but closely related line of work evaluates cultural adaptability and alignment. NormAd (Rao et al., 2025) measures whether models can adapt to diverse cultural norms across 75 countries using situational vignettes, showing that even state-of-the-art systems perform far below human levels, especially in abstract or implicit norm settings. CulFiT (Feng et al., 2025) introduces a training paradigm and evaluation dataset (GlobalCultureQA) for open-ended cultural QA, using fine-grained reward modeling to enhance cultural alignment. Similarly, Makieval (Zhao et al., 2025) proposes a framework for evaluating cultural awareness across 13 languages, 19 countries and regions, and six culturally salient topics, assessing openended generations by linking cultural entities in model outputs to Wikidata. While these approaches move beyond factual accuracy toward normative and generative evaluation, they often reduce outcomes to categorical labels (e.g., acceptable vs. unacceptable, aligned vs. misaligned) or optimize models toward a predefined target norm, rather than characterizing the full range of plausible cultural positions a model may express. In contrast to prior work, our framework explic-
1. What regional cultural biases do LLMs have? 2. What is the impact of language in the bias? 3. At what stage of training do the bias emerge? To support this analysis, we construct CROQ, a dataset of 31,680 open cultural questions spanning 24 languages, 11 major topics, and 66 subtopics. We then prompt LLMs with a question pattern which forces the model to choose a specific location, therefore unveiling hidden cultural and regional biases. Finally, we use an LLM as-a-judge on the generated answers of each model to extract the geographic reference chosen by each model, which is later manually validated. Our results indicate that regional preferences and model creativity are primarily determined by the input language. Further experiments suggest that these cultural and regional biases are induced predominantly during the post-training or instruction phase. Notably, experiments across all models, languages and topics reveal a strong inclination for countries like Japan. This finding challenges prior assumptions and opens new avenues for research into how post-training shapes model outputs and biases.
2
Related Work
Research on cultural knowledge and bias in LLMs is expanding rapidly (Adilazuarda et al., 2024; Liu et al., 2025; Pawar et al., 2025), producing a variety of benchmarks, knowledge bases, and evaluation frameworks. While these resources provide important insights into multilingual and cultural capabilities, most are designed around closed-form answers and do not directly capture the leanings of models when responding to open-ended, ambiguous cultural questions. Cultural knowledge is being benchmarked through question answering (QA) tasks. CulturalBench (Chiu et al., 2025) introduces humanauthored multiple-choice questions covering 45 world regions and 17 topics, finding that LLMs fall far short of human performance, particularly in under-represented regions. Include (Romanou et al., 2024) compiles nearly 200k exam-style questions in 44 languages, targeting multilingual and regional knowledge. BLEnD (Myung et al., 2024) emphasizes everyday cultural knowledge (e.g., food, family, leisure) across 16 countries and 2
itly targets open cultural questions for which no single correct answer exists, prompting the model to choose among all regions in the world. By eliciting multiple model generations, we evaluate the leaning of LLMs along diverse cultural topics and languages. This complements existing factual and normative benchmarks by shifting the focus from accuracy to distributional analysis, enabling a more nuanced understanding of cultural bias and alignment in LLMs. Moreover, we attempt to cover a wide range of cultural topics, encompassing areas covered in previous work into a single framework, as we detail in the following section.
3
1,320 queries in total. All questions follow the same basic pattern, avoiding direct references to specific countries or cultures, while remaining culturally meaningful. Questions are generated semiautomatically with GPT-5.1 and manually reviewed to remove repetitions or accidental location references. Example questions per topic are shown below. Beliefs: Social: Education: Arts: Food: Geography: Politics: Health: Media: History: Economy:
Dataset Generation: CROQ
We introduce Culture-Related Open Questions (CROQ), a multilingual dataset designed to uncover the cultural tendencies of LLMs through openended yet culturally grounded questions. CROQ consists of culturally relevant prompts spanning 11 broad topics and 66 finer-grained subtopics, constructed consistently across topics to preserve structural uniformity and regional ambiguity. The dataset has been translated into 24 languages, covering a wide range of language families, typologies, and resource levels. It contains 1,320 questions (20 per subtopic) which, after translation, results in a total of 31,680 questions.
What legends explain the land? What is the role of neighbors? What subjects are most valued? What traditional dances exist? What foods are for daily meal? What rivers influence settlements? What political protests are common? What exercise routines are common? What film industries exist? What historical tours exist? What commodity markets exist?
Multilingual Coverage The dataset was initially constructed in English and then automatically translated and post-edited into 23 additional languages to ensure broad coverage across speaker populations, resource levels, and language families (full list in Appendix B). As a starting point, we included the 13 languages from the BLEND dataset (Myung et al., 2024), which range from lowresource (Amharic -am-, Assamese -as-, Azerbaijani -az-, Hausa -ha- and Sundanese -su-) to midand high-resource (Greek -el-, Indonesian -id-, Korean -ko-, Persian -fa-, Arabic -ar-, Chinese -zh-, Spanish -es-, English -en-) languages and provide wide geographical coverage. To further enhance diversity, we added several of the world’s most widely spoken languages with varying degrees of representation in CommonCrawl (Hindi -hi-, French -fr-, Bengali -bn-, Portuguese -pt-, Russian -ru-, German -de-, Japanese -ja-, Swahili -sw-). Additionally, we included all co-official but low-resource languages of Spain (Basque -eu-, Galician -gl-, Catalan -ca-)1 , which have relatively small speaker populations.
Cultural Taxonomy To build the open questions diverse and representative, we construct a taxonomy of 66 cultural subtopics grouped into 11 higher-level domains: Beliefs, Values, and Identity ( ); Social Structure and Daily Life ( ); Knowledge, Communication, and Education ( ); Cultural Expression and the Arts ( ); Food, Drink, and Leisure ( ); Geographic Aspects ( ); Political Aspects ( ); Health and Wellness ( ); Media and Entertainment ( ); History ( ); and Economy and Industry ( ). This taxonomy is intended to capture a broad and representative spectrum of cultural topics, encompassing as many culturally relevant dimensions as possible. It was developed by aggregating the topics from the datasets discussed in the related work and expanding them into a more comprehensive collection (more detail in Appendix A).
4
Evaluation Framework
We propose a novel method to uncover cultural biases in LLMs. We achieve this by prompting models with CROQ and require them to explicitly select a country or region. Subsequently, a secondary model is employed to systematically extract the geographic information inferred from the
Open Question Generation A central design choice in CROQ is the deliberate open-ended nature of each question. For each of the 66 cultural subtopics, we generate 20 questions, resulting in
1
3
We use ISO 639-1:2002 codes for each language.
responses, allowing us to analyze potential cultural biases (see Figure 1). 4.1
in our dataset. The models include: gpt-4o-mini, gemini-2.5-flash, claude-3.5-haiku, llama4-maverick, command-r-08-2024, magistralsmall-2506, deepseek-v3.2-exp, qwen3-next80b-a3b-instruct.
Regional Question Grounding
To probe cultural priors in LLMs, we leverage the set of culture-related open questions (Section 3) and require the model to openly generate responses to these questions. Each question contains a locational placeholder (in region/place) that the model must implicitly resolve when producing an answer. We further prompt the model to be brief and to select a specific location (example below), thereby requiring it to tacitly choose a region or cultural background to ground its response:
4.4
For our subsequent analysis, we devise three metrics that can help analyse the output of each LLM. Diversity. We define diversity as the number of distinct countries or regions referenced across all model outputs. Higher values indicate broader geographic coverage of cultural references. Entropy. We compute normalized entropy over the distribution of referenced places to capture how evenly cultural references are distributed, with higher scores indicating greater balance. Raw counts. We additionally report country or region frequency counts to provide complementary and interpretable evidence.
What values shape family life {in region/place}? Be brief. Choose yourself the place.
Each LLM evaluated receives the same set of 1,320 open questions in different languages (see Section 3) for more details. The prompt used to guide the model varies by language. To reduce prompt variability, all questions follow a standardized template, and no explicit regional cues are provided. This ensures that any regional or cultural assumptions arise from the model’s internal priors rather than prompt design. 4.2
5
RQ1: What Regional Cultural Biases Do LLMs Have?
Models favor input language countries: Table 1 summarizes the regional distribution of country references across models for the 24 languages. Across all models, references overwhelmingly concentrate on own countries (ranging from 43% for Commandr to 78% for GPT), confirming a strong alignment between the output language and its associated regions. This pattern is consistent and robust, suggesting that language choice alone strongly constrains the cultural priors adopted by the models. Exogenous references are dominated by a small set of regions: The data reveal that for non-own country references, all models display a similar pattern of global salience. On average, Japan (preferred country for six out of the eight evaluated models) and the United States are the most frequently referenced nations, followed by India, China, and France. References to other countries are considerably less frequent and vary by model. Differences among models: Although the distribution across mentioned exogenous countries (Japan, the United States, India, and China) is similar, some differences emerge. Five models favor Japan, while two emphasize the US. Command-R stands out as the least Japan-biased model (albeit the most USbiased) and the one most likely to refuse an answer, producing in general the most diverse outputs with the highest entropy. Command-R, DeepSeek, and
Processing Model Responses
To process the open-ended answers generated by the LLMs, we employ a secondary model as-ajudge to interpret the responses and extract the relevant information required to associate each question with a specific set of countries. This process is designed to uncover potential cultural biases by tasking the judge model with identifying up to five regions or places mentioned or implied in each response; if no such information can be reliably determined, the model is prompted to explicitly indicate that no inference is possible (non-answered responses). Following an evaluation of various prompting strategies, we selected the optimal approach, which achieved an accuracy of 98% over 264 evaluation items, the details of which, including the prompts and additional technical information, can be found in Appendix C. 4.3
Analysis Metrics
Comparison of Frontier LLMs
For our evaluation, we compare a wide range of multilingual LLMs, including both frontier closedand open-weight models. All models are accessed through OpenRouter, used with its provider’s default settings and evaluated across the 24 languages 4
Model GPT Gemini Claude Llama Command-r Magistral Qwen DeepSeek
Own (%)
NA
24,763 (.78) 20,172 (.64) 20,585 (.65) 20,775 (.66) 13,707 (.43) 23,040 (.73) 22,572 (.71) 21,907 (.69)
1,136 1,853 2,063 887 4,815 2,485 1,502 877
811 1,493 1,601 2,701 936 1,754 1,861 2,104
944 1,493 1,200 1,074 2,064 1,254 483 1,006
451 673 510 453 567 496 249 541
237 473 560 524 512 560 434 608
289 439 210 587 504 632 191 351
265 344 218 265 389 392 170 415
128 310 138 190 515 264 60 390
150 271 172 173 309 390 158 499
134 164 340 75 209 315 222 308
234 125 349 222 259 339 67 205
53 75 122 61 113 163 58 110
87 240 141 132 142 211 60 267
Div
Ent
82 113 95 102 120 88 79 116
0.40 0.50 0.48 0.49 0.59 0.47 0.40 0.52
Table 1: Frontier Model outputs for 24 languages. Own counts references to countries in which the language is an official language. NA denotes missing responses. Top referenced countries are Japan, USA, India, China, France, Italy, United Kingdom, Germany, Mexico, Brazil, Russia, Egypt. Div and Ent represent the average for diversity and entropy by model. Bold values indicate the highest count for each model, excluding self-references, as well as the highest diversity and entropy values (two right-most columns).
Gemini have higher output diversity (over 113 different countries mentioned), whereas Qwen, GPT, and Magistral show lower diversity and entropy (fewer than 88 countries mentioned in their outputs – more details on individual model responses in Appendix D.2). Topic Analysis: Table 2 presents the top countries mentioned by the models across the 11 general topics in our dataset. Japan and the US remain salient across all categories, followed by other frequently mentioned countries. However, some variation occurs depending on the topic which may reflect global trends. For example, in Geography and Economy, there is a noticeable shift in prominence between the US and Japan, where US takes the lead in number of references, and Politics and History are the only two topics where Japan is relegated to third and fourth positions. Countries such as Greece, China and France stand out in very specific topics, namely Beliefs, Politics and History, respectively. Or other countries such as South Korea, which is barely mentioned in most topics, emerge as the third most mentioned country in Media and Entertainment. Mentions of other countries also fluctuate across topics, but their overall counts remain substantially lower.
Topic
Top Countries
1. Beliefs 2. Social 3. Education 4. Arts 5. Food 6. Geography 7. Politics 8. Health 9. Media 10. History 11. Economy Table 2: Top 10 countries most frequently referenced by all Frontier API models for each topic in the taxonomy. 24 languages, excluding own country mentions.
in frontier model outputs, with a strong concentration on a small set of dominant regions.
6
RQ2: What is the Impact of Language in Regional Cultural Bias?
Salient regions are overrepresented across all languages: As established in the previous section, models strongly associate languages with their primary regions or countries of origin. To quantify this effect and explore its variation, we conducted a per-language analysis. Table 3 shows that beyond self-referential patterns, the models exhibit a stable hierarchy of non-own country references. Japan and the United States are the most frequently referenced countries across languages. Notably, these references appear even in languages with limited cultural or geographic ties to these countries. This finding aligns with prior observations that LLMs
Summary of findings and discussion. All eight frontier models exhibit a clear bias toward languages referring to their own regions of origin. When associations to own language-region pairs are isolated, this bias becomes more pronounced. Consequently, all models consistently favour Japan or the United States, while regions such as India and China receive comparatively less emphasis, and references to other regions are negligible. This pattern suggests an uneven regional representation 5
Diversity vs CommonCrawl Coverage
Own en zh hi es ar fr bn pt ru id de ja ko fa sw ha az am su el as ca gl eu
31% 48% 70% 66% 69% 56% 79% 73% 36% 76% 70% 59% 59% 57% 78% 51% 75% 77% 79% 70% 83% 72% 75% 78%
2,893 2,064 94 843 504 832 73 479 1,423 411 485 * 1,064 473 230 30 123 43 167 330 41 213 314 132
* * 367 1460 642 * 104 * 92 588 130 173 396 182 85 557 215 236 151 * 59 442 167 161 657 265 431 246 129 133 237 169 203 736 434 219 1,308 574 663 544 263 194 389 85 79 99 30 44 212 81 64 106 17 19 146 49 76 307 161 135 53 * 42 339 170 224 297 129 143 144 48 66
Div
Ent
157 142 56 136 108 149 68 114 146 78 122 114 112 117 89 49 78 38 62 106 41 121 113 68
0.69 0.57 0.14 0.65 0.71 0.54 0.20 0.42 0.60 0.26 0.38 0.36 0.47 0.48 0.50 0.10 0.43 0.13 0.11 0.36 0.30 0.49 0.37 0.37
CommonCrawl share (%) (log scale)
tend to overrepresent culturally salient entities. 101 100
Language family Indo-European Sino-Tibetan Austronesian Afro-Asiatic Turkic Niger Congo Other
hi
bn eu
10 2 10 3
ja de zh ru es fr pt ko ar fa el ca
id
10 1
en
az
gl sw
am
as
4 × 101
ha
su 6 × 101
102
Diversity (log scale)
Figure 2: Correlation between CommonCrawl data share and output diversity (mean of models) across 24 languages (log-scale).
non-responses) and lower diversity scores. This pattern suggests a link between training data volume and cultural framing: for languages with less comprehensive corpora, models default to safer, more self-referential, or less diverse outputs (more detail in Appendix D.1). Pretraining data and diversity metrics correlate: To quantify the relationship between output diversity and language resource availability, we perform a statistical analysis correlating output diversity with the amount of CommonCrawl data available for each language, used as a proxy for pretraining resources (see Figure 2). Across all frontier models, output diversity correlates positively with CommonCrawl coverage. Spearman’s correlation score rs (0.843) reveals a strong rank-based association; being statistically significant (p-value < 0.001). Analyses at the individual model level further confirm this trend for the majority of systems (see Appendix D.3). Together, these results indicate that higher-resource languages benefit from more varied cultural representations, whereas lower-resource languages tend to elicit more selfreferential or constrained outputs, reflecting reduced diversity in generated content.
Table 3: Model outputs for the 24 languages. Own: percentage of references to countries in which the language is an official language; Flags: references to each country (top 4 displayed); * mark references to own countries; Div and Ent represent the average for diversity and entropy for all models. Colours indicate the top three and bottom three results.
Bias is consistent on exogenous references across languages: When isolating non-own-country references, a consistent global hierarchy emerges across all languages. Japan and the United States remain the most frequently cited countries, followed by India, China, and France. This suggests that crosscountry mentions are driven more by a model’s inherent priors of cultural salience than by contextual relevance to the specific language, underscoring a systematic overrepresentation of a narrow set of culturally dominant entities. Low-resource languages produce more selfreferential outputs: A key factor moderating this bias is the resource level of the language. Comparing languages by the volume of available training data reveals a clear pattern. Higher-resource languages (e.g., en, zh, es, fr, ru) show a lower proportion of own-country references and higher output diversity, with a significant skew toward globally dominant regions like Japan and the United States. Lower-resource languages (e.g., su, as, am, ha, eu), demonstrate a markedly higher proportion of own-country references (or an increase in
Summary of findings and discussion. Overall, these findings demonstrate that language models exhibit strong language-aware referencing behavior. However, once self-references are excluded, models consistently rely on the same narrow set of countries across languages. A particularly notable result is the disproportionate prominence of Japan, which emerges as one of the most frequently referenced exogenous countries across nearly all languages. 6
The United States is also highly salient, although this is less surprising given results from previous work and its global influence and the dominance of English as a high-resource language. Taken together, these patterns suggest that the composition of training data, either during pretraining or posttraining, plays a central role in shaping models’ culturally biased referencing behavior.
7
Diversity
RQ3: At What Stage of LLM Training Do Regional Cultural Biases Emerge?
To identify the training stage at which regional cultural biases emerge, we extend our analysis to open-weight base and instruction-tuned LLMs. By comparing models before and after post-training, we aim to disentangle biases learned during largescale pretraining from those introduced or amplified during supervised fine-tuning and alignment. 7.1
Entropy
Model
Base SFT Inst. Base SFT Inst.
Llama-3.1-8B Llama-3.1-70B gemma-2-9b gemma-2-27b Mistral-7B Qwen2.5-7B Qwen2.5-32B Qwen2.5-72B
109 117 95 80 91 147 142 144
-
174 174 136 154 160 150 118 127
0.78 0.79 0.70 0.77 0.81 0.71 0.65 0.66
-
0.66 0.65 0.71 0.73 0.63 0.62 0.61 0.60
Olmo-3-7B Olmo-3-7B-Think OLMo-2-7B OLMo-2-32B
114 114 102 88
145 121 151 151
124 127 158 161
0.74 0.74 0.84 0.77
0.66 0.66 0.63 0.65
0.64 0.68 0.64 0.71
Table 4: Diversity and normalized entropy scores for base, SFT (for OLMo) and instruct model variants. For each model, the highest diversity and entropy are bolded.
7.2
Findings
Figure 3 contrasts the distributions of country references produced by base and instruction-tuned models, while Table 4 shows the diversity and entropy scores for each model variant. Cultural distributions are more balanced on Base Models: Base models (Figure 3, left) exhibit a more balanced geographic coverage. While the United States remains prominent, substantial references are also made to Japan, India, China, and several European countries. This pattern is consistent across architectures and model sizes, suggesting that pretraining alone yields a comparatively diffuse set of cultural associations. This is further reflected on the entropy scores, which are consistently lower after instruction-tuning for all models. Cultural leaning is stronger after InstructionTuning: Despite providing a higher number of countries overall, instruction-tuned models (Figure 3, right) display a pronounced concentration of references to a narrow set of regions. Across all examined model families, instruction tuning sharply increases alignment with the United States and Japan while reducing references to most other countries. This convergence toward culturally dominant regions occurs even in models developed outside Western contexts, indicating that post-training induces a homogenization of cultural perspectives rather than merely reflecting model origin. Supervised fine-tuning as a primary driver of cultural bias: To further assess whether instruction tuning itself introduces or amplifies cultural bias, we analyze OLMo models, which provide
Experimental Setting
Comparison Models. To assess whether instruction tuning itself introduces or amplifies bias, we analyze open models providing both base and instruct variants across multiple enterprises and parameter scales, using default settings. This evaluation is performed only on English and includes Llama-3.1 (8B, 70B) (Dubey et al., 2024), Qwen2.5 (7B, 32B, 72B) (Yang et al., 2024), Gemma-2 (9B, 27B) (Team et al., 2024) and Mistral-7B-v0.3 (7B) (Jiang et al., 2023). In addition, we analyze OLMo-2 (7B, 32B) (Team, 2025a) and OLMo-3 (7B, 7B-Think) (Team, 2025b), which provide base, supervised fine-tuned (SFT), and fully aligned instruction models. This allows us to disentangle the effects of instruction tuning from those of subsequent supervised finetuning. Prompting. For this analysis, due to model availability, we restrict all experiments to English prompts. Given the nature of base models that are explicitly trained to predict the next tokens, we rephrase the question prompt utilised in our experiments so the sentence can be completed rather than answered. Everything else remains unchanged and both base and instruct models are given the same questions as input.2 2
In Appendix G, we provide the prompt and an analysis of the impact of the prompt in base and instruct models.
7
Mentions by BASE Models (en)
Mentions by INSTRUCT Models (en)
600
Llama-8B
Llama-8B 500
Llama-70B Gemma-9B
Gemma-9B
400
Gemma-27B
500
Llama-70B 400
Gemma-27B 300
Mistral-7B Qwen-7B
Qwen-7B
200
Qwen-32B
200
Qwen-32B
100
Qwen-72B
300
Mistral-7B
100
Qwen-72B
JP US IN CN FR IT GB DE MX BR RU EG Country
JP US IN CN FR IT GB DE MX BR RU EG Country
Figure 3: Number of outputs from Base and Instruct models of Llama-3.1, Gemma-2, Mistral and Qwen2.5 variants over the English questions. Countries included: Japan (JP), USA (US), India (IN), China (CN), France (FR), Italy (IT), UK (GB), Germany (DE), Mexico (MX), Brazil (BR), Russia (RU), Egypt (EG).
base, supervised fine-tuned (SFT), and instructionaligned variants. In Figure 4 we observe that the most substantial shift in country reference distributions occurs during supervised fine-tuning: SFT sharply increases concentration on a small number of dominant regions (most notably the United States and Japan) while reducing entropy across all model outputs (see Table 4). Subsequent instruction alignment only marginally mitigates these effects and does not recover the more balanced distributions observed in base models. These results indicate that alignment-induced cultural bias primarily originates from supervised fine-tuning data rather than pretraining, and that later instruction tuning largely preserves, rather than corrects, these learned cultural priors (see Appendix F for further details).
600
india china
france italy
uk germany
Olmo-3-7B
Olmo-3-7B-Think OLMo-2-7B
mexico brazil
OLMo-2-32B
500 400 300 200 100 0
Bas. SFT
Ins. Bas. SFT
Ins. Bas. SFT
Ins. Bas. SFT
Ins.
Figure 4: Number of outputs mentioning top-referenced countries from the Base, SFT, and Instruct variants of the named OLMo models.
suring the impact of certain decisions is hard given all the cultural aspects that are involved in model responses. In this paper, we created a dataset of open question based on a taxonomy of eleven cultural domains and sixty-six subtopics. We used this dataset to test eight frontier LLMs by prompting them in twenty-four languages. The results show an inclination of models to provide cultural examples related to regions for which the prompted language is an official language. Not only that, but there is a similarity on the type of example that models choose, preferring countries such as Japan, United States or India overwhelmingly with respect to other choices across the world. When analysing the potential causes, the results indicate that this concentration of answers emerge during post-training. While this is a stage traditionally thought to guide model into providing more diverse and unbiased responses, it appears to also increase this cultural bias into a different direction. For future work, it would be interesting to think of new methods to take this be-
Summary of findings and discussion. Overall, base models distribute references more evenly across regions, exhibiting lower concentration of regions despite limited diversity. Instruction-tuned models, by contrast, show substantially more concentrated cultural biases. These findings suggest that post-training, rather than pretraining, plays a decisive role in shaping the dominant cultural perspectives expressed by LLMs (more detail and results in Appendix E). For a better understanding, it would be interesting to further check the instructions provided to LLMs that can potentially amplify these biases. Finally, the results from OLMO suggests that this behaviour starts appearing after supervised fine-tuning.
8
japan usa
Conclusions
Ensuring cultural awareness and diversity in LLMs is an open research question in NLP. However, mea8
Ethical considerations
haviour into account, in order to better understand the post-training effects in model output diversity as a whole, in particular when it comes to different languages and regions.
In this work, we employ LLMs to semiautomatically generate candidate questions (GPT5.1, Sec. 3) for the CROQ dataset, and also as-ajudge (gpt-4o-mini, Sec. 4.2) to extract specific information from LLMs’ answers. We recognize that this methodology introduces potential biases present in the model. To mitigate this, all generated questions underwent manual review to eliminate repetitions, accidental location references, and other artifacts that could artificially skew cultural representations. This human-in-the-loop approach was essential given that our research explicitly examines cultural biases in LLM outputs. Using LLMs in the creation of our dataset required careful oversight to avoid encoding the same biases we try to unveil with our work. The cultural taxonomy underlying CROQ was developed by aggregating topics from existing benchmarks and expanding them into a comprehensive framework of 11 domains and 66 subtopics. Although we designed this taxonomy to capture a broad spectrum of culturally relevant dimensions, we acknowledge that cultural phenomena exist on a continuum and cannot be exhaustively categorized. This structural limitation means our analysis, while broad, remains grounded in particular cultural frameworks rather than offering a universal view of human culture. In addition, we made deliberate efforts to ensure broad linguistic coverage by including 24 languages spanning high-, mid-, and low-resource categories, encompassing six distinct language families and providing geographical coverage across continents, trying to capture how cultural biases in LLMs manifest across different linguistic and resource contexts, with particular attention to voices often marginalized in NLP. However, we recognize that 24 languages, while substantial, represent only a small fraction of the world’s ~7,000 languages and the diverse cultures they embody. Our selection, while considered and intentional, necessarily privileges certain linguistic communities and regional perspectives. Therefore, in conducting this research, we are involuntarily reinforcing the underrepresentation of many communities and languages that remain outside the scope of large-scale computational work. We do not present this limitation as merely a technical constraint; rather, it reflects deeper structural inequalities in language technology research. We
Limitations Our study has several limitations that should be considered when interpreting the results. 1) The comparison between base and instructiontuned models is conducted only in English, which limits the extent to which these findings generalize to other languages. In addition, base and instruction-tuned models differ inherently in training objectives and interaction style, which may influence generation behavior independently of the factors analyzed here. Nevertheless, this comparison provides a first step toward analyzing the intersection between base and instruction-tuned models, which we plan to explore further in future work. 2) All experiments are conducted using default generation settings, and we do not explore the robustness of our results to alternative decoding parameters. We deliberately chose this approach in order to focus on the settings employed by most users, enabling us to compare across different languages and models – an additional exhaustive exploration of decoding parameter configurations would incur in high generation and analysis costs. 3) We rely primarily on an automatic judge rather than human evaluation for all languages. While we validate the judge on a subset of languages, automatic evaluation may fail to capture subtle linguistic or cultural distinctions. Nevertheless, given that our expertise is limited to a small number of languages, conducting large-scale human evaluation would not be feasible, making automatic evaluation a necessary choice which we believe that it does not affect to the reliability of the results as a whole. 4) Our evaluation covers only 24 languages, representing a limited subset of global linguistic diversity. We further restrict some analyses to official languages of countries, which may not accurately reflect real-world language use, as many widely spoken languages extend beyond official or national boundaries. Moreover, referring to countries as proxies for cultural contexts may lead to oversimplifications. 5) Finally, we do not assess the impact of the observed behaviors on downstream tasks, leaving open the question of how these findings translate to applied or real-world scenarios. 9
encourage future work to extend beyond national language boundaries and to center the perspectives and needs of underrepresented communities in designing cultural evaluation frameworks. Our results, particularly the concentration of model outputs towards culturally dominant regions like Japan and the United States, have implications for global users of these systems. Models trained to favour certain cultural perspectives may provide inadequate or inappropriate responses for users from other backgrounds. We hope this work contributes to awareness among both model developers and users about the need for more balanced and inclusive cultural training.
Vered Shwartz, and Yejin Choi. 2025. CulturalBench: A robust, diverse and challenging benchmark for measuring LMs’ cultural knowledge through humanAI red-teaming. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 25663– 25701, Vienna, Austria. Association for Computational Linguistics. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Ruixiang Feng, Shen Gao, Xiuying Chen, Lisi Chen, and Shuo Shang. 2025. CulFiT: A fine-grained cultural-aware LLM training paradigm via multilingual critique data synthesis. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 22413–22430, Vienna, Austria. Association for Computational Linguistics.
Acknowledgements Jose Camacho-Collados is supported by a UKRI Future Leaders Fellowship.
Md Arid Hasan, Maram Hasanain, Fatema Ahmad, Sahinur Rahman Laskar, Sunaya Upadhyay, Vrunda N Sukhadia, Mucahid Kutlu, Shammur Absar Chowdhury, and Firoj Alam. 2024. Nativqa: Multilingual culturally-aligned natural query for llms. arXiv preprint arXiv:2407.09823.
References Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Shivdutt Singh, Alham Fikri Aji, Jacki O’Neill, Ashutosh Modi, and Monojit Choudhury. 2024. Towards measuring and modeling “culture” in LLMs: A survey. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 15763–15784, Miami, Florida, USA. Association for Computational Linguistics.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, and 1 others. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825. Nayeon Lee, Chani Jung, Junho Myung, Jiho Jin, Jose Camacho-Collados, Juho Kim, and Alice Oh. 2024. Exploring cross-cultural differences in English hate speech annotations: From dataset construction to analysis. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 4205–4224, Mexico City, Mexico. Association for Computational Linguistics.
Shane Arora, Marzena Karpinska, Hung-Ting Chen, Ipsita Bhattacharjee, Mohit Iyyer, and Eunsol Choi. 2025. CaLMQA: Exploring culturally specific longform question answering across 23 languages. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11772–11817, Vienna, Austria. Association for Computational Linguistics. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, and 12 others. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
Cheng Li, Mengzhuo Chen, Jindong Wang, Sunayana Sitaram, and Xing Xie. 2024. Culturellm: Incorporating cultural differences into large language models. In Advances in Neural Information Processing Systems, volume 37, pages 84799–84838. Curran Associates, Inc. Chen Cecilia Liu, Iryna Gurevych, and Anna Korhonen. 2025. Culturally aware and adapted NLP: A taxonomy and a survey of the state of the art. Transactions of the Association for Computational Linguistics, 13:652–689.
Bram Bulté and Ayla Rigouts Terryn. 2025. Llms and cultural values: The impact of prompt language and explicit cultural framing. Computational Linguistics, pages 1–85.
Junho Myung, Nayeon Lee, Yi Zhou, Jiho Jin, Rifki Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Perez-Almendros, Abinew Ali Ayele, and 1 others. 2024. Blend: A benchmark for llms on everyday knowledge in diverse cultures and languages.
Yu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov,
10
OLMo Team. 2025a. 2 olmo 2 furious. arXiv preprint arXiv:2501.00656.
Advances in Neural Information Processing Systems, 37:78104–78146. Tuan-Phong Nguyen, Simon Razniewski, Aparna Varde, and Gerhard Weikum. 2023. Extracting cultural commonsense knowledge at scale. In Proceedings of the ACM Web Conference 2023, WWW ’23, page 1907–1917, New York, NY, USA. Association for Computing Machinery.
OLMo Team. 2025b. arXiv:2501.00656.
Olmo 3.
arXiv preprint
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, and 1 others. 2024. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115.
Siddhesh Pawar, Junyeong Park, Jiho Jin, Arnav Arora, Junho Myung, Srishti Yadav, Faiz Ghifari Haznitrama, Inhwa Song, Alice Oh, and Isabelle Augenstein. 2025. Survey of cultural awareness in language models: Text and beyond. Computational Linguistics, pages 1–98.
Raoyuan Zhao, Beiduo Chen, Barbara Plank, and Michael A Hedderich. 2025. Makieval: A multilingual automatic wikidata-based framework for cultural awareness evaluation for llms. arXiv preprint arXiv:2505.21693.
Abhinav Sukumar Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, and Maarten Sap. 2025. NormAd: A framework for measuring the cultural adaptability of large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 2373–2403, Albuquerque, New Mexico. Association for Computational Linguistics.
A
Topics and subtopics in the Taxonomy
We construct a comprehensive taxonomy comprising 66 fine-grained cultural subtopics, which are organized into 11 higher-level cultural domains. The taxonomy was designed to capture a broad and diverse range of cultural phenomena while maintaining sufficient granularity to support detailed analysis. Each higher-level domain aggregates thematically related subtopics, enabling both coarsegrained and fine-grained examination of cultural dimensions. The complete taxonomy, including definitions and hierarchical relationships between domains and subtopics, is presented in Table 5.
Angelika Romanou, Negar Foroutan, Anna Sotnikova, Zeming Chen, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Maheshwary, Micol Altomare, Mohamed A Haggag, Alfonso Amayuelas, and 1 others. 2024. Include: Evaluating multilingual language understanding with regional knowledge. arXiv preprint arXiv:2411.19799. Weiyan Shi, Ryan Li, Yutong Zhang, Caleb Ziems, Sunny Yu, Raya Horesh, Rogério Abreu De Paula, and Diyi Yang. 2024. CultureBank: An online community-driven knowledge base towards culturally aware language technologies. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 4996–5025, Miami, Florida, USA. Association for Computational Linguistics. Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, Andre Martins, Leshem Choshen, Daphne Ippolito, and 4 others. 2025. Global MMLU: Understanding and addressing cultural and linguistic biases in multilingual evaluation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 18761–18799, Vienna, Austria. Association for Computational Linguistics.
🌱 🏠🧠 🎭 🍽 🌍 🏛 🩺📺 📜💰
Figure 5: Taxonomy coverage for Related Work datasets. Percentage of subtopics covered in each higher-level domain proposed in our taxonomy for the named datasets.
We conducted a semi-automatic evaluation of our proposed categories and subcategories, and we analyzed the datasets introduced in related work. Figure 5 illustrates the extent to which each general topic in our taxonomy is covered by existing datasets. The analysis shows that many
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, and 1 others. 2024. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118.
11
Domains
Subtopics
1. Beliefs, Values, and Identity
Religion & Spirituality: Beliefs and practices about the divine, morality, and afterlife. Values & Ethics: Core beliefs about right/wrong, community, and individual roles. Symbolism: Visual or material symbols that carry cultural meaning. Mythology & Folklore: Narratives that explain origins, values, or natural phenomena. Gender Roles: Cultural expectations related to gender identity and expression.
2. Social Structure and Daily Life
Family & Social Roles: Structure of family and communal relationships. Community Life: How people relate to neighbors and groups. Manners & Etiquette: Everyday norms for respectful behavior and interaction. Work Culture: Attitudes and habits related to labor, time, and productivity. Conflict Resolution: How societies handle disputes or disagreements.
3. Knowledge, Communication, and Education
Language & Communication: Spoken, written, and non-verbal systems of communication. Education Systems: Ways knowledge is transmitted and valued. Technology & Tools: Use of tools and innovation in daily and cultural life. Knowledge Transmission: How wisdom and skills are preserved across generations. Science & Philosophy: Intellectual traditions shaping discovery and thought.
4. Cultural Expression and the Arts
Music: Rhythmic and melodic expression of culture Dance: Movement as performance, ritual, or celebration Art: Visual cultural expression (painting, sculpture, etc.) Literature & Storytelling: Written and oral traditions conveying culture and values Theater & Performance: Live dramatic cultural expression Film & Cinema: Storytelling through modern moving-image media Crafts & Handicrafts: Traditional material culture and artisanal skills Jewelry & Body Decoration: Ornamentation for identity, beauty, or tradition Architecture & Housing: Building styles and structures shaped by culture and climate
5. Food, Drink, and Leisure
Food & Cuisine: Traditional dishes, cooking styles, and dietary customs Traditional Beverages: Culturally significant drinks Recreation & Leisure: Activities for enjoyment and relaxation Sports & Games: Physical activities and competitions enjoyed culturally Life-Cycle Rituals: Ceremonies marking stages of life Festivals & Celebrations: Seasonal, national, or spiritual gatherings expressing shared identity
6. Geographic Aspects
Climate & Environment: How weather and landscape shape cultural practices, housing, and food Topography & Land Use: How people interact with physical geography like mountains, rivers, or plains Rural vs Urban Culture: Differences in lifestyle, work, and values between city and countryside Regional Identity: Sub-national cultural traits tied to a specific region Borders & Territory: How borders influence cultural blending or division Migration & Diaspora: Movements shaping identity Relationship with Nature: Beliefs and practices related to land, animals, and ecology
7. Political Aspects
Government Systems: Political structures influencing law, rights, and daily life Law & Legal Traditions: Cultural expectations of justice, rights, and punishment Civic Values & Participation: How people engage with politics, voting, and national identity Freedom of Expression: Cultural and political limits on speech or art Colonial History & Influence: Impact of colonization on culture, identity, and systems Power & Hierarchy: How authority is structured culturally and politically Human Rights & Social Movements: How cultures address equity, justice, and activism
8. Health & Wellness
Health Practices: Traditional, spiritual, or modern approaches to healing and care Public Health: Collective strategies for health and safety Wellbeing & Lifestyle: Balancing mental, physical, and social health Birth & Reproductive Health: Fertility, childbirth, and parenting norms and practices
9. Media & Entertainment
Local Media: Print, radio, and TV shaping cultural narratives Digital Media: Online platforms influencing communication, identity, and activism Entertainment Industries: Mass production of cultural content News & Information: Flow of journalism and narratives in society Popular Culture: Icons, trends, and shared cultural references Gaming & Interactive Media: Games, esports, and digital interactivity
10. History
Historical Events: Major turning points shaping collective identity Historical Figures: Leaders, thinkers, and cultural icons Cultural Memory: How societies remember, teach, and reinterpret their past Colonialism & Resistance: Struggles of domination and liberation National Narratives: Shared stories of origin, destiny, and identity Heritage & Preservation: Protecting, curating, and interpreting cultural pasts
11. Economy & Industry
Key Industries: Dominant sectors shaping work and identity Trade & Exchange: Movement of goods, ideas, and culture Labor & Employment: Work systems, professions, and class structures Wealth & Inequality: How resources and opportunities are distributed Economic Growth & Development: Paths to modernization and sustainability Globalization & Cultural Economy: Markets influencing culture and identity
Table 5: Overview of the taxonomy’s domains (11) and subtopics (66), with brief definitions for each category.
datasets address only a limited subset of the general topics defined in our taxonomy. Although NativQA (Hasan et al., 2024) and CaLMQA (Arora et al., 2025) cover all general topics, their cover-
age across subtopics is uneven, with notable imbalances within individual topics. Our goal is not to propose a definitive or exhaustive taxonomy of cultural topics, but rather to cover as broad a range 12
of cultural topics as possible, in order to prompt LLMs with maximal cultural diversity. % of Common Crawl content
B
Selected languages analysis
Table 6 shows the coverage of languages in our work, detailing the distribution of the 24 languages included in our dataset. For each language, the table reports its relative proportion within the Common Crawl snapshot CC-MAIN-2025-43, together with an estimate of the total number of speakers (in millions), derived from publicly available Wikipedia statistics. In addition, we provide the corresponding language family for each language, as well as the ISO 639-1:2002 two-letter language code used consistently throughout the paper and accompanying resources. This information offers a concise yet comprehensive view of the linguistic diversity and representativeness of the dataset analyzed in this work. Code
Language
Family
en zh hi es ar fr bn pt ru id de ja ko fa sw ha az am su el as ca gl eu
English Chinese Hindi Spanish Arabic French Bengali Portuguese Russian Indonesian German Japanese Korean Persian Swahili Hausa Azerbaijani Amharic Sundanese Greek Assamese Catalan Galician Basque
Indo-European Sino-Tibetan Indo-European Indo-European Afro-Asiatic Indo-European Indo-European Indo-European Indo-European Austronesian Indo-European Other Other Indo-European Niger–Congo Afro-Asiatic Turkic Afro-Asiatic Austronesian Indo-European Indo-European Indo-European Indo-European Other
Speakers (M)
CC %
1,528 1,184 609 558 335 312 284 267 253 252 134 126 82 127 87 94 24 60 32 13 24 9 2.4 0.8
44.8292 5.1598 0.1930 4.3672 0.6549 4.4605 0.1093 2.0696 6.1083 0.9505 6.0060 5.2018 0.7754 0.7814 0.0104 0.0027 0.0624 0.0033 0.0015 0.4898 0.0025 0.1741 0.0279 0.0306
Legend
8
Languages (excluding English)
Equal representation (approx) Indo-European Sino-Tibetan Austronesian Afro-Asiatic Turkic Niger Congo Other
6
de ru ja fr es
zh
4
pt
2
eu
0
ca
gl
ko fa el az su sw as am ha
101
100
id
ar hi bn
102
103
Total speakers (millions, log scale)
Figure 6: Distribution of the selected 23 languages (English excluded for a more clear visualisation) by share of data in CommonCrawl (CC-MAIN-2025-43) and estimated number of speakers based on Wikipedia sources.
C
Prompts for LLM-as-a-Judge
To process the open-ended answers produced by the LLMs, we employ a secondary model as a judge to interpret the responses and extract the relevant information. The objective is to associate each question with a set of countries in order to uncover potential cultural biases. Therefore, the judge model is tasked with identifying up to five regions or places mentioned or implied in each answer. If no such information could be reliably determined, the model is prompted to explicitly indicate that no inference is possible (non-answered responses). Basic eval. Inf.(2) Mid.(3)
Thorough eval. Exp.(1) Inf.(2) Mid.(3)
Lang.
Exp.(1)
en es ca eu
93.94 89.39 93.94 84.85
92.42 90.91 96.97 96.97
100 96.97 100 96.97
92.42 77.27 56.06 59.09
89.39 78.79 46.97 62.12
100 96.97 96.97 86.36
Avg.
90.53
94.32
98.48
71.21
69.32
95.08
Table 7: Evaluation Accuracy results for LLM as-ajudge in Basic and Thorough evaluations. System prompt types: (1) Explicit, (2) Infer and (3) Middle.
Table 6: Distribution of the selected 24 languages (ISO 639-1:2002) included on our dataset by share of data in CommonCrawl (CC-MAIN-2025-43) and estimated number of total speakers (in millions) based on Wikipedia sources.
We evaluated three prompting strategies for the judge model: (1) Explicit, in which only regions explicitly mentioned in the text were extracted; (2) Infer, which allowed the model to extract regions that were either explicitly stated or reasonably inferred; and (3) Middle, a balanced approach that prioritized explicit mentions while permitting limited, non-speculative inference when necessary (see Figures 7 and 8). For the evaluation, we conducted a manual assessment of 264 items (one per subtopic across
Figure 6 illustrates the relationship between the selected languages (excluding English), showing the correlation between number of speakers and their share in CommonCrawl. The broad dispersion across the plot demonstrates that CROQ covers a wide range of linguistic scenarios, spanning more than six distinct language families. 13
four languages), covering both high-resource (en, es) and low-resource (ca, eu) settings. We considered two evaluation setups: a basic evaluation, in which the judge model extracted only the country or region mentioned in the output of another model, and a thorough evaluation, in which the model additionally extracted the specific region when explicitly mentioned. To account for finergrained geographic specificity, we report results using the thorough evaluation setting (see Table 7). The manual evaluation showed that the (3) Middle prompting strategy yielded the most accurate and consistent results. Based on this assessment, we adopted this prompt in all subsequent analyses, using gpt-4o-mini from OpenRouter with default settings as the underlying judge model.
D
distribute references across countries. The right figure reports country references aggregated by language, illustrating how country mentions vary across the linguistic dimension. D.1
Analysis by language
The results in Table 9 illustrate consistent patterns in how frontier models associate languages with countries. Across all 24 languages, references to countries where the language is an official language (Own) dominate the model’s outputs. This behavior suggests a heavy reliance on language–country cooccurrence signals, likely reflecting strong correlations within the training data. While such associations are linguistically plausible, their prevalence indicates limited decoupling between a language and its primary national context. Beyond these self-referential patterns, the model exhibits a stable hierarchy regarding non-official country references. Japan and the United States are the most frequently referenced nations across all languages, followed by India, France, and China. Importantly, these references appear even in languages with limited cultural or geographic ties to these nations, suggesting that cross-country mentions are driven more by general global prominence than by language-specific relevance. This finding aligns with prior observations that multilingual models tend to overrepresent culturally salient entities regardless of the linguistic context. Examining the Own and NA columns reveals distinct behaviors based on language resource availability. The Own count is highest for lowresource languages such as Assamese, Sundanese, and Bengali, indicating a strong tendency for the model to ground these languages almost exclusively in their home countries. In contrast, highresource languages like English, Chinese, and Russian show comparatively lower Own counts, reflecting the model’s ability to distribute references more broadly. Additionally, while NA (missing response) counts are generally low, spikes in languages like Hausa and Hindi indicate occasional gaps in model generation for specific low-resource contexts. The O+N column, which aggregates Own and NA percentages, further highlights this divide. Languages such as Bengali, Sundanese, Amharic, Basque, and Hindi exhibit the highest O+N values, reflecting a concentration of outputs that are either biased toward the home country or result in nonresponses—traits characteristic of low-resource processing. Conversely, high-resource languages
Frontier models outputs for 24 languages
As an initial analysis, we examined the outputs of frontier models across seven widely spoken languages and identified a pronounced concentration in their regional references. Specifically, we analyzed the top three most frequently mentioned countries or regions for each language-model pair (see Table 8). Models overwhelmingly anchor their references to the primary country associated with each language, such as the U.S. for English, China for Chinese, France or Canada for French, and Japan for Japanese. This shows that models correctly internalize the conventional linguistic–national associations. Secondary references vary somewhat across models, but still follow intuitive cultural or regional links: for example, Arabic prompts frequently bring up Egypt, Saudi Arabia, and the UAE, while Spanish prompts often reference Mexico, Spain, and Peru. Figure 9 presents the outputs (excluding models’ own references) generated by the gpt-4omini, gemini-2.5-flash, claude-3.5-haiku, llama-4maverick, command-r-08-2024, magistral-small2506, qwen3-next-80b-a3b-instruct, and deepseekv3.2-exp model variants across the complete dataset covering 24 languages. The evaluation spans content associated with twelve countries: Japan (JP), the United States (US), India (IN), China (CN), France (FR), Italy (IT), the United Kingdom (GB), Germany (DE), Mexico (MX), Brazil (BR), Russia (RU), and Egypt (EG). The left figure reports country references aggregated by model, highlighting how different model variants 14
You are a precise geographical context classifier. Your task is to determine which nation, country or state(s) a given text is explicitly referring to. Follow these rules strictly: 1. Do not infer nations, countries or states based on general knowledge, stereotypes, or assumptions beyond what is stated in the text. 2. If multiple nations, countries or states are mentioned, list all of them. Not more than 5. Keep your answers concise and limited to the nation, country or state names. Translate the final answers into English.
Figure 7: Chosen System prompt, Middle (3), given to the Judge model You are a geographical context classifier. Identify the nation, country, or state(s) explicitly mentioned in the text. Rules: 1. Only list countries/states that appear in the text. Do not infer. 2. If multiple countries/states are mentioned, list up to 5. 3. Only list nations, countries, or states; skip regions or cities unless needed for clarity. Answer with country names only, in English, separated by commas.
You are a geographical context classifier. Identify the nation, country, or state(s) the text refers to. Rules: 1. If no country is mentioned, infer the most likely one from context. 2. If multiple countries/states are mentioned, list up to 5. 3. Only list nations, countries, or states; skip regions or cities unless needed for clarity. Answer with country names only, in English, separated by commas.
Figure 8: System prompt given to the Judge model (top: Explicit (1) ; bottom: Infer (2) ) Models Language English (en) Chinese (zh) Spanish (es) Arabic (ar) French (fr) Russian (ru) Japanese (ja)
Table 8: Top three countries most frequently referenced by model for each high-resourced language.
like English, French, and Russian display lower O+N values. This indicates a more diversified output distribution, consistent with languages that possess richer cross-country associations in the training corpus. Overall, a clear trend emerges: low-resource languages tend to yield highly concentrated or missing responses (high O+N), while high-resource languages yield more balanced references reflecting broader cultural associations.
ability. In this context, diversity measures how many unique countries are referenced at all, while normalized entropy captures the evenness of the distribution of mentions across those referenced entities. High-resource languages (e.g., English, Chinese, French, Russian) exhibit both high diversity and moderate-to-high entropy, indicating that they not only reference a wider array of nations but also distribute their focus more uniformly. In contrast, low-resource languages (e.g., Hindi, Bengali, Amharic, Sundanese) demonstrate low diversity and entropy; they reference fewer unique countries,
Finally, the analysis of the relationship between normalized diversity and entropy scores may indicate a correlation with language resource avail15
Mentions by Model (24 lang)
GPT
Mentions by Language (8 Frontier Models)
2500
Gemini 2000
Claude Llama
Countries
1500
CommandR 1000
Magistral Qwen
500
Deepseek
JP US IN CN FR IT GB DE MX BR RU EG
JP US IN CN FR IT GB DE MX BR RU EG Countries
2500 2000 1500 1000 500 0
en zh hi es ar fr bn pt ru id de ja ko fa sw ha az am su el as ca gl eu Language
Figure 9: Outputs (without own references) from gpt-4o-mini, gemini-2.5-flash, claude-3.5-haiku, llama-4-maverick, command-r-08-2024, magistral-small-2506, qwen3-next-80b-a3b-instruct and deepseek-v3.2-exp variants over the full dataset of 24 languages. Countries: Japan (JP), USA (US), India (IN), China (CN), France (FR), Italy (IT), UK (GB), Germany (DE), Mexico (MX), Brazil (BR), Russia (RU), Egypt (EG). Data from Tables 1 and 9 respectively.
Lang.
Own (%)
NA
O+N
JP
US
IN
CN
FR
IT
GB
DE
MX
BR
RU
EG
Div
Ent
en zh hi es ar fr bn pt ru id de ja ko fa sw ha az am su el as ca gl eu
3,223 (.31) 5,110 (.48) 7,430 (.70) 6,935 (.66) 7,325 (.69) 5,956 (.56) 8,219 (.79) 7,719 (.73) 3,831 (.36) 8,055 (.76) 7,359 (.70) 6,213 (.59) 6,191 (.59) 6,015 (.57) 8,196 (.78) 5,385 (.51) 7,960 (.75) 8,088 (.77) 8,331 (.79) 7,393 (.70) 8,732 (.83) 7,622 (.72) 7,967 (.75) 8,266 (.78)
192 258 1,706 152 499 136 1,189 358 173 870 71 375 1,580 722 144 3,270 253 1,133 995 253 140 158 88 903
.32 .51 .87 .67 .74 .58 .89 .76 .38 .85 .70 .62 .74 .64 .79 .82 .78 .87 .88 .72 .84 .74 .76 .87
2,893 2,064 94 843 504 832 73 479 1,423 411 485 5,584* 1,064 473 230 30 123 43 167 330 41 213 314 132
2,028* 1,460 104 588 396 557 151 442 657 246 237 736 1,308 544 389 99 212 106 146 307 53 339 297 144
695* 642 7,423* 130 182 215 2,415* 167 265 129 169 434 574 263 85 30 81 17 49 161 2,900* 170 129 48
367 4,781* 92 173 85 236 59 161 431 133 203 219 663 194 79 44 64 19 76 135 42 224 143 66
265 407 37 110 197 4,656* 35 147 199 59 145 267 508 266 106 25 43 13 48 173 18 306* 135 674*
223 193 13 128 82 225 17 146 214 24 304 98 227 101 55 13 34 19 14 124 6 103 58 37
323* 319 38 48 136 85 41 94 118 35 48 144 229 107 156 21 47 9 15 141 17 31 77 39
277 210 21 100 63 162 20 124 196 46 6,927* 98 214 124 72 16 48 6 32 132 8 45 68 40
309 165 14 2,852* 59 95 13 85 71 29 62 133 123 55 19 3 23 12 10 46 2 142 259 38
314 202 11 186 64 137 16 6,199* 143 48 53 193 160 89 48 9 10 5 7 36 4 50 292 15
57 80 14 23 21 44 10 30 3,827* 10 45 21 97 59 17 1 119 8 6 34 4 17 23 15
104 119 11 58 1,645* 104 13 50 90 30 39 47 124 105 106 37 29 22 18 83 6 51 21 13
157 142 56 136 108 149 68 114 146 78 122 114 112 117 89 49 78 38 62 106 41 121 113 68
0.69 0.57 0.14 0.65 0.71 0.54 0.20 0.42 0.60 0.26 0.38 0.36 0.47 0.48 0.50 0.10 0.43 0.13 0.11 0.36 0.30 0.49 0.37 0.37
Avg.
6,980 (.66)
651
.72
577
414
197
170
153
102
87
92
77
82
33
56
99
0.40
Table 9: Frontier Model outputs for 24 languages. Own: counts references to countries in which the language is an official language. NA: denotes missing responses. O+N: percentage of the sum of Own and NA references relative to the total for the language. Referenced countries: Japan (JP), USA (US), India (IN), China (CN), France (FR), Italy (IT), United Kingdom (GB), Germany (DE), Mexico (MX), Brazil (BR), Russia (RU), Egypt (EG). Div: diversity scores. Ent: normalized entropy values. Bold values indicate the highest count for each language, excluding self-references. An asterisk (*) marks references to own countries.
and those references are heavily concentrated on a single entity (typically own countries) rather than being spread out. This suggests that the model’s capacity to generate broad, evenly distributed cultural references is heavily conditioned by the availability of the underlying training data. D.2
languages and eight frontier models. A key aspect of this metric is that it measures model diversity across different regions per language, meaning that for each language in our set, we calculate how many unique countries or regions are represented in the model’s cultural references. A higher score indicates that a model’s outputs for a given language draw from a broader, more globally varied set of cultural contexts, rather than being concentrated in one or two regions. Averaged over languages, we observe large
Measuring Cultural Diversity in Model Outputs
Table 10 reports diversity scores, the number of distinct geographic regions referenced across 24 16
D.3
model-level differences: CommandR achieves the highest overall diversity, followed by DeepSeek and Gemini, while Mistral, GPT, and Qwen exhibit substantially lower scores. Language-level variation is also pronounced. High-resource languages such as English, French, Russian, and Spanish consistently receive broadest cultural coverage across models. In contrast, several low-resource or regionally concentrated languages (including Amharic, Assamese, Hausa, Hindi, and Bengali) remain well below the global mean, indicating narrower cultural framing in model outputs.
We analyzed the correlation between the percentage of CommonCrawl (CC) data for each language and the diversity of model outputs across all Frontier models. Table 11 reports Pearson’s r and Spearman’s rs correlations, along with their associated p-values.
Beyond aggregate trends, we find substantial cross-model variance within the same language, suggesting that cultural diversity is not determined by language alone. For example, Chinese ranges from 98 (Qwen) to 188 (CommandR), Bengali from 20 (Qwen) to 116 (LLaMA), and Hindi from 18 (Claude) to 107 (DeepSeek). Such differences often exceed average inter-language gaps, underscoring the impact of model-specific training and alignment choices. Overall, these results show that cultural diversity in multilingual generation is highly uneven, favoring high-resource languages, but also that these disparities are model-dependent rather than unavoidable.
Language 160 127 33 112 89 143 36 91 141 72 94 90 73 98 68 34 77 19 41 87 31 102 96 55
171 165 69 131 126 154 75 134 156 118 113 127 120 142 109 74 109 53 77 100 30 115 120 115
159 125 18 127 88 135 28 105 147 89 150 129 145 91 87 31 56 43 64 109 22 140 109 94
135 157 58 141 101 130 116 78 149 56 118 127 99 101 110 86 79 74 30 101 97 126 88 83
170 188 52 149 137 174 115 143 179 53 137 130 99 167 134 34 77 43 100 162 30 171 163 64
145 137 60 121 95 143 64 117 120 56 119 94 149 115 85 11 61 11 82 91 30 97 89 28
152 98 47 134 110 137 20 109 118 62 132 92 75 88 46 35 55 25 26 75 23 105 100 36
165 143 107 169 115 173 90 139 158 115 112 123 137 136 75 86 109 33 77 127 64 114 140 73
157 142 56 136 108 149 68 114 146 78 122 114 112 117 89 49 78 38 62 106 41 121 113 68
ALL
82
113
95
102
120
88
79
116
99
Model
Pearson r p
Spearman rs p
Qwen Gemini DeepSeek GPT Magistral Claude Llama Command-r Avg_all
0.529 0.483 0.425 0.577 0.426 0.452 0.383 0.329 0.498
0.808 0.798 0.775 0.774 0.770 0.760 0.650 0.629 0.843
0.008 0.017 0.038 0.003 0.038 0.027 0.064 0.116 0.013
1.80 × 10−6 3.00 × 10−6 8.90 × 10−6 9.14 × 10−6 1.07 × 10−5 1.63 × 10−5 5.84 × 10−4 9.87 × 10−4 2.31 × 10−7
Table 11: Correlation between CommonCrawl percentage and output diversity for Frontier models. Pearson’s r and Spearman’s rs are reported along with their pvalues.
Across all models combined (Avg_all), we observed a Pearson correlation of 0.498 (p = 0.013), indicating a moderate positive linear relationship, and a Spearman correlation of 0.843 (p < 10−6 ), showing a very strong rank correlation. These results suggest that languages with a higher proportion of CC data tend to exhibit greater diversity in model outputs. When examining individual models, the correlations varied but largely supported this trend. For example, GPT-4o-mini showed the strongest Pearson correlation (0.577, p = 0.003) and a high Spearman correlation (0.774, p < 10−5 ), indicating both linear and rank-based relationships between data availability and diversity. Qwen3-next-80b-a3binstruct also displayed high correlations (Pearson 0.529, Spearman 0.808), while Gemini-2.5-flash and Claude-3.5-haiku exhibited moderate correlations. Deepseek-v3.2-exp and Magistral-small2506 presented slightly lower but still statistically significant correlations, reinforcing the overall pattern. Two models, however, showed weaker correlations. LLaMA-4-maverick had a Pearson correlation of 0.383 (p = 0.064) and Spearman 0.650 (p = 5.84 × 10−4 ), while Command-R-08-2024 showed Pearson 0.329 (p = 0.116) and Spearman 0.629 (p = 9.87 × 10−4 ). Although the Spearman correlations remain significant, the lower Pearson
AVG
en zh hi es ar fr bn pt ru id de ja ko fa sw ha az am su el as ca gl eu
Correlation between CC and output diversity
Table 10: Diversity scores by language for different Frontier API models. ALL: per-model average across languages. AVG: per-language average. Highlighted results: above average results.
17
correlations suggest that the linear relationship is less pronounced for these models, possibly reflecting model-specific factors or limitations in handling low-resource languages. Overall, these results indicate that the amount of possibly available training data for each language may be an important factor influencing output diversity. Languages with less CC data tend to yield less diverse outputs, suggesting a potential reduction in creative or varied responses when data is scarce. This pattern is robust across most models, highlighting the influence of data coverage on multilingual model behavior.
in normalized entropy, indicating a loss of cultural diversity. Notably, this trend holds across different architectures and parameter scales, including models developed in non-Western contexts (e.g., Qwen with China), suggesting that instruction tuning exerts a homogenizing effect that overrides differences introduced during pretraining. Diversity metrics corroborate qualitative trends. The observed redistribution is quantitatively supported by the diversity and entropy metrics. While diversity scores (Div) increase for instruction-tuned models due to higher total counts concentrated in fewer regions, entropy consistently decreases, reflecting a more peaked and less balanced distribution.
E Open-Weight Base and Instruct Models Analysis Over English Dataset
Possible implications. Taken together, these results demonstrate that instruction tuning systematically reduces cultural diversity in model outputs, steering responses toward a limited set of culturally dominant perspectives, particularly those associated with the United States and Japan. This finding has important implications for the deployment of instruction-tuned models in cross-cultural or global applications, where preserving diverse cultural viewpoints may be critical.
Table 12 presents a comparative analysis of cultural leanings exhibited by base and instruction-tuned Open Weighted LLMs when responding to openended cultural questions in English. Responses are categorized by their alignment with references to specific countries, allowing us to examine how pretraining and instruction tuning affect the geographic distribution of model outputs. We additionally report diversity (Div) and normalized entropy (Ent) scores to quantify the spread and balance of cultural references.
F
Base models exhibit broader cultural distributions. Across all model families, base models demonstrate relatively more balanced distributions of cultural references. While the United States is frequently the most represented region, other countries (particularly Japan, India, China, and several European nations) appear with substantial frequency. This pattern is reflected in higher normalized entropy values, indicating a comparatively more diverse set of cultural alignments than their instructed counterparts.
Models’ Analysis Across Post-training Stages
Table 13 reports a controlled comparison between base, supervised fine-tuned (SFT), and instructionaligned (Inst.) variants of open-weight models that provide multiple alignment stages. The analysis focuses exclusively on English prompts, enabling us to isolate the effect of alignment from multilingual prompting effects. By including OLMo-2 (7B, 32B) and OLMo-3 (7B, 7B-Think), which expose intermediate SFT checkpoints, the table allows us to disentangle the impact of supervised fine-tuning from that of final instruction alignment.
Instruction tuning induces strong cultural concentration. In contrast, instruction-tuned variants exhibit a pronounced shift toward a narrow set of cultural references. Across all examined model families, instruction tuning results in a sharp increase in references to Japan and the United States, often making these the two dominant regions by a large margin. On average, references to Japan and the US increase in instruction-tuned models, while references to India, China, Russia, Egypt, and Latin American countries decline substantially. This concentration is accompanied by a consistent reduction
Base models: higher entropy and weaker geographic priors. Across all architectures and scales, base models exhibit the highest entropy values and comparatively balanced country distributions. While some geographic preferences are already present (notably toward the US), these preferences remain moderate and are spread across a broader set of countries. For example, OLMo-2-7B Base achieves the highest entropy in the table, indicating minimal concentration on any single country. However, base models also show higher rates of 18
Model Llama-3.1-8B Llama-3.1-70B gemma-2-9b gemma-2-27b Mistral-7B Qwen2.5-7B Qwen2.5-32B Qwen2.5-72B Average
NA
JP
US
IN
CN
FR
IT
GB
DE
MX
BR
RU
EG
Div
Ent
Base Inst. Base Inst.
354 14 783 13
164 407 237 417
144 391 194 333
200 121 199 125
85 25 116 35
87 18 85 15
106 22 61 9
30 55 36 42
62 14 80 7
42 10 37 15
82 32 85 19
54 9 51 2
17 14 40 15
109 174 117 174
0.78 0.66 0.79 0.65
Base Inst. Base Inst.
30 19 28 9
124 269 100 250
317 217 202 285
169 104 128 58
80 18 71 23
53 33 38 16
26 41 43 27
18 69 75 30
45 29 31 26
43 14 23 31
26 55 33 49
12 2 24 1
25 29 29 14
95 136 80 154
0.70 0.71 0.77 0.73
Base Inst.
52 14
98 233
146 580
93 86
73 34
41 18
38 45
41 57
46 14
45 22
28 58
40 1
31 11
91 160
0.81 0.63
Base Inst. Base Inst. Base Inst.
265 16 248 12 182 10
98 321 183 489 180 343
385 304 399 247 606 386
98 74 131 58 150 106
32 310 35 140 67 192
51 15 36 37 31 31
64 50 19 18 28 22
34 20 39 24 55 15
35 21 15 17 29 29
18 5 31 11 34 6
34 11 15 25 40 16
9 5 8 4 6 1
7 10 5 7 25 11
147 150 142 118 144 127
0.71 0.62 0.65 0.61 0.66 0.60
Base Instruct
243 13
148 341
299 343
146 92
70 97
53 23
48 29
41 39
43 20
34 14
43 33
26 3
22 14
116 149
0.73 0.65
Table 12: Comparison of named base and instruct models over English questions (data used for Figure 3). NA: denotes missing responses. Referenced countries: Japan (JP), USA (US), India (IN), China (CN), France (FR), Italy (IT), United Kingdom (GB), Germany (DE), Mexico (MX), Brazil (BR), Russia (RU), Egypt (EG). Bold values indicate the highest count for each country per model. Div: diversity scores. Ent: normalized entropy values.
Diversity and entropy: misleading signals. While diversity scores (Div) often increase after SFT and instruction tuning, this increase does not correspond to more balanced outputs. Instead, higher diversity frequently co-occurs with lower entropy, reflecting broader but highly uneven country coverage dominated by a few regions.
missing or underspecified responses (NA), particularly in OLMo-3, reflecting lower controllability rather than alignment-induced bias. Supervised fine-tuning sharply amplifies cultural bias. The transition from Base to SFT introduces the largest shift in geographic skew across all models. SFT variants consistently display dramatic increases in counts for a small number of dominant countries (the United States and Japan) accompanied by substantial drops in normalized entropy. For instance, OLMo-2-7B SFT increases US-associated outputs, while entropy drops. Similar patterns appear in both OLMo-3 variants, where Japan becomes the dominant country after SFT. These results indicate that supervised alignment data strongly reinforces cultural priors present in instruction datasets, amplifying bias beyond what is observed in pretraining alone.
G
Prompt Analysis for RQ3
G.1
Prompting for Base Models
Since foundational Large Language Models (LLMs) are originally trained with a selfsupervised next-token prediction objective, they are not inherently designed to perform specific QA tasks without guidance. Unlike instruction-tuned models, base models often lack the alignment necessary to interpret a raw query as a command to provide a factual answer. To address this limitation, we employ In-Context Learning (ICL), a method that leverages explicit reasoning over background knowledge to guide the model’s inference process (Brown et al., 2020). By providing the model with relevant context and demonstrations within the input window, we effectively transform the generative completion task into a structured QA format, allowing the base model to infer the desired output pattern without the need for parameter updates. We operationalize this approach using the
Instruction tuning moderates but does not eliminate alignment bias. Instruction-tuned models partially attenuate the extreme peaks introduced during SFT, but do not recover the diversity of base models. In several cases, entropy remains close to or below SFT levels, suggesting that dominantcountry preferences become stabilized during alignment. Overall, instruction tuning acts as a smoothing mechanism rather than a corrective one. 19
Model Olmo-3-7B
NA
JP
US
Base 608 35 186 SFT 85 395 212 Inst. 41 478 171
IN CN 43 97 89
Base 608 35 186 43 Olmo-3-7B-Think SFT 88 453 213 105 Inst. 200 386 289 179
FR IT GB DE MX BR RU EG Div
Ent
27 25 18 26 109 64 19 86 78
28 39 43
13 30 39
12 28 44
11 22 11
7 26 27
9 114 0.74 10 145 0.66 5 124 0.64
27 40 52
25 18 18 32 32 46
28 32 44
13 22 36
12 32 47
11 22 29
7 2 2
9 114 0.74 13 121 0.66 25 127 0.68
OLMo-2-7B
Base 109 64 97 SFT 252 130 520 Inst. 53 163 566
84 67 79
76 29 34
46 57 54 40 65 62
38 52 31
35 13 20
30 18 20
29 35 29
26 4 11
35 102 0.84 13 151 0.63 18 158 0.64
OLMo-2-32B
Base 155 61 216 119 SFT 33 361 390 51 Inst. 44 221 368 81
54 20 32
38 44 62 36 74 40
34 34 42
34 29 33
28 16 18
26 31 29
17 2 8
9 88 0.77 16 151 0.65 6 161 0.71
Table 13: Comparison of Base, SFT, and Instruct variants over English questions. NA: denotes missing responses. Referenced countries: Japan (JP), USA (US), India (IN), China (CN), France (FR), Italy (IT), United Kingdom (GB), Germany (DE), Mexico (MX), Brazil (BR), Russia (RU), Egypt (EG). Bold values indicate the highest count for each country per model. Div: diversity scores. Ent: normalized entropy values.
Question: Who wrote '1984'? Answer: George Orwell. Question: What is the chemical symbol for water? Answer: H2O. Question: What is the dominant religion (in region/place)? Answer: I'll choose {Country1}. The dominant religion in {Country1} is {Country1_Religion}.
Figure 11: Outputs from the Base and Instruct variants of Olmo3-7B, Olmo2-7B and Olmo2-32B models. The Instruct models are shown in two modes: Ins.(NC), which uses the same prompt as the Base model, and Ins., which uses the standard chat template.
Question: What sacred texts are studied (in region/place)? Answer: I'll choose {Country2}. In {Country2}, the sacred texts studied depend primarily on the religious community, as the country has a diverse religious landscape. Question: {Input Question} Answer:
the model’s generation, ensuring it focuses on the relevant cultural entity.
Figure 10: Prompt given to the Base models in order to follow instructions.
To maintain a rigorous evaluation, we implement a post-processing filter once the judge model extracts the mentioned regions or countries from the output. We specifically identify instances where the LLM simply replicates the example regions provided in the few-shot prompt rather than generating an original response to the cultural query. These cases are classified as "Non-Answered," as they represent a failure of the model to generalize beyond the immediate context. Consequently, all reported counts and statistics in our results are derived exclusively from countries that were not introduced as examples in the prompt, ensuring our findings reflect the model’s internal cultural associations.
prompt structure illustrated in Figure 10. To ensure the model adheres to the correct output format, the prompt includes both general QA demonstrations and task-specific examples that explicitly condition the model to respond with a country or geographic region. The template is constructed with strategic placeholders to facilitate this: two slots are reserved for preselected countries—including a specific attribute for religion in the first instance to demonstrate cultural reasoning, while a final placeholder is designated for the target culture-related question. This few-shot setup serves to constrain 20
G.2
Impact of the prompt
To assess the role of prompt formulation in eliciting instruction-following behavior, we conduct controlled experiments in which prompts designed for base models are applied to instruction-tuned models. We examine whether such prompt mismatches lead to systematic differences in generation, focusing on the instruction-tuned variants of OLMo-2 (7B, 32B) and OLMo-3 (7B). To evaluate how prompting affects Base versus Instruct models, we prompt the Instruct variants with the same input as the Base models (Ins.(NC)) across Olmo3-7B, Olmo2-7B and Olmo2-32B variants (see Figure 11). The 32B model consistently shows more stable, yet also more biased, behavior than the 7B models, suggesting greater robustness of bigger models to prompt-formatting differences. Across settings, the USA and Japan frequently receive higher values across all settings, reflecting biases similar to those observed in prior models. Overall, while the prompt can influence responses, the main conclusion holds: cultural biases remain largely consistent, particularly toward Japan, the USA, and India, reinforcing the hypothesis that such biases stem from the models’ training data.
21