CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords Yifan Wang1 , Junyu Lu2 , Qifan Wang3 , Shun Zhang1∗ , Chaozhuo Li4 , Jiahao Liu5 , Zhijun Cao1 , Lingbin Bu6 , Fanliang Bu1
arXiv:2609.21722v1 [cs.CL] 18 Sep 2026
1
People’s Public Security University of China 2 Dalian University of Technology 3 Meta AI 4 Beijing University of Posts and Telecommunications 5 Meituan 6 Beijing Police College
Abstract Chinese social media has generated a vast and continually evolving lexicon of internet buzzwords whose meanings are often non-literal and deeply rooted in local cultural and pragmatic contexts. Existing research has primarily focused on interpreting these buzzwords within Chinese, leaving largely unexplored whether LLMs can transfer such culturally grounded knowledge across languages and accurately convey the intended meanings in English. This cross-lingual capability is also critical for safety, as harmful expressions may obscure their offensive content through culture-specific homophony, euphemism, irony, or coded language. In this paper, we investigate the ability of advanced LLMs to understand Chinese internet buzzwords across languages. To this end, we introduce CIBuzzBench, the first benchmark for cross-lingual Chineseto-English understanding of Chinese internet buzzwords. CIBuzzBench comprises 3,001 Chinese internet buzzwords annotated with English meaning explanations, English equivalents, category labels, and harmfulness labels. Based on these annotations, we design three evaluation tasks: Meaning Explanation, Cross-lingual Equivalent Matching, and Culturally Grounded Harmfulness Detection. We evaluate representative state-of-the-art proprietary and Chinese LLMs under both English- and Chinese-prompting settings. Our results show that LLMs continue to struggle with the cross-lingual understanding of Chinese internet buzzwords, particularly in finegrained non-literal interpretation, robust equivalent matching under option perturbations, and calibrated harmfulness detection. These findings highlight the persistent challenges posed by culturally grounded language phenomena for multilingual LLMs and safety-oriented evaluation. The dataset and code are available at https://github.com/SuperYFan/CIBuzzBench.
Introduction Chinese internet buzzwords serve as compact carriers of culturally and pragmatically situated meaning. They arise through a variety of mechanisms, including phonetic wordplay, abbreviations, references to popular culture, community-specific slang, figurative language, and experience-based labeling. Their interpretation is often determined less by literal lexical content than by platform∗
Corresponding author: [email protected]. [email protected]; [email protected]; [email protected]; [email protected]; [email protected]; [email protected]; [email protected]; [email protected].
Figure 1: Illustration of CIBuzzBench tasks with two examples. Each Chinese internet buzzword is paired with an English meaning explanation, an equivalent-selection question, and a harmfulness label. specific contexts and shared background knowledge (Yang 2023; Kulkarni and Wang 2017). Consequently, these expressions can be challenging even for native Chinese speakers, as their meanings are frequently non-compositional, contextdependent, and subject to rapid change. The challenge is further amplified in cross-lingual transfer to English: a model must not only recover the intended meaning in Chinese but also identify an idiomatic English expression that preserves the source expression’s semantics, tone, and pragmatic force. Among multiple plausible candidates, a literal translation or a semantically related expression may appear reasonable while nevertheless failing to serve as an appropriate English counterpart. This capability is important for both cross-cultural communication and content safety (Deng et al. 2022; Lu et al. 2023; Xiao et al. 2024; Bai et al. 2025). English explanations and equivalents can help non-Chinese speakers understand
Chinese online discourse and reveal whether LLMs can transfer culturally specific meanings across languages (Ma et al. 2025). This is especially relevant to safety, as harmful meanings are often conveyed through homophones, euphemisms, sarcasm, and other indirect forms that are difficult to detect through surface matching (Zhang et al. 2024; Guo et al. 2025; Li et al. 2025). Although existing resources provide a foundation for Chinese language and safety research, they do not directly test whether LLMs can identify the most appropriate English equivalent rather than a literal or contextually misleading alternative. These challenges ultimately hinge on a model’s ability to preserve the intended meaning, pragmatic force, and safety-relevant implications of a buzzword during cross-lingual interpretation. Therefore, systematically evaluating LLMs’ cross-lingual understanding of Chinese internet buzzwords is essential for assessing their ability to interpret culturally grounded language and reliably transfer such knowledge across languages. To address this gap, we introduce CIBuzzBench, the first benchmark for Chinese-to-English cross-lingual understanding of Chinese internet buzzwords. CIBuzzBench contains 3,001 Chinese buzzwords, each annotated with a Chinese explanation, an English meaning explanation, a category label, a harmfulness label, an English equivalent, and three carefully designed English distractors. It defines three complementary tasks: Meaning Explanation, Equivalent Selection, and Harmfulness Detection. As illustrated in Figure 1, these tasks evaluate whether LLMs can express non-literal meanings in English, select the intended English counterpart among distractors, and identify harmfulness. We evaluate representative LLMs under English- and Chinese-prompting settings and observe persistent limitations across all three tasks. Models often miss fine-grained non-literal meanings, confuse intended English equivalents with literal translations or pragmatically adjacent alternatives, and overlook culturally implicit harmfulness. In Equivalent Selection, errors concentrate on literal and pragmaticneighbor distractors, suggesting that partial understanding of a Chinese expression does not guarantee precise alignment with an appropriate English counterpart. Performance also varies by formation mechanism: quotation-based expressions are especially challenging for Equivalent Selection, whereas homophonic and stylistic expressions are more difficult for Harmfulness Detection. Our contributions are as follows: • We introduce CIBuzzBench, the first benchmark dedicated to Chinese-to-English cross-lingual understanding of Chinese internet buzzwords, with 3,001 entries across six categories and a particular focus on mapping them to English counterparts that preserve meaning, tone, and pragmatic function. • We formulate three tasks, Meaning Explanation, Equivalent Selection, and Harmfulness Detection, to evaluate whether LLMs can explain Chinese internet buzzwords in English, select appropriate English counterparts, and detect harmful usage in cross-lingual settings. • We benchmark representative LLMs under English and Chinese prompts, revealing persistent errors in non-literal interpretation, sensitivity to prompt language and option
order, distractor discrimination, and culturally grounded harmfulness detection.
Related Work Chinese Internet Buzzword Understanding Recent studies examine whether large language models (LLMs) can understand Chinese internet buzzwords and related online expressions, whose meanings often emerge from user-generated content, quotations, homophony, abbreviations, stylistic imitation, and community-specific pragmatic conventions (Sravanthi et al. 2024). CHEER (Huang et al. 2025) constructs a dataset of Chinese internet buzzwords with definitions and user-generated contexts and evaluates whether LLMs can generate accurate Chinese definitions from such evidence. CHIME (Xie et al. 2025) evaluates Chinese Internet meme explanation, including meaning explanation, origin identification, example generation, and contextual meme selection. More broadly, benchmarks on emerging internet concepts and informal language, such as SLANG (Mei et al. 2024) and OpenSub-Slang (Sun et al. 2024), show that these expressions remain challenging because their meanings shift across time, communities, and usage contexts. Existing work, however, mainly focuses on monolingual understanding through Chinese definition generation, meme explanation, or contextual slang detection.
Cross-Lingual Pragmatic Interpretation Cross-lingual language understanding has been widely studied through machine translation, multilingual question answering, and cultural knowledge evaluation (Hu et al. 2020; Ruder et al. 2021; Shi et al. 2024; Chiu et al. 2025). These studies advance multilingual generalization and culturally grounded reasoning but do not fully capture the difficulty of transferring non-literal, pragmatically loaded expressions across languages, especially when cultural knowledge is implicit. ChID (Zheng, Huang, and Sun 2019) formulates Chinese idiom understanding as cloze-style reading comprehension, while CHENGYU-BENCH (Fu et al. 2025) evaluates idiom connotation, usage appropriateness, and open-ended contextual completion. SlangDIT (Liang et al. 2025) studies slang detection, cross-lingual explanation, and contextaware translation, emphasizing intermediate interpretation rather than direct form-to-form transfer. Chinese internet buzzwords pose a related but more dynamic challenge because their meanings are often not compositionally recoverable from surface forms but arise from homophony, quotation, abbreviation, metaphor, sarcasm, stylistic devices, and online pragmatic conventions. Simply translating a buzzword’s literal form may therefore preserve its surface structure while losing its intended online meaning. Compared with prior work, CIBuzzBench targets Chineseto-English pragmatic understanding of internet buzzwords rather than monolingual Chinese explanation, relatively fixed idiom comprehension, or sentence-level slang translation. Its central task asks models to select an English counterpart that preserves meaning, tone, and pragmatic function among controlled literal, cultural-mismatch, and pragmaticneighbor distractors. Meaning Explanation and Harmfulness
Detection complement this task by assessing English semantic transfer and safety-relevant interpretation.
CIBuzzBench CIBuzzBench is a sense-anchored lexical-entry benchmark for evaluating whether LLMs can interpret Chinese internet buzzwords in English-facing settings. Each entry centers on one Chinese internet buzzword and one documented online sense specified by its Chinese explanation, together with a category label, English meaning explanation, English equivalent, and harmfulness label. We therefore evaluate the annotated buzzword sense rather than the use of the surface string. As shown in Figure 2, the construction process contains two stages: data collection and annotation. The first stage builds a source set of Chinese buzzwords and Chinese explanations, while the second stage converts these entries into task-specific annotations through LLM-assisted generation and human verification.
Label
Count
Percent
Category Stylistic device Quotation Experience Slang Homophonic pun Abbreviation
792 602 592 481 398 136
26.4% 20.1% 19.7% 16.0% 13.3% 4.5%
Harmfulness Non-harmful Harmful
2,574 427
85.8% 14.2%
Table 1: Label distribution in CIBuzzBench. The benchmark contains 3,001 Chinese internet buzzwords with category and harmfulness annotations.
Annotation Process
Data Collection We manually collect Chinese internet buzzwords and their corresponding Chinese explanations from public Chinese buzzword encyclopedia websites,1 and also consult existing Chinese buzzword, meme, and toxicity resources (Huang et al. 2025; Xie et al. 2025; Bai et al. 2025). We carefully select entries one by one and retain only expressions whose Chinese explanations are sufficiently clear to support crosslingual annotation. A surface string may have different meanings across domains or contexts, so we treat the collected Chinese explanation as the source-language semantic anchor and annotate the sense it documents. After manual collection and organization, CIBuzzBench contains 3,001 Chinese internet buzzwords. To characterize the linguistic and pragmatic mechanisms behind these buzzwords, we follow Xie et al. (2025) and assign each entry to one of the following six category labels: • Experience: buzzwords derived from individuals summarizing personal experiences or situations. • Quotation: buzzwords originating from historical stories, public events, movie plots, TV shows, games, livestreams, novels, or celebrity quotes. • Stylistic device: buzzwords crafted with rhetorical techniques such as metaphor, euphemism, irony, or sarcasm. • Homophonic pun: buzzwords created by replacing original characters or phrases with forms of similar or identical sounds. • Slang: buzzwords based on widely recognized colloquial expressions specific to a particular time, community, platform, or social context. • Abbreviation: buzzwords formed by shortening proper nouns or general phrases, including morpheme reductions, initialisms, and simplified spellings. These labels describe the dominant mechanism by which a buzzword acquires its meaning rather than its topical domain. Table 1 reports the category and harmfulness distribution. 1 https://gengbaike.cn/; https://yougengbaike.com/index.html; https://regengbaike.com/; https://hizdm.net/; https://ttseed.cn/.
As shown in Figure 2, task annotation starts from the manually collected Chinese buzzword–explanation pairs. GPT-5.5 is used as an annotation assistant to generate initial candidates for the three tasks: an English meaning explanation for Meaning Explanation, one English equivalent with three distractors for Equivalent Selection, and a harmfulness label for Harmfulness Detection. These candidates are only drafts, even though GPT-5.5 is also included in the model evaluation. In CIBuzzBench, an English equivalent is defined as an English expression synonymous with the documented Chinese buzzword sense. The drafts are then checked by three task-specific reviewers, all NLP graduate students, who are responsible for Meaning Explanation, Equivalent Selection, and Harmfulness Detection respectively. Annotator backgrounds and the training and iterative annotation procedures are detailed in Appendix A.1 and A.2, respectively. Reviewers reject annotations that are inaccurate, overly literal, ambiguous, pragmatically mismatched, or inconsistent with the Chinese explanation; rejected cases are re-analyzed and re-annotated. For polysemous expressions, the final gold annotations are tied to the documented sense, and alternative senses are treated as potential sources of model confusion rather than as changes to the label. Accepted cases are further confirmed by a seven-member group, and unresolved cases are decided by group discussion and vote. For Harmfulness Detection, a buzzword is labeled harmful only when the annotated common online usage carries offensive or attacking force toward a person, group, or identity; neutral discourse markers, selfdeprecation, non-targeted jokes, and negative affect alone are labeled non-harmful. Thus, GPT-5.5 is used as a productivity tool rather than an authority for benchmark labels, and the task datasets are human-confirmed. Independent human validation on stratified samples showed high consistency for Meaning Explanation, a 95.3% gold-selection rate and 91.7% three-way exact agreement for Equivalent Selection, and substantial agreement for Harmfulness Detection (Fleiss’ κ = 0.6920), with details provided in Appendix A.3.
Figure 2: Construction process of CIBuzzBench. In the data collection stage, we collect Chinese internet buzzwords and Chinese explanations from Chinese buzzword websites and assign category labels. In the task annotation stage, GPT-5.5 assists in generating annotations for the three tasks, which are then checked by task-specific reviewers; rejected cases are re-analyzed and re-annotated, while accepted cases are confirmed by the group and finalized as task datasets.
Task Design CIBuzzBench defines three tasks for Chinese-to-English cross-lingual understanding. The zero-shot prompt templates for the three tasks are provided in Appendix B.1. Meaning Explanation asks a model to produce a concise English explanation of the non-literal meaning of a Chinese internet buzzword. It evaluates whether a model can infer the intended Chinese meaning and express it naturally in English rather than relying on literal translation. Equivalent Selection presents the Chinese buzzword with four English options, including one English equivalent, defined as an English expression synonymous with the documented sense, and three controlled distractors. We design the distractors to probe three common confusion types: a literal or surface-form option that follows the wording of the Chinese expression, a keyword or cultural-mismatch option that is topically related but semantically wrong, and a pragmaticneighbor option that has a similar communicative function but differs from the annotated sense. Harmfulness Detection asks a model to infer from the Chinese buzzword alone whether its documented common online sense is harmful or non-harmful. We include it as a safety-oriented application of cross-lingual buzzword understanding, evaluating lexical and culturally grounded recognition of offensive or attacking meanings encoded indirectly through slang, homophony, euphemism, quotation, or irony.
Experiments Experimental Setup We conduct four groups of experiments. The main zeroshot evaluation uses six representative state-of-the-art API
models: GPT-5.5 (OpenAI 2026), Claude Opus 4.8 (Anthropic 2026), Gemini 3.1 Pro (Google DeepMind 2026), DeepSeek-V4-Pro (Xu et al. 2026), Qwen3.7-Max (Qwen Team 2026), and GLM-5.2 (Z.ai 2026). We then evaluate 2-shot prompting for GPT-5.5 and Gemini 3.1 Pro Preview, LoRA fine-tuning for Qwen3-8B (Qwen Team 2025) and GLM-4-9B (Z.ai 2025), and a Qwen model-size comparison using Qwen3-4B, Qwen3-8B, and Qwen3-14B. The zeroshot and model-size experiments use the full 3,001-entry benchmark, while the few-shot and fine-tuning experiments use a 4:1 train–test split with 2,401 training entries and 600 test entries. All tasks are evaluated under English and Chinese prompt settings, while the input buzzword is Chinese. We use deterministic decoding with temperature 0.0. For API-based runs, top-p and top-k are left at provider defaults; for local Qwen model-size and SFT runs, sampling is disabled with temperature 0.0, top-p is set to 1.0, and top-k is not used. Additional details, including prompt templates, data split, and results for the few-shot, fine-tuning, and model-size experiments, are provided in the Appendix. Local evaluation scripts and metric computation are run on an NVIDIA Tesla V100 GPU (32GB). Meaning Explanation asks each model to generate a concise English explanation of the non-literal meaning of a Chinese internet buzzword. We report BLEU, ROUGE-L, and BERTScore F1 as supplementary measures of reference similarity. Because valid English explanations can vary substantially in wording, our primary measures are Human and LLM-judge scores of semantic and pragmatic equivalence on a 0–5 scale. The LLM judge evaluates the full benchmark. For human evaluation and judge validation, we randomly
Model
BLEU ROUGE-L BERT-F1 Human LLM
Model
Macro F1
Harmful F1
English Prompt GPT-5.5 Claude Opus 4.8 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.7-Max GLM-5.2 Average
4.01 2.94 3.22 3.05 4.02 2.85 3.35
22.33 19.42 19.95 19.73 20.20 17.10 19.79
88.35 87.07 87.72 87.60 87.98 86.67 87.57
3.23 3.09 4.10 3.11 3.50 3.02 3.34
3.69 3.49 4.12 3.57 3.94 2.98 3.63
English Prompt GPT-5.5 Claude Opus 4.8 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.7-Max GLM-5.2 Average
76.23 71.09 80.45 75.36 74.00 72.94 75.01
60.71 49.47 66.17 58.94 53.94 52.85 57.01
Chinese Prompt GPT-5.5 Claude Opus 4.8 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.7-Max GLM-5.2 Average
4.05 3.58 4.07 2.20 4.66 3.73 3.72
22.92 20.94 21.37 18.63 22.54 21.70 21.35
88.60 87.83 88.25 87.21 88.43 88.30 88.10
3.13 3.03 3.90 3.06 3.84 3.61 3.43
3.62 3.36 4.04 3.46 3.95 3.72 3.69
Chinese Prompt GPT-5.5 Claude Opus 4.8 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.7-Max GLM-5.2 Average
77.41 70.41 79.86 74.68 76.10 78.39 76.14
62.32 48.61 64.96 56.50 58.04 62.35 58.80
Table 2: Meaning Explanation results under English and Chinese prompts. LLM-judge scores use the full benchmark, while Human scores use 200 sampled predictions per model and prompt language. Average rows report unweighted means across the six models. Values in bold indicate the best results, whilst those underlined indicate the second-best.
Table 4: Harmfulness Detection Macro F1 and harmful-class F1 under English and Chinese prompts. Average rows report unweighted means across the six models. Values in bold indicate the best results, whilst those underlined indicate the second-best.
Model
measures and interpret BLEU, ROUGE-L, and BERT-F1 as supplementary reference-similarity measures. The results show that models can often produce explanations related to the broad topic or affective tone of a buzzword, but still struggle to express the annotated online sense precisely in English. The main errors are not merely awkward wording; they often involve missing the non-literal stance, target, pragmatic force, or usage condition that distinguishes a buzzword from a literal translation or a generic paraphrase. Gemini 3.1 Pro performs best under both Human and full-benchmark LLM-judge evaluation, while Qwen3.7-Max is the strongest alternative under the LLM judge. Across the 400 sampled predictions per model, with 200 under each prompt language, Human and LLM-judge scores show substantial agreement, with an overall average QWK of 0.75; detailed per-model results are reported in Appendix C. For Equivalent Selection, Table 3 shows that models can identify an appropriate English synonym in the constrained multiple-choice setting. By removing the burden of openended generation and providing candidates, this task more directly evaluates whether models can distinguish the gold equivalent from literal, cultural-mismatch, and pragmaticneighbor distractors. Gemini 3.1 Pro and GPT-5.5 perform most consistently, but the remaining errors indicate that recovering the broad Chinese meaning does not always lead to precise semantic and pragmatic alignment in English. The average Macro F1 is higher under Chinese prompts than under English prompts, with the largest increases observed for Claude Opus 4.8, Qwen3.7-Max, and GLM-5.2, while DeepSeek-V4-Pro shows a small decrease. For Harmfulness Detection, Table 4 tests whether culturally grounded buzzword understanding supports safety decisions across instruction languages. Harmful-class F1 remains lower than Macro F1, showing that harmful senses are more
GPT-5.5 Claude Opus 4.8 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.7-Max GLM-5.2 Average
English Prompt (F1)
Chinese Prompt (F1)
87.12±0.66 79.94±1.81 88.03±0.36 79.66±0.48 82.38±0.26 82.27±0.48 83.23±0.26
87.83±0.18 83.17±2.41 88.41±0.49 78.91±0.54 84.80±0.42 84.40±0.23 84.59±0.50
Table 3: Equivalent Selection Macro F1 under English and Chinese prompts. Average rows report unweighted means across the six models. Values in bold indicate the best results, whilst those underlined indicate the second-best. sample 200 predictions from each model under the English prompt and another 200 under the Chinese prompt, giving 400 predictions per model, and obtain Human and LLMjudge scores on the same samples. The scoring rubric and full LLM-judge prompt are summarized in Appendix B.2. Equivalent Selection is a four-way multiple-choice task with one gold English equivalent and three controlled distractors. We report macro F1. To account for option-order effects, we generate five shuffled versions using seeds 1111, 2222, 3333, 4444, and 5555, evaluate models on the same versions, and report mean and standard deviation across seeds. Harmfulness Detection is a binary classification task in which models predict whether a Chinese internet buzzword is harmful or non-harmful. Given the class imbalance, we report macro F1 and harmful-class F1.
Main Results For Meaning Explanation, we use the Human and LLM-judge columns in Table 2 as the primary semantic-equivalence
Model
Prompt ME Human
ES F1
HD F1
3.13 3.03 4.10 3.70
87.62±1.13 88.19±0.40 88.39±0.73 88.24±0.62
78.83 77.20 84.13 84.98
Few-shot Prompting GPT-5.5 en GPT-5.5 zh Gemini 3.1 Pro en Gemini 3.1 Pro zh
4.02 4.01 4.21 4.17
88.98±1.39 89.45±0.85 89.89±0.32 90.18±0.65
79.78 78.47 84.82 84.78
LoRA Fine-tuning Qwen3-8B en Qwen3-8B zh GLM-4-9B en GLM-4-9B zh
2.60 2.81 2.85 2.77
86.74±0.45 86.55±1.00 86.15±0.82 84.81±1.06
65.92 69.35 72.48 71.05
Zero-shot Prompting GPT-5.5 en GPT-5.5 zh Gemini 3.1 Pro en Gemini 3.1 Pro zh
Table 5: Zero-shot prompting, few-shot prompting, and LoRA fine-tuning results. ME: Meaning Explanation; ES: Equivalent Selection; HD: Harmfulness Detection. Values in bold indicate the best results within each setting, whilst those underlined indicate the second-best. Category
Entries
LLM Score
Stylistic device Quotation Experience Slang Homophonic pun Abbreviation
792 602 592 481 398 136
3.34 3.53 4.00 4.19 3.32 3.81
Table 6: Full-benchmark category-level Meaning Explanation scores averaged across models and prompt languages.
Figure 3: Category-level error rates for Equivalent Selection and Harmfulness Detection, averaged across evaluated models and prompt languages.
Figure 4: Average distribution of wrong answers by distractor type in Equivalent Selection. Percentages are computed over incorrect predictions and averaged across evaluated models and option-shuffle seeds.
tuned open models also perform competitively on Equivalent Selection, suggesting that supervised adaptation helps learn task format and option discrimination. However, lower Meaning Explanation and Harmfulness Detection scores indicate that limited supervised data does not fully compensate for weaker base-model knowledge and pragmatic reasoning.
Error Analysis difficult to identify. Averaged across models, both metrics are higher under Chinese prompts, with the largest increases observed for GLM-5.2 and Qwen3.7-Max. Claude Opus 4.8, Gemini 3.1 Pro, and DeepSeek-V4-Pro instead obtain higher scores under English prompts, showing that the observed prompt-language pattern varies across models. We also examine Qwen model-size effects on the full benchmark, with detailed experimental results and analysis provided in Appendix D. Scaling brings clearer gains for Equivalent Selection and Harmfulness Detection than for Meaning Explanation.
Few-Shot and Fine-Tuned Experiment We further examine few-shot prompting and supervised finetuning under a 4:1 train–test split, with 2,401 training entries and 600 test entries. Table 5 reports zero-shot and 2-shot prompting results for GPT-5.5 and Gemini 3.1 Pro on the same test split, where the demonstrations for 2-shot prompting are sampled from the training split, together with LoRA fine-tuning results for Qwen3-8B and GLM-4-9B. Detailed task-specific results and experimental settings are provided in Appendix E. Compared with zero-shot prompting, 2-shot prompting has significantly improved performance. Fine-
Table 6 reports full-benchmark category-level Meaning Explanation scores, with complete per-model results provided in Appendix F.1. Homophonic pun and Stylistic device receive the lowest LLM-judge scores. These expressions require models to recover sound-based transformations, rhetorical or figurative force, and culture-specific source references before expressing the intended sense in English. Slang and Experience receive higher scores because their meanings more often permit direct pragmatic paraphrases. Figure 3 reports full-benchmark category-level error rates for Equivalent Selection and Harmfulness Detection after averaging over models and prompt languages, while Appendix F.2 reports the per-model patterns. Quotation is the hardest category for Equivalent Selection, with Homophonic pun and Stylistic device closely following, whereas Slang is the easiest. Harmfulness Detection shows a different pattern, with Homophonic pun and Stylistic device producing the highest error rates because harmful force can be hidden through phonetic substitution, euphemism, irony, or metaphor rather than an explicit toxic word. Figure 4 reports the distribution of distractor types among incorrect Equivalent Selection predictions. Across prompt languages, errors are concentrated in literal and pragmaticneighbor options, while cultural-mismatch options account
Figure 5: Case study comparing gold annotations and representative model predictions across the three tasks. for a smaller share. This pattern reveals two main sources of confusion. Models may follow the surface wording without recovering the intended online sense, or identify the broad communicative function while missing the finer semantic and pragmatic boundaries of the English equivalent. Compared with English prompts, the model-averaged distribution under Chinese prompts contains fewer literal errors but more pragmatic-neighbor errors. This shift suggests that Chinese instructions may reduce reliance on surface-form correspondence, while leaving models more likely to confuse English expressions that share a broad communicative function but differ in fine-grained meaning, tone, or usage. Appendix F.3 provides the corresponding per-model distributions.
Case Study The aggregate errors become clearer when the three tasks are viewed together. Figure 5 presents two cases with different failure sources. In Case 1, yue lao ci zhi le does not merely mean that romantic fate is hopeless; its annotated sense expresses a broader shift from believing in destined love to prioritizing money and material security. GLM-5.2 instead explains it as the resignation of a love deity, and all models choose the pragmatic-neighbor option Cupid quit rather than money over love. This shows that models can map the cultural figure Yue Lao to an English counterpart while still missing the meme’s socioeconomic stance. Case 2 reveals a different form of inconsistency. 4000+ is a coded malicious curse wishing death upon someone’s entire family, and GPT-5.5 correctly recovers this meaning in Task 1. However, GPT-5.5 and DeepSeek-V4-Pro select the literal option four thousand plus in Task 2, while Claude Opus 4.8 chooses the overly general pragmatic neighbor curse word. Four models also classify the expression as non-harmful despite its explicitly hostile meaning. Together, the cases show that accurate explanation in one setting does not guarantee stable equivalent selection or harmfulness detection, particularly when the in-
tended sense depends on cultural references or coded surface forms. More examples are provided in Appendix G.
Conclusion We introduced CIBuzzBench, a benchmark for Chineseto-English cross-lingual understanding of Chinese internet buzzwords. Built on 3,001 annotated entries, CIBuzzBench evaluates whether LLMs can explain non-literal meanings in English, select appropriate English counterparts, and detect harmful usage. Experiments with representative LLMs show that current models still struggle with culturally grounded pragmatic interpretation, especially when meanings depend on quotation, homophony, stylistic indirection, or finegrained distractor distinctions. Future work can extend CIBuzzBench with temporal updates, richer usage contexts, and additional target languages to study how online meanings evolve and transfer across cultures. Another direction is to develop evaluation and training methods that better separate literal form, pragmatic intent, cultural reference, and safety-relevant offensiveness, so that LLMs can interpret internet language more accurately without over- or under-detecting harmful usage.
Ethics Statement CIBuzzBench is intended for research on cross-lingual understanding and safety evaluation. Its entries are collected from public Chinese buzzword websites and existing research datasets, without private user profiles or conversational metadata. Because some entries contain offensive, derogatory, sexually objectifying, threatening, or otherwise harmful meanings, we use them only for diagnostic evaluation and report aggregate results. Harmfulness labels reflect common online usage and should not be treated as contextfree moderation decisions; the benchmark should not be used to generate abusive content or to moderate users without additional context-specific review.
References Anthropic. 2026. Claude Opus 4.8. https://www.anthropic. com/. Model documentation and release information. Bai, Z.; Yang, L.; Yin, S.; Lu, J.; Zeng, J.; Zhu, H.; Sun, Y.; and Lin, H. 2025. STATE ToxiCN: A Benchmark for Span-level Target-Aware Toxicity Extraction in Chinese Hate Speech Detection. In Findings of the Association for Computational Linguistics: ACL 2025, 10206–10219. Association for Computational Linguistics. Chiu, Y. Y.; Jiang, L.; Lin, B. Y.; Park, C. Y.; Li, S. S.; Ravi, S.; Bhatia, M.; Antoniak, M.; Tsvetkov, Y.; Shwartz, V.; and Choi, Y. 2025. CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs’ Cultural Knowledge Through Human-AI Red-Teaming. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 25663–25701. Association for Computational Linguistics. Deng, J.; Zhou, J.; Sun, H.; Zheng, C.; Mi, F.; Meng, H.; and Huang, M. 2022. COLD: A Benchmark for Chinese Offensive Language Detection. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 11580–11599. Association for Computational Linguistics. Fu, Y.; Huang, Z.; Yang, L.; Lu, Y.; and Dai, Z. 2025. CHENGYU-BENCH: Benchmarking Large Language Models for Chinese Idiom Understanding and Use. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2355–2366. Association for Computational Linguistics. Google DeepMind. 2026. Gemini 3.1 Pro Preview. https://ai. google.dev/gemini-api/docs/models. Model documentation. Guo, H.; He, J.; Ma, J.; Na, H.; Wang, Z.; Zhang, H.; Chen, Q.; Wang, W.; Shi, Z.; Shen, T.; and Chen, L. 2025. Lost in Pronunciation: Detecting Chinese Offensive Language Disguised by Phonetic Cloaking Replacement. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, 2538–2550. Suzhou, China: Association for Computational Linguistics. Hu, J.; Ruder, S.; Siddhant, A.; Neubig, G.; Firat, O.; and Johnson, M. 2020. XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, 4411–4421. PMLR. Huang, C.; Luo, J.; Wang, X.; Lei, W.; and Lv, J. 2025. Can Large Language Models Understand Internet Buzzwords Through User-Generated Content. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12916–12941. Association for Computational Linguistics. Kulkarni, V.; and Wang, W. Y. 2017. TFW, DamnGina, Juvie, and Hotsie-Totsie: On the Linguistic and Social Aspects of Internet Slang. arXiv preprint arXiv:1712.08291. Li, X.; Zhou, Y.; Zhao, L.; Li, J.; and Liu, F. 2025. Impromptu Cybercrime Euphemism Detection. In Proceedings of the 31st International Conference on Computational Linguistics, 9112–9123. Abu Dhabi, UAE: Association for Computational Linguistics.
Liang, Y.; Meng, F.; Wang, J.; and Zhou, J. 2025. SlangDIT: Benchmarking LLMs in Interpretative Slang Translation. arXiv preprint arXiv:2505.14181. Lu, J.; Xu, B.; Zhang, X.; Min, C.; Yang, L.; and Lin, H. 2023. Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and Benchmarks. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16235–16250. Association for Computational Linguistics. Ma, B.; Li, Y.; Zhou, W.; Gong, Z.; Liu, Y. J.; Jasinskaja, K.; Friedrich, A.; Hirschberg, J.; Kreuter, F.; and Plank, B. 2025. Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8679–8696. Vienna, Austria: Association for Computational Linguistics. Mei, L.; Liu, S.; Wang, Y.; Bi, B.; and Cheng, X. 2024. SLANG: New Concept Comprehension of Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 12558–12575. Association for Computational Linguistics. OpenAI. 2026. GPT-5.5. https://openai.com/. Model documentation and release information. Qwen Team. 2025. Qwen3: Think Deeper, Act Faster. https: //qwenlm.github.io/blog/qwen3/. Model documentation and release information. Qwen Team. 2026. Qwen3.7: The Agent Frontier. https: //qwen.ai/. Model documentation and release information. Ruder, S.; Constant, N.; Botha, J.; Siddhant, A.; Firat, O.; Fu, J.; Liu, P.; Hu, J.; Garrette, D.; Neubig, G.; and Johnson, M. 2021. XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 10215–10245. Association for Computational Linguistics. Shi, W.; Li, R.; Zhang, Y.; Ziems, C.; Yu, S.; Horesh, R.; Paula, R. A. D.; and Yang, D. 2024. CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies. In Findings of the Association for Computational Linguistics: EMNLP 2024, 4996–5025. Association for Computational Linguistics. Sravanthi, S.; Doshi, M.; Tankala, P.; Murthy, R.; Dabre, R.; and Bhattacharyya, P. 2024. PUB: A Pragmatics Understanding Benchmark for Assessing LLMs’ Pragmatics Capabilities. In Findings of the Association for Computational Linguistics: ACL 2024, 12075–12097. Bangkok, Thailand: Association for Computational Linguistics. Sun, Z.; Hu, Q.; Gupta, R.; Zemel, R.; and Xu, Y. 2024. Toward informal language processing: Knowledge of slang in large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 1683–1701.
Xiao, Y.; Hu, Y.; Choo, K. T. W.; and Lee, R. K.-W. 2024. ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 5815–5831. Association for Computational Linguistics. Xie, Y.; Wang, C.; Ma, Z.; and Miao, F. 2025. Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme Explanation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 17073–17094. Association for Computational Linguistics. Xu, A.; Lin, B.; Xue, B.; Wang, B.; Xu, B.; Wu, B.; Zhang, B.; Lin, C.; Dong, C.; Ling, C.; et al. 2026. Deepseek-v4: Towards highly efficient million-token context intelligence. arXiv preprint arXiv:2606.19348. Yang, Y. 2023. Analysis of the top ten Chinese Internet buzzwords from the perspective of sociolinguistics. ASEAN Journal of Applied Languages, 2: 67–78. Z.ai. 2025. GLM-4-9B-0414. https://huggingface.co/ THUDM/GLM-4-9B-0414. Model card. Z.ai. 2026. GLM-5.2. https://z.ai/. Model documentation and release information. Zhang, H.; Gao, H.; Hu, Q.; Chen, G.; Yang, L.; Jing, B.; Wei, H.; Wang, B.; Bai, H.; and Yang, L. 2024. ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models. arXiv preprint arXiv:2410.18491. Zheng, C.; Huang, M.; and Sun, A. 2019. ChID: A Largescale Chinese IDiom Dataset for Cloze Test. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 778–787. Association for Computational Linguistics.
Appendix A A.1
Annotation Team and Quality Control Annotator Backgrounds and Diversity
The seven-member annotation and verification team consists of master’s students trained in natural language processing, linguistics, and Chinese language studies, with research experience spanning hate speech, offensive-language detection, lexical semantics, Chinese pragmatics, and crosslingual analysis. All members are familiar with Chinese internet buzzwords and regularly use major Chinese online platforms. The team comprised three women and four men, with four members from northern China and three from southern China. This composition provided variation in regional, dialectal, cultural, and online-community backgrounds, reducing reliance on a single regional or community-specific interpretation.
A.2
Training and Iterative Annotation
Before formal annotation, the team received task-specific training on sense anchoring, non-literal meaning, crosslingual transfer, contextual synonymy, distractor construction, and harmfulness criteria, followed by a shared pilot batch used to refine the guidelines and resolve boundary cases. During full annotation, task-specific reviewers checked GPT-5.5-generated drafts against the documented Chinese sense, while the broader group examined accepted and disputed cases. Inaccurate, overly literal, ambiguous, or pragmatically shifted annotations were revised, and guideline updates triggered re-examination of affected entries until no unresolved issues remained.
A.3
Task-Specific Quality Control and Agreement
Meaning Explanation. The Chinese explanations used as semantic anchors were collected from established online Chinese buzzword dictionary and encyclopedia websites, and we retained only entries whose explanations were sufficiently clear to identify a documented online sense. The gold English explanations were checked for preservation of the core non-literal meaning, target, stance, pragmatic force, and usage conditions. For independent validation, three annotators rated 300 stratified random entries on the 0–5 semanticequivalence scale. The gold English explanations received an average score of 4.95 out of 5. Across the 300 sampled entries, 99.8% of the 900 ratings were at least 4, the three annotators assigned identical scores to 89.3% of the entries, and all remaining disagreements were within one point. These results indicate high consistency in the Meaning Explanation annotations. Equivalent Selection. In CIBuzzBench, an English equivalent is defined as an English expression synonymous with the documented Chinese buzzword sense. Literal translations, topical associations, and expressions with a related but semantically different speech act are not treated as synonyms, and the gold must be the uniquely synonymous option within each four-option set. All option sets underwent a ambiguity audit during construction. For independent validation, three
annotators re-annotated 300 stratified random entries, with 50 entries from each category, without access to the final labels. Each annotator selected the most synonymous option from four randomly ordered candidates. Across the 900 independent judgments, the annotators selected the gold option in 95.3% of cases, and a majority selected it for 96.7% of the sampled entries. All three annotators made the same selection for 91.7% of the entries and unanimously selected the gold option for 90.7%. These results indicate high interannotator consistency and show that the gold equivalents were consistently preferred over the controlled distractors. Harmfulness Detection. Harmfulness labels were primarily reviewed by team members working on hate-speech and offensive-language detection and apply to the documented sense rather than every possible use of the surface string. A sense is harmful when it attacks, demeans, sexualizes, threatens, or otherwise targets a person, group, or identity, while negative sentiment, self-deprecation, non-targeted jokes, complaints, and neutral discussion markers alone are insufficient. For independent validation, three annotators reannotated 400 randomly sampled entries without access to the final labels. The sample was balanced between 200 harmful and 200 non-harmful cases, with category-proportional sampling within each label. Fleiss’ kappa was 0.6920, indicating substantial agreement and consistent application of the harmfulness criteria.
B B.1
Prompt Templates
Zero-shot Prompt Templates
Figure 6 lists the exact English and Chinese prompt templates used in our zero-shot evaluation. The two prompt-language settings keep the task definition, output constraint, and Chinese input fixed, so differences between them mainly reflect instruction-language effects rather than changes in task content.
B.2
Meaning Explanation Scoring Rubric
Figure 7 shows the complete prompt used for LLM-judge scoring of Meaning Explanation. The prompt provides the Chinese buzzword, its Chinese explanation, the gold English explanation, and the model prediction, and asks the judge to output only an integer score from 0 to 5.
C
Human–LLM Judge Agreement
For each evaluated model, we randomly sample 200 Meaning Explanation predictions under the English prompt and another 200 under the Chinese prompt, giving 400 predictions per model. Human annotators and the LLM judge score the same predictions using the 0–5 rubric. Table 7 reports quadratic weighted kappa (QWK) separately for the 200 paired scores in each prompt-language setting. QWK penalizes larger score gaps more strongly than smaller ones. To quantify sampling uncertainty, we perform 10,000 paired bootstrap resamples within each model–prompt setting, recompute QWK for each resample, and use the 2.5th and
Figure 6: Prompt templates used for the three CIBuzzBench tasks. [TERM] denotes the Chinese internet buzzword and [OPTIONS] denotes the four English options in Equivalent Selection. Model
English Prompt
Chinese Prompt
GPT-5.5 Claude Opus 4.8 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.7-Max GLM-5.2
0.7739 [0.6883, 0.8070] 0.7212 [0.6902, 0.7465] 0.7447 [0.7097, 0.7889] 0.7284 [0.6584, 0.8062] 0.7252 [0.6859, 0.7510] 0.8527 [0.7930, 0.9022]
0.7906 [0.7179, 0.8527] 0.7347 [0.6817, 0.7849] 0.6845 [0.6596, 0.7307] 0.6883 [0.6515, 0.7420] 0.7690 [0.7038, 0.8277] 0.7814 [0.7125, 0.8356]
Average
0.7577 [0.7343, 0.7757]
0.7414 [0.7272, 0.7703]
Table 7: Quadratic weighted kappa between human and LLM-judge scores for Meaning Explanation. Brackets report 95% confidence intervals from 10,000 paired bootstrap resamples. The Average row reports the mean QWK across models, with intervals computed from the bootstrap distribution of the model mean.
Model
English Prompt Qwen3-4B 0.73 Qwen3-8B 1.86 Qwen3-14B 1.47
12.09 17.00 15.77
86.46 87.35 86.65
1.33 1.84 1.97
1.90 2.16 2.20
Chinese Prompt Qwen3-4B 1.08 Qwen3-8B 2.09 Qwen3-14B 1.85
14.97 18.77 17.81
86.78 87.81 87.35
1.35 1.79 2.08
1.92 2.15 2.36
Table 8: Qwen model-size results for Meaning Explanation on the full benchmark.
D
97.5th percentiles as the 95% confidence interval. The average QWK is 0.7577 under English prompts and 0.7414 under Chinese prompts, with an overall mean of 0.7496. The model-level intervals remain well above zero across all settings, supporting stable agreement between human evaluation and the LLM judge.
BLEU ROUGE-L BERT-F1 Human LLM
Qwen Model-Size Experiment Results
For the model-size experiment, we evaluate Qwen3-4B, Qwen3-8B, and Qwen3-14B on the full 3,001-entry benchmark under the same zero-shot prompt templates and decoding settings used in the main experiments. We report both English- and Chinese-prompt results. Equivalent Selection is evaluated over five option-shuffle seeds (1111, 2222, 3333, 4444, 5555). For Meaning Explanation, automatic metrics are computed on the full benchmark, while Human and LLM-judge scores are computed on the sampled evaluation subset.
Figure 7: LLM judge prompt used for semantic-equivalence scoring in Meaning Explanation.
Model Qwen3-4B Qwen3-8B Qwen3-14B
English Prompt (F1)
Chinese Prompt (F1)
45.18±0.47 54.59±0.40 56.68±0.51
45.41±0.39 54.91±0.43 56.98±0.63
Table 9: Qwen model-size results for Equivalent Selection Macro F1 on the full benchmark. Values are mean ± standard deviation over five option-shuffle seeds (1111, 2222, 3333, 4444, 5555).
Tables 8–10 show that increasing Qwen model size improves performance across all three tasks. For Meaning Explanation, both Human and LLM-judge scores increase steadily from 4B to 14B under English and Chinese prompts, indicating stronger sense-level interpretation even though the
Model
Macro F1
Harmful F1
English Prompt Qwen3-4B Qwen3-8B Qwen3-14B
28.77 58.62 59.09
25.90 36.51 37.52
Chinese Prompt Qwen3-4B Qwen3-8B Qwen3-14B
22.83 62.55 63.31
26.04 38.51 40.73
Table 10: Qwen model-size results for Harmfulness Detection Macro F1 and harmful-class F1 on the full benchmark.
automatic reference-based metrics fluctuate. Equivalent Selection also improves consistently with scale. Harmfulness
Split
Entries
Model
Train Test
2,401 600
Few-shot Prompting GPT-5.5 88.98±1.39 Gemini 3.1 Pro 89.89±0.32
89.45±0.85 90.18±0.65
LoRA Fine-tuning Qwen3-8B GLM-4-9B
86.55±1.00 84.81±1.06
Table 11: Train–test split for few-shot prompting and LoRA fine-tuning. Model
English Prompt (F1)
Chinese Prompt (F1)
86.74±0.45 86.15±0.82
BLEU ROUGE-L BERT-F1 Human LLM
Table 13: Detailed Equivalent Selection Macro F1 for fewshot prompting and LoRA fine-tuning. Values are mean ± standard deviation over five option-shuffle seeds (1111, 2222, 3333, 4444, 5555).
Few-shot, English Prompt GPT-5.5 4.91 23.66 Gemini 3.1 Pro 3.08 20.51
88.75 88.00
4.02 4.21
3.76 4.16
Few-shot, Chinese Prompt GPT-5.5 6.19 26.20 Gemini 3.1 Pro 3.70 21.98
89.29 88.44
4.01 4.17
3.67 4.09
Model
LoRA, English Prompt Qwen3-8B 4.41 GLM-4-9B 4.22
20.59 20.87
88.06 88.01
2.60 2.85
2.17 2.27
Few-shot, English Prompt GPT-5.5 79.78 Gemini 3.1 Pro 84.82
64.68 72.19
LoRA, Chinese Prompt Qwen3-8B 4.32 GLM-4-9B 4.17
20.72 20.84
88.09 88.03
2.81 2.77
2.25 2.24
Few-shot, Chinese Prompt GPT-5.5 78.47 Gemini 3.1 Pro 84.78
56.76 70.52
LoRA, English Prompt Qwen3-8B GLM-4-9B
65.92 72.48
40.76 51.70
LoRA, Chinese Prompt Qwen3-8B GLM-4-9B
69.35 71.05
47.34 49.33
Table 12: Detailed Meaning Explanation results for few-shot prompting and LoRA fine-tuning. Detection shows the largest gain from 4B to 8B, followed by further improvements in Macro F1 and harmful-class F1 at 14B. Overall, model scaling strengthens cross-lingual explanation, equivalent discrimination, and harmfulness recognition, although substantial room for improvement remains.
E E.1
Harmful F1
Table 14: Detailed Harmfulness Detection Macro F1 and harmful-class F1 for few-shot prompting and LoRA finetuning.
Few-Shot and Fine-Tuned Experiment Train–Test Split
For the few-shot prompting and LoRA fine-tuning experiments, we split CIBuzzBench at the entry level using a 4:1 train–test ratio. The training split is used to sample demonstrations for few-shot prompting and to train LoRA adapters, while all reported few-shot and fine-tuned results are evaluated on the held-out test split. The split sizes are shown in Table 11. This setup ensures that demonstrations and adapter training do not use the same entries that appear in the reported test results.
E.2
Macro F1
Experiment Results
For few-shot prompting and LoRA fine-tuning, we use the train–test split in Table 11. Few-shot prompting evaluates GPT-5.5 and Gemini 3.1 Pro with two demonstrations sampled from the training split for each task and prompt language. LoRA fine-tuning trains task- and prompt-languagespecific adapters for Qwen3-8B and GLM-4-9B-0414 on the training split and evaluates them on the test split. Decoding settings follow the main experiments, and Equivalent Selection is evaluated over five option-shuffle seeds (1111, 2222, 3333, 4444, 5555). For Meaning Explanation, automatic metrics are computed on the 600-entry test split, while Human and LLM-judge scores are computed on the sampled evaluation subset.
Tables 12–14 show that few-shot prompting mainly benefits the API models, especially in Meaning Explanation and Equivalent Selection, where task demonstrations help constrain the response format and clarify the expected crosslingual mapping. The LoRA fine-tuned open models reach competitive Equivalent Selection performance, but their Meaning Explanation and Harmfulness Detection scores remain lower than the few-shot API models. This suggests that supervised adaptation improves task format learning and option discrimination, while sense-level explanation and harmfulness calibration still depend strongly on base-model knowledge and pragmatic reasoning.
F F.1
Error Analysis
Meaning Explanation Category Scores
Table 15 reports full-benchmark Meaning Explanation scores for every evaluated model, prompt language, and buzzword category. Each cell averages the 0–5 LLM-judge scores over all entries in that category. Homophonic pun and Stylistic device generally receive the lowest scores, followed by Quotation, showing that phonetic transformations, figurative force, and source-dependent meanings remain difficult to express precisely in English. Gemini 3.1 Pro obtains the strongest scores across most categories, with Qwen3.7-
Model
Abbrev. Exper. Homoph. Quota. Slang Stylistic
English Prompt GPT-5.5 Claude Opus 4.8 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.7-Max GLM-5.2
3.90 3.54 4.51 3.71 4.16 2.85
4.00 3.90 4.40 3.93 4.29 3.46
3.36 2.73 4.03 3.11 3.78 2.33
3.52 3.36 3.98 3.44 3.81 2.98
4.26 4.28 4.44 4.14 4.34 3.33
3.38 3.18 3.78 3.27 3.56 2.77
Chinese Prompt GPT-5.5 Claude Opus 4.8 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.7-Max GLM-5.2
3.71 3.40 4.40 3.65 4.15 3.73
3.89 3.73 4.33 3.77 4.26 4.06
3.30 2.67 3.88 3.23 3.85 3.56
3.41 3.24 3.94 3.26 3.86 3.57
4.24 4.21 4.39 4.04 4.34 4.23
3.34 2.99 3.71 3.12 3.57 3.35
Table 15: Full-benchmark category-level LLM-judge scores for Meaning Explanation. Scores range from 0 to 5, with higher values indicating better semantic and pragmatic equivalence. Exper., Homoph., and Quota. denote Experience, Homophonic pun, and Quotation. Max generally the next strongest. Prompt-language effects are not uniform; GLM-5.2 improves substantially under Chinese prompts, whereas several other models change only slightly or decline.
F.2
Category-Level Error Patterns
Figure 8 shows category-level classification error rates for Equivalent Selection and Harmfulness Detection under every evaluated model and prompt language. Darker cells indicate higher error rates, making it possible to inspect whether the aggregate hardest categories are broadly shared or driven by particular model–prompt settings. For Equivalent Selection, Slang consistently has among the lowest error rates, whereas Quotation, Homophonic pun, and Stylistic device are generally more difficult. The category gaps are especially pronounced for Claude Opus 4.8 and DeepSeek-V4-Pro, while GPT-5.5 and Gemini 3.1 Pro show lower and more even error rates across categories. Chinese prompts are associated with lower errors across nearly all categories for Claude Opus 4.8 and GLM-5.2 and across most categories for GPT-5.5 and Qwen3.7-Max. The changes are smaller or mixed for Gemini 3.1 Pro and DeepSeek-V4-Pro, showing that the prompt-language pattern varies by model. Harmfulness Detection exhibits a different category ordering. Homophonic pun has the highest or near-highest error rate in most model–prompt settings, and Stylistic device is also comparatively difficult, whereas Quotation is usually less error-prone than in Equivalent Selection. This contrast suggests that source quotations particularly complicate crosslingual equivalent matching, while phonetic substitutions and indirect rhetorical forms make harmful force harder to recognize.
F.3
Distractor Error Patterns
Figure 9 reports the distribution of wrong choices in Equivalent Selection for every model and prompt language.
The per-model distributions reveal distinct error profiles. Pragmatic-neighbor options account for the largest share of errors for GPT-5.5, Claude Opus 4.8, and Gemini 3.1 Pro under both prompt languages. DeepSeek-V4-Pro and Qwen3.7Max are more strongly affected by literal distractors, as is GLM-5.2 under the English prompt. Cultural-mismatch options form the smallest error share in every model–prompt setting, indicating that broad topical or cultural associations are less misleading than surface-form correspondences and fine-grained pragmatic similarities. Across all six models, the Chinese-prompt distribution contains a smaller share of literal errors and a larger share of pragmatic-neighbor errors, with particularly clear shifts for Gemini 3.1 Pro, Qwen3.7-Max, and GLM-5.2. This pattern suggests that Chinese instructions may reduce reliance on literal correspondence, after which the remaining difficulty lies more often in distinguishing English expressions with similar communicative functions but different meanings, tones, or usage conditions. Because these percentages are conditioned on incorrect predictions, they describe a change in error composition rather than an increase in the absolute number of pragmatic-neighbor errors.
G
Additional Case Studies
Figure 10 presents six Equivalent Selection cases under the Chinese-prompt setting, one from each CIBuzzBench category. For jiao fu wen xue, most models choose the pragmatic neighbor soft husband trope, capturing the husband-related theme but missing the affectionate wife-guy romance framing. For wo xuan lan se yao wan and ru dian, all models select literal options instead of recovering the expressions’ online functions of refusing a forced choice and marking something as destined to become a classic. The predictions for mei liur similarly favor a cultural mismatch or pragmatic neighbor over the intended judgment boring as hell. For kou qu, all models recover the source action of vomiting but choose vomit rather than the reaction-like English equivalent eww. Finally, only Claude and Gemini recover the euphemistic sexual sense of diy, while the other models follow its conventional literal expansion do it yourself. These cases show that partial source-language understanding does not guarantee selection of an English expression with the same meaning, tone, and pragmatic function.
Figure 8: Category-level error rates for Equivalent Selection and Harmfulness Detection under all evaluated models and English and Chinese prompts. The darker the colour, the higher the error rate.
Figure 9: Distribution of wrong answers over controlled distractor types for all evaluated models under English and Chinese prompts in Equivalent Selection.
Figure 10: Equivalent Selection case studies under Chinese prompts across the six CIBuzzBench categories, showing the gold equivalent, controlled distractors, and representative model predictions.