Cross-National Information Attacks: A Two-Decade Analysis of Troll Behavior in Korea Jaehong Kim1,2,* Hyeonseung Kim1,* Alice Oh1 Thorsten Holz2 Wonjae Lee1,†
Jiseon Kim1,2 Meeyoung Cha2,1,†
1 Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea
arXiv:2606.22785v1 [cs.SI] 22 Jun 2026
2 Max Planck Institute for Security and Privacy (MPI-SP), Bochum, Germany
Abstract Coordinated foreign influence operations pose a growing threat to online platforms, but detecting state-linked troll activity and tracking its evolution remain challenging. This paper presents an explainable machine learning framework for theory-guided detection and longitudinal analysis of suspected trolling within Korean online news comment sections. Our hierarchical model classifies comments along three dimensions central to influence campaigns: foreign origin, moral-emotional framing, and target country. To support explainability, it also extracts brief span-level textual evidence that provides human-interpretable rationales. We apply the approach to 112M South Korean news comments authored by 4M users over nearly 20 years, identifying 23,998 accounts exhibiting behavior consistent with coordinated manipulation. Analyzing these accounts, we find that they predominantly rely on morally condemning rhetoric rather than direct promotion of foreign-aligned narratives; this rhetoric receives significantly higher user engagement. Among the highestengagement comments, the moral condemnation most frequently targets domestic political figures (e.g., presidents or party leaders) on both the left and the right, potentially amplifying polarization. Our framework supports transparent platform governance through explainable, evidence-based moderation. These observed rhetorical and engagement patterns can inform how platforms and observatories prioritize defenses and intervene before harmful narrative-target combinations achieve widespread reach.
1
Introduction
As information operations increasingly unfold online, adversaries can scale campaigns that target public cognition rather than physical territory, amplifying coordinated narratives to large audiences [4, 10]. Prior research has documented troll * Equal contribution. † Corresponding authors.
campaigns that disrupt elections, distort public health discourse, and exploit politically sensitive issues to intensify polarization across multiple democracies [2, 3, 8, 20, 25]. From a security perspective, such operations threaten the integrity of democratic debate and motivate defenses that enable reliable detection and auditable analysis over time [33, 41, 43]. Despite progress in machine learning methods for detection [17, 40, 50], two gaps still hinder effective defense against influence operations. First, many systems produce binary or coarse labels with limited evidence explaining why content or accounts are flagged, restricting auditability and operational use. Enforcement actions such as warnings, down-ranking, or account restrictions are frequently contested by users and difficult for moderators to justify without an evidence-based rationale [24, 34]. Second, we lack systematic evidence on how rhetorical strategies achieve platform visibility and thus reach large audiences. While prior work describes organizational and operational structures [39], it remains unclear which content-level tactics reliably surface through ranking mechanisms at scale. Theory suggests that ‘condemning’ target countries may be more effective than overt propaganda [20], but large-scale, longitudinal validation remains scarce. This paper presents a three-stage framework that connects explainable troll detection to longitudinal analysis of influence strategies, enabling both large-scale deployment and systematic auditing. Figure 1 depicts an overview of our approach. The framework is evaluated on a newly assembled corpus of 112 million comments spanning nearly two decades from the most widely used news portal in South Korea, providing a uniquely long observational window into the evolution of foreign-influence rhetoric. At a high level, we propose the following three steps: (1) Candidate pool collection: Starting from a verified seed set of 70 troll accounts publicly released by the Institute for National Security Strategy in Korea [21], we expand to a candidate pool via two complementary traces: users connected through the follow network and users who commented on the same articles as the seeds. This process leads to 112 million
Candidate Pool Collection
Troll Detection
Strategy Analysis
Content-Level Detection
Follow Network 3.7K Users 16M Comments
20 Years 112M Comments 4.3M Users
Suspected Comment
Longitudinal Trend Analysis
👍👎
User-Level Aggregation
Known Trolls
Engagements and Visibility
User-Level Detection … Comments on the Same Articles 4M Users 97M Comments
Feature Engineering
Suspected Database 4M Comments 24K Users
Targeted Entities 🎯
Figure 1: Framework for analyzing suspected troll strategies, where explainable AI connects detection to subsequent strategy analysis. (i) Candidate pool collection (Section 2): expand from known trolls via follow networks and shared-article commenting. (ii) Troll detection (Sections 3 and 4): run an explainable content detector, then aggregate outputs for user-level detection. (iii) Strategy analysis (Section 5): analyze temporal trends, engagement patterns, and targeted entities. comments spanning 2006–2025, suitable for detecting coordinated influence beyond the initially identified accounts. (2) Troll detection: We run an explainable content-level detector that predicts hierarchical labels grounded in cognitive warfare research and social-psychological theory, and outputs spanlevel rationales as textual evidence. We fine-tune an LLM for this supervision and distill it into a lightweight model with roughly 0.1 billion parameters, enabling cost-effective inference over the full corpus. We then aggregate commentlevel outputs into user-level features to identify suspected troll accounts. (3) Strategy analysis: We quantify how rhetorical strategies of these accounts evolve over time, how they translate into engagement and visibility under the platform’s ranking dynamics, and which entities are repeatedly targeted. This analysis helps identify which tactics surface and persist in public discourse and provides an empirical basis for defensive prioritization. Across five presidential administrations in South Korea during the measurement period, one pattern emerges gradually yet persistently: suspected troll accounts increasingly post comments expressing the Condemning Korea narrative, which receive higher user engagement within the target country than direct promotion of foreign agendas. We also find that the influx of troll-like users increases around election periods and during major diplomatic tensions, indicating responsiveness to political context. In fact, Condemning Korea comments are more likely than praising comments to surface in top-ranked threads. Evidence further indicates that political leaders from both right-leaning and left-leaning camps are repeatedly targeted by moral condemnation, suggesting an emphasis on intensifying polarization rather than promoting a single faction. Beyond detection, our framework provides actionable evidence for mitigation and response. Hierarchical predictions with span-level rationales can support evidence-based
moderation workflows, and our analysis offers empirical guidance for prioritizing defensive attention toward high-visibility rhetoric. We discuss these implications for platform systems and defensive prioritization. We make the following contributions. 1. We release nearly a two-decade dataset of 112 million South Korean news comments authored by 4 million distinct users.1 By linking publicly identified troll accounts to their interaction contexts, the dataset enables large-scale longitudinal analysis of influence activity. 2. We design an explainable troll detector that produces hierarchical labels and span-level rationales. Grounded in cognitive warfare and social-psychological theory, the model uses knowledge distillation for fast inference over 112 million comments, enabling auditability and supporting transparent analysis beyond black-box detection. 3. Our data reveal influence strategies, including sharp activity spikes around elections and a growing preference for condemning rhetoric over direct promotion of foreign-aligned narratives. We show that condemning comments targeting local politicians are disproportionately likely to appear in top-ranked threads, providing actionable guidance for defensive prioritization. These contributions advance our understanding of foreign influence within online news comment sections. To support future research, we share our dataset of news comments, verified seeds, and model-detected troll accounts (detailed in the Open Science section).2 This repository provides the empirical foundation for important future directions, such as formalizing adversarial threat models that account for multi-faceted political agendas and the strategic deployment of moral-emotional rhetoric. 1 Due to the nature of this research, examples contain biased content. 2 https://doi.org/10.5281/zenodo.20257085
2 2.1
Foreign Influence Dataset Naver News Platform
This is South Korea’s dominant news platform, which is the primary news source used by 63% of the Korean adult population [28, 35]. Similar to Google News, it operates a platformcentric ecosystem where major media outlets publish articles directly on the site. Users can post comments beneath each article, like or dislike other users’ comments, and sort comments by vote-based ranking. Comment Ranking and Visibility. Visibility is determined by a like ratio, calculated as likes divided by total engagement (likes plus dislikes). This mechanism elevates comments that attract broad agreement while reducing the visibility of comments that provoke strong backlash or controversy. Because top-ranked comments receive disproportionate exposure, the highest positions in a thread offer an efficient channel for broadcasting narratives [22]. Since readers rely on these comments as heuristic cues for the prevailing opinion climate [16], the comment section has become a high-stakes arena for influence operations seeking to manipulate public discourse [21]. Follow Network. Naver introduced a follower-following feature that allows users to subscribe to specific comment authors and view their activity. Followed comments receive priority placement on article pages (up to 100 comments, shown in reverse chronological order) and are aggregated in a dedicated "Followed Comments" feed. The follow relationship is asymmetric: if a user blocks another, the blocked user cannot follow, reply to, or upvote their comments.
2.2
Data Collection
The Institute for National Security Strategy in Korea identified and released in 2024 a set of 70 foreign-actor accounts that persistently promoted narratives aligned with specific foreign interests within Naver News comment sections [21]. We treat these accounts as known trolls and use them as seeds for our data collection strategy. In contrast to this ‘gold set,’ our findings are classified as ‘suspected’ state-linked activity, acknowledging that these results are model-generated predictions rather than manually verified instances. Troll Activity. We first collected all comments posted by the 70 known trolls, yielding 356,378 comments across 262,875 articles with an average of 5,091 comments per account. This unusually high activity rate indicates sustained influence efforts and suggests the use of automation. Candidate Pool Collection. To detect additional troll accounts, we constructed a large candidate pool using two complementary approaches. First, we retrieved the complete follower and following lists of all known trolls and collected comments written by users directly connected to them, yielding 16,085,636 comments from 3,703 users. Second, we collected
Foreign State-Suspected? Level 1
Rationale Span Context
Yes
User Name: K쿠 * * Writing Region List: [‘KR’, ‘MNS’] Article Title:
Says he will plan economic policy for 10 years instead of 5.
Moral Emotion? Level 2
Level 3
Comment
Praising
Condemning
Praising Target?
Condemning Target?
MNS
Partner
Rival
Pig dogs feeding on trash news. America is the truly violent country.
Korea
Troll Comment
Figure 2: A framework for explainable comment-level troll detection. It performs multi-level classification (foreign state suspicion, moral emotion, and target country) while extracting level-specific rationales to justify the final troll prediction. MNS denotes a major neighboring state of South Korea.
all comments posted under the 262,875 articles where known trolls had been active, producing 97,765,429 comments from 4,046,948 unique users. This dual approach captures both the social network surrounding known trolls and the broader user population engaged with troll-targeted content. After merging these datasets with the known troll corpus and removing duplicates, we obtained 112,658,554 comments authored by 4,047,831 users as our target dataset for troll detection, covering the period from April 2006 to March 2025. Summary statistics are in Appendix Table 6.
3
Explainable Content-Level Troll Detection
We present an explainable troll detection model that combines a hierarchical labeling scheme with a scalable knowledge distillation pipeline. We begin by conducting human annotation, defining key influence-messaging cues (such as foreign state origin, moral emotions, and targeted countries) and annotating comments with short span-level rationales to create goldstandard data. We then fine-tune GPT-4.1 on these gold labels to generate hierarchical predictions with supporting spans and apply it to annotate 50,000 candidate comments drawn from our 112 million comment corpus. Finally, we apply knowledge distillation by training lightweight open source models on the GPT-annotated data, enabling fast and cost-efficient inference while preserving accurate hierarchical classifications and span-based explanations.
3.1
Data Categorization
Design Rationale We designed a three-level hierarchical labeling framework to operationalize the detection of foreign influence operations on Naver News, grounded in research on
Table 1: Inter-annotator agreement across hierarchy.
cognitive warfare and social psychology. Inspired by explainable AI methods developed for hate speech detection [23, 51], our framework jointly assigns class labels and their supporting span-level rationales to ensure the explainability of the classification process. Figure 2 provides an overview of the annotation process. Level 1: Foreign state-suspected origin. According to NATO, cognitive warfare involves the weaponization of public opinion by external actors to influence policy and destabilize public institutions [4]. This definition implies that the attacking entity must originate outside the target country. The first level therefore identifies whether the comment author is likely of foreign state-suspected origin, based on linguistic, cultural, and behavioral markers including writing region metadata, username patterns, grammatical patterns, and culturally-specific expressions. Level 2: Moral emotions. The second level identifies moral emotions that have been empirically linked to online polarization [7, 47]. Prior research shows that trolls aim to intensify division and emotional conflict [3, 8, 42]. According to Brady et al. [6], moral emotions such as other-condemning (anger, disgust, contempt) and other-praising (admiration, elevation) play central roles in this process. Other-condemning rhetoric establishes in-group versus out-group boundaries by attacking others’ moral integrity, while other-praising rhetoric reinforces in-group cohesion by elevating the moral status of one’s own side. Empirical studies in Korea show that these emotions are strong predictors of online political polarization [26], and research on foreign propaganda strategies indicates that cognitive-warfare campaigns often rely on these two rhetorical pathways [20]. Level 3: Target country. The third level identifies the country toward which the moral emotion is directed. Although mapping these expressions to a specific nation can be challenging due to text ambiguity (requiring an "Unknown" category), this classification is essential for tracking foreign influence. Pinpointing the national target enables a fine-grained analysis of which states are condemned or praised, revealing the strategic focus of the underlying campaigns. Following established geopolitical classifications in East Asia [5], we define five target categories, assigned whenever a comment directs moral emotion toward a country’s leaders, citizens, or sociopolitical issues. • Trolled region: South Korea • MNS (Major Neighboring State) [Anonymized]: suspected troll origin of foreign state-linked actors (⋆The country label is anonymized to keep the analysis focused on technical patterns rather than geopolitical conflict.) • Partner: states aligned with the origin like Russia • Rival: states in strategic opposition like the US or Japan
Labeling Level Foreign state-suspected Moral emotion Condemning target country Praising target country Troll comment label
Krippendorff’s α 0.994 0.512 0.625 0.706 0.813
• Unknown: no specific country can be identified Identifying troll content. We use these three levels hierarchically as operational criteria for identifying troll-like influence messaging. A comment is labeled as suspected troll content when all three conditions are met: (1) the commenter is located in foreign state-suspected, as inferred from an IP-based region marker and non-native linguistic features (Level 1); (2) the comment expresses a moral emotion (Level 2); and (3) the identified target aligns with the actor’s ideological stance (Level 3). In practice, condemning rhetoric directed at South Korea or rivals is labeled as troll content, whereas praising rhetoric directed at a major neighboring state or its partner is likewise labeled as troll content.
3.2
Annotation
Human Labels We compiled a sample of 1,500 instances from the known troll corpus and the candidate pool for human annotation. This included 825 comments from known trolls to ensure coverage of their characteristic messaging patterns, and 675 comments from users who commented on the same articles. This comparative design captures accounts engaging with identical content and topics but potentially exhibiting distinct stances or behaviors. To mitigate emotional bias, an off-the-shelf moral emotion classifier [26] was used to ensure that at least 100 comments per group contained othercondemning and other-praising expressions. The remaining samples were randomly selected. Each comment was presented alongside its metadata, including the username, writing region, and the corresponding news article title. Annotation was conducted using the Label Studio platform [45] by three authors. Annotators assessed three hierarchical levels for each comment, highlighting specific text spans supporting their judgments at each level. Within this framework, the foreign state-suspected designation was treated as binary, whereas moral emotions and target countries were evaluated as multi-label tasks. Only samples that achieved majority agreement across annotators for both class labels and span rationales were retained, yielding a gold label dataset of 1,452 comments. We tested annotation reliability using Krippendorff’s α. Table 1 shows near-perfect agreement for foreign state-suspected (α = 0.994), as metadata such as writing regions provided
Table 2: Performance comparison of fine-tuned GPT and lightweight models on class- and span-level tasks (Macro F1). Bold indicates the best in each column. Foreign statesuspected
Moral emotion
Condemning target
Praising target
Troll
Class
Span
Class
Span
Class
Span
Class
Span
Class
GPT (100) GPT (200) GPT (300) GPT (400)
0.9940 0.9950 0.9950 0.9960
0.8388 0.9089 0.9466 0.9767
0.7749 0.7805 0.7776 0.7715
0.7740 0.7810 0.7890 0.7902
0.5855 0.4998 0.5174 0.5374
0.4510 0.4712 0.4878 0.4991
0.6931 0.4747 0.5930 0.5202
0.4019 0.4296 0.4309 0.3914
0.9468 0.9508 0.9478 0.9458
0.7178 0.6991 0.7206 0.7143
KcBERT KLUE-RoBERTa KcELECTRA
0.9970 0.9960 0.9960
0.8075 0.6308 0.8175
0.8459 0.8544 0.8806
0.5958 0.5636 0.6304
0.6891 0.6505 0.7119
0.4026 0.3907 0.4417
0.6160 0.6093 0.6892
0.3934 0.3384 0.4042
0.9190 0.9170 0.9279
0.6963 0.6612 0.7222
explicit cues. Agreement for moral emotion was moderate (α = 0.512) and the lowest across our categories due to its subjective nature, although it remains comparable to or higher than existing emotion-labeling benchmarks [11, 13]. The final troll comment label achieved substantial agreement (α = 0.813).
3.3
Model / Shot
GPT Labels We divided the 1,452 annotated samples into 400, 52, and 1,000 instances for training, validation, and testing. Unlike in-context learning—which relies on a model’s prior knowledge and a few illustrative examples—we finetuned the model to update its internal parameters based on human-annotated knowledge [9]. Fine-tuning was conducted using OpenAI’s gpt-4.1-2025-04-14 model, the most advanced fine-tuning API available at the time of experimentation (August 2025) [1]. We trained models with subsets of 100, 200, 300, and 400 samples. The best performance was achieved with 300 samples, resulting in the highest overall F1-score on classification and span-labeling tasks (Average F1 = 0.7206). Detailed results are presented in Table 2. To scale annotation beyond the human-labeled samples, we employed a two-stage pipeline. First, we used a pilot ELECTRA-based binary classifier [30] trained on the same 1,452 samples to filter the 112M comment corpus. Because troll comments are relatively rare, we sampled 40,000 comments predicted as troll and 10,000 predicted as non-troll, yielding 50,000 candidates for hierarchical annotation. Second, we applied the fine-tuned GPT model to these candidates to produce hierarchical labels, including both class and span annotations. We removed samples where predicted spans did not match or exceeded text boundaries, retaining 49,745 highquality samples. Minor span offset discrepancies (e.g., oneor two-character shifts), commonly observed in similar studies [12, 18], were corrected by realigning each span to the closest valid text segment. The distribution of samples across hierarchical classes is provided in Appendix Table 7.
Average F1
Knowledge Distillation
We improve scalability and reduce computational cost by distilling the hierarchical representations learned from the GPT-annotated dataset into lightweight models. We used models pretrained on Korean corpora, including KcBERT [29], KLUE-RoBERTa [38], and KcELECTRA [30], each with approximately 0.1B parameters. A total of 49,745 GPTannotated samples were used for model training and validation, with 80% allocated for training and 20% for validation. We built multiple classifiers for each subtask. Binary classification is used to detect final troll content and foreign state-suspected content. Multi-label classification is applied to moral emotions and their corresponding targets, as a single comment can express multiple emotions and reference multiple targets. Span prediction is formulated as a tokenlevel multi-class classification problem by converting spans into BIO tags, where each token is assigned one of the predefined tags. The model consists of 18 classifiers: 2 for binary classification (troll content and foreign state-suspected), 3 for multi-label classification (moral emotions, condemning targets, and praising targets), and 13 for multi-class classification using BIO tagging (1 for a foreign state-suspected span, 2 for moral emotion spans, and 10 for moral emotion target spans). Figure 3 provides an illustrative example of the final distilled model’s predictions on a human-annotated sample, highlighting token-level rationale spans and their associated labels (with opacity indicating predicted probabilities). The predicted outputs can be compared against the span-level ground truth used for evaluation. All three models were trained using the same hyperparameters: the AdamW optimizer [31], a learning rate of 2e-5, weight decay of 1e-2, and 100 epochs, with early stopping applied to prevent overfitting. Binary cross-entropy was used as the loss function for binary and multi-label classification tasks, while cross-entropy loss was used for multi-class classification. Model performance was evaluated using Macro F1. The KcELECTRA-based model achieved the best performance in most categories and was selected as the final model. The
Figure 3: Illustrative output from the final model, with rationale-related spans highlighted. Each span corresponds to a predicted label: Foreign state-suspected (filled orange), Other-condemning (filled green), Other-praising (filled sky blue), Condemning Korea (outlined navy), and Praising MNS (outlined gold). Span opacity indicates the predicted probability for that label. “MNS” refers to a major neighboring state. Examples translated from Korean using GPT-5; original Korean outputs in Figure 10 in the Appendix. evaluation was conducted on 1,000 human-annotated test samples, and performance results are shown in Table 2. Following model selection, inference was performed on 356,378 comments from known trolls and 112 million comments from the candidate pool. To assess the model’s robustness to surfacelevel wording changes, we used an LLM to paraphrase the test samples and observed comparable F1 scores, suggesting that the model captures the target concepts rather than relying primarily on surface phrasing. Details are provided in Appendix D.
4
User-Level Troll Detection
We identify suspected troll users by combining outputs from our explainable content-level classifier with user-level behavioral signals informed by prior work [17, 27, 40]. We represent each user with a multi-dimensional feature set that captures both substantive messaging and behavioral patterns.
4.1
Non-Troll Ground Truth
While we have a set of 70 known troll accounts, user-level modeling requires a robust ground truth set of non-troll users. To construct this set while controlling for topical context and mitigating class imbalance, we randomly selected 100 candidate users who commented on the same articles as known trolls and sampled 20 comments per user (2,000 comments total) for manual review. Three researchers independently reviewed the sampled comments and made a binary, user-level judgment, using the three-level hierarchy in Section 3.1 as coding guidelines (Level 1: foreign state-suspected origin; Level 2: moral emotions; Level 3: target country). We applied a highly conservative criterion: an account was designated as a non-troll only if all three annotators unanimously agreed that none of the 20 evaluated comments met our troll classification thresholds. This strict procedure
yielded 81 confirmed non-troll users to serve as the negative class for our user-level model.
4.2
Feature Engineering
Aggregated Explainable Features. For each user, we aggregate the continuous probability outputs from our contentlevel classifier rather than binarized class labels. For each comment c, the classifier produces probabilities for features f aligned with the three hierarchical levels; we then compute the user-level mean across the set of comments authored by user u: 1 (f) (f) (1) p̄u = ∑ pc , |Cu | c∈C u (f)
where Cu denotes the set of comments authored by u, pc is the probability that comment c exhibits feature f (e.g., Foreign State-Suspected probability, Praising MNS probability), (f) and p̄u is the user-level mean probability. This aggregation minimizes information loss and captures how strongly and consistently a user exhibits each rhetorical characteristic across their commenting history. We compute 14 features: • Foreign state-suspected origin (Level 1): User-level mean of the classifier’s per-comment probabilities for this label. • Moral emotions (Level 2): User-level means of the classifier’s per-comment probabilities for other-condemning and other-praising. • Target country (Level 3; 10 features): User-level means of per-comment probabilities for condemning and praising each target category (i.e., Korea, MNS, rivals, partners, and Unknown). • Troll comment probability: User-level mean of the classifier’s per-comment troll probability.
(a) Top 5 Global Feature Importance
(b) Local SHAP Summary of Top 5 Features
Figure 4: Global feature-importance and SHAP analyses for the user-level troll-detection model. (a) The top five features ranked by mean absolute SHAP value, averaged across 10-fold cross-validation. (b) SHAP summary plots for the same features, showing the distribution and direction of their effects across users. Higher frequency of following ties to known trolls and higher Praising MNS probability shift predictions toward the Troll class, whereas higher average likes shifts predictions toward Non-trolls. Behavioral and Metadata Features. Following prior methods [27, 40], we incorporate 10 behavioral and metadata features commonly associated with influence operations. Together with our content-based signals, this yields in total 24 user-level features used for classification, including: • Exact match overlap: Proportion of a user’s comments exactly matching known troll posts.3 • Account lifespan: Temporal duration (in hours) between a user’s first and last observed comments. • Average comment length: Mean character count per comment authored by the user. • Engagement statistics (3 features): Mean likes, dislikes, and replies received per comment. • Self-duplication ratio: Proportion of repetitive content, defined as 1 − (unique comments/total comments). • Following known trolls: Frequency of ties or outdegree to known trolls. • Followed by known trolls: Frequency of incoming ties from known trolls. See Appendix C for network details. • Comment volume: Total comments per user.
4.3
Evaluation
We evaluate user-level troll detection using a total of 151 users (70 known trolls and 81 non-troll users). For each user, we extract 24 features and train four supervised classifiers: 3 To ensure meaningful content matching, we restrict our analysis to com-
ments meeting a minimum length threshold (≥10 characters and ≥3 tokens).
SVM, Random Forest, LightGBM, and XGBoost. Using 10fold cross-validation, the models achieve F1 scores ranging from 0.91 to 0.94: SVM reaches 0.91, Random Forest and LightGBM reach 0.93, and XGBoost performs best with 0.94. Feature Importance Analysis. To understand the model’s decision-making process, we analyze feature importance using SHAP values derived from the XGBoost classifier (Figure 4(a)), which quantify each feature’s contribution to the model output in a unified, model-agnostic manner[32]. The five most influential features are the frequency of following ties to known trolls, self duplication ratio, the user-level mean praising MNS probability, the user-level mean Foreign StateSuspected probability, and average likes. The local SHAP summary (Figure 4(b)) further illustrates how feature values shift predictions toward the Troll or Nontroll class, revealing both the directionality and heterogeneity of feature effects across users. In particular, higher values of content-based features, especially praising MNS and Foreign State-Suspected probabilities, tend to push predictions toward the troll class. Overall, model-detected trolls exhibit higher values in network and content-based features, whereas higher average likes are more characteristic of non-troll users.
4.4
Prediction
We applied the trained XGBoost user-level detector to the merged candidate pool described in Section 2.2, comprising 112M comments authored by 4,047,680 users. The candidate pool was constructed by merging (i) the article-based candidate pool of users who commented on the same articles as known trolls and (ii) the follow-network candidate pool of users who follow or are followed by known trolls. For inference, we ensembled the 10 models from 10-fold cross-validation by averaging predicted probabilities. The model initially flagged 25,381 users as suspected trolls. We
Table 3: Illustrative examples for each target country category in comments by model-detected trolls, translated using GPT-5. Target Country
Example
Condemning Korea (i.e., Trolled Region)
A demon race, the Koreans, and a demon state, South Korea. They have maintained a conscription-slavery system for 60 years that serves no purpose other than enslaving their own citizens. These Koreans loudly complain about forced labor from 70 years ago that they never personally experienced, yet remain silent about the ongoing forced labor imposed by their own government on Korean men today.
Condemning Rival (i.e., Opposing State)
The United States is getting desperate over its debt. Are they going to print more dollars again? haha. What about inflation? haha. In 2008, MNS bought U.S. bonds and pulled them through the crisis. So what now? This time MNS will not step in.
Praising MNS (Major Neighboring State)
Korea does not seem suited for American-style democracy. It should learn from MNS-style socialism. Why does Korea keep acting up? Within 10 years, MNS will become the world’s number one economy. For 5,000 years, MNS was the center of the world, and now it is returning to its rightful place. MNS is a big country, while Korea is a small country. If Korea continues like this, what will happen? The United States is far away, but MNS is close. MNS and Korea should be good friends.
Praising Partner (i.e., Aligned State)
That comedy-idiot and his underlings all ran off and hid in civilian residential areas. Don’t show mercy to civilians—just grind the whole city into dust. Stay strong, mighty Russia. Putin, you are a hero.
excluded 1,383 users who had zero troll comments, leaving 23,998 users (0.59% of 4,047,680) in the final set of model-detected accounts. We refer to these accounts as modeldetected trolls, but they should be interpreted as algorithmic predictions rather than independently verified state actors. Among these users, 23,993 appeared in the article pool, 986 appeared in the follow-network pool, and 981 appeared in both pools. Only 5 model-detected trolls were found exclusively via the follow-network pool, whereas 23,012 were found exclusively in the article pool. This distribution suggests that, although proximity in the follow network to known trolls is highly informative, as reflected in the featureimportance results in Figure 4(a), the detector does not rely solely on the follow graph. It also incorporates signals from the explainable troll-content classifier and additional userlevel behavioral features. Table 3 shows the translated examples for each target country category in comments posted by model-detected trolls. The Condemning Korea example frames conscription as “state-led slavery,” while Condemning Rival criticizes U.S. dollar dominance and portrays the United States as dependent on MNS. In contrast, Praising MNS depicts the major neighboring state (i.e., troll origin) as a role model for Korea, emphasizing MNS’s socialist system and economic rise. Finally, Praising Partner derogates Ukrainian leadership as a “comedy-idiot” while glorifying Putin as a “hero.” Together, these examples illustrate the narrative styles produced by model-detected trolls. We further examine these patterns through systematic behavioral analysis in the following section.
4.5
Characterizing Model-Detected Trolls
To examine whether model-detected trolls exhibit troll-like behavioral patterns, we conduct comparative analyses of 23,998 model-detected trolls and 4 million non-troll users. Time-series Evaluation. Prior work [40] validates troll detection by demonstrating that flagged accounts exhibit stronger temporal synchronization with known trolls than baseline users do. Following this strategy, we evaluate temporal synchronization by comparing the activity time series of known trolls, model-detected trolls, and non-troll users. To capture temporal concentration, for each group we construct an activity share series, where each time interval is expressed as the percentage of the group’s total comments over the full observation period. We then compute Pearson correlations over a common observation window. Model-detected trolls exhibit consistently strong alignment with known trolls across all temporal resolutions, with daily r = 0.8144, weekly r = 0.8354, and monthly r = 0.8464. In contrast, non-troll users show weaker alignment, with daily r = 0.6509, weekly r = 0.6829, and monthly r = 0.6996. Some baseline correlation for non-troll users is expected because our article-based candidate pool conditions on overlapping news exposure (i.e., commenting on the same articles as known trolls), which can increase temporal synchronization. However, model-detected trolls show a consistently higher degree of synchronization. To test whether the correlation between model-detected trolls and known trolls is significantly larger than that between non-troll users and known trolls, we perform a statistical comparison of dependent correlations that share the same reference series. We apply a Fisher r-to-z transformation and use a Williams test for comparing correlations with an over-
53.18
Table 4: User-level average percentages computed from content-level predictions across six classes for model-detected trolls and non-troll users. Model-detected trolls consistently exhibit higher prediction rates across all six classes considered important within the hierarchical detection framework. “MNS” refers to a major neighboring state. Class Troll Comments Foreign State-Suspected Condemning Korea Condemning Rival Praising MNS Praising Partner
Model-Detected Trolls 67.19% 81.25% 57.16% 1.85% 1.11% 0.004%
lapping variable [49]. The difference is statistically significant at all resolutions (daily z = 67.18, weekly z = 28.54, monthly z = 14.98, all p < 0.001), indicating that model-detected trolls remain significantly more synchronized with known trolls than non-troll users, beyond what is expected from shared news exposure. Content-Level Signals. We used predictions from an explainable content-level model across six troll-related classes. Table 4 shows that model-detected trolls exhibit higher average prediction scores than non-troll users across all trollrelated classes. We conducted a statistical analysis to examine whether user-level prediction scores differ between modeldetected trolls and non-troll users. Levene’s tests indicated violations of the equal variance assumption across all six classes (p < 0.001), motivating the use of Welch’s t-tests, which are robust to heteroscedasticity. Prediction scores were significantly and consistently elevated for model-detected trolls across all classes (Troll Comments: t = 212.07; Foreign State-Suspected: t = 220.69; Condemning Korea: t = 185.14; Condemning Rival: t = 42.84; Praising MNS: t = 25.88; Praising Partner: t = 3.41; all p < 0.001) and remained significant after Bonferroni correction (α = 0.0083), indicating robustness across multiple tests. These results indicate that the user-level model effectively captures systematic variation between model-detected trolls and non-troll accounts within the predicted troll-related classes (see details on user-level traits in Appendix B). Span-Level Linguistic Evidence. We examined the prevalence of MNS language expressions at the span level. For comparison, we removed the Writing Region token, which appeared frequently in all groups. All measures were computed from tokens extracted within spans predicted as Foreign StateSuspected. As shown in Figure 5, model-detected trolls heavily use culture-specific markers and jargon associated with the major neighboring state, whereas non-troll users rarely exhibit this pattern. Chi-squared tests confirmed significant differences in MNS language frequency and diversity across groups (p < 0.001). Post-hoc one-tailed Z-tests for propor-
Percentage (%)
40
60
54.88
30 20
0
40
21.02
10
Non-Trolls 22.35% 31.24% 18.21% 0.53% 0.40% 0.0004%
84.78 80
50
5.24
0
Frequency
Known Troll
27.68
20
Diversity
Model-Detected Troll
Non-Troll
Figure 5: MNS language frequency and diversity by user group, computed from tokens within spans predicted as Foreign State-Suspected (excluding the Writing Region token). Frequency: % of MNS-language tokens among all tokens in these spans. Diversity: % of unique MNS-language tokens among MNS-language tokens in these spans.
tions further showed that both known and model-detected troll groups exhibit significantly higher frequency (Known vs. Non-Troll: Z = 31.20; Model-Detected vs. Non-Troll: Z = 136.67) and diversity (Known vs. Non-Troll: Z = 13.55; Model-Detected vs. Non-Troll: Z = 62.42), all with p < 0.001 compared to non-troll users. These results indicate that modeldetected trolls tend to exhibit higher frequency and diversity of MNS language usage, suggesting clear linguistic differences from non-troll users. These time-series, content-level, and span-level analyses indicate that model-detected trolls differ systematically from the broader user baseline. Accounts flagged by the model exhibit temporal activity patterns that align more closely with known trolls than with typical non-troll users, have higher average user-level prediction scores across the six troll-related classes, and display distinctive linguistic indicators at the span level. Further comparative observations regarding these model-detected trolls and baseline users are provided in Appendix B.
5
Strategy Analysis
For strategy analysis, we combine the verified seed set of 70 known trolls with the 23,998 model-detected accounts, yielding a consolidated pool of 24,068 suspected troll accounts. To focus on influence-related behavior, we restrict the analysis to comments classified as troll comments (4,123,962), rather than all comments authored by these users (15,054,072). This restriction helps isolate campaign-relevant messaging from broader, non-strategic discussion (e.g., memes). We characterize campaign-level strategies along three dimensions: (i) longitudinal trend analysis, (ii) comment visibility by rhetoric, and (iii) targeted entities in the most visible comments.
(a) Evolution of Suspected Troll Users' Rhetoric across Administrations
(b) Influx of Suspected Troll Users across Major Elections
Figure 6: Temporal patterns of troll activity from 2006 to 2025. (a) Evolution of comment volume across different rhetorical categories. Background colors indicate governing administrations (blue = left-leaning; red = right-leaning), with presidents’ last names shown at the top. The unshaded period in 2017 represents the transitional period following impeachment before the new administration. (b) Monthly influx of newly identified troll-like users, calculated based on each user’s first comment date. Red vertical lines mark presidential elections, while black dashed lines indicate legislative and local elections.
5.1
Longitudinal Trend Analysis
20-Year Rhetorical Asymmetry. As shown in Figure 6 (a), we analyze the longitudinal trajectory of comments authored by suspected troll accounts over a 20-year period (monthly counts; log scale). A primary observation is the marked asymmetry in rhetorical emphasis: regardless of administration changes, suspected troll accounts posted substantially more condemning content than praising content. “Condemning Korea” remains the dominant narrative throughout the entire period. Given that Naver News is a major domestic news platform, this pattern suggests efforts to influence domestic discourse through adversarial framing rather than overt promotion of external narratives. All rhetorical categories show upward trends from 2017, reaching their maximum volumes around 2018. In 2018, “Condemning Korea” peaks at 675,107 comments, followed by “Condemning Rival” (20,254), “Praising MNS (major neighboring state)” (10,345), and “Praising Partner” (108). This surge temporally coincides with two major developments:
(1) the 2017 institutionalization of military-civil fusion and cognitive-domain operations by MNS, which has been discussed as enabling more coordinated psychological and informational campaigns [44], and (2) the escalating diplomatic conflict between Korea and MNS over the deployment of the US THAAD missile defense system, which intensified to the point of economic retaliation. Election-Oriented Entry Patterns. Motivated by prior evidence that coordinated influence activity intensifies around elections [19, 46], we examine whether and to what extent suspected trolls join the platform around elections. Figure 6 (b) plots the influx of newly detected troll users, computed based on their first observed comment date, with major presidential, legislative, and local elections annotated. New account activity increases around these pivotal political cycles. The largest spike occurs around the May 2017 presidential election following President Park’s impeachment, with 635 newly appearing suspected troll accounts. To evaluate the statistical significance of this pattern, we
(a) Distribution of Like Ratio
(b) Effect of Condemning Korea Probability on Predicted Like Ratio
Figure 7: Like ratio by rhetorical strategy. (a) Distribution of like ratios across strategies with the neutral threshold (like = dislike) shown as a dashed line; Condemning Korea is the only strategy whose mean exceeds this threshold. Error bars denote 95% confidence intervals around the mean. (b) Predicted like ratio from a fractional logit model as a function of the Condemning Korea probability, with user-clustered standard errors and 95% confidence intervals. analyze user influx at a weekly resolution, balancing the aggregation bias of monthly counts against the sparsity of daily data. Election periods are defined as weeks within ±30 days of any major election and are compared against non-election weeks using the Mann-Whitney U test. The average weekly influx of flagged troll-like users is 34.19 during election periods compared to 22.61 during non-election periods (+51.23%; p < 0.01). A large cohort of non-troll users (4M) also shows a higher influx during election periods (4,828.11 vs. 3,893.42; +24.01%; p < 0.05), suggesting that elections generally attract new platform participation. To assess whether suspected trolls concentrate disproportionately around elections beyond this general influx, we compute penetration density, defined as the weekly fraction of newly appearing users classified as trolls. Penetration density increases from 0.82% in non-election periods to 0.91% during election periods (+10.98%; p < 0.05, Mann-Whitney U). These results indicate that troll-like accounts show elevated entry rates around elections, consistent with election-oriented mobilization patterns documented in prior work [19, 46].
5.2
Comment Visibility by Rhetoric Strategy
We analyzed which rhetorical strategies employed by suspected troll accounts were most likely to achieve visibility among Naver News users. On this platform, comment visibility is determined by the like ratio, calculated as likes (likes + dislikes)
(2)
with higher ratios elevating comments to the top of the thread. Comments with zero engagement (both likes and dislikes equal to zero) were assigned a like ratio of 0, as they remain invisible at the bottom of threads, similar to comments receiving only dislikes.
Cross-Strategy Statistical Comparison. Figure 7 (a) shows the like ratios associated with different rhetorical strategies. Among troll strategies, condemning rhetoric achieved higher like ratios on average than praising rhetoric. Condemning Korea was the only strategy to exceed the neutral threshold of 0.5 (mean = 0.514), indicating it received more likes than dislikes on average and was most likely to achieve top-thread visibility. In contrast, Condemning Rival (mean = 0.422), Praising MNS (mean = 0.362), and Praising Partner (mean = 0.241) all fell below the 0.5 mark. Because like-ratio distributions were non-normal across all categories (Shapiro-Wilk tests, p < 0.001), we applied a Kruskal-Wallis test, which revealed significant differences among categories (H = 13101.53, p < 0.001). Post-hoc Dunn tests with Bonferroni correction confirmed that all pairwise contrasts were significant (p < 0.001). These results indicate that the visibility of troll content varied significantly by the emotional target of the message, with condemning rhetoric consistently outperforming praising rhetoric in terms of user engagement. This pattern holds across political administrations (see Figure 11 in the Appendix). Condemning Korea ranked first in four out of the five periods, with the sole exception occurring under the Lee administration, where Condemning Rival showed the highest like ratio. Across all regimes, praising rhetoric persistently received lower engagement than condemning rhetoric. Impact of Rhetorical Intensity on Visibility. To further examine whether the intensity of each rhetorical strategy influenced visibility, we conducted a regression analysis. Following prior work [26], we operationalized rhetorical intensity as the softmax probability from our ELECTRA-based contentlevel classifier’s output logits. This probability represents the model’s confidence in classifying a comment into a specific rhetorical category, which we interpret as the strength of that
Table 5: Fractional Logit estimates: Impact of rhetoric intensity on Like Ratio (N=4,123,692). All models control for year and month fixed effects. Significance levels: ∗∗∗ p < .001 ( positive , negative ). Independent Variable (Intensity)
Like Ratio (Coef.)
Condemning Korea
0.1086∗∗∗
Condemning Rival
-0.5093∗∗∗
Praising MNS
-1.0888∗∗∗
Praising Partner
-0.2072
rhetorical strategy in the content. We performed a regression with the four rhetorical intensity scores (Condemning Korea, Condemning Rivals, Praising MNS, Praising Partners) as independent variables and the like ratio as the dependent variable. Because the dependent variable ranges between 0 and 1 with substantial mass at both boundaries (30.44% at 0 and 21.32% at 1), we employed a Fractional Logit Model [36]. Year and month fixed effects were included to account for temporal and seasonal variability. Table 5 shows the effects of rhetorical intensity on comment like ratios. Among the strategies, only Condemning Korea significantly boosted engagement (β = 0.1086, P < 0.001), directly enhancing content visibility. This relationship is illustrated in Figure 7(b): as the intensity of Condemning Korea increases, it crosses the neutral threshold (where likes equal dislikes) at a probability of 0.45. In contrast, both Condemning Rival (β = −0.5093, P < 0.001) and Praising MNS (β = −1.0888, P < 0.001) significantly suppressed like ratios. While Praising Partner also yielded a negative coefficient (β = −0.2072), its effect was not statistically significant. The negative coefficient for Praising MNS suggests that this narrative strategy triggers intense audience antipathy, rendering such content the least likely to achieve organic visibility.
5.3
Targeted Entities in Condemning Korea
We conducted a deeper analysis of the Condemning Korea strategy, which showed the highest engagement potential over the 20-year period and the greatest likelihood of achieving topthread visibility. Our explainable content-level troll detection framework provides span-level rationales that explicitly mark who the Condemning Korea rhetoric targets, enabling us to identify the frequently targeted entities. Figure 8 presents the ten most frequent target spans in Condemning Korea comments with the highest like ratio (=1). Seven of the top ten targets are political leaders. The most frequently mentioned figures are Moon Jae-in (liberal-leaning former president; 16,651 mentions), Lee Jae-myeong (liberalleaning presidential candidate; 13,522 mentions), Yoon Sukyeol (conservative president; 9,887 mentions), and Cho Kuk (liberal-leaning politician; 6,314 mentions). The targets span
Figure 8: Top-10 targets in Condemning Korea comments with the highest like ratio (=1). Bars show the frequency of target spans extracted by the content-level detector. Colors denote target type: political figures by ideological camp (left-leaning vs. right-leaning), and non-figure targets in gray. Text inside each bar shows original Korean spans with English translations on the left.
both liberal and conservative camps, consistent with prior work suggesting that troll activity tends to amplify polarization [3]. Targeting high-salience political figures may be associated with higher visibility and engagement in comment threads. Appendix Table 11 shows that presidents, major candidates, and other elite-associated figures frequently appear among the top-10 target spans in high-visibility Condemning Korea comments across administrations. The incumbent president consistently appears among the top-10 targets regardless of ideological orientation. This trend is more pronounced in the Moon and Yoon administrations, where political figures account for a larger share of the most frequent target spans. These targets range from political leaders (acting as representatives of parties and the state) to political parties and the nation itself, aligning with core attributes of moralized content as discourse oriented toward units larger than the individual (e.g., society or culture) [15]. The concentration of these collective-level targets among the most visible comments is notable, consistent with prior findings that moralized messaging captures attention and can spread through online networks [7, 37]. By achieving high visibility, these comments disproportionately expose audiences to rhetoric centered on systemic or collective concerns rather than private matters, which prior research links to more polarized discourse [7].
6 6.1
Discussion Findings and Stakeholder Reporting
Over the course of nearly two decades, our analysis reveals that Condemning Korea is both the dominant and most amplified narrative in influence campaigns from potentially the major neighboring state. This pattern aligns with cognitive-warfare research suggesting that negative, antagonistic rhetoric is particularly effective at capturing user attention [10]. To support immediate verification and mitigation, the list of 24,068 suspected troll accounts has been shared with Naver News and the Institute for National Security Strategy in South Korea, the organization that released the initial troll label dataset used as our seeds. These shared data provide platform stakeholders and policy observers with the actionable insights required to address systemic account irregularities and uphold the integrity of public discourse.
6.2
Implications for Online Platforms
The proposed explainable content-level detector provides a concrete mechanism for strengthening platform governance. Moderation decisions such as warnings, down-ranking, or account restrictions are frequently contested because affected users receive little insight into why an action was taken [24, 34]. By producing hierarchical label predictions, our model can surface specific textual evidence underlying moderation outcomes, thereby supporting moderation and review workflows and making enforcement decisions more transparent, interpretable, and contestable. Another consideration is deployment at scale. Although LLMs are used during training through knowledge distillation, our inference relies on a smaller model (≈ 0.1B) while maintaining strong predictive performance. This lightweight design is suitable for production environments, where latency and cost constraints typically limit the use of large models. Our findings suggest that adversarial adaptation may not be cost-free, even if the model’s decision logic becomes known. In this setting, the cues that drive comment visibility are closely aligned with the rhetorical signals captured by our theory-grounded, explainable detector. To evade detection, attackers would need to weaken these rhetorical cues; however, because our regression analysis shows that these cues are associated with visibility (Table 5), evasion may come at the expense of reduced reach and influence [7, 26]. In other words, evasion attempts can create a trade-off between bypassing detection and maintaining strategic effectiveness.
6.3
visibility [14, 43], limited defensive resources should be allocated to timely and targeted interventions. Our findings suggest that prioritization should account for both when to monitor and what to monitor. The concentration of suspected troll entry around major elections indicates that election periods are especially important windows for heightened scrutiny. At the same time, condemning moralemotional rhetoric has grown in both prevalence and visibility over two decades, and condemnation of political figures across the political spectrum was particularly likely to gain exposure. These patterns suggest that not all suspicious content carries the same amplification risk, as engagement-based ranking dynamics can preferentially amplify moral-emotional expression [7, 37]. Platforms and observatories can allocate additional review capacity during major election windows and prioritize the verification and, where appropriate, debunking of suspicious messages that employ condemning rhetoric and target political figures, enabling intervention before such messages achieve widespread reach.
Implications for Defensive Prioritization
Effective responses to information operations require more than detection; they demand strategic decisions about intervention and resource prioritization [43, 48]. As harmful narratives may become harder to correct once they gain widespread
7
Conclusion
This study presents a scalable framework for detecting adversarial accounts within online news comment sections. Our approach offers explainable rationales via a hierarchical contentlevel classifier that evaluates foreign state–suspected origin, moral emotions, and target entities of influence operations. By aggregating these outputs into user-level behavioral features from longitudinal data, we track how underlying narrative strategies evolved and gained visibility on a South Korean platform over nearly twenty years. Our findings reveal a persistent pattern: rhetoric focused on moral condemnation dominates troll activity and consistently achieves higher visibility than content praising foreign agendas. The number of new troll-like accounts increases around elections, indicating a strategic responsiveness to major political events. Moreover, span-level evidence shows that domestic political leaders from both liberal and conservative camps are repeatedly targeted. This cross-partisan targeting demonstrates that these operations prioritize intensifying polarization and societal distrust rather than advancing a specific partisan position. Our findings offer a unique empirical view into the realworld evolution of influence operations within South Korean news comments. They hold significant implications for platform governance. Because individual news readers lack this global view, they cannot easily discern which isolated comments are part of a coordinated foreign influence campaign. The use of span-level rationales provides interpretable evidence for content moderation, while the strategy analysis indicates which narratives are most likely to surface and persist. Together, these insights can inform platform policy and help prioritize defensive attention toward the most engagementeffective tactics.
Ethical Considerations Data Collection and Privacy. All data were collected from the public-facing Naver News platform in accordance with the platform’s terms of service, consisting of public comments and associated metadata. To minimize the risk of user re-identification, usernames are partially masked in all publicfacing outputs, including the text examples in this paper and the open repository. Furthermore, computational analyses were conducted entirely on secure institutional infrastructure. We also strictly focus on aggregate behavioral patterns rather than individual profiles. Accordingly, no attempts were made to infer sensitive personal attributes, such as political affiliation, demographic characteristics, or real-world identities. Data Removal. Although the public release is partially masked, we retain the full usernames on access-restricted internal infrastructure, which lets us process removal requests. The Zenodo release includes a bilingual Korean-English guide explaining how individuals can locate their comments using the visible username prefix and comment text to submit a removal request, if needed. For account removal, individuals can provide evidence of account ownership, such as a screenshot of their Naver profile displaying the full username. For comment removal, individuals can specify the relevant comment URL. Upon verification, the corresponding rows will be removed from our maintained research dataset and all future public releases. Stakeholders and Potential Harms. This research has operational implications for suspected troll accounts, general news commenters, civic organizations, and the platform itself. For model-detected accounts, the primary risk involves misclassification and subsequent reputational harm if they are incorrectly labeled as state actors. To mitigate this risk, we consistently use the terms “suspected” and “model-detected,” report aggregate behavioral patterns rather than individual accusations, and withhold unmasked account lists from the public release. We mask all general commenter usernames across public outputs to prevent unwarranted associations. For civil-society groups, we report prevalence and trends rather than actor attribution or operational recommendations against specific countries. For the platform, we used only its public API within the terms of service and rate limits and shared our findings for follow-up verification. Researcher Well-being. The research required repeated exposure to hostile, politically charged comments, which can impose a psychological burden on researchers. All annotation and close reading were conducted by research team members rather than external or paid annotators. To mitigate potential harm, we used pre-annotation risk briefings with opt-out provisions, regular check-ins, team-based review of the most difficult content rather than solo review, access to institutional counseling, and pause requests without exception.
Research Justification and Public Interest. We conducted this study after carefully evaluating potential risks and determining that the public utility of tracking long-term influence operations on Naver significantly outweighed those concerns. We framed this work as defensive measurement research designed to support digital governance, independent scrutiny, and civil-society awareness. To mitigate potential misuse, we avoid attributing model-detected accounts to real-world identities or independently verified state actors, releasing only aggregate analyses and partially masked data. These safeguards allow the findings to inform defensive strategies while minimizing the risks of harassment, retaliation, or the direct targeting of individual users.
Open Science To support reproducibility and future research, our complete dataset is publicly available on Zenodo at https://doi.org/ 10.5281/zenodo.20257085. This repository includes (1) a near 20-year corpus of 112 million news comments with partially masked user identifiers to protect privacy and (2) a bilingual Korean-English administrative guide detailing the user data removal and opt-out verification process.
Acknowledgments We thank Carmela Troncoso and researchers at the MPI-SP for their valuable feedback. This work was supported by the Hyundai Motor Chung Mong-Koo Foundation, the IITP grant (RS-2024-00441762), and the NRF grant (RS-202200165347) funded by the Korean government (MSIT).
References [1] Josh Achiam, Steven Adler, Sandhini Agarwal, et al. GPT-4 technical report, 2023. [2] Meysam Alizadeh, Jacob N Shapiro, Cody Buntain, and Joshua A Tucker. Content-based features predict social media influence operations. Science Advances, 2020. [3] Ahmer Arif, Leo Graiden Stewart, and Kate Starbird. Acting the part: Examining information operations within# BlackLivesMatter discourse. Proceedings of the ACM on Human-Computer Interaction, 2018. [4] A. Bernal, C. Carter, I. Singh, K. Cao, and O. Madreperla. Cognitive warfare: An attack on truth and thought. Technical report, NATO Innovation Hub and Johns Hopkins University Applied Physics Laboratory, 2020. [5] Jude Blanchette, Ryan Hass, and Lily McElwee. Building International Support for Taiwan. JSTOR, 2024.
[6] William J Brady, Molly J Crockett, and Jay J Van Bavel. The MAD model of moral contagion: The role of motivation, attention, and design in the spread of moralized content online. Perspectives on Psychological Science, 2020. [7] William J Brady, Julian A Wills, John T Jost, Joshua A Tucker, and Jay J Van Bavel. Emotion shapes the diffusion of moralized content in social networks. Proceedings of the National Academy of Sciences, 2017. [8] David A Broniatowski, Amelia M Jamison, SiHua Qi, Lulwah AlKulaib, Tao Chen, Adrian Benton, Sandra C Quinn, and Mark Dredze. Weaponized health communication: Twitter bots and Russian trolls amplify the vaccine debate. American Journal of Public Health, 2018. [9] Tom Brown, Benjamin Mann, Nick Ryder, et al. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems, 2020. [10] Bernard Claverie and François Du Cluzel. “cognitive warfare”: The advent of the concept of “cognitics” in the field of warfare. In Cognitive Warfare: The Future of Cognitive Dominance. NATO Collaboration Support Office, 2022. [11] Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. GoEmotions: A dataset of fine-grained emotions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020. [12] Yifan Ding, Michael Yankoski, and Tim Weninger. Spanoriented information extraction: A unified framework. ACM SIGKDD Explorations Newsletter, 2025. [13] Anjalie Field, Chan Young Park, Antonio Theophilo, Jamelle Watson-Daniels, and Yulia Tsvetkov. An analysis of emotions and the prominence of positivity in# BlackLivesMatter tweets. Proceedings of the National Academy of Sciences, 2022. [14] Fabio Giglietto, Laura Iannelli, Augusto Valeriani, and Luca Rossi. ‘fake news’ is the invention of a liar: How false information circulates within the hybrid news system. Current Sociology, 2019. [15] Jonathan Haidt. The moral emotions. In Handbook of Affective Sciences. Oxford University Press, 2003. [16] Jiyoung Han. Commenters and lurkers: Navigating the two-step flow of communication in online news discourse. New Media & Society, 2025.
[17] Hans W. A. Hanley, Deepak Kumar, and Zakir Durumeric. Specious sites: Tracking the spread and sway of spurious news stories at scale. In IEEE Symposium on Security and Privacy, 2024. [18] Maram Hasanain, Fatema Ahmad, and Firoj Alam. Large language models for propaganda span annotation. In Findings of the Association for Computational Linguistics: EMNLP, 2024. [19] Philip N Howard, Bharath Ganesh, Dimitra Liotsiou, John Kelly, and Camille François. The IRA, social media and political polarization in the united states, 2012– 2018. Technical report, Project on Computational Propaganda, University of Oxford, 2018. [20] Tzu-Chieh Hung and Tzu-Wei Hung. How China’s cognitive warfare works: a frontline perspective of Taiwan’s anti-disinformation wars. Journal of Global Security Studies, 2022. [21] Institute for National Security Strategy (INSS). Current state of foreign influence operations: Examples of internet and media misuse. https://www.inss.re.kr/ en/News/bbs/news_en_view.do?nttId=41037284, 2024. [22] Jiwan Jeong, Jeong-han Kang, and Sue Moon. Identifying and quantifying coordinated manipulation of upvotes and downvotes in Naver News comments. In Proceedings of the International AAAI Conference on Web and Social Media, 2020. [23] Younghoon Jeong, Juhyun Oh, Jongwon Lee, Jaimeen Ahn, Jihyung Moon, Sungjoon Park, and Alice Oh. KOLD: Korean offensive language dataset. In Conference on Empirical Methods in Natural Language Processing, 2022. [24] Shagun Jhaver, Amy Bruckman, and Eric Gilbert. Does transparency in moderation really matter? user behavior after content removal explanations on Reddit. Proceedings of the ACM on Human-Computer Interaction, 2019. [25] Maryanne Kelton, Michael Sullivan, Emily Bienvenue, and Zac Rogers. Australia, the utility of force and the society-centric battlespace. International Affairs, 2019. [26] Jaehong Kim, Chaeyoon Jeong, Seongchan Park, Meeyoung Cha, and Wonjae Lee. How do moral emotions shape political participation? a cross-cultural analysis of online petitions using language models. In Findings of the Association for Computational Linguistics, 2024. [27] Klim Kireev, Yevhen Mykhno, Carmela Troncoso, and Rebekah Overdorf. Characterizing and detecting propaganda-spreading accounts on Telegram. In Proceedings of the 34th USENIX Conference on Security Symposium, 2025.
[28] Korea Press Foundation. 2025 Media Users in Korea (2025 언론수용자 조사). Technical report, Korea Press Foundation, Seoul, South Korea, December 2025. Available at https://www.kpf.or.kr/front/ research/consumerDetail.do?seq=600224.
[40] Mohammad Hammas Saeed, Shiza Ali, Jeremy Blackburn, Emiliano De Cristofaro, Savvas Zannettou, and Gianluca Stringhini. Trollmagnifier: Detecting statesponsored troll accounts on Reddit. In IEEE symposium on security and privacy, 2022.
[29] Junbum Lee. KcBERT: Korean comments BERT. In Proceedings of the 32nd Annual Conference on Human and Cognitive Language Technology, 2020.
[41] Daniel Thilo Schroeder, Meeyoung Cha, Andrea Baronchelli, Nick Bostrom, Nicholas A Christakis, David Garcia, Amit Goldenberg, Yara Kyrychenko, Kevin Leyton-Brown, Nina Lutz, et al. How malicious AI swarms can threaten democracy. Science, 2026.
[30] Junbum Lee. KcELECTRA: Korean comments ELECTRA. https://github.com/Beomi/KcELECTRA, 2021. [31] Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019. [32] Scott M. Lundberg and Su-In Lee. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems, 2017. [33] Muhammad Shujaat Mirza, Labeeba Begum, Liang Niu, Sarah Pardo, Azza Abouzied, Paolo Papotti, and Christina Pöpper. Tactics, threats & targets: Modeling disinformation and its mitigation. In Network and Distributed System Security Symposium, 2023. [34] Sarah Myers West. Censored, suspended, shadowbanned: User interpretations of content moderation on social media platforms. New Media & Society, 2018. [35] Nic Newman, Arguedas Ross Arguedas, Craig T Robertson, Rasmus Kleis Nielsen, and Richard Fletcher. Digital news report 2025. Reuters Institute for the Study of Journalism, 2025. [36] Leslie E Papke and Jeffrey M Wooldridge. Econometric methods for fractional response variables with an application to 401 (k) plan participation rates. Journal of Applied Econometrics, 1996. [37] Seongchan Park, Jaehong Kim, Hyeonseung Kim, Heejin Bin, Sue Moon, and Wonjae Lee. Moral outrage shapes commitments beyond attention: Multimodal moral emotions on YouTube in Korea and the US. In Proceedings of the ACM Web Conference, 2026. [38] Sungjoon Park, Jihyung Moon, Sungdong Kim, et al. KLUE: Korean language understanding evaluation. In 35th Conference on Neural Information Processing Systems Track on Datasets and Benchmark, 2021. [39] Ruben Recabarren, Bogdan Carbunar, Nestor Hernandez, and Ashfaq Ali Shafin. Strategies and vulnerabilities of participants in Venezuelan influence operations. In 32nd USENIX Security Symposium, 2023.
[42] Almog Simchon, William J Brady, and Jay J Van Bavel. Troll and divide: the language of online polarization. PNAS Nexus, 2022. [43] Kate Starbird. Disinformation’s spread: bots, trolls and all of us. Nature, 2019. [44] Ian Sullivan. How China fights in large-scale combat operations. Technical report, U.S. Army Training and Doctrine Command (TRADOC G-2), 2025. [45] Maxim Tkachenko, Mikhail Malyuk, Andrey Holmanyuk, and Nikolai Liubimov. Label Studio: Data labeling software. https://github.com/ HumanSignal/label-studio, 2020–2025. [46] U.S. Senate Select Committee on Intelligence. Report on Russian active measures campaigns and interference in the 2016 U.S. election, volume 2: Russia’s use of social media, with additional views. Technical report, United States Senate, 2019. [47] Jay J Van Bavel, Claire E Robertson, Kareena Del Rosario, Jesper Rasmussen, and Steve Rathje. Social media and morality. Annual Review of Psychology, 2024. [48] Claire Wardle and Hossein Derakhshan. Information disorder: Toward an interdisciplinary framework for research and policymaking. Council of Europe, 2017. [49] Evan J Williams. The comparison of regression variables. Journal of the Royal Statistical Society: Series B (Methodological), 1959. [50] Madelyne Xiao and Jonathan Mayer. SoK: Machine learning for misinformation detection. In 34th USENIX Security Symposium, 2025. [51] Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. Predicting the type and target of offensive posts in social media. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2019.
Troll Dataset Summary
ment overlap with known trolls and content-level prediction patterns.
We define the 70 accounts publicly released by the Institute for National Security Strategy in South Korea [21] as known trolls, accounts flagged by our user-level detector as modeldetected trolls, and all remaining accounts after excluding these two groups as non-troll users. Table 6: Troll dataset summary by troll category. The total number of comments by known trolls increased from an initial 356,378 to include an additional 3,799 comments identified from the candidate pool. The total number of non-troll users increased from the initial 81 users by adding 4,023,682 additional users. Troll Category
Users
Comments
Known Troll Model-Detected Troll Non-Troll
70 23,998 4,023,763
360,177 14,698,055 97,600,322
Total
4,047,831
112,658,554
Table 7: Hierarchical distribution of classes in the training data (N = 49, 745). The table presents the count and proportion of each class within its respective level. For Levels 2 and 3 (multi-label), only applicable classes are listed. “MNS” refers to a major neighboring state. Level
Class
Count
Proportion
Level 1
Foreign State-suspected Not Foreign State-suspected
21,627 28,118
0.43 0.57
Other-Condemning Other-Praising
17,825 1,608
0.92 0.08
Condemning Korea Condemning MNS Condemning Rivals Condemning Partners Condemning Others
12,849 3,184 2,614 112 349
0.67 0.16 0.14 0.01 0.02
Praising Korea Praising MNS Praising Rivals Praising Partners Praising Others
89 1,424 86 76 22
0.05 0.84 0.05 0.05 0.01
Troll Non-Troll
17,751 31,994
0.36 0.64
Level 2
Level 3
Final
B
Comparative Analysis of Troll Groups
We analyze the differences between model-detected trolls and non-troll users, showing that model-detected trolls are more similar to known trolls than non-troll users in terms of com-
Comment Overlap with Known Trolls. Exact comment overlap with known trolls is included as an input feature, but it contributes negligibly in our model (ranked 24th out of 24 by mean absolute SHAP; mean(|SHAP|) = 0), making it unlikely to be a primary driver of detection. We therefore report this analysis to contextualize the prevalence and qualitative nature of duplicated texts, rather than as an independent validation signal. Using the same criteria as the Exact match overlap with the known trolls feature (comments with ≥10 characters and ≥3 tokens), we measure exact-text overlap with comments authored by known trolls. Figure 9 shows that model-detected trolls exhibited substantially higher overlap with known trolls: 3.88% (930 of 23,998) of model-detected troll users shared comments with at least one known troll user, compared to 0.17% (6,941 of 4 million) of non-troll users. Model-detected trolls also overlapped with more known trolls and shared more matching comments on average than non-troll users (Mann-Whitney U test, p < 0.001).
7.31
6.25
6 4
***
Model-Detected Troll *** Non-Troll
8 Percentage (%)
A
3.88
2
rlap
Ove
U
sers
sers
pU
erla
. Ov Avg
0.22
0.22
0.17
0
rlap
ve g. O
m
Com
ents
Av
Figure 9: Overlap with known trolls by troll group. Overlap Users refers to the percentage of users in each group whose comments overlap with those of known trolls. Average Overlap Users and Average Overlap Comments refer to the average number of known troll users and comments, respectively, that overlap with each user in a group. Asterisks (***) above each class indicate a significant difference in means based on the Mann-Whitney U test (p < 0.001).
Content-Level Probabilities. Appendix Table 8 shows average user-level probabilities, conducted using the same method as in Table 4. Model-detected troll users exhibited significantly higher probabilities across all six classes (Troll Comments: t = 187.19; Foreign State-suspected: t = 195.42; Condemning Korea: t = 184.25; Condemning Rivals: t = 48.89; Praising MNS: t = 27.98; Praising Partners: t = 27.19; all p < 0.001).
Table 8: User-level average percentages computed from content-level probabilities across six classes for modeldetected troll and non-troll users. Model-detected trolls consistently exhibit higher probabilities across all six classes that are considered important within the hierarchical detection framework. “MNS” refers to a major neighboring state. Class
Model-Detected Troll
Non-Troll
Troll Comments Foreign State-suspected Condemning Korea Condemning Rivals Praising MNS Praising Partners
57.66% 75.20% 55.20% 1.48% 0.59% 0.14%
18.27% 27.89% 17.43% 0.45% 0.20% 0.06%
C
Comparative Follow Relationships
We compare the follow relationships of 70 known troll accounts and 81 non-troll accounts based on their collected followers and followings. The known troll accounts have 4,309 followers and 2,375 followings, whereas the non-troll accounts have 1,313 followers and 55 followings. A detailed comparison is presented in Table 9. Activity Known trolls are far more active in forming follow relationships, with on average 2.7 times more followers and 35 times more followings than non-troll accounts. The substantially higher number of followings points to an outward-facing strategy for reaching other users, communities, or potentially influential accounts. In contrast, non-troll users maintain far fewer followings, suggesting that troll following patterns are less consistent with ordinary social ties and more consistent with strategic outreach. Network Known trolls form a cohesive network through mutual following. On average, each troll account follows and is followed by two other known trolls, with some mutually connected to up to nine. Reciprocal following occurs 138 times more often among known trolls than among non-troll accounts. Non-troll accounts show no follow relationships with known trolls. This pattern indicates that known troll accounts operate as an internally connected cluster rather than as isolated accounts, a structure that may support coordinated activity or amplification. Reciprocity Known trolls also exhibit imbalanced follow relationships. We define the follow ratio as the number of followings divided by the number of followers plus one. This ratio is five times higher for known trolls than for non-troll accounts. This outbound-heavy pattern suggests a strategy of reaching many users, which may help create follow-backs and reciprocal connections.
D
Robustness to Paraphrasing
Table 10: Robustness to LLM paraphrasing. BLEU and SBERT cosine similarity are computed between each original test comment and its paraphrase. Macro-F1 is averaged across classes. Performance remains close to the original, and lexical overlap is near zero. Model
Condition
BLEU
SBERT
Macro-F1
Original
–
–
–
0.8411
GPT-5 GPT-5 GPT-5-mini GPT-5-mini
base meta base meta
0.013 0.010 0.016 0.016
0.8148 0.8016 0.8274 0.8120
0.8209 0.8540 0.8088 0.8404
A content-level detector may exploit actor- or wordingspecific surface signatures rather than the target concepts, which could inflate performance. To test for this, we adapted the LLM-bypassing protocol and paraphrased the 1,000 human-annotated test comments using GPT-5 and GPT-5mini under two conditions [27]. The base condition provides only the article title and the comment, requesting a rewrite in the style of a native Korean commenter that preserves meaning, stance, and tone. The meta-informed condition additionally provides the annotated moral emotion and the condemning and praising targets, asking the model to preserve these attitudes while rewriting. Because span-level ground truth is not preserved under paraphrasing, we evaluate at the class level and report macro-F1 averaged across classes. The paraphrases substantially alter surface form, with a mean BLEU of 0.0134 relative to the originals, while preserving semantic content, with a mean SBERT cosine similarity of 0.8140. Despite this rewriting, macro-F1 remains comparable to the original 0.8411 across all four settings (0.8088–0.8540; Table 10). This suggests that the content-level detector does not rely primarily on memorized surface patterns, but instead captures the intended moral-emotional and target-country concepts.
Table 9: Follow network analysis of known troll and non-troll users: Comparative metrics Followers Metric
Average Min Max
Followings
Reciprocal Follows
Follow Ratio
Known Troll Follow
Known Troll
NonTroll
Known Troll
NonTroll
Known Troll
NonTroll
Known Troll
NonTroll
Known NonKnown NonTroll Troll Troll Troll Followers Followers Followings Followings
63.3676 2 472
23.4464 1 238
34.1324 0 499
0.9821 0 10
2.4559 0 28
0.0179 0 1
0.5954 0 7.9677
0.1166 0 1.7500
2.2353 0 11
0 0 0
2.1324 0 27
0 0 0
Figure 10: Original Korean version of the model output shown in Figure 3, with rationale-related spans highlighted. Each span corresponds to a predicted label: Foreign state–suspected (filled orange), Other-condemning (filled green), Other-praising (filled sky blue), Condemning Korea (outlined navy), and Praising MNS (outlined gold). Span opacity indicates the predicted probability for that label. “MNS” refers to a major neighboring state.
Figure 11: Like ratio by rhetoric across Korean administrations. Across five administrations, Condemning Korea yields higher like ratios than praising rhetoric, ranking first in four of five administrations; the only exception is the Lee administration, where Condemning Rival ranks first. Praising MNS and Praising Partner are consistently lower than condemning rhetoric across regimes. Praising Partner is not shown for the Roh and Lee administrations due to the absence of data. Error bars indicate mean ± 95% CI; the dashed line marks 0.5 (neutral: likes equal dislikes).
Table 11: Top 10 most frequent target spans in high-visibility Condemning Korea comments (Like Ratio = 1.0) by administration. Politically salient figures, including elected leaders, candidates, and other politically connected individuals, are shown in bold with colored backgrounds. Blue indicates left-leaning figures, and red indicates right-leaning figures. Roh
Lee
Park
Moon
Yoon
(Apr, 2006 – Feb, 2008)
(Feb, 2008 – Feb, 2013)
(Feb, 2013 – May, 2017)
(May, 2017 – May, 2022)
(May, 2022 – Mar, 2025)
Korea
Korea
Korea
Moon Jae-in
Lee Jae-myung
한국
한국
한국
문재인
이재명
Roh Moo-hyun
Rep. of Korea
Park Geun-hye
Mun Jae-ang (Derogatory)
Yoon (Hanja)
노무현
대한민국
박근혜
문재앙
尹
Our country
Our country
Rep. of Korea
Pres. Moon (Hanja)
Yoon Suk-yeol
우리나라
우리나라
대한민국
文대통령
윤석열
Rep. of Korea
LG Electronics
Park (Hanja)
Cho Kuk
Democratic Party
대한민국
LG
朴대
조국
민주당
Kim Dae-jung
Moon Jae-in
Choi Soon-sil
Democratic Party
Pres. Yoon
김대중
문재인
최순실
민주당
윤 대통령
Samsung
Korea (+ Particle)
Pres. Park (Hanja)
Moon (Hanja)
Kim Keon-hee
삼성
한국은
朴대통령
文
김건희
Lee Myung-bak
Jeolla-do
Hell Joseon
Korea
Moon Jae-in
이명박
전라도
헬조선
한국
문재인
Chosun Ilbo
Lee Myung-bak
Korea (+ Particle)
Yoon Suk-yeol
Rep. of Korea
조선일보
이명박
한국은
윤석열
대한민국
Korea (+ Particle)
Korea (short)
Sewol Ferry
Rep. of Korea
Han Dong-hoon
한국은
한
세월호
대한민국
한동훈
South Korea
Journalist
Our country
Lee Jae-myung
Yoon
남한
기자
우리나라
이재명
윤
* Note: Hanja denotes Chinese characters used in Korean; Particle refers to Korean grammatical markers (e.g., -eun/-neun).