Conceptio › Archive › arXiv CS
arXiv CSopen access

SoK: Analysis of Privacy Risks and Mitigation in Online Propaganda Detection through the PROMPT Framework

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

arXiv:2604.17788v1 [cs.CR] 20 Apr 2026

SoK: Analysis of Privacy Risks and Mitigation in Online Propaganda Detection through the PROMPT Framework Dhiman Goswami

Al Nahian Bin Emran

George Mason University Fairfax, Virginia, United States of America [email protected]

George Mason University Fairfax, Virginia, United States of America [email protected]

Md Hasan Ullah Sadi

Sanchari Das

George Mason University Fairfax, Virginia, United States of America [email protected]

George Mason University Fairfax, Virginia, United States of America [email protected]

Abstract Online propaganda detection pipelines expose measurable privacy risks at multiple stages including data collection, feature extraction, and model inference. We conduct a structured analysis of 162 peer-reviewed studies and formalize the problem using the Propaganda Risk Online Mitigation and Privacy-preserving Tactics (PROMPT) framework. PROMPT models risks 𝑅 and mitigation strategies 𝑆 through a mapping M : 𝑅 → 𝑆 guided by a utility function 𝛼 · PrivacyGain(𝑠 𝑗 ) − 𝛽 · PerfLoss(𝑠 𝑗 ) − 𝛾 · Cost(𝑠 𝑗 ), with tunable (𝛼, 𝛽, 𝛾) enabling stakeholders to balance privacy, accuracy, and deployment costs. To assess practical adoption, we introduce a compliance score that quantifies the alignment of existing methods with GDPR, CCPA etc. requirements. Our evaluation shows that many widely used pipelines remain non-compliant, particularly in metadata handling and user-level aggregation. We further present empirical fine-tuning experiments on transformer-based encoders and decoders under synthetic perturbation, demonstrating a monotonic privacy–utility trade-off: with 𝑞 = 0.05 performance decreased by 1–2% F1 , while at 𝑞 = 0.20 the reduction reached 13–14%. These results establish quantitative baselines for privacy costs in propaganda detection. Our contributions include a formal risk-to-defense mapping, a compliance-oriented auditing metric, and experimental evidence of privacy–performance trade-offs, providing a technical foundation for building regulation-compliant and privacy-aware detection systems.

CCS Concepts • Human-centered computing → Usability testing; • Security and privacy → Privacy protections.

Keywords Propaganda Detection, privacy, anonymization, Privacy Preserving NLP Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. ASIA CCS ’26, Bangalore, India © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 979-8-4007-2356-8/26/06 https://doi.org/10.1145/3779208.3807832

ACM Reference Format: Dhiman Goswami, Al Nahian Bin Emran, Md Hasan Ullah Sadi, and Sanchari Das. 2026. SoK: Analysis of Privacy Risks and Mitigation in Online Propaganda Detection through the PROMPT Framework. In ACM Asia Conference on Computer and Communications Security (ASIA CCS ’26), June 1–5, 2026, Bangalore, India. ACM, New York, NY, USA, 18 pages. https://doi.org/10.1145/3779208.3807832

1

Introduction

In the digital age, individuals increasingly rely on platforms such as Meta (formerly Facebook), X (formerly Twitter), YouTube, and major news outlets like BBC, CNN, and Sky News to share and consume information [126]. While these platforms enable global connectivity and real-time access to news, they also act as fertile grounds for the rapid spread of propaganda. Manzoor et al. define propaganda as “a deliberate, conscious, malicious, and cunning effort by a group, organization, or individual to control and influence public beliefs and actions through selected truth, mass communication, and personal contact” [107]. Prior studies further show that propaganda spreads faster and more widely than factual information on social media [130, 168, 169, 178, 197]. The vast scale of user-generated content and its rapid circulation make effective detection increasingly challenging [77, 97]. Propaganda transcends linguistic and platform boundaries [127, 160], while human cognitive biases and engagement-driven algorithms exacerbate its reach [111, 113]. To address these challenges, researchers employ machine learning (ML) models such as BERT, RoBERTa, and GPT-based architectures to detect common techniques including framing, persuasion, and disinformation [95, 133]. Progress has been enabled by curated datasets capturing political, military, and commercial disinformation [69], such as AG-News, ArAIEval, the DARPA Twitter Bot Challenge, and Kaggle’s fake news dataset. Multilingual corpora like EUvsDisinfo highlight statesponsored campaigns, e.g., pro-Kremlin narratives [100]. Yet, despite their utility, such resources raise serious privacy concerns [47], as large-scale data collection introduces risks of surveillance, reidentification, and misuse [156, 161]. Detection techniques that exploit behavioral or biometric signals may further infringe on individual rights [156]. While privacy-preserving approaches such as DP, anonymization, and secure multi-party computation have been explored [103, 162], most studies prioritize accuracy, leaving privacy and compliance under-addressed [3, 91, 110, 134, 158, 162, 181].

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

To examine this gap, we analyze 162 peer-reviewed studies spanning datasets, ML methodologies, and privacy-preserving technologies. Building on this analysis, we propose the Propaganda Risk Online Mitigation and Privacy-preserving Tactics (PROMPT) framework, which maps privacy threats (e.g., metadata leakage, behavioral inference) to mitigation strategies (e.g., DP, federated learning, secure multi-party computation) and integrates regulatory principles and best practices to guide the design of performant, privacy-preserving, and compliant detection systems. Our study is structured around the following research questions: • RQ1: What privacy risks arise during data collection, model training, and deployment in propaganda detection? • RQ2: How well do current propaganda detection methods incorporate privacy-preserving techniques, and how effectively do they mitigate risks across different data modalities, platforms, and languages? • RQ3: What privacy-preserving methods can be integrated into propaganda detection systems to protect user data while maintaining detection performance? Our work offers the following key contributions: • Here we report on the first SoK on the intersection of propaganda detection and privacy, based on a systematic analysis of 𝑁 = 162 peer-reviewed studies covering datasets, detection methods, and privacy-preserving techniques. Let 𝐹 = {𝑑, 𝑚, 𝑝} denote features for datasets (𝑑), methods (𝑚), and privacy techniques (𝑝). The adoption ratio is defined as |𝑃𝑥 | , 𝑥 ∈ {𝑑, 𝑚, 𝑝}, 𝑁 providing a reproducible baseline to compare uptake across areas (e.g., metadata use, privacy-technique adoption). • We develop the PROMPT framework as a mapping M : 𝑅 → 𝑆, from risk categories 𝑅 = {𝑟 1, . . . , 𝑟𝑘 } (e.g., re-identification, metadata leakage) to mitigation strategies 𝑆 = {𝑠 1, . . . , 𝑠 ℓ } (e.g., DP, SMPC, FL). Defense selection is posed as utility optimization: h i 𝑠 ∗ (𝑟𝑖 ) ∈ arg max 𝛼 PrivacyGain(𝑠 𝑗 ) − 𝛽 PerfLoss(𝑠 𝑗 ) , 𝐶𝑥 ≜

𝑠 𝑗 ∈𝑆

with tunable (𝛼, 𝛽) reflecting deployment priorities. • For 𝑀 evaluated pipelines, we define the compliance score 𝑀

CompScore ≜

1 ∑︁ 1[Compliant(Method𝑖 )] , 𝑀 𝑖=1

and compile a best-practice dictionary B linking regulatory clauses (e.g., GDPR data minimization, PbD) to technical controls (e.g., DP noise calibration, 𝑘-anonymity, metadata filtering), offering actionable guidance. • We empirically observe a monotonic privacy–utility trade-off with increasing perturbation rate 𝑞, i.e., 𝐹 1 (𝑞) ≤ 𝐹 1 (0) and 𝐹 1 (𝑞 2 ) ≤ 𝐹 1 (𝑞 1 ) for 𝑞 2 ≥ 𝑞 1 in our setup.

2

Goswami et al.

Propaganda has been categorized into techniques such as whataboutism, straw man, red herring, bandwagon, reductio ad Hitlerum, exaggeration or minimization, thought-terminating cliché, causal oversimplification, appeal to fear or prejudice, black-and-white fallacy, name-calling or labeling, appeal to authority, loaded language, flag-waving, repetition, slogans, and doubt [33, 41, 177, 186]. These categories underpin computational approaches, which are benchmarked at span, sentence, and document levels. Early studies relied on traditional ML models such as Logistic Regression, SVM, Naïve Bayes, Decision Trees, and TF-IDF classifiers. Li et al. [102] applied a TF-IDF Logistic Regression with linguistic features to reduce identifiable data use, while Aggarwal and Sadana [4] employed SVM with oversampling. Da San Martino et al. [41] integrated n-grams with LSTM and RoBERTa for multi-label detection, and infrastructure-level features were also leveraged to identify disinformation websites [81]. Transformers advanced detection further. Kaczyński and Przybyła [87] proposed a BERT-based multi-task framework with credibility assessment, while Li et al. [104] fused RoBERTa and ResNet50 for multimodal meme detection, raising privacy concerns in image-text fusion. Fadel et al. [52] showed BERT and USE ensembles outperform traditional ML, and Fouad and Weeds [54] applied AraBERT with augmentation for Arabic propaganda, emphasizing risks in under-resourced contexts. Broader surveys highlight methodological and ethical challenges in misinformation research [205]. Other DL architectures also contributed. Blaschke et al. [26] combined Bi-LSTM with BERT embeddings for token-level detection while minimizing metadata use. Gupta et al. [65] proposed a CNN-LSTM hybrid with segmentation and entity recognition to mitigate stylistic profiling. Gundapu and Mamidi [62] explored multimodal meme detection with CNN-LSTM, while Alhabashi et al. [14] extended this with ResNet-50 and MARBERT for Arabic memes, stressing re-identification risks. Patil et al. [131] applied RNN-based sequence learning with BERT embeddings to news propaganda. User studies further reveal that misinformation warnings on video-sharing platforms may not always alter perceptions, complicating trust in detection systems [64]. Recently, LLMs such as GPT-3.5, GPT-4, XLM-RoBERTa, and mT5 have been explored. Szwoch et al. [180] reported GPT-4’s inconsistencies and false positives, while Hamilton et al. [67] introduced GPT-assisted annotation for scalable, privacy-preserving labeling. Bagdasaryan et al. [20] warned of “propaganda-as-a-service”via trigger phrases, and Morio et al. [121] showed ensembles of XLNet, RoBERTa, and GPT-2 improve precision while reducing privacy exposure. Aldabbas et al. [13] introduced MultiProp, a cross-lingual LLM framework leveraging augmentation and meta-learning. Despite these advances, significant privacy risks remain, including re-identification, metadata exposure, and profiling, particularly in cross-lingual and multimodal contexts involving large-scale user data and sensitive annotations.

Background

In this section, we outline foundational concepts and techniques for propaganda detection on online platforms, emphasizing privacy risks when ML and Deep Learning (DL) models process usergenerated content through data collection, model design, and NLPbased computation.

3

Method

We follow the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines [112] for conducting evidencebased literature review. Figure 1 provides the overview of our study.

SoK: Analysis of Privacy Risks and Mitigation in Online Propaganda Detection through the PROMPT Framework

Identification

Database Searching DBLP, ACM, IEEE, Springer, Elsevier ( 290 )

Screened Records Screening ( 457 )

Eligibility

Full-Text Accessed for Eligibility ( 171 )

Selected

Studies Included in the Survey ( 162 )

Additional Search Google Scholar ( 167 )

Excluded Records (Title Screening, Duplicate Removal) ( 286 )

Full-Text Article Excluded (not directly relevant) (9)

Figure 1: PRISMA Diagram: Stepwise Literature Collection, Filtering, and Selection Process.

3.1

Paper Identification

We systematically searched high-impact academic databases, including DBLP, ACM Digital Library, IEEE Xplore, Springer, and Elsevier, alongside top NLP, privacy, and HCI conferences (𝐴∗ and 𝐴 venues in CORE ranking 1 ). To ensure comprehensive coverage, we also retrieved relevant works from Google Scholar.

3.2

Paper Screening

We designed structured queries to identify studies at the intersection of propaganda detection and privacy concerns. The primary keywords included: propaganda, online propaganda, propaganda privacy, propaganda mitigation, with platform-specific terms such as propaganda in social media, propaganda in news media. To capture privacy-preserving approaches, we included terms such as differential privacy, federated learning, anonymization in propaganda detection. Despite rigorous filtering, potential biases in keyword and venue selection may have led to the exclusion of some studies. We included papers published until January 2025. The final dataset provides a comprehensive analysis of privacy risks and mitigation strategies in propaganda detection.

3.3

Paper Selection

We retrieved 457 research papers and applied a multi-stage filtering process aligned with best practices in systematic literature reviews. Title and Abstract Screening: First, we removed duplicates and irrelevant studies, retaining 286 unique paper. We then conducted manual abstract screening for relevance, excluding papers that lacked discussions on privacy risks, mitigation techniques, or ethical considerations. Three independent researchers conducted the screening and conducted a thematic analysis to obtain an interannotator agreement of 0.86 and a Cohen’s Kappa score of 0.72. In this phase we excluded 115 papers, leaving 171 for full-text review. Full-Text Review and Final Selection: We assessed methodological rigor, empirical validation, and contributions to privacyaware propaganda detection. Studies focusing on DP, federated learning, anonymization, and adversarial robustness were prioritized. Papers lacking substantial discussions on privacy, empirical 1 https://portal.core.edu.au/conf-ranks/

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

analysis, or clear methodologies were excluded. Full-text screening involved three researchers, with a fourth researcher resolving disputes in ambiguous cases. In this phase we excluded nine papers, resulting in a final dataset of 162 papers.

4

PROMPT Framework

We propose the PROMPT framework that provides a structured methodology for analyzing propaganda detection systems under privacy, security, and compliance constraints. It integrates six core components: Propaganda Types and Techniques, Propaganda Analysis, Propaganda Privacy and Security Concerns, Propaganda Mitigation, Regulatory Compliance, and Ethical Considerations. Unlike prior descriptive taxonomies, PROMPT formalizes the connection between propaganda techniques, computational risks, and privacypreserving defenses. It draws on and extends prior work: the CrossCultural Privacy Framework [190] that informs PROMPT’s treatment of legal and cultural constraints, NLP-PRISM [60] that describes privacy vulnerabilities with framework integration of different tasks related to social media, the OVERRIDE framework [143] that motivates its data handling mechanisms, and PRAF [157] that reinforces the integration of regulatory compliance and ethical oversight. Together, these elements define a comprehensive methodology that links detection pipelines to measurable privacy risks and mitigation strategies. Figure 2 illustrates the PROMPT architecture, and we detail its components below. Propaganda Types and Techniques: We categorize propaganda into political, religious, social, economic, cultural, and public health domains each defined by recurring rhetorical and linguistic strategies [109]. Let 𝑇 = {𝑡 1, 𝑡 2, . . . , 𝑡𝑘 } denote the set of techniques (e.g., name-calling, repetition, flag-waving). For each domain 𝐷 𝑗 , |𝑇𝐷 |

we define a coverage score, Coverage(𝐷 𝑗 ) = |𝑇 𝑗| , where 𝑇𝐷 𝑗 ⊆ 𝑇 is the set of techniques observed in that domain and |𝑇 | is the number of total number of domains defined by [42]. This formulation enables quantitative comparison across domains, providing measurable features that can be embedded into detection models. For example, political propaganda exhibits high technique diversity (|𝑇𝐷 𝑗 | large), while public health propaganda relies more heavily on a smaller subset such as exaggeration or fear appeals. Within PROMPT, these techniques serve as the input layer of the framework, linking observed strategies to privacy and security risks 𝑅, which are later mapped to mitigation strategies 𝑆 via the PROMPT risk-to-defense function M : 𝑅 → 𝑆. By grounding qualitative categories in formal representations, the framework allows reproducible benchmarking and facilitates the design of machine learning pipelines that explicitly consider adversarial robustness and privacy-aware constraints. Propaganda Analysis: Propaganda analysis in PROMPT is modeled as a multi-stage pipeline where each stage exposes distinct computational and privacy risks. Let P = {𝑝 1, 𝑝 2, . . . , 𝑝𝑛 } denote the set of propaganda instances. The pipeline can be expressed as C

S

A

D

𝑝𝑖 −−→ 𝑥𝑖 −−→ 𝑦𝑖 −−−→ 𝑧𝑖 −−−→ 𝑜𝑖 , where C is data collection and preprocessing, S is storage and security, A is analysis and computational modeling, and D is distribution and networking. Each stage is associated with measurable risks

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

Goswami et al.

Propaganda Risk Online Mitigation and Privacy-preserving Tactics (PROMPT) Propaganda Types & Techniques Political Name calling or labeling, Repetition, Appeal to fear/prejudice, Causal oversimplification, Black-andwhite fallacy, dictatorship, Whataboutism, straw man, red herring, Bandwagon, reductio ad hitlerum

Religious

Social

Loaded language, Appeal to authority, Thoughtterminating cliché, Black-andwhite fallacy, dictatorship, Repetition

Name calling or labeling, Flagwaving, Slogans, Whataboutism, straw man, red herring, Bandwagon

Economic Exaggeration or minimization, Doubt, Appeal to authority, Causal oversimplification, Slogans

Public Health

Cultural Loaded language, Flagwaving, Slogans, Thoughtterminating cliché

Exaggeration or minimization, Doubt, Appeal to fear/prejudice, Appeal to authority, Repetition

Propaganda Analysis Data Collection and Preprocessing

Data Storage and Security

Propaganda Privacy & Security Concerns

Data Analysis and Computational Risk

Data Distribution and Network Risk

Annotation, Surveillance, Facial Recognition, Consent, Anonymity, LLM Annotation, Translated Data, User Generated Content, Cross Lingual Transfer Learning, Data Augmentation

Data Security, Reidentification, Metadata Privacy Risk, Multimodal Data Privacy

Misclassification, Algorithmic Bias, Misuse, False Information, LLM Privacy Violation, Image Processing Privacy Risk, Label Overprediction, Irrelevant Span Detection, Random Truncation, Undefined Content, Threat

Social Media Manipulation, SEO, Clickbait, Information Warfare, Deepfake,Centrality Attack, Community Clustering, Bridge Node Exploitation, Sybil Attack

Differential Privacy, Federated Learning, Synthetic Data Generation, Homomorphic Encryption, Human-in-the-loop, PBD

SMPC, Blockchain-Based Storage, KAnonymity and L-Diversity, Metadata Privacy Filters, PETs

PPML, Adversarial Machine Learning Defenses, XAI, Privacy-Preserving Image Processing, PETs, PbD

Fact Checking Integration, BlockChain Based Verification, AI-based Content Moderation, Graph Anomaly Detection

Propaganda Mitigation

GDPR, CCPA, PIPL, PDPA, HIPAA

Informed Consent & Transparency, Bias & Discrimination, Surveillance & Privacy, Data Ownership & Control, Cross-Border Data Ethics

ISO 27001, NIST, PCI DSS, GDPR, HIPPA, FedRAMP

Encryption & Data Protection, Government & Corporate Surveillance Risks, Data Breach Accountability, Right to Be Forgotten

EU AI Act, IEEE AI Ethics, Basel III, FSB AI Guidelines

Algorithmic Fairness & Bias, Automated Decision-Making Risks, Data Manipulation & Misinformation, Commercial Exploitation of User Data

CISA, FARA, Honest Ads Act, DSA, GDPR, UNESCO Digital Platform Governance

Censorship, Surveillance, Transparency in Social Engineering, Responsible AI

Regulatory Compliance

Ethical Considerations

Figure 2: PROMPT Framework Highlighting Privacy Risks and Mitigation Strategies 𝑅 = {𝑟𝐶 , 𝑟 𝑆 , 𝑟 𝐴 , 𝑟 𝐷 } and potential mitigations 𝑆 = {𝑠𝐶 , 𝑠𝑆 , 𝑠𝐴 , 𝑠𝐷 } proposed by us and defined later in the framework. Data Collection and Preprocessing (C) involves acquiring multimodal propaganda artifacts (text, images, video) and transforming them into feature representations 𝑥𝑖 . Risks include surveillance, lack of consent, and re-identification. Data Storage and Security (S) ensures persistence and access control for feature sets 𝑥𝑖 → 𝑦𝑖 , where threats include metadata leakage, multimodal correlation attacks, and unauthorized retrieval. Data Analysis and Computational Risk (A) applies machine learning classifiers 𝑓𝜃 (𝑦𝑖 ) = 𝑧𝑖 , with vulnerabilities including algorithmic bias, adversarial manipulation, and misclassification. Data Distribution and Network Risk (D) models how propaganda outputs 𝑧𝑖 spread across networks to yield observable influence 𝑜𝑖 , with threats such as Sybil attacks, community clustering, and coordinated information warfare. This structured representation transforms analysis from a descriptive sequence into a formal pipeline with explicit risk points, allowing integration with PROMPT’s mapping M : 𝑅 → 𝑆 to evaluate and select mitigations for each stage. Propaganda Privacy Evaluation: Privacy and security challenges in propaganda detection can be represented as attack surfaces that expose user data and model integrity. We define a risk vector 𝑅® = ⟨𝑟 ident, 𝑟 meta, 𝑟 bias, 𝑟 net ⟩, where each component denotes

identity leakage, metadata exposure, algorithmic bias, and network manipulation based on the four core steps (C, S, A, D) . The overall magnitude of vulnerability can be approximated as the 𝐿2 norm √︃ ® 2 = 𝑟2 + 𝑟2 + 𝑟2 + 𝑟2 . ∥ 𝑅∥ meta net ident bias Identity and Consent Risks occur when annotation, LLM translation, or cross-lingual transfer enable re-identification of individuals without proper consent [173]. These risks are measurable by record-linkage accuracy against anonymized datasets. Metadata Risks emerge from insecure storage or multimodal correlation [120]. They can be quantified using the mutual information 𝐼 (𝑀; 𝑈 ) between metadata 𝑀 and user identity 𝑈 . Algorithmic Risks involve misclassification, adversarial perturbations, and censorship errors [95]. Bias can be evaluated with fairness gaps such as ΔEO = |𝑃 (𝑦ˆ = 1 | 𝑦 = 1, 𝑔 = 0) − 𝑃 (𝑦ˆ = 1 | 𝑦 = 1, 𝑔 = 1)| , where 𝑔 is a sensitive attribute and the fairness gap is the absolute value of the difference between the probability of classification with and without sensitive content exposure. Network Risks stem from Sybil attacks, deepfakes, and coordinated campaigns [24]. Their severity can be modeled by an amplification ratio 𝐴=

Reachpropaganda , Reachbaseline

SoK: Analysis of Privacy Risks and Mitigation in Online Propaganda Detection through the PROMPT Framework

which measures the inflation of dissemination relative to organic spread through network. By associating each risk with a measurable quantity such as linkage accuracy, mutual information, fairness gaps, or amplification ratios, PROMPT provides a foundation for auditing vulnerabilities with quantitative metrics instead of qualitative descriptions. Propaganda Mitigation: PROMPT models mitigation as a multi-objective optimization problem that balances privacy, accuracy, and system cost. Let 𝑅 = {𝑟 1, . . . , 𝑟𝑘 } denote risks and 𝑆 = {𝑠 1, . . . , 𝑠 ℓ } available defenses such as DP, federated learning, SMPC, or adversarial ML defenses. For each pairing (𝑟𝑖 , 𝑠 𝑗 ) we define a mitigation utility: 𝑈 (𝑟𝑖 , 𝑠 𝑗 ) = 𝛼 · PrivacyGain(𝑠 𝑗 ) − 𝛽 · PerfLoss(𝑠 𝑗 ) − 𝛾 · Cost(𝑠 𝑗 ), where 𝛼, 𝛽, 𝛾 are deployment-specific weights. Optimal mitigation is obtained as 𝑠 ∗ (𝑟𝑖 ) = arg max 𝑈 (𝑟𝑖 , 𝑠 𝑗 ). 𝑠 𝑗 ∈𝑆

This abstraction covers a wide spectrum of techniques. For example, DP increases PrivacyGain but induces higher PerfLoss, while blockchain-based verification introduces computational Cost. Integration of fact-checking, explainable AI, and human-in-the-loop review can be treated as constraints that improve contextual reliability rather than raw utility. By formalizing mitigations as optimization over 𝑈 (𝑟𝑖 , 𝑠 𝑗 ), PROMPT provides a structured approach to selecting defenses that align with both security requirements and system-level trade-offs. Regulatory Compliance: PROMPT incorporates compliance as a measurable score that evaluates how detection pipelines adhere to legal and industry standards. Let L = {ℓ1, . . . , ℓ𝑚 } be relevant obligations (e.g., GDPR consent, CCPA opt-out, ISO 27001 security control). For a given method 𝑓 , compliance is defined as 𝑚 1 ∑︁ 1[𝑓 satisfies ℓ 𝑗 ], CompScore(𝑓 ) = 𝑚 𝑗=1 where 1 is the indicator function. This allows comparison of pipelines by their degree of alignment with privacy laws, security standards, and AI governance frameworks. Beyond static measurement, compliance can be treated as a constraint in system design, requiring that selected defenses 𝑆 maximize utility 𝑈 (𝑟𝑖 , 𝑠 𝑗 ) subject to CompScore(𝑓 ) ≥ 𝜏, where 𝜏 is a regulatory threshold. For our analysis, we experimentally find and set to 0.319. This formulation elevates regulatory texts such as GDPR, CCPA, and the EU AI Act from descriptive guidelines into quantitative constraints that directly shape technical system choices. Ethical Considerations: We operationalize fairness as the proportion of safeguards relative to total ethical factors (safeguards + risks). Formally, let |safeguards| |risks| 𝐹= , 𝐵= . |safeguards| + |risks| |safeguards| + |risks| An ethically sound system minimizes 𝐵 while ensuring 𝐹 ≥ 𝜏 where 𝜏 is an empirical fairness threshold inferred from the corpus (see Table 5). In practice, this framing corresponds to enforcing informed consent during data collection, limiting surveillance, preserving user ownership of generated content [78], and implementing encryption and audit protocols at the storage level [25]. At the model level, bias mitigation and explainability methods

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

reduce discriminatory outcomes [37], while at the distribution level, safeguards include transparency reporting, countermeasures against censorship, and responsible governance of automated moderation [165].

4.1

Threat Model

In PROMPT we assume adversaries A characterized as follows: Adversary Types: (i) Platform providers with full access to raw user data and model pipelines, (ii) external observers or third parties with black-box or API-level access, and (iii) insiders or malicious communities with partial access to data, annotations, or system outputs. Capabilities: Adversaries may leverage data access, model querying, training data manipulation, and auxiliary information such as public profiles, external datasets, or cross-platform signals. Attack Surface: Data collection, storage, analysis, and distribution stages of the pipeline. Privacy Violations: Membership inference, re-identification, attribute inference and user profiling, metadata leakage, poisoned training data, adversarial examples, deepfakes, and coordinated network attacks. Each threat 𝑎𝑖 ∈ A is assigned likelihood 𝑃 (𝑎𝑖 ) and impact 𝐼 (𝑎𝑖 ). Í The expected system risk is 𝑖 𝑃 (𝑎𝑖 ) · 𝐼 (𝑎𝑖 ). Mitigation integrates technical defenses such as DP, FL, HE, privacy-preserving ML, blockchain verification, and fact checking, combined with compliance obligations under GDPR, CCPA, PIPL, and the EU AI Act. This approach reduces both technical and regulatory exposure.

4.2

Framework Integration

PROMPT integrates six layers into a pipeline: types and techniques, analysis, privacy and security concerns, mitigation, compliance, and ethics, as illustrated in Figure 2. Let X denote the propaganda inputs, 𝑅® the quantified risk vector, and 𝑆 the set of candidate defenses. The framework defines a mapping M : 𝑅® → 𝑆 that selects defenses to minimize risk while satisfying compliance and fairness. Each component of 𝑅® corresponds to measurable vulnerabilities - reidentification risk, metadata leakage, algorithmic bias, and network manipulation, which are explicitly linked to defenses such as DP, secure multiparty computation, or adversarially robust learning. This integration is not only conceptual (Figure 2) but also computational, where mitigation is expressed as an optimization problem balancing privacy gain, accuracy loss, and deployment cost. We operationalize this process in Algorithm 1, which details the steped procedure for estimating risks, filtering feasible defenses, scoring candidates by marginal utility, and iteratively updating the system until residual risks fall below tolerance. We now operationalize PROMPT’s formal definitions on the collected corpus of 162 studies. Specifically, the normalized risk components 𝑟 ident, 𝑟 meta, 𝑟 bias, 𝑟 net and the utility function 𝑈 (𝑠 𝑗 ) = 𝛼 · PrivacyGain(𝑠 𝑗 ) − 𝛽 · PerfLoss(𝑠 𝑗 ) −𝛾 · Cost(𝑠 𝑗 ) are instantiated with empirical evidence from our thematic coding and compliance analysis. This grounding enables direct comparison between the theoretical guarantees of PROMPT and the observed fairness, bias,

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

and performance trade-offs in transformer-based propaganda detection. We now turn to the Results section to quantify these outcomes.

5 Results 5.1 Propaganda Types and Techniques Table 1: Evaluation of Propaganda Types and Corresponding Technique Coverage

Propaganda Types Political Religious Social Economic Cultural Public health

Propaganda Technique (𝐶𝑜𝑣𝑒𝑟𝑎𝑔𝑒 (D 𝑗 )) 0.788 0.286 0.429 0.286 0.286 0.357

Propaganda detection involves analyzing types and techniques across political, religious, public health, social, economic, and cultural domains (Appendix Table 9). This representation facilitates quantitative comparison of coverage and diversity across domains. With the analysis of PROMPT framework, we find propaganda technique coverage across domains D 𝑗 (Table 1) using the formula from the Section 4. Political propaganda exhibits the highest technique coverage at 0.788, indicating greater methodological attention in detection pipelines. In contrast, religious, economic, and cultural propaganda each reflect relatively limited coverage (0.286), highlighting underexplored dimensions. Social propaganda (0.429) and public health propaganda (0.357) show moderate coverage, suggesting emerging but still incomplete methodological support. As highlighted by Da San Martino et al. [41] and Stefan-Yurii [177], political propaganda employs strategies such as name-calling, repetition, black-and-white fallacy, and whataboutism to manipulate public opinion. Moreover, it includes state-sponsored efforts, electoral influence, and military-driven campaigns to sway public opinion or consolidate power [89, 136]. Similarly, religious propaganda, as discussed by Ahmad et al. [8], uses appeals to authority, loaded language, and thought-terminating clich’es to influence belief. In the context of public health, Polonijo et al. [135] emphasize how exaggeration, minimization, and appeals to fear shape crisis narratives. Meanwhile, social propaganda, as described by Gundapu et al. [62], often utilizes flag-waving, slogans, and straw man arguments to polarize communities. Moreover, it involves persuasion-based techniques, false narratives, framing, bias-driven messaging, and multilingual campaigns designed to influence public sentiment and behavior [67, 129]. Additionally, economic and cultural propaganda, as noted by Muthukumar et al. [122], relies on causal oversimplifications, slogans, and appeals to identity-based biases. Moreover, cultural propaganda includes regional and multi-modal approaches such as Arabic memes and disinformation campaigns, which use both text and images to spread ideological messages [1, 212]. Finally, cross-domain techniques in propaganda detection and analysis include multilabel, multiview, and levels of analysis - such as span, sentence, and fragment-level - often employing computational models to identify and track the spread of propaganda across platforms [33, 186].

Goswami et al.

Beyond conventional strategies, modern propaganda techniques have evolved to exploit digital platforms. As highlighted by Hangloo et al., social media manipulation, SEO strategies, and clickbait mechanisms are now commonly used to amplify misleading narratives. Furthermore, the increasing prevalence of deepfake content and information warfare tactics exacerbates the challenge of identifying propaganda [68]. Given the growing complexity of these techniques, it is imperative to develop computational detection strategies that can adapt to domain-specific propaganda characteristics.

5.2

Propaganda Analysis & Evaluation

Propaganda analysis covers a comprehensive examination of how persuasive content is created, distributed, and detected in different contexts. This sub-section explores the linguistic and regional variations in this research, the digital platforms where misleading narratives proliferate, and the datasets that enable systematic study. It also reviews benchmark competitions that drive innovation and compares the computational methodologies-from traditional machine learning to cutting-edge transformer-based and multimodal approaches-used to identify and mitigate propaganda related content. By integrating these, we aim to provide a cohesive framework for understanding challenges and advancements in propaganda detection and analysis. The studies spanned across 23 languages, 18 online platforms, mapping key sources of propaganda. Moreover, 41 datasets were created, 5 benchmark competitions were organized and 79 computational methodologies were found in the literature. The necessity of a systematic propaganda privacy evaluation is clear, as only 5% of surveyed studies explicitly address privacy concerns in detection pipelines [41, 89, 123, 173, 174, 179]. The PROMPT framework operationalizes this evaluation across four critical stages : Data Collection and Preprocessing, Storage and Security, Analysis and Computational Risk, and Distribution and Network Risk highlighting where vulnerabilities emerge and how they accumulate. We used the formula from Section 4 and shown the results in Table 2. Data Collection and Preprocessing: Data collection for propaganda detection, involving textual and multimedia content from social media and news, raises concerns about surveillance, consent, and anonymity. Solopova et al. present surveillance risks from automated monitoring [173], while Kellner et al. highlight consent issues from collecting user data without approval [89]. Benzmüller et al. emphasize anonymity risks in multilingual analysis [174]. Annotation practices may introduce systemic bias, as shown in Syed et al.’s analysis [179], while Da San Martino et al. report risks from annotation workflows themselves [41]. LLM-assisted annotation raises further concerns, with Nabhani et al. cautioning that tools like GPT-4 may inadvertently expose sensitive content [123]. Linguistic diversity shapes privacy risks during collection: in Europe, propaganda detection spans English [41], French [123], German, Dutch [58], Italian, Spanish, Polish, Greek, Russian, Georgian [101], Romanian [173], Bulgarian, North Macedonian [106], Czech [76], Lithuanian [149], Serbian, and Ukrainian [177]. Conflicts such as the Russian–Ukrainian war accelerated multilingual annotation and transfer learning with human-in-the-loop checks.

SoK: Analysis of Privacy Risks and Mitigation in Online Propaganda Detection through the PROMPT Framework

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

Table 2: Evaluation of Risk Factors of Propaganda Detection. Risk Factors Cross-Lingual Transfer Learning Risk, Multimodal Data Privacy Violation, Community Clustering Exploitation, Image Processing Privacy Risk, User-Generated Content Risk, Social Media Manipulation, Irrelevant Span Detection Image Processing Privacy Risk, User-Generated Content Risk, Bridge Node Exploitation, Metadata Privacy Risk LLM Privacy Violation, Data Security Breach, Label Overprediction, Information Warfare, Facial Recognition, Random Truncation, Threat Generation, Undefined Content, False Information, Centrality Attack, Misclassification, Misuse Re-identification, Consent Violation, Algorithmic Bias, SEO Manipulation, Deepfake Attacks, Surveillance, Sybil Attack, Clickbait Normalized Cumulative Risk

In North America, research focuses on English [41], with preprocessing aimed at normalization, metadata tagging, and robustness. In South America, Spanish corpora [101] emphasize factchecking and sentiment alignment. Asia and the Middle East introduce further challenges, with Arabic [186], Hebrew, Hindi [123], Urdu [8], Chinese [213], and code-mixed Hinglish/Tenglish [62]. Low-resource African languages rely on transfer learning and crowdsourcing [123]. Online platforms compound risks at this stage: Twitter [186], Facebook [104], Instagram, Pinterest [212], YouTube [62], WhatsApp [62], Telegram [174], Skype [51], Al Jazeera, BBC Arabic, CNN Arabic [154], Reddit [30], the dark web [99], Taobao [207], SinaWeibo [203], and LinkedIn all introduce vulnerabilities. Datasets also contribute to privacy risks: BABE [175], PTC [109], CheckThat [58], PROPANEWS [82], AG-News [8], TweetSpin [193], DIPROMATS [116], ArPro [71], ARATWEET [115], Ar-DAD [145], HProp-News [31], Memotion [62], and the Meme Propaganda Techniques Corpus [35]. Competitions such as SemEval [109], WANLP [12], IberLEF [118, 119], NLP4IF [41], and ArAIEval [72, 73] illustrate annotation challenges but often lack transparency and consent. Among 29 privacy risks (|P | = 29), we found 7 concerns related to this specific step (C = 7), yielding a component 𝑟 ident = 0.241.

Data Storage and Security: Once collected, propaganda data presents long-term storage risks. Moral et al. highlight threats from storing diplomatic tweets [118], Modzelewski et al. show re-identification risks [116], and Moreno et al. document metadatarelated exposures [120]. Mittal et al. stress that multimodal storage may disclose cultural and personal traits [115], while Semantha et al. warn of breaches from ignoring Privacy by Design [162]. Auñón et al. note misuse risks when datasets are redistributed without PETs [19]. Regional frameworks shape these risks: GDPR in Europe and North America enforces anonymization, Asia and the Middle East require encryption for sensitive content, while weaker regulations in South America and Africa heighten retention and repurposing risks. Datasets include ARATWEET [115], ArPro [71], FigNews [172], FaceForensics++ [153], and DeeperForensics [85]. Benchmark competitions (WANLP [12], IberLEF [119], ArAIEval [73]) exacerbate issues by redistributing multilingual corpora without consistent anonymization. Multimodal detection magnifies vulnerabilities, as biometric and contextual signals persist long after collection. Out of the 29 identified privacy risks (|P | = 29), only 4 (S = 4) related to this stage have been addressed in existing studies, with the associated risk 𝑟 meta calculated as 0.138.

𝑟 ident 𝑟 meta 0.241 -

𝑟 bias -

𝑟 net -

-

0.138 0.414 -

-

-

-

∥ 𝑅® ∥ 2 0.241 0.138 0.414

0.345 0.345 0.606

Data Analysis and Computational Risk: Analyzing propaganda introduces risks of misclassification, label overprediction, span errors, and synthetic misinformation. PROMPT emphasizes algorithmic bias, adversarial vulnerability, explainability deficits, and misinformation amplification as key threats. Early methods (BoW, TF-IDF, LDA) [41, 102, 120, 145] lacked nuance. Later classifiers (SVM, LR, NB, Decision Trees) [4, 50, 106, 109, 124, 131, 155] and ensembles (RF, GBDT, XGBoost) [43, 154] improved robustness but sacrificed interpretability. Deep learning (LSTM/GRU [17, 93, 131, 184], CNNs [41, 109], Capsule/Multi-Granularity [115, 122]) enhanced performance but added risk. Transformers (BERT, RoBERTa, mBERT, XLM-R) [41, 61, 63, 104], domain-specific variants (BERTweet, DeHateBERT [29, 34]), and generative models (BART, XLNet [82]) advanced contextual analysis but increased bias, hallucination, and synthetic misinformation. Vision-based models (ResNet, VGG, Inception) [63], multimodal architectures (CLIP, VisualBERT, UNITER [35]), and large LLMs (GPT-2/3/4, LLaMA, Mixtral [67, 121, 123, 151]) further magnified privacy concerns due to memorization, regurgitation, and inference leakage. Tools like VADER, SHAP, Botometer [135, 136], and graph/ reinforcement learning methods [187] extended analytical reach but remained limited by bias and efficiency trade-offs. Of the total 29 privacy risks (|P | = 29), literature coverage exists for just 12 (A = 12) at this step, corresponding to a risk component 𝑟 bias of 0.414.

Data Distribution and Network Risk: At distribution, propaganda risks are amplified by network manipulation. Wijenayake et al. show clustering enables exploitation [202], Bessi et al. document coordinated inauthentic behavior [24], and Williamson et al. highlight platform vulnerabilities [203]. Nabhani et al. extend these risks to multilingual contexts [123], while Wu et al. show recommendation-driven amplification [204]. Modzelewski et al. caution about linguistic transfer privacy risks [116], and Krak et al. point to deepfake-driven narrative shaping [95]. Solopova et al. warn of surveillance in distribution channels [173]. Goodman [59] and Hartmann [70] highlight regulatory and historical perspectives on influence manipulation. Together, these findings show that distribution amplifies manipulation and privacy risks. Among the 29 potential privacy risks (|P | = 29), only 10 (D = 10) are considered in current research for this phase, resulting in a risk score 𝑟 net = 0.345. Operationalizing these risk factors (Table 2) shows that cumu® 2 = 0.606, lative risks across C, S, A, D dimensions result in |∥ 𝑅∥ highlighting substantial vulnerabilities. This systematic evaluation

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

Goswami et al.

underscores the need for privacy-preserving ML, adversarial defense, and explainable AI to maintain ethical, secure, and culturally adaptive propaganda detection.

5.3

Propaganda Mitigation

Table 3: Evaluation of Propaganda Mitigation (𝛼 = 𝛽 = 0.5, 𝛾 = 0) Privacy Analysis C S A D

PrivacyGain(𝑆 𝑗 ) 0.421 0.526 0.312 0.263

PerfLoss(𝑆 𝑗 ) 0.189 0.216 0.147 0.164

U(𝑟𝑖 , 𝑠 𝑗 ) 0.116 0.155 0.083 0.049

Various studies propose privacy-preserving mitigation strategies. However, our analysis reveals that these strategies appear in only 5% of the papers on propaganda detection (Appendix Table 11). The evaluation of mitigation across different stages (C, S, A, D) using the formula from Section 4 is shown in Table 11, where the optimization factors 𝛼 = 𝛽 = 0.5 provide equal weight to privacy and performance trade-off and 𝛾 = 0 was kept to show minimal impact of cost in the evaluation. From the results, S yields the highest PrivacyGain (0.526) but also a comparatively higher PerfLoss (0.216), while C strikes a more balanced trade-off between PrivacyGain (0.421) and PerfLoss (0.189). Conversely, D provides the lowest PrivacyGain (0.263) and utility (0.049), highlighting the uneven distribution of benefits across mitigation stages. Semantha et al. [162] highlight the importance of DP, federated learning, and homomorphic encryption to protect user data. Additionally, secure multiparty computation (SMPC) and blockchainbased storage help prevent unauthorized access and ensure data integrity [19]. These methods are crucial in safeguarding data during collaborative propaganda analysis, particularly on social media [24, 197]. Moreover, adversarial machine learning (ML) defenses play a vital role in countering misinformation manipulation [136]. By identifying and neutralizing attacks, these defenses strengthen the resilience of detection systems. Privacy-preserving image processing techniques are also essential, particularly in mitigating facial recognition risks in multimodal propaganda detection [122]. Furthermore, human-in-the-loop models enhance contextual understanding and reduce false positives [110], while Explainable AI (XAI) frameworks improve transparency for regulators and researchers [166, 204]. In addition, graph anomaly detection and blockchain verification provide robust solutions to combat misinformation spread [125, 210]. These methods help identify coordinated misinformation campaigns and ensure data authenticity. Furthermore, privacyenhancing technologies (PETs), such as K-anonymity, L-diversity, and metadata privacy filters, add layers of protection to sensitive user data [116]. These strategies work together to safeguard both privacy and security in propaganda detection systems. Lastly, the integration of these techniques is crucial for ensuring the reliability and privacy of data used in detecting propaganda. By combining various approaches, researchers can enhance the effectiveness of detection systems while protecting against malicious manipulation and preserving user privacy [11, 68].

5.4

Regulatory Compliance

Propaganda detection systems must operate within privacy and accountability regulations, yet only 7% of the surveyed papers explicitly address necessary compliance requirements (Appendix Table 12) [10, 189, 191, 203, 208]. In Europe, regulations such as the GDPR and the Digital Services Act emphasize consent, data minimization, and transparency while establishing rules for harmful content and platform liability [189]. In the U.S., the Foreign Agents Registration Act (FARA) mandates disclosure of certain foreignlinked influence campaigns [208], demonstrating the intersection of propaganda detection with legal oversight. Key challenges include the risk of misclassification: false positives may suppress legitimate speech and raise due-process concerns, whereas false negatives allow harmful content to persist. For instance, bot-detection errors have led to wrongful account suspensions, highlighting the importance of human review and appeal mechanisms [203]. Recent approaches, such as geopolitically informed models [191] and ensemble methods like JUSTDeep [10], further introduce fairness and accountability considerations, underscoring the need for compliance frameworks that document data sources, model features, evaluation procedures, and error handling. Operationalizing the compliance evaluation (Table 4) by using the formula From Section 4, we found a measurable gap between detected safeguards (𝐼 = 0.500) and observed risks (𝜏ˆ = 0.319), yielding a shortfall Δ = −0.125. This indicates that existing protections are insufficient to offset risks such as privacy breaches, misinformation exposure, and targeted profiling. The minimal corrective measure involves introducing additional safeguards, with transparency and reporting identified as the most impactful. Incorporating these measures increases 𝐶𝑜𝑚𝑝𝑆𝑐𝑜𝑟𝑒 (𝑓 ) to 0.500, thereby satisfying the regulatory constraint and demonstrating how measurable improvements can bridge the compliance gap in current propaganda detection research.

5.5

Ethical Considerations

Ethical concerns revolve around data sensitivity, freedom of expression, and the risk of manipulation, yet only 6% of surveyed works explicitly address such issues (Appendix Table 13) [9, 23, 36, 38, 95, 191]. Prior studies highlight fairness, transparency, and user agency as essential safeguards to prevent biases, misinformation, and overpolicing of viewpoints [9, 23, 36, 38, 95], showing the importance of privacy-preserving and accountable AI-driven moderation. The application of the fairness–bias constraint reveals a clear gap between protective measurements and risks in current propaganda detection research. While ethical concerns such as data sensitivity, freedom of expression, and potential misuse are widely acknowledged, only a fraction of studies translate these into quantifiable safeguards. As shown in Table 5, merely 6% of papers address ethical dimensions, focusing primarily on fairness and transparency in cross-lingual augmentation [9], handling of linguistic nuances [95], or rhetorical structures linked to offensive language [36]. By instantiating the formula from Section 4, we obtain 𝐹 = 0.375 for fairness coverage and 𝐵 = 0.500 for risk coverage. The inferred threshold 𝜏ˆ = 0.429 (3/7) indicates that fairness safeguards must reach at least 42.9% of categories to offset the risks observed. Since

SoK: Analysis of Privacy Risks and Mitigation in Online Propaganda Detection through the PROMPT Framework

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

Table 4: Compliance evaluation with an explicit policy threshold 𝜏ˆ𝑐 . 𝐶𝑜𝑚𝑝𝑆𝑐𝑜𝑟𝑒 (𝑓 )

𝜏ˆ𝑐

Δ = 𝐶𝑜𝑚𝑝𝑆𝑐𝑜𝑟𝑒 (𝑓 ) − 𝜏ˆ𝑐

Constraint

Safeguards Detected

Risks Detected

GDPR,

Privacy breaches,

Recommended Fix

CCPA,

Misinformation exposure,

HIPAA

Targeted profiling

Add consent logging, DPIAs, and transparency reporting 0.194

0.319

-0.125

Violated

(expected Δ+0.306) ⇒ 𝐶𝑜𝑚𝑝𝑆𝑐𝑜𝑟𝑒 (𝑓 ) ≈ 0.500 ≥ 𝜏ˆ𝑐

Table 5: Fairness bias constraint evaluation for the thematic corpus. Values use a single consistent definition of 𝜏ˆ = safeguards 3 safeguards+risks = 7 = 0.429. 𝐹 (Fairness)

𝐵 (Risk)

𝜏ˆ

Δ = 𝐹 − 𝜏ˆ

Constraint

Safeguards Detected

Risks Detected

Recommended Fix

Surveillance, Algorithmic Fairness & Bias, 0.375

0.500

0.429

-0.054

Violated

Bias/Discrimination,

Transparency and reporting (+0.125 to 𝐹 )

Data Breach,

⇒ 𝐹 = 0.500 ≥ 𝜏ˆ

Compliance, Informed Consent Misinformation/Manipulation

ˆ the ethical constraint is violated, leaving a measurable short𝐹 < 𝜏, fall of Δ = −0.054. This finding empirically validates the descriptive observation that ethical safeguards remain underdeveloped relative to the risks posed by surveillance, misinformation, and bias. Our analysis shows that the minimal corrective step is the introduction of additional safety. Among the missing categories, transparency/reporting is identified as the most impactful, as it counters misinformation, censorship, accountability deficits. Incorporation of this increases 𝐹 to 0.500, thereby satisfying the fairness constraint ˆ This confirms prior theoretical calls for transparency and (𝐹 ≥ 𝜏). fairness [9, 95] and also demonstrates concretely how a single safeguard can close the measurable ethics gap in current research.

5.6

Table 6: Evaluation of Cross-Domain Propaganda Detection Datasets with Transformer Finetuning Base Model

BERT

Evaluation of Cross-Domain Propaganda Detection Dataset

Notation. Throughout the experiments, 𝑞 denotes the synthetic perturbation rate (fraction of samples with label flips and characterlevel edits). Differential privacy, when discussed conceptually, is parameterized separately by 𝜀 DP and is not interchangeable with 𝑞. We conducted a series of fine-tuning experiments on an encoderonly model (BERT) and a decoder-only model (GPT-2) using the Cross-Domain Propaganda Detection dataset [199] (Train (10755), dev (3585), test (3585) - 60/20/20 split, binary label). This choice provides sufficient coverage of modern text-based deep learning methodologies, as encoders capture contextualized representations optimized for classification tasks [45], while decoders leverage generative and sequence modeling capabilities to support classification through contextual prediction and language modeling [141]. Together, they reflect dominant paradigms in computational approaches to NLP (Table 6). For the encoder, direct fine-tuning of BERT yielded a baseline F1 score of 0.89. Multilingual back translation using Xhosa, Twi, Lao, Pashto, and Yoruba reduced the score slightly to 0.85, indicating some robustness trade-offs. Named Entity Recognition (NER) masking preserved performance, with BERT maintaining its baseline of 0.89. Synthetic perturbation was then introduced by flipping 5% of labels and applying character-level edits to 5% of text (denoted 𝑞 = 0.05); higher perturbation levels (𝑞 ∈ {0.10, 0.15, 0.20}) caused progressively larger F1 degradation.

GPT-2

Procedure Direct Fine-tune Adversarial Defense (Back Translation) NER Masking Perturbation (q = 0.05) Perturbation (q = 0.10) Perturbation (q = 0.15) Perturbation (q = 0.20) BT, NER, Perturbation (q = 0.05) BT, NER, Perturbation (q = 0.10) BT, NER, Perturbation (q = 0.15) BT, NER, Perturbation (q = 0.20) Direct Fine-tune Adversarial Defense (Back Translation) NER Masking Perturbation (q = 0.05) Perturbation (q = 0.10) Perturbation (q = 0.15) Perturbation (q = 0.20) BT, NER, Perturbation (q = 0.05) BT, NER, Perturbation (q = 0.10) BT, NER, Perturbation (q = 0.15) BT, NER, Perturbation (q = 0.20)

Dev F1 0.89 0.84 0.87 0.83 0.77 0.71 0.67 0.74 0.69 0.64 0.62 0.90 0.83 0.87 0.84 0.78 0.73 0.68 0.74 0.71 0.66 0.64

Test F1 0.89 0.85 0.89 0.88 0.86 0.81 0.75 0.83 0.80 0.75 0.64 0.90 0.86 0.90 0.88 0.86 0.81 0.78 0.83 0.80 0.76 0.68

For the decoder, direct fine-tuning of GPT-2 established a baseline F1 of 0.90. Applying multilingual back translation reduced performance slightly to 0.86. NER masking maintained the original score of 0.90, showing negligible impact. With synthetic perturbation at 𝑞 =0.05, GPT-2 reached 0.88, with higher 𝑞 levels again causing progressive decline. Under ensemble fine-tuning with 𝑞 =0.05, GPT-2 achieved an F1 of 0.83; increasing 𝑞 led to sharper accuracy losses. Privacy Gain vs. Utility Trade-off. While the results demonstrate utility degradation as 𝑞 increases, in the context of the threat model, synthetic perturbation can be interpreted as reducing an adversary’s ability to perform concrete attacks such as membership inference, re-identification, and attribute inference by introducing uncertainty in both labels and textual features. For example, increasing 𝑞 would be expected to lower membership inference accuracy by weakening correlations between training instances and

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

Goswami et al.

model predictions, and to hinder re-identification by obfuscating identifiable patterns. A complete evaluation of privacy gain would therefore involve measuring adversarial success rates under these attacks (e.g., attack accuracy or AUC under standard setups such as shadow-model-based membership inference [171]) as a function of q, alongside utility metrics such as F1, thereby characterizing the privacy–utility trade-off.

5.7

Quantifying Adversarial Risk

We instantiate the threat model on representative adversarial actions extracted from the corpus. Each action 𝑎𝑖 is assigned a likelihood 𝑃 (𝑎𝑖 ) representing how frequently it is expected to occur in practice and an impact 𝐼 (𝑎𝑖 ) capturing the severity of the privacy breach if realized. Both values are normalized to the range [0, 1] to ensure comparability across heterogeneous risks. Table 7 summarizes the assigned values and the resulting weighted contributions, with metadata re-identification and model inversion emerging as the most critical threats. The aggregate risk score of 𝑅 = 0.98 indicates that, even under conservative estimates, adversarial actions pose a substantial cumulative risk to propaganda detection pipelines. This quantification provides a concrete basis for prioritizing safeguards in later stages of the PROMPT framework. Table 7: Likelihood and Impact estimates for Representative Adversarial Actions in Propaganda Detection Adversarial Action 𝑎𝑖

𝑃 (𝑎𝑖 )

𝐼 (𝑎𝑖 )

𝑃 (𝑎𝑖 ) · 𝐼 (𝑎𝑖 )

Metadata Re-identification Cross-lingual Inference Leakage Model Inversion Attack Data Poisoning in Training

0.6 0.4 0.3 0.2

0.7 0.5 0.8 0.6

0.42 0.20 0.24 0.12

Aggregate Risk 𝑅

6

0.98

Discussion

In this study, we analyze propaganda detection and mitigation with attention to privacy, security, and ethical shortcomings. Prior work highlights how political propaganda leverages rhetorical tactics such as name-calling and repetition [41], extends to understudied languages and platforms with resources like DIPROMATS and H-Prop-News [116], and raises concerns about surveillance and re-identification [173]. Privacy-preserving methods including differential privacy and federated learning have been advocated [162], alongside compliance with GDPR and ISO 27001 [195] and safeguards for fairness and transparency [9]. By situating our analysis within the PROMPT framework, we address these challenges and demonstrate how privacy and ethics can be embedded in propaganda detection, guiding the evaluation framework used to answer our research questions. Answer to RQ1: Propaganda detection in online platforms poses significant privacy risks due to its reliance on large-scale usergenerated content. These systems collect, store, process vast amounts of data, including text, images, and user interactions, which often contain PII, metadata, and behavioral patterns [124]. As a result, users become vulnerable to tracking, profiling, and re-identification,

even when anonymization techniques are applied. Advanced inference methods can still expose anonymized data, leading to unauthorized surveillance and data exploitation [174]. Furthermore, machine learning models trained on user data are susceptible to model inversion attacks, where adversaries can reconstruct original inputs, potentially revealing sensitive information [116]. Algorithmic profiling presents additional concerns, as it may categorize users based on ideological stances, leading to discrimination, censorship, or unintended biases in content moderation [95]. Inadequate encryption and weak access controls further exacerbate privacy threats, increasing the likelihood of data breaches. Moreover, the diversity of global legal frameworks complicates regulatory compliance, creating inconsistencies in data protection enforcement across jurisdictions [51]. Beyond technical risks, ethical concerns arise, particularly in politically sensitive contexts. Automated propaganda detection mechanisms could be misused to suppress dissenting voices or disproportionately target marginalized groups, undermining freedom of expression and democratic discourse [72]. We have found 0.606 normalized cumulative privacy risk factor across all the steps. Answer to RQ2: Various techniques assess and mitigate privacy risks in propaganda detection; however, each has significant limitations. In particular, privacy risk assessment often involves differential privacy, federated learning, and adversarial testing to evaluate data exposure and model vulnerabilities. For instance, DP measures leakage risk by adding noise to training data, which helps reduce exposure. Nevertheless, this approach degrades model accuracy, making it a less optimal solution in certain contexts [56]. Similarly, federated learning minimizes centralized data collection by decentralizing training, thereby reducing direct access to sensitive data. Yet, it remains vulnerable to gradient inversion attacks, which can reconstruct sensitive information from model updates [118].In addition to these methods, traditional anonymization techniques, such as entity removal and content masking, assess privacy by identifying sensitive text. However, these techniques often fail to prevent re-identification, especially when attackers utilize auxiliary data. Furthermore, they tend to reduce detection accuracy, thereby limiting their overall effectiveness [128]. Likewise, adversarial testing is used to evaluate how models handle manipulated inputs; yet, it struggles to keep up with evolving attack strategies, making it an imperfect solution [154]. Beyond text-based methods, network-based assessments are employed to detect manipulation patterns, such as Sybil and centrality attacks. While these techniques enhance detection capabilities, metadata exposure remains a significant concern [11, 95]. Additionally, multimodal detection, which integrates text, audio, and visual data, improves accuracy; however, it also introduces privacy risks through metadata cross-referencing, potentially exposing user-sensitive information [39, 204]. Finally, large language models (LLMs) pose further challenges in assessing privacy risks due to their black-box nature, making their security vulnerabilities harder to evaluate. Although these techniques help mitigate privacy risks to some extent, they require trade-offs between privacy protection and propaganda detection effectiveness. We observe a privacy–utility trade-off (0.049–0.155) under synthetic perturbation (𝑞), with up to a 7% performance drop in privacy-preserved transformer finetuning.

SoK: Analysis of Privacy Risks and Mitigation in Online Propaganda Detection through the PROMPT Framework

Answer to RQ3: Privacy-preserving strategies play a crucial role in enhancing the security and integrity of propaganda detection models. One effective approach is applying differential privacy at the feature extraction stage, which perturbs high-risk features while maintaining detection performance [212]. Additionally, secure federated learning, when combined with SMPC and homomorphic encryption, ensures that model updates remain protected from inference attacks [98]. Moreover, privacy-preserving techniques such as tokenization-based de-identification, adversarial text sanitization help mask sensitive details while preserving linguistic coherence, making them valuable for mitigating privacy risks [123]. To further enhance transparency and accountability, XAI methods, including SHAP and LIME, facilitate risk assessment and bias detection, thus improving trust in the detection models [54]. At the same time, strengthening adversarial robustness through gradient masking and perturbation defenses fortifies models against manipulation attempts, ensuring their reliability in adversarial settings [41]. Furthermore, automated privacy audits play a key role in complying with regulations such as GDPR [57] and CCPA [98], helping organizations align with legal requirements. In politically sensitive contexts, integrating Human-in-the-Loop (HITL) mechanisms offers an additional layer of oversight, reducing risks associated with automated decision-making [212]. Finally, the PROMPT framework serves as a robust solution for online propaganda analysis. Through its six-step propaganda analysis, PROMPT provides a comprehensive approach, reinforcing privacy-preserving measures and improving security and ethical integrity. By PROMPT analysis, we have found a lack of 0.306 in compliance score and 0.125 in fairness threshold which states that existing regulatory compliance and ethical aspect should be more agile for privacy and security preservation in this domain.

7

Implications

Grounding our synthesis of 162 studies in the PROMPT framework, we identify system-level implications for deploying privacy-aware and regulation-aligned propaganda detection. Current pipelines ex® 2 = 0.606 (Tables 4, 5), limited hibit a normalized privacy risk of ∥ 𝑅∥ adoption of privacy-enhancing technologies, and measurable compliance and fairness deficits. These observations demonstrate that existing systems inherit structural weaknesses across all pipeline stages rather than isolated failures at individual modules. As a result, end-to-end redesign is necessary to align detection practices with emerging regulatory, ethical, and technical expectations.

7.1

Data Collection and Preprocessing (C)

Data acquisition continues to drive the highest concentration of privacy exposure, particularly in multilingual and multimodal environments where inference pathways expand rapidly. Re-identification, covert profiling, and consent violations occur when pipelines aggregate sensitive linguistic cues or cross-reference behavioral metadata. Automated entity redaction, provenance tracking, data minimization, and dataset nutrition labels reduce downstream exposure and establish transparent, auditable data lineage [40]. Federated curation and differential privacy adapters replace centralized retention in high-sensitivity contexts, thereby reducing attack surface and limiting the blast radius of potential breaches [48]. Large language

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

model-based annotation workflows require formal leakage analysis because prompts, intermediate reasoning traces, and labels frequently encode identifiable information [74]. Treating annotation as a controlled computational activity rather than an informal step promotes reproducibility, meets regulatory expectations for traceability, and reduces the likelihood of unintentional data leaks.

7.2

Storage and Security (S)

Metadata correlation, multimodal linkability, and unrestricted redistribution increase long-term privacy risk even when primary content appears sanitized. Encryption at rest, granular access control, and retention limits form a defensible baseline, while metadata minimization and controlled hashing restrict inference channels that enable re-identification [201]. Machine-readable usage policies, audit logs, and Data Protection Impact Assessments elevate pipeline transparency and equip external auditors with verifiable evidence of compliance. Integrating compliance scoring into development and release workflows, with automatic blocking when scores fall below the threshold 𝜏ˆ𝑐 , enforces regulatory alignment and prevents the deployment of systems that violate GDPR, CCPA, or the EU AI Act. These mechanisms transform compliance from a post hoc audit activity into a continuous, measurable design constraint.

7.3

Analysis and Modeling (A)

Model development remains disproportionately centered on accuracy, leaving privacy, robustness, and fairness underexamined. This imbalance produces pipelines that perform well under benchmark conditions but introduce substantial risk during real-world deployment. Reporting protocols that pair accuracy with robustness under perturbation, cross-group fairness gaps, and explicit privacy–utility trade-offs provide a multidimensional view of system behavior and reveal risks that single-metric evaluation conceals [170]. Privacypreserving training techniques such as DP-SGD, federated optimization, encrypted inference, and privacy-preserving multimodal joins enhance protection in high-risk domains [140]. Explainability techniques supply structured evidence for debugging, model validation, and regulatory assessment, enabling investigators to trace harmful outputs and identify structural biases. Mapping each component of 𝑅® to a mitigation strategy in 𝑀 : 𝑅 → 𝑆 embeds principled defense selection into the development cycle, replacing ad hoc adjustments with systematic risk reasoning.

7.4

Distribution and Networks (D)

Propagation networks amplify both privacy risk and societal harm, especially when platforms rely on graph-based features to detect coordinated behavior. These features expose relational patterns that enable adversarial inference and user correlation. Noise injection into graph embeddings, subgraph sampling, and private anomaly detection maintain sensitivity to coordinated manipulation while reducing the exposure of relational structure [152]. Transparency reporting, user-notification channels, and structured appeal pathways strengthen procedural accountability and reduce false-positive harms during moderation. Ensuring that deployment occurs only when fairness and compliance conditions satisfy 𝐹 ≥ 𝜏ˆ

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

prevents the release of models that disproportionately affect vulnerable communities or fail to satisfy jurisdictional regulatory thresholds. These controls operationalize fairness and compliance as measurable gates rather than aspirational ideals.

Goswami et al.

these dynamics highlights the need for governance frameworks that jointly address privacy, accountability, and the geopolitical implications of controlling information flows.

8 7.5

Operational Usage of PROMPT

The PROMPT framework functions as an operational playbook that unifies risk quantification, mitigation selection, and compliance evaluation into a single iterative cycle. It establishes a structured procedure in which practitioners identify dominant risks, map them to feasible defenses, and refine the pipeline until residual risk falls within organizational tolerance. This cycle parallels mature security engineering workflows by linking qualitative threat modeling with quantitative measures that guide decision-making at each stage of system development. Operational use of PROMPT begins with estimating the risk vector 𝑅® and identifying components that exceed acceptable thresholds. The mapping 𝑀 : 𝑅 → 𝑆 then selects mitigation strategies that reduce risk while respecting performance and cost constraints. As updates are incorporated, practitioners ® 2, CompScore, 𝐹 ) across successive releases to evaluate track (∥𝑅∥ how technical changes influence privacy exposure, regulatory alignment, and fairness guarantees. Documenting marginal utility effects 𝑈 (𝑟𝑖 , 𝑠 𝑗 ) provides an auditable record that links each intervention to measurable reductions in risk, creating transparent justification for design choices [164]. This structure elevates privacy, fairness, and compliance from aspirational goals to operational metrics embedded directly into development and governance pipelines. Rather than applying controls after deployment, teams integrate them throughout the lifecycle, from dataset design to distribution. The result is a reproducible process that exposes trade-offs early, supports independent verification, and strengthens organizational accountability. By enforcing continuous monitoring and transparent reporting, the PROMPT playbook enables propaganda detection systems to function as measurable, trustworthy, and regulation-aligned socio-technical infrastructures that maintain resilience under evolving adversarial and policy pressures [16].

7.6

State Power and Information Asymmetry

Beyond system-level considerations, propaganda detection operates within broader geopolitical and institutional power asymmetries. State actors and large platforms possess disproportionate access to data, computational resources, and regulatory influence, enabling large-scale monitoring, narrative shaping, and cross-border information control. While privacy-preserving mechanisms in PROMPT reduce risks such as re-identification, membership inference, and user profiling, they may also limit transparency and accountability when applied in high-stakes moderation or governance contexts. This creates an inherent trade-off between protecting individual privacy and enabling oversight of powerful actors engaged in coordinated information operations. In this setting, the deployment of privacy-aware propaganda detection must be carefully balanced to avoid reinforcing existing asymmetries. Mechanisms such as auditable pipelines, transparency reporting, and independent evaluation become critical to ensure that privacy protections do not obscure harmful large-scale manipulation. Situating PROMPT within

Limitations and Future Work

In this study, we surveyed a broad spectrum of privacy-preserving techniques and highlighted their role in propaganda detection, but we did not evaluate their effectiveness in operational deployments. We focused our corpus on top-tier conferences and journals, which allowed us to capture state-of-the-art research while inevitably leaving out some relevant work from other venues and industry reports. In future work, we will expand our dataset to cover non-English and under-represented sources, with particular attention to lowresource contexts. We also plan to conduct empirical experiments that measure the integration costs and performance trade-offs of privacy-enhancing technologies when deployed at scale.

9

Conclusion

Online propaganda poses risks that demand detection pipelines which are accurate, privacy-preserving, and regulation-compliant, yet advances in ML and DL, while improving detection accuracy, introduce new vulnerabilities such as metadata exposure, adversarial inference, compliance gaps. In this work, we present the first systematic evaluation of privacy, fairness, and compliance in propaganda detection, analyzing 162 peer-reviewed publications through the PROMPT framework, which we developed. Our analysis showed that 81% of systems rely on sensitive user data while only 18% adopt privacy-preserving technologies, yielding a nor® 2 = 0.606. We further demonstrated that malized corpus risk of ∥ 𝑅∥ fairness coverage consistently falls short of the required threshold and that compliance deficits of 0.306 leave current pipelines ethically and legally fragile. Empirical experiments established concrete baselines for the privacy–utility trade-off, with privacy constraints causing up to 22% performance degradation. Finally, our thematic mapping exposed uneven attention to propaganda techniques and a lack of controls and measurements in multilingual and multimodal settings.By operationalizing risk vectors, compliance scores, and fairness thresholds, we turned ethical and legal considerations into quantifiable requirements, establishing a rigorous basis for evaluating and improving propaganda detection systems.

Ethical Considerations This SoK use only publicly available datasets and examples, without collecting new social media data, interacting with users, or processing personally identifiable information. Given the sensitivity of propaganda tied to political, cultural, religious, and demographic identities, we avoid deanonymization, real-world adversarial attacks, and any design of manipulation or disinformation systems; all attack vectors (e.g., re-identification, metadata leakage, adversarial misinformation) are discussed conceptually.

Acknowledgments We would like to acknowledge the Data Agency and Security (DAS) Lab at George Mason University. This work is partially funded by Google Research Awards. The opinions expressed in this work are solely those of the authors.

SoK: Analysis of Privacy Risks and Mitigation in Online Propaganda Detection through the PROMPT Framework

References [1] Reem Abdel-Salam. 2023. rematchka at ArAIEval Shared Task: Prefix-Tuning & Prompt-tuning for Improved Detection of Propaganda and Disinformation in Arabic Social Media Content. In Proceedings of ArabicNLP. Association for Computational Linguistics, Singapore (Hybrid), 536–542. [2] Vlad Achimescu and Dan Sultanescu. 2020. Feeding the troll detection algorithm: Informal flags used as labels in classification models to identify perceived computational propaganda. First Monday (2020). [3] Andrick Adhikari, Sanchari Das, and Rinku Dewri. 2023. Evolution of Composition, Readability, and Structure of Privacy Policies over Two Decades. Proceedings of PETS 3 (2023). [4] Kartik Aggarwal and Anubhav Sadana. 2019. NSIT@ NLP4IF-2019: Propaganda detection from news articles using transfer learning. In Proceedings of NLP4IF. [5] Pir Noman Ahmad, Jiequn Guo, Nagwa M AboElenein, Qazi Mazhar ul Haq, Sadique Ahmad, Abeer D Algarni, and Abdelhamied A. Ateya. 2025. Hierarchical graph-based integration network for propaganda detection in textual news articles on social media. Scientific Reports 15 (2025). [6] Pir Noman Ahmad and Khalid Khan. 2023. Propaganda Detection And Challenges Managing Smart Cities Information On Social Media. EAI Endorsed Transactions on Smart Cities 7 (2023). [7] Pir Noman Ahmad, Adnan Muhammad Shah, and KangYoon Lee. 2023. Propaganda Detection in Public Covid-19 Discussion on Social Media. (2023). [8] Pir Noman Ahmad, Liu Yuanchao, Khursheed Aurangzeb, Muhammad Shahid Anwar, and Qazi Mazhar ul Haq. 2024. Semantic web-based propaganda text detection from social media using meta-learning. Service Oriented Computing and Applications (2024). [9] Vicent Ahuir, Lluís-Felip Hurtado, Fernando García-Granada, and Emilio Sanchis. 2023. ELiRF-VRAIN at DIPROMATS 2023: Cross-lingual Data Augmentation for Propaganda Detection.. In Proceedings of SEPLN. [10] Hani Al-Omari, Malak Abdullah, Ola AlTiti, and Samira Shaikh. 2019. JUSTDeep at NLP4IF 2019 task 1: Propaganda detection using ensemble deep learning models. In Proceedings of NLP4LF. 113–118. [11] Muhammad Al-Qurishi, Majed Alrubaian, Sk Md Mizanur Rahman, Atif Alamri, and Mohammad Mehedi Hassan. 2018. A prediction system of Sybil attack in social network using deep-regression model. Future Generation Computer Systems 87 (2018). [12] Firoj Alam, Hamdy Mubarak, Wajdi Zaghouani, Giovanni Da San Martino, and Preslav Nakov. 2022. Overview of the WANLP 2022 shared task on propaganda detection in Arabic. Proceedings of WANLP (2022). [13] Farizeh Aldabbas, Shaina Ashraf, Rafet Sifa, and Lucie Flek. 2025. MultiProp Framework: Ensemble Models for Enhanced Cross-Lingual Propaganda Detection in Social Media and News using Data Augmentation, Text Segmentation, and Meta-Learning. In Proceedings of AbjadNLP. [14] Yasser Alhabashi, Abdullah Alharbi, Samar Ahmad, Serry Sibaee, Omer Nacar, Lahouari Ghouti, and Anis Koubaa. 2024. ASOS at ArAIEval Shared Task: Integrating Text and Image Embeddings for Multimodal Propaganda Detection in Arabic Memes. In Proceedings of ArabicNLP. [15] Tariq Alhindi, Jonas Pfeiffer, and Smaranda Muresan. 2019. Fine-Tuned Neural Models for Propaganda Detection at the Sentence and Fragment levels. Proceedings of EMNLP-IJCNLP (2019). [16] Majed Alshammari and Andrew Simpson. 2018. A model-based approach to support privacy compliance. Information & Computer Security 26, 4 (2018), 437–453. [17] Anastasios Arsenos and Georgios Siolas. 2020. NTUAAILS at SemEval-2020 Task 11: Propaganda detection and classification with biLSTMs and ELMo. In Proceedings of SemEval. [18] Joseph Attieh and Fadi Hassan. 2022. Pythoneers at WANLP 2022 Shared Task: Monolingual AraBERT for Arabic Propaganda Detection and Span Extraction. In Proceedings of WANLP. [19] JM Auñón, D Hurtado-Ramírez, L Porras-Díaz, B Irigoyen-Peña, S Rahmian, Yusra Al-Khazraji, J Soler-Garrido, and Alexander Kotsev. 2024. Evaluation and utilisation of privacy enhancing technologies—A data spaces perspective. Data in Brief 55 (2024). [20] Eugene Bagdasaryan and Vitaly Shmatikov. 2022. Spinning language models: Risks of propaganda-as-a-service and countermeasures. In 2022 IEEE Symposium on Security and Privacy (SP). IEEE, 769–786. [21] Vít Baisa, Ondřej Herman, and Aleš Horák. 2019. Benchmark dataset for propaganda detection in Czech newspaper texts. In Proceedings of RANLP. 77–83. [22] Arash Barfar. 2022. A linguistic/game-theoretic approach to detection/explanation of propaganda. Expert Systems with Applications 189 (2022). [23] Luka Barišic, Marko Pisacic, and Filip Radovic. [n. d.]. To Context or Not to Context? Analysis of De-contextualized Word Embeddings in Propaganda Detection Task. ([n. d.]). [24] Alessandro Bessi and Emilio Ferrara. 2016. Social bots distort the 2016 US Presidential election online discussion. First monday 21 (2016). [25] Corneliu Bjola. 2018. The Ethics of Countering Digital Propaganda. Ethics & International Affairs 32 (2018).

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

[26] Verena Blaschke, Maxim Korniyenko, and Sam Tureski. 2020. CyberWallE at SemEval-2020 Task 11: An analysis of feature engineering for ensemble models for propaganda detection. Proceedings of SemEval (2020). [27] Marco Casavantes, Manuel Montes-y Gómez, Delia Irazú Hernández Farías, Luis Carlos González-Gurrola, and Alberto Barrón-Cedeño. 2023. PropaLTL at DIPROMATS: Incorporating Contextual Features with BERT’s Auxiliary Input for Propaganda Detection on Tweets.. In Proceedings of SEPLN. [28] Marco Casavantes, Manuel Montes-y Gómez, Luis Carlos González, and Alberto Barróñ-Cedeno. 2023. Propitter: A Twitter Corpus for Computational Propaganda Detection. In Proceedings of MICAI. [29] Marco Casavantes, Manuel Montes-y Gómez, Delia Irazú Hernández-Farías, Luis Carlos González, and Alberto Barrón-Cedeño. 2024. PropaLTL at DIPROMATS 2024: Cross-lingual Data Augmentation for Propaganda Detection on Tweets. In Proceedings of SEPLN. [30] Danilo Cavaliere, Mariacristina Gallo, and Claudio Stanzione. 2023. Propaganda Detection Robustness Through Adversarial Attacks Driven by eXplainable AI. In Proceedings of XAI. [31] Deptii Chaudhari and Ambika Vishal Pawar. 2023. Empowering propaganda detection in resource-restraint languages: a transformer-based framework for classifying hindi news articles. Big Data and Cognitive Computing 7 (2023). [32] Aniruddha Chauhan and Harshita Diddee. 2020. PsuedoProp at SemEval-2020 Task 11: Propaganda span detection using BERT-CRF and ensemble sentence level classifier. In Proceedings of SemEval. [33] Tanmay Chavan and Aditya Manish Kane. 2022. ChavanKane at WANLP 2022 Shared Task: Large Language Models for Multi-label Propaganda Detection. In Proceedings of WANLP. [34] Tanmay Chavan and Aditya Manish Kane. 2022. ChavanKane at WANLP 2022 Shared Task: Large Language Models for Multi-label Propaganda Detection. In Proceedings of WANLP. [35] Pengyuan Chen, Lei Zhao, Yangheran Piao, Hongwei Ding, and Xiaohui Cui. 2024. Multimodal visual-textual object graph attention network for propaganda detection in memes. Multimedia Tools and Applications 83 (2024). [36] Alexander Chernyavskiy, Dmitry Ilvovsky, and Preslav Nakov. 2024. Unleashing the Power of Discourse-Enhanced Transformers for Propaganda Detection. In Proceedings of EACL. [37] Evan Crothers, Nathalie Japkowicz, and Herna L Viktor. 2019. Towards ethical content-based detection of online influence campaigns. In Proceedings of MLSP. [38] Jose Cuadrado, Elizabeth Martinez, Juan Cuadrado, Juan Carlos Martinez-Santos, and Edwin Puertas. 2024. VerbaNex AI at DIPROMATS 2024: Enhancing Propaganda Detection in Diplomatic Tweets with Fine-Tuned BERT and Integrated NLP Techniques. (2024). [39] Jian Cui, Lin Li, Xin Zhang, and Jingling Yuan. 2023. Multimodal Propaganda Detection Via Anti-Persuasion Prompt enhanced contrastive learning. In Proceedings of ICASSP. [40] Rachel Cummings and David Durfee. 2020. Individual sensitivity preprocessing for data privacy. In Proceedings of the Fourteenth Annual ACM-SIAM Symposium on Discrete Algorithms. SIAM, 528–547. [41] Giovanni Da San Martino, Alberto Barron-Cedeno, and Preslav Nakov. 2019. Findings of the NLP4IF-2019 shared task on fine-grained propaganda detection. In Proceedings of NLP4IF. [42] Giovanni Da San Martino, Alberto Barrón-Cedeno, and Preslav Nakov. 2020. Evaluation of propaganda detection tasks. Proceedings of SemEval (2020). [43] Jiaxu Dao, Jin Wang, and Xuejie Zhang. 2020. YNU-HPCC at SemEval-2020 task 11: LSTM network for detection of propaganda techniques in news articles. In Proceedings of SemEval. [44] Daryna Dementieva, Igor Markov, and Alexander Panchenko. 2020. SkoltechNLP at SemEval-2020 Task 11: Exploring unsupervised text augmentation for propaganda detection. In Proceedings of SemEval. [45] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL. [46] Dimas Sony Dewantara and Indra Budi. 2020. Combination of lstm and cnn for article-level propaganda detection in news articles. In Proceedings of ICIC. [47] Soumia Zohra El Mestari, Gabriele Lenzini, and Huseyin Demirci. 2024. Preserving data privacy in machine learning systems. Computers & Security 137 (2024), 103605. [48] Ahmed El Ouadrhiri and Ahmed Abdelhadi. 2022. Differential privacy for deep and federated learning: A survey. IEEE access 10 (2022), 22359–22380. [49] Jan Ellermann. 2016. Terror won’t kill the privacy star–tackling terrorism propaganda online in a data protection compliant manner. In ERA Forum, Vol. 17. Springer, 555–582. [50] Vlad Ermurachi and Daniela Gifu. 2020. UAIC1860 at SemEval-2020 Task 11: Detection of propaganda techniques in news articles. In Proceedings of SemEval. [51] Alessandra Fabrocini. 2021. Electoral propaganda and privacy: the italian Data Protection Authority lays down rules for the advertising campaign. European Journal of Privacy Law & Technologies (2021). [52] Ali Fadel, Ibraheem Tuffaha, and Mahmoud Al-Ayyoub. 2019. Pretrained ensemble learning for fine-grained propaganda detection. In Proceedings of NLP4IF.

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

139–142. [53] Miguel Fernández, Maximiliano Ojeda, Lilly Guevara, Diego Varela, Marcelo Mendoza, and Alberto Barrón-Cedeño. 2024. VICTOR VECTORS@ DIPROMATS 2024: Propaganda Detection with LLM Paraphrasing and Machine Translation. (2024). [54] Mary Fouad and Julie Weeds. 2024. SussexAI at ArAIEval Shared Task: mitigating class imbalance in arabic propaganda detection. In Proceedings of ArabicNLP. [55] Paul Franklin, Donald Cooper, Jan Danel, and Tiger Hu. 2020. Russian Facebook Propaganda Detection with Classification Models. (2020). [56] Kamel Gaanoun and Imade Benelallam. 2022. SI2M & AIOX Labs at WANLP 2022 Shared Task: Propaganda Detection in Arabic, A Data Augmentation and Name Entity Recognition Approach. In Proceedings of WANLP. [57] Eric Goldman. 2020. An introduction to the california consumer privacy act (ccpa). Santa Clara Univ. Legal Studies Research Paper (2020). [58] Paweł Golik, Arkadiusz Modzelewski, and Aleksander Jochym. 2024. DSHacker at CheckThat! 2024: LLMs and BERT for check-worthy claims detection with propaganda co-occurrence analysis. Faggioli et al.[22] (2024). [59] Ellen P Goodman and Lyndsey Wajert. 2017. The Honest Ads Act Won’t End Social Media Disinformation, but It’s a Start. Available at SSRN 3064451 (2017). [60] Dhiman Goswami, Jai Kruthunz Naveen Kumar, and Sanchari Das. 2026. NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey. In Findings of the Association for Computational Linguistics: EACL 2026. 1519–1541. [61] Dmitry Grigorev and Vladimir Ivanov. 2020. Inno at SemEval-2020 Task 11: Leveraging Pure Transformer for Multi-Class Propaganda Detection. Proceedings of SemEval (2020). [62] SUNIL GUNDAPU. 2022. Automatic Detection of Negativity in User-Generated Social Media Content: Building Models for Fake News, Sentiment Analysis, Propaganda Techniques, Offensive and Sarcastic Content Identification. Ph. D. Dissertation. International Institute of Information Technology Hyderabad. [63] Sunil Gundapu and Radhika Mamidi. 2022. Detection of propaganda techniques in visuo-lingual metaphor in memes. arXiv preprint arXiv:2205.02937 (2022). [64] Chen Guo, Nan Zheng, and Chengqi Guo. 2023. Seeing is not believing: a nuanced view of misinformation warning efficacy on video-sharing social media platforms. Proceedings of CHI 7, CSCW2 (2023), 1–35. [65] Pankaj Gupta, Khushbu Saxena, Usama Yaseen, Thomas Runkler, and Hinrich Schütze. 2019. Neural Architectures for Fine-Grained Propaganda Detection in News. Proceedings of EMNLP-IJCNLP (2019). [66] Kyle Hamilton. 2021. Towards an ontology for propaganda detection in news articles. In Proceedings of ESWC. [67] Kyle Hamilton, Luca Longo, and Bojan Bozic. 2024. GPT Assisted Annotation of Rhetorical and Linguistic Features for Interpretable Propaganda Technique Detection in News Text.. In Proceedings of WWW. [68] Sakshini Hangloo and Bhavna Arora. 2022. Combating multimodal fake news on social media: methods, datasets, and future perspective. Multimedia systems 28 (2022). [69] Sheetal Harris, Hassan Jalil Hadi, Naveed Ahmad, and Mohammed Ali Alshara. 2024. Fake News Detection Revisited: An Extensive Review of Theoretical Frameworks, Dataset Assessments, Model Constraints, and Forward-Looking Research Agendas. Technologies 12 (2024). [70] Mareike Hartmann, Yevgeniy Golovchenko, and Isabelle Augenstein. 2019. Mapping (Dis-)Information Flow about the MH17 Plane Crash. In Proceedings of NLP4IF. [71] Maram Hasanain, Fatema Ahmad, and Firoj Alam. 2024. Can GPT-4 Identify Propaganda? Annotation and Detection of Propaganda Spans in News Articles. In Proceedings of LREC-COLING. [72] Maram Hasanain, Firoj Alam, Hamdy Mubarak, Samir Abdaljalil, Wajdi Zaghouani, Preslav Nakov, Giovanni Da San Martino, and Abed Freihat. 2023. ArAIEval Shared Task: Persuasion Techniques and Disinformation Detection in Arabic Text. In Proceedings of ArabicNLP. [73] Maram Hasanain, Md. Arid Hasan, Fatema Ahmad, Reem Suwaileh, Md. Rafiul Biswas, Wajdi Zaghouani, and Firoj Alam. 2024. ArAIEval Shared Task: Propagandistic Techniques Detection in Unimodal and Multimodal Arabic Content. In Proceedings of ArabicNLP. [74] Feng He, Tianqing Zhu, Dayong Ye, Bo Liu, Wanlei Zhou, and Philip Yu. 2024. The emerged security and privacy of llm agent: A survey with case studies. Comput. Surveys (2024). [75] Vitalij Hein. 2023. Propaganda detection in Russian and American news coverage about the war in Ukraine through text classification. Ph. D. Dissertation. Technische Universität Wien. [76] Ondřej Herman, Vít Baisa, and Aleš Horák. 2020. Propaganda detection tool. (2020). [77] Aleš Horák, Vít Baisa, and Ondřej Herman. 2021. Technological approaches to detecting online disinformation and manipulation. Challenging online propaganda and disinformation in the 21st century (2021), 139–166. [78] Benjamin D Horne, Dorit Nevo, and Susan L Smith. 2023. Ethical and safety considerations in automated fake news detection. Behaviour & Information Technology (2023).

Goswami et al.

[79] Wenjun Hou and Ying Chen. 2019. CAUnLP at NLP4IF 2019 shared task: Contextdependent BERT for sentence-level propaganda detection. In Proceedings of NLP4IF. 83–86. [80] Xiaolong Hou, Junsong Ren, Gang Rao, Lianxin Lian, Zhihao Ruan, Yang Mo, and Jianping Shen. 2021. FPAI at SemEval-2021 task 6: BERT-MRC for propaganda techniques detection. In Proceedings of SemEval. [81] Austin Hounsel, Jordan Holland, Ben Kaiser, Kevin Borgolte, Nick Feamster, and Jonathan Mayer. 2020. Identifying disinformation websites using infrastructure features. In Proceedings of FOCI. [82] Kung-Hsiang Huang, Kathleen Mckeown, Preslav Nakov, Yejin Choi, and Heng Ji. 2023. Faking Fake News for Real Fake News Detection: Propaganda-Loaded Training Data Generation. In Proceedings of ACL. [83] Ahmed Samir Hussein, Abu Bakr Soliman Mohammad, Mohamed Ibrahim, Laila Hesham Afify, and Samhaa R El-Beltagy. 2022. NGU CNLP atWANLP 2022 Shared Task: Propaganda Detection in Arabic. In Proceedings of WANLP. [84] Tawfik Jelassi. 2023. Towards an Internet of Trust—UNESCO’s Guidelines for the Governance of Digital Platforms. In Proceedings of LMDE. [85] Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy. 2020. Deeperforensics-1.0: A large-scale dataset for real-world face forgery detection. In Proceedings of CVPR. [86] Yunzhe Jiang, Cristina Gârbacea, and Qiaozhu Mei. 2020. UMSIForeseer at SemEval-2020 Task 11: Propaganda detection by fine-tuning BERT with resampling and ensemble learning. In Proceedings of SemEval. [87] Konrad Kaczyński and Piotr Przybyła. 2021. HOMADOS at SemEval-2021 Task 6: Multi-task learning for propaganda detection. In Proceedings of SemEval. [88] Siddharth Kelkar, Srinivasa Ravi, Sheela Ramanna, and Anand Kumar Madasamy. 2024. Multimodal Propaganda Detection in Memes with Tolerance-Based Soft Computing Method. In Proceedings of IJCRS. [89] Ansgar Kellner, Lisa Rangosch, Christian Wressnegger, and Konrad Rieck. 2019. Political elections under (social) fire? Analysis and detection of propaganda on twitter. arXiv preprint arXiv:1912.04143 (2019). [90] Ansgar Kellner, Christian Wressnegger, and Konrad Rieck. 2020. What’s all that noise: analysis and detection of propaganda on Twitter. In Proceedings of EuroSec. [91] Supriya Khadka and Sanchari Das. 2026. SoK: Understanding the Pedagogical, Health, Ethical, and Privacy Challenges of Extended Reality in Early Childhood Education. arXiv preprint arXiv:2602.12749 (2026). [92] Moonsung Kim and Steven Bethard. 2020. TTUI at SemEval-2020 Task 11: Propaganda detection with transfer learning and ensembles. In Proceedings of SemEval. [93] Iurii Krak, Volodymyr Didur, Maryna Molchanova, Olexander Mazurets, O Zalutska, E Manziuk, and Olexander Barmak. 2024. Method for Political Propaganda Detection in Internet Content Using Recurrent Neural Network Models Ensemble. In Proceedings of CEUR. [94] Iurii Krak, Maryna Molchanova, Volodymyr Didur, Olena Sobko, Olexander Mazurets, and Olexander Barmak. 2025. Method of semantic features estimation for political propaganda techniques detection using transformer neural networks. In Proceedings of CEUR. [95] Iu V Krak, VO Didur, MO Molchanova, OV Mazurets, OV Sobko, OO Zalutska, and OV Barmak. 2024. Method for political propaganda detection in internet content using neural network natural language processing tools. PROBLEMS IN PROGRAMMING (2024). [96] Michael Kranzlein, Shabnam Behzad, and Nazli Goharian. 2020. Team DoNotDistribute at SemEval-2020 Task 11: Features, finetuning, and data augmentation in neural models for propaganda detection in news articles. Proceedings of SemEval (2020). [97] Jai Kruthunz Naveen Kumar, Aishwarya Umeshkumar Surani, Harkirat Singh, and Sanchari Das. 2025. Privacy Discourse and Emotional Dynamics in Mental Health Information Interaction on Reddit. arXiv preprint arXiv:2512.15945 (2025). [98] Sahinur Rahman Laskar, Rahul Singh, Abdullah Faiz Ur Rahman Khilji, Riyanka Manna, Partha Pakray, and Sivaji Bandyopadhyay. 2022. CNLP-NITS-PP at WANLP 2022 Shared Task: Propaganda Detection in Arabic using Data Augmentation and AraBERT Pre-trained Model. In Proceedings of WANLP. [99] Mark Last. 2023. Online Propaganda Detection. In Machine Learning for Data Science Handbook: Data Mining and Knowledge Discovery Handbook. [100] João A Leite, Olesya Razuvayevskaya, Kalina Bontcheva, and Carolina Scarton. 2024. EUvsDisinfo: A Dataset for Multilingual Detection of Pro-Kremlin Disinformation in News Articles. In Proceedings of CIKM. [101] Mikhail Lepekhin and Serge Sharoff. 2023. Ftd at semeval-2023 task 3: News genre and propaganda detection by comparing mono-and multilingual models with fine-tuning on additional data. In Proceedings of SemEval. [102] Jinfen Li, Zhihao Ye, and Lu Xiao. 2019. Detection of propaganda using logistic regression. In Proceedings of NLP4IF. [103] Li Li, Yuxi Fan, Mike Tse, and Kuo-Yi Lin. 2020. A review of applications in federated learning. Computers & Industrial Engineering 149 (2020). [104] Peiguang Li, Xuan Li, and Xian Sun. 2021. 1213Li at SemEval-2021 task 6: detection of propaganda with multi-modal attention and pre-trained models. In Proceedings of SemEval-2021.

SoK: Analysis of Privacy Risks and Mitigation in Online Propaganda Detection through the PROMPT Framework

[105] Oleksandr Lytvyn. 2024. Enhancing Propaganda Detection with Open Source Language Models: A Comparative Study. In Proceedings of MEi: CogSci. [106] Tarek Mahmoud and Preslav Nakov. 2024. Bertastic at semeval-2024 task 4: Stateof-the-art multilingual propaganda detection in memes via zero-shot learning with vision-language models. In Proceedings of SemEval. [107] Samia Manzoor, Aasima Safdar, and Beenish Zaheen. 2019. Propaganda Revisited: Understanding Propaganda in the Contemporary Communication Oriented World. Global Regional Review 4, 3 (2019), 317–324. [108] Norman Mapes, Anna White, Radhika Medury, and Sumeet Dua. 2019. Divisive language and propaganda detection using multi-head attention transformers with deep learning BERT-based language models for binary classification. In Proceedings of NLP4IF. [109] G Martino, Alberto Barrón-Cedeno, Henning Wachsmuth, Rostislav Petrov, and Preslav Nakov. 2020. SemEval-2020 task 11: Detection of propaganda techniques in news articles. Proceedings of SemEval (2020). [110] Giovanni Da San Martino, Stefano Cresci, Alberto Barrón-Cedeño, Seunghak Yu, Roberto Di Pietro, and Preslav Nakov. 2020. A Survey on Computational Propaganda Detection. In Proceedings of IJCAI. [111] Aimi Nadrah Maseri, Azah Anir Norman, Christopher Ifeanyi Eke, Atif Ahmad, and Nurul Nuha Abdul Molok. 2020. Socio-technical mitigation effort to combat cyber propaganda: A systematic literature mapping. IEEE Access (2020). [112] Matthew DF McInnes, David Moher, Brett D Thombs, Trevor A McGrath, Patrick M Bossuyt, Tammy Clifford, Jérémie F Cohen, Jonathan J Deeks, Constantine Gatsonis, Lotty Hooft, et al. 2018. Preferred reporting items for a systematic review and meta-analysis of diagnostic test accuracy studies: the PRISMA-DTA statement. Jama 319 (2018). [113] Chris Meserole. 2018. How misinformation spreads on social media—And what to do about it. Brookings Institute 9 (2018). [114] Elena Mikhalkova, Nadezhda Ganzherli, Anna Glazkova, and Yuliya Bidulya. 2020. UTMN at SemEval-2020 task 11: A kitchen solution to automatic propaganda detection. In Proceedings of SemEval. [115] Shubham Mittal and Preslav Nakov. 2022. IITD at the WANLP 2022 Shared Task: Multilingual Multi-Granularity Network for Propaganda Detection. Proceedings of WANLP (2022). [116] Arkadiusz Modzelewski, Paweł Golik, and Adam Wierzbicki. 2024. Bilingual propaganda detection in diplomats’ tweets using language models and linguistic features. Proceedings of SEPLN (2024). [117] Salar Mohtaj and Sebastian Möller. 2022. TUB at WANLP22 shared task: Using semantic similarity for propaganda detection in Arabic. In Proceedings of WANLP. [118] Pablo Moral, Jesús M Fraile, Guillermo Marco, Anselmo Peñas, and Julio Gonzalo. 2024. Overview of DIPROMATS 2024: Detection, characterization and tracking of propaganda in messages from diplomats and authorities of world powers. Procesamiento del lenguaje natural 73 (2024). [119] Pablo Moral, Guillermo Marco, Julio Gonzalo, Jorge Carrillo-de Albornoz, and Iván Gonzalo-Verdugo. 2023. Overview of DIPROMATS 2023: automatic detection and characterization of propaganda techniques in messages from diplomats and authorities of world powers. Procesamiento del lenguaje natural 71 (2023). [120] Marco Emanuel Casavantes Moreno, Manuel Montes-Y-Gómez, Luis Carlos González Gurrola, Alberto Barrón Cedeno, and Puebla Santa Marıa de Tonantzintla. 2022. A Multidimensional Analysis of Text for Automated Detection of Computational Propaganda in Twitter Technical Report: CCC-22-003. (2022). [121] Gaku Morio, Terufumi Morishita, Hiroaki Ozaki, and Toshinori Miyoshi. 2020. Hitachi at SemEval-2020 task 11: An empirical study of pre-trained transformer family for propaganda detection. In Proceedings of SemEval. [122] A Muthukumar, M Thanga Raj, R Ramalakshmi, A Meena, and P Kaleeswari. 2024. Fake and propaganda images detection using automated adaptive gaining sharing knowledge algorithm with DenseNet121. Journal of Ambient Intelligence and Humanized Computing 15 (2024). [123] Sara Nabhani, Claudia Borg, Khalid Al Khatib, and Kurt Micallef. 2025. Integrating Argumentation Features for Enhanced Propaganda Detection in Arabic Narratives on the Israeli War on Gaza. In Proceedings of Nakba-NLP. [124] Preslav I Nakov, Giovanni Da San Martino, and YU Seunghak. 2023. Explainable propaganda detection. US Patent App. 17/855,248. [125] Iva Nenadić. 2019. Unpacking the" European approach" to tackling challenges of disinformation and political manipulation. Internet policy review 8 (2019). [126] Nic Newman, Richard Fletcher, Craig T Robertson, A Ross Arguedas, and Rasmus Kleis Nielsen. 2024. Reuters Institute digital news report 2024. Reuters Institute for the study of Journalism. [127] Lynnette HX Ng and Araz Taeihagh. 2021. How does fake news spread? Understanding pathways of disinformation spread through APIs. Policy & Internet 13 (2021). [128] Vitaliia-Anna Oliinyk, Victoria Vysotska, Yevhen Burov, Khrystyna Mykich, and Vitor Basto-Fernandes. 2020. Propaganda Detection in Text Data Based on NLP and Machine Learning.. In MoMLeT+ DS. [129] Rashmikiran Pandey, Mrinal Pandey, and Alexey Nazarov. 2022. In Proceedings of ICAC3N.

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

[130] Indravadan Patel, Hien Nguyen, Eugene Belyi, Yonas Getahun, Sarah Abdulkareem, Philippe J Giabbanelli, and Vijay Mago. 2017. Modeling information spread in polarized communities: Transitioning from legacy media to a Facebook world. In Proceedings of SoutheastCon. [131] Rajaswa Patil, Somesh Singh, and Swati Agarwal. 2020. Bpgc at semeval-2020 task 11: Propaganda detection in news articles with multi-granularity knowledge sharing and linguistic features based ensemble learning. Proceedings of SemEval (2020). [132] Vladimir Pezo, Lovro Kovacic, and Lucija Domic. 2024. Is It All About Framing? Impact of Framing on Propaganda Detection. Text Analysis and Retrieval Course Project Reports (2024). [133] Jakub Piskorski, Nicolas Stefanovitch, Giovanni Da San Martino, and Preslav Nakov. 2023. Semeval-2023 task 3: Detecting the category, the framing, and the persuasion techniques in online news in a multi-lingual setup. In Proceedings of SemEval. [134] Vyoma Harshitha Podapati, Divyansh Nigam, and Sanchari Das. 2025. SoK: a systematic review of context-and behavior-aware adaptive authentication in mobile environments. In Proceedings of HAISA. [135] Bruno Polonijo, Sabrina Šuman, and Ivan Šimac. 2021. Propaganda detection using sentiment aware ensemble deep learning. In Proceedings of MIPRO). [136] Manita Pote. 2024. Computational Propaganda Theory and Bot Detection System: Critical Literature Review. arXiv preprint arXiv:2404.05240 (2024). [137] Piotr Przybyła and Konrad Kaczyński. 2023. Where Does It End? Long Named Entity Recognition for Propaganda Detection and Beyond. In Proceedings of SEPLN. [138] Antonio Purificato, Roberto Navigli, et al. 2023. Apatt at semeval-2023 task 3: The sapienza nlp system for ensemble-based multilingual propaganda detection. In Proceedings of SemEval. [139] Dora Pušelj and Tena Škalec. 2020. Propaganda in Press: Challenges of Automatic Detection. Text Analysis and Retrieval Course Project Reports (2020). [140] Xidi Qu, Qin Hu, and Shengling Wang. 2020. Privacy-preserving model training architecture for intelligent edge computing. Computer Communications 162 (2020), 94–101. [141] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1 (2019). [142] Mohamed Ibrahim Ragab, Ensaf Hussein Mohamed, and Walaa Medhat. 2025. Multilingual Propaganda Detection: Exploring Transformer-Based Models mBERT, XLM-RoBERTa, and mT5. In Proceedings of NakbaNLP. [143] Kasturi Rangan Raghavan, Supriyo Chakraborty, Mani Srivastava, and Harris Teague. 2012. Override: A mobile privacy framework for context-driven perturbation and synthesis of sensor data streams. In Proceedings of the SenSys. [144] Mayank Raj, Ajay Jaiswal, Ankita Gupta, Sudeep Kumar Sahoo, Vertika Srivastava, Yeon Hyang Kim, et al. 2020. Solomon at SemEval-2020 task 11: Ensemble architecture for fine-tuned propaganda detection in news articles. Proceedings of SemEval (2020). [145] M Thanga Raj, Muthukumar Arunachalam, R Ramalakshmi, Kottaimalai Ramaraj, Meena Arunachalam, and P Kaleeswari. 2023. A Review on the Detection of Deep Fake and Propaganda Videos and Images-based Voice and Facial Manipulation using AI Techniques. In Proceedings of ICACRS. [146] Malavikka Rajmohan, Rohan Kamath, Akanksha P Reddy, and Bhaskarjyoti Das. 2022. Emotion enhanced domain adaptation for propaganda detection in Indian social media. In Proccedings of ICICV. [147] Eshrag Ali Refaee, Basem Ahmed, and Motaz Saad. 2022. AraBEM at WANLP 2022 Shared Task: Propaganda Detection in Arabic Tweets. In Proceedings of WANLP. [148] Julian Richards. [n. d.]. The Use of Discourse Analysis in Propaganda Detection and Understanding. In Routledge Handbook of Disinformation and National Security. Routledge. [149] Ieva Rizgelienė and Gražina Korvel. 2024. Comparative Analysis of Various Data Balancing Techniques for Propaganda Detection in Lithuanian News Articles. In Proceedings of DB&IS. [150] Francisco-Javier Rodrigo-Ginés, Jorge Carrillo-de Albornoz, and Laura Plaza. 2023. Hierarchical Modeling for Propaganda Detection: Leveraging Media Bias and Propaganda Detection Datasets.. In Proceedings of SEPLN. [151] Nathan Roll and Calbert Graham. 2024. Greybox at semeval-2024 task 4: Progressive fine-tuning (for multilingual detection of propaganda techniques). In Proceedings of SemEval. [152] Rodrigo Roman, Jianying Zhou, and Javier Lopez. 2013. On the features and challenges of security and privacy in distributed internet of things. Computer networks 57, 10 (2013), 2266–2279. [153] Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. 2019. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of CVPR. [154] BC RADOSLAV SABOL. 2022. Propaganda Detection using Stylometric Text Analysis. (2022).

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

[155] Radoslav Sabol and Aleš Horák. 2023. Augmenting Stylometric Features to Improve Detection of Propaganda and Manipulation. Proceedings of RASLAN (2023). [156] Tahereh Saheb. 2023. Ethically contentious aspects of artificial intelligence surveillance: a social science perspective. AI and Ethics 3 (2023). [157] Suleiman Saka and Sanchari Das. 2024. Evaluating Privacy Measures in Healthcare Apps Predominantly Used by Older Adults. arXiv preprint arXiv:2410.14607 (2024). [158] Suleiman Saka and Sanchari Das. 2025. SoK: Reviewing Two Decades of Security, Privacy, Accessibility, and Usability Studies on Internet of Things for Older Adults. arXiv preprint arXiv:2512.16394 (2025). [159] Abdelrhman Saleh, Ramy Baly, Alberto Barrón-Cedeno, Giovanni Da San Martino, Mitra Mohtarami, Preslav Nakov, and James Glass. 2019. Team QCRI-MIT at SemEval-2019 task 4: Propaganda analysis meets hyperpartisan news detection. Proceedings of SemEval (2019). [160] Hasan Saleh. 2023. Beneath the Surface: Exploring the Dark Web and its Societal Impacts. (2023). [161] Fátima C Carrilho Santos. 2023. Artificial intelligence in automated detection of disinformation: A thematic analysis. Journalism and Media 4 (2023). [162] Farida Habib Semantha, Sami Azam, Bharanidharan Shanmugam, and Kheng Cher Yeo. 2023. Pbdinehr: A novel privacy by design developed framework using distributed data storage and sharing for secure and scalable electronic health records management. Journal of Sensor and Actuator Networks 12 (2023). [163] Uzair Shah, Md Rafiul Biswas, Marco Agus, Mowafa Househ, and Wajdi Zaghouani. 2024. MemeMind at ArAIEval shared task: generative augmentation and feature fusion for multimodal propaganda detection in Arabic memes through advanced language and vision models. In Proceedings of ArabicNLP. [164] Sakib Shahriar, Sonal Allana, Seyed Mehdi Hazratifard, and Rozita Dara. 2023. A survey of privacy risks and mitigation strategies in the artificial intelligence life cycle. IEEE Access 11 (2023), 61829–61854. [165] Kyarash Shahriari and Mana Shahriari. 2017. IEEE standard review—Ethically aligned design: A vision for prioritizing human wellbeing with artificial intelligence and autonomous systems. In Proceedings of IHTC). [166] Chengcheng Shao, Pik-Mai Hui, Lei Wang, Xinwen Jiang, Alessandro Flammini, Filippo Menczer, and Giovanni Luca Ciampaglia. 2018. Anatomy of an online misinformation network. PloS one 13 (2018). [167] Mohamad Sharara, Wissam Mohamad, Ralph Tawil, Ralph Chobok, Wolf Assi, and Antonio Tannoury. 2022. Arabert model for propaganda detection. In Proceedings of WANLP. [168] Filipo Sharevski, Jennifer Vander Loop, and Sanchari Das. 2025. Social Media Misinformation and Voting Intentions: Older Adults’ Experiences with Manipulative Narratives. Proceedings of CSCW 9, 2 (2025), 1–29. [169] Filipo Sharevski, Jennifer Vander Loop, Peter Jachim, Amy Devine, and Sanchari Das. 2024. ’Debunk-it-yourself’: health professionals strategies for responding to misinformation on TikTok. In Proceedings of NSPW. 35–55. [170] Reza Shokri and Vitaly Shmatikov. 2015. Privacy-preserving deep learning. In Proceedings of the 22nd ACM SIGSAC conference on computer and communications security. 1310–1321. [171] Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017. Membership Inference Attacks Against Machine Learning Models. In 2017 IEEE Symposium on Security and Privacy, SP 2017, San Jose, CA, USA, May 22-26, 2017. IEEE Computer Society, 3–18. doi:10.1109/SP.2017.41 [172] Marwa Solla, Hassan Ebrahem, Alya Issa, Harmain Harmain, and Abdusalam Nwesri. 2024. Sahara Pioneers at FIGNEWS 2024 shared task: Data annotation guidelines for propaganda detection in news items. In Proceedings of ArabicNLP. [173] Veronika Solopova, Viktoriia Herman, Christoph Benzmüller, and Tim Landgraf. 2024. Check News in One Click: NLP-Empowered Pro-Kremlin Propaganda Detection. In Proceedings of EACL. [174] Veronika Solopova, Oana-Iuliana Popescu, Christoph Benzmüller, and Tim Landgraf. 2023. Automated multilingual detection of pro-kremlin propaganda in newspapers and telegram posts. Datenbank-Spektrum 23 (2023). [175] Timo Spinde, Manuel Plank, Jan-David Krieger, Terry Ruas, Bela Gipp, and Akiko Aizawa. 2021. Neural Media Bias Detection Using Distant Supervision With BABE-Bias Annotations By Experts. In Findings of EMNLP. [176] Kilian Sprenkamp, Daniel Gordon Jones, and Liudmila Zavolokina. 2023. Large language models for propaganda detection. arXiv preprint arXiv:2310.06422 (2023). [177] Malyk Stefan-Yurii. 2024. Cyrillic text classification using semantic capabilities of LLM Embeddings: A propaganda detection use case. (2024). [178] Aishwarya Surani and Sanchari Das. 2026. Co-designing MESA-Bot: Enhancing Accessibility, Privacy, Security, and Trust in a Mental Health Chatbot for Older Adults. In Proceedings of CHI. [179] Liyakathunisa Syed, Abdullah Alsaeedi, Lina A Alhuri, and Hutaf R Aljohani. 2023. Hybrid weakly supervised learning with deep learning technique for detection of fake news from cyber propaganda. Array 19 (2023). [180] Joanna Szwoch, Mateusz Staszkow, Rafal Rzepka, and Kenji Araki. 2024. Limitations of Large Language Models in Propaganda Detection Task. Applied Sciences

Goswami et al.

14 (2024). [181] Faiza Tazi, Suleiman Saka, Shradha Neupane, Ethan Myers, Sanchari Das, Lorenzo De Carli, and Indrakshi Ray. 2025. A Multi-Dimensional Analysis of IoT Companion Apps: a Look at Privacy, Security and Accessibility. IEEE Transactions on Services Computing (2025). [182] Junfeng Tian, Min Gui, Chenliang Li, Ming Yan, and Wenming Xiao. 2021. Mind at semeval-2021 task 6: Propaganda detection using transfer learning and multimodal fusion. In Proceedings of SemEval. [183] Lin Tian, Xiuzhen Zhang, Maria Myung-Hee Kim, and Jennifer Biggs. 2023. Efficient Text-based Propaganda Detection via Language Model Cascades.. In Proceedings of SEPLN. [184] Matthew Tiessen. 2023. Privacy, Propaganda, and Digital ID: Why Our Delicate Values Must Be Deliberately Defended. Washington University Review of Philosophy 3 (2023). [185] Paras Tiwari and R Eswari. 2023. An LSTM based Propaganda Detection System for News Articles. In Proceedings of CICTN. 728–733. [186] Gaurav Singh Tomar. 2022. AraProp at WANLP 2022 Shared Task: Leveraging Pre-Trained Language Models for Arabic Propaganda Detection. In Proceedings of WANLP. [187] Andrea Tundis, Gaurav Mukherjee, and Max Mühlhäuser. 2020. Mixed-code text analysis for the detection of online hidden propaganda. In Proceedings of ARES. [188] Andrea Tundis, Ahmed Ali Shams, and Max Mühlhäuser. 2023. From the detection towards a pyramidal classification of terrorist propaganda. Journal of Information Security and Applications 79 (2023). [189] Aina Turillazzi, Mariarosaria Taddeo, Luciano Floridi, and Federico Casolari. 2023. The digital services act: an analysis of its ethical, legal, and social implications. Law, Innovation and Technology 15 (2023). [190] Blase Ur and Yang Wang. 2013. A cross-cultural framework for protecting user privacy in online social media. In Proceedings of WWW. [191] Antoni Valls, Björn Komander, and Jesús Cerquides. 2024. GeopoliticallyInformed multiModal BERT for propaganda detection in political tweets. In Proceedings of CEUR. [192] Ekansh Verma, Vinodh Motupalli, and Souradip Chakraborty. 2020. Transformers at SemEval-2020 Task 11: Propaganda fragment detection using diversified bert architectures based ensemble learning. In Proceedings of SemEval. [193] Prashanth Vijayaraghavan and Soroush Vosoughi. 2022. TWEETSPIN: Finegrained propaganda detection in social media using multi-view representations. In Proceedings of NAACL. [194] George-Alexandru Vlad, Mircea-Adrian Tanase, Cristian Onose, and DumitruClementin Cercel. 2019. Sentence-level propaganda detection in news articles with transfer learning and BERT-BiLSTM-capsule model. In Proceedings of NLP4IF. [195] Paul Voigt and Axel Von dem Bussche. 2017. The eu general data protection regulation (gdpr). A Practical Guide, 1st Ed., Cham: Springer International Publishing 10 (2017). [196] Vorakit Vorakitphan, Elena Cabrio, and Serena Villata. 2022. Protect: A pipeline for propaganda detection and classification. In Proceedings of CLiC-it. [197] Soroush Vosoughi, Deb Roy, and Sinan Aral. 2018. The spread of true and false news online. science 359 (2018). [198] Hannah C Waight. 2021. Chinese Propaganda Detection. (2021). [199] Liqiang Wang, Xiaoyu Shen, Gerard de Melo, and Gerhard Weikum. 2020. CrossDomain Learning for Classifying Propaganda in Online Contents. In Proccedings of Truth and Trust Online Conference. [200] Ruize Wang, Duyu Tang, Nan Duan, Wanjun Zhong, Zhongyu Wei, Xuan-Jing Huang, Daxin Jiang, and Ming Zhou. 2020. Leveraging Declarative Knowledge in Text and First-Order Logic for Fine-Grained Propaganda Detection. In Proceedings of EMNLP. [201] Adeela Waqar, Asad Raza, Haider Abbas, and Muhammad Khurram Khan. 2013. A framework for preservation of cloud users’ data privacy using dynamic reconstruction of metadata. Journal of Network and Computer Applications 36, 1 (2013), 235–248. [202] Senuri Wijenayake, Joanne Gray, Asangi Jayatilaka, Louise La Sala, Nalin Arachchilage, Ryan M Kelly, and Sanchari Das. 2025. Advancing Interdisciplinary Approaches to Online Safety Research. In Proceedings of OZCHI. [203] William Williamson III and James Scrofani. 2019. Trends in detection and characterization of propaganda bots. In Proceedings of HICSS. [204] Hanqian Wu, Xinwei Li, Lu Li, and Qipeng Wang. 2022. Propaganda Techniques Detection in Low-Resource Memes with Multi-Modal Prompt Tuning. In Proceedings of ICME). [205] Madelyne Xiao and Jonathan Mayer. 2023. SoK: Machine Learning for Misinformation Detection. arXiv e-prints (2023), arXiv–2308. [206] Yunze Xiao and Firoj Alam. 2023. Nexus at ArAIEval Shared Task: Fine-Tuning Arabic Language Models for Propaganda and Disinformation Detection. Proceedings of ArabicNLP (2023). [207] Lidong Xing, Nannan Hou, Zhiqing Zhang, Ke Li, and Fangxu Meng. 2025. Research on false propaganda detection technology based on LLM and BERT. In Proceedings of ICMIC.

SoK: Analysis of Privacy Risks and Mitigation in Online Propaganda Detection through the PROMPT Framework

PROMPT framework integration and defense selection is shown Algorithm 1.

•

•

• •

•

•

•

• • • •

• •

Doubt

Repetition

Slogans

• •

•

•

•

•

•

Flag-waving

• •

Loaded Language

•

Name calling, labeling

• •

Appeal to authority

•

Black-and-white fallacy

•

Appeal to fear/prejudice

Exaggeration, minimization

•

Thought-terminating cliché

Propaganda Types Political [41, 42, 177] Religious [5] Social [62, 67, 129] Economic [122] Cultural [1, 122, 212] Public Health [135]

Bandwagon, Reductio ad Hitlerum

Propaganda Techniques

Causal oversimplification

Artefacts

Table 9: Propaganda Types and Techniques.

Whataboutism, Straw man, Red herring

[208] Hye Young You. 2020. Foreign Agents Registration Act: a user’s guide. Interest Groups & Advocacy 9 (2020). [209] Seunghak Yu, Giovanni Da San Martino, Mitra Mohtarami, James Glass, and Preslav Nakov. 2021. Interpretable Propaganda Detection in News Articles. In Proceedings of RANLP. [210] Savvas Zannettou, Tristan Caulfield, Emiliano De Cristofaro, Michael Sirivianos, Gianluca Stringhini, and Jeremy Blackburn. 2019. Disinformation warfare: Understanding state-sponsored trolls on Twitter and their influence on the web. In Proceedings of WWW. [211] Liudmila Zavolokina, Kilian Sprenkamp, Zoya Katashinskaya, Daniel Gordon Jones, and Gerhard Schwabe. 2024. Think fast, think slow, think critical: designing an automated propaganda detection tool. In Proceedings of CHI. [212] Mohamed Zaytoon, Nagwa M El-Makky, and Marwan Torki. 2024. AlexUNLPMZ at ArAIEval Shared Task: contrastive learning, llm features extraction and multi-objective optimization for arabic multi-modal meme propaganda detection. In Proceedings of ArabicNLP. [213] Wenshan Zhang and Xi Zhang. 2022. Cross-Lingual Propaganda Detection. In Proceedings of IEEE BigData.

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

• •

•

• •

•

Algorithm 1 PROMPT Integration and Defense Selection

Security Concerns

C [142, 173]

Clickbait Misuse

Privacy Analysis

• • • • • • •

S [89, 120]

• •

• •

A [95, 183]

• • • • • • • • • • •

D [123, 203]

•

• • • • • • • • •

•

Table 11: Propaganda Mitigation Techniques.

A [110, 122] D [116, 204]

• •

•

•

• •

•

•

FL

•

HITL

DP

XAI

•

HE

•

Privacy by Design (PbD)

•

Blockchain-Based Storage

•

Graph Anomaly Detection

Fact-Checking Integration

•

Metadata Privacy Filters

•

•

•

•

Synthetic Data Generation

•

K-Anonymity and L-Diversity

•

AI-Based Content Moderation

•

Blockchain-Based Verification

•

Privacy-Preserving Image Processing

•

Privacy-Enhancing Technologies (PETs)

C [116, 136, 204] S [19, 162]

Secure Multi-Party Computation (SMPC)

Privacy Analysis

Adversarial Machine Learning Defenses

Mitigation Techniques Privacy-Preserving Machine Learning (PPML)

Detailed propaganda detection survey results of Language, Platform, Datasets, Benchmark Competition, Computational Methodology are shown in Table 8. Statistical results are shown in Tables 9, 10, 11, 12, 13.

Table 10: Security Concerns of Propaganda Detection.

Cross-Lingual Transfer Learning Risk Multimodal Data Privacy Violation Community Clustering Exploitation Image Processing Privacy Risk User-Generated Content Risk Social Media Manipulation Irrelevant Span Detection Bridge Node Exploitation Metadata Privacy Risk LLM Privacy Violation Data Security Breach Label Overprediction Information Warfare Facial Recognition Random Truncation Threat Generation Undefined Content False Information Centrality Attack Misclassification Re-identification Consent Violation Algorithmic Bias SEO Manipulation Deepfake Attacks Surveillance Sybil Attack

Require: Inputs X, techniques 𝑇 , stages ( C, S, A, D ), mitigations 𝑆, obligations L, thresholds 𝜏𝑐 , 𝜏𝑒 , budget 𝐵, tolerance 𝜖, min improv. 𝜂, max iters 𝐾 Ensure: Defenses S★ , updated model 𝑓 ′ , audit Audit 1: Featurize from 𝑇 : compute Coverage(𝐷 𝑗 ) and embed 2: Init risk 𝑟 ident ← LinkageAcc, 𝑟 meta ←𝐼 (𝑀; 𝑈 ), 𝑟 bias ← ΔEO , 𝑟 net ←𝐴 𝑅® ← ⟨𝑟 ident , 𝑟 meta , 𝑟 bias , 𝑟 net ⟩, 𝐶 ← 0, 𝑓 ′ ← 𝑓 , S★ ← ∅, 𝑘 ← 0 3: while ∥ 𝑅® ∥ 2 > 𝜖 and 𝑘 < 𝐾 do 4: 𝑘 ←𝑘 +1 5: Filter feasible candidates Cand ← {𝑠 ∈ 𝑆 : 𝐶 + Cost(𝑠 ) ≤ 𝐵, CompScore( 𝑓 ′ ) ≥ 𝜏𝑐 , 𝐹 ( 𝑓 ′ ) ≥ 𝜏𝑒 } 6: if Cand = ∅ then break // no feasible action 7: end if 8: Score each 𝑠 ∈ Cand 9: for all 𝑠 ∈ Cand do 10: Δ𝑅 (𝑠 ) ← ∥ 𝑅® ∥ 2 − ∥ 𝑅® − Mitigate(𝑠 ) ∥ // r isk reduction 11: 𝑈 (𝑠 ) ← 𝛼 Δ𝑅 (𝑠 ) − 𝛽 PerfLoss(𝑠 ) − 𝛾 Cost(𝑠 ) 12: end for  13: 𝑠 ★ ← arg max𝑠 ∈Cand 𝑈 (𝑠 ), Δ𝑅 (𝑠 )/max{1, Cost(𝑠 ) } 14: if 𝑈 (𝑠 ★ ) ≤ 0 or Δ𝑅 (𝑠 ★ ) < 𝜂 then break // no beneficial action 15: end if 16: Apply 𝑠 ★ , update 𝑅® ← 𝑅® − Mitigate(𝑠 ★ ), 𝐶 ← 𝐶 + Cost(𝑠 ★ ), S★ ← S★ ∪ {𝑠 ★ } 17: Update compliance and fairness 1 ∑︁ 1[ 𝑓 ′ satisfies ℓ ] CompScore( 𝑓 ′ ) ← | L | ℓ ∈L 𝐹 ( 𝑓 ′ ) ← fairness metric value 18: 𝑆 ← 𝑆 \ {𝑠 ★ } 19: end while ® S★, 𝐶, CompScore( 𝑓 ′ ), 𝐹 ( 𝑓 ′ ), violated constraints} 20: Audit ← { 𝑅, 21: return S★, 𝑓 ′ , Audit

•

• •

• •

•

ASIA CCS ’26, June 1–5, 2026, Bangalore, India

Goswami et al.

Table 8: Survey Result of the Propaganda : Language, Platform, Datasets, Benchmark Competitions, Computational Methodology Reported for all 162 Papers. Key Points Language

Survey Results Asia: Chinese [8, 198, 207, 213], Hindi [4, 14, 15, 17, 31, 54, 62, 86, 92, 101, 123, 142, 163, 172, 192], Urdu [6, 8], Code-Mixed (Hinglish, Tenglish-Telugu-English, Hindi-English mix) [62] Middle East and Africa: Arabic [1, 4, 8, 12–14, 17, 18, 33, 34, 54, 56, 58, 71, 83, 86, 92, 98, 99, 106, 115, 117, 123, 142, 147, 148, 151, 163, 167, 172, 186, 192, 206, 212], Hebrew [4, 13, 14, 17, 36, 54, 86, 89, 90, 92, 101, 123, 132, 138, 142, 163, 172, 173, 192] Europe and America: English [2, 4, 4–6, 8–10, 13–15, 17, 22, 23, 26, 27, 29, 30, 32, 35, 36, 38, 39, 41, 43, 44, 46, 50, 52–55, 58, 61–63, 65–67, 79, 80, 86, 87, 92, 96, 99, 101, 102, 104–106, 108, 109, 111, 114, 116, 118–121, 123, 124, 128, 129, 131, 132, 135–139, 142, 144, 146, 148, 150, 151, 159, 163, 172–174, 176, 179, 180, 182, 183, 185, 187, 188, 191–194, 196, 200, 204, 209, 211, 213], Bulgarian [106, 151, 177], Czech [21, 76, 154, 155], Dutch [58], French [4, 13, 14, 17, 36, 54, 86, 92, 101, 123, 132, 138, 142, 163, 172–174, 192], German [13, 36, 89, 90, 101, 132, 138, 173], Georgian [101, 132], Greek [101, 132], Italian [13, 36, 51, 101, 132, 138, 173, 196], Lithuanian [149], North Macedonian [106, 151], Polish [13, 36, 101, 132, 138, 180], Romanian [173, 174], Russian [4, 13, 36, 55, 75, 89, 90, 99, 101, 132, 138, 173, 174, 177], Serbian [177], Spanish [9, 27, 29, 38, 53, 99, 101, 116, 118, 119, 132, 150, 191], Ukrainian [93–95, 173, 174, 177] Platform Social Media: Twitter [1, 5, 6, 9, 10, 12, 13, 18, 27, 28, 30, 33, 34, 49, 53, 56, 89, 90, 98, 99, 111, 117–120, 136, 145–147, 150, 167, 180, 186, 191, 193, 203, 206, 213], Facebook [5, 14, 30, 35, 39, 51, 55, 62, 80, 87, 99, 104, 106, 111, 123, 142, 163, 182, 204, 212], Instagram [14, 62, 212], Pinterest [14, 212], YouTube [49, 62] News Media: Al Jazeera [12], BBC Arabic [12], CNN Arabic [12], Sky News Arabia [12], Czech News Websites [154, 155] Discussion Forums: Reddit [30], Dark Web Forum [99] Messaging: WhatsApp [51, 62], Telegram [174], Skype [51] Other Digital Platforms: Taobao [207], Sina-Weibo [203] Datasets English: AG-News [8], BABE [150], COVID-19 Fake News Dataset [62], CheckThat [58], MH17 Tweets Dataset [10], DARPA Twitter Bot Challenge [203], IRA Dataset [99], Kaggle Fake News Dataset [99], MBIC [150], PTC (Propaganda Techniques Corpus) [4, 5, 8, 26, 41, 52, 66, 67, 102, 108–110, 121, 131, 137, 144, 180, 194, 200], PROPANEWS Dataset [82], PONC [180], ProSoul [6], ProText [5, 6, 8], QProp [5, 8, 110], THSP 17 [110], TweetSpin [193] Arabic: ARATWEET [115], Ar-DAD [145], ArPro [71], IED [6] Multilingual: DIPROMATS [9, 10, 27, 29, 38, 53, 116, 118, 119, 150, 183, 191], MultiProp [13], H-Prop-News Dataset [31], Czech Propaganda Benchmark Dataset [154], Chinese and English Propaganda Detection Dataset [213], Sogou [8], Propitter [28], SentiMix [62] Multimodal: Celeb-DF [145], DeeperForensics, Deepfake Forensics [145], Deepfake TIMIT [145], FaceForensics [145], FaceForensics++ [145], UADFV [145], Vid-TIMIT [145], Memotion Analysis Dataset [62], Meme Propaganda Techniques Corpus [35], FigNews [142, 172] Benchmark SemEval [4, 17, 26, 32, 43, 44, 50, 61, 80, 86, 87, 92, 96, 101, 104, 106, 109, 114, 121, 131, 138, 144, 151, 159, 182, 192], WANLP [12, 18, 33, 34, 56, 83, 98, 115, 117, 147, 167, 186], Competition IberLEF [9, 27, 29, 116, 150, 183], NLP4IF [4, 6, 10, 15, 41, 67, 79, 196], ArAIEval [1, 14, 54, 123, 163, 206, 212] Computational Statistical ML: BoW [41], Decision Trees [99, 124, 155, 203], Gradient-Boosted Decision Trees (GBDT) [155], KNN [99, 119], LDA [99, 111, 120, 145], LIWC [15], Logistic Regression [2, Methodology 99, 109, 114, 120, 128, 131, 144, 159, 187], Naive Bayes [4, 21, 50, 99, 154], PCA [145], Random Forest [55, 109, 154, 203], Recursive Binary Split (RBS) classification tree [55], SVM [4, 21, 50, 99, 106, 111, 144, 145, 173, 174, 187, 193, 213], TF-IDF [4, 7, 38, 102, 128], XGBoost [43, 131, 145, 213], Linear Regression [174] Deep Learning: Bi-LSTM [4, 17, 26, 31, 43, 44, 194], CNN [5, 31, 35, 41, 44, 46, 65, 109, 111, 120, 131, 145, 146, 187, 192],CapsuleNet [122], ConvoNet [135], DenseNet121 [122], ELMo [4, 4, 8, 17, 192], GANs [145], GRU [4, 93, 184],LSTMs [17, 131], Multi-Granularity Network [115], RNN [6, 95, 111, 120, 129, 145],SCST [82],SpanNER [137] Transformers: ALBERT [5, 61, 63, 138, 204], AraBERT [1, 12, 18, 34, 54, 56, 98, 117, 123, 167], AraELECTRA [34], AraGPT [1], BART [82], BERT [2, 4, 5, 12, 15, 23, 31, 35, 39, 41, 44, 52, 61, 63, 65, 75, 86, 87, 96, 102, 108, 109, 120, 137, 138, 144, 147, 174, 177, 192, 194, 196, 204, 207, 209], BERTweet [27, 29], DeHateBERT [34], DistilBERT [30, 61, 138], HerBERT [138], MARBERT [12, 34, 58], mBERT [115, 142, 213], mT5 [142], RoBERTa [4–6, 8, 12, 34, 38, 58, 61, 63, 75, 88, 104, 109, 119, 131, 138, 142, 144, 150, 154, 196, 213], RoBERTuito [27, 29], Robeczech [154], TwHIN-BERT [191], XLM-R [12, 34, 38, 58, 83, 109, 115, 142, 150, 213], XLNet [5, 61, 63, 109, 131, 138] Vision: CLIP [35], DALLE2 [163], InceptionV3 [63], NF-ResNet50 [204], ResNet-152 [35, 63], ResNet101 [39], ResNet50 [14, 104, 204], UNITER [63], VGG-19 [63], VLM [106], ViLBERT [63], VisualBERT [63, 204], YOLO-CNN [145] LLMs: GPT-2 [13, 121, 131], GPT-3 [67, 177, 180], GPT-3.5 [67, 177, 180], GPT-4 [2, 67, 71, 123, 177, 180, 211], GROVER [82], Mistral [105], Mixtral [151],LLaMa2 [151], Chain of Thought [176], LM-BFF [204] Other Techniques: Botometer [136], HDSF [82], LatexPRO [200], MViTO-GAT [35], OCR [182, 187], Reinforcement Learning [110], VADER [135]

Table 12: Regulatory Compliance of Propaganda Detection.

Table 13: Ethical Considerations of Propaganda Detection.

C [9, 38, 191] S [95, 203]

•

A [23, 70]

•

•

D [70, 173]

•

•

• • •

•

•

•

•

•

• •

• •

Right to Be Forgotten

•

•

Bias & Discrimination

•

Surveillance & Privacy

•

Data Breach Accountability

•

•

Algorithmic Fairness & Bias

•

Censorship

•

Responsible AI

•

Cross-Border Data Ethics

•

Encryption & Data Protection

Privacy Analysis •

Data Ownership & Control

•

•

•

Automated Decision-Making Risks

•

Informed Consent & Transparency

•

Transparency in Social Engineering

•

•

DSA

PIPL •

•

Commercial Exploitation of User Data

•

•

GDPR

FARA

CISA

PDPA

NIST •

CCPA

HIPAA

FedRAMP

PCI DSS

•

Government & Corporate Surveillance Risks

• •

•

Data Manipulation & Misinformation

D [59, 95, 173]

•

Ethical Considerations

•

S [95, 189, 203, 208] A [70, 173, 191]

ISO 27001

EU AI Act

Basel III

Honest Ads Act

•

IEEE AI Ethics

C [9, 23, 38, 84]

FSB AI Guidelines

Privacy Analysis

UNESCO-Digital Platform Governance

Regulatory Compliances

• •

•

•

Record · ID 120425 · SHA-256 e566c8292c880918
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.