Taxonomy of Risks on Automated Fact-Checking Systems Considering its Propagation Jun Yajima ∗1 , Tatsuya Oka1 , and Takao Okubo2
arXiv:2606.25645v1 [cs.CR] 24 Jun 2026
1
2
Fujitsu Limited Institute of Information Security
Abstract
1
Introduction
In recent years, posting and spreading fake news on social network services (SNS) is a major social problem. For example, at the 2016 U.S. presidential election, false rumors about a pizza restaurant spread and led to a shooting incident [1]. In Japan, whenever a massive earthquake strikes, posts claiming that the earthquake was artificially triggered, posts that stoke anxiety by sharing footage of past tsunamis even though no tsunami has actually occurred, and posts about earthquake prediction lacking scientific basis often go viral[2]. These posts are made for a variety of reasons, including cases where the goal is to generate ad revenue through page views, potential disinformation campaigns originating overseas, and instances where those sharing the content genuinely believe the misinformation and spread it. To avoid being misled by fake news, factchecking is essential. Fact-checking is carried out by fact-checking organizations such as PolitiFact[3] in the United States, Full Fact[4] in the United Kingdom, and the Japan Fact-Checking Center[5] in Japan. While fact-checking organizations accurately verify the truthfulness of major posts, not all posts on SNS are fact-checked, and social media users cannot immediately judge the truthfulness of every post they see. Therefore, automated factchecking systems which can perform fact-checking automatically are expected to serve as a complementary tool to the fact-checking conducted by fact-checking organizations. Research into automated fact-checking technology is ongoing, and technological development is progressing. For example, the development project of automated factchecking systems in Japan is being conducted[6].
In recent years, the posting of fake news including disinformation and misinformation on social networking services (SNS) has become a social problem. To combat this fake news, fact-checking that is the process of assessing the veracity of posts on SNS has become increasingly important. While factchecking is currently performed by fact-checking organizations, it is difficult to fact-check all posts on SNS. Therefore, the use of automated factchecking systems is effective. Recent automated fact-checking systems utilize artificial intelligence and large language models, so there are risks of incorrect judgments and posting incorrect results on social media which can lead to the spread of misinformation or to engage in defamation. In this paper, as a first step toward enabling the safe use of automated fact-checking systems, we categorize the specific risks on automated fact-checking systems. In this categorizing, we consider a threestage risk propagation: risk factors, hazardous situations, and harm. Our analysis revealed that 32 specific risks exist in automated fact-checking systems. In this paper, we utilize the categorized risks as analytical cues (guide words) to present the risk assessment of the automated fact-checking system DEFAME. This assessment result indicates that risks that cannot be derived using STRIDE, a conventional IT security risk assessment method can be derived using our guide words.
1
Applications of automated fact-checking systems include use by analysts in fact-checking organizations, use by corporate and other organizational users to identify dis/misinformation targeting their organizations, and use by social media users to factcheck posts they encounter. Since automated fact-checking systems are typically implemented using artificial intelligence (AI) and large language models (LLMs), there are concerns about various risks associated with these technologies. For example, users seeking to generate ad revenue might input sensational disinformation into an automated fact-checking system and repeatedly tweak the text little by little until the system classifies the claim as “true.” If they can obtain their disinformation that is deemed “true”, they use it as an endorsement to try to generate a large number of impressions. For example, take the false claim that “Yesterday’s earthquake was artificially triggered.” By cleverly manipulating (or modifying) the input text and feeding it into the system repeatedly, the attacker continues until a “true” result is obtained. Due to the variability and hallucinations in AI and LLM results, if even a single instance among the vast amount of input data yields a “true” result, there is a risk that disinformation–such as “The claim that the earthquake was artificial is correct. This automated factchecking systems have verified it as true”– will be posted and spread using the fact-checking results as an endorsement. From the perspective of developers/providers of automated fact-checking systems, it is necessary to identify the risks associated with such systems in advance and consider appropriate measures before releasing them. The following are existing technologies and initiatives regarding the risks of automated factchecking systems. Pennycook et al. analyzed fake news from a psychological perspective, examining how fake news comes to be believed and which types of topics are more likely to be believed[7]. This paper points out that automated fact-checking faces misjudgment. However, harm of such risks in normal use or abuse of automated fact-checking systems are not described. STRIDE[8] is frequently used to identify IT security risks, but we argue that identifying risks associated with recent machine learning, generative AI, and risks specific to automated fact-checking are difficult. Because STRIDE was established and widely used before
these technologies became widespread. Nakov et al. introduces initiatives in automated fact-checking and describes challenges related to its realization [9]. Guo et al. provides a survey of automated fact-checking, offering a technical overview of the field and introducing relevant datasets [10]. However, these papers do not address the potential risks or misuse associated with automated fact-checking. The Japan AI safety institute (AISI) published a report[11] summarizes attacks on AI systems and their consequences. While this information is useful, the descriptions are limited to direct damage such as output information leakage and model malfunction; they do not address the indirect harm suffered by users who receive the results of a malfunctioning model, nor do they cover further harm resulting from the misuse of those results. This is insufficient from the perspective of comprehensively covering the risks of automated fact-checking. The Poynter website[12] lists problems in fact-checking, but these problems are organized from the perspective of human fact-checkers and do not address the risks of automated fact-checking. Tanaka et al.’s categorization of generative AI risks[13] organizes generative AI risks into 20 categories and can be used as a reference for our own risk categorization. The AISI Japan’s AI Safety Evaluation Perspectives Guidelines [14] list categories of AI safety and are a useful reference. However, it does not provide detailed information on the risks involved in automated verifying the truthfulness of mis/disinformation. So, we took this guide only as a reference information. This paper identifies the risks specific to automated fact-checking systems with the aim of helping developers/providers systematically understand their risks. This paper excludes general IT security and AI security from its scope and focuses on discussing risks specific to automated factchecking systems. In this analysis, we consider a three stage chronological sequence–risk factors, hazardous situations, and harm–and organize the findings into tables of risks. After that, to confirm the appropriateness of the tables, we use each of these identified risks as analytical cues (guide words) to extract risks associated with the automated fact-checking system DEFAME and discuss the appropriateness of result. Our analysis clarified that we could categorize the risks on automated fact-checking systems into 32 types and clarified 2
these risks could be used as guide words to identify risks in automated fact-checking systems (e.g., DEFAME.) This paper is organized as follows. In Section 2, we introduce fake news and explain the importance of automated fact-checking systems and the risks associated with them. In Section 3, we review prior research related to our analysis. In Section 4, we categorize the risks of automated fact-checking systems and describe a fault tree of the logical relationships between each risk. In Section 5, we present the results of an analysis of DEFAME’s risks based on the risk classification results. In Section 6, we discuss the risk classification presented in this paper. In Section 7, we provide a summary and future works.
Since ad revenue can sometimes be earned based on the number of views or clicks (impressions) of content on SNS, fake content may be created for financial gain. • Political motives (attacks on the social system through incitement of the masses) Fake content may be created with the aim of damaging the reputation of an opponent to benefit a supported political party or candidate during political events such as elections. • Attacking indivisuals or organizations Fake content may be created to damage the reputation of disliked celebrities or companies and undermine their credibility. • Hedonistic motives
1.1
Contributions of this paper
Fake content may be created simply to derive satisfaction from spreading disinformation.
The contributions of this paper are as follows.
Some well-known examples of fake news include 1. Identified 32 risks specific to automated factthe following: checking systems as a basic study of risks on them • The Pizzagate Incident Suspicions of a pizza shop was involved in criminal activity led to a shooting incident[1].
2. Organized the identified risks considering chronological sequence (risk factors, hazardous situations, harm) and represented them into a fault tree (enabling the tracing of logical relationships)
• The lion hoax during the Kumamoto Earthquake in Japan A hoax claiming that a lion had escaped from a zoo was posted, causing severe anxiety among disaster victims. The poster was arrested for obstruction of business by fraudulent means[15].
3. A case study demonstrated that risks can be systematically derived on a Data Flow Diagram (DFD) through risk assessment using the 32 identified risks as guide words
4. A comparison with STRIDE demonstrated These examples show that the spread of fake news that social impacts (defamation, spreading of has serious negative effects on society and making disinformation, etc.) and other risks can be countermeasures are necessary. assessed that cannot be assessed by STRIDE.
2.2
Fact-checking
2
Fake news and automated Fact-checking is one of the measures taken to combat fake news. Fact-checking is used for verifact-checking system fying the truthfulness of social media posts and
2.1
Fake news
other content by consulting the reliability of evidence information obtained from official and/or other sources. Fact-checking is conducted by organizations such as PolitiFact, FullFact, and Japan Fact Check Center (JFC), and findings of several impactive posts
The main purposes of generating and spreading fake news are as follows: • Financial motives 3
on SNS are published. While fact-checking by these organizations is reliable, the scope of their checks is limited to posts that are presumed to have a significant impact or that have been widely shared. Therefore, from the perspective of social media users, it is difficult to verify the truthfulness regarding of all posts they encounter. One solution to this problem is to use automated fact-checking systems that perform fact-checking via IT systems. Automated fact-checking systems automatically verify the truthfulness of inputted posts from various angles and output fact-checking results. Well-known examples of automated fact-checking systems are DEFAME[16] and FacTool[17].
3
Related works
3.1
STRIDE
STRIDE is a technique widely used in security risk assessments in the IT field[8]. In STRIDE, creating a data flow diagram (DFD) and assessing risks on the DFD. Each letter – S, T, R, I, D, and E – represents a specific type of risk. These types of risks are used as guide principles (guide words) for assessing where and what kinds of risks may be hidden at the data boundaries and entities depicted in the DFD. The meanings of each letter in STRIDE are as follows: • S: Spoofing
2.3
Potential risks on automated fact-checking systems
• T: Tampering • R: Repudiation
While automated fact-checking systems are effective against fake news, we think they have risks specific to them. For example, since automated factchecking systems are typically implemented using artificial intelligence (AI) and large language models (LLMs), there is a risk of misjudgment caused by various technical problems (e.g., lack of accuracy, training data quality, or hallucination.) If an automated fact-checking system mistakenly classifies mis/disinformation as “true” due to technical problems, it could mislead users; in addition, if users post the results on social media, this could lead to the spread of mis/disinformation. Furthermore, there is a risk that attackers seeking to spread their disinformation could repeatedly input fake news made by them into the automated factchecking system until it yields a “true” result and then combine them into a single post stating that “This news is true. Truthfulness is guaranteed by fact-checking”. In this scenario, the automated fact-checking system would effectively be endorsing the attacker’s fake news and facing a risk that the system could inadvertently facilitate the spread of disinformation. The automated fact-checking systems are subject to various specific risks. In this paper, we categorize the risks associated with automated fact-checking systems by referencing existing risk assessment techniques and risk classifications. We then use these categorized risks as guiding principles to identify the risks of the automated factchecking system DEFAME and verify the validity of our risk classification.
• I: Information disclosure • D: Denial of service • E: Elevation of privilege STRIDE is a security risk assessment technique for IT systems that has been in use for many years. Automated fact-checking systems mostly be implemented using machine learning (ML) and LLMs, ML and LLMs have only recently been put into practical use, making it difficult for traditional technique STRIDE to identify risks associated with them. Furthermore, STRIDE cannot identify risks specific to automated fact-checking systems easily. This is discussed in Section 6.1.
3.2
Risk classification for machine learning
In 2021, the authors classified the types of risks in ML as part of an initiative independent of this study (not published). In this classification, we added guide words that were missing from STRIDE for the risk assessment of ML systems and removed guide words that were deemed unnecessary. The results of this reorganization of the guide words are as follows. • Abuse • Adversarial Tampering 4
• Denial of Service
not provide detailed information on the risks involved in automated verifying the truthfulness of mis/disinformation. So, we use this guide as only reference information. Their evaluation perspectives on AI safety are as follows.
• Plagiarism • Unauthorized Use • Poisoning
• Control of Toxic Output
• Information Disclosure • Inference
• Prevention of Misinformation, Disinformation and Manipulation
• Spoofing
• Fairness and Inclusion • Addressing to High-risk Use and Unintended Use
These can be used in place of STRIDE as guide words for risk assessment of ML systems. However, since this reorganization was conducted in 2021, it cannot capture the risks associated with LLMs and generative AI –such as ChatGPT, whose use has expanded since then.
• Privacy Protection • Ensuring Security • Explainability
3.3
Risk Taxonomy of generative AI
• Robustness
Tanaka et al. categorized the risks associated with generative AI into 20 types[13]. This categorization organizes the risks of generative AI with reference to ISO/IEC Guide 51[18]. In Tanaka et al.’s paper, they organize risks considering the chronological sequence from their origin to their impact. Although these risks are specific to generative AI and cannot be directly applied to risks of automated fact-checking systems, we decided to use them as a reference because automated fact-checking systems mostly utilize generative AI. When structuring the risks of automated fact-checking systems, we also consider the chronological sequence, as in ISO/IEC Guide 51. The specific risk items identified by Tanaka et al. will be introduced in Section 4.1 as explanation of correspondence between Tanaka’s risk items and our work of risks of automated fact-checking systems.
• Data Quality • Verifiability
4
Risks on Automated FactChecking Systems
4.1
Identifying Risks on Automated Fact-Checking Systems
As described in Section 3.1, using STRIDE alone is difficult to identify automated fact-checking systems’ specific risks. The reasons are discussed in Section 6.1. We organize the risks specific to automated fact-checking systems by referencing the risk issues described in Tanaka et al.’s paper [13], as introduced in Section 3.3. Specifically, we con3.4 Evaluation Perspecties on AI sidered the correspondence between the 20 risk issues identified in their paper and how they appear Safety in the automated fact-checking systems. The folThe Japan AISI has published a guide to evalua- lowing outlines the relationship and consideration tion perspectives on AI safety based on the con- to organizing risks in automated fact-checking syscept of AI safety, which addresses various risks as- tems based on the risk issue in their paper. We also sociated with AI[14]. This guide lists 10 items re- considered the ML risk classification in Section 3.2 lated to AI safety. This guide includes an item on and the Japan AISI evaluation guide in Section 3.4 preventing mis/disinformation. However, it does as reference information. 5
1. Hallucination
7. Privacy In appropriate provision of privacy information used in fact-checking, such as supporting evidence and assessment results, may result in privacy violations. This also involves information disclosure by adversarial attacks.
We classified “hallucinations” as ”misjudgments” in a broad sense. A misjudgment of AI occurs when some kind of misjudgment factors (e.g., hallucinations, training data bias) arise in the normal operation of fact-checking or it intentionally occurred by adversarial attacks. This can lead to the spread of misinformation, defamation, and the generation of impressions on social media for the purpose of generating ad revenue.
8. Copyright infringement In case of the evidence or assessment results contain copyrighted material, which may result in copyright infringement. 9. Exploitation of workers during model creation
2. Potential for risky emergent behaviors
Workers who remove harmful data may suffer physical and mental harm as a result.
This issue relates to misjudgment in normal use or intentional misjudging by adversarial attacks.
10. Cybersecurity This relates to an attack designed to intentionally cause misjudgments or privacy disclosure. It has the potential to lead to data leaks, the spread of mis/disinformation, defamation, and the generation of impressions for the purpose of generating ad revenue. By repeatedly accessing the system, attackers could steal training data or hijack models, or they could inject their own data into the training data to cause misjudgment.
3. Harmful content This relates to the system producing harmful output. For example, this involves producing various harmful outputs, such as violating laws or encouraging criminal activity. Alternatively, the automated fact-checking system that conducts malicious judgments or decisions that favor specific individuals or organizations. This could result in unfair judgments, which in turn could lead to the spread of mis/disinformation, defamation, and the generation of impression-based revenue.
11. Economic impacts Use of automated fact-checking systems will lead to staff cuts among fact-checkers and cause financial hardship for those affected
4. Harm of representation, allocation, and quality of service
12. Acceleration This relates to acceleration of competition with rival systems. As technological competition intensifies, there is a risk that companies may report quality as being better than it actually is, even when there is no factual basis for such claims. This raises the risk of incorrect judgments due to inadequate quality of system. This also relates to overreliance.
We think that this issue relates to the situation that automated fact-checking systems may cause misjudgment. We also think this issue relates to item of harmful content. 5. Disinformation and influence operations This is similar to cases of misjudgment and could lead to the spread of mis/disinformation, defamation, and the pursuit of page views for advertising revenue.
13. Environmental and financial cost There are concerns about power consumption and its cost when using AI.
6. Overreliance
14. Spreading misinformation
Users believe the system’s results are always correct, leading them to make incorrect judgments or spread mis/disinformation when misjudgment occurs.
If the system makes a misjudgment in normal use or case of adversarial attacks, misinformation is spread. 6
15. Increasing sophistication and ease of crime
judgment on social media, thereby spreading misBy having the system determines the validity information (Fr) (harm). On the other hand, Table of the crime methods, making criminal meth- 2 lists risks associated with scenarios where the auods easier to commit crimes and lead to more tomated fact-checking system is attacked. For example, if an attacker inputs malicious data (AMI) sophisticated methods. (risk factor) or poisons the evidence information on 16. Proliferation of conventional and unconven- the web (AEP) (risk factor), the system may make a misjudgment (AMj) (Hazardous situation), and tional weapons then the attacker posts this result on social media, Entering instructions to investigate the it could lead to a chronological risk scenario where method of making weapons into fact-checking the attacker generates monetary income by gaintools could contribute to the proliferation of ing impressions on social media (AIm) (harm) and weapons. This could lead to human damage spreads disinformation (AFr) (harm). or legal violations. 17. Illegal surveillance and censorship
4.3
We don’t think this issue relates to automated fact-checking systems’ risks
Fault tree
The tables of risks on automated fact-checking, organized in Section 4.2 also show the chronologi18. Lack of transparency of training data cal relationships between risks, including risk facA lack of explanation regarding the evidence or tors, hazardous situations, and harm. Figure 1 ilresults may lead to misunderstandings among lustrates these chronological relationships as relausers. tionships in which cause and harm influence each other as a fault tree (FT)[19]. The risks on upper 19. Interactions with other systems nodes of Figure 1 arise when lower nodes are trigThis issue relates that defamation will be gered. For example, the risk of defamation arises posted on SNS when results of development when the layer of hazardous situation, namely, system differ from other fact-checking systems misjudgment, unprecise explanation, or prohibited/unfairness/inappropriate output are triggered, 20. No rights (copyrights or patents) for AI and the hazardous situations are triggered when No copyright can be claimed over the judgment one of the risk factors are triggered. In other words, results. the risks on the upper layer cannot be triggered unless one of the risks on the lower layer is triggered. 4.2 Taxonomy of Risks on Auto- This figure helps to understand the logical structure of the risks that connect the risk factors to the mated Fact-checking Systems harm via the hazardous situations. We summarize the risks of automated fact-checking systems considered in Section 4.1 in Tables 1 and 2. We organize risks chronologically into three stages– “Risk Factor,” “Hazardous situation,” and “Harm”– by referring ISO/IEC guide 51. Table 1 lists risks that may occur during normal use of an automated fact-checking system. For example, a possible chronological risk scenario is as follows: if the AI used in the system makes a misjudgment due to errors caused by any factors such as inaccuracies or hallucinations in AI (Mc) (risk factor), it results in a non-malicious misjudgment (Mj) (Haz- Figure 1: Fault tree on automated fact-checking ardous situation) of final judgment of the system; systems users who trust the result then post the incorrect 7
Phase of risk Risk factor
Risk factor Risk factor Risk factor Hazardous situation Hazardousn situation Hazardous situation Hazardous situation Harm Harm Harm
Table 1: Risks of automated fact-checking systems in normal use Risk name Explanation Misconstruction (Mc) Misconstruction of system (design, implementation, or errors caused by AI) (Insufficient training data, data bias, hallucinations, etc.) Input data (Ip) Risks from input data (entry of fabricated data, or prohibited data, etc.) External model (Em) Risks from external models or external data Normal operation (No) Risks in normal operations Misjudgment (Mj) Unintentional misjudgments or malfunctions Unprecise explanation (Ue) Prohibited / inappropriate output (Po) Unfairness output (Uf)
Harm Harm
Defamation (Dm) False rumors (Fr) Maximize reach on social media (Mr) Information disclosure (Id) Over confidence (Oc)
Harm Harm Harm Harm
Infringement of rights (Ir) Legal violation (Lv) Encouraging crime (Ec) Real-world damage (Rd)
Discrepancies, gaps, or lack of transparency in the explanation of evidence or results Generate prohibited or appropriate outputs Generate biased results (favoring specific groups or individuals) Leads to defamation Spreading Mis/disinformation Generating impressions for the purpose of ad revenue Information disclosure, privacy violations Believing results unquestioningly, over-confidence on the system Infringement, Plagiarism Violations of or encouragement of laws and regulations Encouraging crime Actual harm (physical, financial, material, or emotional damage)
5
Evaluation on actual auto- 5.2 DEFAME mated fact-checking system DEFAME is an open-source automated fact-
5.1
Overview
checking system[16]. Based on the characteristics of the target claim, this system uses an LLM to autonomously select and utilize multiple entities such as web search and image geolocation, to gather evidence for determining the truthfulness of the claim. After that, LLM analyzes this evidence and presents the fact-checking result with the evidence for its conclusion. DEFAME consists of six steps: Plan, Execute, Summarize, Develop, Judge, and Justify. When the data to be evaluated is input, the Plan step determines a strategy of which tools (Web Search, Image Search, Reverse Image Search, Geolocation, etc.) to use for fact-checking. Following the determined strategy, the Execute step collects evidence using the selected tools. Next, in
In this section, to verify the validity of the tables of risks summarized in Section 4.2, we assess the risks associated with an actual automated fact-checking system. We use each of the identified risks as a guide word. Specifically, we identify the risks on the Dynamic Evidence-based FAct-checking with Multimodal Experts (DEFAME)[16], one of the wellknown automated fact-checking systems. Similar to the risk assessment using STRIDE, we draw a Data Flow Diagram (DFD) and assess whether the risks summarized in Section 4.2 exist at each entity and data boundary on the DFD, thereby observing the validity of the risk taxonomy. 8
Table 2: Risks of automated fact-checking systems by adversarial attacks Risk name Explanation Adversarial Malicious Malicious input or entering prohibited input Input (AMI) Risk factor Adversarial Evidence Tampering with data used as evidence to make Poisoning (AEP) it malicious Risk factor Adversarial Attacks (AA) Launch adversarial attacks against AI systems to cause misjudgment or data leaks Hazardous Adversarial Leads to misjudgment situation Misjudgment (AMj) Hazardous Adversarial Unprecise Leads to misunderstand of explanation situation explanation (AUe) Hazardous Adversarial Prohibited / Leads to a prohibited or inappropriate output situation inappropriate output (APo) Hazardous Adversarial Unfairness Leads to a unfairness output situation output (AUf) Harm Adversarial Defamation (ADm) Leads to defamation Harm Adversarial False rumors (AFr) Spreading Mis/disinformation Harm Adversarial Maximize Generating impressions for ad revenue Harm reach on social media (AMr) Harm Adversarial Information Leads to information disclosure, or privacy Harm disclosure (AId) violations Harm Adversarial Infringement Leads to infringement, or plagiarism Harm of rights (AIr) Harm Adversarial Legal Leads to violations of laws and regulations or violation (ALv) encouraging such violations Harm Adversarial Encouraging Leads to the encouragement of crime Harm crime (AEc) Harm Adversarial Real-world Leads to actual harm (physical, financial, Harm damage (ARd) material, or emotional damage) Phase of risk Risk factor
9
the Summarize step, the results from each tool are summarized, and in the Develop step, the claim and evidence are integrated. In the Judge step, the system determines which category the claim belongs to, and depending on the situation, the process from first step may be repeated up to three times. In the Justify step, key findings and relevant evidence are summarized and presented in a format that is easy for humans to understand.
5.3
Risk Assessment Case Study
5.3.1
Assessment Procedure
the risks listed in Table 1 and Table 2 occur. The result of the risk identification is shown in Figure 3. In this figure, risks that occur in normal use (from Table 1) are shown in italics, and risks resulting from adversarial attacks (from Table 2) are underlined. Additionally, for comparison, the result of the STRIDE assessment is shown in bold. The meanings of each abbreviation are shown in Table 1, Table 2, and Section 3.1. Among the risks identified in this paper, those related to “harm” include cases that third parties are affected when DEFAME users post fact-checking results on social media. For this reason, we decided to illustrate the posting on social media and the harmed third parties of the posts in the lower-left corner of Figure 2, as shown in Figure 3.
1. Draw a data flow diagram (DFD) Draw a DFD for the assessment target automatic fact-checking system. Various types of evidence collection tools, judgment function, entities such as users, functions, tools, and boundaries of data exchange between entities should be depicted on DFD. The DFD we drew for DEFAME is shown in Figure 2. DEFAME judgment process is executed in the DEFAME pipeline box. Since the evidence collection executed in the execute step uses external tools, they should be placed outside of the system boundary (on the right side of Figure 2). The LLMs used by the DEFAME pipeline and external tools are generally developed externally vendors. So, they should be placed outside of the boundary. Similarly, the search engine executed by the external tools should also be placed outside of the boundary.
Figure 2: Data Flow Diagram (DFD) on DEFAME 2. Identifying risks at entity and data boundaries on the DFD For each entity and data boundary in the DFD created in the previous step, identify which of 10
5.4
Observation
Looking at the result of the case study conducted in Section 5.3, we found that among the risks categorized in this paper, the risk factors are located in the DEFAME pipeline, external tools, ML/LLM, and search engines. Furthermore, we found that hazardous situations exist in the DEFAME pipeline. This indicates that hazardous situation arises in the DEFAME pipeline after an incident occurs derived by risk factors that are in the system (including DEFAME pipeline, external tools, external ML models, external web search, or their boundaries.) Regarding actual harm, we found that they occur among DEFAME users, on social media if users post results on SNS, and among third parties by posting of the results. On the other hand, for STRIDE, we found that threats other than Repudiation occur at all entities and data boundaries except for users. This indicates that, since DEFAME is composed of general IT systems, general IT threats exist in the system. By this result, we can summarize that the risks cannot be identified by STRIDE such as the harm of the spread of defamation and mis/disinformation can be identified by using our taxonomy.
Figure 3: Result of risk assessment on DEFAME
6
Discussion
6.1
Risk Assessment using STRIDE 6.2 vs Ours
both aspects.
As shown in Section 5.4, the identified risks using STRIDE and ours are differ. The main reason is that the target risks of each guide word are different. The differences between the STRIDE guide words and our guide words are shown in Table 3. The target of STRIDE guide words is to identify general security risks on IT systems. In other words, identifying risks not associated with general security risks is difficult only using STRIDE guide words. On the other hand, our guide words’ target is to identify automated fact-checking systems specific risks. By only using our guide words, though we can identify automated fact-checking system specific risks, it is difficult to identify general security threats of the systems. For example, a threat that tampering system output label and posting it on social media corresponds to the “T: tampering” in STRIDE. Since this threat can be identified using STRIDE, it is not included in our guide words. By using both methods, it is possible to effectively assess risks of automated fact-checking systems for 11
Comprehensiveness of risks
In this paper, we have identified risks specific to automated fact-checking systems in relation to the risks of generative AI described in Tanaka et al.’s paper [13] and shown them in Table 1 and Table 2. In this mapping process, we also identified risks with reference to the Japan AISI’s guide [14] and the list of ML risks described in Section 3.2. Therefore, if there are any lacks in the risks of Tanaka et al.’s paper, there will also be lacks in the risks listed in this paper. If there are risks specific to automated fact-checking systems that are unrelated to generative AI, then the list of risks provided would be incomplete. However, we believe that the generative AI risks identified in Tanaka et al.’s paper sufficiently cover the risks of ML systems that utilize generative AI. Because the risks pointed out in Tanaka et al.’s paper also include the risks of the ML used to implement generative AI, and we consider that they broadly cover the risks of ML systems implemented using ML and generative AI. Even if there is no guarantee that Tanaka et al.’s
Target system Identified risks
Table 3: Difference of our risk taxonomy and STRIDE STRIDE Risk taxonomy of this paper General IT systems including Autometed fact-checking systems automated fact-checking systems General security threats Harm on users, and the risk factors and hazardous situations behind them
paper truly covers all risks associated with ML, including generative AI, the list of risks identified in this paper is still considered useful for developers and distributors of automated fact-checking systems when considering the risks specific of the systems they are developing or distributing.
example corresponds to a specific scenario in which disinformation such as “Yesterday’s earthquake was artificially triggered” is cleverly manipulated (by altering the input text) to be recognized as “true.” If the automated fact-checking system mistakenly returns a “true” result, the attacker uses this result as an endorsement to spread the disinformation, claiming, “The theory that yesterday’s earthquake 6.3 Magnitude of Risk was an artificial one is fact. Even fact-checking has This paper has identified various risks specific to confirmed it to be true.” We believe the tables of automated fact-checking systems. However, when risks in this paper serve as a useful reference when applying countermeasures, there are cases where deriving such scenarios. We would like to explore the magnitude of the risk must also be considered. methods as a future work that derive scenarios easIn general, the magnitude of risk is derived by mul- ier, for example, realize automation. tiplying the effect of risk and likelihood of risk [20]. Since the risks identified in this study are those in 6.5 Validity of the Case Study which the “harm” of the risk actually materializes, their effect can be derived by each type of harm. In this case study, we identified risks on the autoOn the other hand, the likelihood of risks cannot mated fact-checking system DEFAME. Even if risks be derived easily, because likelihood of each risk is were identified for a system other than DEFAME, derived by risk factors, but the likelihood of each the risks are expected to be similar to those idenrisk factor depends on the structure of system. So, tified in this case study. In other words, risks rederivation of magnitude is not easy. Deriving con- lated to risk factors and hazardous situations can crete methods of measures of magnitude is left as be identified in the system, external entities, and future work. their boundaries, and risks related to harm can be identified at users and social media platforms. Because we think that the structures of risks on the 6.4 Derivation of Risk Scenarios other systems are the almost same as DEFAME, Based on this study, we were able to derive the se- namely, risks arise at input data, internal AI malquence of events starting from the risk factors, lead- functions, or adversarial attacks, and the harm is ing to hazardous situations, and resulting in harm. suffered by users or third parties via social media. From the perspective of countermeasures against While this case study is qualitative and limited in these risks, deriving concrete risk scenarios can be scale, it demonstrates that the proposed taxonomy desirable. A risk scenario is a more concrete repre- can be directly applied to real-world automated sentation of a sequence of risks, such as when a dis- fact-checking system designs. information is intentionally entered into the system to cause a misjudgment (AMI), then the system 6.6 Countermeasures Against Risks generates an incorrect result (AMj) (The inputted disinformation is “TRUE” by the misjudgment), To address the risks associated with automated and the result is posted on social media, leading to fact-checking identified in this paper, the following the spread of inputted disinformation (AFr). This three approaches can be considered. Even though 12
these countermeasures are used, all risks may not be eliminated on the system, however, those risks are expected to significantly reduce by using them. From the perspective of developers and distributors, understanding the potential risks in their systems and taking countermeasures against them are important for developing and distributing safe and secure systems.
of these risks in user consent forms and contract terms is another potential countermeasure. This countermeasure is expected to encourage users to pay closer attention to data entry and the understanding of results, and to be more careful when posting results on social media. Additionally, if the system finds a user who is attacking the system, these consent and contract can be used as reason for banning that user.
1. Improving the accuracy of each entity’s judgment Improving the accuracy of the automated factchecking system itself, and that of the external tools helps reduce the likelihood of risks occurring. Therefore, from a developer’s perspective, improving accuracy serves as a risk mitigation. This measure is effective for both risk factors and hazardous situations lurking within the system. 2. Use of measures against adversarial attacks Risks in Table 2 are caused by adversarial attacks. Therefore, using countermeasures against adversarial attacks helps reduce the probability of risks occurring. For example, implementing attack detection functions, improving robustness through adversarial training, and using countermeasures against poisoning are all effective risk mitigation approaches. These measures help protect against both entering adversarial inputs and attacks for entities in the system. 3. Notification of risks to users and/or acceptance of risks under the contract This paper has identified the risks of automated fact-checking systems. We think that presenting these risks to users and ensuring they understand risks before using the system can help mitigate these risks. For example, simply stating, “Because this system uses AI, there is a possibility of misjudgments,” may reduce the risk of users blindly accepting incorrect results and spreading them on social media. This countermeasure can prevent the spread of mis/disinformation. The identified risks in Table 1 and 2 can be used for determining what kind of notifications are effective for users. Similarly, including a mention 13
7
Conclusion
In this paper, we categorized the risks on automated fact-checking systems. By utilizing existing risk assessment techniques and risk classifications, we categorized the risks considering the chronological sequence of risk occurring, namely, risk factors, hazardous situations, and harm. We identified four risk factors, four hazardous situations, and nine harm for normal use of the systems. For risks caused by adversarial attacks, we identified three risk factors, four hazardous situations, and eight harm. As a result, we found 32 types of risks totally associated with automated fact-checking systems. Using the identified risks as guide words, we conducted a case study to identify risks on DEFAME. As a result, we were able to identify risks on the DFD that could not be identified using STRIDE. We discussed the comprehensiveness of the risks, the magnitude of risks, derivation of risk scenarios, the validity of the case study, and countermeasures against the risks. The tables of risks identified in this paper are expected to be useful for developers and distributors of automated fact-checking systems when considering the risks of the systems they are developing and distributing. Since this case study was conducted virtually, to verify whether the identified risks actually manifest in real systems, and automated risk scenario derivation are marked as future works.
Acknowledgement Some of the findings in this paper are based on results obtained from a project, JPNP22007, commissioned by the New Energy and Industrial Technology Development Organization (NEDO).
References
[12] Baybars Örsek, “Opinion — Structural problems with tech platforms prevent fact-checkers [1] BBC, “The saga of ’Pizzagate’: The fake from focusing on harm and virality,” https: story that shows how conspiracy theo//www.poynter.org/commentary/2024/ ries spread,” https://www.bbc.com/news/ structural-problems-with-tech-platfor blogs-trending-38156985 ms-prevent-fact-checkers-from-focusin g-on-harm-and-virality/ [2] Japan Broadcasting Corporation (NHK), “Caution urged over Japan quake fake posts,” [13] H. Tanaka, M. Ide, J. Yajima, S. Onodera, K. https://www3.nhk.or.jp/nhkworld/en/ Munakata, N. Yoshioka, “Taxonomy of Gennews/backstories/4758/ erative AI Applications for Risk Assessment,” 2024 IEEE/ACM 3rd International Confer[3] PolitiFact, https://www.politifact.com/ ence on AI Engineering - Software Engineering for AI (CAIN 2024), 2024. [4] Full Fact, https://fullfact.org/ [5] Japan Fact-check Center, factcheckcenter.jp/
[14] Japan AI Safety Institute, “Guide to Evaluation Perspectives on AI Safety (version 1.10),” https://aisi.go.jp/assets/ pdf/ai_safety_eval_v1.10_en.pdf
https://www.
[6] Fujitsu limited press release, “ Fujitsu to combat fake news in collabo[15] Japan Today, “Man arrested for ration with leading Japanese organizaposting false tweet claiming lion on tions,” https://info.archives.global. the loose after Kumamoto quake,” fujitsu/global/about/resources/news/ https://japantoday.com/category/crime/ press-releases/2024/1016-01.html man-arrested-for-posting-false-tweetclaiming-lion-on-the-loose-after-kuma [7] Gordon Pennycook, David G. Rand, “The moto-quake Psychology of Fake News,” Trends in Cognitive Sciences.
[16] T. Graun, M. Rothermel, M. Rohrbach, A. Rohrbach, “ DEFAME: Dynamic [8] Microsoft Corporation, “Microsoft Threat Evidence-based FAct-checking with Modeling Tool, ” https://learn. Multimodal Experts,” arXiv, https: microsoft.com/ja-jp/azure/security/ //arxiv.org/abs/2412.10510 develop/threat-modeling-tool-threats [17] I-C. Chern, S. Chern, S. Chen, W. Yuan, K. [9] Preslav Nakov, David Corney, Maram Feng, C. Zhou, J. He, G. Neubig, P. Liu, “FacHasanain, Firoj Alam, Tamer Elsayed, Tool: Factuality Detection in Generative AI – Alberto Barrón-Cedeño, Paolo Papotti, A Tool Augmented Framework for Multi-Task Shaden Shaar, Giovanni Da San Martino, and Multi-Domain Scenarios,” arXiV, 2023, “Automated Fact-Checking for Assisthttps://arxiv.org/abs/2307.13528. ing Human Fact-Checkers,”arXiv,2021, https://arxiv.org/abs/2103.07769 [18] International Organization for Standardization (ISO), “ISO/IEC Guide 51: 2014, Safety [10] Zhijiang Guo, Michael Schlichtkrull, Andreas aspects – Guidelines for their inclusion in stanVlachos, ”A Survey on Automated Factdards, ” https://www.iso.org/standard/ Checking,” Transactions of the Association for 53940.html Computational Linguistics, Volume 10. [19] Phillip J. Brooke, Richard F. Paige, “Fault [11] Japan AI Safety Institute, “Known Attrees for security system design and analysis,” tacks and Their Impacts on AI Systems,” Computer & Security, Volume 22, Issue 3. https://aisi.go.jp/assets/pdf/Known_ Attacks_and_Their_Impacts_on_AI_ [20] International Organization for StandardizaSystems_V2_EN.pdf tion (ISO), “IEC Guide 31010:2019, Risk 14
management – Risk assessment techniques, ” https://www.iso.org/standard/72140. html
15