ConceptioArchivearXiv CS
arXiv CSopen access

Defeater Cards: Characterizing and Managing Safety Assurance Case Defeaters

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
softwarearchitecturesoftwareengineeringtesting
software engineering, software architecture, testing

arXiv:2606.11462v1 [cs.SE] 9 Jun 2026

Defeater Cards: Characterizing and Managing Safety Assurance Case Defeaters Usman Gohar Dept. of Computer Science Iowa State University Ames, Iowa

Michael C. Hunter Dept. of Computer Science Iowa State University Ames, Iowa

Salil Purandare Dept. of Computer Science Iowa State University Ames, Iowa

Jordan J. Rios Dept. of Computer Science Iowa State University Ames, Iowa

Myra B. Cohen Dept. of Computer Science Iowa State University Ames, Iowa

Robyn R. Lutz Dept. of Computer Science Iowa State University Ames, Iowa

Abstract—Safety assurance cases provide structured justifications that safety-critical systems meet their safety requirements. Recently, the notion of defeaters has emerged as a rigorous means of challenging the validity of safety arguments. Examples of defeaters might include overly strict claims, unreliable evidence, or reasoning gaps. However, defeaters remain ad hoc, lack structured support for critical reflection, are inconsistently described, are difficult to review, and lack documentation standards. To address this, we propose Defeater Cards, a new structured documentation artifact for systematically characterizing, reasoning about, and managing defeaters in safety cases. Drawing on a literature survey and thematic analysis, we identify documentation criteria that inform the card’s structure, based on the 5W1H framework. Defeater Cards are designed to support informed analysis and evolution, improve traceability and auditability, and enable the reuse of defeater knowledge across systems and product variants. We demonstrate their applicability through two crossdomain case studies, showing how they expose hidden assumptions, surface reasoning gaps, and support ongoing safety assurance case evolution. To support adoption and community reuse, we also release an open-source repository of defeater cards as a baseline upon which researchers and practitioners can build and describe lessons learned. Index Terms—Defeaters, Safety assurance case, Safety requirements identification, Safety requirements validation, Evolving systems

I. I NTRODUCTION Demonstrating the safety and reliability of safetycritical cyber-physical systems (CPS), such as autonomous vehicles and medical devices, is often a prerequisite for their deployment [1]–[3]. Safety assurance cases support this process by providing structured arguments that link safety requirements to supporting evidence, justifying that a system can perform safely in its intended environment [4]–[6]. They serve both as internal artifacts to guide development and as formal justifications reviewed by regulators for certification [7],

[8]. Over the past decades, extensive research has supported the construction of assurance cases, through formal notations such as Goal Structuring Notation (GSN) [9] and through the release of various tools [7], [8], [10]. In practice, however, assurance cases are susceptible to obstacles [11] to the soundness of their claims that the safety goals of the deployed product are adequately satisfied. These obstacles, termed defeaters, are any factors, conditions, or events that weaken or invalidate the safety claims made about the system [12]–[15]. Defeaters arise from infeasible safety requirements, incorrect assumptions, incomplete evidence, contextual gaps, or unforeseen system interactions. If left unaddressed, they compromise the trustworthiness of the assurance case, leading to unwarranted overconfidence in system safety and resulting in failures [16]. For example, consider an assurance case for a small Uncrewed Aircraft System (sUAS) with the requirement, “The sUAS can safely fly in current wind conditions”. This could be challenged by the defeater: “Unless unexpected wind shear or turbulence occurs at higher altitudes” [17]. To date, prior work has focused primarily on identifying and mitigating defeaters during initial assurance case development through semantic analysis, formal reasoning [18]–[21], and human-in-the-loop LLM approaches [22], [23]. However, defeaters emerge throughout the safety case lifecycle, including during regulatory audits and ongoing maintenance. In particular, audits play an essential role in (1) surfacing overlooked argument gaps, (2) validating assumptions, and (3) ensuring that emerging vulnerabilities are addressed before deployment [24]. The effectiveness of these stages heavily depends on process transparency, access to contextual information, and the rationale behind the defeaters. This aligns with our experience in our work on safety-critical systems, where project risk and cost grow when defeaters are in-

adequately documented, and the system or its operational context evolves [25]–[27]. Yet, current practices for documenting defeaters remain limited and fragmented across notations: processes are often ad hoc [7], [14], and defeaters are inconsistently tracked [4] and poorly documented [19], [28]. Recently, there have also been growing concerns about the rigor and validity of the defeater process [4], [29], with research emphasizing that documentation must ensure reproducibility and auditability to enable critical review [5], [30]. However, these principles are inadequately supported by current defeater documentation. Despite repeated calls, the lack of consensus on documentation best practices across notations further impedes adoption and standardization, a challenge echoed in adjacent domains for evaluating systems (e.g., AI safety [31], [32]). As cyber-physical systems increasingly integrate AI and operate in dynamic environments (e.g., drones), where assumptions and risks require continuous reevaluation [33], a systematic approach to identifying, reasoning about, and documenting defeaters becomes essential for safety management, certification, and maintenance. As a step toward that, we propose Defeater Cards: a lightweight, standardized framework for documenting defeaters across six evaluation dimensions, based on a literature review of current challenges and best practices. Inspired by documentation strategies in AI (Data Cards [34], Model Cards [35]) and Requirements Engineering (RE) (Snow Cards [36]), Defeater Cards explicitly capture contextual information and design rationale, support comprehensive analysis, informed review, and reuse across product variants, properties unsupported by current fragmented methods. Our approach is particularly well-suited to emerging modular certification efforts, such as certificates for sUAS, where safety assurance must be distributed across components (e.g., drone type, pilot, flying conditions) and reused in different contexts [17], [37]. To our knowledge, this is the first work addressing this gap in defeater documentation that currently impedes safety-case audits and long-term maintenance. Overall, this work makes four key contributions: • We identify 16 major documentation gaps and challenges, formalized as criteria, in current practice through a structured literature survey. • We propose an audit-friendly standardized artifact, the Defeater Card, to characterize and reason about defeaters. • Initial validation of our framework for two CPS case studies, and develop a dataset of defeaters and integrate them with an existing taxonomy to support use by researchers. The rest of the paper is organized as follows: In Section II, we provide motivation and related work. Section III outlines our design methodology and the

CX1: Environmental conditions, BVLOS parameters, multisensor DAA configuration.

G1: The sUAS can safely complete the BVLOS mission in the specified environmental conditions

G1.1: The specific sUAS has enough charge in its battery to complete the mission.

E1: Post-flight telemetry from n = 100 BVLOS missions.

D2: Unless system is not calibrated correctly

Goal

Strategy

S1: Argue over sUAS battery, environment and Detect and Avoid (DAA) subsystem. G1.2: The Detect-and-Avoid (DAA) subsystem reliably identifies and classifies all nearby airborne hazards during BVLOS flight D1: Unless sensor disagreement under partial occlusion or glare occurs leading to false or missed detections. D3: Unless the test missions did not cover rare mixed-visibility conditions

Context

Evidence

G1.3: The specified environmental conditions are within sUAS’s capabilities.

E2: Add soft angular limitsto the gimbal control system to prevent optical sensor alignment within critical solar angles Resolved

Defeater

Undeveloped

Fig. 1. sUAS Safety Assurance Case fragment with example defeaters

results of our survey, while Section IV introduces Defeater Cards. Section V discusses our two case studies. Finally, we discuss the implications of our work in Section VI and Section VII gives concluding remarks.

II. M OTIVATION & R ELATED W ORK A. Current Practices and Challenges Defeaters represent conditions, assumptions, or evidence gaps that challenge the validity of safety claims in an assurance case, playing a critical role in evaluating safety-critical systems [12], [13] Unlike Fault Tree Analysis, which identifies events that can contribute to hazards, or FMECA, which evaluates failure consequences, defeaters specifically target potential weaknesses in assurance arguments [6]. Figure 1 shows a fragment of an sUAS assurance case that we developed with example defeaters (red boxes). This fragment relates to new US Federal Aviation Administration regulations for flying Beyond Visual Line of Sight (BVLOS). While defeater analysis strengthens the robustness of assurance cases, its practical application remains limited by the lack of systematic methods for documenting how defeaters are identified, evaluated, and resolved [5], [28]. The lack of transparency creates further risks, particularly the potential to game the audits and project false confidence in the assurance case. One barrier to meaningful transparency is a lack of standards for reporting [38], [39], as the current landscape of defeater analysis is largely ad hoc and disjointed across different notations and frameworks [14], [40]. Finally, defeaters typically emerge through dialectic reasoning among developers, auditors, and domain experts, yet current documentation provides little structured support for capturing the rationale and contextual information needed. This often makes it hard to review and audit defeater-based safety assurance cases.

2

These challenges are amplified by a shift from static, controlled systems [33] to dynamic, autonomous, AIdriven systems operating in uncertain environments, where assurance cases may require ongoing updates, e.g., as new field data emerges [3], [22]. Current approaches provide inadequate support for ongoing audit and maintenance [4], [29], leaving reviewers without critical metadata on system capabilities, limitations, and the rationale for assumptions. Therefore, a more systematic and standardized approach to identifying and documenting defeaters can help practitioners anticipate issues throughout the product lifecycle. Defeater Cards attempt to close this gap by unifying notation and reporting key background information, design choices, and justifications.

What: Identification & Evaluation

Why? Justification & Rationale

When: Temporal and Lifecycle Relevance

Where: Context and Scope

Who: Stakeholders, Expertise & Contributors

How: Mitigation and Response

Fig. 2. Structural overview of the Defeater Card: Categorizing identification (5Ws) and mitigation (1H) through the 5W1H framework [59].

maintenance and audit, with standardized documentation as a critical first step. III. D EFEATER C ARDS F RAMEWORK D ESIGN To inform the structure and content of Defeater Cards, we first systematically identify documentation gaps and challenges through a literature review, distill them into high-level documentation criteria, and then operationalize these principles via the 5W1H framework. We primarily investigate two research questions: • RQ1: What defeater documentation gaps and best practices exist in safety assurance cases? • RQ2: Can the Defeater Cards framework be applied across diverse safety-critical domains? We follow the established Design Science Research methodology [60], [61]: (1) Problem identification (Section II), (2) Objectives through identifying gaps (RQ1), (3) Design & development by deriving criteria, and creating Defeater Cards, and (4-5) Demonstration and Evaluation through case studies (RQ2).

B. Related Work Documentation Frameworks. Structured documentation templates have supported transparency and reproducibility to address similar challenges in requirements engineering and adjacent domains. For instance, Snow Cards [36] standardize requirements elicitation, while Model Cards [35] and Data Cards [34] document the limitations and risks of AI models and the characteristics of datasets, respectively. In this work, we build on this line of efforts to introduce Defeater Cards for safety assurance cases. Evaluating and Assessing Assurance cases. Prior research has proposed tools [7], [41], [42] and methods [43]–[45] to assess assurance case validity, typically focusing on structural soundness [46] or quantifying confidence and uncertainty [47]–[49]. Other works explore whether LLMs could help review assurance cases [50], [51]. More recently, defeaters have gained traction as a complementary approach. Several works integrate defeaters into assurance frameworks [52], [53] or use semantic analysis [18]–[21] and LLMs [22], [23] to identify logical inconsistencies and latent challenges. However, these efforts focus primarily on identification during the initial assurance case development, with limited attention to documentation practices that support audit and maintenance throughout the lifecycle. Modern safety-critical systems, particularly those that are dynamic and AI-driven, require continuous reevaluation of assurance cases [33], [54], [55]. For instance, new evidence, shifting contexts, or invalidated assumptions can surface new defeaters or alter existing ones. Knowledge turnover and domain expertise loss compound these challenges, making it difficult to revisit earlier defeater reasoning. While continuous assurance frameworks [56], [57] monitor systems to trigger case updates, these require pre-specified triggers and do not support contextual changes [58]. In this work, we argue that defeater-based assurance should be viewed through the lens of lifecycle

A. Methodology. Identifying Relevant Studies: We employed systematic review methodologies [14], [67] to identify documentation challenges and gaps. We searched Google Scholar, ACM Digital Library, IEEE Xplore, and arXiv using the search string: ("assurance case" OR "safety case") AND ("defeater" OR "defeasible reasoning" OR "counter-argument" OR "assurance case weakeners" OR "deficits"), prioritizing papers on documentation, evaluation, or maintenance practices. We also performed backward and forward snowballing on seminal papers [4], [5], [40] to identify additional relevant work. To broaden our perspective, we also included a subset of documentation practices from AI evaluation and auditing [31], [38], software engineering design rationale [65], [65], and RE (Snow Cards [68]) that share conceptual parallels. Finally, we drew on our team’s experience building defeasible assurance cases. This process yielded N = 46 papers, which are available in the artifact. Thematic Analysis: We analyzed the literature using thematic analysis [69] and open coding [14], [70]. Two authors independently coded the data and reconciled

3

TABLE I TAXONOMY OF D EFEATER D OCUMENTATION C RITERIA DERIVED FROM LITERATURE AND CROSS - DOMAIN BEST PRACTICES , ORGANIZED BY THEME AND MAPPED TO D EFEATER C ARD DIMENSIONS . Documentation Criteria

Representative Sources

Defeater Card

Provenance (R1) Documents the defeater’s discovery source and methodology (e.g., red-teaming, formal analysis) (R2) Records the triggering event or condition that surfaced the defeater (R3) Links the defeater to the specific version and lifecycle phase in which it was identified

[4], [14], [31] [33], [62] [20], [35]

What What + When Metadata + When

Auditability (R4) Documents the background and expertise of contributors (human or automated) (R5) Maintains audit trails of contributor inputs and challenges

[31], [63], [64] [4], [30]

Who Metadata

Justification & Scope (R6) Provides rationale for identifying, prioritizing, or dismissing a potential defeater (R7) States underlying logical premises and assumptions supporting the defeater claim (R8) Assesses and document severity, likelihood, and safety-critical implications (R9) Defines the operational context and environmental boundaries in which the defeater applies (R10) Specifies the system component or subsystem affected by the defeater

[4], [19], [65] [19], [35], [52] [4], [31], [62] [19], [28], [35] [28], [52], [62]

Why Why Why Where Where

Mitigation Rigor & Robustness (R11) Specifies mitigation strategies with boundary conditions and limitations (R12) Documents monitoring arrangements and re-evaluation triggers for unresolved or evolving defeaters (R13) Records residual risks remaining after mitigation (R14) Reports the reliability and validity of proposed mitigations and supporting evidence

[4], [24], [38] [33], [54], [58] [28], [53], [62] [28], [38], [66]

How How How How

Dialectic Reasoning (R15) Preserves challenge-response threads, rebuttals, and counterfactual scenarios (R16) Tracks the resolution status and decision-making process across review iterations

[4], [13], [52] [4], [40], [62]

Why + Comments How + Comments

discrepancies to establish a shared set of themes. Using an expert-based affinity diagramming approach, we synthesized these themes with our prior experience building safety cases to translate abstract codes into formal documentation criteria. These criteria were then mapped to the 5W1H dimensions [59] (shown in Figure 2) and iteratively refined through application to two example systems.

B. RQ1: Documented Criteria, Gaps and Best Practices Our analysis identified five central themes that serve as documentation criteria for Defeater Cards, summarized in Table I. These can also be interpreted as normative requirements for Defeater Cards. Each of these has been mapped to the specific defeater card dimension. 1) Provenance & Traceability: The credibility of a defeater depends on knowing precisely where it came from and how it was identified. Provenance and traceability establish this foundation by linking each defeater to its discovery process, triggering conditions, and the relevant system context. The discovery source and methodology (R1) indicate whether the identification was systematic or opportunistic, affecting completeness and reproducibility [14], [31]. Recording the triggering event or condition (R2) provides the rationale for why the defeater surfaced at a specific point [4], [33]. Linking the defeater to the version and lifecycle phase (R3) preserves traceability as the system evolves [33], [35]. Together, these help prevent the loss of critical audit and reasoning information. 2) Auditability: Existing works on defeaters predominantly focus on practitioners and initial defeater analysis of safety cases. An overlooked use case for defeaters is their use as an approach to review or audit an already developed safety case. However, current documentation practices do not adequately support auditability. This has been repeatedly reported as a critical shortcoming of current defeater analysis and documentation frameworks [4]. Similar concerns have been raised in analogous processes such as red-teaming [31] and AI documenta-

Threats to Validity: There are several threats to validity. First, the documentation criteria vary in granularity. Some are specific, reflecting well-established practices, while others are broader to allow flexibility across domains, contexts, and use cases. Nevertheless, the criteria is intended to inform relevant design choices rather than to provide an exhaustive list. Second, the literature on defeater documentation practices is still nascent and limited, which constrains the size of the study. We mitigate this by incorporating established principles from adjacent domains and frameworks, grounding criteria in empirically validated practices. Third, the Defeater Card reflects current documentation gaps and practices, as well as our subjective analysis of challenges. Future evaluations of defeater cards and new safety assurance standards may surface considerations not yet anticipated. The modular and flexible design is intended to support extensions. Finally, due to space limitations, only two case studies are discussed; broader empirical validation and user studies remain future work. We view the case studies as demonstrating initial feasibility and applicability, similar to prior work [35], [60].

4

tion frameworks [30], [38]. In contrast to provenance, auditability concerns the process surrounding it, who conducted it, their expertise (R4), and whether the defeater process can be independently reviewed and verified (R5). This directly affects the credibility of the raised defeaters, as well as providing important signals of coverage and rigor [14]. For instance, model documentation literature [30], [35] identifies contributor expertise as critical for surfacing potential bias in public-facing systems, a concern echoed in AI fairness research [63], [64]. Therefore, standardized defeater documentation should provide enough contextual information to support review and audit. 3) Justification & Scope: A key feature of any robust and credible evaluation is providing justifications for any design and methodological choices. However, current practices in defeater analysis often leaves these justifications implicit, relying only on subjective interpretation. To ensure logical rigor and prevent frivolous challenges, it is imperative that defeaters are accompanied by a clear rationale. This justification should be grounded in scientific evidence, such as the experimental proxies and search coverage [38] used to validate the failure, or standardized auditing, mapping the challenge to specific regulatory or ethical benchmarks [63], [71]. Assurance 2.0 [19], [52] argues that a challenge without an explicit rationale (R6) and underlying premises and assumptions (R7) cannot be meaningfully evaluated or rebutted and hence, is simply unverifiable. Similarly, it requires documentation of severity and implications (R8) to contextualize why a defeater was considered significant and to avoid defeater gaming. Next, the operational context and boundaries to which the defeater applies (R9) are equally critical. Rushby [19] argues that a challenge without a defined scope is not falsifiable, a position consistent with operational design domain frameworks in autonomous systems [28]. Finally, identifying the affected component (R10) is essential for localized, targeted mitigation. This prevents scoping errors, where a failure is either ignored or incorrectly over-generalized, ensuring the defeater is evaluated only within its valid technical boundaries. 4) Mitigation Rigor & Robustness: In defeater analysis, documenting the rigor of mitigation is critical: evidence for resolving a defeater must clearly indicate how the mitigation is complete and robust. Safety case guidance [4], [62] stresses that mitigation strategies should include explicit boundary conditions and limitations (R11), particularly in modern AI-driven systems where mitigation approaches are often probabilistic or algorithmic [31]. To support safety assurance case maintenance, monitoring arrangements and re-evaluation triggers (R12) (if available) should be included, as systems and contexts evolve [72]. Next, regulatory standards and frameworks often require that residual risks (R13) be

recorded. Finally, a meta-evaluation of the mitigation is helpful for addressing the growing concerns about the lack of detail regarding its reliability and validity (R14). 5) Dialectic Reasoning: This is a central theme in Assurance 2.0 [5] and involves evaluating claims through structured consideration of opposing perspectives [73]. Multiple studies have called for its inclusion in assurance documentation [5], [29], yet fragmented practices across notations (e.g., GSN, EA [9], [13]) often prevent systematic adoption. Critically, much of the insight from earlier reviews is lost if challenge-response threads, rebuttals, and counterfactual scenarios (R15) are not preserved [4]. Rejected arguments can be as informative as accepted ones, providing essential audit evidence that the final artifact alone cannot convey. Equally important is tracking resolution status and decision-making across review iterations (R16) [62], ensuring that the evolution of the argument remains visible and contestable. Our 5W1H-based framework provides a structured method that directly enables practitioners to critically reflect on defeaters. IV. D EFEATER C ARDS : R EPRESENTATION AND S TRUCTURE In this section, we present the six dimensions of our framework for defeater analysis and documentation. Like prior efforts [35], [64], they are not exhaustive; rather, they are designed to guide critical reflection and capture relevant details for stakeholders across the defeater lifecycle. A. Metadata The metadata section captures basic information such as unique identifiers, providing foundational context for versioning, traceability, and cross-referencing within and across systems. Most fields can be auto-filled via tool support, reducing documentation overhead and errors. Defeater ID: A unique identifier assigned to each defeater for tracking and reference. Defeater Tag: Keywords or labels categorizing the defeater (e.g., data bias, invalid assumption) to facilitate filtering and search, typically internal to organizations. Defeater Type: The high-level classification of the defeater, such as logical, evidence-based, or contextual, based on established taxonomies [14], [74]. This is particularly useful for stakeholders as it supports a systematic approach to defeaters. Version: Which version of the defeater is this? And how has it evolved? This is critical for reviewers to understand the lifecycle of defeaters [29]. Date: When the defeater was identified, allowing stakeholders to assess whether original assumptions remain valid as capabilities and environments evolve [35]. Status: The current state of the defeater (e.g., open, mitigated, residual), as not all can be mitigated [13].

5

non-obvious propagation paths. Defeater Description: A concise explanation of the defeater, summarizing the potential issue, failure, or counterargument it represents. Source: What is the source of the defeater? For example, was the defeater identified via expert review, incident reporting (e.g., Drones incident database, AI incident database), or tooling? Documenting the source allows reviewers to assess the challenge’s credibility and limitations as different sources have distinct biases, coverage gaps, and reliability profiles that affect defeater validity [4]. Understanding the source helps evaluate whether the analysis relied on diverse perspectives or systematic methods, and supports identification of broader patterns across systems [28]. This also helps distinguish between empirically grounded and hypothetical defeaters.

Defeater Card Metadata: • Defeater ID | Defeater tag | Defeater Type | Version | Date | Status What? (Identification and Evaluation) | What is being evaluated? (e.g., what argument, evidence, etc.) Affected Claim/Evidence Node Defeater Description • Source Why? (Justification and Rationale) | Why is this important to the validity or credibility of the argument or evidence? • •

Rationale Underlying Assumptions • Severity/Potential Impact Who? (Stakeholders, Expertise & Contributors) | Who are the auditors/reviewers? (e.g., expertise coverage, teams etc.) • •

Automated Human (Individual or Team) – Expertise/Background When? (Temporal & Lifecycle Relevance) | When does the defeater emerge? (e.g., after deployment, under stress conditions, etc.) • Event/Trigger • Defeater Frequency • Phase/Lifecycle Where? (Context and Scope) | Where within the system or operational context does this defeater apply or manifest? • •

C. The Why? (Justification and Rationale) Assurance case evaluation is often undermined by superficial arguments and "confidence-inflating" defeaters [14], [75], a trend mirrored in ML research [76], [77]. Even quantitative confidence measures can be manipulated to create an illusion of rigor [75]. To counteract this, defeaters should be accompanied by well-justified rationales, and details on severity, impact, and inclusion criteria [4], [28], [29]. Such transparency is especially critical when a defeater is dismissed, as it allows reviewers to audit the underlying reasoning. Rationale: This field captures the explicit rationale and justification for including (or later dismissing) a defeater. Examples include unsupported inference steps, known limitations of the evaluation method, or mismatches between the evidence scope and the claim. Underlying Assumption: Unlike system-level assumptions, these are analytical or contextual premises used to interpret evidence and assess defeater plausibility. They encompass beliefs regarding data representativeness, evidence reliability, and operational conditions [35]. Explicitly documenting these assumptions ensures reasoning remains transparent, allowing reviewers to trigger a reassessment if the underlying context shifts. Severity/Potential Impact: Each defeater is characterized by its potential risk to the argument. Severity measures the degree of challenge to the reasoning (e.g., a minor gap vs. a critical failure), while Impact assesses consequences such as failure propagation or increased uncertainty. Likelihood distinguishes between theoretical concerns and probable weaknesses. Together, these provide a multidimensional profile for systematic evaluation [28], which can be captured via severitylikelihood matrices, scenario-based scoring [78], [79] or domain-specific scales [80].

System/Component Operational Context How? (Mitigation & Response) | How is this defeater detected, monitored, or mitigated? • •

Monitoring & Mitigation Reliability and Validity • Limitations (e.g., Sociotechnical, disagreements, project constraints) • Residual Risks Additional Comments: • Notes • •

Fig. 3. Proposed Defeater Card template with sections based on 5Ws, 1H framework to support dialectic reasoning.

B. The What? (Identification and Evaluation) The first question focuses on what is being evaluated, identifying which part of the assurance case is being challenged and providing the necessary context. Affected Node(s): This captures the assurance case node impacted by the defeater. A single defeater can affect multiple nodes and trigger secondary defeaters elsewhere in the argument [52], revealing hidden dependencies and interaction chains, and serving as a proxy for severity (see Why? dimension). In practice, practitioners may begin with the directly challenged node and expand iteratively; automated tools [23], [51] can assist in identifying

6

ensuring the safety case remains valid throughout the system’s life, not just at a single point in time.

D. The Who? (Stakeholders, Tools, and Expertise) To support audit and review, defeater cards should document who participated in the defeater process, including humans (individuals or teams) and automated methods (tools, LLM judges, etc.) [4], [38], [63] Human (Individual and Team): Documenting their domain expertise and experience enables reviewers to assess coverage, identify blind spots, and evaluate the reliability of their reasoning. In complex or human-AI systems, diverse perspectives are essential to surfacing edge cases and unusual scenarios [31]. To maintain privacy, this should focus strictly on professional background and expertise rather than personal identifiers. Automated: Identification increasingly relies on automated methods such as semantic analysis, formal methods, and LLM frameworks [20], [22], [51]. However, every method has inherent limitations, such as training biases in LLMs [81] or the specification dependency of formal methods, which can affect the credibility of the analysis [50]. Disclosing the specific tools used ensures transparency and provides reviewers with the necessary context to identify potential gaps and limitations.

F. The Where? (Context and Scope) This section captures the context in which a defeater emerges along two axes: the specific system component affected and the operational context and scope in which it emerges. By enforcing explicit scoping, the framework avoids selective reporting and supports comprehensive coverage analysis [14]. System/Component: Identifies the part of the system (hardware, software module, subsystem, or process) where the defeater is relevant, ensuring reviewers understand its locus of impact. Documenting this context localizes impact, accounts for context dependence, and reveals coverage gaps where critical components or scenarios remain unexamined [28], [82]. Operational Context: Many modern safety-critical system, such as autonomous vehicles, operate across heterogeneous and dynamic environments. This field specifies the environmental, configuration, or usage conditions under which the defeater occurs, clarifying scope and applicability. This is often relevant for runtime adaptive systems.

E. The When? (Temporal and Lifecycle Relevance) Defeaters manifest at different stages with varying temporal characteristics, from immediate risks to latent vulnerabilities triggered by specific conditions. This dimension maps "when" a defeater emerges, and is essential for designing monitoring strategies and maintaining assurance as systems evolve [33]. Event/Trigger: Event/Trigger: This field identifies conditions that activate a defeater, such as sensor drift, software updates, or environmental shifts. It also captures speculative or future-dated defeaters, such as anticipated regulatory changes. Current frameworks typically lack mechanisms for revisiting such speculative defeaters as conditions change [7]. Documenting both concrete triggers and latent conditions enables runtime monitoring, distinguishes between immediate and future risks, and ensures that speculative concerns are revisited as operational contexts shift. Defeater Frequency: This captures whether a defeater is one-time, periodic, continuous, or event-driven. Frequency affects resource allocation: high-frequency threats may require automated monitoring or streamlined review. In socio-technical contexts, frequency should be recorded alongside relevant observation windows. Phase/Lifecycle Stage: Specifies when a defeater is relevant (e.g., design vs. operation). This context determines the type of evidence available to address the challenge and identifies necessary regulatory checkpoints. Explicitly mapping lifecycle stages potentially reveals cross-phase dependencies, e.g., design assumptions that may transition into operational monitoring requirements,

G. The How? (Mitigation and Response) This dimension captures how defeaters are monitored, mitigated, or otherwise managed. Current frameworks often overlook (1) the limitations of proposed mitigations and (2) their reliability and validity [38]. Explicit documentation of these considerations captures developers’ insights into current or potential future remedies and helps reviewers assess their adequacy and rigor. Monitoring and Mitigation: This field primarily records the mitigation and monitoring strategies to handle the defeater. This can include discussion of computational methods, making changes to the design of the system, making changes to the operation of the system, e.g., by limiting the conditions under which the system is used, making changes to the assurance argument, e.g., adding an independent source of evidence or generating additional evidence for the confidence argument. In instances where mitigation is infeasible, e.g., due to epistemic constraints or conditional risks during runtime or residual risks, monitoring is required to flag violations and then follow up with relevant procedures [33]. Reliability and Validity: This field elicits from the defeater card user any analytic or experiential information or regarding the likelihood that the intended mitigation will perform failure-free as expected. It also encourages reasoning about whether the indicated mitigation for this defeater will perform as required, and how that might be confirmed or demonstrated.

7

Limitations (e.g., Sociotechnical, disagreements, project constraints) This field encourages the stakeholders to explicitly consider and report limitations of the proposed approach(es). These might also include disagreements between reviewers and/or the team, which are important signals in a dialectic framework [5]. Residual Risks: Not all defeaters can be mitigated completely, in which case residual risks remain. Acknowledging and highlighting these supports transparency, can prevent unforeseen failures, and enables appropriate monitoring plans to be put in place.

cards that apply to other parts of the sUAS system, e.g., configurations, battery, etc. B. Molecular Programming Our second example, presented in Figure 5, provides a defeater card for a partial assurance case of a molecular program, [84], with safety-critical applications such as targeted drug delivery [27]. In a molecular program, computational logic is encoded directly into molecules. As such, defeaters often arise from unmodeled biochemical interactions, inaccurate assumptions about molecular kinetics, or simulation-to-wet-lab discrepancies [85]. This case study highlights several observations. First, it demonstrates that the defeater card framework extends naturally to emerging computational domains, beyond software-intensive systems. Second, the Why? and Where? sections surface assumptions about molecular kinetics that were implicit in the original assurance argument but not explicitly recorded, underscoring how our framework functions as an elicitation tool, as well as a documentation artifact. More broadly, this example illustrates that defeater cards can surface gaps earlier in the lifecycle, before deployment, especially important in safety-critical domains [86].

H. Additional comments This field allows free-form discussion of additional concerns by the different stakeholders involved. V. RQ2: C ASE S TUDIES We now present two case studies that apply the Defeater Cards framework to diverse safety-critical domains. We chose two very different domains that have appeared in the safety case literature for the purpose of generality. These examples demonstrate the practical applicability of Defeater Cards and their support for transparent, standardized, and rigorous documentation of defeaters for safety case evaluation. The first is for an sUAS scenario, and the second is for an embedded nanodevice that uses molecular programming. We have also developed and made available an open-source repository with twelve additional defeater cards from seven safetycritical systems in diverse domains as a baseline for researchers and practitioners.

VI. L ESSONS LEARNED We describe below some lessons learned from our experience of developing, using, and reviewing defeater cards. We focus on those lessons that may be useful to projects wanting to use defeater cards and on lessons that point the way forward. Toward better understanding of defeaters. Developing the defeater card reveals limits to a developer’s knowledge of the operational context. We noted that several defeaters in the literature are too high-level to be useful, basically of the form “Unless it doesn’t work." The defeater card 5W+1H format encourages users to more thoroughly reason about and report how that breakage might happen, what effects it might have, and how much priority should be given to preparing for its occurrence. Toward improved maintenance of safety assurance cases in dynamic scenarios. The defeater cards’ organized clarity is especially beneficial when a person other than the originator needs to update the card, as when a potential defeater later is operationalized, triggering the need for a new safety requirement (often foreseen in the card’s mitigation field). We thus recommend, as an initial use case for defeater cards, those dynamic safety assurance cases where developers and maintainers are separate teams. Further, defeater cards can reduce the impact of personnel turnover-induced knowledge loss [87]. Additionally, because defeater cards describe potential future risks to satisfying the safety requirements, they help position a project to anticipate and adapt to likely changes.

A. Small Uncrewed Aerial Systems (sUAS) For our first example, we use the defeater D1 from the sUAS assurance case in Figure 1, which challenges the claim "G1.2: The Detect-and-Avoid (DAA) subsystem reliably identifies all nearby airborne hazards during BVLOS flight." This is presented in Figure 4. In contrast to the original defeater, we can see the contextual details that our framework provides. The Why? section reveals that the assurance argument assumes multi-sensor fusion provides redundancy, yet under sensor miscalibration, fusion amplifies correlated errors rather than suppressing them, undermining the very claim it was meant to support. The How? section documents not only the mitigation but its limitations: mitigation is unavailable on all airframes, and residual missed-intruder risk remains unresolved. Practitioners could, and have (e.g., GSN 3.0 [83]), extended the assurance case itself to capture such detail, but prior work has shown that this quickly becomes unwieldy and degrades readability [19], [28]. Moreover, our framework provides sufficient structure to support critical reflection in the first place. Our artifact contains additional defeater

8

Defeater Card Metadata Defeater ID Defeater Tag Defeater Type: Version: Date: BVLOS-DF-01 Sensor Fusion Confidence Evidence Validity 1.0 March 2025 What? (Identification and Evaluation) | What is being evaluated? (e.g., what argument, evidence, etc.)

Status: Open

Affected Node: Claim C3.2 – "The Detect-and-Avoid (DAA) subsystem reliably identifies and classifies all nearby airborne hazards during BVLOS flight" • Defeater Description: During BVLOS operations, the DAA system fuses inputs from ADS-B, radar, and optical sensors. Under certain conditions (e.g., partial occlusion, or radar clutter), sensor disagreement lowers fusion confidence, causing the system to alternate between “avoid” and “ignore,” leading to false or missed obstacle detections. • Source: Detected in post-flight analysis of simulated airspace trials and confirmed in real-flight telemetry during BVLOS corridor testing (Arizona desert, 2024). Why? (Justification and Rationale) | Why is this important to the validity or credibility of the argument or evidence? •

Rationale: The assurance argument presumes that multi-sensor fusion provides redundancy and thus improves reliability. However, when the calibration of individual sensors is compromised, fusion amplifies correlated errors. This defeater exposes an unacknowledged vulnerability: the confidence metric itself becomes an unreliable proxy for detection quality. • Underlying Assumptions: Sensor noise distributions are stationary across environments. • Severity/Potential Impact: Moderate-to-High. In two out of eight test flights, confidence dips led to delayed evasive maneuvers (>1.5 s latency), violating the 2-s detection-to-action requirement for BVLOS safety envelopes Who? (Stakeholders, Expertise & Contributors) | Who are the auditors/reviewers? (e.g., expertise coverage, teams etc.) •

Automated: N/A Human: Expertise/Background: Safety assurance, ML/AI Perception, BVLOS regulatory compliance. When? (Temporal & Lifecycle Relevance) | When does the defeater emerge? (e.g., under stress conditions, etc.) • •

Event/Trigger: Occurs when operating in mixed-visibility environments e.g., near-horizon glare or partial cloud cover. Defeater Frequency: Appeared in 15% of BVLOS test missions (n = 100). • Phase: Emerged in Operational Testing after certification prototype. Where? (Context and Scope) | Where within the system or operational context does this defeater apply or manifest? • •

System/Component: Detect-and-Avoid subsystem, Multi-Sensor Fusion Layer, Confidence Estimator module. Operational Context: Daytime BVLOS operations in mixed-visibility environments How? (Mitigation & Response) | How is this defeater detected, monitored, or mitigated? • •

Monitoring & Mitigation: Add soft angular limits to the gimbal control system to prevent optical sensor alignment within critical solar angles (calculated pre-flight) • Reliability & Validity: Statistical variance analysis across repeated flights (N = 50) under changing light conditions remained within 95% confidence bands. Cross-validation with manual labeling achieved Cohen’s K = 0.82 (strong agreement) for true conflicts. • Limitations: Not available on all airframes • Residual Risks: Residual missed-intruder risk remains and must be handled by conservative alerting and contingency procedures (e.g., quick descent, landing in place if possible). Multiple sensor degradation still a risk. •

Fig. 4. Example Defeater Card for an sUAS Assurance Case.

Enabling better reviews and audits. Reviews and audits of safety assurance cases can benefit from defeater cards’ scaffolding of needed information. We found that the defeater cards’ structure and guidance resulted in users writing cards that reviewers could read and understand. Contextual explanations and assumptions may be unknown or non-intuitive to the reviewer or auditor, and making these explicit both facilitates and enhances evaluation of the safety case. Balancing practicality and completeness. Structured documentation improves transparency, consistency, and traceability [87]. However, achieving comprehensive coverage requires significant time and resources, creating an inherent trade-off between completeness and feasibil-

ity. The appropriate balance will vary across domains, regulatory contexts, and organizational settings. Accordingly, defeater cards prioritize documenting critical contextual information without aiming for exhaustiveness, while remaining general enough for cross-domain use. Future study is needed to examine how stakeholders will interpret and apply defeater cards in practice. Residual risk of defeater hacking in practice. Despite the structured approach introduced by defeater cards, the risk of defeater hacking remains, i.e., strategically omitting, misclassifying, or selectively including defeaters to create an illusion of robustness [75]. Practioners may focus on low-impact issues while overlooking systemic concerns, reducing defeater analysis to a compliance

9

Defeater Card Metadata Defeater ID Defeater Tag Defeater Type: Version: Date: LC-02 DNA Temporal Logic Logical 3.0 Aug 2025 What? (Identification and Evaluation) | What is being evaluated? (e.g., what argument, evidence, etc.)

Status: Mitigated

Affected Node: G1.1 – "The molecular circuit reliably encodes input event order within specified kinetic tolerances" Defeater Description: The argument claims that DNA strand-displacement temporal logic circuits (TL-circuits) maintain reliable ordering of inputs under varying reaction intervals. However, deviations in toehold kinetics and input concentration shift output timing, causing temporal misalignment between expected and actual reaction order. Description based on [84]. • Source: Drift observed during laboratory experiments. Why? (Justification and Rationale) | Why is this important to the validity or credibility of the argument or evidence? • •

Rationale: The kinetic tolerances specified were derived from idealized conditions that do not account for concentration variability or toehold sensitivity observed. • Underlying Assumptions: Assumes that toehold shortening uniformly improves robustness without side-effects on false activation rates. • Severity/Potential Impact: High. Crosstalk directly violates the circuit’s temporal discrimination function. Who? (Stakeholders, Expertise & Contributors) | Who are the auditors/reviewers? (e.g., expertise coverage, teams etc.) •

Automated: Visual DSD simulator & MATLAB post-processing detected kinetic divergence. Human: Molecular computing researcher (circuit design), bioengineer (experimental validation), When? (Temporal & Lifecycle Relevance) | When does the defeater emerge? (e.g., under stress conditions, etc.) • •

Event/Trigger: Observed during phase two validation of two-input AND/OR circuits when input delay exceeded 20 min Defeater Frequency: Observed consistently with longer toeholds (6–7 nt) across three independent runs • Phase/Lifecycle: Verification and replication phase of molecular computing testbed Where? (Context and Scope) | Where within the system or operational context does this defeater apply or manifest? • •

System/Component: Logic gate strands Operational Context: DNA self-assembly How? (Mitigation & Response) | How is this defeater detected, monitored, or mitigated? • •

Monitoring & Mitigation: Used dual-channel fluorescence tracking to detect kinetic lag between logic modules. Optimized toehold lengths to 5–7 nt and normalized input concentrations (100–300 nM) to stabilize reaction rates. • Reliability and Validity: Re-evaluation at 10–30 min input intervals yielded consistent fluorescence peaks within ± 5% of predicted timing, demonstrating reproducibility under controlled conditions. • Limitations: Mitigation does not address ionic-strength variability or long-term storage effects; re-calibration required for field-deployable settings. • Residual Risks: Reaction fidelity remains sensitive to ambient temperature and reagent degradation; cumulative timing errors may compound in cascaded circuits; fluorescence readout may introduce bias in kinetic analysis. •

Additional Comments: • Notes: See [84] and its Supporting Information. Fig. 5. Example Defeater Card for a Molecular Program [84] Assurance Case.

exercise [16], [88]. Addressing this could involve formal completeness criteria, automated elicitation from logs and failure histories [33], and third-party assessments.

cases.

Toward structured documentation to report defeaters. The proposed framework provides practitioners and safety analysts with a structured approach to systematically document, analyze, and mitigate defeaters while supporting auditability. Contextual explanations and assumptions may be unknown or non-intuitive to the reviewer or auditor, and making these explicit supports evaluation of the safety case. As well, use of defeater cards may enhance LLM-based approaches [22], [23], [50], [51] by providing the contextual information that they require for automated evaluation of safety assurance

We introduce Defeater Cards, a structured framework for documenting defeaters in safety assurance cases. Through a systematic review of the literature and our experience developing and analyzing defeaters, we identified key challenges that informed the design of our framework. Defeater Cards are designed for use during assurance case development to standardize documentation of defeater reasoning. They proactively help capture and characterize potential obstacles to the soundness of the safety arguments. By structuring defeater reasoning, Defeater Cards improve transparency, support

VII. C ONCLUSION

10

updates and auditing, and strengthen the validity and rigor of safety arguments. We demonstrate applicability through two cross-domain case studies. We also release a public repository with templates and example cards across seven systems. We envision Defeater Cards as living artifacts that evolve through community adoption, domain-specific extensions, and empirical validation.

[19] J. Rushby, X. Xu, M. Rangarajan, and T. L. Weaver, “Understanding and evaluating assurance cases,” NASA, no. 20160000772, 2015. [20] F. U. Muram and M. A. Javed, “ATTEST: Automating the review and update of assurance case arguments,” Journal of Systems Architecture, vol. 134, p. 102781, 2023. [21] T. Yuan, S. Manandhar, T. Kelly, and S. Wells, “Automatically detecting fallacies in system safety arguments,” in Principles and Practice of Multi-Agent Systems: International Workshops: IWEC 2014, and CMNA XV and IWEC 2015. Springer, 2016, pp. 47– 59. [22] U. Gohar, M. C. Hunter, R. R. Lutz, and M. B. Cohen, “CoDefeater: Using LLMs to find defeaters in assurance cases,” in Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering, 2024, p. 2262–2267. [23] T. Viger, L. Murphy, S. Diemert, C. Menghi, J. Joyce, A. D. Sandro, and M. Chechik, “AI-Supported Eliminative Argumentation: Practical Experience Generating Defeaters to Increase Confidence in Assurance Cases,” in ISSRE. IEEE, 2024. [24] F. Königstorfer and S. Thalmann, “Ai documentation: A path to accountability,” Journal of Responsible Technology, vol. 11, p. 100043, 2022. [25] R. R. Lutz, A. Patterson-Hine, S. Nelson, C. R. Frost, D. Tal, and R. Harris, “Using obstacle analysis to identify contingency requirements on an unpiloted aerial vehicle,” Requir. Eng., vol. 12, no. 1, pp. 41–54, 2007. [26] R. R. Lutz, J. H. Lutz, J. I. Lathrop, T. H. Klinge, E. R. Henderson, D. Mathur, and D. A. Sheasha, “Engineering and verifying requirements for programmable self-assembling nanomachines,” in ICSE, 2012. [27] R. R. Lutz, “Requirements engineering for safety-critical molecular programs,” in RE. IEEE, 2022, pp. 302–308. [28] R. Bloomfield, G. Fletcher, H. Khlaaf, L. Hinde, and P. Ryan, “Safety case templates for autonomous systems,” arXiv preprint arXiv:2102.02625, 2021. [29] A. C. W. Group et al., “Assurance case guidance: Challenges, common issues and good practice, version 1,” Safety Critical Systems Club, 2021. [30] T. Puhlfürß, J. Butzke, and W. Maalej, “Model cards revisited: Bridging the gap between theory and practice for ethical ai requirements,” in IEEE RE, 2025, pp. 280–291. [31] M. Feffer, A. Sinha, W. H. Deng, Z. C. Lipton, and H. Heidari, “Red-teaming for generative AI: Silver bullet or security theater?” in AIES, 2024, pp. 421–437. [32] [Online]. Available: https://www.anthropic.com/news/ evaluating-ai-systems [33] E. Denney, G. Pai, and I. Habli, “Dynamic safety cases for through-life safety assurance,” in ICSE, vol. 2. IEEE, 2015, pp. 587–590. [34] M. Pushkarna, A. Zaldivar, and O. Kjartansson, “Data cards: Purposeful and transparent dataset documentation for responsible AI,” in ACM FAccT, 2022, pp. 1776–1826. [35] M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru, “Model cards for model reporting,” in ACM FAccT, 2019, pp. 220–229. [36] J. Robertson and S. Robertson, “Volere,” Requirements Specification Templates, 2000. [37] M. Hunter, U. Gohar, M. Cohen, R. Lutz, and J. Cleland-Huang, “A family-based approach to safety cases for controlled airspaces in small uncrewed aerial systems,” in AIAA AVIATION FORUM AND ASCEND 2024, 2024, p. 4626. [38] L. Staufer, M. Yang, A. Reuel, and S. Casper, “Audit cards: Contextualizing ai evaluations,” arXiv preprint arXiv:2504.13839, 2025. [39] P.-J. Courtois and D. L. Parnas, “Documentation for safety critical software,” in Proceedings of 1993 15th International Conference on Software Engineering. IEEE, 1993, pp. 315–323. [40] J. Rushby, “The interpretation and evaluation of assurance cases,” Comp. Science Laboratory, SRI International, Tech. Rep. SRICSL-15-01, 2015.

ACKNOWLEDGMENTS This work was funded by grant 80NSSC23M0058 from the National Aeronautics and Space Administration (NASA) and by NSF CCF 1900716 and CCF-2211589. R EFERENCES [1] RTCA, “Software considerations in airborne systems and equipment certification,” DO-178C, 2011. [2] R. Palin, D. Ward, I. Habli, and R. Rivett, “ISO 26262 safety cases: Compliance and assurance,” 2011. [3] P. Koopman, U. Ferrell, F. Fratrik, and M. Wagner, “A safety standard approach for fully autonomous vehicles,” in SAFECOMP. Springer, 2019, pp. 326–332. [4] Assurance Case Working Group, “Assurance case guidance,” ACWG, Tech. Rep. SCSC-159, 2021. [5] R. Bloomfield and J. Rushby, “Assurance 2.0: A manifesto,” arXiv preprint arXiv:2004.10474, 2020. [6] J. Knight, Fundamentals of Dependable Computing for Software Engineers. CRC Press, 2012. [7] M. Maksimov, S. Kokaly, and M. Chechik, “A survey of toolsupported assurance case assessment techniques,” ACM Computing Surveys (CSUR), vol. 52, no. 5, pp. 1–34, 2019. [8] E. Denney, G. Pai, and J. Pohl, “AdvoCATE: An assurance case automation toolset,” in SAFECOMP. Springer, 2012, pp. 8–21. [9] T. Kelly and R. Weaver, “The goal structuring notation–a safety argument notation,” in DSN workshop on assurance cases. Citeseer, 2004. [10] S. Varadarajan, R. Bloomfield, J. Rushby, G. Gupta, A. Murugesan, R. Stroud, K. Netkachova, and I. H. Wong, “Clarissa: Foundations, tools & automation for assurance cases,” in DASC, 2023, pp. 1–10. [11] A. van Lamsweerde and E. Letier, “Handling obstacles in goaloriented requirements engineering,” IEEE TSE, vol. 26, pp. 978– 1005, 2000. [12] E. Denney, G. Pai, and I. Habli, “Towards measurement of confidence in safety cases,” in ESEM. IEEE, 2011, pp. 380–383. [13] J. B. Goodenough, C. B. Weinstock, and A. Z. Klein, “Eliminative induction: A basis for arguing system confidence,” in ICSE, 2013, pp. 1161–1164. [14] U. Gohar, M. C. Hunter, M. B. Cohen, and R. R. Lutz, “A taxonomy of real-world defeaters in safety assurance cases,” in 2025 IEEE/ACM Workshop on Multi-disciplinary, Open, and RElevant Requirements Engineering (MO2RE). IEEE Computer Society, Apr. 2025, pp. 3–9. [15] E. Letier and A. van Lamsweerde, “Obstacle analysis in requirements engineering: Retrospective and emerging challenges,” IEEE Trans. Software Eng., vol. 51, no. 3, pp. 795–801, 2025. [16] C. Cave, “An independent review into the broader issues surrounding the loss of the RAF Nimrod MR2 Aircraft XV230 in Afghanistan in 2006,” The Stationary Office, Tech. Rep, 2006. [17] M. C. Hunter, U. Gohar, S. Purandare, K. Kjeer, R. Lutz, B. A. Duncan, J. Cleland-Huang, and M. Cohen, “SafeCert: Towards automated safety-case generation for risk assessment in small uncrewed aerial vehicles,” in AIAA AVIATION FORUM AND ASCEND 2025, 2025, p. 3169. [18] A. Murugesan, I. H. Wong, R. J. Stroud, J. Arias, E. Salazar, G. Gupta, R. Bloomfield, S. Varadarajan, and J. Rushby, “Semantic analysis of assurance cases using s(CASP),” in ICLP Workshops, 2023.

11

[41] E. Denney and G. Pai, “Tool support for assurance case development,” ASE, vol. 25, no. 3, pp. 435–499, 2018. [42] A. Di Sandro, G. Selim, R. Salay, T. Viger, M. Chechik, and S. Kokaly, “Mmint-a 2.0: tool support for the lifecycle of modelbased safety artifacts,” in ACM/IEEE MODELS, 2020, pp. 1–5. [43] L. Cyra and J. Gorski, “Support for argument structures review and assessment,” Reliability Engineering & System Safety, vol. 96, no. 1, pp. 26–37, 2011. [44] P. Mayo, “Structured safety case evaluation: a systematic approach to safety case review,” in International Conference on System Safety. IET, 2006, pp. 10–pp. [45] R. K. Panesar-Walawege, M. Sabetzadeh, L. Briand, and T. Coq, “Characterizing the chain of evidence for software safety cases: A conceptual model based on the IEC 61508 standard,” in ICST. IEEE, 2010, pp. 335–344. [46] T. Chowdhury, A. Wassyng, R. F. Paige, and M. Lawford, “Systematic evaluation of (safety) assurance cases,” in SAFECOMP Springer, 2020, pp. 18–33. [47] L. Duan, S. Rayadurgam, M. P. Heimdahl, A. Ayoub, O. Sokolsky, and I. Lee, “Reasoning about confidence and uncertainty in assurance cases: A survey,” in SEHC, FHIES 2014. Springer, 2017, pp. 64–80. [48] Y. Idmessaoud, D. Dubois, and J. Guiochet, “Confidence assessment in safety argument structure-quantitative vs. qualitative approaches,” IJAR, vol. 165, p. 109100, 2024. [49] J. B. Goodenough, C. B. Weinstock, and A. Z. Klein, “Toward a theory of assurance case confidence,” Pittsburgh, PA: Software Engineering Institute, Carnegie Mellon University, 2012. [50] M. Graydon and S. Lehman, “Examining proposed uses of llms to produce or assess assurance arguments, NASA/TM–20250001849,” Tech. Rep., 2025. [51] G. Yu, M. Sivakumar, A. B. Belle, S. Ghari, S. Wang, and T. C. Lethbridge, “LLMs as judges: Toward the automatic review of GSN-compliant assurance cases,” arXiv preprint arXiv:2511.02203, 2025. [52] R. Bloomfield, K. Netkachova, and J. Rushby, “Defeaters and eliminative argumentation in Assurance 2.0,” arXiv preprint arXiv:2405.15800, 2024. [53] S. Diemert, J. Goodenough, J. Joyce, and C. Weinstock, “Incremental assurance through eliminative argumentation,” Journal of System Safety, vol. 58, no. 1, pp. 7–15, 2023. [54] A. Agrawal, S. Khoshmanesh, M. Vierhauser, M. Rahimi, J. Cleland-Huang, and R. Lutz, “Leveraging artifact trees to evolve and reuse safety cases,” in IEEE ICSE, 2019. [55] M. A. Javed, F. U. Muram, H. Hansson, S. Punnekkat, and H. Thane, “Towards dynamic safety assurance for industry 4.0,” Journal of Systems Architecture, vol. 114, p. 101914, 2021. [56] P. Schleiss, F. Carella, and I. Kurzidem, “Towards continuous safety assurance for autonomous systems,” in IEEE ICSRS, 2022, pp. 457–462. [57] O. Jaradat, I. Bate, and S. Punnekkat, “Facilitating the maintenance of safety cases,” in Current Trends in Reliability, Availability, Maintainability and Safety: An Industry Perspective. Springer, 2015, pp. 349–371. [58] C. Cârlan, B. Gallina, and L. Soima, “Safety case maintenance: a systematic literature review,” in SAFECOMP. Springer, 2021, pp. 115–129. [59] S. Waisbord, “The 5Ws and 1H of digital journalism,” in Definitions of Digital Journalism (Studies). Routledge, 2021, pp. 38–45. [60] M. Ceci, D. Bianculli, and L. C. Briand, “Defining a model for content requirements from the law: An experience report,” in RE. IEEE, 2024, pp. 18–30. [61] K. Peffers, T. Tuunanen, M. A. Rothenberger, and S. Chatterjee, “A design science research methodology for information systems research,” Journal of management information systems, vol. 24, no. 3, pp. 45–77, 2007. [62] C. Hobbs, S. Diemert, and J. Joyce, “Driving the development process from the safety case,” Safety-Critical Systems Club, 2024.

[63] S. Costanza-Chock, I. D. Raji, and J. Buolamwini, “Who audits the auditors? recommendations from a field scan of the algorithmic auditing ecosystem,” in ACM FAccT, 2022, pp. 1571–1583. [64] T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. D. Iii, and K. Crawford, “Datasheets for datasets,” Communications of the ACM, vol. 64, no. 12, pp. 86–92, 2021. [65] J. E. Burge and D. C. Brown, “Rationale-based support for software maintenance,” in Rationale management in software engineering. Springer, 2006, pp. 273–296. [66] C. M. Holloway and P. J. Graydon, “Explicate ’78: Assurance case applicability to digital systems,” Federal Aviation Administration, Atlantic City International Airport, NJ, Technical Report DOT/FAA/TC-17/67, January 2018. [67] S. Keele et al., “Guidelines for performing systematic literature reviews in software engineering,” Technical report, ver. 2.3 EBSE technical report, Tech. Rep., 2007. [68] S. Robertson and J. Robertson, Mastering the Requirements Process: Getting Requirements Right, 3rd ed. Addison-Wesley Professional, 2012. [69] V. Clarke and V. Braun, Thematic Analysis. Dordrecht: Springer Netherlands, 2014, pp. 6626–6628. [70] M. Vaismoradi, H. Turunen, and T. Bondas, “Content analysis and thematic analysis: Implications for conducting a qualitative descriptive study,” Nursing & health sciences, vol. 15, no. 3, pp. 398–405, 2013. [71] A. Birhane, R. Steed, V. Ojewale, B. Vecchione, and I. D. Raji, “Ai auditing: The broken bus on the road to ai accountability,” in IEEE SaTML, 2024, pp. 612–643. [72] N. AI, “Artificial intelligence risk management framework (ai rmf 1.0),” URL: https://nvlpubs. nist. gov/nistpubs/ai/nist. ai, pp. 100–1, 2023. [73] D. N. Walton, The new dialectic: Conversational contexts of argument. University of Toronto Press, 1998. [74] W. S. Greenwell, J. C. Knight, C. M. Holloway, and J. J. Pease, “A taxonomy of fallacies in system safety arguments,” in 24th ISSC, 2006. [75] P. J. Graydon and C. M. Holloway, “An investigation of proposed techniques for quantifying confidence in assurance arguments,” Safety science, vol. 92, pp. 53–65, 2017. [76] M. L. Head, L. Holman, R. Lanfear, A. T. Kahn, and M. D. Jennions, “The extent and consequences of p-hacking in science,” PLoS biology, vol. 13, no. 3, p. e1002106, 2015. [77] P. Ganesh, U. Gohar, L. Cheng, and G. Farnadi, “Different horses for different courses: Comparing bias mitigation algorithms in ml,” in Workshop on Algorithmic Fairness Through the Lens of Metrics and Evaluation. PMLR, 2025, pp. 96–118. [78] C. Novelli, F. Casolari, A. Rotolo, M. Taddeo, and L. Floridi, “AI risk assessment: a scenario-based, proportional methodology for the ai act,” Digital Society, vol. 3, no. 1, p. 13, 2024. [79] L. Koessler and J. Schuett, “Risk assessment at AGI companies: A review of popular risk assessment techniques from other safetycritical industries,” arXiv preprint arXiv:2307.08823, 2023. [80] H. Dezfuli, A. Benjamin, C. Everett, G. Maggio, M. Stamatelatos, R. Youngblood, S. Guarro, P. Rutledge, J. Sherrard, C. Smith et al., “NASA risk management handbook,” Tech. Rep., 2011. [81] P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y. Cao, L. Kong, Q. Liu, T. Liu et al., “Large language models are not fair evaluators,” in ACL, 2024, pp. 9440–9450. [82] R. Hawkins, I. Habli, D. Kolovos, R. Paige, and T. Kelly, “Weaving an assurance case from design: a model-based approach,” in 2015 IEEE HASE. IEEE, 2015, pp. 110–117. [83] SCSC Assurance Case Working Group, “Goal structuring notation community standard version 3,” Safety-Critical Systems Club, CA, USA, GSN Community Standard SCSC-141C, 2021. [84] A. P. Lapteva, N. Sarraf, and L. Qian, “DNA strand-displacement temporal logic circuits,” Journal of the American Chemical Society, vol. 144, no. 27, pp. 12 443–12 449, 2022. [85] T. Tun, R. Lutz, B. Nakayama, Y. Yu, D. Mathur, and B. Nuseibeh, “The role of environmental assumptions in failures of dna nanosystems,” in First International Workshop on Complex Faults

12

and Failures in Large Software Systems (COUFLESS). IEEE Press, 2015, p. 27–33. [86] N. G. Leveson, Safeware: System Safety and Computers. Addison-Wesley, 1995. [87] E. Aghajani, C. Nagy, M. Linares-Vásquez, L. Moreno, G. Bavota, M. Lanza, and D. C. Shepherd, “Software documentation: the practitioners’ perspective,” in ICSE, 2020, pp. 590–601. [88] H. Pasman and W. Rogers, “How trustworthy are risk assessment results, and what can be done about the uncertainties they are plagued with?” Journal of Loss Prevention in the Process Industries, vol. 55, pp. 162–177, 2018.

13

Related documents

Record · ID 271933 · SHA-256 1297d1bbca6158d0
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.