Conceptio › Archive › arXiv CS
arXiv CSopen access

Beyond Fixed Fault Models: Comparing LLM-Based and Rule-Based Fault Injection in OpenStack

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

Beyond Fixed Fault Models: Comparing LLM-Based and Rule-Based Fault Injection in OpenStack Giuseppe De Rosa∗† , Pietro Liguori∗ , Domenico Cotroneo‡ ∗ University of Naples Federico II, Naples, Italy † IMT School for Advanced Studies Lucca, Lucca, Italy

‡ University of North Carolina at Charlotte, Charlotte, NC, USA

arXiv:2609.08681v1 [cs.SE] 8 Sep 2026

{giuseppe.derosa20, pietro.liguori}@unina.it, [email protected] Abstract—Software Fault Injection (SFI) supports testing of cloud systems by introducing software defects and observing their manifestation. Rule-based injectors such as ProFIPy provide controlled and reproducible source-level mutations but require fault patterns to be encoded manually. Large Language Models (LLMs) offer a data-driven alternative by generating context-dependent software faults. We compare two code LLMs, Qwen2.5-Coder and DeepSeek-Coder, with ProFIPy in OpenStack’s Nova and Cinder services. On shared injection targets, activation and observablefailure rates are comparable, but operational profiles differ: LLMgenerated faults produce more Catastrophic outcomes on Nova, whereas ProFIPy produces more Silent and Multi-component effects. The sampled LLM outputs also differ in how they manifest failure, while showing greater agreement in their propagation scope. These findings show that LLM-based fault injection extends the behavioral coverage of fixed fault models without establishing general superiority, and that practical adoption still requires controlled generation, runtime validation, system-level oracles, and reproducible experimental provenance. Index Terms—software fault injection, large language models, OpenStack, cloud, mutation testing

I. I NTRODUCTION Cloud platforms must remain available and preserve a consistent state despite software defects, partial failures, and complex interactions among distributed services. This is particularly challenging for Infrastructure-as-a-Service platforms such as OpenStack [1], where a single operation may involve APIs, message queues, databases, hypervisors, storage backends, and asynchronous workflows. A defect may remain local, propagate across services, or produce an apparently successful response while leaving the infrastructure in an inconsistent state. Software Fault Injection (SFI) evaluates robustness by deliberately introducing defects and observing whether the resulting errors are activated, detected, tolerated, contained, or exposed as failures [2], [3]. Source-level SFI targets residual faults: implementation defects that escape testing and code review, remain latent in deployed software, and surface only under specific execution conditions [4], [5]. Their effects are not always immediately visible. For example, an OpenStack volume-attachment request may return a successful API response even though the volume is unusable. Previous experiments detected such non-fail-stop failures only through independent state assertions [6]. Assessing cloud robustness, therefore, requires both representative faults and complementary oracles able to capture client-visible errors, internal failures, and inconsistent system state. The value of an SFI campaign largely depends on its fault model. Arbitrary syntactic mutations may not reflect defects

made by developers and may lead to misleading conclusions [2]. Rule-based injectors address this problem through catalogs derived from empirical bug studies and defect taxonomies [7], [8]. They encode recurring patterns, such as omitted calls, incorrect parameters, wrong return values, and missing control flow, as transformations over the program’s abstract syntax tree. ProFIPy applies this approach to OpenStack source code [6], [9], [10]. Its catalog provides valid, reproducible, and precisely controlled mutations. However, the explored fault space is bounded by manually designed operators, and context-dependent defects that do not match an existing template may remain unexplored. Code Large Language Models (LLMs) [11], [12] offer a generative alternative. After learning from historical residual bugs, they can use the surrounding function context to generate implementation-specific faults rather than instantiate fixed transformations. This may reduce operator-authoring effort and broaden the explored fault space, but generated outputs may be invalid, equivalent, off-target, or incompatible with the target runtime [13]. They therefore require explicit validation. Existing work on learned mutation and LLM-based fault generation has mainly examined code-level properties, such as syntactic validity, similarity to historical defects, and mutation score [14]–[16]. These properties do not establish whether an accepted fault exposes meaningful resilience weaknesses in a deployed distributed system. It remains unclear whether generative fault models reveal operational behaviors beyond those captured by fixed catalogs or primarily produce immediate, easily observable crashes. We investigate this question through a controlled comparison of generative and catalog-based source-level fault injection in OpenStack. Rather than asking whether an LLM can produce a syntactically different function, we examine whether its accepted faults extend the operational behavior space observed under a common deployment, workload, and set of failure oracles. Here, extension denotes additional behaviors observed in the evaluated campaign. We study the Nova compute and Cinder block-storage services of OpenStack Pike. Qwen2.5-Coder and DeepSeek-Coder are fine-tuned on PyResBugs [17] and compared with ProFIPy using the same testbed, workload, coverage measurements, and failure oracles [10]. Since validation produces different accepted target sets, we report both campaign-level results and a locationcontrolled analysis over functions shared by all three injectors. We address the following research questions:

RQ1 (Failure behavior). How do generative and rule-based Existing work, therefore, emphasizes syntax, similarity to historical changes, mutant survival, or mutation score. Those fault models differ in the failures they expose? • RQ2 (Failure propagation). How do generative and rule-based metrics characterize generation and testing utility, not the faults differ in their propagation across system components? operational consequences of activating a generated fault in a deployed system. • RQ3 (Generative diversity). How much does the choice of the generative model affect the operational behavior of This distinction separates our study from LLM mutation testing. injected faults? Our goal is to deploy each accepted mutation in a multi-service Our results show that the two approaches expose complemen- cloud and observe its effects across components and user-visible tary behaviors. On shared Nova targets, accepted LLM-generated state, rather than evaluate the system’s test suite. faults cause more service-level crashes, whereas ProFIPy To our knowledge, no study compares catalog-driven and produces more Silent failures, in which an operation appears generative source-level injection in one deployed distributed successful despite an inconsistent resource state. Throughout system under a common workload, deployment, and oracle the campaign, ProFIPy faults propagate more frequently among configuration. We fill this gap by deploying accepted mutations components, whereas LLM-generated effects are predominantly from models fine-tuned on residual Python bugs and testing local or result in immediate service disruption. Qwen and whether they extend the observed severity, visibility, and DeepSeek also produce different failure outcomes at the same propagation space of a fixed catalog. locations, although their system-level impact is more strongly III. M ETHODOLOGY shaped by the target service and injection point. Overall, This section details the methodology (Figure 1) underlying generative fault models broaden contextual exploration while our comparison. We evaluate the selected, fine-tuned LLMs and complementing the control of fixed operator catalogs. ProFIPy over a shared pool of target functions, using the same II. BACKGROUND & R ELATED W ORK testbed adopted to evaluate ProFIPy [10] as workload, and four A. Software Fault Injection failure oracles to classify the observed outcomes. Software-implemented fault injection operates at levels from A. Dataset processor state to source code and service interactions [3], [18]. PyResBugs [17] is the dataset used to fine-tune the evaluated Low-level corruption commonly models transient faults [19], whereas source transformations emulate persistent implementa- LLMs and generate the faults. It contains 5,007 hand-validated tion defects. We use the latter and observe their effects through an residual bugs mined from dozens of open-source Python projects. end-to-end workload. Source-level injection resembles mutation Each record pairs a fault-free function with the faulty version testing because both create modified program versions, but their that developers left in production, allowing models to learn real goals differ [3], [20]. Mutation testing assesses a test suite’s rather than synthetic defects. Its G-SWFIT [7] and ODC [8] ability to kill mutants, whereas dependability studies examine labels cover ProFIPy’s fault families: wrong parameters, omitted activation, detection, containment, reporting, and propagation. A calls, incorrect returns, and missing handlers. We train only on code pairs to teach the models how to useful testing mutant is therefore not necessarily representative generate software faults without any further natural-language of a residual production fault. Representativeness depends on the transformation, location, instructions. The model, therefore, learns a direct clean-to-faulty activation conditions, context, and failure behavior [2], [21], [22]. transformation. We verified that neither the target functions nor ODC classifies defect semantics [8], while G-SWFIT derives their bug-fixing commits occur in the fine-tuning set. The leakage check compares the normalized target source and operators such as omitted calls, wrong parameters, and missing control flow from field defects [7]. Catalogs offer control and associated commit identifiers with every fine-tuning record. A repeatability but may miss context-dependent faults; ProFIPy shared project or function name alone is not considered leakage because common names can occur in unrelated modules; a match makes them programmable [10]. Other approaches include distributed systems-based, that requires the same target code or fixing commit. instead inject runtime failures: FATE and DESTINI target storage B. Fault Generation and I/O [23]; Molly uses data lineage [24]; Filibuster explores We fine-tune Qwen2.5-Coder-32B-Instruct [11] and DeepSeekservice calls [25]; and Legolas selects locations using program states [26]. They primarily model exceptions, delays, failed calls, Coder-33B [12] with QLoRA [28], [29]. Both are codemessage loss, and storage faults rather than implementation specialized models of comparable size, so the comparison does defects. They complement the source-level faults examined here. not conflate generation strategy with a general-purpose model. Both use 4-bit NF4 weights, double quantization, bfloat16 B. Learned and Generative Fault Models computation, rank r = 16, LoRA α = 32, and dropout 0.05 Data-driven methods learn transformations from historical on the query, key, value, output, gate, up, and down projections, defects. Tufano et al. [16] reversed a model trained on bug trained for 2–3 epochs with early stopping on validation loss, fixes to generate defective code. Neural Fault Injection uses a batch size of 1 with 8-step accumulation, a learning rate of natural-language prompts [27]; PyResBugs pairs residual Python 2×10−4 , 4,096-token sequences, 8-bit paged AdamW, a cosine bugs with fixes and descriptions [17]; and LLMorpheus applies schedule with 3% warm-up, and seed 42. For each generation, each target code block receives the LLMs to JavaScript mutation testing [14]. These methods extend fixed catalogs but can produce invalid, identical instruction “Inject a single and realistic software fault equivalent, runtime-incompatible, or unrepresentative changes. into the function,” with a system message requiring syntactically •

LLM injector Qwen / DeepSeek function-context input, fine-tuned on residual bugs

Python 2.7 filter

OpenStack Pike testbed

Fault-free method (Nova / Cinder)

Nova • Cinder

Oracles & CRASH classification API • assertions logs • coverage

ProFIPy injector fixed, hand-written fault catalog

Fig. 1. Experimental harness. Qwen and DeepSeek generate candidate mutations, whereas ProFIPy applies a controlled operator. Accepted Python 2.7-compatible mutations run on the same testbed and are classified from four oracles.

valid Python, a changed body, and only the faulty function in one code block without explanation; keeping the prompt fixed across targets avoids introducing variation from prompt wording. Sampling uses nucleus sampling with temperature 0.4, top-p = 0.95, up to 1,024 new tokens, and up to three attempts per target. Each candidate then passes through two validation stages before being accepted. The first checks that the output is not empty or whitespace-equivalent and that its AST contains a function definition. This distinguishes a usable function from prose, an empty response, or an unchanged copy of the input. The second stage filters for Python 2.7 compatibility, since targets must run under OpenStack Pike. Neither stage verifies semantic equivalence, signature preservation, or whether the rewrite introduces exactly one independent semantic edit; they establish only that a candidate compiles and is deployable, leaving operational behavior to be determined later by the workload and oracles. The first candidate to pass both stages enters the campaign, and targets that fail all three attempts are dropped rather than resampled indefinitely. C. Measurement and Analysis OpenStack studies report inconsistent state, delayed reporting, non-fail-stop behavior, and cross-service propagation [6], [30]. A client-visible API response alone can therefore misrepresent the final resource state, while a local service log may miss effects that travel through the control plane. These findings motivate four oracles: API, state assertions, logs, and coverage. The API oracle records HTTP 4xx/5xx responses; in-guest assertions test functional state, including volume writes and reads; Syslog entries at WARNING or higher reveal internal errors; and coverage.py [31] records whether mutated code executes. The oracles are complementary: API responses capture client-visible failures but cannot reveal an unusable resource reported as created. Assertions expose internal silent failures, logs capture errors that are not propagated to clients, and coverage distinguishes activated faults. We used these oracles in sequence to identify the failure outcome. A fault may be dormant, i.e., the workload could never exercise it. Thus, coverage first determines whether the target executes (i.e., verifies the fault activation). If it does, the API, assertion, and log observations distinguish an absorbed fault from a manifest failure by examining, respectively, client-level and system-level states.

TABLE I CRASH MODES , CLOUD MANIFESTATIONS , AND DETECTING ORACLES . N O R ESTART OCCURRED . Class

Cloud manifestation

Oracle

Catastrophic

Service daemon crashes at startup; workload blocked Operation hangs and requires a restart (not observed) Operation fails immediately with an API error HTTP 2xx success, but the resulting resource is unusable or inconsistent Failure is detected late; an assertion fails before a subsequent API error Target is dormant, or activation yields no oracle-visible failure

Log, API

Restart Abort Silent Hindering No failure

Timeout API error Assertion Assertion, API Coverage/none

Outcomes are classified using the Koopman CRASH scale [32]: Catastrophic, Restart, Abort, Silent, and Hindering. We adapt these classes to cloud services by defining each in terms of its observable effect, since a distributed service can fail in ways a standalone program cannot (e.g., an operation that returns success while leaving persistent state inconsistent). Runs where the fault is never triggered (dormant) and runs where it is triggered but produces no oracle-visible failure (activated) both fall under No-failure; we report them separately in the results to distinguish untested code from code that tolerated the fault. Classification is based on the end-to-end manifestation observed during the workload. Concretely: a crash of a service daemon at start-up is Catastrophic; an operation rejected immediately with an API error is Abort; an HTTP 2xx response followed by an invalid or inconsistent resource is Silent; and a failure caught by a late assertion, followed by a subsequent API error, is Hindering. Restart is reserved for hangs or timeouts requiring recovery; this class did not occur in our campaign. Table I summarizes each class, its cloud-specific manifestation, and the oracle used to detect it. Finally, an independent dimension complements the CRASH scale. IMPACT records propagation: the number of components affected by the faulty run, other than the original component where the fault is injected. Total for a system-wide effect, Multi when an effect reaches another component, Local when an

TABLE II S CENARIO INVENTORY; ONE MUTATED TARGET PER SCENARIO . Service

ProFIPy

DeepSeek

Qwen

Total

Nova Cinder

23 20

21 17

22 19

66 56

Total

43

38

41

122

activated run remains confined to the injected component, and Dormant when the target is not executed. Local, therefore, includes contained runs with no observed failure. Keeping CRASH and IMPACT separate lets us differentiate severity and propagation. D. Experimental Design

TABLE III CRASH- MODE PERCENTAGES FOR THE FULL SUITES ( SIZES IN PARENTHESES ; PF: P RO FIP Y, DS: D EEP S EEK , Q W: Q WEN ). N O FAILURE COMBINES DORMANT AND ACTIVATED / NO - OBSERVED RUNS . Nova Failure mode Catastrophic Abort Silent Hindering No failure

Cinder

PF

DS

Qw

PF

DS

Qw

(23)

(21)

(22)

(20)

(17)

(19)

4.3 17.4 30.4 17.4 30.4

28.6 19.0 9.5 4.8 38.1

22.7 0.0 9.1 5.0 18.2 0.0 18.2 20.0 31.8 75.0

0.0 11.8 0.0 23.5 64.7

0.0 21.1 5.3 26.3 47.4

A. RQ1: How do generative and rule-based fault models differ in the failures they expose?

Table III and Figure 2 summarize the CRASH outcomes For comparability with ProFIPy, the testbed runs OpenStack Pike on CentOS 7 and Python 2.7. It comprises a Controller that produced by the three injectors. We first examine the full hosts RabbitMQ, MariaDB, Keystone, Glance, and the Nova and experimental suites and then compare only the functions Cinder control services; a Compute node runs nova-compute successfully targeted by all three injectors. Across the full suites, accepted LLM-generated mutations over KVM/QEMU. We inject into Nova compute and driver activate in 77/79 runs, compared with 39/43 ProFIPy mutations. managers and Cinder’s volume manager and driver, which implement the compute and block-storage operations exercised An observable failure occurs in 44/79 LLM runs and 21/43 by the workload. These services also use RabbitMQ and shared ProFIPy runs. These aggregate values describe the behavior of control-plane workflows, making them suitable for observing the respective experimental suites, but part of the difference may result from the functions selected by each injector rather than propagation. All suites run the workload [33] used in the ProFIPy study [6], from the fault-generation approach itself. Comparing the 37 functions shared by all three injectors [10] for comparability. Through REST APIs, it uploads CirrOS, reduces this target-selection effect. Activation is similarly high provisions a private network, router, and floating IP, boots a VM, and creates and attaches a Cinder volume. It checks SSH for ProFIPy (34/37), DeepSeek (36/37), and Qwen (37/37), reachability, reads and writes the attached block device, verifies while observable failures occur in 19/37, 18/37, and 22/37 runs, respectively. Thus, on common targets, the LLM-generated a hard reboot, and deletes all resources. A fault is activated when the workload reaches its injection faults do not show a substantial activation advantage. However, point and dormant otherwise; coverage records this relative to the types of failures remain different. Among the 20 shared Nova a fault-free run. Activation is not guaranteed: some faults may functions, ProFIPy produces one Catastrophic failure, compared introduce triggering conditions that the workload cannot meet. with six for DeepSeek and five for Qwen. Conversely, ProFIPy Also, we highlight that activation does not imply a failure. produces seven Silent failures, compared with two for DeepSeek OpenStack may catch an injected error, roll back an operation, and four for Qwen. A contrast in failure profiles, therefore, or return a controlled response without corrupting the observed persists even when the target functions are held constant. On Nova, the main difference concerns the severity and state. An activated, no-observed-failure run executes but leaves observability of the failures. DeepSeek and Qwen produce no trace in the API, assertion, or log oracles; it may be absorbed, substantially more Catastrophic outcomes than ProFIPy: 28.6% irrelevant to this workload, or semantically equivalent. The and 22.7%, respectively, compared with 4.3%. These faults often experiment does not distinguish these causes. A manifest failure prevent a service from completing its start-up sequence, making executes and is flagged by at least one oracle, after which it is the failure immediately visible. ProFIPy instead produces the classified on the CRASH severity scale. The campaign has 122 one-target scenarios: 66 on Nova and largest proportion of Silent failures, at 30.4%, compared with 56 on Cinder (Table II). Each deploys one mutated function in 9.5% for DeepSeek and 18.2% for Qwen. In these cases, an one source file. A fault-free checkout is performed after every operation reports success even though the resulting resource injection to avoid environmental drift or resources left by earlier is invalid or unusable. Such failures are particularly relevant runs; the idea is to clean the environment before continuing, so because the platform continues serving requests without exposing an explicit error. carryover is not attributed to the next mutation. Cinder exhibits a different failure profile. None of the injectors produces a Catastrophic outcome, and the distributions are IV. R ESULTS dominated by Hindering failures (20.0–26.3%) and runs with We compare the three injectors along three dimensions: the no observed failure (47.4–75.0%). Cinder’s exception-handling activation and failure behavior of their faults, the propagation of paths frequently contain the injected error, allowing the workload their effects across components, and the variability between the either to continue or to fail later through an explicit operationtwo generative models. level error. Consequently, the contrast between the injectors is

Key Takeaway: On the same target functions, LLMgenerated faults cause more visible crashes and workflowspecific failures, while ProFIPy reveals more silent corruptions. Differences in overall activation rates are partly due to the different functions targeted by each suite.

B. RQ2: How do generative and rule-based faults differ in their propagation across system components? Table IV shows that LLM-generated and rule-based faults differ not only in severity, but also in how far their effects spread. Accepted LLM-generated faults more often remain Local or produce a Total impact, whereas ProFIPy faults more often propagate to another component without causing a system-wide outage. This pattern is clearer on Nova. DeepSeek and Qwen produce Total effects in 28.6% and 22.7% of the runs, respectively, compared with 4.3% for ProFIPy. These outcomes largely correspond to start-up crashes that block the end-to-end workload. ProFIPy instead produces more Multi-component effects: 26.1%, compared with 9.5% for DeepSeek and 13.6% for Qwen.

100

75

% of suite

less pronounced than on Nova. The No failure category includes both faults that never activate and faults that activate without producing an observable effect. On Nova, the dormant faults account for four ProFIPy runs, one DeepSeek run, and one Qwen run; no dormant fault occurs on Cinder. After excluding these cases, the numbers of activated faults with no observed failure are 3, 7, and 6 on Nova, and 15, 11, and 9 on Cinder, for ProFIPy, DeepSeek, and Qwen, respectively. These activated faults may have been absorbed by exception handling, may affect behavior outside the exercised workload, or may be semantically equivalent to the original code. Silent failures are distinct from these cases because they both activate and produce an observable inconsistency in the state. Although the LLM-generated suites contain fewer Silent failures overall, they expose failures tied to specific end-to-end workflows. For example, one Qwen mutation causes a volume-attachment request to return an HTTP 2xx response even though the attached volume is unusable. This inconsistency is detected only by the subsequent in-guest write/read assertion. The differences between Nova and Cinder show that the manifestation of failure depends not only on the injected fault but also on the architecture of the target service. Nova exposes a broad severity range, including start-up crashes and silent resource corruption. Cinder more often contains exceptions or delays their effects, shifting the distribution toward Hindering outcomes and runs with no observed failure. Service architecture and exception handling, therefore, mediate the observable impact of all three fault models under the same workload and Oracle configuration. Overall, accepted LLM-generated faults broaden the observed failure profile by producing more Catastrophic and workflowspecific outcomes. ProFIPy exposes more Silent failures and propagated effects. Moreover, the similar activation rates on shared functions show that the higher aggregate activation of the LLM suites cannot be interpreted as a generator-only effect.

50

25

0

N-PF

N-DS

N-Qw

C-PF

C-DS

Catastrophic

Abort

Silent

Hindering

No failure

C-Qw

Fig. 2. CRASH distributions for the full suites (N: Nova, C: Cinder; PF: ProFIPy, DS: DeepSeek, Qw: Qwen).

Cinder shows the same tendency without Nova’s large number of Total outcomes. ProFIPy produces Multi-component effects in 25.0% of the runs, compared with 11.8% for DeepSeek and 5.3% for Qwen, while Local outcomes account for 75.0%, 88.2%, and 89.5%, respectively. Among manifest Cinder failures, all five ProFIPy cases affect another component. By contrast, four of six DeepSeek failures and eight of ten Qwen failures remain Local. These differences are consistent with the structure of the injected mutations. ProFIPy typically applies small rule-based changes, often limited to one expression or statement. Such mutations can alter a value, condition, or return path while preserving execution, allowing the resulting inconsistent state to travel through REST calls, RabbitMQ messages, or shared control-plane workflows. LLM-generated mutations are often broader and may span multiple statements or modify control flow more substantially. They are therefore more likely to directly disrupt the injected service, leaving less opportunity for an intermediate state to propagate across components. Severity and propagation also show a correlation, even though they measure different properties. Catastrophic failures often have Total impact because the loss of a service blocks the complete workload. Silent and Hindering failures are more often compatible with Multi-component propagation because execution continues long enough for an inconsistent state to reach another service. However, one dimension does not determine the other: a severe failure may remain Local, while a subtle mutation may propagate across multiple components. The Qwen Cinder case with Total impact but no daemon crash further shows that a system-wide effect does not necessarily imply a Catastrophic failure. CRASH, therefore, captures how the workload fails, whereas IMPACT captures how far the effect spreads. Key Takeaway: ProFIPy’s small rule-based mutations more often preserve execution and propagate across components, whereas broader LLM-generated mutations more often remain local or cause system-wide disruption. Severity and propagation are correlated, but capture distinct failure properties.

TABLE IV IMPACT PERCENTAGES . Nova

Cinder

Scope

PF

DS

Qw

PF

DS

Qw

Total Multi Local Dormant

4.3 26.1 52.2 17.4

28.6 9.5 57.1 4.8

22.7 0.0 13.6 25.0 59.1 75.0 4.5 0.0

0.0 11.8 88.2 0.0

5.3 5.3 89.5 0.0

Key Takeaway: Model choice affects how a fault manifests, and the two models show a different mutation strategy. DeepSeek and Qwen often differ in CRASH outcome while producing similar IMPACT scopes, primarily due to the injection location. V. D ISCUSSION The evaluation suggests that LLMs help where effort is spent on fault-model construction. Rule-based injectors require fault patterns to be identified, formalized, and implemented as operators before a campaign can begin. An LLM can instead derive context-dependent mutations directly from the target function. This capability broadens the explorable fault space and makes application-specific mutations easier to generate. However, context-specific mutation alone does not guarantee the representativeness of the injected fault. Fault realism must be assessed in relation to both the defect being modeled and the execution environment in which it is activated. The path, therefore, still requires control over mutation scope, evidence that the modified code executes, representative workloads, failure oracles, and sufficient provenance to reproduce the run. LLMs reduce operator-authoring effort, but they do not eliminate the engineering required to establish fault representativeness.

C. RQ3: How much does the choice of the generative model affect the operational behavior of injected faults? DeepSeek and Qwen exhibit different failure profiles. On Nova, DeepSeek causes 6/21 Catastrophic failures, with a 95% Wilson interval of 14–50%, while Qwen causes 5/22, with an interval of 10–43%. The wide and overlapping intervals indicate uncertainty. Descriptively, DeepSeek produces a larger Catastrophic share, whereas Qwen produces a more varied distribution and the largest Abort share on Cinder (21.1%). To separate model choice from injection location, we compare the faults generated at the 20 Nova and 17 Cinder locations shared by both models. The two models produce the same CRASH outcome in only 55% of the shared Nova locations and 71% of the shared Cinder locations. Thus, even at the same A. From Fixed to Data-Driven Fault Models From our experience, an LLM should be treated as a generator injection point, they often induce different failure manifestations. Agreement is substantially higher for IMPACT: 80% on Nova of faults based on a fault model inferred from the training data. and 88% on Cinder. This suggests that the injection location The inspected mutations show why this distinction matters. constrains how far a fault can spread more strongly than it Generated changes may use the local context to alter exception constrains the exact way in which the workload fails. The handling, interfaces, branches, or several related statements. This model influences whether a location produces, for example, an flexibility can express defects that would require a dedicated Abort, Silent, or Catastrophic outcome, while the surrounding rule-based operator, but it can also introduce multiple semantic architecture often determines whether the effect remains Local changes or replace too much of the original implementation. For fault-injection purposes, broader changes are not necesor reaches other components. The higher agreement on Cinder is consistent with its stronger sarily better. A mutation that immediately prevents a service error containment. Its exception-handling paths absorb or localize from starting may expose a valid robustness weakness, but it many mutations, reducing the observable differences between exercises a different behavior from a small defect that preserves the models. At highly vulnerable locations, the surrounding code execution and allows an inconsistent state to propagate. A can dominate the outcome entirely: both models may produce practical LLM-based injector should support explicit control a system-wide failure despite generating different mutations. over mutation granularity and campaign intent. For example, the Model-specific differences become more visible at locations generation process could request a single semantic change, avoid where execution can continue, and the injected error can follow unrelated rewrites, or favor execution-preserving defects when the objective is to study error propagation. A realistic campaign different control-flow or state-propagation paths. Manual inspection also suggests different mutation styles. should deliberately enable more and reproducible configurations Some accepted DeepSeek faults make small interface-level to replicate fault realism and a thorough evaluation. Figure 3 illustrates this. The LLM uses the surrounding changes, such as modifying parameters or call arguments, which can trigger failures at import, start-up, or call time. exception-handling context to produce an architecture-specific Qwen faults more often include larger body rewrites, additional change. The asynchronous request has already been acknowlbranches, try/except blocks, or reimplemented logic, which edged when the compute-side attachment fails, leaving a volume can preserve execution and produce a wider range of later failures. that appears attached but is unusable. ProFIPy instead applies These observations are qualitative because we did not classify an AST transformation that removes calls whose names match volume [10]. The operator is precise and reproducible, but it every generated mutation using a systematic taxonomy. Overall, similar validity and activation rates do not imply does not account for the semantic role of each matched call. The practical advantage of the LLM is that the exceptionequivalent operational behavior. At the same target location, the handling rewrite does not need to be anticipated and implemented two models can generate valid and activated faults that produce as a dedicated operator. Its limitation is that the generated change different CRASH outcomes. However, their higher agreement is not automatically atomic, representative, or reproducible. on IMPACT shows that the architecture and injection location Hence, validation must also extend beyond parsing or compiremain major determinants of the resulting system-level effect. lation. It should distinguish at least four properties: (i) whether

# Fault-free try: return self._attach_volume(context, instance, driver_bdm) except Exception: with excutils.save_and_reraise_exception(): bdm.destroy() # LLM: exception-handler rewrite - with excutils.save_and_reraise_exception(): - bdm.destroy() + bdm.destroy() + raise

TABLE V K EY REQUIREMENTS FOR GENERATIVE FAULT INJECTION . Concern

Lesson

Recommended action

Fault model

LLMs reduce manual operator authoring but may generate broad or compound changes A plausible source change is not necessarily a realistic fault Compilable mutations may be dormant or produce no observable effect Generated faults depend on the model and sampling configuration Generated faults support exploration but offer limited repeatability

Constrain mutation scope and preserve relevant interfaces

Representativeness Operational validity

# ProFIPy: call-removal operator

Reproducibility - return self._attach_volume(context, instance, driver_bdm) + return

Campaign lifecycle Fig. 3. Fault-free volume-attachment handler and source changes introduced by the LLM and ProFIPy.

the mutation is syntactically valid and deployable; (ii) whether it preserves the intended mutation scope and target interface; (iii) whether the modified code is activated by the workload; (iv) whether activation produces an observable operational effect. Table V summarizes these observations. B. Implications for Developers and Researchers For tool developers, the main implication is that generation should be embedded in a staged fault-injection pipeline. The model should propose a mutation, while deterministic components enforce structural constraints, check runtime compatibility, deploy the modified program, verify activation, and collect system-level observations. This division retains the flexibility of generative models without delegating the validity of the experiment entirely to the model. Runtime feedback can further improve this process. Dormant mutations can be regenerated at covered locations, while changes that consistently cause immediate initialization failures can be balanced with prompts or constraints that favor executionpreserving behavior. Failure observations can also be used to select mutations according to campaign objectives, such as service recovery, silent-state corruption, exception containment, or cross-component propagation. The relevant optimization target is therefore not merely the probability of producing compilable code, but the probability of producing an activated and operationally informative fault. For researchers, the results highlight the need to evaluate generative fault models at multiple levels. Code-level properties such as syntax, similarity to historical defects, or mutation score describe the generated artifact but not its behavior in a deployed system. Evaluations should additionally measure activation, failure visibility, severity, propagation, and final system state. Comparisons should control for injection location whenever possible because the target code and surrounding architecture can dominate the resulting behavior. Generated faults also require stronger provenance than deterministic operators. Each run should retain the target location, prompt, generation attempts, validation decisions, deployment outcome, coverage, workload observations, and

Ground generation in real bug data and review fault semantics Validate deployment, measure activation, and use end-to-end oracles Record generation settings, validation decisions, and source diffs Promote relevant generated patterns into deterministic operators

failure classifications. Without these records, differences due to generation variability cannot be distinguished from those due to the testbed or workload. A hybrid lifecycle can combine exploration and reproducibility by using LLMs to discover application-specific mutations without requiring a dedicated operator for each defect, then minimizing, reviewing, and encoding recurrent or operationally relevant mutations as deterministic operators. In this way, the rule-based catalog becomes an evolving and reproducible core informed by generative exploration. VI. T HREATS TO VALIDITY Construct validity. The CRASH scale required adaptation to distributed cloud services, and the distinction between Silent and Hindering outcomes depends on oracle coverage. Start-up crashes are also easier to detect than delayed state corruption. We mitigate these threats through explicit cloudlevel definitions, complementary API, log, state, and coverage oracles, and a separate propagation classification to show their representativeness and failure outcomes. Internal validity. The Python 2.7 compatibility filter may bias the accepted LLM suites toward legacy-compatible faults. Campaign-level differences may also reflect different target locations; we therefore analyze the locations shared by all injectors, although they are not a random sample. Generation is stochastic, and only the first valid candidate within three attempts is evaluated, so differences between Qwen and DeepSeek may also reflect sampling. Given the limited sample sizes, comparisons are primarily descriptive; Wilson intervals quantify uncertainty around individual proportions but do not establish differences between injectors. External validity. The study covers two services, one OpenStack version, one workload, and 122 partly overlapping scenarios. Results may differ for other services, workloads, concurrency levels, platforms, or fault locations. We also evaluate only two code-specialized models, one training dataset, one generation configuration, and one rule-based injector. The findings should therefore be interpreted as evidence from the evaluated campaign rather than as universal superiority of either paradigm.

VII. C ONCLUSION AND F UTURE W ORK This study shows that LLM-based fault injection can complement fixed-model SFI by generating context-dependent mutations without requiring a dedicated operator for each fault pattern. In the evaluated OpenStack campaign, LLM-generated faults broadened the observed failure space, while ProFIPy provided greater control, repeatability, and more execution-preserving and propagating failures. The contribution is therefore not generatorlevel superiority, but evidence that LLMs can reduce fault-model authoring effort and extend fixed catalogs. Realistic generative fault injection still requires controlled mutation scope, deployment and activation validation, representative workloads, system-level oracles, and reproducible provenance. Future work will evaluate broader settings, incorporate runtime feedback into generation, and promote relevant generated faults into deterministic operators, toward an adaptive pipeline. R EFERENCES [1] OpenInfra Foundation, “OpenStack: Open source cloud computing infrastructure,” Project website, 2026, accessed: 2026-07-13. [Online]. Available: https://www.openstack.org/ [2] R. Natella, D. Cotroneo, J. A. Durães, and H. S. Madeira, “On fault representativeness of software fault injection,” IEEE Transactions on Software Engineering, vol. 39, no. 1, pp. 80–96, Jan. 2013. [3] R. Natella, D. Cotroneo, and H. S. Madeira, “Assessing dependability with software fault injection: A survey,” ACM Computing Surveys, vol. 48, no. 3, pp. 44:1–44:55, Feb. 2016. [4] G. De Rosa and P. Liguori, “Will it break in production? metric-driven prediction of residual defects in python systems,” in 2026 56th Annual IEEE International Conference on Dependable Systems and Networks (DSN). IEEE, 2026, pp. 772–785. [5] D. Cotroneo, G. De Rosa, C. Improta, and B. G. Varriale, “What makes software bugs escape testing? evidence from a large-scale empirical study,” in 2026 56th Annual IEEE International Conference on Dependable Systems and Networks (DSN). IEEE, 2026, pp. 136–149. [6] D. Cotroneo, L. De Simone, P. Liguori, R. Natella, and N. Bidokhti, “How bad can a bug get? an empirical analysis of software failures in the OpenStack cloud computing platform,” in Proceedings of the 27th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. New York, NY, USA: Association for Computing Machinery, 2019, pp. 200–211. [7] J. A. Durães and H. S. Madeira, “Emulation of software faults: A field data study and a practical approach,” IEEE Transactions on Software Engineering, vol. 32, no. 11, pp. 849–867, Nov. 2006. [8] R. Chillarege, I. S. Bhandari, J. K. Chaar, M. J. Halliday, D. S. Moebus, B. K. Ray, and M.-Y. Wong, “Orthogonal defect classification—a concept for in-process measurements,” IEEE Transactions on Software Engineering, vol. 18, no. 11, pp. 943–956, Nov. 1992. [9] H. Marques, N. Laranjeiro, and J. Bernardino, “Injecting software faults in Python applications: The OpenStack case study,” Empirical Software Engineering, vol. 27, no. 1, pp. 1–33, 2022, article 20. [10] D. Cotroneo, L. De Simone, P. Liguori, and R. Natella, “ProFIPy: Programmable software fault injection as-a-service,” in 2020 50th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). IEEE, 2020, pp. 364–372. [11] B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Dang, A. Yang, R. Men, F. Huang, X. Ren, X. Ren, J. Zhou, and J. Lin, “Qwen2.5-Coder technical report,” arXiv preprint arXiv:2409.12186, 2024. [Online]. Available: https://arxiv.org/abs/2409.12186 [12] D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. K. Li, F. Luo, Y. Xiong, and W. Liang, “DeepSeek-Coder: When the large language model meets programming—the rise of code intelligence,” arXiv preprint arXiv:2401.14196, 2024. [Online]. Available: https://arxiv.org/abs/2401.14196 [13] C. Improta, R. Tufano, P. Liguori, D. Cotroneo, and G. Bavota, “Quality in, quality out: Investigating training data’s role in AI code generation,” in 2025 IEEE/ACM 33rd International Conference on Program Comprehension (ICPC). IEEE, 2025, pp. 454–465. [14] F. Tip, J. Bell, and M. Schäfer, “LLMorpheus: Mutation testing using large language models,” IEEE Transactions on Software Engineering, vol. 51, no. 6, pp. 1645–1665, Jun. 2025.

[15] B. Wang, M. Deng, M. Chen, C. Yang, Y. Lin, M. Harman, M. Papadakis, and J. M. Zhang, “Boosting LLMs for mutation generation,” Proceedings of the ACM on Software Engineering, vol. 3, no. FSE, pp. 3368–3391, Jun. 2026, article FSE149. [16] M. Tufano, C. Watson, G. Bavota, M. Di Penta, M. White, and D. Poshyvanyk, “Learning how to mutate source code from bug-fixes,” in 2019 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 2019, pp. 301–312. [17] D. Cotroneo, G. De Rosa, and P. Liguori, “PyResBugs: A dataset of residual Python bugs for natural language-driven fault injection,” in 2025 IEEE/ACM Second International Conference on AI Foundation Models and Software Engineering (Forge). IEEE, 2025, pp. 146–150. [18] M.-C. Hsueh, T. K. Tsai, and R. K. Iyer, “Fault injection techniques and tools,” Computer, vol. 30, no. 4, pp. 75–82, Apr. 1997. [19] M. Cinque, D. Cotroneo, G. De Rosa, L. De Simone, and G. Farina, “Cosmos: A fault injection framework to assess hardware-assisted hypervisors,” IEEE Transactions on Dependable and Secure Computing, 2025. [20] Y. Jia and M. Harman, “An analysis and survey of the development of mutation testing,” IEEE Transactions on Software Engineering, vol. 37, no. 5, pp. 649–678, 2011. [21] P. Costa, J. G. Silva, and H. Madeira, “Practical and representative faultloads for large-scale software systems,” Journal of Systems and Software, vol. 103, pp. 182–197, May 2015. [22] R. Just, D. Jalali, L. Inozemtseva, M. D. Ernst, R. Holmes, and G. Fraser, “Are mutants a valid substitute for real faults in software testing?” in Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering. New York, NY, USA: Association for Computing Machinery, 2014, pp. 654–665. [23] H. S. Gunawi, T. Do, P. Joshi, P. Alvaro, J. M. Hellerstein, A. C. Arpaci-Dusseau, R. H. Arpaci-Dusseau, K. Sen, and D. Borthakur, “FATE and DESTINI: A framework for cloud recovery testing,” in 8th USENIX Symposium on Networked Systems Design and Implementation (NSDI 11). Boston, MA: USENIX Association, Mar. 2011. [Online]. Available: https://www.usenix.org/conference/nsdi11/ fate-and-destini-framework-cloud-recovery-testing [24] P. Alvaro, J. Rosen, and J. M. Hellerstein, “Lineage-driven fault injection,” in Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data. New York, NY, USA: Association for Computing Machinery, 2015, pp. 331–346. [25] C. S. Meiklejohn, A. Estrada, Y. Song, H. Miller, and R. Padhye, “Servicelevel fault injection testing,” in Proceedings of the ACM Symposium on Cloud Computing. New York, NY, USA: Association for Computing Machinery, 2021, pp. 388–402. [26] H. Wu, J. Pan, and P. Huang, “Efficient exposure of partial failure bugs in distributed systems with inferred abstract states,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). Santa Clara, CA: USENIX Association, Apr. 2024, pp. 1267–1283. [Online]. Available: https://www.usenix.org/conference/nsdi24/presentation/wu-haoze [27] D. Cotroneo and P. Liguori, “Neural fault injection: Generating software faults from natural language,” in 2024 54th Annual IEEE/IFIP International Conference on Dependable Systems and Networks—Supplemental Volume (DSN-S). IEEE, 2024, pp. 23–27. [28] T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “QLoRA: Efficient finetuning of quantized LLMs,” in Advances in Neural Information Processing Systems, vol. 36. Curran Associates, Inc., 2023. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2023/ hash/1feb87871436031bdc0f2beaa62a049b-Abstract-Conference.html [29] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations, 2022. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9 [30] X. Ju, L. Soares, K. G. Shin, K. D. Ryu, and D. Da Silva, “On fault resilience of OpenStack,” in Proceedings of the 4th ACM Symposium on Cloud Computing. New York, NY, USA: Association for Computing Machinery, 2013, pp. 2:1–2:16. [31] N. Batchelder and contributors, “coverage.py: Code coverage measurement for Python,” Software documentation and source code, 2026, accessed: 2026-07-13. [Online]. Available: https://coverage.readthedocs.io/ [32] P. J. Koopman, Jr., J. Sung, C. P. Dingman, D. P. Siewiorek, and T. Marz, “Comparing operating systems using robustness benchmarks,” in Proceedings of the 16th Symposium on Reliable Distributed Systems (SRDS). IEEE Computer Society, 1997, pp. 72–79. [33] DESSERT Lab, “OpenStack fault injection environment,” Software artifact, 2019, accessed: 2026-07-13. [Online]. Available: https: //github.com/dessertlab/OpenStack-Fault-Injection-Environment

Record · ID 668102 · SHA-256 c291485d5cb81281
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.