Conceptio › Archive › arXiv CS
arXiv CSopen access

PatchyBFT: Automating Diversification of Fault-Tolerant Systems using LLMs

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

PatchyBFT: Automating Diversification of Fault-Tolerant Systems using LLMs Arne Vogel

Christian Berger

Rüdiger Kapitza

[email protected] FAU Erlangen-Nürnberg Erlangen, Germany

FAU Erlangen-Nürnberg Erlangen, Germany [email protected]

FAU Erlangen-Nürnberg Erlangen, Germany [email protected]

arXiv:2609.18512v2 [cs.DC] 17 Sep 2026

Abstract Fault-tolerant agreement protocols fail if replicas share a common flaw that simultaneously affects more replicas than the tolerable threshold. Therefore replicas should ideally fail independently, which can be achieved through diversification. However, in practice, often the same protocol implementation is shared by all replicas which is not surprising given that the provision of multiple diverse implementations is difficult and highly laborious. This poses a major risk, as a shared protocol implementation is a prime candidate for common bugs due to its complexity. With PatchyBFT, we demonstrate how, given a reference implementation, Large Language Models (LLMs) can be utilised for the automated and scalable generation of code that compiles, passes tests, and crucially differs semantically/binary-wise, that can replace code in the reference implementation, thereby significantly reducing diversification costs. We demonstrate the feasibility of diversification of replication protocol implementations using LLMs by diversifying three implementations: PBFT, HotStuff, and Raft, showing how up to 65% of the codebase can be diversified.

CCS Concepts • Computer systems organization → Dependable and faulttolerant systems and networks; Reliability; Redundancy; • Software and its engineering → Software creation and management; Software verification and validation.

Keywords Byzantine Fault Tolerance, Crash Fault Tolerance, State Machine Replication, Software Diversity, Large Language Models

1

Introduction

Fault-tolerant agreement protocols have become increasingly crucial for scalable web services [104], cloud infrastructures [24], and, more recently, distributed ledger systems [10, 11, 42, 64]. Depending on the situation, these protocols can be crash or Byzantine faulttolerant (BFT) and are designed to handle 𝑓 faults in 𝑛 = 2𝑓 + 1 or 𝑛 = 3𝑓 + 1 replicas, respectively, while still providing correct results [18]. In practice, replicas must be fault-independent with respect to common failure modes, such as shared vulnerabilities in replica implementations, deployment in the same region, or operation by the same operator, to achieve this theoretical guarantee. In this work, we focus on diversifying the implementations of agreement protocols as otherwise a shared bug in the system could easily violate the limit of 𝑓 faulty replicas. One of the core methods for providing fault-independent software for replicas is N-version programming [5]. It requires the implementation of multiple, diverse

versions of a system based on a common specification, ideally by different development teams using different programming languages or development methodologies, to minimize the risk of commonmode errors. As a result, N-version programming is considered prohibitively resource-intensive and is typically not applied to common IT services, even though their unavailability or corruption can lead to a poor user experience and significant revenue losses. Since the seminal work on making Byzantine Fault Tolerance practical for everyday IT services [18], the question of how to achieve fault independence in practice without adopting full N-version programming has remained an open topic especially relevant today with distributed ledgers managing billions in value [10, 77, 95]. One line of work is opportunistic N-version programming, which builds on software heterogeneity for well-established APIs [20]. As an example, diversification at the operating system level is a direction that has been proposed [37, 81], as many systems provide a POSIXcompliant system interface so that the protocol implementation and the replicated application can run on a diverse set of operating systems that share almost no bugs. Opportunistic N-version programming has also been explored at the replicated application or service level. Examples include using different relational databases, as SQL is a common standard with various implementations [38, 91]. Finally, more general diversification mechanisms can be applied, such as address space layout randomization [84] or introducing diversification during compilation [47] such as function inlining, outlining, splitting, control flow flattening, or system call mapping randomization among others [21, 53], which have been used in distributed systems such as Spire [6]. Despite all these previous works, diversification at the protocol level, which is at the heart of replicated systems, has largely been abandoned. This is not surprising, as implementing a fault-tolerant agreement protocol involves highly concurrent code featuring complex communication logic and requires the correct use of various cryptographic methods. Thus, even providing a single correct implementation that offers good performance is already a significant challenge. As a result, achieving diversification at the level of the Byzantine fault-tolerant agreement protocol via N-version programming has so far only been achieved for very few widely used protocols such as Ethereum [14, 34]. LLMs currently change the way code is written in industry, with 80% of professional developers already using them, 50% even daily [86]. They are used for code understanding [63], to generate new code [58], explore the design space of programs [102], and even for N-version programming to combat compiler bugs (limited to pure functions without side effects, not suitable for most agreement protocol implementation functions) [74].

Vogel et al.

In this paper, we propose PatchyBFT, which enables the highly automated and scalable diversification of Byzantine Fault Tolerant (BFT) and Crash Fault Tolerant (CFT) agreement protocol implementations using LLMs. This is achieved by using multiple LLMs to generate diversified implementations based on an original version. However, naive code generation is insufficient, so we must address three challenges: (1) At what level of abstraction should code diversification occur and how much context should be provided to an LLM to drive the generation? (2) How can we ensure that the diversified code maintains functional equivalence? (3) How can we validate that diversification actually achieves representational and binary difference rather than merely syntactic variation? To address these challenges, we designed and implemented PatchyBFT, which automates the diversification of Rust-based distributed protocols. Rust is a modern systems programming language widely adopted in the systems research community [9, 39, 78, 94]. Despite the security properties of Rust [22, 71, 87], diversification is still relevant for defending against implementation-level vulnerabilities that escape Rust’s compile-time safety guarantees [1, 2, 43, 60]. To highlight the benefits of PatchyBFT, we automatically diversified Themis [79] (implementing the PBFT algorithm [18, 78]), hotstuff_rs [66] (implementing the HotStuff algorithm [101]), and Openraft [25], a CFT replication protocol implementation. Contributions. PatchyBFT makes diversification practical for BFT and CFT protocols. We demonstrate how to use LLMs to generate protocol changes at function level based on reference implementations. For validation of these changes, we introduce safeguards for functional and diversification correctness. As a proof of concept, we additionally highlight that LLMs can go beyond simple diversification and even fix bugs in BFT protocols, without having been informed of the specific bug. Further contributions include: • Automated protocol diversification: We demonstrate how to use LLMs to generate representation/binary different implementations of complex replication protocols, with up to 65% of the original code being automatically changed. • Functional validation framework: We show how to validate the generated code for functional equivalence with unit tests and protocol conformance tests against the reference implementation, ensuring changes do not break the protocol. • Representational/binary validation framework: We introduce a new methodology for how to ensure that diversified code is representationally and binary different, using clone detection and binary difference validation. The rest of the paper is structured as follows: In Section 2 we provide background information and related work on BFT systems, LLMs and Rust. This is followed by an investigation of using LLMs to create patches in Section 3. Next, we describe the design of PatchyBFT and describe how PatchyBFT fits into the lifecycle of distributed systems in Section 4 followed by implementation details in Section 5. We evaluate PatchyBFT in Section 6 and show future work in Section 7. Finally, we conclude in Section 8.

2

Background & Related Work

Diversification of Byzantine Fault-Tolerant protocols. BFT protocols can tolerate up to 𝑓 faulty replicas in a system of 3𝑓 + 1 replicas [18, 30]. This is a theoretical limit that requires faultindependent implementations in practice. Otherwise, an attacker can exploit the same bug in all replicas, easily exceeding the theoretical threshold of 𝑓 faulty replicas. Similarly, CFT protocols have a threshold of 2𝑓 + 1 where up to 𝑓 crashes can be tolerated. Testing techniques are proposed to limit faults, but are unable to give comprehensive guarantees [8], and while formal models [44, 54, 98] can give comprehensive correctness guarantees, they are laborious (3.7 person years reported) [44] and can still contain bugs through assumptions made for the formal specification [35]. Nversion programming is argued to create fault-independent implementations [5]. However, N-version programming does not scale and due to its high additional development costs, it is rarely used in practice. In Table 1, we compare PatchyBFT with existing work that tackles the issue of fault-independent implementations. Even for Ethereum, which handles billions of dollars in transactions, only a handful of independent implementations exist [14, 34]. The implementations are independently developed and maintained; however, despite the risk of billions of dollars, these independent implementations are rarely utilized [34]. Recent research works try to address this issue for Ethereum by using Trusted Execution Environments (TEEs) for verifiable client diversity and a reward protocol that incentivizes diverse clients [75]. This approach could be integrated into PatchyBFT in permissionless setups, but does not address the issue of how to practically get diversified implementations. BASE [72] provides an abstraction layer that enables the use of independent service implementations, such as different NFS servers. Lazarus [37] monitors vulnerability databases, automatically quarantines vulnerable replicas from the system, and patches it with available patches before allowing it back into the system. Similarly, FOREVER [81] uses evolutions of the underlying system (open ports, authentication mechanisms) and updates for the application to diversify replicas. Works such as BASE, Lazarus, and FOREVER lighten the burden of diversification from the developer. Garcia et al. have analyzed whether operating systems share vulnerabilities and identified that using diverse operating systems has benefits for intrusion-tolerant systems [36]. Still, these techniques do not provide fault-independent implementations for the BFT protocol itself. With Proactive Obfuscation [73], the authors propose using obfuscation techniques such as address reordering, stack padding, or system call randomization to diversify replicas. Although this approach indeed diversifies the protocol implementation, these techniques can only obfuscate vulnerabilities but cannot remove them; we consider them therefore as an orthogonal approach that can be combined with PatchyBFT. Recent works, such as SplitBFT [61], propose a mechanism to ease diversification: the consensus protocol is split into multiple parts (preparation, confirmation, and execution) that can be implemented independently. In addition to non-standard hardware (requiring TEEs), SplitBFT still requires N-version programming. While important, the previously proposed

PatchyBFT : Automating Diversification of Fault-Tolerant Systems using LLMs

techniques do not meet the goal of automated protocol diversification: they either overlook the BFT protocol implementation itself or provide only limited diversification.

function a(

tokenizer

Work

LLM

probabilities

func tion a (

)

0.327

i32

0.293 0.194

f32 bool

0.123

Diversification Automated Code at protocol level generation diversification

SplitBFT [61] Ethereum [14, 34] FOREVER [81] Proactive obfuscation [73] Platania et al. [68] BASE [72] Lazarus [37]

✓ ✓ p

p p ✓

✓ ✓ p

✓

✓

p

✓ p p

✓ ✓ ✓

p ✓ p

PatchyBFT

✓

✓

✓

Table 1: Comparison of PatchyBFT to existing work. To the best of our knowledge, PatchyBFT is the first work that provides diversified implementations for the core protocol of a BFT system in a scalable way.

Rust is a systems programming language designed to be safe and fast [59]. It features a strict type system and ownership model to achieve this safety, which has been shown to decrease the number of vulnerabilities [22, 71, 87]. These safety guarantees have led to the widespread adoption of Rust in distributed protocols and systems programming [9, 39, 78, 94]. But Rust is not infallible. With unsafe, blocks of code can be marked that cannot be fully validated by the Rust compiler (e.g., dereferences of arbitrary pointers are allowed). This is necessary because the ownership model can be too restrictive for specific use cases. However, this can lead to bugs. And even without unsafe code, Rust bugs can still occur and lead to security vulnerabilities [7, 43, 55, 60, 69, 99]. Hassnain et al. give concrete examples of such security vulnerabilities in safe Rust code [43] and Meneely et al. show that only up to 58.2% of C vulnerabilities would have been fixed in a Rust port of the same code [60]. Therefore, even for Rust-based agreement protocols, diversification is needed. Still, the strong type system and ownership model eliminate many potential bugs and are beneficial for distributed protocols [78]. For us, the type system and ownership model are also beneficial in identifying that patches are correct, as type conversion errors are caught early by the compiler. Large Language Models (LLMs) are statistical models designed to predict text based on previous text [15, 65, 70]. They have been used for text generation, summarization, translation, and code generation [15]. This code generation capability is widely used by software developers, with more than 84% of developers using it [86] and 43% of developers already somewhat trusting the output [85]. An LLM is prompted with a set of text that the LLM uses to generate the next piece of text. Figure 1 shows an example of such a prompt (function a() and the subsequent tokens generated by the LLMs, along with their probabilities. Each token has a probability of being the next token in the sequence, which is controlled by the temperature parameter. The higher the temperature, the more random the output, i.e., the LLM is more likely to generate tokens with lower probabilities [70]. LLMs can use Chain-of-Thought (CoT)

Figure 1: LLM code generation example. The LLM takes a prompt as input (e.g., function a() and returns tokens along with their probabilities.

to reason about tasks before generating outputs [52, 97]. Recent LLMs explicitly generate intermediate reasoning steps, sometimes delimited by tags like <think></think>, before a final answer. LLMs have been used for various software engineering tasks, such as understanding code [63], generating new code [58], and exploring program design spaces [102]. Huynh et al. [46] show a 45% success rate of LLMs patching vulnerabilities in a dataset of C/C++ vulnerabilities. Peng et al. [67] see similar results (up to 47%) for vulnerability patching across 5 programming languages. The works of Huynh et al. and Peng et al. show LLMs can fix vulnerabilities in existing code, but do not offer a framework to validate diversification or how to ensure changes are functionally correct. Ron et al. [74] investigated with Galápagos if LLMs can be used for automatic N-version programming. They implemented automatic validation for the correctness and equivalence of the generated code using off-the-shelf format equivalence checking tools such as alive2 [57], or Kani [92]. Galápagos verifies the equivalence of the generated code but limits the functions it can diversify to pure functions, which have no side effects. This is not suitable for most (if not all) BFT implementations that are mostly implemented with stateful functions. PatchyBFT instead provides an integrated method for BFT and CFT systems. Liu et al. [56] and Du et al. [32] both introduced a benchmark for the correctness of LLM generated code. Liu et al. generate test cases using LLMs and mutation-based strategies, while Du et al. wrote 1889 Python benchmarks. These benchmarks are used to evaluate the correctness of generated code by checking the output of the LLM against the fixed expected test case output, but they do not allow verification of any specific function required for generic BFT implementations.

3

Motivation: LLMs Can Create Patches

To show the feasibility of our approach, we first conduct a proofof-concept experiment that evaluates whether LLMs could remove potential bugs from a reference implementation during diversification. Although other work has shown this for different systems and languages [46, 67, 100], we were interested in verifying the capability for two classes of bugs related to our work: faults in distributed algorithms, and security vulnerabilities of Rust. For the faults in distributed systems we obtained an unpublished bug in Themis1 . For Rust vulnerabilities, we used two real-world security bugs, CVE-2017-1000430 [1] and CVE-2019-16140 [2]. This also underscores the fact that security-relevant bugs still pose a risk to distributed protocols, despite the previously mentioned security guarantees for Rust. For this proof-of-concept we diversified the 1 from private correspondence with the authors

Vogel et al.

4

Human fix

if message.destination != self.own_id { return self.decode(src); } if message.source != self.peer_id { return self.decode(src); }

LLM fix

if message.destination != self.own_id { return Err(io::Error::new(InvalidData)); } if message.source != self.peer_id { return Err(io::Error::new(InvalidData)); }

Listing 1: Fix for a decoder bug in Themis where messages were not properly validated.

Qwen

Original System

Deepseek Mistral

Diversified System

Build & Test

applied

Diversification Potential Patches

...

Execution

Verified Patches

Validation with Safeguards

Figure 2: PatchyBFT system overview: PatchyBFT generates sets of patches which are validated before using the set of validated patches for diversification. 1.3% Exploit Chance

3.7% Exploit Chance

15%

Risk 1 0.8

Bug Rate

10%

0.6 0.4

5%

Fault Independence

%

0%

0

10

%

80

%

60

%

40

20

0%

20

0%

% 40 % 60 % 80 % 10 0%

0.2 0%

vulnerable function as described in the design (Section 4) without hints of the vulnerabilities. We then manually checked whether the vulnerability survived diversification, or whether the LLM fixed the underlying vulnerability. Listing 1 details an unpublished impersonation bug in Themis’s PBFT implementation. Messages were not properly checked. As such, a Byzantine replica could impersonate other replicas. The fix was to add checks for the source and destination of the message. In the human fix, an additional optimization is implemented (self.decode(src)) that recursively decodes the rest of the message buffer. The LLM did not detect this because it simply terminates processing on impersonation bugs, which effectively fixes the issue. The Rust vulnerabilities were also fixed during diversification. For the first bug, differing from the fix by the library’s maintainers, the LLM did not change the function signature. For the second bug, the LLM fixed the bug and even removed the need for the unsafe block that introduced the vulnerability, with the drawback of initializing memory twice. These examples show that LLMs are able to fix security vulnerabilities in real-world code, including distributed protocols. This is supported by the works of Huynh et al. [46] and Peng et al. [67], which show LLMs successfully patching security vulnerabilities with 45% and 33%–47% probability, respectively. The fixes demonstrate that LLMs go beyond the existing automatic code diversification approaches. LLMs are able to fix underlying problems when prompted to diversify code. Additionally, this proof of concept shows that the changes go beyond simple copy-pasting from the training set of the LLMs. For the two Rust fixes, the condition not to change the function signature made the LLMs generalize beyond the training set. Furthermore, since the Themis bug is previously unpublished, it is simply not in any LLM training set. Knight and Leveson [51] showed in their seminal work that humans make correlated mistakes when creating multiple versions of a program. While it is unknown which datasets are used for LLMs [93], different LLM providers use different training and reinforcement techniques [96], and there is research on interpretability models which can disable parts of the training set during inference [41] and we have anecdotally seen in this experiment that LLMs can generalize beyond their training data; still, it is an open question whether different LLMs can create fully fault-independent code. Because of this, we focus on ensuring that diverse implementations are representationally different and the compiled binary differs from the original implementation.

Fault Independence

Figure 3: Survival of a whole deployment vs. the bug introduction rate & fault independence against a limited attacker. Even at a 10% chance of introducing a bug, the system can remain safe even against a strong attacker.

PatchyBFT

In short, PatchyBFT takes an implementation of a BFT (or CFT) protocol and automatically generates diversified variants of it (Section 4.2). We do not naively trust the generated code (Section 4.1). Instead, we validate the generated code using various safeguards. These validation steps ensure that the generated code is diversified and integrates without manual work into the existing codebase which we verify with extensive stress testing (Section 4.3). Finally, we show how PatchyBFT fits into the lifecycle of distributed systems in Section 4.4. In essence, we generate sets of patches, validate them for functional equivalence and ensure they are diverse

through safeguards, before using the set of validated patches for the diversification of the system.

4.1

System and Attacker Model

In reality, a bug in a system does not automatically make it vulnerable to an attacker [48, 80]. Therefore, in this work we assume that the attacker is not omnipotent in exploiting vulnerabilities, but rather has a probability of exploiting vulnerabilities. This assumption aligns with the works of Sousa et al. [82, 83], who assume a

PatchyBFT : Automating Diversification of Fault-Tolerant Systems using LLMs

minimum inter-failure time and that the attacker cannot compromise nodes instantaneously or at an arbitrarily high rate. Our system is designed to diversify a BFT protocol implementation. For most current BFT deployments, a single implementation is used for all replicas. This implementation may contain bugs that an attacker could exploit to compromise more than 𝑓 replicas. To mitigate this risk, the goal of PatchyBFT is to generate diversified variants of the BFT implementation using LLMs. Code generated by an LLM is not always correct in the sense that it exhibits the same functional behavior. It may also have other issues and, therefore, needs to be validated before it can be used. We assume that the original protocol implementation includes a series of tests to validate it during development. These tests are considered the ground truth and correct, but they are not comprehensive; they cannot be used as an oracle to determine if any code is correct or not. We use them to validate the generated code and aim to identify and reject all instances where an LLM generated incorrect code. In rare exceptional cases in which safeguards fail to detect issues, the BFT protocol can tolerate bugs specific to individual replica instances rather than affecting all replicas in the absence of diversification, as shown in Figure 3. Furthermore, approaches such as proactive recovery [19, 20] can be used to strengthen resilience further. Implications of the attacker model As a result of our attacker model, even if diversification potentially adds (dependent) faults to diversified functions, it does not automatically allow an attacker to exploit the whole system. This perspective of considering the safety of the entire system and not just single functions is shown in Figure 3. The figure gives the probability that the whole system, not a single replica, is vulnerable, depending on the rate at which faults are introduced during diversification, the probability that bugs are shared during diversification, and the exploit chance, i.e., the probability that an attacker will exploit any given bug. For the probability of an attacker exploiting a bug, we used numbers from the literature which give probabilities of 1.3% [48] and 3.7% [80]. Each replica has 200 functions in which a vulnerability would be severe enough to overtake the replica. The analysis shows that even in a scenario in which 3.7% of bugs are found and exploitable, the diversification will still protect the system even if up to 10% of the functions introduce faults and at least 90% fault independence exists between them. This gives us the assurance that as long as we can keep the exploit rate low enough, potential bugs in diversification will not risk the whole system. Note that while our work cannot show fault-independence, we show representational (using clone detection) and binary diversification, since even with full fault correlation between replicas the system can remain safe as seen in the lower left quadrants of Figure 3. With recent advancements in LLMs, LLMs have been used to identify long standing vulnerabilities in projects such as Firefox, OpenBSD, or ffmpeg [4, 16]. We imagine that this capability can identify bugs in implementations and patches and thus reduce exploit probabilities for whole systems when used before deployment of systems. Further, recent scientific works use LLMs to identify and fix security vulnerabilities with fuzzing [103] which could further reduce vulnerabilities from implementations.

4.2

Code Generation

To motivate the code generation design, we first answer questions about how to use LLM for diversification. At what level should we diversify? There are multiple levels at which we could diversify the code: a single line at a time, at the function level, at the file level, or at the project level. We decided to diversify at the function level. Changing one line at a time is too fine-grained and would leave little room for meaningful diversification. On the other extreme, changing the entire project at once would be challenging for LLMs as it reaches the limits of how much code can be generated. We decided against diversifying at the file level and instead diversifying at the function level. Otherwise, a change might affect multiple functions at once, creating functional dependencies between patches. This would mean that for the application of patches we could no longer freely mix and match patches for different functions from different LLMs. For sufficient context on the codebase, we provide the LLM with additional information about the function (e.g., function signatures and struct definitions). Should we explicitly prompt the LLM with the baseline code or just the intended functionality? This way, we would not risk bugs from the baseline code being copied. Unfortunately, we found that this approach does not work for two reasons. Firstly, the original code lacks sufficient comments and explanations to explain its intended functionality clearly (under-specification problem). Secondly, even when manually providing more context for this approach, the LLM would generate code that would not fit into the existing codebase. How to create multiple diversified versions of the code? We use multiple LLMs to generate diversified code. This way we use LLMs implemented with different algorithms and trained with different training data. While LLMs can share public training data, e.g., data from GitHub, each LLM is trained with a different focus, e.g., explainability or usability. LLMs are fine-tuned to fit this focus, which is often a proprietary process unique to each LLM. In doing this, we diversify the risk of a single LLM potentially generating code with the same bug. We not only generate diversified code, but we also use diversified LLMs to generate the code. For each newly generated variant of the code, we validate that it is diversified compared to the original code and the previously generated versions. More specifically, we check if the code is not a clone [76] and if the generated binary differs from any previous code. Prompts. We prompt the LLM with the full-function body that we want to diversify. Our general prompt is shown in Listing 4. Additionally, we provide the function signature and struct definitions used in the function. This additional context allows the LLM to generate more meaningful code. An example excerpt of the context is shown in Listing 2. With this context, the LLM can access struct definitions not previously seen in the original function. The evaluation (Section 6) demonstrates that this approach produces more successful patches. The prompt limits what the LLM will generate. It ensures that only code is returned, with no explanations or comments on the code. This way, we can ensure that the generated code integrates into the existing codebase without requiring any post-processing. Without these limitations in the prompt, we found that the LLM would often deviate from the task generating output other than code.

Vogel et al.

Diversification Prompt > You are a senior Rust engineer that will create a diversified function implementation. You will be given a function from a practical byzantine fault tolerant (PBFT) implementation. Consider if the function is implemented correctly based on the function name and the PBFT specification. Write correct code, if there is a bug in the code provided you will fix the bug in the output. Your goal is to write correct and safe code. Avoid unsafe code if possible, do not use unwrap. Write defensive code. You can add more checks and assertions if you think they would be useful. The code you generate will have: * the same function signature * return the same type The code will have to work as a drop in replacement, only the internals of the function can change. The function you diversify will be provided in <function> </function> tags. You will output the diversified function in <output></output> tags. The output should only contain valid code. Do not add “‘rust in the output. You only return the newly created rust code, that can compile without warnings. No other text, explanations or other information. Just Rust Code! Here is the function <function> {FUNCTION} </function>. For context here are all the functions the function calls <called> {CONTEXT_FUNCTION} </called>. For context here are all the data structures the function uses <structures> {CONTEXT_DATA} </structures>. Task: Provide an alternative and safe Rust implementation with the same function signature and functionality as the given code $ snippet in <function> tags and output it with <output> tags!

Figure 4: General PBFT prompt for the LLM. For other protocols, protocol-specific terms are replaced accordingly.

Dynamically generated LLM context enum ViewState { Regular, ViewChange { new_view: u64, _timer: Timer, ... pub struct OrderingLog { current_view: Slots<OrderInstance>, old_views: Vec<Slots<OrderInstance>>, }

Listing 2: The context for the LLM is dynamically generated from the diversified function. We provide definitions of any structs or enums used in the function, as well as function signatures.

4.3

Code Safeguards

Before using any generated code, we need to validate that the code behaves as expected and is diversified. Therefore, the generated code must pass through multiple safeguards before being used (see Figure 2). In particular, we validate that the code is functionally correct as well as an actual diversification of the code. For the functional tests we build the code to check syntactical correctness, run unit tests, and run system tests together with unpatched replicas, before finally stress testing the patches by running fully patched implementations together. Further, to validate diversification we use clone detection on the level of the source code as well as validate that the generated binary code for the function differs from the unpatched implementation. Build & Test. The initial validation is the build-and-test safeguard. This safeguard is as simple as it sounds; we try to build the code and run all tests. After this safeguard, we know the code is syntactically correct, and all tests continue to pass. As with conventional development, high test coverage helps catch obvious errors early on. Diversification. After verifying that the code is syntactically correct, we validate that it is diversified. For this we use clone

detection [76, 105] and compare the generated binary code for the function. Clone detection is widely researched with four types of clones considered: Type 1: Identical code except for whitespace and comments Type 2: Syntactically identical code except for changes in variable/function names, types, layout and comments Type 3: Statements can be changed, added or removed in addition to Type 1 and Type 2 changes Type 4: Code that performs the same computation but is implemented through different syntactic variants The goal of PatchyBFT is to achieve type 4 clones, that is, code that implements the same functionality but through different syntactic means. For example, consider a reference implementation of the Fibonacci function. A type 3 clone might switch around additions, add helper variables, or change types of variables, which is obviously not a significant change and is something we want to avoid. A type 4 clone, on the other hand, might re-implement the function from an iterative to a recursive form. But ensuring that a patch is not a clone is not enough. Compiler optimizations can generate the same machine code even for some type 4 different code. E.g., Rust’s zero-cost abstractions trade compile-time effort for turning higherlevel language features into efficient machine code [50]. Therefore, we also compare the machine code of the original code with the generated assembly of the diversified code. For this, we compile the code once unpatched and with the generated patch. We then compare the generated binary for the function we diversified. We accept the patch only if the function’s assembly differs between the two binaries. Execution Check. The execution check is divided into two phases. First, we introduce a fast and simple check where a single patch is verified alone with unmodified replicas; second, we validate multiple patches together to identify faults that only occur if multiple patches interact with each other. This is inspired by software engineering practices where stabilization branches are used to test multiple patches together for a release [17] and work that shows that testing multiple patches together achieves more cost effective bug detection [62]. In the first execution check, we execute a single replica with the diversified function together with multiple unchanged replicas for a fixed period. We set this period to be long enough for requests to be committed, checkpoints to be created, and (after a deliberate crash of the leader) a view change to happen. Here, we verify that the replica behaves as expected, e.g., it does not crash and participates in the consensus process. In the second phase, the stress test phase, once we have enough patches to get diverse fully patched replicas we test the fully diversified replicas. Again, we ensure that a view change occurs by deliberately crashing the leader. If we identify faults in this step, they could be caused by a single failed patch or by multiple patches that fail together. Patches that fail on their own should have been identified before, but, as the runtime is parameterizable, they might only be identified with the additional time spent testing in this test with four fully diversified replicas. For these single failing patches, we narrow down the potential set of patches by repeatedly bisecting the set into two halves, which we test independently until one patch remains. In the second case, where multiple patches only fail

PatchyBFT : Automating Diversification of Fault-Tolerant Systems using LLMs

in conjunction, we can not identify them with the bisecting routine as we can’t ensure that the set of responsible patches remains in each bisection. Instead, we first identify all the patch-replica pairs that are present in every faulty round, e.g., if we repeat the stress test 100 times and 4 rounds fail, we then find all the patch-replica pairs present in all 4 faulty executions. We use patch-replica pairs, as it matters which patch was on which replica, e.g., it matters if the patch was on the leader replica rather than on the follower replica. Next, we create k-subsets for 𝑘 = 2, .., 10 from the set of the potential patch-replica sets identified and check for all k-subsets whether any specific k-subset occurs in a round without fault. We repeat this with increasing 𝑘 until we find sets of patches which were never part of a round without fault. We execute these subsets again, discarding them from a final deployment if a fault occurs during execution again. After passing these safeguards, we have ensured that the code has been successfully diversified, and we have high confidence that it will behave as expected. We can now use the diversified function.

4.4

How to use PatchyBFT & Lifecycle Management

Software is never fully finished, as new requirements and bugs are constantly found and addressed. We imagine that before the initial deployment of a distributed system PatchyBFT is used to create diversified implementations of the system resulting in some set of patches. There are two scenarios for changes that affect our system, (1) function local changes and (2) more global changes. For (1) local changes, users of PatchyBFT could identify the functions that have changed and discard all previously created patches for that function, making new ones while retaining the patches for all other functions. For (2) global changes, e.g., to some widely used data structures, we recommend discarding all previously created patches. PatchyBFT decreases the cost of diversification by orders of magnitude compared to manually creating diversified implementations (see Section 6), and thus discarding all previously created patches before recreating new ones is feasible. Proactive recovery of replicas can prevent the accumulation of faults by proactively replacing replicas with new versions [19]. Using the cost efficiency of PatchyBFT, it can also be used to continuously diversify a system at runtime. Replicas can be shut down and then restarted with new patches applied.

5

Implementation

We have implemented our approach for the Rust-based BFT implementations PBFT Themis [79], HotStuff [66] and the CFT Raft implementation Openraft [25]. PatchyBFT is not limited to these protocol implementations, and it can be applied to any Rust project which has unit tests (optional) and a system test that verifies that the system is working as expected (e.g., is making progress). As an anecdote, once implemented for Themis, we were able to integrate Openraft and HotStuff within ≈5 hours of work. PatchyBFT makes no special requirements on the LLM for the diversification. For inference we used infrastructure provided by the GWDG [31]. We extracted functions and added context to functions using the parsing tool TreeSitter [90]. Inference on the LLMs was done sequentially, but could also be parallelized.

Themis & Openraft & HotStuff. We have implemented and evaluated PatchyBFT for three different distributed algorithms: PBFT, Raft, and HotStuff. With this we evaluate PatchyBFT for PBFT as a classical BFT protocol, HotStuff as a new blockchain protocol and Raft as a CFT protocol, for each limiting the diversification to the protocol itself. For more comprehensive diversification, all the code could be diversified. Safeguards. We implemented the aforementioned safeguards to validate patches. Each safeguard was executed and the results were stored, i.e., even if the “test” safeguard failed, it was still checked if the patch is a clone or not and whether or not the binary is different. The build & test safeguards are simple cargo build and cargo test commands. For clone detection we used the work of Zhu et al. called MSCCD [105]. With MSCCD we detect type 3 clones. For this we used the default configuration of MSCCD with a detection threshold of 0.7. This can be tuned where a higher threshold increases accuracy, but reduces recall. It should be noted that this is a conservative safeguard, e.g., if 8 lines out of 55 were copied, then the whole function is marked as a clone. Additionally, it should be noted that MSCCD can have both false positives as well as false negatives [105], potentially identifying non-clones as clones. After the clone comparison, we use Bean [45], a binary analysis tool for the Executable and Linkable Format (ELF) [23], to compare the binaries. Here, we compare the generated code for the diversified functions, examining the original and diversified binaries. A single-bit change does not necessarily result in a change in the binary as measured by Bean. Bean skips all NOP instructions, alternative encodings of the same operand (8-bit versus 32-bit displacement) result in the same hash. The hash is position-independent, with branch and RIP-relative targets resolved relative to the function start. The (mangled) name, binding, and section are excluded from the hash. This makes Bean robust to single-bit changes. If a hash over the function’s binary code is different, we consider the diversification successful. For the execution safeguard, we execute the patched code with three unpatched replicas. We expect the overall system to make progress (even if we deliberately crash one of the replicas), i.e., the system should process client requests, and replicas should not crash. If this is the case, we consider the patch successful. For this we execute the replicas and client on one machine. We use one client configured to oversaturate the replicas. Finally, for the stress test safeguard, we used Shadow [49], a discrete-event network simulator with a virtual clock and a deterministic, seeded scheduler, and executed four replicas, each fully diversified from the set of patches that passed all previous safeguards. We configure Shadow to inject faults after a configurable time; its deterministic execution and reproducible fault injection are what let us identify problematic patches, as we detail in Section 6. Diversification of the Implementation. After all safeguards including the stress test, we have a set of patches that are functionally equivalent but are type 4 clones. These patches are then applied to the codebases for diversified replicas. For each function that was diversified, we take a random patch from any of the LLMs and apply it to the codebase. This way, we ensure that the diversified replicas are distinct from one another and that we do not rely on any particular LLM.

Vogel et al.

6

Evaluation

In this section, we evaluate the effectiveness of our approach. We answer the following questions: Qeff How effective is our approach in diversifying code? Qs.g. At what safeguards do patches fail? Qunq How many new (unique) patches are created if we request multiple variants of the same LLM? Qctx What is the effect of specific prompts on the results? Qloc What percentage of a codebase can we diversify? Qstt How do individually verified patches behave in a (longterm) stress test with fully diversified implementations? System Configuration. For the diversification we use the open models of Mistral [88], Qwen3 [89], DeepSeek-R1 [29], as well as the commercial model of Sonnet [3]. We ran the experiment on CloudLab [33] on the d6515 machine type unless otherwise stated. The machines have 32 cores (AMD 7452 at 2.35GHz), 128GB of RAM (8x 16 GB 3200MT/s RDIMMs), and 1TB of disk space. For inference we used inference servers from the GWDG [31]. We have limited our evaluation to functions with at least 10 Lines of Code (LoC). This was done, since with fewer LoC, there is little opportunity for diversification. Diversification of PBFT, HotStuff, and Raft. Across all experiments we have generated 4945 patches for our evaluation. Not all of these patches are successful patches that can be used in production. Figure 5 shows the results of the safeguards on these patches observed. It should be noted that these probabilities are not independent. A patch which correctly builds also has a high chance of passing the tests and passing the runtime check ( 1201 1224 ≈ 99.5% for 454 PBFT, 885 ≈ 89.2% for Raft, and ≈ 97.0% for HotStuff). Similarly, 992 468 the clone detection and binary difference checks are dependent on each other: if one succeeds, the respective other one also passes with a high probability (not as easy to calculate2 ). DeepSeek and Qwen3 automatically generate Chain-of-Thought (CoT) (as outlined in the background Section 2), while Mistral does not. This affects their code generation: the reasoning models explored more diverse solutions, while the non-reasoning model generated more conservative outputs closer to the input. Figure 5 illustrates this trade-off: Mistral produces less diverse code (lower binary/clone safeguard pass rates) but achieves higher correctness (build/test/run pass rates). For comparison we have also investigated the results of a state of the art commercial model against the combined results of the open models. As we see in Figure 6 the commercial model achieves better results across the board, especially for the diversification safeguards. As a fraction, 6.5% of all patches generated by open weight models pass all safeguards compared to 19.4% for the commercial model. For the rest of the evaluation we will use the open-weight models. This gives us a lower bound on the effectiveness of PatchyBFT and also enables the reproducibility of the results. In production we imagine the use of state-of-the-art models to yield even better results than shown here. As runtime failures are quite rare (that is, for patches that build, i.e., if a patch builds then it also has a high chance of passing tests 2 All patches that run can also be built, but patches with a different binary can come

from detected clones (subset of the LoC), and identical binaries can come from type 4 different patches with compiler optimizations.

- self.requests.get(&digest) .and_then(|e| e.sequence) .is_some() + let digest = H::hash_request(&request); + let is_duplicate_request = self.requests.contains_key(&digest);

Listing 3: Excerpt of a modification to the is_duplicate function that caused the program to misbehave.

and passing the runtime safeguard), we highlight one that caused PBFT to misbehave. Listing 3 shows an excerpt of a modification to the is_duplicate function that caused the program to misbehave, which was counted as a crash. In the original code, the function verifies that the request has a sequence number. This is ignored by the patch, which only checks if the request hash is already in the set of requests. At runtime, this causes the replicas to timeout after a view change. This timeout is interpreted as a crash and thus failure of the safeguard. We calculated the cost by multiplying the cost per input and output token per LLM with the average number of tokens used. The cost of generating patches is extremely low, much cheaper than manually implementing them. One single successful patch costs less than 10 cents. As an anecdote, the cost of running all experiments for this paper, including the evaluation and development of the idea with commercial model providers, was less than 50 dollars. PatchyBFT can generate diversified patches

| Qeff Qs.g.

If a patch builds, it has a high chance of passing all safeguards. While state-of-the-art models achieve better results, open-weight models offer a reproducible lower bound In Figure 7 we see the success rate of the safeguards broken down by the lines of code per function diversified. This combines the patches of all LLMs for each protocol. In general, we see no difference in the success rate for the safeguards based on the lines of code. Only for clone detection, we see the success rate decrease with increasing lines of code. This is explained by the fact that in a large function, there are more chances that parts of the function are a clone. As explained in Section 5 we used the default configuration of MSCCD, where even a subset of copied lines marks the function as a clone even if the remaining changes are substantial. For the binary difference, we do not see this effect, as in the remaining 55 changed lines there might be substantial changes. Thus, our configuration might be considered the strict baseline, which could be relaxed based on practical considerations. LLMs can generate patches for large functions | Qeff The success rate of patches decreases with complexity, though this is mostly explained by the specific clone detection settings used. Still, even for big functions with more than 60 lines of code there are successful patches generated. Temperature. We found that the temperature used for LLM has little effect on the number of successful patches generated. Figure 8 shows the number of successful patches generated for different temperatures. We generated patches for PBFT with temperatures 0.0, 1.0, and 2.0. In Figure 8 we do not see significant differences

PatchyBFT : Automating Diversification of Fault-Tolerant Systems using LLMs

80 60 40 20

100

100

deepseek-r1 mistral-large-instruct qwen3-30b-a3b-thinking-2507

80

Patches Passing (%)

deepseek-r1 mistral-large-instruct qwen3-30b-a3b-thinking-2507

Patches Passing (%)

Patches Passing (%)

100

60 40 20

80

deepseek-r1 mistral-large-instruct qwen3-30b-a3b-thinking-2507

60 40 20

0 Build Test Binary Clone Runs

0 Build Test Binary Clone Runs

0 Build Test Binary Clone Runs

(a) PBFT

(b) HotStuff

(c) Raft

Figure 5: Safeguard results for PBFT, HotStuff, and Raft. Each safeguard was tested, even if previous safeguards failed.

60 40 20 0

Build

Test

Binary

Clone

Runs

Figure 6: Results for PBFT with the averaged results of the open models Qwen3, DeepSeek, and Mistral compared to the commercial model Sonnet 4.5 from Anthropic.

between the temperature and the success rate between the different safeguards. This is surprising, especially for higher temperatures, i.e., more “creative” LLMs. We would have expected the clones to decrease and the binary difference to increase with increasing temperature. More research is needed to explain this effect. Temperature: little influence on success rate

| Qeff Qs.g.

Temperature only has a slight effect in either direction on the success rate of safeguards. Ablation study. What are the effects of the prompt on the results? To investigate this, we removed parts of the prompt seen in Listing 4 and showed the results for PBFT. First, we removed the context (data structure, function signatures) from the prompt. We can see in Figure 10a that the success rate of the functional equivalence safeguards (build, test, runs) drops significantly. This is explained by the fact that without the context about the codebase the LLMs hallucinate, e.g., which fields a struct might have. With these hallucinations the code does not compile, or if it does, functions are not used correctly, which results in failed test and run safeguards. Quite paradoxically, these hallucinations also explain the increase in patches that pass through the clone safeguard. The hallucinated code does not exist in the code base, so even though it does not compile, for this test it is a success as it is not a clone.

Barrier Success Rate (%)

80

100

Commercial Open

80

Build Test

Binary Different No Clone

Runs

60 40 20 010-19 20-29 30-39 40-49 50-59 60+ Lines of Code (LoC) Bins

(a) HotStuff: we see the build and runs line overlap as the runtime check is less strict than the test safeguard.

100

Barrier Success Rate (%)

Patches Passing (%)

100

80

Build Test

Binary Different No Clone

Runs

60 40 20 010-19 20-29 30-39 40-49 50-59 60+ Lines of Code (LoC) Bins

(b) Raft: the uptick for the binary difference in the 5059 bin is explained by the small sample size with 9/18 succeeding.

Figure 7: Safeguard results broken down by lines of code bins. Even for big functions with 60+ lines of code we see patches being generated. PBFT is omitted for space reasons, it shows a similar trend.

This also highlights the need for holistic safeguards considering not only diversification but also at functional equivalence. Secondly, we investigated if we need to nudge the LLMs to create diversified code. In Figure 10b we show the results of removing the call to diversify the code from the prompt. As expected without

Vogel et al.

deepseek-r1-temp-0.0 deepseek-r1-temp-1.0 deepseek-r1-temp-2.0 mistral-large-instruct-temp-0.0 mistral-large-instruct-temp-1.0 mistral-large-instruct-temp-2.0 qwen3-30b-a3b-thinking-2507-temp-0.0 qwen3-30b-a3b-thinking-2507-temp-1.0 qwen3-30b-a3b-thinking-2507-temp-2.0

Patches Passing (%)

100 80 60 40 20 0

Build

Test

Binary Different

No Clone Different

Runs

binary-different binary-different + no-clone

1

2 3 4 5 Number of Unique Patches per Function

Patches Passing (%)

50 40 30 20 10 0

Figure 9: Running PatchyBFT multiple times for PBFT results in several patches with all-to-all unique binary + no clone.

the specific constraint for diversified code the LLMs do not create diversified code. More patches fail at the binary check and at the clone detection. This highlights the importance of the prompt on the results. Prompts significantly influence the results | Qctx We have shown how adding context about the code base reduced hallucinations of LLMs making them a more reliable tool for diversification. Total changes. Next, we want to evaluate how much of the original code we can diversify. To illustrate this, we consider the number of functions that could be diversified and the total lines of code that this change affects. In Figure 9 we show that running PatchyBFT multiple times produces multiple unique patches for functions. For this, we generated 12 patches for any PBFT function and tested for patches that passed all safeguards whether or not the binary is different between all patches or if there was a clone. The figure details that we can generate up to 7 unique patches for one function that are not clones and are binary different from each other. This shows us that running PatchyBFT multiple times not only generates a diversified version of the function once, but can also be used to continue generating new patches. We imagine this to be useful for deployments using techniques such as Proactive Recovery [19]. This indicates that we can obtain multiple patches for most functions.

100 80 60

Baseline Ablation 64%

58%

62%

54%

63% 64%

40

64%

52%

30% 31%

20 0

7

Build

Test

Binary No Clone Different

Runs

(a) Without context about data structures and functions within the codebase the LLMs hallucinate about non existing elements in the codebase.

Patches Passing (%)

Functions wwith x Patches

Figure 8: Safeguard results with three different temperature settings for PBFT. Temperature has little effect on the results.

100 80 60

Baseline Ablation 64% 65%

62% 65%

40

37%

20 0

64% 66%

63% 30% 12%

Build

Test

Binary No Clone Different

Runs

(b) Without the prompt to diversify the implementation we see a drop in patches passing the diversification safeguards.

Figure 10: Effects of removing the context and the assignment of diversification from the prompt (baseline) in Listing 4.

Figure 11 shows how many lines of code remain from the original code after applying patches. We managed to replace up to 65% of the lines of code with just two runs of PatchyBFT. In the experiment, we ran PatchyBFT twice to create two batches of patches, for each function we generated 6 patches, two with each LLM. For Raft we saw an increase of 73.8% of unique patches between batch one and two, 14.7% for PBFT, and 32.1% for HotStuff. The comparably low code change percentage for Raft is explained by the programming style of OpenRaft. The library is modelled with extremely large match (Rusts switch-case) constructions [26–28], including

PatchyBFT : Automating Diversification of Fault-Tolerant Systems using LLMs

Code Changed (%)

100

deepseek (Batch 1) deepseek (Batch 2) mistral (Batch 1) mistral (Batch 2) qwen3 (Batch 1) qwen3 (Batch 2) Total Change (Batch 1) Total Change (Batch 2)

80 60 40 20 0 Hotstuff PBFT

Raft

Figure 11: Lines of code replaced from the reference implementation using patches.

unfiltered filtered 600

800 1000 1200 1400 1600 1800

Client measured throughput

Figure 12: Client measured throughput for diversified Themis with patches passing all previous safeguards.

nine functions with more than 100 LoC. None of these 100+ LoC functions was successfully diversified. We can diversify large parts of code

| Qloc Qunq

We can generate multiple patches for most functions, even with a few runs of PatchyBFT. Up to 65% of the lines of code could be replaced with validated patches. Stress test barrier & fault injection. With the stress test barrier, we check if individually validated patches still work together once every replica runs a different, fully diversified implementation even when we inject faults (like a leader crash). The impact of an injected fault can be timing-dependent: The failure of the replicated system might only occur under specific message orderings and only when particular patches interact. On a real testbed such faults could behave as Heisenbugs [40]: they may appear in one run and never recur. This makes it difficult to attribute them to a specific patch or set of patches. We therefore run this experiment in Shadow [49], a discrete-event network simulator that drives the replica processes with a virtual clock and a seeded, deterministic scheduler. Two of its properties are essential to our fault-injection methodology: First, Shadow is deterministic: a given configuration and seed reproduce the exact same sequence of events on every run. This is what makes our fault-attribution procedure sound. We can fix the seed to reproduce a failing schedule exactly. Second, Shadow lets us inject a deliberate leader crash at an identical point in virtual time across every run and every diversified configuration. The view-change path is thus triggered under identical conditions, so the only variable that differs between a

failed and a fault-tolerant execution (in which replicas still make progress despite the leader crash) is the set of applied patches and this is precisely what our attribution requires. Note, we use Shadow only to detect and localize faults for this experiment; later, we measure absolute throughput and latency separately through a system evaluation in a real testbed. We applied patches from the whole set of patches to the PBFT implementation, 95 patches in total, with different patches for each replica depending on a seed. We tested four replicas together, shutting the leader down after 20 seconds, and evaluated the throughput of the system. If the throughput was below a threshold we identified as faulty (e.g., when the system stopped after the view change), then the patches together are identified as faulty. Keeping the seed for Shadow fixed, we repeated this experiment 100 times, each with a different seed for which patches to apply for each execution. We initially identified three faulty patches while executing four differently diversified implementations using Shadow. These were identified by bisecting the set of patches during those executions. We identified that these functions were not tested so far, accordingly we extended the test corpus for these functions. This allowed us to retroactively identify these patches as faulty in the test phase. Next, we reran the Shadow experiment with those three patches filtered out. The results can be seen in Figure 12. The figure shows the client measured throughput. Using the extended testing safeguard and the three faulty patches removed, we see three anomalous throughput numbers. We identified this as three executions where a set of patches caused the view-change to fail. We used the procedure described in the design to identify 17 potential 3-element sets of patch-replica pairs responsible for the fault, which we then tested again in isolation, resulting in 10 verified subsets and 4 excluded patches. Rerunning Shadow with a total of 7 (3 individually identified + 4 of the 3-set patches) of 1224 patches from the unfiltered execution removed from the stress test, we see that all 100 executions succeed in the filtered execution. Performance. With the set of patches that have passed the stress test, we conducted a performance evaluation with fully diversified replicas in a testbed. For this we used c6525-25g machines (16core AMD 7302P at 3.00GHz with 128GB ECC Memory (8x 16 GB 3200MT/s RDIMMs), two 480 GB 6G SATA SSDs and two dual-port Mellanox ConnectX-5 25Gb NICs (PCIe v4.0)) on Cloudlab with 4 replicas and one client, each on their own machine. For replicas, we used a batch size of 100, with a batch timeout of 10ms, all replicas responding to the client, each with 1000 bytes. There is one client with configurable concurrent requests, varying request size, and 300 seconds of benchmark time, measuring throughput and latency. The results can be seen in Figure 13. For the baseline, we used undiversified Themis, for the diversified measurement, we used 10 different seeds, showing the mean of all measurements. As can be seen in Figure 13 the difference between the baseline and the diversified execution is minimal. To be precise, the geometric mean of the absolute differences between the baseline and the diversified execution for each measurement is 0.25% for throughput and 0.53% for the latency. Note on Figure 13: the first three data points had concurrent request of (1, 8, 64) lower than the batch size (100), resulting batches not filling up and only triggered on the batch timeout (10 ms), resulting in increased latency and low throughput.

Vogel et al.

baseline, 0 B payload baseline, 128 B payload diversified, 0 B payload diversified, 128 B payload

Latency (ms)

80 60

fn handle_request( mut request: Message<Request>, mut sender: Sender<RawMessage<ClientTag>>, mut recv: Receiver<Message<Response>>, peers: usize, + faults: usize, ) -> impl Future<Output = Result<Message<Response>>> { - let mut quorum = Quorum::new(peers / 3 + 1, peers); + let mut quorum = Quorum::new(faults + 1, peers); // LLM CoT output: // Actually, the PBFT paper says that the quorum size is f+1 // So if `faults` is f, then quorum size should be faults+1. // ... // Since we are not changing the function signature, we must // use the same quorum parameters. So we will keep // `Quorum::new(peers / 3 + 1, peers)`.

40 20 0

10000

20000

30000

40000

Throughput (requests/s)

40000 30000 20000

Throughput (baseline) Throughput (diversified) Latency (baseline) Latency (diversified)

10000 0

0

5

10

12 10 8 6 4 2 0

Latency (ms)

Throughput (requests/s)

Figure 13: Client measured throughput vs. latency with fully diversified replicas.

15

Elapsed time (h)

20

Listing 4: Human only fix for a wrong quorum in Themis. The LLM noticed the wrong quorum size but was not allowed to fix the bug.

failure in 100 fully diversified executions (Figure 12). Even large functions can be diversified (Figure 7), with up to 7 unique new implementations for one function (Figure 9), and while temperature has little effect (Figure 8), prompt engineering affects the success rate of diversification (Figure 10). Diversification has little effect on performance (Figure 13: 0.25% for throughput, 0.53% for the latency), and runs without any issues in a 24 hour test (Figure 14).

7 Figure 14: Client measured throughput and latency over a 24 hour stress test. Showing a rolling average of 60 seconds.

Stress testing identifies remaining faults | Qstt Patches cannot be evaluated individually; some issues only occur in cross-interactions between patches. With bisecting and common k-set statistics we can identify these without the need for combinatorially many tests. Using this, we achieve 100% success rate during stress tests. Twenty-four-hour test. As a final test, we ran Themis PBFT for 24 hours (24.28h = 87400 seconds, to be precise) in a fully diversified configuration on the five c6525-25g machines setup of the stress test, four for the replicas and one for the client. The results can be seen in Figure 14. The plot shows a 60 second rolling average of client measured throughput (min. 32657, median 37773, and max. 41238) and latency (min. 11ms, median 13ms, and max. 15ms). We observe no fault occurring over the duration of the experiment, processing 3,297,273,515 total requests. We believe that the slight decrease in performance (both for the baseline and the diversified replicas) over time is caused by thrashing [12, 13]. Test of time | Qstt Qeff In a final test running Themis for 24 hours, we did not observe any fault introduced by fully validated patches ordering 3,297,273,515 requests. Summary. We have shown that LLMs can diversify up to 65% with two batches of diversification, without observing a single

Discussion

Bugs that span multiple functions. PatchyBFT focuses on diversifying code for one function at a time. This enables us to address bugs at this level (see Section 3). However, it makes it unlikely that we can fix bugs that span multiple functions. This is shown in Listing 4, where LLMs were unable to fix another bug we obtained for Themis because of the limitation of having to maintain the function signature. In the Chain-of-Thought (CoT) of the LLM we saw that it correctly identified the problem, but since the solution (as seen by the human fix) required a changed signature, it did not implement the fix. For issues that span multiple functions, a promising area for future research is the simultaneous generation of code for multiple functions (e.g., connected functions in the call graph). Concurrency bugs. Concurrency bugs commonly span multiple functions with shared state. Fortunately, by using Rust, we can use the type system to prevent many concurrency bugs. Using static analysis tools, we could identify which functions concurrently access shared state and generate code for these functions together. Proactive Recovery. Orthogonal to the methodology of how to automate diversification using LLMs with PatchyBFT, we can envision a use case for proactive recovery as proposed by Castro and Liskov [19]. In practice, proactively replacing replicas with newly diversified ones could be advantageous to keep the exploitation rate of a system low, as discussed in the attacker model (§4.1).

8

Conclusion

If a replicated system runs the same implementation on every replica, then it is only as safe as that implementation. Thus, a single common bug can lead to a system failure. Diverse protocol implementations can remove a single point of failure, but producing

PatchyBFT : Automating Diversification of Fault-Tolerant Systems using LLMs

them by hand has been prohibitively expensive. In this paper, we propose PatchyBFT which tackles this problem: PatchyBFT uses several LLMs to diversify individual functions into versions that are functionally equivalent but representationally and binary-different, by validating each through build-and-test, clone detection, binary comparison, and a stress test that deterministically injects faults. Our evaluation shows that the overall approach is practical. If a patch builds, it will have a high chance of passing every safeguard, and while commercial models perform best, open-weight models alone already provide a reproducible lower bound. Apart from this, we found that supplying code-base context can further improve the success rate by reducing hallucination. Without developer effort, PatchyBFT diversifies up to 65% of a BFT codebase with up to 7 distinct variants per function. As some faults might only surface in the cross-interaction of patches, we localize and remove them with bisection and common 𝑘-set statistics rather than combinatorial testing. After filtering, we observed no failure across 100 fully diversified executions and in a 24-hour run.

References [1] 2017. CVE-2017-1000430. Available from MITRE, CVE-ID CVE-2017-1000430.. https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2017-1000430 [2] 2019. CVE-2019-16140. Available from MITRE, CVE-ID CVE-2019-16140.. https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2019-16140 [3] Anthropic. 2024. Meet Claude | Anthropic. https://www.anthropic.com/claude. https://www.anthropic.com/claude [4] Anthropic. 2026. Partnering with Mozilla to Improve Firefox’s Security. https: //www.anthropic.com/news/mozilla-firefox-security. Accessed: 2026-04-22. [5] A. Avizienis. 1985. The N-Version Approach to Fault-Tolerant Software. IEEE Transactions on Software Engineering SE-11, 12 (1985), 1491–1501. https://doi. org/10.1109/TSE.1985.231893 [6] Amy Babay, Thomas Tantillo, Trevor Aron, Marco Platania, and Yair Amir. 2018. Network-Attack-Resilient Intrusion-Tolerant SCADA for the Power Grid. In 2018 48th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). 255–266. https://doi.org/10.1109/DSN.2018.00036 [7] Yechan Bae, Youngsuk Kim, Ammar Askar, Jungwon Lim, and Taesoo Kim. 2021. Rudra: Finding Memory Safety Bugs in Rust at the Ecosystem Scale. In Proceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles (Virtual Event, Germany) (SOSP ’21). Association for Computing Machinery, New York, NY, USA, 84–99. https://doi.org/10.1145/3477132.3483570 [8] Shehar Bano, Alberto Sonnino, Andrey Chursin, Dmitri Perelman, Zekun Li, Avery Ching, and Dahlia Malkhi. 2022. Twins: BFT Systems Made Robust. In 25th International Conference on Principles of Distributed Systems (OPODIS 2021) (Leibniz International Proceedings in Informatics (LIPIcs), Vol. 217), Quentin Bramas, Vincent Gramoli, and Alessia Milani (Eds.). Schloss Dagstuhl – LeibnizZentrum für Informatik, Dagstuhl, Germany, 7:1–7:29. https://doi.org/10.4230/ LIPIcs.OPODIS.2021.7 [9] Jeb Bearer, Benedikt Bünz, Philippe Camacho, Binyi Chen, Ellie Davidson, Ben Fisch, Brendon Fish, Gus Gutoski, Fernando Krell, Chengyu Lin, Dahlia Malkhi, Kartik Nayak, Keyao Shen, Alex Xiong, Nathan Yospe, and Sishan Long. 2024. The Espresso Sequencing Network: HotShot Consensus, Tiramisu Data-Availability, and Builder-Exchange. Cryptology ePrint Archive, Paper 2024/1189. https://eprint.iacr.org/2024/1189 [10] Christian Berger, Signe Schwarz-Rüsch, Arne Vogel, Kai Bleeke, Leander Jehl, Hans P. Reiser, and Rüdiger Kapitza. 2023. SoK: Scalability Techniques for BFT Consensus. In 2023 IEEE International Conference on Blockchain and Cryptocurrency (ICBC). 1–18. https://doi.org/10.1109/ICBC56567.2023.10174976 [11] Alysson Bessani, Eduardo Alchieri, João Sousa, André Oliveira, and Fernando Pedone. 2020. From Byzantine Replication to Blockchain: Consensus is Only the Beginning. In 2020 50th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). 424–436. https://doi.org/10.1109/DSN48063.2020. 00057 [12] Alysson Bessani, Marcel Santos, João Felix, Nuno Neves, and Miguel Correia. 2013. On the Efficiency of Durable State Machine Replication. In 2013 USENIX Annual Technical Conference (USENIX ATC 13). USENIX Association, San Jose, CA, 169–180. https://www.usenix.org/conference/atc13/technical-sessions/ presentation/bessani [13] Alysson Bessani, João Sousa, and Eduardo E.P. Alchieri. 2014. State Machine Replication for the Masses with BFT-SMART. In 2014 44th Annual IEEE/IFIP International Conference on Dependable Systems and Networks. 355–362. https:

//doi.org/10.1109/DSN.2014.43 [14] bitfly gmbh. 2024. Ethereum Mainnet Statistics. https://ethernodes.org/. https: //ethernodes.org/ [15] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. arXiv:2005.14165 [cs.CL] https://arxiv.org/abs/2005.14165 [16] Nicholas Carlini, Newton Cheng, Keane Lucas, Michael Mooreand, Milad Nasr, Vinay Prabhushankar, and Winnie Xiao. 2026. Assessing Claude Mythos Preview’s cybersecurity capabilities. https://red.anthropic.com/2026/mythospreview/. Accessed: 2026-04-22. [17] Marco Castelluccio, Le An, and Foutse Khomh. 2019. An empirical study of patch uplift in rapid release development pipelines. Empirical Software Engineering 24 (2019), 3008–3044. https://doi.org/10.1007/s10664-018-9665-y [18] Miguel Castro and Barbara Liskov. 1999. Practical byzantine fault tolerance. In OSDI, Vol. 99. 173–186. https://www.usenix.org/legacy/publications/library/ proceedings/osdi99/full_papers/castro/castro_html/castro.html [19] Miguel Castro and Barbara Liskov. 2002. Practical byzantine fault tolerance and proactive recovery. ACM Trans. Comput. Syst. 20, 4 (nov 2002), 398–461. https://doi.org/10.1145/571637.571640 [20] Miguel Castro, Rodrigo Rodrigues, and Barbara Liskov. 2003. BASE: Using abstraction to improve fault tolerance. ACM Trans. Comput. Syst. 21, 3 (Aug. 2003), 236–269. https://doi.org/10.1145/859716.859718 [21] Frederick B Cohen. 1993. Operating system protection through program evolution. Comput. Secur. 12, 6 (1993), 565–584. https://all.net/books/tech/evolve.pdf [22] Harry Coker. 2024. Back to the Building Blocks: A Path Toward Secure and Measurable Software. Technical Report. White House Office of the National Cyber Director (ONCD). [23] TIS Committee. 2024. Tool Interface Standard (TIS) Portable Formats Specification. https://refspecs.linuxfoundation.org/elf/TIS1.1.pdf. https://refspecs. linuxfoundation.org/elf/TIS1.1.pdf [24] James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson Hsieh, Sebastian Kanthak, Eugene Kogan, Hongyi Li, Alexander Lloyd, Sergey Melnik, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Christopher Taylor, Ruth Wang, and Dale Woodford. 2013. Spanner: Google’s Globally Distributed Database. ACM Trans. Comput. Syst. 31, 3, Article 8 (Aug. 2013), 22 pages. https://doi.org/10.1145/2491245 [25] Databend Labs. 2024. Openraft. https://github.com/databendlabs/openraft. [26] Databend Labs. 2024. Openraft. https://github.com/databendlabs/openraft/ blob/a86e6e75e98361ce5cc38d89c3bee9121dff1396/openraft/src/core/raft_core. rs#L2018. [27] Databend Labs. 2024. Openraft. https://github.com/databendlabs/openraft/ blob/a86e6e75e98361ce5cc38d89c3bee9121dff1396/openraft/src/core/raft_core. rs#L1644. [28] Databend Labs. 2024. Openraft. https://github.com/databendlabs/openraft/ blob/a86e6e75e98361ce5cc38d89c3bee9121dff1396/openraft/src/core/raft_core. rs#L1491. [29] DeepSeek. 2025. DeepSeek-R1. https://github.com/deepseek-ai/DeepSeek-R1 [30] Tobias Distler. 2021. Byzantine Fault-tolerant State-machine Replication from a Systems Perspective. ACM Comput. Surv. 54, 1, Article 24 (Feb. 2021), 38 pages. https://doi.org/10.1145/3436728 [31] Ali Doosthosseini, Jonathan Decker, Hendrik Nolte, and Julian M. Kunkel. 2024. Chat AI: A Seamless Slurm-Native Solution for HPC-Based Services. arXiv:2407.00110 [cs.DC] https://arxiv.org/abs/2407.00110 [32] Mingzhe Du, Anh Tuan Luu, Bin Ji, Qian Liu, and See-Kiong Ng. 2024. Mercury: A Code Efficiency Benchmark for Code Large Language Models. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. https://openreview.net/forum?id=vyraA7xt4c [33] Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson, Kirk Webb, Aditya Akella, Kuangching Wang, Glenn Ricart, Larry Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and Prabodh Mishra. 2019. The Design and Operation of CloudLab. In 2019 USENIX Annual Technical Conference (USENIX ATC 19). USENIX Association, Renton, WA, 1–14. https://www.usenix. org/conference/atc19/presentation/duplyakin [34] Ether Alpha. 2024. Client Diversify | Ethereum. https://clientdiversity.org/. https://clientdiversity.org/ [35] Pedro Fonseca, Kaiyuan Zhang, Xi Wang, and Arvind Krishnamurthy. 2017. An Empirical Study on the Correctness of Formally Verified Distributed Systems. In Proceedings of the Twelfth European Conference on Computer Systems (Belgrade, Serbia) (EuroSys ’17). Association for Computing Machinery, New York, NY, USA, 328–343. https://doi.org/10.1145/3064176.3064183

Vogel et al.

[36] Miguel Garcia, Alysson Bessani, Ilir Gashi, Nuno Neves, and Rafael Obelheiro. 2011. OS diversity for intrusion tolerance: Myth or reality?. In 2011 IEEE/IFIP 41st International Conference on Dependable Systems & Networks (DSN). 383–394. https://doi.org/10.1109/DSN.2011.5958251 [37] Miguel Garcia, Alysson Bessani, and Nuno Neves. 2019. Lazarus: Automatic Management of Diversity in BFT Systems. In Proceedings of the 20th International Middleware Conference (Davis, CA, USA) (Middleware ’19). Association for Computing Machinery, New York, NY, USA, 241–254. https://doi.org/10.1145/ 3361525.3361550 [38] Ilir Gashi, Peter Popov, Vladimir Stankovic, and Lorenzo Strigini. 2004. On Designing Dependable Services with Diverse Off-the-Shelf SQL Servers. In Architecting Dependable Systems II, Rogério de Lemos, Cristina Gacek, and Alexander Romanovsky (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 191–214. https://link.springer.com/content/pdf/10.1007/978-3-540-25939-8_9? pdf [39] Adam Gągol, Damian Leśniak, Damian Straszak, and Michał Świętek. 2019. Aleph: Efficient Atomic Broadcast in Asynchronous Networks with Byzantine Nodes. In Proceedings of the 1st ACM Conference on Advances in Financial Technologies (Zurich, Switzerland) (AFT ’19). Association for Computing Machinery, New York, NY, USA, 214–228. https://doi.org/10.1145/3318041.3355467 [40] Jim Gray. 1986. Why Do Computers Stop and What Can Be Done About It?. In Proc. 5th Symp. on Reliability in Distributed Software and Database Systems. [41] Guide Labs Team. 2026. Steerling-8B: The First Inherently Interpretable Language Model. https://www.guidelabs.ai/post/steerling-8b-base-model-release/ Accessed: 2026-05-05. [42] Suyash Gupta, Jelle Hellings, Sajjad Rahnama, and Mohammad Sadoghi. 2019. An In-Depth Look of BFT Consensus in Blockchain: Challenges and Opportunities. In Proceedings of the 20th International Middleware Conference Tutorials (Davis, CA, USA) (Middleware ’19). Association for Computing Machinery, New York, NY, USA, 6–10. https://doi.org/10.1145/3366625.3369437 [43] Muhammad Hassnain and Caleb Stanford. 2024. Counterexamples in Safe Rust. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering Workshops (Sacramento, CA, USA) (ASEW ’24). Association for Computing Machinery, New York, NY, USA, 128–135. https://doi.org/10. 1145/3691621.3694943 [44] Chris Hawblitzel, Jon Howell, Manos Kapritsos, Jacob R. Lorch, Bryan Parno, Michael L. Roberts, Srinath Setty, and Brian Zill. 2015. IronFleet: proving practical distributed systems correct. In Proceedings of the 25th Symposium on Operating Systems Principles (Monterey, California) (SOSP ’15). Association for Computing Machinery, New York, NY, USA, 1–17. https://doi.org/10.1145/ 2815400.2815428 [45] Bernhard Heinloth. 2024. Bean - Binary Explorer & Analyzer. https://gitlab.cs. fau.de/luci-project/bean. https://gitlab.cs.fau.de/luci-project/bean [46] Larry Huynh, Yinghao Zhang, Djimon Jayasundera, Woojin Jeon, Hyoungshick Kim, Tingting Bi, and Jin B. Hong. 2025. Detecting Code Vulnerabilities using LLMs. In 2025 55th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). 401–414. https://doi.org/10.1109/DSN64029.2025. 00047 [47] Todd Jackson, Babak Salamat, Andrei Homescu, Karthikeyan Manivannan, Gregor Wagner, Andreas Gal, Stefan Brunthaler, Christian Wimmer, and Michael Franz. 2011. Compiler-Generated Software Diversity. Springer New York, New York, NY, 77–98. https://doi.org/10.1007/978-1-4614-0977-9_4 [48] Jay Jacobs, Sasha Romanosky, Benjamin Edwards, Idris Adjerid, and Michael Roytman. 2021. Exploit Prediction Scoring System (EPSS). Digital Threats 2, 3, Article 20 (July 2021), 17 pages. https://doi.org/10.1145/3436242 [49] Rob Jansen, Jim Newsome, and Ryan Wails. 2022. Co-opting Linux Processes for High-Performance Network Simulation. In 2022 USENIX Annual Technical Conference (USENIX ATC 22). USENIX Association, Carlsbad, CA, 327–350. https://www.usenix.org/conference/atc22/presentation/jansen [50] Steve Klabnik, Carol Nichols, and Chris Krycho. 2026. The Rust Programming Language. https://doc.rust-lang.org/book/ch00-00-introduction.html? highlight=zero%20cost#people-who-value-speed-and-stability [51] John C. Knight and Nancy G. Leveson. 1986. An experimental evaluation of the assumption of independence in multiversion programming. IEEE Transactions on Software Engineering SE-12, 1 (1986), 96–109. https://doi.org/10.1109/TSE. 1986.6312924 [52] Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large Language Models are Zero-Shot Reasoners. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 22199–22213. https://proceedings.neurips.cc/paper_files/paper/2022/file/ 8bb0d291acd4acf06ef112099c16f326-Paper-Conference.pdf [53] Per Larsen, Andrei Homescu, Stefan Brunthaler, and Michael Franz. 2014. SoK: Automated Software Diversity. In 2014 IEEE Symposium on Security and Privacy. 276–291. https://doi.org/10.1109/SP.2014.25 [54] Mohsen Lesani, Christian J. Bell, and Adam Chlipala. 2016. Chapar: certified causally consistent distributed key-value stores. In Proceedings of the 43rd

Annual ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages (St. Petersburg, FL, USA) (POPL ’16). Association for Computing Machinery, New York, NY, USA, 357–370. https://doi.org/10.1145/2837614.2837622 [55] Zhuohua Li, Jincheng Wang, Mingshen Sun, and John C.S. Lui. 2021. MirChecker: Detecting Bugs in Rust Programs via Static Analysis. In Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security (Virtual Event, Republic of Korea) (CCS ’21). Association for Computing Machinery, New York, NY, USA, 2183–2196. https://doi.org/10.1145/3460120.3484541 [56] Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and LINGMING ZHANG. 2023. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associates, Inc., 21558–21572. https://proceedings.neurips.cc/paper_files/paper/2023/file/ 43e9d647ccd3e4b7b5baab53f0368686-Paper-Conference.pdf [57] Nuno P. Lopes, Juneyoung Lee, Chung-Kil Hur, Zhengyang Liu, and John Regehr. 2021. Alive2: bounded translation validation for LLVM. In Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation (Virtual, Canada) (PLDI 2021). Association for Computing Machinery, New York, NY, USA, 65–79. https://doi.org/10.1145/3453483.3454030 [58] Noble Saji Mathews and Meiyappan Nagappan. 2024. Test-Driven Development and LLM-based Code Generation. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering (Sacramento, CA, USA) (ASE ’24). Association for Computing Machinery, New York, NY, USA, 1583–1594. https://doi.org/10.1145/3691620.3695527 [59] Nicholas D. Matsakis and Felix S. Klock. 2014. The rust language. In Proceedings of the 2014 ACM SIGAda Annual Conference on High Integrity Language Technology (Portland, Oregon, USA) (HILT ’14). Association for Computing Machinery, New York, NY, USA, 103–104. https://doi.org/10.1145/2663171.2663188 [60] Andrew Meneely, Aiden Green, Tyler Jaafari, Matthew Fluet, and Brandon Keller. 2025. "Just Use Rust": A Best-Case Historical Study of Open Source Vulnerabilities in C. In 2025 IEEE/ACM 3rd International Workshop on Software Vulnerability Management (SVM). 25–32. https://doi.org/10.1109/SVM66695. 2025.00008 [61] Ines Messadi, Markus Horst Becker, Kai Bleeke, Leander Jehl, Sonia Ben Mokhtar, and Rüdiger Kapitza. 2022. SplitBFT: Improving Byzantine Fault Tolerance Safety Using Trusted Compartments. In Proceedings of the 23rd ACM/IFIP International Middleware Conference (Quebec, QC, Canada) (Middleware ’22). Association for Computing Machinery, New York, NY, USA, 56–68. https://doi.org/10.1145/3528535.3531516 [62] Armin Najafi, Peter C. Rigby, and Weiyi Shang. 2019. Bisecting commits and modeling commit risk during testing. In Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (Tallinn, Estonia) (ESEC/FSE 2019). Association for Computing Machinery, New York, NY, USA, 279–289. https: //doi.org/10.1145/3338906.3338944 [63] Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024. Using an LLM to Help With Code Understanding. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (Lisbon, Portugal) (ICSE ’24). Association for Computing Machinery, New York, NY, USA, Article 97, 13 pages. https://doi.org/10.1145/3597503.3639187 [64] Ray Neiheiser, Miguel Matos, and Luís Rodrigues. 2021. Kauri: Scalable BFT Consensus with Pipelined Tree-Based Dissemination and Aggregation. In Proceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles (Virtual Event, Germany) (SOSP ’21). Association for Computing Machinery, New York, NY, USA, 35–48. https://doi.org/10.1145/3477132.3483584 [65] OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2024. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL] https://arxiv.org/abs/2303.08774 [66] ParallelChain Lab. 2025. Performant Byzantine Fault Tolerant State Machine Replication (BFT SMR) in Rust. https://github.com/parallelchain-io/hotstuff_rs. [67] Jinjun Peng, Leyi Cui, Kele Huang, Junfeng Yang, and Baishakhi Ray. 2025. CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation. In 2025 IEEE/ACM International Workshop on Large Language Models for Code (LLM4Code). 33–40. https://doi.org/10.1109/LLM4Code66737. 2025.00009 [68] Marco Platania, Daniel Obenshain, Thomas Tantillo, Ricky Sharma, and Yair Amir. 2014. Towards a Practical Survivable Intrusion Tolerant Replication System. In 2014 IEEE 33rd International Symposium on Reliable Distributed Systems. 242–252. https://doi.org/10.1109/SRDS.2014.16 [69] Boqin Qin, Yilun Chen, Zeming Yu, Linhai Song, and Yiying Zhang. 2020. Understanding memory and thread safety practices and issues in real-world Rust programs. In Proceedings of the 41st ACM SIGPLAN Conference on Programming Language Design and Implementation (London, UK) (PLDI 2020). Association for Computing Machinery, New York, NY, USA, 763–779. https://doi.org/10.1145/ 3385412.3386036

PatchyBFT : Automating Diversification of Fault-Tolerant Systems using LLMs

[70] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9. [71] Alex Rebert and Christoph Kern. 2024. Secure by Design: Google’s Perspective on Memory Safety. Technical Report. Google Security Engineering. [72] Rodrigo Rodrigues, Miguel Castro, and Barbara Liskov. 2001. BASE: using abstraction to improve fault tolerance. SIGOPS Oper. Syst. Rev. 35, 5 (Oct. 2001), 15–28. https://doi.org/10.1145/502059.502037 [73] Tom Roeder and Fred B. Schneider. 2010. Proactive obfuscation. ACM Trans. Comput. Syst. 28, 2, Article 4 (July 2010), 54 pages. https://doi.org/10.1145/ 1813654.1813655 [74] Javier Ron, Diogo Gaspar, Javier Cabrera-Arteaga, Benoit Baudry, and Martin Monperrus. 2025. Galápagos: Automated N-Version Programming with LLMs. ACM Trans. Softw. Eng. Methodol. (Dec. 2025). https://doi.org/10.1145/3785363 Just Accepted. [75] Javier Ron, Zheyuan He, and Martin Monperrus. 2025. Proving and Rewarding Client Diversity to Strengthen Resilience of Blockchain Networks. Distrib. Ledger Technol. (Oct. 2025). https://doi.org/10.1145/3773288 Just Accepted. [76] Chanchal Kumar Roy and James R Cordy. 2007. A survey on software clone detection research. Queen’s School of computing TR 541, 115 (2007), 64–68. https://research.cs.queensu.ca/TechReports/Reports/2007-541.pdf [77] Gloire Rubambiza, Shiang-Wan Chin, Mueed Rehman, Sachille Atapattu, José F. Martínez, and Hakim Weatherspoon. 2023. Comosum: An Extensible, Reconfigurable, and Fault-Tolerant IoT Platform for Digital Agriculture. In 2023 USENIX Annual Technical Conference (USENIX ATC 23). USENIX Association, Boston, MA, 197–214. https://www.usenix.org/conference/atc23/presentation/ rubambiza [78] Signe Rüsch, Kai Bleeke, and Rüdiger Kapitza. 2019. Themis: An Efficient and Memory-Safe BFT Framework in Rust: Research Statement. In Proceedings of the 3rd Workshop on Scalable and Resilient Infrastructures for Distributed Ledgers (Davis, CA, USA) (SERIAL ’19). Association for Computing Machinery, New York, NY, USA, 9–10. https://doi.org/10.1145/3366611.3368144 [79] Signe Rüsch. 2023. BFT Framework and PBFT Implementation. https://github. com/ibr-ds/themis. [80] Carl Sabottke, Octavian Suciu, and Tudor Dumitras. 2015. Vulnerability Disclosure in the Age of Social Media: Exploiting Twitter for Predicting Real-World Exploits. In 24th USENIX Security Symposium (USENIX Security 15). USENIX Association, Washington, D.C., 1041–1056. https://www.usenix.org/conference/ usenixsecurity15/technical-sessions/presentation/sabottke [81] Paulo Sousa, Alysson Neves Bessani, and Rafael R. Obelheiro. 2008. The FOREVER service for fault/intrusion removal. In Proceedings of the 2nd Workshop on Recent Advances on Intrusiton-Tolerant Systems (Glasgow, United Kingdom) (WRAITS ’08). Association for Computing Machinery, New York, NY, USA, Article 5, 6 pages. https://doi.org/10.1145/1413901.1413906 [82] Paulo Sousa, Nuno Ferreira Neves, and Paulo Veríssimo. 2006. Proactive resilience through architectural hybridization. In Proceedings of the 2006 ACM Symposium on Applied Computing (Dijon, France) (SAC ’06). Association for Computing Machinery, New York, NY, USA, 686–690. https://doi.org/10.1145/ 1141277.1141435 [83] Paulo Sousa, Nuno Ferreira Neves, Paulo Verissimo, and William H. Sanders. 2006. Proactive Resilience Revisited: The Delicate Balance Between Resisting Intrusions and Remaining Available. In 2006 25th IEEE Symposium on Reliable Distributed Systems (SRDS’06). 71–82. https://doi.org/10.1109/SRDS.2006.37 [84] Brad Spengler. 2001. PaX: The Guaranteed End of Arbitrary Code Execution. https://grsecurity.net/PaX-presentation.pdf [85] Stack Overflow. 2024. 2024 Developer Survey : AI. https://survey.stackoverflow. co/2024/ai. https://survey.stackoverflow.co/2024/ai [86] Stack Overflow. 2025. 2025 Developer Survey. https://survey.stackoverflow.co/ 2025. https://survey.stackoverflow.co/2025 [87] Jeff Vander Stoep. 2024. Eliminating Memory Safety Vulnerabilities at the Source. https://security.googleblog.com/2024/09/eliminating-memory-safetyvulnerabilities-Android.html [88] Mistral AI team. 2024. Mistral Large 2. https://mistral.ai/news/mistral-large2407 [89] Qwen Team. 2025. Qwen3-Coder: Agentic Coding in the World. https://qwen.ai/ blog?id=qwen3-coder [90] Tree Sitter Authors. 2025. An incremental parsing system for programming tools. https://github.com/tree-sitter/tree-sitter. [91] Ben Vandiver, Hari Balakrishnan, Barbara Liskov, and Sam Madden. 2007. Tolerating byzantine faults in transaction processing systems using commit barrier scheduling. In Proceedings of Twenty-First ACM SIGOPS Symposium on Operating Systems Principles (Stevenson, Washington, USA) (SOSP ’07). Association for Computing Machinery, New York, NY, USA, 59–72. https: //doi.org/10.1145/1294261.1294268 [92] Alexa VanHattum, Daniel Schwartz-Narbonne, Nathan Chong, and Adrian Sampson. 2022. Verifying dynamic trait objects in rust. In Proceedings of the 44th International Conference on Software Engineering: Software Engineering in Practice (Pittsburgh, Pennsylvania) (ICSE-SEIP ’22). Association for Computing

Machinery, New York, NY, USA, 321–330. https://doi.org/10.1145/3510457. 3513031 [93] Alexander Wan, Kevin Klyman, Sayash Kapoor, Nestor Maslej, Shayne Longpre, Betty Xiong, Percy Liang, and Rishi Bommasani. 2025. The 2025 Foundation Model Transparency Index. arXiv:2512.10169 [cs.AI] https://arxiv.org/abs/ 2512.10169 [94] Jitao Wang, Bo Zhang, Kai Wang, Yuzhou Wang, and Weili Han. 2024. BFTDiagnosis: An automated security testing framework with malicious behavior injection for BFT protocols. Computer Networks 249 (2024), 110404. https: //doi.org/10.1016/j.comnet.2024.110404 [95] Xin Wang, Sisi Duan, James Clavin, and Haibin Zhang. 2022. BFT in Blockchains: From Protocols to Use Cases. ACM Comput. Surv. 54, 10s, Article 209 (Sept. 2022), 37 pages. https://doi.org/10.1145/3503042 [96] Zhichao Wang, Kiran Ramnath, Bin Bi, Shiva Kumar Pentyala, Sougata Chaudhuri, Shubham Mehrotra, Zixu, Zhu, Xiang-Bo Mao, Sitaram Asur, Na, and Cheng. 2026. Reinforcement Learning for LLM Post-Training: A Survey. arXiv:2407.16216 [cs.CL] https://arxiv.org/abs/2407.16216 [97] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 24824–24837. https://proceedings.neurips.cc/paper_files/paper/2022/file/ 9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf [98] James R. Wilcox, Doug Woos, Pavel Panchekha, Zachary Tatlock, Xi Wang, Michael D. Ernst, and Thomas Anderson. 2015. Verdi: a framework for implementing and formally verifying distributed systems. SIGPLAN Not. 50, 6 (June 2015), 357–368. https://doi.org/10.1145/2813885.2737958 [99] Hui Xu, Zhuangbin Chen, Mingshen Sun, Yangfan Zhou, and Michael R. Lyu. 2021. Memory-Safety Challenge Considered Solved? An In-Depth Study with All Rust CVEs. ACM Trans. Softw. Eng. Methodol. 31, 1, Article 3 (Sept. 2021), 25 pages. https://doi.org/10.1145/3466642 [100] Boyang Yang, Zijian Cai, Fengling Liu, Bach Le, Lingming Zhang, Tegawendé F. Bissyandé, Yang Liu, and Haoye Tian. 2025. A Survey of LLM-based Automated Program Repair: Taxonomies, Design Paradigms, and Applications. arXiv:2506.23749 [cs.SE] https://arxiv.org/abs/2506.23749 [101] Maofan Yin, Dahlia Malkhi, Michael K. Reiter, Guy Golan Gueta, and Ittai Abraham. 2019. HotStuff: BFT Consensus with Linearity and Responsiveness. In Proceedings of the 2019 ACM Symposium on Principles of Distributed Computing (Toronto ON, Canada) (PODC ’19). Association for Computing Machinery, New York, NY, USA, 347–356. https://doi.org/10.1145/3293611.3331591 [102] J.D. Zamfirescu-Pereira, Eunice Jun, Michael Terry, Qian Yang, and Bjoern Hartmann. 2025. Beyond Code Generation: LLM-supported Exploration of the Program Design Space. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 153, 17 pages. https://doi.org/10.1145/3706598. 3714154 [103] Yuntong Zhang, Jiawei Wang, Dominic Berzin, Martin Mirchev, and Abhik Roychoudhury. 2026. Fixing Security Vulnerabilities with Agentic AI in OSSFuzz. In Proceedings of the 48th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). ACM, Rio de Janeiro, Brazil. https: //doi.org/10.1145/3786583.3786880 [104] Wenbing Zhao. 2007. BFT-WS: A Byzantine Fault Tolerance Framework for Web Services. In 2007 Eleventh International IEEE EDOC Conference Workshop. 89–96. https://doi.org/10.1109/EDOCW.2007.6 [105] Wenqing Zhu, Norihiro Yoshida, Toshihiro Kamiya, Eunjong Choi, and Hiroaki Takada. 2022. MSCCD: grammar pluggable clone detection based on ANTLR parser generation. In Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension (Virtual Event) (ICPC ’22). Association for Computing Machinery, New York, NY, USA, 460–470. https: //doi.org/10.1145/3524610.3529161

Record · ID 978403 · SHA-256 86e7fd1159b9ef50
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.