Detecting Soft Errors in Parallel Software with LLM-tuned Instruction Duplication Yafan Huang
Guanpeng Li
University of Iowa Iowa City, IA, USA [email protected]
University of Florida Gainesville, FL, USA [email protected]
A Parallel A Parallel
loop
Serial Computation
KEYWORDS
Interleaved regions!
Software-directed Reliability, Soft Error Detection, Compiler Code Transformation, Parallel Programs, Large Language Models
1
loop
Synchronization
2.0 1.5 1.0
Lud Heartwall BFS Kmeans
0.5 0.0
Thread (1 -> 16)
Runtime Overhead
arXiv:2609.19531v1 [cs.DC] 17 Sep 2026
Serial Initialization
Normalized Exec. Time
multiplication) [40], limiting their general applicability. In contrast, instruction duplication, a software-directed approach, has emerged as a promising solution [47, 51, 60, 65, 74, 85]. By duplicating instructions at compile time and checking value mismatches at runtime, this technique can effectively detect soft errors in a lightweight, algorithm-agnostic manner.
ABSTRACT We propose PaRID (PaRallel Instruction Duplication), a softwaredirected soft error detection framework that requires only compiletime efforts for multithreading parallel programs. PaRID addresses two key challenges: supporting parallel programs with mixed serial and parallel regions and minimizing performance overhead without relying on costly dynamic profiling. It combines parallel-aware code transformation with LLM-tuned performance modeling, guided by eight generalizable findings from an offline characterization study, to enable fast soft error detection in parallel applications. Evaluation on NPB benchmarks shows that PaRID reduces protection overhead from 162.79% to 59.84% on average and achieves up to 5× speedup while maintaining full error detection effectiveness.
490 Lud 420 Heartwall 350 BFS Kmeans 280 210 214.6161.8 135.2 140 70 8.2 0 Serial Program
Figure 1: Two challenges for performing efficient instruction duplications for multithreading parallel programs. C-1 (left): interleaved serial and parallel regions; C-2 (middle): thread count impacts differently on runtime performance across programs; C-2 (right): existing serial instruction duplication impacts differently on runtime overhead for programs.
INTRODUCTION
With the increasing scale of computing infrastructures, error resilience is becoming a critical concern in not only modern highperformance computing systems [34, 59], but also industrial data centers operated by companies such as Google [39] and Meta [27]. A recent report from Frontier [10], the first exascale system deployed at Oak Ridge National Laboratory, highlights this growing challenge: despite its advancements in energy efficiency, memory hierarchy, and massive concurrency, Frontier yet exhibits 2–3× lower error resilience compared to early terascale systems [15]. Among the various sources of faults, soft errors, also known as transient hardware faults, are particularly insidious [7, 77]. These errors often originate as an unexpected bit-flip in hardware components and can silently propagate through the system stack, ultimately corrupting application outputs without any observable warning [27]. In many real-world applications, producing incorrect computation results can have catastrophic consequences. In smoothed particle hydrodynamics simulation [63], which is a numerical method used to predict extreme flood events, a single bit-flip soft error can propagate through interpolation to ∼100 particles [16], leading to incorrect predictions of flood impact zones, exposing communities to life-threatening risks [62]. As such, soft error detection becomes an essential need. Although various detection techniques have been proposed, many suffer from notable drawbacks. Hardware-based schemes impose high energy costs [49], while algorithm-specific checks are tailored to certain computation patterns (e.g., matrix
In the past, instruction duplication has been studied across various scenarios, spanning CPU and GPU architectures [26, 47, 50, 60, 74, 89] and covering both selective and full duplication [19, 43, 51, 58]. However, a critical gap remains before instruction duplication can be widely adopted in real-world settings: its applicability in multithreading parallel CPU environments. This setting is vital in accelerating many workloads, such as stencil computations in seismic imaging [72]. Achieving effective and efficient instruction duplication for parallel programs is non-trivial and requires addressing two fundamental challenges. C-1: Mixed Serial and Parallel Regions. Unlike GPU programs, which are fully parallelized within GPU kernels, or traditional serial CPU programs, multithreading parallel CPU applications often consist of a mix of serial and parallel regions that may be nested or interleaved during execution [22, 23]. Figure 1(a) illustrates this with the core function lud_omp() for Lud program from Rodinia [18]. As seen, parallel regions, serial regions, and synchronization points are interleaved across this function. An effective duplication strategy must correctly handle this hybrid structure by ensuring threadsafe duplication and managing inter-thread orchestrations. To date, none of the existing duplication techniques thoroughly address these requirements [26, 37, 43, 51, 58]. C-2: Minimizing Performance Loss with Detection. In parallel software, efficiency is often the top priority–parallelism is introduced explicitly to accelerate execution [78]. However, combining instruction duplication
Conference’17, July 2017, Washington, DC, USA 2022. ACM ISBN 978-x-xxxx-xxxx-x/YY/MM. . . $15.00 https://doi.org/10.1145/nnnnnnn.nnnnnnn 1
Conference’17, July 2017, Washington, DC, USA
Yafan Huang and Guanpeng Li
with parallelism introduces significant uncertainty in runtime performance. First, the sensitivity of performance to thread count varies widely across programs; a configuration that performs well for one may be suboptimal for another. In Figure 1(b), increasing the thread count from 4 to 16 yields significant speedups for Lud, Heartwall, and Kmeans, but results in slower execution for BFS. Second, even within the same program, different parallel regions may react differently to thread count—some scale efficiently, while others degrade quickly [29]. Additionally, the overhead introduced by instruction duplication is highly program-specific, ranging from 8.25% in BFS to over 200% in LU (see Figure 1(c)), and becomes even more elusive in parallel programs. Existing approaches often rely on trial-and-error to identify optimal thread settings [61]. However, profiling such information demands substantial real-time effort, an impractical burden in production, where workloads routinely run from hours to months [76]. Repeated executions with different thread counts can be prohibitively expensive, making traditional tuning methods infeasible for large-scale applications. To tackle these challenges, we propose PaRID (PaRallel Instruction Duplication). Given an arbitrary parallel program code, PaRID uses two independent steps to generate the duplicated executable binary. While Parallel-aware Code Transformation instruments duplicated and other functional instructions safely for parallel programs, LLMtuned Performance Modeling automatically locates the best thread settings, no matter how large the workload is, with prompt-refined large language model (LLM) inferences, both of which require only compile-time efforts without dynamic profiling. To the best of our knowledge, this is not only the first work to comprehensively propose an effective instruction duplication for parallel programs but also the first work to predict dynamic program features with LLM in software error resilience. Some key results of PaRID can be found as follows.
fault model is consistent with existing works that study soft error detection mechanisms [47, 51, 54, 65, 73, 74]. We utilize LLVM compiler [52] for performing instruction duplication and fault injection due to its standard optimizer and support for multiple programming languages and hardware platforms. LLVM features a human-readable intermediate representation (IR) with instruction-like characteristics, enabling users to perform various transformations. Additionally, LLVM supports optimizations for OpenMP [22] starting from v11 [1]. Using LLVM compiler and IR to explore program error resilience has been widely adopted in existing reliability studies [31, 42, 47, 50, 51, 60, 79, 87].
2.2
The Principle of Instruction Duplication
Instruction duplication is a soft error detection technique that requires only compiler-level program-agnostic code transformation. Figure 2(a) illustrates the protection workflow of instruction duplication. After the compiler compiles the source code into an instruction format, instruction duplication instruments replicas, generating protected code. This protected code is finally compiled into an executable binary for deployment. Instruction duplication is a flexible software-level solution that requires only static code transformation at the compiler level, ensuring this technique is program-agnostic and does not require expensive dynamic profiling. As a result, instruction duplication has been actively studied over the past two decades [47, 51, 58, 60, 65, 74] and successfully applied in real-world safety-critical missions, such as the faulttolerant systems used in Stanford’s Advanced Research and Global Observation Satellite (ARGOS) project [56]. %0 = load ptr %adr
Source Code
• PaRID ’s Parallel-Aware Code Transformation exhibits high compatibility across all 15 Rodinia benchmarks and 8 NPB benchmarks, under diverse thread counts and input sizes. • We summarize 8 Findings from an offline thread-sensitivity study to guide LLMs in selecting optimal thread settings. These findings are generalizable across different machines. • Without thread tuning from LLM, instruction duplication for parallel programs incurs an average execution time of 162.79% over serial execution. PaRID reduces this to 59.84%, achieving up to ∼5× speedup (on NPB EP). • All benefits require only compile-time efforts and do not sacrifice error detection effectiveness.
%0'= load ptr %adr
Compile
Code at Instruction-level Duplicate
Protected Code at Instruction-level Compile
Executable Binary (a) Protection Workflow of Instruction Duplication
%1 = add %0 , 1 %0 = load ptr %adr %1 = add %0 , 1 %2 = mul %1 , 2
%1'= add %0', 1 %2 = mul %1 , 2 %2'= mul %1', 2
Original Inst.
%3 = icmp %2, %2'
Duplicated Inst.
%2==%2'
Functional Inst.
Normal Execution
%2!=%2'
Reporting Errors
(b) Transforming Original Code to Protected Code by Instruction Duplication (Duplicate Stage in (a))
Figure 2: Illustration of instruction duplication. Figure 2(b) further details a code example to explain the principle of instruction duplication. Given a basic block1 with three instructions, instruction duplication inserts replicas that use different registers (e.g. %0 to %0’), forming a separate but functionally identical dataflow. A comparison instruction (e.g. icmp) is added at the end of the two dataflows to verify consistency. If the values match, execution proceeds to the next basic block with normal execution. Otherwise, an error is reported, triggering recovery mechanisms such as checkpoint/restart [44]. In this work, we target full protection – duplicating all eligible instructions – whereas our approach can be easily extended to support selective protection [51, 58].
2 BACKGROUND 2.1 Fault Model and Compiler Platform In this work, we focus on soft errors that occur in the processor’s computational pipeline (i.e., datapath units), including arithmetic logic units, pipeline stages, and load/store units. We exclude errors in memory and cache units from our scope, as they can be effectively mitigated by techniques such as error-correcting codes (ECC) [8, 20, 67, 82] and parity checks [69]. Similarly, we exclude errors that corrupt control-flow units (e.g. program counter) since software signatures [30, 64] can handle them. Soft errors manifesting at the program level are considered single bit-flips, as the occurrence rate of multiple bit-flips is relatively low in existing systems [14]. Our
1 Basic block is a compiler concept and consists of multiple sequentially executed
instructions. Basic blocks form the vertices in control-flow graph. 2
Detecting Soft Errors in Parallel Software with LLM-tuned Instruction Duplication
3
loop parallelized using OpenMP and examine its corresponding instruction-level representation. The source code example is shown in Figure 5(a), and its instruction format (in control-flow graph) can be found in Figure 5(c). As we can see, while this simple loop has only 5 basic blocks in serial implementation, even with a canonical loop format, its OpenMP version results in 12 basic blocks. After characterizing each instruction in detail, we find three main reasons. (1) Loop Transformation: In order to execute programs in parallel, OpenMP transformed the loop into a canonical form suitable for runtime scheduling, where loop bounds, strides, and iteration variables are managed explicitly, resulting in more instructions and more basic blocks. (2) Support for Edge Cases: OpenMP must deal with edge cases like empty loops or loops with unusual bounds, adding conditional branches to check for these cases. (3) Parallel Execution Support: OpenMP handles shared and thread-private variables using different strategies, including parallel, serial, and atomic computations, often relying on external libraries to ensure correct execution. Although (1) and (2) introduce additional instructions, they always execute in a thread-safe manner and have minimal impact on instruction duplication. Consequently, a key requirement for instruction duplication in OpenMP is to handle (3) properly.
INITIAL STUDY
The goal is to propose an effective and efficient instruction duplication technique for parallel software, with only compiletime effort. To ensure effectiveness, instructions must be safely duplicated across mixed serial and parallel regions, addressing C1. To ensure efficiency, the system must provide thread configurations with low runtime overhead, not only from the program itself but also from the duplication, addressing C-2. Importantly, the proposed method should avoid reliance on dynamic profiling (i.e., features that require program execution), which is often impractical in production where applications may run for extended periods [16, 42, 72]. In this section, we conduct an initial study to better understand these challenges; the key insights derived from this study directly inform the design of our proposed solution. In this work, we adopt OpenMP [22] as the parallel interface. OpenMP is an annotation-based parallelization API for multiple programming languages, including C/C++ and Fortran, making it popular for parallel computing tasks in both academia and industry [11, 53]. A recent work [46] also reports that OpenMP occupies 45% of recent parallel analyses. Figure 3 shows an example of paralleling a loop using an OpenMP directive #pragma omp parallel. for(int i=0; i<N; i++) expensive_computation1();
for(int i=0; i<N; i++) expensive_computation1();
for(int i=0; i<N; i++) expensive_computation2();
#pragma omp parallel for for(int i=0; i<N; i++) expensive_computation2();
Serial Region
#pragma omp barrier
Parallel Region
Insight I: The key requirement for instruction duplication in OpenMP programs is managing parallel computation patterns, such as shared variables and atomic operations.
void omp_example(){ int array[2048];
Figure 3: Parallel a for loop using OpenMP directive.
3.1
#pragma omp parallel for for(int i=0; i<2048; i++) array[i] = i * 2;
Static Analysis of OpenMP Programs
Parallel
Num. of Insts.
Num. of BB
Serial
3000 2500 2000 1500 1000 500 0
Serial
entry
Compile omp.precond.then
Parallel Region
Unlike serial code, OpenMP programs consist of interleaved parallel and serial regions, making existing instruction duplication techniques inadequate [42, 58]. We demystify such discrepancy by performing static analysis on eight programs from the Rodinia benchmark suite [18]. Specifically, we compare the number of basic blocks and instructions in serial and OpenMP-parallelized versions, compiled using the same LLVM toolchain. As shown in Figure 4, the parallel versions consistently introduce more complexity. On average, basic block count increases by 26%, from 134.13 to 169.25, with a similar trend observed in instruction count (from 970.35 to 1324.38). These additional instructions often involve parallel constructs, such as synchronization, data sharing, and thread management [71], constraining existing instruction duplication in parallel programs. 400 320 240 160 80 0
Conference’17, July 2017, Washington, DC, USA
cond.true
}
(a) A toy example of using OpenMP directive to parallel a for loop
cond.end
void lud_omp(float *a, int size){ ...
omp.inner. for.cond
for(/*loop condition*/) #pragma omp parallel for for(/*loop condition*/) for(/*loop condition*/) #pragma omp simd for(/*loop condition*/) ... #pragma omp parallel for for(/*loop condition*/) ... for(/*loop condition*/) #pragma omp simd for(/*loop condition*/) #pragma omp simd for(/*loop condition*/) ...
Parallel }
(b) Parallel Rodinia Lud benchmark with OpenMP
e d S e d S p s r r p s r r KNNNeedl Lu BF ckprmo eanhfindelefilte KNNNeedl Lu BF ckprmo eanhfindelefilte Ba K Pat Partic Ba K Pat Partic Figure 4: The number of static basic blocks (left) and instructions (right) for 8 Rodinia serial/parallel programs [18].
cond.false
omp.inner. for.body omp.body. continue omp.inner. for.inc omp.inner. for.end omp.loop. exit omp.precond.end
(c) Mapping Parallel Region from (a) into LLVM IR Control-flow Graph
Figure 5: Illustration of paralleling programs with OpenMP directives and how OpenMP parallel regions reflect at LLVM. In real-world scenarios, users may utilize diverse OpenMP APIs in complex ways, such as frequent synchronization, master/slave thread divergence, and nested directives. Figure 5(b) illustrates the parallel region of the Rodinia [18] Lud benchmark. As seen, multiple
To further understand how OpenMP parallelization leads to increased instruction-level complexity, we begin with a simple for 3
Conference’17, July 2017, Washington, DC, USA
Yafan Huang and Guanpeng Li
directives are applied at different loop levels, alongside synchronization barriers, making it difficult to distinguish parallel and serial regions at the source code level. However, after we compile this code into instructions, we have two key observations. (1) LLVM compiler only treats the outermost OpenMP directive as a separate function call, making it clear to identify parallel regions. (2) Nested directives follow similar computation strategies (e.g. synchronization barrier and atomic operations) to its outermost one.
Figure 6 presents the results. At first glance, except Backprop and Particlefilter, the LLM struggles with accurate runtime prediction for most programs. For instance, at 8 threads, KNN has an actual normalized runtime of 1.01 (i.e., no speedup), but the LLM incorrectly predicts 0.147 (suggesting an 8× speedup). However, the LLM consistently gets it right while predicting trends of performance change from 8 to 16 threads. Interestingly, the LLM also provides rationales alongside its predictions. For example, in BFS, it correctly anticipates a slowdown when increasing from 8 to 16 threads, citing the overhead of implicit barriers in loop-based parallelism. The only exception is Pathfinder. We find that Pathfinder’s parallel region contains only a short 7-line for loop, with most of the execution time dominated by serial code and output operations, which are difficult for LLMs to reason about. However, once this context is explicitly included, LLM can produce accurate predictions: consistent performance across thread counts within Pathfinder.
Insight II: We distinguish parallel and serial regions in OpenMP programs by identifying the outermost directives. In this work, we define an outermost OpenMP directive as an OpenMP Kernel for simplicity.
Dynamic Prediction for Programs via LLM
Recall that the other challenge is performance-aware error detection. Thread count has a significant impact on the runtime behavior of parallel programs, even without any fault tolerance mechanism (see Figure 1(b)). When combined with instruction duplication, this impact becomes even more unpredictable [43]. A promising error detection technique must therefore introduce replicas in a way that maintains both fault coverage and minimal runtime overhead, ideally by selecting the best thread configuration. While dynamic profiling could help identify such configurations, it is often expensive in real-world workloads [16, 42, 72]. To avoid this cost, we seek a solution that relies solely on compile-time analysis. Recently, large language models (LLMs) have shown strong potential for code analysis, such as compiler optimization [21] and dynamic profiling [17], making them a promising tool in our scenario. To evaluate the potential of LLMs for dynamic analysis in parallel programs, we conduct an initial study using Phi-3.5-mini-instruct, a lightweight LLM with 3.82 B parameters released by Microsoft [3]. Compared to its larger counterpart, Phi-3.5-MoE, it achieves competitive performance on code tasks (61.5 vs. 70.7) while requiring 15× fewer parameters, making it feasible for local inference on a single GPU. In our setup, we extract the code corresponding to each program’s parallel region as the initial prompt. We then provide runtime data for 1, 2, and 4 threads, and ask the model to predict runtimes for 8 and 16 threads. Initially, we found that predictions based on normalized runtime performed poorly, so we switched to raw runtime values (in seconds). For clarity, we still report results in normalized format (with the serial runtime set to 1). We reuse the 8 parallel programs from Rodinia, as in Section 3.1.
Insight III: While LLMs may not predict exact runtimes, they can infer performance trends from code based on prior knowledge. Furthermore, prompt refinement significantly improves prediction quality in challenging cases.
4
OUR SOLUTION: PARID
In this work, we propose PaRID (PaRallel Instruction Duplication), a soft error detection framework for OpenMP-based parallel software. PaRID operates entirely at compile-time, requiring only static code transformation and a few lightweight, device-side LLM inferences, while maintaining high error detection effectiveness and low runtime overhead. Given an arbitrary parallel program, PaRID performs two independent steps to generate a protected executable binary. First, after compiling the source code to LLVM IR, PaRID applies Parallel-aware Code Transformation (➊) to produce replicainstrumented IR that is compatible with both serial and parallel regions, addressing C-1. Second, based on an offline characterization study, PaRID utilizes learned prompts to guide the LLM-tuned Performance Modeling (➋) in inferring optimal thread settings for each OpenMP kernel, avoiding expensive dynamic profiling and addressing C-2. The entire PaRID pipeline is fully automated through Python scripts, requiring no user intervention, and can be seamlessly integrated into production workflows. In the following context, we will explain both ➊ and ➋ in detail. Compile to IR Input Original Parallel Program Code
Original Predicted
1.8 1.5 1.2 0.9 0.6 0.3 0.0
Normalized Exec. Time
Normalized Exec. Time
1.2 0.9 0.6 0.3 0.0
Original Predicted
with Learned Prompt (Generated Offline)
1 Parallel-aware Code Transformation
2
LLM-tuned Performance Modeling
IR Code with Duplications Output Compile
3.2
Protected Executable Binary
Thread Setting for Parallel Regions
Figure 7: Workflow of PaRID framework.
e d S e d S p s r r p s r r KNNNeedl Lu BF ckpromeanhfinde lefilte KNNNeedl Lu BF ckpromeanhfinde lefilte Ba K Pat Partic Ba K Pat Partic Figure 6: Original vs LLM predicted parallel program execution time for thread count 8 (left) and 16 (right).
4.1
Parallel-aware Code Transformation (➊)
Figure 8 illustrates how Parallel-aware Code Transformation in PaRID instruments duplication along with functional instructions 4
Detecting Soft Errors in Parallel Software with LLM-tuned Instruction Duplication in a thread-safe manner. Given the original program code in IR, PaRID first performs Identifying of Parallel Regions by localizing each OpenMP kernel (as described in Insight II, Section 3.1). At the instruction level, this is achieved by matching the keyword omp_outlined in function names. For serial regions, PaRID downgrades to traditional instruction duplication, directly instrumenting both duplicated and functional instructions2 . For parallel regions, PaRID conducts Annotation of Eligible Instructions and Dataflow Analysis for Parallel Regions. These two steps enable PaRID to properly handle parallel computation patterns (as mentioned in Insight I, Section 3.1). Finally, with target dataflows contained by candidate instructions, PaRID instruments the replicas and functional instructions, generating the protected IR. We implement the entire code transformation within a single LLVM pass, ensuring the usability. Original Program Code
Parallel Regions
Identifying Parallel Regions
Candidate Instructions
Annotating Eligible Instructions
Target Dataflows
Dataflow Analysis for Parallel
Serial Regions
requires verifying error occurrences. Instead, they act as control mechanisms to ensure proper execution order and mutual exclusion among threads [22]. Duplicating such calls could interfere with the correct functioning of the OpenMP runtime.
Compile
Figure 9: Atomic operations at LLVM IR. Note that some atomic operations in OpenMP programs reflect as an atomicrmw instruction. One example is shown in Figure 9. In this case, multiple threads perform read-modify-write operations (add %value) to the same shared address (%sum), where return register denotes the value of shared register before the computation. We do not duplicate such instructions because doing so would result in mismatched execution orders between two copies and a false error report. Instead, we apply a pre-checking mechanism for all operands involved in the atomicrmw instruction, ensuring values passed in is error-free. Specifically, an additional comparison instruction is inserted to validate the operands before the atomic operation executes. While this approach slightly breaks the dataflow and incurs minor additional overhead, it is justified because atomicrmw operations are relatively infrequent in most workloads [18]. Most importantly, it ensures no loss of fault coverage effectiveness in PaRID.
Protected Program Code Inserting Functional Instructions
Duplicated Instructions
Instruction Duplication
Conference’17, July 2017, Washington, DC, USA
Key Designs in PaRID
Figure 8: Workflow of Parallel-aware Code Transformation in PaRID, where the input and output are both in LLVM IR. After identifying parallel regions, the design of PaRID for handling OpenMP parallel regions focuses on two key components: Annotating Eligible Instructions and Dataflow Analysis for Parallel Regions. The former prunes ineligible instructions and appropriately annotates and processes OpenMP-specific instructions, whereas the latter extracts program dataflows in a thread-safe manner. Together, these components facilitate the subsequent steps of instruction duplication and insertion of functional instructions for OpenMP.
4.1.2 Dataflow Analysis for Parallel Region. In PaRID, we perform duplication and instrument a functional comparison instruction (icmp) at the end of each static data dependency sequence (SDDS), defined as a chain of instructions where the output of one instruction is used as an operand by another [54, 68]. This approach avoids performing comparisons after every pair of duplicated instructions, which would introduce extra jumps and disrupt the program’s control flow. In OpenMP programs, implementing this requires PaRID to analyze program dataflow and ensure that each SDDS is identified in a thread-safe manner. The key strategy in this step is to analyze instruction operand usage while also extracting each SDDS immediately before accessing a shared region, particularly within critical sections, where only one thread executes simultaneously.
4.1.1 Annotating Eligible Instructions. In addition to computational instructions (identical to those in the serial implementation), LLVM generates extra parallel-supporting instructions in each OpenMP kernel. As noted in Section 3.1, instructions related to loop transformations (e.g., %.omp.stride) and edge cases supporting (e.g., %omp.precond) are thread-safe. Among the remaining instructions, we identify two types that are not eligible for duplication. (1) Write to Shared Region. Parallel access to a shared array is typically translated into store operations targeting independent indices. In PaRID, these store operations are excluded from duplication. We reckon this exclusion is justified, as the values being stored are already duplicated and verified in thread-preserved regions prior to storing. Indiscriminately duplicating these write-to-memory instructions could disrupt memory consistency or lead to unnecessary overhead, especially when handling shared memory. Furthermore, our fault model assumes that memory errors are corrected by hardware mechanisms such as ECC, rendering additional duplication redundant for these operations. (2) OpenMP Runtime Calls. We also prune such function calls, such as kmpc_barrier for synchronization and kmpc_critical for managing critical regions, for duplication. These runtime calls do not directly impact computation or dataflow, which
Static Data Dependence Sequence 1 #pragma omp parallel for for(int i=0; i<n; i++){ ... int tmp1 = i*(*shared1); *shared1 = tmp1 + 1; int tmp2 = 4*(*shared2); *shared2 = tmp2 + 2; ... }
(a) Source Code
Static Data Dependence Sequence 2
... %1 = load ptr %i %2 = load ptr %shared1 %3 = mul %1 , %2 %4 = add %3 , 1 store %4, ptr %shared1 %5 = load ptr %shared2 %6 = mul %5 , 4 %7 = add %6 , 2 store %7, ptr %shared2 ...
(b) Instruction Set
%1
%2
%5
%3
%6
%4
%7
Accessing Shared Region1
Accessing Shared Region2
(c) Dataflows
Figure 10: Illustrating Dataflow Analysis for Parallel Regions. Figure 10 illustrates how PaRID recognizes SDDSs in an OpenMP parallel for loop. After compiling the source code (Figure 10(a)) into its instruction set representation (Figure 10(b)), the basic block is divided into two distinct SDDSs, each concluding with a store operation to write data to a shared region. At the dataflow level (Figure 10(c)), PaRID will duplicate each SDDS and compare the
2 Functional instructions in instruction duplication (see Figure 2) refer to the ones
inserted to handle mismatch detection (e.g., comparisons and conditional branches), error reporting (e.g., function calls), and optional recovery triggering [43, 51, 58]. 5
Conference’17, July 2017, Washington, DC, USA
Yafan Huang and Guanpeng Li
results between %4 and %4’ as well as %7 and %7’ in later instrumentation steps, ensuring thread safety while effectively detecting errors. Since LLVM IR is in SSA form [52], operand tracing and SDDS extraction can be done efficiently using def-use chains [48].
4.2
the serial execution of the unprotected program (set to 1.0). Unlike prior works [43, 47, 51], we do not use the traditional overhead metric (i.e., Protection/Original under the same thread count), as changes in thread configuration can naturally reduce runtime. In such cases, a naive overhead comparison may misleadingly suggest improved performance even when duplication introduces cost. All experiments are conducted on a 28-core machine.
LLM-tuned Performance Modeling (➋)
Original Code Identifying Parallel Regions
Thread Settings Code with Parallel Regions Noted
Learned Prompts (i.e. Findings) LLM Engine
4.2.2 Thread Sensitivity Study. Figure 12 presents the results of PaRID Thread Sensitivity Study, demonstrating PaRID ➊ can be seamlessly integrated for parallel programs. We report results for 14/15 benchmarks; Pathfinder is excluded, as its runtime is dominated by serial computations (as shown in Section 3.2). All benchmarks use the standard input provided by Rodinia [2]. For each benchmark, we report normalized execution time under thread counts {1, 2, 4, 8, 16, 28, 32}, where 28 corresponds to the machine core count. We include 32 threads to observe the impact of context switching (i.e., oversubscription [45]). Rather than identifying the "best" thread count for each benchmark, which is highly machinespecific, we extract generalizable findings from the data, guiding LLM inference in a portable and system-agnostic manner. Finding 1: If a benchmark has minimal OpenMP regions and large serial sections, then it will exhibit consistent runtime across thread settings. Besides Pathfinder, it is also observed in B+Tree, where only two for loops are parallelized, while ∼2,500 lines of complex loop structures remain in the serial region. As a result, PaRID exhibits nearly identical execution time across thread counts. While long serial code may challenge customer-side LLM inference, this property can be easily identified through simple static analysis using LLVM passes. In contrast, Hotspot also shows consistent runtime across threads, but the cause is different: its standard input is too small, making the runtime dominated by serial preprocessing and postprocessing. However, such input size issues are less relevant in real-world workloads; we explore this further in Section 5. Finding 2: If the OpenMP kernel is embarrassingly parallel (e.g., for loop with no dependencies), then higher thread counts can be applied for PaRID. This is a commonly observed pattern, seen in Heartwall, LavaMD, Lud, Kmeans, and Particlefilter, all of which achieve peak performance with 16 threads. These benchmarks follow a data-parallel model [38], distributing loop iterations evenly across threads without requiring explicit thread boundaries. Figure 13 shows two examples of embarrassingly parallel OpenMP kernels. In Lud, although the code features more complex or interleaved control flow, PaRID’s Identifying Parallel Regions step can still easily recognize these as embarrassingly parallel kernels. Finding 3: If thread count nears the core count, then PaRID ’s per-thread workload causes the runtime to increase (especially from complicated control-flow). This behavior is consistently observed in 13/15 benchmarks. Recall in Section 3.1, we find that OpenMP naturally introduces additional control-flow and parallel-support instructions. To detect soft errors, PaRID ➊ further inserts replicated and functional control-flow operations, resulting in performance degradation, especially approaching the hardware core limit. Finding 4: If the OpenMP kernel has frequent synchronizations, then PaRID runs best with moderate threads. This behavior is observed in KNN. When an OpenMP kernel includes frequent barriers and critical regions, thread-level waiting becomes unavoidable.
Offline Thread Stage Sensitivity Study
Figure 11: Workflow of LLM-tuned Performance Modeling. Recall that parallelism is inherently introduced to accelerate program execution [35, 78]. As such, PaRID must reduce the runtime overhead introduced by error detection mechanisms (i.e., ➊) to preserve the performance benefits of parallelism. Both thread count (see Figure 1(b)) and instruction duplication overhead (see Figure 1(c)) impact runtime differently across programs and regions. A natural solution is to identify the optimal thread setting for each parallel region, reducing performance loss. Additionally, to ensure usability in real-world settings, this tuning should be done entirely at compile time. Inspired by Insight III in Section 3.2, we leverage the potential of LLMs to infer dynamic performance characteristics from code, especially when guided by well-constructed prompts. PaRID adopts this strategy through LLM-Tuned Performance Modeling, which automatically selects the best thread configuration for each OpenMP kernel in the input program. The idea of this step is as follows: We transform a set of widely used parallel benchmarks using ➊ in PaRID, conduct a detailed threadlevel characterization study, and extract key findings to serve as prompts for guiding LLM inference. Figure 11 illustrates the workflow of LLM-Tuned Performance Modeling in PaRID. In the offline stage, we perform a one-time Thread Sensitivity Study across 15 Rodinia benchmarks [18] to identify performance trends under different thread counts. From this study, we extract several generalizable findings, which are encoded as prompt components for the LLM. In the online stage, given a new input program, PaRID first applies Identifying Parallel Regions (as in ➊) to locate each OpenMP kernel. Then, the LLM Engine analyzes each kernel and recommends a thread configuration based on the previously learned prompt strategies. By doing so, PaRID shifts the expensive dynamic analysis to the offline phase, making online prediction both lightweight and accurate. We describe the details of the Thread Sensitivity Study. 4.2.1 Setups before Characterization. We use 15 benchmarks from the Rodinia suite [18] (v3.1) [2], which provides both serial and parallel implementations and covers diverse computation patterns across multiple domains. There are 19 benchmarks in total, and we try to include all of them. However, we exclude 4 benchmarks due to incompatibility with our toolset. For example, current LLVMOpenMP project [1] does not support the required cross-machine dependencies in Leukocyte. Rodinia is particularly well-suited for our study, as most of its benchmarks are lightweight and contain either a single OpenMP kernel or multiple structurally similar kernels, simplifying analysis. To evaluate the impact of instruction duplication, we report normalized runtime, where the baseline is 6
Detecting Soft Errors in Parallel Software with LLM-tuned Instruction Duplication 4.0 3.0 2.0 1.0 0.0 1 2 4 8 16 28 32
1.5
Conference’17, July 2017, Washington, DC, USA
0.0 1 2 4 8 16 28 32
1.5 1.0 0.5 0.0 1 2 4 8 16 28 32
7.5 5.0 2.5 0.0 1 2 4 8 16 28 32
3.0 2.0 1.0 0.0 1 2 4 8 16 28 32
(a) Hotspot3D
(b) Srad
(c) LavaMD
(d) KNN
1.5 1.2 0.9 0.6 0.3 0.0 1 2 4 8 16 28 32
1.5 1.0 0.5 0.0 1 2 4 8 16 28 32
2.5 2.0 1.5 1.0 0.5 0.0 1 2 4 8 16 28 32
(h) BFS
(i) B+Tree
(j) Backprop
1.0 0.5
3.0
1.0
2.0
0.5
1.0
0.0 1 2 4 8 16 28 32
0.0 1 2 4 8 16 28 32
(e) Needle
(f) Hotspot
(g) Lud
1.5 1.0 0.5 0.0 1 2 4 8 16 28 32
2.5 2.0 1.5 1.0 0.5 0.0 1 2 4 8 16 28 32
2.5 2.0 1.5 1.0 0.5 0.0 1 2 4 8 16 28 32
2.5 2.0 1.5 1.0 0.5 0.0 1 2 4 8 16 28 32
(k) Myocyte
(l) Kmeans
(m) Heartwall
(n) Particlefilter
Figure 12: Thread Sensitivity Study using Rodinia benchmarks with PaRID ➊ transformation. For each benchmark, x-axis is the number of threads; y-axis is the normalized runtime. Pathfinder is excluded since it is dominated by CPU computations. #pragma omp parallel for for(x=0; x<Nparticles; x++){ weights[x] *= exp(likelihood[x]); }
per-thread workload exacerbates the issue. In such cases, using fewer than the core-limit threads yields better performance.
#pragma omp parallel for for(x=0; x<Nparticles; x++){ u[x] = u1 + x/((double)Nparticles); }
4.2.3 Prompt Construction. Beyond natural language, the LLM Engine requires four core components to form the complete prompt: (1) Machine Specification: The LLM Engine needs access to CPU specifications, particularly the core limit, to help produce generalizable recommendations. (2) Parallel Code Region: Each OpenMP kernel is provided as input for the LLM to infer its best-suited thread configuration. (3) Learned Prompts (Findings): These are offline-refined insights derived from our thread sensitivity study. Most are encoded as if-then rules [70, 84] (except Finding 7) to improve LLM interpretability. (4) Serial Code Regions: The LLM is also given a broader context of the program, including serial regions, which is essential for applying Finding 1 and Finding 8. However, in real-world scenarios, the entire codebase may be prohibitively large (e.g., thousands of lines in NPB [12]), exceeding the capacity of lightweight LLMs. To address this, we offload parts of the analysis to static LLVM passes (such as counting OpenMP kernels for Finding 8) to reduce inference workload while preserving key semantic information.
Figure 13: Illustrating two embarrassingly parallel OpenMP kernels (by #pragma omp parallel for) in Particlefilter. With increased per-thread workload from duplication, a moderate thread count achieves better hardware utilization. Finding 5: If the OpenMP kernel uses many variables, then PaRID causes high register and memory pressure. Moderate threads avoid slowdown. This is observed in Myocyte and Needle. For example, Myocyte’s function, ecc(), uses over 600 scalar variables. Since PaRID nearly doubles the number of registers required per thread for performing error detection, this leads to register spills and memory stalls. Therefore, even if these kernels may have balanced computations, a moderate thread number yields better performance. Finding 6: If the algorithm needs power-of-2 partitioning, then use power-of-2 thread counts. Irregular counts in PaRID cause more control-flow and overhead. This is observed in Backprop, where performance degrades significantly with 28 threads but remains acceptable with 16 and 32 threads. The algorithm computes partial derivatives across layers, each sized as a power of two, and relies on even partitioning. Using 28 threads introduces boundary handling issues, which increase control-flow complexity under PaRID and result in higher overhead. For such algorithms, thread counts should align with power-of-2 partitioning to avoid performance penalties. Finding 7: PaRID is sensitive to oversubscription. When the OpenMP kernel has load imbalance, runtime spikes when threads exceed core count. While oversubscribing threads is a common technique to improve performance in OpenMP programs [86], we find that PaRIDtransformed programs are consistently sensitive to this practice across all benchmarks. In particular, when the OpenMP kernel contains imbalanced control flow, such as an outermost if-else with imbalanced workloads, performance degrades significantly. Finding 8: If a program has multiple OpenMP kernels, then PaRID performs worse near core-limit. This is observed in Particlefilter, where PaRID’s performance degrades starting around 18 threads (not shown in Figure 12). Although all 10 OpenMP kernels in this benchmark are embarrassingly parallel, they are interleaved with serial regions, introducing implicit synchronization overhead. Under high thread counts, this overhead is amplified, and PaRID’s
5
EVALUATION
We provide evaluation setups and evaluate PaRID in this section.
5.1
Evaluation Setups
5.1.1 Benchmarks. We select all eight original benchmarks in NAS Parallel Benchmarks (NPB) suite [13] (V3.0 [12]). Details are shown in Table 1. These programs are specifically designed to evaluate parallel performance and include OpenMP support. Compared to other suites [18, 36, 81], NPB programs have larger workloads, tunable input classes, and diverse computation/memory patterns, such as irregular memory access in CG and embarrassingly parallel in EP, making them more reflective of real-world workloads. 5.1.2 Platform. We conduct experiments on an Intel Xeon CPU (2.1 GHz) with 28 cores and 32 GB of RAM, running Ubuntu 20.04 OS. The pass for PaRID is implemented using LLVM v15.0, which includes the standard optimizer and the Clang compiler. For parallel programs, we use OpenMP v5.0 (i.e. with _OPENMP value 201811). We use Phi-3.5-mini-instruct LLM (3.82 B) from Microsoft [3] as the LLM Engine for PaRID. LLM Inference is conducted using an NVIDIA A100 GPU (40 GB, 108 SMs) with CUDA 12.6. 7
Conference’17, July 2017, Washington, DC, USA Name
Description
IS EP CG MG FT BT SP LU
Integer sorting and ranking. Embarrassingly parallel Gaussian stats. Conjugate gradient linear system solver. Multi-Grid on a sequence of meshes. Discrete 3D fast Fourier Transform. Block Tri-diagonal solver. Scalar Penta-diagonal solver. Lower-Upper Gauss-Seidel solver.
Yafan Huang and Guanpeng Li
# Kernels
SLOC
2 2 14 10 7 9 7 8
1,117 681 1,334 1,690 1,682 4,126 3,482 4,039
0.4 0.8 1.6 1.2 0.3 0.6 1.2 0.9 0.2 0.4 0.8 0.6 0.1 0.2 0.4 0.3 0.0 4 8 16PaRID 0.0 4 8 16PaRID 0.0 4 8 16PaRID 0.0 4 8 16PaRID (a) IS
5.1.3 Evaluation Methodology. We evaluate PaRID from two perspectives: execution time and fault coverage. In this work, we prioritize execution time, as OpenMP parallelism is explicitly introduced to accelerate program execution [78]. A key design goal of PaRID is to use ➋ to mitigate the overhead introduced by ➊ through optimal thread tuning. Consistent with Section 12, we report normalized runtime, relative to the serial version of the original, unprotected program (normalized to 1.0). Evaluation settings for fault coverage are discussed separately in Section 5.3. For comparison, our baseline is the default thread configuration used in the NPB suite [12]. We also test alternative thread settings and confirm that this default yields reasonable performance. To the best of our knowledge, there are no existing instruction duplication techniques specifically designed for parallel programs. Additionally, we implement a serial instruction duplication baseline, representing state-of-the-art fault tolerance methods in non-parallelized programs [41, 42, 55, 60, 74]. This allows us to highlight the effectiveness of our parallel-aware duplication strategy in PaRID ➊.
(e) FT
Normalized Runtime
4.41
1.20
0.76
Baseline
0.96 0.13
0.88
0.55
0.54 0.45
1.18
0.71
1.22
PaRID
0.60
(f) BT
(g) SP
(h) LU
better than {1, 2, 28, 32}. This is also observed in Figure 12 (Section 4.2.2). For IS, EP, CG, MG, and FT, PaRID successfully selects the optimal thread configuration. Specifically, IS and CG contain frequent synchronization, both inter-kernel and intra-kernel, so PaRID selects moderate thread counts. In contrast, EP and FT are embarrassingly parallel and benefit from using more threads without reaching the core limit. Notably, in BT, SP, and LU, PaRID outperforms all fixed-thread settings. For example, in SP, PaRID achieves 20.16% faster execution than the best fixed setting (8 threads). These benchmarks represent larger workloads from Computational Fluid Dynamics (CFD) applications, each containing diverse computational patterns [5]. In BT, for instance, flux difference calculations are embarrassingly parallel, while L2-norm computations involve frequent synchronization [80]. PaRID’s per-kernel tuning enables it to adaptively assign optimal thread counts to these heterogeneous regions, leading to superior performance on large, complex workloads. Although fixed-thread settings can achieve comparable performance, they are not practical in production environments. Since real-world applications often run for hours to months [76], exhaustively tuning thread counts through trial-and-error methods incurs prohibitive cost. In contrast, PaRID determines these settings at compile time, offering performance portability without any expensive dynamic profiling efforts. Note that PaRID demonstrates high compatibility with OpenMPbased parallel programs, even in the presence of complex interthread communication patterns, where none of the existing instruction duplication methods can handle effectively [24, 26, 37, 42]. We also evaluate state-of-the-art implementations of serial instruction duplication techniques for comparison. On average, their normalized execution time is 2.44, significantly higher than PaRID’s 0.59, highlighting the effectiveness of our parallel-aware design in ➊. Ablation Study for Learned Prompts. We conduct an ablation study by disabling the Learned Prompts (i.e., eight Findings) in the ➋ stage of PaRID workflow. For each benchmark, we apply PaRID with and without Learned Prompts and measure the normalized execution time. The results are shown in Figure 16. Without prompt refinement, PaRID achieves comparable performance only on IS, EP, and FT. For the remaining benchmarks, the Learned Prompts significantly enhance the LLM Engine’s thread selection, resulting in up
Execution Time with PaRID
3.68
(d) MG
Figure 15: Overall execution time: PaRID vs a fixed-thread setting. For each subfigure, x-axis indicates PaRID or thread settings, and y-axis is normalized runtime. The red line highlights the normalized runtime of PaRID.
We evaluate PaRID’s execution time from following perspectives. 5.0 4.0 3.0 2.0 1.0 0.0
(c) CG
0.8 1.0 0.8 1.4 1.2 0.8 0.6 0.6 1.0 0.6 0.8 0.4 0.4 0.6 0.4 0.4 0.2 0.2 0.2 0.2 0.0 4 8 16PaRID 0.0 4 8 16PaRID 0.0 4 8 16PaRID 0.0 4 8 16PaRID
Table 1: Benchmark details. "# Kernel" indicates the number of OpenMP kernels; "SLOC" indicates Source Lines of Code.
5.2
(b) EP
0.35 0.19
IS EP CG MG FT BT SP LU Figure 14: Overall execution time: PaRID vs baseline.
Overall Execution Time. We first report the overall execution time comparison between PaRID and the baseline method. In the baseline setup, the NPB suite automatically uses the maximum available core count for execution, whereas PaRID applies LLMinferred thread counts for each OpenMP kernel. We use the ClassA input (a standard test size) from the NPB suite [12]. Figure 14 presents the normalized runtime results. In PaRID, runtime varies from 0.12 in EP to 1.19 in IS, while in the baseline method, it ranges from 0.30 in LU to 4.40 in CG. By selecting thread settings at compile time with LLM, PaRID achieves an average 1.74× speedup over the baseline, with a maximum speedup of 4.94× on EP. In addition to the baseline method, we also compare PaRID with fixed-thread settings. A fixed-thread setting of 𝑁 means that all OpenMP kernels in the program are executed with 𝑁 threads at runtime. Figure 15 presents the results for fixed-thread counts of {4, 8, 16}, which consistently achieve top performance across benchmarks, 8
Detecting Soft Errors in Parallel Software with LLM-tuned Instruction Duplication
Normalized Runtime
to ∼9× speedup (in LU) compared to PaRID without prompts. This performance gap arises because, without these prompts, the LLM tends to make overly simplistic assumptions: assigning 32 threads to kernels without barriers (assuming they are embarrassingly parallel), and 16 threads to those with synchronization. These default choices often fail to characterize dynamic behaviors introduced by instruction duplication. This result demonstrates the rationale for the design of the offline Thread Sensitivity Study in PaRID. 5.0 4.0 3.0 2.0 1.0 0.0
4.30
a moderate thread count to avoid the performance drop observed in the baseline method between Class A and B. Similar observations are also observed for fixed-thread settings. 4.0 3.5 3.0 2.5 2.0 1.5 1.0
0.96 0.13 0.13
1.98 0.55
0.43 0.45
1.17
0.71
0.7 PaRID Baseline 0.6 0.5 0.4 0.3 S W A B
0.19
IS EP CG MG FT BT SP LU Figure 16: Impact of Learned Prompts for PaRID.
(e) FT Name
Class S
Class W
Class A
Class B
Class C
IS EP CG MG FT BT SP LU
Keys=16 M=24 NA=1400 323 643 12/60 12/100 12/50
Keys=20 M=25 NA=7000 643 1283 24/200 36/400 33/300
Keys=23 M=28 NA=14000 2563 128 × 2562 64/200 64/400 64/250
Keys=25 M=30 NA=75000 2563 512 × 2562 102/200 102/400 102/250
Keys=27 M=32 NA=150000 5123 5123 162/200 162/400 162/250
B
C
1.2 1.0 0.8 0.6 0.4 0.2
(a) IS
1.97 0.60
PaRID Baseline
S W A
PaRID w/o Learnt Prompts PaRID w/ Learnt Prompts 2.21
1.39 1.20
Conference’17, July 2017, Washington, DC, USA
PaRID Baseline
S W A
B
(b) EP
C
1.6 1.4 1.2 1.0 0.8 0.6 0.4
(c) CG
PaRID Baseline
S W A (f) BT
PaRID 5.5 Baseline 4.5 3.5 2.5 1.5 C 0.5 S W A B
B
1.7 PaRID Baseline 1.4 1.1 0.8 C 0.5 S W A B (g) SP
C
1.6 1.4 1.2 1.0 0.8 0.6 0.4
PaRID Baseline
S W A
B
C
0.6 PaRID 0.5 Baseline 0.4 0.3 0.2 C 0.1 S W A B
C
(d) MG
(h) LU
Figure 17: PaRID with different inputs, where x- and y-axis for each figure mean input classes and normalized runtime. Entire PaRID Workflow Time. Executing PaRID incurs negligible compile-time overhead, as the only expensive component Thread Sensitivity Study is performed offline. The resulting Findings are generalizable and can be reused across different programs at runtime. The primary runtime cost of PaRID lies in LLM inference, which is lightweight: from less than 1 second for IS to ∼15 seconds for BT, even without leveraging inference acceleration techniques such as FlashAttention. Importantly, PaRID workflow is also a onetime cost. Once instruction duplication is inserted, PaRID enables efficient error detection during parallel execution, demonstrating its practicality for real-world scenarios.
Table 2: Problem size for different NPB input classes [12, 13]. X in MG and FT indicates grid size. For A/B in BT, SP, and LU, A and B mean problem size and iteration number.
Input Sensitivity Study. To evaluate PaRID’s adaptability in production, where a program may be executed with arbitrary inputs, we test PaRID on NPB benchmarks using varying input classes. The NPB suite provides multiple input classes to represent different workload sizes [12]. We focus on five input classes (S, W, A, B, and C) for all NPB benchmarks. Input details are in Table 2. Classes S and W are small workloads, whereas Classes A, B, and C represent standard problem sizes, with each successive class increasing the workload size by ∼4×. For example, Class C in LU corresponds to a 1623 grid executed over 250 iterations. Such numerical details along with explanations are inserted at Prompt Construction stage along with (1) Machine Specification for guiding LLM inference. Results are shown in Figure 17. Across all input classes, PaRID consistently outperforms the baseline method, up to 6.55% speed up in CG. For 3 large CFD workloads, it delivers up to 126.39% speedup in LU, 103.73% in BT, and 104.49% in SP. This improvement stems from the fact that the baseline routinely uses the maximum number of threads, which becomes increasingly suboptimal as workload sizes grow. This trend is observed in IS (from Class W to A), EP (from Class B to C), MG (from Class A to B), and LU (from Class A to B). With larger inputs, the per-thread floating-point workload becomes more intensive, potentially leading to degraded performance due to increased pressure on core resources. In contrast, PaRID dynamically adjusts its thread strategy based on input characteristics. For instance, in LU, the OpenMP kernel responsible for computing erhs() becomes increasingly memory- and computation-intensive for inputs larger than Class B. PaRID identifies this shift and selects
5.3
Fault Coverage
We evaluate the error detection effectiveness of PaRID in this section, using LLFI [6, 57] as our fault injection tool. The choice of LLFI is motivated by two key reasons. First, compared to assembly-level tools, LLFI provides accurate insights into program error propagation behaviors, particularly in terms of SDC rates [66]. Second, LLFI (v15 [6]) is fully compatible with our toolset and can be seamlessly extended to support fault injection for OpenMP programs. To simulate faults in the computational elements of the processor as defined by our fault model (Section 2.1), our fault injection methodology can be explained as follows. For each fault injection trial, we randomly sample a site (i.e., a bit in a register involved in execution) and inject a single-bit-flip fault using LLFI [6, 57]. After program execution is finished, we compare the results between current trial and fault-free execution to verify correctness under PaRID protection. This approach mirrors methodologies used in prior works [54, 75, 87]. For each benchmark, the injection process is repeated 1,000 times, yielding an error bar of <3.1% for 95% confidence intervals. Fault coverage is assessed by measuring the silent data corruption (SDC) rate, as it is the most severe failure outcome, often bypassing recovery mechanisms [44]. Our results show that PaRID consistently reduces the SDC rate to 0%, achieving full detection across all benchmarks, regardless of whether they are serial or OpenMP parallel programs or the thread count used. 9
Conference’17, July 2017, Washington, DC, USA
Yafan Huang and Guanpeng Li
Parallel-aware Code Transformation (➊) in PaRID also supports a selective duplication mechanism [42, 51, 58]. PaRID enables both instruction-level and kernel-level selective duplication: the former duplicates only vulnerable instructions, while the latter duplicates instructions only in vulnerable OpenMP kernels. Unlike serial programs, quantifying the fault coverage and runtime of selective duplication in multithreaded parallel applications requires non-trivial analysis. We consider this as our future work.
6.3
6 DISCUSSION 6.1 Other Hardware Platforms In our evaluation (Sections 4.2.2 and 5.1.2), we primarily use an Intel Xeon CPU with 28 cores. We observe that processor configuration directly impacts the optimal thread count for reducing execution time when applying PaRID. For example, in Section 5, a thread count of 16 reflects aggressive parallelism on this machine, while 4 and 8 threads represent more moderate configurations. However, on platforms with significantly more cores (a CPU with stronger parallel capability), these same thread counts become relatively conservative. To ensure the generalizability of our Findings, we avoid specifying exact thread counts in this work. We further evaluate PaRID on a separate system with an AMD EPYC processor (128 cores). Except for machine specifications, all the rest prompts remain the same. In this setting, PaRID continues to outperform the baseline method, which defaults to using the full core count (128 threads). For instance, in the EP benchmark, PaRID assigns 64 threads to each of the two OpenMP kernels, achieving a 5.37× speedup (measured by normalized runtime). These results demonstrate PaRID’s compatibility across different hardware platforms.
6.2
Extending PaRID Protection to Other ISAs
In this work, we perform PaRID code transformation and fault injection at the LLVM IR level. While this is well-justified and supported by existing works [9, 42, 47, 51, 54], extending PaRID to other instruction set architectures (ISAs) remains a valuable direction for future exploration. Recall from Section 4.1, the key components of PaRID ➊ include (1) Identifying Parallel Regions, (2) Annotating Eligible Instructions, and (3) Dataflow Analysis for Parallel. When applying PaRID to other instruction sets, such as the x86 ISA, the transformation process remains largely consistent, with some slight differences. Specifically, while components (2) and (3) are unaffected, the outermost OpenMP directive is not explicitly translated into a function call in the x86 ISA. Instead, OpenMP kernels execute immediately following an OpenMP runtime function, kmpc_fork_call, which serves as a clear marker for identifying parallel regions. This ensures that OpenMP subroutines can still be recognized as distinct entities in the assembly code. Moreover, due to the flexibility of LLVM design, IR code with PaRID protection can be compiled naturally into other ISA formats. In a nutshell, while PaRID is implemented at the LLVM IR level, its design is generic and not restricted to LLVM.
7 RELATED WORKS 7.1 Instruction Duplication Instruction duplication has been extensively explored over the past two decades [25, 47, 58, 60, 65, 74]. This technique operates independently of the program’s algorithm, making it entirely programagnostic. Reis et al. introduced SWIFT [74] and SWIFT-R [73], which leverage instruction-level redundancy to detect errors and recover correct program execution. Laguna et al. [51] employed a machine learning model to predict instruction vulnerability, enabling selective instruction duplication to achieve high fault coverage and low runtime overhead. This work also discussed instruction duplication on process-level parallelism (with MPI). Didehban et al. [24] duplicated instructions at the microarchitecture-level for reducing SDC in ARM CPUs. Huang et al. [43] characterized the runtime overhead variations of instruction duplication from both compilerlevel and microarchitecture-level perspectives. Although effective, their approach targets only serial programs and neglects addressing multithreading parallel programs, which is crucial in production.
Threats to Validity
One major concern for deploying PaRID in production settings is the reliability of the LLM Engine. Although large language models, especially high-end models such as GPT-4 [4] and Deepseek R1 [32], demonstrate strong inference capabilities, the accuracy of smaller device-scale models remains an open question. For example, LLMs with similar parameter sizes to the Phi-3.5-mini-instruct used in this work have unknown capabilities in predicting dynamic program features. To investigate this, we replace the original LLM in PaRID with deepseek-coder-7b-instruct-v1.5 [33], a 6.91B codefocused LLM that is publicly available on Huggingface. We run local inference using the same prompt structure defined in PaRID. Initially, the model outputs a Python script for predicting performance via linear regression, rather than directly responding with thread recommendations. However, after refining the phrasing in the natural language portion of the prompts, while keeping all eight Findings unchanged, the model produces thread decisions consistent with those reported in Section 5. This result suggests that the LLM engine in PaRID is replaceable but requires tuning of prompt languages. It is also worth noting that even state of the art LLMs, due to limited domain knowledge and training data, cannot infer instruction duplication overhead. The eight Findings remain essential for steering the model toward accurate predictions.
7.2
Error Resilience on Parallel Architectures
Reliability research has also focused on exploring error propagation and redundancy-based detection in parallel architectures [47, 50, 60, 87–89]. Yim et al. [89] incorporated duplication along with checksums and accumulation-based range checking for lightweight protection for GPGPU programs. Kuvaiskii et al. [50] adopted AVX, a SIMD instruction set for Intel CPU, to conduct fast error detection and execution recovery. Mahmoud et al. [60] performed code transformations at the SASS level to enable efficient and effective error detection on NVIDIA GPUs. Yang et al. [87] selectively copies unreliable threads for efficient error resilience for GPGPU applications. Fang et al. [28] extended LLFI to support OpenMP parallel programs but focused solely on studying error propagation behaviors rather than implementing an effective error detection 10
Detecting Soft Errors in Parallel Software with LLM-tuned Instruction Duplication mechanism. Wang et al. [83] leveraged thread-level redundancy for soft error detection. Compared with existing works, PaRID not only supports multithreaded parallel programming models but also ensures competitive execution time through LLM-inferred thread tuning, making it better suited for production environments.
8
Conference’17, July 2017, Washington, DC, USA
on a supercomputer. In SC’16: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 645–655. [15] Keren Bergman, Shekhar Borkar, Dan Campbell, William Carlson, William Dally, Monty Denneau, Paul Franzon, William Harrod, Kerry Hill, Jon Hiller, et al. 2008. Exascale computing study: Technology challenges in achieving exascale systems. Defense Advanced Research Projects Agency Information Processing Techniques Office (DARPA IPTO), Tech. Rep 15 (2008), 181. [16] Aurélien Cavelan, Rubén M Cabezón, and Florina M Ciorba. 2019. Detection of silent data corruptions in smoothed particle hydrodynamics simulations. In 2019 19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID). IEEE, 31–40. [17] Bodhisatwa Chatterjee, Neeraj Jadhav, Sharjeel Khan, and Santosh Pande. 2024. Phaedrus: Exploring Dynamic Application Behavior with Lightweight Generative Models and Large-Language Models. arXiv preprint arXiv:2412.06994 (2024). [18] Shuai Che, Michael Boyer, Jiayuan Meng, David Tarjan, Jeremy W Sheaffer, SangHa Lee, and Kevin Skadron. 2009. Rodinia: A benchmark suite for heterogeneous computing. In 2009 IEEE international symposium on workload characterization (IISWC). Ieee, 44–54. [19] Zhi Chen, Alexandru Nicolau, and Alexander V Veidenbaum. 2016. SIMD-based soft error detection. In Proceedings of the ACM International Conference on Computing Frontiers. 45–54. [20] Chiachen Chou, Prashant Nair, and Moinuddin K Qureshi. 2015. Reducing refresh power in mobile devices with morphable ECC. In 2015 45th Annual IEEE/IFIP International Conference on Dependable Systems and Networks. IEEE, 355–366. [21] Chris Cummins, Volker Seeker, Dejan Grubisic, Mostafa Elhoushi, Youwei Liang, Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Kim Hazelwood, Gabriel Synnaeve, et al. 2023. Large language models for compiler optimization. arXiv preprint arXiv:2309.07062 (2023). [22] Leonardo Dagum and Ramesh Menon. 1998. OpenMP: an industry standard API for shared-memory programming. IEEE computational science and engineering 5, 1 (1998), 46–55. [23] Bronis R de Supinski, Thomas RW Scogland, Alejandro Duran, Michael Klemm, Sergi Mateo Bellido, Stephen L Olivier, Christian Terboven, and Timothy G Mattson. 2018. The ongoing evolution of openmp. Proc. IEEE 106, 11 (2018), 2004–2019. [24] Moslem Didehban and Aviral Shrivastava. 2016. nZDC: A compiler technique for near zero silent data corruption. In Proceedings of the 53rd Annual Design Automation Conference. 1–6. [25] Moslem Didehban and Aviral Shrivastava. 2018. A compiler technique for processor-wide protection from soft errors in multithreaded environments. IEEE Transactions on Reliability 67, 1 (2018), 249–263. [26] Moslem Didehban, Hwisoo So, Prudhvi Gali, Aviral Shrivastava, and Kyoungwoo Lee. 2023. Generic soft error data and control flow error detection by instruction duplication. IEEE Transactions on Dependable and Secure Computing 21, 1 (2023), 78–92. [27] Harish Dattatraya Dixit, Sneha Pendharkar, Matt Beadon, Chris Mason, Tejasvi Chakravarthy, Bharath Muthiah, and Sriram Sankar. 2021. Silent data corruptions at scale. arXiv preprint arXiv:2102.11245 (2021). [28] Bo Fang, Karthik Pattabiraman, Matei Ripeanu, and Sudhanva Gurumurthi. 2014. Evaluating the error resilience of parallel programs. In 2014 44th Annual IEEE/IFIP International Conference on Dependable Systems and Networks. IEEE, 720–725. [29] Karl Fürlinger, Michael Gerndt, and Jack Dongarra. 2007. Scalability analysis of the SPEC OpenMP benchmarks on large-scale shared memory multiprocessors. In Computational Science–ICCS 2007: 7th International Conference, Beijing, China, May 27-30, 2007, Proceedings, Part II 7. Springer, 815–822. [30] Olga Goloubeva, Maurizio Rebaudengo, M Sonza Reorda, and Massimo Violante. 2003. Soft-error detection using control flow assertions. In Proceedings 18th IEEE Symposium on Defect and Fault Tolerance in VLSI Systems. IEEE, 581–588. [31] Qiang Guan, Xunchao Hu, Terence Grove, Bo Fang, Hailong Jiang, Heng Yin, and Nathan DeBadeleben. 2020. Chaser: An enhanced fault injection tool for tracing soft errors in mpi applications. In 2020 50th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). IEEE, 355–363. [32] Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948 (2025). [33] Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, et al. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming–The Rise of Code Intelligence. arXiv preprint arXiv:2401.14196 (2024). [34] Luanzheng Guo, Dong Li, Ignacio Laguna, and Martin Schulz. 2018. Fliptracker: Understanding natural error resilience in hpc applications. In SC18: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 94–107. [35] John L Gustafson. 1988. Reevaluating Amdahl’s law. Commun. ACM 31, 5 (1988), 532–533. [36] Matthew R Guthaus, Jeffrey S Ringenberg, Dan Ernst, Todd M Austin, Trevor Mudge, and Richard B Brown. 2001. MiBench: A free, commercially representative
CONCLUSION
In this work, we first conduct an initial study to understand how parallel programs are reflected at the instruction level through static analysis, and to assess how effectively device-scale LLMs can capture dynamic program features through dynamic analysis. Based on these insights, we propose PaRID, an LLM-tuned code transformation framework that achieves full compatibility with OpenMP-based parallel programs while providing optimal thread tuning through LLM inference. We also introduce a Thread Sensitivity Study and summarize eight generalizable Findings. Our evaluation demonstrates that PaRID (1) maintains compatibility with OpenMP programs, (2) efficiently reduces execution time through kernel-specific thread tuning, (3) consistently adapts to varying inputs, and (4) preserves full soft error detection effectiveness.
REFERENCES [1] [n. d.]. LLVM/OpenMP Documentation. https://openmp.llvm.org/ [2] [n. d.]. Rodinia Benchmark Suite v3.0. https://rodinia.cs.virginia.edu/ [3] Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219 (2024). [4] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023). [5] Asif Afzal, Zahid Ansari, Ahmed Rimaz Faizabadi, and MK Ramis. 2017. Parallelization strategies for computational fluid dynamics software: state of the art review. Archives of Computational Methods in Engineering 24, 2 (2017), 337–363. [6] Udit Kumar Agarwal, Abraham Chan, and Karthik Pattabiraman. 2022. Lltfi: Framework agnostic fault injection for machine learning applications (tools and artifact track). In 2022 IEEE 33rd International Symposium on Software Reliability Engineering (ISSRE). IEEE, 286–296. [7] Dimitris Agiakatsikas, George Papadimitriou, Vasileios Karakostas, Dimitris Gizopoulos, Mihalis Psarakis, Camille Belanger-Champagne, and Ewart Blackmore. 2023. Impact of voltage scaling on soft errors susceptibility of multicore server cpus. In Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture. 957–971. [8] Irina Alam and Puneet Gupta. 2022. COMET: On-die and In-controller Collaborative Memory ECC Technique for Safer and Stronger Correction of DRAM Errors. In 2022 52nd Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). IEEE, 124–136. [9] Abdul Rehman Anwer, Guanpeng Li, Karthik Pattabiraman, Michael Sullivan, Timothy Tsai, and Siva Kumar Sastry Hari. 2020. Gpu-trident: efficient modeling of error propagation in gpu programs. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1–15. [10] Scott Atchley, Christopher Zimmer, John Lange, David Bernholdt, Veronica Melesse Vergara, Thomas Beck, Michael Brim, Reuben Budiardja, Sunita Chandrasekaran, Markus Eisenbach, et al. 2023. Frontier: exploring exascale. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1–16. [11] Eduard Ayguadé, Nawal Copty, Alejandro Duran, Jay Hoeflinger, Yuan Lin, Federico Massaioli, Xavier Teruel, Priya Unnikrishnan, and Guansong Zhang. 2008. The design of OpenMP tasks. IEEE Transactions on Parallel and Distributed systems 20, 3 (2008), 404–418. [12] David Bailey et al. 2010. The NAS parallel benchmarks. Technical Report. Technical Report RNR-94-007, NASA Ames Research Center,(March 1994). [13] David H Bailey et al. 1991. The NAS parallel benchmarks—summary and preliminary results. In Proceedings of the 1991 ACM/IEEE Conference on Supercomputing. 158–165. [14] Leonardo Bautista-Gomez, Ferad Zyulkyarov, Osman Unsal, and Simon McIntoshSmith. 2016. Unprotected computing: A large-scale study of dram raw error rate 11
Conference’17, July 2017, Washington, DC, USA
Yafan Huang and Guanpeng Li
embedded benchmark suite. In Proceedings of the fourth annual IEEE international workshop on workload characterization. WWC-4 (Cat. No. 01EX538). IEEE, 3–14. [37] Zhengyang He, Yafan Huang, Hui Xu, Dingwen Tao, and Guanpeng Li. 2023. Demystifying and mitigating cross-layer deficiencies of soft error protection in instruction duplication. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1–13. [38] W Daniel Hillis and Guy L Steele Jr. 1986. Data parallel algorithms. Commun. ACM 29, 12 (1986), 1170–1183. [39] Peter H Hochschild, Paul Turner, Jeffrey C Mogul, Rama Govindaraju, Parthasarathy Ranganathan, David E Culler, and Amin Vahdat. 2021. Cores that don’t count. In Proceedings of the Workshop on Hot Topics in Operating Systems. 9–16. [40] Kuang-Hua Huang and Jacob A Abraham. 1984. Algorithm-based fault tolerance for matrix operations. IEEE transactions on computers 100, 6 (1984), 518–528. [41] Yafan Huang, Sheng Di, Zhaorui Zhang, Xiaoyi Lu, and Guanpeng Li. 2024. Versatile Datapath Soft Error Detection on the Cheap for HPC Applications. In SC24: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1–15. [42] Yafan Huang, Shengjian Guo, Sheng Di, Guanpeng Li, and Franck Cappello. 2022. Mitigating silent data corruptions in hpc applications across multiple program inputs. In SC22: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1–14. [43] Yafan Huang, Zhengyang He, Lingda Li, and Guanpeng Li. 2023. Characterizing Runtime Performance Variation in Error Detection by Duplicating Instructions. In 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 730–741. [44] Joshua Hursey, Jeffrey M Squyres, Timothy I Mattox, and Andrew Lumsdaine. 2007. The design and implementation of checkpoint/restart process fault tolerance for Open MPI. In 2007 IEEE International Parallel and Distributed Processing Symposium. IEEE, 1–8. [45] Costin Iancu, Steven Hofmeyr, Filip Blagojević, and Yili Zheng. 2010. Oversubscription on multicore processors. In 2010 IEEE International Symposium on Parallel & Distributed Processing (IPDPS). IEEE, 1–11. [46] Tal Kadosh, Niranjan Hasabnis, Timothy Mattson, Yuval Pinter, and Gal Oren. 2023. Quantifying openmp: Statistical insights into usage and adoption. In 2023 IEEE High Performance Extreme Computing Conference (HPEC). IEEE, 1–7. [47] Charu Kalra, Fritz Previlon, Norm Rubin, and David Kaeli. 2020. Armorall: Compiler-based resilience targeting gpu applications. ACM Transactions on Architecture and Code Optimization (TACO) 17, 2 (2020), 1–24. [48] Ken Kennedy. 1978. Use-definition chains with applications. Computer Languages 3, 3 (1978), 163–179. [49] Eric P Kim and Naresh R Shanbhag. 2010. Soft N-modular redundancy. IEEE Trans. Comput. 61, 3 (2010), 323–336. [50] Dmitrii Kuvaiskii, Oleskii Oleksenko, Pramod Bhatotia, Pascal Felber, and Christof Fetzer. 2016. Elzar: Triple modular redundancy using intel avx (practical experience report). In 2016 46th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). IEEE, 646–653. [51] Ignacio Laguna, Martin Schulz, David F Richards, Jon Calhoun, and Luke Olson. 2016. Ipas: Intelligent protection against silent output corruption in scientific applications. In Proceedings of the 2016 International Symposium on Code Generation and Optimization. 227–238. [52] Chris Lattner and Vikram Adve. 2004. LLVM: A compilation framework for lifelong program analysis & transformation. In International symposium on code generation and optimization, 2004. CGO 2004. IEEE, 75–86. [53] Jan-Patrick Lehr, Michael Halkenhäuser, Dhruva Chakrabarti, Saiyedul Islam, Dan Palermo, and Ron Lieberman. [n. d.]. ompTest–Unit Testing with OMPT. ([n. d.]). [54] Guanpeng Li, Karthik Pattabiraman, Siva Kumar Sastry Hari, Michael Sullivan, and Timothy Tsai. 2018. Modeling soft-error propagation in programs. In 2018 48th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). IEEE, 27–38. [55] Minli Liao, Sam Ainsworth, Lev Mukhanov, and Timothy M Jones. 2025. A Deep Technical Review of nZDC Fault Tolerance. In Proceedings of the 34th ACM SIGPLAN International Conference on Compiler Construction. 104–116. [56] Michael N Lovellette, KS Wood, DL Wood, Jim H Beall, Philip P Shirvani, Namsuk Oh, and Edward J McCluskey. 2002. Strategies for fault-tolerant, space-based computing: Lessons learned from the ARGOS testbed. In Proceedings, IEEE Aerospace Conference, Vol. 5. IEEE, 5–5. [57] Qining Lu, Mostafa Farahani, Jiesheng Wei, Anna Thomas, and Karthik Pattabiraman. 2015. Llfi: An intermediate code-level fault injection tool for hardware faults. In 2015 IEEE International Conference on Software Quality, Reliability and Security. IEEE, 11–16. [58] Qining Lu, Karthik Pattabiraman, Meeta S Gupta, and Jude A Rivers. 2014. SDCTune: A model for predicting the SDC proneness of an application for configurable protection. In Proceedings of the 2014 international conference on compilers, architecture and synthesis for embedded systems. 1–10. [59] Robert Lucas, James Ang, Keren Bergman, Shekhar Borkar, William Carlson, Laura Carrington, George Chiu, Robert Colwell, William Dally, Jack Dongarra,
et al. 2014. Doe advanced scientific computing advisory subcommittee (ascac) report: top ten exascale research challenges. Technical Report. USDOE Office of Science (SC)(United States). [60] Abdulrahman Mahmoud, Siva Kumar Sastry Hari, Michael B Sullivan, Timothy Tsai, and Stephen W Keckler. 2018. Optimizing software-directed instruction replication for gpu error detection. In SC18: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 842–854. [61] Aniruddha Marathe, Peter E Bailey, David K Lowenthal, Barry Rountree, Martin Schulz, and Bronis R de Supinski. 2015. A run-time system for power-constrained HPC applications. In High Performance Computing: 30th International Conference, ISC High Performance 2015, Frankfurt, Germany, July 12-16, 2015, Proceedings 30. Springer, 394–408. [62] Domenica Mirauda, Raffaele Albano, Aurelia Sole, and Jan Adamowski. 2020. Smoothed particle hydrodynamics modeling with advanced boundary conditions for two-dimensional dam-break floods. Water 12, 4 (2020), 1142. [63] Joe J Monaghan. 2005. Smoothed particle hydrodynamics. Reports on progress in physics 68, 8 (2005), 1703. [64] Nahmsuk Oh, Philip P Shirvani, and Edward J McCluskey. 2002. Control-flow checking by software signatures. IEEE transactions on Reliability 51, 1 (2002), 111–122. [65] Nahmsuk Oh, Philip P Shirvani, and Edward J McCluskey. 2002. Error detection by duplicated instructions in super-scalar processors. IEEE Transactions on Reliability 51, 1 (2002), 63–75. [66] Lucas Palazzi, Guanpeng Li, Bo Fang, and Karthik Pattabiraman. 2019. A tale of two injectors: End-to-end comparison of ir-level and assembly-level fault injection. In 2019 IEEE 30th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 151–162. [67] Minesh Patel, Jeremie S Kim, Taha Shahroodi, Hasan Hassan, and Onur Mutlu. 2020. Bit-exact ECC recovery (BEER): Determining DRAM on-die ECC functions by exploiting DRAM data retention characteristics. In 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 282–297. [68] Paul M Petersen and David A Padua. 1996. Static and dynamic evaluation of data dependence analysis techniques. IEEE Transactions on Parallel and Distributed Systems 7, 11 (1996), 1121–1132. [69] James S Plank and Michael G Thomason. 2004. A practical analysis of low-density parity-check erasure codes for wide-area storage applications. In International Conference on Dependable Systems and Networks, 2004. IEEE, 115–124. [70] Litao Qiao, Weijia Wang, and Bill Lin. 2021. Learning accurate and interpretable decision rule sets from neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 4303–4311. [71] Michael J Quinn. 1994. Parallel computing theory and practice. McGraw-Hill, Inc. [72] Eric Raut, Jie Meng, Mauricio Araya-Polo, and Barbara Chapman. 2020. Evaluating performance of OpenMP tasks in a seismic stencil application. In OpenMP: Portable Multi-Level Parallelism on Modern Systems: 16th International Workshop on OpenMP, IWOMP 2020, Austin, TX, USA, September 22–24, 2020, Proceedings 16. Springer, 67–81. [73] George A Reis, Jonathan Chang, and David I August. 2007. Automatic instructionlevel software-only recovery. IEEE micro 27, 1 (2007), 36–47. [74] George A Reis, Jonathan Chang, Neil Vachharajani, Ram Rangan, and David I August. 2005. SWIFT: Software implemented fault tolerance. In International symposium on Code generation and optimization. IEEE, 243–254. [75] Behrooz Sangchoolie, Karthik Pattabiraman, and Johan Karlsson. 2017. One bit is (not) enough: An empirical study of the impact of single and multiple bitflip errors. In 2017 47th annual IEEE/IFIP international conference on dependable systems and networks (DSN). IEEE, 97–108. [76] David E Shaw, Peter J Adams, Asaph Azaria, Joseph A Bank, Brannon Batson, Alistair Bell, Michael Bergdorf, Jhanvi Bhatt, J Adam Butts, Timothy Correia, et al. 2021. Anton 3: twenty microseconds of molecular dynamics simulation before lunch. In Proceedings of the international conference for high performance computing, networking, storage and analysis. 1–11. [77] Premkishore Shivakumar, Michael Kistler, Stephen W Keckler, Doug Burger, and Lorenzo Alvisi. 2002. Modeling the effect of technology trends on the soft error rate of combinational logic. In Proceedings International Conference on Dependable Systems and Networks. IEEE, 389–398. [78] David Skinner. 2005. Performance monitoring of parallel scientific applications. (2005). [79] Chad Spensky, Aravind Machiry, Nathan Burow, Hamed Okhravi, Rick Housley, Zhongshu Gu, Hani Jamjoom, Christopher Kruegel, and Giovanni Vigna. 2021. Glitching demystified: analyzing control-flow-based glitching attacks and defenses. In 2021 51st Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). IEEE, 400–412. [80] Harold S Stone. 1975. Parallel tridiagonal equation solvers. ACM Transactions on Mathematical Software (TOMS) 1, 4 (1975), 289–307. [81] John A Stratton, Christopher Rodrigues, I-Jui Sung, Nady Obeid, Li-Wen Chang, Nasser Anssari, Geng Daniel Liu, and Wen-mei W Hwu. 2012. Parboil: A revised benchmark suite for scientific and commercial throughput computing. Center for Reliable and High-Performance Computing 127, 7.2 (2012). 12
Detecting Soft Errors in Parallel Software with LLM-tuned Instruction Duplication [82] Timothy Tsai, Nawanol Theera-Ampornpunt, and Saurabh Bagchi. 2012. A study of soft error consequences in hard disk drives. In IEEE/IFIP International Conference on Dependable Systems and Networks (DSN 2012). IEEE, 1–8. [83] Cheng Wang, Ho-seop Kim, Youfeng Wu, and Victor Ying. 2007. Compilermanaged software-based redundant multi-threading for transient fault detection. In International Symposium on Code Generation and Optimization (CGO’07). IEEE, 244–258. [84] Siyuan Wang, Zhongyu Wei, Yejin Choi, and Xiang Ren. 2024. Can llms reason with rules? logic scaffolding for stress-testing and improving llms. arXiv preprint arXiv:2402.11442 (2024). [85] Jiesheng Wei, Anna Thomas, Guanpeng Li, and Karthik Pattabiraman. 2014. Quantifying the accuracy of high-level fault injection techniques for hardware faults. In 2014 44th Annual IEEE/IFIP International Conference on Dependable Systems and Networks. IEEE, 375–382. [86] Yonghong Yan, Jeff R Hammond, Chunhua Liao, and Alexandre E Eichenberger. 2016. A proposal to OpenMP for addressing the CPU oversubscription challenge. In OpenMP: Memory, Devices, and Tasks: 12th International Workshop on OpenMP, IWOMP 2016, Nara, Japan, October 5-7, 2016, Proceedings 12. Springer, 187–202. [87] Lishan Yang, Bin Nie, Adwait Jog, and Evgenia Smirni. 2021. Enabling software resilience in gpgpu applications via partial thread protection. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 1248–1259. [88] Lishan Yang, Bin Nie, Adwait Jog, and Evgenia Smirni. 2021. Sugar: Speeding up gpgpu application resilience estimation with input sizing. Proceedings of the ACM on Measurement and Analysis of Computing Systems 5, 1 (2021), 1–29. [89] Keun Soo Yim, Cuong Pham, Mushfiq Saleheen, Zbigniew Kalbarczyk, and Ravishankar Iyer. 2011. Hauberk: Lightweight silent data corruption error detector for gpgpu. In 2011 IEEE International Parallel & Distributed Processing Symposium. IEEE, 287–300.
13
Conference’17, July 2017, Washington, DC, USA