arXiv:2604.11094v1 [cs.SE] 13 Apr 2026
E2E-REME: Towards End-to-End Microservices Auto-Remediation via Experience-Simulation Reinforcement Fine-Tuning Lingzhe Zhang†
Yunpeng Zhai†
Tong Jia∗
Peking University; Key Laboratory of Data Intelligence and Security Beijing, China [email protected]
Alibaba Group China [email protected]
Peking University; Key Laboratory of Data Space Technology and System Beijing, China [email protected]
Minghua He
Chiming Duan
Zhaoyang Liu
Peking University; Key Laboratory of Data Intelligence and Security Beijing, China [email protected]
Peking University; Key Laboratory of Data Intelligence and Security Beijing, China [email protected]
Alibaba Group China [email protected]
Bolin Ding
Ying Li∗
Alibaba Group China [email protected]
Peking University; Key Laboratory of Data Intelligence and Security Beijing, China [email protected]
Abstract
Keywords
Contemporary microservice systems continue to grow in scale and complexity, leading to increasingly frequent and costly failures. While recent LLM-based auto-remediation approaches have emerged, they primarily translate textual instructions into executable Ansible playbooks and rely on expert-crafted prompts, lacking runtime knowledge guidance and depending on large-scale general-purpose LLMs, which limits their accuracy and efficiency. We introduce End-to-End Microservice Remediation (E2E-MR), a new task that requires directly generating executable playbooks from diagnosis reports to autonomously restore faulty systems. To enable rigorous evaluation, we build MicroRemed, a benchmark that automates microservice deployment, failure injection, playbook execution, and post-repair verification. We further propose E2E-REME, an end-to-end auto-remediation model trained via experience-simulation reinforcement fine-tuning. Experiments on public and industrial microservice platforms, compared with nine representative LLMs, show that E2E-REME achieves superior accuracy and efficiency.
Auto-Remediation, Reinforcement Fine-Tuning, Microservices
CCS Concepts • Software and its engineering → Maintaining software. † Equal contribution. ∗ Corresponding author.
This work is licensed under a Creative Commons Attribution 4.0 International License. FSE Companion ’26, Montreal, QC, Canada © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2636-1/26/07 https://doi.org/10.1145/3803437.3805206
ACM Reference Format: Lingzhe Zhang† , Yunpeng Zhai† , Tong Jia∗ , Minghua He, Chiming Duan, Zhaoyang Liu, Bolin Ding, and Ying Li∗ . 2026. E2E-REME: Towards Endto-End Microservices Auto-Remediation via Experience-Simulation Reinforcement Fine-Tuning. In 34th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (FSE Companion ’26), July 5–9, 2026, Montreal, QC, Canada. ACM, New York, NY, USA, 12 pages. https://doi.org/10.1145/3803437.3805206
1
Introduction
Modern microservice systems have become increasingly complex due to dynamic interactions and rapidly evolving runtime environments [68]. This rising complexity inevitably leads to more frequent and harder-to-predict system failures. Such failures can be extremely costly: according to industry analysis, large enterprises experience an average direct loss of exceed $1,000,000 for every hour of IT downtime [16]. Given the substantial operational and financial impact, modern cloud-native systems urgently require the capability not only to detect failures but also to automatically remediate them in a timely and reliable manner [5, 12, 23, 39, 61, 67]. As a result, autoremediation—which autonomously identifies appropriate corrective actions and executes them with minimal human intervention—has become a critical component for ensuring resilient and cost-efficient operations at scale [11, 15, 19, 22, 53, 59, 62–64]. With the rapid advancement of large language models (LLMs), researchers have increasingly explored leveraging their strong reasoning and code-generation capabilities [14, 18, 34, 35, 44, 52, 60, 65, 66, 71] for microservice remediation [56]. A practical and industryadoptable approach to address microservice auto-remediation is to use LLMs to generate ansible playbooks that can be automatically
FSE Companion ’26, July 5–9, 2026, Montreal, QC, Canada
executed to repair faulty services [40]. Compared with traditional shell scripts, Ansible playbooks serve as structured, declarative specifications for operational procedures and offer a higher-level abstraction with clearer structure, stronger readability, and improved reusability and maintainability [13]. These advantages have made playbooks a widely adopted mechanism for implementing automated operational procedures in large-scale microservice systems, and thus a natural target for LLM-driven auto-remediation [41]. Following this trend, existing work on LLM-driven microservice remediation can be broadly categorized into two groups: methods and benchmarks. Methods focus on generating executable ansible playbooks from human-written instructions. For example, WisdomAnsible [36] fine-tunes CodeGen to produce remediation scripts, MAPE-Ansible [41] leverages GPT-4 and LLaMa-2 70B in a MAPEK loop architecture, and WCA-Ansible [40] is pre-trained from scratch on natural language, source code, and Ansible data. Benchmarks support these studies by providing curated collections of prompts and playbook templates for automation tasks. KubePlaybook [31] contains 130 natural language prompts for generating automation-focused remediation scripts, while Andromeda [32] provides structural representations of over 125,000 Ansible roles, along with more than 800,000 concrete changes between role versions extracted from the underlying Git repositories. However, existing methods and benchmarks still face key limitations when applied in real-world scenarios: • Task-level: These approaches typically rely on human-crafted prompts authored by experienced SREs, where the LLM only translates textual instructions into executable code. Such designs depend heavily on manual intervention, lack iterative feedback from the runtime environment, and fail to achieve end-to-end automation from failure diagnosis to system recovery. • Method-level: The generation of Ansible playbooks critically depends on the current runtime state of the microservice system. Without accurate and up-to-date system state guidance, the generated remediation scripts may be suboptimal or even incorrect. Moreover, representative methods such as MAPE-Ansible [41] rely on very large, closed-source models (e.g., GPT-4 and LLaMa2 70B), which require substantial computational resources and time for inference, limiting both scalability and efficiency. To address the task-level challenge, we propose a new task, End-to-End Microservice Remediation (E2E-MR). E2E-MR aims to directly generate executable ansible playbooks from a given diagnosis report and autonomously recover the faulty system. As illustrated in Figure 1, unlike previous approaches that rely on human-crafted prompts authored by experienced SREs based on diagnostic reports, E2E-MR establishes a closed-loop remediation pipeline, in which LLMs translate diagnostic insights into concrete repair actions that can be automatically executed within the microservice environment. For evaluation and structured comparison of this challenging new task, we introduce MicroRemed 1 , a benchmark designed to assess LLMs’ capabilities in end-to-end microservice remediation. MicroRemed automatically deploys a real microservice system and continuously injects diverse failures. For each injected failure, it generates a corresponding diagnosis report based on the target 1 The benchmark is available at https://github.com/LLM4AIOps/MicroRemed.
Lingzhe Zhang et al. Execute Remediation
Diagnosis Report Microservice Systems
The CartService is suffering from high CPU load.
Failure Diagnosis Failure
LLM
SREs Write an ansible playbook ...: 1) Get a list of deployments ... 2) Use kubectl command ... 3) Use ... to ... output of step 2 ...
Ansible PlayBook
- name: Remediate CPU load with ... hosts: k8s_nodes tasks: - name: Get all deployments with ... shell: | kubectl get deployments …
Failure Diagnosis Diagnosis Report The CartService is suffering from high CPU load.
E2EREME Ansible PlayBook
Figure 1: Previous microservice remediation workflow compared with the end-to-end microservice remediation pipeline proposed in this paper. component and failure type, which is then provided to the LLM under evaluation. The LLM produces an Ansible playbook, which is executed automatically, and the system subsequently verifies whether the injected failure has been successfully repaired. MicroRemed supports unlimited rounds of random failure injection and verification, allowing for extensive stress testing and iterative evaluation. Moreover, to facilitate fair and structured comparison, we categorize remediation targets into three difficulty levels—easy, medium, and hard—based on the complexity and interdependency of the underlying failure scenarios. To address the method-level challenges, we propose E2E-REME, an end-to-end auto-remediation model for microservices via Experience-Simulation Reinforcement Training. E2E-REME2 is designed to mirror how human SREs handle failures: continuously acquiring fresh runtime signals, reasoning over candidate repair strategies, and refining their decisions before executing a final remediation plan. To operationalize this process, E2E-REME employs a lightweight multi-agent workflow, which we refer to as ThinkRemed. It serves as an architectural scaffold that structures remediation into probing, executing, and verifying-refinement steps. Training E2E-REME for the end-to-end microservice remediation task is accomplished via an Experience-Simulation Reinforcement Training pipeline tailored to microservice operations. The pipeline consists of three stages: (1) Expert-Guided Supervised Fine-Tuning (SFT), which provides foundational remediation behaviors; (2) Simulation-Based RFT, which exposes the model to synthetic failure scenarios and teaches it to act within simulated environments; and (3) Reality-Anchored RFT, which further optimizes the model using feedback collected from real remediations. We evaluate E2E-REME using MicroRemed, integrated with two widely adopted microservice systems—Train-Ticket [72] and Online-Boutique [7]—as well as a self-developed lightweight system, Simple-Micro. Experimental results show that E2E-REME surpasses nine representative LLMs, achieving up to an average 49.32% higher accuracy while maintaining competitive inference efficiency. We 2 The model is available at https://modelscope.cn/models/ZhangLingzhe/E2E-REME.
E2E-REME
further validate its performance under realistic industrial workloads and microservice environments, which similarly confirm the effectiveness and robustness of E2E-REME. In summary, the contributions of our work are as follows: • We introduce the task of end-to-end microservice remediation (E2E-MR), which requires LLMs to directly generate executable Ansible playbooks from diagnosis reports and autonomously repair faulty systems. To support systematic evaluation, we build MicroRemed, a challenging benchmark that automates microservice deployment, failure injection, playbook execution, and post-repair verification. • To address the challenges of E2E-MR, we propose E2E-REME, an end-to-end auto-remediation model for microservices. E2EREME operates within a lightweight multi-agent workflow, ThinkRemed, which structures the remediation process into probing, execution, and verification–refinement steps. The model is trained via a tailored Experience-Simulation Reinforcement Training pipeline. • We conduct extensive experiments on three microservice systems and compare E2E-REME against nine representative LLMs. Results demonstrate its superior accuracy and efficiency. Additional validation under realistic industrial-style workloads further confirms the effectiveness and robustness of E2E-REME.
2 Background 2.1 Microservice Auto-Remediation Modern microservice architectures decompose applications into large numbers of loosely coupled services that interact through lightweight APIs. While this design improves scalability and development agility, it also increases operational complexity: faults may arise from configuration errors, resource contention, cascading failures, or inconsistent service states. As a result, automated remediation has become essential for maintaining system reliability. Auto-remediation refers to the process of automatically generating a repair plan, and executing the necessary actions to restore system health. In practice, these repair actions often involve operations such as restarting services, modifying configurations, cleaning corrupted state, or redeploying components. Ansible is widely used for such tasks because it provides a declarative automation framework that can execute system-level and application-level operations across distributed environments. Its playbook-based design enables LLMs or automation agents to produce actionable repair procedures that can be directly executed without human intervention. An Ansible playbook is a YAML-based script defining hosts, tasks, and their execution conditions. Figure 2 is a simplified example illustrating how a remediation workflow can handle high CPU load by automatically scaling a service. This example shows how Ansible playbook operationalizes autoremediation by providing a structured, declarative interface that bridges diagnosis and repair—allowing LLM-based systems to generate actionable and directly executable recovery procedures.
FSE Companion ’26, July 5–9, 2026, Montreal, QC, Canada
--- name : Mitigate high CPU load hosts : microservice_nodes become : yes tasks : - name : Check CPU usage shell : " top - bn1 | awk -F '[ , ]+ ' '/ Cpu /{ print $3 + $5 } '" register : cpu
1 2 3 4 5 6 7 8 9
- name : Scale service if CPU > 80% shell : kubectl scale deploy my - service -- replicas =4 when : cpu . stdout | float > 80
10 11 12 13
- name : Notify monitoring shell : " curl http :// monitor / api / notify -d ' scaled '"
14 15
Reinforcement Fine-Tuning
Reinforcement Fine-Tuning (RFT) adapts language models to decision-making tasks by optimizing model behavior using reward
Figure 2: An Ansible Playbook for CPU scaling or preference feedback, rather than relying solely on supervised instruction–response pairs [3, 73]. Unlike supervised fine-tuning (SFT), which teaches models to imitate expert demonstrations, RFT enables models to explore action spaces, evaluate long-term consequences, and self-correct through iterative interaction with an environment. RFT methods can be broadly categorized into two families: (1) Reward-based policy optimization. These approaches assign scalar rewards to model-generated actions or trajectories and optimize the policy to maximize expected reward. Group-based variants, such as Group Relative Policy Optimization (GRPO) [42], enhance stability by comparing multiple model completions under the same context, enabling the model to learn nuanced distinctions among candidate actions. Such methods are particularly effective for structured reasoning and tool-use scenarios where reward signals derive from execution correctness, efficiency, or safety. (2) Preference-based optimization. In many real-world decision-making tasks, explicit numeric rewards are difficult to define, while human operators can reliably express preferences between paired model outputs. Direct Preference Optimization (DPO) [37] provides a scalable solution by learning directly from such pairwise preferences, aligning model behavior with human judgments without requiring reinforcement learning rollouts or handcrafted reward models. Both families offer complementary strengths: reward-based methods facilitate exploration and environment-driven learning, whereas preference-based methods enable fine-grained alignment with human operational expertise. Despite rapid progress, the application of RFT to system operations—particularly microservice auto-remediation—remains limited. Compared with static preference or tool-use settings, autoremediation introduces uniquely challenging characteristics: dynamic runtime states, multi-step causal dependencies, safety-critical actions, and scarce human supervision. These factors motivate a combined RFT strategy in E2E-REME, leveraging reward-driven exploration in simulated environments together with preferencedriven alignment using real-world operator corrections.
3 2.2
Benchmark Construction
We present the construction of the MicroRemed benchmark in this section. We begin with an overview of the task definition and the underlying design principles (§3.1), followed by the architecture
FSE Companion ’26, July 5–9, 2026, Montreal, QC, Canada
Lingzhe Zhang et al.
Auxiliary Context Chaos
Runtime Envs
Config Failure Injection
Action Constraint Candidate Remediation LLM
Failure Report Target Service
Recovery Generate Ansible PlayBook
Evaluation
Failure Microservice Systems
Failure Type
Execute Execution Remediate Engine
Microservice Systems
Status Verification
Figure 3: MicroRemed Benchmark Pipeline: the benchmark launches a real microservice; Failure Injection injects faults and produces a Failure Report; the Failure Report together with Auxiliary Context is provided to the Candidate Remediation LLM which generates an Ansible Playbook; the Execution Engine executes the playbook; Status Verification checks remediation success; Evaluation and Recovery restores the system for the next run. of the MicroRemed benchmark (§3.2) and the evaluation protocol (§3.3). Finally, we describe the overall composition of MicroRemed (§3.4).
3.1
Design Principles
Existing microservice remediation approaches typically depend on human-crafted prompts designed by experienced SREs, where LLMs merely translate natural language instructions into executable scripts such as Ansible playbooks. This paradigm lacks autonomy and generalization, as it relies heavily on explicit human reasoning rather than the model’s understanding of the system state. To address this limitation, we introduce the task of End-to-End Microservice Remediation (E2E-MR), which aims to evaluate an LLM’s ability to autonomously generate executable remediation plans given only structured diagnostic information. Unlike conventional prompt-based generation, E2E-MR emphasizes a direct remediation process that transforms diagnostic reports into actionable repair operations.
controlled failures, and interacts dynamically with running services. This design enables the benchmark to capture real-time behaviors, system dynamics, and contextual dependencies that static datasets cannot represent. • Execution-based Evaluation. Evaluation is not determined by linguistic or structural similarity of generated outputs, but by execution outcomes. Each generated playbook is executed within the microservice environment, and the benchmark verifies success by assessing whether the system has been fully recovered to its normal operational state. • Comprehensive Scalability. Built on these foundations, the benchmark is designed to be method-scalable, LLM-scalable, failure-scalable, and system-scalable. It supports diverse LLMbased remediation methods, allows plug-and-play replacement of remediation models, accommodates various failure scenarios, and can be easily extended to new microservice systems with minimal configuration effort.
3.2 ∗ 𝑓𝜃 : (S𝑡𝑎𝑟𝑔𝑒𝑡 , T𝑓 𝑎𝑖𝑙 , C𝑎𝑢𝑥 ) → 𝑝 , ∗ 𝑝 = arg max U E (𝑝, S𝑓 𝑎𝑖𝑙 ) = S𝑛𝑜𝑟𝑚𝑎𝑙 𝑝∈P
(1)
Formally, the E2E-MR task can be formulated as Equation 1, where 𝑓𝜃 is the candidate remediation LLM parameterized by 𝜃 , S𝑡𝑎𝑟𝑔𝑒𝑡 denotes the failed microservice, T𝑓 𝑎𝑖𝑙 the failure type, and C𝑎𝑢𝑥 auxiliary contextual information. P is the space of executable playbooks, E represents the execution environment, and U (·) measures the utility of successful recovery. The goal is to generate an optimal playbook 𝑝 ∗ that maximizes the likelihood of recovering the system state S𝑓 𝑎𝑖𝑙 to S𝑛𝑜𝑟𝑚𝑎𝑙 . Therefore, to design a benchmark for the E2E-MR task, we adhere to the following design principles: • Dynamic Execution Benchmark. Unlike most LLM benchmarks that collect static data to form fixed datasets, the proposed benchmark is designed as a live and interactive execution environment. It actively launches real microservice systems, injects
Architecture
Based on the above design principles, we develop MicroRemed. The overall architecture of MicroRemed is illustrated in Figure 3. MicroRemed actively launches real microservice systems and performs Failure Injection to introduce controlled faults. According to the injected target service and failure type, it generates a Failure Report, which—together with a set of Auxiliary Contexts—is provided to the Candidate Remediation LLM to produce an executable Ansible Playbook. The playbook is then executed by an Execution Engine to carry out automated remediation. After execution, a Status Cerification module checks whether the issue has been successfully resolved. Finally, the Evaluation and Recovery stage assesses the remediation outcome and restores the microservice system to its original state, thereby enabling reproducible and iterative experimentation. The Failure Injection module introduces faults into the system through two complementary approaches: chaos injection and configuration injection. For resource-related or runtime failures (e.g., CPU stress, memory pressure, or network latency), MicroRemed
E2E-REME
FSE Companion ’26, July 5–9, 2026, Montreal, QC, Canada
Probe Agent
Auxiliary Context Runtime Envs
Failure
Root Cause Analysis
Action Constraint Root Cause Service
Think Generate
Execution Agent
E2E-REME (Coordinator)
Failure Category
Remediate Microservice Systems
Ansible PlayBook N
Time Window Time
ExperienceSimulation RFT
Experience Pool
Y
Normal?
Judge
Verification Agent
Figure 4: Runtime pipeline of E2E-REME. The model acts as a coordinator within a multi-agent workflow, ThinkRemed, which organizes the remediation process through the coordination of probing, execution, and verification agents. adopts chaos injection, which dynamically perturbs the runtime environment using Chaos Mesh [30] to emulate realistic fault conditions. For configuration-related failures (e.g., incorrect environment variables or service dependency misconfigurations), the system applies configuration injection, which directly modifies specific configuration files or environment settings to trigger controlled failures. The Status Verification module resembles traditional anomaly detection in purpose but differs fundamentally in mechanism. While anomaly detection infers abnormality from large volumes of complex runtime data, status verification performs targeted validation of whether a specific injected failure has been fully remediated. For example, if a CPU-stress failure was injected into service A, status verification will exclusively inspect the CPU metrics of service A to confirm recovery. This targeted design ensures 100% verification accuracy, a level of precision unattainable by general anomaly detection approaches.
our benchmark we include seven representative types of failures and three real-world microservice systems. Failure Types. As shown in Table 1, MicroRemed includes seven representative failures across three categories: resource-level (CPU, memory, I/O saturation), network-level (network loss, network delay), and application-level (pod failure, configuration error). No.
Category
Failure Types
1 2 3
Resource-Level
CPU Saturation Memory Saturation IO Saturation
4 5
Network-Level
Network Loss Network Delay
6 7
Application-Level
Pod Failure Configuration Error
Table 1: Benchmark statistics on failure types
3.3
Evaluation Protocol
MicroRemed supports comprehensive evaluation from multiple perspectives, including performance, efficiency, and resource utilization. Specifically, we adopt the following metrics to quantify the effectiveness of candidate remediation LLMs: Remediation Accuracy (RA) — measures the proportion of failures that are successfully repaired, reflecting the overall performance of the model. Average Remediation Latency (ARL) — evaluates the temporal efficiency of each successful remediation cycle, encompassing both reasoning and execution delays. Average Token Consumption (ATC) — quantifies the language-model cost efficiency, representing the average number of tokens consumed to achieve a successful remediation.
3.4
Benchmark Composition
Although MicroRemed is designed with comprehensive scalability and supports extensible failure types and microservice systems, in
Microservice Systems. MicroRemed integrates three microservice systems. Among them, two widely used benchmarks—TrainTicket [72] and Online-Boutique [7]—are well recognized for emulating realistic production environments. In addition, we include a self-developed lightweight system, Simple-Micro, designed to enable controlled experiments and facilitate fine-grained analysis. Difficulty Levels. Although MicroRemed supports arbitrary combinations of injected failures, we define three standardized difficulty levels—easy (23 cases), medium (49 cases), and hard (80 cases)—to enable fair and structured comparison across remediation methods. Each level corresponds to a curated set of failure combinations that vary in fault diversity, dependency complexity, and recovery difficulty.
4
E2E-REME
To address the end-to-end microservice auto-remediation task, we propose E2E-REME, a model designed to emulate how human SREs diagnose and repair failures. As shown in Figure 4, E2E-REME operates by continuously gathering fresh runtime signals, reasoning
FSE Companion ’26, July 5–9, 2026, Montreal, QC, Canada
4.1
ThinkRemed: A Multi-Agent Microservice Auto-Remediation Framework
Figure 4 illustrates the operational workflow of ThinkRemed. When a microservice system experiences a failure, a state-of-theart root cause analysis identifies the faulty service and the corresponding failure category. This information, together with auxiliary context (e.g., runtime environment and action constraints), is provided as input to E2E-REME, which acts as the Coordinator within ThinkRemed. The Coordinator first receives the auxiliary context C0 and failure report R 0 , and adaptively determines whether to invoke the Probe Agent to gather additional runtime information from the system. The probe agent executes a series of system state queries and returns the corresponding results. Once sufficient information is collected, the Coordinator synthesizes a candidate Ansible playbook 𝑝𝑡 . The generated playbook 𝑝𝑡 is then sent to the Execution Agent, which attempts to remediate the faulty microservice system and records the execution outcomes for subsequent reflection. After execution, the Verification Agent evaluates the remediation result, producing a binary outcome 𝑣𝑡 ∈ {0, 1} indicating success or failure. It is important to note that this Verification Agent differs from the Status Verification used in the benchmark. In the benchmark’s simulated failure environment, Status Verification can directly access low-level system information to determine the repair outcome. In contrast, the Verification Agent operates in a live system setting and relies on state-of-the-art anomaly detection methods to assess whether the remediation was successful. If the remediation fails, the system enters a reflection phase, and control returns to the Coordinator for iterative refinement based on the feedback. To ensure timely remediation and accommodate LLM context limitations, the iteration loop is bounded by a maximum trial budget 𝑇max . 𝑝𝑡 = 𝑓𝜃 (R𝑡 , C𝑡 , I𝑡 ), 𝑠 𝑡 +1 = E (𝑝𝑡 , 𝑠𝑡 ), 𝑣𝑡 = V (𝑠𝑡 +1 ), (2) (R , C ) = U (R , C , 𝑠 ) 𝑡 +1 𝑡 +1 𝑡 𝑡 𝑡 +1 if 𝑣𝑡 = 0 and 𝑡 < 𝑇max In summary, the iterative process of ThinkRemed can be formalized as Equation 2, where 𝑓𝜃 denotes the Coordinator’s reasoning policy, E the execution operator, and V the verification predicate. This multi-agent, iterative workflow enables E2E-REME to continuously reason, act, and refine its remediation strategies in response to dynamic microservice system states.
② Simulation-Based RFT
① Expert-Guided SFT fsyn MicroRemed Benchmark
fsyn ts yn asy Oracle ✓ Re me n Teacher d
③ Reality-Anchored RFT Experience Replay Buffer a
-
MicroRemed Benchmark
tsim asim E2E-REME rsim
+ al re
asim
tsim
l
frea
freal SREs
fsim
l area
freal areal a
+ real
Offline
Real-World Deployment
Online
fsim
Rollouts
Verification
over potential repair strategies, and iteratively refining decisions before executing a final remediation plan. To operationalize this workflow, E2E-REME adopts a lightweight multi-agent framework, ThinkRemed, which structures the remediation process into three agents: probing (collecting runtime evidence), execution (applying candidate repairs), and verification–refinement (validating and adjusting actions based on system feedback). Built on top of this framework, we further train E2E-REME using a tailored ExperienceSimulation Reinforcement Training (RFT) pipeline, enabling the model to learn robust, action-oriented remediation behaviors aligned with real microservice operational dynamics.
Lingzhe Zhang et al.
Multi-Criteria Remed Grader Safety-Aware Action Penalizer
Figure 5: Overall framework of Experience-Simulation RFT
4.2
Experience-Simulation RFT
Experience-Simulation RFT consists of three stages: (1) ExpertGuided SFT, (2) Simulation-Based RFT, and (3) Reality-Anchored RFT. The first two stages—Expert-Guided SFT and Simulation-Based RFT—are used to train E2E-REME in an offline setting, while RealityAnchored RFT continuously fine-tunes the model after it is deployed online. 4.2.1 Expert-Guided SFT. Given that lightweight models exhibit limited reasoning ability, weak tool-use proficiency, and difficulty in directly generating executable Ansible playbooks, the goal of this stage is to teach the model the fundamentals of (i) structured reasoning, (ii) tool-calling behaviors, and (iii) valid Ansible playbook construction. To achieve this, we employ an Oracle Teacher Model and run it on the MicroRemed Benchmark. For each synthetic failure instance 𝑓𝑠𝑦𝑛 , the teacher model is allowed to interact with the environment. If the remediation attempt succeeds, we extract the reasoning trace 𝑡𝑠𝑦𝑛 , and the final executable Ansible playbook 𝑎𝑠𝑦𝑛 . These elements, paired with the original failure description, form the supervision tuples {𝑓𝑠𝑦𝑛 , 𝑡𝑠𝑦𝑛 , 𝑎𝑠𝑦𝑛 }. LSFT = E ( 𝑓syn ,𝑡syn ,𝑎syn ) − log 𝜋𝜃 (𝑡 syn, 𝑎 syn | 𝑓syn ) (3) The Expert-Guided SFT phase thus optimizes the model to reproduce both the teacher’s reasoning process and its generated repair actions. Formally, the SFT objective can be expressed as Equation 3, where denotes the policy of E2E-REME. 4.2.2 Simulation-Based RFT. While SFT enables the model to imitate expert behaviors and acquire essential formatting and tool-use patterns, it cannot teach the model to reason, explore, or self-correct beyond the expert demonstrations. To equip E2E-REME with these capabilities, we design a Simulation-Based RFT stage. In this stage, the SFT-initialized E2E-REME is deployed on the MicroRemed benchmark to generate full rollouts. For each simulated episode, the model produces a reasoning trace 𝑡 sim , an Ansible playbook 𝑎 sim , and receives a remediation outcome indicating whether the system has been successfully repaired. A reward is then assigned to the rollout using a combination of (1) a Multi-Criteria Remediation Grader and (2) a Safety-Aware Action Penalizer. The Multi-Criteria Remediation Grader provides the primary reward signal. It assigns a high reward for successful remediation and further incorporates several auxiliary criteria, including: structural
E2E-REME
FSE Companion ’26, July 5–9, 2026, Montreal, QC, Canada
correctness of the generated playbook, successful execution of individual tasks, absence of execution errors, and token-efficiency of the reasoning trace (encouraging concise reasoning when remediation succeeds). Complementing this, the Safety-Aware Action Penalizer imposes substantial penalties for unsafe actions. Even though action constraints are provided as auxiliary context, lightweight models occasionally generate playbooks that may lead to harmful system-wide side effects. To prevent such behaviors, any playbook containing unsafe operations receives a large negative penalty, regardless of whether the remediation happens to succeed. 𝑅 = 𝛼 ·I[success] + 𝛽 ·𝑟 struct + 𝛾 ·𝑟 exec + 𝛿 ·𝑟 eff − 𝜆·I[unsafe] (4) In summary, the overall reward for a simulated rollout can be defined as Equation 4, where I[·] is the indicator function, 𝑟 struct measures playbook structural validity, 𝑟 exec evaluates execution correctness, 𝑟 eff encourages concise reasoning, and I[unsafe] flags unsafe or system-risky actions. The coefficients 𝛼, 𝛽, 𝛾, 𝛿, 𝜆 control the relative influence of each signal. To optimize E2E-REME under this reward model, we adopt Group Relative Policy Optimization (GRPO), which stabilizes credit assignment across long reasoning–action traces. Given a batch of rollouts {(𝑓𝑖 , 𝑎𝑖 , 𝑅𝑖 )}, GRPO updates the policy by maximizing advantageweighted likelihood ratios within each rollout group. Formally, the optimization objective of this stage is given in Equation 5. Here, 𝜋𝜃 denotes the parametrized policy of E2E-REME after SFT initialization, G(𝑖) denotes the group of rollouts originating from the same simulated failure instance, and the term in parentheses introduces a group-wise baseline to reduce gradient variance and stabilize training.
action 𝑎 +real while decreasing the likelihood of the model-generated − . action 𝑎 real (
− LDPO = − log 𝜎 𝛽 Δ𝜃 (𝑎 +real ) − Δ𝜃 (𝑎 real ) Δ𝜃 (𝑎) ≜ log 𝜋𝜃 (𝑎 | 𝑓real ) − log 𝜋 ref (𝑎 | 𝑓real )
(6)
Formally, this preference-driven learning process is captured in Equation 6. Here, 𝜋ref denotes a frozen reference policy (typically the model obtained after Simulation-Based RFT), 𝛽 controls the sharpness of preference separation, and 𝜎 (·) is the logistic sigmoid function. This stage enables E2E-REME to continuously internalize human expertise, ensuring that the model not only improves after deployment but also conforms to real-world safety conventions, operational best practices, and implicit SRE decision criteria.
5
Evaluation
In this section, we first introduce the implementation of E2E-REME, followed by the experimental setup. We then evaluate E2E-REME in terms of the following four research questions: • RQ1: How accurately does E2E-REME perform microservice remediation compared to baseline LLMs? • RQ2: What is the inference efficiency of E2E-REME—measured by runtime and token consumption? • RQ3: How does each component in E2E-REME contribute to the final remediation accuracy? • RQ4: How well does E2E-REME perform under realistic industrial workloads and microservice environments?
5.1
Implementation & Setting
Through Simulation-Based RFT, E2E-REME learns to explore, self-correct, and refine its reasoning strategies beyond expert demonstrations, enabling more robust and autonomous remediation behaviors.
5.1.1 Implementation. We implement our algorithm using AgentEvolver [50], a self-evolving agent reinforcement-learning framework proposed by Alibaba Tongyi Lab. Unless otherwise stated, we set the training reward parameters to 𝛼 = 1, 𝛽 = 0.1, 𝛾 = 0.1, 𝛿 = 0.5, and 𝜆 = 2. The maximum retry number of ThinkRemed is set to 𝑇max = 1. We adopt Qwen3-8B as the backbone LLM for all fine-tuning stages.
4.2.3 Reality-Anchored RFT. During real-world deployment, each encountered failure instance 𝑓real is first handled by E2E-REME, − . If the producing a model-generated remediation playbook 𝑎 real model-generated action does not successfully remediate the system, the failure is escalated to human operators. Site Reliability Engineers (SREs) then provide a corrective playbook 𝑎 +real , which reflects an expert-preferred remediation strategy under the same condi− , tions. This naturally yields a pairwise preference signal: 𝑎 +real ≻ 𝑎 real indicating that the SRE action should be preferred over the model’s attempt. Rather than relying on manually engineered scalar rewards or heuristic scoring functions, Reality-Anchored RFT employs Direct Preference Optimization (DPO) to directly align E2E-REME with the remediation preferences demonstrated by SREs. Formally, let 𝜋𝜃 denote the policy of E2E-REME. The DPO objective encourages the model to increase the likelihood of generating the expert-preferred
5.1.2 Experimental Setup. E2E-REME is trained on a CentOS 8 server equipped with 24 Intel(R) Xeon(R) CPUs (2.90GHz), 400GB RAM, and four NVIDIA A800 GPUs, each with 80GB of memory. The MicroRemed benchmark microservices are deployed across three machines, each equipped with 16 Intel(R) Xeon(R) CPUs (2.50GHz) and 64GB RAM. To comprehensively evaluate the end-to-end microservice remediation capability of current LLMs, we examine a total of nine representative models, encompassing both closed-source and opensource variants. For fairness and consistency, all LLMs are executed within the ThinkRemed framework. Closed-Source LLMs: Qwen3-Plus, Qwen3-Max, and Qwen3Flash [47]. Open-Source LLMs: QwQ-32B, Qwen3-Next-80B-A3V, Qwen3235B-A22B, DeepSeek-V3.2-Exp [21], Kimi-K2 [45], and GLM4.5 [49].
LGRPO = −
∑︁ 𝑖
1 © ª 𝑅𝑗 ® log 𝜋𝜃 (𝑎𝑖 | 𝑓𝑖 ) · 𝑅𝑖 − |G(𝑖)| 𝑗 ∈ G (𝑖 ) ¬ « ∑︁
(5)
FSE Companion ’26, July 5–9, 2026, Montreal, QC, Canada
LLM Backbone
Lingzhe Zhang et al.
Train-Ticket Accuracy (%)
Latency (s)
Online-Boutique Accuracy (%)
Latency (s)
Simple-Micro Accuracy (%)
Latency (s)
Easy Med Hard Easy Med Hard Easy Med Hard Easy Med Hard Easy Med Hard Easy Med Hard Closed-Sourced LLMs Qwen3-Plus 47.83 30.61 31.17 79.83 77.44 81.22 43.48 43.75 31.58 53.35 83.77 59.35 47.83 36.17 38.03 71.26 75.79 78.18 Qwen3-Max 47.83 28.57 30.77 69.87 86.38 12.36 39.13 37.50 25.32 34.82 46.49 48.20 30.43 22.92 17.91 37.02 39.73 52.43 Qwen3-Flash 21.74 16.33 13.16 43.24 84.96 290.7 34.78 30.61 21.33 42.33 77.26 98.64 22.72 14.58 8.86 50.52 61.12 65.33 Open-Sourced LLMs QwQ-32B 17.39 10.20 7.89 157.4 194.1 183.2 Qwen3-Next 13.04 6.12 5.06 23.33 20.41 29.76 Qwen3-235B 39.13 34.69 33.78 83.34 92.20 73.54 DeepSeek-V3.2 8.70 16.33 11.54 148.3 155.1 121.5 Kimi-K2 21.74 20.00 29.49 90.84 81.56 101.1 GLM-4.5 21.74 20.41 27.63 189.2 112.1 108.1 E2E-REME
26.09 22.45 15.58 109.9 137.3 155.5 17.39 8.33 6.76 141.5 188.7 195.7 17.39 17.02 17.72 22.57 22.73 26.64 21.74 28.57 19.35 24.83 34.58 32.83 39.13 34.69 33.33 74.57 55.52 74.27 34.78 36.73 32.39 82.49 66.65 73.20 31.82 21.28 22.78 63.1 63.1 60.1 21.74 29.17 20.00 129.6 106.4 98.6 22.73 26.53 30.38 75.87 79.07 83.85 47.83 44.89 43.75 90.82 79.28 76.84 43.48 43.75 39.47 135.1 126.1 130.8 43.48 36.73 30.38 127.8 126.3 132.8
82.61 77.78 70.83 39.65 42.33 41.52 95.65 89.80 78.75 39.44 38.31 40.56 91.30 87.50 72.15 49.65 53.42 55.33
Table 2: Remediation Accuracy (left columns) and Latency (right columns) across closed-source and open-source LLM backbones
5.2
RQ1: Remediation Accuracy
We first compare the remediation accuracy of E2E-REME against other LLMs on the E2E-MR task. As illustrated in Table 2, the lightblue columns on the left summarize the accuracy under the easy, medium, and hard level settings for each microservice environment. Overall, except for E2E-REME, Qwen3-Plus achieves the strongest performance among all evaluated models, followed by Qwen3-235B. At the microservice level, Train-Ticket emerges as the most challenging environment, followed by Simple-Micro. Notably, even under the easiest difficulty level, all standalone LLMs fail to exceed 50% accuracy, underscoring the difficulty and rigor of the MicroRemed benchmark. In contrast, E2E-REME consistently outperforms the best baseline, Qwen3-Plus, by 56.48%, 48.50%, and 42.97% across the three microservice environments, demonstrating its clear superiority in end-to-end microservice remediation.
5.3
RQ2: Inference Efficiency
We next compare the remediation latency of E2E-REME with other LLMs. As shown in the light-gray columns of Table 2, the latency reflects the end-to-end time of a full remediation cycle, including model reasoning, system probing, action execution, and recovery verification. Overall, Qwen3-Next exhibits consistently low latency across all environments, indicating a lightweight reasoning pipeline and efficient prompt handling. However, when considered together with its accuracy, this speed advantage comes at the expense of insufficient reasoning depth and unstable remediation performance, rendering it largely impractical for real-world use. In contrast, E2EREME achieves the second-lowest latency among all evaluated
Figure 6: Latency–accuracy trade-off of various large language models on the Online-Boutique microservice models—incurring only modest overhead compared to Qwen3-Next (higher by 40.48%, 39.19%, and 41.76% across the three microservices, respectively), while remaining substantially faster than Qwen3Flash (lower by 70.52%, 45.79%, and 10.49%). More importantly, E2E-REME delivers consistently high accuracy, striking a favorable balance between efficiency and effectiveness and making it suitable for practical end-to-end microservice remediation. To provide a clearer comparison, we further plot the latency–accuracy trade-off in Figure 6, where both latency and accuracy are averaged over the three difficulty levels (Easy, Medium, and Hard) on the Online-Boutique microservice. Each point represents
E2E-REME
Backbone
FSE Companion ’26, July 5–9, 2026, Montreal, QC, Canada
Train-Ticket
Online-Boutique
Easy Medium Hard Easy Medium Hard Closed-Sourced LLMs Qwen3-Plus Qwen3-Max Qwen3-Flash
3363 3583 2948
4988 2145 3651
4359 18822 127299 108053 3256 4821 5366 5758 3891 3490 3003 3378
Open-Sourced LLMs QwQ-32B 4147 5918 5936 2003 3784 Qwen3-Next 14490 13267 13091 12190 9387 Qwen3-235B 4858 8369 6669 10624 15749 DeepSeek-V3.2 5195 7295 7785 5855 6004 6452 4964 6728 4810 6314 Kimi-K2 GLM-4.5 11264 9492 10652 11270 11991
2951 24385 30381 6454 7793 10692
E2E-REME(ours) 1715
1648
1844
1735 1552
2223
Table 3: Average Token Consumption per remediation results a model, where the x-axis denotes average inference latency (lower is better) and the y-axis indicates accuracy. The plot highlights the superiority of E2E-REME, which lies closest to the upper-left region. Relative to the second-best model, Qwen3-Max, E2E-REME achieves 57.71% higher accuracy while reducing latency by 8.65%, respectively. We further compare the average token consumption of each model, as shown in Table 3. The reported values include both input and output tokens, thereby reflecting the total reasoning and generation workload for each remediation process. The results align with the latency findings, though a few exceptions provide additional insight. For example, Qwen3-Plus and Qwen3-Next consume significantly more tokens without proportional increases in latency—mainly due to unnecessary probing steps that generate overly long command outputs. Aside from these cases, the results further corroborate the efficiency of E2E-REME: compared with the second most efficient model, Qwen3-Flash, its token consumption is lower by 49.53% and 45.06%, respectively.
5.4
RQ3: Ablation Study
To assess the contribution of each training stage in E2E-REME, we conduct an ablation study on the Online-Boutique microservice. The results are summarized in Table 4. Starting from the raw Qwen38B backbone, the model achieves only 30.43 %, 20.83%, and 15.56% accuracy across the three difficulty levels, confirming that the base model lacks the specialized knowledge required for effective endto-end remediation.
Method Qwen3-8B +SFT +Sim-RFT +Real-RFT
Accuracy (%)
Latency (s)
Easy Medium Hard
Easy
Medium Hard
30.43 34.78 86.96 95.65
359.35 263.72 36.37 39.44
319.24 250.08 50.47 38.31
20.83 24.49 85.71 89.80
15.56 22.50 69.62 78.75
343.13 285.43 51.84 40.56
Table 4: Ablation study on the Online-Boutique microservice
Introducing SFT yields a substantial improvement, boosting accuracy by 3.66%–6.94% across different levels. This demonstrates that supervised instruction tuning on high-quality remediation trajectories provides essential task grounding and significantly enhances the model’s action planning ability, while also reducing inference latency by 21.78% due to fewer unnecessary probing steps. Adding Sim-RFT yields the largest performance gain, delivering an additional 53.61% improvement. The gains indicate that synthetic experience simulation effectively enriches the model’s exposure to diverse failure-recovery patterns, enabling more robust reasoning in unseen scenarios. Finally, Real-RFT provides an additional but smaller improvement of +7.24% on accuracy. The modest gain is primarily because the failure types and microservices used in RealRFT training differ from those in the Online-Boutique evaluation environment. Nevertheless, even with this mismatch, Real-RFT still contributes to more consistent decision-making and a small reduction in execution time (-14.69%). We further validate the effectiveness of each component in ThinkRemed. To isolate the impact of our training pipeline, we perform this analysis using Qwen3-Plus—the strongest LLM backbone aside from E2E-REME—instead of our trained model. Method
Train-Ticket
Online-Boutique
Easy Medium Hard
Easy Medium Hard
ThinkRemed 47.83
30.61
31.17 43.48
43.75
31.58
w/o Probe 43.48 w/o Reflection 43.48 w/o P. & R. 39.13
34.69 28.57 33.33
30.38 39.13 26.92 34.78 20.51 30.43
40.43 36.17 35.42
30.38 25.32 20.51
Table 5: ThinkRemed’s ablation study As shown in Table 5, we evaluate three variants: removing the probe agent, removing reflection, and removing both. Overall, both components contribute positively. For example, in the Train-Ticket microservice (easy level), removing either probe or reflection reduces accuracy by 13.05%. Across all settings, reflection plays a more critical role than probing: removing reflection results in an average 5.53% accuracy drop, whereas removing the probe agent decreases accuracy by only 1.66%.
5.5
RQ4: Industrial Evaluation
To assess the practical effectiveness of E2E-REME under realistic industrial workloads and microservice environments, we conduct an evaluation on three production microservices paired with their real operational traffic. Unlike the benchmark experiments—which directly measure accuracy and latency—this study focuses on remediation time reduction, a metric that more faithfully captures the efficiency gains experienced by Site Reliability Engineers (SREs). For each incident, SREs manually recorded the time required to complete remediation both with and without E2E-REME, and the relative reduction was used as the final metric. We compare E2E-REME against two representative industrial automation systems, MAPE-Ansible [41] and WCA-Ansible [40]. As shown in Table 6, E2E-REME consistently achieves the largest reduction in remediation time across all three microservices. While MAPE-Ansible (powered by GPT-5) achieves reductions between
FSE Companion ’26, July 5–9, 2026, Montreal, QC, Canada
Method
Lingzhe Zhang et al.
Remediation Time Reduction (%) Microservice A Microservice B Microservice C
MAPE-Ansible WCA-Ansible E2E-REME(ours)
67.21 73.76 81.33
55.83 60.90 79.58
48.95 56.74 76.03
Table 6: Evaluation on 3 realistic industrial workloads and microservices: Percentage reduction in remediation time
48.95% and 67.21%, and WCA-Ansible provides moderate improvements of up to 73.76%, E2E-REME delivers substantially stronger reductions—improving over the second-best WCA-Ansible by 7.57%, 18.68%, and 19.29%, respectively. These findings indicate that even without being explicitly finetuned on the target industrial system, E2E-REME can generate robust and reliable remediation strategies that meaningfully accelerate human operational workflows. This demonstrates the method’s strong potential for practical adoption in production environments. That said, fully realizing the vision of end-to-end microservice remediation—ideally reducing SRE intervention by 99%—will require future work on broader environment robustness, safety mechanisms, and tighter integration with production automation pipelines.
6
Threats to Validity
We discuss the limitations of our work from two perspectives: the benchmark and the methodology. Benchmark. Although the MicroRemed benchmark provides sufficient challenges for evaluating end-to-end microservice remediation, the currently supported failure types remain limited—covering only seven of the most common categories. In real-world systems, failure modes are far more diverse and continuously evolving [46, 54, 55]. Nevertheless, the design of MicroRemed inherently supports extensibility; new failure types can be integrated seamlessly. The main challenge lies in the need to implement corresponding fault injection and detection mechanisms when introducing additional failure types. Methodology. While E2E-REME demonstrates strong performance on the end-to-end microservice auto-remediation task, its long-term stability and generalization still warrant further investigation. Although we include evaluations in realistic industrial environments, the inherent unpredictability of LLMs means that there is always a non-negligible risk of generating incorrect actions that could destabilize or even break the cluster. Therefore, fully deploying such systems in production requires additional safeguard mechanisms—such as action verification, safety filters, or fail-safe rejection modules—to ensure robust and reliable operation.
7 Related Work 7.1 Software Remediation Software remediation, as the next step beyond failure diagnosis, has long been studied as a generation problem. Existing work can be broadly grouped into two categories: mitigation solution generation and remediation script generation.
Mitigation Solution Generation. These approaches focus on generating actionable mitigation strategies for detected anomalies, often leveraging large corpora of historical incident reports. Toufique et al. [1] conduct the first large-scale study evaluating LLMs for root-cause analysis and mitigation in production incidents. Drishti et al. [6] integrate signals from the entire software development lifecycle and apply retrieval-augmented in-context learning to enhance mitigation quality. Pouya et al. [9] model the natural workflow of on-call engineers using three LLM agents—hypothesis formation, hypothesis testing, and mitigation planning. Remediation Script Generation. These methods aim to produce executable scripts or code snippets that directly automate repair actions. Xpert [17] generates tailored KQL queries for incident investigation. ShellGPT [43] fine-tunes GPT models for shell command recommendation. Wisdom-Ansible [36] and MAPE-Ansible [41] generate Ansible playbooks using fine-tuned or GPT-4-based MAPE-K architectures. WCA-Ansible [40] further pretrains a domain-specific model on natural language, source code, and Ansible data to enhance playbook generation. Our work falls into this category, but differs in aiming for end-to-end microservice remediation without relying on human-written playbooks or static domain knowledge.
7.2
LLM-based Failure Management
Large language models have recently been applied to enhance anomaly detection, failure diagnosis, and automated remediation in complex systems [56, 57]. Existing efforts can be broadly grouped into foundation models, fine-tuning-based approaches, and promptdriven methods. Foundation models for system data. A number of works aim to build LLM-style foundation models for time-series or log data. Representative examples include Lag-Llama [38], Timer [29], and TimesFM [4], which unify forecasting, imputation, and detection under transformer architectures. PreLog [20] and KAD-Disformer [48] extend this line to log parsing and multivariate anomaly detection. These models provide domain-specialized priors but are typically limited to a single modality. Fine-tuning general LLMs. Another line of work fine-tunes general-purpose LLMs for IT operations. Examples include UniTime [26], AnomalyLLM [25], and LogLM [27], which adapt GPT/LLaMA-style architectures for time-series forecasting, anomaly detection, and log analytics. RAG4ITOps [69] and OWL [8] further incorporate retrieval or adapter tuning for interactive diagnosis. These methods improve task specialization, but require extensive labeled data or careful adaptation. Our E2E-REME also falls into this category, but differs by integrating multi-stage reinforcement fine-tuning for end-to-end remediation. Prompt-based methods. Prompt-based solutions avoid heavy fine-tuning by leveraging instruction design, chain-of-thought reasoning, or retrieval. RCACopilot [2], Xpert [17], and LasRCA [10] design multi-step prompts for diagnosis and anomaly detection. LogGPT [28], LSTPrompt [24], and LM-PACE [51] improve interpretability through CoT or task decomposition. Retrievalaugmented approaches such as RAGLog [33], LogRAG [70] and XRAGLog [58] enhance log reasoning through case retrieval. While flexible, prompt-based methods often suffer from unstable generation and limited end-to-end automation.
E2E-REME
8
Conclusion
In this paper, we introduce the task of end-to-end microservice remediation. To enable systematic evaluation, we construct MicroRemed, a challenging benchmark that automates microservice deployment, failure injection, playbook execution, and post-repair verification. To tackle this task, we propose E2E-REME, an end-toend auto-remediation model for microservices based on experiencesimulation reinforcement fine-tuning. Experimental results show that MicroRemed presents substantial challenges for existing LLMs, while E2E-REME achieves superior accuracy and efficiency on both MicroRemed and realistic industrial microservice environments.
Acknowledgments This work is supported by Key RD Project of Guangdong Province, China (No.2020B010164003).
References [1] Toufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann, Xuchao Zhang, and Saravan Rajmohan. 2023. Recommending root-cause and mitigation steps for cloud incidents using large language models. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 1737–1749. [2] Yinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang, Xin Gao, Liu Shi, Yunjie Cao, Xuedong Gao, Hao Fan, Ming Wen, et al. 2024. Automatic root cause analysis via large language models for cloud incidents. In Proceedings of the Nineteenth European Conference on Computer Systems. 674–688. [3] Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems 30 (2017). [4] Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou. 2024. A decoderonly foundation model for time-series forecasting. In Forty-first International Conference on Machine Learning. [5] Chiming Duan, Minghua He, Pei Xiao, Tong Jia, Xin Zhang, Zhewei Zhong, Xiang Luo, Yan Niu, Lingzhe Zhang, Yifan Wu, et al. 2025. LogAction: Consistent Crosssystem Anomaly Detection through Logs via Active Domain. arXiv preprint arXiv:2510.03288 (2025). [6] Drishti Goel, Fiza Husain, Aditya Singh, Supriyo Ghosh, Anjaly Parayil, Chetan Bansal, Xuchao Zhang, and Saravan Rajmohan. 2024. X-lifecycle Learning for Cloud Incident Management using LLMs. arXiv preprint arXiv:2404.03662 (2024). [7] Google Cloud Platform. 2025. Online Boutique: A Cloud-First Microservices Demo Application. https://github.com/GoogleCloudPlatform/microservicesdemo. Accessed: October 15, 2025. [8] Hongcheng Guo, Jian Yang, Jiaheng Liu, Liqun Yang, Linzheng Chai, Jiaqi Bai, Junran Peng, Xiaorong Hu, Chao Chen, Dongfeng Zhang, et al. 2023. Owl: A large language model for it operations. arXiv preprint arXiv:2309.09298 (2023). [9] Pouya Hamadanian, Behnaz Arzani, Sadjad Fouladi, Siva Kesava Reddy Kakarla, Rodrigo Fonseca, Denizcan Billor, Ahmad Cheema, Edet Nkposong, and Ranveer Chandra. 2023. A Holistic View of AI-driven Network Incident Management. In Proceedings of the 22nd ACM Workshop on Hot Topics in Networks. 180–188. [10] Yongqi Han, Qingfeng Du, Ying Huang, Jiaqi Wu, Fulong Tian, and Cheng He. 2024. The Potential of One-Shot Failure Root Cause Analysis: Collaboration of the Large Language Model and Small Classifier. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering. 931– 943. [11] Minghua He, Chiming Duan, Pei Xiao, Tong Jia, Siyu Yu, Lingzhe Zhang, Weijie Hong, Jin Han, Yifan Wu, Ying Li, et al. 2025. United we stand: Towards end-toend log-based fault diagnosis via interactive multi-task learning. arXiv preprint arXiv:2509.24364 (2025). [12] Minghua He, Tong Jia, Chiming Duan, Pei Xiao, Lingzhe Zhang, Kangjin Wang, Yifan Wu, Ying Li, and Gang Huang. 2025. Walk the Talk: Is Your Logbased Software Reliability Maintenance System Really Reliable? arXiv preprint arXiv:2509.24352 (2025). [13] Lorin Hochstein and Rene Moser. 2017. Ansible: Up and Running: Automating configuration management and deployment the easy way. " O’Reilly Media, Inc.". [14] Weijie Hong, Yifan Wu, Lingzhe Zhang, Chiming Duan, Pei Xiao, Minghua He, Xixuan Yang, and Ying Li. 2025. CSLParser: A Collaborative Framework Using Small and Large Language Models for Log Parsing. In 2025 IEEE 36th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 61– 72. [15] Xiaosong Huang, Hongyi Liu, Yifan Wu, Lingzhe Zhang, Tong Jia, Ying Li, and Zhonghai Wu. 2025. UDA-RCL: Unsupervised Domain Adaptation for Microservice Root Cause Localization Utilizing Multimodal Data. IEEE Transactions on
FSE Companion ’26, July 5–9, 2026, Montreal, QC, Canada
Services Computing (2025). [16] Information Technology Intelligence Consulting (ITIC). 2024. ITIC 2024 Global Server Hardware,Server OS Reliability Report. Annual Report. ITIC. [17] Yuxuan Jiang, Chaoyun Zhang, Shilin He, Zhihao Yang, Minghua Ma, Si Qin, Yu Kang, Yingnong Dang, Saravan Rajmohan, Qingwei Lin, et al. 2024. Xpert: Empowering incident management with query recommendations via large language models. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering. 1–13. [18] Sathvik Joel, Jie Wu, and Fatemeh Fard. 2024. A survey on llm-based code generation for low-resource and domain-specific programming languages. ACM Transactions on Software Engineering and Methodology (2024). [19] Yuyuan Kang, Xiangdong Huang, Shaoxu Song, Lingzhe Zhang, Jialin Qiao, Chen Wang, Jianmin Wang, and Julian Feinauer. 2022. Separation or not: On handing out-of-order time-series data in leveled lsm-tree. In 2022 IEEE 38th International Conference on Data Engineering (ICDE). IEEE, 3340–3352. [20] Van-Hoang Le and Hongyu Zhang. 2024. Prelog: A pre-trained model for log analytics. Proceedings of the ACM on Management of Data 2, 3 (2024), 1–28. [21] Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437 (2024). [22] Hongyi Liu, Xiaosong Huang, Mengxi Jia, Lingzhe Zhang, Tong Jia, Zhonghai Wu, and Ying Li. 2025. AAAD: Asynchronous Inter-Variable Relationship-Aware Anomaly Detection for Multivariate Time Series. In 2025 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1–6. [23] Hongyi Liu, Yinping Ma, Xiaosong Huang, Lingzhe Zhang, Tong Jia, and Ying Li. 2025. ORA: Job Runtime Prediction for High-Performance Computing Platforms Using the Online Retrieval-Augmented Language Model. In Proceedings of the 39th ACM International Conference on Supercomputing. 884–894. [24] Haoxin Liu, Zhiyuan Zhao, Jindong Wang, Harshavardhan Kamarthi, and B Aditya Prakash. 2024. LSTPrompt: Large Language Models as Zero-Shot Time Series Forecasters by Long-Short-Term Prompting. In Findings of the Association for Computational Linguistics ACL 2024. 7832–7840. [25] Shuo Liu, Di Yao, Lanting Fang, Zhetao Li, Wenbin Li, Kaiyu Feng, XiaoWen Ji, and Jingping Bi. 2024. Anomalyllm: Few-shot anomaly edge detection for dynamic graphs using large language models. arXiv preprint arXiv:2405.07626 (2024). [26] Xu Liu, Junfeng Hu, Yuan Li, Shizhe Diao, Yuxuan Liang, Bryan Hooi, and Roger Zimmermann. 2024. Unitime: A language-empowered unified model for crossdomain time series forecasting. In Proceedings of the ACM Web Conference 2024. 4095–4106. [27] Yilun Liu, Yuhe Ji, Shimin Tao, Minggui He, Weibin Meng, Shenglin Zhang, Yongqian Sun, Yuming Xie, Boxing Chen, and Hao Yang. 2024. Loglm: From task-based to instruction-based automated log analysis. arXiv preprint arXiv:2410.09352 (2024). [28] Yilun Liu, Shimin Tao, Weibin Meng, Jingyu Wang, Wenbing Ma, Yuhang Chen, Yanqing Zhao, Hao Yang, and Yanfei Jiang. 2024. Interpretable online log analysis using large language models with prompt strategies. In Proceedings of the 32nd IEEE/ACM International Conference on Program Comprehension. 35–46. [29] Yong Liu, Haoran Zhang, Chenyu Li, Xiangdong Huang, Jianmin Wang, and Mingsheng Long. 2024. Timer: generative pre-trained transformers are large time series models. In Proceedings of the 41st International Conference on Machine Learning. 32369–32399. [30] Chaos Mesh. 2025. A powerful chaos engineering platform for kubernetes. URL: https://chaos-mesh.org (2025). [31] Zakeya Namrud, Komal Sarda, Marin Litoiu, Larisa Shwartz, and Ian Watts. 2024. Kubeplaybook: A repository of ansible playbooks for kubernetes autoremediation with llms. In Companion of the 15th ACM/SPEC International Conference on Performance Engineering. 57–61. [32] Ruben Opdebeeck, Ahmed Zerouali, and Coen De Roover. 2021. Andromeda: A dataset of Ansible Galaxy roles and their evolution. In 2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR). IEEE, 580– 584. [33] Jonathan Pan, Wong Swee Liang, and Yuan Yidi. 2024. Raglog: Log anomaly detection using retrieval augmented generation. In 2024 IEEE World Forum on Public Safety Technology (WFPST). IEEE, 169–174. [34] Leyi Pan, Zheyu Fu, Yunpeng Zhai, Shuchang Tao, Sheng Guan, Shiyu Huang, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Felix Henry, et al. 2025. OmniSafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models. arXiv preprint arXiv:2508.07173 (2025). [35] Leyi Pan, Shuchang Tao, Yunpeng Zhai, Zheyu Fu, Liancheng Fang, Minghua He, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Aiwei Liu, et al. 2025. d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models. arXiv preprint arXiv:2512.09675 (2025). [36] Saurabh Pujar, Luca Buratti, Xiaojie Guo, Nicolas Dupuis, Burn Lewis, Sahil Suneja, Atin Sood, Ganesh Nalawade, Matt Jones, Alessandro Morari, et al. 2023. Automated code generation for information technology tasks in yaml through
FSE Companion ’26, July 5–9, 2026, Montreal, QC, Canada
large language models. In 2023 60th ACM/IEEE Design Automation Conference (DAC). IEEE, 1–4. [37] Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems 36 (2023), 53728–53741. [38] Kashif Rasul, Arjun Ashok, Andrew Robert Williams, Arian Khorasani, George Adamopoulos, Rishika Bhagwatkar, Marin Biloš, Hena Ghonia, Nadhir Hassen, Anderson Schneider, et al. 2023. Lag-llama: Towards foundation models for time series forecasting. In R0-FoMo: Robustness of Few-shot and Zero-shot Learning in Large Foundation Models. [39] Youcef Remil, Anes Bendimerad, Romain Mathonat, and Mehdi Kaytoue. 2024. Aiops solutions for incident management: Technical guidelines and a comprehensive literature review. arXiv preprint arXiv:2404.01363 (2024). [40] Priyam Sahoo, Saurabh Pujar, Ganesh Nalawade, Richard Genhardt, Louis Mandel, and Luca Buratti. 2024. Ansible lightspeed: A code generation service for it automation. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering. 2148–2158. [41] Komal Sarda, Zakeya Namrud, Marin Litoiu, Larisa Shwartz, and Ian Watts. 2024. Leveraging large language models for the auto-remediation of microservice applications: An experimental study. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering. 358–369. [42] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300 (2024). [43] Jie Shi, Sihang Jiang, Bo Xu, Jiaqing Liang, Yanghua Xiao, and Wei Wang. 2023. ShellGPT: Generative Pre-trained Transformer Model for Shell Language Understanding. In 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 671–682. [44] Aditi Singh, Abul Ehtesham, Saket Kumar, and Tala Talaei Khoei. 2025. Agentic retrieval-augmented generation: A survey on agentic rag. arXiv preprint arXiv:2501.09136 (2025). [45] Kimi Team, Yifan Bai, Yiping Bao, Guanduo Chen, Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, et al. 2025. Kimi k2: Open agentic intelligence. arXiv preprint arXiv:2507.20534 (2025). [46] Zexin Wang, Jingjing Li, Quan Zhou, Haotian Si, Yuanhao Liu, Jianhui Li, Gaogang Xie, Fei Sun, Dan Pei, and Changhua Pei. 2025. A Survey on AgentOps: Categorization, Challenges, and Future Directions. arXiv preprint arXiv:2508.02121 (2025). [47] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388 (2025). [48] Zhaoyang Yu, Changhua Pei, Xin Wang, Minghua Ma, Chetan Bansal, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang, Xidao Wen, Jianhui Li, et al. 2024. Pre-trained kpi anomaly detection model through disentangled transformer. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 6190–6201. [49] Aohan Zeng, Xin Lv, Qinkai Zheng, Zhenyu Hou, Bin Chen, Chengxing Xie, Cunxiang Wang, Da Yin, Hao Zeng, Jiajie Zhang, et al. 2025. Glm-4.5: Agentic, reasoning, and coding (arc) foundation models. arXiv preprint arXiv:2508.06471 (2025). [50] Yunpeng Zhai, Shuchang Tao, Cheng Chen, Anni Zou, Ziqian Chen, Qingxu Fu, Shinji Mai, Li Yu, Jiaji Deng, Zouying Cao, et al. 2025. AgentEvolver: Towards Efficient Self-Evolving Agent System. arXiv preprint arXiv:2511.10395 (2025). [51] Dylan Zhang, Xuchao Zhang, Chetan Bansal, Pedro Las-Casas, Rodrigo Fonseca, and Saravan Rajmohan. 2024. LM-PACE: Confidence estimation by large language models for effective root causing of cloud incidents. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering. 388–398. [52] Lingzhe Zhang, Liancheng Fang, Chiming Duan, Minghua He, Leyi Pan, Pei Xiao, Shiyu Huang, Yunpeng Zhai, Xuming Hu, Philip S Yu, et al. 2025. A survey on parallel text generation: From parallel decoding to diffusion language models. arXiv preprint arXiv:2508.08712 (2025). [53] Lingzhe Zhang, Tong Jia, Weijie Hong, Mingyu Wang, Chiming Duan, Minghua He, Rongqian Wang, Xi Peng, Meiling Wang, Gong Zhang, et al. 2026. RuntimeSlicer: Towards Generalizable Unified Runtime State Representation for Failure Management. arXiv preprint arXiv:2603.21495 (2026). [54] Lingzhe Zhang, Tong Jia, Mengxi Jia, Ying Li, Yong Yang, and Zhonghai Wu. 2024. Multivariate log-based anomaly detection for distributed database. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 4256–4267.
Lingzhe Zhang et al.
[55] Lingzhe Zhang, Tong Jia, Mengxi Jia, Hongyi Liu, Yong Yang, Zhonghai Wu, and Ying Li. 2024. Towards close-to-zero runtime collection overhead: Raftbased anomaly diagnosis on system faults for distributed storage system. IEEE Transactions on Services Computing (2024). [56] Lingzhe Zhang, Tong Jia, Mengxi Jia, Yifan Wu, Aiwei Liu, Yong Yang, Zhonghai Wu, Xuming Hu, Philip Yu, and Ying Li. 2025. A Survey of AIOps in the Era of Large Language Models. Comput. Surveys (2025). [57] Lingzhe Zhang, Tong Jia, Mengxi Jia, Yifan Wu, Hongyi Liu, and Ying Li. 2025. ScalaLog: Scalable Log-Based Failure Diagnosis Using LLM. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5. [58] Lingzhe Zhang, Tong Jia, Mengxi Jia, Yifan Wu, Hongyi Liu, and Ying Li. 2025. XRAGLog: A resource-efficient and context-aware log-based anomaly detection method using retrieval-augmented generation. In AAAI 2025 Workshop on Preventing and Detecting LLM Misinformation (PDLM). [59] Lingzhe Zhang, Tong Jia, Xinyu Tan, Xiangdong Huang, Mengxi Jia, Hongyi Liu, Zhonghai Wu, and Ying Li. 2025. E-log: Fine-grained elastic log-based anomaly detection and diagnosis for databases. IEEE Transactions on Services Computing (2025). [60] Lingzhe Zhang, Tong Jia, Kangjin Wang, Weijie Hong, Chiming Duan, Minghua He, and Ying Li. 2025. Adaptive Root Cause Localization for Microservice Systems with Multi-Agent Recursion-of-Thought. arXiv preprint arXiv:2508.20370 (2025). [61] Lingzhe Zhang, Tong Jia, Kangjin Wang, Mengxi Jia, Yong Yang, and Ying Li. 2024. Reducing events to augment log-based anomaly detection models: An empirical study. In Proceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement. 538–548. [62] Lingzhe Zhang, Tong Jia, Mingyu Wang, Weijie Hong, Chiming Duan, Minghua He, Rongqian Wang, Xi Peng, Meiling Wang, Gong Zhang, et al. 2026. Efficient Failure Management for Multi-Agent Systems with Reasoning Trace Representation. arXiv preprint arXiv:2603.21522 (2026). [63] Lingzhe Zhang, Tong Jia, Yunpeng Zhai, Leyi Pan, Chiming Duan, Minghua He, Mengxi Jia, and Ying Li. 2026. Agentic Memory Enhanced Recursive Reasoning for Root Cause Localization in Microservices. arXiv preprint arXiv:2601.02732 (2026). [64] Lingzhe Zhang, Tong Jia, Yunpeng Zhai, Leyi Pan, Chiming Duan, Minghua He, Pei Xiao, and Ying Li. 2026. Hypothesize-Then-Verify: Speculative Root Cause Analysis for Microservices with Pathwise Parallelism. arXiv preprint arXiv:2601.02736 (2026). [65] Lingzhe Zhang, Yunpeng Zhai, Tong Jia, Chiming Duan, Siyu Yu, Jinyang Gao, Bolin Ding, Zhonghai Wu, and Ying Li. 2025. ThinkFL: Self-Refining Failure Localization for Microservice Systems via Reinforcement Fine-Tuning. arXiv preprint arXiv:2504.18776 (2025). [66] Lingzhe Zhang, Yunpeng Zhai, Tong Jia, Xiaosong Huang, Chiming Duan, and Ying Li. 2025. Agentfm: Role-aware failure management for distributed databases with llm-driven multi-agents. In Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering. 525–529. [67] Ling-Zhe Zhang, Xiang-Dong Huang, Yan-Kai Wang, Jia-Lin Qiao, Shao-Xu Song, and Jian-Min Wang. 2024. Time-tired compaction: An elastic compaction scheme for LSM-tree based time-series database. Advanced Engineering Informatics 59 (2024), 102224. [68] Shenglin Zhang, Sibo Xia, Wenzhao Fan, Binpeng Shi, Xiao Xiong, Zhenyu Zhong, Minghua Ma, Yongqian Sun, and Dan Pei. 2024. Failure diagnosis in microservice systems: A comprehensive survey and analysis. ACM Transactions on Software Engineering and Methodology (2024). [69] Tianyang Zhang, Zhuoxuan Jiang, Shengguang Bai, Tianrui Zhang, Lin Lin, Yang Liu, and Jiawei Ren. 2024. RAG4ITOps: A Supervised Fine-Tunable and Comprehensive RAG Framework for IT Operations and Maintenance. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track. 738–754. [70] Wanhao Zhang, Qianli Zhang, Enyu Yu, Yuxiang Ren, Yeqing Meng, Mingxi Qiu, and Jilong Wang. 2024. LogRAG: Semi-Supervised Log-based Anomaly Detection with Retrieval-Augmented Generation. In 2024 IEEE International Conference on Web Services (ICWS). IEEE, 1100–1102. [71] Zeyu Zhang, Quanyu Dai, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Jieming Zhu, Zhenhua Dong, and Ji-Rong Wen. 2025. A survey on the memory mechanism of large language model-based agents. ACM Transactions on Information Systems 43, 6 (2025), 1–47. [72] Xiang Zhou, Xin Peng, Tao Xie, Jun Sun, Chao Ji, Wenhai Li, and Dan Ding. 2018. Fault analysis and debugging of microservice systems: Industrial survey, benchmark system, and empirical study. IEEE Transactions on Software Engineering 47, 2 (2018), 243–260. [73] Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593 (2019).