ConceptioArchivearXiv CS
arXiv CSopen access

Multi-Perspective Agentic Program Repair via Code Property Graphs and Temporal Execution Graphs

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
softwarearchitecturesoftwareengineeringtesting
software engineering, software architecture, testing

Multi-Perspective Agentic Program Repair via Code Property Graphs and Temporal Execution Graphs Zhili Huang, Ling Xu* , and Hongyu Zhang

arXiv:2607.12605v1 [cs.SE] 14 Jul 2026

Chongqing University, Chongqing, China {[email protected], [email protected], [email protected]}

APR methods use multi-turn interaction, test feedback, repairingredient retrieval, and tool invocation to support fault analysis and iterative patch refinement [9], [21]. More recent agentbased methods further enable LLMs to explore repositories, retrieve relevant context, execute tests, and validate patches within a unified repair workflow [22]. Despite these advances, two limitations remain. The first concerns the use of dynamic execution information. Runtime evidence, including variable states, branch outcomes, method calls, return values, and exception contexts, provides direct observations of program behavior under failing tests and can support the diagnosis of complex logical bugs [23], [24]. However, raw execution traces are often large, fine-grained, and highly repetitive. Existing methods commonly convert execution logs into text and provide them directly to LLMs [25], [26], causing context bloat and possibly burying critical failure evidence in redundant information. When traces exceed the model’s context window, retaining only partial execution fragments may lose the call paths and state-propagation process required to explain the failure. The second limitation concerns how heterogeneous evidence is used during repair reasoning. Existing methods may incorporate static program structures, runtime information, error messages, and test feedback [21], [22], [25], [27]. However, these sources are often combined into a shared context and processed through a single reasoning path to derive a root-cause hypothesis. Candidate patches are then generated and ranked through repeated sampling [28], [29]. Although this process can produce syntactically different I. I NTRODUCTION patches, the candidates may still rely on the same suspicious loAutomated program repair (APR) aims to automatically cations, root-cause hypotheses, and repair direction. Therefore, generate patches for software bugs [1]–[5]. Traditional APR patch-level diversity does not necessarily lead to diverse fault techniques, including search-based [6], constraint-based [7], and analyses or repair strategies. A repair framework should instead template-based [8] approaches, have been effective for recurring exploit the complementary information provided by different bug patterns. However, their reliance on predefined search evidence sources and explore distinct root-cause hypotheses spaces, constraints, or repair patterns limits their ability to adapt before candidate-patch generation. to diverse bugs and complex repair scenarios [9]. LearningTo address these limitations, we present CT-Repair, an based approaches, particularly neural machine translation agentic APR framework that combines queryable static and (NMT)-based APR approaches [10]–[20], instead formulate dynamic program representations with multi-perspective reasonprogram repair as a translation task. These methods learn from ing. Given a defective program and its failing tests, CT-Repair bug-fix pairs (BFPs) [17] to translate buggy code into repaired organizes program evidence at both static and dynamic levels. code. For static analysis, it uses Joern [30] to construct a Code PropLarge language models (LLMs) have recently become erty Graph (CPG) that captures program structure, control flow, an important foundation for APR because of their code data dependencies, and call relations. For dynamic analysis, understanding and generation capabilities. Existing LLM-based CT-Repair instruments the program to collect execution events Abstract—Large language models (LLMs) have improved automated program repair (APR), but two limitations remain. First, raw execution traces are often too large and repetitive to serve as effective model context. Second, repeated patch sampling may produce different implementations without yielding distinct root-cause hypotheses or repair strategies. We present CT-Repair, an agentic APR framework representing static and dynamic evidence as queryable Code Property Graph (CPG) and Temporal Execution Graph (TEG). CT-Repair applies a three-stage filtering pipeline to construct compact TEGs. Three finite-state-machineguided agents analyze each bug from static, dynamic, and hybrid perspectives and independently produce evidence-grounded repair strategies. A strategy-guided generation procedure instantiates these strategies as candidate patches and uses validation feedback to refine the most promising strategy. We evaluate CT-Repair on 854 Java bugs from Defects4J v3.0. In the mixed-model configuration, CT-Repair correctly repairs 489 bugs. Under a controlled GPT-5.4-mini configuration, it repairs 388 bugs, 19 and 30 more than ReinFix and RepairAgent, respectively. The union of the three evidence perspectives repairs 99 more bugs than the strongest individual perspective. The filtering pipeline also compacts runtime evidence, with execution filtering narrowing the candidate method scope by 94.85% on average and behavior filtering further reducing retained runtime records by 55.97%. These results show that structured runtime evidence and multi-perspective reasoning can improve repair effectiveness without relying solely on a larger patch-generation budget. Index Terms—automated program repair, large language models, agentic software engineering, dynamic analysis, multiagent reasoning.

triggered by failing tests. It then uses a three-stage filtering data are available at: https://anonymous.4open.science/r/ pipeline to remove redundant information and constructs a CT-Repair-6D29/. Temporal Execution Graph (TEG) from the remaining runtime II. M OTIVATION evidence. Rather than directly feeding long execution logs into the model, CT-Repair exposes the CPG and TEG through query This section examines the limitations of existing LLM-based interfaces, allowing agents to retrieve evidence relevant to their APR methods in evidence organization and repair reasoning. current fault hypotheses on demand. First, we examine whether raw dynamic execution traces can Based on these queryable representations, CT-Repair shifts be used directly as model context. Second, we analyze why repair diversity from candidate-patch generation to the rea- repeated patch sampling may introduce implementation-level soning stage. The framework runs three agents in parallel, variation without producing sufficiently diverse root-cause each using a different evidence perspective. The static-oriented hypotheses. agent analyzes program structure and dependencies, while the dynamic-oriented agent examines execution paths and runtime A. Dynamic Execution Evidence: Informative but Difficult to states. The hybrid agent combines static and dynamic evidence Use to explain how a defect propagates through the program. Guided Recent LLM-based APR methods use dynamic execution by a finite-state machine (FSM), the agents independently information to characterize runtime behavior triggered by perform fault understanding, evidence collection, and repair- failing tests. However, execution traces are commonly sestrategy generation. They derive distinct root-cause hypotheses rialized as linear text and directly included in the model and corresponding repair strategies before patch generation. context, rather than organized as structured representations CT-Repair then uses a round-robin mechanism to instantiate of program behavior. To assess the practicality of using raw these strategies as candidate patches and evaluates their quality traces in this manner, we conduct a motivational study on 17 using compilation and test results. If no plausible patch is Defects4J projects. Because collecting and processing complete produced in the current round, the framework selects the most execution traces is expensive, we proportionally sample 100 promising strategy and refines it using validation feedback in bugs according to project-level bug distributions. As shown the next iteration. in Table I, each bug produces 2.38 million runtime events We evaluate CT-Repair on 854 real-world Java bugs from on average. Of these events, only 13.15K occur in methods Defects4J v3.0. In the mixed-model configuration, CT-Repair modified by the corresponding developer patches, accounting correctly repairs 489 bugs. Under a controlled setting where for 0.55% on average. In addition, 99.95% of the runtime CT-Repair and the compared methods use GPT-5.4-mini as the events are repetitive under our event-signature definition. base model, CT-Repair correctly repairs 388 bugs, 19 more These results reveal a core tension in representing dynamic than the strongest baseline, ReinFix, corresponding to a relative execution evidence. Complete traces preserve interprocedural improvement of 5.15%. These results indicate that CT-Repair call paths, branch outcomes, and state propagation, but their improves repair effectiveness under the controlled single-model size and redundancy make them difficult to use directly as setting while achieving its highest repair coverage in the mixed- model context. By contrast, retaining only a localized subset of model configuration. events yields a more compact representation but may discard In summary, this paper makes the following contributions: the key execution context needed for diagnosis. Therefore, • We propose CT-Repair, an agentic APR framework that in- dynamic execution evidence requires a structured, compact, troduces reasoning-level diversity through static, dynamic, and queryable organization that reduces redundant records and hybrid evidence perspectives. Three perspective- while preserving the temporal relations and interprocedural specific agents independently derive evidence-grounded context needed for diagnosis, and supports repair agents in root-cause hypotheses and repair strategies before patch retrieving hypothesis-relevant evidence on demand. generation. A round-robin generation mechanism then B. From Patch-Level Diversity to Reasoning-Level Diversity instantiates these strategies as candidate patches. • We design a Temporal Execution Graph (TEG) to orgaRepeated candidate-patch sampling does not necessarily nize dynamic execution evidence. A three-stage filtering lead to diverse fault diagnoses. As illustrated on the left side process removes unexecuted methods, structurally simple of Fig. 1, a common LLM-based APR workflow combines methods, and records disconnected from valid execution source code, dynamic execution evidence, error messages, flows. Query interfaces allow agents to retrieve evidence and test feedback into a shared context, and then samples relevant to their current fault hypotheses on demand. multiple candidate patches [28], [29]. This process may • We evaluate CT-Repair on 854 real-world Java bugs from produce syntactically different patches while preserving the Defects4J v3.0. The evaluation examines repair effec- same underlying root-cause hypothesis. Consider a null-pointer tiveness, runtime-evidence compression, multi-perspective failure. Different samples may introduce an explicit null check, complementarity, and component contributions under rewrite the corresponding condition, or invoke a utility method controlled model and patch-generation budgets. We also that performs an equivalent check. Although these patches encourage future researchers to leverage our approach to differ at the code level, they share the same hypothesis that develop more powerful APR tools. Our source code and a null check is missing before the call. Repeated sampling

Fig. 1. Comparison between patch-level sampling in existing APR workflows and reasoning-level diversity in CT-Repair.

TABLE I P ROJECT- LEVEL STATISTICS OF RAW DYNAMIC EXECUTION TRACES .

Project

#Bugs

Full Developer-Patch Event Repetitive Events Method Events Ratio Event Ratio

Chart Cli Closure Codec Collections Compress Csv Gson JacksonCore JacksonDatabind JacksonXml Jsoup JxPath Lang Math Mockito Time

4 5 14 4 5 6 4 4 4 10 3 7 4 7 10 5 4

666.40K 16.85K 206.50M 17.12K 19.09K 59.30K 8.14K 6.03K 4.06M 1.08M 4.43K 19.17M 1.66M 517.81K 4.26M 26.57K 52.79K

867 431 10.71K 11.70K 99 1.65K 150 1.16K 1.09M 862 384 158.33K 11.72K 15.82K 10.65K 350 996

0.13% 2.56% 0.01% 68.34% 0.52% 2.78% 1.84% 19.15% 26.82% 0.08% 8.68% 0.83% 0.70% 3.06% 0.25% 1.32% 1.89%

98.72% 90.56% 99.98% 95.81% 97.75% 92.72% 89.47% 89.24% 99.92% 97.50% 77.27% 99.99% 99.35% 99.80% 99.93% 96.46% 95.62%

Bug Avg.

100

2.38M

13.15K

0.55%

99.95%

Note: Event Ratio denotes the share of events in developer-modified functions. Repetitive Event Ratio is 1 − Nu /N , where Nu and N denote unique event signatures and total runtime events.

therefore explores alternative implementations of the same repair direction, rather than alternative explanations of the failure. As shown on the right side of Fig. 1, our key idea is to build relatively independent reasoning paths from static, dynamic, and hybrid perspectives instead of merging all information into one process. The static-oriented agent focuses on code structure and dependencies, the dynamic-oriented agent on

actual execution paths and runtime states, and the hybridoriented agent on aligning structure with behavior. Each agent independently formulates and validates a root-cause hypothesis before producing a repair strategy. As a result, candidate patches can be generated from different fault explanations and modification directions, rather than only from stochastic variations of a shared reasoning process. CT-Repair thus introduces diversity at the reasoning stage while retaining implementation-level diversity through subsequent strategyguided patch generation. III. A PPROACH A. Overview This section introduces the proposed CT-Repair framework. The framework generates candidate repairs through four components organized into two stages, as shown in Fig. 2. The first stage constructs queryable static and dynamic evidence and uses multi-perspective agents to derive repair strategies. The second stage instantiates these strategies as candidate patches, validates the patches, and uses validation feedback to refine promising strategies. In the first stage, CT-Repair constructs a CPG for the target codebase and a TEG from the executions triggered by failing tests. Query interfaces over the two graphs provide static and dynamic evidence on demand. Three finite-statemachine-guided agents then analyze the same bug from static, dynamic, and hybrid perspectives. Each agent independently formulates and validates a root-cause hypothesis before deriving a corresponding repair strategy. In the second stage, CT-Repair uses strategy-based roundrobin scheduling to instantiate each repair strategy with multiple LLMs. The resulting candidate patches are compiled and tested.

Fig. 2. Overall workflow of CT-Repair.

TABLE II S TATIC AND DYNAMIC ANALYSIS TOOLS AVAILABLE TO CT-R EPAIR AGENTS . Source

Tool Name

Type

Description

CPG

static identify variable static trace method usage static analyze method details static analyze method control flow static get method javadoc static get class structure static get class javadoc static get imports

Line-Level Method-Level Method-Level Method-Level Method-Level Class-Level Class-Level File-Level

Finds variable definitions, references, and types in a file Retrieves static call sites of a target method Extracts parameters, returns, calls, assignments, and locals Extracts control structures within a target method Retrieves the Javadoc comment of a target method Retrieves fields and declared methods of a target class Retrieves the Javadoc comment of a target class Retrieves import declarations from a target file

TEG

dynamic method macro dynamic method execution summary dynamic sequences info dynamic chronological flow dynamic call context dynamic class execution paths

Method-Level Method-Level Invocation-Level Invocation-Level Invocation-Level Class-Level

Summarizes return variants, callers, and callees Groups executions by arguments, returns, and sequences Retrieves runtime events for selected executions Traces timestamp-ordered events from one execution Retrieves upstream and downstream runtime call context Summarizes runtime paths among methods in a class

If no plausible patch is found, CT-Repair ranks the repair strategies based on the validation outcomes of their candidate patches. It then returns feedback from the best-performing patch of the highest-scoring strategy to the originating agent for further refinement. B. CPG and TEG Construction

Method-level queries retrieve parameters, return information, assignments, control structures, and static call sites. Class-level queries provide declared fields, methods, and related Javadoc documentation. File-level queries retrieve import declarations that indicate dependencies on external components. These queries allow agents to examine the structural context of suspicious code and narrow the search space for potential repair locations. TEG Construction. The TEG captures the runtime behavior observed during failing executions. CT-Repair first narrows the dynamic tracing scope and instruments retained methods at the bytecode level. The instrumentation records runtime events, including method calls, parameters and return values, variable updates, and branch outcomes. CT-Repair then organizes these events into a TEG, where nodes represent method-call instances and edges represent actual calls and temporal order. Each event keeps its timestamp, sequence number, and call index, allowing agents to query specific methods, call instances, and runtime states along the actual execution trace.

This stage organizes static program structure and runDynamic Trace Filtering. Raw execution traces contain time behavior into two complementary graph representations. many records that provide limited additional evidence for fault Supplying an LLM with complete repository code or raw diagnosis or are repeatedly generated by structurally simple execution logs can introduce substantial irrelevant and repetitive methods. CT-Repair applies a three-stage filtering pipeline information, making relevant evidence harder to identify in to control the size of the resulting TEG. First, execution long contexts [31]. CT-Repair therefore constructs a CPG and filtering (EF) uses JaCoCo coverage to exclude methods that a TEG and exposes query interfaces over both graphs. These are not executed by any failing test. Second, structural filtering interfaces allow the agents to retrieve evidence relevant to their (SF) analyzes the AST and excludes empty methods, trivial current hypotheses without processing the complete program getters and setters, and simple delegators. The semantics of or execution trace in every model call. Table II summarizes these methods can generally be recovered from source code, the static and dynamic analysis tools available to the agents. whereas instrumenting frequently invoked instances of them CPG Construction. The CPG captures static structures and may generate large numbers of repetitive runtime events. EF potential dependencies of the target codebase. CT-Repair uses and SF therefore jointly determine the scope of bytecode Joern [30] to construct a unified representation of abstract instrumentation. After the failing tests have been executed, syntax, control flow, data flow, and method-call relations. CT- behavior filtering (BF) removes isolated records and call Repair exposes CPG queries at four levels of granularity. Line- relations that cannot be connected to a valid invocation chain level queries locate variable definitions, references, and types. in the reconstructed execution. Such records lack sufficient

caller-callee or temporal context for analyzing fault propagation. CT-Repair constructs the final TEG from the remaining records, retaining the observed call relations, event order, and runtime states of the failing executions. C. FSM-Guided Multi-Perspective Repair Reasoning

tainties. If the strategy is incomplete, the agent follows refine_fix_suggestion and revises it. Once the strategy satisfies the output schema, the agent emits fix_suggestion_ready, transitions to done, and forwards the strategy to patch generation. Summary-Based State Management. CT-Repair uses summary-based state transfer to control context growth during multi-step reasoning. At each state transition, it ends the current LLM interaction and passes a structured state summary to the next state. The summary records the state-specific findings, unresolved issues, and selected transition action. The interaction history and tool outputs associated with each state are stored separately. If the summary lacks sufficient detail, the agent can retrieve relevant records from earlier states on demand. This design avoids repeatedly including the full reasoning history in subsequent model calls while preserving access to previously collected evidence.

Using the query interfaces over the CPG and TEG, CT-Repair uses an FSM to guide three agents in analyzing bugs from different evidence perspectives and generating repair strategies. Multi-Perspective Root-Cause Reasoning. CT-Repair runs three agents in parallel. All agents receive the same basic bug context, including the buggy code, triggering tests, and failure messages, but prioritize different evidence. Agent-S primarily uses CPG-based tools to examine program structure, static call relations, and control- and data-flow dependencies. Agent-D primarily uses TEG-based tools to inspect observed execution paths, call sequences, branch outcomes, and runtime states. Agent-H uses both tool sets to relate runtime anomalies to program structure and trace state propagation across methods. D. Strategy-Based Round-Robin Patch Generation Each agent independently formulates a root-cause hypothesis After the three agents produce their repair strategies, CTand derives one initial repair strategy. This design introduces Repair instantiates each strategy as concrete candidate patches. diversity at the reasoning stage rather than relying only on A strategy is represented as ⟨Root Cause, Suggestion⟩, where repeated patch sampling from a shared analysis context. CT- Root Cause provides an evidence-supported explanation of Repair supports both single-model and mixed-model configu- the failure and Suggestion specifies the target locations and rations. In the mixed-model configuration, different LLMs are intended changes. CT-Repair supplies the strategy together assigned to the three agents, and the model-to-perspective with the required bug context to the patch-generation models. assignments are rotated across bugs to avoid consistently Candidate generation is therefore conditioned on an explicit associating one perspective with a particular model [32]. diagnosis and modification plan rather than on the raw bug FSM-Guided Reasoning. CT-Repair uses an FSM to context alone. structure the analysis process and regulate tool use. As shown As shown in Fig. 2, CT-Repair uses strategy-guided roundin Fig. 3, the FSM contains three analysis states. Each state robin scheduling to cycle through multiple patch-generation defines an analysis objective, a structured output format, and models for each strategy. If a strategy is processed by r models valid transitions. The tools available to an agent are determined and each model generates at most k candidates, the strategy jointly by its evidence perspective and current state. An agent produces at most r × k patches. This procedure introduces advances after producing the required state output; otherwise, implementation-level variation while keeping the underlying it follows a self-loop and continues the current analysis. root-cause hypothesis and repair direction fixed. • understand_fault. The agent examines triggering Each candidate patch retains an identifier for the strategy tests, failure messages, and buggy code to characterize from which it was generated. CT-Repair can therefore associate the failure, identify suspicious locations, and formulate compilation and test outcomes with the corresponding rootan initial root-cause hypothesis. If the analysis remains cause hypothesis and modification plan. These associations are incomplete, the agent follows reanalyze_fault and subsequently used to evaluate the strategies and select one for stays in the current state. Otherwise, it transitions to further refinement. collect_evidence. E. Patch Validation and Strategy Refinement • collect_evidence. The agent invokes the tools Patch Validation. CT-Repair validates the candidate patches available to its evidence perspective and uses the retrieved evidence to support, revise, or reject the generated in the current iteration in batches. Each patch current hypothesis. If further evidence is needed, is applied to the buggy program and evaluated through it follows need_more_evidence and continues compilation and testing. A patch that compiles and passes the analysis. Once the evidence sufficiently supports all available tests is marked as PLAUSIBLE. Only when no a root-cause explanation, the agent transitions to PLAUSIBLE patch exists, CT-Repair records the validation outcome of each candidate and proceeds to patch scoring and generate_fix_suggestion. strategy selection. • generate_fix_suggestion. The agent derives a structured repair strategy from the supported rootPatch Scoring and Strategy Refinement. Let F0 be the cause hypothesis and collected evidence. The strat- set of tests that fail on the original buggy program, and let Fp egy identifies the target methods, describes the in- denote the set of tests that fail after applying patch p. The set tended changes, and records potential risks or uncer- F0 − Fp represents the original failing tests failures eliminated

Fig. 3. FSM-guided multi-perspective repair reasoning.

by p, whereas Fp − F0 represents new failures introduced by the patch. The score of patch p is defined as follows:   −1, Score(p) = +1,   |F0 −Fp | − |F0 |

|Fp −F0 | |F0 |+|Fp −F0 | ,

if p fails to compile, if p passes all tests, otherwise.

(1)

1 X Score(p). |Ps |

B. Benchmarks

For a compilable but non-plausible patch, the first term rewards the elimination of original test failures, while the second term penalizes newly introduced failures. The score therefore provides more fine-grained feedback than a binary compilation or test outcome. Let Ps denote the candidate patches generated from strategy s. CT-Repair evaluates the strategy using the mean score of its candidate patches: Q(s) =

RQ2: How effectively does TEG reduce execution data? RQ3: Do multi-perspective agents improve reasoning diversity and repair complementarity? • RQ4: How does each major component contribute to CT-Repair’s effectiveness? •

(2)

p∈Ps

If Q(s) ≤ 0 for every strategy, CT-Repair terminates the repair process because the current validation results provide no positive signal for further refinement. Otherwise, it selects the strategy with the highest Q(s) and identifies the highest-scoring patch generated from that strategy. CT-Repair extracts feedback from this patch, including its compilation status, remaining failing tests, newly introduced failures, and code changes. This feedback is returned to the agent that generated the selected strategy, enabling it to re-examine the root cause and refine the repair suggestion. Subsequent iterations focus only on this agent’s strategy, concentrating the limited budget on the most promising repair direction. IV. E XPERIMENT S ETUP A. Research Questions Our study addresses the following research questions (RQs): • RQ1: How does CT-Repair compare with state-of-the-art APR methods?

Defects4J. We conduct the primary evaluation on the widely used Defects4J benchmark [33]. We use Defects4J v3.01 , which contains 854 real bugs from 17 Java projects. Following prior studies [21], [34], we classify bugs according to the number of methods modified by the developer patch. The benchmark contains 483 single-function bugs and 371 multi-function bugs. Defects4J provides buggy and fixed versions together with test suites, enabling reproducible evaluation and fair comparison with existing APR techniques. C. Compared Techniques To evaluate repair performance, we compare CT-Repair with two recent and representative LLM-based agentic APR baselines. RepairAgent [22]: proposes an autonomous LLM-based APR agent that combines dynamic prompting, an FSM, and repair tools to collect bug information, search for repair ingredients, generate patches, and validate them. ReinFix [21]: presents an LLM-based APR method based on repair-ingredient search. During reasoning, it retrieves internal repair ingredients such as variable definitions to support rootcause analysis. During patch generation, it retrieves external repair ingredients from historical bug fixes to improve patchgeneration accuracy. For fair comparison, we reproduce both baselines and rerun them under a unified experimental setting. Specifically, we replace their original large language models with the same 1 https://github.com/rjust/defects4j/tree/v3.0.0

model backbones used by CT-Repair to reduce the influence of model-capability differences. All methods are evaluated on the same Defects4J bug set and use the same patch-validation workflow and decision criteria for plausible and correct patches. D. Evaluation Metrics We adopt the standard evaluation metrics used in prior APR studies. A plausible patch is defined as a patch that passes all provided unit tests, while a correct patch is a plausible patch that is semantically equivalent to the developer’s fix. To determine correctness, two authors independently inspected all plausible patches and resolved disagreements through adjudication until consensus was reached, ensuring that the correctness metric reflects semantic equivalence rather than mere test passing. E. Implementation CT-Repair is implemented with LangGraph and uses three language models: GPT-5.4-mini, Gemini-3-flash, and DeepSeekv4-flash. These models are used for both multi-perspective agent reasoning and patch generation. We set the temperature of reasoning agents to 0 for stable reasoning and the patch-generation temperature to 1 to increase candidate-patch diversity. We also disable the model reasoning budget to control additional reasoning overhead and ensure consistency across model settings. In fault localization (FL), to avoid additional biases introduced by the FL tool, we follow recent works [21], [22] under conditions of perfect fault localization. CT-Repair uses three parallel repair agents, each generating at most one repair strategy per bug from a different perspective. In the first iteration, each strategy is instantiated by three LLMs, yielding at most 3 × 3 candidate patches. If no plausible patch is found, CT-Repair refines only the best-performing strategy and instantiates it with the same three LLMs, generating three additional patches. Thus, the maximum patch budget for one bug is 3 × 3 + 3 candidate patches. To prevent ineffective loops, we limit the number of LangGraph node transitions in a single run to 50. In each FSM analysis state, an agent can call analysis tools at most four times. These limits control tool-call cost and runtime while improving experimental reproducibility. V. E VALUATION A. RQ1: Repair Effectiveness To address RQ1, we evaluate CT-Repair under different basemodel configurations. In single-model settings, all three agents use the same LLM. In the mixed-model configuration, the three agents use different LLMs, with model assignments rotated across bugs. This design avoids binding specific informational perspectives to any single model. We then compare CT-Repair with both ReinFix and RepairAgent under the same GPT-5.4mini setting, considering both repair effectiveness and sampling budget. Repair Effectiveness and Complementarity across Models. As shown in Table III, CT-Repair achieves strong repair effectiveness under all three single-model settings. When using Gemini-3-flash, DeepSeek-v4-flash, and GPT-5.4-mini as the

TABLE III R EPAIR RESULTS ( CORRECT FIXES / PLAUSIBLE FIXES ) AND COST ACROSS LLM SETTINGS . LLM Gemini-3-flash DeepSeek-v4-flash GPT-5.4-mini Mixed-model

Multi-function Single-function 124/156 103/141 93/115 130/162

336/395 307/364 295/346 359/411

All 460/551 410/505 388/461 489/573

Avg. Tokens (K) Avg. Time (s) 153.99 212.83 134.78 178.35

112.78 204.94 305.19 158.90

base LLM, CT-Repair correctly repairs 460, 410, and 388 bugs. The mixed-model configuration achieves the best overall result, generating 573 plausible patches, including 489 correct ones. Compared with the best-performing single-model setting using Gemini-3-flash, the mixed-model configuration repairs 29 additional bugs (489 vs. 460), corresponding to a 6.30% improvement, and achieves gains of 19.27% and 26.03% over DeepSeek-v4-flash and GPT-5.4-mini, respectively. These results indicate that different LLMs provide complementary bug understanding and patch generation capabilities, enabling CT-Repair to repair bugs missed by individual single-model configurations. Effectiveness and Efficiency Trade-off. Different model configurations show different trade-offs between repair effectiveness and computational cost. Gemini-3-flash achieves the shortest average runtime of 112.78 s, while GPT-5.4-mini has the lowest average token consumption of 134.78K. The mixedmodel configuration achieves the best repair effectiveness with an average cost of 178.35K tokens and 158.90 s per bug, both within the ranges observed across the three single-model settings. Comparison with Representative LLM-based agentic APR baselines. As shown in Table IV, under the same GPT-5.4mini setting, CT-Repair correctly repairs 388 bugs and generates 461 plausible patches, outperforming both ReinFix and RepairAgent. Compared with ReinFix, CT-Repair repairs 19 more bugs (388 vs. 369, +5.15%). Compared with RepairAgent, CTRepair repairs 30 more bugs (388 vs. 358, +8.38%). In addition, CT-Repair uses a sampling budget of only 3 × 3 + 3, which is smaller than ReinFix’s 3×3×5 and RepairAgent’s 117 attempts. Overall, CT-Repair not only achieves better repair effectiveness but also outperforms representative existing methods with fewer samples, further validating the effectiveness of its framework design. Answer to RQ1. CT-Repair achieves the best overall repair effectiveness in the mixed-model configuration and outperforms ReinFix and RepairAgent under the same GPT-5.4-mini setting with a smaller sampling budget. B. RQ2: Dynamic Information Compression To answer RQ2, we evaluate whether CT-Repair’s threestage filtering mechanism reduces the size and redundancy of dynamic execution information. Specifically, we compare the scales of methods, events, records, and traces before and after each filtering stage. Among the 854 bugs, 827 support unfiltered-trace collection; the remaining 27 are excluded from

#Bugs –

CT-Repair 3×3+3

ReinFix 3×3×5

RepairAgent 117

Chart Cli Closure Codec Collections Compress Csv Gson JacksonCore JacksonDatabind JacksonXml Jsoup JxPath Lang Math Mockito Time

26 39 174 18 28 47 16 18 26 110 6 93 22 61 106 38 26

19/20 18/22 49/63 10/12 13/14 23/26 12/14 11/11 12/14 48/61 1/1 46/49 6/7 42/48 55/71 17/18 6/10

18/19 16/19 46/59 13/15 11/15 25/28 10/10 11/12 12/14 45/59 2/2 43/45 2/6 40/46 50/67 18/18 7/9

14/17 15/19 40/53 12/14 16/22 18/25 13/13 13/13 11/12 43/48 2/2 44/49 8/11 33/42 55/71 10/13 11/12

Total

854

388/461

369/443

358/436

Removed Methods Single-function

Remaining Methods

4.88%

95.12%

(466 bugs)

3,995.77 −→ 194.92 methods Multi-function

5.52%

94.48%

(361 bugs)

3,846.45 −→ 212.24 methods Overall

5.15%

94.85%

(827 bugs)

3,930.59 −→ 202.48 methods

0%

20%

40%

60%

80%

100%

Percentage of Methods

Fig. 4. Reduction of executed methods after execution filtering.

RQ2 because their unfiltered traces are too large for stable collection. Execution Filtering. As shown in Fig. 4, EF reduces the average number of executed methods from 3,930.59 to 202.48, removing 94.85% of the original methods. Single-function and multi-function bugs show similar reductions of 95.12% and 94.48%, respectively. These results indicate that failing tests exercise only a small fraction of repository methods, and EF substantially narrows subsequent dynamic analysis by removing unexecuted paths. Structural Filtering. Building on EF, we evaluate how SF affects dynamic-information size. As shown in Table V, SF reduces the average number of executed methods from 202.48 to 168.22 (16.9% reduction), runtime events from 1,305.28K to 816.26K (37.5% reduction), and trace size from 506.33 MB to

10^7 80% 10^6

14 bugs account for 80% of removed runtime events

10^5

60%

10^4

Positive-saving bugs: 601 / 827 Non-positive-saving bugs: 226 / 827

40%

10^3

Cumulative share

Project Sampling Times

100%

10^8

Runtime events removed

TABLE IV R EPAIR RESULTS ( CORRECT FIXES / PLAUSIBLE FIXES ) FOR CT-R EPAIR AND BASELINES ON D EFECTS 4J V 3.0 UNDER GPT-5.4- MINI .

10^2 20% 10^1 0% 600

10^0 0

100

200

300

400

500

Bug rank sorted by runtime-event reduction

Fig. 5. Ranked distribution of runtime-event reductions introduced by structural filtering. TABLE V R EDUCTION ACHIEVED BY STRUCTURAL FILTERING . Metric

EF only

EF + SF

Reduction

Executed methods per bug Runtime events per bug (K) Trace volume per bug (MB)

202.48 1305.28 506.33

168.22 816.26 332.57

−16.9% −37.5% −34.3%

332.57 MB (34.3% reduction). Runtime events and trace size decrease more sharply than the number of methods because the low-value methods removed by SF, such as getters and simple delegates, are limited in number but frequently invoked. As shown in Fig. 5, SF reduces runtime events for 601 of 827 bugs, while only 14 bugs account for 80% of the removed events. This indicates that frequent invocations of low-value methods can cause extreme trace growth for some bugs. Without SF, such redundant events would obscure behaviors related to fault propagation and reduce the usability of dynamic information. SF therefore reduces the average trace size while preventing extreme traces from degrading dynamic analysis. Behavior Filtering. After EF and SF, BF further removes redundant records at the behavioral level. As shown in Table VI, BF reduces retained records from 473.87K to 208.64K (55.97% reduction) and trace size from 247.18 MB to 133.75 MB (45.89% reduction). Across bugs, BF reduces retained records for 824 of 827 bugs; 660 bugs show reductions of at least 25%, including 274 with reductions of at least 50% and 75 with reductions above 75%. These results show that BF consistently removes redundant dynamic records across most bugs rather than benefiting only a few extreme traces. Answer to RQ2. TEG progressively removes irrelevant and redundant dynamic information through EF, SF, and BF, producing compact and queryable runtime evidence for repair reasoning. C. RQ3: Reasoning Diversity and Repair Complementarity To answer RQ3, we analyze whether multi-perspective agents can introduce diversity and complementarity into repair on all 854 bugs. At the reasoning level, we compare the reasoning diversity of CT-Repair and ReinFix. To avoid bias from method names and presentation order in LLM-based evaluation, we anonymize the two methods, randomly shuffle their order, and ask three LLM judges to independently evaluate each

TABLE VI E FFECTIVENESS OF BEHAVIOR FILTERING . Bug Type Single-function Multi-function ALL

Records (K)

#Bugs

TABLE VIII PATCH - LEVEL DIVERSITY OF INTRA - AGENT AND INTER - AGENT CANDIDATES .

Trace Size (MB)

EF + SF

EF + SF + BF

Reduction

EF + SF

EF + SF + BF

Reduction

495.52 445.92 473.87

211.94 204.37 208.64

−57.23% −54.17% −55.97%

276.89 208.82 247.18

149.26 113.72 133.75

−46.09% −45.54% −45.89%

466 361 827

TABLE VII R EASONING - LEVEL EVIDENCE DIVERSITY COMPARISON BETWEEN CT-R EPAIR AND R EIN F IX . Metric

LLM Judge

CT-Repair

ReinFix

Difference

Evidence cluster count ↑

GPT-5.5 Gemini-3.5-flash DeepSeek-V4-Pro

1.695 ± 0.008 1.674 ± 0.004 1.683 ± 0.015

1.253 ± 0.012 1.143 ± 0.005 1.048 ± 0.004

+0.442 +0.531 +0.635

Perspective diversity score ↑

GPT-5.5 Gemini-3.5-flash DeepSeek-V4-Pro

2.172 ± 0.013 2.235 ± 0.014 2.327 ± 0.025

1.401 ± 0.014 1.240 ± 0.004 1.094 ± 0.009

+0.771 +0.995 +1.233

Duplicate evidence count ↓

GPT-5.5 Gemini-3.5-flash DeepSeek-V4-Pro

1.259 ± 0.008 1.281 ± 0.004 1.271 ± 0.015

1.641 ± 0.012 1.751 ± 0.005 1.845 ± 0.004

−0.382 −0.470 −0.574

comparison three times, reporting the mean and standard deviation. At the patch level, we use the exact duplication rate to measure the proportion of identical patches and normalized AST distance to measure structural differences between patches, and compare patch diversity within the same agent (Intra-Agent) and across different agents (Inter-Agent). At the repair level, we report the bugs correctly repaired by Agent-S, Agent-D, and Agent-H, as well as their union and overlap, to assess whether multi-perspective reasoning expands repair coverage. Reasoning diversity of multi-perspective Agents. As shown in Table VII, CT-Repair produces more diverse evidence than ReinFix under all three LLM judges. Compared with ReinFix, CT-Repair increases the number of evidence clusters by 0.44–0.64 and the perspective diversity score by 0.77–1.23, while reducing repeated evidence by 0.38–0.57. The consistent evaluations show that Agent-S, Agent-D, and Agent-H derive distinct root-cause analyses and repair strategies from different information sources rather than repeatedly sampling the same repair process. Diversity of the candidate patch space. As shown in Table VIII, for single-function bugs, the exact duplication rate decreases from 30.68% to 22.83%, while the median normalized AST distance increases from 0.0397 to 0.0681, a 71.5% improvement. For multi-function bugs, the duplication rate decreases from 28.61% to 19.43%, while the median normalized AST distance increases from 0.0405 to 0.0794, a 96.0% improvement. These results show that multi-perspective reasoning transforms distinct fault understandings into a broader candidate patch space, with a stronger effect on multi-function bugs involving cross-method dependencies. Repair complementarity across agent perspectives. As shown in Table IX, Agent-S, Agent-D, and Agent-H correctly repair 390, 350, and 377 bugs, respectively, while their union reaches 489. Compared with the best individual agent, Agent-S, combining the three perspectives repairs 99 additional bugs, a 25.38% improvement. The agents uniquely repair 53, 29, and 35 bugs, respectively, and jointly repair 256 bugs. These

Setting

Comparison

Dup. Rate ↓

AST Distance ↑

AST Gain ↑

Single-function

Intra-Agent Inter-Agent

30.68% 22.83%

0.0397 0.0681

– +71.5%

Multi-function

Intra-Agent Inter-Agent

28.61% 19.43%

0.0405 0.0794

– +96.0%

Note: Dup. Rate is the exact duplicate rate of patch pairs. AST Distance is the median normalized AST distance between patches. AST Gain measures the relative improvement of Inter-Agent over Intra-Agent. Arrows indicate better directions.

TABLE IX R EPAIR OVERLAP AND COMPLEMENTARITY AMONG EVIDENCE - PERSPECTIVE AGENTS . Subset

Agent-S

Agent-D

Agent-H

#Fixes

S only D only H only S ∩ D only S ∩ H only D ∩ H only S∩D∩H

• ◦ ◦ • • ◦ •

◦ • ◦ • ◦ • •

◦ ◦ • ◦ • • •

53 29 35 30 51 35 256

Total

390

350

377

489

Note: • indicates that the agent repairs the subset, and ◦ otherwise.

results show that the three perspectives are not redundant; each covers distinct bug scenarios, and their combination transforms diverse fault understandings into complementary repair strategies, substantially expanding repair coverage. Answer to RQ3. Multi-perspective agents improve reasoning diversity, reduce patch duplication, increase structural differences, and expand repair coverage beyond repeated sampling within a single perspective. D. RQ4: Contribution of Main Components To answer RQ4, we conduct a component-wise ablation study to evaluate each major component of CT-Repair. Due to the high computational cost of performing a full ablation on the complete benchmark, we primarily conduct this study on the 395 bugs from Defects4J v1.2. Using the full configuration as the baseline, we separately remove CPG, TEG, Agent-S, AgentD, Agent-H, and iterative refinement, and compare changes in the number of correctly repaired bugs. Ablation Results. As shown in Table X, removing any component degrades CT-Repair’s repair performance. Removing CPG and TEG reduces correctly repaired bugs from 245 to 216 and 224, corresponding to drops of 11.8% and 8.6%, respectively. This indicates that static structural information and dynamic execution information jointly support fault understanding, with CPG capturing code structure and dependencies and TEG providing runtime state changes and fault propagation information. Removing Agent-S, Agent-H, and Agent-D reduces correctly repaired bugs to 215, 223, and 231, corresponding to drops of 12.2%, 9.0%, and 5.7%, respectively. Although the three perspectives contribute differently, each independently improves the overall repair capability. Finally, disabling iterative refinement reduces correctly repaired bugs

TABLE X A BLATION STUDY OF CT-R EPAIR COMPONENTS ON D EFECTS 4J V 1.2. Multi-function

Single-function

All

∆ Impact (%)

w/o TEG w/o CPG w/o Agent-S w/o Agent-D w/o Agent-H w/o Iteration

53 52 52 54 54 57

171 164 163 177 169 176

224 216 215 231 223 233

−8.6% −11.8% −12.2% −5.7% −9.0% −4.9%

CT-RepairFull

65

180

245

Method

to 233, a drop of 4.9%, showing that validation-feedback-based strategy adjustment further improves patch quality. Answer to RQ4. The ablation study shows that CT-Repair’s effectiveness comes from the combined contributions of CPG, TEG, multi-perspective agents, and iterative refinement. VI. T HREATS TO VALIDITY

runtime environments, this runtime comparison merely reflects differences in overall cost magnitude. VII. R ELATED W ORK APR research has entered the era of LLMs [35]. AlphaRepair [29] and GAMMA [36] use LLMs through cloze-style generation or template-based prompting. ChatRepair [9] and Self-Debugging [37] further feed test failures or execution feedback back to the model to iteratively refine patches. Researchers have further introduced dynamic execution information into LLM-based APR. TraceFixer [38] adds local variable values and expected execution states during CodeT5 fine-tuning; LDB [26] tracks basic-block-level intermediate variable values and uses LLMs to validate program states step by step; DynaFix [25] instruments programs to collect variable states, control-flow paths, and call stacks for iterative repair. These studies show that static and dynamic evidence can support LLM-based repair. However, existing methods mostly provide such evidence as textual context, whereas CTRepair organizes it into queryable CPGs and TEGs to support hypothesis-driven evidence retrieval. Recent studies have shifted from one-shot patch generation to agentic repair based on feedback and tool interaction. Existing agent-based methods can be broadly grouped into single-agent tool-augmented repair and multi-agent collaborative repair. Among single-agent methods, ReinFix [21], RepairAgent [22], SWE-agent [39], AutoCodeRover [40], and SpecRover [41] support fault localization, root-cause analysis, and patch generation through repair-ingredient retrieval, repository navigation, code search, context retrieval, or reproduction tests. Among multi-agent methods, UniDebugger [42], MAGIS [43], SWEDebate [44], and TraceRepair [27] guide the repair workflow and validate or refine patches through role specialization, debate mechanisms, or runtime constraints. These methods mainly emphasize tool use, workflow collaboration, and patch selection. In contrast, CT-Repair focuses on deriving diverse repair strategies from distinct evidence perspectives, allowing agents to analyze faults independently from static, dynamic, and hybrid evidence before patch generation and thereby expanding the repair-strategy space.

Data Leakage. LLM training data may contain correct patches for benchmark bugs, affecting evaluation objectivity. To mitigate this risk, we reproduced ReinFix and RepairAgent using the same base model in Section V-A, where CT-Repair still achieves better repair effectiveness. The ablation study in Section V-D further verifies the consistent contributions of all main components. Nevertheless, we cannot completely rule out data leakage. In the future, we plan to further evaluate CTRepair on real-world bugs collected after the models’ trainingdata cutoff dates. Evaluation Assumptions. Our evaluation relies on two assumptions that may affect external validity. First, following recent APR studies [9], [21], [34], CT-Repair is evaluated under perfect fault localization. This setting aligns our evaluation with existing baselines and avoids additional bias from different FL tools, but may overestimate CT-Repair’s performance in realistic end-to-end repair scenarios. Second, Section V-C relies on LLM judges to assess reasoning diversity, potentially introducing bias. To mitigate this risk, we anonymize method names, use multiple LLM judges, and repeat the experiments. We also manually inspected 100 sampled cases and found strong agreement between human labels and aggregated LLM judgments (Cohen’s κ = 0.82). Future work will combine CTRepair with realistic FL techniques and conduct larger-scale human annotation. VIII. C ONCLUSION Repair Costs. LLM-based repair inevitably incurs additional inference costs. Due to the mixed-model configuration, CTWe presented CT-Repair, an agentic APR framework that Repair’s monetary cost is difficult to compare directly, so we combines queryable static and dynamic program representations focus on average token consumption. As shown in Table III with multi-perspective reasoning. CT-Repair uses a CPG to of Section V-A, CT-Repair consumes 178.35K tokens per bug represent structural dependencies and a TEG to organize on average, lower than ChatRepair (467K), AdverIntent-Agent runtime behavior after multi-stage filtering. Static-oriented, (438K), and RepairAgent (270K). Notably, both CT-Repair dynamic-oriented, and hybrid-oriented agents independently and RepairAgent use an FSM to guide repair, while CT-Repair derive root-cause hypotheses and repair strategies, thereby reduces token consumption by 33.9%. This suggests that the improving the diversity of both reasoning processes and summary-based state management mechanism introduced in candidate patches. Extensive experiments show that CT-Repair Section III-C effectively reduces cross-state context redundancy. effectively compresses redundant dynamic information, broadCT-Repair also reduces the average runtime from 920.0 s for ens the repair-strategy space, and improves the program repair RepairAgent to 158.9 s. Given the different base models and capability of LLMs while maintaining reasonable repair costs.

IX. DATA AVAILABILITY To support independent verification and replication, we provide anonymized research artifacts for CT-Repair at https: //anonymous.4open.science/r/CT-Repair-6D29/. R EFERENCES [1] L. Gazzola, D. Micucci, and L. Mariani, “Automatic software repair: a survey,” in Proceedings of the 40th International Conference on Software Engineering, ser. ICSE ’18. New York, NY, USA: Association for Computing Machinery, 2018, p. 1219. [Online]. Available: https://doi.org/10.1145/3180155.3182526 [2] K. Huang, Z. Xu, S. Yang, H. Sun, X. Li, Z. Yan, and Y. Zhang, “Evolving paradigms in automated program repair: Taxonomy, challenges, and opportunities,” ACM Comput. Surv., vol. 57, no. 2, Oct. 2024. [Online]. Available: https://doi.org/10.1145/3696450 [3] C. Le Goues, M. Pradel, and A. Roychoudhury, “Automated program repair,” Commun. ACM, vol. 62, no. 12, p. 56–65, Nov. 2019. [Online]. Available: https://doi.org/10.1145/3318162 [4] M. Monperrus, “Automatic software repair: A bibliography,” ACM Comput. Surv., vol. 51, no. 1, Jan. 2018. [Online]. Available: https://doi.org/10.1145/3105906 [5] Q. Zhang, C. Fang, Y. Ma, W. Sun, and Z. Chen, “A survey of learning-based automated program repair,” ACM Trans. Softw. Eng. Methodol., vol. 33, no. 2, Dec. 2023. [Online]. Available: https://doi.org/10.1145/3631974 [6] C. Le Goues, T. Nguyen, S. Forrest, and W. Weimer, “Genprog: A generic method for automatic software repair,” IEEE Transactions on Software Engineering, vol. 38, no. 1, pp. 54–72, 2012. [7] H. D. T. Nguyen, D. Qi, A. Roychoudhury, and S. Chandra, “Semfix: Program repair via semantic analysis,” in 2013 35th International Conference on Software Engineering (ICSE), 2013, pp. 772–781. [8] K. Liu, A. Koyuncu, D. Kim, and T. F. Bissyandé, “Tbar: revisiting template-based automated program repair,” in Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis, ser. ISSTA 2019. New York, NY, USA: Association for Computing Machinery, 2019, p. 31–42. [Online]. Available: https://doi.org/10.1145/3293882.3330577 [9] C. S. Xia and L. Zhang, “Automated program repair via conversation: Fixing 162 out of 337 bugs for $0.42 each using chatgpt,” in Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, ser. ISSTA 2024. New York, NY, USA: Association for Computing Machinery, 2024, p. 819–831. [Online]. Available: https://doi.org/10.1145/3650212.3680323 [10] Z. Chen, S. Kommrusch, M. Tufano, L.-N. Pouchet, D. Poshyvanyk, and M. Monperrus, “Sequencer: Sequence-to-sequence learning for endto-end program repair,” IEEE Transactions on Software Engineering, vol. 47, no. 9, pp. 1943–1959, 2021. [11] N. Jiang, T. Lutellier, and L. Tan, “Cure: Code-aware neural machine translation for automatic program repair,” in 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE), 2021, pp. 1161–1173. [12] Y. Li, S. Wang, and T. N. Nguyen, “Dlfix: context-based code transformation learning for automated program repair,” in Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering, ser. ICSE ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 602–614. [Online]. Available: https://doi.org/10.1145/3377811.3380345 [13] Y. Li, S. Wang, and T. N. Nguyen, “Dear: a novel deep learning-based approach for automated program repair,” in Proceedings of the 44th International Conference on Software Engineering, ser. ICSE ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 511–523. [Online]. Available: https://doi.org/10.1145/3510003.3510177 [14] T. Lutellier, H. V. Pham, L. Pang, Y. Li, M. Wei, and L. Tan, “Coconut: combining context-aware neural translation models using ensemble for program repair,” in Proceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis, ser. ISSTA 2020. New York, NY, USA: Association for Computing Machinery, 2020, p. 101–114. [Online]. Available: https://doi.org/10.1145/3395363.3397369

[15] X. Meng, X. Wang, H. Zhang, H. Sun, and X. Liu, “Improving fault localization and program repair with deep semantic features and transferred knowledge,” in Proceedings of the 44th International Conference on Software Engineering, ser. ICSE ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 1169–1180. [Online]. Available: https://doi.org/10.1145/3510003.3510147 [16] X. Meng, X. Wang, H. Zhang, H. Sun, X. Liu, and C. Hu, “Templatebased neural program repair,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), 2023, pp. 1456–1468. [17] M. Tufano, C. Watson, G. Bavota, M. D. Penta, M. White, and D. Poshyvanyk, “An empirical study on learning bug-fixing patches in the wild via neural machine translation,” ACM Trans. Softw. Eng. Methodol., vol. 28, no. 4, Sep. 2019. [Online]. Available: https://doi.org/10.1145/3340544 [18] H. Ye, M. Martinez, and M. Monperrus, “Neural program repair with execution-based backpropagation,” in Proceedings of the 44th International Conference on Software Engineering, ser. ICSE ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 1506–1518. [Online]. Available: https://doi.org/10.1145/3510003.3510222 [19] Q. Zhu, Z. Sun, Y.-a. Xiao, W. Zhang, K. Yuan, Y. Xiong, and L. Zhang, “A syntax-guided edit decoder for neural program repair,” in Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ser. ESEC/FSE 2021. New York, NY, USA: Association for Computing Machinery, 2021, p. 341–353. [Online]. Available: https://doi.org/10.1145/3468264.3468544 [20] Q. Zhu, Z. Sun, W. Zhang, Y. Xiong, and L. Zhang, “Tare: Type-aware neural program repair,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), 2023, pp. 1443–1455. [21] J. Zhang, K. Huang, J. Zhang, Y. Liu, and C. Chen, “Repair ingredients are all you need: Improving large language model-based program repair via repair ingredients search,” 2025. [Online]. Available: https://arxiv.org/abs/2506.23100 [22] I. Bouzenia, P. Devanbu, and M. Pradel, “Repairagent: An autonomous, llm-based agent for program repair,” in 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), 2025, pp. 2188–2200. [23] L. Wu, Y. Pei, Z. Yang, K. Li, Z. Lu, H. Tan, X. Lyu, J. Li, Y. Chen, P. Xue, K. Zheng, and D. Hao, “Debugrepair: Enhancing llm-based automated program repair via self-directed debugging,” 2026. [Online]. Available: https://arxiv.org/abs/2604.19305 [24] M. Haque, P. Babkin, F. Farmahinifarahani, and M. Veloso, “Towards effectively leveraging execution traces for program repair with code LLMs,” in Proceedings of the 4th International Workshop on Knowledge-Augmented Methods for Natural Language Processing, W. Shi, W. Yu, A. Asai, M. Jiang, G. Durrett, H. Hajishirzi, and L. Zettlemoyer, Eds. Albuquerque, New Mexico, USA: Association for Computational Linguistics, May 2025, pp. 160–179. [Online]. Available: https://aclanthology.org/2025.knowledgenlp-1.17/ [25] Z. Huang, L. Xu, C. Liu, W. Sun, X. Zhang, Y. Lei, M. Yan, and H. Zhang, “Dynafix: Iterative automated program repair driven by execution-level dynamic information,” 2025. [Online]. Available: https://arxiv.org/abs/2512.24635 [26] L. Zhong, Z. Wang, and J. Shang, “Debug like a human: A large language model debugger via verifying runtime execution step by step,” in Findings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V. Srikumar, Eds. Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 851–870. [Online]. Available: https://aclanthology.org/2024.findings-acl.49/ [27] J. Wu, T. Wu, M. Zhang, Y. Dong, and B. Shen, “Runtime execution traces guided automated program repair with multi-agent debate,” 2026. [Online]. Available: https://arxiv.org/abs/2604.02647 [28] C. S. Xia, Y. Wei, and L. Zhang, “Automated program repair in the era of large pre-trained language models,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), 2023, pp. 1482–1494. [29] C. S. Xia and L. Zhang, “Less training, more repairing please: revisiting automated program repair via zero-shot learning,” in Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ser. ESEC/FSE 2022. New York, NY, USA: Association for Computing Machinery, 2022, p. 959–971. [Online]. Available: https://doi.org/10.1145/3540250.3549101 [30] F. Yamaguchi, N. Golde, D. Arp, and K. Rieck, “Modeling and

discovering vulnerabilities with code property graphs,” in 2014 IEEE Symposium on Security and Privacy, 2014, pp. 590–604. [31] N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang, “Lost in the middle: How language models use long contexts,” Transactions of the Association for Computational Linguistics, vol. 12, pp. 157–173, 2024. [Online]. Available: https://aclanthology.org/2024.tacl-1.9/ [32] C. Snell, J. Lee, K. Xu, and A. Kumar, “Scaling llm test-time compute optimally can be more effective than scaling parameters for reasoning,” in International Conference on Learning Representations, Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu, Eds., vol. 2025, 2025, pp. 10 131–10 165. [Online]. Available: https://proceedings.iclr.cc/paper files/paper/2025/ file/1b623663fd9b874366f3ce019fdfdd44-Paper-Conference.pdf [33] R. Just, D. Jalali, and M. D. Ernst, “Defects4j: a database of existing faults to enable controlled testing studies for java programs,” in Proceedings of the 2014 International Symposium on Software Testing and Analysis, ser. ISSTA 2014. New York, NY, USA: Association for Computing Machinery, 2014, p. 437–440. [Online]. Available: https://doi.org/10.1145/2610384.2628055 [34] X. Yin, C. Ni, S. Wang, Z. Li, L. Zeng, and X. Yang, “Thinkrepair: Self-directed automated program repair,” in Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, ser. ISSTA 2024. New York, NY, USA: Association for Computing Machinery, 2024, p. 1274–1286. [Online]. Available: https://doi.org/10.1145/3650212.3680359 [35] Q. Zhang, C. Fang, Y. Xie, Y. Ma, W. Sun, Y. Yang, and Z. Chen, “A systematic literature review on large language models for automated program repair,” ACM Trans. Softw. Eng. Methodol., Mar. 2026, just Accepted. [Online]. Available: https://doi.org/10.1145/3799693 [36] Q. Zhang, C. Fang, T. Zhang, B. Yu, W. Sun, and Z. Chen, “Gamma: Revisiting template-based automated program repair via mask prediction,” in 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE), 2023, pp. 535–547. [37] X. Chen, M. Lin, N. Schaerli, and D. Zhou, “Teaching large language models to self-debug,” in International Conference on Learning Representations, B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun, Eds., vol. 2024, 2024, pp. 8746–8825. [Online]. Available: https://proceedings.iclr.cc/paper files/paper/2024/ file/2460396f2d0d421885997dd1612ac56b-Paper-Conference.pdf [38] I. Bouzenia, Y. Ding, K. Pei, B. Ray, and M. Pradel, “Tracefixer: Execution trace-driven program repair,” 2023. [Online]. Available: https://arxiv.org/abs/2304.12743 [39] J. Yang, C. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press, “Swe-agent: Agent-computer interfaces enable automated software engineering,” in Advances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., vol. 37. Curran Associates, Inc., 2024, pp. 50 528–50 652. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/ 2024/file/5a7c947568c1b1328ccc5230172e1e7c-Paper-Conference.pdf [40] Y. Zhang, H. Ruan, Z. Fan, and A. Roychoudhury, “Autocoderover: Autonomous program improvement,” in Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, ser. ISSTA 2024. New York, NY, USA: Association for Computing Machinery, 2024, p. 1592–1604. [Online]. Available: https://doi.org/10.1145/3650212.3680384 [41] H. Ruan, Y. Zhang, and A. Roychoudhury, “Specrover: Code intent extraction via llms,” in 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), 2025, pp. 963–974. [42] C. Lee, C. S. Xia, L. Yang, J.-t. Huang, Z. Zhu, L. Zhang, and M. R. Lyu, “UniDebugger: Hierarchical multi-agent framework for unified software debugging,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng, Eds. Suzhou, China: Association for Computational Linguistics, Nov. 2025, pp. 18 237–18 266. [Online]. Available: https://aclanthology.org/2025.emnlp-main.921/ [43] W. Tao, Y. Zhou, Y. Wang, W. Zhang, H. Zhang, and Y. Cheng, “Magis: Llm-based multi-agent framework for github issue resolution,” in Advances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., vol. 37. Curran Associates, Inc., 2024, pp. 51 963–51 993. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/ 2024/file/5d1f02132ef51602adf07000ca5b6138-Paper-Conference.pdf

[44] H. Li, Y. Shi, S. Lin, X. Gu, H. Lian, X. Wang, Y. Jia, T. Huang, and Q. Wang, “Swe-debate: Competitive multi-agent debate for software issue resolution,” 2025. [Online]. Available: https://arxiv.org/abs/2507.23348

Record · ID 366329 · SHA-256 e5b3362e46e84e68
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.