ConceptioArchivearXiv CS
arXiv CSopen access

EvoRAG: Making Knowledge Graph-based RAG Automatically Evolve through Feedback-driven Backpropagation

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
databasesdatamanagementsqlstorage
databases, sql, data management, storage

LLM LLM

KG

Query

EvoRAG: Making Knowledge Graph-based RAG Automatically Evolve through Feedback-driven Backpropagation Zhenbo Fu1 , Yuanzhe Zhang1 , Qiange Wang1 , Hao Yuan1 , Yuehao Xu1 , Enze Yi2 , Yanfeng Zhang1 , Ge Yu1 1 School of Computer Science and Engineering, Northeastern University, Shenyang 110819, China; 2 Northeast Electric Power Research Institute of State Grid Liaoning Electric Power Supply Co, Ltd, Shenyang, Liaoning,

arXiv:2604.15676v1 [cs.DB] 17 Apr 2026

Human Feedback

ABSTRACT

110004, China Prompt {fuzhenbo,zhangyz,yuanhao}@stumail.neu.edu.cn,[email protected] {wangqg,zhangyf,yuge}@mail.neu.edu.cn,[email protected] Feedback

EvolveRAG

QueryKnowledge Graph-based Retrieval-Augmented Generation (KG-

Prompt

RAG) has emerged as a promising paradigm for enhancing LLM reasoning by retrieving multi-hop paths from KGs. However, existing KG-RAG frameworks often underperform in real-world scenarios because Refine the pre-captured knowledge dependencies are not + tailored to the downstream task or its evolving requirements. These frameworks struggle to adapt to task-specific requirements and KG lack mechanisms to filter low-contribution knowledge during generation. We observe that feedback on generated responses offers effective supervision for improving KG quality, as it directly reflects user expectations and provides insights into the correctness and usefulness of the output. However, a key challenge lies in effectively linking response-level feedback to triplet-level contribution evaluation and knowledge updates in the KG. In this work, we propose EvoRAG, a self-evolving KG-RAG framework that leverages the feedback over generated responses to continuously refine the KG and enhance reasoning accuracy. EvoRAG introduces a feedback-driven backpropagation mechanism that attributes feedback to retrieved paths by measuring their utility for response and propagates this utility back to individual triplets, supporting fine-grained KG refinements towards more adaptive and accurate reasoning. Through EvoRAG, we establish a closed loop that couples feedback, LLM, and graph data, continuously enhancing the performance and robustness in real-world scenarios. Experimental results show that EvoRAG improves reasoning accuracy by 7.34% over state-of-the-art KG-RAG frameworks. The source code has been made available at https://github.com/iDC-NEU/EvoRAG.

1

INTRODUCTION

Retrieval-Augmented Generation (RAG) [17, 35, 101] empowers Large Language Models (LLMs) to improve response quality by leveraging external knowledge, and has been widely adopted in various domains [80, 83, 94]. Among RAG paradigms, Knowledge Graph-based RAG (KG-RAG) [60, 97, 104] has gained increasing attention for its ability to transform the textual corpus into structured knowledge graphs (KGs) [104], capturing rich semantic information and entity-level relations. Given a query 𝑞, the core idea of KG-RAG is to retrieve a relevant knowledge subgraph (KSG) from the KG and feed it together with 𝑞 to the LLM for response generation. The retrieved KSG is typically organized into a sequence

EvoRAG Query

Response

Refinement

Prompt

Q

R

Prompt

+ LLM KG (a) Conventional KG-RAG

LLM KG (b) KG-RAG with Feedback Backprop.

Figure 1: Comparison of conventional KG-RAG and EvoRAG. EvoRAG introduces a backpropagation mechanism that propagates response-level feedback to individual triplets, enabling continuous KG refinements.

of reasoning paths, ordered chains of triplets that capture multihop semantic connections, where each triplet follows the form <ℎ𝑒𝑎𝑑 𝑒𝑛𝑡𝑖𝑡𝑦, 𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛, 𝑡𝑎𝑖𝑙 𝑒𝑛𝑡𝑖𝑡𝑦>. Although KG-RAG frameworks have demonstrated promising results by leveraging the structural information, they often underperform in real-world scenarios due to a mismatch between the pre-captured dependencies in KG and the task-specific requirements [97, 104]. This mismatch primarily stems from two structural limitations of KG-RAG frameworks: insufficient adaptability to downstream reasoning tasks and limited dynamicity in handling knowledge freshness and reliability. These factors are overlooked in recent KG refinement studies [58, 68], which mainly focus on detecting factual errors and semantic inconsistencies. Firstly, adaptability refers to the ability to organize and retrieve knowledge in a manner that supports real-time queries. Naively applying RAG with existing KG frameworks often retrieves information that is semantically correct but contributes little to the query. For example, they may retrieve irrelevant triplets or overlook long-range dependencies required for complex reasoning [10, 24, 98]. This leads to ineffective reasoning and poor knowledge utilization. Secondly, dynamicity refers to the ability to detect and eliminate outdated or erroneous information. Existing KG-RAG frameworks are generally built on static KGs. They lack mechanisms for continuously detecting and eliminating invalid knowledge, such as outdated relations or KSGs that are no longer required by the downstream applications [27, 95], leading to degraded reasoning performance. These limitations reveal a data management problem: how to continuously refine KG to improve adaptability and support dynamic evolution.

Split

 Corpus 1 KG Construction

1 KG Construction

We observe that feedback derived from generated contexts, such as signals on correctness, consistency, and user satisfaction, provides a valuable data source for continuously improving the KG. In LLM-based systems, such feedback is commonly obtained by leveraging LLMs to evaluate generated responses [32, 34, 47, 50, 59, 63, 75, 86]. Similar feedback may also come from human judgments [2, 69, 76, 84, 90] or comparisons with available ground truth [23, 78, 100]. Such feedback directly reflects reasoning effectiveness and can serve as a guidance signal for continuously refining the KG to improve task performance. However, a mismatch in granularity between feedback and KG refinement makes such integration difficult: feedback assesses overall response quality over multiple retrieved paths, whereas KG refinement operates at the level of individual triplets, leading to a mismatch as feedback reflects the combined effect of multiple interacting triplets rather than localized signals for each discrete and reusable knowledge units. As a result, feedback needs to be correctly attributed to reusable triplets across multiple reasoning paths and queries. Unlike traditional machine learning systems that explicitly construct a computation graph during the forward pass and rely on backpropagation to propagate gradients to model parameters, feedback-driven KG refinement lacks an inherent mechanism to propagate feedback to underlying triplets, requiring carefully designed algorithms to establish this association. In this work, we present EvoRAG, a self-evolving KG-RAG framework that utilizes feedback to continuously refine the KG, enhancing the accuracy and relevance of the reasoning process, as illustrated in Figure 1. To realize this, EvoRAG introduces a feedbackdriven backpropagation mechanism that establishes a connection between response-level feedback and triplet-level knowledge updates through two key stages. Firstly, EvoRAG attributes the feedback to individual reasoning paths by measuring their utility, i.e., how much each path contributes to the response reflected in the feedback. Reasoning paths are used as intermediates because they directly guide the LLM in generating the response, while also providing a natural bridge to the underlying triplets from which they are constructed. Secondly, EvoRAG propagates the path utility to update the contribution score of each involved triplet. These scores are maintained as learnable parameters and continuously updated across interactions, allowing EvoRAG to prioritize task-relevant knowledge. Through EvoRAG, we construct a closed-loop mechanism that tightly couples feedback, LLM reasoning, and triplet-level KG updates, continuously improving its accuracy in real-world queries. In summary, our primary contributions are as follows:

Insert

Offline Data Processing

(Eva, HasBrother, Bob) (Bob, LivesIn, Niva) (Bob, OwnsPet, Dog) (Bob, Like, Golf) (Bob, WorksWith, Tom) (Tom, WorksAt, Google)

Graph DB

Online KGRAG Service

Text

Triplets

2 KSG Retrieval Retriever Bone

Query: Where does Eva’s brother work?

Dog

Eva

2020

Bob

3 Response Generation Eva’s brother works at Zelo.

Prompt

Alice

Outdated fact

Extract

Niva

Zelo

Tom

LLM

Google

Pruning

Figure 2: The overall workflow of KG-RAG.

2 PRELIMINARY 2.1 KG-RAG RAG. Large language models (LLMs) [1, 56, 73, 74, 92] have recently demonstrated remarkable potential in handling complex tasks [20, 41, 45, 66, 67, 93]. However, they are susceptible to hallucinations when generating answers for queries that require information beyond their knowledge [12, 19, 26, 29, 52, 64, 65, 102]. To address these limitations, Retrieval-Augmented Generation (RAG) [17, 36, 38, 97, 104] has emerged as a promising approach, enhancing LLMs with external knowledge by retrieving relevant information, improving the accuracy and relevance of the generated responses. Among existing RAG methods, KG-RAG [7, 10, 11, 15, 21, 24, 30, 37, 39, 40, 42, 51, 89] has attracted considerable attention due to its ability to leverage structured knowledge [8, 16, 77]. KG-RAG organizes external corpus in a KG, enabling a more comprehensive understanding by effectively leveraging interconnections between pieces of knowledge, making it well-suited for complex, reasoningintensive tasks. KG. Knowledge Graph (KG) encodes a wide range of knowledge in the form of triplets: 𝐺 = {(𝑒 1, 𝑟, 𝑒 2 )|𝑒 1, 𝑒 2 ∈ 𝐸, 𝑟 ∈ 𝑅}, where 𝐸 denotes the set of nodes (entities) and 𝑅 denotes the set of directed edges that signify relations between those entities. The triplet (𝑒 1, 𝑟, 𝑒 2 ) is the fundamental building unit of knowledge in a KG, where 𝑒 1 denotes the subject entity, 𝑟 denotes the predicate, and 𝑒 2 denotes the object entity. Reasoning path refers to the sequence of entities and relations 𝑟𝑖 𝑟1 𝑟2 traversed from a starting entity: 𝐿 = 𝑒 0 −→ 𝑒 1 −→ · · · − → 𝑒𝑖 , where 𝑒 0 is the starting entity, 𝑒𝑖 ∈ 𝐸 is the 𝑖-th entity, and 𝑟𝑖 ∈ 𝑅 is the 𝑖-th relation. Such paths reveal higher-order dependencies and provide interpretable evidence for downstream tasks, e.g., link prediction, question answering, and recommendation.

• We propose a feedback-driven backpropagation mechanism that utilizes feedback to drive the evolution of the KG-RAG framework, enhancing adaptability and improving overall accuracy. • We introduce an effective mechanism that resolves the challenge of mapping response-level feedback to triplet-level updates, enabling fine-grained KG refinement and retrieval improvement. • We develop EvoRAG, a self-evolving KG-RAG framework that adapts to dynamic tasks and requirements. Experimental results demonstrate that EvoRAG improves accuracy by 7.34% compared to the state-of-the-art KG-RAG frameworks.

The workflow of KG-RAG. As illustrated in Figure 2, the KGRAG frameworks typically consist of three key components: KG construction, KSG retrieval, and response generation. (➊) KG construction. KG construction refers to the process of mining structured knowledge from large-scale external corpora, which can be performed using specialized information extraction tools, such as OpenIE [54], or by leveraging LLMs. Specifically, the external corpus is divided into multiple text chunks, each of which undergoes entity recognition and relation extraction to construct 2

KRAG

EvoRAG (ours)

20 15 10 5 0

to downstream reasoning tasks and lack of dynamicity in maintaining knowledge validity. Firstly, the underlying KGs are inherently query-agnostic, as they are constructed offline without considering the specific query intent. This often leads to two common issues: (1) Irrelevant facts introduce noise. As shown in Figure 2, when the user asks “Where does Eva’s brother work?”, the retriever returns unrelated facts such as (𝐵𝑜𝑏, 𝐿𝑖𝑣𝑒𝑠𝐼𝑛, 𝑁𝑖𝑣𝑎), which do not contribute to the answer. (2) Failing to capture long-range dependencies. For example, answering that Eva’s brother works at Google

25

Incorrect Rate (%)

Incorrect Rate (%)

25

#IF

#LP

#OI

20 15 10 5 0

(a) RGB dataset

#IF

#LP

#OI

(b) Multihop dataset

𝐻𝑎𝑠𝐵𝑟𝑜𝑡ℎ𝑒𝑟

𝑊 𝑜𝑟𝑘𝑠𝐴𝑡

𝑇𝑜𝑚 −−−−−−−→ 𝐺𝑜𝑜𝑔𝑙𝑒. However, pre-defined retrieval strategy (e.g., 2-hop retrieval) fails to capture such multi-hop dependencies, causing critical information to be omitted. Secondly, existing KG-RAG systems typically rely on static KGs and lack mechanisms for timely updates [7, 24, 37, 42, 51]. As a result, outdated or invalid information may persist and mislead the LLM. For example, the KG may still contain (𝐵𝑜𝑏,𝑊 𝑜𝑟𝑘𝑠𝐴𝑡, 𝑍𝑒𝑙𝑜), even though the company ceased operations in 2020. In real-world deployments, a KG-RAG system typically serves a large number of concurrent users who interact with a shared knowledge graph. Multiple users often issue queries focused on the same regions of the graph, such as popular entities, ongoing events, or trending topics. Consequently, the system frequently receives highly similar or even identical queries within the same local subgraph. This concentrated access pattern amplifies the aforementioned issues: irrelevant facts and missed long-range dependencies repeatedly affect many users, and outdated triplets continue to mislead multiple responses. To systematically quantify the impact of the above limitations on KG-RAG reasoning, we analyze the erroneous responses in the RGB [9] and MultiHop [72] datasets. We categorize the errors into four types: (1) irrelevant facts (IF), (2) long reasoning paths (LP), (3) outdated information (OI), and (4) other errors caused by missing knowledge in the KG or hallucinations by the LLM (not the focus of this work). To provide a baseline, we adopt a state-of-the-art instance from prior work [6], referred to as KRAG. Firstly, KRAG constructs the retrieval subgraph by expanding from the query entities to include all entities within two hops. Then, a semantic model (BAAI/bge-large-en-v1.5 [82]) is employed to select reasoning paths within the subgraph that are most similar to the query as the retried results. As shown in Figure 3, IF, LP, and OI together account for over half of all errors, with average proportions of 17.9%, 20.7%, and 11.9%, respectively, making them the primary bottlenecks that fundamentally limit the effectiveness of KG-RAG in real-world reasoning tasks. To address these critical limitations, we propose EvoRAG, which incorporates a feedback-driven backpropagation mechanism to explicitly target and mitigate these error sources, resulting in substantial improvements in reasoning accuracy. In the following sections, we present the design and implementation of EvoRAG in detail.

triplets, such as (𝑇𝑜𝑚,𝑊 𝑜𝑟𝑘𝑠𝐴𝑡, 𝐺𝑜𝑜𝑔𝑙𝑒). These extracted triplets are then inserted into a graph database to form a complete KG, which in turn supports subsequent online RAG services. (➋) KSG retrieval. KSG retrieval refers to the process of extracting relevant reasoning paths from a KG for a given query, which typically consists of three steps: Query Entity Recognition, Subgraph Extraction, and Path Retrieval [6]. Firstly, Query Entity Recognition operates on the input query to identify query entities and align them with corresponding nodes in the KG. This step typically employs a semantic similarity model to extract entities whose surface forms or contextual embeddings closely match the query expressions. Secondly, Subgraph Extraction operates on the entire KG to extract a subgraph related to the query entities. The goal of this step is to reduce the search space and improve the effectiveness and efficiency of retrieval. Thirdly, Path Retrieval operates on the extracted subgraphs to collect the top𝑀 reasoning paths starting from the query entities and connecting them to potential target answers (e.g., one of the reasoning paths 𝐻𝑎𝑠𝐵𝑟𝑜𝑡ℎ𝑒𝑟

𝑊 𝑜𝑟𝑘𝑠𝐴𝑡

for the entity “𝐸𝑣𝑎” is 𝐸𝑣𝑎 −−−−−−−−−→ 𝐵𝑜𝑏 −−−−−−−→ 𝑍𝑒𝑙𝑜). Both Subgraph Extraction and Path Retrieval can be implemented using graph algorithms (e.g., neighbor expansion, personalized PageRank) or semantically enhanced models (e.g., embedding models, LLMs). By integrating these approaches, existing KG-RAG implementations can be broadly covered within a unified formulation [7, 10, 11, 14, 24, 30, 37, 39, 40, 42, 51, 70, 98]. (➌) Response generation. Response generation refers to the process of producing the response conditioned on both the user query and the reasoning paths. Specifically, the retrieved paths are combined with the query to form an augmented prompt, which is subsequently fed into an LLM for generation. Unlike traditional RAG, which ranks individual text segments by semantic similarity, KG-RAG links entities and relations to compose multi-hop evidence spanning multiple passages, which enables cross-passage inference and yields more coherent and better grounded responses.

2.2

𝐶𝑜𝑙𝑙𝑒𝑎𝑔𝑢𝑒

requires a 3-hop reasoning path: 𝐸𝑣𝑎 −−−−−−−−−→ 𝐵𝑜𝑏 −−−−−−−−→

Figure 3: Proportion of error types in KRAG and EvoRAG. #IF, #LP, and #OI represent three error types: irrelevant facts, long reasoning paths, and outdated information, respectively.

Motivation

KG-RAG has shown great potential in enhancing generation quality by leveraging the structural information of external knowledge graphs. However, in real-world applications, there exists a notable gap between the predefined dependencies encoded in the KG and the adaptive and dynamic requirements of user queries. This gap results in two fundamental limitations: lack of adaptability

3

SYSTEM OVERVIEW

We propose EvoRAG, a self-evolving KG-RAG framework that continuously improves reasoning effectiveness through fine-grained knowledge refinement and retrieval optimization. As illustrated in Figure 4, the system operates in an iterative loop where users 3

Feedback §4.1

Feedback-based Path Evaluation

Query

§4.2

Path-level Utility Gradient Backpropagation §4.3 Feedback-driven Backpropagation

Hybrid Prioritybased Retrieval §5.2

Guide Triplet-level Score Relation-centric KG Evolution §5.1

triplets, ensuring structural coherence and knowledge quality. Then, EvoRAG adopts a hybrid priority-based retrieval strategy that integrates semantic relevance with feedback-derived contribution scores. This hybrid mechanism balances query-dependent relevance and query-independent reliability, guiding retrieval toward both contextually appropriate and historically effective knowledge.

Response

Knowledge Graph 160

20

+

145

65

Feedback-guided KG Management

4

FEEDBACK-DRIVEN BACKPROPAGATION

In this section, we first introduce the sources of feedback, followed by the feedback-based path evaluation for computing path-level utility, and finally introduce the gradient backpropagation mechanism that propagates this utility to individual triplets to update the contribution scores.

Figure 4: EvoRAG system overview.

continuously submit queries. For each query, EvoRAG retrieves multiple reasoning paths from the KG as contextual knowledge, which are then fed into an LLM to generate responses. To further enhance reasoning quality, EvoRAG introduces a feedback-driven backpropagation mechanism (see Section 4) that transforms coarse-grained feedback into fine-grained supervision over individual data triplets. This mechanism enables the system to iteratively refine both the underlying knowledge and the retrieval process. We implement this mechanism in two stages. In the first stage, we propose a feedback-based path evaluation module that propagates response-level feedback to the retrieved reasoning paths. The output is a path-level utility score that quantifies its overall quality and the unique contribution to the generated response. Instead of directly evaluating individual triplets, we adopt reasoning paths as the intermediate unit. This design is motivated by the observation that triplets do not function in isolation. During inference, responses are generated through multi-hop reasoning over paths composed of multiple triplets, where the contribution of each triplet is inherently context-dependent, shaped by other triplets within the same path as well as the semantics of the input query. Consequently, directly assigning utility to individual triplets may lead to misalignment, as a triplet that is beneficial in one path can be misleading in another. Reasoning paths therefore provide a more appropriate intermediate abstraction, bridging high-level feedback and fine-grained knowledge units. In the second stage, we propose a gradient backpropagation module that further propagates path-level utility to the constituent triplets. Specifically, the module computes the gradients of the utility expectation over paths with respect to path utility and distributes them to triplets based on their roles within each path. To enable consistent supervision across queries, we associate each triplet with a learnable contribution score, which serves as the foundation for fine-grained reasoning supervision and allows the KG to gradually evolve toward task-specific relevance. This score is defined as a dimensionless and non-negative coefficient that estimates the expected influence of a triplet on response quality. In practice, all scores are initialized to 100 and dynamically updated over time based on accumulated feedback signals. With the updated contribution scores, EvoRAG provides feedbackguided KG management, enabling both KG evolution and retrieval optimization (see Section 5). Specifically, EvoRAG implements relationcentric KG evolution, which includes relation fusion and suppression. Relation fusion introduces shortcut edges between frequently co-occurring high-utility triplets to enhance long-range reasoning effectiveness, while relation suppression downweights low-utility

4.1

Feedback

Overview of feedback and its sources. In interactive systems, responses are often accompanied by signals that indicate their quality with respect to task objectives, such as factual correctness, reasoning soundness, or relevance to user intent. We refer to such signals as feedback. In modern LLM-based systems, feedback is commonly obtained by using LLMs to evaluate generated responses and produce quality assessments [32, 34, 47, 50, 59, 63, 75, 86]. More generally, similar feedback may also arise from other sources. For example, human assessments may appear as ratings, binary judgments, or survey responses [2, 69, 76, 84, 90]; and when ground truth is available, comparison with reference answers provides an objective measure of correctness [23, 78, 100]. In this work, we primarily rely on feedback derived from LLMbased evaluation, following prior studies that use LLMs to assess response quality [32, 47, 50, 63]. Specifically, given the ground-truth answers available in our datasets, the evaluator LLM is prompted to assess each generated response and assign a scalar satisfaction score from 1 to 5, providing a reliable and objective assessment. A score of 1 indicates complete dissatisfaction (e.g., irrelevant or factually incorrect responses), while a score of 5 indicates full satisfaction (e.g., responses that are correct, relevant, and well aligned with the query intent), with intermediate scores (2–4) reflecting varying levels of partial adequacy. To demonstrate the generality of EvoRAG, we further evaluate it using other kinds of feedback, including human judgments and ground-truth based feedback using F1 scores, as shown in Section 6.4. Challenge. Although unified feedback scoring provides a consistent evaluation signal, a fundamental challenge lies in its granularity mismatch. Feedback is only available at the response level, offering coarse-grained assessments of overall quality, whereas effective KG refinement requires fine-grained identification of the specific triplets that influence reasoning and retrieval. This mismatch limits precise knowledge adjustment and constrains overall system performance. To address this challenge, our framework translates coarse responselevel feedback into fine-grained supervision through a two-step decomposition process. Feedback is first transformed into path-level utility via a path evaluation module, which captures the contribution of entire reasoning trajectories. The utility is then propagated to constituent triplets through gradient-based backpropagation process to update their contribution scores. 4

𝑼𝒊

Path Evaluation

Response & Feedback -Response-

Bob

Tom

+0.7

Eva

Bob

Dog

0

-Feedback-

Eva

Bob

Zelo

-0.8

0

Eva

Bob

Niva

0

𝜕𝑝𝑎𝑡ℎ 𝜕𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛

+19 Eva

+5

Bob

2.4

Dog

1.4

Zelo

0.3

Niva

1.4

+0 -11 +0

Gradient Propagation

Tom

𝜕𝐿𝑜𝑠𝑠 𝜕𝑝𝑎𝑡ℎ

is low; otherwise, the update is suppressed. This design prevents erroneous utility updates caused by spurious grounding or contradictory evidence, ensuring reliable path-level attribution.

𝓛 = −𝒍𝒐𝒈(𝔼(𝑼))

Eva

Based on the context, Eva’s brother works at Zelo.

4.3

0.36 Loss

Figure 5: Feedback-driven Backpropagation. We first use an evaluation function to compute the utility of each reasoning path based on feedback, and then transform it into gradients that are propagated to individual triplets.

4.2

4.3.1 Forward Computation. In the forward retrieval phase, each query retrieves multiple reasoning paths from the KG, which is composed of a sequence of triplets. Each triplet 𝑡 is associated with a selection probability (𝑃 (𝑡) ∈ [0, 1]), which reflects its semantic relevance to the query and the contribution score. Formally, the 𝑃 (𝑡) can be defined as:

Feedback-based Path Evaluation

To enable effective learning from response-level feedback, we define a utility evaluation function 𝑓 that assigns utility to each reasoning path by analyzing its contribution to the generated response. The evaluation considers the semantics of the input query, the retrieved paths, the generated response, and the feedback score. Specifically, when feedback indicates a correct response, 𝑓 rewards paths that semantically support the answer; when the response is incorrect, it attributes negative utility to paths that may have misled generation. Instead of assigning binary judgments, the evaluation function produces continuous utility in a bounded range (i.e., [−1, 1]), enabling nuanced assessment of each path’s contribution. We formalize the path utility evaluation as: 𝑈 {𝐿} = 𝑓 (𝑞, 𝑅𝑞 , 𝐿, 𝐹𝑆),

Gradient Backpropagation

Once the utility of each reasoning path is evaluated, the next step is to propagate this utility back to individual triplets, refining their contribution scores so that future retrievals are more likely to select high-utility paths. To derive backpropagation, we first need to introduce how the contribution score influences the path selection strategy during forward computation.

𝑃 (𝑡) = (1 − 𝛼)𝑆𝑟 (𝑡) + 𝛼𝑆𝑐 (𝑡),

(2)

where 𝑆𝑟 (𝑡) ∈ (0, 1] measures the semantic similarity between the triplet 𝑡 and the query 𝑞, providing a query-specific score that aligns path selection with the current query. In contrast, 𝑆𝑐 (𝑡) ∈ [0, 1] is the learnable contribution score by normalization, which accumulates feedback across multiple queries and rounds, thereby decoupling triplet assessment from any single query. The parameter 𝛼 ∈ [0, 1] is a learnable trade-off between 𝑆𝑟 and 𝑆𝑐 , allowing the system to balance the current query’s guidance with accumulated knowledge from previous feedback. Based on these probabilities, the priority of a reasoning path (𝐿𝑖 ∈ 𝐿) can be defined as follows:

(1)  exp

where 𝑞 is the user query, 𝑅𝑞 is the response, 𝐿 is the set of retrieved reasoning paths, and 𝐹𝑆 is the feedback score. In practice, 𝑓 can be instantiated by a range of semantic models. In this work, we use an LLM as a constrained path-level scoring function. Given the retrieved paths, the query, the generated response, and the feedback score, the LLM assigns a utility score to each path based on its relevance, following the LLM-as-a-judge paradigm [91, 103]. This process is analogous to RAG guided by the external knowledge: the LLM evaluates paths using the provided context instead of producing new responses, and is therefore minimally affected by hallucination. To capture the multi-faceted nature of path quality under feedback, we decompose path evaluation into three complementary dimensions. Supportiveness measures whether a path provides evidence that supports or undermines the generated response, conditioned on response correctness as determined by feedback. Higher Supportiveness increases path utility, while lower values decrease it. Fidelity and Conflict serve as auxiliary metrics for cross-validation. Fidelity measures how much a path contributes to the response, with higher values meaning greater contribution. Conflict measures whether a path contradicts the response, with lower values meaning less contradiction. The LLM assigns a score in [−1, 1] to each dimension based on the query, retrieved paths, generated response, and feedback. Importantly, a path’s utility is updated based on Supportiveness only when Fidelity is high and Conflict

𝑃 (𝐿𝑖 ) = Í

1 Í 𝑡 ∈𝐿𝑖 log 𝑃 (𝑡 ) |𝐿𝑖 |

𝐿 𝑗 ∈𝐿 exp





1 Í 𝑡 ∈𝐿 𝑗 log 𝑃 (𝑡 ) |𝐿 𝑗 |

,

(3)

where 𝐿𝑖 denotes a reasoning path and |𝐿𝑖 | is its hop length (i.e., the number of triplets it contains). The length normalization takes the log-average of triplet probabilities to eliminate the inherent bias of the multiplicative formulation toward shorter paths, enabling fair comparison across reasoning paths of different hop lengths. The Formula 3 not only ensures relevance to the current query, but also enables generalization to unseen queries. When a new query arrives, the similarity score 𝑆𝑟 is freshly calculated at query time, ensuring specificity to the user’s intent and filtering out irrelevant knowledge, where the contribution score 𝑆𝑐 transfers the accumulated feedback on triplet quality to guide the selection of the reasoning paths. In this way, the model adapts to unseen queries by letting 𝑆𝑟 provide query-time specificity and 𝑆𝑐 provide cross-query transferability, yielding path priorities that balance immediate relevance with long-term utility. Loss function. We formulate an objective that maximizes the expected utility of the retrieved reasoning paths. Formally, the loss function is defined as follows: ! 𝑈 (𝐿𝑖 ) + 1 L = − log(E[𝑈 (𝐿)]) = − log 𝑃 (𝐿𝑖 ) , 2 𝐿 ∈𝐿 ∑︁ 𝑖

5

(4)

where 𝐿 denotes the set of all retrieved reasoning paths. Optimizing expected utility encourages the system to adapt this preference with observed effectiveness, guiding future retrievals toward reasoning paths that consistently yield high-quality responses.

according to the utility of the paths they participate in. By standard results in stochastic convex optimization [4], with a properly chosen learning rate, the iterative updates of 𝑆𝑐 (𝑡) converge to a stationary point. At convergence, the induced distribution over paths aligns with their utility: paths with consistently high utility receive a larger selection probability, while harmful or low-utility paths are gradually down-weighted. This ensures that the retrieval process adaptively emphasizes informative reasoning paths, reflecting the accumulated feedback across queries. Noise tolerance. Feedback-driven backpropagation is inherently noise-tolerant to erroneous feedback and LLM scoring errors in two folds. First, contribution scores are updated cumulatively across queries, allowing occasional errors to be averaged over time. Second, path utility is updated by supportiveness (the feedback-driven dimension) only when fidelity is high and conflict is low; otherwise, the update is suppressed. This cross-validation mechanism prevents spurious feedback or LLM misjudgments from distorting path evaluation, while reinforcing consistently supported paths.

4.3.2 Backward Computation. We compute the gradient of the loss L with respect to the contribution score 𝑆𝑐 (𝑡) as follows:

∇𝑆𝑐 (𝑡 ) L = −

∑︁ 𝛼 2E[𝑈 (𝐿)] 𝑡 ∈𝐿

Î

𝑔∈𝐿𝑖 𝑃 (𝑔)

𝑃 (𝑡)

𝑉𝑖 ,

(5)

𝑖

𝑉𝑖 =

∑︁ 𝑃 (𝐿𝑖 ) © ª 𝑃 (𝐿 𝑗 )𝑈 (𝐿 𝑗 ) ® , ­𝑈 (𝐿𝑖 ) − |𝐿𝑖 | 𝐿 𝑗 ∈𝐿 « ¬

(6)

where 𝑉𝑖 denotes the deviation of a path’s utility from the expected utility across all paths. The gradient ∇𝑆𝑐 (𝑡 ) L quantifies how the triplet 𝑡 contributes to reasoning paths with above- or belowaverage utility. If 𝑡 frequently appears in high-utility paths (i.e., Í paths where 𝑈 (𝐿𝑖 ) > 𝐿 𝑗 ∈𝐿 𝑃 (𝐿 𝑗 )𝑈 (𝐿 𝑗 )), the gradient becomes negative, reducing the loss and thereby increasing the contribution score 𝑆𝑐 (𝑡). Conversely, if 𝑡 tends to appear in low-utility paths, the score will be reduced accordingly. Similarly, the gradient of 𝛼 can be computed as follows:

∇𝛼 L =

∑︁ 𝑆𝑟 (𝑡) − 𝑆𝑐 (𝑡) ∑︁ 𝑡 ∈𝑇

2E[𝑈 (𝐿)]

𝑡 ∈𝐿𝑖

4.3.4 Complexity Analysis. We analyze the time and space complexity of EvoRAG by considering two components. For feedbackdriven backpropagation, contribution scores are propagated along each reasoning path and aggregated to update triplets. The computation scales with the number of paths 𝑃, the average path length 𝐻 , and the number of triplets 𝑇 , yielding a time complexity of O (2𝑃 +𝐻 2 𝑃 +𝑇 ). The memory is dominated by storing path utilities and triplet-level contribution scores, yielding a space complexity of O (𝑃 +𝑇 ). For path evaluation, all reasoning paths are processed in a single LLM call. Let 𝐶 LLM (𝐿) denote a black-box cost function of an LLM call with input length 𝐿, abstracting away model architecture and deployment details. Since the 𝐿 is proportional to the length of all paths, i.e., 𝐿 = O (𝐻𝑃), the overall time complexity of path evaluation is O (𝐶 LLM (𝐻𝑃)), while the space overhead is dominated by the prompt and intermediate LLM states.

Î

𝑔∈𝐿𝑖 𝑃 (𝑔)

𝑃 (𝑡)

𝑉𝑖 ,

(7)

where 𝑇 denotes the set of all triplets. Finally, we apply gradient descent to update both the contribution score and the parameter 𝛼: 𝑆𝑐 (𝑡) = 𝑆𝑐 (𝑡) − 𝜂∇𝑆𝑐 (𝑡 ) L,

(8)

𝛼 = 𝛼 − 𝜂∇𝛼 L,

(9)

where 𝜂 is the learning rate. This optimization process ensures that triplets associated with helpful reasoning paths become more prominent, while those consistently tied to unproductive reasoning are gradually suppressed.

5

4.3.3 Accuracy Analysis. To evaluate the effectiveness of our feedbackdriven backpropagation, we analyze both its convergence and its behavior under noisy feedback. The convergence analysis provides theoretical guarantees for stable and correct updates, while the noise tolerance analysis considers the impact of occasional erroneous feedback on the updates.  Convergence analysis. Since 𝑈 (𝐿) + 1 /2 is in the range of Í [0, 1] and 𝐿𝑖 ∈𝐿 𝑃 (𝐿𝑖 ) = 1, the expected utility E[𝑈 (𝐿)] is bounded within (0, 1], which ensures that the loss L is non-negative and upper-bounded. This boundedness prevents gradient explosion and provides a stable optimization objective. Moreover, L is convex with respect to the path distribution, implying that the objective function admits a unique global optimum over probability distributions [5]. In our framework, the contribution score 𝑆𝑐 (𝑡) only affects L through the softmax-based path probability 𝑃 (𝐿𝑖 ) (Formula 3). The softmax mapping is smooth and strictly monotonic, which guarantees that the gradient ∇𝑆𝑐 (𝑡 ) L exists and is Lipschitz-continuous [53]. This gradient updates the contribution score of each triplet 6

FEEDBACK-GUIDED KG MANAGEMENT

This section presents the feedback-guided KG management mechanism of EvoRAG, which involves relation-centric KG evolution and hybrid priority-based retrieval involving similarity and contribution score for improving retrieval quality.

5.1

Relation-centric KG Evolution

We refine the KG based on feedback-derived contribution scores that reflect how each triplet supports reasoning. Instead of updating entities, EvoRAG adopts a relation-centric evolution strategy. Since entities are shared across multiple triplets, modifying them would introduce cascading and ambiguous effects. In contrast, each relation uniquely defines the semantic intent of a triplet and determines whether the connection between two entities is meaningful, which aligns precisely with what the contribution score measures. The KG evolution is performed after each feedback iteration, which consists of multiple batches of user queries. After aggregating triplet-level scores, we calculate the global mean 𝜇 and standard deviation 𝜎, which define refinement thresholds: triplets with scores greater than 𝜏ℎ𝑖𝑔ℎ = 𝜇 + 𝜎 are considered high quality, while those consistently below 𝜏𝑙𝑜𝑤 = 𝜇 − 𝜎 throughout iterations are considered low quality. These thresholds provide a data-driven basis,

+

Original KG Low-Scoring Triplets Across Iterations

𝒢 (ℎ𝑖 )

𝒢 (0)

𝒢 (ℎ𝑗 ) 70

70

71 150

25

140

130

71 Fusing Continuous High-score Paths

Relation Fusion

145

70

Hidden in Retrieval i

70 15

i+1 i+2 i+3

for each

(ℎ)

3 𝑇𝑠𝑡𝑎𝑟𝑡 ← {𝑡 |𝑡 ∈ G and 𝑆𝑐

(𝑡 ) ≥ 𝜇 + 𝜎 }; parallel for 𝑡𝑠𝑡𝑎𝑟𝑡 ∈ 𝑇𝑠𝑡𝑎𝑟𝑡 do 5 𝐹𝑟𝑜𝑛𝑡𝑖𝑒𝑟 ← [𝑡𝑠𝑡𝑎𝑟𝑡 ]; 6 for ℎ𝑜𝑝 = 1 → 𝐻 and 𝐹𝑟𝑜𝑛𝑡𝑖𝑒𝑟 ≠ [ ] do 7 𝑁 𝑒𝑥𝑡 𝐹𝑟𝑜𝑛𝑡𝑖𝑒𝑟 ← [ ]; 8 for each 𝑡𝑐𝑢𝑟𝑟 ∈ 𝐹𝑟𝑜𝑛𝑡𝑖𝑒𝑟 do (ℎ) 9 𝑁 ← NbrTriplet(𝑡𝑐𝑢𝑟𝑟 ); // sorted by 𝑆𝑐 to prioritize high-contribution relations 10 for each 𝑡𝑛𝑏𝑟 ∈ 𝑁 do 11 𝑆¯ ← Score(𝑡𝑠𝑡𝑎𝑟𝑡 , 𝑡𝑛𝑏𝑟 ); // avg. path score 12 if 𝑆¯ ≥ 𝜇 + 𝜎 and ∄(𝑡𝑠𝑡𝑎𝑟𝑡 .ℎ𝑒𝑎𝑑, 𝑟, 𝑡𝑛𝑏𝑟 .𝑡𝑎𝑖𝑙 ) ∈ 𝐾𝐺 then ¯ 13 𝑆ℎ𝑜𝑟𝑡𝑐𝑢𝑡 .push(((𝑡𝑠𝑡𝑎𝑟𝑡 .ℎ𝑒𝑎𝑑, 𝑟 ∗ , 𝑡𝑛𝑏𝑟 .𝑡𝑎𝑖𝑙 ), 𝑆)); 14 𝑁 𝑒𝑥𝑡 𝐹𝑟𝑜𝑛𝑡𝑖𝑒𝑟 .push(𝑡𝑛𝑏𝑟 );

Iteration

Relation Suppression

4

Figure 6: Illustration of relation-centric KG evolution, where triplets with low contribution scores are progressively downweighted in retrieval probability. ensuring that relation fusion and pruning are guided by statistically significant differences rather than random fluctuations. Guided by these thresholds, the KG is evolved through two operations: Relation Fusion and Relation Suppression. As illustrated in Figure 6, high-quality relations are reinforced by adding shortcut edges that connect the endpoints of multiple hops of high-quality paths, while persistently low-quality relations are suppressed via reduced contribution scores. Together, these operations balance adaptivity with stability, ensuring that the KG evolves in alignment with feedback while maintaining structural coherence. • Relation fusion. This operation strengthens high-quality relations by abstracting a shortcut edge that connects the endpoints of a multi-hop path, effectively reducing reasoning depth. For a path 𝐿𝑖 = (𝑡 1, 𝑡 2, . . . , 𝑡𝑘 ) whose triplets satisfy 𝑆𝑐 (𝑡𝑖 ) > 𝜏ℎ𝑖𝑔ℎ , EvoRAG abstracts a shortcut edge 𝑟ˆ: ˆ 𝑒𝑘 ), (𝑒 1, 𝑟 1, 𝑒 2 ), (𝑒 2, 𝑟𝑖 , 𝑒 3 ), . . . , (𝑒𝑘 −1, 𝑟𝑘 , 𝑒𝑘 ) ⇒ (𝑒 1, 𝑟,

(ℎ)

Input: IKGI+1 G (ℎ)I+2 at iteration ℎ, contribution scores 𝑆𝑐 I+3 triplet 𝑡 appearing in G, max hop 𝐻 Iteration Output: Updated KG G (ℎ+1) (ℎ) 1 Compute global mean 𝜇 and std 𝜎 of 𝑆𝑐 ; 2 Initialize 𝑆ℎ𝑜𝑟𝑡𝑐𝑢𝑡 = [ ];

15

160

145

75

18

150

40

160

60

40 Algorithm 251: Feedback-driven KG Evolution.

69

else

15

break;

16

𝐹𝑟𝑜𝑛𝑡𝑖𝑒𝑟 ← 𝑁 𝑒𝑥𝑡 𝐹𝑟𝑜𝑛𝑡𝑖𝑒𝑟 ;

17

(ℎ+1) ← (𝐺 (ℎ) ∪ 𝑆ℎ𝑜𝑟𝑡𝑐𝑢𝑡 ); 18 𝐺 19

return 𝑇 (ℎ+1)

It is worth noting that the current evolution module focuses on refining relations rather than adding or removing entities. Entity modification corresponds to factual correction, which requires external knowledge verification beyond the scope of feedback-based optimization. In contrast, our contribution score mechanism adjusts the structural importance of existing relations to improve reasoning behavior. Thus, entity-level updates and relation-centric evolution address orthogonal aspects, and future work will explore their integration for joint factual and behavioral adaptation.

(10)

where the 𝑟ˆ is assigned a label and a score: the label is recommended by the LLM based on the semantics of the multi-hop path, and the score is set to the path’s average contribution. • Relation Suppression. We progressively suppress low-quality triplets based on their long-term contribution patterns through two strategies. Low-contribution triplets are softly deprioritized during retrieval. As defined in Formula 2, retrieval is jointly guided by semantic similarity and contribution score, which reduces the retrieval probability of low-score triplets without immediately discarding them. Triplets that are semantically relevant to other queries may still be retrieved and can regain contribution once they prove useful again (Section 5.2). Algorithm 1 summarizes the KG evolution process. Given the KG G (ℎ) , we first compute the mean 𝜇 and standard deviation 𝜎 of triplet contribution scores 𝑆𝑐(ℎ) (𝑡) over all triplets 𝑡 ∈ G (ℎ) (line 1). Then, we start from triplets whose scores exceed the threshold (line 3), and perform parallel BFS triplets to identify candidate multi-hop paths (lines 4-17). A shortcut edge is created between the endpoints of a path if (i) the path’s average contribution score exceeds the threshold and (ii) no existing edge connects the endpoints in the KG (lines 12-14). To improve efficiency, neighbors are traversed in descending order of 𝑆𝑐(ℎ) (line 9), allowing early termination within each BFS layer when remaining neighbors fall below the threshold (line 16). This process ensures that the KG gradually refines its structure, strengthening consistently useful relations while suppressing noisy or unreliable ones.

5.2

Hybrid Priority-based Retrieval

EvoRAG introduces a hybrid priority-based retrieval module that jointly leverages semantic relevance and feedback-derived triplet contribution. Building on existing KG-RAG retrieval methods [6, 15, 21] that emphasize short-term semantic relevance, our approach incorporates long-term feedback to progressively adjust retrieval priorities according to user needs and knowledge reliability. When a query 𝑞 arrives, EvoRAG first encodes it into an embedding vector and retrieves the top-𝑁 relevant entities (𝐸𝑞 ) as starting nodes based on cosine similarity. The 𝑘-hop neighborhoods of these entities are traversed to construct a high-recall subgraph that preserves potentially useful multi-hop relations. This step ensures comprehensive coverage of potentially useful knowledge. Within the extracted subgraph, EvoRAG performs a hybrid ranking of candidate reasoning paths. Each triplet 𝑡 is associated with two complementary scores: (1) a relevance score 𝑆𝑟 (𝑡) that measures semantic similarity to 𝑞, and (2) a contribution score 𝑆𝑐 (𝑡) that captures its historical utility across iterations. These scores are 7

integrated into a unified path priority 𝑃 (𝐿𝑖 ) (Eq. 3), which adaptively balances short-term relevance with long-term reliability. For each starting entity 𝑒 ∈ 𝐸𝑞 , EvoRAG retains the top-𝑀 paths with the priorities, thereby suppressing those dominated by low-quality triplets and emphasizing reusable, high-confidence knowledge. Importantly, EvoRAG remains compatible with various retrieval backends. While 𝑆𝑟 (𝑡) currently follows a similarity-based formulation [6, 15, 21], it can be replaced with other strategies such as ToG [70] or DALK [37]. The key innovation lies in the integration of 𝑆𝑐 (𝑡), which injects feedback-driven triplet quality into the retrieval process, forming a flexible and adaptive hybrid retrieval paradigm. Empirically, setting 𝑁 = 10 and 𝑀 = 10 yields the best performance, achieving high response accuracy. Larger values bring marginal gains while increasing retrieval cost (see Section 6.9). This hybrid retrieval design improves reasoning accuracy by filtering low-contribution triplets and reducing unnecessary context, resulting in more efficient and focused LLM input.

Table 1: Dataset description and the corresponding sizes of constructed KGs.

6 EXPERIMENTAL EVALUATION 6.1 Experimental Setup

Baseline methods. EvoRAG improves the generation quality of KG-RAG frameworks by dynamically refining the knowledge graph based on feedback. To evaluate effectiveness, we compare it with two categories of baselines.

Dataset RGB [9] Multihop (MTH) [72] HotpotQA (HPQ) [85]

Query 300 816 600

Type Single-hop Multi-hop Multi-hop

Entity 54,544 30,953 76,280

Triplets 74,394 26,876 74,942

Table 3 shows that the generated feedback achieves an agreement ratio of 93.38% with the ground truth, indicating that it serves as a reliable proxy for evaluation. To demonstrate the robustness and applicability of EvoRAG, we also evaluate it under diverse feedback settings. We first introduce partially noisy feedback to assess its tolerance to unreliable signals (Section 6.5). We then consider alternative feedback sources, including human expert judgments to reflect practical scenarios without explicit ground truth, and ground-truth-based feedback (e.g., F1 score) as an ideal reference (Section 6.4). The results show that EvoRAG maintains stable effectiveness under various feedback reliability and availability settings.

Environments. The experiment is conducted on a GPU server equipped with 2 Intel(R) Xeon(R) Silver 4316 CPUs, 503GB DRAM, and 2× NVIDIA RTX A6000 (48 GB) GPUs. The server runs Ubuntu 20.04 OS (Linux kernel 5.15.0) with GCC-9.4.0, CUDA 11.3 with driver version 560.35, and PyTorch 1.13.0 backend.

• KG-RAG framework. We compare EvoRAG with three representative KG-RAG baselines, including Microsoft GraphRAG (MRAG) [15], LightRAG (LRAG) [21], and KRAG. MRAG and LRAG are two variants that integrate textual chunks with KGs, providing fine-grained knowledge. MRAG applies community detection algorithms to summarize subgraph information, providing abstract representations. LRAG adopts a dual-level retrieval to combine entity-level and relation-level information. In addition, we adopt KRAG, the most effective KG-RAG instance identified in Lego-GraphRAG [6] (see Section 2.2). • KG-RAG with KGR methods. Since EvoRAG refines the underlying KG, we compare it with KG-RAG frameworks augmented by KG refinement (KGR) methods. We integrate several representative KGR methods into KRAG, including deep learning-based methods such as TransE [3], RotatE [71], and CAGED [99], and an LLM-based method, LLM_sim [13]. These methods aim to identify and correct noisy or irrelevant triplets.

Datasets. We evaluate the effectiveness of EvoRAG on three realworld datasets, as summarized in Table 1. RGB [9] is constructed from news articles to evaluate reasoning capabilities of RAG tasks. MultiHop (MTH) [72] focuses on multi-hop queries that require integrating information from multiple news documents. HotpotQA (HPQ) [85] is a multi-hop QA dataset derived from Wikipedia, where answering each query requires combining evidence from two paragraphs. We use all 300 English-language reasoning queries from the RGB dataset and all 816 reasoning queries from the MTH dataset. For HPQ, we follow prior work [27] and randomly select 600 queries for evaluation due to the high cost of processing the full dataset. Training and test set construction. To evaluate EvoRAG under realistic online-service conditions, we simulate a multi-user scenario where numerous queries concentrate on the same local regions of the knowledge graph. Based on this setting, we construct the training and test sets as follows. The test set directly uses the original queries listed in Table 1, representing user queries in these hotspot regions. The training set is constructed by generating additional query–answer pairs within the subgraphs retrieved based on the test queries. For each subgraph, we sample alternative reasoning paths (up to two hops) that differ from the paths used to form test queries, avoiding data leakage. The corresponding target entities are used as ground-truth answers, producing diverse yet localized queries that simulate multi-user access to the same hotspot regions. During training, the generated queries are issued over multiple iterations to simulate real-time user interactions, allowing the system to collect feedback and update the KG incrementally. The model’s performance is then evaluated on the test set.

KG construction. Since the datasets contain only raw texts, we construct KGs following prior work [15, 21]. Texts are split into 512-token chunks, from which GPT-4o-mini extracts entities and relations using predefined prompts. Table 1 reports KG statistics. MRAG and LRAG further build higher-level textual representations, while KRAG and EvoRAG operate directly on triplet-based KGs. Implementation details. For each query, KRAG extracts 𝑁 = 10 query entities and collects their 2-hop neighbors to form a candidate subgraph. From this subgraph, up to 𝑀 = 10 reasoning paths per query entity are selected. EvoRAG builds upon KRAG by introducing a feedback-driven backpropagation mechanism that learns a contribution score for each triplet (with a learning rate of 0.5) during online operation to guide the retrieval over the progressively refined KG. For KGR methods, TransE and RotatE identify noisy triplets using the constraint ||(ℎ + 𝑟 − 𝑡)|| < 𝛾, where ℎ and 𝑡 denote the embeddings of head and tail entities, and 𝑟 is the relation embedding. We set 𝛾 = 0.1 following [13]. To ensure fair

Feedback generation. By default, we use LLM-generated feedback produced by Qwen2.5-32B (Section 4.1) in our main experiments. 8

Table 2: Performance comparison with KG-RAG and KGR methods. MRAG, LRAG, and KRAG are KG-RAG baselines. TransE, RotatE, CAGED, and LLM_Sim are KGR methods applied to refine the KG, whose effects are evaluated using KRAG due to its operation at the triplet-level retrieval granularity. Best results are highlighted in bold, and worst results in red. ↑ indicates the accuracy improvement over the worst result. Existing KG-RAG frameworks KRAG with various KGR methods Dataset Metric EvoRAG MRAG LRAG KRAG TransE RotatE CAGED LLM_Sim #ACC 75.67% 76.00% 71.00% 67.67% 74.33% 73.33% 67.33% 84.00% ↑8.00-16.67 RGB #EM 47.33% 47.00% 42.33% 42.67% 47.33% 45.67% 43.33% 56.67% ↑9.34-14.34 #F1 68.69% 64.99% 64.21% 60.61% 68.47% 67.02% 62.95% 75.40% ↑6.71-14.79 #ACC 75.61% 76.20% 74.02% 50.12% 76.47% 76.72% 72.06% 80.26% ↑3.54-30.14 MTH #EM 69.61% 70.43% 71.94% 50.37% 72.55% 74.02% 71.45% 78.55% ↑6.00-28.18 #F1 75.86% 74.87% 74.87% 52.23% 77.08% 77.83% 73.27% 80.80% ↑2.97-28.57 #ACC 38.83% 44.83% 39.00% 27.67% 36.33% 29.17% 32.83% 48.16% ↑3.33-20.49 HPQ #EM 25.50% 25.67% 24.83% 17.83% 24.50% 20.17% 21.50% 35.84% ↑10.17-18.01 #F1 37.36% 40.47% 41.32% 29.34% 38.51% 32.21% 36.07% 46.55% ↑6.08-17.21

Total Feedback 300 816 600

Correct Feedback 288 783 529

Proportion 96% 95.96% 88.17%

comparison, we adopt Qwen2.5-32B as the unified LLM backbone across all methods. It is used for path-level evaluation in EvoRAG and response generation in EvoRAG and all baselines.

6.2

Problematic Triplets Ratio (%)

Accuracy (%)

40

90 20

60

30

85 15

55

20

80 10

50

10

75

5

45

70

0

0

0

1

2

3 4 5 Iteration

6

(a) RGB dataset

7

8

Accuracy (%)

Dataset RGB MTH HPQ

Problematic Triplets Ratio (%)

Table 3: Alignment ratio between generated feedback and ground truth.

0

1

2

3 4 5 Iteration

6

7

8

40

(b) HPQ dataset

Figure 7: The problematic triplets ratio and response accuracy across iterations. Each iteration involves answering all queries and propagating corresponding feedback.

Overall Comparison

We compare EvoRAG with various KG-RAG frameworks and with KRAG enhanced by different KGR methods, using accuracy (ACC), exact match (EM), and F1 as evaluation metrics. ACC measures the proportion of queries for which the response contains the groundtruth answer. EM indicates whether the response exactly matches the ground-truth answer. F1 evaluates the token-level overlap by computing the harmonic mean of accuracy and recall. EvoRAG is trained on the training set, enabling it to adaptively refine the KG. Table 2 reports the experimental results. Compared to KG-RAG frameworks, EvoRAG achieves an average improvement of 7.34% in ACC, 9.84% in EM, and 7.29% in F1 score across three datasets. These gains primarily stem from the ability of EvoRAG to adapt retrieval to downstream query and dynamically correct noisy or incomplete knowledge. Existing KG-RAG frameworks rely on static KGs, where irrelevant or noisy knowledge may suppress useful reasoning paths. This prevents them from retrieving high-quality context and ultimately limits performance. In contrast, EvoRAG leverages a feedback-driven backpropagation mechanism that propagates response-level feedback to individual knowledge triplets. This mechanism addresses cases where baselines miss correct answers. EvoRAG reweights suppressed but useful paths upward and diminishes the influence of misleading ones. Accumulated across queries, these adjustments enable the system to recover answers systematically overlooked by static methods. In addition, the KG evolution mechanism in EvoRAG enhances the KG’s adaptability to RAG tasks, allowing it to evolve continuously during online service by suppressing low-quality triplets and generating new relations. Compared to KGR-based KGRAG methods, EvoRAG achieves an average improvement of 13.80% in ACC, 12.74% in EM, and 11.28% in

F1 score across three datasets. Traditional DL-based KGR methods, such as TransE, RotatE, and CAGED, rely on supervised learning to detect noisy triplets. These methods require fine-grained labels indicating the correctness of individual triplets, which are typically unavailable in KG-RAG scenarios. LLM_sim attempts to overcome this limitation by leveraging LLMs to evaluate and refine triplets based on internal knowledge. However, it suffers from hallucinations, especially when encountering knowledge beyond the model’s training data, and may introduce new noise or mistakenly modify correct facts. Therefore, traditional KGR methods are not suitable for the KG-RAG framework and may even significantly degrade accuracy. In contrast, EvoRAG directly leverages the real-time feedback, which provides more direct and task-relevant supervision. Although this feedback may occasionally be noisy and coarse-grained, EvoRAG leverages it through a feedback-driven backpropagation mechanism that aggregates feedback across queries and propagation the feedback into a reliable basis for knowledge refinement. In this way, EvoRAG continually adapts the knowledge graph to actual usage patterns, enabling sustained improvements in reasoning performance over time.

6.3

The Improvement of KG Quality

KG evolution is performed once per iteration, where each iteration processes a set of training queries (equal in size to the test set) in batches, with contribution scores updated after each batch. To evaluate its impact on retrieval quality, we track training accuracy and the number of problematic triplets retrieved per query. Problematic 9

100

80 60 40

RGB

MTH Dataset

HPQ

baseline

10% EF

20% EF 100

80 60 40

RGB

MTH Dataset

LLM-only EvoRAG

75 50 25 0

HPQ

Feedback without Graph Evolution Static Graphs without Feedback

RGB

MTH Dataset

EvoRAG w/o CV

100 Accuracy (%)

HF

Accuracy (%)

Accuracy (%)

GF

Accuracy (%)

LF

100

80 60 40

HPQ

EvoRAG w/ CV

RGB

MTH Dataset

HPQ

Figure 8: Accuracy across feed- Figure 9: Noise tolerance anal- Figure 10: Ablation study of Figure 11: Ablation study on back sources. LF, GF, and HF ysis. 10% and 20% erroneous feedback mechanism. the use of Fidelity and Conflict denote LLM, ground-truth (F1), feedback (EF) is injected durfor cross-validation (CV). and human feedback. ing each iteration. MRAG

6.4

10.0

4

7.5 5.0 2.5 0.0

EvoRAG (ours)

3 2 1

RGB

MTH Dataset

HPQ

0

RGB

MTH Dataset

HPQ

(b) TTFT

Figure 12: Comparison of EvoRAG with KG-RAG frameworks in terms of prompt length and time-to-first-token (TTFT).

LLM-only, LLM with static KG (SG), LLM+SG with feedback (FB), and LLM+SG+FB with graph evolution (i.e., EvoRAG). As shown in Figure 10, SG outperforms the LLM-only by 23.42% in accuracy, demonstrating KG-RAG’s ability to supplement missing knowledge. FB further improves accuracy by 6.6% by prioritizing reliable triplets via feedback, while EvoRAG adds 3.07% by updating the KG to prune low-utility triplets and reinforce correct long-range reasoning. These results demonstrate the effectiveness of the feedback mechanism in improving reasoning accuracy.

Impact of Different Feedback Sources

6.6.2 Ablation Study of Cross-validation Mechanism. We evaluate path utility updates with and without Fidelity and Conflict. As shown in Figure 11, incorporating these cross-validation constraints improves the average accuracy by 2.24% over the variant without cross-validation, confirming their effectiveness in filtering unreliable or contradictory paths.

Feedback Tolerance Analysis

To evaluate the impact of noisy feedback on EvoRAG, we simulate erroneous feedback by randomly selecting a fraction of the collected feedback and inverting the correctness. In the noise injection, we randomly select 10% and 20% of the feedback scores excluding neutral scores of 3, and flip them: scores 1 − 2 are converted to 4 − 5, and scores 4 − 5 are converted to 1 − 2. This process simulates imperfect or noisy feedback, reflecting realistic scenarios where evaluation judgments may be inaccurate or inconsistent. The experimental results are reported in Figure 9. We observe that EvoRAG remains tolerant under noisy feedback. Even with 10% and 20% erroneous feedback, accuracy drops only slightly, by 1.15% and 2.43%, respectively. This resilience stems from the cumulative update of contribution scores across multiple queries and reasoning paths. Since each triplet is evaluated from diverse contexts, occasional erroneous feedback has only a transient impact, while consistent signals dominate over time, enabling accurate convergence despite noise.

6.6

KRAG 5

(a) Token length

To evaluate the applicability of EvoRAG, we compare its accuracy under different feedback sources. Specifically, we consider LLMgenerated feedback, ground-truth based feedback (using F1 score), and human feedback, where users assign satisfaction scores to each response. As shown in Figure 8, EvoRAG achieves comparable performance across all settings, with an average difference of 0.75%. This suggests that the performance gains mainly stem from the availability of feedback signals rather than their specific source.

6.5

LRAG

TTFT (s)

Token Length (K tokens)

triplets defined in Section 2.2 as irrelevant, outdated, incorrect, or long-path connections, are manually annotated. Figure 7 shows that as iterations increase, the ratio of problematic triplets steadily decreases while accuracy improves and stabilizes after approximately 6 iterations across all datasets. This indicates that EvoRAG can effectively incorporate useful feedback within a few iterations. On RGB, EvoRAG suppresses around 83.01% of problematic triplets, achieving larger gains than HPQ due to the dataset’s high level of noisy facts. Performance improvements are most significant in early iterations, where EvoRAG can rapidly identify and down-weight misleading triplets. As iterations progress, the performance gradually converges as the system stabilizes.

6.7

Token Cost Comparison

EvoRAG reduces prompt length by down-weighting noisy, outdated, or irrelevant triplets during retrieval, thereby lowering token cost and improving system efficiency, as token consumption directly determines inference latency and resource overhead. To quantify these benefits, we evaluate prompt token cost and time-to-firsttoken (TTFT), which captures the latency of LLM responses. As shown in Figure 12, compared to MRAG, LRAG, and KRAG, EvoRAG achieves an average reduction of 4.6× in prompt length, resulting in consistently lower TTFT. MRAG and LRAG often retrieve not only reasoning paths but also large amounts of associated text chunks, significantly increasing prompt length and prefill latency. KRAG retrieves raw triplets but lacks mechanisms to suppress semantically redundant or low-quality triplets. In contrast, EvoRAG achieves higher accuracy with shorter prompt length, demonstrating that it retains essential knowledge while successfully filtering out redundant or outdated information.

Ablation Study

6.6.1 Ablation study of Feedback Mechanism. To evaluate the contribution of the feedback mechanism, we compare four settings: 10

KRAG

Backward

40

70

10

4 2

5 10 20 Batch Size

(a) RGB dataset

40

1

2

5 10 20 Batch Size

40

80

Runtime Performance Analysis

30

6 8 10 12 14 16 (a) Varying M on RGB

4

6 8 10 12 14 16 (b) Varying M on HPQ

Figure 14: Impact of hyperparameters on accuracy. 𝑁 indicates the number of retrieved entities; 𝑀 indicates the number of retrieval paths per entity.

The backpropagation mechanism in EvoRAG introduces additional overhead compared to traditional KG-RAG frameworks. To mitigate this overhead, we exploit a key observation: the prompts used in forward and backward propagation share a common prefix, including the same query and retrieved reasoning paths, and differ only in the task-specific suffix. Building on this insight, we adopt a KV-cache reuse strategy in vLLM [33], enabling the model to reuse cached computations for the shared prefix, without compromising effectiveness. We evaluate the runtime of EvoRAG under varying batch sizes, measuring both forward and backward propagation. As shown in Figure 13, KV-cache reuse effectively reduces the overhead of backpropagation, with its relative cost decreasing as batch size increases. At a batch size of 20, backpropagation accounts for only 23.90% of the total runtime. This efficiency gain arises from improved utilization under larger batches, where shared-prefix reuse and system-level scheduling amortize the cost of repeated computations. Overall, these results demonstrate that the overhead introduced by EvoRAG is well-controlled, making it suitable for deployment in real-world online scenarios.

60

Accuracy (%)

Accuracy (%)

60

50

40

30

100 Forward Backward

80

55 50 45 40

MRAG LRAG KRAG EvoRAG

KG-RAG Frameworks (a) Performance comparison

60 40 20

0%

10%

20%

0

1

Error Feedback Ratio (b) Noise tolerance analysis

60

5

10

15

20

Batches (c) Runtime analysis

25

60 EvoRAG KRAG

Accuracy (%)

55

EvoRAG KRAG

55 Figure 15: Evaluation on 5,000 queries with large KG (890,389 50 entities, 1,043,360 triplets). (a)50 Accuracy comparison. (b) 45 45 Noise tolerance analysis. (c) Performance analysis. 40

4

6

8

10

12

(d) Varying N

14

16

40

4

6

8

10

12

(e) Varying M

14

16

Table 4: Proportions of added and removed triplets in retrieval results after KG evolution. Proportion of Triplets Added Removed

Sensitivity Analysis of Entity and Path Numbers

RGB 24.41% 28.97%

MTP 21.06% 38.23%

HPQ 13.46% 17.31%

LRAG, and KRAG by 8.25%, demonstrating that the feedback mechanism remains effective at scale. Figure 15(b) shows that each 10% increase in erroneous feedback reduces accuracy by less than 1%, indicating strong noise tolerance due to cumulative aggregation across queries. Figure 15(c) shows that while the larger KG increases overall runtime (mainly in forward propagation due to retrieval overhead), backpropagation remains lightweight, accounting for only 15.35% of total runtime at batch size 20.

For KG-RAG frameworks, the number of entities (𝑁 ) and reasoning paths per entity (𝑀) directly impact the amount of knowledge retrieved and overall accuracy. We analyze the sensitivity of them by varying one parameter at a time while keeping the other fixed at 10. The experimental results are shown in Figure 14. We observe that increasing either 𝑁 or 𝑀 generally improves accuracy, as more retrieved knowledge provides a richer context for generation. However, the gains diminish when 𝑁 and 𝑀 reach 10, where additional paths contribute marginally to performance, indicating that a limited number of paths are sufficient to enhance generation. Notably, accuracy is more influenced by the number of entities than by the number of paths. This is because adding entities introduces more diverse and potentially relevant information.

6.10

6 8 10 12 14 16 (d) Varying N on HPQ

40

70 4

6.9

4

50

(b) MTH dataset

Figure 13: Comparison of forward and backward propagation time of EvoRAG under different batch sizes.

6.8

30

6 8 10 12 14 16 (c) Varying N on RGB

Accuracy (%)

1

0

Accuracy (%)

0

Time (s)

20

20

Accuracy (%)

40

80

30

EvoRAG (ours)

50

Accuracy (%)

60

Batch Time (s)

Batch Time (s)

40

Accuracy (%)

Forward

6.11

Effectiveness Analysis of Feedback Mechanism

In this section, we analyze how feedback-driven backpropagation affects the KG and retrieval results, and present a case study on representative queries.

Scalability on Large Dataset

6.11.1 Changes in Retrieval Results. We analyze the impact of KG evolution on retrieval by comparing the retrieved triplets before and after evolution. As shown in Figure 4, 28.17% of low-contribution triplets are removed, while 19.64% of high-contribution triplets are newly introduced. This shows that KG evolution shifts retrieval

We evaluate EvoRAG in a large-scale setting with 5,000 queries sampled from HotpotQA on a KG containing 890,389 entities and 1,043,360 triplets. Figure 15 summarizes the results. As shown in Figure 15(a), EvoRAG improves the average accuracy over MRAG, 11

Clark Gable

Query: Which American pre-Code drama released in 1932 starred the actor known as “the King of Hollywood”?

MRAG

LRAG

KRAG 100

100

Red Dust types as American pre-Code Dram Film  Clark Gable starred in (1932)  The King of Hollywood refers to Clark Gable starred in (1932) Red Dust  Clark Gable starred in (1939) Gone with the Wind types as American pre-Code Dram Film Gone with the Wind  The King of Hollywood known as Clark Gable starred in

Wrong!



Retrieval Results with Feedback:

 Clark Gable performed in(1932) Strange Interlude types as American pre-Code Dram Film  The King of Hollywood refers to Clark Gable starred in (1932) Red Dust  Clark Gable starred in (1939) Gone with the Wind types as American pre-Code Dram Film Gone with the Wind  The King of Hollywood known as Clark Gable starred in

EvoRAG (ours)

LLM Red Dust

Strange Interlude

80 60

The king of Hollywood  Clark gable 40 → Appeared between 1924 and 1926 The king of Hollywood  Clark Gable → Starring in Red Dust in 1932 American pre-Code Dram Film  Red Dust  Clark Gable

Correct!



Accuracy (%)

Wrong Knowledge

Original Retrieval Results:

Accuracy (%)

Refers to

RGB

MTH Dataset

HPQ

(a) LLaMA-3.1-70B (4-bit)

80 60 40

RGB

MTH Dataset

HPQ

(b) GPT-4o-mini

Figure 18: Accuracy comparison of different KG-RAG frameworks across different LLM backends.

Figure 16: Comparison of retrieved paths of a query over HPQ before and after 8 feedback iterations.

King of Hollywood  Clark Gable →  The Strange Interlude in 1932

w/o FB

90 80 70 60

60

Accuracy (%)

Accuracy (%)

100

MRAG

LRAG

(a) RGB dataset

pre-Code Dram Film   American Strange Interlude  Clark Gable

indicate that the performance gains of EvoRAG are robust to the choice of LLM backend.

w/ FB

king of Hollywood  Clark Gable  The → Starring in Red Dust in 1932

King of Hollywood  Clark Gable →  The Gone with the Wind

50

7

40

30

RELATED WORK

American Dram Film  Gone with the  Wind  Clark Gable

MRAG

KG-RAG framework. KG-RAG frameworks can be broadly grouped into two categories. One line of work represents knowledge as textual chunks and uses graph structures mainly for indexing interchunk relations [15, 21, 22, 43, 46, 49, 62]. While effective for summarization, these methods are less suitable for complex reasoning. Another line explicitly constructs and leverages KGs to support multi-hop reasoning over structured triplets [7, 10, 11, 24, 30, 31, 37, 39, 40, 42, 44, 51, 61, 70, 98]. However, most of these approaches emphasize retrieval and prompt design, without fully exploiting the KG’s capacity for complex reasoning, which is critical for accurate responses [104].

LRAG

(b) HPQ dataset

Figure 17: Accuracy comparison of MRAG and LRAG on two datasets with Feedback-driven Backpropagation (FB).

away from repeatedly selecting unhelpful relations and toward incorporating triplets that have proven useful. 6.11.2 Case study. Figure 16 shows a case study for a query from the HPQ dataset. In the original retrieval, path contains a clear factual error, but is ranked higher due to its high semantic similarity to the query, causing the correct path to Strange Interlude to be ranked lower. After feedback continuously indicates an incorrect answer in 8 iterations, EvoRAG suppresses this noisy triplet by reducing its contribution score and reinforces correct triplets. As a result, the correct path becomes more salient in retrieval, leading to a correct answer in subsequent queries.

KG refinement. KG refinement (KGR) focuses on enhancing the factual accuracy and utility of KGs by removing redundant triplets, correcting incorrect facts, and adding missing information. Rulebased methods [18, 25, 58, 87] rely on logical constraints but scale poorly, while DL-based approaches [3, 28, 48, 55, 68, 71, 81] typically operate offline and require high-quality training data. Recently, LLM-based methods [13, 27, 79] assess semantic plausibility or generate facts, but remain detached from downstream tasks.

6.12

KG validation. KG validation (KGV) [2, 69, 76] assesses triplet correctness by relying on external evidence and expert-curated annotations, which often require task-specific pipelines and substantial expert involvement, thereby limiting scalability.

The Effectiveness of Feedback-driven Backpropagation in MRAG and LRAG

Since MRAG and LRAG retrieve text chunks and use graphs only as indices, we adapt our framework to operate at the chunk level to support feedback-driven backpropagation. Specifically, we define contribution scores for text chunks and incorporate them into retrieval ranking alongside semantic similarity. As shown in Figure 17, the integration improves the average accuracy of MRAG and LRAG by 2.83% and 1.92%, respectively. The gains are smaller than those achieved by the triplet-level framework, as each text chunk aggregates multiple facts and thus provides coarser-grained feedback signals.

6.13

Feedback-driven model optimization. These approaches [57, 88, 96] optimize model outputs in a feedback-driven manner via fine-tuning or reinforcement learning, yet operate solely at the model level without updating the KG.

8

CONCLUSION

We propose EvoRAG, a self-evolving KG-RAG framework that leverages real-time feedback to continuously refine the KG and improve reasoning accuracy. EvoRAG introduces a feedback-driven backpropagation mechanism that connects response-level feedback and triplet-level knowledge updates, which attributes feedback to individual reasoning paths and propagates it to adjust the scores of involved triplets. This establishes a closed loop, where reasoning feedback drives KG refinement, and the evolving KG improves future reasoning. The experimental results demonstrate that EvoRAG is both effective and robust, offering a scalable solution for adaptive and scalable KG maintenance in real-world applications.

Performance with Various LLM Backends

We evaluate the generality of EvoRAG across different LLM backends by replacing Qwen2.5-32B with Llama-3.1-70B (4-bit) and GPT-4o-mini for all LLM-dependent components in EvoRAG and the baselines. As shown in Figure 18, EvoRAG consistently outperforms existing RAG methods, improving average accuracy by 8.5% with Llama-3.1-70B and by 7.6% with GPT-4o-mini. These results 12

REFERENCES

[23] Peixuan Han, Adit Krishnan, Gerald Friedland, Jiaxuan You, and Chris Kong. 2025. Self-Aligned Reward: Towards Effective and Efficient Reasoners. arXiv preprint arXiv:2509.05489 (2025). [24] Xiaoxin He, Yijun Tian, Yifei Sun, Nitesh Chawla, Thomas Laurent, Yann LeCun, Xavier Bresson, and Bryan Hooi. 2024. G-retriever: Retrieval-augmented generation for textual graph understanding and question answering. Advances in Neural Information Processing Systems 37 (2024), 132876–132907. [25] Yan Hong, Chenyang Bu, and Xindong Wu. 2021. High-quality noise detection for knowledge graph embedding with rule-based triple confidence. In PRICAI 2021: Trends in Artificial Intelligence: 18th Pacific Rim International Conference on Artificial Intelligence, PRICAI 2021, Hanoi, Vietnam, November 8–12, 2021, Proceedings, Part I 18. 572–585. [26] Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232 (2023). [27] Manzong Huang, Chenyang Bu, Yi He, and Xindong Wu. 2025. How to Mitigate Information Loss in Knowledge Graphs for GraphRAG: Leveraging Triple Context Restoration and Query-Driven Feedback. arXiv preprint arXiv:2501.15378 (2025). [28] Shaoxiong Ji, Shirui Pan, Erik Cambria, Pekka Marttinen, and Philip S Yu. 2021. A survey on knowledge graphs: Representation, acquisition, and applications. IEEE transactions on neural networks and learning systems 33, 2 (2021), 494–514. [29] Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. Comput. Surveys 55, 12 (2023), 1–38. [30] Bowen Jin, Chulin Xie, Jiawei Zhang, Kashob Kumar Roy, Yu Zhang, Zheng Li, Ruirui Li, Xianfeng Tang, Suhang Wang, Yu Meng, et al. 2024. Graph Chain-ofThought: Augmenting Large Language Models by Reasoning on Graphs. (2024), 163–184. [31] Mingyu Jin, Haochen Xue, Zhenting Wang, Boming Kang, Ruosong Ye, Kaixiong Zhou, Mengnan Du, and Yongfeng Zhang. 2024. ProLLM: protein chain-ofthoughts enhanced LLM for protein-protein interaction prediction. bioRxiv (2024), 2024–04. [32] Pei Ke, Bosi Wen, Andrew Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, et al. 2024. Critiquellm: Towards an informative critique generation model for evaluation of large language model generation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 13034–13054. [33] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles. 611–626. [34] Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Ren Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, et al. 2024. RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback. In International Conference on Machine Learning. 26874–26901. [35] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems 33 (2020), 9459– 9474. [36] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems. [37] Dawei Li, Shu Yang, Zhen Tan, Jae Baik, Sukwon Yun, Joseph Lee, Aaron Chacko, Bojian Hou, Duy Duong-Tran, Ying Ding, et al. 2024. DALK: Dynamic Co-Augmentation of LLMs and KG to answer Alzheimer’s Disease Questions with Scientific Literature. , 2187–2205 pages. [38] Peizheng Li, Chaoyi Chen, Hao Yuan, Zhenbo Fu, Hang Shen, Xinbo Yang, Qiange Wang, Xin Ai, Yanfeng Zhang, Yingyou Wen, and Ge. Yu. 2025. NeutronRAG: Towards Understanding the Effectiveness of RAG from a Data Retrieval Perspective. Companion of the 2025 International Conference on Management of Data (SIGMOD-Companion ’25) (2025). [39] Shilong Li, Yancheng He, Hangyu Guo, Xingyuan Bu, Ge Bai, Jie Liu, Jiaheng Liu, Xingwei Qu, Yangguang Li, Wanli Ouyang, et al. 2024. GraphReader: Building Graph-based Agent to Enhance Long-Context Abilities of Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2024. 12758–12786. [40] Zhuoqun Li, Xuanang Chen, Haiyang Yu, Hongyu Lin, Yaojie Lu, Qiaoyu Tang, Fei Huang, Xianpei Han, Le Sun, and Yongbin Li. 2024. Structrag: Boosting knowledge intensive reasoning of llms via inference-time hybrid information structurization. arXiv preprint arXiv:2410.08815. [41] Zongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang, Dong Chen, Shuai Wang, and Cuiyun Gao. 2023. Cctest: Testing and repairing code completion

[1] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023). [2] Tyler Bikaun, Michael Stewart, and Wei Liu. 2024. CleanGraph: Humanin-the-loop Knowledge Graph Refinement and Completion. arXiv preprint arXiv:2405.03932 (2024). [3] Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko. 2013. Translating embeddings for modeling multi-relational data. Advances in neural information processing systems 26 (2013). [4] Léon Bottou, Frank E Curtis, and Jorge Nocedal. 2018. Optimization methods for large-scale machine learning. SIAM review 60, 2 (2018), 223–311. [5] Stephen P Boyd and Lieven Vandenberghe. 2004. Convex optimization. Cambridge university press. [6] Yukun Cao, Zengyi Gao, Zhiyang Li, Xike Xie, S. Kevin Zhou, and Jianliang Xu. 2025. LEGO-GraphRAG: Modularizing Graph-Based Retrieval-Augmented Generation for Design Space Exploration. Proceedings of the VLDB Endowment 18, 10 (2025), 3269–3283. [7] Boyu Chen, Zirui Guo, Zidan Yang, Yuluo Chen, Junze Chen, Zhenghao Liu, Chuan Shi, and Cheng Yang. 2025. PathRAG: Pruning Graph-based Retrieval Augmented Generation with Relational Paths. arXiv preprint arXiv:2502.14902 (2025). [8] Chaoyi Chen, Dechao Gao, Yanfeng Zhang, Qiange Wang, Zhenbo Fu, Xuecang Zhang, Junhua Zhu, Yu Gu, and Ge Yu. 2023. NeutronStream: A Dynamic GNN Training Framework with Sliding Window for Graph Streams. Proceedings of the VLDB Endowment 17, 3 (2023), 455–468. [9] Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking large language models in retrieval-augmented generation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 17754–17762. [10] Zhongwu Chen, Chengjin Xu, Dingmin Wang, Zhen Huang, Yong Dou, Xuhui Jiang, and Jian Guo. 2024. Rulerag: Rule-guided retrieval-augmented generation with language models for question answering. arXiv preprint arXiv:2410.22353 (2024). [11] Kewei Cheng, Nesreen K Ahmed, Theodore Willke, and Yizhou Sun. 2024. Structure guided prompt: Instructing large language model in multi-step reasoning by exploring graph structure of the text. arXiv preprint arXiv:2402.13415 (2024). [12] Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf, Dominic Culver, Rui Melo, Caio Corro, Andre FT Martins, Fabrizio Esposito, Vera Lúcia Raposo, Sofia Morgado, et al. 2024. Saullm-7b: A pioneering large language model for law. arXiv preprint arXiv:2403.03883 (2024). [13] Na Dong, Natthawut Kertkeidkachorn, Xin Liu, and Kiyoaki Shirai. 2025. Refining Noisy Knowledge Graph with Large Language Models. In Proceedings of the Workshop on Generative AI and Knowledge Graphs (GenAIK). 78–86. [14] Yuxin Dong, Shuo Wang, Hongye Zheng, Jiajing Chen, Zhenhong Zhang, and Chihang Wang. 2024. Advanced RAG Models with Graph Structures: Optimizing Complex Knowledge Reasoning and Text Generation. In 2024 5th International Symposium on Computer Engineering and Intelligent Communications (ISCEIC). 626–630. [15] Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2024. From local to global: A graph rag approach to queryfocused summarization. arXiv preprint arXiv:2404.16130 (2024). [16] Zhenbo Fu, Xin Ai, Qiange Wang, Yanfeng Zhang, Shizhan Lu, Chaoyi Chen, Chunyu Cao, Hao Yuan, Zhewei Wei, Yu Gu, et al. 2025. NeutronTask: Scalable and efficient multi-GPU GNN training with task parallelism. Proceedings of the VLDB Endowment 18, 6 (2025), 1705–1719. [17] Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 2 (2023). [18] Siddhant Garg, Goutham Ramakrishnan, and Varun Thumbe. 2021. Towards robustness to label noise in text classification via noise modeling. In Proceedings of the 30th ACM international conference on information & knowledge management. 3024–3028. [19] Yingqiang Ge, Wenyue Hua, Kai Mei, Juntao Tan, Shuyuan Xu, Zelong Li, Yongfeng Zhang, et al. 2023. Openagi: When llm meets domain experts. Advances in Neural Information Processing Systems 36 (2023), 5539–5568. [20] Daya Guo, Canwen Xu, Nan Duan, Jian Yin, and Julian McAuley. 2023. Longcoder: A long-range pre-trained language model for code completion. In Proceedings of International Conference on Machine Learning. 12098–12107. [21] Zirui Guo, Lianghao Xia, Yanhua Yu, Tu Ao, and Chao Huang. 2024. LightRAG: Simple and Fast Retrieval-Augmented Generation. arXiv preprint arXiv:2410.05779 (2024). [22] Bernal Jiménez Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. 2024. Hipporag: Neurobiologically inspired long-term memory for large language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. 13

systems. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). 1238–1250. [42] Lei Liang, Mengshu Sun, Zhengke Gui, Zhongshu Zhu, Zhouyu Jiang, Ling Zhong, Yuan Qu, Peilong Zhao, Zhongpu Bo, Jin Yang, et al. 2024. KAG: Boosting LLMs in Professional Domains via Knowledge Augmented Generation. arXiv preprint arXiv:2409.13731 (2024). [43] Xun Liang, Simin Niu, Sensen Zhang, Shichao Song, Hanyu Wang, Jiawei Yang, Feiyu Xiong, Bo Tang, Chenyang Xi, et al. 2024. Empowering large language models to set up a knowledge retrieval indexer via self-learning. arXiv preprint arXiv:2405.16933 (2024). [44] Haochen Liu, Song Wang, Yaochen Zhu, Yushun Dong, and Jundong Li. 2024. Knowledge Graph-Enhanced Large Language Models via Path Selection. (2024), 6311–6321. [45] Junling Liu, Chao Liu, Peilin Zhou, Renjie Lv, Kang Zhou, and Yan Zhang. 2023. Is chatgpt a good recommender? a preliminary study. arXiv preprint arXiv:2304.10149 (2023). [46] Wei Liu, Ailun Yu, Daoguang Zan, Bo Shen, Wei Zhang, Haiyan Zhao, Zhi Jin, and Qianxiang Wang. 2024. Graphcoder: Enhancing repository-level code completion via code context graph-based retrieval and language model. arXiv preprint arXiv:2406.07003 (2024). [47] Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: NLG evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 conference on empirical methods in natural language processing. 2511–2522. [48] Shiheng Ma, Jianhui Ding, Weijia Jia, Kun Wang, and Minyi Guo. 2017. Transt: Type-based multiple embedding representations for knowledge graph completion. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2017, Skopje, Macedonia, September 18–22, 2017, Proceedings, Part I 10. 717–733. [49] Shengjie Ma, Chengjin Xu, Xuhui Jiang, Muzhi Li, Huaren Qu, and Jian Guo. 2024. Think-on-graph 2.0: Deep and interpretable large language model reasoning with knowledge graph-guided retrieval. arXiv e-prints (2024), arXiv–2407. [50] Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2023. Self-refine: Iterative refinement with self-feedback. Advances in neural information processing systems 36 (2023), 46534–46594. [51] Costas Mavromatis and George Karypis. 2024. Gnn-rag: Graph neural retrieval for large language model reasoning. arXiv preprint arXiv:2405.20139 (2024). [52] Seyed Mahed Mousavi, Simone Alghisi, and Giuseppe Riccardi. 2024. Is Your LLM Outdated? Benchmarking LLMs & Alignment Algorithms for TimeSensitive Knowledge. arXiv preprint arXiv:2404.08700 (2024). [53] Yurii Nesterov. 2013. Introductory lectures on convex optimization: A basic course. Vol. 87. Springer Science & Business Media. [54] Christina Niklaus, Matthias Cetto, André Freitas, and Siegfried Handschuh. 2018. A Survey on Open Information Extraction. In Proceedings of the 27th International Conference on Computational Linguistics. 3866–3878. [55] Pouya Ghiasnezhad Omran, Kewen Wang, and Zhe Wang. 2019. An embeddingbased approach to rule learning in knowledge graphs. IEEE Transactions on Knowledge and Data Engineering 33, 4 (2019), 1348–1359. [56] OpenAI. 2024. https://openai.com/blog/chatgpt. [57] Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35 (2022), 27730–27744. [58] Heiko Paulheim. 2016. Knowledge graph refinement: A survey of approaches and evaluation methods. Semantic web 8, 3 (2016), 489–508. [59] Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, et al. 2023. Check your facts and try again: Improving large language models with external knowledge and automated feedback. arXiv preprint arXiv:2302.12813 (2023). [60] Boci Peng, Yun Zhu, Yongchao Liu, Xiaohe Bo, Haizhou Shi, Chuntao Hong, Yan Zhang, and Siliang Tang. 2024. Graph Retrieval-Augmented Generation: A Survey. arXiv preprint arXiv:2408.08921 (2024). [61] Diego Sanmartin. 2024. Kg-rag: Bridging the gap between knowledge and creativity. arXiv preprint arXiv:2405.12035 (2024). [62] Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna, Anna Goldie, and Christopher D Manning. 2024. Raptor: Recursive abstractive processing for tree-organized retrieval. In The Twelfth International Conference on Learning Representations. [63] Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. Advances in neural information processing systems 36 (2023), 8634– 8652. [64] Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. 2023. Large language models encode clinical knowledge. Nature 620, 7972 (2023), 172–180.

[65] Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, et al. 2023. Towards expert-level medical question answering with large language models. arXiv preprint arXiv:2305.09617 (2023). [66] Shamane Siriwardhana, Rivindu Weerasekera, Elliott Wen, Tharindu Kaluarachchi, Rajib Rana, and Suranga Nanayakkara. 2023. Improving the domain adaptation of retrieval augmented generation (RAG) models for open domain question answering. Transactions of the Association for Computational Linguistics 11 (2023), 1–17. [67] Dan Su, Yan Xu, Genta Indra Winata, Peng Xu, Hyeondey Kim, Zihan Liu, and Pascale Fung. 2019. Generalizing question answering system with pre-trained language model fine-tuning. In Proceedings of the 2nd workshop on machine reading for question answering. 203–211. [68] Budhitama Subagdja, D Shanthoshigaa, Zhaoxia Wang, and Ah-Hwee Tan. 2024. Machine learning for refining knowledge graphs: A survey. Comput. Surveys 56, 6 (2024), 1–38. [69] Jingwei Sun, Zhixu Du, and Yiran Chen. 2024. Knowledge Graph Tuning: Realtime Large Language Model Personalization based on Human Feedback. arXiv preprint arXiv:2405.19686 (2024). [70] Jiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang, Chen Lin, Yeyun Gong, Lionel Ni, Heung-Yeung Shum, and Jian Guo. 2024. Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph. [71] Zhiqing Sun, Zhi-Hong Deng, Jian-Yun Nie, and Jian Tang. 2019. RotatE: Knowledge Graph Embedding by Relational Rotation in Complex Space. In International Conference on Learning Representations. [72] Yixuan Tang and Yi Yang. 2024. MultiHop-RAG: Benchmarking RetrievalAugmented Generation for Multi-Hop Queries. arXiv:2401.15391 [73] Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. 2024. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295 (2024). [74] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023). [75] Prapti Trivedi, Aditya Gulati, Oliver Molenschot, Meghana Arakkal Rajeev, Rajkumar Ramamurthy, Keith Stevens, Tanveesh Singh Chaudhery, Jahnavi Jambholkar, James Zou, and Nazneen Rajani. 2024. Self-rationalization improves llm as a fine-grained judge. arXiv preprint arXiv:2410.05495 (2024). [76] Stefani Tsaneva, Danilo Dessì, Francesco Osborne, and Marta Sabou. 2025. Knowledge graph validation by integrating LLMs and human-in-the-loop. Information Processing & Management 62, 5 (2025), 104145. [77] Qiange Wang, Yanfeng Zhang, Hao Wang, Chaoyi Chen, Xiaodong Zhang, and Ge Yu. 2022. Neutronstar: distributed GNN training with hybrid dependency management. In Proceedings of the 2022 International Conference on Management of Data. 1301–1315. [78] Yiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren, Liyuan Liu, Baolin Peng, Hao Cheng, Xuehai He, Kuan Wang, Jianfeng Gao, et al. 2025. Reinforcement learning for reasoning in large language models with one training example. arXiv preprint arXiv:2504.20571 (2025). [79] Yanbin Wei, Qiushi Huang, Yu Zhang, and James Kwok. 2023. KICGPT: Large Language Model with Knowledge in Context for Knowledge Graph Completion. In Findings of the Association for Computational Linguistics: EMNLP 2023. 8667– 8683. [80] Nirmalie Wiratunga, Ramitha Abeyratne, Lasal Jayawardena, Kyle Martin, Stewart Massie, Ikechukwu Nkisi-Orji, Ruvan Weerasinghe, Anne Liret, and Bruno Fleisch. 2024. CBR-RAG: case-based reasoning for retrieval augmented generation in LLMs for legal question answering. In International Conference on Case-Based Reasoning. 445–460. [81] Han Xiao, Minlie Huang, and Xiaoyan Zhu. 2016. TransG: A Generative Model for Knowledge Graph Embedding. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2316– 2325. [82] Shitao Xiao, Zheng Liu, Peitian Zhang, and N Muennighof. 2023. C-pack: packaged resources to advance general Chinese embedding. 2023. arXiv preprint arXiv:2309.07597 (2023). [83] Ran Xu, Wenqi Shi, Yue Yu, Yuchen Zhuang, Bowen Jin, May Dongmei Wang, Joyce Ho, and Carl Yang. 2024. RAM-EHR: Retrieval Augmentation Meets Clinical Predictions on Electronic Health Records. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 754–765. [84] Yifei Xu, Tusher Chakraborty, Emre Kiciman, Bibek Aryal, Srinagesh Sharma, Songwu Lu, and Ranveer Chandra. 2025. RLTHF: Targeted Human Feedback for LLM Alignment. In Forty-second International Conference on Machine Learning. [85] Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 14

Conference on Empirical Methods in Natural Language Processing. [86] Ziyi Ye, Xiangsheng Li, Qiuchi Li, Qingyao Ai, Yujia Zhou, Wei Shen, Dong Yan, and Yiqun Liu. 2024. Beyond scalar reward model: Learning generative judge from preference data. arXiv preprint arXiv:2410.03742 (2024). [87] Kun Yi and Jianxin Wu. 2019. Probabilistic end-to-end noise correction for learning with noisy labels. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 7017–7025. [88] Yue Yu, Zhengxing Chen, Aston Zhang, Liang Tan, Chenguang Zhu, Richard Yuanzhe Pang, Yundi Qian, Xuewei Wang, Suchin Gururangan, Chao Zhang, et al. 2025. Self-generated critiques boost reward modeling for language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 11499–11514. [89] Hao Yuan, Xin Ai, Qiange Wang, Peizheng Li, Jiayang Yu, Chaoyi Chen, Xinbo Yang, Yanfeng Zhang, Zhenbo Fu, Yingyou Wen, et al. 2025. DepCache: A KV Cache Management Framework for GraphRAG with Dependency Attention. Proceedings of the ACM on Management of Data 3, 6 (2025), 1–29. [90] Hongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. 2023. Rrhf: Rank responses to align language models with human feedback. Advances in Neural Information Processing Systems 36 (2023), 10935– 10950. [91] Mert Yuksekgonul, Federico Bianchi, Joseph Boen, Sheng Liu, Pan Lu, Zhi Huang, Carlos Guestrin, and James Zou. 2025. Optimizing generative AI by backpropagating language model feedback. Nature 639, 8055 (2025), 609–616. [92] Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. 2022. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414 (2022). [93] Biao Zhang, Barry Haddow, and Alexandra Birch. 2023. Prompting large language model for machine translation: A case study. In Proceedings of International Conference on Machine Learning. 41092–41110. [94] Boyu Zhang, Hongyang Yang, Tianyu Zhou, Muhammad Ali Babar, and XiaoYang Liu. 2023. Enhancing financial sentiment analysis via retrieval augmented large language models. In Proceedings of the fourth ACM international conference on AI in finance. 349–356. [95] Fangyuan Zhang, Zhengjun Huang, Yingli Zhou, Qintian Guo, Zhixun Li, Wensheng Luo, Di Jiang, Yixiang Fang, and Xiaofang Zhou. 2025. EraRAG: Efficient and Incremental Retrieval Augmented Generation for Growing Corpora. arXiv

preprint arXiv:2506.20963 (2025). [96] Han Zhang, Yu Lei, Lin Gui, Min Yang, Yulan He, Hui Wang, and Ruifeng Xu. 2024. Cppo: Continual learning for reinforcement learning with human feedback. In The Twelfth International Conference on Learning Representations. [97] Qinggang Zhang, Shengyuan Chen, Yuanchen Bei, Zheng Yuan, Huachi Zhou, Zijin Hong, Junnan Dong, Hao Chen, Yi Chang, and Xiao Huang. 2025. A Survey of Graph Retrieval-Augmented Generation for Customized Large Language Models. arXiv preprint arXiv:2501.13958 (2025). [98] Qinggang Zhang, Junnan Dong, Hao Chen, Daochen Zha, Zailiang Yu, and Xiao Huang. 2024. Knowgpt: Knowledge graph based prompting for large language models. Advances in Neural Information Processing Systems 37 (2024), 6052–6080. [99] Qinggang Zhang, Junnan Dong, Keyu Duan, Xiao Huang, Yezi Liu, and Linchuan Xu. 2022. Contrastive knowledge graph error detection. In Proceedings of the 31st ACM international conference on information & knowledge management. 2590–2599. [100] Qizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma, Fenglu Hong, Vamsidhar Kamanuru, Jay Rainton, Chen Wu, Mengmeng Ji, Hanchen Li, et al. 2025. Agentic context engineering: Evolving contexts for self-improving language models. arXiv preprint arXiv:2510.04618 (2025). [101] Penghao Zhao, Hailin Zhang, Qinhan Yu, Zhengren Wang, Yunteng Geng, Fangcheng Fu, Ling Yang, Wentao Zhang, Jie Jiang, and Bin Cui. 2024. Retrievalaugmented generation for ai-generated content: A survey. arXiv preprint arXiv:2402.19473 (2024). [102] XUJIANG ZHAO, JIAYING LU, CHENGYUAN DENG, C ZHENG, JUNXIANG WANG, TANMOY CHOWDHURY, L YUN, HEJIE CUI, ZHANG XUCHAO, TIANJIAO ZHAO, et al. 2023. Beyond One-Model-Fits-All: A Survey of Domain Specialization for Large Language Models. arXiv preprint arXiv 2305 (2023). [103] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems 36 (2023), 46595–46623. [104] Yingli Zhou, Yaodong Su, Youran Sun, Shu Wang, Taotao Wang, Runyuan He, Yongwei Zhang, Sicong Liang, Xilin Liu, Yuchi Ma, et al. 2025. In-depth Analysis of Graph-based RAG in a Unified Framework. arXiv preprint arXiv:2503.04338 (2025).

15

Related documents

Record · ID 31351 · SHA-256 20d4ed76e68b541b
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.