R ECTIFY: An Interactive Workbench for Post-Evaluation RAG Diagnosis, Repair, and Verification Keerthana Murugaraj, Salima Lamsiyah, Martin Theobald University of Luxembourg Esch-sur-Alzette, Luxembourg Correspondence: [email protected]
arXiv:2609.16764v1 [cs.SE] 15 Sep 2026
Abstract
highlight that RAG evaluation must account for both retrieval and generation behavior, including relevance, faithfulness, answer quality, robustness, and benchmark design aligned with real user needs (Gao et al., 2023b; Yu et al., 2025). Yet evaluation alone does not close the debugging loop. Although recent diagnostic frameworks provide actionable per-case recommendations (Cohen et al., 2025), developers are still often left to decide which failures share a common root cause, which pipeline component should be changed, and whether a proposed change actually improves the affected examples. This post-evaluation step is especially important because treating every failed case as an isolated error makes repair slow, inconsistent, and difficult to audit. We present R ECTIFY, an interactive workbench for RAG failure diagnosis, repair, and optional verification. R ECTIFY uses RAGVue (Murugaraj et al., 2026) as its primary evaluator and turns diagnostic evaluation outputs into a post-evaluation repair workflow. It takes evaluated cases, groups recurring failures into four actionable families: retrieval, grounding, generation, and abstention, and maps them to fine-grained repair slices. Each slice is converted into an editable repair card that a developer can inspect, approve, or reject before any change is tested. Approved repairs can optionally be verified in a sandbox on the affected cases, with before and after deltas and provenance logs supporting auditability. Our key contributions are as follows:
Retrieval-Augmented Generation (RAG) evaluators can identify failures such as weak retrieval, poor grounding, incomplete answers, and unsupported generation, but they rarely help developers decide what to repair next. We present R ECTIFY, an interactive Streamlit workbench that turns evaluated RAG cases into auditable repair workflows. R ECTIFY filters cases that do not require repair, routes remaining failures into actionable families and finegrained repair slices, and generates editable repair cards that developers can approve, reject, or verify through sandbox reruns. On a controlled RAG benchmark, R ECTIFY surfaces interpretable failure profiles across BM25, dense, and hybrid retrieval: BM25 mainly triggers noisy-retrieval repairs, while dense and hybrid retrieval leave smaller sets of multi-part underretrieval and underused-evidence cases. Additional analyses show that pre-filtering reduces unnecessary repair candidates and that slicelevel routing yields more targeted repair cards than broad family-level diagnosis. R ECTIFY is publicly available as an open-source Streamlit workbench 1 for helping developers turn evaluation results into inspectable repair decisions.
1
Introduction
RAG has become a common design pattern for building language-model systems that answer questions over external documents (Lewis et al., 2020). This makes RAG useful in settings such as enterprise question answering, scientific search, legal and historical archives, and customer-support systems. (Wiratunga et al., 2024; Murugaraj et al., 2025; Xu et al., 2024) However, RAG pipelines remain difficult to debug: a low-quality answer may arise from failures in retrieval, grounding, answer synthesis, or abstention (Es et al., 2024; Murugaraj et al., 2026; Ru et al., 2024). Recent surveys 1
• We present R ECTIFY, an open-source Streamlit workbench for post-evaluation RAG debugging that turns evaluator outputs into case exploration, failure diagnosis, repair-card review, provenance logging, and sandbox verification. • We introduce a deterministic failure-routing scheme that filters non-actionable cases and maps remaining failures to four macro fami-
§ Code Repository
1
1. Unified Case Representation
2. Failure Families & Repair Slices
3. Repair-Card Generation
4. Human Approval & Provenance
5. Optional Sandbox Verification
answers, contexts, scores
pre-filters; 4 families; 23 slices
template-based repair hypotheses
inspect, edit, approve/reject
rerun cases; report deltas
provenance log scope, notes, timestamp
Figure 1: R ECTIFY post-evaluation workflow. Evaluated cases are normalized, routed into failure families and repair slices, converted into editable repair cards, reviewed by a developer, and optionally verified through sandbox reruns with provenance logging.
From Evaluation to Guidance and Repair. Recent approaches move from evaluation toward developer guidance. RAGXplain converts evaluation scores into per-case natural-language explanations for improving RAG pipelines (Cohen et al., 2025). RAGGY provides composable RAG primitives and an interactive interface for real-time pipeline debugging (Romero Lauro et al., 2026). ARAGOG compares RAG configurations such as reranking, multi-query retrieval, maximal marginal relevance, HyDE, and sentence-window retrieval (Eibich et al., 2024). Doctor-RAG studies failureaware repair for agentic RAG by localizing failures in reasoning trajectories and repairing the diagnosed point (Jiao et al., 2026). These systems make evaluation more actionable, but they do not focus on dataset-level post-evaluation failure grouping, human approval of repair cards, and measured before–after verification for standard RAG pipelines.
lies and 23 fine-grained repair slices, enabling cluster-level repair decisions rather than caseby-case inspection. • We design editable repair cards that connect each failure slice to inspectable configurationlevel interventions over retrieval, chunking, reranking, prompting, abstention, and generation while keeping developers in control of approval and scope. • We evaluate R ECTIFY on a controlled RAG benchmark across BM25, dense, and hybrid retrieval, showing interpretable retriever-specific repair agendas, fewer unnecessary repair candidates after pre-filtering, and more targeted cards than family-level routing.
2
Related Work
RAG Evaluation and Benchmarks. Recent work has developed metrics and benchmarks for evaluating RAG systems beyond end-task accuracy. RAGAS introduced reference-free metrics for faithfulness, answer relevancy, context precision, and context recall (Es et al., 2024). ARES trains lightweight judges for context relevance, answer faithfulness, and answer relevance using synthetic data and limited human annotation (Saad-Falcon et al., 2024). RAGChecker separates retrieval and generation behavior through fine-grained diagnostic metrics (Ru et al., 2024). RAGVue provides diagnostic and explainable reference-free evaluation across retrieval quality, answer relevance and completeness, strict faithfulness, and calibration (Murugaraj et al., 2026). Complementary benchmarks study hallucination, robustness, citation-supported generation, and actionable evaluation labels, including RAGTruth, RGB, ALCE, and RAGBench (Niu et al., 2024; Chen et al., 2024; Gao et al., 2023a; Friel et al., 2024). These works expose important quality signals, but they do not by themselves define a human-approved repair workflow.
In contrast to the existing works, R ECTIFY addresses the gap between RAG evaluation and pipeline revision. It is neither another evaluator nor an automatic self-repair system: it starts from already evaluated cases and treats proposed changes as repair hypotheses that require developer inspection and approval. Given evaluator outputs, R EC TIFY groups recurring failures into repair slices, generates editable configuration-level repair cards, optionally verifies approved repairs on the affected cases, reports before and after deltas, and records decisions in a provenance log. This positions R EC TIFY as a post-evaluation workbench for auditable, human-controlled RAG diagnosis and repair.
3
The R ECTIFY System
R ECTIFY operates after a RAG pipeline has produced answers and an evaluator has scored them. Figure 1 summarizes the post-evaluation workflow. 2
3.1
Unified Case Representation
Evaluated RAG case
R ECTIFY first normalizes each evaluated RAG case into a shared schema. Each case contains the question, generated answer, retrieved contexts, optional expected answer, evaluator scores, diagnostic fields, and available metadata. This common representation allows R ECTIFY to apply the same workflow across cases: failure routing, repair-card generation, human review, provenance logging, and optional sandbox verification. In the current implementation, RAGVue is the primary evaluator because it provides diagnostic signals for retrieval quality, grounding, answer completeness, abstention behavior, and response quality. These signals are used to assign cases to failure families and fine-grained repair slices.
Correct abstention? no yes
Gold answer available? yes
no
Answer close + grounded?
yes
no
No repair needed
no
Any failure trigger? yes Layer 1: Macro family Abstention → Retrieval → Grounding → Generation
3.2
Failure Families & Repair Slices Record secondary family if another condition also fires
R ECTIFY uses a two-layer taxonomy to convert evaluator outputs into repairable failure patterns. Before routing begins, two pre-filters remove cases that do not require repair. First, a correct-abstention filter excludes unanswerable cases when the model appropriately refuses to answer. Second, a goldanswer filter excludes answerable cases when the generated answer is sufficiently close to the gold answer and grounded in the retrieved context. The routing process is illustrated in Figure 2.
Layer 2: Repair slice 23 slices within selected family
Generate repair card
Figure 2: Failure routing in R ECTIFY.
3.3
Cases not removed by these pre-filters are assigned to a primary macro failure family in the following priority order: abstention, retrieval, grounding, and generation. This ordering is conservative: unsupported confident answers are handled first, and retrieval is checked before grounding because faithfulness is difficult to interpret when evidence is weak or noisy. Generation is used when retrieval and grounding are adequate, but the final answer remains unsatisfactory. When multiple failure conditions hold, R ECTIFY also records a secondary family to support compound repairs. Within the selected macro family, R ECTIFY assigns the case to a fine-grained repair slice. The current taxonomy defines 23 slices across the four families, with each slice linked to a repair-card template. This repair step is more specific than family-level diagnosis alone: for example, different retrieval failures may require reranking, broader evidence retrieval, or chunking changes, while grounding failures may require more evidence-constrained prompting. The full failure taxonomy is presented in Table 1.
Repair-Card Generation
For each populated repair slice, R ECTIFY creates a repair card: a structured, editable proposal that describes the diagnosed failure, the affected cases, and the suggested configuration-level intervention. Card generation is deterministic. Given the assigned repair slice and affected case IDs, R EC TIFY selects the corresponding template and fills in the target pipeline stage, proposed change, expected benefit, expected tradeoff, and repair scope. Repair cards turn recurring failure patterns into concrete repair hypotheses. For example, a noisyretrieval slice may suggest enabling reranking, while an underused-evidence slice may suggest a more evidence-constrained prompt. For compound failures, R ECTIFY can combine interventions across stages, such as reranking together with grounded prompting. 3.4
Human Approval & Provenance
R ECTIFY keeps the human at the decision point. The user can inspect the repair slice, review af3
Macro family
Slice
Name
Short interpretation
Abstention
A1 A2 A3
Confident unsupported answer Partial-evidence overconfidence Ambiguous forced answer
Model answers confidently despite no supporting evidence. Model answers fully when evidence only partially supports the claim. Model picks one interpretation instead of flagging ambiguity.
Retrieval
R1 R2 R3 R4 R5 R6 R7
Partial coverage Noisy retrieval Evidence ignored Fragmented evidence Multi-part under-retrieval Distractor-dominated Sparse evidence
Retrieved chunks cover only part of what the question requires. Retrieved chunks contain irrelevant or distracting content. Relevant chunks are retrieved but not used in the answer. Relevant information is split across chunks, none sufficient alone. Retriever returns chunks for only one part of a multi-faceted question. High-scoring irrelevant chunks crowd out relevant ones. Corpus lacks sufficient content to answer the question.
Grounding
G1 G2 G3 G4 G5 G6 G7
Temporal misattribution Entity substitution Unsupported causal bridge Omitted qualifier Broken multi-hop Unsupported synthesis Citation drift
Correct fact assigned to the wrong time period. Correct relationship stated with the wrong entity name. Model infers a causal link not stated in the context. Context is hedged but the answer states the claim as absolute. Error introduced while chaining reasoning steps across chunks. Model combines facts across chunks in an unsupported way. Answer uses correct content but attributes it to the wrong entity.
Generation
S1 S2 S3 Q1 Q2 Q3
Partial aspect coverage Shallow summarization Underused evidence Rambling Poor structure Internal inconsistency
Answer addresses some but not all aspects of the question. Answer is too surface-level and misses important context details. Relevant context is retrieved but not incorporated in the answer. Answer is verbose, unfocused, or contains unnecessary content. Answer is hard to follow due to disorganized presentation. Answer contradicts itself within the same response.
Table 1: R ECTIFY failure taxonomy: four macro families and 23 fine-grained repair slices.
fected cases, edit the proposed parameters, choose the repair scope, and approve or reject the card. Each approval or rejection is stored in a provenance log containing the repair-card identifier, slice type, affected cases, approved parameters, scope, user notes, timestamp, expected benefit, and expected tradeoff. This log makes the repair process auditable: it records what was proposed, what was approved or rejected, which cases were affected, and under which configuration the repair was considered or tested. 3.5
and start the interface with a standard command such as streamlit run streamlit_app.py. The local UI supports loading evaluation files, inspecting diagnosed cases, reviewing repair cards, and generating reports without writing code. This interface is intended for practitioners who prefer a point-and-click workflow while keeping data and credentials on their own machine. Selected screenshots of the interface are provided in Appendix A, and the full walkthrough is included in the repository and demo video.
Optional Sandbox Verification
After approval, a repair card can be verified in a sandbox before it is accepted as useful. The sandbox applies the approved configuration patch to the affected cases, reruns the RAG pipeline, reevaluates the new outputs, and compares them with the original evaluation results. In our experiments, this verification uses RAGVue as the primary evaluator. When sandbox verification is run, the delta report summarizes whether each affected case improved, remained unchanged, or regressed. R ECTIFY reports aggregate counts, improvement rate, average metric deltas, and per-case before and after outputs. This makes repair proposals empirically checkable rather than merely plausible.
4
Evaluation
4.1
Experimental Setup
We evaluate R ECTIFY on a controlled 100-question synthetic RAG benchmark; dataset construction and question statistics are described in Appendix B. We test three retriever configurations over the same corpus and question set: BM25 keyword retrieval (Robertson and Zaragoza, 2009), dense retrieval with all-MiniLM-L6-v2 embeddings (Reimers and Gurevych, 2019; Wang et al., 2020), and a hybrid BM25+dense retriever. For all configurations, Mistral-7B (Jiang et al., 2023) is used as the generator model, served locally via Ollama2 . The generated outputs are then evaluated with 12 RAGVue (Murugaraj et al., 2026) metrics using the same Mistral-7B model as the judge.
Local Streamlit Application. For interactive use, R ECTIFY provides a Python-based Streamlit interface that exposes the main workflow through a local browser application. Users can clone the repository
2
4
https://ollama.com/
Measure
BM25
Dense
Hybrid
Case accounting Unanswerable prefiltered Gold-answer filter
16 62
16 75
16 71
Subtotal: filtered before diagnosis
78
91
87
No repair family triggered Final repair agenda
6 16
4 5
7 6
Repair-slice breakdown R2: Noisy retrieval R5: Multi-part under-retrieval S3: Underused evidence
7 3 6
0 4 1
0 1 5
Table 2: R ECTIFY case accounting and repair-slice breakdown across BM25, Dense, and Hybrid retrieval. The upper block shows filtering and routing outcomes; the lower block shows the final repair agenda.
4.2
Pattern across re- Count Interpretation trievers No repair in all retrievers
81
Fails in at least one retriever
19
Same S3 slice in ≥2 retrievers Same R5 slice in ≥2 retrievers
3
R2 only with BM25
7
Other one-off mixed failures
7
or
2
The question is handled correctly across BM25, Dense, and Hybrid retrieval Questions used to analyze repair-slice consistency Evidence is retrieved but underused by the generator The question needs evidence from multiple documents BM25 introduces noisy keyword-matched chunks Failure depends on retriever setting or slice interaction
Table 3: Repair-slice consistency across BM25, Dense, and Hybrid retrieval.
policy changes for S3. This demonstrates that R EC -
Results
TIFY surfaces interpretable diagnostics that are di-
We focus the evaluation on three main questions, described below. Additional ablations on the prefiltering stage and taxonomy depth are reported in Appendix E.
rectly tied to repair actions. Are repair slices consistent across retrievers? We compare how R ECTIFY labels the same 100 questions under BM25, Dense, and Hybrid retrieval. If the same question receives the same repair slice under multiple retrievers, we treat it as a stable failure pattern. If a slice appears only under one retriever, the failure is more likely tied to that retrieval configuration. Table 3 shows that 81 questions need no repair under any retriever. Among the 19 questions that fail at least once, repeated S3 and R5 labels reveal stable problems: S3 means that retrieved evidence is available but underused, while R5 means that the question requires evidence from multiple documents. In contrast, seven R2 cases appear only with BM25, indicating keyword-matching noise that disappears with dense or hybrid retrieval. This suggests that R ECTIFY’s slices are not arbitrary labels: they help separate stable failure patterns from retriever-specific errors.
How do failure profiles differ across retriever configurations, and does R ECTIFY surface consistent and interpretable diagnostics? Table 2 summarizes the main R ECTIFY outputs for the three retriever configurations: how cases are filtered before diagnosis, how many cases remain in the final repair agenda, and which repair slices are triggered. The results show that retriever choice substantially changes the repair agenda. BM25 leaves 16 cases requiring repair, while dense and hybrid retrieval reduce this to 5 and 6 cases, respectively. This indicates that many failures are retrieval-driven: semantic retrieval resolves cases where BM25 retrieves lexically plausible but weakly relevant evidence. Hybrid retrieval does not clearly dominate dense retrieval in the final repair agenda, but it still preserves complementary lexical and semantic signals. The repair-slice breakdown shows why aggregate evaluator scores (Table 4) alone are not sufficient. Under BM25, the main repair need is R2 noisy retrieval. With dense and hybrid retrieval, R2 drops to 0 cases, and the remaining failures shift toward R5 multi-part under-retrieval and S3 underused evidence. These failure types can all lead to weak faithfulness, completeness, or answer relevance, but they require different repairs: reranking for R2, broader evidence retrieval for R5, and prompt-level or abstention-
Does R ECTIFY reduce debugging effort? We estimate debugging effort under three workflows. In a raw metric-scan workflow, a developer inspects every case where a core metric falls below a threshold. With pre-filtering only, the developer inspects the post-filter-failing cases individually, without any grouping. With R ECTIFY, the same pre-filters are applied first, and the remaining failures are then grouped into repair cards, so the developer reviews one card per repair slice and makes one repair decision per cluster. Figure 3 shows that the raw metric 5
Estimated repair decisions across workflows A: Raw metric scan B: Pre-filters only C: Rectify (ours)
100 89
Decisions for developer
comes smaller, and the dominant failure modes shift: BM25 mainly suffers from noisy retrieval, while dense and hybrid retrieval expose more specific failures, such as multi-part under-retrieval and underused evidence. This shift would be difficult to interpret from aggregate metrics alone, but becomes visible through R ECTIFY’s slice-level diagnosis. The results also show why repair should be treated as a review-and-verification workflow rather than an automatic patching step. Pre-filtering reduces unnecessary repair workload, fine-grained slicing produces targeted repair cards, and sandbox verification exposes both successful repairs and cases where a proposed intervention should be rejected or revised. Overall, R ECTIFY contributes a post-evaluation workflow for RAG development: diagnose recurring failures, propose inspectable repairs, keep the developer in control, and verify changes with before and after evidence.
80
75
74
60 40 20
16
97% 0
3
BM25
5 97%
2
DENSE
6 97%
2
HYBRID
Figure 3: Estimated debugging effort across three workflows: raw metric scan, pre-filtering only, and R ECTIFY with pre-filtering plus repair-slice grouping. Effort is measured as the number of case-level or card-level repair decisions.
Limitations and Future Work
scan requires 74-89 individual case inspections per retriever configuration. The pre-filters reduce this to 5-16 post-filter cases. R ECTIFY’s taxonomy then groups those cases into only 2-3 repair cards while still covering the full final repair agenda. Under our decision-count estimate, R ECTIFY reduces the number of repair decisions by 97%. This reduction is important because it changes the debugging unit. Instead of making many ad hoc case-level judgements, the developer reviews a small number of cluster-level repair hypotheses: for example, enabling reranking for R2 noisy retrieval, increasing top-k for R5 multi-part under-retrieval, or adjusting prompting for S3 underused evidence. Finally, because sandbox verification is optional in the workflow, we report detailed sandbox results separately in Appendix D. These results illustrate how approved repair cards can be checked through before and after deltas before a developer accepts or rejects a repair.
5
R ECTIFY is a local, interactive workbench for postevaluation RAG repair, not an automatic production optimizer. Its current implementation uses a deterministic taxonomy with 23 repair slices. This design makes routing decisions reproducible and auditable, while also making the taxonomy easy to extend as new domains, retrievers, or generation failure patterns introduce additional repair needs. The pre-filtering stage can use benchmark annotations such as gold answers and answerability labels when they are available, but these annotations are not required for the core workflow: without them, R ECTIFY still performs diagnosis from evaluator signals and simply skips annotation-based filtering. Sandbox verification reruns affected cases and reevaluates the new outputs, so developers can apply it selectively to high-impact repair cards or use it as a final check before accepting a proposed configuration change. Future work will expand R ECTIFY along three directions: evaluating it on larger real-world knowledge bases, extending the repair taxonomy to additional RAG architectures and application domains, and studying how the workflow transfers across evaluators with different diagnostic signals. We also plan to test additional generator-judge combinations to further assess the stability of the observed repair profiles.
Conclusion
We presented R ECTIFY, an interactive workbench prototype for moving from RAG evaluation to diagnosis, repair, and verification. Our experiments show that R ECTIFY surfaces actionable failure profiles across retriever configurations. As retrieval quality improves, the repair agenda be6
Ethics Statement
Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral-7b. ArXiv, abs/2310.06825.
R ECTIFY is a developer-facing workbench for postevaluation RAG diagnosis and repair. It does not automatically modify or deploy RAG systems; repair cards require human review, approval, and optional sandbox verification. Our experiments use a synthetic benchmark with fictional entities and do not involve private or personally identifiable data. In real deployments, users should ensure that input documents, evaluation files, model outputs, and connected services comply with relevant privacy, licensing, and data-governance requirements. R EC TIFY ’s suggestions should be treated as decision support rather than guaranteed fixes, especially in high-stakes domains.
Shuguang Jiao, Chengkai Huang, Shuhan Qi, Xuan Wang, Yifan Li, and Lina Yao. 2026. Doctor-RAG: Failure-aware repair for agentic retrieval-augmented generation. arXiv preprint arXiv:2604.00865. Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledgeintensive NLP tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS’20). Curran Associates Inc.
References
Keerthana Murugaraj, Salima Lamsiyah, Marten During, and Martin Theobald. 2025. Topic-rag for historical newspapers: Enhancing information retrieval in humanities research through topic-based retrievalaugmented generation. Computational Humanities Research, 1:e15.
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking large language models in retrieval-augmented generation. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence. AAAI Press.
Keerthana Murugaraj, Salima Lamsiyah, and Martin Theobald. 2026. RAGVUE: A diagnostic view for explainable and automated evaluation of retrievalaugmented generation. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 512–526. Association for Computational Linguistics.
Dvir Cohen, Lin Burg, and Gilad Barkan. 2025. RAGXplain: From explainable evaluation to actionable guidance of RAG pipelines. arXiv preprint arXiv:2505.13538. Matouš Eibich, Shivay Nagpal, and Alexander FredOjala. 2024. ARAGOG: Advanced RAG output grading. arXiv preprint arXiv:2404.01037.
Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, KaShun Shum, Randy Zhong, Juntong Song, and Tong Zhang. 2024. RAGTruth: A hallucination corpus for developing trustworthy retrieval-augmented language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10862– 10878. Association for Computational Linguistics.
Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024. RAGAs: Automated evaluation of retrieval augmented generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, pages 150–158. Association for Computational Linguistics.
Nils Reimers and Iryna Gurevych. 2019. SentenceBERT: Sentence Embeddings using siamese BERTNetworks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP’19), pages 3982–3992.
Robert Friel, Masha Belyi, and Atindriyo Sanyal. 2024. RAGBench: Explainable benchmark for retrievalaugmented generation systems. arXiv preprint arXiv:2407.11005. Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023a. Enabling large language models to generate text with citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6465–6488. Association for Computational Linguistics.
Stephen Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: Bm25 and beyond. Found. Trends Inf. Retr., 3(4):333–389.
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2023b. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997.
Quentin Romero Lauro, Shreya Shankar, Sepanta Zeighami, and Aditya Parameswaran. 2026. RAG Without the Lag: Enabling "What-If" Analysis for Retrieval-Augmented Generation Pipelines. In Proceedings of the 2026 CHI Conference on Human
7
Factors in Computing Systems (CHI’26). Association for Computing Machinery. Dongyu Ru, Lin Qiu, Xiangkun Hu, Tianhang Zhang, Peng Shi, Shuaichen Chang, Cheng Jiayang, Cunxiang Wang, Shichao Sun, Huanyu Li, Zizhao Zhang, Binjie Wang, Jiarong Jiang, Tong He, Zhiguo Wang, Pengfei Liu, Yue Zhang, and Zheng Zhang. 2024. RAGCHECKER: a fine-grained framework for diagnosing retrieval-augmented generation. In Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS’24). Curran Associates Inc. Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2024. ARES: An automated evaluation framework for retrieval-augmented generation systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 338–354. Association for Computational Linguistics. Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020. Minilm: deep selfattention distillation for task-agnostic compression of pre-trained transformers. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS’20). Curran Associates Inc. Nirmalie Wiratunga, Ramitha Abeyratne, Lasal Jayawardena, Kyle Martin, Stewart Massie, Ikechukwu NkisiOrji, Ruvan Weerasinghe, Anne Liret, and Bruno Fleisch. 2024. Cbr-rag: case-based reasoning for retrieval augmented generation in llms for legal question answering. In International Conference on CaseBased Reasoning, pages 445–460. Springer. Zhentao Xu, Mark Jerome Cruz, Matthew Guevara, Tie Wang, Manasi Deshpande, Xiaofeng Wang, and Zheng Li. 2024. Retrieval-augmented generation with knowledge graphs for customer service question answering. In Proceedings of the 47th international ACM SIGIR conference on research and development in information retrieval, pages 2905–2909. Hao Yu, Aoran Gan, Kai Zhang, Shiwei Tong, Qi Liu, and Zhaofeng Liu. 2025. Evaluation of retrievalaugmented generation: A survey. In Big Data, pages 102–120. Springer Nature Singapore.
8
A
Additional Screenshots
Example UI screenshots are shown in Figure 4, and the full walkthrough is available in the repository and the demo video.
B
Dataset
We construct a controlled synthetic corpus with 30 short documents about 10 fictional companies and associated entities. The documents cover company profiles, persons, products, events, and thematic summaries and encode factual relations such as founding year, headquarters, acquisitions, partnerships, and flagship products. We generate 100 questions across eight types: factoid (25), multipart (20), multi-hop (16), temporal (13), comparison (10), explicitly unanswerable (10), multi-hop unanswerable (4), and temporal unanswerable (2). In total, 84 questions are answerable, and 16 are unanswerable. Answerable questions include gold answers and known relevant document identifiers; unanswerable questions test whether the system correctly abstains when the required evidence is absent from the corpus.
Dense
Hybrid
0.627 0.450 0.757 0.391 0.648 0.465 0.655 0.999 0.971 0.990 0.975 0.924
0.745 0.567 0.823 0.474 0.775 0.570 0.778 1.000 0.981 0.970 0.993 0.909
0.741 0.563 0.830 0.465 0.762 0.596 0.762 1.000 0.987 0.990 0.991 0.913
Slice
Repair
n
Improved
Unchanged
Regressed
R2 R5
Enable reranker Increase top-k 3→6
7 8
7 5
0 2
0 1
Table 5: Sandbox verification results for two representative repair slices.
cases, leaves two unchanged, and regresses one. These results support the review-and-verification loop in R ECTIFY. Repair cards are testable hypotheses, not guaranteed fixes: sandbox deltas help developers decide whether a proposed repair improves, leaves unchanged, or regresses the affected cases before accepting it.
E
Additional Ablations
E.1
Pre-filters Reduce Unnecessary Repair Candidates
R ECTIFY applies two pre-filters before failure routing, as described in Section 3.2. These filters are not intended as a replacement for evaluation; they reduce unnecessary repair work before the remaining cases are diagnosed and grouped into repair slices. Table 6 shows the effect of these filters. A raw metric scan flags 89, 74, and 75 cases for BM25, dense, and hybrid retrieval, respectively. When the failure taxonomy is applied without prefilters, the candidate set becomes 68, 54, and 55 cases. With the abstention and gold-answer filters enabled, the final repair agenda drops to 16, 5, and 6 cases. This corresponds to an 82–93% reduction compared with raw metric scanning and a 77–91% reduction compared with taxonomy-based routing without pre-filters. The number of active failure slices also decreases from eight to three, two, and two, producing a smaller and more focused repair agenda.
Full RAGVue Metric Scores
Table 4 reports the full set of RAGVue metric scores used as evaluator signals in the crossretriever analysis.
D
BM25
Strict faithfulness Retrieval relevance Retrieval coverage Answer completeness Answer relevance Context utilization Multi-hop faithfulness Coherence Clarity Negative rejection Answer conciseness Implicit contradiction
Table 4: Full RAGVue metric scores for BM25, Dense, and Hybrid retrieval over the same 100-question benchmark. Bold marks the best value in each row.
Why synthetic? We use a small controlled corpus because R ECTIFY is evaluated as a repair workbench, not as an open-domain QA system. Repair evaluation requires more control than gold answers alone: we need known relevant documents, controlled unanswerable cases, and known corpus coverage gaps to distinguish retrieval failures, missing evidence, poor evidence use, and correct abstention. The synthetic setup also lets us create targeted failure conditions, such as noisy retrieval, missing multi-hop evidence, underused evidence, and unsupported answers. We acknowledge that this corpus is narrower than real enterprise settings, and future work will evaluate R ECTIFY on larger real-world knowledge bases.
C
RAGVue metric
Sandbox Verification Results
Table 5 illustrates sandbox verification for two representative repair slices. For R2 noisy retrieval, applying reranking to seven BM25 cases improves all seven. For R5 multi-part under-retrieval, increasing top-k from 3 to 6 improves five of eight 9
(a) Home
(b) Repair card (c) Delta explorer
Figure 4: Selected screenshots of the local Python-based R ECTIFY Streamlit interface.
Measure
BM25
Dense
Hybrid
Raw metric scan Family routing, no pre-filters R ECTIFY with pre-filters
89 68 16
74 54 5
75 55 6
Reduction vs. raw scan Reduction vs. no pre-filters
82% 77%
93% 91%
92% 89%
evidence retrieval, such as increasing top-k. Finegrained slicing separates these cases into more specific repair cards and keeps S3 underused-evidence cases separate from retrieval failures. This makes the repair agenda more actionable: developers review slightly more specific cards, but each card corresponds to a clearer repair hypothesis.
Table 6: Ablation of R ECTIFY’s pre-filtering stage. Raw metric scan counts broadly flagged cases, while family routing applies the taxonomy before disabling the abstention and gold-answer filters. Measure
BM25
Dense
Hybrid
2 3 4 1
2 3 1 5
2 3 – Yes
2 3 – Yes
Family-only routing Repair cards Avg. cases per card Cases merged into retrieval card Cases merged into generation card
2 8 10 6
Fine-grained slices Repair cards Avg. cases per card R2/R5 separated S3 isolated from retrieval
3 5 Yes Yes
Table 7: Ablation of taxonomy depth over the post-filter repair agenda. Fine-grained slices separate failures that require different repair actions.
E.2
Fine-Grained Slices Produce More Targeted Repair Cards
We ablate the second layer of R ECTIFY’s taxonomy by comparing fine-grained repair slices with a family-only routing baseline. Family-only routing groups failures into broad families such as retrieval or generation, while fine-grained routing separates them into repair slices such as R2 noisy retrieval, R5 multi-part under-retrieval, and S3 underused evidence. Table 7 shows the practical effect of this distinction. Under BM25, family-only routing merges 10 retrieval-family cases into a single card, although these cases require different interventions: R2 points to reranking, while R5 points to broader 10