ConceptioArchivearXiv CS
arXiv CSopen access

Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing Wenhao Yuan1 , Chenchen Lin2 , Jian Chen1 , Jinfeng Xu1 , Xuehe Wang2 , Edith Cheuk Han Ngai1,* 1 The University of Hong Kong, 2 Sun Yat-sen University [email protected], [email protected]

Abstract

arXiv:2604.08401v1 [cs.AI] 9 Apr 2026

In large language model (LLM) agents, reasoning trajectories are treated as reliable internal beliefs for guiding actions and updating memory. However, coherent reasoning can still violate logical or evidential constraints, allowing unsupported beliefs repeatedly stored and propagated across decision steps, leading to systematic behavioral drift in long-horizon agentic systems. Most existing strategies rely on the consensus mechanism, conflating agreement with faithfulness. In this paper, inspired by the vulnerability of unfaithful intermediate reasoning trajectories, we propose Self-Audited Verified Reasoning (SAV E R), a novel framework that enforces verification over internal belief states within the agent before action commitment, achieving faithful reasoning. Concretely, we structurally generate persona-based diverse candidate beliefs for selection under a faithfulnessrelevant structure space. To achieve reasoning faithfulness, we perform adversarial auditing to localize violations and repair through constraint-guided minimal interventions under verifiable acceptance criteria. Extensive experiments on six benchmark datasets demonstrate that our approach consistently improves reasoning faithfulness while preserving competitive end-task performance.

1

Introduction

Large Language Models (LLMs) are increasingly deployed as autonomous agents that plan, reason, and act over extended horizons. Beyond generating answers, LLM agents maintain internal reasoning trajectories for guiding tool invocation, action commitment, and memory updates across decision steps. With the widespread adoption of reasoning paradigms, such as Chain-of-Thought (CoT) (Wei et al., 2022), trajectories are generally regarded as interpretable representations of the agent’s internal * Corresponding author.

Example: Circular Reasoning in Agent Reasoning Q: What science fantasy young adult series, told in first person, has a set of companion books narrating the stories of enslaved worlds and alien species? Gold A: Animorphs.

Reasoning Step 1 (Cue extraction). The question highlights several clues: (i) a science fantasy young adult series, (ii) told in first person, and (iii) it has companion books that narrate stories about enslaved worlds and alien species. Step 2 (Pattern association). The phrase”companion books narrating the stories of enslaved worlds and alien species” reminds me of The Hork-Bajir Chronicles, which features the Hork-Bajir species and themes of enslavement. Step 3 (Membership assertion). The Hork-Bajir Chronicles is a companion book (side story / spinoff) that belongs to the Animorphs franchise. Step 4 (Backward identification). If the companion book referenced by the question is The HorkBajir Chronicles, then the target series that “has” this companion book must be Animorphs. Step 5 (Conclusion). Therefore, the answer is Animorphs.

Generation Pred A: Animorphs.

Figure 1: Demonstration of unfaithful agent reasoning. The agent outputs the correct answer ‘Animorphs’, but its multi-step reasoning process is logically invalid, as an unverified intermediate assumption “The phrase . . . reminds me of The Hork-Bajir Chronicles" is used to derive the conclusion that it already presupposes. This failure mode differs fundamentally from unfaithful CoT, where the reasoning is merely an explanatory artifact, while unfaithful reasoning in the agent determines the following behavior and final decision.

state. However, coherent reasoning traces are fragile for decision-making (Lam et al., 2025). Agents may generate seemingly fluent and structured reasoning, yet violate logical or evidential constraints, reflecting a lack of faithful reasoning (Zhao et al., 2025; Xu et al., 2025b). Such violations are difficult to diagnose from final-task success alone, since correct outcomes can arise from chance, redundancy, or downstream correction, masking the underlying reasoning failure (Chang et al., 2025; Kim et al., 2024), as shown in Figure 1. Unlike singleturn Question Answering (QA), where reasoning can be post hoc and disposable, the agent’s rea-

soning outputs are repeatedly used, amplified, and written into memory (An et al., 2025; Jiang et al., 2025; Tang et al., 2025). Consequently, unfaithful belief states (e.g., unsupported inferences or hidden assumptions) can propagate, bias decisions, and trigger costly actions in closed-loop agent systems (Chakraborty et al., 2025). The risk is not merely incorrect answers, but systematic behavioral drift driven by unfaithful internal beliefs. In agentic systems, existing methods have been adopted to manage uncertainty before committing internal reasoning states to actions, such as selfconsistency (Wan et al., 2025; Xie et al., 2024), multi-agent debate (Liang et al., 2026, 2024b), which maintain multiple candidate reasoning trajectories and rely on consensus-based aggregation for acted belief determination (Zhang and Xiong, 2025). Nevertheless, they rest on the problematic premise that consensus is faithfulness. In practice, multiple sampled trajectories frequently share the same implicit assumptions or inference templates, resulting in structurally correlated yet unfaithful belief states that are repeatedly selected, further reinforced by majority voting, and committed to memory (Ke et al., 2025). Additionally, most existing methods interact with reasoning at the level of surface text rewriting (Shu et al., 2024), without identifying the logical constraints that the specific reasoning step violates, and verifiable acceptance criteria for committing corrected belief states. These limitations reveal that current LLM agents lack an objective of ensuring reasoning faithfulness before action commitment, raising a key question: How can an LLM agent verify the reasoning faithfulness without relying on final-task accuracy or consensus? To address these challenges, we propose SelfAudited Verified Reasoning (SAV E R), a novel framework for enhancing reasoning faithfulness in LLM agents. Rather than relying on final-task outcomes, SAV E R explicitly models the faithfulness of intermediate reasoning steps. To mitigate correlated failure patterns in belief generation, the agent generates a persona-conditioned coalition to elicit structurally diverse candidate belief states and reduce repeated unfaithful templates. It then selects beliefs in a faithfulness-relevant structure space via a quality-aware diversity kernel and k-DPP sampling, followed by adversarial auditing that localizes violations into auditable diagnostics. Finally, SAV E R introduces a constraint-guided minimal counterfactual repair protocol that edits only local-

ized failure slices under verifiable acceptance criteria, iterating audit–repair until the belief passes all checks before being committed to actions or memory. Our key contributions are as follows: • We reveal the overlooked issue of reasoning faithfulness in LLM agents and identify the challenges of verifying intermediate beliefs before action commitment. • We introduce SAV E R, a novel self-auditing framework that verifies and repairs intermediate reasoning trajectories in agents. • We conduct extensive experiments on multiple public datasets to demonstrate the effectiveness of our approach in improving reasoning faithfulness.

2

Related Work

2.1

Faithful Reasoning in LLMs

LLMs have shown strong performance on reasoning tasks. Existing work has explored the prompting strategies, most notably CoT prompting (Wei et al., 2022; Kojima et al., 2022). Extensions such as program-based reasoning (Chen et al., 2023; Zhang et al., 2024; Jiang et al., 2024; Wang et al., 2024) and least-to-most prompting (Zhou et al., 2023; Arora et al., 2023) further structure these steps and improve task accuracy (Ma et al., 2025). However, improved reasoning performance does not necessarily imply faithful reasoning: empirical studies show that generated chains of thought may rely on spurious signals, and intervening on them has limited impact on model predictions, suggesting that such explanations are post-hoc rationalizations (Lyu et al., 2023; Turpin et al., 2023). Motivated by these concerns, a growing body of work has focused on evaluating reasoning faithfulness in LLMs (Luo et al., 2024; Paul et al., 2024; Luo et al., 2025). Existing approaches include counterfactual interventions on reasoning traces (Paul et al., 2024; Ding et al., 2024; Joshi et al., 2024), causal probing methods (Chi et al., 2024; Roy et al., 2025), and targeted diagnostics for chain-of-thought faithfulness (Li et al., 2025). Beyond evaluation, several methods improve faithful reasoning through verification (Dougrez-Lewis et al., 2025) or self-consistency (Wan et al., 2025). Despite these advances, most studies on faithful reasoning focus on LLMs. In agentic settings, however, reasoning trajectories function as persistent belief states to guide downstream decisions,

Belief Generation

Audit

Belief Selection

Repair

Selected Beliefs

Input X

Final Belief

Candidate Beliefs Best Belief Selection

Single Agent

Internal Reasoning Coalition

Feature Mapping

Stress-testing Strategies

Quality Score

Repaired Beliefs

Missing Assumption Invalid Precondition Unjustified Inference k-DPP Sampling

Circular Reasoning Contradiction

Repaired Trajectory

Overgeneralization

Termination Condition Repair Objective

Loop Belief States

Selected Beliefs

Violation Instance Set

Repair Constraint

Figure 2: An overview of our proposed SAV E R framework. This diagram illustrates the overall closed-loop workflow, highlighting the end-to-end process from belief generation and selection to auditing and iterative repair.

allowing erroneous beliefs to accumulate, propagate, and trigger costly actions over long horizons. LLM-centric methods to faithful reasoning are insufficient for agentic decision-making, underscoring the need for agent-specific mechanisms. 2.2

Faithfulness-Aware Reasoning in Agents

An emerging research perspective frames LLMs as agents that perform multi-step reasoning through planning, tool use, and interaction with external environments (Hong et al., 2025; Chen et al., 2025; Xu et al., 2025a). These frameworks expose explicit reasoning trajectories to support complex decision-making in interactive settings (Yao et al., 2023; Yang et al., 2024). However, explicit reasoning does not guarantee faithfulness: empirical studies demonstrate that agent-generated plans or tool-use rationales may not causally determine outcomes, but instead serve as plausible post-hoc justifications (Barez et al., 2025). Despite these observations, existing approaches for mitigating unfaithful agent reasoning remain post-hoc or outcome-driven. Many methods improve robustness by sampling multiple reasoning trajectories or applying external critics, without explicitly verifying whether intermediate reasoning steps are supported by the agent’s available evidence at the time of decision-making (Liang et al., 2024a; Kostka and Chudziak, 2025). Moreover, numerous approaches operate at surface-level trajectory rewriting or consensus aggregation, which

is insufficient to identify structurally correlated yet unsupported belief states that can repeatedly pass heuristic checks (Fu et al., 2025; Grötschla et al., 2025). Consequently, current agentic frameworks lack a mechanism for auditing and verifying internal belief states before action, allowing unfaithful reasoning to be propagated and stored in memory.

3

Methodology

In this section, we introduce the SAV E R framework for faithful reasoning in agentic systems. The complete workflow is shown in Figure 2. We first formalize reasoning faithfulness in § 3.1, and then describe persona-conditioned belief generation in § 3.2 and structure-aware belief selection in § 3.3. We present our reasoning audit mechanism in § 3.4 and introduce the repair procedure that iteratively corrects unfaithful beliefs in § 3.5. 3.1

Modeling Faithful Reasoning in Agent

We consider an agent that generates a multi-step internal reasoning trajectory and then commits to external actions. While improving interpretability, it also introduces the challenge of reasoning faithfulness: intermediate reasoning steps may not be fully supported by the information available during decision-making, motivating an explicit formulation of reasoning faithfulness. Given an input task x, we model the agent’s internal reasoning process as a sequence of discrete steps τ = (s1 , . . . , sL ), where L denotes the

length of the reasoning trajectory and each step sl represents a local claim, inference, assumption, or intermediate conclusion produced by the agent. To quantify whether a reasoning step is justified under the information available to the agent, we introduce the support function Γ(sl | x, Hl , El ) ∈ [0, 1], where the reasoning history Hl = (s1 , . . . , sl−1 ) and El represents the accessible evidence at step l, including retrieved documents, tool outputs, or environment observations. Thus, we define the trajectory-level unfaithfulness rate as L

U (τ ) =

1X I[Γ(sl | x, Hl , El ) < ϵ], L

(1)

l=1

where I[·] denotes the indicator function and a reasoning step is considered unfaithful if its support score falls below a predefined threshold ϵ. 3.2

Persona-Conditioned Belief Generation

Unfaithful reasoning in agents typically arises from repeatable and structurally correlated failure patterns, making naively sampled reasoning traces unreliable. Rather than improving final answer accuracy by generating more traces, our goal is to promote structural diversity among candidate belief states, so that distinct reasoning modes and failure triggers are explicitly exposed. For the given input x, we instantiate an internal reasoning coalition consisting of M reasoning personas by A = {a1 , a2 , . . . , aM }, which models a single LLM agent’s internal cognitive diversity, where each persona corresponds to a distinct structural reasoning bias, e.g., assumption-first vs. evidence-first. Each persona ai is instantiated via persona-specific instruction constraints and reasoning templates. Let Y = C × R denote the belief space, where C is the claim space and R is the space of reasoning trajectories. Persona ai then produces belief yi = G(x; ai ) ∈ Y. We denote yi = (ci , ri ), where the persona’s final claim or decision ci ∈ C and ri = {si,1 , . . . , si,Li } ∈ R denotes the full reasoning trajectory with Li steps and is treated as a candidate belief state. 3.3

Structure-Aware Belief Selection

To select diverse subsets of belief states, we define a structural feature mapping ϕ : ri 7→ ϕ(ri ) ∈ Rd , designed as a proxy for reasoning faithfulnessrelevant structure. We decompose it as ϕ(ri ) = [g(ri ), p(ri ), v(ri ), s(ri )]⊤ ,

(2)

where granularity features g(ri ) quantify step granularity and potential skipping risk; assumptive features p(ri ) reflect how assumptions are introduced and managed, capturing missing/implicit premises and whether assumptions are properly scoped; verification features v(ri ) measure verification behaviors; structural-type features s(ri ) describe the global organization of reasoning. We introduce a lightweight quality scoring function q(yi ; x), which provides coarse filtering of candidate belief states and removes traces that violate minimal usability constraints (e.g., nonsensical steps or internally inconsistent conclusions). Then, we define a quality-aware diversity kernel matrix I ∈ RM ×M with entries Iij = exp(β q̃i ) exp(β q̃j )κ(ϕ(ri ), ϕ(rj )),

(3)

where β controls the quality weighting strength, q̃i denotes a normalized q(yi ; x), and κ(·, ·) is a structural similarity kernel applied to ϕ(ri ). Given candidate reasoning outputs generated by M personas, the agent selects K belief states for auditing. We adopt a k-Determinantal Point Process (k-DPP) to sample a subset S of size K: P(S) ∝ det(IS ), S ⊆ {1, . . . , M },

(4)

where IS denotes the principal submatrix of I defined in Eq. (3) indexed by S. The determinant det(IS ) favors subsets with structurally complementary belief representations in the induced feature space, thereby encouraging the selection of beliefs that exhibit distinct reasoning patterns. As a result, the sampled set avoids allocating auditing capacity to multiple beliefs that violate reasoning faithfulness in similar ways, and instead increases coverage over diverse unfaithful reasoning modes. 3.4

Adversarial Reasoning Audit

Diversity alone does not guarantee reasoning faithfulness, as beliefs may still contain logically unsupported or unjustified steps. To prevent unfaithful beliefs from being committed to actions, we introduce an adversarial reasoning auditing procedure to examine belief states, identify faithfulness-violating reasoning steps, and produce structured diagnostic signals for subsequent repair. Notably, the auditor interrogates the belief state rather than generating or aggregating alternative answers. We perform reasoning auditing by applying a collection of complementary stress-testing strategies to each selected belief yi , i ∈ S. Given

the reasoning trajectory ri and input x, the auditor operates under an observable context that aggregates stated assumptions, verified intermediate facts, and admissible evidence (e.g., retrieved documents or tool outputs). Each stress-testing strategy audits ri from distinct logical perspectives and produces structured audit evidence following a fixed schema to ensure auditability and comparability across beliefs (detailed in Appendix A.2). According to the auditing outcomes, we categorize the faithfulness violations into a type set: T = {Missing_Assumption, Invalid_Precondition, Unjustified_Inference, Circular_Reasoning, Contradiction, Overgeneralization}, which captures the common failure modes in agentic settings, thereby enabling downstream corrective actions (detailed in Appendix A.1). For each belief trajectory ri , the auditor outputs i a violation instance set V(ri ) = {(ti,j , li,j )}m j=1 , where mi = |V(ri )| denotes the number of violation instances detected in trajectory ri , ti,j ∈ T denotes the faithfulness violation type, and li,j ∈ {1, . . . , Li } indexes the reasoning step at which the violation is triggered, distinguishing between globally unfaithful beliefs and those that fail only at specific steps. The auditing process operationalizes this notion by identifying steps si,l for which Γ(si,l | x, Hi,l , Ei,l ) < ϵ and mapping each instance to a concrete violation type t ∈ T . In this way, the violation instance set V(ri ) can be viewed as a discrete, structured instantiation of support assessment. To represent the belief’s faithfulness failure characteristics, we summarize the auditing outcome as an unfaithfulness profile h(ri ) = [ht (ri )]t∈T , where ht (ri ) counts the number of violations of type t in trajectory ri . 3.5

Constraint-Guided Belief Repair

Auditing alone does not improve reasoning faithfulness, and full regeneration of new reasoning traces will generally break the causal link between critique and correction, and make it hard to guarantee the removal of originally identified failure. Consequently, we adopt a minimal counterfactual intervention principle that only edits the localized failure slices identified, while keeping unaffected steps stable to preserve auditability and prevent unnecessary drift. Specifically, for each violation instance (ti,j , li,j ) ∈ V(ri ), the auditor returns structured evidence and an acceptance criterion in a fixed schema (detailed in Appendix A.2), converting subjective critique into faithfulness con-

straints through explicit acceptance criteria (see Appendix B). Given the audit output V(ri ), we define a repair constraint set Θi = {θi,1 , . . . , θi,mi }, where each constraint θi,j encodes a prescribed correction and an explicit criterion for verifying. Let ri denote the original belief-specific reasoning trajectory to be repaired, and r̃i the repaired trajectory, computed by solving r̃i = arg min Lcons (r; Θi ) + λ∆(r, ri ), r

(5)

where Pmi the constraint violation loss Lcons (r; Θi ) = j=1 1[¬Sat(r, θi,j )] penalizes failures to satisfy the acceptance criteria implied by Θi , Sat(r, θ) encodes a concrete and verifiable condition specifying when a violation is resolved, and the minimal edit cost ∆(r, ri ) measures the deviation between the repaired and original trajectories, enforcing minimal intervention. Correcting one violation can expose additional latent violations. We thereby iterate auditing and repair until no violation instances remain, i.e., V(r̃i ) = ∅, the repaired belief is then committed to action and update memory, preventing the propagation of unfaithful reasoning in long-horizon agentic decision-making. After repairing the audited subset {yi }i∈S , the agent selects a belief for execution by maximizing the quality score q(·; x) while penalizing residual unfaithfulness reflected by h(·): X i⋆ = arg max(q(ỹi ; x) − α wt ht (r̃i )), (6) i∈S

t∈T

where ỹi = (c̃i , r̃i ) denotes the repaired belief, wt is a predefined type-dependent severity weight, and α controls the trade-off between superficial quality and verified reasoning faithfulness.

4

Evaluation

4.1

Experimental Setup

Datasets We evaluate our method on six benchmark datasets with various reasoning settings. Multi-hop QA integrates information from multiple sources and performs multi-step reasoning. We adopt three benchmarks for this task: HotpotQA (Yang et al., 2018), 2WikiMHQA (Ho et al., 2020), and MuSiQue (Trivedi et al., 2022). Evidence-sensitive QA focuses on answering questions or verifying claims where correctness critically depends on whether sufficient and appropriate evidence supports the conclusion, prone to unsupported assumptions and unjustified inference steps.

HotpotQA

2WikiMHQA

MuSiQue

EM ↑

F1 ↑

EM ↑

F1 ↑

EM ↑

F1 ↑

EM ↑

F1 ↑

EM ↑

F1 ↑

EM ↑

LLaMA-3.1-8B

VANILLA LM C OT MAD S ELF -R EFINE B-2 SAV E R(Ours)

34.8 38.3 43.1 40.8 42.9 43.7

43.6 47.5 51.2 48.9 51.3 52.6

39.6 44.2 47.9 46.3 46.7 47.7

47.3 50.7 55.4 53.3 54.4 55.5

23.8 27.1 30.9 28.7 31.0 31.8

34.6 36.8 40.8 37.8 40.6 42.5

29.5 32.7 36.6 33.5 35.9 37.1

38.2 43.4 46.9 43.0 44.9 47.8

29.4 33.2 36.3 34.1 36.7 37.2

37.7 42.3 45.2 43.6 46.2 45.7

53.7 55.9 60.7 57.6 60.6 61.1

LLaMA-3.2-3B

VANILLA LM C OT MAD S ELF -R EFINE B-2 SAV E R(Ours)

30.6 34.4 37.4 34.9 36.0 38.3

39.1 43.6 45.8 44.1 45.7 47.5

31.6 35.0 39.8 36.8 39.9 40.0

39.4 43.6 47.2 44.5 46.9 48.6

11.7 15.2 18.4 17.1 18.3 18.6

20.8 23.4 28.1 26.4 27.5 28.3

25.1 29.5 33.4 31.2 34.4 33.9

34.8 38.6 42.4 39.2 42.9 43.8

24.4 29.2 33.5 29.9 32.1 33.2

34.6 37.8 41.2 39.4 40.7 41.9

49.3 52.4 55.6 53.8 54.3 56.4

Qwen-2.5-7B

VANILLA LM C OT MAD S ELF -R EFINE B-2 SAV E R(Ours)

33.9 38.6 42.5 39.4 41.2 43.1

42.3 46.6 50.9 48.6 50.6 51.2

38.4 42.3 46.9 44.1 47.2 47.7

46.0 49.3 54.8 52.2 55.1 55.8

21.9 26.3 30.8 27.2 29.1 30.6

30.5 34.7 39.3 36.4 39.0 39.4

28.2 33.0 36.2 34.6 35.5 36.8

36.9 41.7 44.6 43.7 45.3 45.9

28.4 32.8 35.3 33.8 35.1 35.6

36.5 42.5 44.8 42.9 43.3 44.1

52.7 56.3 60.1 58.2 60.7 60.9

Models

Methods

NQ

Quoref

FEVER

Table 1: The overall evaluation results of SAV E R and other baseline methods on six benchmarks. The bestperformed method is marked by bold and the runner-up performing method is marked by underline.

We consider Natural Questions (NQ) (Karpukhin et al., 2020) and FEVER (Thorne et al., 2018) in this category. Local reasoning tasks resolve referential dependencies within a single context, serving as a baseline for evaluating reasoning faithfulness under minimal structural uncertainty. We choose Quoref (Dasigi et al., 2019) for evaluation. Baselines We compare our method against stateof-the-art baselines. We adopt Vanilla LM (Brown et al., 2020) as a direct generation baseline to answer questions without explicit reasoning. To elicit step-by-step reasoning, we include CoT (Wei et al., 2022), where the model produces a rationale before answering. We consider deliberation-based inference with Multi-Agent Debate (MAD) (Liang et al., 2024b), which aggregates multiple agents’ discussions to form the final answer. For iterative self-improvement, we adopt Self-Refine (Madaan et al., 2023), where the model alternates between generating and revising based on self-critique. Finally, we include Best-of-2 (B-2) (Papineni et al., 2002) to produce two candidate outputs and select the final answer. Evaluation Metrics We utilize two complementary categories of metrics for evaluation. For Tasklevel Performance, we report Exact Match (EM) and token-level F1, following standard evaluation protocols for QA and verification tasks. For Reasoning Faithfulness, as correct final answers do not necessarily imply faithful reasoning, we additionally evaluate faithfulness at the trajectory

level. Based on the audit results, we compute the following faithfulness metrics: (i) Average Violations (Avg Viol), the mean number of detected faithfulness violations per reasoning trajectory; (ii) Violation-Free Rate (VFR), the proportion of trajectories that contain no detected violations; (iii) Unfaithful Step Rate (USR), the fraction of reasoning steps within a trajectory that are flagged as unfaithful; (iv) Post-Repair Residual (Post-Res) measures the remaining violation rate after the audit–repair procedure is applied. Implementation Details We conduct our experiments on Qwen 2.5-7B (Bai et al., 2025), LLaMA3.1-8B, and LLaMA-3.2-3B (Dubey et al., 2024; Touvron et al., 2023). All models are used in the zero-shot inference setting, without task-specific fine-tuning. We default person number M = 4, select K = 2 candidates for auditing. We define β = 1.0, ϵ = 0.5. The audit-repair process is iterated for at most 10 rounds. Faithfulness evaluation is performed with the same auditing protocol for all methods. Violation statistics are computed based on the final reasoning trajectories produced by each method. We run our models on four NVIDIA RTX 4090 GPU devices. 4.2

Main Results

Table 1 reports the overall performance of SAV E R and baselines across six benchmarks under three backbone models, where SAV E R achieves consistently competitive evaluation results. On multi-

HotpotQA Methods

2WikiMHQA

MuSiQue

Avg Viol ↓ VFR ↑ Post-Res ↓ USR ↓ Avg Viol ↓ VFR ↑ Post-Res ↓ USR ↓ Avg Viol ↓ VFR ↑ Post-Res ↓ USR ↓

VANILLA LM C OT MAD SAV E R(Ours)

2.65 1.98 1.33 0.37

7.43% 24.89% 36.74% 81.36%

– – – 0.05

46.41% 27.36% 23.94% 9.12%

2.83 2.21 1.81 0.56

6.58% 17.41% 32.78% 72.34%

– – – 0.08

53.19% 32.11% 28.82% 13.84%

3.25 2.91 2.16 0.83

5.34% 13.26% 26.17% 69.38%

– – – 0.11

62.63% 37.58% 36.51% 19.73%

Table 2: Reasoning Faithfulness Evaluation on Multi-hop QA Benchmarks under LLaMA-3.1-8B. 2WikiMHQA

USR(%)

MAD SAVeR

40 20

USR(%)

HotpotQA 60

MAD SAVeR

60 40 20

1

2

3 Iteration

4

5

1

2

60

4

5

NQ MAD SAVeR

40

MAD SAVeR

60 USR(%)

USR(%)

MuSiQue

3 Iteration

40 20

20 1

2

3 Iteration

4

5

1

2

20 10 1

2

3 Iteration

4

5

FEVER MAD SAVeR

4

5

MAD SAVeR

40 USR(%)

USR(%)

Quoref 30

3 Iteration

20

1

2

3 Iteration

4

5

Figure 3: Audit-Repair (A-R) dynamics on HotpotQA. For SAV E R, iterations correspond to one audit-repair cycle. For MAD, iterations denote debates to reduce inconsistencies without verifiable acceptance criteria.

hop QA benchmarks (HotpotQA, 2WikiMHQA, and MuSiQue), SAV E R demonstrates clear improvements over standard prompting methods and iterative refinement baselines, indicating its effectiveness in handling multi-step reasoning tasks. On evidence-sensitive and single-hop benchmarks (NQ, Quoref, and FEVER), SAV E R also performs competitively, suggesting the superiority of enforcing reasoning verification on performance. We further observe that the performance gains of SAV E R are stable across different model scales. We demonstrate the reasoning faithfulness evaluation on three multi-hop QA benchmarks as summarized in Table 2. SAV E R consistently achieves substantially lower Avg Viol and USR, along with markedly higher VFR, compared to all baselines across datasets, indicating a significant reduction in unfaithful intermediate reasoning. In contrast, CoT and MAD alleviate unfaithfulness to a limited extent but still retain a large proportion of violationprone steps. Moreover, the low Post-Res values of SAV E R indicate that the audit-repair (A-R) process

effectively resolves detected violations. Figure 3 illustrates the evolution of USR across iterative refinement for SAV E R and MAD on six benchmarks under LLaMA-3.1-8B. Across all datasets, SAV E R exhibits a faster and more stable reduction in USR, consistently converging to substantially lower unfaithfulness levels than MAD within a small number of iterations. In contrast, while MAD gradually reduces USR through successive debate rounds, a considerable fraction of unfaithful steps persists even after multiple iterations, indicating that explicitly auditing and repairing localized reasoning failures is more effective than debate-based refinement in preventing the accumulation of unfaithful intermediate reasoning.

4.3

Case Studies and Discussions

We present a representative case study to illustrate how unfaithful reasoning arises in agentic question answering and how SAV E R mitigates such failures through explicit auditing and repair. As shown in Figure 4, different personas generate diverse belief candidates, including assumption-driven numerical estimation and evidence-first extraction. During auditing, SAV E R localizes these failures to specific reasoning slices. In our case, one belief commits a numerical estimate (“3,500-4,000 → 3,700”) without explicit evidence, which is flagged as an unjustified inference. Another belief identifies the relevant entity correctly but fails to provide a citable sentence linking the arena to its seated capacity, violating required preconditions. Heuristic guessing is replaced with evidence-bound extraction, and missing preconditions are satisfied by explicitly attaching verifiable evidence sentences. The repaired beliefs are then re-audited to ensure all acceptance criteria are met before commitment. As a result, the agent produces a final answer that is fully grounded in cited evidence, preventing the accumulation of unfaithful beliefs in long-horizon reasoning.

Task x Q: The arena where the Lewiston Maineiacs played their home games can seat how many people (seated capacity)? Belief Generation

Belief Selection

Audit

Persona coalition A={a1,a2,a3,a4} outputs beliefs yi=(ci,ri): •y1 (assumption-first): “Similar junior-hockey arenas seat 3,500–4,000 → capacity ≈ 3,700.” •y2 (evidence-first): ”Retrieve the home arena name → extract the exact seated-capacity integer from an evidence sentence → then answer.” •y3 (entity-disambiguation): ”Identify the exact homearena entity X → verify ’Maineiacs ↔ X’ → then extract the capacity value.” •y4 (audit persona): ”Do not commit any number unless the capacity integer is explicitly cited from evidence.”

After k-Determinantal Point Process (k-DPP) sampling, select an audit subset S={y1,y3}: •Select y1 to cover the high-risk “range → integer guess” failure template. •Select y3 to cover the entity-grounded path (but requiring evidence binding).

Audit y1: •Failing slice: ”3,500–4,000 → 3,700”. •Probe: ”Is there an evidence sentence explicitly stating ’3,700 seated capacity’?” → No. •Output: Unjustified Inference.

Re-audit after Repair

Repair

•y1: no heuristic leap; capacity integer is evidencebound → pass. •y3: missing preconditions satisfied with citable evidence → pass.

Repair for y1 (Unjustified Inference): •Remove range→integer guessing; the final capacity integer must be copied from an explicit evidence sentence. •Minimal fix: replace the failing slice with evidence-bound extraction: ”Retrieve the home arena name → locate the arena page/tool output → extract the seated-capacity integer N from an evidence sentence.”

Evidence sentence states: ”Seating capacity: 3,677 (seated).” Commit: Answer = 3,677 seated. Safe memory write: store only the evidence-grounded fact (home arena X → 3,677 seated), not the guessing template.

Audit y3: •Failing slice: Missing required citable capacity evidence. •Probe: provide citable evidence for (1) “home arena = X” and (2) “X seated capacity = N” → Missing •Output: Invalid Precondition.

Repair for y3 (Invalid Precondition): •Minimal completion (no erroneous slice to rewrite): attach the missing evidence by producing two citable sentences: (1) Lewiston Maineiacs’ home arena = X; (2) X seated capacity = N (integer). •Acceptance criteria: both sentences explicitly contain X and the integer N.

Figure 4: A case study on a multi-hop factual query to correct unjustified inference. The agent initially proposes plausible capacity values based on arena similarity and entity identification, but without explicit evidence. SAV E R audits belief candidates, flags unsupported numerical guesses and missing citable capacity statements, and repairs them by enforcing evidence-bound extraction of the exact seated-capacity integer from retrieved sources. After iterative re-auditing, only the verified capacity value is committed to the final answer and written to agent memory. HotpotQA

2WikiMHQA

Methods

EM ↑

F1 ↑

Avg Viol ↓

VFR ↑

Post-Res ↓

USR ↓

EM ↑

F1 ↑

Avg Viol ↓

VFR ↑

Post-Res ↓

USR ↓

w/o Persona w/o k-DPP w/o Auditing w/o Repair SAV E R

43.2 43.3 43.8 44.0 43.7

52.4 52.2 52.8 52.9 52.6

0.49 0.64 1.37 1.56 0.37

74.55% 71.47% 42.65% 33.68% 81.36%

0.06 0.08 – – 0.05

11.97% 15.86% 26.74% 37.63% 9.12%

47.6 47.5 47.8 48.1 47.7

55.3 55.2 55.7 55.6 55.5

0.78 0.86 1.76 1.83 0.56

65.98% 61.78% 38.95% 29.17% 72.34%

0.10 0.14 – – 0.08

19.24% 20.52% 29.17% 39.84% 13.84%

Table 3: Ablation Study on the HotpotQA and 2WikiMHQA Benchmarks under LLaMA-3.1-8B.

4.4

Ablation Studies

In Table 3, we present the ablation study on HotpotQA and 2WikiMHQA, examining the contribution of each component in SAV E R. Removing any module consistently degrades reasoning faithfulness, while having marginal effects on EM and F1, indicating the effectiveness on the intermediate reasoning quality. Specifically, removing persona generation leads to a noticeable increase in Avg Viol and USR, suggesting the significance of structured reasoning diversity for exposing distinct failure modes. Disabling the k-DPP-based belief selection further exacerbates unfaithfulness, highlighting the role of structure-aware diversity in preventing correlated reasoning errors. More severe degradation is observed when auditing or repair is removed: both settings result in substantially higher Avg Viol and USR. These results confirm that audit and constraint-guided repair are essential

for effectively reducing unfaithful reasoning.

5

Conclusion

In this work, we studied the agent reasoning faithfulness, where coherent reasoning can still violate logical or evidential constraints, and such unfaithful beliefs may propagate and accumulate in agentic systems, leading to systematic behavioral drift. We propose SAV E R, a framework that explicitly verifies intermediate belief states before action commitment. SAV E R generates diverse candidate beliefs, selectively inspects them at the trajectory level, and corrects localized reasoning failures under explicit acceptance criteria, enabling the agent to prevent unsupported inferences from being written to memory. Extensive experiments across multiple benchmarks demonstrate that SAV E R substantially improves reasoning faithfulness while maintaining competitive end-task performance.

Limitations This work exhibits several limitations worth noting. Firstly, extra computational overhead is introduced by maintaining multiple candidate belief states and performing iterative A-R cycles. Although SAV E R limits audit to a small, structurally diverse subset, the A-R loop remains more expensive than single-pass prompting or lightweight refinement strategies, particularly in tasks with short reasoning chains. Secondly, strict faithfulness enforcement may be unnecessary in simple scenarios and could introduce redundant reasoning operations. While SAV E R is designed to localize and minimally correct unsupported reasoning steps, it currently lacks an explicit mechanism to adapt verification depth to task difficulty. Future work could explore adaptive auditing policies that condition verification on uncertainty or task complexity, enabling agents to dynamically trade off reasoning faithfulness and efficiency.

Ethical Considerations This work aims to improve the faithfulness of internal reasoning in agents by auditing and repairing unsupported intermediate beliefs before they are committed to actions or memory. All experiments are conducted on publicly available datasets, and no additional collection of personal or sensitive data is involved. All models and data are used in accordance with their intended purposes and licenses. Although SAV E R reduces the risk of propagating unfaithful reasoning, it does not guarantee correctness in high-stakes applications such as medical, legal, or safety-critical settings. The auditing and repair procedures rely on the underlying LLM, and biases present in the base model may still affect verification outcomes. Thus, human oversight remains necessary when deploying.

GenAI Usage Disclosure This work is entirely original and was conducted by the authors. Generative AI tools were not used to produce any content of the work; they were used solely to assist with language refinement and improve clarity and quality of the text.

References Kaikai An, Fangkai Yang, Liqun Li, Junting Lu, Sitao Cheng, Shuzheng Si, Lu Wang, Pu Zhao, Lele Cao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang,

and Baobao Chang. 2025. Thread: A logic-based data organization paradigm for how-to question answering with retrieval augmented generation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 18300– 18319. Association for Computational Linguistics. Aseem Arora, Shabbirhussain Bhaisaheb, Harshit Nigam, Manasi Patwardhan, Lovekesh Vig, and Gautam Shroff. 2023. Adapt and decompose: Efficient generalization of text-to-sql via domain adapted leastto-most prompting. In Proceedings of the 1st GenBench Workshop on (Benchmarking) Generalisation in NLP, pages 25–47. Association for Computational Linguistics. Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, and 8 others. 2025. Qwen2.5-vl technical report. Preprint, arXiv:2502.13923. Fazl Barez, Tung-Yu Wu, Iván Arcuschin, Michael Lan, Vincent Wang, Noah Siegel, Nicolas Collignon, Clement Neo, Isabelle Lee, Alasdair Paren, and 1 others. 2025. Chain-of-thought is not explainability. Preprint, alphaXiv, page v1. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, and 12 others. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc. Neeloy Chakraborty, Melkior Ornik, and Katherine Driggs-Campbell. 2025. Hallucination detection in foundation models for decision-making: A flexible definition and review of the state of the art. ACM Computing Surveys, 57(7):1–35. Ching Chang, Yidan Shi, Defu Cao, Wei Yang, Jeehyun Hwang, Haixin Wang, Jiacheng Pang, Wei Wang, Yan Liu, Wen-Chih Peng, and Tien-Fu Chen. 2025. A survey of reasoning and agentic systems in time series with large language models. arXiv preprint arXiv:2509.11575. Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2023. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. Transactions on Machine Learning Research. Zhaoling Chen, Robert Tang, Gangda Deng, Fang Wu, Jialong Wu, Zhiwei Jiang, Viktor Prasanna, Arman Cohan, and Xingyao Wang. 2025. Locagent: Graphguided llm agents for code localization. In Proceedings of the 63rd Annual Meeting of the Association

for Computational Linguistics (Volume 1: Long Papers), pages 8697–8727. Association for Computational Linguistics. Haoang Chi, He Li, Wenjing Yang, Feng Liu, Long Lan, Xiaoguang Ren, Tongliang Liu, and Bo Han. 2024. Unveiling causal reasoning in large language models: Reality or mirage? In Advances in Neural Information Processing Systems, volume 37, pages 96640–96670. Curran Associates, Inc. Pradeep Dasigi, Nelson F Liu, Ana Marasović, Noah A Smith, and Matt Gardner. 2019. Quoref: A reading comprehension dataset with questions requiring coreferential reasoning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5925–5932, Hong Kong, China. Association for Computational Linguistics. Bowen Ding, Qingkai Min, Shengkun Ma, Yingjie Li, Linyi Yang, and Yue Zhang. 2024. A rationalecentric counterfactual data augmentation method for cross-document event coreference resolution. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 1112–1140. Association for Computational Linguistics. John Dougrez-Lewis, Mahmud Elahi Akhter, Federico Ruggeri, Sebastian Löbbers, Yulan He, and Maria Liakata. 2025. Assessing the reasoning capabilities of llms in the context of evidence-based claim verification. In Findings of the Association for Computational Linguistics: ACL 2025, pages 20604–20628. Association for Computational Linguistics. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, A. Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony S. Hartshorn, Aobo Yang, Archi Mitra, A. Sravankumar, A. Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, and 510 others. 2024. The llama 3 herd of models. Dayuan Fu, Keqing He, Yejie Wang, Wentao Hong, Zhuoma Gongque, Weihao Zeng, Wei Wang, Jingang Wang, Xunliang Cai, and Weiran Xu. 2025. Agentrefine: Enhancing agent generalization through refinement tuning. In The Thirteenth International Conference on Learning Representations. Florian Grötschla, Luis Müller, Jan Tönshoff, Mikhail Galkin, and Bryan Perozzi. 2025. Agentsnet: Coordination and collaborative reasoning in multi-agent llms. arXiv preprint arXiv:2507.08616. Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps. In Proceedings of the 28th International Conference on Computational Linguistics, pages 6609– 6625, Barcelona, Spain (Online). International Committee on Computational Linguistics.

Sirui Hong, Yizhang Lin, Bang Liu, Bangbang Liu, Binhao Wu, Ceyao Zhang, Danyang Li, Jiaqi Chen, Jiayi Zhang, Jinlin Wang, Li Zhang, Lingyao Zhang, Min Yang, Mingchen Zhuge, Taicheng Guo, Tuo Zhou, Wei Tao, Robert Tang, Xiangtao Lu, and 9 others. 2025. Data interpreter: An LLM agent for data science. In Findings of the Association for Computational Linguistics: ACL 2025, pages 19796–19821. Association for Computational Linguistics. Yue Jiang, Qin Chao, Yile Chen, Xiucheng Li, Shuai Liu, and Gao Cong. 2024. UrbanLLM: Autonomous urban activity planning and management with large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 1810–1825. Association for Computational Linguistics. Zhuohang Jiang, Pangjing Wu, Xu Yuan, Wenqi Fan, and Qing Li. 2025. Qa-dragon: Queryaware dynamic rag system for knowledge-intensive visual question answering. arXiv preprint arXiv:2508.05197. Abhinav Joshi, Areeb Ahmad, and Ashutosh Modi. 2024. Cold: Causal reasoning in closed daily activities. In Advances in Neural Information Processing Systems, volume 37, pages 5145–5187. Curran Associates, Inc. Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick SH Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781. Association for Computational Linguistics. Zixuan Ke, Fangkai Jiao, Yifei Ming, Xuan-Phi Nguyen, Austin Xu, Do Xuan Long, Minzhi Li, Chengwei Qin, PeiFeng Wang, silvio savarese, Caiming Xiong, and Shafiq Joty. 2025. A survey of frontiers in LLM reasoning: Inference scaling, learning to reason, and agentic systems. Transactions on Machine Learning Research. Survey Certification. Minsoo Kim, Jongyoon Kim, Jihyuk Kim, and Seungwon Hwang. 2024. QuBE: Question-based belief enhancement for agentic LLM reasoning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21403–21423. Association for Computational Linguistics. Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems, volume 35, pages 22199–22213. Curran Associates, Inc. Adam Kostka and JarosĹ Chudziak. 2025. Towards cognitive synergy in llm-based multi-agent systems: integrating theory of mind and critical evaluation. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 47.

Man Ho Lam, Chaozheng Wang, Jen-tse Huang, and Michael Lyu. 2025. Codecrash: Exposing llm fragility to misleading natural language in code reasoning. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, volume 38. Jiachun Li, Pengfei Cao, Yubo Chen, Jiexin Xu, Huaijun Li, Xiaojian Jiang, Kang Liu, and Jun Zhao. 2025. Towards better chain-of-thought: A reflection on effectiveness and faithfulness. In Findings of the Association for Computational Linguistics: ACL 2025, pages 10747–10765. Association for Computational Linguistics. Dayong Liang, Xiao-Yong Wei, and Changmeng Zheng. 2026. Multi-agent undercover gaming: Hallucination removal via counterfactual test for multimodal reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence. Sirui Liang, Baoli Zhang, Jun Zhao, and Kang Liu. 2024a. ABSEval: An agent-based framework for script evaluation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 12418–12434. Association for Computational Linguistics. Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu. 2024b. Encouraging divergent thinking in large language models through multi-agent debate. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 17889–17904, Miami, Florida, USA. Association for Computational Linguistics. Linhao Luo, Yuan-Fang Li, Gholamreza Haffari, and Shirui Pan. 2024. Reasoning on graphs: Faithful and interpretable large language model reasoning. In The Twelfth International Conference on Learning Representations. Linhao Luo, Zicheng Zhao, Gholamreza Haffari, YuanFang Li, Chen Gong, and Shirui Pan. 2025. Graphconstrained reasoning: Faithful reasoning on knowledge graphs with large language models. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 41540–41565. PMLR. Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch. 2023. Faithful chain-ofthought reasoning. In The 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (IJCNLPAACL 2023). Qianou Ma, Weirui Peng, Chenyang Yang, Hua Shen, Kenneth Koedinger, and Tongshuang Wu. 2025. What should we engineer in prompts? training humans in requirement-driven llm use. ACM Transactions on Computer-Human Interaction, 32(4).

Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023. Self-refine: Iterative refinement with self-feedback. In Advances in Neural Information Processing Systems, volume 36, pages 46534–46594. Curran Associates, Inc. Kishore Papineni, Salim Roukos, Todd Ward, and WeiJing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318, USA. Association for Computational Linguistics. Debjit Paul, Robert West, Antoine Bosselut, and Boi Faltings. 2024. Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 15012– 15032. Association for Computational Linguistics. Amartya Roy, N Devharish, Shreya Ganguly, and Kripabandhu Ghosh. 2025. Causal-LLM: A unified oneshot framework for prompt- and data-driven causal graph discovery. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 8259–8279. Association for Computational Linguistics. Lei Shu, Liangchen Luo, Jayakumar Hoskere, Yun Zhu, Yinxiao Liu, Simon Tong, Jindong Chen, and Lei Meng. 2024. Rewritelm: An instruction-tuned large language model for text rewriting. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 18970–18980. Xiangru Tang, Tianyu Hu, Muyang Ye, Yanjun Shao, Xunjian Yin, Siru Ouyang, Wangchunshu Zhou, Pan Lu, Zhuosheng Zhang, Yilun Zhao, Arman Cohan, and Mark Gerstein. 2025. Chemagent: Self-updating memories in large language models improves chemical reasoning. In The Thirteenth International Conference on Learning Representations. James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. Fever: a large-scale dataset for fact extraction and verification. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819. Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Rodriguez Aurelien, Joulin Armand, Grave Edouard, and Lample Guillaume. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.

Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. Musique: Multihop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10:539–554. Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2023. Language models don’t always say what they think: Unfaithful explanations in chain-ofthought prompting. In Advances in Neural Information Processing Systems, volume 36, pages 74952– 74965. Curran Associates, Inc. Guangya Wan, Yuqi Wu, Jie Chen, and Sheng Li. 2025. Reasoning aware self-consistency: Leveraging reasoning paths for efficient llm sampling. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 3613–3635. Association for Computational Linguistics. Zilong Wang, Hao Zhang, Chun-Liang Li, Julian Martin Eisenschlos, Vincent Perot, Zifeng Wang, Lesly Miculicich, Yasuhisa Fujii, Jingbo Shang, Chen-Yu Lee, and Tomas Pfister. 2024. Chain-of-table: Evolving tables in the reasoning chain for table understanding. In The Twelfth International Conference on Learning Representations. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc. Zhihui Xie, Jizhou Guo, Tong Yu, and Shuai Li. 2024. Calibrating reasoning in language models with internal consistency. In Advances in Neural Information Processing Systems, volume 37, pages 114872– 114901. Curran Associates, Inc. Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. 2025a. A-mem: Agentic memory for LLM agents. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. Yige Xu, Xu Guo, Zhiwei Zeng, and Chunyan Miao. 2025b. SoftCoT: Soft chain-of-thought for efficient reasoning with LLMs. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 23336– 23351. Association for Computational Linguistics. John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024. Swe-agent: Agent-computer interfaces enable automated software engineering. In Advances in Neural Information Processing Systems, volume 37, pages 50528–50652. Curran Associates, Inc.

Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 conference on empirical methods in natural language processing, pages 2369–2380. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations. Shaowei Zhang and Deyi Xiong. 2025. Debate4MATH: Multi-agent debate for fine-grained reasoning in math. In Findings of the Association for Computational Linguistics: ACL 2025, pages 16810–16824. Association for Computational Linguistics. Tianhua Zhang, Jiaxin Ge, Hongyin Luo, Yung-Sung Chuang, Mingye Gao, Yuan Gong, Yoon Kim, Xixin Wu, Helen Meng, and James Glass. 2024. Natural language embedded programs for hybrid language symbolic reasoning. In Findings of the Association for Computational Linguistics: NAACL 2024, pages 4131–4155. Association for Computational Linguistics. Chengshuai Zhao, Zhen Tan, Pingchuan Ma, Dawei Li, Bohan Jiang, Yancheng Wang, Yingzhen Yang, and Huan Liu. 2025. Is chain-of-thought reasoning of llms a mirage? a data distribution lens. arXiv preprint arXiv:2508.01191. Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, and Ed H. Chi. 2023. Least-to-most prompting enables complex reasoning in large language models. In The Eleventh International Conference on Learning Representations.

Record · ID 2649 · SHA-256 3d8608b28c32520b
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.