Conceptio › Archive › arXiv CS
arXiv CSopen access

Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments João Meneses dos Santos1 , Arlindo L. Oliveira1,2 1

Instituto Superior Técnico, Universidade de Lisboa, Portugal 2

INESC-ID, Lisboa, Portugal

arXiv:2609.19128v1 [cs.AI] 16 Sep 2026

Abstract

action unavailable in the current state, or continues navigating despite no measurable progress. These failures are difficult to solve with larger prompts alone because they occur at the interface between language generation and environment transition. This work studies this problem through a dual-process language-agent architecture. SwiftSage (Lin et al., 2024) provides a natural baseline because it separates Swift, a fast System 1-style action proposer, from Sage, a slower System 2style planner whose outputs are executed through an action buffer. This design improves efficiency over methods that query a large model at every timestep, but it leaves two important limitations in long-horizon tasks. First, the agent lacks persistent episodic memory: it cannot selectively reuse salient prior experience across episodes. Second, it does not systematically validate actions immediately before execution or intervene in a bounded way when behavior stagnates. We address these limitations with two modular extensions. The Adaptive Memory Module (AMM) adds salience-gated episodic writing and triggerdriven retrieval. It records compact episodes after informative transitions, such as success, positive progress, near misses, or invalid-action failures, and retrieves relevant memories only at selected recovery and planning points. The Self-Reflection Module (SRM) adds bounded execution-time control. It validates actions immediately before they reach the environment, monitors post-step trajectory signals for stagnation, and invokes a constrained Critic only when corrective intervention is justified. The modules occupy different causal interfaces. AMM is informational: it changes what the controller can remember and reuse, but it never chooses the final action. SRM is control-oriented: it changes what is allowed to reach the environment and when corrective actions are inserted. This separation makes the resulting ablation study in-

Language agents remain brittle in interactive environments, where success requires longhorizon state tracking, valid action execution, and recovery from failed steps. We extend SwiftSage, a dual-process agent that combines a fast action proposer with a slower planner, using two modular cognitive extensions: an Adaptive Memory Module (AMM) for saliencegated episodic storage and trigger-driven retrieval, and a Self-Reflection Module (SRM) for bounded execution-time validation and corrective intervention. Both modules are implemented as feature-flagged extensions over the same execution substrate, enabling controlled ablations on ScienceWorld. Across four configurations—baseline, baseline+AMM, baseline+SRM, and the full system—the full system achieves the best mean final score (64.62), success rate (43.17%), and successfulstep efficiency (19.33 steps), while SRM is the strongest standalone contributor. The results suggest that execution-time control is the dominant bottleneck in this setting, while episodic memory becomes most useful once the runtime loop is stabilized.

1

Introduction

Large language models can generate fluent text and solve many short-horizon reasoning problems, but these abilities do not automatically yield reliable agentic behavior. Interactive environments require an agent to maintain state over time, choose actions that are valid in the current world, decompose goals into executable subgoals, recover from unexpected observations, and avoid locally plausible but unproductive loops. The central problem is therefore not only what an agent knows, but how it coordinates fast action proposal, slower deliberation, memory, and execution-time monitoring. A common failure pattern in such environments is that a reasonable high-level plan degrades at execution time: the agent repeats a stale action, proposes an 1

terpretable as cooperation between evidence and control rather than as an opaque redesign of the baseline agent. We evaluate the system on ScienceWorld (Wang et al., 2022), an interactive text benchmark requiring agents to perform elementary science tasks through grounded sequential actions. The evaluation compares four configurations under a shared runtime substrate: baseline, baseline+AMM, baseline+SRM, and the full system. The results show that the full system achieves the best aggregate performance, while SRM is the strongest standalone extension. AMM alone yields smaller gains, but its contribution is more coherent once SRM stabilizes execution. Overall, the findings support a precise claim: in this setting, execution-time control is the dominant lever for improving interactive language agents, while episodic memory is most useful when inserted as bounded evidence into an already controlled runtime loop.

2

2023). These approaches show that feedback can improve LLM behavior, but they also motivate boundedness: unconstrained self-correction can increase cost or degrade performance when feedback is unreliable (Huang et al., 2023). SRM follows this lesson by making reflection trigger-gated and execution-facing rather than always on. AMM is motivated by Complementary Learning Systems, where rapid episodic learning and slower consolidation play distinct roles (Marr, 1971; Mcclelland et al., 1995), as well as by memoryaugmented generation and cognitive-agent memory systems. Retrieval-augmented generation conditions model outputs on external information (Lewis et al., 2020), while MemGPT treats the context window as a scarce resource and explicitly moves information between transient context and persistent storage (Packer et al., 2023). CoALA similarly frames language agents in terms of memory, actions, and decision procedures (Sumers et al., 2023), and long-horizon language-agent systems such as Generative Agents show how memory, reflection, and planning can support coherent behavior over extended interactions (Park et al., 2023). AMM adapts these ideas to ScienceWorld by storing compact episodic records and retrieving them only under operational triggers, instead of using memory as continuous prompt expansion or as an alternative planner. The closest architectural predecessor is SwiftSage (Lin et al., 2024). It already instantiates fast and slow thinking in ScienceWorld: Swift proposes local actions using an efficient model, while Sage performs higher-level planning and grounding through an action buffer. ScienceWorld is especially suitable for this comparison because it evaluates whether science knowledge can be transformed into valid procedures, not merely whether a model can state the right answer. Our contribution is to preserve the SwiftSage substrate while adding two missing mechanisms: persistent episodic reuse and just-in-time execution control. This framing also separates our work from approaches that simply add more reasoning calls. We ask whether memory and reflection improve an already dual-process controller when inserted at bounded, causally interpretable interfaces.

Motivation and Related Work

Dual-process theory distinguishes fast, automatic cognition from slower deliberative reasoning (Kahneman, 2011). In language-agent design, this distinction is useful as an engineering abstraction rather than a claim of cognitive equivalence: fast pathways support cheap local action proposal, while slower pathways support planning, verification, and recovery. Chain-of-thought prompting and related methods make deliberation more explicit in static reasoning tasks (Wei et al., 2022; Press et al., 2023; Khot et al., 2023; Zhou et al., 2023; Wang et al., 2023; Zhang et al., 2023), but interactive agents additionally need to decide when to deliberate, how to ground plans in current state, and how to prevent invalid actions from consuming environment steps. A second relevant line of work augments LLMs with actions, tools, and feedback, including affordance-grounded action selection in SayCan (Ahn et al., 2022). ReAct interleaves reasoning traces with environment-facing actions (Yao et al., 2023), Reflexion uses verbal feedback from prior attempts (Shinn et al., 2024), Self-Refine iteratively revises model outputs (Madaan et al., 2024), and CRITIC verifies and corrects outputs using external tools (Gou et al., 2024). Toolformer and ART further show that tool calls can be learned or orchestrated in multi-step reasoning pipelines (Schick et al., 2023; Paranjape et al.,

3

Method

The proposed system preserves the SwiftSage control loop and inserts AMM and SRM at precise 2

runtime interfaces. At each timestep, the controller executes the next buffered action when available; otherwise it queries Swift, escalating to Sage under baseline conditions such as invalid actions, noprogress behavior, or the need for deliberate planning. AMM can augment selected Swift, Sage, and Critic prompts with retrieved memories. SRM validates actions before execution and can inject bounded corrective actions into the same buffer used by Sage.

memories as hints rather than authority: current observations, inventory, admissible actions, and runtime constraints always dominate. Retrieved memories are deduplicated, filtered by operational type, and truncated using a fixed compression order before injection. If retrieval, formatting, or prompt-structure checks fail, execution falls back to the unmodified baseline prompt. AMM therefore changes the evidence available to the controller without changing the final execution channel.

3.1

3.2

Adaptive Memory Module

AMM addresses the absence of persistent experience reuse. It has two hook families: a post-step write hook and pre-decision retrieval hooks. The write hook runs after the environment returns an observation and score transition. It receives the task, previous state, executed action, resulting observation, score change, and recent history, then builds a candidate episodic record. This record is stored only if a salience gate detects an informative transition, such as terminal success, positive score change, near-miss progress, explicit invalid-action feedback, or an avoidance-worthy failure. Stored memories are compact semi-structured records, not raw transcripts. Each memory includes fields such as task, local state, recent context, executed action, resulting observation, score transition, and a type tag. The tag distinguishes success, nearmiss, and avoidance-oriented memories. This representation improves retrieval targeting and makes prompt injection safer, because retrieved evidence is already concise, typed, and easy to filter or truncate. Retrieval is trigger-driven. T1 corresponds to Swift failure: when Swift fails to produce a valid action, AMM retrieves related episodes and retries Swift with memory-conditioned context. T4 corresponds to System 2 planning: when the baseline invokes Sage and memory planning is enabled, AMM retrieves success and near-miss episodes for deliberative planning. T2 and T3 are implemented as stagnation and repeated-invalid-action retrieval/caching hooks, but in the evaluated AMM-only configuration they do not directly modify model inputs. This conservative design avoids injecting weakly grounded negative evidence into the fast pathway. Prompt augmentation is bounded and explicitly delimited. Swift receives only a small memory block in recovery mode, while Sage may receive a slightly larger block because it is already the deliberative pathway. Prompts instruct the model to treat

Self-Reflection Module

SRM targets execution-time failures: actions may be plausible in language but invalid, stale, redundant, or ineffective in the current environment. It consists of Gate–1 validation, post-step stagnation detection, and bounded Critic intervention. Gate–1 is applied immediately before an action reaches the environment. It receives the proposed action, current valid-action set, recent state descriptors, and runtime constraints. It normalizes the action, checks admissibility, applies deterministic repair when the mismatch is minor and safe, and drops the action if it remains invalid or violates constraints. Gate–1 is source-agnostic: Swift actions, Sage-buffered actions, and Critic-generated actions all pass through the same pre-execution control surface. After each executed step, SRM updates deterministic diagnostics over the recent trajectory. These diagnostics track repeated observations, repeated actions with no effect, invalid-action loops, excessive navigation without state diversity, and verblevel loops that do not change the effective state. When these signals cross a threshold, SRM emits a structured stagnation report summarizing the failure pattern, actions to avoid, and relevant runtime constraints. The Critic is invoked only when stagnation is detected and safeguards permit a call. Its prompt contains the task, current state, recent trajectory, stagnation report, runtime constraints, and, in the full system, optional AMM evidence. The Critic must output a short executable action list. Its output is parsed, filtered, and inserted into the same FIFO buffer used by Sage; it is never executed directly. Buffered Critic actions still pass through Gate–1. Calls are bounded by budgets and cooldowns, making SRM a controlled repair mechanism rather than an always-on second planner. This is important because the Critic is powerful but risky: if used too often, it can interrupt otherwise valid trajectories 3

Figure 1: Full-system architecture. SwiftSage remains the action-selection substrate; AMM adds salience-gated episodic writing and trigger-driven retrieval; SRM adds Gate–1 validation, stagnation detection, and bounded Critic correction before actions reach ScienceWorld.

or replace local action selection with unnecessary deliberation. SRM therefore treats the Critic as a last-mile repair mechanism, while Gate–1 provides the always-on safeguard. 3.3

retrieved memories provide optional evidence; and Gate–1 remains the final safeguard. This design enables a clean comparison between baseline, baseline+AMM, baseline+SRM, and full system. It also makes the composition interpretable: AMM should help when the agent needs relevant prior evidence, especially in recovery and planning regimes; SRM should help when the main failure is execution validity or local stagnation. The full system tests whether these mechanisms are complementary or redundant.

Full-System Composition

The full system composes AMM and SRM without changing the shared execution substrate. AMM writes salient memories after environment steps and retrieves bounded evidence under its trigger regimes. SRM applies Gate–1 before execution and monitors stagnation after execution. Their main interaction occurs during reflective correction: when SRM invokes the Critic, AMM may provide retrieved episodic evidence as supporting context. The hierarchy remains explicit: current valid actions and constraints dominate; stagnation diagnosis identifies the local failure mode;

4

Experimental Setup

We evaluate on ScienceWorld (Wang et al., 2022), using tasks 0–29 and up to 10 natural-language variations per task, for up to 271 evaluation episodes per configuration. Each episode runs until com4

Table 1: Aggregate performance across configurations. Final is task-level macro-averaged final score; Success is macro-averaged task success; Steps@Succ is computed only over successful episodes. Configuration Final

Succ.

Steps

Baseline +AMM +SRM Full

23.99% 26.94% 41.33% 43.17%

24.48 24.03 19.76 19.33

51.67 53.83 64.33 64.62

5

Results and Analysis

Table 1 shows that the full system obtains the best aggregate performance, with a mean final score of 64.62, success rate of 43.17%, and Steps@Succ of 19.33. Baseline+SRM is close behind in score, reaching 64.33, while baseline+AMM reaches 53.83 compared with 51.67 for the baseline. Relative to the baseline, the full system improves mean final score by 25.1%, SRM alone by 24.5%, and AMM alone by 4.1%. The central pattern is therefore that SRM accounts for most of the standalone improvement, while AMM provides a smaller positive aggregate effect. The grouped results in Table 2 refine this picture. On Short tasks, baseline+SRM performs best, suggesting that just-in-time execution filtering is especially valuable when trajectories are short and mistakes consume a large fraction of the budget. On Medium and Long tasks, the full system obtains the highest grouped score, although on Long tasks it is nearly tied with SRM. This indicates that memory is most useful in the harder regimes where prior episodes can support recovery or planning, but only after execution has been stabilized. Success and efficiency show the same trend. Success increases from 23.99% in the baseline to 26.94% with AMM, 41.33% with SRM, and 43.17% in the full system. Successful trajectories also become shorter in SRM-enabled settings: Steps@Succ falls from 24.48 in the baseline to 19.76 with SRM and 19.33 in the full system. Thus, the strongest configurations do not merely collect more partial reward; they complete more episodes and require fewer steps when they succeed. Mechanism-level logs explain why. AMM is active in memory-enabled configurations, writing 4.48 memories per episode in AMM-only and 4.89 in the full system. Most writes are near-miss records, meaning the memory store primarily captures partial progress rather than only terminal successes. However, memory is used less frequently once SRM is enabled: T1 triggers fall from 17.13 to 12.15, T4 triggers from 19.13 to 11.77, Swift injections from 30.81 to 23.03, and Sage injections from 19.13 to 11.77. The full system therefore writes slightly more experience but needs fewer memory-conditioned recovery attempts, consistent with SRM reducing the number of unstable trajectories. SRM shows the complementary profile. SRMonly records 9.77 stagnation reports and 1.35 effec-

pletion or a fixed step budget. All configurations use the same environment interface, valid-action grounding, one env.step per timestep, Swift/Sage arbitration, and action-buffer execution. When SRM is enabled, actions additionally pass through Gate–1 and stagnation-triggered Critic actions may be inserted into the buffer. When AMM is enabled, memory writing and trigger-driven retrieval are active. AMM-enabled configurations are evaluated with a populated memory agent, because the module is designed to reuse prior episodes. The earlier weaker-memory protocol is treated as a sensitivity condition rather than the main architecture-level comparison. Protocol note. We do not claim a new ScienceWorld SOTA; our focus is a controlled ablation over a reproduced SwiftSage-style substrate under a shared local-model/runtime setting. Differences from Lin et al. (Lin et al., 2024) include the local Qwen2.5-7B-Instruct-1M Sage/Critic runtime and the populated-memory protocol used for AMMenabled configurations. We report three primary metrics: final score at termination, task success, and Steps@Succ, the mean number of steps conditioned on successful completion. To avoid overweighting tasks with more variations, scores are first aggregated at the task level and then macro-averaged. We also group tasks by oracle trajectory length: Short (0 < ∗Len ≤ 20), Medium (20 < ∗Len ≤ 50), and Long (∗Len > 50). Mechanism-level logs track AMM writes and injections, SRM drops and Critic calls, and controller-usage proxies such as Swiftexecuted steps, buffer-executed steps, and System 2 call rates. These logs are essential because final scores alone cannot distinguish whether improvement comes from better memory-conditioned planning, fewer invalid actions, shorter successful trajectories, or reduced reliance on expensive System 2 calls. 5

Table 2: Grouped ScienceWorld results. Values are mean final score. Parentheses show relative change over the baseline within each group. Group

∗Len

Baseline

+AMM

+SRM

Full

Short Medium Long

11.76 28.58 94.30

57.74 45.20 50.93

66.43 (+15.1%) 47.16 (+4.3%) 47.78 (-6.6%)

81.50 (+41.2%) 48.42 (+7.1%) 61.64 (+21.0%)

80.06 (+38.7%) 49.72 (+10.0%) 61.69 (+21.1%)

Overall

49.26

51.67

53.83 (+4.1%)

64.33 (+24.5%)

64.62 (+25.1%)

tive Critic calls per episode; the full system records 9.35 and 1.29. Gate–1 drops are much more frequent: 20.03 per episode in SRM-only and 17.48 in the full system, mostly because proposed actions are not in the current valid-action set. This implies that SRM’s main contribution is not frequent reflective replanning, but continuous filtering of invalid or stale actions before they consume environment steps. Controller usage further supports this interpretation. Direct Swift-executed steps increase from 44.08% in the baseline to 49.16% with AMM, 57.13% with SRM, and 62.01% in the full system. Buffer-executed steps fall from 55.92% to 37.99%, and System 2 call rates fall from 24.6 per episode in the baseline to 14.1 in the full system. Better performance therefore does not come from more deliberation. It comes from making execution more reliable, so fewer invalid actions reach the environment and fewer costly System 2 interventions are required. Sensitivity analyses reinforce this conclusion. Evaluating AMM with populated memory improves score from 51.47 in a weaker-memory regime to 53.83. Increasing the SRM Critic budget from three to six calls worsens performance, reducing score from 64.33 to 52.95. Disabling Swiftlevel T1 memory injection in the full system lowers score from 64.62 to 63.33, with the strongest negative effect on Long tasks. These results suggest that memory is useful but should remain targeted, and that reflection must be bounded rather than expanded indiscriminately.

6

immediately before environment interaction, where invalid or stale actions would otherwise consume steps. In contrast, AMM changes the information available to the model but cannot itself ensure that a proposed action is admissible or useful under the current state. Second, memory is useful but conditional. AMM alone improves the aggregate score only modestly, despite frequent writing and retrieval. This does not imply that episodic memory is irrelevant; rather, it suggests that memory is not the first bottleneck when execution remains unstable. The full system writes slightly more memories but uses memory-conditioned recovery and planning less often than AMM-only. This pattern suggests that SRM reduces the number of unstable states requiring memory intervention, allowing retrieved episodes to be used more selectively. Third, more reflection is not automatically better. The Critic-budget sensitivity analysis shows that expanding reflective intervention can degrade performance. This supports the design choice of bounded reflection: Gate–1 should remain the cheap alwayson safeguard, while the Critic should be reserved for trajectories that exhibit explicit stagnation. The controller-usage statistics point in the same direction. Stronger configurations use fewer System 2 calls and fewer buffer-executed steps, not more. They win by making the fast path safer and the slow path more targeted. These findings refine the cognitive motivation behind the architecture. The contribution is not that memory and reflection should be added everywhere, but that each mechanism should be placed at a causal interface that matches its role. AMM supplies evidence for recovery and planning. SRM governs execution and corrective intervention. SwiftSage remains the shared fast/slow substrate. This separation is what makes the ablation interpretable and what prevents the full system from becoming a monolithic prompt expansion.

Discussion

The results support three claims about modular dual-process agents. First, execution-time control is the dominant bottleneck in this ScienceWorld setting. Baseline+SRM nearly matches the full system in final score and produces most of the success-rate and efficiency gains. This is consistent with SRM’s position in the runtime loop: it acts 6

7

Conclusion

ences into generalized procedural knowledge. A second limitation concerns SRM. The control layer is intentionally interpretable and rulebounded, but several choices remain hand-crafted, including stagnation thresholds, action-filtering rules, focus-related constraints, fixed Critic budgets, and deterministic repair policies. Gate–1 is effective as a conservative safeguard, but its repairs are shallow when an invalid action cannot be cleanly normalized or mapped to an admissible alternative. Learned verifiers, confidence-aware escalation, richer affordance models, and adaptive reflective budgets may improve this component while preserving the boundedness that proved important in the current experiments. A third limitation is evaluation scope. The experiments are restricted to ScienceWorld, a text-based simulated science environment, and to the implementation choices inherited from the underlying SwiftSage-style runtime. Although ScienceWorld is well aligned with the target failure modes, the results may not transfer directly to other interactive environments, multimodal settings, robotic action spaces, or different model backbones. In addition, AMM-enabled configurations are evaluated with a populated memory agent, which is appropriate for testing experience reuse but does not fully characterize cold-start learning, warm-up dynamics, memory aging, compaction, or forgetting. Finally, the analysis remains partly correlational at the prompt level. Mechanism-level logs show when memories are written, retrieved, and injected, and when SRM drops actions or invokes the Critic, but they do not fully isolate every causal promptlevel factor behind individual successes or failures. More fine-grained causal tracing, counterfactual replay, and controlled prompt ablations would strengthen the interpretability of future evaluations.

We introduced two modular cognitive extensions for a SwiftSage-style dual-process language agent: AMM for salience-gated episodic reuse and SRM for bounded execution-time control. Evaluated on ScienceWorld, the full system achieves the best aggregate score, success rate, and successful-step efficiency, while SRM is the strongest standalone contributor. AMM alone is beneficial but modest; its clearest role appears in combination with SRM, where memory can support recovery and planning after execution has been stabilized. The broader implication is that cognitively inspired agent extensions are most useful when inserted at the right causal interface. Adding memory is not sufficient if the agent still executes invalid or stale actions, and adding reflection is harmful when it is unbounded. In this setting, robust interactive behavior emerges from a disciplined composition: memory provides evidence, reflection controls execution, and the baseline fast/slow controller remains the shared substrate for interpretable ablation.

8

Limitations

A first limitation concerns the memory subsystem. AMM currently relies on compact semi-structured episodic records stored in an external memory agent and consumed through bounded prompt injection. This design is appropriate for controlled integration and inspection, but it likely constrains the standalone value of memory: stored traces are useful as situated evidence, yet their representation, retrieval ranking, and conversion into concrete action improvements remain relatively simple. Future work should study richer episodic formats, stronger retrieval and reranking, consolidation across related traces, and more targeted prompt-conditioning strategies for Swift, Sage, and the Critic. A related limitation is that AMM remains primarily episodic. The system does not yet implement a mature semantic consolidation layer that abstracts across episodes into reusable skills, procedures, or task-general regularities. From a Complementary Learning Systems perspective, the current implementation covers the rapid episodic side more than the slower semantic side. This helps explain why AMM alone is beneficial but modest: the memory store can recall concrete past experiences, but it does not yet reliably transform repeated experi-

9

Ethical Considerations

This work studies agents in a simulated educational science environment and does not involve human subjects or private user data. The main risks are indirect: techniques that improve autonomous action selection, recovery from failure, and long-horizon persistence could be transferred to less controlled settings. For that reason, the proposed modules emphasize bounded intervention, explicit action validation, logging, and environment-grounded constraints rather than unconstrained autonomous planning. The system should not be interpreted as safe 7

for deployment in open-ended real-world environments without additional oversight, safety evaluation, and domain-specific constraints.

Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders, and Joseph E. Gonzalez. 2023. MemGPT: Towards LLMs as operating systems. arXiv preprint arXiv:2310.08560. Bhargavi Paranjape, Scott Lundberg, Sameer Singh, Hannaneh Hajishirzi, Luke Zettlemoyer, and Marco Tulio Ribeiro. 2023. ART: Automatic multistep reasoning and tool-use for large language models. arXiv preprint arXiv:2303.09014.

References Michael Ahn, Anthony Brohan, Natalie Brown, Yevgen Chebotar, Oscar Cortes, and 1 others. 2022. Do as i can, not as i say: Grounding language in robotic affordances. In Conference on Robot Learning.

Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, pages 1–22.

Zonglin Gou, Zongqi Shao, Yuxuan Gong, Yifan Shen, and 1 others. 2024. CRITIC: Large language models can self-correct with tool-interactive critiquing. In International Conference on Learning Representations.

Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah Smith, and Mike Lewis. 2023. Measuring and narrowing the compositionality gap in language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 5687–5711. Association for Computational Linguistics.

Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2023. Large language models cannot self-correct reasoning yet. arXiv preprint arXiv:2310.01798. Daniel Kahneman. 2011. Thinking, fast and slow. New York: Farrar, Straus and Giroux.

Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36:68539–68551.

Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2023. Decomposed prompting: A modular approach for solving complex tasks. In The Eleventh International Conference on Learning Representations.

Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36.

Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, and 1 others. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33:9459– 9474.

Theodore Sumers, Shunyu Yao, Karthik R Narasimhan, and Thomas L Griffiths. 2023. Cognitive architectures for language agents. Transactions on Machine Learning Research.

Bill Yuchen Lin, Yicheng Fu, Karina Yang, Faeze Brahman, Shiyu Huang, Chandra Bhagavatula, Prithviraj Ammanabrolu, Yejin Choi, and Xiang Ren. 2024. SwiftSage: A generative agent with fast and slow thinking for complex interactive tasks. Advances in Neural Information Processing Systems, 36.

Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté, and Prithviraj Ammanabrolu. 2022. Scienceworld: Is your agent smarter than a 5th grader? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11279–11298.

Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, and 1 others. 2024. Self-Refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36.

Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.

David C. Marr. 1971. Simple memory: a theory for archicortex. Philosophical transactions of the Royal Society of London. Series B, Biological sciences.

Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824– 24837.

James Mcclelland, Bruce Mcnaughton, and Randall O’Reilly. 1995. Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory. Psychological review, 102:419–57.

Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing reasoning and acting in language

8

models. In: Proceedings of the 11th International Conference on Learning Representations. Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2023. Automatic chain of thought prompting in large language models. In The Eleventh International Conference on Learning Representations. Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, and Ed H. Chi. 2023. Least-to-most prompting enables complex reasoning in large language models. In The Eleventh International Conference on Learning Representations.

9

A

Reproducibility Notes

The main paper reports the architecture-level findings. This appendix provides implementation details, complete tables, and sensitivity checks needed to interpret the reported results. The evaluation uses ScienceWorld tasks 0–29 with up to ten natural-language variations per task, yielding up to 271 evaluation episodes per configuration. All configurations share the same ScienceWorld interface, valid-action grounding, one-env.step-per-timestep execution semantics, Swift/Sage arbitration, and FIFO actionbuffer execution. Reported metrics are computed over evaluation episodes only and are aggregated task-first before macro-averaging across groups.

Memory initialization and leakage control. AMM is evaluated as an experience-reuse mechanism rather than as a cold-start learner. The protocol initializes the memory agent with a warm-up stage of two variations per task, i.e., 60 episodes in total; these warm-up episodes are used only to populate memory and are excluded from all reported evaluation metrics. The main AMM comparison uses a populated memory agent, matching the operating condition of the full system, while the weaker-memory regime is reported separately in Table 8. This distinction is central to interpreting AMM: the results measure retrieval from an initialized episodic store, not learning from an empty memory. During an episode, the online score is determined by ScienceWorld before any post-step AMM write can affect later behavior; post-step writes may contribute to later retrieval but do not retroactively affect the score of the transition that generated them.

Runtime constants and model settings. The executable substrate keeps Swift as the fast action proposer and Sage as the slow planner. Swift follows the SwiftSage-style Flan-T5-large behavior-cloning component, while Sage and the SRM Critic use Qwen2.5-7B-Instruct-1M served locally through vLLM. AMM stores salient episodes in Letta archival memory. The active memory-consumption caps are three retrieved episodes for Swift T1 recovery and five retrieved episodes for Sage T4 planning; Critic outputs are parsed as short corrective action lists of up to five actions and are inserted into the same FIFO buffer as Sage plans. The main SRM setting uses a maximum Critic budget of three calls per episode; the six-call setting is reported only as a sensitivity variant in Table 8.

Aggregation and table organization. Tables 3 and 4 are task-level macro-averages: variations are averaged within each task, and task values are then averaged within oracle-length groups. Tables 5 and 6 are task-balanced per-episode mechanism summaries. Table 7 reports configuration-level usage proxies, with System 2 calls corresponding to Sage calls in Baseline/AMM and to Sage+Critic calls in SRM/Full. The appendix is organized in the same order as the main analysis: full scores, success and efficiency, mechanism activity, usage proxies, and sensitivity analyses. Each block includes a short interpretation paragraph immediately after the corresponding table.

B

Score Results by Task

Table 3 expands the grouped result table from the main paper with the full per-task score breakdown. 10

Table 3: Full ScienceWorld score table. Values are mean final score over variations. ∗Len is the average oracle trajectory length used to define Short, Medium, and Long groups. Parentheses in grouped rows show relative change over the baseline within each group. Task (Type)

∗Len

Baseline

+AMM

+SRM

Full

1-1 (L) 1-2 (L) 1-3 (L) 1-4 (L) 2-1 (M) 2-2 (M) 2-3 (L) 3-1 (S) 3-2 (M) 3-3 (M) 3-4 (M) 4-1 (S) 4-2 (S) 4-3 (S) 4-4 (S) 5-1 (L) 5-2 (L) 6-1 (M) 6-2 (S) 6-3 (M) 7-1 (S) 7-2 (S) 7-3 (S) 8-1 (M) 8-2 (S) 9-1 (L) 9-2 (L) 9-3 (L) 10-1 (L) 10-2 (L)

107.7 78.6 88.9 75.2 21.4 35.2 65.0 13.6 20.8 25.6 29.0 14.6 8.8 12.6 14.6 69.5 79.6 33.6 15.1 23.0 7.0 7.0 8.0 40.0 16.3 97.0 84.9 123.1 130.1 132.1

71.56 42.22 38.44 51.22 84.70 39.80 54.60 48.40 36.40 68.00 64.80 60.80 92.50 45.80 75.80 26.10 24.60 29.75 27.78 11.33 85.00 55.00 63.10 26.80 23.25 59.00 62.00 73.00 41.70 66.70

55.78 61.22 38.44 69.44 75.80 39.70 56.10 64.40 47.00 70.70 70.40 75.00 100.00 59.10 90.00 23.90 23.40 35.75 27.78 10.33 80.00 70.00 74.80 27.60 23.25 54.00 51.50 55.50 50.60 33.50

71.67 58.67 50.33 68.89 74.30 40.00 84.00 71.40 49.00 67.50 73.60 100.00 100.00 92.50 100.00 10.80 70.60 30.50 30.00 10.67 95.00 100.00 93.30 41.80 32.75 66.00 55.00 62.00 66.70 75.00

61.89 47.67 46.00 66.00 72.70 40.20 84.20 75.40 36.80 70.50 70.30 98.30 100.00 90.80 100.00 11.30 85.40 37.00 37.78 28.44 95.00 100.00 93.30 41.80 10.00 68.00 59.00 69.00 75.10 66.70

Short (S) Medium (M) Long (L) Overall

11.76 28.58 94.30 49.26

57.74 45.20 50.93 51.67

66.43 (+15.1%) 47.16 (+4.3%) 47.78 (-6.6%) 53.83 (+4.1%)

81.50 (+41.2%) 48.42 (+7.1%) 61.64 (+21.0%) 64.33 (+24.5%)

80.06 (+38.7%) 49.72 (+10.0%) 61.69 (+21.1%) 64.62 (+25.1%)

Reading note. Improvements are not uniformly distributed. AMM is competitive on selected tasks, but the largest and most consistent gains appear in SRM-enabled runs. The full system is strongest overall because it preserves SRM’s execution-time control while allowing memory to support recovery and planning on longer trajectories.

C

Success and Successful-Step Efficiency

Table 4 reports completion and efficiency metrics using the same oracle-length grouping as the main score analysis. Table 4: Task success and efficiency, macro-averaged within oracle-length groups. Success is the fraction of successful episodes; Steps@Succ is the mean number of steps conditioned on successful completion. Baseline Group Short (S) Medium (M) Long (L) Overall

+AMM

+SRM

Full

Succ.%

Steps

Succ.%

Steps

Succ.%

Steps

Succ.%

Steps

32.96 11.94 25.86 23.99

11.77 20.75 37.37 24.48

46.59 11.94 20.69 26.94

10.78 44.13 39.96 24.03

64.77 11.94 40.52 41.33

8.09 31.00 29.70 19.76

73.90 8.96 39.70 43.17

9.75 26.67 31.89 19.33

Reading note. The full system has the best overall completion rate and the best overall Steps@Succ. Short tasks show the clearest completion gain, while Long tasks show that SRM-enabled systems complete substantially more episodes than the baseline and require fewer steps when they succeed. Medium tasks are less stable: the full system improves final score but not success rate.

D

Mechanism-Level Activity

Tables 5 and 6 report the mechanism-level activity that supports the interpretation of the aggregate results. 11

Table 5: AMM activity by oracle-length group. Values are per-episode task-balanced averages for memory writes, retrieval triggers, and Swift/Sage memory injections. Config.

Group

Writes

Success

Near-miss

Avoid.

T1

T4

Swift inj.

Sage inj.

AMM AMM AMM AMM

Short Medium Long Overall

3.63 4.48 5.18 4.48

1.38 0.86 0.65 0.95

1.79 2.87 3.37 2.71

0.47 0.75 1.16 0.82

18.43 21.58 13.09 17.13

20.12 26.05 13.69 19.13

33.62 40.64 21.91 30.81

20.12 26.05 13.69 19.13

Full Full Full Full

Short Medium Long Overall

3.96 5.55 5.23 4.89

1.96 0.90 0.85 1.14

1.92 3.22 3.74 2.99

0.37 1.44 0.67 0.78

10.91 19.04 8.60 12.15

10.33 16.74 9.41 11.77

21.00 34.64 16.99 23.03

10.33 16.74 9.41 11.77

Reading note. AMM writes are dominated by near-miss memories in both memory-enabled configurations, so the store primarily captures partial progress rather than only terminal successes. Once SRM is enabled, AMM writes slightly more but retrieves and injects less often, suggesting that execution control reduces the number of unstable states requiring memory-conditioned recovery.

Table 6: SRM activity by oracle-length group. Values are per-episode task-balanced averages for stagnation reports, effective Critic calls, main Gate–1 drop reasons, and total Gate DROP counts. Reason-code columns are non-exclusive and therefore do not sum to total drops. Config.

Group

Stagn.

Critic

Not valid

NOOP

Target unseen

Gate DROP

SRM SRM SRM SRM

Short Medium Long Overall

10.47 13.11 6.96 9.77

0.67 2.20 1.36 1.35

7.45 25.48 20.74 17.57

1.02 1.38 2.08 1.54

0.00 3.37 1.56 1.52

13.75 27.00 25.06 20.03

Full Full Full Full

Short Medium Long Overall

8.61 14.40 6.61 9.35

0.72 1.76 1.45 1.29

7.50 22.89 18.44 15.98

1.30 1.49 3.23 2.12

0.10 1.63 1.23 0.97

8.80 24.49 20.05 17.48

Reading note. The dominant SRM mechanism is Gate–1 filtering rather than frequent Critic replanning. Effective Critic calls remain close to one per episode overall, while Gate drops are much larger and are mostly caused by actions absent from the current valid-action set. This supports the claim that execution-time validation is the main source of SRM’s improvement.

E

System Usage and Cost Proxies

Table 7 summarizes controller-usage proxies that complement the score, success, and mechanism-level analyses. Table 7: System usage and cost proxies by configuration. System 2 means Sage calls in Baseline/AMM and Sage+Critic calls in SRM/Full. Config.

Swift %

Buffer %

S1 calls

S2 calls

Baseline AMM SRM Full

44.08 49.16 57.13 62.01

55.92 50.84 42.87 37.99

40.8 60.9 29.4 54.4

24.6 19.1 15.7 14.1

Reading note. The best-performing systems do not improve by increasing deliberation. System 2 calls decrease from the baseline to the full system, while the share of direct Swift-executed steps increases. This is consistent with the interpretation that SRM makes the fast path safer and the slow path more selective.

F

Sensitivity Analyses

Table 8 reports targeted sensitivity analyses. These rows are not additional primary architectures; they test whether the conclusions depend on memory initialization, Critic budget, or Swift-level T1 memory injection. 12

Table 8: Sensitivity analysis of selected implementation choices. Results are grouped final scores. Variant AMM weaker-memory regime AMM populated memory SRM Critic max = 6 SRM Critic max = 3 Full without Swift T1 memory injection Full system

Overall

Short

Medium

Long

Role

51.47 53.83 52.95 64.33 63.33 64.62

60.45 66.43 62.70 81.50 79.23 80.06

45.69 47.16 42.10 48.42 53.88 49.72

47.85 47.78 52.05 61.64 56.36 61.69

Sensitivity to memory initialization. Main AMM comparator. Sensitivity to reflective-intervention budget. Main SRM comparator. Sensitivity to Swift-level AMM recovery. Main full-system comparator.

Reading note. The Critic-budget comparison is the strongest warning against unbounded reflection: increasing the maximum number of Critic calls from three to six lowers performance in every oracle-length group. The memory-initialization comparison shows that AMM benefits from a more mature episodic store, while the T1 ablation indicates that Swift-level memory recovery contributes most clearly on Long tasks.

13

Record · ID 965399 · SHA-256 07ca56a4afa0ef65
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.