2026-09-22
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Peng Xia1,2* , Rujun Han1 , Zifeng Wang1 , Yanfei Chen1 , Yufan Zhang1 , Yoonho Lee3 , Chengsong Huang4 , Han Yu1 , Zhongying CuiZhu1 , Yifei Ming1 , Huaxiu Yao2 , Burak Gokturk1 , Tomas Pfister1 and Chen-Yu Lee1
arXiv:2609.24972v1 [cs.LG] 21 Sep 2026
1
Google Cloud AI Research, 2
UNC-Chapel Hill, 3
Stanford University, 4
Washington University in St. Louis
An LLM agent’s capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution.
github.com/google-research/rrsi
regularized-rsi.com
1. Introduction Modern LLM agents are systems rather than standalone models (Lopopolo, 2026; Rajasekaran, 2026). A frozen backbone model is wrapped in a harness of prompts, control flow, tool interfaces, memory and context management. Agent harness decides whether the same model reads the right file before editing it, recovers from a failed command, manages efficient working context, and writes its findings into the deliverables. Much recent progress in agent products came from harness engineering rather than from new model weights (Karten et al., 2026a; Weng, 2026; Zhang and Khattab, 2026). However, this engineering relies on manual efforts, where humans inspect failed trajectories and tweak the scaffold by hand, so progress is limited by how many trajectories an engineer can read. Recent methods automate this loop by using LLMs to optimize harness components from task feedback (Chen et al., 2026; Karten et al., 2026b; Lee et al., 2026a,b; Lin et al., 2026a; Lou et al., 2026; Nie et al., 2026; Niklaus, 2026; Zhang et al., 2026a,e). Such iterative harness evolution provides a practical form of recursive self-improvement (RSI) (RSI-Exam Team, 2026; Team et al., 2026; Wang et al., 2025; Zhang et al., 2026b) at the agent-system level, where feedback from the current system is used to improve the harness that shapes its subsequent behavior. However, as illustrated in Figure 1 (a), test-time harness evolution repeatedly proposes and selects edits using feedback from a finite evolve set, creating an adaptive overfitting risk: evolve-set performance may improve without corresponding gains on unseen tasks. Recent studies observe substantial gaps between evolution and held-out performance, and show that apparent improvements can arise from task-specific fitting or * This work was done while Peng was a Student Researcher at Google Cloud AI Research. Corresponding author(s): [email protected], {rujunh, chenyulee}@google.com
RRSI : Regularized Recursive Self-Improvement of Agent Harnesses