S KILL A DAM : S TABLE AND E FFICIENT S KILL E VOLUTION FOR AGENTS Gaoyuan Li1 Meihao Fan1 Yizhe Liu1 Shaolei Zhang1∗ Ju Fan1 Siyi Wang2 Jiaheng Hou2 Xudong Weng2 Honghan Tian2 Zang Li2 2 Renmin University of China Tencent {logey04,fmh1art,liuyizhe2004,zhangshaolei98,fanj}@ruc.edu.cn {skylasywang,marvinhou,steveweng,abeltian,gavinzli}@tencent.com
arXiv:2609.08944v1 [cs.AI] 8 Sep 2026
1
A BSTRACT Agent skills provide a lightweight way to equip frozen language-model agents with domain knowledge and procedural guidance, yet obtaining high-quality skills remains costly and difficult to scale. Expert-written skills require substantial human effort. Recent skill self-evolution methods automate an iterative loop that uses execution feedback to revise skills, but their heuristic update strategies often yield unstable optimization and low iteration efficiency. We identify two challenges in realizing stable and efficient skill self-evolution. Direction Stability requires effective corrections to accumulate rather than be overwritten by iteration-local feedback. Update Adaptivity requires the scope of each revision to reflect the consistency of recent case-level improvements. We introduce S KILL A DAM, an Adam-inspired framework for optimizing discrete and non-differentiable skill documents. As a functional analogue of Adam’s first moment, an optimization memory records identified problems and the outcomes of prior solution attempts to stabilize the update direction. As a functional analogue of Adam’s second moment, a volatility-driven edit budget tracks the history-weighted variation of recent case-level improvements and adaptively controls the update magnitude. Across seven benchmarks that span short- and long-horizon tasks, S KILL A DAM achieves state-of-the-art performance with more stable optimization dynamics. It also obtains stronger skills with substantially fewer optimization iterations and lower cost than prior methods. Code repository: https://github.com/ruc-datalab/SkillAdam.
1
I NTRODUCTION
Large language model (LLM) agents have emerged as a powerful paradigm for solving complex real-world tasks through multi-step planning and tool use. Representative tasks include web interaction (Wang et al., 2025a), travel planning (Zhang et al., 2026b), data preparation (Fan et al., 2026; Deng et al., 2026), and data discovery and analysis (Zhang et al., 2025; 2026c; Liu et al., 2026). However, LLM agents often struggle in domain-specific scenarios because their general-purpose capabilities cannot satisfy long-tail domain requirements. For example, planning a multi-city trip requires reasoning about visa regulations, transportation dependencies, and airline-specific booking policies. To address this challenge, recent agent frameworks introduce Agent Skills, defined by Anthropic as modular packages of instructions and supporting resources that equip agents with specialized capabilities (Anthropic, 2025). However, high-quality Skills are both essential and difficult to obtain. SkillsBench shows that carefully curated Skills can substantially improve agent performance (Li et al., 2026), highlighting the importance of high-quality Skills. In contrast, SkillAxe reports that Skills generated directly by LLMs often remain ineffective without further refinement (Gautam et al., 2026), suggesting that automatically constructing high-quality Skills remains a challenging problem. A common solution is to rely on domain experts to manually author Skills based on their expertise, but this process is time-consuming and labor-intensive. Another line of work employs language models to generate ∗
Corresponding author.
1
Initial Skill
Trajectory-Informed Initialization
what to revise? preset rules
Random Initialization
Volatility-Driven Edit Budget
case-level changes
Initial Skill
edit budget
high
Optimization Memory
how much to revise? generate a complete new skill
issue issue issue issue issue
Final Skill
(a) Existing SGD-like methods.
Optimal Skill (Final Skill)
Optimization Direction
Skill Quality
Optimal Skill
low
(b) S KILL A DAM (ours).
Figure 1: Conceptual comparison of iterative skill optimization strategies.
or refine Skills using manually designed prompts or heuristic rules. Although these methods reduce the burden of manual Skill authoring, they still rely heavily on human-designed prompts and task-specific heuristics, limiting their ability to generalize across domains. To reduce human intervention, recent studies have investigated automated Skill construction from interaction experience. AutoManual incrementally updates structured rules and compiles them into an instruction manual (Chen et al., 2024). Agent Skill Induction learns verified programmatic Skills from web interactions (Wang et al., 2025a), while Trace2Skill synthesizes a unified Skill from a diverse pool of execution traces (Ni et al., 2026). More recently, several studies have formulated automated Skill construction as skill self-evolution. Some methods iteratively refine a Skill document (Gautam et al., 2026; Alzubi et al., 2026; Yang et al., 2026), while others evolve Skill packages or maintain Skill repositories (Zhang et al., 2026a; Ouyang et al., 2026). These approaches reduce manual effort by refining Skills over multiple iterations. As illustrated by the iteration-local optimization path in Figure 1(a), however, existing methods still lack reliable control over what to revise and how much to revise in each iteration. To address this limitation, we propose S KILL A DAM, a stable and efficient framework that aims to preserve effective corrections across iterations while adapting each revision to the reliability of recent evidence. Realizing these properties raises two challenges. The first is Direction Stability. Since each iteration observes feedback from only a limited set of cases, successive revisions may focus on different problems and undo one another. For example, one revision may instruct the agent to minimize the total cost by selecting the cheapest feasible itinerary. A later revision may add more attractions to produce a richer travel plan, but these additions can increase the cost and violate the earlier budget constraint. New revisions must therefore remain consistent with effective corrections accumulated in earlier iterations. The second challenge is Update Adaptivity. The appropriate edit scope should depend on how consistently a recent revision affects the evaluated cases. If one revision improves some cases but degrades others, a broad subsequent edit risks overwriting useful guidance and should therefore be constrained. If the gains are consistent across cases, a broader edit is better supported. These challenges resemble those in stochastic optimization, where each update is based on partial evidence and the appropriate step size depends on the scale of recent signals. Adam (Kingma & Ba, 2015) provides two complementary design principles: aggregate historical update signals to stabilize direction, and rescale the effective step size using accumulated signal magnitude. S KIL L A DAM adopts these principles functionally in the discrete skill space rather than applying Adam numerically. As illustrated in Figure 1(b), S KILL A DAM first constructs a trajectory-informed initial skill from execution trajectories and their evaluation feedback. During self-evolution, an Evolving Issue Tracker records identified problems, their current status, and the outcomes of prior solution attempts, allowing each new update to consider both current feedback and accumulated evidence. In parallel, a volatility-driven edit budget measures how unevenly a candidate changes case-level performance and controls the allowable scope of the next modification. The tracker determines what should be revised, while the budget determines how much may be revised. Together with trajectory-informed 2
2
1 Rollout Sample from Dataset
Parameter 𝜃!"#
Mini-batch 𝐵!
1 − 𝛽#
Prediction 𝑜! = 𝑓 𝐵! ; 𝜃!"#
Case 1 Case … Case k
Moment First Moment Estimation 𝑚 = 𝛽 𝑚 + 1 − 𝛽 𝑔 , 𝑚ˆ = 𝑚! ! # !"# # ! ! !
Loss ℒ! = ℒ 𝑜! , 𝐵!
Gradient 𝑔! = 𝛻$ ℒ! 𝜃
Second Moment 𝑣! = 𝛽% 𝑣!"# + 1 − 𝛽% 𝑔!% , 𝑣ˆ! =
𝑣! 1 − 𝛽%!
3
Parameter Update
Δ𝜃! = −𝛼
Adam Parameter 𝜃!
𝑚 ˆ!
𝑣ˆ! + 𝜖 𝜃! = 𝜃!"# + Δ𝜃!
(a) Adam inspiration. Adam aggregates first- and second-moment estimates to stabilize the update direction and adapt the effective step size. 1 Rollout
Sample
2 Case k
Case k
Current Skill 𝑆!%& Satisfy clothing constraints. Choose cheapest cart. Verify constraints
… …
Optimization Memory Evolving Issue Tracker 𝑀!
Mini-batch 𝐵!
Case k
Case 1 Case … Case k
Case k
…
…
Trajectories 𝒯! …
…
×𝑘 in total
…
3 Skill Update
Moment Estimation
ℱ!'($
Constraints missed Cheaper cart missed Items mismatched
ℱ!"#$$
Volatility-driven Edit Budget
Patch Generator
SkillAdam Not Improved
SkillPatch 𝑔! - Verify constraints + Map each… + Keep…
Apply on 𝑆!"#
Improved Skill 𝑆!
Updated Skill 𝑆!
Feedbackℱ!"#$$ …
Case-level Changes
×𝑘 in total
Edit Budget
Satisfy clothing constraints. Choose cheapest cart. Map each requirement to one cart item. Keep requirements separate
(b) S KILL A DAM framework. S KILL A DAM uses the Evolving Issue Tracker and a volatility-driven edit budget to provide analogous control over the direction and magnitude of skill updates in the discrete skill space.
Figure 2: Functional correspondence between Adam and S KILL A DAM.
initialization, these components produce the more coherent and efficient optimization path shown in Figure 1(b). Our contributions are summarized as follows: • We propose S KILL A DAM, an Adam-inspired framework for stable and efficient skill selfevolution. • We introduce an optimization memory and a volatility-driven edit budget as functional analogues of Adam’s first- and second-moment mechanisms, respectively stabilizing the optimization direction and adapting the update magnitude. • We conduct extensive experiments on seven benchmarks, where S KILL A DAM achieves state-ofthe-art performance while requiring substantially fewer optimization iterations and reducing overall optimization cost.
2
R ELATED W ORK
Agent Skills. We follow Anthropic’s official definition of Agent Skills as modular packages of instructions and supporting resources that equip an agent with specialized capabilities (Anthropic, 2025). Because Skills are stored independently of model parameters, the same package can be distributed to compatible agents without retraining. Earlier work had already externalized reusable capabilities, although it did not share this package definition. Voyager stores executable programs in a library that can be retrieved for later tasks (Wang et al., 2023). AutoManual learns structured rules through interaction and compiles them into a readable instruction manual (Chen et al., 2024). Agent Skill Induction learns and verifies programmatic skills for web agents (Wang et al., 2025a). SkillsBench provides a common evaluation of package-based Agent Skills and shows that curated Skills improve agent performance across diverse tasks (Li et al., 2026). ExpeL and Agent Workflow Memory instead retain natural-language experience or reusable workflows from past executions (Zhao et al., 2023; Wang et al., 2025b). Our work focuses on Skills whose primary artifact is a naturallanguage instruction document and studies their iterative optimization under task evaluation. Prompt Optimization. Automatic prompt optimization studies how to improve discrete language artifacts without updating model parameters. OPRO proposes new instructions from previously evaluated candidates and their scores (Yang et al., 2024). ProTeGi turns error feedback into tex3
tual gradients and applies search to select prompt edits (Pryzant et al., 2023). TextGrad propagates natural-language feedback through a computation graph to optimize prompts and other textual components (Yuksekgonul et al., 2024). GEPA uses reflective feedback from rollouts to evolve prompts (Agrawal et al., 2026). ERM retains feedback from earlier attempts to support exemplarguided prompt optimization (Yan et al., 2025). They show that evaluation signals can guide discrete textual updates without changing model parameters. Their optimization targets are prompts or components of language-model programs, while S KILL A DAM optimizes a reusable Agent Skill used across task instances. Skill Self-Evolution. Recent work directly automates Agent Skill construction and revision. Trace2Skill analyzes a broad pool of executions and consolidates trajectory-local lessons into a unified skill directory (Ni et al., 2026). SkillAxe iteratively diagnoses and refines LLM-authored skill documents with structured evaluation signals (Gautam et al., 2026). EvoSkill discovers and revises skills through failure analysis, then retains validated candidates through Pareto selection (Alzubi et al., 2026). CoEvoSkills jointly evolves a Skill Generator and a Surrogate Verifier to construct multi-file Skill packages without ground-truth test content (Zhang et al., 2026a). SkillOS trains a curator that updates an external skill repository from accumulated experience (Ouyang et al., 2026). SkillOpt is closest to our setting because it applies bounded textual edits to one skill document and accepts an update only when validation performance improves (Yang et al., 2026). Their optimization targets differ. Some optimize one document, while others construct packages or maintain a repository. Across these settings, existing methods do not jointly maintain persistent optimizer states for the direction and magnitude of successive revisions.
3
P RELIMINARIES
3.1
S KILL O PTIMIZATION P ROBLEM
We consider a skill as a Markdown-formatted natural-language instruction that guides an agent’s behavior in a target domain. In this work, we focus on Skills whose primary artifact is such a structured natural-language instruction document. Let S denote the discrete space of possible skills and D = {d1 , . . . , dN } denote a task dataset. Given a skill S ∈ S and a task d ∈ D, the domain evaluator returns structured case-level evaluation feedback E(S, d) = (E(S, d), C(S, d)) , (1) m where E(S, d) ∈ R is an m-dimensional vector of domain-specific metrics, and C(S, d) denotes optional diagnostic information, such as error descriptions or judge rationales. The metric vector E supports numerical aggregation and acceptance decisions, whereas C is retained in F as languagespace evidence for patch generation and issue tracking. We use F(S, B) = {E(S, d) | d ∈ B}
(2)
to denote the evaluation feedback collected over a task batch B ⊆ D. Let Φ be a task-dependent aggregation function that maps task-level metric vectors to a scalar objective. We define J(S) = Φ ({E(S, d) | d ∈ D}) ,
S ∗ = arg max J(S). S∈S
(3)
The target domain determines the precise forms of E and C, and its evaluation protocol determines Φ. Equation 3 aggregates only E because J is numerical; the diagnostic content C remains available through F to the LLM-based update and memory functions. 3.2
A DAM O PTIMIZATION
In differentiable optimization, Adam (Kingma & Ba, 2015) maintains exponential moving averages of the first and second moments of stochastic gradients. Let ∇t = ∇θ Lt (θt−1 ) denote the stochastic gradient of the mini-batch loss at iteration t. Adam updates mAdam = β1 mAdam t t−1 + (1 − β1 )∇t ,
(4)
Adam vtAdam = β2 vt−1 + (1 − β2 )∇2t ,
(5)
4
Here, mAdam and vtAdam are the first- and second-moment estimates. The coefficients β1 , β2 ∈ t [0, 1) are exponential decay rates. The square is applied element-wise. Since both moment estimates are initialized at zero, Adam applies bias correction: m b Adam = t
mAdam t , 1 − β1t
vbtAdam =
vtAdam . 1 − β2t
(6)
The parameters are then updated as m b Adam , θt = θt−1 − α p t vbtAdam + ϵ
(7)
where α > 0 is the base learning rate and ϵ > 0 is a small constant that prevents division by zero and improves numerical stability. The first-moment estimate aggregates gradient information across iterations to stabilize the update direction, while the second-moment estimate rescales the base learning rate according to the recent squared-gradient magnitude, yielding an adaptive effective step size. The ratio α αteff = p (8) Adam vbt +ϵ can be interpreted as the element-wise effective step size.
4
M ETHOD
To enable stable and efficient skill self-evolution in a discrete skill space, we propose S KILL A DAM with two Adam-inspired mechanisms: an Evolving Issue Tracker that stabilizes the update direction and a volatility-driven edit budget that adapts the update magnitude. We describe the framework and its components below. 4.1
T HE S KILL A DAM F RAMEWORK
Framework Overview. Figure 2 presents the S KILL A DAM cycle. It begins with rollout, after which moment estimation informs the skill update. Before iterative optimization begins, S KIL L A DAM constructs an initial skill S0 from execution trajectories and their evaluation feedback. At iteration t, the current skill is executed on a sampled mini-batch, and the optimizer states carried from the previous iteration, Mt−1 and σt , guide the generation of a candidate modification. After the candidate has been evaluated, the Evolving Issue Tracker compares the rollout and validation outcomes and evolves the structured issue list that constitutes the optimization memory Mt . The same paired feedback is used to update the improvement-volatility estimate. The three blocks in Figure 2 describe functional roles. They do not prescribe a strict within-iteration execution order. In particular, the memory and volatility states obtained from the validation outcome of iteration t guide the skill update at iteration t + 1. The complete operational order is given in Algorithm 1, while Table 1 summarizes the functional correspondence between Adam and S KILL A DAM. Rollout. At iteration t, S KILL A DAM samples a mini-batch Bt ⊂ D and executes the agent with the current skill St−1 : Tt = Rollout (St−1 , Bt ) ,
(9)
where Tt denotes the collected execution trajectories. The domain evaluator then produces the corresponding case-level evaluation feedback Ftroll = F (St−1 , Bt ) ,
(10)
as defined in Equation 2. The trajectories record how the agent behaves during execution, while Ftroll contains the resulting metric scores and diagnostic information. Together, (Tt , Ftroll ) provide the iteration-specific optimization signal used to generate the next skill update. 5
Moment Estimation. S KILL A DAM maintains two persistent optimizer states across iterations. The optimization memory is instantiated as an Evolving Issue Tracker (EIT), denoted by Mt , while the volatility estimate Vt controls the magnitude of subsequent updates. Both states are refreshed after the candidate skill has been evaluated and are then carried into the next iteration. The Evolving Issue Tracker is represented as a collection of structured issue entries: Mt = {Ij | j ∈ Jt } ,
Ij = (pj , zj , Aj ) ,
(11)
where Jt is the set of issue identifiers recorded through iteration t. The pair (pj , zj ) describes an error pattern and its current status. The set Aj stores previous solution attempts with their observed outcomes. The tracker therefore maintains each issue and its resolution history. After the candidate skill at iteration t has been evaluated, the tracker directly evolves its state: Mt = UEIT Mt−1 , Tt , Ftroll , gt , Ftval , at . (12) The LLM-based update function UEIT compares the rollout and validation outcomes to maintain the issue records. It links observed failures to existing issues when possible and creates entries when needed. It also records each attempted modification with its outcome and reopens an issue when the same failure recurs. The Evolving Issue Tracker serves as a functional analogue of Adam’s first-moment estimate: mAdam t
←→
Mt .
(13)
Both states integrate current evidence with information accumulated across previous iterations. Adam aggregates numerical gradients. The Evolving Issue Tracker instead accumulates issue histories and solution outcomes to stabilize the update direction. S KILL A DAM additionally estimates how consistently the candidate affects the evaluated cases. Let roll val Et,i and Et,i denote the case-level feedback for the current and candidate skills, respectively, on case di ∈ Bt . The case-level improvement is val roll δt,i = s Et,i − s Et,i , di ∈ Bt , (14) where s(·) extracts the benchmark-specific scalar score used to measure improvement. The mean case-level improvement is 1 X δt = δt,i . (15) |Bt | di ∈Bt
S KILL A DAM estimates the current improvement volatility using X 2 1 Vbt = δt,i − δ t , |Bt | ≥ 2. |Bt | − 1
(16)
di ∈Bt
When fewer than two valid case-level comparisons are available, we set Vbt = 0. A high value indicates that the candidate produces substantially different effects across cases, while a low value indicates more consistent effects. The history-weighted volatility estimate is updated as Vt = β2 Vt−1 + (1 − β2 )Vbt ,
V0 = 0.
The edit budget for the next iteration is then computed by Vt σt+1 = max bmin , bbase 1 − clip , 0, 1 , Vmax
(17)
(18)
where bbase is the base edit budget and bmin is the minimum allowable budget. The volatility saturation threshold is Vmax . High volatility yields a smaller edit budget, while low volatility permits a broader modification. The budget used at iteration t, σt , is determined from the state accumulated through iteration t − 1. This mechanism provides the following correspondence: vtAdam
←→
αteff
Vt ,
←→
σt+1 .
(19)
The optimization memory therefore determines which problems should guide the subsequent update, while the volatility-driven edit budget determines how extensively the skill may be modified. 6
Skill Update. At iteration t, the LLM-based patch generator proposes a skill modification conditioned on the current rollout evidence and the optimizer states carried from the previous iteration: gt = GLLM St−1 , Tt , Ftroll , Mt−1 , σt . (20) Here, Tt and Ftroll describe the behavior observed in the current iteration, Mt−1 provides crossiteration guidance on which problems should be addressed, and σt controls the allowable scope of the modification. Applying the generated modification to the current skill produces a candidate: Set = Apply (St−1 , gt ) . The candidate is evaluated on the same mini-batch used for the current-skill rollout: Ftval = F Set , Bt .
(21)
(22)
Consequently, Ftroll and Ftval provide directly comparable case-level feedback for the current and candidate skills. A benchmark-specific acceptance gate determines whether the candidate should replace the current skill: at = Gθ St−1 , Set , Ftroll , Ftval ∈ {0, 1}. (23) The skill is updated as ( St =
Set , St−1 ,
at = 1, at = 0.
(24)
The Evolving Issue Tracker uses each attempted modification and its evaluated outcome to evolve from Mt−1 to Mt . The same paired feedback is used to compute Vbt and Vt . The updated skill and optimizer state are carried into the next iteration, closing the optimization loop shown in Figure 2. 4.2
A LGORITHM AND A DAM C ORRESPONDENCE
Algorithm 1 presents the operational order of S KILL A DAM. The current skill first produces rollout trajectories and evaluation feedback on a sampled mini-batch. The optimizer states carried from the previous iteration then guide the generation of a candidate skill. After the candidate has been evaluated on the same cases, S KILL A DAM decides whether to accept it and uses the resulting comparison to evolve the Evolving Issue Tracker and update the volatility estimate. These updated states guide the next iteration. Figure 2 provides a block-level comparison between Adam and S KILL A DAM, while Table 1 makes the correspondence between their optimization quantities explicit. The analogy is functional, not numerical. S KILL A DAM does not compute gradients in the discrete skill space. It constructs languagespace states that serve the same optimization roles. The rollout feedback Ftroll plays a loss-like role by evaluating the current behavior, while the generated patch gt serves as the gradient-like local update signal in the skill space. The Evolving Issue Tracker Mt accumulates issue histories and solution outcomes across iterations, analogous to Adam’s first moment. Similarly, Vt aggregates recent improvement variability and determines the next edit budget σt+1 .
5
E XPERIMENTS
5.1
B ENCHMARKS
We evaluate S KILL A DAM on seven benchmarks that cover knowledge work and interactive planning. Six benchmarks contribute one evaluation slice each. DeepPlanning contributes Shopping Levels 1–3 and Travel EN, giving ten slices in total. Each benchmark uses its native task environment, tool interface, and evaluator. 7
Algorithm 1 S KILL A DAM: Adam-Inspired Skill Self-Evolution Require: Task dataset D, evaluator E, mini-batch size k, maximum iterations Tmax , EMA coefficient β2 , budget parameters bbase , bmin , Vmax , and random seed ξ Ensure: Optimized skill STmax 1: (T0 , F0 ) ← CollectInitializationData(D, E, ξ) 2: S0 ← InitializeLLM (T0 , F0 ) 3: M0 ← ∅ 4: V0 ← 0 5: σ1 ← bbase 6: for t = 1, 2, . . . , Tmax do 7: Bt ← Sample(D, k, ξ + t), Tt ← Rollout(St−1 , Bt ), Ftroll ← F(St−1 , Bt ) 8: gt ← GLLM (St−1 , Tt , Ftroll , Mt−1 , σt ), Set ← Apply(St−1 , gt ) 9: if Set = null then 10: St ← St−1 , Mt ← Mt−1 , Vt ← Vt−1 , σt+1 ← σt 11: continue 12: end if 13: Ftval ← F(Set , Bt ), at ← Gθ (St−1 , Set , Ftroll , Ftval ) 14: if at = 1 then 15: St ← Set 16: else 17: St ← St−1 18: end if 19: Mt ← UEIT Mt−1 , Tt , Ftroll ,gt , Ftval , at 20: {δt,i }di ∈Bt ← ∆ Ftroll , Ftval 21: Vbt ← Var ({δt,i }di ∈Bt ) 22: Vt ← β2 Vt−1 + (1 −jβ2 )Vbth im Vt 23: σt+1 ← max bmin , bbase 1 − clip Vmax , 0, 1 24: end for 25: return STmax Table 1: Core functional correspondence between Adam and S KILL A DAM. The correspondence is functional, with evaluation feedback and a language-space patch serving the roles of the loss and gradient. Adam
S KILL A DAM
Functional role
Parameters θt−1 Prediction ot Loss Lt Gradient ∇t First moment mAdam t Second moment vtAdam Effective step size αteff Apply ∆θt Updated parameter θt
Current skill St−1 Execution trajectories Tt Evaluation feedback Ftroll Skill patch gt Evolving Issue Tracker Mt Historical volatility Vt Edit budget σt+1 Apply and gate gt Accepted skill St
Represent the optimized state Record the execution output Evaluate the current behavior Provide the local update signal Stabilize the update direction Accumulate update variability Adapt the update magnitude Update the optimized state Carry the state forward
We classify five benchmarks as short-horizon, as they require fewer than ten tool calls per task on average. SearchQA (Dunn et al., 2017) evaluates open-domain question answering with noisy retrieved context. SpreadsheetBench (Ma et al., 2024) requires an agent to modify spreadsheets from natural-language instructions. OfficeQA (Opsahl-Ong et al., 2026) evaluates question answering over historical U.S. Treasury Bulletins. DocVQA (Mathew et al., 2021) evaluates visual question answering over document images. LiveMathematicianBench (LMB; also abbreviated as LiveMath in the tables) (He et al., 2026) evaluates reasoning over mathematical theorems and proof sketches. The two long-horizon benchmarks require at least ten tool calls per task on average. ALFWorld (Shridhar et al., 2021) contains text-based embodied household tasks. DeepPlanning (Zhang et al., 2026b) evaluates multi-step shopping and travel planning. 8
Metrics. We report the standard primary metric for each benchmark. SearchQA, OfficeQA, and LMB use Exact Match. SpreadsheetBench uses hard task success, which requires the complete spreadsheet-editing task to pass. DocVQA uses ANLS-hard. ALFWorld uses episode goalcompletion rate. DeepPlanning reports Shopping Case Accuracy and Travel Case Accuracy. DPShopping pools the three shopping levels by case count, with 25, 25, and 10 test cases. DP-Travel contains 60 test cases. We define DP-Avg = (DP-Shopping + DP-Travel)/2 using the unrounded domain-level accuracies. It is not an equal average of the four DeepPlanning slices. All main results are percentages. 5.2
BASELINES
We compare S KILL A DAM with seven baselines. NoSkill runs the target agent without an injected skill. HumanSkill uses an expert-authored skill, and LLMSkill uses a skill written in one LLM call. The optimization baselines are Trace2Skill (Ni et al., 2026), TextGrad (Yuksekgonul et al., 2024), GEPA (Agrawal et al., 2026), and SkillOpt (Yang et al., 2026). For the five short-horizon benchmarks, baseline results are taken from the GPT-5.5 no-harness setting reported by SkillOpt. For ALFWorld, the NoSkill result is taken from SkillOpt, and we reproduce the SkillOpt result in our environment. For DeepPlanning, we reproduce both NoSkill and SkillOpt in the task environment used for S KILL A DAM. Within each benchmark, the compared methods use the same target-agent configuration, test cases, and evaluator. S KILL A DAM and SkillOpt also start from the same initial skill, which is constructed from a fixed set of baseline execution trajectories. SkillOpt retains its original train and selection split together with its native selection, slow-update, and optimizer-memory procedures. We run SkillOpt for four epochs under this protocol. 5.3
S ETUP
Common Protocol. All experiments use a no-harness, direct-chat setting. The agent interacts with the native task interface and evaluator of each benchmark without an additional orchestration layer. Skills are injected as natural-language instructions, and the target model remains frozen during optimization. We use GPT-5.5 for the six benchmarks outside DeepPlanning. These runs use medium reasoning effort, a temperature of 1.0, and a maximum output length of 16,384 tokens. DeepPlanning uses Claude Sonnet 4.5 with a temperature of 0.0 and the same output limit. Travel EN uses a separate frozen model to convert the generated plan into the required format before evaluation. Initialization and Optimization. S KILL A DAM and SkillOpt share an initial skill constructed from fixed baseline execution trajectories. For the six benchmarks outside DeepPlanning, S KIL L A DAM merges the original train and selection partitions into one optimization pool. SkillOpt retains the original split. The test partition is unchanged and is used only for final evaluation. DeepPlanning uses an odd-even split by case identifier. Odd-numbered cases are used for optimization, and even-numbered cases are reserved for testing. We use seed 42 for local case sampling, optimization-pool shuffling, and other controlled random operations. On the six benchmarks outside DeepPlanning, S KILL A DAM makes one pass through the optimization pool and uses a benchmark-specific stopping rule. DeepPlanning samples optimization batches at random and uses task-specific stopping criteria. At each iteration, S KILL A DAM evaluates the current skill and its proposed revision on the same sampled cases. Their execution trajectories and evaluation feedback are passed to the optimizer states described in Section 4.1. Acceptance Protocol. The acceptance gate uses the benchmark’s primary and auxiliary metrics. A proposed revision is accepted when at least one designated metric reaches its improvement threshold and every protected metric remains within its regression boundary. An auxiliary metric can therefore support acceptance when the primary metric is unchanged. S KILL A DAM does not use an additional validation set for this decision. SkillOpt follows its native selection and slow-update rules. Evaluation.
Each test case is evaluated once with the target-agent configuration specified above. 9
Table 2: Main results on five short-horizon benchmarks. Baselines follow the GPT-5.5 no-harness results reported by SkillOpt. Scores are percentages. Bold and underlining mark the best and secondbest values. Method
SearchQA
Spreadsheet
OfficeQA
DocVQA
LiveMath
NoSkill HumanSkill LLMSkill Trace2Skill TextGrad GEPA SkillOpt
77.7 81.8 80.9 82.4 81.4 84.8 87.3
41.8 72.9 43.2 49.6 41.1 73.6 80.7
33.1 66.9 51.7 65.7 42.0 63.9 72.1
78.8 90.1 89.6 90.6 87.2 89.1 91.2
37.6 38.4 40.0 52.0 49.2 43.2 66.9
S KILL A DAM
87.5
81.1
72.1
92.3
67.7
Table 3: Main results on ALFWorld with GPT-5.5 and DeepPlanning with Claude Sonnet 4.5. DPAvg averages the unrounded DP-Shopping and DP-Travel accuracies. Scores are percentages. Bold and underlining mark the best and second-best values.
5.4
Method
ALFWorld
DP-Shopping
DP-Travel
DP-Avg
NoSkill SkillOpt
83.6 87.3
31.7 41.7
0.0 1.7
15.8 21.7
S KILL A DAM
89.6
45.0
11.7
28.3
M AIN R ESULTS
We compare S KILL A DAM with all baselines on two categories of benchmarks, i.e., short-horizon benchmarks and long-horizon benchmarks. The results are recorded in Table 2 and Table 3. Results on Short-Horizon Benchmarks. As illustrated in Table 2, S KILL A DAM achieves the overall best performance among all baselines, with four strict wins and one tie on five benchmarks. Specifically, S KILL A DAM achieves remarkably better performance compared with HumanSkill and LLMSkill, with an average improvement of 14.45% and 31.20% respectively. This is because S KIL L A DAM adopts iterative skill optimization instead of one-shot skill generation. Thus, it can optimize the skill based on the evaluation result of the immediate agent rollouts, which produces higherquality skills. In addition, compared with the second-best methods, i.e., SkillOpt, S KILL A DAM still shows better performance. For example, S KILL A DAM achieves higher accuracy on DocVQA and LiveMath, with an improvement of 1.21% and 1.20%. This improvement stems from two technical designs of S KILL A DAM, i.e., Optimization Memory and Volatility-driven Edit Budget. With these designs, we can optimize the skills more stably. Thus, we can produce better skills than SkillOpt. Results on Long-Horizon Benchmarks. Table 3 reports the long-horizon results. As illustrated in the table, the advantages of S KILL A DAM are more prominently demonstrated. Specifically, while both NoSkill and SkillOpt nearly fail on DP-Travel, S KILL A DAM achieves an accuracy of 11.7%. Moreover, compared with SkillOpt, S KILL A DAM improves DP-Avg from 21.7% to 28.3%, a gain of 6.7 percentage points computed from the unrounded accuracies. The results demonstrated that our proposed evolution algorithm can work better on more challenging tasks. In summary, S KILL A DAM achieves the best overall performance compared with other baselines, demonstrating the effectiveness of the proposed skill evolution strategy. 5.5
A BLATION S TUDIES
We conduct a cumulative ablation study on DeepPlanning to examine the two core mechanisms. Starting from the full framework, we first remove the volatility-driven edit budget and then additionally remove the optimization memory. Thus, each row removes one additional component from the configuration above it. All variants follow the same training and evaluation protocol. 10
Table 4: Cumulative ablation on DeepPlanning. M and B denote the optimization memory and volatility-driven edit budget. Each row removes one additional component. Scores are percentages, and bold marks the best value in each column. Configuration S KILL A DAM (M + B) −B −B, −M NoSkill
DP-Shopping L1
L2
L3
All
52.0 48.0 56.0 40.0
36.0 36.0 20.0 24.0
50.0 40.0 30.0 30.0
45.0 41.7 36.7 31.7
DP-Travel
DP-Avg
11.7 1.7 1.7 0.0
28.3 21.7 19.2 15.8
Table 5: Cross-model transfer from GPT-5.5 to GPT-5.4-mini. Score denotes target-model task performance. Retention measures the percentage of source-model performance preserved after transfer. Both metrics are percentages. Bold marks the better result between SkillOpt and S KILL A DAM. Method
Metric
SearchQA DocVQA Spreadsheet ALFWorld LiveMath OfficeQA Average
NoSkill
Score
75.9
71.4
36.1
73.1
14.7
22.1
48.9
SkillOpt SkillOpt
Score Retention
80.8 92.6%
90.4 99.1%
61.4 76.1%
66.4 76.1%
28.2 42.2%
51.2 71.0%
63.1 76.2%
S KILL A DAM Score S KILL A DAM Retention
78.9 90.2%
91.2 98.8%
65.0 80.1%
82.8 92.4%
33.1 48.9%
55.8 77.4%
67.8 81.3%
As illustrated in Table 4, the full S KILL A DAM achieves the best overall result, increasing DP-Avg from 19.2% without both mechanisms to 28.3%. Specifically, with Optimization Memory fixed, adding the Volatility-driven Edit Budget improves DP-Avg by 6.7 percentage points, with the largest gain occurring on DP-Travel. This result supports the importance of Update Adaptivity in longhorizon planning. Specifically, when a revision produces inconsistent effects across cases, reducing the next edit scope can avoid damaging previously correct constraints, whereas consistent improvements permit broader updates. Optimization Memory provides a complementary benefit. Without the edit budget, adding memory raises DP-Avg from 19.2% to 21.7%, with gains on Shopping L2 and L3 despite a decrease on L1. The concentration of gains on the more difficult levels suggests that recording issue states and prior solution outcomes helps preserve and reconcile multiple corrections across iterations, thereby maintaining a more stable optimization direction. However, because this is a cumulative ablation, it does not independently isolate the effect of memory when the edit budget is enabled or the interaction between the two mechanisms. 5.6
C ROSS -M ODEL T RANSFER
We evaluate whether the optimized skills remain effective when transferred to a different agent backbone. Specifically, the skills produced by S KILL A DAM and SkillOpt using GPT-5.5 are deployed verbatim on GPT-5.4-mini without further optimization or adaptation. We conduct this evaluation on six benchmarks supported by both backbones. Table 5 reports the task score obtained on GPT-5.4-mini and the corresponding retention ratio. Let ssrc and stgt denote the scores obtained on GPT-5.5 and GPT-5.4-mini, respectively. We define the retention ratio as Retention = stgt /ssrc . A higher retention ratio indicates that a larger fraction of the skill’s source-model performance is preserved after transfer. No-Skill results on GPT-5.4-mini are included as reference performance and are taken from Yang et al. (2026). As illustrated in Table 5, S KILL A DAM achieves higher target-model scores on five of the six benchmarks and higher retention ratios on four of the six benchmarks. Compared with SkillOpt, S KIL L A DAM improves the average target-model score from 63.1% to 67.8% and the average retention ratio from 76.2% to 81.3%. The advantage remains after excluding ALFWorld, indicating that the improvement is not driven by a single benchmark; SearchQA is the only target-score exception. 11
Overall score (1–5)
5 SkillAdam (ours)
SkillOpt
Avg Score of SkillAdam (ours)
Avg Score of SkillOpt
4
3
2
1
ALFWorld
DocVQA
SearchQA
Spreadsheet
OfficeQA
LMB
DP/L1
DP/L2
DP/L3
DP/Travel
Figure 3: Overall skill-quality scores on a 1–5 scale across ten benchmark slices. Each value averages the six dimension medians from three independent calls to the same GPT-5.5 judge. Horizontal markers show the mean across slices.
These results show that the skills optimized by S KILL A DAM transfer more effectively across agent backbones. A plausible explanation is that Optimization Memory consolidates recurring problems and evaluated solution attempts across iterations, encouraging the final skill to capture task-level procedures rather than source-model-specific wording. The transfer experiment evaluates the complete framework, however, and therefore does not isolate which component is responsible for the gain. 5.7
S KILL Q UALITY A NALYSIS
Beyond task performance, we evaluate the optimized skills with a frozen GPT-5.5 judge. The comparison covers ten benchmark slices, including six individual benchmarks and four DeepPlanning slices. Each skill is evaluated through three independent calls to the same judge model. The input does not reveal the method name. Each call assigns a score from 1 to 5 on six dimensions. Task Alignment and Non-Obvious Insight measure relevance. Constraint Handling and Decision Framework measure operational guidance. Cognitive Load and Inference Efficiency measure usability. For each dimension, we take the median of the three scores. The overall score for a skill is the mean of these six medians. We then average the overall scores across the ten benchmark slices. As illustrated in Figure 3, S KILL A DAM obtains a higher skill-quality score than SkillOpt on all ten benchmark slices, increasing the mean score from 3.30 to 3.88. Because every paired comparison favors S KILL A DAM, the improvement is consistent across tasks rather than being driven by a small number of benchmarks. This result complements the task-performance results in Section 5.4. The benchmark metrics measure whether the agent completes the task, whereas the judge evaluates whether the skill provides relevant, operational, and usable guidance. The agreement between the two evaluations suggests that S KILL A DAM improves not only downstream execution but also the quality of the skill document itself. This improvement is consistent with Optimization Memory preserving useful constraints and the Volatility-driven Edit Budget limiting uncontrolled revisions. However, the judge experiment evaluates the complete framework, and Figure 3 aggregates all six dimensions; it therefore isolates neither an individual component nor a single quality dimension. 5.8
T RAINING C OST
We compare optimization-phase API use for S KILL A DAM and SkillOpt on the four DeepPlanning slices. Both methods use Claude Sonnet 4.5 and optimize over the same cases. The counts cover the optimization and iteration phase. They exclude initial-skill generation and final test evaluation. Unrelated smoke tests are also excluded. Token counts come from the raw counters returned by the API and do not represent a dollar cost. We report input tokens, output tokens, and API requests together with the resulting DeepPlanning performance. 12
Table 6: Optimization-phase API use and DeepPlanning test performance. API-use deltas are relative reductions. Performance deltas are absolute percentage-point differences. Metric
SkillOpt
S KILL A DAM
Delta
Input Tokens Output Tokens Total Tokens API Requests
222.1M 4.5M 226.6M 9,071
72.5M 1.5M 74.0M 2,830
−67.3% −66.8% −67.3% −68.8%
Resulting test performance (higher is better) DP-Shopping DP-Travel DP-Avg
(b) Accuracy vs. Token Cost Test Case Accuracy (%)
Test Case Accuracy (%)
(a) Accuracy vs. Iteration 60
50
40
30 0
1
2
3
4
5
6
7
8
9
10
60
50
40
30 0
Iteration Seed
SkillAdam (ours)
+3.3 pp +10.0 pp +6.7 pp
45.0% 11.7% 28.3%
41.7% 1.7% 21.7%
10
20
30
40
50
Cumulative Tokens (M) SkillOpt
Accepted
Rejected
Slow-update
Figure 4: Optimization dynamics on DeepPlanning Shopping Level 1.
As illustrated in Table 6, S KILL A DAM reduces total token consumption by 67.3% and API requests by 68.8% compared with SkillOpt, while improving DP-Avg from 21.7% to 28.3%. Thus, the lower optimization cost is not achieved by sacrificing the quality of the final skill. The two methods consume a similar number of tokens per request, and S KILL A DAM is slightly higher on this measure. Therefore, the reduction does not come from shorter individual calls, but from requiring fewer optimization requests and modification attempts. This result is consistent with the two technical designs of S KILL A DAM: Optimization Memory avoids repeatedly rediscovering previously identified failures, while the Volatility-driven Edit Budget reduces broad, weakly supported revisions. Measured by DP-Avg per million optimization tokens, S KILL A DAM is approximately four times as efficient as SkillOpt.
6
A NALYSES
We analyze S KILL A DAM and SkillOpt on DeepPlanning Shopping Level 1. Both methods use Claude Sonnet 4.5 and start from the same initial skill, which has 40% test case accuracy. One iteration denotes one minibatch-level modification attempt. S KILL A DAM compares the current skill with its proposed revision on the same sampled optimization cases and applies its multi-metric acceptance gate. SkillOpt follows its native selection and slow-update rules. The markers in Figure 4 therefore record the decision made by each method’s own update protocol. The figure reports the post-hoc test accuracy of every evaluated modification. Panel (a) uses the minibatch-level iteration index, and Panel (b) uses cumulative token consumption. Test performance is shown only for analysis and does not determine whether a modification is accepted. A modification with a high test score can still be rejected by the cases and metrics used in the corresponding update protocol. 13
6.1
S TABILITY
As illustrated in Figure 4(a), S KILL A DAM finds a strong revision in the first iteration, and all subsequent accepted skills remain above the initial performance. In contrast, SkillOpt’s accepted regular updates fluctuate around or below its starting performance, while its later high-scoring candidates are not retained. This comparison shows that S KILL A DAM produces a more stable accepted optimization path. The result is consistent with the design of Optimization Memory. By retaining previously identified issues, their status, and the outcomes of prior solutions, S KILL A DAM can incorporate new feedback without repeatedly overwriting useful corrections. The acceptance gate further prevents insufficiently supported revisions from replacing the current skill. The rejected S KILL A DAM revision at iteration 8 nevertheless has a high post-hoc test score. This does not contradict the stability result because acceptance is determined only by sampled optimization cases and protected metrics, while the test score is used solely for retrospective analysis. Since the figure reports one run and memory operates together with the gate, it characterizes the overall optimization behavior rather than isolating the causal effect of memory alone. 6.2
E FFICIENCY
As illustrated in Figure 4(b), S KILL A DAM reaches an accepted skill with 56% test accuracy after approximately 4M tokens, whereas SkillOpt first produces a comparably strong candidate after approximately 20M tokens and does not accept it. S KILL A DAM therefore reaches and retains a stronger skill with about one fifth of the token cost in this run. Combined with the similar token cost per request in Table 6, the result shows that the efficiency gain comes from fewer unproductive modification attempts rather than cheaper individual calls. This behavior is consistent with the two core mechanisms: the Volatility-driven Edit Budget restricts broad edits when recent case-level effects are inconsistent, and Optimization Memory prevents repeated rediscovery of earlier failures. Together, they allow S KILL A DAM to identify useful revisions with fewer iterations.
7
C ONCLUSION
We introduced S KILL A DAM, a gradient-inspired framework for iterative skill self-evolution that adapts the two principles behind Adam to the skill space. An optimization memory accumulates structured records across iterations to provide stability, and a volatility-driven edit budget scales each update to the reliability of the improvement evidence to provide adaptivity. Across seven benchmarks that span short-horizon and long-horizon agentic tasks, S KILL A DAM achieves state-ofthe-art performance at substantially lower training cost. Its skills also transfer better across models.
R EFERENCES Lakshya A. Agrawal, Shangyin Tan, Dilara Soylu, Noah Ziems, Rishi Khare, Krista Opsahl-Ong, Arnav Singhvi, Herumb Shandilya, Michael J. Ryan, Meng Jiang, Christopher Potts, Koushik Sen, Alexandros G. Dimakis, Ion Stoica, Dan Klein, Matei Zaharia, and Omar Khattab. GEPA: Reflective prompt evolution can outperform reinforcement learning. In International Conference on Learning Representations (ICLR), 2026. URL https://arxiv.org/abs/2507. 19457. Oral presentation. Salaheddin Alzubi, Noah Provenzano, Jaydon Bingham, Weiyuan Chen, and Tu Vu. EvoSkill: Automated skill discovery for multi-agent systems. arXiv preprint arXiv:2603.02766, 2026. URL https://arxiv.org/abs/2603.02766. Anthropic. Introducing Agent Skills, October 2025. URL https://claude.com/blog/ skills. Minghao Chen, Yihang Li, Yanting Yang, Shiyu Yu, Binbin Lin, and Xiaofei He. AutoManual: Constructing instruction manuals by LLM agents via interactive environmental learning. In Advances in Neural Information Processing Systems, vol14
ume 37, 2024. URL https://papers.nips.cc/paper_files/paper/2024/hash/ 0142921fad7ef9192bd87229cdafa9d4-Abstract-Conference.html. Chao Deng, Shaolei Zhang, Ju Fan, and Xiaoyong Du. DataEvolver: Automatic data preparation for large language models through multi-level self-evolving. arXiv preprint arXiv:2606.07001, 2026. URL https://arxiv.org/abs/2606.07001. Matthew Dunn, Levent Sagun, Mike Higgins, V. Ugur Guney, Volkan Cirik, and Kyunghyun Cho. SearchQA: A new Q&A dataset augmented with context from a search engine. arXiv preprint arXiv:1704.05179, 2017. URL https://arxiv.org/abs/1704.05179. Meihao Fan, Ju Fan, Yuxin Zhang, Shaolei Zhang, Xiaoyong Du, Jie Song, Peng Li, Fuxin Jiang, Tieying Zhang, and Jianjun Chen. DeepPrep: An LLM-powered agentic system for autonomous data preparation. Proceedings of the VLDB Endowment, 19(11):3371–3384, 2026. doi: 10.14778/ 3836663.3836695. URL https://www.vldb.org/pvldb/vol19/p3371-fan.pdf. Srishti Gautam, Arjun Radhakrishna, and Sumit Gulwani. SkillAxe: Sharpening LLM-authored agent skills through evaluation-guided self-refinement. arXiv preprint arXiv:2606.10546, 2026. URL https://arxiv.org/abs/2606.10546. Linyang He, Qiyao Yu, Hanze Dong, Baohao Liao, Xinxing Xu, Micah Goldblum, Jiang Bian, and Nima Mesgarani. LiveMathematicianBench: A live benchmark for mathematician-level reasoning with proof sketches. arXiv preprint arXiv:2604.01754, 2026. URL https://arxiv. org/abs/2604.01754. Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015. URL https://arxiv.org/abs/ 1412.6980. Xiangyi Li et al. SkillsBench: Benchmarking how well agent skills work across diverse tasks. arXiv preprint arXiv:2602.12670, 2026. URL https://arxiv.org/abs/2602.12670. Yizhe Liu, Shaolei Zhang, and Ju Fan. DA-Studio: An agentic system for end-to-end data analysis. Proceedings of the VLDB Endowment, 19(12):4766–4769, 2026. doi: 10.14778/3827998. 3828117. URL https://www.vldb.org/pvldb/vol19/p4766-liu.pdf. Zeyao Ma, Bohan Zhang, Jing Zhang, Jifan Yu, Xiaokang Zhang, Xiaohan Zhang, Sijia Luo, Xi Wang, and Jie Tang. SpreadsheetBench: Towards challenging real world spreadsheet manipulation. In Advances in Neural Information Processing Systems (NeurIPS), 2024. URL https://arxiv.org/abs/2406.14991. Spotlight. Minesh Mathew, Dimosthenis Karatzas, and C.V. Jawahar. DocVQA: A dataset for VQA on document images. In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2021. URL https://arxiv.org/abs/2007.00398. Jingwei Ni, Yihao Liu, Xinpeng Liu, Yutao Sun, Mengyu Zhou, Pengyu Cheng, Dexin Wang, Erchao Zhao, Xiaoxi Jiang, and Guanjun Jiang. Trace2Skill: Distill trajectory-local lessons into transferable agent skills. arXiv preprint arXiv:2603.25158, 2026. URL https://arxiv. org/abs/2603.25158. Krista Opsahl-Ong, Arnav Singhvi, Jasmine Collins, Ivan Zhou, Cindy Wang, Ashutosh Baheti, Owen Oertell, Jacob Portes, Sam Havens, Erich Elsen, Michael Bendersky, Matei Zaharia, and Xing Chen. OfficeQA Pro: An enterprise benchmark for end-to-end grounded reasoning. arXiv preprint arXiv:2603.08655, 2026. URL https://arxiv.org/abs/2603.08655. Siru Ouyang, Jun Yan, Yanfei Chen, Rujun Han, Zifeng Wang, Bhavana Dalvi Mishra, Rui Meng, Chun-Liang Li, Yizhu Jiao, Kaiwen Zha, Maohao Shen, Vishy Tirumalashetty, George Lee, Jiawei Han, Tomas Pfister, and Chen-Yu Lee. SkillOS: Learning skill curation for self-evolving agents. arXiv preprint arXiv:2605.06614, 2026. URL https://arxiv.org/abs/2605. 06614. 15
Reid Pryzant, Dan Iter, Jerry Li, Yin Lee, Chenguang Zhu, and Michael Zeng. Automatic prompt optimization with “gradient descent” and beam search. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 7957–7968. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.emnlp-main.494. URL https://aclanthology.org/2023.emnlp-main.494/. Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. ALFWorld: Aligning text and embodied environments for interactive learning. In International Conference on Learning Representations (ICLR), 2021. URL https://arxiv. org/abs/2010.03768. Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023. URL https://arxiv.org/abs/2305.16291. Zora Zhiruo Wang, Apurva Gandhi, Graham Neubig, and Daniel Fried. Inducing programmatic skills for agentic tasks. In Conference on Language Modeling, 2025a. URL https: //openreview.net/forum?id=lsAY6fWsog. Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig. Agent workflow memory. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pp. 63897–63911. PMLR, 2025b. URL https://proceedings.mlr.press/v267/wang25bx.html. Cilin Yan, Jingyun Wang, Lin Zhang, Ruihui Zhao, Xiaopu Wu, Kai Xiong, Qingsong Liu, Guoliang Kang, and Yangyang Kang. Efficient and accurate prompt optimization: The benefit of memory in exemplar-guided reflection. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, pp. 753–779. Association for Computational Linguistics, 2025. doi: 10.18653/v1/2025.acl-long.37. URL https://aclanthology.org/2025.acl-long. 37/. Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, and Xinyun Chen. Large language models as optimizers. In International Conference on Learning Representations, 2024. URL https://arxiv.org/abs/2309.03409. Yifan Yang, Ziyang Gong, Weiquan Huang, Qihao Yang, Ziwei Zhou, Zisu Huang, Yan Li, Xuemei Gao, Qi Dai, Bei Liu, Kai Qiu, Yuqing Yang, Dongdong Chen, Xue Yang, and Chong Luo. SkillOpt: Executive strategy for self-evolving agent skills. arXiv preprint arXiv:2605.23904, 2026. URL https://arxiv.org/abs/2605.23904. Mert Yuksekgonul, Federico Bianchi, Joseph Boen, Sheng Liu, Zhi Huang, Carlos Guestrin, and James Zou. TextGrad: Automatic “differentiation” via text. arXiv preprint arXiv:2406.07496, 2024. URL https://arxiv.org/abs/2406.07496. Hanrong Zhang, Shicheng Fan, Henry Peng Zou, Yankai Chen, Zhenting Wang, Jiayu Zhou, Chengze Li, Wei-Chieh Huang, Yifei Yao, Kening Zheng, Xue Liu, Xiaoxiao Li, and Philip S. Yu. CoEvoSkills: Self-evolving agent skills via co-evolutionary verification. arXiv preprint arXiv:2604.01687, 2026a. URL https://arxiv.org/abs/2604.01687. Shaolei Zhang, Ju Fan, Meihao Fan, Guoliang Li, and Xiaoyong Du. DeepAnalyze: Agentic large language models for autonomous data science. arXiv preprint arXiv:2510.16872, 2025. URL https://arxiv.org/abs/2510.16872. Yinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu, Yang Su, Lianghao Deng, Xudong Guo, Chenxu Lv, and Junyang Lin. DeepPlanning: Benchmarking long-horizon agentic planning with verifiable constraints. arXiv preprint arXiv:2601.18137, 2026b. URL https://arxiv.org/ abs/2601.18137. Yuxin Zhang, Ju Fan, Meihao Fan, Shaolei Zhang, and Xiaoyong Du. CoDA-Bench: Can code agents handle data-intensive tasks? In Proceedings of the 43rd International Conference on Machine Learning, 2026c. URL https://arxiv.org/abs/2606.15300. 16
Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. ExpeL: LLM agents are experiential learners. arXiv preprint arXiv:2308.10144, 2023. URL https: //arxiv.org/abs/2308.10144.
17