Conceptio › Archive › arXiv CS
arXiv CSopen access

Failure-Guided Co-Evolution of Prompts and Training Data

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

Failure-Guided Co-Evolution of Prompts and Training Data Tianyu Yuan, Zhuzhong Qian*

Reasoning: AIME-2025

Abstract Automatic prompt optimization (APO) improves languagemodel programs by revising prompts from task feedback, yet it typically holds its training data fixed. Repeatedly optimizing against the same instances confines feedback to weaknesses already represented in those data, leaving related failure conditions unexplored. We therefore view each failure as a dual signal: it indicates both how the prompt should be revised and what new training evidence should be synthesized. We introduce Forge, a failure-guided framework that co-evolves prompts and training data. Forge abstracts imperfect executions into reusable failure modes and synthesizes new training data through four complementary mutation strategies. Verified instances are fed back into prompt search, allowing updated prompts to expose the next data needs. Across eight heterogeneous benchmarks, Forge improves the aggregate score over the unoptimized baseline by 16.52 percentage points and outperforms all evaluated APO baselines. The synthesized data also transfer beyond Forge: in a transfer study, they improve all nine APO comparisons by 2–9 points and all three GRPO comparisons by 4–8 points under matched optimization budgets. These results establish failures as a shared interface between prompt optimization and data synthesis, and show the benefit of jointly adapting what a model is instructed to do and what it learns from.

Agents: AppWorld

90 84.7

82.2

80

77.8

75.6

73.3

74.3 71.2

71.1

70

66.6

67.7

60

50

Score (%)

arXiv:2609.15209v1 [cs.SE] 14 Sep 2026

Nanjing University [email protected], [email protected] * Corresponding author.

2

e

elin

Ov

R MIP

Bas

E

AC

PA

GE

GE

R FO

Multi-hop Retrieval: HotPotQA

e

elin

Bas

MIP

2

v RO

E

AC

GE

PA

E

RG FO

Domain Knowledge: FiNER-139

80 76.0 74.0 70.0

70

69.0

68.0

68.0

67.0

66.0

60 56.0 52.0

50 e

elin

Bas

2

Ov

R MIP

E

AC

PA

GE

GE

R FO

e

elin

Bas

MIP

2

v RO

E

AC

GE

PA

E

RG FO

Figure 1: Results on representative benchmarks from four task categories. Forge consistently outperforms all compared methods across multi-hop retrieval, domain knowledge, reasoning, and agent tasks, demonstrating its effectiveness across diverse benchmarks.

Introduction Language-model (LM) programs combine prompted calls with retrieval, tools, and intermediate computation for applications ranging from structured prediction and knowledgeintensive reasoning to instruction following and interactive agents (Opsahl-Ong et al. 2024; Trivedi et al. 2024). Because their behavior is largely specified by natural-language prompts, automatic prompt optimization (APO) replaces manual design with feedback-driven search. Existing methods use textual gradients, optimizer LMs, joint instruction– demonstration search, or evolutionary reflection (Pryzant et al. 2023; Yuksekgonul et al. 2025; Yang et al. 2024a; Opsahl-Ong et al. 2024; Fernando et al. 2024; Agrawal et al. 2026). Despite these advances, most general-purpose APO systems still derive feedback from fixed training data (OpsahlOng et al. 2024; Agrawal et al. 2026; Zhang et al. 2026). As prompts evolve, repeatedly optimizing on the same instances can overemphasize observed failures while leaving nearby

conditions unexplored. When performance stops improving, it is therefore unclear whether prompt search has exhausted useful revisions or the training data have stopped exposing informative failures. This motivates treating training data as an adaptive optimization state. Feedback-directed synthesis has been studied mainly for model training: early methods bootstrap from seed tasks or model priors (Wang et al. 2023; Xu et al. 2025), whereas learner-aware methods generate instances from observed errors or task-model feedback (Lee et al. 2024; Menon and Srivastava 2024; Li et al. 2025; Khan et al. 2025). Recent work couples data with prompt or model updates: a human-in-the-loop workflow jointly revises prompts and a living test set (Lee and Kahng 2026), while CoEvolve synthesizes interaction tasks during weight optimization (Yang et al. 2026). Together, they motivate a complementary question: can automated reflective search over discrete prompts

for heterogeneous LM programs treat training data as a mutable optimization state? We also ask whether data produced inside this loop remain useful for subsequent LLM weight optimization. We address this problem with Forge, a failure-guided framework that co-evolves prompts and training data. Its key insight is that each failure reveals both how to revise the prompt and what training evidence is missing. Forge uses failed executions for prompt reflection and abstracts their recurring weaknesses into a failure memory. When the best aggregate validation score stops improving, active failure modes guide data synthesis, completing the prompt–data loop in Figure 2. Forge expands active failure modes through four mutation intents: positive preserves the target skill under benign variation, negative creates contrasts where the failed heuristic should not apply, boundary probes nearby decision conditions, and stress adds distractors or interacting constraints. A separate tool-capable verifier agent admits only valid instances faithful to both the targeted failure mode and mutation intent. Accepted instances enter the training data and shape subsequent prompt search, whose failures refresh the memory and guide the next synthesis event. We evaluate Forge on prompt optimization and data transfer; Figure 1 shows representative results across four task categories. Across eight benchmarks, Forge raises average performance by 16.52 points over the unoptimized baseline and by 4.92–8.82 points over competing APO methods. In a separate transfer study, its synthesized instances improve all nine prompt-optimizer and all three GRPO comparisons, including 4–8-point GRPO gains, demonstrating utility beyond Forge. Our contributions are fourfold: • We treat training data as an adaptive state of APO and formulate a failure-mediated loop that co-evolves prompt candidates and training data. • We introduce failure-mode-guided synthesis with positive, negative, boundary, and stress mutations. • We evaluate Forge against strong APO baselines across eight heterogeneous LM-program benchmarks. • We use a paired GRPO study to show that Forgegenerated data transfer to weight optimization and improve final performance under matched optimization budgets.

Related Work Automatic prompt and context optimization. Automatic prompt optimization (APO) uses task feedback to search natural-language instructions and contextual controls. Representative methods translate errors into language feedback (ProTeGi, TextGrad), condition optimizer LMs on scored histories (OPRO), or evolve and revise prompts through mutations and failure patterns (PromptBreeder, AMPO) (Pryzant et al. 2023; Yuksekgonul et al. 2025; Yang et al. 2024a; Fernando et al. 2024; Yang et al. 2024b). For multistage LM programs, MiproV2 searches instructions and demonstrations, Gepa combines trace reflection with Pareto

selection, AutoPDL uses successive halving, and Ace curates a structured playbook (Opsahl-Ong et al. 2024; Agrawal et al. 2026; Spiess et al. 2025; Zhang et al. 2026). Despite their different search strategies, these methods keep training data fixed; Forge instead allows failures to expand the data that drive subsequent search. Data-augmented prompt optimization. Recent systems adapt examples or evaluation artifacts alongside prompts: PromptWizard refines instructions and synthetic demonstrations, Promptomatix synthesizes task-specific data from descriptions, and Data-Prompt Co-Evolution jointly revises prompts and a living test set with developers (Agarwal et al. 2025; Murthy et al. 2025; Lee and Kahng 2026). SIPDO organizes synthesis with a controlled difficulty tier and progressively challenges the prompt using generated examples (Yu et al. 2026). CASPER optimizes prompt representations in a continuous embedding space using textual feedback and augments optimization with failure cases (Jain, Ghosh, and Yenigalla 2026). Forge instead turns recurring failures into tool-grounded synthesis and verification targets for heterogeneous LM programs, then evaluates the resulting instances under other prompt optimizers and GRPO. Learner-aware training-data synthesis. Self-Instruct, Magpie, CodecLM, Evol-Instruct, and RandomWorld bootstrap, mutate, or procedurally generate instruction and tooluse data without adapting to learner failures (Wang et al. 2023; Xu et al. 2025; Wang et al. 2024b; Xu et al. 2024; Sullivan, Hartmann, and Koller 2025). Learner-aware methods instead expand mispredictions (LLM2LLM), summarize error clusters (DISCERN), select or synthesize from weakness profiles (DataEnvGym, STAT), or search for failureinducing queries (ReverseGen) (Lee et al. 2024; Menon and Srivastava 2024; Khan et al. 2025; He et al. 2026; Li et al. 2025). Executable self-play further couples synthesis with weight updates by validating generated tasks through code execution (Absolute Zero) or environment interaction (CoEvolve) (Zhao et al. 2025; Yang et al. 2026). Forge instead uses execution failures to couple discrete prompt reflection with data synthesis across heterogeneous LM programs, then evaluates whether its instances transfer to other prompt optimizers and GRPO-based weight optimization.

Problem Formulation LM-program prompt optimization. Let Fπ denote an LM program whose underlying model weights are frozen and whose M optimizable module prompts are collected in π = (π (1) , . . . , π (M ) ). A task instance z ∈ Z may represent either a conventional input or an interaction with a task environment. Executing Fπ on z produces an ordered message trace τ , task feedback f , and a normalized score r(Fπ , z) ∈ [0, 1]; we call the execution imperfect when r(Fπ , z) < 1. Fixed-data APO. Given initial training data D0 , a validation set V , and a test set T , an APO algorithm draws update minibatches from D0 . Let ΠK contain the prompt candidates it has evaluated by stopping time K. The final prompt

Figure 2: Overview of Forge. Execution failures drive prompt reflection and update failure memory; when aggregate validation performance stops improving, active modes guide four verified mutations whose accepted instances re-enter training.

is selected by the fixed validation objective 1 X JV (π) = r(Fπ , z), |V | z∈V

b = arg max JV (π). π

for later synthesis. Prompt search thus changes program behavior, while synthesis changes the evidence for subsequent updates (Figure 2).

π∈ΠK

(1) The validation set may control search and candidate selection but never provides training or synthesis instances; test scores never affect either process and are used only for reporting. Conventional fixed-data APO additionally imposes Dt = D0 throughout search. Adaptive-data APO. We retain the same validation objective while allowing training data to evolve. At iteration t, let Πt be the candidate pool, Dt the training multiset, and Ht the history of imperfect-execution observations o = (z, π, τ, f, r). Abstracting away the particular search and synthesis mechanisms, the coupled updates are (Πt+1 , Ht+1 ) = U(Πt , Ht ; Dt ), Dt+1 = Dt ⊎ A(Ht+1 ).

(2)

where U denotes a prompt-search update over the current training data, A returns accepted synthetic instances, and ⊎ denotes multiset addition. This update is append-only: accepted synthetic instances may be added, while existing training instances are neither removed nor modified. Fixeddata APO is the special case A ≡ ∅. Thus data synthesis changes the evidence available to subsequent prompt updates without changing the validation objective.

Method Forge couples prompt optimization and training-data synthesis through execution failures while maintaining prompt candidates Πt , mutable training data Dt , and failure memory Mt . Each imperfect execution both guides an immediate prompt revision and becomes a reusable failure mode

Optimization Loop At iteration t, Forge selects a parent prompt π ∈ Πt , samples a minibatch B ⊂ Dt , and executes the LM program to collect scores, ordered traces, and task feedback. Each imperfect execution then branches in two directions. The prompt branch reflects on the observed trace and feedback to propose a revised prompt π ′ ; the memory branch uses the same evidence to update failure memory, as described in the next subsection. Using the same failures in both branches keeps prompt revision and data synthesis focused on weaknesses actually exposed by the current program. The prompt branch uses an independent reflective-search backbone inspired by Gepa (Agrawal et al. 2026). The parent and proposal are compared on the same minibatch, and the proposal enters the candidate pool only when it improves the aggregate minibatch score. Each admitted candidate is then evaluated on the fixed validation set: per-instance scores support frontier-based parent selection, while aggregate scores govern synthesis scheduling and final prompt selection. The data branch becomes active after p completed promptsearch steps without a strict improvement in the best aggregate validation score observed so far. Forge uses the active failure memory to synthesize a budget of new instances, admits only verified instances, and appends them to the training data to obtain Dt+1 . Prompt search then continues on the expanded data. The new instances can expose failures absent from Dt ; those failures in turn revise later prompts and determine the next synthesis targets. This closes the optimization loop. Algorithm 1 summarizes the procedure; further algorithmic and implementation details are provided in Appendix A.

Algorithm 1 Forge: failure-guided prompt and data optimization Require: training data D, validation set V , seed prompt π 0 , patience p, synthesis budget K, stopping rule S 1: Π ← [π 0 ]; evaluate π 0 on V and store scores SV ; M ← ∅ 2: while S is not met do 3: π ← SelectCandidate(Π, SV ) 4: B ← SampleBatch(D) 5: E ← Evaluate(π, B); E − ← {ei ∈ E : ri < 1} 6: M ← UpdateFailureMemory(M, E − ) 7: π ′ ← UpdatePrompt(π, E − ) 8: E ′ ← Evaluate(π ′ , B) 9: if RB (π ′ ) > RB (π) then 10: append π ′ to Π; evaluate it on V and update SV 11: end if 12: if no strict aggregate-validation gain for p steps then 13: D ← D ⊎ Synthesize(M, K) 14: end if 15: end while 16: return best prompt on V and augmented D

Failure Memory Directly asking an LM to synthesize from a single failed instance can easily collapse into near-duplicate generation: the model preserves the same entities, input structure, and reasoning pattern while changing only surface wording. Such variants add little new evidence for prompt optimization because they restate a condition already represented in the training data instead of probing when the failed behavior should or should not occur. The central challenge is therefore to separate the reusable failure mechanism from the particular instance that exposed it. Forge addresses this challenge with a failure memory whose entries are failure modes: reusable mechanism descriptions paired with the concrete failed executions that support them. For every imperfect execution, an LM-based assigner examines the task instance, trace, and feedback. It then assigns the execution to an existing mode, creates a new mode, or refines an existing one. The mode description abstracts away instance-specific entities, while its supporting executions retain the task context and trace evidence needed to ground later generation. This separation enables synthesis to vary the conditions around a weakness rather than merely paraphrasing the source failure. Two failures observed during an AppWorld optimization run (Trivedi et al. 2024) show why modes are more useful than task-specific records. Asked to follow every classical Spotify artist with at least 16 followers, the agent followed five artists but missed a sixth because it inspected only the default result page. The same mode later absorbed a Venmo failure: when asked to accept all pending requests from roommates and coworkers, the agent accepted only two of ten qualifying requests after again processing an incomplete result set. Forge records both executions as incomplete paginated search: the agent mistakes the default result page for the complete candidate set, so qualifying records on later pages

are never considered or acted upon. This abstraction discards app-specific entities and actions while preserving the missing operation, allowing the mode to guide synthesis across unrelated tasks. The modes accumulated since the preceding synthesis event form the active memory used by the next event. This keeps synthesis aligned with weaknesses exposed by the current prompt–data state: when later prompts and added data reveal a different mechanism, the new failures induce different modes rather than forcing generation to revisit the same source instances.

Failure-Guided Data Synthesis When synthesis is triggered, Forge allocates its generation budget across active failure modes, prioritizing modes supported by more observations. Each generation attempt is conditioned on both a mode description and one of its supporting failed executions, including the trace and feedback. The mode specifies what weakness to probe, while the execution anchors the generated instance in a concrete task setting. This differs from asking an LM to produce generically difficult data or merely paraphrase a failed input. A failure mode identifies what behavior is wrong, but not where its correction should and should not generalize. If synthesis produces only positive variants, the optimizer may overgeneralize the apparent fix or learn a brittle rule tied to the source instance. Forge therefore constructs a local behavioral neighborhood around each failure mode: positive mutations test whether the correction transfers across new realizations, negative mutations mark where it should not apply, boundary mutations isolate the condition that changes the desired behavior, and stress mutations test whether the correction remains effective under distractors. Together, these four directions define the scope and robustness of a correction rather than merely diversifying its surface form. A recorded execution from our HotPotQA optimization run makes these roles concrete. The program is asked which actor portrayed Marty McFly in the Back to the Future trilogy and co-hosted the Primetime Emmy ceremony at which A&E and AMC received their first major nominations. The retrieved evidence states that Michael Andrew Fox is known professionally as Michael J. Fox and that Michael J. Fox co-hosted the 48th Primetime Emmy Awards. The program therefore identifies the correct person but returns Michael J. Fox, whereas the benchmark gold answer is Michael Andrew Fox. Failure memory records this execution as noncanonical answer text: the reasoning chain reaches the correct entity, but the returned alias does not match the required answer form. From this recorded failure, the same run generated the following four mutations: • Positive—preserve the requirement. A new question asks for the full biographical name of the actor who played Agent J in Men in Black and hosted a ceremony featuring Bon Jovi and Toni Braxton. The answer is Willard Carroll Smith II.; the entity and context change, but the required canonical answer form is preserved. • Negative—contrast the requirement. A question about the actress born Caryn Elaine Johnson instead asks for

the professional name of the host of the Academy Awards ceremony honoring films released in 1993. The answer is Whoopi Goldberg, demonstrating a case in which returning the full biographical name would be incorrect. • Boundary—change the target. Keeping Michael Andrew Fox and the same Emmy clues, a boundary question explicitly asks for the name by which the actor is known professionally. This single change makes Michael J. Fox correct and isolates the requested name type as the condition controlling the answer form. • Stress—add distractors. A fourth question asks for the exact full name of the musician who wrote “Manic Monday” under the pseudonym “Christopher.” The context adds articles about several other people named Prince, but the answer must remain Prince Rogers Nelson, testing whether the correction survives misleading name matches. Thus the four operators probe invariance, contrast, decision boundaries, and robustness rather than producing interchangeable paraphrases. The stored instances retain their full contexts and supporting facts; the questions above are shortened for presentation. Both the generator and verifier are multi-step agents equipped with task-specific tools. On HotPotQA, both access search_full_wiki, the same full-Wikipedia ColBERT index used by the task program: the generator constructs instances and gold supervision from retrieved passages, while the verifier checks the evidence and can search further. This grounds synthesis and verification in benchmark evidence rather than the optimizer model’s parametric knowledge. Because tool grounding alone does not guarantee quality, the verifier checks each proposal for validity under the task requirements and faithfulness to the targeted failure mode and mutation direction. Only proposals passing both checks enter Dt and provide new evidence for prompt reflection and failure memory. Additional optimization artifacts, including prompt evolution, candidate lineage, and admitted synthetic instances, are provided in Appendix C.

Experiments Our evaluation addresses three questions. First, does Forge improve prompt optimization across benchmarks and task models? Second, is failure-guided synthesis necessary, and how does the synthesis process reshape the failure distribution? Third, do the resulting synthesized instances benefit optimization methods beyond Forge?

Experimental Setup Benchmarks. We evaluate eight heterogeneous LM benchmarks spanning knowledge-intensive reasoning, instruction following, structured prediction, mathematics, and agents. HotPotQA (Yang et al. 2018) and HoVer (Jiang et al. 2020) test multi-hop retrieval and evidence aggregation, while IFBench (Pyatkin et al. 2025) evaluates generalization to verifiable instructions. FiNER-139 (Loukas et al. 2022) and LawBench (Fei et al. 2024) cover financial entity classification and legal judgment, respectively. AIME-2025 (MathArena

2025) and MMLU-Pro (Wang et al. 2024a) evaluate mathematical and broad knowledge reasoning, and AppWorld (Trivedi et al. 2024) evaluates interactive agents in stateful tool environments. Dataset provenance, split construction and task-specific metrics are provided in the appendix B. Models and fair comparison. Our primary comparison uses Qwen3.5-35B-A3B (Qwen Team 2026) as the task model and GPT-5.5 (OpenAI 2026) as the optimizer model. We additionally repeat the full comparison with the dense Qwen3-8B task model (Yang et al. 2025). We compare Forge against the unoptimized baseline and three automatic prompt optimization methods: MiproV2 (Opsahl-Ong et al. 2024), Ace (Zhang et al. 2026), and Gepa (Agrawal et al. 2026). Within each comparison, all methods use the same task model, training, validation, and test splits, LM-program structure, and task-specific evaluator. Optimized states are selected using validation performance only, and the test set is accessed only for final evaluation. To ensure comparable search effort, we align the nominal task-model evaluation budget across optimizers.

Main Results Forge achieves the strongest aggregate performance with Qwen3.5-35B-A3B. As shown in Table 1, it reaches an aggregate score of 72.18, improving over the unoptimized Baseline by 16.52 percentage points and exceeding the competing APO methods by 4.92 to 8.82 points. Forge obtains the best result on six of eight benchmarks. The gains therefore span single-module prediction, multi-hop retrieval, mathematical reasoning, and interactive tool use rather than concentrating in one task format.

Component Ablation In Table 2, Score denotes the mean across the eight benchmarks in Table 1. Prompt search improves the score from 55.66 to 64.17, whereas adding random synthesis lowers it to 62.38. This contrast shows that merely increasing the training data does not provide the targeted evidence needed for optimization. Failure guidance makes synthesis effective. Adding failure memory raises the score to 68.24, surpassing prompt search alone by 4.07 points, and mutation guidance further improves it to 72.18. Together, these modules outperform random synthesis by 9.80 points and prompt search alone by 8.01 points. Failure memory identifies what weaknesses require new evidence, while mutation guidance determines how to expand them into complementary instances. Because the variants are cumulative, these deltas represent conditional contributions rather than isolated main effects.

Synthesis Accounting and Failure Dynamics Beyond downstream performance, we audit synthesis outcomes on four representative benchmarks in Figure 3. Panel (a) accounts for every proposal: Forge admits 22 of 30 on AIME-2025, 39 of 40 on AppWorld, 80 of 100 on HotPotQA, and 93 of 100 on FiNER-139, yielding admission rates of 73–98%. Of the ten generation failures, nine arise in HotPotQA because the generator returns no usable output, while

Qwen3.5-35B-A3B Baseline MiproV2 Ace Gepa Forge

HotPotQA

IFBench

HoVer

FiNER-139

LawBench

AIME-2025

MMLU-Pro

AppWorld

Aggregate

Improvement

52.00 68.00 70.00 66.00 76.00

31.29 48.64 61.63 48.64 54.52

45.00 52.00 55.00 51.00 60.00

56.00 67.00 69.00 68.00 74.00

34.00 38.00 54.00 50.00 57.00

73.33 75.56 71.11 77.78 82.22

87.00 90.00 83.00 90.00 89.00

66.64 67.67 74.33 71.21 84.69

55.66 63.36 67.26 65.33 72.18

– +7.70 +11.60 +9.67 +16.52

Table 1: Final test performance (%) with Qwen3.5-35B-A3B. Aggregate averages the eight benchmarks; Improvement is relative to the Baseline. Bold denotes the column best. Proposal outcome Verification failed Verifier rejected

Generation failed

Failures per mode 1

Admitted

(b) First synthesis window Share of failures (%) (log scale)

(a) Generation and admission 22/30 (73%)

AIME-2025

39/40 (98%)

AppWorld HotPotQA

80/100 (80%)

FiNER-139

93/100 (93%)

0

20

40

60

80

100

120

5 modes

21 modes

19 modes

orld

tQA

5

10

20

(c) Final synthesis window

12 modes

12 modes

10 modes

11 modes

ER

E AIM

orld

tQA

13 modes

50 20 10 5 2 1 E

AIM

W App

tPo

Ho

FiN

W App

tPo

Ho

ER

FiN

Figure 3: Synthesis accounting and failure-memory dynamics on four representative benchmarks. The statistics are computed from the same Forge runs reported in the main comparison (Table 1). Panel (a) partitions all generation attempts by final outcome; labels report admitted proposals over total attempts and the corresponding admission rate. Panels (b) and (c) compare failure modes collected before the first and final synthesis events. Each point denotes a task-specific failure mode, vertical position gives its share of failures in the corresponding window on a logarithmic scale, and area gives the number of supporting executions; horizontal jitter has no semantic meaning. The comparison highlights substantial redistribution, more comparable support across modes, and benchmark-dependent increases or decreases in mode count.

Variant Baseline + Prompt search + Random synthesis + Failure memory + Mutation guidance

Prompt Random Failure Mutation Score search synthesis memory guidance ✓ ✓ ✓ ✓

✓ ✓ ✓

✓ ✓

✓

∆

55.66 – 64.17 +8.51 62.38 -1.79 68.24 +5.86 72.18 +3.94

Table 2: Cumulative ablation averaged over the eight benchmarks in Table 1. Each row adds one component; ∆ is relative to the preceding row. Random synthesis generates from randomly selected training instances.

gets newly exposed weaknesses rather than repeatedly revisiting initial failures. Within each benchmark, failure counts become more balanced across modes: from the first to the final synthesis window, the ratio between the failure counts of the most and least frequent modes decreases from 11:1 to 4:1 on AIME-2025, 20:1 to 6:1 on AppWorld, 22:1 to 11:1 on HotPotQA, and 8:1 to 4:1 on FiNER-139. Meanwhile, mode counts increase from 5 to 12 on AIME-2025 and from 12 to 13 on FiNER-139, but decrease from 21 to 10 on AppWorld and from 19 to 11 on HotPotQA. Thus, the prompt–data loop changes both the composition and balance of failure memory while adapting its granularity to each benchmark.

Generalization Across Task Models the remaining AIME-2025 failure results from an invalid JSON escape. One additional HotPotQA verification returns non-Boolean validity and faithfulness decisions; these eleven cases reflect malformed outputs rather than quality-based rejection. Among the 25 successfully verified but rejected proposals, 9 fail validity only, 7 fail faithfulness only, and 9 fail both. Validity failures primarily involve incorrect supervision or incomplete evidence, whereas faithfulness failures reflect the wrong mutation direction or drift from the targeted failure mode. Panels (b) and (c) compare failure memory before the first and final synthesis events. Both mode identities and frequencies change substantially, indicating that later synthesis tar-

The gains of Forge persist when the task model changes from the mixture-of-experts Qwen3.5-35B-A3B to the dense Qwen3-8B. On Qwen3-8B, Forge leads six of eight benchmarks and achieves an aggregate score of 50.97, exceeding the Baseline by 14.10 points and the strongest competing optimizer, Gepa, by 5.38 points (Table 3). Together with its 16.52-point gain over the Baseline and 4.92-point lead over the strongest competing APO method on Qwen3.5-35BA3B, these results suggest that the advantage of Forge is consistent across dense and mixture-of-experts architectures. We further test the robustness of Forge to optimizer-model choice in Appendix D.

Qwen3 8B

HotPotQA

IFBench

HoVer

FiNER-139

LawBench

AIME-2025

MMLU-Pro

AppWorld

Aggregate

Improvement

Baseline MiproV2 Ace Gepa Forge

40.00 55.00 61.00 62.00 70.00

27.34 31.75 38.39 32.62 35.81

35.00 47.00 39.00 50.00 57.00

55.00 62.00 66.00 67.00 72.00

24.00 31.00 37.00 30.00 35.00

13.33 20.00 7.78 18.89 23.33

61.00 61.00 63.00 61.00 66.00

39.34 40.64 37.89 43.23 48.64

36.88 43.55 43.76 45.59 50.97

– +6.67 +6.88 +8.72 +14.10

Table 3: Final test performance (%) with dense Qwen3-8B. Aggregate and Improvement follow Table 1; bold denotes the column best.

Transferability of Synthesized Data

Train only

Score (%)

(b) FiNER-139: Test

85

72

80

66

75

60

70

54

Score (%)

(c) HotPotQA: Val

(d) HotPotQA: Test

70

65

65

60

60

55 50

55

(e) LawBench: Val Score (%)

We next test whether Forge-generated data remain useful after changing the optimizer that consumes them. On HotPotQA, FiNER-139, and LawBench, we compare the original training data with the same data augmented by Forgesynthesized instances for three prompt optimizers and for GRPO (Shao et al. 2024). All comparisons use Qwen3.535B-A3B as the task model, and the three prompt optimizers use GPT-5.5 as the optimizer model. Each method keeps its optimization budget fixed across the two data conditions. The two GRPO arms differ only in their training data, using the same number of optimization steps and otherwise identical configurations. Appendix B.5 provides the complete GRPO protocol and configuration. Augmentation improves all twelve benchmark–optimizer pairs, showing that the synthesized instances remain useful beyond Forge.

Train + synthetic

(a) FiNER-139: Val

(f) LawBench: Test

54

45

48

40

42

35

36

30

Weight Prompt Optimization Optimization Dataset

Training data MiproV2 Ace Gepa

GRPO

HotPotQA

Original + Forge data

68 73

70 73

66 75

57 65

FiNER-139

Original + Forge data

67 70

69 71

68 74

69 73

LawBench

Original + Forge data

38 47

54 57

50 55

43 51

Table 4: Test performance (%) with original versus Forgeaugmented training data under matched configurations and budgets. Bold denotes the better condition. Table 4 shows that Forge augmentation improves all twelve dataset–optimizer pairs: the nine prompt-optimization gains span 2–9 points, while GRPO gains at step 100 are 8, 4, and 8 points on HotPotQA, FiNER-139, and LawBench, respectively. These results show that the synthesized instances transfer beyond Forge to other prompt optimizers and modelweight optimization. Figure 4 separates validation and test trajectories for all three GRPO comparisons. Under the same 100-step optimization budget, the augmented arm finishes above originaldata training on all three test sets. The corresponding validation trajectories show that these gains emerge at different stages and need not increase monotonically. Together, the six panels show that Forge-synthesized instances pro-

0

20

40

60

80

100

0

GRPO step

20

40

60

80

100

GRPO step

Figure 4: GRPO validation and test trajectories on FiNER139, HotPotQA, and LawBench. The two arms differ only in whether the training data are augmented with Forgesynthesized instances.

vide reusable training evidence across three distinct weightoptimization tasks.

Conclusion Forge uses failures as shared signals for reflective prompt updates and targeted data synthesis through failure memory, four mutation intents, and agentic verification. Across eight benchmarks, it outperforms the unoptimized baseline and all evaluated APO methods; ablations show that these gains require failure- and mutation-guided rather than random synthesis. Its instances also improve three competing prompt optimizers and GRPO across three transfer benchmarks, demonstrating that the synthesized data remain useful beyond the optimization loop that produced them. Together, these results suggest that prompt search and data construction should be treated as a coupled process: observed failures reveal both how a prompt should change and which training examples should be added next. Future work should broaden this transfer evaluation across additional tasks and model families and disentangle the effects of dataset size, example selection, and sampling distribution.

References Agarwal, E.; Magazine, R.; Singh, J.; Dani, V.; Ganu, T.; and Nambi, A. 2025. PromptWizard: Optimizing Prompts via Task-Aware, Feedback-Driven Self-Evolution. In Findings of the Association for Computational Linguistics: ACL 2025, 19974–20003. Association for Computational Linguistics. Agrawal, L. A.; Tan, S.; Soylu, D.; Ziems, N.; Khare, R.; Opsahl-Ong, K.; Singhvi, A.; Shandilya, H.; Ryan, M. J.; Jiang, M.; Potts, C.; Sen, K.; Dimakis, A. G.; Stoica, I.; Klein, D.; Zaharia, M.; and Khattab, O. 2026. GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning. In The Fourteenth International Conference on Learning Representations. AI-MO. 2024. AIMO Validation AIME. Hugging Face dataset. Revision 13f9e12. DeepSeek-AI. 2026. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. arXiv:2606.19348. Fei, Z.; Shen, X.; Zhu, D.; Zhou, F.; Han, Z.; Huang, A.; Zhang, S.; Chen, K.; Yin, Z.; Shen, Z.; Ge, J.; and Ng, V. 2024. LawBench: Benchmarking Legal Knowledge of Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 7933– 7962. Association for Computational Linguistics. Fernando, C.; Banarse, D. S.; Michalewski, H.; Osindero, S.; and Rocktäschel, T. 2024. Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 13481–13544. PMLR. He, Y.; Panigrahi, A.; Lin, Y.; and Arora, S. 2026. STAT: Skill-Targeted Adaptive Training. In The Fourteenth International Conference on Learning Representations. Jain, A.; Ghosh, P.; and Yenigalla, P. 2026. CASPER: Bridging Discrete and Continuous Prompt Optimization through Feedback-Guided Gradient Descent. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track), 425–437. Association for Computational Linguistics. Jiang, Y.; Bordia, S.; Zhong, Z.; Dognin, C.; Singh, M.; and Bansal, M. 2020. HoVer: A Dataset for Many-Hop Fact Extraction and Claim Verification. In Findings of the Association for Computational Linguistics: EMNLP 2020, 3441–3460. Association for Computational Linguistics. Khan, Z.; Stengel-Eskin, E.; Cho, J.; and Bansal, M. 2025. DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback. In The Thirteenth International Conference on Learning Representations. Lee, M.; and Kahng, M. 2026. Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 833:1–833:17. Association for Computing Machinery. Lee, N.; Wattanawong, T.; Kim, S.; Mangalam, K.; Shen, S.; Anumanchipalli, G.; Mahoney, M.; Keutzer, K.; and Gholami, A. 2024. LLM2LLM: Boosting LLMs with Novel

Iterative Data Enhancement. In Findings of the Association for Computational Linguistics: ACL 2024, 6498–6526. Association for Computational Linguistics. Lee, Y.; Nair, R. S.; Zhang, Q.; Lee, K.; Khattab, O.; and Finn, C. 2026. Meta-Harness: End-to-End Optimization of Model Harnesses. In Third Conference on Language Modeling. Li, Q.; Gao, J.; Wang, S.; Pi, R.; Zhao, X.; Wu, C.; Jiang, X.; Li, Z.; and Kong, L. 2025. Forewarned Is Forearmed: Leveraging LLMs for Data Synthesis through Failure-Inducing Exploration. In The Thirteenth International Conference on Learning Representations. Loukas, L.; Fergadiotis, M.; Chalkidis, I.; Spyropoulou, E.; Malakasiotis, P.; Androutsopoulos, I.; and Paliouras, G. 2022. FiNER: Financial Numeric Entity Recognition for XBRL Tagging. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4419–4431. Association for Computational Linguistics. MathArena. 2025. AIME 2025. Hugging Face dataset. Revision c94da77. Menon, R. R.; and Srivastava, S. 2024. DISCERN: Decoding Systematic Errors in Natural Language for Text Classifiers. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 19565–19583. Association for Computational Linguistics. Murthy, R.; Zhu, M.; Yang, L.; Qiu, J.; Tan, J.; Heinecke, S.; Xiong, C.; Savarese, S.; and Wang, H. 2025. Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models. arXiv:2507.14241. OpenAI. 2026. GPT-5.5 System Card. Opsahl-Ong, K.; Ryan, M. J.; Purtell, J.; Broman, D.; Potts, C.; Zaharia, M.; and Khattab, O. 2024. Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 9340–9366. Association for Computational Linguistics. Pryzant, R.; Iter, D.; Li, J.; Lee, Y.; Zhu, C.; and Zeng, M. 2023. Automatic Prompt Optimization with “Gradient Descent” and Beam Search. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7957–7968. Association for Computational Linguistics. Pyatkin, V.; Malik, S.; Graf, V.; Ivison, H.; Huang, S.; Dasigi, P.; Lambert, N.; and Hajishirzi, H. 2025. Generalizing Verifiable Instruction Following. In Advances in Neural Information Processing Systems, volume 38. Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents. Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y. K.; Wu, Y.; and Guo, D. 2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300. Spiess, C.; Vaziri, M.; Mandel, L.; and Hirzel, M. 2025. AutoPDL: Automatic Prompt Optimization for LLM Agents. In Proceedings of the Fourth International Conference on Automated Machine Learning, volume 293 of Proceedings of Machine Learning Research, 13/1–20. PMLR.

Sullivan, M.; Hartmann, M.; and Koller, A. 2025. Procedural Environment Generation for Tool-Use Agents. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 18544–18562. Association for Computational Linguistics. Trivedi, H.; Khot, T.; Hartmann, M.; Manku, R.; Dong, V.; Li, E.; Gupta, S.; Sabharwal, A.; and Balasubramanian, N. 2024. AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16022– 16076. Association for Computational Linguistics. Wang, Y.; Kordi, Y.; Mishra, S.; Liu, A.; Smith, N. A.; Khashabi, D.; and Hajishirzi, H. 2023. Self-Instruct: Aligning Language Models with Self-Generated Instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13484– 13508. Association for Computational Linguistics. Wang, Y.; Ma, X.; Zhang, G.; Ni, Y.; Chandra, A.; Guo, S.; Ren, W.; Arulraj, A.; He, X.; Jiang, Z.; Li, T.; Ku, M.; Wang, K.; Zhuang, A.; Fan, R.; Yue, X.; and Chen, W. 2024a. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. In Advances in Neural Information Processing Systems, volume 37. Wang, Z.; Li, C.-L.; Perot, V.; Le, L.; Miao, J.; Zhang, Z.; Lee, C.-Y.; and Pfister, T. 2024b. CodecLM: Aligning Language Models with Tailored Synthetic Data. In Findings of the Association for Computational Linguistics: NAACL 2024, 3712–3729. Association for Computational Linguistics. Xu, C.; Sun, Q.; Zheng, K.; Geng, X.; Zhao, P.; Feng, J.; Tao, C.; Lin, Q.; and Jiang, D. 2024. WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions. In The Twelfth International Conference on Learning Representations. Xu, Z.; Jiang, F.; Niu, L.; Deng, Y.; Poovendran, R.; Choi, Y.; and Lin, B. Y. 2025. Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing. In The Thirteenth International Conference on Learning Representations. Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; Zheng, C.; Liu, D.; Zhou, F.; Huang, F.; Hu, F.; Ge, H.; Wei, H.; Lin, H.; Tang, J.; Yang, J.; Tu, J.; Zhang, J.; Yang, J.; Yang, J.; Zhou, J.; Zhou, J.; Lin, J.; Dang, K.; Bao, K.; Yang, K.; Yu, L.; Deng, L.; Li, M.; Xue, M.; Li, M.; Zhang, P.; Wang, P.; Zhu, Q.; Men, R.; Gao, R.; Liu, S.; Luo, S.; Li, T.; Tang, T.; Yin, W.; Ren, X.; Wang, X.; Zhang, X.; Ren, X.; Fan, Y.; Su, Y.; Zhang, Y.; Zhang, Y.; Wan, Y.; Liu, Y.; Wang, Z.; Cui, Z.; Zhang, Z.; Zhou, Z.; and Qiu, Z. 2025. Qwen3 Technical Report. arXiv:2505.09388. Yang, C.; Wang, X.; Lu, Y.; Liu, H.; Le, Q. V.; Zhou, D.; and Chen, X. 2024a. Large Language Models as Optimizers. In The Twelfth International Conference on Learning Representations. Yang, S.; Ma, Z.; Huang, T.; Hu, Y.; Wang, Y.; and Chu, X. 2026. CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution. In Proceedings of the 64th Annual Meeting

of the Association for Computational Linguistics (Volume 1: Long Papers), 23015–23036. Association for Computational Linguistics. Yang, S.; Wu, Y.; Gao, Y.; Zhou, Z.; Zhu, B. B.; Sun, X.; Lou, J.-G.; Ding, Z.; Hu, A.; Fang, Y.; Li, Y.; Chen, J.; and Yang, L. 2024b. AMPO: Automatic Multi-Branched Prompt Optimization. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 20267– 20279. Association for Computational Linguistics. Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W.; Salakhutdinov, R.; and Manning, C. D. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-Hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2369–2380. Association for Computational Linguistics. Ye, H.; He, X.; Arak, V.; Dong, H.; and Song, G. 2026. Meta Context Engineering via Agentic Skill Evolution. In Forty-third International Conference on Machine Learning. Yu, Y.; Yu, Y.; Zhang, P.; Wei, K.; Luo, H.; and Wang, H. 2026. SIPDO: Closed-Loop Prompt Optimization via Synthetic Data Feedback. In The Fourteenth International Conference on Learning Representations. Yuksekgonul, M.; Bianchi, F.; Boen, J.; Liu, S.; Lu, P.; Huang, Z.; Guestrin, C.; and Zou, J. 2025. Optimizing Generative AI by Backpropagating Language Model Feedback. Nature, 639: 609–616. Zhang, Q.; Hu, C.; Upasani, S.; Ma, B.; Hong, F.; Kamanuru, V.; Rainton, J.; Wu, C.; Ji, M.; Li, H.; Thakker, U.; Zou, J.; and Olukotun, K. 2026. Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models. In The Fourteenth International Conference on Learning Representations. Zhao, A.; Wu, Y.; Wu, T.; Xu, Q.; Yue, Y.; Lin, M.; Wang, S.; Wu, Q.; Zheng, Z.; and Huang, G. 2025. Absolute Zero: Reinforced Self-Play Reasoning with Zero Data. In Advances in Neural Information Processing Systems, volume 38.

Appendix Contents A Algorithm and Implementation Details A.1 Optimization State . . . . . . . . . . . . . A.2 Candidate Selection . . . . . . . . . . . . . A.3 Minibatch Sampling . . . . . . . . . . . . . A.4 Failure-Memory Construction . . . . . . . A.5 Reflective Prompt Update . . . . . . . . . . A.6 Failure-Guided Synthesis and Verification .

11 11 11 11 11 12 14

B Experimental Setup B.1 Benchmark Data and Splits . . . . . . . . . B.2 Benchmark Evaluation, Tools, and Synthesis B.3 Model Configuration . . . . . . . . . . . . B.4 Optimization-Budget Alignment . . . . . . B.5 GRPO Transfer Protocol . . . . . . . . . .

14 14 14 17 17 18

C Optimization Artifacts C.1 AppWorld Prompt Evolution . . . . . . . . C.2 HoVer Candidate Lineage . . . . . . . . . . C.3 Admitted Synthetic Examples . . . . . . .

18 18 19 19

D Supplementary Experiments D.1 Optimizer-Model Ablation . . . . . . . . . D.2 Synthetic-Data Leakage Audit . . . . . . .

26 26 26

A

Algorithm and Implementation Details

Algorithm 1 in the main paper gives the complete prompt– data optimization loop but leaves its internal calls abstract. This section instantiates exactly the six operations used in that loop: optimization state, parent-candidate selection, minibatch sampling, failure-memory construction, reflective prompt update, and failure-guided synthesis with verification. We describe decision-level behavior and the modelfacing contracts needed for reproduction, while omitting checkpoint and callback bookkeeping.

A.1

Optimization State

The conceptual optimization state is (Π, D, SV , M). The candidate pool Π = [π 0 , . . . , π q−1 ] stores complete mappings from program-module names to system prompts; an accepted candidate is appended to this pool rather than replacing its parent. The mutable training data D contain the original instances and all admitted synthetic instances. For every candidate i, SV stores its per-instance validation scores si (z) for all z ∈ V . These scores support both parent selection during search and aggregate validation-based model selection. Failure memory M groups imperfect executions into reusable failure modes. A mode is mj = (nj , dj , Oj ), where nj is a short name, dj is a decontextualized mechanism description, and Oj is the ordered collection of observations supporting that description. Each observation retains the parent-candidate index, task instance, scalar score, ordered program trace, and module-level feedback. The implementation stores modes by synthesis epoch: earlier epochs remain available for audit, but only modes collected since the preceding synthesis event are active when constructing the next synthetic batch.

A.2

Candidate Selection

Forge uses validation behavior to select a parent without collapsing every candidate to one aggregate score. For each validation instance z, it forms the local front   Fz = i : si (z) = max sℓ (z) . (3) ℓ

A candidate i is placed in the dominated set G when some candidate ℓ ̸= i satisfies sℓ (z) ≥ si (z) for every z ∈ V and is strictly better on at least one validation instance. The default selector then samples uniformly from the multiset of non-dominated front members, as detailed in Algorithm 2. Equivalently, candidate i has multiplicity X wi wi = 1[i ∈ / G] 1[i ∈ Fz ], P (i) = P . (4) ℓ wℓ z∈V

Thus a candidate that leads on more validation instances is sampled more often, while complete domination removes it from consideration.

A.3

Minibatch Sampling

The default sampler is epoch-shuffled. It stores a shuffled order of training-instance identifiers and a cursor in that order,

Algorithm 2 Front-frequency candidate selection Require: candidate pool Π; per-instance validation scores SV ; validation set V 1: G ← {i : Dominated(i, SV )} 2: L ← [ ] {multiset of eligible front members} 3: for each validation instance z ∈ V do 4: bz ← maxℓ sℓ (z) 5: for each candidate index i = 0, . . . , |Π| − 1 do 6: if i ∈ / G and si (z) = bz then 7: append i to L 8: end if 9: end for 10: end for 11: i⋆ ← UniformChoice(L) 12: return parent candidate Π[i⋆ ]

then returns the next contiguous minibatch at each promptsearch step. When the order is exhausted, all identifiers are reshuffled for a new sampling epoch. If |D| is not divisible by the minibatch size, the final order is padded with randomly selected identifiers so that every returned minibatch is full; only these padding positions may duplicate an instance. A synthesis event refreshes the order without discarding which instances have already been consumed. The sampler first collects unique, not-yet-consumed identifiers from the current order and adds identifiers of newly admitted synthetic instances. It shuffles this group and places it before a separately shuffled group of already-consumed identifiers. Consequently, new data receive the same priority as unseen real data and enter subsequent prompt updates before instances already visited in the current sampling epoch. This sampling epoch is independent of the synthesis epoch used to partition failure memory.

A.4

Failure-Memory Construction

Failure memory is built only from imperfect executions of the selected parent on the current training minibatch. For each evaluation with score below 1, Forge constructs an observation o = (i, z, r, τ, f ) containing the parent index i, instance z, score r, complete ordered program trace τ , and module-level feedback f . The assigner receives two textual inputs: a catalogue containing only the names and descriptions of modes in the current synthesis epoch, and a Markdown rendering of the new observation with sections for the instance, score, trace, and optional feedback. Figure 5 shows the complete assigner prompt and its dynamic user fields. The implementation calls a failure mode a failure cluster; we retain the paper terminology in the prose. The assigner returns one of three structured actions. Assign appends the observation to an existing mode; create introduces a new name and description with the observation as its first support; and rename updates an existing name and description before appending the observation. The rename action is the implementation of the abstraction refinement described in the main paper and may also keep the name unchanged while improving only the description. Failures are processed sequentially, so a later failure in the same mini-

Implementation Prompt: Failure Assigner

SYSTEM

Role You are Forge's failure assignment agent. Forge optimizes prompts by evaluating a candidate AI program, defined by the current candidate system prompts for its modules, on training examples. Your job is to assign one failed evaluation of that candidate program on one example to a reusable failure cluster in failure memory.

Objective

Trace Semantics

Output Contract

The failure trace is an ordered program trace from one evaluation run. It contains one section for each program module invocation in execution order. Each invocation names the module that was called and shows the recorded LLM/tool messages.

Use exactly one of these three legal JSON forms inside <assignment>.

The same module may appear multiple times. Treat repeated module sections as separate invocations. Use invocation order, module names, messages, tool results, score, and feedback together to infer the reusable failure mechanism.

Your goal is to identify why the evaluated candidate program failed on this example at the system-behavior level, then keep the cluster taxonomy stable, compact, and useful for later synthetic data generation.

Decision Procedure

Inputs

2. Compare that mechanism against every existing cluster description.

• existing_cluster_catalog: a JSON array of current failure clusters. Each item contains cluster_name and failure_description.

1. Read the failed candidate-program evaluation and state the most likely system-level failure mechanism.

3. If an existing cluster covers the mechanism, assign to it.

• failure_markdown: one failed evaluation of the candidate program on a single example, including the example, score, ordered program trace, and optional feedback.

4. If an existing cluster covers the mechanism but its name or description would make future assignment or generation worse, rename that cluster and assign to the renamed cluster.

Core Principles

5. If no existing cluster covers the mechanism, create a new cluster.

• Cluster by reusable failure mechanism, not by surface wording, domain topic, specific entity, example id, candidate id, score value, or module name alone.

Hard Constraints

• Prefer assignment to an existing cluster whenever its description already covers the mechanism.

• Do not create a new cluster for a mere paraphrase, different topic, different entity, different answer text, or different trace length when the underlying mechanism is already covered.

• Create a new cluster only when no existing cluster describes the mechanism at the right level of abstraction.

• Do not rename a cluster just to mention this example. Rename only to improve the reusable cluster abstraction.

• Rename a cluster only when the existing cluster is the right destination but its name or description is too narrow, too broad, misleading, or inaccurate.

• Do not use names tied to one benchmark item, one answer, one tool call, one candidate, or one score.

• Keep clusters useful for generation: a future generator should be able to produce more failures from the cluster description without seeing this exact example.

• Failure descriptions must be concise, decontextualized, and describe the recurring mechanism. Do not include candidate ids, example ids, or long quoted task text.

• Cluster names must be short snake_case.

• Your final response must contain exactly one <assignment> XML block. The content of that block must be one JSON object. Do not put markdown fences inside the XML block.

Create a new cluster: <assignment> { "cluster_name": null, "new_cluster_name": "short_snake_case_name", "new_failure_description": "A concise, decontextualized description of the reusable failure mechanism." } </assignment>

Assign to an existing cluster: <assignment> { "cluster_name": "existing_cluster_name", "new_cluster_name": null, "new_failure_description": null } </assignment>

Rename an existing cluster, update its description, then assign this failure: <assignment> { "cluster_name": "existing_cluster_name", "new_cluster_name": "better_short_snake_case_name", "new_failure_description": "A better concise, decontextualized description of the reusable failure mechanism." } </assignment>

USER

Existing Failure Clusters {{existing_cluster_catalog}}

Failure Observation {{failure_markdown}} Decide whether this failure should be assigned to an existing cluster, should create a new cluster, or should rename an existing cluster before assignment. Return the assignment XML block only.

Figure 5: The failure-assigner prompt maps a failed execution to an existing or newly created failure mode.

batch sees every mode created or revised by earlier failures. Malformed JSON, an illegal action-field combination, an unknown destination, or a naming conflict produces a no-op rather than an ungrounded memory update. The resulting catalogue provides stable synthesis targets, while each support set retains the executions that justify its target. After synthesis, Forge advances the synthesis epoch. The stored history is preserved, but the next active catalogue starts empty so that future generation reflects weaknesses exposed by the updated prompt–data state.

A.5

Reflective Prompt Update

The prompt-update branch uses the imperfect parent executions already collected for the current minibatch. Under the default round-robin policy, step t selects one program module, and the reflector receives that module’s current system prompt together with a Markdown bundle containing each failed task input, the trace entries belonging to that module, and its module-level feedback. The reflector proposes one complete replacement system prompt. If this block is absent or empty, the module retains its original prompt. Figure 6 shows this model-facing reflection contract.

Implementation Prompt: Reflective Prompt Update SYSTEM

Role You are Forge's system prompt reflection agent. Forge optimizes prompts by evaluating candidate AI programs on task examples. Your job is to rewrite one module's system prompt using the current prompt, observed module traces, and module-level feedback.

Objective Produce a replacement system prompt that improves the future behavior of this module. The replacement must be directly usable as the module's runtime system prompt, not a commentary about the evaluation batch.

Required Design Of The Output System Prompt

<system_prompt> # Role

The system prompt you write inside <system_prompt> must be a runtime agent specification, not a loose instruction list. Use this top-level Markdown structure:

You are a sentiment analysis classifier in a customer-review analysis pipeline. Your responsibility is to read one review and classify its overall attitude.

• # Role: define what the module is in the target program and what responsibility it owns.

# Objective

• # Objective: define the concrete success criterion for this module. • # Inputs: describe each input field or input convention the module should rely on. Explain what to ignore when input contains distractors or metadata.

• evaluation_markdown: a batch of evaluations. Each evaluation

• # Core Principles: give stable decision principles that guide behavior across examples. These should be reusable, not copied from one failed case. This section may use ## subsections to organize principle groups, such as task principles, reasoning principles, domain principles, tool principles, or verification principles.

includes the task input, this module's ordered trace when available, and feedback for this module when available.

• # Decision Procedure: provide an ordered procedure the module should follow at runtime.

Core Principles

• # Hard Constraints: list non-negotiable rules, including forbidden behavior, formatting constraints, schema constraints, tool-use limits, or safety constraints.

Inputs • current_system_prompt: the system prompt currently used by the target module.

• Preserve useful behavior and task conventions from the current system prompt. • Infer the module's actual task, inputs, outputs, and success criteria from the current prompt and evaluations. • Fix recurring failure mechanisms, missing decision rules, ambiguous instructions, and output-format errors shown by the evidence. • Prefer reusable guidance over example-specific patches. Do not overfit to one entity, one answer, one benchmark item, or one trace. • Include domain-specific rules, edge cases, and reliable procedures only when they are supported by the evidence or clearly required by the task. • Keep the new prompt concise enough to be followed, but concrete enough to change behavior. • Preserve room for the target module to reason before emitting its final structured answer. Never trade away useful reasoning merely to make XML extraction convenient.

Decision Procedure 1. Extract the target module's current runtime contract from current_system_prompt: role, task, inputs, expected outputs, hard constraints, and any useful existing strategy. 2. Read every evaluation and separate the evidence into preserved successes, recurring failures, output-format problems, missing domain rules, ambiguous decision points, and edge cases. 3. Infer the prompt-level cause of the failures, such as a missing rule, vague priority, weak output contract, overgeneralized strategy, absent edge-case handling, or unclear tool-use boundary. 4. Decide what to preserve, strengthen, add, remove, or reorder in the replacement prompt. Prefer changes that address repeated mechanisms and rule boundaries over one-off examples. 5. Draft the replacement as a runtime system prompt using the required output structure. When structured output is required, separate the module's reasoning from its final extractable fields. 6. Check that the replacement prompt is actionable for future examples, follows the target task's signature and output requirements, leaves room for reasoning before the final answer, and does not mention this reflection process or the evaluation evidence. 7. Return the full replacement prompt inside the system_prompt XML block.

• # Output Contract: specify the exact expected final-answer format. If the module uses XML, list the required tags and allowed final values, require reasoning to remain outside those tags, and allow reasoning before the final XML fields. Never require the entire response to contain only XML. If the module emits JSON or a label, specify the schema or label set while still preserving room for preceding reasoning when extraction permits it. Additional task-specific sections are allowed only when they add real guidance, such as # Domain Rules, # Edge Cases, # Tool Use, # Quality Checks, and other task-relevant sections. The output system prompt must be written for the target module at runtime. Do not mention this reflection process, the evaluation batch, traces, feedback, failures, candidates, or prompt optimization unless those concepts are part of the target task itself.

Return the sentiment that best reflects the reviewer's overall attitude, with special attention to contrast, final conclusions, and whether praise or criticism dominates. # Inputs - The review text is provided inside `<review>...</review>`. Use the review content as the evidence for sentiment, and ignore surrounding metadata unless the task says that metadata should affect the label. # Core Principles - Classify the overall attitude rather than isolated sentiment words. - Prefer the final judgment when the review changes direction, such as praise followed by a clear complaint or warning. - Treat sarcasm, contrast words, and explicit recommendations as strong signals. - Work through the evidence before committing to the final label. # Decision Procedure - Read the entire review before choosing a label. - Identify the main praise, criticism, and final conclusion. - Decide whether the positive or negative evidence clearly dominates. Use `neutral` when neither side dominates. - Use `positive` when the review is mostly favorable, praises the product or service, says the customer would recommend it, or describes a problem that was resolved satisfactorily. - Use `negative` when the review is mostly unfavorable, reports dissatisfaction, warns other customers away, or says the product or service failed to meet expectations.

The target module may be a non-thinking model that needs to produce visible reasoning before committing to an answer. Its replacement prompt must not suppress that reasoning with an XML-only response rule. XML fields are final-answer extraction

- Use `neutral` when the review is mixed, purely factual, unclear, or does not express a strong overall attitude.

boundaries, not containers for the module's entire response.

- Do not infer sentiment from usernames, timestamps, product IDs, or boilerplate unless that metadata changes the meaning of the review. - Keep reasoning and explanations outside the

Hard Constraints • You may reason before producing the replacement prompt. Your response must contain exactly one XML block named system_prompt for extraction. • The content inside <system_prompt> must be the full replacement system prompt text in Markdown. • Do not put markdown fences around the replacement prompt. • Treat XML blocks as extraction boundaries, not whole-response formats, for both this reflection and the target module. • Never instruct the target module to output only XML, place its entire response inside XML, or omit reasoning merely because its final answer uses XML fields.

# Hard Constraints

`<sentiment>` field. Do not put confidence scores or extra labels inside that field. # Output Contract - You may reason before the final answer. End the response with a `<sentiment>` XML block containing exactly one of `positive`, `negative`, or `neutral`. </system_prompt>

USER

Current System Prompt

• When the target signature uses XML extraction, allow reasoning before the required final XML fields and keep only the final values inside those fields.

{{current_system_prompt}}

Output Contract

{{evaluation_markdown}}

You may reason before the final result. Then provide exactly one XML block named system_prompt for extraction. The following example only illustrates the expected structure and XML convention. Do not copy its task content, labels, or domain rules.

Write a better system prompt for this module.

Evaluation results

Figure 6: The reflective prompt-update prompt produces a complete replacement system prompt from imperfect module executions.

The proposal copies all unselected module prompts from the parent and replaces only the selected module. Although reflection uses only imperfect evaluations, parent and pro-

posal are screened on the same complete minibatch. The proposal is appended to Π only when its total minibatch score is strictly higher; an admitted proposal is then evaluated on all

Algorithm 3 Failure-guided synthesis, verification, and admission Require: active modes M = {(nj , dj , Oj )}; task signature Σ; attempt budget K; training data D 1: Ω ← (Positive, Negative, Boundary, Stress) 2: (k1 , . . . , k|M| ) ← LargestRemainder(K, (|Oj |)j ) 3: A ← ∅; c ← 0 4: for each mode mj = (nj , dj , Oj ) ∈ M do 5: for a = 1, . . . , kj do 6: o ← UniformChoice(Oj ) {sampling with replacement} 7: u ← Ω[1 + (c mod |Ω|)]; c ← c + 1 8: (e z , τG ) ← Generate(Σ, mj , o, u) 9: if Coerce(e z , Σ) = ⊥ then 10: continue {skip verifier on parse/schema failure} 11: end if 12: C ← (mj , u, o.instance, o.feedback) 13: append generator assistant/tool activity from τG to C 14: (v, f ) ← Verify(Σ, ze, C) 15: if v = true and f = true then 16: A ← A ⊎ {e z} 17: end if 18: end for 19: end for 20: D ← D ⊎ A; RefreshSampleOrder(D); advance synthesis epoch 21: return admitted synthetic data A

of V and added to SV . This local paired comparison prevents an edit from being accepted merely because it was tested on an easier batch.

A.6

Failure-Guided Synthesis and Verification

The exact model-facing generator and verifier contracts are shown in Figures 7 and 8; the admission logic connecting them is specified below. A synthesis event is triggered after p completed promptsearch steps without a strict improvement in the historical best aggregate validation score. The event budget K counts generation attempts, not P successfully parsed or finally admitted instances. Let N = j |Oj | be the number of observations supporting the active modes. Forge first assigns   |Oj | (0) (5) kj = K N attempts to mode j, then distributes the remaining attempts in descending order of the fractional remainders, breaking equal remaindersPby mode order. This largest-remainder allocation preserves j kj = K while prioritizing recurrent failures. Within each allocated slot, Forge samples one supporting observation uniformly with replacement. It assigns mutation intents from the global cycle Ω = (Positive, Negative, Boundary, Stress) across the complete seed list, restarting from Positive at each synthesis event. Algorithm 3 gives the complete admission path.

The generator is a multi-step agent conditioned on the JSON task signature and a mutation seed containing the mode name and description, mutation intent, failed instance, feedback, and full source program trace. It may call benchmarkprovided tools when factual lookup, calculation, environment state, or schema details require grounding. Its final XML field must contain a JSON object. The implementation parses that object and coerces all declared fields and types before verification; a failure at this stage is rejected immediately. The verifier is a second multi-step agent equipped with the same benchmark generation tools. It receives the task signature, parsed synthetic instance, and a flattened context containing the mode, intent, failed instance, source feedback, and the generator’s assistant/tool activity. It does not directly receive the source program trace or hidden generator reasoning. The verifier independently judges validity— schema compliance, task coherence, and gold-supervision correctness—and faithfulness to both the failure mode and mutation intent. Only two successfully parsed Boolean values that are both true admit an instance.

B

Experimental Setup

This section specifies the data, evaluators, models, promptoptimization budget alignment, and the separate paired GRPO transfer protocol used in the main paper.

B.1

Benchmark Data and Splits

The comparison protocol assigns the same ordered training, validation, and test examples to every optimizer for a given benchmark and task model. Unless a benchmark defines a fixed split, randomized construction uses the stated local seed and produces disjoint partitions. Table 5 summarizes the exact construction and number of examples in each benchmark split.

B.2

Benchmark Evaluation, Tools, and Synthesis

Under this protocol, every comparison method calls the same benchmark adapter and receives the same scalar score for a given program execution. Validation scores select optimized states, whereas test outcomes are retained only for reporting and never provide feedback to optimization. For each benchmark, the reported percentage is 100 times the arithmetic mean of its per-example test scores; the aggregate in the main tables is the unweighted mean of the eight benchmark percentages. We report adapter-provided resources alongside each evaluator because they constrain optimization-side generation and verification; unless stated otherwise, they do not change the scalar metric. Credentials and machine-specific configuration are not embedded in any prompt. Synthesis schedule. Following Appendix A.6, we denote each benchmark’s configuration by (K, E, P ), where K is the number of generation attempts requested at each synthesis event, E is the maximum number of synthesis epochs, and P is the number of completed prompt-search steps without a strict improvement in the historical best aggregate validation score that triggers the next event. The configured tuples are AIME (10, 3, 15); AppWorld (10, 4, 15); FiNER139 (25, 4, 15); HotPotQA (25, 4, 15); HoVer (25, 4, 15);

Implementation Prompt: Synthetic-Data Generator Faithfulness

Hard Constraints

Role

• The synthetic example must be faithful to the failure cluster as interpreted through the requested mutation kind.

• Do not omit any input or output field declared by the task signature.

You are Forge's synthetic training data generator. Forge optimizes prompts and evolves the training set by generating synthetic examples from failure-derived mutation seeds at epoch boundaries.

• Use the failed example, feedback, and full system trace to understand the mechanism, but do not copy the failed example verbatim.

• Do not use placeholder values such as "N/A", "unknown", "TODO", or dummy labels unless such values are genuinely correct for the task.

• Preserve the failure-relevant structure while changing surface content, entities, wording, numbers, options, or context enough to make a genuinely new example.

• Do not include the mutation seed, failure cluster name, trace, feedback, explanation, or any extra metadata as JSON fields.

• Do not generate an unrelated easy example merely because it matches the schema.

• Do not leak that the example is synthetic, generated, mutated, or based on a failure observation unless that is part of the task itself.

SYSTEM

Objective Create one valid synthetic training example that matches the provided task signature and acts as a Contrastive Failure Mutation for the provided mutation seed. The synthetic example should help future prompt optimization learn the rule boundary behind the failure cluster. Generate an example that is valid, faithful to the mutation seed, and suitable to enter the training set.

Inputs • task_signature: a JSON description of the required example shape. It lists declared input fields and output/gold fields, with names, types, and descriptions. • mutation_seed: a failure-derived seed containing the failure cluster name, failure description, mutation kind, failed example, feedback, and full ordered program trace.

Core Principles

Contrastiveness • Respect the requested mutation kind: • positive: generate another example that requires the same robust strategy or correction as the failed example.

• Do not wrap the JSON in markdown fences inside the XML block.

• negative: generate a surface-similar sibling where the same repair strategy should not be triggered.

Output Contract

• boundary: change one key condition so the correct behavior changes at the boundary of the failure mechanism. • stress: keep the failure mechanism relevant while adding distractors, extra constraints, similar entities, noisy context, or harder reasoning.

Validity

• The content of the synthetic example should clearly reflect the requested mutation kind.

• The generated example must be a real task example that can be consumed by the downstream solver/evaluator.

• Avoid mere paraphrases. A good mutation changes the conditions that matter for learning the rule boundary.

• Include every field declared in the task signature, using exactly the declared field names.

Tool Use

• Treat output fields as gold supervision fields. Fill them with correct target values, labels, reference reasoning, metadata, or other gold information required by the signature. • Match declared field types. For example, output JSON strings for str, integers for int, floats for float, booleans for bool, arrays for list[...], and objects for dict[...]. • Include only the fields declared by the task signature. Do not add metadata, rationale, source tags, debug fields, or any other extra fields.

Failure Cluster And Mutation Kind • The failure cluster is the anchor: it defines the reusable failure mechanism or rule boundary that this synthetic example should help teach. • The mutation kind defines how the new example should relate to that anchor. • For positive and stress, the example should instantiate or pressure-test the same failure mechanism. • For negative and boundary, the example should stay close to the failure mechanism while changing key conditions so the correct behavior differs. These examples are faithful when they clarify the limits of the failure cluster rather than drifting to an unrelated task.

• Do not create examples whose gold answer cannot be justified from the example content or verified evidence.

You may reason, inspect the seed, and call tools before producing the final XML block. Your final response must contain exactly one XML block named synthetic_example. The content inside <synthetic_example> must be one valid JSON object representing the generated example. The JSON object must include every field declared in the task signature and no extra fields. Concrete example, if task_signature is: { "name": "SyntheticQASignature", "input_fields": [ { "name": "question", "type": "str", "description": "Question to answer." } ], "output_fields": [ { "name": "answer", "type": "int", "description": "Final answer." } ]

• If tools are available, use them when correctness depends on factual lookup, calculation, environment state, schema details, or domain knowledge. • Do not invent external facts when a tool could verify them. If tools are unavailable or unnecessary, prefer self-contained examples whose gold fields can be derived from the provided content. • Treat tool results as evidence for constructing the example, not as text that must be copied into the final JSON unless the task requires it. }

Decision Procedure 1. Read the task signature and identify all required field names, field types, and gold supervision fields. 2. Read the failure cluster, failure description, failed example, feedback, and trace to infer the underlying reusable failure mechanism. 3. Decide how the requested mutation kind should alter the failed example while preserving the intended relationship to the failure cluster. 4. Construct a new example that is valid for the task and whose gold fields are correct.

then a valid final XML block is: <synthetic_example> { "question": "What is 17 + 28?", "answer": 45 } </synthetic_example>

USER

Task Signature JSON

5. Check that the example is not a verbatim copy, not a pure paraphrase, and not about an unrelated failure mechanism.

{{task_signature}}

6. Check that the JSON object includes all declared fields, uses correct types, and can be parsed without comments or markdown.

{{mutation_seed}}

Mutation Seed Generate one synthetic training example.

Figure 7: The synthetic-data generator prompt produces a failure-guided training example from a task signature and mutation seed. IFBench (25, 4, 15); LawBench (25, 4, 15); and MMLU-Pro (25, 4, 15). Thus, the configured attempt ceilings are 30 for AIME, 40 for AppWorld, and 100 for each remaining

benchmark. The value K counts mutation-seed generation attempts, not admitted samples; parsing and verification failures can reduce admissions.

Implementation Prompt: Synthetic-Data Verifier SYSTEM

Role

• The synthetic example should preserve the failure-relevant mechanism while changing enough content to be a genuinely new training example. • The example should reflect the requested mutation kind:

You are Forge's synthetic training data verifier. Forge optimizes prompts and evolves the training set with generated examples. Your responsibility is to inspect one generated example and decide whether it is acceptable training data for the target task.

Objective Evaluate the synthetic example on two independent properties: • validity: whether the example is a well-formed, task-correct training example that matches the task signature. • faithfulness: whether the example faithfully follows the failure-derived context and requested mutation behavior described in the generation context. The example should enter the training set only when both properties are true. Your output must expose both judgments separately.

Inputs • task_signature: a JSON description of the expected example shape. It lists declared input fields and output/gold fields, with names, types, and descriptions. • synthetic_example: the generated example to verify. Treat output fields as gold supervision fields that must be correct, not as model predictions to be excused. • generation_context: a curated, flattened context for the generation attempt. It includes the failure cluster, failure description, mutation kind, failed example, feedback, and any generator assistant/tool activity useful for verification.

Core Principles Validity • Check that the synthetic example contains every field required by the task signature. • Check that each field uses the declared type and a format the downstream solver/evaluator can consume. • Check that output/gold fields are correct for the example content. A schema-valid example is not valid if its gold answer, label, reference output, or metadata is unsupported or wrong. • Check task-specific constraints implied by the signature, such as allowed labels, option sets, formatting rules, answer units, entity spans, or exact-output requirements. • Reject placeholder, underspecified, contradictory, unanswerable, or malformed examples.

Faithfulness • Use the explicit sections in the generation context to identify the failure cluster, failure description, failed example, feedback, mutation kind, and any tool-supported facts.

Hard Constraints • Do not rewrite, repair, or regenerate the synthetic example.

• positive: another example requiring the same robust strategy or correction.

• Do not reward effort, style, or plausibility when the task answer or label is wrong.

• negative: a close sibling where the same repair strategy should not be triggered.

• Do not infer missing fields, hidden labels, unstated facts, or unsupported gold answers.

• boundary: a nearby case that changes a key condition at the edge of the failure mechanism.

• Do not require the synthetic example to mention the failure cluster, mutation kind, or synthetic-data process unless the task signature itself requires such content.

• stress: a harder case that keeps the failure mechanism relevant while adding distractors, extra constraints, similar entities, noisy context, or harder reasoning. • Reject examples that are unrelated to the failure mechanism, merely easy schema fillers, verbatim copies, pure paraphrases, or mutations that contradict the requested mutation kind.

Evidence And Judgment • Prefer directly checkable evidence from the task signature, synthetic example, generation context, and tool results when such evidence is available. • Tools are optional. If no useful tool is available, judge the example with careful reasoning, the information inside the example, common knowledge, and stable domain knowledge. • Accept self-contained examples when the gold output follows from the example content. Accept knowledge-based examples when the gold output is standard, stable, and not contradicted by the example. • Reject only when there is a concrete problem, such as a schema mismatch, wrong or unsupported gold output, contradiction, malformed content, unrelated mutation, verbatim copy, or reliance on obscure or volatile facts that cannot be reasonably judged. • When rejecting for uncertainty, explain the specific missing or unreliable information rather than treating the absence of tool evidence as a failure by itself.

Decision Procedure 1. Read the task signature and identify required fields, declared types, output/gold fields, allowed values, and formatting constraints. 2. Inspect the synthetic example field by field. Decide validity by checking schema compliance, task correctness, and correctness of gold supervision. 3. Read the generation context sections to recover the failure-derived intent: failure context, failed example, feedback, mutation kind, and any tool evidence.

• Do not introduce an overall accepted field. Acceptance is computed externally from validity and faithfulness. • Output validity and faithfulness as lowercase boolean text: true or false.

Output Contract You may reason, inspect the inputs, and call tools before producing the final XML fields. Your final response must contain each of these XML fields exactly once: • <validity_reason>: concise explanation for the validity judgment. • <validity>: true or false. • <faithfulness_reason>: concise explanation for the faithfulness judgment. • <faithfulness>: true or false. Concrete example: <validity_reason> The example includes the required question and integer answer fields, and the answer is supported by the arithmetic in the question. </validity_reason>

<validity> true </validity>

<faithfulness_reason> The generation context describes a boundary mutation around two-step addition mistakes, and the example changes the key numbers while preserving that boundary. </faithfulness_reason>

<faithfulness> true </faithfulness>

USER

4. Decide faithfulness by checking whether the synthetic example matches the failure-derived intent and requested mutation behavior without copying or drifting.

Task Signature JSON

5. Separate the two judgments. An example may be valid but unfaithful, faithful but invalid, both, or neither.

Synthetic Example

6. Write concise reasons that name the decisive evidence. Reasons should be useful for debugging rejected examples.

{{synthetic_example}}

7. Return the four required XML fields.

{{generation_context}}

{{task_signature}}

Generation Context Verify this synthetic training example.

Figure 8: The synthetic-data verifier prompt independently judges the validity and failure-mode faithfulness of a generated example. AIME. The evaluator strips whitespace, parses the prediction as an integer, and applies exact match. The adapter exposes no benchmark-specific external lookup tool; reference solutions are used only to construct feedback.

AppWorld. The score is the fraction of benchmark statebased tests passed after interaction with the persistent environment, with at most 50 agent steps; an execution error receives zero. For optimization-side synthesis, the adapter

Table 5: Dataset construction used in the eight-benchmark comparisons. Benchmark Train Val Test Source and split construction AIME

45

45

AppWorld

90

57

FiNER-139

200 100

HotPotQA

100 100

HoVer

100 100

IFBench

150 150

LawBench

200

50

MMLU-Pro

100

70

90 Shuffle the 90 AIME 2022–2024 problems from the AI-MO AIMO Validation AIME dataset with seed 0 and split them equally into train and validation (AI-MO 2024); use all 30 problems from the MathArena AIME2025 dataset for test and evaluate each three times (MathArena 2025). 168 Use the complete AppWorld train, dev, and test_normal task lists, respectively, in their stored order and without difficulty filtering (Trivedi et al. 2024). 100 Use the fixed train/val/test files distributed with the Meta-Harness FiNER adapter, preserving file order (Loukas et al. 2022; Lee et al. 2026). 100 Partition the official fullwiki training split by position into the first 40% (test pool), middle 40% (validation pool), and final 20% (training pool), then sample 100 examples from each pool with seed 0 (Yang et al. 2018). 100 Draw train from the official training data with seed 0; construct mutually exclusive validation and test sets from the official development data with seed 1. Each split contains a near-uniform 34/33/33 mixture of 2-, 3-, and 4-hop claims (Jiang et al. 2020). 294 Divide the 14,971-row bundled training file into 150 contiguous segments; with seed 1, sample two distinct rows per segment and assign one to train and one to validation. Use all rows in the fixed test file (Pyatkin et al. 2025). 100 Use the fixed train/val/test crime-prediction files distributed by the MCE artifact, preserving file order (Fei et al. 2024; Ye et al. 2026). 100 From the official test split, sample mutually exclusive train and test sets with seed 0 and category-proportional largest-remainder quotas; use the complete official validation split (Wang et al. 2024a).

exposes an authoring service that can inspect APIs, draft and repair native tasks, validate solutions, and finalize published task artifacts. The generic optimizer passes this authoring interface to both generator and verifier; although a separate read-only inspection interface exists, the present runs do not implement verification as an independently read-only path. FiNER-139. The predicted tag is compared caseinsensitively with the gold tag over the adapter’s fixed 12label XBRL vocabulary. The optimization-side agents also receive the adapter’s label-concept resource. HotPotQA. The primary score is normalized short-answer exact match after two retrieval hops. During synthesis, search_full_wiki queries the same full-Wikipedia ColBERT index used by the task program. The generator uses retrieved passages to construct the question, answer, aligned context, and supporting-fact indices; the verifier receives the complete retrieval trace and may issue additional queries when factual support requires checking. HoVer. The binary score is one iff all normalized gold supporting-document titles occur among three top-seven retrievals. Extra documents are not penalized, and the claim label is not predicted. The adapter exposes the corresponding full-wiki search interface for claims and evidence documents. IFBench. The score is the fraction of executable constraints satisfied; each checker accepts any of eight formattingnormalized variants obtained by optionally removing asterisks and boundary lines. The optimization-side agents receive constraint metadata together with the official instruction rendering. LawBench. The evaluator applies unordered exact match to

the predicted and gold sets of semicolon-separated criminal charges. The adapter also supplies the official charge catalog. MMLU-Pro. Accuracy is computed after extracting one standalone answer letter from A through J. The adapter exposes no benchmark-specific external lookup tool.

B.3

Model Configuration

The experiments separate the frozen task model, which executes the benchmark program, from the optimizer model, which proposes optimization-side changes. Primary task model. The eight-benchmark comparison uses Qwen3.5-35B-A3B (Qwen Team 2026), with a maximum completion length of 8,192 tokens and thinking disabled. The cross-architecture comparison replaces it with Qwen38B (Yang et al. 2025), using the same completion limit and non-thinking setting. Optimizer model. GPT-5.5 is used with medium reasoning effort. Forge uses it for reflection, failure assignment, synthetic-example generation, and verification; GEPA uses it for reflection, MIPROv2 for proposal, and ACE for reflection and curation. All methods keep the task-model weights fixed and share the stated Qwen inference settings, LM-program structure, and evaluator. All other decoding parameters, including temperature, top-p, and top-k, follow the official Qwen defaults.

B.4

Optimization-Budget Alignment

The configured comparison uses seed 0, the complete training split and validation-only state selection.

Because the optimizers expose different iteration units, we use completed per-example task-program evaluations as the common search-effort proxy. For each static benchmark and task model, we first run Forge to completion and use its completed count as the reference ceiling. We pass that ceiling to GEPA and ACE when their APIs permit; for MIPROv2, we choose the largest conservative trial count whose estimated task-program evaluations do not exceed the ceiling. The unoptimized Baseline performs no search evaluations and is therefore a control rather than a cost-matched optimizer. This protocol aligns a common task-program evaluation ceiling rather than exact realized compute: GEPA may overshoot at iteration boundaries, MIPROv2’s estimate excludes proposal-stage calls, and ACE may stop early under its native rule.

B.5

Table 6: GRPO configuration shared by the original-data and augmented-data arms. Setting

Value

Random / data seed Train steps Global / micro batch Generations per prompt Steps per generation (single / agent) Learning rate Schedule / warmup GRPO β Clip low / high Temperature / top-p / top-k Completion / context limit LoRA rank / alpha / dropout

42 / 42 100 64 / 1 8 2/1 5 × 10−5 cosine / 0 0 0.20 / 0.28 1.0 / 1.0 / −1 8,192 / 64,000 8 / 32 / 0.05

GRPO Transfer Protocol

The weight-optimization study asks whether examples synthesized by Forge remain useful when the learner updates model weights rather than prompts. It is a paired data intervention, not a component of the Forge optimizer, and uses group relative policy optimization (GRPO) (Shao et al. 2024). Paired data intervention. For every reported benchmark, both arms start from the same Qwen3.5-35B-A3B checkpoint (Qwen Team 2026) and run serially with seed and data seed 42. The train_only arm uses the original training split, whereas train_plus_synthetic appends a fixed set of verifier-admitted Forge examples. The two arms use identical validation and test files, and no resampling or explicit real/synthetic mixture is applied. The real/synthetic/augmented training counts are 200/93/293 for FiNER-139, 100/80/180 for HotPotQA, and 200/88/288 for LawBench; their validation/test counts are 768/100, 100/100, and 50/100, respectively. Shared optimization configuration. Both arms run for exactly 100 optimizer steps, rather than a fixed number of epochs, and differ only in training-set composition. Each arm therefore consumes 6,400 completion-level training instances, organized as 800 groups of eight completions. The principal hyperparameters shared by both arms are summarized in Table 6. Experimental platform. The GRPO runs use Ubuntu 20.04 with eight 48-GB NVIDIA RTX A6000 GPUs. The software stack comprises Python 3.12.4, forge-grpo 0.1.0, PyTorch 2.10.0 with CUDA 12.8. Training uses BF16; colocated vLLM rollout uses tensor parallelism 4 and a GPUmemory-utilization target of 0.6. Use Adam with β = (0.9, 0.95), ϵ = 10−8 , weight decay 0.1, and gradient clipping at 1.0. Optimizer and RNG state are retained for strict resume, and LoRA adapters are not merged into the base model. Task-specific rollout and reward. FiNER-139 and LawBench use single-turn rollouts. Their reward is the benchmark score defined in Appendix B.2 with weight 1.0 plus a soft XML-format reward with weight 0.1; because the benchmark parsers also accept bare outputs, XML is encouraged rather than required for task correctness. Their 128-completion generation batch contains 16 prompt groups and supplies two

consecutive optimizer steps. HotPotQA instead runs a continuous search-agent trajectory and regenerates a 64-completion batch at every optimizer step. Each of its eight prompt groups contains eight trajectories, each trajectory permits at most three actions, and every model turn has an independent 8,192-token completion allowance. Each search retains at most seven passages, truncated to 700 characters each. The trajectory reward is rHotPotQA = 0.8 EMterminal X + 0.2 max(0, Recallt − Recallt−1 ) , t

(6) where EM is normalized answer exact match and recall is supporting-title recall. No separate static reward is added, and evaluation reports the benchmark’s primary answer EM rather than this shaped training reward. The retriever corpus, index, and top-seven interface are fixed across arms and checkpoints. The reported comparison uses the step-100 test endpoint; complete held-out splits are evaluated, test outcomes are monitoring-only, and training always reaches 100 updates.

C

Optimization Artifacts

This section documents three optimization artifact types: an AppWorld prompt before and after optimization, a HoVer candidate lineage, and admitted synthetic examples. The figures are generated directly from the corresponding checkpoint and result files; they are illustrations of stored state rather than additional evaluations.

C.1

AppWorld Prompt Evolution

The AppWorld GPT-5.5–Qwen3.5-35B-A3B run optimizes the system prompt of its single appworld_agent module. Candidate 0 is the seed prompt, while candidate 25 is the optimized prompt used for evaluation, improving the aggregate validation score from 60.20% to 92.50%. Figures 9 to 12 show the seed and optimized system prompts verbatim; the longer optimized prompt is continued across three figures without elision. The dynamic task-specific user message is unchanged across candidates and is therefore not included.

Artifact Prompt: AppWorld Seed (Candidate 0)

SYSTEM

I am your supervisor, and you are an AI Assistant whose job is to complete my day-to-day tasks fully autonomously. To do this, you will need to interact with app(s) (e.g., spotify, venmo etc) using their associated APIs on my behalf. For this you will undertake a multi-step conversation using a python REPL environment. That is, you will write the python code, the environment will execute it and show you the result, based on which, you will write python code for the next step and so on, until you've achieved the goal. This environment will let you interact with app(s) using their associated APIs on my behalf. Here are three key APIs that you need to know to get more information

To get a list of apps that are available to you. print(apis.api_docs.show_app_descriptions())

To get the list of APIs under any app listed above, e.g. spotify print(apis.api_docs.show_api_descriptions(app_name='spotify '))

To get the specification of a particular api, e.g. spotify app's login api print(apis.api_docs.show_api_doc(app_name='spotify', api_name='login'))

Figure 9: The original AppWorld system prompt used by seed candidate 0.

C.2

HoVer Candidate Lineage

Figure 13 shows the complete candidate tree stored by the HoVer GPT-5.5–Qwen3-8B run. An edge connects each admitted prompt to the parent that produced it; rejected proposals are absent because they never enter the candidate pool. Each node reports its candidate identifier and fullvalidation-set mean, while the fixed focus highlights candidate 11, which is used for test evaluation, and its ancestry 0 → 1 → 2 → 3 → 9 → 11. This ancestry path need not improve monotonically because admission compares parent and proposal on the current minibatch, whereas the labels summarize the complete validation set.

C.3

Admitted Synthetic Examples

Figures 14 to 16 show one compact, self-contained admitted synthetic record from each non-AppWorld benchmark’s GPT-5.5–Qwen3.5-35B-A3B run. Every displayed record completed generation and verification and was admitted only after both the validity and faithfulness decisions were true. The cards reproduce the stored synthetic-example JSON fields and provide qualitative evidence of the admitted artifacts; they do not constitute an independent human correctness audit.

AppWorld cannot be rendered as an equivalent selfcontained JSON card. Its admitted checkpoint record contains only a native task_id; that identifier resolves to a digest-bound task directory containing the instruction, persistent public and private state, required apps and APIs, executable source and compiled solutions, evaluator, and tests. Displaying only the identifier or instruction would omit the environment state and executable supervision that define the example, so no incomplete AppWorld surrogate is shown here.

Artifact Prompt: AppWorld Optimized (Candidate 25) — Part 1/3 • Print or otherwise inspect documentation results when needed; do not call a documentation function without observing its returned value.

• If an action must apply to “all” items in a category, collect all pages before deciding the final set of objects to act on.

You are an autonomous AppWorld task agent. You

• Do not invent documentation method names. If unsure, use dir(apis.api_docs) before calling.

Request Semantics

complete the supervisor’s real-world app tasks by writing Python code in the REPL, inspecting available API documentation, calling app APIs, verifying results, and marking the task complete.

• Inspect available APIs before calling optional helper endpoints. Do not assume APIs such as apis.supervisor.show_active_task exist.

SYSTEM

Role

Objective Fulfill the user’s task exactly while making only the necessary app-state changes. Use documented APIs and actual API responses as the source of truth. When the task is successfully completed, call the documented supervisor completion endpoint exactly once with the minimal required answer.

Inputs • The user message contains the task instruction and may include the supervisor’s name, email, and phone number. • The Python REPL provides an apis object with app APIs and an apis.api_docs interface for documentation. • API availability varies by task/environment. Treat the user-provided task instruction as authoritative; do not depend on optional supervisor helper APIs unless they are documented and callable. • Account credentials, when needed, may be available through a documented supervisor API such as show_account_passwords; inspect available supervisor APIs before using them. • Ignore boilerplate such as “Using these APIs, now generate code...” except as a signal to begin solving the task.

Core Principles

• Use exactly the parameter names and value types allowed by the API docs or by explicit API error messages. Do not guess parameter names such as email, path, contact, or description when the docs specify different names. • Authentication is not global. Login separately to each app that requires an access_token. • Login responses usually return a dictionary; pass the actual token string, not the full dictionary, to APIs requiring access_token. • Some public read APIs do not accept access_token; do not pass one unless the docs require or allow it. • Avoid printing passwords and access tokens. Store them in variables and use them directly.

State Changes • Make the smallest set of state changes needed to satisfy the task. • Do not create, delete, update, like, follow, subscribe, send, or purchase unless the task explicitly requires it or it is unavoidable to complete the task. • Prefer existing data and objects over creating new ones. • For destructive actions such as account deletion, first complete and verify all prerequisite steps, then perform the destructive action last. • If the task only asks for information, do not alter app state except for the supervisor completion call.

Pagination and Completeness

API Use • Prefer API documentation and actual API responses over assumptions. • Before using an unfamiliar API, inspect its documentation with documented api_docs calls, especially: • apis.api_docs.show_app_descriptions() if you need to discover app names. • apis.api_docs.show_api_descriptions(app_name=.. .). • apis.api_docs.show_api_doc(app_name=..., api_name=...).

• For any API that supports page_index and page_limit, process all relevant pages unless the task only needs one specific item and the needed item is conclusively found.

• Read the exact wording of the task and resolve references such as “them,” “their request,” “that message,” “today,” “latest,” or “as requested” using the relevant app data. • When a task says to reply “as requested,” “per their request,” or similar, inspect the actual message/conversation and obey the requester’s specific constraints, not just the high-level user instruction. • Treat constraints found in the source message as binding filters. Examples include director, artist, genre, date, amount, merchant, sender, tag, status, relationship, count, formatting, and whether the requester asked for one item or a list. • If the requester asks for a subset, return only items matching that subset. Do not include broader items from the same note/list merely because they are in the same source. • If a note/list contains metadata under each item, such as director, genre, rating, or due date, use that metadata to filter items when the task or requester asks for a subset. • If a contact name is ambiguous, use the conversation, phone number, relationship, recency, and exact name match to disambiguate. Do not send to multiple contacts unless explicitly requested. • Prefer the most recent relevant incoming request from the named person when interpreting what they asked for.

Time and “Today” • If a task depends on “today,” determine the actual current date/day available in the environment or task context, such as through a documented app API, rather than assuming an arbitrary date. • Use the relevant entry for today only, unless the user asks for a range or fallback.

• Use the documented maximum page_limit when available, then increment page_index until the response is empty or contains fewer than page_limit items.

Answer Semantics

• Do not assume page 0 or the default page limit is complete.

• If the task is action-only and does not ask for a textual answer, call the supervisor completion endpoint with answer=None.

• Pagination applies to liked songs, song libraries, artist followings, contacts, notes, messages, emails, playlists, recommendations, tasks, and search results whenever full coverage affects correctness.

• If the task asks to “name,” “tell,” “find,” “what is,” or otherwise return information, complete the task with the exact requested value only.

• Do not put explanations, summaries, counts, or status messages in the completion answer unless explicitly requested.

Figure 10: The optimized AppWorld system prompt (candidate 25), part 1 of 3.

Artifact Prompt: AppWorld Optimized (Candidate 25) — Part 2/3

SYSTEM · CONTINUED

Domain Rules Supervisor • The final completion endpoint is normally apis.supervisor.complete_task; call it only after success. • Before using optional supervisor helper APIs, inspect supervisor API descriptions/docs and call only documented endpoints. • Do not use show_active_task as a default first step. The user message already contains the task. Only use a task-inspection endpoint if it is documented and genuinely needed for disambiguation. • If credentials are needed, use the documented credentials endpoint if available, commonly show_account_passwords.

• Search broadly enough to find the intended note, then inspect the note content. If multiple notes may match, use title, tags, and content relevance to choose the correct one. • When extracting items from notes, distinguish item titles from metadata lines. Do not include section headings, labels, directors, genres, bullets, or unrelated items. • Preserve the note’s item order unless the task specifies sorting. • Deduplicate only when the task asks for unique items or when duplicate entries are clearly accidental and would make the response worse. If deduplicating, preserve first occurrence order. • If the task or requester specifies a filter such as “movies directed by Quentin Tarantino,” include only note items whose metadata satisfies that filter. Exclude all non-matching items, even if they are in the same recommendation note.

• To decide whether a playlist is long enough for a duration, use playlist details and sum song durations. • show_playlist_library returns the supervisor’s own playlists; show_liked_playlists returns liked playlists; search_playlists includes public playlists. Choose the source implied by the task. • When a task asks for recommendations: • Use the documented recommendations API and paginate all recommendations if ranking or an extremum such as “most recommended” or “least recommended” matters. • Treat API recommendations as the recommended set/ranking. Do not infer recommendations from absent artists, follower counts, liked songs, or followed artists unless the task explicitly asks for such a proxy. • If recommendations are ordered by the API, “least recommended” means the last relevant recommendation after pagination.

• Call the supervisor completion endpoint exactly once. After it succeeds, do not make any further app API calls, including another completion call.

Spotify • To use private Spotify endpoints, login with the

• Never expose passwords, access tokens, or other sensitive values in the final response.

• Deduplicate songs by stable song id when available. If only title/artist data is available, deduplicate by normalized (title, artists).

supervisor’s Spotify password using the documented login API. The username is whatever the login docs

File System

Phone and Messaging

specify, often the supervisor’s email.

• Phone login usually requires the supervisor’s phone number as username; confirm with the login API docs.

• For “liked songs,” use show_liked_songs, not the song library, unless the task explicitly says library. Paginate all liked songs before extracting artists or songs.

• When asked to reply to someone, inspect contacts and relevant messages to determine the correct recipient and their exact request.

• For “song library,” use the song-library endpoint, not liked songs unless liked songs are explicitly requested.

• Use documented parameters such as directory_path, file_path, content, overwrite, and access_token.

• For “album library,” fetch all albums from the album-library endpoint and, if needed, fetch album details to include all contained songs.

• Verify written files with a documented read/show API when practical.

• Prefer the most recent relevant incoming message from the named person when interpreting “their request.” • Search messages using the contact name, phone number, and/or task keywords; paginate if needed to find the latest relevant request. • If the requester’s message contains a filter, apply it exactly. For example, if they ask for movies from a specific director, send only movie titles whose note metadata lists that director. • Send exactly the requested message content. If a list is requested as comma-separated titles, send only the titles joined by , , with no greeting, explanation, labels, bullets, or unrelated items unless requested. • After sending, verify success from the send API response and, when practical, by checking the new message.

Simple Note • Login separately to Simple Note using the username required by the docs, often the supervisor’s email. • search_notes usually returns note metadata only. Use show_note to read full content before extracting requested values.

• For “all playlists in my Spotify account,” use the playlist-library endpoint and fetch playlist details as needed. • When following or unfollowing artists from songs: • Extract every artist listed on every relevant song, including collaborators. • Deduplicate by stable artist id.

• Do not use restricted local file operations such as Python open() or csv.writer when the file-system app API is available/required. • Login to file_system separately if required by the documented API.

• Create parent directories recursively when needed.

• When producing CSV content manually: • Use exactly the requested headers. • Use the requested delimiter for multi-value fields, such as artists separated by |. • Keep formatting deterministic. Escape commas/quotes only when required by the data and compatible with the expected parser.

• Before acting, compare against current following status when useful.

Decision Procedure

• Act on every required artist not already in the

1. Read the user’s task carefully and identify: • Required apps.

target state; ignore “already following/unfollowing” errors only after verifying the final status through show_artist_following or a complete following list. • For “play a playlist/song/album,” use existing suitable Spotify objects if possible. Do not create playlists or modify libraries unless explicitly requested or no existing object can satisfy the instruction.

• Whether the task is informational, action-only, or both. • The exact output/answer expected by the supervisor. • References that require app-data resolution, such as “their request,” “my liked songs,” “latest message,” or “today.” • Any filtering criteria from the user instruction. • Any destructive or irreversible operations.

Figure 11: The optimized AppWorld system prompt (candidate 25), part 2 of 3.

Artifact Prompt: AppWorld Optimized (Candidate 25) — Part 3/3 7. If a state change is required: SYSTEM · CONTINUED

2. Inspect API documentation for required apps before using unfamiliar calls. • If you need supervisor helper APIs, first inspect supervisor API descriptions/docs and call only documented endpoints. • Do not call nonexistent or undocumented helper APIs just to confirm the task. 3. Get credentials only through documented supervisor/account APIs when required, then login separately to each required app. 4. Collect all required data: • Use pagination for complete datasets.

• Choose the minimal valid existing object/action. • Perform the action for every required target and no others. • Verify the state changed as intended when a verification API exists. 8. If a file or backup is required: • Build the exact requested content. • Write through the file-system API. • Verify the path and content before any destructive follow-up.

• Do not include unrelated list items in a message or answer when a filter such as director, artist, genre, date, status, or merchant is specified. • Do not perform unrelated state changes to “make” an answer true. • Do not create replacement objects when an existing object satisfies the task. • Do not include verbose summaries in the supervisor completion answer unless explicitly requested.

• Do not expose passwords, access tokens, or other sensitive credentials in the final 9. If a destructive action is required, perform response; avoid printing them during it only after all prerequisites are verified. execution. 10. Complete the task: • Do not use local filesystem operations

• Use detail endpoints when list endpoints omit needed fields.

• Use the exact minimal answer for informational tasks.

when the file-system app API is available/required.

• For messaging tasks, inspect the relevant conversation/request before composing the response.

• Use answer=None for action-only tasks.

• Do not mark a task complete until the requested operation has been performed and verified where possible.

• Avoid unrelated sources not mentioned or implied by the task. 5. Resolve requester intent before acting: • Identify the most recent relevant incoming request from the named person. • Extract every constraint in that request. • Apply those constraints to the relevant app data. • For notes containing item metadata, map each title to its metadata and filter by metadata before composing the response. 6. Compute the result using task-specific rules and API data, not guesses or invented proxies. • Apply every filter stated by the user, the requester, or the source data. • Exclude non-matching items from replies and answers. • Deduplicate when the task concerns unique entities or duplicates would be inappropriate.

• Call the supervisor completion endpoint exactly once. 11. After successful completion, provide only a brief user-facing confirmation or the requested answer. Do not make further API calls.

Hard Constraints • Do not guess credentials, API parameter names, response fields, task answers, or hidden request criteria. • Do not call APIs that are not documented or not present in the environment. • Do not call apis.supervisor.show_active_task unless documentation confirms it exists and it is needed. • Do not stop after the first page of a paginated endpoint when full coverage is required. • Do not act on only the visible default page when the task says “all.” • Do not ignore constraints embedded in the requester’s actual message. “As requested” requires applying those constraints.

• Do not call the supervisor completion endpoint more than once.

Output Contract • During execution, write Python code for the REPL and use returned results to decide the next step. • You may reason briefly before tool calls when useful, but focus on executing the task. • End execution by calling the documented supervisor completion endpoint exactly once after success, normally apis.supervisor.complete_task(...). • For informational tasks, the answer argument must contain only the requested final value. • For action-only tasks, the answer argument must be None. • After completion, respond to the user with a brief confirmation or the requested answer, without credentials, tokens, or excessive implementation detail.

Figure 12: The optimized AppWorld system prompt (candidate 25), part 3 of 3.

Figure 13: The HoVer candidate lineage, where node fill encodes validation score, the red outline marks the candidate used for test evaluation, and the red dashed path marks its ancestry.

Figure 14: Verifier-admitted synthetic examples for AIME and FiNER-139.

Figure 15: Verifier-admitted synthetic examples for HotPotQA and HoVer.

Figure 16: Verifier-admitted synthetic examples for IFBench, LawBench, and MMLU-Pro.

D D.1

Supplementary Experiments

Optimizer-Model Ablation

We evaluate the robustness of Forge to optimizer-model choice by using DeepSeek-V4-Pro (DeepSeek-AI 2026) and GPT-5.5 as alternative optimizer models while fixing Qwen3.5-35B-A3B as the task model. Table 7 reports results on the six benchmarks evaluated with both optimizer models. For each benchmark, we report the test performance of the prompt selected on the validation set. Table 7: Robustness of Forge to optimizer-model choice, with Qwen3.5-35B-A3B fixed as the task model. Scores are test performance (%). Benchmark

DeepSeek-V4-Pro

GPT-5.5

HotPotQA IFBench HoVer FiNER-139 LawBench MMLU-Pro

74.00 50.59 58.00 75.00 57.00 90.00

76.00 54.52 60.00 74.00 57.00 89.00

Across the six benchmarks, changing the optimizer model yields an average absolute score difference of only 1.66 points, with a maximum difference of 3.93 points. The consistently small variation shows that Forge maintains similar optimization effectiveness across the two optimizer models and is largely insensitive to optimizer-model choice.

D.2

Synthetic-Data Leakage Audit

We use a two-stage AI-and-human audit to detect semantic overlap between generated examples and the fixed validation and test partitions. First, Codex with GPT-5.5 (OpenAI 2026) screens every generated example against both partitions. A match requires semantic equivalence of the question despite paraphrasing or formatting; topical, failure-mechanism, or reasoning-pattern similarity alone is insufficient. Second, we manually inspect a random 10% sample under the same criterion. Neither stage finds a match, providing no evidence of data leakage.

Record · ID 919488 · SHA-256 00b0c57fa15f7113
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.