Conceptio › Archive › arXiv CS
arXiv CSopen access

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

2026-08-17

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Haozhe Liu1 * , Tian Ye1 * , Sensen Gao2 † , Qihang Cao2 † , Yitong Li1 † , Mingchen Zhuge Duomin Wang1 , Ruihua Zhang1 , Ping Luo1 , Jiawang Bian2 , Lei Zhu1 , Ligeng Zhu1 , Enze Xie1 , Song Han1 3 1 *

2

NVIDIA

†

Blog Post

As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement. We take an RSI-inspired approach at the harness layer, scaling auto-research loops across increasingly numerous and diverse environments for harness rollouts. At this scale, the process yields reusable improvements that transfer beyond their development setting, moving automated harness discovery toward production-level outcomes. Four mechanisms survive selection and form SoL-Pi, spanning action execution, context compaction, observation handling, and delegated reading. On the 51-task EdgeBench evaluation, SoL-Pi achieves performance comparable to Pi across GPT-5.6 Sol and Opus 5 while reducing recorded token traffic by 44.7–49.0% and API cost by about one third. In other words, estimated hourly savings are $8.75–$13.50 relative to native Codex and Claude Code harnesses, and $4.36–$5.71 relative to Pi. (a) SoL-Pi: SoL-Pi:Scaling ScalingAuto-Research Auto-Research Loop Loop

us methods methods ous methods u ou m m m hod hod hod vious methods methods methods Previous methods hod ouv ou memhod

Previous Previous methods methods (a) Previous methods ab aSoL-Pi P aSoL-Pi P P vPSoL-Pi vSoL ou ou m m m hod hod hod (a) Previous Previous Previous methods methods methods (a) (a) (a) SoL-Pi SoL-Pi SoL-Pi harness: harness: harness: Scaling Scaling Scaling Auto-Research Auto-Research Auto-Research (b) (b) (b) SoL-Pi SoL-Pi harness: harness: harness: Scaling Scaling Auto-Research Auto-Research Auto-Research (b) (b) (b) SoL-Pi SoL-Pi SoL-Pi harness: Loop harness: harness: Scaling Scaling Scaling Auto-Research Auto-Research Auto-Research Loop Loop (a) SoL-Pi SoL-Pi harness: harness: Scaling Scaling Scaling Auto-Research Auto-Research Loop Loop SoL SoL P P harne P harne Sca Sca Sca ng ng ng Au Au Au o(b) oSoL-Pi Re obSoL-Pi Re Re earch earch Loop Loop (b) SoL-Pi SoL-Pi harness: harness: harness: Scaling Scaling Auto-Research Auto-Research Auto-Research (b) (b) SoL-Pi SoL-Pi harness: Loop harness: Loop harness: Scaling Scaling Scaling Auto-Research Auto-Research Loop Loop Loop b SoL b SoL SoL Pvou harne Pharness: harne harne Sca Sca Sca ng ng Au Au Au ong b(b) obRe oAu SoL bRe SoL earch SoL P earch P harne P harne harne Sca Sca Sca ng ng ng Au Au Au oAu oRe oRe earch earch (a) Previous methods aSoL PSoL vP ou mScaling hod (a) SoL-Pi harness: Scaling Auto-Research Loop Pharne harne Sca ng oearch Re earch harness: Scaling Auto-Research harness: Loop Scaling Auto-Research b P ha nScaling S ang ng Au oRe SoL aearch P hLoop ha Loop nLoop SAuto-Research aSca ng oRe Rearch aLoop hLoop LoopLoop aSoL-Pi Prev ou me hod SoL P harne Sca Au oearch Re earch Loop b(b) SoL harne Sca ng Au obR Re SoL P harne Loop ng Au o Re earch

enchmark Benchmark Benchmark tasks tasks tasks chmark nchmark tasks Benchmark tasks B n hm nhm ktasks k kkk tasks Benchmark Bhm nk k hm k B n hm k k

GitHub GitHub GitHub Issue–PR Issue–PR pairs pairs pairs GitHub GitHub Issue–PR Issue–PR pairs pairs GitHub Issue–PR pairs G G Hub G Hub Hub uIssue–PR u Issue–PR u PR PR PR p p p GitHub pairs Hub u PR Benchmark Benchmark tasks tasks Benchmark Benchmark Benchmark B B nBenchmark B nPublic hm hm hm ktasks ktasks ktasks kk k Benchmark GGHub unB ppktasks Public nPR hm ktasks

Run Run Run harness harness harness Run harness Run Run Run harness harness harness

R

Harness Harness Harness

Harness Harness Harness Harness Harness Harness H H H n nH nHarness n Agent Agent HAgent n Agent Agent Agent AgentA A A A

A nvironment ronment Environment Action Action Action Environment Action nment ronment Action Action vironment Action m m m mA A A A m

A

Research Researchenvironments environments

Next Next Next round round round Next Next round round Next round N N N

m

Merged m Merged Merged Merged AI proposes AIAI proposes AI proposes proposes

Merged Merged M Merged M M M

A

AIAI proposes AI proposes proposes Match PRs AIAI proposes AI proposes proposes

Propose Propose H Propose CPropose

refactors refactors in in parser parser Rust C++ Java refactors in parser Rust C++ Java Rust Rust Rust C++ C++ C++ Java Java Java m Merged Merged Merged Merged Rust Rust C++ C++ Java Rust C++ Java Java Merged Merged M Merged M M M

Rust C++ C++ Java Java Rust Rust Rust C++ C++ Java Java Propose Propose Propose P P P P

AI AI proposes proposes A AI proposes prop improvement A A A improvement improvement improvement R C Match PRs DataWeb Web Data Data Data Web Web M M MLMLMLML improvement m m change mimprovement mimprovement mA mpApply mchange m Apply Apply Apply change p change AutoLoop Open science app pipeline science app pipeline science science science appapp app pipeline pipeline DAutoLoop Wpipeline Data Data Web Web ML ML Web ML benchmark Re-run benchmark Data Web MLM D D D W W W M M M Re-run Re-run Re-run benchmark benchmark benchmark Re-run Re-run Re-run benchmark benchmark benchmark AutoLoop Open Open ARe-run A Apply Apply change change Apply Apply change change Apply change Apply change A A A A A A AutoLoop AutoLoop AutoLoop Open Open Open AutoLoop AutoLoop AutoLoop Open Open Open AutoLoop mp m n mpOpenm n science Environments science appappapp pipeline pipeline pipeline mscience Re-run Re-run Re-run benchmark benchmark benchmark Re-run benchmark benchmark benchmark Environments m m m mAu m m m Environments Environments Environments Environments Environments Environments Au Lp nRe-run pnOp AutoLoop AutoLoop Open Open AutoLoop AutoLoop Open Open D WOp AutoLoop Open AutoLoop Open Au Au L Environments L Lp pOp Op Op npnOp n Re-run Au Au Au LAu L Lp L pOp pOp nnn M A

m

Score improves Score Score Score improves improves improves

Score Score Score improves mimproves mimproves m m

Environments A En Environments Environments Environments En En En nm nm nm n nnm n n Au Keep L change p Op n Keep Keep Keep change change change En nm n

Score improves Kchange Score Score Score improves improves improves Keep Keep change change Keep K K K

K

InfraDocs Docs CLI CLI CLICLI Infra Infra Infra Docs Docs En nm Environments Environments Environments En En nm nm nm nsite nsite nsitensite tool scripts mEn tool tool tool scripts scripts scripts

Au p Op D n CLI CLI CLI Infra InfraL Docs Docs CLI Infra Infra Docs D D Docs D tool tool scripts scripts nm sitesitesite En n tool scripts C

Score Score Score improves mimproves mimproves m m

D

Propose Propose Propose Propose H CC m Propose Propose Propose P P P

L w C

Rust Rust Rust C++ C++ Java C++ Java Java Test candidates Test Test Test candidates candidates candidates

ObservationPack ObservationPack ObservationPack Test Test Test Test candidates candidates candidates ObservationPack ObservationPack ObservationPack O O O Test Test candidates Tcandidates Test candidates T T T

L w C

Sy em

P p P p Test Test candidates Tcandidates Test candidates T T T R CE E EE BData AData DML A BBB B C C COptimization Optimization Optimization A AData AData B BBWeb C C CC DML DML DML AAA DOptimization DD EPerformance EE E Web Web Web Evidence-Preserving Evidence-Preserving Evidence-Preserving Evidence-Preserving science science science appappapp pipeline pipeline pipeline A D B DM EAI AE B C Op D Reducer Reducer Reducer Reducer WC Research Evidence-Preserving Evidence-Preserving Evidence-Preserving E E

CLI Infra Infra Docs Docs CLI CLI CLI Infra Infra Docs Docs tool scripts scripts sitesitesite site tool tool tool scripts scripts

CLI CLI CLI Infra Infra Infra Docs D Docs D Docs DD tool tool scripts scripts sitesitesite tool scripts

C

Cost Performance Performance Cost Cost Cost Performance Performance

Cost Cost Cost

Retain Retain Retain Retain

Performance Performance Performance mmm

Harness (Higher Cost) m

Average score ↑ GPT-5.6 AverageAPI score cost ↑ (USD) ↓ GPTSol API cost (USD) Average ↓ score ↑ 5Average 6Average So 6 So API GPT-5.6 Sol Sol GPT-5.6 GPT-5.6 Sol Average Average Average score score score ↑ ↑ ↑ re GPT Average API score API score API score cost ↑cost ↑cost (USD) ↑(USD) (USD) ↓re ↓ GPT ↓uGPT API API cost cost cost (USD) Average (USD) Average (USD) Average ↓ ↓score ↓score score ↑↑↑ GPT-5.6 Sol GPT-5.6 Sol GPT 5GPT 56He 56So 6So So GPT 5 56 56So65So So b d ou EdgeBench ou EdgeBench uGPT

Research AI Math R Math Math Math Reducer R Reducer R Reducer R

D

≥

O TOp Optimization Optimization aaon on A BBB B C C COp Dmm Optimization on on AAA DOp Dm Ea EaE E Performance

T AD B C DM E E EE app pipeline Ascience Ascience AD Bapp Bapp C C Cpipeline DM DM DM science science app pipeline pipeline DB W Data Data Web Web ML ML Data Web ML D W W W

Harness (Higher Cost)

mEdgeBench m results (b) Held-out EdgeBench results Held-out result (b) (b) (b) Held-out Held-out Held-out EdgeBench EdgeBench EdgeBench results results -out d-out ut EdgeBench EdgeBench EdgeBench result result result bdSol H dEdg ouEdg Edg n hu u uGPT-5.6 GPT-5.6 Sol GPT-5.6 Sol dEdgeBench ou Edg Bre nre ubGPT-5.6 (b) Held-out Held-out Held-out EdgeBench EdgeBench EdgeBench results tu ut result result b(b) H bH H dou dou ou Edg B BnBnhB nhresults hresults u out result GPT-5.6 GPT-5.6 Sol Sol GPT-5.6 GPT-5.6 Sol Sol Sol ou EdgeBench EdgeBench re uhu u(b)

ex

Auto-Research Loop Auto-Research Loop Next Next Next round round round Next Next round round Next round N N N

Next round Next roundN Nto GitHub GitHub GitHub Issue–PR Issue–PR Issue–PR pairs pairs pairs GitHub GitHub Issue–PR Issue–PR pairs pairs G G Hub GitHub uIssue–PR uu PR PR p p pairs GHub PR pw GitHub Issue–PR pairs How How How to to build to build build How How How toH to build build build Four Four Four retained retained N N How How to to build build How How tow to build build Four Four retained GHub PR pw H H How to build H wH How w to build H w How build How to build Fretained FuFour F uuretained n ndn nd dd Four retained Hto w H w Furetained G Hub Hub a umore u PR pH w H wmechanisms FmUnseen umUnseen n d a more efficient efficient a more a more aamore efficient efficient efficient mechanisms mechanisms Unseen Unseen Unseen Unseen Unseen UnseenUn aa more more more efficient efficient a more a more efficient efficient mechanisms mechanisms mechanisms a more efficient more efficient mechanisms a efficient more efficient a more efficient m m m ffi ffi ffi m m m ffi ffi ffi m m m h h n h n m B n hm k k Unseen Unseen Unseen Unseen Unseen Unseen Un een een een Un Un een een een m ffi mto reduce ffi m h n n Un mUn n Un n repositories How GitHub GitHub GitHub Issue Issue Issue GitHub GitHub GitHub Issue Issue Issue GitHub Issue GitHub Issue GitHub GitHub Issue Issue GitHub GitHub Issue Issue repositories How to reduce m ffi m ffi m h n m harness? harness? harness? harness? harness? harness? G G Hub GitHub Hub u Issue u G G Hub GitHub Hub u Issue u G G Hub uRun G GHub harness? Un een Un een HubRun uharness Hub u harness? uharness? harness? harness? harness? Benchmark Benchmark Benchmark Benchmark Benchmark Benchmark research harness? harness? Run harness harness harness Benchmark Benchmark Run harness Diverse Run Run harness Benchmark Benchmark Benchmark Benchmark Benchmark Benchmark Diverse research harness B n hma k B n hma k token cost? GLoop Hub u Run G Hub u Execution Diverse Diverse Diverse research research research Diverse Diverse Diverse research research research trace Four retained mechanisms Auto-Research Auto-Research Auto-Research Loop Loop fails fails fails Auto-Research Auto-Research Auto-Research Loop Loop Loop fails fails fails Diverse research Diverse research Action Action Action Fusion Fusion Fusion Auto-Research Loop fails Auto-Research Loop fails token cost? Diverse Diverse research research Diverse Diverse research research Diverse research Diverse research D D D h h h D D D h h h Execution trace Auto-Research Auto-Research Loop Loop fails fails Auto-Research Auto-Research Loop Loop fails fails environments Action Action Action Fusion Fusion Fusion Auto-Research Loop fails R Auto-Research Loop Benchmark Four retained Benchmark mechanisms A A A A A A D h fails D h A A A A Amulti-file environments to to handle to handle handle multi-file multi-file multi-file refactors refactors refactors to to handle to handle handle multi-file multi-file refactors refactors refactors environments environments environments environments environments environments tomulti-file multi-file refactors to handle multi-file refactors Idea pool environments Harness Harness Dnm h Read files environments Dnm h to to handle handle multi-file refactors refactors to to handle handle multi-file multi-file refactors refactors Harness to handle multi-file refactors to handle multi-file refactors environments environments environments environments mhandle m m m m Scientific Scientific Scientific Scientific Scientific Scientific A Rm A R environments environments n n n nm nm n n n n n n nm nm n n n Scientific Scientific A mHarness m Idea pool Harness Harness n nm n n nm n Harness H H H n n n Freeze Freeze Freeze Scientific Scientific Scientific Scientific H n Scientific Scientific Freeze Read files S S n nfi S n nnfinfififi S nnfiFreeze fifiFreeze SS Open Open Open Open Open Open Open Freeze Freeze Freeze Freeze Freeze Freeze Freeze Freeze Freeze ActionSFusion m Agent n nmm n Open n nm n Open Open Open Open Open Open O O O O O O Freeze Freeze Freeze Freeze Agent Agent H n Freeze Freeze O OEdit code Agent F F F F F F F F Action Fusion Sc en fic Sc en fic Agent Agent Find bugs Agent Edit code A A A A O Observe F Online Online Online Context Context Context Harness Harness Harness Harness Harness Harness Harness Harness Harness Harness Harness Harness SoL-Pi SoL-Pi SoL-Pi Python TypeScript GoGoGoGo Harness Harness Harness SoL-Pi OPython FO FOnline Context Find Run tests Python Python Python TypeScript TypeScript TypeScript Go Python Python TypeScript TypeScript TypeScript Python TypeScript Python TypeScriptGoGoGo GoHarness Merged Merged Merged PR PR PRbugs Merged Merged Merged PR PR PR Observe Merged PRAAction Merged PR Online Online Online Context Context Context Harness Harness Harness Harness Harness Harness Harness Harness SoL-Pi SoL-Pi SoL-Pi O C C C Harness Harness Harness Harness SoL H H H H H H H H H H H H SoL SoL P PPP TypeScript H H (Higher HO H SoL Compact Compact Compact Environment Environment Environment Action Action Python Python TypeScript TypeScript TypeScript Go Go Go Python Python Python TypeScript TypeScript TypeScript Go Go Go Run tests Gd GPR GPR GCost) G G Environment Action Python Merged Merged Merged Merged (Higher (Higher (Higher Cost) Cost) Cost) (Lower (Lower (Lower Cost) Cost) (Higher (Higher Cost) Cost) Cost) (Lower (Lower (Lower Cost) Cost) Cost) G G (Higher Cost) (Lower Cost) (Higher Cost) (Lower Cost) Merged Merged Propose M M M d d PR d PR PR M M M d d PR PR M d PR M d Harness Harness Harness Harness Online Context Inspect logs Fix: Fix: handle Fix: handle handle multi-file multi-file multi-file Fix: Fix: handle Fix: handle handle multi-file multi-file multi-file Compact Compact Compact Environment Environment Action Action Fix: handle Fix: handle multi-file C C mCm mSystems (Higher Cost) Cost) (Lower Cost) Cost) (Higher Cost) Cost) Cost) Cost) Environment Action Compact mmulti-file mm m A A A A Cost) (Lower Cost) Cost) (Lower Cost)H O(Lower H(Higher H(Higher HHC C C L(Lower Lw Lw C C H(Higher H(Higher H C C L(Lower Lw w C C C HC HSystems P Lw wC CH H C C LwCw C Harness Harness Harness Propose Systems Systems Systems Systems Systems Systems Ha Ha Ha n nnn SoL Ha G G Inspect logs Fix: Fix: handle handle multi-file multi-file Fix: Fix: handle multi-file multi-file M M Fix: handle multi-file Fix: handle multi-file m m m m m md PR m d PR refactors in parser refactors inm parser refactors refactors refactors in in parser in parser parser refactors refactors refactors inhandle in parser in parser parser Compact Systems Systems Systems Systems Systems Systems Sy Sy mmm m Sy Sy Sy mm Sy Sy mm refactors refactors refactors in in parser in parser parser

AI proposes AIAI proposes AI proposes proposes IA AI proposes proposes Aimprovement AI proposes A improvement improvement improvement mprovement improvement m m mimprovement m m Apply change m p Apply Apply Apply change change change A pm

p nge e change

MIT

Core contributors.

Code

arXiv:2609.20519v1 [cs.AI] 17 Sep 2026

3

NTU

Equal contribution.

Ha nObservationPack

Cost E Ma Math Math Ma Math h hhhRetain Ma R Ma

Cost Performance Performance Cost Cost Cost Performance Performance

Next round

R Retain Retain Retain R R R

Cost Cost Cost

Retain Retain Retain Retain

Performance Performance Performance mmm

R

Next round m

Reducer

m za E on

Cost

Sy em

ObservationPack

Evidence-Preserving Optimization Optimization Optimization Optimization Reducer Evidence-Preserving Optimization Optimization Op m on Optimization Op Op Op mm m a aaon aon on ≥

Ma h Retain

Math Math Math Math

Merge &Op freezem za

Ma Math Math Ma Ma Math Ma h hhh SoL-Pi Harness

on

Merge & freeze

Lower token cost Harness SoL-Pi Ma h

Lower token cost

R Retain Retain Retain R R R

R

Claude Opus Claude Opus Claude Claude Claude Opus Opus Opus 5 5 55 Claude Claude Claude Opus Opus Opus 5 555 Average API score cost ↑ (USD) ↓ API cost (USD) ↓ C aud Opu 5 C aud Opu 55cost Claude Claude Opus Opus Claude Claude Opus Opus Average Average Average score score score cost ↑cost ↑cost (USD) ↑(USD) (USD) ↓↓↓ C API cost cost (USD) (USD) (USD) ↓↓↓ Claude Opus Claude Opus C C aud C aud aud Opu Opu Opu 5API 5API 5API C aud C aud aud Opu Opu Opu 5API 5API

Claude Claude Claude Claude Codex Codex CodexClaude 34.7 34.7 $1,787 $1,787 Claude 43.7 43.7 $2,535 $2,535 Claude Claude Claude Claude Claude Claude Claude Claude Claude Code AP Code Code Code Codex Codex CodexCodex Codex Codex A Codex Codex CodexClaude 34.7 34.7 34.7 $1,787 $1,787 $1,787 43.7 43.7 43.7 43.7 $2,535 $2,535 $2,535 $2,535 $2,535 AP U ↓D↓$1,787 A U $1,787 D$1,787 A score AP U D$2,535 AP U D Code Code Code Code Code Code Code Code Codeaud Code C aud Opu Average score API score API score cost ↑cost ↑cost (USD) ↑(USD) API API API cost cost cost (USD) Average Average Average ↓score score ↑ ↑43.7 ↑43.7 Average Average API score score cost ↑cost ↑cost (USD) ↑(USD) ↓Code API API API cost cost (USD) ASo A Average APi API AP AP AP U U(USD) DUD↓D$1,339 AP AP AP A U(USD) A U(USD) D A U D↓D↓score A Average AOpu A API AP AP AP U U(USD) DUD↓D↓ Code AP AP AP U(USD) U(USD) DUD↓D↓ ↓ GPT 5 6Average 5API 5cost Pi So Pi Pi 44.8 Pi 44.8 GPT 5 6 $1,339 44.8 Pi C C Pi $1,741 44.8 $1,741 C C Pi PiC Pi Pi Pi Pi Pi Pi Pi C C PiC PiC Pi C Pi Pi Pi Pi PiC PiC Pi Pi Pi C 44.8 44.8 44.8 $1,339 $1,339 44.8 $1,339 44.8 44.8 $1,339 $1,339 $1,339 44.8 44.8 44.8 $1,741 $1,741 $1,741 44.8 44.8 44.8 $1,741 $1,741 $1,741 C Claude Claude Claude Claude Claude Claude Claude Claude Claude Claude Claude Claude C C C C C C C C C C Codex Codex Codex Codex Codex Codex 34.7 34.7 34.7 34.7 $1,787 $1,787 $1,787 43.7 43.7 43.7 43.7 $2,535 $2,535 $2,535 $2,535 Codex Codex Codex C 34.7 34.7 42.0 $1,787 $1,787 $1,787 43.7 43.7 $2,535 $2,535 C C C C C C C C C Code Code Code Code Code Code Code Code SoL-Pi SoL-PiA SoL-Pi SoL-Pi SoL-Pi SoL-Pi SoL-Pi 42.0 $894 $894 42.2 $1,158 42.2 $1,158 Code Code Code Code C C C C C C C C C C C AP USD SoL-Pi AP A USD 42.2 A 42.2 USD AP USD SoL-Pi SoL-Pi SoL-PiSoL-Pi SoL-Pi SoL-Pi SoL-Pi SoL-PiSoL-Pi SoL-Pi SoL-Pi $894 SoL-Pi SoL-Pi SoL-PiSoL-Pi SoL-Pi SoL-Pi SoL-Pi SoL-Pi SoL-Pi 42.0 42.0 42.0 $894 $894 $89442.0 42.0 42.0 $894 $894 42.2 42.2 $1,158 $1,158 $1,158 42.2 42.2AP $1,158 $1,158 $1,158 Pi Pi Pi Pi Pi Pi Pi Pi Pi Pi Pi Pi Pi Pi Pi Pi Pi Pi Pi Pi Pi 44.8 44.8 $1,339 $1,339 44.8 44.8 $1,339 $1,339 44.8 44.8 $1,741 $1,741 44.8 44.8 $1,741 $1,741 44.8 $1,339 44.8 $1,339 44.8 $1,741 44.8 $1,741

34.7 Ascore Average score score ↑ ↑34.7 ↑34.7 A Average A Average A

Pi

Pi

A

C 50.0% cost saved 50.0% cost saved 50.0% 50.0% 50.0% cost cost cost saved saved saved 50.0% 50.0% 50.0% cost cost cost saved saved saved 50 0% cos saved 50 0% cos saved m m 50.0% 50.0% cost cost 50.0% 50.0% cost cost xdex r or Claude Claude Code. Code. Cost Cost scales scales vary vary bySavings by model. Savings model. vs.vs. Codex Codex or or Claude Claude Code. Code. Cost Cost scales scales vary vary bysaved by model. model. 50.0% cost 50.0% cost or Claude Code. Cost scales vary by Savings model. vs. Codex or Claude Code. Cost scales vary by model. 50 50 50 0% 0% 0% cos cos cos saved saved 50 50 50 0% 0% 0% cos cos cos saved saved saved m m m m m m

C C or Claude Code. Cost scales vary by model. C gs vs. Codex or Claude Code. Cost scales vary by Savings model. vs. Codex C $894 SoL-Pi SoL-Pi SoL-Pi SoL-Pi SoL-Pi SoL-Pi SoL-Pi 42.0 42.0 $894 $894 42.0 Codex xdex or or Claude or Claude Claude Code. Code. Code. Cost Cost Cost scales scales scales vary vary vary by42.0 Savings by model. Savings by model. Savings model. vs.vs. Codex vs. Codex Codex or or Claude or Claude Claude Code. Code. Code. Cost Cost Cost scales scales scales vary vary vary by42.0 by model. by model. model. SoL-Pi SoL-Pi SoL-PiSoL-Pi SoL-Pi $894 $894 42.0 $894

m

50 0% cost saved m

C 54.3% cost saved 54.3% cost saved Csaved SoL-Pi SoL-Pi $1,158 $1,158 42.2 42.2 $1,158 $1,158 54.3% 54.3% 54.3% cost cost cost saved saved saved 54.3% 54.3% 54.3% cost cost saved saved SoL-Picost $1,158 42.2 $1,158 54 3% cos saved 54 3% cos saved 54.3% 54.3% cost cost 54.3% 54.3% cost cost 54.3% cost 54.3% cost 54 54 54 3% 3% 3% cos cos cos saved saved saved 54 54 54 3% 3% 3% cos cos cos saved saved saved

C C CSoL-Pi C SoL-Pi SoL-Pi SoL-PiSoL-Pi SoL-Pi 42.2 42.2 42.2

50 0% cost saved

54 3% cost saved

54 3% cost saved

F gure 1 SoL-P d scovers a more token-effic ent harness through automated research (a) SoL-P : Sca ng Auto-Research Loop Prepared research env ronmen s supp y asks o an AI runn ng he base harness A research AI nspec s s execu on races proposes cand da e changes and fi ers he dea poo hrough capab y and effic ency ga es Four re a ned mechan sms are n egra ed and refined n o SoL-P before he harness s frozen for eva ua on on unseen benchmarks The race deas and ga e symbo s are schema c capab y s checked w h n fixed o erances and he d-ou resu s never feed back n o search (b) Examp e resu s on EdgeBench average score and API cos for he na ve harnesses (Codex w h GPT-5 6 So and C aude Code w h Opus 5) P and he comp e e four-mechan sm SoL-P harness SoL-P reduces API cos by 50 0% re a ve o Codex on GPT-5 6 So and by 54 3% re a ve o C aude Code on Opus 5

© 2026 NV D A A

gh s ese ved

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

1. Introduction Advances in foundation models enable agents to tackle increasingly open-ended tasks over longer horizons with less supervision [1, 2, 3, 4]. This shift supports applications such as autonomous research, software engineering agents, self-evolving personal assistants, and early forms of recursive self-improvement (RSI) [5, 6, 7, 8]. As agents operate over longer horizons, task-level token efficiency becomes a first-order systems concern [9, 10]. Existing efficiency work has primarily focused on lowering the cost per token through faster attention kernels and serving infrastructure [11, 12], model compression techniques such as quantization [13, 14], or the use of cheaper models [15, 16]. In this paper, we explore an orthogonal direction: improving token use through the agent harness that mediates interactions between the model and its environment. Harness-level optimization can improve efficiency without additional model training, complementing infrastructureand model-level approaches [17, 18]. However, optimizing a harness is difficult in practice. Since tool use, context management, verification, delegation, recovery, and termination are tightly coupled, a change that is locally beneficial may cause downstream failures or shift token costs to later stages of execution. In practice, harness development often requires substantial human effort to inspect long execution traces, identify recurring failure modes, and translate these observations into code changes. This process is costly and difficult to scale across tasks and environments. To accelerate this process, we adopt an RSI-inspired approach in which an AI optimizer iteratively improves the agent harness for token efficiency. In the SoL-Pi: Scaling Auto-Research Loop workflow (Figure 1(a)), the research AI observes execution traces from a separate agent running the base harness, proposes candidate changes, and tests them in prepared research environments. Capability and efficiency checks determine which candidates are retained, while development results guide subsequent iterations. Recent work has demonstrated the feasibility of automated harness improvement. Meta-Harness searches over executable harness programs and evaluates their transfer to held-out datasets and models [17], while Recursive Harness SelfImprovement (RHI) iteratively refines prompt-level specifications of the agent loop for individual tasks [19]. However, a recent study using held-out tasks finds that evolved harnesses can overfit the tasks used during search and provide only marginal gains on unseen tasks [20]. These findings motivate a clear separation between search feedback and final evaluation [21]. To discover transferable efficiency improvements, we introduce SoL-Pi, a system for discovering harness improvements that transfer beyond the tasks used during search. SoL-Pi organizes autonomous research as a broad-to-deep funnel that separates candidate development from held-out validation and scales through isolated search lineages. Its design is guided by three principles: • Breadth and depth. Breadth expands hypothesis coverage, while depth repeatedly implements, reviews, and hardens promising candidates. • Independent validation. Held-out evidence is evaluated only after a candidate is frozen and never returns to search, preventing validation failures from being patched into task-specific solutions. • Scalable orchestration. Isolated, disposable lineages let the funnel expand across more ideas and environments without coupling failures across candidates. Guided by these principles, we scale AI-led auto-research across roughly ∼150 proposed directions and ∼500 executable environments, comprising more than 3,000 runs and more than 60,000 agent–environment interactions. The search yields four mechanisms that together form SoL-Pi. On EdgeBench [22], SoL-Pi achieves performance comparable to Pi across GPT-5.6 Sol and Opus 5 while reducing recorded token traffic by 44.7–49.0% and API cost by about one third. Beyond these results, SoL-Pi suggests that the lasting value of RSI may lie in a search process that scales across public environments to discover reusable improvements.

2. Method 2.1. Harness Auto-Research for Token Efficiency Our search pipeline frames harness improvement as an RSI-inspired search for reusable efficiency mechanisms, starting from a broad pool of agent-generated hypotheses. The research agent analyzes execution trajectories from the base harness to identify recurring sources of overhead, builds an idea pool of candidate harness changes, and tests their effects in development environments. To qualify for acceptance, a mechanism must improve efficiency beyond a single development environment while preserving the agent’s ability to complete the required work. 2

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

SEARCH / DEVELOPMENT

Trajectory Rollouts

Map–Reduce Analysis

FREEZE

revise

Mechanism Proposal

Candidate Implementation

next iteration

Independent Review

Development Validation

HELD-OUT

Held-Out Evaluation no feedback to search

Figure 2 | Development feedback and held-out evaluation remain separate throughout the search. Execution trajectories guide mechanism proposals and implementation, while independent review and development validation inform revisions. Held-out evaluation occurs only after the candidate is frozen, and its results never feed back into the search loop. Before experimentation begins, capability metrics, acceptable tolerances, and efficiency metrics are fixed and remain unchanged throughout the search. These metrics and tolerances are strictly isolated from the optimizing agent’s control to prevent it from gaming the acceptance criteria. Candidate selection applies two sequential gates: every capability metric must remain within its predeclared tolerance, and the candidate must improve at least one declared efficiency metric. Among candidates that pass both gates, the pipeline retains the nondominated results under the declared metrics. Since each mechanism is explored independently, integration remains part of the SoL-Pi: Scaling Auto-Research Loop workflow in Figure 1(a): we combine retained mechanisms into SoL-Pi and refine its hyperparameters and implementation while preserving capability. We reserve EdgeBench for final validation, keeping it isolated from the search process. The harness and acceptance rule are frozen before evaluation. Held-out results never feed back into the Auto-Research Loops: a failed validation rejects the candidate without triggering further optimization. 2.2. Broad-to-Deep Harness Search Our search pipeline allocates research effort in two stages: an outer stage explores a broad pool of mechanism hypotheses, and an inner stage develops selected hypotheses independently. The outer search begins with 152 proposed directions across six proposal families: context, progress, tools, delegation, prompt and policy, and improvement and evaluation. Before rollout budgets are assigned, Oracle Analysis examines existing development trajectories to identify avoidable work in the base harness. Each selected direction must identify a concrete source of overhead and propose a harness change to address it. Independent development allows unpromising directions to terminate without affecting other experiments. The proposal families classify hypotheses by their origin rather than constrain where changes are implemented. For example, ObservationPack originates from two context hypotheses but ultimately modifies the observation boundary. This distinction helps organize the search while leaving the architecture of each mechanism open. Each independent search follows the conventional autoresearch cycle: propose a change, implement it, run a fixed experiment, inspect the result, and retain, revise, or discard the candidate [23]. We extend this cycle with an iterative implementation loop based on the Ralph Loop [24]. The implementer refines the candidate to meet an explicit completion criterion, then an independent reviewer checks it before evaluation. Failed reviews trigger revision. Each iteration may yield multiple exploration trajectories. Independent analyzers examine one trajectory each for repeated actions, context growth, large observations, or sparse diagnostic signals. A reducer combines their findings into a candidate-level summary that guides the next proposal. Figure 2 summarizes the workflow and its isolation boundary. We run each independent search as a disposable instance of a shared skill template containing a minimal research loop and operating instructions. Each experiment copies the template, sets its parameters, and runs to completion, retaining the candidate and evidence while discarding modified orchestration code. Search breadth comes from running more isolated loops; depth comes from repeated refinement within each loop. These counts describe the scope of our search; they do not establish a scaling law. 2.3. Search Environments Our search pipeline uses separate environments for mechanism discovery and transfer evaluation. The search set comprises 535 executable environments: 495 repository tasks derived from GitHub issue–pull request pairs and 40 synthetic tasks with executable success verifiers. Together, they support automatically evaluated search grounded in real 3

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

software changes and open-ended problem solving. Repository-derived environments. Each environment pairs a GitHub issue with its pre-fix repository state and offline dependencies, using the accepted patch and change history as a reference trajectory. The pull request and regression test are hidden from the agent. We retain only environments whose test fails before the patch and passes afterward. Verifier-driven environments. We first generate an executable verifier that defines success, then construct a task environment around it. This allows multiple valid solution paths without requiring a reference trajectory. The 40 synthetic environments primarily use a Terminal-Bench-2-style verifier interface [25], extending the search beyond repository histories while preserving automatic evaluation. Figure 3 summarizes both construction paths and the isolation boundary between mechanism search and held-out evaluation.

REPOSITORY-DERIVED issue–PR

|

pre-fix repo

VERIFIER-DRIVEN verifier

|

task env.

FREEZE

495 environments |

HELD-OUT

fail–pass test

535 search environments

Mechanism search

EdgeBench

40 environments |

multiple solutions

no feedback to search

Figure 3 | Two environment families support mechanism discovery while keeping EdgeBench held out. Repositoryderived environments use pre-fix repositories and hidden regression tests verified to fail before the accepted patch and pass afterward. Verifier-driven environments allow multiple solution paths under executable success criteria. EdgeBench remains isolated from search and is used only after the harness is frozen. 2.4. Discovered Harness Mechanisms The search yields four reusable mechanisms that pass capability-constrained selection: Action Fusion, Online Context Compact, ObservationPack, and Evidence-Preserving Reducer. They target action execution, context management, observation storage, and delegated reading, respectively, reducing repeated work while preserving information needed for later decisions. Figure 4 shows their execution paths and fallback boundaries. Action Fusion. Base Pi often edits a file and then issues a separate command to test, build, or run it. Action Fusion combines both actions into one tool request and returns their outcomes in a single observation, eliminating an intermediate model round trip. Commands that require inspecting the mutation result remain separate. Online Context Compact. Online Context Compact uses plan-step completion to reconsider when to compact the context. The agent maintains its plan through update_plan. At each completion boundary, the harness estimates remaining model requests from the observed requests between completed steps and the number of unfinished steps. It caps this estimate by the requests that would fill the current context window at the observed growth rate. The cost gate compares projected input savings with the estimated extra cost of rewriting the prompt cache. Later compactions also account for unrecovered rewrite costs and require a larger savings margin. At these boundaries, the harness invokes Pi’s native compaction when this gate passes or when context usage approaches the window limit, provided compaction can shorten the context. ObservationPack. Large tool outputs can recur in later requests even when little of their content remains relevant. ObservationPack locally archives results exceeding the threshold (10 KiB) and sends them in full for the next two provider requests. From the third request onward, it substitutes a stable handle, the original size, and a short excerpt of complete head and tail lines. The agent can retrieve exact pages through the handle as needed. Smaller results remain unchanged.

4

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

(a) Action Fusion

(b) Online Context Compact Context length

API call

API call

API call

Baseline

API call 2

1.0×

1.3×

API call 3

API call 4

1.6×

1.9×

Baseline Read result Choose run

Edit / Write

Receive output

Keep full history

API call

API call

3→2

Action Fusion

API calls File changed + command output

Edit / Write + then_run

1 call saved

(c) ObservationPack Baseline

API call 1

Large tool result > 10 KB text

ObservationPack

Stored original

API call 1

API call 2

Subtask complete

1.0×

1.3×

Compact if savings > cost

Online Compact

API call 3

API call 4

0.55×

0.75×

Cost gate

(d) Evidence-preserving reducer API call 2

Full result in every API call

API call 3

API call 4

Build / test result > 4 KB

Archive exact original

Small model

Candidate receipt

Low-cost API call

Higher input cost

Select evidence API call 1

API call 2

API call 3

API call 4

Stable handle + 1 KB excerpt Large result

API call 1

First 2 calls: full result

Lower input cost

Deterministic verifier Schema · hash Status Exact quotes Smaller receipt

Pass

Verified receipt

Fail

Original result

Frontier API call

On demand: recall exact original chunk

Figure 4 | The four retained mechanisms act at different points in the agent–environment loop. (a) Action Fusion combines a mutation and its follow-up command into one request, reducing API calls from three to two. (b) Online Context Compact evaluates compaction at subtask completion and applies it only when projected savings exceed the rewrite cost. (c) ObservationPack sends large results in full for the first two provider requests, then replaces them with a stable handle and a 1 KB excerpt. Exact original chunks remain available on demand. (d) Evidence-Preserving Reducer uses a low-cost model to compact results, with deterministic verification and fallback to the original on failure. Grey marks baseline traffic, green marks mechanism paths and savings, and red marks excess cost or verification failure. Evidence-Preserving Reducer. Evidence-Preserving Reducer compresses build and test logs of at least 4 KiB from a predefined set of commands. File reads and search results bypass the reducer. The harness archives the exact output and asks a lower-cost model to extract key evidence into a compact receipt. A deterministic verifier checks the receipt’s schema, source hash, exit status, exact quotes, and size. The harness falls back to the original log if verification fails, credentials are suspected, or the receipt provides no size reduction. The reducer processes tool results before ObservationPack projects the model context; ObservationPack recognizes the reducer’s receipt marker and skips those results to preserve the verified evidence. The auxiliary model extracts evidence, while the main agent retains responsibility for diagnosis and action selection. The four mechanisms support different stages of the agent’s workflow. When the agent edits code, Action Fusion combines the edit with a follow-up command. When the environment returns output, Evidence-Preserving Reducer extracts verified evidence from build and test logs, while ObservationPack avoids repeatedly sending large results in full. When the agent completes a plan step, Online Context Compact checks whether compacting the accumulated context would save tokens. These mechanisms therefore address complementary sources of overhead. We evaluate the combined harness end to end to verify their joint effect on capability and efficiency. 2.5. Implementation Details and Backend Setup We implement all four mechanisms as extensions to Pi. Action Fusion adds an optional follow-up command to file-mutation tools and returns both outcomes in one observation. Online Context Compact checks the cache-cost gate at plan-step completion and also supports compaction near the context limit. The gate estimates cache-rewrite overhead from the context size and the cache-write/read price ratio; it does not separately price the summarization call. ObservationPack archives large outputs locally and provides stable handles for exact retrieval. Evidence-Preserving Reducer uses GPT-5.6 Luna at high to extract evidence and verifies the resulting receipt before passing it to the main 5

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

agent. Before held-out evaluation, we freeze the mechanism source, configuration, metrics, capability tolerances, and acceptance rule. All EdgeBench outputs remain outside the search loop. Of EdgeBench’s 51 public tasks, 11 are used for one-way acceptance of frozen candidates; the remaining 40 are reserved for final evaluation of generalization.

3. Experimental Evaluation We evaluate SoL-Pi on EdgeBench [22], Terminal-Bench 4 [26], IMO 2026 [27], and a kernel-optimization benchmark [28]. We report average score, token traffic (in billions), and API cost. Token efficiency is measured as API cost per unit of aggregate task score. Sec. 3.1 compares SoL-Pi with native and third-party harnesses and tests whether it transfers from GPT-5.6 Sol to Opus 5 without modification. Sec. 3.2 reports Terminal-Bench 4 completion rates and IMO 2026 Lean 4-verified problem counts, together with costs for both. Sec. 3.3 evaluates SoL-Pi in a coordinated agent swarm. Sec. 3.4 analyzes individual and combined mechanism effects and activation patterns across backends and configurations. Sec. 3.5 traces the development and selection of Action Fusion. 3.1. Overall Comparison EdgeBench currently releases 51 of its 134 tasks as open source, and our evaluation uses this public set [22]. Table 1 compares eight evaluated configurations under GPT-5.6 Sol and lists each primary backend; the EdgeBench official GPT-5.5 row is an unranked score-only reference. We report two SoL-Pi operating points. SoL-Pi [Efficiency] is the fixed complete four-mechanism stack, reported as the token-efficiency-oriented point. SoL-Pi [Performance] is the single-mechanism configuration with the highest average score, selected separately for each backend from Table 4; under GPT-5.6 Sol, it corresponds to ObservationPack. The Efficiency point uses a total of 1.10 B tokens, 49.0% fewer than Pi, while retaining 93.7% of Pi’s average score (42.0 vs. 44.8). Its token cost is 33.2% lower than Pi’s. The Performance point raises average score from 44.8 to 47.2, a 5.3% gain, while reducing token traffic by 6.1% and improving token efficiency by 9.8%. Table 1 | Harness comparison on EdgeBench. SoL-Pi [Efficiency] is the fixed complete-stack point selected for token efficiency, while SoL-Pi [Performance] is the GPT-5.6 Sol single-mechanism configuration with the highest average score in Table 4. Best and next-distinct values among the eight measured rows are bold and underlined. The EdgeBench official GPT-5.5 @2h checkpoint (31.2) is unranked [29]; dashes denote unreported token traffic and token cost. Harness

Token Cost ($)↓

Recorded Token Traffic (B)

Backend Input↓

Cache R.↓ Cache W.↓

Output↓

Total↓

Avg. Score↑

Token Eff. ($/score)↓

EdgeBench official @2h

GPT-5.5

–

–

–

–

–

–

31.2

–

Codex [30] OpenSquilla [31] Oh-My-Pi [32] OpenCode [33] Oh-My-Opencode [34] Pi [35]

GPT-5.6 Sol GPT-5.6 Sol GPT-5.6 Sol GPT-5.6 Sol GPT-5.6 Sol GPT-5.6 Sol

0.0053 0.0001 0.1135 0.0012 0.0044 0.0011

3.0287 1.2533 2.0448 2.2625 2.4373 2.1326

0.0145 0.0776 0.0484 0.2865 0.1286 0.0141

0.0052 0.0044 0.0168 0.0164 0.0121 0.0059

3.0537 1.3353 2.2235 2.5668 2.5825 2.1538

1,787 1,243 1,832 3,422 2,678 1,339

34.738 24.506 26.921 29.552 38.523 44.833

1.0086 0.9945 1.3347 2.2704 1.3633 0.5855

SoL-Pi [Efficiency] SoL-Pi [Performance]

GPT-5.6 Sol GPT-5.6 Sol

0.0009 0.0002

1.0605 2.0005

0.0316 0.0160

0.0061 0.0056

1.0990 2.0224

894 1,271

42.003 47.208

0.4174 0.5280

Note. API prices change over time; all token costs in this paper use the API prices as of August 17, 2026.

Table 2 compares the native harness, Pi, SoL-Pi [Efficiency], and SoL-Pi [Performance] on GPT-5.6 Sol and Opus 5. To test transfer, we apply SoL-Pi, developed with GPT-5.6 Sol, to Opus 5 without further search or adaptation. On Opus 5, it retains 94.3% of Pi’s average score while reducing token traffic by 44.7% and API cost by 33.5%, relative to Pi’s point estimates. These results, alongside the activation patterns in Sec. 3.4, provide preliminary evidence of transfer to an unseen LLM backend. Figure 1(b) summarizes the native-harness, Pi, and complete-stack SoL-Pi results on EdgeBench across the two backends. These bars are example results from our broader benchmark evaluation; Terminal-Bench 4 and IMO 2026 are reported in Sec. 3.2.

6

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

Table 2 | SoL-Pi configurations and transfer across models on EdgeBench. SoL-Pi [Efficiency] combines all four mechanisms and transfers from GPT-5.6 Sol to Opus 5 without further search or adaptation. SoL-Pi [Performance] uses the highest-scoring single mechanism for each model in Table 4: ObservationPack for GPT-5.6 Sol and Action Fusion for Opus 5. Within each model block, best values are bold and next-distinct values are underlined. Recorded Token Traffic (B)

Harness Input↓

Cache R.↓

Cache W.↓

Output↓

Token Cost ($)↓

Avg. Score↑

Token Eff. ($/score)↓

Total↓

GPT-5.6 Sol (Search Backend) Codex [30] Pi [35]

0.0053 0.0011

3.0287 2.1326

0.0145 0.0141

0.0052 0.0059

3.0537 2.1538

1,787 1,339

34.738 44.833

1.0086 0.5855

SoL-Pi [Efficiency] SoL-Pi [Performance]

0.0009 0.0002

1.0605 2.0005

0.0316 0.0160

0.0061 0.0056

1.0990 2.0224

894 1,271

42.003 47.208

0.4174 0.5280

Opus 5 (Held-Out Backend) Claude Code [36] Pi [35]

0.0002 0.0000

1.8116 2.3203

0.1701 0.0348

0.0226 0.0145

2.0045 2.3697

2,535 1,741

43.689 44.756

1.1377 0.7625

SoL-Pi [Efficiency] SoL-Pi [Performance]

0.0000 0.0000

1.2626 2.0520

0.0352 0.0352

0.0123 0.0144

1.3101 2.1016

1,158 1,605

42.224 50.482

0.5376 0.6235

Table 3 | Harness comparison on Terminal-Bench 4 and IMO 2026. Terminal-Bench 4 uses 63 CPU-only tasks; the IMO evaluation covers six problems. All IMO runs use GPT-5.6 Sol (xhigh). Costs are in USD, and best values within each benchmark are bold. Terminal-Bench 4

Harness

Codex [30] Pi [35] SoL-Pi

IMO 2026

Solved Tasks (out of 63)↑

Total Model Cost ($)↓

Cost / Solved Task ($)↓

Pass (out of 6)↑

Total Model Cost ($)↓

Cost / Passed Problem ($)↓

18 18 15

272.35 286.45 211.12

15.13 15.91 14.07

5 3 3

114.47 75.95 62.69

22.89 25.32 20.90

Note. GPU-dependent Terminal-Bench 4 tasks were excluded due to infrastructure limits. The IMO evaluation uses AxiomMath/IMO2026’s formal statements and follows Humanfia’s verification setup [37, 38]. Each problem is capped at 150 minutes to limit unproductive looping, and costs include all model activity within this budget.

3.2. Evaluation on Terminal-Bench 4 and IMO 2026 We also compare Codex, Pi, and SoL-Pi on 63 CPU-only tasks from Terminal-Bench 4 [26]. Table 3 reports the number of solved tasks, total model cost, and cost per solved task, with all costs reported as API costs. Codex and Pi each solve 18 tasks, while SoL-Pi solves 15. Compared with Pi, SoL-Pi reduces total model cost by 26.3% ($211.12 vs. $286.45) and cost per solved task by 11.6% ($14.07 vs. $15.91). These results suggest that the harness’s efficiency gains generalize beyond EdgeBench, with lower total cost. We further evaluate the three harnesses on IMO 2026 [27] using GPT-5.6 Sol (xhigh), requiring each solution to be formalized and verified in Lean 4. SoL-Pi passes three of six problems at a total model cost of $62.69, achieving the lowest cost per passed problem ($20.90), compared with $22.89 for Codex and $25.32 for Pi. 3.3. Efficient Agent Swarms We evaluate SoL-Pi in a multi-agent kernel-optimization experiment measured in simulated machine cycles [28]. We compare three configurations in one two-hour run each: a single Codex agent, a Codex coordinator with 20 Pi baseline workers, and a Codex coordinator with 20 SoL-Pi workers using all four mechanisms. The single agent and coordinators use GPT-5.6 Sol at xhigh; all workers use GPT-5.6 Luna at xhigh. Every run starts from the same frozen starter, which requires 147,734 cycles, with fresh sessions and no solutions or notes from earlier runs. In both swarms, workers form five groups of four, with independent workspaces and a shared evidence board within each group. Workers exchange notes to reproduce or combine promising findings, while the coordinator relays findings across groups. The shared best result is updated only when the coordinator submits an immutable candidate snapshot and independent verification confirms a strict improvement.

7

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

Main Coordinator [Codex-Sol-xhigh] + 20 workers [SoL-Pi] Single Codex-Sol-xhigh Main Coordinator [Codex-Sol-xhigh] + 20 workers [Pi-baseline]

Coordinator

Codex · GPT-5.6 Sol · xhigh

Group 02

Group 03

Group 04

Group 05

5 6 7 8

9 10 11 12

13 14 15 16

17 18 19 20

Group board

Group board

Group board

Group board

Group board

Cycles ↓ (log scale)

Group 01 1 2 3 4

Verified progress

Tasks + frontier ↓ · findings ↑

Immutable snapshots

0

1k

80

Main-agent acceptance

Frozen tests + cycles

Strict gain → commit + JSON

(a) Swarm architecture

20 workers [SoL-Pi]

1,333

1,366

Single 20 workers Codex- [Pi-baseline] Sol-xhigh

$60.11

$82.12 $39.20

100k 0

Verification

1,127

Model cost (USD) ↓

10k

Snapshots + compact evidence

Candidate pool

Final cycles ↓ 1,500

0

30

60

90

Cumulative time (min)

120

20 workers [SoL-Pi]

Single 20 workers Codex- [Pi-baseline] Sol-xhigh

(b) Verified search progress and two-hour results

Figure 5 | Agent swarm architecture and two-hour search results. (a) Group boards support local collaboration, and the coordinator shares findings across groups and submits candidates for independent verification. (b) Filled points show correct submitted candidates, open circles mark accepted commits, and step lines track the best accepted result. The cost chart shows cumulative API cost. Figure 5(b) shows that the SoL-Pi swarm reaches 1,127 cycles at $60.11, compared with 1,333 cycles at $39.20 for the single agent and 1,366 cycles at $82.12 for the Pi baseline swarm. Under the same execution budget, the SoL-Pi swarm reduces API cost by 26.8% relative to the Pi baseline swarm, while the single agent remains least expensive. All final candidates pass the official correctness check. The SoL-Pi swarm and single agent pass all eight speed thresholds, whereas the Pi baseline swarm passes seven, missing the final threshold of fewer than 1,363 cycles. These results suggest that the value of SoL-Pi may extend beyond individual agents: a more efficient harness could help agent swarms and multi-agent systems turn a fixed budget into more effective collective exploration. 3.4. Learned Mechanisms and Backend Behavior We assess the standalone contribution of each independently learned mechanism through an add-one evaluation. Table 4 compares the Pi baseline, four variants that each add a single mechanism to Pi, and the complete SoL-Pi stack under both model backends. Every component reduces the total token count under both backends. ObservationPack records the highest average score under GPT-5.6 Sol, while Action Fusion does so under Opus 5. The complete stack represents the efficiency-oriented point and has the lowest total token count and token cost in both backend blocks. Figure 6 shows that mechanism activation varies across backends. Trigger rate is the fraction of tasks on which a mechanism activates, while trigger intensity is the mean number of activations per triggered task. Both are lower under Opus 5, which may reflect the harness being optimized exclusively on GPT-5.6 Sol trajectories. Nevertheless, every configuration evaluated under Opus 5 improves token efficiency on its triggered tasks, and the complete stack preserves a similar aggregate score–efficiency trade-off across the two backends (Table 2). To assess the effects of combining mechanisms, Figure 7 compares each standalone configuration with the full stack. ObservationPack becomes more selective in the full stack, potentially due to overlap with Evidence-Preserving Reducer on observation-heavy trajectories. For every mechanism, the full stack achieves a larger token-efficiency gain than the standalone configuration on their respective triggered-task subsets. This pattern is consistent with complementarity in the evaluated stack, although comparisons on each configuration’s own triggered-task subset do not isolate interaction effects.

8

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

Table 4 | Single-component evaluation on EdgeBench. Each component row adds one mechanism to Pi; SoL-Pi [Efficiency] combines all four. Light green and † identify the configuration selected as SoL-Pi [Performance] for each backend; dark green highlights SoL-Pi [Efficiency]. Token cost and token efficiency are computed using fixed API prices. Best values are bold, and the next-best distinct values are underlined. Recorded Token Traffic (B)

Configuration Input↓

Cache R.↓

Cache W.↓

Output↓

Token Cost ($)↓

Avg. Score↑

Token Eff. ($/score)↓

Total↓

GPT-5.6 Sol Pi Baseline [35] + Action Fusion + Online Context Compact + Evidence-Preserving Reducer + ObservationPack †

0.0011 0.0011 0.0008 0.0051 0.0002

2.1326 1.8718 1.2602 1.9089 2.0005

0.0141 0.0180 0.0217 0.0181 0.0160

0.0059 0.0060 0.0055 0.0055 0.0056

2.1538 1.8968 1.2881 1.9375 2.0224

1,339 1,235 935 1,200 1,271

44.833 46.664 41.993 44.630 47.208

0.5855 0.5190 0.4365 0.5274 0.5280

SoL-Pi [Efficiency]

0.0009

1.0605

0.0316

0.0061

1.0990

894

42.003

0.4174

Opus 5 Pi Baseline [35] + Action Fusion † + Online Context Compact + Evidence-Preserving Reducer + ObservationPack

0.0000 0.0000 0.0000 0.0000 0.0000

2.3203 2.0520 1.8618 1.8685 1.3960

0.0348 0.0352 0.0388 0.0315 0.0356

0.0145 0.0144 0.0145 0.0131 0.0102

2.3697 2.1016 1.9152 1.9131 1.4418

1,741 1,605 1,537 1,456 1,176

44.756 50.482 49.155 43.405 47.047

0.7625 0.6235 0.6130 0.6578 0.4899

SoL-Pi [Efficiency]

0.0000

1.2626

0.0352

0.0123

1.3101

1,158

42.224

0.5376

Cache Reuse and Total Cost. Shortening context can reduce prompt-cache reuse when it changes a previously cached prefix [39]. Cached input also incurs a cost, so preserving a long prefix is not always the cheapest choice over an entire task. Online Context Compact and ObservationPack trade some prefix reuse for less repeated input. Table 4 shows this trade-off: with GPT-5.6 Sol, the complete stack reduces cache-read traffic from 2.1326 B to 1.0605 B tokens, while cache-write traffic increases from 0.0141 B to 0.0316 B. Despite the additional cache-write traffic, total model cost falls from $1,339 to $894. Every single-mechanism configuration also lowers cost per score point relative to Pi under both backends. These results support evaluating the full task cost alongside task quality, rather than cache reuse alone. GPT-5.6 Sol

80

92.2% 80.4%

78.4% 76.5%

68.6%

60 39.2%

40

33.3% 23.5%

20

0

Avg. triggers / triggered task (log scale)

Triggered tasks (%)

100

(b) Trigger Intensity 100

(c) Token Efficiency Gain

70.58

50

20

13.54

12.75 8.59

10

5.40

5

3.42 2.04 1.59

2 1

Action Fusion

Online Evidence- Observation Context Preserving Pack Compact Reducer

Action Fusion

Online Evidence- Observation Context Preserving Pack Compact Reducer

Triggered-task token efficiency gain (%)

(a) Task Trigger Rate

Opus 5

40

30

35.8%

24.7% 25.5% 20.1%

20.9% 23.3%

20 10.2%

10

7.2%

0 Action Fusion

Online Evidence- Observation Context Preserving Pack Compact Reducer

Figure 6 | Backend-dependent activation of the retained mechanisms on EdgeBench. Panels show trigger rate, trigger intensity on a logarithmic scale, and token-efficiency gain.

9

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

Enabled Alone

(a) Task Trigger Rate

(b) Trigger Intensity 100

92.2% 92.2% 92.2%

(c) Token Efficiency Gain

70.58 55.26

Triggered-task token efficiency gain (%)

100

All Mechanisms Enabled

50 80.4%

78.4%

Avg. triggers / triggered task (log scale)

Triggered tasks (%)

80

56.9%

60 41.2% 39.2%

40

20

20 12.75

10

8.59

7.67

5 2.04

2.40 1.66

2

0

40

30

Online Evidence- Observation Context Preserving Pack Compact Reducer

28.7%

24.7% 20.9%

20.1%

20

10.2%

10

1 Action Fusion

29.0%

29.0%

27.4%

0 Action Fusion

Online Evidence- Observation Context Preserving Pack Compact Reducer

Action Fusion

Online Evidence- Observation Context Preserving Pack Compact Reducer

Figure 7 | Merge behavior of independently explored mechanisms under GPT-5.6 Sol (xhigh). Paired bars compare standalone and full-stack trigger rate, trigger intensity, and token-efficiency gain for each mechanism. Gains use each configuration’s own triggered-task subset and corresponding disabled baseline, so comparisons are descriptive. ObservationPack becomes more selective in the full stack, while every mechanism shows a larger token-efficiency gain than in its standalone configuration, a pattern consistent with complementarity in the evaluated stack. (c) Task Score

(b) Action Fusion Trigger Rate

Run independent final checks Result · release behavior confirmed Decision · ready

03 PROMPT OPTIMIZATION

84.2 87.0

80

77.8

80.8

81.4 79.7 77.8

74.1

70 1

2

3

4

5

6

7

8

9

10

1

2

3

4

Search batch

80

(e) Model Turns

6

7

8

9

10

(f) Total Tokens (M) 34

1,400 1,386 Saved 149 turns (10.8%)

85.1

Saved 3.74 M tokens (11.5%)

32 1,300

60

5

Search batch

(d) Adjacent Action Share 100

Share of next actions (%)

Measure missed adjacent actions Result · 12.3% direct headroom Decision · optimize

85

75

28.3%

20%

1-6 · 6 iterations

0 · 1 iteration

69.9% 66.7%

60.9%

58.5% 50.9%

40%

Stabilize Action Fusion execution Result · 87.7% usage · 0 invalid calls Decision · stable baseline

01 ORACLE ANALYSIS

62%

60%

85.3

78.9%

76.9%

80%

7-24 · 18 iterations

Remove Iteration 8; refine prompt and schema Result · Search batch 10 performed best overall Decision · selected · 100% trigger · 87.0 task score

02 BASELINE BUILD

88.9

90

Task score

26-27 · 2 iterations

Trigger rate (%)

(a) Optimization Process 04 FINAL VALIDATION

100%

100%

32.38

30

40

1,237

1,200

28.64

28

20 14.9 0

1,100 bash

others

26 Baseline

All triggered

Baseline

All triggered

Figure 8 | Action Fusion from opportunity to retention. Panel (a) summarizes the 27 recorded iterations across four stages. Stage 03 records 18 prompt-optimization explorations, whereas panels (b)–(c) show ten retained steps: some steps evaluate multiple candidates in parallel, and only the best result from each parallel batch is retained. Panel (d) reports observed adjacent-action composition; panels (e)–(f) project model turns and tokens under full triggering. 3.5. From Observation to a Retained Mechanism We use Action Fusion as a case study of how a candidate was discovered and retained. As shown in Figure 8, the recorded lineage spans 27 iterations across four stages: oracle analysis, baseline construction, prompt and tool-schema optimization, and final held-out validation. The oracle-analysis panels (d)–(f) identify repeated adjacent actions and project an 11.5% token reduction under full triggering, motivating a dedicated Action Fusion lineage. During baseline construction, prompt-only triggering proved

10

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

unreliable. The agent therefore extended the tool schema to expose the fused action directly, establishing a stable baseline with no invalid calls. It then refined the prompt and schema on development tasks, introducing trigger rate alongside task score as an intermediate acceptance metric. The final configuration was selected using both metrics, frozen, and retained after held-out validation. This case illustrates how a lineage-local auto-research loop can stabilize a mechanism’s interface and triggering behavior. The agent’s introduction of trigger rate also demonstrates that auto-research can develop mechanism-specific intermediate metrics to guide optimization alongside end-task performance.

4. Related Work 4.1. Agent Harnesses and Automated Agent Design An agent’s behavior depends not only on its underlying model, but also on the harness that presents state, exposes actions, and processes feedback. SWE-agent showed that changing the agent–computer interface around a fixed model can materially affect repository-level software-engineering performance [2]. Subsequent work has examined interface design, context management, tool use, and delegation as system-level choices that shape agent behavior [40, 41]. These design choices become increasingly consequential in long-horizon tasks, where accumulated interaction histories and environment observations can increase context size and execution overhead [9, 22]. Automated agent design explores several optimization targets. GEPA improves prompts through reflection on execution trajectories [42], while ADAS formulates agent design as code search [43]. AFlow and AgentSquare search over workflow structures and modular agent components, respectively [44, 45]. Related research on recursive self-improvement has theoretical roots in the Gödel Machine [46]. Recent systems explore executable forms of self-modification: the Darwin Gödel Machine iteratively modifies and evaluates agent code [7], while Hyperagents also make the meta-level improvement procedure itself editable [8]. Our approach draws on this line of research to search for reusable efficiency improvements in harness mechanisms while keeping the underlying model fixed. 4.2. Automated Harness Optimization Recent methods optimize harness code and configurations using execution feedback. AutoHarness synthesizes environment-specific code guards and, in some settings, complete code policies [47]. Recursive Harness SelfImprovement (RHI) refines prompt-level specifications of the agent loop for individual tasks [19]. Meta-Harness searches over executable harness programs using prior code, scores, and execution traces, maintaining a Pareto frontier over task performance and context cost [17]. AHE evolves modular coding-harness components from trajectory evidence and evaluates their transfer across tasks and models [18]. MemoHarness stores execution experience and uses it to adapt harness configurations to individual test cases at inference time [48]. The extent to which the improvements discovered generalize remains an important evaluation question. Wang et al. report limited gains on held-out tasks for the methods they evaluate and highlight the risk of overstating improvements when search and evaluation tasks overlap [20]. Our search pipeline explores candidates across diverse executable development environments, using a broad-to-deep funnel to refine promising changes while keeping final held-out evaluation separate from candidate development. The search targets token efficiency while checking that candidates maintain task performance. SoL-Pi combines four mechanisms selected through this process. Its held-out evaluation provides preliminary evidence that RSI-inspired search at the harness layer can discover mechanisms that generalize beyond their development environments. 4.3. Context and Token-Efficient Agents Long-horizon agents can incur substantial token overhead as interaction histories and environment observations accumulate. LLM-as-Code limits context accumulation by moving deterministic control flow into executable code [49]. Other methods reduce the information retained during execution: AgentDiet removes redundant and outdated trajectory content [50], while ACON optimizes the compression of observations and interaction histories [51]. Context-Folding and AgentFold enable agents to manage their working context through trajectory folding and compression [10, 52], and Context as a Tool exposes context maintenance as an explicit agent action [53]. Agentic Context Engineering (ACE) focuses on accumulating and refining reusable strategies in context through execution feedback [54]. These approaches provide mechanisms for controlling context growth and organizing execution. Our approach studies automated discovery and selection across multiple harness components, seeking reusable changes that reduce token usage while maintaining

11

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

task performance. SoL-Pi combines the selected mechanisms across action execution, context compaction, observation handling, and delegated reading.

5. Conclusion We introduce SoL-Pi, an RSI-inspired auto-research system for discovering and combining efficient agent-harness mechanisms. On EdgeBench, the best-performing candidates improve model performance by 5.3–12.8% and token efficiency by 9.8–18.2%, while the complete stack reduces token traffic by 44.7–49.0% and token cost by about one third at comparable performance. The complete stack’s consistent results across GPT-5.6 Sol and Opus 5 demonstrate strong cross-model generalization within the evaluated setting, positioning SoL-Pi as a preliminary step toward scalable RSI systems. 5.1. Limitations and Future Directions Pre-Training the Harness. Generalizing harness artifacts discovered through RSI remains a persistent challenge. Our results provide preliminary evidence that scaling auto-research loops can mitigate this challenge. Analogous to pretraining, the harness is exposed to many tasks and updated from the resulting trajectories. We hypothesize that scaling both executable environments and the diversity of research ideas can yield sustained gains; we call this long-term research direction pretraining the harness. We will continue to explore this direction. Multi-Backend Training. Treating the harness as trainable suggests a complementary path: learning from multiple LLM backends. Our current harness was updated from trajectories generated by a single LLM; on the other evaluated backend, its mechanisms trigger less often and less intensively, although they still improve token efficiency when triggered. Training and validating the harness across multiple backends may improve robustness while preserving task quality and efficiency. Recursive Efficient Improvement. Efficiency may also become recursive: a more efficient harness could lower the cost of the auto-research used to build its successor. We plan to use SoL-Pi as the starting harness for the next research cycle, where lower per-run costs could let a fixed budget cover more executable environments, trajectories, and research ideas. In this view, efficiency is both an outcome of harness research and a resource for expanding the search that follows, so a more efficient harness may help discover an even more efficient one. We call this possibility recursive efficient improvement; it is a long-term research vision rather than a compounding effect demonstrated by the present study. Search Coverage and Cost. Running complete auto-research loops in our environment is computationally expensive, making controlled comparisons of search breadth and depth under a fixed budget particularly challenging. We view scaling laws along these dimensions as a promising research direction and leave their systematic investigation to future work.

References [1] Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021. [2] John Yang, Carlos Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems, 37: 50528–50652, 2024. [3] Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, et al. Measuring ai ability to complete long tasks. arXiv preprint arXiv:2503.14499, 352, 2025. [4] Wilson Lin. Towards self-driving codebases. Cursor research blog, 2026. self-driving-codebases. Accessed 2026-08-17.

URL https://cursor.com/blog/

12

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

[5] OpenAI. Introducing GPT-5.3-Codex. OpenAI technical report, 2026. introducing-gpt-5-3-codex/. Accessed 2026-08-17.

URL https://openai.com/index/

[6] Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search. arXiv preprint arXiv:2504.08066, 2025. [7] Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange, and Jeff Clune. Darwin gödel machine: open-ended evolution of self-improving agents. In International Conference on Learning Representations, volume 2026, pages 104223–104294, 2026. [8] Jenny Zhang, Bingchen Zhao, Wannan Yang, Jakob Foerster, Jeff Clune, Minqi Jiang, Sam Devlin, and Tatiana Shavrina. Hyperagents. arXiv preprint arXiv:2603.19461, 2026. [9] Anthropic. Effective harnesses for long-running agents. Anthropic Engineering, 2025. URL https://www.anthropic. com/engineering/effective-harnesses-for-long-running-agents. Accessed 2026-08-29. [10] Rui Ye, Zhongwang Zhang, Kuan Li, Huifeng Yin, Zhengwei Tao, Yida Zhao, Liangcai Su, Liwen Zhang, Zile Qiao, Xinyu Wang, et al. Agentfold: Long-horizon web agents with proactive context management. arXiv preprint arXiv:2510.24699, 2025. [11] Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning. In International Conference on Learning Representations, volume 2024, pages 35549–35562, 2024. [12] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles, pages 611–626, 2023. [13] Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: Accurate and efficient post-training quantization for large language models. In International conference on machine learning, pages 38087–38099. PMLR, 2023. [14] Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323, 2022. [15] Lingjiao Chen, Matei Zaharia, and James Zou. FrugalGPT: How to use large language models while reducing cost and improving performance. arXiv preprint arXiv:2305.05176, 2023. URL https://arxiv.org/abs/2305.05176. [16] Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M. Waleed Kadous, and Ion Stoica. RouteLLM: Learning to route LLMs with preference data. arXiv preprint arXiv:2406.18665, 2024. URL https://arxiv.org/abs/2406.18665. [17] Yoonho Lee, Roshen Nair, Qizheng Zhang, Kangwook Lee, Omar Khattab, and Chelsea Finn. Meta-harness: End-to-end optimization of model harnesses. arXiv preprint arXiv:2603.28052, 2026. [18] Jiahang Lin, Shichun Liu, Chengjun Pan, Lizhi Lin, Shihan Dou, Zhiheng Xi, Xuanjing Huang, Hang Yan, Zhenhua Han, Tao Gui, et al. Agentic harness engineering: Observability-driven automatic evolution of coding-agent harnesses. arXiv preprint arXiv:2604.25850, 2026. [19] Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Matei Zaharia, and Yujin Tang. Recursive harness self-improvement. arXiv preprint arXiv:2607.15524, 2026. [20] Yike Wang, Huaisheng Zhu, Zhengyu Hu, Yige Yuan, Zhengyu Chen, Shakti Senthil, Hannaneh Hajishirzi, Yulia Tsvetkov, Pradeep Dasigi, and Teng Xiao. Rethinking the evaluation of harness evolution for agents. arXiv preprint arXiv:2607.12227, 2026. [21] Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overoptimization. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 10835–10866. PMLR, 2023. URL https://proceedings.mlr.press/v202/gao23h.html. [22] ByteDance Seed. EdgeBench: Scaling laws of environment learning, 2026. URL https://edge-bench.org/. Accessed 2026-08-17. [23] Andrej Karpathy. autoresearch: The experiment loop. GitHub repository, 2026. URL https://github.com/karpathy/ autoresearch/blob/master/program.md. Accessed 2026-08-17. [24] Anthropic. Ralph wiggum plugin: Ralph loop command. GitHub repository, 2026. URL https://github.com/ anthropics/claude-code/blob/main/plugins/ralph-wiggum/commands/ralph-loop.md. Accessed 2026-08-17.

13

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

[25] Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini, et al. Terminal-Bench: Benchmarking agents on hard, realistic tasks in command line interfaces. arXiv preprint arXiv:2601.11868, 2026. URL https://arxiv.org/abs/2601.11868. [26] Ryan Marten. Terminal-Bench 4.0. https://www.tbench.ai/news/terminal-bench-4-0, 2026. Accessed September 14, 2026. [27] International Mathematical Olympiad. IMO 2026 Problems. https://www.imo-official.org/problems/2026/, 2026. Accessed September 14, 2026. [28] Anthropic. Anthropic’s original performance take-home. performance_takehome. Accessed September 12, 2026.

https://github.com/anthropics/original_

[29] ByteDance Seed. EdgeBench official leaderboard. Official GitHub repository, 2026. URL https://github.com/ ByteDance-Seed/EdgeBench. Official @2h score, commit a87350a; accessed 2026-08-25. [30] OpenAI. Codex CLI. GitHub repository, 2026. URL https://github.com/openai/codex. accessed September 15, 2026. [31] TokenRhythm. OpenSquilla. GitHub repository, 2026. URL https://github.com/TokenRhythm/opensquilla. accessed September 15, 2026. [32] can1357. Oh My Pi. GitHub repository, 2026. URL https://github.com/can1357/oh-my-pi. accessed September 15, 2026. [33] Anomaly. OpenCode. GitHub repository, 2026. URL https://github.com/anomalyco/opencode. accessed September 15, 2026. [34] code-yeongyu. Oh My OpenCode. GitHub repository, 2026. URL https://github.com/code-yeongyu/ oh-my-openagent. Repository now named oh-my-openagent; accessed September 15, 2026. [35] Earendil Works. Pi: Coding Agent Toolkit. GitHub repository, 2026. URL https://github.com/earendil-works/ pi. Formerly pi-mono; accessed September 15, 2026. [36] Anthropic. Claude Code. GitHub repository, 2026. URL https://github.com/anthropics/claude-code. accessed September 15, 2026. [37] Humanfia. Humanize Olympic Agents (HOA). https://humanfia.ai/projects/hoa, 2026. Accessed September 17, 2026. [38] Axiom Math. AxiomProver at IMO 2026. https://github.com/AxiomMath/IMO2026, 2026. Accessed September 17, 2026. [39] Anthropic. Prompt Caching. Claude API documentation, 2026. URL https://platform.claude.com/docs/en/ build-with-claude/prompt-caching. Accessed September 17, 2026. [40] Xuying Ning, Katherine Tieu, Dongqi Fu, Tianxin Wei, Zihao Li, Yuanchen Bei, Jiaru Zou, Mengting Ai, Zhining Liu, Ting-Wei Li, et al. Code as agent harness. arXiv preprint arXiv:2605.18747, 2026. [41] Elias Lumer, Sahil Sen, Kevin Paul, and Vamse Kumar Subbiah. Recursive agent harnesses. arXiv preprint arXiv:2606.13643, 2026. [42] Lakshya A Agrawal, Shangyin Tan, Dilara Soylu, Noah Ziems, Rishi Khare, Krista Opsahl-Ong, Arnav Singhvi, Herumb Shandilya, Michael J Ryan, Meng Jiang, et al. Gepa: Reflective prompt evolution can outperform reinforcement learning. In International Conference on Learning Representations, volume 2026, pages 8479–8565, 2026. [43] Shengran Hu, Cong Lu, and Jeff Clune. Automated design of agentic systems. In International Conference on Learning Representations, volume 2025, pages 21344–21377, 2025. [44] Jiayi Zhang, Jinyu Xiang, Zhaoyang Yu, Fengwei Teng, Xionghui Chen, Jiaqi Chen, Mingchen Zhuge, Xin Cheng, Sirui Hong, Jinlin Wang, et al. Aflow: Automating agentic workflow generation. In International Conference on Learning Representations, volume 2025, pages 34040–34077, 2025. [45] Yu Shang, Yu Li, Keyu Zhao, Likai Ma, Jiahe Liu, Fengli Xu, and Yong Li. Agentsquare: Automatic llm agent search in modular design space. In International Conference on Learning Representations, volume 2025, pages 3841–3865, 2025. [46] Jürgen Schmidhuber. Gödel machines: Fully self-referential optimal universal self-improvers. In Artificial general intelligence, pages 199–226. Springer, 2007.

14

SoL-Pi : Recursively Scaling Auto-Research Loops for Efficient Agent Harness

[47] Xinghua Lou, Miguel Lázaro-Gredilla, Antoine Dedieu, Carter Wendelken, Wolfgang Lehrach, and Kevin P Murphy. Autoharness: improving llm agents by automatically synthesizing a code harness. arXiv preprint arXiv:2603.03329, 2026. [48] Yue Huang, Wenjie Wang, Han Bao, Yuchen Ma, Xiaonan Luo, Yi Nian, Haomin Zhuang, Zheyuan Liu, Yue Zhao, and Xiangliang Zhang. Memoharness: Agent harnesses that learn from experience. arXiv preprint arXiv:2607.14159, 2026. [49] Junjia Qi, Zichuan Fu, Jingtong Gao, Wenlin Zhang, Hanyu Yan, Xian Wu, and Xiangyu Zhao. Llm-as-code agentic programming for agent harness. arXiv preprint arXiv:2606.15874, 2026. [50] Yuan-An Xiao, Pengfei Gao, Chao Peng, and Yingfei Xiong. Reducing cost of llm agents with trajectory reduction. Proceedings of the ACM on Software Engineering, 3(FSE):1241–1263, 2026. [51] Minki Kang, Wei-Ning Chen, Dongge Han, Huseyin A Inan, Lukas Wutschitz, Yanzhi Chen, Robert Sim, and Saravan Rajmohan. Acon: Optimizing context compression for long-horizon llm agents. arXiv preprint arXiv:2510.00615, 2025. [52] Weiwei Sun, Miao Lu, Zhan Ling, Kang Liu, Xuesong Yao, Yiming Yang, and Jiecao Chen. Scaling long-horizon llm agent via context-folding. arXiv preprint arXiv:2510.11967, 2025. [53] Shukai Liu, Bo Jiang, Jian Yang, Yizhi Li, Jinyang Guo, Xianglong Liu, and Bryan Dai. Context as a tool: Context management for long-horizon swe-agents. In Findings of the Association for Computational Linguistics: ACL 2026, pages 20604–20617, 2026. [54] Qizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma, Fenglu Hong, Vamsidhar Kamanuru, Jay Rainton, Chen Wu, Mengmeng Ji, Hanchen Li, et al. Agentic context engineering: Evolving contexts for self-improving language models. In International Conference on Learning Representations, volume 2026, pages 86069–86100, 2026.

15

Record · ID 978472 · SHA-256 7f3910c43a4927c7
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.