IEEE TRANSACTIONS ON SERVICES COMPUTING
1
AgentLoop: Runtime Control of Slot-closed Execution Loops for Tool-augmented LLM Agents
arXiv:2609.33315v1 [cs.DC] 27 Sep 2026
Wanyi Zheng, Minxian Xu, Senior Member, IEEE, Kan Hu, Kejiang Ye, Senior Member, IEEE, and Chengzhong Xu, Fellow, IEEE
Abstract—Tool-augmented large language model (LLM) agents are becoming an important execution unit in service computing, but existing agent loops still lack explicit runtime signals for assessing task completion. The challenge lies in the fact that an agent may continue reasoning or invoking services even after the runtime context has stopped changing, while evidence already collected remains unsynthesized into a complete answer, which leads to inefficiency in resource usage. To address these challenges, this paper presents AgentLoop, which provides runtime control of slot-closed execution loops for tool-augmented agents. Slot closure means that the information slots required by a request have been covered by sufficient runtime evidence, and that unresolved slots are explicitly identified before the loop stops. AgentLoop converts open-ended agent iteration into state-driven execution control: it maintains a compact runtime state, uses model-assisted structured verification to check answer completeness and missing evidence, and applies bounded stability and low-gain signals over neighboring LLM/tool rounds before selecting one of three actions: Continue Invocation, Answer Synthesis, or Terminate Iteration. AgentLoop does not change the underlying model or tool set, and instead adds a fine-grained stopping criterion at execution time. Experiments show that AgentLoop reduces redundant execution and context growth, with total token cost reduced by up to 88.44% and average service invocations reduced by up to 76.85% against baselines. The ablation study further shows that the slot-centered control path plays a central role, since disabling it increases execution depth and substantially reduces accuracy. Overall, the results suggest that efficient tool-augmented agents can benefit from explicit runtime signals for deciding when further LLM/tool iterations no longer add useful context or supported evidence. Index Terms—Large language model agents, tool-augmented agents, agent orchestration, runtime control, agent loop control.
I. I NTRODUCTION
L
LM agents extend generation with intermediate reasoning, planning, and external actions. Reasoning methods
This work is supported by National Key R&D Program of China (No.2026YFE0199800), National Natural Science Foundation of China under Grant 62572462,Guangdong Science and Technology Cooperation Project under Grant 2025A0505020065, Key Research and Development and Technology Transfer Program of Inner Mongolia Autonomous Region (2025YFHH0110) and Shenzhen Science and Technology Program under Grant ZDYJ20251211121533004. W. Zheng is with Southern University of Science and Technology, Shenzhen University of Advanced Technology, and Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences (Email: [email protected]). M. Xu (corresponding author), K. Hu and K. Ye are with Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences (Email: [email protected], [email protected], [email protected]). C. Xu is with Institute of AI and Brain Sciences, Department of Computer Science, University of Macau, Macau, China (Email: [email protected]).
Request
Agent Controller Shared Runtime State
❌
Write traces
LLM Inference Engine
Tool Execution Layer
Observation Module
Execution Layer
Reflection Gate
✅ Response
Fig. 1.
General execution structure of LLM agents.
such as Chain-of-thought, Self-consistency, Tree of Thoughts, and Plan-and-solve organize internal problem solving [1]– [4]. Tool-augmented methods such as ReAct, MRKL, and Toolformer connect these states to external services [5]–[7]. Together, they establish the basic execution pattern of an agent, in which the system interprets a request, invokes services, observes results, and produces an answer. Fig. 1 illustrates this execution structure. The agent controller does more than issue a single model call. It repeatedly coordinates model inference, tool execution, observation, state writing, and reflection. This structure makes LLM agents flexible, but it also creates a control problem. The system must decide when the shared runtime state contains enough evidence to answer and when another reasoning-action round is still useful. In service computing, this pattern corresponds to service discovery, service invocation, and result integration. Application programming interface (API)-oriented systems such as Gorilla, ToolLLM, ToolBench, API-Bank, TaskMatrix.AI, and RestGPT improve tool selection, parameter generation, and service composition [8]–[11]. Agent platforms and workflow systems further support open-ended tasks, multi-agent collaboration, and execution optimization [12]–[14]. Related cloudnative and distributed serving work provides the deployment foundation for scalable LLM systems [15], [16]. These advances make tool use increasingly capable, but they also make runtime traces longer and service evidence more diverse. From the perspective of service computing, LLM agents can be viewed as adaptive service orchestration systems, where foundation models act as reasoning services and external tools represent dynamically invoked computational services.
IEEE TRANSACTIONS ON SERVICES COMPUTING
Therefore, efficient agent execution requires not only service selection and composition, but also runtime control mechanisms that determine whether additional service invocations contribute meaningful information [17]. This reinforces that agent reliability depends not only on selecting useful actions, but also on controlling how context accumulates across iterations [18]. The remaining systems problem is task-completion awareness. Existing agents commonly stop through a fixed budget, a predefined workflow boundary, or model self-judgment. These mechanisms do not reliably distinguish necessary exploration from repeated LLM reasoning rounds or repeated service calls, and they do not always separate missing evidence from evidence that has already been obtained but not synthesized. Recent evidence on agentic coding shows the scale of this cost. Agentic coding tasks have been shown to consume about 4.17M tokens per task on average, compared with 3.39k tokens for code chat and 1.19k tokens for single-turn code reasoning [19]. Repeated runs on the same task can vary substantially in total tokens, and higher token use does not necessarily improve accuracy [19]. Consequently, an agent may continue after the useful context or tool evidence has stabilized, or return tool fragments instead of a complete service result. This issue affects both service quality and resource use. Each unnecessary round adds model inference, possible tool access, context assembly, and queueing cost, while repeated reasoning or repeated observations may contribute little new information. A practical controller should therefore make two related judgments: whether another LLM-driven reasoningaction round is likely to update the runtime context, and whether another tool invocation is likely to close an information gap or resolve a conflict. To address the aforementioned challenges, this paper proposes AgentLoop. Unlike existing stopping strategies that rely on predefined budgets or model-generated confidence, AgentLoop performs requirement-aware runtime control by jointly reasoning about information coverage, evidence availability, and marginal utility of future service interactions. AgentLoop uses structured runtime verification to check answer completeness, missing evidence, tool need, confidence, and failure mode, and then applies bounded stability and low-gain1 signals over neighboring LLM/tool rounds. Based on these signals, AgentLoop suppresses redundant states and selects among Continue Invocation, Answer Synthesis, and Terminate Iteration. AgentLoop is an execution-time layer and does not require changing the base model or tool interfaces. The key contributions are as follows: We formulate runtime completion control for toolaugmented LLM agents as a joint decision problem considering requirement coverage, evidence availability, context stability, and execution budget. • We design AgentLoop, a lightweight runtime control architecture consisting of state management, slot verification, redundancy suppression, and iteration control.
2
•
We conduct extensive experiments across multiple agent benchmarks and LLM backbones, demonstrating substantial reductions in execution cost while maintaining answer quality. II. R ELATED W ORK
Existing work has improved reasoning, tool use, and orchestration, but task-completion awareness is still rarely treated as a runtime control problem. We group related studies by their main control target and compare them with AgentLoop from the perspective of stopping, context and evidence modeling, and redundant-round control. A. Reasoning Enhancement and Reflective Generation Chain-of-thought, Self-consistency, Tree of Thoughts, Planand-solve, and Plan-and-act improve intermediate reasoning, search, and action organization [1]–[4], [23]. Reflexion, Selfrefine, and SE-Agent further introduce critique and trajectory revision [20], [21], [24]. These methods help agents reason or revise answers, but their completion signals are still mainly generated by the model itself. They rarely maintain an explicit runtime state that can determine whether another reasoning round changes useful context, whether a missing fact requires another service call, or whether existing evidence only needs to be synthesized. In addition, the stopping rule in these methods is usually either a fixed iteration budget or the model’s own judgment that the answer is already good enough, and neither of them separates the case where the context has stopped changing from the case where the task still has an unfilled requirement. B. Tool-augmented Language Models ReAct [5], MRKL [6], Toolformer [7], Gorilla [8], AvaTaR [25], ToolLLM [9], ToolBench, API-Bank [10], ART [26], Chameleon [27], and LLMCompiler [14] improve tool selection, API grounding, composition, or parallel execution. Tool-ecosystem studies further optimize tool construction and instructions [28], [29]. Their main objective is to make tool access more accurate and scalable. However, successful tool access is not equivalent to task completion: these systems usually store tool returns or traces, but do not explicitly check whether another LLM/tool iteration would update useful context, add new supported evidence, or repeat an already stable state. In practice this appears as continued invocation on top of an already stable tool result, where later calls return the same content and add only latency and context length.
•
1 In this paper, the terms low-gain and low-value are used interchangeably to describe unnecessary or redundant service invocations that contribute limited additional evidence or context to the runtime state.
C. Agent Evaluation and Workflow Orchestration AgentBench [30] and GAIA [31] expose agent capability and grounded-task failure modes. Conversational Information Gain (CIG) measures how utterances advance deliberative dialogues by updating a semantic memory state, providing a related but evaluation-oriented view of information gain [32]. AutoGen, AFlow, ADAS, GPTSwarm, Agent Workflow Memory, evolving orchestration, and Agentix improve conversation, workflow construction, graph search, memory, or serving [13],
IEEE TRANSACTIONS ON SERVICES COMPUTING
3
TABLE I C APABILITY- BASED COMPARISON OF AGENT L OOP WITH REPRESENTATIVE REASONING , TOOL - USE , AND WORKFLOW METHODS .
Method
CoT [1], Self-consistency [2], ToT [3] Reflexion [20], Self-refine [21] ReAct [5], MRKL [6], Toolformer [7] Gorilla [8], ToolLLM [9], API-Bank [10] AutoGen [13] LangGraph [22] AgentLoop(ours)
Task-level State
Context and Evidence
Execution Control
Req. Slots
Answer Ready
Context Tool Gap LLM/Tool Redundancy Budget State Evidence Attribution Utility Control Guard
−− −− −− −− −− −− ✓
△ ✓ △ −− △ △ ✓
✓ ✓ ✓ −− ✓ ✓ ✓
−− −− ✓ ✓ △ △ ✓
−− −− −− −− −− −− ✓
−− △ △ −− △ △ ✓
−− −− −− −− ✓ ✓ ✓
−− −− −− ✓ ✓ −− ✓
Note: ✓means explicit design target, △means partial or indirect support, and −−means not a primary part. Req. Slots: the request is decomposed into information items that the final answer must cover. Answer Ready: whether the current draft already organizes the supported items into a deliverable answer. Context State: whether runtime context is kept as an explicit state rather than an unstructured trace. Tool Evidence: whether tool returns are verified as support for the required items, not only stored and replayed. Gap Attribution: whether an unmet requirement is attributed to missing evidence, to evidence that has been obtained but not yet synthesized, or to an already stable state. LLM/Tool Utility: whether the marginal value of one more LLM or tool round is estimated before that round is issued. Redundancy Control: whether rounds that reproduce an already stable state are suppressed. Budget Guard: whether execution depth, latency, and failure state are guarded as a runtime budget.
[33]–[39]. These systems constrain or optimize execution at the platform level. Their stopping conditions are often tied to nodes, branches, workflow boundaries, or budgets, while an open tool node may still lack a local rule for distinguishing unchanged LLM context, missing evidence, and evidence that has been collected but not yet synthesized. Table I summarizes this distinction as a capability matrix. The table is meant to show that existing methods usually cover only part of the control stack, while AgentLoop combines request-level state representation with evidence-aware control of both LLM and tool iterations. For instance, AutoGen and LangGraph do not decompose requests into items that the final answer to cover (request slot), and do not consider the missing evidence can lead to unmet requirement (gap attribution). The distinction is therefore not that AgentLoop replaces reasoning, tool-use, or workflow frameworks. It adds a runtime control layer that can be embedded into these systems, closes request-level information slots, distinguishes missing evidence from unsynthesized evidence, and suppresses low-gain LLM or tool iterations before they expand the execution trace. III. P ROBLEM M OTIVATION We summarize the need for runtime completion control through three observations. We use low-value execution risk here as a request-level operational indicator for repetitive or low-yield execution: the agent has already invoked tools or produced repeated LLM outputs, but the final state still fails to add useful supported information. Operationally, this indicator covers post-tool failure and repeated model outputs in the recorded trace, and it functions as a direct control signal for repetitive runtime behavior rather than a learned utility estimate. This motivation study uses the same unified GPU-node and runtime configuration as the main experiments, with cache disabled and the same request order within each backbone. a) Observation 1: Tool invocation by itself does not guarantee that request-level evidence is complete. Fig. 2(a) shows that the low-value execution pattern is clearly visible in both workflow-based and conversation-based baselines. The overall risk reaches 17.1% for LangGraph and 37.6% for AutoGen, with post-tool failure instances making up most
of the count. Fig. 2(b) further shows that the risk rises on ToolBench and Open-Agent-Trace, where service composition and evidence integration are more demanding. The central issue is therefore not only whether an agent can invoke a service, but whether it can recognize, verify, and synthesize the returned evidence. b) Observation 2: Low-gain LLM behavior is a persistent runtime pattern rather than an isolated failure case. Fig. 2(c) shows that low-gain LLM calls remain close to total LLM calls across the Open-Agent-Trace request sequence. The low-gain share reaches 49.82% under request-level aggregation and 53.08% when three-request windows are averaged. Despite the different averaging orders, both measurements indicate that low-gain execution is a recurring runtime state. This motivates online monitoring at the agent-iteration level, so AgentLoop should detect both LLM rounds that no longer update useful context and tool invocations that no longer add supported evidence. c) Observation 3: Similar LLM-call depth can hide very different ratios of productive and low-gain execution. Fig. 3 shows that LangGraph and AutoGen have similar average LLM-call depth on ToolBench and Open-Agent-Trace, yet their low-gain components differ substantially. On OpenAgent-Trace, AutoGen has a 49.82% aggregate low-gain share compared with 16.38% for LangGraph. Thus, LLM-call count alone is insufficient for evaluating execution efficiency, and the controller must also judge whether each reasoning-action round changes the useful runtime context and whether associated tool calls add supported evidence.
IV. AGENT L OOP D ESIGN Based on the aforementioned observations, to address incomplete evidence integration and repeated low-gain execution, we design AgentLoop as a runtime slot-closed control layer for tool-augmented LLM agents. It represents request requirements as slots and uses runtime context and tool evidence to decide whether to continue invocation, synthesize an answer, or terminate the iteration.
IEEE TRANSACTIONS ON SERVICES COMPUTING
4 AutoGen / Open-Agent-Trace
Tool Use Not Converted
Redundant Output Only
15.6%
LangGraph
LangGraph
Both
41.2% 6.2%
GAIA 34.0%
AutoGen 0.0%
10.0%
37.6%
20.0% 30.0% Rate over All Requests
9.3% 30.8%
Open-AgentTrace
40.0%
91.9%
0%
20%
(a) Overall risk.
Fig. 2.
AutoGen
28.5%
ToolBench
17.1%
Mean LLM Calls / 3 Requests
6
40%
60%
80%
4 3 2 1 0
100%
Total Calls Low-Gain Calls
5
0
100
(b) Dataset risk.
200 300 Request Index
400
(c) LLM-call trend.
Motivation overview.
A. System Overview AgentLoop adds a runtime control layer around an existing tool-augmented agent. Given the request, available tools, current answer draft, and execution trace, it chooses Continue Invocation, Answer Synthesis, or Terminate Iteration, while leaving the base model and external tools unchanged.
Agent Execution Trace
Runtime State Manager
Control Signals
State
Slot State Verifier
Verification Signals
Iteration Controller
Low-Gain Calls 3.0
Low-Gain Avg LLM Calls / Request Share (%) Low-Gain Avg LLM Calls / Request Share (%) Low-Gain Avg LLM Calls / Request Share (%)
3.0
Read
Shared Runtime State
2.20
2.20
2.20
0.5 1.0 0.5 0.31 0.0 0.5 0.0
3.0
Feedback
2.30 2.5 3.02.5
2.30
2.20
2.30
2.0 2.52.0
2.20
Fig. 4.
Method schematic of AgentLoop.
1.0 1.51.0
0.31
0.55
0.55
0.5 1.00.5 0.38 0.0 0.50.0
1.10 0.38
1.10
1.10
0.55 0.38 0.31 50% 50% 50 0.0 50 50 0.0 50 25% 25% 16% 16% 15% 15% 50% 50 50 0 0 0 0 25% LangGraph LangGraph AutoGen AutoGen LangGraph LangGraph AutoGenAutoGen 16% 15% 0
State Update and Feedback
2.20
1.5 2.01.5
1.0 1.5 1.0
Execution Results
Other Calls 3.0
2.5 3.0 2.5 2.10 2.10 2.0 2.5 2.0 2.10 1.5 2.0 1.5
Execution Decision
Read / Update Write
Low-Gain Low-Gain Calls Calls Other Calls Other Calls
Fig. 3.
500
(a)LangGraph ToolBench.
AutoGen
0
(b) Open-Agent-Trace. LangGraph AutoGen
LLM-call composition on ToolBench and Open-Agent-Trace.
Fig. 4 summarizes the state-mediated data flow: the execution trace is compressed into shared runtime state, checked by the Slot State Verifier, and consumed by the Iteration Controller to produce the next execution decision. Given only the request, the current answer draft, and the compressed evidence summary, the verifier performs three checks: 1) whether the answer draft already covers the information that the request requires; 2) whether the covered requirements are actually supported by the evidence obtained so far, rather than asserted by the model; and 3) if a requirement is still unmet, whether the gap is attributable to evidence that has not been retrieved, to evidence that has been retrieved but not yet organized into the answer, or to a tool state that is already stable or self-contradictory. The resulting execution feedback updates the shared state before the next agent round, and the detailed component responsibilities are shown in Fig. 5. The control principle is that AgentLoop separates three cases that are often conflated: 1) missing evidence may require another tool invocation; 2) stable evidence with an incomplete response requires Answer Synthesis; and 3) stable context with sufficient, supported evidence permits Terminate Iteration. Redundancy Suppression is therefore a decision signal over both LLM rounds and tool calls rather than a rule that forbids repeated use of a tool.
Fig. 5 shows the closed-loop architecture and its three main components. The Runtime State Manager extracts slot state from historical requests and tool traces, writes a compact Runtime State Object into the Slot State Store, and exposes the resulting Slot State Context to the online loop. In the implementation, this runtime state is a compact summary rather than a full trace dump: it keeps the latest request context, the latest tool evidence, and the control flags needed by the next round. Using the request, available tools, and slot context, the Slot State Verifier checks Slot Requirements, Evidence Verification, and Gap Attribution. The Iteration Controller reads the execution trace and runtime state, applies Redundancy Suppression, Budget Guarding, and a low-gain stability heuristic, and selects the next execution decision based on LLM-context changes and new tool evidence. V. S YSTEM M ODELING AND A LGORITHM D ESIGN In this section, we introduce the system modeling and algorithm design in AgentLoop. A. Online Runtime Mechanism After runtime stage t, AgentLoop represents the state observed by its controller as Xt = (q, Yt , Call t , Ēt , Vt , Ht , Bt ).
(1)
Here, q is the user request, Yt is the current answer draft, and Call t is the set of tool calls issued in the current or most recent round. Ēt is the compressed summary of tool states and returned evidence. Operationally, it keeps only a bounded textual representation of each new tool result, such as the tool name, error flag, key finding, and unresolved fields,
IEEE TRANSACTIONS ON SERVICES COMPUTING
5
① Runtime State Manager Slot State Extraction
Historical Requests AgentLoop
Runtime State Object {
}
"slot_state": ["Entity", "Attribute", "Evidence"], "tool_state": ["Search", "Visit", "Calc"], "stop_policy": "slots_covered", "confidence": 0.91
Slot State Store
Tool Traces
Slot State Context
Decision Signal
Slot State Injection Slot Verification Unit New Request Tool AgentLoop Execution Available Tools
Slot Requirements
Execution Trace
Redundancy Suppression
Evidence Verification
Budget Guarding
Runtime State
Gap Attribution
② Slot State Verifier
Fig. 5.
Iteration Controller
Information Gain Estimation
Continue Invocation Answer Synthesis Terminate Iteration
③ Iteration Controller
Overview of the AgentLoop runtime control architecture.
rather than the full raw trace. Vt is the structured verifier output, Ht is the history of dialogue messages, tool calls, tool results, and control events, and Bt contains runtime boundaries and configuration such as iteration limits, reflection limits, timeout constraints, similarity thresholds, and feature switches. Equation (1) is a state abstraction rather than a separately instantiated Python class. It corresponds to the implementation fields that maintain model context, current model results, previous calls, previous execution results, and the tool-state cache. The verifier reads the request, current answer, and compressed evidence state and returns Vt = Verify(q, Yt , Ēt ) = (at , Mt , nt , κt , ft ).
(2)
The binary variable at is the answer_complete flag. Mt is the list of missing evidence or unresolved requirements, and nt is the need_tool flag. The confidence value κt ∈ [0, 1] is the verifier confidence. The failure state ft records conditions such as ok, answer_incomplete, tool_set_unstable, tool_conflict, and low_confidence. Together, Mt , nt , and ft implement Gap Attribution: an unmet requirement is attributed to missing evidence when nt = 1, to evidence that has already been obtained but not yet synthesized when nt = 0 and ft = answer_incomplete, and to an already stable or contradictory tool state when ft is tool_set_unstable or tool_conflict. In other words, the slot state is a fixedschema verifier record generated from the current request, the current answer, and the compressed evidence summary, rather than a rule-based slot extractor. Historical requests and tool traces are retained in Ht and Ēt as execution context, but the verifier itself consumes only the current round state. When evidence is missing or contradictory, the failure labels guide the controller toward another tool round or answer integration. The verifier is a structured auxiliary-model output used by the controller. It is not an independent module named SlotGen. The online closure test is therefore binary:
The indicator I[·] takes value one when its condition is true and zero otherwise. Thus, Ct = 1 means that the verifier regards the answer as complete and does not request another tool round. It is a direct closure signal, not just a weighted coverage score. To compare adjacent tool-call plans, AgentLoop uses the mixed text similarity and the resulting call-state similarity in one definition: ρ(x, y) = 0.55 Lex(x, y) + 0.45 Jac(Token(x), Token(y)), 1, Call t−1 = Call t = ∅, 0, NM t = 0 ∨ XOR t = 1, Stcall = m X 1 ρ(arg t−1,j , arg t,j ), otherwise. m j=1
(4) Here, Lex is normalized lexical similarity, Token returns the token set of a text, and Jac is Jaccard similarity. NM t is one when the tool-name multisets match, and XOR t is one when exactly one call set is empty. The parameter text arg t,j is the normalized argument of the jth aligned call and m is the number of aligned calls. The implementation groups calls by tool name and sorts their normalized JSON arguments before comparison. The mixed score is a lexical and token-set comparison, not an embedding similarity. The corresponding result-state similarity is where Result t is the set of tool results in round t, res t,j is an aligned normalized result text, and k is the number of aligned results. Mismatch t is one when either result set is empty, tool grouping or multiplicity differs, or an error state differs. The predicate is error records whether a result represents a tool error. The comparison checks these fields before comparing compressed result text.
Stresult =
1, 0,
Result t−1 = Result t = ∅, Mismatch t = 1,
k 1X ρ(res t−1,j , res t,j ), otherwise. k j=1
Ct = I[at = 1 ∧ nt = 0].
(3)
(5)
IEEE TRANSACTIONS ON SERVICES COMPUTING
6
Let NameMatcht indicate equality of tool-name multisets. Let ExactCallt and ExactResultt indicate equality of the normalized call and result signatures. The stability gate is Gstab = I NM t = 1 t ∧ (ExactCallt = 1 ∨ EnableCallSim = 1 ∧ Stcall ≥ τcall ) ∧ (ExactResultt = 1 ∨ EnableResultSim = 1 ∧ Stresult ≥ τresult ) .
(6) Here, NM t is the name-match variable defined above. ExactCallt and ExactResultt indicate equality of normalized call and result signatures. EnableCallSim and EnableResultSim are configuration switches. The default thresholds are fixed runtime settings rather than theoretical constants. The gate means that the neighboring action and result states are stable enough that further replanning has low expected value. The implementation uses discrete branches rather than a continuous utility optimizer. A unified abstraction of these branches is need_tool_round, verifier_retry, rollback_retry, rollback_retry, πt = integrate_answer_retry_break, stop_replanning, skip_decode_output_break, skip_decode_output, keep_decode,
nt = 1 ∨ Mt ̸= ∅, ft = tool_conflict, ft = tool_set_unstable, ft = low_confidence ∧ κt < τretry , ft = answer_incomplete ∧ nt = 0 ∧ Mt = ∅, = 1 ∧ ft = ok, Ct = 1 ∧ Gstab t Ct = 1 ∧ κt ≥ τgate , κt ≥ τgate , otherwise.
(7) Here, πt is a mathematical abstraction of the code’s control branches. It is not the name of a standalone function in the implementation. The branches are evaluated from top to bottom. Unresolved evidence, conflicting tool states, and unstable tool states take priority over stability-based stopping. The thresholds τretry and τgate are fixed runtime settings for low-confidence retry and decode skipping. In particular, = 1 only signals that the neighboring tool state is Gstab t stable. It triggers stop_replanning only when the closure signal is active and the verifier state is ok. In Algorithm 1, the branches of Equation (7) are grouped by whether the bounded loop stays active: branches that request another model or tool round map to CI, the answer-integration branch maps to AS, and the closure branches map to TI. The state in Fig. 5 is compressed by the Runtime State Manager, checked by the structured verifier, and consumed by the Iteration Controller. If evidence is missing, nt = 1 can request another tool round. If evidence is already available but at = 0, answer integration is preferred. If the action and result signatures are stable and the answer state is closed, the controller stops replanning. This separation prevents offline diagnostic metrics from being mistaken for online control inputs. B. Trace Diagnostics and Algorithms The next four equations define offline diagnostics from completed request traces. Let ui be the number of tool calls for request i, and let ci ∈ {0, 1} be its final correctness flag, where one means correct and zero means incorrect. The posttool failure flag is pi = I[ui > 0 ∧ ci = 0].
(8)
Algorithm 1 AgentLoop Closed-loop Execution Control Require: q: request; T : tool set; B = (Tmax , Θ, Ω): runtime budget Ensure: Y ∗ : final answer 1: Tmax : maximum number of rounds; Θ : thresholds; Ω : similarity switches; A = {CI, AS, TI} 2: CI := Continue Invocation; AS := Answer Synthesis; TI := Terminate Iteration 3: Y : answer draft; Ē : evidence summary; H : execution history; Call t , Result t : round states 4: Vt = (at , Mt , nt , κt , ft ) : verifier state; at , nt ∈ {0, 1}; κt ∈ [0, 1]; Mt : missing set; ft : failure state 5: Ct : closure; Stcall , Stresult : similarities; : stability gate; πt ∈ A : control action Gstab t 6: I[·] ∈ {0, 1}; SynthesizeEvidence(Y, Ē) : answer synthesis; ReturnGroundedAnswer(Y, Ē) : grounded answer 7: // Initialization 8: (Y, Call −1 , Result −1 ) ← (∅, ∅, ∅); (Ē, H, t) ← (∅, ∅, 0) 9: // Online execution and state update 10: while t < Tmax do 11: (Y, Call t ) ← GenerateReasoningCalls (q, Y, Ē, H) 12: (H, Result t , Ē) ← ExecuteToolsUpdateState (T , Call t , H) 13: Vt ← VerifyAnswerState(q, Y, Ē) = (at , Mt , nt , κt , ft ) 14: Ct ← I[at = 1 ∧ nt = 0] 15: (Stcall , Stresult , Gstab )← t ComputeStateStability Vt , Call t−1 , Call t , Result t−1 , Result t 16: // Control decision and action dispatch 17: πt ← SelectControlAction (Ct , Vt , Gstab , B) t 18: if πt = CI then 19: (Call t−1 , Result t−1 , t) ← (Call t , Result t , t + 1) 20: else if πt = AS then ∗ 21: Y ← SynthesizeEvidence(Y, Ē) 22: return Y ∗ 23: else if πt = TI then 24: Y ∗ ← ReturnGroundedAnswer(Y, Ē) 25: return Y ∗ 26: end if 27: end while 28: Y ∗ ← ReturnGroundedAnswer(Y, Ē) 29: return Y ∗
This flag does not measure the tool error rate. It marks a request that used at least one tool but still ended incorrectly. Let inter repeati indicate repeated output across adjacent answer rounds and let intra repeati indicate duplicated content within the final answer. The repetition-risk flag is ri = I[inter repeati = 1 ∨ intra repeati = 1].
(9)
The combined low-value execution marker is zi = I[pi = 1 ∨ ri = 1].
(10)
This is a trace-based operational definition built for execution analysis. It is not a Shannon entropy, KL divergence, or strict information-gain estimate, and it serves as a practical control indicator for repetitive or low-yield execution. For N requests, the low-gain request rate (LGR) is PN LGR =
i=1 zi
N
.
(11)
IEEE TRANSACTIONS ON SERVICES COMPUTING
7
Algorithm 2 Low-value Execution Risk Detection Require: T = {(ui , ci , inter repeati , intra repeati )}N i=1 : completed traces N Ensure: {zi }i=1 , LGR = ng /N , LGTS = ug / max(1, u) 1: ui ∈ N0 ; ci , inter repeati , intra repeati ∈ {0, 1}; I[·] ∈ {0, 1} 2: ng : flagged requests; ug : flagged calls; u: total calls 3: (ng , ug , u) ← (0, 0, 0) 4: // Risk detection 5: for i = 1, . . . , N do 6: pi ← I[ui > 0 ∧ ci = 0], ri ← I[inter repeati = 1 ∨ intra repeati = 1] 7: zi ← I[pi = 1 ∨ ri = 1] 8: (ng , ug , u) ← (ng + zi , ug + zi ui , u + ui ) 9: end for 10: // Metric aggregation 11: LGR ← ng /N ; LGTS ← ug / max(1, u) 12: return {zi }N i=1 , LGR, LGTS
The low-gain tool-call share (LGTS) is PN i=1 ui zi , PN u > 0, PN i=1 i LGTS = i=1 ui 0, otherwise.
(12)
Here, LGR is request-level, while LGTS attributes tool calls to flagged requests. Both are computed after the trace is complete and are used for diagnosis rather than online stopping. In that sense, they summarize workload-level risk patterns instead of estimating the marginal value of any individual tool call. Algorithm 1 describes the online closed-loop execution procedure. Lines 1–8 define the runtime interface, control parameters, action abbreviations, state variables, and initialization. The online execution stage performs one model/tool round, updates the compressed evidence, verifies answer completeness, and compares neighboring call/result states to obtain the control signals (lines 9–15). The controller then selects and dispatches one of the three actions: continue invocation, synthesize the available evidence, or terminate with an evidencegrounded answer (lines 16–25). The bounded loop closes at line 26, and lines 27–28 provide a fallback answer when the round limit is reached. Thus, the algorithm separates evidence acquisition from answer synthesis and avoids extending the loop when further execution is unlikely to add useful information. Algorithm 2 provides the offline diagnostic counterpart of the online controller. The input records, variable domains, counters, and initial values are specified in lines 1–3. The riskdetection stage derives the post-tool failure flag pi , repetitionrisk flag ri , and combined low-value execution flag zi , while simultaneously accumulating flagged requests and tool calls (lines 4–9). The final stage aggregates LGR and LGTS and returns the request-level labels and diagnostic metrics (lines 10–12). Unlike Algorithm 1, this procedure is post-hoc and characterizes workload-level execution risk rather than selecting the next online action. C. Experimental Metrics and Complexity The execution graph closure rate (GCR) is an evaluation metric rather than an online control input: GCR =
N i 1 X h oki =1∧NonEmptyOutputi =1 I ∧(NoTool . (13) i =1∨NoPendingToolCalli =1) N i=1
Here, oki indicates successful execution, NonEmptyOutputi indicates that the final answer is nonempty, NoTooli indicates that no tool was used, and NoPendingToolCalli indicates that the final message has no pending tool call. GCR measures whether the execution graph returns to a final answer state. It is not an accuracy metric. At the event level, the stable tool-set stop rate (STSSR) is PJ j=1 I[actionj = stop replanning] STSSR = . (14) max(1, J) Here, J is the number of stability-gate checks and actionj is the action emitted at check j. This event-level metric should be distinguished from a request-level posterior stable-stop rate that may be computed by a separate analysis script. Complexity Analysis. Let Tmain be the actual number of main-loop tool rounds, with Tmain ≤ Tmax , and let Tref be the number of extra reflection tool rounds. Let Rmax be the maximum reflection rounds, Pmax the maximum calls in one tool round, Larg and Lres the average argument and result lengths, and Imax the maximum answer-integration or recheck count. Let Cplan be one planning-call cost, Ctool wall one round of tool wall-clock cost, Cverify one verifier-call cost, Cintegrate one answer-integration cost, and Ccontrol the local cost of control, signature construction, and gating. Let H be the context and history space, and let U be the number of different call keys in the tool cache. Then Time(AgentLoop) = O (Tmain + Tref )[Cplan + Ctool wall + Pmax (Larg + Lres )] + Rmax [Cverify + Ccontrol ] + Imax [Cintegrate + Cverify ] , ! H + U (Larg + Lres ) Space(AgentLoop) = O . +Pmax (Larg + Lres )
(15) The tool calls in one round may run concurrently, so Ctool wall is close to the wall time of the slowest tool. Total resource consumption still accumulates the cost of all tool calls. LLM inference and external service execution are black-box service costs and should not be reduced to ordinary string-length O(n) terms. AgentLoop’s added local work mainly comes from argument and result normalization, sorting, signature construction, and similarity computation. If the tool-state cache is disabled, the cache term containing U can be removed, but the history context H remains. VI. P ERFORMANCE E VALUATION Datasets and Tasks. The evaluation covers thousands of requests across three datasets. ToolBench emphasizes multiAPI selection and composition, GAIA emphasizes grounded QA, and Open-Agent-Trace emphasizes realistic multi-step workflows. ToolBench [9] contains 356 tool-intensive requests focused on API selection and multi-service composition . GAIA [40] denotes the evaluated subset of GAIA-traces. The original collection contains 1,204 grounded questionanswering trace requests, from which we select 1,000 requests with complete task and trace information for evaluation. Open-Agent-Trace [41] contains 499 requests representing realistic multi-step agent workflows .
IEEE TRANSACTIONS ON SERVICES COMPUTING LangGraph AgentLoop FC-Reflection
LangGraph FC-Reflection
AutoGen
• LangGraph
♦ AgentLoop
1
0.5
0.0
100
1
(d) Qwen: ToolBench.
10 Latency (s)
0.0
10
100 Latency (s)
(e) Qwen: GAIA.
(f) Qwen: Open-Agent-Trace.
Framework Wait
Model/Control Model/Control
Execution ToolTool Execution
6.3 5
0
LangGraph
FCReflection
40
Framework Wait
3.8
16.6 8.7
10 FCReflection
10.0
3.6
2
LangGraph
Model/Control
FCReflection
Execution (b)Tool Llama: GAIA.
Framework Wait
AutoGen AgentLoop
(d) Qwen: ToolBench.
26.4
20 7.3 0
LangGraph
FCReflection
7.2 5.1
2.5 LangGraph
FCReflection
AutoGen AgentLoop
Model/Control ToolOpen-Agent-Trace. Execution Framework Wait (c) Llama: 44.3
40 23.4
Framework Wait
9.8 8.0
5.0
0.0
AutoGen AgentLoop
Tool Execution
7.5
45.5 22.2
LangGraph
4.2
3.2
33.6
30
0
4
0
AutoGen AgentLoop
Model/Control (a) Llama: Tool Execution ToolBench.
20
Latency (s)
8.9
Model/Control
Latency (s)
9.2
Latency (s)
Latency (s)
10
Framework Wait Framework waiting
Latency (s)
Tool Execution 10.7
Latency (s)
100
0.5
End-to-end (E2E) latency ECDFs.
Model/Control
Fig. 7.
(c) Llama: Open-Agent-Trace. 1.0 ECDF
ECDF
ECDF
0.5
10 Latency (s)
10 100 AgentLoop FC-Reflection Latency (s) LangGraph AutoGen
(b) Llama: GAIA. 1.0
0.0
0.5
0.0
10 100 AgentLoop FC-Reflection Latency (s) LangGraph AutoGen
(a) Llama: ToolBench. 1.0
8
1.0
0.5
0.0
10 100 AgentLoop FC-Reflection Latency (s) LangGraph AutoGen
AgentLoop AutoGen
⋆ AutoGen
ECDF
0.5
0.0
Fig. 6.
▲ FC-Reflection
LangGraph FC-Reflection
1.0 ECDF
ECDF
1.0
AgentLoop AutoGen
AutoGen AgentLoop
40 25.8
21.0
20
0
9.9 LangGraph
(e) Qwen: GAIA.
FCReflection
AutoGen AgentLoop
(f) Qwen: Open-Agent-Trace.
End-to-end latency breakdown across model/control processing, tool execution, and framework waiting.
All methods use the same datasets, concurrency setting, timeout boundary, and cache-disabled configuration. We evaluate Llama-3-8B-Instruct and Qwen3.6-35B-A3B under the same agent settings. Baselines and Metrics. We compare AgentLoop with AutoGen [13], FC-Reflection [7], [20], and LangGraph [22]. Evaluation metrics include average E2E latency, FTR, average iteration depth, tool calls per request, tokens per request, ECDFs, average latency, effective-token redundancy, and answer accuracy. Execution success measures whether the process completes, while answer accuracy measures whether the final response is supported by evidence, so the two are reported separately rather than merged into a single score. Experimental Configurations. Experiments were conducted on a unified GPU node with four NVIDIA Tesla V100 PCIe GPUs, each with 32 GB of memory. All methods used the same hardware platform and model service. Inference used a unified local model service. The main experiments used Llama-3-8B-Instruct, and the cross-model setting used Qwen3.6-35B-A3B with GPTQ Int4 and vLLM FP16. The request rate was 1 request/s, the timeout was 180s, and cache was disabled. For each dataset and method, we used the identical request order and runtime parameters. The experimental
results analysis are as below: a) Latency Comparison: To compare latency distribution, Fig. 6 demonstrates that AgentLoop generally shifts requests toward lower E2E latency, especially on ToolBench and Open-Agent-Trace. In the Llama setting, GAIA retains a heavier long tail, whereas the Qwen setting shows a broader advantage across all three datasets. AgentLoop reaches 8.787s average latency on Llama ToolBench and 15.144s P90 latency on Llama Open-Agent-Trace; on Qwen, its ToolBench average is about 5.85s and the largest average reduction appears on GAIA. On ToolBench, the average latency is reduced by more than 70% relative to the stronger baselines. These gains result from slot-closure and stability checks that suppress repeated reasoning-action rounds, rather than from faster external tools. The remaining GAIA tail also shows that completion control cannot eliminate requests whose evidence is intrinsically difficult to obtain. The same control effect appears in the breakdown and FTR results. Fig. 7 shows that the clearest reduction is in model/control processing, confirming that AgentLoop avoids repeated planning while leaving tool execution unchanged. The advantage is therefore a coordination gain: the controller reduces unnecessary rounds before they become additional
IEEE TRANSACTIONS ON SERVICES COMPUTING LangGraph
FC-Reflection FC-Reflection AgentLoop
10.0 7.5 5.0 2.5 0.0
ToolBench
GAIA
20
10
Open-AgentTrace
0
GAIA
Open-AgentTrace
10 ToolBench
GAIA
Open-AgentTrace
(a) Average FTR comparison.
AutoGenAgentLoop AgentLoop AutoGen P90 FTR Latency (s)
Avg FTR Latency (s)
20
Fig. 9.
ToolBench
(b) P90 FTR comparison.
AutoGen LangGraphAgentLoop FC-Reflection
0
Dataset
5
Average and FC-Reflection P90 first-token response latency for Llama. LangGraph LangGraph FC-Reflection
30
TABLE II C ONTROLLER - SIDE CONTROL METRICS OF AGENT L OOP.
15
(a) Average FTR comparison.
Fig. 8.
FC-Reflection
AutoGen AgentLoop AgentLoop AutoGen P90 FTR Latency (s)
Avg FTR Latency (s)
LangGraph LangGraph AutoGen
9
75 50 25 0
ToolBench
GAIA
Open-AgentTrace
(b) P90 FTR comparison.
Average and P90 first-token response latency for Qwen.
model and framework waiting time. Figs. 8 and 9 further show earlier first-token responses on most datasets, with the clearest improvement against LangGraph; Qwen achieves the best FTR on all datasets. Thus, the latency benefit is shared across both backbones and includes both common-case and tail-response improvements. In the Llama results, the gains are strongest on workloads with multi-step service composition, where repeated planning contributes a larger share of the total delay. The weaker GAIA result is consistent with requests that require more evidence acquisition before closure. b) Token and Execution Cost: To compare token and execution cost, Fig. 10 demonstrates a consistent reduction in input and output tokens, comp baselines. AgentLoop achieves this mainly through shorter reasoning-action loops and a compact runtime state instead of repeatedly reassembling the full trace. The benefit is larger on longer traces because, once the required slots are covered, carrying the complete interaction history contributes overhead rather than useful state. The reduction is substantial against LangGraph on Llama, particularly on ToolBench and GAIA, and carries over to Qwen, where AgentLoop has the lowest input and output token costs on all three datasets. The two token views also separate prompt-side savings from response-side savings. Their consistent direction suggests that the reduction is not caused by truncating only generated answers, but by avoiding repeated context construction and unnecessary loop expansions. This is important for long-horizon service workflows, where inputtoken accumulation can dominate the total cost. It also shows that the control state is not merely a stopping signal, but a compact representation that reduces the amount of repeated evidence passed into later model calls. Fig. 12 combines execution depth, LLM calls, tool calls, and answer accuracy. A shorter bar group with a higher accuracy point is the desired efficiency-quality pattern. AgentLoop keeps depth moderate while preserving high accuracy in the Llama panels, staying near a depth of three on ToolBench and retaining the higher accuracy curve with a shorter chain on Open-Agent-Trace. In the Qwen panels, it is not always the most accurate method, but consistently shortens execution by making continuation conditional on unresolved slots and
Slot Closure
Stable Hit
Stable Stop
ToolBench 98.6% 97.3% 96.9% GAIA 81.6% 78.4% 77.9% Open-Agent-Trace 95.6% 91.4% 91.0% Note: The table shows whether the controller marks required slots as closed, detects stable evidence or action states, and converts stability into stopping or answer synthesis. All three quantities are recorded by the controller itself and describe its internal decision behavior, not the correctness of the answer.
positive marginal utility; on GAIA, it can lead in accuracy while using far fewer steps. The result indicates selective stopping rather than depth reduction alone. The comparison also clarifies that execution depth should not be interpreted as a quality objective by itself. AgentLoop continues when the slot state or evidence remains unresolved, but avoids spending additional depth after the state has become sufficiently stable. This explains why its efficiency gain can coexist with competitive accuracy. Although the absolute depth and accuracy vary between Llama and Qwen, the same conditional-continuation principle is visible in both settings. This behavior is especially important for service-computing workloads, where excessive depth increases not only model cost but also external invocation latency and scheduling pressure. By linking continuation to unresolved slots, AgentLoop reduces these costs while preserving the execution capacity needed for genuinely multistep requests. c) Redundancy Suppression and Slot Closure: To examine the effects of runtime controller, Table II reports Slot Closure, Stable Hit, and Stable Stop as controller-side control metrics. Let N be the number of requests, let Vi denote the slot-verification events recorded for request i, and let Gi denote the stability-gate checks recorded for request i. Slot Closure is the fraction of requests whose last verification event reports a closed runtime state, that is, Ct = 1 in Equation (3): N i 1 X h (i) (i) I Vi ̸= ∅ ∧ alast = 1 ∧ nlast = 0 . SlotClosure = N i=1 (16) Stable Hit is the fraction of requests in which the stability gate was reached at least once, so it measures gate coverage rather than the outcome of the comparison: N
StableHit =
1 X I[Gi ̸= ∅] . N i=1
(17)
Stable Stop is the fraction of requests in which stability was converted into an early stop of replanning: StableStop = N1
PN
i=1 I[∃ g ∈ Gi :
actiong = stop replanning] .
(18) All three metrics use the request count N as denominator. Since the gate exits replanning once it fires, StableStop ≤ StableHit, and their gap indicates reached gates that did not trigger early stopping. Stable Stop is request-level, whereas STSSR in Equation (14) is event-level. Slot Closure is 98.6% on ToolBench, 81.6% on GAIA, and 95.6% on Open-AgentTrace; Stable Hit and Stable Stop are also close, indicating that reached gates usually produce early stops. These
IEEE TRANSACTIONS ON SERVICES COMPUTING
ToolBench
GAIA
400 200
Open-AgentTrace
0
(a) Llama: input tokens.
Fig. 10.
FC-Reflection
600
AutoGen Tokens / Request
2000
LangGraph AutoGen
AutoGen AgentLoop
ToolBench
GAIA
Open-AgentTrace
(b) Llama: output tokens.
FC-Reflection AgentLoop
AgentLoop
6000
Tokens / Request
LangGraph
4000
0
10
LangGraph FC-Reflection
AutoGen AgentLoop
Tokens / Request
Tokens / Request
LangGraph FC-Reflection
4000 2000 0
ToolBench
GAIA Open-AgentTrace
LangGraph AutoGen
FC-Reflection AgentLoop
ToolBench
GAIA Open-AgentTrace
1000
500
0
(c) Qwen: input tokens.
(d) Qwen: output tokens.
Input and output token cost across Llama and Qwen. Non-Redundant
Estimated Redundant Part
TABLE III M EASURED SLOT- VERIFICATION OVERHEAD OF AGENT L OOP.
Reduction Ratio
80 923
1500
60
46.3% 1000 500
38.8%
33.6% 242
289
1070
479 0
20
455
ToolBench
0 Open-Agent-Trace
GAIA
Tool Calls
♦ Iter Depth
• Accuracy
40% 20% Lang FC- Auto Agent Graph Refl. Gen Loop
0%
60%
4
40%
2
20%
0
60% 40%
2
20% Lang FC- Auto Agent Graph Refl. Gen Loop
0%
20% 0%
(e) Llama: Open-Agent-Trace.
Calls / Depth (count)
40%
Accuracy (%)
Calls / Depth (count)
80%
Lang FC- Auto Agent Graph Refl. Gen Loop
40%
2
20% Lang FC- Auto Agent Graph Refl. Gen Loop
0%
(d) Qwen: GAIA.
60%
0
60%
4
0
100%
2
80%
6
(c) Llama: GAIA. 4
0%
100%
8 Calls / Depth (count)
80%
4
Lang FC- Auto Agent Graph Refl. Gen Loop
(b) Qwen: ToolBench. 100% Accuracy (%)
Calls / Depth (count)
(a) Llama: ToolBench. 6
80%
6
Accuracy (%)
1
8
Accuracy (%)
60%
2
100% Calls / Depth (count)
80%
Accuracy (%)
Calls / Depth (count)
100% 3
0
Avg. Verification Latency (ms)
Control-side effective-token redundancy for Llama. LLM Calls
0
Slot Verification Time Share
ToolBench 3.61% 317.4 GAIA 0.40% 129.9 Open-Agent-Trace 1.72% 206.6 Note: Time share is the total slot-verification time divided by the total end-to-end latency of the same requests, and latency is the per-request verification time, that is, the verification time accumulated within a request and averaged over all requests.
100%
8
80%
6
60%
4
40%
2
20%
0
Lang FC- Auto Agent Graph Refl. Gen Loop
Accuracy (%)
Fig. 11.
40
Dataset
Reduction Ratio (%)
Tokens / Request
100 2000
0%
(f) Qwen: Open-Agent-Trace.
Fig. 12. Execution depth, LLM calls, tool calls, and answer accuracy across Llama and Qwen. Each row pairs the two backbones on the same dataset, with Llama in the left column and Qwen in the right column. Horizontal-axis method labels are wrapped over two lines, where FC-Refl. abbreviates FCReflection.
values describe controller behavior rather than final-answer correctness, which is evaluated separately. In a stratified handlabeled sample of 96 requests aligned with verifier decisions, precision, recall, and F1 were 0.840, 0.913, and 0.875, respectively. Because the sample over-represents difficult cases,
these values characterize verifier behavior on hard requests rather than the average request. The lower closure rate on GAIA is consistent with its longer-tail behavior: more requests finish with unresolved evidence even though the controller still detects stable states in a substantial fraction of cases. The close Stable Hit and Stable Stop values on all datasets indicate that the stability signal is usually converted into an action instead of remaining only an internal observation. This separation also prevents the controller-side metrics from being mistaken for accuracy measurements: they explain why the loop stops, while the final answer quality is evaluated by the accuracy results above. This distinction is important because an early stop can be desirable only when it is supported by slot closure or stable evidence, not when it simply reflects an exhausted budget or an arbitrary workflow boundary. Table III instantiates the Cverify term of Equation (15). Slot verification accounts for at most 3.61% of E2E latency, with per-request verification time between 129.9 ms and 317.4 ms. ToolBench has the highest share (3.61%) because its shorter chains provide less amortization, while GAIA has the lowest share (0.40%) because its longer requests amortize verification overhead. Thus, the control layer remains a second-order cost relative to model inference and tool execution, although the fixed cost is more visible for very short tool chains. The tradeoff is favorable when one verification step prevents several later reasoning or invocation rounds; it is less pronounced when the task needs only a short tool chain. Fig. 11 summarizes the Llama redundancy analysis for the Llama backbone, where the goal is to estimate how much of the token budget is attributable to repeated or low-value context under the heuristic decomposition. We can note that up to 46.3% unnecessary tokens can be reduced in GAIA. d) Ablation Study: To evaluate the slot-centered control path, Fig. 13 compares the full system with AgentLoop w/o Slot and shows whether slot control improves efficiency without damaging answer quality. The ablation indicates fewer unproductive loops and higher final answer quality with slot-
IEEE TRANSACTIONS ON SERVICES COMPUTING
ToolBench Accuracy Open-Agent-Trace Accuracy 7.4
8
1.0 6.9
6 4
3.4
0.6
4.2
0.4
2 0
0.8
Accuracy (%)
Avg Iteration Depth
ToolBench Open-Agent-Trace
11
0.2 AgentLoop Depth
AgentLoop w/o Slot Depth
0.0
Fig. 13. Ablation of the slot-centered control path by comparing full AgentLoop with AgentLoop w/o Slot.
centered control. The ablation keeps the same base model, tool environment, and dataset setting, but disables the slot-centered control path as a group, including the reflection score gate, convergenceround test, and answer integration/recheck step. It therefore measures the contribution of the coordinated slot-control path rather than the isolated effect of one switch. Without this path, average depth increases from about 3.4 to 7.43 and accuracy decreases from about 0.938 to 0.612 on ToolBench; on Open-Agent-Trace, depth increases from about 4.23 to 6.86 and accuracy decreases from about 0.949 to 0.584. These results show that coupling slot completion with Iteration Controller decisions improves the efficiency-quality trade-off. The direction is consistent across both datasets: removing slotcentered control allows additional rounds without providing the same answer-quality support. The ablation thus attributes the improvement to explicit request-state tracking and its connection to the controller, rather than to a generic reduction in the number of calls.
VIII. C ONCLUSION This paper presented AgentLoop, a runtime slot-closed execution control mechanism for tool-augmented LLM agents. The central view is that open-ended agent iteration can be converted into explicit runtime-state control: slot closure provides a task-level completion signal, while context stability, evidence stability, budget pressure, and low-value heuristics guide execution-level decisions. By coordinating Runtime State Manager, Slot State Verifier, and Iteration Controller, AgentLoop represents user requests as slot states, verifies evidence, attributes gaps, suppresses redundant context and evidence states, guards budget, and selects Continue Invocation, Answer Synthesis, or Terminate Iteration. Experiments on thousands of requests from three datasets show that AgentLoop improves the balance between execution efficiency and answer quality when runtime context and tool evidence have room to stabilize. Across the two model backbones, AgentLoop reduces redundant context growth and highdepth execution, with total token cost reduced by up to 88.44% and average depth reduced by up to 76.85% against high-cost baselines. In practice, this saving is largely driven by fewer loop expansions and less trace carryover. The ablation further shows the role of slot-centered control: disabling it increases average depth from 3.4 to 7.43 on ToolBench and from 4.23 to 6.86 on Open-Agent-Trace, while reducing accuracy from 0.938 to 0.612 and from 0.949 to 0.584. Future work will explore lighter slot verification, finer-grained budget control, domain-specific slot extraction, and stability evaluation under concurrent service workloads. R EFERENCES
VII. D ISCUSSION AgentLoop is intended as a runtime-state control layer rather than a replacement for existing agent frameworks. Its main value is most visible when an agent must coordinate multiple reasoning-action rounds and decide whether further execution will add useful information. In such settings, slot closure provides a task-level completion signal, while stability and low-value heuristics provide execution-level control signals. These signals complement model self-judgment, tool-call success, and workflow boundaries. The results also define a practical deployment boundary. For tool-intensive or workflow-style requests, the additional state checks can reduce repeated context growth and move execution toward answer synthesis once the required information is stable. This makes AgentLoop suitable as an embedded control component inside broader service-computing agent systems. Several limitations follow from how the control layer is currently built. The stability test that decides whether a state has converged is computed from lexical and set-level overlap, so a semantically equivalent paraphrase is still counted as a change and can permit one further round. The ablation in Fig. 13 disables the slot-centered path as a group, which means the reported gap measures the combined effect of slot verification, the reflection score gate, and answer integration rather than the marginal value of any single component.
[1] J. Wei, X. Wang, D. Schuurmans, et al., “Chain-of-thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems (NeurIPS), 2022, pp. 24 824–24 837. [2] X. Wang, J. Wei, D. Schuurmans, et al., “Self-consistency improves chain of thought reasoning in language models,” in International Conference on Learning Representations (ICLR), 2023. [3] S. Yao, D. Yu, J. Zhao, et al., “Tree of thoughts: Deliberate problem solving with large language models,” in Advances in Neural Information Processing Systems (NeurIPS), 2023, pp. 11 809–11 822. [4] L. Wang, W. Xu, Y. Lan, et al., “Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models,” in Annual Meeting of the Association for Computational Linguistics (ACL), 2023, pp. 2609–2634. [5] S. Yao, J. Zhao, D. Yu, et al., “ReAct: Synergizing reasoning and acting in language models,” in International Conference on Learning Representations (ICLR), 2023. [6] E. Karpas, O. Abend, Y. Belinkov, et al., “MRKL systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning,” arXiv preprint arXiv:2205.00445, 2022. [7] T. Schick, J. Dwivedi-Yu, R. Dessı̀, et al., “Toolformer: Language models can teach themselves to use tools,” in Advances in Neural Information Processing Systems (NeurIPS), 2023, pp. 68 539–68 551. [8] S. G. Patil, T. Zhang, X. Wang, et al., “Gorilla: Large language model connected with massive apis,” in Advances in Neural Information Processing Systems (NeurIPS), 2024, pp. 126 544–126 565. [9] Y. Qin, S. Liang, Y. Ye, et al., “ToolLLM: Facilitating large language models to master 16000+ real-world apis,” in International Conference on Learning Representations (ICLR), 2024. [10] M. Li, Y. Zhao, B. Yu, et al., “API-bank: A comprehensive benchmark for tool-augmented LLMs,” in Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023, pp. 3102–3116. [11] Y. Song, W. Xiong, D. Zhu, et al., “RestGPT: Connecting large language models with real-world restful apis,” arXiv preprint arXiv:2306.06624, 2023. [12] T. Xie, F. Zhou, Z. Cheng, et al., “Openagents: An open platform for language agents in the wild,” in Conference on Language Modeling (COLM), 2024.
IEEE TRANSACTIONS ON SERVICES COMPUTING
[13] Q. Wu, G. Bansal, J. Zhang, et al., “Autogen: Enabling next-gen llm applications via multi-agent conversation,” in Conference on Language Modeling (COLM), 2024. [14] S. Kim, S. Moon, R. Tabrizi, et al., “An llm compiler for parallel function calling,” in International Conference on Machine Learning (ICML), 2024, pp. 24 370–24 391. [15] M. Xu, J. Wu, S. Song, et al., “Cloud-native and distributed systems for efficient and scalable large language models – a research agenda,” arXiv preprint arXiv:2604.17227, 2026. [16] J. Liao, M. Xu, W. Zheng, et al., “DOPD: A dynamic PD-disaggregation architecture for maximizing goodput in LLM inference serving,” IEEE Transactions on Services Computing, vol. 19, no. 2, pp. 1134–1147, 2026. [17] M. Dadhich, “Agentic context management: Solving long-context failures in llm agents,” arXiv preprint arXiv:2607.21503, 2026. [18] X. Li, R. Ming, M. Chu, et al., “ACM: Agentic context management for long horizon tasks,” arXiv preprint arXiv:2607.23809, 2026. [19] L. Bai, Z. Huang, X. Wang, et al., “How do ai agents spend your money? analyzing and predicting token consumption in agentic coding tasks,” arXiv preprint arXiv:2604.22750, 2026. [20] N. Shinn, F. Cassano, A. Gopinath, et al., “Reflexion: Language agents with verbal reinforcement learning,” in Advances in Neural Information Processing Systems (NeurIPS), 2023, pp. 8634–8652. [21] A. Madaan, N. Tandon, P. Gupta, et al., “Self-refine: Iterative refinement with self-feedback,” in Advances in Neural Information Processing Systems (NeurIPS), 2023, pp. 46 534–46 594. [22] LangChain, “Langgraph documentation,” Online, 2026, accessed: 2026-0519. [23] L. E. Erdogan, N. Lee, S. Kim, et al., “Plan-and-act: Improving planning of agents for long-horizon tasks,” in International Conference on Machine Learning (ICML), 2025, pp. 15 419–15 462. [24] Y. Guo, J. Lin, H. Wang, et al., “SE-Agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents,” in Advances in Neural Information Processing Systems (NeurIPS), 2025, pp. 128 912–128 939. [25] S. Wu, S. Zhao, Q. Huang, et al., “AvaTaR: Optimizing llm agents for tool usage via contrastive reasoning,” in Advances in Neural Information Processing Systems (NeurIPS), 2024, pp. 25 981–26 010. [26] B. Paranjape, S. Lundberg, S. Singh, et al., “ART: Automatic multistep reasoning and tool-use for large language models,” arXiv preprint arXiv:2303.09014, 2023. [27] P. Lu, B. Peng, H. Cheng, et al., “Chameleon: Plug-and-play compositional reasoning with large language models,” in Advances in Neural Information Processing Systems (NeurIPS), 2023, pp. 43 447–43 478. [28] G. Wölflein, D. Ferber, D. Truhn, et al., “Llm agents making agent tools,” in Annual Meeting of the Association for Computational Linguistics (ACL), 2025, pp. 26 092–26 130. [29] B. Wu, E. Meij, and E. Yilmaz, “A joint optimization framework for enhancing efficiency of tool utilization in llm agents,” in Findings of the Association for Computational Linguistics (ACL Findings), 2025, pp. 22 361–22 373. [30] X. Liu, H. Yu, H. Zhang, et al., “AgentBench: Evaluating llms as agents,” in International Conference on Learning Representations (ICLR), 2024. [31] G. Mialon, C. Fourrier, C. Swift, et al., “Gaia: a benchmark for general ai assistants,” 2023. [Online]. Available: https://arxiv.org/abs/2311.12983 [32] M.-B. Chen, J. H. Lau, and L. Frermann, “CIG: Measuring conversational information gain in deliberative dialogues with semantic memory dynamics,” in Annual Meeting of the Association for Computational Linguistics (ACL), 2026, pp. 47 702–47 725. [33] J. Zhang, J. Xiang, Z. Yu, et al., “AFlow: Automating agentic workflow generation,” in International Conference on Learning Representations (ICLR), 2025. [34] S. Hu, C. Lu, and J. Clune, “Automated design of agentic systems,” in International Conference on Learning Representations (ICLR), 2025. [35] M. Zhuge, W. Wang, L. Kirsch, et al., “GPTSwarm: Language agents as optimizable graphs,” in International Conference on Machine Learning (ICML), 2024, pp. 62 743–62 767. [36] Z. Z. Wang, J. Mao, D. Fried, et al., “Agent workflow memory,” in International Conference on Machine Learning (ICML), 2025, pp. 63 897–63 911. [37] Y. Dang, C. Qian, X. Luo, et al., “Multi-agent collaboration via evolving orchestration,” in Advances in Neural Information Processing Systems (NeurIPS), 2025, pp. 182 975–183 009. [38] M. Luo, X. Shi, C. Cai, et al., “Agentix: An efficient serving engine for llm agents as general programs,” in USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2026, pp. 2443–2459. [39] H. Bai, M. T. Islam, M. Xu, et al., “ORACL: Optimized reasoning for autoscaling via chain of thought with LLMs for microservices,” IEEE Transactions on Services Computing, pp. 1–14, 2026. [40] S. Bose, “gaia traces,” Hugging Face Dataset, 2025, accessed: 2026-08-27. [Online]. Available: https://huggingface.co/datasets/shamikbose89/gaia traces [41] J. Simon, “Open agent traces: Synthetic multi-agent workflow datasets,” Hugging Face dataset, 2026.
12
Wanyi Zheng Wanyi Zheng received her Bachelor’s degree from Hebei University in 2025. She is currently pursuing a Master’s degree at the Southern University of Science and Technology and is also a joint-training student at the Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences. Her research interests include LLMbased agent systems and intelligent service systems, with a primary focus on agent runtime optimization, efficient tool use, and system-level optimization for LLM-powered agents.
Minxian Xu (Senior Member, IEEE) received the Ph.D. degree from the University of Melbourne, Melbourne, VIC, Australia, in 2019. He is currently an Associate Professor with Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences. His research interests include cloud-native systems for LLM inference and AI infrastructure. He has co-authored more than 90 peer-reviewed papers in prominent journals and conferences, including ACM CSUR, IEEE TSC, TC, TMC, TAAS, TCC, TOIT, ICSOC and ICWS, thes work attracted 7,000+ citations. He serves as the Associate Editor of IEEE TSC.
Kan Hu received his Master’s degree from the University of Chinese Academy of Sciences in 2025. He has been awarded a PhD scholarship at the University of Melbourne and is yet to commence his doctoral studies. His research interests include LLM-based agent systems, and he particularly concentrates on optimizing large-scale multi-agent systems by leveraging cloud computing paradigms—especially in terms of resource elasticity, load balancing, and efficient task coordination, aiming to enhance system performance.
Kejiang Ye (Senior Member, IEEE) received the B.S. and Ph.D. degrees from Zhejiang University. He is currently a Professor and the Director of the Research Center for Cloud Computing, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences. He was a Postdoctoral Research Associate with Carnegie Mellon University, Pittsburgh, PA, USA. His research interests include digital technology and systems, such as cloud computing, big data, and industrial Internet. He is a Distinguished Member of the China Computer Federation.
Chengzhong Xu (Fellow, IEEE) received the Ph.D. degree in computer science and engineering from The University of Hong Kong, Hong Kong, in 1993. He is currently the Dean of the Faculty of Science and Technology and the Interim Director of the Institute of Collaborative Innovation, University of Macau. His research interests include parallel and distributed computing, with an emphasis on resource management for performance, reliability, availability, power efficiency, and security. His work spans servers and cloud datacenters, wireless embedded devices, and edge AI systems, with applications in smart city and autonomous driving. He has authored two research monographs and more than 600 papers, which have received more than 22,000 citations with an H-index of 77, and have been cited in more than 300 international patents.