Conceptio › Archive › arXiv CS
arXiv CSopen access

Beyond State Machines: Executing Network Procedures with Agentic Tool-Calling Sequences

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

Beyond State Machines: Executing Network Procedures with Agentic Tool-Calling Sequences Purna Sai Garigipati∗† , Onur Ayan∗ , Kishor Chandra Joshi† , Xueli An∗ ∗ Heisenberg Research Center, Huawei Technologies Duesseldorf GmbH, 80992 Munich, Germany

Email: {purna.sai.garigipati,onur.ayan,xueli.an}@huawei.com † Eindhoven University of Technology, Eindhoven, The Netherlands

arXiv:2605.02584v1 [cs.NI] 4 May 2026

Email: {k.c.joshi}@tue.nl

Abstract—Agentic AI will be an essential enabling technology for designing future mobile communication systems, which could provide flexible and customized services, automate complex network operations, and drive autonomous decision-making across the network. This work studies how Large Language Model (LLM)-based network AI agents can be utilized to execute network procedures expressed as sequences of tool invocations. We investigate four approaches, which differ in how the agent obtains the procedure and in how execution is distributed between the agent and the underlying tools. We evaluated the latency and execution correctness across these approaches using a User Equipment (UE) IP allocation procedure as a case study. Furthermore, we conduct a stress test to examine how many sequential procedural steps an LLM agent can reliably execute before failure. Our results show that approaches relying on iterative agent-side reasoning incur higher latency and are more prone to execution errors, while approaches where the procedure is encapsulated within a single tool, which internally orchestrates the required steps by invoking other tools, reduce latency by limiting repeated reasoning. The stress-test results further show that the model with advanced tool-calling capability maintains reliable execution over longer procedures than the other evaluated models; however, all models exhibit reliability degradation as procedure length increases, revealing clear execution limits in multi-step tool-based workflows. To systematically analyze failures in procedure execution, we introduce a procedurespecific error taxonomy that categorizes deviations in multi-step procedural execution. Index Terms—Large Language Model (LLM), Agentic AI, Mobile Communication Networks, Procedure Execution

I. I NTRODUCTION Agents empowered by Large Language Models (LLMs) have introduced a new paradigm of autonomous systems capable of reasoning, planning, and interacting with external tools to accomplish complex tasks. Such systems extend beyond single-shot inference and operate through iterative decisionmaking and tool invocation. This paradigm shift has gained significant attention because it enables complex workflows to be executed without explicitly hard-coded control logic, instead relying on the model inference to determine the sequence of actions required to achieve a given objective. Recent work shows that agentic approaches are being actively explored in the context of next-generation networks, particularly in 6G. Several studies focus on the Radio Access Network (RAN), where agentic frameworks have been proposed for real-time control, resource management, and op-

timization [1]–[3]. In addition, recent efforts consider end-toend intelligence across both RAN and core networks, integrating monitoring, policy control, and cross-layer optimization using agentic approaches [4]–[6]. Collectively, these works indicate a shift toward agent-driven network automation. In this context, we study the use of LLM-based agents to execute network procedures. Typically, a network procedure consists of a sequence of dependent operations that must be executed in a strict order to achieve a target system state. Traditionally, such procedures are implemented as scripts or workflows, which require explicit development, testing, and deployment. Supporting variability in network conditions often requires extensive conditional logic, making these implementations difficult to scale and maintain. Moreover, such implementations are tightly coupled to predefined states and can fail when the observed network state deviates from expected conditions. In contrast, an agent-based approach enables procedures to be specified at a higher level using natural language descriptions. Given a set of standardized tools, different procedures can be dynamically composed by varying the sequence of tool invocations based on the task and intermediate outcomes, without requiring explicit reprogramming. This flexibility is important in scenarios where a network operator or an application needs to execute a new procedure that is not predefined in existing specifications. This raises a fundamental question: can an LLM-based agent reliably execute telecom-grade procedures? To answer this, we study the agent’s ability to perform sequential tool invocation, focusing on whether the correct step ordering is maintained, how the execution behaves across repeated runs, what types of procedural violations occur, and how the reliability changes as the length of the procedure increases. We further analyze different mechanisms for providing the procedure to the agent, reflecting realistic deployment scenarios. Although agentic systems enable flexible execution, previous work shows that LLM-based agents are prone to failures during multi-step reasoning and tool interaction, and existing studies have proposed taxonomies to characterize such failures, covering aspects such as planning errors, tool invocation issues, and execution inconsistencies [7]–[9]. However, these taxonomies are designed for open-ended or generalized domains and do not address the strict constraints of sequential

tool execution. We therefore define a procedure-specific error taxonomy tailored to this setting, enabling precise analysis of procedural execution correctness for agentic systems. The main contributions of this paper are summarized as follows: • Procedural Execution Approaches: We present four approaches for delivering procedural logic to LLM agents and characterize how procedure placement impacts endto-end latency and execution correctness. • Scalability Limits of Sequential Tool Execution: We identify the execution limits of LLM agents by showing how performance degrades as the number of sequential steps increases. • Procedure-Specific Error Taxonomy: We define an error taxonomy tailored to sequential tool execution, enabling precise classification of failures in multi-step agentic workflows. The remainder of this paper is organized as follows. Section II formulates the procedural execution model, defines the evaluation metrics, and presents the execution approaches and error taxonomy. Section III describes the experimental scenarios and presents the evaluation results. Section IV concludes the paper. II. M ETHODOLOGY This section presents the procedural execution model considered in this work, the performance metrics used for evaluation, the four execution approaches, and the error taxonomy adopted for analyzing failures. A. Definition of a Procedure as Tool Sequence We consider a task execution setting in which an LLMbased agent interacts with a tool server that exposes a total number of m tools, denoted by T = {τ1 , τ2 , . . . , τm }. Given a user intent i, the agent must identify and execute the matching procedure Pi with Pi ∈ P where P denotes the set of procedures available at the agent. We define procedure Pi as an ordered sequence of tool calls, where each required tool τi,j belongs to the available toolset T : Pi = (τi,1 , τi,2 , . . . , τi,k ),

(1)

Here, the index i represents the procedure that matches the user’s intent from a set of possible procedures P. The second index represents the step number within the procedure. For example, τi,1 is the first tool occurring in the procedure Pi followed by τi,2 up to the last tool τi,k . The observed execution produced by the agent is similarly represented as an ordered sequence of tool calls: Oi = (τ̂i,1 , τ̂i,2 , . . . , τ̂i,k̂ ),

(2)

where τ̂i,j denotes the tool invoked at step j during execution of procedure i, and k̂ is the number of executed steps. In this formulation, Pi represents the correct procedure (i.e., ground-truth), while Oi denotes the sequence of tool calls executed by the agent.

B. Evaluation Metrics We evaluate procedural execution using latency and execution correctness. The total latency cost C(i) is defined as the end-to-end execution time required to process intent i, including both the LLM reasoning and tool invocation: C(i) =

Nllm X j=1

Lllm j +

k̂ X

Ltool j ,

(3)

j=1

where Nllm is the number of LLM reasoning steps, k̂ is the tool number of tool invocations executed, and Lllm denote j and Lj the latency of the corresponding reasoning and tool-execution step, respectively.1 Execution correctness is measured using a binary reliability metric R(Pi ). For a single run, it evaluates whether the observed sequence of actions Oi perfectly matches the expected procedure Pi : ( 0, if k ̸= k̂ or τi,j ̸= τ̂i,j for at least one j (4) R(Pi ) = 1, otherwise By this definition, a score of 0 indicates that the observed sequence Oi differs from Pi in length (k ̸= k̂) or composition (τi,j ̸= τ̂i,j ), which implies an erroneous procedure execution as categorized later in Section II-D. Averaging R(Pi ) over repeated runs yields the execution correctness rate for a given model and approach. C. Procedural Execution Approaches The four approaches considered in this work are shown in Fig. 1. They share the same basic agent–tool interaction setting, but differ in where the procedure is defined and how the sequential execution is carried out, particularly in the distinction between iterative agent-driven execution and toolencapsulated execution. • A1: Agent-Embedded Procedure: The set of procedures P is available to the agent as its system prompt. After receiving intent i, the agent must first parse the prompt to identify and extract the matching procedure Pi ∈ P. Once identified, the agent reasons over this specific sequence and invokes the required tools step by step. For a procedure of length k, this leads to approximately Nllm ≈ k + 1 reasoning steps,2 where the final step corresponds to response summarization. • A2: Server-Provided Procedure: The agent first retrieves the procedure information P from an external database or repository, parses the specific procedure Pi for the intent, and then executes it sequentially in the same manner as A1. This additional retrieval step increases the number of reasoning steps to Nllm ≈ k + 2. • A3: User/Agent-Provided Procedure: The specific procedure Pi is explicitly specified within the incoming 1 Transmission latency is negligible and is therefore excluded from the latency cost equation. 2 Values are approximate as they represent the optimal execution path; in practice, model errors or tool failures often cause the agent to deviate, leading to additional reasoning turns or early termination.

(a)

(b)

User/Agent

LLM Agent

MCP Server

System prompt 𝓅

User/Agent

LLM Agent Fetch procedure 𝓅 Return 𝓅

Send intent 𝑖

Send intent 𝑖 Parsing procedure 𝑃

A1

𝑁 ≈𝑘+1 Reasoning turns

Send summarized response

Procedure DB

MCP Server

Parsing procedure 𝑃

START LOOP For step 𝑗 = 1 TO 𝑘 Invoke Tool 𝜏̂ ,

𝑁 ≈ 𝑘+2 Reasoning turns

START LOOP For step 𝑗 = 1 TO 𝑘 Invoke Tool 𝜏̂ ,

Return result 𝑟 ,

Return result 𝑟 ,

Agent processes feedback 𝑟 , for next turn END LOOP

Send summarized response

Agent processes feedback 𝑟 , for next turn END LOOP

A2

(c)

(d)

User/Agent

LLM Agent

Send intent 𝑖 with procedure 𝑃 = (𝜏 , , 𝜏 , , …, 𝜏 , )

MCP Server

User/Agent Send intent 𝑖

Parsing procedure 𝑃 START LOOP For step 𝑗 = 1 TO 𝑘

A3

LLM Agent

𝑁 ≈ 𝑘+1 Reasoning turns

Invoke Tool 𝜏̂ ,

Send summarized response

Agent processes feedback 𝑟 , for next turn END LOOP

A4

MCP Server Send single call for Encapsulated Tool 𝜏

𝑁 ≈2 Reasoning turns (trigger call+summary)

Return result 𝑟 ,

Tool 𝜏 internally executes procedure 𝑃 Return final result of 𝑃

Send summarized response

Fig. 1. Comparison of four procedural execution approaches. (a) A1 embeds the procedure within the agent, (b) A2 retrieves the procedure from an external database, (c) A3 receives the procedure in the input prompt, and (d) A4 encapsulates the procedure within a single tool. The figure highlights the difference between iterative multi-step execution (A1–A3) and single-call execution (A4).

prompt, which may originate from a user or another agent. The LLM agent parses the sequence and executes it iteratively, resulting in a reasoning overhead comparable to A1, i.e., Nllm ≈ k + 1. • A4: Tool-Encapsulated Procedure: The entire logic of Pi is encapsulated inside a single tool (τE ). After receiving the user intent, the agent selects and invokes this tool once, and the internal procedure is executed within the tool implementation. Thus, the number of LLM reasoning steps is reduced to approximately Nllm ≈ 2, corresponding to one trigger call and one final summarization step. This shifts execution complexity from the LLM to deterministic tool logic. D. Error Taxonomy for Tool-Based Procedures Whenever a run fails, i.e., when R(Pi ) = 0, the observed deviation is categorized to identify the failure mode. We adopt a strict error taxonomy because exact procedural execution is essential in structured network tasks. 1) Wrong Tool: A Wrong Tool error occurs when the agent fails to correctly invoke the required tool for a given step in the procedure. This can manifest in three forms: • Tool Outside Procedure: The agent invokes a tool τ̂i,j that does not appear in Pi , i.e., a tool that does not belong to the intended procedure. • Wrong Tool Name: The agent intends to use the correct tool but invokes it using an incorrect or hallucinated name.

Wrong Parameters: The agent invokes the correct tool but provides incorrect, missing, or invalid input arguments. This corresponds to incorrect tool invocation and represents a critical deviation, as it introduces actions outside the intended execution logic. Due to this severity, any execution containing both Wrong Tool and Duplicate Tool errors is strictly classified as a Wrong Tool error. 2) Duplicate Tool: A Duplicate Tool error is recorded when a valid tool τ ∈ Pi is invoked multiple times unnecessarily within Oi , without progressing the execution state. This typically reflects a reasoning loop and leads to an increase in both the number of reasoning steps Nllm and, subsequently, the latency cost C(i). 3) Premature Stop: A Premature Stop error occurs when the execution terminates before completing all k steps of the procedure, even though the executed tools are in the correct order up to the stopping point. In other words, Oi is a proper prefix of Pi with k̂ < k. 4) Wrong Order: A Wrong Order error is recorded when all required tools are present in Oi , but the sequence does not match the prescribed ordering of Pi , i.e., Oi ̸= Pi despite containing the same elements. 5) No Tool Calls: A No Tool Calls error is defined when the observed sequence is empty, i.e., Oi = ∅, indicating that no tool invocation was performed despite the requirement for procedural execution. •

User/Agent

LLM Agent Request

Scenario A: UE IP Allocation Encapsulated tools

Invoke Tool Return result

IP allocation tools

Scenario B: Scalability Stress Test Analytics tools

𝜏

𝜏

𝜏

𝜏

𝜏

𝜏

𝜏

𝜏

𝜏

𝜏

𝜏

𝜏

𝜏

MCP Server 2

MCP Server 3

Response MCP Server 1

……

𝜏

Fig. 2. Overview of the experimental setups. Scenario A illustrates the UE IP Allocation workflow across two MCP servers: MCP Server 2 provides the IP allocation tools used in approaches A1–A3, while MCP Server 1 hosts encapsulated tools used in approach A4. The highlighted tools in Scenario A indicate the subset of tools executed for the representative request. Scenario B utilizes MCP Server 3, providing a pool of m = 100 network analytics tools to stress-test sequence stability for procedures of length up to k = 50.

III. E XPERIMENTAL E VALUATION To evaluate the performance and stability of LLM-based procedure execution, we designed two experimental scenarios, as shown in Fig. 2. Scenario A compares the four approaches introduced in Section II-C, while Scenario B serves as a stress test to evaluate how the correctness of sequential tool execution is affected by the target procedure length k. A. Scenario A: UE IP Allocation Scenario A simulates a simplified User Equipment (UE) IP allocation procedure. At a high level, the procedure first authorizes the UE and the requested session type, then checks whether a static IP is already assigned. If no static IP for the given UE is found, the network allocates an IP address using the appropriate DHCP server. Finally, the assigned IP is recorded in the registry. We adapted the logical guidelines of 3GPP standards [10], [11] into a set of executable network tools, where the orchestration procedure is executed through LLM-driven tool invocation. As shown in Scenario A of Fig. 2, MCP Server 2 hosts the IP allocation tools used directly in approaches A1, A2, and A3, where the agent executes the procedure step-by-step. These tools are defined as follows: 1) UE Authorization (τauth ): Validates the UE ID and the requested session type. 2) Static IP Retrieval (τstatic ): Checks whether a static IP is pre-assigned for the UE. 3) Dynamic IPv4 Allocation (τdhcpv4 ): Dynamically allocates an IPv4 address. 4) Dynamic IPv6 Allocation (τdhcpv6 ): Dynamically allocates an IPv6 address. 5) IP Assignment (τregistry ): Finalizes IP assignment by notifying the UE and recording the allocation in the network registry. The agent is provided with procedure information through the approaches described in Section II-C. In this setup, the procedures cover IP address allocation for IP PDU sessions,

including IPv4, IPv6, and IPv4v6 cases. The correct toolcall sequence depends on both the request and the outputs of intermediate tool calls. For example, consider an IPv4 address allocation request for a UE. The first step is authorization. If authorization fails, the correct procedure has length k = 1: Pi = (τauth ). If the UE is authorized, the next step checks whether a static IP address is configured. In Scenario A, the evaluated request corresponds to a UE with an available static IP address. Therefore, the correct procedure has length k = 3: Pi = (τauth , τstatic , τregistry ). That is, the agent must authorize the UE, retrieve the static IP address, skip the dynamic IP address allocation steps (i.e., τdhcpv4 and τdhcpv6 ), and finalize the IP assignment. Otherwise, if no static IP address is available, the procedure continues with the appropriate DHCPv4 or DHCPv6 tool before finalization. Thus, for approaches A1, A2, and A3, intermediate reasoning is required because the agent must use each tool output to determine the next action. MCP Server 1 hosts a set of encapsulated tools used in approach A4. Each encapsulated tool corresponds to a particular procedure and deterministically invokes the required lower-level tools that constitute the procedure one after another. In this setup, τE 1 corresponds to the UE IP allocation procedure, while the other encapsulated tools correspond to different procedures. Thus, in A4, the agent no longer has to invoke multiple tools to perform Pi . Instead, it only selects one of the encapsulated tools located at the MCP Server 1 and subsequently invokes it. Note that this implies a single inference step to identify the right tool at the beginning, while in A1, A2, and A3 multiple inference steps during the procedure execution are necessary. The nature of the input intent i varies by approach. For A1, A2, and A4, the intent specifies only the desired outcome

D /DWHQF\'LVWULEXWLRQSHU$SSURDFK

E /DWHQF\'LVWULEXWLRQSHU0RGHO 

/DWHQF\&RVWC(i) V

   $

$

$SSURDFK

$

1XPEHURI5XQV

([HFXWLRQ&RUUHFWQHVV5DWHR(Pi)



   

 

$

F ([HFXWLRQ&RUUHFWQHVV5DWHE\$SSURDFK 0RGHO





 

%

%

G (UURU&RPSRVLWLRQE\$SSURDFK 







 









 







1XPEHURI5XQV

/DWHQF\&RVWC(i) V



0RGHO

%

%

H (UURU&RPSRVLWLRQE\0RGHO 















 



 

$

$

$

$

$SSURDFK



$

$

$

$SSURDFK

$

0RGHO % %



%

%

0RGHO

%

%

(UURU7\SH % %

&RUUHFW([HFXWLRQ 'XSOLFDWH7RRO

3UHPDWXUH6WRS 1R7RRO&DOOV

7RRO2XWVLGH3URFHGXUH

:URQJ3DUDPHWHUV

Fig. 3. Evaluation of the UE IP Allocation procedure (Scenario A). The top row displays the end-to-end latency cost C(i) categorized by (a) approach and (b) model size. The bottom row illustrates (c) the execution correctness rate, representing the average of the reliability metric R(Pi ), alongside the distribution of specific error types across (d) approaches and (e) models. Note that 0.8B, 3B, 9B, and 35B refer to the Qwen-0.8B, Qwen-Coder-3B, Qwen-9B, and Qwen-35B models, respectively.

and the request parameters, such as the target ue_id and session_type, without explicitly providing the tool-call sequence. For A3, the input is explicit, where we assume that a client-side agent associated with the user provides the procedure Pi directly in the request, i.e., the required tools and their order are specified as part of the input. We evaluated four LLMs to assess the impact of model scale and tool-calling capability: • Qwen 3.5: Models evaluated at different parameter scales (0.8B, 9B, and 35B) to study the impact of model size on sequence stability, hereafter referred to as Qwen-0.8B, Qwen-9B, and Qwen-35B. • Qwen3-Coder-Next: A model with advanced tool-calling capability and approximately 3B active parameters per token during inference, hereafter referred to as QwenCoder-3B. To ensure statistical significance, the IPv4 allocation request was executed 50 times per model for each approach, yielding 200 independent runs per approach in Scenario A. The results for Scenario A are summarized in Fig. 3. Latency Cost: Fig. 3(a) shows that latency C(i), defined in (3), varies across approaches primarily due to differences in the number of reasoning turns. Approaches A1 and A3 exhibit similar latency, as both require comparable multi-step reasoning. Approach A2 incurs the highest latency due to the additional procedure retrieval step. In contrast, Approach

A4 achieves the lowest latency by reducing LLM interaction to a single trigger call. From a model perspective, Fig. 3(b) indicates that latency increases with model size, with Qwen35B exhibiting the highest delay, while Qwen-0.8B shows the lowest latency but a wider spread, which is consistent with the unstable executions discussed below. Execution Correctness and Errors: Execution correctness R(Pi ) across approaches and models is shown in Fig. 3(c). For this short procedure (k = 3), Qwen-Coder-3B, Qwen-9B, and Qwen-35B consistently achieve high execution correctness rates, while Qwen-0.8B performs poorly. The distribution of error types further explains these outcomes. Across approaches, Fig. 3(d) shows that A4 achieves the best overall approach-level performance, with the highest number of correct executions and the lowest number of errors. This indicates that encapsulating the procedure inside a deterministic tool reduces step-by-step execution failures. Errors in A4 are mainly due to Duplicate Tool errors, which occur primarily for the Qwen-0.8B model and indicate repeated invocation of the encapsulated tool rather than failures in the internal execution of the procedure. Fig. 3(e) further shows that most errors occur for Qwen-0.8B, which frequently exhibits Duplicate Tool, Premature Stop, and Tool Outside Procedure errors. In contrast, Qwen-Coder-3B achieves the highest number of correct executions and provides the best trade-off between latency and correctness. Qwen-9B and Qwen-35B

/DWHQF\&RVWC(i) V

([HFXWLRQ&RUUHFWQHVV5DWHR(Pi)

D /DWHQF\DQG([HFXWLRQ&RUUHFWQHVV5DWH 











 

%/DWHQF\ %/DWHQF\ %/DWHQF\ %&RUUHFWQHVV %&RUUHFWQHVV %&RUUHFWQHVV





 











3URFHGXUH/HQJWKk 7RRO&DOOV

E (UURU&RPSRVLWLRQE\3URFHGXUH/HQJWKDQG0RGHO



 

1XPEHURI5XQV





 

 

  









(UURU7\SH

&RUUHFW([HFXWLRQ 'XSOLFDWH7RRO 3UHPDWXUH6WRS 1R7RRO&DOOV





 

 

 

















3URFHGXUH/HQJWKk 7RRO&DOOV



%DU2UGHU/HIW %_&HQWHU %_5LJKW % 1RWH2QO\WKH%PRGHOLVSUHVHQWHGIRUOHQJWKVk

30

Fig. 4. Scalability stress test results (Scenario B). Panel (a) illustrates the trade-off between mean latency cost C(i) and execution correctness rate R(Pi ) as the target procedure length k increases. Panel (b) details the composition of procedural errors at varying procedure lengths. Note that 3B, 9B, and 35B refer to the Qwen-Coder-3B, Qwen-9B, and Qwen-35B models, respectively.

0.8B model was excluded from this phase, as the results from Scenario A showed that its capabilities were insufficient for extended logic loops. The results for Scenario B are summarized in Fig. 4. Sequence Scalability: Fig. 4(a) shows how execution correctness rate and latency evolve as the procedure length k increases. Latency increases with sequence length, reflecting the growing number of reasoning steps required for longer procedures. Qwen-9B and Qwen-35B degrade rapidly with increasing k, with execution correctness rates dropping sharply by k = 20. Notably, Qwen-35B exhibits a faster decline than Qwen-9B, indicating that larger parameter size does not necessarily translate to improved robustness in long sequential reasoning tasks. In contrast, Qwen-Coder-3B maintains nearperfect execution correctness up to k = 30. Breaking Point and Failure Analysis: The error composition in Fig. 4(b) explains the observed degradation. As k increases, failures are dominated by Premature Stop errors, indicating that the agent terminates execution before completing the full sequence. This trend is consistent across models but occurs significantly earlier for Qwen-9B and Qwen-35B. Beyond k = 30, even Qwen-Coder-3B, which has advanced tool-calling capability, begins to degrade, with execution correctness dropping substantially by k = 50, indicating a practical upper bound on reliable multi-step execution where maintaining consistency across long interaction histories becomes challenging. IV. C ONCLUSION

also perform well for this short procedure, although Qwen35B shows more errors than Qwen-9B despite its larger size. Finally, we observe that Wrong Order errors and Wrong Tool Name errors do not occur in any of the evaluated configurations. B. Scenario B: Scalability Stress Test Because Scenario A requires only a short tool sequence (k = 3), it does not fully expose the limitations of long multistep reasoning loops. To identify the operational breaking point of these LLMs, we introduced a sequential stress-test setup using MCP Server 3, as shown in Scenario B of Fig. 2. Using approach A1, we expanded the tool pool T inside MCP Server 3 to contain m = 100 network analytics tools (e.g., Average Cell Throughput, AMF Load, and Handover Success Rate). Each tool accepts a target geographic region as input and returns the corresponding KPI. We define the user intent i as a query to analyze the network health of a target region. To fulfill this request, the agent must execute target procedures Pi of varying lengths k ∈ {5, 10, 20, 30, 40, 50}. For a given length k, the agent must sequentially invoke k distinct tools to gather the required KPIs, summarize the collected data, and flag any abnormal metrics. This isolates and stresses the multi-turn reasoning loop Nllm , allowing us to evaluate how well the models maintain sequence stability as the procedure length increases. We performed 30 independent runs per procedure length for each evaluated model. The

In this work, we investigated how LLM-based agents execute network procedures through sequential tool invocations. We compared four execution approaches and showed that approaches relying on iterative agent-side reasoning incur higher latency and are more error-prone, while the toolencapsulated approach achieves lower latency and higher execution correctness by reducing repeated reasoning. Across the evaluations, we also observed that increasing model size alone does not necessarily improve robustness in tool-calling sequences. Finally, the stress-test results showed that the model with advanced tool-calling capability maintains reliable execution over longer procedures than the other evaluated models; however, all models ultimately degrade as the number of sequential tool calls increases, revealing clear breaking points in long procedure execution. As future work, we plan to fine-tune a base model using execution traces labeled with our error taxonomy to evaluate whether this improves long sequential tool execution. We also plan to investigate mechanisms such as agent harness and reusable agent skills to further stabilize complex telecomgrade network procedures. ACKNOWLEDGEMENTS This work has been supported by the European Union’s Horizon Europe MSCA-DN programme through the SCION Project under Grant Agreement No. 101072375.

R EFERENCES [1] H. Navidan et al., “Toward autonomous o-ran: A multi-scale agentic ai framework for real-time network control and management,” 2026, arXiv:2602.14117. [2] E. Bandara et al., “An agentic ai control plane for 6g network slice orchestration, monitoring, and trading,” 2026, arXiv:2602.13227. [3] C. Feng, A. Zhang, G. Min, Y. Huang, T. Q. S. Quek, and X. You, “Towards 6g native-ai edge networks: A semantic-aware and agentic intelligence paradigm,” 2025, arXiv:2512.04405. [4] Y. Han, H. Ko, N. Ko, T. Taleb, and Y. Chen, “Toward e2e intelligence in 6g networks: An ai agent-based ran-cn converged intelligence framework,” 2026. [Online]. Available: https://arxiv.org/abs/2602.23623 [5] G. Jiang, K. Wang, X. Chen, and Y. Huang, “Agentic ai empowered intent-based networking for 6g,” 2026, arXiv:2601.06640. [6] Y. Xiao et al., “Sanet: A semantic-aware agentic ai networking framework for cross-layer optimization in 6g,” 2025, arXiv:2512.22579. [7] K. Zhu et al., “Where llm agents fail and how they can learn from failures,” 2025, arXiv:2509.25370. [8] M. B. Shah, M. M. Morovati, M. M. Rahman, and F. Khomh, “Characterizing faults in agentic ai: A taxonomy of types, symptoms, and root causes,” 2026, arXiv:2603.06847. [9] C. Winston and R. Just, “A taxonomy of failures in tool-augmented llms,” in 2025 IEEE/ACM International Conference on Automation of Software Test (AST), 2025, pp. 125–135. [10] 3GPP, “System architecture for the 5G System (5GS),” Technical Specification (TS) 23.501 V19.7.0, 2026. [11] ——, “Non-Access-Stratum (NAS) protocol for the 5G System (5GS); Stage 3,” Technical Specification (TS) 24.501 V19.6.0, 2026.

Record · ID 155196 · SHA-256 677d8d3d765a4e09
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.