Beyond State Machines: Executing Network Procedures with Agentic Tool-Calling Sequences Purna Sai Garigipati∗† , Onur Ayan∗ , Kishor Chandra Joshi† , Xueli An∗ ∗ Heisenberg Research Center, Huawei Technologies Duesseldorf GmbH, 80992 Munich, Germany
Email: {purna.sai.garigipati,onur.ayan,xueli.an}@huawei.com † Eindhoven University of Technology, Eindhoven, The Netherlands
arXiv:2605.02584v1 [cs.NI] 4 May 2026
Email: {k.c.joshi}@tue.nl
Abstract—Agentic AI will be an essential enabling technology for designing future mobile communication systems, which could provide flexible and customized services, automate complex network operations, and drive autonomous decision-making across the network. This work studies how Large Language Model (LLM)-based network AI agents can be utilized to execute network procedures expressed as sequences of tool invocations. We investigate four approaches, which differ in how the agent obtains the procedure and in how execution is distributed between the agent and the underlying tools. We evaluated the latency and execution correctness across these approaches using a User Equipment (UE) IP allocation procedure as a case study. Furthermore, we conduct a stress test to examine how many sequential procedural steps an LLM agent can reliably execute before failure. Our results show that approaches relying on iterative agent-side reasoning incur higher latency and are more prone to execution errors, while approaches where the procedure is encapsulated within a single tool, which internally orchestrates the required steps by invoking other tools, reduce latency by limiting repeated reasoning. The stress-test results further show that the model with advanced tool-calling capability maintains reliable execution over longer procedures than the other evaluated models; however, all models exhibit reliability degradation as procedure length increases, revealing clear execution limits in multi-step tool-based workflows. To systematically analyze failures in procedure execution, we introduce a procedurespecific error taxonomy that categorizes deviations in multi-step procedural execution. Index Terms—Large Language Model (LLM), Agentic AI, Mobile Communication Networks, Procedure Execution
I. I NTRODUCTION Agents empowered by Large Language Models (LLMs) have introduced a new paradigm of autonomous systems capable of reasoning, planning, and interacting with external tools to accomplish complex tasks. Such systems extend beyond single-shot inference and operate through iterative decisionmaking and tool invocation. This paradigm shift has gained significant attention because it enables complex workflows to be executed without explicitly hard-coded control logic, instead relying on the model inference to determine the sequence of actions required to achieve a given objective. Recent work shows that agentic approaches are being actively explored in the context of next-generation networks, particularly in 6G. Several studies focus on the Radio Access Network (RAN), where agentic frameworks have been proposed for real-time control, resource management, and op-
timization [1]–[3]. In addition, recent efforts consider end-toend intelligence across both RAN and core networks, integrating monitoring, policy control, and cross-layer optimization using agentic approaches [4]–[6]. Collectively, these works indicate a shift toward agent-driven network automation. In this context, we study the use of LLM-based agents to execute network procedures. Typically, a network procedure consists of a sequence of dependent operations that must be executed in a strict order to achieve a target system state. Traditionally, such procedures are implemented as scripts or workflows, which require explicit development, testing, and deployment. Supporting variability in network conditions often requires extensive conditional logic, making these implementations difficult to scale and maintain. Moreover, such implementations are tightly coupled to predefined states and can fail when the observed network state deviates from expected conditions. In contrast, an agent-based approach enables procedures to be specified at a higher level using natural language descriptions. Given a set of standardized tools, different procedures can be dynamically composed by varying the sequence of tool invocations based on the task and intermediate outcomes, without requiring explicit reprogramming. This flexibility is important in scenarios where a network operator or an application needs to execute a new procedure that is not predefined in existing specifications. This raises a fundamental question: can an LLM-based agent reliably execute telecom-grade procedures? To answer this, we study the agent’s ability to perform sequential tool invocation, focusing on whether the correct step ordering is maintained, how the execution behaves across repeated runs, what types of procedural violations occur, and how the reliability changes as the length of the procedure increases. We further analyze different mechanisms for providing the procedure to the agent, reflecting realistic deployment scenarios. Although agentic systems enable flexible execution, previous work shows that LLM-based agents are prone to failures during multi-step reasoning and tool interaction, and existing studies have proposed taxonomies to characterize such failures, covering aspects such as planning errors, tool invocation issues, and execution inconsistencies [7]–[9]. However, these taxonomies are designed for open-ended or generalized domains and do not address the strict constraints of sequential
tool execution. We therefore define a procedure-specific error taxonomy tailored to this setting, enabling precise analysis of procedural execution correctness for agentic systems. The main contributions of this paper are summarized as follows: • Procedural Execution Approaches: We present four approaches for delivering procedural logic to LLM agents and characterize how procedure placement impacts endto-end latency and execution correctness. • Scalability Limits of Sequential Tool Execution: We identify the execution limits of LLM agents by showing how performance degrades as the number of sequential steps increases. • Procedure-Specific Error Taxonomy: We define an error taxonomy tailored to sequential tool execution, enabling precise classification of failures in multi-step agentic workflows. The remainder of this paper is organized as follows. Section II formulates the procedural execution model, defines the evaluation metrics, and presents the execution approaches and error taxonomy. Section III describes the experimental scenarios and presents the evaluation results. Section IV concludes the paper. II. M ETHODOLOGY This section presents the procedural execution model considered in this work, the performance metrics used for evaluation, the four execution approaches, and the error taxonomy adopted for analyzing failures. A. Definition of a Procedure as Tool Sequence We consider a task execution setting in which an LLMbased agent interacts with a tool server that exposes a total number of m tools, denoted by T = {τ1 , τ2 , . . . , τm }. Given a user intent i, the agent must identify and execute the matching procedure Pi with Pi ∈ P where P denotes the set of procedures available at the agent. We define procedure Pi as an ordered sequence of tool calls, where each required tool τi,j belongs to the available toolset T : Pi = (τi,1 , τi,2 , . . . , τi,k ),
(1)
Here, the index i represents the procedure that matches the user’s intent from a set of possible procedures P. The second index represents the step number within the procedure. For example, τi,1 is the first tool occurring in the procedure Pi followed by τi,2 up to the last tool τi,k . The observed execution produced by the agent is similarly represented as an ordered sequence of tool calls: Oi = (τ̂i,1 , τ̂i,2 , . . . , τ̂i,k̂ ),
(2)
where τ̂i,j denotes the tool invoked at step j during execution of procedure i, and k̂ is the number of executed steps. In this formulation, Pi represents the correct procedure (i.e., ground-truth), while Oi denotes the sequence of tool calls executed by the agent.
B. Evaluation Metrics We evaluate procedural execution using latency and execution correctness. The total latency cost C(i) is defined as the end-to-end execution time required to process intent i, including both the LLM reasoning and tool invocation: C(i) =
Nllm X j=1
Lllm j +
k̂ X
Ltool j ,
(3)
j=1
where Nllm is the number of LLM reasoning steps, k̂ is the tool number of tool invocations executed, and Lllm denote j and Lj the latency of the corresponding reasoning and tool-execution step, respectively.1 Execution correctness is measured using a binary reliability metric R(Pi ). For a single run, it evaluates whether the observed sequence of actions Oi perfectly matches the expected procedure Pi : ( 0, if k ̸= k̂ or τi,j ̸= τ̂i,j for at least one j (4) R(Pi ) = 1, otherwise By this definition, a score of 0 indicates that the observed sequence Oi differs from Pi in length (k ̸= k̂) or composition (τi,j ̸= τ̂i,j ), which implies an erroneous procedure execution as categorized later in Section II-D. Averaging R(Pi ) over repeated runs yields the execution correctness rate for a given model and approach. C. Procedural Execution Approaches The four approaches considered in this work are shown in Fig. 1. They share the same basic agent–tool interaction setting, but differ in where the procedure is defined and how the sequential execution is carried out, particularly in the distinction between iterative agent-driven execution and toolencapsulated execution. • A1: Agent-Embedded Procedure: The set of procedures P is available to the agent as its system prompt. After receiving intent i, the agent must first parse the prompt to identify and extract the matching procedure Pi ∈ P. Once identified, the agent reasons over this specific sequence and invokes the required tools step by step. For a procedure of length k, this leads to approximately Nllm ≈ k + 1 reasoning steps,2 where the final step corresponds to response summarization. • A2: Server-Provided Procedure: The agent first retrieves the procedure information P from an external database or repository, parses the specific procedure Pi for the intent, and then executes it sequentially in the same manner as A1. This additional retrieval step increases the number of reasoning steps to Nllm ≈ k + 2. • A3: User/Agent-Provided Procedure: The specific procedure Pi is explicitly specified within the incoming 1 Transmission latency is negligible and is therefore excluded from the latency cost equation. 2 Values are approximate as they represent the optimal execution path; in practice, model errors or tool failures often cause the agent to deviate, leading to additional reasoning turns or early termination.
(a)
(b)
User/Agent
LLM Agent
MCP Server
System prompt 𝓅
User/Agent
LLM Agent Fetch procedure 𝓅 Return 𝓅
Send intent 𝑖
Send intent 𝑖 Parsing procedure 𝑃
A1
𝑁 ≈𝑘+1 Reasoning turns
Send summarized response
Procedure DB
MCP Server
Parsing procedure 𝑃
START LOOP For step 𝑗 = 1 TO 𝑘 Invoke Tool 𝜏̂ ,
𝑁 ≈ 𝑘+2 Reasoning turns
START LOOP For step 𝑗 = 1 TO 𝑘 Invoke Tool 𝜏̂ ,
Return result 𝑟 ,
Return result 𝑟 ,
Agent processes feedback 𝑟 , for next turn END LOOP
Send summarized response
Agent processes feedback 𝑟 , for next turn END LOOP
A2
(c)
(d)
User/Agent
LLM Agent
Send intent 𝑖 with procedure 𝑃 = (𝜏 , , 𝜏 , , …, 𝜏 , )
MCP Server
User/Agent Send intent 𝑖
Parsing procedure 𝑃 START LOOP For step 𝑗 = 1 TO 𝑘
A3
LLM Agent
𝑁 ≈ 𝑘+1 Reasoning turns
Invoke Tool 𝜏̂ ,
Send summarized response
Agent processes feedback 𝑟 , for next turn END LOOP
A4
MCP Server Send single call for Encapsulated Tool 𝜏
𝑁 ≈2 Reasoning turns (trigger call+summary)
Return result 𝑟 ,
Tool 𝜏 internally executes procedure 𝑃 Return final result of 𝑃
Send summarized response
Fig. 1. Comparison of four procedural execution approaches. (a) A1 embeds the procedure within the agent, (b) A2 retrieves the procedure from an external database, (c) A3 receives the procedure in the input prompt, and (d) A4 encapsulates the procedure within a single tool. The figure highlights the difference between iterative multi-step execution (A1–A3) and single-call execution (A4).
prompt, which may originate from a user or another agent. The LLM agent parses the sequence and executes it iteratively, resulting in a reasoning overhead comparable to A1, i.e., Nllm ≈ k + 1. • A4: Tool-Encapsulated Procedure: The entire logic of Pi is encapsulated inside a single tool (τE ). After receiving the user intent, the agent selects and invokes this tool once, and the internal procedure is executed within the tool implementation. Thus, the number of LLM reasoning steps is reduced to approximately Nllm ≈ 2, corresponding to one trigger call and one final summarization step. This shifts execution complexity from the LLM to deterministic tool logic. D. Error Taxonomy for Tool-Based Procedures Whenever a run fails, i.e., when R(Pi ) = 0, the observed deviation is categorized to identify the failure mode. We adopt a strict error taxonomy because exact procedural execution is essential in structured network tasks. 1) Wrong Tool: A Wrong Tool error occurs when the agent fails to correctly invoke the required tool for a given step in the procedure. This can manifest in three forms: • Tool Outside Procedure: The agent invokes a tool τ̂i,j that does not appear in Pi , i.e., a tool that does not belong to the intended procedure. • Wrong Tool Name: The agent intends to use the correct tool but invokes it using an incorrect or hallucinated name.
Wrong Parameters: The agent invokes the correct tool but provides incorrect, missing, or invalid input arguments. This corresponds to incorrect tool invocation and represents a critical deviation, as it introduces actions outside the intended execution logic. Due to this severity, any execution containing both Wrong Tool and Duplicate Tool errors is strictly classified as a Wrong Tool error. 2) Duplicate Tool: A Duplicate Tool error is recorded when a valid tool τ ∈ Pi is invoked multiple times unnecessarily within Oi , without progressing the execution state. This typically reflects a reasoning loop and leads to an increase in both the number of reasoning steps Nllm and, subsequently, the latency cost C(i). 3) Premature Stop: A Premature Stop error occurs when the execution terminates before completing all k steps of the procedure, even though the executed tools are in the correct order up to the stopping point. In other words, Oi is a proper prefix of Pi with k̂ < k. 4) Wrong Order: A Wrong Order error is recorded when all required tools are present in Oi , but the sequence does not match the prescribed ordering of Pi , i.e., Oi ̸= Pi despite containing the same elements. 5) No Tool Calls: A No Tool Calls error is defined when the observed sequence is empty, i.e., Oi = ∅, indicating that no tool invocation was performed despite the requirement for procedural execution. •
User/Agent
LLM Agent Request
Scenario A: UE IP Allocation Encapsulated tools
Invoke Tool Return result
IP allocation tools
Scenario B: Scalability Stress Test Analytics tools
𝜏
𝜏
𝜏
𝜏
𝜏
𝜏
𝜏
𝜏
𝜏
𝜏
𝜏
𝜏
𝜏
MCP Server 2
MCP Server 3
Response MCP Server 1
……
𝜏
Fig. 2. Overview of the experimental setups. Scenario A illustrates the UE IP Allocation workflow across two MCP servers: MCP Server 2 provides the IP allocation tools used in approaches A1–A3, while MCP Server 1 hosts encapsulated tools used in approach A4. The highlighted tools in Scenario A indicate the subset of tools executed for the representative request. Scenario B utilizes MCP Server 3, providing a pool of m = 100 network analytics tools to stress-test sequence stability for procedures of length up to k = 50.
III. E XPERIMENTAL E VALUATION To evaluate the performance and stability of LLM-based procedure execution, we designed two experimental scenarios, as shown in Fig. 2. Scenario A compares the four approaches introduced in Section II-C, while Scenario B serves as a stress test to evaluate how the correctness of sequential tool execution is affected by the target procedure length k. A. Scenario A: UE IP Allocation Scenario A simulates a simplified User Equipment (UE) IP allocation procedure. At a high level, the procedure first authorizes the UE and the requested session type, then checks whether a static IP is already assigned. If no static IP for the given UE is found, the network allocates an IP address using the appropriate DHCP server. Finally, the assigned IP is recorded in the registry. We adapted the logical guidelines of 3GPP standards [10], [11] into a set of executable network tools, where the orchestration procedure is executed through LLM-driven tool invocation. As shown in Scenario A of Fig. 2, MCP Server 2 hosts the IP allocation tools used directly in approaches A1, A2, and A3, where the agent executes the procedure step-by-step. These tools are defined as follows: 1) UE Authorization (τauth ): Validates the UE ID and the requested session type. 2) Static IP Retrieval (τstatic ): Checks whether a static IP is pre-assigned for the UE. 3) Dynamic IPv4 Allocation (τdhcpv4 ): Dynamically allocates an IPv4 address. 4) Dynamic IPv6 Allocation (τdhcpv6 ): Dynamically allocates an IPv6 address. 5) IP Assignment (τregistry ): Finalizes IP assignment by notifying the UE and recording the allocation in the network registry. The agent is provided with procedure information through the approaches described in Section II-C. In this setup, the procedures cover IP address allocation for IP PDU sessions,
including IPv4, IPv6, and IPv4v6 cases. The correct toolcall sequence depends on both the request and the outputs of intermediate tool calls. For example, consider an IPv4 address allocation request for a UE. The first step is authorization. If authorization fails, the correct procedure has length k = 1: Pi = (τauth ). If the UE is authorized, the next step checks whether a static IP address is configured. In Scenario A, the evaluated request corresponds to a UE with an available static IP address. Therefore, the correct procedure has length k = 3: Pi = (τauth , τstatic , τregistry ). That is, the agent must authorize the UE, retrieve the static IP address, skip the dynamic IP address allocation steps (i.e., τdhcpv4 and τdhcpv6 ), and finalize the IP assignment. Otherwise, if no static IP address is available, the procedure continues with the appropriate DHCPv4 or DHCPv6 tool before finalization. Thus, for approaches A1, A2, and A3, intermediate reasoning is required because the agent must use each tool output to determine the next action. MCP Server 1 hosts a set of encapsulated tools used in approach A4. Each encapsulated tool corresponds to a particular procedure and deterministically invokes the required lower-level tools that constitute the procedure one after another. In this setup, τE 1 corresponds to the UE IP allocation procedure, while the other encapsulated tools correspond to different procedures. Thus, in A4, the agent no longer has to invoke multiple tools to perform Pi . Instead, it only selects one of the encapsulated tools located at the MCP Server 1 and subsequently invokes it. Note that this implies a single inference step to identify the right tool at the beginning, while in A1, A2, and A3 multiple inference steps during the procedure execution are necessary. The nature of the input intent i varies by approach. For A1, A2, and A4, the intent specifies only the desired outcome