Research Article
1
Toward Agentic Optical Networks: A Vision of LLM Agent-Driven Autonomous Lifecycle Management YAO Z HANG1,+ , S HENGNAN L I1,+ , Y UCHEN S ONG1 , Y IDI WANG1 , Y UE PANG2 , W ENBIN C HEN1 , X IAOTIAN J IANG3 , X IAO L UO1 , M EIXIA F U4 , M IN Z HANG1 , YONGLI Z HAO1 , S HANGUO H UANG1 , A LAN PAK TAO L AU5 , AND DANSHI WANG1,*
arXiv:2609.32226v1 [cs.NI] 26 Sep 2026
1 State Key Laboratory of Information Photonic and Optical Communication, Beijing University of Posts and Telecommunications, Beijing, 100876, China 2 China Telecom Cloud Network Operating System R&D Center, Beijing, 102299, China 3 State Key Laboratory of Optical Fiber and Cable Manufacture Technology, China Telecom Research Institute, Beijing, China 4 School of Automation and Electrical Engineering, Institute of Industrial Internet, University of Science and Technology Beijing, Beijing, 100083, China 5 Photonics Research Institute, Department of Electrical Engineering, The Hong Kong Polytechnic University, Kowloon, Hong Kong, SAR, China + These authors contributed equally to this work. * [email protected]
Compiled September 29, 2026
As optical networks continue to expand in scale, complexity, and service diversity, the implementation of automation has become essential for ensuring agility, efficiency, and reliability in lifecycle management (LCM) of optical networks. Large language model (LLM) Agent, distinguished by its progressively sophisticated capabilities in logical reasoning, adaptive decision-making, complex problem solving, and multi-task orchestration, presents great opportunities to advance network automation beyond traditional AI techniques. Nevertheless, the application of LLM Agent in optical networks remains in its early exploratory stage, challenged by the lack of multi-task coordination, high computational demands, data dependence, and reliability concerns. In this paper, we envision a conceptual roadmap toward Agentic Optical Networks (AONs) by integrating LLM Agents throughout the LCM with high-level autonomy. First, we trace the evolution from manual operations to AI-empowered frameworks and distill key technologies in Agent, providing actionable insights into leveraging its strengths for addressing practical network automation challenges. A core contribution of this paper is the proposal of a hierarchical multi-Agent framework, which is specifically developed to manage every phase in LCM of AONs, including planning, deployment, operation, maintenance, upgrade, and decommission, thereby enabling more cohesive and comprehensive automation throughout the entire lifecycle. In addition, future directions and underlying challenges are also discussed at the intersection of LLM and optical networks. By aligning the LLM Agent with the specialized requirements of AONs, this work aims to explore the potential for the evolution of optical networks moving from task-level semi-automatic execution toward lifecycle-level full autonomy. http://dx.doi.org/10.1364/ao.XX.XXXXXX
1. INTRODUCTION The lifecycle management (LCM) of optical networks represents a systematic framework to govern the entire lifespan of optical infrastructure, involving six critical phases: planning, deployment, and operation, maintenance, upgrade, and decommissioning [1]. Serving as the backbone of modern telecommunications, optical networks are engineered as large-scale complex systems, consisting of cross-connected optical fibers, expanding nodes, and massive optical components. Traditional LCM methodologies, constrained by manual interventions and reactive operations, struggle to address complex LCM of modern optical networks, which requires proactive and intelligent governance [2]. Pro-
pelled by the breakthroughs in artificial intelligence (AI), the paradigm is currently evolving toward autonomous LCM [3] characterized by closed-loop autonomy. Throughout the evolution of optical networks, AI techniques have progressively contributed distinct capabilities across different levels of autonomy. Since the 2010s, basic machine learning (ML) methods have facilitated data-driven and efficient solutions for tasks primarily involving classification and clustering [4]. Meanwhile, advanced deep learning (DL) techniques have gained significant attention in optical networks [5, 6] owing to their powerful learning capabilities, enabling the high-precision modeling of complex nonlinear relationships by mapping input
Research Article
data to desired outputs [7]. However, both ML and DL models are typically tailored for specific tasks and exhibit limited generalization across diverse problem domains. Specifically, it is necessary to retrain or fine-tune separate models for different applications, such as performance prediction, parameter identification, and configuration optimization. Furthermore, these models generally lack the ability to comprehend or analyze the broader network context, which is an essential cognitive function required for adaptive decision-making [8]. As a result, these techniques facilitate foundational levels of automation in optical networks but fall short of supporting high-level automation. Advancing toward fully autonomous optical networks necessitates the integration of more advanced AI techniques with general intelligence, thereby enabling models to understand, reason, and make decisions based on a comprehensive view of network state and operational intent. In recent years, generative AI (GenAI) has experienced rapid advancements [9] marked by the emergence of large language models (LLMs) that exhibit remarkable capabilities in generating coherent text, image, code, and even structured solutions. With powerful reasoning and cognitive capabilities, LLM can handle a wide range of tasks through prompt-based interaction and in-context learning, thereby creating innovative opportunities for AI applications in specialized domains [10]. In optical networks, initial efforts have explored LLM Agents to solve multiple tasks [11–14], including performance optimization, fault diagnosis, resource orchestration, and device control. By leveraging their reasoning and cognitive capabilities, LLM Agents aim to overcome the fragmentation and scalability limitations of conventional AI solutions, thereby facilitating more integrated, autonomous, and adaptive capabilities for future Agentic Optical Networks (AONs). Predictably, building upon LLM capabilities, the paradigm of optical network is evolving toward Agent-based systems. Nevertheless, the application of LLM Agent in optical networks is still in its early exploratory stage. Several key challenges remain to be addressed before its potential can be fully harnessed across the LCM process [15, 16]. First, the deployment of Agents faces practical constraints, including the high computational cost, large model sizes, and substantial token consumption, which limit their scalability and adaptability in real-world scenarios. Second, while techniques such as prompt engineering, context engineering, and harness engineering have shown promise in adapting to specific tasks, a universally efficient methodology for optical network’s LCM has yet to be established. Third, given the diverse and complex tasks throughout the entire lifecycle of optical networks, it is difficult for a single monolithic model to fully meet all functional and operational requirements. Therefore, there is a pressing need for a multi-agent collaborative framework to support modular deployment and facilitates coordinated execution across diverse tasks. Furthermore, addressing issues such as hallucination and reliability is essential to ensure safe and trustworthy automation in network operations. These challenges highlight the necessity of developing Agentic frameworks and elaborating collaborative workflows that are specifically tailored to the LCM of AONs. In this paper, we present a comprehensive roadmap for integrating LLM Agent into the full lifecycle of AONs. The main contributions are as follows: 1) From a network-centric perspective, we review the application of LLMs in optical networks and summarize representative related works, highlighting the potential of LLM Agents to enhance reasoning, decision-making, and automation capabilities. 2) We introduce several core technolo-
2
gies in LLM-driven Agent development, providing the optical network community with guidance on leveraging these capabilities for automation tasks. 3) As the main contribution, a hierarchical multi-Agent framework is proposed to implement the autonomous LCM of AONs, including network planning, deployment, operation, maintenance, upgrade, and decommission phases. This organization extends Agent-based automation from individual operational and management tasks toward the full network lifecycle, while providing clear task ownership and modular scalability as network functions and lifecycle requirements evolve. 4) Finally, several key challenges at the intersection of LLM Agents and optical networks are outlined. By systematically aligning LLM technologies with the specific demands of AONs, we provide a vision of Agent-enhanced automation throughout the entire lifecycle, paving the way toward fully self-driving optical networks that will underpin the next generation of intelligent connectivity.
2. APPLICATIONS OF LLMS IN OPTICAL NETWORKS Currently, LLMs, as a representative form of GenAI, mark a significant leap forward in natural language processing (NLP), transforming how the language is understood and generated. The rapid advancement of LLM technologies also presents promising opportunities to automate numerous tasks in the field of optical networks [11]. Recent efforts to integrate LLMs into optical networks, as illustrated in Table 1, can be broadly categorized into three progressive stages. In the initial phase, research focused on leveraging LLMs’ capabilities in context understanding and generation to assist in basic log and alarm processing tasks. For instance, AlarmGPT [19] and instruction-tuned LLaMA [20] demonstrated how LLMs could automatically summarize alarms, parse logs, and generate structured outputs, thereby reducing reliance on manual intervention. These approaches primarily treated LLM as intelligent text processors embedded in optical network toolchains. As exploration deepened, researchers began incorporating domainspecific knowledge into prompts and utilizing LLMs for workflow decomposition and task chaining. For example, in quality of transmission (QoT) estimation and network planning tasks, LLMs were employed to involve physical models, analyze simulation results, suggest parameter values, in a loop-like manner [21, 26]. Fault management tasks also benefited from LLMs’ ability to reason across symptoms and suggest candidate root causes [22]. More recently, with the growing demand for autonomous operation in optical networks, LLMs have begun to take on Agent-like roles. Instead of merely assisting with textual processing or isolated task reasoning, they are increasingly integrated into broader network systems, enabling interaction with realtime data, tools, and control environments [12]. One direction involves coupling LLMs with digital twin (DT) platforms, allowing them to perceive dynamic network states and participate in closed-loop decision-making. For instance, Sun et al. [13] explored an AI Agent for optical network management, where the LLM interprets simulation outputs and responses operational adjustments. Similarly, Liu et al. [14] demonstrated how LLM-powered Agent could continuously monitor network and adapt responses. These efforts underscore a transition from data interpretation to proactive execution and feedback for complex tasks. Another emerging trend is the integration of LLMs with network control and management interfaces, such as software-
Research Article
3
• Background • Motivations • Contributions
Section 1 Introduction
Section 3 • •
Word embedding Transformer
⑨ AI Agent • •
Technical Foundations of LLM Agent
• •
• •
① - ③ The construction of LLMs
Prompt elements Prompt techniques
⑤ Reinforcement
⑥ RAG • •
PEFT: LoRA/QLoRA Domain dataset
• •
Knowledge Compression
RLHF Reward model
④ - ⑨ Key technologies for domain enhancement
LLM Agent for Optical Network Full-lifecycle Management
Fine-tune LoRA
Prompt
Network Director (Highest level)
• •
Close/Open-source Applicability
⑦ Prompt engineering
Tools System design
Section 4 & 5
• •
Next-token prediction Zero-shot/Few-shot
⑧ AI workflow
Harness With DT
④ Fine-tuning
③ Model size
② Model pre-training
① Model architecture • •
Section 2 Applications of LLM in optical networks (Related Works)
RAG
Pre-train
Digital twin of optical network (DTON)
Chain of thought
Agentic AI
Administrative Agents
Data base
Division Agents (for each lifecycle)
Knowledge library
AI Experts (in each division)
A O N
Tool kit
...
Hierarchical multi-agent framework for LCM
① Planning • • • •
Requirement analysis Topology design Resource selection ① Planning Route planning
② Deployment • • • •
② Deployment
Configuration Initialization Build DTON QoT estimation
③ Operation
⑥ Decommission
④ Maintenance
⑤ Upgrade
• • • • DC
Monitoring Inspection QoT assurance Dynamic operation
DC
Access network
Metro network
• • • •
Core network
Alarm definition Failure diagnose Solve and repair Report generation
⑤ Upgrade
Full-lifecycle of optical network
③ Operation • • • •
④ Maintenance
DC
Demand analysis Upgrade strategy Optimization Validation
⑥ Decommission • • • •
Report confirm Removal/Retention Service switch Physical operation
Physical-layer optical network Section 6 Outlook
• Hallucination risks • Model selection and deployment
• Token costs • Long-horizon memory
Fig. 1. A comprehensive overview of the structure and key topics covered in this work.
defined network (SDN) controllers or network management system (NMS). This includes efforts to map user intents to structured network commands via programmable interfaces [25, 36], as well as architectures that integrate LLMs directly into service orchestration pipelines and configuration workflows [23, 24]. Further developments have demonstrated frameworks for endto-end task execution [28], where LLM Agent not only interpret operational goals but also coordinate resource allocation and execute network updates in response to environmental changes. Building upon these single-Agent foundations, recent breakthroughs have increasingly shifted towards deploying locallyhosted, open-source LLMs to address data privacy and latency concerns in real-world scenarios. For example, field trials have demonstrated AI Agents powered by local open-source models achieving full LCM in elastic optical networks, executing autonomous service provisioning and failure recovery within minutes [30]. Moreover, structured human expertise has been
explicitly integrated into the prompts [31, 32] to enhance domain adaptability and reliability of these Agents. As network scenarios become more complex, the paradigm is rapidly evolving from single-Agent assistants to collaborative multi-Agent systems. In these architectures, complex workflows are decomposed and assigned to specialized Agents that communicate and collaborate to prevent the accumulation of errors [34]. Similarly, multi-Agent frameworks have been applied to specific component-level tuning, such as autonomously control Raman amplifiers [35], or for strategy generation with DT interaction [33]. These multi-Agent strategies, closely aligned with the vision of fully autonomous, intent-driven networks [37, 38], represent a transformative leap toward self-operating infrastructures. Overall, these explorations highlight the feasibility of applying LLMs to various network tasks, and showcase the potential of Agent to enhance efficiency, reduce human workload, and
Research Article
4
Table 1. Representative works integrating LLMs with optical networks Applications Framework of LLM-driven optical network
Capabilities and scope
Method Type
Ref
LLM applications, challenges, and opportunities for intelligent network operations, prediction, and analysis
GPT-based prompt
[11]
Capabilities and limitations of applying LLMs to optical networks
Conceptual
[15–18]
GPT-based prompt
[19]
Alarm/Log analysis
LLM-enabled professional alarm/log Q&A, analysis, and reasoning
QoT estimation Failure management
LLaMA fine-tune
[20]
LLM-assisted QoT estimation and DT-based multi-task network management
GPT-based prompt
[21]
LLM-based fault analysis and management fused multi-mode data
GPT-based prompt
[22]
Prompt
[23, 24]
Configuration and control
Practical pipeline for automated configuration in SDN network and testbeds
Simulation assistant
System simulation and performance evaluation, with natural language
Network tomation
operations
Qwen fine-tune
[25]
GPT-based prompt
[26]
au- LLM-assisted network operation workflows with task automation and tool interac- LLaMA fine-tune tion
Field-trial demonstration
Field-trials for LLM-assisted lifecycle management with interaction with network management and operational tools
Performance optimization
Integrates human expertise to enhance efficiency and reliability in AI Agent
Autonomous multi-task col- Multi-Agent collaboration for task decomposition, coordination, and tool calling laboration
support decision-making in complex, dynamic optical systems. From the perspective of lifecycle coverage, existing studies span individual network functions, selected operational stages, and increasingly multiple lifecycle phases, including recent fieldtrial demonstrations. In terms of Agent organization and task orchestration, existing approaches range from single LLM-based assistants and task-specific Agents to multi-Agent schemes for task decomposition, coordination, and collaborative execution. Regarding system interaction, prior studies explored integration of LLMs and Agents with DTs, NMSs, and domain-specific tools to different extents. These efforts should be viewed as complementary steps in the evolution of LLM-enabled optical network automation, collectively illustrating a progression from task-level assistance toward increasingly coordinated and autonomous network management. From this perspective, rather than replacing existing taskspecific or workflow-oriented approaches, this work adopts a lifecycle-oriented perspective and organizes Agent capabilities around the major phases of an optical network. We aim to synthesize these developments from a system-level architectural perspective and to further explore how heterogeneous Agent capabilities can be organized with network management systems and domain-specific tools to progress from task-level LLM assistance toward full-lifecycle autonomous management. Moving forward, further improvements can be achieved by leveraging advanced techniques such as instruction tuning, reinforcement learning (RL), fine-tuning, tool augmentation, retrievalaugmented generation (RAG), and harness engineering, all of which are discussed in the following sections. These techniques offer the promise of building more robust, reliable, and domainaligned LLM Agents tailored to the unique demands of optical networks.
3. TECHNICAL FOUNDATIONS OF LLM AGENT Building upon the recent advances in LLMs, this section delves into the underlying technical foundations that equip LLM to be Agent to address various complex tasks. We focus on the core architectures, particularly transformer, as well as the mechanisms that enable LLMs to generalize across domains. This technical overview provides the necessary background for understanding
Prompt, RAG
[13, 27, 28] [12, 14, 29, 30]
AI Agent, GPT-4o
[31, 32]
Multi-Agent framework
[33–35]
how LLMs can be adapted and applied to autonomous optical network management in subsequent sections. A. The construction of LLM
Following the seminal work by Vaswani et al. [39], Transformerbased models have become central to modern NLP and beyond due to their scalable and context-aware modeling capability. In NLP, input text is first divided into discrete units called tokens, which may correspond to words, subwords, or characters depending on the tokenizer. Each token is then mapped to a dense numerical vector through an embedding layer, providing a continuous representation that can be processed by the neural network. The core innovation of Transformer, self-attention, enables each token to dynamically evaluate the relevance of other tokens in the same input sequence and aggregate their information into a context-aware representation. While multi-head attention performs multiple attention operations in parallel, allowing the model to capture different semantic and structural patterns simultaneously [40, 41]. The performance and reliability of LLMs are fundamentally shaped by the data trained on, which must balance linguistic diversity and domain relevance [42]. Large-scale general text sources support broad language understanding, while domain-specific resources, such as optical network logs, standards, protocols, and technical literature, improve reasoning in specialized contexts [43]. Data preprocessing, including quality filtering, deduplication, and sensitive information removal, further enhances representation quality and model alignment with practical usage [44, 45]. Pre-training enables LLMs to learn linguistic structure and reasoning patterns from large unlabeled corpora. Most stateof-the-art LLMs, including GPT [46] and LLaMA [47], adopt autoregressive decoder-only architectures that predict tokens sequentially. As representations propagate through successive layers, they accumulate increasingly contextual information [48]. Importantly, well-pretrained models often exhibit strong zeroshot and few-shot learning abilities [49]. In zero-shot learning, the model performs a task without task-specific examples being provided in the prompt, relying on it pretrained knowledge and the task instruction. In few-shot learning, a small number of task examples are provided in the prompt to guide the model toward
Research Article
5
(a) Fine-tune for domain-specific conditions
(b) Reinforcement learning in environment policy model (trainable)
query
Domain data/info
Fine-tune
scores
A1 A2 A3 ...
(d)
Prompt engineering
(e)
KL
outputs
retrieval
query reward model (frozen)
reference model (frozen)
group computation
user
r1 r2 r3 ...
AI workflow using toolkit
prompt
(c) RAG using domain knowledge library
LLM2
Tool1
Exit
LLM3
Tool2
LLM5
LLM4
Exit
LLM Knowledge library
answer
(f) Harness engineering-enhanced AI Agent Context injection prompts, memory, skills, conversation
instruction
input data output indicator
LLM1
CoT deep reasoning
context
Domain toolkit
Out
Action
Control compaction, orchestration, iteration loops
Model Reasons → decides writes
reads
Persist Harness
filesystem, git, progress files
calls tools, MCPs feedback results
Observe & Verify logs, test results
Fig. 2. Key technologies for domain enhancement. (a) Fine-tune (b) Reinforcement learning (c) RAG (d) Prompt engineer (e) AI workflow (f) Harness engineering-enhanced AI Agents.
the desired behavior, without updating its parameters. These capabilities allow a single pretrained model to adapt to different tasks through instructions or examples, providing an important foundation for its application to heterogeneous optical-network management tasks. Current pretrained LLMs can be broadly categorized into closed-source and open-source models. Closed-source models such as OpenAI’s GPT, Google’s Gemini, and Anthropic’s Claude, are characterized by extremely large training corpora and parameter scales, providing state-of-the-art general and even domain-specific reasoning capabilities, including advanced cross-modal understanding and generation. They are simple to use and highly effective for solving complex problems. In optical network scenarios, such models are therefore more suitable for non-sensitive, high-level tasks, for instance, conceptual architecture design, early-stage algorithm prototyping, or exploratory analysis of network behaviors when real operational data is not involved. Open-source models such as Qwen, DeepSeek, LLaMA, and GLM, support domain customization through fine-tuning and reinforcement learning. These models are particularly suitable for optical network scenarios that require customization and strict data governance, such as building performance estimation assistants, generating device-level configurations, supporting troubleshooting assistants, or enabling closed-loop control modules that must run within operator environments with proprietary data. B. Key technologies of LLM for vertical domain adaptation
1) Fine-tuning for domain task enhancement While pre-trained foundation models exhibit strong general reasoning capabilities, they may lack the specialized knowledge required for optical networks, where engineering decisions depend on field monitoring performances, physical-layer constraints, and device-specific standards, that rarely appear in general pretraining corpora. Therefore, fine-tuning becomes the key mechanism to bridge this gap, by further training models in domain-specific datasets, enabling models to internalize optical network concepts, operational patterns, and engineering heuristics [50], as illustrated in Fig. 2 (a). Several fine-tuning
strategies support domain adaptation with different computational and data requirements. Full-parameter fine-tuning updates the entire model but is resource intensive, whereas partial and parameter-efficient fine-tuning (PEFT) methods [51] offer more practical solutions when labeled data and computational resources are limited. Techniques such as low-rank adaptation (LoRA) and Quantized LoRA (QLoRA) reduce training overhead while maintaining strong adaptation performance [52, 53]. 2) Reinforcement learning for preference alignment After pre-training and fine-tuning, LLMs gain general generation ability and domain adaptability, yet they may still produce factually incorrect or misaligned outputs. RL therefore serves as a critical post-training stage for aligning model behavior with human preferences and task objectives. Reinforcement learning from human feedback (RLHF) [54] improves response quality, safety, and alignment by incorporating human or learned preference signals into iterative policy optimization, contributing to the success of systems such as ChatGPT and GPT-4 [55]. As a reward-driven learning paradigm, RL is particularly valuable for optical network O&M tasks, where decisions often involve trade-offs rather than deterministic answers. For example, feasible solutions in power optimization or fault recovery may satisfy technical constraints but differ in robustness, operational risk, or service impact. Such engineering preferences are difficult to encode through supervised datasets alone but can be naturally incorporated through reward-based learning [56]. 3) Retrieval-augmented generation (RAG) Although fine-tuning and RL improve domain adaptation and behavioral alignment, LLMs may still generate hallucinated or incomplete responses when tasks require precise technical knowledge or rigorous physical reasoning. In optical networks, many problems, such as interpreting nonlinear interference (NLI) or reasoning under multi-parameter constraints, depend on specialized knowledge that cannot be reliably encoded into model parameters alone. In this case, RAG addresses this limitation by grounding model outputs in external knowledge sources, improving response accuracy and relevance [57]. Effective RAG depends less on complex architectures than on a well-structured knowledge foundation capable of providing relevant and reli-
Research Article
able information when needed [58, 59]. For optical networks, this requires a structured knowledge library where physical theories, engineering guidelines, standards, and operational documents are explicitly organized and retrievable, rather than relying solely on the model’s internal knowledge. To ensure efficient use, knowledge should be carefully categorized, structured, and stored in vector form for fast retrieval. When a task is encountered, relevant knowledge is retrieved and incorporated into the LLM prompt, enabling context-aware and domain-grounded reasoning. 4) Prompt engineering By bridging the gap between user intent and model understanding [60], whether the model is pre-trained, fine-tuned, or reinforced, leveraging prompt engineering can unlock the full potential of LLMs, particularly in domain-specific applications. A well-constructed prompt typically includes four essential elements [61]: instruction, context, input data, and output indicator. The instruction defines the task clearly and steers the model’s reasoning toward user intents. Context supplements the instruction with domain knowledge, external data, or operational constraints, to enhance the model’s accuracy. Input data provides task-specific information, enabling the model to handle particular needs effectively. The output indicator specifies the desired format or structure, ensuring usability in downstream workflows. Beyond prompt elements, prompt techniques further enhance the reasoning and analytical capabilities of LLMs [43]. Chain-of-thought (CoT) prompting supports step-by-step reasoning by decomposing complex problems into sequential subtasks [62]. This capability is particularly valuable in optical networks, where many operational tasks are inherently multi-step and cannot be solved reliably through a single direct response. 5) AI Workflow with toolkit To enhance the reliability and efficiency of LLMs in task execution, workflow-based systems have emerged as a foundational paradigm. In this framework, complex tasks are decomposed into well-defined sub-steps, where LLMs operate within predesigned execution pipelines and interact with external modules through standardized interfaces. Unlike early LLM applications that focused primarily on language generation, workflow systems emphasize task-oriented automation, with the LLM acting as an intelligent coordinator within a structured execution process. As illustrated in Fig. 2 (e), workflow execution typically involves task decomposition, action planning, execution, and result aggregation, ensuring controllability and reproducibility for stable tasks. Many optical-network operations naturally align with this paradigm. For example, network planning can be decomposed into topology parsing, traffic demand analysis, routing and spectrum assignment (RSA), QoT estimation, and configuration generation. Similarly, operational tasks such as fault localization can be structured into sequential stages including alarm and telemetry collection, anomaly correlation analysis, root-cause inference, fault localization, and recovery recommendation generation. By organizing these steps into modular pipelines, workflow can coordinate domain-specific models and computational modules to accomplish complex engineering tasks more reliably and systematically. Workflow-based execution is particularly suitable when the task objectives, execution steps, required tools, and decision criteria can be specified in advance. In such cases, deterministic workflows provide advantages in repeatability, efficiency, and controllability, and there is no need to introduce autonomous Agent reasoning beyond the predefined procedure.
6
6) Harness engineering-enhanced Agent While workflow systems offer strong reliability, their predefined structures limit adaptability in dynamic and open-ended environments. To address this limitation, LLM Agents have evolved beyond static tool orchestration toward closed-loop, context-aware, and self-adaptive systems, as illustrated in Fig. 2 (f). An Agent is a self-directed entity that leverages LLMs for planning, reasoning, memory management, and tool invocation, enabling iterative action refinement through environmental feedback. Rather than replacing domain-specific tools or deterministic workflows, an Agent provides a flexible reasoning and orchestration layer above them. Given a high-level objective, the Agent can determine which information and capabilities are required, select and invoke appropriate tools or workflows, interpret their intermediate results, and adapt the subsequent execution according to the current context. Compared with workflows, which rely on predefined execution pipelines, Agents dynamically reason and adjust their behaviors at runtime, making them better suited for tasks requiring flexibility and adaptation [63]. Recent advances have further expanded the Agent paradigm through harness engineering, which provides a structured framework for orchestrating and operationalizing Agent behavior. Harness engineering further strengthens Agent reliability by organizing context management, tool interaction, persistent memory, observation, verification, and recovery mechanisms. These capabilities support closed-loop execution in which an Agent can evaluate intermediate results, refine its actions, and coordinate with other Agents through multi-Agent or Agentto-Agent (A2A) collaboration [64, 65]. Unlike static workflows, harness-enhanced Agents can dynamically determine task decomposition and execution strategies at runtime, making them well suited for complex optical network O&M scenarios.
4. LLM-CENTRIC FULL-LIFECYCLE MANAGEMENT FOR AGENTIC OPTICAL NETWORK The full LCM of optical networks encompasses multiple phases, including planning, deployment, operation, maintenance, upgrade, and decommissioning, typically spanning several decades, as illustrated in Fig. 3. As the foundational layer of modern communication infrastructure, enhancing the level of autonomy in optical networks is critical for improving the reliability of cross-domain traffic delivery while reducing the operational complexity associated with large-scale network management. In this context, LLMs, with their strong capabilities in intent understanding and logical reasoning, have the potential to support a wide range of O&M tasks. More importantly, LLMs can be envisioned as the cognitive core of AONs, continuously evolving alongside the network and permeating all lifecycle phases. Consequently, future AONs are expected to emerge as unified and intelligent systems, jointly constructed by physical infrastructure, DT, control planes, and multi-Agent systems. To systematically enable LLM-centric LCM in AONs, it is essential to first establish a comprehensive understanding of the tasks and workflows across different lifecycle phases. Accordingly, this section summarizes the major responsibilities of each phase, followed by the proposed hierarchical multi-Agent framework for full-lifecycle automation. A. The full lifecycle of optical networks
1) Planning The planning phase represents the initial stage of the optical network lifecycle, which involves defining network re-
Research Article
7
Fig. 3. Diagrammatic representation of the key tasks corresponding to each phase of the optical network lifecycle, covering plan-
ning, deployment, operation, maintenance, upgrade, and decommission. quirements, forecasting future capacity needs, and creating a blueprint for the physical and logical architecture [66]. This phase includes analyzing bandwidth demands, coverage targets, and service types, selecting appropriate technologies, designing fiber routes and redundancy schemes, and planning node placement across core, aggregation, and access layers. Equipment selection must balance performance, cost, and supply chain robustness, while operators often adopt multi-vendor strategies to mitigate vendor lock-in and enhance network flexibility. Based on candidate fiber routes and equipment options, optical amplifiers (OAs) placement and optical transponder units (OTUs) allocation can be jointly optimized with fiber and equipment selection. The corresponding decisions may include the locations and quantities of OAs and the allocation, types, and capacities of OTUs, subject to fiber-loss, QoT, cost, reliability, and deployment constraints. Building upon the selected resources, a preliminary network topology is constructed, enabling subsequent performance evaluation based on both fiber characteristics and device parameters. This evaluation provides the foundation for downstream configuration tasks, including wavelength planning driven by initial bandwidth demands. The outputs of the planning phase therefore comprise an optimized resource allocation and deployment scheme, initial configuration parameters for key optical devices, such as OAs, WSSs, and OTUs, as well as a unified IP addressing plan. This phase, often referred to as the Day 0 stage of the optical network lifecycle, establishes the baseline for all subsequent deployment and operational processes. 2) Deployment Once the planning phase is completed, the optical network transitions into the deployment phase, during which the designed infrastructure is physically instantiated and initialized for operation. This phase encompasses fiber connectivity establishment, equipment installation and configuration, as well as initial system performance optimization. Building upon the established physical connectivity and baseline configurations, the NMS subsequently assumes control, enabling real-time mon-
itoring and data collection to support the construction of DT [3, 67]. Based on the acquired network state information, optimization algorithms can be applied to validate and refine system performance, ensuring that service requirements are satisfied. Meanwhile, key deployment-stage data, which including fiber parameters, configuration files, begin-of-life (BoL) performance margins, and physical device location information, are systematically recorded to provide a reliable reference for subsequent operation stages. Upon completion of these processes, the network becomes service-ready, marking the transition to the Day 1 stage of the optical network lifecycle. 3) Operation Following deployment, the optical network transitions into the operation phase, during which it continuously carries and manages live IP traffic. This phase is inherently dynamic and long-term [5], involving three tightly coupled functional processes: adaptive optical channel management, continuous network health assessment, and real-time performance optimization. As service demands evolve over time, optical channels must be dynamically provisioned, adjusted, or released to accommodate traffic fluctuations. In parallel, network health is continuously evaluated through the analysis of performance metrics, sensor readings, and device status information. Based on the observed network state, performance optimization mechanisms are further triggered to maintain service quality [68]. Compared to the planning and deployment phases, the operation phase exhibits significantly higher complexity and persists over a longer period while continuously interacting with live traffic and facing the stringent requirement for ultra-reliable execution, where even minor disruptions may lead to substantial service impact. This necessitates decision-making processes to be not only accurate but also risk-aware and globally coordinated. In this context, autonomous operation emerges as a critical enabler, capable of reducing human-induced errors while rapidly adapting to unexpected network conditions through holistic, system-level awareness. 4) Maintenance
Research Article
Maintenance is performed concurrently with the operation phase and represents the most enduring stage of the optical network lifecycle, playing a critical role in sustaining long-term network reliability and availability. This phase focuses on the timely detection, localization, and mitigation of network faults arising from diverse sources [3, 5]. To address such failures, maintenance activities rely on a combination of network monitoring and coordinated field operations. In NMS, alarms are generated to facilitate fault analysis and location, after which maintenance personnel collaborate with field technicians to perform recovery actions, including fiber repair and module replacement. The efficiency of fault detection and resolution directly impacts the overall availability and continuity of services in the network. However, in large-scale optical networks, traditional manual maintenance mode faces significant challenges in terms of scalability, response time, and knowledge consistency, particularly in the presence of staff turnover and increasingly complex system configurations. In this context, autonomous maintenance is anticipated to be a key enabler to accelerate fault recovery, reduce operational overhead, and alleviate reliance on human expertise. 5) Upgrade The upgrade phase of optical networks encompasses both technological evolution and network scale expansion, enabling the system to adapt to the long-term growth in service demand and advances in transmission and control capabilities. Such upgrades may involve capacity expansion, equipment replacement, spectrum reconfiguration, and updates to control and management functions as network requirements evolve. Despite occupying a relatively small portion of the entire lifecycle, the upgrade phase is inherently high-risk, as any misconfiguration or maloperation may lead to performance degradation and even service disruption in live networks. Consequently, upgrade operations require meticulous planning and evaluation, typically involving extensive simulations and performance assessments on the DT platform, and in many cases, physical testing within controlled laboratory environments. In this context, the integration of high-level autonomy with DT technologies is expected to significantly improve both the efficiency and reliability of upgrade processes by enabling predictive analysis, risk-aware decision-making, and pre-deployment validation. 6) Decommission The final phase of the optical network lifecycle is decommissioning, which marks the systematic retirement of network infrastructure as it becomes obsolete or no longer economically viable. This phase is typically triggered by several factors, including large-scale technology transitions, progressive system aging, and declining service demand that renders existing systems inefficient [69]. Decommissioning is inherently complex, involving large-scale service migration, resource reclamation, and coordinated asset management across multiple network layers. Among these processes, traffic cutover operations are particularly critical, as they must be executed with minimal disruption to ongoing services. Consequently, ensuring a safe and seamless transition requires precise planning, real-time monitoring, and tightly coordinated execution. In this context, autonomous network capabilities play a pivotal role by enabling automated service migration, reducing operational risks, and maintaining service continuity throughout the decommissioning process. Across these phases, the network evolves from design and physical realization to long-term operation, evolution, and eventual retirement. The diversity of objectives, constraints, and operational conditions across the lifecycle motivates a phase-
8
oriented organization of Agent responsibilities, as discussed in the following subsection. B. Hierarchical multi-Agent framework
To effectively facilitate the autonomous LCM of optical networks, we propose a hierarchical multi-Agent framework comprising the four levels: 1) a Network Director overseeing global operations, 2) three Administrative Agents providing cross-cutting capabilities for task coordination, memory management, and security support, including Task Coordinator, Memory Container, and Security Manager, 3) six Division Agents responsible for different phases of the network lifecycle, including planning, deployment, operation, maintenance, upgrade, and decommission phases, and 4) multiple AI Experts within each division executing specific functions. Each Agent is designed to autonomously execute tasks within its domain while collaborating across layers, ensuring seamless and intelligent optical network management, as depicted in Fig. 4. At the highest level of the framework, the Network Director serves as the central decision-making entity, orchestrating the entire multi-Agent system. It acts as the primary interface between human operators and the AI Agents, maintaining a global view of network objectives and operational states. It is initially customized with basic knowledge of optical network operation and is responsible for controlling the overall task execution process of network automation. Rather than directly executing all domain-specific tasks, the Network Director interprets high-level intents, allows the system to maintain a unified network-wide objective while delegating specialized tasks to dedicated Agents. At the secondary level, three Administrative Agents support the Network Director: Task Coordinator, Memory Container, and Security Manager, each responsible for a critical aspect of autonomous network governance. These Agents provide crosscutting capabilities that are shared by multiple lifecycle divisions and therefore are not tied to any particular network phase. The Task Coordinator is responsible for the distribution of tasks across divisions and optimizing the task execution process. The Memory Container functions as a knowledge repository, storing historical data, learned experiences, decision logs, and feedback reports, thereby preserving contextual information across different lifecycle phases and supporting informed decision-making. The Security Manager ensures network reliability and safety through proactively monitoring anomalies and faults, detecting potential threats, and pre-validating configuration strategies to protect the system from risks. For actions that may affect the physical network, the Security Manager can validate strategies from a global perspective before execution and triggering human confirmation for high-risk operations. Separating these shared functions from individual Division Agents avoids duplicated coordination, memory, and security mechanisms and enables consistent policies and contextual information to be maintained throughout the network lifecycle. At the division level, six Division Agents designed for different phases of the optical network: Planning, Deployment, Operation, Maintenance, Upgrade, and Decommission, are responsible for managing the full lifecycle of the optical network. This phase-oriented organization assigns a clear ownership of tasks according to their lifecycle context, since different phases have distinct objectives, constraints, information requirements, and decision processes. Each Division Agent acts as a middle manager that interprets task targets from administrative layers, formulates phase-specific strategies, and orchestrates task ex-
Research Article
9 Task instruction
User/Operator
Planning Agent
...
Response
Optical Network
Network Director
Task Coordinator
Memory Container
Security Manager
Coordinate tasks among divisions Optimize the task execution process
Storing historical data, learned experiences, decision logs, and feedback reports
Detect threats, and ensure security
Operation Agent
Deployment Agent
...
Infrastructure Configuration Field Selector Scheduler Engineer
Feedback
... Performance Tuner
Dynamic Operator
... Automated Inspector
Alarm Handler
Upgrade Agent
Maintenance Agent
... Abnormal-data Handler
Upgrade Planner
Decommission Agent
... Validation Specialist
Retirement Planner
Migration Supporter
Network management system (NMS) Optical Network
Fig. 4. Hierarchical multi-Agent framework for optical network full-lifecycle management. The architecture consists of four layers
of intelligent Agents. At the top layer, the Network Director Agent interacts with the user/operator, handling both intent input and feedback aggregation from lower layers. The second layer comprises three administrative Agents, each responsible for a major domain of network lifecycle management. The third layer includes primary Agents that oversee the automation of specific lifecycle phases. At the bottom layer, sub-Agents are assigned to execute concrete tasks.
Table 2. Representative knowledge, network data, and tools for Agents across different lifecycle phases of AONs. Lifecycle phase
Representative knowledge and data
Representative tools and interfaces
Planning
Network topology; traffic demands and forecasts; equipment specifications; deployment constraints; equipment databases;
RSA tool; QoT estimation; resource planning; DT platform;
Deployment
Equipment inventory; device capabilities; configuration tem- NMS API; device configuration interfaces; commissioning and plates; installation records; diagnostic tools;
Operation
service statistics; topology and network state; QoT measurements; alarm information; action logs;
Real-time telemetry APIs; QoT estimation; RSA tool; DT platform; NMS API; performance-analysis scripts;
Maintenance
Alarm rules; historical cases; maintenance records; OTDR traces; equipment state information; troubleshooting procedures;
Alarm analysis tools; fault-diagnosis tools; OTDR APIs; DT platform; NMS API; recovery and reconfiguration tools;
Upgrade
Capacity requirements; traffic growth trends; equipment databases; spectrum utilization; upgrade constraints;
Capacity optimization tool; RSA tool; QoT estimation; DT platform; resource planning; configuration and validation tools;
Decommission
Service inventory; SLA requirements; equipment and asset records; retirement procedures;
QoT and performance verification; service migration tools; NMS API; asset-management systems; configuration tools
ecution within its domain. It also validates the outputs of its sub-Agents against task objectives and constraints, forming a local closed loop in which unsatisfactory results are refined or reexecuted before being passed to the upper level. These Agents ensure that every lifecycle phase, from initial planning to network decommission, is intelligently coordinated and seamlessly integrated into the overall management process. At the expert level, AI Experts operate under the supervision of their respective Division Agents to execute specific sub-tasks efficiently. The Planning Agent focuses on network design, resource allocation, and feasibility analysis. It can direct the Infrastructure Selector to choose appropriate components and instruct the Configuration Scheduler to implement initial settings. The Deployment Agent ensures seamless infrastructure rollout, commanding the Field Engineer for physical installation and the Performance Tuner for system verification and stability testing. The Operation Agent maintains real-time network functionality, with the Automated Inspector executing anomaly detection and the Dynamic Operator managing traffic and resource allocation according to the actual requirements. The Maintenance Agent ensures reliability by directing the Alarm Handler to repair problems, and the Abnormal-data Handler to prevent failures. The
Upgrade Agent enhances performance, instructing the Upgrade Planner to implement improvements and the Validation Specialist to verify network stability. The Decommission Agent manages network retirement, overseeing the Retirement Planner for equipment removal and the Service Switcher for seamless data and service transitions. Experts shown here are representative rather than exhaustive; additional task-specific Experts, such as Device Manager and Physical Decommission Agent, can be instantiated according to specific network requirements and may be added as the network evolves. This division-to-expert structure separates lifecycle-level coordination from task-level execution, allowing each Expert to specialize in a specific function and interact with the corresponding models, knowledge bases, and network tools. As a whole, all Agents are capable of inter-Agent communication via A2A protocols [70] and interact synergistically to ensure seamless coordination across the entire network lifecycle. The Planning, Deployment, and Decommission Agents are activated when their corresponding lifecycle phases are initiated: the Planning Agent defines network architecture and resource allocation, the Deployment Agent manages infrastructure rollout and initial setup, and the Decommission Agent oversees retirement and mi-
Research Article
(a)
10
Optical network to be planned
Toll Station Expressway West Node 1
?
Which vendor’s equipment and how many sets?
Which fiber segments and optical amplifier sites?
?
Node3
Node1
We plan to build a new 4-node ring optical network. Please select appropriate equipment, fiber segments, and amplifier sites based on the following requirements, and provide equipment parameter configurations. Requirements: the four nodes are [node1, node2, node3, node4]; C band system with a single-channel rate of 400Gb/s; any two nodes have dual-route redundancy protection, and the latency must be less than 0.5 ms; the initial bandwidth between any two node is 2.4Tb/s.
①
Equipment: Vendor N, 6 sets of OTN-C equipment Fiber segments: Node1-> Expressway West-> Toll Station-> … ->Node2 -> Optics Valley Road ->… Node3-> Information port->… Node4 Amplifier sites: East Toll Station, Auto repair shop, … , Victory Mall
(c)
Name Infrastructure Selector Goal Output the equipment, fiber segments and amplifier sites that meets the requirements
Planning
LLM
I need to provide the equipment type and quantity, specific fiber segments and amplifier sites according to the requirements, while considering cost and stability.
Memory
Name
Parameter Scheduler
Goal
Generate a set of initial parameter configurations and provide BoL margin
③
Self-iteration
④
Planning
Action
Equipment selection : Vendor N,
LLM
6 sets of C-band OTN…
×
Fiber segments selection :
Goal Achieved? √
Node1-> Expressway West-> Toll Station->… Amplifier site selection: East Toll Station…
LLM
Call Tools Call resource selection tools to
I need to build the topology first, then perform QoT estimation based on the characteristic of equipment, determine the best RWA scheme, and finally provide BoL margin.
Vendor and specification•
Origin and destination •
Amplifier sites •
Location
×
LLM
Goal Achieved? √
Call Tools parameters, and BoL margin calculation
Tool Kit
Tool kit
Knowledge Library
Equipment selection Cost
Fiber segments resource •
OTU1-node1-191.4THz, 135GBaud … OA-node1-east: gain 21dB … WSS1-node1-Media Channel: MC 1 … BoL margin: OTU-1 node1 4.5dB
RAG
Knowledge Library Equipment information •
Parameter configuration:
Obtain a set of initial equipment
Query infrastructure resource infos
RAG
⑤ Action
Memory
obtain a list of optical network infrastructure
Query infrastructure resource infos
Out
In
Out
Self-iteration
Parameter Scheduler
④ Task2: Provide parameter configurations based on requirements and [equipment, fiber segments, amplifier sites], and evaluate BoL margin of each service channel.
(b)
In
OTU1-node1-191.4THz, 135GBaud,+3dBm Tx power … OA-node1-east: gain 21dB … WSS1-node1-Media Channel: MC 1 add port 1 191325 191475 … BoL margin: OTU1-node1 4.5 dB, OTU2-node1 3.6dB … ⑤
Planning Agent
③
Gain 23dB Tilt 1.5dB OTU1 Node4 Tx +3dBm 191.4THz…
The planning report is as follows: ✓ Selected equipment, fiber segments, amplifier sites ✓ Initial parameter configurations The optimal configurations and BoL margin are as follows:
⑥
Task1: Select a set of [equipment, fiber segments, amplifier sites] that meets requirements. ②
②
Planned network
Fiber segments
Victory Mall
Infrastructure Selector
Optical amplifier sites
EastToll Station
?
What are the configuration values for each parameter of the equipment?
Node2
•
Length, fiber loss, rental cost Electric power
• • •
Vendor and Supply chain health analysis Performance evaluation Cost
• •
Segments combination Latency and cost calculation
Fiber segments selection Amplifier sites selection •
Electric power evaluation
Equipment information • •
Vendor and specification • Type and Characteristic
Fiber segments information • •
Origin and destination Length, fiber loss
Topology build Port rate and interface type
• •
Node relative position Link building
Routing and spectrum assignment • •
Parameter configuration optimization & QoT estimation
Latency SNR considered
BoL margin
Fig. 5. Schematic of multi-Agent collaboration in the planning phase of the optical network, (a) coordination among Planning Agent, Infrastructure Selector, and Parameter Selector, (b)-(c) functional workflow of the Infrastructure Selector and the Parameter Scheduler: based on the input, the LLM plans the task, retrieves necessary information via RAG from the knowledge base, and invokes tools to carry out a sequence of operations. Once the task is completed and the objective is met, it returns the results to the higher-level Agent.
gration. In contrast, the Operation, Maintenance, and Upgrade Agents remain continuously active throughout the network’s operational phase, ensuring real-time monitoring, maintenance, and performance optimization. These three Agents form the core of network management, working closely with the Task Coordinator, Memory Container, and Security Manager to maintain efficiency, adaptability, and security. For example, when the Operation Agent detects abnormal performance metrics, it first attempts self-correction before escalating the issue to the Task Coordinator, which then assigns the task to the Maintenance Agent for immediate troubleshooting or to the Upgrade Agent for hardware adjustments and system enhancements. Although all these multi-level Agents are powered by LLMs as the core technology, we can flexibly select models of different scales based on task complexity and capability requirements, enabling hierarchical deployment across the cloud, edge, and device. Core control and complex task processing can be handled by large models in the cloud, while latency-sensitive and resource-constrained tasks are efficiently executed by relatively small models deployed at the edge or devices, optimizing com-
putational resource allocation while ensuring high-level intelligence. In practical deployment, each lifecycle Agent is grounded in a combination of domain knowledge, network-state information, and executable tools. Knowledge and data provide the contextual information required for Agent reasoning, while domain-specific tools provide validated computational, simulation, monitoring, and control capabilities. The Agent coordinates these heterogeneous resources according to the task and the current network state. Representative knowledge sources, network data, and tools for the six lifecycle phases are summarized in Table 2.
5. AUTONOMOUS LIFECYCLE MANAGEMENT OF AGENTIC OPTICAL NETWORKS The AONs extend beyond single task automation toward coordinated and lifecycle intelligence. In this section, we present the architecture of LLM Agent-driven autonomous LCM of optical networks, encompassing the detailed workflow of the six phases from planning to decommission. Instead of focusing on individual task automation, we emphasize how multiple specialized
Research Article
Agents can be organized and orchestrated to support all stages in a coherent and scalable manner. By introducing how Agent capabilities can be integrated into optical networks, we aim to provide a structured foundation that can inform initial system design and deployment for future AONs. A. Planning phase
In the initial planning phase of the optical network lifecycle, the primary task is to perform demand analysis and topology design based on factors such as the operator’s service development plan, capital investment strategy, and network reliability requirements. This process determines the fundamental attributes of the new network, including node geographic locations, network capacity, redundancy protection schemes, latency requirements between nodes, and initial bandwidth provisioning. As this process constitutes the top-level design of the network, it involves diverse and dynamically changing information elements, and it is necessary to combine information from a knowledge base and utilize various tools for Agents. To address these challenges, based on the above hierarchical multi-Agent framework, it is designed to deploy three Agents for the planning phase. First, the Infrastructure Selector Agent is responsible for selecting infrastructure resources. Second, the Parameter Scheduler Agent generates configuration parameters. Third, the Planning Agent serves as the orchestrator, coordinating task reception, information integration, and downward command delegation, functioning as a communication bridge within this multi-Agent system. These roles do not imply that every step requires LLM reasoning. Deterministic database queries, rule-based filtering, QoT estimation, and numerical optimization remain preferable for well-defined inputs and constraints because they provide reliable and reproducible outputs. The added value of LLM-based Agents lies at the orchestration level, particularly in interpreting high-level or incomplete operator intents, translating contextual preferences into selection criteria and optimization objectives, reasoning over trade-offs among performance, cost, supplychain robustness, and reliability, and dynamically selecting and sequencing the required tools. Accordingly, the Infrastructure Selector uses LLM reasoning to synthesize heterogeneous information and establish resource-selection priorities, whereas the Parameter Scheduler invokes deterministic tools to generate reproducible configurations and performance estimates. For stable tasks with fixed requirements and predefined workflows, a purely deterministic implementation may be sufficient. Agentbased orchestration becomes valuable when requirements are incomplete or evolving, multiple tools must be coordinated, or intermediate results require the workflow to be adapted. In this phase, Network Director notifies the Task Coordinator to process the user-defined task and assign the optical network planning task to the Planning Agent. Taking a simple four-node ring topology as an example, the collaboration process among these Agents can be illustrated as depicted in Fig. 5 (a). First, the Planning Agent receives key information about the network, including the names of four nodes, the per-wavelength rate of a C-band system, and the requirement for dual-path redundancy protection. Based on these requirements, the Planning Agent delegates the task of resource selection to the Infrastructure Selector. The workflow of the Infrastructure Selector Agent is depicted in Fig. 5(b). After understanding the Planning Agent’s intent, the Infrastructure Selector retrieves information on fiber segments related to the specified node locations and corresponding amplifier sites using RAG from the infrastructure knowledge
11
base. Meanwhile, it retrieves equipment information that meets user requirements, including vendor names, technical specifications, and pricing. Next, the Infrastructure Selector enters the Action phase and invokes resource-selection and deploymentoptimization tools to select equipment and fiber segments and, where multiple candidate solutions are available, to determine the locations and quantities of OAs and the allocation of OTUs. Next, the Infrastructure Selector enters the Action phase, invoking resource selection tools to choose from the three categories of resources: equipment, fiber segments, and amplifier sites. Equipment selection must consider multiple criteria such as supply chain robustness, cost, and technical performance, with specific priorities depending on the operator’s strategic preferences. Once a feasible set of resources is selected and deemed to meet user requirements, the Infrastructure Selector sends the result to the Planning Agent. If the selected resources fail to meet expectations, the Infrastructure Selector must iterate on the selection process. Following this, the Planning Agent integrates user requirements with the selected infrastructure resources and instructs the Parameter Scheduler to generate initial configuration parameters and estimate BoL margin for service channels. Upon receiving instructions from the Planning Agent, Parameter Scheduler initiates its workflow by retrieving detailed resource information via RAG from the knowledge base. It then proceeds to invoke a series of tools: first, a topology construction tool used to simulate the four-node ring network, followed by a QoT estimation tool and device configuration optimization algorithm. These tools determine parameters such as OA gain/tilt and WSS channel attenuation settings. The Parameter Scheduler then completes RSA according to latency and SNR requirements. Finally, the BoL margin of the service channels is obtained. If the margin values meet or exceed the minimum threshold, the Parameter Scheduler submits the generated configuration back to the Planning Agent. After that, Planning Agent consolidates the final planning results into a planning report, which includes a detailed network topology and configuration summary. This report is returned to the Task Coordinator, which guides the subsequent deployment phase. This example illustrates how three Agents can collaborate to efficiently complete infrastructure selection, parameter pregeneration, and performance estimation for the planning phase. It is important to note that all Agents are based on LLM and rely heavily on well-designed prompts and harness engineering to effectively execute planning, action reasoning, RAG queries, tool invocation, and result evaluation. Additionally, to ensure optimal infrastructure selection by the Infrastructure Selector Agent, the completeness and standardization of information in the infrastructure knowledge base is critical. Similarly, for the Parameter Scheduler Agent, beyond knowledge base quality, the accuracy and reliability of tools in the toolset library are equally essential. In this framework, LLM-based Agents act as orchestrators rather than directly replacing deterministic numerical optimizers. They translate planning requirements into objectives and constraints, retrieve candidate resources and uncertainty information, invoke appropriate deterministic or robust optimization tools, and verify and iterate the returned solutions. Robust or rolling-horizon optimization can also be invoked when traffic demand, component aging, failure risks, or resource availability evolve over time. Therefore, OA and OTU deployment optimization can be implemented as a specialized computational module while remaining embedded as a tool-enabled subtask of the planning phase. The resulting de-
Research Article
12
(a)
Please complete the optical network deployment according to the planning report.
⑨ ①
Deployment completed
Task1: Guide the fiber team to complete the fiber connection and the equipment team to complete installation and initialization according to the planning report. ②
Field Engineer
④
The fiber splicing between the four nodes has been completed, and all equipment have been initialized.
③
Fiber team • Splicing • Patching
Guidance & reporting
Equipment team • Installation • Initialization
Network performance optimization has been completed. Essential initial information, such as fiber length, loss, and amplifier gain, BER, has been recorded in the database. ⑧ ⑤
Deployment Agent
Task 2: Perform optical network performance tuning and record the initial network information in the database.
Node 2
10101010110101010101010101010 10010111101010101010101010101 01010101011111010101010101010 Node3 10101010101010010001111111111 Node1 10001010100100101010111001011 01010111100000101001010010100 Node4
OA
Config
Node 3
Node 1 Fiber splicing
Data
Node 4
NMS
Field optical network
(b)
Planning
②
LLM
I need to direct fiber and equipment team to complete the on-site installation, and the check the network connectivity.
Memory
Query fiber and equipment resource information.
RAG Infrastructure Knowledge Base Fiber segments information • Origin and destination • Length, fiber loss
Equipment knowledge • Installation manual • Command line manual
Self-iteration
Out
Action
LLM
④
Fiber platform : Start fiber connection Goal Equipment platform: Put on the Achieved? shelf, power on, and initialize the configuration NMS platform: The equipment is connected and the IP address is correct… Call external interaction platform interface tools to guide fiber and equipment teams, and managed equipment
Call Tools External interaction platform interface tools Fiber platform • Issuance of optical fiber connection instructions Equipment platform • Issuance of equipment installation instructions NMS platform • Equipment management integration • Configure initial parameters
⑦
Network initial data storage
Digital twin
(c)
Name Field Engineer Goal Guide fiber and equipment team to complete the field engineering according to the planning report
In
⑥ Verification & Feedback
Node2
OA
Performance Tuner
Name Performance Tuner Goal Complete optical network performance tuning and initial data storage
In ⑤
Planning
LLM
I need to optimize the optical network performance and record the initial network information
Memory
Query fiber and equipment resource information.
RAG Infrastructure Knowledge Base Equipment information • Vendor and specification • Type and Characteristic
Fiber segments information • Origin and destination • Length, fiber loss
Out
Self-iteration
⑧
Action
LLM
Parameter optimization: Launch power: otu1 +3.5dBm… EDFA gain/tilt: OA1 25.6dB/1.5dB… WSS voa: [1, 2.2, 2.0, 1.5…] BER query:otu1 BER 2.3e-3…
Goal Achieved?
Network initial data storage Call parameter configuration tools to optimize network performance
Call Tools Parameter configuration Tools Parameter optimization • Launch power • EDFA gain/tilt, WSS channel VOA Digital twin interface • Parameter verification • Performance feedback Database interface API
Fig. 6. Schematic of multi-Agent collaboration in the deployment phase of the optical network, (a) coordination among Deploy-
ment Agent, Field Engineer, and Performance Tuner, (b) functional workflow of the Field Engineer, (c) functional workflow of the Performance Tuner. ployment plan is then physically implemented in the subsequent deployment phase. B. Deployment phase
In the deployment phase of the optical network lifecycle, the primary objective is to implement the network topology and equipment configurations generated during the planning phase and to perform initial performance tuning. This phase is characterized by frequent interactions between the network operation center (NOC) and field engineers, as well as communications between the NMS and physical devices. To efficiently manage this complicated task, three types of AI Agents are assigned to this phase: a top-level Deployment Agent, a Field Engineer Agent responsible for field coordination, and a Performance Tuner Agent responsible for network performance optimization. The Deployment Agent receives deployment tasks from the Task Coordinator, along with the planning report produced by the Planning Agent. Following the typical workflow of the deployment phase, the Deployment Agent first instructs the Field Engineer Agent to carry out on-site tasks such as fiber connection and equipment installation. Here, the Field Engineer Agent can be embodied and equipped by human with mobile terminals, or potentially through robotic systems for assisted or automated operations. Then Deployment Agent assigns the Performance Tuner to execute optical network performance tuning and store initial network data. The workflow of Field Engineer
is illustrated in Fig. 6 (b). Upon receiving instructions, the Field Engineer can access relevant fiber and equipment information through RAG, and interacts with the fiber and equipment platform to coordinate with the fiber and equipment teams, respectively, for field implementation. Since the Field Engineer can interface directly with human field engineers, the platforms may support multimodal interactions including text, voice, image, and video, similar to features provided by modern instant messaging tools. During the fiber connection process, Field Engineer guides the fiber team to connect fiber segments based on the planning report. Some segments are connected via splicing, while others may require patch cords. It is essential that Field Engineer instructs the fiber team to use equipment such as optical time domain reflectometer (OTDR) to ensure fiber attenuation and splice loss meet engineering standards. Additionally, using installation manuals retrieved from the equipment knowledge base, Field Engineer guides the equipment team in installing devices at optical amplifier sites and ROADM nodes. This includes completing fiber patching and initializing the equipment to enable proper inter-device communication. Field Engineer may then invoke the NMS to incorporate the equipment into the management system and deploy the OTU and optical equipment configurations generated during planning. Once these steps are completed, the optical channel (OCH) can be established, and the network reaches a near-optimal operational state. At this point, telemetry data, such as network status and transmission
Research Article
13
Upper Agents Requirement
NMS
Telementry
APIs
Add channel
Report/Response
Drop channel
Requirement for dynamic adjustment
LLM
Operation Agent
Response Re-allocated traffic and network performance Self-iteration
In
Out
Planning
Action
1. Requirement analysis 2. Performance prediction 3. Configuration
1. Plan and rehearsal 2. Feasibility analysis 3. Deploy and response
× √
LLM
Call Tools
Memory
Involve APIs and tools for traffic re-allocation
Query historical logs and solutions
RAG
Tool Base
Knowledge Base Expertise to solve operational requirements • Choices for performance optimization (by priority) → 1. power control → 2. EDFA configuration → 3. Re-route → 4. Report cannot opt ... • Steps for service dynamic adjustment 1. Add channels: →RWSA →QoT estimation →config →check → 2. Drop channels: →require →QoT estimation →config →check → • Requires hardware support →Report to upper agents for assistance ...
• • •
APIs of NMS • Performance data • Configuration APIs APIs of DTON Calculation tools • Power optimization • EDFA configuration • Route algorithm • Spectrum allocation • QoT estimation tools • ...
?
Automatic inspection target
Adjust configuration and improve performance based on the traffic requirements
Dynamic Operator
How is the network state and does it need to be inspected?
Report
Automated Inspector
Requirement: “Drop Och-3 on Path (N1→ N2 → N3) and establish a new 400Gb/s Och-4 between N1 and N4 .” Operation Agent: Task decomposition: Step 1: Retrieve APIs of NMS for data collection Step 2: Retrieve the algorithm for RWSA of new service Step 3: Retrieve APIs of DTON for simulating the dynamic changes after add/drop Step 4: Evaluate the network performance and analyze feasibility of add/drop Step 5: Retrieve APIs of NMS for configuration Step 6: Check the performance and generate a response
Dynamic Operator: Action trace: [ 14:32:05 ] NMS Telemetr y Status: ACTIV E Central Frequency: 193.25
THz
S ervice ID: svc - 400 G- 003 Route: NODE - 1 → NODE- 2 → NODE- 3 Current GSNR: 20. 42 dB
[ 14:32:07 ] Service Managemen t Request received: DELETE svc - 400 G- 003 [ 14:32:10 ] Configuration Manage r S ervice svc - 400 G- 003 removed successfully Spectrum resources released: [ 193.20 , 193.30 ] TH z [ 14:32:12 ] Provisioning Request S ervice ID: svc - 400 G- 004 Selected frequency: 194.10 THz Selected route: NODE- 1 → NODE- 3 → NODE- 4 [ 14:33:08 ] QoT Estimation Predicted GSNR: 19.83 dB Required GSNR: 16.00 dB QoT margin: + 3.83 dB Status : FEASIBLE [ 14:33:21 ] NMS Configuration API [ 14:33:31 ] NMS Telemetry Status: ACTIVE Central Frequency: 194.10 THz
Provisioning configuration to device s Service ID: svc - 400 G- 004 Route: NODE - 1 → NODE- 3 → NODE- 4 Current GSNR : 19.67 dB
Operation Agent: Och-3 released, Och-4 added at ...
Fig. 7. An illustration of multi-Agent collaboration during the operation phase of AONs, along with the functional workflow of the
Dynamic Operator and a detailed description of its behavior in a channel adding scenario. performance, can be collected to construct an DT of the deployed physical network, as shown in Fig. 6 (a). Since actual fiber parameters may slightly deviate from those assumed during planning, fine-tuning of device configurations is necessary to achieve optimal power flatness and SNR. These adjustments, performed by the Performance Tuner, can be verified within the DT before being issued to the physical network. The optimization parameters mainly includes launch power profile, EDFA gain/tilt, and WSS channel attenuation settings. Once performance metrics converge to their optimal values, the optimized configuration is deployed to the network devices. Finally, the fiber parameters, equipment configurations, and transmission performance data are stored in the network initial database, which serves as the starting point of the lifecycle-wide database for subsequent operation and maintenance phases. This phase further demonstrates the critical importance of high-quality knowledge and tool libraries in enabling Agent efficiency and accuracy. For example, the Performance Tuner relies heavily on the accuracy of the DT and the effectiveness of optimization algorithms. Therefore, careful design and continuous maintenance of both the knowledge base and toolset are essential throughout the entire network lifecycle. C. Operation phase
In the operation phase, the primary focus shifts to ensuring the network’s ongoing stability, efficiency, and reliability, as well as adapting to real-time changes in traffic and condition. This phase is essential for maintaining high-quality service throughout the network’s lifecycle. The Operation Agent plays a central
role in managing this phase, orchestrating network tasks, monitoring optical performance, and taking corrective actions when necessary. Unlike the Planning Agent and Deployment Agent, which operate only at the initial phases of the network’s lifecycle, the Operation division remains continuously active, overseeing the daily functioning of the network and ensuring that it runs smoothly throughout its operational lifetime. To efficiently manage the various tasks required during this phase, the Operation Agent works in close collaboration with two specialized sub-Agents: Dynamic Operator and Automated Inspector. Dynamic Operator is responsible for the continuous adjustment of network parameters to accommodate fluctuations in network traffic, bandwidth demands, and other dynamic variables. This Agent makes real-time decisions to balance network traffic and improve performance, ensuring that each part of the network is operating at peak efficiency. It plays a crucial role in continuously optimizing network performance, ensuring service quality, and dynamically adjusting network configurations based on evolving demands. One of Dynamic Operator’s primary functions is network performance enhancement to guarantee high-quality service provisioning. It continuously evaluates network performance by analyzing key indicators such as latency, power, SNR, and BER. Based on changes in traffic demand or new service requirements, it can autonomously adjust network configurations and expand capacity as needed. Through NMS, Dynamic Operator primarily executes software-level adjustments, including launch power optimization, device parameter configuration, and routing adjustments. As illustrated in Fig. 7, a representa-
Research Article
14
Telementry
NMS APIs
Report/Response
Requirement
Upper Agents Telementry
NMS
?
APIs
Requirement for dynamic adjustment
Automatic inspection target
Adjust configuration and improve performance based on the report
Perform a daily network operation inspection and report any possible anomalies
Dynamic Operator
Response Optimized network performance
Operation Agent
Report
Automated Inspector
No event or report anomalies
Operation Agent: “Perform once network inspection and report any potential anomalies.” Step 1: Retrieve APIs to query performance indicators in the current network state Step 2: Retrieve tools to detect abnormal power deviations and identify affected elements Step 3: Analyze any abnormal events based on historical logs and expertise
In
Automated Inspector: Action trace: [ 15:41:05 ] NMS Telemetry OCH- 03 : output power min - 5.6 [ 15:41:12 ] Anomaly Detection [ 15:41:15 ] Alarm
dBm
OCH- 01 : output power min - 4.2 dBm OCH- 04 : output power min - 45.0 dBm OCH- 04 : output power min - 45.0 dBm XX - OLA- 024 : LOS (Loss of Signal)
LLM OCH
- 02 : output power min - 4.6 dBm OCH- 05 : output power min - 4.7 dBm Power level below normal operating range.
√
LLM
Involve APIs and tools to assist in analysis
RAG
Monitoring indicators for covert problems •
OCH- 04 : output power min −45.0 dBm Service status: DEGRADED Primary path: NODE - 1 > NODE - 3 > NODE - 4 Secondary path: NODE
OCH- 04 : output power min - 5. 1 dBm path: NODE - 1 > NODE - 3 > NODE - 4 - CLEARED
Analyze... and record...
Knowledge Base
Action trace:
Primary
Inspection scheduling
Call Tools
• - 1 > NODE - 2 > NODE- 4
• •
[ 15: 44 : 37 ] QoT Estimation Predicted GSNR after protection switching: 18.92 dB Required GSNR: 16.00 dB QoT margin: + 2.92 dB Protection path status: FEASIBLE [ 15:45:01 ] NMS Configuration API OCH- 04 rerouted to: Secondary path: NODE - 1 > NODE - 2 > NODE - 4 [ 15:45:16 ] NMS Telemetry Working path: SECONDARY Current GSNR: 18.76 Db
Action
Memory
Operation Agent: “Optimize the configuration and verify the performances.”
[ 15: 42: 07] NMS Telemetry Workin g path: PRIMARY
Out ×
Planning
Query indicator descriptions and solutions
Step 1: Retrieve APIs of NMS for data collection Step 2: Check path switching strategy for optimization, protecting the och transmission Step 3: Evaluate the network performance and analyze feasibility of the path switching Step 4: Retrieve APIs of NMS for configuration Step 5: Check the performance and generate a response
Dynamic Operator:
Self-iteration
Service status: ACTIVE Secondary path: NODE - 1 > NODE - 2 > NODE - 4
Alarm: LOS
Operation Agent: Report results, update memory.
Service performance (Power, SNR, BER, channel/ bandwidth, spectrum utilization, profile flatness ...) Fiber anomaly (Length, insertion loss, fast fiber loss anomaly, co-trench/co-cable identification ...) Equipment health (Transceiver, WSS, amplifier ...) Environmental sensor (Temperature, humidity ...)
Expertise for prediction and resolution • •
Indicator threshold • Step-by-step solutions Historical decision-making ...
Tool Base •
•
APIs of NMS • Performance query • Spectrum query • Route query • Device query • Module query • ... Calculation tools • Power flatness • Spectrum utilization • ...
Fig. 8. Schematic of functional workflow of the Automated Inspector in operation phase, along with a detailed description of its
behavior in a network anomaly inspection scenario. tive dynamic add/drop task demonstrates the execution process from an upper-level requirement to concrete network actions. The Operation Agent first decomposes the requirement into API retrieval, RSA, DT-based performance evaluation, and configuration verification steps. The Dynamic Operator then invokes the corresponding NMS and computational tools and produces an executable action trace. In the demonstrated case, an existing 400 Gb/s service is released and a new service is provisioned on an alternative path and frequency, followed by GSNR evaluation and NMS configuration. The returned network status confirms that the new service is active and that the required transmission performance is satisfied. During the execution, the Dynamic Operator follows the plan generated by the Operation Agent while checking the returned network state after each action. In the demonstrated trace, NMS telemetry is first retrieved to determine the current service and route state, followed by service release and provisioning requests. The resulting configuration status and predicted GSNR are then checked to verify that the modified network remains operational. Finally, the Dynamic Operator returns the execution status and updated network state to the Operation Agent for further coordination. Additionally, Dynamic Operator assists in service quality assurance, establishing traffic scheduling and priority management to ensure that network resources are allocated based on predefined service level agreement (SLA), which guarantees the real-time and emergency communications. By enforcing intelligent scheduling policies, Dynamic Operator can maintain the required quality of service (QoS) levels across diverse scenarios. Furthermore, Dynamic Operator supports software and system upgrades, ensuring smooth transitions without service disruption. The Automated Inspector is the real-time monitoring and
diagnostics Agent in optical networks, dedicated to ensuring long-term stability through continuous surveillance, predictive maintenance, and environmental awareness. Unlike the Dynamic Operator, which focuses on active optimization, the Inspector monitors key metrics such as optical power, GSNR, and bandwidth utilization to detect QoT fluctuations and anomalies. Integrated within NMS, it offers centralized visibility and automated reporting. It performs routine inspections on power flatness and spectrum usage to optimize GSNR and spectral efficiency, while also diagnosing physical infrastructure by evaluating fiber attenuation, equipment health, and identifying potential faults. Additionally, it monitors environmental conditions like temperature and humidity, enabling proactive risk mitigation and maintaining resilient network operations. Automated Inspector can autonomously analyze anomalies and generate structured reports for further action. Detected issues are categorized and sent to the Operation Agent to determine the response strategy. If performance tuning is needed, it delegates execution to the Dynamic Operator for real-time adjustments like power tuning or path reconfiguration. The Operation Agent also periodically initiates automated inspections, during which the Inspector retrieves monitoring indicators and methodologies from its knowledge base, invoking tools and APIs for comprehensive checks. As illustrated in Fig. 8, the Automated Inspector receives an inspection request from the Operation Agent, retrieves the required NMS performance indicators and analysis methods, and detects abnormal optical-power conditions from the returned telemetry. The action trace shows an abnormal power condition below the normal operating range. The Inspector reports these observations to the Operation Agent, which then determines whether the issue can be resolved within the current operation domain. For the demonstrated case, the
Research Article
15 Fault
Fiberfailure Software failure
Devicefailure Powerfailure
Fiber degradation Fiber interruption IP Conflict Protocol Mismatch Port Misconfiguration Card failure Optical module failure Switch/Router failure Overheating Power outage failure
Network Management System
Port Status Perform data diagnosis tasks
maintenance personnel
Action
Query historical logs and solutions
RAG
Knowledge Library Alarm information
Expert experience
• • • •
Correlation between alarms Previous processing examples
Out
1. Call alarm system 2. root cause analysis 3. Summary output
× √
LLM
Call Tools Involve APIs and tools for alarm analysis
Tool Kit Interface Tools Alarm processing tools • Query alarm • Alarm compression • Query fault • Root alarm analysis algorithm • Interaction • Provide handling with other agents suggestions
Fault diagnosis report
In
Out
The same processing framework as the other two agents
RAG
In LLM
1. Query performance data 2. Process performance data 3. Provide a conclusion
Knowledge Library
Memory
Expert experience • Previous maintenance cases • Professional Knowledge Handbook
Tool Kit
Query historical logs and solutions
RAG
Knowledge Library Performance indicators information
Interface Tools • Communicate with other Agents • Check the repair progress
• • •
Voltage
Bandwidth utilization
SNR
Delay
Current
Self-iteration Planning
Call Tools
Tx/Rx power Link loss
Maintenance Agent
1. Need to query alarms 2. Determine root alarm 3. Provide a conclusion
Alarm definition • Alarm level Alarm type • Potential impact
Database
Instruction issuance
Self-iteration
Memory
Temperature
BER
Abnormal-data Handler
Alarm-Handler
Perform alarm diagnosis tasks
Planning
LLM
store in
Power fluctuate failure
In
Performance data
influence
generate
Definition of performance indicators Threshold of performance indicators Possible impacts of exceeding performance indicators
Action
Out
1. Call data base 2. Performance analysis 3. Summary output
× √
LLM
Call Tools Involve APIs and tools for data analysis
Tool Kit Interface Tools • Query performance data Algorithms & Models • Performance statistical methods • Performance analysis algorithm • Fault prediction model • Performance prediction model
Fig. 9. Schematic of multi-Agent collaboration in the maintenance phase of the AONs. Faults trigger alarms and cause anomalies
in performance indicators, which are detected by the Alarm Handler and Abnormal-data Handler. By correlating alarms with performance data, the Maintenance Agent identifies the fault and assists maintenance personnel in resolution. A fault diagnosis report is generated once the fault is resolved and the system returns to normal. Operation Agent generates a follow-up optimization request and instructs the Dynamic Operator to retrieve the relevant configuration interfaces, evaluate the optimization strategy, and apply the resulting adjustment. After configuration, the performance is re-evaluated and the updated network state is reported back to the Operation Agent. D. Maintenance phase
The goal of the maintenance phase is to quickly identify and resolve faults, thereby minimizing service interruptions and reducing economic losses as much as possible. With the continuous expansion of optical networks, the number and variety of devices and components are also growing, leading to an increasing overall complexity of the network. In the event of a failure, a large number of alarms are triggered and massive abnormal performance data are generated. Under such a complex environment, timely and accurate fault localization and repair becomes critical for ensuring optical network stability. Firstly, we introduce several common types of faults and the key network indicators that are typically affected when a fault occurs, as shown in Fig. 9. First, one of the most common optical network faults are optical fiber failures, including both gradual fiber degradation and abrupt fiber breaks. Degradation tends to occur over time and can be detected through specific performance indicators. In contrast, fiber breaks are usually sudden, often resulting from external mechanical forces or environmental changes that alter the fiber’s physical properties and rapidly degrade optical transmission. Second, the software malfunctions arise from incompatibilities between software and hardware components, leading to issues such as IP conflicts, protocol mismatches, or port misconfigurations. Third, the equipment malfunctions refer to the performance degradation of devices or modules due to factors such as aging, manufactur-
ing defects, or environmental stress. Typical examples include failures in boards, optical modules, or overheating of network components. Fourth, the power failures frequently occur in data centers, which are caused by power outages, voltage instability, or power distribution faults. These faults often lead to abnormal or fluctuating values in key network performance indicators, directly or indirectly impacting QoS. The network indicators we monitor include temperature, voltage, current, optical power, BER, link loss, SNR, delay, port status, bandwidth, and others. Through telemetry, these indicators reflect the real-time network state and help alert maintenance personnel to potential risks. To support this, threshold-based alarms are used to alert operators when performance metrics exceed predefined limits. Although most alarms today are still threshold-driven, the concept is flexible and can incorporate multi-dimensional conditions rather than relying solely on single-metric triggers. The overall framework of the maintenance phase and the roles of the involved Agents are depicted in Fig. 10. When a fault occurs, the network management system generates alarms, while the fault may also cause abnormal changes in the performance data stored in the database. The Maintenance Agent first assigns diagnostic tasks to two specialized Agents: the Alarm Handler and the Abnormal-data Handler. The Alarm Handler retrieves and correlates alarm lists to distinguish root alarms from derived alarms. Meanwhile, the Abnormal-data Handler retrieves telemetry data and examines deviations from normal performance baselines. The two Agents exchange their intermediate conclusions and jointly narrow down the fault location and possible root cause. These diagnostic activities are supported by a shared knowledge base and tool base. The knowledge base contains alarm definitions and propagation rules, performance baselines and alarm thresholds, device manuals and troubleshooting guides,
Research Article
16
influence
Link Down fault (unknown for us)
Performance data
analyzing
Generate alarms
analyzing
conclusion share
store in
send alarm lists Network Management System
send data Abnormal-data Handler
call for alarm lists Alarm-Handler Task distribution repaire
confirm
Database
Task distribution
Instruct-Agent
confirm generate the fault handling report
Guide Feedback
instruction issuance
Knowledge Base • • • • •
call for data
Data Access Tools
Alarm definitions and propagation rules Performance baselines and alarm thresholds Device manuals and troubleshooting guides Historical fault cases and maintenance records Repair SOPs and safety constraints
• • •
Alarm query Performance telemetry query Topology and service-impact query
Task distribution
call for alarm lists and get feedback
generate the fault handling report
call for data and get feedback
Tool Base Diagnosis and Handling Tools • • •
Alarm correlation and anomaly analysis OTDR and optical-power testing Work-order and report generation
Analyzing and conclusion share Analyzing and conclusion share
Repaire
confirm
Fault Handling Report
Multi-Agent Action Trace [10:02:15] NMS: Link NE-2–NE-3 DOWN [10:02:31] Alarm-Handler: NE-3/Rx: LOS (Critical) Downstream nodes: LOF/AIS alarms detected [10:02:35] Abnormal-data Handler: Rx power: -40.0 dBm (Baseline: -14.7 dBm) Upstream Tx power: +1.6 dBm (Normal) Traffic rate: 100 Gb/s → 0 Gb/s [10:03:02] Shared Conclusion: Possible fiber discontinuity near NE-3 [10:03:22] Instruct-Agent: Issue inspection and repair instructions
confirm instruction issuance
Fault: Link NE-2–NE-3 interruption Root Cause: Lose fiber connector at the NE-3 ODF Evidence • LOS alarm at the receiving port • Rx power dropped to −40.0 dBm • Upstream Tx power remained normal Repair Action: Connector inspected, cleaned, and reconnected Recovery Result • Rx power recovered to −14.8 dBm • LOS/LOF alarms cleared • 100-Gb/s service restored
Fig. 10. Illustration of Agent collaboration in a specific fault diagnosis and recovery scenario during the maintenance phase. The
figure demonstrates the whole process, including fault triggering, alarm and performance anomaly detection, Agent-driven root cause analysis, on-site repair guidance, and final result consolidation. historical fault cases and maintenance records, and repair SOPs and safety constraints. The tools are divided into two groups. Data access tools provide alarm, performance telemetry, topology, and service-impact queries. Diagnosis and handling tools support alarm correlation and anomaly analysis, OTDR and optical-power testing, and work-order and report generation. After receiving the diagnostic results, the Maintenance Agent checks the consistency of the conclusions and determines the appropriate repair procedure based on the retrieved manuals, historical cases, and operational constraints. It then sends the fault location and repair instructions to field maintenance personnel. After the repair is completed, the Maintenance Agent coordinates with the Alarm Handler and Abnormal-data Handler to verify that the alarms have cleared and that the performance indicators have returned to their normal ranges. Finally, it consolidates the diagnostic evidence, repair actions, and recovery results into a structured fault-handling report. The conclusionsharing links in the figure represent inter-Agent communication, which can be implemented through A2A-compatible interfaces. Next, we illustrate the fault localization and recovery process through the representative link-down case shown in Fig. 10. At 10:02:15, the network management system reports that the link between NE-2 and NE-3 is down. The Maintenance Agent responds by assigning diagnostic tasks to the Alarm Handler and Abnormal-data Handler. The Alarm Handler queries the network management system and identifies a critical loss-ofsignal (LOS) alarm at the receiving port of NE-3, together with downstream loss-of-frame (LOF) and alarm-indication-signal (AIS) alarms. In parallel, the Abnormal-data Handler retrieves the relevant performance data. It finds that the received opti-
cal power has dropped from a baseline of approximately -14.7 dBm to -40.0 dBm, while the upstream transmit power remains normal at approximately +1.6 dBm. The traffic rate has also decreased from 100 Gb/s to zero. By correlating the alarm and performance evidence, the two diagnostic Agents infer that the failure is likely caused by a fiber discontinuity near NE-3. The Maintenance Agent then consults the relevant troubleshooting procedures and issues inspection and repair instructions to the field personnel. The on-site inspection confirms that a loose fiber connector at the NE-3 optical distribution frame (ODF) caused the interruption. The connector is inspected, cleaned, and reconnected. After the repair, the received optical power recovers to approximately -14.8 dBm, the LOS and LOF alarms are cleared, and the 100 Gb/s service is restored. Finally, the Maintenance Agent records the root cause, diagnostic evidence, repair action, and recovery status in a structured fault-handling report. E. Upgrade phase
The ever-growing network traffic necessitates continuous upgrades in optical networks. During routine network management, the Upgrade Agent analyzes real-time data and historical trends to predict capacity saturation at different parts of networks. When C-band utilization approaches 80%, the Upgrade Agent generate upgrade reports to operator, cross-referencing external context such as SLAs. Meanwhile, the operator is also monitoring such links whose capacity is approaching its limit. To prevent potential failures and accommodate new traffic, the operator can instruct the Upgrade Agent to promote this link to C+L-band transmission, as illustrated in Fig. 11.
Research Article
17
Observation: Traffic between Node 1 and Node 2 continues to grow, more capacity outside the C-band required. Requirements: The upgraded transmission system should cover L-band, ultimate 6THz bandwidth and additional 60 channels are expected, the budget is around 1M. Task breakdown: 1. Pre-calculate the device performance baseline for the link. 2. Push the baselines to Implementation & Validation Agent for device selection.
Upgrade Agent
Upgrade to L-band!
Orders Traffic perception
Operator
C-band
C-band
①
5. Negotiate the upgrade time window with Business Management Agent, who provides a service priority list. 6. Push upgrade time to Implementation & Validation Agent, guide on-site personnel to install device and configuration.
C-band
Upgrade Planner
Freq
Node2 … Validation Specialist
device
Self-iteration
The gain of EDFA should >18dB. Output power is predefined 1.6dBm/ch.
Memory
Query upgrade device information
Knowledge Library Equipment information • Vendor and costs Transponder information • Speed and formats Amplifier sites
Service migration from C to L-band
Performance Optimizer
③
Action
Devices meet baselines. Giving devices install and configuration guides.
RAG
New service upload
④
Self-iteration Planning
In
Service level list
L-band
Business Manager
…
②
Upgrade time window
⑤
Node1
3. Push the list to Performance Optimization Agent for configurations and margin. 4. Feedback to Implementation & Validation Agent for possible device list updates.
?
Upgrade
?
?
New business established
Planning
Out
Call Tools
Tool Kit Equipment selection • Vendor and Supply chain health analysis • Performance evaluation • Cost and installation guide
Calculate the performance of the link from the device list.
RAG
Action
Invoke GN model for NLI calculation.
Memory
Upgrade optimization info.
Knowledge Library Topology • Link connection • Spectrum grid • Power range
Call Tools
Tool Kit Spectrum allocation GN/ SRS model Three-step optimization
Fig. 11. Schematic of multi-Agent collaboration in the upgrade phase of the AONs. The figure depicts how the Upgrade Agent predicts capacity saturation and coordinates with sub-Agents, including the Upgrade Planner, Validation Specialist, Performance Optimizer, and Business Manager, to carry out a C-to-C+L band upgrade. It visualizes each Agent’s role in planning, device validation, performance optimization, and scheduling, ensuring a seamless and efficient network upgrade process.
Upon receiving the upgrade order, the Upgrade Agent summarizes the upgrade requirements, including bandwidth expansion, capacity needs, budget constraints, and other relevant factors. These requirements are then forwarded to the sub-Agent Upgrade Planner, which generates a step-by-step upgrade plan. Each step invokes specialized sub-Agents to execute specific tasks efficiently. In the upgrade from C-band to C+L-band system, the Upgrade Planner first determines the necessary devices and performance baselines for each device. For example, the output power of an L-band transponder should be around 0 dBm, while the saturation power of an L-band amplifier must exceed 23 dBm. These baselines are then passed to the Validation Specialist, which selects appropriate devices based on its infrastructure knowledge, containing records of available devices of different types, alongside vendor documentation validated against standards like ITU-T. Within the Validation Specialist, a structured decisionmaking cycle is followed, consisting of planning, action, and the memory to provide necessary information. Once the required devices are selected, the Performance Optimizer is activated to calculate the optimal configurations for the upgraded L-band transmission link. This sub-Agent also follows a similar cycle, where it breaks down the task into sub-steps and formulates a CoT process to determine the best approach. Different actions are selected based on computational tools from the memory, and results are continuously analyzed for further optimization. Moreover, there is self-iteration of derived strategy from the outputs to the inputs. For example, when the L-band devices are selected, the Performance Optimizer can invoke the three-step QoT optimization algorithm [71] to find optimal configurations for the C- and L-band EDFAs. As the C-band transmission performance will also be affected by the upgraded L-band devices and traffic, fast optimization tools are necessary to mitigate transient interference during the upgrade.
After determining the optimal configurations, all relevant information is transferred to the Business Manager, which schedules an appropriate upgrade time window to minimize network disruptions. Additionally, the Business Manager ensures that critical services maintain high performance and that C-band optical performance remains within acceptable limits. It also identifies new services and existing services that need to be migrated to the L-band. Once these steps are completed, the C+L-band transmission system is successfully upgraded, ensuring seamless scalability and uninterrupted service delivery. F. Decommission phase
With the emergence of new services, advancements in communication technologies, and shifts in communication capacity distribution across different nodes, parts of the transmission system may need to be decommissioned, or specific links with minimal services may be removed. Taking link decommissioning as an example, potential decommissioning candidates are first identified during the operational phase based on traffic and service conditions. Once a decommissioning request is initiated, the upper-level orchestration layer activates the Decommission Agent, which then continuously monitors the relevant traffic and service status within the optical network, as illustrated in Fig. 12. For an identified candidate, the Decommission Agent evaluates the SLA and generates a comprehensive report for the operator. This report assesses the necessity and feasibility of decommissioning, considering factors such as the importance of remaining services, operational costs, and potential savings from decommissioning. Once the report is submitted, the operator makes the final decision on whether to proceed with decommissioning. Once the decision is confirmed, the Decommission Agent forwards the decommissioning request to the Retirement Planner, which then breaks down the process into logical, feasible steps, delegating specialized tasks to sub-Agents for efficient
Research Article
18
Observation: Traffic between Node 1 and Node 2 continues to decrease, most capacity of this link is wasted.
Remove the link!
Decommission Orders Agent Traffic perception
Requirements: Decommissioning this link and switch the remained services to another route.
Operator
Li n
Reduced traffic
Task breakdown:
Freq
Node3
Node4
Retirement Planner
3. Push the decommissioning plan to Service Switch Agent for coordinate the switch time windows.
5. Push to Physical Decommission Agent and perform physical operations, including fiber disconnection, device power-off, and asset recovery.
! sionNode2
Node1
2. Develop a decommissioning plan, including hardware removal, spectrum release and find service switch routes.
?
m is
Freq
1. Generate a report on the feasibility of link decommission and evaluate the impact of decommission on redundancy.
4. Perform service switching to alternative route, including rerouting configuration, and verify post-switch performance (QoT, latency).
k
m de c o
Self-iteration
Planning In
RAG
Self-iteration Action
I need to develop a decommissioning plan, should find new route, …
Device Manager
Service Switcher
Find a route meet Out spectrum allocation, capacity, xxx.
Call Tools
Memory
Decide switch time, service priorities, and interruption time
Memory
Query decommission information Knowledge Library
Action
Planning I need to switch remaining services to a new route, considering SLA…
Query decommission information Tool Kit Knowledge Library
• Network redundancy • Topology and RSA • Power and electric
• QoT prediction • RSA problem
• Service level agreement (SLA) • Network plane
Tool Kit • QoT predicter • Latency computer
Fig. 12. Schematic of multi-Agent collaboration in the decommission phase of the AONs. The figure depicts how the Decommis-
sion Agent monitors traffic and initiates decommissioning workflows when link utilization drops. It highlights the roles of the Retirement Planner, Service Switcher, and Device Manager in evaluating SLA impact, rerouting services, coordinating switching windows, and physically decommissioning network elements, ensuring minimal disruption and efficient resource reclamation. execution. First, the Retirement Planner generates a detailed report assessing alternative routing capacity, SLA compliance, and the impact of decommissioning on network redundancy to ensure that decommissioning does not compromise service quality and that sufficient backup routes exist. It then formulates a structured decommissioning plan covering hardware removal/retention, spectrum release, and service rerouting. Following a structured cycle, planning, action, and self-iteration, the plan is further broken down into sub-tasks using CoT reasoning. The Agent identifies suitable alternative service routes using network knowledge and computational tools, ensuring these routes provide adequate capacity, available spectrum slots, and compliance with latency constraints. Once the plan is developed, it is pushed to the Service Switcher for implementation. The Service Switcher coordinates the switching time window to ensure minimal network disruption, reroutes services to alternative paths, and conducts post-switch performance validation, including QoT and latency checks. After successful service switching, the process moves to the Device Manager, which executes the physical decommissioning operations, including fiber disconnection, device power-off, and asset recovery management, such as optical module removal and inventory registration. After the retirement actions and postdecommissioning validation are completed, the Decommission Agent concludes the active decommissioning process and becomes inactive until another decommissioning task is initiated. Upon completion of these steps, the link decommissioning process is finalized, ensuring efficient resource optimization while maintaining network stability.
6. OUTLOOK Despite the promising potential of the multi-Agent technique in AON, there are still challenges that needs to be addressed before
the full implementation in practice. A. Hallucination risks
Hallucination [72, 73] refers to cases where LLMs generate fluent and coherent outputs that are factually incorrect, logically inconsistent, or contextually irrelevant. Hallucinations may arise from insufficient domain knowledge during pre-training, noisy data correlations, or inference-stage factors such as ambiguous prompts and limited context windows [74]. This issue poses a major challenge for the reliable deployment of LLM Agents in optical networks, where high accuracy and faithful consistency are essential. In optical networks, hallucinations often manifest in the form of incorrect parameter suggestions, nonexistent device models, or logically flawed configuration schemes [75]. Such errors are particularly critical in tasks such as QoT estimation, failure localization, and control plane automation, where precision and consistency are paramount. Mitigating hallucinations requires domain-aware validation and system-level safeguards rather than relying solely on model improvements [76]. In the proposed framework, generated results are first supervised within the corresponding lifecycle division. The Division Agent evaluates the outputs of its subordinate AI Experts against task objectives, network constraints, and expected quality, and triggers refinement and re-execution when necessary. This local closed loop continues until an acceptable result is obtained or a predefined iteration limit is reached, after which unresolved tasks are escalated to the upper level or human operators. For actions that may affect the physical network, the Security Manager provides an additional safety gate by checking generated configurations or commands against operational and security constraints. Generated decisions should be verified against network state, topology, and operational policies before execution. In particular, DTON can serve as a validation sandbox to evaluate generated strategies and configurations against
Research Article
physical-layer constraints and real-time network conditions before deployment. High-risk or service-affecting actions can further require human-in-the-loop confirmation before execution. After deployment, network states and service performance are continuously monitored, and corrective actions or rollback to a previously validated configuration can be initiated if unacceptable outcomes are detected. Besides, hallucination mitigation should be regarded as a joint responsibility of both LLM Agents and deterministic network infrastructures to ensure trustworthy operation. B. Model selection and deployment
In practical deployment, the hierarchical multi-Agent framework requires that Agents can be adapted to the complexity, latency requirement, data sensitivity, and computational demand of individual tasks. High-level Agents, for example, the Network Director, are responsible for global objective interpretation, cross-domain information integration, long-horizon planning, task decomposition, and coordination among multiple Agents. Cloud infrastructure can host large models and shared knowledge services for computationally intensive tasks. In contrast, Agents for operation, monitoring, security, and other latencysensitive tasks can be smaller language models, deployed at the edge or locally at network sites, particularly when rapid responses or local processing of operational data are required. A hierarchical escalation mechanism can also be adopted, where a lightweight local model handles routine requests and invokes a more capable remote model only when the task is ambiguous, complex, or high-risk. Moreover, considering that some tasks require faster, more accurate, and complex calculations, we design and recommend using workflows with tools instead of calling LLM to accelerate task implementation and ensure accuracy. C. Token costs
Token length constraints present a fundamental limitation for applying LLMs in optical networks, where decision-making often relies on large volumes of structured and unstructured data. Practical tasks may require the model to simultaneously process multi-span transmission parameters, per-channel configurations, network topologies, historical monitoring records, and real-time telemetry data, leading to rapidly increasing context size and token consumption as network scale grows. This challenge becomes more severe in full LCM tasks, where decisions frequently depend on information accumulated across multiple stages. Addressing this issue requires efficient context management [77], including hierarchical information abstraction, skill-based function calling, retrieval-based context selection, and memory-augmented architectures. In our proposed multi-Agent framework, this issue can be mitigated through the dedicated Memory Container, which serves as an intermediate layer for context compression and management. By leveraging efficient token compression, structured information encoding, and selective retrieval mechanisms, the Memory Container can distill large-scale network data into compact and task-relevant representations, thereby reducing token consumption while maintaining essential global context. The context requirements are also closely coupled with model size and computational resource allocation. D. Long-horizon memory
Long-horizon memory and stateful reasoning remain critical challenges for applying LLMs to optical network LCM, where
19
decisions spans extended time horizons, and depend on evolving network states, historical configurations, and accumulated performance data. However, conventional LLMs operate in a largely stateless manner, relying only on the current context window and lacking persistent awareness of prior interactions or long-term system evolution. This limitation is particularly evident in sequential tasks such as fault diagnosis, dynamic reconfiguration, and upgrade planning, where earlier actions directly influence subsequent decisions. Addressing this challenge requires external memory and state management mechanisms, including structured repositories, memory-augmented architectures, and DT-assisted synchronization. In our framework, the Memory Container further acts as a persistent state management module that supports storage, update, and retrieval of network state information across different lifecycle stages. By maintaining evolving network states, it enables temporally consistent and context-aware reasoning over long operational horizons. In addition, persistent network knowledge and short-term task contexts can be managed separately so that only the information required for the current reasoning step is provided to the LLM. The context requirements are also closely coupled with model size and computational resource allocation, with more demanding or high-priority tasks receiving greater computing resources and lightweight tasks being handled by smaller models or edge resources. Such coordinated management of context, model scale, and computing resources can reduce context and computational costs while maintaining the responsiveness and reasoning capability required for optical network lifecycle management.
7. CONCLUSION This paper has systematically envisioned the LLM Agent-driven evolution of optical networks toward AONs. First, the technical foundations of LLM and Agent are illustrated, with particular emphasis on domain adaptation techniques and performance evaluation considerations in optical networking scenarios. From a system-level perspective, this work provided a structured decomposition of the optical network lifecycle into six key phases, and analyzed the associated tasks, data requirements, and operational constraints in each stage. Based on these insights, we presented a hierarchical multi-Agent framework that organizes heterogeneous Agents across different functional layers, enabling coordinated perception, reasoning, and action. This framework offers a unified perspective for integrating LLMdriven intelligence into optical network management, bridging the gap between high-level decision-making and low-level operational execution. Beyond architectural design, an important takeaway of this work lies in highlighting the role of LLM Agent as a potential enabler of closed-loop, adaptive, and context-aware control in optical networks. By coupling LLM reasoning capabilities with domain knowledge, real-time data, and external tools, Agent systems can move toward more flexible and scalable management strategies compared to conventional pipeline-based solutions. In evolution of optical networks toward AONs, this work aims to present a structured and forward-looking perspective on leveraging LLM Agent for the autonomous LCM. Rather than offering a definitive solution, it seeks to outline a conceptual architecture and provide technical insights for researchers and developers to explore the evolving landscape of AONs. It is expected that continued efforts in model development, system integration, and real-world validation will be necessary to fully realize the
Research Article
potential of AONs.
FUNDING National Natural Science Foundation of China (62522104 and U24B20133).
REFERENCES 1.
2. 3. 4. 5.
6.
7.
8. 9. 10. 11. 12.
13.
14.
15. 16.
17. 18. 19.
20.
21.
22.
23.
Y. Song et al., “Lifecycle management of optical networks with dynamicupdating digital twin: a hybrid data-driven and physics-informed approach,” IEEE J. on Sel. Areas Commun. (2025). M. Lam and J. van der Lande, “Quantifying the benefits of optical network automation,” Tech. rep., Analysys Mason (2024). D. Wang et al., “A review of machine learning-based failure management in optical networks,” Sci. China Inf. Sci. 65, 211302 (2022). R. Gu et al., “Machine learning for intelligent optical networks: A comprehensive survey,” J. Netw. Comput. Appl. 157, 102576 (2020). F. Musumeci et al., “An overview on application of machine learning techniques in optical networks,” IEEE Commun. Surv. & Tutorials 21, 1383–1408 (2018). J. Lu et al., “Performance comparisons between machine learning and analytical models for quality of transmission estimation in wavelengthdivision-multiplexed systems,” J. Opt. Commun. Netw. 13, B35–B44 (2021). D. Wang and M. Zhang, “Artificial intelligence in optical communications: from machine learning to deep learning,” Front. Commun. Networks 2, 656786 (2021). X. Sui et al., “A review of optical neural networks,” IEEE Access 8, 70773–70783 (2020). Y. Cao et al., “A comprehensive survey of ai-generated content (aigc): A history of generative ai from gan to chatgpt,” arXiv:230304226 (2023). L. Bariah et al., “Large generative ai models for telecom: The next big thing?” IEEE Commun. Mag. 62, 84–90 (2024). D. Wang et al., “When large language models meet optical networks: paving the way for automation,” Electronics. 13, 2529 (2024). Y. Song et al., “Synergistic interplay of large language model and digital twin for autonomous optical networks: Field demonstrations,” IEEE Commun. Mag. (2025). C. Sun et al., “Experimental demonstration of local ai-agents for lifecycle management and control automation of optical networks,” J. Opt. Commun. Netw. 17, C82–C92 (2025). X. Liu et al., “First field trial of llm-powered ai agent for lifecycle management of autonomous driving optical networks,” in OFC, (2025), pp. Th1A–2. D. Wang et al., “Large language model for optical network automation: Prospects and challenges,” in OFC, (2025), pp. Th1A–4. Y. Zhang et al., “Ai agent for autonomous optical networks: architectures, technologies, and prospects [invited tutorial],” J. Opt. Commun. Netw. 18, A159–A178 (2026). S. Cruzes, “Revolutionizing optical networks: The integration and impact of large language models,” Authorea Prepr. (2024). J. Wang, “Generative ai for network operations,” in OFC, (2025), p. Workshop. Y. Wang et al., “Alarmgpt: an intelligent alarm analyzer for optical networks using a generative pre-trained transformer,” J. Opt. Commun. Netw. 16, 681–694 (2024). Y. Pang et al., “Large language model-based optical network log analysis using llama2 with instruction tuning,” J. Opt. Commun. Netw. 16, 1116–1132 (2024). Y. Zhang et al., “Gpt-enabled digital twin assistant for multi-task cooperative management in autonomous optical network,” in OFC, (Optica Publishing Group, 2024), pp. Th1G–4. Y. Weijie et al., “Spatio-temporal knowledge graph with large language model for comprehensive fault management in optical networks,” in ECOC, (2024), pp. W2A–102. N. D. Cicco et al., “Open implementation of a large language model pipeline for automated configuration of software-defined optical networks,” in ECOC, (2024), pp. W3E–1.
20
24. A. Zhou et al., “Large language model-driven ai agent in sdn controller towards intent-based management of optical networks,” in ECOC, (2024), pp. W3E–2. 25. C. Wang et al., “Llm-centric transport network configuration management framework and demonstration,” in OFC, (2025), pp. W3E–4. 26. X. Jiang et al., “Opticomm-gpt: a gpt-based versatile research assistant for optical fiber communication systems,” Opt. Express 32, 20776– 20796 (2024). 27. C. Sun et al., “First experimental demonstration of full lifecycle automation of optical network through fine-tuned llm and digital twin,” in ECOC, (2024), pp. Th3B–6. 28. A. Abishek et al., “End-to-end transport network digital twins with cloudnative SDN controllers and generative AI,” J. Opt. Commun. Netw. 17, C70–C81 (2025). 29. Y. Zhang et al., “First field-trial demonstration of l4 autonomous optical network for distributed ai training communication: An llm-powered multi-ai-agent solution,” arXiv:250401234 (2025). 30. H. Huang et al., “Field trial of llm-based autonomous network management with ai-agent in real-time 400g/800g elastic optical network,” in 2025 ECOC, (IEEE, 2025), pp. 1–4. 31. Q. Qiu et al., “Expertise-guided llm agent realizing autonomous optical power optimization in field-deployed networks,” in 2025 ECOC, (IEEE, 2025), pp. 1–4. 32. Y. Zhang et al., “Design and evaluation of an llm-based agent for qot estimation and performance optimization in optical networks,” IEEE Open J. Commun. Soc. (2025). 33. Y. Hao et al., “Multi-agent llm-powered ai for autonomous optical power commissioning of oms links,” in 2025 ECOC, (IEEE, 2025), pp. 1–4. 34. Y. Zhang et al., “Generative ai-driven hierarchical multi-agent framework for zero-touch optical networks,” IEEE Commun. Mag. (2025). 35. S. Xiang et al., “Optima: Collaborative multi-agent framework for modelling and controlling raman amplifier in intelligent optical networks,” in 2025 ECOC, (IEEE, 2025), pp. 1–4. 36. V. Ricard et al., “Applying digital twins to optical networks with cloudnative sdn controllers and generative ai,” in ECOC, (2024), pp. M2E–4. 37. H. Arunachalam et al., “Generative ai for optical networking: From design automation to self-operating networks,” in 2026 OFC, (Optica Publishing Group, 2026), pp. 1–3. 38. E. Ayush Maheshwari, “Autonomous networks and agentic ai – shaping the future of connectivity,” in 2026 OFC, (Optica Publishing Group, 2026), pp. 1–3. 39. A. Vaswani et al., “Attention is all you need,” Adv. neural information processing systems 30 (2017). 40. D. Bahdanau et al., “Neural machine translation by jointly learning to align and translate,” arXiv:14090473 (2014). 41. A. Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv:201011929 (2020). 42. G. Dong et al., “Baichuanseed: Sharing the potential of extensive data collection and deduplication by introducing a competitive large language model baseline,” arXiv:240815079 (2024). 43. W. X. Zhao et al., “A survey of large language models,” arXiv:230318223 1 (2023). 44. G. I. Meadows et al., “Localvaluebench: A collaboratively built and extensible benchmark for evaluating localized value alignment and ethical safety in large language models,” arXiv:240801460 (2024). 45. G. Penedo et al., “The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only,” arXiv:230601116 (2023). 46. A. Radford et al., “Improving language understanding by generative pre-training,” OpenAI Blog (2018). 47. H. Touvron et al., “Llama: Open and efficient foundation language models,” arXiv:230213971 (2023). 48. J. Devlin et al., “Bert: Bidirectional encoder representations from transformers,” arXiv:181004805 p. 15 (2018). 49. A. Radford et al., “Language models are unsupervised multitask learners,” OpenAI blog 1, 9 (2019). 50. Z. Han et al., “Parameter-efficient fine-tuning for large models: A comprehensive survey,” arXiv:240314608 (2024).
Research Article
51. L. Wang et al., “Parameter-efficient fine-tuning in large models: A survey of methodologies,” arXiv:241019878 (2024). 52. X. Wang et al., “Lora ensembles for large language model fine-tuning,” arXiv:231000035 (2023). 53. T. Dettmers et al., “Qlora: Efficient finetuning of quantized llms,” Adv. neural information processing systems 36, 10088–10115 (2023). 54. L. Ouyang et al., “Training language models to follow instructions with human feedback,” Adv. neural information processing systems 35, 27730–27744 (2022). 55. J. Achiam et al., “Gpt-4 technical report,” arXiv:230308774 (2023). 56. Y. Liu et al., “Developing a domain-specific llm for optical networks: A reinforcement learning-based fine-tuning framework,” IEEE Transactions on Netw. Serv. Manag. 23, 3655–3678 (2026). 57. S. Zhao et al., “Retrieval augmented generation (rag) and beyond: A comprehensive survey on how to make your llms use external data more wisely,” arXiv:240914924 (2024). 58. C. Jeong, “A study on the implementation method of an agent-based advanced rag system using graph,” arXiv:240719994 (2024). 59. A. N. T. Dieu et al., “The enhanced context for ai-generated learning advisors with advanced rag,” in 2024 18th International Conference on Advanced Computing and Analytics (ACOMPA), (2024), pp. 94–101. 60. P. Liu et al., “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM computing surveys 55, 1–35 (2023). 61. L. Giray, “Prompt engineering with chatgpt: a guide for academic writers,” Annals biomedical engineering 51, 2629–2633 (2023). 62. J. Wei et al., “Chain-of-thought prompting elicits reasoning in large language models,” Adv. neural information processing systems 35, 24824–24837 (2022). 63. F. Amigoni et al., “Anthropic agency: a multiagent system for physiological processes,” Artif. Intell. Medicine 27, 305–334 (2003). 64. S. Hong et al., “Metagpt: Meta programming for multi-agent collaborative framework,” arXiv:230800352 3, 6 (2023). 65. Q. Wu et al., “Autogen: Enabling next-gen llm applications via multiagent conversations,” in First conference on language modeling, (2024). 66. R. Ramaswami and K. Sivarajan, Optical networks: a practical perspective (Elsevier, 2001). 67. Q. Zhuge et al., “Building a digital twin for intelligent optical networks [invited tutorial],” J. Opt. Commun. Netw. 15, C242–C262 (2023). 68. Y. Pointurier, “Design of low-margin optical networks,” J. Opt. Commun. Netw. 9, A9–A17 (2016). 69. Y. Takita et al., “Towards seamless service migration in network reoptimization for optically interconnected datacenters,” Opt. Switch. Netw. 23, 241–249 (2017). 70. Google, “Agent2agent (a2a) protocol,” https://github.com/google/A2A (2025). 71. Y. Song et al., “Efficient three-step amplifier configuration algorithm for dynamic c+ l-band links in presence of stimulated raman scattering,” J. Light. Technol. 41, 1445–1453 (2022). 72. H. Ye et al., “Cognitive mirage: A review of hallucinations in large language models,” arXiv:230906794 (2023). 73. L. Huang et al., “A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,” ACM Transactions on Inf. Syst. 43, 1–55 (2025). 74. V. Adlakha et al., “Evaluating correctness and faithfulness of instructionfollowing models for question answering,” Transactions Assoc. for Comput. Linguist. 12, 681–699 (2024). 75. X. Wang et al., “Wireless hallucination in generative ai-enabled communications: Concepts, issues, and solutions,” arXiv:250306149 (2025). 76. S. Tonmoy et al., “A comprehensive survey of hallucination mitigation techniques in large language models,” arXiv:240101313 6 (2024). 77. Z. Zhang et al., “A survey on the memory mechanism of large language model-based agents,” ACM Transactions on Inf. Syst. 43, 1–47 (2025).
21