ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

Foreign object detection in power transmission lines using SESYOLO.

Duan P et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed systems architecture

Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Sci Rep . 2026 Mar 3;16:11807. doi: 10.1038/s41598-026-41080-7 Search in PMC Search in PubMed View in NLM Catalog Add to search Foreign object detection in power transmission lines using SESYOLO Pingting Duan Pingting Duan 1 Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE, Minzu University of China, Beijing, 100081 China 2 School of Information Engineering, Minzu University of China, Beijing, 100081 China Find articles by Pingting Duan 1, 2, # , Xuran Zhang Xuran Zhang 3 School of Information Science and Technology, Fudan University, Shanghai, 200438 China Find articles by Xuran Zhang 3 , Xiao Liang Xiao Liang 1 Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE, Minzu University of China, Beijing, 100081 China 2 School of Information Engineering, Minzu University of China, Beijing, 100081 China 4 Institute of National Security, Minzu University of China, Beijing, China Find articles by Xiao Liang 1, 2, 4, ✉ , Yuanjia Cui Yuanjia Cui 5 Huairou District Bureau of Economy and Information Technology of Beijing Municipality, Beijing, 101499 China Find articles by Yuanjia Cui 5 , Hetian Wang Hetian Wang 6 BII Transportation Technology(Beijing) Co., Ltd., Beijing, 100101 China Find articles by Hetian Wang 6 Author information Article notes Copyright and License information 1 Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE, Minzu University of China, Beijing, 100081 China 2 School of Information Engineering, Minzu University of China, Beijing, 100081 China 3 School of Information Science and Technology, Fudan University, Shanghai, 200438 China 4 Institute of National Security, Minzu University of China, Beijing, China 5 Huairou District Bureau of Economy and Information Technology of Beijing Municipality, Beijing, 101499 China 6 BII Transportation Technology(Beijing) Co., Ltd., Beijing, 100101 China ✉ Corresponding author. # Contributed equally. Received 2024 May 8; Accepted 2026 Feb 17; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13066553  PMID: 41775836 Abstract The complex environment of power transmission lines renders foreign object attachment to electrical equipment a frequent cause of faults. However, existing object detection frameworks struggle to accurately identify the types of foreign objects. This manuscript aims to address the scarcity of samples of foreign objects on power transmission lines and develop a high-precision, low-latency detection algorithm. Artificial Intelligence Generated Content (AIGC) is employed to generate high-quality images, providing abundant training data. The newly introduced Spatial and Channel Reconstruction Convolution (SCConv) reduces redundant computations and promotes representative feature learning. Additionally, the reconstructed Efficient Reparameterized Generalized-FPN (Efficient RepGFPN) effectively exchanges high-level semantic information and low-level spatial information without adding extra computational burden, which is advantageous for handling foreign objects of varying sizes. The newly designed Squeeze and Excitation Detect (SE-Detect) enables the extraction of richer feature information with fewer parameters. These enhancements are specifically designed to improve the detection of small and irregular foreign objects under complex background interference, which is a common challenge in aerial inspection scenarios. WIoU with better balance sample quality is selected as the loss function for training. Finally, a distillation schema is introduced to improve performance to a higher level. In conclusion, the improved model achieves a 9% increase in the [email protected] and a 9.1% improvement in the recall rate when compared to YOLOv8. Particularly noteworthy is the [email protected] of 93.9% in bird’s nest detection. These results confirm the algorithm’s effectiveness in scenarios involving small targets and cluttered visual environments. Keywords: Power Transmission Lines, Foreign Object Detection, YOLOv8, Feature Fusion Subject terms: Computer science, Scientific data, Computational science Introduction The power industry is a crucial part of the national economy, especially the stable operation of transmission lines, which is vital for the quality and safety of electricity transmission in the power grid 1 , 2 . The environment around transmission lines is complex, often affected by foreign objects such as bird nests, plastic trash, kites and balloons. These objects can cause short circuits, leading to faults in the power system, and pose economic and safety risks 3 , 4 . According to data from the State Grid, incidents caused by these foreign objects rank second only to lightning and external force incidents among all causes of power outages 5 . Therefore, timely detection and handling of these foreign objects are crucial for ensuring the safety of transmission lines. However, traditional manual inspections are inefficient and fail to meet the demands of smart grids 6 . Therefore, utilizing drones for inspecting power transmission lines and advancing intelligent inspection technology not only saves time and resources but also enhances detection efficiency, offering greater advantages over manual inspection 7 – 10 . Recent lightweight CNN breakthroughs show significant promise for resource-limited industrial vision applications. Examples include: efficient semantic feature extraction enabling high-accuracy detection of ancient mural elements with low computational cost 11 ; hybrid attention mechanisms achieving 92.9% precision in real-time instrument monitoring 12 ; and specialized lightweight CNNs attaining 99.6% accuracy in biometric recognition at 16.5 ms latency 13 . These developments collectively demonstrate the viability of optimized deep learning for industrial inspection systems. Although recent advances in deep learning have significantly improved object detection performance, existing detectors such as YOLOv5, YOLOv8, and Faster R-CNN still face key limitations in power transmission line inspection tasks. First, many foreign objects–such as bird nests, balloons, and kites–are small in size and irregular in shape, making them difficult to detect after the typical downsampling operations used in convolutional networks. Second, the inspection scenes captured by UAVs often include complex structured backgrounds (e.g., cables, towers, and vegetation), which can confuse detectors and lead to a high false-positive rate. Finally, public datasets on such foreign objects are scarce and imbalanced, complicating training and leading to poor generalization. To overcome these challenges, we propose SESYOLO, an enhanced YOLO-based detector tailored for real-time foreign object detection under aerial inspection constraints. Our model incorporates targeted modules, including SCConv for spatial-channel enhancement, Efficient RepGFPN for feature fusion, SE-Detect for lightweight attention, WIoU for robust loss weighting, and a distillation mechanism to boost accuracy without increasing inference cost. This study successfully constructs the FOD24 dataset, containing 2817 images, through techniques such as Hue Saturation Value (HSV) enhancement, random blur and noise addition, and simulated weather conditions. Additionally, utilizing the AIGC platform effectively addressed the issue of sample imbalance. This dataset provides comprehensive and authentic data support for foreign object detection on power transmission lines. SCConv is introduced into the backbone network, optimizing convolution operations and integrating spatial and channel information to enhance the model’s feature extraction capabilities, thus providing richer information for feature fusion. The neck network is revamped using Efficient RepGFPN, enabling efficient multi-scale feature fusion and optimizing feature interaction. This reduces upsampling operations, cutting computational overhead and model latency, ultimately boosting target detection accuracy and real-time performance. The Squeeze and Excitation attention mechanism (SE) is incorporated to construct the SE-Detect head, modeling and recalibrating the relationships between feature channels to enhance detection performance without increasing computational load. The bounding box regression loss function incorporates the WIoU v3 mechanism, using an innovative dynamic non-monotonic strategy to optimize gradient gain distribution, successfully enhancing the model’s localization precision and generalization ability. Employing a distillation strategy that integrates feature, response, and relational knowledge enhances detection accuracy by enabling the student network to extract more information from the teacher network. Related work Researchers employ conventional algorithms for the detection of foreign object intrusions in power transmission lines. C. Chen et al. utilize adaptive filtering techniques to eliminate terrain interference and employ Euclidean clustering algorithms to segment the detection results, thereby constructing a recognition model for foreign object intrusions in power lines 14 . L. Cheng et al. introduce an algorithm based on distance estimation focused on accurately calculating the coordinates of foreign objects 15 . Concurrently, S. Jiao et al. integrate the Euclidean distance method with a predictive region drift strategy to enhance the accuracy of foreign object detection 16 . Despite these advancements, traditional algorithms have limitations in recognizing types of foreign objects, noise resistance, and the diversity of target segmentation, increasing the complexity of detection. Consequently, numerous scholars explore the use of machine learning methods such as Multilayer Perceptrons (MLP) and Support Vector Machines (SVM) for the detection of foreign objects. S. Z. Wu et al. propose a cascaded structure based on MLP 17 , while F. Mahdi Elsiddig Haroun et al. enhance the feature set of SVMs using satellite imagery to improve detection efficiency 18 . X. Ye et al. employ a Particle Swarm Optimization-enhanced SVM for detection purposes 19 . Although machine learning models require extensive feature engineering, their data mining capabilities may be inferior to those of deep learning models. However, the superior high-level feature extraction and end-to-end solutions provided by deep learning present new possibilities for the detection of foreign objects in power transmission lines. In the field of deep learning, Liang et al. develop a method for detecting foreign objects in power transmission lines based on the Faster R-CNN framework 20 . Despite the high accuracy of the two-level networks, the processing speed is slow and not suitable for real-time detection. In contrast, single-stage networks such as SSD and the YOLO series demonstrate higher detection speeds and efficiency. Li et al. replace the backbone network of YOLOv3 with Mobilenetv2 to reduce the number of parameters 21 , albeit at the expense of some detection rate. Song et al. enhance performance by integrating k-means clustering and DIoU NMS with YOLOv4 22 , 23 . Huang et al. utilize YOLOv5s in conjunction with Ghost convolution and KL divergence loss 24 , 25 . Liu et al. incorporate attention mechanisms and ASPP modules into YOLOX to improve detection accuracy 26 . Meanwhile, Yu et al. combine hyperparameter optimization and SPD convolution with YOLOv7 to enhance the accuracy of detecting small targets 27 , though this results in slower detection speeds. Given the constraints on computational resources and the necessity for rapid processing in practical deployments, the choice is made to utilize YOLOv8n. This model balances minimal parameter and computational demands with high accuracy and speed. Deep learning approaches rely heavily on large quantities of high-quality data. However, the challenges of class imbalance and data scarcity in the detection of foreign objects in power transmission lines limit both model training and accuracy. Consequently, research shifts towards refining algorithms and enhancing data. Algorithm refinement includes techniques such as zero-shot learning and transfer learning 28 , 29 , while data augmentation methods encompass image enhancement, GAN-based techniques 30 , and generative model approaches 31 Recent object detection models such as YOLOv12, YOLO-MS 32 , RT-DETR 33 , and NanoDet 34 have demonstrated impressive performance across various benchmarks. YOLOv12 introduces architectural enhancements and anchor-free detection strategies that improve accuracy, while YOLO-MS focuses on multi-scale fusion and improved generalization 32 . RT-DETR adopts transformer-based architectures to support end-to-end object detection with high precision 33 . However, many of these approaches significantly increase model size and computational complexity, making them less suitable for rsource-constrained scenarios such as UAV-based foreign object inspection. Our work specifically targets lightweight, real-time detection under embedded deployment constraints, where model interpretability, inference speed, and energy efficiency are prioritized. A comprehensive comparison with these advanced models is reserved for future work, once deployment-ready versions of these detectors become more accessible. Nevertheless, we acknowledge their importance and have updated the related work section to better reflect recent developments. Proposed method In the field of foreign object detection on power transmission lines, main challenges encompass interference from the background environment with foreign object feature recognition. Additionally, there is difficulty in extracting valid feature information from small targets due to noise. Moreover, the diversity of foreign object shapes reduces detection accuracy. To address these issues, the YOLOv8 algorithm is employed as the base model in this study. Through advanced feature extraction and fusion techniques, a balance between detection accuracy and processing speed is effectively achieved. Upon feeding image data into the network, feature extraction is initially performed by the model, followed by feature fusion, ultimately enabling recognition and classification of foreign objects. In response to the special requirements of foreign object detection on power transmission lines, targeted enhancements are made to the YOLOv8 algorithm in this study, aiming to construct a more suitable and efficient detection algorithm for such applications. These improvements enhance both computational efficiency and practical applicability. In the field of foreign object detection on power transmission lines, the main challenges include interference from complex backgrounds, difficulty in extracting features from small or camouflaged objects, and the large shape variability of foreign intrusions. To address these issues, we adopt YOLOv8n as a lightweight base framework due to its balance of speed and accuracy. However, the original YOLOv8n does not fully meet the domain-specific requirements in its default configuration. Therefore, we introduce several targeted modifications to adapt YOLOv8n to this application. First, SCConv is used to improve spatial-channel feature representation for subtle, noise-affected targets. Second, we reconstruct the feature fusion mechanism to enhance semantic flow across scales, optimized for power line imagery. Third, an SE-based detection head is deployed to recalibrate attention on critical regions, mitigating the background interference problem. These modules are jointly optimized under real-time constraints, ensuring compatibility with UAV-based edge computing environments. Unlike simple module stacking, each enhancement is structurally integrated to match the UAV deployment constraints and the high false-positive risk in power grid scenarios. The resulting network, SGSYOLO, achieves robust detection performance while maintaining low inference cost. The overall network structure is shown in Fig. 1 . Fig. 1. Open in a new tab The structure of SESYOLO. Backbone Feature extraction is the process of progressively mining for extracting information from image data layer by layer. Convolutional Neural Networks(CNNs) are widely used in computer vision due to their ability to capture representative features. However, CNNs require substantial computational resources, partly because convolution layers extract redundant features 35 . The standard convolution operation, with its fixed kernels and receptive fields, struggles with the diverse challenges of detecting foreign objects on transmission lines. To address these issues, the proposal involves introducing SCConv before the SPPF layer in YOLOv8, aiming to enhance model accuracy by optimizing the feature extraction process 36 . In our work, SCConv is not directly adopted but integrated into the YOLOv8n backbone prior to the SPPF layer, serving as a lightweight enhancement to spatial-channel feature encoding for small-object detection. Rather than replacing the entire convolutional pipeline, we selectively insert SCConv at a critical bottleneck to balance representational richness and computational efficiency. This tailored use improves the model’s sensitivity to subtle structural features of foreign objects while preserving YOLOv8n’s speed advantages. To improve feature representation while maintaining computational efficiency, we introduce a modified SCConv module into the YOLOv8n backbone before the SPPF layer. SCConv is inspired by architectures that decouple spatial and channel-wise information to reduce feature redundancy. Previous designs have explored spatial/channel attention separately or jointly, but often introduce significant computation overhead. In contrast, we incorporate a lightweight version of this approach suited for real-time applications. SCConv focuses on reducing redundancy in feature maps across spatial and channel dimensions, thereby enhancing the model’s representation learning capacity. Specifically, for the intermediate input features X in the bottleneck residual block, spatial features are first refined into through the SRU operation, followed by refining channel features to Y using the CRU operation. The SCConv module capitalizes on the spatial and channel redundancies among features, allowing seamless integration into any CNN architecture to reduce redundancy between intermediate feature maps and improve the CNN’s feature representation. The structure of this module is illustrated in Fig. 2 . Fig. 2. Open in a new tab The architecture of SCConv integrated with Spatial Reconstruction Unit (SRU) and Channel Reconstruction Unit (CRU). In the SCConv module, all parameters are concentrated in the transformation stage. Consequently, the reduction in theoretical memory usage is analyzed. The parameters of the standard convolution can be calculated as follows: 1 In this context, k represents the kernel size, and and respectively denote the number of input and output feature channels. The parameters of the proposed SCConv module include: 2 In the proposed module, denotes the split ratio, r represents the squeeze ratio, and g is the group size in the Grouped Weighted Convolution operation.The variables and represent the number of input and output feature channels, respectively. Here, a comparison demonstrating the performance of the proposed is presented. In experiments, typical settings include , , , and . Under conditions where , the parameter count can be reduced to one-fifth, while the model achieves better performance than the standard convolution. Introducing before the SPPF layer in YOLOv8 can fully harness its capabilities to optimize feature extraction. The SPPF layer, as a crucial part of the backbone network, heavily relies on the quality of its input features for the overall model performance. By reducing the redundancy in the spatial and channel dimensions of the input features, enhances the quality of features received by the SPPF layer, thereby enhancing the accuracy of the entire model. In our design, we downscale the transformation dimensions and reduce convolutional grouping complexity compared to canonical SCConv structures, ensuring compatibility with YOLOv8n’s minimal latency requirements. The modified SCConv preserves performance while enabling integration into a highly lightweight backbone. Neck Efficient RepGFPN overcomes the limitations of unidirectional information flow by offering a more effective feature fusion strategy 37 . This network leverages the strengths of the Generalized FPN to enhance both accuracy and efficiency, with specific adjustments detailed in Fig. 3 . While inspired by generalized pyramid structures, our implementation redesigns the neck to better suit UAV-based real-time inspection. Specifically, we compress the lateral pathways, modify upsampling ratios, and use fewer fusion operations to reduce redundant propagation. These adjustments enable more efficient cross-scale semantic learning while maintaining low overhead, which is essential for small-object detection in aerial imagery. It utilizes a flexible channel dimension design tailored to the computational complexity and information representation differences among various scale feature maps. This design enables different channel counts per feature scale, improving model accuracy under computational constraints and overcoming performance limitations of fixed channel configurations. Moreover, Efficient RepGFPN optimizes feature interaction mechanisms by eliminating unnecessary upsampling steps present in the Generalized FPN, reducing computational costs and model latency while maintaining high accuracy. In terms of feature fusion strategy, it discards the traditional 3x3 convolution-based method in favor of CSPNet, significantly boosting model precision. Further enhancements are made by integrating re-parameterization techniques with connections from the Effective Layer Aggregation Network 38 , optimizing the CSPNet approach without adding significant computational overhead. Fig. 3. Open in a new tab The structure of Efficient RepGFPN. Overall Efficient RepGFPN achieves an excellent balance between real-time performance and accuracy through its flexible channel design, optimized feature interactions, and efficient fusion modules. This model not only efficiently extracts information across different scales but also significantly enhances target detection precision and efficiency through improved feature fusion and interaction mechanisms. It offers an effective and precise new solution for real-time object detection applications. Head To significantly enhance detection performance without sacrificing accuracy or increasing computational load, a new detection head, SE-Detect, based on the Squeeze-and-Excitation attention mechanism 39 , is introduced in this study. Rather than directly applying a standard SE module, our SE-Detect head is specifically redesigned to balance efficiency and performance. We adopt a simplified excitation structure with fewer fully connected layers and lower channel squeeze ratios, optimized for detecting faint or occluded targets with minimal computational burden. This custom design enables effective enhancement of salient features in cluttered environments while maintaining inference speed. The core concept of SENet is to re-calibrate features by explicitly modeling the dependencies between feature channels, as illustrated in Fig. 4 . Fig. 4. Open in a new tab Squeeze and excitation network structure. By integrating the SE network module into the Bottleneck network, this study enable the network to focus on the texture and color features of small targets by assigning high weights, while applying low weights to ignore non-critical information such as the background.This method continuously optimizes by re-calibrating the model weights and focuses on important target information across multiple paths, dynamically re-calibrating channel features. Loss function The study employs the WIoU v3 mechanism 40 , which utilizes a dynamic non-monotonic mechanism for precise assessment of anchor box quality. This significantly enhances the model’s capability to handle medium-quality anchor boxes and improves both localization accuracy and detection capability. WIoU v3 introduces the outlier parameter , which is designed to assess the quality of anchor boxes. Using a non-linear focusing factor r constructed based on values, WIoU v3 achieves quality weighting on the basis of WIoU v1. When values are low, indicating superior anchor box quality, it assigns a lower r value, thereby assigning a lower weight to high-quality anchor boxes in the loss function. Conversely, higher values suggest poorer anchor box quality, leading to the use of reduced gradient gains to mitigate the adverse gradient effects introduced by low-quality anchor boxes. Employing this refined gradient gain adjustment mechanism, WIoU v3 dynamically balances the weighting of the loss function for anchor boxes of varying quality. This encourages the model to focus more on average quality samples, thus significantly enhancing the model’s overall performance. 3 4 5 where is a function that assigns lower weights to high-quality anchor boxes when values are low, and mitigates the adverse gradient effects from low-quality anchor boxes when values are high. Feature distillation In the realm of model enhancement, Knowledge Distillation (KD) emerges as a highly effective method, as highlighted by Hinton et al. 41 . Challenges associated with applying KD to the YOLO series of models, such as the complexity of hyperparameter tuning, are overcome. This is achieved by employing a feature-based distillation approach for transferring dark knowledge, which mitigates the influence of noisy features. This approach can simultaneously extract recognition and localization information from intermediate feature maps, thereby crucially improving model performance as it captures both recognition and location information from these maps 42 . To swiftly verify and select the most suitable distillation method for SESYOLO, a series of experiments are conducted. It is found that Contrastive Weighted Distillation (CWD) is more appropriate for the model architecture, while the Mimicking method performs poorly compared to Masked Generative Distillation (MGD) due to differences in feature representation, affecting adaptability and generalizability. The proposed distillation strategy consists of two phases: In the first phase, the student model is distilled by the teacher model within a strong mosaic domain for 290 epochs. Faced with a challenging and enhanced data distribution, the student model extracts useful information more effectively under the guidance of the teacher. The second phase involves fine-tuning the student model in a non-mosaic domain for 10 epochs without further distillation, preventing the teacher model from leading the student into an unfamiliar domain, thereby impairing the student’s performance. Although long-term distillation could mitigate this negative effect, it is costly. Thus, SESYOLO adopts a balanced approach aimed at fostering the independence of the student model. Furthermore, in the SESYOLO, the distillation process includes two innovative improvements: the use of an alignment module, consisting of a linear projection layer that adjusts the resolution of student features to match the teacher features. This module aids the student features in better mimicking the teacher features and enhances the distillation effect by reducing the impact of true value differences. Additionally, when calculating the KL loss, the standard deviation of each channel is used as a temperature coefficient, further optimizing the distillation process. With this meticulously designed distillation strategy, SESYOLO achieves significant performance improvements while maintaining cost efficiency. Experiments Dataset The dataset of foreign objects on power transmission lines is collected using high-precision positioning technology and intelligent analysis and diagnostic techniques. It includes four common types of foreign objects: bird nests, kites, balloons, and plastic bags. Additionally, insulators are included as a separate category to enhance the algorithm’s identification capability, totaling 1401 images. Simulated overexposure and underexposure effects are achieved using HSV enhancement. Random-sized kernels are applied to add blur and noise, while images with rain, snow, and fog effects are created by merging images of varying gradient granularity. Specifically, these enhancement techniques are selectively applied to 20% of the images in the original dataset, generating 940 enhanced images, significantly enriching the dataset content. However, to address sample imbalance issues, this study utilizes the AIGC-based application platform for dataset augmentation, adding 486 images. Finally, a dataset consisting of 2817 images of foreign objects is constructed, with 2502 images for training, 157 images for validation, and 158 images for testing. Specifically, 1,401 images are real UAV captures. To increase data diversity and mitigate class imbalance, we apply two types of traditional augmentation: HSV-based visual enhancement (470 images) and scene compositing-based augmentation (470 images), resulting in 940 enhanced images. In addition, we synthesize 476 images using AIGC techniques (text-to-image generation and image inpainting), which are distributed across all five object categories. These synthetic samples are manually filtered to ensure visual realism and contextual accuracy. Only the training and validation sets include augmented or synthetic images. The test set consists of 158 real UAV images, with manual annotation, to ensure unbiased evaluation of detection performance. Although the test set contains only 158 samples, each class is well represented and all images underwent strict annotation and quality filtering. While small in number, the high precision of labeling and balance across classes ensure its robustness for evaluation. In summary, the dataset generation process combines multiple strategies: (1) real-world image acquisition via UAV patrols; (2) classical image augmentation using HSV and geometric transforms; (3) synthetic image generation using AIGC models followed by manual selection. Each strategy is designed to enhance dataset balance and model generalization. In terms of availability, due to data sensitivity and confidentiality agreements with the State Grid Corporation of China, the raw image data and labeled dataset are not publicly available. The dataset was collected under restricted access conditions as part of a national grid safety inspection project. As a result, data sharing is currently prohibited. However, upon request and subject to formal approval, portions of the dataset may be considered for academic research collaboration. Evaluation metric To comprehensively evaluate the detection performance of the proposed improved model, a range of evaluation metrics have been employed. These metrics encompass accuracy, recall rate, [email protected], [email protected], model parameter count, model size, and detection speed. Some of the evaluation metrics utilize the following parameters in their formulas: TP (true positives, predicted as positive samples and are indeed positive), FP (false positives, predicted as positive but are actually negative), and FN (false negatives, predicted as negative but are actually positive). Precision: Accuracy measures the proportion of correctly predicted samples to the total number of samples, indicating the precision of the model. 6 Recall: It also known as true positive rate, denotes the proportion of actual positive samples that are predicted as positive by the model, reflecting the model’s ability to recognize positive samples. 7 mAP: Mean Average Precision is the average of precision values, combining information from both accuracy and recall, used to measure the model’s performance at different recall levels. [email protected] represents the mAP value at an IoU threshold of 0.5, while [email protected] represents the average mAP over IoU thresholds ranging from 0.5 to 0.95 (usually with a step size of 0.05). mAP is a comprehensive metric for evaluating algorithmic detection performance. In this study, it refers to the average detection accuracy of all categories of transmission line anomalies. A higher mAP value indicates better overall performance of the algorithm. 8 FPS: Frames Per Second is used in object detection to measure the speed at which a system processes images, specifically referring to the number of image frames processed per second. In practical applications, the processing time per frame may include image preprocessing time, model inference time, and post-processing time, among others. The calculation formula for FPS is based on the time taken to process each frame. 9 Experimental results Ablation experiment A series of ablation experiments are conducted. These aim to comprehensively evaluate the impact of different components on experimental performance and validate the effectiveness of the improvement strategies adopted in this study. Despite a slight increase in the number of parameters and computational load, the optimized model has achieved a 2.4% improvement in [email protected] over the initial prototype, as detailed in Table 1 . Table 1. Ablation experiments on n-size. Baseline SCConv Efficient RepGFPN SE-Detect WIoU [email protected] [email protected]:.95 Model size Params FLOPs YOLOv8n 84.1 64.1 6.3MB 3.2M 8.1G 85.5 63.3 6.5MB 3.0M 8.2G 85.6 66.9 6.9MB 3.1M 8.4G 85.8 66.8 6.3MB 2.9M 8.1G 83.2 62.3 6.3MB 2.9M 8.1G 85.9 65.0 10.8MB 5.0M 19.3G 84.0 62.3 6.5MB 2.9M 9.3G 83.9 64.3 6.5MB 2.9M 9.3G 86.4 66.3 7.9MB 3.6M 12.0G 86.5 66.2 7.9MB 3.6M 12.0G Open in a new tab From the ablation results, several performance trends can be observed. SCConv, when added individually, leads to a noticeable increase in both precision and recall, indicating its effectiveness in enhancing spatial and channel feature representation, especially for small and irregular foreign objects. Efficient RepGFPN improves recall more significantly, due to its improved multiscale semantic fusion, which is crucial for detecting objects at varying scales in aerial imagery. The SE-Detect head contributes moderately to performance by refining salient channel-wise features with minimal overhead. When combined, the modules produce compounding benefits. The full SESYOLO architecture outperforms all partial variants, suggesting strong complementarity among the modules. These results demonstrate that the improvements are not only effective in isolation but also work synergistically to enhance robustness and generalization. Overall, the ablation study confirms that each component plays a distinct and beneficial role in improving the model’s accuracy and practical deployment value. Distillation experiment The distillation experiment results in Table 2 indicate that a temperature coefficient of 1 yields the best performance, with a 3.1% increase in [email protected] value. Table 2. Effect of distillation on [email protected]. tau Feature loss type Feature loss ratio Before distillation After distillation 10 cwd 1.0 86.5 84.5 5 cwd 1.0 86.5 87.6 2 cwd 1.0 86.5 88.6 1 cwd 1.0 86.5 89.6 Open in a new tab Comparative experiment To assess the performance advantages of the SESYOLO model, a comparative analysis is conducted with other object detection models from the YOLO series. The comparative experimental results in Table 3 demonstrate that the proposed network architecture not only ensures efficient computational speed but also exhibits superior detection performance. Notably, optimal performance is achieved on both the [email protected] and [email protected] metrics, reflecting high accuracy in object detection tasks. Compared to existing YOLO detection models, the architecture presents a significant advantage in precision, markedly enhancing the accuracy of task execution. Additionally, with a recall rate of 81.2%, the missed detection rate is effectively reduced, improving reliability in practical application scenarios. In addition to YOLOv8n, we compared our model with YOLOv5-s and YOLOv8-s to ensure fairness in terms of parameter count and model size. All FPS tests were conducted using single-image inference (batch size = 1) on an RTX 4060 GPU using PyTorch 1.12 and CUDA 11.6. Table 3. Comparative experiment. Model [email protected] [email protected]:.95 Precision Recall Params FLOPs FPS (RTX 4060) YOLOv5-n 79.4 54.1 87 69.2 1.9M 4.5G 169 YOLOv5-s 83.3 56.1 82.9 81.4 6.6M 15.8G 136 YOLOv5-m 84.4 59.7 93.9 79.6 19.9M 47.9G 72.9 YOLOv6-n 71.1 45.9 74.4 67.2 4.3M 11.1G 102 YOLOv6-s 72.7 47.8 80.9 62.9 17.2M 44.2G 108 YOLOv6-m 73.9 52.0 76.7 68.0 34.3M 82.2G 47 YOLOv7 77.2 53 79.2 73.6 34.8M 103.2G 2 YOLOv8-n 80.6 54.4 82.1 72.1 3.2M 8.1G 176 YOLOv8-s 81.7 55.7 81.9 72.0 10.6M 28.4G 112 YOLOv8-m 81.8 57.2 85.5 76.7 24.6M 78.7G 58 YOLOv9 84.0 60.6 81.4 83.5 48.3M 236.7G 30 OURS 89.6 67.7 83.4 81.2 3.6M 12.0G 142 Open in a new tab A detailed comparison of the performance between the YOLOv8n model and the proposed method is conducted, as shown in Table 4 . The proposed method, compared to the YOLOv8n model, achieves an improvement of 3 to 5 percentage points in mAP across various detection categories. Particularly in the task of detecting foreign objects on power transmission lines, precision is increased to 93.9% for the bird’s nest category, significantly improving detection efficiency. This enhancement provides a robust technical foundation for deployment in real-world application scenarios. Table 4. Comparison of mAP for different categories. Category YOLOv8n Ours [email protected] [email protected]:.95 [email protected] [email protected]:.95 Nest 89.6 52.0 93.9 57.1 Balloon 78.2 56.7 91.6 63.2 Kite 77.9 60.1 81.1 65.3 Plastic bag 84.3 78.0 87.5 77.0 Insulator 90.5 73.7 93.8 75.8 Open in a new tab Heatmap analysis GradCAM is an effective visualization tool used to generate heatmaps 43 . Through the reverse propagation mechanism of GradCAM, the model’s output class confidence can be converted into gradient values, visualizing the gradient intensity of feature maps in the heatmap. Deeper red shadows indicate areas the model focuses on more, while deeper blue shadows indicate areas of lower attention. As depicted in Fig. 5 , the experimental outcomes distinctly reveal the disparities in feature attention among various object detection models. Observations from the results indicate that the YOLOv8n model demonstrates a deficiency in allocating attention to diminutive objects and exhibits a relatively low sensitivity to objects situated at a distance. In stark contrast, the model introduced in this study has shown exceptional performance in mitigating background noise, with a pronounced concentration of attention on the central regions of objects. This refined attention allocation mechanism significantly enhances the accuracy of the model in bounding box prediction tasks, thereby substantially improving the overall detection performance. Furthermore, this improvement not only bolsters the model’s robustness but also provides a robust technical foundation for performance optimization in practical applications. Fig. 5. Open in a new tab ( a ) Original images; ( b ) heat maps of YOLOv8n; ( c ) heat maps of our model. Conclusions To enhance the accuracy of foreign object detection in power line inspection images, the SESYOLO model is developed in this study, based on the YOLOv8n framework. The SCConv module is integrated within the backbone network, effectively merging spatial and channel information, thereby significantly boosting feature extraction capabilities. Additionally, the model is redesigned to achieve multi-scale feature fusion, effectively reducing computational costs and model latency. In the detection layer, a detection head based on the SE attention mechanism is employed to optimize the integration of shallow and deep feature maps. Concurrently, the WIoU v3 mechanism is incorporated into the bounding box regression loss function to balance gradient gains between high-quality and low-quality samples, substantially enhancing the model’s localization accuracy and generalization ability. To further improve model performance, a knowledge distillation loss function is designed and knowledge distillation algorithms are applied. Experimental results demonstrate the exceptional performance of the SESYOLO model. Notably, the model requires only 7.9 MB of memory and achieves an [email protected] of 89.6%, while maintaining a detection speed of 142 frames per second, fulfilling the requirements for real-time detection of foreign objects on power lines. Acknowledgements The authors would like to thank the financial support from the Fundamental Research Funds for the Central University under Grant 2022QNYL32 and 2024ZLQN40, and in part by the Beijing Science and Technology Planning Project under Grant Z231100001723002. (Corresponding authors: Xiao Liang. Co-first authors: Pingting Duan, Xuran Zhang.) Author contributions P.D. conceptualized and designed the overall study, executed experiments, analyzed results, and supervised the project. X.Z. contributed substantially to the algorithm design phase, including proposing lightweight structural strategies and optimizing feature integration for small-object detection. X.Z. also performed critical ablation experiments and drafted key sections of the manuscript, particularly the Proposed Method and Experiments. X.L. provided methodological guidance, ensured experimental rigor, and improved the clarity and structure of the manuscript during revision. Y.C. was responsible for dataset collection and preliminary data analysis. H.W. assisted in language editing and manuscript formatting. Data availability The raw and processed data required to reproduce these results are available by contacting the author—Xiao Liang([email protected]). Code availability The custom code developed for this study was used to generate the results reported in the manuscript. The complete code has been provided to the journal for editorial and peer review purposes. Due to confidentiality and intellectual property restrictions, the source code is not publicly available. Researchers interested in accessing the code may contact the corresponding author directly. Declarations Competing interests The authors declare no competing interests. Footnotes Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. P. Duan: These authors contributed equally to this work. References 1. Rong, S., He, L., Du, L., Li, Z. & Yu, S. Intelligent detection of vegetation encroachment of power lines with advanced stereovision. IEEE Trans. Power Deliv. 36 , 3477–3485 (2020). [ Google Scholar ] 2. Wang, L., He, Y. & Li, L. A single-terminal fault location method for hvdc transmission lines based on a hybrid deep network. Electronics 10 , 255 (2021). [ Google Scholar ] 3. Butt, O. M., Zulqarnain, M. & Butt, T. M. Recent advancement in smart grid technology: Future prospects in the electrical power network. Ain Shams Eng. J. 12 , 687–695 (2021). [ Google Scholar ] 4. Jing, Z. et al. Reliability analysis of distribution network operation based on short-term future big data technology. In Journal of Physics: Conference Series . 1584, 012027 (IOP Publishing, 2020). 5. Wen, X. et al. High-risk region of bird streamer flashover in 110 kv composite insulators and design for bird-preventing shield. Int. J. Electr. Power Energy Syst. 131 , 107010 (2021). [ Google Scholar ] 6. Wu, M. et al. Improved yolox foreign object detection algorithm for transmission lines. Wirel. Commun. Mobile Comput. 2022 , 5835693 (2022). [ Google Scholar ] 7. Koshelev, V. I. & Kozlov, D. Wire recognition in image within aerial inspection application. In 2015 4th Mediterranean Conference on Embedded Computing (MECO) , 159–162 (IEEE, 2015). 8. Zou, K., Jiang, Z. & Zhang, Q. Research progresses and trends of power line extraction based on machine learning. In 2021 2nd International Symposium on Computer Engineering and Intelligent Communications (ISCEIC) , 211–215 (IEEE, 2021). 9. Wu, Y. et al. A novel line position recognition method in transmission line patrolling with uav using machine learning algorithms. In 2018 IEEE International Symposium on Electromagnetic Compatibility and 2018 IEEE Asia-Pacific Symposium on Electromagnetic Compatibility (EMC/APEMC) , 491–495 (IEEE, 2018). 10. Qi, S. Cleaning system based on autonomous patrol of uav and intelligent detection of foreign matters. In 2020 7th International Forum on Electrical Engineering and Automation (IFEEA) , 675–678 (IEEE, 2020). 11. Shen, J. et al. An algorithm based on lightweight semantic features for ancient mural element object detection. npj Herit. Sci. 13 , 70. 10.1038/s40494-025-01565-6 (2025). [ Google Scholar ] 12. Shen, J., Liu, N., Sun, H., Li, D. & Zhang, Y. An instrument indication acquisition algorithm based on lightweight deep convolutional neural network and hybrid attention fine-grained features. IEEE Trans. Instrum. Meas. 73 , 1–16. 10.1109/TIM.2023.3346488 (2024). [ Google Scholar ] 13. Shen, J. et al. Finger vein recognition algorithm based on lightweight deep convolutional neural network. IEEE Trans. Instrum. Meas. 71 , 1–13. 10.1109/TIM.2021.3132332 (2022). [ Google Scholar ] 14. Chen, C., Yang, B., Song, S., Peng, X. & Huang, R. Automatic clearance anomaly detection for transmission line corridors utilizing uav-borne lidar data. Remote Sens. 10 , 613 (2018). [ Google Scholar ] 15. Cheng, L. & Wu, G. Obstacles detection and depth estimation from monocular vision for inspection robot of high voltage transmission line. Cluster Comput. 22 , 2611–2627 (2019). [ Google Scholar ] 16. Jiao, S. & Wang, H. The research of transmission line foreign body detection based on motion compensation. In 2016 First International Conference on Multimedia and Image Processing (ICMIP) , 10–14, 10.1109/ICMIP.2016.14 (2016). 17. Wu, S., Kan, M., He, Z., Shan, S. & Chen, X. Funnel-structured cascade for multi-view face detection with alignment-awareness. Neurocomputing 221 , 138–145. 10.1016/j.neucom.2016.09.072 (2017). [ Google Scholar ] 18. Mahdi Elsiddig Haroun, F., Mohamed Deros, S. N., Bin Baharuddin, M. Z. & Md Din, N. Detection of vegetation encroachment in power transmission line corridor from satellite imagery using support vector machine: A features analysis approach. Energies . 14 (2021). 19. Ye, X., Wang, D., Zhang, D. & Hu, X. Transmission line obstacle detection based on structural constraint and feature fusion. Symmetry 10.3390/sym12030452 (2020). [ Google Scholar ] 20. Liang, H., Zuo, C. & Wei, W. Detection and evaluation method of transmission line defects based on deep learning. IEEE Access 8 , 38448–38458. 10.1109/ACCESS.2020.2974798 (2020). [ Google Scholar ] 21. Li, H. et al. An improved yolov3 for foreign objects detection of transmission lines. IEEE Access 10 , 45620–45628. 10.1109/ACCESS.2022.3170696 (2022). [ Google Scholar ] 22. Song, Y. et al. Intrusion detection of foreign objects in high-voltage lines based on yolov4. In 2021 6th International Conference on Intelligent Computing and Signal Processing (ICSP) , 1295–1300, 10.1109/ICSP51882.2021.9408753. (2021). 23. Hui, Z. et al. Intelligent bird’s nest hazard detection of transmission line based on retinanet model. In Journal of Physics: Conference Series , vol. 2005, 012235 (IOP Publishing, 2021). 24. Huang, Y. et al. Real-time detection method for transmission line faults applying edge computing and improved yolov5s algorithm. Elect. Power Constr. 44 , 91–99 (2023). [ Google Scholar ] 25. Li, H., Dong, Y., Liu, Y. & Ai, J. Design and implementation of uavs for bird’s nest inspection on transmission lines based on deep learning. Drones 6 , 252 (2022). [ Google Scholar ] 26. Liu, B., Huang, J., Lin, S., Yang, Y. & Qi, Y. Improved yolox-s abnormal condition detection for power transmission line corridors. In 2021 IEEE 3rd International Conference on Power Data Science (ICPDS) , 13–16 (IEEE, 2021). 27. Yu, C. et al. Foreign objects identification of transmission line based on improved yolov7. IEEE Access (2023). 28. Zhao, Z. et al. Applications of unsupervised deep transfer learning to intelligent fault diagnosis: A survey and comparative study. IEEE Trans. Instrum. Meas. 70 , 1–28 (2021).33776080 [ Google Scholar ] 29. Su, H., Xiang, L., Hu, A., Xu, Y. & Yang, X. A novel method based on meta-learning for bearing fault diagnosis with small sample learning under different working conditions. Mech. Syst. Signal Process. 169 , 108765 (2022). [ Google Scholar ] 30. Sudharsan, R. & Ganesh, E. A swish rnn based customer churn prediction for the telecom industry with a novel feature selection strategy. Connect. Sci. 34 , 1855–1876 (2022). [ Google Scholar ] 31. Chen, Z., Yang, J., Feng, Z. & Zhu, H. Railfod23: A dataset for foreign object detection on railroad transmission lines. Sci. Data 11 , 72 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 32. Tian, Y., Ye, Q. & Doermann, D. Yolov12: Attention-centric real-time object detectors. Preprint at arXiv:2502.12524 (2025). 33. Kong, Y., Shang, X. & Jia, S. Drone-detr: Efficient small object detection for remote sensing image using enhanced rt-detr model. Sensors 24 , 5496 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 34. Miao, D., Wang, Y., Yang, L. & Wei, S. Foreign object detection method of conveyor belt based on improved nanodet. IEEE Access 11 , 23046–23052 (2023). [ Google Scholar ] 35. Russakovsky, O. et al. Imagenet large scale visual recognition challenge. Int. J. Comput. Vis. 115 , 211–252 (2015). [ Google Scholar ] 36. Li, J., Wen, Y. & He, L. Scconv: spatial and channel reconstruction convolution for feature redundancy. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 6153–6162 (2023). 37. Xu, X. et al. Damo-yolo: A report on real-time object detection design. Preprint at arXiv:2211.15444 (2022). 38. Wang, C.-Y., Bochkovskiy, A. & Liao, H.-Y. M. Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 7464–7475 (2023). 39. Hu, J., Shen, L. & Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition , 7132–7141 (2018). 40. Tong, Z., Chen, Y., Xu, Z. & Yu, R. Wise-iou: bounding box regression loss with dynamic focusing mechanism. Preprint at arXiv:2301.10051 (2023). 41. Hinton, G., Vinyals, O. & Dean, J. Distilling the knowledge in a neural network. Preprint at arXiv:1503.02531 (2015). 42. Huang, T. et al. Masked distillation with receptive tokens. Preprint at arXiv:2205.14589 (2022). 43. Selvaraju, R. R. et al. Grad-cam: Visual explanations from deep networks via gradient-based localization. In 2017 IEEE International Conference on Computer Vision (ICCV) , 618–626, 10.1109/ICCV.2017.74 (2017). Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Data Availability Statement The raw and processed data required to reproduce these results are available by contacting the author—Xiao Liang([email protected]). The custom code developed for this study was used to generate the results reported in the manuscript. The complete code has been provided to the journal for editorial and peer review purposes. Due to confidentiality and intellectual property restrictions, the source code is not publicly available. Researchers interested in accessing the code may contact the corresponding author directly. Articles from Scientific Reports are provided here courtesy of Nature Publishing Group ACTIONS View on publisher site PDF (2.1 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 1015 · SHA-256 7bce5b38d47922d7
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.