IEEE Internet Computing
From Pixels to Semantics: Edge AI for UAV-Based Critical Infrastructure Inspection Reza Farahani *, Distributed Systems Group, TU Wien, Austria Naser Hossein Motlagh, University of Helsinki, Finland Zoha Azimi, University of Klagenfurt, Austria Christian Timmerer, University of Klagenfurt, Austria
arXiv:2609.18448v1 [cs.DC] 16 Sep 2026
Lorenzo Carnevale, University of Messina, Italy Sasu Tarkoma, University of Helsinki, Finland Schahram Dustdar, Distributed Systems Group, TU Wien, Austria and ICREA, Barcelona, Spain ∗ Corresponding author: [email protected]
Abstract—Critical infrastructure assets such as bridges, tunnels, dams, and power line networks require timely and scalable inspection. While conventional manual inspection remains costly and hazardous, unmanned aerial vehicle (UAV)-based inspection has emerged as an efficient alternative for monitoring difficult-to-access structures. Existing UAV inspection pipelines have evolved from cloud-centric offline processing toward edge-based perception using lightweight object detectors such as YOLO for realtime defect localization. This article explores the transition toward fully edge-native semantic inspection powered by lightweight vision language models (VLMs), where UAVs move beyond object detection toward contextual structural understanding. It categorizes existing UAV inspection architectures, identifies their key system challenges and architectural requirements, and experimentally assesses the feasibility of semantic edge intelligence on NVIDIA Jetson UAV-class hardware using the COCO-Bridge dataset. The evaluation integrates a fine-tuned YOLO-26M for object localization and a lightweight SmolVLM-256 for semantic reasoning. Finally, it outlines future directions toward agentic, autonomous, trustworthy, and collaborative semantic UAV inspection across the edge-cloud continuum.
Introduction The aging and growing scale of critical infrastructure systems demand inspection solutions that are faster, safer, and more scalable than conventional manual procedures. Recent advances in unmanned aerial vehicles (UAVs), commonly known as drones, high-resolution sensing, and artificial intelligence (AI) are rapidly transforming infrastructure monitoring by enabling automated visual inspection across largescale and difficult-to-access environments. As shown in Fig. 1, UAV-based inspection is increasingly applied
XXXX-XXX © 2026 IEEE Digital Object Identifier 10.1109/XXX.0000.0000000 September
Published by the IEEE Computer Society
across diverse infrastructure domains, e.g., facade and tower assessment, road surface monitoring, tunnel inspection, and power-line analysis. These scenarios involve heterogeneous structural anomalies, such as cracks, corrosion, seepage, displacement, rutting, and component degradation that require accurate visual assessment under dynamic operational conditions. The combination of autonomous aerial mobility and modern imaging platforms enables UAVs to capture detailed visual information with significantly higher flexibility and coverage than traditional inspection approaches. However, continuous UAV missions also generate massive streams of high-resolution visual data, while only a small fraction of captured frames typically contains inspection-relevant information. EffiIEEE Internet Computing
1
IEEE Internet Computing
Building and Tower Inspection
Road Surface Monitoring
(1) Facade Crack, (2) Degradation, (3) Breakage
(1) Pothole, (2) Rutting, (3) Crack
Tunnel Inspection (1) Seepage, (2) Deformation, (3) Displacement
Power Line Inspection (1) Degradation, (2) Missing Component, (3) Corrosion
FIGURE 1: Representative UAV-based infrastructure inspection scenarios and related structural anomalies. ciently filtering redundant data and extracting and interpreting relevant visual evidence, therefore, remain key challenges for scalable, real-time inspection pipelines. Rather than serving solely as airborne sensing platforms, recent UAVs incorporate onboard perception and AI inference capabilities to process inspection data during flight. This shift toward onboard intelligence is changing how inspection data are processed, transmitted, and interpreted across edge-cloud environments. This article analyzes the evolution of UAV-based infrastructure inspection from cloud-centric visual analysis to onboard edge perception and fully edge-native semantic inspection. Beyond existing paradigms, we investigate the feasibility of deploying lightweight semantic edge intelligence directly on UAV-class hardware through experiments on an NVIDIA Jetson platform using a real-world bridge inspection dataset.” Specifically, we fine-tune and evaluate a lightweight YOLO detector for structural component localization together with a compact vision language model (VLM) for contextual defect interpretation and semantic inspection reporting. Finally, we identify key architectural requirements, operational trade-offs, and outline future directions for intelligent UAV inspection systems.
From Cloud-Centric Inspection to Semantic Edge Intelligence Fig. 2 categorizes UAV-based infrastructure inspection pipelines according to the placement and capabilities of AI inference within the inspection workflow. These paradigms differ not only in where inference is performed, but also in the level of semantic information extracted, ranging from raw visual data acquisition to semantically enriched inspection findings. 2
Cloud-centric visual inspection, shown in Fig. 2(a), represents the earliest UAV inspection paradigm with no on-board AI capabilities. The UAV primarily serves as an airborne sensing platform, capturing high-resolution imagery along with telemetry and positioning data (e.g., GPS/IMU). The complete visual stream is transmitted to cloud or ground infrastructure, where offline processing stages such as image enhancement, segmentation, feature extraction, and defect detection are executed. The resulting outputs typically include detected anomalies, bounding boxes, class labels, and inspection reports. Following this approach, Li et al. [1] proposed a UAVassisted bridge inspection framework using Faster Region-based Convolutional Neural Network (R-CNN) for crack detection, where all captured imagery is processed offline after the inspection mission. Gwon et al. [2] addressed image quality assessment for UAV-based bridge inspection using a CNN-based classifier that filters blurred and degraded images before structural analysis. State-of-the-art limitations: While they benefit from powerful remote computation and largescale model execution, they introduce communication overhead, delayed feedback, and dependency on reliable connectivity due to the continuous transmission of high-resolution visual streams. Onboard edge perception, illustrated in Fig. 2(b), shifts defect detection directly onto UAV-side edge hardware, reducing bandwidth consumption and improving inspection responsiveness. Lightweight deep learning models such as YOLO enable real-time infer-
From Pixels to Semantics: Edge AI for UAV-Based Critical Infrastructure Inspection
September 2026
IEEE Internet Computing
ence during flight, allowing the UAV to perform defect localization and region proposal generation directly on captured frames. Instead of transmitting complete visual streams, it typically extracts and transmits only relevant regions of interest (ROIs), along with metadata, for cloud-side storage and management. This transition shifts the role of UAVs from passive imaging devices to active edge-perception platforms. Following this approach, Ji et al. [3] proposed a YOLOv8-based framework for steel bridge weld defect detection, optimized using lightweight feature extraction and attention mechanisms. Ruggieri et al. [4] enhanced YOLO11 with attention modules for reinforced concrete bridge inspection, improving detection precision under low-resolution conditions while reducing inference latency. Li et al. [5] introduced an optimized real-time detection transformer (RT-DETR) architecture for railway obstacle intrusion detection with reduced parameter count and faster inference, while Ding et al. [6] improved RT-DETR for power line inspection by enhancing small-object detection and removing conventional post-processing stages. Rong et al. [7] employed PL-YOLOv8 for power line inspection, where only vegetation encroachment events exceeding predefined thresholds are transmitted with GPS metadata. State-of-the-art limitations: Compared to cloudcentric pipelines, onboard edge perception reduces unnecessary data transfer and enables faster inspection feedback. However, such approaches remain largely detection-centric, with outputs limited to localized anomalies and predefined defect categories. Semantic edge intelligence with VLMs, depicted in Fig. 2(c), integrates lightweight VLMs directly into the UAV edge pipeline to move beyond object-level detection toward contextual understanding of structural conditions. Through multimodal reasoning during flight, such systems enable anomaly interpretation, naturallanguage description, and visual question answering. Unlike conventional detection pipelines that output only bounding boxes and confidence scores, VLM-based inspection produces semantically enriched findings that combine localized anomalies with contextual interpretation, enabling actionable infrastructure defect reports. Existing research already demonstrates the potential of semantic AI for UAV systems. Li et al. [8] employed a VLM-driven framework for power line inspection, where semantic scene understanding and navigation reasoning are performed through VLM and large language model (LLM) integration. Zhang et al. [9] combined VLM reasoning with segmentation models September 2026
for semantic road crack assessment and automated report generation, while Chen et al. [10] proposed a Contrastive Language-Image Pre-training (CLIP)based framework for semantic bridge inspection guidance through human-UAV interaction. State-of-the-art limitations: These approaches remain largely cloud-centric, while the lightweight and resource-aware deployment of semantic VLM reasoning directly on UAV-side edge hardware remains largely unexplored. This transition introduces new architectural challenges and system requirements, discussed in the next section.
System Challenges for Semantic UAV Infrastructure Inspection Enabling fully edge-based VLM-driven UAV inspection introduces fundamentally different requirements and challenges compared to cloud-centric or detection-only pipelines. Unlike conventional object detectors that primarily localize predefined defects, semantic edge intelligence requires multimodal reasoning, contextual interpretation, and natural-language generation directly on resource-constrained UAVs. This transition introduces new challenges, summarized in Table 1. Edge Hardware and Sensing. Semantic UAV inspection requires integrating high-resolution imaging sensors, embedded accelerators, and energy-efficient communication modules that support real-time multimodal inference during flight. Since structural anomalies such as cracks, corrosion, and spalling often occupy only a small fraction of captured frames, stable high-resolution sensing and precise image acquisition become critical for reliable semantic interpretation. At the same time, UAVs remain constrained by limited battery capacity, onboard memory, thermal envelopes, and payload budgets, making lightweight and resource-aware AI deployment essential. Low-latency Multimodal Reasoning Unlike offline cloud processing, semantic edge inspection requires tightly coupled perception and reasoning pipelines operating directly during flight. This introduces strict latency, memory, and energy constraints for executing VLM inference on UAV-side hardware. The system must continuously balance local processing, selective offloading, and communication overhead according to runtime conditions such as battery level, wireless link quality, and computational load. Consequently, model compression, quantization, pruning, and adaptive inference strategies become central requirements for practical deployment.
From Pixels to Semantics: Edge AI for UAV-Based Critical Infrastructure Inspection
3
IEEE Internet Computing
(a) Cloud-centric inspection.
(b) Onboard edge-based inspection.
(c) Edge-native semantic inspection.
FIGURE 2: Comparison of UAV infrastructure inspection paradigms according to AI placement and inspection understanding, progressing from cloud-centric visual analysis to edge-native semantic intelligence. TABLE 1: Key requirements, challenges, and optimization techniques for semantic UAV infrastructure inspection. Requirement
Challenge
Optimization Techniques
Edge sensing and hardware
High-resolution semantic inspection under limited battery, memory, thermal, and payload constraints.
Lightweight VLMs, TensorRT, quantization, pruning, embedded accelerators.
Low-latency reasoning
multimodal
Low-latency multimodal reasoning during flight-time operation.
Adaptive inference, selective offloading, PEFT/LoRA, dynamic model scaling.
Hierarchical semantic processing
Continuous VLM inference over all frames is computationally inefficient.
YOLO/RT-DETR cascades, ROI filtering, key-frame triggering.
Mission-aware operation
Inference delay increases hover time, propulsion energy, and mission duration.
Joint flight-inference scheduling, adaptive frame sampling, energy-aware planning.
Multimodal training data
Limited infrastructure image-text datasets and semantic annotations.
Synthetic captioning, weak supervision, instruction tuning.
Cross-structure generalization
Defect appearance varies across materials, geometry, and environments.
Domain adaptation, continual learning, and structure-aware fine-tuning.
Trustworthy autonomy
Incorrect semantic decisions may affect mission planning.
Confidence estimation, XAI, RAG grounding, and human verification.
Hierarchical Semantic Processing. Continuous semantic reasoning over every captured frame is computationally inefficient and often operationally unnecessary. Practical UAV inspection pipelines thus require hierarchical processing architectures, in which lightweight detectors such as YOLO or real-time detection transformer (RT-DETR) [5] first identify candidate regions or key frames, triggering semantic VLM reasoning only when a detailed contextual assessment is required. Such hierarchical pipelines are essential for balancing semantic richness, responsiveness, and onboard energy consumption. Mission-aware Scheduling and Flight Constraints. Unlike static edge systems, UAV inspection pipelines operate under continuous mobility and strict flight-time constraints. Inspection quality is thus 4
tightly coupled with flight speed, hover duration, camera viewpoint, and inference latency. Continuous VLM inference during flight may introduce processing backlogs, increased hovering time, and additional propulsion energy consumption, directly reducing mission endurance. Consequently, semantic UAV inspection requires mission-aware scheduling strategies that jointly coordinate frame acquisition, inference triggering, and flight control under latency-energy constraints. Multimodal Training Data. Training semantic inspection models requires aligned multimodal datasets that combine infrastructure imagery with descriptive textual annotations. Unlike conventional detection datasets, semantic inspection demands contextual descriptions of structural conditions, defect severity, material degradation, and environmental context.
From Pixels to Semantics: Edge AI for UAV-Based Critical Infrastructure Inspection
September 2026
IEEE Internet Computing
However, publicly available multimodal infrastructure datasets remain extremely limited, particularly for UAV inspection scenarios. Generating high-quality semantic annotations is additionally costly and domain-specific, making data scarcity a major bottleneck for robust VLM-driven inspection systems. Generalization across Infrastructure Types. Infrastructure assets exhibit substantial variability in geometry, materials, degradation patterns, and environmental conditions. Defects such as cracks, corrosion, seepage, and displacement often appear differently across concrete, steel, and composite structures. Consequently, semantic inspection models must generalize across diverse infrastructure domains while maintaining reliable contextual reasoning. This requires structure-aware adaptation, domain-specific fine-tuning, and robust multimodal reasoning capabilities beyond fixed-category defect detection. Trustworthiness and Explainability. Infrastructure inspection often supports safety-critical decisionmaking, requiring that semantic outputs be reliable, interpretable, and verifiable. Unlike conventional object detectors, VLMs may generate hallucinated or inconsistent descriptions that are difficult to validate automatically. Ensuring trustworthy reasoning, explainable semantic outputs, confidence-aware reporting, and human-verifiable inspection findings is essential for practical deployment in UAV inspection systems.
Feasibility Study: Edge-Native Semantic UAV Bridge Inspection Evaluation Setup. To assess the feasibility of fully edge-based semantic inspection, we used an NVIDIA Jetson Orin representative of UAV-class edge hardware, equipped with an 8-core CPU, 16 GB RAM, 1024 CUDA cores, and 32 Tensor cores. We performed offline model fine-tuning on a server with a 128core Intel Xeon Gold CPU and two NVIDIA Quadro GV100 GPUs. We employed the COCO-Bridge dataset [11], containing 774 bridge inspection images and more than 2500 annotated structural components, including bearings, gusset plate connections, cover plate terminations, and out-of-plane stiffeners. For defect localization, we fine-tuned YOLO26M [12] using the Ultralytics framework with partial backbone freezing for 50 epochs at 640 × 640 resolution using AdamW (5 × 10−4 ) and a batch size of 16. For semantic inspection, we adopted SmolVLM-256 [13], a lightweight 256M-parameter VLM optimized using PEFT [14] and LoRA [15] applied to both the vision encoder and text decoder. We conducted fine-tuning via Hugging Face TRL (SFTTrainer) with fp16 precision September 2026
(r = 16, α = 32, dropout 0.15, learning rate 2 × 10−4 , effective batch size 8). Since publicly available multimodal bridge inspection datasets remain limited, semantic captions describing structural surface conditions were automatically generated using GPT-4o. Component Detection vs. Semantic Inspection. As illustrated in Table 2, YOLO26M accurately localizes bridge components such as gusset plate connections, bearings, and stiffeners using bounding-box detection, but provides limited information about structural condition. Conversely, the fine-tuned SmolVLM-256 generates detailed inspection-oriented descriptions directly from captured imagery, identifying oxidation, staining, flaking, discoloration, and localized surface degradation patterns. Compared to the generic outputs produced by the pre-trained model, domain-specific fine-tuning substantially improves semantic inspection quality and contextual structural interpretation. Edge Resource and Energy Analysis. To quantify operational overhead on UAV-class hardware, we monitored resource utilization using tegrastats [16] with 1 ms sampling intervals, recording CPU/GPU utilization, memory usage, inference latency, and total system power consumption. Computational energy was estimated over the inference interval and contextualized using the DJI M210 V2 UAV battery profile (349.2 Wh, 24-minute endurance). As shown in Fig. 3(b), YOLO26M completes inference in approximately 0.14 s, while SmolVLM-256 requires nearly 2.8 s. Although YOLO26M exhibits slightly higher instantaneous GPU utilization and peak power draw (Fig. 3(a)), its short execution window results in significantly lower cumulative energy consumption. In contrast, SmolVLM-256 sustains elevated GPU and power utilization over a longer interval, primarily affecting inspection responsiveness while introducing additional computational energy overhead (Figs. 3(c)-3(d)). To evaluate cumulative operational impact, we further modeled inspection over a 10 s, 30 fps video sequence (300 frames). YOLO26M processes the sequence in approximately 42 s, consuming only 0.03 % of the UAV energy budget. In contrast, continuous SmolVLM-256 inference requires nearly 14 min while consuming approximately 0.6 % of the battery capacity. Although computational energy remains modest relative to the UAV battery budget, inference latency is the dominant bottleneck, making continuous frameby-frame VLM reasoning impractical. These results demonstrate that lightweight VLM-based semantic inspection is feasible on modern UAV hardware, but efficient deployment requires hierarchical pipelines in which lightweight detectors selectively trigger semantic reasoning only for candidate regions or key frames.
From Pixels to Semantics: Edge AI for UAV-Based Critical Infrastructure Inspection
5
IEEE Internet Computing
Bearing
(a)
(b)
(c)
Pre-trained
Bearing
The surface of the iron is dirty and rusty.
The rusty appearance of the components is consistent with a rusty surface.
The concrete is gray and the metal is black.
The components show signs of rust, as evidenced by the orange and brown discoloration and rough texture. There are also areas where the gray metal appears to be painted.
The components’ surfaces show significant staining and flaking, particularly at the edges and near the connections. The edges have orange and brown discoloration, indicating oxidation.
The surfaces of the concrete and metal components appear to have a consistent texture, with no visible signs of discoloration or flaking. The concrete has a smooth finish with no visible cracks or pitting.
Toward Autonomous Semantic UAV Inspection Systems The results presented in this article indicate that lightweight VLMs can enable semantic infrastructure inspection directly on UAV-class edge hardware. However, the progression from component-level detection toward contextual structural understanding also introduces a broader transition in how future UAV inspection systems will operate across the edge-cloud continuum. Beyond isolated defect localization, nextgeneration inspection pipelines are expected to increasingly incorporate adaptive reasoning, collaborative intelligence, multimodal sensing, and contextaware understanding of infrastructure during flight. Agentic semantic inspection, where lightweight VLMs are integrated with planning and decisionmaking modules to enable closed-loop inspection workflows. Instead of following static trajectories and continuously transmitting visual streams, future UAVs may dynamically adapt viewpoint selection, hover duration, sampling frequency, and inspection priorities based on semantic uncertainty and detected structural conditions. Such pipelines can reduce unnecessary sensing and communication overhead, improve inspection efficiency, and reduce energy consumption. Semantic digital twins, where UAVs continuously update semantically enriched digital representations of bridges, tunnels, and power line assets using contextual structural observations collected during flight. 6
Out of Plane Stiffener
Out of Plane Stiffener
Gusset Plate Connection
Fine-tuned
Bridge Image Dataset
TABLE 2: Comparison of semantic inspections by pre-trained and fine-tuned domain-adapted SmolVLM-256.
Beyond isolated defect detection, such systems can support longitudinal degradation tracking, predictive maintenance, and condition-aware infrastructure management across edge-cloud environments. Collaborative edge intelligence, where multiple UAVs jointly perform hierarchical sensing, cooperative perception, and distributed semantic reasoning across the computing continuum. Lightweight distilled models may execute locally on UAVs, while larger foundation models are selectively invoked at nearby edge servers or cloud infrastructure. Such architectures can enable scalable, large-area inspection while preserving realtime responsiveness and communication efficiency. Multimodal semantic inspection, where visual reasoning is combined with complementary modalities such as thermal imaging, LiDAR, vibration sensing, acoustic analysis, and wireless sensing. Integrating sensing with semantic reasoning can improve defect interpretation, environmental awareness, and robustness under challenging operational conditions. Trustworthy and adaptive reasoning, where future semantic inspection systems incorporate uncertainty estimation, retrieval-grounded reasoning, explainable outputs, and human-in-the-loop verification to improve reliability in safety-critical inspection tasks. In parallel, continual learning, synthetic semantic data generation, and adaptive multimodal fine-tuning will become increasingly important for supporting diverse infrastructure types and evolving degradation patterns.
From Pixels to Semantics: Edge AI for UAV-Based Critical Infrastructure Inspection
September 2026
IEEE Internet Computing
RAM (%)
Inference time (s)
80 60 40 20 0 YOLO26m
SmolVLM-256
(a) CPU, GPU, and RAM utilization.
SmolVLM-256
10
2.5
8
2
6
1.5 1
4
0.5
2 0
0 YOLO26m
SmolVLM-256
(b) Inference latency and power consumption.
SmolVLM-256
YOLO26m
100
YOLO26m
11
80
Power (W)
GPU Utilization (%)
Power (W)
Power (W)
GPU (%)
Inference Time (s)
Resource Usage (%)
CPU (%)
60 40 20
10 9 8 7 6
0 -2 -1 0 1 2 3 4 Time Relative To Inference Start (s)
(c) Temporal GPU utilization during inference.
-2 -1 0 1 2 3 4 Time Relative To Inference Start (s) (d) Temporal power consumption during inference.
FIGURE 3: Resource utilization and inference overhead of YOLO26M and SmolVLM-256 on UAV edge hardware.
Conclusion This article presented the transition from cloud-centric visual processing toward fully edge-native semantic inspection powered by lightweight VLMs. It categorized existing UAV inspection architectures, discussed their associated system challenges and operational tradeoffs, and demonstrated the feasibility of semantic edge intelligence on UAV-class hardware using fine-tuned YOLO26M localization and lightweight SmolVLM-256 reasoning. The results confirmed that semantic UAV inspection is technically feasible on modern UAV-class edge platforms, although scalable deployment requires hierarchical, resource-aware, and trustworthy semantic processing pipelines.
REFERENCES 1. R. Li, J. Yu, F. Li, R. Yang, Y. Wang, and Z. Peng, “Automatic Bridge Crack Detection Using Unmanned Aerial Vehicle and Faster R-CNN,” Construction and Building Materials, vol. 362, p. 129659, 2023. September 2026
2. G.-H. Gwon, J. H. Lee, I.-H. Kim, and H.-J. Jung, “CNN-based Image Quality Classification Considering Quality Degradation in Bridge Inspection Using an Unmanned Aerial Vehicle,” IEEE Access, vol. 11, pp. 22 096–22 113, 2023. 3. W. Ji, S. Liu, L. Deng, J. Li, Y. Liu, and Z. Xiong, “WDI-YOLO: A Lightweight Steel Bridge Weld Defect Detection Algorithm Using UAV Images,” Journal of Constructional Steel Research, vol. 235, p. 109833, 2025. 4. S. Ruggieri, A. Cardellicchio, A. Nettis, V. Renò, and G. Uva, “Using Attention for Improving Defect Detection in Existing RC Bridges,” IEEE Access, vol. 13, pp. 18 994–19 015, 2025. 5. P. Li, Y. Peng, S.-M. Wang, and C. Zhong, “Improved RT-DETR Framework For Railway Obstacle Detection,” IEEE Access, 2025. 6. H. Ding, C. Zhou, and C. Lian, “An Improved RTDETR Model for UAV-Based Power Line Inspection Under Various Weather Conditions,” in China Automation Congress. IEEE, 2024, pp. 941–945. 7. S. Rong, L. He, S. F. Atici, and A. E. Cetin, “Ad-
From Pixels to Semantics: Edge AI for UAV-Based Critical Infrastructure Inspection
7
IEEE Internet Computing
vanced YOLO-based Real-time Power Line Detection for Vegetation Management,” IEEE Transactions on
tributed multimedia, agentic edge AI, and Green AI. Contact him at: [email protected]
Power Delivery, 2025. 8. G. Li, C. Wang, Q. Sun, Z. Chen, and H. Wang, “VLM-PI: Power Inspection System Based on Visual Large Models,” in 28th International Conference on Computer Supported Cooperative Work in Design (CSCWD). IEEE, 2025, pp. 2152–2157. 9. H. Zhang, X. Pan, M. Asgarinejad, and T. Song, “Multimodal Large Language Model-driven Framework For Road Crack Assessment,” Automation in Construction, vol. 185, p. 106872, 2026. 10. Z. Chen, Y. Zou, V. A. González, J. Ingham, and L. M. Wotherspoon, “Bridge Inspection Using a Multimodal Vision Language Model,” in Proceedings of The Sixth International Conference on Civil and Building Engineering Informatics, vol. 22, 2025, pp. 578–588. 11. E. Bianchi, A. L. Abbott, P. Tokekar, and M. Hebdon, “COCO-bridge: Structural Detail Data Set for Bridge Inspections,” Journal of Computing in Civil Engineering, vol. 35, no. 3, p. 04021003, 2021. 12. Ultralytics, “Ultralytics yolo26,” https://docs.ultralytics. com/models/yolo26/, accessed on 10-May-2026. 13. A. Marafioti, O. Zohar, M. Farré, M. Noyan, E. Bakouch, P. Cuenca, C. Zakka, L. B. Allal, A. Lozhkov, N. Tazi et al., “SmolVLM: Redefining Small and Efficient Multimodal Models,” arXiv preprint arXiv:2504.05299, 2025. 14. Z. Han, C. Gao, J. Liu, J. Zhang, and S. Q. Zhang, “Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey,” arXiv preprint arXiv:2403.14608, 2024. 15. E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al., “LoRA: Low-Rank Adaptation of Large Language Models.” ICLR, 2022. 16. NVIDIA, “Tegrastats Utility,” https://docs.nvidia. com/drive/drive_os_5.1.6.1L/nvvib_docs/index.html# page/DRIVE_OS_Linux_SDK_Development_Guide/ Utilities/util_tegrastats.html, accessed: 2026-06-15.
Reza Farahani is a Univ.Ass. Postdoctoral Researcher at the Distributed Systems Group (DSG), TU Wien, and a Lecturer at the University of Klagenfurt, Austria, where he received his Ph.D. in Computer Science in 2023. He has recently coordinated the Austrian EdgeAI-Drone project and participated in the EUfunded ENFIELD Exchange Scheme on Green drone AI. Previously, he contributed to the Christian Doppler Laboratory ATHENA and EU Graph-Massivizer project, co-leading WP5 on serverless orchestration and largescale edge-cloud testbeds. His research interests include distributed systems, serverless computing, dis8
Naser Hossein Motlagh is a Senior Researcher at the University of Helsinki, Finland. He received his D.Sc. in Networking Technology at Aalto University, Finland, in 2018. His research interests include the Internet of Things, wireless sensor networks, edge AI, environmental sensing, and unmanned aerial and underwater vehicles. Contact him at: [email protected] Zoha Azimi is a Ph.D. candidate at the University of Klagenfurt, Austria. She received her M.Sc. in artificial intelligence from the University of Bologna, Italy. Her research interests include AI, energy-efficient multimedia systems, and agentic AI. Contact her at: [email protected] Christian Timmerer is a Full Professor with the University of Klagenfurt, Austria, and the Director of the Christian Doppler Laboratory ATHENA. He is also a CoFounder and Chief Innovation Officer of Bitmovin. His research interests include multimedia systems, adaptive streaming, and drone communication. Contact him at: [email protected] Lorenzo Carnevale is an Assistant Professor at the University of Messina, Italy. His research interests include edge intelligence and AI for resource-constrained environments. He has led technical activities in Horizon Europe projects on natural-disaster applications and distributed TinyAI. Contact him at [email protected] Sasu Tarkoma is a Professor with the University of Helsinki and the University of Oulu, Finland. He received the Ph.D. degree in computer science from the University of Helsinki, in 2006. His research interests include mobile computing, data sciences, and AI. Contact him at: [email protected] Schahram Dustdar is a Full Professor of Computer Science and Head of the Distributed Systems Group at TU Wien, Austria and ICREA Professor in Barcelona, Spain. He is serving as an Associate Editor for several IEEE and ACM journals and is EiC of Computing (Springer). His distinctions include the ACM Distinguished Scientist and Speaker and IBM Faculty Awards. He is member of the Academia Europaea. Contact him at: [email protected].
From Pixels to Semantics: Edge AI for UAV-Based Critical Infrastructure Inspection
September 2026