ConceptioArchivearXiv CS
arXiv CSopen access

Look Once, Beam Twice: Camera-Primed Real-Time Double-Directional mmWave Beam Management for Vehicular Connectivity

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
knowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Look Once, Beam Twice: Camera-Primed Real-Time Double-Directional mmWave Beam Management for Vehicular Connectivity Avhishek Biswas∗,1 , Apala Pramanik∗,1 , Eylem Ekici2 , Mehmet C. Vuran1 1

arXiv:2605.05071v1 [cs.NI] 6 May 2026

School of Computing, University of Nebraska–Lincoln, Lincoln, NE, USA {abiswas3, apramanik2}@huskers.unl.edu, [email protected] 2 Electrical and Computer Engineering, The Ohio State University, Columbus, OH, USA [email protected]

Abstract—Millimeter-wave (mmWave) frequencies promise multi-gigabit connectivity for vehicle-to-everything (V2X) networks, but face challenges in terms of severe path loss and mobility-related beam misalignment. Reliable V2X connectivity requires fast, double-directional beam alignment. However, existing methods suffer from high training overhead and limited generalization to unseen scenarios. This paper presents VIsion-based BEamforming (V I B E), a hybrid model-based, closed-loop, learning architecture for real-time double-directional mmWave beam management primed by camera sensing. V I B E fuses machine learning, model-based reasoning, and closed-loop RF feedback to balance beam-pair establishment latency with link quality. V I B E bypasses exhaustive training overhead and accelerates link establishment by leveraging camera observations to reduce the beamsearch space. Lightweight beam refinement and offset tracking mechanisms adaptively refine beams in response to dynamic application requirements. V I B E is implemented and evaluated across online indoor/outdoor testbeds, public datasets, and realtime vehicular experiments, demonstrating strong generalization capabilities, making it suitable for real-time V2X communication. Comparisons with 5G NR hierarchical beamforming show that V I B E consistently maintains lower outage rates. Furthermore, V I B E outperforms state-of-the-art end-to-end ML models for beam selection when evaluated on public datasets and achieves outage rates as low as 1.1–1.4%. The results show that a hybrid model-based, closed-loop learning architecture is better suited for real-world mmWave vehicular connectivity than end-to-end trained ML models. For reproducibility, we publish our code to https://github.com/UNL-CPN-Lab/Look-Once-Beam-Twice. Index Terms—mmWave, 6G, Beamforming, V2X networks

I. I NTRODUCTION The large-scale deployment of 5G and the evolution toward 6G enable ultra-high data rates, low latency, and ubiquitous connectivity for intelligent transportation systems and V2X networks [1]–[5]. To support this vision, 3GPP Release 20 emphasizes integrated sensing, beam-based mobility, predictive handovers, and AI-driven decision-making for high-mobility environments [6]. To this end, millimeter-wave (mmWave) frequencies (30–300GHz) are promising candidates for next generation vehicular networks, supporting high-rate V2X links [7]. mmWave frequencies enable compact antenna arrays at the BS and UE, allowing highly directional beams to overcome severe path loss [8]. However, beam misalignment can cause over 10dB power loss [9], leading to rapid performance degradation under mobility. Measurements from commercial 5G networks ∗ Avhishek Biswas and Apala Pramanik contributed equally to this work.

Dir - Omni

BS

-98dBm

Double - Directional

SL - Omni

UE

BS

UE

BS

UE

-24dBm

Fig. 1. Double-directional links improve worst-case received power by 21.8dB and 49.1dB compared to directional-to-omnidirectional (Dir-Omni) and sector-level directional to omnidirectional (SL-Omni) links, improving potential cell size (via Remcom Wireless InSite [11]).

show that only a small subset of available beams is used in practice, with beam refinement largely base station(BS)-centric and inconsistent across operators [10]. Emerging 5G-integrated vehicular platforms increasingly incorporate beam-steerable antenna arrays at both the BS and UE, as shown by recent prototypical and experimental systems [1], [3]. This enables double-directional beamforming at link access, improving link budget and coverage over single-directional links [12]. The impact of utilizing double-directional links during the link access stage is illustrated in Fig. 1, using FR2 urban coverage simulations at 60 GHz in Wireless InSite [11]. Existing mobile mmWave beam management typically assumes either (a) directional-toomnidirectional (Dir–Omni) links, where the BS uses directional beams and the UE is omnidirectional, or (b) sectorlevel directional-to-omnidirectional (SL–Omni) links, where wide BS sector beams (e.g., 5G NR SSBs) reduce acquisition delay. While effective for reducing initial access latency, both approaches effectively reduce mmWave cell size. Despite these gains, exhaustive double-directional beam search incurs high latency and computational cost, scaling as O(N 2 ) [7], which is prohibitive for V2X applications [13]. Reduced-complexity methods, such as hierarchical, coded, and compressed sensing– based beamforming [14], mitigate this overhead but adapt poorly to rapid channel variations, limiting the coverage gains of double-directional access under mobility. To meet stringent V2X timing and reliability requirements, prior work has focused on reducing beam alignment latency

using onboard sensors and end-to-end learning models [7]; however, these approaches often fail to generalize across environments and mobility conditions. In this paper, we address the problem of low-latency double-directional beam alignment under SNR constraints. We propose V I B E: VIsionbased BEamforming, a lightweight and adaptive beamforming framework for vehicular communication. Unlike end-to-end learning approaches that directly map raw sensory inputs to beam decisions, V I B E uses a hybrid closed-loop design. The framework combines model-driven beam selection with online SNR feedback. Initial beam estimation is decoupled from runtime adaptation allowing for a low-overhead coarse beam decision using camera priming and radio coordinate projection. The beam pair is then refined online through a lightweight, iterative SNR-driven feedback. This design reduces beam management overhead compared to 5G NR and learning-based models. It also improves link reliability. Faster beam alignment and lower outage are achieved under different SNR constraints, therefore not requiring retraining and large-scale labeled RF datasets. We validate VIsion-based BEamforming through dynamic indoor and outdoor experiments and benchmark V I B E against current 5G beam management standard [10] and state-of-theart methods on public datasets. Key Contributions. Our contributions are as follows: • We present V I B E , a practical, double-directional beam alignment framework that is hardware-agnostic and requires no offline RF training, operating seamlessly with diverse camera setups. • V I B E introduces a hybrid closed-loop adaptation mechanism that combines iterative beam refinement, offset tracking, and a bespoke learning framework, enabling real-time adaptation to SNR dynamics in mobile environments. • We evaluate V I B E across online indoor/outdoor testbeds, public datasets, and multiple camera configurations, demonstrating impressive generalization capabilities, making it suitable for real-time V2X communication, while significantly reducing beam search overhead. • We publicly release the datasets, trained models, and associated code.1 This includes a novel vision model trained on RF street furniture and evaluated in diverse urban scenes. The remainder of the paper is organized as follows. Related work is discussed in Section II, followed by an overview of V I B E in Section III. The beam alignment problem and V I B E design are presented in Sections IV and V, respectively. Evaluation results and conclusions appear in Sections VI-A1 and VII. II. R ELATED W ORK Commercial 5G mmWave networks exhibit coverage and performance limits despite dense deployments. In a 2.23 km2 Chicago area with over 34 base stations, coverage reached 1 https://github.com/UNL-CPN-Lab/Look-Once-Beam-Twice

only 35%, with throughput degrading under distance and mobility [15] and inefficient spectrum use due to slow interband switching [16]. Beam management further constrains performance by using a few beams (typically ≤ 5 of 36), relying on single-directional gNB-driven refinement, and switching beams infrequently (0.6–2.8 s) [10]. Coordinated Tx/Rx steering with narrow beams can instead deliver up to 14 dB SINR gains and 200 Mbps higher throughput [10], motivating double-directional beamforming. Double-directional beamforming is computationally expensive due to exhaustive transmit–receive beam pair scanning. Prior work reduces this cost using LBC-based discovery [17], Kolmogorov model–based learning [14], and compressed sensing or structured search, but these methods rely on static sparsity and degrade under mobility. To mitigate beam alignment overhead under mobility, recent work leverages side information from sensors such as radar, LiDAR, inertial sensors, and cameras. Among these, cameras are particularly attractive due to their low cost and rich spatial information, and have proven effective for line-of-sight beam prediction [18]. However, most vision- and sensor-assisted beam prediction methods remain BS-centric, offline, and non-adaptive, including CNN-based sector prediction using fixed mappings and indoor data [19], image-based beam inference trained on synthetic datasets [20], and learning-driven mobility-aware approaches based on sequence modeling or multi-modal fusion, such as CNN+GRU proactive handover [21], vision–position fusion for top-k prediction [22], and LiDAR/GNSS-based recurrent tracking [23]. Moreover, BS-side camera deployment raises privacy concerns [24], whereas modern vehicles already integrate multiple cameras for perception and driver assistance [25], making UEside vision a practical and privacy-preserving alternative. While these deep learning–based methods improve prediction accuracy, they operate as black-box models, lack online RF validation, and are typically evaluated offline, limiting robustness and generalization. Related efforts incorporating inertial sensors, encoder–decoder architectures, semantic awareness, or geometric scene reconstruction further reduce training overhead or input dimensionality [26]–[30], but remain largely scenario-specific, and restricted to controlled settings; notably, the semantic-aware approach in [29], which replaces raw images with object masks or bounding boxes, serves as a key baseline in our evaluation. More recent multicamera and multimodal frameworks improve beam search efficiency and tracking robustness [31], [32], while UE-centric methods such as Omni-CNN and FLASH-and-Prune reduce complexity by focusing on SL–Omni links [33], [34], albeit at the cost of reduced cell size and limited adaptability. In summary, existing double-directional beamforming approaches reduce overhead at the expense of adaptability or leverage sensing without real-time RF feedback. In Section VI-A1, we evaluate offline end-to-end ML-based methods, which fail to generalize under mobility and dynamic channels, and lack closed-loop adaptation. These gaps motivate a realtime, sensing-driven, and feedback-enabled framework for robust beam alignment in dynamic vehicular environments.

Camera Priming 1

UE

2

Camera Coordinate System

Ow BS

BS

UE

Iterative Beam Refinement

Beam Initialization

Radio Coordinate Projection 3

UE

UE

Offset Tracking

4

5

BS

BS

UE

Oc

World Coordinate System

BS

Fig. 2. Overview of V I B E.

III. OVERVIEW We present an overview of V I B E, a camera-primed realtime double-directional link establishment framework for V2X mmWave networks in Fig. 2. V I B E combines machine learning, model-based reasoning, and closed-loop RF feedback to balance beam alignment latency and link quality. V I B E distinguishes itself through three innovative approaches: Double-directional Link Establishment. As emerging vehicular prototypes increasingly integrate dedicated 5G/mmWave radios and antenna arrays [1], [3], a key challenge emerges: beam acquisition protocols limit cell size despite the feasibility of double-directional links in V2X networks. V I B E addresses this gap with a rapid double-directional link establishment approach that leverages camera-priming to reduce the beam search space. UE-centric Design. Existing sensor-based beam acquisition solutions primarily assume BS–centric sensors (e.g., cameras, LiDAR, radar) [22], [35]–[37], raising privacy concerns that hinder real-world deployment [24]. In contrast, V I B E adopts a UE–centric design that assumes cameras already deployed in modern vehicles for perception and driver assistance [25], [38]. Consequently, V I B E mitigates privacy concerns, as camera data remains within the vehicle, consistent with current deployment trends. Section VI-A1 further shows that V I B E can be readily deployed in a BS-centric manner while still outperforming existing approaches in delay and performance. Real-time Focus. V I B E is designed for real-time operation. Rather than relying on end-to-end machine learning workflows with excessive delays, we introduce a hybrid modelbased, closed-loop learning architecture for real-time operation. V I B E is implemented and evaluated in live vehicle-toinfrastructure experiments. We show that V I B E can maintain signal-to-noise ratio (SNR) requirements within reasonable delays compared to the state-of-the-art. V I B E consists of five components: (1) Camera Priming, which detects the BS and provides an initial direction estimate to reduce beam search; (2) Radio Coordinate Projection, which maps camera coordinates to radio coordinates using a model-based approach, enabling adaptation across cameras and preserving privacy; (3) Beam Initialization, which converts the direction estimate into UE and BS beam indices compatible with discrete beambooks; (4) Iterative Beam Refinement, which performs a fast local sweep to meet SNR

Mobile Vehicle (MV)

MCS Base Station (BS) User Equipment (UE) WCS

Fig. 3. The V2X communication scenario.

requirements; and (5) Offset Tracking, which maintains residual angle corrections to further reduce beam-pair establishment delay. Experimental results show that this hybrid model-based, closed-loop architecture generalizes effectively under mobility. IV. P RELIMINARIES We consider an uplink mmWave V2I communication scenario as shown in Fig. 3, where a mobile vehicle (MV), equipped with a radio and a camera, communicates with a base station (BS). We define four coordinate systems that are leveraged throughout the paper: World coordinate system (WCS), mobile coordinate system (MCS), radio coordinate system (RCS), and the camera coordinate system (CCS). A. System Setup Within the WCS, the global position of an object (e.g., the mobile vehicle) is defined by its location pm|W = [xm|W , ym|W ], where XW points East and YW points North, and its heading (yaw) angle ψm|W , w.r.t. the YW axis. Similarly, the UE, BS, and the camera are defined as (pue|W , ψue|W ), (pbs|W , ψbs|W ), and (pc|W , ψc|W ), respectively. Within the MCS, (pue|M , ψue|M ) and (pc|M , ψc|M ) denote the positions and orientations of the UE and camera, respectively, w.r.t. the vehicle. For both the UE and the BS radios, RCS is used to represent the beamforming angles Θue|R and Θbs|R w.r.t. their boresight, respectively. CCS will be utilized to represent the images observed by the camera in the following.

B. Channel Model

BS

Assuming both BS and the UE are equipped with uniform linear arrays (ULAs) of NB and NU antennas, respectively, the received signal at the BS is given by [39]: H

y(t) = wbs (t) H(t)wue (t)s(t) + n(t), NU ×1

(1)

NB ×1

where wue (t) ∈ C and wbs (t) ∈ C are the UE transmit precoder and BS receive combiner, respectively, H(t) ∈ CNB ×NU is the uplink channel matrix, s(t) ∈ C is the transmitted signal, and n(t) ∼ CN (0, σ 2 ) is the complex Gaussian noise with zero mean and variance σ 2 . Accordingly, SNR is given by: H wbs (t)H(t)wue (t) γ(wue , wbs ) = σ2

2

.

(2)

C. Problem Definition Assumptions. We assume predefined and fixed beambooks B = {b1 , ..., b|B| } and U = {u1 , ..., u|U | } at the BS and UE, respectively, with overlapping beams of fixed width. Furthermore, the beambook indices are defined as kbs ∈ {1, ..., |B|} and kue ∈ {1, ..., |U|}, and the beambook beam angles are (kbs ) (kue ) denoted as Θbs|R and Θue|R . The BS and UE operate under LoS conditions with aligned, parallel boresights. Extension to non-line-of-sight conditions is considered out of scope and constitutes our future work. The boresight assumptions could be easily relaxed through existing pose estimation solutions [40]. The channel matrix H(t) is time-varying due to environmental dynamics and mobility, and the optimal ∗ ∗ , kbs are unknown. transmit/receive beamforming indices kue Additionally, the packet size per beam search is fixed, the SNR constraint, γth , is application-specific and given. Problem. Accordingly, our goal is to design an online beam-pair selection policy: π : (C, ZP , M) −→ (kue , kbs ) , such that γ(U[kue ], B[kbs ]) ≥ γth , where the inputs to the policy are camera observations, C, pilot signal measurements, ZP , and the internal memory carried over from previous iterations, M. Exhaustive beam sweeping is time-prohibitive in mobile scenarios. Our goal is therefore to use camera observations and accumulated memory to reduce the beam search space while maintaining acceptable link quality under mobility. V. V I B E : VI SION - BASED BE AMFORMING In this section, we present V I B E (Fig. 2) and describe its five components as illustrated in Fig. 2.

Zw Ow World Yw Coordinate System

Xw

Camera Focal Plane Focal Point (xcf,ycf ) f Z=

Yi

BS Bounding Box Xi Xc,j

Optical Axis

UE LoS Angle

Zc

Oc

Xc

Yc

Camera Coordinate System

Fig. 4. Pinhole camera model for LoS angle estimation [41].

A. Camera Priming V I B E reduces the beam-pair search space via camera priming by detecting the BS in the UE camera view. As shown in Fig. 4, an object detection model trained on street radio furniture processes the images and outputs bounding boxes, B, as  J B = (xc,j|C , yc,j|C , ℓj , cj ) j=1 , (3) where (xc,j|C , yc,j|C ) denotes the center pixel coordinates of the j-th bounding box w.r.t. the camera coordinate system, ℓj is the predicted class label (e.g., “radio”), and cj ∈ [0, 1] is the corresponding confidence score. The total number of detected objects is denoted by J. We assume a single BS is visible in the image, which is reasonable given typical deployment densities. The BS coordinates are then projected into radio coordinates, as described next. B. Radio Coordinate Projection Upon detection of the BS, the horizontal pixel of the bounding-box center, xc,j , represents the azimuthal displacement of the BS w.r.t. the camera optical axis (Fig. 4). Accordingly, the estimated LoS angle in the CCS is [41]:   (xc,j|C (t) − xcf ) · P , (4) θ̂ue|C (t) = atan2 f where xcf is the focal point abscissa, P is the pixel pitch (in meters), and f is the focal length of the camera. Next, we project this estimation first to the MCS and then to the RCS. The camera and the radio are mounted on the MV with yaws, ψc|M and ψr|M , respectively. Then, the LoS angle is projected into the RCS as: θ̂ue|R (t) = θ̂ue|C (t) + ψc|M − ψr|M .

(5)

This estimation is utilized to initialize the beam pairs. It is important to note that the camera-aided beam estimation is subject to noise from measurement errors, calibration drift, and limited resolution, introducing angular error in the estimated LoS direction. εc (t) = θue|R (t) − θ̂ue|R (t),

with |εc (t)| ≤ δc ,

(6)

where δc denotes the maximum sensor-induced angular deviation under expected operating conditions. Since both the

radio and the camera are mounted on the vehicle, the vehicle heading does not affect this projection. Commercial mmWave

C. Beam Initialization The azimuth angle estimate, θ̂ue|R (t), from the RCS projection is then quantized to find the closest beam index in the UE beambook: kue (t) = arg

min

(k)

k∈{1:|U |}

Θue|R − θ̂ue|R (t) .

(7)

The UE transmits the index kue (t) (or equivalently (kue (t)) Θue|R ) to the BS over a sub-6 GHz control link. Assuming parallel boresights, the BS estimates the beamforming angle at the opposite azimuth: (k

(t))

ue θ̂bs|R (t) = Θue|R

+π ,

(8)

and quantizes its prediction similarly. Note that this quantization introduces angular mismatches that are bounded by the half beam spacing of each UE and BS beambooks, as we address next. Algorithm 1 V I B E-MA: Beam Selection with Iterative Beam Refinement and Offset Tracking Require: Predicted beam bpred , SNR Th. γth , Guard δ 1: if bpred ̸= None then 2: γ ← MeasureSNR(bpred ); nbeams ← 1 3: if γ ≥ γth then return (bpred , γ, nbeams ) 4: end if 5: if hist[] is not empty then 6: badj ← bpred + mean(hist[]) 7: γ ← MeasureSNR(badj ); nbeams ← 2 8: if γ ≥ γth then return (badj , γ, nbeams ) 9: else return LocalBeamRefinement(badj ) 10: end if 11: else return LocalBeamRefinement(bpred ) 12: end if 13: end if 14: Function LocalBeamRefinement(bc ) 15: b∗ ← None; γ ∗ ← −∞; nbeams ← 0 16: for δ = 1 to δmax do 17: for d ∈ {−1, +1} do 18: b ← bc + dδ; γ ← MeasureSNR(b) 19: nbeams ← nbeams + 1 20: if γ ≥ γth then 21: Append dδ to hist[] return (b, γ, nbeams ) 22: else if γ > γ ∗ then b∗ ← b; γ ∗ ← γ 23: end if 24: end for 25: end for 26: return (b∗ , γ ∗ , nbeams )

D. Iterative Beam Refinement and Offset Tracking Camera-primed beam initialization suffers from calibration drift, noise, and beam codebook quantization, while mobility

Streetlight

Radio 5G BS

Fig. 5. V I B E-YOLOR model sample detections of mmWave base stations and infrastructure in diverse urban scenes at Lincoln, Nebraska,USA.

introduces temporal drift that degrades SNR—effects often missed by offline methods. To address this, we designed a fast local beam sweep with refinement and offset tracking, with two variants: V I B E-MA, which uses a moving average of past offsets, and V I B E-MLP, a lightweight neural network for direct correction. V I B E-MA. The procedure is shown in Algorithm 1. It first checks whether the SNR from beam initialization exceeds the threshold γth ; if so, the predicted beam is accepted. If no offset history exists (e.g., at initialization), the UE performs local refinement around the predicted beam using an alternating +1, −1, +2, −2, . . . search. For each candidate, the UE signals the beam angle to the BS over a sub-6 GHz link and measures the SNR. The search terminates once a beam meets the threshold, and the resulting offset ô relative to the prediction is stored. If no beam satisfies the threshold, the beam with the highest SNR is selected, and the offset history is not updated. The UE maintains an offset history using a moving average of past offsets to correct the predicted beam, yielding badj . If badj fails to meet the threshold, the algorithm falls back to local beam refinement. V I B E-MLP. In addition to the rule-based moving average, we evaluate a learned adaptation method, V I B E-MLP, which directly predicts the beam offset using a trained black-box model. The network consists of three fully connected layers with LayerNorm, ReLU activations, and dropout, and outputs a single offset value trained using Smooth L1 loss and the Adam optimizer. At runtime, if the initial SNR falls below the threshold, V I B E-MLP infers the corrective offset. E. Implementation To implement the V I B Eframework, we develop an object detection pipeline for BS identification for beam initialization using YOLOv11 [42], pre-trained on MS-COCO and fine-tuned on four curated datasets: indoor mmWave radios, commercial mmWave antennas (TG Sounders [43]), deployed 5G small cells, and urban streetlights emulating real-world mmWave deployments [10]. This enables detection of four additional classes beyond COCO, improving adaptability across scenarios (Fig. 5). The resulting detector, combined with the initial beam estimation in Section V-A, is referred to as V I B EYOLOR and serves as an internal baseline. Building on this detection capability, we evaluate the runtime efficiency of the V I B E pipeline. Fig. 6 shows an average end-to-end latency of 0.231s. Image processing accounts for 0.075s (40%), beam configuration at the UE and BS for 0.050s

UE

BS

Request UE Beam Index Capture Image Detect BS Estimate Beam Index Send UE Beam Index Set UE Beam Send UE Beam Index Estimate BS Beam Set BS Beam Send ACK Stabilize Beam-pair Measure SNR 0.5%(1ms)

0.5%(1ms)

40.0%(75ms) 0.5%(1ms)

0.5%(1ms)

26.6%(50ms)

0.5%(1ms)

26.5%(50ms)

18.6%(35ms)

8.0%(15ms)

0.5%(1ms)

Fig. 6. Timing breakdown of the V I B E pipeline (Percentages indicate perstep time, EC: Embedded computer at UE).

(26.5%), and beam stabilization with SNR measurement for another 0.050s (26.5%). These stages contribute over 90% of the total latency and are primarily hardware dependent, making them key targets for optimization. VI. E VALUATIONS This section presents a comprehensive evaluation of the proposed double-directional beamforming solutions. We first analyze V I B E and its internal baselines in controlled indoor experiments (Section VI-A), then compare V I B E against stateof-the-art methods on public datasets using outage, coverage, and beam alignment time. Finally, we conduct real-time outdoor experiments to assess latency and outage under dynamic channel conditions (Section VI-C). A. Indoor Evaluations 1) Experiment Setup: Indoor experiments are conducted inside the Cyber Physical Networking Lab, Schorr Center, University of Nebraska-Lincoln, using a controlled setup with a fixed BS and a motorized UE to emulate vehicle mobility (Fig. 7). Both nodes employ Sivers Semiconductors 60 GHz EVK06002 phased-array front-ends in the n263 FR2 band with USRP B200-mini SDRs for baseband processing. -90

Boresight ω Beamforming θ Angle

+90

LoS +90 UE on rotating platform

-90 Base Station

Fig. 7. Indoor testbed at CPN Lab, University of Nebraska-Lincoln, USA.

Each phased-array provides 64 analog beams spanning ±45◦ in 1.5◦ steps, with half-power beamwidth of 6◦ in

azimuth and 18◦ in elevation. The UE is rotated through 180◦ at angular speeds of 0.25, 1, and 4◦ /s to emulate vehicle motion. Camera-priming is evaluated by either a 60◦ narrow field of view (NFOV) Intel RealSense camera or a 90◦ wide FOV (WFOV) Luxonis OAK-D camera. We evaluate V I B E-YOLOR, V I B E-MLP, and V I B E-MA against ground truth measurements from exhaustive double-directional beampair sweeps. Based on the ground truth SNR distributions, three SNR thresholds are defined at the 80th , 90th , and 95th percentiles, for consistent evaluation. In the evaluations, outage probability is measured against ground truth, where an outage is recorded if the algorithm fails to select any beam pair exceeding the SNR threshold. The beam alignment time, Tb , is measured using clock_gettime() from algorithm start until beam selection, at which point the SNR is recorded. Since latency depends on hardware and implementation, the reported delays serve as baseline measurements for fair comparison rather than fundamental limits.

30 Outage (%)

EC

20

VIBE-MA (Offline) VIBE-MA (Online) VIBE-YOLOR (Offline) VIBE-YOLOR (Online)

10 0

Q₀.₈₀ Q₀.₉₀ Q₀.₉₅ Quantile Based Threshold Fig. 8. Offline vs. online evaluations.

2) Evaluation Results: Offline vs. Online Evaluations. Recent mobile mmWave studies rely on offline evaluations with live images but pre-collected SNR, which omit fast fading and hardware delays. To show this effect, we compare this offline setting with online evaluation, where SNR is measured in real time. In Fig. 8, we report outage for V I B E-YOLOR and V I B E-MA across Q0.80 , Q0.90 , and Q0.95 . For V I B E-YOLOR, outage at Q0.95 increases from 4.7% offline to 33.1% online, showing that offline results overstate reliability. In contrast, V I B E-MA outage increases by less than 1.8 pp across all thresholds. Offset tracking further reduces outage by 0.8 pp at 0.25◦ /s and 0.5 pp at 4◦ /s, demonstrating robustness under both slow and fast rotations. SNR Adaptation. In Fig. 9, we present a sample realtime performance of V I B E-MA with WFOV camera, UE rotating at 1◦ /s, where the dashed lines are the SNR thresholds. It can be observed that, V I B E-MA dynamically adapts to the SNR criteria, consistently maintaining higher SNR levels while achieving outages of 13.7% (Q0.80 ), 11.7% (Q0.90 ), 0% (Q0.95 ) on average. This showcases V I B E-MA’s capability to adapt to different SNR thresholds.

VIBE-MA(Q₀.₉₅) SNR Th. Q₀.₉₅

SNR (dB)

26

VIBE-MA(Q₀.₉₀) SNR Th. Q₀.₉₀

VIBE-MA(Q₀.₈₀) SNR Th. Q₀.₈₀

TABLE I Indoor outage and beam alignment time (Tb ) under varying SNR thresholds and rotation speeds.

22 YOLOR

MLP

NFOV Out. Tb (s) (%)

NFOV Out. Tb (s) (%)

NFOV Out. Tb (s) (%)

WFOV Out. Tb (s) (%)

0.25

4.1

0.09

3.8

0.27

0.3

0.22

6.2

0.13

1.00

5.0

0.09

4.0

0.22

2.7

0.22

13.7

0.68

4.00

5.6

0.09

5.4

0.22

5.4

0.22

27.2

1.77

0.25

4.9

0.09

4.0

0.22

3.5

0.22

2.7

0.13

1.00

5.9

0.09

3.6

0.23

4.5

0.23

11.7

1.17

4.00

24.5

0.09

7.8

0.25

6.0

0.26

25.0

1.42

0.25

22.2

0.09

4.1

0.26

1.8

0.25

4.4

0.40

1.00

33.1

0.09

3.4

0.33

3.6

0.35

0.0

2.72

4.00

64.1

0.09

10.0

0.31

11.1

0.50

28.5

1.10

18 SNR Th.

14 10

Q0.80

45

30

15

0

15 −

30 −

45

6 Boresight Angle (°)

Q0.90 Fig. 9. SNR performance of V I B E-MA with WFOV camera at an angular speed of 1◦ /sec.

Comparison with 5G NR. In Fig. 10, we compare V I B EMA with 5G NR under increasing rotation speeds. 5G NR shows high SNR outages. Outage probability exceeds 50% across all quantiles. The beam alignment time remains in the range of 7.7−8.0 s. When beam switching is deferred until the best beam pair is identified (rotation speed = 0), outage reduces to 14%, 16%, and 23% for Q0.8 , Q0.9 , and Q0.95 , respectively. Outage increases to nearly 100% as rotation speed increases. In contrast, V I B E-MA achieves much lower beam alignment

Fig. 10. Beam alignment time (Tb ) and outage probability comparison of V I B E-MA with 5G NR

time and outage. At the highest rotation speed, V I B E-MA maintains Tb = 0.5 s for Q0.95 with outage remaining below 12%. These results show V I B E-MA’s robustness to mobilityinduced angular dynamics. As low-latency beamforming is critical for sustaining connectivity in mobile mmWave scenarios, conventional 5G NR hierarchical beamforming struggles under mobility. Internal Baselines. Finally, we provide a comprehensive comparison in Table I under different SNR thresholds and speeds, and different cameras in an online setting. In majority of the cases, V I B E-MA achieves the lowest outage, reducing outage by up to 29.5pp (e.g., from 33.1% to 3.6% at Q0.95 and 1◦ /s) albeit with an increase in alignment time from 0.09s to 0.35s. V I B E-YOLOR maintains a constant beam alignment time, which may be desirable in low-speed and low SNR threshold conditions. While V I B E-MLP occasionally outperforms V I B E-MA (e.g., 3.6% outage at Q0.90 , 1◦ /s), the differences are marginal and inconsistent, emphasizing V I B E-MA’s overall generalizability. Compared to V I B E-MLP, which is trained on prior indoor offset data, V I B E-MA reduces

Q0.95

Speed (deg/s)

MA

outage by up to 1.8 pp while achieving similar beam alignment times, highlighting the benefit of its hybrid closed-loop design for meeting real-time SNR thresholds. Furthermore, V I B EMA is hardware agnostic. When tested with a WFOV camera, V I B E-MA achieves 0% outage at Q0.95 , 1◦ /s, confirming its robustness across different sensing configurations. At higher speeds (4◦ /s), the WFOV configuration incurs a 17.4 pp higher outage than NFOV (11.1%) due to reduced angular resolution. Coarser quantization under fast motion increases beam uncertainty and corrective search time, raising Tb from 0.50 s (NFOV) to 1.10 s (WFOV). Overall, V I B E-MA reduces outage by up to 53pp and 3.5pp compared to V I B E-YOLOR and V I B E-MLP, respectively. Although this incurs a modest increase in alignment time, it remains suitable for real-time operation (Section VI-C). The closed-loop hybrid design is hardware agnostic and robust in real-time scenarios. B. State-of-the-art Comparisons 1) Experiment Setup: To evaluate robustness and crossscenario generalization, we compare V I B E-YOLOR and V I B E-MA against two state-of-the-art baselines: MobileNet+LeNet (MNet–LeNet) [29], trained on Scenario 7 of [35], and ResNet-50 [22], trained on Scenario 6 of [35]. Both baselines are evaluated using standard top-k beam prediction accuracy (e.g., top-1 and top-3), with generalization tested on unseen Scenario 9 [35]. To ensure a fair comparison, we also evaluate V I B E in BS-centric configurations, demonstrating its adaptability beyond UE-centric operation. It is important to note that YOLOR was not trained in any of these scenarios, making all of them unseen. Model performance is evaluated in terms of outage, which is based on the number of instances where the predicted beam power falls under the received power threshold, and beam alignment time, which accounts for image inference and beamforming delay derived from indoor measurements. We deploy the opensource models, as is, locally and all evaluations are conducted on an NVIDIA A2000 GPU (12GB VRAM) using quantilebased normalized received power threshold of (Q0.80 , Q0.90 ,

9

6,7

V I B E-MA

MNet + LeNet [29]

(.

Scenario

TABLE II Outage probability and beam alignment time for V I B E-MA against MNet+LeNet [29] (trained on Scen. 7) and ResNet-50 [22] (trained on Scen. 6). Scen. 9 serves as an unseen environment for all models. ResNet-50 [22]

Norm. Out. Tb Top-1 Top-2 Top-3 Tb Top-1 Top-2 Top-3

Tb

(%)

(s)

Pr Th. (s)

(%)

(%)

(%) (s)

(%)

(%)

(%)

Q0.80

1.4 0.23

44.1

33.6

26.1 0.17

0.8

0.1

0.0 0.56

Q0.90

1.3 0.23

57.0

52.1

47.2 0.17

5.7

0.3

0.0 0.56

Q0.95

1.3 0.26

73.5

63.6

58.4 0.17

53.8

15.8

1.5 0.56

Q0.80

1.0 0.20

46.7

37.1

32.1 0.17

79.8

63.4

50.5 0.40

Q0.90

1.1 0.23

70.5

55.8

49.1 0.17

91.1

83.4

75.6 0.41

Q0.95

1.1 0.23

84.1

70.6

65.3 0.17

95.1

91.1

86.6 0.40

and Q0.95 ) because SNR information was unavailable in the datasets. 2) Evaluation Results: The results for V I B E-MA, MNetLeNet, and ResNet-50 across both seen and unseen scenarios are shown in Table II. The first three rows report V I B E-MA results averaged across Scenarios 6 and 7, on which V I B EMA is not trained. In the bottom rows, we report performance on Scenario 9, which remains unseen for all models. When evaluated on the scenarios of baseline models, V I B E-MA consistently achieves low outage with an average beam alignment time below 0.26s. Unlike the baselines, V I B E-MA maintains stable performance across thresholds and scenarios. Despite not being trained on the same datasets, V I B E-MA maintains a very low outage of 1.3%-1.4%. On Scenario 7, V I B E-MA achieves 24.7pp to 57.1pp lower outage than MNet-LeNet as threshold increases from Q0.80 to Q0.95 . On Scenario 6, V I B EMA performs comparably to ResNet-50, which achieves 0% Top-3 outage at lower thresholds. However, at Q0.95 , ResNet50 outage increases to 1.5%, while V I B E-MA remains at 1.3%, with 143% faster beam alignment. Generalization. In the unseen Scenario 9, both baselines struggle to generalize: MNet-LeNet Top-3 outage exceeds 65%, and ResNet-50 leads to 85.6% outage. On the other hand, V I B E-MA sustains outage of only 1.1%, with up to 69.9pp lower outage and 79% lower latency than ResNet-50. These results show that V I B E-MA can generalize across environments and thresholds, while highlighting the limitations of black-box ML models trained on specific scenarios. In Fig. 11, we compare the coverage percentage [100 − (Outage %)] of V I B E-MA and V I B E-YOLOR with the TopK predictions from MNet-LeNet under varying thresholds. V I B E-MA consistently results in the highest coverage (98.6%-98.9%), in all thresholds and even under challenging conditions that are not within its training set. Compared to the strongest baseline, MNet-LeNet Top-3, V I B E-MA reduces outage by up to 69.3pp. The coverage gains of V I B E-MA are partly attributed to V I B E-YOLOR, which encapsulates the first three stages of V I B E-MA. V I B E-YOLOR outperforms



 () ()

 %







( ()

Fig. 11. Coverage percentage [100-(Outage %)] across thresholds in seen and unseen scenarios.

MNet–LeNet Top-1 by 9.5–13.7 pp and achieves coverage comparable to Top-2 using a single beam decision. Although Top-3 attains up to 11.5 pp lower outage than V I B E-YOLOR, it requires an additional beam selection step and still underperforms V I B E-MA. The generalization gap can be observed in Fig. 11 when MNet-LeNet outage is compared for seen and unseen scenarios, where V I B E-MA is virtually unaffected. Furthermore, the coverage of V I B E-MA is not affected as the threshold increases, where other methods suffer at higher quantiles. The results highlight the importance of closed-loop feedback and the generalization ability of V I B E-MA. Overall, V I B EMA achieves a strong balance between low outage and low latency in BS-centric settings, despite not being trained on any evaluation scenarios. This highlights the suitability of hybrid model-based, closed-loop architectures over end-to-end ML for real-world mmWave vehicular connectivity. C. Outdoor Evaluations 1) Experiment Setup: The outdoor experiment is conducted on the University of Nebraska-Lincoln campus, using a fixed BS and a mobile UE moving along an 80-m straight path, as shown in Fig. 12 (top). To emulate a worst-case V2I scenario, the UE boresight is oriented perpendicular to the road, causing rapid beam angle variations during motion. The UE detects urban streetlights—representing mmWave BS deployments in U.S. cities [10]—using a 60◦ NFOV Intel RealSense camera, and performs real-time beam selection and refinement. Experiments are conducted for V I B E-YOLOR, V I B E-MA, and V I B E-MLP at angular velocities of 1.6◦ /s, 8.0◦ /s, and 12.8◦ /s. 2) Evaluation Results: In Figs. 13, we show the CDFs of the margin from SNR threshold (11dB, 17dB, and 23dB). More specifically, we plot the CDF of γth − γ (i.e., F (γth − γ) = P (γth −γ ≤ x)), which essentially shows the probability that the achieved SNR is above or equal to the SNR threshold, while normalizing different SNR thresholds to x = 0 on the plot. When the vehicle is traveling at the angular velocity of 8.0◦ /s, V I B E-YOLOR meets the SNR threshold in only 33.9%, 30.0%, and 18.5% of cases for SNR thresholds of 11dB, 17dB, and 23dB, respectively, showing a steep decline as the SNR threshold increases. V I B E-MLP performs slightly better for the highest SNR threshold but worse for others as compared

SNR Th. 11dB

UE End

80m

Mobile UE

BS

Fig. 12. Outdoor Evaluation Testbed at University of Nebraska-Lincoln, USA.

to V I B E-YOLOR, achieving 29.4%, 27.8%, and 21.0%. In contrast, V I B E-MA achieves substantially higher real-time reliability, meeting the SNR threshold in 78.5%, 89%, and 73.7% of cases. This performance can further be enhanced through advanced techniques such as adaptive modulation and coding, offering the potential for full connectivity in mmWave V2X networks. Practical Considerations. Some performance metrics, including latency and coverage, are hardware-dependent and can be improved through further engineering, which is beyond the scope of this paper. Specifically, the end-to-end latency of V I B E is constrained by the inference speed of the vision pipeline and the beam switching rate of the phased-array hardware, which together bound the maximum vehicular speed the system can support. As a result, V I B E is evaluated at angular velocities corresponding to vehicle speeds of 1, 5, and 8 mph—representative of low-speed urban and parking scenarios. These are not fundamental limitations of the V I B E framework, but rather artifacts of the current prototype hardware. Scalability to higher speeds is achievable along two independent axes. First, extending camera range increases the distance at which a BS becomes visible, which reduces the rate of angular change experienced during approach and thereby relaxes latency requirements. In our current setup, a camera with a 3 m focal length enables BS detection up to 16 m; replacing this with existing ADAS-grade cameras that support ranges up to 100 m [44] would enable operation at speeds of approximately 50 mph. Second, inference latency can be independently reduced by deploying on more capable edge platforms: YOLOv11x latency drops from 75ms on a Jetson Orin Nano to 20 ms on an NVIDIA DRIVE AGX [45], directly translating to higher supportable speeds. Together, these improvements suggest a clear and practical path toward highwayspeed deployment using commercially available hardware. Summary. V I B E-MA, a hybrid model-based, closed-loop learning architecture, consistently outperforms competing methods. Indoors, V I B E-MA reduces outage by up to 53pp and 3.5pp over internal baselines. Against state-of-the-art baseline methods on unseen datasets, V I B E-MA outperforms MNet-LeNet by achieving 47.7pp lower outage. Compared to

CDF

BS UE Start



(a)YOLOR

SNR Th. 17dB (b)MLP

SNR Th. 23dB (c)MA

  



















Margin from SNR Threshold Fig. 13. Outdoor evaluations: CDF of margin from SNR threshold (x = 0): (a) V I B E-YOLOR, (b) V I B E-MLP, and (c) V I B E-MA (8.0◦ /s or 5mph).

ResNet-50, V I B E-MA has 69.9pp lower outage and is 73.9% faster. Finally, in outdoor, real-time trials, V I B E-MA lowers outage by up to 59pp compared to its internal baselines. These results establish V I B E-MA as a robust, scene-agnostic solution for a reliable beam alignment. VII. C ONCLUSIONS We present V I B E, which combines visual sensing with lightweight online correction to enable fast beam acquisition and reliable link maintenance without large-scale RF training data. Extensive indoor, outdoor, and cross-scenario evaluations show consistently low outage and strong robustness in unseen environments, where 5G NR hierarchical beamforming and black-box ML models struggle. V I B E provides a practical, low-cost alternative to radar, LiDAR, and GNSS using camera priming. Limitations and Future Work. We plan to further improve V I B E timing performance through hardware and software optimizations for high-velocity operation. The UE–BS coordination can be extended to arbitrary BS alignments using pose estimation. While this work focuses on line-of-sight scenarios, V I B E can be extended to non-line-of-sight settings using historical beam measurements and structural awareness. Camera impairments in high-mobility environments can be mitigated through multi-camera configurations or predictive tracking. Impact. Beyond performance gains, V I B E confines visual sensing to on-vehicle cameras, mitigating privacy concerns of BS-mounted sensors. Our results demonstrate that cameraguided adaptive beamforming is practical for resilient, realtime V2X connectivity. Broad adoption of double-directional links from channel access can further increase cell size, reduce BS density, and lower handover frequency. ACKNOWLEDGMENT This work was supported in part by the National Science Foundation (NSF) under Grants 2030141, 2030272, and 2112471. The authors would also like to thank E. Biswas and S. Shin for their assistance with conducting the outdoor experiments. R EFERENCES [1] Audi, “Audi of America, Verizon partner to bring 5G to vehicle lineup,” https://media.audiusa.com/releases/511, Feb. 2022.

[2] P. Lipscombe, “Verizon to build 5G test track in Germany with Audi,” https://www.datacenterdynamics.com/en/news/verizon-to-build-5g-tes t-track-in-germany-with-audi/, Mar. 2024. [3] Samsung, “mmWave 5G TCU is enabling new in-vehicle experiences,” https://www.samsung.com/global/business/networks/insights/press-relea se/0111-mmwave-5g-tcu-is-enabling-new-in-vehicle-experiences/, Jan. 2021. [4] NTT Corp., “NTT Corp., NTT DOCOMO and NEC demonstrate distributed MIMO technology for high-frequency 6G communications in automobiles and trains,” https://group.ntt/en/newsrelease/2025/03/25/25 0325a.html, Mar. 2025. [5] W. Chen et al., “5G-advanced toward 6G: Past, present, and future,” IEEE Journal on Selected Areas in Communications, vol. 41, no. 6, pp. 1592–1619, Mar 2023. [6] X. Lin, “A tale of two mobile generations: 5G-advanced and 6G in 3GPP release 20,” IEEE Communications Standards Magazine, pp. 1–9, Jun 2025. [7] J. Tan et al., “Beam alignment in mmWave V2X communications: A survey,” IEEE Communications Surveys & Tutorials, vol. 26, no. 3, pp. 1676–1709, Aug. 2024. [8] D. Lockie and D. Peck, “High-data-rate millimeter-wave radios,” IEEE Microwave Magazine, vol. 10, no. 5, pp. 75–83, Aug 2009. [9] V. Va, J. Choi, and R. W. Heath, “The impact of beamwidth on temporal channel variation in vehicular channels and its implications,” IEEE Trans. Vehicular Technology, vol. 66, no. 6, pp. 5014–5029, Jun. 2017. [10] Y. Feng et al., “Vivisecting beam management in operational 5G mmWave networks,” Proc. ACM CoNEXT, vol. 3, no. CoNEXT2, pp. 1–26, Jun 2025. [11] Remcom, “Wireless Insite 3D Wireless Prediction Software,” https://ww w.remcom.com/wireless-insite-propagation-software/, 2024. [12] C. K. Anjinappa and I. Guvenc, “Millimeter-wave V2X channels: Propagation statistics, beamforming, and blockage,” in IEEE VTC-Fall, Aug 2018. [13] O. Abari et al., “Millimeter wave communications: From point-to-point links to agile network connections,” in Proc. ACM HotNets Workshop, Nov 2016. [14] Q. Duan, T. Kim, and H. Ghauch, “KM learning for millimeter-wave beam alignment and tracking: Predictability and interpretability,” IEEE Access, vol. 9, pp. 117 204–117 216, Aug 2021. [15] A. Narayanan et al., “A comparative measurement study of commercial 5G mmWave deployments,” in Proc. IEEE INFOCOM, May 2022. [16] Y. Liu and C. Peng, “A close look at 5G in the wild: Unrealized potentials and implications,” in Proc. IEEE INFOCOM, May 2023. [17] Y. Shabara, C. E. Koksal, and E. Ekici, “Beam discovery using linear block codes for millimeter wave communication networks,” IEEE/ACM Trans. on Networking, vol. 27, no. 4, pp. 1446–1459, May 2018. [18] Q. Xue et al., “A survey of beam management for mmWave and THz communications towards 6G,” IEEE Communications Surveys & Tutorials, vol. 26, no. 3, pp. 1520–1559, Feb 2024. [19] B. Salehi et al., “Machine learning on camera images for fast mmWave beamforming,” in Proc. IEEE MASS, Dec 2020. [20] M. Alrabeiah, A. Hredzak, and A. Alkhateeb, “Millimeter wave base stations with cameras: Vision-aided beam and blockage prediction,” in Proc. IEEE VTC, Nov 2019. [21] G. Charan, M. Alrabeiah, and A. Alkhateeb, “Vision-aided 6G wireless communications: Blockage prediction and proactive handoff,” IEEE Trans. Vehicular Technology, vol. 70, no. 10, Oct 2021. [22] G. Charan et al., “Vision-position multi-modal beam prediction using real millimeter wave datasets,” in Proc. IEEE WCNC, Nov 2021, pp. 2727–2731. [23] A. Oliveira et al., “Machine learning-based mmWave MIMO beam tracking in V2I scenarios: Algorithms and datasets,” in Proc. IEEE LATINCOM, Dec 2024, pp. 1–5. [24] The Washington Post, “Police secretly monitored New Orleans with facial recognition cameras,” https://www.washingtonpost.com/busin ess/2025/05/19/live-facial-recognition-police-new-orleans/, May. 2025. [25] Markets and Markets, “Advanced Driver Assistance Market Size, Share, & Analysis,” https://www.marketsandmarkets.com/Market-Reports/dri ver-assistance-systems-market-1201.html/, May 2025. [26] A. Zhou, X. Zhang, and H. Ma, “Beam-forecast: Facilitating mobile 60 GHz networks via model-driven beam steering,” in Proc. IEEE INFOCOM, May 2017, pp. 1–9.

[27] S. Jiang and A. Alkhateeb, “Computer vision aided beam tracking in a real-world millimeter wave deployment,” in Proc. IEEE GLOBECOMM Workshops, Dec 2022. [28] G. Charan et al., “Camera based mmWave beam prediction: Towards multi-candidate real-world scenarios,” IEEE Transactions on Vehicular Technology, vol. 74, no. 4, pp. 5897–5913, Dec 2024. [29] S. Imran, G. Charan, and A. Alkhateeb, “Environment semantic aided communication: A real world demonstration for beam prediction,” in Proc. IEEE ICC Workshops, Jun 2023. [30] M. Arnold et al., “Vision-assisted digital twin creation for mmwave beam management,” in IEEE International Conference on Communications, Jun 2024. [31] B. Lin et al., “Multi-camera views based beam searching and BS selection with reduced training overhead,” IEEE Trans. Communications, vol. 72, no. 5, pp. 2793–2805, Jan 2024. [32] T. Zhang, J. Liu, and F. Gao, “Vision aided beam tracking and frequency handoff for mmWave communications,” in Proc. IEEE INFOCOM Workshops, Jul 2022. [33] B. Salehi et al., “Omni-CNN: A modality-agnostic neural network for mmwave beam selection,” IEEE Trans. Vehicular Technology, vol. 73, no. 6, pp. 8169–8183, Jan 2024. [34] B. Salehi et al., “FLASH-and-prune: Federated learning for automated selection of high-band mmwave sectors using model pruning,” IEEE Trans. Mobile Computing, vol. 23, no. 12, pp. 11 655–11 669, May 2024. [35] A. Alkhateeb et al., “DeepSense 6G: A large-scale real-world multimodal sensing and communication dataset,” IEEE Communications Magazine., vol. 61, no. 9, pp. 122–128, Sep. 2023. [36] U. Demirhan and A. Alkhateeb, “Radar aided 6G beam prediction: Deep learning algorithms and real-world demonstration,” in Proc. IEEE WCNC, Apr 2022. [37] S. Jiang, G. Charan, and A. Alkhateeb, “LiDAR aided future beam prediction in real-world millimeter wave V2I communications,” IEEE Wireless Communication Letters, vol. 12, no. 2, pp. 212–216, May 2022. [38] M. M. G. Reports, “AUTOMOTIVE CAMERA MARKET OVERVIEW,” https://www.marketgrowthreports.com/market-reports/a utomotive-camera-market-100220#:∼:text=AUTOMOTIVE%20CAME RA%20MARKET%20TRENDS,-the-art%20imaging%20technologies/, December 2025. [39] A. Alkhateeb et al., “Channel estimation and hybrid precoding for millimeter wave cellular systems,” IEEE Journal of Selected Topics in Signal Processing, vol. 8, no. 5, pp. 831–846, Jul 2014. [40] J. Nilsson, J. Fredriksson, and A. C. Ödblom, “Reliable vehicle pose estimation using vision and a single-track model,” IEEE Trans. on Intelligent Transportation Systems, vol. 15, no. 6, pp. 2630–2643, May 2014. [41] P. Sturm, “Pinhole camera model,” in Computer Vision: A Reference Guide, R. Kimmel, M. M. Bronstein, and P. Favaro, Eds. Springer, Apr. 2021, pp. 983–986. [42] M. Hussain and R. Khanam, “YOLOv11: An overview of the key architectural enhancements,” 2024. [Online]. Available: https: //arxiv.org/abs/2410.17725 [43] A. Shkel, A. Mehrabani, and J. Kusuma, “A Configurable 60GHz phased array platform for multi-link mmWave channel characterization,” in IEEE ICC Workshops, Jun 2021. [44] Otobrite, “oToGuard level 2+ all-in-one ADAS,” https://www.otobrite.c om/product/otoguard#:∼:text=Level%200∼2+%20ADAS%20functions ,LCA%2C%20LKA%2C%20and%20more., 2024. [45] NVIDIA-Technical-Blog, “How drive agx, cuda and tensorrt achieve fast, accurate autonomous vehicle perception,” https://developer.nvidia.com/blog/how-drive-agx-cuda-and-tensorrtachieve-fast-accurate-autonomous-vehicle-perception/, Oct 2019.

Record · ID 158553 · SHA-256 70b61408f4537f32
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.