ConceptioArchivearXiv CS
arXiv CSopen access

Adaptive Learning Strategies for AoA-Based Outdoor Localization: A Comprehensive Framework

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Received XX Month, XXXX; revised XX Month, XXXX; accepted XX Month, XXXX; Date of publication XX Month, XXXX; date of current version XX Month, XXXX. Digital Object Identifier 10.1109/TMLCN.2022.1234567

arXiv:2605.05055v1 [cs.LG] 6 May 2026

Adaptive Learning Strategies for AoA-Based Outdoor Localization: A Comprehensive Framework Bac Trinh-Nguyen12 , Graduate Student Member, IEEE, Sara Berri1 , Member, IEEE, Sin G. Teo2 , Tram Truong-Huu3 , Senior Member, IEEE and Arsenia Chorti14 , Senior Member, IEEE 2

1 ETIS UMR 8051, CY Cergy Paris University, ENSEA, CNRS, FR Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR), SG 3 Singapore Institute of Technology (SIT), SG 4 Barkhausen Institut gGmbH, Dresden, DE

Corresponding author: Bac Trinh-Nguyen (email: [email protected]). B. Trinh-Nguyen has been partially supported by the EC through the Horizon Europe/JU SNS project ROBUST-6G (Grant Agreement no. 101139068), the Horizon Europe COST Action Project 6G-PHYSEC, the CNRS IPAL Project CONNECTING. S. Berri have been partially supported by the EC through the Horizon Europe/JU SNS project ROBUST-6G (Grant Agreement no. 101139068), the EU HORIZON MSCA-SE TRACE-V2X project (Grant No. 101131204), the ANR-PEPR 5G Future Networks projects. A. Chorti has been partially supported by the EC through the Horizon Europe/JU SNS project ROBUST-6G (Grant Agreement no. 101139068), the Horizon Europe COST Action Project 6G-PHYSEC, the CNRS IPAL Project CONNECTING, the ENSEA SRV project RETRO, the CYU TalCyb Chair in Cybersecurity and by the French government under the France 2030 ANR program “PEPR Networks of the Future” (ref. ANR-22-PEFT-0008 HISEC and ANR-22-PEFT-0009 FOUNDS). The authors would like to thank S. Wesemann, G. Kaltbeitzel, D. Wiegner, M. Kinzler, S. Merk and S. Woerner from Nokia for realizing the channel measurements and sharing the data.

ABSTRACT

Localization in 5G / B5G and 6G networks is essential for important use cases such as intelligent transportation, smart factories, and smart cities. Although deep learning has enabled improving localization accuracy, depending on the deployment scenario and the effort required for dataset collection campaigns on a given infrastructure, the training process for localization models can vary significantly. Furthermore, with respect to feature selection, recent works have demonstrated the robustness of angle-of-arrival (AoA)-based localization. In view of these two points, we propose an adaptive framework for AoA-based localization that consists of two alternative learning strategies, each suited either for “large” or “small” training datasets. The proposed framework is evaluated on a real, massive multiple-input multiple-output (mMIMO) orthogonal frequency division multiplexing (OFDM) outdoor channel state information (CSI) dataset. First, we investigate offline learning when large training datasets are available; we propose a hierarchical framework that first distinguishes between line of sight (LoS) and non line of sight (NLoS) regions and then moves to more fine grained localization in the respective region. This approach provides high-performance localization through accumulated batch retraining and an integrated hyperparameter optimization mechanism, achieving 100% accuracy in distinguishing LoS and NLoS regions, 99.82% accuracy for predefined trajectories classification in the LoS region and approximately 98% accuracy for those in the NLoS region. Second, when only a small training dataset is available, an online learning framework is proposed, using incremental tree-based and ensemble-based models for handling streaming data and continuously updating mode, as well as an online few-shot learning model for rapidly initializing new classes from a limited labeled support set. The aggregated Mondrian Forest (AMF) achieves an accuracy of approximately 94% for the classification of trajectories in both LoS and NLoS regions, with very low forgetting rates ranging from 0.0248 to 0.0427. These results showcase that robust localization in outdoor wireless environments is achievable with low-latency, and, demonstrate that high accuracies can be achieved incrementaly during network operation by exploiting online learning, alleviating the need for large dataset collection campaigns. INDEX TERMS 6G, wireless localization, channel state information, angle of arrival, machine

learning, continual learning, online learning, few-shot learning, generative models. VOLUME ,

1

I. INTRODUCTION

L

OCALIZATION in 5G / B5G and 6G networks is an essential capability due to its critical role in various applications such as smart industry, autonomous agents, and intelligent transportation systems. In modern multiple input multiple output (MIMO) systems, channel state information (CSI) is measured between transmitting and receiving devices and can be used for localization tasks [1], [2]. Recent works have discussed the robustness of CSI features against noise and adversarial attacks and the potentially important role of robust localization in enhancing security and trustworthiness [3], [4] of wireless systems. Among different trusted features for localization, earlier works relied on time of arrival (ToA) and time difference of arrival (TDoA). Recently, the works in [5]–[8] have shown that angular-based positioning techniques such as angle-ofarrival (AoA) and angle-of-departure (AoD) not only provide a scalable and cost-effective approach with significant potential for future applications [9], but also, importantly, cannot be forged in digital array MIMO systems and can serve as robust physical layer features for localization and location-based authentication. However in real-world wireless systems, the fluctuations of estimated AoA values due to noise, multipath propagation and interference are generally non negligible. Although this complicates their direct use for localization, AoA estimates still provide discriminative patterns that can be exploited as powerful features in machine learning (ML)-based localization methods. Combining AoA features with supervised ML techniques has been shown to achieve strong localization performance in both indoor and outdoor scenarios [10], [11], reconfirming that ML and deep learning (DL) have facilitated enhanced localization accuracy. In the following, we present an overview of several localization learning approaches, including offline learning, online learning, and few-shot learning, highlighting current literature gaps and how our proposed framework aims at closing them. We then also briefly review our early results on the proposed AoA-based localization offline-learning framework, followed by the presentation of our main contributions on a comprehensive framework of apative learning for AoA-based outdoor localization. A. OFFLINE LEARNING-BASED LOCALIZATION

In recent years, ML techniques have been widely adopted for radio frequency (RF)-based localization to overcome the limitations of traditional approaches, including sensitivity to environmental dynamics, hardware imperfections, and synchronization constraints [12]. For example, a method based on k -nearest neighbors (KNN) was shown in [13] to achieve high localization accuracy using CSI in indoor environments. The work of Yang et al. [14] further provided a comprehensive survey of

ML-driven indoor localization, highlighting the use of supervised learning techniques such as support vector machines (SVMs), KNN and neural networks within offline fingerprinting frameworks. Beyond indoor scenarios, MobIntel [15] demonstrated the effectiveness of ML-based methods for passive outdoor localization using received signal strength indicator (RSSI) measurements. In addition, multipathassisted CSI fingerprinting approaches have been proposed in [16], achieving meter-level accuracy in outdoor localization settings. These works validated the strong potential of ML-based techniques in diverse RF localization environments. However, these approaches do not distinguish between LoS and NLoS propagation environments, which have well-known and distinct impacts on achievable localization accuracy due to their differing multipath structures, instead relying on blackbox learning to predict location. B. ONLINE LEARNING-BASED LOCALIZATION

Online learning, also known as continual or incremental learning, aims to update models sequentially with incoming data streams while mitigating catastrophic forgetting and avoiding complete retraining from scratch. As a result, it can reduce computational cost and resource consumption compared to batch retraining approaches. Due to mobility, fading, and environmental changes in wireless and CSI-based systems, continual learning has been applied to some tasks, such as channel estimation and action recognition [17]–[20]. Elastic weight consolidation (EWC) [21] is one of regularization-based methods that constrains updates on parameters identified as important for previous tasks. Memory-aware synapses (MAS) [22] and replay-based methods such as iCaRL [23] retain a subset of past samples and add them during the training process on new tasks to mitigate forgetting. These methods typically rely on backpropagation, stored samples, and epochbased retraining on historical data, which may limit their applicability in streaming and resource-constrained wireless scenarios. In contrast, online and streaming learning focuses on single-pass, incremental model updates when samples arrive sequentially, while storage and retraining are treated as optional supporting mechanisms rather than core requirements. Within this paradigm, tree-based incremental models such as the hoeffding tree (HT) [24] leverage statistical guarantees to grow decision trees from streaming data, while adaptive variants like the hoeffding adaptive tree (HAT) [25] incorporate drift detection to handle non-stationary environments. Ensemble methods improve robustness and accuracy by combining multiple incremental models, such as adaptive random forests (ARF) [26], streaming random patches (SRP) [27], and aggregated mondrian forests (AMF) [28],

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ 2

VOLUME ,

which exploit model diversity and drift awareness. These ensemble and streaming tree-based approaches are well suited to wireless localization scenarios where CSI distributions evolve over time in dynamic environments. In this work, we utilize incremental models implemented through the River library [29] to enable continual adaptation under streaming data conditions in an AoAbased online learning framework. We rely on these builtin configurations without introducing an external drift detection module. However, the use of these online learning approaches in AoA-based localization is still limited, particularly in terms of balancing adaptation capability, computational efficiency, and reliability of localization decisions under varying propagation conditions.

emerged as effective solutions for synthesizing additional training data, with generative adversarial networks (GANs) [37] and variational autoencoders (VAEs) [38] being among the most widely used approaches to generating realistic synthetic samples. To further enhance class-conditional generation, conditional variants, such as conditional GANs (CGANs) [39] and conditional VAEs (CVAEs) [40] incorporate label information during the generation process. These generative models have been applied to enrich the CSI datasets and improve generalization under limited data conditions [41], [42]. Previous studies have shown that generative data augmentation can enhance model robustness to environmental variations and measurement noise, particularly in offline training scenarios.

C. FEW-SHOT LEARNING

Few-shot learning aims to enable models to perform new tasks using only a small number of labeled examples. Meta-learning, or learning to learn, provides an effective solution for few-shot learning by training models across multiple tasks to acquire transferable knowledge that allows rapid adaptation to unseen tasks with limited data [30]. Existing meta-learning approaches for few-shot learning can be categorized into metric-based methods such as prototypical networks (ProtoNet) [31], memorybased methods, and learning-based methods such as model-agnostic meta-learning (MAML) [32]. In the context of wireless sensing and localization, several recent works have explored meta-learning techniques for fast adaptation across environments. Fewshot learning was investigated for Wi-Fi-based indoor positioning using CNN and meta-learning techniques, demonstrating that meta-learning can achieve competitive performance under limited data conditions [33]. MetaLoc [34] introduced the MAML-based fingerprinting localization framework that learned meta-parameters from multiple well-calibrated domains. Cui et al. [35] proposed ProFi-Net, a prototype-based few-shot Wi-Fi gesture recognition method that extended Prototypical Networks with feature-level attention and gradually increased noise-based query augmentation to improve robustness and generalization. Few-shot learning for AoA estimation using prototypical networks [36] demonstrated strong performance under limited data and domain shifts, highlighting the potential of few-shot approaches for adaptive localization systems. Nevertheless, although previous works have explored the use of few-shot learning techniques, such as ProtoNet, in wireless communication tasks, their application to angular-based localization remains limited. D. DATA AUGMENTATION

Due to the high cost of collecting labeled wireless data, data scarcity remains a major challenge in training robust localization models. Generative models have VOLUME ,

E. Proposed Adaptive Localization Framework

Although prior studies have explored offline learning, online learning, and few-shot learning for wireless localization, important gaps still remain. Existing offline approaches typically assume relatively fixed environments and sufficient labeled training data, which may limit their applicability in dynamic deployment conditions. Online learning approaches offer the ability to incrementally update models as new data arrive, but their potential for AoA-based localization has not been sufficiently investigated. Similarly, few-shot learning has shown promise for rapid adaptation with limited labeled samples, yet it is often studied in isolation rather than as part of a broader adaptive localization strategy. Moreover, previous studies considered offline learning, continual learning, few-shot learning as isolated solutions. Therefore, there is still a need for an adaptive localization framework that can support both offline and online learning paradigms under different scenarios while leveraging AoA-based representations for practical realworld deployment. In our previous work [43], we extracted AoA features from CSI measurements using multiple signal classification (MUSIC) [44] and estimation of signal parameters using rotational invariance techniques (ESPRIT) [45]. We employed a hierarchical classification strategy that first discriminates between LoS and NLoS regions, followed by two region-specific classifiers to identify predefined trajectories within each region. In [43], we systematically assessed six baseline machine learning models: k -nearest neighbors (KNN), random forest (RF), logistic regression (LR), gradient boosting machine (GBM), extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), as well as the stacking ensemble model [46]. Although the baseline localization results were already strong, in the current work we achieve further performance gains through hyperparameter optimization of classifiers, where accuracy is used as the optimization objective. 3

Bac TN et al.: Adaptive Learning Strategies for AoA-Based Outdoor Localization: A Comprehensive Framework

Despite the near perfect performance, offline supervised learning requires a sufficiently large and representative labeled dataset and typically requires full retraining when the channel characteristics, user trajectories, or LoS / NLoS conditions vary over time, non-stationary data distributions. The models trained in an offline manner are unable to fully capture all these variations and therefore require periodic retraining or when performance monitoring modules detect degradation below a predefined threshold. During this process, all collected data must be stored and labeled prior to retraining, which increases system latency and resource consumption. Even more importantly, offline learning relies on the availability of large labeled datasets, which might be impractical in many radio access network (RAN) deployment scenarios and limits its adaptability to dynamic environments and newly observed classes. To address these shortcoming, we investigate two online learning approaches, including continual tree- and ensemble treebased methods, as well as continual few-shot learning, to address the buring issue of limited data availability. We adopt a metric-based few-shot learning strategy inspired by ProtoNet for scenarios in which only a small number of labeled samples are available at a time. The proposed approach enables rapid initialization of newly observed trajectories or regions from limited labeled data, which is important for practical wireless localization in limitedlabel streaming environments. While prior studies typically treat offline learning, continual learning, and few-shot learning as isolated approaches, this work proposes a unified framework for AoA-based localization that systematically leverages these paradigms across different operational phases of the network. The proposed framework aligns naturally with the deployment lifecycle of future 6G RAN systems. In the initial phase, where labeled data are scarce, online learning combined with few-shot techniques enables rapid model initialization and early adaptation. As additional data become available at the gNodeBs and are aggregated through cloud or virtual RAN pooling, continual learning facilitates progressive model refinement under streaming conditions. In later stages, once sufficient labeled data have been collected, offline training can be used to derive a more accurate and stable model. Overall, this results in a scalable, adaptive, and robust localization solution suited for dynamic wireless environments. The contributions of this paper are summarized as follows: • We design an offline AoA-based localization framework based on high-accuracy hierarchical classifiers. In the first stage, a binary classifier distinguishes LoS from NLoS regions with 100% accuracy. In the second stage, region-specific classifiers achieve 4

99.82% accuracy in the LoS region and approximately 98% accuracy in the NLoS region after hyperparameter tuning. • Next, we propose an online learning framework that leverages incremental tree-based and ensemblebased classifiers to process streaming data, along with continual few-shot learning to quickly initialize new classes with limited labeled samples, thereby enabling the adaptive operation under dynamic conditions. The results demonstrate that AMF is the most effective online learner for AoA-based localization, achieving about 94% accuracy in both LoS and NLoS regions, with very low forgetting rates between 0.0248 and 0.0427. • In addition, we propose a conditional variational autoencoder (CVAE) for data augmentation in the offline setting to simulate potential environmental variations and evaluate the robustness of our proposed methods.

The remainder of the paper is organized as follows. Section II presents our proposed adaptive localization framework. Section III presents the experimental setup and analysis of experimental results. Finally, Section IV concludes the paper and discusses future directions. II. PROPOSED AOA-BASED LOCALIZATION FRAMEWORK

To enable robust localization in wireless communications, we propose an adaptive AoA-based localization framework that integrates both offline and online learning strategies depending on the deployment scenario while meeting the requirements for low-latency and realtime applications. Fig. 1 illustrates the proposed adaptive localization framework. In the offline learning block, the system is initialized through offline training using AoA features extracted from CSI measurements. When deployed, trained models provide a strong performance baseline with high reliability. To ensure stable deployment, we introduce a parallel model evaluation and replacement mechanism along with the hyperparameter optimization module. The deployed model remains unaffected and operates in inference mode. New AoA features extracted from CSI measurements and labeled samples are accumulated and used to train candidate models. Candidate models that demonstrate better performance than the current model are promoted to deployment. This design ensures reliable localization over time. In scenarios where latency and resource are critical, the online learning approach operates using continual learning and few-shot learning. This is particularly important when only limited data is available at deployment time, such as when a new base station is installed, where large-scale data collection is impractical. The adaptive models can be pre-trained with a small VOLUME ,

Offline Learning Framework CSI Measurements

AoA Feature Extractor

2 Trigger retraining and model selection

Deployed Model (inference + update) Feature Repository

Online Learning Framework CSI Measurements

AoA Feature Extractor

Streaming data

Model block

Main flow

Function block

Model management

Storage block

Generative replay

Deployed Model (inference + update)

Candidate Models Candidate Models

1

Optimize hyperparameters

1

Update model

Candidate Models Candidate Models

Data Augmentation

Evaluation & Monitoring

Evaluation & Monitoring

2

Store in buffer

3

Trigger reinitialization / rehearsal

Replay Buffer

FIGURE 1: Overview of AoA-based localization framework. Offline learning performs batch training and retraining with model selection, while online learning handles streaming data and updates models incrementally.

amount of labeled data to facilitate smooth model initialization, enable gentle parameter adjustment, and help the model generalize faster and more accurately during the inference phase. During the inference phase, the model parameters are updated incrementally as new data samples arrive. To mitigate performance degradation caused by concept drift or catastrophic forgetting, the scheme incorporates a small replay buffer and a monitoring mechanism. The online learning workflow is as follows: 1) The model is updated incrementally as new data arrive. 2) Newly acquired data are stored in a replay buffer. 3) If data drift is detected or the model exhibits catastrophic forgetting, the evaluation module triggers a rehearsal or selective model reinitialization. Since we investigate two approaches within the online learning framework, the details of continual learning and the few-shot learning approach are described in Section II Part B. A. OFFLINE LEARNING LOCALIZATION FRAMEWORK

The hierarchical two-stage classification framework is illustrated in Fig. 2. It consists of two main stages. In the first stage, a binary classifier discriminates between LoS and NLoS regions. In the second stage, two regionspecific multiclass classifiers distinguish fine-grained predefined trajectories (tracks) within each region. These classifiers are trained and evaluated prior to deployment VOLUME ,

in two main phases: the training phase and the inference phase. Training phase. The training phase aims to develop and select models for deployment and operation. CSI measurements are preprocessed and labeled by regions and track identifier (track ID). Next, MUSIC and ESPRIT algorithms are applied to estimate AoAs and construct feature vectors, which are then fed into the classifier block to train a hierarchical classifier including a binary LoS / NLoS classifier, followed by two regionspecific multi-class classifiers: one for LoS tracks and one for NLoS tracks. During training, we evaluate multiple ML algorithms, including LR, KNN, RF, GBM, XGBoost, LightGBM, and a stacking ensemble model (combining the top n best performing ML models, which n ∈ 2, ..., 6. Finally, the best-performing and most robust models are selected for deployment. Inference phase. Incoming CSI samples are processed in real-time to extract AoA features. This will first be classified by the LoS / NLoS classifier to determine to which region the incoming data belong to. The corresponding region-specific track classifier is then used to identify the exact track. If a sample is classified as LoS, the LoS track classifier predicts the track ID. Otherwise, it is forwarded to the NLoS track classifier to predict the track ID. Hyperparameter Optimization. To improve classification performance without altering the overall localization framework, a monitoring and evaluation module is added. This module continuously assesses the performance of the deployed model and triggers ac-

5

Bac TN et al.: Adaptive Learning Strategies for AoA-Based Outdoor Localization: A Comprehensive Framework

The final goal is to identify well-performing configurations that can be used consistently in the deployment phase. For each classifier, we define an objective function that samples a hyperparameter configuration from a predefined search space as described in Table 11 and evaluates generalization performance using 5-fold crossvalidation on the training set. The optimization process aims to maximize the mean cross-validation accuracy across folds. B. ONLINE LEARNING LOCALIZATION FRAMEWORK

In this subsection, we present a data augmentation mechanism in Part 1, which employs generative models to support model training and evaluation under data scarcity conditions. We introduce our online learning framework through several approaches. First, we describe a baseline online learning strategy based on periodic batch retraining with conventional ML models. Next, we present incremental learning in streaming settings using tree-based classifiers and ensemble models (Part 2). Finally, we investigate few-shot learning with prototypical networks as a representative approach to enable rapid adaptation when only a small number of labeled samples are available for new tracks (Part 3).

1) Data Augmentation using Generative Models

FIGURE 2: Architecture of hierarchical two-stage classifiers (Offline Learning Framework).

cumulative batch retraining of candidate models when performance degradation is detected or when sufficient new data have been collected. During this process, model hyperparameters are optimized using the accumulated dataset. A trigger mechanism is used to promote the deployment of a candidate model when the performance of the deployed model degrades and the candidate model achieves superior performance. This process serves as a backup solution in the event of an unexpected model failure. There are several methods for hyperparameter tuning in ML, including manual tuning where hyperparameters are selected heuristically, exhaustive search strategies such as grid search, random search, and Bayesian optimization [47]. In this work, we investigate Bayesian optimization and perform automated hyperparameter tuning using Optuna [48] for the candidate classifiers introduced in Section A. Instead of brute-force combinations of values, Bayesian optimization formulates the tuning process as a sequential optimization problem. It leverages information from previous evaluations to guide the search toward promising regions of the hyperparameter space and determines subsequent trials accordingly. 6

Even with strong fine-tuned classifiers, the current offline approach may still degrade in real-world deployments due to environmental dynamics, delayed or unavailable labels, and the computational cost of retraining models from scratch to adapt with distribution shifts. To address these limitations, we investigate a data augmentation module based on two conditional generative models for feature-level augmentation: conditional generative adversarial network (CGAN) and conditional variational autoencoder (CVAE). This module generates synthetic samples from the original data with two main objectives: • To mimic realistic variations in the AoA feature distribution supporting model evaluation and improving robustness under previously unseen conditions, • To support upsampling of replay buffer during rehearsal or reinitialization processes.

Their architectures are presented in Fig. 3. In both cases, generation is conditioned on a label vector c that encodes the scenario information, including the propagation region (LoS / NLoS) and the corresponding predefined track IDs. Table 12 summarizes the architectures of these two generative models evaluated for feature-level augmentation. Both models are implemented as multi-layer perceptrons (MLPs). Specifically, CGAN consists of an MLP-based generator and an MLP-based discriminator, while CVAE consists of an MLP-based encoder and an MLP-based decoder. VOLUME ,

and the decoder. Given an AoA feature vector x and its conditional label c, the encoder qϕ (z|x, c) maps the concatenated input [x, c] to the parameters of a Gaussian latent distribution µϕ (x, c) and log σϕ2 (x, c). A latent vector is then obtained via the reparameterization trick [51]: z = µ + σ ⊙ ϵ,

ϵ ∼ N (0, I),

(4)

where ⊙ denotes element-wise multiplication. The decoder pθ (x | z, c) takes the latent vector z and conditional label c to reconstruct x̃. Training minimizes the negative evidence lower bound (ELBO) which consists of a reconstruction term, implemented as the mean squared error (MSE) and a Kullback-Leibler (KL) divergence regularization term: LVAE = Lrecon + LKL   = Eqϕ (z|x,c) ∥x − x̃∥22 +

FIGURE 3: Architectures of the Conditional Generative Adversarial Network (CGAN) and the Conditional Variational Autoencoder (CVAE) used for feature-level AoA data augmentation. Both models operate on AoA feature vectors conditioned on class labels to generate / reconstruct synthetic AoA samples.

Conditional Generative Adversarial Network (CGAN). The conditional GAN comprises a generator G and a discriminator / critic D. Let x denote a real sample drawn from the data distribution, conditioned on the label vector c (encoding LoS / NLoS and track ID), the generator synthesizes AoA feature vectors x̃ = G(z, c) from a noise input z , aiming to generate class-consistent samples that resemble real AoA features. The critic D(·, c) assigns higher scores to real samples than to generated ones. Unlike standard GANs, the Wasserstein GAN (WGAN) [49] replaces the cross-entropy with the Wasserstein distance, providing smoother gradients and more stable training. In addition, we apply a gradient penalty (WGAN-GP) [50] to enforce the 1-Lipschitz constraint. The generator G is trained to minimize loss LG = −E[D(x̃, c)],

(1)

and the critic is trained by minimizing LD = −E[D(x, c)] + E[D(x̃, c)] + λgp GP,

(2)

where λgp controls the contribution of the gradient penalty (GP) term i h 2 GP = Ex̂ (||∇x̂ D(x̂, c)||2 − 1) (3) where x̂ denotes an interpolated sample between a real sample x and a generated sample x̃, used to evaluate the gradient penalty. Conditional Variational Autoencoder (CVAE). CVAE is structured by two main components, the encoder VOLUME ,

(5)

β DKL (qϕ (z | x, c) ∥ p(z)) ,

where DKL (·|·) is the Kullback–Leibler divergence and β controls the strength of the KL regularizer. The CVAE and CGAN models are evaluated using a realistic outdoor dataset, with detailed experimental results reported in Section III Part 1. As illustrated in Fig. 1, the trained CVAE-based data augmentation module is connected to the replay buffer in the practical implementation. Based on these results, CVAE produces more stable training behavior and high quality synthetic AoA features in our setting. Therefore, we apply CVAE as the primary data augmentation method for the remainder of the online learning strategy. This design enables data upsampling even when only a small set of samples is available, thus facilitating quick model reinitialization and supporting evaluation after rehearsal using the replay buffer.

2) Continual Learning

In real-world deployments, CSI / AoA data arrive sequentially as a data stream. Labels may be missing, delayed, or even only available sparsely, while environmental dynamics can induce distribution shifts over time. An ideal localization framework should maintain stable performance under non-stationary conditions while minimizing the computational cost of repeated full retraining. Batch Retraining with Conventional ML. We extend the conventional ML pipeline introduced in Section II Part A to a streaming-like setting by periodically retraining models on accumulated data batches (e.g., every day or every week). We consider two batch retraining approaches: • Buffer retraining: model is retrained using only the most recent data window (current buffer). • Cumulative retraining: model is retrained using all samples observed so far. 7

Bac TN et al.: Adaptive Learning Strategies for AoA-Based Outdoor Localization: A Comprehensive Framework

Although the cumulative training approach provides strong performance, it has clear practical limitations: • Batch requirement and latency: the system must wait until enough samples are collected before retraining, leading to delay adaptation. • Lack of streaming updates: conventional offline learners are not designed for incremental updates and typically require full retraining. • Computational and storage cost: since data needs to be stored, this increases retraining time and memory requirements.

These limitations motivate the need for a continual / incremental method that supports online optimization of model parameters without full retraining, adapts to evolving data distributions, and reduces computation and storage. In the following, we evaluate a set of incremental models including tree-based incremental models and ensemble methods. Continual / Incremental Learning. To support deployment in dynamic environments, we adopt a continual learning approach. In contrast to batch training, online learning updates the model incrementally as new samples arrive, enabling adaptation to potential distribution shifts due to environmental changes and user’s mobility. These observations motivate us to move beyond batch retraining toward a learning paradigm that can continuously update as data arrive. Therefore, we formulate AoA-based localization as a data-stream problem and employ continual classifiers that perform lightweight incremental updates. In this work, we evaluate several continual learning models, including: • Hoeffding Tree Classifier (HT): Incremental decision tree for data streams that grows using the Hoeffding bound to decide splits from a limited number of observations, as defined: r R2 ln(1/δ) ϵ= , (6) 2n where n denotes the number of samples observed at a node, R is the range of the chosen split evaluation metric, and δ is the confidence parameter that limits the probability of selecting a suboptimal split. • Hoeffding Adaptive Tree Classifier (HAT): Driftadaptive extension of Hoeffding trees that monitors predictive performance of branches using a drift detector and can replace subtrees when degradation is detected. • Adaptive Random Forest Classifier (ARF): Streaming ensemble of incremental trees designed for nonstationary streams and uses per-tree drift detection to selectively reset or replace underperforming trees. • Streaming Random Patches Classifier (SRP): Online ensemble method that generalizes bagging / random-subspaces for streams. The default base 8

learner is a Hoeffding Tree, though other incremental estimators can be used. • Aggregated Mondrian Forest (AMF): Online forest based on Mondrian-tree partitions, supporting single-pass learning and anytime prediction. Each node maintains a regularized label distribution, and predictions can be aggregated along the path to the leaf. The forest outputs the average of class probabilities across trees. • Gaussian Naive Bayes (GNB): Lightweight incremental probabilistic baseline that maintains a Gaussian per class and feature. Predictions are computed via the summed per-feature log-likelihoods under the Naive Bayes independence assumption, making it fast and memory-efficient. Streaming flow and online update procedure. The continual learning framework operates under a streaming setting in real-time deployment, including an initial warm-up phase and a subsequent online inference and update stage. • Streaming flow: we consider that a data stream xt denotes the AoA feature vector extracted at time t and yt is the associated label. • Warm-up training (model initialization): we initialize each classifier using the first 10% of the dataset that assumes a realistic setting in which only a small labeled dataset ready at the beginning. • Online phase (predict-then-train streaming task): For the 90% of dataset, we will implement the predict-then-train strategy for simulating online learning, each data stream xt will arrive and the classifier tries to predict and learn from this new data:

1) Predict: ŷt = ft−1 (xt ) using the current model ft−1 . 2) Evaluate: update the streaming metrics using (ŷt and yt ). 3) Learn: update the model with the new labeled sample ft ← − update(ft−1 , xt , yt ). This phase mirrors online deployment, where decisions must be made immediately, and the model improves as labeled feedback becomes available. 3) Few-shot Learning

While continual learning supports efficient adaptation to streaming data when the set of trajectories (tracks / classes) is fixed and labels are available regularly, it does not fully address settings in which samples from previously unseen trajectories appear over time. In such cases, continual classifiers may require substantial labeled data to reliably update their decision boundaries, resulting in slow adaptation and potentially unstable updates. To address this gap, we therefore investigate an online VOLUME ,

few-shot meta-learning strategy based on prototypical networks (ProtoNet). ProtoNet is a metric-based method that can form a new class from K labeled samples by constructing class prototypes in the learned embedding space and supports rapid initialization of unseen classes. We first train and evaluate ProtoNet under the standard episodic few-shot meta-learning setup, and then extend it to an online setting by updating class prototypes incrementally as new labeled samples arrive. Support Set and Query Set. In meta-learning, the dataset D is split into two sets: meta-train (Dmeta_train ) and meta-test (Dmeta_test ), corresponding to train and test sets in supervised ML. During the training and testing processes, the meta-train and the meta-test are again divided into train and test sets. Thus, to avoid confusion with the conventional train / test terminology used within each episode (task), we use the terms support set and query set. Finally, the datasets are partitioned as follows: support query Dmeta_train = {Dmeta _train , Dmeta_train },

(7)

support query Dmeta_test = {Dmeta _test , Dmeta_test }.

(8)

where Dmeta_train is used to train the meta-learner model and this model is evaluated on episodes sampled from Dmeta_test . N-way K-shot problem. The training process in metalearning proceeds over episodes (tasks). In each episode T , a support set S and a query set Q are constructed. The support set follows the N -way K -shot setting, which means that it consists of N classes with K labeled samples per class. The query set contains M additional samples for each of the same N classes. Formally, S = {(xn,k , yn ) | n ∈ {1, . . . , N }, k ∈ {1, . . . , K}} , Q=



(9) (x′n,m , yn )

n ∈ {1, . . . , N }, m ∈ {1, . . . , M } . (10)

Episodic training. During each meta-train iteration, we sample an episode Ti = (Si , Qi ) from Dmeta_train . Given an embedding network fθ (·), the learner is adapted using only the support set Si and is evaluated on the query set Qi . Few-shot meta-learning helps model learn to adapt quickly for new tasks with a very few labeled samples, therefore, the learning strategy is different from traditional way that is, instead of learning to classify directly, the model learns how to learn from small datasets. Prototypical Networks (ProtoNet). Prototypical networks (ProtoNet) are a widely used approach for fewshot classification. The key idea of ProtoNet, as illustrated in Fig. 4, is to represent each class by a prototype, defined as the mean of its support examples in a learned embedding space. This prototype represents the center of each class, enabling rapid class initialization from K labeled examples. ProtoNet uses an embedding network VOLUME ,

FIGURE 4: Architecture of prototypical networks (ProtoNet) for AoA-based classification. fθ (·) to map an input x into values in the mapping space: z = fθ (x) ∈ Rd .

(11)

For each class c ∈ {1, ..., N }, ProtoNet computes a prototype as the mean embedding of its support set Sc : X 1 fθ (xi ). (12) pc = |Sc | (xi ,yi )∈Sc

Given a new query sample x, ProtoNet computes the squared Euclidean distance between the query embedding and the class prototype. Class probabilities are obtained by applying a softmax over negative distances:  exp −||fθ (x) − pc ||22 pθ (y = c | x) = PN . (13) 2 ′ c′ =1 exp (−||fθ (x) − pc ||2 ) The embedding parameters θ are learned by minimizing the average negative log-likelihood L = − log ptheta (y = k | x) of the true class c using stochastic gradient descent. Continual Few-shot Learning via ProtoNet. Motivated by the fact that class prototypes can be updated incrementally, we apply a ProtoNet-based solution to handle the practical setting where only a small number of labeled examples per class may be available at any time. ProtoNet handles new observed classes differently from incremental tree-based classifiers. Rather than expanding the classifier structure by creating new branches or removing weak branches. ProtoNet performs classification on an embedding space where each class is represented by a prototype. When labeled samples from a new class arrive, a small labeled support set is passed through the embedding network / encoder to obtain their embedding vectors. The corresponding class prototype is computed as the mean embedding of these support samples. This prototype is then added to the prototype set, enabling continual few-shot adaptation without retraining from scratch. For each new incoming query sample, the encoder maps it into the embedding space and computes its distances to all existing prototypes. The query is assigned to the class of the closest prototype. In the following, we report the experimental setup and results for the proposed AoA-based localization approach. We first present the performance of offline 9

Bac TN et al.: Adaptive Learning Strategies for AoA-Based Outdoor Localization: A Comprehensive Framework

machine learning models before and after improvement through hyperparameter tuning. We then compare generative models for data augmentation. Finally, we analyze the results of online learning approaches with continual tree-based and ensemble models and the fewshot learning approach. III. EXPERIMENTS A. DATASET DESCRIPTION

In this work, we evaluate our proposed approaches on a real world 64-antenna massive MIMO 50-subcarrier OFDM outdoor dataset at FR1 (2.18 GHz) collected at the Nokia campus, Stuttgart, Germany [52]. The dataset includes uplink CSI measurements from a massive MIMO digital antenna array, arranged in 4 rows × 16 columns of single-polarized patch antennas. The antenna array is installed on the roof of a 20 m-high building with a 10 degrees downtilt. The horizontal spacing between adjacent antennas is λ/2, and the vertical spacing is λ, at a central frequency of 2.18 GHz. 50 OFDM subcarriers spanning a 10 MHz bandwidth are recorded per antenna, pilot bursts are transmitted every 0.5 ms. The receiver is a monopole antenna mounted on a cart at an approximate height of 2 meters. The cart follows predefined trajectories (tracks) at an average speed of 3-5 km/h, yielding a spatial sampling interval of 0.5 mm. The size of the dataset for each individual track is reported in Table 2. Due to obstacles such as buildings and trees, the measurement area comprises both LoS and NLoS regions. In this work, we consider five LoS tracks (tracks 6, 9, 10, 11, and 12) and five NLoS tracks (tracks 1, 2, 3, 13, and 20), as shown on the campus map in Fig. 5. B. FEATURE EXTRACTION

Let H ∈ CN ×M ×T denote the CSI tensor, where N = 64 is the total number of antennas arranged in 4 rows and 16 columns, M = 50 is the number of subcarriers and T is the number of consecutive measurement locations along each track, sampled every 0.5 mm. We use the CSI dataset described above to extract AoA vectors across different antenna-array rows and OFDM subcarriers. To do so, we apply the MUSIC and ESPRIT algorithms to estimate the azimuth AoA for each data segment. Since the estimation accuracy of both methods depends on the input length, we study its impact on classification by considering window lengths W ∈ 500, 1000, 2000 consecutive measurements for AoA extraction. This corresponds to CSI spans of 0.25, 0.5, 1 meters for each AoA estimate. To further enlarge the AoA dataset, we slide the window W along each track with a 50% overlap (shift_ratio), which ensures smoother transitions between adjacent frames and increases the number of samples. To obtain a feature vector from a data segment (window) with W samples along each track, we first 10

TABLE 1: Performance of MUSIC and ESPRIT in AoA estimation across track 11. Algorithm

AoA estimation (°) Mean Std Dev

Processing time (s) Total Mean

MUSIC

-35.17

1.71

7.1672

0.0300

ESPRIT

-35.30

1.63

0.2777

0.0012

TABLE 2: Data shapes of the original CSI and estimated AoAs (MUSIC) for LoS tracks (6, 9, 10, 11, 12) and NLoS tracks (1, 2, 3, 13, 20). The dataset generated using ESPRIT follows the same format. AoA data shape by window size

Tracks

CSI data shape

500

1000

2000

1

(60, 50, 120000)

(479, 200)

(239, 200)

(119, 200)

2

(60, 50, 120000)

(479, 200)

(239, 200)

(119, 200)

3

(60, 50, 122000)

(487, 200)

(243, 200)

(121, 200)

13

(60, 50, 122000)

(487, 200)

(243, 200)

(121, 200)

20

(60, 50, 118000)

(471, 200)

(235, 200)

(117, 200)

6

(60, 50, 96000)

(383, 200)

(191, 200)

(95, 200)

9

(60, 50, 120000)

(479, 200)

(239, 200)

(119, 200)

10

(60, 50, 120000)

(479, 200)

(239, 200)

(119, 200)

11

(60, 50, 120000)

(479, 200)

(239, 200)

(119, 200)

12

(60, 50, 92000)

(367, 200)

(183, 200)

(91, 200)

estimate four AoAs independently from the four rows of the uniform linear array. We then exploit the 50 OFDM subcarriers to compute 50 AoA estimates per row, resulting in a 4×50 = 200-dimensional AoA feature vector. As a result, we construct three AoA dataset configurations, summarized in Table 2. Fig. 6 illustrates an AoA estimation example for track 11 using a window length W = 1000, at antenna row 0 and subcarrier 10. The mean AoA estimates produced by MUSIC and ESPRIT are almost identical (−35.17° and −35.30°, respectively), with comparable standard deviations of about 1.71° and 1.63°. The computational costs are reported in Table 1, MUSIC requires 7.1672 seconds of processing time which is approximately 26 times longer than ESPRIT (0.2777 seconds). On average, this corresponds to 30 ms per estimate for MUSIC versus 1.2 ms for ESPRIT. C. OFFLINE LEARNING PERFORMANCE 1) Performance of LoS / NLoS Classifier (Stage 1)

We evaluate the performance of six models: LR, KNN, RF, GBM, LightGBM, XGBoost, while progressively increasing the training-set size from 5% to 80% in steps of 5%. Table 3 reports the first-stage LoS / NLoS classification results using the AoA features estimated by MUSIC. All evaluated models achieved perfect performance with an accuracy of 1. These results indicate that, given sufficient training data, the first-stage classifier preserves perfect accuracy and does not degrade subsequent LoS VOLUME ,

FIGURE 5: Nokia campus in Stuttgart, Germany. The red rectangle denotes the mMIMO antenna array mounted on top of a building, while the lines with arrows represent the trajectories (tracks) and their respective directions. Red solid lines indicate NLoS tracks, while blue dashed lines represent LoS tracks.

Angle (degrees)

−27.5

MUSIC ESPRIT

−30.0 −32.5 −35.0 −37.5 −40.0 0

Model (W)

Acc.

F1

AUC

T (s)

I (ms)

LR (2000)

1

1

1

0.029

0.239

LR (1000)

1

1

1

0.123

0.238

LR (500)

1

1

1

0.315

0.264

KNN (2000)

1

1

1

0.002

72.657

KNN (1000)

1

1

1

0.002

77.847

200

KNN (500)

1

1

1

0.004

83.951

Sample Position (window index)

RF (2000)

1

1

1

0.377

12.485

RF (1000)

1

1

1

0.569

11.896

RF (500)

1

1

1

1.003

11.929

GBM (2000)

0.99

0.99

1

3.439

0.794

GBM (1000)

0.99

0.99

1

7.804

0.817

GBM (500)

0.99

0.99

1

17.426

0.824

LightGBM (2000)

1

1

1

19.498

11.020

LightGBM (1000)

1

1

1

27.386

11.485

LightGBM (500)

1

1

1

40.610

11.850

XGBoost (2000)

1

1

1

26.473

53.174

XGBoost (1000)

1

1

1

28.684

53.324

XGBoost (500)

1

1

1

28.540

52.469

50

100

150

FIGURE 6: Example of AoA estimation for track 11, using antenna row 0 at subcarrier 10, with a window size of 1000.

and NLoS track classifiers in the second stage. Even with only 5% of the training data, all models exceeded the accuracy of 97%, highlighting the strong discriminative power of the AoA-based features. Notably, KNN and RF reached 100% accuracy with just 5% training data (about 50 samples) for all AoA window sizes tested, W ∈ {500, 1000, 2000}. Both models required substantially less training time than the remaining methods, highlighting their efficiency. However, KNN incurs a noticeably higher inference latency than LR, RF, or GBM. For real-time scenarios where inference speed is critical, LR and GBM offer an attractive trade-off, with inference times of 0.24-0.26 ms and 0.78-0.82 ms, respectively. VOLUME ,

TABLE 3: Model performance (accuracy, F1-score, and ROC AUC), training time in seconds (T (s)), and inference time in milliseconds (I (ms)) across different window sizes (W).

2) Performance of Track Classifiers (Stage 2)

The results and the corresponding training and inference times for the ML models are summarized in Table 4. Overall, tree ensemble methods (RF, GBM, XGBoost, and LightGBM) achieve the strongest results for both MUSIC- and ESPRIT-based AoA features. The stacking ensembles (constructed by combining the top n base classifiers, with n ∈ {2, . . . , 6} consistently achieve the 11

Bac TN et al.: Adaptive Learning Strategies for AoA-Based Outdoor Localization: A Comprehensive Framework

TABLE 4: The comparison of accuracy, training time, and inference time across ML models for MUSIC and ESPRIT (window size = 2000), sorted by accuracy. Models

Accuracy (%)

Training (s)

Inference (s)

1023.9869

0.0657

MUSIC Stacking Top-6

97.98

RF

94.30

0.5239

0.0110

GBM

89.46

10.6841

0.0018

XGBoost

89.10

117.3329

0.0212

LightGBM

85.98

103.7215

0.0063

KNN

76.62

0.0020

0.0383

LR

71.09

0.6774

0.0003

Stacking Top-6

95.38

1260.2022

0.0693

RF

92.61

0.7599

0.0108

LR

86.73

0.7763

0.0002

KNN

85.06

0.0020

0.0374

XGBoost

84.69

120.0175

0.0211

GBM

83.41

28.4534

0.0017

LightGBM

82.32

137.4263

0.0044

TABLE 5: Stage 2 region-specific track classification accuracy (%) using AoA features (2000 snapshots). Baseline (settings of [43]) vs Optuna-tuned configurations. ”” indicates configurations not evaluated in the baseline [43].

3) Hyperparameter Optimization

The hyperparameter optimization module is integrated into the offline learning framework and automatically fine-tunes the models using the available dataset from the feature repository. This step helps identify optimal hyperparameters, therefore improving model robustness and stability. Since the Stage-1 LoS / NLoS classifier achieves perfect accuracy on our dataset, we do not further optimize this classifier. Instead, we focus on the Stage-2 region-specific multi-class classifiers. In the Appendix, Table 11 presents the tuned parameters and their search space. Table 13 reports the best hyperparameter configurations identified by the Optuna framework [48] for each classifier in the LoS and NLoS regions and for each AoA estimator (MUSIC and ESPRIT). For each model, Optuna runs 500 trials and selects the configuration that maximizes the mean accuracy of 5-fold cross-validation on the training data. For each 12

ESPRIT Baseline Optuna

Model

LoS

LR KNN SVM RF

71.1 76.6 94.3

74.4 77.72 98.53 99.82

86.7 85.1 92.6

88.4 92.07 99.81 97.05

GBM XGBoost LightGBM

89.5 89.1 86

99.63 97.79 97.98

83.4 84.7 82.3

94.85 93.74 92.27

LR KNN SVM RF GBM XGBoost

50.1 72.4 87.1 71.2 75

54.11 72.2 79.73 95.81 97.99 88.62

60.4 90.3 87.4 68.3 73.7

59.8 91.46 95.65 95.81 96.64 89.11

LightGBM

74.2

88.27

75.4

89.45

ESPRIT

best classification performance, particularly when paired with ESPRIT. In most cases, ESPRIT provides a higher accuracy than MUSIC as the AoA estimator. Stacking the top-6 model has the highest results (98% with MUSIC, 95.4% with ESPRIT). While lightweight models such as LR and KNN offer very fast inference, they yield lower accuracy, reaching 71.1% and 76.6% with MUSIC, respectively. Among individual classifiers, RF provides the most attractive trade-off: it achieves high accuracy (94.3% with MUSIC and 92.6% with ESPRIT) while maintaining a low training time (0.52–0.76 s) and a low inference time (approximately 11 ms).

MUSIC Baseline Optuna

Region

NLoS

Optuna-tuned configuration, we retrain each classifier on the full training set and evaluate it once on the held-out test set. We compare the accuracy of the baseline region-specific track classifier proposed in [43] and the Optuna hyperparameter optimization for both MUSIC and ESPRIT algorithms in Table 5. The results show that fine-tuning improves the performance of most models in both LoS and NLoS regions, with larger gains in NLoS with boosted ensembles (GBM / XGBoost / LightGBM). The best fine-tuned performance is achieved by 99.82% in LoS (RF, MUSIC) and 97.99% in NLoS (GBM, MUSIC), indicating that carefully tuned ensemble methods can improve model performance even under NLoS propagation. While the proposed offline localization framework combined with hyperparameter optimization achieves high accuracy under offline conditions, it does not fully capture the challenges of real-world deployments. In practice, outdoor propagation environments are highly dynamic, which can induce distribution shifts between training and operational data. Moreover, labels are not readily available for incoming CSI measurements, making the labeling process costly and time-consuming. These factors can degrade performance after deployment and motivate the need for models that can continuously update from streaming measurements, learn effectively from a few labeled samples, and remain stable without catastrophic forgetting. VOLUME ,

D. CONTINUAL / INCREMENTAL LEARNING 1) Data Augmentation

VOLUME ,

Original Synthetic

30 20

t-SNE 2

10 0 −10 −20 −30 −60

−40

−20

0

20

40

60

20

40

60

t-SNE 1

(b) CVAE 40

Original Synthetic

30 20

t-SNE 2

In this section, we evaluate the effectiveness of conditional generative models CGAN and CVAE using a real outdoor CSI dataset. We aim to assess training stability, the quality of the AoA features generated, and their impact on localization performance under limited data conditions. We train CGAN and CVAE models on 10 tracks across the LoS and NLoS regions, using 543 and 597 AoA feature vectors, respectively. The AoA features are extracted using MUSIC and ESPRIT with a window length of W = 2000 and 50% overlap. The dataset is split into training and testing sets with an 80 / 20 ratio. During training, the generative models are evaluated at each epoch, and the best-performance models are selected on the basis of validation performance. The t-SNE visualization comparing the distributions of real AoA features estimated by MUSIC and synthetic samples generated by CGAN and CVAE is shown in Fig. 7. Fig. 7(a) shows that CGAN-generated samples form a nearby but partially separated cluster, indicating a distribution mismatch between generated and real features. In contrast, CVAEgenerated samples in Fig. 7(b) largely overlap with the real data, confirming that CVAE captures the dominant structure of the feature space. The architecture of CGAN and CVAE is presented in Table 12 in the Appendix. Moreover, Fig. 8 compares the original and synthetic AoA features for several representative track samples. The results show that CVAE reconstructions better preserve the main feature patterns, whereas CGAN outputs have higher variance and more noise artifacts, suggesting that CVAE captures the dominant structure of the feature space. To further examine CVAE behavior, a three dimensional t-SNE visualization is presented in Fig. 9, illustrating the overall distribution of real and synthetic data as well as details for track 3, track 13 (NLoS region) and track 6, track 12 (LoS region). CVAE reconstructs 100 synthetic samples per class, resulting in a total of 1000 generated samples. Depending on the underlying distribution of each track, the generated samples cluster consistently with the structure of the original data. In addition, to assess the quality of the generated samples and their impact on localization performance, we use the tuned random forest classifier which is trained on real data (Section II Part A) to predict labels for synthetic samples. Five tracks from LoS region are selected and 1000 samples are generated for each class, resulting in a total of 5000 generated samples. The corresponding classification results are reported in Table 6. Since the synthetic data generated by the CGAN exhibit a distribution that deviates significantly from the training and test data, the resulting test samples become difficult for the model to predict. Consequently, the performance

(a) CGAN

10 0 −10 −20 −30 −80

−60

−40

−20

0

t-SNE 1

FIGURE 7: Compares original and synthetic samples distributions between CGAN and CVAE.

of the model decreases significantly, reaching an accuracy of approximately 37%. In contrast, CVAE-based samples introduce only slight deviations from the original data distribution, which better reflect small changes in external conditions. As a result, the model maintains strong performance, achieving an accuracy of approximately 86%. Based on the experimental results, CVAE demonstrates its suitability as a supporting module within the online learning-based localization framework. In particular, CVAE-generated samples enable effective upsampling of small replay buffers, which is critical for rehearsal and reinitialization processes. In addition, synthetic samples provide a useful tool for evaluating model robustness under distributional variations. 13

Bac TN et al.: Adaptive Learning Strategies for AoA-Based Outdoor Localization: A Comprehensive Framework

Original Synthetic

(a) CGAN

AoA

1 0 −1

NLoS (track_2)

AoA

1 0

The experiments are repeated over 10 trials. In Fig. 10, we report the mean accuracy with the corresponding standard deviation of RF and SVM with the two retraining schemes. In general, cumulative retraining achieves higher accuracy as it retains information from previously observed samples. In contrast, buffer retraining creates a wider and more fluctuating standard deviation band, indicating less stable performance over time. Although buffer retraining is more resource-efficient, it may suffer from instability and forgetting when the buffer window is too small.

−1

NLoS (track_13)

(b) CVAE

3) Continual Learning for Streaming Data

In this experiment, six incremental learning models are deployed, including aggregated mondrian forests (AMF), 0 adaptive random forests (ARF), gaussian naïve bayes (GNB), hoeffding adaptive trees (HAT), hoeffding trees −1 (HT), and streaming random patches (SRP). All models are configured using hyperparameter settings adopted NLoS (track_2) from the River benchmark for multiclass classification 1 tasks [53]. A confidence threshold τ = 0.5 is applied, 0 which means that only samples with predicted class probability greater than or equal to 50% are accepted to −1 update the models. To simulate a realistic deployment 0 25 50 75 100 125 150 175 200 NLoS (track_13) scenario for the online learning framework, CVAE-based Feature index data augmentation is applied to generate 1000 synthetic samples per class. The first 10% of the dataset is used FIGURE 8: Compares original and synthetic AoA feafor warm-up training as the initialization of the model. ture patterns between CGAN and CVAE. The remaining 90% of the data is treated as a streaming sequence, where the AoA feature vectors are processed sequentially and fed into the models one sample at a TABLE 6: Evaluation results of a tuned RF classifier on time. synthetic AoA samples (LoS region, 5 tracks). Evaluation metrics. To ensure a fair comparison, we Model Acc. (%) Precision (%) Recall (%) F1-score (%) keep the same AoA feature extraction pipeline for all CGAN 37.40 24.50 37.40 28.38 models and evaluate them under the same training strategies and metrics: CVAE 86.16 88.42 86.16 85.67 AoA

AoA

1

2) Batch Retraining with Conventional ML

To compare buffer-based and cumulative batch retraining strategies in the offline learning scheme, the dataset is divided into 10 sequential batches, each containing 10% of the total samples. This setup mimics a deployment scenario in which the offline learning framework periodically retrains the model as new data become available or retrains the model when predefined retraining criteria are met. We evaluate two retraining approaches: • Buffer retraining: the model is retrained using only a recent data window, e.g., the latest 10% of the dataset. • Cumulative retraining: the model is retrained using all samples observed so far, stored in the feature repository. 14

• Warm-up accuracy: accuracy measured immediately after the warm-up phase (model initialization). • Online accuracy: prequential accuracy computed over the online phase. • Acceptance rate: we apply a confidence threshold τ = 0.5 which means if a confidence probability is over 50% after a prediction, the sample is accepted and used for learning. • Inference latency and update (training) time: to measure computational cost. • Forgetting rate: to measure the ability of classifier against catastrophic forgetting, we compute the forgetting rate (FR) as defined:

FR =

FE , T

(14) VOLUME ,

FIGURE 9: t-SNE projection comparing real and synthetic AoA feature vectors (MUSIC) for the LoS and NLoS regions (dataset size = 1000). Cumulative + RF

Accuracy

1.0

Cumulative + SVM

0.8 0.6 0.4 LoS NLoS

0.2 0.0 0

1

2

3

1.0

Accuracy

4

5

6

7

8

9 0

1

2

3

Buffer + RF

4

5

6

7

8

9

6

7

8

9

Buffer + SVM

0.8 0.6 0.4 0.2 0.0 0

1

2

3

4

5

6

7

8

Episode

9 0

1

2

3

4

5

Episode

FIGURE 10: Comparison of buffer and cumulative batch retraining for RF and SVM (mean accuracy over 10 trials is shown, shaded bands denote standard deviation). where: FE =

T X

1[f (xt−1 ) = yt−1 ∧ f (xt ) ̸= yt ],

(15)

t=1

is the number of forgetting events over time T , that means the model made a correct prediction at time t − 1 but misclassified the sample at time t. Lower FR indicates better memory retention. Fig. 11 illustrates the learning process of the hoeffding adaptive tree classifier applied to AoA features estimated VOLUME ,

using the MUSIC algorithm. There are two phases: an initial warm-up (model initialization) phase and an online inference and update phase. During the warm-up phase (zoomed region in Fig. 11), the model is exposed to streaming data for the first time and incrementally constructs its decision structure. Due to the absence of sufficient statistics in the early stage, initial predictions are unreliable, resulting in noticeable accuracy fluctuations and a temporary performance drop within the first 200 samples. As more observations become available,

15

Bac TN et al.: Adaptive Learning Strategies for AoA-Based Outdoor Localization: A Comprehensive Framework

1.0 Train (553) Online (4989)

Accuracy

0.8

0.6

0.6

Model

0.4

0.0

0.2

0.0

0

0

1000

100

2000 3000 Sample Index

200

300

4000

400

500

5000

FIGURE 11: The hoeffding adaptive tree (HAT) classifier peformance of warm-up training and online inference phases applied to AoA features estimated using MUSIC.

the classification accuracy progressively increases and stabilizes once it reaches approximately 70%. In the online inference phase, the model continues to refine its decision boundaries in an incremental approach, leading to a gradual performance improvement that reaches 77.75% accuracy, without any sudden degradation or catastrophic collapse. This behavior is inherent to hoeffding-based decision trees, which require a sufficient number of samples to statistically validate a split. Until the hoeffding bound is satisfied, candidate splits remain uncertain and sample routing may be suboptimal. Once the bound is met, confident splits are performed, achieving accuracy recovery. After sufficient splits and leaf updates, the model converges to a more stable decision structure, resulting in consistent and robust predictions over time. The incremental learning results for six incremental models under both LoS and NLoS conditions are reported in Table 8. Across all scenarios (LoS / NLoS, MUSIC / ESPRIT), AMF consistently achieves the best online performance, reaching online accuracy around 0.94 while maintaining the lowest forgetting rate (0.0250.043). This indicates strong stability under the prequential predict-then-train strategy and shows that AMF adapts effectively without frequent degradation over time. SRP is the second-best online accuracy in both regions. ESPRIT yields a clear improvement over MUSIC in NLoS (0.753 vs 0.728), implying that the AoA estimator can meaningfully affect online learning robustness in more challenging propagation conditions. In contrast, GaussianNB and Hoeffding-tree variants accept almost all samples (acceptance is approximately 100%) achieve lower online accuracy and higher forgetting rates, suggesting that frequent incremental updates may lead to unstable decision boundaries rather than consistent improvement.

Mean (ms)

Total (s)

Inference

Update

Inference

Update

AMF

1.6925

17.7248

8.4868

88.86 21.4613

0.4 0.2

16

TABLE 7: Computational cost comparison between continual tree-based and ensemble classifiers and the continual ProtoNet-based few-shot approach. Results are averaged across LoS / NLoS and MUSIC / ESPRIT configurations.

ARF

1.6048

4.28

8.048

GNB

2.1538

0.4258

10.7988

2.1355

HAT

2.3275

2.7475

11.67

13.7763

HT

2.2795

1.6395

11.429

8.2205

SRP

10.9113

77.398

54.7395

388.1073

ProtoNet (k=1)

1.05

4.775

0.9703

4.3793

ProtoNet (k=5)

1.125

5.05

0.2545

1.1415

ProtoNet (k=10)

1.225

5.6

0.1563

0.7175

Table 7 demonstrates that incremental tree-based classifiers maintain millisecond-level inference and update latency, supporting real-time deployment. Among them, AMF and ARF provide a balanced trade-off between accuracy and computational cost, while GNB and HT offer the lowest update overhead at the expense of reduced accuracy. In contrast, ensemble methods such as SRP significantly increase computational cost, with a total update time exceeding 380 seconds, limiting scalability for long streaming sequences. Fig. 12 summarizes the performance of six incremental learning models. The bar chart reports the online accuracy, where higher values indicate better predictive performance, while the line chart reports the forgetting rate, where lower values indicate stronger retention and reduced catastrophic forgetting. The performance of the six incremental models, including accuracy after warm-up training, accuracy after online learning phase, acceptance rate, and forgetting rate are presented in the Appendix (Table 14 and Table 15). E. FEW-SHOT LEARNING 1) Prototypical Networks - Standard Episodic Meta-learning

Experiment setup. We evaluate standard ProtoNet for few-shot track identification separately under LoS and NLoS conditions. The AoA feature vectors are derived from CSI using MUSIC and ESPRIT algorithms. Each feature vector is computed over a window of 2000 snapshots, as described in Section III Part B. The LoS set includes tracks {6, 9, 10, 11, 12} and NLoS set includes tracks {1, 2, 3, 13, 20}. For each region, we construct N − way K − shot episodes with K ∈ {1, 2, 4, 8, 12} and N = 3. Using N = 3 allows us to simulate the case of insufficient samples, only a small number of classes and labeled examples are available. ProtoNet consists of an embedding network fθ implemented as a 3-layer MLP as described in Table 16. For each (N, K) episode, VOLUME ,

0.00 ARF

HAT

HT

0.00

GNB

0.50

0.1

0.25 0.00

Online acc.

0.2

ARF

HAT

SRP

HT

ARF

HAT

HT

GNB

NLoS + ESPRIT 0.2

0.75 0.50

0.1

0.25

0.0 SRP

0.0 AMF

1.00

0.75

AMF

0.1

0.25

NLoS + MUSIC

1.00

Online acc.

0.50

0.0 SRP

0.2

0.75

Forget rate

0.25

Online acc.

0.1

Forget rate

0.50

Forget rate

Online acc.

0.2

0.75

AMF

LoS + ESPRIT

1.00

Forget rate

LoS + MUSIC

1.00

0.00

GNB

0.0 AMF

Online accuracy

SRP

ARF

HAT

HT

GNB

Forgetting rate

FIGURE 12: Online Accuracy (column chart) and Forgetting Rate (line chart) comparison across six incremental models in four scenarios (LoS / NLoS × MUSIC / ESPRIT). TABLE 8: Online performance: higher Accuracy (Acc) is better, while a lower Forgetting Rate (FR) is better. AMF

SRP

ARF

GNB

HAT

HT

Acc

FR

Acc

FR

Acc

FR

Acc

FR

Acc

FR

Acc

FR

LoS

MUSIC ESPRIT

0.9432 0.9383

0.0343 0.0427

0.8161 0.8423

0.1239 0.1108

0.823 0.8244

0.1231 0.1201

0.779 0.779

0.1455 0.1455

0.7775 0.7775

0.1433 0.1433

0.7782 0.7782

0.1463 0.1463

NLoS

MUSIC ESPRIT

0.9409 0.9412

0.0248 0.0258

0.7282 0.7534

0.1723 0.1586

0.713 0.6962

0.1838 0.1874

0.6135 0.6135

0.2114 0.2114

0.604 0.604

0.2134 0.2134

0.6153 0.6153

0.2098 0.2098

experiments are repeated over 10 trials to calculate the mean accuracy with 95% confidence interval. We train ProtoNet for 1000 meta-training episodes using Adam with a learning rate 10−3 . Evaluation is performed on 200 meta-test episodes sampled from the held-out split. We use a standard 80 / 20 split of the available samples within each region to form metatrain and meta-test sets prior to episodic sampling, so meta-test can be simulated the case of unseen samples (potentially different track subsets depending on episode construction) within the same region. Table 9 summarizes the few-shot results in terms of mean episodic accuracy and 95% confidence intervals. Across both LoS and NLoS regions, performance improves as the number of shots K increases, indicating that ProtoNet benefits from additional labeled support examples per track. Under LoS, ProtoNet achieves strong performance even with low shot, improving accuracy from 0.7898 ± 0.0069 (1-shot) to 0.9029 ± 0.0111 (4-shot) and then saturating around 0.914-0.915 for K ≥ 8. In contrast, NLoS remains more challenging: accuracy increases from 0.5466 ± 0.0166 (1-shot) to 0.7566 ± 0.0115 (12-shot), with wider confidence intervals, reflecting greater variability across trials. Overall, these results confirm that additional

VOLUME ,

support examples improve performance in both cases and the persistent LoS-NLoS gap suggests that NLoS AoA features shows the difficulty in processing from the NLoS region due to multipath effects. TABLE 9: LoS vs NLoS comparison (ProtoNet, MUSIC, N=3, mean accuracy with 95% CI) and LoS-NLoS w.r.t K. K-shot

LoS acc ± 95% CI

NLoS acc ± 95% CI

Gap (LoS - NLoS)

1

0.7898 ± 0.0069

0.5466 ± 0.0166

0.2432

2

0.8541 ± 0.0145

0.6803 ± 0.0215

0.1738

4

0.9029 ± 0.0111

0.7250 ± 0.0146

0.1779

8

0.9135 ± 0.0123

0.7500 ± 0.0159

0.1635

12

0.9148 ± 0.0136

0.7566 ± 0.0115

0.1582

2) Continual Few-shot Learning via ProtoNet

In this experiment, we leverage CVAE-based data augmentation to expand the dataset to 1000 samples per class. We then construct a sequence of N −way K −shot tasks from these samples, each episode represents a subset of the newly available data and samples are not 17

Bac TN et al.: Adaptive Learning Strategies for AoA-Based Outdoor Localization: A Comprehensive Framework

• accbefore (e) denotes the classification accuracy in i episode i before updating the model with episode e, • accafter (e) denotes the classification accuracy in i episode i after updating the model with episode e.

Table 10 presents the accuracies for different values of K and its corresponding valid episodes E over 10 trials. Overall, LoS consistently outperforms NLoS, while in the NLoS setting, both AoA estimators improve their performance with K , LoS results are more stable across K . Comparing AoA estimators, ESPRIT achieves higher final accuracy than MUSIC under LoS condition (increasing from 0.8425 (K = 1) to 0.8693 (K = 10). The number of valid episodes decreases for larger K because each episode requires more labeled support samples per class, thus, the reported accuracies correspond to the last episode in each configuration. Fig. 13 illustrates the final classification accuracy with respect to K (K = 1, . . . , 10) in both LoS and NLoS regions, using AoA features extracted via the MUSIC and ESPRIT algorithms. Fig. 14 reports the evolution of the mean accuracy over training episodes, averaged over 10 trials (95% CI) for three representative values K (K = 1, 5, 10), AoA features extracted using the MUSIC algorithm. Across all conditions, accuracy increases rapidly in the early 18

1.0 0.8

Mean accuracy

reused across episodes, in contrast to standard episodic meta-training. Each episode uses N = 3 classes, with K support samples per class and Q query samples per class. Evaluation metrics. The performance of continual fewshot learning is assessed using classification accuracy, episode-level forgetting rate to quantify catastrophic forgetting, and computational cost measured through inference latency and update (training) time. Let E denote the number of episodes obtained from the split, an episode e consists of a support set and a query set, and episodes arrive sequentially as e = 1, . . . , E . The model is updated after each episode. The forgetting rate F R(e) of episode e is defined as the performance degradation in previously learned episodes caused by updating the model with episode e. In continual learning with streaming data, forgetting rate captures short-term instability as data arrive sequentially over time. In contrast, for prototypical networks operating in a continual few-shot setting, forgetting is defined at the episode level and measures the degradation of performance on previously observed episodes after sequential updates. For episode e ⩾ 2, the episode-level forgetting rate is defined as: e−1   1 X after FR(e) = max 0, accbefore (e) − acc (e) i i e − 1 i=1 (16) where:

0.6 0.4 LoS-MUSIC LoS-ESPRIT NLoS-MUSIC NLoS-ESPRIT

0.2 0.0 1

2

3

4

5

6

7

8

9

10

K-shot

FIGURE 13: Classification accuracy (mean ± 95% confidence interval) w.r.t K for continual few-shot learning under four scenarios (LoS / NLoS × MUSIC / ESPRIT). episodes and then improves more gradually, indicating progressive stabilization of the embedding space and the prototype estimates. Across four cases presented in Table 10, we observe that increasing support size K improves representation stability and reduces catastrophic forgetting. With a larger number of support samples, the network can compute more reliable prototype estimates, resulting in less noisy embedding updates. For example, in the NLoS + MUSIC configuration, the forgetting rate decreases from 0.0222 at K = 1 to 0.0076 at K = 10, corresponding to approximately a reduction of 65.8%. Moreover, NLoS region consistently has higher forgetting rates than LoS region. This is due to variable and unstable AoA representations, making episodic updates more disruptive to previously learned knowledge. For example, at K = 1, the forgetting rate for MUSIC and ESPRIT under NLoS (0.0222 and 0.0218, respectively) is higher than those under LoS (0.0124 and 0.0080, respectively). In addition, forgetting rates decrease significantly from K = 1 to K = 4 and stabilize for K ≥ 5. A similar stabilization trend is observed in classification accuracy, with no substantial performance gains beyond moderate support sizes. Regarding the inference and update times of this method, Table 7 presents three representative results for K = 1, 5, 10. The mean inference time of ProtoNet remains within 1.05-1.23 ms per episode. Similarly, the mean update time remains approximately 5 ms per episode. Although larger support sizes K slightly increase per-episode computational cost, they reduce the total number of episodes required, resulting in a shorter total update time. These findings indicate that continual ProtoNet-based few-shot learning with a moderate support size is sufficient to stabilize continual fewshot learning while avoiding unnecessary computational overhead. VOLUME ,

TABLE 10: Final accuracy (mean ± 95% CI) vs K . Number of valid episodes E for ProtoNet episodic sampling under N -way K -shot Q-query settings (LoS / NLoS; MUSIC / ESPRIT). LoS

NLoS

MUSIC

ESPRIT

MUSIC

FR

E

mean ± 95% CI

FR

E

mean ± 95% CI

ESPRIT FR

E

mean ± 95% CI

FR

K

N =Q

E

mean ± 95% CI

1

3

906

0.7616 ± 0.0037

0.0124

909

0.8425 ± 0.0031

0.0080

918

0.4371 ± 0.0040

0.0222

922

0.4649 ± 0.0052

0.0218

2

3

452

0.7892 ± 0.0051

0.0076

451

0.8431 ± 0.0057

0.0051

456

0.5186 ± 0.0063

0.0138

454

0.5280 ± 0.0064

0.0137

3

3

302

0.0067

298

0.0046

304

0.0104

306

4

3

256

0.8020 ± 0.0026

0.0061

256

0.8595 ± 0.0033

0.0043

262

0.5934 ± 0.0072

0.0100

261

0.5959 ± 0.0040

0.0104

5

3

224

0.8045 ± 0.0040

0.0071

226

0.8442 ± 0.0036

0.0044

224

0.5934 ± 0.0086

0.0092

227

0.6098 ± 0.0063

0.0094

6

3

179

0.8127 ± 0.0063

0.0057

179

0.8693 ± 0.0014

0.0040

183

0.6138 ± 0.0076

0.0082

180

0.6357 ± 0.0090

0.0085

7

3

162

0.8188 ± 0.0066

0.0061

164

0.8513 ± 0.0045

0.0042

165

0.6262 ± 0.0090

0.0077

164

0.6305 ± 0.0083

0.0083

8

3

151

0.8164 ± 0.0072

0.0059

148

0.8705 ± 0.0043

0.0046

150

0.6150 ± 0.0054

0.0082

149

0.6390 ± 0.0092

0.0082

9

3

137

0.0064

137

0.0050

140

0.0073

136

10

3

126

0.0055

125

0.0040

128

0.0076

129

0.8070 ± 0.0037

0.8247 ± 0.0061 0.8177 ± 0.0058

0.8555 ± 0.0044 0.8693 ± 0.0028

K=1

1.0

LoS Accuracy

0.8455 ± 0.0036

0.5565 ± 0.0086

0.6307 ± 0.0093 0.6424 ± 0.0109

K=5

0.5832 ± 0.0055

0.6219 ± 0.0083 0.6373 ± 0.0135

0.0109

0.0087 0.0081

K = 10

0.8 0.6 0.4 0.2

NLoS Accuracy

1.0 0.8 0.6 0.4 0.2 0

250

500

750

0

Episode

100

Episode

200

0

50

100

Episode

FIGURE 14: Evolution of the mean accuracy over training episodes for ProtoNet under three K -shot settings (K = 1, 5, 10; MUSIC). Solid lines indicate the mean accuracy, while the shaded regions correspond to the 95% confidence intervals. F. Discussion

This section presented the experimental setup and results for AoA-based localization frameworks under both offline and online learning settings. In the offline learning framework, the fine-tuned models achieve high

VOLUME ,

localization performance, with accuracy reaching up to 99.8% for trajectory classification in both LoS and NLoS regions. These results indicate that offline learning can provide a strong and stable baseline, supporting its use as a component in physical-layer authentication or as an

19

Bac TN et al.: Adaptive Learning Strategies for AoA-Based Outdoor Localization: A Comprehensive Framework

input feature for trust evaluation schemes, as proposed in [54]. For the online learning framework, we investigate two learning approaches in simulated dynamic environments generated using CVAE-based data augmentation. The first approach employs incremental learning with treebased classifiers and ensemble models to continuously update the model under streaming data conditions. The second approach explores continual few-shot learning using prototypical networks, enabling rapid adaptation when only limited labeled samples are available. Given the same number of learning samples, the incremental learning approach achieves higher accuracy than the few-shot approach, with average accuracies of 94.32% and 87.05%, respectively. In addition, the computational comparison reported in Table 7 reveals distinct online learning strategies between these two approaches. ProtoNet achieves an inference latency in the range of 1.05-1.23 ms per episode, comparable to incremental tree-based models and significantly faster than ensemble-heavy methods such as SRP. ProtoNet also achieves moderate update cost (approximately 5 ms per episode), performs substantially fewer updates than incremental classifiers, resulting in a lower total adaptation time for moderate support sizes (0.72 s for K = 10). Incremental models update after every incoming sample (about 5000 updates in our setting), enabling immediate adaptation without requiring data buffering. In contrast, ProtoNet aggregates the samples into episodes (900 episodes for K = 1, about 120 episodes for K = 10) and performs fewer but heavier updates. While this reduces the total update (training) time for larger support sizes, it introduces a delay due to episode construction and temporary data storage. In summary, our findings demonstrate that both incremental classifiers and continual ProtoNet-based few-shot learning achieve millisecond-level inference and update costs, making them suitable for real-time or near realtime high-accuracy localization frameworks. Moreover, the consistently strong localization performance achieved under LoS conditions suggests that channel propagation characteristics can be naturally incorporated into trustworthiness assessment. In particular, localization decisions obtained in LoS environments can be treated as more reliable and can serve as trustworthiness metrics for wireless communication tasks.

updates for streaming data. In general, the findings highlight the potential of AoA-based learning for practical low-latency localization. There are several directions for future work. First, further optimization and fine-tuning of continual learning models can be investigated to improve performance and stabilize the learning process. Second, the framework could be extended to more challenging settings to evaluate its robustness under practical impairments such as blockage, calibration mismatch, and interference. Third, the framework could be integrated into physical layer authentication systems or incorporated into trust and reputation management frameworks as a physical layer trust indicator. APPENDIX

TABLE 11: Hyperparameter tuning setup and search spaces Model

Tuned hyperparameters and their search space C ∼ LogUniform[10−5 , 102 ];

LR

solver ∈ {liblinear, lbfgs, saga}; penalty=l2; max_iter=1000 k ∈ [1, 30]; weights ∈ {uniform, distance};

KNN

algorithm ∈ {auto, ball_tree, kd_tree, brute}; leaf_size ∈ [10, 50] kernel ∈ {linear, rbf, poly, sigmoid};

SVM

C ∼ LogUniform[10−5 , 102 ]; γ ∼ LogUniform[10−2 , 1] if non-linear; degree ∈ [2, 5] if poly

RF

n_estimators ∈ [50, 300]; max_depth ∈ [3, 30]; max_features ∈ {sqrt, log2, None}; min_samples_split ∈ [2, 10]; min_samples_leaf ∈ [1, 10]

GBM

n_estimators ∈ [50, 300]; learning_rate ∼ LogUniform[10−3 , 0.3]; max_depth ∈ [3, 10]; subsample ∈ [0.5, 1.0]; max_features ∈ {sqrt, log2, None}

XGB

n_estimators ∈ [50, 300]; learning_rate ∼ LogUniform[10−3 , 0.3]; max_depth ∈ [3, 10]; subsample ∈ [0.5, 1.0]; colsample_bytree ∈ [0.5, 1.0]; gamma ∈ [0, 5]

LGBM

n_estimators ∈ [50, 300]; num_leaves ∈ [20, 150]; learning_rate ∼ LogUniform[10−3 , 0.5]; max_depth ∈ [3, 15]; subsample ∈ [0.5, 1.0]; colsample_bytree ∈ [0.5, 1.0]

IV. CONCLUSIONS

This paper proposed an adaptive AoA-based localization frameworks for wireless environment, consisting of offline learning and online learning strategies to support different deployment conditions. Experimental results on a real-world outdoor mMIMO OFDM CSI dataset showed that the proposed framework can achieve strong localization performance and supports incremental model 20

TABLE 12: Architectures of generative models Model

Network (MLP)

CGAN

G : [z(100), c(20)] → 128 → 256 → 200 D : [x(200), c(20)] → 128 → 64 → 1

CVAE

Encoder: [x(200), c(20)] → 256 → 128 → (µ, log σ 2 ) Decoder: [z(64), c(20)] → 128 → 256 → 200

VOLUME ,

TABLE 13: Best hyperparameters found by Optuna (500 trials, 5-fold cross-validation). Region

AoA

MUSIC

Model

Best hyperparameters

RF

“n_estimators”: 297, “max_depth”: 27, “max_features”: ”log2”, “min_samples_split”: 3, “min_samples_leaf”: 2

AMF

0.6559

0.9409

0.6290

0.0248

SVM

“kernel”: “rbf”, “C”: 58.5476005441, “gamma”: 0.02981714257685

ARF

0.5448

0.7130

0.5391

0.1838

GNB

0.5573

0.6135

0.9994

0.2114

“n_estimators”: 256, “learning_rate”: 0.0585218637465, “max_depth”: 10, “subsample”: 0.770060357471452, “max_features”: “log2” “n_estimators”: 295,

HAT

0.5609

0.6040

0.9982

0.2134

HT

0.5538

0.6153

0.9994

0.2098

SRP

0.5573

0.7282

0.7729

0.1723

LoS

ESPRIT

MUSIC

GBM

NLoS

ESPRIT

TABLE 15: The detailed performance of six incremental models is evaluated in the NLoS condition with MUSIC / ESPRIT, including warm-up training accuracy (Training Acc), online inference accuracy (Online Acc), acceptance rate (Accept. Rate), and forgetting rate.

GBM

“learning_rate”: 0.0500295940890, “max_depth”: 10, “subsample”: 0.8695777940266635, “max_features”: ”log2”

TABLE 14: The detailed performance of six incremental models is evaluated in the LoS condition with MUSIC / ESPRIT, including warm-up training accuracy (Training Acc), online inference accuracy (Online Acc), acceptance rate (Accept. Rate), and forgetting rate.

(a) NLoS + MUSIC Model Training Acc Online Acc Accept. Rate Forgetting Rate

(b) NLoS + ESPRIT Model Training Acc Online Acc Accept. Rate Forgetting Rate AMF

0.6631

0.9412

0.6020

0.0258

ARF

0.5430

0.6962

0.5359

0.1874

GNB

0.5573

0.6135

0.9994

0.2114

HAT

0.5609

0.6040

0.9982

0.2134

HT

0.5538

0.6153

0.9994

0.2098

SRP

0.5573

0.7534

0.7539

0.1586

TABLE 16: Prototypical network architecture specification Component

(a) LoS + MUSIC Model Training Acc Online Acc Accept. Rate Forgetting Rate AMF

0.7559

0.9432

0.7286

0.0343

ARF

0.7269

0.8230

0.7996

0.1231

GNB

0.7360

0.7790

1.0000

0.1455

HAT

0.7215

0.7775

1.0000

0.1433

HT

0.7360

0.7782

1.0000

0.1463

SRP

0.7251

0.8161

0.8962

0.1239

Details

Input feature

x ∈ R200 (AoA feature vector)

Embedding network fθ

MLP 200 → 128 → 64 → 32 (BatchNorm + ReLU + Dropout(0.3))

Prototype pc

mean embedding of samples in class c

Distance

squared Euclidean distance in R32

Classifier

softmax over negative distances

Optimized parameters

θ only (embedding network); prototypes are computed per episode

(b) LoS + ESPRIT Model Training Acc Online Acc Accept. Rate Forgetting Rate AMF

0.7993

0.9383

0.7280

0.0427

ARF

0.7143

0.8244

0.8236

0.1201

GNB

0.7360

0.7790

1.0000

0.1455

HAT

0.7215

0.7775

1.0000

0.1433

HT

0.7360

0.7782

1.0000

0.1463

SRP

0.7071

0.8423

0.8511

0.1108

REFERENCES [1] C.-H. Hsieh, J.-Y. Chen, and B.-H. Nien, “Deep learning-based indoor localization using received signal strength and channel state information,” IEEE Access, vol. 7, pp. 33 256–33 267, 2019. [2] H. Chen, Y. Zhang, W. Li, X. Tao, and P. Zhang, “ConFi: Convolutional neural networks based indoor Wi-Fi localization using channel state information,” IEEE Access, vol. 5, pp.

VOLUME ,

18 066–18 074, 2017. [3] M. Mitev, T. M. Pham, A. Chorti, A. N. Barreto, and G. Fettweis, “Physical Layer Security—From Theory to Practice,” IEEE BITS the Information Theory Magazine, vol. 3, no. 2, pp. 67–79, 2023. [4] S. Gil, M. Yemini, A. Chorti, A. Nedić, H. V. Poor, and A. J. Goldsmith, “How Physicality Enables Cy-Trust: A New Era of Trust-Centered Cyber–Physical Systems,” Proceedings of the IEEE, vol. 113, no. 10, pp. 1121–1154, 2025. [5] T. M. Pham, L. Senigagliesi, M. Baldi, G. P. Fettweis, and A. Chorti, “Machine Learning-Based Robust Physical Layer Authentication Using Angle of Arrival Estimation,” in IEEE GLOBECOM 2023, 2023. [6] M. Srinivasan, L. Senigagliesi, H. Chen, A. Chorti, M. Baldi, and H. Wymeersch, “Aoa-based physical layer authentication in analog arrays under impersonation attacks,” in 2024 IEEE 25th International Workshop on Signal Processing Advances in Wireless Communications (SPAWC), 2024, pp. 496–500. [7] T. M. Pham, L. Senigagliesi, M. Baldi, R. F. Schaefer, G. P. Fettweis, and A. Chorti, “Leveraging angle of arrival estimation against impersonation attacks in physical layer

21

Bac TN et al.: Adaptive Learning Strategies for AoA-Based Outdoor Localization: A Comprehensive Framework

authentication,” IEEE Transactions on Information Forensics and Security, vol. 21, pp. 3226–3239, 2026. [8] S. Skaperas and A. Chorti, “On the robustness of aoa as an authentication feature under spoofing: Fundamental limits from misspecified cramer rao theory,” 2026. [Online]. Available: https://arxiv.org/abs/2603.21219 [9] G. K. Fischer, T. Schaechtle, A. Gabbrielli, J. Bordoy, I. Häring, F. Höflinger, and S. J. Rupitsch, “A systematic survey and comparative analysis of angular-based indoor localization and positioning technologies,” IEEE Communications Surveys & Tutorials, 2025. [10] X. YAO, Z. XU, and F. QIANG, “High-Precision Indoor Localization via Dual-Modal AOA/TOA Fusion with Deep Learning and Particle Filters.” Radioengineering, vol. 34, no. 4, 2025. [11] D. Liu, L. Wu, and Z. Zhang, “Deep Learning-Enhanced Indoor Localization Using Joint AOA-TOA Fingerprints,” in 2025 6th Information Communication Technologies Conference (ICTC), 2025, pp. 184–189. [12] D. Burghal, A. T. Ravi, V. Rao, A. A. Alghafis, and A. F. Molisch, “A Comprehensive Survey of Machine Learning Based Localization With Wireless Signals,” arXiv preprint arXiv:2012.11171, 2020. [13] A. Sobehy, É. Renault, and P. Mühlethaler, “CSI-MIMO: KNearest Neighbor Applied to Indoor Localization,” in IEEE ICC 2020, 2020. [14] T. Yang, A. Cabani, and H. Chafouk, “A Survey of Recent Indoor Localization Scenarios and Methodologies,” Sensors, vol. 21, 2021. [15] F. Bao, S. Mazokha, and J. O. Hallstrom, “Mobintel: Passive Outdoor Localization via RSSI and Machine Learning,” in Proc. IEEE WiMob 2021, 2021, pp. 247–252. [16] S. Chen, J. Fan, X. Luo, and Y. Zhang, “Multipath-Based CSI Fingerprinting Localization With a Machine Learning Approach,” in 2018 Wireless Advanced (WiAd), 2018. [17] M. Akrout, A. Feriani, F. Bellili, A. Mezghani, and E. Hossain, “Continual learning-based MIMO channel estimation: A benchmarking study,” in IEEE ICC 2023, 2023, pp. 2631–2636. [18] Y. Zhang, F. He, Y. Wang, D. Wu, and G. Yu, “CSIbased cross-scene human activity recognition with incremental learning,” Neural Computing and Applications, vol. 35, no. 17, pp. 12 415–12 432, 2023. [19] T. Zhang, Q. Fu, H. Ding, G. Wang, and F. Wang, “CAREC: Continual Wireless Action Recognition with Expansion– Compression Coordination,” Sensors, vol. 25, no. 15, p. 4706, 2025. [20] M. A. Mohsin, M. Umer, A. Bilal, M. A. Jamshed, and J. M. Cioffi, “Continual Learning for Wireless Channel Prediction,” in Proc. 42nd International Conference on Machine Learning, Vancouver, 2025. [21] J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska et al., “Overcoming catastrophic forgetting in neural networks,” Proceedings of the national academy of sciences, vol. 114, no. 13, 2017. [22] R. Aljundi, F. Babiloni, M. Elhoseiny, M. Rohrbach, and T. Tuytelaars, “Memory aware synapses: Learning what (not) to forget,” in Proc. of the European Conference on Computer Vision (ECCV), 2018. [23] S.-A. Rebuffi, A. Kolesnikov, G. Sperl, and C. H. Lampert, “icarl: Incremental classifier and representation learning,” in Proc. IEEE conference on Computer Vision and Pattern Recognition, 2017. [24] G. Hulten, L. Spencer, and P. Domingos, “Mining timechanging data streams,” in Proceedings of the seventh ACM SIGKDD international conference on Knowledge discovery and data mining, 2001. [25] A. Bifet and R. Gavalda, “Adaptive learning from evolving data streams,” in International symposium on intelligent data analysis. Springer, 2009, pp. 249–260. [26] H. M. Gomes, A. Bifet, J. Read, J. P. Barddal, F. Enembreck, B. Pfharinger, G. Holmes, and T. Abdessalem, “Adaptive random forests for evolving data stream classification,” Machine Learning, vol. 106, no. 9, pp. 1469–1495, 2017.

22

[27] H. M. Gomes, J. Read, and A. Bifet, “Streaming random patches for evolving data stream classification,” in 2019 IEEE international conference on data mining (ICDM). IEEE, 2019, pp. 240–249. [28] J. Mourtada, S. Gaïffas, and E. Scornet, “AMF: Aggregated Mondrian forests for online learning,” Journal of the Royal Statistical Society Series B: Statistical Methodology, vol. 83, no. 3, pp. 505–533, 2021. [29] J. Montiel, M. Halford, S. M. Mastelini, G. Bolmier, R. Sourty, R. Vaysse, A. Zouitine, H. M. Gomes, J. Read, T. Abdessalem et al., “River: machine learning for streaming data in Python,” 2021. [30] H. Gharoun, F. Momenifar, F. Chen, and A. H. Gandomi, “Meta-learning approaches for few-shot learning: A survey of recent advances,” ACM Computing Surveys, vol. 56, no. 12, pp. 1–41, 2024. [31] J. Snell, K. Swersky, and R. Zemel, “Prototypical networks for few-shot learning,” Advances in neural information processing systems, vol. 30, 2017. [32] C. Finn, P. Abbeel, and S. Levine, “Model-agnostic metalearning for fast adaptation of deep networks,” in International conference on machine learning. PMLR, 2017, pp. 1126–1135. [33] F. Xie, S. H. Lam, M. Xie, and C. Wang, “Few-shot learning in wi-fi-based indoor positioning,” Biomimetics, vol. 9, no. 9, p. 551, 2024. [34] J. Gao, D. Wu, F. Yin, Q. Kong, L. Xu, and S. Cui, “MetaLoc: Learning to learn wireless localization,” IEEE Journal on Selected Areas in Communications, vol. 41, no. 12, pp. 3831– 3847, 2023. [35] Z. Cui, S. Zhang, K. Lou, and L.-N. Tran, “Profi-net: Prototype-based feature attention with curriculum augmentation for wifi-based gesture recognition,” in Asia-Pacific Web (APWeb) and Web-Age Information Management (WAIM) Joint International Conference on Web and Big Data. Springer, 2025, pp. 191–204. [36] O. Mashaal, E. Mohammed, A. Digby, P. Leone, L. Swersky, A. Eshaghbeigi, and H. Abou-Zeid, “Lightweight and Generalizable AoA Estimation for IoT: A Novel Few-Shot Learning Approach,” in ICC 2025-IEEE International Conference on Communications, 2025. [37] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. WardeFarley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Communications of the ACM, vol. 63, no. 11, 2020. [38] L. Pinheiro Cinelli, M. Araújo Marins, E. A. Barros da Silva, and S. Lima Netto, “Variational autoencoder,” in Variational methods for machine learning with applications to deep networks. Springer, 2021. [39] M. Mirza and S. Osindero, “Conditional generative adversarial nets,” arXiv preprint arXiv:1411.1784, 2014. [40] K. Sohn, H. Lee, and X. Yan, “Learning structured output representation using deep conditional generative models,” Advances in neural information processing systems, vol. 28, 2015. [41] J. Wang, C. Zhao, H. Du, G. Sun, J. Kang, S. Mao, D. Niyato, and D. I. Kim, “Generative AI enabled robust data augmentation for wireless sensing in ISAC networks,” IEEE Journal on Selected Areas in Communications, 2025. [42] X. Feng, K. A. Nguyen, and Z. Luo, “A survey on data augmentation for WiFi fingerprinting indoor positioning,” IEEE Sensors Reviews, vol. 2, pp. 246 – 264, June 2025. [43] B. Trinh-Nguyen, S. Berri, S. G. Teo, T. Truong-Huu, and A. Chorti, “High-accuracy aoa-based localization using hierarchical ml classifiers in outdoor environments,” in GLOBECOM 2025-2025 IEEE Global Communications Conference. IEEE, 2025, pp. 2180–2185. [44] R. Schmidt, “Multiple Emitter Location and Signal Parameter Estimation,” IEEE Trans. Antennas Propag., vol. 34, no. 3, 1986. [45] R. Roy and T. Kailath, “ESPRIT-Estimation of Signal Parameters via Rotational Invariance Techniques,” IEEE Trans. Acoust., Speech, Signal. Process., vol. 37, no. 7, pp. 984–995, 1989.

VOLUME ,

[46] D. H. Wolpert, “Stacked Generalization,” Neural Networks, 1992. [47] L. Owen, Hyperparameter Tuning with Python: Boost your machine learning model’s performance via hyperparameter tuning. Packt Publishing Ltd, 2022. [48] T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A next-generation hyperparameter optimization framework,” in The 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2019, pp. 2623–2631. [49] M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein generative adversarial networks,” in International conference on machine learning. PMLR, 2017, pp. 214–223. [50] I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville, “Improved training of wasserstein gans,” Advances in neural information processing systems, vol. 30, 2017. [51] D. P. Kingma, T. Salimans, and M. Welling, “Variational dropout and the local reparameterization trick,” Advances in neural information processing systems, vol. 28, 2015. [52] M. K. Shehzad, L. Rose, S. Wesemann, and M. Assaad, “MLBased Massive MIMO Channel Prediction: Does It Work on Real-World Data?” IEEE Wireless Communications Letters, vol. 11, no. 4, 2022. [53] T. F. of Online Machine Learning, “Multiclass classification - River — riverml.xyz,” https://riverml.xyz/0.23.0/ benchmarks/{M}ulticlass%20classification/, [Accessed 11-022026]. [54] B. Trinh-Nguyen, S. Berri, S. G. Teo, T. Truong-Huu, and A. Chorti, “A Framework for Global Trust and Reputation Management in 6G Networks,” in 7th International Conference on Machine Learning for Networking (MLN 2024), Reims, France, 2024.

Bac Trinh-Nguyen received his B.Sc. and M.Sc. degrees from the University of Information Technology, VNU-HCM in 2018 and 2021, respectively. He is currently pursuing his Ph.D. within the joint IPAL project, a collaboration between CY Cergy Paris University - specifically the ETIS (Équipe Traitement de l’Information et Systèmes) Laboratory, ICI (Information, Communication, and Images) team - and the Institute for Infocomm Research (I2R), Agency for Science, Technology and Research (ASTAR), Singapore. His research interests include cybersecurity and the application of machine learning to wireless communications, with a focus on physical layer security (PLS), localization, trustworthiness and privacy.

Sara Berri is an Associate Professor (MCF) at the CY Cergy Paris University and a member of the ETIS (Equipe Traitement de l’Information et Systèmes) Laboratory, ICI (Information, Communication, and Images) team, since 2020. Previously, she was a postdoctoral researcher at Télécom Paris (2018-2020), CCN (Cybersecurity for Communication Networks) team of LTCI (Laboratoire Traitement et Communication de l’Information) Laboratory and she defended her PhD thesis in March 2018. She serves as Associate Editor for the IEEE Transactions on Network and Service Management. She is the recipient of the 2025 CY Alliance award: ”For women and science”. Her research encompasses several aspects of network optimization, O-RAN, localization, resource allocation, privacy in location-based services, intelligent transport systems using optimization, machine learning and game theory.

VOLUME ,

Sin G. Teo is a Scientist at Institute for Infocomm Research (I2R), Agency for Science, Technology and Research (A*STAR), Singapore. Currently, he is a principal investigator of several projects that blends the domains of Artificial Intelligence and Cybersecurity, and leads a team of A.I. for Cybersecurity in the Cybersecurity department. He obtained the Ph.D. degree from Monash University, Australia in 2016. His research interests include applied cryptography, data privacy and security, malware and network anomaly classification, federated learning, and deep learning.

Tram Truong-Huu (M’12 – SM’15) is an Associate Professor at the Singapore Institute of Technology (SIT), Infocomm Technology (ICT) Cluster, currently serving as the Programme Leader for SIT Bachelor of Engineering with Honours in Information and Communications Technology majoring in Information Security. He has been a senior computer scientist at the Institute for Infocomm Research (I2R), the Agency for Science, Technology and Research (A*STAR), since May 2019, and held a joint appointment from August 2021 to August 2025. He received his Ph.D. degree in computer science from the University of Nice - Sophia Antipolis (now Côte d’Azur University), France, in December 2010. From January 2011 to June 2012, he held a post-doctoral fellowship at the French National Center for Scientific Research (CNRS), France. He worked at the National University of Singapore as a research fellow from July 2012 and then senior research fellow from January 2017. His research interests focus on federated learning and the application of artificial intelligence to cybersecurity and next-generation networking. He has been a member of the IEEE since 2012 and a senior member since 2015.

Arsenia Chorti is a Professor at the École Nationale Supérieure de l’Électronique et de ses Applications (ENSEA) at the ETIS Lab UMR 8051, Research Fellow of the Barkhausen Institut gGmbH and a Visiting Scholar at Princeton University. Her research spans the areas of wireless communications and wireless systems security for 5G and 6G, with a particular focus on physical layer security. Current research topics include: context aware security, 5G / 6G, integrated sensing and communications, machine learning for communications, semantic and goal oriented communications. She is a Senior IEEE Member, and has served as IEEE Distinguished Lecturer (2024-2025), Associate Editor in Chief of the IEEE ComSoc Best Readings, Member of the IEEE INGR on Security, Chair of the IEEE Focus Group on Physical Layer Security (202124) and member of the IEEE Teaching Awards Committee (201719). She is currently member of various ITU Working Groups including on CGDatasets. She has participated in the reduction of the ITU report M.2516-0 on Future technology trends of terrestrial International Mobile Telecommunications systems towards 2030 and beyond (section on trustworthiness). She has served in the IEEE P1940 Standardization Workgroup on “Standard profiles for ISO 8583 authentication services”. She was selected as one of the “100 Brilliant and Inspiring Women in 6G” for 3 consecutive years between 2024-26.

23

Record · ID 158556 · SHA-256 48255b67ac45ded0
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.