Conceptio › Archive › arXiv CS
arXiv CSopen access

CSI-Agent: LLM-Assisted Few-Shot Adaptation for Cross-Domain Wi-Fi CSI Sensing

Tianya Zhao et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

CSI-Agent: LLM-Assisted Few-Shot Adaptation for Cross-Domain Wi-Fi CSI Sensing Tianya Zhao∗ , Chuan Liu, Xuyu Wang* §

arXiv:2609.31990v1 [cs.AI] 25 Sep 2026

∗

Knight Foundation School of Computing and Information Sciences, Florida International University, Miami, FL 33199, US Emails: [email protected], [email protected], [email protected]

Abstract—Wi-Fi channel state information (CSI) has enabled device-free sensing applications such as human activity recognition. However, CSI sensing models remain brittle in crossdomain deployment, where changes in users or environments can produce incorrect predictions. Existing solutions usually treat this problem as an offline model-design problem, by pretraining a stronger representation or applying one fixed adaptation method to the entire target domain. In practice, labeled target data are scarce and different classes may fail in different ways under the same domain shift. To address this, we propose CSI-Agent, an evidence-seeking LLM agent that reformulates cross-domain CSI adaptation as a deployment-time decisionmaking problem. Rather than processing raw CSI or making sample-level predictions, CSI-Agent summarizes target-domain behavior into sensing-grounded class-level evidence. It establishes a strong target-adaptive default from complementary CSI views and uses an LLM planner to determine whether each class should retain the default or invoke a specialized action. Deterministic verification and bounded execution further reduce unreliable interventions. We evaluate CSI-Agent on four public datasets using five cross-domain splits covering device, user, environment, and compositional shifts. Under 1-shot adaptation, CSI-Agent achieves the best target-domain performance across all splits and improves the average Macro-F1 by about 16% compared to the strongest baseline method. Index Terms—Wi-Fi Sensing, Channel State Information, Large Language Model, Domain Shift.

I. I NTRODUCTION Wi-Fi channel state information (CSI) has emerged as an important sensing modality for ubiquitous and mobile computing [1]–[4]. Wi-Fi sensing is particularly attractive because it requires no dedicated wearable devices, avoids capturing visual content, and can leverage existing commodity wireless infrastructure. By characterizing how wireless signals propagate through an environment, CSI captures multipath variations induced by human motion, body posture, gestures, and interactions with surrounding objects [5]. Recent advances in deep learning have substantially improved the performance of CSI-based applications, including human activity recognition [6], fall detection [7], and gesture recognition [8]. Despite this progress, CSI-based sensing remains fragile under cross-domain deployment [9]. In this work, a domain is characterized by deployment-specific factors that affect the CSI distribution, such as the environment, user, and transceiver configuration. A domain shift occurs when these factors differ between training and deployment, even though the sensing task remains unchanged. For example, a CSI model trained in one room may suffer substantial performance degradation

when deployed in another. This fragility arises because CSI is jointly shaped by human motion, environmental geometry, transceiver configuration, hardware characteristics, and multipath propagation. Consequently, the same activity may produce substantially different CSI patterns across domains, while different activities may become less distinguishable in an unseen domain. Existing approaches address domain shift at different stages of the sensing pipeline. Before deployment, supervised training across multiple domains and self-supervised pretraining aim to learn more transferable representations [10], [11]. During deployment, domain adaptation leverages target-domain observations to align source and target distributions [6], [12], [13], while few-shot adaptation uses a small number of labeled target samples to specialize the model to a new domain [14]– [16]. These approaches have substantially improved crossdomain generalization, but they make different assumptions about the target-domain data available at deployment. Collecting a large labeled CSI dataset for every new deployment is costly and often impractical. Few-shot target-domain adaptation therefore offers a practical compromise: limited labeled examples may be collected, for example, through a brief user-enrollment or site-calibration procedure [17]–[19]. These examples provide valuable evidence about the new domain without requiring full-scale data collection and massive model retraining. After this initial setup, the sensing system should operate autonomously on subsequent unlabeled observations without repeatedly requesting additional annotations. However, existing few-shot adaptation methods typically follow a predetermined adaptation recipe and apply it uniformly across the target domain. Such a fixed treatment overlooks the heterogeneous effects of domain shift and adaptation across classes. A mechanism may improve a large portion of classes, have little effect on others, and further degrade those for which it is poorly matched. Consequently, no single mechanism is uniformly beneficial across the target domain, and its global application may introduce negative transfer. Selecting different actions for different data is also non-trivial, because cross-domain models can remain highly confident even when their predictions are incorrect. These limitations motivate a different view of cross-domain Wi-Fi CSI sensing systems. Rather than treating adaptation solely as the optimization of a fixed model or the execution of a predetermined action, we formulate it as a deploymenttime decision-making problem. Under this formulation, the

sensing system must diagnose specific domain shift behavior, determine whether and how recovery should be performed, and reject actions that may degrade sensing performance. Therefore, the focus shifts from learning a single adapted model to controlling a set of complementary sensing actions under limited target-domain evidence. This deployment-time decision process naturally calls for an agentic controller. Large language model (LLM) agents provide a promising foundation for such a controller because they can integrate heterogeneous structured evidence, coordinate specialized tools, and provide explicit rationales for their decisions [20]–[23]. Rather than applying one single adaptation rule uniformly, an LLM agent can reason about the observed deployment conditions and construct a class-conditioned recovery plan. Moreover, the same highlevel planner can be applied across users, environments, and transceiver configurations without retraining for each new domain shift. Challenges. Realizing this LLM-agent-based approach for CSI sensing presents three key challenges. First, raw CSI measurements are high-dimensional, non-linguistic signals that are unsuitable for reliable interpretation by general-purpose LLMs. The agent must reason from sensing-grounded evidence that preserves the numerical and physical characteristics relevant to domain shift. Second, the available target-domain evidence is limited and uncertain. Labeled data from the new domain are scarce and costly to obtain, while most targetdomain samples remain unlabeled. Different diagnostic signals may also provide conflicting indications about the behavior of individual classes. Therefore, the agent must determine whether adaptation is necessary and which recovery action is appropriate under this uncertainty. Third, an adaptation action that appears promising based on limited evidence may still degrade certain classes or reduce overall sensing performance. Therefore, each proposed action must be evaluated before execution, while a reliable fallback should be retained when the evidence is insufficient. Solution. To address these challenges, we develop CSIAgent, an LLM-agent-based system for cross-domain CSI adaptation at deployment time. First, specialized sensing tools process raw CSI and convert it into structured numerical and physical evidence, including the reliability of limited labeled target data, class-wise activation and false-activation statistics, and the relative behavior of complementary signal views. Second, the LLM agent integrates this evidence with statistics derived from unlabeled target observations to diagnose transfer behavior for each class and construct a class-conditioned adaptation plan. Third, a sensing-grounded verifier evaluates every proposed action using reliability and no-harm criteria and replaces unreliable actions with a safer fallback. In this design, the LLM does not directly classify CSI. Instead, it provides the control intelligence needed to reason under limited targetdomain supervision, coordinate deterministic sensing tools, and produce auditable adaptation decisions. The main contributions of this paper are as follows. • To the best of our knowledge, this is the first work to

investigate LLM-agent-based deployment-time adaptation for cross-domain Wi-Fi sensing. We reformulate adaptation from a fixed model-optimization procedure into a decision-making problem with practical constraints. • We propose CSI-Agent, a sensing-grounded agentic framework that separates high-level reasoning from CSI signal processing. Specialized tools extract structured diagnostic evidence, an LLM agent constructs a classconditioned adaptation plan, and a verifier evaluates the proposed actions using reliability and no-harm criteria before execution. • We evaluate CSI-Agent on four public Wi-Fi CSI datasets using five cross-domain splits covering device, user, environment, and compositional shifts. The results demonstrate that CSI-Agent improves cross-domain sensing performance over fixed adaptation baselines, reduces negative transfer, and produces auditable adaptation decisions. The rest of the paper is organized as follows. Section II provides background and motivation. Section III discusses related work. Section IV illustrates the problem formulation. Section V presents our CSI-Agent design. Section VI evaluates the proposed system. Section VII concludes the paper. II. BACKGROUND AND M OTIVATION A. Channel State Information In OFDM-based Wi-Fi systems, CSI estimates the complex channel response between a transmitter and receiver over multiple subcarriers and antenna links. For subcarrier m at time t, a CSI coefficient can be written as Ht (m) = At (m)ejϕt (m) , where At (m) and ϕt (m) denote the measured amplitude and phase, respectively. The received signal contains multiple propagation paths caused by reflection, diffraction, and scattering. Human motion changes the lengths and strengths of these paths, producing temporal variations in CSI that encode activity-related information. By collecting CSI over consecutive packets, a sensing system obtains a time-frequency representation that can be used for activity recognition [5]. B. Preliminary Study We conduct preliminary experiments on CSI-Bench [24] to further motivate deployment-time adaptation for cross-domain Wi-Fi CSI sensing. The study considers a practical few-shot deployment setting in which a source-trained sensing model and a small labeled target set are available, while the remaining target samples are unlabeled. We first examine whether a fixed adaptation strategy is uniformly reliable across target domains and classes. We then investigate whether reliable LLM planning requires structured, sensing-grounded evidence rather than source-model confidence or raw CSI. Observation 1: Fixed adaptation is target-dependent. We use a standard 1-shot prototypical network (PTN) [16] as a representative fixed few-shot adaptation strategy and compare it with the source ResNet [25]. A natural use of limited target labels is to apply the same adaptation strategy to every target domain. However, Fig. 1(a) shows that this fixed adaptation is not uniformly reliable. The 1-shot PTN improves macro-F1

(a) Target-dependent adaptation

(b) Class-dependent adaptation

Fig. 1. Fixed few-shot adaptation is not uniformly reliable. Dev. 1 and Dev. 2 denote EchoSpot and AmazonPlug, while Comp. 1 and Comp. 2 denote U02/E01 and U05/E05 in CSI-Bench [24], respectively. (a) A fixed 1-shot PTN can either improve or degrade different target domains. (b) Within the same cross-device split, PTN affects activity classes unevenly.

on Dev. 2 and Comp. 1, while reducing it on others. These results show that limited target labels can be valuable, but a predetermined adaptation strategy may also introduce negative transfer when applied uniformly. Observation 2: Adaptation effects are class-dependent. The effect of adaptation also varies across sensing classes within the same target domain. Fig. 1(b) reports the class-wise recall changes from the source model to the 1-shot ProtoNet on two cross-device target domains. Under the same adaptation strategy, some classes improve substantially, while others experience considerable degradation. A single domain-level decision is therefore too coarse to capture heterogeneous classlevel transfer behavior. Deployment-time adaptation should instead diagnose individual classes and allow them to use different adaptation actions.

(a) High Confident Errors

(b) Prompt Scale

Fig. 2. Raw confidence and raw CSI are unsuitable for deployment-time LLM planning. (a) On target domains, source-model predictions with confidence above 0.9 still contain many errors, showing that confidence alone is an unreliable control signal. (b) Directly serializing CSI samples into an LLM prompt leads to a rapidly growing token scale, motivating compact sensinggrounded evidence before LLM planning.

Observation 3. LLM planning requires structured sensing evidence. A natural question is whether readily available deployment signals can be used directly for LLM planning. We examine two intuitive choices: source-model confidence and raw CSI data. As shown in Fig. 2(a), among source predictions with confidence above 0.9, 49.6% to 69.2% are still incorrect. Confidence alone is therefore insufficient for diagnosing transfer failures or selecting adaptation actions. Meanwhile, Fig. 2(b) shows that serializing one CSI sample already requires approximately 112K tokens, while a targetdomain query set requires 530M to 1.4B tokens. Raw CSI is

thus impractical as direct LLM input and does not explicitly expose the class-level behavior needed for adaptation. Motivation. These observations lead to three design requirements for CSI-Agent. First, adaptation should account for target-specific effects, since the same few-shot strategy may improve one target domain while degrading another. Importantly, this does not imply that a strong domain-level adaptation method is broadly ineffective. It may perform well overall while remaining suboptimal for individual classes. Second, adaptation should operate at the class level, as different sensing classes under the same domain shift may require different treatments. Third, LLM planning should be grounded in structured sensing evidence rather than relying solely on readily available deployment inputs, such as raw CSI and source-model confidence. III. R ELATED W ORK A. Cross-Domain Wi-Fi Sensing Cross-domain Wi-Fi sensing systems aim to maintain sensing performance under changes in users, environments, and transceiver configurations [9]. Early studies primarily improve generalization by learning domain-invariant representations or constructing more robust physical features. CrossSense trains a roaming model to generate synthetic samples for new environments, thereby reducing target-site data collection costs [26]. EI employs adversarial learning to remove environment- and user-specific information from activity representations [6], while Widar 3.0 derives a body-coordinate velocity profile from multiple Wi-Fi links for gesture recognition [27]. WiHF derives a domain-independent motion-change pattern for realtime gesture recognition and user identification [28]. Recently, more studies reduce target-domain labeling costs through one-shot or few-shot adaptation. TOSS performs semisupervised domain adaptation by jointly leveraging a small number of labeled target samples and abundant unlabeled target observations [29]. Wi-Learner extracts Doppler features and adapts to a new deployment using one labeled example per gesture [18]. FewSense adapts a pretrained feature extractor using a few target-domain samples and performs similaritybased recognition [15]. MetaFormer further combines metalearning with Transformer representations to adapt using only one labeled target sample per class [30]. OneSense employs an aug-meta learning framework to enable scalable few-shot learning for gesture recognition systems [31]. Although these methods require target-domain supervision, they generally employ a predetermined adaptation procedure after deployment. B. LLM-Based Sensing Recent studies have explored large language models for wireless sensing and signal processing. Wi-Chat incorporates physical knowledge of Wi-Fi sensing into prompts and uses an LLM to perform zero-shot activity recognition [32]. AutoIoT translates natural-language user requirements into interpretable programs for AIoT applications and iteratively improves the generated programs with limited user intervention [33]. SignalLLM adopts an agentic approach to decompose general

signal-processing tasks, construct processing workflows, and invoke external tools [34]. Our work advances these directions by reframing crossdomain adaptation from a predetermined procedure into an evidence-driven and verifiable decision process. CSI-Agent grounds an LLM planner in structured target-domain diagnostics to determine whether each class should retain a strong default or use a specialized sensing action, while deterministic tools handle CSI processing, prediction, verification, and bounded execution.

an LLM agent proposes an action for each class, while a deterministic verifier rejects unreliable proposals. Finally, a bounded executor applies the accepted actions relative to a common support-weighted default and produces the targetdomain predictions. Source Domain

S tar = {(xsi , yis )}Ck i=1 ,

(1)

with k samples per class, which may be collected through a brief enrollment or calibration procedure. The system also observes an unlabeled target set: tar

Q

Nq = {xqj }j=1 ,

(2)

whose ground-truth labels need to be inferred during deployment. Unlike the small episodic query sets used in conventional few-shot learning, Qtar covers every unlabeled sample in the target domain, so Nq is typically much larger than Ck. Instead of learning a single adapted model, we formulate deployment-time adaptation as a decision problem. The system selects a target-domain predictor h from a constrained family H using only the information above. The ideal predictor maximizes target-domain macro-F1:  h∗ = arg max MacroF1 {yjq }, {h(xqj )} , (3) h∈H

where {yjq } denotes

the hidden query labels used only for offline evaluation. During deployment, CSI-Agent selects h using only labeled support behavior and unlabeled query statistics. In this paper, CSI-Agent instantiates H through classconditioned combinations of deterministic sensing actions. V. S YSTEM D ESIGN

A. Overview Fig. 3 presents the overview of CSI-Agent. Given a source sensing model, a small labeled target support set, and an unlabeled target query set, CSI-Agent performs deploymenttime adaptation through four stages. It first constructs a constrained toolbox of complementary deterministic sensing actions. It then summarizes their class-level behavior into sensing-grounded diagnostic cards. Based on this evidence,

Diagnostic Evidence

Source Action 𝑎!"#

Support Evidence

Source Source Dataset 𝒟 !"# Model 𝑓$

Full View 𝑎'())

Target Evidence

Target Domain

Lowpass View 𝑎)*+,&!!

IV. P ROBLEM F ORMULATION In this paper, we consider closed-set Wi-Fi CSI sensing under cross-domain deployment. Let Y = {1, . . . , C} denote the known label space for a sensing task. A labeled source dataset Dsrc is available before deployment and is used to train a source sensing model fθ . At deployment time, fθ is applied to a target domain dtar whose CSI distribution may differ from the source because of new users, environments, or devices, while the label space Y stays the same. For each target domain, the system observes a small labeled support set as follows:

Sensing Action Toolbox

Support Set 𝒮 %&"

Unlabeled Set 𝒬 %&"

Support Fit

Risk Flags

Support-Weighted 𝑎12

Action Fit

Verdict

Outputs

Action Execution

Final Prediction 𝑦) -

Selected Action 𝑎#

Auditable Record

Action Probability 𝑃&!

LLM Rationale Verifier Decision

Diagnostic Card

Quarter-Band View 𝑎-.&/0

Planning and Verification LLM Agent Planner {status, failure_mode, action, reason}

Class Score Vector 𝓏#

Verifier

Fig. 3. Overview of the CSI-Agent framework.

B. Sensing Action Toolbox CSI-Agent uses a constrained set of deterministic sensing actions as follows: A = {asrc , afull , alowpass , aqband , aSW }.

(4)

We use the term action to denote a fixed sensing scorer rather than a free-form program. Each action follows a predefined execution rule and produces a class-probability vector. Source action. The source action directly applies the frozen source classifier fθ to the original CSI sample: Psrc (x) = softmax(fθ (x)).

(5)

This action remains effective when the source representation transfers well to the target domain, but it may degrade substantially under severe domain shifts. View-specific recovery actions. To capture different features, CSI-Agent uses three amplitude-based CSI views V ∈ {full, lowpass, qband}, where v ∈ V indexes a CSI view. The full view retains the original CSI representation. The lowpass view applies a short temporal moving average to suppress unstable high-frequency variations. The quarter-band view, denoted by qband, retains a contiguous quarter of the subcarriers and masks the remaining ones, thereby providing a localized spectral view. Let Tv denote the input transformation associated with view v, and let gθ denote the feature extractor of the source-trained model. The view-specific representation is hv (x) = gθ (Tv (x)).

(6)

For each view v ∈ V and class c, CSI-Agent computes source prototypes as follows: X 1 µvc = src hv (xsrc (7) i ), |Dc | src src src (xi ,yi )∈Dc

where Dcsrc contains the source samples from class c. Then, CSI-Agent estimates a coarse target-domain feature shift without query labels: X 1 X 1 ∆v = hv (x) − src hv (x), (8) |T | |D | src x∈T

x∈D

where T = S ∪ Q denotes the available target-domain batch. The shifted source prototype is µ evc = µvc + ∆v . This shared shift provides coarse alignment for global differences between the source and target feature distributions. However, a single global shift may not align all classes equally well, while estimating reliable class-specific shifts from only a few labeled samples is difficult. CSI-Agent therefore evaluates the resulting scorers at the class level and uses this evidence for subsequent action selection. CSI-Agent uses the shifted prototypes µ evc in two complementary ways. First, it converts the distances between labeled support samples and shifted prototypes into supportside diagnostic probabilities:  exp −∥hv (xsi ) − µ evc ∥22 /τ S s , (9) Pv (xi , c) = P s evc′ ∥22 /τ ) c′ ∈Y exp (−∥hv (xi ) − µ where τ > 0 is a temperature parameter, set to 32 in all experiments. A higher PvS (xsi , c) indicates that the viewspecific representation hv (xsi ) is closer to the shifted prototype of class c relative to the competing class prototypes. These probabilities are used only to evaluate the reliability of view v on the labeled support set. Second, to obtain query predictions while preserving sourcedomain class structure, CSI-Agent fits a ridge classifier. Let ϕ(u) = [u⊤ , 1]⊤ denote the bias-augmented version of feature vector u. For each view v, the classifier is obtained by Wv∗ = arg min W

X

ϕ(hv (xsi ))⊤ W − e⊤ yis

(xsi ,yis )∈S

+λ

C X

2 2

2

2 ϕ(e µvc )⊤ W − e⊤ c 2 + α∥W ∥F ,

c=1

(10) where W maps the augmented feature representation to C class scores, ec ∈ RC is the one-hot vector of class c, λ ≥ 0 controls the prototype anchoring strength, and P α > 20 controls the ridge regularization, with ∥W ∥2F = i,j Wij . In this paper, we set λ = 1 and α = 1. The first term fits the classifier to the limited labeled target samples, while the second term retains the class structure represented by the shifted source prototypes. For an unlabeled query sample xqj ∈ Q, the viewspecific probability vector is  PvQ (xqj ) = softmax ϕ(hv (xqj ))⊤ Wv∗ . (11) Support-weighted action. To reduce reliance on any single CSI view, CSI-Agent combines the three view-specific scorers according to their reliability on the labeled support set. For s each view v, let ybv,i = arg maxc∈Y PvS (xsi , c) denote the

predicted label of support sample xsi . The support reliability of view v is defined as   s rv = max MacroF1 {yis }Ck (12) yv,i }Ck i=1 , {b i=1 , ϵ , where ϵ > 0 is a small constant that prevents a view from receiving an exactly zero weight. The normalized weight of view v is wv = P rv ru . For an unlabeled query sample xqj , u∈V the support-weighted probability vector is X Q wv PvQ (xqj ). (13) PSW (xqj ) = v∈V

The support-weighted action provides a target-adaptive and robust default. It emphasizes views that are more reliable on the labeled target samples while retaining complementary evidence from the other views. C. Sensing-Grounded Diagnostic Evidence For each known class c ∈ Y, CSI-Agent collects diagnostic evidence that summarizes how each action behaves on the labeled support set and the unlabeled query set. Support evidence. For the positive support samples with yis = c, CSI-Agent measures the class-wise recall, the mean probability assigned to class c, and the mean rank of class c. These quantities characterize how reliably an action recognizes the class. For the negative support samples with yis ̸= c, CSIAgent measures the false-activation rate and computes a onevs-rest margin between the average class scores on positive and negative samples. Together, these measurements capture both recognition of class c and interference from competing classes. Unlabeled target evidence. Since the query labels are unavailable, CSI-Agent summarizes only the prediction behavior induced by each action. For action a ∈ A and class c ∈ Y, the primary statistic is the query activation rate:  Nq  1 X Q q ′ QRatea,c = I arg max P (x , c ) = c . (14) a j c′ ∈Y Nq j=1 This quantity represents the fraction of unlabeled target samples assigned to class c by action a. CSI-Agent compares QRatea,c with the activation rate of the support-weighted default, denoted by QRateSW,c , to determine how the candidate action changes the target-domain prediction pattern. In this paper, we use ρc = C1 as a coarse uniform reference. This reference makes the weak assumption that the target classes occur at roughly similar frequencies. It is used only to flag severe under-activation or over-activation, rather than to estimate the true target-domain class distribution, and is never used alone to accept an action. Diagnostic card. CSI-Agent organizes the support evidence and query activation statistics into a structured diagnostic card for each class and candidate action. The card contains four fields: • Support Fit, which indicates whether the action improves, preserves, or worsens class separation on the labeled support samples.

Activation Fit, which indicates how the action changes the target-domain activation of the class relative to the default scorer and the coarse uniform reference. • Risk Flags, which identify increased false activation, aggravated under-activation or over-activation, and conflicts between support and unlabeled target evidence. • Verdict, which summarizes the action as viable, weak, or risky before LLM planning.

•

D. LLM Agent Planning and Verification Given the diagnostic cards constructed in the previous stage, CSI-Agent uses an LLM to propose a sensing action ac ∈ A for each class c. A deterministic verifier then evaluates each proposed action before execution. LLM agent planning. For each class c ∈ Y, the LLM agent receives its diagnostic card together with a set of candidate actions. Although a general-purpose LLM may encode knowledge about wireless sensing, it is not directly grounded in high-dimensional, non-linguistic CSI measurements. CSIAgent therefore does not expose raw CSI to the LLM or ask it to produce sample-level predictions. Instead, the agent diagnoses the observed transfer behavior and proposes a deterministic scorer based on the sensing-grounded evidence. The planner returns a structured output as follows: oc = {status, failure mode, action, reason}.

(15)

The status indicates whether class c appears reliably transferred or requires intervention. The failure mode summarizes whether the class is under-activated, over-activated, confused with competing classes, or poorly separated. The action field records the proposed action ac , while the reason provides an auditable explanation grounded in the diagnostic card. Sensing-grounded verification. The LLM agent provides the primary diagnosis and action proposal. Since the available target-domain evidence is limited, CSI-Agent applies a lightweight deterministic verifier before execution. The verifier checks whether the proposed action is consistent with the diagnosed failure mode and whether it provides sufficient evidence of improvement over the support-weighted default without introducing substantial false activation or degrading support-side class separation. If these conditions are not satisfied, the proposal is rejected and the class retains the support-weighted action. The verifier decision is represented by gc ∈ {0, 1}, where gc = 1 accepts the proposed action ac and gc = 0 activates the default fallback. Candidate filtering limits clearly implausible options before planning, whereas verification evaluates the final LLM proposal before execution. The action ac and gate gc are then passed to the bounded executor. E. Bounded Action Execution Following verification, CSI-Agent applies each accepted action as a bounded adjustment to the common supportweighted default. For query sample xqj and class c, the final class score is h i Q Q zc (xqj ) = PSW (xqj , c)+βgc PaQc (xqj , c) − PSW (xqj , c) , (16)

TABLE I DATASET STATISTICS AND CROSS - DOMAIN INFORMATION . # D OMAINS DENOTES THE NUMBER OF TARGET DOMAINS . Nsrc AND Ntar ARE THE SOURCE - TRAINING AND RAW TARGET- SPLIT SAMPLE COUNTS . Dataset CSI-Bench CSI-Bench WiAR SignFi WiMANS

Shift Type

# Classes

# Domains

Nsrc /Ntar

Cross-Device Compositional Cross-User Cross-Env Cross-Env

5 5 16 50 6

2 2 6 1 2

21,473 / 9,631 21,473 / 18,759 1,279 / 2,400 1,400 / 500 1,317 / 3,762

where β ∈ [0, 1] controls the intervention strength and is set to 0.35. A rejected action retains the support-weighted score, while an accepted action modifies it without necessarily replacing the default completely. Then, the final prediction is ybjq = arg maxc∈Y zc (xqj ). This bounded execution preserves a common scoring reference across classes while limiting negative transfer from uncertain action proposals. VI. E XPERIMENTAL E VALUATION A. Experiment Setup All experiments are conducted on a Linux server with an Intel(R) Xeon(R) Gold 6258R CPU and NVIDIA A100 GPUs with 40GB of memory. All source models are optimized with AdamW at a learning rate of 3 × 10−4 for at most 200 epochs, with early stopping (patience 10) after five epochs. SimCLR uses 200 pretraining and linear-probe epochs. EI and TOSS use learning rates of 10−4 and 5 × 10−5 , respectively. For the main CSI-Agent experiments, the LLM is Qwen3.5-9B. Datasets and splits. We evaluate CSI-Agent on four public Wi-Fi CSI datasets using five cross-domain splits that cover device, user, environment, and compositional shifts, as summarized in Table I. For CSI-Bench [24], we follow the official Cross-Device split and the Compositional split containing the held-out user–environment pairs U02/E01 and U05/E05. For WiAR [35], users a–d form the source domain, while users e, f, 7, 8, 9, and 10 define six target domains. For SignFi [8], we retain the first 50 gesture classes from the original dataset. The lab_150 data form the source domain, while the home_276 data form the target domain. Finally, we use the provided WiMANS [36] Cross-Env split, whose target domains are an empty room and a meeting room. Input size. In this paper, all methods use CSI amplitude only, without phase information. Each sample is represented as a single-channel amplitude map with 500 temporal samples and 112 CSI dimensions for CSI-Bench, WiAR, and SignFi50, and 270 CSI dimensions for WiMANS. Baselines. We compare three categories of cross-domain baselines. Source-side representation learning includes supervised ResNet-18 [25], SimCLR pretraining with a linear probe [10], and EI [6]. These methods improve transferability through supervised, self-supervised, or domain-invariant representation learning without using labeled target data. Unlabeled target-domain adaptation includes DANN [13] and CORAL [37], which align source and target feature distributions using unlabeled target observations. Few-shot target-

TABLE II S OURCE AND ONE - SHOT TARGET PERFORMANCE ( DOMAIN - AVERAGE M ACRO -F1). T HE BEST RESULTS ARE HIGHLIGHTED IN BOLD .

No Labeled Target Support

Split ↓

One-Shot Target Adaptation

ResNet

SimCLR

DANN

CORAL

EI

PTN

FT

TOSS

FewSense

CSI-AGENT

CSI-Bench

Source Device Comp.

0.9810 0.2535 0.2577

0.8539 0.3403 0.3242

0.9802 0.2416 0.2725

0.9845 0.2712 0.2560

0.8596 0.2535 0.2332

0.9503 0.2879 0.2383

0.9608 0.2702 0.2184

0.9069 0.2275 0.2462

0.9618 0.2824 0.2238

0.9196 0.3464 0.3431

WiAR

Source User

0.8037 0.0418

0.7687 0.0334

0.7795 0.0460

0.8040 0.0484

0.7433 0.0443

0.8567 0.1796

0.8417 0.2702

0.7594 0.0882

0.7030 0.2885

0.8595 0.4006

SignFi

Source Env.

0.5984 0.0131

0.7963 0.0079

0.6247 0.0033

0.5863 0.0115

0.6284 0.0066

0.8105 0.2926

0.7994 0.2642

0.6824 0.0036

0.6208 0.2873

0.8780 0.7977

WiMANS

Source Env.

0.9029 0.2687

0.8195 0.2327

0.9137 0.2873

0.9195 0.2821

0.9086 0.3316

0.8503 0.4118

0.8616 0.3896

0.8796 0.4725

0.8736 0.5345

0.8948 0.5571

0.1670

0.1877

0.1701

0.1738

0.1738

0.2820

0.2825

0.2076

0.3233

0.4890

Target Average

domain adaptation methods use the same k-shot support set as CSI-Agent. PTN [16] performs classic prototype-based classification, while Fine-Tune (FT) updates the source model using the labeled target samples. FewSense [15] adapts the learned representation for similarity-based recognition. TOSS [29] further combines limited labeled target samples with pseudolabeled target observations for semi-supervised adaptation. B. Cross-Domain Performance Table II reports the source-domain and one-shot targetdomain performance. Although most methods achieve strong performance on the source domain, their accuracy often drops substantially after deployment. For example, CORAL achieves a Macro-F1 of 0.9845 on the CSI-Bench source domain but only 0.2712 under the Cross-Device split. A similar degradation is observed on WiAR, SignFi, and WiMANS, confirming that strong source-domain recognition does not necessarily translate into robust cross-domain sensing. The relative performance of existing methods also varies considerably across target domains. SimCLR is the strongest baseline on both CSI-Bench splits, whereas FewSense performs best on WiAR and WiMANS, and PTN achieves the highest baseline result on SignFi. Some methods exhibit even more pronounced target dependence. For example, TOSS reaches 0.4725 on WiMANS but only 0.0036 on SignFi. These results are consistent with our preliminary analysis that no fixed adaptation strategy is uniformly effective across different deployment shifts. CSI-Agent achieves the best target-domain performance on all five evaluation splits. On CSI-Bench, it consistently outperforms the strong SimCLR baseline under both device and compositional shifts. The improvements are more substantial under the challenging cross-user and cross-environment settings, where CSI-Agent surpasses the strongest baseline by 0.1121 on WiAR, 0.5051 on SignFi, and 0.0226 on WiMANS. Averaged across the five target splits, CSI-Agent achieves a Macro-F1 of 0.4890, improving upon the strongest baseline method, FewSense, by 0.1657. These consistent results demon-

Source +4.0 -6.0 -5.8 -4.1 -50.2-19.6-37.0-27.6-19.0-52.3-78.4-30.3-24.0 Full +1.3+0.2 -3.1 -4.0 +1.7+1.5 -0.8 -3.6 +1.7+2.4 -6.3 -5.2 +3.1 Low-pass +0.5 -5.1 -8.6 -1.7 +0.0+2.5 -1.3 +5.1 -1.1 +5.1+0.7-12.1-11.5 QBand -3.2 +2.9 -1.5 -2.1 -5.8 -4.2 -4.7 -2.8 +0.0 -1.2 -6.9 -6.4 -6.2 Domain* +0.0+0.6+0.0+0.0+0.0+0.5+0.0 -2.1 +0.4+3.6 -2.1 +0.0+3.3 CSI-Agent +0.9+0.5+0.2 -0.2 +0.1+0.8+0.5+0.4+0.6+2.0+0.6+0.3+3.1 v. p. WiAR SignFiWiMANS nch De h Com CSI-Be CSI-Benc

30 15 0

Macro-F1 Gain Over SW (%)

Dataset ↓

15 30

Fig. 4. Performance impact of fixed sensing actions and agentic action selection across target domains under 1-shot adaptation. Domain* denotes a domain-level variant of CSI-Agent that selects a single action for all classes in each target domain. Each cell reports the Macro-F1 change in percentage points relative to the support-weighted action, which is omitted from the figure. Positive values indicate improvement, whereas negative values indicate performance degradation.

strate the effectiveness of evidence-driven deployment-time adaptation across diverse domain shifts. C. Selective Action Analysis Target-dependent effects of sensing actions. To examine whether the target dependence observed in Section II-B extends to the actual action space of CSI-Agent, Fig. 4 compares each candidate action against the support-weighted default, which provides a strong target-adaptive reference. The performance impact varies substantially across target domains. For example, the source action improves one CSI-Bench device target by 4% but causes severe degradation on SignFi, while the Low-pass action improves some WiAR targets but reduces performance on both WiMANS targets. At the same time, Full and Low-pass provide positive gains on several domains and are generally more stable than reverting to the source classifier, indicating that the toolbox contains useful recovery options. However, no candidate action is consistently beneficial, and retaining the support-weighted default remains preferable in several domains. These results show that target-

1

1-shot 3-shot

5-shot 7-shot

Dev. -0.4 -0.5 +0.0 -0.3 Comp. +0.5 +0.4 +0.0 -1.8 WiAR -2.8 -2.8 -2.3 +0.8 0 1 SignFi +0.7 -0.8 -1.3 -0.5 2 WiMANS -12.6 -12.5 -0.4 -1.4

Dev. Comp. WiAR SignFWi iMANS

LM ery ifier und w/o L w/o Qwu /o Verw/o Bo

Macro-F1 (%)

Macro-F1 (%)

100 80 60 40 20 0

3

Fig. 5. Support-size sensitivity of Fig. 6. Ablation of CSI-Agent under 1CSI-Agent. Each bar reports Macro- shot Adaptation. Each cell reports the F1 with different k labeled target Macro-F1 change relative to the full samples per class. Dev. and Comp. CSI-Agent. Dev. and Comp. denote the denote the CSI-Bench Cross-Device CSI-Bench Cross-Device and Compoand Compositional splits. sitional splits.

dependent effects are not specific to a single few-shot adapter and motivate deployment-time selection. Effectiveness of class-level action selection. The bottom two rows of Fig. 4 compare domain-level and class-level planning under the same support-weighted default. Domain* selects one action for all classes in a target domain, whereas CSI-Agent allows different classes to retain the default or invoke different recovery actions. Domain-level planning provides gains on some targets but remains neutral on many others and causes noticeable degradation in two cases. In contrast, CSI-Agent improves 12 of the 13 target domains, with its only degradation limited to 0.2 percentage points. CSI-Agent does not always match the best fixed action on every individual domain, since it selects actions without query labels and applies conservative class-level corrections rather than optimizing domain-level test performance. Nevertheless, it is substantially more consistent than any fixed action. Moreover, CSI-Agent produces positive gains on several targets where none of the individual candidate actions improves over the support-weighted default, showing that class-specific action combinations can recover complementary benefits that are unavailable to global action selection. D. Support-Size Sensitivity Analysis In practical deployment, k represents the number of labeled target samples collected per sensing class during enrollment. A smaller k reduces the calibration and annotation burden, which is particularly important when target-domain data are difficult to collect. Fig. 5 evaluates CSI-Agent with k ∈ {1, 3, 5, 7} labeled samples. Performance improves consistently across all five evaluation cases as additional labeled support becomes available. The largest gains generally occur from 1-shot to 3-shot, particularly on WiAR and WiMANS, showing that even a small amount of additional target supervision can substantially improve adaptation performance. The improvements become more gradual at larger support sizes, with SignFi approaching saturation. This trend indicates that CSI-Agent makes effective use of scarce supervision while exhibiting the expected diminishing returns as the support set becomes more representative of the target domain. Importantly, no evaluation split shows unstable or non-monotonic behavior as k increases, suggesting that the sensing actions and diagnostic evidence

TABLE III P ERFORMANCE OF CSI-AGENT WITH DIFFERENT LLM PLANNERS UNDER THE SAME ADAPTATION PROTOCOL (M ACRO -F1, %). D EV. AND C OMP. DENOTE THE CSI-B ENCH C ROSS -D EVICE AND C OMPOSITIONAL SPLITS . Planner

k

Dev.

Comp.

WiAR

SignFi

WiMANS

Qwen3.5-9B Llama-3.1-8B GPT-5.4-mini Claude Haiku 4.5

1 1 1 1

34.64 34.70 34.40 33.96

34.31 34.16 34.44 34.44

40.06 39.89 40.03 40.47

79.77 79.77 79.41 79.77

55.71 55.65 54.55 55.50

Qwen3.5-9B Llama-3.1-8B GPT-5.4-mini Claude Haiku 4.5

5 5 5 5

48.10 48.15 48.15 48.10

38.75 38.75 38.75 38.60

63.67 63.46 63.71 63.78

95.90 95.45 95.05 95.45

67.44 67.44 67.59 67.63

benefit consistently from additional support. Overall, CSIAgent remains effective in the low-cost 1-shot setting while continuing to benefit from additional labeled target samples with diminishing marginal gains. E. Ablation Study Fig. 6 evaluates the contribution of the planning, diagnostic, and safety components compared to the full CSI-Agent under the 1-shot setting. The w/o LLM variant replaces the LLM planner with a deterministic support-only rule while retaining the verifier and bounded executor. Its average performance decreases to 45.96% (∆ = −2.93), with substantial degradation on WiAR and WiMANS. This result shows that selecting actions solely from sparse support-side statistics is insufficient under heterogeneous target shifts. The w/o Query variant retains the LLM planner but removes the diagnostic statistics derived from unlabeled query samples. Its average MacroF1 drops to 45.65% (∆ = −3.25), including a pronounced reduction on WiMANS (∆ = −12.55). Together, these results show that reliable action planning requires both query-aware sensing evidence and joint reasoning over class-level diagnostic signals, rather than a fixed support-driven rule. The verifier and bounded executor provide smaller but important safety benefits. Removing the verifier has little effect on the two CSI-Bench splits but degrades WiAR (∆ = −2.3) and SignFi (∆ = −1.3), indicating that evidence-based proposals can still be unsafe under particular shifts. Removing bounded execution causes clear degradation on CSI-Bench Compositional (∆ = −1.8) and WiMANS (∆ = −1.4). These results show that verification and bounded residual execution complement the planner by limiting the impact of unreliable interventions when they occur. F. LLM Planner Robustness Having established that a simple support-driven rule is insufficient, we next examine whether CSI-Agent depends on a particular LLM backbone. We instantiate the planner with Qwen3.5-9B [38], Llama-3.1-8B [39], GPT-5.4-mini [40], and Claude Haiku 4.5 [41], covering both open-weight and proprietary models under the same prompt, evidence, and execution protocol. Table III reports their performance under the 1-shot and 5-shot settings. The average Macro-F1 ranges only from 48.57% to 48.90% at 1-shot and from 62.65% to 62.77%

TABLE IV D EPLOYMENT OVERHEAD PER TARGET DOMAIN UNDER 1- SHOT ADAPTATION . D EV. AND C OMP. DENOTE THE CSI-B ENCH C ROSS -D EVICE AND C OMPOSITIONAL SPLITS .

Configuration ↓

Dev.

TABLE V R EPRESENTATIVE EVIDENCE - GROUNDED DECISION TRACES UNDER 1- SHOT ADAPTATION . E VIDENCE COMPARES THE SUPPORT- WEIGHTED DEFAULT WITH THE PROPOSED ACTION , WHILE Rc REPORTS THE POST- HOC RECALL CHANGE OF THE FINAL BOUNDED OUTPUT.

Comp. WiAR SignFi WiMANS

Deployment latency (s/domain) Source Inference 1.28 2.40 0.14 0.37 Toolbox + Evidence 8.84 16.91 0.83 1.54 CSI-AGENT 20.59 28.60 33.84 91.02 Sample-Level LLM 7,573.69 15,642.78 612.37 802.78 Input prompt tokens (103 /domain) CSI-AGENT 4.81 4.85 15.19 47.67 Sample-Level LLM 1,643.99 3,210.18 132.45 156.38

0.99 7.84 21.36 2,752.70

5.90 643.24

at 5-shot, with no single planner consistently dominating across all target splits. These results show that CSI-Agent is robust across the tested LLM planners. The planner makes class-level decisions over structured diagnostic evidence and a small, constrained action space, rather than performing sample-level CSI recognition. Consequently, once an LLM is capable of following the planning protocol and interpreting the diagnostic signals, greater general model capacity does not necessarily translate into better sensing performance. This design allows different capable LLMs to make comparably effective deployment decisions without tying CSI-Agent to a specific backbone. G. Overhead Analysis CSI-Agent uses the LLM only for class-level deployment planning rather than invoking it separately for individual query samples. Table IV quantifies the resulting latency and input-token cost. Source inference first processes the target split, while the toolbox constructs the candidate scorers and sensing-grounded diagnostic evidence. The complete CSIAgent pipeline additionally includes LLM planning, verification, and bounded execution. The deterministic toolbox and evidence construction introduce moderate overhead beyond source inference. Including class-level LLM planning, the complete pipeline requires 20.59–91.02 s per target domain. The highest cost occurs on SignFi, which contains 50 sensing classes and consequently produces the largest class-level prompt at 47.67K input tokens. This reflects the design of CSI-Agent, whose LLM input scales primarily with the number of classes and candidate actions rather than with the number of individual CSI samples. We further estimate the cost of a counterfactual samplelevel LLM design by constructing compact diagnostic cards for individual queries. Even with these compressed cards instead of raw CSI, sequential sample-level reasoning would require 612.37–15,642.78 s and 132.45K–3.21M input tokens per target domain. Compared with this estimate, CSI-Agent reduces latency by 8.8×–547.0× and input-token cost by 3.3×– 661.9×. These results demonstrate that sensing-grounded aggregation makes LLM-assisted adaptation substantially more scalable than sample-level reasoning.

Case

Decision-Time Evidence

Condensed Rationale

Decision and Outcome

CSI-Bench E05 C0

p: .136 → .474; Rank: 2 → 1; FA: .25 → .25

QBand improves support Accept bounded probability and rank QBand correction. without increasing false Rc : 0 → 1 activation. (∆Rc = +100 pp)

WiAR user-9 C7

q: .047 → .056; p: .314 → .367; Rank: 2 → 1; FA: .400 → .333

Low-pass moves query Accept bounded activation toward the Low-pass correction. uniform reference, Rc : 0 → 1 improves support fit, and (∆Rc = +100 pp) reduces false activation.

SignFi home 276 C0

q: .020 → .020; p: .062 → .089; Rank: 4 → 1; FA: .102 → .020

Low-pass preserves Accept bounded query activation near the Low-pass correction. uniform reference, Rc : 0 → 1 improves support fit, and (∆Rc = +100 pp) reduces false activation. Note. q, p, Rank, and FA denote query class rate, support probability, support rank, and false activation, respectively.

H. Illustrative Decision Traces Table V links decision-time evidence with the post-hoc outcome for the same sensing class. The three selected cases illustrate successful non-default interventions under device, user, and environment shifts. QBand recovers a class under the CSI-Bench Cross-Device split, while Low-pass recovers classes under the WiAR cross-user and SignFi crossenvironment settings. The final outcomes are computed using hidden query labels only after the action plan and execution results have been fixed. These labels are unavailable to the diagnostic card, LLM planner, verifier, and bounded executor. The examples therefore illustrate that evidence-grounded decisions can translate into actual class-level recovery rather than providing plausible rationales alone. The reported rationales are condensed from the LLM outputs and serve as auditable decision references rather than faithful causal explanations. VII. C ONCLUSION This paper presented CSI-Agent, an evidence-seeking LLM agent for deployment-time adaptation in cross-domain Wi-Fi CSI sensing. Instead of applying one fixed adaptation strategy to an entire target domain, CSI-Agent uses compact, sensinggrounded evidence to diagnose class-specific transfer behavior and determine when specialized recovery actions are justified. Comprehensive experiments show that CSI-Agent consistently improves target-domain performance across various domain shifts. The results further demonstrate the benefits of classlevel action selection, query-aware diagnostic evidence, and constrained planning, while showing robustness across different LLM backbones and avoiding costly sample-level LLM inference. These findings suggest that LLM agents can support practical cross-domain sensing when their decisions are grounded in structured sensing evidence and constrained by reliable execution mechanisms.

R EFERENCES [1] M. Li and W. Wang, “Uni-Fi: Integrated multi-task wi-fi sensing,” arXiv preprint arXiv:2601.10980, 2026. [2] T. Zhao, N. Wang, G. Cao, S. Mao, and X. Wang, “Functional data analysis assisted cross-domain wi-fi sensing using few-shot learning,” in ICC 2024-IEEE International Conference on Communications. IEEE, 2024, pp. 4780–4785. [3] X. Wang, C. Yang, and S. Mao, “PhaseBeat: Exploiting csi phase data for vital sign monitoring with commodity wifi devices,” in 2017 IEEE 37th international conference on distributed computing systems (ICDCS). IEEE, 2017, pp. 1230–1239. [4] Y. He, J. Liu, M. Li, G. Yu, J. Han, and K. Ren, “Sencom: Integrated sensing and communication with practical wifi,” in Proceedings of the 29th Annual International Conference on Mobile Computing and Networking, 2023, pp. 1–16. [5] Y. Ma, G. Zhou, and S. Wang, “Wifi sensing with channel state information: A survey,” ACM Computing Surveys (CSUR), vol. 52, no. 3, pp. 1–36, 2019. [6] W. Jiang, C. Miao, F. Ma, S. Yao, Y. Wang, Y. Yuan, H. Xue, C. Song, X. Ma, D. Koutsonikolas et al., “Towards environment independent device free human activity recognition,” in Proceedings of the 24th annual international conference on mobile computing and networking, 2018, pp. 289–304. [7] H. Wang, D. Zhang, Y. Wang, J. Ma, Y. Wang, and S. Li, “RT-Fall: A real-time and contactless fall detection system with commodity wifi devices,” IEEE Transactions on Mobile Computing, vol. 16, no. 2, pp. 511–526, 2016. [8] Y. Ma, G. Zhou, S. Wang, H. Zhao, and W. Jung, “SignFi: Sign language recognition using wifi,” Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 2, no. 1, pp. 1–21, 2018. [9] C. Chen, G. Zhou, and Y. Lin, “Cross-domain wifi sensing with channel state information: A survey,” ACM Computing Surveys, vol. 55, no. 11, pp. 1–37, 2023. [10] T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International conference on machine learning. PmLR, 2020, pp. 1597–1607. [11] Z. Hong, Z. Li, S. Zhong, W. Lyu, H. Wang, Y. Ding, T. He, and D. Zhang, “Crosshar: Generalizing cross-dataset human activity recognition via hierarchical self-supervised pretraining,” Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 8, no. 2, pp. 1–26, 2024. [12] A. Farahani, S. Voghoei, K. Rasheed, and H. R. Arabnia, “A brief review of domain adaptation,” Advances in data science and information engineering: proceedings from ICDATA 2020 and IKE 2020, pp. 877– 894, 2021. [13] Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. March, and V. Lempitsky, “Domain-adversarial training of neural networks,” Journal of machine learning research, vol. 17, no. 59, pp. 1–35, 2016. [14] T. Zhao, X. Wang, and S. Mao, “Cross-domain, scalable, and interpretable rf device fingerprinting,” in IEEE INFOCOM 2024-IEEE Conference on Computer Communications. IEEE, 2024, pp. 2099– 2108. [15] G. Yin, J. Zhang, G. Shen, and Y. Chen, “Fewsense, towards a scalable and cross-domain wi-fi sensing system using few-shot learning,” IEEE Transactions on Mobile Computing, vol. 23, no. 1, pp. 453–468, 2022. [16] J. Snell, K. Swersky, and R. Zemel, “Prototypical networks for few-shot learning,” Advances in neural information processing systems, vol. 30, 2017. [17] F. Wang, T. Zhang, W. Xi, H. Ding, G. Wang, D. Zhang, Y. Cui, F. Liu, J. Han, J. Xu et al., “A survey on wi-fi sensing generalizability: Taxonomy, techniques, datasets, and future research prospects,” IEEE Communications Surveys & Tutorials, 2026. [18] C. Feng, N. Wang, Y. Jiang, X. Zheng, K. Li, Z. Wang, and X. Chen, “Wi-learner: Towards one-shot learning for cross-domain wi-fi based gesture recognition,” Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 6, no. 3, pp. 1–27, 2022. [19] Z. Wang, J. Li, W. Wang, Z. Dong, Q. Zhang, and Y. Guo, “Review of few-shot learning application in csi human sensing,” Artificial Intelligence Review, vol. 57, no. 8, p. 195, 2024. [20] J. Luo, W. Zhang, Y. Yuan, Y. Zhao, J. Yang, Y. Gu, B. Wu, B. Chen, Z. Qiao, Q. Long et al., “Large language model agent: A

survey on methodology, applications and challenges,” arXiv preprint arXiv:2503.21460, 2025. [21] S. Yao, J. Zhao, D. Yu, I. Shafran, K. R. Narasimhan, and Y. Cao, “React: Synergizing reasoning and acting in language models,” in NeurIPS 2022 Foundation Models for Decision Making Workshop, 2022. [22] T. Schick, J. Dwivedi-Yu, R. Dessı̀, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” Advances in neural information processing systems, vol. 36, pp. 68 539–68 551, 2023. [23] K. Huang, X. Yin, H. Huang, and W. Gao, “Modality plug-and-play: Runtime modality adaptation in llm-driven autonomous mobile systems,” in Proceedings of the 31st Annual International Conference on Mobile Computing and Networking, 2025, pp. 543–558. [24] G. Zhu, Y. Hu, W. Gao, W.-H. Wang, B. Wang, and K. Liu, “CSI-Bench: A large-scale in-the-wild dataset for multi-task wifi sensing,” Advances in Neural Information Processing Systems, vol. 38, 2026. [25] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778. [26] J. Zhang, Z. Tang, M. Li, D. Fang, P. Nurmi, and Z. Wang, “CrossSense: Towards cross-site and large-scale wifi sensing,” in Proceedings of the 24th annual international conference on mobile computing and networking, 2018, pp. 305–320. [27] Y. Zhang, Y. Zheng, K. Qian, G. Zhang, Y. Liu, C. Wu, and Z. Yang, “Widar3. 0: Zero-effort cross-domain gesture recognition with wifi,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 11, pp. 8671–8688, 2021. [28] C. Li, M. Liu, and Z. Cao, “WiHF: Enable user identified gesture recognition with WiFi,” in IEEE INFOCOM 2020-IEEE Conference on Computer Communications. IEEE, 2020, pp. 586–595. [29] Z. Zhou, F. Wang, J. Yu, J. Ren, Z. Wang, and W. Gong, “Targetoriented semi-supervised domain adaptation for wifi-based har,” in IEEE INFOCOM 2022-IEEE Conference on Computer Communications. IEEE, 2022, pp. 420–429. [30] B. Sheng, R. Han, F. Xiao, Z. Guo, and L. Gui, “MetaFormer: Domainadaptive wifi sensing with only one labelled target sample,” Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 8, no. 1, pp. 1–27, 2024. [31] L. Zhao, R. Xiao, J. Liu, and J. Han, “One is enough: Enabling one-shot device-free gesture recognition with COTS WiFi,” in IEEE INFOCOM 2024-IEEE Conference on Computer Communications. IEEE, 2024, pp. 1231–1240. [32] H. Zhang, Y. Ren, H. Yuan, J. Zhang, and Y. Shen, “Wi-chat: Large language model powered wi-fi sensing,” arXiv preprint arXiv:2502.12421, 2025. [33] L. Shen, Q. Yang, Y. Zheng, and M. Li, “Autoiot: Llm-driven automated natural language programming for aiot applications,” in Proceedings of the 31st Annual International Conference on Mobile Computing and Networking, 2025, pp. 468–482. [34] J. Ke, Q. Hu, S. Yuan, Y. Xu, and J. Yang, “SignalLLM: A generalpurpose llm agent framework for automated signal processing,” arXiv preprint arXiv:2509.17197, 2025. [35] L. Guo, L. Wang, C. Lin, J. Liu, B. Lu, J. Fang, Z. Liu, Z. Shan, J. Yang, and S. Guo, “Wiar: A public dataset for wifi-based activity recognition,” IEEE Access, vol. 7, pp. 154 935–154 945, 2019. [36] S. Huang, K. Li, D. You, Y. Chen, A. Lin, S. Liu, X. Li, and J. A. McCann, “WiMANS: A benchmark dataset for wifi-based multiuser activity sensing,” in European Conference on Computer Vision. Springer, 2024, pp. 72–91. [37] B. Sun, J. Feng, and K. Saenko, “Correlation alignment for unsupervised domain adaptation,” in Domain adaptation in computer vision applications. Springer, 2017, pp. 153–171. [38] Qwen Team, “Qwen3.5: Towards native multimodal agents,” Feb. 2026. [Online]. Available: https://qwen.ai/blog?id=qwen3.5 [39] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan et al., “The llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024. [40] OpenAI, “Introducing gpt-5.4 mini and nano,” Mar. 2026. [Online]. Available: https://openai.com/index/ introducing-gpt-5-4-mini-and-nano/ [41] Anthropic, “Claude haiku 4.5 system card,” Oct. 2025. [Online]. Available: https://www.anthropic.com/claude-haiku-4-5-system-card

Record · ID 1108668 · SHA-256 6ad2a41f5fa2525d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.