1
Active Inference-Enabled Agentic Closed-Loop ISAC with Long-Horizon Planning Guangjin Pan* , Zhuojun Tian† , Mehdi Bennis‡ , Henk Wymeersch* Department of Electrical Engineering, Chalmers University of Technology, Sweden † Division of Information Science and Engineering, KTH Royal Institute of Technology, Sweden ‡ Centre of Wireless Communications, University of Oulu, Finland
arXiv:2604.19599v1 [eess.SP] 21 Apr 2026
*
Abstract—Wireless agentic systems enable agents to autonomously perceive, reason, and act. However, existing works neglect the tight coupling between sensing and control in closedloop integrated sensing and communication (ISAC) systems. In this paper, we propose an active inference (AIF)-driven wireless agentic system for closed-loop ISAC, which jointly optimizes control and sensing resource allocation via backward–forward message passing on a factor graph. The AIF agent maintains a generative model as a digital twin by integrating a localization model for uncertainty-aware state inference and a localization channel knowledge map (CKM) for approximating observation quality during planning. Simulation results demonstrate that the AIF-enabled agent adaptively allocates sensing resources based on spatially varying channel conditions, achieving superior balance among tracking accuracy, control effort, and sensing resource consumption over baseline strategies. Index Terms—Closed-loop ISAC, active inference, factor graph, message passing, wireless agentic systems
I. I NTRODUCTION The emergence of wireless agentic systems, where intelligent agents autonomously perceive, reason, and act upon the physical environment, is reshaping the design of nextgeneration wireless systems [1]. A prominent application of this paradigm lies in integrated sensing and communication (ISAC) [2], where base stations (BSs) sense targets such as UAVs [3], vehicles [4], or robots [5], and simultaneously deliver control commands to guide their behavior. In such closed-loop ISAC systems, the wireless agentic system forms a perception–cognition–action loop: the BS perceives the target’s state through channel measurements, the intelligent agent within the BS infers the state of the target, plans future actions, and the target executes the resulting control decisions. This paradigm enables the wireless network to evolve from a passive communication infrastructure into an autonomous decision-making system that actively interacts with and shapes its physical environment. A fundamental challenge in closed-loop ISAC is the tight coupling between sensing and control. On the one hand, the quality of state estimation depends on the allocated sensing resources, which directly affects control performance. On the other hand, control decisions determine the target’s future This work was supported in part by the SNS JU project 6G-DISAC under the EU’s Horizon Europe research and innovation Program under Grant Agreement No. 101139130, the Swedish Foundation for Strategic Research (SSF) (grant FUS21-0004, SAICOM), the ERANET CHIST-ERA Project MUSE-COM2, and in part by the Research Council of Finland (former Academy of Finland) Project Vision-Guided Wireless Communication. The computations were enabled by resources provided by the National Academic Infrastructure for Supercomputing in Sweden (NAISS), partially funded by the Swedish Research Council through grant agreement No. 2022-06725.
trajectory, influencing the channel conditions and thus the future sensing quality. However, existing works largely neglect this coupling. Plenty of research focuses on ISAC resource allocation aimed at improving sensing performance [6] or balancing sensing and communication [7], [8]. Other studies incorporate sensing results into closed-loop control, but typically assume fixed sensing quality, thereby ignoring the impact of sensing resource allocation on control performance [9], [10]. Furthermore, wireless agentic systems require a principled framework that unifies perception, reasoning, and action selection under uncertainty. Reinforcement learning (RL)based approaches can learn joint policies, but they operate as black-box optimizers that lack interpretability and do not provide explicit uncertainty quantification over the agent’s internal state [11]. This makes it difficult to understand why a particular sensing configuration is chosen or how state uncertainty drives resource allocation decisions. Active inference (AIF), rooted in the free energy principle from neuroscience [12], offers an interpretable and uncertainty-aware framework for autonomous agents, where state inference and action planning emerge from minimizing a unified free-energy objective that naturally balances goal achievement and uncertainty reduction. In this paper, we propose an AIF-driven wireless agentic system for closedloop ISAC that jointly optimizes control and sensing resource allocation via message passing on a factor graph. The AIFdriven agent maintains a generative model as a digital twin, incorporating a state-transition model, a pretrained localization model for inference, and a localization channel knowledge map for approximating observation quality during planning. Compared with our preliminary work [3], we make two key improvements. First, we replace the Cramér-Rao lower bound (CRLB)-based observation covariance approximation with learned neural network models. The CRLB provides only a theoretical lower bound on localization error, which can be loose in practical scenarios with complex non-linear channelto-position mappings, whereas the learned models capture the actual localization performance. Second, we extend the singlestep myopic planning to long-horizon backward–forward message passing, enabling uncertainty-aware decision-making that anticipates the impact of current sensing and control actions on future performance. II. S YSTEM MODEL As illustrated in Fig. 1, we consider a wireless agentic system for closed-loop sensing and control of a mobile entity (ME) (e.g., a UAV, robot, or autonomous vehicle). The system
2
resource kt : allocating more subcarriers increases the effective bandwidth and reduces uncertainty Σt . The observation model (3) is used in both the inference and planning stages, but the methods of obtaining Σt are different as discussed below. 1) Inference Stage: Localization Model: During inference, the ISAC BS allocates kt subcarriers to acquire the channel frequency response Ht , which depends on the ME’s state st and the sensing configuration kt . A pretrained localization model then maps the channel measurement to the observation and its covariance:
Wireless Agentic System AIF-Driven Agent Generative Model (DT)
Action
Inference VFE Minimization
yt , Σ̂t
q(st)
Planning EFE Minimization
lx , ly , k
Loc Model ∗ kt+1
Σ̂t (k) Loc CKM
Ht Measurement
ISAC BS
u∗ t
loc yt , Σ̂loc t = Fθ (Ht ).
Environment Mobile Entity
Fig. 1. AIF-enabled wireless agentic system for closed-loop ISAC.
comprises an ISAC BS and an AIF-driven agent. The ISAC BS serves as the sensory interface, acquiring channel measurements and extracting position observations via a localization model. The AIF-driven agent serves as the cognitive core, maintaining an internal generative model that acts as a digital twin (DT) of the physical environment. At each time slot, the AIF-driven agent performs two operations: inference, which updates the belief over the ME’s state by minimizing the variational free energy (VFE), and planning, which jointly optimizes the control input and sensing resource allocation by minimizing the expected free energy (EFE). The resulting actions are twofold: the control command is transmitted to the ME, and the sensing configuration is applied to the ISAC BS for the next measurement slot. This forms a closed-loop cycle driven by active inference. A. State Transition Model We consider the ME moving in a horizontal twodimensional plane. The state at time t is defined as st ≜ [lx,t , ly,t , vx,t , vy,t ]⊤ , where lx,t and ly,t denote the position coordinates, and vx,t and vy,t denote the velocity components. With a sampling interval ∆t, the state evolves according to the linear Gaussian model: p(st+1 | st , ut ) = N (st+1 ; Ast + But , Q),
(1)
2
where ut ∈ R is the acceleration command applied to the ME, Q is the process noise covariance, corresponding to dynamics: 1 2 I2 ∆t I2 ∆t I2 A= , B= 2 . (2) 0 I2 ∆t I2 B. Sensing and Observation Model The ME’s state is partially observed through wireless channel measurements collected at the ISAC BS. At each time slot, the BS allocates kt ∈ K subcarriers for localization, where K denotes the set of candidate subcarrier configurations. Each subcarrier has a spacing of Bsc . Therefore, the effective sensing bandwidth at time slot t is kt Bsc . We model the observation as the estimated position: p(yt | st , kt ) = N (yt ; Cst , Σt )
(3)
where C = [I2 , 02×2 ] extracts the 2D position from st , yt ∈ R2 is the position observation, and Σt ∈ R2×2 is the observation covariance. The covariance depends on the sensing
(4)
Since the true channel Ht is available at time t, both yt and loc Σ̂loc t are obtained directly from the measurements, where Σ̂t serves as the estimate of Σt in (3). 2) Planning Stage: Localization CKM Model: During planning, the AIF-driven agent evaluates the observation quality at future positions under candidate sensing resource allocations, for which no real channel measurements exist. To bridge this gap, inspired by digital twin techniques, we proposed to incorporate a localization channel knowledge map (CKM) in the generative model, which predicts the expected observation covariance from position and sensing resource: Σ̂ckm (lx , ly , k) = Fϕckm (lx , ly , k),
(5)
where (lx , ly ) is a queried position, k is the candidate number of subcarriers, and ϕ denotes the model parameters. The predicted Σ̂ckm (lx , ly , k) serves as a surrogate for Σt in (3), enabling the planning stage to approximate the observation model at future time steps. Since the planning stage needs to anticipate the localization accuracy of Fθloc at future positions, Fϕckm is trained on the outputs of Fθloc to capture its positionand subcarrier-dependent performance characteristics. C. Action Model and Objective The AIF-driven agent jointly determines the control input and the sensing resource allocation at each time slot. The action is at = (ut , kt+1 ), where ut drives the ME’s state evolution through (1), and kt+1 configures the ISAC BS for the next observation through (3). The agent selects at to balance control performance, stateestimation accuracy, and sensing-resource consumption via the following cost terms: desired • Observation cost: Given the desired trajectory st 1 desired desired est with yt = Cst , Jt = 2 (yt − ytdesired )⊤ Qgoal (yt − ytdesired ), where Qgoal ⪰ 0 is a weighting matrix. ctrl • Control cost: Jt = 12 u⊤ t Rgoal ut , where Rgoal ⪰ 0 penalizes the control energy to prevent excessive maneuvering. sens • Sensing-resource cost: Jt = 12 αgoal kt2 , where αgoal ≥ 0 controls the tradeoff between sensing accuracy and resource consumption. We consider long-term planning over a horizon of T steps. The overall cost from the current time t to the horizon end is t+T X−1 sens J= (Jτctrl + Jτest (6) +1 + Jτ +1 ). τ =t
3
⑩
D. Problem Formulation via Active Inference In active inference, the cost (6) defines the agent’s goal prior, a distribution encoding preferences over future observations and actions. The generative model maintained by the AIF-driven agent can be factorized as
⑨
× P̃ (yt+1:t+T , ut:t+T −1 , kt+1:t+T ), (7) Qt−1 where Pg,1 = τ +1 | sτ +1 , kτ +1 ) τ =1 p(sτ +1 | sτ , uτ ) p(yQ t+T −1 is the past generative factor, Pg,2 = p(sτ +1 | τ =t sτ , uτ ) p(yτ +1 | sτ +1 , kτ +1 ) characterizes the future state dyQt+T −1 namics and observations, and P̃ = τ =t p̃(yτ +1 , uτ , kτ +1 ) is the goal prior with per-step factor 1 sens p̃(yτ +1 , uτ , kτ +1 ) = exp(−Jτctrl − Jτest (8) +1 − Jτ +1 ). Z The AIF-driven agent operates through two stages: Inference stage. At time t, minimize VFE to update the belief: F1 = Eq(s1:t |y1:t )[log q(s1:t )−logPg,1 (s1:t , y1:t | a1:t−1 )]. (9) Planning stage. Minimize the EFE over the variational distributions q(ut:t+T −1 ) and q(kt+1:t+T ): " t+T −1 X Gt = Eqt log q(sτ +1 , yτ +1 , uτ , kτ +1 | sτ ) τ =t
− log Pg,2 (sτ +1 , yτ +1 | sτ , uτ , kτ +1 ) # − log p̃(yτ +1 , uτ , kτ +1 ) ,
(10)
where qt ≜ q(st+1:t+T , yt+1:t+T , ut:t+T −1 , kt+1:t+T | st ) is the variational predictive distribution over future states, observations and actions. The optimal actions can be obtained as the modes of the marginal variational factors, i.e., u∗τ = arg max q(uτ ) and kτ∗+1 = arg max q(kτ +1 ). III. ACTIVE I NFERENCE -BASED J OINT S ENSING AND C ONTROL We solve (9) and (10) via message passing on the factor graph. At each time step t, the system operates as follows: (i) the ISAC BS collects Ht using kt subcarriers; (ii) Fθloc extracts (yt , Σ̂t ) and the AIF-driven agent updates the belief q(st ) via VFE minimization; (iii) backward and forward message passing yield the beliefs q(ut ) and q(kt+1 ), where Fϕckm provides Σ̂(k) for planning; (iv) u∗t is transmitted to the ME ∗ and kt+1 configures the BS for the next slot. In the following, we detail the inference and planning procedures, and describe the training of the localization model and CKM model. A. Inference Stage The inference stage minimizes the VFE F1 in (9) to obtain the posterior belief over st . Under the linear-Gaussian model, this admits a closed-form solution equivalent to Kalman filtering. At time t, the ISAC BS collects Ht (st , kt ) and obtains (yt , Σ̂t ) via (4). Given the belief q(st−1 ) = s N (st−1 ; m̂st−1 , P̂t−1 ) and control ut−1 , the prediction step s s s gives m̂t|t−1 = Am̂st−1 +But−1 and P̂t|t−1 = AP̂t−1 A⊤ +Q. Combining the transition and observation messages yields the
⑧
⑦ ⑨
⑥
⑥ ⑤④ ③ ① ②
⑧
Pg (s1:t+T , y1:t+T | a1:t+T −1 ) = Pg,1 (s1:t , y1:t | a1:t−1 ) × Pg,2 (st+1:t+T , yt+1:t+T | st , at:t+T −1 )
⑦
(a)
①
⑤④ ③ ① ②
⑩
④
③ ②
(b)
⑤⑦ ⑧ ⑥ ⑨ (c)
Fig. 2. Factor graph for the planning stage: (a) backward pass for computing backward beliefs, (b) forward pass for control u∗τ , and (c) forward pass for sensing resource allocation kτ∗+1 .
posterior q(st ) = N (st ; m̂st , P̂ts ) with s −1 P̂ts = ((P̂t|t−1 )−1 + C ⊤ (Σ̂loc C)−1 , t )
(11)
−1 s m̂st = P̂ts (C ⊤ (Σ̂loc yt + (P̂t|t−1 )−1 m̂st|t−1 ). t )
(12)
Here, ut−1 and kt are determined by the planning stage at time t−1, Coupling planning and inference in a closed loop. B. Planning Stage The planning stage minimizes the EFE in (10) to determine the control uτ and sensing allocation kτ +1 over a planning horizon from t to t + T . We solve this via belief propagation on the factor graph, which is decomposed into a backward pass (encoding future preferences) and a forward pass (extracting optimal actions). The factor graph is illustrated in Fig. 2. The goal prior in (8) factorizes into −1 pu (uτ ) = N (uτ ; 0, Rgoal ),
py (yτ +1 ) = N (yτ +1 ; yτdesired , Q−1 +1 goal ), −1 pk (kτ +1 ) = N (kτ +1 ; 0, αgoal ).
(13) (14) (15)
In the planning stage, the observation covariance is replaced by the predicted covariance Σ̂ckm (lx , ly , k) from (5). For brevity, desired desired we write Σ̂ckm (kτ ) ≜ Fϕckm (lx,τ , ly,τ , kτ ), where the desired desired position (lx,τ , ly,τ ) is taken from sdesired as a proxy for τ the real position at time τ . 1) Backward Pass: As shown in Fig. 2 (a), the backward message from the observation factor to sτ +1 is obtained by marginalizing yτ +1 and kτ +1 over the observation likelihood, observation goal prior, and sensing cost prior (combining 1 – 4 to obtain 5 ): ˆ X µ 5 (sτ +1 )∝ pk(kτ +1 ) N (yτ +1 ; Csτ +1 , Σ̂ckm (kτ +1 ))) kτ +1∈K
× N (yτ +1 ; yτdesired , Q−1 +1 goal ) dyτ +1 .
(16)
For each kτ +1 , the integral over yτ +1 yields ckm N (yτdesired ; Csτ +1 , Λk ) with Λk ≜ Q−1 (kτ +1 )). +1 goal + Σ̂ To extend to the full state dimension, we introduce a diffuse prior σ 2 ≫ 1 on the velocity components, desired giving N (sτ +1 ; µ̄, Σ̄k ) with µ̄ = [ yτ +1 ] and 0 Σ̄k = diag(Λk , σ 2 I2 ). Approximating the mixture over K by moment matching yields µ 5 (sτ +1 ) ∝ N (sτ +1 ; µ̄, Σ̄∗ ) with
4
P P w Λ 0 Σ̄∗ = [ k∈K0 k k σ2 I ] and wk = pk (k)/ k′ ∈K pk (k ′ ). 2 The observation message 5 is then combined with s,bw the backward belief, i.e., 6 , N (sτ +1 ; m̂s,bw τ +1 , P̂τ +1 ), propagated from time τ + 2, yielding µ 7 (sτ +1 ) ∝
s,tot s,tot s,bw −1 −1 ∗ −1 N (sτ +1 ; m̂s,tot τ +1 , P̂τ +1 ) with P̂τ +1 = ( Σ̄ ) +(P̂τ +1 ) ) s,tot −1 s,bw ∗ −1 µ̄+(P̂τs,bw and m̂s,tot +1 ) m̂τ +1). Marginalizing τ +1 = P̂τ +1 ( Σ̄ ) sτ +1 and uτ through the transition factor and control prior based on 7 and 9 , we obtain the backward belief µ (sτ ) ∝ N (sτ ; m̂s,bw , P̂τs,bw ) with τ
10
P̂τs,bw = (A⊤ Sτ−1 A)−1 and m̂s,bw = P̂τs,bw A⊤ Sτ−1 m̂s,tot τ τ +1 , s,tot −1 ⊤ where Sτ = P̂τ +1 + Q + BRgoal B . At the terminal time τ = t + T , the backward belief is initialized as uninformative: s,bw 2 2 m̂s,bw t+T = 0 and P̂t+T = σ0 I with σ0 ≫ 1. The recursion then proceeds from τ = t + T down to τ = t. 2) Forward Pass: The forward pass starts from the posterior belief q(st ) = N (st ; m̂st , P̂ts ) obtained in the inference stage. At each step τ , the control uτ is first optimized, and then the sensing allocation kτ +1 is determined given u∗τ . Since the system operates in a receding-horizon fashion, only the ∗ first-step actions (u∗t , kt+1 ) are executed, and the planning is repeated at the next time slot with updated belief. Optimal control: As shown in Fig. 2 (b), the forward mess,f sage at sτ is µ 8 (sτ ) ∝ q(sτ ) = N (sτ ; m̂s,f τ , P̂τ ), where q(st ) is initialized from the inference-stage posterior (11)–(12) and propagated forward for τ > t. The backward message s,tot s,tot 7 at sτ +1 is N (sτ +1 ; m̂τ +1 , P̂τ +1 ). Fusing 8 and 7 through the transition factor by marginalizing sτ +1 yields the message 10 toward uτ . Multiplying 10 with the control −1 prior 9 , pu (uτ ) = N (uτ ; 0, Rgoal ), the posterior over uτ u u is q(uτ ) = N (uτ ; m̂τ , Pτ ), where P̂τs,pred = AP̂τs,f A⊤ + Q, −1 ⊤ −1 u , and Dτ = P̂τs,pred + P̂τs,tot +1 ,Pτ = (B Dτ B + Rgoal ) s,tot s,f u ⊤ −1 u hatmτ = Pτ B Dτ (m̂τ +1 − Am̂τ ). The optimal control is the mode of q(uτ ), i.e., u∗τ = m̂uτ . Sensing resource allocation: As shown in Fig. 2 (c), given u∗τ , we first fuse the forward messages 1 – 3 , which encode the state estimate and control action, with the backward message 4 , which encodes future planning information, to s,fuse obtain µ 5 (sτ +1 ) ∝ N (sτ +1 ; m̂s,fuse τ +1 , P̂τ +1 ) with s,pred −1 −1 −1 P̂τs,fuse ) + (P̂τs,bw ) , +1 ) +1 = ((P̂τ
(17)
s,fuse s,pred −1 pred −1 s,bw m̂s,fuse ) ŝτ +1 +(P̂τs,bw m̂τ +1 ), (18) +1 ) τ +1 = P̂τ +1 ((P̂τ pred where ŝτ +1 = Am̂s,f + Bu∗τ . Passing 5 through the τ
observation factor and the observation goal prior 6 by marginalizing over sτ +1 and yτ +1 yields µ 8 (kτ +1 ) ∝
ckm desired , Vk ), where Vk = Q−1 (kτ +1 )+ N(C m̂s,fuse τ +1 ; yτ +1 goal +Σ̂ s,fuse ⊤ C P̂τ +1 C . Multiplying 8 with the sensing cost prior 9 , the posterior over kτ +1 is q(kτ +1 ) ∝ pk (kτ +1 )·µ 8 (kτ +1 ), and the optimal sensing allocation is kτ∗+1 = arg maxk∈K q(kτ +1 ), evaluated by enumeration since |K| is small. C. Model Training The proposed algorithm relies on two pre-trained components: the localization model Fθloc used in the inference stage, and the covariance prediction model Fϕckm used in the planning stage. Both are trained offline prior to deployment.
1) Localization Model: The localization model Fθloc employs a ResNet-34 [13] backbone as a shared feature extractor. The extracted features are fed into a position head and variance head, each consisting of a two-layer MLP (64 hidden units, 2 outputs). Given the channel Ht , the position head outputs the location estimate [ˆlx,t , ˆly,t ], while the variance head outputs loc,2 loc,2 loc,2 loc,2 [σ̂x,t , σ̂y,t ], yielding Σ̂loc t = diag(σ̂x,t , σ̂y,t ). Training proceeds in two stages. In Stage 1, the backbone and position head are trained by minimizing the mean squared error (MSE) over the 2D position using data from all k ∈ K. In Stage 2, the backbone is frozen and the |K| variance heads are trained independently by minimizing the negative log-likelihood (NLL) [14]: Lvar =
N batch X X (ld,n − ˆld,n )2 1 loc,2 (log σ̂d,n + ) loc,2 2Nbatch n=1 σ̂d,n d∈{x,y}
where Nbatch is the batch size. We train a separate variance head for each k rather than a single shared head, because different subcarrier configurations lead to different noise levels in the position observations. A shared head would be forced to average over these heterogeneous noise distributions, resulting in variance estimates that are neither accurate for small k (underestimating uncertainty) nor for large k (overestimating uncertainty). Separate heads allow each to specialize in the noise regime of its corresponding k, resulting in well-calibrated covariance estimates across all sensing configurations. 2) Localization CKM Model: During planning, future channels are unavailable, where Fθloc cannot be used to obtain Σ(k). To solve this problem, we train a lightweight MLP Fϕckm (·) with two hidden layers of 128 units each and a 2-dimensional output layer, which predicts the observation covariance from position and sensing configuration: [σ̂xckm,2 , σ̂yckm,2 ] = Fϕckm (lx , ly , k), ckm
(19)
diag(σ̂xckm,2 , σ̂yckm,2 ).
To generate with Σ̂ (lx , ly , k) = training data, we randomly sample positions (lx , ly ) and subcarrier configurations k ∈ K, simulate the corresponding channels, and apply Fθloc to obtain the predicted variance σ̂d2 for each sample. The MLP is then trained by minimizing the MSE between its output and the variance predicted by Fθloc : N batch X X 1 ckm,2 loc,2 2 Lckm = (log σ̂d,n − log σ̂d,n ) . (20) 2Nbatch n=1 d∈{x,y}
IV. N UMERICAL R ESULTS We consider the Sionna [15] street canyon scene. The BS is at (−32, 12, 35) with a 16-element ULA at halfwavelength spacing, operating at fc = 3 GHz with a transmit power of 23 dBm and noise power spectral density of −174 dBm/Hz. The ME has a single isotropic antenna at 1.5 m height. The set of candidate subcarrier configurations is K = {50, 100, 150, 200, 250, 300, 350, 400}, corresponding to effective sensing bandwidths from 12.5 MHz to 100 MHz. The training dataset consists of 20,000 randomly sampled ME positions generated via Sionna ray tracing with up to 3 interaction bounces and diffraction enabled. The dataset is split into 15,000 samples for Stage 1 (location training) and 5,000 samples for Stage 2 (variance training), with non-
Zone 2
40
20
Zone 1
0
X (m)
20
40
400 350 300 250 200 150 100 50
3.0 2.5 2.0 1.5 1.0 0.5
Fig. 3. AIF trajectory colored by the selected k, overlaid on the observation variance map. AIF allocates more subcarriers in high-variance regions (Zone 1) and fewer in low-variance regions (Zone 2).
overlapping splits to prevent overfitting. Both models are trained with the Adam optimizer at a learning rate of 10−4 for 200 epochs. For the closed-loop simulation, the ME follows a straight-line reference trajectory from (−50, 0) m to (50, 0) m at a constant velocity of vref = 1 m/s, yielding N = 1000 time steps with ∆t = 0.1 s. The process noise covariance 2 is Q = σw I4 with σw = 0.01. AIF goal-prior parameters are Qgoal = diag(0.5, 0.5), Rgoal = diag(0.1, 0.1), and αgoal = 1 × 10−6 . The planning horizon is T = 10 unless otherwise stated. All results are averaged over 10 independent runs with different random seeds. Fig. 3 shows the trajectory overlaid on the localization variance map (for k = 200 subcarriers). The trajectory color encodes the subcarrier count k selected by AIF at each time step. In Zone 1, where the predicted observation variance is high, AIF allocates more subcarriers to reduce localization uncertainty. In Zone 2, where the variance is low, AIF reduces the sensing resources. This confirms that the proposed method adaptively balances sensing accuracy and resource consumption based on the spatially varying channel quality. Table I compares the AIF method against two baselines over 10 planning steps, averaged over 10 independent runs: (i) AIF control with sensing resources sampled from the prior pk (k) (prior-k + AIF-u), and (ii) AIF sensing allocation with greedy control that drives the next-step position to the desired target (AIF-k + greedy-u). The proposed AIF achieves the lowest total cost and the lowest average EFE, confirming that jointly optimizing control and sensing via AIF yields the best overall performance. Compared with prior-k + AIF-u, AIF reduces the observation cost by 47% through adaptive sensing allocation. The greedy-u baseline achieves the smallest observation cost but incurs an excessive control cost, as it ignores the control penalty entirely. Fig. 4 shows the average AIF cost versus planning horizon for all three methods. The proposed AIF achieves the lowest cost across all horizons, dropping by 85% from horizon 1 to 10 and saturating thereafter. The prior-k + AIF-u baseline follows a similar trend but converges to a higher cost, confirming the benefit of adaptive sensing allocation. AIF-k + greedyu improves only marginally with longer horizons, as the greedy control limits the agent’s ability to exploit long-horizon planning. These results confirm that jointly optimizing sensing and control with long-horizon planning is essential for closedloop performance. V. C ONCLUSION In this paper, we propose an AIF-driven wireless agentic system for closed-loop ISAC that jointly optimizes control
TABLE I C OST C OMPARISON OF D IFFERENT S TRATEGIES Method
J est
J ctrl
J sens
J
Ave. EFE
AIF (proposed) Prior-k + AIF-u AIF-k + Greedy-u
110.01 206.55 82.12
1.61 3.70 1357.73
39.05 32.77 37.97
150.67 243.02 1477.82
−1.453 -0.695 -0.079
Average AIF Cost
Desired trajectory BS Start End
Observation variance tr ( (k)) (k=200)
12.5 10.0 7.5 5.0 2.5 0.0 2.5 5.0
Number of Selected Subcarriers
Y (m)
5
1.8 AIF (proposed) Prior-k + AIF-u AIF-k + Greedy-u
1.2 0.6 0
0 1
5
10 15 Planning horizon length
20
Fig. 4. Average AIF cost under different planning horizon lengths.
and sensing resource allocation. The AIF agent incorporates a pretrained localization model for uncertainty-aware state inference and a CKM-based covariance prediction model for long-horizon planning. Simulation results confirm that the proposed method adaptively allocates sensing resources based on spatially varying channel conditions, achieving superior balance among tracking accuracy, control effort, and sensing resource consumption over baseline strategies. R EFERENCES [1] K. Dev et al., “Advanced architectures integrated with agentic AI for next-generation wireless networks,” IEEE Communications Standards Magazine, 2025. [2] G. Pan et al., “Semantic communication for rate-limited closedloop distributed communication-sensing-control systems,” arXiv preprint arXiv:2512.19177, 2025. [3] G. Pan et al., “Active inference framework for closed-loop sensing, communication, and control in UAV systems,” arXiv preprint arXiv:2509.14201, 2025. [4] J. Wang et al., “Networking and communications in autonomous driving: A survey,” IEEE Commun. Surveys Tuts., vol. 21, no. 2, pp. 1243–1274, 2019. [5] W. Wu et al., “Goal-oriented semantic communications for robotic waypoint transmission: The value and age of information approach,” IEEE Trans. Wireless Commun., vol. 23, no. 12, pp. 18 903–18 915, 2024. [6] B. Li et al., “Maximizing the value of service provisioning in multiuser ISAC systems through fairness guaranteed collaborative resource allocation,” IEEE J. Sel. Areas Commun., vol. 42, no. 9, pp. 2243–2258, 2024. [7] A. Khalili et al., “Efficient uav hovering, resource allocation, and trajectory design for ISAC with limited backhaul capacity,” IEEE Trans. Wireless Commun., vol. 23, no. 11, pp. 17 635–17 650, 2024. [8] J. Zou et al., “Energy-efficient beamforming design for integrated sensing and communications systems,” IEEE Trans. Commun., vol. 72, no. 6, pp. 3766–3782, 2024. [9] P. Duan et al., “Distributed cooperative LQR design for multi-input linear systems,” IEEE Trans. Control Netw. Syst., vol. 10, no. 2, pp. 680–692, 2022. [10] H. Jin et al., “Co-design of sensing, communications, and control for low-altitude wireless networks,” IEEE Trans. Mobile Comput., no. 01, pp. 1–13, 2025. [11] O. Lockwood et al., “A review of uncertainty for deep reinforcement learning,” in Proc. AAAI‘2022, vol. 18, no. 1, 2022, pp. 155–162. [12] K. Friston, “The free-energy principle: a unified brain theory?” Nature reviews neuroscience, vol. 11, no. 2, pp. 127–138, 2010. [13] G. Pan et al., “AI-driven wireless positioning: Fundamentals, standards, state-of-the-art, and challenges,” IEEE Commun. Surveys Tuts., vol. 28, pp. 4394–4428, 2026. [14] J. Pinto et al., “An uncertainty-aware performance measure for multiobject tracking,” IEEE Signal Processing Letters, vol. 28, pp. 1689– 1693, 2021. [15] “Sionna,” https://nvlabs.github.io/sionna/.