A STRA: Asynchronous Age-Aware Satellite Random Access via Mean-Field Control
arXiv:2605.18282v1 [cs.NI] 18 May 2026
Sayam Chakraborty∗† , Aimin Li∗ , Yiğit İnce∗ , Sajjad Baghaee∗ , Elif Uysal∗ , Fellow, IEEE ∗ Communication Networks Research Group (CNG), EE Dept, METU, Ankara, Turkiye † Dept of Avionics, Indian Institute of Space Science and Technology, Trivandrum, India E-mail: [email protected]; {aimin, yigit.ince, uelif}@metu.edu.tr; [email protected] Abstract—Satellite Internet-of-Things (IoT) enables massive status-update services beyond terrestrial coverage, but grantfree uplink access creates a coupled freshness-control problem: increasing repetition and receiver-side diversity improves a device’s capture-SIC opportunities, yet the resulting population congestion degrades network-wide freshness. Existing AoI-aware random-access models often rely on slot-synchronous collisions, fixed delivery probabilities, or scalar transmit-or-wait decisions and therefore cannot capture asynchronous satellite uplinks with capture and SIC. This paper develops a PHY-aware mean-field framework, termed A STRA (Asynchronous Age-Aware Satellite Random Access), for freshness-driven satellite IoT random access. We build an access model that captures asynchronous arrivals, partial overlaps, capture, and SIC while preserving the dependence of delivery success on each device’s repetitiondiversity action. We then formulate the population interaction as a scalable mean-field MDP in which devices optimize access timing and intensity using only local AoI observations. The resulting system admits a mean-field equilibrium in which individual optimality and endogenous congestion are mutually consistent. We further prove that the optimal equilibrium policy admits an age-threshold structure. Numerical results show that the proposed policy reduces AoI relative to age-independent baselines. Index Terms—Satellite IoT, Random Access, Age of Information, Mean-Field Games
I. I NTRODUCTION Satellite Internet-of-Things (IoT) is becoming a key connectivity option for global monitoring and machine-type communication where terrestrial infrastructure is unavailable or uneconomical [2], [3]. Large populations of low-power ground devices sporadically generate short status updates and access the satellite uplink without centralized scheduling. This grantfree paradigm avoids excessive signaling overhead, but poses a control problem: each device must decide when and how aggressively to transmit while sharing a medium whose congestion is generated endogenously by the population’s own access decisions. Since satellite IoT devices are typically energy-constrained, aggressive replication must be balanced against both freshness and energy expenditure. Detailed proofs and additional results can be found in [1]. Aimin Li contributed equally to this work. This work was supported by the European Union (through ERC Advanced Grant 101122990-GO SPACE-ERC-2023AdG). Yiğit İnce was also supported by Turk Telekom within the framework of the 5G and Beyond Joint Graduate Support Programme, coordinated by the Information and Communication Technologies Authority. Views and opinions expressed are those of the authors only and do not necessarily reflect those of the funding agencies.
Gateway
User 1
User 2
User 𝑁
Fig. 1: Asynchronous satellite IoT uplink with capture-SIC. In each frame (Tf , M slots), N devices transmit d replicas over R resource pools.
Classical satellite random-access designs mainly target throughput and reliability. Slotted ALOHA, Contention Resolution Diversity Slotted ALOHA (CRDSA), Irregular Repetition Slotted ALOHA (IRSA), and coded random-access schemes improve performance through packet repetition and successive interference cancellation (SIC) [4]–[6]. More recent work has extended this line by combining IRSA with power-domain and multi-receiver diversity. In particular, NonOrthogonal Multiple Access (NOMA)-based IRSA uses discrete received-power levels to resolve collisions through signal-to-interference-plus-noise ratio (SINR)-based capture and SIC [7]. Energy-efficient IRSA variants exploit per-replica power diversity to improve both spectral and energy efficiency under SIC decoding [8]. Multi-satellite NOMA-IRSA further shows that additional satellite receivers can reduce packet loss and improve energy efficiency by providing receiver diversity [9]. Together, these studies highlight the importance of replica-level design, capture-SIC, and receiver diversity in satellite IoT random access. However, their design objectives including packet loss, throughput, asymptotic load thresholds, spectral efficiency, and energy efficiency remain incomplete for status-update traffic, in which even a reliably delivered packet may have limited value if it is stale. To capture this limitation, Age of Information (AoI) has emerged as a freshness metric that quantifies the timeliness of the most recently received update [10], [11]. A central
insight from AoI theory is that freshness optimization differs fundamentally from delay or throughput optimization: in some regimes, deliberate waiting can reduce long-term age [12], [13]. This insight has motivated a family of agethreshold policies, under which a device contends only when its local AoI exceeds a threshold δ. Atabay et al. [14], Chen et al. [15], and Yavascan et al. [16] studied such policies; in particular, Yavascan et al. showed that threshold-based access can substantially improve AoI scaling relative to plain slotted ALOHA. Ahmetoglu et al. [17] further improved this scaling via collision-sensing minislots, while Chen et al. [18] generalized the threshold to an age-gain criterion with orderoptimality guarantees. Collectively, these results show that age-aware control can yield substantial freshness gains over age-agnostic policies. SIC-aided protocols such as IRSA have also been studied from an AoI perspective [19], [20], and de Jesus et al. [21] extended age-dependent random access to a two-hop multi-relay topology. However, these works embed SIC in fixed, framesynchronous protocols, so repetition intensity and receiver-side diversity are not modeled as AoI-dependent control variables. At the system level, Zhou and Saad [22] formulated a meanfield game for carrier-sense multiple access (CSMA)-based ultra-dense IoT and proved the existence and convergence of a mean-field equilibrium for AoI-optimal backoff rates. While their framework highlights the potential of mean-field methods for large-scale AoI optimization, it relies on a CSMA model with closed-form transition rates and does not extend to asynchronous capture-SIC satellite uplinks. Despite this progress, three limitations remain unresolved in satellite IoT: (i) Most AoI-aware analyses assume slotsynchronous collision channels, whereas satellite uplinks feature propagation-delay offsets, fractional overlaps, fading, capture, and imperfect SIC. (ii) Repetition-based AoI studies typically optimize fixed access rules rather than AoI-dependent control over both replica count and receiver-side resource diversity. These coupled dimensions create a nontrivial tradeoff among reliability, congestion, and transmission effort that cannot be captured by a scalar access probability. (iii) The frame-level success probability is usually treated as exogenous, even though it is determined endogenously by the population’s access decisions. To address these gaps, we develop A STRA (Asynchronous Age-Aware Satellite Random Access), a mean-field Markov decision process (MDP) framework for AoI-aware satellite IoT random access. In A STRA, each device adapts the number of resource pools and the number of replicas per pool based solely on its local AoI. The main contributions are as follows: • System model. We develop the A STRA random-access model, which captures key satellite-uplink effects, including asynchronous packet arrivals, partial overlaps, capture, and SIC. Unlike conventional AoI random-access models that rely on idealized slot-collision abstractions or exogenously specified delivery probabilities [14], [15], [19], [20], our model preserves the dependence of delivery success on both a device’s access action and the
aggregate population behavior. Mean-field MDP framework. We formulate ASTRA as a mean-field AoI control problem in which each device adapts not only when to transmit but also its repetition level and receiver-side diversity. This extends existing threshold-ALOHA schemes [15]–[17], which mainly optimize the transmit-or-wait decision; fixed-repetition SIC schemes [19], [20], which do not adapt repetition to information freshness; and the mean-field game formulation in [22], which controls only a single backoff rate under CSMA. • Threshold structure and performance gains. We prove that the optimal policy admits an age-threshold structure, rather than assuming such a structure a priori as in [16], [17], [21]. This result shows that simple age-based access remains optimal for a given congestion level even when repetition and receiver-side diversity are adapted jointly. Simulations further show that the proposed policy reduces AoI relative to age-independent policies. •
II. S YSTEM M ODEL We consider a frame-based uplink satellite IoT randomaccess system with N ground devices and R receiver-side resource pools, as illustrated in Fig. 1. Devices sporadically generate status updates and access the satellite link without centralized per-frame scheduling. Each device makes one access decision per frame based on its own AoI. This grantfree and decentralized operation is well suited to massive satellite IoT, but it creates a cross-layer mismatch: access decisions are made at frame boundaries, whereas packet overlap, capture events, and successive interference cancellation (SIC) are determined by continuous-time interactions at the satellite/gateway receiver. A. Resource Pools and Frame Structure We model the satellite/gateway receiver through R parallel resource pools in each frame. A resource pool is a logical random-access pool: packets placed in the same pool contend with one another, and the receiver produces one pool-level decoding outcome after the corresponding satellite, beam, frequency, or code-domain receiver processing. A pool may be implemented by a frequency/code partition, a beam, a singlesatellite observation branch, or a gateway-side observation branch formed from observations of multiple visible satellites. Hence, R denotes the number of logical access pools, not necessarily the number of satellites. This abstraction allows multi-satellite observation diversity to be represented at the pool level through the pool-level decoding model, without requiring the access policy to select individual satellites. Time is divided into frames of duration Tf . Within each frame, each resource pool is partitioned into M logical slots of duration Tf . (1) Ts = M Hence, a frame consists of R parallel receiver-side access pools over a common observation interval, each with its own
3𝑇𝑠
2𝑇𝑠
𝑇𝑠
𝑀𝑇𝑠
Frame Length 𝑇𝑓 = 𝑀𝑇𝑠 𝑇𝑠
𝑇𝑠
Power
errors, and timing uncertainty. The corresponding packet interval at the receiver is
𝑇𝑠 Replica 1
Tp
P1
Replica 2 Tp
P2
Replica 3 Tp
P3 𝛿2 𝛿1
𝛿3
Replica 4 Tp
𝛿4 𝐿34
𝐿12
Iv = [tv , tv + Tp ).
As illustrated in Fig. 2, the residual offsets δv shift packet arrivals away from their nominal slot boundaries. Hence, even replicas with different nominal slot indices may still overlap fractionally within the same pool. D. Rician Fading and SIC Decoding
Time
Fig. 2: Asynchronous packet reception and partial overlaps. Packets assigned to slots arrive with residual offsets δi , leading to continuous-time transmissions of length Tp that may partially overlap. The overlap duration Lij determines the time-averaged interference contribution in SINR calculations.
Each replica experiences a random received power due to the satellite uplink channel. We model small-scale fading by a unit-mean Rician coefficient GRician . The received power of v replica v is therefore Pv = P̄v GRician , v
logical slot structure. Each transmitted replica has physical duration Tp ≤ Ts . B. Replica Repetition and Pool-Diversity Control At the beginning of each frame, a device chooses how aggressively to access the satellite uplink along two coupled dimensions: the number of selected resource pools and the number of replicas transmitted in each selected pool. We define the access action of device n in frame k as an,k ≜ (dn,k , qn,k ) ∈ A,
(2)
where qn,k denotes the number of selected resource pools and dn,k denotes the number of replicas transmitted in each selected pool. The two action components play different physical roles. The variable qn,k captures inter-pool diversity, whereas dn,k captures intra-pool repetition diversity. Increasing either one can improve update delivery reliability, but also increases transmission cost and network congestion [4], [5]. We also include the idle action (0, 0), which allows a device to skip transmission in a frame. Such deliberate waiting is important in a freshness-critical model: when the current AoI is small, deferring access can be preferable to transmitting immediately [12], [16], [23]. Assuming there are maximum D per-pool repetitions, the action space for each user is: A = {(0, 0)} ∪ {(d, q) : d ∈ {1, . . . , D}, q ∈ {1, . . . , R}}. (3) The transmission cost for user n in frame k is modeled as the total number of transmitted replicas, E(an,k ) = dn,k qn,k ,
(6)
∀n ∈ {1, · · · , N }, k ∈ N+ .
(4)
(7)
where P̄v denotes the nominal received power and the Rician K-factor characterizes the relative strength of the line-of-sight component. For two replicas u and v, their overlap length is Luv = |Iu ∩ Iv |. At SIC iteration ℓ, the time-averaged interference seen by replica u is 1 X (ℓ) αv Pv Luv , (8) Iu(ℓ) = Tp v̸=u
(ℓ)
where αv ∈ {1, ϵ} is the residual interference factor. Initially, (0) αv = 1 for all replicas; once replica v is decoded, its residual factor is updated to ϵ. Replica u is decodable if SINR(ℓ) u =
Pu (ℓ) σ 2 + Iu
≥ γth ,
(9)
where σ 2 is the receiver noise power and γth is the capture threshold. Let o n (10) D(ℓ) = u : SINR(ℓ) u ≥ γth denote the set of decodable replicas at iteration ℓ. If D(ℓ) = ∅, the SIC procedure stops. Otherwise, the receiver decodes the highest-power replica among the currently decodable ones. In the present implementation, we adopt the strongest-first rule u⋆ ∈ arg max Pu . u∈D (ℓ)
(11)
A tagged device is declared successful in a frame if at least one of its replicas is decoded in at least one selected pool after the pool-level capture-SIC procedure and gateway-level OR fusion. The resulting frame-level success probability is summarized in the calibrated interface introduced next.
C. Asynchronous Arrival
E. Success Law, AoI Dynamics, and Design Objective
The frame-level access action is evaluated through an asynchronous capture-SIC decoding process at the satellite/gateway receiver. If replica v is assigned to logical slot mv , its receiverside start time is
The success probability of a tagged device depends on two quantities: its own frame-level access action and the aggregate interference generated by the remaining population. To make this dependence explicit, define the empirical per-pool load seen by device n in frame k as X dj,k qj,k e −n (k) ≈ 1 Λ , (12) Tf R
tv = (mv − 1)Ts + δv ,
(5)
where δv ∈ [0, Ts ) is a residual timing offset due to heterogeneous satellite propagation delays, residual synchronization
j̸=n
where dj,k qj,k is the number of replicas transmitted by device j, and the factor 1/R reflects uniform pool selection. Thus, the tagged action an,k represents the device’s own control e −n (k) represents the congestion environment decision, while Λ induced by the other devices. The formal mean-field version of this load descriptor is given in (26). Given a tagged action a = (d, q) and a per-pool load Λ, the asynchronous physical layer is summarized by the calibrated success law p̂(a; Λ) ≜ P(Y = 1 | a, Λ), (13) where Y ∈ {0, 1} is a generic tagged-device frame-level success indicator. The value p̂(a; Λ) is the probability that at least one tagged replica is decoded in at least one selected pool after asynchronous capture-SIC and gateway-level OR fusion. For device n in frame k, let Yn (k) ∈ {0, 1} denote the success indicator. Under the calibrated success law, e −n (k) = Λ = p̂(a; Λ). (14) P Yn (k) = 1 an,k = a, Λ Let ∆n (k) ∈ N+ denote the gateway-side AoI of device n at the beginning of frame k, representing the elapsed time (in frames) since its most recently accepted update. The AoI evolves as ( 1, Yn (k) = 1, ∆n (k + 1) = (15) ∆n (k) + 1, Yn (k) = 0. To accommodate the intrinsic scalability of massive grantfree satellite IoT where centralized per-frame coordination is practically infeasible, we focus on purely distributed access policies. Under this paradigm, each device operates within a decoupled local perfect feedback loop, observing only its own gateway-side AoI without any knowledge of the instantaneous actions, AoI states, or slot choices of neighboring devices. Consequently, we restrict our attention to the class of symmetric stationary AoI-dependent policies: π(∆), ∆ ∈ N+ ,
(16)
For any given symmetric policy π, the long-term average AoI per device is: T −1 N
1 XX ¯ Eπ [∆n (k)], ∆(π) = lim sup T →∞ N T n=1
The design goal is to balance information freshness and transmission effort. Formally, this motivates the following constrained optimization problem: ¯ min ∆(π) s.t.
Ē(π) = B,
where Π denotes the class of symmetric stationary AoIdependent policies and B represents the strictly enforced average replica budget. To establish tractability, we resort to the unconstrained Lagrangian scalarization: ¯ min ∆(π) + η Ē(π), π∈Π
η ≥ 0,
III. M EAN -F IELD MDP The calibrated success law p̂(a; Λ) couples each device’s local access decision with the congestion generated by the population. We first study the representative-device MDP under a fixed load Λ, then impose a self-consistency condition that closes the mean-field loop. A. Representative MDP Under Fixed Load Λ Fix a per-pool congestion intensity Λ. For computation, we use the finite AoI state space D = {1, . . . , ∆max }. The representative device observes ∆ ∈ D and selects an action a ∈ A. Under the calibrated success law, the transition kernel is p̂(a; Λ), ∆′ = 1, PΛ (∆′ | ∆, a) = 1 − p̂(a; Λ), ∆′ = min{∆ + 1, ∆max }. (21) For an energy multiplier η ≥ 0, define the one-stage Lagrangian cost cη (∆, a) = ∆ + ηE(a). (22) The finite-state average-cost Bellman equation is then given by [24]
(17)
ρη (Λ) + Vη (∆; Λ) = min Qη,Λ (∆, a), a∈A
and the corresponding long-term average transmission cost is formulated as: T −1 N
(18)
k=0
Here, the expectation Eπ [·] is taken over the joint probability measure induced by the local randomized action selections, the underlying asynchronous physical-layer randomness, and the network congestion process emerging when all devices independently execute the same policy. Noting that E(a) = dq, the metric Ē(π) explicitly quantifies the average number of transmitted replicas per device per frame, thereby serving as a direct analytical proxy for uplink energy consumption.
(20)
where η controls the AoI-energy tradeoff. Larger η favors conservative access and deliberate waiting, while smaller η favors more aggressive update attempts through stronger repetition and broader pool diversity.
k=0
1 XX Eπ [E(an,k )]. Ē(π) = lim sup T →∞ N T n=1
(19)
π∈Π
(23)
where Qη,Λ (∆, a) ≜∆ + ηE(a) + p̂(a; Λ)Vη (1; Λ) + 1 − p̂(a; Λ) Vη (min{∆ + 1, ∆max }; Λ). (24) A fixed-load best response is any selector br πη,Λ (∆) ∈ arg min Qη,Λ (∆, a). a∈A
(25)
In the implementation, (23) is solved by relative value iteration with reference-state normalization [24].
B. Mean-Field Consistency
D. Structural Properties of the Bellman Equation
For a stationary policy π(∆) and population AoI distribution m(∆), the induced per-pool replica start-time intensity is N −1 X d(π(∆))q(π(∆)) Λ(m, π) = m(∆) . (26) Tf R
In this subsection, we establish three basic structural properties of the single-user Bellman equation under a fixed meanfield load Λ: the existence of an average-cost optimality equation (ACOE), the monotonicity of the relative value function, and the threshold structure of the optimal action. These properties provide the theoretical basis for the thresholdtype policies observed later in the numerical results. Unless otherwise stated, the structural results below are stated for the untruncated AoI dynamics, while ∆max is used only in the finite-state numerical MDP.
∆∈D
In (26), the term N − 1 removes the tagged device from the population count. The remaining factor gives the expected number of replicas that a device using action π(∆) injects into a generic pool. Theorem 1 (Existence of a mean-field fixed point). Fix the energy multiplier η ≥ 0. Under the finite-state and continuity assumptions stated in [1, Appendix A], there exists a stationary mean-field operating point πη⋆ ∈ BRη (Λ⋆η ),
(27)
m⋆η = µ(πη⋆ , Λ⋆η ), Λ⋆η = Λ(m⋆η , πη⋆ ),
(28) (29)
where BRη (Λ) denotes the set of stationary optimal policies under load Λ, and µ(π, Λ) is the stationary distribution of the Markov chain induced by (21) under policy π.
Theorem 2 (ACOE existence with zero-success actions allowed). There exist a scalar ρη , a finite-valued relative value function V : D → R, and a stationary deterministic policy π ⋆ : D → A such that h ρη + V (∆) = min ∆ + ηE(a) + p̂(a; Λ)V (1) a∈A i + 1 − p̂(a; Λ) V (∆ + 1) , ∆ ≥ 1. (30) Moreover, any stationary deterministic minimizer of the righthand side of (30) is average-cost optimal.
■
Proof. The proof is given in [1, Appendix B].
Equations (27)–(29) show the closed-loop nature of the problem. The policy depends on Λ through the success probability, while Λ is induced by the policy and the stationary AoI distribution. Theorem 1 guarantees existence of a relaxed mean-field operating point for the truncated model. The numerical algorithm below searches for such a self-consistent operating point, typically returning a deterministic policy when the Bellman minimizer is unique.
Corollary 1 (Action Dominance). Suppose that
Proof. The proof is given in [1, Appendix A].
C. Numerical Fixed-Point Solution For each η, we compute the mean-field operating point by a nested fixed-point iteration. 1) For a provisional load Λ, solve the representative MDP in (23) to obtain a best-response policy. 2) For the current AoI distribution m, update the load using Λnew = Λ(m, π).
E(aij ) = E(aji ),
3) Under the resulting policy and load, compute the stationary AoI distribution µ, and update m ← (1 − α)m + αµ. The first two steps enforce load consistency, while the third step enforces population consistency. The iteration stops when both the load and AoI distribution residuals are below prescribed tolerances.
(31)
Then action aji is dominated by action aij and cannot appear in the optimal policy. Proof. The proof is given in [1, Appendix E].
■
Theorem 3 (Threshold structure of the optimal policy). Fix the mean-field load Λ and consider the Bellman equation (23). Define h(∆) ≜ V (∆ + 1) − V (1). (32) Then the optimal action satisfies a⋆ (∆) ∈ arg min {ηE(a) − p̂(a; Λ)h(∆)} . a∈A
(33)
Assume that the non-dominated effective actions Aeff (Λ) = {a0 , a1 , . . . , aK }
A damped update is used: Λ ← (1 − β)Λ + βΛnew .
p̂(aij ; Λ) > p̂(aji ; Λ).
■
can be ordered so that E(a0 ) < E(a1 ) < · · · < E(aK ), and p̂(a0 ; Λ) < p̂(a1 ; Λ) < · · · < p̂(aK ; Λ).
(34)
Then the optimal policy is of threshold type in the AoI state ∆. Proof. The proof is given in [1, Appendix D].
■
IV. N UMERICAL R ESULTS We now evaluate A STRA, the proposed mean-field MDP framework. The numerical results are designed to illustrate
100
102
(a)
$(x)=0.5x+0.5x2 $(x)=0.9x+0.1x2
Average AoI
10
$(x)=0.1x+0.9x2 Age-agnostic RA ASTRA (Prop.)
-1
10-2
10-3 0
5
10
15
20
101
25
10-1
100
Mean Energy per Device (Replicas/Frame)
Fig. 3: Success probability versus per-pool congestion intensity Λ for all transmission actions (d, q). Noise variance is σ 2 = 0.5.
140
(b)
120
A. Simulation Setup We consider N = 30 devices, R = 3 resource pools, frame duration Tf = 1, capture threshold γth = 2 and AoI truncation level ∆max = 200. The success interface is calibrated offline using the asynchronous packet-level simulator. In the considered configuration, each frame contains M = 3 logical slots, and each packet has duration Tp = 0.25. The lookup table is computed over the load grid Λ ∈ {0, 2, 4, . . . , 25} and over the action set in (3). For each table entry, the success probability is estimated by Monte Carlo simulation under Rician fading, additive noise, and capture-SIC decoding. To characterize the AoI-energy tradeoff, we sweep the energy weight η over a logarithmic grid and solve the associated mean-field fixed point for each value of η. B. Calibrated Success Interface Fig. 3 shows the calibrated success probability p̂(a; Λ) as a function of the per-pool congestion intensity Λ. As expected, the success probability decreases as the aggregate load increases. The decay is action-dependent: actions with stronger repetition or broader pool usage may provide higher reliability at low or moderate congestion, but they also become more vulnerable as the per-pool replica-start intensity grows. This behavior is precisely why the lambda interface is useful for ASTRA: it captures the physical tradeoff between reliability gain and congestion-induced interference. C. AoI-Energy Tradeoff We compare ASTRA with two age-independent baselines. • IRSA-inspired Baseline: Each active device transmits replicas according to a prescribed replica-degree distribution. In our implementation, we consider three fixed degree distributions over one-replica and two-replica transmissions, namely ΛIRSA (x) = αx + (1 − α)x2 ,
100 Average AoI
three aspects of A STRA: the calibrated success interface, the AoI-energy tradeoff induced by the mean-field policy, and the threshold structure of the resulting optimal actions.
80
60
40
$(x)=0.5x+0.5x2 $(x)=0.9x+0.1x2 $(x)=0.1x+0.9x2 Age-agnostic RA ASTRA (Prop.) 10-1 100 Mean Energy per Device (Replicas/Frame)
Fig. 4: Average AoI versus mean energy per device for the proposed ASTRA scheme and baseline policies. Noise variance is (a) σ 2 = 0.5 and (b) σ 2 = 1.
with α ∈ {0.5, 0.1, 0.9}. Pool selection is fixed to q = 1. To make the comparison energy-consistent, each fixed distribution is mixed with the idle action (0, 0), so that the resulting average replica budget matches the target energy level [1, Appendix G]. These baselines capture standard IRSA-inspired randomized repetition schemes under the same success-probability approximation used for our system, but using AoI-independent decisions. • Age-agnostic Random Access Baseline: All devices use the same stationary randomized policy that is independent of AoI. Specifically, each device selects action ai ∈ A with probability ri , regardless of its current AoI. For each average-energy level, the common mixing vector r is optimized through the linear program in [1, Appendix F] to maximize the resulting average success probability. Fig. 4 reports the resulting AoI–energy tradeoff. The red curve shows the computed ASTRA operating points obtained by sweeping the energy multiplier. The IRSA-inspired baselines are shown as individual markers, while the dotted blue curve gives the optimized age-agnostic randomized baseline. ASTRA achieves a much lower average AoI over the plotted energy range, especially in the low-energy regime. This gain comes from using energy selectively in stale AoI states, rather than spending transmissions independently of freshness.
a
Budget B = 0.20
Selected action (d,q)
(2,2) (2,1)
"=59
(2,2) (2,1)
"=35
(1,2) (1,1)
Budget B = 1.09
b "=127
"=20
(1,2)
"=12
"=17
(1,1)
(0,0)
(0,0) 50
100
AoI state, "
50
100
AoI state, "
Fig. 5: Threshold structures of the equilibrium policies under two representative energy budgets.
D. Optimal Policies Under Representative Energy Budgets Fig. 5 shows the deterministic policies for two energy budgets. Under the tighter budget Fig. 5(a), the policy stays conservative over most AoI states and switches to higherenergy actions only when AoI becomes large, since the energy penalty ηE(a) dominates. Under the larger budget Fig. 5(b), switching thresholds shift leftward, activating stronger actions at smaller AoI values because the effective energy penalty is weaker. Both policies exhibit a clear threshold structure: conservative actions at low AoI, switching to higher-energy actions as AoI grows, consistent with Theorem 3. The absence of action a21 agrees with Corollary 1. V. C ONCLUSION This paper developed A STRA, a mean-field MDP framework for AoI-aware satellite IoT random access under asynchronous capture-SIC decoding. The physical layer is summarized by a calibrated success interface p̂(a; Λ), allowing a tractable frame-level control model. Each device adapts its repetition and pool-diversity action using only its local AoI, with population congestion determined self-consistently. Since this is a novel AoI-dependent asynchronous randomaccess formulation, we compared the A STRA policy with AoI-independent baselines evaluated under the same physical model. The numerical results show that A STRA improves the AoI-energy tradeoff by using conservative actions at small AoI and switching to more aggressive actions when updates become stale, which is consistent with the threshold structure derived from the Bellman equation. R EFERENCES [1] S. Chakraborty, A. Li, Y. İnce, S. Baghaee, and E. Uysal, “A STRA: Asynchronous age-aware satellite random access via mean-field control,” arXiv preprint, 2026. [2] J. A. Fraire, S. Céspedes, and N. Accettura, “Direct-to-satellite IoT – a survey of the state of the art and future research perspectives,” in Proc. Int. Conf. Ad-Hoc, Mobile, Wireless Netw. (ADHOC-NOW), ser. LNCS, vol. 11604. Springer, 2019, pp. 241–258. [3] O. Kodheli et al., “Satellite communications in the new space era: A survey and future challenges,” IEEE Commun. Surveys Tuts., vol. 23, no. 1, pp. 70–109, 2021. [4] E. Casini, R. D. Gaudenzi, and O. del Rio Herrero, “Contention resolution diversity slotted ALOHA (CRDSA): An enhanced random access scheme for satellite access packet networks,” IEEE Trans. Wireless Commun., vol. 6, no. 4, pp. 1408–1419, Apr. 2007.
[5] G. Liva, “Graph-based analysis and optimization of contention resolution diversity slotted ALOHA,” IEEE Trans. Commun., vol. 59, no. 2, pp. 477–487, Feb. 2011. [6] E. Paolini, G. Liva, and M. Chiani, “Coded slotted ALOHA: A graphbased method for uncoordinated multiple access,” IEEE Trans. Inf. Theory, vol. 61, no. 12, pp. 6815–6832, Dec. 2015. [7] X. Shao, Z. Sun, M. Yang, S. Gu, and Q. Guo, “NOMA-based irregular repetition slotted ALOHA for satellite networks,” IEEE Commun. Letters, 2019. [8] E. Recayte, T. Devaja, and D. Vukobratovic, “Energy-efficient irregular repetition slotted ALOHA for IoT satellite systems,” in Proc. IEEE Int. Conf. on Commun. Workshops (ICC Workshops), 2024. [9] E. Recayte and C. Amatetti, “Multi-satellite NOMA-irregular repetition slotted ALOHA for IoT networks,” arXiv preprint arXiv:2601.00341, 2026. [10] S. Kaul, R. Yates, and M. Gruteser, “Real-time status: How often should one update?” in Proc. IEEE INFOCOM, 2012, pp. 2731–2735. [11] R. D. Yates, Y. Sun, D. R. Brown, S. K. Kaul, E. Modiano, and S. Ulukus, “Age of information: An introduction and survey,” IEEE J. Sel. Areas Commun., vol. 39, no. 5, pp. 1183–1210, May 2021. [12] Y. Sun, E. Uysal-Biyikoglu, R. D. Yates, C. E. Koksal, and N. B. Shroff, “Update or wait: How to keep your data fresh,” IEEE Trans. Inf. Theory, vol. 63, no. 11, pp. 7492–7508, Nov. 2017. [13] R. D. Yates and S. K. Kaul, “Status updates over unreliable multiaccess channels,” in Proc. IEEE Int. Symp. Inf. Theory (ISIT), 2017, pp. 331– 335. [14] D. C. Atabay, E. Uysal, and O. Kaya, “Improving age of information in random access channels,” in Proc. IEEE INFOCOM Workshops, 2020, pp. 912–917. [15] H. Chen, Y. Gu, and S.-C. Liew, “Age-of-information dependent random access for massive IoT networks,” in Proc. IEEE INFOCOM Workshops, 2020, pp. 930–935. [16] O. T. Yavascan and E. Uysal, “Analysis of slotted ALOHA with an age threshold,” IEEE J. Sel. Areas Commun., vol. 39, no. 5, pp. 1456–1470, May 2021. [17] M. Ahmetoglu, O. T. Yavascan, and E. Uysal, “MiSTA: An ageoptimized slotted ALOHA protocol,” IEEE Internet Things J., vol. 9, no. 17, pp. 15 484–15 496, Sep. 2022. [18] X. Chen, K. Gatsis, H. Hassani, and S. S. Bidokhti, “Age of information in random access channels,” IEEE Trans. Inf. Theory, vol. 68, no. 10, pp. 6548–6568, Oct. 2022. [19] A. Munari, “Modern random access: An age of information perspective on irregular repetition slotted ALOHA,” IEEE Trans. Commun., vol. 69, no. 6, pp. 3572–3585, Jun. 2021. [20] J. F. Grybosi, J. L. Rebelatto, and G. L. Moritz, “Age of information of SIC-aided massive IoT networks with random access,” IEEE Internet Things J., vol. 9, no. 1, pp. 662–670, Jan. 2022. [21] G. G. M. de Jesus, J. L. Rebelatto, and R. D. Souza, “Age-of-information dependent random access in multiple-relay slotted ALOHA,” IEEE Access, vol. 10, pp. 112 076–112 085, 2022. [22] B. Zhou and W. Saad, “Age of information in ultra-dense IoT systems: Performance and mean-field game analysis,” IEEE Trans. Mobile Comput., vol. 23, no. 5, pp. 4533–4547, May 2024. [23] H. Tang, Y. Chen, J. Wang, P. Yang, and L. Tassiulas, “Age optimal sampling under unknown delay statistics,” IEEE Trans. Inf. Theory, vol. 69, no. 2, pp. 1295–1314, Feb. 2023. [24] M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming. John Wiley & Sons, 1994.
A PPENDIX A P ROOF OF T HEOREM 1 We use the following standard assumptions for the finitestate mean-field MDP. First, the truncated AoI state space D = {1, . . . , ∆max } and the action set A are finite. Second, for each action a ∈ A, the calibrated success interface p̂(a; Λ) is continuous in Λ over a compact interval [0, Λmax ]. Third, the load interval contains all feasible population-induced loads, i.e., 0 ≤ Λ(m, π) ≤ Λmax
for all stationary policies π and distributions m. These conditions hold for the truncated numerical model when the lookup table is interpolated continuously and Λmax is chosen large enough to cover the maximum per-pool replica intensity.
For every state with m⋆η (∆) > 0, define
Proof. We prove the result using a stationary occupationmeasure formulation. For a fixed load Λ, define the transition kernel PΛ (∆′ | ∆, a)
For states with m⋆η (∆) = 0, choose any distribution over A. By construction, x⋆η is optimal for the representative MDP under load Λ⋆η , so πη⋆ ∈ BRη (Λ⋆η ). The flow constraints in (35) imply that m⋆η is stationary under (πη⋆ , Λ⋆η ). Finally, (37) gives Λ⋆η = Λ(m⋆η , πη⋆ ).
according to the AoI dynamics and the calibrated success probability p̂(a; Λ). Let x(∆, a) denote a stationary stateaction occupation measure over D × A. For fixed Λ, the feasible occupation-measure set is X (Λ) = {x ≥ 0 : C0 (x) = 1, C∆′ (x) = 0, ∀∆′ ∈ D}, X X C0 (x) ≜ x(∆, a), ∆∈D a∈A
X
′
X X
′
x⋆η (∆, a) . ⋆ b∈A xη (∆, b)
πη⋆ (a | ∆) = P
Thus (m⋆η , πη⋆ , Λ⋆η ) satisfies (27)–(29), proving the existence of a stationary mean-field fixed point. ■ A PPENDIX B P ROOF OF T HEOREM 2 A. Useful Lemma
x(∆, a)PΛ (∆ | ∆, a).
Define the normalized relative value gap
(35) Because the state and action spaces are finite, X (Λ) is nonempty, compact, and convex. Nonemptiness follows from the existence of a stationary distribution for every finite Markov chain induced by a stationary policy. For fixed Λ, the representative average-cost MDP with multiplier η can be written as the linear program X X min x(∆, a) ∆ + ηE(a) . (36)
H(∆) ≜ V (∆) − V (1).
C∆′ (x) ≜
x(∆ , a) −
a∈A
∆∈D a∈A
x∈X (Λ)
∆∈D a∈A
Let Oη (Λ) denote the set of optimal solutions of (36). Since the feasible set is compact and the objective is linear, Oη (Λ) is nonempty, compact, and convex. Next define the load induced by an occupation measure as d(a)q(a) N −1 X X x(∆, a) . (37) G(x) = Tf R ∆∈D a∈A
This map is linear and hence continuous. Now define the setvalued map Γ(x) = Oη (G(x)). The domain is the compact convex probability simplex over D × A. Moreover, Γ(x) is nonempty, compact, and convex for every x. Because PΛ is continuous in Λ, the feasible occupation-measure correspondence X (Λ) is closed-graph, and by Berge’s maximum theorem the optimal-solution correspondence Oη (Λ) is upper hemicontinuous. Since G(x) is continuous, Γ is also upper hemicontinuous. Therefore, by fixed-point theorem, there exists an occupation measure x⋆η such that x⋆η ∈ Γ(x⋆η ) = Oη (G(x⋆η )). Set Λ⋆η = G(x⋆η ),
m⋆η (∆) =
X
x⋆η (∆, a).
(38)
(39)
Lemma 1 (Monotonicity of the relative value gap). The relative value gap H(∆) is nondecreasing in ∆, i.e., H(∆ + 1) ≥ H(∆),
∀∆ ≥ 1.
Proof. The proof is given in Appendix C.
(40) ■
B. Formal Proof Proof. For α ∈ (0, 1), define the discounted value function "∞ # X π t Vα (∆) = inf E∆ α ∆(t) + ηE(a(t)) . (41) π
t=0
It satisfies the discounted Bellman equation h Vα (∆) = min ∆ + ηE(a) + αp̂(a; Λ)Vα (1) a∈A i + α 1 − p̂(a; Λ) Vα (∆ + 1) .
(42)
Consider the constant policy that always applies ā until the first reset to state 1. Let T be the first reset time. Then T is geometrically distributed with parameter p̂(ā; Λ), so "T −1 # X 1 − p̂(ā; Λ) 1 , E t = . (43) E[T ] = p̂(ā; Λ) p̂(ā; Λ)2 t=0 Using this admissible policy, we obtain the bound 0 ≤ Hα (∆) ≜ Vα (∆) − Vα (1) ∆ + ηE(ā) 1 − p̂(ā; Λ) ≤ + , p̂(ā; Λ) p̂(ā; Λ)2
(44)
which is finite for every fixed ∆ and uniform in α. Hence, for each fixed ∆, the family {Hα (∆)}α∈(0,1) is bounded. Along a sequence αn ↑ 1, we may extract a pointwise limit
a∈A
Hαn (∆) → H(∆),
∀∆ ≥ 1,
(45)
and (1 − αn )Vαn (1) → ρη .
(46)
Subtracting Vα (1) from both sides of (42) yields h (1 − α)Vα (1) + Hα (∆) = min ∆ + ηE(a) a∈A i + α 1 − p̂(a; Λ) Hα (∆ + 1) . (47) Letting αn ↑ 1 in (47), and using (45) together with (46), we obtain h ρη + H(∆) = min ∆ + ηE(a) a∈A i + (1 − p̂(a; Λ))H(∆ + 1) . (48) Since H(∆) = V (∆) − V (1), this is equivalent to (30). Because A is finite, the minimizer of the right-hand side can be chosen as a deterministic function of ∆. Any such stationary deterministic minimizer is average-cost optimal. ■ A PPENDIX C P ROOF OF L EMMA 1
A PPENDIX D P ROOF OF T HEOREM 3 Proof. From (30), separate the action-independent terms: ρη + V (∆) = ∆ + V (∆ + 1) h + min ηE(a) a∈A
i − p̂(a; Λ) V (∆ + 1) − V (1) . (57) Using (32), we obtain ρη + V (∆) = ∆ + V (∆ + 1) h i + min ηE(a) − p̂(a; Λ)h(∆) , a∈A
(58)
which yields (33). Now define ga (h) ≜ ηE(a) − p̂(a; Λ)h.
(59)
For two effective actions ai and aj with i < j, (34) gives
Proof. For α ∈ (0, 1), define the discounted Bellman operator h (Tα W )(∆) = min ∆ + ηE(a) + αp̂(a; Λ)W (1) a∈A i + α 1 − p̂(a; Λ) W (∆ + 1) . (49) Suppose W (∆) is nondecreasing in ∆. For any fixed action a, define Fa (∆) = ∆ + ηE(a) + αp̂(a; Λ)W (1) + α 1 − p̂(a; Λ) W (∆ + 1).
(50)
Then
E(ai ) < E(aj ),
p̂(ai ; Λ) < p̂(aj ; Λ).
(60)
Hence gaj (h) − gai (h) = η E(aj ) − E(ai ) − p̂(aj ; Λ) − p̂(ai ; Λ) h.
(61)
The right-hand side is a strictly decreasing affine function of h. Therefore the two action costs cross at most once, at η E(aj ) − E(ai ) Hij (Λ) = . (62) p̂(aj ; Λ) − p̂(ai ; Λ) Equivalently,
Fa (∆ + 1) − Fa (∆) = 1 + α 1 − p̂(a; Λ)
× W (∆ + 2) − W (∆ + 1)
gaj (h) ≤ gai (h)
(51)
≥ 0,
(52)
so Fa (∆) is nondecreasing for every action a. Therefore, Tα W is also nondecreasing. Starting value iteration from the constant function W0 (∆) ≡ 0, all iterates Wn+1 = Tα Wn (53) are nondecreasing. Since discounted value iteration converges to Vα , the discounted value function Vα (∆) is nondecreasing. Hence Hα (∆) ≜ Vα (∆) − Vα (1) (54) is also nondecreasing. From Theorem 2, along a sequence αn ↑ 1, Hαn (∆) → H(∆),
∀∆ ≥ 1.
(55)
Since each Hαn is nondecreasing and pointwise limits preserve monotonicity, H(∆) is nondecreasing. This proves H(∆ + 1) ≥ H(∆),
∀∆ ≥ 1.
(56) ■
⇐⇒
h ≥ Hij (Λ).
(63)
Thus, when h is small, the lower-energy action is preferred, while for sufficiently large h, the higher-success action is preferred. This is the single-crossing property. By Lemma 1, V (∆) is nondecreasing in ∆, so h(∆) = V (∆ + 1) − V (1)
(64)
is nondecreasing in ∆. Therefore, as ∆ increases, h(∆) crosses the pairwise thresholds Hij (Λ) in order, and the minimizing action can only move from lower-energy/lowersuccess actions to higher-energy/higher-success actions. Hence the optimal policy is of threshold type in the AoI state. ■ A PPENDIX E P ROOF OF C OROLLARY 1 Proof. Recall the effective action objective ga h(∆) = ηE(a) − p̂(a; Λ)h(∆).
(65)
Using (31), we obtain ga21 h(∆) − ga12 h(∆) = η E(a21 ) − E(a12 ) + p̂(a12 ; Λ) − p̂(a21 ; Λ) h(∆) = p̂(a12 ; Λ) − p̂(a21 ; Λ) h(∆) ≥ 0. (66)
Hence ga12 h(∆) ≤ ga21 h(∆) ,
∀ ∆.
(67)
Moreover, the inequality is strict whenever h(∆) > 0. Therefore, action a21 is dominated by action a12 and cannot be selected by the optimal policy. ■
proposed policy are compared under the same average replica budget. We consider three prescribed IRSA-type replica-degree distributions over one-replica and two-replica transmissions: Λα (x) = αx + (1 − α)x2 ,
α ∈ {0.5, 0.1, 0.9}.
(73)
This appendix describes the age-independent randomized baseline used in Fig. 4. Consider a policy that chooses action ai ∈ A with probability ri , independently of the AoI state. Let r = (r1 , . . . , rK )
Equivalently, an active device selects degree d = 1 with probability α and degree d = 2 with probability 1−α. In these IRSA baselines, pool diversity is not used and the number of selected pools is fixed as q = 1. Therefore, the transmission action is either (1, 1) or (2, 1), and the per-frame transmission cost is E(d, q) = dq. (74)
denote the action-mixing vector, where K = |A|. The average energy of this policy is
The mean number of replicas transmitted by an active IRSA device is then
A PPENDIX F AGE -I NDEPENDENT R ANDOMIZED BASELINE
Ē(r) =
K X
d¯α = α · 1 + (1 − α) · 2 = 2 − α. ri E(ai ).
(68)
i=1
Under the lambda approximation, if the baseline is evaluated at average energy c, the induced per-pool load is Λ(c) =
N −1 c . Tf R
(69)
For fixed c, the average success probability of the randomized policy is K X p̄(r; c) = ri p̂(ai ; Λ(c)). (70)
Since the proposed system allows the idle action (0, 0), we match a target average energy budget B by mixing the fixed IRSA transmission rule with the idle action. Let θα (B) denote the probability that a device is active in a frame under the IRSA baseline. To achieve average energy B, we set B B . θα (B) = ¯ = 2−α dα
The best age-independent randomized policy at energy level c is obtained from the linear program p̄ (c) = max r
s.t.
K X i=1 K X
ri p̂(ai ; Λ(c))
(71)
r0,0 (B) = 1 − θα (B),
(α)
(77)
(α) r1,1 (B) = θα (B)α, (α) r2,1 (B) = θα (B)(1 − α),
(78) (79)
with all other action probabilities equal to zero. By construction, the resulting average energy is (α)
(α)
(α)
ĒIRSA (B) = r1,1 (B)E(1, 1) + r2,1 (B)E(2, 1)
ri E(ai ) = c,
i=1 K X
(76)
Thus, the complete action distribution of the IRSA-inspired baseline is
i=1
⋆
(75)
= θα (B) [α · 1 + (1 − α) · 2] ri = 1,
ri ≥ 0,
i=1
The corresponding age-independent randomized baseline is ¯ rand (c) = ∆
= θα (B)(2 − α) = B.
i = 1, . . . , K.
1 . p̄⋆ (c)
(72)
This expression follows from the geometric AoI law induced by a state-independent Bernoulli success process. The baseline is optimal only within the restricted class of AoI-independent randomized policies. Therefore, it is not a lower bound on the performance of AoI-dependent policies. A PPENDIX G E NERGY N ORMALIZATION FOR IRSA- INSPIRED BASELINES This appendix describes how the energy budget is computed for the IRSA-inspired baselines used in the numerical comparison. The purpose is to ensure that the IRSA baselines and the
(80)
Hence, the IRSA baseline is energy-matched to the proposed policy at the same average replica budget. In our mean-field load approximation, the average replica budget B and the per-pool load Λ are related by Λ=
N −1 B, RTf
(81)
or equivalently,
RTf Λ. (82) N −1 When the Monte Carlo success-probability table is indexed by a discrete load variable G, we identify G with Λ and use B=
RTf G. (83) N −1 For a given operating point G, the activity probability in (76) B(G) =
is therefore computed as θα (G) =
B(G) RTf G = . 2−α (N − 1)(2 − α)
(84)
Let p̂((d, q); Λ) denote the calibrated frame-level success probability of a tagged device using action (d, q) under perpool load Λ. Under the above IRSA action distribution, the average success probability of the IRSA-inspired baseline is (α)
(α)
(α)
pIRSA (B) = r1,1 (B)p̂((1, 1); ΛB ) + r2,1 (B)p̂((2, 1); ΛB ) = θα (B) [αp̂((1, 1); ΛB ) + (1 − α)p̂((2, 1); ΛB )] , (85) where ΛB =
N −1 B. RTf
(86)
The idle action contributes zero successful updates and is therefore omitted from (85). Finally, under the Bernoulli frame-level success approximation, the AoI process of this age-agnostic IRSA baseline is a geometric reset process: ( (α) 1, with probability pIRSA (B), ∆(k + 1) = (α) ∆(k) + 1, with probability 1 − pIRSA (B). (87) Thus, the corresponding average AoI is computed as ¯ (α) (B) = ∆ IRSA
1 (α) pIRSA (B)
.
(88)
The construction above is feasible when 0 ≤ θα (B) ≤ 1,
or equivalently
0 ≤ B ≤ 2 − α.
(89)
Operating points outside this range cannot be matched exactly by mixing the fixed IRSA degree distribution with the idle action alone, and are therefore excluded from the IRSAinspired curve.