JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
1
RL-ASL: A Dynamic Listening Optimization for TSCH Networks Using Reinforcement Learning
arXiv:2604.07533v1 [cs.NI] 8 Apr 2026
F. Fernando Jurado-Lasso , Member, IEEE, and J. F. Jurado
Abstract—Time Slotted Channel Hopping (TSCH) is a widely adopted Media Access Control (MAC) protocol within the IEEE 802.15.4e standard, designed to provide reliable and energyefficient communication in Industrial Internet of Things (IIoT) networks. However, state-of-the-art TSCH schedulers rely on static slot allocations, resulting in idle listening and unnecessary power consumption under dynamic traffic conditions. This paper introduces RL-ASL, a reinforcement learning–driven adaptive listening framework that dynamically decides whether to activate or skip a scheduled listening slot based on real-time network conditions. By integrating learning-based slot skipping with standard TSCH scheduling, RL-ASL reduces idle listening while preserving synchronization and delivery reliability. Experimental results on the FIT IoT-LAB testbed and Cooja network simulator show that RL-ASL achieves up to 46% lower power consumption than baseline scheduling protocols, while maintaining near-perfect reliability and reducing average latency by up to 96% compared to PRIL-M. Its link-based variant, RL-ASLLB, further improves delay performance under high contention with similar energy efficiency. Importantly, RL-ASL performs inference on constrained motes with negligible overhead, as model training is fully performed offline. Overall, RL-ASL provides a practical, scalable, and energy-aware scheduling mechanism for next-generation low-power IIoT networks. Index Terms—Energy Efficiency, Internet of Things, Idle Listening, Reinforcement Learning, Time Slotted Channel Hopping.
I. I NTRODUCTION
T
HE rapid evolution of Networked Embedded Systems (NESs) has revolutionized the Internet of Things (IoT), enabling seamless connectivity and intelligent decisionmaking across diverse applications. IoT technologies are integral to critical domains such as industrial automation, healthcare monitoring, and smart grid systems [1]–[3]. These systems require highly reliable, energy-efficient communication protocols to ensure sustained performance and scalability under varying conditions. Among the communication technologies that support IoT applications, the Time Slotted Channel Hopping (TSCH) protocol has emerged as a leading solution for robust and deterministic networking [4]. By combining time synchronization and channel hopping, TSCH achieves resilience against interference and multipath fading, making it particularly wellsuited for Industrial Internet of Things (IIoT) deployments. Communication in TSCH is organized into synchronized time Manuscript received June 9, 2025; revised xx, xx. F. Fernando Jurado-Lasso is an independent researcher, based in Cali, Colombia (e-mail: [email protected]). J. F. Jurado is with the Department of Basic Science, Faculty of Engineering and Administration, Universidad Nacional de Colombia Sede Palmira, Palmira 763531, Colombia (e-mail: [email protected]).
slots, with nodes adhering to predefined schedules that determine when to transmit, receive, or remain idle [5], [6]. This deterministic operation minimizes collisions and provides predictable performance, which is essential for industrial applications requiring high reliability and low latency. Despite these advantages, TSCH networks suffer from a key inefficiency known as idle listening—a condition where nodes keep their radios active during receive (Rx) slots without any actual packet transmission. This behavior, often occurring in networks with sporadic or low traffic, leads to unnecessary energy expenditure. In long-lived IoT deployments with battery-powered or energy-harvesting devices, mitigating idle listening is crucial to extending network lifetime and sustaining performance. The proposed solution targets industrial and environmental monitoring systems where both energy efficiency and responsiveness are critical. Examples include process supervision and anomaly detection in production plants, or safety and condition monitoring in warehouse and logistics infrastructures. In such deployments, sensors must operate for years without maintenance, while still being capable of promptly reporting critical events such as temperature deviations, vibration anomalies, or air-quality alerts. Unlike traditional environmental monitoring systems, these applications require a balanced energy-latency tradeoff: energy must be conserved to extend device lifetime, but latency cannot be excessively sacrificed without degrading responsiveness and system reliability. Battery replacement is often costly or infeasible in these settings, and prolonged reporting delays may cause product degradation or safety risks. Therefore, adaptive listening mechanisms that simultaneously achieve low power consumption and reduced latency are essential for next-generation IIoT systems. Although TSCH schedules are deterministic, deciding whether a node should activate its radio in a given reception slot remains challenging due to uncertainty in effective transmission timing caused by retransmissions, traffic, etc. Static heuristics or explicit signaling often lead to conservative listening strategies with increased latency or protocol overhead. To address this challenge, this paper proposes Reinforcement Learning-based Adaptive Slot Listening (RL-ASL), a RL–based robust, deployment-time adaptive listening policy designed for environments where traffic patterns may vary within a known operational range. Responsiveness is achieved through rapid, local runtime decisions based on a pre-trained policy, while energy efficiency is ensured by minimizing idle listening through learned slot-skipping strategies. RLASL dynamically decides whether a node should listen or skip a slot based on learned traffic and scheduling patterns, thereby reducing idle listening without compromising packet
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
reliability. Importantly, the proposed approach is agnostic to how TSCH schedules are generated. RL-ASL can operate on top of both static schedules and adaptive TSCH schedulers that dynamically add, remove, or relocate cells based on traffic demand or network conditions, as it relies only on locally observable slot-level information at runtime. The learning phase is performed offline using extensive simulations that cover multiple representative traffic patterns, and the resulting policies are aggregated into a single generalized Qtable using Federated Learning (FL). This generalized policy enables nodes to adapt their listening behavior at runtime based solely on locally observed state information, without requiring retraining or redeployment when traffic patterns vary within the trained regime. Here, adaptivity refers to per-slot, runtime decision-making based on observed traffic dynamics, rather than online policy retraining or networkwide reconfiguration after deployment. Scenarios involving highly unpredictable or adversarial traffic changes may require complementary mechanisms, such as centralized reconfiguration or online learning, which are outside the scope of this work. Unlike prior work, RL-ASL is fully implemented in C ONTIKI -NG [7] and evaluated on the FIT IoT-LAB [8] testbed, ensuring both realism and reproducibility. We also provide an independent implementation of PRIL-M [9], the closest state-of-the-art adaptive listening protocol. Our experimental results demonstrate that policies trained on simple simulated topologies remain effective when deployed on larger and structurally different real-world networks, including relay nodes experiencing traffic loads not explicitly seen during training. All our code and experimental artifacts are publicly released to foster further research in adaptive listening for TSCH networks.
A. Contributions This paper makes the following key contributions: 1) Proposing RL-ASL: We introduce a novel RL–based adaptive listening protocol for TSCH networks that significantly reduces idle listening and improves energy efficiency. 2) Formal Modeling and Constraints: We formalize the listen-receive and listen-skip decision constraints that govern adaptive listening behavior, establishing a rigorous framework for protocol design. 3) Generalized Offline Learning: We design an offline training methodology based on diverse traffic patterns and FL, yielding a single Q-table that generalizes across varying network conditions without requiring online retraining. 4) Full Implementation in C ONTIKI -NG: We implement RL-ASL entirely within the C ONTIKI -NG operating system, ensuring compatibility with real-world TSCH deployments and open testbeds. 5) Experimental Evaluation on FIT IoT-LAB: We evaluate RL-ASL on the FIT IoT-LAB platform, comparing its performance with state-of-the-art protocols—particularly PRIL-M—across diverse network topologies and traffic patterns.
2
6) Open-Source Release: We publicly release both our RL-ASL and PRIL-M implementations, enabling reproducibility and future comparative research in adaptive listening mechanisms 1 . The remainder of this paper is organized as follows: Section II reviews related work on energy-efficient communication in TSCH networks. Section III presents the network model and discusses idle listening inefficiencies. Section IV details the design of the RL-ASL protocol. Section V describes implementation details and the experimental setup. Section VI outlines the performance metrics and baseline protocols used for evaluation. Section VII presents and discusses the experimental results. Finally, Section VIII concludes the paper and discusses future research directions. II. R ELATED W ORK Energy-efficient and reliable communication remains a core challenge in low-power wireless networks, especially in IIoT settings where long lifetimes and predictable performance are essential [10]. The TSCH protocol has become a leading solution for such systems thanks to its deterministic time-slotting and channel-hopping mechanisms [11]. However, TSCH networks still suffer from energy waste due to idle listening and limited adaptability to dynamic traffic conditions.
A. Autonomous and Centralized Scheduling Mechanisms Early research in TSCH scheduling focused on improving slot allocation to balance reliability and energy efficiency. Orchestra [12] pioneered autonomous scheduling by letting each node independently assign its transmission and reception cells based on network topology. This approach reduced coordination overhead and improved robustness but did not explicitly address the problem of idle listening, as nodes kept their radios active during all assigned Rx slots regardless of actual traffic. Building upon this, Kim et al. proposed ALICE [13], a link-based scheduling protocol where each node allocates cells dynamically according to its neighbors’ traffic direction and updates schedules every slotframe. ALICE improved upon Orchestra in terms of throughput and energy efficiency, yet nodes still incurred energy losses when listening to unused Rx slots. Similarly, centralized and heuristic-based schedulers—such as those by Khan et al. [14] and Papadopoulos et al. [15]—achieved gains through global coordination or guard-time optimization but at the expense of scalability and autonomy. More recent adaptive schedulers such as TESLA [16] and OST [17] introduced elastic slotframe management, dynamically adjusting the slotframe length and cell allocation based on observed traffic intensity. These methods improve throughput and energy efficiency under fluctuating loads but still require the radio to remain active during allocated Rx slots, leaving idle listening largely unaddressed. 1 The code is available at https://github.com/fdojurado/contiki-ng-rl-asl
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
B. Recent Adaptive and Traffic-Aware Scheduling Recent advances have produced more sophisticated adaptive schedulers that monitor network traffic and topology in real time. For example, A3 [18] employs receiver-side traffic estimation to autonomously adjust the number of Tx/Rx cells in the slotframe, achieving notable adaptability and reduced congestion under dynamic traffic. DT-SF [19] follows a similar principle by adjusting slot allocations according to traffic demands and link reliability metrics. Similarly, TA-RPL [20] extends adaptability beyond the MAC layer by integrating traffic-aware routing with TSCH scheduling to optimize endto-end network performance. It leverages the number of allocated transmission cells as an indicator of both link quality and traffic intensity, enabling routing decisions that balance load and improve bandwidth utilization across the network. While these methods substantially improve scheduling agility, they primarily operate at the slotframe or routing level and do not consider fine-grained energy optimization within allocated Rx slots. As a result, even adaptive schedulers may suffer from residual idle listening when nodes expect but do not receive traffic. C. Learning-Based Scheduling Approaches Reinforcement learning (RL) has recently emerged as a powerful tool for adaptive scheduling in constrained IoT environments. Pratama et al. proposed RL-SF [21] and LLQLSF [22], which use Q-learning to optimize cell allocations and minimize latency across distributed nodes. Similarly, ELISE [23] and HRL-TSCH [24] apply hierarchical or modelbased RL to adapt schedules to changing traffic patterns while maintaining deterministic communication. Although these studies demonstrate that RL can outperform heuristic approaches in adapting to dynamic conditions, they focus on slotframe configuration or global scheduling optimization, not on reducing idle listening at the link level. D. Idle Listening Reduction Mechanisms Several mechanisms have been proposed specifically to mitigate idle listening in TSCH networks. Nsabagwa et al. [25] formulated the scheduling problem as a constraint satisfaction problem (CSP) to minimize idle slots in centralized networks, achieving reduced delay and energy consumption but limited scalability. PRIL-F [26] and its successor PRIL [27] introduced proactive radio-sleep coordination based on the ASN, allowing nodes to skip idle Rx slots. More recently, PRIL-M [9] improved robustness to acknowledgment loss and synchronization errors, demonstrating significant energy gains under periodic traffic. However, PRIL-based methods rely on deterministic traffic patterns and shared synchronization, making them less effective in dynamic or bursty environments. Kalita et al. proposed OASA [28], an on-the-fly adaptive scheduler that adjusts slot allocations according to instantaneous traffic conditions, and TACTILE [29], which distributes slots across multiple channels to reduce desynchronization. While these mechanisms reduce energy consumption by adapting slot allocations, their heuristic nature limits long-term adaptability in non-stationary traffic environments.
3
Notably, PRIL and its variants constitute the existing body of work that explicitly targets adaptive unicast idle listening reduction in TSCH networks. Other recent TSCH proposals primarily address scheduling, routing, or slotframe elasticity and are therefore complementary rather than directly comparable to listening control mechanisms such as RL-ASL. E. Positioning of RL-ASL Despite substantial progress, no prior work directly addresses unicast idle listening in distributed TSCH networks using a learning-based approach. Existing adaptive and RLbased schedulers optimize slot allocation or slotframe parameters but leave per-slot radio activation unmanaged. RL-ASL fills this gap through a lightweight, distributed RL agent that learns per-link traffic patterns and decides whether to listen or skip each Rx slot. Unlike rule-based or centralized schemes, it autonomously adapts to stochastic traffic while maintaining near-perfect reliability. Operating at the granularity of individual receive opportunities, RL-ASL complements existing schedulers by enhancing energy efficiency without requiring global coordination or traffic predictability. In summary, previous studies focus on slotframe- or network-level adaptation, leaving per-slot energy management largely unexplored. RL-ASL bridges this gap by learning fine-grained listening behaviors that adapt to dynamic traffic, advancing link-level energy optimization in TSCH networks. III. N ETWORK M ODEL AND I DLE L ISTENING I NEFFICIENCY This section presents the network model adopted in this work and defines the concept of idle listening inefficiency in TSCH-based networks. A. Network Model We consider a TSCH network composed of N nodes operating in a synchronized slotframe of L time slots, indexed from 0 to L − 1, with each slot of duration τ seconds. The network time is represented by the Absolute Slot Number (ASN), which provides a global synchronization reference for slotframe alignment and coordinated transmissions. Each node n listens for incoming unicast transmissions during one or more designated reception slots per slotframe, denoted by the set Lkn ⊆ {0, 1, . . . , L − 1} in slotframe k. Communication occurs between node n and its set of neighbors Mn . The total number of listening activations for node n during an observation window of K consecutive slotframes is given by K X |Lkn |, k=1
which represents the cumulative number of reception opportunities across time. Nodes are categorized according to their functional role in the data collection process:
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
1) Node Classification: • Root node (G): Collects data from all other nodes. • Relay nodes (Y): Forward data toward the root. • Leaf nodes (Z): Generate and transmit data to relay or root nodes. 2) Packet Transmission and Reception: Packets transmitted from node n to node m are denoted as ψn→m ∈ Ψ, while successfully received packets or those detected through control coordination (e.g., broadcast notifications) are represented as ωm→n ∈ Ω. Each transmission or reception event occurs within a specific slot l of slotframe k. 3) Idle Listening and Energy Waste: Idle listening occurs when a node keeps its radio active in a reception slot but no packet is transmitted to it. This leads to unnecessary energy expenditure without contributing to data reception. We distinguish two cases: (i) Unnecessary listening: The node is awake but no transmission is directed to it. (ii) Necessary listening: A transmission is attempted, even if it fails; this is not considered idle. Formally, we define the idle listening indicator for node n in slotframe k, slot l, as: ( k,l 1 if l ∈ Lkn and ∄ψm→n ∈ Ψ, δn (k, l) = (1) 0 otherwise. This definition ensures that only truly idle listening instances are counted as wasted energy, consistent with TSCH operation principles. 4) Traffic Generation Model: Each node generates packets at random time intervals ∆Tn , modeled as ∆Tn ∼ N (µn , σn2 ), where µn and σn represent the mean and standard deviation of packet generation intervals, respectively. The distribution is truncated at zero to ensure positive intervals. This model reflects practical low-power wireless behavior where jitter is intentionally introduced to mitigate collisions and improve channel fairness, as commonly implemented in operating systems such as C ONTIKI -NG [7] and RIOT [30]. 5) Network Topology and Communication Range: Nodes are placed in a two-dimensional plane, with each node n located at coordinates ρn = (xn , yn ), where xn , yn ∈ [xmin , xmax ] × [ymin , ymax ]. The p Euclidean distance between two nodes n and m is dn,m = (xn − xm )2 + (yn − ym )2 . Successful communication occurs when dn,m ≤ dmax , where dmax denotes the maximum communication range. 6) Slotframe and Channel Management: We adopt the Orchestra scheduling framework to coordinate transmissions across multiple slotframes: (i) time synchronization (beacon) slotframe, (ii) broadcast slotframe, and (iii) unicast slotframe. Our focus is on the unicast slotframe, where receiver nodes listen in a single scheduled Rx slot per slotframe (|Lkn | = 1). The network uses C channel offsets, denoted by C = {0, 1, . . . , C − 1}, mapped to physical frequencies through the TSCH hopping sequence to mitigate interference and enable spatial reuse. IV. RL-ASL: RL FOR ADAPTIVE LISTENING The objective of RL-ASL is to minimize the overall energy cost associated with idle listening while maintaining high delivery reliability—a balance that can be viewed as minimizing
4
Fig. 1. Network model for a TSCH network with five nodes.
a total energy cost Jtotal under reliability constraints. RLASL achieves this adaptively through local Q-learning at each receiver, without global coordination. RL-ASL is a fully distributed, receiver-side policy that uses tabular Q-learning to decide, at every scheduled unicast Rx slot, whether a node should listen or skip that slot. The design couples a compact per-neighbor temporal model (Exponential Weighted Moving Average (EWMA) inter-arrival and variance) with a feature-engineered aggregated state, a probabilistic transmission model (distance-to-nearest Gaussian), and an expected-reward computation that trades energy savings against missed receptions. Nodes learn an ε-greedy tabular policy that is updated online (training mode) or loaded as a frozen table (evaluation mode).
A. Neighbor model and notation For node n and each neighbor m ∈ Mn we maintain a compact statistical descriptor 2 exp Θn,m = asnlast n,m , µ̂n,m , σ̂n,m , asnn,m , κn,m , where: asnlast n,m is the ASN at which n last heard m; • µ̂n,m is an EWMA estimate of the inter-arrival (period) between receptions from m (in slots); 2 • σ̂n,m is an EWMA estimate of the variance of those interarrivals; last exp • asnn,m = asnn,m + µ̂n,m is the next expected ASN for m; • κn,m ∈ Z≥0 counts consecutive predicted-but-missed receptions (predicted skips). •
When node n observes a reception from m at ASN ti , the inter-arrival ∆n,m (ti ) = asnn,m (ti ) − asnn,m (ti−1 ) is used to update the EWMA estimates: µ̂n,m ← (1 − λ)µ̂n,m + λ ∆n,m (ti ), 2 2 2 σ̂n,m ← (1 − λ)σ̂n,m + λ ∆n,m (ti ) − µ̂n,m , with smoothing coefficient λ ∈ (0, 1). On a detected missed event we advance the expectation and increment κn,m : exp asnexp n,m ← asnn,m + µ̂n,m ,
κn,m ← κn,m + 1.
These per-neighbor statistics are the primitives used by the probability model and the state encoding below.
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
B. Probability-of-transmission model (distance-to-nearest) At current ASN a node n estimates the probability that at least one child m ∈ Mn will transmit in that slot. Let exp ∆n,m (a) = a − max asnlast n,m , asnn,m . Denote µ̂n,m and σ̂n,m the EWMA mean and standard de2 viation (square root of σ̂n,m ) of inter-arrivals. To ensure numerical robustness we clamp σ̂n,m as: σ̂n,m ← min α µ̂n,m , max(σmin , β µ̂n,m , σ̂n,m ) , with configuration constants α, β ∈ (0, 1) and σmin > 0. Define the phase and distance-to-nearest: ϕn,m (a) = ∆n,m (a) mod µ̂n,m , dn,m (a) = min ϕn,m (a), µ̂n,m − ϕn,m (a) . We model a per-neighbor instantaneous transmission likelihood as a Gaussian kernel on the distance-to-nearest: d2n,m (a) , pn,m (a) ∈ [0, 1]. (2) pn,m (a) = exp − 21 2 σ̂n,m Assuming conditional independence Qacross neighbors, the probability that no child transmits is m∈Mn (1 − pn,m (a)), so the aggregated probability of at least one transmission is Y pn (a) = 1 − 1 − pn,m (a) , (3) m∈Mn
which the implementation clamps to pn (a) ∈ [εp , 1 − εp ] with εp = 10−3 (approx.) to avoid degenerate expectations. This distance-to-nearest model produces a temporal “bump” slightly before and after expected instants, providing permissive behavior under timing uncertainty. C. State encoding (feature engineering and aggregation) To keep the Q-table compact, RL-ASL maps neighborhood statistics to a small discrete state s ∈ S = {0, . . . , S − 1} via feature engineering and mixed-radix encoding. a) Per-neighbor bins.: For node n with neighbors Mn = {m1 , . . . , mMn }, compute for each neighbor ∆ n,mi (a) bi = binB , i = 1, . . . , Mn , µ̂n,mi where binB (·) maps the normalized elapsed inter-arrival to one of B discrete bins. b) Neighborhood aggregates.: From {bi } and {dn,mi (a)} compute: b = round
cshort =
Mn X
Mn 1 X bi , Mn i=1
I(bi < bth ),
i=1
dmin (a) = min dn,m (a), dbin min = binD dmin (a) , m∈Mn X cnear = min Cmax , I(dn,m (a) ≤ σ̂n,m ) . m∈Mn
Here bth , Cmax are configuration constants, and binD (·) maps the minimum distance to one of D discrete bins.
5
c) Mixed-radix encoding.: We form the aggregated state index via a bijection s = fenc b, cshort , dbin min , cnear , where S = B · C1 · D · C2 (with C1 , C2 the discrete ranges of cshort , cnear ). All components are clipped to their declared ranges so s is guaranteed to satisfy 0 ≤ s < S. This compact integer index is the input to the tabular Q-learner.
D. Action space At each scheduled unicast Rx slot (decision epoch) node i selects at ∈ A = {a(0) , a(1) } = {SKIP RX, DO NOT SKIP RX}. The radio activation indicator ξi,t ∈ {0, 1} is determined by the chosen action: ( 0 if at = SKIP RX, ξi,t = 1 if at = DO NOT SKIP RX. E. Reward shaping and expected-reward computation To explicitly account for the cost of missed transmissions, RL-ASL uses an expected-reward formulation that penalizes skipping reception slots proportionally to the estimated likelihood of an incoming transmission. Let the configured reward/penalty parameter set be R = {Rsucc > 0, Rskip > 0, Cidle < 0, Cmiss < 0}, interpreted as success reward, skip reward, idle-listen cost and miss penalty respectively. Given the aggregated transmission probability pi,t ≡ pi (at ) computed by Eq. (3), the expected immediate reward for node i taking action at is: pn (at ) Cmiss + (1 − pn (at )) Rskip , E[ri,t | at ] = (4) p (a ) R n t succ + (1 − pn (at )) Cidle , The implementation uses this expectation as follows: If at = SKIP RX there is no observation in the current slot: the Q-update uses E[ri,t | at ] immediately. • If at = DO NOT SKIP RX the agent listens; upon slot completion the actual outcome is observed. If a packet arrived then ri,t = Rsucc ; otherwise the recomputed pi,t is used and the update target is E[ri,t | at ]. •
Two heuristic adjustments improve stability: 1) Missed-neighbor penalty: if any neighbor j satisfies ∆n,j (a) ≥ µ̂n,j + ζmiss σ̂n,j 2) Near-transmit penalty: if multiple neighbors are near their expected transmit (within one σ̂), the expected reward for skipping is decreased proportionally to the near count (implementation uses a small logged penalty to bias learning away from skipping).
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
F. Tabular Q-learning and learning dynamics Each node i ∈ N maintains a local tabular action–value function
6
Algorithm 1 RL-ASL: Per-node online decision and learning loop
1: for each scheduled Rx slot of node i at ASN a do 2: if not associated or |Mi | = 0 then 3: listen (ξi,t ← 1); Continue Qi : S × A → R, 4: end if 5: extract features and encode state si,t = fenc (·) 6: select ai,t via ε-greedy on Qi (si,t , ·) parameterized by learning rate α ∈ (0, 1], discount factor 7: if ai,t = SKIP RX then compute pi,t via Eq. (3) γ ∈ (0, 1), and exploration rate εe . At each decision epoch t, 8: 9: compute expected reward r̄i,t via Eq. (4) after observing state si,t , selecting action ai,t , and receiving 10: update Qi (si,t , ai,t ) using r̄i,t or estimating the immediate reward ri,t , the Q-table is updated 11: else 12: activate radio (ξi,t ← 1) according to the standard Q-learning rule: 13: if packet received then 14: ri,t ← Rsucc 15: else Qi (si,t , ai,t ) ← Qi (si,t , ai,t ) recompute pi,t ; ri,t ← E[ri,t | ai,t ] h i 16: 17: end if ′ + α ri,t + γ max Q (s , a ) − Q (s , a ) . i i,t+1 i i,t i,t 18: update Qi (si,t , ai,t ) using ri,t a′ ∈A 19: end if (5) 20: bookkeeping (update counters, episode return, decay εe ) 21: end for
Action selection follows an ε-greedy policy: ( arg maxa∈A Qi (si,t , a), w.p. 1 − εe , ai,t = Uniform(A), w.p. εe , where εe decays multiplicatively per episode as εe+1 = max(ε PTe −1min , ηε εe ). Each node maintains episode statistics Ge = t=0 ri,t and a rolling average Groll over a sliding window of W recent episodes. Whenever Groll exceeds its previous maximum, the table may be checkpointed for debugging or persistence. During training, Q-values are updated online; in evaluation mode, the Q-table is frozen and read-only.
G. Runtime decision and learning loop Algorithm 1 summarizes the per-node logic of RL-ASL. At each scheduled unicast reception slot of node i at ASN a: (i) Feature extraction: the node computes per-neighbor features {bn,m (a)}, aggregates them into (b, cshort , dbin min , cnear ), and encodes the resulting state si,t = fenc (·) as in Section IV-C. (ii) Action selection: an action ai,t ∈ A is drawn from the ε-greedy policy on Qi (si,t , ·). (iii) Decision execution: - If ai,t = SKIP RX, the node keeps the radio off, estimates pi,t by Eq. (3), computes the expected reward E[ri,t | ai,t ] by Eq. (4), and applies the Q-update (5) immediately (using si,t+1 = si,t ). - If ai,t = DO NOT SKIP RX, the node listens; upon slot completion it observes whether a packet was received and sets ri,t = Rsucc if packet received; otherwise ri,t = E[ri,t | ai,t ]. before performing the same Q-update. (iv) Bookkeeping: increment the step counter, update Ge and Groll , and decay εe when the episode ends. H. Complexity and robustness considerations All computations are local and lightweight: the Q-table size is |S| · |A| entries, and per-neighbor statistics require a few floating-point scalars. The only nonlinear operations are √ exp(·), ·, and modular arithmetic for the phase calculation. Numerical stability is ensured through bounded counters, clamping of probabilities pi,t ∈ [εp , 1 − εp ], and limiting σ̂n,m as defined in Section IV-B.
I. Discussion on Mobility Support Mobility-induced topology changes, such as those triggered by upper-layer routing protocols (e.g., RPL), can be naturally accommodated by RL-ASL through its 2-H (see Section V-A). Routing updates in RPL-based networks may take several minutes [31]; during this period, affected nodes revert to standard TSCH behavior, listening to all scheduled Rx slots to maintain network connectivity and discover new neighbors. Once routing converges and a new parent is selected, the node reinitiates the 2H process to synchronize with the new receiver. After successful coordination, RL-ASL resumes normal operation, allowing the RL process to gradually adapt to the updated neighborhood statistics. This approach enables RLASL to preserve connectivity and progressively re-optimize performance under moderate mobility without requiring major protocol modifications. J. Routing Layer Interaction and Adaptation RL-ASL operates at the MAC layer and remains agnostic to the routing protocol above it. This modular design eases integration into existing TSCH stacks, as RL-ASL relies only on local per-neighbor statistics (e.g., reception success, interarrival intervals) and the node’s ASN, which are routingindependent. Our evaluation focused on static topologies to isolate MAC-layer adaptation effects. However, routing-layer dynamics—such as parent changes or topology updates in RPL—can transiently modify neighborhood relationships. In such cases, RL-ASL automatically reinitializes its 2-H sequence and local neighbor statistics, enabling rapid re-synchronization without global reconfiguration. Future work could strengthen routing-MAC interaction by introducing cross-layer signaling. For example, routing events (e.g., DAO updates or parent switches) could trigger a temporary exploration phase where RL-ASL prioritizes listening and accelerates learning for new neighbors. This would preserve energy efficiency while enhancing robustness in mobile or time-varying networks.
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
7
K. Summary
V. I MPLEMENTATION AND E XPERIMENTAL S ETUP This section describes the implementation of RL-ASL in C ONTIKI -NG [7], the experimental environment on the FIT IoT-LAB [8] testbed, and the configurations used to evaluate its performance under diverse network and traffic conditions. A. Integration of RL-ASL in Contiki-NG RL-ASL was implemented natively in C ONTIKI -NG (v5.0), an open-source operating system for low-power and lossy networks (LLNs) with built-in support for IEEE 802.15.4, 6LoWPAN, RPL, and the TSCH MAC protocol. This provides a realistic software stack for evaluating learning-based MAClayer mechanisms. The RL-ASL module extends the Orchestra [12] scheduling framework by introducing a RL-driven adaptive listening mechanism. Each node autonomously decides whether to activate or skip its scheduled unicast receive (Rx) slot, reducing idle listening while preserving reliability. At runtime, neighboring nodes coordinate through a lightweight 2-H mechanism that synchronizes the transmitter and receiver before each transmission attempt. The handshake succeeds if ( Rx 1, if lnTx = lm ∧ cn = cm ∧ ackn = 1, 2-Hn→m = 0, otherwise, Rx where lnTx and lm denote the scheduled slot indices of nodes n and m, cn , cm their channel offsets, and ackn the acknowledgment flag. If a node updates its preferred next-hop, the 2-H sequence is reinitiated, allowing the previous receiver to safely skip its next Rx slot. The Q-learning component operates in two phases: 1) Offline training: performed entirely in the Cooja network simulator [32], using diverse traffic patterns to update Q-values and converge to an optimal slot-skipping policy. 2) On-device inference: executed in real hardware using the pre-trained Q-table, stored directly in firmware.
B. Experimental Platform All experiments were conducted on the FIT IoT-LAB testbed—a large-scale open-access facility for experimentation with low-power wireless systems. The deployment utilized the IoT-LAB M3 platform, which integrates an ARM Cortex-M3 microcontroller, an Atmel AT86RF231 IEEE 802.15.4 radio,
13
14
6
Sink
Leaf
15
2
17 4
5
1
7
3
16
3
11
2
Relay 1
12
5
(a)
3.0 2.5 2.0 1.5 1.0 22 0.5 0.0
M3-2 M3-20 M3-4M3-6M3-8 M3-30 M3-22 M3-10 M3-48 M3-32 M3-34 M3-24 M3-12 M3-14 M3-54 M3-50 M3-36 M3-38 M3-26 M3-16 M3-18 M3-56 M3-58 M3-52 M3-40 M3-42 M3-44 M3-28 M3-60 M3-46 M3-62 M3-64
Z (m)
RL-ASL integrates: (i) a probabilistic temporal transmission model, (ii) compact state aggregation via mixed-radix encoding, and (iii) expected-reward-driven tabular Q-learning to balance energy efficiency and reliability on a per-node basis. The fully distributed design requires no message exchange or synchronization across nodes. Heuristics for missed-neighbor and near-transmit situations improve convergence stability and robustness under realistic TSCH dynamics. Algorithm 1 mirrors the embedded implementation and is optimized for memory- and timing-constrained IoT devices.
10
4
21
8
9
18
19
20
(b)
0 1
Sink Relay Leaf M3-1 Unused M3-3 M3-19 M3-5M3-7 Routing Topology M3-29 M3-21 M3-9M3-11 M3-31 M3-47 M3-33 M3-23 M3-13 M3-35 M3-15 M3-53 M3-49 M3-37 M3-25 M3-17 M3-55 M3-57 M3-51 M3-39 M3-41 M3-43 M3-27 M3-59 M3-45 M3-61 M3-63 0.0
2 3 Y (m 4 5 6 ) 7
8
1.00.5 2.01.5 3.02.5 X (m) 4.03.5
(c)
Fig. 2. Experimental network topologies from Strasbourg IoT-LAB site: (a) Simple 5-node topology used for training and validation, (b) Larger star topology, and (c) Deployment of the star topology in the FIT IoT-LAB.
and onboard sensors for temperature, light, and acceleration. This platform provides a representative low-power embedded system with constrained computational and memory resources, while enabling detailed monitoring of energy, radio, and timing metrics at the hardware level. Each node runs C ONTIKI -NG with full TSCH synchronization and fine-grained energy instrumentation, allowing precise measurement of duty cycle, latency, and packet reliability. The controlled experimental environment and large number of available nodes allow reproducible and scalable evaluation of RL-ASL under realistic operating conditions. C. Network Scenarios and Topologies Two representative network topologies were used to assess RL-ASL: • Simple topology: a compact 5-node network comprising one sink, one relay, and three leaf nodes (Fig. 2a). It supports controlled validation of per-slot decision dynamics and convergence behavior. All RL-ASL training was exclusively performed in this topology. • Star topology: a larger multi-hop deployment derived from the Strasbourg IoT-LAB site (Fig. 2b and 2c). Leaf nodes are positioned up to three hops from the sink, forming a hybrid star–mesh structure for scalability and latency evaluation under realistic link variability. This topology only uses the pre-trained, aggregated (FL) Qtable obtained from the simple topology. Nodes operate with standard TSCH parameters: 10 ms timeslots, 0 dBm transmission power, and network-wide slot alignment via the ASN. D. Traffic Patterns To evaluate adaptability, multiple traffic generation patterns were implemented as C ONTIKI -NG processes. Each node periodically generates packets aligned with its TSCH schedule. The bursty traffic behavior evaluated in this work arises from heterogeneous periodic sources rather than from fully random or erratic traffic. In the heterogeneous traffic pattern, nodes generate packets periodically but with different inter-arrival times. When such flows converge at relay nodes, their superposition creates transient congestion and burst-like reception opportunities, especially in multi-hop topologies. This model
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
8
Jittered
Transmission intervals per node ID
High Traffic
✓
Heterogeneous Traffic
✓
Sparse Traffic
✓
Periodic Traffic
✗
All nodes: 13 s IDs 3, 11, 15, 19: 17 s IDs 4, 12, 16, 20: 30 s IDs 5, 6, 13, 17, 21: 50 s IDs 4, 18, 22: 73 s Alternating IDs: 60 or 73 s IDs 3, 11, 15, 19: 17 s IDs 4, 12, 16, 20: 19 s IDs 5, 13, 17, 21: 23 s IDs 4, 18, 22: 29 s
Reward
Pattern
800
1.00
LR=0.15 LR=0.1 LR=0.05
600 400
0.75
Epsilon
TABLE I T RAFFIC GENERATION PATTERNS AND PER - NODE T X INTERVALS .
0.50
200 Epsilon
0 0
200
400
600
Episode
800
1000
0.25
1200
Fig. 3. Convergence of average reward during Q-learning training in the relay topology for different learning rates α.
TABLE II RL-ASL Q- LEARNING CONFIGURATION PARAMETERS . Parameter
Value
Parameter
Value
Episode length (Te ) Learning rate (α) Min. exploration (εmin ) Reward success (Rsucc ) Idle cost (Cidle ) Terminal reward (success) Inter-arrival bin width (B) Short-bin threshold (bth )
500 0.15 0.05 +1.0 –0.5 +5.0 10 2
Discount factor (γ) Initial exploration (ε0 ) Decay factor (ηε ) Skip reward (Rskip ) Miss penalty (Cmiss ) Terminal penalty (failure) Distance bins (D)
0.9 1.0 0.997 +0.5 –1.0 –5.0 4
reflects common industrial monitoring scenarios in which subsets of sensors temporarily report at higher rates due to local events or configuration differences. Table I summarizes the considered traffic modes and per-node transmission intervals. E. Q-learning Configuration and Hyperparameters Table II summarizes the configuration of the RL-ASL tabular Q-learning agent implemented in C ONTIKI -NG. The design prioritizes simplicity and computational efficiency for real-time execution on constrained IoT motes. These parameters yield a discrete state space of Ns = 10 × (3+1) × 4 × (3+1) = 640 states, each associated with two possible actions (listen or skip), resulting in a Q-table of size 640 × 2. The reward structure promotes energy efficiency by encouraging nodes to skip idle receive slots, while penalizing missed receptions and unnecessary listening. F. Training and On-Device Inference Offline training was conducted exclusively in both the simple topology and the Cooja network simulator. The Qlearning process is executed entirely offline, and the resulting Q-table is embedded in flash memory as a fixed decision policy, motivated by the constraints of low-power industrial IoT devices. After deployment, RL-ASL performs no online learning or exploration; runtime behavior is limited to deterministic table lookups based on locally observable state, ensuring predictable execution and compatibility with safetycritical TSCH systems. Each run simulated 108 ms (approximately 27.8 hours) of virtual network time, updating Q-values under an exploration rate ε decaying from 1.0 to 0.1. Learning rates α ∈ {0.15, 0.1, 0.05} were evaluated; α = 0.15 yielded
the most stable convergence (Fig. 3) and was adopted for deployment. To enhance generalization across heterogeneous traffic conditions, a lightweight FL aggregation was applied to the Qtables trained independently for each traffic pattern in the simple topology. Specifically, individual models Qi were combined via a weighted Federated Averaging (FedAvg) step: X Ei Qglobal = wi Qi , wi = P , j Ej i where Ei denotes the number of episodes used to train model i. The resulting global model was then deployed directly on real IoT-LAB nodes for inference in both topologies. The final global Q-table comprises 640 × 2 = 1280 parameters, each stored as a 32-bit floating-point value, resulting in a memory footprint of approximately 5 kB. Given the limited Flash memory available in low-power motes, this representation can be further optimized through fixed-point quantization. For instance, scaling Q-values by a factor of 10 and storing them as 16-bit integers reduces the footprint to 2.5 kB with negligible accuracy loss, since the maximum observed Q-value magnitude (|Q|max ≈ 33.4) comfortably fits within the signed 16-bit range. The global Q-table was compiled into firmware as a static lookup table. Inference reduces to a table lookup plus two integer comparisons, requiring only a few bytes of RAM for indexing. Flash overhead is platform-dependent: under 10 kB on the TI CC2650 (128 kB Flash / 20 kB RAM) and about 5– 6 kB on the IoT-LAB M3 (STM32F103, 512 kB Flash / 64 kB RAM), comfortably within device constraints. 1) Energy Consumption of Training: All training of RLASL was conducted within the Cooja network simulator, which provides cycle-accurate emulation of C ONTIKI -NG nodes with realistic radio behavior and timing, ensuring reproducibility without real energy expenditure. Training directly on embedded hardware is infeasible for low-power IoT networks, as RL typically requires thousands or millions of interaction steps to converge—resulting in excessive training time and energy use from repeated transmissions and updates. In contrast, simulation in Cooja network simulator enables accelerated training under controlled yet realistic conditions, faithfully modeling network dynamics and interference. Once converged, only the compact pre-trained Q-table is deployed on the motes. During operation, inference consists solely of a table lookup, incurring negligible computational and energy cost, as demonstrated in Section VII.
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
TABLE III S ENSITIVITY OF RL-ASL TO THE REWARD PARAMETER Rskip IN THE SIMPLE TOPOLOGY. Rskip
Latency [ms]
PDR [%]
RDC [%]
0.25 0.5 0.75
202.37 196.38 198.70
99.91 99.97 99.97
1.35 1.32 1.33
9
B. Power and Energy Consumption The average current consumption (κ) of each node is computed using the per-state duty cycle reported by C ONTIKI NG’s built-in Energest module, which tracks time spent in transmit, receive, idle, and low-power states. For node n at time t, X κn (t) = Dn,s (t)Is , (8) s∈S
G. Sensitivity Analysis Table III illustrates the sensitivity of RL-ASL to the reward parameter Rskip in the simple topology. Across the tested range, all three performance metrics exhibit only minor variations. In particular, the PDR remains consistently above 99%, varying between 99.91% and 99.97%. The end-to-end latency ranges from 196.38 ms to 202.37 ms, corresponding to a difference of about 6 ms across all configurations, while the RDC remains low and stable between 1.32% and 1.35%. These results indicate that RL-ASL is largely insensitive to moderate changes in reward weighting and does not rely on carefully tuned values of Rskip to maintain reliable and energyefficient operation within a given deployment scenario. This insensitivity reduces the need for precise reward calibration during offline training and supports the practical deployability of RL-ASL. VI. P ERFORMANCE M ETRICS AND BASELINE P ROTOCOLS This section defines the performance metrics used to evaluate RL-ASL and describes the baseline protocols employed for comparison. The evaluation focuses on quantifying the protocol’s ability to improve reliability, latency, and energy efficiency under varying traffic and topology conditions. A. Performance Metrics We consider four key performance metrics to assess RLASL: PDR, latency, and power consumption. Together, these metrics provide a comprehensive view of the trade-offs between reliability, responsiveness, and energy efficiency. 1) PDR: The PDR quantifies link reliability and is defined as the ratio of successfully received packets to the total packets transmitted. For a unicast link from node n to node m, |Ωn→m | PDRn→m = |Ψ , where Ωn→m and Ψn→m denote the n→m | sets of successfully received and transmitted packets, respectively. The network-wide PDR is computed as the average over all leaf nodes: P |Ωn→G | PDR = Pn∈Z , (6) n∈Z |Ψn→G | where Z is the set of leaf nodes and G denotes the sink. 2) Latency: Latency measures the end-to-end delay experienced by packets. For P a unicast plink np → m, it isp defined p 1 as: Dn→m = |Ωn→m p∈Ωn→m (trx − ttx ), where ttx and trx | represent the transmission and reception timestamps of packet p. The network-wide average latency is then computed as: P P (tprx − tptx ) n∈N P p∈Ωn→G D= . (7) n∈N |Ωn→G |
where Dn,s (t) is the fraction of time in state s and Is is the corresponding current draw. We use the current consumption values specified in the M3 platform datasheet [8]: ICPU = 14.0 mA, ILPM = 0.014 mA, IDEEP LPM = 0.002 mA, ITX = 11.6 mA, and IRX = 12.3 mA, with a supply voltage of V = 3.3 V. These parameters are used consistently to compute all energy-related metrics presented in Section VII. C. Baseline Protocols To contextualize the performance of RL-ASL, we compare it against three representative TSCH scheduling approaches spanning autonomous, link-based, and adaptive paradigms: • Orchestra (Orch.) [12]: A reference autonomous scheduling framework that assigns timeslots based on node roles. We use the receiver-based Orchestra mode with default parameters, which provides a fair baseline for static listening schedules. • Link-based Scheduling (Orch.-LB) [13], [33]: A deterministic link-centric scheduler that allocates timeslots according to link quality metrics. This baseline highlights the gains of RL compared to structured yet non-adaptive scheduling. • PRIL-M [9]: A recent probabilistic adaptive listening protocol that dynamically adjusts reception windows based on traffic. PRIL-M represents the current state of the art in adaptive TSCH listening. We implement RL-ASL as an add-on service that operates on top of the Orchestra and link-based scheduling frameworks—denoted RL-ASL and RL-ASL-LB. As a non-invasive layer, the service dynamically skips or activates receive slots according to the learned policy while preserving the underlying timeslot allocation and coordination logic of the base schedulers. Note that PRIL-M is evaluated only under periodic traffic conditions, as its design relies on piggybacking deterministic next-transmission timing information in packet headers. This mechanism assumes strictly periodic traffic and is therefore not applicable to heterogeneous or non-periodic workloads, for which its underlying assumptions no longer hold. VII. R ESULTS AND D ISCUSSION Experiments were conducted on the FIT IoT-LAB testbed using two network configurations: a simple linear topology and a complex star topology. For each topology, performance was assessed under four traffic patterns—high, heterogeneous, sparse, and periodic—as described in Section V. It is worth noting that RL-ASL is designed for static and low-mobility
80
ASL
B
L-L
RL-
AS RL-
h.
100
LB
Orc
h.Orc
90 80
ASL
RL-
B
L-L
AS RL-
(a)
h.
Orc
100
.-LB
PDR [%]
90
10
PDR [%]
100
PDR [%]
PDR [%]
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
90 80
. B L-L Orch
ASL
RL-
h Orc
AS RL-
(b)
100
.-LB
90 80
L -LB rch. ch.-LB IL-M L-AS ASL O R Or RL-
h Orc
PR
(c)
(d)
10 1 ASL RL-
-LB ch.-LB L-ASL R Or
100 10 1
h.
Orc
RL-
-LB .-LB L-ASL ASL Orch R
(a)
101
Latency [s]
100
101
Latency [s]
101
Latency [s]
Latency [s]
Fig. 4. Comparison of PDR across different traffic patterns for the simple topology: (a-d) PDR under high, heterogeneous, sparse, and periodic traffic patterns.
100 10 1
h.
Orc
-LB ch.-LB L-ASL R Or
ASL
RL-
(b)
101 100 10 1
h.
-LB .-LB ASL Orch. RIL-M P ASL Orch RL-
Orc
RL-
(c)
(d)
0.8
0.6 0
SL
A RL-
RX1.4 UC idle
LB Orch.
L-AS
RL
1.4
B h.-L Orc
1.0
0.8
0.5
0.5 0.0
SL
A RL-
(a)
B
L-L -AS
RL
h.
Orc
1.4
1.4
1.0
LB
1.0 0.5 0.0
SL
h.Orc
A RL-
(b)
1.0
0.7
0.5
LB Orch.
L-AS
RL
Power [mW]
RX non-UC RX UC active 1.0
Power [mW]
CPU TX 1
Power [mW]
Power [mW]
Fig. 5. Latency distributions across different traffic patterns for the simple topology: (a-d) Latency under high, heterogeneous, sparse, and periodic traffic patterns.
LB
1.0
0.5
0.5 0.0
0.6
0.8
1.0
L -LB rch. ch.-LB IL-M L-AS ASL O R Or L R
h.Orc
PR
(c)
(d)
10 1
0.75 1.00 1.25
Power [mW] (a)
80
10
90
0
10 1 0.50 0.75 1.00 1.25
85 80
Power [mW] (b)
100
101 10
95 90
0
10 1 0.50 0.75 1.00 1.25
Power [mW] (c)
85 80
101 10
0
RL-ASL RL-ASL-LB Orch. Orch.-LB PRIL-M
10 1 0.50 0.75 1.00 1.25
100 95 90 85
PDR [%]
85
95
PDR [%]
90
100
101
Latency [s]
95
PDR [%]
100
Latency [s]
10
0
RL-ASL RL-ASL-LB Orch. Orch.-LB
PDR [%]
101
Latency [s]
Latency [s]
Fig. 6. Power consumption breakdown across different traffic patterns for the simple topology: (a-d) Power consumption under high, heterogeneous, sparse, and periodic traffic patterns.
80
Power [mW] (d)
Fig. 7. Trade-off comparison across different traffic patterns for the simple topology: (a-d) Overall performance under high, heterogeneous, sparse, and periodic traffic patterns. The metrics are normalized for comparison.
TSCH deployments, where network topology remains stable for sufficiently long periods to allow receiver-side listening policies to be learned and applied effectively. Such conditions are typical in many industrial, environmental monitoring, and smart infrastructure scenarios. In highly dynamic networks with continuous or fast mobility, frequent re-synchronization and parent changes may limit the time during which adaptive listening can be exploited. To ensure statistical robustness, each protocol was executed three times each lasting 1 hour, and results are reported as the average across all runs. Error bars in all figures represent 95% confidence intervals. We analyze results for each topology and traffic pattern, focusing on key performance metrics: PDR, latency, power consumption, and overall performance (via radar charts).
A. Simple Topology We first analyze the results obtained from the simple topology depicted in Fig. 2a, evaluated under the four traffic patterns. Fig. 4, 5, 6, and 7 summarize the performance across PDR, latency, power consumption, and overall tradeoffs, respectively. 1) Reliability: Fig. 4a-4d show that all protocols achieve near-perfect reliability, with network-wide PDRs close to 100%. PRIL-M achieves a slightly lower PDR (99.5%) in the periodic scenario. In contrast, both RL-ASL and RL-ASLLB maintain a near-perfect PDR across all traffic patterns, demonstrating that adaptive listening does not compromise reliability. This confirms the robustness of our RL design, which effectively balances exploration and exploitation while maintaining link stability.
80
B
L-L
AS RL-
h.
Orc
(a)
LB
h.Orc
ASL
RL-
100 90 80
ASL
RL-
B
L-L
AS RL-
h.
Orc
.-LB
h Orc
(b)
100
PDR [%]
90
11
PDR [%]
100
PDR [%]
PDR [%]
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
90 80
.-LB
h Orc
ASL
RL-
. B L-L Orch
AS RL-
100 90 80
L -LB rch. ch.-LB IL-M L-AS ASL O R Or RL-
PR
(c)
(d)
Fig. 8. Comparison of PDR across different traffic patterns for the Star topology: (a-d) PDR under high, heterogeneous, sparse, and periodic traffic patterns.
2) Latency: Fig. 5a-5d illustrate latency distributions. The main latency differences stem from the underlying TSCH scheduling model. Receiver-based protocols (Orchestra and RL-ASL) exhibit slightly higher delays due to having only one Rx cell per slotframe, which increases the average waiting time. In contrast, link-based variants (Orchestra-LB and RLASL-LB) achieve lower latency since multiple per-link Rx cells reduce queuing delay and enable faster transmissions. Notably, RL-ASL and RL-ASL-LB achieve latency equal to or slightly lower than their Orchestra counterparts despite employing adaptive listening. This confirms that learningbased slot skipping does not disrupt scheduling continuity. RLASL reduces average latency by up to 96% compared to PRILM, while RL-ASL-LB achieves up to 98% latency reduction relative to PRIL-M. Overall, RL-ASL maintains low latency even under heterogeneous and sparse traffic, showcasing its adaptability to non-periodic workloads. 3) Energy Efficiency: Fig. 6a-6d present the power consumption breakdown across CPU, transmission, and reception states. RL-ASL reduces total power consumption on average by over 44% compared to Orchestra across all traffic patterns, validating the effectiveness of adaptive listening in minimizing idle-listening energy waste. Similarly, RL-ASL-LB achieves over 46% lower power consumption than OrchestraLB, demonstrating that per-link scheduling combined with RLdriven listening optimization yields substantial energy savings. The most significant savings occur in the Rx unicast idle state, as adaptive listening effectively suppresses unnecessary listening intervals while preserving delivery performance. Importantly, CPU power remains nearly constant across all protocols, indicating that RL-ASL’s online inference introduces negligible computational overhead—an essential property for embedded devices. Transmission power also remains comparable, confirming that reduced listening does not incur retransmission overhead. PRIL-M achieves slightly lower power in the periodic case due to its traffic awareness; however, RL-ASL matches this efficiency while supporting arbitrary, non-periodic traffic patterns. 4) Overall Performance: Fig. 7a-7d present the overall performance trade-offs involving PDR, latency, and power consumption. Both RL-ASL and RL-ASL-LB consistently achieve the best balance across all metrics and traffic patterns. RL-ASL offers the most energy-efficient operation with slightly higher latency, while RL-ASL-LB provides lower latency at a modest energy cost. Compared to PRIL-M, RLASL delivers comparable power efficiency under periodic traffic but provides lower latency and generalizes effectively to non-periodic and heterogeneous workloads—an essential
capability for real-world IoT deployments.
B. Star Topology We next evaluate the star topology (Fig. 2b), which introduces higher contention and asymmetric traffic. The results, shown in Fig. 8-11, confirm the scalability and robustness of RL-ASL under more complex conditions. 1) Reliability and Latency: Fig. 8a-8d show that RL-ASL and RL-ASL-LB sustain a 100% PDR under all traffic patterns, while PRIL-M’s PDR drops to 83%. This demonstrates the robustness of RL-ASL in handling contention and asymmetric traffic without sacrificing reliability. Latency results (Figs. 9a-9d) further emphasize the benefits of RL. RL-ASL reduces average latency by up to 87% compared to PRIL-M, while RL-ASL-LB achieves reductions of up to 95%. Latency remains stable despite increased contention in the star topology, indicating that RL-ASL effectively manages slot skipping without introducing additional delay. 2) Energy Efficiency and Multi-Metric Trade-offs: Fig. 10a10d show that RL-ASL maintains the lowest total power consumption among all protocols, outperforming PRIL-M even in the periodic traffic scenario. This highlights that RLASL adapts to both temporal and spatial dynamics rather than relying on static periodicity assumptions. Reductions in idle listening power exceed 35% compared to Orchestra and reach up to 56% compared to Orchestra-LB, confirming the effectiveness of adaptive listening in complex topologies. RLASL-LB achieves similar savings, demonstrating that per-link scheduling combined with RL-based listening optimization remains effective under high contention. Finally, Fig. 11a–11d show that both RL-ASL and RL-ASLLB consistently achieve the best overall performance across all evaluated metrics and traffic patterns in the star topology. RL-ASL yields the most energy-efficient operation, at the cost of a slight increase in end-to-end latency, whereas RL-ASLLB shows lower latency with a modest increase in energy consumption. In comparison to PRIL-M, RL-ASL demonstrates substantially improved energy efficiency, lower latency, and higher reliability in the periodic traffic scenario—which represents the most favorable operating condition for PRILM—while also maintaining robust performance under nonperiodic and heterogeneous traffic workloads. These results indicate that, despite being trained on a simple topology, the learned policy generalizes effectively and adapts to more complex network structures and traffic conditions that were not seen during training.
-LB
SL
A RL-
.-LB
h
Orc
ASL RL-
101 100
h. Orc
A RL-
-LB
SL
.-LB
h Orc
(a)
ASL
RL-
101
Latency [s]
100
12
Latency [s]
101
Latency [s]
Latency [s]
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
100
. rch
-LB
SL
O
A RL-
h
Orc
(b)
.-LB
. rch
ASL
RL-
101 100 10 1
O
-M SL LB LB ch. SL- Orch.- Or RL-A PRIL
A RL-
(c)
(d)
0.9
0.7
0
SL
A RL-
RX1.5 UC idle
LB Orch.
B h.-L Orc
L-AS
RL
1.4
1.5 1.0 0.5 0.0
SL
A RL-
(a)
1.0
0.8
0.6
B
L-L -AS
RL
h.
Orc
1.4
1.5
LB
1.0 0.5 0.0
h.Orc
1.0
0.8
0.6
SL
A RL-
LB Orch.
L-AS
RL
(b)
1.5
1.5
Power [mW]
RX non-UC RX UC active 1.0
Power [mW]
1
CPU TX
Power [mW]
Power [mW]
Fig. 9. Latency distributions across different traffic patterns for the Star topology: (a-d) Latency under high, heterogeneous, sparse, and periodic traffic patterns.
LB
1.0
0.0
0.9
0.8
0.6
0.5
1.0
. ASL RIL-M SL-LB Orch rch.-LB P A RLO L R
h.Orc
(c)
(d)
0.75
1.00
1.25
Power [mW] (a)
1.50
80
90
100
85
0.75
1.00
1.25
80
Power [mW] (b)
100
101
95 90
100
85
0.75
1.00
1.25
101
RL-ASL RL-ASL-LB Orch. Orch.-LB PRIL-M
100
80
0.75
Power [mW] (c)
1.00
100 95 90 85
PDR [%]
85
95
PDR [%]
90
100
101
Latency [s]
95
PDR [%]
100
100
Latency [s]
RL-ASL RL-ASL-LB Orch. Orch.-LB
PDR [%]
101
Latency [s]
Latency [s]
Fig. 10. Power consumption across different traffic patterns for the Star topology: (a-d) Power consumption under high, heterogeneous, sparse, and periodic traffic patterns.
80
1.25
Power [mW] (d)
Fig. 11. Trade-off comparison across different traffic patterns for the Star topology: (a-d) Overall performance under high, heterogeneous, sparse, and periodic traffic patterns. The metrics are normalized for comparison.
C. Practical Energy Impact To assess the practical benefit of RL-ASL, we estimate the expected battery Lifetime (LT) from the measured average power consumption. Assuming a 3 V supply and a 220 mAh coin-cell battery (Ebatt ≈ 2.38 kJ), the expected LT in days is: Ebatt L= , where Ebatt = 3 × 220 × 3.6 = 2376 J. Pavg × 86400 Table IV summarizes the measured average power and the corresponding LT for both network topologies. In the simple topology, RL-ASL consumes only 0.56 mW, extending LT to about 174 days—a 77% gain over Orchestra and 180% over Orchestra-LB. Even the more responsive RL-ASL-LB maintains 126 days, still outperforming static scheduling. PRIL-M achieves the highest efficiency under strictly periodic traffic (0.53 mW, 184 days), while RL-ASL generalizes better to non-periodic workloads. In the star topology, RL-ASL (0.65 mW) achieves an estimated 150 days, improving LT by roughly 55% relative to Orchestra (1.01 mW, 97 days) and more than 120% over Orchestra-LB (1.45 mW, 68 days). RL-ASL-LB (0.86 mW) sustains about 115 days, balancing adaptivity and energy use. For larger nodes powered by two AA batteries (≈ 21.6 kJ), RL-ASL could extend LT from roughly 2.4 years (Orchestra)
TABLE IV AVERAGE P OWER AND E STIMATED L IFETIME (3 V, 220 M A H BATTERY ) Protocol Orchestra Orchestra-LB PRIL-M RL-ASL RL-ASL-LB
Simple [mW]
LT [days]
Star [mW]
LT [days]
0.996 1.419 0.529 0.561 0.789
99 70 185 174 124
1.008 1.453 0.751 0.647 0.856
98 68 131 152 114
TABLE V I MPACT OF M ODERATE M OBILITY ON RL-ASL P ERFORMANCE Protocol RL-ASL RPL Baseline
PDR [%]
Radio Duty Cycle [%]
Latency [ms]
93.1 96.7
2.8 3.4
186.97 184.70
to 4.2 years, with RL-ASL-LB sustaining around 3 years—a notable advantage for long-lived IIoT deployments. D. Impact of Moderate Mobility on RL-ASL To assess the behavior of RL-ASL under mobility, we conducted a deliberately simple RPL-based simulation using the Cooja network simulator. The scenario involves a single mobile node moving at a low speed (0.2 m/s) and alternately
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
attaching to one of two potential parents in a tree topology. This experiment uses the heterogeneous traffic pattern described earlier and is intended as a controlled stress test rather than a comprehensive mobility evaluation. As expected, mobility induces repeated parent changes and receiver re-synchronization phases. During these intervals, nodes temporarily revert to standard TSCH operation, limiting the applicability of receiver-side listening optimization. Under these conditions, the RPL baseline achieves a PDR of approximately 96%, whereas RL-ASL attains 93%. Despite this reduction in delivery ratio, RL-ASL significantly reduces the radio duty cycle from 3.4% to 2.8%, corresponding to an energy saving of approximately 18%. This result indicates that RL-ASL continues to reduce idle listening whenever short periods of topology stability are present, even under moderate mobility. E. Summary of Findings Across all experiments, RL-ASL demonstrates consistent superiority in energy efficiency while maintaining perfect reliability and competitive latency. RL-ASL-LB offers slightly lower delays due to its per-link scheduling structure, while RL-ASL provides the best overall energy-delay trade-off. Compared with PRIL-M, RL-ASL achieves comparable performance under periodic traffic but generalizes effectively to non-periodic and heterogeneous scenarios—an essential capability for real-world IoT deployments. A targeted mobility stress test further shows that, while frequent parent changes reduce delivery ratio, RL-ASL continues to achieve substantial idle-listening reductions whenever short periods of topology stability are present. Overall, these results validate the proposed RL-ASL framework as a robust, adaptive, and energy-efficient solution for dynamic TSCH networks, combining the reliability of deterministic scheduling with the flexibility of RL. VIII. C ONCLUSION This work presented RL-ASL, a RL–driven adaptive listening framework for TSCH networks that enhances energy efficiency without compromising reliability or latency. By integrating learning-based slot skipping into standard TSCH scheduling, RL-ASL enables nodes to adapt their listening behavior to traffic dynamics, achieving significant reductions in idle-listening power while preserving synchronization and delivery guarantees. Experimental results from the FIT IoTLAB testbed and Cooja network simulator show that RLASL consistently outperforms state-of-the-art protocols such as Orchestra and PRIL-M across multiple topologies and traffic patterns. It reduces power consumption by up to 46%, maintains near-perfect reliability, and lowers average latency by up to 96% compared to PRIL-M. The link-based variant, RLASL-LB, further improves delay under contention, confirming the scalability of the proposed learning framework. Model training is performed entirely in simulation, while inference on real motes is reduced to a simple table lookup with negligible computational and energy overhead. This makes RL-ASL practical for deployment in low-power embedded networks,
13
bridging the gap between simulation-based learning and realworld operation. A targeted mobility stress test further indicates that, although frequent parent changes reduce delivery ratio, RL-ASL continues to deliver substantial idle-listening reductions whenever short periods of topology stability are present. In summary, RL-ASL demonstrates that RL can be effectively integrated into TSCH scheduling to deliver a robust, adaptive, and energy-aware communication framework for next-generation IoT networks. Future extensions of RL-ASL will explore tighter integration with mobility-aware routing and scheduling mechanisms, enabling coordinated adaptation of parent selection and receiver listening behavior to further enhance performance in mobile and dynamic environments. ACKNOWLEDGMENTS The authors acknowledge the use of AI-based tools for minor language refinements, such as grammar, structure, formatting, and spelling, during manuscript preparation. All intellectual contributions, technical content, and interpretations are solely those of the authors. R EFERENCES [1] D. C. Nguyen, M. Ding, P. N. Pathirana, A. Seneviratne, J. Li, D. Niyato, O. Dobre, and H. V. Poor, “6G Internet of Things: A Comprehensive Survey,” IEEE Internet Things J., vol. 9, no. 1, pp. 359–383, Jan. 2022. [2] M. Soori, B. Arezoo, and R. Dastres, “Internet of things for smart factories in industry 4.0, a review,” IOTCPS, vol. 3, pp. 192–204, Jan. 2023. [3] F. F. Jurado-Lasso, L. Marchegiani, J. F. Jurado, A. M. Abu-Mahfouz, and X. Fafoutis, “A Survey on Machine Learning Software-Defined Wireless Sensor Networks (ML-SDWSNs): Current Status and Major Challenges,” IEEE Access, vol. 10, pp. 23 560–23 592, Feb. 2022. [4] T. Watteyne, M. R. Palattella, and L. A. Grieco, “Using IEEE 802.15.4e Time-Slotted Channel Hopping (TSCH) in the Internet of Things (IoT): Problem Statement,” IETF, Tech. Rep. RFC 7554, May 2015. [5] A. R. Urke, Ø. Kure, and K. Øvsthus, “A Survey of 802.15.4 TSCH Schedulers for a Standardized Industrial Internet of Things,” Sensors, vol. 22, no. 15, pp. 1–34, Dec. 2021. [6] Y. Zhang, H. Huang, Q. Huang, and Y. Han, “6TiSCH IIoT network: A review,” Comput. Netw., vol. 254, p. 110759, Dec. 2024. [7] G. Oikonomou, S. Duquennoy, A. Elsts, J. Eriksson, Y. Tanaka, and N. Tsiftes, “The Contiki-NG open source operating system for next generation IoT devices,” SoftwareX, vol. 18, p. 101089, Jun. 2022. [8] C. Adjih, E. Baccelli, E. Fleury, G. Harter, N. Mitton, T. Noel, R. Pissard-Gibollet, F. Saint-Marcel, G. Schreiner, J. Vandaele, and T. Watteyne, “FIT IoT-LAB: A large scale open experimental IoT testbed,” in IEEE WF-IOT 2015, Milan, Italy, Dec. 2015, pp. 459–464. [9] S. Scanzio, F. Quarta, G. Paolini, G. Formis, and G. Cena, “Ultralow Power and Green TSCH-Based WSNs With Proactive Reduction of Idle Listening,” IEEE Internet Things J., vol. 11, no. 17, pp. 29 076–29 088, Sep. 2024. [10] R. S. Rathore, O. Kaiwartya, K. N. Qureshi, I. T. Javed, W. Nagmeldin, A. Abdelmaboud, and N. Crespi, “Towards Enabling Fault Tolerance and Reliable Green Communications in Next-Generation Wireless Systems,” Appl. Sci., vol. 12, no. 17, p. 8870, Sep. 2022. [11] A. Tabouche, B. Djamaa, and M. R. Senouci, “Traffic-Aware Reliable Scheduling in TSCH Networks for Industry 4.0: A Systematic Mapping Review,” IEEE Commun. Surveys Tuts., vol. 25, no. 4, pp. 2834–2861, Aug. 2023. [12] S. Duquennoy, B. Al Nahas, O. Landsiedel, and T. Watteyne, “Orchestra: Robust Mesh Networks Through Autonomously Scheduled TSCH,” in SenSys ‘15. Seoul, South Korea: ACM, Nov. 2015, pp. 337–350. [13] S. Kim, H.-S. Kim, and C. Kim, “ALICE: Autonomous link-based cell scheduling for TSCH,” in IPSN ’19. Montreal Quebec Canada: ACM, Apr. 2019, pp. 121–132. [14] M. Ojo, S. Giordano, G. Portaluri, D. Adami, and M. Pagano, “An energy efficient centralized scheduling scheme in TSCH networks,” in IEEE ICC Workshops. Paris, France: IEEE, 2017, pp. 570–575.
JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, JUNE 2025
[15] G. Z. Papadopoulos, A. Mavromatis, X. Fafoutis, R. Piechocki, T. Tryfonas, and G. Oikonomou, “Guard Time Optimisation for Energy Efficiency in IEEE 802.15.4-2015 TSCH Links,” in SaSeIoT 2016, InterIoT 2016. Paris, France: Springer, Feb. 2017, pp. 56–63. [16] S. Jeong, J. Paek, H.-S. Kim, and S. Bahk, “TESLA: Traffic-Aware Elastic Slotframe Adjustment in TSCH Networks,” IEEE Access, vol. 7, pp. 130 468–130 483, Sep. 2019. [17] S. Jeong, H.-S. Kim, J. Paek, and S. Bahk, “OST: On-Demand TSCH Scheduling with Traffic-Awareness,” in IEEE INFOCOM. Toronto, ON, Canada: IEEE, Aug. 2020, pp. 69–78. [18] S. Kim, H.-S. Kim, and C.-k. Kim, “A3: Adaptive Autonomous Allocation of TSCH Slots,” in IPSN ’21. New York, NY, USA: ACM, May 2021, pp. 299–314. [19] O. Tavallaie, J. Taheri, and A. Y. Zomaya, “Design and Optimization of Traffic-Aware TSCH Scheduling for Mobile 6TiSCH Networks,” in IoTDI ’21. New York, NY, USA: ACM, May 2021, pp. 234–246. [20] Y. Ha and S.-H. Chung, “Traffic-Aware 6TiSCH Routing Method for IIoT Wireless Networks,” IEEE Internet Things J., vol. 9, no. 22, pp. 22 709–22 722, Nov. 2022. [21] Y. H. Pratama and S. Chung, “RL-SF: Reinforcement Learning based Scheduling Function for Distributed TSCH Networks,” in IEEE ICEIEC. Beijing, China: IEEE, Jul. 2022, pp. 5–8. [22] Y. H. Pratama, S.-H. Chung, and D. Z. Fawwaz, “Low-Latency and QLearning-Based Distributed Scheduling Function for Dynamic 6TiSCH Networks,” IEEE Access, vol. 12, pp. 49 694–49 707, Apr. 2024. [23] F. F. Jurado-Lasso, M. Barzegaran, J. F. Jurado, and X. Fafoutis, “ELISE: A Reinforcement Learning Framework to Optimize the Slotframe Size of the TSCH Protocol in IoT Networks,” IEEE Syst. J., vol. 18, no. 2, pp. 1068–1079, Jun. 2024. [24] F. F. Jurado-Lasso, C. Orfanidis, JF. Jurado, and X. Fafoutis, “HRLTSCH: A Hierarchical Reinforcement Learning-Based TSCH Scheduler for IIoT,” IEEE Trans. Cogn. Commun. Netw., vol. 10, no. 6, pp. 2102– 2118, Dec. 2024. [25] M. Nsabagwa, J. Muhumuza, R. Kasumba, J. S. Otim, and R. Akol, “Minimal Idle-Listen Centralized Scheduling in TSCH Wireless Sensor Networks,” in TSP. Athens, Greece: IEEE, Aug. 2018, pp. 1–5. [26] S. Scanzio, G. Cena, A. Valenzano, and C. Zunino, “Energy Saving in TSCH Networks by Means of Proactive Reduction of Idle Listening,” in ADHOC-NOW. Bari, Italy: Springer, Oct. 2020, pp. 131–144. [27] S. Scanzio, G. Cena, and A. Valenzano, “Enhanced Energy-Saving Mechanisms in TSCH Networks for the IIoT: The PRIL Approach,” IEEE Trans. Ind. Inform., vol. 19, no. 6, pp. 7445–7455, Jun. 2023. [28] A. Kalita and M. Gurusamy, “On-the-Fly Autonomous Slot Allocation in 6TiSCH-Based Industrial IoT Networks,” IEEE Trans. Ind. Inform., vol. 20, no. 7, pp. 9365–9374, Jul. 2024. [29] A. Kalita and M. Khatua, “Autonomous Allocation and Scheduling of Minimal Cell in 6TiSCH Network,” IEEE Internet Things J., vol. 8, no. 15, pp. 12 242–12 250, Aug. 2021. [30] E. Baccelli, O. Hahm, M. Günes, M. Wählisch, and T. C. Schmidt, “RIOT OS: Towards an OS for the Internet of Things,” in INFOCOM WKSHPS, Turin, Italy, Apr. 2013, pp. 79–80. [31] C. Vallati, S. Brienza, G. Anastasi, and S. K. Das, “Improving Network Formation in 6TiSCH Networks,” IEEE Trans. Mob. Comput., vol. 18, no. 1, pp. 98–110, Jan. 2019. [32] F. Osterlind, A. Dunkels, J. Eriksson, N. Finne, and T. Voigt, “CrossLevel Sensor Network Simulation with COOJA,” in IEEE LCN 2006, Tampa, FL, USA, Nov. 2006, pp. 641–648. [33] A. Elsts, S. Kim, H.-S. Kim, and C. Kim, “An Empirical Survey of Autonomous Scheduling Methods for TSCH,” IEEE Access, vol. 8, pp. 67 147–67 165, Mar. 2020.
F. Fernando Jurado-Lasso (GS’18–M’21) received the Ph.D. degree in engineering and the M.Eng. degree in telecommunications engineering from The University of Melbourne, Melbourne, VIC, Australia, in 2020 and 2015, respectively, and the B.Eng. degree in electronics engineering from Universidad del Valle, Cali, Colombia, in 2012. His research focuses on intelligent networked embedded systems, machine learning for low-power and time-synchronized wireless networks, crosslayer scheduling and resource optimization, and Internet of Things (IoT) protocols and architectures.
14
J. F. Jurado received the doctoral degree and M.Sc. degree in physics from Universidad del Valle, Cali, Colombia, in 2000 and 1986, respectively, and the B.Sc. degree in physics from Universidad de Nariño, Pasto, Colombia, in 1984. He is currently a Professor with the Department of Basic Sciences, Faculty of Engineering and Administration, Universidad Nacional de Colombia, Palmira, Colombia. His research interests include nanomaterials, magnetic and ionic materials, nanoelectronics, embedded systems, and the Internet of Things (IoT). He is a Senior Researcher recognized by Minciencias, Colombia, and has been designated as an Emeritus Researcher by Minciencias.