ConceptioArchivearXiv CS
arXiv CSopen access

PoHAR: Understanding Hyperlocal Human Activities with Pollution Sensor Networks

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

PoHAR: Understanding Hyperlocal Human Activities with Pollution Sensor Networks Prasenjit Karmakar, Karthik Reddy, Sandip Chakraborty

arXiv:2605.09434v1 [cs.DC] 10 May 2026

Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur, India. {prasenjitkarmakar52282, vkr2471}@gmail.com, [email protected]

Abstract—Low-cost air quality sensors are becoming ubiquitous in our daily lives as public awareness of air pollution continues to grow, and people take measures to monitor and improve the air they breathe indoors. Besides the standard operation of these sensors, fluctuations in environmental parameters can be leveraged to understand human behavior and activities in indoor spaces. Unlike traditional audio-visual, Radio Frequency, and inertial sensors, air quality sensors are easily scalable to a household, are privacy-preserving, and more economical. Such distributed sensor networks must jointly make decisions to monitor indoor occupants for downstream smart home and healthcare applications. However, due to low processing power, memory, and energy, they often struggle to maintain distributed data consensus and identify activity-affected sensor groups for accurate on-device inference. In this paper, we propose PoHAR framework that implements: (i) a conflict-free replicated data primitive for data sharing, (ii) a hierarchical clustering for ESP32 to detect activity-affected sensor groups with a self-supervised distance metric, and (iii) a leader-based group inference with off-the-shelf ML classifiers, enabling the sensor network to collaboratively detect hyperlocal indoor activities. Our extensive experiments demonstrated ondevice activity detection, achieving 97.41% accuracy for indoor activity and 99.68% for cooking activity, using off-the-shelf ML models with latency below 34 microseconds. Index Terms—Sensor Networks, Embedded ML, HAR.

I. INTRODUCTION In modern indoor environments, accurately detecting and responding to multiple concurrent human activities requires sensor networks that are both intelligent and resource-efficient [1]– [3]. While traditional sensing modalities such as cameras [4], [5], microphones [6], [7], Radio Frequency (RF) signals [8], [9], and inertial wearables [10], [11] have been widely explored for Human Activity Recognition (HAR), their deployment in real-world homes remains limited. Cameras and microphones raise significant privacy concerns, RF requires specialized hardware, and wearables are not scalable. In contrast, lowcost air quality sensors [12], [13] have become ubiquitous in households, driven by growing public awareness of indoor pollution and the increasing availability of consumer-grade air monitors. Beyond their primary purpose of pollutant tracking, fluctuations in air pollutants encode rich environmental signatures that reflect human activities [14], [15] such as cooking, cleaning, ventilation, or movement across rooms, making them a promising, privacy-preserving modality for indoor HAR. However, harnessing the full potential of such a system introduces unique challenges. Indoor pollutants often form localized hotspots and propagate non-uniformly depending on

room structure, ventilation, and the nature of activities [13], [14]. Consequently, sensors placed in different rooms or even different corners of the same room may capture distinct activity-induced signals, making spatial diversity an inherent characteristic of air-quality-based sensing. Prior work on multisensor and multi-view HAR has explored combining data from heterogeneous sensors to achieve richer scene understanding [10], [16], [17]. However, such aggregation is largely centralized, leading to high communication overhead when sharing high-fidelity data such as video, audio, or RF channel measurements with a central processor. Moreover, these methods do not account for localized environmental impact, where only a subset of sensors is affected by a given activity, a common characteristic in indoor air quality sensing. This leads to reduced accuracy and an inability to detect multiple simultaneous activities across different indoor zones. To address spatially varying influence across sensor nodes, researchers in wireless sensor networks have studied distributed clustering and group formation techniques such as HEED [18], DWEHC [19], and DEEC [20]. While these energy-aware and similarity-driven approaches demonstrate the feasibility of decentralized grouping, they struggle to scale in dense deployments because communication overhead scales quadratically with the number of nodes [21]. Moreover, these methods rely on static or hand-crafted distance metrics that fail to capture the latent structure of pollution-induced variations. Recent works [14], [15] show that pollutant spread creates dynamic, activity-dependent sensor clusters, but existing systems lack mechanisms to detect these affected sensor groups at runtime. In this paper, we proposed PoHAR that implements a selfsupervised (SSL) [22] similarity-aware partition-based hierarchical clustering algorithm [23] for ESP32 microcontrollers to detect affected sensor groups using a distributed conflictfree replicated set-data (set-CvRDT) primitive and RAFTbased [24] leader election and group inference for hyperlocal indoor activity recognition. PoHAR is tested with an in-house air quality sensor network deployed in household settings. These sensor devices capture pollutants such as CO2 , VOCs, and particulate matter, along with humidity and temperature. We conducted comprehensive experiments to evaluate PoHAR framework. We investigate SSL embedding quality using tSNE. Further, we evaluated Set-CvRDT through multiple scenarios, including concurrent operations and frequent updates, achieving consistent state convergence with minimal conver-

gence latency (i.e., 90 µs for 5-node consensus), evaluated the on-device clustering (i.e, 13 iterations to reduce 50 nodes to 3 clusters), and leader election with various node configurations and failure scenarios, demonstrating reliability. Finally, we assessed machine learning (ML) models on the leader ESP32, achieving over 97.41% accuracy for household and 97.28% for cooking activity recognition. Our key contributions are: 1) Conflict-free Replicated Set: We designed a distributed set data structure (Set-CvRDT) for consistent data sharing in a distributed sensor network under node failure. 2) Pollution-aware Clustering of Sensor Nodes: We implemented a hierarchical clustering algorithm [23] for ESP32 microcontrollers that groups sensors based on SSL embeddings [22] extracted from pollution measurements. This helps improve activity prediction by removing bias from unrelated sensors. 3) On-device Hyperlocal Prediction: We deploy machine learning models directly on ESP32 leader devices for hyperlocal activity recognition. We achieve over 97.41% accuracy for indoor activity and 97.28% for cooking activity recognition, with a latency of 34 µs. II. R ELATED W ORK We review the related work in three areas: (i) data modalities, (ii) approaches in human activity recognition (HAR) in contrast to air quality data, and (iii) decision making in distributed sensor networks. Details are as follows. A. Data Modalities in Human Activity Recognition In the HAR literature, researchers primarily used privacyinvasive visual data [1], [4], [5] and acoustic signals [2], [6] to identify human activities. Such systems have limited application indoors, such as in smart homes and healthcare settings, due to privacy concerns. Recently, we have seen the use of non-invasive modalities, such as Radio Frequency (RF) [8], [9], [25]–[27], that demonstrate HAR capabilities using point-cloud data. Moreover, inertial wearables [10], [11] are also explored in this context. However, the primary drawback of such approaches is that they lack understanding of human behavior beyond movements unless paired with privacy-invasive visual data. For instance, without the visual feed, such systems can only detect a person stir-frying a food item, disregarding the exact item being cooked. Therefore, we propose using air pollution data [3], [13], [14] to capture environmental changes induced by indoor activities. B. Approaches in Human Activity Recognition With video and audio data, researchers focus on learningbased feature extractors [2], [7], [28], [29] to design classifiers. Recent advancements have shifted towards spatio-temporal approaches [30], [31], along with PointNet [32], [33] and Point-GNN [34], to efficiently process RF-based non-invasive channel state, Doppler, and point-cloud data. Such approaches are very power hungry and require labeled data. We propose using a low-power self-supervised learning (SSL) [22], [35]

framework that leverages readily available air quality data with sparse HAR labels [13]. Moreover, with competitive modalities, limitations persist regarding the RF’s angular resolution, the camera’s FoV, and the microphone’s distance. Towards this, researchers have aggregated viewpoints [10], [16], [17] from multiple sensors to enrich the model’s input data. However, such aggregation leads to significant network load when sharing high-fps audio-visual or RF data with a central server. We propose sharing intermediate low-dimensional SSL-based embeddings from each air quality sensor to a leader node to improve network efficiency. C. Distributed Decision Making Recent studies [14], [15] have shown that air pollution forms local hotspots indoors, which are ventilated, trapped, or spread to adjacent rooms depending on human activities like turning on a fan, opening doors or windows, etc. As a result, proximate air quality sensors [13] are affected by the pollution-generating activity. With the spread of such pollutants, a dynamic sensor cluster can pick up environmental changes. This becomes more evident in the case of simultaneous indoor activities. However, identifying such dynamic clusters at run-time is taxing. HEED [18] introduced a hybrid approach that combines a node’s residual energy with similarity metrics. DWEHC [19] extended this with multilevel clusters and weighted metrics, reducing intra-cluster energy by 40% compared to HEED. DEEC [20] addressed heterogeneous networks using adaptive probabilities based on residual-to-average energy ratios. However, these methods face scalability issues [1], as communication overhead grows quadratically with the number of nodes. In contrast, hierarchical methods [23], [36] scale efficiently. We implemented a partition-based distributed agglomerative hierarchical clustering algorithm [23] for ESP32 microcontrollers on air quality sensors, performing distance-based partitioning and distance-aware merging of intermediate SSL embeddings to dynamically identify the affected sensor cluster. Lastly, each sensor cluster performs a leader-based [24] HAR inference with off-the-shelf ML classifiers. III. METHODOLOGY We illustrate the overall workflow of the PoHAR framework in Fig. 1. The air quality sensors collaborate to identify activityaffected sensor groups and perform hyperlocal activity recognition. As shown in the diagram, a Pollution Sensor Network (PSN) with multiple nodes (C1–C6) selects the leader (i.e., C4) using the (i) RAFT election mechanism [24]. Next, each sensor extracts a (ii) self-supervised learning (SSL) embedding from its local air quality time-series data using a neural network pretrained with time–frequency consistency [22] on unlabelled air quality data. These low-dimensional embeddings are shared across the network using a (iii) Set-CvRDT primitive, ensuring consistent, conflict-free exchange without requiring centralized synchronization. The leader executes the (iv) hierarchical clustering algorithm, which iteratively partitions and merges the SSL embeddings based on distances to identify activityaffected sensor groups. Finally, the identified sensor groups,

Set-CvRDT {S1}

PSN

C1

C5 C1

C4

C1

C3 C1

C2 C2

{SG1}

C5

C5

C3 Sensor data

C4 C3 C6

Leader Election

SSL SSL Netwotk Embedding

Time-Frequency Consistency Network

C3

Itr: 1

C4 C2

C2

C6

Itr: 2

Sharing SSL Embeddings

{SGN}

C5

C4

C6

Clustering

C6

Hyperlocal

Fig. 1: System overview of PoHAR.

each corresponding to a different indoor activity, perform their own leader-based (v) hyperlocal inference using off-the-shelf ML models. For example, one sensor group captures cooking, while another captures a cleaning activity. Thus, PoHAR ensures fine-grained HAR by allowing only the affected sensors to participate in activity inference. Details are as follows.

A={ E1 HEX_ },R={ } HEX_ A={ E1 },R={ } E2

C1

HEX_ A={ E1 },R={ E1 HEX_} E2

HEX_

HEX_

A={ },R={ } E1 UDP

C2

HEX_ A={ E1 },R={ } E2 HEX_

HEX_ A={ E1 },R={ E1 HEX_} E2 HEX_

A={ E1 HEX_ },R={ } UDP

A={ },R={ } CN

M={ E2 HEX_ } E1

E2 HEX_ A={ E1 },R={ } E2

A={ E1 HEX_ },R={ }

HEX_

A={ },R={ }

A. Distributed Leader Election The PoHAR framework relies on a lightweight, robust leader election mechanism to coordinate leader-based clustering and inference across resource-constrained sensor nodes. We implemented the RAFT consensus algorithm [24], which provides a fault-tolerant method for electing a stable leader in a distributed setting. RAFT is well-suited for ESP32 microcontrollers due to its simplicity, deterministic behavior, and low communication overhead. Each pollution sensor maintains a local state as follower, candidate, or leader and participates in periodic heartbeat and timeout-driven elections. A sensor transitions to the candidate state when it does not receive a heartbeat from an existing leader within a randomized timeout window. Candidates then request votes from neighboring sensors, and the sensor that obtains a majority becomes the next leader. RAFT ensures that only a single leader is active at any given time, even in the presence of message loss or intermittent connectivity. The elected leader updates the Set-CvRDT and broadcasts it across the sensor network to share local SSL embeddings for pollution-aware clustering. B. SSL-based Embeddings We implemented a time–frequency consistency (TF-C) based SSL model [22] to generate intermediate embeddings from raw air pollution time-series data at each sensor. Indoor activities such as cooking and sleeping often lead to rapid spikes, gradual rises, or periodic fluctuations. To capture these complementary cues without requiring labeled activity data, each sensor processes its local multivariate air quality stream using a lightweight neural encoder (see Fig. 1) trained with an SSL objective [22]: augmentations of the same signal segment (e.g., jittering, masking, or frequency perturbations) must produce embeddings that remain close in the latent space, while embeddings from unrelated segments should diverge. By enforcing agreement between the time-domain and frequencydomain views of the same pollution window, the model learns

Add

Update

Add

Update

Rem

HEX_ A={ E1 },R={ E1 HEX_} E2 HEX_

Update

Fig. 2: Conflict-free Replicated Set Data Operations.

robust low-dimensional representations that encode semantically meaningful information, suitable for downstream clustering and activity inference under noisy sensor readings and diverse indoor conditions. Algorithm 1 Update Set-CvRDT Protocol for Replicas. Require: Received add-set(Arcv ), rem-set(Rrcv ); Local add-set(A), rem-set(R), main-set(M ) 1: for all item ∈ Arecv do 2: if item ∈ / A & item ∈ / Rrecv then 3: Insert item into M ▷ Update main-set 4: Insert item into A ▷ Track item addition 5: end if 6: end for 7: for all item ∈ Rrecv do 8: if item ∈ M then 9: Remove item from M ▷ Update main-set 10: Insert item into R ▷ Track item removal 11: end if 12: end for

C. Distributed Conflict-free Replicated Set We implemented a conflict-free replicated data type (SetCvRDT) for easy and reliable data sharing among distributed sensors. The set-CvRDT is essential in key areas of the PoHAR framework. The data type is used to share SSL embeddings during clustering. Moreover, we store activity-affected sensorcluster information to perform hyperlocal inference. The setCvRDT has three grow-only sets: add-set (A), rem-set (R), and main-set (M) as shown in the Fig. 2. The figure illustrates consistent propagation of Set-CvRDT elements across multiple sensors (C1-CN) over an unreliable UDP communication

channel. Each sensor begins with empty sets. When a sensor generates a set addition event (e.g., E1 at C1 or E2 at C2), it appends the (element E, unique HEX code) pair to its local add-set and broadcasts the update to others. The HEX code uniquely identifies each element in the replica sets. Due to UDP’s lossy and asynchronous nature, other sensors may receive the updates at different times. Each replica monotonically incorporates updates using the state update protocol in Algorithm 1, ensuring that all add operations eventually converge. When a removal occurs (e.g., removing E1), the originating sensor inserts the event into its rem-set (R) and disseminates the update. Finally, the main-set (M) at each sensor is updated from local and received copies of the addset and rem-set, ensuring all replicas converge to the same set despite message delays, reordering, or loss. The set-CvRDT implementation ensures crucial properties for distributed storage. We ensure all replicas will eventually have a consistent state (Convergence). We ensure that failed or newly added nodes will ultimately reach the same final state as the other replicas (Fault Tolerance). We ensure that the order or grouping of set operations (i.e., addition or removal of items) across replicas does not affect the final state (Commutativity and Associativity). Finally, we assign a unique HEX code to each item in the set to resolve conflicts, ensuring that each item can be reinserted (Reinsertions) whenever required. D. Pollution-aware Clustering We implement pollution-aware sensor clustering using similarity-based partitioning and distance-aware merging within each partition. Each sensor computes its local SSL embeddings to represent its pollution context. When partitioning the network, we group clusters with their top nearest neighbors together. We compute the Euclidean distance of the SSL embeddings as a similarity measure between any two sensors. This forms the edges of the initial graph representation of the sensor network. We only include a list of edges with the shorter distance to threshold θ, and represent all ignored edges as a lower bound bL indicating that their distances are greater than bL . The distance-aware merging algorithm works on each partition and performs merges locally. Whenever it merges a cluster pair, it ensures that the two clusters are mutual nearest neighbors by checking distance bounds (i,e, bL and bU ), thereby guaranteeing the correctness of the clusters. Given an undirected weighted graph G = (C, W ), where C is a set of nodes and W is a set of weighted edges indicating the distances between pairs of items in C, Algorithm 2 initializes each node into its own cluster. We assume each initial node c ∈ C has an associated label denoted by label(c). Here, it’s the sensor IDs. We further assume that the cluster label is the maximum label of the cluster’s node. Let Dist(Ci , Cj ) be the euclidean distance between clusters Ci and Cj , so that we can define weight w(Ci , Cj ). We define L(Ci ) as the list of nearest neighbors of Ci , whose size limit is a configurable parameter. For each Cj ∈ L(Ci ), we define bL (Ci , Cj ) and bU (Ci , Cj ) as the lower and upper bounds of dist(Ci , Cj ) respectively,

C1

C2

C12

C3

C4

C5

C6

C12

C45

C123

C3

C456

C6

C45

C123

C456

C123

C456

Fig. 3: Pollution-aware merging of nodes in the clustering. both initialized to dist(Ci , Cj ). We update the distance between a merged cluster Cij and any other cluster or node Cx as equation 1. Similarly, distance bounds are also updated as per equation 2 and equation 3. Dist(Ci , Cx ) · |Ci | + Dist(Cj , Cx ) · |Cj | |Ci | + |Cj | (1) bL (Ci , Cx ) · |Ci | + bL (Cj , Cx ) · |Cj | (2) bL (Cij , Cx ) = |Ci | + |Cj | bU (Ci , Cx ) · |Ci | + bU (Cj , Cx ) · |Cj | bU (Cij , Cx ) = (3) |Ci | + |Cj |

Dist(Cij , Cx ) =

Algorithm 2 Pollution-aware Clustering of Sensor Nodes Input: Cluster graph G = (C, W ); Threshold θ Output: Final clusters C ∗ 1: while there exists Dist(Ci , Cj ) ≤ θ do 2: P ← P artition(G) 3: C′ ← ∅ 4: for all partitions Ph = (Ch , {L(Ci ) | Ci ∈ Ch }) ∈ P do 5: /∗ Polluiton-aware local merging for Ph ∗/ 6: Gh ⇐ Build a graph from Ph 7: for all Ci ∈ Ch do 8: N N (Ci ) ⇐ nearest neighbor of Ci whose upper bound (bU ) is smaller than the lower bounds (bL ) of other neighbors; null if non-existent 9: end for 10: while true do 11: if ∃(Ci , Cj ) : (N N (Ci ) = Cj ) ∧ (N N (Cj ) = Ci ) ∧ (bU (Ci , Cj ) ≤ θ) then 12: Merge Ci with Cj in Gh , 13: Update Gh , and N N (·) 14: else 15: break 16: end if 17: end while 18: C ′ ← C ′ ∪ Merged clusters in Gh 19: end for 20: C ← Integrate C ′ ▷ Merge overlapping clusters 21: W ← Merge weights based on C 22: end while 23: C ∗ ← C We merge clusters within a partition as long as there exists a pair of mutual nearest neighbors (Ci , Cj ), and bU (Ci , Cj ) ≤ θ

as per lines 11 to 13 in the Algorithm 2 (See Fig. 3). It outputs a set of clusters C ′ . Lastly, distance bounds are used to merge overlapping clusters before the final clusters C ∗ output. E. Hyperlocal Inference Each activity-affected sensor cluster independently performs hyperlocal inference to determine the specific indoor activity occurring within its localized region. To coordinate this process without relying on a central controller, each cluster initiates a second RAFT-based leader-election round [24], ensuring that a single inference leader is selected among the affected sensors. Once elected, the inference leader collects SSL embeddings from all cluster members that already encode the time–frequency characteristics of local air-quality fluctuations, and aggregates them for classification. The leader then executes inference using off-the-shelf ML models (i.e., Decision Trees, Random Forests, Extra Trees, Gaussian Naive Bayes, and lightweight Neural Networks). By assigning inference responsibilities only to the sensors affected by the localized activity, the system achieves fine-grained, energy-efficient HAR while preserving the spatial relevance of the indoor pollutants. IV. E XPERIMENTAL RESULTS Here, we present experimental results with PoHAR framework over our in-house pollution sensor network deployment. A. Implementation Details Sensor Network Deployment: Our study deployed a dense network of low-cost air quality sensors across 30 diverse indoor sites over six-months, covering both summer and winter seasons. These sites include studio apartments, shared classrooms, research laboratories, residential households, and food canteens, providing a representative ecosystem of realworld human activity and pollution dynamics. Across the deployment, the sensing infrastructure collected 89.1 million samples, amounting to 13646 hours of continuous indoor air quality measurements. To contextualize the sensor data with human behavior, we incorporated 3957 manually annotated activity events from 24 active participants (among 46 total occupants). These annotations capture a wide spectrum of indoor activities, including engagement and occupancy (i.e., enter, exit), occupant behavior (i.e., fan on/off, AC on/off), prohibited or impactful practices (i.e., gathering, eating), and cookingrelated activities: annotated across five cooking categories (i.e., boiling, deep-frying, shallow-frying, steaming, and shallowfrying with boiling) and eleven food items (e.g., tea, rice, fish, egg, lentils, leafy vegetables). This large-scale, heterogeneous deployment enables a fine-grained analysis of spatiotemporal indoor pollution patterns shaped by everyday human activities and diverse cooking practices. SSL Model and Off-the-shelf ML Models: To implement the Self-supervised embedding network, we adapted the neural network architecture from the time-frequency consistency model [22]. Further, we have implemented Decision Tree (DT), Random Forest (RF), Extra Trees (ET), Gaussian Naive Bayes (NB), and Multilayer Perceptron (MLP) classifiers to detect

Fig. 4: t-SNE visualization of cooked food items.

hyperlocal activities at the leader nodes with the emlearn [37] Python library. Evaluation Metric: We evaluate the platform using four key metrics. Latency is measured in microseconds, milliseconds, or seconds to capture responsiveness across system operations. Memory usage is tracked in bytes to assess the platform’s footprint on resource-constrained ESP32 devices. Power consumption is measured in milliwatts to quantify energy efficiency during different stages of PoHAR. Finally, model performance is evaluated based on accuracy in predicting indoor activities. These metrics collectively provide a clear view of system efficiency and predictive reliability. B. SSL Embeddings Quality Analysis Figure 4 illustrates the SSL-based latent embeddings of the cooking activities projected onto a two-dimensional space by applying t-SNE to the embeddings extracted from the trained time-frequency consistency (TF-C) model. The plot shows that the TF-C model successfully captured meaningful semantic relationships between different food items, resulting in 24 distinct clusters. Each food item category, including Paratha, Potol, Pakoda, Gourd, Curry, Posto, Chickpeas, Mixed spices, Egg, Papad, Cauliflower, Fish, Dal, Brinjal, Chicken, etc., forms a relatively cohesive cluster in the embedding space. C. Pollution-aware Clustering Effeciency Memory Usage: Memory consumption increases nearly linearly with the number of nodes, rising from approximately 63 KB at 10 nodes to about 228 KB at 50 nodes, as shown in Fig. 5a. This reflects the overhead of maintaining additional cluster structures, distance bounds, and nearest-neighbor mappings as the system scales. The steady growth demonstrates a lightweight memory footprint suitable for PoHAR. Power Usage with Number of Nodes: Average power usage shows a gradual upward trend with increasing nodes, with median USB power consumption ranging from roughly 250 mW (10 nodes) to 360 mW (50 nodes) as per Fig. 5b. Despite the increase, the boxplots indicate stable, bounded power

158KB

150 100 50 0

124KB 63KB

10

20 30 40 Number of Nodes

50

380 360 340 320 300 280 260 240

(a) Memory with nodes

20 Nodes

350

10

20 30 40 Number of Nodes

50

(b) Power usage with nodes

Total Execution Time (ms)

193KB

200

USB Avg Power (mW)

228KB

USB Average Power (mW)

Memory Used (KB)

250

300 250 50 Nodes

350 300 250

0.0

0.1

0.2

0.3 0.4 Time (s)

0.5

0.6

(c) Power draw

350 300 250 200 150 100 50 0

320ms

204ms

82ms

111ms

31ms

10

20 30 40 Number of Nodes

50

(d) Clustering time with nodes

Fig. 5: Pollution-aware clustering performance: (a) memory usage with number of nodes, (b) power usage with number of nodes, (c) power draw while execution, (d) total clustering time with number of nodes.

USB Average Power (mW)

1200

Leader Follower

1000 800 600 400 2

3 4 Number of Nodes

(a) RAFT power usage

5

0.0

2.5

5.0 7.5 10.0 12.5 Time difference (seconds)

15.0

(b) Recovery time on leader failure

Fig. 6: RAFT-based leader election performance: (a) power usage with number of cluster nodes, (b) Recovery time to find a new cluster leader on current leader failure.

behavior, suggesting that clustering computations introduce minimal energy variability in low-power deployments. Real-Time Power Draw During Execution: Power traces in Fig. 5c exhibit periodic bursts correlating with cluster-merge and graph-update operations. For both cases, power fluctuates between 260–360 mW, while for 50 nodes, peaks become slightly less frequent due to increased clustering time at the leader. These patterns confirm that the algorithm performs short, compute-intensive bursts rather than sustained heavy loads, making it compatible with real-time execution. Clustering Time with Increasing Nodes. As observed in the power analysis, execution time scales proportionally with the number of nodes, increasing from 31 ms at 10 nodes to 320 ms at 50 nodes as shown in Fig. 5d. Even at the largest configuration, the total clustering time remains well below one second, highlighting the efficiency of partition-based processing and incremental cluster-graph updates. This characteristic is crucial for pollution-aware systems using multiple distributed sensors. D. Distributed Leader Election Analysis Leader and Follower Power Consumption: Fig. 6a compares the power usage of leader and follower nodes across cluster sizes ranging from two to five devices. The leader consistently consumes more power (approximately 600–650 mW) due to its coordination responsibilities, including heartbeat generation and log management. In contrast, follower nodes operate at a

lower and stable power range (270–300 mW), with negligible variation as the number of nodes increases. This indicates that RAFT imposes minimal incremental overhead as clusters grow, making it efficient for resource-constrained settings. Leader Failure Recovery Time: Fig. 6b presents the temporal distribution of leader recovery times following induced leader failures. The RAFT protocol reliably re-establishes the leader within approximately 3–6 seconds. The narrow spread of the violin plot reflects consistent timeout-triggered re-election across repeated trials, demonstrating RAFT’s robustness and predictable fault tolerance in the sensor network. E. Activity Classification Performance Detection of Indoor Activities Tree-based models exhibited the strongest performance across all metrics, with Random Forest emerging as the best overall model, achieving an accuracy of 97.41% with an average inference time of only 34 µs as shown in TABLE I. The confusion matrix is shown in Fig. 7a. Extra Trees offered comparable accuracy but incurred higher latency, while the Decision Tree remained the fastest (9 µs) with competitive accuracy. In contrast, models relying on floatingpoint operations, such as Gaussian Naive Bayes and the small Neural Network, performed poorly on the ESP32, yielding less than 50% accuracy and significantly longer prediction times, due to limited floating-point compute on ESP32. TABLE I: Model Performance for Indoor Activities. Model Decision Tree Random Forest Extra Trees Gaussian NB Neural Network

Accuracy 93.56% 97.41% 97.36% 48.28% 49.45%

Avg. Time 9 µs 34 µs 62 µs 356 µs 89 µs

Max Time 42 µs 330 µs 463 µs 382 µs 438 µs

Detection of Cooking Activities Across both cooking type and food identification tasks, tree-based models consistently delivered the strongest results. Random Forest emerged as the most effective model overall, achieving accuracy above 99.4% with low inference times of 16–23 µs, offering the best balance between accuracy and computational cost as shown in Fig. 7, TABLE II, and TABLE III. Extra Trees achieved slightly higher accuracy (up to 99.81%) but required considerably more time (43–64 µs), while Decision Tree provided the fastest predictions

0

0

600

0

0

0

762

14

0

0

0

500

0

0

0

11

884

0

20

0

400

0

0

0

0

0

204

0

0

300

0

0

0

0

19

0

165

0

200

0

0

0

0

0

0

0

29

100

1

147

0

0

0

0

0

347

0

0

0

0

0

99

0

0

0

0

0

7

600 400 200 0

ow

0

800

all Sh

(a) Indoor activity

106 0 0 0 0 0 0 0 0 0 0 0 143 0 0 0 0 0 0 0 0 0 0 0 62 0 0 0 0 0 0 2 0 0 0 0 35 0 0 0 0 0 0 0 0 0 0 0 220 0 0 0 1 0 0 1 0 0 0 0 25 0 0 0 0 0 0 0 0 0 0 0 53 0 0 0 0 0 0 0 0 0 0 0 43 0 0 0 0 1 0 0 1 0 0 0 531 0 0 0 0 0 0 0 0 0 0 0 102 0 0 0 0 0 0 0 0 0 0 0 252

500

Le Po Poin afy v Lad pp te e ie P y s d g L get s f Tea Rot Riceotatoeeds ourdentilsable inge Fish Egg s r i Eg g Le Ladie Fi afy s sh ve fing ge er ta Po L bles int en ed tils Po go pp ur ys d ee Po ds tat o Ric e Ro ti Tea

0

0

Fry ing ng ,B oil ing Ste am ing

0

0

Fry i

0

0

ng

31

1

Fry i

0

976

ow

700

0

800

ep

0

all

0

0

De

0

0

Sh

0

0

ow Fry Sha D Ste ing, llow eep am Boi Fry Fry Boi ing ling ing ing ling Bo ilin g

0

0

all

0

0

Sh

0

34

ff Ac on Ea tin g En ter s Ex its Fan off Fan on Ga the r

F F E the an o an o Exit nter Eatin Ac o Ac o r n ff s s g n ff

0

0

Ac o

Ga

64

(b) Cooking type

400 300 200 100 0

(c) Food Item

TABLE II: Model Performance for Cooking Classification. Model Decision Tree Random Forest Extra Trees Gaussian NB Neural Network

Accuracy 97.28% 99.68% 99.81% 47.72% 44.36%

Avg. Time 6 µs 16 µs 43 µs 157 µs 61 µs

Max Time 39 µs 271 µs 339 µs 181 µs 352 µs

TABLE III: Model Performance for Food Item Classification. Model Decision Tree Random Forest Extra Trees Gaussian NB Neural Network

Accuracy 94.17% 99.49% 99.68% 46.52% 45.71%

Avg. Time 7 µs 23 µs 64 µs 178 µs 83 µs

Max Time 38 µs 289 µs 311 µs 209 µs 391 µs

F. On-device Inference Power Analysis Fig. 8a compares the average power consumption of different ML models during on-device inference. Tree-based models (i.e., Decision Tree, Extra Trees, and Random Forest) show nearly identical power usage of 272–274 mW, indicating that their lightweight integer-based computations impose minimal additional load on the ESP32. Gaussian Naive Bayes incurs slightly higher power consumption (283 mW), likely due to repeated floating-point operations. In contrast, the MLP model consumes the highest power at 308 mW, reflecting the computational overhead of neural network inference on constrained hardware.

300 250 200 150 100 50 0

272.05

273.98

272.84

283.06

308.86

350

Random Forest

300 250

Power (mW)

(6–7 µs) with strong accuracy for both tasks. All tree-based models maintained real-time inference capability on ESP32. In contrast, models that relied heavily on floating-point operations, Neural Networks, and Gaussian Naive Bayes performed poorly on the ESP32 platform. Even with a reduced neural network architecture (two 10-neuron hidden layers), accuracy dropped below 50% and inference times increased substantially, largely due to ESP32’s limited floating-point support and the required integer conversion of sensor data. Overall, Random Forest proved to be the optimal choice, delivering high accuracy with efficient computation, enabling real-time cooking activity detection.

Average Power Consumption (mW)

Fig. 7: Classification confusion matrix of Random Forest: (a) indoor activity dataset, (b) cooking type, and (c) food item in the pollution sensor network deployment.

350

Sklearn MLP

300

DT

ET RF NB MLP Machine Learning Model

(a) ML model average power usage

250

0.00

0.02

0.04 0.06 Time (s)

0.08

0.10

(b) ML model Power draw

Fig. 8: On-device power usage for hyperlocal inference. Fig. 8b presents fine-grained power draw traces for Random Forest and the MLP. Both models exhibit periodic spikes corresponding to inference cycles, but their magnitudes and durations differ significantly. Random Forest exhibits short, sharp spikes with a quick return to idle, illustrating its fast execution and low computational burden. The MLP, however, shows longer, higher-power spikes, indicating prolonged processing times and heavier compute demands. These patterns reinforce that tree-based models not only deliver better accuracy but also maintain more energy-efficient real-time performance for hyperlocal inference on ESP32. V. C ONCLUSION We presented PoHAR, a lightweight and scalable framework that enables distributed air-quality sensor networks to collaboratively detect hyperlocal indoor activities under strict hardware constraints. By integrating CRDT-based data sharing, hierarchical clustering with a self-supervised distance metric, and RAFT-style leader-based group inference, our approach effectively addresses the core challenges of on-device processing, consensus maintenance, and activity-aware sensor grouping in low-power deployments. Extensive evaluation on real-world datasets demonstrates that P O HAR achieves high inference accuracy of 97.41% for indoor activities and 99.68% for cooking activities, while maintaining below 34 µs latency and milliwatt-level power usage. The results highlight the

viability of repurposing commodity air-quality sensors for privacy-preserving human activity recognition. Open-sourced at https://github.com/prasenjit52282/PoHAR. ACKNOWLEDGEMENT

The work is supported by the Prime Minister Research Fellowship (IIT/Acad/PMRF/SPRING/2022-23, dated 24 March 2023) and Google Award on Society-Centered AI 2025. R EFERENCES [1] N. Robertson and I. Reid, “A general method for human activity recognition in video,” Computer Vision and Image Understanding, vol. 104, no. 2-3, pp. 232–248, 2006. [2] J. A. Stork, L. Spinello, J. Silva, and K. O. Arras, “Audio-based human activity recognition using non-markovian ensemble voting,” in 2012 IEEE RO-MAN: The 21st IEEE International Symposium on Robot and Human Interactive Communication, pp. 509–514, IEEE, 2012. [3] B. Fang, Q. Xu, T. Park, and M. Zhang, “Airsense: an intelligent homebased sensing system for indoor air quality analytics,” in Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing, pp. 109–119, ACM, 2016. [4] P. C. Ribeiro, J. Santos-Victor, and P. Lisboa, “Human activity recognition from video: modeling, feature selection and classification architecture,” in Proceedings of International Workshop on Human Activity Recognition and Modelling, vol. 61, p. 78, 2005. [5] W. Lin, M.-T. Sun, R. Poovandran, and Z. Zhang, “Human activity recognition for video surveillance,” in 2008 IEEE international symposium on circuits and systems (ISCAS), pp. 2737–2740, IEEE, 2008. [6] S. Ntalampiras and I. Potamitis, “Transfer learning for improved audiobased human activity recognition,” Biosensors, vol. 8, no. 3, p. 60, 2018. [7] S. Cristina, V. Despotovic, R. Pérez-Rodrı́guez, and S. Aleksic, “Audioand video-based human activity recognition systems in healthcare,” IEEE Access, vol. 12, pp. 8230–8245, 2024. [8] A. Sen, A. Das, S. Pradhan, and S. Chakraborty, “Continuous multiuser activity tracking via room-scale mmwave sensing,” in 2024 23rd ACM/IEEE International Conference on Information Processing in Sensor Networks, pp. 163–175, IEEE, 2024. [9] K. Xu, J. Wang, H. Zhu, and D. Zheng, “Evaluating self-supervised learning for wifi csi-based human activity recognition,” ACM Transactions on Sensor Networks, vol. 21, no. 2, pp. 1–38, 2025. [10] Y. Jain, C. I. Tang, C. Min, F. Kawsar, and A. Mathur, “Collossl: Collaborative self-supervised learning for human activity recognition,” Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 6, no. 1, pp. 1–28, 2022. [11] A. Abedin, M. Ehsanpour, Q. Shi, H. Rezatofighi, and D. C. Ranasinghe, “Attend and discriminate: Beyond the state-of-the-art for human activity recognition using wearable sensors,” Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 5, no. 1, pp. 1–22, 2021. [12] Future Market Insights, “Indoor air quality monitor market outlook,” tech. rep., Future Market Insights, 2023. [13] P. Karmakar, S. Pradhan, and S. Chakraborty, “Indoor air quality dataset with activities of daily living in low to middle-income communities,” Advances in Neural Information Processing Systems, vol. 37, pp. 70076– 70100, 2024. [14] P. Karmakar, S. Pradhan, and S. Chakraborty, “Exploring indoor air quality dynamics in developing nations: A perspective from india,” ACM Journal on Computing and Sustainable Societies, vol. 2, no. 3, pp. 1–40, 2024. [15] P. Karmakar, S. Pradhan, and S. Chakraborty, “Exploiting air quality monitors to perform indoor surveillance: Academic setting,” in Adjunct Proceedings of the 26th International Conference on Mobile HumanComputer Interaction, pp. 1–6, 2024. [16] I. Gokarn, Y. Hu, T. Abdelzaher, and A. Misra, “Ra-mosaic: Resource adaptive edge ai optimization over spatially multiplexed video streams,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 21, no. 9, pp. 1–25, 2025. [17] H. Xie, H. Yao, X. Sun, S. Zhou, and S. Zhang, “Pix2vox: Context-aware 3d reconstruction from single and multi-view images,” in Proceedings of the IEEE/CVF international conference on computer vision, pp. 2690– 2698, 2019.

[18] O. Younis and S. Fahmy, “Heed: a hybrid, energy-efficient, distributed clustering approach for ad hoc sensor networks,” IEEE Transactions on mobile computing, vol. 3, no. 4, pp. 366–379, 2004. [19] P. Ding, J. Holliday, and A. Celik, “Distributed energy-efficient hierarchical clustering for wireless sensor networks,” in International conference on distributed computing in sensor systems, pp. 322–339, Springer, 2005. [20] L. Qing, Q. Zhu, and M. Wang, “Design of a distributed energy-efficient clustering algorithm for heterogeneous wireless sensor networks,” Computer communications, vol. 29, no. 12, pp. 2230–2237, 2006. [21] A. Taherkordi, R. Mohammadi, and F. Eliassen, “A communicationefficient distributed clustering algorithm for sensor networks,” in 22nd International Conference on Advanced Information Networking and Applications-Workshops (aina workshops 2008), pp. 634–638, IEEE, 2008. [22] X. Zhang, Z. Zhao, T. Tsiligkaridis, and M. Zitnik, “Self-supervised contrastive pre-training for time series via time-frequency consistency,” Advances in neural information processing systems, vol. 35, pp. 3988– 4003, 2022. [23] Y. Wang, V. Narasayya, Y. He, and S. Chaudhuri, “Pack: An efficient partition-based distributed agglomerative hierarchical clustering algorithm for deduplication,” Proceedings of the VLDB Endowment, vol. 15, no. 6, pp. 1132–1145, 2022. [24] D. Ongaro and J. Ousterhout, “In search of an understandable consensus algorithm,” in 2014 USENIX Annual Technical Conference, pp. 305–319, USENIX Association, 2014. [25] A. D. Singh, S. S. Sandha, L. Garcia, and M. Srivastava, “Radhar: Human activity recognition from point clouds generated through a millimeterwave radar,” in Proceedings of the 3rd ACM Workshop on Millimeterwave Networks and Sensing Systems, pp. 51–56, 2019. [26] P. Zhao, C. X. Lu, B. Wang, N. Trigoni, and A. Markham, “Cubelearn: End-to-end learning for human motion recognition from raw mmwave radar signals,” IEEE Internet of Things Journal, vol. 10, no. 12, pp. 10236–10249, 2023. [27] X. Zeng, Y. Shi, and A. Zhou, “Multi-har: Human activity recognition in multi-person scenes based on mmwave sensing,” in 2022 IEEE 8th International Conference on Computer and Communications (ICCC), pp. 1789–1793, IEEE, 2022. [28] J. Yan, X. Zhang, C. Tan, and D. Li, “Skelformer: An adaptive hierarchical transformer-based approach on skeleton graphs for human action recognition in video sequences,” PloS one, vol. 21, no. 1, pp. 1–20, 2026. [29] T. F. N. Bukht, H. Rahman, M. Shaheen, A. Algarni, N. A. Almujally, and A. Jalal, “A review of video-based human activity recognition: theory, methods and applications,” Multimedia Tools and Applications, vol. 84, no. 17, pp. 18499–18545, 2025. [30] R. B. Rusu and S. Cousins, “3d is here: Point cloud library (pcl),” in 2011 IEEE international conference on robotics and automation, pp. 1–4, IEEE, 2011. [31] C. Choy, J. Gwak, and S. Savarese, “4d spatio-temporal convnets: Minkowski convolutional neural networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 3075–3084, 2019. [32] C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 652– 660, 2017. [33] C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” Advances in neural information processing systems, vol. 30, 2017. [34] W. Shi and R. Rajkumar, “Point-gnn: Graph neural network for 3d object detection in a point cloud,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 1711–1719, 2020. [35] T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International conference on machine learning, pp. 1597–1607, PmLR, 2020. [36] S. Gilpin, B. Qian, and I. Davidson, “Efficient hierarchical clustering of large high dimensional datasets,” in Proceedings of the 22nd ACM international conference on Information & Knowledge Management, pp. 1371–1380, 2013. [37] J. Nordby, M. Cooke, and A. Horvath, “emlearn: Machine Learning inference engine for Microcontrollers and Embedded Devices,” Mar. 2019.

Record · ID 175214 · SHA-256 a693aaed581ca27f
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.