A BIOMETRIC SENSOR NETWORK TO ENABLE REAL-TIME MEASUREMENT OF INDIVIDUAL STUDENT ENGAGEMENT IN STEM LECTURE ENVIRONMENTS
By Ahmed Elsayed
arXiv:2607.28944v1 [cs.CR] 31 Jul 2026
B.S. in Electrical and Computer Engineering, 2016 A Dissertation Submitted to the Faculty of the J.B. Speed School of Engineering of the University of Louisville in Partial Fulfillment of the Requirements for the Degree of Master of Science in Electrical Engineering Department of Electrical and Computer Engineering University of Louisville Louisville, Kentucky December 2025
Copyright 2026 by Ahmed Elsayed
All rights reserved
A BIOMETRIC SENSOR NETWORK TO ENABLE REAL-TIME MEASUREMENT OF INDIVIDUAL STUDENT ENGAGEMENT IN STEM LECTURE ENVIRONMENTS
By
Ahmed Elsayed B.S. in Electrical and Computer Engineering, 2016 Dissertation approved on December 11, 2025 by the following dissertation Committee:
Michael McIntyre, Ph.D., Advisor
Aly Farag, Ph.D.
Thomas Tretter, Ph.D.
John Naber, Ph.D.
ii
ACKNOWLEDGMENTS
I would like to express my deepest gratitude to my mother for her unwavering love, sacrifices, and prayers—her support has been the foundation of everything I have accomplished. I am also profoundly grateful to my brother, sisters, cousins, nephews, and my entire family in Egypt, who have continually encouraged me and made countless sacrifices so I could pursue my academic goals. I extend my sincere appreciation to my friends, my colleagues at the CVIP Lab, and the Egyptian community in Louisville for their companionship, support, and encouragement throughout this journey. I am especially grateful to Dr. Aly Farag for allowing me to join his research group at the CVIP Lab, and to Ahmed Shalaby for his continuous support and for recommending me for an opportunity at Apple. I am also grateful to Ahmed Abdelkawy for the many fruitful discussions and collaborative efforts during our time working together in the CVIP Lab, and to Dr. Michel MacIntyre for his helpful support.
iii
ABSTRACT
A BIOMETRIC SENSOR NETWORK TO ENABLE REAL-TIME MEASUREMENT OF INDIVIDUAL STUDENT ENGAGEMENT IN STEM LECTURE ENVIRONMENTS Ahmed Elsayed December 11, 2025 Student engagement (SE) is a critical predictor of academic performance and retention in STEM education, yet existing measurement approaches are often intrusive, manually intensive, or unsuitable for real-time classroom use. This thesis proposes a novel Biometric Sensor Network (BSN) designed to enable real-time measurement and continuous tracking of individual student engagement in STEM classroom environments. The system enables capturing of behavioral, emotional, and cognitive indicators through camera-based sensing while preserving ethical and privacy constraints. To measure these indicators unobtrusively and ethically, we propose a BSN composed of Student Processing Units (SPUs) that function as distributed sensing nodes. The network is explicitly designed to satisfy five objectives: it must be non-intrusive, non-invasive, non-stigmatizing, real-time, and automatic, while ensuring rigorous protection of student data security and privacy. Each SPU supports two operational modes: (i) a dataset-collection mode, in which raw student video is temporarily recorded to construct a private SE dataset for
iv
model training and validation, and (ii) an analysis mode, in which the SPU performs real-time inference on 10-second video segments without storing or transmitting raw frames. In this analysis role, each SPU enables fully on-device processing—including face detection, gaze estimation, and affective analysis—ensuring that no identifiable video data leaves the device. A secure backend infrastructure manages device authentication, session orchestration, and encrypted data ingestion. The full system integrates hardware design, computer-vision pipelines, wireless networking, security protocols, and session-level data management. A series of evaluations demonstrates that the proposed SPU meets the required performance benchmarks, including sustaining multi-device data bandwidth, maintaining sub-second clock synchronization accuracy, executing moderate CV/ML inference workloads in real time, and achieving more than two hours of battery-powered operation. The system significantly improves upon a previous baseline developed in the CVIP Lab in terms of security, autonomy, compute capability, and data reliability. Overall, this work provides a complete, deployable BSN architecture that enables real-time measurement and tracking of individual student engagement during STEM classroom lectures while adhering to the strict ethical, privacy, and operational requirements of real-world educational environments.
v
TABLE OF CONTENTS
Acknowledgments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
iii
Abstract . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
iv
List of Tables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
viii
List of Figures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
x
Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
1
I.1
Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
1
Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
4
II.1 Student Engagement Definition . . . . . . . . . . . . . . . . . . . . .
4
II.2 Attention, Flow, and Engagement . . . . . . . . . . . . . . . . . . . .
6
II.3 Engagement Measures . . . . . . . . . . . . . . . . . . . . . . . . . .
7
II.4 Real-Time Engagement Measures . . . . . . . . . . . . . . . . . . . .
9
II.5 Literature Review (zReal-time Methods) . . . . . . . . . . . . . . . .
11
Analysis and Design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
13
III.1 Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
13
III.2 System Requirements . . . . . . . . . . . . . . . . . . . . . . . . . . .
24
III.3 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
32
BSN Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
35
vi
IV.1 SPU Hardware . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
36
IV.2 SPU Software . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
43
IV.3 Server Software . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
54
BSN Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
61
V.1 Baseline Comparison . . . . . . . . . . . . . . . . . . . . . . . . . . .
61
Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
72
REFERENCES . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
75
Curriculum Vitae . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
88
vii
LIST OF TABLES
1
SPU Hardware Requirements . . . . . . . . . . . . . . . . . . . . . . . .
25
2
CFU Hardware Requirements . . . . . . . . . . . . . . . . . . . . . . . .
27
3
SPU Software Requirements . . . . . . . . . . . . . . . . . . . . . . . . .
28
4
Security and Data Protection Requirements (Edge Devices: SPUs/CFUs)
29
5
Server Software Requirements . . . . . . . . . . . . . . . . . . . . . . . .
31
6
Security and Data Protection Requirements (Server) . . . . . . . . . . .
32
7
Comparison of candidate single-board computers (SBCs) evaluated for the SPU. Metrics include compute performance, physical dimensions, cost, power requirements, and required peripherals. . . . . . . . . . . . . . . .
8
37
Comparison of camera modules compatible with the Raspberry Pi 5, including interface type, stereo and RGB resolution with maximum frame rates, on-board AI acceleration, power consumption, and approximate cost. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
38
9
Bill of Materials (BOM) for the SPU . . . . . . . . . . . . . . . . . . . .
43
10
Comparison between the baseline [42] and the proposed BSN SPU hardware. The proposed system provides significant improvements in compute capability, wireless connectivity, camera quality, physical security, autonomy, and AI processing capacity with a moderate cost increase. . . . . .
viii
62
11
Comparison of software features between the baseline [42] and the proposed BSN SPU software. The proposed software stack provides major improvements in security and data privacy to support IRB-compliant classroom operation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
12
63
Summary of server load and network-throughput experiment for ten SPUs. All devices successfully transmitted all generated packets with zero loss. Uploads were not real-time but completed within approximately one hour after the session ended, confirming correct buffering, retry logic, and cleanup of local temporary storage. . . . . . . . . . . . . . . . . . . . . .
13
65
Clock synchronization offsets between each SPU and the BSN server measured over a ten-minute interval. The offsets show tight clustering around zero ms, confirming stable NTP-based synchronization across all devices.
68
14
Summary of battery-endurance trials under continuous SPU operation. .
68
15
Inference performance of selected DepthAI pipelines running on the OAKD Lite accelerator. Results indicate that the accelerator provides sufficient throughput for real-time classroom engagement sensing. . . . . . . . . . .
ix
71
LIST OF FIGURES
1
Venn diagram to represent engagement as a malleable construct in which each component is interdependent and cannot be completely disentangled from the other two. The diagram also shows that SE overlaps with related constructs such as attention and flow. . . . . . . . . . . . . . . . . . . . .
5
2
The six primary methods for measuring SE reported in the literature. . .
8
3
The five primary types of signals used as real-time indicators to assess SE as reported in the literature. . . . . . . . . . . . . . . . . . . . . . . . . .
4
9
Design-decision workflow summarizing the selection of SE measurement methods, real-time indicators, and sensing technology leading to the final SPU-based BSN architecture. . . . . . . . . . . . . . . . . . . . . . . . .
5
34
System architecture of the SPU hardware, illustrating the interaction between the SBC, camera module, touch display, wireless network interface, battery management subsystem, and local storage. . . . . . . . . . . . .
6
System-level schematic diagram of the SPU, showing the electronic components and their interconnections. . . . . . . . . . . . . . . . . . . . . .
7
36
40
Mechanical drawings of the SPU enclosure showing key dimensions, mounting features, and component placement. All dimensions are given in millimeters. Refer to Table 9 for the designated parts corresponding to the identifiers shown in the upper-right drawing. . . . . . . . . . . . . . . . .
x
42
8
Fabricated SPU unit showing the assembled enclosure, integrated touch display, camera module, and internal power system. The image illustrates the final physical form factor and layout of the complete sensing node as deployed in classroom environments. . . . . . . . . . . . . . . . . . . . .
9
42
Software architecture flowchart of the SPU, illustrating the five primary subsystems: network registration, device registration, session management, upload logic, and battery monitoring. This diagram summarizes the complete operational workflow executed by the SPU from boot to shutdown. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
10
44
Device registration flow for the SPU. The diagram illustrates the certificatebased onboarding process, including network verification, PSK-based authorization, CSR generation, certificate issuance via STEP CA, and the return of the server-generated public key used for secure local data encryption. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
11
47
Functional block diagram of the SPU recording-session logic, showing the sequential workflow for session control (start/stop based on server signals or user interaction), video capture and encoding, hybrid encryption, local buffering and uploading, and session termination. . . . . . . . . . . . . .
12
48
Functional block diagram illustrating the interaction between the SPU and the BSN backend server. The BSN server stack consists of an NGINX reverse proxy, a Django application server, an on-premises MinIO objectstorage service, and an NTP time synchronization service. The diagram shows how the SPU requests time-limited, upload-only pre-signed URLs and transmits encrypted data to the server over mutually authenticated TLS 1.3 connections. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
xi
51
13
User-interface flow of the SPU software, shown as six sequential screens: (1) splash/loading screen; (2) ULSponsor Wi-Fi registration page shown when the device has no Internet access; (3) device-registration screen for PSK-based certificate provisioning; (4) session-waiting interface displayed after successful onboarding; (5) recording interface shown at session start with the camera disabled pending student consent; and (6) active recording interface displayed once consent is granted. . . . . . . . . . . . . . . .
14
53
Three windows of the developed front-end GUI for classroom calibration. The top-left window is the login page. The bottom window is the calibration page, which supports adding new classrooms and cameras, performing calibration, and visualizing the results. The image in the center of the bottom window shows a live camera stream that has been rectified to remove distortion. The top-right window presents a 3D visualization of a calibrated classroom at the CVIP Lab. Although the classroom contains a single actual screen, additional dummy surfaces (e.g., desk surfaces) were added as objects of interest to test the calibration algorithm. . . . . . . .
15
59
Visual synchronization check of ten SPUs during the clock-alignment experiment. The horizontal axis corresponds to SPU IDs, while the vertical axis represents time samples. Each cell shows the timestamp displayed by an SPU at a given moment. The near-identical timestamps across columns indicate that all SPUs remain tightly synchronized with the server throughout the ten-minute observation period. . . . . . . . . . . . . . . .
xii
67
CHAPTER I INTRODUCTION
I.1
Introduction
Science, technology, engineering, and mathematics (STEM) majors play a crucial role in the economic development of industrialized nations [1, 2]. However, many of these countries continue to face challenges in producing sufficient numbers of STEMtrained graduates [3, 4, 5]. For example, in the United States, projections suggest that the number of STEM degree earners would need to increase by roughly 33% to meet future workforce demands [6]. The major cause of this shortfall is the high attrition rate among students pursuing STEM majors. Studies show that in the United States, 50–60% of students who enter college intending to complete a STEM degree ultimately either switch to a non-STEM field or leave college without earning a degree, despite being academically capable [7, 8, 6]. Poor performance in foundational STEM “gateway” courses is a leading factor in student attrition from STEM majors [9, 7]. Additional contributing factors include a limited sense of belonging, insufficient exposure to authentic research experiences, and the persistent underrepresentation of minority groups in STEM [10, 11, 4, 12]. High attrition rates in STEM majors call for urgent countermeasures, and many strategies have been proposed in the literature. Promising approaches include active learning with peer support [13, 14], corequisite math pathways [15], belonging interventions [16], STEM intervention programs [17], and course-based undergraduate
1
research experiences (CUREs) [18, 19]. Although these initiatives differ in design, they share a common mechanism: fostering student engagement (SE)—behavioral, emotional, or cognitive—as a means of reducing attrition [20, 21]. During the past four decades, SE has been the focus of extensive research, supported by large-scale data sources such as annual engagement surveys conducted across hundreds of higher education institutions [22]. Many studies have shown that higher levels of SE are associated with improved academic performance, higher likelihood of degree completion, lower dropout rates, and reduced negative behaviors within academic settings [23, 24, 25]. These findings underscore the importance of SE in reducing attrition, making it crucial to gain insight into how students engage within courses to improve teaching and learning outcomes and to identify timely interventions for at-risk students. Indeed, engagement has been described as “the holy grail of learning“ [26]. This work addresses the issue of high attrition rates among college students from STEM majors by proposing a novel and effective biometric sensor network that enables real-time measurement and tracking of individual SE during STEM classroom lectures. This BSN is effective in the sense that it can operate in a non-intrusive, non-invasive, non-stigmatizing, automatic, and real-time way while adhering to the regulations and policies imposed by the University of Louisville (UofL) Institutional Review Board (IRB). Furthermore, this BSN leverages recent advances in imaging sensors and edge computing to enable the measurement of SE through state-of-the-art algorithms in computer vision (CV) and machine learning (ML). The organization of this work is as follows. Chapter 2 provides the theoretical background of SE from an educational psychology perspective; it defines SE, describes how it is measured, and reviews recent methods in the literature relevant to this research. Chapter 3 presents the analysis and design of the proposed BSN, including system analysis and key technical design decisions. Chapter 4 describes the hardware and software implementations of the design introduced in Chapter 3. Chapter 5
2
presents the experiments and evaluation, demonstrating the effectiveness of the proposed BSN. Finally, Chapter 6 provides the conclusion and outlines potential future directions.
3
CHAPTER II BACKGROUND
The purpose of this chapter is to provide the necessary background for readers who do not have a foundation in educational psychology. We begin with the adopted definition of SE, followed by a discussion of the distinctions between engagement, flow, and attention, as the latter two constructs—although different from engagement—are closely related to it and are sometimes used interchangeably in the literature. Next, we review the various methods used to measure SE, highlighting the advantages and limitations of each approach.
II.1
Student Engagement Definition
In educational psychology, SE is often one of the most misused and overgeneralized constructs [27]. Therefore, any research that aims to measure SE must begin with a clear and precise definition of what SE is, ensuring that the construct under investigation is conceptually well-grounded before discussing how it can be measured. SE has been widely conceptualized as a meta-construct comprising three interrelated dimensions—behavioral, emotional, and cognitive engagement—as described in the seminal framework proposed by Fredricks et al. [21]. There is no consensus among researchers on a precise definition, and recent studies have proposed additional dimensions such as social engagement [28] and agentic engagement [29].
4
However, it remains unclear whether these represent truly distinct dimensions or whether they can be subsumed under one of the three canonical components. In this work, we adopt the widely cited framework of Fredricks et al. [21], which conceptualizes engagement as a multidi- Figure 1.
Venn diagram to represent
mensional construct consisting of three engagement as a malleable construct in interrelated yet distinct dimensions: be- which each component is interdependent havioral, emotional, and cognitive en- and cannot be completely disentangled gagement. Within this framework, en- from the other two.
The diagram also
gagement is viewed as a malleable con- shows that SE overlaps with related construct in which each component is inter- structs such as attention and flow. dependent and cannot be completely disentangled from the other two. This construct can be best described by the Venn diagram shown in Figure 1. Behavioral Engagement (BE). BE has been defined in multiple ways; in this study, we adopt the definition that frames it as active involvement in learning and academic tasks. This includes behaviors such as effort, persistence, concentration, attention, asking questions, and participating in class discussions [21]. Emotional Engagement (EE). EE refers to students’ affective responses in the classroom, which encompass emotional reactions such as interest, boredom, happiness, sadness, anxiety, and anger [30]. Cognitive Engagement (CE). CE, similar to BE, has been defined in multiple ways, with conceptualizations emerging from research on both school engagement and learning and instruction [21]. In this study, we adopt the definition proposed
5
by Lamborn et al. [31], which emphasizes students’ psychological investment and effort in learning, describing CE as “the student’s psychological investment in and effort directed toward learning, understanding, and mastering the knowledge, skills, or crafts that the academic work is intended to promote.”
II.2
Attention, Flow, and Engagement
Attention and flow are two psychological constructs that, while distinct from SE, are closely related to it and help illuminate the mechanisms underlying students’ involvement in learning. Understanding these constructs is important because they provide insight into how learners allocate mental resources, regulate their focus, and experience immersion during academic activities. Csikszentmihalyi [32] describes attention as a limited form of psychic energy that regulates the stream of consciousness, with its allocation shaping one’s overall life experience. At the neurological level, attention comprises four sub-processes: working memory, top-down sensitivity control, competitive selection, and automatic filtering of salient stimuli [33]. Through selective competition, neural representations derived from salience filters and sensitivity control are chosen to enter working memory, where analysis, decision-making, and planning take place. This framework distinguishes between two types of attention: voluntary (top-down) and involuntary (bottom-up). Voluntary attention relies on sensitivity control, competitive selection, and working memory functioning within a stable feedback loop, whereas involuntary attention is primarily triggered by infrequent, salient stimuli occurring suddenly in time or space (e.g., the blare of an ambulance siren or the flashing lights of a police vehicle). Csikszentmihalyi [32] argues that voluntarily focusing attention on a limited stimulus field is necessary to attain experiences that are subjectively and socially valued, whereas difficulty regulating voluntary attention can contribute to psychopathology. Flow is described as a subjective, intrinsically enjoyable experiential state in which an individual becomes fully absorbed in a task, losing awareness of distractions and
6
even the passage of time [34]. In educational settings, the highest level of SE can be understood as the experience of flow sustained throughout a learning activity. For flow to occur, concentration, interest, and enjoyment must coexist within the activity (e.g., attending a lecture), and thus a measure of engagement can be constructed from these three variables [35]. Although attention and flow are distinct constructs, they are closely connected to SE because they reflect core psychological processes that govern how students participate in learning. Attention determines how cognitive resources are allocated, enabling learners to focus on instructional content, persist in academic tasks, and resist distraction—functions that directly support behavioral and cognitive engagement. Flow, in contrast, represents an optimal experiential state in which attention is fully directed toward the task and is accompanied by heightened interest, enjoyment, and intrinsic motivation. Sustained voluntary attention is a prerequisite for entering the flow state, and the experience of flow itself reflects a high-intensity form of engagement. Together, attention and flow help explain how learners invest effort, regulate emotion, and maintain involvement in academic activities, thereby reinforcing their conceptual connection to the canonical dimensions of SE.
II.3
Engagement Measures
SE assessment in the literature can be classified into six main approaches: student self-report surveys, teacher ratings, observational measures, administrative data, experience sampling methods (ESM), and real-time measurements [36]. The advantages and limitations of each approach are briefly outlined below, with a more detailed discussion devoted to real-time measurement, as our method belongs to this category. Student self-report surveys are the most widely used method. They are low-cost, scalable, and capture students’ subjective perceptions and links between contextual factors and engagement. However, they are limited to emotional and cognitive dimensions, provide only school- or classroom-level insights, treat engagement as a static
7
post-task construct, may not reflect actual behaviors or strategies, and are prone to social desirability bias [36]. Teacher ratings rely on teachers to score students’ engagement across multiple indicators. They are simple, inexpensive, and scalable, but mainly capture behavioral engagement, while cognitive and emotional dimensions are harder to infer. They are also subject to bias from Figure 2. The six primary methods for both student and teacher characteristics measuring SE reported in the literature. [36]. Administrative data, such as attendance and grades, measure behavioral engagement. Collected routinely on all students, they allow longitudinal tracking and early intervention. Yet, they conflate outcomes with indicators of engagement, are biased by student characteristics, and lack standardization across institutions [36]. ESP is rooted in flow research [37]. Students are randomly interpreted via ESM signals to fill out short surveys about their location, activities, behavior, and cognitive and affective responses. It offers time and context-dependent measures of student subjective experience, allowing data on engagement to be collected as it happens. These methods can be applied repeatedly to a large population, allowing for comparison of SE levels across time and context. The cons of this technique include the necessity of a high level of commitment from participants. suffers from participant fatigue, hasty completion exaggeration, and deliberation falsification [36]. Observational measures involve trained coders who systematically watch students and apply a predefined coding scheme to record specific behavioral indicators within a fixed period of time. These measures can capture behavioral engagement at the student, group, or classroom levels, provide descriptions of both engagement and its
8
surrounding context, and are generally less disruptive to classroom activities. However, they also require decisions regarding the sampling period and units of analysis; they are labor-intensive, limited to relatively small populations, susceptible to observer bias, and difficult to generalize across different educational contexts [36].
II.4
Real-Time Engagement Measures
All methods discussed so far are limited by coarse temporal granularity, capturing engagement at low-resolution timescales or lacking automated data collection. These approaches are therefore unable to detect rapid fluctuations in SE or provide continuous, fine-grained measurement during learning activities. In contrast, a more recent and promising approach, real-time measures, has emerged, enabled by advances in CV,
Figure 3. The five primary types of signals used as real-time indicators to assess SE as reported in the literature.
ML, sensor technologies, computer networks, and edge computing. This approach tracks SE dynamics at temporal scales as fine as seconds or minutes by relying on discrete and objective indicators. In the literature, five main types of indicators are commonly used to assess SE: log files [38, 39], eye tracking [40], facial expressions [41, 42], body actions [43, 44], and physiological signals such as electroencephalogram (EEG), heart rate variability (HRV), and galvanic skin response (GSR) [45]. Using real-time measures offers several advantages: it is objective and less prone to student or teacher biases, provides detailed information about how engagement changes over time, can be automated to reduce interruptions, and allows the collection of large amounts of data in a short period [36].
9
However, this approach is not without limitations. Real-time data collection requires technical expertise, is often conducted in small, structured tasks rather than dynamic classroom environments, raises privacy concerns, and can be costly in terms of data processing and storage. Furthermore, there is no consensus on the optimal temporal granularity for measuring engagement, and the resulting data must be carefully analyzed—typically using CV and ML algorithms—which may lack interpretability and require careful validation to produce meaningful results [36]. Log files are outside the scope of this work as they are typically used in e-learning environments. Physiological signals can be applied in classrooms, but often require sophisticated, intrusive, or invasive sensors and are affected by body movements or sweat, making them less suitable for our purposes. Eye tracking records learners’ gaze patterns to reveal visual attention and engagement with specific areas of interest (AOIs), such as the blackboard or teacher [40]. It is closely related to cognitive load (CL) and is considered an important indicator of cognitive engagement [46, 47]. Facial expressions (FE) can be described in terms of facial action units (FACS) [48] and can convey internal human emotions [49]. Many algorithms have been proposed for automatic recognition of facial expressions (FER) from camera images or videos, and comprehensive reviews of this research are available in [50, 51]. FER has been explored as a promising tool to automatically measure students’ emotional engagement [36]. Body action refers to observable movements or postures of the human body, which can be analyzed to infer activities, intentions, or emotional states [52, 53]. In CV, body action recognition is defined as the process of labeling sequences of body movements captured through cameras or sensors into predefined action categories [54, 55]. This approach has been successfully applied to measure student behavioral engagement in classrooms by analyzing students’ actions during lectures [43, 56].
10
II.5
Literature Review (zReal-time Methods)
Many methods have been proposed in the literature to measure one or more dimensions of SE based on real-time indicators. In this section, only methods that rely on visual cues are discussed. Sheng et al. [57] introduced an improved YOLOv8-based detector to localize and classify student behaviors from single classroom images; however, their work stopped at behavior detection and did not link the detected behaviors to validated measures of engagement, nor did it evaluate whether behavior counts or temporal patterns could predict engagement levels. Nezami et al. [58] constructed a private dataset of SE comprising frontal facial images of students interacting with a virtual world environment for science education. Engagement and disengagement labels were assigned by annotating images with both behavioral and emotional cues. Their deep learning model achieved a classification accuracy of 72.38%; however, the approach treats engagement as a static construct that can be inferred from a single image. This assumption overlooks the inherently dynamic and temporal nature of engagement. Furthermore, the method is tailored to e-learning or virtual learning contexts and does not readily generalize to traditional classroom settings, where social interactions, group dynamics, and environmental complexity play a critical role in shaping engagement. Singh et al. [59] addressed key limitations in prior engagement research by introducing EngageNet, a large-scale dataset comprising 31 hours of video from 127 participants captured in diverse real-world conditions. Unlike earlier datasets that primarily conceptualized engagement as a static construct inferred from short clips, EngageNet incorporates behavioral, cognitive, and self-reported measures, providing a more nuanced and multidimensional view of engagement. The authors also established strong baseline models using facial action units, gaze, head pose, and transformer-based architectures, and validated their approach across both EngageNet and the EngageWild dataset to demonstrate generalizability. A key limitation, how-
11
ever, is that the data collection setup assumes students are consistently facing a laptop, which restricts ecological validity; in real classroom settings, students often shift attention across multiple targets such as the blackboard, projector screens, and the lecturer.
12
CHAPTER III ANALYSIS AND DESIGN
This chapter details the theoretical analysis on which we proposed a novel BSN to enable measuring the engagement of individual students in a college STEM classroom while attending a lecture. This BSN, with the right set of algorithms deployed on it, can help mitigate the problem of high attrition rates among STEM majors by providing fast, objective feedback to professors and decision makers, enabling early intervention for students at risk. This analysis serves as the foundation upon which we design an effective BSN. The outcome of this chapter is the specification and requirements of the proposed BSN. This chapter is pivotal to this work: Chapter 4 presents an implementation of the BSN requirements developed here, while Chapter 5 describes the experiments conducted after implementation to evaluate the system’s performance and confirm that it satisfies the specifications and requirements established in this chapter.
III.1
Analysis
Problem: High Attrition Rate Among Students Pursuing STEM Majors Studies show that, in the United States, for example, 50–60% of students who enter college intending to complete a STEM degree ultimately either switch to a non-STEM field or leave college without earning a degree, despite being academically capable [7, 8, 6]. 13
Root Cause: Poor Performance in Foundational STEM Gateway Courses Poor performance in gateway STEM courses is a major factor contributing to student attrition from STEM pathways. These introductory, high-demand courses often determine whether students persist in a STEM major or choose to switch to a nonSTEM field, and in many cases, they can even influence a student’s decision to leave college entirely [9, 7]. Difficulties in mastering foundational material, combined with the pressure and pacing of these courses, create a critical barrier that disproportionately affects students who otherwise might succeed in STEM if provided timely support. Mitigation: Monitor Student Engagement During STEM Classes and Provide Support for Students at Risk A substantial body of research has demonstrated that higher levels of SE are linked to better academic performance, greater likelihood of degree completion, lower dropout rates, and fewer negative behaviors in educational environments [23, 24, 25]. Collectively, these findings highlight SE as a critical factor in reducing attrition and emphasize the need to understand how students engage within courses in order to enhance teaching and learning and to support timely interventions for students at risk. Indeed, engagement has even been described as “The holy grail of learning.” [26]. Goal: An Effective System for Automatically Measuring Individual Student Engagement in Real Time in STEM Classrooms This work aims to propose a system that enables the measurement of individual college students’ engagement effectively while attending a lecture in a STEM classroom, and does so in a manner that is both ethically acceptable and technically viable. To be precise, each term highlighted in bold in the previously stated aim must be clearly defined and operationalized. 14
Measurement Framework A central challenge in designing a system capable of measuring SEt is determining how the abstract construct of engagement can be translated into measurable components. Following the guidance of [26], we structure this analysis around three fundamental questions: 1. What is the construct definition utilized? 2. What is the grain size for time, task, and agent? 3. What are the indicators to be collected? Construct Definition. We adopt the multidimensional framework of [21], which characterizes engagement in terms of behavioral, emotional, and cognitive components. These components are described in detail in Section II.1. Importantly, the system cannot directly measure these dimensions; rather, it can capture only proxy indicators. To map these proxy indicators to meaningful engagement estimates, complementary data—obtained via experience sampling methods (ESM) or systematic observation—must be collected to provide the ground truth for the measures generated by the proposed system. Grain Size. Following [60], time is operationalized at a resolution of 10 seconds; the task is defined as attending a lecture in a STEM classroom; and the agent corresponds to the individual student being assessed. With this construct definition and grain size, our conceptualization aligns closely with that of Symonds et al. [61]. The discussion of indicators is postponed until after the next subsection, as the measurement methods—and therefore the indicators themselves—depend on the objectives established in that section.
15
Effectiveness By an effective system, we mean a system capable of mitigating the problem at hand. This necessitates that the system introduce minimal disruption to the normal flow of the lecture. Otherwise, the system may produce biased measurements if students notice that they are being monitored and consequently alter their behavior. Worse still, the system could contribute to additional disengagement if its presence in the classroom causes students to feel anxious or uncomfortable. Precisely for the system to be effective, it must fulfill five objectives: 1. Non-intrusive
3. Non-stigmatizing
2. Non-invasive
4. Real-time
5. Automatic
We acknowledge that the five objectives just mentioned above are somewhat broad and require greater precision. Therefore, each objective is defined more explicitly as follows. A Non-intrusive system does not disrupt classroom activities, distract students, or alter their natural behavior. More concretely, a system is non-intrusive if it does not intercept the direct line of sight between the student and the professor, the projector, the blackboard, the student’s notes, or the student’s laptop. A non-invasive system requires no physical contact between itself and the student and avoids any procedure that interacts with the body in a medical sense. Concretely, a system is non-invasive if the only physical contact with it occurs when a student voluntarily interacts with it at the beginning of class to provide consent and initiate the system; no sensors may be worn, attached, or applied to the student’s body at any point during operation. A non-stigmatizing system ensures that no student is singled out, labeled, or made to feel judged as a result of the sensing process. Specifically, a system is nonstigmatizing if students are free to choose any available seat in the classroom, and it does not require designated seating or special placement for individual students. 16
Furthermore, no student may be marked or tagged in any way during operation, as such practices would differentiate students from their peers and risk creating feelings of discomfort or stigma. A system is called real-time if it captures, processes, and provides engagementrelated information within a specified time window ∆t. A system is called automatic if, once initialized, it requires no further human intervention during the operation. Engagement Indicators As detailed in Section II.3 and Section II.4, six methods are available in the literature to measure SE. However, not all of these methods align with the five objectives established in the previous subsection for the system to be effective. The goal of this discussion is to determine which methods, and therefore which indicators, are suitable for our system. The student self-report surveys and administrative data methods do not fulfill the real-time objective and are therefore excluded from consideration. The teacher ratings method satisfies neither the real-time nor the automatic objectives and is also unsuitable for our system. The observational measures method fulfills four of the stated objectives; however, it is not automatic. Despite this limitation, it remains a valuable tool for validating the results produced by an automatic, real-time system such as the one we aim to design. An automatic SE measurement system typically requires a labeled dataset (ground truth) for training. Because such data raises privacy concerns, no publicly available high-quality SE datasets exist for college STEM classrooms. This necessitates the collection of our own private dataset, and thus the system must support this capability. It is important to emphasize that dataset collection is a temporary step performed only once for system validation. In this context, observational methods can be used to manually annotate the collected data, which is acceptable since
17
annotation is required only during development and not during deployment. Finally, because data collection is a serious matter, the system must ensure that all collection procedures adhere strictly to the regulations and policies of the University of Louisville Institutional Review Board. The ESM method does not satisfy the real-time, automatic, or non-intrusive objectives and therefore cannot be used in our system. However, similar to observational measures, it may still be employed to collect direct indicators from students immediately before or after class sessions to help validate the results generated by the automatic system. This leaves us with the final method, real-time measures. As the name suggests, this method supports measuring SE on a timescale ranging from a few seconds to a few minutes. Moreover, when combined with appropriate CV and ML algorithms and a high-quality dataset, it can be fully automated. Real-time methods support five types of indicators, as detailed in Section II.4; however, not all of them align with the five objectives stated previously. For example, HRV and EEG require invasive sensors and are therefore unsuitable. Fortunately, three of the five indicators—facial expressions, body actions, and eye gaze—can be collected using noninvasive sensors, such as a camera. The remaining two objectives, non-intrusive and non-stigmatizing, depend on the physical deployment of the system and will be proven later. To conclude this section, the real-time measures method is the most suitable among the available approaches, as it aligns with the five objectives stated in the previous subsection. This method enables the indirect measurement of the three components of engagement using three different camera-based biometric indicators, as follows: 1. Behavioral engagement: via body actions and movement patterns. 2. Emotional engagement: via facial expressions and affective cues. 3. Cognitive engagement: via eye movement patterns and gaze direction.
18
Ethics Ethical considerations are central to the design and deployment of any system intended to measure SE, particularly when it relies on biometric or video-based sensing. Such a system must respect student autonomy, protect privacy, and minimize any potential harm or discomfort arising from its presence in the classroom. This requires clear and transparent communication regarding what data are collected, how they are used, and who has access to them, as well as ensuring that participation is voluntary and informed. Moreover, data collection and storage procedures must adhere strictly to the regulations and policies of the Institutional Review Board (IRB), with safeguards in place to prevent misuse, unauthorized disclosure, or identification of individual students. Ultimately, the ethical foundation of the system must prioritize student well-being and trust, ensuring that the measurement process enhances learning rather than compromising the dignity or rights of those being observed. Environment The proposed system operates within a traditional college STEM classroom—an inperson lecture setting in which students are physically present, and an instructor delivers course material using a projector, whiteboard, or a combination of instructional media. We formalize this environment as the tuple (E, S, P ), where E denotes classroom calibration parameters, S is the set of students, and P (t) represents the professor’s feature vector at time t. The student set is defined as S = { si (t) | 1 ≤ i ≤ NC }, where si (t) is the feature vector describing student i at time t, and NC is the maximum classroom capacity. Both si (t) and P (t) evolve over time, reflecting the dynamic and interactive nature of a live classroom environment. 19
The calibration parameters E capture all fixed properties of the classroom relevant to system operation. These include the 3D geometry of the room, the locations and poses of instructional elements such as the blackboard and projector screen, and the placement and intrinsic parameters of fixed cameras used during development mode. In general, E consists of any static quantities required by the BSN for accurate sensing, alignment, or spatial reasoning throughout a class session. Technical Viability For the proposed system to be technically viable, it must be feasible to deploy, operate, and maintain within the constraints of a typical college STEM classroom. This requires that the system rely on affordable, readily available hardware; run efficiently on resource-constrained edge devices; and integrate robust CV and ML algorithms capable of processing data in real time. The system must also tolerate variability in classroom conditions, including lighting, seating arrangements, and student movement, without significant degradation in performance. Furthermore, the system should be scalable, modular, and compatible with standard networking and data-management practices to support long-term deployment. In essence, technical viability ensures that the system can function reliably in real-world educational environments, not only in controlled laboratory settings. Building on these insights, the next section translates these engagement indicators and objectives into concrete system requirements that guide the design of the proposed BSN. Sensing Technology From the analysis conducted in the previous subsections, we conclude that the sensing technology must be camera-based and that three types of indicators can be captured: facial expressions, eye gaze, and body actions. The next question, therefore, is which
20
topology should be used to deploy these cameras. The answer to this question is addressed in the next subsection. Camera Deployment Topology Accurate capture of facial expressions and eye gaze requires high-resolution images of the student’s face, which in turn necessitates positioning the camera close to the student’s face—typically on the student’s desk, facing the student alongside their laptop or tablet—to record subtle details. Such a camera can operate with a narrow field of view (FOV), and its pose can be freely adjusted to ensure that the student’s upper body, particularly the face, is consistently within view. In contrast, capturing body actions requires cameras to be placed at a greater distance, facing students from above, so that the student’s full upper body remains within the FOV. These cameras can maintain a fixed position and are typically mounted on classroom walls or ceilings, with each camera covering multiple students within its field of view. This distinction yields two possible deployment topologies: a dynamic topology, in which each student is equipped with a device we call a student processing unit (SPU) having a narrow-FOV camera positioned freely across the classroom; and a static topology, in which a fixed number of wide-FOV cameras are deployed per classroom. The question, then, is which topology is most suitable for our system. To address this, we now examine the advantages and disadvantages of each approach. Static Topology The static topology offers advantages in scalability and cost. A single classroom, including those with large capacities, can be covered using only a few cameras. Moreover, this approach does not require specialized hardware; any off-the-shelf IP camera with a suitable field of view and resolution can be used.
21
However, this topology also presents several drawbacks. First, it does not satisfy the non-stigmatizing objective. For example, in a classroom where some students prefer not to be monitored or recorded, it becomes nearly impossible to honor their preferences without restricting their freedom to choose their seats. Second, because the cameras are positioned far from students’ faces, they cannot capture high-quality facial imagery and therefore cannot reliably extract facial expressions or eye gaze. Third, the topology raises privacy concerns since images must be transmitted to a server for processing rather than processed at the edge. Fourth, while camera installation requires only a one-time effort, the fixed placement may affect classroom usability for other courses, as students unaffiliated with the monitoring system may feel uncomfortable due to the camera presence. Despite these limitations, the static topology is well-suited for collecting an SE dataset that can later be annotated using observational protocols. Dynamic Topology The dynamic topology offers several advantages. First and foremost, it satisfies the non-stigmatizing objective, since each student is assigned a dedicated SPU and is therefore free to choose any seat in the classroom. Second, it mitigates privacy concerns because all processing occurs on the edge; no student-identifiable data, such as images or videos, is transmitted to a centralized server for further processing. Furthermore, if a student feels uncomfortable for any reason, the SPU can be turned off or removed entirely. Third, this topology reliably captures two of the engagement indicators, namely facial expressions and eye gaze. Finally, once the class session ends, the SPUs can be collected quickly and stored securely, ensuring that the classroom can be used normally for other courses without affecting students not involved in the monitoring process. Despite its versatility, the dynamic topology also has disadvantages. The system cost scales linearly with classroom size, since each student requires an SPU. This is
22
not a major issue for small and medium-sized classrooms, which are typical in STEM courses. Additionally, the topology requires designing a custom unit—i.e., the SPU. However, this is a one-time effort and can be accomplished using readily available edge-computing devices such as the Raspberry Pi[62] or NVIDIA Jetson[63]. A further limitation is that the SPU has restricted computational capability, which may prevent it from running computationally heavy algorithms. This challenge can be mitigated by selecting efficient algorithms or employing techniques such as quantization and pruning to reduce computational cost. Finally, the dynamic topology cannot reliably capture students’ body actions. In summary, the dynamic topology is the only configuration that satisfies the non-stigmatizing objective and is therefore the topology of choice for deployment. However, during the development phase, the static topology is necessary for capturing and recording our private SE dataset, which will later be annotated by trained coders and used to validate and benchmark the proposed system. The Biometric Sensor Network We are now ready to introduce the concept of a Biometric Sensor Network (BSN), which consists of two types of sensing nodes: dynamic nodes, called Student Processing Units (SPUs), and static nodes, called Classroom Fixed Units (CFUs). Both node types communicate with a central server over a network to coordinate sensing, processing, and data management. To ensure that the BSN satisfies the remaining objective of being non-intrusive, the network technology must be wireless, and each SPU must be battery powered. This prevents cable clutter in the classroom and allows the SPU to be placed or repositioned freely, preserving the natural classroom environment and students’ freedom of movement. In addition, the BSN includes a front-end client in the form of a web-based dashboard, which provides a unified interface to control, monitor, and operate the entire BSN during classroom sessions. In addition, the BSN supports two modes of operation: a development mode and
23
a deployment mode. In development mode, both SPUs and CFUs are used to capture raw video of students for the purpose of constructing a private SE dataset. The CFU footage is later annotated by trained coders to produce engagement labels for each 10-second segment for every student. In deployment mode, CFUs are no longer used. Instead, each SPU runs the necessary algorithms locally to infer SE from 10-second video segments captured by its own camera, and then transmits the inferred SE values to the central server for storage and real-time monitoring.
III.2
System Requirements
Having established the structure, operational modes, and functional objectives of the proposed BSN, we now turn to the system requirements and design considerations necessary to realize this network in practice. The next subsections detail the hardware and software components of the SPUs and CFUs, the communication architecture, and the processing pipeline that enables real-time engagement inference within classroom environments. SPU Hardware Requirements The Student Processing Unit (SPU) serves as the core sensing and computation node in the dynamic topology of the BSN, and its hardware requirements are defined to ensure that it can operate reliably, securely, and in compliance with the non-intrusive, non-invasive, and non-stigmatizing design objectives. Each SPU must integrate a high-quality camera capable of capturing facial expressions and eye gaze, a compute module capable of running lightweight CV/ML inference models in real time, and a touchscreen interface through which students can provide consent at the beginning of each session. To preserve classroom flexibility and avoid cable clutter, the SPU must operate wirelessly and be powered by a battery system that provides safe and uninterrupted operation throughout the class period. In addition, precise clock
24
synchronization is required to align inference windows across devices, and encrypted temporary storage is necessary to support secure buffering before data transmission. The enclosure plays an essential role in both usability and security. It must be compact, non-distracting, and thermally well-ventilated, while also ensuring strong physical security. In particular, external ports such as USB or Ethernet must be physically hidden, covered, or internally routed so that students cannot access them during operation. This prevents unauthorized data extraction, device tampering, or circumvention of system security controls. Collectively, these requirements ensure that the SPU functions as an effective, secure, and ethically responsible sensing device within the proposed BSN. The complete set of hardware requirements for the SPU is summarized in Table 1. Table 1. SPU Hardware Requirements Component
Requirements
Camera Module
High-resolution 720p+ camera, narrow FOV, low-light capable, minimum 15 fps.
Compute Unit
ARM-based single board computer (SBC) (e.g., Raspberry Pi 5 [62]) with optional GPU support for real-time CV/ML inference.
Touchscreen Interface
Integrated touch display for collecting student consent before each session, displaying system status, and enabling basic interaction and control. The touchscreen must automatically turn off within 30 seconds after consent is provided to minimize distraction, and it must be reactivatable by touch to allow the student to end the recording session or power off the device if desired.
25
Component
Requirements
Battery System
Minimum 2 hours of continuous use; Maximum charging time is 2 Hours; Overcharge, discharge, overcurrent, and thermal protection.
Network Interface
Wireless interface (Wi-Fi 5/6) for reliable, low-latency communication with the central server.
Clock Sync.
NTP-based time alignment with jitter less than 500 ms.
Enclosure
Compact, non-obstructive housing with proper thermal ventilation; no distracting lights or visible indicators; the enclosure must mechanically support the full weight of the SPU and be designed with appropriate balance so it can stand stably on a desk without additional mounts, adhesives, supports, or screws.
Physical Security
All external ports (e.g., USB, Ethernet) must be physically hidden or rendered inaccessible during operation, exposing only a power button and a protected batterycharging port to prevent tampering or unauthorized data access.
Local Storage
Supports at least 30 GB of encrypted temporary buffering; automatic deletion after processing and upload.
CFU Hardware Requirements The Classroom Fixed Unit (CFU) serves as the static sensing node within the BSN and is used primarily during the development phase to support large-scale data collection and annotation. Unlike the SPU, which operates in proximity to individual students, the CFU is positioned at fixed locations in the classroom—typically overhead—to capture wide-angle views of multiple students simultaneously. Its primary
26
purpose is to enable the construction of a labeled SE dataset through synchronized video capture that can later be annotated by trained coders. Consequently, the hardware requirements for the CFU emphasize high-resolution wide-field imaging, stable physical installation, reliable network connectivity, and precise time synchronization. These requirements are summarized in Table 2 and ensure that the CFU can operate continuously, accurately, and securely during classroom recording sessions. Table 2. CFU Hardware Requirements Component
Requirements
Wide-FOV Camera
92–120° FOV, overhead mounted; 4K resolution.
Mounting System
Ceiling/wall fixed mount.
Network Interface
Wired (10 Gbs Ethernet) or wireless (Wi-Fi 5/6); consistent streaming performance.
Power Supply
Permanent AC power or Power Over Ethernet (PoE).
Clock Sync.
NTP-based time alignment with jitter less than 500 ms.
SPU Software Requirements The software stack running on each SPU is critical to ensuring that the device operates securely, autonomously, and in real time, in accordance with the system objectives established in the earlier analysis. Because the SPU performs all engagement inference locally—without transmitting raw video—the software must support efficient on-device CV/ML processing, secure communication protocols, and reliable synchronization with the central server. In addition, the software must enforce strict security guarantees, including a hardened operating system, encrypted data handling, and authenticated over-the-air updates, to prevent tampering and protect student privacy. Robust power management and fault-tolerant networking further ensure continuous operation throughout a class session. Table 3 summarizes the key software requirements necessary for the SPU to function as an effective, secure, and privacy-preserving
27
sensing node within the BSN. Table 3. SPU Software Requirements Component
Requirement
Operating System
Hardened 64-bit Linux distribution with minimized attack surface and disabled unnecessary services.
CV/ML Inference En-
Real-time models for face detection, gaze estimation,
gine
and behavioral analysis optimized for edge devices.
On-device Processing
All inference is performed locally; raw video must never be stored or transmitted.
Time Synchronization
NTP-based synchronization with maximum jitter of 500 ms to align 10-second inference windows.
Communication
TLS 1.3 encrypted communication with automatic retry, reconnection, and failover handling.
Power Manager
Battery monitoring, low-power alerts, and controlled safe-shutdown logic.
Wi-Fi Registration
Software must support onboarding SPUs onto the UofL Wi-Fi network using approved campus authentication protocols.
Node Registration
SPUs must obtain a signed SSL certificate from the server via a one-time pre-shared key to enable authenticated operation in designated classrooms.
Automatic
Session
Control
The software must automatically detect active class sessions, enter session mode at start time, and exit session mode upon session completion.
Consent Gatekeeping
The SPU must not activate its camera unless the student provides explicit consent through the touchscreen interface at the beginning of each session.
28
Component
Requirement
System Status Bar
The GUI must display clock, Wi-Fi status, certificate validity, secure-connection status, and battery level.
System Power Con-
The GUI must provide a safe interface for shutdown and
trols
reboot operations.
OTA Updates
Support for authenticated, signed over-the-air updates for both software and firmware.
Security and Data Protection for Edge Nodes The edge devices in the BSN, namely the SPUs and CFUs, must enforce strong security and privacy protections, as they operate in close physical proximity to students and handle potentially sensitive visual data. Security at the edge focuses on minimizing the attack surface, ensuring that devices cannot be tampered with, and guaranteeing that no raw video or identifying biometric information ever leaves the device. To achieve this, the edge nodes must implement encrypted storage, authenticated boot mechanisms, strong device-level authentication, and strict access control. In addition, all processing of engagement indicators must occur locally to preserve privacy, with only nonidentifying features transmitted to the server. These requirements ensure that edge devices remain secure, privacy-preserving, and compliant with institutional regulations. The key security and data-protection requirements for SPUs and CFUs are summarized in Table 4. Table 4. Security and Data Protection Requirements (Edge Devices: SPUs/CFUs) Aspect
Requirement
Encryption
Local storage must use AES-256 encryption; raw video must never be transmitted off-device.
29
Aspect
Requirement
Secure Boot
Device must verify firmware integrity at boot to prevent unauthorized modifications.
Device Authentication
Mutual TLS authentication using hardware-backed certificates.
Data Minimization
Raw video must be deleted immediately after inference; only nonidentifying indicators may be sent to the server.
Access Control
SSH must be disabled or restricted to key-based access; no default or shared credentials.
Privacy by Design
Only behavioral indicators may be extracted; no biometric identification or student re-identification allowed.
Server Software Requirements The central server forms the core coordination and data-management layer of the BSN. It must provide secure and reliable services for authenticating users and devices, ingesting engagement data from SPUs, coordinating active classroom sessions, and supporting real-time monitoring through a web-based dashboard. Because SPUs operate autonomously at the edge, the server is responsible for validating node identity, authorizing access, aggregating engagement estimates, and maintaining system-wide consistency through periodic health checks. In addition, the server must implement robust analytics capabilities, enforce role-based access control, and maintain comprehensive audit logs to support debugging, anomaly detection, and institutional compliance. The key software requirements for the server are summarized in Table 5.
30
Table 5. Server Software Requirements Component
Requirement
Backend API
Provides APIs for user registration and authorization, node registration and authorization, periodic health checks, and secure ingestion of engagement data.
Database
Stores engagement time-series data with AES-256 encryption and automated backup rotation.
Dashboard Interface
Web interface for system monitoring, configuration, and real-time engagement visualization.
Analytics Engine
Computes aggregated engagement metrics per student, class session, and time window.
User Management
Enforces role-based access control (e.g., admin, instructor, researcher).
Logging and Auditing
Maintains comprehensive logs for debugging, anomaly detection, and regulatory compliance.
Security and Data Protection for Server The server infrastructure of the BSN must enforce strong security and privacy guarantees, as it serves as the central point for data aggregation, device identity validation, and long-term storage of engagement estimates. In addition to providing encrypted transport channels and secure storage for engagement data, the server must support an on-premises, access-controlled object storage solution for dataset archival and operate its own certificate authority to issue and manage SPU certificates, including revocation when necessary. Robust identity and access management—with MFA and role-based access control—is required to restrict system operations to authorized personnel. Furthermore, the server must support intrusion detection and incident response mechanisms and provide the ability to isolate compromised nodes
31
to preserve overall system integrity. All security controls must comply with institutional and regulatory requirements such as IRB and FERPA. The principal security and data-protection safeguards implemented at the server level are summarized in Table 6. Table 6. Security and Data Protection Requirements (Server) Aspect
Requirement
Transport Security
All device–server and user–server communications must use TLS 1.3.
Database Protection
Engagement data stored at rest must use AES-256 encryption with automated secure backup rotation.
Identity
&
Access
System must enforce MFA and role-based access control
Management
(RBAC) following least-privilege principles.
Secure Object Storage
Supports on-premises, access-controlled S3-compatible storage (e.g., MinIO) for dataset collection and archival.
Certificate Authority
Issues and manages SPU certificates through an on-
(CA)
premises CA to support authentication and certificate revocation.
Incident Response
Server must support intrusion detection/prevention (IDS/IPS) and isolate or revoke compromised SPUs.
Compliance
All processing and storage must comply with IRB, FERPA, and institutional data-protection policies.
III.3
Summary
Figure 4 summarizes the design process that led to the adoption of an SPU-based BSN. Beginning with six established methods for measuring SE, each was evaluated against the five required system objectives: non-intrusive, non-invasive, nonstigmatizing, real-time, and automatic. Only the real-time measurement approach
32
satisfies these criteria. Within this method, a further evaluation of available indicators showed that facial expressions, eye gaze, and body actions are the only signals that can be captured using non-invasive and non-stigmatizing technology suitable for a live classroom environment. These constraints naturally lead to a camera-based system, which must be portable, wireless, battery-operated, and capable of on-device processing. The resulting design decision converges on the SPU architecture, which integrates a single-board computer, camera module, wireless interface, local storage, battery subsystem, and touch display for consent and control.
33
34
Figure 4. Design-decision workflow summarizing the selection of SE measurement methods, real-time indicators, and sensing technology leading to the final SPU-based BSN architecture.
CHAPTER IV BSN IMPLEMENTATION
This chapter presents the complete hardware and software implementation of the proposed BSN, translating the system requirements derived in the previous chapter into a functional, deployable platform. Whereas the earlier analysis established what the system must achieve and why each design choice is necessary, this chapter describes how the system is realized in practice. The discussion is organized around the two principal components of the BSN: the Student Processing Units (SPUs), which operate as portable edge nodes responsible for real-time engagement inference, and the central server, which coordinates device authentication, session management, data aggregation, and visualization. On the hardware side, the chapter details the SPU system architecture, component selection, integration, and configuration. On the software side, we describe the development of the SPU software stack for on-device CV/ML inference, consent collection, network communication, and secure operation, as well as the server-side backend responsible for data ingestion, analytics, access control, and system monitoring. Together, the hardware and software implementations presented in this chapter constitute a fully functional BSN capable of operating in both development and deployment modes, supporting dataset collection, real-time engagement measurement, and secure classroom integration. This chapter provides all engineering details necessary
to
reproduce,
evaluate,
and 35
extend
the
proposed
system.
IV.1
SPU Hardware
As established in the previous chapter, the SPU hardware must integrate six essential functional components: a singleboard computer (SBC), a camera module, a touch display, a wireless communication interface, a battery management subsystem, and a local storage unit. Fig- Figure 5. System architecture of the ure 5 illustrates the overall architecture SPU hardware, illustrating the interaction of the SPU and the interactions between between the SBC, camera module, touch these components. In this section, we display, wireless network interface, battery identify suitable off-the-shelf hardware management subsystem, and local storage. elements that fulfill these requirements and provide a practical realization of the proposed design. SBC A key design decision concerns the choice of a single-board computer (SBC), as it directly affects the computational throughput, latency performance, power consumption, physical footprint, and available peripheral interfaces. Several commercially available SBCs were evaluated against these criteria, and a summary of their characteristics is provided in Table 7. The comparison accounts for processing capability, device dimensions, cost, typical power usage, and availability of required interfaces such as SD card storage, USB3, Wi-Fi, DSI, and I2 C. Based on this evaluation, the Raspberry Pi 5[62] was selected because it offers an optimal balance between performance, energy efficiency, cost, and ecosystem maturity for the SPU’s operational 36
setting. In addition, the Raspberry Pi 5[62] simplifies the overall hardware design by integrating both a Wi-Fi 6 wireless interface and a native microSD storage slot, thereby eliminating the need for external networking modules and removable storage hardware. SBC Raspberry Pi 5 [62] Raspberry Pi 4B [64] Jetson Orin Nano [63] Orange Pi 5 Plus [65] ROCK 5B [66] BeagleBone Black [67] BeaglePlay [68]
Compute Power 4× Cortex-A76 @ 2.4GHz 20–25 GFLOPS 4× Cortex-A72 @ 1.8GHz 10–15 GFLOPS 6-core ARM + Ampere GPU 40–67 TOPS RK3588 (4×A76 + 4×A55) 0.8 TFLOPS + NPU 6 TOPS RK3588 (4×A76 + 4×A55) 0.8 TFLOPS + NPU 6 TOPS Cortex-A8 @ 1GHz 0.5 GFLOPS 4× Cortex-A53 + PRUs 2 GFLOPS
Size∗ (w×l×h) 85×56×18 mm
Price $45–145
Power∗ 2.5–12W
85×56×17 mm
$35–85
2.7–8W
100×80×30 mm
$249
7–25W
100×64×18 mm
$130–220
5–15W
100×72×18 mm
$150–220
3–15W
86×53×17 mm
$70
1–2W
100×100×20 mm
$100–120
2–5W
Peripherals SD Card, USB3, WiFi 6, DSI, I2C SD Card, USB3, WiFi 5, DSI, I2C SD Card, USB3, WiFi 6, DSI, I2C SD Card, USB3, DSI, I2C SD Card, USB3, DSI, I2C SD Card, I2C SD Card, USB3, WiFi 5, I2C
Table 7. Comparison of candidate single-board computers (SBCs) evaluated for the SPU. Metrics include compute performance, physical dimensions, cost, power requirements, and required peripherals. ∗
Size and power consumption values may vary depending on cooling accessories (e.g., heatsinks, fans), workload intensity, and vendor revisions. Values should be interpreted as typical operating ranges rather than strict guarantees.
Camera Module The next design decision concerns the selection of the camera module to be integrated into the SPU. The Raspberry Pi 5[62] supports two primary camera interfaces: the built-in CSI (Serial Camera Interface) for dedicated camera modules and the USB 3.0 port for higher-bandwidth external vision systems. Both categories offer viable options, and the choice depends on the required resolution, field of view, low-light performance, and the computational demands of real-time engagement inference. While standard Raspberry Pi camera modules provide high-quality imaging with low power consumption and tight hardware integration, they rely entirely on the SBC for image processing. In contrast, modern smart camera modules incorporate on-board accelerators that can execute CV/ML workloads, thereby reducing the com-
37
putational burden on the Pi 5. This is especially important for running face detection, gaze estimation, and expression recognition in real time. After evaluating available options, the DepthAI OAK-D Lite [69] camera module was selected as the preferred choice. It features a 4K RGB sensor capable of capturing 4K video at 30 frames per second (fps) and Full HD (FHD) video at 60 fps. In addition, it includes a mono-stereo depth pair operating at a resolution of 480p with frame rates of up to 120 fps, along with an integrated Intel Myriad X[70] AI accelerator providing up to 4 TOPS (approximately 1.4 TOPS usable for CV/ML workloads). This onboard processing capability substantially offsets the Raspberry Pi’s limited hardware acceleration and enables efficient real-time inference on the SPU. Furthermore, the device is powered directly from the Pi’s USB port, simplifying the SPU hardware design by eliminating the need for additional power rails or external converters. The OAK-D Lite is also considerably more cost-effective than other depth-enabled, accelerator-equipped camera modules, making it an ideal choice for large-scale classroom deployment across SPU nodes. A comparison of possible camera options is provided in Table 8. Camera Module
Interface
Pi Camera Module 3 [71] OAK-D Lite [69] OAK-D S2 [72] OAK-D Pro [73] RealSense D435x [74]
CSI-2 USB 3.0 USB 3.0 USB 3.0 USB 3.0
Stereo Res. / FPS No 480p @ 120 fps 1MP @ 120 fps 1MP @ 120 fps FHD @ 90 fps
RGB Res. / FPS 12MP @ 30 fps 4K @ 30 fps 4K @ 30 fps 4K @ 30 fps FHD @ 90 fps
AI Acc.
Power (W)
Cost ($)
No 1.4 TOPS 1.4 TOPS 1.4 TOPS No
N/A 5–7.5 5–7.5 5–7.5 0.04–2.6
30 149 299 399 314–354
Table 8. Comparison of camera modules compatible with the Raspberry Pi 5, including interface type, stereo and RGB resolution with maximum frame rates, onboard AI acceleration, power consumption, and approximate cost.
Local Storage and Touch Display The SPU requires sufficient local storage to ensure reliable operation and to buffer data during classroom sessions. A 64 GB Class 10 SD card is used, allocating approximately 10 GB for the operating system and software stack, with the remaining
38
56 GB reserved for temporary backup storage. This capacity allows the system to record OAK-D camera streams for at least two hours—well above the typical onehour lecture duration—providing a safety margin in the event of network congestion or bandwidth limitations. Local buffering ensures that no engagement-relevant data is lost during transmission delays, and recordings can be automatically uploaded or deleted depending on the BSN’s mode of operation. For the touch interface, the SPU uses the Raspberry Pi Touch Display 2 [75], which provides a 24-bit RGB panel with a native resolution of 720 × 1280 pixels. The display consumes approximately 2.5 W at maximum brightness, making it suitable for battery-powered portable operation. It is used primarily for consent collection, system status display, and basic user interaction, and automatically turns off during sessions to minimize distraction and reduce power consumption. Battery Management System Given the estimated peak power consumption of the SPU, approximately 12 W for the Raspberry Pi 5, up to 7.5 W for the OAK-D Lite camera, and around 2.5 W for the touch display, the total worst-case load is about 22 W. Moreover, the Raspberry Pi 5 power specification requires that its supply be capable of delivering 5 V at 5 A, which places additional constraints on the power subsystem. The battery management system must therefore satisfy two primary requirements: (i) provide a stable 5 V output with enough current headroom (up to 5 A) to reliably power the Pi 5 and its peripherals, and (ii) supply sufficient battery capacity to support at least one full class session without interruption. To meet these requirements, we employ the Geekworm X1200 UPS HAT [76], which is specifically designed for the Raspberry Pi 5 and supports two 18650 lithiumion cells while providing a regulated 5.1 V output at up to 5 A continuous current. The X1200 interfaces with the Pi through pogo-pin contacts, eliminating the need for bulky connectors and reducing mechanical wear during assembly. The module also
39
integrates over-current, over-discharge, and over-charge protection circuitry, making it a robust and reliable power solution for the SPU. Furthermore, the X1200 offers uninterrupted power capability and includes a fuel-gauge interface accessible over I2 C to enable safe shutdown signaling, which is essential for preventing filesystem corruption and ensuring reliable SPU operation in classroom deployments. For this design, two 3500 mAh cells are used, yielding a nominal energy capacity of Ebat ≈ 2 × 3.7V × 3.5Ah ≈ 25.9Wh. Assuming a conservative 90% power-conversion efficiency, the usable energy delivered at the 5 V rail is approximately 0.9 × 25.9 ≈ 23 Wh. Under a continuous worst-case load of about 22 W, which corresponds to slightly more than one hour of operation. In practice, however, the average power draw is considerably lower: the display is blanked during most of the session, the OAK-D Lite typically consumes around 5 W, and the Raspberry Pi 5 rarely operates at its peak power, drawing less than 5 W under moderate workloads. This brings the total steady-state consumption to under 12 W, which allows the SPU to operate for at least two hours on the available battery capacity. Figure 6 presents the complete electronic system schematic of the SPU, illustrating the interconnections among all selected components, including the Raspberry Pi 5, OAK-D Lite camera, battery management subsystem (X1200), dual 18650 battery pack, local storage, and touch display.
This Figure 6.
System-level schematic dia-
schematic consolidates the hardware degram of the SPU, showing the electronic cisions discussed in the previous subseccomponents and their interconnections. tions and serves as the blueprint for the
40
SPU’s electrical integration and subsequent PCB and enclosure design. SPU Enclosure1 A custom enclosure was designed and fabricated to securely house all SPU electronic components while maintaining a compact and functional form factor. The enclosure was manufactured using the University of Louisville’s 3D-printing facilities. Its design satisfies several critical requirements: (i) it exposes only the USB-C charging port and the power button, ensuring that all unused Raspberry Pi 5 ports remain physically inaccessible for enhanced hardware security; (ii) it provides an internal mounting structure that enables straightforward assembly of the SBC, camera, display, and battery system; and (iii) it incorporates a mechanical rotation-limiting feature for the touch display hinge, preventing the SPU from tipping forward when placed on a desk. The final assembled enclosure has overall external dimensions of approximately 19×16.3×5.7 cm, making it compact enough for unobtrusive placement on a student’s desk. Detailed mechanical drawings—including all dimensions and subcomponent placements—are provided in Figure 7. Table 9 lists all components referenced in the mechanical design. A fully assembled SPU prototype was fabricated using the enclosure, electronic components, and mechanical integration described in the previous subsections. The completed unit demonstrates the final physical form factor, port layout, display mounting mechanism, and overall ergonomics of the device. Figure 8 illustrates the finished SPU assembly as deployed in classroom environments. Having established the complete hardware architecture of the SPU and detailed the design and fabrication of its enclosure, we now turn to the software stack that enables the device to operate as an intelligent sensing node within the BSN. While the hardware provides the necessary physical platform for computation, sensing, and power management, it is the software that governs session control, secure communi1
The enclosure was designed by a contracted freelance mechanical designer based on system specifications provided by the author.
41
Figure 7. Mechanical drawings of the SPU enclosure showing key dimensions, mounting features, and component placement. All dimensions are given in millimeters. Refer to Table 9 for the designated parts corresponding to the identifiers shown in the upper-right drawing.
Figure 8. Fabricated SPU unit showing the assembled enclosure, integrated touch display, camera module, and internal power system. The image illustrates the final physical form factor and layout of the complete sensing node as deployed in classroom environments. cation, on-device inference, and overall system reliability. The next section, therefore, outlines the SPU software architecture and the key components required to meet the
42
Item No.
Description
Qty
Unit Price ($)
1
Raspberry Pi Touch Display 2
1
60
2
OAK-D Lite Camera Module
1
149
3
Raspberry Pi 5 SBC (2.4 GHz, Quad-Core, 4 GB RAM)
1
64.26
4
Geekworm X1200 UPS HAT (5 V / 5 A)
1
43
5
18650 Lithium-Ion Battery Pack (Set of 2)
1
13
6
Raspberry Pi 5 Active Cooler
1
5
7
Right-Angle USB-C Power Cable
1
5
9
SPU Screen Enclosure (3D Printed)
1
7
10
SPU Screen Bracket (3D Printed)
1
3
11
Aluminum Male–Female Threaded Hex Standoff, M2.5
4
0.16
12
Stainless Steel Phillips Pan Head Screw, M2.5, 12mm
4
0.05
13
Stainless Steel Phillips Flat Head Screw, M2.5, 10mm
4
0.05
14
Stainless Steel Phillips Flat Head Screw, M4
2
0.09
15
Multipurpose 304 Stainless Steel Flat Bar
2
0.83
16
M5 Binding Barrel and Chicago Screw (18-8 Stainless Steel)
2
0.32
17
Class 10 MicroSD Card (64 GB)
1
10.25
Table 9. Bill of Materials (BOM) for the SPU functional, security, and real-time processing requirements defined in the previous chapter.
IV.2
SPU Software
The SPU software architecture is composed of five tightly integrated subsystems that together enable secure, autonomous, and real-time operation in classroom environments. As illustrated in Figure 9, these subsystems manage Wi-Fi onboarding, cryptographic device registration, session execution and consent handling, secure data upload, and continuous battery health monitoring. Although not shown explicitly in the flowchart, the SPU GUI also includes a persistent status bar that displays the system clock, battery level, Wi-Fi state, server connectivity, certificate validity, and quick-access controls for settings and power management. Each of the five subsystems is described in detail in the following subsections.
43
44
Figure 9. Software architecture flowchart of the SPU, illustrating the five primary subsystems: network registration, device registration, session management, upload logic, and battery monitoring. This diagram summarizes the complete operational workflow executed by the SPU from boot to shutdown.
Before doing so, we briefly outline the underlying software technologies that support their implementation and ensure consistent, reliable operation across all SPU devices. Software Technology Stack The SPU software is implemented using a modern and lightweight technology stack designed for robustness, portability, and efficient execution on embedded hardware. Python[77] serves as the primary development language due to its readability and rich ecosystem of CV and ML libraries. The user interface is built using PySide6 [78], providing a responsive, touch-friendly GUI suitable for classroom deployment. Camera communication and on-device inference are handled through the DepthAI framework[79], which interfaces seamlessly with the OAK-D Lite to access RGB streams, stereo depth, and neural-network outputs. Secure uploads to the server’s S3-compatible object storage are performed using the boto3 library[80], ensuring encrypted and reliable data transfer. To support maintainable and reproducible development, the project employs a modern continuous integration and continuous delivery (CI/CD) toolchain, including Git with GitHub Actions for automated builds, Docker for environment isolation, pytest[81] for unit and integration testing, and ruff[82] for linting and code-quality enforcement. Python environments are managed using miniforge [83], which enables consistent dependency resolution across development machines and ARM-based deployment targets such as the Raspberry Pi 5. Together, this technology stack provides a stable, scalable, and developer-friendly foundation for implementing, validating, and deploying the SPU software. Network Registration Logic The SPU is designed to operate primarily on the University of Louisville’s ULSponsor wireless network. Relying on this campus-wide onboarding network provides several advantages: (i) all SPU traffic is automatically protected by the university’s Virtual
45
Private Network and firewall infrastructure; (ii) the SPU can be deployed in any classroom on campus without additional configuration; and (iii) no private access points or custom network setups are required. Although the SPU is technically capable of connecting to any Wi-Fi network, this capability is reserved for administrative debugging and testing and is not exposed through the SPU GUI. Manual network changes are performed only by an authorized system administrator. Upon startup, the SPU begins by checking for Internet connectivity. If connectivity is unavailable, it determines whether it is currently associated with the ULSponsor network. When connected to this onboarding SSID, the SPU displays a registration interface that allows an authorized user to authenticate the device and request Wi-Fi access for its MAC address. This process must be performed only once per SPU— as long as the device retains the same MAC address, it remains authorized on the university network. To support this workflow, the system administrator must first create a guest account, which is then used to onboard any new SPU device. If registration is successful, the SPU reloads its network configuration and repeats the connectivity check. If registration fails, the device continues prompting for onboarding until Internet access is available. This subsystem ensures that each SPU can securely join the campus network without requiring per-classroom setup or manual configuration, enabling fully autonomous operation across all deployment locations. Device Registration Logic Before an SPU can participate in classroom deployments, it must be cryptographically authenticated and registered with the BSN server. As illustrated in Figure 10, each SPU is provisioned at manufacture time with the server’s public certificate, ensuring that all subsequent communication occurs only with the legitimate BSN server and preventing man-in-the-middle and spoofing attacks. Upon startup, the SPU checks whether a valid device certificate is already present. If so, the device bypasses registration and proceeds directly to the session-waiting state. If no valid
46
Figure 10. Device registration flow for the SPU. The diagram illustrates the certificate-based onboarding process, including network verification, PSK-based authorization, CSR generation, certificate issuance via STEP CA, and the return of the server-generated public key used for secure local data encryption. certificate exists, the SPU displays the device-registration interface, which prompts the user to enter a time-limited pre-shared key (PSK). This PSK is generated by the system administrator for a specific classroom, is valid for only ten minutes, may be used exactly once, and ensures that each SPU is bound to a single classroom at any given time. Once the PSK is entered, the SPU locally generates a private key and a certificate signing request (CSR). The CSR, along with the PSK and the device identifier, is sent to the BSN server. The server verifies the PSK and, if valid, signs the CSR using the on-premises STEP CA, thereby issuing a device certificate. In addition, the server generates a dedicated asymmetric key pair for secure data protection: the private key is stored exclusively on the server and never leaves it, while the corresponding public key is returned to the SPU. This public key is later used by the SPU to encrypt buffered engagement data stored locally before upload, ensuring that even if an SPU is stolen or compromised, the data cannot be decrypted—only the BSN server, which holds the matching private key, can recover it. After receiving the signed certificate and the server-side public key, the SPU
47
Figure 11. Functional block diagram of the SPU recording-session logic, showing the sequential workflow for session control (start/stop based on server signals or user interaction), video capture and encoding, hybrid encryption, local buffering and uploading, and session termination. stores both credentials securely and transitions into normal operational mode. If PSK verification fails, the device remains on the registration screen until valid credentials are provided. This certificate-based onboarding procedure needs to be performed only once—typically by an administrator or a designated technical staff member—before the SPU is deployed in a classroom. Once completed, the SPU is fully authenticated, cryptographically trusted, and able to store sensitive data securely, enabling it to integrate seamlessly into the classroom sensing workflow. Running Session Logic Once device registration is completed, the SPU transitions into a waiting state in which it periodically polls the BSN server for the activation of a classroom session. The system supports two distinct session types: a record session used during the development phase to collect raw video for dataset construction, and an analysis session used during deployment for real-time engagement inference. When the server signals the start of a session, the SPU presents a consent interface to the student. In compliance with IRB requirements, the camera pipeline remains fully disabled until the student explicitly grants consent. Only after consent is obtained does the SPU enter the active sensing workflow appropriate for the selected session type.
48
For a recording session, the SPU captures video frames from the OAK-D Lite at 25 fps, using a 720p resolution stream from the RGB camera and a 480p stream from the mono stereo pair. Each frame is first encoded on the camera’s hardware accelerator using the H.264/H.265 codec and then encrypted before being written to the local buffer. To minimize computational overhead, the SPU employs a hybrid encryption scheme in which each encoded frame is encrypted using a randomly generated SHA-256–based symmetric key [84], while the symmetric key itself is encrypted using the server’s asymmetric public key. The encrypted frame, encrypted symmetric key, and associated metadata are then written to the local buffer, where they remain only temporarily: as soon as network conditions allow, the SPU uploads each buffered frame to the server’s S3-compatible object storage [85] and deletes the local copy immediately upon confirmed receipt. Each record includes session identifiers, timestamps, camera intrinsics, frame dimensions, and device identifiers, following the structure: { "session_id": ..., "video_id": ..., "frame_counter": ..., "metadata": ..., "encrypted_key": ..., "nonce": ..., "tag": ..., "device_id": ..., "timestamp": ..., "device_timestamp": ..., "timestamp_offset": ..., "width": ..., "height": ..., 49
"type": ... } For an analysis session, the SPU processes video frames on the edge to compute engagement metrics in real time. Frames are grouped into 10-second windows, and for each window, the SPU performs face detection, gaze estimation, and behavioral feature extraction to derive an engagement score. Both the encrypted frames and the resulting engagement metrics are buffered in local storage. The SPU continues this process until the session ends or the student manually stops recording. All buffered data is subsequently transmitted to the server by the upload subsystem, which handles secure delivery and confirmation-based deletion. The overall workflow executed during a recording session is summarized in the functional block diagram shown in Figure 11, which illustrates the control flow governing session activation, consent handling, frame acquisition, encryption, buffering, and termination conditions. Secure Upload Logic The upload subsystem operates continuously in parallel with all other SPU processes once the device possesses a valid certificate. This ensures that any previously unuploaded packets—including those generated during temporary network outages—are eventually delivered reliably. The overall workflow is illustrated in Figure 12, which summarizes the interaction between the SPU, the BSN server, and the on-premises MinIO storage service. At regular intervals, the SPU requests from the BSN server a time-limited, uploadonly pre-signed link to the MinIO object-storage backend [86]. The server verifies the SPU’s certificate and, if valid, issues a pre-signed URL that remains active for 30 minutes. The SPU uses this URL to upload any buffered packets during that window; once the link expires, the SPU automatically requests a new one to maintain continuous authorization without exposing long-lived credentials. 50
Figure 12. Functional block diagram illustrating the interaction between the SPU and the BSN backend server. The BSN server stack consists of an NGINX reverse proxy, a Django application server, an on-premises MinIO object-storage service, and an NTP time synchronization service. The diagram shows how the SPU requests timelimited, upload-only pre-signed URLs and transmits encrypted data to the server over mutually authenticated TLS 1.3 connections. All uploads are performed over a mutually authenticated TLS 1.3 [87] connection. Upon successful receipt of each packet (or batch), the server sends an acknowledgment, after which the SPU deletes the local copy to prevent redundant retransmissions and reduce storage usage. If a packet fails to upload, it remains in the buffer until connectivity is restored or a new upload window becomes available. To avoid overloading the server with large numbers of individual upload requests—especially during recording sessions where each video frame is encrypted independently—the SPU aggregates encrypted frames into ZIP archives of fifty frames each. Every ZIP bundle contains encrypted frames, associated metadata, and the corresponding encrypted symmetric keys. This batching strategy significantly improves throughput, reduces request overhead, and preserves confidentiality and integrity guarantees throughout the upload process.
51
Battery Monitor Logic The SPU continuously monitors its battery voltage and charge percentage through the UPS module’s fuel-gauge interface. When the battery voltage falls below 3 V, the SPU issues a “battery-critical” signal to the session-control subsystem, allowing several seconds for pending uploads to complete and local buffers to flush before initiating a controlled shutdown. This mechanism prevents abrupt power loss, protects the battery from over-discharge, mitigates the risk of filesystem corruption, and ensures that no partially written data remains in an inconsistent state. The battery-monitor logic operates independently of both the recording and upload loops, providing continuous, power-aware supervision throughout the entire session. SPU User Interface Workflow Figure 13 illustrates the complete user-interface flow of the SPU application, shown as a sequence of six screens arranged in a 2×3 layout. Screen 1 displays the loading and splash interface that appears immediately after the SPU boots. If the device lacks Internet connectivity and is associated with the“ULSponsor” onboarding network, Screen 2 is presented, prompting the user to enter valid UL credentials to obtain Wi-Fi access. Once network connectivity is established, the SPU checks whether a valid device certificate is installed. If no certificate is found, Screen 3—the deviceregistration interface—is shown, where the user must enter a time-limited pre-shared key (PSK) to authorize provisioning. Upon successful certificate issuance, the SPU transitions to Screen 4, the sessionwaiting state, where it polls the BSN server for an active classroom session. When a session begins, the SPU displays Screen 5, the recording interface with the camera disabled until student consent is explicitly granted. After the student provides consent, the camera pipeline is activated and the SPU enters Screen 6, the full recording interface used for either dataset collection or engagement-analysis sessions. This sequence ensures secure, ethical, and controlled activation of the SPU in accordance 52
53
Figure 13. User-interface flow of the SPU software, shown as six sequential screens: (1) splash/loading screen; (2) ULSponsor Wi-Fi registration page shown when the device has no Internet access; (3) device-registration screen for PSK-based certificate provisioning; (4) session-waiting interface displayed after successful onboarding; (5) recording interface shown at session start with the camera disabled pending student consent; and (6) active recording interface displayed once consent is granted.
with IRB and system requirements. SPU Operating System Security To ensure that each SPU operates as a secure, tamper-resistant sensing node, the device runs on a hardened Linux-based operating system that has been explicitly configured to minimize its attack surface. All nonessential services and interfaces are disabled, including SSH access, physical TTY terminals, and the standard desktop environment. As a result, even if an attacker connects an external keyboard or mouse to the SPU, the device cannot be interacted with or placed into a standard login shell. Instead of relying on conventional user login workflows, the SPU launches directly into its application runtime under a restricted system user. The graphical interface is served through a minimal X11 environment without any window manager, shell access, or desktop session, thereby preventing privilege escalation through GUI-based interactions. The filesystem is locked down to prevent modification of system binaries, and all software updates are delivered manually. Together, these measures create a tightly controlled execution environment in which the SPU behaves exclusively as an appliance: it performs its sensing, inference, and upload tasks autonomously, while offering no general-purpose computing capabilities to a physical attacker. This hardened OS configuration is essential for maintaining data privacy, preventing unauthorized access, and ensuring the integrity of the BSN deployment across classroom environments.
IV.3
Server Software
Django was selected as the primary backend framework for the BSN because it provides a mature, secure, and highly extensible web-application foundation while allowing the entire backend to be implemented in Python. This minimizes the number of programming languages required across the BSN ecosystem, simplifies development and maintenance, and enables seamless integration with the Python-based SPU soft54
ware stack. Figure 12 presents a high-level block diagram of the BSN server architecture, illustrating how instructors interact with the system through a web client and how SPUs communicate with the backend during data-upload sessions. The backend of the BSN system is implemented as a Django-based web server[88] with a MySQL[89] relational database and a modular, microservice-oriented architecture. Each functional domain of the system—user authentication, classroom configuration, device provisioning, session coordination, and secure data upload—is encapsulated in a dedicated Django app. The server is deployed using Gunicorn as the WSGI application server [90] and NGINX[91] as a reverse proxy, providing scalability, security, and production-grade performance. The backend runs on an Ubuntu 20.04[92] server installation and integrates with two critical on-premises services: (i) a STEP CA[93] instance responsible for issuing and revoking SPU certificates, and (ii) a MinIO S3-compatible object-storage service[86] that manages encrypted data ingestion. All server functionality is exposed through a versioned REST API [94] under the /api/v1/ namespace. The following subsections describe each major backend component and its role within the overall system architecture. Accounts App The accounts app provides user authentication, authorization, and identity management for instructors, researchers, and administrators. Built on top of dj-rest-auth, it exposes endpoints for login, logout, password reset, email verification, and user profile retrieval. Role-Based Access Control (RBAC) ensures that only authorized users can initiate sessions, configure classrooms, or access sensitive data. The primary API endpoints provided by this app are listed below: /api/v1/dj-rest-auth/ /api/v1/dj-rest-auth/user/
55
Classroom Manager App The classroom manager app maintains metadata describing classrooms, their assigned SPUs, and operational configurations. Administrators may retrieve classroom records, update settings, and bind SPUs to specific physical environments. This mapping is essential for IRB compliance and for contextualizing engagement data. The primary API endpoints provided by this app are listed below: /api/v1/classroom/ /api/v1/classroom/<pk>/ Device Authentication App The device auth app implements the certificate-based onboarding workflow used to authenticate each SPU before it enters operational mode. Its responsibilities include: (i) validating time-limited pre-shared keys (PSKs); (ii) receiving Certificate Signing Requests (CSRs); (iii) invoking STEP CA to issue signed device certificates; and (iv) managing health-check telemetry. The primary API endpoints provided by this app are listed below: /api/v1/device/provision/ /api/v1/device/health/ /api/v1/psk/ /api/v1/psk/<pk>/regenerate/ This subsystem ensures that every SPU is cryptographically trusted and securely provisioned prior to participating in the sensing workflow. Session Manager App The session manager app orchestrates the lifecycle of classroom sessions, including start and stop events, session-type configuration (recording or analysis), and realtime polling by SPUs. When a session begins, the server provides the SPU with the 56
necessary session metadata; when it ends, the server signals termination so that the device can finalize encoding, flush buffers, and stop sensing. The app also ensures that only one active session may exist for a given classroom at any time, preventing conflicting operations. Furthermore, sessions are automatically deactivated once their predefined duration has elapsed, ensuring consistent system behavior even if the instructor forgets to stop the session manually. The primary API endpoints provided by this app are listed below: /api/v1/session/start/ /api/v1/session/stop/ /api/v1/spu/session/current/ Storage App The storage app handles integration with the on-premises MinIO server used for encrypted data ingestion. It generates time-limited, upload-only pre-signed URLs that SPUs use to transmit encrypted bundles of data. Before a URL is issued, the server verifies the device’s certificate to ensure proper authentication. The primary API endpoints provided by this app are listed below: /api/v1/spu/storage/presigned-directory/ This approach enforces a zero-trust model: SPUs can only upload, never read or delete objects, and their upload permissions expire automatically (typically in 30 minutes). API Security and Transport Layer Protection All server endpoints are protected using TLS 1.3 to ensure the confidentiality and integrity of all communications between SPUs, user clients, and the backend services. In addition to transport-layer encryption, virtually all SPU-facing APIs require the device to present a valid, server-signed client certificate as part of a mutual TLS 57
(mTLS) handshake. This certificate-based authentication prevents unauthorized or spoofed devices from interacting with the system, blocks man-in-the-middle attacks, and enforces a strict zero-trust communication model. As a result, only properly provisioned SPUs are able to register, join sessions, upload encrypted data, or retrieve operational configuration from the BSN server. Clock Synchronization To ensure temporal consistency across all SPUs during classroom operation, the backend server also functions as the authoritative time source for the BSN. The server hosts a local NTP service[95] to maintain a stable and accurate system clock and exposes this reference to all connected SPUs. During normal operation, each SPU periodically synchronizes its local clock with the server’s timestamp, achieving a typical alignment accuracy within 200–500 ms. This level of precision is sufficient for aligning engagement metrics, session boundaries, and video-frame timestamps across devices, and is essential for both dataset construction and real-time classroom analytics. By centralizing time synchronization through an on-premises NTP server citeRFC5905, the BSN maintains coherent temporal semantics even in the presence of network jitter or device-specific clock drift. Calibration One advantage of using Oak camera is it comes with its cameras are calibrated both for stereo and RGB. This eliminates the need to calibrate each SPU manually, which is a tedious process. For the FCUs we developed a GUI that can be used to help make calibration of these cameras easy. This GUI supports adding new classrooms and cameras, performing calibration, and visualizing the results. Moreover this GUI also supports calibrating objects of interest, such as a blackboard or a projector screen. The calibration approach is simple: a fixed set of fiducial markers, e.g, ARUCO markers [96], placed in a classroom such that each pair of cameras sees at least
58
Figure 14. Three windows of the developed front-end GUI for classroom calibration. The top-left window is the login page. The bottom window is the calibration page, which supports adding new classrooms and cameras, performing calibration, and visualizing the results. The image in the center of the bottom window shows a live camera stream that has been rectified to remove distortion. The top-right window presents a 3D visualization of a calibrated classroom at the CVIP Lab. Although the classroom contains a single actual screen, additional dummy surfaces (e.g., desk surfaces) were added as objects of interest to test the calibration algorithm.
59
one common marker. A pose graph is then created from the cameras and markers and optimized to get each camera pose relative to a common reference. A similar approach can also be used to mark the fixed objects of interest in the classroom, then a pose graph is created and optimized to obtain the pose of each object relative to a common reference in 3D. Each object of interest is modeled as a rectangular surface with two ARUCO markers placed at opposite corners, and it must be fully visible in at least one camera view. Our current system supports this feature; more specifically, we implemented the approach proposed by [97] to calibrate our classroom. Figure 14 shows three windows from the front-end application used for classroom calibration and visualization of the results. To assess the performance of the calibration algorithm, we use the mean re-projection error, which quantifies the average discrepancy between the projected 3D points and their detected 2D locations in the camera images.
60
CHAPTER V BSN EVALUATION
This chapter presents a comprehensive evaluation of the proposed SPU and BSN architecture. The assessment serves two purposes: first, to compare the performance, security, and functional capabilities of the new SPU against the baseline system previously developed in the CVIP Lab and reported in [42]; and second, to demonstrate through a series of empirical tests that the proposed system meets the operational requirements defined in earlier chapters. The evaluation examines several critical dimensions of system performance, including battery-powered runtime (with a target of exceeding two hours of continuous operation), sustained network bandwidth for encrypted video-data transfer, clock synchronization accuracy between SPUs and the BSN server, and the SPU’s ability to execute moderate real-time CV/ML workloads such as face detection, gaze estimation, and emotion classification models. Together, these results validate that the proposed SPU achieves the intended design objectives and constitutes a substantial advancement over the earlier baseline platform. V.1
Baseline Comparison
To contextualize these results, it is necessary to evaluate the proposed BSN against the baseline SPU described in [42], which represents the previous system revision developed in the CVIP Lab. The new SPU introduces major enhancements across both hardware and software. On the hardware side, the design upgrades four key subsystems: the camera (from a basic Pi camera to a 4K DepthAI module with on-board AI 61
Feature
Baseline [42]
Proposed SPU
SBC
Raspberry Pi 4
Raspberry Pi 5
CPU
Quad-core Cortex-A72 @ 1.5 GHz
Quad-core Cortex-A76 @ 2.4 GHz
RAM
2 up to 8 GB LPDDR4
4up to 16 GB LPDDR4X
Storage
32 GB microSD card
64 GB microSD card
Display
Raspberry Pi Touch Display 1 (800×480)
Raspberry Pi Touch Display 2 (720×1280)
Camera
Official Pi Camera Module (no depth, no accelerator)
DepthAI OAK-D Lite 4K RGB, stereo depth 4 TOPS AI accelerator
Enclosure
Off-the-shelf enclosure; ports exposed; no physical tamper protection
Custom 3D-printed enclosure; sealed ports; physical security
In Power
Wired charger; no UPS
7 Ah Battery up to 2 hours up time; safe shutdown support
Touch
10-point capacitive touch
5-point capacitive touch
AI Acc.
None
OAK-D Lite (4 TOPS)
Boot Time
50 Sec
28 Sec
Wi-Fi
Wi-Fi 5
Wi-Fi 6 higher throughput & stability
Size
14 × 22 × 4 mm
16.5 × 19 × 5.5 mm
Cost
$200
$383
Table 10. Comparison between the baseline [42] and the proposed BSN SPU hardware. The proposed system provides significant improvements in compute capability, wireless connectivity, camera quality, physical security, autonomy, and AI processing capacity with a moderate cost increase. acceleration), the power architecture (from wired-only to battery-powered operation), the compute capability (from a Pi 4 to a Pi 5 with an auxiliary neural accelerator), and the enclosure (from off-the-shelf housings with exposed ports to a custom, physically secured enclosure). On the software side, significant advances include a comprehensive data-security and privacy framework, certificate-based authentication, hardened OS configuration, encrypted local storage, and a fully automated sensing and up-
62
Feature
Baseline [42]
Proposed SPU
Operating System
Pi OS with default services enabled
Hardened Pi OS; disabled SSH, TTYs, desktop login;
GUI Framework
Tkinter running inside a full desktop environment
PySide6 (Qt 6) in minimal X11 session; no desktop, no shell; kiosk-secure runtime
Startup Behavior
Requires auto-login to launch application
Direct boot into SPU app; no login required; locked-down execution environment
Security Model
TLS 1.3 only; device certificates (Manual); no encrypted storage
Mutual TLS 1.3, device certificates (STEP CA), encrypted local storage (hybrid AES+RSA)
Data Handling
Raw video stored unencrypted; manual transfer
All buffered data AESencrypted per frame; S3 upload-only URLs; server-side decryption only
Data Upload
No object storage
Automated upload via MinIO; confirmation-based deletion
Inference
Yes
Supports it, Not implemented
Clock Sync.
No synchronized timestamps
NTP-based server–SPU synchronization (200–500 ms)
Session Management
Automatic start/stop
Server-driven sessions; dual modes (recording/analysis); automatic transitions
Network Reg.
Manually connects to Wi-Fi; insecure configuration possible
Automatic via ULSponsor
Device Reg.
Manual
Automatic; PSK-based provisioning.
Table 11. Comparison of software features between the baseline [42] and the proposed BSN SPU software. The proposed software stack provides major improvements in security and data privacy to support IRB-compliant classroom operation. load pipeline. These combined improvements enable a more secure, autonomous, and scalable BSN suitable for real-world classroom deployment, motivating a structured comparison against the earlier baseline.
63
Experiment 1: Server Load and Network Throughput Evaluation This experiment evaluates whether the backend server and network infrastructure can sustain the expected data traffic generated by multiple SPUs during a live classroom session. A test classroom was created and populated with ten registered SPUs. Because a typical lecture lasts approximately 60 minutes, a 70-minute test session was created and activated through the admin portal to provide additional buffer for initialization and shutdown overhead. Theoretical Bandwidth Requirement Each SPU generates three packets per second, with an average packet size of approximately 1.5 MB. Thus, the per-device data rate is
3 × 1.5 MB/s = 4.5 MB/s, which corresponds to
4.5 MB/s × 8 = 36 Mb/s. For a deployment of ten SPUs operating simultaneously, the aggregate upload bandwidth requirement is therefore
10 × 4.5 MB/s = 45 MB/s ≈ 360 Mb/s. This value represents the theoretical worst-case throughput that the server and network infrastructure must sustain during a live classroom session under continuous data generation and upload. On each SPU, a synthetic data generator was deployed to emulate real device output, matching the frame rate, packet size, and metadata structure of actual encrypted SPU payloads. Because the upload subsystem runs continuously, all SPUs
64
immediately began sending synthetic packets through the complete secure upload pipeline using time-limited pre-signed URLs. During the 70-minute session, the server was unable to ingest packets in strict real time at the full theoretical rate; however, the SPUs successfully buffered all generated packets locally without data loss. After the session ended, the SPUs continued uploading their buffered packets, completing the process approximately one hour later. This demonstrates that the system can tolerate transient bandwidth saturation by leveraging local buffering, eventual consistency, and the periodic refresh of upload credentials. To verify correctness, the number of packets generated and uploaded by each SPU was recorded. Local storage usage was measured before and after the session to confirm that all buffered packets were deleted after successful upload. Table 12 summarizes the results. This experiment confirms that, even under heavy multi-device load, the BSN backend can safely and reliably ingest all data without loss by combining edge buffering with secure upload retries. SPU ID SPU–01 SPU–02 SPU–03 SPU–04 SPU–05 SPU–06 SPU–07 SPU–08 SPU–09 SPU–10
Packets Generated 4200 4200 4200 4200 4200 4200 4200 4200 4200 4200
Packets Uploaded 4200 4200 4200 4200 4200 4200 4200 4200 4200 4200
Packet Loss 0 0 0 0 0 0 0 0 0 0
Storage Change (MB) 0 0 0 0 0 0 0 0 0 0
Table 12. Summary of server load and network-throughput experiment for ten SPUs. All devices successfully transmitted all generated packets with zero loss. Uploads were not real-time but completed within approximately one hour after the session ended, confirming correct buffering, retry logic, and cleanup of local temporary storage.
65
Experiment 2: Clock Synchronization Accuracy This experiment evaluates whether all SPUs maintain consistent time alignment with the BSN server, a requirement for synchronizing video frames, engagement metrics, and session boundaries. ten SPUs were powered on, and the clock offset between each device and the server was measured once every ten seconds for a duration of ten minutes via SSH. Table 13 summarizes the results, reporting the number of samples per device, along with the average, minimum, and maximum of the measured offsets. To further validate synchronization visually, the clock widget was remotely launched on ten SPUs, arranged in a grid layout on a laptop display, and recorded over a tenminute period. Representative frames were extracted and inspected to verify that all SPUs’ displayed times advanced in near-perfect unison. Figure 15 illustrates the captured grid view. Although the timestamps displayed across SPUs are nearly identical, minor visual differences of up to approximately 100 ms can occasionally be observed between widgets in Figure 15. These discrepancies do not reflect actual clock misalignment; rather, they arise from the rendering and transmission overhead introduced by the X11 forwarding pipeline used to display all clock widgets remotely on a single laptop screen. The underlying NTP-synchronized device clocks remain tightly aligned, as confirmed by the quantitative jitter measurements reported earlier. Across all devices, the measured offsets fall within a narrow range (mostly within ±3 ms), with negligible jitter for the majority of SPUs. These results confirm that the NTP-based synchronization mechanism achieves the temporal precision required for classroom analytics and multi-device dataset alignment. Experiment 3: Battery Endurance Under Continuous Operation To verify that the SPU can operate on battery power for the full duration of a typical lecture, battery-discharge tests were conducted under realistic sensing conditions. In the first trial, an SPU was fully discharged, then recharged completely while recording 66
67
Figure 15. Visual synchronization check of ten SPUs during the clock-alignment experiment. The horizontal axis corresponds to SPU IDs, while the vertical axis represents time samples. Each cell shows the timestamp displayed by an SPU at a given moment. The near-identical timestamps across columns indicate that all SPUs remain tightly synchronized with the server throughout the ten-minute observation period.
Device ID 1 2 3 4 5 6 7 8 9 10
Samples 131 131 131 131 131 131 131 131 131 131
Avg Offset (ms) -1.2096 -0.8002 -0.0242 -1.4537 -0.0757 1.3071 0.9944 0.4537 0.5046 -3.0781
Min Offset (ms) -2.4546 -0.8002 -0.0242 -1.4537 -0.0757 1.3071 -0.6923 0.4537 0.5046 -3.0781
Max Offset (ms) 0.7432 -0.8002 -0.0242 -1.4537 -0.0757 1.3071 3.2164 0.4537 0.5046 -3.0781
Table 13. Clock synchronization offsets between each SPU and the BSN server measured over a ten-minute interval. The offsets show tight clustering around zero ms, confirming stable NTP-based synchronization across all devices. the total charging time. The full charging process required approximately two hours. A 70-minute session was then initiated, during which the device continuously recorded all three OAK-D Lite camera streams. Battery percentage before and after the session was logged. A second trial repeated the test, this time starting from a 75% charge level. Table 14 reports the session duration, initial and final battery percentages, and the effective battery consumption for both trials. These results confirm that the SPU can sustain more than two hours of continuous sensing, satisfying the system’s operational requirements. Table 14. Summary of battery-endurance trials under continuous SPU operation. Trial
Session Duration (min)
Initial Battery (%)
Final Battery (%)
Consumed (%)
Trial 1 (Full charge)
70
100
56
54
Trial 2 (Partial charge)
70
75
18
58
Experiment 4: On-Device AI Compute Performance This experiment evaluates whether the SPU’s integrated OAK-D Lite accelerator can support the moderate real-time CV/ML workloads required for classroom engagement analysis. Three representative neural network pipelines were selected from the DepthAI model zoo, covering emotion recognition, gaze estimation, and fatigue
68
detection. These tasks reflect the core sensing functions of the BSN. Each model was deployed on the SPU, and the achieved inference frame rate (FPS) was recorded. The results, summarized in Table 15, demonstrate that the OAK-D Lite provides sufficient compute throughput to meet the latency and responsiveness requirements of the proposed system. Emotion Recognition Pipeline The first evaluated model is the Emotion-Recognition-8-ENet network [98], based on the EfficientNet architecture [99]. This model classifies faces into eight emotion categories: anger, contempt, disgust, fear, happiness, neutral, sadness, and surprise. On the OAK-D Lite, the model operates as a two-stage DepthAI pipeline: (i) the YuNet face detector [100] locates faces and produces crops; (ii) the ENet classifier evaluates each crop. The full pipeline achieves approximately 11 FPS using 64×64 pixel face crops, enabling smooth multi-person classroom monitoring. A higher-resolution version of the model (256×256 crops) was also tested and achieved approximately 5 FPS. Gaze Estimation Model The second evaluated model is the DepthAI Gaze Estimation pipeline, which demonstrates multi-stage and multi-input inference on the OAK-D Lite. The processing sequence consists of: • Stage 1: Face Detection and Landmark Extraction using YuNet [100], which provides facial keypoints for cropping the face and eyes. • Stage 2: Head Pose Estimation using the OpenVINO head-pose model [101].
69
• Stage 3: Gaze Direction Estimation using the ADAS gaze-estimation model [102], combining cropped eye images with the estimated head-pose vector. The complete pipeline achieves real-time performance at approximately 10 FPS on the OAK-D Lite, making it suitable for estimating student visual attention during classroom instruction. Fatigue Detection Model The third evaluated model is the DepthAI Fatigue Detection pipeline, which analyzes facial keypoints to estimate fatigue indicators such as eye closure and forward-leaning posture. After detecting and cropping the face from the RGB frame, the MediaPipe Face Landmarker model [103] produces a dense set of 3D facial keypoints. These points are used to derive two key behavioral markers: • Eye Closure: Eye-aspect ratios are computed to detect prolonged eyelid closure. • Forward Leaning: Geometric relationships among facial keypoints are analyzed to infer characteristic posture associated with fatigue. The complete pipeline runs at approximately 10 FPS on the OAK-D Lite, enabling reliable real-time detection during classroom sessions.
70
Model / Pipeline
FPS
Notes
Emotion Recognition (64×64 crops)
11
Two-stage pipeline: YuNet + ENet classifier
Emotion Recognition (256×256 crops)
5
Higher-resolution variant for improved accuracy
Gaze Estimation Pipeline
10
Three-stage pipeline: face detection, head pose, gaze estimator
Fatigue Detection Pipeline
10
MediaPipe Face Landmarker + posture/eye analysis
Table 15. Inference performance of selected DepthAI pipelines running on the OAK-D Lite accelerator. Results indicate that the accelerator provides sufficient throughput for real-time classroom engagement sensing.
71
CHAPTER VI CONCLUSIONS
This thesis presented the design, implementation, and evaluation of a novel Biometric Sensor Network (BSN) to enable real-time measurement of individual student engagement in STEM classroom environments. The system integrates robust hardware, secure software, computer-vision pipelines, wireless networking, and backend infrastructure into a unified platform capable of operating unobtrusively, ethically, and autonomously during live instructional sessions. The Student Processing Unit (SPU), developed as the foundational sensing node in the BSN, was engineered to meet five key system requirements: non-intrusive operation, non-invasive sensing, non-stigmatizing deployment, real-time processing, and fully automated operation. The redesigned SPU introduced several advancements over the previous baseline platform, including a significantly more powerful Raspberry Pi 5 compute module, an OAK-D Lite camera providing 4K RGB imaging, stereo depth estimation, and a 4-TOPS AI accelerator, a battery-powered architecture that supports portable two-hour operation, and a custom 3D-printed enclosure offering enhanced physical security for classroom use. On the software side, the SPU incorporates a hardened operating system with disabled shells, restricted user access, encrypted data handling, certificate-based authentication, and an end-to-end secure data pipeline. A modular software architecture was developed to manage the full device lifecycle, including network registration, device provisioning, session coordination, secure data upload, and power-aware behavior. 72
Complementing the SPU, the backend server provides certificate management (via STEP CA), authenticated REST APIs, a session-control service, and encrypted data ingestion through an on-premises MinIO object-storage cluster. Combined, these components form a scalable, secure, and institution-compliant sensing infrastructure. A series of evaluation experiments demonstrated that the proposed BSN meets its design objectives. The network-throughput experiment showed that the backend can reliably handle concurrent uploads from multiple SPUs without packet loss. Clock synchronization measurements confirmed that SPUs maintain sub-second alignment with the server, sufficient for temporal consistency in engagement analytics. Batteryruntime tests validated that the SPU can sustain over two hours of continuous operation, satisfying the demands of a typical lecture. Finally, inference benchmarks verified that the OAK-D Lite accelerator supports real-time execution of key CV/ML workloads, including emotion recognition, gaze estimation, and fatigue detection. Overall, the contributions of this work establish a robust foundation for scalable, ethical, and real-time engagement sensing in university classrooms. The resulting BSN provides not only a practical tool for empirical research in student engagement but also a flexible platform for future extensions in learning analytics, classroom management, and adaptive instructional technology.
Future Work Although the system developed in this thesis is functional and demonstrates strong empirical performance, several avenues remain for future research and development: • Improved Engagement Models: Future versions may incorporate transformerbased vision models or multimodal fusion techniques to improve robustness and generalization. • Cross-Device Synchronization: Hardware-triggered synchronization or precision time protocol (PTP) could further reduce temporal drift between SPUs.
73
• Scalability Studies: Large-scale deployments involving hundreds of SPUs would provide insights into long-term performance, network load, and maintenance requirements. • Adaptive Battery Management: Predictive power modeling and dynamic frame-rate adjustment could extend battery life during long sessions. • Privacy-Preserving Analytics: Techniques such as federated learning or ondevice distillation could enable richer engagement inference while maintaining strict privacy guarantees. In summary, this work demonstrates that a secure, autonomous, and ethicallycompliant BSN is both technically feasible and practically deployable in real classrooms. The platform developed here opens the door to new research in real-time engagement analytics and has the potential to meaningfully transform how student learning is understood and supported at scale.
74
REFERENCES
[1]
A. P. Carnevale, A. R. Hanson, and B. Cheah, “The economic value of college majors,” (2015).
[2]
J. Rothwell, “The hidden stem economy,” ”” (2013).
[3]
G. J. Beach, “The u.s. technology skills gap: What every technology executive must know to save america’s future,” The U.S. Technology Skills Gap: What Every Technology Executive Must Know to Save America’s Future pp. 1–323 (2024).
[4]
S. Malcom and M. Feder, “Barriers and opportunities for 2-year and 4-year stem degrees: Systemic change to support students’ diverse pathways,” Barriers and Opportunities for 2-Year and 4-Year STEM Degrees: Systemic Change to Support Students’ Diverse Pathways pp. 1–214 (2016).
[5]
D. Kelly, H. Xie, C. W. Nord, F. Jenkins, J. Y. Chan, and D. Kastberg, “Performance of u.s. 15-year-old students in mathematics, science, and reading literacy in an international context. first look at pisa 2012. nces 2014-024.” National Center for Education Statistics (2013).
[6]
S. Olson and D. G. Riordan, “Engage to excel: Producing one million additional college graduates with degrees in science, technology, engineering, and mathematics. report to the president.” Executive Office of the President (2012).
75
[7]
X. Chen, “Stem attrition among high-performing college students: Scope and potential causes,” Journal of Technology and Science Education 5, 41–59 (2015).
[8]
X. Chen, “Stem attrition: College students’ paths into and out of stem fields. statistical analysis report. nces 2014-001.” National Center for Education Statistics (2013).
[9]
L. Aulck, R. Aras, L. Li, C. L’Heureux, P. Lu, and J. West, “Stem-ming the tide: Predicting stem attrition using student transcript data,” KDD 17 (2017).
[10] highschool, “High school benchmarks 2019: National college progression rates.” National Student Clearinghouse (2019). [11] C. S. Rozek, G. Ramirez, R. D. Fine, and S. L. Beilock, “Reducing socioeconomic disparities in the stem pipeline through student emotion regulation,” Proceedings of the National Academy of Sciences of the United States of America 116, 1553–1558 (2019). [12] E. J. Theobald, M. J. Hill, E. Tran, S. Agrawal, E. N. Arroyo, S. Behling, N. Chambwe, D. L. Cintrón, J. D. Cooper, G. Dunster, J. A. Grummer, K. Hennessey, J. Hsiao, N. Iranon, L. Jones, H. Jordt, M. Keller, M. E. Lacey, C. E. Littlefield, A. Lowe, S. Newman, V. Okolo, S. Olroyd, B. R. Peecook, S. B. Pickett, D. L. Slager, I. W. Caviedes-Solis, K. E. Stanchak, V. Sundaravardan, C. Valdebenito, C. R. Williams, K. Zinsli, and S. Freeman, “Active learning narrows achievement gaps for underrepresented students in undergraduate science, technology, engineering, and math,” Proceedings of the National Academy of Sciences of the United States of America 117, 6476–6483 (2020). [13] L. Feng, E. W. Close, C. J. Luxford, J. A. Pierson, A. Olmstead, J. Shim, V. S. Koka, and H. C. Galloway, “Transforming undergraduate stem education: The
76
learning assistant model and student retention and graduation rates,” Research in Higher Education 66, 1–26 (2025). [14] K. A. Clements, K. T. Vallone, T. P. Clements, E. H. Catania, L. L. Claiborne, K. L. Friedman, T. R. Graham, H. J. Johnson, S. R. Starko, T. D. Todd, J. Watkins, and C. J. Brame, “Impacts of learning assistants on student belonging and confidence vary across science disciplines and course contexts,” CBE Life Sciences Education 24 (2025). [15] F. X. Ran and H. Lee, “Does corequisite remediation work for everyone? an exploration of heterogeneous effects and mechanisms,” Education Finance and Policy pp. 1–58 (2025). [16] G. M. Walton, M. C. Murphy, C. Logel, D. S. Yeager, J. P. Goyer, S. T. Brady, K. T. Emerson, D. Paunesku, O. Fotuhi, A. Blodorn, K. L. Boucher, E. R. Carter, M. Gopalan, A. Henderson, K. M. Kroeper, L. A. Murdock-Perriera, S. L. Reeves, T. T. Ablorh, S. Ansari, S. Chen, P. Fisher, M. Galvan, M. K. Gilbertson, C. S. Hulleman, J. M. L. Forestier, C. Lok, K. Mathias, G. A. Muragishi, M. Netter, E. Ozier, E. N. Smith, D. B. Thoman, H. E. Williams, M. O. Wilmot, C. Hartzog, X. A. Li, and N. Krol, “Where and with whom does a brief social-belonging intervention promote progress in college?” Science 380, 499–505 (2023). [17] E. E. Shortlidge, M. J. Gray, S. Estes, and E. C. Goodwin, “The value of support: Stem intervention programs impact student persistence and belonging,” CBE Life Sciences Education 23 (2024). [18] E. Bekkering, “Course-based undergraduate research experiences (cures) for computer science?” Information Systems Education Journal (ISEDJ) 23, 4–15 (2025).
77
[19] C. Broussard, M. G. Courtney, S. L. Dunn, K. Godde, and V. Preisler, “Expanding the cure: the impact of course-based undergraduate research experiences across natural and social sciences,” Frontiers in Education 10, 1593436 (2025). [20] M.-T. Wang and J. Degol, “Staying engaged: Knowledge and research needs in student engagement hhs public access,” Child Dev Perspect 8, 137–143 (2014). [21] J. A. Fredricks, P. C. Blumenfeld, and A. H. Paris, “School engagement: Potential of the concept, state of the evidence,” Review of Educational Research 74, 59–109 (2004). [22] EngReprot, “Engagement insights: Survey findings on the quality of undergraduate education. annual results 2019.” National Survey of Student Engagement (2020). [23] R. M. Carini, G. D. Kuh, and S. P. Klein, “Student engagement and student learning: Testing the linkages,” Research in Higher Education 47, 1–32 (2006). [24] J. A. Fredricks, “Academic engagement,” International Encyclopedia of the Social & Behavioral Sciences: Second Edition pp. 31–36 (2015). [25] M. Adnan, A. Habib, J. Ashraf, S. Mussadiq, A. A. Raza, M. Abid, M. Bashir, and S. U. Khan, “Predicting at-risk students at different percentages of course length for early intervention using machine learning models,” IEEE Access 9, 7519–7539 (2021). [26] G. M. Sinatra, B. C. Heddy, and D. Lombardi, “The challenges of defining and measuring student engagement in science,” Educational Psychologist 50 (2015). [27] R. Azevedo, “Defining and measuring engagement and learning in science: Conceptual, theoretical, methodological, and analytical issues,” Educational Psychologist 50, 84–94 (2015). 78
[28] M. T. Wang, J. Fredricks, F. Ye, T. Hofkens, and J. S. Linn, “Conceptualization and assessment of adolescents’ engagement and disengagement in school: A multidimensional school engagement scale,” European Journal of Psychological Assessment 35, 592–606 (2019). [29] J. Reeve and C. M. Tseng, “Agency as a fourth aspect of students’ engagement during learning activities,” Contemporary Educational Psychology 36, 257–267 (2011). [30] E. A. Skinner and M. J. Belmont, “Motivation in the classroom: Reciprocal effects of teacher behavior and student engagement across the school year,” Journal of Educational Psychology 85, 571–581 (1993). [31] S. Lamborn, F. Newmann, and G. Wehlage, “The significance and sources of student engagement,” Student engagement and achievement in American secondary schools pp. 11–39 (1992). [32] M. Csikszentmihalyi, “Flow and education,” Applications of Flow in Human Development and Education: The Collected Works of Mihaly Csikszentmihalyi pp. 129–151 (2014). [33] E. I. Knudsen, “Fundamental components of attention,” Annual Review of Neuroscience 30, 57–78 (2007). [34] M. Csikszentmihalyi, Flow: The psychology of optimal experience, vol. 1990 (Harper & Row New York, 1990). [35] D. J. Shernoff, M. Csikszentmihalyi, B. Schneider, and E. S. Shernoff, “Student engagement in high school classrooms from the perspective of flow theory,” Applications of Flow in Human Development and Education: The Collected Works of Mihaly Csikszentmihalyi pp. 475–494 (2014).
79
[36] J. A. Fredricks, “The measurement of student engagement: Methodological advances and comparison of new self-report instruments,” Handbook of Research on Student Engagement: Second Edition pp. 597–616 (2022). [37] D. J. Shernoff and S. Twersky, “Flow in schools reexamined: Cultivating engagement in learning from classrooms to educational games,” Handbook of Positive Psychology in Schools: Supporting Process and Practice pp. 283–294 (2022). [38] J. D. Gobert, R. S. Baker, and M. B. Wixon, “Operationalizing and detecting disengagement within online science microworlds,” Educational Psychologist 50, 43–57 (2015). [39] S. Li, J. Zheng, and S. P. Lajoie, “The relationship between cognitive engagement and students’ performance in a simulation-based training environment: an information-processing perspective,” Interactive Learning Environments 31, 1532–1545 (2023). [40] B. T. Carter and S. G. Luke, “Best practices in eye tracking research,” International Journal of Psychophysiology 155, 49–62 (2020). [41] X. Tang, Y. Gong, Y. Xiao, J. Xiong, and L. Bao, “Facial expression recognition for probing students’ emotional engagement in science learning,” Journal of Science Education and Technology 34, 13–30 (2025). [42] I. Alkabbany, A. M. Ali, C. Foreman, T. Tretter, N. Hindy, and A. Farag, “An experimental platform for real-time students engagement measurements from video in stem classrooms,” Sensors 2023, Vol. 23, Page 1614 23, 1614 (2023). [43] A. Abdelkawy, A. Farag, I. Alkabbany, A. Ali, C. Foreman, T. Tretter, and N. Hindy, “Measuring student behavioral engagement using histogram of actions,” Pattern Recognition Letters 186, 337–344 (2024).
80
[44] T. S. Ashwin and R. M. R. Guddeti, “Unobtrusive behavioral analysis of students in classroom environment using non-verbal cues,” IEEE Access 7, 150693– 150709 (2019). [45] P. W. Kim, “Real-time bio-signal-processing of students based on an intelligent algorithm for internet of things to assess engagement levels in a classroom,” Future Generation Computer Systems 86, 716–722 (2018). [46] K. Walter and P. Bex, “Cognitive load influences oculomotor behavior in natural scenes,” Scientific Reports 11, 1–12 (2021). [47] X. Liu and Y. Cui, “Eye tracking technology for examining cognitive processes in education: A systematic review,” Computers & Education 229, 105263 (2025). [48] P. Ekman and W. V. Friesen, “Facial action coding system,” Environmental Psychology & Nonverbal Behavior (1978). [49] P. Ekman, “Facial expression and emotion,” American Psychologist 48, 384– 392 (1993). [50] S. Li and W. Deng, “Deep facial expression recognition: A survey,” IEEE Transactions on Affective Computing 13, 1195–1215 (2022). [51] C. A. Corneanu, M. O. Simón, J. F. Cohn, and S. E. Guerrero, “Survey on rgb, 3d, thermal, and multimodal approaches for facial expression recognition: History, trends, and affect-related applications,” IEEE Transactions on Pattern Analysis and Machine Intelligence 38, 1548–1568 (2016). [52] L. Huang, F. Du, W. Huang, H. Ren, W. Qiu, J. Zhang, and Y. Wang, “Threestage dynamic brain-cognitive model of understanding action intention displayed by human body movements,” Brain Topography 37, 1055–1067 (2024).
81
[53] N. Dael, M. Mortillaro, and K. R. Scherer, “Emotion expression in body action and posture,” Emotion 12, 1085–1101 (2012). [54] J. K. Aggarwal and M. S. Ryoo, “Human activity analysis,” ACM Computing Surveys (CSUR) 43 (2011). [55] R. Poppe, “A survey on vision-based human action recognition,” Image and Vision Computing 28, 976–990 (2010). [56] I. Alkabbany, A. Ali, A. Farag, I. Bennett, M. Ghanoum, and A. Farag, “Measuring student engagement level using facial information,” Proceedings - International Conference on Image Processing, ICIP 2019-September, 3337–3341 (2019). [57] X. Sheng, S. Li, and S. Chan, “Real-time classroom student behavior detection based on improved yolov8s,” Scientific Reports 15, 1–11 (2025). [58] O. M. Nezami, M. Dras, L. Hamey, D. Richards, S. Wan, and C. Paris, “Automatic recognition of student engagement using deep learning and facial expression,” Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 11908 LNAI, 273–289 (2020). [59] M. Singh, X. Hoque, D. Zeng, Y. Wang, K. Ikeda, and A. Dhall, “Do i have your attention: A large scale engagement prediction dataset and baselines,” ACM International Conference Proceeding Series pp. 174–182 (2023). [60] J. Whitehill, Z. Serpell, Y. C. Lin, A. Foster, and J. R. Movellan, “The faces of engagement: Automatic recognition of student engagement from facial expressions,” IEEE Transactions on Affective Computing 5, 86–98 (2014).
82
[61] J. E. Symonds, A. Kaplan, K. Upadyaya, K. S. Aro, B. M. Torsney, E. Skinner, and J. S. Eccles, “Momentary student engagement as a dynamic developmental system,” Journal of Theoretical and Philosophical Psychology (2024). [62] R. P. Ltd., “Raspberry pi 5 product brief,” (2025). [63] NVIDIA Corporation, “Jetson orin nano developer kit,” https://nvidia. com/en-us/autonomous-machines/embedded-systems/jetson-orin/ nano-super-developer-kit/ (2023). Accessed: 2025-02-10. [64] Raspberry Pi Foundation, “Raspberry pi 4 model b,” https://raspberrypi. com/products/raspberry-pi-4-model-b/ (2019). Accessed: 2025-02-10. [65] Orange
Pi,
“Orange
pi
5
plus,”
http://orangepi.org/html/
hardWare/computerAndMicrocontrollers/service-and-support/ Orange-Pi-5-plus.html (2023). Accessed: 2025-02-10. [66] Radxa, “Radxa rock 5b,” https://radxa.com/products/rock5/5b/ (2022). Accessed: 2025-02-10. [67] BeagleBoard.org
Foundation,
“Beaglebone
black,”
https://www.
beagleboard.org/black (2013). Accessed: 2025-02-10. [68] BeagleBoard.org Foundation, “Beagleplay,” https://www.beagleboard.org/ boards/beagleplay (2023). Accessed: 2025-02-10. [69] Luxonis, “Oak-d lite: Compact spatial ai camera,” https://docs.luxonis. com/hardware/products/OAK-D%20Lite. Accessed: 2025-12-06. [70] I. Corporation, “Intel® movidius™ myriad™ x vision processing unit 4gb — specifications,” https://www.intel.com/content/www/us/en/products/ sku/125926/intel-movidius-myriad-x-vision-processing-unit-4gb/ specifications.html. Accessed: 2025-12-09.
83
[71] Raspberry Pi Ltd., “Raspberry pi camera module 3,” https://www. raspberrypi.com/products/camera-module-3/. Accessed: 2025-12-06. [72] Luxonis, “Oak-d s2: Stereo depth ai camera,” https://docs.luxonis.com/ hardware/products/OAK-D%20S2. Accessed: 2025-12-06. [73] Luxonis, “Oak-d pro: Stereo depth camera with active ir,” https://docs. luxonis.com/hardware/products/OAK-D%20Pro. Accessed: 2025-12-06. [74] Intel Corporation,
“Intel realsense depth camera d435x,”
https://
realsenseai.com/products/stereo-depth-camera-d435/. Accessed: 202512-06. [75] Raspberry Pi Ltd., “Raspberry pi touch display 2,” https://raspberrypi. com/products/touch-display-2/ (2024). Accessed: 2025-02-10. [76] Geekworm, “Geekworm x1200 ups hat for raspberry pi 5,” https://wiki. geekworm.com/X1200 (2024). Accessed: 2025-02-10. [77] P. S. Foundation, “Python programming language,” https://www.python. org/. Accessed: 2025-12-09. [78] T. Q. Company, “Qt for python — documentation,” https://doc.qt.io/ qtforpython-6/index.html. Accessed: 2025-12-09. [79] Luxonis, “Depthai — documentation (software v3),” https://docs.luxonis. com/software-v3/depthai/. Accessed: 2025-12-09. [80] A. W. Services, “Boto3 — aws sdk for python documentation,” https: //boto3.amazonaws.com/v1/documentation/api/latest/index.html. Accessed: 2025-12-09. [81] pytest development team, “pytest — documentation,” https://docs.pytest. org/. Accessed: 2025-12-09.
84
[82] A. Software, “Ruff — documentation,” https://docs.astral.sh/ruff/. Accessed: 2025-12-09. [83] C.-F. Community, “Miniforge — a community Conda installer,” https:// github.com/conda-forge/miniforge. Accessed: 2025-12-10. [84] N. I. of Standards and Technology, “Secure hash standard (shs),” FIPS Publication 180-4, U.S. Department of Commerce (2015). Defines SHA-1, SHA-224, SHA-256, SHA-384, and SHA-512. Accessed: 2025-12-09. [85] A. W. Services, “Amazon simple storage service (s3),” https://aws.amazon. com/s3/. Accessed: 2025-12-09. [86] I. MinIO, “Minio — high-performance object storage,” https://www.min.io/. Accessed: 2025-12-09. [87] E. Rescorla, “The transport layer security (tls) protocol version 1.3,” RFC 8446, Internet Engineering Task Force (2018). Accessed: 2025-12-09. [88] D. S. Foundation, “Django web framework,” https://www.djangoproject. com/. Accessed: 2025-12-09. [89] O. Corporation, “Mysql database system,” https://www.mysql.com/. Accessed: 2025-12-09. [90] G. Developers, “Gunicorn — python wsgi http server for unix,” https: //gunicorn.org/. Accessed: 2025-12-09. [91] I. F5 NGINX, “Nginx web server,” https://nginx.org/. Accessed: 2025-1209. [92] C. Ltd., “Ubuntu 20.04 lts (focal fossa),” https://releases.ubuntu.com/20. 04/. Accessed: 2025-12-09.
85
[93] S. Labs, “Step ca — online certificate authority,” https://smallstep.com/ docs/step-ca/. Accessed: 2025-12-09. [94] R. T. Fielding, “Architectural styles and the design of network-based software architectures,” Ph.D. thesis, University of California, Irvine (2000). [95] D. Mills, J. Martin, J. Burbank, and W. Kasch, “Network time protocol version 4: Protocol and algorithms specification,” RFC 5905, Internet Engineering Task Force (2010). Accessed: 2025-12-09. [96] S. Garrido-Jurado, R. Muñoz-Salinas, F. J. Madrid-Cuevas, and M. J. Marı́nJiménez, “Automatic generation and detection of highly reliable fiducial markers under occlusion,” Pattern Recognition 47, 2280–2292 (2014). [97] P. Garcı́a-Ruiz, F. J. Romero-Ramirez, R. Muñoz-Salinas, M. J. Marı́n-Jiménez, and R. Medina-Carnicer, “Large-scale indoor camera positioning using fiducial markers,” Sensors 24 (2024). [98] A. Savchenko, “Facial expression recognition with adaptive frame rate based on multiple testing correction,” in “Proceedings of the 40th International Conference on Machine Learning (ICML),” , vol. 202 of Proceedings of Machine Learning Research, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, eds. (PMLR, 2023), vol. 202 of Proceedings of Machine Learning Research, pp. 30119–30129. [99] M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in “International conference on machine learning,” (PMLR, 2019), pp. 6105–6114. [100] W. Wu, H. Peng, and S. Yu, “Yunet: A tiny millisecond-level face detector,” Machine Intelligence Research 20, 656–665 (2023).
86
[101] OpenVINO Toolkit, “Head pose estimation — adas 0001 — model documentation,” https://docs.openvino.ai/2023.3/omz_models_model_head_ pose_estimation_adas_0001.html (2023). Accessed: 2025-12-08. [102] OpenVINO Toolkit,
“Gaze estimation demo (c++) — open model
zoo documentation,” https://docs.openvino.ai/2023.3/omz_demos_gaze_ estimation_demo_cpp.html (2023). Accessed: 2025-12-08. [103] C. Lugaresi, J. Tang, H. Nash, C. McClanahan, E. Uboweja, M. Hays, F. Zhang, C.-L. Chang, M. Yong, J. Lee et al., “Mediapipe: A framework for perceiving and processing reality,” in “Third workshop on computer vision for AR/VR at IEEE computer vision and pattern recognition (CVPR),” , vol. 2019 (2019), vol. 2019.
87
CURRICULUM VITAE
Ahmed Elsayed
Education • B.Sc. in Electronics and Communications Engineering, Zagazig University, Zagazig, Egypt
Sept. 2011 – Jun. 2016
GPA: Excellent (87.73%) Capstone Project: Talking Gloves — an American Sign Language recognition system enabling communication for deaf individuals.
Professional Positions • Computer Vision Engineer, Apple, Vision Pro Hands Team, Sunnyvale, CA, USA
Jan. 2023 – Feb. 2024
Led refactoring of critical dataset, training, metrics, and evaluation pipelines; redesigned dataset-generation workflows (reducing creation time from 2 days to 30 minutes); developed model-selection tools and label-correction frameworks; optimized user-study labeling processes; contributed extensively to the Sahara repository; resolved FFMPEG frame mismatch bug, improving touch/no-touch label quality.
88
• Graduate Research Assistant, Computer Vision and Image Processing (CVIP) Lab, University of Louisville
Jan. 2022 – Dec. 2023
Projects include: 3D jaw reconstruction using an inter-oral sensor network; real-time student engagement measurement in STEM classrooms (developed a MEAN-stack web application); ATRV autonomous robot upgrades; robotic arm system for autonomous vehicle refueling. • Hardware Engineer (Part-Time), Egyptian Space Agency (EgSA) Oct. 2020 – Dec. 2021 Designed, simulated, and laid out the ZU-CubeSat communication subsystem PCB (UHF band). • Teaching Assistant, Faculty of Engineering, Zagazig University Nov. 2016 – Dec. 2021 Assisted in teaching: Electronic Devices, Analog Circuits, Integrated Circuits, RF Electronics. Led labs in: Electronic Circuit Design, HDL (Verilog/VHDL), PCB, IoT, Arduino, MATLAB. • Hardware & Embedded Software Engineer, Upwork Freelancing Platform May 2018 – Apr. 2021 Delivered PCB designs, embedded systems, IoT prototypes, high-speed digital boards, and DC-DC converters for global clients. • Cofounder & Hardware Engineer, ZagSystems Startup Company May 2018 – Dec. 2021 Designed hardware for IoT-based home automation and data-collection systems.
89
Publications and Presentations 1. Samir Harb; A. Elsayed; M. Yousuf; I. Alkabbany; A. Ali; S. Elshazley, “Accurate Colon Segmentation Using 2D Convolutional Neural Networks With 3D Contextual Information,” 2024 IEEE International Conference on Image Processing (ICIP), Abu Dhabi, United Arab Emirates, 2024, pp. 3212–3218, doi: 10.1109/ICIP51287.2024.10647313. 2. Samir Harb, M. Yousuf, A. Elsayed, A. Ali, S. Elshazly, and A. Farag, “Deep Learning-Based Colon Segmentation for Accurate Colorectal Polyps Detection,” in Cancer Detection and Diagnosis, CRC Press, 1st Edition, 2025, pp. 9.
90