arXiv:2605.20182v1 [cs.LG] 19 May 2026
Atoms of Thought: Universal EEG Representation Learning with Microstates Xinyang Tian∗
Ruitao Liu∗
Ziyi Ye†
Institute for Interdisciplinary Information Sciences, Tsinghua University Beijing, China [email protected]
Institute for Interdisciplinary Information Sciences, Tsinghua University Beijing, China [email protected]
Institute of Trustworthy Embodied AI, Fudan University Shanghai, China [email protected]
Siyang Xue
Xin Wang
Xuesong Chen‡
School of Clinical Medicine, Tsinghua University Beijing, China [email protected]
Beijing Five Seasons Medical Technology Co., Ltd. Beijing, China [email protected]
Beijing Five Seasons Medical Technology Co., Ltd. Beijing, China [email protected]
Abstract Learning universal representations from electroencephalogram (EEG) signals is a cutting-edge approach in the field of neuroinformatics and brain-computer interfaces (BCIs). Conventionally, EEG is treated as a multivariate temporal signal, where time- or frequency-domain features are extracted for representation learning. This paper investigates a simple yet effective EEG representation, i.e., microstates. Microstates represent the building blocks of brain activity patterns at a microscopic time scale. We build a universal microstate tokenizer from a large medical EEG dataset by clustering continuous EEG signals into sequences of discrete microstates. The microstate tokenizer is then adopted universally across a series of downstream tasks, including sleep staging, emotion recognition, and motor imagery classification. Experimental results show that EEG representation learning with microstates outperforms traditional time-domain and frequency-domain features under different models and across different tasks. Further analysis shows that microstates offer greater interpretability and scalability, thereby opening up applications in both cognitive neuroscience and clinical research.
Keywords EEG Analysis, Microstates, Sleep Staging, Emotion Recognition, Motor Imagery Classification Accepted by the 3rd International Workshop on Multimodal and Responsible Affective Computing (MRAC 2025). Version of Record DOI: 10.1145/3746270.3760230.
1
Introduction
Electroencephalogram (EEG) signals provide valuable insights into brain activity, making them indispensable in fields such as clinical medicine, neuroscience, and cognitive psychology [80]. For example, EEG has been widely used in clinical settings to detect certain diseases and anomalies [9, 38, 59], to investigate neural ∗ Both authors contributed equally to this research. † Research was conducted as a Ph.D. student at Tsinghua University. ‡ Corresponding author. Email: [email protected]
mechanisms underlying cognitive processes [40, 56], and to design brain-computer interfaces [26]. Recently, with the maturity of deep learning techniques, integrating AI technology with EEG analysis has become the new paradigm, significantly improving classification performance in downstream tasks [1, 3]. Despite these merits, EEG signals are highly non-linear and non-stationary [62], which pose challenges to extracting effective representations from EEG signals. Conventionally, EEG is treated as multivariate time series data with features extracted in the time and frequency domain for further analysis [62]. Such features come with two major drawbacks. On the one hand, they are susceptible to artifacts and are prone to be confined within a task-specific and subject-specific representation space. Conventional time and frequency domain EEG representations will inevitably incorporate artifacts [58] related to eye blinks, myoelectricity, and the environment. Additionally, time and frequency domain EEG representations vary significantly across subjects and tasks, making it challenging to generalize. This results in suboptimal performance on a single task and degraded generalizability across different tasks [29], and requires huge amounts of task-specific data, which are usually unavailable. On the other hand, time- and frequency-domain features are unable to uncover transient and dynamic information. Conventional methods often struggle to capture highresolution EEG features. Time-domain information, which directly utilizes raw EEG signals, is often considered inefficient due to its low signal-to-noise ratio (SNR) [67]. Frequency-domain information uses a fixed window length, which consequently obscures temporal resolution and results in a certain degree of information loss [14, 62]. To address these challenges, we introduce a novel approach that integrates deep learning with a biologically grounded concept in EEG analysis: EEG microstates [35]. EEG microstates are quasistable discrete patterns of scalp electrical potential that last for brief periods, typically 60-120 milliseconds [43]. While conventional features tend to ignore the physiological and clinical context of EEG signals, EEG microstates are believed to correspond to fundamental and stable cognitive states [20, 43, 76]. Previous researchers have revealed a series of underlying mechanisms of thought and cognition [10, 36, 44, 56]. Building on such results, we leverage EEG
Figure 1: Visualization of Different Representations and Downstream Tasks. Conventional representations mainly reside in the time domain and frequency domain. We propose the microstate representation, which is a universal representation that outperforms other representations under different model structures and across different tasks [8, 15, 34].
2.1
microstates as a discrete and intrinsic representation of brain activity that is more aligned with the underlying neural mechanisms, improving both interpretability and robustness. We validate the effectiveness of EEG microstates across three critical tasks—sleep staging, emotion recognition, and motor imagery [6, 80] and with different models, showing superior performance compared to conventional representations. Moreover, we test the accuracy of EEG representation learning with increasing data size, observing that EEG microstates show greater performance gain than conventional features. Furthermore, we investigate the distribution of EEG microstates across various cognitive functions and present a potential relationship to interpret cognitive functions with EEG microstates. The main contributions of this work are as follows:
2
EEG Microstates in Cognitive Neuroscience
• We introduce EEG microstates as a universal representation of brain activity, bridging the gap between deep learning techniques and neural activity patterns. • We demonstrate the effectiveness of this microstate-based approach in three critical tasks—sleep staging, emotion recognition, and motor imagery classification, and with different model structures. Experimental results indicate that the microstate tokenizer initialized in one task can be generalized to a series of downstream tasks, showcasing its universal applicability and alleviating the impact of data scarcity. • We conduct in-depth analysis showing that EEG microstate is more scalable than time-domain and frequency-domain methods and can serve as an explainable feature linking to various cognitive functions.
EEG microstate analysis was first introduced by Lehmann et al. [35], and has gained significant attention as a promising tool for representing brief, stable patterns of brain activity. Microstates are thought to reflect fundamental cognitive states that the brain switches between, providing valuable insights into the temporal organization of brain function [43]. Studies have shown that various diseases, such as epilepsy, sleep disorder and Alzheimer’s disease, can alter EEG microstates [11, 23, 33, 39, 52, 63]. Recent research has applied microstate analysis to a wide range of cognitive tasks, including emotion, attention, and social abilities [22, 24, 51, 54, 55], demonstrating the effectiveness of microstates in understanding cognitive and pathological states. The most common approach to producing microstates originates from Pascual-Marqui et al. [47]. They used the k-means clustering method to conduct the EEG microstate analysis, which further become the most popular technique for microstate classification. Other studies have introduced alternative methods for microstate analysis, which are based on a series of clustering algorithms [28, 41, 42, 45, 50]. However, most existing research has focused on interpreting microstates based on the physical conditions of subjects, while efforts to learn EEG representations for downstream classification and detection tasks remain limited. Moreover, the interpretability of microstates and their connection to fundamental cognitive states make them a promising candidate for representing EEG signals in contemporary deep-learning models, yet no current studies have tested this potential.
Related Work
2.2
EEG analysis has long been a critical tool in both clinical diagnosis and research, with various representation learning methods to enhance the accuracy of diagnosis. This section elaborates on a variety of techniques developed to extract meaningful information from the brain’s electrical activity, particularly in the medical and deep learning fields.
Representation Learning for EEG Analysis
Machine learning, especially deep learning techniques have been increasingly integrated into EEG analysis to improve the accuracy and efficiency of EEG-based classification tasks. Typically, machine learning models require EEG representations extracted from the raw signals as input, which can be broadly categorized into timeand frequency-domain features. On the one hand, raw EEG itself 2
3.3
can serve as the most straightforward time-domain representation. Al-Hussaini et al. [5] used fixed-length windows of 30s segmented from raw EEG signals during prototype learning for sleep staging, which treated the signals as multivariate time series data. Perslev et al. [48] also used raw EEG signals as their CNN-based model representation for sleep staging. On the other hand, information in the frequency domain is also commonly extracted as EEG representations. V. and Bhattacharyya [64] used multivariate variational mode decomposition (MVMD) to extract spectral information for emotion recognition. Zheng et al. [78] used the Hilbert-Huang transform to analyze scalp EEG signals. It has been shown in [71, 81] that using frequency-domain information improves performance in emotion recognition. Despite the above achievements brought about by deep learning, conventional representations often contain person- or task-specific artifacts [60, 70, 72]. Due to the models’ susceptibility to noise and artifacts, training such models either undermines their performance and generalizability, or requires a huge amount of person- or taskspecific data. To address these challenges, Afzal et al. [1] proposed a novel graphical representation of raw EEG data, which improves seizure detection but is still task-specific. Based on the development of timedomain representations and suitable model structures [17, 46, 74] and inspired by the development in natural language processing (NLP), Gui et al. [27] proposed a vector quantization pre-training method to obtain representations for downstream tasks. Wang et al. [69] also utilized a pre-training paradigm to extract relevant representations by spatio-temporal representation alignment in order to depict the brain. They observed that such representations can be better generalized across downstream tasks, but consume a large amount of computational power and time. Moreover, the input EEG signal of the pre-trained model is still treated as multivariate time series data.
3
Frequency Bands. The frequency-domain representations are extracted based on the frequency power distribution among frequency bands[75]. The frequency domain are divided into several frequency bands, including the 𝛿-band (0.5 ∼ 4Hz), 𝜃 -band (4 ∼ 8Hz), 𝛼-band (8 ∼ 12Hz), 𝜎-band (12 ∼ 16Hz), 𝛽-band (16 ∼ 30Hz) and 𝛾-band (30 ∼ 40Hz). Short-Time Fourier Transform (STFT). Time-frequency transformation can be carried out via numerous methods, namely short-time Fourier transform (STFT), discrete/continuous wavelet transform (DWT/CWT), and empirical mode decomposition (EMD) [4, 32]. Short-time Fourier transform, owing to its straightforwardness and thorough theoretical analysis, is applied in many EEG-related tasks [12, 18, 30, 68]. Consequently, we choose this method as our frequency-domain baseline. Given the raw EEG signal of a single measurement channel 𝑠𝑠𝑖𝑛 ∈ R 𝑓𝑠 𝑇 , we perform short-time Fourier transform with fixed window size 𝑡 𝑤 and overlap ratio 𝑟𝑜 . The length after the short-time Fourier transform will be 𝑓𝑠𝑇 − 𝑓𝑠 𝑡 𝑤 𝑙 𝑓 𝑟𝑒𝑞 = +1 (1 − 𝑟𝑜 )𝑓𝑠 𝑡 𝑤 and if we leave out the margin then the resulting length will be 𝑓𝑠𝑇 = 𝑓 𝑓 𝑟𝑒𝑞𝑇 (1 − 𝑟𝑜 )𝑓𝑠 𝑡 𝑤 1 𝑓 𝑓 𝑟𝑒𝑞 = (1 − 𝑟𝑜 )𝑡 𝑤 𝑙 𝑓 𝑟𝑒𝑞 =
EEG Representations
Using STFT, we obtain a frequency axis 𝐹𝑎 and time axis 𝑇𝑎 and the amplitude of the signal at each frequency 𝑓 ∈ 𝐹𝑎 and time point 𝑡 ∈ 𝑇𝑎 . Note that |𝑇𝑎 | = 𝑙 ′ and hence the output shape is ′ 𝑠 𝑓 ,𝑡,𝑠𝑖𝑛 ∈ R |𝐹𝑎 | ×𝑙 .
This section lists conventional EEG representations in the timedomain and frequency-domain, and our microstate representation. It also elaborate on detailed methods and procedures to construct different representations.
3.1
Band Power Integration. Now that we have obtained the single channel data 𝑠 𝑓 𝑟𝑒𝑞,𝑠𝑖𝑛 ∈ R𝐹 ×𝑙 𝑓 𝑟𝑒𝑞𝑇 where 𝐹 is the frequency resolution, we integrate the rows that correspond to each frequency band to obtain the total power within that band. In this case, the integration result will have shape R𝐵×𝑙 𝑓 𝑟𝑒𝑞𝑇 . By flattening and stacking all channels, the final result has shape 𝑠 𝑓 𝑟𝑒𝑞 ∈ R𝑁 ×𝐵𝑓 𝑓 𝑟𝑒𝑞𝑇 .
Problem formulation
The objective of EEG signal analysis and physical state prediction can be defined as follows: We are given the input raw EEG signal 𝑠 ∈ R𝐶 ×𝑓𝑠 𝑇 where 𝐶 denotes the number of channels, 𝑓𝑠 is the sampling frequency and 𝑇 is the sample duration. The EEG signal analysis aims to predict the physical state of the sampled subject, which can be represented by a sequence of discrete labels 𝑙 = 𝐿 𝑓𝑙 𝑇 where 𝐿 = {𝑎 1, 𝑎 2, . . . , 𝑎𝑚 } is the set of labels and 𝑓𝑙 is the state frequency.
3.2
Frequency-Domain Features Extraction
Raw EEG signals often obscure frequency information, and thus, sometimes directly using it does not produce desirable results. Therefore, a common approach to handling such time-domain signals is to use their corresponding frequency-domain signals as features [65].
3.4
Unsupervised Microstate Tokenizer
Our clustering method is based on k-means [47]. But to guarantee that the clustering model has a sufficient level of generalizability, unlike previous works, we have to fit the clustering model on a huge amount of data. To allow the model to take all data points into consideration without consuming too much memory, we use incremental learning by dividing data into small batches of size 𝑛.
Time-Domain Features Extraction
The most straightforward approach for EEG analysis is to directly utilize the time-domain information features, i.e., the raw EEG signals [62]. In this setting, the raw EEG signal is sliced into fixedlength windows with duration 𝑇𝑤 , and the feature 𝑠𝑡𝑖𝑚𝑒,𝑤 will be of the shape of R𝐶 ×𝑓𝑠 𝑇𝑤 , with the corresponding labels 𝑙 𝑤 ∈ 𝐿 𝑓𝑙 𝑇𝑤 .
3.4.1 Stream Clustering [2]. The original k-means algorithm clusters the data points by initially selecting 𝑘 centers, grouping data points according to their distance to the centers, and computing 3
Figure 2: Pipeline of our Work. Our work can be broken into two parts. The first involves fitting a tokenizer to extract microstates from six EEG channels F3, F4, C3, C4, O1, O2. The second consists of training different models on the microstate signals and performing downstream tasks. The tokenizer of the first stage is independent of the models of the second stage. For each task, the model always includes an embedding layer to convert the discrete microstates into high-dimensional embeddings [8]. • We want to test whether we can model brain activity during sleep data, which can be generalized to downstream tasks during wakefulness
the centroid of each cluster as the new centers. The algorithm terminates when reaching the maximum iterations or the positions of the centers have converged. Streaming k-means follows a fashion similar to classical k-means, but each time it uses only a small batch of data to update the cluster centers. Therefore, we do not need to store the entire dataset in memory, while only need to store a batch. This renders better performance than using only a small portion of data, and reduces memory consumption compared to clustering on all data points. The procedure is shown in Algorithm 1.
Extracting target channels. Since PSG signal involves numerous components like EEG, EOG, and ECG, the number of channels comes with varying sizes due to loss of data or shortage of equipment. Consequently, we filter out all channels except 𝑁 of them that are present in all samples. The generic method will be clustering the data by treating them as 𝑁 dimensional points. In the HSP dataset, only the channels F3, F4, C3, C4, O1, O2 are present across a relatively large amount of subjects, whereas other channels only appears sporadically among very limited number of subjects. Consequently these 6 leads are selected for clustering, since we have to extract a generalizable representation across subjects in order to obtain a truly universal representation. After extracting the necessary channels from the original data 𝑠 ∈ R𝐶 ×𝑓𝑠 𝑇 we obtain the filtered data 𝑠𝑒𝑥𝑡 ∈ R𝑁 ×𝑓𝑠 𝑇 where 𝑁 = 6.
Algorithm 1 Streaming K-Means Initialize cluster centers 𝑐 1, 𝑐 2, . . . , 𝑐𝑘 ∈ R𝐶 𝑖𝑡𝑒𝑟 ← 0 while 𝑖𝑡𝑒𝑟 < 𝑚𝑎𝑥_𝑖𝑡𝑒𝑟 and centers have not converged do Get a new batch 𝑑 1, 𝑑 2, . . . , 𝑑𝑛 ∈ R𝐶 𝑆𝑖 ← {𝑑 𝑗 | arg min𝑡 ∥𝑑 𝑗 − 𝑐𝑡 ∥ 2 = 𝑖} Í 𝑖| 𝑐𝑖 ← |𝑆1𝑖 | 𝑟|𝑆=1 𝑆𝑖 𝑟 end while
Filtering. A bandpass filter with low pass 1Hz and high pass 40Hz is applied to retain the most relevant frequency bands (e.g., delta, theta, alpha, beta). The array shape remains unchanged during this operation.
3.4.2 Fitting the Clustering Model. We adopt the following experimental setup.
Resampling. The HSP dataset comes with different sample frequencies, including 256Hz and 512Hz. We resample the signals to 100Hz for better clustering results. Now the data shape becomes 𝑠𝑟𝑒𝑠 ∈ R𝑁 ×𝑓𝑟𝑒𝑠 𝑇 with 𝑓𝑟𝑒𝑠 = 100Hz.
Dataset. We select the Human Sleep Project (HSP) dataset to fit the clustering model [73]. The dataset includes polysomnography (PSG) data from over 20K subjects and is still growing in size. The reasons why we use EEG data recorded during sleep are as follows: • It is usually difficult and costly to record EEG signals during wakefulness, and these datasets are usually limited in size • Currently the community has abundant sleep data. They can also reflect the cognitive level and consciousness, despite the fact that their being less organized, noisy and has a limited number of channels (usually 6 channels)
Global field power (GFP) peaks extraction. Global field power is computed as the standard deviation of all sensors. The peaks are defined as their local maxima, which have the highest signal-tonoise ratio [43]. Therefore, we extract GFP peaks on the six channels. This produces the final input for the clustering model, which has size 𝑠𝑔𝑓 𝑝 ∈ R𝑁 ×𝑡 where 𝑡 denotes the number of GFP peaks in the sequence R𝑁 ×𝑓𝑟𝑒𝑠 𝑇 . 4
Table 1: Classification accuracies and model parameters for sleep staging under different representations and different model architectures on the Human Sleep Project (HSP) dataset. The highest performance among all representations under a certain model is highlighted in boldface. Representation Acc Raw EEG (Time Domain) STFT (Freqency Domain) Microstates (Ours)
0.710 0.778 0.801
CNN+LSTM Kappa params 0.597 0.690 0.722
707K 692K 687K
3.4.3 Constructing the Microstates. Having fitted clustering model, we can apply it to raw EEG signals of shape 𝑠 ∈ R𝑁 ×𝑓𝑠 𝑇 to obtain the microstate sequence 𝑐 ∈ 𝑆 𝑓𝑠 𝑇 , 𝑆 ∈ {𝑏 1, 𝑏 2, . . . , 𝑏𝑘 } with 𝑘 = 1000. 𝑆 is the set of microstates.
0.786 0.790 0.810
0.793 0.794 0.810
0.702 0.710 0.736
3.2M 3.2M 3.4M
Representation Raw EEG (Time Domain) STFT (Freqency Domain) Microstates (Ours)
Downstream Tasks
In this section, we introduce the experimental setup for testing the performance of our microstate representation and conventional representations. The microstate sequence is produced by the tokenizer trained in the previous section (the fitted KMeans in Figure 2).
4.1
Sleep Net Zero Acc Kappa params 0.713 0.711 0.736
10.9M1 3.2M 3.2M
Table 2: Classification accuracies and model parameters for emotion recognition (CNN-based model, SEED dataset) under different representations. The highest performance among all representations is highlighted in boldface.
Fitting the clustering model. After obtaining the data of shape 𝑠𝑔𝑓 𝑝 ∈ R𝑁 ×𝑡 , we set the number of clusters 𝑘 = 1000 and 𝑛 = 50 for batch size and fit the GFP peaks.
4
Sleep Transformer Acc Kappa params
Accuracy
Kappa
Params
0.846 0.797 0.862
0.769 0.694 0.793
19.1M 19.1M 20.1M
For time- and frequency-domain signals, we change the embedding layer into a convolution layer which functions similarly as the embedding.
Sleep Staging Loss Function. The loss function is set as cross-entropy loss. Suppose that the output of the fully connected layer is (ℎ 1, ℎ 2, ℎ 3, ℎ 4, ℎ 5 ), which are the scores of the five classes. We perform a softmax on the scores and the cross entropy loss is defined as follows:
During different sleep stages, brain activity varies accordingly. This lies the foundation for predicting sleep stages using EEG signals. Dataset. For sleep staging, we use the HSP dataset [73] mentioned in fitting the clustering model to test the performance of our representation.
loss = −
Sleep stages. For sleep staging, 𝐿 consists of the five sleep stages: {W,N1,N2,N3,R}. W corresponds to wake stage, N1, N2 and N3 correspond to different non-rapid eye movement (NREM) stages, and R corresponds to rapid eye movement (REM) stage.
5 ∑︁
𝑝 (𝑖) log Softmax(ℎ𝑖 )
𝑖=1
which is minimized when the correct label 𝑗 has score ℎ 𝑗 significantly larger than the other labels.
4.2
Preprocessing. The EEG signal is filtered and resampled to 𝑓𝑟𝑒𝑠 = 100Hz. For extracting frequency-domain information, we choose 𝑡 𝑤 = 1s and 𝑟𝑜 = 0. Finally, the EEG signal is slices into 𝑇𝑤 = 300s windows.
Emotion Recognition
Emotions are a key part of our physical state, and have a strong connection with brain activity. Consequently, our work involves training a microstate-based model for emotion classification.
Model Architecture. To test the universality of our microstate representation, we adopt the following model structures: CNN+LSTM.[57] This model architecture consists of 3 CNNs and 2 GRUs for extracting spatial and temporal information, and fully connected layers for classification. An embedding layer is added for our microstate representation. Sleep Transformer.[49] Sleep Transformer uses 2 Roformer [61] layers to extract the local and global features before inputting into the final linear layer. An embedding layer and 2 CNNs are added for our microstate representation. Sleep Net Zero.[37] The model consists of a feature extraction unit composed of several residual blocks, a Roformer layer and a linear layer. To adapt to our microstate representation, we add an embedding layer and 4 CNNs, and remove the feature extraction unit.
Dataset. We use the SEED dataset [19, 79] for emotion recognition. The SEED dataset consists of 15 subjects watching video clips that express different emotions. The overall tone is categorized as positive, negative, and neutral. The EEG signal is recorded with 62 channels and at a frequency 200Hz. Emotion Labels. We utilize the overall tone of each movie clip as our labels, and thus 𝐿 consists of positive, negative, and neutral. Preprocessing. We directly use the sample frequency 𝑓𝑠 = 200Hz. We pad all samples to 𝑇𝑤 = 265s. Other configurations are the same as in sleep staging. Model Architecture. The model used here is a CNN-based classifier [31] with similar modifications and cross-entropy loss. It has 5 CNNs and 4 linear layers. 5
Table 3: Classification accuracies, model parameters for motor imagery classification (ResNet, Motor Movement/Imagery dataset) under different representations. The highest performance among all representations is highlighted in boldface. Representation
Accuracy
Kappa
Params
0.362 0.323 0.437
0.149 0.097 0.250
20.3M 21.5M 21.4M
Raw EEG (Time Domain) STFT (Freqency Domain) Microstates (Ours)
4.3
(b) Cohen’s Kappa
(a) Accuracy
Figure 3: Accuracy (left) and Cohen’s Kappa (right) with different representations under Sleep Transformer and different number of samples.
Motor Imagery Classification
Physical movement or imagination is another important factor of human physical status. Therefore our work involves predicting the movement or imagination activity via microstate sequences.
Table 4: Classification accuracies and model parameters for emotion recognition (CNN-based model, SEED dataset) including full channels (62 in total). The highest performance among all representations is highlighted in boldface.
Dataset. We use the Motor Movement/Imagery Dataset [25, 53] for the task of motor imagery classification. The dataset consists of EEG signals sampled from 109 subjects. Each subject underwent 14 trials involving four tasks and two baseline rest sessions. The tasks are as follows: • Task 1: Open and close the left or right fist. • Task 2: Imagine opening and closing the left or right fist. • Task 3: Open and close both fists or both feet. • Task 4: Imagine opening and closing both fists or feet.
Representation Raw EEG (6 channels) Raw EEG (62 channels) STFT (6 channels) STFT (62 channels) Microstates (Ours)
Accuracy
Kappa
Params
0.846 0.854 0.797 0.854 0.862
0.769 0.778 0.694 0.778 0.793
19.1M 19.2M 19.1M 19.1M 20.1M
Movement/Imagery Labels. In this setting, we focus on the onset of movement/imagery and the labels 𝐿 consists of left hand, right hand, both hands and both feet. Preprocessing. In this dataset, each label corresponds to roughly 4s of EEG signals, and hence we choose 𝑇𝑤 = 4s. Other configurations are the same as the above experiments.
From Table 1, we observe that EEG representation with microstates outperforms time domain and frequency features in three different backbone models, including CNN+LSTM, Sleep Transformer, and Sleep Net Zero. Among the results, microstates achieve the highest accuracy of 0.81 using a sleep transformer or sleep net zero. This indicates that EEG microstates have the potential to serve as a universal representation and outperform temporal- and frequency-domain features across tasks and model structures. Similar observations are obtained in the emotion recognition task based on a CNN-based model and on the motor imagery classification task based on a ResNet model. We also record the standard deviation of the performance of microstate representation on Sleep-Net-Zero, which gives an accuracy of 0.808(±1.897 · 10−3 ) and Kappa 0.733(±2.482 · 10−3 ). We further compare the performance of time- and frequencydomain features with microstates. We see that frequency-domain representation performs well on sleep staging, while producing suboptimal results on other tasks. We suspect that this is because sleep staging is highly frequency-associated, while on other tasks, frequency-domain features may be weak due to information loss [14, 62] in raw EEG signals. On the other hand, raw EEG signals are often subject to noise [67] and does not produce optimal results. Compared to time- and frequency-domain features, microstates present a robust performance across datasets and classification models. This demonstrates that the microstate representation obtained from sleep EEG data can be generalized to various critical tasks and different models, serving as a universal representation.
Model Architecture. The model architecture is based on [13], which consists of a CNN and 3 residual blocks followed by 4 linear layers. We use configurations similar to the above.
5
Results and Analysis
We compared our proposed representation with conventional timeand frequency-domain representations across different tasks and under different model configurations.
5.1
Evaluation Metrics
We evaluated our representation with different model structures and across different tasks. The evaluation metrics are the classification accuracy and Cohen’s Kappa.
5.2
Microstates as a Universal Representation
We compared the accuracy and Cohen’s Kappa on the test set under 3 different models and across three key tasks with different representations. Results on sleep staging are shown in Table 1, and results on emotion recognition and motor imagery classification are shown in Table 2 and Table 3, respectively. 1 For Sleep Net Zero, more parameters are used since the input size of raw EEG signals
is (6, 30000) which is six times that of microstates (30000, 6) and 17 times that of frequency-domain signals (6, 1800) . 6
5.5
Figure 4: Accuracy and Cohen’s Kappa under Sleep Net Zero with different number of microstates.
5.3
Results using Full Channel Data
Due to data constraint, the microstate tokenizer is trained on data from only 6 channels. To allow for a comprehensive comparison, we test the performance of the CNN classifier on the SEED dataset using full channels (62 channels in total). The results are shown in Table 4. From the results we see that using EEG signals from the 6 channels can achieve similar results to that of using full data, as is seen from the raw EEG signals that increasing the number of channels does not significantly boost performance. Notice that the performance of frequency-domain representation increases significantly, which we conjecture that it is because the additional channels compensate for the information loss during the time-frequency transformation. Results show that using only 6 channels does not significantly degrade performance, which justifies our clustering on these channels.
5.4
Interpreting Microstates
In this section, we give an analysis of the interpretability of the microstate representation, which in turn leads to its better performance over other representations. As mentioned in [58], one challenge in analyzing EEG signals is that they are highly subject-dependent and vary significantly across different people. Consequently, it is hard to extract effective intersubject representations using conventional time- or frequencydomain information. Microstates solve this issue by providing a coarse-grained discrete representation that groups similar EEG states together. Its clustering-based nature guarantees its capability to extract universal features while retaining the differences. We analyze the proportion of the most 20 frequently-occurring microstates among groups of 30 subjects under W, N3, and R stage. Results in Figure 5 show that under the same sleep stages, the most frequent microstates are common across all subject groups. For example, the microstates 419, 421, 385, 333 occur with high frequency among all subject groups during W and R stage, whereas the microstates 487, 378, 452, 651 occur with high frequency among all groups under N3 stage. This suggests that the microstate representation captures the similarity between subjects, albeit their having different EEG voltages. Hence this in turn prevents the model from being distracted towards personal specific nuances. Furthermore, the most frequent microstates under W and R stages both contain 419 and 161. The state 419 has all its channels below 2.2𝜇V, denoting a state with a weak EEG signal, while the state 161 has its voltage within the interval 4 ∼ 11𝜇V, which is also relatively low. This is consistent with the fact that during W stage, the EEG signal is dominated by 𝛼 waves, which have a low amplitude and high frequency. Also, during R stage, the brain activity is similar to W stage since this is when dreams take place [21]. Consequently, it does not come as a surprise that W and R stages share many microstates in common, indicating a similar brain activity pattern. However, the microstate 378 denotes signals within the interval 10 ∼ 24𝜇V, and 452 has signals within −5 ∼ −21𝜇V. Both of them are relatively strong brain activity. This is again consistent with the fact that during N3 the EEG signal has a larger portion of 𝛿 waves with a larger amplitude [77]. This suggests that microstates are capable of extracting the similarities between sleep stages, while also retaining their differences.
Microstates as a Scalable Representation
We further test the performance of EEG representation learning with microstates across different scales of training data in the sleep staging task under Sleep Transformer on the HSP dataset. We also find that microstates offer a more scalable representation. As shown in Figure 3, microstates do not exhibit strengthened performance when the size of the training data is smaller than 2,000. However, when the number of samples increases, the performance of the microstate representation shows a more pronounced performance gain in comparison to other features. Our experiment demonstrates that the microstate representation is also capable of scaling across the size of the training data. This reveals the potential of EEG representation with microstates, especially using deep learning methods and increased data size. On the other side, we tested the performance of our tokenizer under different number of clusters by selecting different parameters 𝑘 for clustering. We evaluated the model performance on the validation set with different number of microstates. The results are shown in Table 4. From the results we see that the performance increases while the number of microstates increases. Results show that the performance of the microstate representation also scales with increasing number of microstates.
6
Discussions and Conclusion
In this work, we introduce EEG microstates as a clinically grounded approach for integrating deep learning and EEG signal analysis. Our approach improves the representation of brain activity by aligning more closely with the underlying neural mechanisms and cognitive activities, enhancing both clinical and research applications. Experimental results demonstrate the effectiveness of EEG microstates in three critical tasks—sleep staging, emotion recognition, and motor imagery classification and across different models, where it outperforms traditional time-domain and frequencydomain methods. Furthermore, we show that EEG microstates present more performance gain than time- and frequency-domain features when scaling the data size, indicating that EEG microstates can alleviate 7
Figure 5: Visualizing Microstates Distribution. Visualization of the distribution of different microstates among different subjects undergoing different sleep stages. We can see that the microstate representation simutaneously retain the similarity between subjects and between W and R stage, while also preserves the difference between W stage N3 stage. • We only focused on the representation side. Based on the microstate representation, we hypothesize that it is possible to develop a pre-trained model that can generalize to several EEG-related downstream tasks. Above all, we believe that combining deep learning techniques with biologically grounded EEG microstates opens up a portal to future research on improving the accuracy of EEG analysis across different tasks and on uncovering more correlations between microstates and brain activity. Future work might involve reconstructing brain signals with more channels to alleviate the lack of channel data.
the burden of data scarcity and pave the way to more scalable settings. We also show that EEG microstates can provide interpretable insights for EEG analysis and deep learning, offering a promising direction for future research and clinical practice. The adoption of EEG microstates holds significant potential for advancing both cognitive neuroscience and the field of clinical diagnostics. Several limitations guide future work, such as: • We only experimented with limited tasks and limited number of datasets. Particularly, the training of the tokenizer was only performed on sleep data. This is reasonable because the HSP dataset is the largest, but more research can be conducted across tasks in the future.
8
References
[21] Abdeljalil El Hadiri, Lhoussain Bahatti, Abdelmounime El Magri, and Rachid Lajouad. 2024. Sleep stages detection based on analysis and optimisation of non-linear brain signal parameters. Results in Engineering 23 (2024), 102664. doi:10.1016/j.rineng.2024.102664 [22] MohammadReza EskandariNasab, Zahra Raeisi, Reza Ahmadi Lashaki, and Hamidreza Najafi. 2024. A GRU–CNN model for auditory attention detection using microstate and recurrence quantification analysis. Scientific Reports 14 (2024). https://api.semanticscholar.org/CorpusID:269211640 [23] Shenzhi Fang, Chaofeng Zhu, Jinying Zhang, Luyan Wu, Yuying Zhang, Huapin Huang, and Wanhui Lin. 2024. EEG microstates in epilepsy with and without cognitive dysfunction: Alteration in intrinsic brain activity. Epilepsy & Behavior 154 (2024), 109729. doi:10.1016/j.yebeh.2024.109729 [24] Linda Fiorini, Francesco Bossi, and Francesco Di Gruttola. 2024. EEG-based emotional valence and emotion regulation classification: a data-centric and explainable approach. Scientific reports 14, 1 (October 2024), 24046. doi:10.1038/ s41598-024-75263-x [25] A. Goldberger, L. Amaral, L. Glass, J. Hausdorff, P. C. Ivanov, R. Mark, J. E. Mietus, G. B. Moody, Peng C. K., and H. E. Stanley. 2000. PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals. Circulation 101, 23 (2000), e215–e220. Online. [26] Xiaotong Gu, Zehong Cao, Alireza Jolfaei, Peng Xu, Dongrui Wu, Tzyy-Ping Jung, and Chin-Teng Lin. 2021. EEG-Based Brain-Computer Interfaces (BCIs): A Survey of Recent Studies on Signal Sensing Technologies and Computational Intelligence Approaches and Their Applications. IEEE/ACM Transactions on Computational Biology and Bioinformatics 18, 5 (2021), 1645–1666. doi:10.1109/ TCBB.2021.3052811 [27] Haokun Gui, Xiucheng Li, and Xinyang Chen. 2024. Vector Quantization Pretraining for EEG Time Series with Random Projection and Phase Alignment. In Proceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235), Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (Eds.). PMLR, 16731–16750. https://proceedings.mlr.press/v235/gui24a.html [28] Abir Hadriche, Laurent Pezard, Jean-Louis Nandrino, Hamadi Ghariani, Abdennaceur Kachouri, and Viktor K. Jirsa. 2013. Mapping the dynamic repertoire of the resting brain. NeuroImage 78 (2013), 448–462. doi:10.1016/j.neuroimage.2013. 04.041 [29] Jingzhao Hu, Chen Wang, Qiaomei Jia, Qirong Bu, Richard Sutcliffe, and Jun Feng. 2021. ScalingNet: Extracting features from raw EEG data for emotion recognition. Neurocomputing 463 (2021), 177–184. doi:10.1016/j.neucom.2021.08.018 [30] Sunhee Hwang, Kibeom Hong, Guiyoung Son, and Hyeran Byun. 2020. Learning CNN features from DE features for EEG-based emotion recognition. Pattern Analysis and Applications 23, 3 (2020), 1323 – 1335. doi:10.1007/s10044-01900860-w Cited by: 104. [31] Abhishek Iyer, Srimit Sritik Das, Reva Teotia, Shishir Maheshwari, and Rishi Sharma. 2022. CNN and LSTM based Ensemble Learning for Human Emotion Recognition using EEG Recordings. Multimedia Tools and Applications (04 2022). doi:10.1007/s11042-022-12310-7 [32] Smith K. Khare, Victoria Blanes-Vidal, Esmaeil S. Nadimi, and U. Rajendra Acharya. 2024. Emotion recognition and artificial intelligence: A systematic review (2014–2023) and research recommendations. Information Fusion 102 (2024), 102019. doi:10.1016/j.inffus.2023.102019 [33] Domantė Kučikienė, Ravichandran Rajkumar, Katharina Timpte, Jan Heckelmann, Irene Neuner, Yvonne Weber, and Stefan Wolking. 2024. EEG microstates show different features in focal epilepsy and psychogenic nonepileptic seizures. Epilepsia 65 (01 2024). doi:10.1111/epi.17897 [34] Byeong-Hoo Lee, Ji-Hoon Jeong, Kyung-Hwan Shim, and Dong-Joo Kim. 2020. Motor Imagery Classification of Single-Arm Tasks Using Convolutional Neural Network based on Feature Refining. arXiv:2002.01122 [cs.HC] https://arxiv.org/ abs/2002.01122 [35] D. Lehmann, H. Ozaki, and I. Pal. 1987. EEG alpha map series: brain micro-states by space-oriented adaptive segmentation. Electroencephalography and Clinical Neurophysiology 67, 3 (1987), 271–288. doi:10.1016/0013-4694(87)90025-3 [36] D Lehmann, W.K Strik, B Henggeler, T Koenig, and M Koukkou. 1998. Brain electric microstates and momentary conscious mind states as building blocks of spontaneous thinking: I. Visual imagery and abstract thoughts. International Journal of Psychophysiology 29, 1 (1998), 1–11. doi:10.1016/S0167-8760(97)000986 [37] Shuzhen Li, Yuxin Chen, Xuesong Chen, Ruiyang Gao, Yupeng Zhang, Chao Yu, Yunfei Li, Ziyi Ye, Weijun Huang, Hongliang Yi, et al. 2024. SleepNetZero: Zero-Burden Zero-Shot Reliable Sleep Staging with Neural Networks Based on Ballistocardiograms. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 4 (2024), 1–25. [38] Robert Lin, Ren-Guey Lee, Chwan-Lu Tseng, Heng-Kuan Zhou, C. F. Chao, and Joe-Air Jiang. 2006. A NEW APPROACH FOR IDENTIFYING SLEEP APNEA SYNDROME USING WAVELET TRANSFORM AND NEURAL NETWORKS. Biomedical Engineering: Applications, Basis and Communications 18 (2006), 138– 143. https://api.semanticscholar.org/CorpusID:2412588 [39] Hong Liu, Haoling Tang, Wei Wei, Gesheng Wang, Yong Du, and Jianghai Ruan. 2021. Altered peri-seizure EEG microstate dynamics in patients with absence
[1] Arshia Afzal, Grigorios Chrysos, Volkan Cevher, and Mahsa Shoaran. 2024. Rest: Efficient and accelerated eeg seizure analysis through residual state updates. arXiv preprint arXiv:2406.16906 (2024). [2] Charu C. Aggarwal, Philip S. Yu, Jiawei Han, and Jianyong Wang. 2003. - A Framework for Clustering Evolving Data Streams. In Proceedings 2003 VLDB Conference, Johann-Christoph Freytag, Peter Lockemann, Serge Abiteboul, Michael Carey, Patricia Selinger, and Andreas Heuer (Eds.). Morgan Kaufmann, San Francisco, 81–92. doi:10.1016/B978-012722442-8/50016-1 [3] David Ahmedt-Aristizabal, Tharindu Fernando, Simon Denman, Lars Petersson, Matthew J. Aburn, and Clinton Fookes. 2019. Neural Memory Networks for Seizure Type Classification. 2020 42nd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC) (2019), 569–575. https: //api.semanticscholar.org/CorpusID:210966442 [4] Aydin Akan and Ozlem Karabiber Cura. 2021. Time–frequency signal processing: Today and future. Digital Signal Processing 119 (2021), 103216. doi:10.1016/j.dsp. 2021.103216 [5] Irfan Al-Hussaini, Cao Xiao, M. Brandon Westover, and Jimeng Sun. 2019. SLEEPER: interpretable Sleep staging via Prototypes from Expert Rules. arXiv:1910.06100 [cs.LG] https://arxiv.org/abs/1910.06100 [6] Ghita Amrani, Amina Adadi, Mohammed Berrada, Zouhayr Souirti, and Saïd Boujraf. 2021. EEG signal analysis using deep learning: A systematic literature review. In 2021 Fifth International Conference On Intelligent Computing in Data Sciences (ICDS). 1–8. doi:10.1109/ICDS53782.2021.9626707 [7] Kleanthis Avramidis. 2021. Affective Analysis and Interpretation of Brain Responses to Music Stimuli. [8] Anahit Babayan, Miray Erbey, Deniz Kumral, Janis Reinelt, Andrea Reiter, Josefin Röbbig, H. Schaare, Marie Uhlig, Alfred Anwander, Pierre-Louis Bazin, Annette Horstmann, Leonie Lampe, Vadim Nikulin, Hadas Okon-Singer, Sven Preusser, André Pampel, Christiane Rohr, Julia Sacher, Angelika Thoene-Otto, and Arno Villringer. 2019. A mind-brain-body dataset of MRI, EEG, cognition, emotion, and peripheral physiology in young and old adults. Scientific Data 6 (02 2019), 180308. doi:10.1038/sdata.2018.308 [9] Dongmei Bai, Tianshuang Qiu, and Xiaobing Li. 2007. [The sample entropy and its application in EEG based epilepsy detection]. Sheng wu yi xue gong cheng xue za zhi = Journal of biomedical engineering = Shengwu yixue gongchengxue zazhi 24 1 (2007), 200–5. https://api.semanticscholar.org/CorpusID:21621951 [10] Juliane Britz, Dimitri Van De Ville, and Christoph M. Michel. 2010. BOLD correlates of EEG topography reveal rapid resting-state network dynamics. NeuroImage 52, 4 (2010), 1162–1170. doi:10.1016/j.neuroimage.2010.02.052 [11] Verena Brodbeck, Alena Kuhn, Frederic von Wegner, Astrid Morzelewski, Enzo Tagliazucchi, Sergey Borisov, Christoph M. Michel, and Helmut Laufs. 2012. EEG microstates of wakefulness and NREM sleep. NeuroImage 62, 3 (2012), 2129–2139. doi:10.1016/j.neuroimage.2012.05.060 [12] Zheng Chen, Ziwei Yang, Lingwei Zhu, Wei Chen, Toshiyo Tamura, Naoaki Ono, Md Altaf-Ul-Amin, Shigehiko Kanaya, and Ming Huang. 2023. Automated Sleep Staging via Parallel Frequency-Cut Attention. IEEE Transactions on Neural Systems and Rehabilitation Engineering 31 (2023), 1974–1985. doi:10.1109/TNSRE. 2023.3243589 [13] Joseph Y. Cheng, Hanlin Goh, Kaan Dogrusoz, Oncel Tuzel, and Erdrin Azemi. 2020. Subject-Aware Contrastive Learning for Biosignals. ArXiv abs/2007.04871 (2020). https://api.semanticscholar.org/CorpusID:220425132 [14] Zhuoling Cheng, Xuekui Bu, Qingnan Wang, Tao Yang, and Jihui Tu. 2024. EEG-based emotion recognition using multi-scale dynamic CNN and gated transformer. Scientific Reports 14 (2024). https://api.semanticscholar.org/CorpusID: 275117639 [15] Ming Chu and Jingfeng Bi. 2023. Six classes of motor imagery EEG signals in the upper limb. doi:10.21227/8qw6-f578 [16] Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. arXiv:1412.3555 [cs.NE] https://arxiv.org/abs/1412.3555 [17] Jiaxiang Dong, Haixu Wu, Haoran Zhang, Li Zhang, Jianmin Wang, and Mingsheng Long. 2023. SimMTM: A Simple Pre-Training Framework for Masked TimeSeries Modeling. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associates, Inc., 29996–30025. https://proceedings.neurips.cc/paper_ files/paper/2023/file/5f9bfdfe3685e4ccdbc0e7fb29cccf2a-Paper-Conference.pdf [18] Xiaobing Du, Cuixia Ma, Guanhua Zhang, Jinyao Li, Yu-Kun Lai, Guozhen Zhao, Xiaoming Deng, Yong-Jin Liu, and Hongan Wang. 2022. An Efficient LSTM Network for Emotion Recognition From Multichannel EEG Signals. IEEE Transactions on Affective Computing 13, 3 (2022), 1528–1540. doi:10.1109/TAFFC. 2020.3013711 [19] Ruo-Nan Duan, Jia-Yi Zhu, and Bao-Liang Lu. 2013. Differential entropy feature for EEG-based emotion classification. In 6th International IEEE/EMBS Conference on Neural Engineering (NER). IEEE, 81–84. [20] Robert Efron. 1970. The minimum duration of a perception. Neuropsychologia 8, 1 (1970), 57–63. doi:10.1016/0028-3932(70)90025-4 9
epilepsy. Seizure 88 (2021), 15–21. https://api.semanticscholar.org/CorpusID: 232359456 [40] Huisheng Lu, Mingshi Wang, and Hongqiang Yu. 2005. EEG Model and Location in Brain when Enjoying Music. 2005 IEEE Engineering in Medicine and Biology 27th Annual Conference (2005), 2695–2698. https://api.semanticscholar.org/CorpusID: 21783113 [41] Marzia Lucia, Christoph Michel, Stephanie Clarke, and Micah Murray. 2007. Single-subject EEG analysis based on topographic information. International Journal of Bioelectromagnetism www.ijbem.org 9 (01 2007), 168–171. [42] Scott Makeig, Stefan Debener, Julie Onton, and Arnaud Delorme. 2004. Mining event-related brain dynamics. Trends in Cognitive Sciences 8, 5 (2004), 204–210. doi:10.1016/j.tics.2004.03.008 [43] Christoph M. Michel and Thomas Koenig. 2018. EEG microstates as a tool for studying the temporal dynamics of whole-brain neuronal networks: A review. NeuroImage 180 (2018), 577–593. doi:10.1016/j.neuroimage.2017.11.062 Brain Connectivity Dynamics. [44] P. Milz, P.L. Faber, D. Lehmann, T. Koenig, K. Kochi, and R.D. Pascual-Marqui. 2016. The functional significance of EEG microstates—Associations with modalities of thinking. NeuroImage 125 (2016), 643–656. doi:10.1016/j.neuroimage.2015. 08.023 [45] Micah Murray, Denis Brunet, and Christoph Michel. 2008. Topographic ERP Analyses: A Step-by-Step Tutorial Review. Brain topography 20 (07 2008), 249–64. doi:10.1007/s10548-008-0054-5 [46] Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. 2022. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. ArXiv abs/2211.14730 (2022). https://api.semanticscholar.org/CorpusID:254044221 [47] R.D. Pascual-Marqui, C.M. Michel, and D. Lehmann. 1995. Segmentation of brain electrical activity into microstates: model estimation and validation. IEEE Transactions on Biomedical Engineering 42, 7 (1995), 658–665. doi:10.1109/10. 391164 [48] Mathias Perslev, S. Darkner, Lykke Kempfner, Miki Nikolic, Poul Jørgen Jennum, and C. Igel. 2021. U-Sleep: resilient high-frequency sleep staging. NPJ Digital Medicine 4 (2021). https://api.semanticscholar.org/CorpusID:233237814 [49] Huy Phan, Kaare Mikkelsen, Oliver Y. Chén, Philipp Koch, Alfred Mertins, and Maarten De Vos. 2022. SleepTransformer: Automatic Sleep Staging With Interpretability and Uncertainty Quantification. IEEE Transactions on Biomedical Engineering 69, 8 (2022), 2456–2467. doi:10.1109/TBME.2022.3147187 [50] Gilles Pourtois, Sylvain Delplanque, Christoph M. Michel, and Patrik Vuilleumier. 2008. Beyond Conventional Event-related Brain Potential (ERP): Exploring the Time-course of Visual Emotion Processing Using Topographic and Principal Component Analyses. Brain Topography 20 (2008), 265–277. https://api.semanticscholar.org/CorpusID:15084282 [51] Giulia Prete, Pierpaolo Croce, Filippo Zappasodi, Luca Tommasi, and Paolo Capotosto. 2022. Exploring brain activity for positive and negative emotions by means of EEG microstates. Scientific Reports 12 (03 2022), 1–11. doi:10.1038/ s41598-022-07403-0 [52] Asha S.A, Sudalaimani C, Devanand P, Alexander G, Arya Maniyan Lathikakumari, Sanjeev V Thomas, and Ramshekhar N Menon. 2024. Analysis of EEG microstates as biomarkers in neuropsychological processes – Review. Computers in Biology and Medicine 173 (2024), 108266. doi:10.1016/j.compbiomed.2024.108266 [53] G. Schalk, D.J. McFarland, T. Hinterberger, N. Birbaumer, and J.R. Wolpaw. 2004. BCI2000: a general-purpose brain-computer interface (BCI) system. IEEE Transactions on Biomedical Engineering 51, 6 (2004), 1034–1043. doi:10.1109/TBME. 2004.827072 [54] Bastian Schiller, Matthias Sperl, Tobias Kleinert, Kyle Nash, and Lorena Gianotti. 2023. EEG Microstates in Social and Affective Neuroscience. Brain Topography 37 (07 2023), 1–17. doi:10.1007/s10548-023-00987-4 [55] Felix Schlegel, D. Lehmann, Pascal Faber, Patricia Milz, and Lorena Gianotti. 2011. EEG Microstates During Resting Represent Personality Differences. Brain topography 25 (06 2011), 20–6. doi:10.1007/s10548-011-0189-7 [56] Benjamin A. Seitzman, Malene Abell, Samuel C. Bartley, Molly A. Erickson, Amanda R. Bolbecker, and William P. Hetrick. 2017. Cognitive manipulation of brain electric microstates. NeuroImage 146 (2017), 533–543. doi:10.1016/j. neuroimage.2016.10.002 [57] Xiaorui Shao and Chang Soo Kim. 2022. A Hybrid Deep Learning Scheme for Multi-Channel Sleep Stage Classification. Computers, Materials & Continua (2022). https://api.semanticscholar.org/CorpusID:243464227 [58] Xinke Shen, Xianggen Liu, Xin Hu, Dan Zhang, and Sen Song. 2023. Contrastive Learning of Subject-Invariant EEG Representations for Cross-Subject Emotion Recognition. IEEE Transactions on Affective Computing 14, 3 (2023), 2496–2511. doi:10.1109/TAFFC.2022.3164516 [59] In-Ho Song, Doo-Soo Lee, and Sun I Kim. 2004. Recurrence quantification analysis of sleep electoencephalogram in sleep apnea syndrome in humans. Neuroscience Letters 366, 2 (2004), 148–153. doi:10.1016/j.neulet.2004.05.025 [60] Tengfei Song, Suyuan Liu, Wenming Zheng, Yuan Zong, Zhen Cui, Yang Li, and Xiaoyan Zhou. 2021. Variational Instance-Adaptive Graph for EEG Emotion Recognition. IEEE Transactions on Affective Computing 14 (2021), 343–356. https: //api.semanticscholar.org/CorpusID:233621668
[61] Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2023. RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv:2104.09864 [cs.CL] https://arxiv.org/abs/2104.09864 [62] D. Puthankattil Subha, Paul K. Joseph, Rajendra Acharya U., and Choo Min Lim. 2010. EEG Signal Analysis: A Survey. Journal of Medical Systems 34 (2010), 195–212. https://api.semanticscholar.org/CorpusID:1140473 [63] Luke Tait, Francesco Tamagnini, George Stothart, Edoardo Barvas, Chiara Monaldini, Roberto P Frusciante, Mirco Volpini, Susanna Guttmann, Elizabeth J. Coulthard, Jon T. Brown, Nina Kazanina, and Marc Goodfellow. 2019. EEG microstate complexity for aiding early diagnosis of Alzheimer’s disease. Scientific Reports 10 (2019). https://api.semanticscholar.org/CorpusID:209578713 [64] Padhmashree V. and Abhijit Bhattacharyya. 2022. Human emotion recognition based on time–frequency analysis of multivariate EEG signal. Knowledge-Based Systems 238 (2022), 107867. doi:10.1016/j.knosys.2021.107867 Time-frequency analysis of [65] V. Vanitha and P. Krishnan. 2017. EEG for improved classification of emotion. International Journal of Biomedical Engineering and Technology 23, 2-4 (2017), 191–212. arXiv:https://www.inderscienceonline.com/doi/pdf/10.1504/IJBET.2017.082661 doi:10.1504/IJBET.2017.082661 [66] Anna Elisabetta Vaudano, Nicoletta Azzi, and Irene Trippi. 2019. Normal Sleep EEG. Springer International Publishing, Cham, 153–175. doi:10.1007/978-3-03004573-9_10 [67] Neeraj Wagh, Jionghao Wei, Samarth Rawal, Brent M. Berry, and Yogatheesan Varatharajah. 2022. Evaluating Latent Space Robustness and Uncertainty of EEG-ML Models under Realistic Distribution Shifts. arXiv:2209.11233 [eess.SP] https://arxiv.org/abs/2209.11233 [68] Fei Wang, Shichao Wu, Weiwei Zhang, Zongfeng Xu, Yahui Zhang, Chengdong Wu, and Sonya Coleman. 2020. Emotion recognition with convolutional neural network and EEG-based EFDMs. Neuropsychologia 146 (2020), 107506. doi:10. 1016/j.neuropsychologia.2020.107506 [69] Guangyu Wang, Wenchao Liu, Yuhong He, Cong Xu, Lin Ma, and Haifeng Li. 2024. EEGPT: Pretrained Transformer for Universal and Reliable Representation of EEG Signals. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. [70] Jiquan Wang, Sha Zhao, Haiteng Jiang, Shijian Li, Tao Li, and Gang Pan. 2024. Generalizable Sleep Staging via Multi-Level Domain Alignment. arXiv:2401.05363 [eess.SP] https://arxiv.org/abs/2401.05363 [71] Xiao-Wei Wang, Dan Nie, and Bao-Liang Lu. 2011. EEG-Based Emotion Recognition Using Frequency Domain Features and Support Vector Machines. In International Conference on Neural Information Processing. https://api.semanticscholar. org/CorpusID:9355572 [72] Yiming Wang, Bin Zhang, and Yujiao Tang. 2024. DMMR: Cross-Subject Domain Generalization for EEG-Based Emotion Recognition via Denoising Mixed Mutual Reconstruction. In AAAI Conference on Artificial Intelligence. https://api.semanticscholar.org/CorpusID:268678230 [73] M. B. Westover, V. Moura Junior, R. Thomas, S. Cash, S. Nasiri, H. Sun, A. Gupta, J. Rosand, M. Ghanta, W. Ganglberger, U. Katwa, K. Stone, Z. Zhang, G. Ganjoo, T. E. Nassi PhD Candidate, R. Wei, D. Hwang, L. M. Trotti, A. Parekh, E. Meulenbrugge, E. Mignot, R Au, G. Clifford, and D. Rapoport. 2023. The Human Sleep Project (version 2.0). Brain Data Science Platform. https://doi.org/10.60508/qjbv-hg78. [74] Haixu Wu, Teng Hu, Yong Liu, Hang Zhou, Jianmin Wang, and Mingsheng Long. 2022. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. ArXiv abs/2210.02186 (2022). https://api.semanticscholar.org/CorpusID: 252715491 [75] Guowen Xiao, Mengwen Ye, Bowen Xu, Zhendi Chen, and Quansheng Ren. 2021. 4D Attention-based Neural Network for EEG Emotion Recognition. arXiv:2101.05484 [cs.LG] https://arxiv.org/abs/2101.05484 [76] Chen Zechuan and Zhang Kan. 2024. The Principles of Psychology. Springer Nature Singapore, Singapore, 1–2. doi:10.1007/978-981-99-6000-2_1042-1 [77] Xiaoli Zhang, Xizhen Zhang, Qiong Huang, Yang Lv, and Fuming Chen. 2024. A review of automated sleep stage based on EEG signals. Biocybernetics and Biomedical Engineering 44, 3 (2024), 651–673. doi:10.1016/j.bbe.2024.06.004 [78] Jingyi Zheng, Mingli Liang, Sujata Sinha, Linqiang Ge, Wei Yu, Arne Ekstrom, and Fushing Hsieh. 2022. Time-Frequency Analysis of Scalp EEG With HilbertHuang Transform and Deep Learning. IEEE Journal of Biomedical and Health Informatics 26, 4 (2022), 1549–1559. doi:10.1109/JBHI.2021.3110267 [79] Wei-Long Zheng and Bao-Liang Lu. 2015. Investigating Critical Frequency Bands and Channels for EEG-based Emotion Recognition with Deep Neural Networks. IEEE Transactions on Autonomous Mental Development 7, 3 (2015), 162–175. doi:10.1109/TAMD.2015.2431497 [80] Xinliang Zhou, Chenyu Liu, Zhongruo Wang, Liming Zhai, Ziyu Jia, Cuntai Guan, and Yang Liu. 2024. Interpretable and Robust AI in EEG Systems: A Survey. arXiv:2304.10755 [eess.SP] https://arxiv.org/abs/2304.10755 [81] Ning Zhuang, Ying Zeng, li Tong, Chi Zhang, Hanming Zhang, and Bin Yan. 2017. Emotion Recognition from EEG Signals Using Multidimensional Information in EMD Domain. BioMed Research International 2017 (08 2017), 1–9. doi:10.1155/ 2017/8317357 10
Acknowledgments This work is supported by the Ministry of Science and Technology of China STI2030-Major Projects (No. 2021ZD0201900, 2021ZD0201902). The computations in this research were performed using the CFFF platform of Fudan University.
at different frequencies. [4]. For a given frequency 𝑓 , the power is computed as follows: ∫ ∞ 𝑋 (𝑡, 𝑓 ) = 𝑤 (𝑡 − 𝜏)𝑠 (𝜏)𝑒 −𝑖2𝜋 𝑓 𝜏 𝑑𝜏 −∞
1 |𝑋 (𝑡, 𝑓 )| 2 2𝜋 where 𝑤 (𝑡) is a window function. Note that ∫ ∞ ′ 𝑋 (𝑡, 𝑓 ) = 𝑠 (𝜏)𝑒 −𝑖2𝜋 𝑓 𝜏 𝑑𝜏 𝑝 (𝑡, 𝑓 ) =
A
Experimental Setup
This section gives the detailed experimental configuration of our downstream tasks.
−∞
A.1
is the usual Fourier transform, and the window function 𝑤 (𝑡) only has finite support which serves as a short-time weighted sum of the integral. We use the Hann window function defined as follows: 1 2𝜋𝑡 𝑇 2 (1 − cos 𝑇 ) |𝑡 | ≤ 2 𝑤 (𝑡) = 𝑇 0 |𝑡 | > 2 where 𝑇 is the window length. In our setting we set 𝑇 = 𝑡 𝑤 = 1s. The overlap ratio is set to be 𝑟𝑜 = 0. Using the above approach, the processed frequency-domain signals have length 𝑙 ′ where 𝑓𝑟𝑒𝑠𝑇 − 𝑓𝑟𝑒𝑠 𝑡 𝑤 ′ 𝑙 = +1 (1 − 𝑟𝑜 )𝑓𝑟𝑒𝑠 𝑡 𝑤 is the number of windows. Leaving out the margin, we have that 𝑇 𝑙′ = = 𝑓 𝑓 𝑟𝑒𝑞𝑇 (1 − 𝑟𝑜 )𝑡 𝑤 here for 𝑟𝑜 = 0 and 𝑡 𝑤 = 1s, we have 𝑓 𝑓 𝑟𝑒𝑞 = 1Hz. Having calculated the power 𝑝 (𝑡, 𝑓 ) at frequency 𝑓 and time 𝑡, we obtain the spectrogram 𝑃 ∈ R𝐹 ×𝑓 𝑓 𝑟𝑒𝑞𝑇 for each channel, where 𝐹 is the frequency axis and 𝑓 𝑓 𝑟𝑒𝑞𝑇 is the time axis. Next, we apply band integration. Human EEG signal is divided into the following frequency bands: • 𝛿-band: 0.5 ∼ 4Hz • 𝜃 -band: 4 ∼ 8Hz • 𝛼-band: 8 ∼ 12Hz • 𝜎-band: 12 ∼ 16Hz • 𝛽-band: 16 ∼ 30Hz • 𝛾-band: 30 ∼ 40Hz and we combine the powers among 𝐹 within each band. We use simpson integration as our numerical quadrature method, which is defined as ∫ 𝑏 𝑏 −𝑎 𝑎 +𝑏 𝑓 (𝑥)𝑑𝑥 = 𝑓 (𝑎) + 4𝑓 + 𝑓 (𝑏) 6 2 𝑎
Sleep Staging
A.1.1 Dataset. The dataset used is the Human Sleep Project (HSP) dataset [73]. This dataset includes PSG signals from over 20K subjects. Signals are sampled under various frequencies including 256Hz and 512Hz. Each sample includes a night’s sleep of a subject sampled, with sleep stages annotated every 30 seconds. Equivalently, we have the raw EEG signals 𝑠 ∈ R𝐶 ×𝑓𝑠 𝑇 where 𝐶 consists of different classes of channels such as EOG, ECG, and EEG and differs across subjects, 𝑓𝑠 denotes the sample frequency which also differs across subjects, and 𝑇 is the time duration of a night’s sleep, 1 which is typically 6-7 hours. The label frequency is 𝑓𝑙 = 30 Hz. A.1.2
Preprocessing.
Extracting target channels. As for constructing the microstates, we have to filter out a fixed number of channels. To achieve this goal, we extract 𝑁 = 6 channels which is common among all samples. The EEG leads are shown in the following diagram [19, 79]. The channels chosen are F3, F4, C3, C4, O1, O2. After extraction, the EEG signals have shape 𝑠𝑒𝑥𝑡 ∈ R𝑁 ×𝑓𝑠 𝑇 .
Figure 6: The Distribution of EEG Leads Filtering and Resampling. The raw EEG signals are then bandpass filtered between 1Hz and 40Hz, followed by a resampling at 𝑓𝑟𝑒𝑠 = 100Hz. After these procedures, the raw EEG signals now have shape 𝑠𝑟𝑒𝑠 ∈ R𝑁 ×𝑓𝑟𝑒𝑠 𝑇 . Having obtained the resampled data, we construct the representations accordingly.
for a step interval [𝑎, 𝑏]. After the integration, the array shape becomes 𝑠 𝑓 𝑟𝑒𝑞,𝑠𝑖𝑛 ∈ R𝐵×𝑓 𝑓 𝑟𝑒𝑞𝑇 where 𝐵 = 6 is the number of bands. As our final step, we flatten the array for each channel to 𝑠 𝑓 𝑟𝑒𝑞,𝑠𝑖𝑛,𝑓 𝑙𝑎𝑡 ∈ R𝐵𝑓 𝑓 𝑟𝑒𝑞𝑇 and stack the 𝑁 channels together, resulting in shape 𝑠 𝑓 𝑟𝑒𝑞 ∈ R𝑁 ×𝐵𝑓 𝑓 𝑟𝑒𝑞𝑇 .
Constructing microstates. The fitted clustering model is applied to the resampled data 𝑠𝑟𝑒𝑠 ∈ R𝑁 ×𝑓𝑠 𝑇 . The result is a microstate sequence 𝑐 ∈ 𝑆 𝑓𝑟𝑒𝑠 𝑇 where 𝑆 ∈ {𝑏 1, 𝑏 2, . . . , 𝑏𝑘 } is a set of 𝑘 discrete states. Here we let 𝑘 = 1000.
Slicing. We select fixed window size 𝑇𝑤 = 300s. In this case, a microstates sample will have shape 𝑐 𝑤 ∈ 𝑆 𝑓𝑟𝑒𝑠 𝑇𝑤 = 𝑆 30000 , and the raw EEG data will have shape 𝑠𝑡𝑖𝑚𝑒,𝑤 ∈ R𝑁 ×𝑓𝑟𝑒𝑠 𝑇𝑤 = R6×30000 . The frequency-domain representation will have shape 𝑠 𝑓 𝑟𝑒𝑞,𝑤 ∈ R𝑁 ×𝐵𝑓 𝑓 𝑟𝑒𝑞𝑇𝑤 = R6×1800 .
Constructing baseline. We directly use the raw EEG signals for time-domain features. The input shape is 𝑠𝑡𝑖𝑚𝑒 ∈ R𝑁 ×𝑓𝑟𝑒𝑠 𝑇 . To extract frequency information, we use short-time Fourier transform. The major goal is to decompose the signal into powers 11
1 A.1.3 Labels. The label frequency is 𝑓𝑙 = 30 Hz, and hence the 𝑓 𝑇 𝑤 𝑙 label sequence will be of shape 𝑙 𝑤 ∈ 𝐿 = 𝐿 10 where 𝐿 consists of the five sleep stages.
A.2
Emotion Recognition
A.2.1 Dataset. The dataset used is the SEED dataset [19, 79]. The SEED dataset consists of 15 subjects whose EEG signals of 62 channels are recorded when watching movie clips expressing different emotions which are categorized as positive, neutral and negative. There are a total number of 15 trials, during which subjects view episodes with positive, neutral, negative, negative, nuetral, positive, negative, neutral, positive, positive, neutral, negative, neutral, positive, negative emotions. The dataset is filtered between 0 and 75Hz and downsampled to 200Hz. In raw EEG samples have shape 𝑠 ∈ R𝐶 ×𝑓𝑠 𝑇 where 𝑓𝑠 = 200Hz and 𝐶 = 62. 𝑇 is the length of the movie clip which varies between trials. A.2.2
Figure 7: The Valence-Arousal Space
A.3
Motor Imagery Classification
A.3.1 Dataset. We use the Motor Movement/Imagery dataset [25, 53]. The dataset consists of 109 subjects undergoing 14 trials. The 14 trials includes two rest sessions and four tasks. The four tasks are: • Task 1: Open and close the left or right fist. • Task 2: Imagine opening and closing the left or right fist. • Task 3: Open and close both fists or both feet. • Task 4: Imagine opening and closing both fists or feet. Every subject went through two rest sessions and three rounds of successive tasks in the order above. The labels are given during movement roughly every four seconds. There are in total three labels. 𝑇0 corresponds to rest, 𝑇1 corresponds to the onset of moving or imagining moving the left or both fists, and 𝑇2 corresponds to the onset of moving or imagining moving the right fist or both feet. The samples contain 64 channels at 𝑓𝑠 = 160Hz. In this case, the raw EEG signals have shape 𝑠 ∈ R𝐶 ×𝑓𝑠 𝑇 where 𝐶 = 64, 𝑓𝑠 = 160Hz and 𝑇 is the duration of each trial.
Preprocessing.
Extracting target channels. We extract the 6 target channels as above for labeling. The resulting shape is 𝑠𝑒𝑥𝑡 ∈ R𝑁 ×𝑓𝑠 𝑇 where 𝑁 = 6. Constructing microstates and baseline. We do not filter and resample the EEG signals since these are done initially. Applying the clustering model, we obtain the microstate sequence 𝑐 ∈ 𝑆 𝑓𝑠 𝑇 where 𝑆 is the set of 1000 states. The raw signal has shape 𝑠𝑡𝑖𝑚𝑒 ∈ R𝑁 ×𝑓𝑠 𝑇 , and the frequency-domain signal has shape 𝑠 𝑓 𝑟𝑒𝑞 ∈ R𝑁 ×𝐵𝑓 𝑓 𝑟𝑒𝑞𝑇 where 𝑓 𝑓 𝑟𝑒𝑞 = 1Hz and 𝐵 = 6 is the number of bands. Windowing. Since the movie clips are not of the same length, we set 𝑇𝑤 = 265s which is the duration of the longest video, and pad the signals that are shorter. For microstates, a new token is introduced for padding, whereas for the other two representations, we pad zeros. Now the microstate sequence has length 𝑓𝑠𝑇𝑤 = 53000, the raw EEG signals have shape 𝑠𝑡𝑖𝑚𝑒,𝑤 ∈ R𝑁 ×𝑓𝑠 𝑇𝑤 = R6×53000 and the frequency-domain signals have shape 𝑠 𝑓 𝑟𝑒𝑞,𝑤 ∈ R6×1590 .
A.3.2
Preprocessing.
Extracting target channels. Again, the six target channels are extracted and the resulting shape is 𝑠𝑒𝑥𝑡 ∈ R𝑁 ×𝑓𝑠 𝑇 , 𝑁 = 6.
A.2.3 Labels. As mentioned in [7], human emotions can be characterized in the valence-arousal space as in Figure 7. Valence and arousal are two dominant factors categorizing human feelings. Since the SEED dataset only features the valence aspect, the prediction of emotions is focused on the valence component, with labels 𝐿 defined as positive, neutral and negative. Each segmented window corresponds to a single emotion label.
Constructing microstates and baseline. We directly apply the clustering model on the raw EEG and obtain the microstate sequence 𝑐 ∈ 𝑆 𝑓𝑠 𝑇 . The raw EEG signals have shape 𝑠𝑡𝑖𝑚𝑒 ∈ R𝑁 ×𝑓𝑠 𝑇 and the frequency-domain signals have shape 𝑠 𝑓 𝑟𝑒𝑞 ∈ R𝑁 ×𝐵𝑓 𝑓 𝑟𝑒𝑞𝑇 where 𝑓 𝑓 𝑟𝑒𝑞 = 1Hz. Slicing. Since each label lasts for roughly 4s. We set 𝑇𝑤 = 4s. And thus the microstate sequence has length 640, the raw EEG signal has shape 𝑠𝑡𝑖𝑚𝑒,𝑤 ∈ R6×640 where as the frequency-domain signal has shape 𝑠 𝑓 𝑟𝑒𝑞,𝑤 ∈ R6×24 . A.3.3 Labels. We let 𝐿 consists of four labels: left hand, right hand, both hands, both feet. Left hand corresponds to the label 𝑇1 in trials 3, 4, 7, 8, 11, 12, right hand corresponds to the label 𝑇2 in trials 3, 4, 7, 8, 11, 12, both hands corresponds to the label 𝑇1 in trials 5, 6, 9, 10, 13, 14 and both feet corresponds to the label 𝑇2 in trials 5, 6, 9, 10, 13, 14. Each sample corresponds to a single movement label.
12
B
Model Architecture and Training
Generally speaking, the attention mechanism needs a key 𝒌 𝑖 , query 𝒒𝑖 and value 𝒗𝑖 for each input position 𝑖. We can write them as 𝒒𝑖 = 𝑓𝑞 (𝒙 𝑖 , 𝑖)
This section shows the detailed model structures adopted in this work.
B.1
CNN+LSTM [57]
B.1.1
Model Details.
𝒌 𝑖 = 𝑓𝑘 (𝒙 𝑖 , 𝑖) 𝒗𝑖 = 𝑓𝑣 (𝒙 𝑖 , 𝑖) where 𝒙 𝑖 is the word vector at position 𝑖. The attention between position 𝑚, 𝑛 is calculated as
Overview of model structure. The following shows the model structure. The three models have parameters 707K, 692K and 687K respectively. Models are shown in Table 3, Table 4 and Table 5.
𝑒
𝑎𝑚,𝑛 =
CNN and GRU.. The convolution layers are employed to extract the spatial information across channels, and the gated recurrent units (GRUs) are used to extract temporal information. GRU is a simplified version of long short term memory (LSTM) [16]. It consists of two gates—the update gate 𝑧 and the reset gate 𝑟 . At each time step 𝑡, the activation of the update gate 𝒛𝑡 is computed as 𝒛𝑡 = 𝜎 (𝑊𝑧 𝒙 𝑡 + 𝑈𝑧 𝒉𝑡 −1 ) where ℎ𝑡 −1 is the activation of the GRU at time step 𝑡 − 1 and 𝜎 denotes the element-wise sigmoid function. Similarly, the activation 𝒓 𝑡 of the reset gate is computed as
Í𝑁
𝒒𝑇 𝒌 𝑚 √ 𝑛 𝑑
𝑗=1 𝑒
𝒒𝑇 𝒌𝑗 𝑚 √ 𝑑
and since this value is calculated in parallel, we have to incorporate the positional information 𝑖 along with 𝒙 𝑖 into the queries and keys. The main idea of RoFormer is to select a positional embedding such that 𝒒𝑇𝑚 𝒌 𝑛 = 𝑔(𝒙𝑚 , 𝒙 𝑛 , 𝑛 − 𝑚) is a function that depends solely on the input word vector and the relative position between 𝑚, 𝑛. To construct such a positional embedding, let the embedding dimension be 𝑑 which is an even number, then we construct the following matrix
𝒓 𝑡 = 𝜎 (𝑊𝑟 𝒙 𝑡 + 𝑈𝑟 𝒉𝑡 −1 ) Next, the candidate activate ℎ˜𝑡 is computed as
cos 𝑚𝜃 1 © sin 𝑚𝜃 1 0 0 𝑹𝑑Θ,𝑚 = . . . 0 « 0
𝒉˜ 𝑡 = tanh(𝑊 𝒙 𝑡 + 𝒓 𝑇𝑡 𝑈 𝒉𝑡 −1 ) The activation ℎ𝑡 at time step 𝑡 is computed as 𝒉𝑡 = (1𝑇 − 𝒛𝑇𝑡 )𝒉𝑡 −1 + 𝒛𝑇𝑡 𝒉˜ 𝑡 −1 Using this mechanism, the model can selectively consider input at different time steps. Embedding layer. To adapt the model simultaneously to continuous and discrete EEG representations, we use different layers for microstates and conventional representations. For microstates, an embedding layer is adopted to convert discrete microstates into high-dimensional vectors, and for continuous signals, we use a convolution layer, which functions similarly by mapping the input into a high-dimensional latent space. The dimensions are chosen appropriately to guarantee that the model parameters are roughly the same.
− sin 𝑚𝜃 1 cos 𝑚𝜃 1 0 0 .. . 0 0
... ... ... ... .. . ... ...
0 0 0 0 .. . cos 𝑚𝜃 𝑑 2 sin 𝑚𝜃 𝑑 2
0 ª 0 ® ® 0 ® ® 0 ® ® .. ® ® . ® − sin 𝑚𝜃 𝑑 ®® 2 cos 𝑚𝜃 𝑑 ¬ 2
and we have 𝒒𝑚 = 𝑹𝑑Θ,𝑚 𝑾 𝑞 𝒙𝑚 𝒌 𝑛 = 𝑹𝑑Θ,𝑛 𝑾 𝑘 𝒙 𝑘 𝒒𝑇𝑚 𝒌 𝑛 = 𝒙𝑇𝑚 𝑾 𝑇𝑞 𝑹𝑑Θ,𝑛−𝑚 𝑾 𝑘 𝒙 𝑘 = 𝑔(𝒙𝑚 , 𝒙 𝑛 , 𝑛 − 𝑚) Embedding layer. An embedding layer is added before the microstates model. There are 2 extra vectors for padding and classification token. A convolution layer is used instead for continuous signals.
B.1.2 Training Configuration. This section lists the training configurations of the above 3 models in Table 6. The models are trained on an NVIDIA-H20 GPU. The parameters in each case is optimized for performance and memory utilization.
B.2.2 Training Configuration. The models are trained on an NVIDIAH20 GPU. Parameters are optimized for performance and memory utilization.
B.2
Sleep Transformer [49]
B.3
Sleep Net Zero [37]
Model Details.
B.3.1
Model Details.
B.2.1
Overview of model structure. The following shows the model structure of Sleep Net Zero. Parameters are 10.9M, 3.2M and 3.2M. Model details are shown in Table 12, Table 13 and Table 14.
Overview of model structure. The following shows the model structure of Sleep Transformer. Parameters are 3.2M, 3.2M and 3.4M, respectively. Model structures are shown in Table 7, Table 8 and Table 9.
Embedding layer. We adopt an embedding layer with 1002 tokens for microstates. For raw EEG signals, since its input size is significantly larger than the other two representations, we increase the embedding dimension for better performance.
RoFormer. The main part of the model uses an attention-based mechanism to extract temporal features. RoFormer is proposed in [61], which utilizes a novel positional embedding. 13
Table 5: CNN+LSTM for Raw EEG
layer
output
configuration
− Conv1d BatchNorm1d Conv1d MaxPool1d Dropout Conv1d MaxPool1d Dropout Conv1d MaxPool1d Dropout GRU Dropout GRU Dropout reshape Linear
(6, 30000) (1024, 30000) (1024, 30000) (128, 10000) (128, 5000) (128, 5000) (64, 1000) (64, 500) (64, 500) (32, 500) (32, 250) (32, 250) (64, 250) (64, 250) (128, 250) (128, 250) (10, 3200) (10, 5)
− input channels 6, output channels 1024, kernel size 5 padding 2 1024 input channels 1024, output channels 128, kernel size 3 stride 3 kernel size 2, stride 2 𝑝 = 0.25 input channels 128, output channels 64, kernel size 5 stride 5 kernel size 2, stride 2 𝑝 = 0.25 input channels 64, output channels 32, kernel size 3 padding 1 kernel size 2, stride 2 𝑝 = 0.25 input size 32, hidden size 64, 2 layers 𝑝 = 0.25 input size 64, hidden size 128, 2 layers 𝑝 = 0.25 − input features 3200, output features 5
Table 6: CNN+LSTM for Frequency-Domain
layer
output
configuration
− Conv1d BatchNorm1d Conv1d MaxPool1d Dropout Conv1d MaxPool1d Dropout Conv1d MaxPool1d Dropout GRU Dropout GRU Dropout reshape Linear
(6, 1800) (1024, 1800) (1024, 1800) (128, 600) (128, 300) (128, 300) (64, 60) (64, 30) (64, 30) (32, 30) (32, 15) (32, 15) (64, 15) (64, 15) (128, 15) (128, 15) (10, 192) (10, 5)
− input channels 6, output channels 1024, kernel size 5 padding 2 1024 input channels 1024, output channels 128, kernel size 3 stride 3 kernel size 2, stride 2 𝑝 = 0.25 input channels 128, output channels 64, kernel size 5 stride 5 kernel size 2, stride 2 𝑝 = 0.25 input channels 64, output channels 32, kernel size 3 padding 1 kernel size 2, stride 2 𝑝 = 0.25 input size 32, hidden size 64, 2 layers 𝑝 = 0.25 input size 64, hidden size 128, 2 layers 𝑝 = 0.25 − input features 192, output features 5
Overview of model structure. The following shows the model structure of the CNN-based model. Parameters are 19.1M, 19.1M and 20.1M. Model details are shown in Table 16, Table 17 and Table 18.
B.3.2 Training Configuration. The models are trained on an NVIDIAH20 GPU. Parameters in each case are optimized for performance and memory utilization.
B.4
CNN-Based Model for Emotion Recognition [13]
Embedding layer. For microstates, an embedding layer with vocabulary 1001 and dimension 1024 is employed. The extra token is for padding. Convolution layers are used in the place of embedding for the other two representations.
Apart from sleep staging, we show the model used for emotion recognition. B.4.1
Model Details. 14
Table 7: CNN+LSTM for Microstates
layer
output
configuration
− Embedding transpose BatchNorm1d Conv1d MaxPool1d Dropout Conv1d MaxPool1d Dropout Conv1d MaxPool1d Dropout GRU Dropout GRU Dropout reshape Linear
(30000, ) (30000, 512) (512, 30000) (512, 30000) (64, 10000) (64, 5000) (64, 5000) (32, 1000) (32, 500) (32, 500) (16, 500) (16, 250) (16, 250) (32, 250) (32, 250) (64, 250) (64, 250) (10, 1600) (10, 5)
− number of embeddings 1000, dimension 512 − 512 input channels 512, output channels 64, kernel size 3 stride 3 kernel size 2, stride 2 𝑝 = 0.25 input channels 64, output channels 32, kernel size 5 stride 5 kernel size 2, stride 2 𝑝 = 0.25 input channels 32, output channels 16, kernel size 3 padding 1 kernel size 2, stride 2 𝑝 = 0.25 input size 16, hidden size 32, 2 layers 𝑝 = 0.25 input size 32, hidden size 64, 2 layers 𝑝 = 0.25 − input features 1600, output features 5
Table 8: Training Configuration for CNN+LSTM
parameter Raw EEG Freqency-Domain Microstates
batch
optimizer
learning rate
split (train:val:test)
early stop
Adam Adam Adam
10 −4
7:1:2 7:1:2 7:1:2
patience 20 on Kappa patience 20 on Kappa patience 20 on Kappa
64 256 512
10−4 10−4
Table 9: Sleep Transformer for Raw EEG
layer
output
configuration
− Conv1d Conv1d transpose reshape RoFormer slice and reshape RoFormer Linear
(𝑏, 6, 30000) (𝑏, 6, 6000) (𝑏, 6, 3000) (𝑏, 3000, 6) (10𝑏, 300, 6) (10𝑏, 300, 256) (𝑏, 10, 256) (𝑏, 10, 256) (𝑏, 10, 5)
− input channels 6, output channels 6, kernel size 5 stride 5 input channels 6, output channels 6, kernel size 2 stride 2 − − hidden size 256, 2 hidden layers, 4 heads, intermediate size 1024 retrieve only the first along the second dimension and reshape hidden size 256, 2 hidden layers, 4 heads, intermediate size 1024 input features 256, output features 5
B.4.2 Training Configuration. This section lists the training configurations of the above three models. The models are trained on an NVIDIA-H20 GPU. The parameters in each case are optimized for performance and memory utilization.
B.5
B.5.1
Model Details.
Overview of model structure. The following shows the structure of ResNet model. Parameters are 20.3M, 21.5M and 21.4M. Model details are shown in Table 20, Table 21 and Table 22.
ResNet Model for Motor Imagery Classification
ELU.. The ELU activation function is defined as 𝑥 𝑥>0 𝐸𝐿𝑈 (𝑥) = 𝑥 𝛼 (𝑒 − 1) 𝑥 ≤ 0
Finally, we list our model for motor imagery classification. 15
Table 10: Sleep Transformer for Frequency-Domain
layer
output
configuration
− transpose reshape RoFormer slice and reshape RoFormer Linear
(𝑏, 6, 1800) (𝑏, 1800, 6) (10𝑏, 180, 6) (10𝑏, 180, 256) (𝑏, 10, 256) (𝑏, 10, 256) (𝑏, 10, 5)
− − − hidden size 256, 2 hidden layers, 4 heads, intermediate size 1024 retrieve only the first along the second dimension and reshape hidden size 256, 2 hidden layers, 4 heads, intermediate size 1024 input features 256, output features 5
Table 11: Sleep Transformer for Microstates
layer
output
configuration
− Embedding transpose Conv1d Conv1d transpose reshape RoFormer slice and reshape RoFormer Linear
(𝑏, 30000) (𝑏, 30000, 128) (𝑏, 128, 30000) (𝑏, 128, 6000) (𝑏, 128, 3000) (𝑏, 3000, 128) (10𝑏, 300, 128) (10𝑏, 300, 256) (𝑏, 10, 256) (𝑏, 10, 256) (𝑏, 10, 5)
− number of embeddings 1002, dimension 128 − input channels 128, output channels 128, kernel size 5 stride 5 input channels 128, output channels 128, kernel size 2 stride 2 − − hidden size 256, 2 hidden layers, 4 heads, intermediate size 1024 retrieve only the first along the second dimension and reshape hidden size 256, 2 hidden layers, 4 heads, intermediate size 1024 input features 256, output features 5
Table 12: Training Configuration for Sleep Transformer
parameter
batch
optimizer
learning rate
split (train:val:test)
early stop
Raw EEG Freqency-Domain Microstates
64 1000 512
Adam Adam Adam
10 −4 10−4 10−4
7:1:2 7:1:2 7:1:2
patience 20 on Kappa patience 20 on Kappa patience 20 on Kappa
Table 13: ResNetFeatureExtractor
layer
output
configuration
− Conv1d BatchNorm1d ReLU MaxPool1d ResBlocks ResBlocks ResBlocks ResBlocks
(6, 30000) (64, 30000) (64, 30000) (64, 30000) (64, 15000) (512, 3000) (128, 3000) (256, 3000) (512, 3000)
− input channels 6, output channels 64, kernel size 7 padding 3 64 − kernel size 3 stride 2 padding 1 ResBlocks(in=64,out=512,stride=5), see above ResBlocks(in=512,out=128,stride=1), see above ResBlocks(in=128,out=256,stride=1), see above ResBlocks(in=256,out=512,stride=1), see above
Embedding layer. For microstates, an embedding layer with vocabulary 1000 and dimension 1024 is employed. Convolution layers are used in the place of embedding for the other two representations.
B.5.2 Training Configuration. This section lists the training configurations of the above 3 models. The models are trained on an NVIDIA-H20 GPU. Model parameters in each case are optimized for performance and memory utilization. 16
Table 14: Sleep Net Zero for Raw EEG
layer
output
configuration
− ResNetFeatureExtractor transpose RoFormer Linear reshape and mean
(6, 30000) (512, 3000) (3000, 512) (3000, 512) (3000, 5) (10, 5)
− see above − hidden size 512, 2 hidden layers, 2 heads, intermediate size 1024 input features 512, output features 5 compute the mean of every consecutive 300 scores
Table 15: Sleep Net Zero for Frequency-Domain
layer
output
configuration
− Conv1d Conv1d Conv1d Conv1d Conv1d transpose RoFormer Linear reshape and mean
(6, 1800) (640, 1800) (640, 360) (320, 360) (320, 180) (160, 180) (180, 160) (180, 160) (180, 5) (10, 5)
− input channels 6, output channels 640, kernel size 5 padding 2 input channels 640, output channels 640, kernel size 5 padding 5 input channels 640, output channels 320, kernel size 3 padding 1 input channels 320, output channels 320, kernel size 2 padding 2 input channels 320, output channels 160, kernel size 3 padding 1 − hidden size 160, 2 hidden layers, 2 heads, intermediate size 1024 input features 160, output features 5 compute the mean of every consecutive 18 scores
Table 16: Sleep Net Zero for Microstates
layer
output
configuration
− Embedding transpose Conv1d Conv1d Conv1d Conv1d transpose RoFormer Linear reshape and mean
(30000, ) (30000, 512) (512, 30000) (512, 6000) (256, 6000) (256, 3000) (128, 3000) (3000, 128) (3000, 128) (3000, 5) (10, 5)
− number of embeddings 1002, dimension 512 − input channels 512, output channels 512, kernel size 5 padding 5 input channels 512, output channels 256, kernel size 3 padding 1 input channels 256, output channels 256, kernel size 2 padding 2 input channels 256, output channels 128, kernel size 3 padding 1 − hidden size 128, 2 hidden layers, 2 heads, intermediate size 1024 input features 128, output features 5 compute the mean of every consecutive 300 scores
Table 17: Training Configuration for Sleep Net Zero
parameter
batch
Raw EEG Freqency-Domain Microstates
C
96 512 128
optimizer
learning rate
split (train:val:test)
early stop
Adam Adam Adam
10 −4
7:1:2 7:1:2 7:1:2
patience 20 on Kappa patience 20 on Kappa patience 20 on Kappa
10−4 10−4
More Microstate Analysis
C.1
This section provides more analysis and visualization of microstates.
Comparison Between Wake Stage and Rapid Eye Movement (REM) Stage
Humans undergo vivid dreaming processes during REM stage [66]. In turn, EEG signals in REM stage share the same characteristics 17
Table 18: CNN for Raw EEG
layer
output
configuration
− Conv1d Conv1d and ReLU Conv1d and ReLU Conv1d and ReLU MaxPool1d Dropout Conv1d and ReLU Conv1d and ReLU MaxPool1d Dropout flatten Linear Dropout Linear Dropout Linear Dropout Linear
(6, 53000) (1024, 53000) (256, 10600) (128, 2120) (128, 1060) (128, 530) (128, 530) (128, 530) (128, 530) (128, 265) (128, 265) (33920, ) (512, ) (512, ) (256, ) (256, ) (64, ) (64, ) (3, )
− input channels 6, output channels 1024, kernel size 1 input channels 1024, output channels 256, kernel size 5 stride 5 input channels 256, output channels 128, kernel size 5 stride 5 input channels 128, output channels 128, kernel size 2 stride 2 kernel size 2 stride 2 𝑝 = 0.1 input channels 128, output channels 128, kernel size 3 padding 1 input channels 128, output channels 128, kernel size 3 padding 1 kernel size 2 stride 2 𝑝 = 0.1 − input features 33920, output features 512 𝑝 = 0.1 input features 512, output features 256 𝑝 = 0.1 input features 256, output features 64 𝑝 = 0.1 input features 64, output features 3
Table 19: CNN for Frequency-Domain
layer
output
configuration
− reshape Conv2d Conv2d & ReLU Conv2d & ReLU Conv2d & ReLU MaxPool2d Dropout Conv2d & ReLU Conv2d & ReLU MaxPool2d Dropout flatten Linear Dropout Linear Dropout Linear Dropout Linear
(6, 1590) (6, 265, 6) (1024, 265, 6) (256, 265, 6) (128, 265, 6) (128, 265, 6) (128, 265, 3) (128, 265, 3) (128, 265, 3) (128, 265, 3) (128, 265, 1) (128, 265, 1) (33920, ) (512, ) (512, ) (256, ) (256, ) (64, ) (64, ) (3, )
− let the last dimension be the freqency bands input channels 6, output channels 1024, kernel size (1, 1) input channels 1024, output channels 256, kernel size (1, 5) padding (0, 2) input channels 256, output channels 128, kernel size (1, 5) stride (0, 2) input channels 128, output channels 128, kernel size (1, 3) stride (0, 1) kernel size (1, 2) stride (1, 2) 𝑝 = 0.1 input channels 128, output channels 128, kernel size (1, 3) padding (0, 1) input channels 128, output channels 128, kernel size (1, 3) padding (0, 1) kernel size (1, 2) stride (1, 2) 𝑝 = 0.1 − input features 33920, output features 512 𝑝 = 0.1 input features 512, output features 256 𝑝 = 0.1 input features 256, output features 64 𝑝 = 0.1 input features 64, output features 3
with that during wakefulness. We analyze the 30 most frequentappearing microstates during W and REM stages across groups of 30 subjects. The results are as follows: From the above tables we see that the 4 microstates 161, 385, 419 and 421 occur frequently in both R and W stages, with roughly the same ranks. This suggests that the brain undergoes similar activity patterns during these stages.
Further examination of these microstates shows that these microstates have low potential which is below 10𝜇V. This is consistent with the fact that during W stage, brain signals are dominated by 𝛼 waves which have a low potential. Also, this result indicates certain similarities between W stage and R stage since the brain undergoes similar activity.
18
Table 20: CNN for Microstates
layer
output
configuration
− Embedding Conv1d and ReLU Conv1d and ReLU Conv1d and ReLU MaxPool1d Dropout Conv1d and ReLU Conv1d and ReLU MaxPool1d Dropout flatten Linear Dropout Linear Dropout Linear Dropout Linear
(6, 53000) (1024, 53000) (256, 10600) (128, 2120) (128, 1060) (128, 530) (128, 530) (128, 530) (128, 530) (128, 265) (128, 265) (33920, ) (512, ) (512, ) (256, ) (256, ) (64, ) (64, ) (3, )
− number of embeddings 1001, dimension 1024 input channels 1024, output channels 256, kernel size 5 stride 5 input channels 256, output channels 128, kernel size 5 stride 5 input channels 128, output channels 128, kernel size 2 stride 2 kernel size 2 stride 2 𝑝 = 0.1 input channels 128, output channels 128, kernel size 3 padding 1 input channels 128, output channels 128, kernel size 3 padding 1 kernel size 2 stride 2 𝑝 = 0.1 − input features 33920, output features 512 𝑝 = 0.1 input features 512, output features 256 𝑝 = 0.1 input features 256, output features 64 𝑝 = 0.1 input features 64, output features 3
Table 21: Training Configuration for CNN
parameter
batch
optimizer
learning rate
split (train:val:test)
early stop
Raw EEG Freqency-Domain Microstates
128 128 128
Adam Adam Adam
5 × 10 −4 5 × 10−4 10−4
7:1:2 7:1:2 7:1:2
patience 100 on Kappa patience 100 on Kappa patience 100 on Kappa
Table 22: ResNet for Raw EEG
layer
output
configuration
− Conv1d Encoder flatten Classifier
(6, 640) (1024, 640) (128, 640) (81920, ) (4, )
− input channels 6, output channels 1024, kernel size 3 padding 1 see below − see below
Table 23: ResNet for Frequency-Domain
layer
output
configuration
− Conv1d Encoder 2 flatten Classifier 2
(6, 24) (1024, 24) (256, 24) (6144, ) (4, )
− input channels 6, output channels 1024, kernel size 3, padding 1 see below − see below
19
Table 24: ResNet for Microstates
layer
output
configuration
− Conv1d Encoder flatten Classifier
(6, 640) (1024, 640) (128, 640) (81920, ) (4, )
− input channels 6, output channels 1024, kernel size 3, padding 1 see below − see below
Table 25: Encoder Architecture
layer
output
configuration
− Conv1d ResBlock1d ResBlock1d ResBlock1d ELU
(1024, 640) (512, 640) (256, 640) (128, 640) (128, 640) (128, 640)
− input channels 1024, output channels 512, kernel size 13, padding 6 in 512, out 256, kernel 11, see below in 256, out 128, kernel 9, see below in 128, out 128, kernel 7, see below −
Table 26: Encoder 2 Architecture
layer
output
configuration
− Conv1d ResBlock1d ResBlock1d ResBlock1d ELU
(1024, 24) (768, 24) (512, 24) (256, 24) (256, 24) (256, 24)
− input channels 1024, output channels 768, kernel size 13, padding 6 in 768, out 512, kernel 11, see below in 512, out 256, kernel 9, see below in 256, out 256, kernel 7, see below −
Table 27: Classifier Architecture
layer
output
configuration
− Linear and ReLU Linear and ReLU Linear and ReLU Linear
(81920, ) (128, ) (128, ) (64, ) (4, )
− in features 81920, out features 128 in features 128, out features 128 in features 128, out features 64 in features 64, out features 4
Table 28: Classifier 2 Architecture
layer
output
configuration
− Linear and ReLU Linear and ReLU Linear and ReLU Linear
(6144, ) (128, ) (128, ) (64, ) (4, )
− in features 6144, out features 128 in features 128, out features 128 in features 128, out features 64 in features 64, out features 4
20
Table 29: Training Configuration for ResNet
parameter
batch
Raw EEG Freqency-Domain Microstates
128 128 128
optimizer
learning rate
split (train:val:test)
early stop
Adam Adam Adam
5 × 10 −4
7:1:2 7:1:2 7:1:2
patience 100 on Kappa patience 100 on Kappa patience 10 on Kappa
5 × 10−4 2 × 10−6
Table 30: Rank among Subjects under W Stage
microstate
rank among 10 groups of subjects
.. . 160 161 162 .. . 384 385 386 .. . 418 419 420 421 422 .. .
.. . 24 14 917
29 11 827
23 16 892
23 17 872
30 14 814
25 15 775
26 15 839
25 14 813
20 19 807
27 14 858
218 2 640
150 3 654
220 2 531
258 3 402
165 2 634
659 1 608 3 508
597 1 670 2 525
511 1 609 3 494
409 1 607 2 349
587 1 654 3 548
.. . 229 3 338
166 2 721
236 3 589
165 2 677
188 4 543 .. .
701 1 537 2 573
519 1 396 3 434
470 1 645 2 515
667 1 619 3 458
672 1 592 2 496 .. .
Table 31: Rank among Subjects under R Stage
C.2
microstate
rank among 10 groups of subjects
.. . 160 161 162 .. . 384 385 386 .. . 418 419 420 421 422 .. .
.. . 27 15 831
23 15 810
24 18 840
22 21 699
26 14 672
23 14 762
25 21 772
25 16 688
20 22 882
24 21 720
180 3 817
174 8 773
192 3 736
220 3 714
200 3 634
592 1 711 2 332
557 1 753 2 381
594 1 644 2 352
528 1 811 2 325
616 1 722 2 346
.. . 158 5 544
156 4 626
202 7 712
202 3 818
168 4 818 .. .
466 1 589 3 420
639 1 671 2 386
569 2 731 1 368
527 1 581 2 281
.. .
Comparison Between the Wake Stage and the Non-Rapid Eye Movement III Stage
NREM3 stage denotes deep sleep. In this case, the brain activity differs from that in the wake stage.
634 1 631 2 460
21
Table 32: Rank among Subjects under N3 Stage
microstate
rank among 10 groups of subjects
.. . 377 378 379 .. . 451 452 453 .. . 650 651 652 .. .
.. . 625 1 800
396 2 888
672 3 799
703 1 845
134 2 825
737 2 832
748 1 830
577 2 824
358 2 817
771 1 791
755 15 932
822 8 709
871 5 910
614 13 965
782 8 899
68 6 814
109 3 850
170 4 755
135 3 872
120 2 784
.. . 753 18 939
811 7 979
707 10 915
779 10 881
627 10 902 .. .
28 6 913
148 4 908
89 13 793
244 2 786
117 4 929 .. .
Figure 10: Visualizing Microstates 161, 385, 419 and 421
Figure 9: The Model Structure of ResBlock1d with in=N, out=M and kernel=K From the microstates distribution we see that the dominant microstates are different from that of W stage. To further back this observation, we record the rank of microstates 378, 452, 385 across W and N3 stage. Results demonstrate that the microstates that frequently occur during W stage typically occur rarely in N3 stage. This again shows that the brain activity differs considerably between these 2 stages. We further visualize the microstates:
Figure 8: The Model Structure of ResBlocks with in=N, out=M and stride=K Figure 11: Visualizing Microstates 378, 452 and 651
22
and we see that these microstates correspond to a state with a relatively high potential, typically > 10𝜇V. This is consistent with the fact that during N3 stage, brain activity is dominated by 𝛿 waves which has a high amplitude [77]. Nonetheless, EEG signals are oscillating and will not always remain at a high voltage, hence in certain cases low-potential states like the microstates dominating in W stage will also occur with a relatively high frequency.
Table 33: Rank among Subjects under N3 Stage
microstate
microstate 378 .. . 651 .. . 487 .. . 452 .. . .. .
rank 1 .. . 6 .. . 14 .. . 18 .. . .. .
microstate .. . 378 .. . 651 .. . 452 .. . 487 .. .
rank among 10 groups of subjects
378
W N3
132 0
120 1
214 2
146 0
169 1
158 1
109 0
165 1
172 1
163 0
452
W N3
95 17
85 6
166 9
98 9
112 9
115 14
112 7
100 4
143 12
132 7
421
W N3
3 489
2 124
1 26
2 52
2 278
2 161
2 71
2 54
2 34
2 78
microstate 419 421 385 .. . 333 .. . 161 .. . .. .
rank 1 2 3 .. . 7 .. . 14 .. . .. .
rank .. . 2 .. . 4 .. . 7 .. . 11 .. .
microstate .. . 487 378 .. . 452 .. . 651 .. . .. .
rank .. . 2 3 .. . 10 .. . 13 .. . .. .
rank 1 2 3 .. . 5 .. . 15 .. . .. .
microstate 419 421 333 385 .. . .. . 161 .. . .. .
rank 1 2 3 4 .. . .. . 11 .. . .. .
microstate 421 419 333 .. . 385 .. . 161 .. . .. .
rank 1 2 3 .. . 6 .. . 11 .. . .. .
microstate 419 421 385 .. . 333 .. . 161 .. . .. .
rank 1 2 3 .. . 7 .. . 16 .. . .. .
subjects 1 ∼ 30 W subjects 31 ∼ 60 W subjects 61 ∼ 90 W stage stage stage Table 35: Rank among Subjects under N3 Stage
subjects 1 ∼ 30 N3 subjects 31 ∼ 60 N3 subjects 61 ∼ 90 N3 stage stage stage Table 36: Rank among Subjects under N3 Stage
microstate 419 333 421 .. . 385 .. . 161 .. . .. .
microstate 419 385 421 .. . 333 .. . 161 .. . .. .
C.3
rank 1 2 3 .. . 7 .. . 18 .. . .. .
Comparing other Sleep Stages
For sleep stage N1 and N2, the dominant microstates are also 419, 385, 421, 615 and other states found in W stage. This suggests that these microstates capture a class of weak EEG signals that the brain usually switches between. Also, the brain activity in stages N1 and N2 shares certain aspects with that in W stage. We also found that when transforming from stage W through stage N1, stage N2 and finally to stage N3, the frequency of microstate 489 is increasing. The visualization of 489 is as follows:
subjects 1 ∼ 30 R subjects 31 ∼ 60 R subjects 61 ∼ 90 R stage stage stage Table 34: Rank among Subjects under N3 Stage
Figure 12: Visualizing Microstates 489 which also has relatively high potential with all leads between 2 ∼ 14𝜇V. This again shows that from W through N1, N2 to N3, high amplitude brain activity becomes more and more common. 23