Metrics Matter: Comparing Classification and Segmentation Models for Automated Sleep Scoring Using Electrophysiological Signals
Open Access
- Author:
- Umbel, Lily
- Area of Honors:
- Engineering Science
- Degree:
- Bachelor of Science
- Document Type:
- Thesis
- Thesis Supervisors:
- Bruce Gluckman, Thesis Supervisor
Gary L. Gray, Thesis Honors Advisor - Keywords:
- machine learning
ai
ml
sleep scoring
sleep architecture
classification
segmentation
neural engineering - Abstract:
- Machine learning has emerged as a powerful methodology for automating time-intensive tasks in research. One domain this has been applied to within neural engineering is sleep scoring. Traditional sleep scoring is a labor-intensive process involving experts hand scoring electrophysiological signals, including electroencephalogram (EEG) and electromyogram (EMG). This thesis investigates whether classification- or segmentation-based machine learning models perform better in the context of characterizing sleep architecture on EEG/EMG data. Both classes of models were evaluated on three days of electrophysiological recordings from a single APOE4 mouse, using the same band-power feature inputs and evaluation procedures across all models. Classification models included convolutional neural networks (CNN), long short-term memory networks (LSTM), and hidden Markov models (HMM). These models assign discrete sleep stage labels (Wake, REM, NREM) to fixed time epochs. On the other hand, segmentation models—including fully convolutional networks (FCN) and U-Nets—identify transition boundaries continuously across the temporal signal. Additionally, three hybrid models were tested using the same procedure: CNN→U-Net (segmentation-based) and CNN→BiLSTM (classification-based), and CNN→HMM (classification-based). Beyond the standard machine learning metrics (accuracy and macro-F1), this thesis introduces sleep architectural evaluation metrics that quantify fragmentation, bout-length distribution, and state transition times. While standard metrics and confusion matrices suggest comparable performance across models, these architectural metrics reveal fundamental differences. Classification models produce hypnograms with vastly distorted structure, over-predicting sleep stage transitions by 300–700%. Despite classification models achieving accuracies above 97% and F1 scores above 0.93, these standard metrics failed to capture the crucial differences revealed by architectural analysis. In contrast, segmentation models more accurately capture temporal structure: the U-Net was the only pure model whose bout-length distributions were statistically indistinguishable from ground truth. The significance of this finding is that sleep is better modeled as a continuous phenomenon to capture the underlying probabilistic dynamics. This work directly supports the Gluckman Group's goal of robust automated sleep scoring to study the relationship between sleep architecture, APOE4 genotype, and Alzheimer’s disease risk.
Accessible Version in Progress
We're generating an accessible version of this file to meet ADA Title II requirements. This process may take up to one hour. Please return later to access the accessible copy once it's ready.
You can still download the current version by clicking "OK".
What's happening:
An accessible PDF is being generated using Adobe with AI used to generate alternative text (alt text) for images in the PDF.