<oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd"><dc:title>Metrics Matter: Comparing Classification and Segmentation Models for Automated Sleep Scoring Using Electrophysiological Signals</dc:title><dc:creator>Umbel, Lily </dc:creator><dc:subject>machine learning</dc:subject><dc:subject>ai</dc:subject><dc:subject>ml</dc:subject><dc:subject>sleep scoring</dc:subject><dc:subject>sleep architecture</dc:subject><dc:subject>classification</dc:subject><dc:subject>segmentation</dc:subject><dc:subject>neural engineering</dc:subject><dc:coverage>Engineering Science</dc:coverage><dc:relation>B S</dc:relation><dc:description>Machine learning has emerged as a powerful methodology for automating time-intensive tasks in research. One domain this has been applied to within neural engineering is sleep scoring. Traditional sleep scoring is a labor-intensive process involving experts hand scoring electrophysiological signals, including electroencephalogram (EEG) and electromyogram (EMG). This thesis investigates whether classification- or segmentation-based machine learning models perform better in the context of characterizing sleep architecture on EEG/EMG data. Both classes of models were evaluated on three days of electrophysiological recordings from a single APOE4 mouse, using the same band-power feature inputs and evaluation procedures across all models. Classification models included convolutional neural networks (CNN), long short-term memory networks (LSTM), and hidden Markov models (HMM). These models assign discrete sleep stage labels (Wake, REM, NREM) to fixed time epochs. On the other hand, segmentation models—including fully convolutional networks (FCN) and U-Nets—identify transition boundaries continuously across the temporal signal. Additionally, three hybrid models were tested using the same procedure: CNN→U-Net (segmentation-based) and CNN→BiLSTM (classification-based), and CNN→HMM (classification-based). Beyond the standard machine learning metrics (accuracy and macro-F1), this thesis introduces sleep architectural evaluation metrics that quantify fragmentation, bout-length distribution, and state transition times. While standard metrics and confusion matrices suggest comparable performance across models, these architectural metrics reveal fundamental differences. Classification models produce hypnograms with vastly distorted structure, over-predicting sleep stage transitions by 300–700%. Despite classification models achieving accuracies above 97% and F1 scores above 0.93, these standard metrics failed to capture the crucial differences revealed by architectural analysis. In contrast, segmentation models more accurately capture temporal structure: the U-Net was the only pure model whose bout-length distributions were statistically indistinguishable from ground truth. The significance of this finding is that sleep is better modeled as a continuous phenomenon to capture the underlying probabilistic dynamics. This work directly supports the Gluckman Group's goal of robust automated sleep scoring to study the relationship between sleep architecture, APOE4 genotype, and Alzheimer’s disease risk.</dc:description><dc:contributor>Bruce Gluckman, Thesis Supervisor</dc:contributor><dc:contributor>Gary L. Gray, Thesis Honors Advisor</dc:contributor><dc:rights>open_access</dc:rights><dc:date>2026-04-17T17:47:28Z</dc:date><dc:identifier>https://honors.libraries.psu.edu/catalog/10076lku5022</dc:identifier></oai_dc:dc>