Approaching the Human Ceiling in Sleep Staging: A Controlled Comparison of Feature-Engineered and Deep Architectures
DOI:
https://doi.org/10.52465/joiser.v4i3.20Keywords:
Hypnogram, EEG, EOG, Sleeping stages, Deep learningAbstract
Polysomnography (PSG) is the clinical gold standard for sleep assessment, yet manual epoch-by-epoch scoring is labor-intensive and subject to substantial inter-rater variability, which limits its use in large-scale, home-based, and routine clinical practice and makes reliable automated scoring an urgent need. This study suggests and compares an automated sleep-staging model based on two different supervised pipelines: a conventional feature-engineered Random Forest and a state-of-the-art end-to-end Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) network. Based on the recordings of the Sleep-EDF Expanded database, we preprocessed the multichannel signals (EEG and EOG) into 30-second epochs, which were classified into five stages (Wake, N1, N2, N3 and REM). Our findings show that the feature-engineered baseline is very competitive with 95.49% accuracy, but the CNN-LSTM model performs better (96.44% accuracy) and has a much higher capability of classifying the transitional N1 stage which is harder. Because these results are obtained from a single overnight recording with a within-subject train/test split, they are best interpreted as a controlled proof-of-concept upper bound rather than as evidence of clinical-grade or human-level performance; cross-subject (leave-one-subject-out) validation on multiple recordings is required before any such claim can be made.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Information System Exploration and Research

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.


