A Mutual-Information-Guided and ADASYN-Augmented Machine Learning Framework for Early Prediction of Parkinson’s Disease
سال انتشار: 1405
نوع سند: مقاله ژورنالی
زبان: انگلیسی
مشاهده: 110
فایل این مقاله در 12 صفحه با فرمت PDF قابل دریافت می باشد
- صدور گواهی نمایه سازی
- من نویسنده این مقاله هستم
استخراج به نرم افزارهای پژوهشی:
شناسه ملی سند علمی:
JR_MSESJ-8-3_002
تاریخ نمایه سازی: 5 خرداد 1405
چکیده مقاله:
Early detection of Parkinson’s disease (PD) is essential for timely medical intervention and improving patient outcomes. Speech signal analysis offers a non-invasive, cost-effective, and easily deployable diagnostic pathway. However, achieving reliable early prediction remains challenging due to data imbalance, redundant features, and model instability. This study aims to develop an optimized and robust machine learning framework that enhances the predictive accuracy and stability of PD detection from speech data. An optimized machine learning model based on eXtreme Gradient Boosting (XGBoost) was developed for early PD prediction. The model’s hyperparameters were tuned using the Tree-structured Parzen Estimator (TPE), while Mutual Information (MI) was employed to select the most informative features from the speech dataset. To address class imbalance, the Adaptive Synthetic Sampling Approach for Imbalanced Learning (ADASYN) was applied to generate synthetic minority samples. Model performance and stability were evaluated using ten independent runs of Stratified ۱۰-Fold Cross-Validation (SCV). The proposed framework achieved superior predictive performance with an average accuracy of ۹۷.۲۷%, precision of ۹۸.۷۹%, F۱-score of ۹۷.۱۸%, recall of ۹۵.۷۷%, and ROC-AUC of ۹۸.۱۱% across multiple evaluations. Comparative analysis with similar studies demonstrated improved robustness, reliability, and balance between sensitivity and specificity. The integration of MI-based feature selection and ADASYN-based data augmentation significantly enhanced the performance and stability of the XGBoost model for early PD prediction. The proposed model demonstrates strong potential for clinical use as a decision support system, providing a low-cost, non-invasive, and remotely deployable tool for early PD diagnosis using patient speech signals.Early detection of Parkinson’s disease (PD) is essential for timely medical intervention and improving patient outcomes. Speech signal analysis offers a non-invasive, cost-effective, and easily deployable diagnostic pathway. However, achieving reliable early prediction remains challenging due to data imbalance, redundant features, and model instability. This study aims to develop an optimized and robust machine learning framework that enhances the predictive accuracy and stability of PD detection from speech data. An optimized machine learning model based on eXtreme Gradient Boosting (XGBoost) was developed for early PD prediction. The model’s hyperparameters were tuned using the Tree-structured Parzen Estimator (TPE), while Mutual Information (MI) was employed to select the most informative features from the speech dataset. To address class imbalance, the Adaptive Synthetic Sampling Approach for Imbalanced Learning (ADASYN) was applied to generate synthetic minority samples. Model performance and stability were evaluated using ten independent runs of Stratified ۱۰-Fold Cross-Validation (SCV). The proposed framework achieved superior predictive performance with an average accuracy of ۹۷.۲۷%, precision of ۹۸.۷۹%, F۱-score of ۹۷.۱۸%, recall of ۹۵.۷۷%, and ROC-AUC of ۹۸.۱۱% across multiple evaluations. Comparative analysis with similar studies demonstrated improved robustness, reliability, and balance between sensitivity and specificity. The integration of MI-based feature selection and ADASYN-based data augmentation significantly enhanced the performance and stability of the XGBoost model for early PD prediction. The proposed model demonstrates strong potential for clinical use as a decision support system, providing a low-cost, non-invasive, and remotely deployable tool for early PD diagnosis using patient speech signals.
کلیدواژه ها:
نویسندگان
Ghadeer Aqil Ali
Department of Electrical And Computer Engineering, Urmia University, Urmia, Iran.
Leila Sharifi
Assistant Professor, Department of Electrical and Computer Engineering, Urmia University, Urmia, Iran
Parviz Rashidi-Khazaee
Assistant Professor, Department of Information Technology and Computer Engineering, Urmia University of Technology, Urmia, Iran
Hossein Nahid-Titkanlue
Assistant Professor, Department of Industrial Engineering, Payame Noor University, Tehran, Iran.
مراجع و منابع این مقاله:
لیست زیر مراجع و منابع استفاده شده در این مقاله را نمایش می دهد. این مراجع به صورت کاملا ماشینی و بر اساس هوش مصنوعی استخراج شده اند و لذا ممکن است دارای اشکالاتی باشند که به مرور زمان دقت استخراج این محتوا افزایش می یابد. مراجعی که مقالات مربوط به آنها در سیویلیکا نمایه شده و پیدا شده اند، به خود مقاله لینک شده اند :