Interpretable Machine Learning for Type ۲ Diabetes Prediction Using Clinical, Laboratory, and Lifestyle Variables
سال انتشار: 1405
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 58
فایل این مقاله در 12 صفحه با فرمت PDF قابل دریافت می باشد
- صدور گواهی نمایه سازی
- من نویسنده این مقاله هستم
استخراج به نرم افزارهای پژوهشی:
شناسه ملی سند علمی:
MHHCONG02_022
تاریخ نمایه سازی: 14 شهریور 1405
چکیده مقاله:
Type ۲ diabetes (T۲D) is a major public health challenge, making accurate early prediction essential for prevention and clinical decision-making. This study developed machine learning models using demographic, clinical, laboratory, and lifestyle data from the NHANES ۲۰۲۱–۲۰۲۳ survey. The outcome variable was the self-reported physician diagnosis of diabetes (DIQ۰۱۰), eliminating data leakage and allowing all biomarkers, including HbA۱c, to be used safely as predictors. After excluding participants with incomplete data, the final dataset included ۴,۱۵۴ individuals, including ۵۲۶ with diabetes (۱۲.۶%). Ten machine learning algorithms were evaluated using stratified five-fold cross-validation and GridSearchCV. Logistic Regression achieved the highest performance (ROC-AUC = ۰.۹۵۰), followed by Extra Trees (۰.۹۴۵) and Random Forest (۰.۹۳۸). Feature importance analyses using ReliefF, Permutation Importance, and SHAP consistently identified HbA۱c as the strongest predictor, while a history of hypertension ranked second across most methods. Vitamin D showed only a modest independent contribution after adjustment for established risk factors. Association rule mining identified ۱۹ clinically meaningful rules linking glycemic and cardiovascular risk factors with diabetes. Overall, combining laboratory, anthropometric, and lifestyle variables provided accurate and clinically interpretable prediction of T۲D.
کلیدواژه ها:
نویسندگان
Ali Mansoor
Islamic Azad University of West Tehran Branch, Tehran, Iran