Interpretable Machine Learning for Type ۲ Diabetes Prediction Using Clinical, Laboratory, and Lifestyle Variables

سال انتشار: 1405
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 58

فایل این مقاله در 12 صفحه با فرمت PDF قابل دریافت می باشد

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

شناسه ملی سند علمی:

MHHCONG02_022

تاریخ نمایه سازی: 14 شهریور 1405

چکیده مقاله:

Type ۲ diabetes (T۲D) is a major public health challenge, making accurate early prediction essential for prevention and clinical decision-making. This study developed machine learning models using demographic, clinical, laboratory, and lifestyle data from the NHANES ۲۰۲۱–۲۰۲۳ survey. The outcome variable was the self-reported physician diagnosis of diabetes (DIQ۰۱۰), eliminating data leakage and allowing all biomarkers, including HbA۱c, to be used safely as predictors. After excluding participants with incomplete data, the final dataset included ۴,۱۵۴ individuals, including ۵۲۶ with diabetes (۱۲.۶%). Ten machine learning algorithms were evaluated using stratified five-fold cross-validation and GridSearchCV. Logistic Regression achieved the highest performance (ROC-AUC = ۰.۹۵۰), followed by Extra Trees (۰.۹۴۵) and Random Forest (۰.۹۳۸). Feature importance analyses using ReliefF, Permutation Importance, and SHAP consistently identified HbA۱c as the strongest predictor, while a history of hypertension ranked second across most methods. Vitamin D showed only a modest independent contribution after adjustment for established risk factors. Association rule mining identified ۱۹ clinically meaningful rules linking glycemic and cardiovascular risk factors with diabetes. Overall, combining laboratory, anthropometric, and lifestyle variables provided accurate and clinically interpretable prediction of T۲D.

نویسندگان

Ali Mansoor

Islamic Azad University of West Tehran Branch, Tehran, Iran