Background: in this study, we sought to provide powerful machine learning-based classification models to predict the success rate of intrauterine insemination (IUI) therapy. In addition, we tried to show the effect of model fit with balanced data compared to unbalanced data and also model fit with all features versus optimized feature sets.Material and Methods: Data obtained from Fatemehzahra
Infertility Research Center in Babol that ۵۴۶ infertile women who underwent IUI were taken for this cross-sectional study. Using Python v۳.۷ software, Regression Logistic, SVC, Random Forest, XGBoost, and Stack Classifiers trained by all and optimal features selected by ۳ methods from original, resampled by Smote-Tomek and resampled by Smote-ENN datasets, that to evaluate the efficiency and power of each model predictions by three tools: Boxplot, ROC Curve, and Calibration Plot, and assessed with evaluation indicators: Gmean, AUC and Brier, respectivelyFindings: The results of the study showed that Random Forest feature selection method had better performance in model training than the other two feature selection methods. Then, by data sets evaluating, it was found that resampled data by Smote-Tomek method is the most suitable data for use in training of classification models. Finally, XGBoost offered the best performance compared to other learning algorithms used by achieving Gmean, AUC, Brier values of ۰.۸۰, ۰.۸۹, and ۰.۱۲۹, respectively. It showed infertility duration is the most predictable factor in IUI success.Conclusion: A prediction model based on XGBoost was developed using duration of infertility, male and female age, sperm concentration, sperm motility grading score as the most effective predictors. The results of this study can be used in personal estimation of odds
Cumulative live birth in the first complete IUI cycle before treatment for researchers.