Machine Learning-Based Prediction of Antepartum Depression Symptoms: A Prospective Cohort Study
Source: PubMed Central Open Access, NCBI / U.S. National Library of Medicine
Background Antepartum depression is a common mental health issue. This study aimed to develop a tool to predict antepartum depression risk and contribute to improving the screening rate of antepartum depression. Methods This study used a prospective design, including a total of 1,701 mid-pregnancy women, who were followed up until late pregnancy. We comprehensively incorporate predictors from biological, psychological, demography, and obstetrics. Several machine learning algorithms (logistic regression, random forest, support vector machine and extreme gradient boosting) were used to predict antepartum depression. Predictive variables were screened using the least absolute shrinkage and selection operator (LASSO). The models were built based on 70% of the training set and evaluated on the remaining 30% test set using metrics such as accuracy, precision, sensitivity, specificity, F1 score, and area under the receiver operating characteristic curve (AUC). We also ranked the importance of the variables. Results LASSO was used to select the nine variables (out of a total of 42). The importance ranking of the nine variables is as follows: psychological resilience, rumination, stress, social support, body image satisfaction during pregnancy, preparation for neonatal parenting, low-density lipoprotein cholesterol, work status and relationship with parents. The logistic regression achieved 68% accuracy, 25.6% precision, 68.7% sensitivity, 68.4% specificity, 37.3% F1-score, and 76% AU
Abstract
Background Antepartum depression is a common mental health issue. This study aimed to develop a tool to predict antepartum depression risk and contribute to improving the screening rate of antepartum depression. Methods This study used a prospective design, including a total of 1,701 mid-pregnancy women, who were followed up until late pregnancy. We comprehensively incorporate predictors from biological, psychological, demography, and obstetrics. Several machine learning algorithms (logistic regression, random forest, support vector machine and extreme gradient boosting) were used to predict antepartum depression. Predictive variables were screened using the least absolute shrinkage and selection operator (LASSO). The models were built based on 70% of the training set and evaluated on the remaining 30% test set using metrics such as accuracy, precision, sensitivity, specificity, F1 score, and area under the receiver operating characteristic curve (AUC). We also ranked the importance of the variables. Results LASSO was used to select the nine variables (out of a total of 42). The importance ranking of the nine variables is as follows: psychological resilience, rumination, stress, social support, body image satisfaction during pregnancy, preparation for neonatal parenting, low-density lipoprotein cholesterol, work status and relationship with parents. The logistic regression achieved 68% accuracy, 25.6% precision, 68.7% sensitivity, 68.4% specificity, 37.3% F1-score, and 76% AUC in the test set. The random forest achieved 66.1% accuracy, 24.1% precision, 68.5% sensitivity, 65.7% specificity, 35.6% F1-score, and 75.4% AUC in the test set. The support vector machine achieved 65.1% accuracy, 23% precision, 65.7% sensitivity, 65.1% specificity, 34.1% F1-score, and 70.1% AUC in the test set. The extreme gradient boosting achieved 86.8% accuracy, 57.8% precision, 15.7% sensitivity, 98.1% specificity, 24.7% F1-score, and 71.8% AUC in the test set. Conclusion The antepartum depression risk prediction model developed in this study has moderate discriminative ability, which helps achieve early identification and classification management of risk, and provides a basis for targeted interventions.
