Machine learning-based triage of non-red-flag low back pain in national-level collegiate athletes: A cross-sectional study in Taiwan.
Source: PubMed, NCBI / U.S. National Library of Medicine
Non-red-flag low back pain (LBP) is prevalent among athletes, yet field-based triage often hinges on subjective judgment under significant time constraints, or it may be ignored and left untreated. To benchmark multiple machine learning (ML) algorithms using routine field data to support structured triage prioritization for athletes with LBP. A total of 412 symptomatic collegiate athletes were classified into self-management or professional referral groups using an expert-consensus triage framework. Logistic regression (LR), random forest (RF), XGBoost, and LightGBM were compared based on discrimination, calibration, and generalization stability. The final model was assessed with decision curve analysis (DCA) and interpreted using SHapley Additive exPlanations (SHAP), while Youden's index was used to select the optimal ROC-based cutoff. The RF model showed the most balanced performance, with a smaller generalization gap than XGBoost and LightGBM (8% vs >20%). It achieved a test AUROC of 0.723, slightly outperforming L2-regularized logistic regression (0.702), and demonstrated better calibration, with a lower Brier score (0.219 vs 0.239). DCA showed positive net benefit across clinically relevant thresholds of 0.20-0.60, with a net benefit of 0.294 at the Youden-derived threshold of 0.454, compared with 0.202 for the L2-regularized LR model. At this threshold, sensitivity was 93%, and specificity was 53%. SHAP analysis highlighted symptom history, functional measures, and prio
Abstract
Non-red-flag low back pain (LBP) is prevalent among athletes, yet field-based triage often hinges on subjective judgment under significant time constraints, or it may be ignored and left untreated. To benchmark multiple machine learning (ML) algorithms using routine field data to support structured triage prioritization for athletes with LBP. A total of 412 symptomatic collegiate athletes were classified into self-management or professional referral groups using an expert-consensus triage framework. Logistic regression (LR), random forest (RF), XGBoost, and LightGBM were compared based on discrimination, calibration, and generalization stability. The final model was assessed with decision curve analysis (DCA) and interpreted using SHapley Additive exPlanations (SHAP), while Youden's index was used to select the optimal ROC-based cutoff. The RF model showed the most balanced performance, with a smaller generalization gap than XGBoost and LightGBM (8% vs >20%). It achieved a test AUROC of 0.723, slightly outperforming L2-regularized logistic regression (0.702), and demonstrated better calibration, with a lower Brier score (0.219 vs 0.239). DCA showed positive net benefit across clinically relevant thresholds of 0.20-0.60, with a net benefit of 0.294 at the Youden-derived threshold of 0.454, compared with 0.202 for the L2-regularized LR model. At this threshold, sensitivity was 93%, and specificity was 53%. SHAP analysis highlighted symptom history, functional measures, and prior ankle injury as key contributors. This study shows that an interpretable RF-based framework integrating symptom and functional data provides a potentially scalable, probabilistic tool for frontline LBP triage. It delivers acceptable discrimination and improved calibration, serving as a transparent decision-support adjunct to enhance screening safety. Further external validation and implementation studies in diverse athletic cohorts are needed to confirm generalizability. This study was a non-interventional observational study using de-identified secondary data from collegiate athletic records. Prospective registration was therefore not applicable.
