Library
PubMed
research article
Professional

Development and internal validation of a machine-learning-based risk identification model for diabetic kidney disease in patients with type 2 diabetes.

Source: PubMed, NCBI / U.S. National Library of Medicine

Frontiers in endocrinologyLuo Jianquan, Chen Guoxin, Yang JiachengPublished 1/1/2026Last synced 8/16/2026Status: syncedPMID: 42582167DOI: 10.3389/fendo.2026.1888060

Diabetic kidney disease (DKD) is a major microvascular complication of type 2 diabetes mellitus (T2DM) and the leading cause of end-stage renal disease in China. Limited disease awareness and insufficient early screening tools hinder timely intervention for DKD patients. This study aimed to develop and internally validate a non-invasive machine learning-based prediction model for early DKD risk identification among T2DM patients using routine clinical laboratory indicators. A retrospective cross-sectional study was conducted with 602 eligible T2DM patients (457 without DKD, 145 with DKD) recruited from Sihui People's Hospital between January 2023 and June 2024. Subjects were randomly split into an 8:2 training set (n=481) and independent test set (n=121). Baseline clinical characteristics were compared between groups. Univariate and multivariate logistic regression analyses were performed to screen independent DKD predictors. Six machine learning algorithms including logistic regression, XGBoost, random forest, AdaBoost, support vector classifier (SVC), and Gaussian naive Bayes (GNB) were constructed and comprehensively assessed via AUC, accuracy, sensitivity, specificity, calibration curves, decision curve analysis (DCA), and SHapley Additive exPlanations (SHAP) interpretability analysis. Baseline comparisons showed that DKD patients presented worse glycolipid, hepatic, renal, hematological and urinary protein indicators than patients with isolated T2DM, while sex, age, BMI,

Abstract

Diabetic kidney disease (DKD) is a major microvascular complication of type 2 diabetes mellitus (T2DM) and the leading cause of end-stage renal disease in China. Limited disease awareness and insufficient early screening tools hinder timely intervention for DKD patients. This study aimed to develop and internally validate a non-invasive machine learning-based prediction model for early DKD risk identification among T2DM patients using routine clinical laboratory indicators. A retrospective cross-sectional study was conducted with 602 eligible T2DM patients (457 without DKD, 145 with DKD) recruited from Sihui People's Hospital between January 2023 and June 2024. Subjects were randomly split into an 8:2 training set (n=481) and independent test set (n=121). Baseline clinical characteristics were compared between groups. Univariate and multivariate logistic regression analyses were performed to screen independent DKD predictors. Six machine learning algorithms including logistic regression, XGBoost, random forest, AdaBoost, support vector classifier (SVC), and Gaussian naive Bayes (GNB) were constructed and comprehensively assessed via AUC, accuracy, sensitivity, specificity, calibration curves, decision curve analysis (DCA), and SHapley Additive exPlanations (SHAP) interpretability analysis. Baseline comparisons showed that DKD patients presented worse glycolipid, hepatic, renal, hematological and urinary protein indicators than patients with isolated T2DM, while sex, age, BMI, DBP and FPG showed no significant intergroup differences (all P>0.05). Multivariate logistic regression identified glycated hemoglobin (HbA1c), β-microglobulin (βMG), and urine protein (PRO) as independent risk factors, while serum albumin (ALB) and estimated glomerular filtration rate (eGFR) acted as protective factors. Single indicator ROC analysis showed PRO and βMG achieved the highest diagnostic AUC of 0.944. All six machine learning models exhibited excellent discriminative performance with validation AUCs over 0.975. Logistic regression was selected as the optimal model, yielding a test-set AUC of 0.979, sensitivity of 89.7%, and specificity of 92.4%. The model demonstrated favorable calibration and sustained positive net clinical benefit across almost all threshold probabilities in DCA. SHAP analysis ranked HbA1c, ALB, PRO, eGFR, and βMG as the top five predictive features, clarifying the individual risk contribution of each biomarker. The interpretable logistic regression model built on routine non-invasive clinical indicators reliably identifies DKD risk in T2DM patients. Glycemic control and renal injury biomarkers serve as core predictive factors. This tool provides convenient, low-cost early risk stratification for clinical practice, especially for primary care settings. Further external multicenter prospective validation is required to generalize its clinical application.

Educational only
This information is for general education and is not medical advice. Always talk to a licensed U.S. clinician about your situation, medications, or treatment decisions.