HEALTHCARE DATA ANALYTICS FOR EARLY PREDICTION OF CHRONIC DISEASES USING ELECTRONIC HEALTH RECORDS: A MACHINE LEARNING–BASED PREDICTIVE STUDY
Keywords:
Electronic health records; machine learning; chronic disease; predictive analytics; diabetes; hypertension; cardiovascular disease.Abstract
Background: Chronic diseases such as diabetes mellitus, hypertension, and cardiovascular disease impose a substantial burden on healthcare systems. Electronic health records (EHRs) contain routinely collected clinical information that may facilitate early disease-risk prediction through machine-learning approaches.
Objective: To develop and evaluate machine-learning models for early prediction of major chronic diseases using routinely available EHR data.
Methods: A retrospective predictive analytics study was designed using EHR data from 5,286 adults. Disease-specific cohorts were constructed for type 2 diabetes, hypertension, and cardiovascular disease. Demographic, anthropometric, physiological, biochemical, and clinical variables were analyzed. Logistic regression, Random Forest, Support Vector Machine, and XGBoost models were developed using an 80:20 training-test split and five-fold cross-validation. Performance was assessed using AUROC, accuracy, sensitivity, specificity, F1-score, and calibration. SHAP analysis was used to determine predictor importance.
Results: XGBoost demonstrated the highest predictive performance, achieving AUROCs of 0.91 for diabetes, 0.88 for hypertension, and 0.92 for cardiovascular disease. Important predictors included HbA1c, fasting glucose, BMI, blood pressure, age, LDL-C, renal function, and smoking status.
Conclusion: EHR-based machine-learning models may enable early identification of individuals at increased chronic disease risk and support targeted preventive interventions.
