Executive Summary

Technology

December 10, 2025

Development and evaluation of predictive models for heart disease

Author: Luis Enrique Icart Maciel — Advisor: Fábio Lima

Summary prepared by the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege focused on synthesis and writing.

This study sought to develop and evaluate predictive models for heart disease using open data, focusing on identifying the most influential variables. Build machine learning algorithms with robust performance and high interpretability, allowing for clear identification of determinant risk factors. The approach was to create a precise and transparent clinical decision support tool, useful for early patient screening, especially in resource-limited contexts. The emphasis on explainability sought to overcome the “black box” barrier of artificial intelligence, promoting trust and adoption by healthcare professionals.

Cardiovascular diseases (CVD) are the leading cause of global mortality (Islam and Majumder, 2013), with 17.9 million deaths annually (World Health Organization, 2021). The spectrum includes coronary heart disease, cerebrovascular disease, and peripheral arterial disease. Behavioral risk factors such as inadequate diet, sedentary lifestyle, alcohol consumption, and tobacco use are recognized catalysts for the development of these pathologies.

In Brazil, cardiovascular diseases are the leading cause of death, accounting for 20% of deaths in individuals over 30 years old, with higher incidence in the South and Southeast regions (Mansur and Favarato, 2016). The situation is aggravated in low- and middle-income countries, which concentrate more than three-quarters of global CVD deaths, reflecting disparities in access to care (Coffey et al., 2021). Population growth and the prevalence of risk factors are pressuring health systems, making accessible diagnosis a logistical and financial challenge (Dutta et al., 2020).

In this scenario, machine learning emerges as an approach for a more proactive and personalized health model (Sarker, 2024). The algorithms’ ability to analyze large volumes of data and identify complex patterns can enhance cardiovascular risk prediction. The digitalization of the medical sector has facilitated the collection of clinical information, creating a conducive environment for predictive models that optimize clinical decision-making and resource allocation (Aljanabi et al., 2018). However, for their clinical application, the accuracy of the models must be accompanied by interpretability.

The “black-box” nature of many algorithms is a barrier, as healthcare professionals need to understand the factors behind a prediction to trust it (Lundberg et al., 2018). Explainable models, which justify their predictions, are fundamental for the responsible and ethical use of artificial intelligence in diagnostics (Yin and Bingi, 2023), aligning with this study’s objective to balance performance and transparency.

This research has applied characteristics, with an experimental and quantitative design, suitable for investigating cause-and-effect relationships between risk factors and the presence of heart disease (Anderson-Cook, 2005; Babbie, 2020). The study developed and evaluated predictive models using the open dataset “Heart Disease” from the “UCI Machine Learning Repository” (2007). The choice of this dataset, extensively used in previous research (Ding and Sadeghi, 2019; Dhurandhar et al., 2019; Aljanabi et al., 2018; Wang, 2018), ensured the replicability and comparability of the results.

Data preparation was a critical step. Although the original dataset contained 76 attributes, the version used in this study has 14 main variables, with no missing values. An exploratory analysis identified outliers, especially in weight and height. To mitigate their impact, observations with a body mass index below 48 kg (1st percentile) and a height below 148 cm (1st percentile) were removed. This intervention resulted in a final dataset with 66,833 records and 14 columns, used for modeling.

The target variable was defined as binary: presence (1) or absence (0) of heart disease. Predictor variables included demographic (gender, age), anthropometric (height, weight, BMI), clinical (systolic and diastolic pressure, cholesterol and glucose levels), and behavioral (physical activity, alcohol, smoking) data. For modeling, categorical variables were transformed into dummies (n-1), and continuous ones were standardized via z-score, using training data to avoid information leakage. The data were split into 80% for training and 20% for testing, as per standard practice to evaluate generalization (Catania et al., 2022).

Six algorithms were evaluated: Logistic Regression, Decision Trees, Random Forest, XGBoost, AdaBoost, and LightGBM. Logistic Regression was chosen for being an interpretable linear model (Nasarian et al., 2024), and Decision Trees for modeling non-linear relationships (Valente et al., 2021). Random Forest and the boosting algorithms (XGBoost, AdaBoost, LightGBM) were selected for their high predictive performance in complex problems (Imani et al., 2025; Gao et al., 2023).

Hyperparameter optimization was performed with “Randomized Search” (50 iterations) and 5-fold cross-validation. Performance was evaluated by AUC-ROC and Recall, informative metrics in medical scenarios (Richardson et al., 2023). Interpretability was ensured by the SHAP technique (Lundberg and Lee, 2017). The workflow was implemented in Python with Scikit-learn, XGBoost, LightGBM, and SHAP.

Descriptive analysis revealed a mean BMI of 27.5 (overweight) and systolic (126.5 mmHg) and diastolic (81.3 mmHg) blood pressure averages close to the upper limits of normal, suggesting a population at risk. The correlation matrix confirmed expected associations: age showed the strongest positive correlation with heart disease (r = 0.24), followed by BMI (r = 0.19) and weight (r = 0.17). Systolic and diastolic blood pressure showed a strong correlation with each other (r = 0.73). Systolic blood pressure had the highest correlation with the target variable (r = 0.43), followed by diastolic (r = 0.34), indicating that elevated blood pressure levels are strong predictors.

Box plots of systolic and diastolic pressures showed a clear separation between groups with and without heart disease, with diagnosed individuals presenting higher median values and distributions. Analysis of categorical variables, via chi-square test, indicated statistically significant associations (p < 0.05) between the disease and levels of blood pressure, cholesterol, glucose, smoking, alcohol, and physical activity. Individuals with hypertension level 2 showed a prevalence of 80.1% for heart disease, and those with cholesterol well above normal, 76.1%. Physical inactivity was also associated with a higher percentage of disease (53.3%) compared to active individuals (48.6%).

The comparative evaluation of the six algorithms showed that, on the training set, Random Forest achieved the highest AUC-ROC (0.817) and accuracy (0.744). However, on the test set, the XGBoost model stood out with the highest AUC-ROC (0.7996), followed by LightGBM (0.7990) and Random Forest (0.7989). The consistency between training and test results suggests that overfitting was avoided. Based on these results, XGBoost was selected as the final model, offering the best balance between discrimination ability (AUC-ROC) and sensitivity (Recall of 0.6882).

In medical diagnosis, Recall is crucial for minimizing false negatives. The analysis of the density plot of the probabilities predicted by XGBoost revealed a considerable overlap between the group distributions, explaining the difficulty in separating them perfectly and justifying the search for optimization of the classification threshold.

To optimize the model for the clinical scenario, where the detection of positives is prioritized, a Precision-Recall curve analysis was performed to adjust the classification threshold. The default threshold of 0.5 was not ideal. The analysis indicated that a threshold of 0.36 represented a strategic balance point, making the model more sensitive. This change increased the Recall to 0.80, meaning the model now identified 80% of patients with the disease. Consequently, Precision was reduced to 0.60, increasing false positives. This trade-off is clinically justifiable, as it is preferable to investigate healthy patients further rather than fail to diagnose a sick patient.

The interpretability of the XGBoost model was investigated with SHAP analysis, clinically validating the findings. The mean impact plot of SHAP values confirmed systolic blood pressure as the variable with the greatest influence, followed by age and cholesterol status well above normal, aligning with established medical knowledge. The SHAP scatter plot detailed how the values of each feature impact the prediction: high values of systolic blood pressure and age increase the probability of disease (positive SHAP values), while physical activity practice demonstrated a protective role (negative SHAP values).

This ability to analyze the prediction at the individual level transforms the model from a “black box” into a decision support tool. It allows the healthcare professional to understand which specific factors contributed to the calculated risk, facilitating communication with the patient and the development of personalized intervention plans, aligning the power of AI with clinical reasoning.

The comparison with other studies reveals that techniques such as Multilayer Perceptron (MLP) or combinations of Machine Learning with Deep Learning have reported superior metrics, with AUC-ROC of 0.95 and accuracies above 94% (Bhatt et al., 2023; Bharti et al., 2021). Although these studies used similar variables, differences in datasets, populations, and feature engineering strategies may explain the disparity. The performance of the model developed here, although lower, is robust and clinically relevant. The focus on easily obtainable and highly interpretable variables gives it a distinct practical value, especially for primary screening.

In a scenario where cardiovascular diseases overload healthcare systems, machine learning predictive models are a promising strategy. The result demonstrated the feasibility of developing a robust and interpretable XGBoost model, using easily obtainable variables such as age, blood pressure, and BMI.

The choice of these variables increases the potential for application in contexts with limited resources, such as primary care. SHAP analysis validated the model’s logic by highlighting clinically established predictors. The strategic decision to adjust the classification threshold to 0.36 raised the Recall to 0.80, adapting the model to clinical needs where sensitivity is a priority, minimizing the chance of an ill patient not being identified.

Despite limitations such as the use of an international database and the absence of more detailed clinical variables, the study fulfilled its purposes, demonstrating that it is possible to develop and evaluate an explainable machine learning model for the prediction of heart disease, which offers a strategic balance between performance and sensitivity, with direct applicability in outpatient screening scenarios.

References
Aljanabi, M.;Qutqut, M. H.;Hijjawi, M. 2018. Machine learning classification techniques for heart disease prediction: a review. International Journal of Engineering & Technology, 7(4): 5373–5379.
Anderson-Cook, C. M. 2005. Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Journal of the American Statistical Association, 100(470): 708–708.
Babbie, E. R. 2020. The practice of social research. Cengage Au.
Bharti, R.;Khamparia, A.;Shabaz, M.;Dhiman, G.;Pande, S.;Singh, P. 2021. Prediction of Heart Disease Using a Combination of Machine Learning and Deep Learning. Computational Intelligence and Neuroscience, 2021(1): 8387680.
Bhatt, C. M.;Patel, P.;Ghetia, T.;Mazzeo, P. L. 2023. Effective Heart Disease Prediction Using Machine Learning Techniques. Algorithms, 16(2): 88.
Catania, C.;Guerra, J.;Romero, J. M.;Caffaratti, G.;Marchetta, M. 2022. Beyond Random Split for Assessing Statistical Model Performance.
Coffey, S.;Roberts-Thomson, R.;Brown, A.;Carapetis, J.;Chen, M.;Enriquez-Sarano, M.;Zühlke, L.;Prendergast, B. D. 2021. Global epidemiology of valvular heart disease. Nature Reviews Cardiology, 18(12): 853–864.
Dhurandhar, A.;Shanmugam, K.;Luss, R. 2019. Leveraging Simple Model Predictions for Enhancing its Performance. ArXiv.
Ding, N.;Sadeghi, P. 2019. A Submodularity-based Agglomerative Clustering Algorithm for the Privacy Funnel. ArXiv.
Dutta, A.;Batabyal, T.;Basu, M.;Acton, S. T. 2020. An efficient convolutional neural network for coronary heart disease prediction. Expert Systems with Applications, 159: 113408.
Gao, X.;Alam, S.;Shi, P.;Dexter, F.;Kong, N. 2023. Interpretable machine learning models for hospital readmission prediction: a two-step extracted regression tree approach. BMC Medical Informatics and Decision Making, 23(1): 104.
Imani, M.;Beikmohammadi, A.;Arabnia, H. R. 2025. Comprehensive Analysis of Random Forest and XGBoost Performance with SMOTE, ADASYN, and GNUS Under Varying Imbalance Levels. Technologies, 13(3): 88.
Islam, A. K. M. M.;Majumder, A. A. S. 2013. Coronary artery disease in Bangladesh: A review. Indian Heart Journal, 65(4): 424–435.
Lundberg, S. M.;Lee, S.-I. 2017. A Unified Approach to Interpreting Model Predictions. Em Advances in Neural Information Processing Systems. Curran Associates, Inc.
Lundberg, S. M.;Nair, B.;Vavilala, M. S.;Horibe, M.;Eisses, M. J.;Adams, T.;Liston, D. E.;Low, D. K.-W.;Newman, S.-F.;Kim, J.;Lee, S.-I. 2018. Explainable machine-learning predictions for the prevention of hypoxaemia during surgery. Nature Biomedical Engineering, 2(10): 749–760.
Mansur, A. de P.;Favarato, D. 2016. Tendências da Taxa de Mortalidade por Doenças Cardiovasculares no Brasil, 1980-2012. Arquivos Brasileiros de Cardiologia, 107: 20–25.
Nasarian, E.;Alizadehsani, R.;Acharya, U. R.;Tsui, K.-L. 2024. Designing interpretable ML system to enhance trust in healthcare: A systematic review to proposed responsible clinician-AI-collaboration framework. Information Fusion, 108: 102412.
Richardson, E.;Trevizani, R.;Greenbaum, J. A.;Carter, H.;Nielsen, M.;Peters, B. 2023. The ROC-AUC accurately assesses imbalanced datasets. Available at SSRN 4655233.
Sarker, M. 2024. Revolutionizing Healthcare: The Role of Machine Learning in the Health Sector. Journal of Artificial Intelligence General science (JAIGS) ISSN:3006-4023, 2(1): 36–61.
UCI Machine Learning Repository, U. M. L. R. 2007. Heart Disease.
Valente, F.;Henriques, J.;Paredes, S.;Rocha, T.;Carvalho, P. de;Morais, J. 2021. Improving the compromise between accuracy, interpretability and personalization of rule-based machine learning in medical problems.
Wang, T. 2018. Hybrid Decision Making: When Interpretable Models Collaborate With Black-Box Models. ArXiv.
World Health Organization, W. H. O. 2021. Cardiovascular diseases.
Yin, Y.;Bingi, Y. 2023. Using Machine Learning to Classify Human Fetal Health and Analyze Feature Importance. BioMedInformatics, 3(2): 280–298.


Executive summary from the Final Project of the Specialization in Data Science and Analytics of the MBA USP/Esalq

Learn more about the course; click here:

Who edited this article

Most recent

You may also like