January 16, 2026
Comparative analysis of Machine Learning techniques for default prediction
Author: Luciane Berger da Silva — Advisor: Daniel Alvarez Firmino
Summary prepared by the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege focused on synthesis and writing.
This study analyzes and compares the performance of machine learning algorithms in credit default prediction. The research compares linear models (logistic regression) and tree-based models (decision tree, random forest), investigating the effect of preprocessing strategies, such as missing value imputation and class balancing. The objective is to identify the most suitable models, considering performance metrics and practical applicability for financial institutions, especially in imbalanced data scenarios, characteristic of credit risk.
Credit granting is fundamental for the economy, but it involves the risk of default, the non-fulfillment of contractual obligations by the borrower (Baesens et al., 2016). Effective management of this risk is crucial for the sustainability of financial institutions. Predictive analysis tools are indispensable for identifying risk patterns based on historical data, allowing for preventive measures such as adjusting limits or refusing high-risk operations (Lessmann et al., 2015; Pinto et al., 2024).
The literature on credit risk transitions from traditional statistical models to machine learning algorithms. Tree-based methods, such as random forest, demonstrate better performance in capturing non-linear interactions (Aniceto et al., 2020). However, logistic regression maintains its prominence in the financial sector due to its interpretability and regulatory acceptance (Bücker et al., 2020). The challenge is to reconcile the predictive power of advanced models with the need for transparency, motivating research into hybrid approaches (Dumitrescu et al., 2022).
Despite the advances, gaps persist in the application of these techniques. Many studies do not explore the impact of treating imbalanced data, a common characteristic in default databases, where the performing class is the majority. Imbalance can induce algorithms to a bias in favor of the majority class, resulting in models with high overall accuracy but low capacity to identify default cases. This work contributes to the literature by offering a systematic comparative analysis that evaluates different algorithms and investigates how preprocessing strategies, especially class balancing ones, influence predictive performance. By analyzing the results under multiple metrics, the study aims to provide practical insights for financial institutions on how to select and configure models that align with their objectives, whether maximizing performance or minimizing losses from false negatives.
The methodology is an applied quantitative research, which uses supervised machine learning techniques to compare predictive approaches in identifying customers at risk of default. The public database “Give me Some Credit” (Freshcorn, 2011), with 150,000 anonymized records and behavioral variables, was used. Although the North American origin of the database limits generalization to the Brazilian context, its volume and use in academic studies make it a robust environment for validating modeling techniques in imbalanced data scenarios.
Data preprocessing began with the analysis of missing values, present in the income (19.82%) and number of dependents (2.62%) variables. A chi-squared test indicated that the absence was not random. Four treatment strategies were evaluated: record exclusion; median imputation; K-Nearest Neighbors (KNN) imputation; and Multiple Imputation by Chained Equations (MICE). Each approach generated a version of the database for comparative analysis of the impact of missing data treatment on model performance.
The imbalance of the target variable, with only 7% of defaulters, was another challenge. To mitigate the risk of bias in favor of the majority class, three balancing strategies were tested: none (baseline); undersampling, which reduces the majority class; and synthetic minority oversampling technique (SMOTE), which generates synthetic observations of the minority class (He and Garcia, 2009). The comparison evaluated the impact of each on metrics sensitive to imbalance, such as sensitivity and F1-score.
The database was divided into 70% for training and 30% for testing, in a stratified manner. Three classification algorithms were implemented (Kotsiantis, 2007): logistic regression (interpretability), decision tree (non-linear patterns) (Grus, 2021), and random forest (robustness and generalization) (Breiman, 2001). Hyperparameters were optimized with RandomizedSearchCV, GridSearchCV, and 5-fold cross-validation. Performance was evaluated on the test set with a comprehensive set of metrics: Area Under the ROC Curve (AUC-ROC), Area Under the Precision-Recall Curve (AUC-PR), Gini index, accuracy, precision, recall, F1-score, Matthews Correlation Coefficient (MCC), and Brier score, ensuring a multidimensional analysis.
The logistic regression results showed that L2 regularization (Ridge) was the best configuration, which is expected in scenarios with multicollinearity (Hastie et al., 2009). Maintaining all variables was justified by the pursuit of global predictive performance (Fávero and Belfiore, 2017). The strategy of excluding missing records, without balancing or with undersampling, produced the best overall results, with an AUC-ROC of 0.803 and a sensitivity of 0.627. Imputation techniques (median, KNN, MICE) had a marginal impact. The use of SMOTE increased sensitivity to 0.670 but harmed accuracy and precision. However, the combination of MICE with SMOTE achieved the best Brier score (0.168), indicating superior calibration. This highlights a trade-off: data exclusion favors overall performance, while MICE and SMOTE may be preferable when reducing false negatives and calibration are priorities.
The decision tree model, without balancing, exhibited high accuracy (0.933) but null sensitivity, failing to identify defaulters, a documented phenomenon (Breiman et al., 1984). The use of SMOTE promoted gains, raising sensitivity to approximately 0.30. The combination of median imputation with SMOTE presented the best balance, with an AUC-ROC of 0.836 and a Brier score of 0.057. Undersampling, on the other hand, maximized sensitivity (close to 0.77), but with a significant drop in accuracy and a worsening in calibration, reinforcing the trade-off between detecting defaulters and maintaining overall robustness (Thomas et al., 2017).
The random forest was the most robust algorithm, with consistent AUC-ROC around 0.86. For global performance, the combination of median imputation without balancing stood out, with AUC-ROC of 0.861, AUC-PR of 0.384, and Brier score of 0.126. To minimize false negatives, median imputation with undersampling was the most effective, achieving the study’s highest sensitivity (0.795) and the best AUC-ROC (0.864), albeit with losses in accuracy. In contrast, strategies with SMOTE maximized accuracy but reduced sensitivity.
The final comparison, using AUC-PR as the main criterion due to the imbalance, confirmed the superiority of the random forest for overall performance. This model presented the best balance, with the highest AUC-PR (0.384), sensitivity of 0.418, and the lowest Brier score (0.126), indicating excellent discrimination and calibration. Logistic regression outperformed the decision tree in sensitivity, but with losses in other metrics, corroborating studies that point to ensembles as more stable (Isidoros and Arcozzi, 2024). In the analysis focused on minimizing false negatives (sensitivity), tree-based models were superior. The random forest led with a sensitivity of 0.795, followed by the decision tree with 0.779. Both incurred costs in precision and calibration, exemplifying the trade-off in risk management: reducing false negatives mitigates capital losses, while an increase in false positives can lead to the loss of business opportunities.
The relevance of variables and the performance of models must be contextualized. The significance of traditional variables may decrease in economic shock scenarios (Gambacorta et al., 2024), and performance can be improved with alternative data, such as psychometric information, especially for populations with limited credit history (Djeundje et al., 2021). The choice of a model should not be based solely on statistical metrics. Models with high accuracy do not always minimize the regulatory and capital costs associated with errors (Xia et al., 2022). The final decision should integrate statistical analysis with the institution’s strategic objectives, considering the financial costs of each type of error. This study provides a framework for this analysis, demonstrating how different combinations of algorithms and preprocessing align with different risk appetites.
The work demonstrated that there is no universally superior model for default prediction, but rather trade-offs between performance, interpretability, and operational impact. Logistic regression proved to be a competitive and transparent benchmark. Decision trees, sensitive to imbalance, improved with resampling techniques. Random forest was the most balanced and robust model, with the best overall performance in discrimination (AUC-ROC ≈ 0.86) and calibration. The analysis highlighted the critical role of preprocessing: undersampling was more effective in increasing sensitivity, while SMOTE, in some scenarios, favored accuracy, emphasizing the need to align the methodology with strategic objectives.
The study’s limitations include the use of a single North American database and the non-exploration of algorithms such as gradient boosting or neural networks. Future research can expand the analysis to Brazilian databases, incorporate other models, and integrate error cost metrics to quantify the financial impact. It is concluded that the objective was achieved: it was demonstrated that the choice of machine learning algorithm and data preprocessing strategies for default prediction fundamentally depends on the financial institution’s strategic objectives, highlighting a clear trade-off between overall predictive performance and the minimization of losses due to false negatives.
References:
Aniceto, G. F.; Barboza, F. L. C.; Kimura, H. 2020. Credit risk analysis using machine learning classifiers. Brazilian Review of Finance 18(4): 1–28.
Baesens, B.; Roesch, D.; Scheule, H. 2016. Credit risk analytics: measurement techniques, applications and examples in SAS. John Wiley & Sons, Hoboken, NJ, USA.
Breiman, L. 1984. Classification and regression trees. Chapman & Hall/CRC, Boca Raton, FL, USA.
Breiman, L. 2001. Random forests. Machine Learning 45(1): 5–32.
Bücker, M.; Szepannek, G.; Gosiewska, A.; Biecek, P. 2020. Transparency, auditability and explainability of machine learning models in credit scoring. Journal of the Operational Research Society 71(8): 1281–1290.
Carvalho, J. R. 2015. Credit risk analysis: fundamentals, methodologies and applications. Atlas, São Paulo, SP, Brazil.
Djeundje, V . B.; Crook, J.; Calabrese, R.; Hamid, M. 2021. Enhancing credit scoring with alternative data. Expert Systems with Applications 167: 113766.
Dumitrescu, E.; Hué, S.; Hurlin, C.; Tokpavi, S. 2022. Machine learning for credit scoring: improving logistic regression with non-linear decision-tree effects. European Journal of Operational Research 297(3): 1178–1192.
Fávero, L. P.; Belfiore, P. P. 2017. Data analysis manual: statistics and multivariate modeling with Excel®, SPSS® and Stata®. Elsevier, Rio de Janeiro, RJ, Brazil.
Fávero, L. P.; Belfiore, P. P. 2024. Data analysis: exploratory and confirmatory multivariate techniques. Elsevier, Rio de Janeiro, RJ, Brazil.
Freshcorn, B. 2011. Give Me Some Credit: 2011 Competition Data. Available at: https://www.kaggle.com/datasets/brycecf/give-me-some-credit-dataset.
Gambacorta, L.; Huang, Y.; Qiu, H.; Wang, J. 2024. How do machine learning and non-traditional data affect credit scoring? New evidence from a Chinese fintech firm. Journal of Financial Stability 73: 101284.
Grus, J. 2021. Data science from scratch: fundamental concepts with Python. 2nd ed. Alta Books, Rio de Janeiro, RJ, Brazil.
Hastie, T.; Tibshirani, R.; Friedman, J. 2009. The elements of statistical learning: data mining, inference, and prediction. 2nd ed. Springer, New York, NY, USA.
He, H.; Garcia, E. A. 2009. Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering 21(9): 1263–1284.
Hosmer, D. W.; Lemeshow, S.; Sturdivant, R. X. 2013. Applied logistic regression. 3rd ed. John Wiley & Sons, Hoboken, NJ, USA.
ISIDOROS, I.; ARCOZZI, N. 2024. Improved convergence rates for some kernel random forest algorithms. Mathematics in Engineering 6(2): 1-22.
Kotsiantis, S. B. 2007. Supervised machine learning: a review of classification techniques. Informatica 31(3): 249–268.
Lessmann, S.; Baesens, B.; Seow, H. V.; Thomas, L. C. 2015. Benchmarking state-of-the-art classification algorithms for credit scoring: an update of research. European Journal of Operational Research 247(1): 124–136.
Pinto, R. S.; Ywata, A.; Tessmann, R. H.; Lima, F. 2024. Are machine learning models more effective than logistic regressions in predicting bank credit risk? An assessment of the Brazilian financial markets. International Journal of Monetary Economics and Finance 17(1): 1–22.
Thomas, L. C.; Crook, J. N.; Edelman, D. B. 2017. Credit scoring and its applications. 2nd ed. SIAM, Philadelphia, PA, USA.
Xia, Y.; Zhang, C.; Li, Y.; Chen, W. 2022. A comparative study of credit scoring models. Knowledge-Based Systems 235: 107629.
Zou, H.; Hastie, T. 2005. Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 67(2): 301–320.
Executive summary from the Final Course Work of the Specialization in Data Science and Analytics of the MBA USP/Esalq
Learn more about the course; click here: