February 26, 2026
Predictive Modeling of Default with Supervised Machine Learning Algorithms
João Gabriel de Medeiros Luz Pedro; Carlos Nabil Ghobril
Summary prepared by the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege focused on synthesis and writing.
The objective of this research was to analyze and process a dataset from the Central Bank of Brazil (Bacen) to implement, train, and validate Machine Learning models, using the supervised algorithms CatBoost, Random Forest, and XGBoost, in order to measure predictive capacity in identifying default. The complexity of the Brazilian economic scenario, with the increase in family indebtedness, requires financial institutions to have precise tools for credit risk management. The ability to anticipate consumer payment behavior is a pillar for the sector’s sustainability, and the application of data science techniques emerges as a solution to improve the accuracy of analyses.
The landscape of indebtedness in Brazil justifies the urgency of studies in the area. Data from the Survey of Indebtedness and Consumer Default (PEIC) of April 2024 show that 78.5% of Brazilian families had debts, and 12.1% admitted they could not honor their commitments (CNC, 2024). A survey by CNDL and SPC Brasil registered 68.62 million Brazilians in default, corresponding to 41.51% of the adult population (Poder360, 2024). These numbers signal a systemic risk that can affect the stability of financial institutions. A practice that contributes to indebtedness is the extension of payment terms, which compromises future income and increases the probability of default in scenarios of instability (CNC, 2024). Therefore, traditional credit analysis, based on static rules, is insufficient. The need to improve risk analysis mechanisms is imperative to mitigate losses.
This research proposes a machine learning-based approach for default prediction. The use of supervised algorithms allows for the construction of models that learn patterns from historical data, identifying non-linear relationships that would be difficult to detect by conventional statistical methods (Kotsiantis, 2007). One of the technical challenges faced was the presence of a large number of categorical variables, such as ‘occupation’, credit ‘modality’, and ‘uf’, which are crucial for characterizing the customer profile but require specific treatments to be used by algorithms like Random Forest and XGBoost. The choice of preprocessing techniques and feature selection were fundamental steps to ensure the generation of robust predictive models.
The scope of the work covered a complete cycle of a data science project: data extraction and processing, exploratory analysis, training, evaluation, and comparison of multiple Machine Learning models. The validation of the results was carried out through metrics such as accuracy, F1-Score, and the area under the ROC curve (AUC-ROC), ensuring an objective evaluation of each algorithm’s performance. The study sought to identify the most effective model for the proposed scenario, providing a basis for the implementation of decision support systems in financial institutions.
The methodology began with the collection of historical credit data from the Credit Information System (SCR) of Bacen (Central Bank of Brazil, 2024). The initial dataset comprised 12 CSV files, corresponding to the months of 2024. For manipulation and processing, the Apache Spark framework was used, which allowed for the unification and cleaning of the files, including the removal of malformed rows and the specification of the column separator and encoding (UTF-8). The preprocessing step is fundamental, as the effectiveness of algorithms depends on the quality of the input data (Fávero & Belfiore, 2024). After cleaning, numerical formats were standardized, replacing commas with periods as decimal separators and removing thousands separators. The consolidated dataset was stored in the Apache Parquet format, optimized for read operations. The technological infrastructure included the Python language in the Spyder IDE and libraries such as Pandas, Matplotlib, Seaborn, and the frameworks CatBoost, Random Forest, and XGBoost.
Exploratory Data Analysis (EDA) was conducted to understand the dataset structure, composed of 13 categorical and 12 numerical variables. The target variable, overdueaboveof15days, was transformed into a binary variable (0 for compliant, 1 for non-compliant), framing the problem as binary classification. Correlation analysis was performed on the complete dataset and on a subset focused on Individual Customers (PF), as Corporate Customers (PJ) have distinct profiles. The analysis revealed that the proportion of non-compliance was higher in the PF group. Graphical visualizations explored the relationship between non-compliance and variables such as State (UF), indexer, and credit modality, confirming the relevance of these features.
The selection of algorithms was based on their effectiveness in classification problems. Random Forest was chosen for its robustness against overfitting, achieved by aggregating decision trees (Breiman, 2001). XGBoost was selected for its high performance and speed, being an optimized implementation of gradient boosting (Chen & Guestrin, 2016; Friedman, 2001). CatBoost was included for its native ability to handle categorical variables, simplifying preprocessing (Prokhorenkova et al., 2018). Performance evaluation used Accuracy, Precision, Recall, and F1-Score. Additionally, the ROC curve and AUC-ROC calculation were used to assess the discriminative capacity of the models (Fawcett, 2006). To ensure robustness, k-fold cross-validation was employed, which mitigates the risk of performance being dependent on a specific train-test split (Fávero & Belfiore, 2024).
The obtained results revealed distinct performances. The CatBoost Classifier demonstrated superior performance, requiring minimal hyperparameter tuning and simplified preprocessing. Its main advantage was natively handling categorical variables. The initial confusion matrix showed a high accuracy rate, and the AUC-ROC curve presented values close to 1.0 for training and testing, with an AUC of 0.99, signaling excellent discrimination capability and absence of overfitting. The variable importance analysis indicated that the feature carterainadimplidaarrastada had the highest significance. However, its high correlation with the target variable could lead to conceptual data leakage. To investigate this effect, a second CatBoost training was performed removing this variable. The results showed a slight drop in performance, but the variable importance plot became more distributed, with other features like ativo_problematico and modalidade gaining relevance. K-fold cross-validation confirmed the stability and robustness of the results in both versions of the model.
The Random Forest Classifier presented satisfactory results, although inferior to CatBoost. The implementation required more intensive preprocessing, with the removal of variables with high multicollinearity and the transformation of categorical variables via one-hot encoding. This process increased the dataset’s dimensionality, resulting in a longer training time. The confusion matrix and the AUC-ROC curve of Random Forest indicated good predictive power, but with AUC values slightly lower than CatBoost’s. The XGBoost Classifier, using the same preprocessing as Random Forest, demonstrated remarkable performance, outperforming Random Forest in all metrics and approaching CatBoost. Its most evident advantage was processing speed, significantly faster than Random Forest. The XGBoost AUC-ROC curve also showed excellent values, and the model’s robustness was confirmed by cross-validation.
In comparative analysis, CatBoost stood out as the algorithm with the best overall performance. Its advantage lies in the combination of high accuracy with simplicity in the preprocessing of categorical variables. XGBoost emerged as a strong alternative, with performance almost as good as CatBoost’s and superior speed to Random Forest. Random Forest, although robust, proved less suitable for this scenario, due to its sensitivity to the preprocessing of categorical variables and higher computational cost. The discussion of the results shows that the choice of algorithm should adapt to the characteristics of the data and the problem requirements.
The superiority of gradient boosting-based algorithms (XGBoost and CatBoost) over the bagging method (Random Forest) suggests that the sequential construction of trees; each new tree corrects the errors of the previous one, was more effective in capturing default patterns. CatBoost’s performance reinforces the importance of algorithms with optimized mechanisms for categorical features, a common challenge in financial datasets. The need to apply one-hot encoding for Random Forest and XGBoost increased computational complexity and may have diluted the informative power of some variables. The importance of variables revealed by the models offers insights for risk management. Features such as overduedelinquent portfolio and problematic_asset were the strongest predictors, but the models’ ability to extract power from other variables, such as credit modality and indexer, demonstrates the potential of Machine Learning. The identification that certain credit modalities are more associated with default may allow financial institutions to adjust their policies more granularly.
The robustness of the models, confirmed by cross-validation, is crucial for their applicability in the real world. The consistency of CatBoost and XGBoost performance across the different validation folds indicates that the models have learned generalizable patterns, not just noise from the training set. This provides greater reliability to the results and confidence for their implementation in production systems. The research, therefore, validates a rigorous methodological process that can be replicated by financial institutions to develop their own predictive modeling solutions.
This work demonstrated the feasibility and effectiveness of applying Machine Learning models for default prediction using data from the Central Bank of Brazil. The process ranged from rigorous data treatment to the implementation and comparison of three supervised algorithms. Exploratory analysis was fundamental in guiding specific treatments, such as focusing on individual clients. The comparison between models revealed that, for the analyzed dataset, the CatBoost algorithm achieved the best overall results. Its ability to natively handle categorical variables simplified preprocessing and resulted in a model with excellent accuracy and reliability. XGBoost also showed exceptional performance, positioning itself as a competitive alternative. Random Forest, despite generating acceptable results, was the least performant and the most costly in processing time. It is concluded that the objective was achieved: it was demonstrated that boosting algorithms, especially CatBoost, combined with adequate data preprocessing, offer a robust and precise solution for default prediction, potentially enhancing credit risk strategies in financial institutions.
References:
Banco Central do Brasil (BACEN). 2024. Available at: <https://dadosabertos. bcb. gov. br/dataset/scrdata>. Accessed on: 03 Apr. 2025.
Breiman, L. (2001). Random Forests. Machine Learning, 45(1), 5–32. doi: 10.1023/A:1010933404324.
Chen, T.; Guestrin, C. 2016. XGBoost: A scalable tree boosting system. 1st ed. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785-794.
(CNC) Confederação Nacional do Comércio de Bens, Serviços e Turismo. 2024. Consumer Indebtedness and Default Survey (Peic) – April 2024. Available at: <https://portaldocomercio. org. br/publicacoesposts/pesquisa-de-endividamento-e-inadimplencia-do-consumidor-peic-abril-de-2024/:>. Accessed on: 03 Apr. 2025.
Fávero, L. P.; Belfiore, P. 2024. Manual de Análise de Dados: estatística e Machine Learning com EXCEL®, SPSS®, STATA®, R® e Python®. 2nd ed. LTC, Rio de Janeiro, RJ, Brazil.
Fawcett, T. (2006). An Introduction to ROC Analysis. Pattern Recognition Letters, 27(8), 861–874. doi: 10.1016/j. patrec.2005.10.010.
Friedman, J. H. (2001). Greedy Function Approximation: A Gradient Boosting Machine. Annals of Statistics, 29(5), 1189–1232. doi: 10.1214/aos/1013203451.
Kotsiantis, S. B. (2007). Supervised Machine Learning: A Review of Classification Techniques. Informatica, 31(3), 249–268.
Poder360. 2024. Indebtedness in Brazil reaches 41.51% of the adult population in November. Available at:<https://www. poder360. com. br/poder-economia/inadimplencia-no-brasil-atinge-4151-da-populacao-adulta-em-novembro/>. Accessed on: 17 Jun. 2025.
Prokhorenkova, L.; Gusev, G.; Vorobev, A.; Dorogush, A. V.; Gulin, A. 2018. CatBoost: unbiased boosting with categorical features. 1st ed. Advances in Neural Information Processing Systems (NeurIPS)., 31.
Executive summary from the Final Course Work of the Specialization in Data Science and Analytics of the MBA USP/Esalq
Learn more about the course; click here: