Article

October 02, 2026

Classification of defaulting customers using supervised machine learning techniques

Classification of Defaulted Customers Using Supervised Machine Learning Techniques

Felipe Ricardo Huergo Cagol; Ugo Henrique Pereira da Silva

DOI: 10.22167/2675-6528-202602895

Article derived from a Final Course Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.

Summary

The risk of default in credit operations demanded analytical approaches to anticipate losses. This study comparatively evaluated the performance of supervised machine learning models in classifying defaulting customers in credit card operations. The public dataset “Default of Credit Card Clients” from the University of California Irvine was used, with 30,000 observations and class imbalance. The algorithms Logistic Regression, Random Forest, and Extreme Gradient Boosting were employed. The imbalance was addressed by assigning weights to the classes, and model optimization occurred with the RandomizedSearchCV method, prioritizing sensitivity. Cross-validation results indicated that the Extreme Gradient Boosting model presented a higher capacity for identifying the defaulting class and better discriminatory performance, followed by Random Forest and Logistic Regression, with a sensitivity of 0.8250 and an AUC-ROC of 0.7844 for XGBoost. Interpretability analysis, conducted by the Shapley Additive Explanations (SHAP) technique, highlighted the predominance of variables associated with payment behavior, especially the history of delays. It was concluded that tree-based models, particularly boosting techniques, proved to be more suitable for capturing complex patterns in the data, configuring themselves as consistent alternatives for credit risk management.

Keywords: Machine Learning; Credit Card; Classification; Extreme Gradient Boosting; Credit Risk.

1. Introduction

The granting of credit limits in credit card operations represents a strategic decision that directly impacts the financial risk of institutions. Unlike other forms of credit, card operations do not have real guarantees or automatic payment withholding mechanisms, as observed in payroll-deducted credit. Consequently, the options for recovering defaulted amounts are more restricted, often limited to including the customer in credit protection registries.

This market segment is characterized by high competition and limited revenue sources, mainly from card usage fees and revolving credit interest in cases of default. This scenario creates an inherent trade-off between the rigor in credit granting and the need to expand market scale. Therefore, the relevance of quantitative approaches in supporting decision-making in credit risk management is accentuated, aiming to balance delinquency control with the financial return desired by the entity.

Historically, credit risk management has relied on traditional statistical methods, such as Logistic Regression. The choice of these models is justified by their interpretability and low computational cost, as well as being aligned with the requirements for provisioning for doubtful debt (PCLD) and the guidelines of the Basel Committee (BCB, 2021; BCBS, 2004; Hand and Henley, 1997). However, the advancement of computational capacity has driven the adoption of machine learning algorithms in credit data mining, increasing prediction accuracy. Recent studies demonstrate the effectiveness of these models in predicting default, particularly techniques based on gradient boosting (Aarfi et al., 2024; Bentéjac et al., 2021).

Despite the advances, the literature points to methodological limitations arising from the natural imbalance of credit portfolios. In these contexts, prioritizing global accuracy can lead to misleading evaluations, as algorithms tend to neglect the minority class of bad payers (Brown and Mues, 2012). To mitigate this problem, approaches such as the creation of synthetic data, through the Synthetic Minority Over-sampling Technique (SMOTE), are commonly employed (Chawla et al., 2002). However, studies warn that the artificial generation of samples can introduce limitations related to data representativeness, impacting the model’s generalization capability (Marqués et al., 2012).

Additionally, it is crucial to deepen the analysis of the financial impact of prediction errors. In real-world scenarios, the cost of a false negative, which occurs when granting credit to a bad payer, is higher than the cost of a false positive, which corresponds to denying credit to a good payer (Islam et al., 2018; Ling and Sheng, 2011). This is because the financial losses generated by a defaulting customer require the constitution of PCLD, affecting the banking result and Regulatory Capital. Given the need to align predictive models with the risk appetite of financial institutions and to consider the distinct cost of each type of error, this work seeks to analyze the trade-off between precision and sensitivity in granting credit card limits, as well as to evaluate whether ensemble models outperform traditional approaches in handling data imbalance.

Thus, this study aims to analytically investigate the predictive performance and decision mechanics of three supervised learning techniques, Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost), applying cost-sensitive methodologies on a public dataset from the University of California (UCI). It seeks to evaluate how different algorithms behave in the face of structural class imbalance and how adjusting the decision threshold can mitigate losses. Complementarily, it aims to identify factors associated with default through the use of the Shapley Additive Explanations (SHAP) technique, which allows measuring the contribution of each variable to the model’s predictions, as well as the direction and magnitude of its impact on the final outcome.

2. Material and Methods

This study was characterized as an exploratory and quantitative research, with an experimental design, using public secondary data. Its objective was to analytically investigate the predictive performance and decision mechanics of three supervised learning techniques, Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost), applying cost-sensitive methodologies. We sought to evaluate the algorithms’ behavior in the face of structural class imbalance and the impact of adjusting the decision threshold on loss mitigation, in addition to identifying factors associated with default through the Shapley Additive Explanations (SHAP) technique.

Data collection was carried out by importing the public dataset “Default of Credit Card Clients”, made available by the University of California Irvine (UCI). This dataset, referring to credit card operations of a financial institution in Taiwan, was collected in the year 2005. The choice for public data occurred due to confidentiality restrictions and guidelines of the General Law for the Protection of Personal Data (LGPD), avoiding the use of proprietary corporate data.

The dataset used comprised 30,000 instances, characterized by an imbalance between the classes of performing and non-performing. The base was composed of 23 explanatory variables and one response variable. The explanatory variables included four demographic attributes and 19 numerical attributes, while the response variable was binary, indicating performance or non-performance in the subsequent month.

The demographic variables considered were sex, education level, marital status, and age. The quantitative variables included the amount of credit granted, the payment status for the last six months (April to September 2005), the invoice amounts during this period, and the respective payments made. The temporal structure of the data allowed for the observation of customer payment behavior over time, which was relevant for inferring the probability of default.

Based on the original variables, four new metrics were constructed to capture different dimensions of customers’ financial behavior. These were: the average credit limit utilization rate (AVG_UTILIZATION_RATE), the average invoice payment index (AVG_PAY_TO_BILL), the debt variation (DEBT_TREND), and the delay frequency index (DELAY_COUNT). Feature engineering on credit card time series is a practice that expands the explanatory capacity of models (Bahnsen et al., 2016; Lessmann et al., 2015).

The data preprocessing step was fundamental, given the presence of categorical variables and the heterogeneity of scale among the numerical attributes. For the categorical variables, the *one-hot encoding* technique was applied, which encoded them into binary vector format. This transformation was applied to the sex, education level, and marital status variables, avoiding erroneous interpretations of order between categories (James et al., 2021).

For the numerical variables, the standardization method for zero mean and unit variance was used. This methodology ensured that all variables were treated equitably, considering the large scale disparity between them, such as the value of the credit granted in relation to other numerical data (Hastie et al., 2009). Standardization was adjusted based only on the training set, and the validation and test sets were transformed with the same parameters.

The database was partitioned into training (80%) and testing (20%) sets, preserving the original class distribution through stratified sampling. Additionally, the training set was subdivided into internal training and internal validation subsets. Internal training was employed for model fitting, and internal validation was used to analyze the impact of the decision threshold on performance metrics.

The models were evaluated through stratified five-fold cross-validation, applied to the internal training set. In each iteration, the model was trained on part of the data and validated on the remaining subset, ensuring the use of all observations. The reported performances corresponded to the average of the results obtained in the five iterations, accompanied by the standard deviation, which allowed for the evaluation of performance stability (Hastie et al., 2009).

The selection of algorithms was based on the purpose of comparing different machine learning approaches. Logistic Regression was adopted as a benchmark model, given its consolidation and interpretability in the banking sector (Baesens et al., 2003). Random Forest (Breiman, 2001) was selected as a representative of the *bagging* strategy, recognized for its variance reduction capability. XGBoost was chosen based on recent literature (Xu, 2024), which demonstrates the superiority of *boosting* algorithms in credit risk assessment (Chen and Guestrin, 2016).

The optimization of the models sought to identify hyperparameter combinations that maximized predictive performance. For this purpose, the RandomizedSearchCV method was employed, which performs random sampling of hyperparameter combinations from predefined search distributions and intervals for each algorithm (Bergstra and Bengio, 2012). One hundred combinations were evaluated for each model, prioritizing the sensitivity metric, given the importance of identifying defaulting customers.

For Logistic Regression, values of the regularization parameter C and optimization methods *liblinear* and *lbfgs* were evaluated. For Random Forest, intervals were considered for the number of trees, maximum depth, minimum number of samples for splitting and per leaf, and variable selection strategies. For XGBoost, values were defined for the number of estimators, learning rate, depth and minimum weight of observations per node, in addition to regularization and sampling parameters (XGBoost Developers, 2024).

The treatment of class imbalance was performed by assigning differentiated weights to the classes in the objective function, without altering the original dataset (Ling and Sheng, 2011). For XGBoost, the *scale_pos_weight* hyperparameter was used, defined as the ratio between the number of observations in the majority class (performing) and the minority class (non-performing), adjusting the contribution of positive class examples during the minimization of the loss function (Chen and Guestrin, 2016).

The metrics used to evaluate model performance were Accuracy, Precision, Sensitivity, F1-Score, and AUC-ROC. In binary classification scenarios with imbalanced datasets, accuracy alone may not be the most suitable measure, being complemented by metrics derived from the confusion matrix to analyze different types of classification errors (James et al., 2021). The AUC-ROC metric quantified the model’s discrimination power between the performing and non-performing classes (Fawcett, 2006).

The Shapley Additive Explanations (SHAP) technique was used to interpret the individual contribution of the explanatory variables in predicting customer default (Lundberg and Lee, 2017). SHAP allowed the model’s output to be decomposed into contributions attributed to each variable, measuring the marginal impact of each and considering their interactions with other factors. This increased model transparency and aided in identifying the determinants of credit risk.

The decision threshold consisted of the cutoff parameter adopted to calibrate the model according to the financial institution’s risk appetite. Different thresholds were tested on the best-performing model to verify the *trade-off* between sensitivity and precision (Fawcett, 2006; James et al., 2021). This threshold represented the cutoff point from which the probabilities estimated by the model defined the classification of clients as defaulters.

For the development of this work, the artificial intelligence tool Google Gemini was used as support in stages such as grammatical review, improvement of textual clarity, support in identifying bibliographic references, verification of formal formatting aspects, and punctual assistance in structuring code snippets and clarifying technical concepts (Bigaton et al., 2025). The developed code was made publicly available to ensure the reproducibility of the study.

3. Results and Discussion

The predictive performance evaluation of machine learning models, Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost), was carried out through stratified cross-validation, considering metrics such as Accuracy, Precision, Sensitivity, F1-Score, and AUC-ROC. The results obtained, both before and after hyperparameter optimization, revealed significant differences among the algorithms, especially in the ability to identify the minority class of defaulting customers. The optimization aimed to refine the models’ performance, adjusting them to the specific characteristics of the dataset and prioritizing sensitivity, a crucial metric in credit risk contexts, where the cost of a false negative is high (Islam et al., 2018).

Predictive performance of the evaluated models

For Logistic Regression, it was observed that hyperparameter optimization resulted in discrete variations in performance metrics. Before optimization, the model presented a sensitivity of 0.5917 and an F1-Score of 0.5318, with an AUC-ROC of 0.7600. After the adjustment process, sensitivity slightly increased to 0.6023 and F1-Score to 0.5337, while AUC-ROC recorded a small reduction to 0.7477. This behavior suggested that, in the context of the experiment, the model operated close to its predictive capacity limit, indicating that expanding the hyperparameter search space did not generate substantial gains, as expected for linear models (James et al., 2021).

The Random Forest model showed more expressive changes after optimization. Sensitivity, which was initially 0.3537, increased significantly to 0.6456, and the F1-Score rose from 0.4594 to 0.5268. However, this improvement in identifying defaulters was accompanied by a reduction in precision and accuracy. The AUC-ROC, in turn, had a slight increase from 0.7691 to 0.7735. This shift in the balance between false positives and false negatives indicated a prioritization of the defaulter class, a desirable behavior in scenarios where minimizing false negatives is a priority, even at the cost of an increase in errors on the non-defaulter class (Brown and Mues, 2012).

The Extreme Gradient Boosting (XGBoost) also showed a relevant increase in sensitivity after optimization, going from 0.6341 to 0.8250. As with Random Forest, this improvement was accompanied by reductions in precision and accuracy, reflecting the prioritization of identifying delinquent customers, aligned with the adopted optimization criterion. The XGBoost AUC-ROC, which was 0.7837 before optimization, had a small increase to 0.7844. The overall results indicated that the greatest performance gains, especially in terms of sensitivity, occurred in tree-based models, which was expected, given that sensitivity was the prioritized metric in the optimization (Chen and Guestrin, 2016).

The evaluation of the variability of metrics across the cross-validation folds indicated stability in the models, with reduced standard deviations relative to the means, suggesting the consistency of the results. In terms of discriminatory capacity, tree-based models outperformed Logistic Regression. The optimized XGBoost model achieved an AUC-ROC of 0.7805, followed by the optimized Random Forest with 0.7618, while the optimized Logistic Regression presented an AUC-ROC of 0.7308. These findings demonstrated the greater ability of ensemble models to distinguish classes across different decision thresholds (Hastie et al., 2009).

The analysis of the impact of the decision threshold on the performance of the optimized XGBoost model revealed a trade-off between sensitivity, precision, and specificity. Reducing the decision threshold resulted in an increase in sensitivity, expanding the ability to identify defaulting customers, but at the cost of a decrease in precision and specificity. For example, by setting the threshold at 0.10, sensitivity reached 0.9981, with a precision of 0.2230 and a specificity of 0.0120, prioritizing the maximum capture of defaulters. In contrast, a threshold of 0.90 reduced sensitivity to 0.2957, but increased precision to 0.6916 and specificity to 0.9625, reflecting a more expansive stance in credit granting (Fawcett, 2006).

The confusion matrix of the optimized XGBoost model, using a decision threshold of 0.50, revealed that 2,420 performing customers were correctly classified, and 1,107 non-performing customers were also correctly identified. However, 2,253 false positives were observed, where performing customers were erroneously classified as non-performing, and 220 false negatives, where non-performing customers were classified as performing. The sensitivity on the test set was approximately 83.6%, a value consistent with cross-validation (82.5%), indicating good generalization capability of the model. The high number of false positives, however, suggests a conservative stance of the model, prioritizing the mitigation of default risk over the expansion of the approved customer base, which implies an opportunity cost for the institution (James et al., 2021).

Importance of variables in explaining default (SHAP) and relationship with exploratory analysis and derived variables

The interpretability analysis, performed using the Shapley Additive Explanations (SHAP) technique, allowed the identification of the most relevant factors in predicting default. The results showed that variables associated with payment behavior, such as delay history and recent payment status, exerted the greatest marginal contribution to the predicted probability of default. In contrast, the exploratory hypotheses related to demographic variables, such as sex, education level, and marital status, were not confirmed in the context of the estimated model, as these attributes showed low predictive contribution, indicating that the patterns observed in the exploratory analysis were not consistently sustained in the model (Lundberg and Lee, 2017).

Specifically, the frequency of delays and the payment status in September 2005 were the variables with the greatest impact on the classification of delinquent customers. The average credit limit utilization rate and the granted credit amount also stood out as important factors. SHAP analysis revealed a negative association between the granted credit limit and delinquency, suggesting that higher limits are associated with a lower probability of non-payment. On the other hand, a positive association was observed between the degree of limit utilization and the risk of delinquency, where higher credit utilization tends to increase the probability of default (Lessmann et al., 2015).

The interdependence between the variables was a relevant finding, with SHAP dependency plots indicating joint effects. Customers with a high credit limit utilization rate and, simultaneously, lower granted credit values, showed a more pronounced positive contribution to default risk. In contrast, for customers with higher credit limits, the impact of the utilization rate proved more moderate, suggesting a possible risk-mitigating effect. This complex interaction underlines that default is not determined by isolated factors, but rather by a combination of multiple dimensions of the customer’s financial behavior (Lundberg and Lee, 2017).

Additionally, the joint analysis of the variables of delay frequency and payment status in September 2005 suggested that the combination of a history of delays and recent unfavorable payment behavior potentiated the risk of default. The recency of the delay acted as an amplifying factor of the payment history, indicating that the temporal proximity of a non-payment event carries greater weight in predicting risk. These results reinforce the relevance of behavioral variables in determining credit risk, demonstrating that default is intrinsically linked to the dynamics of the client’s financial behavior over time (Bahnsen et al., 2016).

Comparative analysis of the results obtained in this research with scientific publications on this topic

The results obtained in this study demonstrated consistency with empirical evidence reported in the literature, especially in works that used similar datasets. It was observed that tree-based models and ensemble techniques, such as Random Forest and XGBoost, tend to present better discriminatory capacity compared to linear approaches, such as Logistic Regression, in credit risk classification (Yeh and Lien, 2009). This pattern was consistently identified in subsequent studies that highlighted the effectiveness of boosting methods in modeling nonlinear relationships and predicting default (Chen and Guestrin, 2016; Lessmann et al., 2015).

However, the direct comparison of these results with other studies should be interpreted with caution due to methodological differences. Firstly, the validation procedure adopted in this study, which combined repeated evaluation on multiple partitions with final validation on an independent set, provided more robust out-of-sample performance estimates. In contrast, some previous studies employed single train and test partitions, which can introduce greater sensitivity to sample specificities. Secondly, hyperparameter optimization through RandomizedSearchCV allowed for efficient exploration of relevant regions of the search space, being particularly suitable for complex algorithms such as Random Forest and XGBoost (Bergstra and Bengio, 2012).

Finally, the treatment of class imbalance through cost-sensitive approaches, which assign differentiated weights to classes in the loss function, differed from studies that used resampling techniques, such as SMOTE (Chawla et al., 2002). Both strategies are discussed in the literature on imbalanced classification and are associated with different trade-offs between bias and variance, as well as distinct impacts on minority class identification, especially in terms of sensitivity and false negative rate (Ling and Sheng, 2011; Brown and Mues, 2012). These methodological differences help contextualize potential variations in reported results, indicating that direct comparisons based solely on aggregated metrics should be interpreted with caution.

Despite the observed consistency and analytical contributions, the present study presented some limitations that should be considered. Firstly, the dataset used refers to information from 2005, which may not fully reflect current credit behavior, given the evolution of the financial market and consumption patterns. Additionally, external validation was not performed on independent databases, which restricts the ability to generalize the results to other contexts or institutions. Finally, other techniques for handling class imbalance, such as oversampling or undersampling, were not evaluated, which could offer additional perspectives on model performance (Marqués et al., 2012).

Given these limitations, future studies can explore the application of the models to more recent databases, which reflect the current economic and behavioral scenario. The incorporation of external validation, through the application of the models to independent databases from different contexts, institutions, or periods, is also a promising avenue to enhance the robustness and generalization capacity of the results. Additionally, the systematic comparison of different imbalance treatment strategies, including resampling approaches, could enrich the understanding of the models’ adherence to real credit granting contexts, contributing to improved risk management.

In summary, the research demonstrated that ensemble-based models, notably Extreme Gradient Boosting, outperformed Logistic Regression in classifying delinquent customers, especially after hyperparameter optimization that prioritized sensitivity. SHAP interpretability analysis revealed that payment behavior and credit utilization are the most relevant predictors, while demographic variables had a lesser impact. The study also highlighted the trade-off between sensitivity and precision in choosing the decision threshold, emphasizing the importance of aligning the modeling strategy with the institution’s risk appetite. These findings reinforce the potential of machine learning techniques to enhance credit risk management, provided they are applied with appropriate methodological decisions.

4. Conclusion

This study analytically investigated the predictive performance and decision mechanics of supervised learning algorithms, Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost), in classifying delinquent customers in credit card operations, considering class imbalance and decision threshold adjustment. It was found that the Extreme Gradient Boosting model exhibited the best discriminatory performance and highest capacity for identifying the delinquent class, with a sensitivity of 0.8250 and an AUC-ROC of 0.7844, outperforming Random Forest and Logistic Regression. Hyperparameter optimization, which prioritized sensitivity, was crucial for these gains in tree-based models. Interpretability analysis, using the Shapley Additive Explanations (SHAP) technique, highlighted that variables associated with payment behavior, such as history of late payments and recent payment status, were the most relevant predictors of delinquency, while demographic attributes had a lesser impact. A negative association was observed between the granted credit limit and delinquency, and a positive association between credit limit utilization and the risk of non-payment. The main contribution lies in the empirical demonstration of the effectiveness of ensemble models, combined with cost-sensitive strategies and decision threshold adjustment, to improve credit risk management in imbalanced data scenarios.

However, the study presented limitations, such as the use of data from 2005, which may not reflect current credit behavior, and the absence of external validation in independent databases, which restricts the generalization of the results. Additionally, other class imbalance treatment techniques, such as oversampling or undersampling, were not comparatively evaluated. For future studies, it is suggested to apply the models to more recent databases and incorporate external validation in different contexts or institutions. It is also recommended to systematically compare various imbalance treatment strategies, including resampling approaches, in order to enrich the understanding of the models’ adherence to real credit granting contexts and broaden the robustness and generalization capacity of the proposed solutions.

Bibliographic References

Aarfi, S.A.; Ahmed, N.; Syed, K.A. 2024. Predicting credit card default using machine learning: an empirical analysis. American Journal of Intelligent Systems. Disponível em: <https://www.researchgate.net/publication/386019597_Predicting_Credit_Card_Default_Using_Machine_Learning_An_Empirical_Analysis>. Acesso em: 18 out. 2025.

Baesens, B.; van Gestel, T.; Viaene, S.; Stepanova, M.; Suykens, J.; Vanthienen, J. 2003. Benchmarking state-of-the-art classification algorithms for credit scoring. Journal of the Operational Research Society 54(6): 627-635.

Bahnsen, A.C.; Aouada, D.; Stojanovic, A.; Ottersten, B. 2016. Feature engineering strategies for credit card fraud detection / risk analysis. Expert Systems with Applications 51: 134-142.

Banco Central do Brasil [BCB]. 2021. Resolução CMN nº 4.966, de 25 de novembro de 2021. Diário Oficial da União, Brasília, 26 nov. 2021. Seção 1, p. 110.

Basel Committee on Banking Supervision [BCBS]. 2004. International Convergence of Capital Measurement and Capital Standards: A revised framework. Disponível em: <https://www.bis.org/publ/bcbs107.htm>. Acesso em: 14 mar. 2026.

Bentéjac, C.; Csörgo, A.; Martínez-Muñoz, G. 2021. A comparative analysis of gradient boosting algorithms. Artificial Intelligence Review. Disponível em: <https://link.springer.com/article/10.1007/s10462-020-09896-5>. Acesso em: 18 out. 2025.

Bergstra, J.; Bengio, Y. 2012. Random search for hyper-parameter optimization. Journal of Machine Learning Research 13: 281-305.

Bigaton, A.; Velázquez, D.R.T.; Santos, G.D.; Belém, M.J.X.; Beltran, M.P. 2025. Manual de Boas Práticas: Uso da inteligência artificial para o desenvolvimento de trabalhos de conclusão de curso. Pecege, Piracicaba, SP, Brasil.

Breiman, L. 2001. Random forests. Machine Learning 45(1): 5-32.

Brown, I.; Mues, C. 2012. An experimental comparison of classification algorithms for imbalanced credit scoring data sets. Expert Systems with Applications 39(3): 3446-3453.

Chawla, N.V.; Bowyer, K.W.; Hall, L.O.; Kegelmeyer, W.P. 2002. SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research 16: 321-357.

Chen, T.; Guestrin, C. 2016. XGBoost: a scalable tree boosting system. In.: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016. Anais… p. 785-794. Disponível em: <https://arxiv.org/pdf/1603.02754.pdf>. Acesso em: 13 dez. 2025.

Fawcett, T. 2006. An introduction to ROC analysis. Pattern Recognition Letters 27(8): 861-874. Disponível em: https://doi.org/10.1016/j.patrec.2005.10.010. Acesso em: 17 jan. 2026.

Hand, D.J.; Henley, W.E. 1997. Statistical classification methods in consumer credit scoring: a review. Journal of the Royal Statistical Society: Series A 160(3): 523-541.

Hastie, T.; Tibshirani, R.; Friedman, J. 2009. The Elements of Statistical Learning: Data mining, inference, and prediction. 2ed. Springer, New York, NY, USA.

Islam, S.R.; Eberle, W.; Ghafoor, S.K. 2018. Credit Default Mining Using Combined Machine Learning and Heuristic Approach. Tennessee Technological University, Cookeville, TN, USA.

James, G.; Witten, D.; Hastie, T.; Tibshirani, R. 2021. An Introduction to Statistical Learning: With applications in R. 2ed. Springer, New York, NY, USA.

Lessmann, S.; Baesens, B.; Seow, H.V.; Thomas, L.C. 2015. Benchmarking state-of-the-art classification algorithms for credit scoring. European Journal of Operational Research 247(1): 124-136.

Ling, C.X.; Sheng, V.S. 2011. Cost-sensitive learning and the class imbalance problem. In.: Sammut, C.; Webb, G.I. (Ed.). Encyclopedia of Machine Learning. Springer, Boston, MA, USA. p. 231-235.

Lundberg, S.M.; Lee, S.I. 2017. A unified approach to interpreting model predictions. In.: Advances in Neural Information Processing Systems, 30., 2017. Anais… p. 4765-4774. Disponível em: <https://arxiv.org/abs/1705.07874>. Acesso em: 28 dez. 2025.

Marqués, A.I.; García, V.; Sánchez, J.S. 2012. On the use of data sampling techniques to cope with imbalanced credit scoring data sets. Journal of the Operational Research Society 63(8): 1099-1111.

XGBoost Developers. 2024. XGBoost Documentation: Parameters. Disponível em: <https://xgboost.readthedocs.io/en/stable/parameter.html>. Acesso em: 21 jan. 2026.

Xu, T. 2024. Comparative analysis of machine learning algorithms for consumer credit risk assessment. Transactions on Computer Science and Intelligent Systems Research 4: 60-67. Disponível em: <https://doi.org/10.62051/r1m3pg16>. Acesso em: 16 nov. 2025.

Yeh, I.C.; Lien, C.H. 2009. The comparisons of data mining techniques for the predictive accuracy of probability of default of credit card clients. Expert Systems with Applications 36(2): 2473-2480.

Article originating from the Final Course Work of the Specialization in Data Science and Analytics of the MBA USP/Esalq

To learn more about the course, click here and access the MBX Academy platform

You may also like

October 02, 2026

Determinants of supermarket location in São Paulo

A study investigated the determining factors for supermarket location in the state of São Paulo, with the objective of investigating the factors that explain the presence and expansion of these establishments, considering socioeconomic, demographic, and market dimensions. Data from the 2010 and 2022 Demographic Censuses of IBGE and information from the National Registry of Legal Entities of the Federal Revenue of Brazil were used to build a georeferenced database. A Random Forest classification model was applied, adjusted by grid search with cross-validation, prioritizing the recall-macro metric due to the imbalance of the dependent variable, which represented the presence or absence of supermarkets within a 50-meter buffer. The results indicated that supermarket location is strongly associated with demographic, income, and population characteristics in the surrounding area. The analysis of variable importance showed that sociodemographic factors, such as elderly literacy, household income, and the presence of other food establishments, exerted significant influence, especially in the immediate vicinity. The findings reinforced the hypothesis that the spatial distribution of supermarkets is not random, being conditioned by socioeconomic characteristics and the commercial structure of the territory, offering subsidies for business decisions and urban planning.

Keywords: Spatial Analysis; Machine learning; Expansion; Commercial location; Supermarkets.

Neuroscience And Learning In Education

October 02, 2026

Anti-Racist Education: Inclusive Educational Practices and Social Development

Antiracist education, understood as a structuring axis of inclusive education and social development, was investigated in the Brazilian context. The study aimed to identify and analyze, based on legal documents and teachers’ perceptions, educational practices capable of promoting antiracism in school and society, and how the implementation of Laws nº 10.639/03 and nº 11.645/08 contributed to social justice. A qualitative and documentary approach was adopted, with analysis of educational legislation, curricular guidelines, institutional reports, and academic literature. Complementarily, a semi-structured questionnaire was applied to 295 Basic Education teachers. The data were evaluated quantitatively and qualitatively, through thematic content analysis, and validated with bibliographic studies. The results revealed a paradox: despite a robust legal framework, the implementation of antiracist policies proved fragile and sporadic, with a lack of teacher training, adequate teaching materials, and monitoring. Significant educational inequalities between white and black students were found to persist, and most teachers acknowledged the occurrence of racism in schools, but without clear institutional protocols. Neuroscientific analysis showed that racism negatively impacts students’ cognitive and emotional development. It was concluded that antiracist education is central to quality education, requiring political commitment, public investment, and intersectoral articulation. The integration of Neuroscience in teacher training and the production of qualified materials are crucial to strengthen the school’s role in building a more just and inclusive society.

Keywords: Social Development; Antiracist Education; Social Justice; Law 10.639/03; Inclusive Educational Practices.

Neuroscience And Learning In Education

October 02, 2026

Paths of Inclusion: Perceptions of Parents and Teachers on the Schooling of Students with Dual Exceptionality in the Brazilian Context

Dual Exceptionality, characterized by the coexistence of High Abilities/Giftedness and neurodevelopmental disorders, represents a complex phenomenon that challenges traditional identification and schooling models. The study aimed to understand the perceptions of parents or guardians, teachers, and other education professionals regarding the schooling of students with Dual Exceptionality in the Brazilian context, investigating challenges, pedagogical strategies, and possibilities for inclusion based on equity. The research adopted a qualitative, exploratory, and descriptive approach, and collected data through an online, voluntary, and anonymous questionnaire answered by 25 participants. Discursive data were analyzed using thematic content analysis. The results indicated that knowledge about the topic is often built from personal and professional experiences, revealing gaps in systematic training. Difficulties were identified in identifying these students, in teacher training, and in implementing individualized educational plans, pedagogical flexibility, and curriculum enrichment. Socio-emotional repercussions, such as frustration and low self-esteem, were reported. However, some schools demonstrated inclusive practices based on equity, articulating specific needs and potentialities. Although the results do not allow for generalizations, they highlighted the need to strengthen professional training and the articulation between school, family, and specialized services. It was concluded that the inclusion of students with Dual Exceptionality requires practices that simultaneously recognize their difficulties and potentialities, ensuring equitable conditions for participation, learning, and development.

Keywords: Human development; Teacher training; School inclusion; Neurodivergence; Pedagogical practices.

October 02, 2026

Data Transformation into Strategy: Applied Research for Ecotourism Operation Optimization

The growing demand in ecotourism in Minas Gerais has driven the search for business intelligence to transform customer data into strategic information. The study aimed to structure a data science pipeline to collect, segment, and classify the customer base of an ecotourism operation, in order to optimize marketing actions and anticipate market movements. An exploratory, quali-quantitative research was conducted through a case study. 2,777 transactional records from an ecotourism company, referring to January 2024 to December 2025, were used. The methodological process involved automated data collection (Google Sheets API), processing and enrichment (ETL), validation, and creation of RFM (Recency, Frequency, and Monetary Value) attributes. Dimensionality reduction via PCA and K-Means clustering was applied, with the number of clusters defined by the Elbow method and Silhouette Score. The results were validated with DBSCAN and K-Medoids. The results revealed the identification of three behavioral customer segments: “Loyal”, “Low Value”, and “Potential”. The “Loyal” segment represented the highest accumulated economic value, while the “Potential” segment stood out for its high average ticket and potential for conversion into recurrence. The integration of data analysis techniques proved to be a robust and replicable method for generating intelligence in ecotourism. It was concluded that the structured data science pipeline enabled the behavioral segmentation of the customer base, the statistical validation of the groups, and the creation of a predictive system for new buyers, providing subsidies for data-driven strategic decisions and future analyses.

Keywords: Clustering; Business intelligence; Machine Learning; Customer segmentation; Decision making.

October 02, 2026

Sentiment Analysis on Brazilian Banks on Twitter/X: Comparison between Traditional and Digital Institutions

A study analyzed public perception of Brazilian financial institutions on the Twitter/X platform, highlighting the importance of sentiment monitoring on social networks for understanding reputation and customer experience in the banking sector. The objective was to compare user perception of the image and reputation of traditional and digital banks, based on the sentiment patterns identified in the analyzed manifestations, seeking to identify structural differences between these groups. The methodology was based on the analysis of 1,096 tweets collected between November 2022 and June 2023. Two complementary sentiment analysis approaches were used, the sum and the average of labels, to capture the majority sentiment and nuances of perception. Additionally, the Market Profile Model, with indicators of emotional reputation, reputational risk, neutrality, and polarization, and the Banking Clustering Model, which allowed grouping institutions according to perception patterns, were developed. The results indicated a predominance of neutral and negative sentiments, a higher volume of interactions in digital banks, and structural differences in the emotional intensity of perceptions, with greater stability in digital banks and greater polarization in traditional ones. It was concluded that the combination of analytical and statistical techniques contributed to an in-depth understanding of institutional image in the digital environment, demonstrating the importance of data-driven reputation management strategies.

Keywords: Digital banks; Traditional banks; Data modeling; Opinion mining; Social Networks.

October 02, 2026

Optimization of annual budget planning through project management methodologies

The Annual Budget Planning (POA) is a crucial process for translating organizational strategy into operational and financial goals, but it frequently faces deadline pressures, interdepartmental dependencies, and the repetition of habitual expenses. The study aimed to analyze how the combined application of project management practices and Zero-Based Budgeting (OBZ) can optimize the POA. To this end, a case study was developed in the Brazilian operation of a publicly traded company in the beverage sector, using documentary research of its 2023 results report and an anonymous questionnaire applied to 47 respondents. Documentary analysis indicated growth in net revenue, expansion of gross profit and adjusted EBITDA, and contained advancement of selling, general, and administrative expenses, suggesting cost discipline and operational leverage. The complementary survey revealed a high perception of cascading effect on the schedule, strong support for defining cost package owners, and a preference for technical justification of expenses, in addition to demand for controlled flexibility after the baseline definition. It was concluded that structuring the POA as a project, associated with the rigor of OBZ, increased the process predictability, reinforced accountability for expenses, and broadened the coherence between budgetary execution and economic-financial performance.

Keywords: Cost Control; Operational Efficiency; Zero-Based Budgeting; PMBOK; Beverage Sector.

Digital Business

October 02, 2026

Influence of social media on consumer behavior

The study of consumer behavior sought to understand the factors that influence purchasing decisions in the context of increasing digitalization, where the internet is widely used by the Brazilian population. The objective was to analyze how social networks influence consumer purchasing behavior and identify the types of content that generate the most attention. Data collection occurred through an online questionnaire, distributed to the general public, which resulted in 159 valid responses. The data were processed and analyzed using descriptive statistics and variable cross-tabulation to identify trends and correlations. The results revealed that 89.9% of respondents had already made purchases after exposure to content on social networks. Platforms such as TikTok and Pinterest showed the highest conversion rates among their users. It was observed that organic reviews and recommendations from friends or family exerted the greatest influence on purchasing decisions. Furthermore, it was identified that the absence of prior financial planning and the high frequency of exposure to dynamic content on social networks acted as catalysts for recurring purchases. It was concluded that social networks have consolidated themselves as strategic conversion channels, and understanding these mechanisms is fundamental for brands to develop efficient digital marketing strategies, prioritizing transparency and social proof.

Keywords: Consumer behavior; Purchase decision; Content strategy; Digital marketing; Social networks.