Article

October 02, 2026

Predictive modeling for inventory management: increased productivity and efficiency in a food distribution center

Predictive Modeling for Inventory Management: Increased Productivity and Efficiency in a Food Distribution Center

Felipe Augusto Toscani Figueiredo; Dora Yovana Barrios Leal

DOI: 10.22167/2675-6528-202602897

Article derived from a Final Course Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.

Summary

The efficient management of logistics operations in distribution centers is fundamental for competitiveness, especially in labor-intensive activities, which impact operational costs and service level. The study aimed to develop and evaluate different demand forecasting models to optimize inventory and resource management in a food distribution center. For this purpose, historical movement data from a WMS system were used, applying linear regression models, Holt-Winters exponential smoothing, and the LightGBM algorithm. The results indicated that traditional models presented limitations in capturing demand dynamics, with inferior performance to LightGBM. The machine learning algorithm demonstrated superior performance, especially after incorporating greater data granularity and temporal variables, which allowed for better adherence to demand dynamics. The application of the model resulted in a 36.38% reduction in displacement for pallet picking and 31.98% for case picking, totaling an overall reduction of 33.64%. It was concluded that the use of predictive models contributed to the improvement of operational efficiency, allowing for better resource allocation and organization of items in the order picking process.

Keywords: Machine learning; ABC classification; Operational efficiency; Logistics; Demand forecasting.

1. Introduction

The logistics chain comprises all the stages that connect production to the final consumer, involving suppliers, industries, distributors, and points of sale. The efficient management of this process constitutes a strategic pillar for organizations, as it ensures efficiency and competitiveness in an increasingly dynamic and demanding market (Cabral et al., 2023). Within modern supply chains, distribution centers play an essential role in storing, organizing, and moving products, functioning as an intermediate link between production and customer delivery (Finco et al., 2023; Tulli, 2023). In these locations, four main activities stand out: receiving, warehousing, order picking, and shipping (Amorim-Lopes et al., 2020).

Among these activities, the order picking operation, known as picking, is widely recognized as the most laborious and intensive step within warehouse activities, especially in non-automated fragmented picking systems. This essential process for order fulfillment involves separating products according to the quantities and sequences requested by customers, often requiring the operator to move to the storage location to collect materials (Amorim-Lopes et al., 2020; Picoli et al., 2023). Despite technological advancements in warehouse management, many operations still face inefficiencies due to poorly planned layouts, such as long operator routes, bottlenecks in high-traffic areas, underutilization of available space, and delays in picking and shipping stages (Casella et al., 2023; Tulli, 2023). These factors increase operational costs and extend order fulfillment times, potentially compromising customer satisfaction.

Warehouse processes are often supported by a Warehouse Management System (WMS), which manages physical and information flows, making operations more agile and precise. The WMS ensures that all physical transactions are recorded in a database, guaranteeing inventory accuracy and providing essential information for subsequent analyses (Brochado, 2024; Zaman et al., 2023). However, the reliance on fixed resource allocations and historical averages for planning, without a dynamic view of demand, often disregards the inherent volatility of operations, resulting in underutilization of resources and lower efficiency in the picking process.

To optimize inventory and resource management, the use of the ABC Curve is a common approach. This statistical classification of products is based on their relative importance in terms of value and volume, guiding inventory management and allowing efforts and resources to be concentrated on the most significant products, which promotes greater operational efficiency (Santa Ana, 2021; Todeschini and Lélis, 2022; Tulli, 2023). To assist in the ABC curve classification and bring more accuracy to inventory management, demand forecasting emerges as an essential complement. By anticipating product needs, it is possible to align service capacity with customer demands, reducing waste, improving inventory accuracy, and increasing the service level (Garcia and Fernandes, 2024).

The application of different analysis methods, from informal techniques to quantitative time series models and artificial intelligence techniques, can improve demand forecasting (Torres and Corso, 2021). In this context, this work aimed to develop and evaluate different models for demand forecasting, in order to select the one that provides the best estimates to optimize inventory and resource management in a food distribution center, increasing the efficiency of order picking.

2. Material and Methods

The study adopted a quantitative and descriptive approach, of an applied nature, aiming to develop and evaluate predictive models for inventory and resource management. The research was conducted in a 3PL food distribution center, located in the state of São Paulo. The empirical focus was on order picking movements, in pallet picking and case picking modalities, recorded between January 2023 and September 2025.

The database was extracted from the institution’s Warehouse Management System (WMS), which records all picking movements through barcode readings. The set totaled 2,024,472 records, corresponding to 67,705,531 boxes moved, distributed among 497 products from six families. An item master database was merged to link products to their families, and sensitive information was anonymized by codes.

In processing, duplicates were removed and only completed orders were filtered. Date and time columns were converted to daily format, allowing the aggregation of the quantity moved per day, product, family, and picking type. The “visits” variable was created to count the number of pickings. The explanatory variables included date, time, family, product code, and picking type, with the quantity of boxes moved as the response variable. Categorical variables were converted into numerical representations using One-Hot Encoding. The processing was performed in Python, with Pandas, NumPy, and Matplotlib.

For predictive modeling, the database was separated into training (initial 80%) and testing (most recent 20%) sets, respecting the temporal order. During training, Scikit-Learn’s Time Series Cross-Validation (Pedregosa et al., 2011) was employed. Hyperparameter optimization was performed with Scikit-Learn’s Randomized Search (Géron, 2019), which randomly selects parameter combinations to be tested. Linear regression analyses were conducted in Python, using the statsmodels, scikit-learn, statstests, and scipy libraries.

The selection of predictive models sought to contrast traditional statistical methods with machine learning algorithms (Saraiva and Yoshizaki, 2024). Multiple Linear Regression, the LightGBM algorithm, and the Holt-Winters exponential smoothing model were employed. Multiple Linear Regression served as a comparative baseline. LightGBM, based on decision trees and gradient boosting (Géron, 2019), incorporated lag variables and temporal markers. Holt-Winters (Kessous et al., 2024) decomposed the series into level, trend, and seasonality, using the additive version of the model.

To evaluate the models’ performance, the Coefficient of Determination (R²), the Mean Absolute Error (MAE), and the Root Mean Squared Error (RMSE) were selected (Saraiva and Yoshizaki, 2024). Models with higher R² and lower MAE and RMSE values were considered more accurate. The models were trained with daily data, preserving granularity, and the evaluation of predictions was performed after monthly aggregation, to align with logistical resource planning.

After selecting the model with the best accuracy, the ABC Curve was elaborated based on the forecasted demand of the items, considering the ordered cumulative percentage participation. Class A included the top 20% of exit frequency; Class B, the subsequent up to 50%; and Class C, the remaining 50%. The ABC Curve was elaborated separately for items in the ‘FAM_02’ family (air-conditioned area) and other items (ambient area). The need for collaborators and the average displacement between storage positions were estimated, with predefined distances for transitions between classes.

The results of the predictions obtained by the models were compared to a reference estimate. This estimate considered the maintenance of recent historical behavior, using the average observed in the last two months as the expected trend for the future period. This comparison allowed for the evaluation of the gain in precision and the operational impact provided by the predictive models.

3. Results and Discussion

The analysis of monthly demand at the food distribution center, covering the period from 2023 to 2025, revealed distinct patterns of volume and operational stability. It was observed that the year 2024 recorded the highest level of activity, with a median of 1.52 million boxes and a 3.1% growth in the monthly average compared to the previous year. In contrast, 2023 showed the greatest dispersion of data, indicating more unstable demand and, consequently, difficult predictability for resource planning. In 2025, there was a greater concentration of volumes, suggesting a more stabilized operation, although with a retraction of 3.8% in the monthly average, returning to 2023 levels, with 1.43 million boxes moved.

To verify the statistical significance of these annual variations, a one-factor analysis of variance (ANOVA) was applied. The test results (F = 0.65; p = 0.53) indicated that there were no statistically significant differences between the annual monthly averages. This confirmed that, despite the observed fluctuations between years, the average level of monthly demand remained statistically similar throughout the entire period analyzed. However, in contrast to the annual stability, the analysis of demand on a monthly scale revealed high volatility in the volume handled, with frequent fluctuations and no evident seasonal pattern, which increases forecasting complexity.

Monthly demand varied between approximately 1.22 million and 1.66 million boxes, resulting in a range of about 438,000 boxes between the months with the lowest and highest activity. This 35.8% variation between the extremes confirms that the average annual stability hides severe monthly fluctuations, requiring high flexibility from the logistics operation. From a logistical point of view, these variations are crucial for planning operational resources, as changes of this magnitude in the volume handled directly impact the need for labor, the intensity of operational visits, and the level of movement within the warehouse.

The analysis of productivity by picking type revealed an imbalance between the volume moved and the operational effort. Pallet picking, which involves moving complete pallets, concentrated the largest part of the total demand, representing 67.62% of the volume. On the other hand, case picking, which refers to the manual separation of fractional volumes, although representing only 32.38% of the total demand, was responsible for the vast majority of the flow in the warehouse, generating 76.52% of the stock position visits. This scenario highlights that unit picking is the main driver of movement and labor occupation in the operation.

The classification of items through the ABC Curve was elaborated based on the forecasted demand, considering that, in a 3PL logistics operator, the importance of products is more linked to their movement and consumption frequency than to monetary value. For the definition of groups, the criterion of cumulative percentage participation ordered by demand was adopted: Class A was composed of the 20% of items with the highest exit frequency; Class B by the subsequent items up to the limit of 50% of the total; and Class C by the remaining 50%, with low added value. The ABC Curve was elaborated separately for the items of the FAM_02 family, corresponding to the air-conditioned area, due to storage restrictions, and for the other items, stored in ambient area. The FAM_002 family accounts for approximately 30% of the total demand, reflecting its proportional need for positions in the air-conditioned area.

The performance evaluation of the predictive models considered the Coefficient of Determination (R²), the Mean Absolute Error (MAE), and the Root Mean Squared Error (RMSE), metrics consolidated in the literature (Saraiva and Yoshizaki, 2024). R² measures the proportion of the variability of the dependent variable explained by the model, while MAE and RMSE quantify the magnitude of the errors. Models with higher R² and lower MAE and RMSE values are considered more accurate and with better predictive performance, with these metrics being used complementarily in the comparison between approaches.

The linear regression model, although globally significant (Prob(F-statistic) = 0.00), presented an R² of 0.31, indicating a limited capacity to explain the variability of demand at a daily level. Furthermore, several individual coefficients were not statistically significant, suggesting multicollinearity, and the residuals did not adhere to normality, even after the Box-Cox transformation. These results indicate that, despite its simplicity and interpretability, linear regression served mainly as a comparative baseline for more complex models, as observed by Fávero and Belfiore (2024).

The initial LightGBM models, trained on a monthly basis, showed limitations in capturing temporal dynamics at a high aggregation level. However, with the use of more detailed data, incorporating temporal marking variables (lag and moving average) and daily granularity, an expressive improvement in performance was observed. The best result was achieved by LightGBM with daily data, products (One Hot Encoding), lags, and moving average, reaching an R² of 97.80%, MAE of 35,840 boxes, and RMSE of 46,662 boxes. This behavior reinforces the importance of the data aggregation level on predictive accuracy, as discussed by Hyndman et al. (2011), and highlights that greater granularity contributes to better capture of variations and reduction of noise effects in the series.

The Holt-Winters model, a seasonal exponential smoothing technique, was also evaluated. In the monthly version, it achieved an R² of 93.43%, MAE of 40,849.89, and RMSE of 67,929.57. In the daily version, the results were R² of 91.20%, MAE of 72,074.37, and RMSE of 79,854.12. Although Holt-Winters is suitable for time series with trend and seasonality (Kessous et al., 2024), LightGBM, with the incorporation of additional explanatory variables and the ability to capture non-linear relationships, demonstrated superior performance, aligning with studies that validate the robustness of machine learning in high volatility contexts (Saraiva and Yoshizaki, 2024; Schmid et al., 2025).

The analysis of the LightGBM model’s residuals, comparing actual demand and predictions, revealed distinct patterns for each picking type. In case picking, the residuals were predominantly negative, indicating a tendency for systematic overestimation of demand. In pallet picking, however, the residuals showed greater variability and sign alternation, suggesting a more balanced fit, but with underestimation during peak periods, especially in months with higher demand. This observation led to a refinement of the model, considering the hypothesis that different picking types might exhibit distinct demand patterns.

Given this hypothesis, a new experiment was conducted, maintaining the LightGBM settings but separating the data into two distinct sets: one for case picking and another for pallet picking, replacing the use of One-Hot Encoding for the picking type. This approach resulted in an increase in the monthly R² from 97.80% to 98.24%, representing a gain of 0.44 percentage points. Additionally, there was a reduction in MAE by 13.33% and in RMSE by 18.24%, indicating that the separation of models primarily contributed to reducing errors of greater magnitude, associated with periods of higher demand variation, where the previous model had more difficulty in fitting.

The translation of the results in terms of human resources allowed estimating the real labor requirement versus the one estimated by the model. For the calculation, average productivities of 135.59 boxes per operator per hour for case picking and 872.69 boxes per operator per hour for pallet picking were considered, in addition to a weekly workday of 44 hours. The final model’s predictions resulted in an average difference of 0.87 operator in case picking and only 0.06 operator in pallet picking compared to the real requirement.

The positive variation observed in case picking, which is the main driver of movement in the logistics center, can be considered adequate to the operational context, as a safety margin in resource estimation is preferable. In pallet picking, the small underestimation in August 2025 had a lesser impact on labor sizing, as this process does not represent the main operational bottleneck. Currently, the Logistics Center operates with a fixed number of 21 pickers and 9 forklift operators, based on historical demand calculation. The results show that operational demand is not constant, and a predictive model offers greater flexibility in team sizing.

The simulation of the cross-validation matrices between the actual, predicted, and historical ABC classifications demonstrated that the predictive model shows greater adherence to the actual classification compared to the historical method, for both case picking and pallet picking. In all classes, a higher concentration was observed on the main diagonal in the predictive scenario, indicating a significant increase in the correspondence between the classified curves and the actual demand. In case picking, the accuracy percentages in classes A, B, and C were 83.33%, 81.82%, and 97.13%, respectively, surpassing the 75.00%, 69.70%, and 96.31% of the historical method.

For pallet picking, the hit rates were 100.00% for class A, 77.78% for class B, and 98.19% for class C, in contrast to the 50.00%, 44.44%, and 96.99% of the historical method. This data reinforces the effectiveness of the predictive model in optimizing item prioritization and reducing classification errors, which is fundamental for more efficient inventory management and strategic resource allocation.

Based on the street depth in the Logistics Center, estimated at approximately 190 meters, and the average variation in displacement between classes (47.5 m for A and B, 71.25 m for B and C, and 118.75 m for A and C), the results indicated a reduction of 36.38% for pallet picking and 31.98% for case picking in displacement related to zone changes in the forecast-based scenario. This resulted in an estimated total reduction of 33.64% compared to the historical scenario, evidencing the model’s potential to increase the efficiency of the order picking process.

The drastic reduction in travel reinforces the thesis that Artificial Intelligence-based approaches outperform traditional warehouse organization methods. The use of intelligent algorithms to group and allocate products not only reduces picking distance but also improves operational ergonomics and item distribution accuracy, as pointed out by Adamah and Linner (2024). Therefore, the efficiency gains obtained validate the transition to a predictive model as a fundamental strategy for reducing logistical costs and improving productivity.

In summary, the research demonstrated that the application of predictive models, especially the LightGBM algorithm with daily granularity and temporal variables, surpassed traditional methods in demand forecasting in a food distribution center. The model’s higher accuracy allowed for a more precise estimation of labor needs and an ABC classification more aligned with operational reality, resulting in an overall reduction of 33.64% in order picking travel. These findings confirm that the use of predictive models, combined with the reorganization of items based on predicted demand, contributes significantly to the improvement of operational efficiency and the optimization of resource management, directly addressing the study’s objective.

4. Conclusion

The present study aimed to develop and evaluate demand forecasting models to optimize inventory and resource management in a food distribution center, seeking to improve order picking efficiency. It was found that, among the approaches tested, the LightGBM algorithm demonstrated superior performance in capturing demand dynamics, especially after incorporating greater data granularity and temporal variables. Traditional models, such as linear regression and Holt-Winters exponential smoothing, showed limitations in predictive accuracy. The refinement of LightGBM, by separating data by picking type, increased the coefficient of determination (R²) to 98.24% and significantly reduced average errors, indicating a notable adherence to demand variations.

The application of the predictive model resulted in substantial operational contributions. A more precise estimation of labor needs was observed, with an average difference of 0.87 operators for case picking and 0.06 operators for pallet picking compared to actual demand, allowing for greater flexibility in team sizing. Additionally, the ABC classification of items, based on forecasted demand, showed greater adherence to operational reality than the historical method, optimizing product prioritization. Consequently, the reorganization of items in the order picking process, guided by forecasts, generated a 36.38% reduction in travel for pallet picking and 31.98% for case picking, totaling an overall decrease of 33.64%. These findings validate the transition to predictive models as a fundamental strategy for improving operational efficiency and reducing logistics costs in distribution centers.

Bibliographic References

Amorim-Lopes, M.; Guimarães, L., Alves, J.; Almada-Lobo, B. 2020. Improving picking performance at a large retailer warehouse by combining probabilistic simulation, optimization, and discrete-event simulation. International Transactions in Operational Research, 28 (2): 687-715.

Cabral, G.; Fracarolli, J. V.B; Netto, R. V.; Bassi, A. L. 2023. Gestão da cadeia logística e centros de distribuição como otimizar o tempo das entregas. Revista Eletrônica CREARE – Revista das Engenharias, Ciências e Tecnologias 6 (1): 1-15.

Casella, G.; Volpi,

Finco, S.; Ashta, G.; Persona, A.; Zennaro, I. 2023. Investigating different manual picking workstations for robotized and automated warehouse systems: Trade-offs between ergonomics and productivity aspects. Computers & Industrial Engineering 185: 1-16.

Picoli, M. A.; Moreira, R. S.; Ribeiro, A. F.; Araújo, D. L. A.; Lemos, F. K. 2023. Inovação na operação logística: adoção do picking com códigos de barras e rádio frequência na metalúrgica São Raphael. Revista da Micro e Pequena Empresa (RMPE) 17 (1): 133-147.

Tulli, S. K. C. 2023. Warehouse Layout Optimization: Techniques for Improved Order Fulfillment Efficiency. International Journal of Acta Informatica 2 (1): 138-168.

Article originating from the Final Course Work of the Specialization in Data Science and Analytics of the MBA USP/Esalq

To learn more about the course, click here and access the MBX Academy platform

You may also like

October 02, 2026

Determinants of supermarket location in São Paulo

A study investigated the determining factors for supermarket location in the state of São Paulo, with the objective of investigating the factors that explain the presence and expansion of these establishments, considering socioeconomic, demographic, and market dimensions. Data from the 2010 and 2022 Demographic Censuses of IBGE and information from the National Registry of Legal Entities of the Federal Revenue of Brazil were used to build a georeferenced database. A Random Forest classification model was applied, adjusted by grid search with cross-validation, prioritizing the recall-macro metric due to the imbalance of the dependent variable, which represented the presence or absence of supermarkets within a 50-meter buffer. The results indicated that supermarket location is strongly associated with demographic, income, and population characteristics in the surrounding area. The analysis of variable importance showed that sociodemographic factors, such as elderly literacy, household income, and the presence of other food establishments, exerted significant influence, especially in the immediate vicinity. The findings reinforced the hypothesis that the spatial distribution of supermarkets is not random, being conditioned by socioeconomic characteristics and the commercial structure of the territory, offering subsidies for business decisions and urban planning.

Keywords: Spatial Analysis; Machine learning; Expansion; Commercial location; Supermarkets.

Neuroscience And Learning In Education

October 02, 2026

Anti-Racist Education: Inclusive Educational Practices and Social Development

Antiracist education, understood as a structuring axis of inclusive education and social development, was investigated in the Brazilian context. The study aimed to identify and analyze, based on legal documents and teachers’ perceptions, educational practices capable of promoting antiracism in school and society, and how the implementation of Laws nº 10.639/03 and nº 11.645/08 contributed to social justice. A qualitative and documentary approach was adopted, with analysis of educational legislation, curricular guidelines, institutional reports, and academic literature. Complementarily, a semi-structured questionnaire was applied to 295 Basic Education teachers. The data were evaluated quantitatively and qualitatively, through thematic content analysis, and validated with bibliographic studies. The results revealed a paradox: despite a robust legal framework, the implementation of antiracist policies proved fragile and sporadic, with a lack of teacher training, adequate teaching materials, and monitoring. Significant educational inequalities between white and black students were found to persist, and most teachers acknowledged the occurrence of racism in schools, but without clear institutional protocols. Neuroscientific analysis showed that racism negatively impacts students’ cognitive and emotional development. It was concluded that antiracist education is central to quality education, requiring political commitment, public investment, and intersectoral articulation. The integration of Neuroscience in teacher training and the production of qualified materials are crucial to strengthen the school’s role in building a more just and inclusive society.

Keywords: Social Development; Antiracist Education; Social Justice; Law 10.639/03; Inclusive Educational Practices.

Neuroscience And Learning In Education

October 02, 2026

Paths of Inclusion: Perceptions of Parents and Teachers on the Schooling of Students with Dual Exceptionality in the Brazilian Context

Dual Exceptionality, characterized by the coexistence of High Abilities/Giftedness and neurodevelopmental disorders, represents a complex phenomenon that challenges traditional identification and schooling models. The study aimed to understand the perceptions of parents or guardians, teachers, and other education professionals regarding the schooling of students with Dual Exceptionality in the Brazilian context, investigating challenges, pedagogical strategies, and possibilities for inclusion based on equity. The research adopted a qualitative, exploratory, and descriptive approach, and collected data through an online, voluntary, and anonymous questionnaire answered by 25 participants. Discursive data were analyzed using thematic content analysis. The results indicated that knowledge about the topic is often built from personal and professional experiences, revealing gaps in systematic training. Difficulties were identified in identifying these students, in teacher training, and in implementing individualized educational plans, pedagogical flexibility, and curriculum enrichment. Socio-emotional repercussions, such as frustration and low self-esteem, were reported. However, some schools demonstrated inclusive practices based on equity, articulating specific needs and potentialities. Although the results do not allow for generalizations, they highlighted the need to strengthen professional training and the articulation between school, family, and specialized services. It was concluded that the inclusion of students with Dual Exceptionality requires practices that simultaneously recognize their difficulties and potentialities, ensuring equitable conditions for participation, learning, and development.

Keywords: Human development; Teacher training; School inclusion; Neurodivergence; Pedagogical practices.

October 02, 2026

Data Transformation into Strategy: Applied Research for Ecotourism Operation Optimization

The growing demand in ecotourism in Minas Gerais has driven the search for business intelligence to transform customer data into strategic information. The study aimed to structure a data science pipeline to collect, segment, and classify the customer base of an ecotourism operation, in order to optimize marketing actions and anticipate market movements. An exploratory, quali-quantitative research was conducted through a case study. 2,777 transactional records from an ecotourism company, referring to January 2024 to December 2025, were used. The methodological process involved automated data collection (Google Sheets API), processing and enrichment (ETL), validation, and creation of RFM (Recency, Frequency, and Monetary Value) attributes. Dimensionality reduction via PCA and K-Means clustering was applied, with the number of clusters defined by the Elbow method and Silhouette Score. The results were validated with DBSCAN and K-Medoids. The results revealed the identification of three behavioral customer segments: “Loyal”, “Low Value”, and “Potential”. The “Loyal” segment represented the highest accumulated economic value, while the “Potential” segment stood out for its high average ticket and potential for conversion into recurrence. The integration of data analysis techniques proved to be a robust and replicable method for generating intelligence in ecotourism. It was concluded that the structured data science pipeline enabled the behavioral segmentation of the customer base, the statistical validation of the groups, and the creation of a predictive system for new buyers, providing subsidies for data-driven strategic decisions and future analyses.

Keywords: Clustering; Business intelligence; Machine Learning; Customer segmentation; Decision making.

October 02, 2026

Classification of defaulting customers using supervised machine learning techniques

The risk of default in credit operations demanded analytical approaches to anticipate losses. This study comparatively evaluated the performance of supervised machine learning models in classifying defaulting customers in credit card operations. The public dataset “Default of Credit Card Clients” from the University of California Irvine was used, with 30,000 observations and class imbalance. The algorithms Logistic Regression, Random Forest, and Extreme Gradient Boosting were employed. The imbalance was addressed by assigning weights to the classes, and model optimization occurred with the RandomizedSearchCV method, prioritizing sensitivity. Cross-validation results indicated that the Extreme Gradient Boosting model showed a higher capacity for identifying the defaulting class and better discriminatory performance, followed by Random Forest and Logistic Regression, with a sensitivity of 0.8250 and an AUC-ROC of 0.7844 for XGBoost. Interpretability analysis, conducted by the Shapley Additive Explanations (SHAP) technique, highlighted the predominance of variables associated with payment behavior, especially the history of delays. It was concluded that tree-based models, particularly boosting techniques, proved to be more suitable for capturing complex patterns in the data, configuring themselves as consistent alternatives for credit risk management.

Keywords: Machine Learning; Credit Card; Classification; Extreme Gradient Boosting; Credit Risk.

October 02, 2026

Sentiment Analysis on Brazilian Banks on Twitter/X: Comparison between Traditional and Digital Institutions

A study analyzed public perception of Brazilian financial institutions on the Twitter/X platform, highlighting the importance of sentiment monitoring on social networks for understanding reputation and customer experience in the banking sector. The objective was to compare user perception of the image and reputation of traditional and digital banks, based on the sentiment patterns identified in the analyzed manifestations, seeking to identify structural differences between these groups. The methodology was based on the analysis of 1,096 tweets collected between November 2022 and June 2023. Two complementary sentiment analysis approaches were used, the sum and the average of labels, to capture the majority sentiment and nuances of perception. Additionally, the Market Profile Model, with indicators of emotional reputation, reputational risk, neutrality, and polarization, and the Banking Clustering Model, which allowed grouping institutions according to perception patterns, were developed. The results indicated a predominance of neutral and negative sentiments, a higher volume of interactions in digital banks, and structural differences in the emotional intensity of perceptions, with greater stability in digital banks and greater polarization in traditional ones. It was concluded that the combination of analytical and statistical techniques contributed to an in-depth understanding of institutional image in the digital environment, demonstrating the importance of data-driven reputation management strategies.

Keywords: Digital banks; Traditional banks; Data modeling; Opinion mining; Social Networks.

October 02, 2026

Optimization of annual budget planning through project management methodologies

The Annual Budget Planning (POA) is a crucial process for translating organizational strategy into operational and financial goals, but it frequently faces deadline pressures, interdepartmental dependencies, and the repetition of habitual expenses. The study aimed to analyze how the combined application of project management practices and Zero-Based Budgeting (OBZ) can optimize the POA. To this end, a case study was developed in the Brazilian operation of a publicly traded company in the beverage sector, using documentary research of its 2023 results report and an anonymous questionnaire applied to 47 respondents. Documentary analysis indicated growth in net revenue, expansion of gross profit and adjusted EBITDA, and contained advancement of selling, general, and administrative expenses, suggesting cost discipline and operational leverage. The complementary survey revealed a high perception of cascading effect on the schedule, strong support for defining cost package owners, and a preference for technical justification of expenses, in addition to demand for controlled flexibility after the baseline definition. It was concluded that structuring the POA as a project, associated with the rigor of OBZ, increased the process predictability, reinforced accountability for expenses, and broadened the coherence between budgetary execution and economic-financial performance.

Keywords: Cost Control; Operational Efficiency; Zero-Based Budgeting; PMBOK; Beverage Sector.

Digital Business

October 02, 2026

Influence of social media on consumer behavior

The study of consumer behavior sought to understand the factors that influence purchasing decisions in the context of increasing digitalization, where the internet is widely used by the Brazilian population. The objective was to analyze how social networks influence consumer purchasing behavior and identify the types of content that generate the most attention. Data collection occurred through an online questionnaire, distributed to the general public, which resulted in 159 valid responses. The data were processed and analyzed using descriptive statistics and variable cross-tabulation to identify trends and correlations. The results revealed that 89.9% of respondents had already made purchases after exposure to content on social networks. Platforms such as TikTok and Pinterest showed the highest conversion rates among their users. It was observed that organic reviews and recommendations from friends or family exerted the greatest influence on purchasing decisions. Furthermore, it was identified that the absence of prior financial planning and the high frequency of exposure to dynamic content on social networks acted as catalysts for recurring purchases. It was concluded that social networks have consolidated themselves as strategic conversion channels, and understanding these mechanisms is fundamental for brands to develop efficient digital marketing strategies, prioritizing transparency and social proof.

Keywords: Consumer behavior; Purchase decision; Content strategy; Digital marketing; Social networks.