Digital Business
September 30, 2026
Seasonality forecast in agribusiness: Machine Learning model for commercial targets
Seasonality Forecasting in Agribusiness: Machine Learning Model for Commercial Goals
Cleverson Mickosz do Nascimento; Gisela Consolmagno Pelegrini
DOI: 10.22167/2675-6528-202602785
Article derived from a Final Course Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.
Abstract
Brazilian agribusiness faces high uncertainty in defining commercial and financial goals. The study aimed to develop and evaluate a predictive model capable of anticipating sales seasonality patterns in the sector and transforming the results into managerial information to support planning. The research adopted an applied nature, quantitative approach, and exploratory character. Approximately 130,000 historical revenue records, referring to the period from 2019 to 2024, extracted from corporate ERP and CRM systems, and organized in monthly frequency, were used. Three machine learning algorithms were compared: SARIMAX, Random Forest, and XGBoost. Validation employed a chronological split of 80% of the data for training and 20% for testing, evaluating performance by RMSE and MAPE metrics. XGBoost presented the lowest error among the models, with an RMSE of 0.028 on the normalized scale and a MAPE of 7.9%, outperforming Random Forest (RMSE of 0.041; MAPE of 10.6%) and SARIMAX (RMSE of 0.065; MAPE of 14.8%). The results were integrated into interactive dashboards in Power BI, supporting the financial and commercial planning areas. The findings indicated the technical viability of the approach in the studied context, although its continuous application requires data governance, monitoring, and additional temporal validation.
Keywords: artificial intelligence; business planning; demand forecasting; time series; XGBoost.
1. Introduction
The Brazilian agribusiness represents an essential pillar for the national economy, contributing significantly to the Gross Domestic Product (GDP) and the country’s exports (Brasil, 2025; CEPEA/USP, 2025). However, the sector is intrinsically susceptible to high uncertainty, stemming from exogenous factors such as climatic conditions, pest incidence, and market fluctuations. This volatility poses considerable challenges to the strategic and financial planning of companies operating in agribusiness (Yoo; Oh, 2020).
In view of this scenario, the application of advanced artificial intelligence and machine learning techniques has shown promise in mitigating such uncertainties. Recent studies indicate that these approaches can improve predictability in sales projections, even in contexts characterized by strong external variability (Ensafi et al., 2022). Machine Learning models, in particular, have demonstrated relevant improvements in the accuracy of time series forecasting in agricultural environments, often outperforming traditional statistical models (Paul et al., 2022).
Despite the documented advances in the literature on the use of Machine Learning in agribusiness, a significant gap is still observed. Most studies focus on isolated crops or variables, or concentrate on estimating the productivity and prices of specific commodities. There is a lack of approaches that integrate multiple data sources into a predictive model directly aimed at defining commercial goals. Furthermore, developed Machine Learning solutions often remain isolated and are not adequately incorporated into organizational decision-making processes (Liakos et al., 2018).
In this context, the need to improve Machine Learning algorithms becomes evident so that they can deal more effectively with the complex and characteristic seasonal patterns of agriculture (Al-Gunaid et al., 2019). The ability to capture these temporal dynamics is crucial for enhancing the accuracy of sales forecasts in agribusinesses, providing more robust managerial information for planning. Thus, the present research is justified by the search for a predictive model that can transform historical data into actionable insights, supporting decision-making in financial and commercial planning.
The objective of this work is to develop and evaluate a predictive model to anticipate seasonality in agribusiness sales, based on historical data from 2019 to 2024, covering domestic sales and exports. To achieve this objective, supervised learning algorithms with different methodological characteristics are compared, and the results are integrated into interactive dashboards in Power BI. The application seeks to support the financial planning (Financial Planning & Analysis [FP&A]) and commercial areas in defining more assertive and data-driven goals.
2. Material and Methods
The research was designed with an applied nature, quantitative approach, and exploratory character. The applied nature aimed to develop a solution for commercial and financial planning. The quantitative approach allowed for the systematic analysis of patterns in historical sales series. The exploratory character stemmed from the absence of previous Machine Learning studies for commercial goals in agribusiness. The ethical standards of MBA USP/ESALQ were followed, ensuring data confidentiality and the non-identification of the company.
The data used consisted of approximately 130,000 historical invoicing records, covering domestic sales and exports, collected from 2019 to 2024. This data was extracted from corporate systems, specifically from the Enterprise Resource Planning (ERP) Microsoft Dynamics 365 and the Customer Relationship Management (CRM). Integration, via Azure Data Factory, consolidated the data into a SQL-structured “Data Warehouse”, ensuring standardization and consistency. The final “dataset”, with seven attributes, included sales order date, invoice date, product, customer, salesperson, transaction value, and the daily US dollar exchange rate as an exogenous variable.
The data treatment involved cleaning, standardization, and validation steps to ensure the quality and integrity of the dataset. Missing values were handled by temporal interpolation, a technique that estimated the missing values based on the local trend of the series, preserving seasonal patterns. Formatting inconsistencies were corrected, and duplicate records were removed. Categorical variables were encoded using “one-hot encoding”, and numerical values were normalized to the range of 0 to 1 using the Min-Max technique, aiming to optimize the efficiency of Machine Learning algorithms.
For variable engineering, time-lag variables (“lags”) of 3 and 6 months were created. The objective was to capture temporal dependencies and recurring seasonal patterns, characteristic of agribusiness. The data structure was organized at monthly frequency, allowing the identification of seasonal cycles associated with agricultural harvests and the export calendar, providing crucial information for predictive modeling of sales seasonality.
In the modeling stage, three Machine Learning algorithms with distinct methodological natures were employed to compare their predictive capabilities. SARIMAX (Seasonal AutoRegressive Integrated Moving Average with eXogenous variables) was selected as the classic linear benchmark for seasonal time series, incorporating autoregressive components, differencing, moving averages, seasonality, and exogenous variables (Kumar et al., 2022). Random Forest, a “bagging” method based on decision trees, combined predictions from multiple trees to capture non-linear relationships and complex interactions between variables, useful for the heterogeneity of the records.
The third algorithm was XGBoost, a “gradient boosting” approach. Unlike Random Forest, XGBoost builds trees sequentially, seeking to reduce the residual errors from previous steps. This characteristic favored the modeling of complex nonlinear relationships and heterogeneous structures, showing consistent performance in agricultural sales prediction (Mohamed-Amine et al., 2024). The comparison between the algorithms allowed us to evaluate which paradigm best fit the data structure. The development of the models was carried out in Python, using the pandas and scikit-learn libraries.
For model validation, the dataset was chronologically divided into 80% for training and 20% for testing, preserving the temporal structure of the series and avoiding the use of future information. Predictive performance was evaluated using the Root Mean Square Error (RMSE), which penalizes larger magnitude errors, and the Mean Absolute Percentage Error (MAPE), which expresses the error in percentage terms for managerial interpretation. The validation approach employed, although widely adopted, presented as a limitation the absence of multiple temporal validation windows, as pointed out by Cerqueira et al. (2020).
The general methodological flow of the research comprised data collection and processing, variable engineering, modeling with the selected algorithms, and validation. After model evaluation, the predictive results were integrated into interactive dashboards in Power BI. This integration transformed projections into accessible managerial information, supporting the financial and commercial planning areas in analyzing targets by product, customer, and salesperson, as per the study’s objective.
3. Results and Discussion
The analysis of historical billing data, covering the period from 2019 to 2024, revealed recurring patterns of seasonality in agribusiness sales, both in the domestic market and in exports. These patterns are intrinsically linked to agricultural calendar cycles, such as harvest periods and export windows, and demonstrate the sector’s strong dependence on cyclical exogenous factors, as discussed by Yoo and Oh (2020). The identification of these cycles is fundamental to improving the accuracy of projections and, consequently, the strategic and financial planning of companies.
The visual representation of these historical patterns, along with the projection for the year 2025, highlighted consistent seasonal peaks throughout the analyzed years. This observation reinforces the need for predictive models capable of capturing and anticipating these fluctuations, providing a more solid basis for defining commercial goals. The observed historical series and the projection generated by the XGBoost model for the following year demonstrated a remarkable adherence, indicating the model’s ability to replicate and extend these temporal dynamics.
Regarding the comparative evaluation of predictive models, the performance was analyzed on the test set, which corresponded to the final 20% of the data time series. The results demonstrated that the XGBoost algorithm outperformed the other evaluated models, presenting the lowest error values. The Root Mean Square Error (RMSE) of XGBoost was 0.028 on the normalized data scale, while the Mean Absolute Percentage Error (MAPE) reached 7.9%, indicating high predictive accuracy for the studied context.
In comparison, the Random Forest model obtained an RMSE of 0.041 and a MAPE of 10.6%, and SARIMAX registered an RMSE of 0.065 and a MAPE of 14.8%. This performance difference is significant, as the 7.9% MAPE obtained by XGBoost is below the 10% threshold, which is often classified as indicative of accurate forecasting in time series studies (Piekutowska et al., 2021). This result positions the developed model as a robust tool for business planning in agribusiness.
The superiority of XGBoost in this study can be attributed to the hierarchical and fragmented nature of the dataset used. Variables such as product, customer, and seller did not exhibit a uniform distribution across all records, which resulted in a heterogeneous data structure. Decision tree-based algorithms with “boosting”, such as XGBoost, are particularly effective in handling this complexity, unlike linear models such as SARIMAX, whose suitability depends on the specific characteristics of the data and the objectives of the analysis (Arumugam and Natarajan, 2023).
XGBoost’s ability to capture non-linear relationships and complex interactions between variables is consistent with findings by Li et al. (2024), who demonstrated the effectiveness of XGBoost-based models for predicting fluctuations in agricultural sales segmented into multiple product categories. The ability to process complex data structures with lower error corroborates the results obtained, reinforcing the choice of the algorithm for this specific scenario in Brazilian agribusiness.
However, it is important to note that the literature presents variations in the performance results of Machine Learning algorithms in different contexts. Mohamed-Amine et al. (2024), for example, identified the Gradient Boosting Regressor as the best-performing model in a study of phytosanitary product sales forecasting in a single region, where the dataset was more homogeneous. This divergence highlights that the effectiveness of an algorithm is highly dependent on the specific characteristics of the dataset and the problem at hand.
Nevertheless, the research by Elavarasan and Vincent (2020) corroborates the superiority of XGBoost in agricultural contexts with multiple variables, which aligns with the complexity of the present study, encompassing diverse products, customers, and sellers distributed across different markets. XGBoost’s ability to build trees sequentially, aiming to reduce residual errors from previous steps, allows for more precise modeling of non-linear relationships and heterogeneous structures, justifying its superior performance.
As a practical application of the research, the predictions generated by the XGBoost model were integrated into interactive dashboards developed in Power BI software. This integration represents the managerial deployment component of the model, transforming predictive results into an accessible query interface. The dashboards allow for the visualization of historical series, comparison between domestic sales and exports, and detailed analysis of projections by product, customer, and salesperson, facilitating managerial interpretation.
This integration approach brings the predictive model closer to organizational decision-making practices in agribusiness, as highlighted by Liakos et al. (2018) and Brugler et al. (2024), who emphasize the importance of Machine Learning solutions being incorporated into decision-making processes. Providing information in an intuitive and interactive format allows financial and commercial planning areas to effectively use projections for goal setting.
From an organizational point of view, the developed model contributes significantly to decision-making, replacing projections based on subjective criteria or simple historical averages with estimates grounded in patterns identified in the data. In strategic and commercial planning, the ability to anticipate seasonal peaks and troughs, which are common in agribusiness, allows for more precise inventory sizing, more efficient sales force allocation, and assertive calibration of targets by product, customer, and salesperson, all based on quantitative projections and historical data.
In terms of operational gains, the practical adoption of the model tends to reduce the risk of overestimating or underestimating demand, minimizing losses from excess production or stockouts during periods of high demand. This optimization contributes to the efficiency of the supply chain and the overall profitability of the company. The ability to predict with greater accuracy allows for more proactive and less reactive management of market challenges.
However, the continuous corporate implementation of the model presents some limitations that must be considered. Among them, the need for an integrated data infrastructure, which encompasses systems such as ERP, CRM, and a robust Data Warehouse, stands out to ensure the consistency and quality of records. Furthermore, the maintenance and periodic retraining of the model require professionals with skills in data engineering and science, and effective data governance is crucial to ensure the quality and consistency of records over time.
In summary, the study demonstrated the technical feasibility of developing and evaluating a predictive model based on Machine Learning to anticipate seasonality in agribusiness sales. The analysis of approximately 130,000 historical records revealed recurring seasonal patterns, and the XGBoost algorithm stood out with the best predictive performance, with a MAPE of 7.9%, providing a robust quantitative basis for financial and commercial planning. The integration of results into interactive dashboards in Power BI transformed forecasts into actionable managerial information, supporting the definition of more assertive goals and mitigating uncertainties in the sector.
4. Conclusion
The present study aimed to develop and evaluate a predictive model capable of anticipating seasonality in agribusiness sales, transforming the results into managerial information to support commercial and financial planning. The existence of recurring seasonality patterns in the sector’s sales, both domestic and export, was verified over the period from 2019 to 2024, intrinsically linked to agricultural cycles. The comparison between the machine learning algorithms SARIMAX, Random Forest, and XGBoost demonstrated the superiority of XGBoost, which presented the lowest predictive error, with a MAPE of 7.9%. This performance is attributed to its ability to handle the hierarchical and heterogeneous nature of the data, which included multiple products, customers, and salespeople. The main contribution of this work lies in the integration of these forecasts into interactive dashboards in Power BI, offering the financial and commercial planning areas a robust tool for setting more assertive and data-driven goals, replacing subjective projections with quantitative estimates and optimizing inventory management and sales force allocation.
However, the continuous implementation of the model presents limitations, such as dependence on an integrated data infrastructure and the need for specialized professionals for its maintenance and periodic retraining, in addition to effective data governance. Model validation used a single chronological split of the data, which may influence the estimation of the real error in different periods. For future studies, it is recommended to apply multiple temporal validation windows, incorporate additional climatic and macroeconomic variables, and explore hybrid models that combine linear and non-linear approaches. In a corporate context, it is essential to establish continuous monitoring of predictive error and retraining cycles to ensure the model’s relevance and accuracy over time.
Bibliographic References
Al-Gunaid, M.A.; Shcherbakov, M.V.; Trubitsin, V.N.; Shumkin, A.M.; Dereguzov, K.Y. 2019. Analysis of a short-term time series of crop sales based on machine learning methods. In: Conference on Creativity in Intelligent Technologies and Data Science, 2019. Springer International Publishing. p. 189-200.
Arumugam, V.; Natarajan, V. 2023. Time series modeling and forecasting using autoregressive integrated moving average and seasonal autoregressive integrated moving average models. Instrumentation Mesure Métrologie 22(4): 1-8. Disponível em: https://doi.org/10.18280/i2m.220404.
Brasil. Ministério da Agricultura e Pecuária. 2025. Agronegócio brasileiro fecha 2025 com recorde em exportações de US$ 169 bilhões e superávit de US$ 149,07 bilhões. Disponível em: https://www.gov.br/agricultura/pt-br/assuntos/noticias/agronegocio-brasileiro-fecha-2025-com-recorde-em-exportacoes-de-us-169-bilhoes-e-superavit-de-us-149-07-bilhoes.
Brugler, S.; Gardezi, M.; Dadkhah, A.; Rizzo, D.M.; Zia, A.; Clay, S.A. 2024. Improving decision support systems with machine learning: identifying barriers to adoption. Agronomy Journal 116: 1229-1236. Disponível em: https://doi.org/10.1002/agj2.21432.
Centro de Estudos Avançados em Economia Aplicada [CEPEA/USP]. 2025. PIB do agronegócio brasileiro. Disponível em: https://www.cepea.org.br/br/pib-do-agronegocio-brasileiro.aspx.
Cerqueira, V.; Torgo, L.; Mozetič, I. 2020. Evaluating time series forecasting models: an empirical study on performance estimation methods. Machine Learning 109: 1997-2028. Disponível em: https://doi.org/10.1007/s10994-020-05910-7.
Elavarasan, D.; Vincent, D.R. 2020. Reinforced XGBoost machine learning model for sustainable intelligent agrarian applications. Journal of Intelligent & Fuzzy Systems 39(5): 7605-7620. Disponível em: https://doi.org/10.3233/JIFS-200862.
Ensafi, Y.; Amin, S.H.; Zhang, G.; Shah, B. 2022. Time-series forecasting of seasonal items sales using machine learning: a comparative analysis. International Journal of Information Management Data Insights 2(1): 100058. Disponível em: https://doi.org/10.1016/j.jjimei.2022.100058.
Kumar, N.P.; Bhaskar, S.; Srinidhi, S.P.; Shashank, D.; Karanam, S.G. 2022. Machine learning based predictive analytics for agriculture inventory management system. In: International Conference on Cognitive Computing and Information Processing, 4., 2022, Bengaluru. IEEE. p. 1-7. Disponível em: https://doi.org/10.1109/CCIP57447.2022.10058690.
Li, J.; Lin, B.; Wang, P.; Chen, Y.; Zeng, X.; Liu, X.; Chen, R. 2024. A hierarchical RF-XGBoost model for short-cycle agricultural product sales forecasting. Foods 13(18): 2936. Disponível em: https://doi.org/10.3390/foods13182936.
Liakos, K.G.; Busato, P.; Moshou, D.; Pearson, S.; Bochtis, D. 2018. Machine learning in agriculture: a review. Sensors 18(8): 2674. Disponível em: https://doi.org/10.3390/s18082674.
Mohamed-Amine, N.; Abdellatif, M.; Belaid, B. 2024. Artificial intelligence for forecasting sales of agricultural products: a case study of a Moroccan agrarian company. Journal of Open Innovation: Technology, Market, and Complexity 10(1): 100189. Disponível em: https://doi.org/10.1016/j.joitmc.2023.100189.
Paul, R.K.; Yeasin, M.; Kumar, P.; Kumar, P.; Balasubramanian, M.; Roy, H.S.; Paul, A.K.; Gupta, A. 2022. Machine learning techniques for forecasting agricultural prices: a case of brinjal in Odisha, India. PLOS ONE 17(7): e0270553. Disponível em: https://doi.org/10.1371/journal.pone.0270553.
Piekutowska, M.; Niedbała, G.; Piskier, T.; Lenartowicz, T.; Pilarski, K.; Wojciechowski, T.; Pilarska, A.A.; Czechowska-Kosacka, A. 2021. The application of multiple linear regression and artificial neural network models for yield prediction of very early potato cultivars before harvest. Agronomy 11(5): 885. Disponível em: https://doi.org/10.3390/agronomy11050885.
Yoo, T.-W.; Oh, I.-S. 2020. Time series forecasting of agricultural products’ sales volumes based on seasonal long short-term memory. Applied Sciences 10(22): 8169. Disponível em: https://doi.org/10.3390/app10228169.
Article originating from the Final Course Work of Specialization in Digital Business from the MBA USP/Esalq
To learn more about the course, click here and access the MBX Academy platform