Technology
December 10, 2025
Predictive models for inventory management with time series in the implement industry
Author: Mariana Prignacca — Advisor: Vanessa Mesquita Blas Garcia
Summary prepared by the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege focused on synthesis and writing.
This work explores time series models of the Box-Jenkins family and exponential smoothing to forecast the demand for spare parts for agricultural implements, using a univariate, discrete, and stochastic series. The objective is to evaluate whether these approaches can support the optimization of inventory management. Spare parts inventory management is a strategic challenge, as it requires a balance between ensuring high availability to avoid operational disruptions and minimizing excessive storage costs (Dias, 2009; Ballou, 2006). Insufficient inventory can cause production stoppages and loss of customer confidence, while excess immobilizes capital and generates maintenance and obsolescence expenses (Arnold et al., 2011; Silver et al., 1998).
This duality makes inventory control a strategic activity that requires demand forecasting methods and robust replenishment policies (Muckstadt, 2005; Sherbrooke, 2004). In the Brazilian agribusiness, complexity is amplified by challenges such as limited storage capacity, deficiencies in logistics infrastructure, strong production seasonality, and scarcity of consistent historical data, factors that compromise planning efficiency (Appia, 2024; Folha Agrícola, 2025; Rodrigues and Ross, 2020). In the sugarcane implement sector, the demand for spare parts is influenced by crop seasonality, generating demand peaks in specific periods.
Demand forecasting is a pillar for inventory planning and control (Love, 1979). Accurately estimating future needs allows for optimizing inventory levels and supporting purchasing planning. Sales over time form a time series, and analyzing this data allows for the identification of patterns such as trends and seasonal cycles, which are essential for accurate future estimates (Box et al., 2015). Time series analysis seeks to predict future behavior by observing past behaviors, extracting information from chronologically ordered data (Nielsen, 2020).
Classical predictive models such as those from the Box-Jenkins family and the Holt-Winters method are widely used to capture seasonal patterns and trends (Hamilton, 1994; Hyndman and Athanasopoulos, 2018). However, applying these models to databases with few observations is prone to failures such as overfitting, where the model learns random noise instead of the underlying signal structure (Oses, 2023). This challenge is common in companies with little historical data, as models trained with insufficient data fail to generalize to new data (Géron, 2019).
The Box-Jenkins models, such as ARIMA and SARIMA, are robust tools for time series analysis based exclusively on historical data (Makridakis et al., 1998). In contrast, exponential smoothing models directly capture level, trend, and seasonality components, offering good adaptation in volatile demand scenarios, with low values and the presence of zeros, characteristics common in spare parts inventory (Makridakis et al., 2018). ARIMA models require additional steps for stationarity checking and differencing, which increases modeling complexity (Brockwell and Davis, 2002). In this regard, exponential smoothing models, such as Holt-Winters, can be more effective in low-magnitude seasonal series, providing more stable forecasts.
The study was characterized as an exploratory quantitative research, delineated as a case study. The Box-Jenkins methodology and exponential smoothing methods were applied to a secondary database of a company that manufactures implements for the sugar-energy sector, located in the interior of São Paulo. The historical series comprised the period from March 2022 to March 2025, totaling 37 monthly observations. The data were extracted from the company’s Enterprise Resource Planning (ERP) system. As it involved quantitative data, the analyses were interpreted numerically, allowing for the construction of comparative error tables for the evaluation of results (Turrioni and Mello, 2012; Godoy, 1995).
For the forecast, two families of models were compared: Box-Jenkins (ARIMA and SARIMA) and exponential smoothing (Simple Exponential Smoothing – SES, Holt’s method, simple seasonal method, and Holt-Winters). The choice was based on the competitive performance of these methods on series with a limited volume of observations (Box et al., 2015; Hyndman and Athanasopoulos, 2018; Makridakis et al., 2018). All steps were developed in Python (version 3.11.7), using libraries such as Pandas, NumPy, matplotlib, statsmodels, pmdarima, scikit-learn, arch, and scipy, aligned with current practices for time series forecasting (Joseph and Tackes, 2024).
The data preprocessing involved selecting a single product, identified by the code ‘XBFHH’, chosen for having the highest number of transactions. Sales were grouped into a monthly time series, and months with no sales records were filled with the value zero to ensure series continuity (Nielsen, 2020). Stationarity was verified with the Augmented Dickey-Fuller (ADF) and Kwiatkowski-Phillips-Schmidt-Shin (KPSS) tests, with a significance level of 0.05. The series decomposition into trend, seasonality, and residuals was performed to aid in understanding its behavior and in choosing between additive or multiplicative models (Morettin and Toloi, 2006).
To avoid overfitting, the series was divided into training (80% of the data) and testing (remaining 20%) sets, maintaining chronological order (Tomiazzi et al., 2018). Box-Jenkins modeling was conducted automatically, using the auto_arima function to search for the lowest information criterion (AIC and BIC), and manually, guided by the analysis of Autocorrelation Functions (ACF) and Partial Autocorrelation Functions (PACF) (Dickey and Fuller, 1981). Exponential smoothing modeling tested the SES, Holt, and Holt-Winters models. After fitting, a diagnostic analysis of the residuals was performed to verify if they behaved as white noise, using the Ljung-Box test for autocorrelation, the Kolmogorov-Smirnov test for normality (Stephens, 1974), and the ARCH test for heteroscedasticity (Engle, 1982). The performance of each model was quantified by the Root Mean Squared Error (RMSE) and Mean Absolute Error (MAE) metrics (Chai and Draxler, 2014; Willmott et al., 2009).
The analysis of the sales series for product ‘XBFHH’ revealed high volatility and sporadic behavior. With 37 observations, the series presented a standard deviation (7.81) higher than the average sales (5.30 units/month) and a range from 0 to 36 units. The median of 2 units reinforced the concentrated nature of demand. The Augmented Dickey-Fuller (ADF) test resulted in a p-value of 0.51, indicating that the series is non-stationary and justifying the need for differencing in ARIMA models.
The additive decomposition of the series revealed a gradual growth trend starting from 2023. The seasonality component presented a visual pattern, but a statistical test indicated that it was not statistically significant. The residual plot showed a noise component with high variance, signaling that a considerable part of the data’s variability is irregular, which poses a challenge for model accuracy. Non-stationarity confirmed the need for differencing, while the absence of statistical seasonality and strong noise justified the comparison between different models.
The comparative analysis of nine distinct models revealed that the “simple seasonal (additive)” model presented the best predictive performance, with the lowest RMSE (9.13) and MAE (5.92) values. Following it, the “simple exponential smoothing (SES)” model was positioned. This result suggests that, for this short and volatile series, the simpler smoothing approaches were more effective. In contrast, within the Box-Jenkins family, the ARIMA/SARIMA models with manual parameterization achieved competitive accuracy (RMSE of 9.24), close to that of the best smoothing models. The versions with 100% automatic parameterization, however, presented the highest errors, with RMSEs ranging between 11.99 and 15.22.
The superiority of the manual approach was a central finding. The analysis of the FAC and FACP graphs of the training series confirmed non-stationarity, indicating the need for differencing (d=1). After differencing, the analysis of the new graphs, guided by the principle of parsimony, suggested the choice of parameters p=1 and q=1 (Box et al., 2015). In contrast, the automated approach converged to a more complex specification, such as ARIMA(4,1,0), which had inferior predictive performance. This illustrates that, in short and high-noise series, the algorithmic search for a “perfect fit” can lead to overfitting, while the manual approach resulted in a simpler model with greater generalization capacity.
The diagnostic analysis of the residuals showed that all Box-Jenkins models generated independent residuals with constant variance. The divergence occurred in the normality test, which was failed by some of the more accurate models, such as the Manual SARIMA and SES. This result suggests that the non-normality of the residuals may be an intrinsic characteristic of the data series, whose sales peaks behave as outliers, rather than a model fitting failure. The most promising approach consisted of using simpler models or those guided by human analysis.
The results have practical implications. The identification of a model with a Mean Absolute Error (MAE) of approximately 5.92 units offers a quantitative decision support tool. An inventory manager can use the forecast as a baseline, knowing that, on average, the actual demand will be within a radius of approximately 6 units from the predicted value. This information allows for more rational sizing of safety stock, helping to reduce costs and mitigate the risk of stockouts.
It is fundamental to recognize the study’s limitations. The analysis focused on a single product with a short historical series, and the univariate nature of the analysis disregarded external variables that can impact demand, such as customer maintenance schedules or crop data. The absence of these exogenous variables is a possible cause for the variability not captured by the models. These limitations open avenues for future work, such as exploring models that allow the inclusion of external factors, applying the methodology to other products, and using techniques like Roll-Forward Cross Validation.
This study explored the applicability of time series models for forecasting the demand for spare parts. The exploration demonstrated that, for the series in question, simpler approaches, such as the simple additive seasonal model, proved more effective, outperforming more complex models. This finding suggests that, in contexts of scarce and noisy data, a model’s ability to capture a fundamental pattern without overfitting to noise is more critical than its complexity. Additionally, the study highlighted the superiority of manual parameterization of ARIMA models compared to automatic approaches for this scenario, reinforcing the importance of human analysis and the principle of parsimony. The analysis of limitations indicated that demand variability is unlikely to be explained solely by its past behavior. The most promising direction for future work lies in transitioning to multivariate models that allow the inclusion of exogenous variables. It is concluded that the objective was achieved: it was demonstrated that classic time series models, especially exponential smoothing ones, are viable for providing initial quantitative support to inventory management, even in scenarios with limited and volatile data.
References
Appia. 2024. Stock optimization with AI: reduce losses and maximize profitability in Brazilian agribusiness. Available at: <https://www. appia. com. br/post/otimiza%C3%A7%C3%A3o-de-estoque-com-ia-reduza-perdas-e-maximize-a-lucratividade-no-agroneg%C3%B3cio-brasileiro>. Accessed on: Aug 20, 2025.
Arnold, J. R. T.; Chapman, S. N.; Clive, L. M. 2011. Introduction to Materials Management. 7th ed. Pearson, New Jersey, NJ, USA.
Ballou, R. H. 2006. Supply Chain Management: Business Logistics. 5th ed. Bookman, Porto Alegre, RS, Brazil.
Box, G. E. P.; Jenkins, G. M.; Reinsel, G. C.; Ljung, G. M. 2015. Time Series Analysis: Forecasting and Control. 5th ed. Wiley, Hoboken, NJ, USA.
Brockwell, P. J.; Davis, R. A. 2002. Introduction to Time Series and Forecasting. 2nd ed. Springer, New York, NY, USA.
Chai, T.; Draxler, R. R. 2014. Root mean square error (RMSE) or mean absolute error (MAE)? – arguments against avoiding RMSE in the literature. Geoscientific Model Development 7: 1247-1250.
Dias, M. A. P. 2009. Materials Administration: Principles, concepts and management. Atlas, São Paulo, SP, Brazil.
Dickey, D. A.; Fuller, W. A. 1981. Likelihood ratio statistics for autoregressive time series with a unit root. Econometrica 49: 1057-1072.
Engle, R. F. 1982. Autoregressive conditional heteroscedasticity with estimates of the variance of United Kingdom inflation. Econometrica 50(4): 987-1007.
Folha Agrícola. 2025. Smart logistics is the key to overcoming the challenges of Brazilian agribusiness. Available at: https://folhaagricola. com. br/2025/07/23/logistica-inteligente-e-a-chave-para-superar-os-desafios-do-agronegocio-brasileiro/. Accessed on: Aug 20, 2025.
Géron, A. 2019. Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, tools, and techniques to build intelligent systems. 2nd ed. O’Reilly Media, Sebastopol, CA, USA.
Godoy, A. S. 1995. Qualitative research: fundamental types. RAE – Revista de Administração de Empresas 35(3): 20-29.
Hamilton, J. D. 1994. Time Series Analysis. Princeton University Press, Princeton, NJ, USA.
Hodson, T. 2022. Root-mean-square error (RMSE) or mean absolute error (MAE): when to use them or not. Geoscientific Model Development 15: 5481-5489.
Hyndman, R. J.; Athanasopoulos, G. 2018. Forecasting: Principles and practice. 2nd ed. OTexts, Melbourne, Victoria, Australia.
Joseph, M.; Tackes, J. 2024. Modern Time Series Forecasting with Python. 2nd ed. Packt Publishing, Birmingham, UK.
Kien, T. T.; Park, M. J.; Yang, H. S. 2024. Comparative study of time series analysis algorithms suitable for short-term forecasting in implementing demand response based on AMI. Sensors 24(22): 7205.
Love, S. 1979. Inventory Control. McGraw-Hill, New York, NY, USA.
Makridakis, S.; Spiliotis, E.; Assimakopoulos, V. 2018. Statistical and machine learning forecasting methods: concerns and ways forward. Plos one 13(3): 1-26.
Makridakis, S.; Wheelwright, S. C.; Hyndman, R. J. 1998. Forecasting: Methods and Applications. 3rd ed. John Wiley, New York, NY, USA.
Morettin, P. A.; Toloi, C. M. C. 2006. Time Series Analysis. 2nd ed. Blücher, São Paulo, SP, Brazil.
Muckstadt, J. A. 2005. Analysis and Algorithms for Service Parts Supply Chains. Springer, New York, NY, USA.
Nielsen, A. 2020. Practical Time Series Analysis: Prediction with statistics and machine learning. Alta Books, Rio de Janeiro, RJ, Brazil.
Oses, C.; Ward, L.; Ceder, G. 2023. Small data machine learning in materials science. npj Computational Materials 9(1).
Rodrigues, G. S. S. C.; Ross, J. L. S. 2020. The Trajectory of Sugarcane in Brazil: Geographical, Historical, and Environmental Perspectives. EDUFAL, Maceió, AL, Brazil.
Sherbrooke, C. C. 2004. Optimal Inventory Modeling of Systems: multi-echelon techniques. 2nd ed. Springer, New York, NY, USA.
Silver, E. A.; Pyke, D. F.; Peterson, R. 1998. Inventory Management and Production Planning and Scheduling. 3rd ed. Wiley, New York, NY, USA.
Slater, P. 2017. Spare Parts Inventory Management: A complete guide to sparesology®. Industrial Press, New York, NY, USA.
Stephens, M. A. 1974. EDF statistics for goodness of fit and some comparisons. Journal of the American Statistical Association 69(347): 730-737.
Tomiazzi, J. S.; Judai, M. A.; Nai, G. A.; Pereira, D. R.; Antunes, P. A.; Favareto, A. P. A. 2018. Evaluation of genotoxic effects in Brazilian agricultural workers exposed to pesticides and cigarette smoke using machine-learning algorithms. Environmental Science and Pollution Research 25: 1259-1269.
Turrioni, J. B.; Mello, C. F. 2012. Research Methodology in Production Engineering and Operations Management. Editora da UNIFEI, Itajubá, MG, Brazil.
Willmott, C. J.; Matsuura, K.; Robeson, S. M. 2009. Ambiguities inherent in sums-of-squares-based error statistics. Atmospheric Environment 43: 749-752.
Executive summary from the Final Project of the Specialization in Data Science and Analytics from the MBA USP/Esalq
Learn more about the course; click here: