October 09, 2026
Use of Web Scraping and Data Science for Optimization of Metro-Railway Maintenance
Use of Web Scraping and Data Science for Optimization of Metro-Railway Maintenance
Lucas Duarte de Souza; Felipe Pinto da Silva
DOI: 10.22167/2675-6528-202603177
Article derived from a Course Conclusion Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.
Summary
Metro-railway maintenance generates a significant volume of operational records that often remain underutilized for predictive analyses. The main objective of this study was to automate data collection and compare the performance of predictive models to estimate the weekly volume of failures in the São Paulo Metro’s Route Control System (CBI). A virtualized computational infrastructure (Proxmox and LXC) was implemented with a Python/Flask application for continuous web scraping of data from the SAP system, stored in a MySQL database. 3,181 failure records from 2021 to 2024 were processed. Sliding Window Cross-Validation was applied to evaluate the Holt-Winters method and the Facebook Prophet algorithm, after weekly aggregation of data to stabilize the time series, given the high daily volatility. The results indicated that the Holt-Winters model presented the best predictive fit, with a Mean Absolute Error (MAE) of 5.63, being approximately 30% more accurate than Prophet (MAE = 8.08) in the specific context. An increase from approximately 10 to 18 weekly failures was projected for the next cycle. It was concluded that the integration of an automated data pipeline with the Holt-Winters model offers a robust tool for projecting the failure behavior of the CBI system, enabling the transition to predictive maintenance that optimizes resource allocation and enhances operational reliability.
Keywords: Process Automation; Holt-Winters; Failure Prediction; Time series; Signaling.
1. Introduction
The São Paulo Metro is a vital element in the urban mobility infrastructure of the country’s largest metropolis, responsible for transporting approximately three million seven hundred thousand people daily on weekdays (COMPANHIA DO METROPOLITANO DE SÃO PAULO, [202-?]). Ensuring operational efficiency, continuous availability, and user safety intrinsically depends on highly effective maintenance management. This complex technological ecosystem encompasses thousands of interdependent equipment, including energy, telecommunications, and crucially, railway signaling systems.
With the advancement of Industry 4.0 and the increasing digitalization of metro-railway equipment, the maintenance paradigm has transformed. It is no longer perceived merely as a cost center focused on corrective repairs but has become a strategic area, capable of generating predictive intelligence and optimizing asset availability through the analysis of large volumes of historical data, known as Big Data (Lee, Kao, and Yang, 2014). Recent research has demonstrated the development of explainable Machine Learning frameworks that utilize real-time data streams to increase reliability in metro-railway operators (Abreu et al., 2024), and how data mining techniques allow for the strategic prioritization of maintenance planning, overcoming the limitations of fixed preventive schedules (Obal, Bellinello, and Petrovic, 2025).
Despite the significant volume of data on failures and interventions generated and recorded in internal systems, these records often remain underutilized or restricted to descriptive functions, not exploring the full extent of information for the anticipation of recurring failures. This gap between collection and effective analysis is particularly critical in highly complex assets, such as the Route Control System (CBI), which requires a more proactive approach to ensure service continuity and safety.
The relevance of this work lies in its ability to transform raw data into predictive intelligence, enabling the early projection of degradation curves and the identification of seasonal stress patterns in the system. The effectiveness of systematic analysis of historical records for this purpose is strongly corroborated in the literature, where statistical methods allow for precise estimation of future occurrence volumes (Hyndman; Athanasopoulos, 2021), and the use of data mining enables the strategic prioritization of preventive maintenance actions during periods of greatest criticality (Obal, Bellinello e Petrovic, 2025). Applied to the metro-rail context, as evidenced by Abreu et al. (2024), this proactive approach optimizes resource and team allocation, directly translating into reduced operational costs, increased asset availability, and, crucially, enhanced safety levels and passenger satisfaction.
Given the above, the objective of this work was to develop and compare predictive failure models — specifically the Holt-Winters statistical method and the Facebook Prophet algorithm — to determine which approach offers greater accuracy in predicting weekly failures of the CBI system. To enable this study, a virtualized data infrastructure in Proxmox was used, powered by an automated data collection routine (Web Scraping) from the SAP system, ensuring the integrity and continuous updating of the analyzed information.
2. Material and Methods
This research was classified as applied, aiming to generate practical knowledge for the optimization of railway maintenance (GIL, 2019). A quantitative approach was adopted, based on the collection and statistical analysis of numerical data to identify behavioral patterns (PRODANOV; FREITAS, 2013). The work was structured in stages of infrastructure configuration, automated collection, data processing, and comparative statistical modeling.
The empirical object consisted of the failure records of the São Paulo Metro’s Route Control System (CBI). 3,181 failure records were collected, covering the period from January 2021 to December 2024. The data were extracted from the SAP corporate system, which stores information about occurrences and maintenance interventions.
For the acquisition and processing of data, a virtualized computational infrastructure with Proxmox VE hypervisor was implemented. A customized web application in Python/Flask was hosted to automate the data flow and centralize information. The architecture was segregated into distinct LXC containers: one for MySQL, another for the backend and web scraping application, and a third for the exploratory analysis environment (Jupyter Lab).
The data extraction routine, via web scraping, was implemented in the application’s backend using Python, with the Requests and BeautifulSoup libraries. The algorithm simulated user navigation in the SAP system, with automated authentication by encrypted credentials. The process followed an automated ETL (Extract, Transform, Load) flow, requesting occurrence reports, interventions, and service orders, processing HTML tables and binary files (.xls).
The raw data were processed in memory with the Pandas library for data cleansing, typing, and standardization of fields, such as the extraction of Line and Zone. Subsequently, the processed data were inserted and updated in the MySQL database using the SQLAlchemy ORM, with a relational model that ensured integrity and avoided duplicates. An integrated logging system accompanied the entire process to ensure traceability.
The analytical processing occurred in the Jupyter Lab environment, where the data underwent cleaning, date conversion, and resampling steps. Two distinct time series were generated for analysis: one with daily granularity and another with weekly aggregation. The high daily stochastic volatility justified the weekly aggregation for stabilization of the time series and better exposure of the CBI system’s behavioral patterns.
For predictive modeling, the statistical method Holt-Winters and the Facebook Prophet algorithm were selected. Holt-Winters, a Triple Exponential Smoothing, was applied because it is suitable for series with trend and seasonality. The additive method was chosen, given the constancy of the series variance (HYNDMAN; ATHANASOPOULOS, 2021). The smoothing parameters were automatically optimized by the Statsmodels algorithm to minimize the quadratic error, with the seasonal period fixed at 4.
In a comparative manner, the Facebook Prophet algorithm was used, an additive regression model that adapts to data with strong seasonal effects and trend changes. Prophet formulates the problem as a decomposable curve fitting (TAYLOR; LETHAM, 2018), where the seasonality component is modeled by a Fourier Series. Daily seasonality was deactivated, keeping the weekly one active, and a custom seasonality parameter was inserted to capture monthly behavior.
For performance validation and comparison, Sliding Window Cross-Validation (Walk-Forward Validation) was adopted. This technique, implemented via Scikit-Learn’s TimeSeriesSplit, divided the dataset into 5 sequential iterations. In each iteration, the model was trained with a progressive historical window and tested on the immediately future window, avoiding the use of future information in training. The adopted performance metric was the Mean Absolute Error (MAE), due to its direct interpretability in relation to the physical quantity of failures.
3. Results and Discussion
The initial stage of the research focused on data mining to identify and isolate the Route Control System (CBI) as the critical asset for the study, as established in the objective. During the period between January 2021 and December 2024, 3,181 failure records related to this system were accumulated. This robust database allowed for an in-depth analysis of failure behavior, providing the necessary empirical support for the development and comparison of predictive models, aligning with the proposal to transform raw data into operational intelligence, as suggested by Lee, Kao, and Yang (2014).
Prior to the application of predictive algorithms, an Exploratory Data Analysis (EDA) was performed to characterize the historical series of failures. It was observed that the daily granularity of the series exhibited a highly stochastic behavior, with a Coefficient of Variation (CV) greater than 100%. The average daily failures were 2.97, with a standard deviation of 3.12, and a total count of approximately 1,460 days. This high volatility, along with the presence of a significant number of days without failures (zero values), indicated that the daily granularity would be inadequate for predictive modeling, hindering the convergence of regression models.
In contrast, aggregating the data to a weekly granularity demonstrated a considerable stabilization of the time series. In this format, the series spanned approximately 208 weeks, with an average of 16.87 failures per week and a standard deviation of 8.45. The Coefficient of Variation (CV) was reduced to about 50%, indicating moderate volatility. This weekly aggregation eliminated null values and smoothed the variance, more clearly revealing the cyclical trend of the CBI system degradation process. This resampling approach was crucial for exposing the underlying patterns of system behavior, according to the literature on time series (Hyndman; Athanasopoulos, 2021).
Based on the statistical characterization, the predictive models were applied prioritizing the weekly series. In the initial experiment with the Holt-Winters model, application to the daily granularity resulted in a Mean Absolute Error (MAE) of 1.70, which represented about 57% of the average daily failures, confirming low reliability for this frequency. However, when applied to the weekly series, the Holt-Winters model obtained an MAE of 5.63. This result represented a reduction in the relative error to approximately 33%, indicating a robust capacity to capture the asset degradation trend and project the failure volume.
The projection made by the Holt-Winters model indicated an expected increase from approximately 10 to 18 weekly failures in the next cycle, evidencing the model’s predictive capability to anticipate the system’s future behavior. The visualization of the observed behavior and the Holt-Winters forecast demonstrated an adequate fit, with the forecast line closely following the trend of historical data. This finding is fundamental for the transition from reactive maintenance to a proactive approach, optimizing the allocation of resources and teams, as indicated by the literature (Abreu et al., 2024).
Subsequently, the Facebook Prophet algorithm was applied under the same cross-validation methodology with sliding windows. The Prophet model, although recognized for its ability to handle seasonality and trends, showed instability in the initial test windows. After convergence, the model achieved a final Mean Absolute Error (MAE) of 8.08. Graphical analysis of the Prophet’s forecast against the actual data showed greater dispersion and less adherence to the peaks and valleys of the observed time series, especially in the initial phases of the forecast.
The direct comparison between the two predictive models revealed that Holt-Winters presented superior performance. Holt-Winters achieved a MAE of 5.63 and a Root Mean Squared Error (RMSE) of 7.04, being classified as the best-fitting model. In contrast, Facebook Prophet obtained a MAE of 8.08 and an RMSE of 9.73, resulting in an accuracy approximately 43% lower than that of Holt-Winters for this specific dataset. This substantial difference in error metrics validates the choice of Holt-Winters as the main prediction engine for the weekly failure volume of the CBI system.
For the Holt-Winters model, the adopted configuration included additive trend and seasonality components, with the seasonal period fixed at four, reflecting the monthly cyclical variations in the weekly aggregation. The series smoothing parameters, such as level, trend, and seasonality, were not arbitrarily fixed but automatically optimized by the Statsmodels algorithm. This optimization, based on maximum likelihood, ensured the best mathematical fit to the training data, conferring numerical robustness to the model and aligning with best practices in time series modeling (Hyndman; Athanasopoulos, 2021).
Despite Prophet’s underperformance on global error metrics, the decomposition of its components provided important evidence about the seasonal behavior of the failure series. Weekly seasonal analysis, for example, identified Tuesdays, Thursdays, and Saturdays as the days with the highest probability of failure occurrence. In contrast, Sundays and Mondays showed a natural retraction in the volume of failures, indicating operational and system usage patterns that influence the frequency of occurrences.
Additionally, the decomposition of Prophet’s components revealed monthly criticality cycles, with notable peaks on the 4th and 28th days of the month. On an annual level, a clear trend of worsening occurrences was observed during December. These seasonal insights, although not directly used for primary prediction due to the superiority of Holt-Winters, are valuable for diagnosing operational patterns and for strategic maintenance planning, allowing the maintenance team to anticipate periods of higher demand and allocate resources more efficiently (Obal, Bellinello, and Petrovic, 2025).
The virtualised data infrastructure in Proxmox, with the Python/Flask application for continuous web scraping of the SAP system and storage in MySQL, proved effective in collecting and processing the 3,181 failure records. This data pipeline automation is a fundamental pillar for the sustainability of the predictive approach, ensuring that models are fed with updated and integral information. The ability to extract and process data from corporate systems, such as SAP, is a differentiator that allows the practical application of data science techniques in real metro-railway maintenance environments.
In summary, the research results confirmed that the Holt-Winters model, when applied to the weekly failure series of the CBI system, offers the highest predictive accuracy, with a significantly lower Mean Absolute Error compared to Facebook Prophet. The integration of this model with an automated data collection pipeline represents a robust tool for projecting future failure behavior. This prediction capability allows São Paulo Metro to transition to more efficient predictive maintenance, optimizing resource allocation and increasing the operational reliability of the signaling system, directly responding to the study’s main objective.
4. Conclusion
This study aimed to develop and compare predictive models to estimate the weekly failure volume in the Route Control System (CBI) of the São Paulo Metro, using an automated data infrastructure. It was found that the continuous collection of 3,181 failure records, between 2021 and 2024, through web scraping of the SAP system and storage in MySQL, established a robust database. Exploratory analysis revealed the high volatility of the daily time series, which justified aggregating the data into weekly granularity to stabilize failure behavior. In this context, the Holt-Winters model demonstrated the best predictive performance, with a Mean Absolute Error (MAE) of 5.63, outperforming the Facebook Prophet algorithm, which obtained an MAE of 8.08. The projection made by Holt-Winters indicated an increase from approximately 10 to 18 weekly failures in the next cycle, evidencing its capacity to anticipate asset degradation. The main contribution lies in the integration of an automated data pipeline with an effective predictive model, offering a robust tool that allows the São Paulo Metro to transition from reactive maintenance to a predictive approach, optimizing resource allocation and enhancing the operational reliability of the CBI system.
Despite the demonstrated effectiveness, the study faced the limitation of high stochasticity in daily data, which required weekly aggregation and may have mitigated the detection of very short-term failure patterns. Additionally, it was observed that the Facebook Prophet algorithm, although useful for decomposing seasonal components that identified failure peaks on Tuesdays, Thursdays, and Saturdays, and worsening trends in December, presented instability in the initial test windows and lower overall predictive accuracy for this dataset. As a suggestion for future studies, the implementation of an interactive dashboard for maintenance team visualization of predictions is recommended, facilitating decision-making. Furthermore, deepening the analysis of seasonal patterns identified by Prophet can refine strategic maintenance planning, allowing for more assertive interventions during periods of higher criticality.
Bibliographic References
ABREU, D. et al. An explainable machine learning framework for railway predictive maintenance using data streams from the metro operator of Portugal. Scientific Reports, v. 14, n. 1, p. 8189, 2024.
COMPANHIA DO METROPOLITANO DE SÃO PAULO (METRÔ). Tecnologia: Operação. São Paulo, [202-?]. Disponível em: https://www.metro.sp.gov.br/tecnologia/operacao/. Acesso em: 03 fev. 2026.
GIL, A. C. Como elaborar projetos de pesquisa. 6. ed. São Paulo: Atlas, 2019.
HYNDMAN, R. J.; ATHANASOPOULOS, G. Forecasting: principles and practice. 3. ed. [S.I.]: OTexts, 2021. Disponível em: OTexts.com/fpp3.
LEE, J.; KAO, H.-A.; YANG, S. Service innovation and smart analytics for industry 4.0 and big data environment. Procedia CIRP, v. 16, p. 3-8, 2014.
OBAL, T. M.; BELLINELLO, M.; PETROVIC, S. Uso de mineração de dados para priorização no planejamento de manutenção preventiva. Proceeding Series of the Brazilian Society of Computational and Applied Mathematics, v. 9, n. 1, 2025.
PRODANOV, C. C.; FREITAS, E. C. Metodologia do trabalho científico: métodos e técnicas da pesquisa e do trabalho acadêmico. 2. ed. Novo Hamburgo: Feevale, 2013.
TAYLOR, S. J.; LETHAM, B. Prophet: forecasting at scale. The American Statistician, v. 72, n. 1, p. 37-45, 2018.
Article originating from the Final Course Project of the Specialization in Data Science and Analytics of the MBA USP/Esalq
To learn more about the course, click here and access the MBX Academy platform