Software Engineering
October 08, 2026
Data pipeline for stock monitoring considering the fundamentalist methodology
Data Pipeline for Stock Monitoring Considering the Fundamentalist Methodology
Leonardo Gomes Maciel; Arthur Pinheiro de Araújo Costa
DOI: 10.22167/2675-6528-202603142
Article derived from a Final Course Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.
Summary
The Brazilian financial scenario faces challenges such as family indebtedness, low financial literacy, and the decentralization of information for investment analysis. Given this, the research objective was to develop a financial data pipeline, based on good Data Engineering practices, to structure, process, and make relevant information available to support individual investors’ decision-making in stock analysis. A medallion architecture was implemented, using MinIO S3 for bronze, silver, and gold layer storage, with Delta Lake for data governance. Apache Spark was employed for distributed processing, Prometheus for monitoring, and Power BI for analytical visualization. The data were mostly financial statements from the Securities and Exchange Commission (CVM). A theoretical investment portfolio was built based on Benjamin Graham’s (2017) principles, applying selection filters and data validation. The results indicated a positive return of 32.88% for the theoretical portfolio, outperforming the Ibovespa index (4.25%) in the period from 2021 to 2024, although lower than the Selic rate (46.96%). Qualitatively, the pipeline processed voluminous datasets, with significant reductions in redundancies, such as 99.27% in the BPA table and 94.29% in the DRE, after applying filters. Gains in data organization, traceability, and quality were evidenced, enabling structured and more robust financial analyses.
Keywords: Medallion Architecture; Data Engineering; Investments; Data Pipeline.
1. Introduction
The current Brazilian economic scenario is marked by significant challenges, such as the increase in household debt and the persistent lack of financial literacy. In September 2025, Brazil registered 78.2 million people in default, an increase of 5.6% compared to the previous year (Serasa, 2025). This problem is not recent and has already been addressed in previous studies that highlight the need for financial education for individuals’ quality of life (Campos, Teixeira, and Coutinho, 2015). Organizations such as the Organisation for Economic Co-operation and Development (OECD) have promoted financial literacy and capacity programs, aiming to develop essential skills for informed decision-making (Hoffman; Moro, 2012). The OECD’s recommendations cover topics such as personal financial planning, insurance, savings, investments, and indebtedness (Campos; Teixeira; Coutinho, 2015).
In this context, financial investment emerges as a topic of great relevance. Fundamental analysis, as presented by Benjamin Graham (1949) in his work “The Intelligent Investor”, proposes a method for evaluating companies based on their intrinsic value, aiding in the selection of solid assets with growth potential. An investor who adopts the fundamentalist approach seeks to minimize risks by focusing on consolidated assets in the financial market, considering the company’s history and scenario (Freitas, 2020).
Despite the importance of investments, the capital market presents complexities. The Brazilian Stock Exchange (B3) lists more than 400 stocks, including common, preferred, and units (InfoMoney, 2025). However, the fundamental information needed for in-depth analysis is often decentralized, demanding considerable time and effort from investors (Montóia, 2021). To overcome this difficulty and ensure that information is distributed efficiently, the implementation of a data pipeline becomes essential. This concept encompasses the stages from extracting raw data from various sources to loading it into its destination (Densmore, 2021), with the choice of technologies varying according to the context and maturity of the data engineer.
Given the need for technical knowledge encompassing both finance and technology, and to mitigate the problem of information decentralization and support individual investor decision-making, the general objective of this work is to develop an automated data pipeline for fundamental analysis of Brazilian stocks, applying Benjamin Graham’s methodology (2019).
2. Material and Methods
The research conducted was characterized as an action-research, with a quali-quantitative approach, focused on the development of a data pipeline applied to the financial market. The central objective was to structure, process, and make available information to support individual investors’ decision-making in the fundamental analysis of Brazilian stocks. The study was limited to the evaluation of the capital market, without delving into macroeconomic or sectoral microeconomic analyses, according to Benjamin Graham’s methodology (2019).
In the scope of investments, fundamental analysis was adopted, which seeks to evaluate companies based on their intrinsic value. For asset selection, specific criteria were applied, such as the absence of losses in the analyzed period, a product between Price/Earnings (P/E) and Price/Book Value per Share (P/B) lower than 22.5, dividend distribution, and a current ratio not lower than 1 (Silva, 2020; Assaf, 2018). These criteria aimed to identify solid companies with growth potential, aligned with the profile of a defensive investor.
The data engineering methodology involved the construction of a data pipeline, following best practices in the field and the medallion architecture (Strengholt, 2025). This architecture organizes data into three distinct layers: bronze, silver, and gold, aiming to enhance the structure, quality, and validation of information throughout the process. The implementation was carried out using Python 3.11 (Luna, 2025), with the use of Docker containers to ensure portability and reproducibility of the environment.
The data ingestion stage, corresponding to the bronze layer, focused on collecting raw information from various sources. The data was extracted from the Fundus website, through web scraping, and from the Securities and Exchange Commission (CVM), through ZIP file downloads, covering the period from 2020 to 2024. The “ingest.py” module was used, with the “MultiSourceIngestion” superclass and the “FundusIngestion” and “CVMingestion” classes, to manage the collection and initial storage in MinIO S3.
CVM documents included Registration Forms (FCA), Standardized Financial Statements Forms (DFP) – containing Asset Balance Sheet (BPA), Liability Balance Sheet (BPP), Cash Flow Statement (DFC-MD and DFC-MI), Statement of Changes in Equity (DMPL), Statement of Comprehensive Income (DRA), Income Statement (DRE), and Value Added Statement (DVA) – and Reference Forms (FRE). The data were stored in raw format, with original encoding (iso-8859-1) and metadata in JSON, organized hierarchically by source, file type, identifiers, and extraction date.
In the silver layer, raw data underwent treatment and normalization processes to ensure quality and consistency. PySpark was used for distributed processing, connecting Python to Spark. Activities included the removal of duplicate data, handling of null information, standardization of naming and format, and anomaly detection. For data quality validation, the “run_metrics.py” module with the “SilverQualityValidator” class was implemented, which performed integrity and accuracy checks.
The specific treatments applied in the silver layer involved monetary scale conversion, CNPJ normalization, intelligent filtering to keep only consolidated accounts relevant for fundamental analysis, encoding treatment, record deduplication, data type conversion (strings to DoubleType and DateType), space trimming, and empty row removal. These steps were crucial for refining the data and preparing it for more robust analyses, eliminating inconsistencies and redundancies.
The gold layer was dedicated to modeling and making data available for final consumption. The “Star Schema” dimensional modeling (Strengholt, 2025) was adopted, which organizes data into fact and dimension tables. Dimensions included time, company, and action type, while fact tables covered balance sheet, share capital, stock price, dividends, and income statement. For analytical visualization of the processed information, Power BI Desktop (Ehrenmueller-Jensen, 2024) was used.
The main Python tools and libraries employed in the pipeline development included MinIO S3 for object storage, Apache Spark for distributed processing, Prometheus (Jani, 2024) for monitoring and observability, and Delta Lake for data governance, ensuring ACID transactions, versioning, and schema evolution. Other Python libraries, such as `requests`, `python-dateutil`, `lxml`, `pandas`, and `python-dotenv`, were used for data extraction, manipulation, and management.
The construction of the theoretical investment portfolio followed a rigorous filtering process. It began with a universe of 323 companies, of which 91 were listed on the Ibovespa index in 2021. After applying Benjamin Graham’s filters, 47 stocks were considered eligible. From these, 5 stocks were selected, prioritizing sector diversification. Quarterly average quotation data were obtained from the FATO_COTACAO table in the gold layer, and dividends from the Status Invest platform.
For the calculation of the theoretical portfolio’s percentage return, the average quotation of the first quarter of 2021 and the quotation of the last business day of 2024 were considered. Dividends were accounted for based on the last year of the analysis (2024) and redistributed proportionally to the other years to standardize the treatment of income flows. The formula used for the percentage return calculation was: (Average Stock Price in the year of analysis – Stock Price on the last analyzed year + (number of years analyzed * Dividends from the last analyzed year)) / Stock Price on the last analyzed year.
3. Results and Discussion
This section details the results obtained from the development and application of a financial data pipeline, designed to structure, process, and make available crucial information for the fundamental analysis of Brazilian stocks. The study demonstrated the effectiveness of the implemented architecture in both data organization and quality, as well as in supporting individual investor decision-making, according to Benjamin Graham’s methodology (2019). The quantitative and qualitative findings highlight significant gains in traceability, consistency, and analytical capacity, overcoming challenges inherent in the decentralization of information in the capital market.
The pipeline architecture was designed following the medallion pattern, which organizes data into distinct layers: bronze, silver, and gold, within a data lakehouse environment. This approach aims to progressively enhance the structure, quality, and validation of information from its raw form to the refinement stage for analytical consumption (STRENGHOLT, 2025). The implementation used Docker containers to ensure portability and reproducibility, integrating essential data engineering lifecycle tools, such as Python for ingestion, MinIO for object storage, Apache Spark for distributed processing, and Prometheus for observability.
The data ingestion stage, corresponding to the bronze layer, focused on collecting information from two main sources: the Fundamentus website and the Securities and Exchange Commission (CVM). Data from Fundamentus, obtained via web scraping, and from CVM, through ZIP file downloads, were stored in their raw form (“as-is”) in MinIO S3, preserving the original encoding and accompanied by detailed metadata. This metadata, organized in a dictionary, included the source, asset type, extraction date and time, HTTP request technical details, file format and size, and the ingestor version, ensuring traceability and auditing (STRENGHOLT, 2025).
The directory organization in the bronze layer was structured hierarchically, considering the ingestion source, file type, identifiers by action or document year, and the extraction date and time. This approach allowed for efficient management of raw data and the possibility of punctual reprocessing. However, it was observed that access to the financial statements history from the Fundus site became unfeasible due to the implementation of a captcha system, which made its continuous use in the ingestion script impossible.
In contrast, the CVM proved to be a robust and viable source, providing consolidated financial statement data segmented by year and enabling automations. The documents stored in the bronze layer of MinIO included Registration Information (CNPJ, registration date), Registration Forms (FCA) with trading codes, and Standardized Financial Statements Forms (DFP). The latter cover Asset Balance Sheet (BPA), Liability Balance Sheet (BPP), Cash Flow Statements (DFC-MD and DFC-MI), Statement of Changes in Equity (DMPL), Statement of Comprehensive Income (DRA), Income Statement (DRE), and Value Added Statement (DVA), in addition to the Reference Form (FRE) with comprehensive information about the issuer.
The data treatment and distributed processing occurred in the silver layer, where the `CVMSilverProcessor` class utilized PySpark to transform the raw data. This step was fundamental to resolve inconsistencies and redundancies present in the CVM data. Many financial statements, for example, presented duplicate results for the reference year and the previous year, which generated redundant information and inconsistent calculations. The exclusion of this duplicated information in the silver layer resulted in space savings and greater data consistency for future analyses.
The efficiency of the data treatment script in the silver layer was quantified through quality metrics. In the DRE table, out of 32644 original rows in the bronze layer, 1864 rows were maintained in the silver layer, representing a reduction of 94.29%. For the BPA table, out of 60741 rows, 444 were preserved, indicating a reduction of 99.27%. Similarly, in the BPP table, out of 104017 rows, 910 were maintained, resulting in a reduction of 99.13%. These results demonstrate the pipeline’s ability to eliminate redundancies and irrelevant data, optimizing the volume of information for analysis.
Other datasets also showed significant reductions: the volume of securities had a reduction of 51.62% (from 4636 to 2243 lines), and the share capital, of 65.70% (from 3606 to 1237 lines). In contrast, the dividend distribution dataset maintained its full volume (1503 lines), suggesting intrinsic consistency from the origin. The comparative analysis between the bronze and silver layers, therefore, not only highlights the volume reduction but also ensures the reliability of the information presented for subsequent analysis steps.
In the gold layer, for the loading and availability of data to the end user, the dimensional modeling known as “Star Schema” was adopted. This standard is widely used in Business Intelligence to facilitate analytical visualization and the development of dashboards, such as the one implemented in Power BI. The data were organized into dimension tables, such as `DIM_TEMPO` for temporal control, `DIM_EMPRESA` for action categorization, and `DIM_TIPO_ACAO` for financial market codes. The fact tables, in turn, included `FATO_BALANÇO` (current assets and liabilities, net equity), `FATO_CAPITAL_SOCIAL` (quantity of common and preferred shares), `FATO_COTACAO` (monetary values of share purchases), and `FATO_DIVIDENDOS` and `FATO_DRE` for net profit and dividend calculations.
The construction of the theoretical investment portfolio followed Benjamin Graham’s (2019) methodology, focusing on a defensive investor. Initially, out of the 323 companies traded in 2021, according to CVM data, 91 were listed on the Ibovespa in the same year. Applying Graham’s filters, which included criteria such as P/E x P/B lower than 22.5, current ratio not lower than 1, absence of losses, and dividend distribution, resulted in 47 eligible stocks. To diversify the portfolio and mitigate risks, five stocks were selected: ELETROBRAS (AXIA3), PORTO SEGURO AS (PSSA3), MARFRIG (MBRF3), SIMPAR S.A (SIMH3), and BRASILAGRO (AGRO3).
The percentage return of the theoretical portfolio was calculated considering the average stock price in the first quarter of 2021 and on the last business day of 2024, in addition to the dividends distributed in the last year of the analysis (2024), proportionally redistributed for methodological standardization. The formula used was: Return (%) = (Average Stock Price in the year of analysis – Stock Price on the last analyzed year + (number of years analyzed * sum of dividends from the last analyzed year)) / Stock Price on the last analyzed year. This calculation allowed for a consistent evaluation of the portfolio’s performance over the period.
The financial results of the theoretical portfolio, from January 2021 to December 2024, indicated a positive return of 32.88%. In comparison, the Ibovespa index registered a return of 4.25% in the same period, while the accumulated Selic rate reached 46.96%. The theoretical portfolio, therefore, significantly outperformed the Ibovespa, demonstrating the effectiveness of Graham’s methodology in selecting assets with appreciation potential. However, the portfolio’s return was below the Selic rate, which showed superior performance in the analyzed period.
Market volatility from 2021 to 2024, influenced by events such as the COVID-19 pandemic, can explain the difference compared to the Selic rate. The Selic rate, which varied from 1.90% in 2021 to 12.15% in 2024, reflects an atypical scenario of high interest rates, making it a more profitable investment option compared to a stock portfolio. It is important to note that fundamental indicators, while useful, do not guarantee future returns and should be complemented by a broader macroeconomic and microeconomic analysis, in addition to adequate portfolio diversification, avoiding the allocation of 100% of resources in stocks (GRAHAM, 2019).
In summary, the developed data pipeline fully met the proposed objectives, consolidating information from various companies and segmenting it according to the medallion architecture. This allowed for the analysis of both raw and processed data, facilitating the identification of inconsistencies and improving information quality. The theoretical portfolio, built based on Benjamin Graham’s methodology, demonstrated a return superior to Ibovespa, validating the fundamentalist approach as a valuable tool for individual investors, despite being surpassed by the Selic rate in an atypical market period.
4. Conclusion
The present study aimed to develop an automated financial data pipeline for the fundamental analysis of Brazilian stocks, applying Benjamin Graham’s methodology. The successful implementation of a medallion architecture was verified, which utilized MinIO S3 for bronze, silver, and gold layered storage, Apache Spark for distributed processing, Prometheus for monitoring, and Power BI for analytical visualization. The data, mostly financial statements from the CVM, were processed with significant gains in organization, traceability, and quality, evidenced by reductions in redundancies of 99.27% in the BPA table and 94.29% in the DRE. The theoretical investment portfolio, built based on Graham’s principles, presented a positive return of 32.88% in the period from 2021 to 2024, outperforming the Ibovespa index (4.25%) and demonstrating the viability of the fundamentalist approach for individual investors.
Despite the promising results, the study faced limitations, such as the infeasibility of accessing historical data from the Fundus website due to captcha systems, which restricted ingestion sources. Additionally, although the theoretical portfolio outperformed the Ibovespa, its return was lower than the Selic rate (46.96%) in the analyzed period, a scenario influenced by high interest rates and atypical market volatility, reinforcing that fundamental indicators should be complemented by macroeconomic analyses and diversification. For future studies, it is suggested to implement an orchestrator for automated scheduling and incorporate Grafana for metric visualization, aiming to improve observability. It is also recommended to expand the pipeline’s application to other capital markets, such as international stocks, real estate funds, and cryptoassets, and integrate with additional data sources, such as B3.
Bibliographic References
ASSAF NETO, Alexandre. Mercado financeiro. 14. ed. São Paulo: Atlas, 2018. ISBN-978-85-97-01805-9.
CAMPOS, Celso Ribeiro; TEIXEIRA, James; COUTINHO, Cileda de Queiroz e Silva. Reflexões sobre a educação financeira e suas interfaces com a educação matemática e a educação crítica. Educação Matemática em Pesquisa, São Paulo, v. 17, n. 3, p. 556–577, 2015. III Fórum de Discussão: Parâmetros Balizadores da Pesquisa em Educação Matemática no Brasil.
DENSMORE, James. Data pipelines pocket reference: moving and processing data for analytics. 1. ed. Sebastopol: O’Reilly Media, 2021.
EHRENMUELLER-JENSEN, Markus. Data Modeling with Microsoft Power BI. Sebastopol: O’Reilly Media, 2024.
FREITAS, Lucas Catunda de. Uma análise fundamentalista do setor bancário: um estudo de caso com uma seleção de indicadores. 2020. Trabalho de Conclusão de Curso (Graduação em Finanças) Universidade Federal do Ceará, Faculdade de Economia, Administração, Atuária e Contabilidade, Fortaleza, 2020.
GRAHAM, Benjamin. O investidor inteligente. São Paulo: HarperCollins, 2019.
Hoffman; Moro, 2012 [Referência completa não encontrada no documento original]
INFOMONEY. Ações B3: cotação de hoje. Disponível em: https://www.infomoney.com.br/cotacoes/b3/acao/. Acesso em: 17 set. 2025.
JANI, Yash. Unified Monitoring for Microservices: Implementing Prometheus and Grafana for Scalable Solutions. Journal of Artificial Intelligence, Machine Learning and Data Science, v. 2, n. 1, 2024. ISSN 2583-9888. DOI: https://doi.org/10.51219/JAIMLD/yash-jani/206.Disponível em: https://urfpublishers.com/journal/artificial-intelligence. Acesso em: 5 fev. 2026.
LUNA, Pedro Henrique Santiago de. Pipeline de dados para análise epidemiológica de casos sobre transtornos mentais relacionados ao trabalho no Brasil. Trabalho de Conclusão de Curso (Bacharelado em Sistemas de Informação) – Universidade Federal de Pernambuco, Recife, 2025.
Montóia, 2021 [Referência completa não encontrada no documento original]
Serasa, 2025 [Referência completa não encontrada no documento original]
SILVA JÚNIOR, João Fernandes da. Aplicação das estratégias de Benjamin Graham no mercado acionário brasileiro. 2020. Trabalho de Conclusão de Curso (Graduação em Ciências Contábeis) Faculdade de Ciências Contábeis, Universidade Federal de Uberlândia, Uberlândia, 2020.
STRENGHOLT, Piethein. Building medallion architectures. Sebastopol, CA: O’Reilly Media, 2025.
Article originating from the Final Course Work of the Specialization in Software Engineering of the MBA USP/Esalq
To learn more about the course, click here and access the MBX Academy platform