Article

Software Engineering

October 05, 2026

Using Generative AI for Database Selection: A Requirements-Driven Framework for Generating Architecture Decision Records (ADRs)

Using Generative AI for Database Selection: a Requirements-Oriented Framework for Generating Architecture Decision Records (Adrs)

Geovanni de Morais Gava; Luiz Fernando Pereira Nunes

DOI: 10.22167/2675-6528-202602960

Article derived from a Course Conclusion Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.

Summary

The growth of data-intensive applications and the adoption of microservices architecture have amplified the need for polyglot persistence, imposing a high cognitive load on software architects in choosing and justifying database technologies. This work aimed to propose and develop an Artificial Intelligence (AI) agent-based framework to guide technological selection and generate well-founded Architecture Decision Records (ADRs). An experimental and applied methodology was employed to build a technical knowledge base. The Retrieval-Augmented Generation (RAG) technique, along with the LangChain and LangGraph libraries, was used to orchestrate agents and anchor the responses of a Large Language Model (LLM). The framework extracted natural language requirements, enriched them with RAG, and sent them to the LLM, which generated ADRs to assist in evaluating theoretical trade-offs. The results demonstrated that the agent with RAG reduced generic responses, increasing theoretical grounding and traceability. The RAG approach proved its effectiveness against conventional prompts (zero-shot), favoring the generation of ADRs with a lower level of hallucination and a high level of theoretical traceability. It was concluded that the automated tool fulfilled the function of requirement mapping, resulting in empirically grounded technical documents and aiding governance and decision-making in software architecture.

Keywords: Databases; Artificial Intelligence; LangGraph; LLM; RAG.

1. Introduction

The evolution of information systems in recent decades has profoundly transformed the way organizations deal with the large volume of information generated daily. To process this data in a structured way, the adoption of robust technologies is indispensable. A database is understood as a collection of data that typically describes the activities of one or more related organizations (Ramakrishnan and Gehrke, 2008).

Historically, the relational model dominated this scenario, dictating how corporate applications abstracted and manipulated their records. However, modern software architecture demands an increasingly deep understanding of modeling and physical storage structures. There is a gap between the user’s view and physical storage. Ramakrishnan and Gehrke (2008) highlight that, although a data model hides low-level details, it is still closer to how the database management system stores data than to how the user thinks about the company.

Despite the strong consolidation of the relational model, the advent of data-intensive application scenarios, often referred to as Big Data, and the need for horizontal scalability through server clusters have driven the rise of NoSQL databases. This transition opened doors to a new paradigm in software design. As Sadalage and Fowler (2012) explain, we are entering a world of Polyglot Persistence, where companies and even individual applications use multiple technologies for data management. This technological flexibility, however, imposes a heavy cognitive load on software architects and data engineers, as architectural decision-making requires constant balancing between structural guarantees. The clash between classic transactional properties and distributed environments is inevitable, as relational databases use ACID transactions to handle consistency, which inherently conflicts with a cluster environment, leading NoSQL databases to offer a range of options for consistency and distribution (Sadalage and Fowler, 2012).

In daily practice, the selection of persistence technology often neglects these theoretical constraints, such as the CAP Theorem and ACID versus BASE properties. The choice ends up being based on the popularity bias of market tools or on non-standard empirical performance tests, rather than on solid foundations. With the rise of Generative Artificial Intelligence and Large Language Models (LLMs), an opportunity has emerged to intelligently automate and support system design. However, such models suffer from the phenomenon of “hallucinations”, generating technically inaccurate recommendations in highly specific domains when they lack knowledge and technical grounding.

To mitigate this challenge, the Retrieval-Augmented Generation (RAG) technique is commonly adopted, which anchors the AI’s reasoning and text generation in an external bibliographic database focused on data engineering. The use of the RAG technique is supported by literature for resolving failures based on lack of information and outdated models, as it enriches the generation of responses with external and factual sources, mitigating the risk of technical hallucinations and making the final document highly traceable (Huyen, 2025). Furthermore, RAG reduces cost and latency by selecting only the strictly relevant information, optimizing the use of the AI’s native context window token limit (Huyen, 2025). Given this scenario, the present research is justified by the opportunity to unite the advances of orchestrated AI agents to improve governance in software architecture.

The need for robust and theoretically grounded decisions in database selection, combined with the potential of AI agents to overcome the limitations of generic LLM outputs, highlights the relevance of this study. The central objective of this work is to propose and develop an intelligent framework that, equipped with classic database literature, can interpret natural language requirements, evaluate theoretical trade-offs, and autonomously generate a consistent, traceable Architecture Decision Record (ADR) with the fewest possible “hallucinations”, as well as a reduction in market biases.

2. Material and Methods

The research conducted was characterized by its experimental and applied nature, employing application development approaches with agents and Foundation Models. The procedures were carried out in controlled software engineering simulations, aiming at the construction of an intelligent framework. The main methodological focus resided in the utilization of the Retrieval-Augmented Generation (RAG) technique and autonomous agents, with the purpose of overcoming the limitations inherent to context windows and the propensity of Large Language Models (LLMs) to generate untraceable responses.

The adoption of Artificial Intelligence agents was justified by the capacity to combine diverse tools, knowledge, memory, and learning with advanced models. This combination allowed for the resolution of ambiguous and complex problems through multiple reasoning steps, as highlighted by Albada (2025). The framework’s construction was structured in distinct operational stages, detailed below, to ensure the robustness and functionality of the proposed system.

The first stage consisted of data ingestion and processing for the RAG knowledge base. For this, a knowledge base (grounding) was established from the classic literature of database systems, official tool documentation, and specific theories, such as topologies, replication mechanisms, and consistency models, including ACID versus BASE properties and the CAP Theorem. This base was fundamental to anchoring the AI’s reasoning.

The collected texts were fragmented, a process known as chunking, and converted into embeddings. This conversion was performed using a specific encoder model, Ollama with nomic-embed-text. Subsequently, the resulting embeddings were stored in a vector database, utilizing the combination of PostgreSQL and the pgvector extension, enabling indexed search for relevant information.

The Retrieval-Augmented Generation (RAG) technique was employed to mitigate failures arising from a lack of information and outdated models, enriching response generation with external and factual sources. This contributed to reducing the risk of technical hallucinations and increasing the traceability of the final document (Huyen, 2025). Additionally, RAG optimized the use of the AI’s context window token limit by selecting only strictly relevant information, consequently reducing costs and latency (Huyen, 2025).

The second stage involved the implementation of the agent, where the LangChain and LangGraph libraries were used to build the orchestration and state machine of the conversational agent. LangGraph was chosen for its ability to provide a modular orchestration framework based on directed graphs, supporting cyclic and asynchronous workflows, which is crucial for robust agent architectures (Albada, 2025).

The LangGraph modeled the Architecture Decision Record (ADR) generation process as a state graph. In this flow, an ingestor node was responsible for analyzing and validating user input, while a writer node executed the RAG search, evaluated the relevance of the retrieved content, invoked the Large Language Model (LLM), and formatted the final output in markdown and according to the ABNT standard.

The LangChain library acted as an abstraction layer for AI components, offering ready-made and interchangeable parts. This included document loading and splitting functionalities, such as PyPDFLoader for loading PDFs and RecursiveCharacterTextSplitter for splitting text into chunks with overlap. For embeddings, OllamaEmbeddings was used, which abstracted the call to Ollama for generating vectors.

For vector storage, PGVector was employed, which abstracted the logic of insertion, selection, and similarity search in PostgreSQL, eliminating the need to write SQL manually. Common interfaces, defined by langchain-core, ensured flexibility to swap embedding models or vector databases without altering the rest of the system code.

The third stage focused on prompt engineering and retrieval, where system prompts established explicit rules for the AI. The model was instructed to decompose the user’s problem, identify trade-offs based on extracted requirements, the CAP Theorem, and architectural constraints retrieved by the RAG flow, before making the final decision. The complete prompt sent to the LLM was composed of four distinct layers.

The first layer, the System Prompt, defined the LLM’s role as senior architect, the writing rules, and the ADR structure, including mandatory sections such as Context, Decision, Consequences, Alternatives, and References, in addition to formatting and ABNT citation instructions. The second layer, of expanded requirements, consisted of six multiple-choice questions about data structure, nature of searches, operation proportion, volume and scalability, CAP Theorem, and multi-record transactions, whose answers were transformed into readable text.

The third layer, for follow-up context, was optional and used when the user indicated “I’m not sure” to a question in the previous layer. In this case, two Yes/No context questions were asked for clarification, enriching the prompt. The fourth layer, for RAG context, attached relevant chunks found by RAG to the end of the prompt, providing technical knowledge and bibliographic references to support the answers.

The fourth and final stage comprised the evaluation and validation scenarios of the framework. The system was subjected to simulated tests based on three typical System Design patterns: high event ingestion with priority on linear write scalability (AP), distributed financial transactions with priority on strong consistency and durability (CP/ACID), and a product catalog with flexible searches.

To evaluate the effectiveness of the approach, the Architecture Decision Records (ADRs) generated with the support of the RAG flow and agents were compared with the responses obtained by the zero-shot approach, which relied solely on the model’s parametric knowledge. The evaluation was qualitative in nature, focusing on factual consistency and traceability of justifications, verifying coherence with the architecture references of the technologies. Annexes V and VI of the original work contain the ADRs generated for the scenarios with and without the use of RAG, respectively.

3. Results and Discussion

The implementation of the proposed framework demonstrated the viability of using conversational agents, guided by the Retrieval-Augmented Generation (RAG) technique, for executing complex analytical tasks, such as the selection and technical justification of databases. The execution rounds, conducted in controlled validation scenarios, showed significant differences between the recommendations generated by a Large Language Model (LLM) in a zero-shot approach and those enriched with external bibliographic context, provided by RAG. This distinction highlighted the framework’s ability to mitigate biases and hallucinations, promoting more informed and traceable decisions.

The tests were structured into three distinct scenarios, each representing a common system design pattern, allowing for a comprehensive evaluation of the framework’s performance. The comparative analysis focused on factual consistency, depth of justification, and traceability of the theoretical sources used. It was observed that the absence of RAG frequently led to choices based on popularity or superficial analogies, while its application directed the LLM towards more robust solutions aligned with the specific technical requirements of each context.

Event Ingestion (Linear Writing / AP)

In the scenario simulating massive ingestion of telemetry and logs, where data scale could reach Petabytes and availability (AP) was the top priority in the face of partial network failures, the LLM’s zero-shot approach suggested Redis. This recommendation was based on write speed and the simplicity of the key-value structure, but it ignored the physical RAM constraints and persistence for such high volumes, which would make the cost prohibitive and increase the risk of data loss in case of power failure. Alternatives like MySQL, PostgreSQL, and MongoDB were discarded because they did not support the write volume or were not optimized for pure key-value operations.

In contrast, the RAG-enhanced framework, when processing the same requirements, recommended ScyllaDB. The decision was based on ScyllaDB’s ability to handle large data volumes through its column-family and key-value model, optimized for high-speed access by primary key (Sadalage; Fowler, 2012). The agent highlighted ScyllaDB’s “shard-per-core” architecture, ideal for intense write loads (ScyllaDB, 2023), and its linear scalability across multiple servers (Scabora, 2016). Furthermore, the implementation of the masterless model favors availability in the CAP theorem, with adjustable eventual consistency (ScyllaDB, 2023; Sadalage; Fowler, 2012).

The positive consequences of choosing ScyllaDB included native high availability and linear scalability to Petabytes, along with very low write latency due to the absence of global locks (ScyllaDB, 2023). However, disadvantages were identified, such as inadequacy for queries requiring multiple complex JOINs (Scabora, 2016) and the need for query-oriented data modeling. The alternatives considered and discarded by the agent with RAG were Redis, due to the unfeasible RAM cost for massive volumes; PostgreSQL, because its concurrency control limits massive writing (Ramakrishnan; Gehrke, 2008); and Neo4j, as it is optimized for connections and not for high log ingestion (Robinson et al., 2015).

Financial Transactions (Strong Consistency / CP)

For the scenario of distributed financial transactions, which required a high degree of consistency (CP) and guarantee of isolated transactions, the LLM in a zero-shot approach suggested MongoDB. Although the model attempted to justify the choice based on its ability to organize data as tables and support for multi-document transactions, the descriptions regarding fault tolerance and the superiority of the relational model for strict referential integrity were inaccurate, characterizing a hallucination of non-parametric knowledge. MongoDB is not a native relational database, which could cause inefficiency in complex JOINs, and ensuring strict ACID would require configurations that would reduce performance. Alternatives such as Cassandra, SQLite, and Neo4j were discarded for not supporting strict ACID transactions or not being suitable for the financial context.

With the integration of RAG, the framework successfully retrieved fragments from Ramakrishnan and Gehrke (2008) literature on the Two-Phase Locking Protocol (2PL) and ACID properties. Based on this foundation, the agent recommended PostgreSQL, prioritizing “Consistency” over partition failures in the CAP Theorem. The decision was justified by the relational model which guarantees integrity through fixed schemas (Ramakrishnan; Gehrke, 2008), support for complex JOINs and aggregations via SQL, and efficiency with mixed loads under concurrency. PostgreSQL is optimized for high-performance single servers and implements locking protocols that guarantee CP Consistency (Ramakrishnan; Gehrke, 2008), with native ACID guarantees fundamental for financial operations (Hiremath, 2018).

The positive consequences of choosing PostgreSQL included robustness in concurrent transaction control via isolation (Ramakrishnan; Gehrke, 2008) and the maturity of the ecosystem for financial auditing (Hiremath, 2018). As negative points, were highlighted the lower availability in case of severe network failures, compared to AP systems, and the difficulty of horizontal scalability in relation to NoSQL databases (Sadalage; Fowler, 2012). The alternatives discarded by RAG were MongoDB, because the rigid tabular structure and JOINs are natively superior in RDBMS (Ramakrishnan; Gehrke, 2008); Cassandra, for prioritizing availability over strong ACID consistency (Sadalage; Fowler, 2012); and ScyllaDB, for not being focused on multi-table ACID transactions (ScyllaDB, 2023).

Product Catalog (Flexible Searches / AP)

In the scenario of a product catalog with flexible searches, where the priority was availability (AP) and the need to handle flexible data (JSON) and textual searches, the LLM’s zero-shot approach suggested Neo4j. This choice was based on the superficial premise that “products have categories,” which led to an inadequate analogy with graph databases. Although Neo4j is excellent for recommendation systems and intuitive product tree visualization, it is not optimized for large-scale “fuzzy search” textual searches, and its performance can degrade in searches that do not utilize relationships. Alternatives such as Elasticsearch, PostgreSQL, and Cassandra were discarded for being mere search engines, too rigid for dynamic attributes, or complex for textual searches.

The RAG-enhanced framework, in turn, correctly identified MongoDB as the most suitable solution. The decision was based on MongoDB’s ability to store dynamic product attributes in JSON documents (Sadalage; Fowler, 2012) and to offer text indexes for flexible searches (Płuciennik; Zgorzałek, 2017). The agent highlighted MongoDB’s efficiency for massive read workloads, its sharding support for massive data volumes (Sadalage; Fowler, 2012), the possibility of configuration for high availability (AP), and the acceptability of eventual consistency for product catalogs.

The positive consequences of choosing MongoDB included development agility due to the direct mapping between objects and documents (Płuciennik; Zgorzałek, 2017), as well as the ease of expansion into new markets and product categories. However, disadvantages were observed, such as higher disk space consumption than the relational database model due to denormalization (Sadalage; Fowler, 2012) and the lack of optimization for complex relationship queries, typical of graphs (Robinson et al., 2015). The alternatives discarded by RAG were Neo4j, as graph complexity was not necessary for textual catalog searches (Robinson et al., 2015); ScyllaDB, for not having the flexibility of dynamic fields of the document model (ScyllaDB, 2023); and PostgreSQL, due to schema rigidity that hinders variable attributes (Płuciennik; Zgorzałek, 2017).

The quality of the tool’s final output was expressed in the generation of Architecture Decision Records (ADRs) with analytical rigor. The model was guided by explicit instructions inserted into the prompt, limiting its use solely to the extended knowledge base. This allowed the framework to overcome the common instability of synthetic generation, attesting that the software’s intrinsic systemic architecture serves as the primary selection criterion, mitigating reliance on empirical internet benchmarks, which are often inconsistent. The ability to generate traceable and theoretically grounded ADRs represents a significant advancement in software architecture governance and decision-making, validating the effectiveness of the RAG approach for enhancing generative intelligence in specific technical domains.

4. Conclusion

The present study aimed to propose and develop an intelligent framework, based on Artificial Intelligence agents, to guide the selection of database technologies and generate well-founded Architecture Decision Records (ADRs). It was found that the Retrieval-Augmented Generation (RAG) approach, combined with agent orchestration, proved effective in interpreting natural language requirements and evaluating theoretical trade-offs. The results showed a significant reduction in generic responses and “hallucinations” compared to language models in a zero-shot approach, increasing the theoretical grounding and traceability of recommendations. It was observed that the framework was capable of providing robust solutions aligned with specific technical requirements in diverse scenarios, such as massive event ingestion, financial transactions with strong consistency, and product catalogs with flexible searches, recommending ScyllaDB, PostgreSQL, and MongoDB, respectively, with justifications anchored in the literature.

The main contribution of this work lies in the creation of an automated tool that enhances governance in software architecture, by mapping requirements and generating empirically based technical documents. The ability to produce ADRs with analytical rigor and theoretical traceability represents a significant advance, as it mitigates dependence on market biases or inconsistent empirical benchmarks, prioritizing the intrinsic systemic architecture of the software as the primary selection criterion. Thus, the developed framework directly assists software architects in the complex decision-making process, promoting more informed and reliable choices.

Bibliographic References

Abhinav Hiremath; Comparative Analysis of Relational vs NoSQL Databases in Web Applications Dataset. Vol. 06, Issue: 03, March: 2018

ALBADA, Michael. Building Applications with Al Agents: Designing and Implementing Multiagent Systems. 1. ed. Santa Rosa: O’Reilly Media, 2025

HUYEN, Chip. Al Engineering: Building Applications with Foundation Models. 1. ed. [S. I.]: O’Reilly Media, 2024.

PŁUCIENNIK, E.; ZGORZAŁEK, K. The multi-model databases-a review. In: SPRINGER. 2017.

Ramakrishnan, R.; Gehrke, J. Sistemas de gerenciamento de banco de dados. 3. ed. McGraw Hill Brasil, 2008.

ROBINSON, I.; WEBBER, J.; EIFREM, E. Graph databases: new opportunities for connected data. [S.I.]: “O’Reilly Media, Inc.”, 2015.

Sadalage, P. J.; Fowler, M. NoSQL Distilled: A Brief Guide to the Emerging World of Polyglot Persistence. Addison-Wesley, 2012.

SCABORA, L. D. C. Avaliação do Star Schema Benchmark aplicado a bancos de dados NoSQL distribuídos e orientados a colunas. 2016.

ScyllaDB. ScyllaDB Architecture Overview. ScyllaDB Documentation, 2023.

Article originating from the Final Course Work of the Specialization in Software Engineering of the MBA USP/Esalq

To learn more about the course, click here and access the MBX Academy platform

You may also like

Software Engineering

October 09, 2026

O papel da densidade de texto instrutivo na eficiência de uma aplicação web.

O desenvolvimento de aplicações web se conecta à experiência do usuário, e este trabalho investigou como o uso excessivo de textos instrutivos pode retardar a conclusão de tarefas e impactar a eficiência da aplicação. O objetivo foi identificar o impacto da densidade textual do conteúdo instrutivo na eficiência de uma aplicação web, utilizando como principal referência a terceira lei de usabilidade de Krug. A pesquisa, de caráter exploratório e delineamento experimental quantitativo, empregou um teste A/B em uma aplicação web responsiva, onde a única variável controlada foi a densidade textual (alta vs. baixa, definida pela contagem de palavras). Participaram 25 usuários, e os dados foram coletados via Datadog RUM, mensurando tempo de conclusão, erros de submissão e taxa de conversão. Os resultados revelaram que a variante com densidade textual reduzida (variante B) apresentou uma taxa de conversão superior (58,3% contra 33,3% da variante A) e um tempo médio de conclusão significativamente menor (1:38 minutos contra 4:58 minutos da variante A), representando um aumento de 67,12% na eficiência. O teste t de Welch (p=0,042) confirmou que a redução da densidade textual impactou a eficiência. Concluiu-se que a redução da densidade textual afeta a eficiência e a taxa de conversão, reforçando a importância de conteúdo objetivo e conciso. Contudo, a baixa densidade textual, por si só, não garantiu o pleno entendimento, sendo essencial a comunicação clara e objetiva das instruções, validando a relevância do UX Writing.

Palavras-chave: Eficiência; Experiência de usuário; Teste A/B; Texto Instrutivo; Usabilidade.

Software Engineering

October 09, 2026

Implementation of analytics systems in automation environments in the process industry

The digitalization of process plants depends on structured data collection and storage, without which there is no operational visibility. Industrial analytics is the name given to the chain that processes this data from signal acquisition in the field instrument, through time-ordered storage, to its availability for analysis and other systems. Proprietary industrial software currently covers this chain, and licensing and maintenance costs restrict its adoption. The study aimed to architect, implement, and validate a system of this nature, employing current software development techniques and components at no licensing cost. The system, named Sistema de Aquisição e Tratamento de Informações de Processo (SATIP) [Process Information Acquisition and Processing System], was structured in three independent layers, using a pharmaceutical reactor simulator as a data source. The simulator executed an eight-step recipe and generated time series with stochastic variation. The simulator was written in Go, a language adopted for generating self-contained binaries, suitable for execution on edge equipment. Storage used TimescaleDB, and data was made available through a REST interface with a web dashboard. The system processed approximately 414,000 records per execution, with multivariate trends, alarm logging, and temporal correlation between instruments. The architecture proved to be reproducible and without licensing costs, and the identification of degradation required only the readings already stored by the system.

Keywords: Software architecture; Digitalization; IT/OT integration; Predictive maintenance; Time series.

Software Engineering

October 09, 2026

Use of voice in conjunction with large language models as a tool for digital accessibility

Speech is an essential basis for human interaction, and for people with disabilities, it can represent the primary form of communication with the external environment. Given the growing technological influence, voice command identification has emerged as a promising strategy for human-machine interaction. The work explored how voice, in conjunction with Large Language Models (LLMs), can be used efficiently, naturally, and accurately. For this purpose, various design patterns and the Python language were employed, aiming for greater extensibility. Gemini was used as the LLM provider, sending audio directly and leveraging its function calling capability to interact with the device. A system was developed capable of understanding user intent and converting it into actions, whose differential was the computer vision capability based on screenshots and a mesh system for LLM guidance. Tests revealed a user intent comprehension rate of 91.81% and a success rate in execution of 75.45% with the “Flash-3” model (top p 0.5 and top k 5). The proposed system validated the premise that the integration of LLMs into voice interfaces increases the autonomy of users with motor disabilities, fulfilling the purpose of being a modern and effective Assistive Technology. However, questions were raised about the costs of AI and user data security, indicating the need for improvement.

Keywords: Function calling; Human-machine interaction; Voice recognition; Computer vision.

Software Engineering

October 09, 2026

ReasonGuard: A Reasoning Audit Platform for Language Models Based on Structured Thought Decomposition

ReasonGuard was presented, an Artificial Intelligence reasoning auditing platform, developed as an observability middleware between client applications and large language models (LLMs). The work aimed to offer transparency and traceability for AI-based decision-making processes, addressing the gap in operationalizing structured reasoning techniques for auditing purposes. The system intercepted, analyzed, and documented interactions with LLMs through five modules based on the Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT) paradigms. These modules captured reasoning trails, detected structural logical flaws, evaluated response consistency, and generated audit reports targeted at different stakeholder profiles. The platform was implemented with FastAPI (Python) on the backend, React with TypeScript on the frontend, and PostgreSQL as a relational database, following a modularized architecture with multitenancy isolation per user. The results demonstrated the technical feasibility of the proposed approach, with all five modules operational and integrated. It was concluded that ReasonGuard contributes to AI governance by instrumentalizing structured reasoning paradigms as observability and auditing tools, filling a gap in the literature.

Keywords: AI Auditing; Chain-of-Thought; AI Governance; Large Language Models; Algorithmic Transparency.

Software Engineering

October 09, 2026

EngTT: software for road freight based on operational costs and ANTT parameters

Road transport represents the main logistics modality in Brazil, with the composition of freight costs regulated by the National Land Transport Agency (ANTT). Given the absence of a structured methodology for freight calculation and the dependence on isolated spreadsheets in the sector, the EngTT software was developed. The objective was to create a decision support tool that integrated operational variables, calculated the total cost per route, and compared the results with the ANTT’s minimum floor, identifying non-compliance and margin compression scenarios before pricing. The adopted methodology was applied, quantitatively and experimentally, with incremental construction oriented towards the Minimum Viable Product (MVP) concept. The system was structured in three modes of use – Build Visual Route, Batch by spreadsheet, and Scrape routes – sharing a decoupled calculation engine and centralized parameters. The results obtained demonstrated that the software met the proposed objective, showing margin variations and regulatory compliance between the analyzed routes. It was observed that longer routes, such as Ribeirão Preto × Guarujá, presented an increase in cost due to the need for additional driver per diems, while shorter routes, such as Cajamar × Guarujá, showed greater adherence to the ANTT’s regulatory floor. It was concluded that the developed solution is applicable to the road freight transport sector as an effective tool for pricing and operational management.

Keywords: ANTT; Containerized cargo; Software engineering; Road freight; Python.

Software Engineering

October 08, 2026

Data pipeline for stock monitoring considering the fundamentalist methodology

The Brazilian financial scenario faces challenges such as family indebtedness, low financial literacy, and the decentralization of information for investment analysis. Given this, the research objective was to develop a financial data pipeline, based on good Data Engineering practices, to structure, process, and make relevant information available to support individual investors’ decision-making in stock analysis. A medallion architecture was implemented, using MinIO S3 for bronze, silver, and gold layer storage, with Delta Lake for data governance. Apache Spark was employed for distributed processing, Prometheus for monitoring, and Power BI for analytical visualization. The data were mostly financial statements from the Securities and Exchange Commission (CVM). A theoretical investment portfolio was built based on Benjamin Graham’s (2017) principles, applying selection filters and data validation. The results indicated a positive return of 32.88% for the theoretical portfolio, outperforming the Ibovespa index (4.25%) in the period from 2021 to 2024, although lower than the Selic rate (46.96%). Qualitatively, the pipeline processed voluminous datasets, with significant reductions in redundancies, such as 99.27% in the BPA table and 94.29% in the DRE, after applying filters. Gains in data organization, traceability, and quality were evidenced, enabling structured and more robust financial analyses.

Keywords: Medallion Architecture; Data Engineering; Investments; Data Pipeline.

Software Engineering

October 08, 2026

Performance and Total Cost of Ownership of Cloud Databases and On-Premises Infrastructure

This study analyzed the performance of database operations in cloud and on-premises infrastructures, correlating technical behavior with Total Cost of Ownership projection. The research was characterized as a quantitative and experimental study, in which a test system subjected isolated database instances to progressive execution loads, measuring the impact of network latency and computational consumption. For financial analysis, an investment and operational expenses model was developed, diluted over a thirty-six-month cycle. Technical results revealed that the accumulation of internet latency caused severe time degradation in cloud executions, despite the remote infrastructure operating with high processing idleness, registering almost 98% CPU inactivity. In contrast, the on-premises environment achieved superior performance supported by almost instantaneous network communication. Financially, cost consolidation demonstrated an empirical tie between the physical acquisition model and the service subscription model within a three-year horizon, but the on-premises environment proved more advantageous in a sixty-month cycle. It was concluded that the degradation in the cloud was not due to computational capacity, but to the interaction between route latency and the application’s unitary communication pattern. Cloud adoption requires deep optimization of the system architecture to minimize dependence on constant communication with the remote server. Without this modernization, the on-premises infrastructure consolidated as the most viable strategy, ensuring high performance, budgetary predictability, and data sovereignty.

Keywords: Operational expenses (OPEX); Transactional scalability; Network latency; Legacy systems.

Software Engineering

October 08, 2026

Asynchronous Slack-Jira integration via message queue “middleware”: comparison of “cloud computing” solutions

The latency and interoperability between distributed corporate systems constituted the problem investigated, motivated by the costs and fragilities of manual integrations between collaboration and project management platforms. An asynchronous integration “middleware” between Slack and Jira was developed and validated, with the objective of reducing the perceived user response time and ensuring system stability under load. The methodology consisted of building an event-driven software architecture in Python, using the “producer-consumer” and “adapter” patterns to isolate the user interface from “backend” processing. The solution evolved into a cloud-agnostic architecture, based on “serverless” functions, and was subjected to stress tests in local, real network, and production environments on two clouds. The architecture reduced user waiting time from a synchronous estimate of 2,000 milliseconds to a local average of 8.96 milliseconds. Under a load of 50 simultaneous requests, both cloud providers proved viable: Amazon Web Services registered lower average latency in the reception layer (951.35 ms) and double the throughput, while Microsoft Azure executed background processing with a median of 113 ms. The application of “optimistic UI” ensured fluidity of use, and load leveling by queues eliminated the need for infrastructure over-provisioning. The “middleware” consolidated itself as a scalable, resilient, and protected corporate reference model against technological lock-in.

Keywords: Event-driven architecture; Temporal decoupling; Operational efficiency; Information technology service management; System interoperability.