Article

Software Engineering

October 09, 2026

ReasonGuard: A Reasoning Audit Platform for Language Models Based on Structured Thought Decomposition

Reasonguard: A Reasoning Auditing Platform for Language Models Based on Structured Thought Decomposition

Lucas Nastari Ziza; Lucas Cesar Gomes Alvarinho Squillante

DOI: 10.22167/2675-6528-202603200

Article derived from a Course Conclusion Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.

Summary

ReasonGuard was presented, an Artificial Intelligence reasoning auditing platform, developed as an observability middleware between client applications and large language models (LLMs). The work aimed to offer transparency and traceability for AI-based decision-making processes, addressing the gap in operationalizing structured reasoning techniques for auditing purposes. The system intercepted, analyzed, and documented interactions with LLMs through five modules based on the Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT) paradigms. These modules captured reasoning trails, detected structural logical flaws, evaluated response consistency, and generated audit reports targeted at different stakeholder profiles. The platform was implemented with FastAPI (Python) on the backend, React with TypeScript on the frontend, and PostgreSQL as a relational database, following a modularized architecture with multitenancy isolation per user. The results demonstrated the technical feasibility of the proposed approach, with all five modules operational and integrated. It was concluded that ReasonGuard contributes to AI governance by instrumentalizing structured reasoning paradigms as observability and auditing tools, filling a gap in the literature.

Keywords: AI Auditing; Chain-of-Thought; AI Governance; Large Language Models; Algorithmic Transparency.

1. Introduction

Large Language Models (LLMs) are deep learning systems, typically based on the Transformer architecture, trained on voluminous data structures to predict the most probable continuation of a text sequence. From this mechanism of statistical token prediction, complex capabilities of language generation, question answering, summarization, and, more recently, multi-step reasoning emerge. It is precisely this emergent reasoning capability and its application to high-impact decisions that motivates the present work.

The growing adoption of LLMs in decision-making contexts, including areas such as health, law, finance, and corporate governance, has highlighted a significant gap: the visualization of the reasoning processes that underpin the responses generated by these systems. An illustrative example of this problem is the “plausible but fallacious reasoning” error. When asked about the viability of a contract with a clause for automatic termination in case of payment delay, an LLM may correctly conclude that the contract is terminable, but base this conclusion on an unstated and incorrect premise, such as erroneously assuming that any delay triggers the clause. The final answer sounds convincing and is often correct by coincidence, but the underlying reasoning contains a structural flaw that, in a real-world scenario, could lead to a mistaken decision without explicit signs of error for the user. This type of failure, undetectable from a simple reading of the final answer, is the central problem that ReasonGuard seeks to address.

Although LLMs demonstrate expressive capabilities in complex reasoning tasks, the absence of formal auditing and traceability mechanisms compromises their applicability in scenarios requiring accountability and regulatory compliance. From a regulatory standpoint, the global landscape has been moving towards stricter transparency requirements. The European Artificial Intelligence Regulation (European Union, 2024), for example, establishes specific transparency obligations for high-risk AI systems, including documentation, traceability, and explainability requirements. In the academic context, recent research on accountability infrastructure for AI highlights that accountability remains fragile because process transparency is rarely recorded in a durable and auditable manner (Mokander et al., 2025). It is important to demarcate that, as they are neural networks with billions of parameters whose internal processing is not directly interpretable, LLMs remain, in essence, “black boxes”. What ReasonGuard and structured reasoning decomposition techniques offer is not the opening of this black box, but rather the induction of the model to express, in natural language, a step-by-step justification that can be recorded, versioned, and audited. This, therefore, concerns the traceability of the response generation process, a verifiable record of what the model declared as its reasoning, and not full transparency or reproducibility of the model’s internal decision mechanism, a distinction that guides the entire platform design.

The concept of structured reasoning in language models gained prominence with the work of Wei et al. (2022), who demonstrated the Chain-of-Thought (CoT) technique. This technique, by explicating intermediate reasoning steps before the final answer, significantly improves performance on arithmetic, common sense, and symbolic reasoning tasks. Building on this foundation, Yao et al. (2023) proposed the Tree of Thoughts (ToT) framework, which generalizes CoT by allowing the deliberate exploration of multiple reasoning paths, treating each intermediate step as a node in a search tree. Besta et al. (2024) further advanced this by introducing the Graph of Thoughts (GoT), which models the information generated by the LLM as an arbitrary graph of propositions and their dependencies. In parallel, Wang et al. (2023) proposed the Self-Consistency approach, a decoding strategy that explores multiple reasoning paths and selects the most consistent answer among them, elevating CoT performance on benchmarks.

Despite these advances, a specific gap is observed in the literature: the works cited above focus on two distinct and, to date, unintegrated fronts. The first front treats these techniques (CoT, ToT, GoT, and Self-Consistency) as methods to improve LLM performance on reasoning tasks, without concern for the persistent, verifiable, and auditable record of the generated process. The second front, the works on auditing and explainability (Mokander et al., 2023; Zhao et al., 2024; Bilal et al., 2025; Mokander et al., 2025), discuss what should be audited and propose conceptual verification layers, but do not present a reference implementation that systematically captures, structures, and persists, in a production-ready manner, the decomposed reasoning according to these paradigms. Therefore, there is no platform that operates at the intersection of these two fronts: one that uses CoT, ToT, and GoT not to improve the model’s response, but as observability instruments, capturing the decomposed reasoning, persisting it with integrity guarantees, and making it available in formats usable by compliance, legal, and engineering teams. It is this operationalization gap, the absence of a bridge between the reasoning decomposition technique and the auditing practice in production environments, that the present work seeks to fill.

In this context, the present work proposes ReasonGuard, a reasoning auditing platform that instrumentalizes reasoning paradigms as observability mechanisms, operating between client applications and LLM providers. The general objective of this study is to develop and validate ReasonGuard, a reasoning auditing platform for language models based on structured thought decomposition, with the following specific objectives: (i) capture complete reasoning trails with cryptographic integrity guarantee; (ii) detect structural logical flaws such as contradictions, circularity, logical leaps, and hidden premises; (iii) evaluate response consistency through multiple executions with semantic variations; and (iv) generate audit reports targeted at different stakeholder profiles, contributing to the governance and regulatory compliance of AI-based systems.

2. Material and Methods

The research conducted has an applied nature, oriented towards the generation of practical knowledge to solve the lack of transparency and auditability in the reasoning processes of Large Language Models (LLMs). The study was characterized as exploratory-experimental, combining the investigation of reasoning decomposition techniques with the implementation and validation of modules in a controlled environment. The adopted methodological approach was quali-quantitative, encompassing both the structural analysis of the extracted reasoning and the calculation of numerical metrics, such as validity scores and consistency rates.

The development of ReasonGuard followed an iterative and incremental methodology, based on evolutionary prototyping, structured in six sequential phases. Initially, the system’s requirements and architecture were defined, identifying the five main modules and adopting the API Proxy architectural pattern. Next, the backend was implemented with the core modules, followed by the implementation of the frontend with an interactive dashboard. Subsequent phases included integration and authentication with multitenancy isolation, the construction of a demo application in Streamlit, and finally, the optimization and refinement of performance and usability.

The ReasonGuard system architecture was designed as a reverse proxy with observability features, acting as middleware between client applications and LLM providers. This arrangement allowed for transparent interception of requests without the need for client application modification, enabling real-time analysis and independent storage of audit data. For the client, communication remained unchanged, requiring only the redirection of requests to the proxy’s endpoint, which preserved the expected input and output structure.

During the communication flow, the API Proxy recorded the relevant request information, applied the tracking, validation, and analysis mechanisms defined by the platform, and forwarded the call to the model provider. After processing by the provider, the response was returned to the proxy, which analyzed it, associated it with the original request data, calculated metrics, and generated the audit trail before passing it back to the client. In this way, the LLM provider remained responsible for model inference, while the middleware centralized observability, control, and auditing functions.

The ReasonGuard infrastructure was composed of four services orchestrated via Docker Compose. It included a PostgreSQL database for relational persistence, a backend implemented with FastAPI in Python containing the five core modules and nine REST API routes, a frontend developed with React and TypeScript for the interactive dashboard, and a demo application in Streamlit as an integration example. Multitenancy isolation was guaranteed by user filtering in all database queries, ensuring that each user accessed only their own records.

The core modules were developed as autonomous units, each with a unique responsibility, and progressively integrated into the complete system. The Reasoning Tracker, the first module, was based on the Chain-of-Thought (CoT) paradigm. It operated by injecting structured prompts into the LLM requests, inducing the model to decompose its reasoning into premises, inferences, and conclusions. Parsing was performed via regular expressions, and each record was assigned an SHA-256 hash to ensure data integrity.

The Path Analyzer, the second module, extended the concept to multiple paths, implementing a tree search with heuristic evaluation inspired by the Tree-of-Thought (ToT) paradigm. The problem was decomposed into two to four subproblems, for which two to three hypotheses with feasibility scores, ranging from zero to one hundred, were generated. Paths with scores below forty were pruned and documented. Captured metrics included the total nodes explored, the total paths pruned, and the score of the selected path.

The Logical Validator, the third module, transformed the extracted reasoning into a directed graph of propositions, using the NetworkX library, according to the Graph-of-Thought (GoT) paradigm. Each node represented a proposition (premise, inference, conclusion, or hidden premise) and each edge, a logical relation. The module detected four types of logical flaws: contradictions, logical leaps identified as disconnected nodes, undeclared hidden premises, and circularity, the latter detected by cycle-finding algorithms. A validity score, from zero to one hundred, was calculated based on the proportion of problems in relation to the total number of propositions.

The Consistency Checker, the fourth module, applied a repetition testing methodology inspired by the concept of Self-Consistency (Wang et al., 2023). The original query was semantically reformulated into five variations, preserving the original intent. The generated responses were compared pair-wise to identify the permanence of central information and detect divergences. A convergence rate, expressed on a scale of zero to one, was calculated, along with a final confidence score that combined this base rate with penalties for divergences and bonuses for high consistency.

The Audit Trail Generator, the fifth module, consolidated data from all previous modules into multi-format reports. These reports were generated in PDF via ReportLab, Excel via openpyxl, and native JSON, and directed to three stakeholder profiles: compliance (with integrity hashes and chronological trails), legal (with chain of evidence and divergence points), and technical (with complete data, raw structures, and detailed metrics).

The choice of the technological stack was guided by criteria of suitability for the AI domain, ecosystem maturity, and development productivity. The backend was implemented in Python 3.11 with FastAPI 0.109.0, leveraging native support for asynchronous operations, automatic validation via Pydantic, and OpenAPI documentation. The NetworkX 3.2.1 library provided the necessary graph analysis algorithms for the Logical Validator.

The frontend was built with React 18.2.0 and TypeScript 5.3.3, utilizing React Flow 11.10.2 for interactive graph and tree visualization, Recharts 2.10.4 for KPI charts, and Material-UI 5.15.6 as the interface framework. Server state management employed TanStack Query for caching and synchronization. Authentication was managed by Clerk 4.30.0, with support for OAuth (Google) and email/password, while PostgreSQL 14 provided relational persistence with JSONB field support for storing flexible structures.

The validation strategy followed a three-level approach. In functional validation, each module was individually validated for parsing correctness, pattern detection, and metric calculation. Parsing corresponded to the process of interpreting and structuring the received data, correctly extracting fields, reasoning steps, and other relevant information from the model’s responses for subsequent analyses.

Integration validation encompassed the complete functioning flow of the proxy, from receiving the request sent by the client application to forwarding it to the LLM provider. It included processing by the audit modules, persistence of results, and returning the response to the user. End-to-end authentication mechanisms were also verified, including user identification, resource access authorization, and data isolation throughout the execution.

The interface validation focused on the correct presentation of information produced by the system. The rendering of graphs and reasoning trees with real data, the visual organization of nodes and connections, the display of analysis results, and the updating of performance indicators on the dashboard were evaluated. Efforts were made to ensure that changes in processed data were correctly reflected in the KPIs and other visual components of the platform.

3. Results and Discussion

The results obtained in the development of ReasonGuard demonstrate the technical viability of the proposed approach to operationalize structured reasoning techniques as traceability and auditing tools in Artificial Intelligence governance environments. The platform was implemented completely with respect to the specified requirements, with all five core modules fully functional and integrated. It is essential to situate these findings as evidence of a prototype’s functionality, and not as a statistical validation on a production scale, a crucial distinction for the interpretation of its current contribution.

The system architecture was designed to ensure data integrity and traceability. The data model centers on the user entity, linking to it all records generated by the reasoning tracking, path analysis, logical validation, and consistency verification modules. Reasoning chains are stored as ordered sequences of steps, while path analyses are represented by hierarchical structures with parent-child node relationships. Logical validations, in turn, are modeled as graphs of propositions and directed connections, with all these associations implemented by foreign keys to ensure referential integrity in relation to the user.

On the backend, each module operates with its own routes and services, using SQLAlchemy to process requests and persist results. This modular organization allows analyses to be executed independently, but later integrated into the generation of audit reports. The frontend, in turn, was structured into ten pages dedicated to specific functionalities, including an interactive dashboard, an integrated chat, and visualizations for reasoning tracking, path analysis, logical validation, consistency checking, audit reports, and system settings.

The Reasoning Tracker, the first module, implements the interception of requests and the injection of Chain-of-Thought instructions into a structured object. For each processed request, the module records the declared premises, the inferences made, the final conclusion, and a confidence level, ranging from 0 to 100, for each reasoning step. The parsing of this information is performed using regular expressions, and each record receives a SHA-256 hash, which acts as a “fingerprint” to ensure integrity and detect any undue alteration in the content. This feature is essential for compliance and regulatory audit scenarios, where the immutability of the reasoning evidence is irrefutable.

The reasoning tracker’s visualization interface displays in detail the original prompt, integrity hash, and reasoning steps. For example, in a query about an autonomous car’s decision to swerve to save five pedestrians at the expense of the passenger, the system breaks down the reasoning into premises such as “An autonomous car has the ability to make decisions in emergency situations” and “Saving five lives is considered morally preferable to saving one life”, each with 100% confidence. Subsequent inferences, such as evaluating consequences and choosing the option that minimizes loss of life, are also recorded with their respective confidence levels, culminating in the final conclusion based on utilitarian logic.

The Path Analyzer, the second module, extends the concept of reasoning to multiple paths, implementing a tree search with heuristic evaluation inspired by the Tree-of-Thought paradigm. The problem is decomposed into two to four subproblems, for which two to three alternative hypotheses are generated, each with a feasibility score between 0 and 100. Paths with scores below 40 are pruned and documented, allowing the identification of lines of reasoning that, although not presented to the end-user, may be relevant or reveal potential flaws. The confidence index assigned to each trajectory is a self-assessment declared by the LLM itself, not an objective external metric of correctness.

The interactive decision tree visualization, implemented with React Flow, allows the user to navigate through explored nodes, identify pruned paths, and the selected path. In a scenario of employee layoffs, for example, the system can present the hypothesis of “Openly communicate the reasons for layoffs” with 85% self-assessment, and “Offer voluntary resignation packages” with 75%. A hypothesis such as “Keep employees in the dark until dismissal” can be pruned with 30% self-assessment, indicating low viability. The metrics captured by this module include the total number of nodes explored, the total number of pruned paths, and the score of the selected path, providing a comprehensive view of the solution space explored by the LLM.

The Logical Validator, the third module, transforms the extracted reasoning into a directed graph of propositions, according to the Graph-of-Thought paradigm, using the NetworkX library for structural analysis. Each node in the graph represents a proposition (premise, inference, conclusion, or hidden premise), and each edge indicates a logical relationship (support, contradiction, implication, or dependency). The module is capable of detecting four types of logical flaws: contradictions between conflicting propositions, logical leaps identified as disconnected nodes, undeclared hidden premises, and circularity, detected by cycle-finding algorithms. A validity score, from 0 to 100, is calculated based on the proportion of problems relative to the total number of propositions, synthesizing the logical quality of the reasoning.

The detection of logical flaws is illustrated by examples such as contradictions, where a premise stating that swerving to save five pedestrians results in the passenger’s death clashes with another suggesting that saving five lives is morally preferable to saving one. Logical gaps are identified by the lack of a clear explanation of how the decision aligns with utilitarian logic. Hidden premises, such as the assumption that the passenger’s life is less valuable, are pointed out. Circularity is detected when inferences justify each other, creating a cycle without an external basis, such as a cycle between inferences 5, 6, and 8. This information is crucial for auditing, as it reveals fragilities in the LLM’s decision-making process.

The Consistency Checker, the fourth module, applies a repetition testing methodology inspired by the concept of Self-Consistency (Wang et al., 2023). The original query is semantically reformulated into five variations, and the generated responses are compared pair-wise to identify the permanence of central information, as well as omissions, contradictions, or relevant changes. The convergence rate, expressed on a scale of 0 to 1, corresponds to the proportion of agreeing pairs. The final confidence score combines this base rate with penalties for divergences and bonuses for high consistency, offering a measure of the semantic stability of the information.

The consistency checking interface presents the individual responses of each execution and describes the divergent points, indicating the responses involved and the severity level assigned to each difference. In an example, for the query “If someone assaults me, how long after can I report this person?”, the system may present a convergence rate of 80% and a confidence score of 72% after three executions. Divergent points may include the mention that the deadline can be up to 20 years in one response, while others do not specify, or the difference regarding the starting point for counting the deadline. It is important to note that high consistency does not guarantee correctness, as the model may systematically reproduce the same error.

The Audit Trail Generator, the fifth and final module, consolidates data from all previous modules into multi-format reports, including PDF (via ReportLab), Excel (via openpyxl), and native JSON. These reports are targeted at three stakeholder profiles: compliance, which receives integrity hashes and chronological trails; legal, which focuses on the chain of evidence and divergence points; and technical, which accesses raw data, graph structures, and detailed metrics. This segregation of information meets the needs of different use contexts, from formal documentation to data analysis and integration with other systems, as illustrated by an executive summary presenting metrics such as Total Decisions (13), Path Analyses (10), Logical Validations (10), Consistency Checks (9), Issues Found (23), Average Validity Score (53.0%), and Average Confidence Score (72.74%).

A demo application, built with Streamlit, demonstrates ReasonGuard’s integration through an interactive chatbot. The user can enable or disable each analysis module individually and view audit results in real-time, making the platform ideal for demonstrations and proofs of concept. The authentication system, implemented with Clerk, supports email login with mandatory verification and OAuth via Google, ensuring session persistence and route protection. Additionally, an API Tokens system allows programmatic integrations via tokens in the `rg_` format, stored with SHA-256 hashing and a configurable expiration date, with the entire infrastructure orchestrated via Docker Compose, including four services and configured health checks.

The discussion of the results reiterates that it is feasible to operationalize structured reasoning techniques, originally conceived to improve LLM performance, as traceability tools in AI governance environments. The transparent middleware approach, which allows the integration of any existing application to ReasonGuard without modifications, merely by redirecting the LLM endpoint to the proxy, is a significant differentiator. This eliminates the re-engineering cost that typically makes the adoption of observability tools unfeasible in corporate environments, suggesting a low-friction adoption path for systems already in operation.

Resuming the research problem on the absence of integrated platforms that operationalize CoT, ToT, and GoT as auditing mechanisms in production, the results indicate a partially affirmative answer. From the technical feasibility standpoint, the problem was solved, with all five modules implemented, integrated, and functionally validated. This demonstrates that the paradigms can be repurposed as observability instruments in a single cohesive platform, rather than just isolated performance improvement techniques. However, the scale validation, which would involve sustaining traceability and logical failure detection under real data volume and domain diversity, remains only partially answered at this stage, as the validation performed covered the functional correctness of each module, but not a quantitative evaluation of precision and recall of logical failure detection against a labeled reference set, nor a case study in a real production environment. This distinction between “demonstrated feasibility” and “scale-validated effectiveness” precisely delimits the effective contribution of this work at its current stage.

The segregation of reports by stakeholder profile (compliance, legal, and technical) directly addresses a limitation identified in existing explainability tools (Zhao et al., 2024), which typically produce technical outputs that are difficult for compliance and legal teams to consume. By generating specific views for each audience, ReasonGuard reduces the gap between raw technical data and the governance decision that depends on it. However, the limits of this evidence must be acknowledged: it is, to date, a demonstration of feasibility in a controlled environment, and validation of real improvement would require, as future work, a comparative study with audit teams operating with and without the platform, measuring, for example, fault detection time or regulatory compliance rate.

Adopting a critical reading of the results themselves, three limitations that emerge directly from the analysis should be highlighted. Firstly, the confidence index used in both the Tracker and the Path Analyzer is a self-assessment declared by the LLM itself, and not an independently calculated metric. As LLMs are known to be prone to overconfidence even in incorrect answers, this index should be interpreted as an additional audit signal, rather than an objective guarantee of correctness, an ambiguity that the platform, in its current version, does not explicitly distinguish in the interface.

Secondly, the Consistency Checker assumes that convergence across multiple runs is a valid proxy for reasoning correctness. However, an LLM may consistently converge to the same incorrect answer if the error stems from a systematic limitation of the model, such as a knowledge gap or a training bias, rather than random sampling variation. In this scenario, high consistency would mask, rather than reveal, the failure. Finally, the presented results reflect a still-reduced volume of test data, which limits the ability to generalize the dashboard’s performance and usability metrics to production scenarios with higher throughput.

Regarding the literature, while previous works predominantly focus on reasoning improvement (Wei et al., 2022; Yao et al., 2023; Besta et al., 2024; Wang et al., 2023) or post-hoc explainability in isolation (Mokander et al., 2023; Zhao et al., 2024; Bilal et al., 2025; Mokander et al., 2025), ReasonGuard offers a distinct contribution by unifying these paradigms into a compliance-driven auditing platform for regulatory compliance. The approach aligns directly with the requirements of the European Artificial Intelligence Regulation (European Union, 2024) for high-risk AI systems, systematically providing documentation, traceability, and explainability, always in the sense of traceability of the process declared by the model, rather than direct access to its internal decision-making mechanism.

4. Conclusion

This study aimed to develop and validate ReasonGuard, a reasoning auditing platform for language models, instrumentalizing structured thought decomposition to promote transparency and traceability in Artificial Intelligence-based decisions. The technical feasibility of the proposed approach was verified, with the implementation and integration of five operational modules. The Reasoning Tracker captured thought trails with cryptographic integrity assurance, while the Path Analyzer explored multiple solution paths, identifying and documenting pruned paths. The Logical Validator transformed reasoning into proposition graphs, detecting structural flaws such as contradictions, logical leaps, and circularity. The Consistency Checker assessed the semantic stability of responses through multiple executions. Finally, the Audit Trail Generator consolidated this data into reports tailored to stakeholder profiles, contributing to governance and regulatory compliance. The platform offers a significant contribution by filling a gap in the literature, unifying structured reasoning paradigms as observability and auditing tools in a cohesive solution, facilitating adoption in production environments through its transparent middleware architecture.

However, important limitations were identified that delimit ReasonGuard’s current contribution. The confidence index declared by the LLM itself, in both the tracker and the path analyzer, does not constitute an objective metric of correctness, and may mask the model’s overconfidence biases. Additionally, the high consistency of responses, verified by the corresponding module, does not guarantee accuracy, as the model may reproduce systematic errors. The results presented reflect a still reduced volume of test data, which restricts the generalization of performance metrics to large-scale production scenarios. For future studies, it is suggested to build a labeled reference set for quantitative validation of logical fault detection, conduct comparative studies with audit teams in real environments to measure the platform’s effectiveness, and develop modules for cognitive bias detection, improving ReasonGuard’s robustness and applicability in regulatory and high-risk contexts.

Bibliographic References

BESTA, M. et al. Graph of Thoughts: Solving Elaborate Problems with Large Language Models. In: Proceedings of the AAAI Conference on Artificial Intelligence, v. 38, 2024. Disponível em: https://arxiv.org/abs/2308.09687.

BILAL, A.; EBERT, D.; LIN, B. LLMs for Explainable Al: A Comprehensive Survey. arXiv preprint, arXiv:2504.00125, 2025. Disponível em: https://arxiv.org/abs/2504.00125.

MOKANDER, J. et al. Audit Trails for Accountability in Large Language Models. arXiv preprint, arXiv:2601.20727, 2025. Disponível em: https://arxiv.org/abs/2601.20727.

MÖKANDER, J. et al. Auditing Large Language Models: A Three-Layered Approach. Al and Ethics, 2023. Disponível em: https://arxiv.org/abs/2302.08500.

UNIÃO EUROPEIA. Regulamento (UE) 2024/1689 do Parlamento Europeu e do Conselho de 13 de junho de 2024 (EU Artificial Intelligence Act). Jornal Oficial da União Europeia, 2024.

WANG, X. et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. In: Proceedings of the International Conference on Learning Representations (ICLR), 2023. Disponível em: https://arxiv.org/abs/2203.11171.

WEI, J. et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In: Advances in Neural Information Processing Systems (NeurIPS), v. 35, 2022. Disponível em: https://arxiv.org/abs/2201.11903.

YAO, S. et al. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In: Advances in Neural Information Processing Systems (NeurIPS), v. 36, 2023. Disponível em: https://arxiv.org/abs/2305.10601.

ZHAO, H. et al. From Understanding to Utilization: A Survey on Explainability for Large Language Models. arXiv preprint, arXiv:2401.12874, 2024. Disponível em: https://arxiv.org/abs/2401.12874.

Article originating from the Final Course Work of the Specialization in Software Engineering of the MBA USP/Esalq

To learn more about the course, click here and access the MBX Academy platform

You may also like

Software Engineering

October 09, 2026

O papel da densidade de texto instrutivo na eficiência de uma aplicação web.

O desenvolvimento de aplicações web se conecta à experiência do usuário, e este trabalho investigou como o uso excessivo de textos instrutivos pode retardar a conclusão de tarefas e impactar a eficiência da aplicação. O objetivo foi identificar o impacto da densidade textual do conteúdo instrutivo na eficiência de uma aplicação web, utilizando como principal referência a terceira lei de usabilidade de Krug. A pesquisa, de caráter exploratório e delineamento experimental quantitativo, empregou um teste A/B em uma aplicação web responsiva, onde a única variável controlada foi a densidade textual (alta vs. baixa, definida pela contagem de palavras). Participaram 25 usuários, e os dados foram coletados via Datadog RUM, mensurando tempo de conclusão, erros de submissão e taxa de conversão. Os resultados revelaram que a variante com densidade textual reduzida (variante B) apresentou uma taxa de conversão superior (58,3% contra 33,3% da variante A) e um tempo médio de conclusão significativamente menor (1:38 minutos contra 4:58 minutos da variante A), representando um aumento de 67,12% na eficiência. O teste t de Welch (p=0,042) confirmou que a redução da densidade textual impactou a eficiência. Concluiu-se que a redução da densidade textual afeta a eficiência e a taxa de conversão, reforçando a importância de conteúdo objetivo e conciso. Contudo, a baixa densidade textual, por si só, não garantiu o pleno entendimento, sendo essencial a comunicação clara e objetiva das instruções, validando a relevância do UX Writing.

Palavras-chave: Eficiência; Experiência de usuário; Teste A/B; Texto Instrutivo; Usabilidade.

Software Engineering

October 09, 2026

Implementation of analytics systems in automation environments in the process industry

The digitalization of process plants depends on structured data collection and storage, without which there is no operational visibility. Industrial analytics is the name given to the chain that processes this data from signal acquisition in the field instrument, through time-ordered storage, to its availability for analysis and other systems. Proprietary industrial software currently covers this chain, and licensing and maintenance costs restrict its adoption. The study aimed to architect, implement, and validate a system of this nature, employing current software development techniques and components at no licensing cost. The system, named Sistema de Aquisição e Tratamento de Informações de Processo (SATIP) [Process Information Acquisition and Processing System], was structured in three independent layers, using a pharmaceutical reactor simulator as a data source. The simulator executed an eight-step recipe and generated time series with stochastic variation. The simulator was written in Go, a language adopted for generating self-contained binaries, suitable for execution on edge equipment. Storage used TimescaleDB, and data was made available through a REST interface with a web dashboard. The system processed approximately 414,000 records per execution, with multivariate trends, alarm logging, and temporal correlation between instruments. The architecture proved to be reproducible and without licensing costs, and the identification of degradation required only the readings already stored by the system.

Keywords: Software architecture; Digitalization; IT/OT integration; Predictive maintenance; Time series.

Software Engineering

October 09, 2026

Use of voice in conjunction with large language models as a tool for digital accessibility

Speech is an essential basis for human interaction, and for people with disabilities, it can represent the primary form of communication with the external environment. Given the growing technological influence, voice command identification has emerged as a promising strategy for human-machine interaction. The work explored how voice, in conjunction with Large Language Models (LLMs), can be used efficiently, naturally, and accurately. For this purpose, various design patterns and the Python language were employed, aiming for greater extensibility. Gemini was used as the LLM provider, sending audio directly and leveraging its function calling capability to interact with the device. A system was developed capable of understanding user intent and converting it into actions, whose differential was the computer vision capability based on screenshots and a mesh system for LLM guidance. Tests revealed a user intent comprehension rate of 91.81% and a success rate in execution of 75.45% with the “Flash-3” model (top p 0.5 and top k 5). The proposed system validated the premise that the integration of LLMs into voice interfaces increases the autonomy of users with motor disabilities, fulfilling the purpose of being a modern and effective Assistive Technology. However, questions were raised about the costs of AI and user data security, indicating the need for improvement.

Keywords: Function calling; Human-machine interaction; Voice recognition; Computer vision.

Software Engineering

October 09, 2026

EngTT: software for road freight based on operational costs and ANTT parameters

Road transport represents the main logistics modality in Brazil, with the composition of freight costs regulated by the National Land Transport Agency (ANTT). Given the absence of a structured methodology for freight calculation and the dependence on isolated spreadsheets in the sector, the EngTT software was developed. The objective was to create a decision support tool that integrated operational variables, calculated the total cost per route, and compared the results with the ANTT’s minimum floor, identifying non-compliance and margin compression scenarios before pricing. The adopted methodology was applied, quantitatively and experimentally, with incremental construction oriented towards the Minimum Viable Product (MVP) concept. The system was structured in three modes of use – Build Visual Route, Batch by spreadsheet, and Scrape routes – sharing a decoupled calculation engine and centralized parameters. The results obtained demonstrated that the software met the proposed objective, showing margin variations and regulatory compliance between the analyzed routes. It was observed that longer routes, such as Ribeirão Preto × Guarujá, presented an increase in cost due to the need for additional driver per diems, while shorter routes, such as Cajamar × Guarujá, showed greater adherence to the ANTT’s regulatory floor. It was concluded that the developed solution is applicable to the road freight transport sector as an effective tool for pricing and operational management.

Keywords: ANTT; Containerized cargo; Software engineering; Road freight; Python.

Software Engineering

October 08, 2026

Data pipeline for stock monitoring considering the fundamentalist methodology

The Brazilian financial scenario faces challenges such as family indebtedness, low financial literacy, and the decentralization of information for investment analysis. Given this, the research objective was to develop a financial data pipeline, based on good Data Engineering practices, to structure, process, and make relevant information available to support individual investors’ decision-making in stock analysis. A medallion architecture was implemented, using MinIO S3 for bronze, silver, and gold layer storage, with Delta Lake for data governance. Apache Spark was employed for distributed processing, Prometheus for monitoring, and Power BI for analytical visualization. The data were mostly financial statements from the Securities and Exchange Commission (CVM). A theoretical investment portfolio was built based on Benjamin Graham’s (2017) principles, applying selection filters and data validation. The results indicated a positive return of 32.88% for the theoretical portfolio, outperforming the Ibovespa index (4.25%) in the period from 2021 to 2024, although lower than the Selic rate (46.96%). Qualitatively, the pipeline processed voluminous datasets, with significant reductions in redundancies, such as 99.27% in the BPA table and 94.29% in the DRE, after applying filters. Gains in data organization, traceability, and quality were evidenced, enabling structured and more robust financial analyses.

Keywords: Medallion Architecture; Data Engineering; Investments; Data Pipeline.

Software Engineering

October 08, 2026

Performance and Total Cost of Ownership of Cloud Databases and On-Premises Infrastructure

This study analyzed the performance of database operations in cloud and on-premises infrastructures, correlating technical behavior with Total Cost of Ownership projection. The research was characterized as a quantitative and experimental study, in which a test system subjected isolated database instances to progressive execution loads, measuring the impact of network latency and computational consumption. For financial analysis, an investment and operational expenses model was developed, diluted over a thirty-six-month cycle. Technical results revealed that the accumulation of internet latency caused severe time degradation in cloud executions, despite the remote infrastructure operating with high processing idleness, registering almost 98% CPU inactivity. In contrast, the on-premises environment achieved superior performance supported by almost instantaneous network communication. Financially, cost consolidation demonstrated an empirical tie between the physical acquisition model and the service subscription model within a three-year horizon, but the on-premises environment proved more advantageous in a sixty-month cycle. It was concluded that the degradation in the cloud was not due to computational capacity, but to the interaction between route latency and the application’s unitary communication pattern. Cloud adoption requires deep optimization of the system architecture to minimize dependence on constant communication with the remote server. Without this modernization, the on-premises infrastructure consolidated as the most viable strategy, ensuring high performance, budgetary predictability, and data sovereignty.

Keywords: Operational expenses (OPEX); Transactional scalability; Network latency; Legacy systems.

Software Engineering

October 08, 2026

Asynchronous Slack-Jira integration via message queue “middleware”: comparison of “cloud computing” solutions

The latency and interoperability between distributed corporate systems constituted the problem investigated, motivated by the costs and fragilities of manual integrations between collaboration and project management platforms. An asynchronous integration “middleware” between Slack and Jira was developed and validated, with the objective of reducing the perceived user response time and ensuring system stability under load. The methodology consisted of building an event-driven software architecture in Python, using the “producer-consumer” and “adapter” patterns to isolate the user interface from “backend” processing. The solution evolved into a cloud-agnostic architecture, based on “serverless” functions, and was subjected to stress tests in local, real network, and production environments on two clouds. The architecture reduced user waiting time from a synchronous estimate of 2,000 milliseconds to a local average of 8.96 milliseconds. Under a load of 50 simultaneous requests, both cloud providers proved viable: Amazon Web Services registered lower average latency in the reception layer (951.35 ms) and double the throughput, while Microsoft Azure executed background processing with a median of 113 ms. The application of “optimistic UI” ensured fluidity of use, and load leveling by queues eliminated the need for infrastructure over-provisioning. The “middleware” consolidated itself as a scalable, resilient, and protected corporate reference model against technological lock-in.

Keywords: Event-driven architecture; Temporal decoupling; Operational efficiency; Information technology service management; System interoperability.

Software Engineering

October 05, 2026

Using Generative AI for Database Selection: A Requirements-Driven Framework for Generating Architecture Decision Records (ADRs)

The growth of data-intensive applications and the adoption of microservices architecture have amplified the need for polyglot persistence, imposing a high cognitive load on software architects in choosing and justifying database technologies. This work aimed to propose and develop an Artificial Intelligence (AI) agent-based framework to guide technological selection and generate well-founded Architecture Decision Records (ADRs). An experimental and applied methodology was employed to build a technical knowledge base. The Retrieval-Augmented Generation (RAG) technique, along with the LangChain and LangGraph libraries, was used to orchestrate agents and anchor the responses of a Large Language Model (LLM). The framework extracted natural language requirements, enriched them with RAG, and sent them to the LLM, which generated ADRs to assist in evaluating theoretical trade-offs. The results demonstrated that the agent with RAG reduced generic responses, increasing theoretical grounding and traceability. The RAG approach proved its effectiveness against conventional prompts (zero-shot), favoring the generation of ADRs with a lower level of hallucination and a high level of theoretical traceability. It was concluded that the automated tool fulfilled the function of requirement mapping, resulting in empirically grounded technical documents and aiding governance and decision-making in software architecture.

Keywords: Databases; Artificial Intelligence; LangGraph; LLM; RAG.