Software Engineering
October 09, 2026
ReasonGuard: A Reasoning Audit Platform for Language Models Based on Structured Thought Decomposition
Reasonguard: A Reasoning Auditing Platform for Language Models Based on Structured Thought Decomposition
Lucas Nastari Ziza; Lucas Cesar Gomes Alvarinho Squillante
DOI: 10.22167/2675-6528-202603200
Article derived from a Course Conclusion Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.
Summary
ReasonGuard was presented, an Artificial Intelligence reasoning auditing platform, developed as an observability middleware between client applications and large language models (LLMs). The work aimed to offer transparency and traceability for AI-based decision-making processes, addressing the gap in operationalizing structured reasoning techniques for auditing purposes. The system intercepted, analyzed, and documented interactions with LLMs through five modules based on the Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT) paradigms. These modules captured reasoning trails, detected structural logical flaws, evaluated response consistency, and generated audit reports targeted at different stakeholder profiles. The platform was implemented with FastAPI (Python) on the backend, React with TypeScript on the frontend, and PostgreSQL as a relational database, following a modularized architecture with multitenancy isolation per user. The results demonstrated the technical feasibility of the proposed approach, with all five modules operational and integrated. It was concluded that ReasonGuard contributes to AI governance by instrumentalizing structured reasoning paradigms as observability and auditing tools, filling a gap in the literature.
Keywords: AI Auditing; Chain-of-Thought; AI Governance; Large Language Models; Algorithmic Transparency.
1. Introduction
Large Language Models (LLMs) are deep learning systems, typically based on the Transformer architecture, trained on voluminous data structures to predict the most probable continuation of a text sequence. From this mechanism of statistical token prediction, complex capabilities of language generation, question answering, summarization, and, more recently, multi-step reasoning emerge. It is precisely this emergent reasoning capability and its application to high-impact decisions that motivates the present work.
The growing adoption of LLMs in decision-making contexts, including areas such as health, law, finance, and corporate governance, has highlighted a significant gap: the visualization of the reasoning processes that underpin the responses generated by these systems. An illustrative example of this problem is the “plausible but fallacious reasoning” error. When asked about the viability of a contract with a clause for automatic termination in case of payment delay, an LLM may correctly conclude that the contract is terminable, but base this conclusion on an unstated and incorrect premise, such as erroneously assuming that any delay triggers the clause. The final answer sounds convincing and is often correct by coincidence, but the underlying reasoning contains a structural flaw that, in a real-world scenario, could lead to a mistaken decision without explicit signs of error for the user. This type of failure, undetectable from a simple reading of the final answer, is the central problem that ReasonGuard seeks to address.
Although LLMs demonstrate expressive capabilities in complex reasoning tasks, the absence of formal auditing and traceability mechanisms compromises their applicability in scenarios requiring accountability and regulatory compliance. From a regulatory standpoint, the global landscape has been moving towards stricter transparency requirements. The European Artificial Intelligence Regulation (European Union, 2024), for example, establishes specific transparency obligations for high-risk AI systems, including documentation, traceability, and explainability requirements. In the academic context, recent research on accountability infrastructure for AI highlights that accountability remains fragile because process transparency is rarely recorded in a durable and auditable manner (Mokander et al., 2025). It is important to demarcate that, as they are neural networks with billions of parameters whose internal processing is not directly interpretable, LLMs remain, in essence, “black boxes”. What ReasonGuard and structured reasoning decomposition techniques offer is not the opening of this black box, but rather the induction of the model to express, in natural language, a step-by-step justification that can be recorded, versioned, and audited. This, therefore, concerns the traceability of the response generation process, a verifiable record of what the model declared as its reasoning, and not full transparency or reproducibility of the model’s internal decision mechanism, a distinction that guides the entire platform design.
The concept of structured reasoning in language models gained prominence with the work of Wei et al. (2022), who demonstrated the Chain-of-Thought (CoT) technique. This technique, by explicating intermediate reasoning steps before the final answer, significantly improves performance on arithmetic, common sense, and symbolic reasoning tasks. Building on this foundation, Yao et al. (2023) proposed the Tree of Thoughts (ToT) framework, which generalizes CoT by allowing the deliberate exploration of multiple reasoning paths, treating each intermediate step as a node in a search tree. Besta et al. (2024) further advanced this by introducing the Graph of Thoughts (GoT), which models the information generated by the LLM as an arbitrary graph of propositions and their dependencies. In parallel, Wang et al. (2023) proposed the Self-Consistency approach, a decoding strategy that explores multiple reasoning paths and selects the most consistent answer among them, elevating CoT performance on benchmarks.
Despite these advances, a specific gap is observed in the literature: the works cited above focus on two distinct and, to date, unintegrated fronts. The first front treats these techniques (CoT, ToT, GoT, and Self-Consistency) as methods to improve LLM performance on reasoning tasks, without concern for the persistent, verifiable, and auditable record of the generated process. The second front, the works on auditing and explainability (Mokander et al., 2023; Zhao et al., 2024; Bilal et al., 2025; Mokander et al., 2025), discuss what should be audited and propose conceptual verification layers, but do not present a reference implementation that systematically captures, structures, and persists, in a production-ready manner, the decomposed reasoning according to these paradigms. Therefore, there is no platform that operates at the intersection of these two fronts: one that uses CoT, ToT, and GoT not to improve the model’s response, but as observability instruments, capturing the decomposed reasoning, persisting it with integrity guarantees, and making it available in formats usable by compliance, legal, and engineering teams. It is this operationalization gap, the absence of a bridge between the reasoning decomposition technique and the auditing practice in production environments, that the present work seeks to fill.
In this context, the present work proposes ReasonGuard, a reasoning auditing platform that instrumentalizes reasoning paradigms as observability mechanisms, operating between client applications and LLM providers. The general objective of this study is to develop and validate ReasonGuard, a reasoning auditing platform for language models based on structured thought decomposition, with the following specific objectives: (i) capture complete reasoning trails with cryptographic integrity guarantee; (ii) detect structural logical flaws such as contradictions, circularity, logical leaps, and hidden premises; (iii) evaluate response consistency through multiple executions with semantic variations; and (iv) generate audit reports targeted at different stakeholder profiles, contributing to the governance and regulatory compliance of AI-based systems.
2. Material and Methods
The research conducted has an applied nature, oriented towards the generation of practical knowledge to solve the lack of transparency and auditability in the reasoning processes of Large Language Models (LLMs). The study was characterized as exploratory-experimental, combining the investigation of reasoning decomposition techniques with the implementation and validation of modules in a controlled environment. The adopted methodological approach was quali-quantitative, encompassing both the structural analysis of the extracted reasoning and the calculation of numerical metrics, such as validity scores and consistency rates.
The development of ReasonGuard followed an iterative and incremental methodology, based on evolutionary prototyping, structured in six sequential phases. Initially, the system’s requirements and architecture were defined, identifying the five main modules and adopting the API Proxy architectural pattern. Next, the backend was implemented with the core modules, followed by the implementation of the frontend with an interactive dashboard. Subsequent phases included integration and authentication with multitenancy isolation, the construction of a demo application in Streamlit, and finally, the optimization and refinement of performance and usability.
The ReasonGuard system architecture was designed as a reverse proxy with observability features, acting as middleware between client applications and LLM providers. This arrangement allowed for transparent interception of requests without the need for client application modification, enabling real-time analysis and independent storage of audit data. For the client, communication remained unchanged, requiring only the redirection of requests to the proxy’s endpoint, which preserved the expected input and output structure.
During the communication flow, the API Proxy recorded the relevant request information, applied the tracking, validation, and analysis mechanisms defined by the platform, and forwarded the call to the model provider. After processing by the provider, the response was returned to the proxy, which analyzed it, associated it with the original request data, calculated metrics, and generated the audit trail before passing it back to the client. In this way, the LLM provider remained responsible for model inference, while the middleware centralized observability, control, and auditing functions.
The ReasonGuard infrastructure was composed of four services orchestrated via Docker Compose. It included a PostgreSQL database for relational persistence, a backend implemented with FastAPI in Python containing the five core modules and nine REST API routes, a frontend developed with React and TypeScript for the interactive dashboard, and a demo application in Streamlit as an integration example. Multitenancy isolation was guaranteed by user filtering in all database queries, ensuring that each user accessed only their own records.
The core modules were developed as autonomous units, each with a unique responsibility, and progressively integrated into the complete system. The Reasoning Tracker, the first module, was based on the Chain-of-Thought (CoT) paradigm. It operated by injecting structured prompts into the LLM requests, inducing the model to decompose its reasoning into premises, inferences, and conclusions. Parsing was performed via regular expressions, and each record was assigned an SHA-256 hash to ensure data integrity.
The Path Analyzer, the second module, extended the concept to multiple paths, implementing a tree search with heuristic evaluation inspired by the Tree-of-Thought (ToT) paradigm. The problem was decomposed into two to four subproblems, for which two to three hypotheses with feasibility scores, ranging from zero to one hundred, were generated. Paths with scores below forty were pruned and documented. Captured metrics included the total nodes explored, the total paths pruned, and the score of the selected path.
The Logical Validator, the third module, transformed the extracted reasoning into a directed graph of propositions, using the NetworkX library, according to the Graph-of-Thought (GoT) paradigm. Each node represented a proposition (premise, inference, conclusion, or hidden premise) and each edge, a logical relation. The module detected four types of logical flaws: contradictions, logical leaps identified as disconnected nodes, undeclared hidden premises, and circularity, the latter detected by cycle-finding algorithms. A validity score, from zero to one hundred, was calculated based on the proportion of problems in relation to the total number of propositions.
The Consistency Checker, the fourth module, applied a repetition testing methodology inspired by the concept of Self-Consistency (Wang et al., 2023). The original query was semantically reformulated into five variations, preserving the original intent. The generated responses were compared pair-wise to identify the permanence of central information and detect divergences. A convergence rate, expressed on a scale of zero to one, was calculated, along with a final confidence score that combined this base rate with penalties for divergences and bonuses for high consistency.
The Audit Trail Generator, the fifth module, consolidated data from all previous modules into multi-format reports. These reports were generated in PDF via ReportLab, Excel via openpyxl, and native JSON, and directed to three stakeholder profiles: compliance (with integrity hashes and chronological trails), legal (with chain of evidence and divergence points), and technical (with complete data, raw structures, and detailed metrics).
The choice of the technological stack was guided by criteria of suitability for the AI domain, ecosystem maturity, and development productivity. The backend was implemented in Python 3.11 with FastAPI 0.109.0, leveraging native support for asynchronous operations, automatic validation via Pydantic, and OpenAPI documentation. The NetworkX 3.2.1 library provided the necessary graph analysis algorithms for the Logical Validator.
The frontend was built with React 18.2.0 and TypeScript 5.3.3, utilizing React Flow 11.10.2 for interactive graph and tree visualization, Recharts 2.10.4 for KPI charts, and Material-UI 5.15.6 as the interface framework. Server state management employed TanStack Query for caching and synchronization. Authentication was managed by Clerk 4.30.0, with support for OAuth (Google) and email/password, while PostgreSQL 14 provided relational persistence with JSONB field support for storing flexible structures.
The validation strategy followed a three-level approach. In functional validation, each module was individually validated for parsing correctness, pattern detection, and metric calculation. Parsing corresponded to the process of interpreting and structuring the received data, correctly extracting fields, reasoning steps, and other relevant information from the model’s responses for subsequent analyses.
Integration validation encompassed the complete functioning flow of the proxy, from receiving the request sent by the client application to forwarding it to the LLM provider. It included processing by the audit modules, persistence of results, and returning the response to the user. End-to-end authentication mechanisms were also verified, including user identification, resource access authorization, and data isolation throughout the execution.
The interface validation focused on the correct presentation of information produced by the system. The rendering of graphs and reasoning trees with real data, the visual organization of nodes and connections, the display of analysis results, and the updating of performance indicators on the dashboard were evaluated. Efforts were made to ensure that changes in processed data were correctly reflected in the KPIs and other visual components of the platform.
3. Results and Discussion
The results obtained in the development of ReasonGuard demonstrate the technical viability of the proposed approach to operationalize structured reasoning techniques as traceability and auditing tools in Artificial Intelligence governance environments. The platform was implemented completely with respect to the specified requirements, with all five core modules fully functional and integrated. It is essential to situate these findings as evidence of a prototype’s functionality, and not as a statistical validation on a production scale, a crucial distinction for the interpretation of its current contribution.
The system architecture was designed to ensure data integrity and traceability. The data model centers on the user entity, linking to it all records generated by the reasoning tracking, path analysis, logical validation, and consistency verification modules. Reasoning chains are stored as ordered sequences of steps, while path analyses are represented by hierarchical structures with parent-child node relationships. Logical validations, in turn, are modeled as graphs of propositions and directed connections, with all these associations implemented by foreign keys to ensure referential integrity in relation to the user.
On the backend, each module operates with its own routes and services, using SQLAlchemy to process requests and persist results. This modular organization allows analyses to be executed independently, but later integrated into the generation of audit reports. The frontend, in turn, was structured into ten pages dedicated to specific functionalities, including an interactive dashboard, an integrated chat, and visualizations for reasoning tracking, path analysis, logical validation, consistency checking, audit reports, and system settings.
The Reasoning Tracker, the first module, implements the interception of requests and the injection of Chain-of-Thought instructions into a structured object. For each processed request, the module records the declared premises, the inferences made, the final conclusion, and a confidence level, ranging from 0 to 100, for each reasoning step. The parsing of this information is performed using regular expressions, and each record receives a SHA-256 hash, which acts as a “fingerprint” to ensure integrity and detect any undue alteration in the content. This feature is essential for compliance and regulatory audit scenarios, where the immutability of the reasoning evidence is irrefutable.
The reasoning tracker’s visualization interface displays in detail the original prompt, integrity hash, and reasoning steps. For example, in a query about an autonomous car’s decision to swerve to save five pedestrians at the expense of the passenger, the system breaks down the reasoning into premises such as “An autonomous car has the ability to make decisions in emergency situations” and “Saving five lives is considered morally preferable to saving one life”, each with 100% confidence. Subsequent inferences, such as evaluating consequences and choosing the option that minimizes loss of life, are also recorded with their respective confidence levels, culminating in the final conclusion based on utilitarian logic.
The Path Analyzer, the second module, extends the concept of reasoning to multiple paths, implementing a tree search with heuristic evaluation inspired by the Tree-of-Thought paradigm. The problem is decomposed into two to four subproblems, for which two to three alternative hypotheses are generated, each with a feasibility score between 0 and 100. Paths with scores below 40 are pruned and documented, allowing the identification of lines of reasoning that, although not presented to the end-user, may be relevant or reveal potential flaws. The confidence index assigned to each trajectory is a self-assessment declared by the LLM itself, not an objective external metric of correctness.
The interactive decision tree visualization, implemented with React Flow, allows the user to navigate through explored nodes, identify pruned paths, and the selected path. In a scenario of employee layoffs, for example, the system can present the hypothesis of “Openly communicate the reasons for layoffs” with 85% self-assessment, and “Offer voluntary resignation packages” with 75%. A hypothesis such as “Keep employees in the dark until dismissal” can be pruned with 30% self-assessment, indicating low viability. The metrics captured by this module include the total number of nodes explored, the total number of pruned paths, and the score of the selected path, providing a comprehensive view of the solution space explored by the LLM.
The Logical Validator, the third module, transforms the extracted reasoning into a directed graph of propositions, according to the Graph-of-Thought paradigm, using the NetworkX library for structural analysis. Each node in the graph represents a proposition (premise, inference, conclusion, or hidden premise), and each edge indicates a logical relationship (support, contradiction, implication, or dependency). The module is capable of detecting four types of logical flaws: contradictions between conflicting propositions, logical leaps identified as disconnected nodes, undeclared hidden premises, and circularity, detected by cycle-finding algorithms. A validity score, from 0 to 100, is calculated based on the proportion of problems relative to the total number of propositions, synthesizing the logical quality of the reasoning.
The detection of logical flaws is illustrated by examples such as contradictions, where a premise stating that swerving to save five pedestrians results in the passenger’s death clashes with another suggesting that saving five lives is morally preferable to saving one. Logical gaps are identified by the lack of a clear explanation of how the decision aligns with utilitarian logic. Hidden premises, such as the assumption that the passenger’s life is less valuable, are pointed out. Circularity is detected when inferences justify each other, creating a cycle without an external basis, such as a cycle between inferences 5, 6, and 8. This information is crucial for auditing, as it reveals fragilities in the LLM’s decision-making process.
The Consistency Checker, the fourth module, applies a repetition testing methodology inspired by the concept of Self-Consistency (Wang et al., 2023). The original query is semantically reformulated into five variations, and the generated responses are compared pair-wise to identify the permanence of central information, as well as omissions, contradictions, or relevant changes. The convergence rate, expressed on a scale of 0 to 1, corresponds to the proportion of agreeing pairs. The final confidence score combines this base rate with penalties for divergences and bonuses for high consistency, offering a measure of the semantic stability of the information.
The consistency checking interface presents the individual responses of each execution and describes the divergent points, indicating the responses involved and the severity level assigned to each difference. In an example, for the query “If someone assaults me, how long after can I report this person?”, the system may present a convergence rate of 80% and a confidence score of 72% after three executions. Divergent points may include the mention that the deadline can be up to 20 years in one response, while others do not specify, or the difference regarding the starting point for counting the deadline. It is important to note that high consistency does not guarantee correctness, as the model may systematically reproduce the same error.
The Audit Trail Generator, the fifth and final module, consolidates data from all previous modules into multi-format reports, including PDF (via ReportLab), Excel (via openpyxl), and native JSON. These reports are targeted at three stakeholder profiles: compliance, which receives integrity hashes and chronological trails; legal, which focuses on the chain of evidence and divergence points; and technical, which accesses raw data, graph structures, and detailed metrics. This segregation of information meets the needs of different use contexts, from formal documentation to data analysis and integration with other systems, as illustrated by an executive summary presenting metrics such as Total Decisions (13), Path Analyses (10), Logical Validations (10), Consistency Checks (9), Issues Found (23), Average Validity Score (53.0%), and Average Confidence Score (72.74%).
A demo application, built with Streamlit, demonstrates ReasonGuard’s integration through an interactive chatbot. The user can enable or disable each analysis module individually and view audit results in real-time, making the platform ideal for demonstrations and proofs of concept. The authentication system, implemented with Clerk, supports email login with mandatory verification and OAuth via Google, ensuring session persistence and route protection. Additionally, an API Tokens system allows programmatic integrations via tokens in the `rg_` format, stored with SHA-256 hashing and a configurable expiration date, with the entire infrastructure orchestrated via Docker Compose, including four services and configured health checks.
The discussion of the results reiterates that it is feasible to operationalize structured reasoning techniques, originally conceived to improve LLM performance, as traceability tools in AI governance environments. The transparent middleware approach, which allows the integration of any existing application to ReasonGuard without modifications, merely by redirecting the LLM endpoint to the proxy, is a significant differentiator. This eliminates the re-engineering cost that typically makes the adoption of observability tools unfeasible in corporate environments, suggesting a low-friction adoption path for systems already in operation.
Resuming the research problem on the absence of integrated platforms that operationalize CoT, ToT, and GoT as auditing mechanisms in production, the results indicate a partially affirmative answer. From the technical feasibility standpoint, the problem was solved, with all five modules implemented, integrated, and functionally validated. This demonstrates that the paradigms can be repurposed as observability instruments in a single cohesive platform, rather than just isolated performance improvement techniques. However, the scale validation, which would involve sustaining traceability and logical failure detection under real data volume and domain diversity, remains only partially answered at this stage, as the validation performed covered the functional correctness of each module, but not a quantitative evaluation of precision and recall of logical failure detection against a labeled reference set, nor a case study in a real production environment. This distinction between “demonstrated feasibility” and “scale-validated effectiveness” precisely delimits the effective contribution of this work at its current stage.
The segregation of reports by stakeholder profile (compliance, legal, and technical) directly addresses a limitation identified in existing explainability tools (Zhao et al., 2024), which typically produce technical outputs that are difficult for compliance and legal teams to consume. By generating specific views for each audience, ReasonGuard reduces the gap between raw technical data and the governance decision that depends on it. However, the limits of this evidence must be acknowledged: it is, to date, a demonstration of feasibility in a controlled environment, and validation of real improvement would require, as future work, a comparative study with audit teams operating with and without the platform, measuring, for example, fault detection time or regulatory compliance rate.
Adopting a critical reading of the results themselves, three limitations that emerge directly from the analysis should be highlighted. Firstly, the confidence index used in both the Tracker and the Path Analyzer is a self-assessment declared by the LLM itself, and not an independently calculated metric. As LLMs are known to be prone to overconfidence even in incorrect answers, this index should be interpreted as an additional audit signal, rather than an objective guarantee of correctness, an ambiguity that the platform, in its current version, does not explicitly distinguish in the interface.
Secondly, the Consistency Checker assumes that convergence across multiple runs is a valid proxy for reasoning correctness. However, an LLM may consistently converge to the same incorrect answer if the error stems from a systematic limitation of the model, such as a knowledge gap or a training bias, rather than random sampling variation. In this scenario, high consistency would mask, rather than reveal, the failure. Finally, the presented results reflect a still-reduced volume of test data, which limits the ability to generalize the dashboard’s performance and usability metrics to production scenarios with higher throughput.
Regarding the literature, while previous works predominantly focus on reasoning improvement (Wei et al., 2022; Yao et al., 2023; Besta et al., 2024; Wang et al., 2023) or post-hoc explainability in isolation (Mokander et al., 2023; Zhao et al., 2024; Bilal et al., 2025; Mokander et al., 2025), ReasonGuard offers a distinct contribution by unifying these paradigms into a compliance-driven auditing platform for regulatory compliance. The approach aligns directly with the requirements of the European Artificial Intelligence Regulation (European Union, 2024) for high-risk AI systems, systematically providing documentation, traceability, and explainability, always in the sense of traceability of the process declared by the model, rather than direct access to its internal decision-making mechanism.
4. Conclusion
This study aimed to develop and validate ReasonGuard, a reasoning auditing platform for language models, instrumentalizing structured thought decomposition to promote transparency and traceability in Artificial Intelligence-based decisions. The technical feasibility of the proposed approach was verified, with the implementation and integration of five operational modules. The Reasoning Tracker captured thought trails with cryptographic integrity assurance, while the Path Analyzer explored multiple solution paths, identifying and documenting pruned paths. The Logical Validator transformed reasoning into proposition graphs, detecting structural flaws such as contradictions, logical leaps, and circularity. The Consistency Checker assessed the semantic stability of responses through multiple executions. Finally, the Audit Trail Generator consolidated this data into reports tailored to stakeholder profiles, contributing to governance and regulatory compliance. The platform offers a significant contribution by filling a gap in the literature, unifying structured reasoning paradigms as observability and auditing tools in a cohesive solution, facilitating adoption in production environments through its transparent middleware architecture.
However, important limitations were identified that delimit ReasonGuard’s current contribution. The confidence index declared by the LLM itself, in both the tracker and the path analyzer, does not constitute an objective metric of correctness, and may mask the model’s overconfidence biases. Additionally, the high consistency of responses, verified by the corresponding module, does not guarantee accuracy, as the model may reproduce systematic errors. The results presented reflect a still reduced volume of test data, which restricts the generalization of performance metrics to large-scale production scenarios. For future studies, it is suggested to build a labeled reference set for quantitative validation of logical fault detection, conduct comparative studies with audit teams in real environments to measure the platform’s effectiveness, and develop modules for cognitive bias detection, improving ReasonGuard’s robustness and applicability in regulatory and high-risk contexts.
Bibliographic References
BESTA, M. et al. Graph of Thoughts: Solving Elaborate Problems with Large Language Models. In: Proceedings of the AAAI Conference on Artificial Intelligence, v. 38, 2024. Disponível em: https://arxiv.org/abs/2308.09687.
BILAL, A.; EBERT, D.; LIN, B. LLMs for Explainable Al: A Comprehensive Survey. arXiv preprint, arXiv:2504.00125, 2025. Disponível em: https://arxiv.org/abs/2504.00125.
MOKANDER, J. et al. Audit Trails for Accountability in Large Language Models. arXiv preprint, arXiv:2601.20727, 2025. Disponível em: https://arxiv.org/abs/2601.20727.
MÖKANDER, J. et al. Auditing Large Language Models: A Three-Layered Approach. Al and Ethics, 2023. Disponível em: https://arxiv.org/abs/2302.08500.
UNIÃO EUROPEIA. Regulamento (UE) 2024/1689 do Parlamento Europeu e do Conselho de 13 de junho de 2024 (EU Artificial Intelligence Act). Jornal Oficial da União Europeia, 2024.
WANG, X. et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. In: Proceedings of the International Conference on Learning Representations (ICLR), 2023. Disponível em: https://arxiv.org/abs/2203.11171.
WEI, J. et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In: Advances in Neural Information Processing Systems (NeurIPS), v. 35, 2022. Disponível em: https://arxiv.org/abs/2201.11903.
YAO, S. et al. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In: Advances in Neural Information Processing Systems (NeurIPS), v. 36, 2023. Disponível em: https://arxiv.org/abs/2305.10601.
ZHAO, H. et al. From Understanding to Utilization: A Survey on Explainability for Large Language Models. arXiv preprint, arXiv:2401.12874, 2024. Disponível em: https://arxiv.org/abs/2401.12874.
Article originating from the Final Course Work of the Specialization in Software Engineering of the MBA USP/Esalq
To learn more about the course, click here and access the MBX Academy platform