October 07, 2026
Predictive modeling for prioritizing fiscal inconsistencies from the ECD trial balance
Predictive Modeling for Prioritization of Fiscal Inconsistencies from the ECD Trial Balance
João Pedro Apolinário Cardoso; Thiago Gentil Ramires
DOI: 10.22167/2675-6528-202603036
Article derived from a Final Course Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.
Summary
The growing complexity of economic relations and the high volume of data available to the tax administration demanded the development of more efficient tools for selecting taxpayers subject to tax audits. In this context, supervised machine learning models were applied, based on data extracted from Digital Accounting Records (ECD), to prioritize companies in a scenario of limited human resources and incomplete labeling. To this end, a dataset composed of accounting variables derived from the ECD was constructed, and the Logistic Regression and Random Forest models were applied. Logistic Regression was used as a linear reference, while Random Forest was chosen for its ability to capture non-linear relationships and complex interactions between structured variables. More complex models, such as neural networks, were not prioritized, considering the low availability of positive class examples and the objective of adopting a parsimonious and easily operationally applicable approach. The performance of the models was evaluated with a focus on their ability to concentrate positive cases in the top positions of the ranking. The results indicated the superiority of tree-based models, especially Random Forest, in prioritizing taxpayers. The identification of relevant patterns was also observed even in the context of imperfect labels. It was concluded that the proposed approach has the potential to support the selection of taxpayers for audit and contribute to the more efficient allocation of public resources.
Keywords: Risk analysis; Supervised learning; Accounting data; Audit prioritization; Taxpayer selection.
1. Introduction
The digitalization process in the public sector, which intensified from the 2000s onwards, has promoted significant transformations, especially in tax administration. In this context, the Public Digital Bookkeeping System (SPED) was implemented, replacing paper declarations with digital documents. One of the main tools resulting from this movement is the Digital Accounting Bookkeeping (ECD), according to the Manual of Orientation of Layout 9 of the Digital Accounting Bookkeeping (Brazil [RFB], 2024), which has become essential for tax audits.
The ECD requires taxpayers to declare various accounting books, including the trial balance, the Income Statement (DRE), and the balance sheet. Initially, access to this data was restricted to the Brazilian Federal Revenue Service (RFB), but Normative Instruction RFB No. 2.003/2021 expanded access to state tax authorities. This expansion, along with the standardization of the Referential Chart of Accounts (record 1051), has enabled the massive use of ECD data for cross-referencing and comparisons with other fiscal obligations, such as the Digital Fiscal Record (EFD) and electronic invoices (Silva, 2025).
Among the components of the ECD, the trial balance stands out as the richest data source for state tax authorities. It consists of monthly consolidated balances by account, detailing opening and closing balances, as well as debit and credit amounts (Brazil [RFB], 2024). With these records, tax authorities can perform cross-checks to identify anomalies or fiscal inconsistencies, which serve as indicators for selecting companies to be audited.
Despite technological advancements, tax administration still faces significant limitations in human and operational resources. The selection of taxpayers for audit requires a prior risk analysis to prioritize the most relevant cases, given that the absence of assessment does not necessarily guarantee the fiscal regularity of a taxpayer or fiscal year. This reality poses a challenge in the efficient identification of irregularities.
Currently, tax administrations use rigid rule-based fiscal meshes to select taxpayers. However, the volume and complexity of available fiscal data exceed human analytical capacity, as pointed out by the National Federation of Municipal Tax Auditors and Inspectors (FENAFIM, 2017). International literature also documents the low adaptive capacity and greater propensity to biases of fixed rule-based models compared to machine learning approaches (Chan et al., 2022; Battaglini et al., 2025).
This study uses real data from the Digital Accounting Records (ECD) of companies contributing to the State of Goiás, extracted from periodic balance records (1150 and 1155), in addition to information from fiscal audits conducted in previous years. The research focus is on the use of accounting data from the ECD and derived variables, and may incorporate, depending on the analysis, complementary data from the Digital Fiscal Records (EFD) and the Electronic Invoice (NF-e). Registration data, such as the National Classification of Economic Activities (CNAE), are also considered, in order to identify accounting patterns associated with different economic segments.
Given this scenario, it becomes necessary to adopt, in tax administrations, data mining and machine learning mechanisms as a way to support decision-making. The adoption of these techniques, applied to the ECD, will allow the construction of risk scores, assisting in the prioritization of taxpayers for audits, reducing the need for manual analyses, and thus contributing to the rational use of public resources, as evidenced by empirical studies in recent literature (Battaglini et al., 2025; Yang, 2025). In this context, the work applies and evaluates supervised machine learning models with the objective of generating risk scores that assist in prioritizing companies and periods to be selected for tax audit, considering the history of inconsistencies identified in previous years. The research is part of the innovation efforts in data science in the tax area and contributes to the improvement of taxpayer selection processes, as well as to the development of analytical solutions with potential for application in other public administrations.
2. Material and Methods
This study, of an applied nature, employed a quantitative approach to develop and evaluate supervised machine learning models. The central objective was to generate risk scores that would assist in prioritizing companies and periods for tax audit, considering the history of inconsistencies identified in previous years. The research focused on predictive modeling to optimize the selection of taxpayers in a scenario of limited resources.
The unit of analysis consisted of taxpayer companies in the State of Goiás, with data organized by fiscal year. For the construction of the database, real information from the Digital Accounting Record (ECD) was used, specifically the periodic balance records (1150 and 1155) of the trial balance, in addition to data from the Income Statement (DRE), both components of the ECD.
Complementarily, the database incorporated information obtained from tax audits already completed in previous fiscal years, which served to label the response variable. Cadastral data were also included, such as the National Classification of Economic Activities (CNAE), in order to identify patterns of accounting behavior associated with different economic segments.
The response variable (Y) was defined as binary. The value 1 (one) was assigned when the company was fined in a given year or when it showed strong indications of infringement. The value 0 (zero) was assigned to the other cases, which included both companies analyzed without irregularities and those that were not subject to audit, reflecting the scarcity of resources for analyzing all taxpayers.
This definition of the response variable characterizes the problem as a scenario of rare positive class and heterogeneous negative labels. Such a structure is common in tax auditing applications, where the focus is on risk-based prioritization for subsequent specialized human screening and selection (Chan et al., 2022; Malashin et al., 2025).
In the construction of the explanatory variables, it was initially considered to use all accounts from the trial balance. However, to avoid an excessively large dataset, with correlated and redundant variables, the decision was made to create accounting indicators and indices derived from accounts relevant to fiscal auditing present in the trial balance.
These indicators were developed to represent the economic-financial behavior of companies throughout the analyzed fiscal year, aiming to reduce the dimensionality of the dataset and improve model interpretability. The selection of accounts and the construction of indicators were guided by audit hypotheses based on fiscal presumptions from the tax legislation of the State of Goiás.
Indicators involving accounts such as Cash, Banks, and Availability were created to identify deviations in the relationship between them. Indicators associated with the credit balance in the Cash account were also considered, such as the highest annual credit balance, the average credit balances, the number of months the account remained in credit, and its relevance against the company’s total assets.
On the liabilities side, indicators related to Suppliers and Customer Advances accounts were developed, calculating the ratio of these accounts to total liabilities, the behavior of credit and debit movements, the volatility of balances, and the frequency of periods without activity. Additionally, classic indicators such as Average Collection Period (ACP), Average Inventory Period (AIP), and Average Payment Period (APP) were calculated.
The explanatory variables were based on the accounting accounts standardized by the Reference Chart of Accounts (COD_CTA_REF), according to the linkage performed by the taxpayer in the transmission of the ECD. After analyzing each variable, logarithmic transformations in the form log(1 + x) were applied to reduce asymmetry and mitigate the impact of extreme values. Missing values were treated in a way that preserves the economic meaning of each indicator.
For the implementation of the predictive model, supervised machine learning algorithms based on decision trees were evaluated. The main algorithm chosen was Random Forest, due to its ability to capture non-linear relationships, handle complex interactions between variables, and present robust performance in scenarios of high dimensionality and class imbalance (Xu and Kong, 2024; Yang, 2025).
As a basis for comparison, a logistic regression model was used, serving as a linear reference to evaluate the gains obtained by adopting non-linear models. The comparison allowed us to verify the adequacy of choosing Random Forest for the proposed problem.
The evaluation strategy focused on the quality of the ranking, prioritizing top-precision to minimize false positives and increase the efficiency of the selection process, as recommended in the literature (Battaglini et al., 2025; Malashin et al., 2025). The dataset was divided into training (70%) and test (30%) samples through stratified sampling, preserving the original proportion of classes of the response variable.
Categorical variables, such as the two-digit CNAE, were treated through dummy coding (one-hot encoding) in the modeling. Model performance was evaluated based on traditional metrics, such as the Area Under the ROC Curve (AUC), and, primarily, by ranking quality metrics, such as Precision@K (Kar; Narasimhan; Jain, 2015). This choice aligned with the practical objective of prioritizing taxpayers with the highest potential for fiscal irregularity.
3. Results and Discussion
The results presented in this section derive from preliminary tests conducted with supervised machine learning models, aiming at the prioritization of tax audits. The analysis was based on accounting data extracted from the Digital Accounting Records (ECD) and derived variables, as detailed in the methodology. Given the limitation of human resources for conducting audits, the focus of the evaluation was on the models’ ability to concentrate, in the top positions of a ranking, companies with a higher probability of presenting tax discrepancies. To this end, an evaluation oriented towards the Precision@K metric was employed, which measures the proportion of positive cases among the top K ranked observations, complemented by other performance metrics.
The dataset used for the experiments was divided into training and testing samples, maintaining a proportion of 70% for training and 30% for testing, respectively. This division was performed through stratified sampling, ensuring the preservation of the original class proportions of the response variable (Y). Categorical variables, such as the National Classification of Economic Activities (CNAE) in two digits, were treated through dummy coding (one-hot encoding) for their inclusion in the models, ensuring that sectoral characteristics were adequately represented in the predictive analysis.
Model performance
For performance evaluation, different algorithms were selected and compared. Logistic Regression was employed as a linear reference model (*baseline*), while the Random Forest algorithm was tested in three distinct configurations: conservative, main, and deeper, reflecting different levels of complexity. Additionally, a random selection was included as a minimum comparison benchmark, allowing for the quantification of the gains provided by machine learning approaches. This comparison strategy aimed for a comprehensive understanding of the models’ effectiveness in a scenario of complex accounting data and imperfect labeling.
The results obtained in the test sample demonstrated a clear superiority of tree-based models, especially Random Forest, compared to Logistic Regression and random selection. The Area Under the ROC Curve (AUC) metric for the three Random Forest configurations was around 0.90, indicating an excellent global discrimination capacity between classes. In contrast, Logistic Regression achieved an AUC of approximately 0.78, while random selection obtained an AUC of 0.47, evidencing its ineffectiveness in ordering cases.
The deepest configuration of Random Forest showed the best overall performance, with an AUC of 0.91. This result is particularly relevant, as an AUC close to 1.0 indicates that the model is capable of distinguishing with high precision between contributors with and without indications of tax irregularities. The ROC curve of the Random Forest model in the main configuration, with an AUC of 0.90, graphically illustrates this separation capability, showing a significantly higher true positive rate compared to the false positive rate when compared to a random line, which is crucial for audit efficiency.
This superior performance of tree-based models is in line with recent literature, which points to the effectiveness of algorithms such as Random Forest in contexts of complex and multidimensional data, such as those found in tax audits (Chan et al., 2022; Yang, 2025). The ability of these models to capture non-linear relationships and complex interactions between accounting and registration variables, without the need for rigid assumptions about data distribution, contributes significantly to their robustness and predictive power in identifying risk patterns.
Ranking evaluation (Precision@K and Lift)
The performance evaluation of the models, focusing on ranking quality, was carried out using the Precision@K and Lift@K metrics, which are particularly suitable for the tax audit context, where prioritizing the most relevant cases is fundamental. The Precision@K metric measures the proportion of positive cases within the top K positions of the ranking, while Lift@K indicates how many times the model outperforms a random selection in concentrating positive cases. These metrics are crucial for the tax administration, which operates with limited resources and needs to optimize the allocation of audit efforts.
The deepest configuration of Random Forest demonstrated remarkable performance in concentrating positive cases at the top of the ranking. For K=50, a Precision@50 of 0.42 was obtained, meaning that 42% of the 50 companies best ranked by the model were, in fact, positive cases. The Lift@50 for this configuration was approximately 22.42, indicating that the use of the model is about 22 times more effective in identifying positive cases than a random selection among the 50 most prioritized companies. These values highlight the model’s ability to direct auditors to taxpayers most likely to present tax inconsistencies.
The main configuration of the Random Forest, adopted as a reference for this research stage, also presented expressive results, with Precision@50 of 0.34 and Lift@50 of approximately 18.15. This means that 34% of the 50 top-ranked companies by this model were positive cases, and the effectiveness in identifying these cases was about 18 times higher than that of a random selection. The analysis of the Precision@K metric behavior for the main model revealed that the precision is significantly higher in the first positions of the ranking, gradually decreasing as the value of K increases. This pattern is expected and desirable in prioritization models, as the highest risk cases are naturally concentrated at the top, optimizing audit efficiency.
In contrast, Logistic Regression presented a Precision@50 of 0.16 and a Lift@50 of approximately 6.41, evidencing a lower performance than Random Forest in capturing non-linear patterns and prioritizing risk contributors. Random selection, in turn, obtained a Precision@50 of approximately 0.02, serving as a baseline to scale the substantial gain provided by supervised approaches. The superiority of Random Forest in these metrics reinforces the adequacy of using machine learning models for contributor selection, allowing for a more efficient allocation of public resources.
Despite the slightly superior performance of the deeper Random Forest configuration at the top of the ranking, the decision was made to maintain the main configuration as the benchmark model. This decision was strategic, considering the nature of the problem, which involves a rare positive class and heterogeneous negative labels, which do not necessarily indicate tax compliance. An intermediate complexity model, such as the main Random Forest, can offer greater robustness, ranking stability, and better generalization potential, reducing the risk of overfitting to the training data and ensuring more consistent application in real-world audit scenarios.
Analysis of the generated ranking
The analysis of the ranking generated by the Random Forest model (main configuration) on the test sample revealed a significant concentration of positive cases (Y=1) in the first positions. For example, among the first 20 ranked observations, it was found that 12 belonged to class Y=1, indicating that the model has an adequate capacity to order cases associated with fiscal inconsistencies. The first positions, such as the first place with a score of 0.8584 (year 2022, CNAE 46) and the second with 0.8062 (year 2021, CNAE 47), were indeed positive cases, demonstrating the model’s effectiveness in identifying high-risk taxpayers.
However, the presence of some observations classified as Y=0 among the top positions in the ranking was also noted, such as the third place with a score of 0.7784 (year 2022, CNAE 46) and the seventh with 0.7671 (year 2021, CNAE 47). This behavior is consistent with the nature of the adopted response variable, in which a value of 0 (zero) does not necessarily imply the absence of fiscal irregularity. Due to the scarcity of human resources, many taxpayers classified as Y=0 may be cases that have not yet been audited or that, even if audited, have not had irregularities formally recorded, but may still show signs of infringement.
The positioning of these negatively labeled observations at the top of the ranking is, in fact, an indication of the model’s utility as a prioritization tool. These cases may represent potential targets of interest for future audits, as the model assigned them high scores based on accounting patterns that suggest risk, regardless of their enforcement history. Thus, the generated ranking provides objective subsidies for decision-making, assisting in the allocation of resources in tax audit processes and contributing to the rational use of public resources by directing attention to taxpayers who might otherwise go unnoticed.
Score distribution
The analysis of the distribution of scores assigned by the Random Forest model (main configuration) for the positive (Y=1) and negative (Y=0) classes revealed distinct patterns. It was observed that observations with a positive label, i.e., those with identified fiscal inconsistencies, received higher scores on average. In contrast, observations from the negative class tended to receive lower scores. This separation between the distributions of the two classes is a clear indication of the model’s ability to adequately discriminate between cases with and without indications of fiscal irregularity, which is fundamental for the effectiveness of prioritization.
Despite the clear separation, it was possible to identify overlapping regions between the score distributions of the positive and negative classes. This overlap is a reflection of the imperfect nature of the response variable, as discussed previously. Observations classified as negative (Y=0) may include contributors who have not yet been audited or who, even if audited, have not had irregularities formally identified, but who, in fact, present a risk. In these cases, the model assigned higher scores, positioning them in overlapping regions with the positive class, which reinforces the idea that these contributors may be valid targets for future audits.
The existence of this overlap, therefore, does not diminish the model’s validity, but rather highlights the complexity of the labeling problem in tax audits. It suggests that the model is capable of identifying risk patterns even in an environment where historical labels are incomplete or imperfect. This characteristic is particularly valuable for tax administrations, as it allows for the identification of potential irregularities that traditional tax meshes, based on rigid rules and historical labels, might not detect. Thus, the distribution of scores corroborates the model’s usefulness as a decision support tool, capable of refining the selection of taxpayers for audit.
Discussion of Results and Limitations
The preliminary research results indicate that supervised tree-based models, notably Random Forest, using data derived from the Digital Accounting Records (ECD), presented consistent and robust performance. This performance was observed even with the restrictions imposed by the adopted labeling process, which characterizes a scenario of imperfect labels. The comparison with the logistic regression linear model demonstrated the superiority of Random Forest in terms of both AUC and Precision@K, suggesting that the patterns associated with the problem of fiscal inconsistencies involve non-linear relationships and complex interactions between accounting variables, which are well captured by tree-based algorithms.
From an applied point of view, the findings indicate significant potential for increasing efficiency in selecting taxpayers for audit. The model’s ability to rank cases with indications of irregularity in the top positions allows the tax administration to optimize the allocation of its resources, directing audit efforts towards a reduced number of companies with a higher probability of presenting inconsistencies. This contributes to a more rational use of public resources, maximizing the impact of fiscal actions and improving revenue collection.
It is fundamental to observe that the presence of companies with a negative label (Y=0) in the top positions of the ranking should not be interpreted as a model error, but rather as a characteristic of the nature of the adopted label. The label Y=0 includes both companies evaluated without relevant findings and those that have not yet been audited. Thus, the positioning of these observations at the top of the ranking reflects the assignment of high scores to certain accounting profiles that the model identified as high risk, regardless of the available historical classification. This characterizes a typical scenario of classification with imperfect labels (*label noise*), where the model may identify risks not yet confirmed by previous audits.
Despite the promising results, the study presents important limitations that must be considered. The main limitation refers to the quality of the labels used, since the absence of penalties does not necessarily imply tax compliance, introducing uncertainty into the response variable. This imperfection in the labels may influence the model’s ability to generalize to new data. Furthermore, the model was trained based on historical audits, which may reflect biases inherent in the selection process previously adopted by the tax administration, potentially perpetuating existing audit patterns.
Another relevant limitation is that the analysis focused on a specific set of indicators derived from the ECD, not exploring other complementary fiscal bases, such as the Digital Fiscal Record (EFD) or the Electronic Invoice (NF-e), which could provide additional information for risk identification. The incorporation of new databases and the continuous improvement of the labeling process, with the inclusion of new positive labels from recent audits, represent promising avenues for the improvement and robustness of the model in future applications. In summary, the findings demonstrate the potential of machine learning in prioritizing audits, offering a valuable tool for tax administration, even in a challenging data context.
4. Conclusion
This study sought to apply and evaluate supervised machine learning models to generate risk scores, assisting in the prioritization of companies and periods for tax audit, considering the history of inconsistencies. It was found that tree-based models, especially Random Forest, demonstrated superior performance in discrimination and taxpayer prioritization capacity, compared to logistic regression and random selection. It was observed that Random Forest showed a high capacity to concentrate cases with indications of irregularity in the top positions of the ranking, even in a scenario of imperfect labels. The analysis revealed that the model is effective in identifying accounting patterns associated with risk, offering an objective subsidy for tax administration. This approach contributes significantly to the optimization of public resource allocation, directing audit efforts towards taxpayers with a higher probability of presenting tax inconsistencies, thus maximizing the impact of enforcement actions.
Despite the promising results, the study acknowledges important limitations, such as the quality of historical labels, where the absence of an assessment does not guarantee fiscal compliance, introducing uncertainty into the response variable. Additionally, training the model based on past audits may reflect biases inherent in previous selection processes. For future studies, continuous improvement of the labeling process is suggested, including new positive labels from recent audits, and the incorporation of other complementary fiscal databases, such as the Digital Fiscal Record (EFD) and the Electronic Invoice (NF-e). It is also recommended to expand the analyses to evaluate the robustness and stability of the model in different contexts. In summary, the proposed approach represents a valuable tool for decision support in taxpayer selection, with the potential to enhance the efficiency and rationality of public resource management.
Bibliographic References
Battaglini, M.; Guiso, L.; Lacava, F.; Miller, D.; Patacchini, E. 2025. Refining public policies with machine learning: the case of tax auditing. Journal of Econometrics 249: 105847.
Brasil. 2024. Manual de Orientação do Leiaute da Escrituração Contábil Digital (ECD). Versão 9.0. Receita Federal do Brasil [RFB], Brasília, DF, Brasil. Disponível em: https://www.gov.br/receitafederal/pt-br/centrais-de-conteudo/publicacoes/sped/ecd. Acesso em: 11 out. 2025.
Chan, T.; Tan, Y.; Tagkopoulos, I. 2022. Audit lead selection and yield prediction from historical tax data using artificial neural networks. PLOS ONE 17(11): e0278121.
Federação Nacional dos Auditores e Fiscais de Tributos Municipais [FENAFIM]. 2017. Machine Learning: aprendizagem de máquina e mineração de dados no combate à evasão fiscal. In: Prêmio Nacional de Educação Fiscal e Finanças Públicas, 2017, São Paulo, SP, Brasil. Anais… p. 1-10. São Paulo, SP, Brasil. Disponível em: https://www.fenafim.org.br/premio/wp-content/uploads/2019/04/2017_2_Machine_Learning.pdf. Acesso em: 11 out. 2025.
Kar, P.; Narasimhan, H.; Jain, P. 2015. Surrogate functions for maximizing precision at the top. arXiv preprint, arXiv:1505.06813. Disponível em: https://arxiv.org/abs/1505.06813. Acesso em: 01 fev. 2026.
Malashin, I.; Ruohonen, J.; Janes, A.; Iosifidis, G. 2025. Minimizing unnecessary tax audits using multi-objective hyperparameter tuning of XGBoost with focal loss. PLOS ONE 19(3): e0300928.
Silva, A.A. 2025. O uso dos dados agregados da ECD na auditoria contábil tributária. Instituto dos Auditores Fiscais do Estado da Bahia [IAF], Salvador, BA, Brasil. Disponível em: https://iaf.org.br/conteudo/9978/o-uso-dos-dados-agregados-da-ecd-na-auditoria-contabil-tributaria. Acesso em: 11 out. 2025.
Xu, C.; Kong, Y. 2024. Random forest model in tax risk identification of real estate enterprise income tax. PLOS ONE 19(3): e0300928.
Yang, L. 2025. Predictive modeling of tax compliance risks: a comparative study of machine learning approaches. PLOS ONE 20(9): e0331715.
Article originating from the Final Course Work of the Specialization in Data Science and Analytics of the MBA USP/Esalq
To learn more about the course, click here and access the MBX Academy platform