Executive Summary

Innovation

December 10, 2025

Machine learning models for tax classification from textual descriptions

Author: Luis Felipe Dalle Molle — Advisor: Vinicius Rocha Bíscaro

Summary prepared by the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege focused on synthesis and writing.

This study developed and evaluated machine learning models for the automatic classification of Mercosur Common Nomenclature (NCM) codes from IT product descriptions. The research compared the effectiveness of the Naive Bayes, Random Forest, and Support Vector Machine (SVM) algorithms on a real dataset from a company in the sector. The objective was to create a computational solution to mitigate the risks of manual fiscal classification, which presented an error rate of approximately 10%, and to optimize operational efficiency. The investigation sought to validate the technical applicability of the tools and quantify the performance gain compared to the human process, providing an empirical basis for the automation of tax tasks.

The fiscal classification of goods is fundamental for tax compliance and business competitiveness. The correct assignment of the NCM code determines the tax rates for taxes such as the Tax on Industrialized Products (IPI), PIS/Pasep, and COFINS. Errors in this process can result in fines, loss of tax benefits, and administrative rework. In addition to the financial impact, inaccuracy generates operational bottlenecks, such as the retention of goods during customs clearance, causing supply chain delays. Tax complexity and constant regulatory updates make the manual process prone to errors, justifying the search for automated solutions that offer greater accuracy and agility.

The NCM classification system is an eight-digit code based on the Harmonized System of Designation and Codification of Goods (HS), a six-digit international standard adopted by more than 200 countries (Worlds Customs Organization, 2025). Mercosur added two digits for further detail, resulting in over ten thousand possible classifications (Ministry of Finance, 2019). This granularity increases the complexity of the task, requiring deep technical knowledge of the product and legislation. The textual description of the goods complements the code, being essential for the transparency and compliance of the process.

For the IT industry, fiscal classification is strategic. The Brazilian government, through bodies such as the Ministry of Science, Technology and Innovation (MCTI), grants tax benefits linked to a specific list of NCM codes. An error in classification can disqualify the company from the benefit, directly impacting its cost structure and competitiveness. In this context, accuracy in assigning the NCM is a determining factor for the economic viability of operations in the technology sector.

Academic literature demonstrates the success of machine learning in text classification tasks. Algorithms such as Support Vector Machine, Naive Bayes, and Random Forest, combined with the Term Frequency-Inverse Document Frequency (TF-IDF) textual representation, show robust performance in product categorization (Altaheri & Shaalan, 2020). Recent research has evolved to models like large language models (LLMs) for tax classification (Marra de Artiñano et al., 2023) and multimodal approaches that include images (Amel et al., 2024). However, there is a gap in the literature regarding the application of these techniques to the Brazilian NCM system, a field this study aims to fill by analyzing the effectiveness of classic supervised models.

The methodology followed a Natural Language Processing (NLP) workflow. The starting point was a dataset from an IT company in São Paulo, with records from 2018 to 2023. The original database contained 4,601 observations, from which the long description, the short description, the engineering description (all in English), and the NCM code were selected. The process was structured in four stages: data preparation, textual preprocessing, vectorization, and modeling.

Data preparation included the removal of duplicate rows, missing data, and records with invalid HS codes. To address class imbalance, all HS code classes with fewer than four occurrences were excluded, mitigating the risk of bias from rare classes (Zhang & Wallace, 2017). After filtering, the three textual description fields were concatenated into a single variable, named All_Desc, to consolidate the descriptive information for each product.

Text preprocessing was applied to the All_Desc variable to normalize the text and focus on the most informative terms. The steps included converting to lowercase, removing special characters and punctuation, removing stopwords (common words without semantic value for classification), and applying stemming to reduce words to their morphological roots. These transformations are fundamental to creating a clean and efficient vocabulary, improving algorithm performance (Aggarwal & Zhai, 2012).

For the algorithms to process the text, the descriptions were converted into numerical vectors. The Bag-of-Words (BoW) model was used, which represents each text as a term frequency vector (McTear et al., 2016). This representation was refined with the TF-IDF weighting scheme, which assigns greater weight to words that are frequent in a specific document but rare in the general corpus, highlighting discriminative terms (Salton & Buckley, 1988).

Two vectorization configurations were tested: unigrams (individual words) and bigrams (pairs of consecutive words). Modeling was performed with the Naive Bayes, Random Forest, and SVM algorithms. The dataset was divided into 80% for training and 20% for testing, with stratified sampling. Hyperparameter optimization was done with GridSearchCV (Pedregosa et al., 2011), and the final evaluation was based on metrics such as accuracy, precision, recall, and F1-Score (Fávero et al., 2017).

The final database contained 4,279 observations and 50 NCM classes. Exploratory analysis revealed an unequal frequency distribution among the classes. NCM 8471.30.19, corresponding to “portable automatic data processing machines”, was one of the most representative, which is consistent with the company’s portfolio. Word frequency analysis within this class showed a strong association with terms such as “notebook”, “intel”, and “laptop”, indicating the presence of textual patterns explorable by the models.

The comparative evaluation tested the impact of textual representation (unigrams vs. bigrams) and algorithm (Naive Bayes, Random Forest, and SVM). The hypothesis was that bigrams would improve performance by capturing compound technical terms (Jurafsky & Martin, 2023). The algorithms were chosen for their distinct approaches: Naive Bayes, an efficient probabilistic model for sparse data (McCallum & Nigam, 1998); Random Forest, an ensemble method that models complex relationships (Breiman, 2001); and Support Vector Machine, a robust classifier for high-dimensional text classification tasks (Cortes & Vapnik, 1995).

The quantitative results confirmed the superiority of the Support Vector Machine model combined with the representation of text in TF-IDF weighted bigrams. This model achieved an overall accuracy of 96% and a weighted F1-Score of 95.8%. The SVM model with unigrams also performed well, with 95.5% accuracy. In comparison, Random Forest reached 94.5% accuracy with bigrams, while Naive Bayes obtained 92.8% in the same configuration.

The superior performance of SVM with bigrams can be explained by two factors. First, the ability of bigrams to capture technical expressions like “hard drive” or “power supply” provided the model with richer and more contextual features. Second, the nature of the SVM algorithm, which seeks the optimal separation hyperplane with maximum margin, is suitable for high-dimensional and sparse data such as TF-IDF representations, allowing for good generalization and avoiding overfitting (Cortes & Vapnik, 1995).

A detailed analysis by class was performed to ensure that overall performance did not mask failures in specific categories. The SVM model maintained high precision, recall, and F1-Score metrics even for classes with few samples. Majority classes, such as notebooks, presented an F1-Score of 99%, but many less frequent classes also surpassed 90%. The confusion matrix corroborated these findings, showing that most predictions were concentrated on the main diagonal, with few confusions between classes. This multiclass consistency is fundamental to ensuring the reliability of the model across the entire product portfolio.

The main implication of the results is their practical application. The performance of the SVM model, with an error rate of approximately 4% (inferred from 96% accuracy), represents a drastic improvement compared to the 10% error rate of the manual process. This reduction of more than 50% in the risk of classification errors has a direct impact on the company’s financial and operational health. Automation minimizes the probability of tax assessments and the loss of benefits, in addition to freeing up the tax team from a repetitive task. The time saved can be reallocated to higher value-added activities, such as strategic tax planning. The model’s consistency also eliminates the variability of human judgment, strengthening data governance.

The robustness of the final model demonstrates that machine learning is a powerful tool to face the complexity of the Brazilian tax system. The ability to transform textual descriptions into precise and automated fiscal classifications offers a competitive advantage, allowing the company to operate with greater legal certainty and agility. The research provides a compelling business case for the adoption of technological innovation in the fiscal area.

This work demonstrated the design and validation of an automated tax classification system. Starting from a concrete business problem — the high error rate in manual NCM classification —, a data science methodology was applied to develop an effective solution. The results confirmed that classic machine learning algorithms outperform human performance in complex textual classification tasks. The Support Vector Machine model with TF-IDF vectorization on bigrams emerged as the most performant approach, with 96% accuracy.

The contribution of this study is practical, by offering a tool to reduce fiscal risks and increase efficiency, and academic, by filling a gap in the application of NLP to the NCM system, providing a benchmark for future research. As next steps, it is suggested to explore class balancing techniques, such as SMOTE, and to experiment with models based on contextual embeddings, such as BERT. A future study could quantify the return on investment (ROI) of the solution’s implementation.

The research conducted demonstrated that the application of machine learning models, specifically Support Vector Machine with TF-IDF vectorization in bigrams, is an effective approach for the automatic fiscal classification of IT products, significantly outperforming the accuracy of the manual process.

References
AGGARWAL, Charu C.; ZHAI, Chengxiang. Mining Text Data. Springer, 2012.
ALTÁHERI, H.; SHAALAN, K. Automatic Classification of Customs Commodity Codes Using Machine Learning. International Journal of Advanced Computer Science and Applications (IJACSA), v. 11, n. 7, p. 441-447, 2020.
AMEL, A. et al. HSCodeNet: Multi-Modal Prediction of Harmonized System Codes. arXiv preprint, 2024. Available at: <https://arxiv. org/abs/2406.04349>. Accessed on: Sep. 22, 2025.
BREIMAN, Leo. Random Forests. Machine Learning, v. 45, n. 1, p. 5-32, 2001. Available at: <https://www. stat. berkeley. edu/~breiman/randomforest2001. pdf>. Accessed on: Sep. 12, 2025.
CORTES, C.; VAPNIK, V. Support-vector networks. Machine Learning, v. 20, p. 273–297, 1995. Available at: <https://link. springer. com/article/10.1007/BF00994018>. Accessed on: Sep. 12, 2025.
FÁVERO, L. P.; BELFIORE, P.; SILVA, F. L.; CHAN, B. Y. Manual de Análise de Dados: Estatística e Modelagem Multivariada com Excel®, SPSS® e Stata®. 2nd ed. Rio de Janeiro: Elsevier, 2017.
FAZCOMEX Tecnologia para Comércio Exterior LTDA. Consulta NCM. 2025. Available at: <https://www. fazcomex. com. br/ncm/>. Accessed on: Sep. 12, 2025.
JURAFSKY, D.; MARTIN, J. H. Speech and Language Processing. 3rd ed. (draft). 2023. Available at: <https://web. stanford. edu/~jurafsky/slp3/>. Accessed on: Sep. 10, 2025.
MARRA DE ARTIÑANO, I. et al. Classifying Goods with Machine Learning and Large Language Models. Proceedings of the 7th International Conference on Natural Language Processing and Information Retrieval (NLPIR), Tokyo, Japan, 2023.
MCCALLUM, A.; NIGAM, K. A comparison of event models for Naive Bayes text classification. AAAI-98 Workshop on Learning for Text Categorization, 1998.
MCTEAR, M.; CALLEJAS, Z.; GRIOL, D. Conversational Interfaces: Talking to Smart Devices. Springer International Publishing, 2016.
MINISTÉRIO DA FAZENDA. NCM. 2019. Available at: <https://www. gov. br/receitafederal/pt-br/assuntos/aduana-e-comercio-exterior/classificacao-fiscal-de-mercadorias/ncm>. Accessed on: Mar. 20, 2025.

PEDREGOSA, F. et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research, v. 12, p. 2825–2830, 2011. Available at: <https://jmlr. org/papers/v12/pedregosa11a. html>. Accessed on: Sep. 12, 2025.
SALTON, Gerard; BUCKLEY, Christopher. Term-weighting approaches in automatic text retrieval. Information Processing & Management, v. 24, n. 5, p. 513–523, 1988.
WORLDS CUSTOMS ORGANIZATION. Worlds Customs Organization – Official Web Page. Available at: <https://www. wcoomd. org/>. Accessed on: Mar. 20, 2025.
ZHANG, Y.; WALLACE, B. A sensitivity analysis of (and practitioners’ guide to) convolutional neural networks for sentence classification. Cornell University, 2016.


Executive summary from the Final Course Work of the Specialization in Data Science & Analytics of the MBA USP/Esalq

Learn more about the course; click here:

Who edited this article

Most recent

You may also like