Software Engineering
October 07, 2026
Application of Artificial Intelligence for Sentiment Analysis in Text Messages
Application of Artificial Intelligence for Sentiment Analysis in Text Messages
Isabel Priscila Chaves da Silva Rebello; Marcos Jardel Henriques
DOI: 10.22167/2675-6528-202603004
Article derived from a Final Course Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.
Summary
Sentiment analysis in e-commerce customer text messages is crucial for business decision-making. This study implemented artificial intelligence models with the objective of comparing their performance in evaluating these messages. The methodology consisted of applying a dataset of 4664 records in four distinct experiments: GPT 3.5 Turbo with Zero-shot, GPT 3.5 Turbo with Few-shot, BERT Multilingual Sentiment, and BERTimbau with Fine tuning. The results obtained indicated that the specialized model with Fine tuning presented the best evaluation metrics, a more consistent confusion matrix, and the shortest processing time. In contrast, the BERT Multilingual Sentiment model demonstrated difficulties in interpreting texts in the Portuguese language. The experiments with GPT 3.5 Turbo, particularly the Few-shot approach, proved to be promising, configuring themselves as a viable alternative in scenarios where Fine tuning cannot be applied. It was concluded that the Fine tuning strategy in pre-trained models, such as BERTimbau, offers significant advantages in accuracy and efficiency for large-scale sentiment analysis.
Keywords: Sentiment analysis; Fine Tuning; Artificial Intelligence; LLM; Natural Language Processing.
1. Introduction
Sentiment analysis (SA), also known as opinion mining or opinion extraction, has become a concept of great importance for companies, governments, and organizations. Understanding how customers and users perceive products and services is fundamental for strategic decision-making (SÁNCHEZ-RADA AND IGLESIAS, 2019).
Despite its relevance, collecting and analyzing consumer feedback presents significant challenges. Customer-generated content is often composed of unstructured data, such as free-text messages, which can address diverse topics and contain colloquial expressions, spelling errors, and informal language. This data nature complicates the AS process and the effective treatment of information (TAN, 2023).
In the context of e-commerce, sentiment analysis is a powerful tool with vast potential. Companies seeking to maintain competitiveness in the market need to refine and analyze consumer feedback. SA allows for obtaining valuable information about customer opinions, aiding in the detection of areas for improvement, the enhancement of products and services, and strategic adjustments to meet market demands (Soares, 2023). Furthermore, it enables customers to make more informed purchasing choices. With the large volume of unstructured data generated daily, a tool capable of extracting information on a large scale is essential for an agile understanding of consumer perception, generating strategic and operational advantages.
Sentiment analysis is a segment of Natural Language Processing (NLP) that aims to automatically identify and categorize sentiments and emotions expressed in texts (TAN, 2023). Historically, traditional methods like Bag-of-Words and TF-IDF classified texts by word frequency, but ignored context and semantics, limiting their effectiveness. The evolution led to the development of models like Word2Vec and GloVe, which capture semantic meaning, although they still face challenges with unknown words and the need for large datasets (KOKAB ET AL, 2022). More recently, transfer learning-based approaches, such as BERT, and Large Language Models (LLMs), such as GPT 3.5 Turbo, have emerged, demonstrating high generalization and accuracy in SA tasks, including with zero-shot and few-shot strategies (Souza et al., 2022).
Given the complexity and diversity of artificial intelligence models available for sentiment analysis, and considering the importance of customer feedback for Brazilian e-commerce, it becomes crucial to investigate which approach offers the best performance for text messages in Portuguese. The comparison between large language models (LLMs) with zero-shot and few-shot approaches and pre-trained models with fine-tuning, such as BERTimbau, is fundamental to identify the most effective solution in terms of accuracy and processing efficiency for this specific domain.
In this context, the present work seeks to compare the performance of different Artificial Intelligence models in sentiment analysis of text messages from an e-commerce customer.
2. Material and Methods
In order to achieve the proposed objectives, this study was conducted through a comparative and experimental approach. A literature review was carried out to identify promising artificial intelligence models, followed by the definition and application of these models to a dataset of customer text messages. Subsequently, the models were compared through an experimental evaluation and the results obtained were analyzed.
The unit of analysis consisted of a set of 4664 real text messages, collected from customer reviews of an e-commerce. These messages contained feedback on sellers, products, and delivery services, representing unstructured data generated by consumers (WANKHADE ET AL, 2022). The nature of this data allowed observation of the models’ performance when faced with colloquial expressions, spelling errors, and informal language.
The data preprocessing was a fundamental step to structure the texts. This process included the removal of special characters, such as symbols (except numbers) and URLs, and tokenization, which segmented the text into smaller units (PALOMINO, 2022; TAN ET AL, 2023). Additionally, messages with numerical ratings from 0 to 10 were processed to delimit these values, aiding in the model’s assimilation.
After preprocessing, the 4664 text messages were manually labeled into three sentiment categories: positive, negative, and neutral. The labeling was performed by a single annotator, prioritizing messages with explicit textual content. In cases of ambiguity, the negative class was adopted as the tie-breaking criterion. The labeling rules established as positive scores above seven or compliments, as negative scores below six with direct criticism, and as neutral scores between six and seven or messages unrelated to the services.
For sentiment analysis, three distinct artificial intelligence models were employed, evaluated in four experimental configurations. The models included GPT 3.5 Turbo, nlptown/bert-base-multilingual-uncased-sentiment (an adaptation of BERT for multilingual sentiment analysis), and neuralmind/bert-base-portuguese-cased, known as BERTimbau, trained specifically for Brazilian Portuguese (Souza et al., 2022).
The implementation of GPT 3.5 Turbo occurred in two configurations. In the zero-shot approach, the model classified sentiment (positive, negative, or neutral) based solely on the task description provided in the prompt, without prior examples. In the few-shot configuration, a set of examples was provided to the model to guide its behavior, using context to enhance the quality of inferences.
The BERT Multilingual Sentiment model was used in its pre-trained format. The BERTimbau model was applied using the fine-tuning technique, which involved training the model with the labeled dataset. For this process, the dataset was divided into training and testing partitions, with 70% and 20% of the data, respectively, with the validation partition not being used for direct model evaluation.
The hyperparameters employed in the fine-tuning of BERTimbau included the base model neuralmind/bert-base-portuguese-cased (BERTimbau Base), three training epochs, a learning rate of 2×10-5, and a batch size of eight for training and evaluation. The maximum length of 512 tokens was set, and the AdamW optimizer, along with the cross-entropy loss function, were used according to the standards of the Transformers library (Hugging Face).
The performance evaluation of the models was carried out using a set of statistical metrics widely recognized in the classification literature. The calculated metrics included Accuracy, Kappa Coefficient, Macro Precision, Macro Recall, Macro F1, Weighted Precision, Weighted Recall, and Weighted F1. Additionally, the confusion matrix was employed as a graphical tool to visualize the number of correct classifications in each sentiment class.
3. Results and Discussion
The comparative analysis of the performance of different artificial intelligence models in sentiment evaluation of e-commerce customer text messages revealed significant variations in evaluation metrics, classification consistency, and processing efficiency. The experiments were conducted with four distinct approaches: GPT 3.5 Turbo with Zero-shot, GPT 3.5 Turbo with Few-shot, BERT Multilingual Sentiment, and BERTimbau with Fine tuning. The results obtained allowed for the identification of the most effective strategy for the specific context of messages in Portuguese, according to the central objective of this study.
Initially, it was observed that the BERTimbau model with Fine tuning demonstrated the most robust performance across all evaluated statistical metrics. This model achieved an accuracy of 0.9154, a Kappa coefficient of 0.8701, Macro precision of 0.9076, Macro recall of 0.9114, and Macro F1 of 0.9085. The weighted metrics also followed this trend, with weighted precision of 0.9162, weighted recall of 0.9154, and weighted F1 of 0.9149. These consistently high values indicate a superior classification capability and a deep understanding of the contextual and semantic nuances of the Portuguese language, which is crucial for sentiment analysis in e-commerce data.
In contrast, the BERT Multilingual Sentiment model presented the worst performance among all tested approaches. Its metrics were notably low, with an accuracy of 0.5491, Kappa coefficient of 0.3311, Macro precision of 0.5548, Macro recall of 0.5675, and Macro F1 of 0.5126. Weighted metrics were also lower, with weighted precision of 0.5785, weighted recall of 0.5491, and weighted F1 of 0.522. This result suggests that, despite being a model designed for sentiment analysis and supporting multiple languages, it failed to efficiently capture the semantic particularities of Brazilian Portuguese, corroborating observations by Souza et al. (2022) on the limitations of multilingual models for specific languages.
The approaches with GPT 3.5 Turbo occupied an intermediate performance range. The GPT 3.5 Turbo Zero-shot model achieved an accuracy of 0.7562, Kappa of 0.625, Macro precision of 0.7477, Macro recall of 0.7556, and Macro F1 of 0.7512. The weighted metrics were similar, with weighted precision of 0.7551, weighted recall of 0.7562, and weighted F1 of 0.7553. Although it demonstrates the high capability of a Large Language Model (LLM) in performing classifications with minimal information, the absence of specific examples for the task resulted in inferior performance compared to more adapted approaches.
The application of the Few-shot strategy to GPT 3.5 Turbo resulted in a notable improvement compared to the Zero-shot approach. This model achieved an accuracy of 0.8218, Kappa of 0.7311, Macro precision of 0.8153, Macro recall of 0.8303, and Macro F1 of 0.8108. The weighted metrics were also superior, with weighted precision of 0.8431, weighted recall of 0.8218, and weighted F1 of 0.8224. The inclusion of examples in the prompt provided clearer context for the model, enhancing the quality of inferences and highlighting the potential of LLMs when guided with specific data, even without full fine-tuning, as pointed out by Fatouros (2023).
Beyond performance metrics, the processing capacity of the models was a critical factor in the evaluation. Models based on BERT, both Multilingual Sentiment and BERTimbau with Fine Tuning, demonstrated significantly higher efficiency. Multilingual Sentiment BERT processed the data in 2 minutes and 11 seconds, while BERTimbau with Fine Tuning, which used 934 records for testing, completed the task in just 25 seconds. This translates to an average time per record of approximately 0.028 seconds for Multilingual BERT and 0.027 seconds for BERTimbau, indicating a high processing speed for both.
In contrast, experiments with GPT 3.5 Turbo showed considerably longer processing times. The Zero-shot approach took 150 minutes and 1 second to process 4557 records, resulting in an average time of 1.975 seconds per record. The Few-shot approach performed similarly, with 150 minutes and 41 seconds for 4601 records, totaling 1.965 seconds per record. This time difference, about 73 times greater for LLMs per record, highlights a significant operational advantage of specific models in large-scale scenarios, where processing agility is essential for real-time understanding of consumer perception.
Another relevant aspect in the evaluation of GPT 3.5 Turbo models was the occurrence of request errors. The Zero-shot approach registered 2.30% errors, while the Few-shot presented 1.35% errors. These errors indicate that, despite good metrics, part of the data may not be classified, resulting in information loss. The lower error rate in the Few-shot approach can be attributed to the prompt size, which, by containing examples, demands more from the API and the model, but also offers more robust guidance, minimizing request failures compared to the Zero-shot approach, which relies exclusively on the model’s pre-existing knowledge.
Confusion Matrix
The analysis of the confusion matrices provided a detailed view of the types of classification errors made by each model. For the BERT Multilingual Sentiment model, the matrix revealed a high number of incorrect classifications. Although it correctly classified most negative instances (1022), the model made substantial errors in the positive and neutral classes. For example, it classified 740 neutral messages as negative and 552 positive messages as negative. This tendency to classify sentences as negative suggests a model bias when applied to Portuguese data, reinforcing the difficulty in handling the linguistic specificities of the language.
The confusion matrix of the BERTimbau model with Fine Tuning, in turn, demonstrated a high classification capacity in all classes. With a reduced number of errors, the model correctly classified 217 negative, 251 neutral, and 387 positive classifications. The errors were minimal, with only 13 negatives classified as neutral, 2 as positive; 34 neutrals as negative, 16 as positive; and 1 positive as negative, 13 as neutral. This consistency in the confusion matrix highlights the effectiveness of fine-tuning a monolingual model for Portuguese, allowing for precise adaptation to the dataset’s characteristics.
For the GPT 3.5 Turbo Zero-shot model, the confusion matrix indicated that, despite a reasonable number of correct predictions, the model showed a tendency to err between intermediate and extreme classes. It was observed that 240 neutral messages were classified as negative and 31 positive ones were also classified as negative. Furthermore, 338 positive messages were classified as neutral. This confusion suggests that the model, without specific examples, has difficulty differentiating sentiment nuances, classifying positives or negatives as neutral, or vice versa, which can lead to ambiguous interpretations in complex e-commerce contexts.
The confusion matrix of GPT 3.5 Turbo with the Few-shot strategy showed a significant improvement in accuracy compared to the Zero-shot approach. The model correctly classified 1098 negative messages, 980 neutral, and 1703 positive. However, some confusion was still noted when classifying neutral as negative (456) and positive as neutral (231). This confusion, although reduced, indicates that even with examples, the model may face challenges in contexts where the boundaries between sentiments are more subtle. The possibility of correcting these errors through the analysis of problematic sentences and the provision of new examples is an advantage of this strategy, allowing for continuous model improvement.
In summary, the research results demonstrate that the Fine tuning strategy in pre-trained models, such as BERTimbau, offers the highest precision and consistency in sentiment analysis of e-commerce customer text messages in Portuguese, in addition to a significantly lower processing time. The approaches with GPT 3.5 Turbo, especially Few-shot, showed promise and represent a viable alternative in scenarios where fine tuning is not applicable, despite presenting longer processing times and some confusion in classification. The BERT Multilingual Sentiment model, in turn, proved inadequate for the Portuguese language, highlighting the importance of linguistic specificity for the effectiveness of sentiment analysis.
4. Conclusion
This study aimed to compare the performance of different artificial intelligence models in sentiment analysis of customer text messages from an e-commerce platform. It was found that the BERTimbau model with Fine-tuning demonstrated the most robust performance, achieving the best statistical evaluation metrics and a consistent confusion matrix, in addition to presenting the lowest processing time. In contrast, the BERT Multilingual Sentiment model revealed significant difficulties in interpreting texts in Portuguese, resulting in the lowest metrics among the tested approaches. Implementations with GPT 3.5 Turbo, in both Zero-shot and Few-shot modalities, were situated in an intermediate performance range. The Few-shot strategy, in particular, showed a notable improvement compared to Zero-shot, configuring itself as a promising alternative in scenarios where Fine-tuning cannot be applied. However, it was observed that BERT-based models processed the data significantly faster, about 73 times more agile per record, compared to LLM approaches, which also registered request errors, indicating potential data loss.
Despite the promising results, it was identified that the use of Large Language Models such as GPT 3.5 Turbo can imply costs associated with requests and the occurrence of errors that result in data loss, aspects that demand consideration in large-scale applications. For future studies, it is suggested to track the operational costs of requests to GPT 3.5 Turbo in detail for a more in-depth economic feasibility analysis. It is also recommended to investigate strategies to improve model performance by analyzing error patterns in incorrectly classified messages and using anonymized examples, as well as exploring issues related to the General Law for the Protection of Personal Data (LGPD) in the treatment of customer data. The main contribution of this study lies in demonstrating that the Fine-tuning strategy on pre-trained models, such as BERTimbau, offers significant advantages in accuracy and efficiency for sentiment analysis in Portuguese, providing a valuable tool for e-commerce companies in agile monitoring of consumer perception and continuous improvement of products and services.
Bibliographic References
FATOUROS, Georgios; SOLDATOS, John; KOUROUMALI, Kalliopi; MAKRIDIS, Georgios; KYRIAZIS, Dimosthenis. (2023). Transforming sentiment analysis in the financial domain with ChatGPT. Machine Learning with Applications, v. 14, 100508.
KOKAB, Sayyida Tabinda; ASGHAR, Sohail; NAZ, Shehneela. (2022). Transformer-based deep learning models for the sentiment analysis of social media data. Array, v. 14, p. 100157.
PALOMINO, Marco A.; AIDER, Farida. (2022). Evaluating the Effectiveness of Text Pre-Processing in Sentiment Analysis. Applied Sciences, v. 12, n. 17, p. 8875.
SOARES, William Destefani. (2023). Avaliação de produtos baseada em análise de sentimentos aplicada a postagens textuais em redes sociais. 114 f. Trabalho de Conclusão de Curso (Graduação em Sistemas de Informação) – Instituto Federal do Espírito Santo, Campus Cachoeiro de Itapemirim.
SOUZA, F. C.; NOGUEIRA, R. F.; LOTUFO, R. A. (2022). BERT models for Brazilian Portuguese: pretraining, evaluation and tokenization analysis. In: Proceedings of the 9th Brazilian Conference on Intelligent Systems (BRACIS).
SÁNCHEZ-RADA, J. F.; IGLESIAS, C. A. (2019). Social context in sentiment analysis: formal definition, overview of current trends and framework for comparison. Information Fusion, v. 52, p. 344-356.
TAN, Kian Long; LEE, Chin Poo; LIM, Kian Ming. (2023). A Survey of Sentiment Analysis: Approaches, Datasets, and Future Research. Applied Sciences, v. 13, n. 7, p. 4550.
WANKHADE, Mayur; RAO, Annavarapu Chandra Sekhara; KULKARNI, Chaitanya. (2022). A survey on sentiment analysis methods, applications, and challenges. Artificial Intelligence Review, v. 55, p. 5731-5780.
Article originating from the Final Course Work of the Specialization in Software Engineering of the MBA USP/Esalq
To learn more about the course, click here and access the MBX Academy platform