Article

August 11, 2026

CLI tool for translation automation of web application localization files

Matheus Alvarez Crivellari; Eduardo Fernando Mendes

DOI: 10.22167/2675-6528-202601208

Article prepared by the ResumeAI tool, an artificial intelligence solution developed by the Pecege Institute focused on synthesis and writing.

Summary

The internationalization of web applications requires efficient and consistent translation processes, especially in preserving syntactic structures and specific terminology. However, manual translation of localization files incurs high operational costs, and conventional automatic tools lack refined terminological control and guarantees of structural integrity. Given this scenario, the objective was to develop and validate a command-line interface (CLI) tool to automate the translation of internationalization files in front-end web applications. For this purpose, generative artificial intelligence models were used, combined with the Retrieval-Augmented Generation (RAG) technique and the use of glossaries stored in vector databases. The research was characterized as exploratory, with a qualitative approach and a case study design, conducted in a corporate environment. The methodology included a literature review, the development of the tool on the Node.js platform, and its evaluation through metrics such as response time, structural preservation rate, and terminological consistency. The results demonstrated the technical and operational feasibility of the proposed solution, which contributed to the automation of the translation process and to the Software Engineering field, offering a scalable, reproducible, and applicable approach to different technological contexts.

Keywords: Vector database; Artificial intelligence; Internationalization; Retrieval-Augmented Generation [RAG].

1. Introduction

The globalized digital landscape drives the need for software applications to overcome geographical and linguistic barriers. This demand highlights the importance of software internationalization and localization, practices that allow companies to expand their reach and conquer new markets (Ribeiro, 2005). Internationalization, often identified as i18n, ensures that end-users can interact with applications in their native languages. Complementarily, localization, or l10n, adapts content in a culturally and linguistically acceptable manner for each specific region, encompassing everything from text translation to the formatting of numbers, dates, and user interface layout adjustments, as detailed by the Quick-start technical guide (Unicode, 2025). The adoption of these practices is crucial for the competitiveness and relevance of digital products in today’s market.

Despite its strategic relevance, the process of translating internationalization files in software applications, especially web front-end ones, faces considerable challenges. Manual translation, for example, is often associated with high operational costs and prolonged execution times. In agile development environments, where applications are frequently updated, manually maintaining translations for multiple languages generates unsustainable operational effort. Furthermore, reliance on human translators can introduce terminological inconsistencies, particularly in large-scale projects with distributed teams or over extensive development cycles.

Conventional machine translation (MT) tools emerge as an alternative to mitigate these problems. However, they have inherent limitations, such as difficulty in handling ambiguous words, contextual disambiguation, and specific domain knowledge, which includes technical terms and proper names (Naveen and Trojovský, 2024). Examples like Google Translate, although useful for general translations, generally do not offer refined terminological control. This can result in inconsistent translation of technical terms, product names, or brands throughout an application, compromising clarity and user experience.

Additionally, internationalization files contain not only natural language texts but also critical structural elements, such as placeholders and HTML markup. Preserving the integrity of these components is essential for the correct functioning of the application. Traditional automatic tools often fail to ensure this preservation, which can lead to functional errors. Given these complexities, recent approaches have explored the use of hybrid architectures based on artificial intelligence agents and Retrieval-Augmented Generation (RAG) techniques to handle this type of context (Chen et al., 2024). Artificial intelligence (AI), which encompasses the ability of machines to perform actions associated with human intelligence (Copeland, 2026), and Large Language Models (LLM) in particular, offer new perspectives for automated translation (McDonough, 2026).

Despite the advancement of AI technologies, an exploratory analysis of repositories like the Node.js Package Manager (NPM) reveals that the majority of existing packages for automatic translation of internationalization files lack robust guarantees of terminological consistency and structural preservation. This gap highlights the need for a solution that effectively integrates translation automation with mechanisms that ensure content quality and integrity. The Retrieval-Augmented Generation (RAG) technique enhances the output of language models by querying an external and reliable knowledge base, such as glossaries, improving translation accuracy and quality (AWS, 2026). Thus, the present research is justified by the growing demand for multilingual web applications and the need to optimize the translation process, reconciling operational efficiency with high linguistic quality.

The development of a tool that addresses these issues represents a significant contribution to the area of Software Engineering, offering a scalable, reproducible, and applicable approach to various technological contexts, benefiting teams involved in the development of multilingual applications. Given this scenario and the identified problem, the objective of this study was to develop a command-line interface (CLI) capable of automating the translation process of internationalization files in front-end web applications, using artificial intelligence models and the Retrieval-Augmented Generation (RAG) technique, and to compare the translation results with conventional machine translation methods, evaluating aspects such as structural preservation, terminological consistency, and computational resource optimization.

2. Material and Methods

This work was carried out as an exploratory research, aimed at investigating the use of artificial intelligence models and Retrieval-Augmented Generation (RAG) in the translation process of front-end web application internationalization files.

A qualitative approach was adopted, focused on the analysis of the use of artificial intelligence in the translation of internationalization files, seeking to understand how the proposed solution increases terminological consistency, preserves the structural integrity of the files, and reduces the need for human intervention. The research design was configured as a case study, which allowed for the evaluation of the practical application of the proposed tool in a real web development context.

The study was conducted in a medium-sized company, operating in the distance higher education sector, located in Piracicaba, São Paulo. The institution serves approximately 15,000 students annually distributed across Brazil and countries in America, Europe, Africa, and Asia. No personal data was collected, processed, or analyzed, focusing exclusively on the development and technical evaluation of a software tool.

The research began with a comprehensive bibliographic review of books, theses, dissertations, and scientific articles. The objective was to understand the use of artificial intelligence models in the automated translation process, the application of RAG for improving results, the construction of command-line interface (CLI) tools, the publication of packages in the Node.js Package Manager (NPM) registry, and software design patterns. This knowledge underpinned the development of the tool.

The command-line tool was built on the Node.js platform, using the TypeScript language. This choice was made due to its widespread use in modern front-end web development frameworks, such as React and Angular, and for the benefits of TypeScript’s static typing, which provides greater security, early error detection, and scalability to the application (TypeScript, 2026).

The solution architecture was structured in distinct layers, with well-defined responsibilities, following good software development and code organization practices. The layers included interface (CLI), application (orchestration), domain (business rules, such as glossary translation and indexing), infrastructure (communication with external providers), and cross-cutting (contracts and utilities).

In the domain layer, the Facade pattern was implemented to unify different tasks into a simplified high-level interface (Gamma et al., 1994). The infrastructure layer employed the Adapter and Strategy patterns, which allowed for the interchangeable use of different Large Language Models (LLM) providers, embedding models, and vector databases. Dependency Injection principles were applied to promote decoupling between layers.

A glossary was defined in .csv format, containing triplets of terms in Portuguese, English, and Spanish. This format was chosen because it is widely supported by TypeScript libraries and compatible with tabular manipulation tools. The CLI implemented a “glossary-index” command responsible for indexing the glossary terms, mapping them into high-dimensional vector representations and storing them in a vector database.

The transformation of glossary terms into vector representations, known as embedding, was performed using the open source model all-MiniLM-L6-v2, which maps sentences and paragraphs to a dense vector space of 384 dimensions and is made available by the Huggingface platform. As a vector database, the open source LanceDB project was used, installed locally as a Node.js dependency.

The CLI also implemented a “translate” command to perform the translation of JSON files from the source language to the target languages. This command loaded the source file into memory, traversed it recursively extracting the translatable textual segments (sentences), and separated them into batches of ten units for later submission to the artificial intelligence model.

Each sentence in the batch was transformed into a vector, again using the all-MiniLM-L6-v2 model. Subsequently, a vector similarity query was performed in LanceDB, employing the Euclidean distance metric to retrieve the five closest similar terms, which were used as additional context in the prompt sent to the artificial intelligence model.

The prompt was dynamically constructed at runtime, combining predefined translation instructions, relevant glossary terms retrieved via RAG, and the set of sentences to be translated itself. The prompt explicitly defined the translation rules, the use of the glossary when applicable, and the expected output format. The artificial intelligence models used were those provided by the OpenAI (gpt-40-mini) and Google (gemini-2.5-flash-lite and gemini-2.5-pro) platforms.

The artificial intelligence model’s response was received in a structured JSON format, with the translated values in the target language. The received JSON was interpreted and the translations were inserted into their respective positions in the target document’s JSON file. If translations already existed for the respective sentences in the target file, the existing ones were maintained, preserving previous human interventions.

For the evaluation of the developed tool, the following performance metrics were defined: average response time, structural preservation rate, terminological consistency, and scalability and applicability. The average response time was measured in seconds per translated file, obtained from logs automatically generated during the tool’s use.

The structural preservation rate was verified by automatically comparing original and translated files, measuring the occurrence of errors in structural elements such as placeholders and HTML markup. The algorithm loaded the source and destination JSONs, traversed them recursively, and extracted textual segments and sequences of structural elements, comparing them to identify size inconsistencies or positional divergences.

Terminological consistency was evaluated by checking if the terms defined in a glossary and present in the source language appeared translated in the target language exactly according to their established equivalences. The algorithm processed the glossary to identify and count all occurrences of the terms in the source language, eliminating overlaps to avoid duplicate counts, and associated each validated occurrence with the corresponding segment in the target language.

The terminological consistency metric was calculated by the ratio between the number of correctly translated occurrences and the total identified in the source language. Scalability and applicability were observed from controlled tests in different technological contexts, changing the artificial intelligence model provider and verifying the ease of integration of new LLM providers, databases, and embedding mechanisms.

The collected data were organized and analyzed descriptively and comparatively, seeking to identify performance patterns and practical implications for the development of automated translation solutions in Software Engineering. No identification of individuals occurred, and the research was not characterized as involving human beings, according to the definition of CNS Resolution n° 466/2012, thus dispensing with submission to the Research Ethics Committee.

3. Results and Discussion

The results obtained from the application of the proposed methodology revealed the technical and operational viability of the command-line interface (CLI) tool developed for the automation of internationalization file translation in front-end web applications. The analysis focused on evaluating metrics such as the response time of artificial intelligence models, terminological consistency, and structural preservation of the files. All tests were conducted using a source language file in Brazilian Portuguese, containing 1,439 sentences, which were sent to the artificial intelligence models in batches of ten sentences, totaling 144 batches per language. These findings are crucial for verifying the solution’s effectiveness and guiding future improvements, contributing to the field of Software Engineering.

As a baseline for comparison, the same source language file was translated by a conventional machine translation (MT) approach, employing the `i18n-auto-translation` library with Google Translate as the provider. This library, however, did not support the use of glossaries. The conventional translation results indicated a total duration of 96,000 seconds for US English and 113,000 seconds for Spanish from Spain. In terms of quality, the structural preservation rate was 56.19% for English and 55.52% for Spanish, while terminological consistency was 10.19% for English and only 0.49% for Spanish. These data established a performance benchmark for evaluating the solution based on artificial intelligence and RAG.

Average response time

The average response time metric represents the total period, in seconds, required for data submission to the server and return to the application, including translation processing. Performance varied significantly among the evaluated Large Language Models (LLM). The Google gemini-2.5-flash-lite model demonstrated to be the fastest, with an average time of 1.323 seconds per batch for United States English and 1.680 seconds per batch for Spain Spanish, resulting in total times of 190.497 and 241.974 seconds, respectively. In contrast, the Google gemini-2.5-pro presented the slowest performance, with averages of 11.977 seconds per batch for English and 14.440 seconds per batch for Spanish, totaling 1724.710 and 2079.318 seconds. The OpenAI gpt-4o-mini model registered intermediate times of 3.158 seconds per batch for English and 3.306 seconds per batch for Spanish, with totals of 454.778 and 476.020 seconds. It was observed that translation to English was consistently faster than to Spanish in all scenarios. However, all LLM models presented response times superior to the conventional MT approach, which was 96.000 seconds for English and 113.000 seconds for Spanish.

Structural preservation rate

The structural preservation rate assessed the solution’s ability to maintain the integrity of critical structural elements in JSON files, such as placeholders and HTML markup. The algorithm developed for this metric loaded the source and destination JSON files, traversing them recursively to extract and pair textual segments. For each valid pair, the sequences of structural elements were compared, ignoring textual content, and any divergence in size or position was considered a failure. The results demonstrated that the Google gemini-2.5-flash-lite models (for both English and Spanish), Google gemini-2.5-pro (for English and Spanish), and OpenAI gpt-4o-mini (for English) achieved a 100.00% structural preservation rate, with no failures in 110 sentences with structures. The OpenAI gpt-4o-mini model for Spanish obtained 99.76%, with only one failure in structural preservation. This performance represents an improvement of over 43% compared to conventional MT, which recorded rates of 56.19% and 55.52% for English and Spanish, respectively, indicating a significant advancement in ensuring file integrity.

Terminological consistency

The terminological consistency was evaluated by verifying if the terms defined in a glossary, present in the source language, were translated to the target language exactly according to their established equivalences. The glossary used consisted of 249 triplets of terms in Portuguese, English, and Spanish. The algorithm processed the JSON files, identified and counted 298 occurrences of the glossary terms in the source language, eliminating overlaps to avoid duplicate counts. It then checked for correct occurrences of the translations in the target language. The metric was calculated as the ratio between the number of correctly translated occurrences and the total identified in the source language. For United States English, the models showed high consistency: Google gemini-2.5-flash-lite achieved 88.93% (265 out of 298 correct occurrences), Google gemini-2.5-pro reached 89.60% (267 correct), and OpenAI gpt-4o-mini obtained 91.61% (273 correct). For Spain Spanish, the results were Google gemini-2.5-flash-lite with 70.13% (209 correct), Google gemini-2.5-pro with 80.87% (241 correct), and OpenAI gpt-4o-mini with 80.87% (241 correct). In all scenarios, consistency for English was higher than for Spanish. The comparison with conventional MT, which achieved 10.19% for English and 0.49% for Spanish, demonstrates a substantial improvement in terminological consistency with the proposed approach. Although terminological consistency rates above 80% are high, they still reflect known limitations of machine translation and Retrieval-Augmented Generation (RAG) based approaches. Traditional methods often struggle with translating context-dependent and domain-specific terms, resulting in terminological inconsistencies (Naveen and Trojovský, 2024). Recent studies with LLMs and RAG techniques indicate that the use of retrieved context and auxiliary resources, such as glossaries, reduces these inconsistencies, but does not eliminate them completely (Chen et al., 2024). Thus, the obtained results align with the literature, which points to terminological inconsistency as one of the main challenges in the field, reinforcing the relevance of combining RAG with structured glossaries for consistency improvement (Chen et al., 2025).

Scalability and applicability

The solution was structured with a focus on facilitating the evolution of the codebase, allowing for the agile incorporation of new LLM providers, relational databases, and embedding mechanisms. This architecture was designed with low coupling and minimal refactoring needs, which is fundamental for the maintenance and evolution of the system in dynamic development environments. The scalability and applicability feature was verified through the rapid implementation of concrete classes that interact with LLM providers from Google and OpenAI. This was possible only by extending the defined interfaces and implementing communication using the specific libraries of each vendor, demonstrating the tool’s flexibility to adapt to different technological contexts and artificial intelligence providers.

General discussion of results

The analysis of the results revealed that the developed tool, integrating artificial intelligence models and the RAG technique, offers a robust solution for the automation of internationalization file translation. Although LLM models presented superior response times compared to conventional MT, the quality of translations in terms of structural preservation and terminological consistency was significantly improved. The structural preservation rate reached almost 100% in all scenarios with AI models, an increase of over 43% compared to conventional MT, which is crucial to avoid functional failures in applications. Terminological consistency, which exceeded 80% in most cases, also represented an expressive improvement compared to the results of conventional MT, which were below 11%. These findings demonstrate that the proposed approach is effective in reconciling automation with the guarantee of linguistic and structural quality, addressing the gaps identified in traditional machine translation tools.

The layered architecture, based on design patterns such as Facade, Adapter, and Strategy, and Dependency Injection principles, was fundamental to ensuring the solution’s flexibility and scalability. This structure allows the tool to adapt to different artificial intelligence providers and vector database technologies, facilitating its maintenance and evolution. The ability to integrate glossaries through vector databases and the RAG technique was decisive in elevating terminological consistency, one of the main challenges of automated translation. Although terminological inconsistency has not been completely eliminated, the significant reduction observed validates the proposed hybrid approach. The choice of the artificial intelligence model directly impacts response time, allowing a balance between quality and performance according to project needs. In general, the results confirm the viability and potential of the tool to optimize the internationalization process in web applications, offering a scalable and reproducible approach for development teams.

4. Conclusion

The present study sought to develop a command-line interface (CLI) to automate the translation of internationalization files in front-end web applications, employing generative artificial intelligence models, the Retrieval-Augmented Generation (RAG) technique, and glossaries in vector databases, comparing its performance with conventional machine translation (MT) methods. The technical and operational feasibility of the proposed solution was verified, demonstrating significant advances in translation quality. In terms of structural preservation, artificial intelligence models achieved rates close to 100%, exceeding conventional MT performance by more than 43% and ensuring the integrity of critical elements such as placeholders and HTML tags. Additionally, a terminological consistency of over 80% was observed in most scenarios with the AI and RAG approach, representing a substantial improvement compared to conventional MT, which recorded rates below 11%. This hybrid approach contributed to the Software Engineering field by offering a scalable and reproducible solution for automating the translation process, reconciling efficiency with high linguistic and structural quality.

Despite the advances, it was noted that Large Language Models (LLM) presented response times superior to conventional MT, indicating a trade-off between speed and quality that can be adjusted according to project needs. Terminology consistency rates, although high, still reflect the inherent limitations of machine translation and RAG-based approaches for terms highly dependent on context and specific domain, as pointed out by the literature. However, the tool’s architecture, based on design patterns and low-coupling principles, demonstrated high scalability and applicability, allowing for easy integration of new LLM providers and vector database technologies. It is suggested that future studies explore the optimization of AI model response times and improve RAG mechanisms for even more complex terminological contexts, aiming to eliminate residual inconsistencies and expand the tool’s applicability to larger-scale scenarios.

Bibliographic References

Amazon Web Services [AWS]. n. d. O que é RAG (geração aumentada via recuperação)?. Disponível em: <https://aws.amazon.com/what-is/retrieval-augmented-generation>. Acesso em: 26 janeiro 2026.

Chen, H. et al. 2025. mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data. Disponível em <https://aclanthology.org/2025.findings-acl.433>. Acesso em: 26 janeiro 2026.

Chen, X. et al. 2024. CRAT: A Multi-Agent Framework for Causality-Enhanced Reflective and Retrieval-Augmented Translation with Large Language Models. Disponível em <https://arxiv.org/abs/2410.21067>. Acesso em: 27 outubro 2025.

Copeland, B. J. 2026. Enciclopédia Britannica. Disponível em: <https://www.britannica.com/technology/artificial-intelligence>. Acesso em: 26 janeiro 2026.

Gamma, E. et al. (2000). Padrões de Projeto. Soluções reutilizáveis de software orientado a objetos. 1ed. Bookman, Porto Alegre, RS, Brasil.

McDonough, M. 2026. Enciclopédia Britannica. Disponível em: <https://www.britannica.com/topic/large-language-model>. Acesso em: 26 janeiro 2026.

Naveen, P.; Trojovský, P. 2024. Overview and challenges of machine translation for contextually appropriate translations. iScience. Disponível em: <https://www.sciencedirect.com/science/article/pii/S2589004224021035>. Acesso em: 27 outubro 2025.

Ribeiro, G. C. B. 2005. Tradução e localização de software e outros produtos: audiovisual ou multimídia?. Cadernos de tradução. vol2. no16: 231-250. PUC-Rio, RJ, RJ.

TypeScript. n. d. What is TypeScript?. Disponível em: <https://www.typescriptlang.org/>. Acesso em: 18 abril 2026.

Unicode Consortium [UNICODE]. n.d. Technical Quick Start Guide. Disponível em: <https://home.unicode.org/technical-quick-start-guide/>. Acesso em: 30 setembro 2025.

Article originating from the Final Course Work of the Specialization in Software Engineering of the MBA USP/Esalq

To learn more about the course, click here and access the MBX Academy platform

You may also like

October 02, 2026

Determinants of supermarket location in São Paulo

A study investigated the determining factors for supermarket location in the state of São Paulo, with the objective of investigating the factors that explain the presence and expansion of these establishments, considering socioeconomic, demographic, and market dimensions. Data from the 2010 and 2022 Demographic Censuses of IBGE and information from the National Registry of Legal Entities of the Federal Revenue of Brazil were used to build a georeferenced database. A Random Forest classification model was applied, adjusted by grid search with cross-validation, prioritizing the recall-macro metric due to the imbalance of the dependent variable, which represented the presence or absence of supermarkets within a 50-meter buffer. The results indicated that supermarket location is strongly associated with demographic, income, and population characteristics in the surrounding area. The analysis of variable importance showed that sociodemographic factors, such as elderly literacy, household income, and the presence of other food establishments, exerted significant influence, especially in the immediate vicinity. The findings reinforced the hypothesis that the spatial distribution of supermarkets is not random, being conditioned by socioeconomic characteristics and the commercial structure of the territory, offering subsidies for business decisions and urban planning.

Keywords: Spatial Analysis; Machine learning; Expansion; Commercial location; Supermarkets.

Neuroscience And Learning In Education

October 02, 2026

Anti-Racist Education: Inclusive Educational Practices and Social Development

Antiracist education, understood as a structuring axis of inclusive education and social development, was investigated in the Brazilian context. The study aimed to identify and analyze, based on legal documents and teachers’ perceptions, educational practices capable of promoting antiracism in school and society, and how the implementation of Laws nº 10.639/03 and nº 11.645/08 contributed to social justice. A qualitative and documentary approach was adopted, with analysis of educational legislation, curricular guidelines, institutional reports, and academic literature. Complementarily, a semi-structured questionnaire was applied to 295 Basic Education teachers. The data were evaluated quantitatively and qualitatively, through thematic content analysis, and validated with bibliographic studies. The results revealed a paradox: despite a robust legal framework, the implementation of antiracist policies proved fragile and sporadic, with a lack of teacher training, adequate teaching materials, and monitoring. Significant educational inequalities between white and black students were found to persist, and most teachers acknowledged the occurrence of racism in schools, but without clear institutional protocols. Neuroscientific analysis showed that racism negatively impacts students’ cognitive and emotional development. It was concluded that antiracist education is central to quality education, requiring political commitment, public investment, and intersectoral articulation. The integration of Neuroscience in teacher training and the production of qualified materials are crucial to strengthen the school’s role in building a more just and inclusive society.

Keywords: Social Development; Antiracist Education; Social Justice; Law 10.639/03; Inclusive Educational Practices.

Neuroscience And Learning In Education

October 02, 2026

Paths of Inclusion: Perceptions of Parents and Teachers on the Schooling of Students with Dual Exceptionality in the Brazilian Context

Dual Exceptionality, characterized by the coexistence of High Abilities/Giftedness and neurodevelopmental disorders, represents a complex phenomenon that challenges traditional identification and schooling models. The study aimed to understand the perceptions of parents or guardians, teachers, and other education professionals regarding the schooling of students with Dual Exceptionality in the Brazilian context, investigating challenges, pedagogical strategies, and possibilities for inclusion based on equity. The research adopted a qualitative, exploratory, and descriptive approach, and collected data through an online, voluntary, and anonymous questionnaire answered by 25 participants. Discursive data were analyzed using thematic content analysis. The results indicated that knowledge about the topic is often built from personal and professional experiences, revealing gaps in systematic training. Difficulties were identified in identifying these students, in teacher training, and in implementing individualized educational plans, pedagogical flexibility, and curriculum enrichment. Socio-emotional repercussions, such as frustration and low self-esteem, were reported. However, some schools demonstrated inclusive practices based on equity, articulating specific needs and potentialities. Although the results do not allow for generalizations, they highlighted the need to strengthen professional training and the articulation between school, family, and specialized services. It was concluded that the inclusion of students with Dual Exceptionality requires practices that simultaneously recognize their difficulties and potentialities, ensuring equitable conditions for participation, learning, and development.

Keywords: Human development; Teacher training; School inclusion; Neurodivergence; Pedagogical practices.

October 02, 2026

Data Transformation into Strategy: Applied Research for Ecotourism Operation Optimization

The growing demand in ecotourism in Minas Gerais has driven the search for business intelligence to transform customer data into strategic information. The study aimed to structure a data science pipeline to collect, segment, and classify the customer base of an ecotourism operation, in order to optimize marketing actions and anticipate market movements. An exploratory, quali-quantitative research was conducted through a case study. 2,777 transactional records from an ecotourism company, referring to January 2024 to December 2025, were used. The methodological process involved automated data collection (Google Sheets API), processing and enrichment (ETL), validation, and creation of RFM (Recency, Frequency, and Monetary Value) attributes. Dimensionality reduction via PCA and K-Means clustering was applied, with the number of clusters defined by the Elbow method and Silhouette Score. The results were validated with DBSCAN and K-Medoids. The results revealed the identification of three behavioral customer segments: “Loyal”, “Low Value”, and “Potential”. The “Loyal” segment represented the highest accumulated economic value, while the “Potential” segment stood out for its high average ticket and potential for conversion into recurrence. The integration of data analysis techniques proved to be a robust and replicable method for generating intelligence in ecotourism. It was concluded that the structured data science pipeline enabled the behavioral segmentation of the customer base, the statistical validation of the groups, and the creation of a predictive system for new buyers, providing subsidies for data-driven strategic decisions and future analyses.

Keywords: Clustering; Business intelligence; Machine Learning; Customer segmentation; Decision making.

October 02, 2026

Classification of defaulting customers using supervised machine learning techniques

The risk of default in credit operations demanded analytical approaches to anticipate losses. This study comparatively evaluated the performance of supervised machine learning models in classifying defaulting customers in credit card operations. The public dataset “Default of Credit Card Clients” from the University of California Irvine was used, with 30,000 observations and class imbalance. The algorithms Logistic Regression, Random Forest, and Extreme Gradient Boosting were employed. The imbalance was addressed by assigning weights to the classes, and model optimization occurred with the RandomizedSearchCV method, prioritizing sensitivity. Cross-validation results indicated that the Extreme Gradient Boosting model showed a higher capacity for identifying the defaulting class and better discriminatory performance, followed by Random Forest and Logistic Regression, with a sensitivity of 0.8250 and an AUC-ROC of 0.7844 for XGBoost. Interpretability analysis, conducted by the Shapley Additive Explanations (SHAP) technique, highlighted the predominance of variables associated with payment behavior, especially the history of delays. It was concluded that tree-based models, particularly boosting techniques, proved to be more suitable for capturing complex patterns in the data, configuring themselves as consistent alternatives for credit risk management.

Keywords: Machine Learning; Credit Card; Classification; Extreme Gradient Boosting; Credit Risk.

October 02, 2026

Sentiment Analysis on Brazilian Banks on Twitter/X: Comparison between Traditional and Digital Institutions

A study analyzed public perception of Brazilian financial institutions on the Twitter/X platform, highlighting the importance of sentiment monitoring on social networks for understanding reputation and customer experience in the banking sector. The objective was to compare user perception of the image and reputation of traditional and digital banks, based on the sentiment patterns identified in the analyzed manifestations, seeking to identify structural differences between these groups. The methodology was based on the analysis of 1,096 tweets collected between November 2022 and June 2023. Two complementary sentiment analysis approaches were used, the sum and the average of labels, to capture the majority sentiment and nuances of perception. Additionally, the Market Profile Model, with indicators of emotional reputation, reputational risk, neutrality, and polarization, and the Banking Clustering Model, which allowed grouping institutions according to perception patterns, were developed. The results indicated a predominance of neutral and negative sentiments, a higher volume of interactions in digital banks, and structural differences in the emotional intensity of perceptions, with greater stability in digital banks and greater polarization in traditional ones. It was concluded that the combination of analytical and statistical techniques contributed to an in-depth understanding of institutional image in the digital environment, demonstrating the importance of data-driven reputation management strategies.

Keywords: Digital banks; Traditional banks; Data modeling; Opinion mining; Social Networks.

October 02, 2026

Optimization of annual budget planning through project management methodologies

The Annual Budget Planning (POA) is a crucial process for translating organizational strategy into operational and financial goals, but it frequently faces deadline pressures, interdepartmental dependencies, and the repetition of habitual expenses. The study aimed to analyze how the combined application of project management practices and Zero-Based Budgeting (OBZ) can optimize the POA. To this end, a case study was developed in the Brazilian operation of a publicly traded company in the beverage sector, using documentary research of its 2023 results report and an anonymous questionnaire applied to 47 respondents. Documentary analysis indicated growth in net revenue, expansion of gross profit and adjusted EBITDA, and contained advancement of selling, general, and administrative expenses, suggesting cost discipline and operational leverage. The complementary survey revealed a high perception of cascading effect on the schedule, strong support for defining cost package owners, and a preference for technical justification of expenses, in addition to demand for controlled flexibility after the baseline definition. It was concluded that structuring the POA as a project, associated with the rigor of OBZ, increased the process predictability, reinforced accountability for expenses, and broadened the coherence between budgetary execution and economic-financial performance.

Keywords: Cost Control; Operational Efficiency; Zero-Based Budgeting; PMBOK; Beverage Sector.