August 11, 2026
CLI tool for translation automation of web application localization files
Matheus Alvarez Crivellari; Eduardo Fernando Mendes
DOI: 10.22167/2675-6528-202601208
Article prepared by the ResumeAI tool, an artificial intelligence solution developed by the Pecege Institute focused on synthesis and writing.
Summary
The internationalization of web applications requires efficient and consistent translation processes, especially in preserving syntactic structures and specific terminology. However, manual translation of localization files incurs high operational costs, and conventional automatic tools lack refined terminological control and guarantees of structural integrity. Given this scenario, the objective was to develop and validate a command-line interface (CLI) tool to automate the translation of internationalization files in front-end web applications. For this purpose, generative artificial intelligence models were used, combined with the Retrieval-Augmented Generation (RAG) technique and the use of glossaries stored in vector databases. The research was characterized as exploratory, with a qualitative approach and a case study design, conducted in a corporate environment. The methodology included a literature review, the development of the tool on the Node.js platform, and its evaluation through metrics such as response time, structural preservation rate, and terminological consistency. The results demonstrated the technical and operational feasibility of the proposed solution, which contributed to the automation of the translation process and to the Software Engineering field, offering a scalable, reproducible, and applicable approach to different technological contexts.
Keywords: Vector database; Artificial intelligence; Internationalization; Retrieval-Augmented Generation [RAG].
1. Introduction
The globalized digital landscape drives the need for software applications to overcome geographical and linguistic barriers. This demand highlights the importance of software internationalization and localization, practices that allow companies to expand their reach and conquer new markets (Ribeiro, 2005). Internationalization, often identified as i18n, ensures that end-users can interact with applications in their native languages. Complementarily, localization, or l10n, adapts content in a culturally and linguistically acceptable manner for each specific region, encompassing everything from text translation to the formatting of numbers, dates, and user interface layout adjustments, as detailed by the Quick-start technical guide (Unicode, 2025). The adoption of these practices is crucial for the competitiveness and relevance of digital products in today’s market.
Despite its strategic relevance, the process of translating internationalization files in software applications, especially web front-end ones, faces considerable challenges. Manual translation, for example, is often associated with high operational costs and prolonged execution times. In agile development environments, where applications are frequently updated, manually maintaining translations for multiple languages generates unsustainable operational effort. Furthermore, reliance on human translators can introduce terminological inconsistencies, particularly in large-scale projects with distributed teams or over extensive development cycles.
Conventional machine translation (MT) tools emerge as an alternative to mitigate these problems. However, they have inherent limitations, such as difficulty in handling ambiguous words, contextual disambiguation, and specific domain knowledge, which includes technical terms and proper names (Naveen and Trojovský, 2024). Examples like Google Translate, although useful for general translations, generally do not offer refined terminological control. This can result in inconsistent translation of technical terms, product names, or brands throughout an application, compromising clarity and user experience.
Additionally, internationalization files contain not only natural language texts but also critical structural elements, such as placeholders and HTML markup. Preserving the integrity of these components is essential for the correct functioning of the application. Traditional automatic tools often fail to ensure this preservation, which can lead to functional errors. Given these complexities, recent approaches have explored the use of hybrid architectures based on artificial intelligence agents and Retrieval-Augmented Generation (RAG) techniques to handle this type of context (Chen et al., 2024). Artificial intelligence (AI), which encompasses the ability of machines to perform actions associated with human intelligence (Copeland, 2026), and Large Language Models (LLM) in particular, offer new perspectives for automated translation (McDonough, 2026).
Despite the advancement of AI technologies, an exploratory analysis of repositories like the Node.js Package Manager (NPM) reveals that the majority of existing packages for automatic translation of internationalization files lack robust guarantees of terminological consistency and structural preservation. This gap highlights the need for a solution that effectively integrates translation automation with mechanisms that ensure content quality and integrity. The Retrieval-Augmented Generation (RAG) technique enhances the output of language models by querying an external and reliable knowledge base, such as glossaries, improving translation accuracy and quality (AWS, 2026). Thus, the present research is justified by the growing demand for multilingual web applications and the need to optimize the translation process, reconciling operational efficiency with high linguistic quality.
The development of a tool that addresses these issues represents a significant contribution to the area of Software Engineering, offering a scalable, reproducible, and applicable approach to various technological contexts, benefiting teams involved in the development of multilingual applications. Given this scenario and the identified problem, the objective of this study was to develop a command-line interface (CLI) capable of automating the translation process of internationalization files in front-end web applications, using artificial intelligence models and the Retrieval-Augmented Generation (RAG) technique, and to compare the translation results with conventional machine translation methods, evaluating aspects such as structural preservation, terminological consistency, and computational resource optimization.
2. Material and Methods
This work was carried out as an exploratory research, aimed at investigating the use of artificial intelligence models and Retrieval-Augmented Generation (RAG) in the translation process of front-end web application internationalization files.
A qualitative approach was adopted, focused on the analysis of the use of artificial intelligence in the translation of internationalization files, seeking to understand how the proposed solution increases terminological consistency, preserves the structural integrity of the files, and reduces the need for human intervention. The research design was configured as a case study, which allowed for the evaluation of the practical application of the proposed tool in a real web development context.
The study was conducted in a medium-sized company, operating in the distance higher education sector, located in Piracicaba, São Paulo. The institution serves approximately 15,000 students annually distributed across Brazil and countries in America, Europe, Africa, and Asia. No personal data was collected, processed, or analyzed, focusing exclusively on the development and technical evaluation of a software tool.
The research began with a comprehensive bibliographic review of books, theses, dissertations, and scientific articles. The objective was to understand the use of artificial intelligence models in the automated translation process, the application of RAG for improving results, the construction of command-line interface (CLI) tools, the publication of packages in the Node.js Package Manager (NPM) registry, and software design patterns. This knowledge underpinned the development of the tool.
The command-line tool was built on the Node.js platform, using the TypeScript language. This choice was made due to its widespread use in modern front-end web development frameworks, such as React and Angular, and for the benefits of TypeScript’s static typing, which provides greater security, early error detection, and scalability to the application (TypeScript, 2026).
The solution architecture was structured in distinct layers, with well-defined responsibilities, following good software development and code organization practices. The layers included interface (CLI), application (orchestration), domain (business rules, such as glossary translation and indexing), infrastructure (communication with external providers), and cross-cutting (contracts and utilities).
In the domain layer, the Facade pattern was implemented to unify different tasks into a simplified high-level interface (Gamma et al., 1994). The infrastructure layer employed the Adapter and Strategy patterns, which allowed for the interchangeable use of different Large Language Models (LLM) providers, embedding models, and vector databases. Dependency Injection principles were applied to promote decoupling between layers.
A glossary was defined in .csv format, containing triplets of terms in Portuguese, English, and Spanish. This format was chosen because it is widely supported by TypeScript libraries and compatible with tabular manipulation tools. The CLI implemented a “glossary-index” command responsible for indexing the glossary terms, mapping them into high-dimensional vector representations and storing them in a vector database.
The transformation of glossary terms into vector representations, known as embedding, was performed using the open source model all-MiniLM-L6-v2, which maps sentences and paragraphs to a dense vector space of 384 dimensions and is made available by the Huggingface platform. As a vector database, the open source LanceDB project was used, installed locally as a Node.js dependency.
The CLI also implemented a “translate” command to perform the translation of JSON files from the source language to the target languages. This command loaded the source file into memory, traversed it recursively extracting the translatable textual segments (sentences), and separated them into batches of ten units for later submission to the artificial intelligence model.
Each sentence in the batch was transformed into a vector, again using the all-MiniLM-L6-v2 model. Subsequently, a vector similarity query was performed in LanceDB, employing the Euclidean distance metric to retrieve the five closest similar terms, which were used as additional context in the prompt sent to the artificial intelligence model.
The prompt was dynamically constructed at runtime, combining predefined translation instructions, relevant glossary terms retrieved via RAG, and the set of sentences to be translated itself. The prompt explicitly defined the translation rules, the use of the glossary when applicable, and the expected output format. The artificial intelligence models used were those provided by the OpenAI (gpt-40-mini) and Google (gemini-2.5-flash-lite and gemini-2.5-pro) platforms.
The artificial intelligence model’s response was received in a structured JSON format, with the translated values in the target language. The received JSON was interpreted and the translations were inserted into their respective positions in the target document’s JSON file. If translations already existed for the respective sentences in the target file, the existing ones were maintained, preserving previous human interventions.
For the evaluation of the developed tool, the following performance metrics were defined: average response time, structural preservation rate, terminological consistency, and scalability and applicability. The average response time was measured in seconds per translated file, obtained from logs automatically generated during the tool’s use.
The structural preservation rate was verified by automatically comparing original and translated files, measuring the occurrence of errors in structural elements such as placeholders and HTML markup. The algorithm loaded the source and destination JSONs, traversed them recursively, and extracted textual segments and sequences of structural elements, comparing them to identify size inconsistencies or positional divergences.
Terminological consistency was evaluated by checking if the terms defined in a glossary and present in the source language appeared translated in the target language exactly according to their established equivalences. The algorithm processed the glossary to identify and count all occurrences of the terms in the source language, eliminating overlaps to avoid duplicate counts, and associated each validated occurrence with the corresponding segment in the target language.
The terminological consistency metric was calculated by the ratio between the number of correctly translated occurrences and the total identified in the source language. Scalability and applicability were observed from controlled tests in different technological contexts, changing the artificial intelligence model provider and verifying the ease of integration of new LLM providers, databases, and embedding mechanisms.
The collected data were organized and analyzed descriptively and comparatively, seeking to identify performance patterns and practical implications for the development of automated translation solutions in Software Engineering. No identification of individuals occurred, and the research was not characterized as involving human beings, according to the definition of CNS Resolution n° 466/2012, thus dispensing with submission to the Research Ethics Committee.
3. Results and Discussion
The results obtained from the application of the proposed methodology revealed the technical and operational viability of the command-line interface (CLI) tool developed for the automation of internationalization file translation in front-end web applications. The analysis focused on evaluating metrics such as the response time of artificial intelligence models, terminological consistency, and structural preservation of the files. All tests were conducted using a source language file in Brazilian Portuguese, containing 1,439 sentences, which were sent to the artificial intelligence models in batches of ten sentences, totaling 144 batches per language. These findings are crucial for verifying the solution’s effectiveness and guiding future improvements, contributing to the field of Software Engineering.
As a baseline for comparison, the same source language file was translated by a conventional machine translation (MT) approach, employing the `i18n-auto-translation` library with Google Translate as the provider. This library, however, did not support the use of glossaries. The conventional translation results indicated a total duration of 96,000 seconds for US English and 113,000 seconds for Spanish from Spain. In terms of quality, the structural preservation rate was 56.19% for English and 55.52% for Spanish, while terminological consistency was 10.19% for English and only 0.49% for Spanish. These data established a performance benchmark for evaluating the solution based on artificial intelligence and RAG.
Average response time
The average response time metric represents the total period, in seconds, required for data submission to the server and return to the application, including translation processing. Performance varied significantly among the evaluated Large Language Models (LLM). The Google gemini-2.5-flash-lite model demonstrated to be the fastest, with an average time of 1.323 seconds per batch for United States English and 1.680 seconds per batch for Spain Spanish, resulting in total times of 190.497 and 241.974 seconds, respectively. In contrast, the Google gemini-2.5-pro presented the slowest performance, with averages of 11.977 seconds per batch for English and 14.440 seconds per batch for Spanish, totaling 1724.710 and 2079.318 seconds. The OpenAI gpt-4o-mini model registered intermediate times of 3.158 seconds per batch for English and 3.306 seconds per batch for Spanish, with totals of 454.778 and 476.020 seconds. It was observed that translation to English was consistently faster than to Spanish in all scenarios. However, all LLM models presented response times superior to the conventional MT approach, which was 96.000 seconds for English and 113.000 seconds for Spanish.
Structural preservation rate
The structural preservation rate assessed the solution’s ability to maintain the integrity of critical structural elements in JSON files, such as placeholders and HTML markup. The algorithm developed for this metric loaded the source and destination JSON files, traversing them recursively to extract and pair textual segments. For each valid pair, the sequences of structural elements were compared, ignoring textual content, and any divergence in size or position was considered a failure. The results demonstrated that the Google gemini-2.5-flash-lite models (for both English and Spanish), Google gemini-2.5-pro (for English and Spanish), and OpenAI gpt-4o-mini (for English) achieved a 100.00% structural preservation rate, with no failures in 110 sentences with structures. The OpenAI gpt-4o-mini model for Spanish obtained 99.76%, with only one failure in structural preservation. This performance represents an improvement of over 43% compared to conventional MT, which recorded rates of 56.19% and 55.52% for English and Spanish, respectively, indicating a significant advancement in ensuring file integrity.
Terminological consistency
The terminological consistency was evaluated by verifying if the terms defined in a glossary, present in the source language, were translated to the target language exactly according to their established equivalences. The glossary used consisted of 249 triplets of terms in Portuguese, English, and Spanish. The algorithm processed the JSON files, identified and counted 298 occurrences of the glossary terms in the source language, eliminating overlaps to avoid duplicate counts. It then checked for correct occurrences of the translations in the target language. The metric was calculated as the ratio between the number of correctly translated occurrences and the total identified in the source language. For United States English, the models showed high consistency: Google gemini-2.5-flash-lite achieved 88.93% (265 out of 298 correct occurrences), Google gemini-2.5-pro reached 89.60% (267 correct), and OpenAI gpt-4o-mini obtained 91.61% (273 correct). For Spain Spanish, the results were Google gemini-2.5-flash-lite with 70.13% (209 correct), Google gemini-2.5-pro with 80.87% (241 correct), and OpenAI gpt-4o-mini with 80.87% (241 correct). In all scenarios, consistency for English was higher than for Spanish. The comparison with conventional MT, which achieved 10.19% for English and 0.49% for Spanish, demonstrates a substantial improvement in terminological consistency with the proposed approach. Although terminological consistency rates above 80% are high, they still reflect known limitations of machine translation and Retrieval-Augmented Generation (RAG) based approaches. Traditional methods often struggle with translating context-dependent and domain-specific terms, resulting in terminological inconsistencies (Naveen and Trojovský, 2024). Recent studies with LLMs and RAG techniques indicate that the use of retrieved context and auxiliary resources, such as glossaries, reduces these inconsistencies, but does not eliminate them completely (Chen et al., 2024). Thus, the obtained results align with the literature, which points to terminological inconsistency as one of the main challenges in the field, reinforcing the relevance of combining RAG with structured glossaries for consistency improvement (Chen et al., 2025).
Scalability and applicability
The solution was structured with a focus on facilitating the evolution of the codebase, allowing for the agile incorporation of new LLM providers, relational databases, and embedding mechanisms. This architecture was designed with low coupling and minimal refactoring needs, which is fundamental for the maintenance and evolution of the system in dynamic development environments. The scalability and applicability feature was verified through the rapid implementation of concrete classes that interact with LLM providers from Google and OpenAI. This was possible only by extending the defined interfaces and implementing communication using the specific libraries of each vendor, demonstrating the tool’s flexibility to adapt to different technological contexts and artificial intelligence providers.
General discussion of results
The analysis of the results revealed that the developed tool, integrating artificial intelligence models and the RAG technique, offers a robust solution for the automation of internationalization file translation. Although LLM models presented superior response times compared to conventional MT, the quality of translations in terms of structural preservation and terminological consistency was significantly improved. The structural preservation rate reached almost 100% in all scenarios with AI models, an increase of over 43% compared to conventional MT, which is crucial to avoid functional failures in applications. Terminological consistency, which exceeded 80% in most cases, also represented an expressive improvement compared to the results of conventional MT, which were below 11%. These findings demonstrate that the proposed approach is effective in reconciling automation with the guarantee of linguistic and structural quality, addressing the gaps identified in traditional machine translation tools.
The layered architecture, based on design patterns such as Facade, Adapter, and Strategy, and Dependency Injection principles, was fundamental to ensuring the solution’s flexibility and scalability. This structure allows the tool to adapt to different artificial intelligence providers and vector database technologies, facilitating its maintenance and evolution. The ability to integrate glossaries through vector databases and the RAG technique was decisive in elevating terminological consistency, one of the main challenges of automated translation. Although terminological inconsistency has not been completely eliminated, the significant reduction observed validates the proposed hybrid approach. The choice of the artificial intelligence model directly impacts response time, allowing a balance between quality and performance according to project needs. In general, the results confirm the viability and potential of the tool to optimize the internationalization process in web applications, offering a scalable and reproducible approach for development teams.
4. Conclusion
The present study sought to develop a command-line interface (CLI) to automate the translation of internationalization files in front-end web applications, employing generative artificial intelligence models, the Retrieval-Augmented Generation (RAG) technique, and glossaries in vector databases, comparing its performance with conventional machine translation (MT) methods. The technical and operational feasibility of the proposed solution was verified, demonstrating significant advances in translation quality. In terms of structural preservation, artificial intelligence models achieved rates close to 100%, exceeding conventional MT performance by more than 43% and ensuring the integrity of critical elements such as placeholders and HTML tags. Additionally, a terminological consistency of over 80% was observed in most scenarios with the AI and RAG approach, representing a substantial improvement compared to conventional MT, which recorded rates below 11%. This hybrid approach contributed to the Software Engineering field by offering a scalable and reproducible solution for automating the translation process, reconciling efficiency with high linguistic and structural quality.
Despite the advances, it was noted that Large Language Models (LLM) presented response times superior to conventional MT, indicating a trade-off between speed and quality that can be adjusted according to project needs. Terminology consistency rates, although high, still reflect the inherent limitations of machine translation and RAG-based approaches for terms highly dependent on context and specific domain, as pointed out by the literature. However, the tool’s architecture, based on design patterns and low-coupling principles, demonstrated high scalability and applicability, allowing for easy integration of new LLM providers and vector database technologies. It is suggested that future studies explore the optimization of AI model response times and improve RAG mechanisms for even more complex terminological contexts, aiming to eliminate residual inconsistencies and expand the tool’s applicability to larger-scale scenarios.
Bibliographic References
Amazon Web Services [AWS]. n. d. O que é RAG (geração aumentada via recuperação)?. Disponível em: <https://aws.amazon.com/what-is/retrieval-augmented-generation>. Acesso em: 26 janeiro 2026.
Chen, H. et al. 2025. mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data. Disponível em <https://aclanthology.org/2025.findings-acl.433>. Acesso em: 26 janeiro 2026.
Chen, X. et al. 2024. CRAT: A Multi-Agent Framework for Causality-Enhanced Reflective and Retrieval-Augmented Translation with Large Language Models. Disponível em <https://arxiv.org/abs/2410.21067>. Acesso em: 27 outubro 2025.
Copeland, B. J. 2026. Enciclopédia Britannica. Disponível em: <https://www.britannica.com/technology/artificial-intelligence>. Acesso em: 26 janeiro 2026.
Gamma, E. et al. (2000). Padrões de Projeto. Soluções reutilizáveis de software orientado a objetos. 1ed. Bookman, Porto Alegre, RS, Brasil.
McDonough, M. 2026. Enciclopédia Britannica. Disponível em: <https://www.britannica.com/topic/large-language-model>. Acesso em: 26 janeiro 2026.
Naveen, P.; Trojovský, P. 2024. Overview and challenges of machine translation for contextually appropriate translations. iScience. Disponível em: <https://www.sciencedirect.com/science/article/pii/S2589004224021035>. Acesso em: 27 outubro 2025.
Ribeiro, G. C. B. 2005. Tradução e localização de software e outros produtos: audiovisual ou multimídia?. Cadernos de tradução. vol2. no16: 231-250. PUC-Rio, RJ, RJ.
TypeScript. n. d. What is TypeScript?. Disponível em: <https://www.typescriptlang.org/>. Acesso em: 18 abril 2026.
Unicode Consortium [UNICODE]. n.d. Technical Quick Start Guide. Disponível em: <https://home.unicode.org/technical-quick-start-guide/>. Acesso em: 30 setembro 2025.
Article originating from the Final Course Work of the Specialization in Software Engineering of the MBA USP/Esalq
To learn more about the course, click here and access the MBX Academy platform