Software Engineering
October 09, 2026
Use of voice in conjunction with large language models as a tool for digital accessibility
Using Voice Together with Large Language Models as a Tool for Digital Accessibility
Luiz Gustavo Fernandes Salvador; Ariel da Silva Dias
DOI: 10.22167/2675-6528-202603214
Article derived from a Course Conclusion Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.
Summary
Speech is an essential basis for human interaction, and for people with disabilities, it can represent the main form of communication with the external environment. Given the growing technological influence, voice command identification has emerged as a promising strategy for human-machine interaction. The work explored how voice, in conjunction with Large Language Models (LLMs), can be used efficiently, naturally, and accurately. For this purpose, various design patterns and the Python language were employed, aiming for greater extensibility. Gemini was used as the LLM provider, sending audio directly and leveraging its function-calling capability to interact with the device. A system was developed capable of understanding user intent and converting it into actions, whose differential was the computer vision capability based on screenshots and a mesh system for LLM guidance. Tests revealed a user intent comprehension rate of 91.81% and a success rate in execution of 75.45% with the “Flash-3” model (top p 0.5 and top k 5). The proposed system validated the premise that the integration of LLMs into voice interfaces increases the autonomy of users with motor disabilities, fulfilling the purpose of being a modern and effective Assistive Technology. However, questions were raised about the costs of AI and user data security, indicating the need for improvement.
Keywords: Function call; Human-computer interaction; Voice recognition; Computer vision.
1. Introduction
Communication constitutes one of the fundamental bases for the evolution of the human species, allowing not only the exchange of information essential for survival but also the development of social relationships and the facilitation of cooperation (Smith, 2010). Among the most expressive forms of communication, speech stands out for its ability to convey not only a message but also the speaker’s emotion (Loan et al., 2022). By extrapolating the human-to-human relationship, which propelled human society to its current evolutionary stage, it is possible to recognize that technology enables the exploration of voice as a means for human-machine interaction (Seaborn et al., 2021).
The technological advancement of recent decades is remarkable, directly impacting daily life, habits, and social relationships. This impact is evident in the significant increase in average time spent using computers and smartphones, compared to offline hobbies and activities (Vilhelmson et al., 2017). Many of these online interactions are linked to socialization and entertainment, with platforms like YouTube, Instagram, and Snapchat being among the most popular, especially among young people (Dienlin and Johannes, 2020). Additionally, remote education has experienced expressive growth, particularly during and after the COVID-19 pandemic, with Distance Learning (EaD) going from 18.4% of enrollments in 2011 to 62.8% a decade later (INEP, 2022). Such transformations highlight the need to develop more intuitive and integrating interfaces (Hott and Fraz, 2019), making the identification of voice commands a promising strategy for more natural and efficient human-machine interaction (Norda et al., 2024).
In Brazil, approximately 45 million inhabitants have some disability, with 28.8% facing motor difficulties (Biblioteca Virtual em Saúde Ministério da Saúde, n.d.). Individuals in these conditions may have limitations in accessing technology, which can hinder access to information and impact academic and professional life (Muhammad et al., 2015). In response to these issues, several companies seek to develop devices or software aimed at increasing the compatibility of people with disabilities in certain tasks, known as Assistive Technologies (Freitas et al., 2022). However, in some cases, such as patients with spinal atrophy, technologies that merely enhance motor capabilities may not be sufficient, requiring new strategies for the technological inclusion of these individuals (Lunn and Wang, 2008).
The attempt to use voice as data input is not recent, with examples dating back to 1990, such as Apple’s PlainTalk (Kumar, 2020). However, challenges such as the need for specific cadence, language, and terminology persist, reducing naturalness and hindering the learning of these technologies (Rogers et al., 2022). In this context, the use of Large Language Models (LLMs) is suggested as an intermediary to convey user intentions to the system, which would reduce the need for familiarity with the software and allow for more natural human-machine interaction (Mahmood et al., 2025).
Large Language Models (LLMs) are artificial intelligence (AI) models dedicated to interpreting natural language. Their capabilities range from maintaining conversations and analyzing data to interacting with other systems (Chang et al., 2018). Although they have gained prominence recently, research on statistical models that describe and predict natural language responses dates back to 1950, with Shannon (1948) investigating the ability of Word n-gram models to predict or compress language (Minaee et al., 2024). With the advent of the internet and the vast availability of data, it has become feasible to train these models intensively, resulting in greater versatility, accuracy, and effective integration with other systems (Kaddour et al., 2023).
Considering the factors presented, it becomes necessary to develop software capable of recognizing voice commands and using LLMs as tools for the human-machine interface, enabling the use of applications by people with limited mobility and promoting greater integration and accessibility in the digital environment. Thus, this research explored how this technology can be used efficiently, naturally, and accurately.
2. Material and Methods
The research was characterized as a software development study, with a methodological approach focused on systems engineering and the evaluation of its efficiency. The objective was to create a voice-based human-machine interaction system, using Large Language Models (LLMs) to promote digital accessibility for individuals with limited mobility. The development was carried out in the Python language, chosen for its versatility and vast availability of libraries (Udrake, 2023; Dhruv et al., 2021).
The software architecture employed several design patterns to ensure modularity and extensibility. The “Registry Factory” was used to instantiate and return registered classes (Koh, 2020), while “Strategy” defined class behaviors (Refactoring Guru, n.d.; Peace, 2024). Additionally, a “Singleton” container managed variable sharing (Stencel and Wegrzynowicz, 2008), “Decorator” defined common behaviors (Peng, Zhang and Hu, 2021), the “Command Pattern” allowed function execution by classes with a common interface (Nuzzi, 2019), and “Adapter” ensured data compatibility between LLMs (Nuzzi, 2019).
The system was divided into functional modules. The `config` module managed project settings, persisting information in `.ini` files via Python’s `configparser`. Settings included the input type, defined as “Audio”, the LLM provider, specified as “Gemini”, and the LLM model, “Gemini-flash-latest”. The graphical user interface (GUI) was developed with `PyQt6` (Willman, 2020), featuring a background, icon, and text fields for interaction, executed in a separate “Thread”.
The `genai` module was responsible for communication with Artificial Intelligence, using `google.genai` to adapt configurations, convert prompts, send data, and execute tools. Google Gemini was selected as the LLM provider due to its ability to handle diverse data types, detailed documentation, and optimized cost per “token” (Salvator, 2025). The LLM’s “function calling” functionality was essential, allowing the AI to interpret user intent and execute code snippets to interact with the device (Qu et al., 2025; Wang et al., 2026).
A Pydantic model was developed to configure LLMs in an interoperable way, covering parameters such as API key, model, system instructions, available tools, history maintenance, thinking level, verbose mode, maximum number of tokens, `top_p`, `top_k`, temperature, and seed. Additionally, LLM safety settings were configured, which included categories such as harassment, hate speech, sexually explicit and dangerous content, with a default value of “Medium and above”.
The input data collection was performed by the `input_source` module, configured for audio recording via the `SpeechRecognition` library (Amos, n.d.). The process involved microphone calibration for ambient noise and continuous recording until five seconds of user silence, at which point the audio file was saved. The `input_type` module converted audio and image files to formats compatible with the LLM. The tools (`tool`) were developed to perform functions on the device, including actions without direct screen interaction, such as asking the user, pressing keys, and typing text.
For screen interactions, a computer vision system was implemented, as the LLM did not provide exact coordinates. Two tools were created: “capture screen”, which performed a capture of the device’s current screen, as illustrated in the
, and drew an enumerated grid from 0.0 in the upper left corner to 14.14 in the lower right corner, as presented in the
; and “subdivide capture”, which selected a portion of the previous grid, scaled it, and redrew a 15×15 grid [FIGURE 3 and FIGURE 4]. This method allowed the AI to identify the location of objects by grid cells, enabling mouse interaction with various systems.
The algorithm’s execution flow started with obtaining user input via `input_source` (audio), processed by `genai` and converted into an `input_type` for the LLM. The LLM then processed the data and determined which tools (`tool`) would be called. This cycle repeated until the system interruption request, as illustrated in the
.
The software quality was ensured by unit tests, which evaluated individual components and functions of the code. The `Pytest` module (Pytest, n.d.) was used for the execution of these tests, contributing to the early detection of errors and the reduction of rework (Eisty et al., 2025; Pressman, 2011; Sommerville, 2013). The algorithm’s efficiency was measured by a task-based approach, defining real objectives for users (ABNT, 2002), such as sending messages on social networks, opening online videos, and writing paragraphs in text editing software. The choice of these tasks was based on popular activities among young people in virtual environments (Dienlin and Johannes, 2020) and on their relevance for the integration of students with disabilities in distance learning.
3. Results and Discussion
The implementation of the graphical user interface (GUI) resulted in a simplified and objective interaction environment, characterized by a transparent background. The initial interface design was conceived to be intuitive, featuring a central icon, a quiz text, a secondary text for details, and a main text to describe the ongoing action. This structure aims to provide the user with a clear and direct experience of the system’s state and actions, facilitating understanding and use, especially for individuals with limited mobility, as per the project’s central objective.
The developed system operates through a continuous flow that begins with obtaining a voice command from the user, captured by an audio input module. This audio is then processed and converted into a format understandable for the selected Large Language Model (LLM), in this case, Google Gemini. The LLM interprets the user’s intention and determines which tools, or “tools”, should be activated on the device. This cycle of input, processing, and tool execution repeats until the user decides to end the interaction, ensuring fluid and adaptable communication.
To allow artificial intelligence to interact with the computer screen, overcoming the limitation of LLMs in returning exact coordinates, specific tools were developed. The first, named “capture screen”, performs a capture of the device’s current screen and overlays an enumerated grid, starting at index 0,0 in the upper left corner and progressing to 14,14 in the lower right corner. The second tool, “subdivide capture”, allows selecting a tile of this grid, scaling it, and redrawing a new 15 by 15 grid. This approach enables the AI to identify the location of objects on the screen through tiles, instead of precise coordinates, allowing the mouse to move to the center of a specific tile and thus interact with various systems.
System functionality tests were performed based on real tasks that users should try to achieve, following the methodology of unit tests to ensure software quality. The tasks included sending a message on a social network, opening an online video, and writing a paragraph in a text editing software. The choice of these tasks reflected popular activities among young people in virtual environments, according to Dienlin and Johannes (2020), and the proposal to integrate students with disabilities into distance learning, aiming to validate the system’s effectiveness in practical use scenarios.
The test results indicated that, although most attempts were successful, there were instances where artificial intelligence required assistance to complete tasks, demanding additional information on the steps to follow. It was also observed that decreasing the “top p” parameter to 0.5 and “top k” to 5 resulted in greater accuracy in the system’s responses. This parameter optimization is crucial for refining the LLM’s ability to make more assertive decisions and reduce the need for human intervention.
The choice of Gemini’s “Flash-3” model for testing was motivated by usage limitations and the high cost associated with the “Pro-3” model. Furthermore, the screenshot quality was reduced by half of the original, without compromising the effectiveness of the AI’s actions. This strategic decision allowed for a significant number of tests (110 in total) to be conducted at a total cost of R$ 65.37, providing robust data on the system’s performance under controlled and economically viable conditions.
The detailed analysis of the tests, using the “Flash-3” model with `top k 5` and `top p 0.5`, revealed remarkable performance in several substeps. For the task of “Send social media message”, the “Open network” substep achieved nine correct answers and one error, while “Open contact” registered seven correct answers, three errors, and two hallucinations with error. The “Send message” substep reached nine correct answers and one error. In the task of “Open ‘online’ video”, the “Open browser” substep had six correct answers, four errors, and four hallucinations with error, and two hallucinations without error, indicating challenges in initial navigation.
Continuing the analysis of results for the “Flash-3” model with `top k 5` and `top p 0.5`, the “Open video network” sub-step in the “Open ‘online’ video” task obtained ten correct answers and two hallucinations without error. “Search video” and “Open video” registered seven correct answers and three errors each, with three hallucinations with error for “Search video” and one for “Open video”. For the “Write paragraph in editing ‘software'” task, the “Open ‘software'” and “Open new document” sub-steps had eight correct answers and two errors each, with one hallucination with error in both cases. The “Write text” sub-step achieved seven correct answers and three errors, while “Save document” presented five correct answers and five errors, with four hallucinations with error and one without error, evidencing the complexity of interaction with file saving.
In contrast, tests conducted with the “Flash-3” model configured with `top k 15` and `top p 0.85` showed performance variations. In the “Send social media message” task, the “Open network” sub-step achieved six correct answers and four errors, with three hallucinations with errors and four without errors. “Open contact” maintained ten correct answers, with no errors or hallucinations. The “Send message” sub-step registered eight correct answers and two errors, with two hallucinations with errors. This data suggests that altering LLM parameters directly influences the accuracy and the occurrence of hallucinations, impacting interaction robustness.
Continuing with the evaluation of the “Flash-3” model with `top k 15` and `top p 0.85`, the task “Open video ‘online'” showed that the “Open browser” sub-step obtained seven correct answers and three errors, with two hallucinations with error. “Open video network” registered six correct answers and four errors, with three hallucinations with error and one without error. “Search video” also had six correct answers and four errors, with three hallucinations with error and two without error. The “Open video” sub-step presented only three correct answers and seven errors, with five hallucinations with error, indicating inferior performance for these parameters in video navigation and search scenarios.
In the task of “Drafting paragraph in editing ‘software'” with the “Flash-3” model (`top k 15` and `top p 0.85`), the “Open ‘software'” sub-step achieved eight correct answers and two errors, with four hallucinations without error. “Open new document” registered nine correct answers and one error, with one hallucination with error and two without error. “Write text” reached nine correct answers and one error, with one hallucination with error and one without error. Finally, “Save document” had six correct answers and four errors, with four hallucinations with error, reinforcing the persistent difficulty in the saving stage, regardless of the specific LLM parameters.
The user intent comprehension rate was calculated by summing correct answers and hallucinations with errors, and dividing by the total number of tests. For the “Flash-3” model with `top p 0.5` and `top k 5`, the user intent comprehension rate reached 91.81%, with an execution success rate of 75.45%. With parameters `top p 0.85` and `top k 15`, the comprehension rate was slightly higher, 92.72%, but the execution success rate was lower, 70.90%. These results demonstrate the sensitivity of the system’s performance to LLM parameters, indicating that higher intent comprehension does not always translate into greater success in executing actions.
The inclusion of generative artificial intelligences in voice interfaces, as per the proposed solution, grants greater user autonomy, eliminating the need to memorize specific keywords to activate the system, an advancement compared to previous approaches (Avuçlu et al., 2023). The system demonstrated independence from hardware quality and user position, requiring only an internet connection, making it a flexible and accessible solution (Taher et al., 2021). Furthermore, the work eliminates the prior artificial intelligence training step, promoting a more direct and accessible interaction, unlike research that demands this phase (Mulfari et al., 2021).
However, some important considerations were raised. User data security is a critical point, as information is transmitted to a corporate LLM, which may allow its use for AI training, diverging from solutions that employ local models and do not expose user data (Livero and Santos, 2024). The high rate of hallucinations in task execution, in contrast to the data from Mullfari et al. (2021), was mainly attributed to the interaction via direct mouse control, which is more susceptible to errors. Additionally, the high cost of system use and the latency in LLM processing represent challenges that need to be considered for future improvements.
In summary, the work validated the premise that the integration of Large Language Models into voice interfaces represents an effective tool for increasing the autonomy of users with motor disabilities, fulfilling the purpose of being a modern and functional Assistive Technology. The modular architecture and specific tools, such as the grid logic for screen interaction, were crucial for overcoming the limitations of LLMs in handling exact coordinates. Despite challenges related to cost, processing time, and data security, the results demonstrate a high rate of user intention comprehension and considerable success in task execution, indicating the solution’s potential and the need for continuous optimizations.
4. Conclusion
This research explored the use of voice in conjunction with Large Language Models (LLMs) as an efficient, natural, and accurate tool for human-machine interaction, aiming for digital accessibility for people with limited mobility. The development of a system with a simplified graphical interface and a continuous interaction flow was verified, where voice commands were processed by LLMs to trigger tools on the device. To overcome the limitation of LLMs in handling exact coordinates, a computer vision system based on screenshots and an enumerated grid was implemented, allowing artificial intelligence to identify and interact with objects on the screen through tiles. Functionality tests, conducted with real tasks, indicated a user intent comprehension rate of 91.81% and a task execution success rate of 75.45% with the “Flash-3” model (top p 0.5 and top k 5). This approach validated the premise that the integration of LLMs into voice interfaces significantly increases the autonomy of users with motor disabilities, establishing itself as a modern and effective Assistive Technology.
However, some important limitations were identified that deserve attention in future studies. The security of user data, transmitted to a corporate LLM, raised concerns about the use of this information for AI training. A high rate of hallucinations was also observed in task execution, mainly attributed to interaction via direct mouse control, which proved susceptible to errors. Additionally, the high cost of system use and latency in LLM processing represent challenges to be considered. Therefore, future improvements are suggested, including the exploration of local models for greater data privacy, the integration of new tools to mitigate hallucinations, and continuous optimizations to reduce costs and processing time, aiming to enhance the robustness and viability of the solution.
Bibliographic References
Dienlin, T.; Johannes, N. 2020. The impact of digital technology use on adolescent well-being. Dialogues in Clinical Neuroscience 22(2): 135-142.
Instituto Nacional de Estudos e Pesquisas Educacionais Anísio Teixeira [INEP]. 2
Seaborn, K.; Miyake, N.P.; Pennefather, P.; Otake-Matsuura, M. 2021. Voice in Human-Agent Interaction: A Survey. ACM Computing Surveys 54: 1-43.
Smith, E.A. 2010. Communication and collective action: language and the evolution of human cooperation. Evolution and human behavior 31(4): 231-245.
Van Loan T.; Le, T.D.T.; Xuan, T.L.; Castelli, E. 2022. Emotional speech recognition using deep neural networks. Sensors 22(4): 1414.
Vilhelmson, B.; Elldér, E.; Thulin, E. 2018. What did we do when the Internet wasn’t around? Variation in free-time activities among three young-adult cohorts from 1990/1991, 2000/2001, and 2010/2011. New Media & Society 20(8): 2898-2916.
Article originating from the Final Course Work of the Specialization in Software Engineering of the MBA USP/Esalq
To learn more about the course, click here and access the MBX Academy platform