Article

October 07, 2026

Prediction of ceftazidime-avibactam resistance and carbapenem susceptibility in KPC variants

Prediction of Ceftazidime-avibactam Resistance and Carbapenem Susceptibility in Kpc Variants

Jesus Giovani Mamani Pariona; Henrique Raymundo Gioia

DOI: 10.22167/2675-6528-202603032

Article derived from a Final Course Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.

Summary

The growing global dissemination of KPC enzyme variants associated with ceftazidime-avibactam (CZA) resistance has posed a critical threat to antimicrobial therapeutics, especially given the potential restoration of susceptibility to carbapenems. This study aimed to develop and validate predictive models based on machine learning capable of inferring, from mutations in the blaKPC gene, the phenotypes of CZA resistance and carbapenem susceptibility. To this end, a database of 281 variants was constructed, of which 162 had phenotypic validation and were used as the training set. A complete genomic approach was employed, including substitutions, insertions, and deletions, with One-Hot encoding and supervised modeling using algorithms such as Random Forest, Support Vector Machines (SVM), Logistic Regression, and k-Nearest Neighbors (KNN), validated by cross-validation and balancing with SMOTE. The results demonstrated high predictive performance for the SVM model, with an F1-score exceeding 96% for CZA resistance and approximately 82% for carbapenems, highlighting the relevance of structural events (indels) in enzymatic function. A consistent pattern of evolutionary trade-off was observed, where CZA resistance was frequently associated with a loss of activity against carbapenems. Molecular markers such as D179Y and insertions in the Omega loop were not identified as critical determinants, being statistically validated by Fisher’s exact test with Bonferroni correction. It was concluded that machine learning enabled the prediction of complex phenotypes with high accuracy, offering a promising tool for genomic surveillance and clinical decision support in microbiology.

Keywords: machine learning; blaKPC; structural mutations; antimicrobial resistance; genomic surveillance.

1. Introduction

The global dissemination of Klebsiella pneumoniae isolates carrying novel variants of the blaKPC gene, which confer resistance to ceftazidime-avibactam (CZA), represents a growing threat to antimicrobial therapy and global public health (Ding et al., 2023).

Since its introduction, CZA has been one of the few beta-lactam/beta-lactamase inhibitor combinations capable of effectively treating infections caused by carbapenemase producers of the K. pneumoniae (KPC) type. However, the increasing number of reports of CZA-resistant KPC variants highlights the rapid adaptive potential of this pathogen under antibiotic pressure (Niu et al., 2020).

To date, more than 250 clinical variants of KPC have been documented, of which 161 have been phenotypically characterized as CZA-resistant (Ding et al., 2023; Niu et al., 2020). Notably, approximately 107 (58.5%) of these CZA-resistant KPC variants exhibit activity similar to extended-spectrum beta-lactamases (ESBLs), restoring susceptibility to carbapenems and suggesting the possible reintroduction of these agents as therapeutic alternatives (Taracila et al., 2022).

However, many KPC variants remain uncharacterized or exhibit unpredictable susceptibility patterns, which hinders clinical decision-making and surveillance efforts. Previous studies have demonstrated the in vitro and in vivo efficacy of carbapenems against producers of KPC variants exhibiting phenotypic resistance to CZA but restored susceptibility to carbapenems (Pariona et al., 2025). Successful therapeutic outcomes have been reported in patients infected with such variants after treatment with carbapenems (Giddins et al., 2018; Mueller et al., 2019; Oliva et al., 2023; Shields et al., 2017; Vásquez-Ponce et al., 2023).

In contrast, resistance to CZCA, driven by specific mutations in the KPC enzyme, has been associated with therapeutic failure and increased mortality rates (Ding et al., 2022; Nicola et al., 2022; Shi et al., 2022; Sun et al., 2020). This complexity highlights an evolutionary trade-off phenomenon, where the acquisition of resistance to new inhibitors often occurs at the expense of the original carbapenemase activity, restoring susceptibility to this class of antibiotics (Arcari et al., 2022; Barnes et al., 2017; Ding et al., 2022; Giddins et al., 2018; Kim et al., 2022; Mueller et al., 2019; Nicola et al., 2022; Oliva et al., 2023; Pariona et al., 2025; Shi et al., 2022; Shields et al., 2017; Sun et al., 2020; Taracila et al., 2022; Vásquez-Ponce et al., 2023).

Given this complexity, predicting the phenotype of resistance to CZA and carbapenems, based on mutations in the blaKPC gene, represents a critical need for antimicrobial resistance surveillance and infection control. Machine learning (ML) offers a powerful, data-driven approach to modeling the relationship between genetic variation and antimicrobial resistance phenotypes (Kim et al., 2022).

Thus, improving the understanding of resistance mechanisms, enhancing early detection of emerging variants, and supporting precision-guided antimicrobial therapy in clinical microbiology are fundamental aspects in combating the threat of superbugs. In this study, we aim to develop a robust machine learning-based predictive model capable of classifying CZA and carbapenem resistance in KPC variant producers based on their amino acid mutations.

2. Material and Methods

The present study was characterized as a research of a computational nature, with a quantitative and exploratory methodological approach, focused on the development of predictive models. The objective was to infer the phenotypes of resistance to ceftazidime-avibactam (CZA) and susceptibility to carbapenems in variants of the KPC enzyme, based on their mutations in the blaKPC gene. The research was conducted entirely through the analysis of publicly available secondary data.

The genomic database was constructed from the retrieval of variant protein sequences of Klebsiella pneumoniae carbapenemase (KPC) from the National Center for Biotechnology Information (NCBI) Pathogen Detection Reference Gene Database, accessed on January 25, 2026. Only complete, high-quality protein sequences were selected, covering clinical variants from KPC-2 to KPC-287. Manual curation was performed, cross-referencing each variant with validated experimental data from the scientific literature and the Beta-Lactamase Data Resources (BLDB) repository, focusing on the prediction of CZA resistance and susceptibility to carbapenems.

For the supervised training of the algorithms, a binary reclassification protocol based on reported phenotypic behavior was established. For the predictive model of CZ A resistance, variants described as “Inhibitor-Resistant” were classified as “Resistant”, and those with a classic carbapenemase profile as “Susceptible”. For the carbapenem susceptibility model, variants with loss of hydrolytic activity were categorized as “Susceptible”, and those that retained hydrolysis capacity as “Resistant”. Variants with inconclusive or undocumented phenotypic profiles were segregated for a later stage of blind prediction.

Feature engineering utilized a whole-genome approach to convert biological information into processable numerical vectors. Pairwise global alignment of the collected sequences was performed against the KPC-2 variant, used as the wild-type reference, employing the Needleman-Wunsch algorithm, implemented in the Biopython library. The BLOSUM62 substitution matrix and adjusted gap penalty parameters (gap opening = -10; extension = -0.5) were used. A dynamic coordinate mapping algorithm extracted amino acid substitutions at conserved positions, deletion events, and insertion events. The final data transformation employed the One-Hot Encoding technique, generating a high-dimensional sparse matrix. The KPC-2 variant was weighted in the training set as an archetype of CZA sensitivity and carbapenem resistance, calibrating the model intercepts.

The modeling strategy was based on a supervised training and validation scheme, using the dataset with validated phenotyping (Ground Truth). The performance evaluation of all architectures was conducted via 5-fold Stratified Cross-Validation. Hyperparameter selection was performed by Grid Search. The treatment of class imbalance was adjusted to the nature of each algorithm, with `class_weight=’balanced’` for Logistic Regression (associated with L1 and L2 regularization techniques) and synthetic oversampling (SMOTE) for Random Forest, Support Vector Machines (SVM), and k-Nearest Neighbors (KNN).

Four machine learning algorithms were comparatively evaluated: Random Forest, Support Vector Machines (SVM), Logistic Regression, and k-Nearest Neighbors (KNN). The choice of the final models for each clinical outcome (resistance to CZA and susceptibility to carbapenems) was determined by the average F1-Score obtained in cross-validation, aiming to ensure a balance between precision and sensitivity (recall).

All data processing, sequence alignment, and predictive modeling were conducted in the Python programming language (version 3.10+), using the Google Colab development environment. The analyses were based on the Scikit-learn library for machine learning algorithms, Biopython for biological sequence manipulation, Pandas for data structuring, and Imbalanced-learn for class balancing techniques. To ensure strict reproducibility of experiments and synthetic sampling (SMOTE), a random seed of 42 was fixed in all stochastic steps.

The statistical validation of the results was based on the analysis of the mean and standard deviation of the performance metrics (Accuracy, Precision, Recall, and F1-Score) obtained through the 5 folds of cross-validation. To statistically validate the associations between molecular markers (point mutations and indels) and clinical outcomes of resistance, a univariate mutational enrichment study was conducted. The Fisher’s Exact Test was used to calculate the odds ratio (Odds Ratio – OR) and statistical significance (P-value), with Bonferroni correction to adjust P-values and control the false positive rate (Family-Wise Error Rate). Only markers with an adjusted P-value < 0.05 were considered statistically significant. The analyses were processed using the SciPy library (v.1.10) in a Python environment.

The research was conducted exclusively through computational analysis of public data (*in silico* study), not directly involving human beings, animals, or environmental interventions. Therefore, submission to an ethics committee was not necessary. The study was not associated with field data collection, being characterized as an analysis of publicly accessible secondary data. The sequences and phenotypic data used are publicly available in the NCBI Pathogen Detection Reference Gene Database. The scripts used for preprocessing, model training, and interpretation of variables will be made available upon reasonable request to the corresponding author.

3. Results and Discussion

The present study consolidated a comprehensive genomic database, totaling 281 distinct variants of the Klebsiella pneumoniae carbapenemase (KPC) family, retrieved and processed from public repositories. From this total universe, the rigorous manual curation process allowed the stratification of samples into two fundamental groups based on the availability of clinical evidence. A consolidated set of 162 variants, representing 57.7% of the total, had an experimentally known and literature-validated phenotypic profile, constituting the gold standard for training predictive models.

Additionally, a significant contingent of 119 variants, corresponding to 42.3%, whose phenotypes remained unknown, ambiguous, or uncharacterized, was segregated to compose the risk prediction set. This scenario reflects the global gap in the phenotypic characterization of emerging variants, as described by Ding et al. (2023) and Vásquez-Ponce et al. (2023). The descriptive analysis of frequency distribution in the validated set (n=162) revealed distinct patterns of selective pressure for each antibiotic class evaluated, highlighting the complexity of resistance evolution.

Regarding the clinical outcome of resistance to the combination ceftazidime-avibactam (CZA), a massive predominance of resistant variants was observed in the current global scenario. Of the 162 validated variants, 141 samples, or approximately 87.4% of the training set, presented the CZA resistance phenotype. This finding corroborates evidence that the selective pressure exerted by new inhibitors rapidly favors adapted variants, as pointed out by Giddins et al. (2018) and Niu et al. (2020).

In contrast, the group of variants that maintained sensitivity to this inhibitor represented a restricted minority class, accounting for only 21 variants, or 13% of the total. This severe imbalance between classes highlighted resistance to CZA as the dominant phenotype in the recent evolution of the KPC family, requiring the application of data balancing techniques for training machine learning models, in order to avoid biases and ensure the robustness of predictions.

Simultaneously, the characterization of the carbapenem susceptibility profile revealed a more heterogeneous and balanced epidemiological scenario. For this outcome, it was identified that the majority of the analyzed variants evolved towards a loss of hydrolytic function against carbapenems; specifically, 109 variants were classified as susceptible to this class of antibiotics, representing 67.3% of the dataset. This finding aligns with the concept of fitness cost associated with adaptation to inhibitors, a phenomenon widely discussed in the literature.

On the other hand, classic carbapenem resistance was conserved in 46 variants, corresponding to 28.4% of the dataset. Unlike the scenario observed for CZA, the more balanced distribution in this outcome suggests an evolutionary dynamic where the acquisition of new inhibitor resistance mechanisms often occurs at the expense of original carbapenemase activity. This phenomenon is widely described as an evolutionary trade-off in beta-lactamases (Arcari et al., 2022; Barnes et al., 2017; Ding et al., 2022; Giddins et al., 2018; Kim et al., 2022; Mueller et al., 2019; Nicola et al., 2022; Oliva et al., 2023; Pariona et al., 2025; Shi et al., 2022; Shields et al., 2017; Sun et al., 2020; Taracila et al., 2022; Vásquez-Ponce et al., 2023).

Performance of predictive models and validation

The computational benchmarking experiment, conducted under 5-Fold Stratified Cross-Validation and exhaustive hyperparameter optimization (via GridSearchCV), demonstrated that the transition to a complete genomic representation, including substitutions, insertions, and deletions (indels), was decisive for achieving high performance metrics in both phenotypic outcomes. The inclusion of these structural events allowed the models to capture the complexity of alterations in the KPC enzyme.

For the prediction of ceftazidime-avibactam (CZA) resistance, the Support Vector Machine (SVM) algorithm with RBF kernel has consolidated itself as the architecture with the best overall performance. This model achieved an average F1-Score of 95.90% ± 1.75%, with an area under the ROC curve (ROC-AUC) of 92.04% ± 6.28%. Logistic Regression, adjusted with L1 regularization (Lasso, C = 1) and balancing via class_weight=’balanced’, presented a highly competitive performance, achieving an F1-Score of 94.55% ± 1.67% and ROC-AUC of 73.83% ± 9.70%.

The superiority of SVM and Logistic Regression was evident compared to decision tree-based models, such as Random Forest, which achieved an F1-Score of 91.35% ± 3.08%, and k-Nearest Neighbors (KNN), with an F1-Score of 83.67% ± 6.60%. These findings confirm that CZ A resistance is structured under complex decision boundaries, where the SVM’s RBF kernel offered the best spatial delimitation capability, indicating its robustness in identifying resistance patterns.

For the outcome of carbapenem susceptibility, which represents a superior biological challenge due to the functional trade-off between carbapenemase activity loss and avibactam resistance gain, the SVM (RBF) algorithm again emerged as the best architecture. It achieved an F1-Score of 82.71% ± 4.69% and ROC-AUC of 79.56% ± 5.15%, demonstrating its effectiveness in predicting complex phenotypes.

Logistic Regression, optimized with L2 regularization (C = 50), achieved the second best average F1-Score of 79.50% ± 6.15%, remaining equivalent to Random Forest, which registered an F1-Score of 79.41% ± 6.65%. Both significantly outperformed KNN, which presented an F1-Score of 68.67% ± 8.19%. Notably, in the carbapenem outcome, Logistic Regression achieved the highest global discrimination capacity among all architectures, registering a ROC-AUC of 81.73% ± 5.62%, surpassing SVM (79.56%), Random Forest (77.44%), and KNN (72.96%).

Phenotypic prediction, inference reliability in uncharacterized variants, and attribute importance analysis

The application of the optimized models to KPC variants without reported phenotype demonstrated a robust pattern of reliability in inference, with most predictions falling within high-confidence probability intervals. For the ceftazidime-avibactam (CZA) model, the inference showed exceptional stability, with 77.1% (91 variants) reaching the High Confidence threshold (class probability >80%).

Additionally, 22.0% (26 variants) were in the Medium Confidence range (60-80%), and the gray or indeterminate zone (80%), 36.4% (43 variants) presented Medium Confidence (60-80%) and 13.6% (16 variants) remained in the gray zone (1.30). The markers that showed the greatest association tendency were the Pos_103_P and Pos_103_R substitutions for the ceftazidime-avibactam (CZA) resistance outcome, and Pos_178_D for carbapenem susceptibility.

However, even these candidates remained above the statistical significance threshold after correction for multiple comparisons (p_adj > 0.05). This pattern was expected considering the high allelic diversity of the analyzed database. Most KPC variants showed infrequent molecular alterations, with low individual frequency of each substitution or indel, reducing the statistical power to identify isolated associations between a single marker and the phenotype.

Thus, although some positions showed association trends, the heterogeneous distribution of alterations prevented any individual effect from reaching statistical significance after multiple adjustment. This limitation of univariate tests reinforces the need for multivariate approaches, such as supervised machine learning models, which can simultaneously integrate multiple alterations in the protein sequence and capture combinatorial patterns and possible epistatic interactions between different residues of the KPC enzyme, providing a more complete view of resistance mechanisms.

In summary, the results of this study demonstrate that machine learning, particularly the Support Vector Machine algorithm, is highly effective in predicting ceftazidime-avibactam resistance and carbapenem susceptibility phenotypes in KPC variants, based on mutations in the blaKPC gene. Feature importance analysis revealed that structural events, such as insertions and deletions in critical enzyme hotspots, are determinants of enzymatic function and the observed evolutionary trade-off, where resistance to new inhibitors is often associated with the restoration of carbapenem susceptibility. This work offers a promising tool for genomic surveillance and clinical decision support, despite the limitations of univariate statistical tests in identifying isolated markers due to high allelic diversity.

4. Conclusion

The study aimed to develop and validate predictive models based on machine learning capable of inferring, from mutations in the blaKPC gene, the phenotypes of resistance to ceftazidime-avibactam (CZA) and susceptibility to carbapenems. It was found that the Support Vector Machine (SVM) algorithm demonstrated high predictive performance, with an F1-score above 95% for CZA resistance and approximately 82% for carbapenem susceptibility. The inclusion of a complete genomic approach, encompassing substitutions, insertions, and deletions (indels), was crucial for the models’ ability to capture the complexity of alterations in the KPC enzyme. A consistent pattern of evolutionary trade-off was observed, in which CZA resistance was frequently associated with a loss of activity against carbapenems, restoring susceptibility to this class of antibiotics. Attribute importance analysis highlighted structural hotspots, such as the Omega loop and the 267–275 region, as determinants for enzymatic function and for this trade-off phenomenon.

The application of optimized models allowed for reliable prediction of phenotypes in uncharacterized KPC variants, identifying a significant group of 50 variants with a high-confidence “Double Profile”, concurrently resistant to CZA and susceptible to carbapenems. This finding validates the biological hypothesis of the structural fitness cost imposed by adaptation to new inhibitors. Although univariate statistical tests did not identify significant individual markers due to high allelic diversity, the multivariate machine learning approach overcame this limitation by integrating multiple protein sequence alterations. Stratification by probability levels of predictions offers an effective strategy for data screening, prioritizing variants in the gray zone for in vitro phenotypic validation. Thus, machine learning offers a promising tool for genomic surveillance and clinical decision support in microbiology, enhancing the understanding of resistance mechanisms and the early detection of emerging variants.

Bibliographic References

Arcari et al., 2022 [Referência completa não encontrada no documento original]

Barnes, M.D.; Winkler, M.L.; Taracila, M.A.; Page, M.G.; Desarbre, E.; Kreiswirth, B.N.; Shields, R.K.; Nguyen, M.H.; Clancy, C.; Spellberg, B.; Papp-Wallace, K.M.; Bonomo, R.A. 2017. Klebsiella pneumoniae carbapenemase-2 (KPC-2), substitutions at Ambler position Asp179, and resistance to ceftazidime-avibactam: unique antibiotic-resistant phenotypes emerge from ẞ-lactamase protein engineering. mBio 8(5): e00528-17.

Ding, L.; Shen, S.; Chen, J.; Tian, Z.; Shi, Q.; Han, R.; Guo, Y.; Hu, F. 2023. Klebsiella pneumoniae carbapenemase variants: the new threat to global public health. Clinical Microbiology Reviews 36(4): e00008-

Ding, L.; Shen, S.; Han, R.; Yin, D.; Guo, Y.; Hu, F. 2022. Ceftazidime-avibactam in combination with imipenem as salvage therapy for ST11 KPC-33-producing Klebsiella pneumoniae. Antibiotics 11(5): 604.

Article originating from the Final Course Work of the Specialization in Data Science and Analytics of the MBA USP/Esalq

To learn more about the course, click here and access the MBX Academy platform

You may also like

October 07, 2026

Governance in processes: reduction of rework, time, and cost in accounting closing

Corporate governance, applied to the management of organizational processes, has been associated with improved control, compliance, and operational efficiency. The study investigated how corporate governance mechanisms, combined with automation, influenced operational efficiency and internal compliance in critical monthly closing processes in an agribusiness company. An exploratory and descriptive case study was conducted, with a mixed approach, combining participant observation, documentary analysis, and analysis of operational records from systems and internal controls. The unit of analysis comprised critical processes, selected based on volume, operational risk, and relevance to the accounting closing. The results indicated that the application of structural governance mechanisms, such as defining authorities, segregation of duties, standardization of routines, and implementation of structured workflows, was associated with reduced execution time, decreased rework, and greater traceability of activities. It was also observed that the reorganization of processes contributed to reducing the need for overtime, indicating potential for reducing operational costs associated with monthly closing. The findings demonstrated that governance, when incorporated into the design and execution of processes, acts as a structuring element of operations, promoting alignment between control, execution, and organizational performance.

Keywords: Automation; Internal control; Operational cost; Operational efficiency; Organizational processes.

October 07, 2026

Regulatory Governance of Chemical Substances in the Context of the EU-Mercosur Agreement: Challenges and Opportunities for Brazil

The regulatory governance of chemical substances has undergone profound transformations, driven by the European REACH model and the growing relevance of the ESG (Environmental, Social, and Governance) agenda. The impacts of the implementation of Law No. 15.022/2024 (Brazil REACH) and the dynamics of the EU-Mercosur Association Agreement on the competitiveness of the Brazilian chemical industry were analyzed. The research adopted a qualitative approach, based on documentary analysis and a semi-structured interview with an industry expert. The results showed that, although Brazil REACH represents an advance in regulatory convergence with international standards, its effectiveness was limited by gaps in institutional capacity and by the operational uncertainty of regulatory agencies. It was found that recent political uncertainties and debates in the European Parliament increased perceived risk, but did not interrupt the industrial adaptation process, which occurred proactively to meet the demands for access to global markets and sustainable financing. It was observed that regulatory compliance for chemical substances has come to be perceived not only as a cost center, but as a strategic governance asset, essential for the resilience of companies in global value chains.

Keywords: Technical barriers to trade; Institutional capacity; Regulatory convergence; Socio-environmental governance; Chemical substance regulation.

October 07, 2026

Best practices for strengthening trust in the ethics channel

The ethics channel has consolidated itself as a relevant mechanism for identifying irregularities and misconduct within the organizational scope. However, despite its incentive and regulatory obligation, resistance to its use by stakeholders was observed. The study aimed to examine the main difficulties related to the implementation and operation of the channel, as well as to identify measures capable of mitigating these barriers and strengthening its credibility among internal and external audiences. The research was developed in two complementary stages: a theoretical phase, which analyzed specialized doctrine and applicable Brazilian legislation, and an empirical phase, which consisted of conducting semi-structured interviews with Compliance professionals. The results allowed for the identification of essential requirements for the proper functioning of the channel and highlighted best practices, such as clear institutional communication, continuous training, engagement of senior management, structured treatment flow, safeguards for the whistleblower, careful disclosure of results, and permanent monitoring. It was found that the most recurrent weaknesses stemmed from predominantly communicational factors, which can be mitigated by consistent and segmented acculturation strategies. It was concluded that trust in the ethics channel is structuring for its effectiveness, and the identified vulnerabilities can be substantially mitigated by strategic communication, contributing to a culture of integrity that values transparency and the technical treatment of irregularities.

Keywords: Credibility; Reporting; Integrity; Stakeholders; Whistleblower.

October 07, 2026

Predictive modeling for prioritizing fiscal inconsistencies from the ECD trial balance

The growing complexity of economic relations and the high volume of data available to the tax administration demanded the development of more efficient tools for selecting taxpayers subject to tax audits. In this context, supervised machine learning models were applied, based on data extracted from Digital Accounting Records (ECD), to prioritize companies in a scenario of limited human resources and incomplete labeling. To this end, a dataset composed of accounting variables derived from the ECD was constructed, and the Logistic Regression and Random Forest models were applied. Logistic Regression was used as a linear reference, while Random Forest was chosen for its ability to capture non-linear relationships and complex interactions between structured variables. More complex models, such as neural networks, were not prioritized, considering the low availability of positive class examples and the objective of adopting a parsimonious and easily operationally applicable approach. The performance of the models was evaluated with a focus on their ability to concentrate positive cases in the top positions of the ranking. The results indicated the superiority of tree-based models, especially Random Forest, in prioritizing taxpayers. The identification of relevant patterns was also observed even in the context of imperfect labels. It was concluded that the proposed approach has the potential to support the selection of taxpayers for audit and contribute to the more efficient allocation of public resources.

Keywords: Risk analysis; Supervised learning; Accounting data; Audit prioritization; Taxpayer selection.

October 07, 2026

Cultural challenges in the implementation of the budget in a family-owned technology company

The implementation of the budgetary process in a medium-sized family business in the technology sector, transitioning from an informal management model to more structured planning and control practices, was addressed. The objective was to analyze the cultural challenges faced in this process, with an emphasis on perceptions, adaptations, and organizational transformations. A qualitative and descriptive research was conducted through a single case study in a company in Santa Catarina with approximately 600 employees. Data collection included semi-structured interviews with 12 participants (directors, heads, managers, and analysts) and document analysis. The results showed that the budget was recognized as a relevant tool to increase financial clarity, guide decisions, and strengthen strategic planning, promoting greater predictability and internal organization. However, significant cultural challenges were identified, such as limited financial knowledge among managers, resistance to loss of autonomy, and the need to adapt to new routines. The process drove the formalization of information, accountability of leadership, and the gradual replacement of intuitive decisions with data, also acting as a pedagogical tool for organizational learning and the development of financial skills. It was concluded that the effectiveness of the budgetary process is directly related to the organization’s ability to promote cultural changes, leadership engagement, and continuous evolution of management practices, contributing to the professionalization of management in family businesses.

Keywords: Control; Decision; Governance; Planning; Professionalization.

School

October 07, 2026

Business Plan: Easy Spanish

This study developed a business plan for an innovative model of teaching Spanish as a foreign language in Brazil. The objective was to create a flexible and personalized proposal that promoted student autonomy and academic rigor. For this, a multiple case study with a qualitative approach was used, applying unilateral benchmarking of the companies Open English, SMART Academia de Idiomas, and Uber. The methodology followed the guidelines of the Benchmarking Manual. In the initial phase, a Minimum Viable Product (MVP) was proposed with low-cost digital tools for scheduling, communication, and pedagogical management. The results indicated that the project’s viability lies in the combination of proprietary teaching materials, methodological standardization, after-sales support, and strategic use of technology. The model was configured as an intermediary between teachers and students, inspired by Uber’s operational logic, and incorporated the pedagogical rigor of SMART Academia de Idiomas, with level-based classes and continuous assessments. It was concluded that the proposal presents initial viability and differentiation potential, articulating operational flexibility, diversity of teaching profiles, and student autonomy. Implementation via MVP proved adequate for testing acceptance and reducing risks, with potential for international scalability and future search for investors.

Keywords: Benchmarking; Teaching Spanish; Educational innovation; Digital platform; Business plan.

October 07, 2026

Unsupervised Detection of Behavioral Anomalies in Corporate Access to Sensitive Data

The digitalization of corporate processes has increased exposure to internal threats, where employees with legitimate access to systems violate security policies. Based on behavioral theories of fraud, the proactive identification of these threats was investigated. The objective was to propose and evaluate a system for anomaly detection structured in Social Network Analysis (SNA), comparing the effectiveness of the unsupervised algorithms Isolation Forest (IF) and Local Outlier Factor (LOF) in identifying atypical access patterns in a simulated stochastic database. For this purpose, a dataset of approximately 1 million logs, stochastically generated, was used, which modeled probabilistic profiles with overlapping noise and legitimate accesses. The methodology involved modeling employee-client interactions as a weighted bipartite graph by risk, and IF and LOF were applied to topological network metrics and forensic risk scores. The results demonstrated the superiority of the LOF model, which, by focusing on local density, broke the camouflage of anomalies and achieved an average Recall of 81.3% in multiple simulations, surpassing the performance of Isolation Forest (average Recall of 56%). The solution proved to be a viable and scalable approach to reinforce ‘detection perception’, providing managers with a high-risk access base to prioritize precise investigations, in strict compliance with the LGPD.

Keywords: Social Network Analysis; Anomaly detection; Internal fraud; LGPD; Machine Learning.