October 07, 2026
Unsupervised Detection of Behavioral Anomalies in Corporate Access to Sensitive Data
Unsupervised Detection of Behavioral Anomalies in Corporate Access to Sensitive Data
Jeferson Belinello da Silveira; Christian Duarte Caldeira
DOI: 10.22167/2675-6528-202603031
Article derived from a Final Course Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.
Summary
The digitalization of corporate processes has increased exposure to internal threats, where employees with legitimate access to systems violate security policies. Based on behavioral theories of fraud, the proactive identification of these threats was investigated. The objective was to propose and evaluate a system for anomaly detection structured in Social Network Analysis (SNA), comparing the effectiveness of the unsupervised algorithms Isolation Forest (IF) and Local Outlier Factor (LOF) in identifying atypical access patterns in a simulated stochastic database. For this purpose, a dataset of approximately 1 million logs, stochastically generated, was used, which modeled probabilistic profiles with overlapping noise and legitimate accesses. The methodology involved modeling employee-client interactions as a weighted bipartite graph by risk, and IF and LOF were applied to topological network metrics and forensic risk scores. The results demonstrated the superiority of the LOF model, which, by focusing on local density, broke the camouflage of anomalies and achieved an average Recall of 81.3% in multiple simulations, surpassing the performance of Isolation Forest (average Recall of 56%). The solution proved to be a viable and scalable approach to reinforce ‘detection perception’, providing managers with a high-risk access base to prioritize precise investigations, in strict compliance with the LGPD.
Keywords: Social Network Analysis; Anomaly detection; Internal fraud; LGPD; Machine Learning.
1. Introduction
The digitalization of corporate processes, although it has boosted operational efficiency, has significantly expanded organizations’ exposure to internal threats. These threats are characterized by employees who, holding legitimate access to systems, violate security policies. The Verizon Data Breach Investigations Report (2023) highlights that approximately 74% of all security incidents involve the human element. Unlike external cyberattacks, which need to breach infrastructure barriers, internal risk is insidious, as the malicious agent operates within the security perimeter with valid credentials. The Association of Certified Fraud Examiners (ACFE, 2024) describes “occupational fraud” as the use of a professional position for personal enrichment, acting as a silent corrosive factor in organizations (Pinheiro and Silva Cunha, 2003). The financial impact is substantial, with average losses estimated at five percent of companies’ annual revenue, and the ACFE’s global report (2024) recorded, in 1,921 real cases analyzed, losses exceeding 3.1 billion dollars, with a median loss of 145,000 dollars per incident.
In the Brazilian context, the materiality of this risk is frequently evidenced by operations of the Federal Police (PF). In 2025, investigations revealed structured schemes of digital embezzlement operated by employees with active credentials in institutions such as Caixa Econômica Federal (Federal Police, 2025a; 2025c; 2025d). In these local cases alone, the direct loss caused by collaborators exceeds 4.6 million reais. These episodes, which occur systemically in various regions of the country (Federal Police, 2025b), demonstrate that unconditional trust in collaborators is not a viable security policy, reinforcing the urgency and financial relevance of the topic. The complexity and materiality of this risk evolve in parallel with the structural transformations of the market and society. The growing digitalization and the expansion of access to services in the Brazilian financial landscape, according to the Central Bank of Brazil (2025), have increased the volume of interactions and the traffic of sensitive data. This phenomenon, combined with demographic changes such as population aging (IBGE, 2024), diversifies the user base and demands adaptive and scalable solutions, as traditional auditing methods lose effectiveness in the face of high volume and prolonged data exposure.
For the design of effective monitoring systems, it is fundamental to distinguish anomalies generated by errors (inattention or process failure) from those generated by fraud. The ACFE (2024) defines fraud as the deliberate use of a professional position to obtain personal gain or illicit enrichment, resulting in damage to the institution’s resources. The etiological understanding of this behavior is historically based on the Fraud Triangle, which postulates the occurrence of the illicit act by the simultaneous convergence of three factors: pressure, opportunity, and rationalization (Borba and Maragno, 2017). However, in the context of corporate access to complex databases, the literature has evolved to the Fraud Diamond model (Boldt, 2025), which adds a fourth decisive element: technical capability, where the malicious agent exploits their skills and knowledge of the systems to extract data without being detected.
This risk, especially in the banking sector, is aggravated by the requirements of the General Data Protection Law (LGPD) (Brazil, 2018). Unmotivated access by an employee to customer records, statements, or financial histories, even without immediate financial diversion, constitutes a security incident and privacy violation. Traditionally, institutions use sample audits and systems based on deterministic rules, which trigger alerts only when volumetric limits are exceeded. However, as Chandola, Banerjee, and Kumar (2009) demonstrate, modern fraudsters dilute their undue access into apparently normal behaviors, making static systems ineffective in detecting “camouflaged fraud” and generating many false positives. It is in this scenario that Artificial Intelligence, specifically Unsupervised Machine Learning models (Goldstein and Uchida, 2016), becomes crucial, as banking secrecy restricts the use of real databases with labeled frauds (Lerner and Flach, 2024), making stochastic simulations a scientific and ethical path. The use of Graphs and Social Network Analysis (SNA) allows for assessing the risk of employee-client interactions (Akoglu, Tong, and Koutra, 2015), and the implementation of AI in combating internal fraud aims to instill a strong “detection perception” (ACFE, 2024), acting as a preventive control mechanism.
Given the complexity and materiality of internal threats, and the ineffectiveness of traditional approaches, the need to develop and evaluate predictive anomaly detection systems is justified. Thus, the present work aims to propose and evaluate an anomaly detection system structured in Social Network Analysis (SNA), analyzing and comparing the effectiveness of the unsupervised algorithms Isolation Forest (IF) and Local Outlier Factor (LOF) in identifying atypical access patterns in a simulated stochastic database.
2. Material and Methods
The present study was characterized as an applied research with a quantitative approach (Gil, 2008; Richardson, 1999). Its objective was to propose and evaluate a system for detecting behavioral anomalies in corporate access to sensitive data, employing Unsupervised Machine Learning techniques. The choice of this approach was dictated by the restriction of access to real databases with labeled fraud, in compliance with banking secrecy and the General Data Protection Law (LGPD) (Lerner and Flach, 2024).
To overcome the scarcity of real data and ensure reproducibility, a stochastic synthetic database was built, totaling approximately 1 million log records. This simulation modeled probabilistic profiles of legitimate behavior with noise, disguised fraud, and volumetric fraud. A fixed random seed (random seed = 77) was employed to ensure the determinism of distributions and exact replication. The simulated dataset comprised 1,000 employees, 47,980 customers, and 990,570 access records.
The data structure was organized into three interconnected databases: Employees (`id_employee`, `id_branch`), Customers (`id_customer`, `id_branch_agency`, `account_status`) and Access Logs (`id_employee`, `id_customer`, `access_timestamp`, `simulated_fraud`). The `simulated_fraud` variable was not used during training, serving as Ground Truth for validation. The behavioral parameterization of the stochastic simulation established three overlapping profiles: legitimate with noise (98.5% of the dataset), camouflaged fraud, and volumetric fraud. Fifteen employees were calibrated as fraudsters, divided into tactics of Explorers, Necros, and Re-activators.
The temporal metadata (log timestamps) emulated business hours, with 96% of accesses concentrated between 08:00 and 17:59, and 4% distributed as atypical accesses. The methodology adopted Social Network Analysis (SNA), based on Graph Theory, to model interactions between employees and clients (Akoglu, Tong, and Koutra, 2015). Employees and clients were represented as “Nodes”, and access logs as “Edges” in a weighted bipartite graph. To extract topological attributes, Weighted Degree and PageRank were used (Page et al., 1999).
To mitigate the “Volume Paradox” and the “semantic blindness” of PageRank in heterogeneous networks, an Attribute Engineering based on a Weighted Forensic Scoring Heuristic was implemented. Exponential weights were assigned to accesses according to criticality: deceased client accounts received a weight of 10; inactive accounts and cross-branch accesses received a weight of 5; and regular accesses received fractional weights (0.3 and 0.5). Derived variables included `total_acessos`, `v_fal`, `v_ina`, `v_cro`, and `Pontuação Total` (sum of volumetrics multiplied by business weights).
For anomaly detection, the performance of two unsupervised algorithms was compared: Isolation Forest (IF) (Liu, Ting, and Zhou, 2008) and Local Outlier Factor (LOF) (Breunig et al., 2000). IF focuses on isolating global anomalies, while LOF adopts a local density-based approach, sensitive to camouflaged behavioral deviations (Chandola, Banerjee, and Kumar, 2009). The contamination hyperparameter was adjusted to 1.5%. The traditional split of data into training and testing was not performed, and the algorithms operated in a “blind” scenario (zero-shot learning).
To evaluate the generalization capacity, a cross-validation was performed through multiple simulations (Monte Carlo). The stochastic seed of the data generator was sequentially altered, creating five independent scenarios, each generating a new network topology. Computational development was carried out in Python, using the NetworkX library for graph modeling and scikit-learn for Machine Learning algorithms. Large Language Models (LLMs) were employed as computational support for code structuring and optimization, in accordance with the Best Practices Manual (Pecege, 2025).
3. Results and Discussion
The investigation of atypical access patterns in sensitive corporate environments revealed crucial insights into the effectiveness of unsupervised anomaly detection algorithms. The results obtained from a simulated stochastic database, with approximately 1 million records, allowed for the validation of the proposed methodology and comparison of the performance of Isolation Forest (IF) and Local Outlier Factor (LOF). The study demonstrated that modeling interactions as a bipartite graph, weighted by risk, combined with a hybrid Machine Learning architecture, is a promising approach to identify camouflaged internal threats, overcoming the limitations of traditional methodologies based on deterministic rules and simple volumetrics.
Exploratory Data Analysis (EDA) was fundamental to understanding the complexity of the simulated environment and validating the representativeness of the inserted anomalies. The histogram of the Weighted Degree metric, for example, showed a concentration of legitimate employees in lower risk scores, while the box-plot demonstrated that the application of forensic scores elevated the risk median for fraudulent profiles. This initial distinction, however, did not translate into obvious linear separability, confirming the need for more sophisticated detection approaches.
The correlation matrix of the derived variables revealed high correlations between volumetric attributes, such as Total Accesses and Total Score, which was expected due to the algebraic derivation of the risk score from volumetrics and forensic weights. However, the topological metric PageRank maintained a substantially lower correlation with absolute volume metrics. This finding is crucial, as the introduction of a non-linear topological variable, such as PageRank, was designed to break determinism and allow Machine Learning models to differentiate an employee with high legitimate workload from a disguised fraudster, addressing what has been termed the “Volume Paradox”.
The scatter plot visually illustrated the study’s central premise: behavioral overlap. Crossing PageRank with Weighted Degree revealed a dense spatial overlap between the atypical actions of legitimate employees (noise) and those of fraudsters. This representation confirmed that linear decision boundaries, typical of static rule-based systems, would be ineffective in separating risk classes. The dispersion and overlap of the data statistically justified the adoption of unsupervised algorithms, capable of interpreting local density anomalies, as highlighted by Chandola, Banerjee, and Kumar (2009).
The topological processing and visualization stage of the network (SNA) validated the scalability of graph modeling. The bipartite network, composed of 48,980 nodes (employees and clients) and approximately 990,000 interaction records consolidated into weighted edges, was successfully instantiated in memory. The extraction of topological metrics, such as Weighted Degree and PageRank, proved to be computationally feasible, allowing for the assessment of each collaborator’s risk influence based on the weight of connections and the atypical centrality of interactions.
However, it was recognized that traditional PageRank has limitations in heterogeneous networks, suffering from “semantic blindness” by not differentiating the qualitative nature of nodes or the business risk associated with connections. To mitigate this limitation and adapt the model to the reality of forensic auditing, a hybrid architecture was adopted. The PageRank score was used as a topological variable within a multidimensional space, concatenated with a weighted Forensic Score that assigns specific weights based on regulatory risk and the client’s account status. This approach allowed for highly contextualized anomaly detection, crossing structural centrality with the legal severity of accesses.
In the comparative performance of unsupervised models, the Isolation Forest (IF) algorithm, parameterized with 2,000 estimators, obtained a Recall of 0.60, Precision of 0.60, and F1-Score of 0.60. Although IF isolated profiles with extreme global anomalies, generating six false positives, it failed to identify 40% of the more cautious fraudsters. This result suggests that, by focusing on global isolation, IF may struggle to detect fraud that camouflages itself within normal behavior, according to the Fraud Diamond theory (Boldt, 2025) which emphasizes the fraudster’s ability to exploit system knowledge to avoid detection.
In contrast, the Local Outlier Factor (LOF) model, using Manhattan distance and evaluating the 25 nearest neighbors, demonstrated significantly superior performance. The LOF achieved a Recall of 0.80, Precision of 0.80, and an F1-Score of 0.80. Detailed analysis revealed that the LOF managed to capture twelve out of fifteen simulated fraudsters, with only three false positives among the fifteen generated alerts, resulting in a Precision of 80%. This performance indicates that the LOF, by focusing on local density, was more effective in breaking the camouflage of anomalies, which manifest as deviations from the immediate topological neighborhood, and not necessarily as global anomalies.
The superiority of LOF is attributed to its ability to identify novel local behavioral deviations, which are often mistaken for legitimate activities, overcoming evasion tactics known as “Low and Slow Attacks” (Chandola, Banerjee, and Kumar, 2009). This characteristic is particularly relevant in internal fraud scenarios, where malicious actors seek to dilute their undue actions within a massive volume of legitimate transactions. LOF’s ability to focus on local density allowed the model to break through this camouflage, identifying atypical access patterns that Isolation Forest could not isolate with the same effectiveness.
To consolidate the comparative analysis and evaluate the models’ generalization capability, cross-validation was performed through multiple Monte Carlo simulations. By sequentially altering the stochastic seed of the data generator, five independent scenarios were created, each with a new network topology and distribution of operational noise and malicious agents. This approach mitigated the bias of a single sample and tested the robustness of the algorithms under different conditions.
The cross-validation results confirmed the robustness of the architecture based on LOF’s local density. The model presented an average Recall of 81.3% in detecting the Fraud class, reaching peaks of 86.6% correct identification of hidden threats in most scenarios. In contrast, the Isolation Forest demonstrated high sensitivity to network variations, with its effectiveness dropping to less than 50% in Scenario 1 and stagnating at an overall average of 56%. This divergence in performance metrics reflects the theoretical limitations of volume-based algorithms, which focus on global anomalies and allow for the evasion of opportunistic frauds.
Achieving an average detection rate (Recall) above 81% in a scenario with extreme class imbalance (only 1.5% anomalies) is of particular scientific and operational relevance. In the Machine Learning literature applied to auditing, evaluating the performance of strictly unsupervised algorithms is challenging, as they lack prior knowledge of labels during the training phase (Goldstein and Uchida, 2016). The fact that the LOF model, anchored in graph topology, sustained this high performance throughout cross-validation demonstrates a highly competitive level of effectiveness compared to academic standards for point anomaly detection.
From a practical and business standpoint, this result resolves the need outlined in the study’s introduction. The high capture rate, combined with a low rate of false alarms, translates into a promising forensic screening tool. In a real corporate environment, governed by the rigors of the LGPD (Brazil, 2018), this assertiveness optimizes the human team’s investigation efforts and avoids the organizational friction of false accusations against high-productivity employees. This level of precision fulfills the objective of creating a strong “detection perception” (ACFE, 2024), deterring internal threats without operationally overloading the institution’s incident response systems.
A strategic differential of this Social Network Analysis (SNA) based architecture lies in its extreme parameterization and architectural flexibility. By modeling access interactions as a weighted graph, the system enables the assignment of dynamic “weights” to edges, calibrated according to granular operational characteristics, such as time of day, volume, or account status, and customer data sensitivity. This methodological plasticity reflects the state of the art in the transition from static cybersecurity systems to data-oriented architectures (Mahajan, 2024).
Unlike obsolete systems based on volumetric rules, the solution based on Machine Learning is inherently self-adaptive, acting in real-time to check continuous signals of behavioral changes (Soares, 2020). Should fraud typologies evolve, the institution does not need to restructure the central algorithms; it is enough to adjust the graph’s weighting function. This adaptation capability ensures a continuous improvement mechanism, allowing the “detection perception” to evolve in symmetry with the sophistication of threats, culminating in an intelligent, dynamic forensic system strictly aligned with the privacy requirements of LGPD.
In summary, the results of this research confirm that the proactive detection of behavioral anomalies in corporate access to sensitive data is viable and effective through a hybrid approach that integrates Social Network Analysis and unsupervised Machine Learning algorithms. The superiority of Local Outlier Factor in identifying camouflaged fraud, combined with the robustness demonstrated in multiple simulations, offers a powerful tool to reinforce internal security, optimize investigations, and ensure compliance with data protection regulations, directly responding to the objective of proposing and evaluating an anomaly detection system.
4. Conclusion
This study aimed to propose and evaluate an anomaly detection system structured in Social Network Analysis (SNA), comparing the effectiveness of the unsupervised algorithms Isolation Forest (IF) and Local Outlier Factor (LOF) in identifying atypical access patterns in a simulated stochastic database. It was found that the digitalization of corporate processes increased exposure to internal threats, where employees with legitimate access violate security policies, and that traditional auditing methods proved ineffective against disguised fraud. The methodology involved modeling employee-customer interactions as a weighted risk bipartite graph, using a dataset of approximately 1 million stochastically generated logs to mimic anomalous behaviors and operational noise. The results demonstrated the superiority of the LOF model, which, by focusing on local density, broke the camouflage of anomalies and achieved an average Recall of 81.3% in multiple simulations, surpassing the performance of Isolation Forest, which obtained an average Recall of 56%. This hybrid approach, which concatenated topological network metrics with forensic risk scores, proved crucial for contextualizing detection and overcoming the “semantic blindness” of traditional PageRank, differentiating legitimate high-volume access from fraudulent actions.
The study’s main contribution lies in proposing a viable and scalable approach for proactive insider threat detection, enhancing “detection perception” and providing managers with a high-risk access base to prioritize precise investigations, in strict compliance with LGPD. The solution represents a transition from static rule-based systems to self-adaptive, data-driven architectures capable of evolving with fraud typologies. However, the research used a simulated stochastic database due to banking secrecy and LGPD restrictions, which requires validation in real environments for full implementation. For future studies, it is suggested to implement a feedback loop mechanism, where the results of human investigations on generated alerts feed back into the model, allowing artificial intelligence to continuously adjust its topological sensitivity and learn new fraud patterns, mitigating false positives and adapting to evolving threats.
Bibliographic References
Association of Certified Fraud Examiners [ACFE]. 2024. Fraude Ocupacional 2024: Um Relatório para as Nações. Disponível em: Relatório ACFE 2024 para as Nações. Acesso em: 20 set. 2025.
Pinheiro, G.J.; Silva Cunha, L.R. 2003. A importância da auditoria na detecção de fraudes. Contabilidade Vista & Revista 14(1): 31-47.
Polícia Federal. 2025a. PF prende em flagrante funcionário da Caixa por peculato digital no Rio de Janeiro. Disponível em: https://www.gov.br/pf/pt-br/assuntos/noticias/2025/09/pf-prende-em-flagrante-funcionario-da-caixa-por-peculato-digital-no-rio-de-janeiro. Acesso em: 11 out. 2025.
Polícia Federal. 2025c
Verizon. 2023. Relatório de Investigações de Violação de Dados de 2023: Visão geral do setor público. Disponível em: https://www.verizon.com/business/resources/Ta5a/reports/2023-dbir-public-sector-snapshot.pdf. Acesso em: 03 set. 2025.
Article originating from the Final Course Work of the Specialization in Data Science and Analytics of the MBA USP/Esalq
To learn more about the course, click here and access the MBX Academy platform