Article

Software Engineering

October 08, 2026

Performance and Total Cost of Ownership of Cloud Databases and On-Premises Infrastructure

Performance and Total Cost of Ownership of Cloud Databases and On-Premises Infrastructure

José Antônio Elias da Silva Júnior; Manoel Flavio Leal

DOI: 10.22167/2675-6528-202603055

Article derived from a Course Conclusion Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.

Summary

This study analyzed the performance of database operations in cloud and local infrastructures, correlating technical behavior with Total Cost of Ownership projection. The research was characterized as a quantitative and experimental study, in which a test system subjected isolated database instances to progressive execution loads, measuring the impact of network latency and computational consumption. For financial analysis, an investment and operational expense model was developed, diluted over a thirty-six-month cycle. Technical results revealed that the accumulation of internet latency caused severe time degradation in cloud executions, despite the remote infrastructure operating with high processing idleness, registering almost 98% CPU inactivity. In contrast, the local environment achieved superior performance supported by almost instantaneous network communication. Financially, cost consolidation demonstrated an empirical tie between the physical acquisition model and the service subscription model within a three-year horizon, but the local environment proved more advantageous in a sixty-month cycle. It was concluded that the degradation in the cloud was not due to computational capacity, but to the interaction between route latency and the application’s unitary communication pattern. Cloud adoption requires deep optimization of the system architecture to minimize dependence on constant communication with the remote server. Without this modernization, the local infrastructure consolidated itself as the most viable strategy, ensuring high performance, budgetary predictability, and data sovereignty.

Keywords: Operating expenses (OPEX); Transactional scalability; Network latency; Legacy systems.

1. Introduction

The evolution of information technology has been marked by a significant transition in infrastructure models. Historically, organizations relied on on-premises architectures, which, despite being widely used, presented inherent complexities and high costs associated with maintaining their own infrastructure. More recently, the cloud computing paradigm has emerged as a scalable alternative for outsourcing the management of this infrastructure.

The decision between adopting a cloud infrastructure or maintaining an on-premises structure has become a central dilemma for modern software companies, requiring the re-evaluation of traditional repositories and architectures in the face of new technological demands (Noor et al., 2024). From a financial perspective, studies such as Kumar (2024) indicated that migrating to the cloud can generate savings of 30% to 40% in infrastructure costs in specific scenarios. However, for applications sensitive to network response time (latency), on-premises implementations have shown to achieve up to 42% superior performance (Kumar, 2024), creating a complex trade-off between cost and speed.

In operational practice, the pursuit of the best cost-benefit ratio reveals that the choice between models goes far beyond simple server allocation. Maintaining on-premises infrastructure requires a high initial investment (capital expenditure), but can prove financially advantageous and predictable in the long term. Furthermore, maintaining one’s own servers guarantees total sovereignty over data traffic, a crucial factor in high-volume operations such as massive database conversions during system migration (Holmnäs, 2024). On the other hand, many organizations have failed when migrating to the cloud by ignoring the “hidden costs” of performance. A legacy application, originally optimized for local networks, can become prohibitively expensive in the cloud due to the need to increase the consumption of computational resources (commercially measured in vCores or DTUs) simply to compensate for code inefficiencies and network communication latency.

Recent research in systems architecture has shifted the focus from processing power to the cost of communication between components. Zhang et al. (2025) demonstrated that, in databases with disaggregated storage, remote access latency becomes the dominant degradation factor. Xu et al. (2025) identified synchronization between geographically distributed nodes as the main bottleneck in long-distance networks. Comparative literature between cloud and on-premises infrastructure, however, focuses either on cost aggregation (Pillai, 2024) or on the performance of isolated infrastructure (Husain et al., 2024), without separating the computational effort of the database engine from the waiting imposed by the application’s communication pattern. It thus remains undetermined whether the slowness attributed to the cloud stems from the remote location of the database or the access architecture, a distinction with opposite consequences.

It is evident, therefore, that there is no definitive or generalizable solution for all companies; each context requires careful analysis. Given this scenario, this work is justified by the need to provide empirical data to assist managers in directing their technological strategies assertively. The present study aims to analyze the capacity and performance of database operations in both infrastructures, correlating these technical metrics with the Total Cost of Ownership. The central purpose is not to point to one technology as absolutely superior, but rather to explore its advantages and disadvantages based on empirical data, culminating in the development of a decision matrix that assists managers in directing their technological strategies assertively.

2. Material and Methods

The present study was characterized as a research with a strictly quantitative approach and experimental design. The central objective of the experiment was to subject database infrastructures, allocated in distinct physical media, to progressive transactional loads. The impact of network latency and the internal processing capacity of each environment were measured, aiming to correlate these technical metrics with the Total Cost of Ownership, according to the established general objective.

To ensure the empirical validity of the study, the design focused on isolating variables. It was ensured that the computational effort of the database engine could be chronometrically separated from the time required for packet transport through the public internet and the local network. The experiment’s architecture was strictly segregated, separating the development environment from the execution environment, and the test topology consisted of three independent and isolated instances.

The first instance operated as a Client Machine, configured as a dedicated virtual machine. It ran the “Windows 10” Operating System, provisioned with 4 GB of RAM and 2 logical processing cores, based on the AMD Ryzen 5 5500U architecture. This environment was exclusively and solely responsible for executing the performance measurement application, without the presence of heavy development tools in the background.

The second instance represented the corporate physical environment, named “On-Premise”. It was configured as a virtual machine under the “Windows Server” operating system, possessing hardware specifications identical to the client machine, dedicated to hosting the “SQL Server 2019 Express” database engine. To ensure the fidelity of the local scenario, the virtual network adapters of both machines were configured in bridge mode (“Bridge”).

With this parameterization, each instance operated autonomously in the infrastructure, receiving independent IP addresses assigned by the physical router. Communication occurred through real TCP/IP protocol packet traffic over the local area network (LAN). This isolation prevented the use of loopback connections, which would invalidate the precise measurement of latency characteristic of a physical corporate network.

The third instance of the topology represented the remote cloud environment, named “Cloud”. It was constituted by a database “Azure SQL Database”, provisioned in the virtual core (“vCore”) based purchase model, allocated under the General Purpose service tier. The computation was configured in “Serverless” mode, using the “Standard-series (Gen 5)” structural hardware generation.

The dynamic scaling of processing was bounded to operate with a minimum of 0.5 and a maximum of 2 “vCores”. This configuration was intentionally chosen to represent an input baseline, reflecting the initial scenario of adoption by real companies seeking concept validation or cost elasticity.

The stress testing engine consisted of a console application developed in C# and compiled on the .NET 8.0 platform. Access to data and transactional communication between the client application and the two database instances were mediated by the low-level object-relational mapper Dapper (DAPPERLIB, 2024).

The choice of Dapper was justified by its very high performance. By directly mapping native instructions to objects in memory, Dapper eliminated the overhead of query translation, adding the least possible processing time to data conversion.

A critical methodological factor for the experiment’s validity was the rigorous standardization of the data mass and network traffic packets (“payloads”). The structure of the tables and the synthetic data generated in memory were scaled with fixed sizes and strictly identical character padding. To mitigate this risk, the data generation library “Bogus” (BOGUS, 2024) was used.

This approach ensured that every TCP/IP packet sent from the client to the server, regardless of whether it was in the local environment or in the cloud, had exactly the same weight in bytes. This isolated route latency as the sole temporal variable of the experiment, eliminating any possibility of bias in the results due to discrepancies in the volume of information exchanged in each iteration.

To avoid ambiguity in reading the results, a strict operational definition of the term transaction was adopted. Each transaction corresponded to a single command submitted individually to the bank with automatic confirmation, which is equivalent to a single network round trip. Explicit transactional blocks grouping multiple commands were not used.

Thus, “a load of 128 transactions” designated 128 independent commands, submitted sequentially, each awaiting confirmation from the previous one. This pattern is predominant in legacy record-oriented systems, which constituted the object of this research. The composition of each operation, including the submitted command, the relational validations required from the engine, and the number of round trips per load unit, was detailed.

In all operations, the description field was filled with exactly 5,000 characters, so that the trafficked package had identical weight in both environments. In modifications and deletions, the prior query was executed only once per batch, outside the timer, so that the measured time strictly corresponded to the unit writes.

To simulate the processing effort consistent with real-world business applications, the main table (“TransacaoFinanceira”) was modeled with strong relational bindings. The algorithm generated numeric identifiers that act as foreign keys for the User, Category, and Bank tables. This modeling ensured adherence to ACID properties (Atomicity, Consistency, Isolation, and Durability).

The testing strategy evaluated the three fundamental database persistence operations: Insertion, Modification, and Deletion. To simulate real operational scenarios and maintain test fairness, the modification and deletion operations were not executed blindly, following the “fetch and process” architectural pattern.

Firstly, the application performed a complex query that deterministically returned the identifiers to be manipulated. Then, the submission of alteration requests occurred unitarily within loops. The test batteries followed a geometric growth in powers of base two.

Each cycle started execution with a single unit operation, successively doubling the load (2, 4, 8, 16, 32, and 64), until reaching the saturation limit stipulated at 128 sequential transactions. For the orchestration of data collection and statistical extraction, the “BenchmarkDotNet” precision library (BENCHMARKDOTNET, 2024) was used.

Each geometric scenario was fully repeated four times for each operation, attesting that the infrastructure behavior did not have a punctual or random character. A vital architectural precaution was adopted regarding the “Cold Start” phenomenon, intrinsic to “Serverless” architectures and to the Just-In-Time compilation of .NET.

Before the official timer of each battery was triggered, the global preparation method of BenchmarkDotNet executed natural warm-up routines. These routines consisted of pre-opening database connections, performing count queries, and replenishing exhausted data masses. This control ensured the dynamic allocation of vCores in the cloud, the establishment of active connection pooling, and the uploading of indexes to RAM.

With this mitigation, it was ensured that the final telemetry strictly captured the cost of transport latency, exempting the measurement of the cloud infrastructure’s wake-up lethargic time. To ensure experiment reproducibility, the software architecture developed for measuring response times was isolated into specific classes, using the attributes of the BenchmarkDotNet library.

The update class, named “UpdateBenchmark”, was developed to measure response times, incorporating attributes from the BenchmarkDotNet library for data mass control and mitigation of the Cold Start phenomenon. The global preparation method verified the count of active transactions in the database before the official timer.

The execution of database instructions was delegated to the Dapper package, with the parameterized update instruction submitted atomically by the “Execute()” command, confirming the “row-by-row” test model. The synthetic data generation algorithm, implemented with the Bogus library, was not purely random but contained within restrictive parameters. The description field was instructed to create alphanumeric sequences with exactly 5,000 characters, ensuring that every TCP/IP packet had the same byte weight.

To ensure the repetition of exclusion batteries while maintaining statistical rigor without the need to recreate the infrastructure at each cycle, a differential reset algorithm (“ReporMassaDiferencial”) was developed, integrated into the “benchmark” library lifecycle. Before starting the timer to measure the exclusion time, the “Setup” method performed an active record count.

When it was identified that the table volume was below the pre-established threshold of ten thousand records, the system dynamically calculated the exact difference and requested the repository to insert the missing quantity. This mechanism operated as a state stabilizer, ensuring that all validation rounds of the deletion operation found the database in the same volumetric and index fragmentation conditions.

For the financial analysis, a Total Cost of Ownership (TCO) projection was developed for an initial lifecycle of thirty-six months. The choice of this period was based on standard accounting practices for the depreciation of technology assets. The mathematical methodology of Equivalent Monthly Cost was used, which dilutes capital investment over the months of operation, adding it to recurring operational expenses.

The composition of initial investment and operational expenses, including hardware infrastructure costs, electricity tariffs, cloud service pricing, and real estate rental, was based on market data (DELL, 2026; ENEL, 2026; MICROSOFT, 2026; QUINTOANDAR, 2026).

3. Results and Discussion

The analysis of the empirical data collected during the experiment revealed crucial insights into the performance and cost of database operations in cloud and on-premises infrastructures. The central objective was to go beyond the mere presentation of metrics, promoting a critical discussion on how the choice of infrastructure impacts the transactional behavior of the database and, consequently, the financial and architectural viability of projects. The results were structured to progressively answer the initial hypotheses of the study, confronting the laboratory findings with recent academic literature to validate the proposed conclusions.

The technical evaluation was carried out through a rigorous telemetry process, which captured response time and physical resource consumption under progressive request loads. By subjecting the database engines to execution cycles ranging from unit transactions to intensive batches of 128 operations, it was possible to isolate the actual computational effort of the data engine from the time required for packet transport. This segregation was fundamental to identifying whether the bottlenecks were physical hardware limitations or restrictions imposed by the network topology, allowing for a deeper understanding of the observed phenomena.

Transactional Latency Analysis in Cloud Environment

The technical performance evaluation in the cloud environment, subjected to unitary and progressive loads (from 1 to 128 transactions), revealed a notable empirical phenomenon: high network latency obscured the processing time differences that would normally exist between Insert, Modify, and Delete operations. The cloud environment presented virtually overlapping time curves for these operations. This behavior indicated that the internal processing of the database engine consumed a negligible fraction of the total time compared to the remote communication time, resulting in a delay close to 2,600 milliseconds at the maximum load of 128 transactions.

Almost the entire time measured in this scenario resulted from packet transit (round trip) on the public network. This empirical behavior is strongly supported by contemporary literature, where Cho’ponov (2026) identifies a structural and unavoidable “latency rate,” stipulated between 10 and 50 milliseconds per network hop. Verbitski et al. (2017) also point out that the central restriction for high-throughput processing has shifted from computation and storage to the network. Dividing the total time of the experiment’s maximum load (approximately 2,500 milliseconds) by the 128 submitted iterations, an average latency of 19.5 milliseconds per operation in the cloud was observed, a value that aligns with the critical degradation range predicted for remote environments.

A technical factor that actively contributed to the observed time stacking was the nature of the communication protocol used by the database engine, the “Tabular Data Stream” (TDS). Husain et al. (2024) point out that, in Microsoft SQL Server solutions operating as a service, each unit request requires an acknowledgment before the next instruction in the code is processed. On local networks, this round-trip time is negligible, but in the cloud, the time spent on packet encapsulation and long-distance route transposition becomes the majority component of the operation, overshadowing the platform’s actual processing capacity.

From the perspective of distributed software engineering, Yang et al. (2023) explain that this slowness is not a configuration failure, but rather the architectural price required to maintain strong data consistency. By the CAP Theorem, when partitioning the system across the internet, the remote database must ensure that each write instruction is consolidated across storage nodes before successfully responding to the client application. Consequently, the system code is forced to wait for transit confirmation at each loop iteration, generating the severe temporal accumulation effect observed.

Transactional Latency Analysis in Local Environment (“On-Premise”)

In contrast to the remote scenario, the analysis of the local environment corroborated the research premise from an inverse perspective. Free from the long-distance traffic penalty, the local infrastructure not only responded substantially faster, finishing the maximum batch of 128 transactions in the range of 250 to 410 milliseconds, but also revealed the true computational cost of each operation. In the isolated scenario of almost zero network latency, it became visually possible to notice the divergence of the Exclusion curve from the Insertion curve, a behavior that was previously masked by exogenous waiting on the internet.

The visibility into the internal effort of the database engine is validated by Cho’ponov (2026), who argues that the physical proximity between the application and the server eliminates the additional network load, allowing the infrastructure to deliver its full nominal performance. While in the cloud latency consumes most of the runtime, in the local environment the measured time is a faithful representation of algorithmic efficiency and disk write speed. As observed in the experiment, the absence of external hops allowed for a reduction of approximately 90% in the total runtime for the load of 128 transactions.

Additionally, Husain et al. (2024) reinforce that local Microsoft SQL Server solutions offer processing predictability that is difficult to replicate in database-as-a-service environments without massive investments in redundant hardware layers. The collected data demonstrate this stability, in which the growth of response time occurs linearly and strictly proportionally to the increase in load. Fluctuations and peak waits characteristic of public internet traffic are not observed, ensuring constant execution regardless of external connectivity factors. This factor is identified by Pillai (2024) as critical for sectors operating with massive data volumes and requiring low operational latency as a non-negotiable business requirement.

Global Comparative Analysis and Impact on User Experience

The magnitude of the penalty imposed by the unit request model on cloud infrastructure becomes categorically evident when consolidating the global average of operations. An immediate divergence of performance curves is observed right in the first load cycle. In the unit transaction (1 cycle), the local architecture processed the request in an average of less than 7 milliseconds, while the cloud demanded approximately 50 milliseconds, an initial difference that establishes the efficiency benchmark of each environment. This divergence is not an isolated event, but the empirical materialization of what Cho’ponov (2026) defines as “latency rate”.

The author argues that the transformation of simple internal system calls into external network requests inserts a temporal cost that progressively accumulates with each new network hop. As the code submitted for testing executes repetitive iterations, this initial latency of 50 milliseconds multiplied throughout the execution loop, resulting in the performance abyss observed at the load limit of 128 transactions. From a strategic and business perspective, this temporal degradation directly compromises the “user-perceived performance” (Järvinen, 2025), which highlights end-user dissatisfaction when the software architecture requires multiple round trips to the remote server to complete a single logical task.

The data validate the thesis that migrating an application to the cloud without proper refactoring to minimize constant database interaction will result in chronic slowness, regardless of the contracted processing power. This scenario accurately illustrates the risk associated with the migration strategy known as “Lift-and-Shift”. Kansara (2024) warns that simply moving on-premises systems to the cloud, without prior adaptation of the communication architecture, often turns the elasticity of cloud infrastructure into an unsustainable operational bottleneck. Legacy code that assumes the latency of a local network, when operating on the public internet, makes the system technically inefficient and financially risky.

CPU Utilization and Idle Analysis

Beyond temporal degradation, the technical and financial viability of a database architecture is directly linked to the efficiency in resource consumption, especially computational processing. Telemetry during the execution of the maximum load (128 transactions) revealed that the local environment engaged an average of 14.2% of its processor capacity, while the remote cloud server operated with an idle rate close to 98%, consuming only 1.3% of CPU. This demystifies the hypothesis that the cloud’s slowness would be associated with an exhaustion of its computational capacity.

The consumption of only 1.3% of CPU in the cloud proves that the database resolved the write or read request in milliseconds, and spent the rest of the test’s 2.5 seconds in a state of absolute inactivity, just waiting for the next packet to arrive via the internet. From the perspective of software engineering in the Microsoft SQL Server architecture, this underutilization reveals that the platform spent most of its runtime stuck in the wait event technically known as “ASYNC_NETWORK_IO”. This event occurs when the database engine finishes internal processing almost instantly but is forced to suspend operations because the client application, separated by public network latency, is slow to consume the data or send the next instruction.

This finding corroborates the analyses by Husain et al. (2024), who observe that hardware provisioning in database-as-a-service (DBaaS) solutions often becomes underutilized due to severe external network input/output (I/O) bottlenecks. In contrast, the local infrastructure, with negligible latency on the local area network (LAN), allowed packets to arrive uninterruptedly, causing the database to work in a continuous flow. Without the “ASYNC_NETWORK_IO” block, the physical server finalized the entire block of operations quickly, justifying the full use of the acquired resource and demonstrating superior architectural efficiency for sequential request systems.

From a financial perspective, this cloud idleness translates into a severe hidden cost. Brown (2025) highlights that serverless architectures, such as the Azure SQL Database edition used in this experiment, are marketed with the promise of granular billing and high financial efficiency. However, the author warns about the trade-offs generated when code is not optimized for this billing model. As the cloud bill is calculated by active computation time (billed per second), submitting single iterations forces the database to keep the session open for long periods to resolve a minuscule workload. Brown’s (2025) thesis is therefore validated: the on-demand billing model becomes highly burdensome if the application does not use batch processing, turning network-induced idleness into a direct financial liability for the organization.

Decomposition of Time between Processing and Waiting

Once the performance difference was established, the research sought to answer a fundamental causal question: does the degradation stem from the bank’s remote location or the application’s communication architecture? To resolve the issue, the total time was decomposed into two mutually exclusive parts: the time the engine was effectively processing, obtained by the product of the total time and the average processor utilization percentage measured by telemetry, and the residual time waiting for communication. Under a load of 128 transactions, the effective processing time remained in the same order of magnitude in both environments, with 44.0 milliseconds on the local server versus 33.2 milliseconds on the remote instance.

The result is conclusive: in the computational work performed, the cloud database was no slower than the local one. The entirety of the difference is concentrated in the waiting portion, which increased from 266.0 milliseconds in the local environment to 2,516.8 milliseconds in the cloud. It is concluded that the remote infrastructure does not, in itself, constitute the cause of the degradation. The determining factor is the interaction between the route latency, which fixes the cost of each round trip, and the application’s unitary communication pattern, which determines how many round trips occur. Latency acts as a multiplier and the architecture as an amplifier; in isolation, neither explains the result.

A second finding reinforces this interpretation: the ratio between cloud and local environment times remained practically constant across the entire scale, from the unit transaction to the batch of 128. If the degradation were due to exhaustion of the contracted capacity, an increasing ratio would be expected, with progressive divergence of the curves. The constancy indicates a multiplicative and structural penalty, proportional to the number of round trips, and not resource saturation. The decomposition allows estimating the effect of an architectural change that preserved the contracted infrastructure: by replacing the 128 unit round trips with a single batch submission, the expected time in the cloud would be on the order of 53 milliseconds, compared to the approximately 2,550 measured. This is an analytical projection derived from the collected data, and its order of magnitude supports the thesis that the cloud was not slow, but the application was conservative.

The finding dialogues with recent literature. Zhang et al. (2025) observed that, in disaggregated storage, remote access latency overlaps with processing capacity, penalizing writes most of all, a behavior identical to that measured here. Xu et al. (2025) demonstrated that the bottleneck of geographically distributed banks resides in synchronization over long-distance networks, not in the processing of each node. The convergence indicates a structural phenomenon of distributed architectures, where network latency and application communication patterns interact to determine performance.

Load Regimes and Contextual Superiority of Each Architecture

Once the origin of the degradation was established, it became possible to qualify in which contexts each architecture is superior. Decomposition allows any workload to be classified according to the dominant portion of its runtime. A network-bound regime is one in which waiting for communication exceeds effective processing, observed in both environments of this experiment. In this regime, performance is governed by the product of round-trip latency and the number of round trips, and contracting more capacity does not yield appreciable gain, as the additional resource remains idle waiting. This regime is characteristic of unitary and legacy transactional systems, where local infrastructure is recommended because time is governed by the number of round trips.

A processing-limited regime is one in which computational effort exceeds communication time. Analytical workloads and batch processing fall into this category, as they amortize a single network latency over a large volume of work. Here, cloud elasticity becomes an effective advantage, and the local environment is penalized by the rigid ceiling of acquired hardware. The practical implication is that infrastructure choice should not precede the characterization of the application’s predominant regime. A decision matrix was developed to consolidate this analysis, recommending the cloud for batch processing, analytical workloads, seasonal demands with pronounced peaks, and applications refactored for bundled submission, where a network latency is amortized over a large volume or asset idleness is avoided.

The decision matrix also indicates the local environment for constant volumes in long cycles and the requirement for data sovereignty, as internal traffic is not charged and no third parties are involved. The matrix highlights that no architecture is absolutely superior; superiority is contextual and determined by the workload regime, a premise already supported by Tan et al. (2019) who demonstrated that the choice of a cloud database depends on the match between the architecture and the workload profile, rather than a single performance ranking. The recurring mistake in migrations consists of transposing to the cloud an application limited by network without altering its communication pattern, paying for elasticity without being able to consume it, which converges with Kansara’s (2024) warning regarding the direct transfer of legacy systems.

Total Cost of Ownership [TCO] Decomposition

To quantify the financial impact of the architecture choice and materialize the effects of latency, a Total Cost of Ownership (TCO) projection was developed for an initial lifecycle of thirty-six months. The choice of this period is based on standard accounting practices for the depreciation of technology assets. To enable an equitable comparison between the on-premises physical infrastructure acquisition model and the cloud services subscription model, the mathematical methodology of Equivalent Monthly Cost was used, which dilutes capital expenditure (CAPEX) over the months of operation, adding it to recurring operational expenses (OPEX).

The projection of the Equivalent Monthly Cost revealed a technical tie and strict financial parity in the three-year scenario, with the local environment costing R$ 2,233.30 and the cloud subscription R$ 2,216.00. This empirical finding contradicts much of the common sense in the technology market, which often associates the adoption of cloud computing with a drastic and automatic reduction in infrastructure expenses. Leis and Kuschewski (2021) demonstrate, in the same vein, that performance and cost in the cloud vary by orders of magnitude depending on the contracted configuration, so that there is no economy dissociated from scaling.

Pillai (2024), when analyzing data storage and processing strategies in financial institutions and traditional sectors, argues that the costs of local infrastructure, although requiring a high initial capital investment, are amortized in a highly predictable and secure manner in the long term. The data validate Pillai’s thesis, demonstrating that, by judiciously diluting electricity, real estate, and physical deployment expenses, the architectural decision depends not on supposed immediate savings in the cloud, but rather on the maturity of the application code and the rigor in controlling the monthly budget. The initial investment costs for the local environment (equipment and licenses) were R$ 24,399.00 and deployment (physical and logical) of R$ 9,200.00, with a recurring monthly operational cost of R$ 1,300.00. For the cloud, there was no initial investment, but the recurring monthly operational cost was R$ 2,216.00.

Scenario Study, Scalability and Long-Term Projection

To deepen the validation of the financial model and test the resilience of both architectures, a scenario study was designed, extending the infrastructure’s lifecycle to 60 months (five years), a timeframe frequently adopted by the corporate market for total server park renewal. By recalculating the dilution of the initial local investment (R$ 33,599.00) over this new period and adding the recurring operational cost (R$ 1,300.00), the Equivalent Monthly Cost of the physical environment drops drastically to R$ 1,859.98. In contrast, the cloud service subscription model maintains its static floor of R$ 2,216.00, assuming an optimistic scenario of no inflation or provider price adjustments. In this extended horizon, the local environment consolidates itself as the most financially efficient option, proving that the operational longevity of the physical asset significantly rewards the immobilized capital.

However, the most critical factor for the business lies in simulating a stress scenario, such as an abrupt 20% increase in data traffic. In a local infrastructure, the cost is fixed; subjecting the server to a higher load will primarily result in a marginal degradation of internal response time, without any change in the monthly bill charged to the company. In the cloud, the behavior is diametrically opposite and financially aggressive. Järvinen (2025) concludes that, in unoptimized remote architectures, cloud costs scale punitively. When facing an increased load, the platform configured for automatic scalability will attempt to compensate for software inefficiency by provisioning maximum computing levels (“scale-up”).

An order of magnitude for the impacts of this reactive scalability is reported by Cho’ponov (2026), who estimates an increase of up to 132% in monthly TCO when systems maintain strong couplings and high request frequency over the network. Adopted for illustrative purposes, the cloud cost of R$ 2,216.00 would be close to R$ 5,000.00 per month if the application forced the database to scale to the processing ceiling to mitigate the impact of latency. The value constitutes a sensitivity scenario, not an empirically validated projection. Finally, it is imperative to add data egress costs to the projection. Pillai (2024) warns that many companies are attracted by the low cost of entry into the cloud, but become hostages to high tariffs when transacting large volumes of data back to local networks. In the evaluated local model, internal traffic is free and unlimited. This sovereignty over information transit, combined with protection against reactive scalability, gives local infrastructure an incomparable strategic advantage, protecting corporate cash flow and showing that cloud success requires deep and mandatory optimization of system code.

Comparison with Market Alternatives and Break-Even Point

The identified parity refers to two specific models: proprietary acquisition and bank subscription as a service at the entry layer. The market, however, offers intermediate modalities whose omission would impoverish the analysis. Among the market alternatives for relational database hosting, the following stand out: on-demand managed service, managed service with reservation, self-managed virtual machine, dedicated server or colocation, and proprietary on-premises infrastructure. The most relevant differentiation lies in licensing and traffic. In managed offerings, the engine license is embedded in the tariff, which favors organizations without prior licenses, while companies with amortized licenses tend to pay twice for the same right of use. Reserved capacity, on the other hand, brings the cloud closer to the acquisition model, but suppresses the elasticity that constitutes its main attraction.

Applying the values from this research, the break-even point between the two models is obtained. With an initial investment of R$ 33,599.00 and a recurring cost of R$ 1,300.00 per month in the local environment, versus R$ 2,216.00 for the subscription, the monthly operational savings are R$ 916.00, which places the break-even point at approximately 36.7 months. This explains the parity observed in the thirty-six-month cycle: the adopted horizon coincides almost exactly with the moment when the investment pays for itself, which is why the equivalent monthly costs converge. The break-even point offers an objective criterion: horizons shorter than thirty-six months, or with uncertainty about continuity, favor the subscription, which avoids capital immobilization; longer horizons favor acquisition.

Two external evidences corroborate the reading. Flexera (2025) found that 27% of the declared spending on infrastructure and platform as a service corresponds to unused or poorly sized resources, a proportion compatible with the idleness measured here; the waste does not stem from provider overpricing, but from capacity that the architecture cannot consume. Convergently, the repatriation carried out by the company 37signals, with projected savings exceeding ten million dollars in five years (Hansson, 2024), illustrates the same mechanism on an industrial scale. For the corporate market and technology management, these findings generate immediate practical implications. It is evident that the migration strategy based on the simple direct transfer of legacy systems is technically flawed and prone to risks. Monolithic systems, designed assuming the almost null latency of a local network, cannot be transposed to the cloud without a profound refactoring of their communication logic.

The success of cloud computing requires fostering a culture of excellence in engineering and financial operations, necessitating the modernization of systems to event-driven architectures and batch processing. In scenarios where this refactoring is unfeasible or where informational sovereignty and budgetary predictability are priorities, on-premises infrastructure remains a sovereign and recommended technical choice. In summary, the results demonstrate that performance degradation in the cloud stems not from computational capacity, but from the interaction between network latency and the application’s unitary communication pattern, directly impacting Total Cost of Ownership and long-term viability, consolidating on-premises infrastructure as the most viable strategy for legacy systems without optimization.

4. Conclusion

This study analyzed the performance and Total Cost of Ownership of database operations in cloud and on-premises infrastructures. It was found that performance degradation in the cloud was not due to computational capacity, but to the interaction between route latency and the application’s unitary communication pattern, which forced the database engine to operate with almost 98% CPU idleness. In contrast, the on-premises environment achieved superior performance, supported by near-instantaneous network communication, revealing the true computational cost of operations. Financially, an empirical tie was observed in the Total Cost of Ownership over a thirty-six-month horizon, but the on-premises environment consolidated as the more advantageous option over a sixty-month cycle and in stress scenarios, due to budgetary predictability and data sovereignty. The main contribution of this work lies in the development of a decision matrix that assists managers in directing their technological strategies, highlighting that cloud adoption requires deep optimization of the system architecture to minimize dependence on constant communication with the remote server.

Despite the robustness of the tests, this work acknowledges its methodological limitations, as the scope focused on the market-standard relational database paradigm, under entry-level hardware configurations, and used static-sized data packages. It is suggested that future research replicate this model using non-relational technologies, investigating whether eventual consistency mitigates latency effects. It is also recommended to quantify the labor cost required for the modernization of legacy systems, integrating this value into the overall calculation of migration projects. Without this modernization, on-premises infrastructure consolidated as the most viable strategy for legacy systems, ensuring high performance, budgetary predictability, and data sovereignty.

Bibliographic References

DAPPERLIB. Dapper: a simple object mapper for .Net. [S. I.]: Stack Exchange, 2024. Disponível em: https://github.com/DapperLib/Dapper. Acesso em: 2

Holmnäs, F. 2024. Migration of Customer Data from an On-premise Database to Cloud Infrastructure. Dissertação (Mestrado em Engenharia de Computação) – Faculty of Science and Engineering, Åbo Akademi University, Finlândia. Disponível em: https://www.doria.fi/bitstream/handle/10024/190527/holmnas_fredrik.pdf?sequence=2&isAllowed=y. Acesso em: 25 jan. 2026.

Husain, M. E.; Hussain, I.; Tanweer, S.; Khan, I. R. 2024. Transitioning from Data Centers to Cloud: An In-depth Analysis of Microsoft SQL Server’s Role in DBaaS and On-Premise Solutions. ICIMMI 2023: Proceedings of the 5th International Conference on Information Management & Machine Intelligence. Disponível em: https://dl.acm.org/doi/10.1145/3647444.3652491. Acesso em: 31 mar. 2026.

Kumar, S. 2024. Cloud vs. on-premises: choosing the right data architecture for scalable, secure solutions. International Journal for Multidisciplinary Research 6(6): 1-9. Disponível em: https://www.researchgate.net/publication/394789475_Cloud_vs_On_Premises_Choosing_the_Right_Data_Architecture_for_Scalable_Secure_Solutions. Acesso em: 25 jan. 2026.

Noor, I.; Tariq, S.B.; Shabbir, A.; Aksa, M. 2024. Into the future with cloud: a comparison with on-premises data warehouse. IEOM Society International, Dubai. Disponível em: https://ieomsociety.org/proceedings/2024dubai/541.pdf. Acesso em: 25 jan. 2026.

Pillai, P. 2024. Cloud vs. On-Premise Data Warehousing: A Strategic Analysis for Financial Institutions. Journal of Computer Science and Technology Studies 6(1): 1-12. Disponível em: https://www.al-kindipublisher.com/index.php/jcsts/article/view/9396. Acesso em: 31 mar. 2026.

Xu, D.; Li, T.; Sun, Z.; Chen, Z.; Zhou, W.; Zhang, Y.; Lu, W.; Du, X. 2025. Performant Synchronization in Geo-Distributed Databases. arXiv. Disponível em: <https://arxiv.org/abs/2511.22444>. Acesso em: 17 ago. 2026.

Zhang, G.; Tang, X.; Chang, Q.; Zhang, H.; Hwang, K.; Li, Y.; Huang, R.; Wang, T.; Zhang, W.; Zhang, M.; Chen, Q.; Hou, X.; Wang, Q. 2025. A Low Latency Cache for Cloud RDBMs. VLDB 2025 Workshop: DATAI. Disponível em: <https://www.vldb.org/2025/Workshops/VLDB-Workshops-2025/DATAI/DATAI25_3.pdf>. Acesso em: 17 ago. 2026.

Article originating from the Final Course Work of the Specialization in Software Engineering of the MBA USP/Esalq

To learn more about the course, click here and access the MBX Academy platform

You may also like

Software Engineering

October 09, 2026

O papel da densidade de texto instrutivo na eficiência de uma aplicação web.

O desenvolvimento de aplicações web se conecta à experiência do usuário, e este trabalho investigou como o uso excessivo de textos instrutivos pode retardar a conclusão de tarefas e impactar a eficiência da aplicação. O objetivo foi identificar o impacto da densidade textual do conteúdo instrutivo na eficiência de uma aplicação web, utilizando como principal referência a terceira lei de usabilidade de Krug. A pesquisa, de caráter exploratório e delineamento experimental quantitativo, empregou um teste A/B em uma aplicação web responsiva, onde a única variável controlada foi a densidade textual (alta vs. baixa, definida pela contagem de palavras). Participaram 25 usuários, e os dados foram coletados via Datadog RUM, mensurando tempo de conclusão, erros de submissão e taxa de conversão. Os resultados revelaram que a variante com densidade textual reduzida (variante B) apresentou uma taxa de conversão superior (58,3% contra 33,3% da variante A) e um tempo médio de conclusão significativamente menor (1:38 minutos contra 4:58 minutos da variante A), representando um aumento de 67,12% na eficiência. O teste t de Welch (p=0,042) confirmou que a redução da densidade textual impactou a eficiência. Concluiu-se que a redução da densidade textual afeta a eficiência e a taxa de conversão, reforçando a importância de conteúdo objetivo e conciso. Contudo, a baixa densidade textual, por si só, não garantiu o pleno entendimento, sendo essencial a comunicação clara e objetiva das instruções, validando a relevância do UX Writing.

Palavras-chave: Eficiência; Experiência de usuário; Teste A/B; Texto Instrutivo; Usabilidade.

Software Engineering

October 09, 2026

Implementation of analytics systems in automation environments in the process industry

The digitalization of process plants depends on structured data collection and storage, without which there is no operational visibility. Industrial analytics is the name given to the chain that processes this data from signal acquisition in the field instrument, through time-ordered storage, to its availability for analysis and other systems. Proprietary industrial software currently covers this chain, and licensing and maintenance costs restrict its adoption. The study aimed to architect, implement, and validate a system of this nature, employing current software development techniques and components at no licensing cost. The system, named Sistema de Aquisição e Tratamento de Informações de Processo (SATIP) [Process Information Acquisition and Processing System], was structured in three independent layers, using a pharmaceutical reactor simulator as a data source. The simulator executed an eight-step recipe and generated time series with stochastic variation. The simulator was written in Go, a language adopted for generating self-contained binaries, suitable for execution on edge equipment. Storage used TimescaleDB, and data was made available through a REST interface with a web dashboard. The system processed approximately 414,000 records per execution, with multivariate trends, alarm logging, and temporal correlation between instruments. The architecture proved to be reproducible and without licensing costs, and the identification of degradation required only the readings already stored by the system.

Keywords: Software architecture; Digitalization; IT/OT integration; Predictive maintenance; Time series.

Software Engineering

October 09, 2026

Use of voice in conjunction with large language models as a tool for digital accessibility

Speech is an essential basis for human interaction, and for people with disabilities, it can represent the primary form of communication with the external environment. Given the growing technological influence, voice command identification has emerged as a promising strategy for human-machine interaction. The work explored how voice, in conjunction with Large Language Models (LLMs), can be used efficiently, naturally, and accurately. For this purpose, various design patterns and the Python language were employed, aiming for greater extensibility. Gemini was used as the LLM provider, sending audio directly and leveraging its function calling capability to interact with the device. A system was developed capable of understanding user intent and converting it into actions, whose differential was the computer vision capability based on screenshots and a mesh system for LLM guidance. Tests revealed a user intent comprehension rate of 91.81% and a success rate in execution of 75.45% with the “Flash-3” model (top p 0.5 and top k 5). The proposed system validated the premise that the integration of LLMs into voice interfaces increases the autonomy of users with motor disabilities, fulfilling the purpose of being a modern and effective Assistive Technology. However, questions were raised about the costs of AI and user data security, indicating the need for improvement.

Keywords: Function calling; Human-machine interaction; Voice recognition; Computer vision.

Software Engineering

October 09, 2026

ReasonGuard: A Reasoning Audit Platform for Language Models Based on Structured Thought Decomposition

ReasonGuard was presented, an Artificial Intelligence reasoning auditing platform, developed as an observability middleware between client applications and large language models (LLMs). The work aimed to offer transparency and traceability for AI-based decision-making processes, addressing the gap in operationalizing structured reasoning techniques for auditing purposes. The system intercepted, analyzed, and documented interactions with LLMs through five modules based on the Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT) paradigms. These modules captured reasoning trails, detected structural logical flaws, evaluated response consistency, and generated audit reports targeted at different stakeholder profiles. The platform was implemented with FastAPI (Python) on the backend, React with TypeScript on the frontend, and PostgreSQL as a relational database, following a modularized architecture with multitenancy isolation per user. The results demonstrated the technical feasibility of the proposed approach, with all five modules operational and integrated. It was concluded that ReasonGuard contributes to AI governance by instrumentalizing structured reasoning paradigms as observability and auditing tools, filling a gap in the literature.

Keywords: AI Auditing; Chain-of-Thought; AI Governance; Large Language Models; Algorithmic Transparency.

Software Engineering

October 09, 2026

EngTT: software for road freight based on operational costs and ANTT parameters

Road transport represents the main logistics modality in Brazil, with the composition of freight costs regulated by the National Land Transport Agency (ANTT). Given the absence of a structured methodology for freight calculation and the dependence on isolated spreadsheets in the sector, the EngTT software was developed. The objective was to create a decision support tool that integrated operational variables, calculated the total cost per route, and compared the results with the ANTT’s minimum floor, identifying non-compliance and margin compression scenarios before pricing. The adopted methodology was applied, quantitatively and experimentally, with incremental construction oriented towards the Minimum Viable Product (MVP) concept. The system was structured in three modes of use – Build Visual Route, Batch by spreadsheet, and Scrape routes – sharing a decoupled calculation engine and centralized parameters. The results obtained demonstrated that the software met the proposed objective, showing margin variations and regulatory compliance between the analyzed routes. It was observed that longer routes, such as Ribeirão Preto × Guarujá, presented an increase in cost due to the need for additional driver per diems, while shorter routes, such as Cajamar × Guarujá, showed greater adherence to the ANTT’s regulatory floor. It was concluded that the developed solution is applicable to the road freight transport sector as an effective tool for pricing and operational management.

Keywords: ANTT; Containerized cargo; Software engineering; Road freight; Python.

Software Engineering

October 08, 2026

Data pipeline for stock monitoring considering the fundamentalist methodology

The Brazilian financial scenario faces challenges such as family indebtedness, low financial literacy, and the decentralization of information for investment analysis. Given this, the research objective was to develop a financial data pipeline, based on good Data Engineering practices, to structure, process, and make relevant information available to support individual investors’ decision-making in stock analysis. A medallion architecture was implemented, using MinIO S3 for bronze, silver, and gold layer storage, with Delta Lake for data governance. Apache Spark was employed for distributed processing, Prometheus for monitoring, and Power BI for analytical visualization. The data were mostly financial statements from the Securities and Exchange Commission (CVM). A theoretical investment portfolio was built based on Benjamin Graham’s (2017) principles, applying selection filters and data validation. The results indicated a positive return of 32.88% for the theoretical portfolio, outperforming the Ibovespa index (4.25%) in the period from 2021 to 2024, although lower than the Selic rate (46.96%). Qualitatively, the pipeline processed voluminous datasets, with significant reductions in redundancies, such as 99.27% in the BPA table and 94.29% in the DRE, after applying filters. Gains in data organization, traceability, and quality were evidenced, enabling structured and more robust financial analyses.

Keywords: Medallion Architecture; Data Engineering; Investments; Data Pipeline.

Software Engineering

October 08, 2026

Asynchronous Slack-Jira integration via message queue “middleware”: comparison of “cloud computing” solutions

The latency and interoperability between distributed corporate systems constituted the problem investigated, motivated by the costs and fragilities of manual integrations between collaboration and project management platforms. An asynchronous integration “middleware” between Slack and Jira was developed and validated, with the objective of reducing the perceived user response time and ensuring system stability under load. The methodology consisted of building an event-driven software architecture in Python, using the “producer-consumer” and “adapter” patterns to isolate the user interface from “backend” processing. The solution evolved into a cloud-agnostic architecture, based on “serverless” functions, and was subjected to stress tests in local, real network, and production environments on two clouds. The architecture reduced user waiting time from a synchronous estimate of 2,000 milliseconds to a local average of 8.96 milliseconds. Under a load of 50 simultaneous requests, both cloud providers proved viable: Amazon Web Services registered lower average latency in the reception layer (951.35 ms) and double the throughput, while Microsoft Azure executed background processing with a median of 113 ms. The application of “optimistic UI” ensured fluidity of use, and load leveling by queues eliminated the need for infrastructure over-provisioning. The “middleware” consolidated itself as a scalable, resilient, and protected corporate reference model against technological lock-in.

Keywords: Event-driven architecture; Temporal decoupling; Operational efficiency; Information technology service management; System interoperability.

Software Engineering

October 05, 2026

Using Generative AI for Database Selection: A Requirements-Driven Framework for Generating Architecture Decision Records (ADRs)

The growth of data-intensive applications and the adoption of microservices architecture have amplified the need for polyglot persistence, imposing a high cognitive load on software architects in choosing and justifying database technologies. This work aimed to propose and develop an Artificial Intelligence (AI) agent-based framework to guide technological selection and generate well-founded Architecture Decision Records (ADRs). An experimental and applied methodology was employed to build a technical knowledge base. The Retrieval-Augmented Generation (RAG) technique, along with the LangChain and LangGraph libraries, was used to orchestrate agents and anchor the responses of a Large Language Model (LLM). The framework extracted natural language requirements, enriched them with RAG, and sent them to the LLM, which generated ADRs to assist in evaluating theoretical trade-offs. The results demonstrated that the agent with RAG reduced generic responses, increasing theoretical grounding and traceability. The RAG approach proved its effectiveness against conventional prompts (zero-shot), favoring the generation of ADRs with a lower level of hallucination and a high level of theoretical traceability. It was concluded that the automated tool fulfilled the function of requirement mapping, resulting in empirically grounded technical documents and aiding governance and decision-making in software architecture.

Keywords: Databases; Artificial Intelligence; LangGraph; LLM; RAG.