Software Engineering
October 08, 2026
Performance and Total Cost of Ownership of Cloud Databases and On-Premises Infrastructure
Performance and Total Cost of Ownership of Cloud Databases and On-Premises Infrastructure
José Antônio Elias da Silva Júnior; Manoel Flavio Leal
DOI: 10.22167/2675-6528-202603055
Article derived from a Course Conclusion Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.
Summary
This study analyzed the performance of database operations in cloud and local infrastructures, correlating technical behavior with Total Cost of Ownership projection. The research was characterized as a quantitative and experimental study, in which a test system subjected isolated database instances to progressive execution loads, measuring the impact of network latency and computational consumption. For financial analysis, an investment and operational expense model was developed, diluted over a thirty-six-month cycle. Technical results revealed that the accumulation of internet latency caused severe time degradation in cloud executions, despite the remote infrastructure operating with high processing idleness, registering almost 98% CPU inactivity. In contrast, the local environment achieved superior performance supported by almost instantaneous network communication. Financially, cost consolidation demonstrated an empirical tie between the physical acquisition model and the service subscription model within a three-year horizon, but the local environment proved more advantageous in a sixty-month cycle. It was concluded that the degradation in the cloud was not due to computational capacity, but to the interaction between route latency and the application’s unitary communication pattern. Cloud adoption requires deep optimization of the system architecture to minimize dependence on constant communication with the remote server. Without this modernization, the local infrastructure consolidated itself as the most viable strategy, ensuring high performance, budgetary predictability, and data sovereignty.
Keywords: Operating expenses (OPEX); Transactional scalability; Network latency; Legacy systems.
1. Introduction
The evolution of information technology has been marked by a significant transition in infrastructure models. Historically, organizations relied on on-premises architectures, which, despite being widely used, presented inherent complexities and high costs associated with maintaining their own infrastructure. More recently, the cloud computing paradigm has emerged as a scalable alternative for outsourcing the management of this infrastructure.
The decision between adopting a cloud infrastructure or maintaining an on-premises structure has become a central dilemma for modern software companies, requiring the re-evaluation of traditional repositories and architectures in the face of new technological demands (Noor et al., 2024). From a financial perspective, studies such as Kumar (2024) indicated that migrating to the cloud can generate savings of 30% to 40% in infrastructure costs in specific scenarios. However, for applications sensitive to network response time (latency), on-premises implementations have shown to achieve up to 42% superior performance (Kumar, 2024), creating a complex trade-off between cost and speed.
In operational practice, the pursuit of the best cost-benefit ratio reveals that the choice between models goes far beyond simple server allocation. Maintaining on-premises infrastructure requires a high initial investment (capital expenditure), but can prove financially advantageous and predictable in the long term. Furthermore, maintaining one’s own servers guarantees total sovereignty over data traffic, a crucial factor in high-volume operations such as massive database conversions during system migration (Holmnäs, 2024). On the other hand, many organizations have failed when migrating to the cloud by ignoring the “hidden costs” of performance. A legacy application, originally optimized for local networks, can become prohibitively expensive in the cloud due to the need to increase the consumption of computational resources (commercially measured in vCores or DTUs) simply to compensate for code inefficiencies and network communication latency.
Recent research in systems architecture has shifted the focus from processing power to the cost of communication between components. Zhang et al. (2025) demonstrated that, in databases with disaggregated storage, remote access latency becomes the dominant degradation factor. Xu et al. (2025) identified synchronization between geographically distributed nodes as the main bottleneck in long-distance networks. Comparative literature between cloud and on-premises infrastructure, however, focuses either on cost aggregation (Pillai, 2024) or on the performance of isolated infrastructure (Husain et al., 2024), without separating the computational effort of the database engine from the waiting imposed by the application’s communication pattern. It thus remains undetermined whether the slowness attributed to the cloud stems from the remote location of the database or the access architecture, a distinction with opposite consequences.
It is evident, therefore, that there is no definitive or generalizable solution for all companies; each context requires careful analysis. Given this scenario, this work is justified by the need to provide empirical data to assist managers in directing their technological strategies assertively. The present study aims to analyze the capacity and performance of database operations in both infrastructures, correlating these technical metrics with the Total Cost of Ownership. The central purpose is not to point to one technology as absolutely superior, but rather to explore its advantages and disadvantages based on empirical data, culminating in the development of a decision matrix that assists managers in directing their technological strategies assertively.
2. Material and Methods
The present study was characterized as a research with a strictly quantitative approach and experimental design. The central objective of the experiment was to subject database infrastructures, allocated in distinct physical media, to progressive transactional loads. The impact of network latency and the internal processing capacity of each environment were measured, aiming to correlate these technical metrics with the Total Cost of Ownership, according to the established general objective.
To ensure the empirical validity of the study, the design focused on isolating variables. It was ensured that the computational effort of the database engine could be chronometrically separated from the time required for packet transport through the public internet and the local network. The experiment’s architecture was strictly segregated, separating the development environment from the execution environment, and the test topology consisted of three independent and isolated instances.
The first instance operated as a Client Machine, configured as a dedicated virtual machine. It ran the “Windows 10” Operating System, provisioned with 4 GB of RAM and 2 logical processing cores, based on the AMD Ryzen 5 5500U architecture. This environment was exclusively and solely responsible for executing the performance measurement application, without the presence of heavy development tools in the background.
The second instance represented the corporate physical environment, named “On-Premise”. It was configured as a virtual machine under the “Windows Server” operating system, possessing hardware specifications identical to the client machine, dedicated to hosting the “SQL Server 2019 Express” database engine. To ensure the fidelity of the local scenario, the virtual network adapters of both machines were configured in bridge mode (“Bridge”).
With this parameterization, each instance operated autonomously in the infrastructure, receiving independent IP addresses assigned by the physical router. Communication occurred through real TCP/IP protocol packet traffic over the local area network (LAN). This isolation prevented the use of loopback connections, which would invalidate the precise measurement of latency characteristic of a physical corporate network.
The third instance of the topology represented the remote cloud environment, named “Cloud”. It was constituted by a database “Azure SQL Database”, provisioned in the virtual core (“vCore”) based purchase model, allocated under the General Purpose service tier. The computation was configured in “Serverless” mode, using the “Standard-series (Gen 5)” structural hardware generation.
The dynamic scaling of processing was bounded to operate with a minimum of 0.5 and a maximum of 2 “vCores”. This configuration was intentionally chosen to represent an input baseline, reflecting the initial scenario of adoption by real companies seeking concept validation or cost elasticity.
The stress testing engine consisted of a console application developed in C# and compiled on the .NET 8.0 platform. Access to data and transactional communication between the client application and the two database instances were mediated by the low-level object-relational mapper Dapper (DAPPERLIB, 2024).
The choice of Dapper was justified by its very high performance. By directly mapping native instructions to objects in memory, Dapper eliminated the overhead of query translation, adding the least possible processing time to data conversion.
A critical methodological factor for the experiment’s validity was the rigorous standardization of the data mass and network traffic packets (“payloads”). The structure of the tables and the synthetic data generated in memory were scaled with fixed sizes and strictly identical character padding. To mitigate this risk, the data generation library “Bogus” (BOGUS, 2024) was used.
This approach ensured that every TCP/IP packet sent from the client to the server, regardless of whether it was in the local environment or in the cloud, had exactly the same weight in bytes. This isolated route latency as the sole temporal variable of the experiment, eliminating any possibility of bias in the results due to discrepancies in the volume of information exchanged in each iteration.
To avoid ambiguity in reading the results, a strict operational definition of the term transaction was adopted. Each transaction corresponded to a single command submitted individually to the bank with automatic confirmation, which is equivalent to a single network round trip. Explicit transactional blocks grouping multiple commands were not used.
Thus, “a load of 128 transactions” designated 128 independent commands, submitted sequentially, each awaiting confirmation from the previous one. This pattern is predominant in legacy record-oriented systems, which constituted the object of this research. The composition of each operation, including the submitted command, the relational validations required from the engine, and the number of round trips per load unit, was detailed.
In all operations, the description field was filled with exactly 5,000 characters, so that the trafficked package had identical weight in both environments. In modifications and deletions, the prior query was executed only once per batch, outside the timer, so that the measured time strictly corresponded to the unit writes.
To simulate the processing effort consistent with real-world business applications, the main table (“TransacaoFinanceira”) was modeled with strong relational bindings. The algorithm generated numeric identifiers that act as foreign keys for the User, Category, and Bank tables. This modeling ensured adherence to ACID properties (Atomicity, Consistency, Isolation, and Durability).
The testing strategy evaluated the three fundamental database persistence operations: Insertion, Modification, and Deletion. To simulate real operational scenarios and maintain test fairness, the modification and deletion operations were not executed blindly, following the “fetch and process” architectural pattern.
Firstly, the application performed a complex query that deterministically returned the identifiers to be manipulated. Then, the submission of alteration requests occurred unitarily within loops. The test batteries followed a geometric growth in powers of base two.
Each cycle started execution with a single unit operation, successively doubling the load (2, 4, 8, 16, 32, and 64), until reaching the saturation limit stipulated at 128 sequential transactions. For the orchestration of data collection and statistical extraction, the “BenchmarkDotNet” precision library (BENCHMARKDOTNET, 2024) was used.
Each geometric scenario was fully repeated four times for each operation, attesting that the infrastructure behavior did not have a punctual or random character. A vital architectural precaution was adopted regarding the “Cold Start” phenomenon, intrinsic to “Serverless” architectures and to the Just-In-Time compilation of .NET.
Before the official timer of each battery was triggered, the global preparation method of BenchmarkDotNet executed natural warm-up routines. These routines consisted of pre-opening database connections, performing count queries, and replenishing exhausted data masses. This control ensured the dynamic allocation of vCores in the cloud, the establishment of active connection pooling, and the uploading of indexes to RAM.
With this mitigation, it was ensured that the final telemetry strictly captured the cost of transport latency, exempting the measurement of the cloud infrastructure’s wake-up lethargic time. To ensure experiment reproducibility, the software architecture developed for measuring response times was isolated into specific classes, using the attributes of the BenchmarkDotNet library.
The update class, named “UpdateBenchmark”, was developed to measure response times, incorporating attributes from the BenchmarkDotNet library for data mass control and mitigation of the Cold Start phenomenon. The global preparation method verified the count of active transactions in the database before the official timer.
The execution of database instructions was delegated to the Dapper package, with the parameterized update instruction submitted atomically by the “Execute()” command, confirming the “row-by-row” test model. The synthetic data generation algorithm, implemented with the Bogus library, was not purely random but contained within restrictive parameters. The description field was instructed to create alphanumeric sequences with exactly 5,000 characters, ensuring that every TCP/IP packet had the same byte weight.
To ensure the repetition of exclusion batteries while maintaining statistical rigor without the need to recreate the infrastructure at each cycle, a differential reset algorithm (“ReporMassaDiferencial”) was developed, integrated into the “benchmark” library lifecycle. Before starting the timer to measure the exclusion time, the “Setup” method performed an active record count.
When it was identified that the table volume was below the pre-established threshold of ten thousand records, the system dynamically calculated the exact difference and requested the repository to insert the missing quantity. This mechanism operated as a state stabilizer, ensuring that all validation rounds of the deletion operation found the database in the same volumetric and index fragmentation conditions.
For the financial analysis, a Total Cost of Ownership (TCO) projection was developed for an initial lifecycle of thirty-six months. The choice of this period was based on standard accounting practices for the depreciation of technology assets. The mathematical methodology of Equivalent Monthly Cost was used, which dilutes capital investment over the months of operation, adding it to recurring operational expenses.
The composition of initial investment and operational expenses, including hardware infrastructure costs, electricity tariffs, cloud service pricing, and real estate rental, was based on market data (DELL, 2026; ENEL, 2026; MICROSOFT, 2026; QUINTOANDAR, 2026).
3. Results and Discussion
The analysis of the empirical data collected during the experiment revealed crucial insights into the performance and cost of database operations in cloud and on-premises infrastructures. The central objective was to go beyond the mere presentation of metrics, promoting a critical discussion on how the choice of infrastructure impacts the transactional behavior of the database and, consequently, the financial and architectural viability of projects. The results were structured to progressively answer the initial hypotheses of the study, confronting the laboratory findings with recent academic literature to validate the proposed conclusions.
The technical evaluation was carried out through a rigorous telemetry process, which captured response time and physical resource consumption under progressive request loads. By subjecting the database engines to execution cycles ranging from unit transactions to intensive batches of 128 operations, it was possible to isolate the actual computational effort of the data engine from the time required for packet transport. This segregation was fundamental to identifying whether the bottlenecks were physical hardware limitations or restrictions imposed by the network topology, allowing for a deeper understanding of the observed phenomena.
Transactional Latency Analysis in Cloud Environment
The technical performance evaluation in the cloud environment, subjected to unitary and progressive loads (from 1 to 128 transactions), revealed a notable empirical phenomenon: high network latency obscured the processing time differences that would normally exist between Insert, Modify, and Delete operations. The cloud environment presented virtually overlapping time curves for these operations. This behavior indicated that the internal processing of the database engine consumed a negligible fraction of the total time compared to the remote communication time, resulting in a delay close to 2,600 milliseconds at the maximum load of 128 transactions.
Almost the entire time measured in this scenario resulted from packet transit (round trip) on the public network. This empirical behavior is strongly supported by contemporary literature, where Cho’ponov (2026) identifies a structural and unavoidable “latency rate,” stipulated between 10 and 50 milliseconds per network hop. Verbitski et al. (2017) also point out that the central restriction for high-throughput processing has shifted from computation and storage to the network. Dividing the total time of the experiment’s maximum load (approximately 2,500 milliseconds) by the 128 submitted iterations, an average latency of 19.5 milliseconds per operation in the cloud was observed, a value that aligns with the critical degradation range predicted for remote environments.
A technical factor that actively contributed to the observed time stacking was the nature of the communication protocol used by the database engine, the “Tabular Data Stream” (TDS). Husain et al. (2024) point out that, in Microsoft SQL Server solutions operating as a service, each unit request requires an acknowledgment before the next instruction in the code is processed. On local networks, this round-trip time is negligible, but in the cloud, the time spent on packet encapsulation and long-distance route transposition becomes the majority component of the operation, overshadowing the platform’s actual processing capacity.
From the perspective of distributed software engineering, Yang et al. (2023) explain that this slowness is not a configuration failure, but rather the architectural price required to maintain strong data consistency. By the CAP Theorem, when partitioning the system across the internet, the remote database must ensure that each write instruction is consolidated across storage nodes before successfully responding to the client application. Consequently, the system code is forced to wait for transit confirmation at each loop iteration, generating the severe temporal accumulation effect observed.
Transactional Latency Analysis in Local Environment (“On-Premise”)
In contrast to the remote scenario, the analysis of the local environment corroborated the research premise from an inverse perspective. Free from the long-distance traffic penalty, the local infrastructure not only responded substantially faster, finishing the maximum batch of 128 transactions in the range of 250 to 410 milliseconds, but also revealed the true computational cost of each operation. In the isolated scenario of almost zero network latency, it became visually possible to notice the divergence of the Exclusion curve from the Insertion curve, a behavior that was previously masked by exogenous waiting on the internet.
The visibility into the internal effort of the database engine is validated by Cho’ponov (2026), who argues that the physical proximity between the application and the server eliminates the additional network load, allowing the infrastructure to deliver its full nominal performance. While in the cloud latency consumes most of the runtime, in the local environment the measured time is a faithful representation of algorithmic efficiency and disk write speed. As observed in the experiment, the absence of external hops allowed for a reduction of approximately 90% in the total runtime for the load of 128 transactions.
Additionally, Husain et al. (2024) reinforce that local Microsoft SQL Server solutions offer processing predictability that is difficult to replicate in database-as-a-service environments without massive investments in redundant hardware layers. The collected data demonstrate this stability, in which the growth of response time occurs linearly and strictly proportionally to the increase in load. Fluctuations and peak waits characteristic of public internet traffic are not observed, ensuring constant execution regardless of external connectivity factors. This factor is identified by Pillai (2024) as critical for sectors operating with massive data volumes and requiring low operational latency as a non-negotiable business requirement.
Global Comparative Analysis and Impact on User Experience
The magnitude of the penalty imposed by the unit request model on cloud infrastructure becomes categorically evident when consolidating the global average of operations. An immediate divergence of performance curves is observed right in the first load cycle. In the unit transaction (1 cycle), the local architecture processed the request in an average of less than 7 milliseconds, while the cloud demanded approximately 50 milliseconds, an initial difference that establishes the efficiency benchmark of each environment. This divergence is not an isolated event, but the empirical materialization of what Cho’ponov (2026) defines as “latency rate”.
The author argues that the transformation of simple internal system calls into external network requests inserts a temporal cost that progressively accumulates with each new network hop. As the code submitted for testing executes repetitive iterations, this initial latency of 50 milliseconds multiplied throughout the execution loop, resulting in the performance abyss observed at the load limit of 128 transactions. From a strategic and business perspective, this temporal degradation directly compromises the “user-perceived performance” (Järvinen, 2025), which highlights end-user dissatisfaction when the software architecture requires multiple round trips to the remote server to complete a single logical task.
The data validate the thesis that migrating an application to the cloud without proper refactoring to minimize constant database interaction will result in chronic slowness, regardless of the contracted processing power. This scenario accurately illustrates the risk associated with the migration strategy known as “Lift-and-Shift”. Kansara (2024) warns that simply moving on-premises systems to the cloud, without prior adaptation of the communication architecture, often turns the elasticity of cloud infrastructure into an unsustainable operational bottleneck. Legacy code that assumes the latency of a local network, when operating on the public internet, makes the system technically inefficient and financially risky.
CPU Utilization and Idle Analysis
Beyond temporal degradation, the technical and financial viability of a database architecture is directly linked to the efficiency in resource consumption, especially computational processing. Telemetry during the execution of the maximum load (128 transactions) revealed that the local environment engaged an average of 14.2% of its processor capacity, while the remote cloud server operated with an idle rate close to 98%, consuming only 1.3% of CPU. This demystifies the hypothesis that the cloud’s slowness would be associated with an exhaustion of its computational capacity.
The consumption of only 1.3% of CPU in the cloud proves that the database resolved the write or read request in milliseconds, and spent the rest of the test’s 2.5 seconds in a state of absolute inactivity, just waiting for the next packet to arrive via the internet. From the perspective of software engineering in the Microsoft SQL Server architecture, this underutilization reveals that the platform spent most of its runtime stuck in the wait event technically known as “ASYNC_NETWORK_IO”. This event occurs when the database engine finishes internal processing almost instantly but is forced to suspend operations because the client application, separated by public network latency, is slow to consume the data or send the next instruction.
This finding corroborates the analyses by Husain et al. (2024), who observe that hardware provisioning in database-as-a-service (DBaaS) solutions often becomes underutilized due to severe external network input/output (I/O) bottlenecks. In contrast, the local infrastructure, with negligible latency on the local area network (LAN), allowed packets to arrive uninterruptedly, causing the database to work in a continuous flow. Without the “ASYNC_NETWORK_IO” block, the physical server finalized the entire block of operations quickly, justifying the full use of the acquired resource and demonstrating superior architectural efficiency for sequential request systems.
From a financial perspective, this cloud idleness translates into a severe hidden cost. Brown (2025) highlights that serverless architectures, such as the Azure SQL Database edition used in this experiment, are marketed with the promise of granular billing and high financial efficiency. However, the author warns about the trade-offs generated when code is not optimized for this billing model. As the cloud bill is calculated by active computation time (billed per second), submitting single iterations forces the database to keep the session open for long periods to resolve a minuscule workload. Brown’s (2025) thesis is therefore validated: the on-demand billing model becomes highly burdensome if the application does not use batch processing, turning network-induced idleness into a direct financial liability for the organization.
Decomposition of Time between Processing and Waiting
Once the performance difference was established, the research sought to answer a fundamental causal question: does the degradation stem from the bank’s remote location or the application’s communication architecture? To resolve the issue, the total time was decomposed into two mutually exclusive parts: the time the engine was effectively processing, obtained by the product of the total time and the average processor utilization percentage measured by telemetry, and the residual time waiting for communication. Under a load of 128 transactions, the effective processing time remained in the same order of magnitude in both environments, with 44.0 milliseconds on the local server versus 33.2 milliseconds on the remote instance.
The result is conclusive: in the computational work performed, the cloud database was no slower than the local one. The entirety of the difference is concentrated in the waiting portion, which increased from 266.0 milliseconds in the local environment to 2,516.8 milliseconds in the cloud. It is concluded that the remote infrastructure does not, in itself, constitute the cause of the degradation. The determining factor is the interaction between the route latency, which fixes the cost of each round trip, and the application’s unitary communication pattern, which determines how many round trips occur. Latency acts as a multiplier and the architecture as an amplifier; in isolation, neither explains the result.
A second finding reinforces this interpretation: the ratio between cloud and local environment times remained practically constant across the entire scale, from the unit transaction to the batch of 128. If the degradation were due to exhaustion of the contracted capacity, an increasing ratio would be expected, with progressive divergence of the curves. The constancy indicates a multiplicative and structural penalty, proportional to the number of round trips, and not resource saturation. The decomposition allows estimating the effect of an architectural change that preserved the contracted infrastructure: by replacing the 128 unit round trips with a single batch submission, the expected time in the cloud would be on the order of 53 milliseconds, compared to the approximately 2,550 measured. This is an analytical projection derived from the collected data, and its order of magnitude supports the thesis that the cloud was not slow, but the application was conservative.
The finding dialogues with recent literature. Zhang et al. (2025) observed that, in disaggregated storage, remote access latency overlaps with processing capacity, penalizing writes most of all, a behavior identical to that measured here. Xu et al. (2025) demonstrated that the bottleneck of geographically distributed banks resides in synchronization over long-distance networks, not in the processing of each node. The convergence indicates a structural phenomenon of distributed architectures, where network latency and application communication patterns interact to determine performance.
Load Regimes and Contextual Superiority of Each Architecture
Once the origin of the degradation was established, it became possible to qualify in which contexts each architecture is superior. Decomposition allows any workload to be classified according to the dominant portion of its runtime. A network-bound regime is one in which waiting for communication exceeds effective processing, observed in both environments of this experiment. In this regime, performance is governed by the product of round-trip latency and the number of round trips, and contracting more capacity does not yield appreciable gain, as the additional resource remains idle waiting. This regime is characteristic of unitary and legacy transactional systems, where local infrastructure is recommended because time is governed by the number of round trips.
A processing-limited regime is one in which computational effort exceeds communication time. Analytical workloads and batch processing fall into this category, as they amortize a single network latency over a large volume of work. Here, cloud elasticity becomes an effective advantage, and the local environment is penalized by the rigid ceiling of acquired hardware. The practical implication is that infrastructure choice should not precede the characterization of the application’s predominant regime. A decision matrix was developed to consolidate this analysis, recommending the cloud for batch processing, analytical workloads, seasonal demands with pronounced peaks, and applications refactored for bundled submission, where a network latency is amortized over a large volume or asset idleness is avoided.
The decision matrix also indicates the local environment for constant volumes in long cycles and the requirement for data sovereignty, as internal traffic is not charged and no third parties are involved. The matrix highlights that no architecture is absolutely superior; superiority is contextual and determined by the workload regime, a premise already supported by Tan et al. (2019) who demonstrated that the choice of a cloud database depends on the match between the architecture and the workload profile, rather than a single performance ranking. The recurring mistake in migrations consists of transposing to the cloud an application limited by network without altering its communication pattern, paying for elasticity without being able to consume it, which converges with Kansara’s (2024) warning regarding the direct transfer of legacy systems.
Total Cost of Ownership [TCO] Decomposition
To quantify the financial impact of the architecture choice and materialize the effects of latency, a Total Cost of Ownership (TCO) projection was developed for an initial lifecycle of thirty-six months. The choice of this period is based on standard accounting practices for the depreciation of technology assets. To enable an equitable comparison between the on-premises physical infrastructure acquisition model and the cloud services subscription model, the mathematical methodology of Equivalent Monthly Cost was used, which dilutes capital expenditure (CAPEX) over the months of operation, adding it to recurring operational expenses (OPEX).
The projection of the Equivalent Monthly Cost revealed a technical tie and strict financial parity in the three-year scenario, with the local environment costing R$ 2,233.30 and the cloud subscription R$ 2,216.00. This empirical finding contradicts much of the common sense in the technology market, which often associates the adoption of cloud computing with a drastic and automatic reduction in infrastructure expenses. Leis and Kuschewski (2021) demonstrate, in the same vein, that performance and cost in the cloud vary by orders of magnitude depending on the contracted configuration, so that there is no economy dissociated from scaling.
Pillai (2024), when analyzing data storage and processing strategies in financial institutions and traditional sectors, argues that the costs of local infrastructure, although requiring a high initial capital investment, are amortized in a highly predictable and secure manner in the long term. The data validate Pillai’s thesis, demonstrating that, by judiciously diluting electricity, real estate, and physical deployment expenses, the architectural decision depends not on supposed immediate savings in the cloud, but rather on the maturity of the application code and the rigor in controlling the monthly budget. The initial investment costs for the local environment (equipment and licenses) were R$ 24,399.00 and deployment (physical and logical) of R$ 9,200.00, with a recurring monthly operational cost of R$ 1,300.00. For the cloud, there was no initial investment, but the recurring monthly operational cost was R$ 2,216.00.
Scenario Study, Scalability and Long-Term Projection
To deepen the validation of the financial model and test the resilience of both architectures, a scenario study was designed, extending the infrastructure’s lifecycle to 60 months (five years), a timeframe frequently adopted by the corporate market for total server park renewal. By recalculating the dilution of the initial local investment (R$ 33,599.00) over this new period and adding the recurring operational cost (R$ 1,300.00), the Equivalent Monthly Cost of the physical environment drops drastically to R$ 1,859.98. In contrast, the cloud service subscription model maintains its static floor of R$ 2,216.00, assuming an optimistic scenario of no inflation or provider price adjustments. In this extended horizon, the local environment consolidates itself as the most financially efficient option, proving that the operational longevity of the physical asset significantly rewards the immobilized capital.
However, the most critical factor for the business lies in simulating a stress scenario, such as an abrupt 20% increase in data traffic. In a local infrastructure, the cost is fixed; subjecting the server to a higher load will primarily result in a marginal degradation of internal response time, without any change in the monthly bill charged to the company. In the cloud, the behavior is diametrically opposite and financially aggressive. Järvinen (2025) concludes that, in unoptimized remote architectures, cloud costs scale punitively. When facing an increased load, the platform configured for automatic scalability will attempt to compensate for software inefficiency by provisioning maximum computing levels (“scale-up”).
An order of magnitude for the impacts of this reactive scalability is reported by Cho’ponov (2026), who estimates an increase of up to 132% in monthly TCO when systems maintain strong couplings and high request frequency over the network. Adopted for illustrative purposes, the cloud cost of R$ 2,216.00 would be close to R$ 5,000.00 per month if the application forced the database to scale to the processing ceiling to mitigate the impact of latency. The value constitutes a sensitivity scenario, not an empirically validated projection. Finally, it is imperative to add data egress costs to the projection. Pillai (2024) warns that many companies are attracted by the low cost of entry into the cloud, but become hostages to high tariffs when transacting large volumes of data back to local networks. In the evaluated local model, internal traffic is free and unlimited. This sovereignty over information transit, combined with protection against reactive scalability, gives local infrastructure an incomparable strategic advantage, protecting corporate cash flow and showing that cloud success requires deep and mandatory optimization of system code.
Comparison with Market Alternatives and Break-Even Point
The identified parity refers to two specific models: proprietary acquisition and bank subscription as a service at the entry layer. The market, however, offers intermediate modalities whose omission would impoverish the analysis. Among the market alternatives for relational database hosting, the following stand out: on-demand managed service, managed service with reservation, self-managed virtual machine, dedicated server or colocation, and proprietary on-premises infrastructure. The most relevant differentiation lies in licensing and traffic. In managed offerings, the engine license is embedded in the tariff, which favors organizations without prior licenses, while companies with amortized licenses tend to pay twice for the same right of use. Reserved capacity, on the other hand, brings the cloud closer to the acquisition model, but suppresses the elasticity that constitutes its main attraction.
Applying the values from this research, the break-even point between the two models is obtained. With an initial investment of R$ 33,599.00 and a recurring cost of R$ 1,300.00 per month in the local environment, versus R$ 2,216.00 for the subscription, the monthly operational savings are R$ 916.00, which places the break-even point at approximately 36.7 months. This explains the parity observed in the thirty-six-month cycle: the adopted horizon coincides almost exactly with the moment when the investment pays for itself, which is why the equivalent monthly costs converge. The break-even point offers an objective criterion: horizons shorter than thirty-six months, or with uncertainty about continuity, favor the subscription, which avoids capital immobilization; longer horizons favor acquisition.
Two external evidences corroborate the reading. Flexera (2025) found that 27% of the declared spending on infrastructure and platform as a service corresponds to unused or poorly sized resources, a proportion compatible with the idleness measured here; the waste does not stem from provider overpricing, but from capacity that the architecture cannot consume. Convergently, the repatriation carried out by the company 37signals, with projected savings exceeding ten million dollars in five years (Hansson, 2024), illustrates the same mechanism on an industrial scale. For the corporate market and technology management, these findings generate immediate practical implications. It is evident that the migration strategy based on the simple direct transfer of legacy systems is technically flawed and prone to risks. Monolithic systems, designed assuming the almost null latency of a local network, cannot be transposed to the cloud without a profound refactoring of their communication logic.
The success of cloud computing requires fostering a culture of excellence in engineering and financial operations, necessitating the modernization of systems to event-driven architectures and batch processing. In scenarios where this refactoring is unfeasible or where informational sovereignty and budgetary predictability are priorities, on-premises infrastructure remains a sovereign and recommended technical choice. In summary, the results demonstrate that performance degradation in the cloud stems not from computational capacity, but from the interaction between network latency and the application’s unitary communication pattern, directly impacting Total Cost of Ownership and long-term viability, consolidating on-premises infrastructure as the most viable strategy for legacy systems without optimization.
4. Conclusion
This study analyzed the performance and Total Cost of Ownership of database operations in cloud and on-premises infrastructures. It was found that performance degradation in the cloud was not due to computational capacity, but to the interaction between route latency and the application’s unitary communication pattern, which forced the database engine to operate with almost 98% CPU idleness. In contrast, the on-premises environment achieved superior performance, supported by near-instantaneous network communication, revealing the true computational cost of operations. Financially, an empirical tie was observed in the Total Cost of Ownership over a thirty-six-month horizon, but the on-premises environment consolidated as the more advantageous option over a sixty-month cycle and in stress scenarios, due to budgetary predictability and data sovereignty. The main contribution of this work lies in the development of a decision matrix that assists managers in directing their technological strategies, highlighting that cloud adoption requires deep optimization of the system architecture to minimize dependence on constant communication with the remote server.
Despite the robustness of the tests, this work acknowledges its methodological limitations, as the scope focused on the market-standard relational database paradigm, under entry-level hardware configurations, and used static-sized data packages. It is suggested that future research replicate this model using non-relational technologies, investigating whether eventual consistency mitigates latency effects. It is also recommended to quantify the labor cost required for the modernization of legacy systems, integrating this value into the overall calculation of migration projects. Without this modernization, on-premises infrastructure consolidated as the most viable strategy for legacy systems, ensuring high performance, budgetary predictability, and data sovereignty.
Bibliographic References
DAPPERLIB. Dapper: a simple object mapper for .Net. [S. I.]: Stack Exchange, 2024. Disponível em: https://github.com/DapperLib/Dapper. Acesso em: 2
Holmnäs, F. 2024. Migration of Customer Data from an On-premise Database to Cloud Infrastructure. Dissertação (Mestrado em Engenharia de Computação) – Faculty of Science and Engineering, Åbo Akademi University, Finlândia. Disponível em: https://www.doria.fi/bitstream/handle/10024/190527/holmnas_fredrik.pdf?sequence=2&isAllowed=y. Acesso em: 25 jan. 2026.
Husain, M. E.; Hussain, I.; Tanweer, S.; Khan, I. R. 2024. Transitioning from Data Centers to Cloud: An In-depth Analysis of Microsoft SQL Server’s Role in DBaaS and On-Premise Solutions. ICIMMI 2023: Proceedings of the 5th International Conference on Information Management & Machine Intelligence. Disponível em: https://dl.acm.org/doi/10.1145/3647444.3652491. Acesso em: 31 mar. 2026.
Kumar, S. 2024. Cloud vs. on-premises: choosing the right data architecture for scalable, secure solutions. International Journal for Multidisciplinary Research 6(6): 1-9. Disponível em: https://www.researchgate.net/publication/394789475_Cloud_vs_On_Premises_Choosing_the_Right_Data_Architecture_for_Scalable_Secure_Solutions. Acesso em: 25 jan. 2026.
Noor, I.; Tariq, S.B.; Shabbir, A.; Aksa, M. 2024. Into the future with cloud: a comparison with on-premises data warehouse. IEOM Society International, Dubai. Disponível em: https://ieomsociety.org/proceedings/2024dubai/541.pdf. Acesso em: 25 jan. 2026.
Pillai, P. 2024. Cloud vs. On-Premise Data Warehousing: A Strategic Analysis for Financial Institutions. Journal of Computer Science and Technology Studies 6(1): 1-12. Disponível em: https://www.al-kindipublisher.com/index.php/jcsts/article/view/9396. Acesso em: 31 mar. 2026.
Xu, D.; Li, T.; Sun, Z.; Chen, Z.; Zhou, W.; Zhang, Y.; Lu, W.; Du, X. 2025. Performant Synchronization in Geo-Distributed Databases. arXiv. Disponível em: <https://arxiv.org/abs/2511.22444>. Acesso em: 17 ago. 2026.
Zhang, G.; Tang, X.; Chang, Q.; Zhang, H.; Hwang, K.; Li, Y.; Huang, R.; Wang, T.; Zhang, W.; Zhang, M.; Chen, Q.; Hou, X.; Wang, Q. 2025. A Low Latency Cache for Cloud RDBMs. VLDB 2025 Workshop: DATAI. Disponível em: <https://www.vldb.org/2025/Workshops/VLDB-Workshops-2025/DATAI/DATAI25_3.pdf>. Acesso em: 17 ago. 2026.
Article originating from the Final Course Work of the Specialization in Software Engineering of the MBA USP/Esalq
To learn more about the course, click here and access the MBX Academy platform