Article

Software Engineering

October 09, 2026

Implementation of analytics systems in automation environments in the process industry

Implementation of Analytics Systems in Automation Environments in the Process Industry

Luan Oswaldo Gobo; Lucas Cesar Gomes Alvarinho Squillante

DOI: 10.22167/2675-6528-202603174

Article derived from a Course Conclusion Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.

Summary

The digitalization of process plants depends on structured data collection and storage, without which there is no operational visibility. Industrial analytics is the name given to the chain that processes this data from signal acquisition in the field instrument, through time-ordered storage, to availability for analysis and other systems. Proprietary industrial software currently covers this chain, and licensing and maintenance costs restrict its adoption. The study aimed to architect, implement, and validate a system of this nature, employing current software development techniques and components at no licensing cost. The system, named Sistema de Aquisição e Tratamento de Informações de Processo (SATIP) [Process Information Acquisition and Processing System], was structured in three independent layers, using a pharmaceutical reactor simulator as a data source. The simulator executed an eight-step recipe and generated time series with stochastic variation. The simulator was written in Go, a language adopted for generating self-contained binaries, suitable for execution on edge equipment. Storage used TimescaleDB, and data availability was provided through a REST interface with a web dashboard. The system processed approximately 414,000 records per execution, with multivariable trends, alarm logging, and temporal correlation between instruments. The architecture proved reproducible and without licensing costs, and the identification of degradation required only the readings already stored by the system.

Keywords: Software architecture; Digitization; IT/OT integration; Predictive maintenance; Time series.

1. Introduction

Industry 4.0, a concept publicly presented in 2011 at the Hannover Fair, represents the German strategy for modernizing manufacturing (Hermann et al., 2016). This production paradigm is organized around cyber-physical systems, which are machines and installations capable of exchanging information and self-managing autonomously. The design principles include interoperability, information transparency, operator assistance, and decision decentralization (Hermann et al., 2016). A five-level reference architecture highlights the connection of equipment and the conversion of signals into information as precursors to any layer of cognition or autonomous decision-making (Lee et al., 2015).

Process data is a common and central element in these examples, and its full utilization depends on a complete cycle. This cycle includes signal acquisition in the field instrument, transmission over an industrial network, storage in a repository that preserves temporal ordering, and analysis that transforms the raw series into an interpretable indicator. No single software category covers this entire path, with supervisory systems for real-time operation, historians for long-term storage, and analytics platforms for aggregated reading. General-purpose databases have limitations in supporting the time series storage stage, which has driven the development of specialized systems (Jensen et al., 2017). The interruption of any stage compromises the others, as a signal without ordered storage does not support historical analysis, and a history without a query tool does not support operational decision-making.

The maintenance of the integrity of this cycle is justified by gains that precede sophisticated analytical applications. Automatic data acquisition allows for continuous recording of the plant’s behavior, even during periods without direct supervision. The recorded readings reflect the actual measurement of the instrument, without rounding or manual estimations, which meets explicit requirements of good manufacturing practices. From these conditions, deviations are identified the moment they occur, allowing for correction within the batch and reducing operational losses.

Failure anticipation is one of the clearest gains. In process plants, the unscheduled shutdown of critical equipment can interrupt the ongoing batch and, in pharmaceutical environments, frequently leads to the discarding of processed material. Continuous monitoring of variables such as temperature, pressure, and vibration can identify gradual deviations that precede failures and would go unnoticed in spot inspections, such as the progressive increase in time to reach a thermal setpoint (Kahveci et al., 2022). Maintenance, therefore, becomes oriented by the actual condition of the asset, not by a calendar. In regulated environments, such as pharmaceutical production, process data not only describes the batch but proves it, being essential to demonstrate that the product was manufactured under validated conditions. Annex 1 of the European good manufacturing practices, for example, requires continuous particle monitoring in Grade A areas (European Commission, 2022). In this context, failure in data acquisition or loss of records is not an operational inconvenience, but an absence of evidence, which prevents batch release.

However, digitalization does not face obstacles due to technological lag. Pharmaceutical plants already use high-precision instrumentation and modern control systems, driven by regulation. The difficulty lies in the confidentiality of process parameters, which restricts the use of public cloud platforms, and in the cost of the infrastructure needed to consolidate distributed instruments in a manufacturing unit, which often makes projects unfeasible even before initial installation.

Although technically mature alternatives exist, they remain underutilized in industrial practice (Pietrasik et al., 2024). The reason is less technical and more institutional, with commercial offerings concentrated among a few suppliers and vertical integration between layers, leading to the perception that these proprietary platforms are the only way forward. However, the industrial analytics cycle, from ordered acquisition and storage to availability, can be solved with current software engineering techniques and components at no licensing cost. The objective of this work was to architect, implement, and validate a complete industrial analytics system, from data acquisition to availability for analysis, and thus verify whether this need can be met with current software development techniques, instead of proprietary tools.

2. Material and Methods

The research was characterized as applied and exploratory in nature, conducted through the implementation and validation of a software system. The objective was to verify the feasibility of building an industrial analytics solution with current software development techniques, without resorting to proprietary platforms. To this end, each layer of the system was implemented with components at no licensing cost and widely available programming techniques.

Due to the restriction of access to real data from regulated industrial processes, the research used simulated data as the primary source. The process plant studied was a fictitious representation of a general-purpose pharmaceutical reactor. The developed system, named Process Information Acquisition and Processing System (PIAPS), was structured in three independent layers. Data acquisition focused on reading from the simulated controller, acting as a functional substitute for the instrument-PLC set, generating coherent time series and events.

The simulated plant was detailed by a process and instrumentation diagram (P&ID), according to the ANSI/ISA-5.1 standard, specifying a 1,000-liter reactor, seven valves, two motors, and four analog transmitters. The modeled process was a batch mixing and heating operation, with an eight-step sequential recipe. The physical behavior of the tank was described by volume, product temperature, and gas-space pressure, updated every second, based on the energy balance (Seborg et al., 2016). A stochastic variation was added to the quantities to simulate realistic conditions. The transmitters implemented deadband logic and alarm limits at up to four levels (LL, L, H, and HH).

The choice of data repository was central, opting for TimescaleDB, an extension of PostgreSQL 16 specialized in time series (Jensen et al., 2017; Pietrasik et al., 2024). This solution offers automatic time partitioning, temporal aggregation functions, and columnar compression, while maintaining SQL compatibility and operating in a container on the local network. The database schema included tables for equipment types, instrument metadata, users, and a main hypertable for history. Three ingestion patterns were implemented: LogOnChange for state signals, LogAlways for transmitters with dead bands, and batch writing via PostgreSQL’s native binary protocol.

The simulator and the application programming interface (API) were developed in Go, chosen for its native concurrency model, single binary compilation, static typing, and use of the pgx driver for batch writing. The visualization panel (dashboard) was implemented as a web application with Next.js 14 and React 18. Communication between the interface and the database occurred via a REST API in Go with the Chi router. Infrastructure orchestration was performed with Docker Compose. The analysis interface displayed a multi-series line chart for simultaneous visualization of equipment and metrics, and a process page showed the interactive P&ID diagram. Readings were stored in a hypertable and pivoted in the visualization layer for display.

To demonstrate anomaly detection, the simulation was executed in eight successive batches, with progressive reduction of the global heat transfer coefficient, simulating fouling. As a degradation indicator, the average product heating rate, calculated in the range of 25 to 50 °C, was adopted. This rate, expressed in degrees per minute, was used because it is proportional to the heat exchange coefficient and independent of the absolute cycle duration.

3. Results and Discussion

The execution of the simulation of the Process Information Acquisition and Treatment System (SATIP) generated time series that demonstrated coherence with the expected behavior of a pharmaceutical batch process. During validation, 413,398 records were processed in the history hypertable, covering transmitter readings, equipment states, and alarm events. This data was crucial to verify the system’s ability to faithfully record plant operations and to allow for the unambiguous reconstruction of the production cycle, confirming the viability of the proposed architecture for industrial digitalization.

The analysis of the process behavior, as recorded by SATIP, had the primary purpose of validating the integrity of the data acquisition, storage, and availability chain. The focus was not on the industrial process itself, but rather on the system’s functionality in capturing and organizing information. The durations of the simulated recipe steps were defined by model parameters, not representing a specific industrial process, but ensuring that each phase was clearly distinguishable in the data history, essential for evaluating the software solution.

Process variables behavior

The behavior of the TK-01 tank variables throughout the batch production cycle proved consistent with the expectations of the physical model and the recipe. The analysis of time series for level, product temperature, and headspace pressure allowed for the identification of recipe phases through curve inflections. Filling in steps one and two, for example, manifested as a continuous ramp in tank level, while the temperature drop in step two, the rise in step three, the pressurization in step four, the emptying in step six, and the cleaning cycle in step seven were clearly observable, confirming the consistent operation of the controller and the physical model.

The volume profile of tank TK-01 demonstrated non-monotonic behavior, requiring careful data interpretation. Initially, the volume started at 200 liters and reached 750 liters by the end of stage two, reflecting the filling process. In stage six, the transfer pump reduced the volume to 362.8 liters, and the transition occurred due to timing, not a minimum level. Subsequently, in stage seven, the admission of the cleaning solution raised the volume to 486.9 liters. This distinction is fundamental, as a volume increase at the end of the cycle could be mistakenly interpreted as a resumption of production, when, in fact, it corresponds to the sanitization cycle.

The product temperature, measured by transmitter TIT-01, showed an instructive behavior at the beginning of the cycle. Starting from 25 °C, the temperature decreased to 21.6 °C in stage two, before starting to rise in the heating stage. This initial drop was attributed to the admission of Product B at 15 °C, which, upon mixing with Product A, shifted the resulting temperature downwards, according to the enthalpy balance of the admitted streams. It is relevant to note that the minimum of 21.6 °C remained above the alarm limit L (20 °C), indicating that the event, although expected, did not generate an alarm record, validating the control logic and process modeling.

Thermal response and distinction between temperature transmitters

The analysis of the time series of temperature transmitters TIT-02 (steam jacket) and TIT-01 (product) revealed distinct signatures, despite measuring the same quantity at different points in the reactor. The steam valve XV-06 remained open between 18 and 75 minutes, during which the steam jacket was heated. This distinction is crucial for understanding the thermal dynamics of the system and the SATIP’s ability to record and differentiate these responses, providing detailed data for process analysis.

The TIT-02 transmitter, which measures the steam jacket temperature, showed an abrupt rise to 140 °C when the XV-06 valve was opened, with a time constant of approximately 50 seconds. After the valve was closed, the jacket cooled down slowly to ambient temperature (25 °C), with a time constant of about 200 seconds. This asymmetry in thermal response, with rapid heating and slow cooling, is a characteristic behavior of first-order systems and reflects the difference in heat exchange mechanisms: steam condensation is more efficient in heat transfer than passive loss to the environment. The fluctuation observed during the heating plateau, consistent with the instrument’s 1.0 °C dead band, arises from the stochastic variation incorporated into the model, simulating real measurement noise.

On the other hand, the TIT-01 transmitter, which measures product temperature, showed distinct behavior, with a rate of change inversely proportional to the mass contained in the tank, according to equation (1) of the original TCC. With approximately 750 kg of product, the system’s thermal inertia prevented a response as rapid as that of the jacket. The product gradually heated up and maintained a high temperature for a prolonged period even after the steam supply ceased. This difference is visible in the transfer phase, where the jacket had already returned to ambient temperature while the product was still close to 60 °C. This distinction has a direct practical implication: in a real plant, the deviation between the two readings can indicate the efficiency of heat exchange, suggesting, for example, an increase in wall resistance due to fouling, a phenomenon that will be explored in the anomaly analysis.

Historical log of alarms

The SATIP system demonstrated the ability to effectively record alarm history, as illustrated by the behavior of the PIT-01 pressure transmitter. The alarm state is stored as a time series independent of the instrument’s value series, with a record logged at each simulation cycle while the alarm condition remains active. The configured limits for alarm levels H (1.5 bar) and HH (1.8 bar) are stored in the equipment’s metadata, within the JSONB column of the equipment table, avoiding unnecessary replication at each cycle. These limits were intentionally set below the target pressure for stage four (2.0 bar) to ensure that all normal batches triggered alarm logging, allowing for the validation of the system’s functionality.

The consequence of this organization is the ability to retrospectively reconstruct any alarm condition that occurred during the batch. In the case of PIT-01, alarm H was recorded from 83 minutes onwards, upon crossing 1.5 bar, and alarm HH from 88 minutes onwards, upon crossing 1.8 bar, both persisting until the end of the batch. This functionality is a direct requirement of good manufacturing practices regulations, as it allows for the determination of the exact duration for which each alarm level remained active, essential for traceability and auditing purposes in regulated environments. The system, therefore, offers a robust tool for monitoring and documenting critical events in the process.

Temporal correlation between distinct instruments

SATIP’s ability to correlate variables from different instruments along the same temporal axis was demonstrated by the simultaneous visualization of the TIT-01, TIT-02, LIT-01, and PIT-01 series in a single graph. Although the quantities expressed distinct natures, the choice of a shared vertical axis was deliberate to allow comparison of behavior over time. Contextual reading, which displays the values of all instruments at a given instant, complements the visualization, focusing on temporal evolution rather than the comparison of absolute scales. This functionality is a differentiator that elevates the system from a simple repository of isolated series to an analytics platform.

The mechanism that enables this time series overlap resides in the data architecture. All readings are stored in a single hypertable, in long format, where each row associates a timestamp, an equipment identifier, the metric name, and its corresponding value. The programming interface returns, for each selected instrument, the respective time-ordered series. The visualization layer then performs a pivot operation, iterating through the received series and building a timestamp-indexed structure, in which each instrument contributes a column. The result is a wide-format table, with one row per instant and one column for each equipment-metric pair.

The exact joining of data is guaranteed because the data generator registers all equipment with the same timestamp at each simulation cycle, eliminating the need for interpolation or resampling. For each second of the process, there is a corresponding value for each instrument. The graphics library renders one line per column along the shared time axis, and contextual reading traverses the entire table row, allowing the simultaneous display of all instruments’ values at a given instant. This functionality is analogous to the multivariate trend query offered by proprietary industrial historians, but it is obtained in SATIP without licensing costs.

Correction identified in the view layer

During development, the construction of the join key in the interface revealed a defect in the initial version. The key was derived from the timestamp through a formatting function that retained only the time of day, discarding the date, and the final ordering was performed by lexicographical comparison of this textual representation. Although the behavior was correct for series contained within the same day, in simulations that crossed midnight, the portion after the turn of the day was ordered before the initial portion. This resulted in an artificial discontinuity in the graph and an inverted sequence on the temporal axis, compromising the clarity of the data visualization.

The implemented correction consisted of adopting the full timestamp as the joining and sorting key, applying formatting to display only the time of day on the axis labels. This episode illustrates a common class of defects in time series systems, where the visual representation can inadvertently contaminate the data sorting logic. The experience reinforces the recommendation to strictly separate the internal representation of data from its presentation in the interface, a fundamental principle of software engineering that ensures the integrity and correct interpretation of information.

Degradation detection from history

The practical utility of an industrial analytics system transcends the simple visualization of the current batch, finding its greatest value in the comparison of batches over time to identify trends and anomalies. To demonstrate this capability, the simulation was executed in eight successive batches, with a progressive reduction in the overall heat transfer coefficient. This scenario mimics the gradual fouling of the reactor jacket wall due to residue deposition, a common phenomenon in industrial processes that operate in repeated cycles, and which can lead to the degradation of equipment performance.

The choice of an adequate indicator for degradation detection was justified by the nature of the process. The recipe setpoint for product temperature remained fixed at 60 °C in all batches, and the controller advanced to the next stage only after reaching this temperature. Therefore, degradation did not alter the final temperature reached, but rather the *speed* at which the product reached this setpoint. The average heating rate of the product, calculated in the range of 25 to 50 °C, an interval traversed by all batches before reaching the setpoint, was adopted as the indicator. Expressed in degrees per minute, this quantity is directly proportional to the heat transfer coefficient and independent of the absolute cycle duration, making it a robust estimator of heating efficiency.

The results demonstrated that the average heating rate decreased monotonically over the eight batches. Starting from 0.812 °C min⁻¹ in the clean condition (first batch), the rate reduced to 0.526 °C min⁻¹ in the most degraded condition (eighth batch), which represents a total reduction of 35.2%. Although the drop between consecutive batches varied between 4.7% and 7.5%, magnitudes compatible with operational variability and which, in isolation, would not allow a definitive conclusion, the degradation trend became clearly established over the series of batches, evidencing the system’s capacity to identify subtle performance patterns.

The correspondence between the indicator and the underlying physical quantity was remarkable. The reduction imposed on the heat exchange coefficient between the first and eighth batch was 35.0%, while the reduction observed in the heating rate was 35.2%. The average ratio between the relative rate and the fouling factor, calculated batch by batch, was 0.999, with a standard deviation of 0.003. This confirms that the heating rate behaves as a direct estimator of the overall heat transfer coefficient, allowing the degree of jacket fouling to be quantified without the need for direct measurements or production interruptions for inspection, a significant advance for asset management.

The detection of this degradation did not require additional instrumentation, using only the readings from TIT-01 already stored by the system, without vibration sensors, steam flow meters, or any physical alteration to the plant. It is important to note that no conventional alarm would have been triggered, as the final product temperature remained the same in all batches, and no temperature readings exceeded the configured limits for TIT-01. The degradation manifested in the derivative of the variable, and its identification depended on the comparison between batches, illustrating the fundamental distinction between supervision, which observes the instant, and analytics, which observes the trend over time.

This type of indicator is essential for sustaining condition-based maintenance. The progressive drop in the heating rate serves as an early warning of the imminent inability to reach the setpoint, a condition that would halt production. By identifying this trend, proactive chemical cleaning of the jacket can be scheduled for a pre-defined shutdown window, optimizing operation and minimizing losses. This allows a transition from reactive or calendar-based maintenance to a more efficient, predictive approach, aligned with the principles of Industry 4.0.

Comparison with solutions available in the market

The current market offers two main routes for the registration and analysis of process data. The first is the proprietary industrial historian, coupled with supervisory systems and complemented by a visualization layer. This is a traditional solution, technically robust and well-integrated into the automation ecosystem, but its licensing and maintenance costs are scaled for large installations, making it unfeasible for smaller plants, which cannot justify the investment and underutilize a large part of the resources. The second, more recent route, involves sensors with simplified installation that measure variables such as vibration, temperature, and electrical consumption directly on the asset, transmitting the data via wireless network to cloud platforms, where algorithms generate alerts and recommendations.

SATIP positions itself as a third route, addressing distinct needs. While the solution based on coupled sensors is ideal for monitoring rotating assets, whose conditions manifest in housing vibration and temperature, and whose variable of interest is independent of the processed product, the proprietary historian is suitable for large-scale plants that demand formal commercial support. SATIP, in turn, occupies an intermediate range, offering, without licensing costs, the essential functions of a historian: time-ordered history, multivariate trend consultation, alarm logging, and correlation between instruments. It operates on process variables already measured by existing instrumentation and remains entirely on the local network, making it applicable to plants for which the other two routes are disproportionate: the first due to cost and the second due to the requirement of transmitting revenue parameters through external servers, which is incompatible with confidentiality requirements in regulated environments.

It is fundamental to delimit the scope of this comparison. The present work aimed to establish the technical feasibility of building the complete industrial analytics chain using current software development techniques, and not to demonstrate the superiority of SATIP compared to existing commercial products. Commercial solutions offer additional attributes such as contractual support, documented validation, redundancy, and a wide range of ready-made connectors for equipment from various manufacturers. These are factors that an internally developed system would need to acquire over time and which, in regulated environments, carry significant weight in the purchasing decision, not being the focus of this research.

Architecture Adequacy

From the software engineering perspective, SATIP’s decoupled three-layer architecture has proven to be adequate and robust. The clear separation between the data generator, the database, and the visualization interface allows each component to be replaced or modified independently, without impacting the other layers. For example, the data generator can be replaced by a real data collection service, the database can be migrated to a managed instance, and the interface can be swapped for another visualization tool, as long as it consumes the same access points. This modularity is a direct consequence of opting for open formats and protocols at each boundary, answering the work’s central question about the feasibility of replacing proprietary platforms with general-purpose components, which required project discipline in defining the boundaries between layers.

With a structured and accessible history, SATIP’s infrastructure allows for natural and valuable extensions. It is possible to calculate overall equipment effectiveness (OEE) indicators, which combine availability, performance, and production quality into a single percentage, using the start and end stamps of each recipe phase. Furthermore, integration with enterprise resource planning (ERP) systems for data exchange on production orders and raw material consumption, and with manufacturing execution systems (MES) for batch traceability, constitutes a logical and high value-added extension. These integrations are facilitated by the open architecture and data organization, expanding the system’s potential for industrial management.

Limitations

The limitations of the present study deserve explicit mention to contextualize the findings. Firstly, the data used are simulated. The physical behavior of the reactor was modeled using simplified equations and parameters chosen to produce plausible time series, but not to faithfully represent a specific pharmaceutical process. Validation with real data would require access to a production environment and compliance with the strict regulations for validating applicable computerized systems, which was not the focus of this work.

Secondly, the implemented database schema is a simplified version. Features essential for a production environment at scale, such as automatic compression of historical data, the creation of continuous aggregates to pre-calculate hourly and daily averages, and retention policies differentiated by data category, were deliberately kept out of the project’s scope. These optimizations would be indispensable for large-scale operation and for the efficient management of the volume of data continuously generated.

Finally, the anomaly scenario reproduced in the study used a single degradation mode, with deterministic evolution. In contrast, real time series in industrial environments present measurement noise, significant batch-to-batch variability, and the simultaneous occurrence of multiple failure modes. These complex conditions would require the use of more robust and sophisticated statistical anomaly detection methods than the simple batch-to-batch comparison, which was sufficient for the purposes of this work, but not for generalization in a real production context.

In summary, the Process Information Acquisition and Treatment System (SATIP) demonstrated the feasibility of architecting, implementing, and validating a complete industrial analytics chain using current software development techniques and components without licensing costs. The results showed the system’s capability to record and provide coherent time series, support multivariate trends, accurately record alarms, and crucially, detect progressive equipment degradation from historical data, without the need for additional instrumentation. This confirms that the objective of meeting the digitalization need with open, rather than proprietary, solutions was fully achieved, offering a robust and accessible alternative for the process industry.

4. Conclusion

The study aimed to architect, implement, and validate a complete industrial analytics system, from data acquisition to availability for analysis, and thus, verify if this need can be met with current software development techniques, instead of proprietary tools. The Process Information Acquisition and Treatment System (SATIP) demonstrated the capacity to process approximately 414 thousand records from a pharmaceutical reactor simulator, evidencing its aptitude to faithfully record operations and allow unambiguous reconstruction of production cycles. It was verified that the system supported the visualization of multivariate trends, precise recording of alarms with configurable limits, and temporal correlation between distinct instruments, essential functionalities for operational visibility. The main practical contribution of this work lies in demonstrating the technical feasibility of building a complete industrial analytics chain, using exclusively open-source components and software engineering techniques. This offers a robust and accessible alternative for medium-sized plants, overcoming the cost and confidentiality licensing barriers often associated with proprietary solutions.

The SATIP’s capacity to detect progressive equipment degradation from historical data, such as reactor jacket scaling, by analyzing the average product heating rate, was also identified. This finding is crucial for predictive maintenance, as it allowed for the identification of performance trends without additional instrumentation and without triggering conventional alarms, distinguishing trend analysis from mere point-in-time supervision. However, the study had limitations, such as the use of simulated data with a single deterministic degradation mode, which does not reflect the complexity of real time series with noise and multiple failure modes. The database schema was also implemented in a simplified version, without optimizations for automatic historical data compression, creation of continuous aggregates, or large-scale retention policies. For future studies, it is suggested to validate the system with real process data in a production environment, integrate it with enterprise resource planning and manufacturing execution systems, incorporate robust statistical methods for anomaly detection in complex scenarios, and evolve the schema for scaled production use.

Bibliographic References

European Commission [EC]. 2022. EudraLex – The Rules Governing Medicinal Products in the European Union – Volume 4 – Good Manufacturing Practice – Annex 1: Manufacture of Sterile Medicinal Products. European Commission, Bruxelas, Bélgica.

Hermann, M.; Pentek, T.; Otto, B. 2016. Design principles for Industrie 4.0 scenarios. In: Hawaii International Conference on System Sciences, 2016, Koloa, HI, EUA. Anais… p. 3928-3937.

Jensen, S.K.; Pedersen, T.B.; Thomsen, C. 2017. Time series management systems: A survey. IEEE Transactions on Knowledge and Data Engineering 29(11): 2581-2600.

Kahveci, S.; Alkan, B.; Ahmad, M.H.; Ahmad, B.; Harrison, R. 2022. An end-to-end big data analytics platform for IoT-enabled smart factories: A case study of battery module assembly system for electric vehicles. Journal of Manufacturing Systems 63: 214-223.

Lee, J.; Bagheri, B.; Kao, H.A. 2015. A cyber-physical systems architecture for Industry 4.0-based manufacturing systems. Manufacturing Letters 3: 18-23.

Pietrasik, M.; Wilbik, A.M.; Grefen, P.W.P.J. 2024. The enabling technologies for digitalization in the chemical process industry. Digital Chemical Engineering 12: 100161.

Seborg, D.E.; Edgar, T.F.; Mellichamp, D.A.; Doyle III, F.J. 2016. Process Dynamics and Control. 4ed. John Wiley & Sons, Hoboken, NJ, EUA.

Article originating from the Final Course Work of the Specialization in Software Engineering of the MBA USP/Esalq

To learn more about the course, click here and access the MBX Academy platform

You may also like

Software Engineering

October 09, 2026

O papel da densidade de texto instrutivo na eficiência de uma aplicação web.

O desenvolvimento de aplicações web se conecta à experiência do usuário, e este trabalho investigou como o uso excessivo de textos instrutivos pode retardar a conclusão de tarefas e impactar a eficiência da aplicação. O objetivo foi identificar o impacto da densidade textual do conteúdo instrutivo na eficiência de uma aplicação web, utilizando como principal referência a terceira lei de usabilidade de Krug. A pesquisa, de caráter exploratório e delineamento experimental quantitativo, empregou um teste A/B em uma aplicação web responsiva, onde a única variável controlada foi a densidade textual (alta vs. baixa, definida pela contagem de palavras). Participaram 25 usuários, e os dados foram coletados via Datadog RUM, mensurando tempo de conclusão, erros de submissão e taxa de conversão. Os resultados revelaram que a variante com densidade textual reduzida (variante B) apresentou uma taxa de conversão superior (58,3% contra 33,3% da variante A) e um tempo médio de conclusão significativamente menor (1:38 minutos contra 4:58 minutos da variante A), representando um aumento de 67,12% na eficiência. O teste t de Welch (p=0,042) confirmou que a redução da densidade textual impactou a eficiência. Concluiu-se que a redução da densidade textual afeta a eficiência e a taxa de conversão, reforçando a importância de conteúdo objetivo e conciso. Contudo, a baixa densidade textual, por si só, não garantiu o pleno entendimento, sendo essencial a comunicação clara e objetiva das instruções, validando a relevância do UX Writing.

Palavras-chave: Eficiência; Experiência de usuário; Teste A/B; Texto Instrutivo; Usabilidade.

Software Engineering

October 09, 2026

Use of voice in conjunction with large language models as a tool for digital accessibility

Speech is an essential basis for human interaction, and for people with disabilities, it can represent the primary form of communication with the external environment. Given the growing technological influence, voice command identification has emerged as a promising strategy for human-machine interaction. The work explored how voice, in conjunction with Large Language Models (LLMs), can be used efficiently, naturally, and accurately. For this purpose, various design patterns and the Python language were employed, aiming for greater extensibility. Gemini was used as the LLM provider, sending audio directly and leveraging its function calling capability to interact with the device. A system was developed capable of understanding user intent and converting it into actions, whose differential was the computer vision capability based on screenshots and a mesh system for LLM guidance. Tests revealed a user intent comprehension rate of 91.81% and a success rate in execution of 75.45% with the “Flash-3” model (top p 0.5 and top k 5). The proposed system validated the premise that the integration of LLMs into voice interfaces increases the autonomy of users with motor disabilities, fulfilling the purpose of being a modern and effective Assistive Technology. However, questions were raised about the costs of AI and user data security, indicating the need for improvement.

Keywords: Function calling; Human-machine interaction; Voice recognition; Computer vision.

Software Engineering

October 09, 2026

ReasonGuard: A Reasoning Audit Platform for Language Models Based on Structured Thought Decomposition

ReasonGuard was presented, an Artificial Intelligence reasoning auditing platform, developed as an observability middleware between client applications and large language models (LLMs). The work aimed to offer transparency and traceability for AI-based decision-making processes, addressing the gap in operationalizing structured reasoning techniques for auditing purposes. The system intercepted, analyzed, and documented interactions with LLMs through five modules based on the Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT) paradigms. These modules captured reasoning trails, detected structural logical flaws, evaluated response consistency, and generated audit reports targeted at different stakeholder profiles. The platform was implemented with FastAPI (Python) on the backend, React with TypeScript on the frontend, and PostgreSQL as a relational database, following a modularized architecture with multitenancy isolation per user. The results demonstrated the technical feasibility of the proposed approach, with all five modules operational and integrated. It was concluded that ReasonGuard contributes to AI governance by instrumentalizing structured reasoning paradigms as observability and auditing tools, filling a gap in the literature.

Keywords: AI Auditing; Chain-of-Thought; AI Governance; Large Language Models; Algorithmic Transparency.

Software Engineering

October 09, 2026

EngTT: software for road freight based on operational costs and ANTT parameters

Road transport represents the main logistics modality in Brazil, with the composition of freight costs regulated by the National Land Transport Agency (ANTT). Given the absence of a structured methodology for freight calculation and the dependence on isolated spreadsheets in the sector, the EngTT software was developed. The objective was to create a decision support tool that integrated operational variables, calculated the total cost per route, and compared the results with the ANTT’s minimum floor, identifying non-compliance and margin compression scenarios before pricing. The adopted methodology was applied, quantitatively and experimentally, with incremental construction oriented towards the Minimum Viable Product (MVP) concept. The system was structured in three modes of use – Build Visual Route, Batch by spreadsheet, and Scrape routes – sharing a decoupled calculation engine and centralized parameters. The results obtained demonstrated that the software met the proposed objective, showing margin variations and regulatory compliance between the analyzed routes. It was observed that longer routes, such as Ribeirão Preto × Guarujá, presented an increase in cost due to the need for additional driver per diems, while shorter routes, such as Cajamar × Guarujá, showed greater adherence to the ANTT’s regulatory floor. It was concluded that the developed solution is applicable to the road freight transport sector as an effective tool for pricing and operational management.

Keywords: ANTT; Containerized cargo; Software engineering; Road freight; Python.

Software Engineering

October 08, 2026

Data pipeline for stock monitoring considering the fundamentalist methodology

The Brazilian financial scenario faces challenges such as family indebtedness, low financial literacy, and the decentralization of information for investment analysis. Given this, the research objective was to develop a financial data pipeline, based on good Data Engineering practices, to structure, process, and make relevant information available to support individual investors’ decision-making in stock analysis. A medallion architecture was implemented, using MinIO S3 for bronze, silver, and gold layer storage, with Delta Lake for data governance. Apache Spark was employed for distributed processing, Prometheus for monitoring, and Power BI for analytical visualization. The data were mostly financial statements from the Securities and Exchange Commission (CVM). A theoretical investment portfolio was built based on Benjamin Graham’s (2017) principles, applying selection filters and data validation. The results indicated a positive return of 32.88% for the theoretical portfolio, outperforming the Ibovespa index (4.25%) in the period from 2021 to 2024, although lower than the Selic rate (46.96%). Qualitatively, the pipeline processed voluminous datasets, with significant reductions in redundancies, such as 99.27% in the BPA table and 94.29% in the DRE, after applying filters. Gains in data organization, traceability, and quality were evidenced, enabling structured and more robust financial analyses.

Keywords: Medallion Architecture; Data Engineering; Investments; Data Pipeline.

Software Engineering

October 08, 2026

Performance and Total Cost of Ownership of Cloud Databases and On-Premises Infrastructure

This study analyzed the performance of database operations in cloud and on-premises infrastructures, correlating technical behavior with Total Cost of Ownership projection. The research was characterized as a quantitative and experimental study, in which a test system subjected isolated database instances to progressive execution loads, measuring the impact of network latency and computational consumption. For financial analysis, an investment and operational expenses model was developed, diluted over a thirty-six-month cycle. Technical results revealed that the accumulation of internet latency caused severe time degradation in cloud executions, despite the remote infrastructure operating with high processing idleness, registering almost 98% CPU inactivity. In contrast, the on-premises environment achieved superior performance supported by almost instantaneous network communication. Financially, cost consolidation demonstrated an empirical tie between the physical acquisition model and the service subscription model within a three-year horizon, but the on-premises environment proved more advantageous in a sixty-month cycle. It was concluded that the degradation in the cloud was not due to computational capacity, but to the interaction between route latency and the application’s unitary communication pattern. Cloud adoption requires deep optimization of the system architecture to minimize dependence on constant communication with the remote server. Without this modernization, the on-premises infrastructure consolidated as the most viable strategy, ensuring high performance, budgetary predictability, and data sovereignty.

Keywords: Operational expenses (OPEX); Transactional scalability; Network latency; Legacy systems.

Software Engineering

October 08, 2026

Asynchronous Slack-Jira integration via message queue “middleware”: comparison of “cloud computing” solutions

The latency and interoperability between distributed corporate systems constituted the problem investigated, motivated by the costs and fragilities of manual integrations between collaboration and project management platforms. An asynchronous integration “middleware” between Slack and Jira was developed and validated, with the objective of reducing the perceived user response time and ensuring system stability under load. The methodology consisted of building an event-driven software architecture in Python, using the “producer-consumer” and “adapter” patterns to isolate the user interface from “backend” processing. The solution evolved into a cloud-agnostic architecture, based on “serverless” functions, and was subjected to stress tests in local, real network, and production environments on two clouds. The architecture reduced user waiting time from a synchronous estimate of 2,000 milliseconds to a local average of 8.96 milliseconds. Under a load of 50 simultaneous requests, both cloud providers proved viable: Amazon Web Services registered lower average latency in the reception layer (951.35 ms) and double the throughput, while Microsoft Azure executed background processing with a median of 113 ms. The application of “optimistic UI” ensured fluidity of use, and load leveling by queues eliminated the need for infrastructure over-provisioning. The “middleware” consolidated itself as a scalable, resilient, and protected corporate reference model against technological lock-in.

Keywords: Event-driven architecture; Temporal decoupling; Operational efficiency; Information technology service management; System interoperability.

Software Engineering

October 05, 2026

Using Generative AI for Database Selection: A Requirements-Driven Framework for Generating Architecture Decision Records (ADRs)

The growth of data-intensive applications and the adoption of microservices architecture have amplified the need for polyglot persistence, imposing a high cognitive load on software architects in choosing and justifying database technologies. This work aimed to propose and develop an Artificial Intelligence (AI) agent-based framework to guide technological selection and generate well-founded Architecture Decision Records (ADRs). An experimental and applied methodology was employed to build a technical knowledge base. The Retrieval-Augmented Generation (RAG) technique, along with the LangChain and LangGraph libraries, was used to orchestrate agents and anchor the responses of a Large Language Model (LLM). The framework extracted natural language requirements, enriched them with RAG, and sent them to the LLM, which generated ADRs to assist in evaluating theoretical trade-offs. The results demonstrated that the agent with RAG reduced generic responses, increasing theoretical grounding and traceability. The RAG approach proved its effectiveness against conventional prompts (zero-shot), favoring the generation of ADRs with a lower level of hallucination and a high level of theoretical traceability. It was concluded that the automated tool fulfilled the function of requirement mapping, resulting in empirically grounded technical documents and aiding governance and decision-making in software architecture.

Keywords: Databases; Artificial Intelligence; LangGraph; LLM; RAG.