Data science in the diagnosis and treatment of heart diseases

Column

Data Science Analytics

Innovation

Health

Technology

September 02, 2024

Data science in the diagnosis and treatment of heart diseases

Predictive models, sentiment analysis, and deep learning algorithms are already applied to increase diagnostic accuracy

In recent years, data science has proven to be an indispensable tool in the field of health, enabling significant advances in the diagnosis, prognosis, and treatment of diseases cardiovascular. The integration of advanced machine learning and artificial intelligence (AI) techniques allows for the analysis of large volumes of clinical data, aiding in the identification of patterns and risk factors that were previously challenging to detect.

The predictive models, sentiment analyses, and algorithms for deep learning are already applied to increase diagnostic accuracy, which allows personalizing the treatment of these diseases, causing a transformative impact in the healthcare field.

Recent studies discuss the relevance of these innovative approaches in the reduction of mortality and in the improvement of medical treatment for cardiac patients, emphasizing the importance of data science in modern medicine. Research has focused on creating predictive models to improve the diagnosis and management of heart-related diseases, using a variety of machine learning techniques.

Filipe & Silva (2024), for example, used a technique called Synthetic Minority Over-sampling Technique (SMOTE) to balance the data and predict deaths from heart failure, using five different classification models. Balancing the data is important because many models perform better when the data is balanced. SMOTE is a method that creates new samples to help train the model.

This work demonstrated the effectiveness of a statistical technique called binary logistic regression, which presented the best balance between accuracy and sensitivity. This method is used to predict the probability of an event occurring, such as the presence or absence of a disease, based on a set of variables. This suggests its potential for integration into health systems.

Similarly, Bani Hani & Ahmad (2024) used the Chi-square Automatic Interaction Detector (CHAID) model to predict mortality among Jordanian men with heart attacks, achieving remarkable accuracy and identifying critical variables such as pulse oximetry and systolic blood pressure. These studies highlight the potential of predictive machine learning-based approaches to improve diagnostic accuracy and enable more precise and personalized clinical intervention.

Deep learning

Beyond classical predictive approaches, advanced deep learning techniques are applied in the analysis of medical signals captured by electrocardiograms (ECGs), aiming to improve the early detection and precise diagnosis of diseases. Deep learning is a subfield of AI that uses artificial neural networks to process large amounts of data and identify complex patterns.

Insook et al. (2024), for example, proposed the use of a convolutional neural network (CNN) – a type of deep neural network especially effective in processing visual and signal data – combined with digital filters to improve the classification of cardiovascular diseases from ECG signals, achieving an average accuracy of 98.6%. Digital filters, such as Butterworth, are used to improve signal quality before being analyzed by the neural network. The study demonstrates the relevance of signal pre-processing techniques and the role of deep neural networks in reducing data bias, increasing prediction accuracy.

Similarly, Bhangale et al. (2024) explored the use of ECG images applied to machine learning models, such as Random Forest, for the detection of anomalies. The work demonstrated that Random Forest outperformed other models in terms of robustness and ability to avoid overfitting, making it a promising approach for early diagnosis of heart diseases. Overfitting occurs when a machine learning model fits the training data too well but fails to generalize to new data. In other words, the model “learns” the details and noise of the training data but cannot make accurate predictions on data it has not seen before. These examples highlight the capability of deep learning techniques in evaluating complex biomedical signals, providing resources for non-invasive diagnosis.

Hybrid Models

The use of hybrid models and advanced machine learning techniques has proven to be an effective approach to address the complex challenges in the diagnosis and prognosis of heart diseases. He et al. (2024), for example, developed a predictive model that combines perivascular adipose tissue (PVAT) characteristics with machine learning algorithms to predict adverse events. The study demonstrated that the inclusion of imaging data, such as attenuation indices and PVAT fat volume, significantly increased the accuracy of predictive models, providing a more precise and effective risk stratification.

Similarly, Hossain et al. (2024) introduced the PrecisionCardio approach, which uses six different machine learning models, including Support Vector Machine (SVM) and Random Forest, to predict the trajectory of heart failure with high accuracy. These studies demonstrate how the integration of advanced techniques and hybrid models can improve the understanding and prediction of complex conditions, providing a robust foundation for personalized medicine.

The risk stratification and personalized treatment are areas in which machine learning has shown great potential in managing heart diseases. Wei et al. (2024) explored the use of artificial intelligence algorithms, such as Random Forest and Bagged CART, to predict the risk of acute kidney injury in patients with acute myocardial infarction, demonstrating that these models can be integrated into clinical systems to support critical medical decisions.

Similarly, Liu et al. (2024) investigated sublingual microcirculatory dysfunction as an early indicator for assessing cardiovascular risk in patients with type 2 diabetes, using machine learning to identify significant correlations between microcirculation and the stage of cardiovascular-renal-metabolic syndrome. These studies demonstrate how machine learning techniques can not only enhance diagnostic accuracy but also enable a more personalized approach to patient treatment, adjusting interventions according to the individual risk profile.

Future perspectives

As data science advances, its application in healthcare, particularly concerning heart disease, is expected to become increasingly sophisticated and integrated. The combination of advanced machine learning techniques, the analysis of large volumes of biomedical data, and treatment personalization represent the current state of the art, but future trends indicate an even greater expansion of these capabilities.

The professionals working with these data can be either technology specialists with health knowledge or physicians with technology knowledge. However, to achieve the best results, it is essential that both professionals work together. Technology specialists, such as data scientists and machine learning engineers, bring technical skills to data analysis and modeling. On the other hand, physicians and healthcare professionals provide the clinical knowledge necessary to interpret the results and apply the findings in the medical context. The collaboration between these professionals is fundamental to ensure that the developed solutions are both technically robust and clinically relevant.

In the coming years, the integration between explainable artificial intelligence (XAI) – which makes AI models more transparent and interpretable, helping to explain how and why AI reached certain conclusions – and the use of multimodal data – which combines genomic, imaging, clinical, and lifestyle information – promises to further transform medical practice. More transparent and interpretable predictive models will allow for better communication between doctors and patients, promoting personalized medicine (Yang et al., 2024).

Furthermore, access to artificial intelligence technologies in different geographical regions and the interoperability of global health systems can contribute to a significant reduction in mortality from heart disease worldwide. In summary, data science is at the center of a revolution in cardiovascular health, with a promising future that should bring even more impactful innovations in patient care.

To access the references of this text click here

Who wrote this column

José Erasmo Silva

José Erasmo Silva é professor, formado em Matemática e Administração, com mais de 25 anos de experiência em gestão empresarial e de pessoas. É mestre e doutor em Administração, com foco em Finanças, e especialista em Data Science e Analytics e em Finanças e Controladoria. Realizou pós-doutorado na Universidade Federal da Bahia (UFBA). Atualmente, atua como professor orientador no MBA em Data Science, Inteligência Artificial e Analytics da USP/Esalq e leciona na EEP/FUMEP e na rede estadual de ensino de São Paulo.

You may also like

October 02, 2026

Determinants of supermarket location in São Paulo

A study investigated the determining factors for supermarket location in the state of São Paulo, with the objective of investigating the factors that explain the presence and expansion of these establishments, considering socioeconomic, demographic, and market dimensions. Data from the 2010 and 2022 Demographic Censuses of IBGE and information from the National Registry of Legal Entities of the Federal Revenue of Brazil were used to build a georeferenced database. A Random Forest classification model was applied, adjusted by grid search with cross-validation, prioritizing the recall-macro metric due to the imbalance of the dependent variable, which represented the presence or absence of supermarkets within a 50-meter buffer. The results indicated that supermarket location is strongly associated with demographic, income, and population characteristics in the surrounding area. The analysis of variable importance showed that sociodemographic factors, such as elderly literacy, household income, and the presence of other food establishments, exerted significant influence, especially in the immediate vicinity. The findings reinforced the hypothesis that the spatial distribution of supermarkets is not random, being conditioned by socioeconomic characteristics and the commercial structure of the territory, offering subsidies for business decisions and urban planning.

Keywords: Spatial Analysis; Machine learning; Expansion; Commercial location; Supermarkets.

Neuroscience And Learning In Education

October 02, 2026

Anti-Racist Education: Inclusive Educational Practices and Social Development

Antiracist education, understood as a structuring axis of inclusive education and social development, was investigated in the Brazilian context. The study aimed to identify and analyze, based on legal documents and teachers’ perceptions, educational practices capable of promoting antiracism in school and society, and how the implementation of Laws nº 10.639/03 and nº 11.645/08 contributed to social justice. A qualitative and documentary approach was adopted, with analysis of educational legislation, curricular guidelines, institutional reports, and academic literature. Complementarily, a semi-structured questionnaire was applied to 295 Basic Education teachers. The data were evaluated quantitatively and qualitatively, through thematic content analysis, and validated with bibliographic studies. The results revealed a paradox: despite a robust legal framework, the implementation of antiracist policies proved fragile and sporadic, with a lack of teacher training, adequate teaching materials, and monitoring. Significant educational inequalities between white and black students were found to persist, and most teachers acknowledged the occurrence of racism in schools, but without clear institutional protocols. Neuroscientific analysis showed that racism negatively impacts students’ cognitive and emotional development. It was concluded that antiracist education is central to quality education, requiring political commitment, public investment, and intersectoral articulation. The integration of Neuroscience in teacher training and the production of qualified materials are crucial to strengthen the school’s role in building a more just and inclusive society.

Keywords: Social Development; Antiracist Education; Social Justice; Law 10.639/03; Inclusive Educational Practices.

Neuroscience And Learning In Education

October 02, 2026

Paths of Inclusion: Perceptions of Parents and Teachers on the Schooling of Students with Dual Exceptionality in the Brazilian Context

Dual Exceptionality, characterized by the coexistence of High Abilities/Giftedness and neurodevelopmental disorders, represents a complex phenomenon that challenges traditional identification and schooling models. The study aimed to understand the perceptions of parents or guardians, teachers, and other education professionals regarding the schooling of students with Dual Exceptionality in the Brazilian context, investigating challenges, pedagogical strategies, and possibilities for inclusion based on equity. The research adopted a qualitative, exploratory, and descriptive approach, and collected data through an online, voluntary, and anonymous questionnaire answered by 25 participants. Discursive data were analyzed using thematic content analysis. The results indicated that knowledge about the topic is often built from personal and professional experiences, revealing gaps in systematic training. Difficulties were identified in identifying these students, in teacher training, and in implementing individualized educational plans, pedagogical flexibility, and curriculum enrichment. Socio-emotional repercussions, such as frustration and low self-esteem, were reported. However, some schools demonstrated inclusive practices based on equity, articulating specific needs and potentialities. Although the results do not allow for generalizations, they highlighted the need to strengthen professional training and the articulation between school, family, and specialized services. It was concluded that the inclusion of students with Dual Exceptionality requires practices that simultaneously recognize their difficulties and potentialities, ensuring equitable conditions for participation, learning, and development.

Keywords: Human development; Teacher training; School inclusion; Neurodivergence; Pedagogical practices.

October 02, 2026

Data Transformation into Strategy: Applied Research for Ecotourism Operation Optimization

The growing demand in ecotourism in Minas Gerais has driven the search for business intelligence to transform customer data into strategic information. The study aimed to structure a data science pipeline to collect, segment, and classify the customer base of an ecotourism operation, in order to optimize marketing actions and anticipate market movements. An exploratory, quali-quantitative research was conducted through a case study. 2,777 transactional records from an ecotourism company, referring to January 2024 to December 2025, were used. The methodological process involved automated data collection (Google Sheets API), processing and enrichment (ETL), validation, and creation of RFM (Recency, Frequency, and Monetary Value) attributes. Dimensionality reduction via PCA and K-Means clustering was applied, with the number of clusters defined by the Elbow method and Silhouette Score. The results were validated with DBSCAN and K-Medoids. The results revealed the identification of three behavioral customer segments: “Loyal”, “Low Value”, and “Potential”. The “Loyal” segment represented the highest accumulated economic value, while the “Potential” segment stood out for its high average ticket and potential for conversion into recurrence. The integration of data analysis techniques proved to be a robust and replicable method for generating intelligence in ecotourism. It was concluded that the structured data science pipeline enabled the behavioral segmentation of the customer base, the statistical validation of the groups, and the creation of a predictive system for new buyers, providing subsidies for data-driven strategic decisions and future analyses.

Keywords: Clustering; Business intelligence; Machine Learning; Customer segmentation; Decision making.

October 02, 2026

Classification of defaulting customers using supervised machine learning techniques

The risk of default in credit operations demanded analytical approaches to anticipate losses. This study comparatively evaluated the performance of supervised machine learning models in classifying defaulting customers in credit card operations. The public dataset “Default of Credit Card Clients” from the University of California Irvine was used, with 30,000 observations and class imbalance. The algorithms Logistic Regression, Random Forest, and Extreme Gradient Boosting were employed. The imbalance was addressed by assigning weights to the classes, and model optimization occurred with the RandomizedSearchCV method, prioritizing sensitivity. Cross-validation results indicated that the Extreme Gradient Boosting model showed a higher capacity for identifying the defaulting class and better discriminatory performance, followed by Random Forest and Logistic Regression, with a sensitivity of 0.8250 and an AUC-ROC of 0.7844 for XGBoost. Interpretability analysis, conducted by the Shapley Additive Explanations (SHAP) technique, highlighted the predominance of variables associated with payment behavior, especially the history of delays. It was concluded that tree-based models, particularly boosting techniques, proved to be more suitable for capturing complex patterns in the data, configuring themselves as consistent alternatives for credit risk management.

Keywords: Machine Learning; Credit Card; Classification; Extreme Gradient Boosting; Credit Risk.

October 02, 2026

Sentiment Analysis on Brazilian Banks on Twitter/X: Comparison between Traditional and Digital Institutions

A study analyzed public perception of Brazilian financial institutions on the Twitter/X platform, highlighting the importance of sentiment monitoring on social networks for understanding reputation and customer experience in the banking sector. The objective was to compare user perception of the image and reputation of traditional and digital banks, based on the sentiment patterns identified in the analyzed manifestations, seeking to identify structural differences between these groups. The methodology was based on the analysis of 1,096 tweets collected between November 2022 and June 2023. Two complementary sentiment analysis approaches were used, the sum and the average of labels, to capture the majority sentiment and nuances of perception. Additionally, the Market Profile Model, with indicators of emotional reputation, reputational risk, neutrality, and polarization, and the Banking Clustering Model, which allowed grouping institutions according to perception patterns, were developed. The results indicated a predominance of neutral and negative sentiments, a higher volume of interactions in digital banks, and structural differences in the emotional intensity of perceptions, with greater stability in digital banks and greater polarization in traditional ones. It was concluded that the combination of analytical and statistical techniques contributed to an in-depth understanding of institutional image in the digital environment, demonstrating the importance of data-driven reputation management strategies.

Keywords: Digital banks; Traditional banks; Data modeling; Opinion mining; Social Networks.

October 02, 2026

Optimization of annual budget planning through project management methodologies

The Annual Budget Planning (POA) is a crucial process for translating organizational strategy into operational and financial goals, but it frequently faces deadline pressures, interdepartmental dependencies, and the repetition of habitual expenses. The study aimed to analyze how the combined application of project management practices and Zero-Based Budgeting (OBZ) can optimize the POA. To this end, a case study was developed in the Brazilian operation of a publicly traded company in the beverage sector, using documentary research of its 2023 results report and an anonymous questionnaire applied to 47 respondents. Documentary analysis indicated growth in net revenue, expansion of gross profit and adjusted EBITDA, and contained advancement of selling, general, and administrative expenses, suggesting cost discipline and operational leverage. The complementary survey revealed a high perception of cascading effect on the schedule, strong support for defining cost package owners, and a preference for technical justification of expenses, in addition to demand for controlled flexibility after the baseline definition. It was concluded that structuring the POA as a project, associated with the rigor of OBZ, increased the process predictability, reinforced accountability for expenses, and broadened the coherence between budgetary execution and economic-financial performance.

Keywords: Cost Control; Operational Efficiency; Zero-Based Budgeting; PMBOK; Beverage Sector.