How data science is (re)programming health

Column

Innovation

January 27, 2026

How data science is (re)programming health

Technology transforms diagnosis, prevention, and management in public and private health systems

The integration between data science and health is not just a technological trend, but one of the most promising and complex movements of contemporary digital transformation. As healthcare systems face challenges such as population aging, the increase in chronic diseases, and budgetary overload, the need for data-driven solutions to improve efficiency, diagnostic accuracy, and personalized care grows.

The abundance of clinical, genomic, environmental, and behavioral data demands approaches capable of handling large volumes of information, variety, and speed — which transforms machine learning techniques, neural networks, text mining, and predictive analysis into strategic tools. More than automating routines, it is about redesigning how medical knowledge is produced, clinical decisions are made, and public policies are formulated.

Major scientific journals such as “The Lancet Digital Health”, “Nature Medicine”, and “JAMA” have frequently pointed out that the responsible adoption of data science can generate concrete impacts in reducing inequalities, anticipating outbreaks, and optimizing treatments — provided it is accompanied by ethical criteria, algorithmic transparency, and rigorous validation.

Diagnosis and prediction

Data science tools have significantly expanded the ability to detect early patterns in medical exams, enabling faster and more assertive diagnoses. Deep learning models, for example, already surpass human specialists in tasks such as interpreting retinal images for the detection of diabetic retinopathy or analyzing tomographies for the identification of pulmonary nodules. The power of these techniques lies in the ability to correlate thousands of variables in seconds, revealing subtle signs that would go unnoticed in traditional methods. This creates a new paradigm, in which medical decision-making is assisted by intelligent systems, capable of combining statistical precision with clinical sensitivity.

One of the most emblematic examples of this potential comes from the research published in “Nature” (McKinney et al., 2020), in which an artificial intelligence system developed by Google Health was trained with millions of mammograms to identify breast cancer. The model achieved superior performance to that of radiologists in several parameters, reducing both false negatives and false positives. The proposal is not to replace the human professional, but to provide an algorithmic second opinion, with a direct impact on patient survival and resource allocation.

Public health and epidemiology

In the collective sphere, data science has revolutionized epidemiological surveillance, allowing for the identification of outbreaks, prediction of hospital demands, and guidance of preventive strategies. The use of predictive models fed by real-time data—such as hospital records, online searches, and meteorological data—has made it possible to anticipate epidemic behaviors weeks in advance. Furthermore, the integration of sociodemographic, mobility, and social network data allows for an understanding of how structural factors affect disease propagation, strengthening more targeted public policies.

During the Covid-19 pandemic, Cheng (2020) and collaborators published an article presenting the development and validation of a machine learning model based on random forest to predict the transfer of hospitalized Covid-19 patients to intensive care units within 24 hours, using routine electronic health record data, such as vital signs, laboratory tests, nursing assessments, and electrocardiograms.

From a retrospective cohort of 1,987 patients admitted to non-intensive care units of a large New York hospital system, between February and April 2020, the authors structured daily time series and applied class balancing and cross-validation techniques to deal with the low event rate and clinical data heterogeneity. The model achieved consistent performance as a screening tool, with approximately 73% sensitivity, 76% specificity, and an AUC around 0.80 in the test set, with respiratory rate, white blood cell count, oxygen saturation, inflammatory markers, hemodynamic parameters, and indicators of renal function standing out as the most relevant predictors (Cheng, 2020).

The results suggest that predictive approaches based on machine learning can support clinical prioritization and operational planning hospital in health crisis contexts, by allowing early identification of patients at higher risk of deterioration, although the authors highlight limitations related to low positive predictive value, the use of data from a single center, and the need for model refinement to increase accuracy and generalization (Cheng, 2020).

Personalized medicine

Another promising frontier is so-called personalized medicine, which uses Data Science to integrate genetic, clinical, and environmental information to offer tailored treatments. Instead of standardized therapies, algorithms identify which patient groups respond best to certain medications, which genetic mutations are linked to disease progression, and which drug interactions should be avoided. This makes medical practice more precise, reducing waste and increasing the chances of therapeutic success.

A study published in “Cell” (Khera et al., 2018) exemplifies this potential by proposing a polygenic score — calculated from millions of genetic variants — capable of estimating the risk of having a myocardial infarction or developing diseases such as type 2 diabetes and prostate cancer.

The article demonstrates that large-scale genomic polygenic scores can identify individuals with elevated risk for common diseases at levels comparable to those conferred by rare monogenic mutations, overcoming a historical limitation of clinical genetics. Using data from large genome-wide association studies and validation in over 400,000 participants from the UK Biobank, the authors developed and tested polygenic risk scores for five high-impact public health diseases: coronary artery disease, atrial fibrillation, type 2 diabetes, inflammatory bowel disease, and breast cancer (Khera et al., 2018).

The results show that a substantial fraction of the population — for example, about 8% in the case of coronary artery disease — presents a three-fold or greater risk, a proportion much higher than that of carriers of classic monogenic mutations associated with the same level of risk. The authors highlight that the risk increases sharply in the upper tail of the score distribution and that these predictors have consistent discriminatory performance, with AUCs ranging approximately between 0.63 and 0.81, depending on the disease. The study argues that polygenic stratification can enable more precise strategies for prevention, screening, and allocation of health resources, while also highlighting challenges such as risk communication, integration with clinical and environmental factors, and the current limitation of applicability in non-European populations (Khera et al., 2018).

This convergence between data science and health represents not just a technological innovation, but a paradigmatic shift in how healthcare is understood, managed, and promoted. By transforming raw data into knowledge, intelligent systems allow for the anticipation of adverse events, personalization of therapies, and optimization of resources at levels previously unattainable by traditional approaches.

However, the enthusiasm for these innovations must be accompanied by a critical and ethical stance, especially in the face of risks such as algorithmic bias, digital exclusion, opacity in decision-making criteria, and inadequate use of sensitive information. Rigorous scientific validation, social control, data governance, and the development of regulatory frameworks are indispensable conditions for the benefits of data science to translate into real gains for the community.

Health, as a social right and a field of high complexity, demands that each technological advance be accompanied by responsibility, transparency, and commitment to the common good. Thus, more than sophisticated algorithms, it is human intelligence — collective, empathetic, and ethical — that must guide the use of these tools towards a fairer, more efficient, and people-centered health system.

To access the references of this text click here

Who wrote this column

Renato Máximo Sátiro

Doutor em Administração pela UFG, professor e orientador no curso de Data Science, Inteligência Artificial e Analytics, do MBA USP/Esalq. Administrador de Empresas na Saneago e pesquisador em grupos de pesquisa da UFG e da UnB, com foco em IA, políticas públicas e acesso à Justiça. Desenvolve projetos em machine learning, deep learning, modelos estatísticos, algoritmos e ética na IA, domínio de ferramentas R, Python, Gretl, SPSS e Stata.

You may also like

October 02, 2026

Determinants of supermarket location in São Paulo

A study investigated the determining factors for supermarket location in the state of São Paulo, with the objective of investigating the factors that explain the presence and expansion of these establishments, considering socioeconomic, demographic, and market dimensions. Data from the 2010 and 2022 Demographic Censuses of IBGE and information from the National Registry of Legal Entities of the Federal Revenue of Brazil were used to build a georeferenced database. A Random Forest classification model was applied, adjusted by grid search with cross-validation, prioritizing the recall-macro metric due to the imbalance of the dependent variable, which represented the presence or absence of supermarkets within a 50-meter buffer. The results indicated that supermarket location is strongly associated with demographic, income, and population characteristics in the surrounding area. The analysis of variable importance showed that sociodemographic factors, such as elderly literacy, household income, and the presence of other food establishments, exerted significant influence, especially in the immediate vicinity. The findings reinforced the hypothesis that the spatial distribution of supermarkets is not random, being conditioned by socioeconomic characteristics and the commercial structure of the territory, offering subsidies for business decisions and urban planning.

Keywords: Spatial Analysis; Machine learning; Expansion; Commercial location; Supermarkets.

Neuroscience And Learning In Education

October 02, 2026

Anti-Racist Education: Inclusive Educational Practices and Social Development

Antiracist education, understood as a structuring axis of inclusive education and social development, was investigated in the Brazilian context. The study aimed to identify and analyze, based on legal documents and teachers’ perceptions, educational practices capable of promoting antiracism in school and society, and how the implementation of Laws nº 10.639/03 and nº 11.645/08 contributed to social justice. A qualitative and documentary approach was adopted, with analysis of educational legislation, curricular guidelines, institutional reports, and academic literature. Complementarily, a semi-structured questionnaire was applied to 295 Basic Education teachers. The data were evaluated quantitatively and qualitatively, through thematic content analysis, and validated with bibliographic studies. The results revealed a paradox: despite a robust legal framework, the implementation of antiracist policies proved fragile and sporadic, with a lack of teacher training, adequate teaching materials, and monitoring. Significant educational inequalities between white and black students were found to persist, and most teachers acknowledged the occurrence of racism in schools, but without clear institutional protocols. Neuroscientific analysis showed that racism negatively impacts students’ cognitive and emotional development. It was concluded that antiracist education is central to quality education, requiring political commitment, public investment, and intersectoral articulation. The integration of Neuroscience in teacher training and the production of qualified materials are crucial to strengthen the school’s role in building a more just and inclusive society.

Keywords: Social Development; Antiracist Education; Social Justice; Law 10.639/03; Inclusive Educational Practices.

Neuroscience And Learning In Education

October 02, 2026

Paths of Inclusion: Perceptions of Parents and Teachers on the Schooling of Students with Dual Exceptionality in the Brazilian Context

Dual Exceptionality, characterized by the coexistence of High Abilities/Giftedness and neurodevelopmental disorders, represents a complex phenomenon that challenges traditional identification and schooling models. The study aimed to understand the perceptions of parents or guardians, teachers, and other education professionals regarding the schooling of students with Dual Exceptionality in the Brazilian context, investigating challenges, pedagogical strategies, and possibilities for inclusion based on equity. The research adopted a qualitative, exploratory, and descriptive approach, and collected data through an online, voluntary, and anonymous questionnaire answered by 25 participants. Discursive data were analyzed using thematic content analysis. The results indicated that knowledge about the topic is often built from personal and professional experiences, revealing gaps in systematic training. Difficulties were identified in identifying these students, in teacher training, and in implementing individualized educational plans, pedagogical flexibility, and curriculum enrichment. Socio-emotional repercussions, such as frustration and low self-esteem, were reported. However, some schools demonstrated inclusive practices based on equity, articulating specific needs and potentialities. Although the results do not allow for generalizations, they highlighted the need to strengthen professional training and the articulation between school, family, and specialized services. It was concluded that the inclusion of students with Dual Exceptionality requires practices that simultaneously recognize their difficulties and potentialities, ensuring equitable conditions for participation, learning, and development.

Keywords: Human development; Teacher training; School inclusion; Neurodivergence; Pedagogical practices.

October 02, 2026

Data Transformation into Strategy: Applied Research for Ecotourism Operation Optimization

The growing demand in ecotourism in Minas Gerais has driven the search for business intelligence to transform customer data into strategic information. The study aimed to structure a data science pipeline to collect, segment, and classify the customer base of an ecotourism operation, in order to optimize marketing actions and anticipate market movements. An exploratory, quali-quantitative research was conducted through a case study. 2,777 transactional records from an ecotourism company, referring to January 2024 to December 2025, were used. The methodological process involved automated data collection (Google Sheets API), processing and enrichment (ETL), validation, and creation of RFM (Recency, Frequency, and Monetary Value) attributes. Dimensionality reduction via PCA and K-Means clustering was applied, with the number of clusters defined by the Elbow method and Silhouette Score. The results were validated with DBSCAN and K-Medoids. The results revealed the identification of three behavioral customer segments: “Loyal”, “Low Value”, and “Potential”. The “Loyal” segment represented the highest accumulated economic value, while the “Potential” segment stood out for its high average ticket and potential for conversion into recurrence. The integration of data analysis techniques proved to be a robust and replicable method for generating intelligence in ecotourism. It was concluded that the structured data science pipeline enabled the behavioral segmentation of the customer base, the statistical validation of the groups, and the creation of a predictive system for new buyers, providing subsidies for data-driven strategic decisions and future analyses.

Keywords: Clustering; Business intelligence; Machine Learning; Customer segmentation; Decision making.

October 02, 2026

Classification of defaulting customers using supervised machine learning techniques

The risk of default in credit operations demanded analytical approaches to anticipate losses. This study comparatively evaluated the performance of supervised machine learning models in classifying defaulting customers in credit card operations. The public dataset “Default of Credit Card Clients” from the University of California Irvine was used, with 30,000 observations and class imbalance. The algorithms Logistic Regression, Random Forest, and Extreme Gradient Boosting were employed. The imbalance was addressed by assigning weights to the classes, and model optimization occurred with the RandomizedSearchCV method, prioritizing sensitivity. Cross-validation results indicated that the Extreme Gradient Boosting model showed a higher capacity for identifying the defaulting class and better discriminatory performance, followed by Random Forest and Logistic Regression, with a sensitivity of 0.8250 and an AUC-ROC of 0.7844 for XGBoost. Interpretability analysis, conducted by the Shapley Additive Explanations (SHAP) technique, highlighted the predominance of variables associated with payment behavior, especially the history of delays. It was concluded that tree-based models, particularly boosting techniques, proved to be more suitable for capturing complex patterns in the data, configuring themselves as consistent alternatives for credit risk management.

Keywords: Machine Learning; Credit Card; Classification; Extreme Gradient Boosting; Credit Risk.

October 02, 2026

Sentiment Analysis on Brazilian Banks on Twitter/X: Comparison between Traditional and Digital Institutions

A study analyzed public perception of Brazilian financial institutions on the Twitter/X platform, highlighting the importance of sentiment monitoring on social networks for understanding reputation and customer experience in the banking sector. The objective was to compare user perception of the image and reputation of traditional and digital banks, based on the sentiment patterns identified in the analyzed manifestations, seeking to identify structural differences between these groups. The methodology was based on the analysis of 1,096 tweets collected between November 2022 and June 2023. Two complementary sentiment analysis approaches were used, the sum and the average of labels, to capture the majority sentiment and nuances of perception. Additionally, the Market Profile Model, with indicators of emotional reputation, reputational risk, neutrality, and polarization, and the Banking Clustering Model, which allowed grouping institutions according to perception patterns, were developed. The results indicated a predominance of neutral and negative sentiments, a higher volume of interactions in digital banks, and structural differences in the emotional intensity of perceptions, with greater stability in digital banks and greater polarization in traditional ones. It was concluded that the combination of analytical and statistical techniques contributed to an in-depth understanding of institutional image in the digital environment, demonstrating the importance of data-driven reputation management strategies.

Keywords: Digital banks; Traditional banks; Data modeling; Opinion mining; Social Networks.

October 02, 2026

Optimization of annual budget planning through project management methodologies

The Annual Budget Planning (POA) is a crucial process for translating organizational strategy into operational and financial goals, but it frequently faces deadline pressures, interdepartmental dependencies, and the repetition of habitual expenses. The study aimed to analyze how the combined application of project management practices and Zero-Based Budgeting (OBZ) can optimize the POA. To this end, a case study was developed in the Brazilian operation of a publicly traded company in the beverage sector, using documentary research of its 2023 results report and an anonymous questionnaire applied to 47 respondents. Documentary analysis indicated growth in net revenue, expansion of gross profit and adjusted EBITDA, and contained advancement of selling, general, and administrative expenses, suggesting cost discipline and operational leverage. The complementary survey revealed a high perception of cascading effect on the schedule, strong support for defining cost package owners, and a preference for technical justification of expenses, in addition to demand for controlled flexibility after the baseline definition. It was concluded that structuring the POA as a project, associated with the rigor of OBZ, increased the process predictability, reinforced accountability for expenses, and broadened the coherence between budgetary execution and economic-financial performance.

Keywords: Cost Control; Operational Efficiency; Zero-Based Budgeting; PMBOK; Beverage Sector.