Innovation
January 27, 2026
How data science is (re)programming health
Technology transforms diagnosis, prevention, and management in public and private health systems

The integration between data science and health is not just a technological trend, but one of the most promising and complex movements of contemporary digital transformation. As healthcare systems face challenges such as population aging, the increase in chronic diseases, and budgetary overload, the need for data-driven solutions to improve efficiency, diagnostic accuracy, and personalized care grows.
The abundance of clinical, genomic, environmental, and behavioral data demands approaches capable of handling large volumes of information, variety, and speed — which transforms machine learning techniques, neural networks, text mining, and predictive analysis into strategic tools. More than automating routines, it is about redesigning how medical knowledge is produced, clinical decisions are made, and public policies are formulated.
Major scientific journals such as “The Lancet Digital Health”, “Nature Medicine”, and “JAMA” have frequently pointed out that the responsible adoption of data science can generate concrete impacts in reducing inequalities, anticipating outbreaks, and optimizing treatments — provided it is accompanied by ethical criteria, algorithmic transparency, and rigorous validation.
Diagnosis and prediction
Data science tools have significantly expanded the ability to detect early patterns in medical exams, enabling faster and more assertive diagnoses. Deep learning models, for example, already surpass human specialists in tasks such as interpreting retinal images for the detection of diabetic retinopathy or analyzing tomographies for the identification of pulmonary nodules. The power of these techniques lies in the ability to correlate thousands of variables in seconds, revealing subtle signs that would go unnoticed in traditional methods. This creates a new paradigm, in which medical decision-making is assisted by intelligent systems, capable of combining statistical precision with clinical sensitivity.
One of the most emblematic examples of this potential comes from the research published in “Nature” (McKinney et al., 2020), in which an artificial intelligence system developed by Google Health was trained with millions of mammograms to identify breast cancer. The model achieved superior performance to that of radiologists in several parameters, reducing both false negatives and false positives. The proposal is not to replace the human professional, but to provide an algorithmic second opinion, with a direct impact on patient survival and resource allocation.
Public health and epidemiology
In the collective sphere, data science has revolutionized epidemiological surveillance, allowing for the identification of outbreaks, prediction of hospital demands, and guidance of preventive strategies. The use of predictive models fed by real-time data—such as hospital records, online searches, and meteorological data—has made it possible to anticipate epidemic behaviors weeks in advance. Furthermore, the integration of sociodemographic, mobility, and social network data allows for an understanding of how structural factors affect disease propagation, strengthening more targeted public policies.
During the Covid-19 pandemic, Cheng (2020) and collaborators published an article presenting the development and validation of a machine learning model based on random forest to predict the transfer of hospitalized Covid-19 patients to intensive care units within 24 hours, using routine electronic health record data, such as vital signs, laboratory tests, nursing assessments, and electrocardiograms.
From a retrospective cohort of 1,987 patients admitted to non-intensive care units of a large New York hospital system, between February and April 2020, the authors structured daily time series and applied class balancing and cross-validation techniques to deal with the low event rate and clinical data heterogeneity. The model achieved consistent performance as a screening tool, with approximately 73% sensitivity, 76% specificity, and an AUC around 0.80 in the test set, with respiratory rate, white blood cell count, oxygen saturation, inflammatory markers, hemodynamic parameters, and indicators of renal function standing out as the most relevant predictors (Cheng, 2020).
The results suggest that predictive approaches based on machine learning can support clinical prioritization and operational planning hospital in health crisis contexts, by allowing early identification of patients at higher risk of deterioration, although the authors highlight limitations related to low positive predictive value, the use of data from a single center, and the need for model refinement to increase accuracy and generalization (Cheng, 2020).
Personalized medicine
Another promising frontier is so-called personalized medicine, which uses Data Science to integrate genetic, clinical, and environmental information to offer tailored treatments. Instead of standardized therapies, algorithms identify which patient groups respond best to certain medications, which genetic mutations are linked to disease progression, and which drug interactions should be avoided. This makes medical practice more precise, reducing waste and increasing the chances of therapeutic success.
A study published in “Cell” (Khera et al., 2018) exemplifies this potential by proposing a polygenic score — calculated from millions of genetic variants — capable of estimating the risk of having a myocardial infarction or developing diseases such as type 2 diabetes and prostate cancer.
The article demonstrates that large-scale genomic polygenic scores can identify individuals with elevated risk for common diseases at levels comparable to those conferred by rare monogenic mutations, overcoming a historical limitation of clinical genetics. Using data from large genome-wide association studies and validation in over 400,000 participants from the UK Biobank, the authors developed and tested polygenic risk scores for five high-impact public health diseases: coronary artery disease, atrial fibrillation, type 2 diabetes, inflammatory bowel disease, and breast cancer (Khera et al., 2018).
The results show that a substantial fraction of the population — for example, about 8% in the case of coronary artery disease — presents a three-fold or greater risk, a proportion much higher than that of carriers of classic monogenic mutations associated with the same level of risk. The authors highlight that the risk increases sharply in the upper tail of the score distribution and that these predictors have consistent discriminatory performance, with AUCs ranging approximately between 0.63 and 0.81, depending on the disease. The study argues that polygenic stratification can enable more precise strategies for prevention, screening, and allocation of health resources, while also highlighting challenges such as risk communication, integration with clinical and environmental factors, and the current limitation of applicability in non-European populations (Khera et al., 2018).
This convergence between data science and health represents not just a technological innovation, but a paradigmatic shift in how healthcare is understood, managed, and promoted. By transforming raw data into knowledge, intelligent systems allow for the anticipation of adverse events, personalization of therapies, and optimization of resources at levels previously unattainable by traditional approaches.
However, the enthusiasm for these innovations must be accompanied by a critical and ethical stance, especially in the face of risks such as algorithmic bias, digital exclusion, opacity in decision-making criteria, and inadequate use of sensitive information. Rigorous scientific validation, social control, data governance, and the development of regulatory frameworks are indispensable conditions for the benefits of data science to translate into real gains for the community.
Health, as a social right and a field of high complexity, demands that each technological advance be accompanied by responsibility, transparency, and commitment to the common good. Thus, more than sophisticated algorithms, it is human intelligence — collective, empathetic, and ethical — that must guide the use of these tools towards a fairer, more efficient, and people-centered health system.
| To access the references of this text click here |
Who wrote this column
Renato Máximo Sátiro








