Health
December 10, 2025
Logistic Regression Model for Classification of Death in HIV Patients in Brazil
Author: Fabiano Gomes de Almeida — Advisor: João Vitor Matos Gonçalves
Summary prepared by the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege focused on synthesis and writing.
The objective of this research was to develop and validate a logistic regression model to identify factors associated with mortality in adult HIV patients in Brazil, using data from the Notifiable Diseases Information System (SINAN). The analysis sought to determine the risk predictors and generate an individualized score applicable to new patients, serving as a tool to direct preventive actions, optimize therapeutic resources, and improve clinical management in the Unified Health System (SUS). The intention is to provide quantitative subsidies for patient stratification based on the probability of death, enabling personalized interventions for profiles of greatest vulnerability.
The relevance of the study is anchored in the challenge that the HIV/AIDS epidemic represents for public health. Despite advances in access to diagnosis and antiretroviral therapy (ART), mortality rates are still considerable, especially in specific populations and regions with access barriers (Oliveira et al., 2020). The heterogeneity of the epidemic in Brazil requires understanding the demographic, clinical, and behavioral factors that influence clinical outcomes. In this context, the application of statistical models to large governmental databases, such as those from DATASUS, is a fundamental approach to transforming data into strategic intelligence for health decision-making (Souza, 2018).
Logistic Regression, a supervised machine learning technique, is suitable for modeling the relationship between predictor variables and a binary outcome, such as the occurrence of death. Unlike descriptive analyses, a predictive model allows for anticipating scenarios and identifying non-apparent patterns, strengthening the healthcare system’s ability to act proactively (Santos, 2019). This work aligns with trends in evidence-based medicine and data-driven health management, seeking to translate epidemiological complexity into a practical tool.
The database used, from the HIVA archive of SINAN, consolidates mandatory notification information for HIV cases in adults. The variety of variables, covering sociodemographic, clinical, and exposure characteristics, allowed for the construction of a robust multifactorial model. The analysis of this data makes it possible to confirm known risk factors and identify new associations specific to the Brazilian context, contributing to the knowledge about the dynamics of HIV mortality in the country (Lima & Costa, 2021).
This study seeks to strengthen the bridge between research and clinical and management practice. By generating a validated model with high discrimination power, the research offers an instrument that can be integrated into monitoring processes, assisting health teams in the early identification of individuals who require intensified attention. The ability to generate a risk score quantifies the vulnerability of each patient, allowing for more efficient resource allocation and contributing to reducing mortality and improving the quality of life of people living with HIV in Brazil.
The methodology began with the collection and preparation of data extracted from the HIVA file of SINAN, comprising HIV notifications in adults from 2015 to 2023. The initial dataset contained 76 variables, including notification information, demographics (sex, age, race, education), exposure history, comorbidities, clinical manifestations, and case evolution (death and date).
The target audience was defined as adult patients with HIV attended in the public healthcare system, and the outcome of interest was mortality. The database was filtered by the EVOLUCAO variable, selecting records for “Alive”, “Death by AIDS”, or “Death by other causes”. The response variable (target) was constructed with a predictive horizon of 12 months from the first day of the notification year. This period proved to be the most balanced between the volume of events and the maturation time for evaluation. The binary response variable was assigned a value of 1 if death occurred within 12 months and 0 otherwise. The data were segmented into training and testing sets (crops 2015-2020), out-of-time validation (2021-2022), and production (2023).
Exploratory data analysis was fundamental for variable preparation. Univariate analysis assessed category frequency, the prevalence of missing data, and inconsistencies. For quantitative variables, descriptive statistics were calculated. Bivariate analysis investigated the relationship between each potential predictor and the response variable, using the calculation of relative risk (ratio between the probability of non-death and death). Based on these values, categories of qualitative variables were grouped to optimize predictive power and ensure monotonicity. For quantitative variables, ranges based on quantiles were created, also grouped by relative risk. The process resulted in 25 candidate variables for modeling (4 quantitative and 21 qualitative).
For the model construction, the logistic regression algorithm was employed (Fávero & Belfiore, 2024). The 25 selected variables were converted into “dummies”, with one reference category omitted for each original variable to avoid multicollinearity. The selection of the most relevant variables was automated by the “stepwise” procedure, which seeks a balance between complexity and explanatory power. The absence of multicollinearity was verified by the Variance Inflation Factor (VIF). Model validation was performed using the Area Under the ROC Curve (AUROC) and the Kolmogorov-Smirnov (KS) test to assess discrimination power, and the Population Stability Index (PSI) to monitor performance over time.
The data preparation started from 410,291 records (2015-2023). 39 out of the original 76 variables were discarded due to a high incidence of missing data (>99%) or lack of variability. After filtering the EVOLUCAO field, the dataset was consolidated into 389,661 records. The definition of the response variable, TARGET12, with a 12-month horizon, ensured a sufficient volume of events, allowing the use of the 2021 and 2022 harvests as an “out of time” sample. The final dataset for analysis totaled 389,038 records.
The bivariate exploratory analysis, on 272,628 training and testing records, refined the set of predictors. Of the initial 35 variables, 10 were excluded due to low representativeness, lack of discriminatory power, or counterintuitive behavior. The remaining 25 underwent feature engineering, including the creation of time variables, binary indicators (gestation), and regrouping of categories by relative risk. The five variables with the highest discriminatory power showed a direct correlation between relative risk and mortality rates: categories with “Good” risk had mortality below the general average (1.56%), while those with “Bad” risk exhibited higher rates. The age variable (vn01_IDADE), for example, demonstrated a decreasing order of relative risk (higher chance of survival) with decreasing age, validating its consistency.
The modeling used binary logistic regression on the 25 variables. After creating “dummies”, the “stepwise” procedure on the training set (190,839 records) selected 24 statistically significant “dummies”, representing 18 original variables. All presented a p-value lower than 0.01. Multicollinearity analysis confirmed robustness, with all VIF values below the threshold of 10, ensuring the independence of the predictors.
The model validation demonstrated its high performance. The congruence analysis confirmed that the coefficients’ signs were aligned with theoretical expectation; for example, advanced age and comorbidities showed positive coefficients. The descriptive analysis of the score showed clear separation between the “death” and “no death” groups, with significantly higher mean and median scores for the former group in all samples (training, testing, and “out of time”).
The discrimination power was quantified by accuracy indicators. The Area Under the ROC Curve (AUROC) reached 0.84 in the training sample, 0.8394 in the test, and 0.827 in the “out of time”, values that indicate a very good classification performance (Pereira, 2013). The Kolmogorov-Smirnov (KS) test presented values of 0.5287 (training), 0.5265 (test), and 0.5068 (“out of time”), classifying the model as excellent in its separation capacity (Sicsú, 2010). The consistency of the results between the samples evidences the absence of “overfitting” and the model’s generalization capacity.
The analysis of the score distribution by deciles reinforced the predictive capacity, with an increasing ordering of the mortality rate in deciles of higher scores. For the 2023 production crop, the Population Stability Index (PSI) was 0.0014, a value considered very low (Dataconomy, 2025), confirming that the new patient population maintained a risk profile similar to that of the development population, attesting to the model’s stability.
The results allowed the identification of risk niches with high precision, even with an overall mortality rate of 1.61%. Patients with probable transmission through injectable drug use had a 2.94 times higher chance of death than non-death, with a mortality rate 171.2% above the average. The presence of symptoms such as persistent cough or pneumonia increased the chance of death by 5.88 times, with a mortality rate 437.2% higher than the average. The diagnosis of anemia, lymphopenia, or thrombocytopenia increased the chance of death by 6.25 times, with a mortality rate 460.3% above the average.
The model also identified protective factors. Patients up to 24 years old showed a 3.59 times higher chance of survival, with a mortality rate 71.7% lower than the average. Other factors associated with death included absence of declaration in critical variables, age over 42 years, low education level, prolonged time until confirmatory test, and the federative unit. The lowest mortality rates were observed in pregnant women (0.20%), young people (0.44%), and individuals with complete high school or higher education (0.64%), highlighting the importance of sociodemographic factors and prenatal care.
This study achieved its objectives by developing a robust logistic regression model with high predictive capacity for factors associated with mortality in adult HIV patients in Brazil. The analysis of the HIVA database from SINAN allowed for the quantification of the impact of risk factors and their consolidation into an individualized risk score. The tool demonstrated excellent discrimination performance (AUROC and KS) and temporal stability (out-of-time and PSI analysis). The results provide subsidies for public health management, allowing for patient stratification and the targeting of actions towards groups of greatest vulnerability, optimizing resource allocation.
The practical implications are significant. The model can be implemented as an early warning system in healthcare services, assisting clinical teams in identifying patients who require intensive follow-up. The identification of factors such as injection drug use, anemia, and persistent cough as high-risk predictors reinforces the need for multidisciplinary approaches. The research highlights the potential of health data analysis to improve epidemiological surveillance and evidence-based decision-making. It is concluded that the objective was achieved: it was demonstrated that the developed logistic regression model is a robust tool with high predictive power for classifying the risk of death in adult HIV patients in Brazil, identifying critical factors for the formulation of more effective health policies.
References:
Bueno, L. M. 2011. Credit analysis: model evaluation measures and application of fuzzy theory in decision making. Final Graduation Work. Department of Statistics, University of Brasília, Brasília, DF, Brazil. Available <https://bdm. unb. br/bitstream/10483/3508/1/2011_LoreMartinsBueno. pdf>. Accessed on: August 08, 2025.
Dataconomy. 2025. Population Stability Index (PSI). Dataconomy PT. Available at: <https://pt. dataconomy. com/2025/04/18/indice-de-estabilidade-da-populacao-psi/>.
Fávero, L. P.; Belfiore, P. 2024. Data analysis manual. LTC, Rio de Janeiro, RJ, Brazil.
Gomes, R. & Ferreira, J. C. 2022. Public Health Data Analysis: Methods and Applications. Health Publishing House, São Paulo, SP, Brazil.
Lima, A. C. & Costa, M. S. 2021. HIV/AIDS Epidemiology in Brazil: A Time Series Analysis. Cadernos de Saúde Pública, v. 37, n. 5, e00123420.
Oliveira, C. S. de et al. 2020. Epidemiological profile of AIDS in Brazil using DATASUS information systems. Brazilian Journal of Clinical Analysis, v. 52, n. 1, p. 35–42, 2020. Available at: <https://www. rbac. org. br/artigos/perfil-epidemiologico-da-aids-no-brasil-utilizando-sistemas-de-informacoes-do-datasus>
Pereira, M. B. 2013. Estimation of Sensitivity, Specificity, and ROC Curve. Master’s Dissertation. Department of Mathematics and Applications, School of Sciences, University of Minho, Braga, Portugal. Available at: <https://repositorium. sdum. uminho. pt/handle/1822/29401>
Santos, E. M. 2019. Applied Statistical Modeling. Academic Publishing House, Belo Horizonte, MG, Brazil.
Sicsú, A. L. 2010. Credit Scoring: Development, Implementation, and Monitoring. Blucher, São Paulo, SP, Brazil.
Souza, F. L. 2018. The Role of Epidemiological Surveillance in the Unified Health System (SUS). Brazilian Journal of Epidemiology, v. 21, e180001.
Executive summary from the Final Project of the Specialization in Data Science and Analytics from the MBA USP/Esalq
Learn more about the course; click here: