When the data changes, has reality changed?

Column

Data Analysis

September 30, 2026

When the data changes, has reality changed?

Analysis of nearly 20 million faculty-years between 2020 and 2025 shows how changes in population composition and record production can alter the interpretation of a historical series

Between 2020 and 2025, the number of formal teaching positions identified in RAIS (Annual Relation of Social Information), from the Ministry of Labor, increased from about 2.58 million to 4.03 million. In the same period, there was an expressive transformation in the composition of these positions: statutory ones, which were 63.1% of the total in 2020, fell to 40.2% in 2025, while non-permanent or temporary public positions rose from 16.8% to 43.8%.

The data on sick leave also draws attention. The proportion of employment relationships with registered leave increased from 4.07% in 2020 to 10.54% in 2022, and then decreased, practically returning to the starting point: 4.08% in 2025.

At first glance, it would be tempting to conclude that, after a strong increase until 2022, there was a significant improvement in the following years. When the data are examined by type of employment, however, a result that is difficult to ignore emerges. Among non-permanent or temporary public employment, the registered rate goes from 8.47% in 2022 to only 0.41% in 2025.

What happened? Was there really a reduction of this magnitude in dismissals, or did the way in which this phenomenon came to be recorded change? This question illustrates a central problem in data analysis: not every change observed in a time series necessarily represents an equivalent change in the phenomenon we intend to measure. Before interpreting a result, it is necessary to understand how the data were produced.

Behind the numbers

The results were obtained from RAIS microdata, considering five occupational families related to teaching. The analysis covers all Brazilian states and tracks the characteristics of formal employment relationships, such as occupation, age, sex, type of employment, and records of leave due to illness.

There is an important distinction: the unit analyzed is not necessarily a person, but a work relationship. The same professor can have more than one relationship and, therefore, appear more than once in the database. For this reason, 4.03 million records in 2025 do not mean 4.03 million professors, but approximately 4.03 million teaching positions.

This difference seems merely technical, but it is essential. Administrative databases are originally produced for registration and management purposes. When they are used for data analysis, it is necessary to understand the meaning of each variable and, mainly, what each record represents.

Structure Transformation

Between 2020 and 2025, the number of teaching positions considered in the analysis increased by approximately 56%, from 2.58 million to 4.03 million. More important than the absolute growth, however, is observing how the composition of these positions has changed. In 2020, statutory public positions accounted for 63.1% of the analyzed records, while non-tenured or temporary public positions represented 16.8%. In 2025, the percentages changed to 40.2% and 43.8%, respectively.

If these numbers were observed in isolation, it would be possible to interpret the period as evidence of a rapid transformation in the structure of teaching work. But precisely in this interval, an important change occurred in the production of RAIS itself.

Data source

Starting from the base year of 2023, RAIS compliance is now measured solely by extracting events reported in eSocial for all declarant groups, including public agencies. The RAIS 2023 Technical Note records a significant break in the historical series and recommends that the results from that year not be directly compared with previous years, due to the transition process in data capture methods.

This does not mean that subsequent data are invalid or that all observed changes stem from eSocial. It means that the comparison requires caution. Instead of asking only why the numbers changed, it is necessary to ask how much of this change belongs to the phenomenon studied and how much may be related to the way it began to be recorded.

Attrition Data

In 2020, 4.07% of the analyzed links showed a record of leave. The rate rises to 6.83% in 2021 and reaches 10.54% in 2022, the highest value in the series. After that, it recedes to 8.25% in 2023, 5.50% in 2024, and 4.08% in 2025. There are two temptations in this reading. The first would be to automatically associate the 2022 peak with the effects of the pandemic. The second would be to interpret the subsequent drop as evidence of improvement in teachers’ health conditions. The data used here, in isolation, do not allow us to support either of these causal conclusions. What can be affirmed is more restricted: the leave indicator recorded in the database increased sharply until 2022 and fell in the following three years.

Among statutory public positions, the registered rate of leave increased from 4.63% in 2020 to 10.68% in 2022. It then decreased to 8.19% in 2023, 7.55% in 2024, and 6.59% in 2025. Among non-effective or temporary public positions, the trend is much more pronounced: 2.44% in 2020, 8.47% in 2022, 6.47% in 2023, 1.53% in 2024, and only 0.41% in 2025.

A reduction from 8.47% to 0.41% could produce an attractive headline. But we do not know, from these numbers, if teachers with non-effective appointments became less ill. What we do know is that the frequency with which absence is recorded in these appointments has dropped drastically. The difference is fundamental.

To investigate whether the general rate drop could be explained simply by the change in the composition of the links, the variation between 2022 and 2025 was decomposed into two parts: one associated with the change in the participation of different types of links and another related to the changes in the rates recorded within these groups.

The aggregate rate fell from 10.54% to 4.08%, a reduction of 6.46 percentage points. Of this total, approximately 0.92 percentage points are associated with the change in the composition of employment relationships, while 5.54 percentage points result from changes in the rates recorded within groups. In proportional terms, about 14% of the reduction is associated with composition, and approximately 86% with within-group changes.

Mathematically, the decomposition shows that the aggregate drop did not occur solely because the share of a certain type of link increased. But a statistical technique can correctly decompose what is recorded in the database and still not solve a measurement problem. If the way data is produced changed during the period, the decomposition will also reflect this change.

In 2020, the aggregate dropout rate was 4.07%. In 2025, it was 4.08%. Looking only at the two extremes, practically nothing seems to have changed. The decomposition tells another story: the change in the composition of employment relationships contributed approximately -1.15 percentage points to the rate, while changes in internal rates contributed in the opposite direction, with approximately +1.16 percentage points. The two movements practically canceled each other out.

Thus, 4.07% and 4.08% are almost identical numbers produced by quite different structures. Aggregate indicators are useful, but they can hide important transformations in the analyzed population.

What this case teaches

The RAIS case shows that a data analysis does not start with the algorithm. It starts with understanding the base. We could have observed the drop between 2022 and 2025, applied statistical techniques, and ended the analysis with an apparently consistent conclusion. The calculations would be correct. The problem would lie in the interpretation.

This problem is not exclusive to RAIS. A new filling rule, system migration, a classification change, the inclusion of new users, or a change in the mandatory nature of a certain field can create a break that, in a graph, looks exactly like a change in behavior.

Therefore, some questions should precede any more sophisticated model: did the observed population remain comparable? Did the definitions of the variables remain the same? Did the coverage change? Was there any alteration in the system that produces the records?

More data also do not necessarily mean more information. With millions of links, small differences can be calculated with enormous precision. But numerical precision and validity of interpretation are different things. If a variable starts to be recorded in a different way, millions of records can reproduce this difference with enormous consistency. Before asking “which model should I use?”, often the most important question is “what exactly is this data measuring?”.

The value of an unexpected result

The result of 0.41% could have been treated as a discovery. Instead, it served as a warning. Results that are very different from what was expected do not need to be discarded, nor should they be immediately turned into conclusions. They may indicate a relevant phenomenon, a processing error, a population change, a methodological alteration, or a characteristic of the source that we do not yet understand.

In this case, the analysis began as an exploration of millions of faculty links and ended up revealing something broader about the work with data itself: sometimes, the main finding is not in the number found, but in the reason why it should be examined with caution.

RAIS allows us to see important transformations in the Brazilian formal labor market and offers a wealth of information difficult to reproduce from other sources. The analysis identified expressive growth in teaching positions, relevant changes in their composition, and large fluctuations in records of sick leave between 2020 and 2025.

These results remain informative. What changes is the way to interpret them. The transition to eSocial shows why historical series should not be treated merely as sequences of numbers. Each point in a series is the result of rules, systems, definitions, and administrative processes that also have a history.

For those who work with Data Science, one of the main conclusions of this exercise is that, before explaining why an indicator has changed, one should verify if the same phenomenon continues to be measured in the same way.

An algorithm can find patterns in millions of records. Knowing whether these patterns represent the phenomenon we want to understand remains a human task.

To access the references of this text click here

Who wrote this column

José Erasmo Silva

José Erasmo Silva é professor, formado em Matemática e Administração, com mais de 25 anos de experiência em gestão empresarial e de pessoas. É mestre e doutor em Administração, com foco em Finanças, e especialista em Data Science e Analytics e em Finanças e Controladoria. Realizou pós-doutorado na Universidade Federal da Bahia (UFBA). Atualmente, atua como professor orientador no MBA em Data Science, Inteligência Artificial e Analytics da USP/Esalq e leciona na EEP/FUMEP e na rede estadual de ensino de São Paulo.

You may also like

October 02, 2026

Determinants of supermarket location in São Paulo

A study investigated the determining factors for supermarket location in the state of São Paulo, with the objective of investigating the factors that explain the presence and expansion of these establishments, considering socioeconomic, demographic, and market dimensions. Data from the 2010 and 2022 Demographic Censuses of IBGE and information from the National Registry of Legal Entities of the Federal Revenue of Brazil were used to build a georeferenced database. A Random Forest classification model was applied, adjusted by grid search with cross-validation, prioritizing the recall-macro metric due to the imbalance of the dependent variable, which represented the presence or absence of supermarkets within a 50-meter buffer. The results indicated that supermarket location is strongly associated with demographic, income, and population characteristics in the surrounding area. The analysis of variable importance showed that sociodemographic factors, such as elderly literacy, household income, and the presence of other food establishments, exerted significant influence, especially in the immediate vicinity. The findings reinforced the hypothesis that the spatial distribution of supermarkets is not random, being conditioned by socioeconomic characteristics and the commercial structure of the territory, offering subsidies for business decisions and urban planning.

Keywords: Spatial Analysis; Machine learning; Expansion; Commercial location; Supermarkets.

Neuroscience And Learning In Education

October 02, 2026

Anti-Racist Education: Inclusive Educational Practices and Social Development

Antiracist education, understood as a structuring axis of inclusive education and social development, was investigated in the Brazilian context. The study aimed to identify and analyze, based on legal documents and teachers’ perceptions, educational practices capable of promoting antiracism in school and society, and how the implementation of Laws nº 10.639/03 and nº 11.645/08 contributed to social justice. A qualitative and documentary approach was adopted, with analysis of educational legislation, curricular guidelines, institutional reports, and academic literature. Complementarily, a semi-structured questionnaire was applied to 295 Basic Education teachers. The data were evaluated quantitatively and qualitatively, through thematic content analysis, and validated with bibliographic studies. The results revealed a paradox: despite a robust legal framework, the implementation of antiracist policies proved fragile and sporadic, with a lack of teacher training, adequate teaching materials, and monitoring. Significant educational inequalities between white and black students were found to persist, and most teachers acknowledged the occurrence of racism in schools, but without clear institutional protocols. Neuroscientific analysis showed that racism negatively impacts students’ cognitive and emotional development. It was concluded that antiracist education is central to quality education, requiring political commitment, public investment, and intersectoral articulation. The integration of Neuroscience in teacher training and the production of qualified materials are crucial to strengthen the school’s role in building a more just and inclusive society.

Keywords: Social Development; Antiracist Education; Social Justice; Law 10.639/03; Inclusive Educational Practices.

Neuroscience And Learning In Education

October 02, 2026

Paths of Inclusion: Perceptions of Parents and Teachers on the Schooling of Students with Dual Exceptionality in the Brazilian Context

Dual Exceptionality, characterized by the coexistence of High Abilities/Giftedness and neurodevelopmental disorders, represents a complex phenomenon that challenges traditional identification and schooling models. The study aimed to understand the perceptions of parents or guardians, teachers, and other education professionals regarding the schooling of students with Dual Exceptionality in the Brazilian context, investigating challenges, pedagogical strategies, and possibilities for inclusion based on equity. The research adopted a qualitative, exploratory, and descriptive approach, and collected data through an online, voluntary, and anonymous questionnaire answered by 25 participants. Discursive data were analyzed using thematic content analysis. The results indicated that knowledge about the topic is often built from personal and professional experiences, revealing gaps in systematic training. Difficulties were identified in identifying these students, in teacher training, and in implementing individualized educational plans, pedagogical flexibility, and curriculum enrichment. Socio-emotional repercussions, such as frustration and low self-esteem, were reported. However, some schools demonstrated inclusive practices based on equity, articulating specific needs and potentialities. Although the results do not allow for generalizations, they highlighted the need to strengthen professional training and the articulation between school, family, and specialized services. It was concluded that the inclusion of students with Dual Exceptionality requires practices that simultaneously recognize their difficulties and potentialities, ensuring equitable conditions for participation, learning, and development.

Keywords: Human development; Teacher training; School inclusion; Neurodivergence; Pedagogical practices.

October 02, 2026

Data Transformation into Strategy: Applied Research for Ecotourism Operation Optimization

The growing demand in ecotourism in Minas Gerais has driven the search for business intelligence to transform customer data into strategic information. The study aimed to structure a data science pipeline to collect, segment, and classify the customer base of an ecotourism operation, in order to optimize marketing actions and anticipate market movements. An exploratory, quali-quantitative research was conducted through a case study. 2,777 transactional records from an ecotourism company, referring to January 2024 to December 2025, were used. The methodological process involved automated data collection (Google Sheets API), processing and enrichment (ETL), validation, and creation of RFM (Recency, Frequency, and Monetary Value) attributes. Dimensionality reduction via PCA and K-Means clustering was applied, with the number of clusters defined by the Elbow method and Silhouette Score. The results were validated with DBSCAN and K-Medoids. The results revealed the identification of three behavioral customer segments: “Loyal”, “Low Value”, and “Potential”. The “Loyal” segment represented the highest accumulated economic value, while the “Potential” segment stood out for its high average ticket and potential for conversion into recurrence. The integration of data analysis techniques proved to be a robust and replicable method for generating intelligence in ecotourism. It was concluded that the structured data science pipeline enabled the behavioral segmentation of the customer base, the statistical validation of the groups, and the creation of a predictive system for new buyers, providing subsidies for data-driven strategic decisions and future analyses.

Keywords: Clustering; Business intelligence; Machine Learning; Customer segmentation; Decision making.

October 02, 2026

Classification of defaulting customers using supervised machine learning techniques

The risk of default in credit operations demanded analytical approaches to anticipate losses. This study comparatively evaluated the performance of supervised machine learning models in classifying defaulting customers in credit card operations. The public dataset “Default of Credit Card Clients” from the University of California Irvine was used, with 30,000 observations and class imbalance. The algorithms Logistic Regression, Random Forest, and Extreme Gradient Boosting were employed. The imbalance was addressed by assigning weights to the classes, and model optimization occurred with the RandomizedSearchCV method, prioritizing sensitivity. Cross-validation results indicated that the Extreme Gradient Boosting model showed a higher capacity for identifying the defaulting class and better discriminatory performance, followed by Random Forest and Logistic Regression, with a sensitivity of 0.8250 and an AUC-ROC of 0.7844 for XGBoost. Interpretability analysis, conducted by the Shapley Additive Explanations (SHAP) technique, highlighted the predominance of variables associated with payment behavior, especially the history of delays. It was concluded that tree-based models, particularly boosting techniques, proved to be more suitable for capturing complex patterns in the data, configuring themselves as consistent alternatives for credit risk management.

Keywords: Machine Learning; Credit Card; Classification; Extreme Gradient Boosting; Credit Risk.

October 02, 2026

Sentiment Analysis on Brazilian Banks on Twitter/X: Comparison between Traditional and Digital Institutions

A study analyzed public perception of Brazilian financial institutions on the Twitter/X platform, highlighting the importance of sentiment monitoring on social networks for understanding reputation and customer experience in the banking sector. The objective was to compare user perception of the image and reputation of traditional and digital banks, based on the sentiment patterns identified in the analyzed manifestations, seeking to identify structural differences between these groups. The methodology was based on the analysis of 1,096 tweets collected between November 2022 and June 2023. Two complementary sentiment analysis approaches were used, the sum and the average of labels, to capture the majority sentiment and nuances of perception. Additionally, the Market Profile Model, with indicators of emotional reputation, reputational risk, neutrality, and polarization, and the Banking Clustering Model, which allowed grouping institutions according to perception patterns, were developed. The results indicated a predominance of neutral and negative sentiments, a higher volume of interactions in digital banks, and structural differences in the emotional intensity of perceptions, with greater stability in digital banks and greater polarization in traditional ones. It was concluded that the combination of analytical and statistical techniques contributed to an in-depth understanding of institutional image in the digital environment, demonstrating the importance of data-driven reputation management strategies.

Keywords: Digital banks; Traditional banks; Data modeling; Opinion mining; Social Networks.

October 02, 2026

Optimization of annual budget planning through project management methodologies

The Annual Budget Planning (POA) is a crucial process for translating organizational strategy into operational and financial goals, but it frequently faces deadline pressures, interdepartmental dependencies, and the repetition of habitual expenses. The study aimed to analyze how the combined application of project management practices and Zero-Based Budgeting (OBZ) can optimize the POA. To this end, a case study was developed in the Brazilian operation of a publicly traded company in the beverage sector, using documentary research of its 2023 results report and an anonymous questionnaire applied to 47 respondents. Documentary analysis indicated growth in net revenue, expansion of gross profit and adjusted EBITDA, and contained advancement of selling, general, and administrative expenses, suggesting cost discipline and operational leverage. The complementary survey revealed a high perception of cascading effect on the schedule, strong support for defining cost package owners, and a preference for technical justification of expenses, in addition to demand for controlled flexibility after the baseline definition. It was concluded that structuring the POA as a project, associated with the rigor of OBZ, increased the process predictability, reinforced accountability for expenses, and broadened the coherence between budgetary execution and economic-financial performance.

Keywords: Cost Control; Operational Efficiency; Zero-Based Budgeting; PMBOK; Beverage Sector.