The importance of data organization in the AI era

Column

Artificial Intelligence

October 22, 2025

The importance of data organization in the AI era

Because data quality and context are a competitive differentiator in the AI era

Today, artificial intelligence occupies the center of strategic conversations in organizations. More than an individual perception, its constant presence in events, training, and corporate meetings is confirmed by data. The Artificial Intelligence Index Report 2025, published by Stanford University — considered the most reliable and comprehensive report on AI in the world — confirms the expressive increase in the relevance of the topic. In its 8th edition, the document shows that the adoption of technology has grown rapidly, and private investments have surpassed the mark of 252 billion dollars globally. Artificial intelligence is no longer an experimental resource but has become a structural part of decisions, products, and strategies.

This advance is evident in the following graph, which shows that 78% of organizations already use AI in at least one function, while 71% have adopted generative AI solutions. The leap between 2023 and 2024 highlights the acceleration of the adoption of these technologies, which, in addition to automating processes, now interact naturally and create.

Respondents who say that their organization uses AI in at least one function, 2017–2024
Source: McKinsey & Company Survey, 2024.

Despite the popularity of the topic, many companies are still not preparing their internal environments to extract the best from the technology. Looking at this scenario with pragmatism, investing in structure, organization, and clarity about the data that feeds the models, can represent a real competitive advantage in the coming years. It is essential to develop AI literacy : understand the fundamentals of the technology to participate critically in decisions. This text does not seek to cover all aspects of artificial intelligence, but focuses on an often-overlooked step: the quality of the inputs, especially the data.

In simplified terms, AI models are statistical algorithms that process probabilities based on large volumes of data. In the case of LLMs (Large Language Models), such as those behind popular tools like ChatGPT or Gemini, we are talking about billions — or even trillions — of parameters fed by massive databases. Understanding this scale is challenging: an 8-billion-parameter model is considered small, while DeepSeek, with 80 billion, is already classified as large (and there are much larger models).

But larger models do not necessarily mean better models. Size is linked to the capacity to cover more contexts, not to the intrinsic quality of the responses. A model trained with billions of data points may know as much about recipes as about programming, but this does not guarantee accuracy in every domain. This is precisely where both the potential and the challenge lie: the quality of the data used ultimately defines the quality of the responses that these models are capable of delivering.

Imagine training a model only with gastronomy data, leaving out information about programming. It can become extremely accurate in recipes, but incapable of responding about code, and vice versa. This selectivity reveals what is becoming increasingly relevant: models fed with contextualized, curated, and structured data.

A large part of the so-called “hallucinations” generated by models occurs precisely due to the direct influence of inputs on the output. Sam Altman, CEO of OpenAI, recently stated: “People have a very high degree of trust in ChatGPT, which is interesting, because AI hallucinates. It should be the technology you trust the least.” Excessive trust, combined with a low understanding of how models work, increases the risk of strategic errors and decisions based on superficial responses.

There is much talk about prompt engineering — the discipline of formulating text commands to obtain more assertive responses. But, before the prompt, it is the quality and preparation of the data that determine the model’s efficiency. It is in this context that SLMs (Small Language Models), smaller and more specific models, trained with segmented bases, gain strength. Instead of trying to cover all contexts, they are designed to act with depth in delimited domains.

Lucas de Oliveira and Felipe Silva, specialists in technology and machine learning, highlight that, in addition to presenting greater accuracy, these models are also more energetically efficient. One example is the Phi-3-mini, a 3.8 billion parameter model that achieved performance equivalent to PaLM (540B) in MMLU tests — a 142-fold leap in efficiency.

However, using generic models, without contextual training or a proprietary base, increases the risk of inconsistent results and even exposure of sensitive information. The mere act of sending confidential files, spreadsheets, or texts to public AI tools already represents a risk to data privacy and security, as this information can be incorporated into the training processes of large models.

I talk about the importance of AI literacy because, when we address the topic, different layers of discussion emerge. A recurring term in these conversations is tuning — or, simply put, training models. This process consists of providing data so that the model learns and adapts to a given context. The quantity and quality of this information directly influence the outcome: it is possible to offer a reduced volume of data and obtain a small and specialized model, or to work with a gigantic mass and create broad and generalist models. In essence, training is the way to organize and direct a model’s behavior for a specific context — like the cited example of an AI specialized in gastronomy and without programming context.

Each organization possesses unique data about its operations, customers, and markets. Treating this information systematically — organizing it in a structured and standardized manner — creates a balanced dataset that facilitates usage strategies, whether in traditional statistical models or in prediction algorithms and artificial intelligences aimed at creating scenarios and strategic plans.

These structured datasets are called datasets: organized collections of information ready for analysis. From internal datasets, it is possible to build solutions capable of acting as contextualized support agents, aligned with the reality and language of each business. More than a technical investment, it is a cultural shift: preparing people, processes, and workflows to deal with data strategically and consciously.

In this sense, documenting well has never been so important. Meeting minutes, records, task traceability, and process standardization form a knowledge base that, in the future, can feed internal AI models with context and precision. Organizations that already had this culture have an advantage: their data is more equalized and, therefore, can be used. Those that did not maintain this care need to act now — even if reactively —, creating the habit of recording and organizing from now on. The good news is that AI tools themselves can help in this process.

Videoconferencing platforms already offer transcription features; generative models help apply patterns in notes and reports; and other tools are capable of summarizing large volumes of information. The use of these AIs assists in preparing datasets for training new models — or simply provides additional context for prompts, increasing the accuracy of the results.

Predictive models, when well-structured, help anticipate risks and opportunities. But for this to happen, something even more fundamental is needed: good inputs and AI-literate people. Working in a post-generative AI world is not just about creating good prompts or mastering tools. It’s understanding that models reflect the data they receive: if the input is confusing, biased, or disorganized, the output will be too. In essence, artificial intelligence is an information translation technology, and its use requires responsibility and critical thinking about the data that feeds the system.

Therefore, organizing information consciously, structured, and ethically is perhaps the biggest challenge — and the biggest competitive differentiator — for organizations in the coming years. The assertiveness of models stems from the quality of inputs. And digital maturity will be increasingly measured by the ability to understand, protect, and value one’s own data. Open and external data are available to everyone, but the data that organizations generate and have the privilege to access will be their true differentiator. A piece of data may even exist, but it only becomes valuable when it is organized and ready to be processed.

To access the references of this text click here.

Who wrote this column

Lucas Tangi

Lucas é Design Manager no Pecege, formado em tecnologia e especialista em gestão de equipes criativas. Com ampla experiência liderando equipes de design, é também palestrante, professor e consultor. Já participou de projetos em consultorias de tecnologia, venture builders, ODS e na amazônia brasileira. Entusiasta e pesquisador de futuros, dedica-se à inovação e à criação de soluções com alto impacto social e econômico.

You may also like

October 05, 2026

Prediction of default on tax debt installments

The installment payment of tax debts constitutes a relevant fiscal recovery instrument, allowing taxpayers to regularize their obligations in installments, while exposing the tax administration to the risk of cancellation due to non-compliance. The objective was to develop and evaluate machine learning models for predicting the cancellation of ICMS installments, aiming to support proactive collection strategies. A database of historical installment plans from the Secretariat of Economy of the State of Goiás was used. Decision Tree, Random Forest, and XGBoost algorithms were trained and compared, with hyperparameter optimization via GridSearchCV and evaluation by confusion matrices and ROC curves. XGBoost showed superior performance, with an AUC of 0.869 for the model that included the variable Number of Installments and 0.778 without it, highlighting the centrality of this feature. Applied to the active portfolio of R$ 2.752 billion, the model identified that 92.9% of the total value presented a risk of cancellation equal to or greater than 50%, with R$ 1.488 billion concentrated in extreme risk installments, a result consistent with the historical behavior of cancellations in the initial phases of agreements. The results demonstrated that the adoption of predictive models in the management of tax installments enables the transition from a reactive stance to a preventive approach, with the potential to substantially increase fiscal recovery.

Keywords: Tax administration; Machine learning; Binary classification; Fiscal recovery; XGBoost.

October 05, 2026

Interactive system for generation and comparison of predictive models of monthly rural credit concessions

The ability to predict the future volume of rural credit concessions is essential for efficient resource allocation and setting disbursement targets. The work aimed to develop an interactive web platform to generate, analyze, and compare predictive models of time series of rural credit concessions, covering the period from March 2011 to January 2026. Econometric techniques (ARIMA, SARIMA, SARIMAX) and linear regression were confronted with machine learning methods (Random Forest, XGBoost). The methodology included collecting monthly data on rural credit concessions, macroeconomic variables, and agricultural commodity indicators, followed by exploratory analysis and platform development in a client-server architecture with Python and web technologies. The platform’s application to rural credit forecasting in three scenarios (total, individual, and corporate) evaluated by temporal cross-validation revealed that no technique proved universally superior. Linear regression models with seasonal lags showed the most consistent results across all folds, maintaining stable performance even in periods of level shift, where decision tree algorithms registered significant degradation. The corporate scenario showed low predictability in all models. The results demonstrated that the platform fulfilled its objective, revealing that algorithmic complexity does not guarantee predictive superiority and that the choice of model should be guided by the series’ characteristics and the application context.

Keywords: Agribusiness; Forecasting; Machine learning; Predictive modeling; Time series.

October 05, 2026

Music festivals as a platform for brand engagement with Gen Z consumers

Music festivals have consolidated themselves as complex experience ecosystems, where the convergence between physical entertainment and digital narratives redefines brand positioning strategies. In the post-pandemic scenario, the events sector showed a significant recovery, establishing itself as a strategic platform for engaging with Generation Z, an audience that prioritizes authenticity and shared experiences over traditional advertising formats. The study aimed to analyze how brand activations in these environments impacted the engagement of consumers born between 1995 and 2010. The methodology was characterized by a quantitative and descriptive research, conducted through the application of a structured questionnaire that obtained the participation of 163 respondents. The results showed that experience marketing strategies, especially those that integrated aesthetic attributes and digital sharing potential, presented the highest averages of positive perception. It was found that brand trust was strengthened after successful physical interactions, revealing a symbiosis between the emotional environment of the event and the young person’s digital journey. In contrast, the perception of authenticity mediated by digital influencers obtained the lowest agreement index. It was concluded that Generation Z’s engagement was enhanced by hybrid strategies that allowed consumers to take on the role of protagonist in building the brand narrative, prioritizing the authenticity of direct experience.

Keywords: Consumer behavior; Cultural consumption; Transmedia strategies; Digital influencers; Experience marketing.

Compliance And Esg

October 05, 2026

Compliance to mitigate and prevent theft in construction sites: Integrity program applied to civil construction

Thefts at construction sites represent a significant challenge for the civil construction industry, generating financial, operational, and reputational impacts. In this context, integrity and compliance programs have emerged as management tools to strengthen governance and mitigate risks. The study aimed to propose a compliance program applied to civil construction, focused on preventing and mitigating thefts at construction sites. A qualitative approach, with quantitative support, was adopted, developed in two stages. Firstly, a field survey was conducted between February and March 2026, applying an electronic questionnaire to industry professionals. Subsequently, a fictitious case study was developed, based on the author’s professional experiences, integrating the research results with risk management and governance practices. The results highlighted the importance of implementing compliance programs in the sector and indicated that the combination of control mechanisms, structured processes, team training, reporting channels, continuous monitoring, and strengthening of an ethical culture can reduce vulnerabilities. It was concluded that the adoption of integrity practices enhances organizations’ preventive capacity and improves governance and risk management in the sector.

Keywords: Compliance; Civil construction; Theft; Risk management; Corporate governance.

October 05, 2026

Multi-signal panel for optimizing vulnerability prioritization in cybersecurity

The growing proliferation of cyber vulnerabilities has rendered prioritization based solely on static severity metrics inadequate. A multi-signal analytical panel was developed and evaluated to optimize cyber vulnerability prioritization, integrating Common Vulnerability Scoring System (CVSS), Exploit Prediction Scoring System (EPSS), criteria derived from Stakeholder-Specific Vulnerability Categorization (SSVC), and the Known Exploited Vulnerabilities (KEV) catalog as the supervised target variable. 137,854 vulnerability records from NVD (2022–2025) were consolidated, with 607 confirmed KEVs (0.44%), characterizing a classification problem with a highly imbalanced class. A supervised Random Forest model was trained with ten non-circular variables and evaluated using metrics suitable for imbalance. The model achieved an AUC-ROC of 0.9869 and AUC-PR of 0.4959, outperforming isolated EPSS on both metrics (AUC-ROC=0.9383 and AUC-PR=0.3231). Discrepancy analysis with the deterministic SSVC approach revealed that 83.4% of vulnerabilities were over-prioritized. Of the 4,584 CVEs classified as P0, 4,029 represented high-imminent-risk vulnerabilities not confirmed as exploited, flagged by the model. It was concluded that the multi-signal approach goes beyond querying the KEV catalog, identifying vulnerabilities with high future exploitation potential and guiding proactive remediation prioritization.

Keywords: Exploit prediction; Risk management; Vulnerability management; Known Exploited Vulnerabilities (KEV); Random Forest.

Neuroscience And Learning In Education

October 05, 2026

Meditation: a tool that collaborates with the teacher in the classroom

The study aimed to identify meditative practices, their benefits, limitations, and application possibilities in the school context, with a scientific focus and based on existing publications, seeking to expand strategies for promoting mental health and well-being of students and teachers. An exploratory and descriptive research was conducted, with a qualitative approach, based on a literature review of scientific articles, books, and documents published in the last ten years, and on the researcher’s experience reports. The methodology included the scientific definition of meditation and the analysis of studies on meditation and mindfulness in school settings, incorporating contributions from neuroscience, but avoiding biological reductionism of learning. As a main result, a Booklet of Meditative Practices for Teachers was developed, conceived as a complementary pedagogical support material. The findings indicated that meditation can contribute to the reduction of symptoms of stress, anxiety, and depression, in addition to fostering attention, self-reflection, empathy, and coexistence. It was concluded that the booklet offers simple and adaptable guidelines for the classroom, serving as a complementary tool that does not replace pedagogical interventions or public policies. The sustainability of these practices requires articulation with the pedagogical project, institutional support, and teacher training, integrating neuroscience, pedagogy, and practical experience for an education more attentive to the cognitive, emotional, social, and existential dimensions of students.

Keywords: School environment; Mindfulness; Meditation; Neuroscience; Mental health.