Artificial Intelligence
October 22, 2025
The importance of data organization in the AI era
Because data quality and context are a competitive differentiator in the AI era

Today, artificial intelligence occupies the center of strategic conversations in organizations. More than an individual perception, its constant presence in events, training, and corporate meetings is confirmed by data. The Artificial Intelligence Index Report 2025, published by Stanford University — considered the most reliable and comprehensive report on AI in the world — confirms the expressive increase in the relevance of the topic. In its 8th edition, the document shows that the adoption of technology has grown rapidly, and private investments have surpassed the mark of 252 billion dollars globally. Artificial intelligence is no longer an experimental resource but has become a structural part of decisions, products, and strategies.
This advance is evident in the following graph, which shows that 78% of organizations already use AI in at least one function, while 71% have adopted generative AI solutions. The leap between 2023 and 2024 highlights the acceleration of the adoption of these technologies, which, in addition to automating processes, now interact naturally and create.

Source: McKinsey & Company Survey, 2024.
Despite the popularity of the topic, many companies are still not preparing their internal environments to extract the best from the technology. Looking at this scenario with pragmatism, investing in structure, organization, and clarity about the data that feeds the models, can represent a real competitive advantage in the coming years. It is essential to develop AI literacy : understand the fundamentals of the technology to participate critically in decisions. This text does not seek to cover all aspects of artificial intelligence, but focuses on an often-overlooked step: the quality of the inputs, especially the data.
In simplified terms, AI models are statistical algorithms that process probabilities based on large volumes of data. In the case of LLMs (Large Language Models), such as those behind popular tools like ChatGPT or Gemini, we are talking about billions — or even trillions — of parameters fed by massive databases. Understanding this scale is challenging: an 8-billion-parameter model is considered small, while DeepSeek, with 80 billion, is already classified as large (and there are much larger models).
But larger models do not necessarily mean better models. Size is linked to the capacity to cover more contexts, not to the intrinsic quality of the responses. A model trained with billions of data points may know as much about recipes as about programming, but this does not guarantee accuracy in every domain. This is precisely where both the potential and the challenge lie: the quality of the data used ultimately defines the quality of the responses that these models are capable of delivering.
Imagine training a model only with gastronomy data, leaving out information about programming. It can become extremely accurate in recipes, but incapable of responding about code, and vice versa. This selectivity reveals what is becoming increasingly relevant: models fed with contextualized, curated, and structured data.

A large part of the so-called “hallucinations” generated by models occurs precisely due to the direct influence of inputs on the output. Sam Altman, CEO of OpenAI, recently stated: “People have a very high degree of trust in ChatGPT, which is interesting, because AI hallucinates. It should be the technology you trust the least.” Excessive trust, combined with a low understanding of how models work, increases the risk of strategic errors and decisions based on superficial responses.
There is much talk about prompt engineering — the discipline of formulating text commands to obtain more assertive responses. But, before the prompt, it is the quality and preparation of the data that determine the model’s efficiency. It is in this context that SLMs (Small Language Models), smaller and more specific models, trained with segmented bases, gain strength. Instead of trying to cover all contexts, they are designed to act with depth in delimited domains.
Lucas de Oliveira and Felipe Silva, specialists in technology and machine learning, highlight that, in addition to presenting greater accuracy, these models are also more energetically efficient. One example is the Phi-3-mini, a 3.8 billion parameter model that achieved performance equivalent to PaLM (540B) in MMLU tests — a 142-fold leap in efficiency.
However, using generic models, without contextual training or a proprietary base, increases the risk of inconsistent results and even exposure of sensitive information. The mere act of sending confidential files, spreadsheets, or texts to public AI tools already represents a risk to data privacy and security, as this information can be incorporated into the training processes of large models.
I talk about the importance of AI literacy because, when we address the topic, different layers of discussion emerge. A recurring term in these conversations is tuning — or, simply put, training models. This process consists of providing data so that the model learns and adapts to a given context. The quantity and quality of this information directly influence the outcome: it is possible to offer a reduced volume of data and obtain a small and specialized model, or to work with a gigantic mass and create broad and generalist models. In essence, training is the way to organize and direct a model’s behavior for a specific context — like the cited example of an AI specialized in gastronomy and without programming context.
Each organization possesses unique data about its operations, customers, and markets. Treating this information systematically — organizing it in a structured and standardized manner — creates a balanced dataset that facilitates usage strategies, whether in traditional statistical models or in prediction algorithms and artificial intelligences aimed at creating scenarios and strategic plans.
These structured datasets are called datasets: organized collections of information ready for analysis. From internal datasets, it is possible to build solutions capable of acting as contextualized support agents, aligned with the reality and language of each business. More than a technical investment, it is a cultural shift: preparing people, processes, and workflows to deal with data strategically and consciously.
In this sense, documenting well has never been so important. Meeting minutes, records, task traceability, and process standardization form a knowledge base that, in the future, can feed internal AI models with context and precision. Organizations that already had this culture have an advantage: their data is more equalized and, therefore, can be used. Those that did not maintain this care need to act now — even if reactively —, creating the habit of recording and organizing from now on. The good news is that AI tools themselves can help in this process.
Videoconferencing platforms already offer transcription features; generative models help apply patterns in notes and reports; and other tools are capable of summarizing large volumes of information. The use of these AIs assists in preparing datasets for training new models — or simply provides additional context for prompts, increasing the accuracy of the results.
Predictive models, when well-structured, help anticipate risks and opportunities. But for this to happen, something even more fundamental is needed: good inputs and AI-literate people. Working in a post-generative AI world is not just about creating good prompts or mastering tools. It’s understanding that models reflect the data they receive: if the input is confusing, biased, or disorganized, the output will be too. In essence, artificial intelligence is an information translation technology, and its use requires responsibility and critical thinking about the data that feeds the system.
Therefore, organizing information consciously, structured, and ethically is perhaps the biggest challenge — and the biggest competitive differentiator — for organizations in the coming years. The assertiveness of models stems from the quality of inputs. And digital maturity will be increasingly measured by the ability to understand, protect, and value one’s own data. Open and external data are available to everyone, but the data that organizations generate and have the privilege to access will be their true differentiator. A piece of data may even exist, but it only becomes valuable when it is organized and ready to be processed.
| To access the references of this text click here. |
Who wrote this column
Lucas Tangi








