October 02, 2026
Determinants of supermarket location in São Paulo
Determinants of Supermarket Location in São Paulo
Fernando Lima Trambacos; Hugo Bampi
DOI: 10.22167/2675-6528-202602900
Article derived from a Final Course Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.
Summary
A study investigated the determining factors for supermarket location in the state of São Paulo, with the objective of examining the factors that explain the presence and expansion of these establishments, considering socioeconomic, demographic, and market dimensions. Data from the 2010 and 2022 Demographic Censuses of IBGE and information from the National Registry of Legal Entities of the Federal Revenue of Brazil were used to construct a georeferenced database. A Random Forest classification model was applied, adjusted by grid search with cross-validation, prioritizing the recall-macro metric due to the imbalance of the dependent variable, which represented the presence or absence of supermarkets within a 50-meter buffer. The results indicated that supermarket location is strongly associated with demographic, income, and population characteristics in the surrounding area. The analysis of variable importance showed that sociodemographic factors, such as elderly literacy, household income, and the presence of other food establishments, exerted significant influence, especially in the immediate vicinity. The findings reinforced the hypothesis that the spatial distribution of supermarkets is not random, being conditioned by socioeconomic characteristics and the commercial structure of the territory, offering subsidies for business decisions and urban planning.
Keywords: Spatial Analysis; Machine learning; Expansion; Commercial location; Supermarkets.
1. Introduction
Retail constitutes one of the most relevant sectors of the Brazilian economy, standing out both for its impact on the Gross Domestic Product (GDP) and for its social contribution in generating jobs and income. In 2023, the segment generated a net operating revenue of 7.3 trillion reais (IBGE, 2023), consolidating itself as one of the main drivers of national economic activity. In addition to its macroeconomic expressiveness, the sector excels in its capacity to absorb labor, employing approximately 10.6 million people in 2024 (MTE, 2024), which represents a significant portion of the Brazilian workforce. This characteristic gives retail great importance for economic and social development, not only due to the volume transacted but also because of its widespread presence in practically all municipalities in the country, reaching diverse social strata and regions of the territory.
From a theoretical and practical point of view, retail is directly associated with household consumption behavior, functioning as an essential link between production and final demand. Its proximity to the final consumer means that its operational dynamics sensitively reflect variations in purchasing power, consumption trends, urban transformations, and regional development policies. In this sense, retail transcends the economic sphere, configuring itself as a social and territorial phenomenon whose understanding is fundamental to analyzing the organization of urban space and Brazil’s regional dynamics.
Among the various segments that make up the retail sector, the supermarket segment holds a prominent position, being the largest of them. Supermarkets possess characteristics that make them strategic for the economy and society. This business format concentrates a great diversity of products, especially food items, meeting a basic and universal need of the population. Their relevance in daily supply gives supermarkets a unique role, making them structuring elements of mass consumption and agents of urban space organization. The establishment and expansion of these businesses not only impact the circulation of people and goods but also directly influence real estate development and the configuration of commercial centralities in cities.
Despite the sector’s relevance, academic literature has not yet thoroughly addressed the determinants of supermarket location in Brazil. The absence of systematic studies on the topic represents a significant gap, as understanding the variables that explain the installation and expansion of this segment can generate valuable information for different audiences. This gap prevents a more complete analysis of the urban and economic dynamics associated with this type of commerce.
For private managers, understanding these factors constitutes a strategic input for decisions on territorial expansion and investment allocation. For urban and economic public policy makers, understanding these elements aids in city planning and in promoting more equitable access for the population to essential consumer goods. This work seeks to fill this gap, focusing on the state of São Paulo, through the integration of official databases, such as the 2010 and 2022 Demographic Censuses from the Brazilian Institute of Geography and Statistics (IBGE) and the National Registry of Legal Entities from the Federal Revenue of Brazil, using modern data analysis techniques, especially the random forest methodology. The general objective of this study is to investigate the factors that explain the presence and expansion of supermarkets in the territory of the state of São Paulo, considering socioeconomic, demographic, and market dimensions.
2. Material and Methods
The methodology of this study was characterized by a quantitative approach, combining the collection, integration, and analysis of secondary data with the use of advanced modeling techniques. The objective was to investigate the factors that explain the presence and expansion of supermarkets in the territory of the state of São Paulo, as outlined in the introduction. The procedures were structured to capture and analyze sociodemographic, income, and commercial establishment presence information at different spatial scales.
For the construction of the database, official public sources were used. The main ones were the Demographic Censuses of 2010 (IBGE, 2011) and 2022 (IBGE, 2023), which provided sociodemographic and income data. Additionally, the database of the National Registry of Legal Entities (CNPJ) from the Federal Revenue of Brazil (RFB, 2025) was employed, which enabled the identification of the location and profile of Brazilian companies.
The collection and organization of data occurred in sequential stages. Initially, data from the 2010 and 2022 Demographic Censuses were obtained and aggregated at the census tract level. From these bases, variables representative of sociodemographic characteristics, such as population, households, household types, gender, age group, and literacy, and income characteristics of the population, including the number of households by income bracket and average household income, were selected.
In the subsequent stage, addresses of active CNPJs were extracted from the Brazilian Federal Revenue database (RFB, 2025). Specific retail and service segments were considered, such as food, wholesale, butcher shops and fish markets, pharmacies, and hyper/supermarkets, among others. These addresses were geocoded using the ArcGIS geolocation API (ESRI, 2025), allowing for precise identification of their spatial location in the state of São Paulo.
Subsequently, a theoretical grid of 1,242,215 points, spaced every 50 meters, was created, covering all census tracts classified as urban in the 2022 Demographic Census (IBGE, 2023) in the state of São Paulo. For each of these points, buffers were constructed, which are circular areas with radii of 50, 500, 1,000, 2,000, 3,000, 4,000, and 5,000 meters, with the purpose of capturing information from the immediate and extended surroundings of each theoretical location.
Census information was aggregated to each buffer, with proportionalization for partially contained sectors. For business data, CNPJs were counted for each commercial segment within each buffer. The final database, with 1,029 variables, synthesized sociodemographic, income, and commercial establishment presence information for each theoretical point and different buffers. The dependent variable was defined as binary (1 for existence, 0 for absence of supermarkets in the 50-meter buffer), adjusted for a maximum value of 1, given the improbability of multiple distinct establishments in such a restricted area and the study’s focus on presence or absence.
For the data analysis, a Random Forest classification machine learning method was used. This algorithm, based on decision trees, was employed to estimate the probability of supermarket presence, aggregating individual predictions for greater robustness and overfitting reduction. The Random Forest’s ability to capture non-linear relationships and handle large volumes of data was a determining factor in its selection.
In order to optimize the model’s performance, a grid search procedure with cross-validation (three folds, cv = 3) was performed. The combination of hyperparameters that maximized the recall-macro metric was sought, chosen due to the strong imbalance of the dependent variable. This metric ensured the correct identification of areas with supermarket occurrences, prioritizing the reduction of false negatives.
The hyperparameters tuned in the grid search included: `n_estimators` (50, 100, 150), `max_depth` (6, 10, 14), `min_samples_split` (10, 50, 100), and `min_samples_leaf` (5, 20, 50). Additionally, to mitigate the effects of data imbalance, the `sample_weight` procedure was adopted, adjusting the weights of the observations. Less frequent categories received proportionally higher weights, giving greater relative importance to the presence of supermarkets.
3. Results and Discussion
The research revealed significant insights into the factors determining the location of supermarkets in the state of São Paulo, confirming the hypothesis that their distribution is not random, but systematically influenced by specific territorial characteristics. The application of the Random Forest classification model allowed for the identification and interpretation of crucial socioeconomic, demographic, and market factors. These findings offer a comprehensive understanding of the complex interaction of variables that shape the spatial patterns of supermarket presence, contributing to both academic literature and practical decisions in urban planning and retail expansion strategies. The model’s ability to capture these intricate relationships underscores its utility for the study’s central objective.
The Random Forest model optimization process, performed through a grid search with three-fold cross-validation, identified a hyperparameter configuration that maximized the recall-macro metric. This configuration included a maximum tree depth of six, a minimum of five observations per leaf, a minimum of ten observations for splitting, and fifty estimator trees. The choice of recall-macro as the primary metric was essential due to the class imbalance in the dataset, where the absence of supermarkets was significantly more frequent. This approach ensured that the model prioritized the correct identification of areas with supermarkets, minimizing false negatives and aligning with the objective of mapping potential installation sites.
The model’s performance evaluation on the training set demonstrated an accuracy of 0.7162, with a high macro-recall of 0.8058, indicating a good ability to correctly identify classes associated with the presence of supermarkets. Although the macro-precision was relatively low, at 0.5152, reflecting the inherent difficulty in distinguishing between categories in a context of strong imbalance, the weighted F1-score of 0.8249 suggested good overall performance. This weighted metric offers a balanced assessment of the model’s performance, considering both precision and recall, which is fundamental in scenarios with unequal classes.
On the test set, the results remained consistent with those observed during training, presenting an accuracy of 0.7151 and a weighted F1-score of 0.8243. This stability between the training and test sets indicates an absence of relevant overfitting, demonstrating the model’s generalization capability to previously unobserved data. There was a slight reduction in macro-recall to 0.7971 on the test set, signaling a marginally greater difficulty in correctly identifying all classes of supermarket occurrences outside the training sample. However, the overall consistency of the metrics reinforces the model robustness for the exploratory and analytical purposes of the study.
The analysis of the confusion matrix for the test set illustrated the distribution of the model’s classifications. For the majority class, which represents the absence of supermarkets (class 0), the model correctly classified 263,148 observations, while 105,727 were false positives, meaning areas without a supermarket erroneously predicted as having one. For the minority class, indicating the presence of supermarkets (class 1), the model correctly identified 3,338 observations, but classified 452 as false negatives, meaning areas with a supermarket predicted as not having one. This predominance of classifications in the majority class, even with methodological adjustments, reflects the difficulty in discriminating less frequent events in a context of strong data imbalance.
Despite the limitations in precision for the minority class, the overall results confirm the adequacy of Random Forest as an exploratory and explanatory tool for the phenomenon studied. The performance indicators, especially the recall-macro, demonstrate that the model is capable of capturing consistent patterns associated with the location of supermarkets. This capability is particularly valuable for estimating the probabilities of occurrence of these establishments in the urban territory, in line with the study’s objectives. The methodology employed, with the adjustment of sample weights, contributed to mitigating the effects of imbalance, allowing for a more focused analysis on identifying areas with the greatest potential for supermarket installation.
Importance of variables
The analysis of variable importance measures, one of the main results of the Random Forest model, provided robust evidence on the determining factors for supermarket location. The metrics indicated that variables related to the sociodemographic profile and human capital of the population are among the most relevant, exerting significant influence on the occurrence of these establishments. This relevance is particularly accentuated for the data observed in the immediate buffers, of 50 and 500 meters, suggesting that the characteristics of the immediate surroundings are crucial for the installation decision. The identification of these most important variables allows us to understand which attributes of the territory are most predictive of the presence of supermarkets, aligning directly with the objective of investigating the explanatory factors.
Among the variables of greatest importance, the literacy of people aged 60 or over, according to the 2022 Census, stood out as the most relevant individual variable in the 50-meter buffer. This finding suggests that the presence of an older population with a higher level of education in the immediate surroundings is a significant factor for the location of supermarkets. This correlation may indicate that these areas have a more stable demand for daily consumer products, as well as a greater capacity for purchase planning and access to information, making them more attractive to the supermarket sector. The presence of more developed human capital in the immediate neighborhood appears to be an important attraction for the establishment of these businesses.
The variable associated with the presence of other food sector establishments also proved highly relevant, occupying the second position in importance in the 500-meter buffer and a prominent position in the 50-meter buffer. This finding suggests that supermarkets tend to be located in areas already consolidated as commercial hubs, benefiting from economies of agglomeration. The coexistence with other food businesses may indicate a complementarity of activities and a greater consumer flow, which enhances the success of the new venture. This pattern reinforces the idea that location is strategic and seeks commercial synergies, optimizing access to an already established customer base.
Variables related to household income also stood out among the most important, with emphasis on the total income of class C1 in the 50-meter buffer. Furthermore, income indicators and the number of households in classes C2 and DE also featured as explanatory variables, albeit with less relative weight. The predominance of these variables on restricted spatial scales indicates that the supermarket segment prioritizes areas with a critical mass of consumers and purchasing power compatible with its business model. This strategic choice aims to ensure a flow of customers with adequate consumption capacity, optimizing market potential in the immediate vicinity of the facility and ensuring the economic viability of the venture.
The relevance of variables associated with the density of people and households was equally notable. The number of permanent private households from the 2010 Census and the number of occupied permanent private households from the 2022 Census, both within the 50-meter buffer, demonstrated high importance. Specifically, the population aged 60 years or older in the 50-meter buffer also proved to be a relevant factor. These findings suggest that a higher concentration of potential consumers in the vicinity of supermarkets generates a more regular demand for proximity shopping, making these areas more attractive for the establishment of new businesses and for the sustainability of the business.
The analysis of the importance of the variables revealed that the characteristics of the immediate surroundings, specifically in the 50 and 500-meter buffers, exert a preponderant influence in explaining the location of supermarkets. This suggests that expansion and installation decisions are highly sensitive to micro-local conditions, such as the sociodemographic profile of the neighborhood, income level, and existing commercial structure. The model’s ability to identify these patterns at different spatial scales, with emphasis on proximity, reinforces the understanding that the spatial distribution of supermarkets is a complex phenomenon, guided by a strategic combination of factors aimed at optimizing access to the consumer market.
In summary, the empirical results obtained reinforce the central hypothesis of the study that the location of supermarkets in the state of São Paulo is not a random event. On the contrary, the presence of these establishments is strongly conditioned by a combination of sociodemographic characteristics, such as elderly literacy and population density, in addition to household income levels and the commercial structure of the surrounding area, especially the presence of other food establishments. The predominance of the immediate surroundings in explaining these patterns underscores the importance of a detailed analysis of local conditions for strategic location decisions, providing valuable subsidies for urban planning and the retail sector.
4. Conclusion
This study investigated the determining factors for supermarket location in the state of São Paulo, considering socioeconomic, demographic, and market dimensions. It was found that the spatial distribution of these establishments does not occur randomly, being strongly conditioned by specific territorial characteristics. The analysis of findings, obtained through a Random Forest classification model, showed that sociodemographic factors, such as literacy among people aged 60 or over, household income, especially from class C1, and the density of households and population in the immediate vicinity, exerted significant influence. It was also observed that the presence of other food sector establishments nearby also proved to be a relevant factor, suggesting that supermarkets tend to be located in areas already consolidated as commercial hubs, benefiting from economies of agglomeration and an established consumer flow. The predominance of characteristics of the immediate surroundings, in the 50 and 500-meter buffers, underscores the sensitivity of location decisions to micro-local conditions.
The methodology employed, based on random forest and adjusted to prioritize the macro recall metric, proved adequate for handling the high dimensionality of the data and the non-linear relationships between the explanatory variables. Although the model presented limitations in terms of precision for less frequent classes, its overall performance and stability between the training and testing bases indicated sufficient robustness for the exploratory and analytical purposes of the study. The results offer a substantial practical contribution, providing valuable subsidies for companies in the supermarket sector in guiding territorial expansion strategies and identifying areas with greater market potential. For the public sector, the evidence produced contributes to the understanding of the relationship between the distribution of essential services and the socio-spatial structure of cities, potentially supporting urban planning policies and the promotion of more equitable access to consumer goods for the population.
Bibliographic References
ESRI. ArcGIS Geocoding service | ArcGIS REST APIs. Disponível em: https://developers.arcgis.com/rest/geocode/. Acesso em: 01/10/2025.
Instituto Brasileiro de Geografia e Estatística [IBGE]. 2010. Censo Demográfico 2010: resultados gerais da população e domicílios. IBGE, Rio de Janeiro, IBGE, RJ. Disponível em: https://www.ibge.gov.br/estatisticas/sociais/populacao/9662-censo-demografico-2010.html. Acesso em: 21/10/2025.
Instituto Brasileiro de Geografia e Estatística [IBGE]. 2023. Censo Demográfico 2022: resultados preliminares/população. IBGE, Rio de Janeiro, RJ. Disponível em: https://www.ibge.gov.br/estatisticas/sociais/populacao/27069-censo-demografico-2022.html. Acesso em: 21/10/2025.
Ministério do Trabalho e Emprego [MTE]. 2025. RAIS 2024: Sumário Executivo (Parcial). Ministério do Trabalho e Emprego, Brasília, DF. Disponível em: https://www.gov.br/trabalho-e-emprego/pt-br/assuntos/estatisticas-trabalho/rais/rais-2024/rais-2024-parcial/sumario-executivo_rais-2024-parcial.pdf. Acesso em: 21/10/2025.
Receita Federal do Brasil [RFB]. 2025. Cadastro Nacional da Pessoa Jurídica (CNPJs) [base de dados]. Receita Federal do Brasil, Brasília, DF. Disponível em: https://www.gov.br/receitafederal/pt-br/assuntos/orientacao-tributaria/cadastros/cnpj. Acesso em: 21/10/2025.
Article originating from the Final Course Work of the Specialization in Data Science and Analytics of the MBA USP/Esalq
To learn more about the course, click here and access the MBX Academy platform