Article

October 02, 2026

Data Transformation into Strategy: Applied Research for Ecotourism Operation Optimization

Transforming Data into Strategy: Applied Research for Ecotourism Operation Optimization

Fernanda Karla Pinto Morais; Edilson José Rodrigues

DOI: 10.22167/2675-6528-202602896

Article derived from a Final Course Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.

Summary

The growing demand in ecotourism in Minas Gerais has driven the search for business intelligence to transform customer data into strategic information. The study aimed to structure a data science pipeline to collect, segment, and classify the customer base of an ecotourism operation, in order to optimize marketing actions and anticipate market movements. An exploratory, quali-quantitative research was conducted through a case study. 2,777 transactional records from an ecotourism company, referring to January 2024 to December 2025, were used. The methodological process involved automated data collection (Google Sheets API), processing and enrichment (ETL), validation, and creation of RFM (Recency, Frequency, and Monetary Value) attributes. Dimensionality reduction via PCA and K-Means clustering was applied, with the number of clusters defined by the Elbow method and Silhouette Score. Results were validated with DBSCAN and K-Medoids. The results revealed the identification of three customer behavioral segments: “Loyal”, “Low Value”, and “Potential”. The “Loyal” segment represented the highest accumulated economic value, while the “Potential” stood out for its high average ticket and potential for conversion into recurrence. The integration of data analysis techniques proved to be a robust and replicable method for generating intelligence in ecotourism. It was concluded that the structured data science pipeline enabled the behavioral segmentation of the customer base, statistical validation of the groups, and the creation of a predictive system for new buyers, providing subsidies for data-driven strategic decisions and future analyses.

Keywords: Clustering; Business intelligence; Machine Learning; Customer segmentation; Decision making.

1. Introduction

In recent years, ecotourism has consolidated itself as a segment that transcends simple leisure, aligning direct contact with nature with environmental preservation and the valorization of local culture. This sector, supported by the guidelines of the Ministry of Tourism, not only generates socioeconomic benefits but also influences the entire national tourism chain. Minas Gerais, in particular, stands out as one of the main ecotourism hubs in the country, housing Brazil’s only mountain range, the Espinhaço, recognized by Unesco as a Biosphere Reserve, which makes it a strategic setting for the development of this activity (Miranda, 2025).

Following the expressive growth of ecotourism, tourism agencies assume a crucial consulting role for consumers, organizing trips that meet their audience’s needs (Silva, 2016). However, intense competition demands a deep understanding of the market and customers. In this context, data cease to be fragmented information with low analytical value (Olsen, 2015) and become a strategic asset. They require systematic processes of treatment, modeling, and analysis, capable of generating sustainable competitive advantages and transforming informational intelligence into a decisive factor for organizational performance (Fávero e Belfiore, 2024).

The present study focuses on an ecotourism company in Minas Gerais that offers various outdoor experiences, from day trips to expeditions and immersions. Currently, the managers of this company face the challenge of organizing and optimizing the vast volume of data generated by the business. The transition from empiricism-based management to a data-driven approach is essential to deepen the understanding of the customer profile and to comprehend factors such as seasonality and market movements that can influence strategic decisions.

Knowing the customer is fundamental for the success of marketing actions, from defining the target market to choosing the most appropriate approaches for each segment (Solomon, 2016). In ecotourism, this need is even more pressing, as planning a trip is, for the customer, a memorable moment, often surpassing the memories of the experience itself (Silva, 2016). Furthermore, the current scenario requires companies to use their digital social networks as strategic channels to leverage sales and strengthen relationships with the public (Destro, 2023).

Given this scenario, the research is justified by the need to transform raw data into actionable intelligence to optimize the ecotourism operation, and the objective of this work is to structure the data pipeline of the company in question, enabling the identification of customer groups through their transactional records and, through the application of algorithms, optimize strategic marketing actions, increasing their chances of success, in addition to seeking ways to anticipate movements related to the market, seasonality, and issues related to customer behavior.

2. Material and Methods

The exploratory, quali-quantitative, and applied research was conducted as a case study with descriptive elements (Gil, 2002). The study analyzed transactional records from an ecotourism company in Minas Gerais, concerning negotiations from January 2024 to December 2025. The methodology structured a data science pipeline to identify customer groups, optimizing marketing actions and anticipating market movements and behavior.

The primary dataset (Google Sheets) contained 2,777 transactional sales records. Variables included anonymized transaction date, anonymized email and full name, gender, anonymized CPF, date of birth, phone number, pickup location, Instagram (optional), and the purchased event. Processing followed LGPD (Law nº 13.709/2018), with anonymization of direct identifiers. The use of information was authorized by company management for academic purposes, preserving confidentiality and integrity.

The Extract, Transform, and Load (ETL) process was implemented in Python (Spyder/Anaconda). Libraries such as Gspread, Google-auth, and Sys were used for centralization, authentication, and execution control. Pandas (Caetano, 2025) processed and enriched the data, removing non-numerics, validating CPFs, creating status columns, and recovering incomplete records (backfill), according to methodological guidelines (Barbieri, 2020).

In feature engineering, columns were converted to appropriate data types, and the RFM (Recency, Frequency, and Monetary Value) technique was applied. RFM ranked the customer base into quantitative categories, standardizing purchasing behavior. RFM was aggregated to an operational measure of Lifetime Value (LTV), defined as the accumulated monetary value of customer transactions.

For data preparation and transformation, Scikit-learn was used for Z-score standardization (StandardScaler) and for Principal Component Analysis (PCA), configured to preserve 95% of the explained variance (Netto and Maciel, 2021). Standardization ensured that variables contributed equally to the distance calculation (Barr et al., 2024), and PCA eliminated redundancies.

The definition of the ideal number of clusters (K) was performed with the K-Means algorithm (Scikit-Learn), the Elbow method, and the Silhouette Score index (Sharda et al., 2025). Matplotlib and Seaborn assisted in visualization (Netto and Maciel, 2021; Kyran, 2024). The integrated analysis pointed to segmentation into three groups (K=3), representing the best balance between cohesion and separation of clusters (Gomes, 2024; Pierson, 2019).

To validate the results, the DBSCAN and K-Medoids algorithms were applied. DBSCAN was used as a support method to validate the structure of the groups found by K-Means (Faceli et al., 2025). K-Medoids defined real data points (medoids) as cluster centers, providing greater robustness to outliers and noise (Goldschmidt et al., 2015), and its results converged with those of K-Means, validating the structure of three clusters.

The definition of the commercial names of the clusters occurred by dynamic mapping of the statistical profiles of the groupings. At the end of each K-Means round, the average of the Total_Spent variable was calculated for each cluster and, in ascending order, the labels ‘Low Value’, ‘Potential’, and ‘Loyal’ were assigned. This process aligned the business meaning of the segments with the quantitative profiles, according to the area of application (Faceli et al., 2025).

3. Results and Discussion

The analysis of the ecotourism company’s transactional data began with the exploration of relationships between variables, revealing important customer behavioral patterns. A correlation of 0.98 was observed between Total_Spending and Num_Trips, indicating that customers with higher spending volume tend to take more trips. This strong linear dependence, exceeding the 0.90 threshold established by Fávero and Belfiore (2024), confirmed the expected severe multicollinearity, as total spending is intrinsically linked to travel frequency and average ticket price. Additionally, the Recency_Days variable showed a moderate negative correlation of -0.27 with spending and travel quantity variables, suggesting that customers who traveled more recently tend to spend more and have a higher travel frequency.

To mitigate the effects of multicollinearity and optimize the application of clustering algorithms, the Principal Component Analysis (PCA) technique was employed as a preprocessing step. The PCA configuration was set to preserve 95% of the original explained variance of the data, resulting in the generation of four principal components (PC1 to PC4). This transformation ensured that the new components were uncorrelated with each other, guaranteeing orthogonality and mathematical independence of the information. This procedure eliminated redundancies and made the data more suitable for distance-based algorithms, contributing to the stability and interpretability of subsequent results, as highlighted by Netto and Maciel (2021).

The definition of the ideal number of clusters (K) was performed by applying the Elbow Method and the Silhouette Score jointly. The Elbow Method indicated an inflection point in the inertia curve between K=3 and K=4, suggesting diminishing returns in adding new groups from K=5 onwards. The Silhouette Score, in turn, a metric that quantifies the internal cohesion and separation between clusters (Rousseeuw, 1987), pointed to K=2 with the highest coefficient (0.467), followed by K=3 (0.454). Despite the small difference, the choice of K=3 was justified by its greater practical utility for personalizing marketing strategies, offering a superior balance between group cohesion and distinction between them.

The final segmentation, using the K-Means algorithm with K=3, demonstrated a clear spatial separation of customer groups, each with well-defined characteristics. Cluster 0 concentrated the largest part of the customer population, corresponding to a segment with a higher population volume. Cluster 1 encompassed high-engagement and accumulated spending customers, while Cluster 2 brought together customers with a distinct behavioral pattern, characterized by high recency and average ticket size. Dimensionality reduction via PCA, which maintained 95% of the explained variance in the first two components, ensured that the visual representation of the clusters preserved the essential structure for data interpretation.

To validate the robustness of the segmentation obtained with K-Means, the DBSCAN and K-Medoids algorithms were applied. DBSCAN, parameterized with eps=1.5 and min_samples=3, identified three clusters and classified some observations as noise. However, the resulting segmentation was asymmetric, with Cluster 0 concentrating the majority of observations and Clusters 1 and 2 being smaller, in addition to dispersed noise. Although the number of clusters coincided with K=3, the difference in composition and distribution limited the interpretation of a direct qualitative convergence, which is consistent with the limitations of DBSCAN in handling heterogeneous data distributions, according to Faceli et al. (2025).

The K-Medoids algorithm, applied with K=3 on the data transformed by PCA, also confirmed the three-cluster structure. Unlike K-Means, which uses arithmetic means as centroids, K-Medoids selects actual data points (medoids) as cluster centers, which provides greater robustness to the presence of outliers and noise (Goldschmidt et al., 2015). The distribution of groups in K-Medoids, although with a slightly different composition than K-Means (Cluster 0: n=341; Cluster 1: n=99; Cluster 2: n=224, compared to Cluster 0: n=67; Cluster 1: n=555; Cluster 2: n=42 in K-Means), maintained a similar spatial separation, reinforcing the stability of the choice of K=3.

The statistical characterization of the clusters, using boxplots of the original variables on an unstandardized scale, allowed the translation of quantitative profiles into business intelligence. The “Loyal” segment (Cluster 0) stood out as having the highest accumulated economic value, with an average age of 41 years, five acquired trips, a historical LTV of R$994.00, and an average recency of 290 days. Its average ticket was R$203.00. This profile, which demonstrates an active and frequent relationship with the company, is strategic for loyalty actions and frequency reward programs, aiming to maximize its profitability.

The “Low Value” segment (Cluster 1) represented the largest part of the customer base (83.6%), concentrating buyers of occasional experience, predominantly single purchases. This group presented an average age of 37 years, the lowest average ticket in the base (R$181.00), reduced LTV, and high recency, around 420 days. The wide interquartile dispersion in recency suggests low engagement and a higher propensity for abandonment (churn). Converting a portion of this audience into recurring customers represents a significant gain in the operation’s historical LTV, making it a priority for reactivation strategies.

Finally, the “Potential” segment (Cluster 2) stood out for the highest average ticket among the groups, reaching R$294.00, a value approximately 45% higher than the “Loyal” segment. This group presented a younger average age (33 years), low purchase frequency (average of 1.14 trips), and high average recency, around 480 days. The expressive internal heterogeneity of the average ticket, with outliers above R$300.00, indicates that part of the group has a considerably higher spending power than the median. This profile is susceptible to upsell campaigns, post-sale relationship actions, and communications targeted at higher-value experiences, with great long-term conversion potential.

In summary, the structuring of the data science pipeline and the application of clustering algorithms allowed for the identification of three behavioral customer segments — Loyal, Low Value, and Potential — in the ecotourism operation. This segmentation, validated by multiple techniques, offers an in-depth understanding of customer profiles and purchasing behavior, providing concrete subsidies for optimizing strategic marketing actions and anticipating market movements, aligning directly with the study’s central objective of transforming raw data into actionable intelligence.

4. Conclusion

The present study aimed to structure a data science pipeline to collect, segment, and classify the customer base of an ecotourism operation in Minas Gerais, with the goal of optimizing marketing actions and anticipating market movements. It was found that the integration of data analysis techniques, from automated collection (Google Sheets API) to clustering, constituted a robust and replicable method. The application of Principal Component Analysis (PCA) mitigated multicollinearity and optimized the data for the algorithms. Through K-Means, with K=3 defined by the Elbow and Silhouette Score methods, three behavioral customer segments were identified: “Loyal”, “Low Value”, and “Potential”. The “Loyal” segment represented the highest accumulated economic value, while the “Potential” segment stood out for its high average ticket and potential for conversion into recurrence. The segmentation was statistically validated by DBSCAN and K-Medoids, reinforcing the stability of the groups and the coherence of their profiles. This approach enabled the transformation of raw data into actionable intelligence, providing concrete subsidies for data-driven strategic decisions and the creation of a predictive system for new buyers.

Despite the contributions, the study presented limitations, such as the dependence on the quality of data collected via digital forms, which may contain undetected typing errors, and the volume of the validated unique customer base, which restricted the ability to generalize to minority segments. The analysis was also limited by the available transactional variables, preventing a denser understanding of customer behavior through demographic, health, or marketing data. Furthermore, the single case study nature restricts the generalization of identified patterns to other ecotourism operations and long-term observations. For future work, it is suggested to expand the database with new periods to increase representativeness and identify new clustering opportunities. The implementation of a real-time relationship scoring system, data mining through Natural Language Processing (NLP) on customer feedback, the incorporation of the Analytic Hierarchy Process (AHP) or AHP-Gaussian Method into the RFM model for differentiated weighting of dimensions, and the application of algorithms such as Random Forest and LightGBM to create a predictive classification system for new customers, map variable relevance, and estimate the future value of each customer are recommended.

Bibliographic References

FÁVERO, L. P.; BELFIORE, P. 2024. Manual de Análise de Dados: Estatística e Machine Learning com Excel®, SPSS®, Stata®, R® e Python®. 2. ed., LTC. Rio de Janeiro, Brasil.

MIRANDA, T. Espinhaço: A Cordilheira do Brasil em Minas Gerais. Governo do Estado de Minas Gerais, 26 jun. 2025. Disponível em: https://minasgerais.com.br/pt/blog/artigo/espinhaco-a-cordilheira-do-brasil-em-minas-gerais. Acesso em: 25 jan. 2026.

OLSEN, W.. 2015. Coleta de Dados. Penso. Porto Alegre, Brsil. Disponível em: https://app.minhabiblioteca.com.br/reader/books/9788584290543/. Acesso em: 06 fev. 2026.

SILVA, G.G. 2016. Um comparativo entre o planejamento de uma viagem pelo guia PMBOOK e método convencional. Trabalho para obtenção de título de especialista em Gestão de Projetos. Escola Superior de Agricultura “Luiz de Queiroz” (ESALQ), Campus Piracicaba, Universidade de São Paulo (USP). Paulínia, São Paulo, Brasil.

SOLOMON, M.R..

Article originating from the Final Course Work of the Specialization in Data Science and Analytics of the MBA USP/Esalq

To learn more about the course, click here and access the MBX Academy platform

You may also like

October 02, 2026

Determinants of supermarket location in São Paulo

A study investigated the determining factors for supermarket location in the state of São Paulo, with the objective of investigating the factors that explain the presence and expansion of these establishments, considering socioeconomic, demographic, and market dimensions. Data from the 2010 and 2022 Demographic Censuses of IBGE and information from the National Registry of Legal Entities of the Federal Revenue of Brazil were used to build a georeferenced database. A Random Forest classification model was applied, adjusted by grid search with cross-validation, prioritizing the recall-macro metric due to the imbalance of the dependent variable, which represented the presence or absence of supermarkets within a 50-meter buffer. The results indicated that supermarket location is strongly associated with demographic, income, and population characteristics in the surrounding area. The analysis of variable importance showed that sociodemographic factors, such as elderly literacy, household income, and the presence of other food establishments, exerted significant influence, especially in the immediate vicinity. The findings reinforced the hypothesis that the spatial distribution of supermarkets is not random, being conditioned by socioeconomic characteristics and the commercial structure of the territory, offering subsidies for business decisions and urban planning.

Keywords: Spatial Analysis; Machine learning; Expansion; Commercial location; Supermarkets.

Neuroscience And Learning In Education

October 02, 2026

Anti-Racist Education: Inclusive Educational Practices and Social Development

Antiracist education, understood as a structuring axis of inclusive education and social development, was investigated in the Brazilian context. The study aimed to identify and analyze, based on legal documents and teachers’ perceptions, educational practices capable of promoting antiracism in school and society, and how the implementation of Laws nº 10.639/03 and nº 11.645/08 contributed to social justice. A qualitative and documentary approach was adopted, with analysis of educational legislation, curricular guidelines, institutional reports, and academic literature. Complementarily, a semi-structured questionnaire was applied to 295 Basic Education teachers. The data were evaluated quantitatively and qualitatively, through thematic content analysis, and validated with bibliographic studies. The results revealed a paradox: despite a robust legal framework, the implementation of antiracist policies proved fragile and sporadic, with a lack of teacher training, adequate teaching materials, and monitoring. Significant educational inequalities between white and black students were found to persist, and most teachers acknowledged the occurrence of racism in schools, but without clear institutional protocols. Neuroscientific analysis showed that racism negatively impacts students’ cognitive and emotional development. It was concluded that antiracist education is central to quality education, requiring political commitment, public investment, and intersectoral articulation. The integration of Neuroscience in teacher training and the production of qualified materials are crucial to strengthen the school’s role in building a more just and inclusive society.

Keywords: Social Development; Antiracist Education; Social Justice; Law 10.639/03; Inclusive Educational Practices.

Neuroscience And Learning In Education

October 02, 2026

Paths of Inclusion: Perceptions of Parents and Teachers on the Schooling of Students with Dual Exceptionality in the Brazilian Context

Dual Exceptionality, characterized by the coexistence of High Abilities/Giftedness and neurodevelopmental disorders, represents a complex phenomenon that challenges traditional identification and schooling models. The study aimed to understand the perceptions of parents or guardians, teachers, and other education professionals regarding the schooling of students with Dual Exceptionality in the Brazilian context, investigating challenges, pedagogical strategies, and possibilities for inclusion based on equity. The research adopted a qualitative, exploratory, and descriptive approach, and collected data through an online, voluntary, and anonymous questionnaire answered by 25 participants. Discursive data were analyzed using thematic content analysis. The results indicated that knowledge about the topic is often built from personal and professional experiences, revealing gaps in systematic training. Difficulties were identified in identifying these students, in teacher training, and in implementing individualized educational plans, pedagogical flexibility, and curriculum enrichment. Socio-emotional repercussions, such as frustration and low self-esteem, were reported. However, some schools demonstrated inclusive practices based on equity, articulating specific needs and potentialities. Although the results do not allow for generalizations, they highlighted the need to strengthen professional training and the articulation between school, family, and specialized services. It was concluded that the inclusion of students with Dual Exceptionality requires practices that simultaneously recognize their difficulties and potentialities, ensuring equitable conditions for participation, learning, and development.

Keywords: Human development; Teacher training; School inclusion; Neurodivergence; Pedagogical practices.

October 02, 2026

Classification of defaulting customers using supervised machine learning techniques

The risk of default in credit operations demanded analytical approaches to anticipate losses. This study comparatively evaluated the performance of supervised machine learning models in classifying defaulting customers in credit card operations. The public dataset “Default of Credit Card Clients” from the University of California Irvine was used, with 30,000 observations and class imbalance. The algorithms Logistic Regression, Random Forest, and Extreme Gradient Boosting were employed. The imbalance was addressed by assigning weights to the classes, and model optimization occurred with the RandomizedSearchCV method, prioritizing sensitivity. Cross-validation results indicated that the Extreme Gradient Boosting model showed a higher capacity for identifying the defaulting class and better discriminatory performance, followed by Random Forest and Logistic Regression, with a sensitivity of 0.8250 and an AUC-ROC of 0.7844 for XGBoost. Interpretability analysis, conducted by the Shapley Additive Explanations (SHAP) technique, highlighted the predominance of variables associated with payment behavior, especially the history of delays. It was concluded that tree-based models, particularly boosting techniques, proved to be more suitable for capturing complex patterns in the data, configuring themselves as consistent alternatives for credit risk management.

Keywords: Machine Learning; Credit Card; Classification; Extreme Gradient Boosting; Credit Risk.

October 02, 2026

Sentiment Analysis on Brazilian Banks on Twitter/X: Comparison between Traditional and Digital Institutions

A study analyzed public perception of Brazilian financial institutions on the Twitter/X platform, highlighting the importance of sentiment monitoring on social networks for understanding reputation and customer experience in the banking sector. The objective was to compare user perception of the image and reputation of traditional and digital banks, based on the sentiment patterns identified in the analyzed manifestations, seeking to identify structural differences between these groups. The methodology was based on the analysis of 1,096 tweets collected between November 2022 and June 2023. Two complementary sentiment analysis approaches were used, the sum and the average of labels, to capture the majority sentiment and nuances of perception. Additionally, the Market Profile Model, with indicators of emotional reputation, reputational risk, neutrality, and polarization, and the Banking Clustering Model, which allowed grouping institutions according to perception patterns, were developed. The results indicated a predominance of neutral and negative sentiments, a higher volume of interactions in digital banks, and structural differences in the emotional intensity of perceptions, with greater stability in digital banks and greater polarization in traditional ones. It was concluded that the combination of analytical and statistical techniques contributed to an in-depth understanding of institutional image in the digital environment, demonstrating the importance of data-driven reputation management strategies.

Keywords: Digital banks; Traditional banks; Data modeling; Opinion mining; Social Networks.

October 02, 2026

Optimization of annual budget planning through project management methodologies

The Annual Budget Planning (POA) is a crucial process for translating organizational strategy into operational and financial goals, but it frequently faces deadline pressures, interdepartmental dependencies, and the repetition of habitual expenses. The study aimed to analyze how the combined application of project management practices and Zero-Based Budgeting (OBZ) can optimize the POA. To this end, a case study was developed in the Brazilian operation of a publicly traded company in the beverage sector, using documentary research of its 2023 results report and an anonymous questionnaire applied to 47 respondents. Documentary analysis indicated growth in net revenue, expansion of gross profit and adjusted EBITDA, and contained advancement of selling, general, and administrative expenses, suggesting cost discipline and operational leverage. The complementary survey revealed a high perception of cascading effect on the schedule, strong support for defining cost package owners, and a preference for technical justification of expenses, in addition to demand for controlled flexibility after the baseline definition. It was concluded that structuring the POA as a project, associated with the rigor of OBZ, increased the process predictability, reinforced accountability for expenses, and broadened the coherence between budgetary execution and economic-financial performance.

Keywords: Cost Control; Operational Efficiency; Zero-Based Budgeting; PMBOK; Beverage Sector.

Digital Business

October 02, 2026

Influence of social media on consumer behavior

The study of consumer behavior sought to understand the factors that influence purchasing decisions in the context of increasing digitalization, where the internet is widely used by the Brazilian population. The objective was to analyze how social networks influence consumer purchasing behavior and identify the types of content that generate the most attention. Data collection occurred through an online questionnaire, distributed to the general public, which resulted in 159 valid responses. The data were processed and analyzed using descriptive statistics and variable cross-tabulation to identify trends and correlations. The results revealed that 89.9% of respondents had already made purchases after exposure to content on social networks. Platforms such as TikTok and Pinterest showed the highest conversion rates among their users. It was observed that organic reviews and recommendations from friends or family exerted the greatest influence on purchasing decisions. Furthermore, it was identified that the absence of prior financial planning and the high frequency of exposure to dynamic content on social networks acted as catalysts for recurring purchases. It was concluded that social networks have consolidated themselves as strategic conversion channels, and understanding these mechanisms is fundamental for brands to develop efficient digital marketing strategies, prioritizing transparency and social proof.

Keywords: Consumer behavior; Purchase decision; Content strategy; Digital marketing; Social networks.