October 02, 2026
Data Transformation into Strategy: Applied Research for Ecotourism Operation Optimization
Transforming Data into Strategy: Applied Research for Ecotourism Operation Optimization
Fernanda Karla Pinto Morais; Edilson José Rodrigues
DOI: 10.22167/2675-6528-202602896
Article derived from a Final Course Work (TCC), with content based on the student’s original work and adapted to the editorial format of the E&S Magazine with the support of the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege for textual synthesis and organization.
Summary
The growing demand in ecotourism in Minas Gerais has driven the search for business intelligence to transform customer data into strategic information. The study aimed to structure a data science pipeline to collect, segment, and classify the customer base of an ecotourism operation, in order to optimize marketing actions and anticipate market movements. An exploratory, quali-quantitative research was conducted through a case study. 2,777 transactional records from an ecotourism company, referring to January 2024 to December 2025, were used. The methodological process involved automated data collection (Google Sheets API), processing and enrichment (ETL), validation, and creation of RFM (Recency, Frequency, and Monetary Value) attributes. Dimensionality reduction via PCA and K-Means clustering was applied, with the number of clusters defined by the Elbow method and Silhouette Score. Results were validated with DBSCAN and K-Medoids. The results revealed the identification of three customer behavioral segments: “Loyal”, “Low Value”, and “Potential”. The “Loyal” segment represented the highest accumulated economic value, while the “Potential” stood out for its high average ticket and potential for conversion into recurrence. The integration of data analysis techniques proved to be a robust and replicable method for generating intelligence in ecotourism. It was concluded that the structured data science pipeline enabled the behavioral segmentation of the customer base, statistical validation of the groups, and the creation of a predictive system for new buyers, providing subsidies for data-driven strategic decisions and future analyses.
Keywords: Clustering; Business intelligence; Machine Learning; Customer segmentation; Decision making.
1. Introduction
In recent years, ecotourism has consolidated itself as a segment that transcends simple leisure, aligning direct contact with nature with environmental preservation and the valorization of local culture. This sector, supported by the guidelines of the Ministry of Tourism, not only generates socioeconomic benefits but also influences the entire national tourism chain. Minas Gerais, in particular, stands out as one of the main ecotourism hubs in the country, housing Brazil’s only mountain range, the Espinhaço, recognized by Unesco as a Biosphere Reserve, which makes it a strategic setting for the development of this activity (Miranda, 2025).
Following the expressive growth of ecotourism, tourism agencies assume a crucial consulting role for consumers, organizing trips that meet their audience’s needs (Silva, 2016). However, intense competition demands a deep understanding of the market and customers. In this context, data cease to be fragmented information with low analytical value (Olsen, 2015) and become a strategic asset. They require systematic processes of treatment, modeling, and analysis, capable of generating sustainable competitive advantages and transforming informational intelligence into a decisive factor for organizational performance (Fávero e Belfiore, 2024).
The present study focuses on an ecotourism company in Minas Gerais that offers various outdoor experiences, from day trips to expeditions and immersions. Currently, the managers of this company face the challenge of organizing and optimizing the vast volume of data generated by the business. The transition from empiricism-based management to a data-driven approach is essential to deepen the understanding of the customer profile and to comprehend factors such as seasonality and market movements that can influence strategic decisions.
Knowing the customer is fundamental for the success of marketing actions, from defining the target market to choosing the most appropriate approaches for each segment (Solomon, 2016). In ecotourism, this need is even more pressing, as planning a trip is, for the customer, a memorable moment, often surpassing the memories of the experience itself (Silva, 2016). Furthermore, the current scenario requires companies to use their digital social networks as strategic channels to leverage sales and strengthen relationships with the public (Destro, 2023).
Given this scenario, the research is justified by the need to transform raw data into actionable intelligence to optimize the ecotourism operation, and the objective of this work is to structure the data pipeline of the company in question, enabling the identification of customer groups through their transactional records and, through the application of algorithms, optimize strategic marketing actions, increasing their chances of success, in addition to seeking ways to anticipate movements related to the market, seasonality, and issues related to customer behavior.
2. Material and Methods
The exploratory, quali-quantitative, and applied research was conducted as a case study with descriptive elements (Gil, 2002). The study analyzed transactional records from an ecotourism company in Minas Gerais, concerning negotiations from January 2024 to December 2025. The methodology structured a data science pipeline to identify customer groups, optimizing marketing actions and anticipating market movements and behavior.
The primary dataset (Google Sheets) contained 2,777 transactional sales records. Variables included anonymized transaction date, anonymized email and full name, gender, anonymized CPF, date of birth, phone number, pickup location, Instagram (optional), and the purchased event. Processing followed LGPD (Law nº 13.709/2018), with anonymization of direct identifiers. The use of information was authorized by company management for academic purposes, preserving confidentiality and integrity.
The Extract, Transform, and Load (ETL) process was implemented in Python (Spyder/Anaconda). Libraries such as Gspread, Google-auth, and Sys were used for centralization, authentication, and execution control. Pandas (Caetano, 2025) processed and enriched the data, removing non-numerics, validating CPFs, creating status columns, and recovering incomplete records (backfill), according to methodological guidelines (Barbieri, 2020).
In feature engineering, columns were converted to appropriate data types, and the RFM (Recency, Frequency, and Monetary Value) technique was applied. RFM ranked the customer base into quantitative categories, standardizing purchasing behavior. RFM was aggregated to an operational measure of Lifetime Value (LTV), defined as the accumulated monetary value of customer transactions.
For data preparation and transformation, Scikit-learn was used for Z-score standardization (StandardScaler) and for Principal Component Analysis (PCA), configured to preserve 95% of the explained variance (Netto and Maciel, 2021). Standardization ensured that variables contributed equally to the distance calculation (Barr et al., 2024), and PCA eliminated redundancies.
The definition of the ideal number of clusters (K) was performed with the K-Means algorithm (Scikit-Learn), the Elbow method, and the Silhouette Score index (Sharda et al., 2025). Matplotlib and Seaborn assisted in visualization (Netto and Maciel, 2021; Kyran, 2024). The integrated analysis pointed to segmentation into three groups (K=3), representing the best balance between cohesion and separation of clusters (Gomes, 2024; Pierson, 2019).
To validate the results, the DBSCAN and K-Medoids algorithms were applied. DBSCAN was used as a support method to validate the structure of the groups found by K-Means (Faceli et al., 2025). K-Medoids defined real data points (medoids) as cluster centers, providing greater robustness to outliers and noise (Goldschmidt et al., 2015), and its results converged with those of K-Means, validating the structure of three clusters.
The definition of the commercial names of the clusters occurred by dynamic mapping of the statistical profiles of the groupings. At the end of each K-Means round, the average of the Total_Spent variable was calculated for each cluster and, in ascending order, the labels ‘Low Value’, ‘Potential’, and ‘Loyal’ were assigned. This process aligned the business meaning of the segments with the quantitative profiles, according to the area of application (Faceli et al., 2025).
3. Results and Discussion
The analysis of the ecotourism company’s transactional data began with the exploration of relationships between variables, revealing important customer behavioral patterns. A correlation of 0.98 was observed between Total_Spending and Num_Trips, indicating that customers with higher spending volume tend to take more trips. This strong linear dependence, exceeding the 0.90 threshold established by Fávero and Belfiore (2024), confirmed the expected severe multicollinearity, as total spending is intrinsically linked to travel frequency and average ticket price. Additionally, the Recency_Days variable showed a moderate negative correlation of -0.27 with spending and travel quantity variables, suggesting that customers who traveled more recently tend to spend more and have a higher travel frequency.
To mitigate the effects of multicollinearity and optimize the application of clustering algorithms, the Principal Component Analysis (PCA) technique was employed as a preprocessing step. The PCA configuration was set to preserve 95% of the original explained variance of the data, resulting in the generation of four principal components (PC1 to PC4). This transformation ensured that the new components were uncorrelated with each other, guaranteeing orthogonality and mathematical independence of the information. This procedure eliminated redundancies and made the data more suitable for distance-based algorithms, contributing to the stability and interpretability of subsequent results, as highlighted by Netto and Maciel (2021).
The definition of the ideal number of clusters (K) was performed by applying the Elbow Method and the Silhouette Score jointly. The Elbow Method indicated an inflection point in the inertia curve between K=3 and K=4, suggesting diminishing returns in adding new groups from K=5 onwards. The Silhouette Score, in turn, a metric that quantifies the internal cohesion and separation between clusters (Rousseeuw, 1987), pointed to K=2 with the highest coefficient (0.467), followed by K=3 (0.454). Despite the small difference, the choice of K=3 was justified by its greater practical utility for personalizing marketing strategies, offering a superior balance between group cohesion and distinction between them.
The final segmentation, using the K-Means algorithm with K=3, demonstrated a clear spatial separation of customer groups, each with well-defined characteristics. Cluster 0 concentrated the largest part of the customer population, corresponding to a segment with a higher population volume. Cluster 1 encompassed high-engagement and accumulated spending customers, while Cluster 2 brought together customers with a distinct behavioral pattern, characterized by high recency and average ticket size. Dimensionality reduction via PCA, which maintained 95% of the explained variance in the first two components, ensured that the visual representation of the clusters preserved the essential structure for data interpretation.
To validate the robustness of the segmentation obtained with K-Means, the DBSCAN and K-Medoids algorithms were applied. DBSCAN, parameterized with eps=1.5 and min_samples=3, identified three clusters and classified some observations as noise. However, the resulting segmentation was asymmetric, with Cluster 0 concentrating the majority of observations and Clusters 1 and 2 being smaller, in addition to dispersed noise. Although the number of clusters coincided with K=3, the difference in composition and distribution limited the interpretation of a direct qualitative convergence, which is consistent with the limitations of DBSCAN in handling heterogeneous data distributions, according to Faceli et al. (2025).
The K-Medoids algorithm, applied with K=3 on the data transformed by PCA, also confirmed the three-cluster structure. Unlike K-Means, which uses arithmetic means as centroids, K-Medoids selects actual data points (medoids) as cluster centers, which provides greater robustness to the presence of outliers and noise (Goldschmidt et al., 2015). The distribution of groups in K-Medoids, although with a slightly different composition than K-Means (Cluster 0: n=341; Cluster 1: n=99; Cluster 2: n=224, compared to Cluster 0: n=67; Cluster 1: n=555; Cluster 2: n=42 in K-Means), maintained a similar spatial separation, reinforcing the stability of the choice of K=3.
The statistical characterization of the clusters, using boxplots of the original variables on an unstandardized scale, allowed the translation of quantitative profiles into business intelligence. The “Loyal” segment (Cluster 0) stood out as having the highest accumulated economic value, with an average age of 41 years, five acquired trips, a historical LTV of R$994.00, and an average recency of 290 days. Its average ticket was R$203.00. This profile, which demonstrates an active and frequent relationship with the company, is strategic for loyalty actions and frequency reward programs, aiming to maximize its profitability.
The “Low Value” segment (Cluster 1) represented the largest part of the customer base (83.6%), concentrating buyers of occasional experience, predominantly single purchases. This group presented an average age of 37 years, the lowest average ticket in the base (R$181.00), reduced LTV, and high recency, around 420 days. The wide interquartile dispersion in recency suggests low engagement and a higher propensity for abandonment (churn). Converting a portion of this audience into recurring customers represents a significant gain in the operation’s historical LTV, making it a priority for reactivation strategies.
Finally, the “Potential” segment (Cluster 2) stood out for the highest average ticket among the groups, reaching R$294.00, a value approximately 45% higher than the “Loyal” segment. This group presented a younger average age (33 years), low purchase frequency (average of 1.14 trips), and high average recency, around 480 days. The expressive internal heterogeneity of the average ticket, with outliers above R$300.00, indicates that part of the group has a considerably higher spending power than the median. This profile is susceptible to upsell campaigns, post-sale relationship actions, and communications targeted at higher-value experiences, with great long-term conversion potential.
In summary, the structuring of the data science pipeline and the application of clustering algorithms allowed for the identification of three behavioral customer segments — Loyal, Low Value, and Potential — in the ecotourism operation. This segmentation, validated by multiple techniques, offers an in-depth understanding of customer profiles and purchasing behavior, providing concrete subsidies for optimizing strategic marketing actions and anticipating market movements, aligning directly with the study’s central objective of transforming raw data into actionable intelligence.
4. Conclusion
The present study aimed to structure a data science pipeline to collect, segment, and classify the customer base of an ecotourism operation in Minas Gerais, with the goal of optimizing marketing actions and anticipating market movements. It was found that the integration of data analysis techniques, from automated collection (Google Sheets API) to clustering, constituted a robust and replicable method. The application of Principal Component Analysis (PCA) mitigated multicollinearity and optimized the data for the algorithms. Through K-Means, with K=3 defined by the Elbow and Silhouette Score methods, three behavioral customer segments were identified: “Loyal”, “Low Value”, and “Potential”. The “Loyal” segment represented the highest accumulated economic value, while the “Potential” segment stood out for its high average ticket and potential for conversion into recurrence. The segmentation was statistically validated by DBSCAN and K-Medoids, reinforcing the stability of the groups and the coherence of their profiles. This approach enabled the transformation of raw data into actionable intelligence, providing concrete subsidies for data-driven strategic decisions and the creation of a predictive system for new buyers.
Despite the contributions, the study presented limitations, such as the dependence on the quality of data collected via digital forms, which may contain undetected typing errors, and the volume of the validated unique customer base, which restricted the ability to generalize to minority segments. The analysis was also limited by the available transactional variables, preventing a denser understanding of customer behavior through demographic, health, or marketing data. Furthermore, the single case study nature restricts the generalization of identified patterns to other ecotourism operations and long-term observations. For future work, it is suggested to expand the database with new periods to increase representativeness and identify new clustering opportunities. The implementation of a real-time relationship scoring system, data mining through Natural Language Processing (NLP) on customer feedback, the incorporation of the Analytic Hierarchy Process (AHP) or AHP-Gaussian Method into the RFM model for differentiated weighting of dimensions, and the application of algorithms such as Random Forest and LightGBM to create a predictive classification system for new customers, map variable relevance, and estimate the future value of each customer are recommended.
Bibliographic References
FÁVERO, L. P.; BELFIORE, P. 2024. Manual de Análise de Dados: Estatística e Machine Learning com Excel®, SPSS®, Stata®, R® e Python®. 2. ed., LTC. Rio de Janeiro, Brasil.
MIRANDA, T. Espinhaço: A Cordilheira do Brasil em Minas Gerais. Governo do Estado de Minas Gerais, 26 jun. 2025. Disponível em: https://minasgerais.com.br/pt/blog/artigo/espinhaco-a-cordilheira-do-brasil-em-minas-gerais. Acesso em: 25 jan. 2026.
OLSEN, W.. 2015. Coleta de Dados. Penso. Porto Alegre, Brsil. Disponível em: https://app.minhabiblioteca.com.br/reader/books/9788584290543/. Acesso em: 06 fev. 2026.
SILVA, G.G. 2016. Um comparativo entre o planejamento de uma viagem pelo guia PMBOOK e método convencional. Trabalho para obtenção de título de especialista em Gestão de Projetos. Escola Superior de Agricultura “Luiz de Queiroz” (ESALQ), Campus Piracicaba, Universidade de São Paulo (USP). Paulínia, São Paulo, Brasil.
SOLOMON, M.R..
Article originating from the Final Course Work of the Specialization in Data Science and Analytics of the MBA USP/Esalq
To learn more about the course, click here and access the MBX Academy platform