Article

Technology

March 25, 2025

Data management in scientific research: ethical approaches and integrity practices

DOI: 10.22167/2675-6528-20240052
E&S 2025, 6: e202400052

Ana Paula Lopes Marinho e Alex Nunes de Almeida

With the increasing availability of digital data and the use of Artificial Intelligence (AI) and Machine Learning (ML) methods in data science, a new paradigm for information management and security emerges. Scientific production, based on the collection, manipulation, and sharing of data by researchers[1], makes Data Governance (DG) essential. The quality and transparency of information must be improved, ensuring adequate techniques for replication[2], especially in systems that use Large Language Models (LLM), such as ChatGPT®[3, 4]. However, the implementation of good DG practices faces significant challenges, especially in the need to ensure compliance with data protection laws and the promotion of transparent actions[2].

The application of AI and ML has automated several areas, but faces risks of statistical biases due to the lack of diversity in models, generating distrust in relation to its results[4]. This study aimed to improve good practices in GD and research integrity, mapping tools that assist researchers in managing scientific information. Thus, the initial proposal of a comprehensive model for a Data Management Plan (DMP) was presented, integrating global methods and procedures in a unified approach.

The implementation of good DG practices is challenging, as researchers face difficulties in complying with legislation that ensures data reproducibility and integrity, which are fundamental to meeting rigorous legal requirements and epistemic criteria[5,6]. Some platforms, paid or free, offer DPG systems that allow researchers to fill in information, especially in funded projects[7]. However, to ensure data quality and transparency, the use of these tools needs to be expanded[3]. To discuss the importance of DG in research, relating ethical values and integrity practices, the literature suggests analyzing how researchers acquire and value knowledge, perform tasks, and assign responsibilities[1].

The central question of the study is: how to effectively implement a PGD, considering ethical values and integrity practices? It seeks to advance discussions on good GD practices, highlighting how PGD can ensure the quality and transparency of information, and encouraging the use of tools for managing scientific data, aiming for the dissemination and reproducibility of knowledge.

The methodology was structured in two main stages to balance the breadth of scientific knowledge in order to communicate the essence of the research. In the first, a bibliometric review was carried out, divided into three phases (Table 1), to identify trends, patterns, and gaps in specific areas of knowledge[8] . This preliminary search was fundamental to establish a solid literature base, refining inclusion and exclusion criteria and identifying relevant keywords and sources, ensuring that the research was based on studies with significant impact in the field. However, the search in the Scopus® database revealed that there are few studies on Data Science focusing on integrity in PGD.

Tabela 1. Resultados consolidados da pesquisa bibliográfica

PhasesPreliminary studiesScopus StudiesStudies via SnowballFinal sample total
Bibliographic research results66517
Source: Elaborated by the authors

Faced with this difficulty, the “snowball” sampling strategy (snowball) was used to expand the reference base, connecting researcher networks and identifying additional research[9], facilitating access to interconnected data as per the project’s needs. For the snowball strategy, the Research Rabbit® platform was used as an exploratory tool, as described in the review protocol (Table 2). The choice meets ethical and regulatory requirements.

Tabela 2. Protocolo de pesquisa da revisão sistemática – “snowball sampling”

ActivitiesDescription
Keyword selectionThe research had two sets of keywords: Research Integrity; Data Management Plan
Access to articles and reading of AbstractsIdentification of research topic-relevant articles related to the search terms.
Exploration of References and CitationsBy choosing the references and citations from the initial articles, new sources were identified from the initial research, combining other new ones from representative data.
Search ExpansionBased on the new references and citations found, the search on the platform was expanded to include new additional sources.
Analysis and synthesisEvaluation of information and synthesis. The “snowballing” process was repeated to expand the research and increase the crashing of the topic.
Registration and management of referencesThroughout the process, the references found were saved on the Research Rabbit® platform. This facilitated the organization and tracking of the articles used in the TCC.
Source: Elaborated by the authors based on Naderifar and Ghaljaie[9].

Works by Cardoso et al.[7], Ministry of Health[10], Doorn[3], Engelhardt[11], and Digital Government[12] were included. To update the data, the second phase used the Scopus database (Alam et al.[6]; Bolland & Grey[13]; Bouter[14]; Carcausto-Calla et al.[15]; Valkenburg et al.[1]; Van den Hoven et al.[16]). Subsequently, the search was expanded with the snowball sampling[9] technique to obtain qualitative samples and identify gaps in research integrity (Beugelsdijk & Meyer[2]; McGovern et al.[4]; Hosseini et al.[5]; Wilkinson et al.[17]; Van den Eynden et al.[18]).

In the second stage, scientific GP repositories and platforms, national and international, regulated and recommended by editors committed to the integrity of this research, were mapped. A search was conducted on Google Scholar® and consultations were made with university journals to verify submission policies, data integrity, and transparency on their websites. Business and academic case studies were analyzed that highlight errors in KM, especially when AI algorithms, based on mathematical formulas, are mistakenly considered infallible. These cases, conducted by specialists, underscore the need for the use of KM tools to ensure the replicability of studies.

Restricted access to open resources is a significant challenge that hinders the dissemination and validation of results[5,2]. The lack of open accessibility impairs replicability, creating barriers to research validation[4], as, even with automated checks, failures persist, allowing irresponsible practices such as falsification or plagiarism of information[1]. Table 3 highlights the limitations and potentialities faced in accessing and replicating data in computational science.

Education on integrity faces a gap in the understanding of ethics and reproducibility, highlighting the need for specific training[2, 6, 14]. Studies by Carcausto-Calla et al.[15] and Van den Hoven et al.[16] point to positive results with structured training that uses active methodologies, such as workshops, case studies, simulation games, and even Freirean pedagogy, promoting critical autonomy and research integrity.

Tabela 3. Limitações e potencialidades de acesso a recursos de acesso aberto

LimitationsAuthor(s)
Access to free, high-performance infrastructures for conducting research and creating PhDsMcGovern et al.[4]; Beugelsdijk et al.[2]
Development of simple, transparent AM applications and with plagiarism prevention verification infrastructureValkenburg et al.[1]
The use of LLMs presents dilemmas regarding the application of transparency. Accountability is essential to ensure integrity, legal rigor for data replicability.Hosseini et al.[5]; Beugelsdijk et al.[2]
False and plagiarized scientific knowledge entering databases.Valkenburg et al.[1]
The culture of shared knowledge (Creative Commons) is growing, but we still face challenges with public domain licensing and costs to release research access.Beugelsdijk et al.[2]
The main limitation to the public understanding of the scientific process is the lack of training, adequate infrastructure, consistent reporting, clear procedures, and attention to publication ethics.Alam et al.[6]; Bouter[14]
Researchers need help from their research institutions to optimize the functioning of their moral compass.Bouter[14]
There are few opportunities for interested communities to interact in order to collaboratively seek challenging solutions to research integrity problems.  Alam et al.[6]
PotentialitiesAuthor(s)
The use of LLMs as a tool to facilitate text writing or editing is a practice that promotes inclusion and diversity.Hosseini et al.[5]
Most scientific journals have adopted good research practices, incorporating review processes and codes of ethics that promote data access and research transparency, known in English as Data Access and Research Transparency (DART).Beugelsdijk et al.[2]
Ethics Committees are responsible for creating, disseminating, and approving research protocols. A practice widespread in many countries, including Brazil.Beugelsdijk et al.[2]; Ministry of Health[10]
The Data Protection Law proposes and encourages researchers to develop a DPG with the details of the types of data collected, how they were analyzed and stored.Beugelsdijk et al.[2]
Training programs developed using active methodologies for integrity and research development for the scientific community.Carcausto-Calla et al.[15]
Evidence shows that the trainings strengthened normative concepts, promoted critical autonomy, openness, self-awareness, and courage to deal with integrity issues in practice.Van den Hoven et al.[16]
Source: Prepared by the authors

In academic publications, there is a growing adoption of transparency and data access policies, with many journals integrating the Committee on Publication Ethics (COPE). However, as Beugelsdijk et al.[2] and Bolland et al.[13] point out, even though COPE does not define specific methodological standards, this change aims for a more ethical research culture, but the use of “Creative Commons” resources may face financial challenges. The General Data Protection Law in Brazil and the corresponding legislation in the European Union, enacted in 2018, reinforce the need for a PGD, which is evaluated by the Research Ethics Committees (CEP). The PGD must respect the specific policies of the journal, CEP, and data protection laws.

In Table 4, it is observed that research integrity is seen as an individual responsibility of the researchers[1] which requires care in all stages, emphasizing transparency and security in information management, regardless of the technologies employed[2,6,16], as the commitment to integrity demands responsible conduct from researchers.

Tabela 4. Síntese teórica sobre a integridade na pesquisa

AuthorsConcept of integrity in research
Valkenburg et al.[1]Research integrity is viewed by the authors as a matter of individual researchers’ responsibility regarding their respective research activities.
Beugelsdijk et al.[2]Transparency is a basic condition in scientific research that stems from the data production stage to the analysis results, from data production to analysis results, accumulating credibility in the development of theory about the phenomenon.
Alam et al.[6]According to the authors, integrity in research consists of adding social value and benefits to the research. The application of integrity in research is marked by the consistent adherence to honesty, responsibility, professional courtesy, fairness, and good management.
Van den Hoven et al.[16]Understands that research integrity is the effect of stimulating empowerment towards responsible research conduct.
Source: Elaborated by the authors

Most scientific journals require authors to submit separate documents, such as tables and figures, detailing the database. They also recommend making the data available in legal repositories, accessible via “Digital Object Identifier” (DOI), allowing citations as formal bibliography[2]. The DOI identification of an article connects it to the researcher or research organization. Some platforms cross-reference data by requiring login with the “Open Researcher and Contributor ID” (ORCID) identifier, a non-profit initiative that provides a unique and permanent identifier for researchers and academic collaborators[19]. Used to validate the inclusion of information associated with the DOI, this approach promotes transparency in research, although not all systems offer this integration[6].

Searched platforms and repositories

Repositories providing DOIs were consulted, with emphasis on those recommended by reviewers and that offer a standardized and globally accessible metadata interface for data export[8,20]. For each repository listed, a brief explanation of its main function in the context of data replicability is presented (Table 5).

Tabela 5. Repositórios indicados por revisores que fornecem DOI

Repository nameAccess dataFunction
DataCite Commonshttps://repositoryfinder.datacite.orgAllows the researcher to conduct research, access and reuse datasets
Dryadhttps://datadryad.org/stashData repository that offers storage and sharing of scientific data  
Harvard Dataversehttps://dataverse.org/Allows researchers to share, preserve, and cite their data
Mendeleywww.elsevier. com/solutions/mendeleyReference management and academic social network. It presents document organization and research network functions
ROR – Research Organization Registryhttps://ror.org/Global registry of unique identifiers for research organizations
Source: Prepared by the authors

Several tools for the creation of RMPs, or Data Management Plans (DMP) in English, were investigated. Below, we list the tools identified in this research, accompanied by a brief explanation (Table 6).

Tabela 6. Ferramentas para elaboração de Plano de Gestão de Dados (DMPtools)

NameAccess dataFunction
DCC – Digital Curation Centrehttps://www.dcc.ac.uk/Digital curation center offering resources and guidance for research KM
DMPtoolhttps://dmptool.org/authOnline tool that assists researchers in the development of GD plans. Tool used by several universities
DMPTuuli https://www.dmptuuli.fi/Similar to DMPtool, but specific for Finnish researchers
Mendeleyhttps://data.mendeley.com/Although best known as a reference management service. It offers features for research KM
USGS – Science for a Changing Worldhttps://www.usgs.gov/data-managementUnited States government agency dedicated to the management and analysis of geospatial and environmental data
Source: Elaborated by the authors

Data Management Plan Model

The PGD is an essential document that organizes and details research development. It should include the description of data collection protocols, risk and benefit assessment, and the necessary procedures to ensure the confidentiality and privacy of information, in addition to retention practices[2]. Effective data management ensures its curation, benefiting the scientific community and complying with the governance guidelines established by regulatory bodies[12].

A well-crafted DMP ensures proper data archiving, defining guidelines for its preservation and facilitating its use by future researchers. According to Wilkinson et al.[17], a DMP converts metadata into datasets capable of receiving a DOI identifier. To this end, strategic decisions, such as the use of computational tools, are essential.

Data governance should follow the FAIR principles — Findable, Accessible, Interoperable, and Reusable — to strengthen data transparency and integrity, as well as promote academic collaboration[11, 8]. Complementarily, the CARE principles (Collective, Authority, Responsibility, and Ethics), introduced by Souza et al.[21], focus on the ethical management of data, especially for vulnerable populations. Created at the International Data Conference, the CARE principle was originally aimed at the treatment of indigenous data, but it extends to all populations on which scientific studies are based.

Following Van den Eynden[21] and Beugelsdijk et al.[2], the implementation of good DM practices requires a meticulous and organized approach. The RDM model must align with FAIR and CARE principles (Figure 1) and be developed following these steps:

1) Description of data, metadata standards and interoperability: the process requires detailed information about the data used in the research, ensuring interoperability. The data must be available on a compatible platform that allows its interaction.
2) Legal, ethical restrictions, affiliations, and authors: researchers must comply with laws and regulations on data use and sharing, such as data protection, copyright, and intellectual property, also considering limitations from ethics committees and ethical issues related to participant privacy and consent.
3) Data preservation, sharing, and reuse policy: involves the preservation, sharing, and reuse of data, based on storage practices that ensure long-term access. This includes digital systems, regular backups, adequate documentation, and preservation standards to prevent losses and inadequate access.
4) Formats and standards: a well-structured PGD must indicate file formats and data structure, essential information to ensure interoperability and portability between systems and platforms[3, 2,10].
5) Repositories: specific locations store and organize data systematically and securely, allowing authors to structure them in a way that facilitates access, sharing, and permanent preservation.

Figura 1. Etapas para desenvolvimento do Plano de Gestão de Dados
Fonte: Elaborado pelos autores com base em Van den Eynden[18] e Beugelsdijk et al.[18,2].

The PGD model proposed in this study covers five essential stages for its development, aligned with the three phases of the development cycle of a Final Course Project (TCC). Although aimed at Data Science Analytics students, it can be applied in other disciplines that require a PGD. Considering the complexity of the databases used, a robust computational system that ensures agility, efficiency, and security is indispensable. The use of integrated platforms or repositories for data collection, creation, and sharing is also fundamental. The model’s guidelines, illustrated in Figure 2, aim to guide researchers in adopting careful and responsible practices, promoting the credibility of studies and contributing to the advancement of knowledge in their fields.

Research Cycle

1) Research project: involves the initial activities for the project conception, ensuring that data integrity and reliability are defined from the outset.
2) Preliminary results: represent a crucial milestone, in which it is fundamental to ensure the application of the methodology and materials established in the project.
3) Final text: although it is the final stage, this activity is crucial for reviewing patterns and formats, estimating the volume of data generated, and specifying the tools for its processing.

Figure 2. Research cycle
Source: Prepared by the authors.

In each of the three stages, essential activities for comprehensive data management were incorporated. This integration into the research cycle ensured the application of CARE and FAIR principles, promoting a model that aligns ethical and responsible actions. Thus, the researcher can conduct their activities transparently, describing the database in detail and clearly communicating the reliability of the study’s results. This study presented a model to facilitate the implementation of PGD in GD research and other academic work, emphasizing the importance of integrity in the three phases of the research cycle: conception, development, and final writing. By incorporating ethical practices from the beginning to the conclusion of the project, the responsible conduct of researchers and the need for rigorous data management are reinforced. This approach not only elevates the quality of academic work but also promotes dialogue on transparency in data dissemination and reuse.

References

[1] Valkenburg, G.; Dix, G.; Tijdink, J.; Rijcke, S. 2021. Expanding research integrity: A cultural-practice perspective. Science and engineering ethics 27(10):1-23. https://doi.org/10.1007/s11948-021-00291-z.

[2] Beugelsdijk, S.; Van Witteloostuijn, A.; Meyer, K.E. 2020. A new approach to data access and research transparency (DART). Journal of International Business Studies 51:887-905. https://doi.org/10.1057/s41267-020-00323-z.

[3] Doorn, P.K. 2018. Science Europe Practical Guide to the International Alignment of Research Data Management. https://doi.org/10.5281/zenodo.4915862.

[4] McGovern, A.; Ebert-Uphoff, I.; Gagne, D.J.; Bostrom, A. 2022. Why we need to focus on developing ethical, responsible, and trustworthy artificial intelligence approaches for environmental science. Environmental Data Science, 1, p.e6. DOI: 10.1017/eds.2022.5.

[5] Hosseini, M.; Resnik, D.B.; Holmes, K. 2023. The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscripts. Research Ethics 19(4):449-465. https://doi.org/10.1177/17470161231180449.

[6] Alam, S.; Burley, R.; Graf, C.; Meadows, A.;  Mejias, G.; Pattinson, D. 2024. (Re?) Building trust in research integrity. Information Services & Use 44(1):1-4. DOI: 10.3233/ISU-230200.

[7] Cardoso, J.; Pereira, F.; Moreira, J.M. 2024. Serviço de Repositório Digital de Dados de Investigação. Revista Científica Da UEM: Série Letras E Ciências Sociais, 4(1). Disponível em: http://196.3.97.23/revista/index.php/lcs/article/view/223. Acesso em: 21 jan. 2025.

[8] Gottardi, T.; Medeiros, C.B.; Reis, J.C., 2021. Semantic Search to Foster Scientific Findability: A Systematic Literature Review. Journal of Information and Data Management 12(5). https://doi.org/10.5753/jidm.2021.1919.

[9] Naderifar, M.; Goli, H.; Ghaljaie, F. 2017. Snowball sampling: A purposeful method of sampling in qualitative research. Strides in development of medical education, 14(3). https://doi.org/10.5812/sdme.67670.

[10] Ministério da Saúde [MS]. 2023. Conselho Nacional de Saúde. Comitê de Ética em Pesquisa. Disponível em: < https://conselho.saude.gov.br/comites-de-etica-em-pesquisa-conep?view=default>. Acesso em: 11 de abril de 2024.

[11] Engelhardt, C. 2022. D7. 4 How to be FAIR with your data. Disponível em: https://library.oapen.org/handle/20.500.12657/54460. Acesso em: 21 jan. 25.

[12] Governo Digital. 2024. Governança de Dados. Disponível em: https://www.gov.br/governodigital/pt-br/infraestrutura-nacional-de-dados/governancadedados/governanca-de-dados. Acesso em: 11 de abril de 2024.

[13] Bolland, M.J.; Avenell, A.; Grey, A. 2024. Publication integrity: what is it, why does it matter, how it is safeguarded and how could we do better?. Journal of the Royal Society of New Zealand, 55(1541):1-20. DOI: 10.1080/03036758.2024.2325004.

[14] Bouter, L. 2020. What research institutions can do to foster research integrity. Science and engineering ethics, 26(4):2363-2369. DOI: 10.1007/s11948-020-00178-5.

[15] Carcausto-Calla, W.; Zapata, N.A.; Cueva, F.E.I.; Morales, S.A.G.; Rios, A.R. 2023. Mapping of Empirical Studies on Research Integrity in University Institutions 12(5):9-22. https://doi.org/10.36941/ajis-2023-0122.

[16] Van den Hoven, M.; Lindemann, T.; Zollitsch, L.; Prieß-Buchheit, J. 2023. A Taxonomy for Research Integrity Training: Design, Conduct, and Improvements in Research Integrity Courses. Science and engineering ethics, 29(14):1-21.  https://doi.org/10.1007/s11948-022-00425-x.

[17] Wilkinson, M.D.; Dumontier, M.; Aalbersberg, I.J.; Appleton, G.; Axton, M.; Baak, A.; Mons, B. 2016. The FAIR Guiding Principles for scientific data management and stewardship. Scientific data, 3(1):1-9. https://doi.org/10.1038/sdata.2016.18.  

[18] Van den Eynden, V.; Corti, L.; Woollard, M.; Bishop, L.; Horton, L. 2011. Managing and sharing data; a best practice guide for researchers. Disponível em: https://dam.ukdataservice.ac.uk/media/622417/managingsharing.pdf. Acesso em: 21 jan. 2025.

[19] ORCID. (2024). ORCID – Connecting Research and Researchers. Disponível em: https://orcid.org/. Acesso em: 11 de abril de 2024.

[20] Laender, A.H.F.; Medeiros, C.M.B.; Cendes, I.L.; Barreto, M.L.; Van Sluys, M.A.; Almeida, U.B.D. 2020. Abertura e gestão de dados: desafios para a ciência brasileira. Disponível em: https://www.abc.org.br/wp-content/uploads/2020/09/ABC-Abertura-e-Gest%C3%A3o-de-Dados-desafios-para-a-ci%C3%AAncia-brasileira.pdf. Acesso em: 21 jan. 2025.

[21] Souza, L.P.; Sousa, R.S.C.; Löw, M.M.; Barros, T.H.B. 2023, August. Promovendo a justiça epistêmica: uma análise dos princípios CARE na gestão de dados de pesquisa em relação aos povos indígenas. In Anais do Workshop de Informação, Dados e Tecnologia-WIDaT 6. https://doi.org/10.22477/vi.widat.73.

How to cite

Marinho A.P.L.; Almeida A.N. Data management in scientific research: ethical approaches and integrity practices. Revista E&S. 2025; 6: e20240052.

About the authors

Ana Paula Lopes Marinho – Master in Administration. State Center for Technological Education Paula Souza. Praça Coronel Fernando Prestes, 30, Bom Retiro, 1124-060, São Paulo, São Paulo, Brazil.
Alex Nunes de Almeida – Doctor in Sciences. Pecege. Rua Cezira Giovanoni Moretti, 580, Santa Rosa, 13414-157, Piracicaba, São Paulo, Brazil

You may also like