Research data: types, management, FAIR and data plans

Last update: February 27
  • What is meant by research data and what materials does the scientific community include and exclude?
  • How research data is classified according to format, origin, nature, and level of processing.
  • What does research data management consist of, and what are the FAIR principles that guide its opening and reuse?
  • What should a data management plan aligned with European open science requirements contain?

research data

In the day-to-day work of any scientific project , research data is the most delicate and, at the same time, the most valuable material generated. It's not just numbers in a spreadsheet or recordings stored on a hard drive: it's the foundation that allows the academic community to verify, reproduce, and validate research results. Without it, articles and conclusions largely remain mere assertions.

As open science gains traction, effective management of research data has become essential : organizing, documenting, preserving, and, where possible, sharing it with other researchers. This shift not only responds to demands from funders like the European Commission, but also to a scientific culture that champions open access, transparency, and reuse to accelerate new discoveries.

What exactly is research data?

When we talk about research data, we are referring to all the factual material recorded during the research process that the scientific community recognizes as necessary to validate the results obtained . In other words, these are the materials that support a scientific argument, a theory, or an experimental test.

These data can be facts, observations, measurements, or experiences generated or collected during a study, and they only acquire their true meaning within the context of the project itself . What is data for a particle physics research group may be irrelevant for an art history team, and vice versa; therefore, the discipline greatly influences how data are conceived and used, depending on the methodological approach.

A definition adopted by international organizations such as the OECD states that research data is any material recorded during research , accepted by the relevant scientific community, and used to certify the results achieved. This concept emphasizes two key aspects: recording (anything not documented is invalid) and recognition by the scientific field to which the project belongs.

It is important to emphasize that data alone is not yet information. According to several specialized authors, data becomes information when it is combined and processed using a method that allows for the discovery of patterns and relationships within the phenomenon being studied . In other words, the dataset is the raw material; analysis is what gives it meaning.

The collection of data generated and gathered during the execution of a project is usually called a dataset . This dataset can be very heterogeneous, mixing formats, sources, and levels of processing, but it is understood as a coherent unit linked to the research project that produced it.

What is included and what is not included as research data

The concept of research data is broad and encompasses a wide variety of formats, media, and content . Generally speaking, research data includes, among other things:

  • Laboratory notebooks and field notebooks, where observations, protocols followed, incidents and results of experiments or campaigns are recorded.
  • Primary research data, on paper or in digital format, which collect measurements, surveys, anonymized clinical records, etc.
  • Questionnaires and forms, along with the associated responses.
  • Audio tapes, voice recordings, interviews and other sound recordings collected for the study.
  • Videos and movies that document experiments, field observations, clinical sessions, or any other relevant process.
  • Models developed during the research, both conceptual and computational or statistical.
  • Photographs, scientific images (microscopies, x-rays, satellite images, etc.) and slides.
  • Digital objects and specific files generated by specialized instruments or software.
  • Algorithms, scripts, and software code employees to generate, process, analyze or visualize data.
  • Databases where the collected records are stored and structured.
  • Metadata and metadata schemas that describe the data, its origin and its internal structure.
  • Software configurationssimulation parameters and, in general, all the information necessary to reproduce the results.
  • Test checks and responses, including test results, clinical trials, or psychometric tests.

However, not everything generated during a project is considered final research data. Elements such as personal laboratory notes , draft articles, preliminary unrefined analyses, future plans, informal communications with colleagues, or physical objects (biological samples, specimens, archaeological vessels, test animals, etc.) are typically excluded; these are treated as materials, but not as reusable data.

Also typically excluded are incomplete or partial datasets that are not directly used to support published results, as well as very provisional working notes that do not reach a sufficient level of refinement for the community to recognize them as part of the research data corpus.

Types of research data according to different criteria

Research data can be classified in many ways. One of the most common distinguishes between quantitative and qualitative data . Quantitative data is usually numerical, measurable, and amenable to statistical analysis; qualitative data is based on discourses, images, texts, or behaviors, and is interpreted using methods such as content analysis, discourse analysis, participant observation, and so on.

Another perspective differentiates data based on its form of representation: it can be numerical, descriptive, or visual . Numerical data is stored as series of numbers (for example, physical measurements or results of closed surveys); descriptive data is expressed in natural language or other symbolic systems; visual data includes photos, videos, maps, diagrams, or graphs produced during the study.

It may interest you:  What educational exhibits are there and what can be learned?

Depending on their nature, data can be classified as qualitative or quantitative , a fundamental distinction in many disciplines. This distinction is often accompanied by different methodologies and archiving methods: a set of interviews is not preserved in the same way as a database with thousands of numerical observations.

If we look at the level of processing , it is usually separated into:

  • Raw or primary dataThese are the original records, with minimal processing, exactly as they were obtained during the collection phase. They constitute the raw material for the research.
  • Secondary or processed dataThis refers to data that has been digitized, cleaned, corrected, translated, transcribed, validated, verified or anonymized, but has not yet been transformed into final results (graphs, already interpreted models, etc.).
  • Data analyzedThese are the result of applying analytical techniques to primary or secondary data. They include models, tables, graphs, interpretive texts, and other products that serve to draw conclusions and support decision-making.

Regarding their source , three main groups can be distinguished:

  • Experimental dataGenerated in controlled environments (laboratories, test benches, clinical trials) through the manipulation of variables and the recording of results. A typical example would be chromatography or spectroscopic measurement.
  • Observational data: collected without directly intervening on the phenomenon, as happens in surveys, cohort studies, observation of animal or human behavior, environmental monitoring, etc.
  • Computational or simulation data: obtained through numerical models, simulations, algorithms or calculation tools that generate data from input parameters and defined rules.

Classification by technical format is also important , as it determines how the data can be stored and reused. The data can be:

  • Textual (for example, Word documents, PDFs, RTFs, and similar).
  • Numerical (Excel spreadsheets, CSV files, statistical software files, etc.).
  • Multimedia (JPEG and PNG images, MPEG video files, WAV audio recordings, among others).
  • Structured (databases in XML, SQL, MySQL, PostgreSQL, etc.).
  • Software code (Java, C, Python and other languages).
  • Software or discipline-specific formats (3D CAD models, mesh fabrics, statistical model files, proprietary formats of scientific instruments).

In practice, a single project usually works with a combination of several of these types , which requires designing management strategies adapted to each format and processing phase.

Why research data management (RDM) is key

Research data management, often referred to by its acronym RDM , encompasses the entire data lifecycle : from planning and collection to archiving and long-term preservation. It includes tasks such as organizing, documenting, securely storing, publishing, and reusing the data generated or used in a project.

Effective data management enables researchers to work more efficiently , avoiding duplication, facilitating teamwork, and allowing other groups to reuse data in subsequent studies. Furthermore, it helps researchers meet the requirements of funding agencies, publishers, and universities , which increasingly demand greater clarity regarding data usage.

The main benefits of proper management include:

  • Compliance with funding requirementsMany programs, such as Horizon 2020 or Horizon Europe, require data management plans and publication in open access where possible.
  • Greater transparency and traceabilityGood data documentation facilitates the validation of results by the scientific community.
  • Improved protection and securitySystematic management reduces the risk of loss, corruption, or unauthorized access to sensitive data.
  • FAIR Data (findable, accessible, interoperable, reusable): applying these principles ensures that data can be located, accessed, combined and incorporated into new studies.
  • Saving time and resourcesAvoiding repeating experiments or fieldwork because data was lost or it is not known how to interpret it represents a tangible improvement in research efficiency.

Many universities, libraries, and research support services currently offer guides, infographics, and best practice materials on data management, providing guidance on technical, organizational, and legal aspects: from how to name files to how to comply with the GDPR in projects involving personal data.

FAIR principles: make data findable, usable, and reusable

In 2016, the FAIR Principles were published in Nature's Scientific Data journal , marking a turning point in the way scientific data is managed. Their impact has been so significant that the European Commission incorporated them as a reference in Horizon 2020 projects and maintains them as a standard in Horizon Europe.

The FAIR principles are not a rigid law or standard, but rather a set of measurable qualities that a data publication should possess to be Findable, Accessible, Interoperable, and Reusable. The goal is for both people and machines to be able to locate, understand, and reuse data in the best possible way.

(F) Findable: data that can be found

To make data discoverable, it's not enough to simply store it on a personal hard drive. It must be deposited in appropriate locations , such as trusted repositories or scientific journals that accept datasets, and accompanied by rich and structured metadata.

The FAIR principles place particular emphasis on the use of persistent identifiers , such as DOIs for datasets, ORCIDs for authors, and RoRs for institutions. These identifiers ensure that data remains searchable even if the URL or hosting server changes.

It may interest you:  Transformative education: keys to changing schools and society

In practice, making data findable involves defining good keywords, clear descriptions, and metadata standards accepted by the relevant scientific community, so that search engines and data catalogs can index them correctly.

(A) Accessible: accessible data

Accessibility doesn't mean everything has to be completely open, but rather that data and its metadata can be obtained in a clear and regulated manner . The principle of "as open as possible, as closed as necessary" is recommended.

This means that everything possible should be made open , clearly indicating the access conditions (e.g., immediate open access, temporary embargo, restricted access upon request, etc.). In any case, even when the data cannot be fully published (for ethical, legal, or confidentiality reasons), the metadata should remain accessible so that the community knows the dataset exists.

Accessibility also involves depositing, identifying, and correctly describing data in a repository that offers standard download and access mechanisms through open and well-documented protocols.

(I) Interoperable: data that is understood between systems

For data and its metadata to be combined with other datasets and used in different contexts, interoperability is essential . This is achieved by using open community standards, controlled vocabularies, and well-defined metadata schemas.

Interoperability allows information to flow between people, institutions, and machines without needing to be manually rewritten or converted each time. For example, using open formats like CSV or XML instead of proprietary files makes it easier for different tools to read.

It is also recommended to use metadata schemas appropriate to the data type (e.g., Dublin Core, DataCite, Darwin Core, etc.), link the data to other relevant resources, and avoid closed formats that limit reuse.

(R) Reusable: data reusable by third parties

The final FAIR pillar refers to reuse. For a dataset to be reused in new studies, its origin, conditions of use, and context of generation must be clearly defined . Otherwise, even if it is technically accessible, it will be of little use.

Good practices to promote reuse include the use of appropriate open licenses (such as Creative Commons or data-specific licenses), a detailed description of how the data was collected and processed, and the adoption of metadata and formats widely accepted in the discipline.

There are various assessment tools and services that allow you to verify the extent to which a dataset complies with the FAIR principles, offering recommendations to improve its findability, accessibility, interoperability and reuse.

Open Science, Open Data and research data repositories

Research data falls squarely within the Open Science movement , which promotes free access to both scientific publications and the data on which they are based (Open Data). The idea is that when research is publicly funded, the resulting data should be openly available unless there are well-justified reasons to restrict it.

Open data allows it to be used, reused, and redistributed, extending its impact beyond the original project. Furthermore, its proper publication ensures better preservation, dissemination, and visibility , which typically translates into more citations, collaborations, and research opportunities.

In response to these needs, a growing number of universities and research centers are offering dedicated data repositories , where datasets linked to institutional projects are hosted. However, the landscape is highly diverse, with numerous repositories specializing in different disciplines, countries, data types, or formats.

Given this highly fragmented ecosystem, the research community faces the challenge of locating the right repository to store or find the data it needs. To help with this task, re3data.org was created, an international registry of research data repositories that compiles metadata from thousands of specialized repositories.

Thanks to re3data.org, researchers, funding agencies, libraries, and publishers can explore repositories and filter them by discipline, country, content type, format, license, language, and other criteria. The registry identifies nearly 2.000 data repositories, making it one of the most comprehensive catalogs available.

The re3data.org project began as a joint initiative of several German organizations, funded by the German Research Foundation. It later integrated the DataBib catalog to avoid duplication, in a merger driven by DataCite, an international non-profit organization focused on improving data citations . It also collaborates with open science projects such as BioSharing and OpenAIRE.

It is no coincidence that publishers, research institutions, and funders cite re3data.org in their policies as a reference tool for identifying appropriate repositories. The European Commission, for example, mentions it in its guidelines on open access to scientific publications and research data within the framework of Horizon 2020.

Data management plans in European projects

Within the European Union, the Commission stipulates that all projects funded under Horizon 2020 (and subsequently Horizon Europe) must develop a Data Management Plan (DMP). Furthermore, the data generated is expected to be shared as openly as possible and to adhere to the FAIR principles.

It may interest you:  Educational transformation in Spain half a century after Franco

The DMP is a living document that explains what data will be generated, how it will be managed during the project , and what will be done with it once the project is complete. It typically follows official templates, such as the "Horizon Europe Data Management Plan Template," and is updated throughout the project's lifecycle.

A typical data management plan usually includes the following fundamental sections, which can be adapted to the nature of the project:

General project information

This first section contains the basic project data: project title and identifier , brief description, coordinating institution, funding agency, principal investigator with their identifier (e.g., ORCID), contact information and reference to the different versions of the management plan that are generated.

1. Data Summary

This section provides a general overview of the data that will be used or produced in the project. It details the type and format of the data , its purpose, approximate size, origin (newly generated data, data reused from other sources, etc.), and its expected usefulness to the community.

2. FAIR Data

This part of the plan explains how the FAIR principles will be implemented in the project. It is usually broken down into several subsections to address each dimension separately: findability, accessibility, interoperability, and reusability.

The section on findability indicates the identifiers that will be used (e.g., DOI for datasets), the keywords and descriptors, as well as the metadata rules that will be applied to optimize the location of the data.

The accessibility subsection describes the repository or repositories where the data will be stored , whether it will have a persistent identifier, which data will be publicly accessible and which will remain closed, detailing embargo periods and any usage restrictions . It also typically mentions whether the metadata for closed data will remain accessible.

The section dedicated to interoperability sets out the vocabularies, standards, formats and methodologies that will be used to facilitate the exchange and interoperability of data with other systems and projects, indicating how the standards of the corresponding scientific community are respected.

Finally, the subsection on data reuse documents the origin of the data and provides the information necessary to validate the results and enable reuse. This typically includes references to README files, code documentation, protocols, and, very importantly, the usage licenses under which the data will be published.

3. Other research findings

Beyond the raw data, many projects generate other reusable products : software, workflows, experimental protocols, new materials, samples, or tools. This section analyzes which FAIR principles can also be applied to these results and how they will be managed and shared.

The idea is to ensure that, as far as possible, all generated products that may be of interest to the community are visible, accessible and reusable , always respecting issues of intellectual property, confidentiality and ethics.

4. Resource allocation

Data management is not free: it requires time, infrastructure, and, in many cases, specialized support. Therefore, the DMP must detail what resources will be allocated to make the data FAIR . This includes both direct costs (storage, archiving, technical staff) and indirect costs associated with security and preservation.

This section also clarifies who will be responsible for data management within the team: it can be a specific person, a group, or a combination of research staff and technical support.

5. Data security

Ensuring data is stored securely is a priority, especially when dealing with sensitive or confidential information. This section of the plan addresses issues such as storage in trusted repositories , backups, data loss mitigation strategies, and access control measures.

It also describes how long-term preservation will be managed , specifying, for example, which archival repositories will be used and for how long the data will be preserved after the project is completed.

6. Ethics and legal aspects

Many projects handle personal, clinical, or other sensitive data. This section details the ethical and legal issues that may affect data sharing and disclosure, including compliance with data protection regulations, anonymization processes, and confidentiality agreements.

When research involves personal data, it is essential to explain how informed consent will be obtained , how long the data will be kept, under what conditions it can be reused, and how the privacy of the participants will be guaranteed.

7. Other relevant issues

Finally, the plan may incorporate any other element related to data management, such as the application of national, institutional or sectoral policies , alignment with funding agency guidelines or integration with the institution's open science strategies.

In summary, it becomes clear that research data is much more than a collection of files stored in a folder: it forms the foundation upon which the credibility of modern science rests. Understanding what constitutes data, how it is classified according to its origin, format, and processing, and how its management should be planned following principles such as FAIR and the requirements of European programs, allows any research team to work more efficiently, meet funding requirements, and, above all, contribute to a more open, reusable, and sustainable science.

Related articles:
What are the scope and limitations of an investigation: A comprehensive analysis