Responsible data management can help accelerate research and improve collaboration in rare diseases.

The FAIR principles: making rare disease data work better

What does FAIR mean?

In 2014, a group of scientists, academics, and industry representatives, led by Barend Mons, met in Leiden to discuss an emerging concept called FAIR principles. Their work laid the foundation for a 2016 Scientific Data paper, with ERDERA data expert Mark D. Wilkinson as first author. Today, this paper is considered the canonical reference for the FAIR principles.

FAIR is a set of guiding principles for managing and reusing scientific data. The aim is to keep data useful beyond their original purpose so they can be discovered, accessed and reused in a responsible way.

FAIR applies to data management across disciplines and is widely recognised in research policy and funding frameworks, including European research programmes.

FAIR focuses on making data usable for both humans and machines, in line with the increasing role of computational analysis in modern research.

FAIR emphasises clear rules that support responsible reuse while respecting ethical, legal and societal constraints. It does not require data to be open or freely accessible to everyone.

Where FAIR applies?

The FAIR principles work across three connected elements, which can be understood through the analogy of a dictionary.

data

Data are like words in a dictionary. Some words can have more than one meaning depending on the context. For example, the word bank can refer to a financial institution, the side of a river, or a group of fish. In the same way, data are pieces of information that may be difficult to interpret on their own. In research, data include datasets, registries, measurements, observations or digitally described samples

Infrastructure is everything that brings words and definitions together in a useful and organised way. It includes not only the dictionary itself, but also the rules, standards and systems that keep information consistent and easy to find. In a FAIR ecosystem, infrastructure works much like the Internet: shared protocols, standards and services allow data and metadata to be exchanged and understood across different systems. Catalogues, repositories and platforms help users discover information, but they are only one part of the wider infrastructure that makes data searchable, connected and reusable.

In a well-organised dictionary, each word is not only defined but also uniquely identified, for example through a reference number that allows it to be found unambiguously. This ensures that its meaning remains consistent across contexts. This combination of clear definitions, unique identifiers and structured organisation is what allows data to be understood and reused not only by people, but also by machines.

FAIR works best when all three elements are in place. For example, one FAIR principle requires that both data and their descriptions are listed in searchable systems. This shows why infrastructure is essential: without it, even well-described data would remain invisible.

FAIR principles one by one

The different elements of FAIR are connected, but each serves a distinct purpose. Each principle covers a different aspect of data reuse, so the principles can be applied in a practical way for different types of data and research contexts.

Data are findable when they can be easily located and their relevance can be recognised.

In a well-organised dictionary, each word is not only defined but also uniquely identified, for example through a reference number that allows it to be found unambiguously. This ensures that its meaning remains consistent across contexts. This combination of clear definitions, unique identifiers and structured organisation is what allows data to be understood and reused not only by people, but also by machines.

FAIR works best when all three elements are in place. For example, one FAIR principle requires that both data and their descriptions are listed in searchable systems. This shows why infrastructure is essential: without it, even well-described data would remain invisible.

Like a dictionary entry that allows a word to be correctly used in new sentences, well-described data can be understood and applied in new contexts over time.

Data are accessible when there is a clear and reliable way to access or request them.

This principle does not require all data to be openly available. Instead, it requires that the process for accessing data is clearly described and uses standard procedures. Even when data cannot be shared directly, information about them should remain available so others can discover that they exist and understand how access may be requested. In health research, access often needs to be managed carefully to protect privacy and sensitive information.

Similarly, a dictionary may not be available directly on the shelf. Users might need to request it through a librarian or follow a specific procedure before consulting it. What matters is that the process is clearly explained, so people know how access can be requested.

Data are interoperable when they can be connected, combined and understood across different systems.

This is possible because they use shared standards, common formats and agreed vocabularies to describe information in a consistent way. By using the same terms and definitions to represent the same concepts, data from different sources can be interpreted correctly by both people and computers, reducing misunderstandings and making it easier to work with information from multiple locations.

This is especially important in rare diseases, where progress depends on bringing together small pieces of information from many places and disciplines.

For example, a condition such as Duchenne muscular dystrophy may be recorded in several registries, hospitals or research projects. The problem is that one source may refer to “Duchenne muscular dystrophy”, another to “DMD”, and another to a standard identifier. If each source uses the same identifier and vocabulary to describe this disease, the data can be connected and analysed together. Shared vocabularies help ensure that all of them are recognised as the same disease. This helps researchers compare results across countries and build a more complete picture of the condition.

In the dictionary analogy, this is what ensures that words are described using a shared structure and vocabulary, and have a consistent meaning and definition, so they can be understood and used across different contexts. This ensures that words have the same meaning wherever they appear, allowing information from different sources to be linked and understood together.

Data are reusable when they remain useful beyond their original purpose.

They are accompanied by clear explanations that help others understand where the data come from and how they can be used responsibly. This allows data to support new research questions in the future, reducing the need to collect similar data again. As a result, research efforts are not lost to repeating work that has already been done, and progress can continue over time.

Like a dictionary entry that allows a word to be correctly used in new sentences, well-described data can be understood and applied in new contexts over time.

Why FAIR data matters

The FAIR principles emerged in response to a growing challenge in research: researchers were generating large amounts of valuable data, but it was difficult to find, understand or reuse once the project ended. This reduced the impact of research investments and slowed scientific progress, particularly in rare disease research, where data are scarce and widely dispersed.

FAIR addresses this by improving how data are organised and prepared from the start. When data are easier to discover, access, combine and reuse, researchers can build on existing knowledge instead of starting again. Importantly, FAIR is designed to make data understandable and reusable not only for humans, but also for machines, enabling automated data discovery, integration and analysis at scale. This reduces duplication, supports collaboration across borders and disciplines, and helps research move forward more efficiently.

At the same time, FAIR promotes transparent and ethical data practices, with clear conditions for access and reuse that respect privacy, ethics and data ownership. This clarity helps foster trust among researchers, data holders and patients, and supports research systems that handle health data responsibly. Over time, FAIR approaches ensure that research data continue to generate value beyond a single project, benefiting both science and society.

How ERDERA supports FAIR data

In ERDERA, FAIR principles shape how rare disease data are handled from the start.

Researchers and data holders are supported in preparing their data so it can be described clearly, connected to other resources and reused beyond a single project. Making data FAIR is a collaborative process between data holders and FAIR service providers. By contributing FAIR data, researchers help enable larger and more powerful studies that combine information from multiple sources, ultimately advancing rare disease research and improving patients’ lives.

This support is provided through the Data Services Hub, which brings together expertise, shared methods and technical tools, while ensuring that data remain under the control of those who collect them.

By working this way, ERDERA helps reduce fragmentation, supports long term use of research results and makes collaboration across countries and disciplines easier in rare disease research.

Key takeaway:
FAIR principles help make data easier to find, understand, connect and reuse, enabling better collaboration and more efficient research.

You might also be interested in

Ana Rath, Data Services Co-Lead in ERDERA, explains how the consortium is helping researchers and data holders make rare disease data more findable, interoperable and reusable while protecting privacy and keeping data holders in control.
ERDERA contributed to the first Journée nationale FrBioNet, bringing a European rare disease perspective to discussions with the French biobanking community.
Held in Riga on 9–10 June, the workshop brought National Mirror Group experts, researchers, clinicians and policymakers together to exchange practical lessons on how national rare disease registries can better support research and alignment across countries.
A new online learning series for ERN professionals, clinicians, researchers, and stakeholders.