Real-World Data: how the EMA's data quality framework is reshaping regulatory decision-making

What the new guidance means for data quality, fitness for use, and the role of AI and synthetic data in regulatory readiness

Key takeaways

  • The European Medicines Agency (EMA) has developed an operational framework for assessing the quality of Real-World Data (RWD) used for regulatory purposes.
  • Data quality is assessed across three levels: processes, datasets, and context of use.
  • The principle of fitness for use becomes central: a dataset is considered adequate only if it is suitable for addressing a specific regulatory question.
  • Checklists, metrics, and maturity models make data quality a documentable and verifiable process.
  • For the life sciences sector, the importance of data governance, traceability, interoperability, and secure data reusability continues to grow.
  • Artificial intelligence and synthetic data can support the practical implementation of the framework’s principles by improving the quality, completeness, and secure reusability of healthcare datasets.

Regulatory context

The use of RWD is playing an increasingly important role in generating evidence to support regulatory decision-making. Electronic health records, disease registries, administrative databases, prescription data, and other observational data sources are being used across multiple stages of the medicinal product lifecycle, from pharmacovigilance and post-authorisation studies to supporting safety, effectiveness, and monitoring assessments.

This evolution is also changing the way Real-World Evidence (RWE) is evaluated. It is no longer sufficient to assess only the results generated by a study; it has become essential to understand and document the quality of the data from which that evidence is derived.

To address this need, on 30 October 2023, the European Medicines Agency (EMA) adopted the Data Quality Framework for EU Medicines Regulation (EMA/326985/2023), a general framework that establishes the principles, terminology, and quality dimensions applicable to data used in medicines regulation. The implementation guidance specifically addressing RWD (EMA/503781/2024), adopted on 16 March 2026, translates these principles into operational guidance for assessing the quality, reliability, and fitness of observational datasets for specific regulatory use cases.

The framework is also part of a broader European initiative aimed at making health data more reliable, interoperable, and securely reusable. This approach is aligned with initiatives such as DARWIN EU and the European Health Data Space (EHDS), which focus on developing infrastructures, standards, and governance models for the secondary use of health data to support research, innovation, and evidence-based regulatory decision-making.

DARWIN EU provides the federated infrastructure through which the European regulatory network can generate evidence based on RWD, while the EHDS establishes the future legal and operational framework for the secondary use of health data. TEHDAS2 contributes to translating the EHDS framework into operational procedures, formats, and guidance for Health Data Access Bodies (HDABs), data holders, and data users. Although they serve different purposes and operate in distinct domains, the EMA framework and the EHDS converge on common requirements, including documentation of data provenance, measurable data quality, process maturity, interoperability, and data fitness for the intended use.

Why a dedicated framework for Real-World Data?

Real-world data differ substantially from data collected in traditional clinical trials. The EMA identifies four key distinguishing characteristics:

  • Heterogeneity of data sources;
  • Secondary use of data beyond their original purpose of collection;
  • Complexity of data governance;
  • Challenges related to the reproducibility of evidence.

Real-world data are generally collected for clinical, healthcare, or administrative purposes and are only subsequently reused for research, regulatory assessments, or policy and governance decisions. As a result, they may contain missing data, coding inconsistencies, variability in data collection processes, different levels of data currency, and information that is not always fully aligned with the specific regulatory question being addressed.

For this reason, data quality cannot be assessed solely by examining the contents of a dataset. It is equally important to consider the context in which the data were generated, transformed, documented, governed, and ultimately used.

The three data quality determinants

The EMA framework proposes a methodology structured around three complementary levels, referred to as Data Quality Determinants.

Foundational determinants

Foundational determinants relate to the systems and processes through which data are generated. They encompass aspects such as data provenance, governance, the Quality Management System (QMS), quality control procedures, data transformations, documentation, and metadata management.

This level provides an understanding of how the data were produced, which processes influenced their quality, and the level of process maturity achieved by the organization managing them. To support this assessment, the framework introduces checklists and maturity models that enable organizations to describe the development and robustness of their quality processes in a structured and consistent manner.

Intrinsic determinants

Intrinsic determinants concern the characteristics that can be observed directly within the dataset itself. Data quality is described through five key dimensions:

  • Reliability
  • Extensiveness and completeness
  • Coherence
  • Timeliness
  • Relevance

In the implementation guidance dedicated to RWD, intrinsic quality metrics are used to characterize the dataset’s reliability, extensiveness and completeness, coherence, and timeliness. Relevance, however, is closely linked to the overall assessment of fitness for use, as the adequacy of a dataset’s quality can only be determined in relation to a specific regulatory question and the analytical approach adopted to address it. The framework also provides examples of quality metrics and documentation practices designed to make assessments more transparent, consistent, and comparable across use cases.

Question-specific determinants

Question-specific determinants introduce the principle of fitness for use. The quality of a dataset is not an absolute property but depends on the specific regulatory question the dataset is intended to answer.

A dataset may be suitable for a pharmacovigilance study but not for a comparative effectiveness assessment. Similarly, a data source may be adequate for describing an observational trend but insufficient to support a regulatory decision requiring a higher degree of granularity, completeness, or representativeness.

This principle is one of the cornerstones of the framework: data quality must always be assessed in relation to the intended context of use, the objectives of the study, and the level of evidence required.

From Principles to Practice: The Framework’s Operational Tools

The EMA document goes beyond establishing general principles by introducing practical tools to support implementation of the framework. These include:

  • Checklists for characterizing data systems, processes, and governance;
  • A methodology for identifying and documenting data quality metrics;
  • A maturity model for assessing the development of quality management processes;
  • A structured methodology for evaluating fitness for use;
  • Practical examples illustrating how data quality should be documented in regulatory submissions.

A key feature of the framework is the EMA’s decision not to define universal quality thresholds or prescribe a mandatory set of quality metrics. Instead, data quality assessments should be proportionate to the specific regulatory use case and documented in a transparent manner.

This approach provides flexibility while placing greater responsibility on stakeholders to justify their methodological choices, demonstrate the suitability of the selected dataset, and ensure that the processes supporting its use are fully traceable.

Are your RWD datasets ready for EMA and EHDS requirements?
👉 Contact our experts to get your data ready

Implications for the Life Sciences sector

The EMA’s data quality framework for Real-World Data (RWD) redefines responsibilities across the entire data value chain.

For data holders, documenting data provenance, metadata, collection processes, data transformations, and quality controls becomes increasingly important. Simply making data available is no longer sufficient; organizations must be able to demonstrate its reliability, traceability, and suitability for potential regulatory uses.

For study sponsors, Contract Research Organizations (CROs), and researchers, the framework requires a more structured assessment of the suitability of the selected dataset. The choice of data source should be justified in light of the study design, target population, study endpoints, availability of relevant variables, and any inherent limitations.

For regulatory authorities, the framework provides a harmonized approach for assessing not only the analyses submitted but also the quality of the underlying data on which those analyses are based.

As a result, data quality becomes a demonstrable, measurable, and documentable attribute rather than an implicit assumption. This requires investment in data governance, interoperability, quality management processes, and technologies capable of improving both the quality and the reusability of healthcare datasets.

How Aindo can support the implementation of the framework

Aindo has adopted an integrated system for data governance, preparation, generation, and validation that is fully aligned with the principles of the EMA framework. The objective is not merely to improve datasets from a technical perspective, but to ensure that the processes applied are fully documented, the quality of the results is measurable, and their suitability for the intended use can be clearly demonstrated.

Within this framework, synthetic data generation plays a central role in addressing one of the key challenges associated with RWD: their reuse for purposes different from those for which they were originally collected. Synthetic datasets can preserve the statistical properties and meaningful relationships contained in the original data while simultaneously reducing the risks associated with sharing and processing information relating to identifiable individuals.

Synthetic data generation is supported by data preparation and enhancement capabilities that create the conditions necessary for reliable data reuse:

  • Data integration harmonizes heterogeneous data sources, improving consistency, interoperability, and traceability.
  • Data imputation supports the management of missing values, enhancing dataset completeness.
  • Data rebalancing corrects imbalances in the representation of populations and subgroups, reducing the risk that such biases could compromise the suitability of the data for addressing the regulatory question.

Building on these capabilities, Aindo’s platform translates the three levels of Data Quality Determinants defined by the EMA into processes, controls, and metrics that can be applied throughout the entire synthetic data preparation, generation, and validation lifecycle.

At the level of foundational determinants, Aindo has established its processes within a formalized governance and quality management system. Its ISO 9001:2015-certified Quality Management System and Europrivacy certification provide verifiable evidence—through audit activities—of responsibilities, procedures, data protection measures, and quality controls. These certifications are complemented by structured documentation that enables the reconstruction of data provenance, applied transformations, system configurations, and the various stages of the data generation and validation process.

While these certifications do not replace the assessment of an individual synthetic dataset nor automatically demonstrate fitness for use, they constitute important process-level evidence supporting the foundational determinants defined by the EMA framework.

At the level of intrinsic determinants, the validation of synthetic datasets combines statistical fidelity, consistency, and information coverage metrics with record proximity and privacy assessment metrics, including Distance to Closest Record (DCR), Nearest Neighbour Distance Ratio (NNDR), and Membership Inference Attack (MIA) testing. Although these metrics are not direct equivalents of the EMA quality dimensions, their combined application can provide evidence of reliability, consistency, extensiveness and completeness, while also demonstrating the level of privacy protection achieved.

At the level of question-specific determinants, Aindo applies a decision-making process that links the selection of the generative model and validation strategy to the principle of fitness for use.

The choice between statistical approaches, traditional machine learning models, and generative AI models takes into account the structure and characteristics of the data, the available sample size, the intended analyses, the specific regulatory question, and the required level of privacy protection.

Consequently, the assessment goes beyond determining whether the generated data are statistically similar to the original data. It also verifies that the synthetic dataset preserves the characteristics required for the specific intended use.

Conclusions

The EMA’s data quality framework for Real-World Data represents a significant evolution in the European regulatory approach. Data quality is no longer regarded as a static property of a dataset but as the outcome of documented processes, measurable metrics, and assessments performed within the context of a specific regulatory use case.

For organizations operating in the life sciences sector, this means strengthening data governance, traceability, interoperability, and quality management processes. Demonstrating the fitness for use of datasets is becoming a fundamental requirement for the use of real-world data in generating regulatory evidence.

The transition from methodological principles to operational implementation is already underway. The EHDS Regulation requires Member States to notify the European Commission of the designated Health Data Access Bodies (HDABs) by 26 March 2027 and, by the same date, foresees the adoption of technical specifications for dataset descriptions and the Data Quality and Utility Label. This label will include elements such as documentation and metadata, completeness, accuracy, validity, timeliness, consistency, maturity of quality management processes, coverage, and data transformations. The general framework governing secondary use of health data and the issuance of data permits will become applicable on 26 March 2029.

Although the EMA framework is not formally incorporated into the EHDS Regulation and the two instruments serve different purposes, the convergence between the EMA’s Data Quality Determinants and the EHDS data quality and utility requirements points in a clear direction: data provenance, process maturity, quality metrics, coverage, transformations, and fitness for use will need to be demonstrated—not merely declared.

Within the same ecosystem, DARWIN EU operationalizes the federated generation of evidence based on RWD, while the EHDS establishes the legal and organizational framework for the secondary use of health data, including the documentation and assessment of healthcare dataset quality.

For organizations preparing for the EHDS environment, the window for adapting processes, systems, and documentation is already open. Entering the first fully operational procedures without a well-established capability to demonstrate fitness for use may result in delays, requests for additional evidence, or the inability to unlock the value of data in a timely manner.

In this context, artificial intelligence and synthetic data can play a strategic role only when embedded within a formalized governance and quality management system, supported by recognized certifications, verified through audit activities, and validated against the specific intended use.

Join Us

Want to learn more or work with us? We'd love to hear from you.