
Building External Control Arms with Synthetic Data
A rigorous, transparent, and regulatory-ready approach to accelerating clinical research with synthetic control arms
What the new guideline means for regulatory questions, context of use, risk, evaluation, and documentation of evidence generated through modelling and simulation
Model-Informed Drug Development (MIDD) involves the use of computational modelling and simulation methods to integrate nonclinical data, clinical data, prior information and knowledge about the medicinal product, disease and target population. The resulting evidence can contribute to development planning and to decisions made by pharmaceutical companies, regulatory authorities and other stakeholders.
Approaches covered by the guideline include population pharmacokinetics and pharmacodynamics, physiologically based pharmacokinetic and biopharmaceutics models, exposure-response analyses, model-based meta-analyses, quantitative systems pharmacology and toxicology, agent-based models, disease progression models and artificial intelligence (AI) and machine learning (ML) approaches. These methods may be used individually or in combination.
The final text of the ICH M15 guideline was adopted by the EMA’s Committee for Medicinal Products for Human Use (CHMP) in January 2026 and became effective on 23 July 2026. In June 2026, the US Food and Drug Administration (FDA) also published its final guidance based on the same harmonised text.
ICH M15 is not a technical manual for building a particular type of model. It does not prescribe a specific algorithm, software package, or universal metric. Rather, it defines the principles for determining whether a model, the data used and the resulting outcomes are sufficiently reliable and relevant to contribute to a specific decision.
The guideline applies to both established methods and emerging approaches and should be used in conjunction with relevant ICH guidelines, recognised scientific standards and accepted practices for the methodology concerned. Its aim is to establish a common language among modelling experts, clinicians, statisticians, data scientists, quality specialists, regulatory functions and regulatory assessors.
Early planning plays a central role. Integrating the MIDD strategy from the early stages of development makes it possible to generate the necessary data, define evaluation criteria prospectively and engage with regulatory authorities before results become available.
This approach is already being put into practice. In May 2026, EMA launched a pilot procedure providing scientific advice and protocol assistance for programmes in which MIDD evidence is expected to play a significant role in regulatory decision-making. Eligibility is assessed in part on the basis of the question of interest, context of use and expected model impact.
The quality of a model cannot be assessed in isolation.
The same model may be used for an internal exploratory analysis, to complement already robust clinical evidence, or as the main source of information supporting a decision on dose, population, study design, or therapeutic indication.
The model may be technically identical across these different scenarios, but the required level of confidence changes substantially. The greater the weight of the model in the decision and the more serious the consequences of a potential error, the more rigorous its evaluation should be.
ICH M15 therefore shifts the focus from the generic question “Is the model valid?” to a more specific one: “Are the model, the data and the resulting outcomes appropriate for answering this question, in this context of use and given these potential decision consequences?”
The guideline also distinguishes between model outcomes and evidence supporting MIDD. Model outcomes include predictions, simulations and related conclusions. They become MIDD evidence only when application of the assessment framework, including technical model evaluation, supports their appropriateness for contributing to the answer to the question of interest.
Regulatory credibility is therefore not an absolute property of a model. It is a context-dependent conclusion based on the relationship between the question, context of use, role of the model, other available evidence, consequences of a wrong decision and the robustness of the evaluation performed.
At the core of ICH M15 are six elements that should be described in the MIDD evidence assessment table both during planning and when evidence is submitted to regulatory authorities. For the four elements subject to classification (model influence, consequence of wrong decision, model risk and model impact), a low, medium or high rating should be accompanied by a rationale.
The question of interest is the question that the MIDD strategy is intended to help answer. It should be explicitly stated, reflect the information needed at the relevant stage of development, and support multidisciplinary assessment and regulatory decision-making.
The question may be broader than the intended use of an individual model. When a strategy is intended to answer different questions, the guideline recommends preparing separate assessment tables. This ensures that each decision remains clearly linked to the relevant models, data and technical criteria.
The context of use describes the role and scope of the model used to address the question. It should clarify:
Therefore, it is not sufficient simply to name the methodology. The context of use should explain what the model will do, the boundaries within which it will be used, and how its outcomes will be integrated with the other available information.
Model influence represents the intended weight of the model outcomes in the decision, taking into account the contribution of other evidence.
It should be described and classified as low, medium, or high. When model outcomes are the sole source of information supporting the decision, model influence should generally be considered high. When relevant data and evidence from other sources are available, for example, clinical trials, nonclinical studies or evidence generated in clinical practice, model influence may be lower, depending on the actual weight assigned to the model.
Consequence of wrong decision refers to the potential negative effect of an incorrect decision, for example on patient safety or treatment efficacy. The classification should consider both the severity of the possible consequences and the likelihood that they will occur, based on all information available at the time of assessment.
Model risk represents the contribution of model outcomes to a potential wrong decision and to the undesirable consequences that may result.
It is derived by combining model influence and the consequences of a wrong decision. If both are low, model risk will generally be low. If both are high, model risk will be high. When the two assessments differ, the overall judgement may be driven by the more influential consideration, but the rationale should be documented.
The guideline does not impose a numerical formula or a rigid matrix. It requires a reasoned, transparent and understandable classification. The resulting model risk is primarily used to determine the depth and rigour of the technical evaluation.
Importantly, model risk is not an intrinsic risk of modelling or simulation. It always exists in relation to a specific question and to the model’s contribution to a particular decision.
Model impact reflects the extent to which the proposed MIDD strategy differs from existing regulatory standards or, where no formal standard exists, from regulatory expectations applicable to the question of interest.
Model impact should also be classified as low, medium, or high and accompanied by a rationale. An established use already addressed in specific guidelines may have limited impact. An innovative use intended to replace evidence normally obtained through traditional studies may instead have high impact.
Model influence and model impact are not synonymous. Model influence measures the weight of the model in the decision; model impact measures the extent to which the strategy departs from regulatory standards or expectations. Therefore, a model may have high influence but limited impact, or introduce a highly innovative methodology while contributing alongside several other sources of evidence.
Four additional components complement the six key elements and accompany the strategy from planning through evidence submission.
During planning, the sponsor should define the technical criteria against which the model and its outcomes will be evaluated. These criteria should be specific to the question of interest and commensurate with model risk. The same level of evaluation is therefore not required for every analysis, but the selected level should be justified.
The appropriateness of the proposed MIDD strategy should also be described: why the use of that model, those data and that combination of evidence is suitable for answering the question. The rationale should consider the context of use, the weight assigned to the model, other sources of information and the ability of the technical criteria to demonstrate the reliability of the outcomes.
Once the analyses have been completed, the following should be reported:
The conclusion included in the assessment table represents the sponsor multidisciplinary team’s assessment. It does not correspond to the outcome of the subsequent review by the regulatory authority.
ICH M15 divides technical model evaluation into two main areas: verification and validation with applicability assessment. The depth of evaluation should at least meet recognised scientific standards for the methodology used and should be commensurate with model risk.
The guideline does not require verification, validation and applicability assessment to be performed by an independent third party. The relevant technical activities may be performed internally or entrusted to qualified parties; however, the drug developer submitting the MIDD strategy remains responsible for defining the technical criteria, ensuring appropriate evaluation and documentation, and integrating the results into the multidisciplinary assessment of the evidence.
Verification should establish that:
In practical terms, presenting final plots and performance indicators is no longer sufficient. It should be possible to reconstruct the code used, versions, data transformations, analytical steps and the conditions under which the result was obtained.
Validation considers the overall comparison between the model, data, prior information and available knowledge. Applicability assessment, by contrast, determines whether the data and the model are appropriate for each intended use.
Among other aspects, the guideline calls for:
Method-specific issues should also be considered, including selection bias in model-based meta-analyses, knowledge gaps in mechanistic models and overfitting in AI/ML approaches.
External validation using independent data is encouraged and, depending on the question of interest, context of use and risk, may increase confidence in the results or, in some cases, become essential for the proposed application. Independence in this context refers to the data used for validation and does not, in itself, imply the use of an external validation body.
Applicability and the overall appropriateness of the strategy remain distinct concepts. Applicability concerns the technical suitability of the data and model for the intended use. Appropriateness concerns the ability of the overall MIDD strategy to answer the regulatory question.
The guideline introduces a documentation structure linking planning, conduct of analyses, evaluation and submission to regulatory authorities.
The assessment table is the central communication tool. It should concisely report the six key elements, technical criteria, rationale for the appropriateness of the strategy and, once the analyses have been completed, the evaluation summary and conclusion on the MIDD evidence.
The table should be used during interactions with regulatory authorities, updated as the plan evolves, and included in the most appropriate sections of the regulatory documentation. Supporting analysis documents should be cross-referenced.
The Model Analysis Plan (MAP) prospectively documents each planned analysis. Predefinition should take place before accessing the data or conducting the analysis, as appropriate to the context of use.
The plan generally includes an introduction, objectives, data and methods. It should also describe the planned model-evaluation activities and the technical criteria that will be used to assess the results. Having a prospectively defined plan can make discussions with regulatory authorities more effective, helping to avoid technical criteria and methodological decisions being established only after the results have been observed.
The Model Analysis Report (MAR) documents the results of each analysis submitted to regulatory authorities.
The report normally includes an executive summary, rationale, objectives, data and methods, results, discussion, conclusions and appendices. It should describe model development and evaluation, predictions or simulations, uncertainty, comparison against the technical criteria and any changes from the plan.
Where a MAP has been prepared, it should be appended to the report: any deviation from the analyses specified in the plan should be described and justified in the report.
Traceability also extends to the entire analytical chain. Data, programming scripts, definition files and other electronic documents used to generate the evidence should be submitted or made available for regulatory review.
ICH M15 changes the way model-informed strategies need to be designed, governed and communicated.
For sponsors, regulatory strategy cannot be added retrospectively once the model has been completed. The question of interest, context of use, influence, risk, impact, and technical criteria should be defined while the development programme can still be adapted and the necessary data generated.
For quantitative pharmacology experts, statisticians and data scientists, evaluation no longer concerns predictive performance alone. Code verifiability, reproducibility, version control, robustness, sensitivity to assumptions, clinical relevance of metrics and documentation of deviations become central considerations.
For data owners and data management teams, decisions relating to source selection, exclusions, transformations, imputations and bias management become part of the regulatory evidence framework. Data provenance and traceability should make it possible to reconstruct the path from the original data to the result used in decision-making.
For quality and regulatory functions, there is a greater need to define responsibilities, controls, review procedures and document-retention rules. The conclusion regarding the appropriateness of the evidence is explicitly multidisciplinary and cannot be delegated solely to the team that developed the model.
For contract research organisations (CROs) and other specialised service providers, it becomes essential to ensure that data, code, software, configurations and methodological decisions are transferable, verifiable and available for review.
Finally, when model impact is high, early engagement with regulatory authorities becomes particularly important. EMA’s pilot procedure specifically targets programmes in which MIDD evidence is expected to play a substantial role and enables in-depth interaction between model-development experts and regulatory assessors.
ICH M15 does not automatically confer regulatory validity on a technology, platform, model or dataset. It requires that the data used, the models developed, the outcomes generated and the related evaluation processes be appropriate for the specific question of interest, transparently documented and evaluated with a level of rigour commensurate with risk.
Meeting this bar in practice tends to require a specific kind of organisational capability: one that connects data acquisition, harmonisation, synthetic data generation and AI/ML model development within a single, auditable pipeline, rather than treating these as separate, loosely connected services. This is the same structural logic ICH M15 asks for between question of interest, model role, and documented evaluation — the framework rewards continuity and traceability across the chain, not isolated technical excellence at one step.
In concrete terms, this capability spans acquiring data from heterogeneous sources and systems, mapping them to common structures, completing or rebalancing available information, generating synthetic cohorts, and evaluating them against analytical utility, potential bias and privacy protection. The resulting datasets support statistical analyses, simulations and model development, and can facilitate collaboration across organisations and countries while reducing the need to transfer personal data directly.
Several aspects of this capability map directly onto ICH M15’s requirements:
A causal-inference case study illustrates what this kind of end-to-end demonstration can look like. A synthetic cohort was generated from a real NHANES dataset examining statin use and all-cause mortality, and an identical observational causal-inference workflow — naive comparison, G-computation, IPTW, and TMLE — was run on both the real and the synthetic data. The synthetic cohort reproduced not only the marginal statistical distributions but the underlying confounding structure and the resulting causal conclusion at every analytical step, including the reversal of a confounded naive estimate. This is the kind of applicability evidence ICH M15 asks for: not a single similarity score, but preservation of the clinically relevant relationships a specific analytical question depends on.
The regulatory role of synthetic data and AI or ML models should, however, be determined on a case-by-case basis. A synthetic dataset used exclusively to verify the functioning of code has a very different influence on decision-making from a synthetic cohort that contributes to dose selection, population definition or the replacement of part of the clinical evidence.
The same distinction applies to models. An algorithm used for internal exploration, hypothesis generation or analytical prioritisation will generally have a different level of influence from a model used to predict a clinical outcome, identify a population, simulate a treatment or directly support a regulatory decision. As model influence and the consequences of a potential error increase, so should the rigour of the evaluation, the strength of the supporting evidence and the depth of the documentation.
Evaluation should therefore be commensurate with the actual role assigned to the data or model. When synthetic data, generative models or predictive models directly contribute to a decision, demonstrating overall statistical similarity or reporting a single performance metric is not sufficient. What matters is the preservation of clinically relevant relationships, outcomes, subgroups, distribution tails, and rare events. For AI and ML models, calibration, robustness to changes in data and assumptions, transportability to other populations, uncertainty in predictions, performance across subgroups, and the risk of systematic bias also need to be considered.
The relationship between synthetic data and model development also needs to be described precisely. Synthetic data may be used to train a model, supplement underrepresented populations, test the analytical process, or assess model behaviour under specific scenarios. However, a synthetic dataset generated from the same data used to develop or train the model cannot automatically be considered an independent source for external validation — even when it appears formally distinct from the original data. Independence of the information source and separation between development and validation need to be demonstrated, not assumed.
The combined use of real and synthetic data requires the provenance of the information to remain identifiable, the contribution of different sources to be distinguishable, and the data used at each stage of development, fine-tuning, testing and validation to be reconstructable. This transparency is essential for correctly interpreting model performance and for assessing its applicability to the specific context of use.
Formal quality management, information security and data protection systems — such as ISO 9001, ISO 27001 and Europrivacy certifications — together with infrastructure designed around traceability, documentation, human oversight and risk-proportionate controls, can provide process-level evidence relating to responsibility, control, security, reproducibility and verifiability. These elements do not, however, replace the evaluation of a specific dataset or model, nor do they automatically determine its regulatory acceptability.
ICH M15 represents a significant step in the evolution of model-informed drug development. The guideline does not introduce a universal checklist of requirements for declaring any model valid. Rather, it establishes a pathway for demonstrating that the data, model, outcomes and level of evaluation are appropriate for a specific decision.
The pathway is clear: define the question of interest; specify the context of use; assess model influence, consequences of wrong decision, model risk and model impact; prospectively define the technical criteria; verify and validate the model; document results, limitations and deviations; and integrate the conclusions through a multidisciplinary team.
For life sciences organisations, this means embedding the MIDD strategy within the overall governance of drug development and strengthening collaboration across clinical, quantitative, regulatory, technology, data management and quality functions.
The entry into effect of the European guideline and the launch of EMA’s pilot procedure show that the transition from principles to implementation is already underway. In this context, AI and synthetic data can play a strategic role only when embedded within governed, traceable and risk-proportionate processes. They are not a shortcut around regulatory standards, but tools that can contribute to generating more accessible, reproducible and useful evidence, provided their suitability for the specific intended use is demonstrated.
As the European Health Data Space moves from regulation into operational infrastructure, the capability that matters most will be the ability to connect data governance, synthetic data generation and evidence generation within a single, auditable technical framework — rather than offering these as separate, loosely connected services.

A rigorous, transparent, and regulatory-ready approach to accelerating clinical research with synthetic control arms

What the new guidance means for data quality, fitness for use, and the role of AI and synthetic data in regulatory readiness

Why synthetic data represents a technological key to enable the secondary use of clinical data in Italian healthcare and accelerate medical research today.
What we look for
Beyond your technical background, we look for:
What we offer
💡 Growth & Impact: join a fast-growing company where you’ll lead strategic projects, shape solutions and see the tangible impact of your work
🌴 Flexibility & Wellbeing: hybrid or fully remote work, ticket restaurant and health insurance
🤝 Collaborative Culture: work in autonomous teams with highly talented colleagues, in a supportive, innovative and ethical environment
How to apply
To apply, please send your CV and a motivation letter to [email protected], with the subject “Spontaneous Application - [Your Area of Expertise]”.
Aindo is an Equal Opportunity-Affirmative Action Employer – Minority / Female / Disability / Veteran / Gender Identity / Sexual Orientation / Age.