The Comparator Is Written Before the First Patient
A single-arm trial does not remove the comparator. It relocates it into a natural-history study, registry, or external-control dataset, whose design may later determine the credibility of the treatment effect, the structure of the economic model, and the uncertainty carried into reimbursement negotiations.
The comparator has not disappeared
Rare-disease developers often describe a pivotal study as “single arm”, as though the absence of randomisation also removes the need for a comparator. It does not. The comparison is simply made elsewhere—against patients drawn from a natural-history study, a disease registry, historical records or a hybrid retrospective–prospective cohort.
This relocation highlights how early decisions matter. Recognizing the importance of decisions about cohort entry, observation timing, and severity classification can help researchers and stakeholders feel valued and underscore their critical role in ensuring credible study outcomes. These initial choices ultimately influence study results and price negotiations.
Much of the future value argument may therefore be fixed before the first treated patient enters the pivotal study.
Four evidence objects that should not be confused
The science becomes clearer when four related but distinct objects are separated:
| Evidence object | Its scientific job |
| Natural-history study | Describes disease phenotype, progression, prognostic factors, clinical milestones and burden over time. |
| External-control cohort | Selects from that evidence a population intended to represent what would have happened to the treated patients without the intervention. |
| Causal comparative analysis | Defines the estimand, aligns time zero and eligibility, addresses confounding and estimates a treatment effect. |
| HTA and economic evidence | Translates untreated progression, survival, quality of life, resource use and uncertainty into comparative value. |
A natural-history dataset does not become a credible comparator merely because it exists. It becomes one only when it can estimate the untreated outcome for patients sufficiently similar to those entering the pivotal study, from the same clinical decision point, using outcomes that are measured in a compatible way.
| The question to askNot “Do we have natural-history data?” but “Can these data estimate the untreated outcome for the patients in our pivotal study, from the moment at which they become eligible for treatment?” |
What the current evidence tells us
Germany: acceptance remains exceptional

A 2025 analysis of German benefit assessments examined 295 single-arm or externally controlled assessment cases. The G-BA accepted the evidence for an added-benefit conclusion in 36 cases (12.2%). Among 93 orphan-drug assessments based on single-arm evidence, the evidence was accepted in only nine (9.7%). The authors found that acceptance was associated with features such as dramatic effects, severe rare disease, lack of alternatives and paediatric indications, while the reasoning was not always explicit. [1]
The important scientific point is not simply that German evidentiary standards are high. It is that a single-arm dataset must do more than show change: it must support a clinically interpretable contrast at the level of the G-BA-defined population and outcome.
Libmeldy: severity can divide the benefit conclusion
The assessment of atidarsagene autotemcel (Libmeldy) illustrates why the untreated course must be stratified before the evidence is pooled. In the 2021 G-BA decision, presymptomatic children received a hint of major added benefit, while the early-symptomatic group received a hint of non-quantifiable added benefit. The product and broader evidence programme were the same; what changed was the disease stage and the ability of the comparative evidence to quantify effect within that stratum. [2]
The long-term clinical publication subsequently compared 39 treated patients with 49 untreated patients and reported outcomes separately for presymptomatic late-infantile, presymptomatic early-juvenile and early-symptomatic early-juvenile disease. That separation is not a presentational detail. It is central to the scientific interpretation of treatment effect. [3]
England: openness does not remove the causal problem
A review of 64 NICE submissions using real-world data to estimate relative treatment effects found that disease registries and electronic health records were common data sources and that more than one third of submissions used naïve or unadjusted comparisons. The authors called for clearer causal questions, prespecified analysis plans and designs that emulate the trial that would ideally have been conducted. [4]
The NICE assessment of sebelipase alfa for Wolman disease shows the other side of the problem. The company used data from 21 untreated people to inform the comparator. NICE accepted that the likely QALY gain justified the maximum severity weighting, but the committee still scrutinised the uncertainty in survival, utilities, extrapolation and the commercial arrangement. A dramatic untreated prognosis can make a weak comparison less decisive to the central conclusion; it does not make the comparison methodologically strong. [5]
France: the comparison should be anticipated, not retrofitted

HAS guidance states that when real-world data are intended to provide an external control for a non-randomised clinical trial, the comparison should be anticipated and scheduled in advance so that it forms part of a deductive evidential strategy. The same guidance highlights the need to plan the collection of patient-relevant outcomes, utility values and resource use where they will be required for assessment. [6]
A 2025 cross-country review of 175 oncology external-control analyses reached a similarly cautious conclusion: none was accepted without restrictions. The proportion not rejected was highest in England, lower in France and zero in the included German and Norwegian analyses. The recurring objections concerned data quality, heterogeneity, bias, study design and statistical methods. The disease area was oncology, but the methodological warning applies directly to rare-disease development. [7]
| The evidence is not saying “do not use external controls”It is saying that their credibility is designed into the programme early—or lost early. Statistical sophistication cannot compensate for incompatible patients, incompatible clocks or incompatible outcomes. |
How to design the comparator before protocol lock
1. Define the target trial and the estimand before selecting the database
Begin by describing the hypothetical randomised trial that would answer the question under ideal conditions. Define the treatment-eligible population, the untreated or standard-care strategy, the treatment decision point, the start and duration of follow-up, the outcomes, the intercurrent events and the treatment effect to be estimated.
This target-trial logic prevents the available registry from dictating the scientific question merely because certain variables happen to have been collected. A broad registry may remain valuable for epidemiology, budget impact and generalisability while still being unsuitable as the primary external control.
2. Align eligibility before attempting statistical adjustment
The pivotal and external populations should use the same operational definitions of diagnosis, genotype, age, disease stage, prior treatment, functional status, organ involvement and irreversible damage. The eligibility filter deserves more attention than the eventual length of the covariate list.
Matching, weighting and regression can address measured imbalance within an area of genuine overlap. They cannot manufacture comparability when treated and untreated patients occupy different stages of disease or enter follow-up for different clinical reasons.
3. Protect time zero
The index date is one of the most consequential elements of an external comparison. In the trial it may be treatment, enrolment or confirmation of eligibility. In the natural-history cohort it may be tempting to use diagnosis, first symptom, referral, registration or the first complete assessment. Those dates are not interchangeable.
The defensible approach is to identify the date on which each external patient would have met the pivotal study’s eligibility criteria. Baseline variables should be measured using information available at or before that point, and follow-up should begin from the same treatment decision point used in the treated cohort. This reduces lead-time, immortal-time and selection bias.
| Practical ruleA sophisticated adjustment model will not rescue an analysis whose clock starts at different stages of disease. |
4. Treat severity as a potential effect modifier
In many rare diseases, severity is not simply a baseline covariate. Presymptomatic, early-symptomatic and advanced disease may represent qualitatively different trajectories, as may rapidly progressing and attenuated genotypes. Treatment effect may depend on whether irreversible neurological, muscular or organ damage has already occurred.
Define clinically credible strata before the outcomes are known, and ensure that the natural-history study can describe prognosis, progression, major events, survival, quality of life and resource use within each stratum. A pooled average can be statistically stable and clinically misleading.
5. Use the same measurement architecture on both sides
An outcome is not harmonised merely because both datasets use the same clinical label. “Motor decline”, “loss of ambulation”, “cognitive impairment” and “disease progression” may differ according to instrument, threshold, assessor, timing and missing-data rules.
For every material endpoint, predefine:
the instrument, version, language and scoring rules;
the clinically meaningful threshold and event definition;
the assessment schedule and permissible visit window;
assessor training, adjudication and source-data requirements;
the handling of missed visits, death and other intercurrent events.
Recent work on rare-disease trial readiness similarly emphasises longitudinal phenotyping, clinically relevant outcomes, PROMs and the effects on patients and caregivers. [8]
6. Collect HTA-native outcomes while the untreated population is observable
Regulatory and HTA requirements overlap, but they are not identical. Clinical development may focus appropriately on survival, function, biomarkers and safety. The later economic model will also require untreated quality of life, resource use, care requirements and transitions between disease states.
The natural-history protocol should prospectively consider:
preference-based health-related quality of life using an instrument relevant to the intended jurisdictions;
caregiver quality of life and time where the burden is material and the methods can be justified;
hospital, intensive-care, outpatient, community and social-care use;
supportive treatments, equipment, home adaptations and informal care;
time to clinically and economically important state transitions.
The essential principle is to collect utilities across the untreated severity range before the relevant states disappear through mortality, treatment or changing standards of care. A small vignette study assembled after launch is usually a weaker substitute for direct, prospectively planned collection.
7. Preserve contemporaneity and define “untreated” precisely
Historical controls may have worse outcomes because diagnosis, screening, ventilation, nutritional support, infection control or specialist care have improved. “Untreated” rarely means receiving no care; it means receiving the relevant standard of supportive care without the investigational disease-modifying intervention.
Record calendar period, country, centre, diagnostic route, supportive treatment and changes in care. Where historical records provide long follow-up but lack modern measures, a hybrid retrospective–prospective design may be preferable. Such designs can combine duration and event history with contemporary structured assessments, but the two components must be examined for differences in prognosis, measurement and care. [9]
8. Pre-specify the analysis before the treated results are known
Complete and, where possible, register the protocol and statistical analysis plan before the pivotal results are available. Specify the primary external cohort, estimand, index date, severity strata, confounder-selection approach, adjustment method, missing-data strategy, censoring rules and sensitivity analyses.
This is not only a safeguard against selective analysis. It demonstrates that the comparator was built to answer a clinically meaningful question rather than selected because it produced a favourable result. Regulatory reviews of single-arm programmes repeatedly identify non-contemporaneous controls, subjective outcomes and imbalances in prognostic variables as central weaknesses. [10]
9. Stress-test the comparator before the pivotal protocol is locked
Run a blinded feasibility analysis and answer the following before enrolling the pivotal study:
How many external patients meet every pivotal eligibility criterion?
How many remain in each prespecified severity stratum?
Can a common index date be established without using future information?
Are the major outcomes measured in a genuinely compatible way?
Which prognostic variables are structurally missing rather than merely incomplete?
Is follow-up long enough to observe the events that will drive benefit and the economic model?
Are individual patient data and source-data verification available?
Is there sufficient overlap to support the planned adjustment method?
A registry containing 200 patients may yield an effective external control of 18 after eligibility, timing and measurement are aligned. Discovering that before protocol lock is useful. Discovering it after the pivotal study has reported is a structural failure of the evidence programme.
10. Seek regulatory and HTA advice while the design can still change
For designated orphan medicines, EMA protocol assistance provides product-specific scientific advice. Under the EU HTA Regulation, Joint Scientific Consultations allow developers to discuss the evidence plan while clinical studies are being designed. The questions should be concrete: Is the external-control design acceptable in principle? Are the strata and index date appropriate? Are the outcomes, utilities and sensitivity analyses fit for the intended comparative assessment? [11] [12]
The question should never be merely, “Do you accept real-world evidence?” Agencies assess a design and a decision problem, not an evidence category.
The minimum viable comparator dataset
| Domain | Minimum information | Why it matters |
| Diagnosis | Diagnostic criteria, genotype, route and date of diagnosis | Defines phenotype and diagnostic era |
| Timing | First symptom, eligibility date, index date and follow-up start | Creates a common time zero |
| Severity | Prespecified clinical, functional and prognostic strata | Prevents incompatible trajectories being pooled |
| Standard of care | Supportive treatments, procedures and changes over time | Defines what “untreated” means |
| Clinical outcomes | Same definitions, instruments and windows as the pivotal study | Supports a credible relative-effect estimate |
| Quality of life | Preference-based patient and justified caregiver measures | Populates QALYs and severity states |
| Resource use | Hospital, community, equipment and informal-care requirements | Populates cost and budget-impact models |
| Data provenance | Source, centre, period, completeness and validation | Demonstrates fitness for purpose |
| Analysis | Estimand, confounders, missingness, censoring and sensitivities | Limits bias and selective analysis |
Common failure modes
Collecting only what clinicians routinely record
Routine records can be clinically rich while omitting the exact eligibility variables, assessment windows and economic outcomes needed for comparative analysis.
Adding the comparator after the pivotal trial
The cohort and analysis may then appear—and sometimes become—selected in response to the observed treatment result.
Pooling every disease stage
A larger cohort can erase severity-dependent prognosis and treatment effect.
Starting follow-up at different clinical moments
No statistical method fully repairs a comparison in which treated patients start at treatment eligibility and controls start years earlier at diagnosis.
Measuring quality of life only after treatment
The untreated denominator then rests on mapping, small vignette studies or assumptions rather than direct evidence.
Treating missingness as a technical nuisance
Missing genotype, age at first symptom, baseline severity or standard-of-care history can make eligibility alignment impossible.
Assuming a dramatic outcome makes design irrelevant
A large effect can dominate the clinical conclusion, but uncertainty migrates into utilities, extrapolation, commercial arrangements and price.
The real strategic error
The greatest mistake is not choosing the wrong matching algorithm. It is allowing the natural-history study, pivotal trial and economic model to be designed as separate projects by separate teams.
Clinical development may collect what is required for registration. Epidemiologists may describe the disease broadly. Health economists may arrive later and discover that no preference-based quality-of-life data exist in the untreated population. Market-access teams may then attempt to construct a comparator from datasets that were never intended to support causal inference. At that point, the missing evidence is no longer a protocol amendment. It is a structural limitation of the launch.
The natural-history programme should therefore be designed jointly by clinical development, epidemiology, statistics, health economics, market access, clinicians and patient representatives. Each discipline sees a different part of the future decision. Recent HTA scholarship similarly argues that rare-disease evidence challenges extend across natural history, endpoints, comparative effectiveness, economic modelling and decision uncertainty rather than residing in a single methodological compartment. [13]
Conclusion
Rare-disease development will always involve uncertainty. Small populations, heterogeneous phenotypes and ethical constraints can make conventional randomisation impractical. No external-control method can recreate randomisation after the fact.
But uncertainty should not be confused with avoidable imprecision. A well-designed natural-history programme can establish a common treatment decision point, preserve meaningful severity strata, harmonise measurement, document changing standards of care and collect the outcomes required for both clinical and economic assessment.
You may negotiate reimbursement years after launch. But the comparator against which value is judged was probably written before your first patient was treated.
References and source links
All links below are direct DOI, journal or official agency links. No tracking or ChatGPT redirect links are used.
11. European Medicines Agency. Scientific advice and protocol assistance.
14. National Institute for Health and Care Excellence. NICE real-world evidence framework. ECD9.