Abstract
Objective:
When developing new medicines for children, the potential to extrapolate from adult data to reduce the experimental burden in children is well recognised. However, significant assumptions about the similarity of adults and children are needed for extrapolations to be biologically plausible. We reviewed the literature to identify statistical methods that could be used to optimise extrapolations in paediatric drug development programmes.
Methods:
Web of Science was used to identify papers proposing methods relevant for using data from a ‘source population’ to support inferences for a ‘target population’. Four key areas of methods development were targeted: paediatric clinical trials, trials extrapolating efficacy across ethnic groups or geographic regions, the use of historical data in contemporary clinical trials and using short-term endpoints to support inferences about long-term outcomes.
Results:
Searches identified 626 papers of which 52 met our inclusion criteria. From these we identified 102 methods comprising 58 Bayesian and 44 frequentist approaches. Most Bayesian methods (n = 54) sought to use existing data in the source population to create an informative prior distribution for a future clinical trial. Of these, 46 allowed the source data to be down-weighted to account for potential differences between populations. Bayesian and frequentist versions of methods were found for assessing whether key parameters of source and target populations are commensurate (n = 34). Fourteen frequentist methods synthesised data from different populations using a joint model or a weighted test statistic.
Conclusions:
Several methods were identified as potentially applicable to paediatric drug development. Methods which can accommodate a heterogeneous target population and which allow data from a source population to be down-weighted are preferred. Methods assessing the commensurability of parameters may be used to determine whether it is appropriate to pool data across age groups to estimate treatment effects.
Keywords
1 Introduction
Extrapolation has been defined as extending data and conclusions available from studies conducted in a ‘source population’ to make or support inferences for a ‘target population’. 1 Extrapolating from existing data, also commonly referred to as bridging or borrowing strength, is common in drug development. Examples include incorporating historical data into the analysis of contemporary clinical trials2–4 and, more controversially, using information on a drug’s short-term effect to draw conclusions about its long-term effect. 5 Alternatively, one may seek to test the efficacy of a medicine in a new geographic region when data are available confirming it is beneficial for patients from another locality. In such cases, it may suffice to conduct a smaller ‘bridging’ study in the new region that will collect efficacy and safety data to support the extrapolation of data from other localities to this site. 6
For extrapolations to be appropriate, source and target populations should be similar in terms of the key parameter(s) of interest. Extrapolations are ‘complete’, in the sense that existing data obviate the need to collect data from the target population, when there is strong prior opinion that differences between populations are small. Such opinion may be informed by pre-clinical work or experiences of developing related drugs or treating related patient groups. When there is greater uncertainty about the biological plausibility of similarities, ‘partial’ extrapolations may be more acceptable. A partial strategy would stipulate that existing data in the source population be complemented by supportive data in the target population generated by a reduced drug development programme. This reduced programme would be targeted to fill in gaps in existing knowledge or to verify similarities about which there is most uncertainty. To illustrate how an extrapolation strategy might be selected, suppose that data from the standard of care arm of several historical trials are available to inform the design and analysis of a new study. If investigators are confident that the standard of care has changed little over time and response rates have been stable, the historical data may be used as the control arm of the new (single-arm) trial. Otherwise, the historical data may be used to augment data from the new study, which would be designed as a randomised controlled trial (RCT) but would allocate fewer patients to control. Making full use of existing data can have important implications for the efficiency and feasibility of drug development in difficult to study populations such as rare diseases or groups where there are ethical and practical barriers to trial recruitment.
The use of extrapolation to facilitate the development of safe and effective medicines for children has received much attention.7–10 Adult data are often available at the time development of a new medicine begins in children. Moreover, trials in children can be more challenging to conduct due to practical constraints on available sample sizes and pharmacokinetic sampling.
11
There is also a common perception that recruitment into paediatric trials will be challenging, although this has been contradicted by recent research finding that parents and practitioners are willing to enter children into trials.
12
Dunne et al.
7
discuss the paediatric study decision tree8,10 shown in Figure 1, which is an algorithmic approach to determining which additional data are needed in children to support paediatric licensing decisions. The level of extrapolation is determined by whether adults and children can be assumed to be similar in terms of key characteristics, such as disease progression and the pharmacokinetic–pharmacodynamic (PK–PD) relationship of the drug. While this framework clearly identifies scenarios in which different extrapolation strategies are appropriate, it neither accommodates uncertainty about extrapolation assumptions nor allows for differences between age groups of children. To capture the heterogeneity of growth, development and pharmacokinetics in the population, the ICH E11 guideline
10
suggests one possible age grouping: preterm newborn infants, term newborn infants (0–27 days), infants and toddlers (28 days to 23 months), children (2–11 years) and adolescents (12–16/18 years, dependent on region). Batchelor and Marriott
13
state that there may be age-related changes in drug pharmacokinetics caused by anatomical and physiological differences between younger and older children and adults. However, Stephenson
14
notes that adults’ and children’s responses to many drugs have much in common. The European Medicines Agency (EMA)
1
has proposed a general framework for extrapolation allowing for the incorporation of uncertainty about assumptions. This framework stipulates that an extrapolation concept, containing explicit hypotheses on expected differences between populations, should inform the development of an extrapolation plan. This plan will detail which additional data will be generated in the target population, and these data should, in turn, be used to verify the extrapolation concept.
Paediatric study decision tree: image reproduced from Food and Drug Administration.
8

This paper describes the findings of a systematic review conducted to identify statistical methods that can be used to optimise extrapolations in paediatric drug development. We sought methods relevant for using data from a source population to support inferences for a target population. To provide focus for the literature search, we restricted our attention to publications developing methods in the context of four applications in which extrapolations are common, namely, paediatric clinical trials, trials extrapolating efficacy across ethnic groups or geographic regions, the use of historical data in contemporary clinical trials and the use of short-term endpoints to support inferences about long-term outcomes. The rest of the paper proceeds as follows. Section 2 outlines the strategy used to identify relevant papers and methods which are briefly summarised in Section 3. In Section 4, we give a detailed account of the methods found, grouped according to four common approaches. We conclude in Section 5 with a discussion of the suitability of these methods for making extrapolations in paediatric drug development.
2 Methods
Articles were identified by searching the Science Citation Index Expanded (SCI-EXPANDED) database of the Web of Science. Searches were restricted to English language papers listed on Web of Science prior to 31st January 2014 in the following categories: biology, mathematical and computational biology, mathematics (applied, interdisciplinary applications), medical informatics, research and experimental medicine, pediatrics, and statistics and probability. Preliminary searches were also made of other databases (JSTOR, PubMed) but no additional relevant articles were found. Separate searches of the SCI-EXPANDED database were made to identify potentially relevant papers proposing statistical methods for: (a) incorporating historical data into contemporary clinical trials; (b) using data on short-term endpoints to support inferences on long-term outcomes; (c) paediatric clinical trials and (d) bridging clinical trials. Since there was considerable overlap between the search terms needed to identify papers on the last two topics, these were combined so that a total of three separate searches were made. Search terms can be found in the web-based materials accompanying this manuscript (Supplementary Appendix A). We searched for papers containing these search terms either in the title, abstract or keywords.
Articles identified using this search strategy were then screened, first by title and then by abstract. At each stage the following types of manuscripts were excluded: (a) conference proceedings, (b) reports of clinical trials, (c) reports of meta-analyses or evidence synthesis analyses and (d) papers unrelated to medical statistics (returned because one search term, ‘bridge’, occurs in many contexts). A full text review of the remaining articles was then performed. At this stage manuscripts were excluded if they did not consider statistical methods, if they used source population data only to inform the design of a future trial or if they considered trials using a historical control arm without consideration of possible differences between populations. From each paper we extracted details of all statistical methods relevant for extrapolating data from a source population to support inferences for a target population. Methods for establishing whether data from source and target populations are consistent were regarded as relevant, assuming that if commensurability is established it would be appropriate to analyse data pooled across populations. A data extraction form (Supplementary Appendix B) was completed for each statistical method and the number of methods extracted from each paper was recorded. When identical methods were found in more than one paper, we recorded the method as it appeared in the earliest publication. Papers presenting only duplicate methods were excluded from the review. Data were extracted by one author (IW) seeking guidance from others (LVH, TJ) where necessary.
3 Results
Searches identified 52 papers satisfying the stated inclusion/exclusion criteria as summarised in Figure 2, from which we extracted 102 methods. A single method was extracted from each of 34 papers. Of the remaining papers, eight presented two methods each, while 10 presented three or more methods each.
Flow diagram of systematic review results.
Methods can be categorised into four main areas: (i) paediatric drug development (five of 102 methods), (ii) use of historical data in contemporary clinical trials (48 of 102), (iii) bridging trials extrapolating efficacy data between ethnic groups or geographic regions (43 of 102) and (iv) the use of short-term data to support inferences on long-term outcomes (six of 102). This is displayed in Figure 3. All five methods in category (i) considered extrapolating information from an adult source population to support inferences about children. Of the 48 methods in category (ii), 25 sought to extrapolate from a historical control group to support conclusions about control response rates in a contemporary patient group. Of the 43 methods in category (iii), 14 took as the target population an unstudied patient group in a new geographic region and sought to borrow strength from existing data on patients in another geographic region for whom the treatment had already been shown to be efficacious. One further method in this category evaluated the consistency of data in two ethnic groups of patients. The remaining 28 methods in category (iii) were proposed to assess the consistency of treatment effects across regions of a multi-regional clinical trial (MRCT).
Plot showing distribution of methods across four main areas.
Of the 102 methods, 100 expected data from the source and target populations to make inferences about key parameters in the latter group, and as such are appropriate for making partial extrapolations. An example of a method that did not expect data from the target population, Nedelman et al. 15 suggest that a necessary condition for using adult efficacy data to support conclusions about the efficacy of oxcarbazepine as a monotherapy for children with epilepsy, is that PK–PD relationships should be similar in adults and children receiving oxcarbazepine as an add-on therapy.
None of the methods found considered extrapolating safety data across populations. Instead all methods expected either efficacy or PD data (100 of 102) or PK data (two of 102). In the context of paediatric drug development, this may be due to the fact that the paediatric study decision tree stipulates that safety data must be collected in children regardless of one’s confidence in extrapolation assumptions. Most methods (100 of 102) sought to make comparisons between treatments while two methods were proposed in the context of dose-finding trials.
4 Thematic analysis of methods for extrapolation
Methods were first classified according to the type of statistics used, that is Bayesian or frequentist statistics. Categories were then refined to form three broad groups of approaches, namely Bayesian methods using existing data to create an informative prior distribution for a parameter of a target population, Bayesian and frequentist methods assessing the commensurability of parameters of source and target populations, frequentist methods synthesising data across populations using a joint model or weighted test statistic. Further details of the extrapolation methods are given below.
In all descriptions of methods, we will index parameters and data from the source (target) population by a subscript S (T). Therefore, xS (xT) will denote data from a source (target) population which depends on an unknown parameter θS (θT). When θS and θT are assumed equal, we will refer to their common value as θ. When several datasets are available from a source population, we will let H denote the total number of datasets available and nhS denote the size of dataset h,
4.1 Bayesian methods
Searches identified 58 Bayesian methods from 25 papers.2–4,16–37 Of these, 54 methods2–4,16–33 sought to create an informative prior for θT while four34–37 assessed the consistency of treatment effects or PK responses between the source and target populations.
4.1.1 Using existing data in a source population to create a prior for θT
All methods in this category sought to augment data from a future trial in the target population (xT) with existing data from one or more studies in the source population (xS). For example, θT and θS could be response rates on the standard of care available to patients in a new and historical trial, respectively. In this setting, differences between θT and θS may arise due to differences between trial protocols, advances in medical care or demographic shifts in the patient population over time. More generally, the source data will be useful for learning about θT only if the clinical effects of treatments in the source and target populations patients are similar. Of the 54 methods which used xS to create an informative prior for θT, most proposed discounting these data to account for potential differences. Thirty-one methods2–4,16–23 considered differences between θT and θS, and formulated priors for θT which when updated with emerging data from the new trial adaptively weight xS according to the commensurability of xS and xT. Fifteen methods adopted a fixed non-adaptive approach to down-weight xS. Eight methods did not down-weight xS at all, so that the final posterior distribution for θT would attribute equal weight to the source and target population data.
Most approaches in this category were proposed for incorporating data from a historical trial into a contemporary study. One approach which has received much attention is the power prior and 10 variations on this were found.2,3,16,17 Power priors are formed by raising the likelihood of the historical data to a power
It has been noted that the hierarchical power prior in equation (1) violates the likelihood principle since it omits the normalising constant for a0.16,38 Modifying equation (1) to incorporate the normalising constant
A similar Bayesian model for xS and xT is assumed to derive the commensurate prior (CP) for θT.
3
Again modelling conditional prior opinion on θT as
Hobbs et al.
3
adapt the CP in equation (3) for the case of normally distributed data to propose a location commensurate prior (LCP), assuming historical patient responses have mean μS and variance
Meta-analytic predictive (MAP) priors are an approach to combining data across several heterogeneous source populations to formulate an informative prior for θT. The use of historical control data potentially allows for the randomisation of fewer contemporary patients to control in a future RCT. Two methods developed this approach,4,20 synthesising data from the control arms of several historical trials in a Bayesian random-effects meta-analysis to derive the posterior predictive distribution for the parameter of interest in the control group of a new study. The MAP prior is then updated using Bayes theorem when data from the new trial become available. We classify methods4,20 as adaptive approaches to down-weighting data from the source population since the MAP prior can be approximated as a mixture of conjugate distributions78 and have heavier tails than a simple conjugate prior. Thus, in the event of a prior-data conflict, the historical data will eventually be discarded from the posterior analysis of the new trial.
When deriving the MAP prior, meta-analytic models are formulated assuming parameters of the historical and contemporary datasets are exchangeable. Suppose there are H historical trials generating estimates
Cuffe
21
considers a new RCT extrapolating from a single historical study to support inferences for the expected response on control. Responses from nS (historical) and nT (contemporary) control patients are summarised by the sample means xS and xT, respectively. These statistics are assumed to follow a Bayesian random-effects model
Mixture priors are another approach for using existing data to create an informative prior distribution for θT. Two methods22,23 use mixture priors to augment data from a future clinical trial in a new geographic region with data, xS, from an area that has previously been studied. These methods set the prior for the treatment effect in the new region as
Fifteen methods2,3,23–31 used existing data from a source population to formulate an informative prior for θT, down-weighting these data in a non-adaptive, pre-specified manner. The power prior can be considered in this category if a0 in equation (1) is taken to be a fixed constant and Hobbs et al. 3 refer to this approach as the conditional power prior (CPP). Six methods2,3,23–25 propose power priors with fixed a0. Ibrahim and Chen 2 propose a variation on this approach for the case that historical data are from a single trial and patient responses follow an arbitrary regression model. Neither paper discusses how to choose a0.2,3 De Santis 24 defines a geometric prior, raising the likelihood of data from a single historical trial to a power a0 = r/nS, where r is a constant specified by the analyst. The author also modifies this approach to weight different historical datasets by different fractions when they differ in their relevance to the new trial. De Santis 24 illustrates how the geometric prior can be used to inform early stopping decisions in a new Bayesian clinical trial. Rietbergen et al. 25 consider the CPP incorporating data from several historical studies, assigning data from each study a weight elicited from expert opinion. Gandhi et al. 23 consider the CPP for the purposes of incorporating existing binary data from a geographic region in which a drug has been shown to be effective into the analysis of a bridging trial conducted in a new region. The authors recommend performing sensitivity analyses to explore the impact on inferences of different choices of weights. Hobbs et al. 3 also provide a variation on the CP described in the previous subsection which treats τ as fixed.
Schoenfeld et al.
26
augment data from a clinical trial in children with data from a completed adult trial, assuming parameters of adult and paediatric data are samples from a normal population distribution with mean
Chen et al.
27
derive a Bayesian empirical prior distribution for a treatment effect θT in a specific local region of a MRCT which borrows strength from data from other trial sites. The prior
Six other methods in this category of approach shift the location and/or inflate the standard error of an estimate of θS to create an informative prior for θT while discounting the source population data.28–31 For example, French et al. 30 formulate a normal prior distribution for θT with mean equal to the MLE of θS obtained from xS, and standard deviation equal to four times the standard error of the MLE; the authors propose using this prior for the Bayesian interim monitoring of a trial which will terminate with a conventional frequentist analysis. Whitehead et al. 31 consider Bayesian sample size calculations, using historical placebo data to create an informative prior distribution for the expected response on placebo in the new trial. This prior is normally distributed, with the mean taken to be the mean response from the historical placebo group and precision chosen to reflect how many patients the prior should represent.
Eight methods23,29,30,32,33 used data from a source population to create an informative prior distribution for θT without any down-weighting. Thus, once available, data from the target and source populations are pooled to derive a posterior distribution for θT.
4.1.2 Assessing consistency between source and target populations
Four Bayesian methods were proposed to assess the consistency of parameters in source and target populations.34–37 Pei and Hughes
34
seek to assess whether candidate doses for adults and children result in similar percentages of patients experiencing low levels of a drug; inferences are made testing whether the proportion of children recording PK levels below a quantile estimated from adult data is non-inferior or equivalent to a design value. Tsou et al.
35
use Bayesian most plausible prediction
40
to assess the consistency of treatment effect estimates generated by a new clinical trial comparing an experimental treatment (E) with control (C) in a new geographic region, and reference studies which have demonstrated the advantage of E versus C in an original geographic region, under the assumption of normally distributed treatment effect estimates. The difference between treatment group sample means for the bridging trial,
4.2 Frequentist methods
Forty-four frequentist methods were identified15,34–36,41–66 of which 11 methods41–51 synthesised data from source and target populations in a joint model, three methods52–54 combined data across populations through a weighted test statistic and 30 methods15,34–36,48–51,55–66 proposed criteria to assess the consistency of estimates of key parameters in different populations.
4.2.1 Joint model incorporating data from source and target populations
Five methods41–45 proposed using short-term data to support inferences about a long-term endpoint assuming simple models to relate observations on different outcomes. In this setting, θT and θS could represent long- and short-term treatment effects, or characterise the distribution of the two endpoints. Several authors extrapolate from short-term data to inform early stopping decisions for sequential trials. Hampson and Jennison 41 seek to increase the efficiency of group sequential tests (GSTs) monitoring a long-term outcome by incorporating data on a correlated short-term endpoint so as to increase the Fisher information available for θT at each interim analysis. MLEs of θT are found maximising the joint likelihood of xS and xT assuming pairs of responses on the same patient follow a bivariate normal distribution. No assumption is made about the form of the relationship between the short- and long-term responses other than that they are correlated. The authors derive optimal designs and show that incorporating data on a highly correlated short-term endpoint can reduce the expected sample size of a trial by around 5% of the fixed sample size when the time to availability of the short-term endpoint is at least half that of the long-term endpoint. A similar problem is considered by Galbraith and Marschner, 42 who incorporate into GSTs repeated measurements of a continuous endpoint taken at an arbitrary number of follow-up times. The vector of repeated measurements for each individual is assumed to follow a multivariate normal distribution, with correlations between the measurements being exploited to improve estimation and inference associated with the long-term measurement. Marschner and Becker 43 increase the interim information available for a long-term response probability by incorporating data on a short-term binary endpoint, deriving the MLE of the long-term response rate from the joint likelihood of the combined dataset. The values of the short- and long-term endpoints may be associated; however, a patient’s short-term response does not necessarily determine their long-term response.
Stallard 44 uses observations on short- and long-term endpoints to support early stopping and treatment selection decisions in a seamless Phase II/III clinical trial. Responses on the same patient are assumed to follow a bivariate normal distribution, fitted using the double regression method of Engel and Walstra. 67 Wüst and Kieser 45 also consider bivariate normal outcomes and derive a more precise estimator of the variance of the long-term outcome incorporating short- and long-term data. Using this improved estimator to inform blinded sample size adjustments at an interim analysis reduces the variability of the final trial sample size when compared to using long-term data alone.
Six methods46–51 synthesise data from source and target populations using a frequentist random-effects model. Thall and Simon 46 combine historical and contemporary control data via a univariate random-effects meta-analysis while Arends et al. 47 model short-term and long-term outcomes from trials using a multivariate random effects model. Chen et al. 48 and Ko 49 use a random effects model to accommodate heterogeneity between regions and test for an overall treatment effect. Liu et al. 50 use a random effects model to test for similarity or non-inferiority between treatment effects in different regions. Ko 51 models survival data from different regions using a proportional hazards model with frailties to allow patients in different regions to have varying underlying hazards of experiencing an event.
4.2.2 Combining data across populations in a weighted test statistic
Three methods52–54 propose making final inferences about the efficacy of a new treatment in a new geographic region on the basis of a test statistic combining information from the source and target populations. Suppose ZT and ZS are standardised test statistics comparing mean responses on a new treatment and placebo in a new and original region, respectively. For reasonable sample sizes, ZT and ZS follow at least approximately standard normal distributions. Lan et al.
52
propose a weighted Z statistic for testing efficacy across regions,
4.2.3 Assessing the consistency of data from source and target populations
Thirty methods were proposed to assess the consistency of data from different populations. Chen et al.
55
survey nine methods in their systematic review for testing the commensurability of a treatment effect across regions of a MRCT, of which we extracted eight. These methods comprised ‘Global methods’ assessing consistency based on a test statistic combining data across all trial regions, ‘multivariate quantitative’ methods assessing consistency by considering all pairwise differences between region-specific effect estimates and ‘multivariate qualitative methods’ assessing whether patients from all trial regions can benefit from a new treatment. All eight methods assumed patient responses to be normally distributed. Let Δ
j
be the difference in mean response on treatments E and C in trial region j, for
One Global method is Cochran’s Q statistic
68
for testing the null hypothesis
Global test statistics can also be used to test for a qualitative interaction between the treatment effect and trial regions. The Gail–Simon test
74
of
Multivariate qualitative methods reviewed by Chen et al.
55
include testing
Hsiao et al. 65 propose two-stage designs for bridging trials. The trial begins recruiting patients from the original region. If efficacy in this region is confirmed at the interim analysis, the trial proceeds to recruit patients from the new region in Stage 2. Otherwise the trial terminates early for lack of benefit. On conclusion of the trial, data accumulated from both regions are pooled and analysed to test a one-sided null hypothesis of no treatment effect. If the result of Stage 1 is similar to the pooled result of Stage 2, the result from the new region is declared consistent with that from the original region and we conclude that the new treatment is effective in both localities.
Cai et al. 66 propose evaluating the similarity of data from clinical trials performed in different ethnic populations using a ‘distribution adjusted mean’. This method assumes that there is a covariate Y prognostic for the primary endpoint which differs in distribution between the two ethnic groups. If Y is continuous, its domain can be partitioned into intervals and the relative frequency of each interval in the target population is recorded. These frequencies are then used to calculate the weighted average response in the source population, averaging across the mean responses for each interval of Y. This adjusted mean response is then compared with the unadjusted mean for the target population to assess the consistency of response between the populations.
Nedelman et al. 15 develop a method comparing children and adults receiving a new drug as an add-on therapy, with the aim of using these data to support inferences about children receiving the drug as monotherapy. If the PK–efficacy relationship is similar for adults and children receiving add-on therapy, this is taken to support an assumption of similar relationships for adults and children receiving monotherapy. Separate linear models are fitted to the PK–efficacy data from adults and children, and model parameters are compared to establish whether there are differences between age groups.
Chow et al. 36 apply the ‘reproducibility probability’ method 77 to bridging studies, calculating the reproducibility probability as the power of the bridging study to detect a treatment effect equal to the estimated effect from the reference study which itself produced a significant result. If the reproducibility probability exceeds a critical value (determined by a regulatory agency) then the bridging study may be considered unnecessary, that is clinical data from the original region can be completely extrapolated to the new region to support claims of efficacy.
5 Discussion
This systematic review summarises statistical methods relevant for extrapolating data from a source population to a target population and has captured a wide range of methodology. Several of the approaches identified are potentially applicable for making extrapolations to support paediatric drug development. In this context, adult data, pre-clinical data and data on children receiving treatment for related conditions may all be available at the time development of a medicine begins in children. Thus, methods which can harness existing data to derive informative prior distributions for key parameters in children are particularly appealing. However, we speculate that down-weighting existing data would be more acceptable in this setting to account for potential differences between populations. Therefore, the applicability of those eight methods which give comparable weight to historical and contemporary data is likely to be limited unless there is a strong prior rationale for similarities. Alternatively, the methods identified by this review for assessing the consistency of parameters of source and target populations may be used as objective criteria for determining when it is appropriate to pool data from adults and children, or indeed pool data across different age groups of children.
When there is some prior understanding of the factors that may explain differences between populations, a weight for the existing data may be pre-specified. Otherwise Bayesian approaches such as the power prior, CP, mixture prior or MAP prior, which adaptively down-weight existing data, may be preferred. One criticism that has been made of MAP priors is that the posterior predictive distribution for θT given historical data must be typically derived using Markov Chain Monte Carlo. Therefore, since the prior is not available analytically, it cannot be easily reproduced by others unless they have access to the historical data combined in the meta-analysis. To overcome this challenge, Schmidli et al. 78 propose representing the MAP prior as a mixture of a small number of conjugate prior distributions which can be easily recorded and shared.
In Section 1 it was noted that there may be differences between age groups of children. Twenty-five methods2,15,18,55,61,66 identified by this review can accommodate a heterogeneous target population because key parameters are taken to be parameters of (semi-)parametric models capable of adjusting for baseline demographics. Several methods proposing a joint model for data from the source and target populations assume only that data from different populations are correlated. However, this is unlikely to be the case for paediatric drug development when source and target data will typically be observations on different patients. In this case multivariate meta-analytic models, as used by Arends et al., 47 are potentially more relevant since they can capture correlations between parameters of different populations. Future research will consider tailoring these models to support extrapolations in paediatric trials.
Several papers were identified by our literature search which, although they did not contain statistical methods, are relevant for discussion. Manolis et al. 9 discuss the role of modelling and simulation in paediatric investigation plans (PIPs), which are documents pre-specifying what studies will be conducted to support development of a medicine for children. The authors review positive PIP opinions (summarising key elements of PIPs supported by the EMA) and find that population PK models are the most frequently referenced modelling approach, while exposure-response and dose-response models are rarely cited: modelling and simulation, when proposed, is typically used to support dose predictions, study optimisation and data analysis. Khalil and Läer 79 review physiologically based pharmacokinetic (PBPK) models as applied to paediatric drug development, where parameters of PBPK models for children may be extrapolated from another species or age group.
Other methods not included in the systematic review were found proposing ways for using data from a source population to support inferences for a target population. Reif et al. 80 fit a population PK model to data from an adult Phase I trial and use this model to design clinical trial simulations needed to devise a sparse PK sampling schedule for children. De Santis 81 consider using a design prior borrowing information from historical data to plan a clinical trial, for instance to inform sample size selections. Additionally, 12 methods included in the review17,20,26,31,35,37,46,50,52,54,63 use source data to inform the design (through sample size calculations) and analysis of a prospective trial in the target population. In addition, four methods19,24,30,65 use source and target data to inform mid-study adaptations to the study in the target population.
Software was available for few of the 102 methods identified by this review. Computer syntax was included in a main paper or accompanying supplementary material for nine methods20,25,30,31,33,34,37; code was stated as available upon request from the corresponding author of one method 60 ; syntax for another method 16 was included in a related commentary article. 82 The strategy used to identify available software is described in Supplementary Appendix C, while the results are listed in Supplementary Appendix D.
This systematic review has aimed to be a comprehensive overview of methods for extrapolation. However, one limitation is that we chose to focus our literature searches on the four application areas listed in Section 2 and by doing so may have missed other relevant methods. Another limitation is that one author extracted the data so independent reviews of all papers were not performed.
Footnotes
Acknowledgements
We thank two reviewers for their thoughtful and supportive comments which helped to improve the manuscript.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported in part by grants from the National Institute for Health Research (NIHR-RMOFS-2013-03-05, IW; NIHR-CDF-2010-03-32, TJ), the Medical Research Council (MR/J014079/1, LVH; MR/M013510/1, IW) and the Medical Research Council North-West Hub for Trials Methodology Research (MR/K025635/1).
Supplement Material
Supplementary material is available for this article online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
