Abstract
Decision-analytic models must often be informed using data that are only indirectly related to the main model parameters. The authors outline how to implement a Bayesian synthesis of diverse sources of evidence to calibrate the parameters of a complex model. A graphical model is built to represent how observed data are generated from statistical models with unknown parameters and how those parameters are related to quantities of interest for decision making. This forms the basis of an algorithm to estimate a posterior probability distribution, which represents the updated state of evidence for all unknowns given all data and prior beliefs. This process calibrates the quantities of interest against data and, at the same time, propagates all parameter uncertainties to the results used for decision making. To illustrate these methods, the authors demonstrate how a previously developed Markov model for the progression of human papillomavirus (HPV-16) infection was rebuilt in a Bayesian framework. Transition probabilities between states of disease severity are inferred indirectly from cross-sectional observations of prevalence of HPV-16 and HPV-16–related disease by age, cervical cancer incidence, and other published information. Previously, a discrete collection of plausible scenarios was identified but with no further indication of which of these are more plausible. Instead, the authors derive a Bayesian posterior distribution, in which scenarios are implicitly weighted according to how well they are supported by the data. In particular, we emphasize the appropriate choice of prior distributions and checking and comparison of fitted models.
Keywords
Introduction
Building a decision-analytic model to represent the history of disease and treatment usually involves choices that are based on uncertain information. The magnitude of this uncertainty can be quantified, but improperly quantifying uncertainty can lead to biased decisions as well as wrongly allocating resources for further research.1,2 Uncertainties in models are usually best characterized probabilistically: by parameterizing the model flexibly, 3 characterizing each parameter by a data-derived distribution, and simulating the resulting probabilistic model to produce a distribution for model outputs. One approach to this is described by the umbrella term of Bayesian evidence synthesis, which is a statistical framework for explicitly modeling several related and connected sources of data, and naturally incorporates the uncertainty in model parameters.
In particular, when the current evidence is weak or only indirectly related to the main parameters, it may not be straightforward to accurately represent it in the model. A crude approach to calibrating a model against indirect data is to informally adjust the parameters until predictions of key outcomes visually appear to fit observed data. For example, recent models of human papillomavirus (HPV) infections4–6 adjust parameters governing the transmission and natural history of HPV until models produce estimates of HPV prevalence that are similar to data. A more quantitative approach involves sampling model parameters from plausible ranges, comparing observed data with outputs from the model, and retaining a subset of parameters that satisfy some arbitrary standard of fit to the data, resulting in a range of scenarios, usually with no further indication of which are more plausible. This approach has been used, for example, in models of hepatitis C 7 and HPV.8–11 Vanni and others 12 reviewed the choices involved in model calibration, including the appropriate measure of fit to the data, the algorithm to search for the best-fitting parameters, the standard of fit required to deem a set of parameters plausible, and methods for weighting the retained scenarios in probabilistic sensitivity analysis. The uncertainty about these choices can substantially affect the model outputs.13,14
Calibration as Bayesian Evidence Synthesis
These uncertainties can be accounted for by considering model calibration as a problem of Bayesian evidence synthesis. The main features of such an approach are as follows.
Multiple indirectly related data sets are assumed to have been generated from probability distributions (“statistical models”) with parameters that are related to each other. Consider the simplified example in Figure 1. There are 2 data sets (A and B), each providing direct information about different parameters, which are known functions of the parameters of interest to decision making. This is an example of a directed acyclic graph, explained in more detail in “Building a Bayesian Model for Evidence Synthesis.”
This graph forms the basis for computing the posterior probability distribution of the parameters. This represents the updated state of evidence, from combining prior beliefs and observed data using Bayes’ theorem. We describe this process in more detail in “Building a Bayesian Model for Evidence Synthesis.”
The uncertainty about the parameters of interest, expressed through this distribution, is then propagated through the model to generate uncertainty distributions for the results required for policy making, such as life years gained, infections prevented, or expected costs (Figure 1).
The distributions of the results give appropriate weights to each parameter value, according to how much evidence there is in the data, thus providing a natural way to calibrate the model against the data.

Illustration of decision model calibration as a directed acyclic graph. “Nodes” at the start of an arrow (parents) are assumed to “generate” the nodes at the end of the arrow (children). Bold-bordered boxes represent observed data. The white boxes are quantities that are defined as explicit functions of other parameters. The dashed boxes represent quantities that cannot be defined this way and thus are given prior distributions (see “Directed Acyclic Graphs”).
This approach is variously referred to as multiparameter evidence synthesis,15,16 generalized evidence synthesis, 17 or comprehensive decision modeling. 18 This provides a theoretically grounded solution to the choices discussed by Vanni and others. 12 The measure of fit to the data is implicit in Bayes’ theorem, and there are standard methods for computing the posterior distribution. There is no need to define an arbitrary standard of fit, and calibration is unified with probabilistic sensitivity analysis. For example, Goubar et al. 19 and Presanis et al. 20 synthesized multiple survey data sets in this way to estimate the prevalence of HIV infection. Whyte and others 21 used these methods to calibrate the parameters of a decision model for colorectal cancer natural history, and in a model for transmission of HPV, Bogaards and others 22 used similar methods to estimate transmissibility and resistance to infection.
In “Building a Bayesian Model for Evidence Synthesis,” we outline the key steps of building a comprehensive Bayesian model for evidence synthesis, calibration, and expressing uncertainty, and we show how they can be applied to rebuilding a recently published model for the progression of HPV infections, 11 introduced in the next section. By improving the characterization of uncertainty in this model, we reflect more closely the extent of the current evidence and the potential value of further research. A more subtle benefit is that we reduce biases in the model outputs, since they are a nonlinear function of the uncertain inputs. 2 The posterior also tells us how well the data confirm or modify our prior beliefs about the input parameters. We emphasize careful assessment of the fit of the model, potential conflicts between different sources of evidence, and the appropriate choice and influence of the prior distribution. Finally, we discuss the strengths and challenges of this approach.
HPV Progression Model
Infection with HPV type 16 or 18 is associated with about 70% of cervical cancers. To evaluate the long-term benefits of cervical screening and vaccination against HPV, estimates of the natural history of HPV-related disease from initial infection to invasive cancer are required. A model has been developed 11 to estimate progression rates of HPV-related disease, through grades of cervical intraepithelial neoplasia (CIN) to cancer. This was combined with transmission and economic models to evaluate policies for HPV vaccination in the United Kingdom.10,23
For the purpose of our tutorial, we investigate only the progression component. A discrete-time Markov model is used to represent the natural history of a single HPV-16 infection. This has a monthly cycle and 9 states, illustrated in Figure 2. The parameters we aim to estimate are the 5 progression and 3 regression probabilities

Markov model for natural history of human papillomavirus 16 (HPV-16) infection—states of cervical neoplasia and permitted monthly transitions with associated probabilities. Progression to treatment, death, and hysterectomy is allowed from all states up to precancer. CIN, cervical intraepithelial neoplasia.
Human Papillomavirus (HPV) Model Parameters and Associated Data Sources
See Figure 3 for how these are connected in a graphical model. CIN, cervical intraepithelial neoplasia; NHS, National Health Service.
In Jit and others, 11 point estimates of the transition probabilities were calculated from data. The only expression of uncertainty was a discrete set of 54 alternative scenarios for the transition probabilities; otherwise, these and many other uncertain parameters were assumed to be fixed. Here we reimplement this model as a Bayesian evidence synthesis. The scenarios and the fixed parameters are replaced by uncertain parameters with smooth prior distributions, which are updated to posterior distributions conditionally on the data. This implicitly weights each of the scenarios by how well they are supported by the evidence and improves the characterization of uncertainty.
Building a Bayesian Model for Evidence Synthesis
Directed Acyclic Graphs
In a Bayesian evidence synthesis, the “model” consists of not only a mathematical representation of the disease and/or treatment process (Figure 2 in this example) but also the network of statistical models and relationships that connect the parameters of that process with observed data and prior information. The basis of this approach is to build a directed acyclic graph or graphical model to represent these relationships, shown in Figure 3 for the HPV example. Quantities at the start of an arrow are assumed to “generate” the quantities at the end of the arrow; therefore, we call these parents and children, respectively. There are 3 basic types of quantity (or “node”). Data generated from a statistical model, whose parameters are the node’s parents. These are shown as bold boxes in Figure 3. Unknown parameters defined as deterministic functions of other parameters. These are shown as white boxes in Figure 3. Parameters with no parents and thus no arrows directed toward them in the graph. These are shown as dashed boxes in Figure 3. In other words, while they may generate data or be used in the definitions of child parameters, they are not defined themselves as functions of further parameters. These must be given prior distributions representing beliefs about their plausible values, as explained further in “Specifying Priors for Parameters.” Or if uncertainty about them is negligible, they may be assumed to be constant, for example, mortality rates that are estimated from full-population data in the HPV application.

Directed graphical model for evidence synthesis to estimate transition probabilities between states of human papillomavirus 16 (HPV-16) infection. Notation as in Figure 1. CIN, cervical intraepithelial neoplasia.
Example. In the ARTISTIC data, there are
Example. The DNA test for HPV-16 used in ARTISTIC has 100% sensitivity but is not always clinically relevant due to cross-reactivity between HPV types. We therefore express the diagnosed prevalence
Once nodes have been defined, they can be used in the definitions of other uncertain quantities required by the model. In this example, the true HPV-16 prevalence

Prevalences of diagnosed human papillomavirus 16 (HPV-16) infection among women in ARTISTIC (2006) data and extrapolation to younger women—observed and model posterior 95% credible intervals. Since the specificity is more than 99%, this also illustrates the approximate trajectory of the true prevalence,
Including Indirect Data
Indirect data can be included by adding extra nodes and arrows in the graph to define how the data are generated from a statistical model and how the parameters of that model are related to other parameters.
Example. The transition probabilities between grades of HPV-16–related neoplasia are largely informed by age-specific counts
Thus, knowing the
The remainder of the graphical model for the HPV example is set out in detail in the online supplement. Briefly, the transition probabilities to diagnosed (squamous cell) cervical cancer are informed by observed incidences from 2004 cancer registry data, adjusted to represent HPV-16–related cases in the screened population. The transition probabilities between states of neoplasia are assumed to generate monthly CIN state prevalences and cancer incidences in a hypothetical open cohort of women infected with HPV-16. These prevalence and incidence parameters, after various adjustments, are assumed to generate the observed data. Thus, all evidence informing the model has been included in the graph as explicit data, parameters with prior distributions, or constant parameters.
Thus, while decision models are typically described as being “populated” with values derived from data, the Bayesian approach also makes the data and its analysis an integral part of the model.
Specifying Priors for Parameters
Prior distributions must be chosen for quantities with no parents in the graph, just as in standard probabilistic decision modeling. 2 These may be vague or based on substantive information but must represent our beliefs prior to observing the data included in the graph.
Vague priors
If there are sufficient data already in the model to give a precise estimate of a parameter, then a vague prior within plausible ranges is reasonable. With sufficient data, the exact choice of prior will not be influential, although sensitivity analysis is advisable if the choice is uncertain or suspected to be influential (see “Model Checking and Sensitivity Analysis”).
Example. The prevalences of cytological states
Parameters should be transformed to a natural scale before being given a prior, to enable beliefs to be expressed intuitively.
Example. Since the parameters

Prior distributions (thin lines) and posterior distributions (thicker strips with darkness proportional to probability density and 95% credible limits) for cytological screening accuracies, human papillomavirus 16 (HPV-16) test specificity, and parameters of the model for extrapolating HPV-16 prevalence. CIN, cervical intraepithelial neoplasia.
Informative priors
If there are no direct data to inform a parameter, informative priors could be derived from published literature or from expert beliefs, ideally formally elicited. 27 When the priors are updated to posteriors, we can assess how much the data confirm or modify our substantive beliefs.
Example. We enhance the original analysis of Jit and others
11
by incorporating prior knowledge about the transition probabilities between grades of neoplasia in the presence of HPV-16 (

Prior and posterior distributions of monthly transition probabilities between states of human papillomavirus 16 (HPV-16) infection (note some axis limits are extended) compared with scenarios from Jit and others. 11 The darkness of each strip is proportional to the posterior density, fading to white at zero density, with 95% credible limits. CIN, cervical intraepithelial neoplasia.
Example. For the sensitivities and specificities of cytological screening and the false-positive rate of the HPV-16 test,
Computing the Posterior Distribution
Once all quantities in the graph have been defined or given priors, the joint posterior distribution of all unknowns will be calculated, which may include quantities of direct interest to a decision maker, such as the incremental net benefit of an intervention. The decision then allows for the parameter uncertainty in all model inputs and includes the evidence from the calibration data. In other words, probabilistic sensitivity analysis and model calibration are performed simultaneously.
The graph shows how the underlying parameters of interest—in Figure 3, the transition probabilities
Markov chain Monte Carlo
Bayes’ theorem states that the joint posterior distribution (probability density) Initial values are chosen for all New values are then sampled from the full-conditional distributions
Through the distributions of the children (the likelihood), any data directly or indirectly generated by Each iteration consists of 1 sample from the full set of nodes
The BUGS language and software
The BUGS language
32
allows graphical models to be expressed in a text format. In brief, each node’s definition, as a deterministic or random function of its parent parameters, corresponds to a language statement. For example, the definitions of
The software then constructs the graphical model internally and chooses and implements appropriate methods to draw random numbers from the full-conditional distribution of each node. An extensive guide to Bayesian modeling using the BUGS language and software is given by Lunn et al., 29 and Welton et al. 16 give a practical guide to its use in health decision modeling. The full BUGS model code representing the definitions in this example is provided in an online supplement. For this application, we use the JAGS software for BUGS language interpretation and computation. 33 Implementation in WinBUGS or OpenBUGS 32 would have been equally feasible.
Model Checking and Sensitivity Analysis
Although a Bayesian graphical model is a natural framework for evidence synthesis, it can still involve many assumptions that should be questioned. A model can be assessed by checking and comparing the fit of model predictions to the data used to build it,
34
then improving the model if necessary (“internal”
35
or “dependent”
36
validation). When more than 1 data set informs a quantity of interest, any potential inconsistency or conflict must be investigated.16,20 We recommend that any relevant data are included in the graphical model, rather than held back for external validation. Judgment is needed here, since modifying the model to accommodate increasingly less relevant evidence will make it increasingly cumbersome and prone to misspecification. If it is uncertain whether some data are relevant, perhaps due to differences in population characteristics or clinical practice, sensitivity analysis should be undertaken. When the evidence informing a particular part of the model is weak, so that different reasonable choices of prior distribution may affect the results, these should be compared in sensitivity analysis. Sensitivity analyses may also show the relative influence of the prior and the data on the conclusions. Relative fit between alternative Bayesian models can be compared with the posterior mean deviance (as in Presanis and others
20
) or the deviance information criterion (DIC), which is an estimate of the ability to predict a replicate data set.37,38
Note that the Bayesian approach does not eliminate the need for discrete sensitivity analyses for the parts of the model that are poorly informed by data, but the number of scenarios can be greatly reduced if these are replaced wherever possible by smooth priors, as in the HPV example.
In the next sections, we explain how these checks were carried out in the example and led to refinements of the model. In Figure 6, the prior and posterior distributions of the 5 state progression and regression probabilities are illustrated under the model as described, with 4 other assumptions explained below. This also shows the estimates from the scenarios considered in Jit and others
11
—note that there are some extreme scenarios that are implausible and thus have negligible weight in the posterior. The only substantial posterior correlation
Model Checking and Accommodating Conflicts
First, in the HPV example, we graphically compare the posterior distribution for the cytological state prevalences

Prevalences of cervical dysplasia among women in ARTISTIC (2006) data by human papillomavirus 16 (HPV-16) diagnosis—observed data and posterior distributions (median and 95% credible interval [CI]) under base case and conflicted models.
We accommodate this conflict by assuming that the odds of being in states CIN1 or higher for an HPV-16–negative woman in the ARTISTIC data is a constant multiplier of the corresponding odds in the national data. This constant is estimated as part of the model. This produces a much better fit in Figure 7 (model labeled base case). This supposes that women selected for ARTISTIC have slightly different prevalences of non–HPV-16–related neoplasia from the general population. The prevalences of HPV-16–related disease might also be different, but we assume these are the same given that the adjustment we have made produces an adequate fit. All subsequent results we present are from this model unless stated otherwise. Both the progression probability to CIN1 and the corresponding regression rate are changed after resolving this conflict (Figure 6). Other data sets are fitted reasonably well by the posterior distributions (not shown).
Note that the posterior distributions for
Sensitivity to Data Inclusion
There is also doubt about whether the national screening data are worth including at all, given that state prevalences by HPV-16 presence are observable from ARTISTIC alone, and HPV-16 presence is not recorded in the national data. Although including it may improve the precision of inferences from ARTISTIC alone, there is the risk of other implicit, conflicting information “feeding back” through the graph and affecting other parameter estimates, for example, the HPV-16 prevalences. We therefore compare the results under a third model (labeled no NHSCCSP data), where these data and the HPV-16–negative data from ARTISTIC (which are only required to understand the contribution of non–HPV-16–related dysplasia to the prevalences of dysplasia in the national data) are excluded. The estimated transition probabilities do not substantially change (see Figure 6). However, the posterior distribution of the maximum HPV-16 prevalence (Figure 4) is shifted upward by a couple percentage points, accounting better for the observations from around age 20 years. This suggests that the national data are feeding back to this part of the model.
Since we would expect the national screening program to better represent the distribution of states than ARTISTIC, which is a smaller, regionally biased sample of women consenting to participate, our base case results are those that do include the national data. In general, if the model is used to evaluate a policy for the population, any relevant full population data should be included, particularly since the highly selected populations typically included in randomized trials may not represent the patients of interest.
Sensitivity to Prior Assumptions
As well as the 2 alternative models that use the national screening data differently, we performed a further 2 sensitivity analyses. The first investigates how much the posterior distributions of the transition probabilities were affected by the strong priors compared with the data. We downweighted the contribution of the prior distributions by making them substantially weaker, arbitrarily multiplying their variances by 10. Note that completely vague Uniform(0,1) priors on all probabilities were not practicable, since they resulted in MCMC convergence failure, but more important, they would not have really represented our prior beliefs in this case—for instance, values over 0.5 for the monthly progression probability would be deemed highly unrealistic beforehand. Under the weaker priors, the posterior distributions of the CIN1→CIN2 and CIN2→CIN1 probabilities became more widely dispersed. The posterior means of the other parameters were unchanged, and the posterior variances were only slightly inflated, suggesting that the stronger prior merely adds precision to the study data rather than shifting estimates away from the data. The exception is the regression rate from CIN1 to (CIN-free) HPV-16, where the study data appear to support lower values of this probability than the prior. The base case model compromises between prior and data through Bayes’ theorem.
The other key uncertain parameters describe the accuracy of the cytological screening tests. The information about these came from a review of a very heterogeneous range of studies. Therefore, we performed another analysis in which all sensitivities and specificities are fixed at 100%, one of the extreme scenarios considered in Jit and others.
11
The posterior distributions for many of the transition probabilities (Figure 6) changed, confirming that they are sensitive to the screening accuracies. In particular, the posterior variances are smaller when these probabilities are fixed. The fit of this model is also much poorer than the base case, judging from DIC and plots (not shown). Information about the screening accuracies also comes, very indirectly, from the data (through Figure 3). This results in posteriors (Figure 5) that are reasonably precise compared with the diffuse priors and indicate the strength of evidence in the data for each value within the prior bounds. There is a moderate posterior correlation
Discussion
Bayesian graphical modeling is a useful framework for including all policy-relevant evidence in a decision model, even evidence that is only indirectly related to the main parameters. We have outlined the key steps of this approach and demonstrated how we used it to obtain posterior distributions for progression and regression probabilities of HPV-16–related cervical neoplasia. The posterior gives a more accurate reflection of the available evidence than the scenario analyses used in previous HPV models, by also including the evidence about how plausible each scenario was and reducing bias due to ignoring some parameter uncertainties. A single posterior distribution is also easier to interpret.
These methods are complex, but necessarily so in the HPV example, and only require a similar amount of programming to the original implementation. The principles of Bayesian statistical modeling are widely applicable in health policy evaluation.16,17 The BUGS language and software can provide the posterior distribution of any Bayesian model in principle, although models with very large numbers of unknowns, such as this one, require custom extensions to the software to be written in more low-level programming languages. Whyte and others 21 also describe a similar implementation of a Bayesian health economic model using Visual Basic within Excel.
The HPV model could be extended as in Jit and others10,23 to incorporate an infection and economic model to evaluate policies for HPV vaccination. This was originally based on more than a thousand alternative scenarios that met a certain standard of fit but with no measure of relative plausibility between them. A Bayesian reimplementation to formally characterize this uncertainty would be expected to have a similar computational burden to the original approach, as was the case for the progression-only component. However, this is likely to raise more uncertainties due to weak evidence (e.g., about transmission, infection, and natural clearance) and potential conflicts with other data sources in the model. It is therefore unclear how much benefit would be gained from greater statistical formality in this example. As a graphical model becomes more complex, there is greater potential for erroneous information from one part of the model to indirectly give bias in another part. There is ongoing research on controlling the propagation of information in graphs through restrictions variously termed “cutting feedback” 39 and “modularization.” 40
No statistical method can eliminate uncertainty in a model, since evidence is always limited. A decision model should synthesize all relevant evidence, but judgments must always be made about what evidence is sufficiently relevant or strong. Potential conflicts between 2 sources of data on the same quantity should be investigated and explained, leading to refinements in the model. 20 In general, for more complex models, careful validation and sensitivity analysis become more important (see “Model Checking and Sensitivity Analysis”). This was demonstrated in the HPV example, where, after examining plots of model fit, some model assumptions were relaxed to accommodate the population screening data.
Weak evidence about some component of a Bayesian model will result in sensitivity of the results to the prior distribution. Conversely, if prior sensitivity is detected, this indicates areas where stronger data or further research are required. In the HPV model, for example, there was substantial uncertainty around the accuracy of cervical screening. Although the prior that we used was based on the best available information, sensitivity analysis showed that these parameters were influential. Although a posterior distribution gives a better representation of the data than a range of scenarios, a limited number of sensitivity analyses are still useful to show how the results would be affected if the evidence were to change. If this model had been employed for decision making, formal “value of information” methods (see, e.g., Welton and others 41 ) could be used to predict for which parameters more research would give the greatest expected benefits.
Footnotes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
