Abstract
Keywords
Medical decision making has shifted from a provider-centered model, in which doctors have the responsibility of making patient decisions, to a shared, patient-centered model in which patients are more involved in decisions about their medical care.1–4 Effective use of health statistics by patients is therefore becoming increasingly important and a prerequisite for informed decision making. However, communication and interpretation of these statistics are typically fraught with problems.5–7
A specific difficulty involved in the interpretation of medical statistics is the estimation of posterior probabilities, such as p(A|B). According to Bayes’ theorem,
In the case of medical diagnosis, “A” refers to the patient’s true disease status (has disease or does not have disease), and “B” refers to the patient’s test result (positive or negative). Two posterior probabilities are of particular interest in medicine: the positive predictive value (PPV), which is the probability that the patient has the disease given a positive test result, and the negative predictive value (NPV), which is the probability that the patient does not have the disease given a negative test result.
Gigerenzer and others 7 reported a study in which obstetric-gynecologic physicians were tested on their ability to estimate PPVs in the context of breast cancer screening. Most respondents vastly overestimated the PPV of a positive mammogram as being greater than 80%, when in fact it was 10%. This type of overestimation may result in overdiagnosis, potentially harmful follow-up diagnostics such as biopsies, and unnecessary escalation of patient stress. 7
In light of the difficulties associated with processing risks framed as probabilities, frequency-based risk formats have received extensive study. For example, Galesic and others 6 compared the effect of natural-frequency and conditional-probability formats on younger and older adults’ comprehension of medical test results. Natural frequencies (e.g., “50 out of 10,000 people have insulin-dependent diabetes”) elicited more accurate estimates than did conditional probabilities (e.g., “The probability that a person has insulin-dependent diabetes is 0.5%”). This benefit was observed for both younger and older adults, as well as those of high and low numeracy level. However, approximately 45% of participants in the natural-frequency condition still made significant estimation errors, suggesting that interpretation of test results remains challenging even when probabilities are expressed in terms of natural frequencies.
In addition to a reliance on natural frequencies, a great deal of research has focused on visualization tools such as icon arrays, whose effectiveness at improving risk comprehension may depend on the target group’s graph literacy and numeracy level.8–10 As reviewed next, a relatively less explored strategy for improving risk communication involves “experienced probabilities,” that is, exposure to probability distributions over time, rather than via static numerical or graphical displays.
Descriptive versus Experience-Based Risk Formats
Decisions from description are based on explicit summaries of a priori probability information. 11 Decisions from experience are based on statistical probabilities and may be influenced by recency, exploration-exploitation tradeoff, and other cognitive factors.12,13 An example serves to illustrate the distinction: One may base the decision to carry an umbrella on the probability of rain in the weather forecast (description) or on the rainfall pattern in recent days (experience). Within the risky-choice literature, the “description-experience gap”14–16 refers to the finding that people choose as if they overweight small described risks and underweight small experienced risks. Applied to the context of Bayesian inference in medical diagnosis, then, it is possible that the commonly observed overestimation of small posterior probabilities (e.g., the PPV of a diagnostic test) is a by-product of the descriptive format in which medical probabilities are typically communicated. 17
Although experience-based risk formats are increasingly being studied in the medical context,18,19 only 1 study to date has investigated how description- and experience-based risk formats affect Bayesian inference in the medical context. Fraenkel and colleagues 20 examined lung cancer risk comprehension and lung cancer screening preferences among patients attending an outpatient pulmonary practice. Patients were randomly assigned to 1 of 3 different formats representing risk information about low-dose computed tomography (LDCT) scans: 1) an experience format, involving a series of slides showing LDCT scans of 250 patients in random order, in addition to numerical information; 2) icon arrays and numbers; or 3) numbers only. Participants were tested for their objective knowledge of risk of lung cancer screening and their preference for screening. Contrary to the authors’ hypothesis, results showed that icon arrays accompanied by numbers produced more accurate knowledge than the experience format with numbers and the numbers-only format. In addition, those in the experience format with numbers showed an increased endorsement of screening (even though screening produces a relatively high rate of false-positives). Overall, these findings are thus not consistent with the idea that experienced probabilities necessarily improve risk comprehension or medical decisions among patients. However, the authors note that this may have been due to the complexity and number of scans presented at relatively high speed, which may have prevented effective encoding. In summary, existing evidence of the effectiveness of experienced risks in the medical context is modest and in need of replication.
Age Differences in Medical Decision Making
The study by Fraenkel and colleagues 20 included middle-aged and older adults but did not systematically examine the impact of age on risk comprehension or choice predisposition. In light of current demographic trends and increased rates of health problems among older adults, a closer examination of age differences is critical.
Normal aging is associated with cognitive slowing and with reductions in working memory, attentional control, and episodic long-term memory 21 —cognitive processes likely involved in Bayesian reasoning. There is also evidence of cohort differences in numeracy and statistical literacy,22,23 with older adults sometimes faring more poorly in these domains than younger adults. On the other hand, research on frequency processing24–26 has shown that encoding of frequency information is relatively automatic and that sensitivity to frequencies is largely preserved in old age. There is also some prior evidence of preserved sensitivity to statistical regularities 27 and relatively intact experience-based probability judgments in older adults. 28 Overall, the pattern of preserved and impaired abilities in aging suggests that frequency-based, experiential probability formats may be particularly well suited for enhancing risk comprehension in older adults.
The Current Study
The objective of the current study was to compare the impact of probability format (description v. experience) on Bayesian inference in younger and older adults. In the description condition, modeled after Galesic and others, 6 participants read a statistical summary of the relevant probabilities of a diagnostic test (i.e., base rate of disease, as well as conditional probabilities such as the sensitivity and false-alarm rate of the diagnostic test). However, in departure from Galesic et al., 6 disease names were fictitious to minimize the influence of prior knowledge—potentially different for younger and older adults—on performance. In the experience condition, modeled after Fraenkel and others, 20 participants viewed the joint distribution of health status and test result. Specifically, participants viewed a slideshow in which they saw a series of patient profiles. Individual patients were shown on the computer screen one at a time. Each patient was characterized by his or her health status (does or does not have disease) and test result (positive or negative test result). The group of patients was a representative sample of the population, in the sense that the relative frequencies of the different disease/test result combinations in the sample equaled the population probabilities. During the slideshow, participants thus had the opportunity to update their prior beliefs about the population on the basis of experienced instances.
We predicted that the experience format would result in more accurate estimates of posterior probabilities (PPV and NPV) in comparison to the description format. Furthermore, age differences in estimation accuracy were expected to be greater in the descriptive format compared with the experience format. This hypothesis was based on the evidence for age-related declines in aspects of fluid intelligence, including speed, working memory, and executive functions. 21 These factors have been shown to play a role in text comprehension 29 and numerical reasoning, 30 both of which likely affect processing of written passages such as those used in the current study and other similar work.6,18 In contrast, there is little support for a role of working memory in experience-based tasks.31,32 In addition to testing hypotheses about the impact of probability format on posterior probability estimates, we also aimed to examine whether probability format would have an impact on participants’ self-reported confidence as well as their preference to make their own medical decisions or to rely on a physician for making decisions for them.
Method
Participants
All participants gave written informed consent for the study, which was approved by the Research Ethics Board at Ryerson University in Toronto, Canada. The final sample included 80 younger adults (
Design
The study employed a 2 × 2 × 2 mixed design, with age group (younger v. older) and format (description v. experience) as between-subjects factors and disease type (polykronisia, zymbosis) as a within-subjects factor. Half of the participants in each age group were randomly assigned to either the descriptive format or the experience format, respectively. On each trial of the task, participants received information about a fictitious disease and its diagnostic test before providing a series of probability judgments. Two trials were administered, corresponding to 2 fictitious diseases, polykronisia and zymbosis (see Table 1). Trial order was random.
Disease Properties
Note: PPV = positive predictive value; NPV = negative predictive value; disease/positive test = number of patients with disease who get positive test result; frequency = number of patients (out of a total of 100) representing each combination of disease status (has disease/does not have disease) and test result (negative/positive).
Stimuli and Apparatus
E-Prime 2.0 (Psychology Software Tools, Inc.) was used for stimulus presentation and response collection on a 16.0-inch LCD display running 32-bit Windows 7 Enterprise Edition. Viewing distance was approximately 50 cm. Text instructions for the descriptive format appeared in black against a white background, and task stimuli for the experience format appeared in red and blue 18-point Times New Roman font against a white background.
Procedure
In the description format, participants read passages that included information about the prevalence of the disease as well as the sensitivity and the false-positive rate of the test (see the supplemental appendix). All probabilities were presented in natural frequency format (e.g., “2 out of every 100 people”), as frequencies are easier to understand than probabilities. 6 Participants had up to 7 min to read the statistical summary. Once time was up or they indicated their readiness, participants continued to the test phase. They were asked to estimate disease prevalence, sensitivity, false-positive rate, specificity, PPV, and NPV using a natural frequency response format for each response. Only the estimates of PPV and NPV are relevant to the objectives of the current study.
In the experience format (see Figure 1), participants viewed a series of slides showing a representative sample of 100 fictitious patients who had undergone a screening test for the disease. Patients were presented one at a time for 3 s, separated by a 1-s blank white screen. Including an introductory screen, the slide show of 100 patients lasted approximately 7 min, making the overall exposure time similar to the maximum time available to participants in the description format. Patients were characterized by a combination of disease status (disease or no disease) and test result (positive or negative). The words “Has Disease” and “Positive Test Result” were presented in red font, whereas the words “Does Not Have Disease” and “Negative Test Result” were presented in blue font. The colored font served as an additional cue to increase the salience of the different combinations of disease status and test result. The number of patients representing the combinations of disease status (has disease v. does not have disease) and test result (positive v. negative) differed for the two diseases (see Table 1). Participants were prohibited from taking written notes during the slideshow and were discouraged from using rote memorization or other mnemonic techniques. Instead, they were instructed to simply pay attention to the information on the screen. The test phase was the same as in the description condition.

Schematic of the slideshow used in the experience format. Participants view a sequence of 100 patient cases representing combinations of disease status (has disease/does not have disease) and test result (negative/positive). The frequency of each combination is shown in Table 2.
Sample Characteristics
Note: MMSE = score on the Mini-Mental State Examination 33 ; numeracy = score on a scale that included the 11-item numeracy scale 34 and 1 coin-toss item 35 ; positive mood and negative mood: scores on the Positive and Negative Affect Schedule 36 ; depression, anxiety, and stress: scores on the Depression Anxiety Stress Scales (DASS-2137). Standard deviations are shown in parentheses.
Significant age difference (P < 0.01).
Significant format effect (P < 0.01).
After the computer task, participants completed a battery of background measures and cognitive tests, as well as a self-assessment questionnaire, which prompted participants to rate, using a 5-point scale, their level of confidence and comfort working with numbers, the difficulty of the estimation task, and their belief that their estimates were close to the correct answers. Participants were also asked to indicate on a 10-point scale whether they would prefer to rely on themselves (lowest value on scale: 1) or on a physician (highest value on scale: 10) in making final decisions about their medical care.
Data Analysis
In a first step, the PPV and NPV estimation errors were determined as the absolute (unsigned) difference between participants’ estimates and the corresponding correct values. For example, if the correct PPV was 20% and the participant’s estimate was 14%, then the estimation error was 6%. We used unsigned, rather than signed, errors in light of the restricted range of possible errors for PPVs and NPVs. We also ran the analyses on mean squared errors. To foreshadow, these analyses yielded the same pattern of results (direction and statistical significance levels of the main effects and interactions), so for simplicity, we report only the absolute errors.
Separate mixed analyses of variance (ANOVAs) were conducted on PPV and NPV estimation errors, with age group (younger v. older) and format (description v. experience) as between-subjects factors and disease (polykronisia v. zymbosis) as a within-subject factor. In light of violations of normality in the data, we also conducted Mann-Whitney U tests on the estimation errors. Each U test assessed differences between the levels of one factor (e.g., younger v. older adults), collapsing across the levels of the other factors. To foreshadow, the results of these pairwise nonparametric tests mirrored the main effects found in the ANOVAs and are therefore not reported separately.
Results
Participant Characteristics
Demographic, cognitive, and affective characteristics are shown in Table 2. A series of between-subjects ANOVAs on these characteristics, with factors age group (young v. old) and format (description v. experience), revealed several significant effects of age group, some significant effects of format, and no significant interactions.
The age effects largely mirrored typical patterns for cognitive and affective measures reported in the psychological literature on healthy aging.21,38 Even though older adults (
Unexpectedly, the ANOVAs also revealed 2 significant effects of format. Depression scores were higher for participants in the description condition (
Task Performance
PPV
There were no significant main effects of age or disease on PPV estimation errors (see Table 3). However, there was a significant effect of format, F(1, 156) = 131.02, P < 0.01, ηp2 = 0.46, indicating that estimation errors were significantly smaller in the experience format (
Posterior Probability Estimation Errors and Self-Assessment Responses
Note: PPV error = absolute difference between estimated and true positive predictive value; NPV error = absolute difference between estimated and true negative predictive value; confidence = rated confidence in working with numbers (on a scale of 1–5); difficulty = rated difficulty of the estimation task (on a scale of 1–5); belief in accuracy = belief that estimates were correct (on a scale of 1–5); self v. physician = preference for self- v. physician-made medical decisions (on a scale of 1–10).
Because absolute estimation errors do not indicate whether the errors reflect over- or underestimation, we also examined the distributions of raw (signed) PPV estimation errors (see panels A and B in Figure 2). The distributions showed a tendency to underestimate the true PPVs in the experience format and a tendency to overestimate PPVs in the description format. In addition, errors were tightly clustered around the true value in the experience format, whereas they ranged more widely in the description format, for both diseases.

Visualization of the distribution of estimation errors for both the positive predictive value (PPV) and the negative predictive value (NPV), shown separately for the two fictitious diseases (zymbosis, polykronisia). Zero on the x-axis indicates no error (i.e., the estimate was identical to the actual PPV or NPV). Negative values indicate underestimation of the actual PPV or NPV, and positive values indicate overestimation of the actual PPV or NPV. Two older adults in the description condition gave out-of-range estimates (e.g., 111/100), which were not included here.
NPV
There were no significant main effects of age or disease on NPV estimation errors nor any significant interactions (see Table 3). However, there was a significant effect of format, F(1, 156)= 23.31, P < 0.01, ηp2 = 0.13, indicating that estimation errors were significantly smaller in the experience format (
Self-assessment
ANOVAs on self-assessment responses (see Table 3) with factors age group (young v. old) and format (description v. experience) revealed significant main effects of format but no effects of age group and no significant interactions. Participants in the description format (
Influence of Numeracy
In light of prior reports of numeracy effects on Bayesian inference in the medical context, 6 we examined nonparametric correlations between numeracy and PPV and NPV estimation errors, separately for each participant group. None of the correlations reached significance, Ps > 0.05.
Discussion
The goal of this study was to assess the effect of different probability formats on estimates of the PPVs and NPVs of medical test results in healthy younger and older adults. We hypothesized that an experience format, involving sequential encoding of representative patient cases, would result in more accurate estimates of predictive values than a description format involving verbal summaries of relevant statistics. 6 Consistent with this hypothesis, we found a significant format effect on estimation errors, which were significantly smaller in the experience format, compared with the description format. Younger and older adults showed similar effects of probability format (and similar overall performance levels). There was no evidence for a relationship between numeracy and estimation accuracy. Finally, despite their superior performance on the estimation task and higher self-reported belief in the accuracy of their estimates, participants in the experience condition did not indicate a stronger preference for making their own medical decisions than participants in the description condition.
Description-Experience Gap in Bayesian Inference
The results in our description condition are in line with prior studies that have shown Bayesian inference to be difficult when relevant information is presented descriptively.6,7 A direct comparison between our results and those of Galesic and others 6 is not possible because the authors 6 did not report exact PPV estimates and did not assess NPV estimates. However, our results were qualitatively similar to those of Galesic and others, 6 as we observed significant estimation errors in the description condition, in both age groups.
The current results depart somewhat from those reported by Fraenkel and others, 20 who presented patients with information about lung cancer screening and found that a descriptive format (icon arrays) produced better comprehension and choice preference outcomes than an experience format. However, as noted by Fraenkel and others, 20 the slideshow employed in their study may have been too complex to permit effective encoding (250 slides at a rate of 1 s/slide, each slide featuring text, an unfamiliar CT scan, and 1 of 3 colors that represented patient disease status and test result). In contrast, in the current study, participants viewed 100 slides at a rate of 3 s/slide, and each slide included only 2 pieces of information (patient disease status and test result). It is possible that encoding of the relative frequencies of specific disease/diagnosis combinations in Fraenkel and others’ 20 experience condition was hindered by the fast presentation rate and the high attentional demands of the slideshow, which may have undermined effective encoding. In this sense, the experience format employed in the current study may have provided more “description” than that in the study by Fraenkel and others. 20
Consistent with observations in the risky-choice literature showing that rare outcomes affect choices as if they are overweighted in decisions from description, 14 participants in the description condition tended to overestimate the true PPVs. In contrast, participants in the experience condition tended to underestimate the true PPVs (although to a far lesser extent). However, it should be noted that we use the term estimation loosely. Posterior probability judgments can be made using strategies and heuristics (e.g., anchoring) that fall short of a strict definition of probability estimates.39,40 Future research should adopt a more fine-grained approach to shed light on the specific processes younger and older adults use to arrive at posterior probability judgments.
An interesting question is whether this description-experience gap in probabilistic inference affects subsequent attitudes and decisions. Responses to the self-assessment questions revealed a description-experience gap in participants’ confidence in their own estimates but not in their preference to rely on a physician. Future studies should examine whether subjective decision-making competence can be enhanced by providing participants with informative feedback following experience-based training trials.
The Role of Age and Numeracy
Contrary to our second hypothesis, we observed no significant age differences in the accuracy of estimated PPVs and NPVs. On the assumption that processing the verbal summaries in the description condition would tap cognitive abilities such as working memory and executive control, abilities known to decline with age, we had expected older adults to perform more poorly than younger adults in this condition. However, the null effect of age on Bayesian inference in the description condition mirrors prior findings by Galesic and colleagues, 6 who also found no significant age differences in predictive-value estimation from description, despite an age difference in numeracy. It should be noted that a common limitation in both studies is the reliance on convenience samples of healthy older adults with relatively high levels of education. More diverse samples may be needed to demonstrate age-related declines in the ability to process described probabilities.
Limitations and Future Directions
The current study had several limitations. First, despite random assignment of participants to the description and experience conditions, there were unexpected differences between these groups. Specifically, participants in the description condition scored higher on measures of depression and stress compared with those in the experience condition. Since the cognitive and affective measures were assessed at the end of the session, after the probability estimation task, the group differences in negative affect may have resulted from the experimental manipulations. However, without preexperiment baseline measures, this possibility could not be tested directly. In future studies, it would be important to assess pre- and posttask affect to examine whether exposure to described probabilistic information induces stress and negative mood. Given the detrimental effect of stress on cognitive function, including decision0making performance, 41 this question has obvious clinical relevance.
A second limitation concerns the differences between description and experience formats. Although the 2 format conditions were matched on many relevant aspects (e.g., response modalities used during the test phase), there were also several differences that may have affected the results. For example, the time spent reading the verbal summaries in the description condition was self-paced (with an upper limit), whereas the slideshow in the experience condition was experimenter paced. By necessity, there were also differences in the physical properties of the stimuli (verbal summaries and slideshows composed of words in varying colored font), and it is unclear which of these may have affected performance. In addition, participants in the description condition were presented with marginal (base-rate) and conditional probability information, whereas participants in the experience condition were presented with joint distributions (i.e., disease status and test result). Therefore, participants did not have access to the same information. To increase the similarity of information across conditions, it would be useful in future studies to add a 2 × 2 table with the disease status and test result in the description condition to present joint distributions that were also provided in the experience condition. Pinpointing the “active ingredient” underlying the format effect on Bayesian inference will require more fine-grained manipulations of encoding conditions in future studies.
Another limitation of the current study design involved the use of only 2 fictitious diseases. A deeper understanding of the mechanisms underlying the “experience advantage,” as well as its practical applicability, will require testing a broader range of denominators (e.g., 500 v. 100 patients in slideshow), disease prevalences, test sensitivities, and test specificities. To increase the ecological validity of the current findings, it will also be important to establish their replicability with real diseases and among patients facing actual medical decisions.
Conclusion
This study is the first to provide evidence for a significant increase in comprehension of medical test results (positive and negative predictive values) following exposure to experienced probabilities. It supports Hogarth and Soyer’s 17 suggestion that sequential observation of representative instances can serve as an attractive complement to description-based methods and that this approach holds promise for younger as well as older decision makers.
Footnotes
Acknowledgements
We thank Dr. Pete Wegier for his comments and suggestions throughout this project. We also thank Ryan Marinacci and Ryan S. Williams for their assistance with data entry and management.
Financial support for this study was provided in part by a grant from the Natural Sciences and Engineering Research Council (DG No. 358797 to JS), by the Canada Research Chair program (JS), and an Early Researcher Award from the Ontario Ministry of Research and Innovation (JS). The funding agreement ensured the authors’ independence in designing the study, interpreting the data, writing, and publishing the report.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
