Abstract
Background/aims
Considerable human and financial resources are typically spent to ensure that data collected for clinical trials are free from errors. We investigated the impact of random and systematic errors on the outcome of randomized clinical trials.
Methods
We used individual patient data relating to response endpoints of interest in two published randomized clinical trials, one in ophthalmology and one in oncology. These randomized clinical trials enrolled 1186 patients with age-related macular degeneration and 736 patients with metastatic colorectal cancer. The ophthalmology trial tested the benefit of pegaptanib for the treatment of age-related macular degeneration and identified a statistically significant treatment benefit, whereas the oncology trial assessed the benefit of adding cetuximab to a regimen of capecitabine, oxaliplatin, and bevacizumab for the treatment of metastatic colorectal cancer and failed to identify a statistically significant treatment difference. We simulated trial results by adding errors that were independent of the treatment group (random errors) and errors that favored one of the treatment groups (systematic errors). We added such errors to the data for the response endpoint of interest for increasing proportions of randomly selected patients.
Results
Random errors added to up to 50% of the cases produced only slightly inflated variance in the estimated treatment effect of both trials, with no qualitative change in the p-value. In contrast, systematic errors produced bias even for very small proportions of patients with added errors.
Conclusion
A substantial amount of random errors is required before appreciable effects on the outcome of randomized clinical trials are noted. In contrast, even a small amount of systematic errors can severely bias the estimated treatment effects. Therefore, resources devoted to randomized clinical trials should be spent primarily on minimizing sources of systematic errors which can bias the analyses, rather than on random errors which result only in a small loss in power.
Keywords
Introduction
The rising costs of clinical research have for some time been a matter of major concern.1,2 The clinical development of new therapies, especially drugs and medical devices, is a complex enterprise that cannot be re-engineered solely on the basis of costs. The safety and well-being of participating patients need to be carefully monitored. Equally importantly, clinical trials must yield reliable results, as these may impact the treatment of future patients. A great deal of effort and time is usually spent in making sure that the data collected in clinical trials are completely error-free. Among the specific procedures used to ensure freedom from errors are source data verification, double data entry, audits, and other data-management activities that aim at cleaning and validating the data. While these procedures make intuitive sense, the costs involved in ensuring 100% error-free data become exorbitant in large multicenter trials such as those that are typically required for the approval of new therapies. 3 These costs can only be justified, economically and ethically, if they are likely to have an impact on patient safety or the trial results. 4 Yet, the evidence base for the efficiency of data monitoring is relatively scarce. 5
Several authors have studied the impact of specific types of errors on the outcome of randomized clinical trials. For time-to-event endpoints, Korn and colleagues have investigated the impact of measurement error in the timing of events and concluded that such errors were unlikely to bias the hazard ratio between the treatments being compared. 6 It is not surprising, therefore, that the cost-effectiveness of blinded independent central review of progression-free survival in cancer trials has been questioned,7,8 except in situations for which the measurement of time to progression is likely to be biased rather than just subject to random error. 9 Moreover, a strategy has been proposed to decrease the costs and time associated with blinded independent central review by reducing the number of images reviewed in randomized trials with larger size and larger treatment effects. 10 As stressed by George and colleagues,11,12 one of the key strengths of randomized clinical trials, especially when large, is the robustness of their conclusions to minor deviations from ideal practice. Such is not the case for non-randomized trials that are carried out at earlier stages of clinical development, so in the remainder of this article, we focus exclusively on randomized trials.
Despite previous efforts by some authors, there is a general lack of appreciation and of quantification of the types and volumes of data errors which are likely to affect the results of randomized trials. We have carried out simulations aimed at quantifying the impact of hypothetical data errors on the outcome of actual clinical trials, two in ophthalmology 13 and one in oncology. 14 The two ophthalmology trials were analyzed and reported as one study, hereafter referred to as one trial. We chose these trials because they were of relatively similar size, but in different indications and with contrasting outcomes: the ophthalmology trial showed a highly significant benefit for the experimental treatment, 13 while the oncology trial showed no difference in response rates between the randomized treatments. 14
Methods
Ophthalmology trial
This trial accrued 1186 patients with age-related macular degeneration worldwide (NCT00321997). 13 The experimental agent was pegaptanib, a single-stranded nucleic acid that blocks the activity of vascular endothelial growth factor, a pro-angiogenic molecule implicated in the pathogenesis of the disease. Patients were randomized to receive intravitreous injections of pegaptanib (at a dose of 0.3, 1, or 3 mg) or sham injections (with a syringe applied on the surface of the eye to simulate the pressure of an injection) every 6 weeks over a period of 1 year. Treatment efficacy was assessed through the visual acuity of the patients, quantified by the number of letters correctly read on a standardized visual acuity chart read from a distance of 2 m. The primary efficacy endpoint in the trial and for our analysis was the proportion of patients losing fewer than 15 letters of visual acuity at 1 year. In our simulations, we considered the dose that was approved by the Food and Drug Administration (FDA) in this indication, that is, 0.3 mg of pegaptanib administered as 6-weekly intra-ocular injections. At this dose of pegaptanib, 70% of the 294 patients lost fewer than 15 letters from baseline to 1 year in the experimental group, versus 55% of the 296 patients in the control group (difference in proportions, 15%; 95% confidence interval (CI), 7.0% to 22.4%; p = 0.0001).
Oncology trial
This trial accrued 736 patients with metastatic colorectal cancer in The Netherlands (NCT00208546). 14 The experimental treatment concerned the addition of cetuximab, a monoclonal antibody against the epidermal growth factor receptor, to a standard regimen consisting of bevacizumab, a monoclonal antibody against the vascular endothelial growth factor, plus capecitabine and oxaliplatin chemotherapy. Patients were randomized to receive this standard regimen alone or combined with cetuximab. Although the primary efficacy endpoint in the trial was progression-free survival, the endpoint of interest in our simulation was the response rate, as defined by Response Evaluation Criteria in Solid Tumors (RECIST). 15 Briefly, the response rate was the proportion of patients in whom the baseline sum of the longest diameters of all target lesions (assessed using computed tomography scans every 9 weeks) was reduced at any time after baseline by more than 30%. Although the response rate is seldom used as primary endpoint in pivotal phase III trials in oncology, it is accepted for approval by regulatory agencies, 16 and it has been used as primary endpoint in nearly 40% of published phase III trials in some tumor types (e.g. breast cancer 17 ). Here, we have chosen to analyze the response rate in order to have a binary endpoint in both trials. In order to assess the impact of errors in tumor-size measurements on response rate, our simulations were carried out on the 506 patients who had a baseline and at least two tumor-size assessments during follow-up. Among these patients, the response rate was 56.6% for the 256 patients in the experimental group and 57.2% for the 250 patients in the control group (difference in proportions, −0.6%; 95% CI, −9.2% to 8.1%; p = not significant (NS)). Due to the restriction of the analysis to those 506 patients, these response rates differ from the ones reported originally (52.7% vs 50%; p = NS). 14
Simulations of trial results
We used individual patient data from both trials. In order to assess the impact of data errors on the results of these trials, we used the actual patient data from each trial and added errors to a variable proportion of patients. The proportion of patients to whom errors were added varied from 0% (the actual trial results, without added errors) to 50%. The errors were added to the measurements used to calculate the endpoint of interest in each trial. The errors were sampled from a normal distribution with a mean of zero and standard deviation equal to the standard deviation of the visual acuity measurements collected in the ophthalmology trial and to half of the patient-specific baseline tumor size in the oncology trial.
For each simulation scenario, we generated 1000 trial results, for the results not to be unduly affected by the play of chance. Specifically, 1000 simulations yielded a standard error of less than 2% on estimated proportions (such as the average treatment effect across all simulations). For each trial result, two sets of simulations were performed, one in which errors were added as described above (i.e. random errors), and the other in which errors of the same magnitude were added with their sign chosen to favor one of the randomized groups (i.e. systematic errors). In the ophthalmology trial, the systematic errors favored the sham group in order to assess how much error was required to eliminate the treatment effect observed originally, 13 while in the oncology trial, the errors added favored the experimental group in order to assess how much error was required to create a counterfactual treatment effect in comparison with the original results. 14
Outcomes of interest
For each simulated ophthalmology trial result, the treatment effect was estimated as the difference in the proportions of patients achieving the primary endpoint between the pegaptanib group and the sham group. A Cochran–Mantel–Haenszel test was used to compare these proportions, with stratification—as in the original analysis—for baseline visual acuity, type and size of macular lesion, and whether the patient had received prior photodynamic therapy. 13 For each simulated oncology trial result, the treatment effect was estimated as the difference in response rates between the cetuximab group and the control group. A chi-square test was used to compare response rates, as in the original analysis of the trial. 14
The median and interquartile range of the estimated treatment effects and of the p-values for each simulated trial result were used to assess whether the estimation of treatment effect was biased, and the extent to which statistical significance was impacted, by the addition of random and systematic errors to increasing proportions of patients in both trials.
Results
Random errors
When random errors were added to the endpoints of interest, there was little bias in the estimated treatment effect, both for the ophthalmology trial (Figure 1) and for the oncology trial (Figure 2). Moreover, the p-value did not change qualitatively in either trial, regardless of the proportion of patients with added errors (Figures 3 and 4). In the ophthalmology trial, the estimated treatment effect was reduced by less than 15% even when random errors were added to 50% of the patients (Figure 1); the median p-value remained well below statistical significance (p < 0.001 in all cases; Figure 3). In the oncology trial, the estimated treatment effect remained close to zero even when random errors were added to 50% of the patients (Figure 2); the median p-value remained non-statistically significant (p > 0.05 in all cases; Figure 4).

Estimated treatment effects of pegaptanib on the proportion of patients with age-related macular degeneration losing fewer than 15 letters of vision acuity at 1 year for increasing percentages of patients with errors (upper 5 boxplots, random errors; lower 5 boxplots, systematic errors). The boxplots consist of a box, representing the median value and interquartile range, and whiskers extending to the extreme values over 1000 simulated trials. The shaded area corresponds to the 95% confidence interval of the observed treatment effect. 13

Estimated treatment effects of adding cetuximab on the response rate of patients with advanced colorectal cancer for increasing percentages of patients with errors (upper 5 boxplots, random errors; lower 5 boxplots, systematic errors). The boxplots consist of a box, representing the median value and interquartile range, and whiskers extending to the extreme values over 1000 simulated trials. The shaded area corresponds to the 95% confidence interval of the observed treatment effect. 14

p-values of Cochran–Mantel–Haenszel tests comparing pegaptanib and control for increasing percentages of patients with errors (upper 5 boxplots, random errors; lower 5 boxplots, systematic errors). The boxplots consist of a box, representing the median value and interquartile range, and whiskers extending to the extreme values over 1000 simulated trials. The shaded area corresponds to significant p-values.

p-values of chi-square tests comparing treatment with or without cetuximab for increasing percentages of patients with errors (upper 5 boxplots, random errors; lower 5 boxplots, systematic errors). The boxplots consist of a box, representing the median value and interquartile range, and whiskers extending to the extreme values over 1000 simulated trials. The shaded area corresponds to non-significant p-values.
Systematic errors
In contrast, when systematic errors were added to the outcomes of interest, the treatment effect was severely biased (Figures 1 and 2), and the p-values changed substantially even for small proportions of patients with added errors (Figures 3 and 4). In the ophthalmology trial, for which the original results showed a significant treatment effect in favor of pegaptanib, the simulated treatment effect dropped to zero and became progressively negative (i.e. in favor of control) when systematic errors were added, achieving statistical significance when errors were added to 50% of the patients (Figure 1); the median p-value became spuriously non-statistically significant (p > 0.05) when systematic errors were added to 10%–20% of the patients (Figure 3). In the oncology trial, the estimated treatment effect grew to become as large as a 25% difference in response rates when systematic errors were added to 50% of the patients (Figure 2); the median p-value spuriously reached statistical significance when systematic errors were added to 10%–20% of the patients (p < 0.05; Figure 4).
Conclusion
Our simulations clearly illustrate the different impact of two types of errors that can occur in clinical trials, called “random” and “systematic” for simplicity. Random errors are independent of the treatment group. Such errors typically include measurement errors, due to the limited precision of the instruments or to the subjectivity of the readers who measure the outcomes of interest; errors due to sloppiness, such as transcription errors form source documents to the case report form; and most cases of data fabrication, where investigators make up values merely to avoid having to report some data as being missing. In trials that are carried out in double-masked fashion, all data errors are by definition random if the masking is effective, that is, if there is no way for the patients or the investigators to guess what treatment an individual patient was allocated to.
Our results should not be seen as an encouragement for sloppiness or for the saving of resources that aim at identifying and eliminating systematic errors that can seriously bias the treatment effects. Instead, they should be seen as another justification for using randomized experiments, which are inherently more robust to random errors and poor data quality than non-randomized ones. Of note, randomized trials are similarly robust to other types of errors not studied here (e.g. in covariates used for stratification or for model adjustments). Randomized trials may, however, be subject to systematic errors and biases, especially in non-masked or ineffectively masked trials if the endpoint under consideration is subjective and potentially influenced by knowledge of the treatment allocation, or as a result of flaws in trial design, for instance if patients with incomplete therapy are excluded, or if the follow-up is different in the treatment and in the control group. Systematic errors are also observed in some cases of data falsification, where data values are intentionally modified to create or enhance a treatment difference.
The results of our simulations can be predicted from statistical theory, but they demonstrate using actual data that a surprisingly large amount of errors can be tolerated in randomized trials if, and only if, the errors are independent of treatment group. Of note, the standard deviation of the errors added in the simulations was extremely large. In practice, measurement errors are never that large. When such errors are added at random, that is, without regard to treatment group, they tend to cancel each other (i.e. their average tends to zero), which results in no bias and a trivial increase in the variance of the estimated effect and the p-values. Therefore, the impact of random errors would only be a concern, if at all, in trials much smaller than those considered here. In contrast, the simulations also demonstrate that systematic errors added to a surprisingly small proportion of patients are sufficient to produce a major bias in the estimated treatment effect, with a corresponding impact on statistical significance. In the example of the ophthalmology trial, a highly significant treatment effect was wiped out by adding systematic errors to a small proportion of patients, while in the oncology trial, a non-existent treatment effect was artificially created by adding systematic errors to a small proportion of patients.
Our study was carried out in the context of binary endpoints. This context has been extensively studied in the statistical and epidemiological literature. Bross 18 showed that misclassification in 2 × 2 tables does not affect the validity of the significance test to compare two similarly affected groups, but tends to reduce the power of the test. Here, however, we study the impact of random errors applied to the measurements used to calculate the binary outcome of interest, which is a step remote from the misclassification problem. When random errors are added to the measurements, they produce no net misclassification on average. When systematic errors are added, they produce a so-called non-differential misclassification, which results in biased estimates of treatment effects. 19
Although our results were obtained with binary endpoints, which are in broad use in clinical trials because of their straightforward interpretation, similar results could easily be generated with other commonly used types of endpoints. In the trial in age-related macular degeneration, for instance, an analysis of the change in visual acuity at 1 year (a continuous outcome usually measured on a 0–100 point scale) led to qualitatively similar conclusions as the analysis of the proportion of patients losing fewer than 15 letters of visual acuity at 1 year (data not shown). Other authors have studied the impact of measurement errors on time-to-event endpoints such as progression-free survival in the advanced breast cancer trial. 6 These endpoints raise other types of issues because errors can bear on the measurements made over time (e.g. tumor size) as well as on the timing of an event defined based on these measurements (e.g. progression of disease).
Our results suggest that in randomized clinical trials, every effort should be made to avoid systematic errors, while there is no justification for devoting major resources to the detection and correction of random errors. Unfortunately, in reality, a large amount of resources are currently devoted to avoid both types of errors. 20 For example, it is still common for industry-sponsored trials to aim for 100% source data verification during on-site monitoring visits.5,21 This is a costly and largely useless activity to ensure the reliability of the trial results, since in properly masked trials, the randomized treatment comparison is extremely robust to most of the errors that such intensive data verification might uncover.12,22 A recent FDA Guidance for Industry makes it clear that there is no regulatory requirement to perform 100% data verification, hence the high cost of this activity can no longer be justified even in trials of new drugs conducted by the pharmaceutical industry. 23 As a matter of fact, the FDA Guidance explicitly suggests the use of alternative, less costly, and more efficient methods for ensuring the quality of the data and for reducing risks in clinical research, especially through the use of centralized monitoring. A similar approach is advocated in a position paper from the European Medicines Agency, where the emphasis is placed on quality by design rather than on onerous and ineffective data checks. 24 Labor-intensive activities can then be appropriately focused on checking patient safety, as well as making sure that no systematic errors have occurred due to procedural or human factors, since these could have the potential of causing a major bias in the comparison of the randomized treatments.25,26 All in all, alternatives to verification of 100% of source data are now available under the general umbrella of risk-based monitoring. These alternatives include targeted (or “for cause”) on-site monitoring, as well as different types of centralized monitoring, including statistical monitoring. 27
Arguably, a challenge for investigators and sponsors is that, during a given trial, it may be difficult to differentiate between random and systematic errors, especially because of the blinding of the treatment arms. It may be worthwhile to invest more resources in earlier stages of a trial, perhaps using extensive data monitoring, to determine that errors are chiefly random, and then implement less costly quality-assurance strategies going forward. Central statistical monitoring can be more effective in this respect than on-site monitoring, since the latter is based on a case-by-case review, whereas the former uses the totality of the data. It is important to stress that most systematic errors can be prevented through proper design. 28 Double blinding is often used to minimize the potential for systematic errors, but alternatives are available for trials that cannot be conducted in double-masked manner (as is frequently the case in oncology). In such open-label trials, endpoints can be ascertained by independent investigators, patient exclusions can be adjudicated by an independent panel of experts, centers with delinquent or suspicious data can be eliminated from the analysis after review by the trial steering committee, and so on. 25 Even with these precautions in place, systematic errors can still creep in the conduct of a trial, and therefore, it is desirable to scrutinize all incoming trial data statistically in order to discover strange patterns that could be the mark of systematic errors or even fraud in some of the participating centers.27,29,30 Such statistical monitoring, followed by in-depth audits only in those centers flagged as statistically aberrant, may be an effective way of vastly reducing the cost of clinical research while at the same time improving its quality.3,12
Footnotes
Acknowledgements
The authors are grateful to Eyetech Pharmaceuticals and Pfizer Inc. for permission to use data from the ophthalmology trial.
Declaration of conflicting interests
M.B., P.S., E.Q. and E.C. are employees of IDDI. M.B. and E.Q. hold stock of IDDI. M.B. holds stock of CluePoints.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Trial registration
Trial registration numbers: NCT00208546 and NCT00321997.
