Abstract
Background
Composite time-to-event endpoints are beneficial for assessing related outcomes jointly in clinical trials, but components of the endpoint may have different censoring mechanisms. For example, in the PRagmatic EValuation of evENTs And Benefits of Lipid-lowering in oldEr adults (PREVENTABLE) trial, the composite outcome contains one endpoint that is right censored (all-cause mortality) and two endpoints that are interval censored (dementia and persistent disability). Although Cox regression is an established method for time-to-event outcomes, it is unclear how models perform under differing component-wise censoring schemes for large clinical trial data. The goal of this article is to conduct a simulation study to investigate the performance of Cox models under different scenarios for composite endpoints with component-wise censoring.
Methods
We simulated data by varying the strength and direction of the association between treatment and outcome for the two component types, the proportion of events arising from the components of the outcome (right censored and interval censored), and the method for including the interval-censored component in the Cox model (upper value and midpoint of the interval). Under these scenarios, we compared the treatment effect estimate bias, confidence interval coverage, and power.
Results
Based on the simulation study, Cox models generally have adequate power to achieve statistical significance for comparing treatments for composite outcomes with component-wise censoring. In our simulation study, we did not observe substantive bias for scenarios under the null hypothesis or when the treatment has a similar relative effect on each component outcome. Performance was similar regardless of if the upper value or midpoint of the interval-censored part of the composite outcome was used.
Conclusion
Cox regression is a suitable method for analysis of clinical trial data with composite time-to-event endpoints subject to different component-wise censoring mechanisms.
Keywords
Introduction
Composite time-to-event endpoints are a popular choice for assessing related outcomes jointly in clinical trials, as they can offer higher power compared to assessing outcomes individually as long as the effects of the intervention on the components are similar. However, components of the endpoint may have different ascertainment mechanisms, producing a mixture of censoring types. Within a composite endpoint, some outcomes may be right censored, which occurs when the exact time of censoring occurs. A common example of a right-censored outcome is death because the exact time of death is typically known. However, some outcomes in composite endpoints may be interval censored, which occurs when an event happens within a known time interval, but the exact event time within the interval is unknown. This can typically occur when a test or measure is collected at fixed intervals, such as annual assessments.
There are many examples of composite time-to-event endpoints in current clinical trials in a variety of clinical areas. For instance, there is a long history of cardiovascular composite outcomes in hypertension and other cardiovascular disease trials, integrating events such as myocardial infarction, stroke, heart failure, and cardiovascular death. 1 More recent trials focused on older adults have selected patient-centered outcomes such as the development of dementia and disability, which, unlike cardiovascular disease events, are not necessarily tied to an acute medical procedure or hospitalization. For example, the primary endpoint for both the Aspirin in Reducing Events in the Elderly (ASPREE) 2 and the ongoing PRagmatic EValuation of evENTs And Benefits of Lipid-lowering in oldEr adults (PREVENTABLE) trials is a composite outcome of survival free of dementia and persistent disability. 3 For such an outcome, one component is right censored (all-cause mortality), whereas dementia and persistent disability are assessed at follow-up visits with a pre-specified frequency and are subject to interval censoring. This type of scenario has been referred to as component-wise censoring or dual censoring and arises with respect to several different composite outcomes, including progression-free survival in cancer patients4–7 and the progression of chronic kidney disease. 8 It can also occur in cardiovascular disease trials when including both myocardial infarction and “silent” myocardial infarction as derived from serial electrocardiograms. 9
Several approaches have been proposed for analyzing composite time-to-event outcomes with component-wise censoring, including regression models, illness death multi-state models, and multiple imputation. The latter approach involves using multiple imputation to fill in times from interval-censored events, subsequently using standard regression or multi-state model techniques for right-censored data.10,11 A drawback of multiple imputation is that the endpoint depends on a stochastic mechanism, thereby producing inconsistent results across different imputation runs. Another limitation is that it may be difficult to a priori specify variables to be included in an imputation model, in addition to functional forms, as it is unknown which variables will be related to the probability of observing an outcome event, particularly those observed post-randomization. In addition, variables for imputation may be associated with the trial’s primary outcome, which might result in biased treatment effect estimates. 12 Multi-state models allow for the analysis of transitions across discrete states;8,13–19 however, a limitation of this approach is that a treatment effect estimate is produced for each transition path. This is often not ideal for randomized trials, where it is desirable to have a single measure of treatment effect so as not to suffer losses in efficiency driven by multiplicity. In several recent papers, novel semi-parametric regression approaches for composite endpoints with component-wise censoring have been proposed,20–25 but these methods are not currently well-established or available in standard statistical software.
Given the limitations of multiple imputation and multi-state models, the ubiquitous Cox 26 regression model is typically selected for the primary analysis for composite endpoints with component-wise censoring in clinical trials. Although this is a well-accepted method for right-censored time-to-event endpoints, it is unclear how defaulting to the use of Cox regression performs under a component-wise censoring scheme, in which a composite outcome is determined based on time to the first event occurrence from two or more outcomes. An open question remains about exactly how to calculate the composite outcome for use in Cox regression when at least one outcome is right-censored and therefore exactly measured and other outcome(s) occur within a known interval of time. For interval-censored outcomes, it is unknown if the midpoint of the interval or the upper value of the interval should be used to calculate the composite outcome. A recent paper by Eaton and Zabor 27 compared Cox regression models with multi-state models for component-wise censoring, but this paper focused on differential visit schedules and used relatively small sample sizes of 100, 200, and 400. A simulation study suggested that Cox model estimates had adequate power and low bias when the upper value of the interval-censored component was used in the composite outcome. However, it is unclear if these results are applicable for large clinical trials with thousands of participants and rarer events, such as PREVENTABLE, compared to the motivating example in cancer in Eaton’s recent paper.
The goal of this article was to conduct a simulation study to investigate the performance of Cox regression under different scenarios for composite endpoints with component-wise censoring for a large, randomized trial. We simulated data by varying the strength and direction of the association between treatment and outcome for right- and interval-censored endpoints, the proportion of events within the composite outcome arising from right-and interval-censored endpoints, and the method for including interval-censored outcomes in the Cox model (upper value and midpoint of the interval). Under these scenarios, we estimated the bias, confidence interval coverage, and power for the treatment comparison to achieve statistical significance, as estimated via Cox regression.
The remainder of this article is organized as follows. In the following section, we describe the motivating clinical trial, PREVENTABLE, and in section “Methods,” we describe Cox models and present the data simulation setup. In section “Results,” we present results from the simulation study. Finally, in section “Discussion,” we discuss implications of the simulation results in general, and specifically for the primary analysis of the PREVENTABLE trial.
Motivating clinical trial
PREVENTABLE is a pragmatic clinical trial funded by the National Institute on Aging designed to assess the efficacy of statins in older adults (https://preventabletrial.org). While there is evidence favoring the use statins in preventing cardiovascular disease, 28 evidence is more limited in older adults who are at risk for a number of competing diseases including cardiovascular disease, cancer, chronic kidney disease, and dementia. 29 The trial plans to recruit 20,000 community-dwelling adults aged 75 years or older without clinically evident cardiovascular disease, significant disability, or dementia. Participants are randomly assigned to treatment with a moderate intensity statin (atorvastatin 40 mg daily) or placebo, with planned median follow-up of approximately 4 years. The trial is being conducted across the National Patient-Centered Clinical Research Network (PCORnet) 30 and the Veterans Affairs system, and began recruiting participants in September of 2020.
The primary endpoint for the trial is survival free of persistent disability and dementia. As previously described, this is a composite endpoint, with one outcome that is right censored (all-cause mortality) and two outcomes subject to interval censoring (disability and dementia). Mortality is ascertained from electronic health record data, Medicare claims, and the National Death Index. Disability and dementia are ascertained via annual telephone assessments and are interval censored because the exact times to these events are unknown, but the events occurred at some point between assessments. The composite primary endpoint in PREVENTABLE is defined as the time that the first event occurs out of the different outcomes. While planning the statistical analysis plan for the trial, we discovered that although many trialists have applied Cox regression to composite outcomes subject to interval or mixed censoring schemes,2,31,32 including ourselves, 33 there is a paucity of literature that establishes its performance in terms of power, confidence interval coverage, and bias for the treatment effect estimate. The next section describes Cox models and the simulation study designed to evaluate its performance in the setting of component-wise censoring for composite endpoints, patterned after the PREVENTABLE trial.
Methods
Cox regression is a powerful semi-parametric method for analyzing right-censored time to event outcomes with one or more predictor variables. The model is written as follows:
where
We aimed to mimic the PREVENTABLE clinical trial setup within the data simulation study to evaluate Cox regression for composite outcomes with component-wise censoring. We generated a total of 1000 data sets containing 20000 subjects for each scenario. The simulation algorithm is depicted in Figure 1. We simulated equal proportions for treatment allocation of subjects based on a binomial distribution with probability 0.5. Next, we generated total follow-up times for subjects from a uniform distribution ranging from 3 to 5 years. Then, we generated a right-censored outcome based on a Weibull distribution with proportional hazards using the simsurv R package. 36 The hazard function for the Weibull distribution was
where

Simulation study algorithm.
Several evaluation metrics were analyzed from the data simulation. The primary outcome of interest was empirical power to achieve statistical significance for the treatment comparison, defined as the proportion of the 1000 simulation runs in which the treatment coefficient from the Cox model was significant at the 0.05 level. Secondary outcomes of interest were bias and coverage, for scenarios where the treatment coefficient was the same for the interval- and right-censored components of the composite outcome. The average bias for the coefficients for the right- and interval-censored outcomes was calculated across the 1000 simulation runs from the Cox model using the composite outcome. We also calculated the empirical 95% coverage probability, defined as the proportion of the 1000 simulation runs in which the true value of the coefficient was contained within the 95% confidence interval from the Cox model using the composite outcome. We compared empirical power for simulated scenarios when using the upper value or midpoint value for the interval-censored component of the composite outcome via Bland–Altman plots with the R package blandr. 39
We investigated different scenarios by varying simulation parameters. We varied the regression coefficients for both the right- and interval-censored outcomes over a range of values: −0.3, −0.15, 0, 0.15, and 0.3, corresponding to hazard ratios of 0.74, 0.86, 1, 1.16, and 1.35, which allowed for analysis of differing strengths and directions of associations between treatment and outcome. We also varied the proportion of events arising from the right- and interval-censored outcomes, with scenarios with equal numbers coming from the right- and interval-censored outcomes as well as scenarios with a majority (75%) arising from one or the other. This was achieved with different Weibull parameter specifications in the simulation. For approximately equal allocation, we used λ = 0.03 and γ = 0.6 for the right-censored outcome and λ = 2 and γ = 7 for the interval-censored outcomes. For the scenario with approximately 75% of events from the right-censored outcome and 25% from the interval-censored outcomes, we used λ = 0.03 and γ = 0.3 for the right-censored outcome and λ = 2 and γ = 6 for the interval-censored outcomes. For the scenario with approximately 25% of events from the right-censored outcome and 75% from the interval-censored outcomes, we used λ = 0.03 and γ = 0.8 for the right-censored outcome and λ = 2 and γ = 10 for the interval-censored outcomes. We chose these Weibull parameter values because we aimed to have an event rate similar to that expected for PREVENTABLE, around 10%–15% based on data from SPRINT and ASPREE.2,33 These simulations were run under two general scenarios: one using the upper endpoint of the interval and one using the midpoint of the interval in subjects that had the interval-censored outcomes, because it is unclear how to best include the interval-censored outcomes within the composite endpoint.
Simulations were conducted using R software version 4.1.2. Code for the simulation and analysis, as well as running additional simulations based on user inputted parameters, is available on Github (https://github.com/speiser10/Component-wise-censoring-paper). We have included a documentation file that includes instructions for how to use the code in order to generate simulations for different trial scenarios.
Results
Simulation results comparing the upper value of the interval for the interval-censored outcome and the midpoint of the interval are presented in Table 1. Based on random scatter in a band about 0 in the Bland–Altman plots, power was similar whether the upper value or midpoint of the interval-censored outcomes was used in the models for the different simulated scenarios (Figure 2). Overall, the plots in Figure 2 did not indicate that using the midpoint or the upper value within the composite endpoint is preferable. In the following results presentation, we discuss results for the simulations that used the upper value of the interval-censored outcomes in the calculation of the composite event.
Empirical power from 1000 simulations of each scenario.
Power is presented based on the one estimate of the treatment effect from the Cox model for the composite outcome.
B_right: beta coefficient for the right-censored component; B_int: beta coefficient for the interval-censored component; upper value: upper value of the interval component is used in the composite outcome; midpoint: midpoint value of the interval component is used in the composite outcome.

Bland–Altman plots comparing differences in performance metrics for models with the upper value and midpoint value of the interval-censored part of the composite outcome.
Evaluation of power
Power to achieve statistical significance for comparing treatments was impacted by the proportion of events from the right- and interval-censored outcomes (Table 1). First, we will present results for scenarios in which about half of the events came from the right-censored outcome and half came from the interval-censored outcomes (Table 1). Power to achieve statistical significance for comparing treatments was greater than 96% when the direction of the association was the same (i.e. positive values for both coefficients from the right- and interval-censored outcomes or negative values for both coefficients). In the case where the coefficients for both censoring schemes were 0, the power was approximately 5%. When one coefficient was 0 and the other was either 0.15 or −0.15, power ranged from 40% to 56%; however, power increased to above 85% for the cases where one coefficient was 0 and the other was either 0.3 or −0.3. When the direction of the association was opposite for the right- and interval-censored outcomes, power for achieving statistical significance for comparing treatments was the lowest for the moderate coefficient values (0.15 and −0.15) and slightly higher if one of the coefficients was either 0.3 or −0.3. For example, when the coefficient from the right-censored outcome was −0.3 and the coefficient for the interval-censored outcomes was either 0.15 or 0.3, power was approximately 30%, whereas with the right- and interval-censored outcome coefficients of 0.15 and −0.15, respectively, the power was 9%.
Next, we will present results for scenarios in which more events came from the right-censored outcome than the interval-censored outcomes. When the coefficient for the right-censored outcome was strong (either 0.3 or −0.3), power was generally above 95%, even if the direction of the coefficient for the interval-censored outcomes was opposite of the right-censored component. The only exception to this was when the coefficient for the right-censored outcome was −0.3 and the interval-censored outcomes were 0.3, which achieved power of 61%. When the coefficient for the right-censored outcome was moderate (either 0.15 or −0.15), power was higher than 70% when the coefficient for the interval-censored outcomes was 0 or had the same direction as the right-censored outcome. For scenarios where the coefficient for the right-censored outcome was 0, power was the highest when the coefficient for the interval-censored outcome was 0.3 or −0.3, although the power was only 60% and 42%, respectively.
Comparable results were observed for scenarios in which more events came from the interval-censored outcomes than the right-censored outcome. The highest power values were achieved when the coefficient for the interval-censored outcomes was strong (either 0.3 or −0.3). One exception to this is the scenario where the coefficients for the interval- and right-censored outcomes were −0.3 and 0.3, respectively, which had power of 35%. When the coefficient for the interval-censored outcomes was moderate (either 0.15 or −0.15), power was higher than 63% when the coefficient for the right-censored outcome was 0 or had the same direction as the interval-censored outcomes. For scenarios where the coefficient for the interval-censored outcomes was 0, power was the highest when the coefficient for the right-censored outcome was 0.3 or −0.3, in which the power was 83% and 65%, respectively.
Evaluation of bias and coverage probability
Bias for the estimated treatment effect for the coefficients was close to 0 in simulated scenarios where the right- and interval-censored coefficient values were the same (Table 2). Across the different simulation scenarios, the coverage probability for the treatment effect was close to 95% when the coefficients for the right- and interval-censored outcomes were the same (Table 2).
Bias and coverage from 1000 simulations of each scenario where the beta coefficients for treatment effect from the right- and interval-censored components are equal.
Bias and coverage are presented based on the one estimate of the treatment effect from the Cox model for the composite outcome. Beta: beta coefficient for the right-censored and interval-censored components; upper value: upper value of the interval component is used in the composite outcome; midpoint: midpoint value of the interval component is used in the composite outcome.
Discussion
We conducted a simulation study to quantify the performance of Cox regression models under component-wise censoring, where some events for a composite outcome arise from a right-censored outcome and some arise from interval-censored outcomes. We simulated several scenarios for a large, randomized trial by varying the direction and strength of association between treatment and outcome and the proportion of events attributed to each censoring scheme. Also, we included the interval-censored portion of the composite endpoint by including either the upper value of the interval or the midpoint of the interval and concluded that both of these methods to incorporate the interval-censored outcomes into the endpoint performed similarly. Under simulations assuming no treatment effect, the use of composite time-to-event outcomes subject to component-wise censoring within Cox regression did not significantly increase the Type 1 error rate. Regardless of the proportion of events coming from each censoring scheme, power to detect a treatment effect was high when the direction of the association was the same for the right and interval outcomes. We did not observe significant bias in treatment effect estimates when the relative treatment effect on each component outcome was assumed to be the same. Results from the simulation presented in this article support the primary analysis plan of using Cox regression for the PREVENTABLE composite endpoint.
We did not observe a clear preference for how to include the interval-censored outcomes within the overall composite, in terms of selecting the midpoint of the interval versus the upper value. Previous studies have similarly been mixed with respect to this choice. Based on the Cox–Aalen model, simulations in Boruvka and Cook 40 indicated better performance with selecting the midpoint under scenarios with mixed censoring. Conversely, simulations focusing on component-wise censoring and different visit schedules in Eaton and Zabor 27 suggest selecting the last-known disease-free state. However, both of these studies considered smaller sample sizes (up to N = 400 or 500) and therefore may not be applicable for larger clinical trials such as PREVENTABLE.
Use of an imputed event time for interval-censored components of a composite outcome should be interpreted cautiously given that we did not consider a full range of scenarios that varied the length of the interval between measurements, the interval length across participants, or distributions for event times. Previous work on accommodating interval-censored data within the Cox model has demonstrated bias when interval censoring is ignored, with the bias increasing as a function of width of the interval and the distribution of event times. 41 Similarly, we mentioned in the introduction the issue of bias from utilizing a composite outcome in general, independent of the censoring process. 35 We did not observe substantive bias in our simulations for scenarios under the null hypothesis or when the treatment had a similar relative effect on each component outcome. The consideration of bias is more challenging to define when there is variability in a treatment effect across components of a composite endpoint, as it is a function of the dependence between events, their stochastic ordering, the censoring distribution, and also a function of time. 35
The considerations around bias stress the importance of evaluating treatment effects on the individual components of a composite outcome. Some trials are designed such that when one event within the composite outcome occurs, the participant ceases to be followed for other events. For example, when a person is adjudicated with dementia, follow-up for other outcomes such as disability or mortality may not continue. This is not an issue in the motivating PREVENTABLE trial, as follow-up for dementia or disability still occurs after the occurrence of one of these events, and mortality is ascertained passively via electronic health records, Medicare claims, and the National Death Index. However, other trials may not be designed in this manner. Within any trial, the desire to assess the homogeneity of relative effects across components of a composite outcome will need to be balanced with the feasibility and cost of full ascertainment of all outcome events. A second consideration for composite outcome endpoints is that components may reflect competing outcomes, such that observing one event censors any other events. Thus, the treatment may affect the censoring distribution. For right-censored outcomes, this may not be a problem, but this may cause estimates for interval-censored outcomes to be biased, as we saw in our simulation. These are some examples of how investigators need to carefully consider composite time-to-event outcomes when designing clinical trials.
This study has many strengths. Under a variety of simulated scenarios for a large clinical trial, we were able to provide a thorough evaluation of a well-established method (Cox regression) in the setting of component-wise censoring, a common method for time-to-event outcomes but previously performance was unknown in this setting. We simulated data to mimic the PREVENTABLE trial, which allowed us to support the analysis plan for Cox regression with the composite outcome of death, disability, and dementia. To generalize this for other clinical trial parameters (e.g. sample size and event rate), we provide code in a Github repository so that other researchers can conduct a similar simulation study tailored to their needs.
There are some limitations to the study that should be considered. One limitation is that data were generated under the assumption of proportional hazards, which may not be appropriate for all scenarios. In this study, we generated a right-censored outcome and an interval-censored outcome, each with proportional hazards, and combined them into a composite outcome. Although each outcome individually meets the assumption, there is no guarantee that the composite endpoint will similarly satisfy proportional hazards. Therefore, it is essential to check this assumption for the composite outcome. The degree to which proportional hazards for the composite outcome may have been violated may have impacted the results in our simulated scenarios. A future study could investigate this in more detail. In addition, results may not generalize for all trials, but we provide code so that this study can be replicated for different settings. Future work could evaluate other scenarios not included in this study, such as varying the length of intervals between measurements, and the event rate as the trial progresses (constant event rate, more events early, more events late), or comparing statistical efficiency versus approaches that directly account for interval censoring.
Despite these limitations, this study provides empirical justification for using Cox regression for composite time-to-event endpoints under certain scenarios. Cox models had high power, low bias, and high coverage when the strength and direction of the association between treatment and outcome were the same for each censoring scheme; however, power was still fairly high in cases where the direction of the association was the same for the right- and interval-censored portions of the endpoint. Results from this simulation study provide practical guidance for when it is appropriate to use Cox regression in the setting of component-wise censoring for large clinical trials.
Footnotes
Acknowledgements
The authors thank Drs Dave Reboussin and Mike Miller for excellent discussions about the topic in this paper and for sharing their clinical trial expertise with us.
Data availability statement
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The study was supported by the National Institute on Aging, PRagmatic EValuation of evENTs And Benefits of Lipid-lowering in oldEr adults (PREVENTABLE) trial (grant no. U19 AG065188) and the Wake Forest Older Americans Independence Center (grant no. P30 AG021332).
