Abstract
The log-rank test is widely used to test difference in event time distribution between treatment groups. However, if subjects are not randomly assigned to treatment groups, which is often the case in observation studies, the log-rank test is not asymptotically correct for detecting group survival difference due to the imbalance of confounding variables between groups. We develop a class of modified weighted log-rank tests and Renyi-type tests for two-sample survival comparison under non-random treatment assignment. The new tests can also account for non-random censoring that depends on baseline covariates. The proposed methods involve building working models for treatment assignment, cause-specific hazard of dependent censoring, and the time to event. We prove that, when either the models for treatment assignment and dependent censoring or the model for the event time is true, the new tests are asymptotically correct, i.e. being doubly robust. Numerical experiments demonstrate the tests’ double-robustness property in finite samples of realistic sizes, and also show that the doubly robust log-rank test is at least as powerful as the regular log-rank test when the treatment assignment is random and there is no dependent censoring. An application to a kidney transplant data set illustrates the utility of the proposed methods.
1 Introduction
The log-rank test is the most popular method for two-sample survival comparison, and it is the most powerful nonparametric test when the two hazard functions differ in a proportional fashion. Its variants, weighted log-rank tests (Fleming and Harrington, 1 Chapter 7) and the Renyi-type tests, 2 are powerful to detect various types of hazard difference other than proportionality, such as early departure and crossing. For all these tests to be valid, the independent censoring assumption (Condition (3.7) in Fleming and Harrington 1 ) must be satisfied in each of the two samples. When the tests are used to compare two treatments with respect to a survival endpoint, the treatment assignment must be random for having a causal interpretation of the tests’ results.
In observational studies, whether or not a subject receives certain treatment often depends on some variables of the subject, e.g. age, health condition and insurance coverage. These variables are also often influential to the event time of interest. Thus they are so-called confounders for evaluating the treatment effect. Xie and Liu 3 developed an inverse-probability-of-treatment weighted (IPTW) log-rank test to deal with confounding in two-sample survival comparison. However, as shown in Li, 4 the IPTW test is conservative due to not accounting for the variability of the estimated weights. The independent censoring assumption is also violated in many applications. For instance, in studies with long follow-ups, there is usually a considerable percentage of censoring due to loss to follow-up rather than administrative reasons, and the lost subjects often differ from the remaining group in some baseline characteristics such as age, education, marital status, physical fitness, and socioeconomic status that have effects on the even time of interest. Robins and Finkelstein 5 proposed an inverse-probability-of-censoring weighted (IPCW) log-rank test to deal with dependent censoring in two-sample survival comparison. Unfortunately, their test also does not account for the variability of the estimated weights, and thus is expected to be conservative, similar to the IPTW log-rank test.
For the situation where both non-random treatment assignment and dependent censoring are present, Schaubel and Wei 6 and Zhang and Schaubel 7 proposed methods for estimating the cumulative treatment effect on time-to-event outcomes. Specifically, Schaubel and Wei 6 developed double inverse-weighted (inverse-probability-of-treatment-and-censoring weighting) estimators for the ratio of cumulative hazards, relative risk and difference in restricted mean lifetime; Zhang and Schaubel 7 developed a doubly robust estimator for the difference in restricted mean lifetime, which is consistent when either the even time model given adjustment covariates is correctly specified or the treatment assignment and dependent censoring models given adjustment covariates are correctly specified. For the same situation, Li 4 proposed a class of adjusted weighted log-rank tests and their Renyi-type versions using the idea of double inverse weighting. 6 Li’s 4 methods can test for the difference between treatment-specific hazards across an arbitrary time window, while the methods of Schaubel and Wei 6 and Zhang and Schaubel 7 can only infer the cumulative treatment effect from the time origin to certain time point.
In this paper, we develop a class of doubly robust weighted log-rank tests and their Renyi-type versions for two-sample survival comparison in the presence of confounders and dependent censoring. Viewing treatment assignment and censoring as a monotone coarsening (Tsiatis, 8 Chapter 8), the development of the tests is similar to constructing doubly robust estimators for parameters in semiparametric models with coarsened data (Tsiatis, 8 Chapter 10). In particular, the tests involve building working models for treatment assignment, dependent censoring and the event time given a set of adjustment covariates. Double robustness here means that the proposed tests are asymptotically correct when either the even time model is true or the treatment assignment and dependent censoring models are true. The new tests are not only more robust than the adjusted weighted log-rank tests, 4 which require both the assumed models for treatment assignment and dependent censoring to hold, but also expected to be more powerful when all the three working models are true, as implied by Theorem 10.1 of Tsiatis. 8 Li 4 showed via simulations that the adjusted weighted log-rank tests are as powerful as the regular log-rank test in the absence of confounders and dependent censoring. We expect that the doubly robust tests also enjoy this property and may be more powerful than the log-rank test if the event time working model utilizes some covariates that are informative of the survival experience.
The rest of the article is organized as follows. In Section 2, we develop a class of doubly robust weighted log-rank statistics and give their asymptotic distributions, based on which we then develop the doubly robust weighted log-rank and Renyi-type tests and establish their double robustness property. As by-products, a doubly robust cumulative hazard estimator and a doubly robust survival estimator are also developed. The former has a similar form to the doubly robust estimator in Zhang and Schaubel, 7 but is derived under a weaker assumption. Section 3 presents numerical experiments to assess the Type-I errors and powers of the doubly robust tests in finite samples. The proposed methods are applied to a kidney transplant data set in Section 4. We conclude the article with some discussion on future research directions. The proof of the asymptotic distribution for the doubly robust weighted log-rank statistic is delegated to a Web Appendix.
2 Methods
2.1 Data and assumptions
We aim to compare the population-average survival between two treatments labelled 0 and 1. Specifically, letting
A standard weighted log-rank test builds on increments of the Nelson–Aalen estimator for the treatment-specific cumulative hazard. If there were no censoring and
Throughout the article, we make the coarsening at random (CAR) assumption
However, it will be convenient to continue with this abuse of notation, and it will not invalidate the exposition of the methods in the sequel. Under monotone coarsening, the CAR assumption is equivalent to that
2.2 A Doubly robust estimator for the treatment-specific survival function
To compare treatment-specific survivals, one usually plots estimated survival curves at first. In the presence of non-random treatment allocation or dependent censoring, the popular Kaplan–Meier estimator no longer consistently estimates the treatment-specific survival function
Under monotone coarsening and the CAR assumption,
For the cause-specific hazard of censoring
If model (7) is correct,
For the hazard of event time
If model (8) is correct,
Define
Thus
Plugging equations (9) to (14) into equation (4), we obtain
Theorem 1 shows that, if either models (5) and (7) are correct or model (8) is correct, Set Theorem 1
The uniform consistency of the treatment-specific survival estimator
Under the CAR assumption, the regularity conditions (a)–(k) in the Web Appendix and the assumption that either models (5) and (7) are correct or model (8) is correct, One can consistently estimate the covariance function Corollary 1
2.3 Doubly robust weighted log-rank tests
We propose a class of doubly robust weighted log-rank tests based on the stochastic process of cumulative weighted differences between the doubly robust estimators
Theorem 2 in the following shows the asymptotic representation of the doubly robust weighted log-rank statistic (17), based on which we will develop the doubly robust weighted log-rank tests and Renyi-type tests. Set Theorem 2
The proof of Theorem 2 is delegated to the Web Appendix. Its basic idea is to write
Theorem 1, Corollary 1 and Theorem 2 are proved under the condition that the censoring time C is merely a dependent censoring time. If there is also an independent censoring time that is independent of the event time and the dependent censoring time given the treatment, one should just model the cause-specific hazard of the dependent censoring time given the treatment and the baseline covariates, and Theorem 1, Corollary 1 and Theorem 2 still hold. The proofs with independent censoring are very similar to those without independent censoring and thus omitted.
Remark 1
The doubly robust weighted log-rank test for
Hence, the doubly robust weighted log-rank test is asymptotically correct.
To compare the treatment-specific hazards over an arbitrary subset
2.4 Doubly robust Renyi-type tests
In view of the original Renyi-type test,
2
we define the doubly robust Renyi-type test statistic as
It is hard to derive the asymptotic distribution of Calculate Generate m independent sets of independent standard normal variables Compute the simulation-based p-value: Reject H0 at level α if
Theorem 2, the multiplier central limit theorem, and the law of large numbers in conjunction imply that this doubly robust Renyi-type test has asymptotic significance level α.
Following the development of the doubly robust weighted log-rank test for
3 Numerical experiments
We conducted numerical experiments to check the Type-I errors and powers of the proposed tests under non-random treatment assignment and dependent censoring. In all the simulations, the nominal significance level for any test is 0.05 and the number of Monte Carlo repetitions is 1000. We used similar simulation senarios to Zhang and Schaubel
7
and Li.
4
For every Monte Carlo sample, we generated three baseline covariates X1, X2 and X3 each from a standard normal distribution truncated at –0.5 and 0.5. The correlation between untruncated X1 and X3 equals 0.2, and all other pairwise correlations are 0. The treatment variable A was simulated from Bernoulli with probability
The treatment-specific hazard plots for the Null:
Nearly Proportional
Early Departure
Crossing

To show double robustness, we considered the doubly robust tests for which only the coarsening model or the event time model is correctly specified. We also investigated the empirical Type-I errors of the log-rank test and the adjusted log-rank tests for which only the treatment assignment model or the dependent censoring model is true. For the treatment assignment model used in the doubly robust tests and the adjusted log-rank tests, the correct model was fitted using covariates X1 and X2, and the incorrect model was fitted using X1 only. For the dependent censoring model used in the doubly robust tests and the adjusted log-rank tests, the correct model was fitted using covariates X1 and X2, and the incorrect model was fitted using X2 only. For the event time model used in the doubly robust tests, the correct model was fitted using covariates X1, X2 and X3, and the incorrect model was fitted using X1 and X3. We use ‘TTF’ and ‘FFT’ to, respectively, mean that only the treatment assignment and dependent censoring models are true and that only the event time model is true, for the doubly robust tests. Similarly, ‘TF’ and ‘FT’, respectively, mean that only the treatment assignment model is true and that only the dependent censoring model is true, for the adjusted weighted log-rank tests.
Type-I error rates and powers of the doubly robust log-rank (DRLR), Prentice–Wilcoxon (DRPW) and Renyi-type log-rank (DRRLR) tests.
Type-I error rates of the log-rank test (LR) and the adjusted log-rank tests (ALR) under non-random treatment assignment and dependent censoring (in the
Powers of the doubly robust (DR) log-rank (LR), Prentice–Wilcoxon (PW) and Renyi-type log-rank tests with all the working models being true and the adjusted log-rank (LR), Prentice–Wilcoxon (PW) and Renyi-type log-rank tests in the
In the Web Appendix B.1, we compared the Type-I errors and powers of the doubly robust tests versus the log-rank test in the absence of non-random treatment assignment and dependent censoring. The simulation results show that the doubly robust tests maintain the nominal Type-I error even if they make unnecessary adjustment for confounding and dependent censoring, and that they are at least as powerful as the log-rank test. In the Web Appendix B.2, we evaluated the finite sample performance of the doubly robust survival estimator
4 Analysis of the SRTR data
Kidney-alone (KA) transplant and simultaneous kidney-pancreas (SPK) transplant are two treatment options for end-stage renal disease (ERSD) in Type I diabetes patients. Compared to KA transplant, SPK has the advantage of treating both ERSD and the diabetes, but is a more complicated surgery, likely leading to more post-operative complications. It is of clinical interest to compare the post-transplant survival experience between these two treatments. To do that, we apply the proposed methods to a data set of Type I diabetes patients with ERSD who had a SPK or KA transplant at age
We include only the subjects who received the transplant for the first time in the analysis. Subjects lost to follow-up or surviving through 31 May 2016 without observed graft failure were censored. We assume that the choice of transplant type and the censoring are non-random, and consider a common set of baseline covariates in the three working models for the transplant type, the censoring hazard and the survival outcome. These covariates are age at transplant, gender, race, blood type, pretransplant time on dialysis, and donor age, same as in Zhang and Schaubel
7
and Li.
4
The final sample contains 1930 KA and 1636 SPK transplant recipients without missing values in the outcome or covariates. The longest common follow-up time between the two groups of transplant recipients is 13.58 years. Seventy-eight percent of the subjects in the sample were censored. The fitted logistic model for transplant type shows that age at transplant, race and donor age are significant at 0.05 level. This model fits the data decently as reflected by a c-statistic of 0.841. In the fitted cause-specific hazard model for censoring, age at transplant is significant for KA recipients, and age at transplant, race and donor age are significant for SPK recipients. The Cox–Snell residual plots (Figure 1 in the Web Appendix) show that the model fits the censoring time data well. In the fitted Cox model for the event time, age at transplant and race are significant for KA recipients, and race and donor age are significant for SPK recipients. The Cox–Snell residual plots (Figure 2 in the Web Appendix) show that the model fits the event time data well except at the right tail of the event time distribution. The overlapping of some significant covariates among the three models suggests the existence of confounding and dependent censoring.
The doubly robust estimates (with 95% pointwise asymptotic confidence intervals based on the log-log transformation) of average survival function (left), conditional survival function given no failure for four years (middle), and conditional survival function given no failure for six years (right) of graft failure time for KA recipients (solid line) and SPK recipients (dashed line).
The two-sided p-values for testing the KA versus SPK difference in the graft failure hazard over 0 to 4, 4 to 6, and 6 to 10 years after transplant.
DRLR: doubly robust log-rank; DRPW: doubly robust Prentice–Wilcoxon; DRRLR: doubly robust Renyi-type log-rank; DRRPW: doubly robust Renyi-type Prentice–Wilcoxon; ALR: adjusted log-rank; LR: log-rank.
5 Discussion
We developed a set of doubly robust weighted log-rank tests and Renyi-type tests for two-sample survival comparison under confounding and dependent censoring. Through asymptotic arguments and simulations, we showed that the new tests have the nominal Type-I error rate and are consistent when either the working models for treatment assignment and dependent censoring are true or the working model for the event time is true. Our simulations also showed that the doubly robust tests are more powerful than the adjusted weighted log-rank tests 4 if all the assumed models for treatment assignment, dependent censoring and the event time are true, and that they maintain the nominal Type-I error and are at least as powerful as the log-rank test in the absence of confounding and dependent censoring.
The proposed tests apply to competing risks data for comparing population-average cause-specific hazards between two treatment groups. They also apply to left-truncated survival data. Furthermore, they apply to survival data with both dependent and independent censoring, where one should just model the cause-specific hazard of dependent censoring given the covariates, as demonstrated in the simulations.
We conclude by pointing out some future research directions. First, it is meaningful to extend our two-sample tests to multi-sample tests. This can be done by using a multinomial logistic regression model for treatment assignment. Second, either or both of the Cox models for dependent censoring and event time could be replaced by other hazard models, e.g. the additive hazards model. 14 Third, we presume that the covariates needed to satisfy the CAR assumption are time-independent. It would be ideal to relax this assumption to allow time-varying covariates with (almost) known trajectories (e.g. under frequent monitoring). Theoretically, this can be achieved using time-dependent Cox models for dependent censoring and event time, but it is not easy to program the resulting tests, as the plug-in estimator for the asymptotic variance of the new doubly robust log-rank statistic is rather involved. Bootstrap might be helpful to reduce the programming effort here. Finally, how to construct a doubly robust test for comparing treatment-specific survivals in sub-groups would be an interesting topic in this precision medicine era. This subgroup analysis is not just to apply our methods exclusively to the sub-group data, but rather to borrow information from the other subjects in the sample to enhance the efficiency of the doubly robust tests under some reasonable model assumptions.
Supplemental Material
Supplemental material for Doubly robust weighted log-rank tests and Renyi-type tests under non-random treatment assignment and dependent censoring
Supplemental Material for Doubly robust weighted log-rank tests and Renyi-type tests under non-random treatment assignment and dependent censoring by Chenxi Li in Statistical Methods in Medical Research
Footnotes
Acknowledgements
The kidney transplant data used here have been supplied by the Minneapolis Medical Research Foundation (MMRF) as the contractor for the Scientific Registry of Transplant Recipients (SRTR). The interpretation and reporting of these data are the responsibility of the author and in no way should be seen as an official policy of or interpretation by the SRTR or the U.S. Government.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
