We propose a causal mediation approach to semi-competing risks under left truncation sampling by considering an intermediate event as a mediator and a terminal event as an outcome. We focus on the causal relationship from exposure to the terminal outcome in relation to the intermediate event. In particular, we study the direct effect, the effect of exposure on the terminal event that is not through the intermediate event, and the indirect effect—the effect of exposure on the terminal event that is mediated through the intermediate event. We propose nonparametric and semiparametric methods, both accounting for left truncation. The nonparametric estimator can be viewed as a model-free time-varying Nelson–Aalen estimator that is robust to model misspecification. The semiparametric estimator calculated with the Cox proportional hazards model enjoys flexibility in adjusting for potential confounders as covariates. The asymptotic properties for both estimators, including uniform consistency and weak convergence, were established using the martingale theorem and functional delta method. The finite sample performance of the proposed estimators was evaluated through extensive numerical studies that investigated the influences of left truncation, confounding, and sample size. The utility of the proposed methods was illustrated using a hepatitis study.
Motivated by the community-based prospective cohort study REVEAL,1 we consider the natural history of chronic hepatitis B. Scholars have established that, relative to the average person, hepatitis B patients have an increased incidence of liver cirrhosis and a higher mortality rate; however, we are particularly interested in whether hepatitis B affects mortality mediated by increasing the incidence of liver cirrhosis as a function of age. In developing an analytical methodology for this question, two statistical topics must be addressed: Semi-competing risks and left truncation. When two events are of interest, semi-competing risks are frequently encountered in biomedical research, where the intermediate event may be masked by the terminal event, but not vice versa.2 In our motivating example, the fact that liver cirrhosis incidence may be masked by mortality, but having liver cirrhosis remains subject to mortality, represents a typical semi-competing risks problem. The challenge of left truncation arises from the design of most cohort studies, including the REVEAL cohort, where the participants were recruited at certain time points, and individuals who had not survived up to the time of recruitment were not included in the study. This methodology may go beyond the scope of the motivation study. We propose developing novel methods of quantifying the mediation effect of an exposure or intervention on a time-to-event outcome through a time-to-event mediator in the presence of semi-competing risks and left truncation.
Current approaches to semi-competing risks can be classified into several categories. Approaches in one category seek to characterize the joint survivor function of the times to intermediate and terminal events. For example, Day et al.3 considered the predictive hazard ratio4 for the dependence of two outcomes under the assumption that the ratio is constant, and the joint survivor function corresponds to the Clayton copula model,2,5 which is also equivalent to gamma frailty models. Approaches in another category apply multistate modeling,6 where semi-competing risks data are modeled for transitions between different states: healthy (i.e., without any intermediate or terminal event), diseased (i.e., having an intermediate event), and dead (i.e., having a terminal event). Wang7 proposed a nonparametric estimator for the survival function of the sojourn time between states; by incorporating a gamma frailty in the transition hazards of the multistate model, Xu8 derived a new nonparametric maximum likelihood estimator for the copula parameter under the Clayton copula model. Approaches in another category focus on the survivor’s average causal effect (SACE; Frangakis and Rubin9 and Zhang and Rubin10). The SACE only concerns the effect on intermediate event among those who would have been survivors of terminal events, regardless of their exposure. The first two categories successfully characterized the dependence and transition probabilities of the events, but the causal interpretations of these statistical quantities remain elusive, whereas the third category focuses on a different causal effect than mediation. To this end, we propose an alternative approach to semi-competing risks using causal mediation modeling.
Causal mediation analysis provides a general framework for characterizing the mediation mechanism linking exposure and outcome. Causal mediation models originated from social science literature11,12 and were theoretically consolidated by a potential outcome framework.13 Causal mediation models have been rigorously developed for counterfactual definitions, identifiability assumptions, and estimators.14–17 Causal mediation models decompose the effect of the exposure on the outcome into an indirect effect (IE): Effect on the outcome mediated through an intermediary factor (termed a “mediator”); and a direct effect (DE), which is not through the mediator. Conceptually, formulating the semi-competing risks using a causal mediation model, where the intermediate and terminal events are the mediator and outcome, respectively, is straightforward. Although causal mediation analyses of time-to-event outcomes are available,18–21 previously published methods require a fully observable mediator. In semi-competing risks data, the intermediate event is masked by either conventional right censoring or a terminal event. Our recent work has proposed the nonparametric estimator for mediation analyses accounts for the time to an intermediate event.22 The invited discussions by Fulcher et al.23 and Stensrud et al.24 pointed to several issues, including the ill-defined time to the intermediate event after the occurrence of the terminal event, cross-world independence, and Markov assumption. While these issues have been discussed and clarified in the rejoinder,25 we propose a new estimand by adopting interventional approach26,27 to address the former two, and an analytic approach to evaluate and weaken the Markov assumption. Moreover, we found that the proposed estimand is closely related to the separable effects of Stensrud et al.,28 which is further discussed. However, the current mediation analyses for survival outcomes do not account for the bias resulting from the left truncation. Therefore, our proposed approach attempts to bridge causal mediation modeling and bivariate survival analyses under left truncation.
Left truncation can be construed as a mechanism of missing data, particularly when the time scale of interest is age.29 Under the left-truncation phenomenon, researchers tend to recruit individuals with slower-than-usual progression of a disease (overall, a healthier condition than non-participants), which leads to biased sampling of the underlying population. For the motivation study on hepatitis, we were interested in the effect as a function of age because of its biological implications. From 1991 to 1992, individuals aged 30–65 years were recruited to participate in the REVEAL cohort. Consequently, those who died before the relevant age in the same recruitment period did not have any opportunity to participate in the study. Therefore, an analysis that does not account for left truncation usually leads to a substantial overestimation of the survival time.30 To adjust for the bias caused by the left truncation, we adopted the conditional likelihood function approach,31,32 where the likelihood is calculated by conditioning on the observed left truncation time. This approach enabled us to focus more on the times to liver cirrhosis and mortality without imposing fully parametric assumptions on the truncation time and was thus robust to mis-specified distributions of the truncation time.
Assumptions and estimands
For subject where and is the sample size, let , , , and be the terminal event, intermediate event, left truncation times, exposure, and confounders, respectively, in the underlying population. With these random variables, we further denote the stochastic processes of the intermediate and terminal events as and where is an indicator function. We made the following statistical assumptions.
: both time to the underlying left truncation and time to censoring are independent of the times to the underlying intermediate and terminal events given the underlying exposure and confounders.
is the end time of study. If , then 8,33,34: if the intermediate event has not occurred prior to the terminal event, it will never occur by the end of the study.
For , and , where is the cdf of the censoring time.
: all possible combinations of given can occur with nonzero probabilities.
For , : the hazard of the terminal event at time only depends on the status of the intermediate event right before , not on its history.
These assumptions are used to ensure the independence mechanism for left truncation and right censoring: 32,35,36 (Assumption (1.1)), the sequence of intermediate and terminal events (Assumption (1.2)), and stochasticity for the survival and censoring times, as well as the occurrence of joint combinations of exposure and intermediate event status (Assumptions (1.3) and (1.4)). In addition, our approach depends on a Markov assumption (Assumption (1.5)) that is assumed in the multistate and frailty approaches.8,37 The Markov assumption states that given the same status of the intermediate event at time and the fact that subjects with the same survive until , the risk of having a terminal event at is independent of when the intermediate event occurs.
Note that the Markov assumption is evaluated in the data application, and the evaluation was conducted using the Cox model, which includes the duration since the occurrence of the intermediate event as a covariate. Such diagnostic analysis supports the plausibility of the Markov assumption (details in the Web Appendix). In addition to assumption diagnostics, inclusion of the duration-related covariates may be used as an analytic approach to satisfy or weaken the Markov assumption, which can easily be implemented by our semiparametric approach (details are in the “Semiparametric estimator” section). In addition, the Markov assumption and Assumption (2.1) in our study are related to the assumption required for causal interpretations of the Kaplan–Meier estimates of the time-varying covariate by Sjölander38 (details are provided in the Discussion section).
We introduce potential outcome notations for intermediate and terminal event processes. We denote as the setting of , is the counterfactual process of at had been set to . is the counterfactual process of at , had and been set to and 0. is the counterfactual process of at had been set to , and is the counterfactual process for at had , and been set to (not necessarily equal to ), and 0, respectively. To identify the proposed counterfactual hazard defined in the “Direct and indirect effects” section, under the interventional approach,26,27 we make the following causal assumptions.
: there is no unmeasured confounding for the association of and given .
: there is no unmeasured confounding for the association of and given .
is a random draw from the distribution of .
We stress that the above assumptions do not include the cross-world independence assumption, and are much weaker than those used by Huang.22 Assumptions (2.1) and (2.2) can be illustrated by using a single world intervention graph (SWIG; Richardson and Robins;39Figure 1). For , we denote the time-specific error terms: for , for , , , and for , and the following confounders , , , and need to be adjusted to ensure that
Therefore, outcome-outcome (-) confounder (), mediator-outcome confounder (), exposure-outcome confounder () and exposure-mediator confounder () need to be adjusted to ensure Assumptions (2.1) and (2.2). We propose sensitivity analyses to evaluate Assumptions (2.1) and (2.2) in the Web Appendix. Moreover, as mentioned in equations (1) and (2), we use and to describe the potential intervention values for and , respectively, received from exposure, which can be set at distinct levels.
Single world intervention graph (SWIG) that describe the intervention of and in as well as intervention of , and in .
Direct and indirect effects
Our key estimand is denoted by
where is a random draw from an identical probabilistic distribution of , using an idea analogous to the interventional approach.26,27 We express the counterfactual hazard without conditioning the covariates to simplify the notations, but it is trivial to stratify by throughout the development. Under Assumptions (2.1)–(2.3) prove that (detailed in the Appendix)
where = and . The identification formula is identical to that by Huang,22 but is obtained under weaker assumptions without cross-world independence. Using the counterfactual hazard, we can easily obtain the cumulative counterfactual hazard:
and the counterfactual survival probabilities
where .
In the context of the definitions by Pearl15 and Robins,16 DE is defined as the effect of on the terminal event not mediated by the intermediate event, as , where
IE is defined as the effect of on the terminal event mediated by affecting the intermediate event, as , where
Additionally, the total effect (TE) can be obtained using . Note that the effect measure is not limited to the cumulative hazard difference. Because our estimation focuses on the cumulative hazard, other effect measures such as the cumulative hazard ratio, for , for and for ; survival probability ratios, , and , can be easily obtained.
Correspondence of separable effect
Separable effect has recently been proposed to identify causal effects in the setting of competing risks.28 In the separable effect approach, the treatment can be decomposed into and , where affects without any direct effect on and affects without any direct effect on .
We denote the potential values for separable exposures and as and , respectively, to avoid confusion with and in the proposed method. Although both and affect the mediator , and and affect the outcome , their influence is on different effect measures. In our proposed method, affects the prevalence of the mediator, and affects the hazard of the outcome conditional on the mediator, as shown in (3). In the separable effect, and affect cause-specific hazard of the mediator and outcome, respectively, as will be shown in the following dismissible component conditions.
To establish a connection with Stensrud et al.,28 we derived the counterfactual hazard function under competing risks, using the assumptions specified by Stensrud et al.28 The counterfactual counting processes of and in the separable model are and , respectively. The key to identification of separable effect is the two dismissible component conditions, which can be expressed as:
Thus, with the above notations and assumptions, we define the following counterfactual hazard function of the separable effect method, denoted by , which can be estimated using the following second equality under the above dismissible component conditions (details in the Web Appendix):
where
and . On the other hand, our proposed estimand under competing risks leads to the following identification formula under Assumptions (2.1)–(2.3)
where can be expressed as
Therefore, although the causal mechanisms differ between the proposed approach and the separable effects approach, their identifications are similar, as indicated by equations (6) and (7). Furthermore, treating and , the expression for equation (6) is similar to (7). The two equations are identical in setting . Moreover, both (6) and (7) can be viewed as weighted integrals for with different weight functions and , which describes the probability of being greater than , conditional on being greater than under different intervention mechanisms for the exposure.
Proposed estimators
Here, we construct the estimation for the proposed method based on the likelihood approach (details are in Section S.3 of the Web Appendix). Moreover, from the data with left truncation, we observe only individuals with information on , where for those satisfying . The corresponding processes under left truncation and right censoring are given by and , respectively, where and .
Nonparametric estimator
We consider a nonparametric approach where potential confounding factors are adjusted using stratification. According to the conditional likelihood function mentioned in S.3.1 of the Web Appendix (equation (A8)), the -stratified conditional log-likelihood, denoted as , of the entire data set under right censoring and left truncation is equal to:
Let be the estimated jump of at from this likelihood function. Then, we obtain by solving the root of . Thus, we obtain the nonparametric estimator
in which
For the estimand , according to equation (8), we formalize as
Then, for each , we can estimate by finding the root of equation . It follows that our nonparametric estimators of and are
and , where .
Semiparametric estimator
We further develop model-based semiparametric estimators for and . Given the status of the intermediate event , we propose a Cox proportional hazards model for :
where is the baseline cumulative hazard function and is the regression coefficient for , both conditional on status . Then, according to equation (A8) of the Web Appendix and model (9), the conditional log-likelihood can be formulated as
Therefore, can be obtained by solving , and thus
where and and is explained in the following. We note that the estimator for the cumulative baseline hazard is the Breslow-type estimator. can be obtained by solving the following partial likelihood score:
that depends on the status of .
For , we relate at a given time to using a logistic regression model, where and :
Then, we estimate by using the maximum likelihood estimators which maximizes
The model-based semiparametric approach can efficiently adjust for confounding by simply including potential confounders as covariates in the models (9) and (11). We also note that to estimate DE and IE from the counterfactual hazard (3), only needs to be estimated at time points with .
Asymptotic results
Our nonparametric and semiparametric estimators are based on the counting process of the terminal event, which depends on the status of the intermediate event and . To study the asymptotic properties of the proposed estimators, we first establish the martingality of newly defined processes. Let and denote the generalized notation of the counting processes and compensators, respectively, in both methods: and for the nonparametric method and and for the semiparametric method.
Under assumptions (1.1), (1.2), and (1.5), is a zero-mean martingale under filtration:
We simplify the notation by letting denote the associated parameters in the nonparametric method, or those in the semiparametric method , and and denote the estimator and the theoretical counterpart, respectively. From the properties of the aforementioned martingale and the result of Andersen and Gill,40 one can show that estimators for the counterfactual hazard are asymptotically unbiased and consistent for a given . From propositions II.5.2 and II.5.3 in Andersen et al.,41 we further prove that unbiasedness and consistency hold uniformly.
For nonparametric and semiparametric methods, and cardinality of set goes to infinity,
converges to zero in probability, and
converges to zero, where is the end time of the study.
Applying the triangle inequality, we can prove that the proposed estimators for DE and IE also enjoy uniform asymptotic unbiasedness: Both
converge to 0, and uniform consistency
converge to 0 in probability. To study the weak convergence of the proposed estimators, we first establish that of .
As , converge weakly to a zero-mean Gaussian process denoted by .
The aforementioned result can be verified by considering the joint likelihood function as a special case in Zeng and Lin42 Because the estimand is a function of , the weak convergence results of and can be derived using the functional delta method.43
As and cardinality of set goes to infinity, and converge weakly to zero-mean Gaussian processes, denoted by and , with variances at time
respectively, where and are the derivatives of and at , respectively.
To estimate the pointwise variance in practice, we plug in and , where the forms of and and the limiting expressions for are provided in the Web Appendix.
Theorem 2 also indicates that is both asymptotically unbiased and uniformly consistent. Moreover, converges weakly to zero mean Gaussian process with variance
at time .
By treating and for as parameters. can be obtained by the block diagonal matrix , where denotes the covariance matrix of . is the covariance matrix of and is the covariance matrix of (the details of the expression can be found in the Web Appendix).
Numerical studies
Simulation
We conducted simulation experiments to evaluate the finite-sample performance of the proposed methods. Each experiment was conducted based on a total of 1,000 replicates under the sample size = 500 or 1,000. In particular, the simulation data were generated as follows: the data of , where , the confounder, follow a Bernoulli distribution with and follow a standard normal distribution; the data of follow a uniform distribution over the time interval ; censoring time follows a Weibull distribution with the scale of 2 and shape of 5; the data of follow a Weibull distribution with the scale and shape 1, where and 0 represent the alternative hypothesis (with IE) and null hypothesis (without IE), respectively, and or 0 for the setting with or without the confounding. , where follows a Weibull distribution with the scale and shape 1 where and 0 under and , respectively, and and 0 under settings with and without confounding, respectively. Moreover, we introduced additional simulations to evaluate the performance of the proposed semiparametric approach when the proportional hazards assumption is violated (detailed results illustrated in Figure S3 of the Web Appendix). Specifically, in these simulations, we replaced the Weibull distribution with a log-normal distribution. The new parameters were set with a variance of 1 and means of and for the scale parameters that were originally used in the Weibull scenarios.
To demonstrate the bias resulting only from the left truncation, but not confounding (), Figure 2 illustrates that the averages of the estimates from both nonparametric and semiparametric methods, without accounting for left truncation, deviate from the true values if nonzero effects are present. However, no bias exists under the null hypothesis. As depicted in Figure 3 () and S1 ( in the Web Appendix), our proposed methods properly correct the bias caused by left truncation, and the averages of our method’s DE and IE estimates align accurately with the true values. The performance of the proposed variance estimators was evaluated by comparing the pointwise averages of the estimated asymptotic standard errors (NP_asy and SP_asy) at each time point with those of the empirical standard error calculated by averaging across 1,000 replicates (NP_emp and SP_emp). The averages of the proposed variance estimators match well with their empirical counterparts, and the variances in the semiparametric approach are smaller than those in the nonparametric approach (Figures 3 and S1). Moreover, the confidence intervals (CIs) calculated with the proposed variance estimators in nonparametric and semiparametric approaches achieve proper (approximately 95%) coverage rates for both DE and IE (Table 1 [] and ST1 [, in the Web Appendix]). On the other hand, as illustrated in Figure S3, the estimates from the proposed semiparametric approach align with the true values under the condition. However, under , they may introduce potential bias when the proportional hazards assumption fails. To demonstrate the performance of the proposed semiparametric method in adjusting for confounders, we set . Figure 4 illustrates that without accounting for the confounder , the averages of the estimates deviate from the true values, and the bias can be corrected by adjusting for .
Performance of proposed methods without accounting for the left truncation under =500.
The asymptotic unbiasedness and variance for proposed method under =500.
Average point estimates by the proposed semiparametric method with and without accounting for the confounder under =500.
Results of (I): Direct and indirect effects with pointwise confidence interval of hepatitis B infection in relation to liver cirrhosis incidence estimated using asymptotic and bootstrap method variance.
Results of (II): Direct, indirect (left) and total (right) effects with pointwise confidence interval of hepatitis B infection in relation to liver cirrhosis incidence estimated using asymptotic and bootstrap variance.
Data application
In the REVEAL motivation study, data were collected from a community-based prospective cohort of 23,820 individuals recruited during 1991 and 1992. At enrollment, the status of hepatitis B was determined by hepatitis B surface antigen (HBsAg) from the collected serum samples. Studies have argued that most HBsAg-seropositive Taiwanese residents have hepatitis B perinatally before 3 years of age.44,45 Based on the epidemiological characteristics of Taiwan, we assume that hepatitis B in the REVEAL study occurred at birth.46,47 The REVEAL study was followed to the end of 2008; during the follow-up, liver cirrhosis and death were ascertained through data linkage with the National Cancer Registry and Death Certification Profile, respectively. In other words, it is highly likely that both the start time of left truncation and the exposure-receiving time are at the same time, that is, birth. It is more biologically meaningful to characterize hepatitis B-induced mortality in relation to the incidence of liver cirrhosis as a function of age. However, using age as the time scale is subject to bias from left truncation, where individuals entering the study at an older age may be healthier than those at a younger age. We used the proposed methods to estimate the DE and IE by age and adjusted for bias from the left truncation. To evaluate the performance of the variance estimators, we calculated the CIs using asymptotic variances and bootstrapping. Two analyses were conducted in this study. First, for a better comparison between the proposed nonparametric and semiparametric methods, we minimized potential confounding by focusing on a more homogeneous subpopulation consisting of 4,895 male participants aged less than 50 years at study entry and without any history of alcohol consumption (analysis I). Second, a total of 12,789 participants aged less than 50 years with confounders were included: Sex (male vs. female) history of alcohol consumption (yes vs. no) and alanine transaminase levels at study entry ( unit/liter vs. unit/liter and unit/liter) were analyzed using the semiparametric method (analysis II).
As displayed in Figures 5 and 6, both nonparametric and semiparametric methods in analysis I and the semiparametric method in analysis II (setting sex as male and other covariates as their sample means) provide similar results in point estimates. For the CIs, we found that the pointwise intervals calculated with asymptotic variances were similar to those obtained by bootstrapping with 1,000 replicates (constructed by and quantiles of the estimator for each time point) and that the intervals in the semiparametric method were slightly narrower than those in the nonparametric method. Both methods reveal that hepatitis B increased mortality through liver cirrhosis incidence (particularly after the age of 50 years) and that the effect not through liver cirrhosis was nonsignificant across different ages. The results suggest that most hepatitis B-related mortality occurs due to liver cirrhosis. Such analysis may also have implications for health policies. For example, to reduce preventable deaths among hepatitis B carriers, frequent liver cirrhosis screening after the age of 45 years may be implemented for this high-risk population. Assumption (1.5) is evaluated by comparing the survival curves for those with and and those with and , which supports the plausibility of the assumption (details in the Web Appendix); to evaluate Assumptions (2.1)–(2.3), we conduct sensitivity analyses for the proposed nonparametric and semiparametric methods in the Web Appendix, which supports the validity of our finding. We also applied copula and multistate models to analyze the REVEAL data under analysis I. The Clayton copula model was adopted herein using the estimators of Jiang et al.48 Kendall’s tau of 0.895 suggests a high correlation between the event times of liver cirrhosis and mortality, 95 bootstrap CI for those without hepatitis B, and 0.893 for patients with hepatitis B. The joint survival functions (Figure S4 in the Web Appendix) demonstrate that hepatitis B leads to a relatively lower joint survival probability. For analysis using the multistate model, we analyzed Path A from a healthy state to incidence of liver cirrhosis Path B from a healthy state to mortality, Path C from liver cirrhosis incidence to mortality via the R packages “mstate”.49 Multistate modeling suggests that hepatitis B significantly affects the liver cirrhosis incidence from the healthy state (Path A) with a log-hazard ratio of 2.38 (95 CI ); the effects of hepatitis B on Path B (log-hazard ratio0.12 (95 CI ) and Path C (log-hazard ratio=0.28 (95 CI ) are not significant.
Discussion
Here, we develop novel approaches to analyze semi-competing risks data in the presence of left truncation by causal mediation modeling. Although existing methods focus more on joint survivor functions (copula models) or transition probabilities between states (multistate models), we aim to characterize DE and IE, which may be more meaningful and scientifically interpretable. For statistical inference, the proposed methods are based on conditional likelihood; we calculate the likelihood of intermediate and terminal events conditioning on the truncation time.
To establish the estimand, the random-draw approach is adopted. The main idea of the random draw approach can be realistically expressed by considering the intervene and in an independent random copy (such as an experiment controlled under similar conditions) of to obtain . Thus, would have an identical distribution to but . Thus, as illustrated in Figure 1, we borrow the mechanism of instead of the original in , and then the cross-world and assumptions in Huang22 can be escaped.
Two major differences exist between the proposed nonparametric and semiparametric methods. First, adjustment for confounders is more efficient in the semiparametric method. The semiparametric method is model based and allows covariate adjustment in regression models to account for potential confounding factors. For the nonparametric method, confounders can be adjusted by stratification, which, however only applies to categorical confounders. Although stratum-specific effects may be more robust, they lose efficiency and provide no pooled effect. Second, the semiparametric method depends on the proportional hazard assumption and model assumptions, such as linearity for the included covariates; the nonparametric method is free of these assumptions, but its efficiency is traded off for robustness.
Moreover, if time-varying confounders are present, additional research is needed to fully address these complexities when confounders affected by the exposure are as shown in Figure 7(a). However, when the confounder mechanism follows the structure depicted in Figure 7(b), our semiparametric approach, which utilizes the Cox proportional hazards model, can effectively adjust for time-varying exposure-outcome confounders () by treating them as covariates. Additionally, this approach employs logistic regression to estimate model parameters at each event time, allowing for the adjustment of time-varying exposure-mediator confounders () by incorporating them as covariates.
Dag for time-varying confounding.
In the presence of left truncation, the estimation of DE and IE is unbiased if the effects are null, that is, the exposure is independent of the terminal event (Figure 2). This is because the exposure and survival times are independent in the underlying population, and the independence would still be preserved under biased sampling. However, if the effects are not null (i.e., the exposure is associated with the survival time), either through DE or IE, then biased sampling due to left truncation tends to recruit individuals with longer-than-usual survival. Moreover, the distributions of the exposure and mediator (intermediate event) in these survivors may differ from those in the underlying population.
The occurrence of an event, such as birth or diagnosis of a disease, is usually the start time for survival analysis in the presence of left truncation. The event of birth or disease occurrence can be accurately determined. However, causal inference research concerns the effect of exposure or treatment, which makes the initiation of exposure or treatment a natural time of origin. The start time of a treatment is well defined in a clinical trial setting, but it may not be a trivial task to define when exposure to a pathogen or environmental risk factor occurs in observational studies. In our data application, we exploit the epidemiologic characteristics that most hepatitis B carriers are exposed to the virus at birth.44–47 For observational studies with a less well-defined timing of exposure, careful consideration needs to be made when specifying the time of origin.
The Markov assumption and Assumption (2.1) in our study are related to the assumption required for causal interpretations of the Kaplan–Meier estimates of the time-varying covariate by Sjölander.38 As illustrated in Figure 1, the presence of time-specific variability for : and for and and for , where and are independent errors, but and share a common cause . Under our Assumption (2.1), must be adjusted to ensure for valid causal interpretation. The adjustment of may also satisfy Assumption (1) for Sjölander38 for causal interpretation. Assumption (2.1) may be weaker than Assumption (1) for Sjölander,38 because we still allow time-specific variability and rather than adjusting for all the variability of .
Supplemental Material
sj-pdf-1-smm-10.1177_09622802241313291 - Supplemental material for Causal mediation analysis for time-to-event mediator and outcome in the presence of left truncation
Supplemental material, sj-pdf-1-smm-10.1177_09622802241313291 for Causal mediation analysis for time-to-event mediator and outcome in the presence of left truncation by Jih-Chang Yu and Yen-Tsung Huang in Statistical Methods in Medical Research
Footnotes
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the National Science and Technology Council, Taiwan (grant 112-2118-M-305-005-MY2).
ORCID iDs
Jih-Chang Yu
Yen-Tsung Huang
Supplemental material
Supplemental material for this article is available online.
Appendix A. Prove the equation (1)
References
1.
ChenCJYangHISuJ, et al. Risk of hepatocellular carcinoma across a biological gradient of serum hepatitis b virus dna level. J Am Med Assoc2006; 295: 65–73.
2.
FineJPJiangHChappellR. On semi-competing risks data. Biometrika2001; 88: 907–919.
3.
DayRBryantJLefkopoulouM. Adaptation of bivariate frailty models for prediction, with application to biological markers as prognostic indicators. Biometrika1997; 84: 45–56.
4.
OakesD. Bivariate survival models induced by frailties. J Am Stat Assoc1989; 84: 487–493.
5.
ClaytonDG. A model for association in bivariate life tables and its application to epidemiological studies of familiar tendency in chronic disease epidemiology. Biometrika1978; 65: 141–151.
6.
KalbfleischJDPrenticeRL. The statistical analysis of failure time data. Inc., New Jersy: John Wiley and Sons, 2002.
7.
WangW. Nonparametric estimation of the sojourn time distributions for a multipath model. J R Stat Soc Ser B2003; 65: 921–935.
8.
XuJKalbfleischJDTaiB. Statistical analysis of illness-death processes and semicompeting risks data. Biometrics2010; 66: 716–725.
9.
FrangakisCERubinDB. The defining role of principal stratification and effects for comparing treatments adjusted for post-treatment variables: from treatment noncompliance to surrogate endpoints. Biometrics2002; 58: 191–199.
10.
ZhangJLRubinDB. Estimation of causal effects via principal stratification when some outcomes are truncated by “death”. J Educ Behav Stat2003; 28: 353–368.
11.
BaronRMKennyDA. The moderator-mediator variable distinction in social psychological research: conceptual, strategic, and statistical consideration. J Pers Soc Psychol1986; 51: 1173–1182.
12.
MacKinnonD. Introduction to statistical mediation analysis. NY: Taylor & Francis, 2008.
13.
RubinD. Bayesian inference for causal effects. Ann Stat1978; 6: 34–58.
14.
ImaiKKeeleLYamamotoT. Identification, inference and sensitivity analysis for causal mediation effects. Stat Sci2010; 25: 51–71.
15.
PearlJ. Direct and indirect effects. In: Proceedings of the seventeenth conference on uncertainty and artificial intelligence, 2001, pp.411–420. Morgan Kaufmann, San Francisco.
16.
RobinsJM. Semantics of causal DAG models and the identification of direct and indirect effects. New York, NY: Oxford University Press, 2003.
17.
VanderWeeleTJVansteelandtS. Conceptual issues concerning mediation, intervention and composition. Stat Inference2009; 2: 457–468.
18.
HuangYTCaiT. Mediation analysis for survival data using semiparametric probit models. Biometrics2016; 72: 563–574.
19.
LangeTHansenJV. Direct and indirect effects in a survival context. Epidemiology2011; 22: 575–581.
20.
Tchetgen TchetgenEJ. On causal mediation analysis with a survival outcome. Int J Biostat2011; 7: 33.
21.
VanderWeeleTJ. Causal mediation analysis with survival data. Epidemiology2011; 22: 582–585.
22.
HuangYT. Causal mediation of semicompeting risks. Biometrics2021a; 77: 1143–1154.
23.
FulcherIRShpitserIDidelezV, et al. Discussion on “causal mediation of semicompeting risks” by yen-tsung huang. Biometrics2021; 77: 1165–1169.
24.
StensrudMJYoungJGMartinussenT. Discussion on “causal mediation of semicompeting risks” by yen-tsung huang. Biometrics2021; 77: 1160–1164.
25.
HuangYT. Rejoinder to “causal mediation of semicompeting risks”. Biometrics2021b; 77: 1170–1174.
VansteelandtSDanielRM. Interventional effects for mediation analysis with multiple mediators. Epidemiology2017; 28: 258–265.
28.
StensrudMJYoungJGDidelezV, et al. Separable effects for causal inference in the presence of competing events. J Am Stat Assoc2020; 117: 1–9.
29.
TsaiWYJewellNPWangMC. A note on the product-limit estimator under right censoring and left truncation. Biometrika1987; 74: 883–886.
30.
HuangCYNingJQinJ. Semiparametric likelihood inference for left-truncated and right-censored data. Biostatistics2015; 16: 785–798.
31.
Lynden-BellD. A method of allowing for known observational selection in small samples applied to 3cr quasars. Mon Not R Astron Soc1971; 155: 95–118.
32.
WangMC. Nonparametric estimation from cross-sectional survival data. J Am Stat Assoc1991; 86: 130–143.
33.
HaneuseSLeeKH. Semi-competing risks data analysis: accounting for death as a competing risk when the outcome of interest is nonterminal. Circ: Cardiovasc Qual Outcomes2016; 9: 322–331.
34.
LeeCGilsanzPHaneuseS. Fitting a shared frailty illness-death model to left-truncated semi-competing risks data to examine the impact of education level on incident dementia. BMC Med Res Methodol2021; 21: 1–13.
35.
ChengYJHuangCY. Combined estimating equation approaches for semiparametric transformation models with length-biased survival data. Biometrics2014; 70: 608–618.
LeeKHHaneuseSSchragD, et al. Bayesian semi-parametric analysis of semi-competing risks data: investigating hospital readmission after a pancreatic cancer diagnosis. J R Stat Soc Ser C2015; 64: 253–273.
38.
SjölanderA. A cautionary note on extended kaplan–meier curves for time-varying covariates. Epidemiology2020; 31: 517–522.
39.
RichardsonTSRobinsJM. Single world intervention graphs (swigs): a unification of theh counterfactual and graphical approaches to causality, 2013.
40.
AndersenPKGillRD. Cox’s regression model for counting processes: a large sample study. The Ann Stat1982; 1100–1120.
41.
AndersenPKBorganOGillRD, et al. Statistical models based on counting processes. New York : Springer Science & Business Media, 2012.
42.
ZengDLinD. Maximum likelihood estimation in semiparametric regression models with censored data. J R Stat Soc: Ser B (Stat Methodol)2007; 69: 507–564.
43.
Van der VaartAW. Asymptotic statistics. Cambridge university press, Vol. 3, 2000.
44.
ChenCJWangLYYuMW. Epidemiology of hepatitis b virus infection in the asia-pacific region. J Gastroenterol Hepatol2000; 15: E3–E6.
45.
HsuHYChangMHChenDS, et al. Baseline seroepidemiology of hepatitis b virus infection in children in taipei, 1984: a study just before mass hepatitis b vaccination program in taiwan. J Med Virol1986; 18: 301–307.
46.
StevensCEBeasleyRPTsuiJ, et al. Vertical transmission of hepatitis b antigen in taiwan. N Engl J Med1975; 292: 771–774.
47.
StevensCENeurathRABeasleyRP, et al. Hbeag and anti-hbe detection by radioimmunoassay: correlation with vertical transmission of hepatitis b virus in taiwan. J Med Virol1979; 3: 237–241.
48.
JiangHFineJPChappellR. Semiparametric analysis of survival data with left truncation and dependent right censoring. Biometrics2005; 61: 567–575.
49.
de WreedeLCFioccoMPutterH, et al. MSTATE: an r package for the analysis of competing risks and multi-state models. J Stat Softw2011; 38: 1–30.
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.