Abstract
The problem of misclassification in covariates is ubiquitous in survival data and often leads to biased estimates. The misclassification simulation extrapolation method is a popular method to correct this bias. However, its impact on Weibull accelerated failure time models has not been studied. In this paper, we study the bias caused by misclassification in one or more binary covariates in Weibull accelerated failure time models and explore the use of the misclassification simulation extrapolation in correcting for this bias, along with its asymptotic properties. Simulation studies are carried out to investigate the numerical properties of the resulting estimator for finite samples. The proposed method is then applied to colon cancer data obtained from the cancer registry at Memorial Sloan Kettering Cancer Center.
Keywords
Introduction
Misclassification of binary covariates occurs frequently for survival analysis in epidemiological data. For example, dietary intake is misclassified based on food frequency questionnaires;1,2 diagnosis of dental caries may be misclassified due to geographical differences in clinical practice; 3 socioeconomic status may be misclassified based on recall data as opposed to the use of historical records; 4 prenatal smoking status may be misclassified in gestational studies. 5 It is well known that the analysis based on the misclassified covariates will lead to biased estimation, especially when the misclassification rate is high.
Several methods have been investigated for handling misclassification for survival data analysis, such as probabilistic bias analysis, 6 weighted least squares method, corrected score function method, pseudo-partial likelihood method,7–9 pooled estimation method10,11 the regression calibration approach11,12 and misclassification simulation extrapolation (MC-SIMEX) method. Additionally, purely numerical studies have also been conducted. 13 Among these methods, the MC-SIMEX method gained popularity recently due to its flexibility, easy implementation, robustness and minimal assumption requirement. 14 It is an extension of the simulation extrapolation (SIMEX) method 15 for measurement error in continuous covariates and first proposed by Kuchenhoff et al. 3 for regression models for complete data. Then it is studied for different models in survival data analysis. Bang et al., 11 studied the performance of the MC-SIMEX method in a Cox model. The performance of MC-SIMEX in log-normal accelerated failure time (AFT) models was studied by Slate and Bandyopadhyay 14 and in log-logistic AFT models by Sevilimedu et al. 16 However, it has not been explored for Weibull AFT model although the Weibull distribution is a very popular distribution to model monotonic hazard functions in survival data analysis. In addition, none of the existing works11,14,16 provided details about the asymptotic properties of the MC-SIMEX estimator in the corresponding survival models. Furthermore, they did not investigate the form of the extrapolation function used in the MC-SIMEX method.
Our work is motivated by colon cancer data from the Memorial Sloan Kettering Cancer Center (MSKCC) registry, where the survival times of patients may follow a Weibull distribution. We are interested in the effect of misclassification caused by faulty categorization of pathological nodal (pathn) status, which plays a pivotal role in determining survival in colon cancer patients. In this paper, we explore the performance of the MC-SIMEX 3 method for misclassification correction in Weibull survival data, while also investigating its asymptotic properties and the form of the extrapolation function. These asymptotic properties of the MC-SIMEX estimator and the form of the extrapolation function are applicable to general AFT models, in which the data can follow any specified distributions. In the simulation studies, we also investigate the robustness of the proposed method to different scenarios such as misspecification of survival time distribution, dependent censoring, presence of multiple categories in covariates and the presence of multiple covariates that may be subject to misclassification.
Methodology
Weibull AFT model and misclassification matrix
Let Ti and Ci denote the failure and censoring times respectively. The observations are
When the true covariate Xi is observed, the likelihood function of the
The maximum likelihood estimator of
The use of the observed covariate
Here, we provide an overview of the MC-SIMEX procedure for the estimation of
Simulation component: For a fixed set of values of
Extrapolation component: Let
Properties of the Mc-SIMEX estimator
In this section, we will provide the properties of the MC-SIMEX estimator. They are applicable to general AFT models, in which the survival time can follow any specified distributions. The Weibull AFT model, in which the survival time T is specified to follow Weibull distribution, is one of the general AFT models.
Asymptotic properties of
Under regularity conditions for the maximum likelihood estimator, the MC-SIMEX estimator,
The form of extrapolation function in Mc-SIMEX method
To investigate the form of the extrapolation function in the MC-SIMEX method, we prove the relationship between the naïve parameter
The components of

Figure demonstrating the behavior of
Settings
In this section, we compare the new proposed estimator, referred to as the Weibull-Q MC-SIMEX estimator (Q denoting a quadratic form of extrapolation), with the Cox MC-SIMEX estimator proposed by Bang et al.,
11
Weibull-L MC-SIMEX estimator (L denoting linear form of extrapolation), Weibull-C MC-SIMEX estimator (C denoting cubic form of extrapolation), true estimator which uses the correct value of the covariate and the naïve estimator which does not correct for misclassification for finite samples. We consider a wide range of scenarios. The objective of the first simulation study is to assess the performance of the Weibull-Q MC-SIMEX estimator with different values of scale parameter (
For the proposed Weibull-Q MC-SIMEX estimator, Cox MC-SIMEX estimator, Weibull-L MC-SIMEX estimator and Weibull-C MC-SIMEX estimator, a total of 50 pseudo-datasets are generated (B) at each level of error inflation with
One misclassified covariate
In the first set of simulation studies, we investigate the performance of the Weibull-Q MC-SIMEX estimator in the following AFT model:
Table showing the performance of the true
Sample size is set to 500. Scale parameter is denoted by σ.
From the perspective of bias and empirical variance, the true estimator generally has the smallest value as it uses the true covariate values, while the naïve estimator has the largest value as it uses the wrong covariate values. Both the Weibull-Q MC-SIMEX estimator and the Cox MC-SIMEX estimator have smaller bias and empirical variance than the naïve estimator, which shows that the MC-SIMEX method demonstrates a much-improved performance. The Weibull-Q MC-SIMEX estimator has a smaller bias and empirical variance than the Cox MC-SIMEX estimator, which indicates Weibull-Q MC-SIMEX method has a better correction for the misclassification. The estimated variance is usually smaller than the empirical variance for the Weibull-Q MC-SIMEX method and the Cox MC-SIMEX method, especially when the scale parameter is large. This may be due to the unknown extrapolation function for the variance. In terms of the coverage probability, the true estimator has coverage close to the nominal level as expected. The Weibull-Q MC-SIMEX method always has a better coverage probability than the Cox MC-SIMEX method.
Overall, the Weibull-Q MC-SIMEX method performs better than the Cox MC-SIMEX method in terms of bias, empirical variance, and coverage probabilities. Comparing the Weibull MC-SIMEX estimator with different extrapolation functions, the Weibull-C MC-SIMEX estimator always has the smallest bias and the Weibull-L MC-SIMEX estimator always has the largest bias. The Weibull-C MC-SIMEX estimator always has a much larger empirical variance than the other two estimators. The Weibull-Q MC-SIMEX estimator is better than the Weibull-L MC-SIMEX estimator except for the case with a small-scale parameter. As for the coverage probability, the Weibull-Q MC-SIMEX estimator performs robustly well than the other two estimators, except that it has a slightly smaller coverage than the Weibull-L MC-SIMEX estimator for the case with a small scale parameter. This indicates that the quadratic extrapolation function provides a good compromise from the perspective of bias-variance trade-off, which is supported by the findings in Section 3.2.
In this setting, we explore the performance of the proposed estimator when the distribution of survival time is misspecified as a Weibull distribution while its original distribution is exponential, log-normal or log-logistic. The survival times for the mentioned distributions are generated using model (3) with scale parameter σ = 1. We simulate the log normally distributed survival data by generating the error term from the normal distribution (0,1) or normal distribution (0,2). The log-logistic survival data is simulated by generating the error terms from the logistic distribution (0,1) or logistic distribution (0,2). The censoring times are generated in a similar manner to that described in Section 4.2. The true values of the binary covariate X are generated from a random binomial distribution with a mean of 0.5. The results are given in Table 2.
Table showing the performance of the true
, naïve
, Cox misclassification simulation extrapolation (MC-SIMEX)
, Weibull-Q MC-SIMEX
, Weibull-L MC-SIMEX
and Weibull-C MC-SIMEX
estimators with misclassified covariate X, when the true underlying distribution is misspecified as a Weibull distribution. Sample size is set to 500.
Table showing the performance of the true
For exponential survival data, the Weibull-Q MC-SIMEX method generally outperforms the Cox MC-SIMEX method, the Weibull-L MC-SIMEX method, and the Weibull-C MC-SIMEX method in terms of bias, empirical variance, and coverage probability. However, it has a slightly larger bias than the Weibull-C MC-SIMEX method. The estimated variance is usually smaller than the empirical variance for the MC-SIMEX-related methods.
For log-normal survival data, the performance of the Cox MC-SIMEX estimator is not stable. When the data variance is small, i.e., normal (0,1), it performs very well, that is, it has the smallest bias (even smaller than the true estimator), smaller empirical variance and better coverage probability than the Weibull-Q MC-SIMEX estimator. However, when the data variance is large, i.e., normal (0,2), it performs very poorly, that is, it has the largest bias (even larger than the naïve estimator), much larger empirical variance, and worse coverage probability than the Weibull-Q MC-SIMEX estimator. However, the performance of the Weibull-Q MC-SIMEX method is reasonably well and robust Similar to the results in Section 4.2, the estimated variance is usually smaller than the empirical variance for the Cox MC-SIMEX estimator and the Weibull-Q MC-SIMEX estimator due to the unknown extrapolation function of the variance. The Weibull-Q MC-SIMEX method always has smaller bias than the Weibull-L MC-SIMEX method and comparable bias to the Weibull-C MC-SIMEX method. It has a smaller empirical variance and better coverage probability than the Weibull-L MC-SIMEX method for large samples (n = 1000) and small sample (n = 500) with small variance and the Weibull-C MC-SIMEX method. This again supports the finding in Section 3.2.
For log-logistic survival data, the Cox MC-SIMEX estimator performs worse than for log normal survival data in terms of bias, empirical variance and coverage probability, while the Weibull-Q MC-SIMEX method still performs reasonably well. The Weibull-L MC-SIMEX method has a larger bias than the Weibull-Q MC-SIMEX method, although their performances are similar in terms of empirical variance and coverage probability. Although the Weibull-C MC-SIMEX method has a smaller bias than the Weibull-Q MC-SIMEX method, it has much worse empirical variance and coverage probability.
Overall, for the distributions considered in this section, the Weibull-Q MC-SIMEX method performs reasonably well and in a robust manner. Other methods, such as the Cox MC-SIMEX method, the Weibull-L MC-SIMEX method and the Weibull-C MC-SIMEX method have poor performance in some settings and perform generally worse than the Weibull-Q MC-SIMEX method.
Table showing the performance of the true
, naïve
, Cox misclassification simulation extrapolation (MC-SIMEX)
, Weibull-Q MC-SIMEX
, Weibull-L MC-SIMEX
and Weibull-C MC-SIMEX
estimators with misclassified covariate X, when the censoring mechanism is dependent upon X. Sample size is set to 500.
Table showing the performance of the true
Scale parameter is denoted by σ.
The patterns of the results are similar to those in Section 4.2. It shows that dependent censoring mechanism does not affect the performance of these methods.
In real data analysis, it is often the case that categorical covariates have more than 2 levels, in contrast to what has been described so far. In order to mimic this scenario and evaluate the proposed estimator in such a scenario, we assumed that the covariate X came from a multinomial distribution with three categories, which are 0,1, and 2, each with a probability of 0.33. The misclassification matrix was set to
The resulting AFT model is therefore:
Table showing the performance of the true
, naïve
, Cox misclassification simulation extrapolation (MC-SIMEX)
, Weibull-Q MC-SIMEX
, Weibull-L MC-SIMEX
and Weibull-C MC-SIMEX
estimators when the misclassified covariate X has two categories.
Table showing the performance of the true
Sample size is set to 500. Scale parameter is denoted by σ. Estimated variance 1 is used to denote the estimated variance of the first category of X and estimated variance 2 is used to denote the estimated variance of the second category of X. Notation is defined similarly for other parameters.
The comparisons of the Weibull-Q MC-SIMEX estimator with other estimators are similar to those in Section 4.2 in terms of bias and empirical variance. From the perspective of coverage probability, when sample size is small, i.e., n = 500, the Weibull-Q MC-SIMEX estimator is better than the Cox MC-SIMEX estimator for σ = 0.5; they perform similarly for σ = 1.5 and the Cox MC-SIMEX is better for σ = 2. When n = 1000 (Table s4 in Supplemental material), the Weibull-Q MC-SIMEX estimator has similar coverage probability to Cox MC-SIMEX estimator when σ = 0.5 and better coverage probability for σ = 1.5. Although the Cox MC-SIMEX estimator is still better for σ = 2, the difference of the coverage of these two methods is smaller than that when n = 500. This indicates that when the misspecified covariate has more categories, a larger sample size is required for the Weibull-Q MC-SIMEX estimator to have better coverage probability than the Cox MC-SIMEX estimator.
In this section, we evaluate a setting wherein two binary categorical covariates are subject to misclassification. Therefore, the resulting model is:
Table showing the performance of the true
, naïve
, Cox misclassification simulation extrapolation (MC-SIMEX)
, Weibull-Q MC-SIMEX
, Weibull-L MC-SIMEX
and Weibull-C MC-SIMEX
estimators when there are two misclassified covariates X1 and X2.
Table showing the performance of the true
Sample size is set to 500. Scale parameter is denoted by σ. Estimated variance 1 is used to denote the estimated variance of X1 and estimated variance 2 is used to denote the estimated variance of X2. Notation is defined similarly for other parameters.
Overall, the Weibull-Q MC-SIMEX estimator generally performs better than the Cox MC-SIMEX estimator for Weibull or other survival data with one or more misclassified covariates. It is robust to all scenarios considered in the simulation studies. Moreover, the Weibull-Q MC-SIMEX estimator performs better than the Weibull-L MC-SIMEX estimator and the Weibull-C MC-SIMEX estimator in terms of the bias-variance trade-off, which supports the theoretical findings in Section 3.2.
Real data analysis with one misclassified covariate
We illustrate an application of the Weibull-Q MC-SIMEX method to a colon cancer dataset containing survival information of 1327 patients who suffered from Stage I to Stage III colon cancer, for which they were surgically treated at MSKCC from January 1990 to December 2000. The dataset was obtained from the MSKCC registry. The dataset provides information on variables such as age and prognostic indicators such as T-stage, N-stage, lymphovascular invasion (LVI), and peri-neural invasion (PNI).
21
In colon cancer patients it has been established that the pathn status of a patient is often misclassified and that the true pathn status (a binary covariate) is a function of the number of nodes examined. Therefore, we estimated the true pathn status of a patient using a statistical algorithm proposed by Gonen et al.
22
We then used this true pathn status to estimate the misclassification matrix as
Of the total of 1327 patients who underwent surgery, 12 were eliminated from the dataset due to ambiguous values on the total nodes examined. The median follow-up time in the cohort was 55 months with date of surgery being treated as the starting time point. 163 of the 1315 individuals died due to colon cancer and the rest were lost to follow-up or died of other causes. Disease-specific survival from colon cancer was the primary outcome of interest The distribution of survival times was checked using the fitdistrplus 23 package in R 4.2. 24 The Weibull distribution appeared to fit the data well based on preliminary visual examination of the QQ plots, probability plots, cumulative density plots, and probability density plots of survival times (Figure 2). Median survival time was not reached (Figure 3).

Figure showing probability density function, q–q plot, cumulative distribution function (CDF) and P–P plot assuming that the distribution of survival times follows Weibull distribution.

Km plot of survival times in the colon cancer dataset with the X-axis representing time in months and the y-axis representing survival probability.
We regressed the survival time of individuals on the following covariates: age, T stage, LVI, PNI, and assigned pathn status using an AFT model. The variable pathn status is considered as the covariate which is subject to misclassification. The AFT model takes the form:
Due to lack of proper validation data (as the true pathn status could only be estimated), we do not provide results for the true estimator. For the Weibull-Q MC-SIMEX, Weibull-L MC-SIMEX, Weibull-C MC-SIMEX and Cox MC-SIMEX procedures, we ran a total of 50 replications of pseudo-dataset generation for each level of
Table showing the estimates of the coefficient associated with pathological nodal status obtained by fitting the AFT model to colon cancer data.
MC-SIMEX: misclassification simulation extrapolation; AFT: accelerated failure time.
All the estimators shown in Table 6 indicate the detrimental effect of a positive pathn status on survival. It can be deciphered from the Weibull-Q MC-SIMEX estimator that patients with a negative pathn status have a survival time which is 2
Even though data fits Weibull distribution well as shown in Figure 2, it can be misspecified and the original distribution can be other distributions. Then the commonly used statistical criteria such as AIC, BIC, and loglikelihood are implemented to check all distributions considered in Section 4.3, The results are shown in Table 7.
Table showing the fit statistics of the accelerated failure time (AFT) regression model under different distributional assumptions using AIC, BIC, and log-likelihood.
As there are only minor differences in AIC, BIC, and log-likelihood for the distributions considered in Section 4.3, any distribution for the data is justifiable. Under such conditions, we have demonstrated via simulations in Section 4.3 that a Weibull-Q MC-SIMEX estimator provides the most reliable parameter estimates in comparison to Cox MC-SIMEX estimator, except for log-normally distributed data with lower variance, for which the Cox MC-SIMEX estimator has less bias. We estimated the variance of the log of survival times in the current dataset to be 0.80. Given that this variance is close to that of standard log-normal distribution, it will be fair to say that if this data follows log-normal distribution, Cox MC-SIMEX estimator may have smaller bias than the Weibull-Q MC-SIMEX estimator as shown in the simulation study in Section 4.3. However, since the Weibull-Q MC-SIMEX estimator performs generally well for all distributions, we recommend using the results from Weibull-Q MC-SIMEX method to make conclusions for this data. Because the log-normal distribution may best fit the data based on the statistics in Table 7, the log-normal MC-SIMEX method may provide more accurate results. However, simulations are needed to compare the performance of log-normal MC-SIMEX method with other methods. Future research could investigate the performance of the MC-SIMEX method with different distributions, and therefore yield more guidelines for real dataanalysis.
In this section, we would like to illustrate the application of Weibull-Q MC-SIMEX method for the colon cancer data considered in Section 5.1 with two misclassified covariates. In order to examine the performance of the proposed estimator in this situation, we simulated a misclassified version of the variable LVI using the misclassification matrix
Table showing the estimates of the coefficient associated with pathological nodal status and mLVI.
MC-SIMEX: misclassification simulation extrapolation; mLVI: misclassified version of lymphovascular invasion
The result shows the coefficient estimate for pathn status are very similar to what was seen in Section 5.1. With regard to the coefficient estimates for mLVI, we now already know that the Weibull-Q MC-SIMEX estimator is the most reliable method based on the results of the simulation studies. It concludes that LVI is not a significant covariate associated with survival time of colon cancer patients. Although other methods have different corresponding estimates when compared to that of the Weibull-Q MC-SIMEX estimator, they still reach the same conclusion as that of the Weibull-Q MC-SIMEX estimator.
In this article, we investigate the performance of the Weibull-Q MC-SIMEX estimator in an AFT model in which the survival times followed a Weibull distribution and are subject to right censoring. We demonstrate that under varying scenarios of censoring and misclassification, the Weibull-Q MC-SIMEX estimator is robust and performs better than the naïve estimator and the existing Cox MC-SIMEX estimator.
Additionally, we also provide proof for the relationship between the parameter estimates obtained at each level of error inflation in the presence of a correctly measured confounder and demonstrate graphically and using simulation studies that the quadratic extrapolation function will offer the best fit in a Weibull AFT model. This finding highlights the importance of the choice of the extrapolation function when using the MC-SIMEX method. Colon cancer is one of the leading causes of death in the US. Therefore, it is important that while estimating hazards and time to survival among patients with risk factors, potential misclassification must be corrected in the covariates. Through this study, we show that the Weibull-Q MC-SIMEX estimator could be used to obtain accurate predictions of time to survival, by correcting for misclassification.
Potential future directions include investigating the performance of nonparametric extrapolation techniques such as splines in the MC-SIMEX method, application of the procedure to situations where the correctly measured confounder and the misclassified covariate interact with each other, application of the procedure to a competing risks model and extension of the method to other fields such as mediation analysis.
Supplemental Material
sj-docx-1-smm-10.1177_09622802231168248 - Supplemental material for Misclassification simulation extrapolation method for a Weibull accelerated failure time model
Supplemental material, sj-docx-1-smm-10.1177_09622802231168248 for Misclassification simulation extrapolation method for a Weibull accelerated failure time model by Varadan Sevilimedu, Lili Yu and Hani Samawi in Statistical Methods in Medical Research
Footnotes
Acknowledgments
The authors are grateful to Dr Mithat Gonen, Attending Biostatistician and Chief of Biostatistics Service, MSKCC, for providing us with the colon cancer data.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental material
Supplemental material for this article is available online.
Appendix
Derivation of equation 2:
Derivation of
We define:
Derivation of
Along similar lines, as previously stated in the derivation section for
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
