Abstract
Marginalized two-part random-effects generalized Gamma models have been proposed for analyzing medical expenditure panel data with excessive zeros. While these models provide marginal inference on expected healthcare expenditures, the usual unilateral specification of heteroscedastic variance on one of the two shape parameters for the generalized Gamma distribution in these models fails to encompass important special cases within the generalized gamma modeling framework. In this article, we construct marginalized two-part random-effects models that employ the log-normal, log-skew-normal, generalized Gamma, Weibull, Gamma, and inverse Gamma distributions to delineate the spectrum of nonzero healthcare expenditures in the second part of the models. These marginalized models supply additional choices for analyzing healthcare expenditure panel data with excessive zeros. We review the concepts of marginal effect and incremental effect, and summarize how these effects are estimated. For studies whose primary goal is to make inference on marginal effect or incremental effect of an independent variable with respect to healthcare expenditures, we advocate empirical mean square error criterion and information criteria to choose among candidate models. Then, we use the proposed models in an empirical analysis to examine the impact of the New Cooperative Medical Scheme on healthcare expenditures among older adults in rural China.
Keywords
1 Introduction
Non-negative semicontinuous healthcare expenditure data1–3 collected in health economics and health services research studies are right-skewed with a heavy right tail, heteroscedastic in variance with respect to some covariates, and contain excessive zeros observed from the non-users of health facilities. In addition, the data collected in these studies are often panel data, because the studies are designed to repeatedly collect healthcare expenditures from the same cross section of individuals over time. A number of sophisticated statistical models have been proposed in the literature to analyze this type of expenditure data.
Conditional two-part random-effects models that incorporate latent random effects into the analysis of healthcare expenditure panel data were first reported by Olsen and Schafer 4 and Tooze et al. 5 Olsen and Schafer 4 and Tooze et al. 5 extended the two-part regression models for cross-sectional data to the settings of panel data by including unobserved random coefficients into both parts of the models. Conditional two-part random-effects models are those two-part models with a random-effects logistic or probit model for the probability of observing a panel positive expenditure in Part (I) and a generalized linear random-effects model for the observed panel positive healthcare expenditures in Part (II). The essential assumption on these models is that the random effects from Part (I) and the random effects from Part (II) are correlated. As such, the probability of observing a positive expenditure in Part (I) is correlated to the amount of the positive expenditure at a specific time point. However, Olsen and Schafer 4 and Tooze et al. 5 both adhered to a log-normal regression model in Part (II) of the proposed two-part models. Later, Manning et al. 6 showed that it is more efficient to use the three-parameter generalized Gamma distribution to model the positive healthcare expenditures. The generalized Gamma distribution has been proven superior to the log-normal distribution in Olsen and Schafer 4 and Tooze et al., 5 because it includes the standard Gamma, inverse Gamma, Weibull, exponential, and lognormal distributions as its special cases and consequently can provide much more flexibility in the circumstances where the above distributions cannot adequately characterize the positive expenditures. Manning et al. 6 also extended the traditional generalized Gamma regression model, so that it allows for the existence of substantial heteroscedasticity among the expenditures. Following Manning et al., 6 Liu et al. 7 and Liu et al. 8 constructed several conditional two-part random-effects models for healthcare panel expenditures, in which the lognormal random-effects model developed by Olsen and Schafer 4 and Tooze et al. 5 in Part (II) of the models was extended to be a log-skew-normal or a generalized Gamma random-effects model with a regression model for a scale parameter that allowed for heteroscedastic variance. Conditional two-part random-effects models in Liu et al. 7 and Liu et al. 8 have an inherited drawback that it is inconvenient to make direct marginal inference with respect to the positive or overall panel healthcare expenditures. To overcome this drawback, Zhang et al. 9 developed two types of marginalized two-part random-effects generalized Gamma models (abbreviated as “generalized Gamma m2REMs”; “m2REMs” stands for “marginalized two-part random-effects models”): the Type I generalized Gamma m2REM was constructed for marginal inference on positive healthcare expenditures and the Type II generalized Gamma m2REM was for marginal inference on overall healthcare expenditures.
This article is devoted to developing the marginalized two-part random-effects models, in which the specification of generalized Gamma distribution in Part (II) is replaced by the specification of log-normal, log-skew-normal, Gamma, and inverse Gamma distribution. The motivation to further develop these models is that, because of the particular setup in the heteroscedastic variance model for one of the two shape parameters in the generalized Gamma m2REMs in Zhang et al.,
9
the generalized Gamma m2REMs do not include the m2REMs with other distribution specification as their special cases anymore. The generalized Gamma distribution has been reviewed in Zhang et al.
9
Note that, in the generalized Gamma distribution, there are two shape parameters σ and κ that dominate the variance function of the distribution. In Zhang et al.,
9
the heteroscedastic variance among positive panel healthcare expenditures is characterized by one of the shape parameters through
The scientific contributions of this article are multi-fold. Besides the development of the new m2REMs in Section 2 as the alternatives of the generalized Gamma m2REMs, we also describe in Section 4 how to derive estimates and variance estimates for marginal and incremental effects of a covariate with respect to both overall and positive healthcare panel expenditures within the framework of the m2REMs. Given these multiple choices for modeling healthcare expenditure panel data, we propose an approach to select the best model based upon the observed data; that is, an effect-specific selection approach introduced by Dow and Norton 11 is modified in Section 5 to make it suitable to selecting among the m2REMs. The other contribution of this article is our empirical analysis and re-examination on the impact of the New Cooperative Medical Scheme (NCMS)12,13 on individual healthcare expenditures in China. The New Cooperative Medical Scheme (NCMS) is a health insurance program that was launched in 2003 by the Chinese central government for the rural residents living in China (Chinese citizens are separated into either rural or urban residents by their residency registration). It was reported that the enrollment of the NCMS rose to 97% of rural population in 2011.14,15 Lei and Lin 12 explored the influence of the NCMS on healthcare expenditures among rural residents in China, and did not find that the NCMS decreases individual healthcare expenditures. Here, we investigate the impact of the NCMS on a subset of rural residents who are older than 60 years old in China. A set of panel expenditure data among the elders is extracted from a carefully designed U.S.–China collaborative large-scale survey, the China Health and Nutrition Survey. The data are analyzed using the proposed m2REMs in Section 2. The results show that, even for the elders who lived in rural areas in China, the NCMS is not a significant factor that can decrease their healthcare expenditures.
2 Modeling framework
2.1 Marginalized two-part random-effects log-normal (log-skew-normal) models
Let Yit be a semicontinuous dependent variable, representing the amount of healthcare expenditure of the ith individual (
In equation (2),
This can be solved for Δ
it
using numerical integration combined with the Newton–Raphson or quasi-Newton algorithm. A close-form solution may exist for Δ
it
in some special cases: when
In Part (II) of the Type I log-normal m2REM, a marginally specified model for the mean of the positive healthcare expenditure is formulated as
Therefore, equations (3) to (5) lead to
Equations (1) to (5) together constitute the Type I log-normal m2REM. Here, the random effects ait in equation (2) and bit in (4) are assumed to be correlated and follow a multivariate normal distribution so that the two parts are correlated; that is, we assume generally
The Type I log-normal m2REM aims at providing direct marginal inference with respect to the marginal mean of positive healthcare expenditures over time, but not with respect to the marginal mean of overall healthcare expenditures that includes both the zero expenditures from the non-users of healthcare facilities and the positive expenditures from the users. To achieve the goal of making direct marginal inference on overall healthcare expenditures, we can use the following marginally specified model (8) to replace equation (3) to construct the Type II marginalized two-part random-effects log-normal model (abbreviated as Type II log-normal m2REM). In the Type II log-normal m2REMs, the positive expenditures in Part (II) follow a heteroscedastic log-normal distribution
In equation (8),
Therefore, in the Type II log-normal m2REM, Λ
it
in equation (4) has a form of
Model specification of Type I and Type II m2REMs.
The above Type I and Type II log-normal m2REMs can be extended to Type I log-skew-normal m2REM and Type II log-skew-normal m2REM, in which the positive expenditures in Part (II) follow a heteroscedastic log-skew-normal distribution. These models can be constructed by using a log-skew-normal regression model to replace the log-normal regression model in Part (II) of the Type I and Type II log-normal m2REMs. For a variable Y, if Y follows a log-skew-normal distribution, then
Combined with
To construct the Type II log-skew-normal m2REM, we take equations (1), (2), (4), (5) and (8), but assume
Combined with
Model specification of the Type I and Type II log-skew-normal m2REMs is also summarized in Table 1.
2.2 Marginalized two-part random-effects generalized Gamma (Weibull) models
The generalized Gamma distribution is a continuous probability distribution with one scale parameter and two shape parameters.
6
For a non-negative response variable Y, if Y follows a generalized Gamma distribution with scale parameter μ and shape parameters σ and κ, then
The generalized Gamma distribution is very flexible. It includes Gamma, inverse Gamma, Weibull, exponetial, and log-normal distributions as its special cases (see Table 1 in Manning et al.
6
for details). The Type I and Type II marginalized two-part random-effects models with the positive healthcare expenditures in Part (II) following a heteroscedastic generalized Gamma distribution (abbreviated as Type I and Type II generalized Gamma m2REMs) were introduced by Zhang et al.
9
Part (I) of the Type I generalized Gamma m2REM is constructed by equations (1) and (2) as in the log-normal m2REMs in Section 2.1, but Part (II) of the model takes equation (3) and assumes
This, combined with
For the Type II generalized Gamma m2REM, it can be derived from
This, combined with
Weibull distribution is a special case of the generalized Gamma distribution when κ = 1. Therefore, by fixing κ = 1, the Type I and Type II generalized Gamma m2REMs degenerates to Type I and Type II Weibull m2REMs, respectively. Model specification of the Type I and Type II generalized Gamma and Weibull m2REMs is also summarized in Table 1.
2.3 Marginalized two-part random-effects (inverse) Gamma models
The generalized Gamma distribution (13) includes Gamma and inverse Gamma distributions as its special case. When
When
However, it should be noted that the Type I and Type II generalized Gamma m2REMs do not cover the Type I and Type II Gamma or inverse Gamma m2REMs, in which the positive healthcare expenditures in Part (II) follow a heteroscedastic Gamma or inverse Gamma distribution. The reason is that, in the Type I and Type II generalized Gamma m2REMs, the shape parameter κ or its negative number
Part (I) of the Type I Gamma m2REM is constructed by equations (1) and (2) as in the log-normal m2REMs in Section 2.1, but Part (II) of the model takes equation (3) and assumes
This, combined with
Part (I) of the Type II Gamma m2REM is constructed by equations (1) and (2), but Part (II) of the model takes equation (8) and assumes
This, combined with
To formulate the Type I and Type II inverse Gamma m2REMs, some adjustment on equation (5) is required. Note that, if a non-negative response variable Y follows an inverse Gamma distribution with a density
This, combined with
Part (I) of the Type II inverse Gamma m2REM is also constructed by equations (1) and (2), but Part (II) of the model takes equation (8) and assumes
This, combined with
Model specification of the Type I and Type II Gamma and inverse Gamma m2REMs is summarized in Table 1.
3 Maximum likelihood estimation and simulation studies
For all of the Type I and Type II m2REMs introduced in Section 2, let
In the likelihood function (19), we assume the observed panel data are balanced (i.e. the panel data are collected on every participant in the sample at each time of the study). When the panel data are unbalanced (i.e. a participant’s response can be missing at one follow-up time and then may be measured at a later follow-up time), it is requires to slightly modify the likelihood function (19) to be precise. When the missing responses are missing completely at random (the probability that responses are missing does not depend on either the observed or the missing responses) or missing at random (the probability that responses are missing depends on the observed responses but is unrelated to any missing responses), the full likelihood of the m2REMs is
True values, parameter estimates, standard deviations (SD), and standard errors (SE) obtained in the simulation study for the Type I and Type II m2REMs with various parametric distributions (log-normal, log-skew-normal, Weibull, Gamma, and inverse Gamma distributions) in the positive healthcare expenditures in Part (II).
4 Estimation of marginal and incremental effects
The quantities of interest in the empirical analysis of panel data on healthcare expenditures are often the marginal or increment effects of an independent variable on expected healthcare expenditure at a specific time point.20,21 Zhang et al. 9 derived specific formulas for estimating these effects in the context of Type I and Type II generalized Gamma m2REMs. These formulas can be directly applied to the Type I and Type II m2REMs introduced in Section 2 without modification, and therefore are summarized below.
Let yit be the semicontinuous dependent variable that represents the amount of the healthcare expenditure of the ith individual at the time t. Let
The above marginal effect
The average incremental effect
Obtained by averaging the individual marginal effects are the estimators of average marginal and incremental effects
When any of the Type I m2REMs is used, it is convenient to estimate the marginal and incremental effects of an independent variable with respect to the expected positive healthcare expenditure
It quantifies the marginal change in the expected positive healthcare expenditure with a small amount of increase in xitk while holding other independent variables
The average marginal and incremental effects on the expected positive healthcare expenditures are subsequently given by
Zhang et al. 9 elaborated the procedure of variance estimation of the (average) marginal and incremental effects for conducting inference on these effects, when using either the Type I and Type II m2REMs. Variance estimators for these effects were derived using the delta methods or Taylor series approximations, and the details were reported in Zhang et al. 9 In some empirical analysis of healthcare expenditures, (average) elasticity and semi-elasticity are also quantities of interest. Please refer to Zhang et al. 9 for the specifics of estimation and variance estimation of these quantities.
5 Model comparison
In the m2REMs introduced in Section 2, the log-normal, log-skew-normal, generalized Gamma, Weibull, Gamma, and inverse Gamma distributions are employed to delineate the spectrum of the positive healthcare expenditures in Part (II) of the models. These marginalized models supply multiple choices for analyzing the healthcare expenditure panel data with excess zeros. For the studies in which the primary interest of investigation is to draw inference on the marginal or incremental effect of an independent variable with respect to the expected overall or positive expenditures, the empirical mean square error (MSE) criterion has been proven effective to choose among multiple candidate models.
11
While estimating the (average) marginal and incremental effects is the aim of investigation, the MSEs of the estimators of (average) marginal and incremental effects are the variance of the estimators plus the square of their bias
The MSE criterion selects the model with the minimum MSE as the best model among the candidates. However, in practice, the true effects
The information criteria, such as the Akaike information criterion (AIC) and the Bayesian information criterion, are also advocated here to conduct model selection among the m2REMs. These criteria are not effect-specific and thus may conclude differently with the empirical MSE criterion.
6 Empirical analysis: estimating incremental effects of the New Cooperative Medical Scheme on healthcare expenditures among older adults in China
6.1 Data and variables
A set of panel data on the healthcare expenditures among older (60 years old or older) rural residents in China were analyzed using the m2REMs introduced in Section 2. The primary objective of this empirical analysis is to evaluate the impact of the New Cooperative Medical Scheme (NCMS) program on healthcare expenditures among older adults in rural areas in China. The panel data were taken from the China Health and Nutrition Survey (CHNS), a collaborative project between the Carolina Population Center at the University of North Carolina at Chapel Hill and the National Institute for Nutrition and Health at the Chinese Center for Disease Control and Prevention. The CHNS is designed to investigate the effects of health and nutrition programs implemented by Chinese governments throughout the past 20 years, in order to determine how the social and economic transformation in China impacts the health and nutritional status of the population during China’s rapid economic ascendance. Nine waves of the survey were conducted in 1989, 1991, 1993, 1997, 2000, 2004, 2006, 2009, and 2011. The survey implemented a multistage, random cluster process of data collection and by now it has drawn a sample of approximately 7200 households with more than 30,000 the individuals. The household and individual surveys contain several modules. One of the modules is the health services section that collected information on respondent demographics, health, nutrition, and income. Nine provinces (Liaoning, Heilongjiang, Guangxi, Guizhou, Henan, Hubei, Hunan, Jiangsu, and Shandong) with substantial variation in geography, economic development, and public health facilities had participated in the CHNS by 2011. In this empirical analysis, the five most recent waves from 2000 to 2011 (i.e. 2000, 2004, 2006, 2009, and 2011 waves) were used. The initial sample included 74,619 observations from both urban and rural residents who were 18 years old or older. The criterion to distinguish whether a participant is an urban or a rural resident is based upon the resident registration information. The individuals with an agricultural resident registration were classified as rural residents. Because the NCMS is only available to the residents with an agricultural resident registration, we excluded the observations from urban residents, yielding a sample of 29,895 observations from rural residents who participated in the CHNS from 2000 to 2011. The empirical analysis concerns the healthcare expenditures among the older adults. Thus, 23,454 additional observations from the participants that were younger than 60 years old were excluded, which left 6441 observations for further consideration.
In this empirical analysis, the primary dependent variable is the out-of-pocket healthcare expenditures during the previous four weeks of the survey (denoted by Yit for individual i at the tth survey year) reported by the CHNS older participants. The out-of-pocket healthcare expenditure is the summation of three portions: self-care expense on the illness, injury, chronic or acute disease (e.g. the expenditure from purchases of over-the-counter drugs), expense from utilization of formal healthcare facilities including both inpatient and outpatient services (e.g. the expenditure of registration, treatment, medicines, lab tests, surgeries, hospital beds, etc.), and expense from the utilization of preventive health services (e.g. spending on general physical examination, eye exam, blood tests, blood pressure screening, tumor screening, etc.). The expense from utilization of formal healthcare facilities or preventive health services was calculated by multiplying the reported expenditure before insurance coverage by one minus the percentage of the expenditure paid by insurance. The percentage of health insurance coverage was counted as zero if the participant reported having no health insurance coverage or unknown.
The scientific interest of this empirical analysis is to estimate the incremental effect of the NCMS coverage with respect to the overall and positive out-of-pocket healthcare expenditures among older adults in rural China. As such, the primary independent variable of the analysis is a time-varying three-level indicator of health insurance coverage status with 0 representing no health insurance coverage, 1 representing the NCMS coverage, and 2 representing the coverage from other types of health insurance. When the CHNS inquires about the type of health insurance of the participants, it adopted two slightly different categorization approaches. The 2009 and 2011 questionnaires asked the participants to choose one insurance type, if they were covered by health insurance, from the following options: the NCMS, the Urban Employee Basic Medical Insurance, 15 the Urban Resident Basic Medical Insurance, 15 any commercial medical insurance, government free medical insurance, and other types of health insurance. Thus, these questionnaires clearly distinguished the NCMS from other type of health insurance coverage. However, the 2006 questionnaire provided the following options to the participants who were covered by health insurance: the Cooperative Medical Scheme (CMS), the Urban Employee Basic Medical Insurance, the Urban Resident Basic Medical Insurance, any commercial medical insurance, government free medical insurance, health insurance program for women and children, and other types of health insurance; and the 2000 and 2004 questionnaires provided the following options: the CMS, any commercial medical insurance, government free medical insurance, health insurance program for women and children, worker’s compensation insurance, insurance coverage as family members, unified planning medical service, and other types of health insurance. The CMS referred to in the 2000, 2004, and 2006 questionnaires included both the NCMS and the original CMS 14 (the original CMS, or the old CMS, was the major health insurance program for rural residents in China before early 1980s; due to China’s economic reform, it collapsed as a direct consequence of the lost of welfare funds from collective economy). Since the NCMS was launched in 2003, when the CMS was chosen in the 2004 and 2006 CHNS, it may represent either the old CMS or the NCMS. In order to distinguish between these two programs, we adopted the approach introduced by Lei and Lin 12 and use community level data to define the health insurance status for the individuals that chose the CMS. The community survey component asked local government officials whether the CMS had been implemented in their community. It is true that this survey question to the communities was too vague, because it did not specifically distinguish the original CMS or the NCMS. However, because the pilot implementation of the NCMS began in 2003, the communities must implement the NCMS, instead of the old CMS, in 2004 and 2006 if the local government officials reported absence of CMS in the 2000 CNHS survey or before, but notified the establishment of it in 2004 and 2006. Consequently, the individuals who lived in such communities that offered the NCMS and who chose the CMS in the 2004 and 2006 CHNS questionnaires were identified to be covered by the NCMS health insurance program.
In this empirical analysis, the independent variables on individual characteristics include age, gender, marital status, years of schooling, logarithm of one plus household income, and whether the individual reported at least one type of the chronic diseases listed in the survey questionnaire (diabetes, hypertension, myocardial infarction, or apoplexy). The independent variables also include four dummy variables for five survey waves and eight dummy variables for nine provinces. The out-of-pocket healthcare expenditure and household income, both in terms of Ren Min Bi or RMB (RMB is the official currency of China), were converted to the 2011 price level according to the annual customer price indices (CPI) published by the National Bureau of Statistics of China. 22
Detailed description of the number of observations excluded from the original CHNS data.
6.2 Empirical analysis strategy
The strategy of our empirical analysis is to employ the Type I and Type II m2REMs with various distributions for positive healthcare expenditures in Part (II) and estimate the incremental effect of the NCMS health insurance program on both overall and positive out-of-pocket healthcare expenditures. The Type I m2REMs that was used to analyze the CHNS data take the form of
The Type II m2REMs that was used to analyze the CHNS data take the form of
Type I m2REM (25) and Type II m2REM (26) are specific forms of the proposed Type I and Type II m2REMs in Section 2, where the link function specifications are
6.3 Estimation of fixed effects, heteroscedasticity, and variance components
Parameter estimates, estimated standard errors, and p-values from fitting the Type I and Type II log-skew-normal m2REMs to the CHNS data.
Parameter estimates, estimated standard errors, and p-values from fitting the Type I and Type II log-normal m2REMs to the CHNS data.
Parameter estimates, estimated standard errors, and p-values from fitting the Type I and Type II generalized Gamma m2REMs to the CHNS data.
Parameter estimates, estimated standard errors, and p-values from fitting the Type I and Type II Gamma m2REMs to the CHNS data.
Parameter estimates, estimated standard errors, and p-values from fitting the Type I and Type II inverse Gamma m2REMs to the CHNS data.
6.4 Estimation and comparison of incremental effects of the NCMS
The primary objective of this empirical analysis is to examine the incremental effect and average incremental effect of the NCMS, as well as the coverage of other insurance, on both the positive and the overall healthcare expenditures among the older adults who participated in the CHNS. The estimation and variance estimation procedures summarized in Section 4 were applied to assess these effects and their standard errors (SEs). To remove any interference induced by the insignificant independent variables other than the NCMS and other insurance coverage, the incremental and average incremental effects were estimated and further tested under the reduced model in Tables 4–8.
Estimates, estimated standard errors, and p-values of the average incremental effects of the NCMS and other insurance on the expected overall and positive healthcare expenditure, as well as the AIC and the BIC values, given by the Type II and Type I m2REMs.
By assuming the log-normal Type I and Type II m2REMs were consistent and consequently were used as the pre-specified model, the corresponding empirical MSEs of
Under the Type I log-skew-normal m2REM, the incremental effects (not the average incremental effects) of the NCMS and other insurance coverage on the expected positive healthcare expenditure were estimated as –472.45 (SE: 304.26, p-value: 0.12) and –390.49 (459.19, 0.40), respectively, for the diseased older adults and were estimated as –194.42 (122.57, 0.11) and –160.70 (186.08, 0.39), respectively, for the nondiseased older adults. Using the Type II log-skew-normal m2REM, the incremental effects of the NCMS on the expected overall healthcare expenditure were estimated under the six combinations of marriage status, disease status, and the survey year in 2006 versus other survey years. In 2006, the estimates, as well as the estimated SEs and the p-values are –131.08 (SE: 89.44, p-value: 0.14) for the married and diseased older adults, –26.62 (17.75, 0.13) for the married and nondiseased older adults, –82.39(56.33, 0.14) for the diseased older adults with other marriage status, and –16.73(11.20, 0.14) for the nondiseased older adults with other marriage status. In other survey years, the estimates, as well as the estimated SEs and the p-values are –288.42(199.67, 0.15) for the married and diseased older adults, –58.56 (39.63, 0.14) for the married and nondiseased older adults, –181.28 (125.84, 0.15) for the diseased older adults with other marriage status, and –36.80 (25.01, 0.14) for the nondiseased older adults with other marriage status. None of estimated incremental effects of the NCMS is significant on either the overall or the positive expected healthcare expenditures.
7 Concluding remarks
In this article, we explore a useful array of model specifications that may be used as alternatives to the generalized Gamma m2REMs, including log-normal, log-skew-normal, generalized Gamma, Weibull, Gamma, and inverse Gamma distributions in Part (II) of the m2REMs. These alternative models supply additional choices and needed flexibility for analyzing the healthcare expenditure panel data with excess zeros. We present approaches for deriving marginal inference, via marginal and incremental effects, for this set of models and provide guidance on choosing the best fitting model based upon the empirical MSE and information criteria. In the empirical analysis, using the proposed m2REMs, we examine the influence of the NCMS and other insurance coverage on individual healthcare expenditures among the older adults lived in rural China. Our conclusion is in alignment with the finding in Lei and Lin 12 who did not identify any significant impact of the NCMS on the average medical expenditures for residents in China.
A cautionary note for using the m2REMs should be raised when the panel data are unbalanced and the missingness responses are not missing at random (the probability that responses are missing depends on both the observed and the missing responses). With the assumption of not missing at random, we can develop either selection models or pattern-mixture models, combined with the m2REMs, to address such an issue of non-ignorable missingness. 19 Yet, the developed methods must be valid with respect to the potential missing mechanism and sensitivity analysis may be required.
Footnotes
Acknowledgements
We sincerely thank two anonymous reviewers and Editors for their valuable comments, which had substantially improved this article. This research used data from China Health and Nutrition Survey (CHNS).
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Dr Wei Liu’s research was partially supported by National Natural Science Foundation of China (Project 11601106). We thank the National Institute for Nutrition and Health, China Center for Disease Control and Prevention, Carolina Population Center (P2C HD050924, T32 HD007168), the University of North Carolina at Chapel Hill, the NIH (R01-HD30880, DK056350, R24 HD050924, and R01-HD38700) and the NIH Fogarty International Center (D43 TW009077, D43 TW007709) for financial support for the CHNS data collection and analysis files from 1989 to 2015 and future surveys, and the China–Japan Friendship Hospital, Ministry of Health for support for CHNS 2009, Chinese National Human Genome Center at Shanghai since 2009, and Beijing Municipal Center for Disease Prevention and Control since 2011.
