Abstract
This research note contributes to the discussion of methods that can be used to identify useful auxiliary variables for analyses of incomplete data sets. A latent variable approach is discussed, which is helpful in finding auxiliary variables with the property that if included in subsequent maximum likelihood analyses they may enhance considerably the plausibility of the underlying assumption of data missing at random. The auxiliary variables can also be considered for inclusion alternatively in imputation models for following multiple imputation analyses. The approach can be particularly helpful in empirical settings where violations of missing at random are suspected, and is illustrated with data from an aging research study.
Missing data pervade the social and behavioral sciences. Two state-of-the art methods for the analysis of incomplete data sets are maximum likelihood (ML) and multiple imputation (MI; Schafer & Graham, 2002). Most current applications of ML to missing data are based on the assumption that data are missing at random (MAR). MAR holds if the probability of response is not related to the actually missing values (e.g., Little & Rubin, 2002). MAR represents a relatively restrictive assumption in empirical research that is weaker than the popular missing completely at random mechanism. MAR is not statistically testable since it is a statement about data that are actually not available (unobserved). When MAR is violated (i.e., the probability of response is related to the unobserved/missing values), data are said to be not missing at random (NMAR). Similar to MAR, NMAR is also not statistically testable, but based on substantive considerations one may be in a position to argue in favor of NMAR and against MAR.
When data are NMAR, an inclusive analytic strategy has been recommended for dealing with violations of MAR (Collins, Schafer, & Kam, 2001; Little & Rubin, 2002; Schafer & Graham, 2002). This strategy involves the inclusion of auxiliary variables (AVs) in the analyses of an incomplete data set. AVs, broadly speaking, are measures that can “explain missingness” but are not of intrinsic interest in and of themselves. Furthermore, they are not usually included in models of relevance to a research question if there are no missing data (e.g., Enders, 2010). Their implementation in the analysis of an incomplete data set is achieved either by incorporating them as additional predictors of outcome variables, or as “background” variables using the approach in Graham (2003; the AVs are then only allowed to be correlated among themselves, with the independent manifest variables, and with the error terms pertaining to the dependent observed measures). An additional way of using AVs to handle missing data is through their inclusion in imputation models within the MI approach (cf. Little & Rubin, 2002), but such applications will not be of concern in the remainder of this note.
As discussed in detail elsewhere (e.g., Collins et al., 2001; Graham, 2009), researchers should consider using as AVs those variables that “explain missingness” and contribute mostly to enhancing the MAR assumption plausibility. More concretely, effective AVs are either (a) causes of missingness, (b) correlates of missingness, or (c) correlates of outcome variables with missing data (e.g., Graham, 2009). Given this helpful feature of AVs a natural question that arises is how to identify useful, or effective, AVs, that is, AVs that will enhance substantially the MAR assumption if included in subsequent analyses. The literature on analysis of incomplete data sets does not offer clear-cut and definitive guidelines as yet, however, for responding to this query. In the rest of this note, with the intent to contribute to the development of such guides, we will be concerned with an earlier indicated method that could help identify potentially effective AVs (cf. e.g., Enders, 2010). Accordingly, it has been proposed to consider those candidate AVs, on which there are differences across (a) the group of subjects with missing data on the response and (b) the group of subjects with data present on the outcome. These differences can be detected via testing for mean and variance inequalities of the groups, which under normality would capture, as well known, all their distributional differences. 1
Hypothesis testing may not necessarily be helpful for the AV identification, however, particularly with large samples (when statistical power can be excessively high) or small samples (when power may be too low to detect such group differences). These two arguments are part of the more comprehensive criticism that statistical (“nil”) hypothesis testing has received over the past two decades in the behavioral and social sciences (e.g., Harlow, Mulaik, & Steiger, 1997; this criticism applies to “simple null hypotheses,” like the ones of relevance in the remainder). Their implication for the practice of missing data analysis is that with large samples one may proclaim as significant group differences in means and/or variances that are minor and not even interpretable in substantive terms. This would suggest a practically irrelevant variable to be included as an AV in subsequent analyses, thus unnecessarily complicating the ensuing analytic results and/or numerical process leading to them (e.g., Enders, 2010). In the case of small samples, one may run the danger of missing out on one or more potentially effective AVs, hence forfeiting a chance to conduct at least close-to-MAR analyses subsequently on the original incomplete data set when including that variable(s). Either of these possibilities adds further difficulties to the fundamental problem of dealing with missing data, which stems from the fact that potentially very important information about a studied phenomenon is not available.
To respond to this criticism, the present note focuses on an interval estimation method based on the popular latent variable modeling (LVM) methodology (B. O. Muthén, 2002) that can be used to examine mean and variance differences across the missing data and present data groups for a given outcome measure(s). In addition to point estimates, the confidence intervals (CIs) furnished by the method yield sets of plausible values for the mean and variance group differences on a response variable(s) of interest. With this feature, the intervals provide researchers with additional information about these differences, thus allowing them to reach more dependable conclusions about the utility of candidate AVs (see also Note 1). The procedure discussed below can be routinely used by behavioral and social scientists in preliminary analyses intended to provide insights about appropriate avenues of analyzing subsequently an incomplete data set. The approach is readily used in empirical research settings via the popular software Mplus (L. K. Muthén & Muthén, 2012) and R (Venables, Smith, and The R Development Core Team, 2012).
Background and Notation
In the remainder of the note, we assume that p+ 1 continuous (or approximately continuous) measures are collected on a sample of n persons in a study of a population of interest (p, n > 1). Denote these variables by y1, . . . , yp+1, and without limitation of generality assume that one is interested in the last variable (i.e., yp+1), as an outcome or dependent variable for a research question. We also assume that some or all the p variables are not completely observed, in particular the dependent variable yp+1. (With more than a single-response variable, the following method is applicable to each of them, in the role of yp+1 next.) Symbolize by G0 and G1 the groups of subjects with complete data and with missing data on the outcome of interest, respectively. Denote the means and variances of these variables, in the corresponding groups with present and missing data on yp+1, by µ01, . . . , µ0p, µ11, . . . , µ1p, and σ012, . . . , σ0 p 2 , σ112, . . . , σ1 p 2 , respectively. We note that the groups G0 and G1 are variable specific, depending in general on the response variable, but for simplicity of notation dispense below with adding a variable index to them unless confusion may arise.
As indicated in the extant literature on missing data analysis (e.g., Graham, 2009), useful AVs can be found among those that exhibit mean and variance differences across their pertinent groups G0 and G1 (see also the “Conclusion” section). The rest of this note is concerned with discussing a readily applicable method of interval estimation of these groups’ mean and variance differences:
and
under the assumption
Next we discuss a readily and widely applicable LVM procedure for interval estimation of these two group difference indexes of interest.
A Latent Variable Modeling Procedure for Examining Candidate Auxiliary Variables
To achieve the aims of this note, we adopt a popular LVM framework represented by the following equation (e.g., Bollen, 1989):
where y is a p× 1 vector containing the given set of observed measures y1, y2, . . . , yp, µ is a p× 1 vector of associated intercepts, Λ = [λ mn ] is a p×q matrix of factor loadings, η is a q× 1 vector of latent variables, and ε is a p× 1 vector of uncorrelated error terms (q≤p). For the particular goals of interval estimation of the above group difference indexes (Equations 1 and 2), we will be using a two-group version of a special case of Equation (3) with p = q = 1, λ11 = 1, and ε1 = 0, as well as E(η1) = 0 in the group G0 (cf. L. K. Muthén & Muthén, 2012). This modeling setup with p > 1 is discussed in detail in Raykov and Marcoulides (2010) in the context of examining aspects of group differences via LVM in the presence of missing data using the ML approach to the analysis of incomplete data. In the rest of this note, (a) we specialize these earlier developments to the case p = 1 or relevance here, where in addition the groups are based on whether data are present or not on a prespecified response; and (b) we provide a means of examining group differences in variances in incomplete data sets, whose interval evaluation was similarly not of concern in that prior source.
Point and Interval Estimation of Missing/Present Group Differences in Means
For the model defined in Equation (3) and the discussion immediately following it, the mean group difference Δµ
j
in Equation (1) is represented as the latent mean in group G1 (L. K. Muthén & Muthén, 2012; cf. Raykov & Marcoulides, 2010). Hence, on fitting this special case of Equation(3) to the data from the above two groups associated with response or lack thereof on the outcome variable of interest, an estimate of the pertinent parameter Δµ
j
in Equation (1) becomes available, denoted as
where
This CI can be obtained for any of the variables y1, y2, . . . , yp once a researcher settles on an outcome measure (i.e., yp+1) of interest, and furnishes a range of plausible values for the population difference in means on each of them across the groups of subjects with and without data on that measure. If the CI is entirely on one side of the zero point, this can be seen as evidence in favor of considering yj as a potentially effective AV in subsequent ML- (or MI-) based analyses (see also Note 1). Conversely, if the CI contains the zero point, then pending assessment of the results of the group variance differences (see next subsection) one may suggest that in the analyzed data there may not be strong evidence to consider yj as a possibly useful AV in following analyses (see Note 1).
Point and Interval Estimation of Missing/Present Group Discrepancy in Variances
In addition to mean differences, variance discrepancies across the groups G0 and G1 could be useful in the process of identifying a potentially effective AV. In particular, an interval estimate for an appropriate index of differences in group variances would render a range of plausible values for their discrepancy at large, and could similarly be important when searching for useful AVs.
Under the model in Equation (3) (see also discussion immediately following it), the variance group difference measure
By way of contrast, the group variance ratio in Equation (2), after a suitable monotonic transformation defined below, will approximate considerably better normality under realistic sample sizes. As described next, it is this normality that we will capitalize on when obtaining a CI for the proposed group variance discrepancy (GVD) index in Equation (2). Before we explicate this transformation and procedure for interval estimation of the index
To interval estimate the GVD index
More formally, to obtain a CI for the GVD index
with ln(·) denoting natural logarithm. A large-sample standard error of the transformed GVD estimator,
(j = 1, . . . , p; cf. Browne, 1982). From Equations (6) and (7), a large-sample 100(1 −α)% CI (0 < α < 1) for the transformed GVD index, κ, is obtained as
Finally, to furnish a large-sample 100(1 −α)% CI for the index
Specifically, designating by κl and κu the lower and upper endpoints of the CI in expression (8), respectively, use of Equation (9) furnishes the sought large-sample 100(1 −α)% CI for τ, the GVD index of interest here:
The CI in expression (10) represents a range of plausible population values for the variance discrepancy (2) across the two groups of interest, G0 and G1, for a candidate AV, yj. An important feature of the interval (10) is that because of the particular approach to its construction, none of its endpoints can be negative or include empirically meaningless values. If the CI (10) does not cover 1, it may be suggested that this measure, yj, may be an effective AV (j = 1, . . ., p; see also Note 1). If the CI includes 1, it may be suggested that in the analyzed data there is no evidence for yj having different variances in the groups under consideration.
Appendix B provides an R-function that can be used for the calculation of the CI in expression (10), once its estimate and standard error are supplied from the output associated with fitting Model (3) with Mplus (see also Appendix A for the code to obtain the estimate and standard error).
We illustrate next the discussed procedure for group difference evaluation using empirical data.
Illustration on Data
To demonstrate the utility and applicability of the discussed approach to effective AV identification, we use data from an aging research study by Baltes, Dittmann-Kohli, and Kliegl (1986). In their investigation, the reserve capacity of older adults in the fluid intelligence subabilities Induction (Inductive Reasoning) and Figural Relations was focused on. As part of that study, n = 219 older adults were administered a battery of p = 8 intelligence tests at a baseline and a repeated assessment occasion. These repeated measures were the so-called ADEPT Induction, ADEPT Figural relations, Thurstone’s Standard Induction, Culture-Fair, and Raven’s progressive matrices tests, as well as two perceptual speed and vocabulary tests. (For further details on the tests, see Baltes et al., 1986.) For the method illustration purposes of this section, we consider the repeated assessment’s score on the ADEPT Induction test as an outcome variable of concern (with respect to a research question). On this test, there were n0 = 177 elderly with observed data and n1 = 42 persons with missing values. In addition, for our aims here we will be concerned with the set of eight intelligence test scores at baseline assessment as candidate AVs for inclusion in subsequent analyses of interest carried out on this incomplete data set.
We begin by fitting the model in Equation (3) to the resulting two group data (with present and missing values on the repeated ADEPT Induction score; see previous section of this note). Thereby, we take in turn each of the eight test scores at baseline in the role of the candidate auxiliary variable yj in the preceding section (j = 1, . . . , 8; see Appendix A for Mplus code details). Since this model is saturated, for each of its eight applications here it is associated with perfect fit (and hence there is no need to report its perfect fit indexes explicitly). The resulting estimates of the group mean and variance difference indexes Δµ
j
and
Estimates, Standard Errors, and 95% Confidence Intervals of the Group Mean and Variance Difference Indexes (1) and (2), for Each Baseline Measure (Considered as Candidate Auxiliary Variable).
Note. AI = ADEPT Induction test (at baseline); AF = ADEPT Figural Relations test; TI = Thurstone’s Induction test; CF = Culture-Fair test; RA = Raven’s Advanced Progressive Matrices test; PS1 = first Perceptual Speed test; PS2 = second Perceptual Speed test; VOC = Vocabulary test; S.E. = standard error; CI = confidence interval.
From Table 1, we see that the baseline ADEPT Induction and Thurstone’s Standard Induction measures may be potentially useful AVs (see also Note 1). This is because of two findings that are readily deduced from the third and last column of Table 1 containing the CIs for these measures’ group mean and variance differences, respectively. Specifically, since the CI for the mean group difference on either of these measures is positioned entirely below the zero point that signifies no such difference, it is suggested that the group of elderly with missing values on the repeated assessment ADEPT Induction test is markedly lower on either of these two baseline measures than the group of older adults with data on this follow-up outcome variable. Similarly, the CI for the GVD index on either of these baseline measures is positioned entirely below 1, which as mentioned earlier is the point signifying no variance differences. Hence, it is suggested that elderly who dropped out after baseline are more similar to each other on either of these two baseline measures, than elderly who participated in the follow-up assessment (i.e., provided data on the ADEPT Induction test at the repeated assessment).
We thus conclude that there are considerable differences in means and variances across the two groups in question here, for the baseline ADEPT Induction and Thurstone’s Standard Induction measures, in that dropouts are markedly lower in means and extent of individual differences. These two findings suggest that elderly who were lost to follow up may well have been lower on the underlying Inductive Reasoning subability at baseline, which is tapped into by these two Induction tests. (On the latent structure of the test battery, see Baltes et al., 1986.)
Furthermore, there is evidence in Table 1 that the first Perceptual Speed test and the Vocabulary test exhibit largely the same means and variability across the groups under consideration. This is because the CIs of their mean and variance group difference indexes contain their corresponding critical values indicated earlier in the note (viz., 0 for no group mean differences, and 1 for no group variance differences). In addition, there is some partial evidence that the ADEPT Figural Relations test may be a useful AV—its mean difference index CI does not include the critical value 0 whereas its GVD index CI covers the corresponding critical value of 1. Similarly, there is some partial evidence that the Culture-Fair test, Raven’s test, and the Second Perceptual test may be useful AVs for the same reasons (see also Note 1).
Table 1 also reveals an interesting finding with respect to the differences between the subjects with missing data on the response variable of concern in this illustration, on one hand, and those elderly with data on this repeated ADEPT Induction measure, on the other hand. Specifically, on all eight baseline measures examined from this cognitive aging study, the group of older adults who were not available for the follow-up assessment with the measure had a lower mean (see second column of Table 1). Moreover, on six of these eight initially presented intelligence tests, the mean of the dropout group was significantly lower than the mean of the repeated assessment participants (with no adjustment for multiple testing, which may be argued as not being of relevance when exploring potential AVs; e.g., Enders, 2010; see fourth column of Table 1). Also, on two of these six measures, the dropouts exhibited significantly lower degree of individual differences, that is, were more “compact” in the distribution of their pertinent baseline scores (i.e., more similar to each other). Hence, the outlined procedure in this note is also useful in identifying variables on which dropouts are distinct (as well as the direction of these discrepancies) from the remaining participants in a longitudinal study, both with respect to (a) mean performance and (b) degree of individual differences and similarities on examined measures.
Conclusion
This note intended to contribute to the field of missing data analysis in the social and behavioral sciences. ML and MI are two currently available state-of-the-art approaches for dealing with incomplete data sets (e.g., Schafer & Graham, 2002), which are typically based on the assumption of data MAR. (Although the theory of MI is not based on MAR, current software implementations of MI assume MAR and it is practically exceedingly difficult to develop a dependable imputation model when data are not MAR.) The MAR assumption is not statistically testable, however, and is a relatively restrictive condition that will be violated whenever the probability of missingness is related to the actually missing (unobserved, unavailable) values. Oftentimes in empirical research however, and for instance in longitudinal research, the propensity of missingness depends on the actually missing values, that is, MAR is not fulfilled.
As discussed in more detail in the extant literature, an approach to dealing with violations of MAR consists in the so-called inclusive analytic strategy that has been popularized over the past decade (e.g., Collins et al., 2001; Graham, 2009). Accordingly (a) one extends with AVs the set of predictors in subsequent models fitted to (part of) the incomplete data set of concern or (b) includes AVs as “background” variables (Graham, 2003) in these models, which are only allowed to be correlated among themselves, with the independent observed variables, and with the error terms of dependent manifest variables. Similarly, helpful AVs can be included in imputation models preceding MI-based analyses. Effective AVs contribute to enhancing the plausibility of the MAR assumption, and possibly to statistical power, thus making it more likely that the subsequent ML-based modeling yield essentially unbiased parameter estimates.
In this context, the present note was specifically concerned with discussing a method possibly contributing to identifying useful, or effective, AVs (see also Note 1). We capitalized on an earlier made observation by other authors that such variables could be found among those that are associated with considerable differences across the groups of subjects with present data and with missing data on an outcome variable(s) of concern with regard to a research question. We discussed an LVM approach to point and interval estimation of these groups’ means and variances. The approach is also useful more generally for finding measures on which subjects with missing data on a response variable(s) of concern are different from those with data on the outcome(s). The procedure of this article is to be considered only providing partial information contributing to identification of effective AVs, and should best be used in tandem with a method for point and interval estimation of the correlations between candidate AVs and outcome(s) in the presence of missing data (see Footnote 1). A discussion of such a method goes beyond the confines of this note.
Several limitations of the discussed procedure are worth noting here. One stems from its requirement for large samples. The reason is that the underlying method of estimation is typically ML that has optimal statistical properties with large samples. 2 Another limitation follows from the assumption of (approximately) continuous candidate AVs. We do not see in principle obstacles, however, for extending this procedure to the case of discrete measures considered as candidate AVs. Last but not least, the assumption of normality for these measures needed also to be made for a trustworthy application of the used estimation method, in the context of interest to this note that includes possible violations of MAR in the first instance. To counteract these violations to begin with, useful AVs that are identified on substantive grounds can already be included when fitting the model defined in Equation (3) (see also discussion immediately following it). With violations of normality, application of (a) the increasingly popular bootstrap approach may be recommendable for obtaining CIs for the two-group difference indexes in Equations (1) and (2), or alternatively (b) the robust ML method (unless there is piling of cases at a floor or ceiling effect on a measure used; e.g., L. K. Muthén & Muthén, 2012; see also Appendix A). We encourage future research into the possible robustness of the outlined procedure against single or multiple violations of these assumptions.
In conclusion, this note enriches the arsenal of empirical social and behavioral scientists working with incomplete data sets, by offering a readily and widely applicable procedure that can be a helpful aid in identifying useful AVs that enhance the plausibility of the frequently made assumption of data MAR.
Footnotes
Appendix A
Appendix B
Acknowledgements
Thanks are due to C. K. Enders, B. T. West, and K.-H. Yuan for valuable discussions on missing data analysis and auxiliary variables, as well as to P. B. Baltes, F. Dittman-Kohli, and R. Kliegl for permission to use their data from an aging and fluid intelligence project for illustration purposes. We are also indebted to an anonymous reviewer for valuable and critical comments on an earlier version of the note, which contributed considerably to its improvement.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
