In organizational studies involving multiple levels, the association between a covariate and an outcome often differs at different levels of aggregation, giving rise to widespread interest in “contextual effects models.” Such models partition the regression into within- and between-cluster components. The conventional approach uses each cluster’s sample average of the covariate as a regressor to identify the between-cluster component of the regression. This procedure, however, yields biased estimates of contextual effects unless the cluster sizes are large. Moreover, bias in estimation of such contextual coefficients in turn introduces bias in estimated coefficients of other correlated cluster-level covariates. Missing data further complicate valid inferences. This article proposes an alternative approach that conditions on the latent “true” cluster means of covariates having contextual effects while taking into account ignorable missing data with a general missing pattern at each level. The proposed model may include random coefficients. We compare inferences under different approaches in estimation of a contextual effects model using data from two national surveys of high school achievement.
The contextual effect analysis has long been of interest in analyzing clustered or multilevel data in organizational sciences (Firebaugh, 1978; Hammond, 1973; Lee & Bryk, 1989; Raudenbush & Bryk, 1986, 2002; Raudenbush & Willms, 1995; Willms, 1986). The fundamental idea is that, within an organization, a person-level covariate and the outcome have an association known as the “within-group” association that may be different from the “between-group” association, the association between the organizational means of the covariate and the outcome. There may be several reasons for this difference. The organization mean on the covariate captures characteristics of peers, so that the between-group association may reflect peer effects on the outcome. Moreover, peer composition may be associated with the resources available to the organization or to its normative environment or climate. These factors in turn may be related to the outcome. Yet, all of these organizational features are held constant when one assesses the within-group regression coefficient. In many cases, there is no reason to assume that the between-group and within-group associations are identical or even similar. The contextual effect is formally defined as the difference between the between-group and the within-group regression coefficients and may be of interest in itself (Willms, 1986). Furthermore, it is often important to adjust for the contextual effect when assessing the association between other organizational characteristics and the outcome. For example, in an influential article, Lee and Bryk (1989) examined student educational achievement as a function of high school sector (public vs. Catholic), disciplinary climate, and prior academic background, controlling for the contextual effect of student socioeconomic status (SES).
The current article considers a problem that arises quite generally in contextual analysis: the sample mean of the covariate is generally used to represent the organization mean on the covariate. Unfortunately, the sample mean is an unreliable estimate of the organization mean, and this unreliability will generally lead to bias not only in estimating the contextual effect but also in estimating the association between other organizational covariates and the outcome controlling for the contextual effect. We propose a model that regards the organization mean on the covariate as a latent variable. The approach allows multiple covariates at the person level, any one of which may have a contextual effect, and multiple covariates at the organization level. This general approach also allows the person-level covariates to have random coefficients, another pervasive feature of organizational data, particularly in education (Lee & Bryk, 1989; Raudenbush & Bryk, 2002). Finally, we tackle a pervasive problem in all such analyses based on large-scale survey data: the presence of missing data on the outcome, the covariates, or both that may arise at both person and organization levels.
A Simple Contextual Effect Model
Figure 1 (Raudenbush & Bryk, 2002) displays a simple contextual effect model having a person-level covariate, SES, on the horizontal axis and the outcome, student math achievement, on the vertical axis. The three ellipses represent the joint distribution of these two variables within each of three hypothetical schools: “low SES,” “average SES,” and “high SES” schools with respective mean SES, –1, 0, and 1. Within each school, the association between student SES and math achievement is positive and, for simplicity, the within-school regression coefficients are constant at . Note that the is the expected difference in math achievement between two students who attend the same school and differ by one unit in SES. A second regression line (the heavy line) represents the effect of the school-mean SES on the three school-mean math achievements, , which represents the expected difference in the school-mean achievement scores between two schools that differ by one unit on the school-mean SES. The contextual coefficient is (Raudenbush & Bryk, 2002, p. 141). As Figure 1 shows, is the expected difference in math achievement between two students who have the same SES but who attend schools that differ by one unit on the mean SES. Although the within-school regressions are depicted with a fixed slope in
Figure 1, they may, in reality, be heterogeneous with varying slopes across the schools (Raudenbush & Bryk, 2002, chap. 5, p. 157). Later in this article, we explore such “nonuniform effects models” (Raudenbush & Willms, 1995).
Contextual effect, , associated with attending School 2 versus School 1.
Formalization
We can generalize this simple contextual effect model for student attending school as
where is a univariate response variable, is a person-level covariate having a contextual effect (the SES in
Figure 1), is the school mean on the person-level covariate, and school-specific and student-specific are independent random effects for Level-1 nested within Level-2 . We shall also assume a simple two-level model for the covariate:
The person-level varies around its school mean with random effect having within-school variance while the school mean itself varies around the overall mean with random effect having between-school variance .
The Problem of Bias
Form of the bias
Researchers interested in contextual effects do not specify Model 1. Rather, they estimate a contextual effect model of the same that substitutes the sample mean for the true mean . From , we can readily identify the bias associated with this conventional approach
Under Model 2,
where is known as the “reliability” of as an estimate of (Raudenbush & Bryk, 2002, Chap. 3). This reliability varies from school to school as a function of the within-school sample size . Consequently, the exact form of the bias will depend on the joint distribution of and . However, the form of the bias is instructive in the balanced case when for all so that . Substituting into yields
Equation 3 suggests a solution to the problem of the bias: simply substitute the for the unknown in the . The conditional mean can be computed by a Bayesian approach specifying the joint prior distribution of , and and integrating over these using Markov Chain Monte Carlo (MCMC) methods such as Gibbs sampler and Metropolis-Hastings algorithm (Gelfand & Smith, 1990; Gelman et al., 1995; Gilks, Ricardson, & Spiegelhalter, 1996; Hastings, 1970; Liu, 2001; Metropolis et al., 1953). Alternatively, one may avoid the specification of the priors and substitute the “empirical Bayes (EB)” estimate by the method of maximum likelihood (ML) where , and are the ML estimates (MLE) and . The EB approach is useful when the number of schools, , is large. The Bayes estimates can be obtained from Bayesian software packages such as Winbugs (Spiegelhalter, Thomas, Best, & Lunn, 2002) and MLwiN (Rasbash, Steele, Browne, & Prosser, 2004); most standard packages for hierarchical linear or mixed models will produce as output the values of . These Bayes approaches to bias correction, however, do not extend readily to models with missing data.
The Problem of Missing Data
Reparametrization as a joint distribution of variables subject to be missing
Equation 1 describes the conditional distribution of the response variable given covariates and , while Equation 2 explains the marginal distribution of the covariates. We view the latent as a missing datum. To facilitate estimation using a model-based approach to and subject to be missing, it is useful to reexpress and in terms of the joint distribution of , , and . Let and . It is revealing to express the joint distribution as
from which it clearly follows,
Equation 7 is a one-to-one transformation between the desired Model 1 and the joint distribution of with
In the balanced case, we know that the conventional estimation of results in bias which is also the bias in estimating . From in Equation 8, it is clear that underestimated leads to overestimated .
Although the form of Equation 6 helps clarify the relationship between the joint distribution and its Conditional Form 1, , a more parsimonious form is simply
Estimation of the model that handles missing data
We have seen that the joint distribution of , Equation 6 or Equation 9 is equivalent to the Conditional Forms 1 and 2. The joint distribution 9 is a special case of a multilevel multivariate hierarchical general linear model (HLM). At Level 1 within an organization, we have
for , , and . At Level 2 between organizations, we have
for , and . The combined HLM or mixed model is
Let be an observed-value indicator matrix such that is a two-by-two identity matrix for completely observed, for observed and missing and for missing and observed. The extracts all available data. The complete-data Model 12 can be estimated solely based on its observed-data model
by iterative algorithms such as the EM algorithm and Fisher scoring (Shin & Raudenbush, 2007). Based on the observed Model 13, one may, using all available data, estimate the contextual effects model in one of two ways: directly estimate Model 12 and then transform to Model 1; multiply impute data from and then analyze the multiply imputed data based on Model 1 using widely available software for multilevel analysis by now-standard rules for analyzing multiply imputed data (Rubin, 1987, 1996) as specialized to the multilevel case by Shin & Raudenbush (2007).
A Constrained Model When the Contextual Effect Is Null
Suppose now that the data provide no evidence in favor of the contextual effect so that in Model 1, that is , so that Model 1 becomes
The appropriate joint distribution is obtained by recognizing that, in this case, so that the joint Model 9 becomes
Equation 15 has seven parameters while Equation 9 has eight parameters. Researchers interested in the model having no contextual effects with missing data should take care not to use the apparently appropriate Model 9 but rather to use the constrained Model 15. See Shin and Raudenbush (2007) for details.
Some Extensions of the Simple Model
In this section, we extend the simple Model 1 to a reasonably general contextual effects model that may include Level-1 covariates having contextual effects; Level-1 covariates having no contextual effects; Level-1 covariates having random effects across organizations; and Level-2 covariates.
Including a Level-1 Covariate
Suppose now that we wish to include an additional Level-1 covariate that is presumed not to have a contextual effect. This principle extends to the case of multiple Level-1 covariates, each of which may or may not have a contextual effect. For example, consider a model
where has a fixed effect , are subject to be missing, and all others are the same as they are in Model 1. To set up estimation for this model, one can first write down the full contextual effects model
where is the latent group mean of . From this model, it is straightforward to write down the joint distribution of and then derive the constraint in the covariance structure corresponding to in the conditional Model 17. Details are presented in Appendix A.
Including a Level-2 Covariate
As mentioned in the introduction, a common goal in multilevel analysis of survey data is to estimate the association between an organization-level covariate, , and the outcome, adjusting for the contextual effect of a Level-1 covariate. A simple model would be
where is a Level-2 covariate having a fixed effect . Among prominent examples in the literature, is a measure of the disciplinary climate of the school and is student SES.
Identifying the bias when replaces in Model 18
In all published examples we have found to date, researchers have used the sample average in place of the unknown true . The bias in estimating the contextual effect is then propagated to bias in estimation of whenever and are correlated. Thus, we have
The bias terms are for estimating and for estimating .
Correcting the bias
In parallel with the simplest contextual effect in the previous section, we can eliminate the bias by substituting for in Model 18 or by substituting the EB estimate , where are MLE. However, as in the simplest case, this approach does not extend readily to handle missing data.
Incorporating missing data
To base inference on the latent variable rather than the unreliable sample mean while allowing for missing data, we reparameterize the desired Model 18 in terms of the joint distribution of the outcome and covariates subject to be missing. We then estimate the joint model based on its observed-data model having the general form of Equation 13. Details appear in Appendix B.
Random Coefficients
A random coefficient for a completely observed covariate
Suppose that we are interested in a model with a contextual effect associated with student SES, denoted here as . However, we wish to control for ethnicity as indicated by a binary variable if the student is an ethnic minority student, otherwise, and we believe that the association between minority status and the outcome, mathematics achievement, varies randomly from school to school (c.f., Lee & Bryk, 1989). As is commonly the case, SES is frequently missing while minority status is completely observed. We might then write
for and where .
To estimate the desired Model 23 from incomplete data, we need to write the joint distribution of variables subject to be missing given the variables completely observed. The desired Model 23, , and suggest that such a joint distribution, , has 12 parameters. The joint model may be expressed as
for and . Nonzero yields an extraneous interaction effect between and on . Estimation of the complete-data Model 24 is based on the observed-data model. See Appendix C for details.
A random coefficient for a covariate subject to missingness
Suppose again that we are interested in a model with a contextual effect associated with student SES . As in the Model 23, we wish to control for a Level-1 covariate whose effect varies from school to school. However, this Level-1 covariate is subject to missingness. The desired contextual effect model is Equation 23, which substitutes for the completely observed . This model having eight parameters along with and implies that a joint model of to handle missing data is constrained to have a single effect of on and and has 15 parameters. Therefore, the fully unconstrained joint model should involve more than 15 parameters unlike the joint distribution (Equation 32) in Appendix A. The fully unconstrained joint distribution is not multivariate normal and, hence, the usual factorization under joint normality that leads to the linear conditional model does not apply. Thus, our ML approach is difficult to implement while Bayesian inference via MCMC would be more straightforward.
A Reasonably General Model
With one exception, all cases in Sections 2 and 3 can be represented in the general model below and readily estimated by ML via the EM algorithm or Fisher Scoring as detailed in Appendix C. The exception arises when a covariate subject to missingness has a random coefficient. The generalization applies to multiple covariates having contextual effects, to multiple covariates not having contextual effects, and to multiple covariates having random coefficients. Data can be missing with a general missing pattern at any level. The general contextual effects model is
where is a scalar response, is a vector of covariates having fixed effects , is a vector of covariates having the true latent cluster means and contextual effects , is a vector of covariates having random effects and for Level-1 nested within Level-2 . We assume to be completely observed. Although this approach does not require the presence of an intercept in , many applications do. Thus, let where subscript “D” in stands for completely observed covariates in .
We view as missing. Our strategy is to estimate the joint model of and covariates subject to missingness in given completely observed covariates. We partition so that where Level-1 and Level-2 are and completely observed covariates, respectively, and where Level-1 and Level-2 are and covariates subject to missingness, respectively. Let be a vector of Level-1 variables subject to be missing. We estimate the parameters of the joint distribution constrained to be a one-to-one transformation of the general contextual effects Model 25 and then transform to the desired (Shin & Raudenbush, 2007). See Appendix C for details.
In the next two sections, we compare the differences in the inferences of conventional, latent cluster-mean, and EB contextual effects models. The complete-data conventional and EB models are estimated by HLM 6 (Raudenbush, Bryk, Cheong, & Congdon, 2004). We illustrate how to estimate the desired latent cluster-mean Model 25 from incomplete data. The joint Model 39 in Appendix C is estimated by the authors' C program via ML and translated to the desired model. We use Fisher information to provide an objective comparison of standard errors with those of HLM 6 based on Fisher information. The convergence criterion is the difference in the observed log likelihoods between two consecutive iterations taken to be less than .
Illustrative Example I: Error-Prone Measurement
In this section, we illustrate how the approach described in this article operates to remove bias associated with errors in measurement of the organizational mean of the covariate in contextual effects models. To clarify key principles, we assume no missing data and compare the inferences made using the conventional approach that substitutes the sample mean for the unknown ; the EB approach that substitutes for the unknown ; and the latent cluster-mean approach that uses the latent . We illustrate how our latent variable approach can be used to reduce bias while allowing estimation from incomplete data in the next section.
We compare the inferences of conventional, latent cluster-mean, and EB models with a subset of the High School and Beyond Study of 1980. It has 7,185 students nested within 160 schools with no missing values (Raudenbush & Bryk, 2002). Each school has 14 to 67 students surveyed with a mean of 45 students. The data for analysis are described in
Table 1. MATH is the mathematics achievement test score; SES is the first principal component of family income, mother’s education, and father’s occupation; MINORITY is an indicator taking on a value of unity if the student is Black American or Hispanic, 0 otherwise; and SECTOR is an indicator taking on a value of unity if the school attended is a Catholic school and 0 if the school is a public school (non-Catholic private schools were excluded from the analysis for simplicity).
Subset of the High School and Beyond Study of 1980 for Analysis
Level
Variable
Description
Mean (Standard Deviation)
I
MATH
Math achievement score
12.75 (6.88)
SES
Standardized socioeconomic score
0.00 (0.78)
MINORITY
1 if Black or Hispanic
0.27 (0.45)
II
SECTOR
1 if Catholic, 0 if Public
0.44 (0.50)
The Random-Intercept Model
Students attending a private Catholic school perform 2.8 points (14.2–11.4) better than those attending a public school in the average math achievement scores. The inquiry of interest is whether such a gap reduces after controlling for the contextual effect of SES. Let , , , and in such that the desired contextual effect model is
It took 14 iterations for the joint Model 39 of to converge to ML. Table 2
compares inferences across the conventional, the desired latent cluster-mean, and the EB models.
Conventional, Latent Cluster-Mean, and EB Models With the Effect of SECTOR on MATH Controlling for the Contextual Effect of SES
Covariate
Parameter
Estimate (Standard Error)
Conventional Model
Latent Cluster-Mean Model
EB Model
Intercept
12.13 (0.20)
12.16 (0.20)
12.16 (0.20)
1.23 (0.30)
1.16 (0.31)
1.16 (0.31)
2.19 (0.11)
2.19 (0.11)
2.19 (0.11)
3.14 (0.38)
3.37 (0.41)
3.36 (0.41)
2.31 (0.36)
2.20 (0.36)
2.32 (0.36)
37.02 (0.63)
37.02 (0.63)
37.02 (0.63)
The within-cluster effect and its standard error 2.19 (0.11) are identical between conventional and latent cluster-mean models, as expected. The contextual effect and its standard error 3.14 (0.38) under the conventional model are underestimates relative to their counterparts 3.37 (0.41) under the latent cluster-mean model. A more consequential concern with the conventional model is with the effect of SECTOR after controlling for the contextual effect of SES. The conventional effect 1.23 is an overestimate relative to the latent cluster-mean counterpart 1.16. The intercepts, 12.13 under conventional model and 12.16 under latent cluster-mean model are different although the difference is small due to the standardized SES with the sample mean 0. Unlike the conventional model, the EB model produces effects and standard errors practically indistinguishable from those of the latent cluster-mean model. The Level-2 variance estimates under both conventional and EB models are overestimates compared to their counterpart.
This example is comparatively benign because the school sample sizes in the sample data are comparatively large, ranging from 14 to 67 with the average estimated reliability of socioeconomic status high at . In data sets having smaller and smaller reliability estimates, the biases will be larger. The resulting inferences based on the conventional model may be substantially biased. We may take an extreme example of a school having such that the reliability is 0.26 with . The unreliability of the has a direct influence on the estimated contextual effect based on . Consequently, a conventional analyst may be tempted to drop such an “unreliable” school. On the contrary, the latent cluster-mean model approach estimates the contextual effect based on the underlying , which makes use of “unreliable” as well as “reliable” schools in sample. Thus, the estimated contextual effect based on the latent cluster-mean model is less affected by such “extreme” schools and hence more reliable. If students attending such extreme schools add valid information at an individual level, we recommend that an analyst based on the latent cluster-mean model consider such individuals for analysis as we illustrate in Section 5. Controlling for the contextual effect of socioeconomic status, the math achievement gap between students attending a private Catholic school and students attending a public school reduces 59% to 1.16 points on average.
Bias on Minority Effect
Based on their single-level analysis, Fryer and Levitt (2004) reported suggestive evidence that Black students in their first 2 years of school lose substantial ground in their test scores due to attending lower quality schools relative to White students. We reason that such an impact may carry on after the students enter a high school and that the quality of schools is reflected in their socioeconomic status. Consequently, we hypothesize that such a minority gap will be at least partially explainable by the contextual effect of socioeconomic status. In this subsection, we investigate whether the lower math performance of minority students (Black or Hispanic) relative to that of other ethnic groups can be partially explained by the contextual effect of the socioeconomic status. Minority students perform 4.13 points (13.88–9.75) lower than their counterparts in the average math achievement scores.
The desired latent cluster-mean Model 25 has , , and such that
It took 14 iterations for the joint Model 39 of to converge. The conventional and EB models replace with and , respectively, for , , and . Table 3 displays the estimated conventional, latent cluster-mean, and EB models. The within-cluster effects of the socioeconomic status 1.97 (0.12) and 1.95 (0.11) are close to each other under the conventional and latent cluster-mean models, respectively, whereas the contextual effect and its standard error 2.95 (0.38) under the conventional model are underestimates relative to their counterparts 3.60 (0.44) under the latent cluster-mean model. The minority gap −2.71 (0.24) of the conventional model is also an underestimate compared to −2.87 (0.20) of the latent cluster-mean model, whereas the intercept 13.42 (0.15) under the conventional model is an overestimate compared to its counterpart 13.12 (0.16). The discrepant estimation in the contextual effect affects every estimate in the conventional model to be different from its counterpart in the latent cluster-mean model. However, the effect estimates and their standard errors are practically indistinguishable across the latent cluster-mean and EB models. Again, the Level-2 variance estimates of both conventional and EB models are overestimates relative to their counterpart.
Conventional, Latent Cluster-Mean, and EB Models With the Effect of MINORITY on MATH Controlling for the Contextual Effect of SES
Covariate
Parameter
Estimate (Standard Error)
Conventional Model
Latent Cluster-Mean Model
EB model
Intercept
13.42 (0.15)
13.12 (0.16)
13.12 (0.16)
−2.71 (0.24)
−2.87 (0.20)
−2.87 (0.20)
1.97 (0.12)
1.95 (0.11)
1.95 (0.11)
2.95 (0.38)
3.60 (0.44)
3.59 (0.44)
2.59 (0.38)
2.38 (0.38)
2.51 (0.38)
36.13 (0.61)
36.12 (0.61)
36.13 (0.61)
The latent cluster-mean model shows that, controlling for the contextual effect of socioeconomic status, the minority gap in math achievement reduces by 30.5% from 4.13 to 2.87 on average. The result indicates evidence that the minority gap in math achievement is partially explainable by the differences in school composition.
where having fixed effects for , , and . The conventional Model 4.22 is based on the error-prone . We are most interested in correcting the bias in the estimated and . The conventional model is compared to the corresponding EB model that substitutes for in Model 4.22. To estimate the latent cluster-mean model that replaces with latent , we first compose the joint Model 39 that is the closest transformation of the desired latent cluster-mean model under the model assumption that a covariate having a random coefficient should be completely observed
for and . The EB was used to satisfy the model assumption. Conditioning on and results in a model,
for . The joint Model 28 implies , , , , , and . We compare the Model 29, in place for the latent cluster-mean model, against the conventional Model 4.22 and the EB model.
The estimated conventional Model 4.22, Model 29, and EB model appear in Table 4
. It took 2,222 iterations for the joint Model 28 to converge. The estimates are practically indistinguishable between the conventional model and Model 29, whereas the between-cluster coefficient 5.33 (0.37) under the conventional Model 4.22 is noticeably an underestimate against its counterpart 5.56 (0.39) under the Model 29. The effect of SECTOR, 1.23 under the conventional Model 4.22, is an overestimate compared to 1.16 under Model 29, while interaction effect estimates of the conventional Model 4.22 are underestimates relative to those of Model 29. However, all coefficient estimates under the EB model are indistinguishable from those of the Model 29. The Level-2 variance estimates under conventional and EB models are overestimates relative to those of the Model 29.
Models With SES Having a Random Slope and a Contextual Effect on MATH
Covariate
Parameter
Estimate (Standard Error)
Conventional Model 4.22
Model 29
EB model
Intercept
12.13 (0.20)
12.15 (0.20)
12.15 (0.20)
1.23 (0.30)
1.16 (0.31)
1.16 (0.31)
1.04 (0.30)
1.13 (0.31)
1.13 (0.32)
−1.64 (0.24)
−1.67 (0.24)
−1.67 (0.24)
2.95 (0.15)
2.96 (0.15)
2.96 (0.15)
5.33 (0.37)
5.56 (0.39)
5.56 (0.39)
2.32 (0.36) 0.19 (0.20)
2.01 (0.36) 0.12 (0.20)
2.33 (0.36) 0.19 (0.20)
0.19 (0.20) 0.06 (0.21)
0.12 (0.20) 0.05 (0.21)
0.19 (0.20) 0.05 (0.21)
36.72 (0.63)
36.71 (0.63)
36.72 (0.63)
Illustrative Example II: Missing Data
In this section, we illustrate our latent variable approach with the National Education Longitudinal Study of 1988 that contains missing values. The sample contains 11,232 nationally representative respondents who attended 1,446 schools. The individuals were initially surveyed in 1988 as eighth graders and resurveyed through four follow-up surveys in 1990, 1992, 1994, and 2000. The subjects surveyed also include parents, teachers, and school administrators. The data for analysis are described in
Table 5.
Data for Analysis From National Education Longitudinal Study of 1988
Level
Variable
Description
Mean (Standard Deviation)
Missing (%)
SALARY
Log (weekly salary in dollar amount)
6.41 (0.47)
5,506 (49.0)
I
FEMALE
1 if female
0.53 (0.50)
0 (0)
SES
Base year socioeconomic status
−0.08 (0.78)
0 (0)
MATH
Math standardized score in tens
5.17 (1.02)
407 (3.6)
II
PRIVATE
1 if private high school
0.18 (0.38)
1 (0.1)
The outcome variable is log weekly earnings. We are interested in comparing, without bias, the log weekly earnings of students who attended private schools and public schools, controlling for the contextual effect of socioeconomic status and the effects of gender and math achievement. The 11,232 respondents consist of 9% African Americans, 13% Hispanic, 7% Asian, 1% native American, and 70% White. Among them, 53% are female. Of the 1,446 high schools, 18% are private. The school sizes range from 1 to 312.
The response variable , log of weekly earnings, has 5,506 values missing, typical of a large-scale survey. The contextual effects Model 25 has , , and for , , and such that
Note that, because 49% of the response values are missing and because covariates and are subject to missingness, neither the conventional model nor the EB model can be extended to estimate the desired Model 30. The joint Model 39 to handle missing data has , , , , , , and such that
It took 305 iterations for joint Model 31 to converge to ML. The estimated Model 30 appears in
Table 6
. The effect estimates and their standard errors have been multiplied by a factor of 10 for a better display of the estimated model. Controlling other covariates in the model, female gender (–2.21 [0.12]) is negatively associated while the student’s math achievement (0.54 [0.07]) and the school composition of socioeconomic status (0.74 [0.22]) are positively associated with students' earnings. Controlling for other covariates in the model, attending a private school (0.40 [0.22]) has a modest positive association with students' future earnings.
Effect of Attending a Private School on Student Earnings, Controlling for the Contextual Effect of Socioeconomic Status, and the Effects of Gender and Math Achievement
Covariate
Parameter
Estimate (Standard Error)
Intercept
62.05 (0.38)
FEMALE
−2.21 (0.12)
MATH
0.54 (0.07)
PRIVATE
0.40 (0.22)
SES
0.46 (0.10)
0.74 (0.22)
0.07 (0.02)
1.93 (0.04)
Note: Effects and standard errors have been multiplied by a factor of 10.
Discussion
In this article, we have considered a class of research settings in which persons are nested within clusters such as schools. In these settings, it is common to represent a person-level outcome as a linear function of covariates measured at either the person level or the cluster level. The partial association between each person-level covariate and the outcome can be represented with one or two parameters. The two-parameter representation allows the within-cluster and between-cluster associations to differ. This decision generates the contextual effects model considered in this article. Selecting the one-parameter representation constrains the between-cluster and within-cluster associations to be equal. The choice can be consequential not only in understanding the associations between person-level covariates and the outcome but also in understanding the association between cluster-level covariates and the outcome, given the person-level covariates. Past research reviewed above suggests that contextual effects will arise with some regularity in organizational research. However, in a scenario with multiple person-level covariates, the analyst will appreciate flexibility in selecting the two-parameter representation for some person-level covariates and the one-parameter representation for others.
Two problems have motivated this article. First, the two-parameter representation requires knowledge of the cluster mean of the person-level covariate. However, this mean is typically unknown but rather must be estimated from the data. If the estimate is imprecise as a result of small sample sizes per cluster, the analysis will yield biased estimates of the contextual effect as well as biased estimates of the association between cluster-level covariates and the outcome, given the contextual effect. Second, the analyst will typically encounter missing data at both levels. The two problems are related in that the latent cluster-level mean of the covariate may be regarded as a missing datum.
To solve these problems, we have represented the joint distribution of all variables subject to missingness as a linear function of the completely observed covariates and random effects assumed multivariate normal in distribution. The approach requires the assumption of data ignorably missing (Little & Rubin, 2002; Rubin, 1976). Using this approach, the investigator can decide whether to represent each Level-1 covariate’s partial relationship to the outcome with one or two parameters.
An alternative approach substitutes for the unknown true cluster mean, its conditional expectation of the true cluster mean, given all other covariates and parameter estimates. We have called this the EB approach. We have shown that the EB approach and our latent variable approach produce essentially the same effect estimates when the data are completely observed. However, the EB approach does not extend easily to the case of missing data.
Our approach allows only the completely observed Level-1 covariates to have coefficients that vary randomly over clusters. This limitation can be relaxed by using Bayesian estimation implemented by means of an MCMC algorithm. Implementation of such a full Bayesian approach is beyond the scope of the current study. Another limitation is that the approach assumes normality even for an indicator covariate subject to missingness. Schafer (1997) suggests that this practice poses no problem.
The approach lends itself to an alternative method of analysis not illustrated here. One can generate multiple complete data sets using model-based imputation given the observed data and the parameters where missing data, including the unknown cluster means of covariates having contextual effects, are ignorably missing (Little & Rubin, 2002; Schafer, 1997). Shin and Raudenbush (2007) extended this to the case of two-level data with data ignorably missing at either level.
Footnotes
Notes
Appendix
References
1.
Acevedo-GarciaD.LochnerK. A.OsypukT. L.SubramanianS. V. (2003). Future directions in residential segregation and health research: A multilevel approach. American Journal of Public Health, 93, 215–221.
2.
DempsterA. P.LairdN. M.RubinD. B. (1977). Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society, Series B, 76, 1–38.
3.
DempsterA. P.RubinD. B.TsutakawaR. K. (1981). Estimation in covariance components models. Journal of American Statistical Association, 76, 341–353.
4.
Diez-RouxA. V. (1998). Bringing context back into epidemiology: Variables and fallacies in multilevel analysis. American Journal of Public Health, 88, 216–222.
5.
Diez-RouxA. V. (2001). Investigating neighborhood and area effects on health. American Journal of Public Health, 91, 1783–1789.
6.
Diez-RouxA. V.MerkinS. S.ArnettD.ChamblessL.MassingM.NietoF. J. (2001). Neighborhood of residence and incidence of coronary heart disease. New England Journal of Medicine, 345, 99–106.
7.
Diez-RouxA. V.NietoF. J.CaulfieldL.TyrolerH. A.WatsonR. L.SzkloM. (1999). Neighborhood differences in diet: The Atherosclerosis Risk in Communities (ARIC). Journal of Epidemiology and Community Health, 53, 55–63.
8.
Diez-RouxA. V.NietoF. J.MuntanerC.TyrolerH. A.ComstockG. W.ShaharE. (1997). Neighborhood environments and coronary heart disease: A multilevel analysis. American Journal of Epidemiology, 146, 48–63.
9.
DuncanC.JonesK.MoonG. (1998). Context, composition and heterogeneity: Using multilevel models in health research. Social Science & Medicine, 46, 97–117.
10.
DuncanC.JonesK.MoonG. (1999). Smoking and deprivation: Are there neighbourhood effects?Social Science & Medicine, 48, 497–505.
11.
FirebaughG. (1978). A rule for inferring individual-level relationships from aggregate data. American Sociological Review, 43, 557–572.
12.
FryerR. G.Jr.LevittS. D. (2004). Understanding the Black-White test score gap in the first two years of school. The Review of Economics and Statistics, 86, 447–464.
13.
FullerW. A. (1987). Measurement error models. New York: John Wiley.
14.
GelfandA. E.SmithF. M. (1990). Sampling-based approaches to calculating marginal densities. Journal of American Statistical Association, 85, 398–409.
15.
GelmanA.CarlinJ. B.SternH. S.RubinD. B. (1995). Bayesian data analysis. London: Chapman & Hall.
16.
GilksW. R.RicardsonS.SpiegelhalterD. J. (1996). Markov chain Monte Carlo in practice. London: Chapman & Hall/CRC.
17.
HammondJ. L. (1973). Two sources of error in ecological correlations. American Sociological Review, 38, 764–777.
18.
HastingsW. K. (1970). Monte Carlo sampling methods using Markov chains and their applications. Biometrika, 57, 97–109.
19.
LairdN. M.WareJ. H. (1982). Random-effects models for longitudinal data. Biometrics, 38, 963–974.
20.
LeeV. E.BrykA. S. (1989). A multilevel model of the social distribution of high school achievement. Sociology of Education, 62, 172–192.
21.
LittleR. J. A.RubinD. B. (2002). Statistical analysis with missing data. New York: John Wiley.
22.
LiuJ. S. (2001). Monte Carlo strategies in scientific computing. New York: Springer-Verlag.
23.
LongfordN. T. (1987). A fast scoring algorithm for maximum likelihood estimation in unbalanced mixed models with nested random effects. Biometrika, 74, 817–827.
24.
MagnusJ. R.NeudeckerH. (1988). Matrix differential calculus with applications in statistics and econometrics. New York: John Wiley.
25.
MetropolisN.RosenbluthA. W.RosenbluthM. N.TellerA. H.TellerE. (1953). Equations of state calculations by fast computing machines. Journal of Chemical Physics, 21, 1087–1092.
26.
PickettK. E.PearlM. (2001). Multilevel analysis of neighborhood socioeconomic context and health outcomes: A critical review. Journal of Epidemiology and Community Health, 55, 111–122.
27.
PongS. (1998). The school compositional effect of single parenthood on 10th-grade achievement. Sociology of Education, 71, 24–43.
28.
RasbashJ.SteeleF.BrowneW.ProsserB. (2004). A user’s guide to MLwiN version 2.0. London: Institute of Education.
29.
RaudenbushS. W.BrykA. S. (1986). A hierarchical model for studying school effects. Sociology of Education, 59, 1–17.
30.
RaudenbushS. W.BrykA. S. (2002). Hierarchical linear models. Newbury Park, CA: Sage.
31.
RaudenbushS. W.BrykA. S.CheongY.CongdonR. T. (2004). HLM 6: Hierarchical linear and nonlinear modeling. Lincolnwood, IL: Scientific Software International.
32.
RaudenbushS. W.WillmsJ. D. (1995). The estimation of school effects. Journal of Educational and Behavioral Statistics, 20, 307–335.
33.
RossC. E. (2000). Neighborhood disadvantage and adult depression. Journal of Health and Social Behavior, 41, 177–187.
34.
RubinD. B. (1976). Inference and missing data. Biometrika, 63, 581–592.
35.
RubinD. B. (1987). Multiple imputation for nonresponse in surveys. New York: John Wiley.
36.
RubinD. B. (1996). Multiple imputation after 18+ years. Journal of American Statistical Association, 91, 473–489.
37.
SchaferJ. L. (1997). Analysis of incomplete multivariate data. London: Chapman & Hall.
38.
SchaferJ. L.YucelR. M. (2002). Computational strategies for multivariate linear mixed-effects models with missing values. Journal of Computational and Graphical Statistics, 11, 437–457.
39.
ShinY.RaudenbushS. W. (2007). Just-identified versus over-identified two-level hierarchical linear models with missing data. Biometrics, 63, 1262–1268.
40.
SpiegelhalterD.ThomasA.BestN.LunnD. (2002). WinBUGS user manual version 1.4. Cambridge, UK: MRC Biostatistics Unit.
41.
WillmsJ. D. (1986). Social class segregation and its relationship to pupils' examination results in Scotland. American Sociological Review, 51, 224–241.
42.
WuC. F. J. (1993). On the convergence properties of the EM algorithm. The Annals of Statistics, 11, 95–103.