Abstract
The author introduces methods for the decomposition analysis of multigroup segregation measured by the index of dissimilarity, the squared coefficient of variation, and Theil’s entropy measure. Using a new causal framework, the author takes a unified approach to the decomposition analysis by specifying conditions that must be satisfied to decompose segregation into unexplained and explained components. Here, the unexplained component represents the direct effects of the group variable on the conditional probability of acquiring a social position—such as a residential district in an analysis of residential segregation or an occupation in an analysis of occupational segregation—and the explained component represents indirect effects of the group variable on the outcome through covariates. The major merit of this approach is its ability to control individual-level covariates for the decomposition analysis of segregation. Two methods, one for semiparametric outcome models with the identity link function and the other for semiparametric outcome models with the multinomial logit link function, are introduced in this unified framework. The application of these methods focuses on occupational segregation among racial/ethnic groups. Father’s occupation, subject’s educational attainment, and the region of interview are included as covariates, using data from the General Social Surveys.
Keywords
Introduction
This article is concerned with the decomposition analysis of multigroup segregation, whose representative measures are discussed by Reardon and Firebaugh (2002). In particular, I introduce the decomposition of multigroup segregation measured by the index of dissimilarity (ID), the squared coefficient of variation (SCV), and Theil’s entropy measure. The decomposition method introduced here differs from a decomposition of quantity into within-group and between-group components. Rather, it is a decomposition of segregation measures into the component that represents the direct effect of the group variable X on segregation and the component that represents the indirect effect of X on segregation through covariates

Assumed causal structure.
There is a fundamental difference between the decomposition analysis introduced in this article and that of Reardon and Firebaugh (2002), which is closely related to Theil’s (1972) statistical decomposition analysis. Hence, I will first explain the conceptual difference between the two and the reason why I use a different approach to decomposition analysis from the one used in the Reardon-Firebaugh study. Suppose we have two categorical variables, one indicating ascribed group membership, such as race or gender, and the other indicating social/organizational positions, such as schools or occupations. Segregation indices, such as measures of racial segregation of schools or gender segregation of occupations, measure association between the two categorical variables. The segregation indices in the Reardon-Firebaugh study measure the extent to which population in individuals’ group membership, such as race, is reduced by knowledge of their social positions, such as schools. If knowledge of schools determines the race of individual students perfectly, or equivalently, if no heterogeneity of races exists within schools, schools are completely racially segregated. On the other hand, if knowledge of schools does not determine the race of students at all, or equivalently, if schools have the same racial composition of students, then there is no segregation.
The alternative segregation indices differ in measuring population heterogeneity. For example, when we use the concentration measure for heterogeneity, also called Simpson’s interaction index (Lieberson 1969), we obtain the multigroup extension of the ID (Sakoda 1981), and when we use Theil’s entropy measure of uncertainty for heterogeneity, we obtain the entropy measure of segregation (Reardon and Firebaugh 2002). Note that these segregation indices themselves provide a decomposition analysis. For example, heterogeneity in the race of students is decomposed into between-school variability in race and within-group variability in race. A further extension of this decomposition is straightforward if we have two nested levels of organizational units, such as schools and school districts. Then we can further decompose the between-school component into between–school district and within–school district components. This approach is a variation of the decomposition of variance, with the replacement of variance by a measure of population heterogeneity for a categorical variable. Note that because these segregation measures are association measures, they are symmetric except for the scaling factor (as I will show).
Another decomposition analysis related to the approach taken in this article, which differs from both the Reardon-Firebaugh and Theil traditions, is the decomposition of inequality introduced by Blinder (1973) and Oaxaca (1973) on the basis of regression models. This approach was later extended to a semiparametric method on the basis of the use of propensity score weighting by DiNardo, Fortin, and Lemieux (1996), which we will refer to as the DFL method. These methods make a distinction between the dependent variable, which is an individual outcome, and a key group variable, such as race or gender, and a set of control variables. This decomposition analysis is concerned with the decomposition of inequality in the outcome between groups into the component explained by group differences in the control variable and the remaining unexplained component. In particular, the DFL method eliminates, for the case in which the link function is the identity function in Figure 1, the indirect effect of X on Pr(Y = j) through
It is important to note there is a reversal between explanandum and explanans between the Reardon-Firebaugh method and the DFL method. In Reardon and Firebaugh’s decomposition analysis, the heterogeneity in group membership is explained by heterogeneity of social positions. In the DFL’s decomposition analysis, inequality (or heterogeneity) of social positions is explained by heterogeneity in group membership. This reversal is technically possible for the decomposition of segregation because, as mentioned earlier, segregation measures are symmetric except for the scaling factor. Only the reversed order taken by the latter approach, however, makes it possible to control for individual-level covariates in assessing the association between group membership and social positions.
It may seem that segregation measures, which are descriptive in nature, are not suited to modeling counterfactual situations. This decomposition, however, tries to provide an answer for a question such as, “What would be the extent of residential racial segregation in a counterfactual situation in which the indirect effects of race on the outcome through a given set of covariates were eliminated?” Residential racial segregation includes, in part, the effects of the association of racial groups with family wealth and family structure. The explained component therefore indicates the extent of segregation explained by an indirect-effect mechanism that racial groups are associated with family wealth and family structure, and those covariates of racial groups affect segregation. The unexplained component is the extent of residential racial segregation that would remain in a counterfactual situation in which such indirect effects were eliminated.
Yamaguchi (2017, 2019) introduced the DFL approach and its extension to the decomposition analysis of segregation, and he used the ID in occupational distribution between men and women to decompose the gender segregation of occupations into two components: a component explained by gender differences in human capital characteristics and the remaining unexplained component. Using data from Japan, he found that the unexplained component that remains under the equalization of human capital characteristics between men and women increases, rather than decreases, gender segregation of occupations, and he clarified the underlying mechanism of this paradoxical result. Note that unlike the decomposition of variance, where the control for an additional explanatory variable never increases the unexplained component, the unexplained component may increase under the control for covariates in the decomposition analysis of inequality. It is because the “explained” and “unexplained” components in the decomposition of inequality are actually the component explained by the indirect effect of the group variable due to its association with covariates that affect the outcome, and the remaining direct-effect component.
Because the direct and indirect effects may affect the difference in outcome between groups in the opposite direction, a control for covariates may increase, rather than decrease, inequality. However, use of the terms direct effects and indirect effects without qualifications is also problematic because of their ambiguous causal-analytic connotations. Generally, such effects depend on (1) alternative counterfactual situations that we assume for covariate distributions when they are made independent of the group variable to estimate the “direct effect” of group membership on the outcome and (2) conditions we impose on data to identify the direct effects—conditions that differ, for reasons explained in this article, between segregation measures that require a standardization of the covariate distribution based on the identity function and those that require a standardization of the covariate distribution based on the multinomial logit function. These specifications are made explicit in this article when interpreting results in terms of direct and indirect effects. We will keep the terms explained and unexplained components as terms without causal-analytic connotations.
It is straightforward to extend Yamaguchi’s (2017) method to the decomposition of multigroup segregation on the basis of the ID. However, the extension for the decomposition of the SCV and Theil’s entropy measure generates a distinct methodological issue. If the link function of the causal model of Figure 1 is the identity function, attaining statistical independence between X and
The lack of collapsibility was a known problem in the development of standardization methods of rates and odds in demographic research (Clogg and Eliason 1988), although it was not related to a causal-analytic framework. Not unlike propensity score weighting, the demographic standardization methods use a counterfactual situation in which a categorical control variable has the same distribution, called the standard distribution, across groups. The standardization method thus makes the group variable statistically independent of the control variable, to assess group effects on the outcome independent of the association between the group variable and the control variable. The method was first developed for the standardization of rates by Kitagawa (1955), Schoen (1970), Clogg (1978), and Clogg, Shockey, and Eliason (1990). It was also developed and elaborated for the standardization of odds by Teachman (1977), Clogg and Eliason (1987), Xie (1989), and Yamaguchi (2012).
The methods developed earlier, namely Schoen’s method for rates and Teachman’s methods for odds, were based on use of the discrete uniform distribution, or the equal-probability distribution, as the standard distribution of a single categorical control variable. However, except for special cases (e.g., age distribution of a synthetic cohort without mortality), the assumption of the equal-probability distribution is unrealistic. Furthermore, methods that use the marginal distribution of the control variable as well as the uniform discrete distribution for the standard distribution do not always generate standardized odds ratios between X and Y that are equal to odds ratios derived from the unique effects of X on Y in the multinomial logit model even when interaction effects of X and
Yamaguchi (2012) extended Xie’s method, which used a single categorical control variable, to the use of multiple categorical control variables, then reformulated it in the causal-analytic framework and called the method “log-linear causal analysis.” For Xie’s and Yamaguchi’s methods, the standardized odds ratios do not depend on the choice of the standard distribution of control variables
The first methodological issue discussed in the next section is a reformulation of multigroup segregation indices when the explanandum and explanans are reversed. The reformulation expresses three representative segregation indices so that they characterize the discrepancies between groups in the outcome distribution. Here, “outcome” refers to a school in an analysis of school segregation or an occupation in analysis of occupational segregation. These reformulations show that different segregation indices have some important commonalities, but their specificities require the assumption of different link functions for the decomposition analysis.
One may question the advantages of using a particular segregation measure over another. The differences among alternative indices are mostly technical, however. As described earlier, and also shown by Reardon and Firebaugh (2002), the difference between the ID and Theil’s entropy measure is solely derived from the choice between two alternative measures of population heterogeneity. This article also shows that the SCV is a rescaling of Pearson’s chi-square statistics, and Theil’s entropy measure is a rescaling of the likelihood-ratio statistics. Hence, the decomposition of segregation based on these two measures is also the decomposition of chi-square statistics for the test of independence between group membership and social positions. On the other hand, there is no formal relation between the ID and chi-square statistics, and as a result, the absolute scale of the ID differs greatly from the other two, as will be shown in the application. Some further thoughts on the choice of alternative segregation indices will be discussed in the concluding section.
The counterfactual situation also requires specifications for some alternative assumptions. To use the propensity score adjustment to make the group variable statistically independent of covariates, one needs to make an assumption about the alternative standard distribution of covariates, because the common covariate distribution is not specified by the condition of statistical independence alone. Another assumption relevant for the decomposition of segregation measures is the distinction Yamaguchi (2017) made between the “supply-driven DFL model” and the “demand-driven matching model.” In the DFL model, the marginal distribution of the outcome positions—where the outcome positions imply occupations in an analysis of occupational segregation or residential districts in an analysis of residential segregation—is assumed to change depending on how the covariate distributions of groups change in counterfactual situations. When covariates are individual characteristics, such as characteristics of job applicants in the labor market, the DFL model represents the “supply-driven” determination of the outcome distribution. The marginal distribution of outcome positions in the matching model is assumed to remain unchanged, however, because of the distribution of positional outcomes being determined solely by the demand for the positions, and the assumption that a given counterfactual situation will only change the matching between people and outcome positions. This article shows that this distinction can be applied not only to the decomposition of segregation based on the ID but also to the decomposition of segregation based on Theil’s entropy measure and the SCV. However, we do not know whether the outcome will be generated by the supply-driven DFL model or by the demand-driven matching model. Because the two alternative models represent two extreme situations, we can expect these models give a range of outcomes within which an actual outcome will be realized under the specified counterfactual situation.
With data from the General Social Survey 2000 to 2010, the application focuses on the decomposition of racial/ethnic segregation of occupation on the basis of three segregation measures using (1) father’s occupation, (2) educational attainment, and (3) region as covariates to see how much of the racial/ethnic segregation in occupation can be explained by racial/ethnic differences in father’s occupation alone, how much by racial/ethnic differences in educational attainment alone, and how much jointly by father’s occupation and subject’s educational attainment. I also examine how the explained component changes when we control for racial/ethnic differences in geographic labor market locations.
Methods
The marginal and joint frequencies of observations are denoted by
Multigroup Segregation Measures and Their Reformulations
In this section, I show the reformulation of selected segregation measures, all of which were originally expressed by Reardon and Firebaugh (2002) as a function of
Index of Dissimilarity
The ID for multigroup segregation (Sakoda 1981) can be reformulated as
where
Squared Coefficient of Variation
The SCV C can be reformulated as follows:
Equation (2A) shows that C is a weighted average of
Theil’s Entropy Measure
Theil’s entropy measure can be reformulated as follows:
where
The General Approach to the Decomposition Analysis
Equations (1B), (2B), and (3B) show that the three segregation indices can be generally expressed as
where
One might consider this quantity as the unexplained component of segregation, and
Generally, the unexplained component of segregation should meet the following two necessary conditions for it to reflect the direct effects of X on Y, which would be realized upon elimination of the indirect effects of X on Y through the association of X with covariates
(A) When there is no unique effect of X on Y, or equivalently, if
(B) When there is no unique effect of
Conditions (A) and (B) are necessary conditions for the decomposition when the direct effect of X is absent, and when the indirect effect of X on Y through
Equation (5) satisfies condition (A) because when
(C) For the direct effect of X on segregation to be equal in amount, not by chance, to the observed segregation,
An important question here is whether
To address this issue, let me reexpress equation (5) when
Then, we obtain
Note that
There is, however, a problem. When we realize
Formally, the presence of collapsibility for the identity function, and the lack of it for the multinomial logit function, can be expressed as follows. Let the standardized conditional probability of the outcome, under the imposition of
where thetas are parameters. Then, we obtain
Hence, the effect of X on Y at different covariate values does not vary with covariate value
which holds when the link function is the identity function, does not hold for the multinomial logit function, and depends on a function of
(D) When a segregation measure depends on ratios of outcome probabilities between groups, the direct effect of X on Y must preserve the odds ratios between X and Y that are independent of covariate values when no interaction effects of X and
Condition (D), however, leaves uncertainty about a method to realize a counterfactual situation in which the interaction effects of X and
In the following, I first present an extension of the semiparametric model with the identity function, which Yamaguchi (2017) introduced for decomposition of the ID between two groups, for decomposition of multigroup segregation. Then, I consider additional semiparametric models with the multinomial logit link function, leading to identification of statistics that satisfy conditions (D) and (C) for the other two segregation measures.
I begin by reviewing the method for attaining statistical independence between the group variable and the covariates. I then describe the extension of two models, the supply-driven DFL model and the demand-driven matching model, introduced by Yamaguchi (2017) by assuming the identity link function for multigroup decomposition analysis. I then introduce the decomposition method that assumes the multinomial logit link function.
A Brief Review of the IPT Weighting Methods
To attain statistical independence between the group variable and covariates, I use the method of weighting by the IPT (Robins 1998; Rubin 1985). The following are two practically meaningful options for the standard distribution of covariates:
(1) We may assume all groups will obtain the distribution of covariates for the entire population under the independence between X and
where P(.) indicates the probability measure (Morgan and Winship 2007).
(2) We may assume groups will obtain the covariate distribution of a particular group, say group 1, for which the IPT weights for the mth group are given as
In both cases, the estimation for
Decomposition Method for the ID Based on the Semiparametric Outcome Models with the Identity Link Function
The DFL Model
Suppose we assume the following semiparametric saturated linear probability model with categorical covariates
where
Suppose now we realize, by use of weights
because
Note that we also obtain
which depends only on the weighted average of
As we saw in equation (1A) for the ID, ID is a weighted average of
Condition (D), although it is not necessary for the identity link function, holds when there are no interaction effects of X and
is a function of
Yamaguchi (2017) called this the DFL model because, technically, it is the same as that derived by applying the decomposition method that DiNardo et al. (1996) introduced to the decomposition of differences in the conditional probability.
The Matching Model
Yamaguchi (2017) also introduced an alternative model he called the matching model. The matching model for the multigroup case assumes the equation
where
Similar to the DFL model, we obtain
as the estimate for the unexplained component of
On the Possibility of Applying the DFL and the Matching Model to Alternative Segregation Measures
Regarding the SCV, equation (2A) indicates this segregation measure is the weighted average of
Because of the functional form of Theil’s entropy measure, it is evident that condition (D) is not met for this measure, either. Hence, the decomposition method based on the semiparametric models with the identity link function is applicable only to the ID.
Decomposition Method Based on the Semiparametric Outcome Model with the Multinomial Logit Link Function
In this case, we can use the method of “log-linear causal analysis” that Yamaguchi (2012, 2016) developed. Yamaguchi introduced a method that extends Xie’s (1989) “CD purging” method for the standardization of odds among conditional probabilities of the outcome variable.
I slightly modify an expression of the semiparametric additive effect model described in equation (8). I do so by introducing the intercept that may change under counterfactual joint distribution of the group variable and covariates as follows:
In equation (19),
However, the model of equation (19) requires a prior estimation for the cross-classified expected frequencies that eliminate the interaction effects of X and
Yamaguchi (2012, 2016) was concerned with estimation of the set of adjusted marginal conditional probabilities of Y for the given X, which we denote by
for the model of equation (19), thereby preserving the same unique effects of X on Y while eliminating the covariate effects on Y. Yamaguchi referred to the method as “log-linear causal analysis” and applied the method where covariates are confounders of the effects of X on Y, but here I apply the same method without any connotation of causality for the purpose of purging out the effects of any given set of covariates. Appendix B describes the estimation procedure of
By using this method, we can obtain the estimate of
It is evident that this function, which reflects the effects of X on segregation, does not reflect the effects of
Because equation (20) does not specify function F(.), this decomposition method can be applied to all segregation indices that permit the expression of equation (7) over and beyond the three segregation measures used as examples here. For the ID, however, we may still regard the assumption of the identity link function as preferable because the assumption does not require prior elimination of the interaction effects by the ML estimation.
Steps for Applying Methods
For a given set of covariates, we create IPT weights that generate independence between the group variable and covariates in weighted data. This is the input data for the DFL model.
Using the data generated in step 1, we adjust the marginal distribution of the outcome to be equal to the observed one while maintaining
We apply the DFL model and the matching model to estimate the unexplained component for the ID.
By using the IPT-weighted data for the DFL model as the input data, and using Yamaguchi’s (2012, 2016) covariate-effect purging method, we obtain the two-way table of X and Y that retains the unique effects of X on Y for the IPT-weighted data. We repeat this step using the weighted data of the matching model.
Using the adjusted two-way frequency tables from step 4, we estimate the unexplained component of Theil’s entropy measure and the unexplained component of the SCV for the DFL model and the matching model.
Application
Application of these models focuses on analysis of occupational racial/ethnic segregation, which has not been analyzed as much as occupational gender segregation. Geographic segregation of workplaces among racial/ethnic groups is strong (Gradín, del Rio, and Alonso-Villar 2011; Tomaskovic-Devey et al. 2006), and occupational racial segregation has been studied as part of general occupational segregation by sex and race, especially focusing on macro time trends (Albelda 1986; King 1992; Queneau 2009; Watts 1995), but the relationship between geographic racial/ethnic segregation and occupational racial/ethnic segregation has been little studied. Exceptions are studies by Tomaskovic-Devey et al. (2006) and Gradín et al. (2011). By using industry, establishment size, and region as indicators of workplaces, Tomaskovic-Devey et al. analyzed desegregating time trends for occupational racial/ethnic segregation between and within workplaces. Their major finding is that desegregation occurred mostly within workplaces, rather than between. Gradín et al. analyzed geographic variability in occupational racial/ethnic segregation using state-level data. Relevant to the analysis presented here, they found that segregation tends to be high in states with low numbers of Hispanic and Asian residents. These analyses all use national-, regional-, or workplace-level data, and none include individual-level control variables, as the analysis presented here does.
Data
I use data from the General Social Survey 2000 to 2010 (2000, 2002, 2004, 2006, 2008, and 2010). I do not use data from years after 2010, because of changes in the code of occupation and industry, or from before 2000, because of the lack of consistent availability of the ethnicity variable. The population consists of men and women ages 20 to 69 years who are identified as either white or African-American. Hispanic people are separated out as “Hispanics” regardless of race, which yields three categories of “Hispanics,”“non-Hispanic whites,” and “non-Hispanic blacks.” I refer to the latter two categories as “whites” and “blacks” for simplicity below. This is our group variable. Unweighted sample counts of the three groups are 9,916, 1,993, and 712, respectively, totaling 12,612 for the entire sample. Segregation among these three groups is measured for the distribution of eight occupational categories: (1) managers and administrators, (2) professional and technical workers, (3) sales workers, (4) clerical and administrative support workers, (5) service workers, (6) skilled manual workers, (7) semiskilled manual workers, and (8) unskilled manual workers. The last three categories are distinctions among nonservice manual workers.
Three categorical covariates for controls are used: (1) educational attainment (0 to 11 years, 12 years, 13 to 15 years, 16 years, and more than 16 years), (2) father’s occupation (the eight occupation categories plus a “missing” category), and (3) region (nine categories distinguished by the General Social Survey), represented by the REGION variable.
To adjust for sampling variabilities, samples are weighted by the General Social Survey variable for sampling weights WTSSALL, although the total weighted frequency is equated to the total unweighted frequency in the sample. The test of statistical independence takes this into account, as described later.
The conditional probability distribution of the eight occupations by the race/ethnic variable and the marginal distribution for data weighted by sampling weights are given in Table 1.
Occupational Composition of the Three Groups
Decomposition Analysis
I present decomposition analyses for three segregations: (1) three-group segregation, (2) black-white segregation, and (3) Hispanic-white segregation. I also present the decomposition models using different covariate controls: (1) control for father’s occupation, (2) control for subject’s educational attainment, (4) control for father’s occupation and subject’s educational attainment, and (4) control for the three covariates by adding the region variable. The analysis aims to determine what proportion of occupational racial/ethnic segregation is explained by racial/ethnic differences in occupational family background alone, educational attainment alone, and jointly by the two, and how the extent of occupational racial/ethnic segregation changes when we control for the nine regions. I examine results from the DFL model and the matching model for the three segregation indices.
Tests of Statistical Independence between the Group Variable and Covariates
The decomposition analyses are based on different indices and link functions but are all based on the same set of data adjusted by the IPT weighting. Table 2 presents results of the test of statistical independence between the group variable and each covariate based on the Clogg-Eliason method (Clogg and Eliason 1987) for observed data weighted by the sampling weight and covariate-control models 3 and 4 weighted by the product of the sampling weight and the IPT weight. Models 1 and 2, with a single categorical covariate, each perfectly fit the data; therefore, their test results are not presented. The Clogg-Eliason method is applicable to log-linear models for weighted data by modeling weighted data by parameters while using unweighted sample counts for the chi-square tests of significance. Skinner and Vallet (2010) pointed out, however, that when the “cell weights,” which indicate the ratio of unweighted versus weighted observations for each cell (a combined state of cross-classified data), vary within cells, the Clogg-Eliason method tends to underestimate the standard error and thereby overestimate the chi-square statistics for goodness-of-fit tests. Hence, the Clogg-Eliason method provides a conservative test for statistical independence. In the tests presented in Table 2, observed data use sampling weights, and the IPT-adjusted data use the product of the sampling weight and the IPT weight; therefore, both weights vary within each cell of two-way tables.
Tests of Independence between the Group Variable and Each Covariate for the Unweighted and IPT-Weighted Data
Note: Variable labels indicate X for race/ethnicity, V1 for father’s occupation, V2 for educational attainment, and V3 for region. Results for models 1 and 2 that use one categorical covariate for propensity score estimation fit the data perfectly; therefore, they are not presented in the table. Model 3 uses V1 and V2, and model 4 uses V1, V2, and V3 for the propensity score estimation. IPT = inverse probability of treatment; LL = log-likelihood.
Interaction effects that are included in the propensity-score estimation, in addition to the main effects of V1 and V2, are (1) “college graduates or above”דmanagers for V1” and (2) “college graduates or above”דmissing for V1.”
In addition to the main effects of the three covariates, and interaction effects (1) and (2), the following interaction effects are included in the estimation of propensity scores: (3) “college graduates or above”דeach region dummy,” (4) “father’s occupation missing”דeach region dummy,” and (5) “REGION = California”דeach father’s occupation dummy.”
Table 2 shows that estimates of the propensity scores successfully attain statistical independence for models 3 and 4. The notes with Table 2 describe the interaction effects among covariates that were necessary for each model to predict the propensity score.
Main Results
Table 3 presents the main results for the decomposition of racial/ethnic segregation among white, black, and Hispanic individuals for each of the four covariate-control models. For the six decomposition measures, the table presents the index value, its unexplained component, and its explained component. The relative proportions of the unexplained and explained components are also presented.
Three-Group Segregation
Note: DFL = DiNardo, Fortin, and Lemieux; ID = index of dissimilarity; SCV = squared coefficient of variation.
First, results in Table 3 indicate that the ID and two other measures differ greatly not only in absolute value but also in the estimate for the relative extent of the explained portion, and the ID tends to give a larger proportion of the unexplained. The first fact is well known, but not the second one. Yet because both the Theil’s entropy measure and the SCV are a rescaled value of chi-square statistics, the relative proportions between the explained and unexplained components tend to agree between the two measures. Note that because we use the marginal distribution of the entire sample as the standard distribution of covariates, estimates from the DFL model and the matching model tend to be highly congruent regardless of segregation indices. This characteristic may not hold, however, if we use the covariate distribution of a particular group, other than that of the entire population, as the standard distribution.
Regardless of indices, results in Table 3 indicate the following. First, the racial/ethnic difference in subject’s educational attainment has a much greater explanatory power for occupational racial/ethnic segregation than does racial/ethnic difference in father’s occupation. Second, because the effects of father’s occupation and subject’s educational attainment are largely overlapping, and the latter effect is greater than the former, the extent of racial/ethnic segregation that is jointly explained by racial/ethnic differences in both father’s occupation and subject’s educational attainment is only about 7 percent greater for the entropy measure and the SCV, and only 3.5 percent greater for the ID, than the extent of segregation explained only by racial/ethnic differences in subject’s educational attainment. Third, when we add regions to the control variables, the unexplained portion of segregation increases, rather than decreases, for Theil’s entropy measure and the SCV. The change, however, is very small for the ID.
There is a strong association between region of residence and racial/ethnic group distinctions, as shown in Table 2 in the chi-square test for the observed data. This leads to our third finding: Theil’s entropy measure and the SCV indicate that if people of different race/ethnicity were equally distributed across regions, racial/ethnic segregation would increase rather than decrease, because in regions where minority groups are over-represented, the extent of occupational racial/ethnic segregation is smaller.
This situation may result either because an increase in the population of minority groups in a region tends to increase racial/ethnic equality in occupational opportunity, or because minority groups tend to migrate to regions where more racial/ethnic equality of occupational opportunity exists. The present study cannot determine which is the case, but Tomaskovic-Devey et al. (2006) suggested the former interpretation, because they found that desegregation historically occurs mostly within workplaces.
What does the fact that this tendency is found for the Theil’s measure and the SCV, but not the ID, imply? Compared with the ID, Theil’s entropy measure and the SCV reflect more of the over-representation/under-representation in the outcome categories with smaller sizes. The relative contribution each outcome category makes for the discrepancy in the outcome distribution between a particular group m and the average population is
Relative Contribution of the Discrepancy in Segregation by Occupation (on the Basis of Observed Data)
Note: ID = index of dissimilarity; SCV = squared coefficient of variation.
Results in Table 4 show that compared with the ID, the SCV reflects black individuals’ overrepresentation among service workers and semiskilled manual workers and Hispanic individuals’ overrepresentation among unskilled workers and underrepresentation among managerial workers. Hence, the fact that the counterfactual equalization in regional racial/ethnic distributions increases racial/ethnic effects on segregation by Theil’s entropy measure and the SCV suggests black and Hispanic individuals are more concentrated in regions where their occupational handicaps, because of the above-described over- and underrepresentations in specific occupations, are relatively small. However, these implications are obtained indirectly from analysis of discrepancies between results from the decomposition of Theil’s entropy measure and the squared coefficient variation and results from the decomposition of the ID, and thus will require further analysis of segregation by region for direct confirmation.
Table 5 presents the main results for the decomposition of black-white occupational segregation. Compared with the results for three-group occupational segregation presented in Table 3, the extent to which black-white occupational segregation is explained by father’s occupation becomes greater for both father’s occupation alone and father’s occupation controlling for racial differences in subject’s educational attainment. Again, a control for region tends to increase the racial/ethnic effects on segregation when measured by Theil’s index and the SCV but not for the ID.
Black-White Segregation
Note: DFL = DiNardo, Fortin, and Lemieux; ID = index of dissimilarity; SCV = squared coefficient of variation.
Table 6 presents the main results of the decomposition of Hispanic-white occupational segregation. Compared with results for black-white occupational segregation, the role that father’s occupation plays becomes much smaller for Hispanic-white segregation, and when group differences in educational attainment are taken into account, father’s occupation plays little role, indicating that differences in father’s occupation between white and Hispanic individuals only indirectly affect occupational segregation between the two groups by affecting Hispanic-white differences in educational attainment.
Hispanic-White Segregation
Note: DFL = DiNardo, Fortin, and Lemieux; ID = index of dissimilarity; SCV = squared coefficient of variation.
The extent to which Hispanic-white differences in educational attainment explain Hispanic-white segregation in occupation differs greatly depending on whether we measure by the ID or by one of the other two measures. Using the ID, education explains less of the Hispanic-white segregation than it does black-white segregation, but for both Theil’s entropy measure and the SCV, education explains more of the Hispanic-white segregation than it does black-white segregation. Even though the two groups of measures have different sensitivities to discrepancies (as seen in Table 4), this finding leaves ambiguity regarding whether educational attainment plays a more important role in black-white or Hispanic-white segregation.
The tendency that a control for region increases racial/ethnic effects on segregation when measured by Theil’s index and the SCV, but not the ID, is also found in white-Hispanic segregation—in fact the tendency is stronger. The results from Theil’s index and the SCV are consistent with Gradín et al.’s (2011) finding that racial/ethnic segregation in occupation is greater in states with smaller proportions of Hispanic and Asian residents.
Conclusions and Discussion
This article introduces the decomposition method for multigroup segregation, a method that permits controls for individual-level data. This research owes much to the counterfactual thinking and related weighting method for propensity score adjustment developed by Rubin (1985) in his semiparametric approach to causal analysis, and to the application of Rubin’s method by DiNardo et al. (1996) in the decomposition analysis of inequality. The advancements made in this article are (1) a unified approach to the decomposition of segregation for semiparametric models having either the identity link function or the multinomial logit link function and (2) the incorporation of two methodological advancements made by Yamaguchi (2012, 2016), one relating to the decomposition of the ID and the other on log-linear causal analysis.
The application of the decomposition method in racial/ethnic segregation in occupation integrates two different traditions of sociological research. One is research in racial/ethnic occupational status attainment, where researchers are typically concerned with the effects of occupational origin and educational attainment; the other is research on geographic segregation of workplaces among racial/ethnic groups and its consequences for racial/ethnic inequality. Although the analyses presented here are purely illustrative, findings suggest minority groups are disproportionally present in regions where there is smaller occupational segregation among racial/ethnic groups. Thus, the geographic segregation of racial/ethnic groups is likely to be associated with a lower, not higher, level of inequality in occupational status attainment among racial/ethnic groups.
I would like to add a final remark regarding further applications of the method introduced in this article. Like other log-linear models, the decomposition methods based on the multinomial logit link function separate parameters that affect the marginal distributions of outcomes, such as occupational and school distribution, from parameters that affect the effect of group membership on outcomes. Hence, as in the distinction between structural mobility and circulation (or exchange) mobility in the study of social mobility, we can separate changes in the extent of segregation, and its decompositions into the direct-effect and indirect-effect components, into those due to structural changes, such as changes in occupational or residential distribution, and changes in the effects of group membership on the attainment of social positions, such as occupation or residential district. We can also examine effects of hypothetical changes on the distribution of outcomes on segregation. For the use of the ID with an identity link function, however, we can conduct such an analysis only when we assume a matching model in which parameters for the marginal outcome distribution are separated from parameters for the effects of the group variable on the outcome, but not when we assume the DFL model, which does not have separate parameters to manipulate the marginal distribution of the outcome variable.
Footnotes
Appendix A: Estimation of Parameters for the Matching Model
The DFL method satisfies the following conditions (a) through (d):
This is attained by the following log-linear iterative adjustment method. Let us denote by
We can use the iterative proportional adjustment, across rounds of iterative estimation starting with t = 1, to obtain the final estimate of
Then, the ratio of the final estimate to the initial estimate can be expressed as
where
Appendix B: Method of Purging Covariate Effects for the Model with Multinomial Logit Function
We assume that all covariates
First, we purge the interaction effects from
where
It can be mathematically shown that for conditional probabilities
satisfy equation (10) given in the main text.
Upon elimination of the interaction effects of C and X on Y, the covariate-state-specific conditional probability
With such a specification for
Without loss of generality, we may set weights of equation (2B) to be equal to the relative proportion of C’s states among those that satisfy
Because only the odds ratios, and not the conditional probabilities, preserve the effects of X on Y, we may further adjust the two-way frequencies obtained initially as
