Abstract
Structural equation models have provided a seemingly rigorous method for investigating causal relations in nonexperimental data in the presence of measurement error or multiple measures of putative causes or effects. Methods have been developed for fitting these very complex models globally and obtaining global fit statistics or global measures of their approximation to sample data. Structural equation models are idealizations that can serve only as approximations to real multivariate data. Further, these models are multidimensional, and the approximation is itself multidimensional. Tests of “significance” and global indices of approximation do not provide an adequate basis for judging the acceptability of the approximation. Standard applications of structural models use a composite of two models—a measurement (path) model and a path (causal) model. Separate analyses of the measurement model and the path model provide an informed judgment, whereas the composite global analysis can easily yield unreasonable conclusions. Separating the component models enables a careful assessment of the actual constraints implied by the path model, using recently developed methods. An empirical example shows how the conventional global treatment yields unacceptable conclusions.
An open question remains as to how psychologists should decide whether a psychometric model gives an adequate account of a social science data set. In the following informal essay (note that essay is the appropriate word), my discussion of the question—with a focus on structural equation modeling—is general and necessarily speculative. I approach the topic initially from a personal and historical perspective. It is argued that because in applications, structural equation models are idealizations, they are approximations to the real world. Further, such approximations are multidimensional and cannot be captured in a test of significance or a goodness-of-approximation index. I further show that an informative analysis can and should include separate treatments of the measurement (common factor) model and the path (causal) model that usually form a composite structural equation model. The separation allows the investigator to check the adequacy of the measurement model and to make a detailed examination of the causal hypothesis, thereby fulfilling the two primary objectives of any such investigation.
Preamble
In 1960, when I began work on psychometric theory—doctoral work on nonlinear factor analysis (McDonald, 1967)—psychological science was mainly served by the standard linear models of statistics, accompanied by analysis of variance. Such truly psychometric methods as factor analysis were looked on rather disdainfully by mathematical statisticians because the methods lacked statistical machinery.
At that time, we already had a rigorous account of the theory of maximum likelihood estimation in exploratory and confirmatory factor models from Lawley (1940, 1958). However, the numerical demands that efficient estimation methods placed on such psychometric models were not met by the primitive electronic computers we had. I learned this to my cost in 1958 in a course on programming in machine language for an ILLIAC computer, with a tiny memory, input on five-hole tape, and output on toilet paper.
As reanalysis of some of the data shows, the approximate numerical methods given by Spearman and Jones (1951) and by Thurstone (1947), for example, developed in the era before computers, gave remarkably good results compared with modern estimation methods. The first period of statistical respectability for psychometric models came with the advent of FORTRAN, and other user-friendly scientific programming languages, and rapid developments in the speed and memory of mainframe computers. The crucial step came when Karl Jöreskog applied the variable-metric numerical minimization method given by Fletcher and Powell (1963) to a series of psychometric models (see Gruvaeus & Jöreskog, 1970). These models included Lawley’s maximum likelihood treatment both of the exploratory factor model (Jöreskog, 1966, 1967; Lawley, 1940) and of the confirmatory factor model (Jöreskog, 1969; Jöreskog & Gruvaeus, 1967; Lawley, 1958). The big advance was Jöreskog’s application of variable-metric minimization to Wiley’s (1970, 1973) linear structural relations model—a combination of the standard path model of econometrics with two common factor models. The trade name of the resulting program, LISREL (Jöreskog, 1977; Jöreskog & Van Thillo, 1972), is often used these days as the title of the Wiley model. In turn, the advent of the silicon chip and Pentium processing has allowed increasingly computer-intensive methods, substituting for mathematically intensive methods. Currently, we can fit just about any model, no matter how poorly conceived, and get Bayesian estimates, by Markov chain Monte Carlo methods, giving a new meaning to “MC squared.” The models do not even need to be identified.
However, not long after the first programs enabling an objective, asymptotic chi-square criterion for rejecting, say, a measurement (factor) model or a structural (factor plus causal) model, a problem appeared. Psychologists applying the models to real data found that the chi-square test typically rejected models that they wished to accept, at the levels of significance they had become used to. It further became clear that every model was false and would be rejected if the sample size was large enough. These simple facts can still escape the notice of psychometric theorists, whose training is in mathematical statistics and whose applications are to computer simulations, not to the real world.
Over the last few decades, in response to the perceived problem with tests of significance, psychometric theorists have offered a large number of indices of the fit of a model to data, sometimes with authoritative but quite arbitrary suggestions for their application. It is, in fact, very easy to invent a fit index (see McDonald & Marsh, 1990). However, it is virtually impossible to find objective criteria for their application. (A reviewer of a study by one of my graduate students said confidently that because we do not know how to use any of the fit indices, the gold standard is to quote as many of them as possible.)
This brief historical account must serve to introduce this topic and to motivate an examination of the present state of the field. Focusing on the problem stated at the outset—how to determine the acceptability of a psychometric model—the discussion is confined mainly to the general structural equation model, but the conclusions are somewhat more general.
Psychometric Models and Their Uses
A remarkable amount of statistical theory has been subsumed under a multiple-regression model accompanied by analysis of variance (and, more recently, the general linear model of Nelder & Wedderburn, 1972, with functions linking various forms of response to the regressor variables). Similarly, a remarkable amount of psychometric theory for multivariate data has been subsumed under a general structural equation model, accompanied by analysis of covariance structures—a nonlinear structure implied by the linear equations (see Jöreskog, 1977; Jöreskog & Van Thillo, 1972; McArdle & McDonald, 1984; McDonald, 1978, 1980; Muthen, 1984; and Wiley, 1970, 1973).
Both the linear (regression) model and the (linear) structural model consist of mathematical equations containing random variables. We can think (i.e., I find it convenient to think) of the algebra of those equations as the syntactics—the pure grammar—of the models, containing no reference to the real world of applications. In addition, we need the semantics of the models—rules of their application to the real world. (An analogue in physics is given by Maxwell’s equations. By statements linking their component variables to observations, these equations give an account of electromagnetic phenomena.) Although this is a statement of the obvious, it is notable that not all psychometric models are published along with explicit rules for their application to data. Of course, it may be that the rules of application are deemed self-evident, or learned by a kind of apprenticeship, in the study of examples of their use.
Most formulations of the general structural equation model make it a hybrid of two distinguishable multivariate models—a measurement model and a path model. Exceptions are the covariance structure analysis model (McDonald, 1978, 1980) and the RAM model (McArdle & McDonald, 1984), but the distinction is commonly useful. The measurement model is a set of regression equations such that the partial correlations between the dependent variables are zero when conditioned on the independent variables. (Recall that a regression equation has the property that the error term—regression residual—is uncorrelated with the independent variables in the equation.) The path model consists of a collection of equations relating members of a set of variables to a subset of other variables in the set and a pattern of zero and nonzero covariances between the error terms in the equations. The path equations are not in general regression equations. That is, the error terms need not be uncorrelated with the independent variables. However, under common conditions they are also regressions. 1
Given separate measurement (M) and path (P) models, we can then choose semantic rules—rules of correspondence to applications in the real world. In an application of the measurement model, typically we obtain multiple measurements on a convenient or at best quasi-random sample of subjects. The first rule (M1) is to treat the obtained multiple measurements as realizations of the dependent random variables; these are called indicators. The second rule (M2) is to treat the independent random variables as modeling that which the observed empirical indicators measure in common. The independent variables explain the relations between the empirical indicators, and we therefore call them common factors, latent traits, or latent variables. As a consequence, they are unobservable; that is, their values are not determined by data. Notice that Rule M1 is obvious, although it does draw attention to likely failures of practical sampling procedures. Rule M2 is also obvious but, unfortunately, ambiguous as stated.
A convenient further clarification of Rule M2 is to say that a common factor is a quantitative abstract attribute of the examinees, recognizable from common properties of the indicators, as responded to by the examinees. The regression coefficients in the model (the factor loadings) represent the mean difference in the indicators between populations that differ by one unit in the attribute they measure in common. The error terms—unique components—constitute an interaction between the examinees and the idiosyncratic (specific) properties of the indicators. Many of the dimensions of cognition, personality, attitude, and so on seem to fit this simple version of Rule M2. However, some theorists regard the independent variables in the model as corresponding to common unobserved causes of their indicators. This causal interpretation is more likely to be adopted by psychometric theorists than by applied researchers, who would quickly recognize difficulties in actually using it. The main difficulty is that, in general, causes must be identified independently of their effects; they do not resemble their effects, and they cannot be identified from their effects alone. 2
There are possible limitations to both readings of Rule M2. Readers may check these limitations against their experience of applications. Is it unsurprising that the nine criteria for major depressive disorder in the Diagnostic and Statistical Manual of Mental Disorders (4th ed., American Psychiatric Association, 1994) are unidimensional (Aggen, Neale, & Kendler, 2005)? How should we conceptualize their relationships to the—unobserved—depression syndrome?
The measurement model requires a third rule if it is to be used as a model of generalizability—for example, to estimate the reliability of a lengthened test (see McDonald, 1999). Measurement Rule M3 states that the indicators under analysis are a proper subset of a very large domain of indicators of the same attributes, which could have been written and administered “had we but world enough and time.” McDonald (2003) made a general case for this rule. At least implicitly, it is in common use in test construction, whether accepted or not!
Let us turn now to the question of applying a path model (without common factors) to real data. To repeat, a path model consists of a set of random variables and equations relating each to a subset of the others, together with a pattern of zero and nonzero covariances between the error terms in the equations. Three semantic rules are needed for the application of a path model to a set of measures on a convenient or quasi-random sample of subjects. Path Rule 1 (P1) is that the measures are modeled as the set of random variables in the path model. Path Rule 2 (P2) is that the independent variables in each equation are direct causes of the dependent variable. Conversely, the dependent variable is the direct effect of those causes. In general the coefficient of each independent variable—its directed arc coefficient—represents the change in the dependent variable if a (conjectured) intervention increases that independent variable by one unit, with appropriate control of the others. See Pearl (2000) for a graphical account and McDonald (2002) for an algebraic account of this rule. These accounts also give the conditions under which the path equations become regression equations: that is, conditions under which observing the causal variables is equivalent to manipulating them. Path Rule 3 (P3) is that any correlation between the error terms will represent relations between the variables not accounted for by the causal variables in the network. These are relations that remain after appropriate hypothetical control of the given causal variables. The error terms in the path model are commonly known as disturbances. The term comes from an econometric myth to the effect that an economic system would be determinate if it were not “disturbed” by unexpected events.
Obviously, the interpretation of a path model given in Rules P2 and P3 may not command universal acceptance. The reader may consider what alternative conception of an asymmetric relation between variables would motivate a conjecture about the direction of that relationship. (Clearly, “X is correlated with Y” is equivalent to “Y is correlated with X,” but “X causes Y” is not equivalent to “Y causes X.”) The ability of behavioral scientists to carry out thought experiments, conjecturally manipulating variables and imagining their effects, enables the careful construction of path models. Similarly, their ability to conceptualize an attribute of their subjects as an abstract common property of indicators enables the careful construction of representative indicators. Both processes—imagining interventions and imagining behavior domains—have their limitations and questionable aspects.
When, finally, we combine the measurement model with a path model and have common factors that are causes or effects of other common factors, the rules of application combine in the obvious way, with each common factor identified as a conceptual variable by the common properties of its own indicators (not by indicators of its causes or effects). In a mixed case in which some of the variables in the path model are common factors (identified by multiple indicators) and some are single measures that enter the causal model directly, some arbitrariness of interpretation may arise. McDonald (1999, pp. 391–397) provided an example of three equivalent models in which measured depression can be a cause, an effect, or a fourth indicator of a common factor of “self-punitive attitude” (this is a reanalysis of work by Hull, Lehn, & Tedlie, 1991). The investigator can choose among these only on substantive grounds.
Whatever disagreements may occur over the rules of correspondence, I hope that a clear distinction has been drawn between the mathematical models and the rules that link them to applications in the real world. Let us next consider these models as falsifiable hypotheses, as idealizations, and as approximations to reality. We can then turn to the question of assessing their acceptability.
Models as Hypotheses, as Idealizations, and as Approximations
For many decades, behavioral science research has rested on the Neyman–Pearson hypothesis-testing philosophy. We set up a restrictive null hypothesis—a point hypothesis that implies a sampling distribution, hence a probability for our observations. We hope to reject it at a favorite level of “significance” in order to affirm an alternative, less restrictive hypothesis. 3 The volume edited by Harlow, Mulaik, and Steiger (1997), significantly titled What If There Were No Significance Tests?, serves to mark what we may hope is a switch from that philosophy to the assessment of effect size in the linear model and the use of confidence bounds—which may, of course, include the former null (point) hypothesis. 4
In the case of structural hypotheses, it would appear that we wish to affirm rather than reject the constraints implied by the model. A clear logic for such an affirmation does not seem to have been proposed. Commonly, a restrictive hypothesis can be expressed as a point (null) hypothesis. Thus, in standard theory for structural models, the restrictive model implies that the noncentrality parameter in the chi-square distribution is zero—a single point on the number line. From a Bayesian perspective, all point hypotheses are false a priori—the probability of a point on the number line—and their zero probability cannot be altered by data. We can reasonably expect that they will be rejected at any chosen significance level for a sufficiently large sample size, which may of course exceed the population of our planet. Nonsignificance at a chosen or available sample size does not constitute evidence in favor of a point hypothesis. These statements are truisms, but a brief survey of recent articles in psychometric journals would show that there are some very competent psychometric theorists who have not noticed them yet. This is possibly because these psychometric theorists work with computer simulations rather than real data. Computer simulations can be generated by true point hypotheses. Counterparts of such simulations may not exist in the real world.
Let us accept that, in general, restrictive statistical models are approximations to the reality of their applications to empirical data. There is, further, a sense in which psychometric models such as the common factor (measurement) model and the path (causal) model are idealizations, and their ideal character is a major reason why they are approximations to reality. The rules of application function rather loosely.
The obvious case is the common factor model. Ideally, each item is constituted (a) by a common part that corresponds to the abstractive attribute to be measured on the examinees and (b) by an idiosyncratic component that should correspond to the way in which the item differs uniquely from all the other items. In applications, it is difficult to the point of impossibility to write a set of items that differ essentially by a semantic component unique to each, so that they give uncorrelated interactions with the examinees. The task is further complicated if we seriously use the factor model as a model of generalizability (as I believe we must) and hope to go on writing further items of the same kind, with psychometric properties as good as the initial set written.
Similarly, in developing a causal path model, we imagine that the error terms (disturbances) are the effects of omitted causal variables, which, if found and added, would make the causal system determinate (see McDonald, 1997b). It is necessary to suppose that when further causes are added, any changes in the directed arc coefficients will be negligible. This assumption requires a remarkable amount of cooperation of the unknown with the known. There are a number of other ways in which these models are idealizations, but perhaps the points covered will suffice.
Our psychometric models, as hypotheses, are generally false. They are idealizations and therefore, through the rules of correspondence, approximations to the real world of applications. Let us turn now to the question of assessing the degree of approximation and deciding whether it is adequate.
Assessing Approximation
Beginning with the seminal work of Tucker and Lewis (1973), there has been a proliferation of indices of goodness of approximation–badness of approximation of structural models to multivariate data. (A change of terminology from goodness of fit to the more appropriate term goodness of approximation seems to have become widely accepted.) The many coefficients are not exhaustively surveyed here. Generally, they are functions of the discrepancies between sample covariances or correlations and the fitted (restricted) covariances or correlations implied by the model. Some are also functions of the number of parameters in the model, in relation to the number of covariances, and (sometimes implicitly) the sample size. Some measure goodness and some badness of approximation. Some are absolute indices, and some are relative to a “null” model nested within the model being fitted. McDonald and Marsh (1990) attempted a systematic study of the algebraic properties of the most fashionable indices. They accidentally invented one or two more and indicated ways in which indefinitely more could be invented without any intellectual effort.
Behavioral scientists are conditioned from their undergraduate years to like indices that range (at least in absolute value) from zero to one, where “zero” means awful and “one” means ideal. It all began with Pearson’s product–moment correlation coefficient. Not surprisingly, a number of the invented approximation indices have this property. These indices are often accompanied by a recommendation that approximation is acceptable if the index exceeds .9. This recommendation arises, seemingly, from the fact that the number .9 looks close to the number 1, not from any consideration about applications. The practice has no mathematical or empirical justification.
Given a series of nested models of increasing complexity, most fit indices will change monotonically as the complexity increases—goodness improving and badness decreasing. An obvious example is fitting 1, 2, 3, …, common factors in an exploratory model. Note that complexity has never been defined. It is generally taken to correspond, at least approximately, to the number of fitted parameters, or inversely (as “parsimony”) to the number of degrees of freedom. A number of psychometric theorists became excited over the possibility opened up by Akaike’s (1977) introduction of “an information criterion” (AIC). This index purported to balance badness of fit against model complexity, reaching an optimum value for a model of intermediate complexity and providing an objective criterion for choosing a “best” model (see, e.g., Bozdogan, 1987). The AIC and variants of it are still favored by some idealistic psychometric theorists and, accepting their authority, still used by some investigators. Computer simulations can appear to confirm its utility for model choice, because computer simulations belong to the ideal world of the mathematical model and not to the real world of applications, where the model is only an approximation.
However, McDonald (1989) showed that instead of balancing badness of fit against model complexity, the AIC actually balances badness of fit against sampling error, which reduces to zero in large enough samples. An opposition that should be a property of the population is, in the AIC, a property of the sample. As the sample size increases, a more complex model becomes “best” by the criterion, until the model has zero degrees of freedom. McDonald gave an empirical example in which the AIC behaves just like a test of significance, choosing a number of factors that are just not significant at the 1% level. The choice (for 17 variables) ranged from one factor at a sample size of 75 to nine factors at a sample size of 1,338. It should be clear that the AIC does not serve its intended purpose. McDonald and Marsh (1990) showed that, apart from the AIC, there is more than one way in which the number of parameters or degrees of freedom can be arbitrarily incorporated into various measures of fit and given some sort of justification. The difficulty is in the multiplicity of these inventions and the lack of any clear basis for their application. None provide an objective ground for model choice.
A recent special issue of Personality and Individual Differences (Barrett, 2007; Bentler, 2007; Goffin, 2007; Hayduk, Cummings, Boadu, Pazderka-Robinson, & Boulianne, 2007; Markland, 2007; McIntosh, 2007; Miles & Shevlin, 2007; Millsap, 2007; Mulaik, 2007; Steiger, 2007) contains a collection of articles on goodness of approximation that reflect severe doubts about the applicability of fit indices. These articles do not deny the position stated by McDonald and Ho (2002): There are four known problems with fit indices. First, there is no established empirical or mathematical basis for their use. Second, no compelling reason has been given for choosing a relative fit index over an absolute index…. Third, there is not a sufficiently strong correspondence between alternative fit indices for a decision based on one to be consistent with a decision based on another; the availability of so many could license the choice of the best-looking in an application, though we may hope this does not happen. Fourth, and perhaps most important, a given degree of global misfit can originate from a correctable misspecification giving a few large discrepancies, or it can be due to a general scatter of discrepancies not associated with any particular misspecification. (p. 72)
First, note that the badness of approximation of our ideal model to the data is primarily represented in the discrepancies between the sample covariances and the covariances implied by the fitted model. Most indices are functions of these discrepancies, but the important point is that approximation is itself multidimensional, and a single index cannot possibly capture the wide variations in the sizes and patterns of the discrepancies. It is, however, reasonable to rest our judgments on their individual sizes. It is also reasonable to simplify judgments by turning the discrepancies into the residual correlations that have long been traditional in factor analytic work. The models are intended to account for relationships, and Pearson was not wrong to measure (linear) relationships with correlation coefficients. After we make some allowance for sampling error, the residual correlations in a measurement model will quite often include some that correspond to what Thurstone (1947) called “doublet factors” and Spearman called “shared specifics” or “overlap” (Spearman & Jones, 1951). A property is shared by just two items, making their idiosyncratic components cease to be unique. Depending on the context of the investigation, these can be accounted for by a doublet factor or eliminated by summing them or deleting one. 5 Doublets are common in applications.
Even when one or more global fit indices meet a recommended arbitrary criterion for acceptance, there may be clusters of residuals, which suggest and could support additional factors defined by three or more measures. Importantly, the cluster must correspond to a nameable attribute of the examinees. A general principle emerges—a metacriterion, we might say. We may declare a measurement model to be an acceptable approximation if the residuals do not support a more complex, substantively convincing model. For this purpose, we can still use Thurstone’s (1947) classical criterion for “salience”—a standardized loading greater than .3. A careful examination of the residuals may reveal an unsystematic scatter across the matrix with no troublingly large values. (My students often asked me for a criterion for “troublingly large.” I would offer “>.1,” pointing out that a residual less than .1 would not allow a product of two salient loadings (>.3). However, this cannot be a rule for all contexts. The question is an embarrassing one.)
In the case of a pure path model—without common factors or latent variables—it has recently become clear from the work of Pearl (2000) and McDonald (2002, 2004b) that not only global approximation indices but also residual correlations provide a less than optimal way to judge the adequacy of the approximation. Each path model implies a basic set of constraints, one for each degree of freedom in the model. (The basic set of constraints may logically entail other constraints, but there are only as many independent constraints as there are degrees of freedom.) If we assume uncorrelated disturbances, the constraints take the form of zero partial correlations when variables mediating between causes and their indirect effects are partialed out. For example, if Negative Life Events cause Depression, which causes Suicidal Tendencies, controlling Depression breaks the causal nexus between Negative Life Events and Suicidal Tendencies. The assumption of uncorrelated disturbances implies that partialing out Depression would give a zero correlation between Negative Life Events and Suicidal Tendencies. Checking one partial correlation, which is all that the model implies, is more tightly focused than consulting a goodness-of-approximation index or looking at residual correlations, of which the goodness-of-approximation index is a function. It also directs us to the appropriate thought experiment—whether depression can be controlled by counseling or by pharmaceuticals, giving a possible intervention to break the causal nexus. Further, it draws appropriate attention to the set of equivalent models that imply the same constraint. (On substantively logical grounds, we reject the equivalent model—Suicidal Tendencies cause Depression, which causes Negative Life Events.) Checking basic constraints encourages investigators to examine the substantive details of their model.
When we combine a measurement model with a path model, the first important but neglected point to note is the fact that a competent investigator can easily combine a bad path model with a good measurement model, or a good path model with a bad measurement model, depending on the state of relevant empirical knowledge. Generally the number of degrees of freedom in the measurement model is much larger than in the path model, so in the first case the goodness of approximation of the measurement model can be expected to swamp the badness of approximation of the path model, which appears to happen quite commonly. McDonald and Ho (2002) examined 41 published structural equation model studies, in which only 14 gave results for the measurement model and the full structural model, fitted separately. Investigators who got a good fit to the full model and the measurement model did not take the further step of checking the fit of the path model by differencing the chi-square values obtained. If they had done this step, they would have discovered (just using a goodness-of-approximation index separately for each) that the fit of the path model was bad, contrary to their conclusions. Those who fitted only the full structural equation model never had an opportunity to discover that their path model did not give an acceptable approximation to the data. I conjecture that many published structural equation models are unacceptable for this reason.
These observations by McDonald and Ho (2002) were limited to published chi-squares and computable (but inadequate) goodness-of-approximation indices. They serve to show, at least, that the measurement model and the path model should be examined separately. One quite reasonable way to do this is to fit the measurement model, then fit the path model to the factor correlations, and examine the resulting residual correlations in the path model itself. (This two-step procedure loses efficiency of estimation—probably a negligible loss—but may gain more in detailed information than it loses in efficiency.) However, given what has already been said about the pure path model, the most informative method of all is to test the factor correlations for the extent to which the basic constraints are approximated. This method also has the advantage that it draws attention to the class of equivalent models that imply the same constraints and to the thought experiments that supply meaning to the model. These remarks may gain some concreteness from the following example.
An Empirical Illustration
A correlation matrix obtained by Feist, Bodner, Jacobs, Miles, and Tan (1995) is given in Table 1 . The sample size is 149. The 14 indicators—items measured on Likert-type scales—are intended to measure the following five attributes: Subjective Physical Health Problems, Daily Hassles, World Assumptions, Constructive Thinking, and Subjective Well-Being. For Subjective Physical Health Problems, the indicators are frequency of symptoms, muscular symptoms, and gastrointestinal symptoms. For Daily Hassles, the indicators are time pressures, money worries, and inner concerns. For World Assumptions, the indicators are benevolence of world and people and self as worthy. For Constructive Thinking, the indicators are global constructive thinking, behavioral coping, and emotional coping. For Subjective Well-Being, the indicators are purpose in life, environmental mastery, and self-acceptance.
Feist et al.’s (1995) Data: Correlation Matrix, Means, and Standard Deviations (N = 149)
Feist et al. (1995) fitted two models to these data, one in which Physical Health Problems and Daily Hassles determine Subjective Well-Being through the intermediate variables World Assumptions and Constructive Thinking—the bottom-up model—and one that reverses causality—the top-down model. McDonald (1999, chapter 17) showed that with a slight modification of these models, they were equivalent. Fitting McDonald’s version of the bottom-up model by maximum likelihood gives the results in Figure 1 .

Standardized parameters and standard errors in parentheses for the bottom-up latent variable model fitted to Feist et al.’s (1995) data. FS = frequency of symptoms; GS = gastrointestinal symptoms: MS = muscular symptoms; BWP = benevolence of world and people; SW = self as worthy; PH = Physical Health; WA = World Assumptions; SWB = Subjective Well-Being; PIL = purpose in life; EM = environmental mastery; SA = self-acceptance; DH = Daily Hassles; CT = Constructive Thinking; TP = time pressures; MW = money worries; IC = inner concerns; GCT = global constructive thinking; BC = behavioral coping; EC = emotional coping.
The chi-square is 159.57, with 69 degrees of freedom, p < .001. The root mean square error of approximation (RMSEA) is 0.094, and the goodness-of-fit index (GFI) is .993. It is common to find about this much information, with perhaps half a dozen more goodness-of-approximation indices, in most reports of structural equation model studies. Many investigators would regard this information as sufficient. Most would ignore the test of significance and somehow resolve the contradictions between the goodness-of-approximation indices—in this case, “poor” fit by the RMSEA and “good” fit by the GFI. Many would accept the model because the GFI is almost 1.
A few studies might, as recommended here, publish the residual correlation matrix from this composite model. This matrix is given in Table 2 (results above the diagonal). The reader may decide whether to concur with the judgment that these are acceptably small except possibly for the residuals between self-acceptance and frequency of symptoms and between self-acceptance and gastrointestinal symptoms. Recall that our metacriterion is that the fit is acceptable if the residuals do not justify a more complex, reasonable model. Note for future reference that no off-diagonal residual is an exact zero.
Residual Correlation Matrices: Results for Factor Model Fitted to Feist et al.’s (1995) Data Above the Diagonal and Results for the Bottom-Up Model Below the Diagonal
The next step is to fit the measurement (common factor) model, which results in the factor loadings and unique variances in Table 3 and the factor correlations in Table 4 , χ2 = 138.08(67), p < .001. Differencing the chi-squares for the composite model and the factor model results in χ2 = 21.49(2), p < .001, for the causal model. A test of significance would reject both models. We might also compute separate RMSEAs, or any other indices based on the chi-square, for each component of the composite structural equation model. In this case, the RMSEA for the factor model is 0.085, and for the causal model, it is 0.256. We immediately surmise that the causal model would prove unacceptable on more careful investigation, but in the composite analysis its bad behavior was disguised by the acceptable fit of the measurement model, with its much larger number of degrees of freedom.
Results of a Factor Analysis Model Fitted to Feist et al.’s (1995) Data: Factor Coefficients
Note. Standard errors are shown in parentheses.
Results of a Factor Analysis Model Fitted to Feist et al.’s (1995) Data: Factor Correlations
Note. Standard errors are shown in parentheses.
We then go on to obtain the residual correlations for the factor model, which are also shown in Table 2 (below the diagonal). Note that the residual correlations in the factor model are the estimated partial covariances of the (standardized) indicators when the common factors are partialed out. Curiously, the only suspiciously large residual correlation in the measurement model is –.199 between purpose in life and time pressures, which mildly suggests a direct negative relation, open to further investigation. Note that this residual is small in the composite model, whereas the residuals of self-acceptance with two of the Physical Health symptoms are suspiciously large in the composite model but not in the measurement model, which suggests an unsurprising direct causal link between Physical Health and self-acceptance. It would be more surprising if such a link did not exist, and the effect of Physical Health on self-acceptance was purely mediated by World Assumptions and Constructive Thinking.
As a further step, the path (causal) model was fit to the factor correlation matrix, giving the directed arc coefficients and the disturbance covariance matrix in Table 5 . The parameter estimates from the composite model and from the separate analyses for the measurement and path model are in reasonable agreement, given the sample size. The residual correlations are also given in Table 5. Notice that 13 of the 15 residuals are exact zeros. The 2 nonzero residuals are between the indirectly connected variables. These results are to be expected by theory. An approximation index based on all 15 residuals would be inappropriate, given that 13 of them must be zero. The nonzero residuals are the partial covariances of the indirectly connected common factors when the mediating factors are held constant. If the model were “true,” these would be zero in the population. Because the factors are standardized, the fact that the unconditional correlation between Subjective Well-Being and Physical Health is –.186 while the covariance with World Assumptions and Constructive Thinking constant is .249 shows clearly that the postulated mediators do not function as such, and the path model is unacceptable as an approximate account of the data.
Results for the Bottom-Up Model Applied to the Interfactor Correlations
Alternatively, we may estimate the two partial correlations, which, by the path model, should be zero. These are the correlation between Subjective Well-Being and Physical Health with World Assumptions and Constructive Thinking constant, namely, .585, and the correlation between Subjective Well-Being and Daily Hassles with World Assumptions and Constructive Thinking constant, namely, –.117. These values were computed from the factor correlations by classical methods. They can also be obtained from the partial covariances (dividing by the conditional standard deviations, which are just the standard deviations of the disturbances). This method gives slightly different estimates, namely, .598 and –.098, respectively. The only implications of the path model are that these should be (approximately) zero. Checking these constraints should be the primary objective of the investigation, once we are satisfied that the factor model gives adequate measurements of the attributes (see McDonald, 2002, 2004b).
Whereas the residual correlations in the composite model for self-acceptance–frequency of symptoms and self-acceptance–gastrointestinal symptoms mildly suggested a problem with the composite model, the partial correlations give a more precisely focused account of its defective character. From the composite model, without the separation of the measurement model and the path model, we would not have noticed the exact zero residuals implied by theory, and we could have drawn inappropriate conclusions about the path model.
Conclusion
These remarks can be brought to a close with just a final hint about further generalizations. Fitting multidimensional item-response models involves a replay, sometimes not noticed by specialists in item-response theory, of most of the problems and properties of the common factor model (see McDonald, 2000). However, in item-response theory we have a choice between full-information estimation (which is based on the probability of the response patterns) and bivariate information (which is based on pairwise probabilities or on item correlations; see McDonald, 1982, 1997c, and Muthen, 1984). There seems very little to choose between full-information and bivariate-information methods in respect of efficiency of estimation, essentially because there is little reliable information in the higher order relations between items beyond pairwise probabilities. However, from the present standpoint, to assess the goodness of approximation of an item-response model, we seem to need the residuals—covariances or correlations—as supplied by McDonald’s NOHARM or Muthen’s Mplus, and not supplied by every full-information program. (Full-information programs could, of course, supply them.) This advice is particularly directed to those who still seem to take a Rasch position about item-response theory and disallow its richer possibilities.
There are a number of ways to combine a structural equation model with a linear model for experimental data containing multiple response measures, to give an account of the structure of experimental error or to treat common factors as outcomes. McDonald (1980) showed in principle how to combine the Pothoff–Roy analysis of variance (ANOVA) model with a structural equation model. With further development, the treatment of approximation given here could carry over to such models. However, serious problems of measurement scale may arise in applications, because we would not typically wish to standardize response variables to aid judgment of the fitted structure.
It is natural to ask further how far the treatment of approximation given here can carry over to the linear (regression) model with a single response variable. The linear regression model contains fixed (controlled) experimental variables and/or random explanatory variables, accompanied by ANOVA. It can be treated as a path model, with directed arcs from the explanatory variables to the response variable. However, it is an unrestricted path model, and the question of its approximation to reality does not arise in any obvious way, unless we compare it to a nonlinear model or a model with all possible interactions or, indeed, regard it as an approximate description of a determinate world. McDonald (1997a) showed how the F statistic for interaction or other components in a complex ANOVA model can be used to measure the approximation to the full model given by a restricted model with those components omitted. A treatment along these lines for multilevel (hierarchical linear) models seems quite feasible, with measures of the approximation of a single-level model to two-level data, for example. Further work on this problem may be desirable.
In summary, the general principles suggested here for judging the acceptability of a structural model are as follows: (a) Most structural models are mathematical idealizations. Their rules of application, if explicit, will be seen to link them loosely, as approximations, to real behavioral science data. (b) Tests of significance and global approximation indices do not supply a reasonable foundation for judging the acceptability of a model. (c) Approximation is itself multidimensional, and discrepancies between model and data require specific examination and judgment, with uncertainty that cannot be quantified. A possible metacriterion rests on the idea that a model is acceptable if the pattern of the residuals does not justify a more complex, rationally interpretable model. (d) For the common structural equation model, a composite of a factor model and a causal path model, the analytic sequence should include separate analyses of the component models and an examination of the actual constraints implied by the path model (as illustrated). (e) Judgments of acceptability of a model should be supported always by careful substantive conceptualization—doing good thought experiments. The decision should not be handed over to mathematical theorists who are not first psychologists.
Finally, I offer a concession. Those who are made uncomfortable by these conclusions will do no harm to their work if they obtain and report chi-squares and goodness-of-approximation indices, in addition to carrying out the detailed examination recommended here. Doing so will at least supply a security blanket for their discomfort with the responsibility of making good scientific judgments. Of course, placing confidence bounds on parameter estimates will be an indispensable aid to judgment. However, if the analysis stops at the globally fitted model, with global approximation indices, it is incomplete and uninformative.
Footnotes
Notes
Acknowledgments
This article is based on my invited address to Division 5 of the American Psychological Association (Boston, MA, August 14–17, 2008), given in response to receiving the Samuel J. Messick Award for Distinguished Scientific Contributions. My thanks go to Albert Maydeu-Olivares for assisting with the empirical example used in the article.
The author declared that he had no conflicts of interest with respect to his authorship or the publication of this article.
