Abstract
Factor score regression (FSR) is widely used as a convenient alternative to traditional structural equation modeling (SEM) for assessing structural relations between latent variables. But when latent variables are simply replaced by factor scores, biases in the structural parameter estimates often have to be corrected, due to the measurement error in the factor scores. The method of Croon (MOC) is a well-known bias correction technique. However, its standard implementation can render poor quality estimates in small samples (e.g. less than 100). This article aims to develop a small sample correction (SSC) that integrates two different modifications to the standard MOC. We conducted a simulation study to compare the empirical performance of (a) standard SEM, (b) the standard MOC, (c) naive FSR, and (d) the MOC with the proposed SSC. In addition, we assessed the robustness of the performance of the SSC in various models with a different number of predictors and indicators. The results showed that the MOC with the proposed SSC yielded smaller mean squared errors than SEM and the standard MOC in small samples and performed similarly to naive FSR. However, naive FSR yielded more biased estimates than the proposed MOC with SSC, by failing to account for measurement error in the factor scores.
Keywords
Introduction
In the educational, behavioral, and social sciences, researchers often explore relationships between latent variables such as motivation and intelligence. The most obvious and widely adopted tool of choice to analyze these relationships is structural equation modeling (SEM). SEM is a statistical modeling procedure that estimates the measurement model and the structural relations simultaneously, and is (in combination with maximum likelihood estimation [MLE]) often regarded to be the gold standard (Bollen, 1989). However, there are several disadvantages to (traditional) SEM. For example, a large sample size is a desideratum, and a misspecification of the model might affect the estimation of all parameters (Bentler & Yuan, 1999; Boomsma, 1985; Nevitt & Hancock, 2004).
When substantive interest is primarily in assessing the structural relations between latent variables, a straightforward alternative to SEM is factor score regression (FSR) (Skrondal & Laake, 2001). In FSR, the parameters are estimated in two steps. In a first step, confirmatory factor analysis is used to estimate the parameters of the measurement model, and factor scores are predicted for each latent variable. In a second step, the factor scores are used in place of the latent variables to estimate the (linear) relationships among the latent variables. When the structural model is recursive (i.e. no feedback loops), the analysis in the second step boils down to a series of linear regressions (Devlieger et al., 2016). When the structural model is non-recursive, path analysis can be used (Devlieger & Rosseel, 2017). In this article, we used the term FSR for both settings. Factor scores can be calculated in various ways, leading to different sets of observed factor scores. When using Regression (Thomson, 1934; Thurstone, 1935) or Bartlett factor scores (Bartlett, 1937), these are by construction always a linear combination of the observed indicators measuring the latent variable. Therefore, they still contain measurement error and naively using them in place of latent variables will result in biased regression coefficients (Bollen, 1989; Lastovicka & Thamodaran, 1991; Shevlin et al., 1997). We described other issues that can arise when using this method in the “Discussion” section. In this article, we referred to the above-described method—without any form of bias-correction—as FSR.
It is nevertheless possible to remove the bias in the estimated regression coefficients. One possible approach is based on the work of Croon (2002), and termed the method of Croon (MOC) by Devlieger et al. (2016). The MOC proceeds by removing the measurement error from the observed variance–covariance matrix of the factor scores in order to obtain a consistent estimator of the variance–covariance matrix of the latent variables. This (co)variance matrix can then be used to obtain asymptotically unbiased estimators of the regression coefficients. Devlieger et al. (2016) and Takane and Hwang (2018) showed in their simulation studies that the MOC is a good performing alternative for SEM. As illustrated in Devlieger and Rosseel (2017), a major advantage of the MOC is its robustness against local misspecifications of the model.
Although SEM and the MOC are useful methods, both methods can perform poorly in small samples (
Indeed, it is well-known that SEM and the MOC take into account the measurement error, and therefore give unbiased results. What is less known by the users of both methods is the price to pay for these unbiased results. The bias-correction performed by SEM and the MOC leads to a large variability of the parameter estimates, in particular when the sample size becomes small. This is related to the concept of bias-variance trade-off, which implies that a decreasing bias will lead to an increasing variance and vice versa (Hastie et al., 2006; James et al., 2013). Therefore, in addition to bias, we also focus on the estimators’ variability and mean squared error (MSE) in this article. Cox and Hinkley (1979) stress that an estimate with a small bias and a small variance might be preferable to one with no bias but a large variance.
Small sample corrections (SSCs) have already been presented in the general SEM framework (e.g. Ozenne et al. 2020). However, this is not the case for the MOC. Hence, in this article, we propose to modify the MOC with a SSC. Our modification is based on earlier work reported in the measurement error models literature (Carroll et al., 2006; Fuller, 1987; Wall & Amemiya, 2000). In the context of linear regression with measurement error in the predictors, Fuller (1980) proposed two modifications in order to improve small sample estimation. His aim was to reduce the variability by allowing some bias, resulting in a better MSE accordingly. The goal of this article is to incorporate these two modifications into the MOC, and examine its performance.
The rest of this article is structured as follows. First, we introduced a simple structural equation model and the notation. We then described different estimation methods to analyze the relationships between latent variables: traditional SEM, FSR, and the MOC. Later, information concerning measurement error models and Fuller’s modifications is given. Thereafter, we showed how to integrate these different adjustments into a single small-sample correction for the MOC. Subsequently, two different simulation studies are described and the results are presented. Lastly, we discussed the performance of the estimators and the robustness of their performance in various models and provided some concluding thoughts.
A Simple Structural Equation Model
To facilitate the presentation of the correction techniques developed in this article, we considered a structural equation model with one latent predictor (denoted by

A simple structural equation model.
We used this simple model because it allowed us to present many formulas in an easy to understand scalar form. For general settings, the reader can refer to Appendix A where all results are presented using matrix notation. The joint model consists of two models: a measurement model and a structural model. The measurement model can be written as:
where
The structural model is given by:
where
Without loss of generality, we assumed that the observed variables are centered and therefore,
Maximum Likelihood Estimation in Structural Equation Modeling
When all observed variables are continuous, a common estimator in SEM is MLE. When using MLE, the parameter estimates are found by minimizing the discrepancy function as follows:
where
As mentioned earlier, SEM suffers from convergence issues in small samples. Nonconvergence occurs when the applied optimization method (e.g. quasi-Newton optimization) fails to acquire a solution satisfying certain criteria (Anderson & Gerbing, 1984; Boomsma, 1985). De Jonckere and Rosseel (2022) showed that bounded estimation leads to a major decrease in the nonconvergence rate, without negative effects on the point estimates for the unbounded parameters. Given the focus on small samples in this article, bounded estimation (using standard bounds) is used for all estimators in this article in order to avoid convergence problems. Importantly, applying standard bounds prevents the elements of
Factor Score Regression
FSR is a two-step method. In the first step, confirmatory factor analysis (CFA) is used to estimate the parameters of the measurement model. As with SEM, convergence issues may also occur here. Therefore, we again use MLE with bounded estimation for the CFA. Because the two measurement models (for
where
In the Regression method, the factor score vectors
In the second step, we use the factor scores to estimate the regression coefficient
where the “
The Method of Croon
The source of the bias when using naive FSR is that the variance–covariance matrix of the factor scores differs from the variance–covariance matrix of the latent variables. Hence, the MOC (Croon, 2002) performs a correction to the variance–covariance matrix of the factor scores so that the corrected variance–covariance matrix is consistent for that of the latent variables. For clarity, we call the elements of the initial variance–covariance matrix the (co)variances of the factor scores, in this case denoted by
The variance of a latent variable based on the variance of the factor scores can be obtained by the formula:
Furthermore, the covariance of a latent variable based on the covariance of the factor scores can be obtained by:
Thereafter, it is possible to obtain an unbiased estimate of the regression parameter using the corrected variance in Equation (8) and covariance in Equation (9):
The formulas above simplify when the Bartlett predictor is used to compute the factor scores, because in this case:
Fuller’s Modifications for Small Sample Estimation
We now present the two modifications for the small sample estimator suggested by Fuller (1980). Consider a simple linear regression model with one predictor and one independent variable. Using the notation of Fuller (1987), we can write this as:
where the
is unbiased for
where
will now be biased. The impact of the measurement error is attenuation of the regression coefficient toward zero. However, if an estimate of
Note that
where
with
where
Fuller’s correction involves two modifications. First, a decision based on
Combining Fuller’s Modifications and the Method of Croon
In this section, we showed how the two modifications of Fuller (1980) can be incorporated into the MOC estimator. For the considered model, only
where
This decision based on
where
We call the adjusted estimator in Equation (22) with the two modifications the MOC with a SSC. The two modifications can be employed separately by leaving out the decision based on
If we assume
The rationale of the correction is that the variance of the factor scores is subtracted with only a fraction
Simulation Studies
Two simulation studies were performed. In the first study, we examined the operating characteristics of the MOC with SSC estimator as compared to other estimators. We have considered 12 different conditions varying in sample size and reliability of the indicators. Given the focus on small sample estimation, the following sample sizes were chosen: N = 20, 30, 50, 100, 200, and 2,000. In the second study, we assessed the robustness of the performance of the three specifications of
Simulation Study 1
For the first simulation study, two data-generating models based on Equation (A2) were considered. They differed only in terms of the reliability of the indicators. The first model had a reliability of 0.5 for all indicators, whereas in the second model the reliability was set to 0.8. The conditions corresponded respectively to low and high reliability or, in other words, a condition with a higher and a lower amount of measurement error (Brunner & Austin, 2009; Nunnally & Bernstein, 1994). Both models contained three correlated (latent) predictors (

Simulation model. The variances of the latent predictors varied in both models. For the model with low reliability, these were
For each simulated dataset, we fitted the same model in Figure 2. We then compared seven different estimators: (1) traditional SEM (SEM-MLE), (2) the standard MOC, (3) the
To evaluate each estimator, the following performance criteria were considered: the proportion of simulated datasets where the method converged (from here on named the convergence rate), the relative mean bias, the empirical standard deviation (ESD), and the MSE of the estimators. Table 1 summarizes the performance criteria. We simulated 10,000 samples for each of the 6 × 2 = 12 settings.
Overview of performance criteria. For a given simulated dataset, the estimator is denoted by
Simulation Study 2
In the second simulation study, six data-generating models were considered. They differed in terms of the reliability of the indicators, as well as in the number of latent predictors and indicators. The first three models had a reliability of 0.5 for all indicators, whereas in the last three models, the reliability was set to 0.8. Hence, there was again a condition with low and high reliability (Brunner & Austin, 2009; Nunnally & Bernstein, 1994). Models 1 and 4 both contained three correlated (latent) predictors (
Analogous to the first simulation study, we have compared the same seven estimators. Bounded estimation was again used for all estimators to obtain a higher number of converged solutions. In contrast to the first simulation study, the interest lies here in the robustness of the performance of the different specifications of
Results Simulation Study 1
Based on the bias-variance trade-off, a certain pattern in the results was anticipated. On the one hand, we expected SEM and the MOC to be the best performing estimators with respect to bias, but with the worst regarding variability. On the other hand, we foresaw FSR and the SSC with the largest α-value to have the largest bias, but with the lowest variability. Moreover, we expected the bias and variability to converge toward zero as the sample size grows larger. Except for FSR and the SSC with the largest α-value, we anticipated the bias to remain, even in large sample sizes. Finally, it was expected that the patterns will be less distinct when the reliability is high. Tables with the values of the considered performance criteria for the various estimators and settings can be found in the Supplemental Material.
Convergence rate
Because methods (2) to (7) were adaptations of FSR, they would all either converge or fail to converge for the same dataset. Hence, we grouped these methods together only when assessing convergence rates and compared SEM-MLE and FSR. Table 2 reveals a high overall proportion of successful replications. The lowest convergence rates (99.30% and 99.89%) are detected for SEM-MLE in the conditions with low reliability and the smallest sample sizes. In fact, a converged solution is acquired 9,930 out of 10,000 times when
Overview of convergence rates in percentage. Bounded estimation was used for all estimators. The latent variables were handled separately in the FSR methods, implying that convergence was only achieved if the factor analysis was successful for all four measurement models.
Note. FSR = factor score regression; SEM-MLE = structural equation modeling–maximum likelihood estimation.
Bias
As expected, the results plotted in Figures 3 and 4 display that the mean bias decreases when the sample size grows. Similarly, the differences between the estimators become smaller, but remain larger with low reliability. In contrast to SEM-MLE, the MOC, and

Relative bias in case of low reliability.

Relative bias in case of high reliability.
The findings show identical results for the MOC and
Interestingly, the (relative) bias seems to depend on how large the regression coefficient is. If the true value is small (e.g.
Empirical standard deviation
As expected, the results suggest that the ESD reduces as the sample sizes grows. The differences between the estimators become smaller as well, especially in case of high reliability. The findings reveal no major differences between

Empirical standard deviation in case of low reliability (some values lie out of scale, these can be found in the corresponding table in the Supplemental Material).

Empirical standard deviation in case of high reliability.
The results imply that SEM-MLE is the worst performing estimator in the smaller sample sizes (
Mean squared error
As expected, the findings show decreasing MSE values as the sample size grows larger. Figures 7 and 8 indicate that the differences between the estimators become smaller as well, even more so in case of high reliability.

Mean squared error in case of low reliability (some values lie out of scale, these can be found in the corresponding table in the Supplemental Material).

Mean squared error in case of high reliability (some values lie out of scale, these can be found in the corresponding table in the Supplemental Material).
The overall worst performing method is SEM-MLE, specifically in the conditions with low reliability and small sample sizes. From a sample size of 100 onwards, SEM-MLE performs similar to the other estimators. The MOC shows comparable results to the other FSR estimators. However, its performance is poor in the conditions with low reliably and a sample size of 20 or 30. Hence, the
In most conditions, the best performing method in terms of MSE is FSR, followed by the SSC with the largest α-value. On the other hand, FSR is the worst performing estimator when considering the MSE of
Results Simulation Study 2
As in the first simulation study, the findings show decreasing average MSE values as the sample size grows larger. Figures 9 and 10 indicate that the differences between the estimators become smaller as well, especially in case of high reliability.

Average mean squared error in case of low reliability (some values lie out of scale, more information regarding these values can be found in Table 3).

Average mean squared error in case of high reliability (some values lie out of scale, more information regarding these values can be found in Table 4).
Regarding the best and worst performing methods, the same results as in the first simulation study can be observed. The worst performing methods regarding average MSE are SEM-MLE and the MOC. The best performing method is FSR, followed by the SSC. From a sample size of 100 onwards, the methods perform more alike. Additionally, the results reveal once more the relevance of the
In line with the earlier findings, FSR is the best performing method with respect to average MSE in most conditions (i.e. when
Relative average mean squared error in case of low reliability.
Note. FSR = factor score regression; SEM = structural equation modeling; MLE = maximum likelihood estimation; SSC = small sample correction.
Relative average mean squared error in case of high reliability.
Note. FSR = factor score regression; SEM = structural equation modeling; MLE = maximum likelihood estimation; SSC = small sample correction.
As expected, the largest relative average MSE values can be found for SEM-MLE and
Summary of Simulation Studies
In the first simulation study, the overall convergence rates were high for both SEM-MLE and the FSR methods. This high number of converged solutions suggests a beneficial effect from the use of bounded estimation (in this case standard bounds), as discovered earlier by De Jonckere and Rosseel (2022). The
As mentioned earlier, the SSC performed well in terms of MSE in the considered conditions. The rationale of the SSC is to introduce some bias for less variability. This trade-off has a beneficial effect for the MSE in small sample estimation. Out of the three suggestions, the estimator with
SEM-MLE and the MOC were the worst performing estimators regarding ESD and MSE, particularly in the smaller sample sizes. However, both estimators outperformed FSR and the SSC regarding bias in most conditions. SEM-MLE and the MOC were particularly superior in terms of bias in case of low reliability and larger sample sizes. The results also showed decreasing bias of SEM-MLE and the MOC as the sample size grows. Furthermore, the ESD will also reduce, which will eventually (in lager sample sizes) result in a better performance compared to FSR and the SSC.
For the second simulation study, in each of the considered models, highly similar patterns were present in the results. This illustrates the robustness of the performance of the three α-values for the SSC in models varying in the number of latent predictors and indicators. Furthermore, the estimators performed slightly better in the model where latent variables were measured by more indicators.
Discussion
The goal of this article was to (a) evaluate the performance of (i) standard SEM, (ii) the standard MOC, (iii) naive FSR, and (iv) the MOC with the proposed SSC in small samples sizes and (b) assess the robustness of the performance of the different specification for the SSC in different models. We assessed the estimators in terms of convergence rates, relative mean bias, ESD, and MSE. The comparison was performed in 12 different conditions, with varying sample sizes and reliability of the indicators.
The findings reveal FSR as the best overall performing predictor, followed by the SSC with the largest α-value. Only for the conditions where
Although the FSR and SSC estimators had the best performance in most of the conditions, they do not take the measurement error (entirely) into account. This results in a lower MSE in small samples. However, despite the beneficial effect on the MSE, the consequences of not taking measurement into account should not be neglected. On the one hand, the regression parameters can be over- or underestimated depending on the correlation structure between the latent variables (Cole & Preacher, 2014). On the other hand, an inflation of the type I error rate may occur under certain circumstances (Brunner & Austin, 2009). Moreover, the more complex the model becomes, the worse the consequences of both problems.
When analyzing relationships between latent variables, we encourage researchers to not choose a method ill-considered. Instead, we recommend to consider the sample size, complexity of the model, reliability of the indicators at hand, and to use the best method for the situation. If the sample size is large enough (
Like all simulation studies, our study was not without any limitations. The first limitation of this study is that only certain conditions and models were examined. Although these conditions and models may be commonly found in the literature, different settings may lead to different results. The second limitation is that we have not considered the type I error and the power as performance criteria. Future research should look into hypothesis testing to see how well the SSC and other estimators perform in (small) sample sizes, preferably using a different simulation model to see how general the results obtained in this article are.
Supplemental Material
sj-zip-1-epm-10.1177_00131644221105505 – Supplemental material for A Small Sample Correction for Factor Score Regression
Supplemental material, sj-zip-1-epm-10.1177_00131644221105505 for A Small Sample Correction for Factor Score Regression by Jasper Bogaert, Wen Wei Loh and Yves Rosseel in Educational and Psychological Measurement
Footnotes
Appendix A: Formulas in Matrix Notation
Acknowledgements
The authors would like to thank the editor and two reviewers for their comments on prior versions of this manuscript.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Wen Wei Loh was supported by the Special Research Fund (BOF) of Ghent University postdoctoral fellowship BOF.PDO.2020.0045.01.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
