Abstract
Descriptive fit indices that do not require a formal statistical basis and do not specifically depend on a given estimation criterion are useful as auxiliary devices for judging the appropriateness of unrestricted or exploratory factor analytical (UFA) solutions, when the problem is to decide the most appropriate number of common factors. While overall indices of this type are well known in UFA applications, especially those intended for item analysis, difference indices are much more scarce. Recently, Raykov and collaborators proposed a family of effect-size-type descriptive difference indices that are promising for UFA applications. As a starting point, we considered the simplest measure of this family, which (a) can be viewed as absolute and (b) from which only tentative cutoffs and reference values have been provided so far. In this situation, this article has three aims. The first is to propose a relative version of Raykov’s effect-size measure, intended to be used as a complement of the original measure, in which the increase in explained common variance is related to the overall prior estimated amount of common factor variance. The second is to establish reference values for both indices in item-analysis scenarios using simulation. And the third aim (instrumental) is to implement the proposal in both R language and a well-known non-commercial factor analysis program. The functioning and usefulness of the proposal is illustrated using an existing empirical dataset.
Keywords
Most applications of the factor analysis (FA) model for item-analysis purposes are carried out under the following conditions: (a) a large number of items, (b) not too large samples, and (c) complex structures (e.g., Floyd & Widaman, 1995; Gerbing & Hamilton, 1996; McCrae et al., 1996). These are the usual conditions at the earlier stages, when the purpose is simply to assess the dimensionality of an item pool without imposing any particular structure (Conway & Huffcutt, 2003; Floyd & Widaman, 1995). However, they are also quite usual at later stages when the purpose is to plausibly assess the structure of a sizable item pool that we already know will be complex (Cattell, 1986; Conway & Huffcutt, 2003; Ferrando, 2021; Gerbing & Hamilton, 1996; McCrae et al., 1996).
A highly defensible approach for specifying and fitting an FA solution in the conditions above is to use an unrestricted factor analysis (UFA) solution, that is, one that imposes the minimum arbitrary constraints to identify an initial solution that will be further rotated (Jöreskog, 1977; Lawley & Maxwell, 1963). While this type of solution is more indeterminate than a restricted solution, it is far more flexible and better accommodates the inherently complex structures that most items have (particularly in personality and attitude measurement) (Cattell, 1986). We shall stress that we consider the term unrestricted more appropriate than the more usual term exploratory. The latter has nothing to do with the identification restrictions but refers more to the purpose for which the analysis is carried out (which is misleading: an UFA solution can be used for confirmatory purposes while many heavily post hoc modified “confirmatory” solutions are indeed exploratory).
The appropriateness of an UFA solution in the scenario above should be assessed by using a multifaceted approach that, first and foremost, requires reasoned reflection (e.g., parsimony, meaning, and interpretability of the solution). In more objective terms, (a) the goodness of model-data fit (GOF), (b) the strength and replicability of the solution, and (c) the determinacy and accuracy of the factor score estimates derived from it should be the main facets to be considered (Ferrando & Lorenzo-Seva, 2018; Garrido et al., 2016; Reise et al., 2013). Of these facets, this article focuses on GOF, the most basic, which, in principle, can be assessed by using the three broad types of procedures and/or indices developed for structural models in general (e.g., La Du & Tanaka, 1995; McDonald & Mok, 1995). The first type are the test statistics based on the (central and/or non-central) chi-square distribution. The second type are goodness-of-fit statistics that are explicitly defined through the overall test statistic, but which aim to quantify the degree of misfit of the fitted solution. Indices of these two types depend on the estimation procedure chosen and rely on a formal statistical basis. In contrast, the third type of indices/procedures, which are those considered in this proposal, are non-inferential and descriptive, and do not require a formal statistical basis for model fitting. Therefore, they do not directly depend on any test statistic or on the estimation criteria (e.g., Bentler & Bonett, 1980; La Du & Tanaka, 1995). If it is admitted that the joint nature of model fit assessments is simultaneously descriptive and inferential (La Du & Tanaka, 1995; Tanaka, 1993), then this third type of indices do have a place in GOF assessment, if only as adjuncts to the more inferentially based indices of the first and second type, or as the only alternative when the model is fitted with a procedure that lacks a rigorous statistical basis (Bentler & Bonett, 1980).
While the summary above is applicable to any structural solution, its application to UFA solutions in the item-analysis scenario considered here presents some particularities that should be taken into account. To start with, in the UFA context, a GOF investigation assesses only the “optimal” number of common factors that are needed to adequately reproduce the inter-item correlation matrix (Lawley & Maxwell, 1963; Steiger et al., 1985). So, the GOF investigation can be based on the initial (direct or unrotated) solution because the fit results do not change when this initial solution is further rotated.
Determining the optimal number of common factors is a key decision that remains “an imperfect art,” especially in item analysis (Clark & Bowles, 2018). The starting reason for this is that an UFA solution never holds exactly in the population; it provides, at best, an approximation to the data matrix. So it is pointless to try to determine the “true” number of common factors. Rather, what the analyst should try to determine is the number of factors that are worthwhile to retain (i.e., the “optimal” number). As mentioned earlier, this judgment should be based on multiple criteria; however, in pure GOF terms, two main sources of evidence are needed: acceptable overall fit and no substantial incremental fit when more factors are added (Steiger et al., 1985).
Judgments that deviate from the optimal number of factors can go in both directions: underfactoring and overfactoring, and is worth to summarize the consequences of both types of misspecifications (see, for example, Clark & Bowles, 2018, or Ferrando & Lorenzo-Seva, 2018). Underfactoring tends to produce too broad factors, biased/distorted item parameter estimates, loss of information, and factor score estimates that lack univocal interpretation and which reflect the impact of multiple sources of variance. Overfactoring, on the other hand, might give rise to weak or specific factors of little substantive interest, make theory unnecessarily complex, and yield factor score estimates that are indeterminate and unreliable, and which cannot provide accurate individual measurement.
The GOF recommendations given above become particularly complex in item-analysis UFA applications, which have traditionally been based on methods that lack a strong or a formal statistical basis (but that, even so, can be quite defensible in this scenario; for example, Ferrando & Lorenzo-Seva, 2018; Jöreskog, 1977, 2003). So, the non-inferential procedures and indices discussed above are perhaps the most common in this domain. Well-known examples are parallel analysis (Timmerman & Lorenzo-Seva, 2011), percentage of explained common variance (ECV; Ferrando & Lorenzo-Seva, 2018; Reise et al., 2013, ten Berge & Kiers, 1993), and indices based on the magnitude of the residual covariances such as root mean square residual and goodness of fit index (e.g., Ferrando & Lorenzo-Seva, 2018; McDonald & Mok, 1995; Reise et al., 2013). Now, most of these commonly used indices are overall indices that assess the extent to which a given solution, with a specified number of common factors, fits the data well (first source of GOF evidence). In contrast, descriptive difference or incremental indices (Steiger et al., 1985) that assess the extent to which fitting the most complex of two competing UFA solutions substantially improves model-data fit (second source of GOF evidence) are far scarcer. As an example, in the scenario considered here, consider an application in which we want to assess whether 3 or 4 factors are needed to appropriately explain the inter-item correlations. A reasonable GOF strategy is, in the first place, to separately assess the two competing solutions (i.e., overall fit), and, provided that the fit of both is reasonable, to then proceed to assess whether the difference in fit (or in explained covariation) between both solutions is nontrivial, or, more specifically, whether the more parameterized 4-factor model provides a significantly better fit (or explains the data better) than the more constrained 3-factor model.
Purposes
In a set of recent papers, Raykov and collaborators (Raykov & Calvocoressi, 2021; Raykov et al., 2022) have proposed a class of effect size (ES) measures intended to help choose the most appropriate FA solution among competing alternatives. In other words, a class of descriptive GOF difference indices. More specifically, for nested solutions (when UFA solutions with a different number of factors are compared), Raykov et al. (2022) proposed an ES index, initially intended for confirmatory factor analysis solutions, but which can also be used in the UFA case. We shall focus here on the simplest (unweighted) version of this index when used in the difference-GOF assessment described above for comparing UFA solutions with a different number of factors.
As discussed later, the Raykov et al. measure can be viewed as an absolute measure of improvement. Now, the initial aim of the present note is to propose a relative version of Raykov’s ES index. As explained later, the relative measure we propose has important relations with well-known overall measures of appropriateness and can be computed in an extremely simple way. In spite of the simplicity of what we propose, both the index and the relations obtained are, to the best of our knowledge, new contributions.
A second general aim is to use simulation to assess the behavior of both the original and the proposed indices in the scenario considered here: item-analysis purposes and complex conditions, and next, to derive from the simulation results cutoffs or reference values for both indices.
Finally, the third aim is instrumental: both indices were implemented in a popular factor analysis program. The software computes point estimates, bootstrap confidence intervals (CIs), and threshold values to help to decide if the most parametrized model adds substantial information to the less parametrized factor model. In addition, we also implemented the index in R language.
Background Results and the Existing Effect-Size Index
In canonical or principal-axis form (e.g., Harman & Jones, 1966), the structural UFA model in the population is
where
The modeling in Equations 1 and 2 is quite comprehensive and can be applied to (a) binary variables, (b) graded variables treated as ordered-categorical, and (c) graded or more continuous variables treated as continuous. In case (a), the elements of
Denote now by M0 and M1 the two competing UFA solutions that we want to compare, so that the M0 solution is the simpler and the more restricted. If the comparisons are sequential (e.g., Jöreskog, 1977; Raykov et al., 2022; Steiger et al., 1985), M0 will generally be a solution of type (1) with p factors, whereas M1 will be a solution of the same type with p+1 factors.
Next, let the m items be indexed as j = 1 . . . m. Let
Conceptually, EM is the average item communality reproduced by the UFA solution considered (0 or 1 in our case). Raykov and collaborators defined the index more generally as an average R2 index. However, for a canonical (orthogonal) standardized exploratory factor analysis (EFA) solution such as the one considered here, both formulations are equivalent (e.g., Raykov & Calvocoressi, 2021, Equation 2). As for its interpretation, EMD quantifies the increase in the average ECV that is obtained when the more parameterized solution is fitted. So, (a) has a clear interpretation (it is an ES measure), (b) is purely descriptive and, therefore, unaffected by the process of statistical inference and its outcomes, and (c) can be used with any UFA estimation procedure.
In GOF terms, EMD in Equation 3 can be regarded as an (a) incremental, (b) normed, but (c) absolute difference index (e.g., Bentler & Bonett, 1980; Tanaka, 1993). It is incremental in the sense that it assesses the improvement in fit (in this case in average reproduced common variance) when going from solution M0 (with p factors) to solution M1 (generally with p+1 factors). It is normed because the index can be viewed as a positive difference between two average R2 and so gives values in the 0-1 metric. Finally, it is absolute because it does not depend on a relative comparison to an estimate of the overall
The Present Proposal: A Relative ES Index
In this section, we propose a relative version of the EMD index in which the increase in the average ECV is related to the overall amount of common variance contained in the dataset. As an estimate of this overall common variance (i.e., communality), which can be obtained without specifying a particular solution in p factors (see, for example, ten Berge & Kiers, 1993), we used the reproduced communality based on a large number of factors. For example, the specific number of factors to be considered could be the number that accounts for at least 99% of the common variance available in the data (i.e., that if one or more factors is added, the improvement in the common variance would be less than 1%). In no case, however, would we advise to specify more than the number of factors defined by Ledermann’s (1937) bound, and the reason is to avoid solutions with identification and determinacy problems. Ledermann’s bound when solved for the number of common factors can be interpreted as the maximum number of common factors that can be determined from the observed variables in terms of still leaving a number of degrees of freedom greater than zero, and so, allowing the model to be testable (see Hayashi & Marcoulides, 2006, for a clear discussion). In the same line, ten Berge (1998) noted that the number of common factors that is required to account for the overall common variance of the variables is close to Ledermann’s bound, but is usually too high to be of interest. More important, these solutions will generally not be identified.
Using obvious notation, we shall denote the total communality as
So, in principle, GOLDEN is a relative or normalized version of the EMD index with respect to the average total communality. In other words, it is the ratio between the average increase in communality when going from the M0 to the M1 solution and the average of the total communality in the dataset. Note that, if a dataset consists of “almost perfect” items with very small amounts of measurement error, they will both have similar values. However, in a far more fallible dataset, with a substantial amount of item measurement error, the same increase in ECV will result in a far larger GOLDEN value, as this increase is a more substantial proportion of the amount of common variance in the dataset.
GOLDEN also has interesting relations with the popular Explained Common Variance (ECV) index mentioned above. Indeed, for the two competing solutions, the ECV index is obtained as (Ferrando & Lorenzo-Seva, 2018; Reise et al., 2013; ten Berge & Kiers, 1993)
And using simple algebra, it is found that
So, our proposed relative index is also the difference between the ECV indices of the two competing solutions, with the added bonus that a point estimate can be easily obtained from the separate outputs of these solutions (Ferrando & Lorenzo-Seva, 2018).
Taking Uncertainty Into Account: CIs
So far, the EMD and GOLDEN indices have been discussed at the population level. In empirical research, when one of them is computed, a point estimate is obtained that must be complemented with a CI. In particular, in the present case, we have implemented Bootstrap-based CIs for both indices.
For the type of application considered here, CIs provide immediate and useful information. So, for both EMD and GOLDEN, if the zero value does not lie outside the CI, then there is no point in further trying to interpret the point estimate, as it does not even reach significance. Used in this way, the CI works as a safety mechanism to prevent spurious gains from being interpreted. CIs can also be used for a more lenient or strict use of the proposed cutoffs: for example, requesting the associated CI to be positioned entirely above the proposed cutoff (Lorenzo-Seva & Ferrando, 2021; Raykov et al., 2022). This point is further discussed later.
Simulation Studies
As mentioned earlier, the study in this section focuses solely on graded-response items treated as ordered-categorical variables, which, in our experience, is the most common in applications. Hopefully, the binary and continuous scenarios will be considered in future studies.
The Monte Carlo simulation study was carried out using samples drawn from a true population model. Using an obliquely rotated solution, a population loading matrix was defined. The parameters that were manipulated in the simulation study are as follows:
True number of factors (p): between 2 and 5. The inter-factor correlation values were uniform randomly drawn from the range [−.5; .5].
Number of observed variables (m): for each factor in the loading matrix to be generated, the number of variables in each factor was uniformly and randomly drawn from the range [5; 30]. Therefore, the number of variables does not perfectly correlate with the number of factors. For example, one loading matrix could be defined by 30 variables and two factors, while another could be defined by 25 variables and five factors.
Communality: the communality of each variable was uniformly and randomly drawn from the interval [.60; .95], and the value of salient loadings was uniformly and randomly chosen from the interval [.55; .90]. Non-salient loadings were uniformly and randomly drawn from the interval [−.10; .10].
Sample size: the number of observations was uniformly and randomly drawn from the interval [300; 1,000].
For each population loading matrix, we checked that no Heywood cases were present, and the population correlation matrix was positive definite.
Once the true factor model had been generated, random samples were generated based on the true factor model, and each data sample was discretized to have five response categories. In order to factor analyze the data, the corresponding polychoric correlation matrix was fitted using unweighted least squares (ULS; for example, Jöreskog, 1977). So, the simulation study is solely focused on case (b) as defined above.
Once the sample was ready, we compared the factor solution based on the true number of factors with all the possible solutions with fewer factors, and with the same number of solutions with extra factors. For example, if the true number of factors was 3, the three-factor model was compared with the one-factor model and with the two-factor model; in addition, the three-factor model was also compared with the four-factor model, and with the five-factor model. The comparisons were made using EMD and GOLDEN indices.
After 1,000 replications, we computed the mean and the standard deviations of both indices. Outcomes related to EMD are shown in Table 1. The vertical line differentiates underfactoring, that is, factor solutions based on too few factors (when compared with the true factor model) from overfactoring, that is, solutions with extra factors. Note that, for underfactoring to be detected, the EMD values before the vertical line should be substantial, as increasing the complexity of the solution is expected to add considerable amounts of common variance. On the other hand, for overfactoring to be detected, the EMD values after the vertical line should be close to zero, as the most complex solutions based on extra factors are not expected to add considerable amounts of common variance. The inspection of the threshold values of the various population models (based on a true number of factors between two and six) seems to indicate that a threshold value of .045 should help to decide when the most complex solution does not add a relevant amount of common variance. However, it can be observed that the threshold value is related to the number of factors in the models compared: the more factors there were, the lower the threshold value was in the simulation study. In addition, when the number of true factors was low (two and three), the standard deviation was large for the complex solutions that were not expected to add a considerable amount of common variance.
Averages and Standard Deviations of the EMD Index When Comparing the True Factor Model With Alternative Factor Models
The outcomes for the GOLDEN index are shown in Table 2. Again, before the vertical line, the values of the GOLDEN index should be substantial while after the vertical line they should be close to zero. The inspection of the threshold values of the different population models (with between two and five factors) seems to indicate that, as it occurs with the EMD index, the threshold value is related to the number of factors at hand.
Averages and Standard Deviations of the GOLDEN Index When Comparing the True Factor Solution With Alternative Factor Solutions
In conclusion, while the results in Tables 1 and 2 suggest that “omnibus” reference values could be proposed for both indices, it seems preferable to obtain a specific threshold for each dataset under analysis that takes into account the characteristics of the solution (mainly the number of common factors involved). To obtain this empirical threshold, we generate a “population” dataset in which the less parameterized of the two competing solutions is correct. So, in this population, the most parameterized solution does not add substantial information. Details are provided below.
Once an UFA solution of the type M0 has been fitted to an observed correlation matrix
where
If
produces a “sample” data matrix
Guidelines for Using the Indices
Raykov et al. (2022) proposed obtaining 95% CIs for EMD and considered that substantial evidence of improvement is that both ends of the CI were above 0.10. Moderate evidence being point estimates between 0.05 and 0.10. Weak evidence is that both ends of the CI are below 0.05. In this case, we would clearly choose the less parameterized M0 solution.
Instead of using an arbitrary threshold value that can be used in any situation, we empirically estimated a threshold value that takes into account the characteristics of the dataset analyzed. To obtain the threshold value, we propose the following procedure: (a) take the M0 solution as a population solution in which the model exactly holds; (b) produce a random sample of size N based on the M0 solution; (c) factor analyze the random sample to fit both EFA solutions (i.e., M0 and M1 solutions); and (d) compute EMD and GOLDEN indices and consider the value obtained as the threshold value for each index. The CIs around the sample estimate can then be compared to the empirical threshold value in the way Raykov et al. (2022) suggested, that is, the CI should be completely above the empirical threshold if the gain is to be considered significant.
We also propose assessing the replicability of the decisions obtained from a calibration sample by making further analyses in a second sample, to avoid capitalization on chance. When the sample available is large enough to be split into two subsamples (see, for example, Lorenzo-Seva, 2022), assess whether the decision that was taken in the first subsample is replicated in the second subsample.
Implementation
All the indices studied in this article have been implemented in version 12.04 of the program FACTOR (Ferrando & Lorenzo-Seva, 2018), a well-known, free EFA program that can be downloaded at http://psico.fcep.urv.cat/utilitats/factor/. The implementation uses bootstrap resampling to obtain the CIs, and the intensive resampling explained above to determine the empirical threshold. GOLDEN index has also been implemented in R. In this case, the researcher must update in the code the name of the data file, the number of factors in each factor model, and the number of samples to compute the bootstrap CIs. Interested researchers can obtain the R code from the authors by email.
Illustrative Example
A set of 24 items from the Statistical Anxiety Test (Vigil-Colet et al., 2008) was used for the illustrative example. All the items are rated on a 5-point graded format and are designed to measure three factors: Examination anxiety (eight items), Asking for help anxiety (eight items), and Interpretation anxiety (eight items). In the validation study, a sample of 459 undergraduate students from the first year of a bachelor’s degree in Psychology, between 18 and 60 years old (M = 20.4, SD = 4.2), answered the test.
The inter-item polychoric correlation matrix was found to be positive definite and showed a Kaiser–Meyer–Olkin (KMO) value of .864. This matrix was factor analyzed using robust unweighted least squares (RULS). The GOF results obtained are shown in Table 3. They were all based on fit statistics that are explicitly defined through the overall test statistic (i.e., first and second type above).
GOF Results for the UFA Solutions Considered
Note. GOF = goodness of model-data fit; UFA = unrestricted or exploratory factor analytical; RMSEA = root mean square error of approximation; CFI = comparative fit index; NNFI = non-normed fit index.
In addition, Table 4 provides the RMSEA-D statistic for difference test comparisons. Essentially, this is a modified version of the RMSEA index but based on the chi-square difference test statistic not on the overall statistic. The guidelines for interpreting its values are the same as those of the original index (see Savalei et al., 2023, for details).
GOF Difference Assessment Based on RMSEA-D
Note. GOF = goodness of model-data fit; RMSEA-D = root mean square error of approximation of difference.
Parallel analysis suggested that two common factors should be retained.
On the whole, the results of the GOF assessment made in the validation study can be summarized as follows. On one hand, the overall fit statistics suggest that the solutions with two or more factors have an excellent fit. On the other hand, the root mean square error of approximation of difference (RMSEA-D) incremental index suggests that the fit is significantly improved when going from 1 to 4 common factors. So, the most plausible decision is that the statistical anxiety test structure is multidimensional. However, the overall fit results in Table 3 also suggest that this structure might be reasonably well approximated by a unidimensional solution, and two additional pieces of information point in this direction. First, the inspection of the pattern showed that it exhibited positive manifold (i.e., all the loadings were substantial and positive). Second, the ECV estimate was .789 (bootstrap 95% CI = 0.764 and 0.815), above the 0.70 usual cutoff for judging a unidimensional approximation as reasonable (Raykov & Bluemke, 2021; Reise et al., 2013). Indeed, approximating a multidimensional structure by a unidimensional solution potentially presents the underfactoring problems discussed above: biased “mixture” loading estimates that would reflect the differential influence of the local factors, loss of information, and scores reflecting the influence of multiple sources of variance. The positive manifold pattern and the high ECV estimate, however, suggest that neither the loading biases nor the multiple sources of variance invalidate the use of general scores, which would still have a fairly univocal interpretation in terms of a broad dimension of statistics anxiety, which dominates the most local factors that also determine the item responses (Raykov & Bluemke, 2021). Examination of the item contents (which are available in Vigil-Colet et al., 2008) also suggests the tenability of this interpretation. As for the potential loss of information, this point is discussed below.
We turn now to the proposals in this article, which were applied to compare the solutions most in agreement with the theory from which the test was developed, that is, a single general factor versus a three-factor structure. The indices considered in this article, together with their threshold values and CIs, are displayed in Table 5 and provided in the output of version 12.04 of FACTOR.
Non-Inferential Goodness-of-Fit Indices for Nested Comparisons Based on Non-Linear Factor Analysis
The results in Table 5 are quite clear. In both indices, the zero value falls outside the CI, which allows us to exclude from the outset the idea that the gain in explained variance is an artifact. Next, and also in both indices, the CIs are entirely above the empirical threshold value obtained, as explained above. The interpretation provided by GOLDEN is the clearest: the increase in ECV when going from 1 to 3 factors is between 20% and 24% of the total common variance in the dataset, which appears to be a substantial gain. So, the statistical significance of this comparison (which can be obtained by using the chi-square difference test statistic or the derived RMSEA-D in Table 5) does not seem to reflect a trivial or a minor gain (see Raykov et al., 2022), but a very substantial one. Therefore, while it still seems reasonable to use a total score based on all the test items as an approximation (e.g., to obtain a quick estimate of the overall level of test anxiety of a student), the present results suggest that the loss of information would be considerable.
Discussion
Assessing the fit of an UFA solution intended for item-analysis applications is a complex issue that presents particular problems. Items are generally complex, and item pools may be quite large while samples are generally not. In these conditions, simple FA procedures that lack a rigorous statistical foundation but which are robust can be a viable option (Gerbing & Hamilton, 1996), and descriptive GOF indices that do not require a formal statistical basis are particularly appealing. They can be used when the UFA estimation procedure lacks the bases for obtaining more rigorous fit indices, as well as useful auxiliaries of these indices when they are available.
Evidence for determining the optimal number of factors in an UFA solution requires both overall fit and incremental or difference fit to be assessed (e.g., Tanaka, 1993). Now, while there is an abundance of overall descriptive GOF indices that have been traditionally used in the type of UFA applications considered in this article, virtually no incremental descriptive indices have been proposed in this domain. Therefore, what we propose here fills a gap in the panoply of tools which is available to the applied researcher, and so, its interest seems well justified. We first considered an existing index of this type, studied how it functioned in the conditions on which this proposal is focused, and proposed cutoffs, reference values, and guidelines for its use. Second, we proposed a relative version of this index, obtained interesting results that directly relate it to the difference in ECVs, and also proposed reference values and guidelines for its use in item-analysis scenarios. Overall, we believe that our relative version provides a clearer interpretation than the original index.
In a somewhat simplistic way, the indices considered here can be used in two ways. First, if the UFA fitting procedure has a formal statistical basis, then rigorous test statistic indices as well as GOF indices derived from them can be used for appraising whether increasing the number of factors significantly improves fit (as we did in the illustrative example). If this is so, the indices considered here can be next used as a source of additional information in the same way that any ES measure is used: to appraise the practical relevance of the significant effect which was found. Second, if no formal statistics can be derived from the UFA fitting procedure, then the difference indices considered here become the main source of information regarding incremental fit. If this is the case, we would note that we have also proposed here procedures for judging the significance of the difference effects which were found in the form of CIs.
The results of the simulation study suggested that to propose a single, omnibus cutoff or reference value for both indices, although feasible, would be too simplistic. Instead of this, we followed the initial rationale of Bollen and Stine (1992) and proposed using an empirical threshold that takes into account the characteristics of the solutions that are compared. In principle, the thresholds obtained so far are meaningful and plausible, but much more research is still needed on this issue.
In the same vein as above, and more generally, this is an initial proposal, so it has its share of limitations and points that deserve further study. As far as its limitations are concerned, we have avoided theoretical developments (mainly when deriving CIs and threshold values) and focused on simulation, which is perfectly feasible with the equipment that exists today. And, in terms of further study, further intensive simulation beyond the conditions considered here is clearly warranted. Thus, the simulation study carried out here needs to be extended to the linear UFA model for continuous responses, and also to the extreme, binary case, of ordered-categorical modeling. We fully acknowledge that, in terms of guidelines and reference values, what is proposed here must be considered as tentative and will be updated as more information comes available.
In spite of these shortcomings, we believe that what we propose is of interest, can be widely applied, and will be useful for UFA practitioners who work in item-analysis applications. Both the existing and the new index are simple, easy to interpret (particularly the relative index), and provide a more practical view of what is gained when solutions are fitted with different numbers of factors and not formal inferential statistics exist. Finally, one of the strong points of our proposal is the instrumental one. Everything that is proposed here is implemented as a resource in a free, widely known, and user-friendly UFA program, and R.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This project has been made possible by the support of the Ministerio de Ciencia e Innovación, the Agencia Estatal de Investigación (AEI), and the European Regional Development Fund (ERDF) (PID2020-112894GB-I00).
