Abstract
A sample suffers range restriction (RR) when its variance is reduced comparing with its population variance and, in turn, it fails representing such population. If the RR occurs over the latent factor, not directly over the observed variable, the researcher deals with an indirect RR, common when using convenience samples. This work explores how this problem affects different outputs of the factor analysis: multivariate normality (MVN), estimation process, goodness-of-fit, recovery of factor loadings, and reliability. In doing so, a Monte Carlo study was conducted. Data were generated following the linear selective sampling model, simulating tests varying their sample size (
For decades, research in social sciences has been facing a ubiquitous limitation: making inferences from a restricted, questionably representative sample to an unrestricted population. The range restriction (RR) is a classical statistical problem that implies dealing with a variable which population variance has been reduced due to a selection process. Pearson (1903) pointed this selection as “probably the chief factor in the production of correlation” (p. 2). Since then, methodological research has explored multiple RR corrections and implications. Mainly, on correlation (Ree & Carretta, 2020), regression slopes (Hunter et al., 2006; Mendoza & Mumford, 1987), and validity and reliability (e.g., Fife et al., 2012; Sackett et al., 2002) as well as more complex techniques, like analysis of variance (ANOVA, Fife, 2013) or meta-analysis (Hunter et al., 2006). Even so, practically no literature has focused on its consequences over factor analysis (FA).
Imagine a variable with its whole possible range of values. Now, two different RRs can reduce its variability: direct and indirect (Hunter et al., 2006; Mendoza & Mumford, 1987). Selecting the data starting from a specific value yields a direct range restriction (DRR). For example, personnel selection commonly uses the test-data exclusively from the already selected applicants—whose score in that precise test exceeded certain cutoff value. Selecting the data, not from the variable of interest, but from another variable related to it (Sackett & Yang, 2000), yields an indirect range restriction (IRR). For example, a researcher might want to measure intelligence in the population, but through a college student sample. The restriction does not occur over intelligence, but over the education level, quite related to it. This selection variable combines any possible source of selection, regardless if it was measured or not, objective or not (Hunter et al., 2006). In most of the research contexts, it is rather impossible to empirically identify nor measure it. Thereby, researchers commonly use it as a theoretical variable, which includes every single variable that took part on the sample selection process, from the most significant to the most subtle one.
ConvenienceSampling
The IRR is the most frequent scenario in applied psychological, employment, and educational research (Thorndike, 1949, as cited in Hunter et al., 2006) through the sampling process. The convenience sampling is broadly used in quantitative studies (Etikan et al., 2016) and, as its name indicates, accounts for the researcher’s accessibility in recruitment and administration (Hanel & Vione, 2016). According to these criteria, the handiest participants for most researchers are, as the reader would immediately identify, undergraduate students.
Students significantly differ from the general population in different aspects. For example, they are younger and more educated, both the most influential variables over attitudes (Sears, 1986). However, this could not be a problem if those variables were appropriately controlled (Hanel & Vione, 2016). Also, some social psychologists assume that convenience sampling has no significant impact on inferences, because their phenomena are universal and do not depend on context (Sears, 1986). In other research fields, representativeness may not be of core importance, like in basic psychology focused on theoretical (Mook, 1983; Peterson & Merunka, 2014) or in consumer psychology focused on students as population of interest (Peterson & Merunka, 2014). Therefore, we must rate the value of convenience sampling depending on the purpose (Mook, 1983).
The generalization—inferring conclusions from the sample to the population—is arguably the most pursued goal in applied research. Sears (1986) pointed out that generalizing from participants who represent the tails of the distribution is problematic, referring to the RR problem. Is the extended convenience sampling compromising this goal? Social psychology (e.g., Sears, 1986), developmental psychology (e.g., Nielsen et al., 2017), cognitive psychology (e.g., Murray et al., 2014), organizational behavior research (e.g., Johns, 1991), counseling psychology (e.g., Worthington & Whittaker, 2006), or even psychological research in general (Arnett, 2008; Thalmayer et al., 2021) have reached similar diagnosis: the convenience samples are highly widespread and biasing a great diversity of psychological conclusions. Peterson and Merunka (2014) compared test-data from different college student samples and showed that the replication power with this sampling is not enough to achieve generalization when using FA. The FA seeks to optimize the covariance matrix between items; so, if RR reduces this variability, most likely it will also affect the FA estimates. Generally, few studies have approached this problem over the FA in simulation studies. Only Fife et al. (2012) generated levels of restricted data, following the item response theory, and found that RR affects the estimation of the reliability of test scores. Thus, more simulations with a larger scope are needed to address the impact of RR over the principal FA outputs (normality, estimation, fit, loadings, and reliability recovery).
The Factorial Model Under RR
Specifically for the unidimensional case, an item response is described as follows:
The observed variable (

Example of Distributions of the Latent Factor (
Now, we have seen that the RR can affect directly (DRR) or indirectly (IRR) the observed variable. For a DRR, the selection occurs over
Linear Selective Sampling Model (LSSM)
The LSSM by Tucker and MacCallum (1997), as the common factor model, accounts for a vector of
where
According to these definitions, the covariance matrix reproduced by the factor model,
As a novelty, the LSSM states that the latent factor can be explained by two different sources of variability regarding the selection process from sampling (4). The first source comes exclusively from the vector of selection variables,
In the LSSM, the selection variables are assumed to be linearly associated with the factors. Although other authors conceived this relation as causal, in a way that scores in
Figure 2 shows the path diagram of one of the models simulated in this work. The lines with one arrowhead represent causal relations, in this case from the latent factor to the six observed variables. The line with double arrowheads represents a bidirectional, noncausal association between this factor and the selection variable.

Path Diagram of the Linear Selective Sampling Model Adapted From the Design by Tucker & MacCallum (1997).
Now, the selection process operates at the population level. In other words, the population can be divided into different subpopulations, depending on whether they satisfy certain selection criteria given by
Regarding the assumptions of the LSSM, first, no correlation between discrepancies and selection variables is assumed, hence for any subpopulation
Also, the homoscedasticity assumption implies that the discrepancies of covariance are invariant across subpopulations:
Taking these definitions into consideration, the factor covariance matrix for a subpopulation
Finally, returning to Equation (3), the application of the LSSM to a certain subpopulation
where matrix
Finally, these authors do not make any assumptions regarding the mean vectors of each subpopulation
Every RR implies a loss of variance by definition. However, changes in the mean would depend on the specific restricted range. In the case illustrated in this study, the missing range always corresponds to the left tail of the distribution, thus moving the mean to the right as the restriction increases. However, in this article we will focus on the covariance structure, which is independent from the mean structure (Brown, 2015; Tucker & MacCallum, 1997). Inspired from the comments of one reviewer, we have illustrated another example of IRR which does not alter the mean vectors (i.e., restricting both tails simultaneously). A brief simulation study with this alternative IRR is reported in the appendix.
Simulation Study
For this work, we simulated test-data following the LSSM, under nine selection ratios (i.e., proportion of population selected in
MVN Assessment
Plenty of statistical analyses assume MVN in data. In this work, we focused on FA, which most applied estimation method, ML, requires observed test-data to be multivariate normal. For the moment, literature has checked the statistical power of the available MVN tests under a wide variety of deviations from this MVN: skewed, leptokurtic, platykurtic, etc. (Alpu & Yuksek, 2016; Ebner & Henze, 2020; Mecklin & Mundfrom, 2005; Romeu & Ozturk, 1993). The present simulation study explored the RR, as what we believe is the most extended—but still unexplored—deviation from the MVN.
Although a wide variety of MVN tests is currently available (Mecklin & Mundfrom, 2005), no single method performs better than any other under every deviation of the MVN. Thus, researchers should apply the more suitable test for the specific non-normality problem that their data might be facing (Oppong & Agbedra, 2016). Since the IRR yields skewed observed distributions, we applied the following MVN tests: Mardia’s skewness (MSkew, Mardia, 1970), Henze–Zirkler (HZ, Henze & Wagner, 1997; Henze & Zirkler, 1990), multivariate Royston (Roy, Royston, 1982, 1983), and generalized Shapiro–Wilk (GSW, Villasenor-Alva & González Estrada, 2009), as they have previously shown good performance under skewed data (Alpu & Yuksek, 2016; Mecklin & Mundfrom, 2005; Romeu & Ozturk, 1993). Also, we applied Mardia’s kurtosis test (MKurt) to serve as a comparative against the previously mentioned tests. We expect that, unlike MKurt, the rest of the MVN tests will be sensitive to IRR to some extent.
Estimation Process
As the IRR is a sampling problem, it is expected to affect the estimation process, causing non-convergence and improper solutions or augmenting the number of iterations needed to convergence. An improper solution includes at least one nonpositive error variance, also known as Heywood cases when these estimates are negative (Dillon et al., 1987). Heywood cases result in standardized loadings out of range (>|1|) and are caused not only by model misspecification, but also by sampling problems like small sizes and variances near zero (Kolenikov & Bollen, 2012). Under IRR, the data suffer a reduction of variance, so we expected more non-convergence and improper solutions as the restriction size increases.
Goodness-of-fit
The GOF is the degree in which the observed covariance matrix, our sample data, fits the reproduced matrix with the estimated parameters by the model. The ML method implemented in available software is the most frequently applied and it assumes MVN. Some violations of this assumption can bias the test statistics to assess GOF, although when data are skewed but non-kurtotic (as it may be the IRR case), the ML statistic shows robustness (Chou et al., 1991). Still, under the plausible violation of the MVN assumption, we used two estimation methods: ML and maximum likelihood mean and variance adjusted test statistic (MLMVS, Satterthwaite style 2 ). MLMVS is a robust version of ML for non-normal data that corrects the χ2 test statistic (and its degrees of freedom) by adjusting its mean and variance with the so-called Satterthwaite approach (Satterthwaite, 1941, as cited in Satorra & Bentler, 1994).
We registered the following indices, according to the recommendations of Brown (2015) and Hu and Bentler (1999): Four absolute fit indices, namely chi-square (χ2), relative χ2 (χ2/df ratio), root mean square error of approximation (RMSEA), and standardized root mean square residual (SRMR) and an incremental fit index, namely comparative fit index (CFI). Lei and Lomax (1991) found that χ2 was the index most affected by non-normality, when compared to CFI. Since the loading estimates must be identical with ML and MLMVS, we expected SRMR to remain equal, because it does not depend on the χ2 (Maydeu-Olivares, 2017), unlike the rest of the indices.
Recovery of Factor Loadings
As in many other methods that include regression, the RR attenuates estimated weights or slopes (e.g., Mendoza & Mumford, 1987). Although this particular condition has not been tested in the FA literature yet, we also expected the recovery of our factor loadings to be poorer as the restriction size increases, due to an increasing attenuation in the estimates. In fact, this is what happens when simulating ceiling effects, which implies a restriction on the measurement scale, reducing the variance similarly (but not equally) as in RR (Schweizer et al., 2019).
Recovery of Reliability
Reliability in the factor model,
Here, we worked with a RR that affects
Fife et al. (2012) distinguish between global and local reliability. The global reliability refers to the reliability for the unrestricted population and the local reliability refers to the reliability for the restricted population. A restricted sample will make a more accurate estimation of the local than of the global reliability. These authors only informed that the local reliability was relatively well recovered, but the bias raised as the restriction size increased (Fife et al., 2012). As the applied researcher often wants to generalize to an unrestricted population, in this work we checked the recovery of the global reliability, comparing the restricted sample estimate with the unrestricted population parameter. As in Fife et al. (2012) for the local reliability, we expected the recovery to be worse as the restriction size increases, but more marked for the global reliability.
To evaluate this, we used three reliability coefficients: α (Cronbach, 1951), ω total (McDonald, 1970), and greatest lower bound (GLB, Jackson & Agunwamba, 1977). Trizano-Hermosilla and Alvarado (2016) explored the recovery of reliability and found that GLB shows a positive bias under normality, but less bias than α and ω under skewed data. We expected a similar behavior here, because of the skewed nature of data under IRR. Since we were restricting our simulations to the unidimensional tau-equivalent case, we also expected α and ω to overlap (McNeish, 2018). However, we included ω because Fife et al. (2012) found that both coefficients frequently differed in their tau-equivalent models.
Method
We followed four steps in exploring the effects of RR on the abovementioned analyses. First, we sampled a certain number of examinees from the latent factor
Extraction of Restricted Samples
We simulated a population of 200,000 examinees, from which we first extracted nine restricted subpopulations, each with a different selection ratio from 10% to 90% of IRR, mimicking the procedure that Fife et al. (2012) and Pfaffel et al. (2016) did. This was attempted by cutting off the selection variable (
Means (μ) and Standard Deviations (σ) of the Unrestricted Population (C100) and Each Restricted Subpopulation (From C90 to C10).
Since these subpopulations have different sizes, and we wanted to control this variable, the second step consisted of extracting one random sample from each restricted population. As a consequence, we got to work with nine restricted samples (restricted from the original population) of the same sample size. In addition, we add a 10th sample, with the same sample size, but no IRR, randomly extracted from the complete population (C100).
We conducted these Monte Carlo simulations using the R language (R Core Team, 2019). Concretely, the specific function rnorm from the basic package stats (R Core Team, 2019) and the function mvrnorm from the package MASS (Venables & Ripley, 2002) were used to generate normal and MVN random variables, respectively. The reader can find a reduced example of how to extract these 10 samples in R (and the resulting histograms) in the Supplemental Material. The whole script for this complete sampling procedure, the simulated data, and the Supplemental Material is available at https://osf.io/b2kse.
Generated Test-Data and Simulated Conditions
The test-data were continuously simulated from a LSSM of
where
According to the LSSM,
where
There have been several simulation studies that have accounted for a compensatory effect between the sample size, the loading size, and the number of items (e.g., Ondé & Alvarado, 2020), so here we wanted to control the three of them. Until now, 10 sample conditions were generated varying their selection ratio or restriction size (
Factor Model Analyses
For each simulated sample, we estimated a unifactorial FA with ML and MLMVS, using the R function cfa from the lavaan package (Rosseel, 2012), which by default determines the metric of the factor by fixing the first item’s factor loading to one. In every scenario, the model’s degrees of freedom were positive.
MVN Assessment
For the MVN tests, we used the R function mvn from the MVN package (Korkmaz et al., 2014), to apply the MSkew, MKurt, HZ, and Roy tests of MVN, and mvshapiro_test from goft (Gonzalez-Estrada & Villasenor-Alva, 2017) to apply GSW. For each condition, we calculated the empirical proportion of rejections (EPR), namely the proportion of samples in which the p value of an MVN test is smaller than or equal to the nominal significance level (
Estimation Process
The estimation process was evaluated by obtaining (a) the proportion of non-convergence and improper solutions (i.e., Heywood cases) and (b) the number of iterations to convergence of the rest of the cases.
Goodness-of-fit
Finally, to evaluate the GOF we used the R function fitMeasures from package lavaan (Rosseel, 2012), and collected the non-robust values of χ2, χ2/df, RMSEA, SRMR, and CFI, as well as the robust versions obtained when using the MLMVS correction. Values of χ2/df lower or close to 3 (Kline, 2015), RMSEA lower or close to .06, SRMR lower or close to .08, and CFI greater or close to .95 suggest good fit (Hu & Bentler, 1999).
Recovery of Factor Loadings
To assess the recovery of factor loadings in each condition, we calculated three measures of correspondence between the theoretical loading (
being
And third, the coefficient of congruence (
Values above |10%| of
Recovery of Reliability
We evaluated the recovery of reliability by calculating the % bias: the estimated reliability,
as
Metamodel Analysis
In order to analyze the results of the simulation from an inferential perspective, we conducted the so-called metamodels (Skrondal, 2000) with SPSS version 25 (IBM Corp., 2017). Each metamodel consists of an ANOVA for the dependent variables that we considered most appropriate according to the descriptive results. In this simulation context, it is common to limit the metamodel to the two- or three-way interactions and exclude the higher-order interactions, since the negligible interpretations and the need of parsimony (Skrondal, 2000). Each metamodel included the sample size (
Results
In this section, we comment each output at a time (MVN, estimation process, GOF, recovery of factor loadings, and reliability). First, we checked the assumptions for the ANOVA: the independence of observations was guaranteed by the simulation nature, the potential heteroscedasticity was not expected to affect the F-statistic (Blanca et al., 2018), and the multi-sample sphericity could not be assumed so we applied the Huynh–Feldt’s correction. We conducted four four-way ANOVAs and four five-way ANOVAs. Table 2 summarizes all the effect sizes of these models and will be used to describe the results obtained in the simulation in the following sections.
Effect Sizes (
Note. Effect sizes higher than
Results of the MVN Assessment
The metamodel of MVN tests was performed to study the differences in the

Empirical Proportion of Rejection (EPR) for the Five Multivariate Normality Tests: Mardia’s Skewness (MSkew), Mardia’s Kurtosis (MKurt), Henze–Zirkler (HZ), Multivariate Royston (Roy), and Generalized Shapiro–Wilk (GSW) Across Convenience Samples.
Results of the Estimation Process
In the most suboptimal scenario (loadings of .50, six items, and 200 cases), 14% of the replications resulted in non-convergent solutions and 15% in Heywood cases. Similar suboptimal scenarios showed lower proportions, but in the rest of them the proportions were zero. In the following sections, these cases were excluded from the analyses, because the number of remaining replications was still sufficient for our purposes.
Regarding the number of iterations (
Results of the GOF
The indices χ2, χ2/df, and RMSEA did not vary across selection ratios and constantly indicated good fit. Similarly, the SRMR also showed values under the .06 cutoff point, but this time the selection ratio slightly reduced the fit, as the loading size increased. For all of these indices, the differences between the ML and the MLMVS estimation methods were barely noticeable. Figures for these indices are included in the Supplemental Material.
On the contrary, CFI (Figure 4) was sensitive to the restriction size, so we applied the metamodel to CFI, but also included RMSEA to compare both performances. Table 2 shows that RMSEA did not obtained any large effect size, while CFI quite the contrary. The three-way interaction effect between the estimator (

Comparative Fit Index (CFI) Across Convenience Samples in Each Simulated Scenario.
Results of the Recovery of Factor Loadings
The recovery of factor loadings was similar whether using

Loadings Relative Bias Across Convenience Samples in Each Simulated Scenario.
Results of the Reliability Recovery
Another five-way ANOVA was conducted for the reliability bias, where the

Reliability Bias for Cronbach’s α, ω, and GLB Coefficients Across Convenience Samples in Each Simulated Scenario.
Empirical Example
To illustrate the attenuation effect in the estimation of factor loadings and reliability under IRR in real data, we took a freely available dataset from the Vietnam Experience Study (accessed via https://osf.io/dbn4k). This dataset contains data from 4,462 participants in 19 intelligence tests which total scores were used by Kirkegaard and Nyborg (2021) to calculate the g-factor scores.
For the complete sample, we estimated the unidimensional model—as Kirkegaard and Nyborg (2021) did—and recorded the factor loadings and reliability (α coefficient). Then we extracted 100 random subsamples of 200 participants from the complete sample (C100) and from four restricted samples (C80, C60, C40, and C20). Restriction was generated by selecting the 200 participants whose factor scores (estimated with the R function factor.scores from psych package) exceeded the percentiles 20, 40, 60, and 80. In addition, we simulated data using the same procedure as described in this article replicating the empirical conditions (same loadings as in the complete sample,
Results of Factor Loadings’ Relative Bias (RB), Factor Loadings’ Root Mean Square Error (RMSE), and % Bias of α, With Empirical and Simulated Data, in Each Level of Restriction (C100, C80, C60, C40, and C20). Empirical Data Were Obtained by Kirkegaard & Nyborg (2021)
Discussion
Although the RR is a classical problem, to our knowledge, previous literature has played little attention to exploring how it affects the functioning of FA. Here we tested the IRR: a subtle but extended form of this restriction. When using test-data in applied research from a convenience sample (most commonly conformed by undergraduate students), we are likely to be dealing with an IRR problem (Hunter et al., 2006): we can expect that the latent factor which we intend to measure is restricted in some extent, but our observed data could remain apparently preserved.
MVN Assessment
Multiple analysis techniques—like ML FA—require checking if we can accept that our data are MVN distributed. Nowadays, applied researchers have access to plenty of MVN tests in common software (in this work, we focused on R language), so it is crucial to ensure their robustness against frequent deviations from this MVN. Do our MVN tests have enough statistical power to detect the IRR problem? Well, the answer is, as always, “it depends.” Regardless of the number of items and sample size employed, most MVN tests would suggest to maintain the MVN in data affected by any restriction size. As we expected, MKurt is not sensitive to the restriction, but neither did HZ nor GSW, which had been found to be adequate detecting skewed distributions (Alpu & Yuksek, 2016; Ebner & Henze, 2020; Mecklin & Mundfrom, 2005). This suggests that the IRR problem cannot be easily equated with a skewness problem.
Nonetheless, MSkew and Roy were found to be more sensitive to the IRR as the loading size increases. When loadings are weak, the contribution of the error component to the observed distribution is higher than the contribution of the latent factor. Under IRR, although the factor is restricted, the error remains normally distributed, so when both are combined to form the observed variable, this will tend to be normal as well. This effect is illustrated in a GIF of figures included in the Supplemental Material. Two basic implications are represented: (a) As the IRR grows (C90, C80, . . ., C10), the observed distribution shifts to the right and (b) as the loading size grows, the observed distribution is more skewed.
Finally, the rise in the sensitivity of Roy when increasing the number of items is in line with the results obtained by Alpu and Yuksek (2016). Although these authors did not analyzed test-data, they found that an increase in the number of variables was associated with an increase in the power of Roy. Surprisingly, although MSkew was expected to perform better (Romeu & Ozturk, 1993), when dealing with the IRR problem, its sensitivity drops as the number of items increases.
Estimation Process
As we expected, the higher the IRR, the more non-convergent and improper cases appear, and the more iterations are needed to converge. This effect is more prompt under conditions with small sample, loading, and test sizes. Ondé and Alvarado (2020) found similar results under asymmetry and described the influence of these variables as a “compensatory effect,” which are present in most of our analyses—mainly with the loading size, which is also our most influential variable, along with the restriction size.
Goodness-of-fit
All the fit indices (χ2, χ2/df, RMSEA, SRMR, and CFI) showed an excellent fit and none were influenced by IRR when using ML. This could be interpreted as a sign of robustness under the particular simulated skewness of the IRR, in line with Chou et al. (1991), among others. But it would not be a correct interpretation, now that we have already shown that many other variables are affected (like recovery factor loadings and reliability). On the contrary, when using the MLMVS correction of the mean and variance of the test statistic (as well as its degrees of freedom), CFI differed. The MLMVS index suggested a worse fit as the restriction rised, mainly when loadings were small. It would be interesting for future research to explore why the MLMVS correction detects the IRR specifically in CFI. Maybe this could lead to the development of a procedure for the applied researchers to detect the restriction size they might be facing in their data.
Previously, we have said that these results could lead us to the conclusion of robustness: The generalized good fit across conditions suggests that the models estimated by the FA fit well the restricted data. However, this does not imply that the parameter estimates are even close to the ones on the unrestricted population. Precisely in our work, these estimates were more biased as the IRR increases. Thus, this is good example of how excellent fit does not necessarily implies an excellent approximation to the population’s model.
Given the nature of the IRR problem, when analyzing the GOF we need to take into consideration the distinction of four different matrices (Tucker & MacCallum, 1997). Regarding the observed data, we can define
Recovery of Factor Loadings
One of our most striking results was that factor loadings are more underestimated as the restriction size augments. The work of Ximénez (2006) illustrates that increasing the loading sizes improves the recovery of factor loadings. Similarly, in our study, the loading size significantly interacts with the restriction size, reducing such underestimation when loadings are larger. Also, the variability in recovery is higher when the loadings are low and the restriction is large, suggesting that the estimation under those conditions may lead to unstable and hard to replicate factors (Ondé & Alvarado, 2020).
One potential explanation may be related to the skewness in the IRR data. These results are in line with Lei and Lomax (1991), who found that the non-normality increased the bias of recovery. However, unless these authors, we did not find an effect of the sample size on the recovery bias, meaning—again—that maybe the IRR cannot be reduce to a non-normality problem. The future research should include an analysis of how the IRR and both estimation methods (ML and MLMVS) affect the standard errors of the parameters, as Fife et al. (2012) suggested.
Recovery of Reliability
Our results regarding reliability are very similar to Sackett et al. (2002)’s study, as shown in their Figure 1 (Scenario B). They also obtained an increasing underestimation of the true reliability as a function of the IRR size. These authors found that this effect smoothed as the true reliability was higher, represented in our work as loading size—the larger the loading, the larger the true reliability. This result is in line with the systematic interaction observed between restriction and loading sizes in our simulation.
While Fife et al. (2012) tested the local recovery (between the restricted population’s reliability and the estimated one with the restricted sample) and found a maximum bias of approximately −10% under the largest restriction, we tested the global recovery (between the unrestricted population and the restricted sample) and found a maximum bias of −40%. This means that, if a researcher is using a restricted sample (e.g., graduate students of certain university), they better make inferences of the true reliability of the restricted population (e.g., graduate students, in general) than of an unrestricted population (e.g., general population).
Regarding the three reliability coefficients in this study, we also found a lower underestimation for GLB than for α or ω, as Trizano-Hermosilla and Alvarado (2016) did. Unlike Fife et al. (2012), we did not find differences between α and ω, probably because of the unidimensionality and tau-equivalence of our simulated data conditions.
Implications and Recommendations
Nowadays, the statistical sciences, in general, and Psychology, in particular, are facing a reproducibility crisis: only nearly 40% of the most influential psychological (experimental and correlational) studies have replicated their original results (Open Science Collaboration, 2015). Hanel and Vione (2016) already pointed to the use of student samples as plausible responsible for this crisis. Hence, it is urgent to conduct studies that evaluate whether this convenience sampling procedure, in particular, and the RR problem, in general, are actually influencing the credibility of our science and, if so, to what extent and under which circumstances.
With this particular way of simulating IRR, we have seen the different step-by-step difficulties the researchers might face when applying unidimensional FA with a convenience sample. First, they would check the MVN assumption. If their true loadings are weak, none of the MVN tests here analyzed would detect an IRR problem. As the loadings rise, Roy—and MSkew in a less extent—could get more sensitive, mainly if the number of items and the sample size are enough. Second, regardless of the restriction size, most GOF indices would suggest an excellent fit (provided the properly specified model) with ML (except CFI when correcting with MLMVS), so they would probably conclude that their model is adequate despite third and fourth steps. Third, depending on the IRR size of their data, the underestimation of factor loadings might even reach more than −50% of relative bias. And fourth, this would lead to a poor recovery of the true reliability that can improve when using GLB instead of α or ω. Let us not forget that, for the moment, we do not have any technique to estimate the extent of their IRR problem, and these issues we are describing would be probably ignored by the applied researchers.
Limitations and Potential Extensions
This work aimed to open up a new field of research regarding this RR problem over the FA. Nonetheless, another research lines regarding censored, truncated, missing data might be overlapped with this RR problem. Their differences, if they even exist, must be first traced. As in any other simulation study, the conditions here analyzed are limited; and thus, we must frame our conclusions accordingly. For future research, we find the following extensions as appropriate. First, to broaden the analyses to multidimensional models that is to multivariate selection and test scenarios, with different correlation values between the factor and the selection variables. In doing this, model misspecification could be simulated as well, like Ximénez (2006) did. Second, to assess other manifestations of the RR, like DRR, double-hurdle IRR (Fife et al., 2012), or range enhancement (Dahlke & Wiernik, 2020, i.e., restricting the values in the middle of the distribution). Third, other data formats can be simulated, like Likert-type scales, more frequent in psychological research. In summary, the future generation of selection scenarios needs to respond to the applied research in real conditions, in order to endow these studies of ecological validity.
The reader has probably heard the saying: “Psychology is the science that studies the behavior of Psychology students.” Modeling psychological events is challenging, and so is modeling the sampling process, necessary—if not essential—to build our knowledge about those psychological events. If we do not want to keep studying “the behavior of Psychology students,” we need to keep studying the consequences of using such convenience samples. Currently, the use of online data collection systems is becoming popular, which, potentially, could allow access to representative samples of the general population (e.g., Oltmanns & Widiger, 2020), as well as the use of quota sampling to represent characteristics of the general population in the sample (e.g., Sorrel et al., 2021). Nevertheless, the online sampling, also affected by self-selection bias, could still be considered as convenience sampling, which attracts younger and more affluent people, with personality differences from normative samples (McCredie & Morey, 2018). Surely, the solution to the problems addressed in this article involves running studies with multiple samples that would allow us to conclude on the replicability and, ultimately, the generalizability of the results obtained. In short, the validation process must be understood as a process of accumulation of evidence.
Footnotes
Appendix
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by Grant PID2019-105177GB-C22 from the Ministerio de Ciencia e Innovación of Spain and Grant SI3/PJI/2021-00258 from Consejería de Ciencia, Universidades e Innovación of Comunidad de Madrid, Spain, through the Pluriannual Agreement with Universidad Autónoma de Madrid.
