Abstract
Forced-choice (FC) questionnaires have gained increasing attention as a strategy to reduce social desirability in self-reports, supported by advancements in confirmatory models that address the ipsativity of FC test scores. However, these models assume a known dimensionality and structure, which can be overly restrictive or fail to fit the data adequately. Consequently, exploratory models can be required, with accurate dimensionality assessment as a critical first step. FC questionnaires also pose unique challenges for dimensionality assessment, due to their inherently complex multidimensional structures. Despite this, no prior studies have systematically evaluated dimensionality assessment methods for FC data. To fill this gap, the present study examines five commonly used methods: the Kaiser Criterion, Empirical Kaiser Criterion, Parallel Analysis (PA), Hull Method, and Exploratory Graph Analysis. A Monte Carlo simulation study was conducted, manipulating key design features of FC questionnaires, such as the number of dimensions, items per dimension, response formats (e.g., binary vs. graded), and block composition (e.g., inclusion of heteropolar and unidimensional blocks), as well as factor loadings, inter-factor correlations, and sample size. Results showed that the Maximal Kaiser Criterion and PA methods outperformed the others, achieving higher accuracy and lower bias. Performance improved particularly when heteropolar or unidimensional blocks were included or when the questionnaire length increased. These findings emphasize the importance of thoughtful FC test design and provide practical recommendations for improving dimensionality assessment in this format.
Determining the number of dimensions underlying item response data is a cornerstone of psychometric research, essential for ensuring valid measurement and meaningful interpretation in psychological assessment. Dimensionality assessment is central to exploratory factor analysis (Garrido et al., 2013; Preacher et al., 2013) and underpins the development and validation of assessment tools. Although numerous methods exist, dimensionality detection remains a complex and critical challenge (Auerswald & Moshagen, 2019; Finch, 2020). While extensively studied for some response formats, its application to forced-choice (FC) questionnaires presents unique challenges due to their ipsative nature and complex factorial structure, and remains largely underexplored. This study addresses this gap by systematically evaluating dimensionality assessment methods for this type of data. Therefore, this study aims to (1) assess the performance of dimensionality detection methods in the specific context of FC data, and (2) identify the design factors that are crucial for accurately recovering dimensionality.
FC questionnaires have gained popularity as an alternative to traditional Likert scales, which are widely used but prone to response biases, such as social desirability, acquiescence, and extreme or midpoint response styles (Couch & Keniston, 1960; Paulhus, 1991). These biases distort score validity and inflate convergent validity estimates with other measures affected by the same biases (Soto & John, 2019; Soto et al., 2008). FC questionnaires address these issues by presenting respondents with blocks of statements paired in social desirability, requiring them to rank or rate items relative to one another. To facilitate a distinction, let us refer to block as a set of statements of stimuli in the FC format (generally, by design, each statement in a block measures a different dimension, so blocks are usually multidimensional), whereas by item we denote a single administered statement or stimulus as in traditional Likert response format. Here, we focus specifically on FC questionnaires of blocks consisting of two statements.
Two main response formats exist for pairwise FC comparisons: binary FC (BFC), where respondents select the statement that best represents them, and graded FC (GFC), which allows respondents to position themselves between statements using multiple graded categories (see Figure 1). A limitation of BFC questionnaires is that they produce dichotomous outcomes, often resulting in lower reliability estimates compared to Likert scales (Brown & Maydeu-Olivares, 2017; Zhang et al., 2023). In contrast, GFC questionnaires combine the benefits of improved control over social desirability and clearer differentiation between responses, potentially enhancing reliability (for a comprehensive review of these formats, see Brown & Maydeu-Olivares, 2017, 2018; Zhang et al., 2023).

Example binary and graded forced-choice blocks.
Ipsativity and FC Questionnaire Modeling by Confirmatory Models
Ipsativity is one of the main challenges for scale scores derived from FC questionnaires, as it has led to concerns and recommendations against applying traditional psychometric and reliability analyses to these measures. It refers to the inherent dependency among trait scores, where increases in one trait result in a decrease in others, making cross-individual comparisons inadequate. This dependency arises because endorsing one statement in a block automatically means not endorsing the others in that block. In extreme cases, such as when scores are calculated by summing the number of times a dimension is chosen, the result is a purely ipsative scenario, where the sum of trait scores remains constant across all individuals (Clemans, 1966). This property produces a rank-deficient covariance matrix of scale scores, invalidating the use of standard factor analysis techniques. For instance, the ipsative nature of FC data introduces systematic biases in the intercorrelations among scale scores, with the expected average intercorrelation becoming negative. Such biases can lead to artifactual bipolar factors, where theoretically distinct dimensions collapse onto the same factor but in opposing directions (e.g., Dunlap & Cornwell, 1994). These distortions compromise the validity of the factor structure, limiting the interpretability and utility of the resulting raw score scales.
To address these challenges posed by ipsativity, several item response theory (IRT) models have been proposed, including the Thurstonian IRT model (TIRT, Brown & Maydeu-Olivares, 2011), the Multi-Unidimensional Pairwise Preference model (e.g., MUPP-2PL; Morillo et al., 2016), and the ordinal factor analysis for graded-preference data (Brown & Maydeu-Olivares, 2017). These models enable between-individual comparisons by estimating latent traits while accounting for the dependencies introduced by the FC format. By modeling block properties, such as loadings and thresholds, IRT approaches reduce ipsativity by allowing the relationship between item endorsement and trait scores to vary across blocks. Simulation studies consistently show that IRT-based scoring methods outperform classical scoring approaches in recovering traits (Hontangas et al., 2016; Lee et al., 2022). In the context of pairwise FC comparisons, the TIRT model is equivalent to traditional confirmatory factor analysis (Brown & Maydeu-Olivares, 2011).
Assume that responses to a pairwise block are coded on an ordinal scale from 1 (maximum preference for statement i, displayed on the left) to K (maximum preference for statement j, displayed on the right). In the TIRT model, the latent preference underlying each pairwise response is defined as the unobserved difference in utility between the two statements:
where item utilities follow a linear factor model:
with
where
It is important to note that while the limitations of factor analysis on subscale scores are well-documented, particularly in cases of full ipsativity where one score can be entirely derived from the others (e.g., Clemans, 1966; Dunlap & Cornwell, 1994), factor analysis at the block level does not inherently suffer from the same degree of collinearity issues. This distinction arises because responses to pairwise comparison blocks are conditionally independent, allowing for parameter recovery when appropriate conditions are met. Simulation studies have demonstrated that, under thoughtful test design and adherence to stricter assumptions about the questionnaire’s structure, accurate parameter recovery is achievable (Brown & Maydeu-Olivares, 2011; Morillo et al., 2016). However, these approaches often require more rigorous test design considerations, highlighting the critical role of careful test construction in mitigating ipsativity-related challenges (e.g., Schulte et al., 2021).
Challenges With Confirmatory Approaches
FC IRT models are typically applied within a confirmatory framework, assuming that the dimensional structure of the data is known and stable, often transferred from rating-scale applications via parameter invariance (Morillo et al., 2019; Stark et al., 2005). However, this assumption can be compromised by context effects and shifts in perceived item meaning within blocks (Lin & Brown, 2017). Moreover, TIRT models often impose simple structures that overlook cross-loadings, despite evidence that these are common in complex constructs such as personality (Garrido et al., 2020; Hopwood & Donnellan, 2010; Marsh et al., 2010). When theoretical knowledge is limited, as is often the case in early stages of test development, exploratory approaches offer greater flexibility and fewer restrictive assumptions.
FC Test Design and Dimensionality of Multidimensional Blocks
Despite the utility of IRT and FA models in addressing ipsativity, it is well established that the design of FC questionnaires plays a critical role in recovering item and person parameters. For pairwise BFC and GFC comparisons, the challenges arise because FC blocks are typically bidimensional (i.e., each block measures two dimensions). This contrasts with single-stimulus formats such as Likert-type items, where each item is generally constructed to measure a single dominant trait, and any secondary associations are treated as minor or negligible. In FC blocks, however, each response inherently reflects the comparison between two traits, making the block structurally multidimensional by design. This distinction is not merely conceptual; FC blocks, by their very nature, can increase the risk of collinearity in the factor loading matrix if their design lacks variability, as will be illustrated. This structural feature has important implications for the dimensional sensitivity of the test.
For multidimensional IRT models, the number of dimensions to which a test is sensitive can be lower than the number of dimensions it measures, depending on the test design (Morillo, 2018, p. 73; Reckase, 2009, p. 182). For instance, Morillo (2018) shows that when blocks measure dimensions in the same direction (i.e., maintaining a constant ratio between their item discrimination parameters), the test information matrix becomes singular (not full-rank). This rank deficiency is analogous to that observed in bifactor models derived through Schmid-Leiman transformations, where proportionality constraints on factor loadings result in a rank-deficient loading matrix (Waller, 2018). Consequently, in FC questionnaires, blocks with insufficient variability in their loading patterns may fail to adequately distinguish among dimensions, leading to an underestimation of the true dimensionality of the data (an illustrative example of this rank deficiency can be seen in the Online Supplemental Appendix; OSF, osf.io/6xeks/).
Thus, to mitigate these challenges, careful block design is essential. Brown and Maydeu-Olivares (2011) caution against relying solely on homopolar (i.e., equally keyed) blocks, where items measuring their respective traits are paired in the same direction (either positively or negatively keyed). This type of block concentrates information in specific directions of the multidimensional space. Consequently, blocks provide overlapping information, capturing relative differences between traits (e.g., Openness being higher than Extraversion), but failing to provide accurate absolute measures (Brown & Maydeu-Olivares, 2011). To address this limitation, one proposed strategy is the inclusion of heteropolar (or mixed-keyed) blocks, which pair items measuring traits in opposite directions, providing greater differentiation among traits (Brown & Maydeu-Olivares, 2011; Bürkner et al., 2019; Frick et al., 2021; Sun et al., 2023). However, heteropolar blocks are more challenging to design and implement due to the difficulties associated with matching items on social desirability (Bürkner, 2022; see Graña et al., 2024 for an in-depth review of this debate regarding the use of heteropolar blocks). Another effective strategy is the use of unidimensional blocks, where only one trait is measured per block. These blocks can be constructed by: (1) including single-stimulus items alongside blocks (e.g., Xiao et al., 2017), (2) pairing two items measuring the same dimension (e.g., Stark et al., 2005), or (3) pairing an item with a distractor unrelated to the measured traits (e.g., Lee et al., 2022). In this regard, methods for the optimal assembly of FC questionnaires have been proposed, incorporating these design recommendations to maximize reliability (Kreitchmann et al., 2022). For instance, building item blocks that maximize within-block loading differences has been found to enhance reliability (Bürkner, 2022).
These findings underscore that the challenges associated with ipsativity in FC questionnaires are not purely intrinsic to the response format but are significantly shaped by block design and test construction. To overcome these limitations, careful consideration must be given to item pairing, directional balance within blocks, and the overall test structure. By addressing these design-dependent factors, researchers can improve the accuracy and validity of dimensionality assessments in FC questionnaires.
Exploratory Framework and Dimensionality Assessment
In a fully exploratory framework, a critical first step is to determine the number of underlying dimensions in the data. As mentioned earlier, although dimensionality can be explored through different procedures, their performance with FC data remains untested. This is particularly challenging because the previously mentioned effects (e.g., test design and structure complexity) are expected to influence FC data dimensionality assessment. As a result, dimensionality estimation methods may behave differently when applied to this format. Moreover, these effects are likely to operate along a continuum rather than as categorical shifts, making it especially relevant to examine how robust different dimensionality estimation methods are under varying levels of structural complexity. Establishing whether dimensionality assessment methods can be successfully applied to FC data is a necessary step toward building an exploratory framework and could also provide evidence of the internal structure validity for this type of data. Specifically, it would enable researchers to verify whether theoretical or designed dimensions align with empirical ones.
Several previous studies with single-stimulus items have explored the effects of structural complexity on dimensionality assessment, particularly through the inclusion of cross-loadings, they do not fully capture the unique characteristics of FC questionnaires. For example, Auerswald and Moshagen (2019) investigated the effect of adding a cross-loading of ±0.20 on a second factor for some items in the test. They found that dimensionality detection improved under some conditions, possibly due to increased factor determinacy. However, Li et al. (2020), in a more comprehensive study, examined the impact of cross-loadings under a broader range of conditions by varying both the proportion of affected items (from 10% to 50%) and the magnitude of cross-loadings (from 0.10 to 0.30). Their findings suggest that cross-loadings tend to impair dimensionality detection, particularly when they are frequent or relatively strong. While these scenarios share some similarities with FC formats, the cross-loadings simulated in these studies were still smaller than those typically found in FC blocks. Brandenburg and Papenberg (2024) further increased structural complexity by generating indicators with substantial loadings in two factors. However, even in this case, the cross-loadings were uniformly positive and not structured to reflect the defining features of FC designs. Thus, these studies still fall short of capturing the specific configuration of FC blocks, namely, the systematic pairing of items with opposing loadings, which has direct implications for both their psychometric properties and their dimensional complexity. Moreover, none of these studies examine how FC test design features might influence the performance of dimensionality estimation methods in recovering the intended factor structure.
The current study evaluates five dimensionality assessment methods: Parallel Analysis (PA; Horn, 1965), Empirical Kaiser Criterion (EKC; Braeken & van Assen, 2017), the Hull method (Lorenzo-Seva et al., 2011), Exploratory Graph Analysis (EGA; Golino & Epskamp, 2017), and Kaiser’s eigenvalue-greater-than-one criterion (K1; Kaiser, 1960). PA, EKC, Hull, and EGA were prioritized, among other existing methods, because of their consistently strong performance across multiple simulation studies involving both continuous and categorical variables, as well as across a broad range of factorial conditions (Auerswald & Moshagen, 2019; Cosemans et al., 2022; Golino et al., 2020; Jiménez et al., 2023; Li et al., 2020). PA, one of the most comprehensively studied methods and widely regarded as a gold standard (Goretzko, 2025), determines dimensionality by comparing observed eigenvalues to those derived from simulated datasets based on random, uncorrelated variables. The EKC enhances the traditional K1 criterion by adjusting for eigenvalue bias related to sampling variability and sequential dependency, demonstrating superior accuracy, especially in situations involving correlated factors and shorter scales (Braeken & van Assen, 2017; Li et al., 2020). The Hull method integrates model fit and parsimony through graphical evaluation, offering a strong alternative to purely eigenvalue-based criteria (Auerswald & Moshagen, 2019). EGA is a highly accurate, network-based method that identifies latent dimensions by detecting community structures within regularized partial correlation networks (Cosemans et al., 2022; Golino et al., 2020). The K1 criterion was included due to its continued widespread use, despite consistent evidence of its tendency to overestimate dimensionality, particularly in practical applications with finite samples (Auerswald & Moshagen, 2019; Goretzko & Bühner, 2020). During the implementation of the EKC, we identified a conceptually distinct variant based on the implementation of Auerswald and Moshagen (2019) used in several R packages, which compares each eigenvalue to a recursively defined threshold based solely on the null model. Given its practical relevance and divergence from the original EKC, we included this version in our analyses under the label Maximal Kaiser Criterion (MKC). A more comprehensive description of the six dimensionality assessment methods explored in the present study is available in the Online Supplemental Appendix.
Several alternative dimensionality methods were initially considered but ultimately excluded due to various limitations. The Minimum Average Partial method (Velicer, 1976) has consistently demonstrated lower accuracy compared to the leading methods (Cosemans et al., 2022; Finch, 2025). Methods such as Comparison Data (Ruscio & Roche, 2012) and Revised Parallel Analysis (Green et al., 2011) showed initial promise but have exhibited marked deterioration in performance with more realistic factor structures (Auerswald & Moshagen, 2019; Xia, 2021). The Next Eigenvalue Sufficiency Test (Achim, 2017) is another eigenvalue-based method that has achieved good performance across various studies; however, its main advantages over PA appear to be in conditions with very small samples (N < 300; Caron, 2025a), which are not appropriate for FC models. The Out-of-Sample Prediction Error is a very recent promising alternative that needs further scrutiny with structures that more closely resemble real data (Haslbeck & van Bork, 2024). Finally, advanced machine learning methods such as Factor Forest (Goretzko & Bühner, 2020) and Comparison Data Forest (Goretzko & Ruscio, 2024) were developed specifically for single-stimulus factor structures and might require retraining to effectively handle data characteristic of FC assessments.
Expected Performance Under Forced-Choice Conditions
Although these methods have not been systematically evaluated in the context of FC data, the most comparable scenario studied in the literature involves the presence of cross-loadings, particularly when items load substantially on more than one dimension. Cross-loadings have been found to result in a concentration of variance in the leading eigenvalues, thereby diminishing the relative size of subsequent ones (Li et al., 2020; Zopluoglu & Davenport Jr., 2017). Consequently, PA may show lower accuracy and underestimate factors when cross-loadings are strong or widespread, especially in structures where items load similarly on multiple dimensions (Brandenburg & Papenberg, 2024; Zopluoglu & Davenport Jr., 2017). By contrast, Li et al. (2020) found that the EKC outperformed PA in the presence of cross-loadings. This advantage is likely due to EKC’s adjustment for the variables-to-sample-size ratio and its sequential comparison of eigenvalues, which enhances its robustness against the compression of the eigenvalue structure typically induced by cross-loadings. The Hull method, which relies on model fit, is anticipated to be less directly affected by cross-loadings. However, relevant prior analyses involved only relatively weak cross-loadings (Auerswald & Moshagen, 2019), limiting the generalizability of these findings to complex FC structures. Lastly, EGA assumes that each variable loads on a single dimension, a premise that conflicts with the inherent multidimensionality of FC blocks. Simulation evidence (Brandenburg & Papenberg, 2024) indicates that EGA performs poorly under conditions of substantial cross-loadings, supporting the notion that it may be less suitable for FC applications.
Goals of the Present Study
The goals of the present study are twofold: (1) to assess the performance of dimensionality detection methods in the context of FC data, and (2) to identify the design factors that may help recover dimensionality. This study is conducted under the hypothesis that EKC and PA should exhibit adequate performance, followed by the Hull method, while EGA is expected to perform worse than the first three, and K1 should be the worst-performing method. The MKC, a data-independent variant of EKC, is expected to perform worse than PA due to its increased conservativeness. In addition, some design factors, such as the inclusion of heteropolar or unidimensional blocks, are hypothesized to improve the performance of dimensionality assessment methods. Addressing these two objectives of this study should facilitate the transition to an exploratory modeling framework for FC data.
Method
Design
The independent variables manipulated in this simulation study are presented in Table 1. We manipulated the number of dimensions (D), the number of items per dimension (J), the number of response categories (K), the loadings mean (LM), the average difference in within-block loadings (LD), the correlations among dimensions (RDim), the percentage of heteropolar blocks (H), the inclusion of unidimensional blocks (U), and the sample size (N). Notice that, in FC, the observed variables are blocks (i.e., comparisons between items). Each block consists of two items, each measuring a different dimension. Therefore, a block is associated with the pair of dimensions measured by its two items. If each dimension is evenly paired with the others across the test, the number of blocks involving each dimension is J/2. Depending on the number of dimensions (D) and items per dimension (J), the total number of blocks varied across conditions (D × J/2; e.g., from 18 to 60).
Independent Variables According to the Research Design.
The levels for these independent variables were selected based on several considerations. The inclusion of heteropolar and/or unidimensional blocks represents key test design decisions expected to reduce ipsativity problems in obtained scores (Lee et al., 2022) and, consequently, enhance dimensionality detection. Similarly, it is expected that the number of response categories influences the reliability of the FC measures (Brown & Maydeu-Olivares, 2018; Zhang et al., 2023), making it relevant to explore its potential effect on dimensionality recovery. Following Zhang et al. (2023), three scenarios were considered: 2, 4, and 5 response categories. The levels for the mean of standardized loadings and the average difference of within-block loadings were based on the empirical study of Graña et al. (2024). Note that, from the TIRT model perspective, standardized loadings for blocks are expected to be smaller than those for items. This is because the latent variable
The number of dimensions, the number of items per dimension, and types of blocks (heteropolar and unidimensional) were selected to ensure adequate and realistic questionnaire compositions when combined. As an example, Table 2 shows simulated loadings matrices for a questionnaire with 3 dimensions (D = 3), 12 items per dimension (J = 12), loadings mean of 0.4 (LM = 0.4), and average difference in within-block loadings of 0.2 (LD = 0.2). To begin, we focus on the base condition, in which all the blocks are homopolar. According to the Thurstonian IRT model, homopolar blocks consist of items whose loadings have opposite signs (one positive and one negative or vice versa). This is because the model conceptualizes the latent response as the difference between the latent traits associated with each item in the pair. This formulation aligns with the nature of FC designs: when both items measure dimensions in the same direction, selecting one item in the block reflects an increase in one trait while simultaneously decreasing the other.
Examples of Loading Matrices for Simulated Questionnaires.
Note. Blocks in bold indicate changes from the base condition to 33% heteropolar; italicized blocks represent changes from base condition to one unidimensional block per dimension. Changed values are underlined.
In a condition with heteropolar blocks, one-third of the blocks in these matrices (blocks in bold, e.g., blocks 1 and 2) would be switched to heteropolar; this means the sign of the first or second loading (selected at random) would be changed to generate the heteropolar blocks (both loadings having the same sign). Unlike homopolar blocks, heteropolar blocks represent a scenario where items measure traits in opposite directions. As a result, endorsing one item in the block reflects an increase in both traits (whereas endorsing the other implies a decrease).
In a condition with unidimensional blocks, one loading of the underlined blocks (e.g., block 3) would be changed to 0, ensuring one unidimensional block per dimension. Notice how many other questionnaire characteristics (e.g., blocks per dimension) can be calculated. For example, the number of possible pairs of dimensions is 3 (= D × [D − 1]/2) and the total number of blocks is 18 (= D × J/2), which translates to 6 blocks per pair of dimensions. In a heteropolar blocks condition, there would be 6 heteropolar blocks (one-third of the total, 2 per pair of dimensions), and if unidimensional blocks were included, there would be 3 of them (one per dimension). As another example, in a condition with 5 dimensions and 24 items per dimension, there would be 10 pairs of dimensions, a total of 60 blocks (6 per pair of dimensions), 20 heteropolar blocks if included (2 per pair of dimensions), and 5 unidimensional blocks if included.
Data Generation
For each replication, data were generated from the Thurstonian IRT model following the next steps. First, the population factor correlation matrix (
Simulating the block standardized loadings matrix directly, rather than generating an item loadings matrix followed by block assembly and loadings standardization, was preferred as it aligns more closely with an exploratory framework where no prior information about item loadings is assumed. However, to ensure that the results of this study are generalizable to such scenarios, the simulated FC loadings were transformed into standardized item loadings (i.e., the values items would have if presented in a single-stimulus format, such as Likert) to assess their plausibility. Ignoring the sign, the means of item-level loadings were 0.559 and 0.693, for LM = 0.4 and LM = 0.5 conditions, respectively. The average difference in within-block loadings was 0.189 and 0.259, for LD = 0.2 and LD = 0.3 conditions, respectively.
After the simulation of the loadings matrix, the error correlation matrix (
Dimensionality Assessment Methods
We applied six different dimensionality assessment methods (K1, EKC, MKC, PA, EGA, and Hull). With the exception of the Hull method, which requires Pearson correlations, all the other methods were performed on polychoric correlation matrices, except for the K = 2 case, where tetrachoric correlation matrices were used. These matrices were estimated using the Turbofuns R package, Version 1.0.0 (Zhang et al., 2021). K1 was calculated as the number of eigenvalues from the correlation matrix that exceeded 1. EKC was calculated as the number of eigenvalues greater than their corresponding reference eigenvalues following the original formula proposed by Braeken and van Assen (2017). MKC was calculated as the number of eigenvalues greater than their corresponding reference eigenvalues using the function EMPKC in the package EFA.dimensions, Version 0.1.8.4 (O’Connor, 2024). For the PA method, 1,000 random datasets were generated by independently permuting the values within each variable (column-wise sampling) of the empirical dataset. This procedure preserves the original sample size, number of variables, and marginal distributions, while disrupting the relationships between variables, effectively removing the influence of any underlying latent structure. The reference eigenvalues, representing the null model of no common factors, were defined as the 95th percentile of the eigenvalue distribution obtained from these permuted datasets. For EGA, the GLASSO regularization method (Friedman et al., 2008) and the Louvain algorithm for clustering were applied to the correlation matrix, using the EGAnet package, Version 2.0.8 (Golino & Christensen, 2024). The Hull method was applied to the Pearson correlation matrix using the maximum likelihood extraction method and the CFI fit index as criteria, as implemented in the EFAtools package version 0.4.4 (Steiner & Grieder, 2020). The lower and upper bounds for the number of factors were set to zero and the number of factors suggested by PA plus one, respectively.
Assessment Criteria
To assess the performance of the different dimensionality assessment methods, three dependent variables were computed: the hit rate (HR), the bias, and the mean absolute error (MAE). HR, representing the proportion of correct dimensionality assessments, was computed as:
where I is the indicator function,
A bias of 0 indicates no systematic error, while positive or negative values indicate over- or underestimation of the number of dimensions, respectively. However, a bias close to zero may mask instances where positive and negative errors cancel each other out. To address this limitation, MAE was calculated as the average of absolute differences between the estimated and true number of dimensions:
A MAE of zero reflects perfect accuracy, while higher values indicate greater inaccuracy.
Finally, univariate analyses of variance (ANOVAs) using sum of squares type III were conducted in R using the package afex, Version 1.4-1 (Singmann et al., 2024) to assess the effect of the factors and their interactions on the performance of each method. A total of three ANOVAs were performed (one for each dependent variable), including as independent variables the simulated factors of this study and the dimensionality estimation methods (as a repeated-measures factor, excluding K1 and ECK due to their poor performance overall), and allowing for all possible interactions. Partial eta-squared (
Results
Table 3 shows the overall marginal average results for all dimensionality assessment methods across all conditions and dependent variables. MKC and PA exhibited the best global performance, achieving an HR of 0.8 and an MAE of 0.2, and the Hull method ranked third (e.g., HRHull = 0.75). This pattern held across nearly all levels of the independent variables, with MKC and PA standing out as the top-performing methods. Although both showed slight underestimation (BiasMKC = −0.16; BiasPA = −0.20), they also demonstrated an acceptable bias overall. The only scenario where MKC and PA did not outperform the other methods was in the absence of heteropolar blocks, where EGA showed the best performance in HR, Bias, and MAE (e.g., for H = 0, HRMKC = 0.62; HRPA = 0.60; HREGA = 0.72). Nonetheless, EGA’s overall marginal performance remained poor (e.g., HREGA = 0.62), despite showing a low bias (BiasEGA = 0.13). This value likely reflects a cancellation of over- and underestimation errors rather than a true indicator of accuracy, as indicated by its high MAE (MAEEGA = 0.48). Finally, K1 and EKC yielded the poorest performance (HRK1 = 0.50; HREKC = 0.51) and showed considerable overestimation in the number of factors (BiasK1 = 1.78; BiasEKC = 1.52). Consequently, K1 and EKC will not be considered in subsequent inferential analyses.
Marginal Results.
Note. Numbers in bold indicate the best method in each condition for each dependent variable. D = Number of dimensions; J = Number of items per dimension; K = Number of response categories; LM = Loadings mean; LD = Average difference in within-block loadings; RDim = Correlations among dimensions; H = Heteropolar blocks; U = Inclusion of unidimensional blocks; N = Sample size.
Table 4 shows the main effect sizes from the ANOVAs performed (a representation of the performance of each method considering all factors can be seen in Figures SS1–SS3 in the Online Supplemental Appendix). A relevant effect of the dimensionality assessment method used was found on all dependent variables (
ANOVAs Relevant Effect Size Values.
Note. For interaction effects, only those where
According to the interactions shown in Figure 2, the inclusion of heteropolar blocks consistently resulted in unbiased recovery of dimensionality for MKC, PA, and Hull, irrespective of other design factors. In contrast, the absence of heteropolar blocks led to a decrease in HR, which was aggravated under specific conditions. This decrease was more pronounced when: correlations among dimensions were smaller (RDim × H:

HR for the interactions involving the inclusion of heteropolar blocks.
The EGA method followed a different pattern, as evidenced by two triple interactions. It showed that the effect of the inclusion of heteropolar blocks or the increase in the number of items per dimension depended on the number of dimensions (for HR; Method × D × H:

MAE for the interaction between the number of categories and sample size.
Finally, notice how, in Figures 2 and 3, the overall ordering of the methods held as seen in Table 3, being MKC and PA the best performing with similar results, followed by the Hull method, and then EGA, which obtained better results than the other methods only in very specific conditions related to the absence of heteropolar blocks.
Discussion
Despite the growing interest in FC questionnaires, the systematic evaluation of traditional dimensionality assessment methods in this context remains limited. Confirmatory approaches, such as CFA or IRT models, are commonly employed to validate internal structures but often yield poor fit indices, highlighting their shortcomings. This underscores the need to study dimensionality recovery from an exploratory perspective, particularly addressing the unique challenges of FC data, including their complex factorial structures and the risk of dimensionality loss, which can limit the effectiveness of conventional methods. This simulation study examines how factors related to FC test design affect the performance of commonly used dimensionality assessment methods. The objective is to identify the conditions under which these methods can accurately assess dimensionality in FC questionnaires. By doing so, this work not only advances dimensionality assessment as a source of validity evidence for internal structure but also lays the groundwork for the application of exploratory factor models.
Results show that MKC and PA are generally effective for assessing dimensionality in FC data, achieving a mean HR of 0.80 along with low absolute error. A systematic pattern emerged among the three best-performing methods in this study (MKC, PA, and the Hull method) regarding bias, showing underestimation of the number of dimensions when the methods failed. This underestimation likely reflects a reduced sensitivity to dimensionality when the ratio of loadings is nearly proportional across blocks (Morillo, 2018), as evidenced by the better performance when heteropolar and/or unidimensional blocks are included. For instance, MKC and PA achieved near-perfect HR values (approaching 1) when heteropolar blocks were included, and HR close to 0.86 in the case of unidimensional blocks.
Marginal results for the simulated factors indicate that when no heteropolar blocks were included in the questionnaire, EGA has better results than MKC or PA. This difference was minimal when the number of items per dimension or the within-block loading differences increased. Also, when dimensions were correlated (i.e., RDim = 0.3), all methods performed equally well, indicating that factor correlations positively impact dimensionality recovery for FC data.
Although the main effect sizes associated with sample size only reached the medium magnitude (
Surprisingly, the performance of the original EKC method was unexpectedly poor, yielding results that closely resembled those of the classic Kaiser criterion (e.g., HR values were nearly perfectly correlated, r = .99), and systematically underperformed compared to PA. The pattern of eigenvalues revealed that the reference threshold generated by EKC for the K + 1 component, where K denotes the true number of factors, was often nearly identical to the corresponding empirical eigenvalue, resulting in an overextraction decision. These findings indicate that the favorable performance of EKC in simple factorial structures does not generalize to FC data. Another possible explanation lies in the assumptions underlying the derivation of EKC. EKC is based on the distribution of observed Pearson correlations, which are appropriate for continuous variables but are known to misrepresent the associations among ordinal indicators. In our analysis, as in previous studies (Goretzko & Bühner, 2020), we rely instead on polychoric correlations, which better capture the relationships among the underlying continuous latent variables. However, because eigenvalues derived from polychoric correlation matrices are generally larger than those obtained from Pearson correlations, this leads to reference values that are too low. This interpretation aligns with our finding that EKC performance improves as the number of response categories increases, narrowing the gap between polychoric and Pearson correlations. The issue appears to affect EKC more than MKC, since although both methods are expected to yield a downward-biased first reference value, EKC’s bias accumulates across subsequent comparisons. This is because EKC sequentially subtracts the influence of preceding observed eigenvalues, derived from polychoric correlations, which are systematically larger than their Pearson-based counterparts. As a result, EKC’s second and subsequent reference values become increasingly downward biased, further amplifying the tendency to overextract factors. However, it is important to note that the core limitation does not stem from the choice of correlation type alone. Rather, it reflects a broader mismatch between EKC’s derivation, based on continuous variables, and the categorical nature of the data, which could yield distorted reference values even if Pearson correlations were used.
In summary, to optimize dimensionality detection in FC questionnaires, MKC and PA are generally recommended due to their robust performance across various conditions. Incorporating heteropolar blocks into the test design proves particularly effective in improving the accuracy of these methods. Heteropolar blocks, which pair items measuring traits in opposite directions, provide greater differentiation among traits but are challenging to design and implement, requiring careful attention to social desirability matching (Bürkner, 2022; Graña et al., 2024). Unidimensional blocks, on the other hand, demonstrated a less relevant effect, but may be easier to include and can take the form of single-statement items or blocks where one item does not measure any of the questionnaire’s dimensions of interest. Both strategies have shown significant improvements in dimensionality recovery, particularly under challenging conditions. Although with a smaller effect, designing questionnaires with 4 or 5 response options, rather than only 2, should also improve dimensionality recovery.
In the absence of heteropolar blocks, increasing the number of items per dimension can also enhance the performance of MKC and PA. In addition, leveraging available information to optimize block assembly—for example, by maximizing within-block loading differences—can further improve dimensionality detection. Optimization algorithms, such as those proposed by Kreitchmann et al. (2022, 2023), can help ensure that block designs maximize the reliability and validity of trait recovery. Finally, given its systematic positive impact across methods, increasing sample size remains one of the most accessible and effective strategies for improving dimensionality assessment.
Interestingly, when only homopolar blocks are used, the performance of MKC and PA, as well as the Hull method and EGA, improves in the presence of factor correlations. This finding contradicts prior research suggesting that methods typically perform worse with correlated factors, as orthogonal factors are more easily distinguishable (e.g., Garrido et al., 2013). However, the unique structural properties of FC data likely explain this result. In the absence of heteropolar blocks, factor correlations influence the eigenvalue structure in a way that facilitates dimensionality detection. Specifically, as factor correlations increase, the average block-dimension correlation (i.e., LM computed from the loadings in the structure matrix) decreases, but the differentiation of within-block block-dimension correlation (i.e., LD computed from the loadings in the structure matrix) increases, which can enhance the separation between dimensions. This pattern differs from what is typically observed in Likert-type data, where higher factor correlations tend to blur dimensional distinctions. A more detailed mathematical explanation of this phenomenon, along with illustrative examples, is provided in the Online Supplemental Appendix.
Limitations and Future Research
Several limitations of the present study warrant discussion. First, this research focused exclusively on pairwise FC comparisons. In this context, the TIRT model is equivalent to traditional confirmatory factor analysis (Brown & Maydeu-Olivares, 2011). Extending the dimensionality estimation methods to FC blocks with more than two items, such as triplets or quadruplets, would necessitate decomposing responses into pairwise binary comparisons. However, this approach introduces the challenge of accounting for correlated errors among comparisons within the same block, which cannot be ignored as they may significantly impact the accuracy of dimensionality recovery.
Second, the generalizability of these findings to levels of the factors not included in the study should be approached with caution. While the number of dimensions did not show a significant impact on dimensionality recovery in this study, higher-dimensional structures may present different challenges. These are of particular interest in FC research because such structures appear to mitigate ipsativity (Bürkner, 2022). In addition, exploring the effects of varying proportions of unidimensional and heteropolar blocks could offer valuable insights. For example, a greater proportion of unidimensional blocks is expected to improve performance, while reducing heteropolar blocks may have the opposite effect.
Third, all strategies proposed in this study rely on some prior knowledge about the blocks. For instance, constructing heteropolar and unidimensional blocks requires assumptions about the dimensions being measured and the polarity of the items. Similarly, optimizing block assembly assumes prior knowledge about the magnitude of item loadings. In practice, however, test construction typically involves access to such information. For example, FC questionnaires are often developed based on previous Likert-scale data, providing insights into item properties. Although the effect of inaccuracies in this prior information was not tested in this study, and warrants further investigation.
Fourth, this research focused exclusively on dimensionality recovery, which represents the first step in understanding the structure of FC data. Recovering the detailed factor structure poses additional challenges. For example, traditional EGA is a preferred method when heteropolar blocks are absent, but it assumes that each indicator belongs to only one unique cluster. This limitation still requires further exploration. For factor analysis, the effectiveness of different rotation methods should also be tested to address the specific challenges posed by FC structures.
As a final note, during the course of this work, we observed that in the case of the EKC proposed by Braeken and van Assen (2017) and implemented in the R package semTools (v.0.5-6; Jorgensen et al., 2025), Auerswald and Moshagen (2019) actually used a modified version of the method. This modification, which was not explicitly indicated as such, has since become widespread and is now implemented in other R packages such as EFAtools (v.0.5.0; Steiner and Grieder, 2020), EFA.dimensions (v.0.1.8.4; O’Connor, 2024), and Rnest (v.1.1., Caron, 2025b). For clarity, we have referred to this modified approach as the Maximal Empirical Kaiser method throughout this study. It is important to note that these observations reflect the state of available implementations at the time of writing. While this study examines the performance of both approaches under various conditions, further research is warranted to better understand their differences and implications.
Conclusion
When assessing dimensionality in FC data, MKC and PA emerge as the most accurate methods among those evaluated in this study. To optimize their performance, careful consideration of FC test design is essential. Including heteropolar blocks proves to be the most effective strategy, although it requires the challenging task of social desirability matching (Graña et al., 2024). When heteropolar blocks are not feasible, unidimensional blocks or optimized test assembly, such as maximizing within-block loading differences based on prior knowledge, can significantly enhance dimensionality recovery (Bürkner, 2022; Kreitchmann et al., 2022, 2023). These approaches, however, assume some basic prior understanding of the block structure. In addition, using a graded format (either 4 or 5 categories) is recommended over the traditional binary format. Last, increasing the length of the questionnaire provides a straightforward method to improve dimensionality detection.
Ultimately, thoughtful test design is paramount for successful dimensionality recovery. By carefully tailoring block composition, researchers can accurately determine the internal structure of a questionnaire and validate whether the block assembly process has preserved the intended structure of an item bank.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was funded by MICIU/AEI/10.13039/501100011033, ERDF/EU and the ESF+ under the project “Computerized adaptive tests based on new assessment formats” (reference: PID2022-137258NB-I00 and PREP2022-001047), as well as by the UAM IIC Chair on Psychometric Models and Applications.
Ethical Approval and Informed Consent Statements
Not applicable.
ORCID iDs
Data Availability Statement
Supplemental Material
Supplemental material for this article is available online.
