
Other
Select search scope: search across all journals or within the current journal
1-20 of 23 articles

The growing demand for personalized online learning underscores the necessity for diagnostic assessments that are tailored to the cognitive abilities of individual examinees. The combination of cognitive diagnostic models (CDMs) with computerized multistage testing (MST) holds potential for meeting these educational needs. However, research on the integration of MST with cognitive diagnosis (CD-MST) has been limited, largely due to the challenges in establishing criteria for constructing test modules and defining routing rules prior to administering the test. This study aims to introduce an innovative design approach for CD-MST that employs a strategy of partitioning the skill-space, which encompasses all possible attribute mastery profiles, to address the challenges. By partitioning the skill-space into groups of attribute profiles, distinct modules tailored to each partitioned group can be constructed, ensuring that each examinee is adaptively routed to the most suitable module at each testing stage. Item information functions for CD-MST are also proposed by defining the information conditionally on an attribute profile, in order to quantify an item’s discrimination power for each profile. Furthermore, a strategic approach for automated module assembly in CD-MST is developed to construct modules that maximize information for each attribute profile group while satisfying all practical constraints. Simulation results indicate that the proposed CD-MST improves estimation accuracy compared to traditional linear test and can effectively utilize a wider range of item types from the item bank.
In Automatic Item Generation (AIG),
The study compared the effectiveness of four methods for detecting differential item functioning (DIF) in polytomous multidimensional data with a simple structure: the item response theory likelihood ratio test (IRT-LR), two ordinal logistic regression approaches (using raw scores vs. latent trait estimates as the matching variable), and the multidimensional MIMIC-interaction method. Data were generated under a two-dimensional graded response model with 28 five-category items. Simulation conditions manipulated DIF type (uniform, nonuniform), DIF magnitude (0, 0.3, 0.6), group size ratio (1:1, 3:1), latent trait correlation (ρ = 0, 0.5), and the presence of group impact, yielding 40 conditions with 100 replications each. Across conditions, IRT-LR and both logistic regression approaches generally maintained Type I error within acceptable limits, whereas the MIMIC-interaction model showed inflated Type I error in the presence of impact. All methods demonstrated high power for moderate uniform DIF, but detection rates declined substantially for low DIF and for nonuniform DIF. Logistic regression with latent trait estimates showed the most stable overall performance, combining adequate Type I error control with comparatively high power across conditions. Logistic regression with raw scores demonstrated relatively stronger performance for moderate nonuniform DIF. In contrast, IRT-LR exhibited lower power despite conservative Type I error control. Results suggest that regression-based approaches, particularly logistic regression using latent trait estimates, provide robust performance for DIF detection in multidimensional polytomous assessments under simple structure.
Linear regression and analysis of variance are widely used in applied psychological measurement to estimate group, condition, and covariate effects, yet statistical efficiency and conventional inference can be compromised when outcome variance changes across groups or covariate levels. This article introduces varGuid for R and varguid for Python, open-source implementations of variance-guided regression for linear models. The method estimates a covariate-dependent mean-variance relationship and uses it to iteratively reweight the original mean model. Ordinary analyses use iteratively reweighted least squares, whereas sparse analyses use an iteratively reweighted lasso. Because the original design matrix and outcome scale are retained, regression coefficients and ANOVA contrasts remain directly comparable with conventional effect estimates. Robustness here refers to relaxation of the homoscedasticity assumption rather than resistance to outliers. Under the conditions established for variance-guided regression, the estimator matches the homoscedastic baseline in population predictive quasi-risk when variance is constant and improves on that baseline when variance depends on covariates. The packages accept general linear-model design matrices, including ANOVA-style encodings, and provide baseline and variance-guided predictions, example data, and heteroscedasticity-consistent summaries for non-lasso fits. The Python implementation also supports NumPy and pandas inputs, Patsy formulas, model summaries, and a scikit-learn-compatible estimator. Both implementations are operating-system independent and require no unusual hardware. Source code, documentation, examples, and installable files are available through CRAN, PyPI, GitHub, and Zenodo. The packages provide accessible tools for applying variance-guided regression as a primary or companion analysis when homogeneity of variance is uncertain in routine measurement research and related quantitative applications.
scindex is an R package for analysing inter-rater reliability in binary classification tasks. The package computes Cohen’s κ and Fleiss’ κ and, when ground-truth labels are available, estimates signal detection theory parameters, including sensitivity, specificity, and decision thresholds. It also implements the Strategic Convergence Index (SCI), a measure of convergence in raters’ response criteria.
This study aims to examine the performance of adaptive quadrature (AQ) estimation method for ordinal confirmatory factor analysis (CFA). Specifically, we compared four link functions (complimentary log-log [CLL], logit, log-log, and probit) of the AQ estimation method across varying factor structures, sample sizes, distributional shape of latent trait, and number of quadrature points. The study is conducted via a simulation study and using empirical data. The results demonstrate that the probit link function exhibits superiority across the vast majority of conditions, consistently yielding the highest proper convergence rates and, among successfully converged solutions, the lowest parameter recovery errors, and the best relative fit, whereas the logit generally showed the weakest performance. Additionally, a critical divergence was discovered regarding asymmetric link functions: while the probit link generally provided the best model fit across the positively skewed simulation conditions, the log-log link yielded the best relative model fit for the positively skewed empirical data. Furthermore, the study reveals the complex role of quadrature points in multidimensional spaces. Although using eight quadrature points may be necessary in more complex simulated models, it frequently causes severe estimation failures when applied to sparse real-world data.
Multidimensional forced choice (MFC) test formats are commonly used as an alternative to traditional rating scale formats to reduce aberrant responding, especially faking in high-stakes settings. However, MFC remains susceptible to random responding, particularly in low-stakes settings where respondents may be insufficiently motivated and in high-stakes settings where some assessments may be viewed as less consequential. To ensure the validity of inferences drawn from MFC data, effective methods for detecting random responding are needed. This research contributes to the MFC literature on aberrant responding detection by evaluating the effectiveness of the item response theory (IRT)-based person fit statistic
Choosing suitable estimation methods for cognitive diagnostic models (CDMs) is critical. However, practitioners often face issues like non-convergence, boundary estimates, extreme values, and unstable suboptimal solutions, which affect the accuracy and reliability of parameter estimates. In this study, we compared expectation–maximization (EM), Bayesian modal estimation (BM), their monotonic constraint variants (EMM and BMM), and variational Bayes (VB) methods. A simulation study was conducted, manipulating factors such as sample size (50, 200, 1000), test length (15, 30), item quality (high, low), and attribute distribution (uniform and multivariate normal). The performance was assessed based on the empirical frequency of each issue, the recovery accuracy of the parameters, and the sensitivity to algorithm initialization. The results, analyzed using the generalized deterministic inputs noisy “and” gate model, reveal three main findings. First, an insufficient sample size was identified as a key factor in problems related to parameter estimation. Second, methods that incorporate prior information (BM and VB) exhibited fewer cases of non-convergence and extreme estimates than EM. Third, the sensitivity analysis showed that the stability of solutions was affected by the choice of initial values, emphasizing the need for proper initialization to reduce the risk of becoming trapped in local suboptimal solutions. This systematic comparison demonstrates that no single estimator is universally superior, and the choice depends on practical constraints. Our findings offer evidence-based guidance for selecting context-sensitive methods, thereby improving the validity of CDMs in real-world applications.

Self-report questionnaires are widely used in research and practice. In most applications, the vulnerability of these questionnaires to response biases like faking is ignored. However, especially in high-stakes situations such as personnel selection, measurement can be severely biased when test-takers engage in faking to present themselves more favorably. To separate faking-related variance from substantive trait variance, the Multidimensional Nominal Response Model (MNRM) has been used to reduce systematic bias in trait estimation by allowing for item-specific relations between response categories and social desirability. A critical but untested assumption of this approach is that perceptions of social desirability are homogeneous across test-takers. However, individuals may differ considerably in how they perceive the desirability of the item content. Here, we conducted simulation studies to investigate how violations of this assumption affect the MNRM’s ability to recover substantive trait person parameters. We implemented three distinct manipulations of heterogeneous desirability perceptions and examined their impact on person parameter recovery. Results showed that the MNRM is robust against violations of homogeneous social desirability perceptions as long as test-takers’ faking behavior is aligned with their perceived desirability of the item content. In contrast, when test-takers fake responses in ways that are inconsistent with item-wise desirability perceptions, parameter recovery seems to decline. Implications for practice and possible model extensions are discussed.
Item response theory (IRT) observed and true score equating are often conducted assuming that the latent variable is normally distributed. Although this might be a reasonable assumption for many educational and psychological assessments, not all variables can be approximated by a normal distribution. Under the common-item nonequivalent groups design, the current study examined the impact of latent density misspecification on IRT observed and true score equating. Specifically, equating results provided by two separate calibration estimates based on the Stocking–Lord linking method with normal and uniform weights and three concurrent calibration estimates obtained with different characterizations of the latent densities for the old and new groups were compared using both simulated and real data sets. In general, the concurrent calibration method with the latent densities for the two groups estimated using the empirical histogram method provided equating results with the least amount of error for most of the study conditions. Using normal weights with the Stocking–Lord method generally performed much better than using uniform weights; however, the overall performance of the Stocking–Lord method with normal weights was acceptable only if the latent densities for the two groups were normal distributions or close to normal distributions.
When only summary statistics from published studies are available, the Hunter–Schmidt interval is the standard tool for inference on Spearman’s disattenuated correlation, but it treats reliability estimates as known constants and ignores their sampling variability. We derive a simple delta method variance that accounts for the uncertainty of all estimates while requiring nothing beyond the summaries already at hand. Under bivariate normality of scores and coefficient alpha from a normal parallel model, the corrected interval is asymptotically valid. In simulations it achieves coverage near nominal, while Hunter–Schmidt can undercover substantially when reliability is imprecisely estimated.
Latent structure analysis methods, including latent profile analysis (LPA), latent class analysis (LCA), item response theory (IRT), exploratory factor analysis (EFA), and confirmatory factor analysis (CFA), are widely used in psychological and educational research to model unobserved constructs and identify heterogeneity across individuals. However, applying these methods often requires advanced statistical expertise and the use of multiple specialized software packages with different workflows, which can limit accessibility and increase analytical complexity. This paper introduces
Researchers understand that conducting numerous pairwise comparisons between group means increases the Type I error rate, prompting the use of planned contrasts like orthogonal contrast sets. Implicit to orthogonal contrast sets is the principal assumption that groups are balanced in size. Further, when dealing with complex variables like latent constructs, specialized modeling is necessary. Understanding how violating the assumptions of orthogonal contrasts, specifically under conditions of sample imbalance, can help identify variability in parameter recovery. This study examines the effect of sample size imbalance and modeling approach on the accuracy of latent group mean difference estimates when using orthogonal contrasts. Monte Carlo simulations compared the Multiple Indicators Multiple Causes (MIMIC) and re-parameterized multigroup confirmatory factor analysis models while manipulating sample sizes, group proportions, and effect size. Results suggest declining parameter recovery as group imbalance increased, particularly in small samples, with some estimates falling below acceptable thresholds for power, Type I error, and bias. The MIMIC model consistently produced more accurate estimates, though is replete with implicit measurement assumptions that are seldom tested. These findings suggest that researchers using orthogonal contrasts when comparing groups on a latent variable continuum must (a) be aware of examined group’s sample size proportions and the impact of group size inequalities on estimate accuracy, and (b) carefully consider the costs and benefits of the latent variable modeling approach, including how the model addresses measurement non-invariance.
Survey questionnaires are essential tools in psychological and educational research, as the data they gather directly influence research conclusions and policy decisions. A major challenge in ensuring data quality is identifying aberrant response patterns that can jeopardize research outcomes, as they may introduce errors into subsequent analyses, potentially resulting in flawed theoretical conclusions and misguided practical applications. This study presents a machine learning solution that employs autoencoder neural networks to detect aberrant response patterns in survey data as a computational method. We evaluated the effectiveness of autoencoder neural networks in identifying response anomalies through both simulated and real data. The results indicate that this approach can effectively detect anomalies in responses, providing researchers with more options for their analyses and subsequent conclusions. Ultimately, this enhances the trustworthiness of findings in psychological and educational research.
Differential item functioning (DIF) detection is an important yet understudied problem in computerized adaptive testing (CAT). In this article, we proposed a two-level logistic model to improve DIF detection in CAT by explicitly accounting for nuisance effects arising from CAT-induced structural dependency. First, we conceptualized that adaptive item selection induces systematic dependencies among examinees and items through provisional ability estimates, whereas traditional single-level DIF methods assume independent observations and may yield misleading results in CAT settings. Then, using a numeric example and Monte Carlo simulations, we compared our proposed two-level model with competing single-level models under various CAT conditions, manipulating test length, exposure control, ability estimator, DIF type, and DIF prevalence. Item-level Type-I error and statistical power conditional on joint model convergence were reported for each model. We showed that the proposed two-level model has improved control of spurious DIF and competitive power relative to single-level models, particularly with shorter tests and smaller exposure rates. However, we observed that the model convergence varied systematically across simulated conditions, highlighting that inferential accuracy and convergence reliability are intertwined in complex CAT DIF settings. Through this study, we underscored both the promise of multilevel DIF modeling in CAT and the need for future research to jointly evaluate convergence and inferential performance when assessing DIF models.
babebi is an R package for analysing complete two-time, two-rater pre–post rating designs. The package estimates pre–post effects using a linear model with a rater indicator as covariate and provides adjusted estimates of change, posterior summaries, and BIC-based Bayes factor approximations. It also includes Monte Carlo validation routines calibrated from the observed design to evaluate inferential performance under study-specific conditions.
Item preknowledge refers to the case where examinees have advanced knowledge of test material prior to taking the examination. When examinees have item preknowledge, the scores that result from those item responses are not true reflections of the examinee’s proficiency. Further, this contamination in the data also has an impact on the item parameter estimates and therefore has an impact on scores for all examinees, regardless of whether they had prior knowledge. To ensure the validity of test scores, it is essential to identify both issues: compromised items (CIs) and examinees with preknowledge (EWPs). In some cases, the CIs are known, and the task is reduced to determining the EWPs. However, due to the potential threat to validity, it is critical for high-stakes testing programs to have a process for routinely monitoring for evidence of EWPs, often when CIs are unknown. Further, even knowing that specific items may have been compromised does not guarantee that any examinees had prior access to those items, or that those examinees that did have prior access know how to effectively use the preknowledge. Therefore, this paper attempts to use response behavior to identify item preknowledge without knowledge of which items may or may not have been compromised. While most research in this area has relied on traditional psychometric models, we investigate the utility of an unsupervised machine learning algorithm, extended isolation forest (EIF), to detect EWPs. Similar to previous research, the response behavior being analyzed is response time (RT) and response accuracy (RA).
Romero et al. (2015; see also Wollack, 1997) developed the