
Research article
Select search scope: search across all journals or within the current journal

In many psychological experiments, interaction effects in factorial analysis of variance (ANOVA) designs are often estimated using total scores derived from classical test theory. However, interaction effects can be reduced or eliminated by nonlinear monotonic transformations of a dependent variable. Although cross-over interactions cannot be eliminated by trans formations, the meaningfulness of other interactions hinges on achieving a measurement scale level for which nonlinear transformations are inappropriate (i.e., at least interval scale level). Classical total test scores do not provide interval level measurement according to contemporary item response theory (IRT). Nevertheless, rarely are IRT models applied to achieve more optimal measurement properties and hence more meaningful interaction effects. This paper provides several condi tions under which interaction effects that are estimated from classical total scores, rather than IRT trait scores, can be misleading. Using derived asymptotic expecta tions from an IRT model, interaction effects of zero on the IRT trait scale were often not estimated as zero from the total score scale. Further, when nonzero inter actions were specified on the IRT trait scale, the esti mated interaction effects were biased inward when estimated from the total score scale. Test difficulty level determined both the direction and the magnitude of the biased interaction effects.
Most item selection in computerized adaptive testing is based on Fisher information (or item information). At each stage, an item is selected to maximize the Fisher information at the currently estimated trait level (θ). However, this application of Fisher information could be much less efficient than assumed if the estimators are not close to the true θ, especially at early stages of an adaptive test when the test length (number of items) is too short to provide an accurate estimate for true θ. It is argued here that selection procedures based on global information should be used, at least at early stages of a test when θ estimates are not likely to be close to the true θ. For this purpose, an item selection procedure based on average global information is proposed. Re sults from pilot simulation studies comparing the usual maximum item information item selection with the pro posed global information approach are reported, indicat ing that the new method leads to improvement in terms of bias and mean squared error reduction under many circumstances.

Binary or graded
This study compared three procedures—the Mantel- Haenszel (MH), the simultaneous item bias (SIB), and the logistic regression (LR) procedures—with respect to their Type I error rates and power to detect nonuniform dif ferential item functioning (DIF). Data were simulated to reflect a variety of conditions: The factors manipulated included sample size, ability distribution differences between the focal and the reference groups, proportion of DIF items in the test, DIF effect sizes, and type of item. 384 conditions were studied. Both the SIB and LR proce dures were equally powerful in detecting nonuniform DIF under most conditions. The MH procedure was not very effective in identifying nonuniform DIF items that had disordinal interactions. The Type I error rates were within the expected limits for the MH procedure and were higher than expected for the SIB and LR proce dures ; the SIB results showed an overall increase of approximately 1% over the LR results.
The concepts of reliability and validity and their associated coefficients typically have been restricted to a single measurement occasion. This paper describes dynamic generalizations of reliability and validity that will incorporate longitudinal or developmental models, using latent curve analysis. Initially a latent curve model is formulated to depict change. This longitudinal model is then incorporated into the classical definitions of reli ability and validity. This approach permits the separa tion of constancy or change from the indexes of reli ability and validity. Statistical estimation and hypoth esis testing be achieved using standard structural equations modeling computer programs. These longitu dinal models of reliability and validity are demon strated on sociological psychological data.
Williams & Zimmerman (1996) provided much- needed clarification on the reliability of gain scores. This commentary translates these ideas into recogniz able patterns of change that tend to produce reliable or unreliable gain scores. It also questions the relevance of the traditional idea of reliability to the measurement of change.
The properties of gain scores are linearly deter mined by the properties of their components. Thus, the reliability of a gain is uniquely determined by the reliabilities of the components, the correlation be tween them, and their standard deviations. Reliability is not inherently low, but the components of gains used in many investigations make low reliability likely. Correlations of the difference between two measures and a third variate are also determined uniquely by three correlations and two standard de viations. Raw score standard deviations frequently tell more about the measurement metric and how it is used than about the psychological processes underly ing the measurements. Correlations involving gains/ differences cannot be understood adequately unless the essential sample statistics of the components are known and reported.
The critiques of Collins (1996) and Humphreys (1996) certainly throw light on properties of gain scores and difference scores that have led to controversies in the past. Collins' examples reveal that familiar formulas for the reliability of differences do not adequately reflect the precision of measures of change, because they do not allow for intraindividual change. Some additional examples are provided here, and a similar argument is applied to the reliability of a single test. As Collins im plies, these arguments indeed disclose flaws, not only in the conventional approach to the reliability of gains and differences, but also in the basic concept of reliability in classical test theory.

In this, the first software review for