Abstract
Score linking is widely used to place scores from different assessments, or the same assessment under different conditions, onto a common scale. A central concern is whether the linking function is invariant across subpopulations, as violations may threaten fairness. However, evaluating subpopulation differences in linked scores is challenging because linking error is not independent of sampling and measurement error when the same data are used to estimate the linking function and to compare score distributions. We show that common approaches involving neglecting linking error or treating it as independent substantially overestimate the standard errors of subpopulation differences. We introduce new methods that account for linking error dependencies. Simulation results demonstrate the accuracy of the proposed methods, and a practical example with real data illustrates how improved standard error estimation enhances power for detecting subpopulation non-invariance.
Introduction
Alternate assessments are frequently linked to support score comparability across different administrations or across distinct assessments of the same construct (Battauz & Leôncio, 2023; Deng & Rios, 2022; Leôncio et al., 2023; Liu & Jurich, 2023; Moses, 2022; Wallmark et al., 2023). Changes in administration format, such as transitioning from paper-based to digital delivery, shifting from in-person testing to remote proctoring, or moving from non-adaptive to adaptive test designs, may also require linking across assessments (Anselmi et al., 2023; Castellano et al., 2023, 2024; Chao & Chen, 2023; Fishbein et al., 2018; Jewsbury et al., 2020; Jones et al., 2022; Tian & Choi, 2023; Weirich et al., 2024; Wyse, 2023). Linking may also be applied across different assessments, such as linking national to international tests, or state to national assessments (Lim & Sireci, 2017; Reardon et al., 2021). Looking ahead, emerging assessment paradigms featuring increased personalization, dynamic interactivity, and AI-driven functionalities will necessitate linking procedures for which the fairness and validity across diverse subpopulations should be empirically evaluated (Arslan et al., 2024; Bulut et al., 2024; Hao et al., 2024; Johnson, 2025; Sinharay & Johnson, 2024; Wu et al., 2025).
The gold standard data collection design for linking is the random-groups design, in which the examinees assessed on alternate assessments or administration conditions are randomly equivalent (Dorans, 2018; Kim et al., 2010). Depending on the study design, examinees or other sampling units such as schools may be assigned to multiple conditions, resulting in dependent samples. Alternatively, when no sampling units overlap across conditions, the samples may be independent. Assuming the assessments measure the same underlying construct, linking is used to align the score distributions across groups, typically by equating sample moments or percentiles of the score distributions (Kolen & Brennan, 2014).
Because linking aligns scores at the aggregate level, it is essential to evaluate score comparability for subpopulations to ensure that assessment or administration changes do not systematically benefit or disadvantage any major subpopulation. The extent to which the linking function holds across subpopulations, referred to as subpopulation invariance, can be empirically assessed to evaluate whether the link supports fair and valid comparisons for all groups (Brennan, 2008; Dorans, 2004; Dorans et al., 2008; Huggins, 2014; Liu & Dorans, 2013; Petersen, 2008; Sinharay & Haberman, 2014).
For illustrative purposes, consider scores from a new assessment that have been linearly linked to the metric of a prior assessment, or the target metric. Let
Subpopulation invariance may be evaluated by comparing summary statistics computed for the same subpopulation across the prior and new assessments (e.g., Yin, Bezirhan, et al., 2023). Let
As the linking coefficients Equations (2) and (3) are estimated from sample data, they are subject to sampling variability. This introduces uncertainty into the linked scores
While most approaches for quantifying linking variance focus on individual scores (e.g., Moses & Zhang, 2011; A. A. von Davier, 2013), some methods have been developed to estimate linking variance for summary statistics. Li (2022) proposed a bootstrap estimator for the standard error attributed to equating error on means. Mazzeo et al. (2018) introduced a Monte Carlo estimator for a linking variance component on means and other aggregate statistics. Johnson (1998) derived a first-order delta method approximation for the linking variance on means, and Kawahashi et al. (2024) derived a second-order delta method approximation for equating error affecting both scores and means.
However, as noted by each original author (Johnson, 1998; Kawahashi et al., 2024; Li, 2022; Mazzeo et al., 2018), these methods do not account for dependencies between linking error with sampling and measurement errors. Such dependencies may be particularly consequential when evaluating subpopulation invariance. Specifically, the same scores used to estimate the linking function coefficients are also used to compute the subpopulation estimates for the new assessment,
Less work has addressed the dependencies between linking error and other sources of error. Jewsbury (2024) examined this issue and noted that such dependencies are difficult to model because their structure varies depending on the type of score comparison being made. To facilitate methodological development, Jewsbury proposed classifying score comparisons into categories that share common dependency structures, referred to as scenarios. Two example scenarios include: (1) post-bridge to bridge: comparing an administration of the new assessment to the administration used to estimate the linking coefficients (i.e., the bridge administration) and (2) post-bridge to pre-bridge: comparing an administration of the new assessment after the bridge to an administration of the prior assessment before the bridge.
Jewsbury (2024) proposed a taxonomy of methods that address dependencies in score comparisons. Total-based methods account for all sources of uncertainty simultaneously, component-based methods decompose the variance into independent components corresponding to distinct sources (e.g., independent samples), and difference-based methods correct standard error estimates for linking error. Within the total-based and component-based categories, a further distinction is made between equation-based methods (i.e., analytically derived plug-in estimators) and resampling-based methods (e.g., Monte Carlo, jackknife, or bootstrap), while difference-based methods are always equation-based.
While Jewsbury (2024) derived methods within each class in the taxonomy for several scenarios, their focus was on score comparisons over time, particularly in the context of reporting trends in large-scale assessments. As such, their methods cannot be directly applied to subpopulation invariance, a distinct type of score comparison distinct from the scenarios they considered. However, the principles underlying their taxonomy of methods can be extended to derive applicable methods for subpopulation invariance, which is the focus of this paper.
Beyond issues of linking error and its dependencies, a third issue is whether the new and prior assessments were administered in dependent or independent samples. Samples are dependent if they share units (e.g., schools or individuals). Our proposed component-based method assumes independence, while our proposed total-based method accommodates dependencies. Applying the component-based method to dependent samples may introduce bias due to assumption violation, whereas using the total-based method with independent samples may be inefficient.
Present Paper
In this paper, we introduce novel methods for subpopulation invariance based on Jewsbury’s (2024) taxonomy: total-based resampling, total-based equation, component-based resampling, component-based equation, and a difference-based correction-term method. The three equation-based methods are derived analytically, while the two resampling-based methods are formulated using the jackknife resampling technique. For comparison, we include the conventional method, which ignores linking error, as well as the linking component method, which accounts for linking error but neglects linking error dependencies. These seven methods are detailed in the Theory and Methods section.
In the Simulation section, we evaluate the statistical properties of the seven methods through simulation. In the Practical Example section, we apply these methods to real data to demonstrate the impact of appropriately accounting for linking error and linking error dependencies on inferences.
Theory and Methods
Prior Methods
For comparison with the methods novel to this paper, we consider two alternate classes of methods: first, conventional methods that ignore linking error; and second, linking variance component methods that account for linking error but neglect dependencies between linking error and other sources.
Conventional Methods
A conventional method refers to any variance estimation approach that does not account for linking error, such as standard or default estimators not specifically adapted to incorporate linking variance. By analogy, we define conventional sources of error as those captured by conventional methods; namely, sampling and measurement error. For example, the conventional application of the jackknife (e.g., Wolter, 2007) to estimate the variance of
Note that the conventional method refers to applying a standard estimation procedure to the transformed scores of the new assessment,
The conventional estimator of the variance of
The conventional estimator of the variance of
Linking Component Methods
Second, the linking component approach estimates the linking variance on a statistic separately to the conventional sources of sampling and measurement, defined as the linking (variance) component. By assuming that linking error is independent to conventional sources, the sum of the estimated linking component and the conventionally estimated variance provides an estimate of the total variance (Jewsbury, 2019).
As described earlier, prior research on linking variance almost exclusively assumes independency of linking error and other sources (e.g., Johnson, 1998; Li, 2022; Mazzeo et al., 2018; Moses & Zhang, 2011). While Li (2022) uses a bootstrap estimator and Mazzeo et al. (2018) use a Monte Carlo estimator, we implement a jackknife estimator for consistency with other methods. The jackknife implementation of the linking component approach can be written as
Notice that the linking coefficients are re-estimated within each resample, whereas the untransformed score is held constant. Consequently, the linking component accounts for linking variance but not conventional sources of error, whereas the conventional estimator Equations (6) and (7) accounts for conventional sources of error, but not linking variance.
Equations for the linking component can be found in Jewsbury (2019; see Equations (42) and (48)).
New Methods
In this section, we introduce several new methods under the taxonomy of Jewsbury (2024) that account for linking dependencies with other sources and apply to tests of subpopulation invariance. The derivations follow a similar approach as Jewsbury (2024), but in contrast to their methods that concern differences over time, our methods apply to subpopulation invariance, that is, Equation (4). In addition to the derivations of the equation-based methods, we also denote how resampling-based methods may be implemented with the jackknife.
For greater generality, we derive the equations for composite scores. Composite scores are weighted sums of subscale scores and are used by some assessments such as NAEP (NCES, 2025d). Specifically, a composite score is defined as
In the case of composite scores, we assume that the link has been applied to the subscores. That is,
To avoid separate derivations for each type of statistic that might be evaluated for subpopulation invariance (e.g., mean and standard deviation), we follow Jewsbury (2024) in deriving error variances for the type U estimator, defined generally to include many types of statistical summary functions that satisfy the expression
The general form of the type U estimator allows us to derive variance expressions for many commonly used statistics, such as means and percentiles, without rederiving each separately. Specifically,
In the simulation and practical example, we focus on means for a unidimensional score. As noted, this classifies as type AB, so
Total-Based Approaches
Total-based methods directly evaluate the variance of an estimator while simultaneously accounting for all sources of variance, including linking and conventional sources such as sampling and measurement. Total-based resampling methods involves a resampling method such as the bootstrap or the jackknife, and accounts for linking variance by re-estimating and re-applying the linking function within each resample (for more general discussion, see Jewsbury, 2019, 2024). Total-based equation methods use expressions to estimate the variance of the estimator, considering all sources of error.
Total-Based Resampling
The jackknife implementation of total-based resampling may be written as
Notice that both the linking coefficients and the untransformed score vary across the resamples. This contrasts to the linking component approach where one is held fixed at a time while the other is allowed to vary; see Equations (6) and (8). Both linking and conventional sources of uncertainty are accounted for simultaneously with total-based resampling, inherently accounting for dependencies (Jewsbury, 2024).
Total-Based Equation
The estimator Equation (4) may also be written as
The variance of
Alternatively, the variance of
Component-Based Approaches
Component-based methods use Taylor series-like approximations to estimate components of variance, where each component represents an independent source of variance (see Jewsbury, 2024). For differences on linked scores, component-based methods are theoretically appropriate when the new and prior assessments are administered to independent samples. In this case, a component may be estimated to account for the variance of
Note that the difference between the linking component method (Li, 2022; Mazzeo et al., 2018) and Jewsbury’s (2024) component-based approach is that the linking component method defines a component associated with linking error, regardless of whether it is actually independent of other sources. In contrast, Jewsbury’s (2024) component based approach defines components for each independent source of error. In the context of subpopulation invariance, the linking component is the variance of the linking error on
Applying the component-based approach to subpopulation invariance, when the two samples are independent, produces one component associated with the new assessment sample and one component associated with the prior assessment sample. The former represents both conventional error on
Component-Based Resampling
The jackknife implementation of component-based resampling may be written as
Notice that each component is defined by holding constant all random variables associated with one sample, while re-estimating the random variables associated with the other sample. Consequently, component-based resampling is expected to be more efficient than total-based resampling (Equation (14)) when samples are independent, but biased when samples are dependent.
Component-Based Equation
Component estimators may be defined by the new assessment and the prior assessment separately, such that
The variances of the components are
The component-based approach assumes the total variance is equal to the sum of the variance of the components. That is,
Difference-Based Approaches: The Correction-Term Method
Difference-based methods aim to estimate the difference between the correct variance and the variance returned by a conventional method that neglects linking variance (Jewsbury, 2024, 2025). The primary motivation for developing difference-based approaches is that implementing them may be more convenient and less computationally intensive than other methods, especially if the conventionally estimated variance is already available. Consequently, we develop a difference-based method, the correction-term method for subpopulation invariance, that can be calculated without estimating any additional variances. Necessarily, this means that the correction-term method may be more approximate than other methods.
The estimator assumed by conventional methods is
Therefore, the variance of the estimator assumed by conventional methods is
The variance of the estimator may also be written as
The difference between the variance of the correct estimator and the conventional approach is
The simplified quadratic formula or correction-term method approximates the difference in Equation (18) as a quadratic function of the statistic of interest, with coefficients that can be pre-calculated to apply to any statistic of the same type and any scenario or type of score comparison (here, subpopulation invariance). That is,
To approximate Equation (18) in quadratic form, covariances involving the standardized scores are dropped because these terms are not invariant across all statistics. Furthermore, composite scores are approximated as if the composite scores were linked directly instead of the subscales being linked.
The parameters for the resulting correction term are
The correction-term method in full is
For reference, we define the complete correction-term method as a variant of the correction-term approach that retains the covariance terms involving
The complete correction-term method is defined with the same expressions as the correction-term method except that,
Simulation
The primary purpose of the simulation was to evaluate the statistical performance of the proposed and existing variance estimation methods in terms of bias, root mean square error (RMSE), and confidence interval coverage probability.
A secondary purpose was to evaluate the impact of the use of theoretically invalid methods. Specifically, (1) What is the impact of ignoring linking error by using a conventional variance estimator? (2) What is the impact of accounting for linking error but ignoring linking error dependencies, as in the linking component method? (3) What is the impact of assuming independence samples when samples are dependent?
The simulation followed a random-groups data collection design, in which a new and prior assessment were administered to randomly equivalent samples (Dorans, 2018). To induce non-normality, scores were generated using a linear regression model with dichotomous predictors, analogous to the statistical method used to generate the scores in large-scale assessments (Jewsbury, 2023).
Scores from the new administration were linearly transformed to the prior metric using Equations (1), (2), and (3). Subpopulation invariance was assessed by estimating the focal group mean difference across the two assessments with Equation (4), and several variance estimation methods were applied to estimate the associated standard error.
Each condition was replicated
Method
Conditions
The simulation varied the following conditions:
Data Generating Model
To simulate scores for two assessment samples, we generated the score vector
Letting
For the focal group
Linking coefficients were estimated with Equations (2) and (3) and applied to transform the sample
Variance Estimation Methods
Two existing methods were implemented: (1) the conventional jackknife method that ignores linking error and (2) the linking component jackknife method that accounts for linking error but ignores linking error dependencies. In the independent samples condition, the conventional method was implemented with Equation (6) and the linking component method was implemented with Equations (6) and (8). In the dependent samples condition, the conventional method was implemented with Equation (7) and the linking component method was implemented with Equations (7) and (8).
The five novel methods that account for both linking error and linking error dependencies were also implemented. The total-based resampling method was implemented using the jackknife in Equation (14) for both independent and dependent sample conditions. The total-based equation was implemented using Equation (15). Although the cross-sample covariances in the total-based equation could have been set to zero in the independent samples condition, we estimated all covariances in both independent and dependent sample conditions to align with the total-based resampling method. Similarly, the component-based resampling method was implemented using the jackknife in Equation (16) and the component-based equation method using Equation (17) for both independent and dependent sample conditions.
In the independent samples condition, the correction-term method was implemented with Equations (6) and (20), and the cross-sample covariances were set to zero when computing the correction-term coefficients. In the dependent samples condition, the correction-term method was implemented with Equations (7) and (20).
Outcomes
We evaluated the percentage bias, percentage RMSE and 95% confidence interval coverage probability for all methods. The percentage bias was calculated as
Percentage RMSE was calculated as
The confidence interval coverage probability was calculated as
Results and Discussion
The pattern of results was consistent across score normality and degree of subpopulation non-invariance. Therefore, we illustrate the results with a representative condition: non-normal scores and degree of subpopulation non-invariance
Because the existing methods exhibited substantially higher bias, inflating the y-axis range, differences among the proposed methods are difficult to discern when plotted with the existing methods. For this reason, we include two figures: Figure 1 shows only the proposed methods; Figure 2 includes all methods. Simulation: Statistical Properties of Novel Standard Error Estimators, Subpopulation Non-invariance Effect Size Simulation: Statistical Properties of all Standard Error Estimators, Subpopulation Non-invariance Effect Size 

Validity of Proposed Methods
The new total- and component-based methods performed well (Figure 1). As expected, total-based methods showed negligible bias and confidence interval coverage close to the nominal
Notably, the corresponding resampling and equation pairs performed very similarly with largely overlapping results in the plots (Figure 1). Consequently, the equations can be used to understand and predict the behavior of the resampling methods.
Impact of the Use of Theoretically Invalid Methods
Methods that do not properly account for the linking error, linking error dependencies, or sample dependence yielded poor results (Figure 2). The conventional estimator, which ignores linking error, showed consistently high positive bias and inflated RMSE, especially when the focal group made up a larger portion of the sample.
The linking component estimator, which includes a linking error term but ignores its dependence with the conventional estimate, had even worse bias (Figure 2). Since this estimator adds a positive component to an already positively biased estimate (see Equation (8)), the bias is necessarily increased.
When applied to dependent samples, component-based methods (which assume independence) also showed non-negligible biases (Figure 1). This finding supports the recommendation to use total-based methods for dependent samples. When samples were independent, both total- and component-based methods were approximately unbiased, with the latter showing slightly lower RMSE as previously noted.
Practical Example
We conducted an empirical analysis using data from the 2021 Progress in International Reading Literacy Study (PIRLS; von Davier et al., 2023). PIRLS was selected because it represents one of the most methodologically complex assessments, involving sampling weights, multistage sampling, and plausible values. In addition, its international design enabled replication of the analysis across multiple country samples.
This analysis served two purposes. First, it illustrates the application of the proposed methods to real data and demonstrates how power to detect subpopulation non-invariance improves when standard errors are correctly estimated. Second, it enables an empirical evaluation of the methods by assessing the plausibility of the results and the consistency of the methods in a complex real-world assessment setting.
PIRLS 2021 was the first administration in which countries could choose to administer a computer-based version of the assessment, departing from the prior paper-based assessment. Countries opting into the digital assessment were also assessed with the paper assessment in a randomly equivalent sample, enabling the metrics of the new (digital) and prior (paper) assessments to be linked (Almaskut et al., 2023).
We deviated from the official PIRLS procedure, which linked a pooled sample across all the participating countries (Yin, Fishbein, et al., 2023), and instead performed the linking within each country separately. This allowed us to evaluate statistical power more directly by replicating equivalent analyses across multiple country samples.
Within each country, we examined subpopulation invariance of the mean reading score conditional on student’s responses to the yes-no question, “Do you have any of these things in your home? Your own computer or tablet” (see IEA, 2020 for the specific question). This item was chosen as a likely candidate to exhibit subpopulation non-invariance, and in directions that can be reasonably anticipated (i.e., better performance on digital for students with a computer or tablet, better performance on paper for students without).
The analysis was conducted in R 4.4.2 with base functions (R Core Team, 2024). The code is provided in the Online Supplemental Materials.
Method
Data
The R data for 2021 PIRLS was downloaded from the IEA Data Repository (Fishbein et al., 2024). We identified 26 pairs of matching paper and digital datasets. To simplify the results, we included only samples that represented entire countries, excluding Belgium (Flemish language) and Moscow City. 1 The paired data for the United States, the 27th participant in the digital administration, was absent from the IEA Data Repository, possibly due to a decision not to report digital scores (NCES, 2025e).
Sample sizes for the digital assessments ranged from 3,030 to 8,551, with an outlier of 27,448 (United Arab Emirates). Paper-based samples ranged from 835 to 3,207 participants. The paper and digital samples were designed to be randomly equivalent. In some countries the samples were dependent (Yin, Bezirhan, et al., 2023). However, it is unclear whether the sample dependence could be modeled because the jackknife sampling zones may not have been structured to reflect this dependency (Foy & Almaskut, 2023). Therefore, consistent with the approach used in the PIRLS technical documentation (Yin, Fishbein, et al., 2023), we treated the samples as independent for analysis.
Analysis
For each country, we first unapplied the official PIRLS linking transformation from the digital plausible values using the published linking coefficients (Yin, Bezirhan, et al., 2023). We then estimated country-specific linking coefficients using Equations (2) and (3), and applied them to the digital plausible values via Equation (1).
PIRLS employs standard large-scale assessment methodology, including case weights, replicate weights, and plausible values. Due to its complexity, we refer readers to detailed treatments elsewhere of the basic methodology (e.g., Jewsbury et al., 2024; M. von Davier et al., 2023). Only key details and extensions of the basic methodology will be described here.
Estimates of population parameters, including the population mean and standard deviation estimates in Equations (2) and (3), were obtained using the standard methodology used in PIRLS (M. von Davier et al., 2023). That is, the parameters were estimated using the case weights 2 and for each plausible value, with results averaged across the plausible values.
Following official procedures, we constructed 250 replicate weights per country for the paired jackknife method (Foy & Almaskut, 2023). Within-imputation variances were estimated using these replicate weights for each plausible value and averaged. Between-imputation variances were estimated using standard procedures, and total variance was obtained via Rubin’s rules (Rubin, 1987).
Degrees of freedom for the subpopulation mean estimates
We estimated the subpopulation non-invariance of the mean reading score within levels of the variable “[do you have in your home] your own computer or tablet”
3
using Equation (4). The levels of the variables were labeled yes, no, and missing (specifically, omitted or invalid). Standard errors, confidence intervals, and
Applying the resampling methods (total and component resampling) was complex due to plausible values, which requires two variance estimators: one for within-imputation variance and one for between-imputation variance (Rubin, 1987). We adapted both the within- and between-imputation variance estimators to account for linking variance under the total and component approaches. Full details on these adaptations are provided in the Appendix.
Equation-based methods (total equation, component equation, and correction-term) were implementing by plugging-in estimated parameters and (co)variances into Equations (15), (17), and (20), respectively. The parameters and (co)variances were estimated using standard PIRLS estimation procedures (Foy & Almaskut, 2023). For the total equation and correction-term methods, the cross-sample covariances were set to zero. Consequently all methods with the exception of the total resampling method assumed sample independence. We also implemented the complete correction-term method using Equations (21) and (22) as a reference for the correction-term method.
Results and Discussion
Impact of Accounting for Linking Variance
Given the scope of the analysis—applying seven variance estimation methods across three subpopulations in 24 countries—the full set of results is provided in the Online Supplemental Materials. In the main text, Figure 3 illustrates the increase in power obtained by using valid standard error estimators versus ignoring linking error when testing for subpopulation invariance. To reduce visual complexity, only the component-based methods are shown alongside the conventional method; these were selected based on their superior performance under sample independence in the simulation study. Figure 4 displays a heatmap of significance probability values across countries and subpopulations to assess the empirical consistency of all of the proposed methods. 95% Confidence Intervals of Subpopulation Mean Non-invariance Estimates Based on Alternate Standard Error Estimation Methods. Note. Yes, No, and Missing are subpopulations defined by the response to the question “[do you have in your home] your own computer or tablet.” Darker confidence intervals indicates significance at Heatmap of Significance Probability Values Based on the Proposed Standard Error Estimation Methods. Note. Yes, No, and Missing are subpopulations defined by the response to the question “[do you have in your home] your own computer or tablet”. 

Consistent with the simulation results, the conventional method produced larger confidence intervals relative to the component resampling and component equation methods. This overestimation was more pronounced in the larger subgroup, Own computer at home: yes (mean proportion = 63.8%, SD = 12.8%), compared to the no subgroup (mean proportion = 33.7%, SD = 12.4%). As expected, all three methods yielded similar estimates for the very small missing subgroup (mean proportion = 1.4%, SD = 1.0%).
While truth is unknown with real data, we can consider the results in terms of plausibility. According to the law of total variance, the variance of
Figure 3 visually illustrates this discrepancy for the yes group. With correct standard errors, the cross-country spread in
The results indicate a meaningful increase in power when standard errors are correctly estimated (Figure 3). With the conventional method, none of the yes group comparisons reached statistical significance. However, using the component-based methods,
For the no group, one country showed significant
For the missing subgroup, two countries showed significant results across all three methods, and a third country was significant only under the component-based methods. Again, all significant estimates were positive, suggesting students who did not answer the question performed better on the paper version. Potentially, students who struggle with cognitive assessments administered in digital format may also struggle to respond to contextual questions administered in digital format.
Consistency of the Proposed Methods
Figure 4 displays a heatmap of significance probability values, illustrating the empirical consistency of the proposed methods. The four total and component-based methods produced identical significance outcomes in all but one instance (see Online Supplemental Materials). The only exception was a single result from the total resampling method, and some differences were expected as all other methods assumed sample independence. However, even in that case, the difference in
The correction-term method showed meaningful discrepancies from the other methods (Figure 4), and produced negative variance estimates for 11 countries in the yes group and 1 country in the no group. The poor performance can be attributed to the additional approximations of the correction-term approach where statistic-specific covariances are dropped, as the complete correction term method including the statistic-specific covariances was highly aligned with the other methods.
The underperformance of the correction-term method was more pronounced for the yes group, likely because of the magnitude of the correction. For the United Arab Emirates yes group, the conventional estimate of
Despite the apparent sensitivity of the correction-term method to approximations, the degree of underperformance of the method was not observed in the simulation, even with large correction magnitudes (e.g., 90% and 95% focal group conditions). The discrepancy may reflect the limited degrees of freedom of the PIRLS variance estimators. The paired jackknife procedure applied to imputed data has degrees of freedom lower than the number of jackknife sampling zones (Barnard & Rubin, 1999), which can be as few as 11 (Foy & Almaskut, 2023). The between-imputation variance estimator with five plausible values has only four degrees of freedom (Casella & Berger, 2002).
Discussion
This paper introduced five new methods for evaluating subpopulation invariance within the broader taxonomy of linking variance methods proposed by Jewsbury (2024). These methods were validated through simulation studies, where they demonstrated substantially improved statistical properties compared to existing approaches. In the context of subpopulation invariance, properly accounting for linking variance leads to reduced standard errors and increased statistical power. This increase in power was illustrated through an empirical application using data from the 2021 PIRLS assessment.
In random-groups designs, population invariance is ensured by construction via equating the marginal score distributions across groups (see Equation (1)). Accounting for linking variance in the estimation of standard errors enhances power to detect subpopulation non-invariance by correctly conditioning on the fact that population-level invariance has already been imposed by design (Equation (15)).
The results underscore the critical impact of correct standard error estimation. Using conventional methods that ignore linking variance led to dramatic overestimation of standard errors, resulting in severe loss of power, as seen in both simulation and the PIRLS example. Methods that include linking error but ignore its dependencies with other error sources, as most existing linking methods do (e.g., Johnson, 1998; Kawahashi et al., 2024; Li, 2022; Mazzeo et al., 2018), only exacerbated the overestimation. Similarly, applying methods that assume sample independence to dependent samples (e.g., using component-based methods where total-based methods are appropriate) resulted in notable bias.
Across both the simulation and the empirical analysis, resampling-based and equation-based methods yielded highly consistent results. Given their distinct theoretical foundations, this agreement offers additional validity evidence. Resampling methods, while more computationally intensive, are more flexible and can in principle accommodate non-linear linking functions (e.g., equipercentile) or more complex statistics (e.g., the standardized Root Mean Square Difference from Dorans & Holland, 2000). In contrast, equation-based methods offer greater transparency and interpretability, but are currently limited to linear linking and simpler summary statistics (e.g., means, standard deviations, and percentiles).
The distinction between our total-based and component-based methods centers on the treatment of sample dependence. Total-based methods explicitly accommodate dependence, while component-based methods assume independence. This was reflected in the simulation results: total-based methods outperformed under dependence, whereas component-based methods were slightly more efficient (lower RMSE) under independence.
One unexpected finding was the relative underperformance of the correction-term method, particularly in the empirical analysis. Although this method is intended to offer a computationally efficient approximation, prior work (Jewsbury, 2024) found that correction-term approaches performs comparably to less approximate methods in the context of trend comparisons. The discrepancy may stem from the relative magnitude of the correction term: in trend analyses, the correction tends to be positive and can be small relative to the total variance, whereas in subpopulation invariance analyses, the correction is negative and can exceed the total variance by several multiples (see e.g., Figures 2 and 3). Under such conditions, even modest proportional error in estimating the correction can lead to substantial mis-estimation of the standard error. Accordingly, use of the correction-term method should be accompanied by sensitivity checks against more robust methods.
As demonstrated in the practical example, the proposed methods are applicable even in the highly complex context of large-scale educational assessments. The PIRLS example involved case weights, multistage sampling, and plausible values. In particular, plausible values require that variance be decomposed into within-imputation and between-imputation variances (Jewsbury et al., 2024). Both PIRLS’ paired jackknife estimator for the within-imputation variance and the Monte Carlo estimator for the between-imputation variance were extended to account for linking variance under both the total and component resampling approaches (see Appendix).
While this paper focuses on linking with random-groups data collection designs (Dorans, 2018), other types of linking exist. The methods can also apply to pseudo-equivalent groups linking after the adjustment for non-equivalence has been applied (e.g., Haberman, 2015). However, linking variance with a non-equivalent anchor test design (NEAT) is more distinct (Robitzsch & Lüdtke, 2019; 2024; Xu & von Davier, 2010). For NEAT designs, the double jackknife may be the most analogous to the methods considered here, especially the total resampling method (Haberman et al., 2009).
In sum, this paper demonstrates the critical importance of properly accounting for linking variance when evaluating subpopulation invariance. We proposed several new methods that address this issue, validated them through simulation, and illustrated their utility through an empirical application.
Supplemental Material
Supplemental Material - Standard Error Estimation for Subpopulation Non-invariance
Supplemental Material for Standard Error Estimation for Subpopulation Non-invariance by Paul A. Jewsbury in Applied Psychological Measurement
Footnotes
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The study was funded by ETS.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
Supplemental Material
Supplemental material for this article is available online.
Notes
Appendix
Applying the methods to large-scale assessments such as PIRLS involving case weights, complex samples, and plausible values adds additional complexity. The proposed equation-based estimators (total equation, component equation, and correction term) can be estimated by applying the standard procedures recommended by the large-scale assessment to estimate each (co)variance in the equation (Jewsbury et al., 2024). However, applying our proposed resampling methods is less straightforward. In this Appendix we describe how the resampling-based methods can be applied through extending the existing variance estimators used in large-scale assessments.
Large-scale assessments using plausible values follow Rubin’s rules to estimate parameters and their variances (Rubin, 1987). That is,
In the context of subpopulation invariance, the statistic of interest
In the practical example, we adapt PIRLS’ variance estimators to incorporate linking variance. As is typical in large-scale assessments, PIRLS uses a resampling-based method to estimate the within-imputation variance. These estimators typically take the form:
In PIRLS,
To incorporate linking variance into the within-imputation estimator (Equation (31)), we modify how the replicate estimate
To fully account for linking variance, the between-imputation variance estimator must also be adapted. To express this, we re-write the standard formula as
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
