Abstract
In this study, smoothing and scaling approaches are compared for estimating subscore-to-composite scaling results involving composites computed as rounded and weighted combinations of subscores. The considered smoothing and scaling approaches included those based on raw data, on smoothing the bivariate distribution of the subscores, on smoothing the bivariate distribution of the subscore and weighted composite, and two weighted averages of the raw and smoothed marginal distributions. Results from simulations showed that the approaches differed in terms of their estimation accuracy for scaling situations with smaller and larger sample sizes, and on weighted composite distributions of varied complexity.
Introduction
The creation of a composite score as a weighted combination of subscores is a common practice in large-scale testing, arising when the construct of interest is assessed with a battery of subtests and content areas, or when the test contains subtests with items of different formats (e.g., multiple-choice and constructed response items and subtests). The weights of the subscores used in forming the weighted composite have implications for the measurement characteristics of the composite score (i.e., reliability and validity). Other implications involve the smoothing, scaling, and equating procedures used to maintain consistent reported scores across alternate composite forms, or to express equating results for one of the subscores on the scale of the weighted composite. The interest of this study is the influence of subscores’ weights on the smoothing and scaling procedures used to obtain subscore-to-composite scaling results. The next section describes three subscore weighting choices which result in simple and more complex subscore-to-composite scaling situations commonly encountered in practice. For these three situations, four approaches to smoothing the score distributions and converting the scores of one of the subscores to the weighted composite scale are described and then compared in simulations.
Three Examples of Composite Score Weightings and Distributions
Three examples of weighted composites encountered in testing programs’ tests are illustrated using the single group data from von Davier, Holland, and Thayer (2004, p. 115). These data contain the scores on two 20-point tests, X and Y, as taken by one group of 1,453 examinees. The X and Y data are summarized in Tables 1 to 3 and are used in weighted form to produce three hypothetical rounded and weighted composites as
Descriptive Stats of X, Y, and Composite,
Descriptive Stats of X, Y and Composite,
Descriptive Stats of X, Y and Composite,
Some implications of the three sets of weights for composite scaling procedures are evident in the three weighted composite score distributions from Tables 1 to 3. Figures 1 to 3 plot two forms of the distributions of the composite scores based on the three sets of X and Y weights and the bivariate XY population distribution used in this study (described in the “Method” section). For Figures 1 to 3, the A-graphs show the expected number of examinees at each weighted composite score and the B-graphs show the total number of score combinations of X and Y that produce each rounded and weighted composite score. Plotting the numbers of composite scores based on the weighted score combinations of X and Y (B) along with the numbers of examinees at each composite score (A) is useful for showing the correspondence of the composite score distributions with the popularity of the composite scores as determined by the sets of weights. For the

Distributions of (A) the expected number of examinees and (B) the number of X and Y score combinations resulting in each weighted composite score,

Distributions of (A) the expected number of examinees and (B) the number of X and Y score combinations resulting in each weighted composite score,

Distributions of (A) the expected number of examinees and (B) the number of X and Y score combinations resulting in each weighted composite score,
Four Approaches to X-to-Composite Scaling
The approaches of interest for converting the scores of one of the subscores (i.e., X) to the scale of the weighted composite all involve the estimation of an equipercentile relationship for X and the weighted composite,
where
Raw data
Probably the most direct approach to implementing Equation (1) to produce an X-to-composite scaling is the raw test score data from the group of examinees who take tests X and Y (Raw Data). This approach has the advantage of being relatively simple. However, raw equipercentile results are often not completely “raw.” Equipercentile estimates can be awkward and problematic with only raw data, such as when specific scores of X, Y, and the composite are unobserved in the raw data and the resulting percentile estimates are not unique for every possible score. There are several ad hoc approaches to addressing problems due to unobserved scores for equipercentile conversions, and the one of interest in this study involves estimating the probabilities of X (and similarly for the composite) from the observed data (
Equipercentile results from a Raw Data approach would be expected to be relatively unbiased because whatever structures might occur in the distributions of X and the composite ought to be visible in the observed data (Moses & Holland, 2007). Raw Data results would also be expected to be highly variable due to the relatively strong influence of sampling variability on the estimated probabilities.
Smoothing the bivariate X–composite distribution
In practice, Equation (1) is usually implemented using smoothed probabilities rather than raw probabilities of the X and weighted composite distributions. The smoothing is intended to improve the accuracy of the estimated equipercentile results by reducing the influence of sampling variability but not increasing estimation bias very much (Moses & Holland, 2007). A common smoothing approach involves fitting a loglinear model to the test score distributions (Holland & Thayer, 2000). The type of loglinear model most likely considered for an X-to-composite situation would relate the log of the expected bivariate probabilities as a linear function of the X and composite scores,
where, with maximum likelihood estimation of the
Smoothing the bivariate XY distribution
Another smoothing approach involves fitting a loglinear model to the bivariate distribution of X and Y,
and obtaining the marginal distributions of X and the weighted composite by appropriately aggregating the
Thus, to fit the observed mean, variance, and skew of the unrounded weighted composite and its observed covariance with X, the model in Equation (4) could be extended to include all the X and Y terms in Equations (5) to (8),
Weighted averages of the raw and smoothed data
A final smoothing approach is considered that is an attempt to improve the accuracy of the Raw Data approach in Equation (2). By noting that Equation (2) is essentially a weighted average of the observed probabilities of the marginal distributions and the probabilities from uniform distributions, it is possible to consider a more general Weighted Average approach defined as,
where
(Agresti, 1990; Bishop, Fienberg, & Holland, 1975; Fienberg & Holland, 1973; Moses & Oh, 2009). Equation (10) has been described as a pseudo or empirical Bayes estimate in that standard multinomial assumptions for
Assessing the Four Smoothing and Scaling Approaches for the Three Weighted Composites
The approaches for X-to-composite scaling described in the previous section vary in terms of their simplicity and plausibility. Although the XY smoothing approach would seem to be a more accurate reflection of the population model that produced the subscore data, it can also be more complex to implement and perhaps not appreciably better than the other three approaches in terms of accuracy, simplicity, and efficiency. This study was intended to compare the approaches’ accuracy implications for the three weighted composite situations described in Figures 1 to 3.
Method
The smoothing and scaling approaches of interest were compared in a series of simulations. A loglinear model of the form in Equation (9) was fit to the von Davier et al. (2004)XY data (Tables 1-3) and treated as the population model for the XY distribution. From the XY population model, three composite score distributions were obtained from the three sets of weights for X and Y (Tables 1-3, Figures 1-3), and equipercentile X-to-composite scaling functions based on Equation (1) were calculated for the three composites and treated as three population scaling functions. Finally, samples of 1,000 and 10,000 examinees’XY score combinations were randomly drawn from the XY population model and the four smoothing and scaling approaches were used to estimate the population X-to-composite scaling functions in the sample data. The accuracies of the smoothing and scaling approaches were summarized by averaging the squared deviations of these methods’ sample estimates from the population values.
Population Distributions and Scaling Functions
The population model used to study the smoothing and scaling approaches was chosen based on the suggested loglinear model from von Davier et al.’s (2004) investigations, using I and H values of 3 in Equation (9) to fit the mean, variance, and skewness of X and Y, and also fitting the XY, covariance. Fitting the mean, variance, and skewness of the three unrounded weighted composite score distributions was also desired because it was consistent with how the X and Y distributions were modeled, and so as suggested in Equation (9) the results of Equations (5) to (8) were incorporated by including and fitting the XY2 and X2Y cross-moments. Population X-to-composite scaling functions were obtained by first computing the three weighted composite score distributions from the XY population model using one of the three sets of weights for X and Y and then calculating the three equipercentile X-to-composite scaling functions (Equation 1).
Smoothing and Scaling Approaches
The Raw Data, XComp smoothing, XY smoothing, and Weighted Average smoothing approaches described in the “Introduction” section were assessed in the simulations of this study. To be consistent with the population model, the smoothing models used with the XComp smoothing, XY smoothing, and weighted average approaches were implemented by fitting the means, variances, and skewnesses of the marginal distributions these approaches modeled. The XY smoothing was implemented by fitting Equation (9) with I = H = 3. The XComp smoothing was implemented by fitting Equation (3) with I = H = 3. The two estimated models used with the Weighted Average approach for the marginal distributions of X and the composite were obtained as loglinear models like Equation (11) that fit the means, variances, and skewnesses of the marginal distributions. For the Weighted Average approach several weights were considered and
Simulations and Evaluations
The simulations involved drawing 500 independent and random samples of 1,000 scores and 500 other independent and random samples of 10,000 scores from the XY population model and computing the samples’X-to-composite scaling with the considered smoothing and scaling approaches. The values of 1,000 and 10,000 were selected because they produced results which are illustrative of the approaches’ performance tendencies for data sets with smaller and larger sample sizes. The simulation process was repeated for each of the three X and Y weights and their corresponding X-to-composite scaling functions. Comparisons of the smoothing and scaling approaches’ estimates with the dth random sample of size N,
where
The smoothing and scaling approaches were evaluated in greater detail for the three X-to-composite scaling situations by considering the score-level biases of each approach,
The score-level standard errors of each approach were also considered,
Results
The smoothing and scaling approaches’RMSE results for the three scaling situations and two sample sizes are summarized in Table 4. The results show that scaling functions produced with XY smoothing were the most accurate (i.e., had the smallest RMSE values) for all three scaling situations and the two sample sizes. The superior accuracy of the XY smoothing was to some extent an artifact of the simulations, in that the XY smoothing was based on the actual population model used to generate the XY sample data. The second most accurate smoothing and scaling approach was usually the XComp smoothing, except for the X-to-composite scaling where the weighted composite was obtained with the
Root Mean Squared Error (RMSE) Results.
More detailed comparisons of the performances of the scaling and smoothing approaches are presented in terms of the methods’ score-level biases (Figures 4-6) and score-level standard errors (Figures 7-9). Figures 4 to 6 show that score-level biases do not vary much with respect to sample size, but are relatively large for X-to-composite scalings involving the more complicated weighted composite distributed computed with the

Conditional biases of the smoothing and scaling approaches for the

Conditional biases of the smoothing and scaling approaches for the

Conditional biases of the smoothing and scaling approaches for the

Conditional standard errors of the smoothing and scaling approaches for the

Conditional standard errors of the smoothing and scaling approaches for the

Conditional standard errors of the smoothing and scaling approaches for the
Discussion
Like the distribution of a composite formed as a weighted combination of subscores, the scaling of one of the subscores to the weighted composite has a level of complexity that is largely determined by the weights of the subscores. As described in this study, there are several approaches that could be used to estimate subscore-to-composite scalings. Although somewhat unfamiliar, the XY smoothing approach based on smoothing the subscore distributions and estimating the weighted composite distribution and subscore-to-composite scaling from the smoothed subscore distributions is a direct reflection of the subscore distributions. The XY smoothing approach might also be considered a more plausible data generation model than the other approaches because of its treatment of the weighted composite distribution as a by-product of the subscore data and distributions. Simulations based on treating the smoothed subscore distributions and subscore-to-composite scalings from the XY smoothing approach as population quantities show the unsurprising result that using the XY smoothing approach in simulations results in greater estimation accuracy than other approaches. Other subscore-to-composite scaling approaches that deal with the composite distribution more simply and more directly may be less accurate, though these approaches can be of practical interest due to being more familiar and, possibly, more easily and efficiently implemented.
The subscore-to-composite scaling approach based on smoothing the joint distribution of the subscore and composite is likely to be implemented in practice, as it is the single group scaling method commonly described in equating texts (von Davier et al., 2004, pp. 113-130). This XComp approach produced the second most accurate scaling results for most of the conditions of this study. The XComp approach does especially well when sample sizes are smaller and when the weighted composite is relatively simple. For scaling situations where sample sizes are large and the weighted composite distribution is complex, the XComp approach can be relatively inaccurate because of relatively large estimation bias (Figure 6B). Additional simulations not reported here show that the increased estimation bias of the XComp approach can be partially addressed using indicator functions to model a complex weighted composite distribution (Holland & Thayer, 2000; von Davier et al., 2004), though bias never improves to the level of the XY smoothing approach.
Weighted Averaging was considered in this study as an approach that could potentially improve the estimation accuracy resulting from the Raw Data approach. Because the Weighted Average approach considered in this study deals only with the marginal distributions of the subscore and composite, it is an efficient strategy that avoids bivariate problems such as the identification of score combinations with “structural zeros” (Holland & Thayer, 2000, p. 144) that can be difficult with very complex weighted composite distributions. Although usually less accurate than the XY and XComp approaches, the Weighted Average approach did improve on the estimates obtained from the Raw Data approach, where averages based on 99% smoothed data improved estimation for smaller sample sizes (due to reduced variability, Figures 7-9), and averages based on 1% smoothed data improved estimation for larger sample sizes (due to smaller biases, Figures 4-6).
Most of the smoothing and scaling approaches were shown to produce scaling estimates with improved accuracy relative to scaling with raw data. However, the approaches’ accuracy improvements varied based on sample size, the complexity of the weighted composite score distribution, and the complexity of the subscore-to-composite scaling. These issues suggest that the findings of this study may be more illustrative than generalizable with respect to scaling situations not considered in this study. Going beyond the current study, weighted composites could be created using weights selected for maximizing composite reliability (Woodbury & Lord, 1955), maximizing validity (Woodbury & Novick, 1967), or from more than two subtests that exhibit higher or lower intercorrelations and different standard deviations than those considered in this study. Future studies may find that the accuracies of the subscore-to-composite smoothing and scaling approaches might vary for weighted composites created in other ways. The current study demonstrates a simulation approach that can be useful for assessing smoothing and scaling approaches’ performances for other scaling situations involving weighted composites.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
