Abstract
This study investigates how different rounding rules and ways of providing Angoff standard-setting judgments affect cut-scores. A simulation design based on data from the National Assessment of Education Progress was used to investigate how rounding judgments to the nearest whole number (e.g., 0, 1, 2, etc.), nearest 0.05, or nearest two decimal places for individual items or clusters of items affected cut-scores for individual panelists and a group of panelists across four different pools of items. For the simulated ratings from a group of panelists, the recovery of the cut-scores was examined using the mean and the median. Results showed that rounding to nearest whole number had the potential to produce fairly large statistical biases in cut-score estimates. Biases were less when judgments were simulated across cluster of items. The largest biases were found at the advanced cut-score, but the greatest potential changes in the percentage of students that would be above the cut-score were found for the basic cut-score. Rounding to the nearest 0.05 or nearest two decimals places did not have a large impact on cut-score estimates and had little effect on the percentage of students above the cut-score. Implications for policy and future standard-setting practices are provided.
Introduction
As the need and call to set cut-scores on large-scale assessments has become more pronounced, several new methods for deriving cut-scores have been proposed and implemented. Many of these new methods can be broadly described as variations of the Angoff method (Angoff, 1971). Some variations of the Angoff procedure proposed in recent years include extensions to polytomous items and mixed-format tests (e.g., extended Angoff, Hambleton & Plake, 1995; Angoff with mean estimation, Reckase, 2000), providing judgments for subsets of items instead of individual items (e.g., direct consensus procedure; Sireci, Hambleton, & Pitoniak, 2004), rounding judgments to the nearest whole number instead of two decimal places (e.g., extended Angoff, Hambleton & Plake, 1995; yes/no procedure, Angoff, 1971; Impara & Plake, 1997; item score string approach, Bay, 1998; Ferdous & Plake, 2008; Reckase & Bay, 1999), and the use of bubble sheets where item judgments are recorded in 0.05 increments (Nichols, Twing, Mueller, & O’Malley, 2010; Taube, 1997). However, with the proposal of these methods, many requisite statistical investigations of cut-scores have not been completed.
The failure to examine the statistical functioning of a standard-setting method before it is used operationally means that it is possible that a method may result in statistically biased cut-scores. Bias in this context is defined as the difference between a panelist’s intended cut-score and the estimated value of the cut-score for the panelist over his or her standard-setting judgments (Reckase, 2006). Here, bias is viewed as it is in many parameter recovery studies where the goal is that a cut-score that a panelist intends to set, the intended cut-score, is the number that is estimated for a panelist from his or her standard-setting ratings. Any difference between the estimated value of the cut-score and the intended cut-score constitutes bias. In other areas of statistics, bias is defined as the difference between the true value of the parameter and the expected value of the parameter over replications. Central to this definition of statistical bias is the intended cut-score, which is conceptualized as a hypothetical point on the scale underlying the assessment that is the panelist’s best judgment of a cut-score from his or her interpretation of the performance-level descriptor (PLD; Reckase, 2006). The intended cut-score is the ideal place that a panelist would place his or her cut-score if he or she made no error in translating his or her understanding of the PLD to the scale underlying the assessment.
One can see how the question of whether or not a standard-setting method will yield the cut-score that a panelist or group of panelists intended is an extremely important question. In educational accountability settings, the cut-scores from standard setting often are adopted by the policy board that decides on the final location of the cut-scores that are ultimately used to make decisions about school and student performance. If the adopted cut-scores are biased representations of what the panelists or group of panelists intended when they provided their recommendations this may result in the unintended consequence of misclassifying schools or students into different categories than they should have been classified into based on the panelists’ understandings of the PLDs.
Therefore, the purpose of this study is to examine how well different rules used in Angoff-type standard-setting methods affect cut-scores. Our study expands on the previous literature in several important ways. First, we discuss from a theoretical perspective not only how the concept of an intended cut-score can be applied for an individual panelist but also how it might be utilized for a group of panelists. Second, our study examines several ways of combining Angoff ratings that have not received a lot of attention in the research literature. In particular, we directly compare the impact of panelists rounding judgments to two decimals, the nearest whole number, and the impact of rounding using something similar to bubble sheets where judgments are recorded to the nearest 0.05. We also examine how rounding would affect cut-scores when judgments are rounded for individual items or across cluster of items. In addition, we investigate how well the mean and median work for recovering intended group cut-scores. Since our study simulates data for panelists based on National Assessment of Educational Progress (NAEP) data, we are also able to couch our discussion in a framework familiar to researchers and practitioners.
Angoff Method and Its Derivatives
The original Angoff method (Angoff, 1971) asks panelists to provide a yes/no judgment about whether a minimally competent examinee (MCE)—an examinee that just possesses the necessary skills, knowledge, and ability to meet the standard—would get each item correct, and it was suggested in a footnote that probability judgments could be used as an alternative. Now, the Angoff method most often refers to a process where probability judgments are made. Each of these probability judgments is summed to arrive at the cut-score on the number-correct score scale for an individual panelist. The average or median over panelists is the final group cut-score. When item response theory (IRT) models are employed, the cut-score for a panelist is typically determined by translating the panelist’s cut-score on the number-correct score scale to the θ scale through the IRT test characteristic curve. The original Angoff procedure is among the most researched and applied standard-setting methods (Brandon, 2004; Mehrens, 1995).
A significant number of researchers have questioned whether the task of providing Angoff probability judgments is overly complex (Shepard, 1995; Shepard, Glaser, Linn, & Bohrnstedt, 1993) and whether panelists can carry out the procedure with sufficient accuracy (Impara & Plake, 1998; Plake & Impara, 2001; Schulz, 2006; Shepard, 1995; Shepard et al., 1993). Examining early NAEP standard setting, Shepard and colleagues (Shepard, 1995; Shepard et al., 1993) found that panelists who used the Angoff procedure struggled to rate consistently with item difficulties or across item types. They also suggested that panelists tended to round their ratings and avoided providing ratings at the extremes of the scale. Impara and Plake (Impara & Plake, 1998; Plake & Impara, 2001) showed that a sample of teachers had difficulty using the Angoff procedure and estimating probabilities in general. Schulz (2006) provided an additional argument that the Angoff method can be affected by rater inconsistency. Although some scholars disagree with these sentiments (Cizek, 2001; Hambleton, et al., 2000), especially when panelists receive sufficient feedback (Reckase, 2001), criticisms of the Angoff procedure and other standard-setting methods are still common.
Common Modifications
Many of the new Angoff derivative methods modify the Angoff procedure in similar ways. One common modification is to have panelists provide the most likely score on each item expressed as a whole number of score points (i.e., a score of 0, 1, 2, etc.). For example, the extended Angoff method, which is designed for polytomous item formats, asks panelists to indicate the most likely score as a whole number of points. Similarly, the yes/no method and the item score string approach also ask for whole number judgments. Another recent modification is to have panelists provide their standard-setting judgments on a bubble sheet where recommendations can be given in 0.05 increments. These modifications for providing judgments stand in contrast to the traditional Angoff procedure or the Angoff method with mean estimation where judgments are typically expressed having two decimal points. This raises the question of what the impact of rounding judgments to the nearest whole number or the nearest 0.05 is when compared with rounding to two decimal places.
A second modification is to ask panelists to provide judgments across clusters of items. This modification reduces the time needed to provide judgments as well as the cognitive burden of having to provide judgments for each item. An example of this modification is found in the early rounds of the direct consensus method. In this case, panelists provide judgments of the score a MCE should receive on clusters of items. The clusters of items are typically based on content domains or common stimuli (e.g., reading passages). This modification raises the question of what the impact of providing judgments across a cluster of items is when compared with providing judgments for each individual item.
A third and somewhat different question is how different ways of aggregating the panelist judgments affects group cut-score estimates. The mean is typically used when there is a desire to consider each panelist’s cut-score equally when calculating the group cut-score. The median is typically used when one is concerned about potential outlier panelists having a large influence on the cut-score. However, an examination of whether the mean or median has the potential to produce less statistical bias for a group of panelists has not been fully addressed. Examining these three questions is the focus of the investigations in this article.
Previous Research
A limited amount of research has directly investigated these three questions in the research literature. One study that investigated the impact of rounding to the nearest whole number is the study by Reckase and Bay (1999). Their research investigated the item score string approach and showed that there could be potential biases when the distribution of the p values for the items was not symmetrically distributed around a panelist’s cut-score. However, their study has not been published in a scholarly journal and they did not directly examine cut-scores across a group of panelists.
Another study that examined rounding is the study by Reckase (2006). He examined the impact of rounding to one or two decimal places compared with not rounding ratings under the Rasch and three-parameter logistic (3PL) IRT models for an individual panelist. He found minimal potential biases when rounding to two decimal places with the largest biases found at the extremes of the score scale. He did not examine what the biases would be if the judgments were instead rounded to the nearest whole number or if they were rounded to the nearest 0.05. Similar to Reckase and Bay (1999), Reckase (2006) did not examine how rounding would affect a group of panelists. Examining potential biases across panelists is an important question because it is possible that small biases for individual panelists may cancel out for a group (Schulz, 2006). The modification of rounding ratings for clusters of items has not been directly examined in the research literature.
Several studies have indirectly investigated the use of the mean and the median for computing group cut-scores in Angoff standard setting. Typically, these studies just report group cut-scores. For example, Busch and Jaeger (1990) investigated cut-scores on a National Teacher Examination. They found that cut-scores were very similar, although there were some differences between some panels and rounds. The reason that the mean and the median have probably not been systematically compared is probably due to the fact that their properties are well known. When the distribution of scores is symmetric, the mean is equal to the median. When the distribution has outliers or is not symmetric, the mean and the median will deviate from one another. In this study, we focus on how the mean and the median work for recovering intended cut-scores. Some differences are expected since the distribution of panelists is not symmetric.
Theoretical Framework
The theoretical framework for this study comes from Reckase (2006) and focuses on the first criterion that he suggested for evaluating standard setting. This criterion is based on an intended cut-score and examining the extent to which a standard-setting method would produce that intended cut-score if the panelist carries the method out as it is designed to be implemented. For example, if panelists are instructed to conceptualize and consider the MCE and provide judgments in an Angoff-type standard setting on each item of what they think the MCE should know and be able to do, then the estimated cut-score from these ratings should be equal to and consistent with a single location on the underlying score scale that is equivalent to their intended cut-score.
The intended cut-score is a hypothetical number on the underlying score scale for a panelist that we assume is a number on an IRT θ scale. We use IRT because IRT is typically used to scale many large-scale assessments, including NAEP. The intended cut-score can be viewed as the best attempt of a panelist to define a cut-score on the underlying scale from his or her understanding of the PLD. The intended cut-score is panelist specific and may change depending on the panelist or the context in which panelists provide their judgments. The intended cut-score may also be different across rounds of standard setting. When applying the concept of an intended cut-score to evaluate standard setting, the question asked is whether or not the estimated value over the judgments provided by a panelist is indicative and in line with a single number on the underlying scale that is his or her intended cut-score. Any difference between the intended cut-score and the estimated value for a panelist is defined as bias.
The concept of an intended cut-score can also be extended to a group level. This extension is important since the concern in most standard settings is with whether or not the cut-score estimated for the group of panelists is consistent with the group cut-score that would be calculated if each of the panelists gave judgments in line with his or her intended cut-score. This extension uses the same framework for determining bias in cut-scores for individual panelists proposed by Reckase (2006) and then makes the additional assumption that the panelists provide independent ratings. The assumption of independent ratings is hard to test and may not hold in operational applications, but is usually used when computing statistics in standard setting (Schulz & Mitzel, 2005).
After making this assumption, one can evaluate the cut-score at the group level by aggregating the individual panelists’ intended cut-scores and then comparing this aggregated value to the cut-score obtained when aggregating across the panelists’ cut-scores in the simulation of the standard-setting method. For example, one can evaluate how well the mean works in a particular context by taking the mean of the panelist-intended cut-scores and comparing this aggregated value to the mean of the simulated panelist cut-scores for a particular method. A similar strategy can also be applied with the median.
It is important to point out that the view that standard setting should be evaluated based on the extent to which a standard-setting method results in a panelist’s or group of panelists’ intended cut-scores is not the only approach that has been used for evaluating standard setting. In other contexts, standard setting has been evaluated by collecting evidence focusing on the relationship of stimuli that a panelist is asked to judge with other data. Usually, these evaluations are made without considering the relationship of the judgments provided on these stimuli to a single location on the underlying scale. For example, a common way of evaluating the item-level judgments in the Angoff method is to look at the item ratings provided by panelists and to correlate these judgments with the estimated p values of the items (Brandon, 2004; Goodwin, 1999; Hurtz & Auerbach, 2003; Hurtz & Jones, 2009). The idea is that when panelists are providing consistent and accurate ratings, the correlation between the items and the p values should be high (Goodwin, 1999; Hurtz & Auerbach, 2003; Hurtz & Jones, 2009).
However, without considering how the judgments provided for the standard-setting stimuli directly relate to the underlying scale on which students are scored, it can become challenging to detect certain patterns of errors in standard-setting judgments. For example, it is possible that a panelist systematically judges items to be higher or lower than he or she intended. This is known as a rater being too lenient or too harsh. A systematic uniform shift in which judgments are changed by the same amount would not change the value of the correlation between the item ratings and p values. Panelists may also systematically regress their ratings toward the middle of the scale as has been suggested by Shepard et al. (1993) and Schultz (2006). These errors again would be hard to detect using a correlation between ratings and p values. These errors, however, could be simulated and the recovery of cut-scores examined using the concept of an intend cut-score. Reckase (2006) provides an example of how errors in judgment can be incorporated into his evaluation approach.
The point is that this approach or other approaches that evaluate standard setting without considering the relationship of standard-setting stimuli to the underlying scale do not answer the question that should be of interest in standard setting. This question is how well are panelists able to translate their understanding of the PLD to the underlying score scale used to make decisions about examinees with the standard-setting method? In this article, this question is asked and investigated by examining how well the panelist’s intended cut-score would be recovered across a range of variations of the Angoff method using an IRT-based framework.
To simplify the investigations, in this study we assume that the standard setting is only for a single round. This allows us to avoid modeling complexities that more rounds of standard setting would create or adding assumptions that panelist ratings are not systematically influenced by other panelists or prior ratings. Most standard settings use several rounds where panelists interact, receive feedback, and are moderated by the same facilitator. These aspects can create complex dependencies and require more complicated models and simulations than the ones in this article. We also do not include error in the simulated conditions since an important concern is how the methods would perform in an ideal setting. Hence, this study presents the impact of the Angoff modifications and rounding rules in the ideal situation that panelists understand the standard-setting task within the constraints imposed by the modifications or rounding rules.
Data and Method
NAEP Data
Data in this study were simulated based on panelist distributions and item parameters from the 2005 NAEP 12th grade mathematics Angoff standard-setting pilot study (ACT, Inc., 2005). We use information on the item parameters (2PL, 3PL, or generalized partial credit model parameters) and the panelist distributions for the 20 panelists in the last round of standard setting to simulate data. We also quantify the impact of the potential biases in group cut-score estimates by using data on the percent above cut-score (PAC) available from the standard setting.
The mean, median, and standard deviation of the assumed hypothetical intended cut-scores for the 20 panelists are displayed in Table 1. Table 1 shows that the hypothetical group cut-scores on the underlying IRT θ scale are very similar when using the mean and the median and that the variability of the group of panelists is somewhat less for the proficient cut-score than for the other two cut-scores. The mean and the median in Table 1 were the aggregated cut-scores at the group level that we compared the group cut-score estimates with in the simulations.
Mean, Median, and Standard Deviation of Hypothetical Intended Cut-Scores
Basing the simulation on NAEP is compelling because it facilitates a coherent understanding of how well the modifications would perform on a large-scale assessment used for making important policy decisions. Of course, these inferences can only be drawn in the idealized scenario that the panelists performed their tasks in line with the way that the methods were simulated. If the methods are statistically biased and would have large impacts on cut-scores in these settings, it is questionable whether the methods should be used in other situations as well.
Simulation
A complex simulation design was used to evaluate the different rounding rules for the Angoff method. First, we simulated the panelist’s ratings on each item by assuming that the panelist’s θ cut-score in the last round of NAEP standard setting was his or her intended cut-score. The ratings simulated for each panelist were his or her expected true scores on the items. They were determined by finding the value of the item characteristic curve for the IRT model that was applied on the item at that θ. For items following the 2 PL model, the expected true score was
where θ was the panelist’s intended cut-score, a i was the discrimination parameter for item i, and b i was the difficulty parameter. For items following the 3 PL model, the expected true score was
where θ was the panelist’s intended cut-score, a i was the discrimination parameter for item i, b i was the difficulty parameter, and c i was the pseudo-guessing parameter. For items following the generalized partial credit model, the expected true score was
where
and P ix (θ) denotes the probability of a score of x on the item i assuming the panelist’s intended cut-score θ, m i represented the highest possible score for item i, a i was the discrimination parameter for item i, and b ik was the threshold parameter between category k and category k + 1.
The true scores on each of the items were then used to determine each panelist’s rating in each condition. For the conditions involving ratings for individual items, each of the true scores on the items were rounded to nearest whole number (e.g., 0 or 1 for dichotomous items and 0, 1, 2, 3, etc. for polytomous items), nearest 0.05, or nearest two decimals. For example, if the true score for an item estimated from the assumed intend cut-score was 1.2345, the item would be rounded to 1 using rounding to nearest whole number, 1.25 using rounding to the nearest 0.05, and 1.23 using rounding to nearest two decimal places. These rounded item scores were summed across items to determine the panelist’s estimated cut-score on the number-correct score scale. For the conditions involving ratings at the cluster level, the expected item scores were found for each item and summed across the 23 Teacher Domains in the 2005 NAEP pilot study. These 23 Teacher Domains focus on similar mathematical concepts. The ratings for these 23 clusters were then rounded to the nearest whole number, nearest 0.05, or nearest two decimal places. The rounded sum across the 23 clusters was taken as the panelist’s estimated cut-score on the number-correct score scale. Finally, the Newton–Raphson procedure was used to compute the estimated cut-score in the θ metric for each panelist in each condition by finding the θ value that corresponds to that number-correct score.
This process was repeated for four different pools of items and three different cut-scores. The three different cut-scores were the basic, proficient, and advanced cut-scores. The first pool of items consisted of all 180 items used in the NAEP mathematics pilot study. This pool of items allowed for statistical biases to be examined in the situation in which all the possible items in the NAEP item pool were used in standard setting. In practice, all these items are not used in standard setting and instead panelists are assigned to rate one of two sets of items. The numbers of items in these two item pools were 107 items and 109 items, respectively. Ratings were also simulated for each of these two sets of items. This allowed for statistical biases in actual NAEP-type standard settings to be quantified. The fourth pool of items was simulated by randomly selecting three blocks of items from the 10 blocks in the NAEP item pool. These three blocks resulted in 53 items. This shorter test length is similar to the test lengths used in many state testing programs and allows for statistical biases in these situations to be quantified.
In total, there were three forms of rounding (nearest whole number, nearest 0.05, or nearest two decimal places), two ways of providing judgments (each individual item vs. clusters of items), four item pools (180 items [full pool], 107 items [Pool 1], 109 items [Pool 2], and 53 items [Three Blocks of Items]), and three cut-scores (basic, proficient, and advanced). These conditions were fully crossed producing 72 simulated conditions for each panelist. In addition, the mean and the median were also investigated yielding 72 total comparisons for the recovery of each group cut-score.
Evaluation Criteria
The criteria that we examined as part of this study were related to the statistical biases at the individual and group levels. At the individual level, the statistical bias in a panelist’s intended cut-score is defined as
where θ
j
> was the assumed intended cut-score for panelist j and
where n was the number of panelists and the other terms have the same meaning as in Equation (5). The desire for this statistic is that it is small and close to zero since this indicates that the biases in the individual panelist cut-scores are relatively small.
At the group level, the interest in this study was in statistical bias for the group’s intended cut-score. This can be represented as
when the mean was used to compute the cut-score, and
when the median was used to compute the cut-score. The notation
We also examined what the difference in the PAC would be for the group of panelists if the cut-score was based on the rounded scores in each simulated condition compared with the intended cut-scores. This provides an indication of the practical impact of the rounding of the scores in each condition.
Results
Individual Panelist Level
The results of the evaluation of the individual panelist cut-scores are displayed in Table 2. Several clear patterns are apparent in Table 2. First, the effect of rounding to the nearest whole number compared with rounding the nearest 0.05 or the nearest two decimal places dramatically affected individual panelist cut-scores. In all cases, the average absolute bias was higher when rounding to the nearest whole number compared with rounding to the nearest 0.05 or nearest two decimal places. Rounding to the nearest 0.05 increased the average absolute bias over the nearest two decimal places, although in all cases the differences between the rounding to the nearest 0.05 and nearest two decimal places was less than 0.01. The fact that the average absolute biases were quite small in magnitude for both rounding to the nearest two decimals or 0.05 suggests that both these rounding rules worked fairly well.
Average Absolute Bias in Individual Panelist’s Cut-Scores on the θ Scale
Across the three cut-score placements, the greatest amount of average absolute bias occurred in computing the advanced cut-scores, followed by the basic and proficient cut-scores. In the case of the advanced cut-scores, the distribution of items above and below the 0.5 cutoff used in rounding to the nearest whole number was the least balanced, resulting in greater average absolute bias. Basic cut-scores tended to have more bias than the proficient cut-scores because the proficient cut-scores were less extreme and because there were more items on the scale that were closer to the panelists’ intended cut-scores. This reduced the impact of the rounding.
Rounding at the cluster level also appeared to reduce potential biases compared with rounding for each individual item. The lower level of statistical bias at the cluster level occurred because there were fewer stimuli that were rounded in computing the cut-score. The rounding of fewer stimuli reduced the magnitude of the rounding errors. Similar to when the judgments were rounded for individual items, the average absolute bias was greatest for the advanced and basic cut-scores at the cluster level.
Examining the results across the item pools suggests that, as expected, using different sets and numbers of items affected cut-scores. The most dramatic differences occurred for the basic and advanced cut-scores for the three blocks of items compared with the other item pools. In these cases, the basic and advanced cut-scores had greater average absolute bias. However, this does not suggest that a panelist’s cut-score always will be more accurate with more items. For example, if the fewer test items were closer and more evenly distributed near the location of the panelist’s intended cut-score, then the cut-score could be more accurately recovered with fewer items. In fact, the average absolute bias for proficient cut-score for the three blocks of items is less than the average absolute bias for whole item pool when individual item judgments were rounded to the nearest whole number. This highlights the importance of knowing the location of the panelist’s intended cut-score and its relationship to the rated items. It also suggests that just using a large number of test items will not necessarily ameliorate the effects of rounding. This issue is considered in greater detail in the Discussion section.
Group Panelist Level
Table 3 shows the potential statistical bias in the mean and median cut-score estimates for the group of panelists. Examining Table 3 shows that in some cases the mean is better recovered, whereas in other cases the median is better recovered. This finding is not unexpected since the distribution of panelists was not symmetric and the aggregated individual cut-scores were different in each condition. In such cases, it is common to observe that the mean or median perform differently.
Bias in Mean and Median Group Cut-Score Estimates on θ Scale
Similar to findings for the individual panelists, the amount of bias from using the mean or the median was a function of the location of the panelist’s intended cut-score and the items used. Table 3 suggests that the basic cut-scores were negatively biased while the proficient and advance cut-scores were positively biased when rounding the individual items to the nearest whole number. When rounding to the nearest whole number at the cluster level, the proficient cut-score was negatively biased. The advanced cut-scores produced the greatest amount of bias when compared with the cut-score placements at the other proficiency levels. This suggests that in the case of rounding to the nearest whole number, bias in cut-score estimates did not cancel out for a group of panelists.
The bias for both the mean and the median was less than 0.01 at the group level for both rounding to the nearest 0.05 or the nearest two decimal places in all conditions. This suggests that both forms of rounding work quite well if panelists are able to provide consistent judgments. These findings are not surprising given that the average absolute bias for the individual panelist cut-score estimates was also small in magnitude. Again, more test items did not necessarily mean better recovery of group cut-scores.
The amount of statistical bias in the group cut-scores from rounding to nearest whole number was mitigated somewhat by rounding at the cluster level. This again was a function of rounding after summing all the item scores versus rounding each of the individual item scores (Table 3). This suggests that if panelists are able to perform the standard-setting task accurately group cut-scores will be more accurately recovered when judgments are provided across clusters of items than for individual items.
Impact on Percent Above Cut-Score
To provide greater perspective into the impact of biases observed in Table 3, we examined how the potential biases would change the PAC if the cut-score was placed at the group cut-score estimate instead of at the group’s intended cut-scores. These percentages were found by taking the difference between PAC for the simulated cut-score estimate and the PAC for the intended cut-score after finding the PAC for the closest θ value on the NAEP scale in the pilot study. We choose to round the values to the closest θ value because in practice scores are typically determined from transforming the θ values to scale scores and then rounding those scores. Table 4 shows the impact of these biases on the PAC.
Change in Percent Above Cut-Score for Group Cut-Score Estimates Compared With Group Intended Cut-Scores
In all cases, rounding to the nearest 0.05 (as would be done when using bubble sheets) or rounding to the nearest 0.01 did not change the PAC. This is occurred because relatively small changes in the cut-score on the θ scale has virtually no effect on the cut-score on the reported score scale. These results confirm that when panelists are able to carry out the standard-setting task consistently and in line with the way the methods are designed, rounding to the nearest 0.05 or nearest 0.01 will have a minimal impact on cut-scores.
The impact of rounding to the nearest whole number (as is done in the yes/no approach, the extended Angoff procedure, the item score string approach, or in early rounds of the direct consensus procedure) can have a dramatic impact on the PAC, however. If one rounded to nearest whole number on each of the individual items the differences in PAC when using the mean would range between roughly 5.610 and 13.010 for the basic cut-score, between −3.823 and −4.387 for the proficient cut-score, and between −1.156 and −1.262 for the advanced cut-score. For the median, differences would range between 4.490 and 14.190 for the basic cut-score, between −4.387 and −5.346 for the proficient cut-score, and between −1.156 and −1.343 for the advanced cut-score.
The positive differences in the PAC for the basic cut-scores occurred because the negative biases raise the number of students exceeding the cut-score. The opposite holds true when cut-scores exhibit positive biases. Also, it is important to point out that greater changes in PAC are observed at the basic and proficient cut-scores despite the fact that these cut-scores have less bias on the θ scale than the advanced cut-scores. This occurs for these data since there are more students near the basic and proficient cut-scores than near the advanced cut-scores. The biases for both mean and median are substantially smaller when rounding across clusters of items. In some cases, they can even be zero. However, there are still some situations in which the potential biases can exceed 3% and all changes in PAC are not zero for all cut-score placements.
These data illustrate that the magnitudes of the biases when rounding to the nearest whole number would be large enough that they would significantly affect quantities such as AYP. This suggests that rounding to the nearest whole number does not present a viable alternative in Angoff-type standard settings.
Discussion
The purpose of this study was to investigate several rounding rules used in the Angoff procedure and how these rules affect cut-scores in a simulated setting. Results suggested that rounding to the nearest whole number could severally affect the cut-scores for both individual panelists and a group of panelists. The effect of rounding to the nearest whole number was mitigated somewhat when the rounding was done at the cluster level compared with the individual item due to use of fewer stimuli that were rounded. However, biases still remained.
The results in this article do not suggest that rounding to the nearest whole number will always produce a large amount of statistical bias or change the PAC. Similarly, the results do not suggest that using more test items will necessarily produce less bias. To determine the amount of statistical bias when rounding to the nearest whole number, one needs to know the items that the panelist is asked to rate and their location in relationship to the panelist’s intended cut-score.
To solidify this point, consider two different set of items. The first set of items consists of 10 items that are fit by the Rasch model and are equally spaced from θ = −2 to θ = 2 in an ordered item booklet (OIB; i.e., a booklet of items ordered in terms of their difficulty) with a response probability criterion (i.e., the probability of getting the item correct) of 0.50. The second set of data consist of 20 items that are also fit by the Rasch model and are equally spaced from θ = −1 to θ = 3. Also, assume that the panelist intended cut-score is θ = 0.
Figure 1 displays the two sets of items and the panelist’s intended cut-score of θ = 0 with a dotted line. In the first set of data, there are five items rounded to a score of 1 (i.e., the five items below the dotted line) and five items rounded to a score of 0 (i.e., the five items above the dotted line), resulting in a cut-score on a number-correct scale of 5. The value of the test characteristic curve at the intended cut-score is also 5, meaning when the cut-score is translated back to the θ scale the θ value is 0 and the bias is 0. For the second set of data, there are five items rounded to a score of 1 and 15 items rounded to a score of 0 for a cut-score on the number-correct scale of 5 and θ value of −0.438. This leads to bias of −0.438 on the θ scale.

Locations and potential bias for two sets of Rasch items
This example helps illustrate some of what we observed in our simulation. First, it shows that it is possible to get more bias when rounding to the nearest whole number in the Angoff procedure with more items. Second, the example shows the importance of knowing the panelist’s intended cut-score and the properties of the rated items. If one knew the location of a panelist’s intended cut-score one could select a set of items so that the effect of rounding would roughly cancel out. In particular, one could use the concept of an OIB from bookmark standard setting (Lewis, Mitzel, & Green, 1996; Mitzel, Lewis, Patz, & Green, 2001) and one could design an OIB with a response probability criterion of 0.50 so that roughly half of the items were above the panelist’s cut-score and the other half of the items were below their cut-score. At the cluster level, the OIB would be designed in a similar manner to what is used for polytomous items where the probability of obtaining a score at that level or higher is set equal to the response probability criterion of 0.50. The reason for designing the OIB based on a response probability criterion of 0.50 instead of some other response probability criterion, such as 0.67, is that if roughly half of the items were above and below the panelist’s cut-score in the OIB, then it would allow the effect of rounding to the nearest whole number to roughly cancel out over the full set of items or cluster of items, respectively.
Even though an approach to reduce the effect of rounding to the nearest whole number can be devised in theory, in practice this approach would not work for a couple of reasons. First, it is impossible to know in advance exactly where a panelist intends to set his or her cut-score. This makes it impossible to know in advance how to design an OIB that is targeted to the intended cut-score. Second, even if one knew the panelist’s intended cut-score in practical situations most panelists do not agree about the location of the cut-score. This means that it is almost impossible to design a set of items for the Angoff derivative methods that use rounding to the nearest whole number that will not result in some statistical bias for some panelists.
Further complicating the issue is the fact that in most large-scale testing programs multiple cut-scores need to be set. In fact, every state with an NCLB accountability program sets multiple cut-scores. The essential realization is that since these cut-scores are often at different locations at least one of the cut-scores would produce statistical bias that is practically significant. This occurs because it is impossible to design an OIB with a response probability criterion of 0.50 in which half of the items would be above and below each of the cut-score placements. The greater the disparity in the number of items above and below the desired cut-score placement the greater the potential for statistical bias. This suggests that Angoff derivative methods that use rounding to the nearest whole number present many potential problems for setting cut-scores in large-scale assessment contexts.
Conversely, results from our simulation suggested that cut-scores were well recovered when judgments were rounded to the nearest 0.05 (as is done with bubble sheets) or the nearest two decimal places for individual items or across clusters of items. This suggests that these rounding rules are viable alternatives for setting cut-scores when panelists understand the standard-setting task and provide standard-setting judgments consistent with a single cut-score location. It also suggests that in the context of NAEP common procedures used to set cut-scores would have very little bias if panelists carry them out as they are designed and explained.
However, as some researchers have pointed out, in real standard-setting situations panelists are not completely consistent in their Angoff judgments. The impact and pattern of these inconsistencies in most operational standard settings is unknown, which makes it extremely challenging to know how much potential statistical bias there is in operational standard setting. Our perspective is that the use of sufficient training and feedback as part of the standard setting can serve to ameliorate most of these issues. In the Angoff procedure, the Reckase chart (Reckase, 2001) has proven to be quite effective and helpful for informing panelists about their judgments and reducing rater inconsistency in NAEP. Additional research on the potential biases present in operational standard setting would be a very valuable direction for future research.
Finally, it is important to emphasize one of the unique theoretical parts of this study, which consisted of the further development of a conceptual framework that can be employed to examine the statistical bias at the group level. Examples in this article illustrated how this concept could be applied to evaluate Angoff-type standard-setting methods. We believe that beginning to research how well cut-scores may be recovered using the concept of an intended cut-score at both the individual and group levels is fundamental given the high stakes that are often attached to cut-scores. If a standard-setting method is not able to recover hypothetical intended cut-scores in the situation in which panelists perfectly understood the standard-setting task, the use of the standard-setting method in any other situation would be called into serious jeopardy.
Footnotes
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The author(s) received no financial support for the research, authorship, and/or publication of this article.
