Abstract
Fleiss’s Kappa is an extension of Cohen’s Kappa, developed to assess the degree of interrater agreement among multiple raters or methods classifying subjects using categorical scales. Like Cohen’s Kappa, it adjusts the observed proportion of agreement to account for agreement expected by chance. However, over time, several paradoxes and interpretative challenges have been identified, largely stemming from the assumption of random chance agreement and the sensitivity of the coefficient to the number of raters. Interpreting Fleiss’s Kappa can be particularly difficult due to its dependence on the distribution of categories and prevalence patterns. This paper argues that a portion of the observed agreement may be better explained by the interaction between category prevalence and inherent category characteristics, such as ambiguity, appeal, or social desirability, rather than by chance alone. By shifting away from the assumption of random rater assignment, the paper introduces a novel agreement coefficient that adjusts for the expected agreement by accounting for category prevalence, providing a more accurate measure of interrater reliability in the presence of imbalanced category distributions. It also examines the theoretical justification for this new measure, its interpretability, its standard error, and the robustness of its estimates in simulation and practical applications.
Keywords
Introduction
Cohen (1960) introduced the Kappa coefficient to assess the degree of interrater agreement between two raters or methods evaluating the same subjects. Fleiss (1971) later extended this measure to accommodate multiple raters assigning categorical ratings on a nominal scale. Several Kappa-like agreement measures have been proposed in the literature, including Gwet’s AC1 (2001) and Krippendorff’s Alpha (2004). Interrater agreement reflects the extent to which raters’ judgments are consistent, reproducible, and reliable. Such raters may include clinical psychologists, teachers, social scientists, or professionals from various fields—such as medicine, psychology, sociology, or business—who evaluate subjects on a particular phenomenon or trait using categorical rating scales with any number of categories (Cousineau & Laurencelle, 2017; Dawson, 2017; Kraemer et al., 2002; Moons & Vandervieren, 2022).
Researchers studying aggressive behavior in children may observe and rate instances of aggression within a controlled play environment. Each episode is recorded, and the observers’ ratings are later compared to assess the level of interrater agreement concerning the presence and intensity of aggression. Similarly, when evaluating anxiety, three different methods—self-report questionnaires, clinical interviews, and psychological assessments—may be used to classify adult participants into four levels of anxiety. The goal is to determine the extent of interrater agreement among these methods in accurately classifying anxiety levels. In the context of test development, subject-matter experts evaluate items to assess content validity. They rate the extent to which each item measures the intended construct, and their ratings are compared to determine interrater agreement regarding item validity.
Fleiss’s Kappa offers greater flexibility than Cohen’s Kappa, as it accommodates any number of raters—a feature particularly useful in a variety of research contexts and applications. Both Cohen (1960) and Fleiss (1971) recognized that a certain level of agreement among raters can occur by chance. Cohen defined this chance agreement as the joint product of the marginal probabilities for each category, whereas Fleiss defined it as the square of the sum of the marginal probabilities for each category. Accordingly, both coefficients—Cohen’s Kappa and Fleiss’s Kappa—are calculated by subtracting the expected chance agreement from the observed proportion of agreement.
However, it is important to clarify that the chance agreement in these measures does not assume that raters are assigning categories randomly in the naive sense. Rather, it reflects the level of agreement that would be expected if each rater applied their judgments independently—whether those judgments stem from noise, differing interpretations, or consistent but unaligned understandings of the categories. This point has been emphasized in critiques (e.g., Brennan & Prediger, 1981), which argue that the statistical correction may not reflect real-world rating processes, where raters are typically trained or otherwise expected to apply consistent criteria.
Since the development of Cohen’s Kappa and Fleiss’s Kappa as chance-corrected measures of interrater agreement, several interpretive issues and paradoxes have been identified (Brennan & Prediger, 1981; Delgado & Tibau, 2019; Warrens, 2010; Zec et al., 2017). As noted by Feinstein and Cicchetti (1990a) and Mandrekar (2011), these challenges arise because Kappa statistics are influenced not only by the actual agreement between raters but also by confounding factors such as rater bias and the prevalence of categories. One key criticism is that the “chance agreement” in Kappa reflects the level of agreement expected under the assumption that raters assign categories independently, based on possibly differing interpretations—not necessarily that their assignments are random in a naive sense. Consequently, Kappa can yield deceptively low values even when observed agreement is high, or vary inconsistently across studies with similar rating behaviors.
Therefore, relying solely on Kappa values to assess agreement can be misleading. To enhance interpretation, researchers are encouraged to report additional metrics—such as prevalence indices and bias indices—alongside Kappa, and to clearly state any assumptions regarding the raters’ shared or divergent understandings of the rating categories. This provides a more transparent and nuanced account of the factors influencing interrater agreement.
Moreover, several key factors influence the behavior and interpretation of agreement measures: number of raters, number of categories, category prevalence, variability in prevalence patterns, category definition and interpretability, and rater bias and shared frameworks. For the effect of the number of raters, the effect of adding more raters on agreement is not uniform—it depends on the context. In many real-world settings, increasing the number of raters can lead to more disagreement, simply because each rater brings a unique interpretation, and more opinions create more room for divergence. However, in structured settings with shared training or clearly defined categories, additional raters may actually increase agreement by stabilizing assessments and averaging out individual variability.
When the observed proportion agreement is fixed when comparing interrater agreement among different number of raters, the true proportion agreement and interrater agreement coefficients is expected be smaller for more raters. This is because achieving the same level of observed agreement with more raters generally requires stronger agreement patterns, leading to a relative decrease in true agreement. A well-designed agreement measure should remain interpretable across different numbers of raters and avoid systematic bias (inflation or deflation) due to changes in rater count alone. This claim is examined in greater detail in a subsequent section.
As the number of categories increases, agreement typically becomes more difficult to achieve due to finer distinctions and greater opportunities for disagreement. A good agreement measure should reflect this increase in complexity, but not penalize raters disproportionately for using more nuanced rating schemes.
Moreover, agreement measures should account for the effect of category prevalence. In interpreting the effects of category prevalence on agreement, the quantity
When certain categories dominate, high agreement may arise due to base rates rather than shared understanding. Conversely, when categories are rarely used, observed agreement may be low even if raters are applying similar criteria. A conceptually sound agreement measure should adjust for these base rate effects in a way that distinguishes between structural prevalence and true concordance. For variability in prevalence patterns, when raters differ in how frequently they use categories—whether due to interpretation, bias, or training—agreement measures must consider this variability. Some indices (e.g., Cohen’s Kappa) are sensitive to asymmetric marginal distributions, which may or may not reflect disagreement. Measures that assume symmetrical or identical marginal distributions may misrepresent the true level of concordance.
The clarity or ambiguity of category definitions directly affects agreement. Clear, unambiguous categories foster higher agreement, while vague or overlapping categories increase interpretive divergence. An ideal agreement measure should not assume all categories are equally interpretable or equally likely to be used. Finally, agreement may be influenced by systematic rater biases, training differences, or varying conceptual frameworks. Agreement measures that assume all raters share the same understanding of categories may overstate true consensus. Conversely, assuming fully independent interpretations may understate meaningful agreement. A conceptually robust approach must acknowledge the spectrum between shared and divergent interpretations.
It is well-established that observed proportion agreement is influenced by the prevalence of rating categories. The way in which these categories are defined significantly affects how raters interpret and evaluate the subject matter, which in turn shapes the distribution of responses. Category prevalence directly impacts the calculation of Fleiss’s Kappa by altering the expected agreement under chance. Specifically, Fleiss’s Kappa estimates chance agreement based on the marginal distributions of the raters, which can reflect either independent judgments or systematic differences in category usage (Feinstein & Cicchetti, 1990b; Fleiss et al., 1969, 1979; Vach, 2004).
When category distributions are imbalanced—that is, when one category is substantially more prevalent than others—the expected chance agreement becomes skewed. This can result in inflation or deflation of the Kappa statistic, depending on which category dominates. In such situations, Kappa may fail to reflect the actual level of interrater agreement: high prevalence can artificially increase Kappa, overstating agreement, while low prevalence can suppress it, understating agreement (Byrt et al., 1993).
These distortions highlight the importance of explicitly addressing how prevalence and interpretive alignment among raters influence agreement metrics. This consideration underlies our motivation to explore a prevalence-corrected agreement measure that more accurately captures agreement independent of category imbalance.
To address the criticisms and paradoxes associated with category prevalence and rater bias, several researchers have proposed alternative agreement coefficients to Cohen’s Kappa and Fleiss’s Kappa. Notable examples include Gwet’s AC1 (2001) and Krippendorff’s alpha (2004). These measures aim to offer more stable estimates of agreement, particularly under conditions of prevalence imbalance or minimal bias. However, despite these improvements, all of these coefficients remain chance-adjusted and continue to rely—explicitly or implicitly—on the assumption of random category assignment by raters. As such, they are not immune to the foundational limitations that affect traditional chance-corrected agreement measures.
Taken together, these factors point to the need for agreement measures that are not only mathematically sound but also behaviorally and contextually grounded. Our proposed coefficient aims to better align with these conceptual expectations by adjusting for the systematic effects of category prevalence while preserving the interpretive value of observed agreement.
This paper introduces a novel agreement coefficient that moves beyond traditional chance-adjusted measures, aiming to address several issues and paradoxes associated with existing agreement statistics. It begins by examining the impact of category prevalence on the level of agreement among raters. The paper then defines and estimates the category prevalence-agreement effect, using this to adjust the observed proportion agreement. Building on this adjustment, a new coefficient of agreement for multiple raters or methods on categorical scales is proposed. In addition, the paper discusses the interpretation and significance of this new coefficient, along with its sampling stability, including the estimation of its standard error.
Coefficient Fleiss’s Kappa
Let
To facilitate the calculation of the agreement coefficient, we first count how many times each specific category is assigned by all raters for each subject. This count is defined as
In Fleiss’s Kappa, the overall observed proportion agreement is calculated as the average agreement among
Hence, the overall observed proportion agreement is the proportion of agreeing pairs of assignments out of the total
where
Landis and Koch (1977) provided interpretation guidelines for Fleiss’s Kappa and other Kappa-like agreement coefficients.
Gwet’s AC1 coefficient (2001) utilizes a computation of the observed agreement that is similar to that used in Fleiss’s Kappa. However, it differs in how it estimates the expected agreement by chance. Unlike Kappa, which assumes random assignment of categories based on marginal probabilities, Gwet’s AC1 adjusts for chance agreement using an alternative model that reduces sensitivity to category prevalence and marginal distribution imbalances. This results in a more stable and often more robust estimate of chance-corrected interrater agreement, particularly in cases of highly skewed or unbalanced data. The observed proportion agreement and the correction for chance agreement are defined by Gwet’s AC1 as
Krippendorff’s α statistic (Krippendorff, 2004) is widely used in the field of interrater agreement, particularly in content analysis and related areas where subjective judgment is involved. It accommodates various data types—including nominal, ordinal, interval, and ratio scales—and allows for missing data, making it a versatile and robust measure for assessing the reliability of ratings across multiple coders. The proportion of observed agreement is defined differently as
where
where
All of these interrater agreement statistics are designed to assess agreement for categorical data, and their values typically range from –1 to 1. A value of 1 denotes perfect agreement, a value of 0 reflects agreement no better than chance, and negative values indicate agreement worse than chance.
A New Coefficient of Agreement
The new agreement coefficient calculates observed proportion agreement similarly to Fleiss’s Kappa but introduces a fundamentally different correction mechanism. Whereas Fleiss’s Kappa and many traditional agreement coefficients estimate expected agreement under the assumption that raters assign categories independently and randomly (Fleiss, 1971), the new coefficient explicitly accounts for the systematic influence of category prevalence on observed agreement. Rather than attributing all agreement beyond chance to randomness, this method recognizes that variations in agreement may arise from how raters interpret and apply the categories—especially when those categories differ in clarity, salience, or conceptual distinctiveness.
Category prevalence can reflect more than just rater tendencies; it may also stem from the intrinsic features of the categories themselves, such as clarity of definition, social desirability, or relevance to the subject matter. For example, clearly defined categories are more likely to produce consistent interpretation and higher agreement, while vague or ambiguous categories often reduce agreement due to variability in how they are understood. In addition, the number of categories influences agreement: raters tend to agree more when using fewer categories (e.g., two vs. three or more). By correcting for the effect of category prevalence—not by assuming randomness, but by acknowledging the structured influence of interpretive clarity—the proposed coefficient aims to offer a more valid reflection of interrater agreement in real-world settings.
The quantity
Importantly, high prevalence and high agreement can co-occur not because raters are influenced by each other or by prior category usage (which is typically unknown to them), but because certain categories are more universally understood or more readily mapped onto shared conceptual frameworks. Agreement can be caused by category prevalence when raters appear to agree often, not because they truly interpret their rating the same way, but because one category is so common that they both happen to pick it frequently. This inflates or deflates the observed agreement and can mask the true level of interrater reliability. In this sense, the prevalence-agreement effect reflects a structured, non-random relationship between category properties and rater behavior. Recognizing this effect helps justify the need for agreement measures that adjust for prevalence in a way that acknowledges these systematic influences, rather than attributing them solely to chance.
To quantify this, the prevalence-agreement effect (
When
Hence, it is expected that the extent of true agreement among multiple raters will be higher for categories with greater prevalence, while categories with lower prevalence are likely to produce less true agreement. For each category
The lower limit of
The lower limit of
Conversely, the lower limit of
The upper limit of
When
When
When
The degree of agreement actually attained—free from the influence of category prevalence—is defined as the true proportion agreement, denoted by
Interrater agreement can then be estimated by normalizing the true proportion agreement
The coefficient
When coefficient
Interpretation of Coefficient Lambda
For any number of raters and categories, coefficient
The lower limit of
where
The value
Values of
Moreover,
The chance agreement defined by Fleiss’s Kappa can be reinterpreted within the framework of coefficient Lambda. Since Fleiss’s κ = 0 corresponds to λ =
Table 1 summarizes the interpretation framework for coefficient Lambda. These categories offer a qualitative framework for interpreting the level of agreement or disagreement between raters based on
Framework for Interpreting Coefficient
Standard Error of
Using the estimators of
Hence,
Here, V stands for variance and Cov is for covariance. The three terms on equation (11) are estimated as
where
Under the null hypothesis that the population agreement coefficient
This allows for hypothesis testing and the construction of confidence intervals for
Numerical Examples
Table 2 presents an agreement matrix for a hypothetical example in which three coders (
Agreement Matrix for Three Raters With Five Categories
The data exhibit extreme category prevalence, with category A dominating—having a prevalence of 1.643, while the other categories (B through E) show very low or zero prevalence values (0, .5, .643, .214, respectively). This imbalance in category usage produces a prevalence-agreement effect of
The estimated standard error of
Using the dataset reported in Fleiss (1971) to assess the consistency of psychiatric diagnoses, each of 30 patients was evaluated by six psychiatrists, selected from an original pool of 43. This represents a standardized adjustment from the original study by Sandifer et al. (1968), where patients had been rated by between six and ten psychiatrists. The standardization to six raters per patient ensured uniformity in evaluating interrater agreement. Under this setup, the observed proportion agreement among the psychiatrists was .556, while the expected agreement by chance was calculated at .220, resulting in a Fleiss’s Kappa of .430—commonly interpreted as moderate agreement. However, a prevalence-agreement effect of 0.064 was also identified, indicating that the observed agreement slightly overestimated the true agreement by 0.064. This inflation is attributed to variations in category prevalence across the five diagnostic categories, whose prevalence values were reported as 0.867, 0.867, 1.000, 1.833, and 1.433, respectively. After adjusting for the prevalence-agreement effect, the resulting coefficient
Table 3 presents a hypothetical example of an agreement matrix in which four clinical psychologists (
Agreement Matrix for Four Raters With Six Categories and Estimates of Interrater Agreement
However, the data exhibit variation in category usage, with prevalence values across the six categories of .60, .80, .92, .80, .48, and .40, respectively. This distribution results in a prevalence-agreement effect of
The standard error of
Number of Raters and True Agreement
When the observed proportion agreement is fixed across conditions, the true proportion agreement and interrater agreement coefficients should be smaller for more raters. This is because achieving the same level of observed agreement with more raters generally requires stronger agreement patterns, leading to a relative decrease in true agreement. Among common indices, the coefficient Lambda uniquely reflects this principle.
This can be understood by considering the different agreement patterns that emerge with varying numbers of raters. With two raters, there is only one possible agreement pattern (
Consider three agreement matrices (Table 4) constructed with an equal observed proportion agreement (
Interrater Agreement Coefficients for Two and Three Raters With Equal Observed Proportion Agreement
Although all three matrices yielded the same observed proportion agreement (
Monte Carlo Simulation
Scope
A Monte Carlo simulation study was conducted to evaluate coefficient Lambda and compare it with three interrater agreement coefficients, Fleiss’s Kappa, Gwet’s AC1, and Krippendorff’s Alpha (α). The simulation manipulated three types of agreement patterns: balanced agreement, where categories had equal prevalence rates, and three levels of imbalanced agreement, where category prevalence rates differed (moderate prevalent and large prevalent). For each agreement pattern, data were generated under combinations of the following simulation factors: seven levels of true interrater agreement (those corresponds reliability values of 0, 0.1, 0.2, 0.3, 0.4, 0.5, and .9), three numbers of rater (
Simulation Procedure
The Underlying Variable Approach (UVA; Muthén, 1984) was employed to simulate ordinal rating data for interrater agreement analysis. This method posits that observed discrete ratings arise from latent continuous variables following a standard multivariate normal distribution. The following steps were followed:
Step 1: Rater assessments were modeled as latent variables drawn from a multivariate standard normal distribution (M = 0, SD = 1) and reliability levels (
Step 2: Prevalence patterns were induced by shifting the mean of the latent distributions: M = 0 (symmetric category usage) for balanced patterns, M = 1.0 (skewed toward higher categories) for moderate prevalence, and M = 1.5 (strong skew toward higher categories) for large prevalence.
Step 3: For each agreement pattern, observed scores (
Analyses and Evaluation Criteria
A total of 336 unique data set combinations were generated out of 378 combinations by systematically varying five factors: number of raters (3 levels), true agreement levels (7 levels), prevalent agreement patterns (3 levels), number of categories (2 levels), and sample size (3 levels). It should be noted that no data sets (42 sets) were generated for a sample size of 50 under the third prevalent agreement pattern owing to incompatibility in the data generation process. Each combination was replicated 5,000 times using Monte Carlo simulations. The following analyses and evaluation criteria were then applied to summarize the results:
Estimates and standard errors were computed using the corresponding formulas for coefficient Lambda and for each of the three interrater agreement coefficients across all generated data sets (5,000 × 336). These values were then averaged across the 5,000 replications within each of the 336 conditions. The average of the estimates for each coefficient represents the mean of its sampling distribution and was used to examine how each interrater agreement coefficient responds to the five factors and their interactions. Similarly, the average of the estimated standard errors for each coefficient represents the estimated standard deviation (via the direct formulas) of its sampling distribution, enabling an assessment of how the standard errors are affected by the five factors and their interactions.
The standard deviation of the estimates of coefficient Lambda, as well as those of the other interrater agreement coefficients, was computed across 5,000 replications for each condition using:
The third analysis aimed to evaluate the sensitivity and specificity of coefficient Lambda and to compare these properties with other interrater agreement coefficients across different conditions of the five factors. Specificity was assessed by examining the Type I error rate, which is the probability of rejecting the null hypothesis of interrater chance agreement when it is actually true. A coefficient is regarded as specific if it maintains a Type I error rate below the statistical significance threshold (0.05 in this study). For coefficient Lambda, the null hypothesis of a chance agreement is H0: λ=1/R, whereas for the other coefficients it is H0:
Results
Estimates of Interrater Agreement
Figure 1 presents the average estimates of coefficient Lambda alongside three other interrater agreement coefficients—Fleiss’s Kappa, Gwet’s AC1, and Krippendorff’s Alpha—calculated over 5000 replications under all conditions defined by the five factors. Rows correspond to sample sizes (N = 50, 100, 300), and columns to the number of categories (J = 3 on the left, J = 5 on the right). The x-axis orders conditions hierarchically by number of raters (2, 4, 6), prevalence levels (balanced, moderate, strong), and three true agreement levels (0, 0.5, 0.9). Only three agreement levels are shown to save space. The results for the other four levels were found to be similar. It is important to note that the observed proportion agreements were fixed across different numbers of raters, for all conditions for other factors.

Estimates of interrater agreement for Lambda and other measures for varying number of raters, number of categories, prevalence levels, true agreement levels, and sample size.
Coefficient Lambda and other coefficients showed higher estimates of interrater agreement as true agreement levels increased at all conditions of other factors. Coefficient Lambda and other coefficients were not influenced by sample size as they produced comparable average estimates at the three sample sizes for different conditions for other factors. Also, coefficient Lambda and other coefficients showed similar average estimates when moving from J=3 to J=5 for different conditions for other factors.
Coefficient Lambda exhibited a consistent decrease in average estimates of interrater agreement as the number of raters increased (R = 2, 4, 6) across all conditions. In contrast, the other three interrater agreement coefficients yielded roughly equal average estimates with increasing numbers of raters. This aligns with the earlier discussion on the relationship between interrater agreement and the number of raters when the observed proportion agreement is held constant across varying numbers of raters. Unlike the other coefficients, which fluctuated upward or downward but ultimately averaged out to equality across rater conditions, coefficient Lambda clearly reflects the expected decrease in agreement as the number of raters increases.
With respect to prevalence patterns, the observed proportion agreement increased as category prevalence became stronger, across all conditions of the other factors. When certain categories are more common (i.e., imbalanced prevalence), raters are more likely to agree because they are more frequently assigning the same common category. This naturally boosts the observed proportion agreement, even though the agreement might not reflect true consistency in judgment.
Similarly, the AC1 coefficient showed increasing average estimates of interrater agreement with higher category prevalence. In contrast, coefficient Lambda, Fleiss’s Kappa, and Krippendorff’s Alpha maintained stable average estimates regardless of prevalence levels. Notably, coefficient Lambda consistently yielded larger average estimates of interrater agreement than Fleiss’s Kappa and Krippendorff’s Alpha across all prevalence patterns, particularly when the number of raters was small.
Comparing the values from the interrater agreement measures to the observed proportion agreement reveals distinct patterns. Under balanced agreement conditions, coefficient Lambda produced values that were sometimes higher and sometimes lower than the observed proportion agreement, while under imbalanced conditions, it consistently yielded slightly lower values. In contrast, Fleiss’s Kappa and Krippendorff’s Alpha consistently produced lower values than the observed proportion agreement across both balanced and imbalanced conditions, with the difference being more pronounced under imbalanced conditions. Gwet’s AC1 showed a mixed pattern: it yielded slightly lower values than the observed proportion agreement under balanced conditions, but much higher values under imbalanced conditions, closely approximating the observed proportion agreement.
Under imbalanced agreement conditions, coefficient Lambda, Fleiss’s Kappa, and Krippendorff’s Alpha all yielded lower values compared to their counterparts under balanced agreement conditions across all true levels of interrater agreement. Among these, coefficient Lambda showed the smallest decrease, while Fleiss’s Kappa and Krippendorff’s Alpha exhibited more pronounced reductions. In contrast, under imbalanced agreement conditions, Gwet’s AC1 produced significantly higher values compared to its counterparts under balanced conditions and closely approximated the observed proportion agreement across all levels of true interrater agreement.
Under chance agreement (ρ = 0), coefficient Lambda took a value of 0.5 with
Standard Error and Its Bias
Figure 2 presents the average standard errors of coefficient Lambda alongside three other interrater agreement coefficients—Fleiss’s Kappa, Gwet’s AC1, and Krippendorff’s Alpha—calculated over 5,000 replications under all conditions defined by the five factors. The organization of Figure 2 follows the same structure as Figure 1. Results indicate that the standard error of coefficient Lambda is consistently smaller than that of any other coefficient, reflecting its relatively narrower scale (0 to 1) across all conditions. The average standard errors of Fleiss’s Kappa and Krippendorff’s Alpha were nearly identical, whereas Gwet’s AC1 generally produced smaller standard errors than the other coefficients, including Lambda, as prevalence patterns increased.

Standard errors for interrater agreement for Lambda and other measures for varying number of raters, number of categories, prevalence levels, true agreement levels, and sample size.
Across all coefficients, results showed that the average standard errors systematically decreased with larger sample sizes, more categories, and greater numbers of raters. Conversely, the average standard errors of Lambda and the other coefficients—except Gwet’s AC1—increased with higher true agreement levels and stronger prevalence patterns.
Figure 3 illustrates the bias of the standard errors for coefficient Lambda alongside three other interrater agreement coefficients—Fleiss’s Kappa, Gwet’s AC1, and Krippendorff’s Alpha—across all conditions defined by the five factors. Bias was calculated as the difference between the average estimated standard errors over 500 replications and the empirical standard error of the interrater agreement coefficient estimates. The layout of Figure 3 mirrors that of Figure 1. Results indicate that the bias of the standard errors was generally small for all interrater agreement coefficients, with coefficient Lambda consistently showing the smallest bias across all conditions. Under stronger prevalence patterns, Gwet’s AC1 also demonstrated relatively small bias. Furthermore, the bias of standard errors systematically decreased as sample size increased. For the remaining factors, no clear pattern in the bias of standard errors was observed for any of the coefficients.

Bias of standard errors for interrater agreement for Lambda and other measures for varying number of raters, number of categories, prevalence levels, true agreement levels, and sample size.
Specificity and Sensitivity
Figure 4 presents the sensitivity and specificity of coefficient Lambda alongside three other interrater agreement coefficients—Fleiss’s Kappa, Gwet’s AC1, and Krippendorff’s Alpha—based on 5,000 replications across all conditions defined by the five factors. The layout of Figure 4 follows the same structure as Figure 1. Coefficient Lambda demonstrated acceptable specificity rates (below 0.05) across all conditions involving variations in the number of categories, raters, sample sizes, and prevalence patterns. In terms of sensitivity, Lambda was responsive to changes in true agreement levels, showing gradually increasing power as true agreement increased across all conditions. At each true agreement level, the sensitivity of Lambda decreased when moving from balanced agreement patterns to stronger prevalence patterns, increased when the number of categories rose from three (J = 3) to five (J = 5), increased as the number of raters rose from two (R = 2) to four (R = 4) and six (R = 6), and increased as sample size enlarged from 50 (N=50), to 100 (N=100) and 300 (N=300). Similar results were found with Cohen Kappa for two raters by Cousineau and Laurencelle (2017).

Sensitivity and Specificity and for interrater agreement for Lambda and other measures for varying number of raters, number of categories, prevalence levels, true agreement levels, and sample size.
Fleiss’s Kappa and Krippendorff’s Alpha exhibited trends in sensitivity and specificity that were largely consistent with those of coefficient Lambda across all conditions involving variations in the number of categories, raters, sample sizes, and prevalence patterns. Similar results about Cohen Kappa were also reported by Cousineau and Laurencelle (2017). In contrast, Gwet’s AC1 showed trends similar to Lambda only under the balanced agreement pattern. However, as prevalence patterns became stronger, AC1 lost specificity, yielding inflated Type I error rates above 0.05. This inflation distorted sensitivity estimates, which reached values of 1.0 even at low levels of true agreement.
Conclusion and Discussion
This paper introduces coefficient Lambda (λ) as a novel measure of interrater agreement applicable to multiple raters or methods. It is defined as the ratio of the observed proportion agreement to the overall proportion agreement, with both quantities adjusted for the prevalence-agreement effect. Unlike traditional chance-corrected coefficients, Lambda is grounded in the assumption that a portion of the observed agreement is systematically influenced by variations in category prevalence—not random chance. These prevalence variations may result from differences in category ambiguity, clarity, attractiveness, social desirability, the number of available categories, or other factors linked to the nature of the measured construct and the context in which the ratings occur. By accounting for this prevalence-agreement effect, coefficient Lambda offers a more realistic and interpretable estimate of the true agreement among raters. Coefficient Lambda estimates the extent of interrater true agreement among raters with relative to disagreement. Also, it gives how much interrater agreement among raters represents from the total agreement after correcting both for category prevalence.
Coefficient Lambda addresses the scaling and interpretation of Fleiss’s Kappa and other Kappa-like coefficients including Krippendorff’s Alpha and Gwet’s AC1, which complicates interpretation across different numbers of raters. A well-known concern with Fleiss’s Kappa is that its scale is not fixed between –1 and 1, but rather varies depending on the number of raters
This non-constant scaling creates interpretative challenges: the same Kappa value can imply different levels of agreement depending on the number of raters. Although some researchers suggest disregarding negative Kappa values, these negative estimates are not rare and should not be ignored. Negative values are generally interpreted as agreement worse than chance, which may signal issues in rater training, scale design, or the clarity of categories being rated. Therefore, researchers must exercise caution when interpreting negative Kappa values and consider the study context, including the construct being measured, the characteristics of the rating scale, and the raters’ expertise. According to guidelines from Altman (1991) and McHugh (2012), Kappa values below .60 indicate inadequate agreement, and limited confidence should be placed in study findings based on such results.
Coefficient Fleiss’s Kappa has been widely criticized in the literature for consistently yielding values that are lower than the observed proportion agreement (
Lambda, Fleiss’s Kappa, and Krippendorff’s Alpha all produced lower values in imbalanced conditions compared to balanced ones. Fleiss’s Kappa and Krippendorff’s Alpha correct for chance agreement, and when category prevalence is skewed, the chance agreement increases. As a result, these coefficients discount more of the agreement as being due to chance, leading to lower agreement scores. Among these, Lambda showed the smallest decrease, suggesting it may be less sensitive to category imbalance than Kappa or Alpha.
It is important to note that a coefficient Lambda value lower than the observed proportion agreement (
As the number of raters increased from two to four to six, it is noted that coefficient Lambda showed decreasing estimates of interrater agreement; whereas Fleiss’s Kappa, Krippendorff’s Alpha, and Gwet’s AC1 did not exhibit this pattern. This trend held true across different numbers of categories and varying true levels of agreement. From a conceptual standpoint, interrater agreement tends to decrease as the number of raters increases. This is because achieving consensus among a larger group of individuals is inherently more challenging than among a smaller group. With fewer raters, the likelihood of aligning on the same category or judgment is higher due to less variability in perspectives, interpretations, and decision-making criteria. However, as more raters are introduced, the diversity of viewpoints and potential for disagreement naturally increase, making it more difficult to reach a high level of uniformity in ratings. This phenomenon is particularly evident in complex or subjective tasks, where different raters may interpret criteria or rating scales differently. As a result, interrater agreement measures often reflect lower values when more raters are involved—not necessarily because the quality of the raters has declined, but because the probability of unanimous or near-unanimous agreement decreases with group size.
The proposed coefficient Lambda differs from traditional agreement coefficients in several important ways. First, its values range from 0 to 1, and it does not produce negative values, avoiding the interpretative issues associated with negative Kappa-type coefficients. Second, Lambda maintains a fixed and consistent scale, regardless of the number of raters or categories involved in the agreement study. Third, it provides a uniform interpretation of agreement levels across different study designs and contexts, making it easier to compare findings across diverse applications. These properties make coefficient Lambda a more practical and intuitive tool for measuring interrater agreement, particularly for researchers and practitioners across various scientific disciplines who seek a stable, interpretable, and meaningful metric of agreement.
The interpretation of coefficient Lambda is based on comparing the true proportion of agreement—defined as
In addition, the value of chance agreement for Lambda is given by
In our analysis, we found that the new agreement coefficient is more sensitive to variations in category prevalence than Fleiss’s Kappa and other Kappa-like statistics. This behavior reflects the design of the new coefficient, which explicitly adjusts for the prevalence-agreement interaction, accounting for how certain categories may elicit higher or lower agreement due to clarity, frequency, or social salience. In contrast, Fleiss’s Kappa and other Kappa-like statistics, which adjust for expected agreement purely through marginal probabilities, does not fully capture this interaction and remains less responsive to structural imbalance in the data.
Moreover, coefficient Lambda demonstrated acceptable levels of sensitivity and specificity to chance agreement, comparable to those of Fleiss’s Kappa and Krippendorff’s Alpha, across varying conditions of true agreement levels, number of raters, number of categories, sample size, and prevalence patterns. Future research is recommended to further investigate the sensitivity and specificity of Lambda at different population values, as well as to develop hypothesis-testing procedures for comparing Lambda values obtained from independent samples (e.g., males vs. females) and paired samples (e.g., across time points).
Future research on coefficient Lambda should further investigate its statistical properties across diverse rating conditions, including varying numbers of raters, categories, and prevalence distributions. In particular, studies are needed to examine its robustness in unbalanced settings and its behavior relative to other agreement coefficients when disagreement is systematically structured. Simulation and empirical work could also clarify how Lambda performs in large-scale assessments. Such investigations would contribute to establishing Lambda as a more comprehensive and reliable tool for evaluating interrater agreement.
Footnotes
Appendix
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article. Material preparation, data collection, and analysis were performed by me.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
Ethical Approval and Informed Consent Statements
The manuscript did not require ethics approval since it is a methodology paper that does not have direct contact with human or animal subjects. The manuscript did not require consent to participate since it is a methodology paper that does not involve direct contact with human subjects.
Data Availability Statement
There is no data to share since the study does not require data collection.
