Abstract
In counseling research, reliability is an extremely important concept. Reliability of scores refers to how consistent scores remain across time, instruments, and conditions. The most commonly reported methods of assessing reliability are Cronbach’s alpha and Kuder-Richardson's index of reliability, otherwise known as measures of internal consistency. The purpose of this article is to describe internal consistency reliability coefficients, delineate the history of these coefficients, depict contemporary issues in counseling research, and present recommendations for practice.
Keywords
The concept of reliability in counseling research is important since its evidence is essential, although not sufficient, to establishing evidence of validity. Reliability refers to the degree to which scores yielded by an instrument administered to specific individuals, at a specific point in time, and under certain conditions, are replicable (Hopkins, 1998). In order for scores to be replicable, they must be consistent. Therefore, the most popular definition of reliability is that reliability represents the consistency with which scores measure behaviors or constructs across variations in instrument items, forms of the instrument, occasions of measurement, or other measurement conditions (Vogt, 1999). Although reliability most often is associated with scores yielded by tests or examinations, reliability can index any set of scores generated by any type of measure and across any domain (e.g., cognitive, affective, psychomotor).
Computationally, reliability can be defined as the ratio obtained when the true score variation is divided by the observed score variation (Crocker & Algina, 1986). As such, the reliability coefficient is essentially an effect size measure. Unreliability occurs as a result of measurement errors, which can be either random or systematic (Onwuegbuzie & Daniel, 2004). Random measurement errors are the result of chance occurrences that stem from factors such as variations in the administration of the instrument, fluctuations in the respondent’s mental or psychological state (e.g., levels of anxiety), guessing, socially desirability in responses, satisficing (i.e., the tendency to agree with an item due to the exertion of minimal cognitive effort), acquiescence (i.e., the tendency of participants simply to agree with survey items regardless of content), response set (i.e., the tendency of participants to respond to general feelings about the instrument topic rather than the specific item content), and scoring errors (Onwuegbuzie & Weems, 2004; Reinhardt, 1996; Weems & Onwuegbuzie, 2001; Weems, Onwuegbuzie, & Lustig, 2003; Weems, Onwuegbuzie, Schreiber, & Eggers, 2003). Systematic errors represent errors that consistently affect respondents' scores because of a particular characteristic of an individual respondent, the group of respondents, or the instrument that is independent of the underlying construct. Whereas random errors reduce both the consistency (i.e., reliability) and the utility of scores, by either increasing or attenuating any respondent’s score in an unpredictable manner, systematic errors adversely affect the practical usefulness of the scores. Thus, reliability provides information about the errors of measurement.
In a theoretical sense, reliability is a unitary characteristic of scores (Crocker & Algina, 1986). Consequently, it is somewhat misleading to refer to “types of reliability.” Notwithstanding, there are various ways to estimate reliability. However, by far the most popular method of assessing reliability across all disciplines in the social and behavioral science field and beyond is estimating the reliability of scores on the basis of alternate configurations of the items across a single administration of the instrument (Henson, 2001). The reliability estimate that is computed using this technique routinely is called the coefficient of internal consistency (Cronbach, 1951; Hogan, Benjamin, & Brezinski, 2000) or internal consistency reliability coefficient (Henson, 2001). The widespread appeal of the coefficient of internal consistency among quantitative researchers stems from the fact that it is derived from scores resulting from one administration of a single measure, unlike many other estimates of reliability that necessitate either two or more administrations (i.e., test–retest reliability index, or coefficient of stability), two or more instruments (i.e., alternate forms reliability index, or coefficient of equivalence), both two or more administrations and two or more instruments (i.e., coefficient of stability and equivalence), two or more scorers (i.e., interrater reliability index), or two or more ratings of the same set of observations by the same scorer (i.e., intrarater reliability index). Further, coefficients of internal consistency often are appropriate even when another approach to estimating reliability (e.g., coefficient of stability) is the major focus of a given reliability study.
History of the Internal Consistency Reliability Coefficient
The concept of reliability can be traced back to Spearman’s (1904) seminal article. Spearman delineated in this article that the absolute value of the correlation coefficient between the measurements for any pair of variables must be attenuated when the measurement for either or both variables are affected by what he called accidental variation (i.e., random variation) than otherwise would be the case. Thus, Spearman derived a formula for correcting for the attenuating effect of accidental variation.
Motivated by Pearson’s criticism of his 1904 article, in 1907, Spearman published a proof of his correction-for-attenuation formula. Spearman (1910) and Brown (1910) independently derived the same formula by which to calculate a reliability coefficient from the two halves of a single composite measure. In their formula, Spearman and Brown assumed that the two halves had identical observed score means and variances—what is contemporarily known as being classically parallel. This formula was named “Spearman-Brown formula.” However, despite its popularity, arguments prevailed regarding how reliability should be defined. In particular, whereas Brown defined the reliability coefficient as the correlation between scores on repeated administration of the same instrument, Kelley (1921) defined the reliability coefficient as the coefficient of correlation between comparable tests. Meanwhile, Abelson (1911) used Spearman’s (1910) proof to derive a formula for what was coined by Kelley 5 years later as the index of reliability.
Since 1910, a variety of coefficients have been developed to measure internal consistency. The earliest internal consistency coefficients were based on correlations derived from scores from dichotomous splits of a test into approximately parallel subsets of items, followed by various correction formulae to account for attenuation problems due to using scores on the shortened versions of the test to estimate the reliability of the scores on the total test (Brennan, 2001). Kuder and Richardson (1937) and Cronbach (1951) improved upon these original formulae by developing formulae that averaged out the effects of measurement error due to any particular split-half coefficient. Specifically, Kuder and Richardson developed two reliability indices for dichotomously scored items that were called the Kuder-Richardson 20 formula and Kuder-Richardson 21 formula, or KR-20 and KR-21, respectively, with the latter being derived under the additional assumption that all items on a test are of equal difficulty. Thus, the KR-20 and KR-21 reliability estimates are identical only when all items are equal in difficulty; otherwise, the KR-21 reliability coefficient will be systematically lower than is the KR-20 estimate.
Cronbach’s formula, which is a generalization of KR-20, was developed for items scored on any scale, dichotomous or otherwise, which he called coefficient alpha. KR-20, KR21, and coefficient alpha are based on the assumption that scores on each subset of items from an instrument are perfectly parallel to the scores on any other subset of items from the same test. Alternatively stated, this assumption requires that scores on the part tests used in deriving the estimates are essentially tau equivalent (i.e., forms may have different means and error variances), which is weaker than the classical assumption of parallel forms.
Hoyt (1941) developed an approach to internal consistency using an analysis of variance, wherein an internal consistency reliability coefficient is derived from the variances and covariances among scores. Hoyt’s procedure served as the foundation for generalizability theory, which uses concepts from experimental design to examine the degree to which a series of measurement conditions act individually and in interaction with each another to influence the reliability of scores.
The most recent major development of internal consistency reliability estimates has been provided by Vacha-Haase (1998). Vacha-Haase proposed a method she called reliability generalization (RG), which represents an examination of the variability of score reliability across multiple studies, wherein coded study characteristics (e.g., sample demographics) are used to predict variability in score reliability across studies, in an attempt to determine which sampling conditions most affect internal consistency reliability estimates.
Today, KR-20 and Cronbach’s coefficient alpha are by far the most commonly used indices to determine internal consistency reliability estimates for scores from instruments that are dichotomously scored and that have a specific number of fixed response options, respectively (Henson, 2001). Indeed, Hogan et al. (2000) found that approximately 75% of reported score reliability estimates in the Directory of Unpublished Experimental Mental Measures—published by the American Psychological Association (APA)—represented internal consistency estimates. Both internal consistency formulae (i.e., KR-20 and Cronbach’s coefficient alpha) take the following form:
Contemporary Issues
Currently, there are six major errors made by researchers associated with internal consistency reliability estimates (Onwuegbuzie & Daniel, 2002, 2004). First, many researchers fail to recognize that an internal consistency reliability estimate cannot be generalized to approximate the value of a different type of reliability estimate (Onwuegbuzie & Daniel, 2002). For example, the Dyadic Adjustment Scale (DAS; Spanier, 1976) has four subscales and therefore internal consistency coefficients ranging from .70 to .95. Yet, the stability reliability coefficients ranged from .75 to .87, a much lower and smaller range than the internal consistency coefficients. If one were to state that the stability reliability of the data for the DAS was .95 (using the reported coefficient for the internal consistency), this would be incorrect. Indeed, as shown through this example, internal consistency coefficients overestimate reliability generalizing over occasions due to the contrived replications.
Second, many researchers incorrectly interpret internal consistency estimates, referring to them as representing the reliability of the instrument (Thompson & Vacha-Haase, 2000; Vacha-Haase, 1998). Yet reliability is a property of scores, not of instruments. As stated by Thompson and Synder (1998) in their review of articles published by the Journal of Counseling and Development: Some authors explicitly invoked language asserting that tests are reliable. Examples included the following: “the BES is highly reliable” (Kaminski & McNamara, Vol. 74, p. 289); “Cronbach’s alpha for the DAS is .96” (Conteras, Hendrick, & Hendrick, Vol. 74, p. 410); “weak reliabilities of the Preencounter and Encounter subscales” (Carter & Parks, Vol. 74, p. 488); “reliabilities of the TRIG subscales” (Brown, Richards, & Wilson, Vol. 74, p. 506); and “may be a reliable and valid measure” (Melchert, Hays, Wiljanen, & Kolocek, Vol. 74, p. 642). (p. 439)
Third, and likely stemming from the preceding error, a large proportion of researchers do not report internal consistency estimates (or any other reliability estimates) for data from their own samples, despite strong recommendations to do so by reputable bodies such as the APA Task Force on Statistical Inferences (Wilkinson & Task Force on Statistical Inference, 1999) and authors of the sixth edition of the Publication Manual of the APA (i.e., “Provide information on instruments used, including their psychometric and biometric properties;” American Psychological Association [APA], 2010, p. 31). These recommendations come from decades of researchers not reporting internal consistency estimates for their data at hand. For example, Wilson (1980) found that “Only 37% of the AERJ studies explicitly reported reliability coefficients for the data analyzed … That reliability … is unreported in almost half the published research is … inexcusable at this late date” (pp. 8–9); and yet the “late date” was over two decades ago! Ten years later, Meier and Davis (1990) analyzed three volumes (1976, 1977, and 1987) of the Journal of Counseling Psychology and found that most authors did not report reliability of the data collected. More recently, Thompson and Snyder (1998) analyzed one volume of the Journal of Counseling & Development and found that of the 25 articles published in the volume, 13 reported score reliability from previous studies, 8 reported reliability from both previous studies and the current study, and one reported reliability from only the current study. Without information about score reliability, it is impossible to assess the extent to which threats to internal validity unduly affected both the power of statistical significant tests and the associated effect sizes from the current study.
Fourth, of the relatively few researchers who report internal consistency estimates for their own data, some make the mistake of testing these coefficients for statistical significance using the nil null hypothesis (Daniel, 1998; Onwuegbuzie & Daniel, 2002; Witta & Daniel, 1998). For example, Morrow and Jackson (1993) found examples of inappropriate reporting of statistical significance and reliability estimates: Significance testing of the reported reliability was conducted by Summers, Miller, and Ford (1991), who reported reliabilities (both test-retest and internal consistency estimates) between .24 and .91 … reportedly “significant at the .01 level” (p. 244) … Significant testing of the reliability … is of little use to the reader … it is inappropriate to simply report the reliability as “significant”. (p. 352)
Fifth, when interested in determining the effect of an intervention by comparing scores on the same instrument administered both before and after the intervention phase, some researchers report reliability estimates pertaining only to either the pre-intervention scores or the post-intervention scores (Allen & Yen, 1979). For example, Thrun et al. (2009) investigated HIV care clinics and whether training the providers to counsel their clients regarding prevention of HIV would affect the providers' attitudes, self-efficacy, and counseling behaviors. This longitudinal study included a survey to assess the providers' willingness to discuss prevention, attitudes, self-efficacy, and counseling behavior. These authors reported the “internal consistency reliability (Cronbach’s alpha) of the scales ranged from .74 to .97 … [with] an overall composite score … of .96;” yet they did not report internal consistency estimates for the data each time these data were collected over time. It is more appropriate to estimate the reliability of difference scores, as outlined by Crocker and Algina (1986).
Sixth, and finally, the vast majority of researchers mistakenly assume that the score reliability is invariant across subsamples (Onwuegbuzie & Daniel, 2004). Yet, Onwuegbuzie and Daniel (2004) demonstrated both theoretically and empirically that this is not necessarily the case. Indeed, it is possible for a reliability coefficient from the full sample not only to mask marked differences in subsample reliabilities but also to conceal low score reliabilities generated from scores of one or more of the subgroups. Thus, it is possible to obtain a large reliability estimate for the full sample even when the reliability coefficient of one or more subgroups is unacceptably small. Therefore, for example, Simons, Giorgio, Houston, and Jacobucci (2007) need to include reliability estimates for the overgroup and subgroups. These authors focused their article on the statistically significant differences found between males and females and Caucasian and African American participants, stating “The results further indicated that males have more favorable views … while African-Americans have less favorable attitudes…” (p. 62). Unfortunately, they included only one internal consistency coefficients for the surveys.
Recommendations for Practice
Based on the errors discussed above, we offer the following 12 recommendations for the use of internal consistency reliability estimates in counseling research: Avoid relying on internal consistency estimates derived from induction (i.e., normative) samples as an estimate of reliability of the data at hand. For example, Ashby, Dickinson, Gnilka, and Noble (2011), in their study on hope, perfectionism, and depression among middle school students, appropriately report the internal consistency estimates from past research, as well as from the data at hand: “Kandel and Davies (1986) reported evidence for the validity of the scale and adequate test–retest reliability. Internal consistency reliability for this sample was .80” (p. 133). Always report internal consistency estimates for scores yielded by quantitative instruments, unless the researcher is interpreting “replications” as that arising from an instrument or forms of an instrument administered on different occasions. For example, in their study of an intervention to prevent rape with high school students, Hillenbrand-Gunn, Heppner, Mauch, and Park (2010) appropriately reported the internal consistency estimates along with the test–retest estimates for one of the measures used, “Schewe (2002) reported an alpha coefficient of .87 and test–retest reliability of .73 for the IRMA-SF (Illinois Rape Myth Acceptance-Short Form [Payne, Lonsway, & Fitzgerald, 1999]) based on a sample of 829 males. For the current study, the score alpha coefficients were .83 (male students) and .77 (female students)” (p. 44). Provide adequate supporting data (e.g., sample size and composition, measurement conditions, test form) when reporting internal consistency estimates, so that the estimate might be understood fully. The importance of this practice cannot be overstated. Roberts and Onwuegbuzie (2003) provided cases of how internal consistency reliability estimate can be attenuated by the level of homogeneity of the sample. Furthermore, Vacha-Haase, Kogan, and Thompson (2000) documented that lower score reliability can reflect differences in sample composition. If gain scores are of interest, use information about the variance and reliability of scores on the pretest and posttest to compute an estimate of reliability of the gain scores. For example, Schweisheimer and Walberg (1976), in their experiment of peer counselors in high school, used gain scores, but did not estimate the reliability of the gain scores. These authors state “Sixteen variables were measured both at or near the beginning and at or near the end of the experimental counseling period. The variables and their reliabilities are listed in table 1” (p. 399). The table includes test–retest reliabilities and one internal consistency reliability, but there is not any estimation of the gain scores themselves. Without this information, it is difficult to assess the rigor of the results. Avoid stating that “the test is reliable” or “the test has high internal consistency.” An example that should not be followed is Schönrock-Adema, Van der Molen, and van der Zee (2009), in their study of microcounseling skills training, the authors state, “Several authors have demonstrated the reliability and validity of role-play tests (e.g., Bellack & Hersen, 1988; Smit, 1995; Smit & Van der Molen, 1996)” (p. 248). Instead, these authors could have stated that “several authors have demonstrated the reliability and validity of the data from role-play tests.” Another example comes from Eriksen and McAuliffe (2003) in their article describing the development of the Counseling Skills Scale. These authors state inappropriately, “Even when specific counseling sessions are the focus of analysis and when experts judge counselor behavior, few valid and reliable measures of actual counseling skills exist” (p. 122). Never use statistical significance tests of internal consistency estimates. For example, Brief, Burke, George, Robinson, and Webster (1988) examined whether a personality construct, negative affectivity, would be related to self-report measures of job stress and job strain and whether observed relationships between these stress and strain measures would be inflated markedly by negative affectivity. These researchers developed a table (i.e., Table 1) that contained “Internal consistency reliability coefficients are shown on the diagonal” (p. 196). Referring to all coefficients in this table, these researchers inappropriately stated that they were all “statistical significant at the .01 level” (p. 196). Construct and report confidence intervals (either one-sided or two-sided) around internal consistency estimates to show the effect of error. For example, Helms, Henze, Sass, and Mifsud (2006), in providing strategies for implementing good practices for analyzing, reporting, interpreting, and using reliability data in counseling research, not only demonstrated how to construct 95% two-sided confidence intervals around Cronbach's alpha by hand using real data generated from responses (n = 550) of the White Racial Identity Attitudes Scale (Helms & Carter, 1990), but they also provided citations for authors who present syntax for computing these confidence intervals for the Statistical Package for the Social Sciences (SPSS) and SAS software programs. However, conveniently, confidence intervals around internal consistency estimates easily can be obtained via the most recent versions of SPSS (i.e., from version 16 onward) without using syntax (i.e., by selecting Analyze → Scale → Reliability Analysis → Statistics → Intraclass Correlation Coefficient). In group-comparison studies, disaggregate internal consistency estimates by reporting a score reliability coefficient for each subgroup of interest. Hillenbrand-Gunn, Heppner, Mauch, and Park (2010) appropriately reported the disaggregated internal consistency estimates, by documenting the alpha coefficients for males and females: “For the current study, the alpha coefficients were .83 (male students) and .77 (female students)” (p. 44). Link the value of correlational results in studies back to information about internal consistency of scores on the variables being correlated. For instance, Onwuegbuzie, Roberts, and Daniel (2005) appropriately outlined how displaying disattenuated correlation coefficients alongside their unadjusted correlation coefficients allows the reader to evaluate the impact of unreliability on each bivariate relationship. Further, these methodologists developed what they termed a what if reliability analysis to complement the conventional null hypothesis statistical significance test of bivariate relationships. This what if reliability analysis indicates how the sample size needed to detect a statistically significant bivariate relationship decreases as the observed score reliability coefficient pertaining to the independent and/or dependent measure (theoretically) increases, holding all other factors constant. The what if reliability analysis examples provided in this article “demonstrated how p values and effect sizes have the potential to misrepresent reality when the reliability context is not considered” (p. 238). As noted by Onwuegbuzie et al. (2005), what if reliability analyses “help researchers to interpret their results by considering the extent to which the reliability of scores on one or both variables affects the statistical significance of the bivariate correlation coefficient” (p. 230). Conduct follow-up internal consistency analyses (e.g., examining response patterns, examining sample homogeneity) when these estimates are low, in an attempt to determine possible reasons. For example, Weems and Onwuegbuzie (2001) examined the effect of the arrangement and number of scale step options on reliability estimates. Specifically, they analyzed three different data sets containing counseling-related data (e.g., scores measures of hope, perfectionism, and anxiety) to investigate the effect of (a) using midpoint options on score reliability and score validity, (b) including/excluding midpoint choices on item mean and reliability, and (c) using reverse-code items on scale mean and reliability. In a follow-up study, Onwuegbuzie and Weems (2004) conducted a study to investigate the personality characteristics of respondents who frequently utilize midpoint categories on rating scales. The combined findings of Weems and Onwuegbuzie (2001) and Onwuegbuzie and Weems (2004) indicate that score reliability can be attenuated (a) by providing only a small number of response options, (b) when overselection of midpoint options prevail, (c) when the majority of items invoke an overselection of neutral responses, and (d) when reverse-code (i.e., mixed stem) items are present; and (e) that certain individuals (e.g., those with negative self-perceptions about their levels of creativity, the lowest levels of self-oriented perfectionism) are more likely to select midpoint options. If RG studies have been conducted for the instrument being used, cite these when providing information about the instrument. For example, when using the Myers-Briggs Type Indicator (MBTI; Myers & McCaulley, 1989), authors should cite the Capraro and Capraro (2002) RG study. Conduct RG studies of instruments that are used regularly. There are many published RG studies for instruments that are commonly used in counseling research. For example, Yin and Fan (2000) assessed the reliability generalization (RG) of the Beck Depression Inventory (BDI), Wallace and Wheeler (2002) assessed the RG for the Life Satisfaction Index, and Vacha-Haase, Kogan, Tani, and Woodall (2001) assessed the Minnesota Multiphasic Personality Inventory (MMPI). This is only a sample of RG studies that have been published; as more instruments are routinely utilized in counseling studies, further RG studies should be conducted.
Finally, internal consistency reliability coefficients are not direct measures of reliability; rather, they are theoretical estimates that are based on classical test theory. Therefore, it is important to discuss the reliability values as being the true value of the reliability for the data. Moreover, internal consistency estimates are limited to the extent that they only provide information about reliability in terms of consistency of scores across a given set of items; they do not take into consideration the occasion of measurement, form of the test, or other important measurement conditions that may be addressed via other approaches for estimating score reliability. Despite these limitations, when used appropriately, internal consistency reliability coefficients often provide useful information to researchers in counseling.
Footnotes
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The author(s) received no financial support for the research, authorship, and/or publication of this article.
