Seven approaches to averaging reliability coefficients are presented. Each approach starts with a unique definition of the concept of “average,” and no approach is more correct than the others. Six of the approaches are applicable to internal consistency coefficients. The seventh approach is specific to alternate-forms coefficients. Although the approaches generally produce unequal averages, a Monte Carlo study found little difference among the average reliabilities calculated by the first six approaches. The first three approaches may be especially useful for reliability generalization studies.
Alexander, R. A. (1990). A note on averaging correlations. Bulletin of the Psychonomic Society, 28, 335-336.
2.
Charter, R. A. (2001).It is time to bury the Spearman-Brown “prophecy” formulafor some commonapplications. Educational and Psychological Measurement, 61, 690-696.
3.
Charter, R. A. (2003a). A breakdown of reliability coefficients by test type and reliability method, and the clinical implication of low reliability. The Journal of General Psychology, 130, 290-304.
4.
Charter, R. A. (2003b). Combining reliability coefficients: Possible application to meta-analysis and reliability generalization. Psychological Reports, 93, 643-647.
5.
Corey, D. M., Dunlap, W. P., & Burke, M. J. (1998). Averaging correlations: Expected values and bias in combined Pearsonrs and Fisher'sz transformations. The Journalof General Psychology, 125, 245-261.
6.
Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16, 297-334.
7.
Dunlap, W. P., Silver, N. C., & Bittner, A. C. (1986). Estimating reliability with small samples: Increased precision with averaged correlations. Human Factors, 28, 685-690.
8.
Dunlap, W. P., Silver, N. C., & Phelps, G. R. (1987). A Monte Carlo study of using the first eigenvalue for averaging intercorrelations. Educational and Psychological Measurement, 47, 917-923.
9.
Feldt, L. S. (1965). The approximate sampling distribution of Kuder-Richardson reliability coefficient twenty. Psychometrika, 30, 357-370.
10.
Feldt, L. S., & Brennan, R. L. (1989).Reliability. In R. H. Linn (Ed.), Educationalmeasurement (3rd ed., pp. 105-146). New York: Macmillan.
11.
Feldt, L. S., & Charter, R. A. (2003a). Estimating the reliability of a test split into two parts of equal or unequal length. Psychological Methods, 8, 102-109.
12.
Feldt, L. S., & Charter, R. A. (2003b). Estimation of internal consistency reliability when test parts vary in effective length. Measurement and Evaluation in Counseling and Development, 36, 23-27.
13.
Feldt, L. S., Woodruff, D. J., & Salih, F. A. (1987). Statistical inference for coefficient alpha. Applied Psychological Measurement, 11, 93-103.
14.
Gulliksen, H. (1950). Theory of mental tests. New York: John Wiley.
15.
Henson, R. K., & Thompson, B. (2002). Characterizing measurement error in scores across studies: Some recommendations for conducting “reliability generalization” studies. Measurement and Evaluation in Counseling and Development, 35, 113-126.
16.
Hunter, J. E., & Schmidt, F. L. (1990). Methods of meta-analysis: Correcting error and bias in research findings. Newbury Park, CA: Sage.
17.
Kristof, W. (1963). The statistical theory of stepped-up reliability coefficients when a test has been divided into several equivalent parts. Psychometrika, 28, 221-238.
18.
Kristof, W. (1974). Estimation of reliability and true score variance forms split of test into three arbitrary parts. Psychometrika, 39, 207-225.
19.
Marascuilo, L. A., & Serlin, R. C. (1988). Statistical methods for the social and behavioral sciences.New York: Freeman.
20.
Parker, K. C., Hanson, R. K., & Hunsley, J. (1988). MMPI, Rorschach, and WAIS: A meta-analytic comparison of reliability, stability, and validity. Psychological Bulletin, 103, 367-373.
21.
Paulson, E. (1942). An approximate normalization of the analysis of variance distribution. Annals of Mathematical Statistics, 13, 233-235.
22.
Rulon, P. J. (1939). A simplified procedure for determining the reliability of a test by split-halves. Harvard Educational Review, 9, 99-103.
23.
Schmidt, F. L., Hunter, J. E., & Raju, N. S. (1988). Validity generalization and situational specificity: A second look at the 75% rule and Fisher's z transformation. Journal of Applied Psychology, 73, 665-672.
24.
Silver, N. C., & Dunlap, W. P. (1987). Averaging correlation coefficients: Should Fisher's transformation be used?Journal of Applied Psychology, 72, 146-148.
25.
Strube, M. J. (1988). Averaging correlation coefficients: Influence of heterogeneity and set size. Journal of Applied Psychology, 73, 559-568.
26.
Vacha-Haase, T., Henson, R. K., & Caruso, J. C. (2002). Reliability generalization: Moving toward improved understanding and use of score reliability. Educational and Psychological Measurement, 62, 562-569.
27.
Zar, J. H. (1984). Biostatistical analysis (2nd ed.). Englewood Cliffs, NJ: Prentice Hall.