Abstract
Cognitive assessment of young children contributes to high-stakes decisions because results are often used to determine eligibility for early intervention and special education. Previous reviews of cognitive measures for young children highlighted concerns regarding adequacy of standardization samples, steep item gradients, and insufficient floors for young children functioning at lower levels. The present report extends previous reviews by including measures recently published or revised, nonverbal cognitive assessment tools, and issues specific to assessing bilingual or non-English-speaking children. Sixteen tests were reviewed, including all available measures of cognitive functioning for 2- to 4-year-old children normed in the United States. Test characteristics evaluated included (a) representativeness and recency of standardization data, (b) item bias analysis, (c) psychometric characteristics, and (d) appropriateness for assessing young children with developmental delays and non-English-speaking children. Implications are discussed for clinicians, researchers, and test developers.
For very young children, cognitive tests are frequently used as part of an assessment battery to assess developmental delays, make clinical diagnoses, and conduct research regarding developmental disabilities. These tests may have high-stakes consequences when used in eligibility determinations for early intervention, later developmental disabilities services, and special education. However, theoretical and psychometric difficulties have been identified in reliably measuring cognitive functioning in very young children (Neisworth & Bagnato, 2004). These difficulties are magnified in children with developmental delays and in children with limited English proficiency, and can impact outcomes such as access to services, classroom placements, and interpretation of research findings. Therefore, it is critical for professionals and researchers to carefully evaluate the strengths and limitations in the psychometric properties and standardization samples of available cognitive tests.
Previous reviews of cognitive measures for young children (Bradley-Johnson, 2001; Lichtenberger, 2005) highlighted specific problems in technical adequacy including (a) minimal inclusion of children with atypical development in standardization samples (Bradley-Johnson, 2001), (b) steep item gradients leading to large changes in standard scores based on very small changes in raw scores (Bradley-Johnson, 2001), and (c) insufficient floors for young children functioning at lower levels (Lichtenberger, 2005). The importance of these factors was echoed in publications such as The Division for Early Childhood (DEC) of the Council for Exceptional Children’s recommended standards for the assessment of young children (Neisworth & Bagnato, 2000) and the Standards for Educational and Psychological Testing (American Educational Research Association [AERA], American Psychological Association, & National Council on Measurement in Education, 1999). The DEC’s recommendations emphasized the importance of sensitivity, ensuring that tests accurately assess lower levels of functioning and detect small increments of change. Adequate floors (those which are low enough to differentiate the lowest 1% of functioning) and adequate item gradients (sufficient items so that very small changes in raw scores do not lead to large changes in standard scores) are essential in this context. In his review of preschool instruments, Bracken (1987) proposed criteria to establish minimal adequacy of test floors and item gradients.
When examining the standardization sample for a test, it is important to consider the impact of outdated norms on score interpretation. Flynn (1984, 1998, 2010) documented that over time, individuals perform better on IQ tests than those in prior generations; tests become “easier” and scores increase if tests are not re-normed periodically. This “Flynn Effect” has been found to be particularly pronounced in individuals functioning in the lower ranges of intellectual ability (Kanaya, Scullin, & Ceci, 2003; Teasdale & Owen, 2000; Zhou, Zhu, & Weiss, 2010). The specific causes for the Flynn Effect have been debated (Ceci & Kanaya, 2010; Kaufman, 2010). However, experts in the field have recommended consideration of the recency of test standardization when selecting a cognitive assessment tool, particularly for making high-stakes decisions (Fletcher, Stuebing, & Hughes, 2010).
The representativeness of the standardization sample is also critical. While tests are typically evaluated for representativeness of the general population in characteristics such as race or ethnicity, gender, geographic region, and socioeconomic status (SES), the DEC recommendations and the Standards for Educational and Psychological Testing (AERA et al., 1999) also emphasized the need to include children with atypical development in standardization samples. Neisworth and Bagnato (2004) noted that children with developmental delays have often been excluded from standardization samples, thus precluding the development of standard modifications in testing and valid measurement of domains needed to identify intellectual disability. McFadden (1996) pointed out that exclusion of children with atypical development will result in classification errors when comparing a child with atypical development with a normative sample that excluded the lower end of the normal curve.
For children who are bilingual or do not speak English, the problem of reliably measuring cognitive functioning is amplified (Kester & Peña, 2002; Mindt et al., 2008). The Standards for Educational and Psychological Testing (AERA et al., 1999) and the position statement of the National Association for the Education of Young Children (NAEYC; 2005) on Screening and Assessment of Young English-Language Learners noted the importance of using culturally and linguistically appropriate assessment tools, appropriate translations, and demonstration of equivalence or comparability of a translated test to the English-language version. Few assessment and diagnostic tools have been designed specifically for children exposed to two languages (Peña & Halle, 2011). Despite increased attention to appropriate bilingual and cross-cultural assessment practices (Merenda, 2006; Oakland, Illiescu, Chen, & Chen, 2013; Peña, 2007), many assessors have continued to use inappropriate tools and procedures. Professionals and researchers may try to compensate for a lack of measures specifically designed or normed on non-English-speaking populations by making informal adjustments or accommodations. However, informal translation of test items from English to the child’s native language can lead inadvertently to changes in the psychometric properties and score interpretations of the original test (Hambleton & Patsula, 1998).
Inaccurate assessments may lead to academic, behavioral, and social repercussions for bilingual children. Children may be placed in inappropriate classrooms due to mistaken conclusions about their level of functioning. Language disorders may go undetected if communication difficulties are assumed to be due to a lack of English proficiency without assessment of a child’s primary language proficiency. Moreover, teachers’ expectations may be lowered if English proficiency is under-representing the child’s capabilities. Inaccurate testing may contribute to the disproportional placement in special education programs of English-language learners (Artiles, Rueda, Salazar, & Higareda, 2005).
Because early childhood assessments carry such weight, it is essential to critically evaluate currently available cognitive assessment tools to determine the best available measures for different purposes, and to encourage test developers to address identified problems. In this review, the authors examined the strengths and weaknesses of commonly available tools for assessing cognition in children ages 2 to 4 years. We have updated and extended the reviews provided by Bradley-Johnson (2001) and Lichtenberger (2005) by including nonverbal cognitive assessment tools, measures that were published or revised after those reviews, and considering issues specific to the assessment of bilingual children.
The age group of 2- to 4-year-old children was selected because of the high-stakes impact of testing that occurs during the transition from early intervention to other service systems, the relative lack of training among psychologists to assess this age group, and the unique challenges in developing tests with appropriate psychometric properties for young children. Furthermore, the review includes consideration of the appropriateness of available tests for assessment of children from non-English-speaking or bilingual backgrounds, which is particularly important with young children because they may have had limited or no exposure to English. Whereas the review includes information regarding reliability and validity, the primary focus is on dimensions important to the assessment of young children which are not often addressed directly in test technical manuals or standards for psychological assessment, particularly psychometric issues that arise when assessing young children who have developmental delays.
All of the tests reviewed included a measure of “cognitive functioning.” However, not all of the tests are considered to measure “intelligence,” nor are they all appropriate for use in diagnosing intellectual disabilities. Some test developers, such as the author of the Battelle Developmental Inventory, Second Edition (BDI-II; Newborg, 2005), stated that the test was not designed to be used to diagnose specific developmental disabilities. Although other test manuals may not include such explicit cautions, clinicians using tests as part of an assessment of eligibility or a diagnostic evaluation should consider carefully the constructs measured by the test, and whether the test is appropriate for the proposed use.
Method
A search was conducted to identify all published tests which included a measure of cognitive functioning in children aged 2 to 4 years, and which provided normative data for the U.S. population. Sources included Internet searches, academic journal databases in psychology and medicine, consultation with colleagues, and contacts with major test publishers.
Test Characteristics Evaluated
Table 1 outlines the characteristics evaluated for each test, the criteria used to evaluate the adequacy of the tests, and sources used to determine criteria. Domains evaluated included (a) the standardization sample, (b) item bias analysis, (c) psychometric properties, (d) utility for nonverbal children, and (e) utility for non-English-speaking children. For each test characteristic the reviewers focused only on the age range under review (age 2 years 0 months through 4 years 11 months), even if the test may be used for a wider age range. If the test included measures of other constructs in addition to cognition (e.g., gross motor functioning), the focus was only on those subtests or composites designed to assess cognition.
Criteria for Test Evaluation.
Note. AERA = American Educational Research Association; SEM = standard error of measurement; NAEYC = National Association for the Education of Young Children.
Evaluation of test floors
Bracken’s (1987) recommendations were used to evaluate the adequacy of test floors. The norms tables for each subtest were reviewed to determine the standard score corresponding to a raw score of 1 (“floor standard score”). To evaluate the floor for composite and total test scores, we computed the total raw score resulting if the child obtained a raw score of 1 on every subtest within that composite or total test. Norms tables were used to determine the standard score corresponding to that total raw score. Using Bracken’s criteria, we considered the floor adequate if the corresponding standard scores were at least two standard deviations below the mean (below first percentile).
Evaluation of item gradients
Item gradients refer to the change in standard scores that results from changes in raw scores. The greater the number of items in a given age range or functioning level of the test, the more sensitive the test will be to differences between children’s performances. When a test has large changes in standard scores based on very small changes in raw scores, then it has insufficient item gradients. Bracken’s (1987) recommendations were followed for the evaluation of item gradients. For each cognitive subtest, item gradients were considered adequate if there were at least 3 raw score points within each standard deviation of change in standard score. We only considered errors in item gradients that were below the mean for the subtest, because our focus was on assessment of children with developmental delays. Item gradients were considered adequate if there were no item gradient errors below the mean.
Evaluation of method of test translation
Guidelines for test translation outlined by the Standards for Educational and Psychological Tests (AERA et al., 1999) were used to evaluate the tests that had published translations into languages other than English. Specifically, we reviewed (a) the method used for translation, (b) the availability of normative data for the translated test, and (c) evidence of comparability of the English and the translated version(s) of the test.
Results
The list of tests, test references, and a summary of the findings for key characteristics of each test is provided in Tables 2 and 3.
Characteristics of Cognitive Assessment Measures for 2- to 4-Year-Old Children.
Note. SEM = standard error of measurement; BDI-2 = Battelle Developmental Inventory, Second Edition; CAS-2 = Cognitive Abilities Scale, Second Edition; CAYC = Cognitive Assessment of Young Children; DAS-II = Differential Abilities Scale, Second Edition; KABC-II = Kaufman Assessment Battery for Children, Second Edition; PTI-2 = Pictorial Test of Intelligence, Second Edition; PTONI = Primary Test of Nonverbal Intelligence; SB-5 = Stanford–Binet Intelligence Scales, Fifth Edition; SB-EC = Stanford–Binet Intelligence Scales for Early Childhood; WJ-III = Woodcock–Johnson III Tests of Cognitive Abilities; Batería WM-III = Batería III Woodcock–Munoz; WNV = Wechsler Nonverbal Scale of Ability; WPPSI-IV = Wechsler Preschool and Primary Scale of Intelligence, Fourth Edition.
Characteristics of Standardization Samples.
Note. BDI-2 = Battelle Developmental Inventory, Second Edition; CAS-2 = Cognitive Abilities Scale, Second Edition; CAYC = Cognitive Assessment of Young Children; DAS-II = Differential Abilities Scale, Second Edition; KABC-II = Kaufman Assessment Battery for Children, Second Edition; M-P-R = Merrill-Palmer-R; PTI-2 = Pictorial Test of Intelligence, Second Edition; PTONI = Primary Test of Nonverbal Intelligence; SB-5 = Stanford–Binet Intelligence Scales, Fifth Edition; SB-EC = Stanford–Binet Intelligence Scales for Early Childhood; WJ-III = Woodcock–Johnson III Tests of Cognitive Abilities; Batería WM-III = Batería III Woodcock–Munoz; WNV = Wechsler Nonverbal Scale of Ability; WPPSI-IV = Wechsler Preschool and Primary Scale of Intelligence, Third Edition.
Standardization Samples
Table 2 provides details about the number of children per year in the standardization sample for ages 2 to 4 years, and the year of collection of the standardization sample. All but two of the tests reviewed met the criterion set by Salvia, Ysseldyke, and Bolt (2013) that tests used for high-stakes decisions should be re-normed every 15 years (1998-present); seven tests met the criterion set by Alfonso and Flanagan (2009) of being normed within 10 years (2003-present). The two tests with outdated norms were the Leiter International Performance Scale–Revised (Leiter-R; Roid & Miller, 1997), normed in 1994 to 1995, and the Mullen Scales of Early Learning: AGS Edition (Mullen Scales; Mullen, 1995), normed in 1981 to 1989. The Leiter is in the process of revision with an expected publication date of the third edition in 2013.
Detailed information regarding the standardization samples is provided in Table 3, including the representativeness of the standardization sample as compared with the U.S. Census on demographic variables (usually gender, race/ethnicity, geographic region, and educational level of parent as a proxy for SES). Nine tests documented the inclusion of children with atypical development in their standardization sample and all but one test conducted special group validity studies with children with atypical development in the 2- to 4-year age range.
Psychometric Properties: Reliability and Validity
As detailed in Table 2, all the tests reviewed documented reliability through internal consistency of the composite score(s), and through reporting of the standard error of measurement for scores. Two tests had test–retest reliability coefficients below the threshold of r = .80 recommended for adequate stability of the test scores; these were the Mullen Scales and the Woodcock–Johnson III Tests of Cognitive Abilities (WJ-III; Woodcock, McGrew, & Mather, 2001, 2007). The Batería III Woodcock–Munoz (Batería WM-III; Munoz-Sandoval, Woodcock, McGrew, & Mather, 2005) did not provide test–retest reliability data for the Spanish version of the test.
While validity data are gathered over time as a test is used in practice, information is provided in Table 2 regarding whether several types of validities were documented in the test manual: (a) factor analysis of scores, (b) comparisons of the test to other comparable tests thought to measure similar constructs, (c) comparisons of special groups of children in their performance on the test, and (d) predictive validity, or the test’s ability to predict future functioning. Almost all of the tests reviewed documented comparisons to other tests of similar constructs, and almost all documented special group comparisons. There were two exceptions: (a) special group comparisons on the WJ-III were done only for children aged 5 years and older and (b) the Batería WM-III did not separately examine the validity of the Spanish version of the test. Because a calibration sample was used to make the Batería WM-III comparable with the WJ-III, the manual indicated that the psychometric properties of the Batería WM-III should be considered comparable with the WJ-III. However, it would be preferable for reliability and validity of scores to be separately assessed for the Spanish version of the test, because children’s scores are compared with a different standardization sample from the original WJ-III. Almost all tests documented factor analytic findings for test scores; exceptions were the BDI-II, and analysis done only for ages 5 years and older for the WJ-III and Batería WM-III. Three test manuals included information about predictive validity: the Cognitive Abilities Scale, Second Edition (CAS-2; Bradley-Johnson & Johnson, 2001), the Kaufman Assessment Battery for Children, Second Edition (KABC-II; Kaufman & Kaufman, 2004), and the Mullen Scales. For other tests, the published literature may document predictive validity as tests are used over time.
Psychometric Properties: Item Floors
Table 2 presents the youngest age at which each test demonstrated good or adequate item floors; as detailed in Table 1, tests were considered good with respect to floors if all subtests and composite scores had adequate item floors, and adequate if the composite scores had an adequate floor, but one or more subtests making up the composite had inadequate floors. There were three tests with inadequate composite scores at the lowest age range for the test. The Leiter-R has problematic floors on the composites from ages 2:0 through 2:5. Three out of seven subtests have problematic floors through age 4:11; therefore, examiners should use caution when interpreting “strengths” on those subtests. The Pictorial Test of Intelligence (PTI-2; French, 2001) has an insufficient composite floor for age 3:0 to 3:5. The composite floor is adequate starting at age 3:6, but two of the three subtest floors remain problematic until age 5:0. The WJ-III has inadequate composite floors for age 2:0 through 2:1; while the primary cognitive composite is adequate starting at age 2:2, more than half of the subtests making up the composite have inadequate floors through age 4:4. The Wechsler Preschool and Primary Scale of Intelligence (WPPSI-IV; Wechsler, 2012) has problematic floors for the Working Memory Index composite at ages 2:6 through 3:2; the other composite floors are adequate for the full age range of the test. The Wechsler Nonverbal Scale of Ability (WNV; Wechsler & Naglieri, 2006) has an adequate floor for the four-subtest version of the Full Scale score; however, the test has inadequate floors for two out of the four subtests (Coding and Object Assembly). Therefore, using the two-subtest option for calculation of the Full Scale Score would result in an inadequate floor at ages 4:0 through 4:8 if the combination of those two subtests was used, or if the combination of Coding and Recognition was used.
Psychometric Properties: Item Gradients
As described under “Method” section, errors in item gradients occur when there are fewer than 3 raw score points for each standard deviation change in standard score. Ten of the 16 tests reviewed have adequate item gradients for all subtests included in the cognitive composite scores. Problematic item gradients on all or most subtests were found on the Leiter-R, Stanford–Binet Intelligence Scales (SB-5; Roid, 2003)/Stanford–Binet Intelligence Scales for Early Childhood (SB-EC; Roid, 2005), and WJ-III/Batería WM-III, where small changes in raw scores result in clinically significant changes in scaled scores through age 4:11. The KABC-II has problematic item gradients on three subtests through age 4:11. The WPPSI-IV has problematic item gradients on three subtests used in the calculation of the Full Scale IQ: Picture Memory (ages 2:6 through 3:11), Information (age 2:6 through 2:11), and Object Assembly (age 2:6 through 3:8).
Evaluation of Item Bias
Two primary methods were used by test developers to evaluate item bias: (a) expert review to screen out items that may be offensive, stereotyping, or confusing and (b) statistical analysis of items to examine differences in test responses in different ethnic or racial groups. Nine tests conducted expert review of items, and 14 conducted statistical review of item bias. Only one test, the Mullen Scales, did not report any information about evaluation of item bias.
Translations for Non-English-Speaking Children
Decisions about test selection for a young child exposed to a primary language other than English are complex. As discussed by Ortiz and Ochoa (2005) in the case of school-age children, knowledge about the child’s exposure to and proficiency in languages and the type and extent of bilingual or English-only instruction in school should influence decisions about the best assessment approach. Assessment in English and the home language represents the best approach, when children have been exposed to both languages and appropriate tests are available (NAEYC, 2005).
The Spanish version of the WJ-III, the Batería WM-III, was developed using a carefully documented method of test translation and calibration to the English version, with separate norms based on a sample of native-Spanish-speaking children and adults from eight countries. However, the age range for the calibration sample was not specified in the manual, and may not have included young children. The Differential Abilities Scale, Second Edition, Early Years Spanish Supplement (Elliott, 2012) is a Spanish translation of the Differential Abilities Scale (DAS-II; Elliott, 2007) for ages 2:6 through 6:11. Test translation included a rigorous process of back-translation, expert review, pilot testing, and equivalency testing including a sample of 395 Spanish-speaking children. For translated subtests that did not result in scores equivalent to the English published test, the examiner is directed to convert raw scores to ability scores to allow use of the original DAS-II norms.
Five additional tests included translation of at least some parts of the test in Spanish (see Table 4), but the tests were not evaluated for equivalence or comparability of the translated version to the English version of the test, and norms for Spanish-speaking children were not provided. Three of the Spanish translated tests did not include non-English-speaking children in the standardization sample: BDI-2, KABC-II, and WNV. The KABC-II manual stated that the test is not intended to be administered in Spanish and is designed for use with children who are proficient in English. The Leiter-R is administered entirely nonverbally, so no translation of test instructions is needed. The Primary Test of Nonverbal Intelligence (PTONI; Ehrler & McGhee, 2008) provides translation of test instructions in eight languages, and the WNV provides translation in five languages; the child’s responses on both tests are nonverbal. However, no information about the method of translation was documented in the manual for either the PTONI or the WNV.
Characteristics of Cognitive Assessment Measures for Nonverbal and Non-English-speaking 2- to 4-Year-Old Children.
Note. BDI-2 = Battelle Developmental Inventory, Second Edition; CAS-2 = Cognitive Abilities Scale, Second Edition; LR = language reduced; CAYC = Cognitive Assessment of Young Children; DAS-II = Differential Abilities Scale, Second Edition; ASL = American Sign Language; KABC-II = Kaufman Assessment Battery for Children, Second Edition; NV = nonverbal; ESL = English as a Second Language; M-P-R = Merrill-Palmer-R; PTI-2 = Pictorial Test of Intelligence, Second Edition; PTONI = Primary Test of Nonverbal Intelligence; SB-5 = Stanford–Binet Intelligence Scales, Fifth Edition; SB-EC = Stanford–Binet Intelligence Scales for Early Childhood; WJ-III = Woodcock–Johnson III Tests of Cognitive Abilities; Batería WM-III = Batería III Woodcock–Munoz; WNV = Wechsler Nonverbal Scale of Ability; ELL = English-Language Learner; WPPSI-IV = Wechsler Preschool and Primary Scale of Intelligence, Fourth Edition.
Nonverbal cognitive tests are another option to consider for children who have had primary exposure to a language other than English in the home (Ortiz & Ochoa, 2005). Only one test in the 2- to 4-year-old age range is completely nonverbal in instructions and responses, the Leiter-R. The KABC-II has a nonverbal composite that does not require any verbal response and in which instructions can be given in pantomime. Another option is to consider tests with reduced language demands. There are eight additional tests or composites within tests (see Table 4) that do not require any verbal response and have relatively low receptive language demands. It is important for clinicians to specifically examine the language demands for each item or subtest to ensure that children have the necessary understanding of test instructions before interpreting their performance.
Discussion
This comprehensive review of cognitive assessment measures for 2- to 4-year-old children focused on psychometric and test development issues that directly impact the usefulness of tests for young children, particularly those with developmental delays and/or those whose primary language is not English. While test development standards and test manuals generally focus on establishing reliability and validity, the psychometric considerations in this review highlight additional characteristics that are important when developing appropriate tests for this subpopulation of children. Careful review of normative tables and standardization samples assists examiners in critically evaluating the applicability of specific tests to particular populations of interest in clinical or research settings.
When selecting a test for a specific subpopulation, the following conclusions can be drawn about currently available tests. To assess young children from English-speaking families, there are seven cognitive tests which were rated as good or adequate in all areas: BDI-2, Bayley-III, CAS-2, Cognitive Assessment of Young Children (CAYC; Langley, Fewell, & Maddox, 2010), DAS-II, Merrill–Palmer–Revised (M-P-R; Roid & Sampers, 2004), and WNV. Three additional tests were rated good or adequate in all areas with the exception of problematic item gradients on some subtests: KABC-II, SB-5, and the WPPSI-IV (which has item gradient problems only for children younger than 4 years). The PTI-2 was rated as good or adequate in all areas for ages 3:6 to 4:11; problematic item floors were found for younger children. Finally, the PTONI was rated as good or adequate in all areas with the exception of some minor problems with the representativeness of the standardization sample.
When assessing children from Spanish-speaking families with limited exposure to English, the DAS-II Early Years Spanish Supplement provides a carefully translated test with established equivalence with the English version. The Batería WM-III also provides an appropriately translated Spanish language option. However, the manual provides insufficient information to determine whether children in the younger age range were included in the calibration sample for the Spanish version, and problematic item floors and item gradients limit the test’s appropriateness for younger or lower-functioning children. Examiners who use a translated test such as the BDI-2 or KABC-II would need to compare the resulting scores with norms for the English version of the test, which were collected on an entirely English-speaking standardization sample. In that case, caution should be used in interpreting standard scores, because the normative group may be quite different from the child being tested and the Spanish translation has not been evaluated to determine whether it is of equivalent difficulty to the English test. Use of a nonverbal test or composite, or a test with reduced language demands, is another option for testing young children whose primary language is not English. Tests which were rated as adequate or good in all or most areas and also include a composite score with reduced language demands are the CAS-2, DAS-II, KABC-II, M-P-R, PTONI, SB-5, WNV, and WPPSI-IV. Psychologists should consider the receptive and expressive language demands for each subtest when interpreting results rather than relying on a test manual’s label of nonverbal, a term which is used differently by different publishers.
Because appropriate tests are available for this age group that were normed within the past 15 years, it is recommended that examiners avoid tests with older norms for either research or clinical purposes. Numerous published research studies, particularly in the autism field, continue to use the Mullen Scales, despite normative data that are more than 30 years old, the absence of children with atypical development in the normative sample, inadequate test–retest reliability, and the lack of evaluation of item bias of the test (e.g., Kasari, Gulsrud, Wong, Kwon, & Locke, 2010; Landa & Kalb, 2012; Toth, Munson, Meltzoff, & Dawson, 2006). It is recommended that researchers select one of the comparable developmental tests that cover similar domains of functioning as the Mullen Scales with much more recent norms and more appropriate standardization samples (e.g., BDI-2, Bayley-III, or, for a few more years, the M-P-R). Similarly, in clinical settings, particularly those with large numbers of non-English-speaking children, the Leiter-R continues to be frequently used despite the norms being 18 years old and despite problems with floors and item gradients for young children; publication of the third edition of the Leiter in 2013 may address these issues. In clinical settings, the use of outdated tests can result in children inappropriately being found ineligible for services. In research studies, outdated tests may lead to overestimates of the actual level of functioning of participants in studies (as compared with the current cohort of children). In addition, outdated tests have standardization samples less representative of the current population.
In the years since the previous reviews of cognitive tests for young children (Bradley-Johnson, 2001; Lichtenberger, 2005) were published, there have been important test revisions and additions of new test options for this age group. In particular, revisions of the Bayley, BDI, DAS, and WPPSI, and the addition of the CAYC, PTONI, WNV, and the Spanish DAS, will meet the needs of many clinicians and researchers for measures with adequate psychometric properties to enable appropriate testing. In addition, the current article extends the previous reviews by including all available cognitive measures for this age group, adding discussion of nonverbal cognitive assessment tools, and including consideration of issues unique to the assessment of bilingual or non-English-speaking children.
To address the concerns outlined, it is recommended that test developers pay special attention to adequate floors by providing sufficient basal items to improve accuracy at the low end of scales. In addition, it is critical that sufficient items be included throughout each scale so that item gradients discriminate appropriately between children of different levels of ability. Adequate floors and item gradients at the low end of the scale are critical to evaluate children with developmental delay. Clinicians should be attentive to floor and item gradient issues when selecting a test and when evaluating a child’s performance. For example, for some subtests reviewed, a change of one raw score point led to a change in standard score from the significantly below-average range (i.e., less than the first percentile) to a standard score in the low average range. For some subtests, a child receiving credit for only one item (often an item that includes teaching by the examiner), can score within the low average range. Careful examination of norms tables can inform psychologists when caution is necessary in interpreting scores, or when selecting an alternative test will more accurately measure a child’s functioning.
It is recommended that publishers develop appropriately translated and normed tests for young children in Spanish, in particular. There are 35 million people in the United States who speak Spanish at home as a primary language, and 22% of American children (ages 5-17) speak Spanish at home (U.S. Census Bureau, 2010). Young children may not have been exposed to English yet outside the home; therefore, it is even more important to have tests available in their native language than is the case with older children, who have developed proficiency in English through exposure at school. When developing translated versions of tests, the International Test Commission (2010) provided specific guidelines about appropriate methods of translation and use of translated tests. Finally, it is critical to obtain normative data for translated tests and to establish equivalence to the English version.
This review is limited in its focus on the utility and appropriateness of standardized cognitive assessment measures for children aged 2 to 4 years; the goal is to highlight key issues important to test development in young children, and to provide information important to test selection in an easily accessible format. In actual assessment practice, cognitive assessment measures are one tool among many, and are appropriately used only in the context of clear assessment questions, together with additional assessment methods. In particular, when assessing young children, attention to functioning in the natural context, contributions of professionals from various disciplines, and information provided by family members are essential to developing a clear picture of a child’s strengths and needs.
Footnotes
Acknowledgements
The authors acknowledge the provision of complimentary copies of test manuals from Pro-Ed, Inc. and Stoelting Company.
Authors’ Note
An earlier version of parts of this article was presented at the National Training Institute, Zero to Three, Washington, D.C., December 2011.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
