Abstract
Much has been written about gifted students with learning disabilities, but there have been few large-scale empirical investigations, and the concept has proven controversial. The authors reviewed the available empirical literature on these students, focusing on (a) the criteria by which the students were identified and (b) the students’ performance on standardized tests of ability and achievement. In addition, the test scores of these students were aggregated to determine typical performance levels. A total of 46 empirical articles were reviewed, and major findings included wide variability in identification criteria across studies, frequent reliance on dubious methods of learning disability identification, and a lack of academic impairment among the identified students. Implications for the “gifted/LD” category are discussed.
The professional literature on gifted students with learning disabilities (G/LD students; also called twice-exceptional students) continues to grow, with articles appearing on assessment guidelines as well as intervention approaches. There are even books on the G/LD population for parents and teachers seeking to support their G/LD students (e.g., Baum & Owen, 2004; Weinfeld, Barnes-Robinson, Jeweler, & Shevitz, 2006). In the present article, we review the empirical literature on G/LD students back to their earliest mention to investigate (a) the range of criteria used to identify these students and (b) the cognitive and achievement characteristics of these students. Since the professional literature offers descriptions of this population, it is imperative that we understand which students are given the G/LD label and whether the same students would tend to be identified using different sets of diagnostic criteria. The last such review (Cohen & Vaughn, 1994) concluded that more research was needed since the available literature failed to yield empirically based guidelines for identifying and describing G/LD students. It was our hope that in the intervening years, the G/LD literature would have grown sufficiently to allow generalizations about the criteria used to identify this population as well as the students’ characteristics.
Evolution of the G/LD Concept
In the 1970s, interest arose in gifted children who also had various disability conditions. The Council for Exceptional Children formed a committee on the topic, held two national conferences, and released a fact sheet for educators on how to serve these children (Nielsen, 2002; Whitmore & Maker, 1985). In addition, professional journals began to publish case studies of such individuals (e.g., Meisgeier, Meisgeier, & Werbolo, 1978; Thompson, 1971). More specific interest in G/LD students increased in the 1980s, when Johns Hopkins University held a conference on the topic and published a book (Fox, Brody, & Tobin, 1983) based on the research presented there. Empirical research also began to be published describing the characteristics of these students.
By the late 1980s, there was enough published scholarship on G/LD students for Vaughn (1989) to survey the extant literature in a search for consensus on definition, identification, and intervention with this population. Her conclusions were largely negative; although she acknowledged the “intuitively appealing” nature of the G/LD concept, Vaughn expressed concern that school districts tended to develop their own definitions of the G/LD category, researchers tended to publish case studies of G/LD students rather than studies with large-scale samples, and none of the intervention programs developed at that point reported data on efficacy. A follow-up literature review 5 years later (Cohen & Vaughn, 1994) reached similar conclusions.
Since then, many more articles have been published on G/LD students, including an article on “best practices” for identifying these students (McCoach, Kehle, Bray, & Siegle, 2001), suggesting that research had progressed sufficiently to allow for evidence-based practices. More recently, though, Lovett and Lewandowski (2006) criticized the G/LD assessment guidelines proposed in a variety of sources, arguing that they either were challenged by research evidence (e.g., using IQ subtest profile analysis) or relied on outdated methods of assessing learning problems. However, Lovett and Lewandowski focused on proposed guidelines rather than on empirical studies of G/LD students, leaving unaddressed the question of whether empirical studies used similar inclusion criteria, one topic of the present review.
Two Contested Diagnostic Categories
The issue of inclusion criteria, the way of circumscribing our target population, is a matter of special concern in G/LD scholarship because both giftedness and learning disabilities suffer from disagreement over identification procedures. Although scholars agree that some students have special gifts requiring unique programming and other students have learning problems that require individualized attention, there is a long history over which students should qualify for each of the two labels that make up the G/LD category. In the case of giftedness, there are almost as many definitions as there are scholars (see Sternberg & Davidson, 2005, for a survey of diverse viewpoints). Even when IQ tests are agreed on as a core component of the definition, there is debate over whether the cutoff score for giftedness should be 120, 130, or some other number (L. J. Coleman & Cross, 2005; Minton & Pratt, 2006; Ruf, 2003) or whether the use of any cutoff score should be abandoned (e.g., Richert, 2003). The issue is further complicated by G/LD scholars’ frequent contention that the standards for giftedness should be relaxed when assessing G/LD students since the learning disability is thought to artificially suppress the students’ IQ scores (Krochak & Ryan, 2007; Nielsen, 2002).
Learning disabilities have been the topic of similarly varying perspectives. Until recently, most schools diagnosed LD by calculating the difference between a student’s ability (typically a composite IQ score) and his or her achievement in a subject area (Reschly & Hosp, 2004). If there was a “severe discrepancy” between the student’s ability and achievement, the student was thought to suffer from LD that prevented achievement consistent with his or her ability (Kavale, 2002). However, IQ–achievement discrepancy scores came under attack on several grounds (Stanovich, 1999; Sternberg & Grigorenko, 2002). Critics have noted that these scores have poor reliability over time, fail to distinguish among groups of students needing different instructional programming, and fail to identify students who have lower IQs in addition to LD.
Several diagnostic approaches have come to take the place of IQ–achievement discrepancies. Perhaps the most popular is the response to intervention (RTI) approach (Gresham, 2002) in which students are provided increasingly intense and individualized instructional interventions in a given subject area and an LD diagnosis is made after the student fails to respond to multiple interventions. RTI models have their own critics, however (e.g., Kavale & Spaulding, 2008; Reynolds & Shaywitz, 2009), and still other scholars have proposed LD criteria based on low achievement alone (Stanovich, 1999), neuropsychological deficits (Rourke, 2005), or even rehabilitated versions of IQ–achievement discrepancies (Kavale, Holdnack, & Mostert, 2005). As such, there is still no consensus on how to define LD.
The Present Study
Since the definitions of giftedness and LD each show such range, it is a very real concern that the G/LD category may be too heterogeneous to allow generalizations. Given that literature on G/LD students is now directed at a variety of audiences (teachers, school counselors, school psychologists, parents, etc.), we were interested in determining if any useful generalizations can be made about this population. In the present study, we reviewed all of the available empirical literature on G/LD students to determine the range of criteria used to identify students as G/LD as well as the cognitive and academic characteristics of these students. The only earlier reviews of the G/LD literature have been narrative in nature as well as wider in scope, covering intervention programs and noncognitive characteristics (e.g., emotional disorders, social skills). We sought, instead, to produce quantitative summaries of a smaller body of information.
Method
The research studies included in this review were identified using several steps. First, a search of the ERIC and PsycINFO databases was completed using the following keyword descriptors: gifted and learning disabilities, twice exceptional, and dual exceptionality. The search was not limited by year or by publication format (i.e., dissertations as well as journal articles were included). This process yielded a total of 940 abstracts.
Second, each of the 940 abstracts was reviewed to determine whether each article was an empirical study presenting original data on G/LD students. If an abstract described an empirical study, or if it was ambiguous, the full article was obtained for further review. Of the 940 articles initially retrieved, 49 (5.2%) were found to be empirical, and the full text of these articles was obtained for further review. Of the 49 articles, 3 were excluded based on issues with the authors’ definitions of their samples (i.e., in each of these cases, the samples were not explicitly classified as G/LD, even though the investigators felt that their data were useful in understanding the G/LD population), leaving 46 empirical articles.
Finally, the 46 empirical articles were examined with specific attention to three types of information: (a) the inclusion criteria used to determine giftedness, (b) the inclusion criteria used to determine classification as LD, and (c) IQ and achievement test scores of the participants.
Results
Of the 46 studies (empirical articles), 18 were dissertations, full copies of which were obtained through either electronic databases or interlibrary loan services. All 46 studies reported inclusion or classification criteria. Only 19 of the 46 studies both reported mean test scores and had samples of at least five participants in the G/LD group; test scores from these studies were aggregated. Of these 19 studies, almost all (k = 18) reported mean IQ scores on a Wechsler IQ test (a version of either the Wechsler Intelligence Scale for Children [WISC] or Wechsler Adult Intelligence Scale [WAIS]), and 5 reported achievement test scores for at least one of three clusters (reading, mathematics, or written language) on a version of the Woodcock–Johnson Tests of Achievement. Of the 19 studies, 2 reported scores for performance on additional measures as well.
Criteria for Classification as G/LD
Although all of the studies reported the criteria by which participants were deemed appropriate for the study, the criteria were not always explicitly categorized into evidence of giftedness and evidence of learning disability. We were generally able to categorize things this way, and Table 1 shows the inclusion criteria for each of the 46 studies.
Inclusion Criteria for Empirical Gifted Students With Learning Disabilities (G/LD) Studies.
Note: WISC-R = Wechsler Intelligence Scale for Children–Revised; VIQ = Verbal IQ; PIQ = Performance IQ; FSIQ = Full-Scale IQ; CAS = Cognitive Assessment System; WRMT-R = Woodcock Reading Mastery Tests–Revised; WAIS-R = Wechsler Adult Intelligence Scale–Revised; MR = Mental retardation; TOAL-2 = Test of Adolescent Language–2; IEP = individualized education program; DSM-IV-TR = Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition, Text Revision; CTBS = Canadian Test of Basic Skills; WRAT-R = Wide Range Achievement Test–Revised.
Evidence of giftedness
Most of the studies (k = 36, or 78%) specified a minimum IQ cutoff as at least part of the identification criteria for giftedness. The most common cutoff was 120 (k = 16), followed by 125 (k = 8) and 130 (k = 7). However, these cutoffs sometimes had different meanings since certain investigators required a full-scale (total) composite score above the cutoff, whereas other investigators allowed any of the composite scores (e.g., verbal IQ, performance IQ, etc.) to be above the cutoff to classify the student as gifted.
Several studies relied on prior identification by schools (k = 5) or used other criteria, such as exhibiting traits from a gifted behavior checklist, having high achievement in a subject area, or being nominated by a teacher as gifted.
Evidence of learning disability status
For evidence of LD, many studies (k = 17, 37%) relied on prior documentation, prior classification by schools, or the meeting of official school district or state guidelines, without specifying performance any further. However, many investigators specified an IQ–achievement discrepancy (k = 20, 43%). The most common explicitly stated discrepancy required was one standard deviation (k = 5), although typically the degree of discrepancy was not mentioned. Only 8 of the studies (17%) required that the student show achievement scores below an absolute standard (or at least “difficulty” with an academic subject), and even among these studies that achievement standard was sometimes at or above the average level (e.g., 1 study required that achievement be below the 70th percentile). Finally, 3 studies used discrepancies between different IQ composite scores or subtest scatter as evidence of LD.
Additional evidence of G/LD status
A small number of studies (k = 6) used at least one criterion for inclusion that could not be categorized as having to do with either giftedness or LD. Often, these criteria were vague, such as requiring that the participants exhibit “twice exceptional characteristics” or “asynchronous development,” be a “unique or extreme case,” or have performance “illustrative of the strengths, weaknesses, and variability in skill levels that characterize” G/LD students.
Test Performance of G/LD Samples
Most of the studies (k = 27) reported mean IQ scores of their participants (or reported each participant’s IQ score), whereas fewer studies (k = 13) reported at least one mean achievement score. Of the studies, 20 did not report any test scores at all or included only score ranges. In addition, many of the articles reporting test score averages were based on very small samples (n < 5) and were presented as a multiple case study series rather than as quantitative research per se.
Table 2 shows participants’ average IQ and achievement scores for studies that reported score averages and had at least five participants representing each score average (k = 19). All but one of these studies reported IQ scores from a Wechsler IQ test (a version of the WAIS or WISC), and the most popular achievement test used was the Woodcock–Johnson tests (k = 5).
Test Scores for Empirical Gifted Students With Learning Disabilities Studies.
Note: WJ-III = Woodcock–Johnson III; WISC-R = Wechsler Intelligence Scale for Children–Revised; WAIS-R = Wechsler Adult Intelligence Scale–Revised; WJ-R = Woodcock-Johnson Revised; WRAT-R = Wide Range Achievement Test–Revised; PIAT-R = Peabody Individual Achievement Test–Revised.
Weighted averages of IQ and achievement scores were conducted across studies for the six test scores that were reported by multiple studies (for this purpose, different versions of the same test were combined, as were scores from the WISC and WAIS). Table 3 reports the weighted averages for these six test scores as well as the number of studies and total number of participants on which these averages are based. The average Wechsler Full-Scale IQ was 122.8 (93rd percentile), based on 983 participants across 17 studies. The average Wechsler Performance IQ was very similar (125.9, 95th percentile), whereas the average Wechsler Verbal IQ was somewhat lower (118.6, 89th percentile). The average Woodcock–Johnson achievement scores were each based on between 440 and 442 participants across five studies. The average Mathematics Cluster score was highest (111.1, 77th percentile), followed by the Reading Cluster score (95.8, 39th percentile) and the Writing Cluster score (93.0, 31st percentile).
Weighted Mean Test Scores for Studies With n > 5 and Tests Used in Multiple Studies.
Note: FSIQ = Full-Scale IQ; VIQ = Verbal IQ; PIQ = Performance IQ; WJ = Woodcock–Johnson.
Discussion
In this review, our goal was to synthesize the empirical literature on G/LD students, to determine how these students were identified, and to report how they performed on tests of intelligence and academic skills. Here, we summarize our findings in six general points.
First, there was far less empirical literature than one would expect based on the number of citations in the literature and the apparent acceptance of the G/LD concept. Only approximately 5% of the articles written about G/LD students included data from empirical research, and even many of these empirical reports were case studies and/or had very small sample sizes. Moreover, more than one third of the reports were dissertations, which are not widely accessible. It seems, then, that there is far more written about this population (940 abstracts of articles were found) than there are data about the population.
Second, we found wide variability in the criteria used to identify both components of the G/LD classification. What “G/LD” means varied from study to study, and although individual studies made confident claims about what G/LD students are like based on the sample examined, students who were classified as G/LD in one study may not have been classified the same way in other studies. Even when studies agreed in general terms (e.g., the use of an IQ cutoff to identify giftedness), the precise numerical requirements differed across studies. Clearly, there is no overarching consensus in the G/LD field as to how to identify students who should be classified as G/LD.
Third, although LD criteria varied widely, use of an IQ–achievement discrepancy was common. Almost half of the studies mentioned the discrepancy explicitly, and many other studies specified the use of state regulations, which often included discrepancy requirements at the time (Reschly & Hosp, 2004). The problems of discrepancy scores for classification as LD are well known, and the wide usage of these scores in identifying G/LD students means that our collective knowledge about this population is likely to be based largely on a problematic identification method.
Fourth, very few studies required any kind of true academic impairment (by an achievement score or otherwise) for the LD diagnosis. IQ–achievement discrepancies require academic impairment in a relative sense—impairment relative to a student’s IQ. But only 17% of the studies required low achievement in a normative sense, that is, compared to other students. Even among these studies, some investigators allowed students with average achievement to be classified as LD. For instance, Waldron and Saphire (1990) required only that students score less than the 70th percentile on standardized achievement tests (in addition to the discrepancy). The “LD” label, then, has a very different meaning in the G/LD context, at least in the studies that we examined. Although low achievement has typically been seen as one of the requirements for an LD diagnosis (indeed, even for a referral for evaluation for possible LD), this is not the case in the G/LD population.
Fifth, although most studies required an IQ score above a certain cutoff for giftedness, the IQ scores of the G/LD population were somewhat less than what might be expected. The weighted average full-scale IQ across almost 1,000 G/LD students was 122.8 in the studies we examined, which is not even in the top 5% of the population, assuming a perfectly normal distribution. And since this is the average score, many G/LD students had IQ scores less than this. In addition to wide variability in the criteria used for classification of G/LD cited earlier, we attribute this relatively low performance to the “either–or” inclusionary criteria used by many studies, in which any composite IQ above a certain cutoff would qualify a student as gifted. This type of inclusionary criterion capitalizes on measurement error, allowing unrepresentative scores to classify students.
Sixth and finally, the weighted average achievement scores for G/LD students were all in at least the average range (between 93 and 112). This raises the possibility that the typical G/LD student is not impaired in an absolute sense (i.e., relative to the average same-age student) in academic skills, and some G/LD students are substantially above average academically. Admittedly, it is theoretically possible that many G/LD students had quite low achievement in only one area and that the weighted averages are inflated by high performance in other areas. However, we view this possibility as unlikely for two reasons. First, several of the authors reported test scores for each student in the sample (or maximum and minimum test scores for the sample), and it was not unusual to find G/LD students without any below average achievement scores, even in their area of classification. Second, a recent study by Lovett and Sparks (2010) found that less than 5% of a high-IQ (≥ 120) sample of college students with LD diagnoses met LD criteria that included a normative academic impairment standard. Therefore, we interpret our weighted average academic achievement scores as suggesting a lack of academic impairment in many G/LD students.
Conclusions: Does G/LD Exist?
The aforementioned findings revealed by our examination of the G/LD literature raise the following question: If many G/LD students fail to show academic skill deficits compared to same-age peers, and some G/LD students do not even have generally high IQs, might the G/LD category be a null set? That is, does the G/LD construct refer to an actual group of students? Is it a meaningful notion? In posing the question this way, we are inspired by Stanovich’s (1994) article titled “Does Dyslexia Exist?” in which he noted that questions about the existence of a disorder reduce to issues of measurement and identification of the disorder. In the case of dyslexia, Stanovich noted that if “dyslexia” simply refers to poor reading skills and is identified through a low score on a standardized reading test, dyslexia certainly exists. However, as the measurement and identification procedures change, the question becomes more complex.
So it is with the G/LD concept. If G/LD simply refers to a high IQ score and a low achievement score, G/LD students exist. Indeed, as Lovett and Lewandowski (2006) noted, “[N]o matter how high the IQ score cutoffs and how low the achievement cutoffs are, some children will meet criteria for both giftedness and specific LD” (p. 525). However, problems occur when the cutoffs are not the same across different studies (and, in practice, across schools and school districts), when some diagnosticians do not even use cutoffs, and when some G/LD scholars do not even provide guidelines for identification, arguing that the two conditions of giftedness and specific learning disability can mask each other and prevent the detection of either condition (see McCoach et al., 2001, for critical discussion of the “masking hypothesis” in the G/LD literature).
The key to making the G/LD concept meaningful is selecting one identification method that is used consistently and that has validity evidence. Recent studies of LD identification have shown that small differences in identification methods can yield very different samples of students classified as LD (e.g., Proctor & Prevatt, 2003; Sparks & Lovett, 2009). We would expect this to be even more the case when two sets of criteria (one for giftedness, one for LD) are permitted to vary, and so choosing a consistent identification technique is truly of paramount importance.
How can we develop an appropriate identification technique? Giftedness is the easier of the two conditions to operationalize, and on that topic we make four specific recommendations. First, the identification of giftedness should be restricted to intellectual giftedness, recognizing that although students may excel in nonintellectual areas, those areas are of less relevance to academic settings in which identification occurs. Second, the most reliable measures of intellectual giftedness with the most accumulated validity evidence are standardized IQ tests; therefore, IQ scores should be the primary method of identifying gifted students. Teacher nominations and school grades may lead to referral, but these sources of evidence should not determine identification. Third, the full-scale IQ score should be the default score used in identification; not only does the full-scale IQ generally have the most reliability and validity evidence, but also alternative identification rules in which any high composite score signals giftedness can lead to false positives through measurement error (i.e., by chance, many students will have an unusually high score on one of the composite IQs). Finally, the choice of a specific cutoff may depend on local programming resources and needs, but we recommend a cutoff of at least 120, which would identify less than 10% of a theoretical normal IQ distribution. More liberal cutoffs would allow students whose intelligence is in the average range (albeit the “high average” range) to be described, incorrectly, as unusually intellectually gifted.
The LD diagnosis is the more complicated component of the G/LD concept, but the key element of LD for which we would make the case is academic impairment in a normative, absolute sense. The many problems with IQ–achievement discrepancies are exacerbated in the context of giftedness, where the discrepancy can exist even though all academic skills are in the average range. Even the contemporary defenders of discrepancies for LD diagnosis acknowledge that LD requires low academic skills. As Kavale et al. (2005) noted, even when a discrepancy is used, LD “should be associated only with significantly below-average achievement levels” since “students should be referred only if they are exhibiting signs of academic difficulty” (p. 5). The formal model of LD diagnosis that best acknowledges the importance of impairment was developed by Dombrowski, Kamphaus, and Reynolds (2004; also see Brueggemann, Kamphaus, & Dombrowski, 2008, for further discussion). These authors proposed that the first sign of LD be a low “norm-referenced academic achievement test score.” They proposed a cutoff of a standard score of 85, although even a cutoff of 90 would at least show that the student is functioning in the bottom quartile. Dombrowski et al. also noted the importance of evidence for academic impairment in the real-world classroom (i.e., below average grades or similar evidence). There are additional components to the model (exclusion of other causes of low achievement, etc.), but the academic impairment is central. The model is supported by research on the prognosis and underlying cognitive problems associated with learning disabilities, but even more important for our purposes, it ensures that even gifted students with LD have academic skill problems.
It is interesting that Dombrowski et al. (2004) noted, as an aside, that their model “will virtually eliminate the practice of diagnosing a child with ‘gifted LD’” (p. 369). While advocating for their model, we disagree with the authors on this point. Some gifted students have true academic impairment, and these students would be meaningfully classified as G/LD. For example, a student could achieve a full-scale IQ score of 125 and obtain a reading score of 80. Of course, the size of this “true G/LD” population will be much smaller than the population identified by the criteria currently in wide use, in which LD standards are relaxed to accommodate giftedness and vice versa.
In sum, then, G/LD students do exist, and these students may benefit from services for both gifted students and students with learning disabilities. If reliable and valid methods are used to identify each component of the G/LD classification, and identification procedures for each component are not affected by the suspected presence of the other component, G/LD is a meaningful concept. Regrettably, our examination of the extant G/LD literature showed a lack of attention to these concerns, despite evincing an uncritical acceptance of the G/LD concept. We hope that our findings, which document the consequences of neglecting these concerns, will inspire improved identification procedures and a meaningful G/LD category.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
