Abstract
This article synthesizes reading intervention research studies intended for use with struggling or at-risk students to determine which studies adequately address population validity, particularly in regard to the diverse reading needs of English language learners. An extensive search of the professional literature between 2001 and 2010 yielded a total of 67 reading intervention studies targeting at-risk elementary students. Findings revealed that many current research studies fail to adequately describe the sample, including the accessible and target populations, and to disaggregate their findings based on demographic characteristics. When population validity issues are not addressed, researchers cannot generalize findings to other populations of students, and it becomes unclear what intervention strategies work, especially with English language learner student populations. However, 25 studies did specifically recognize and address the needs of English language learners, indicating more researchers are taking into consideration the diverse needs of other struggling student populations.
Keywords
Special education researchers are engaged in a focused effort to improve the quality and utility of special education research. This is evidenced by the special issue of Exceptional Children dedicated to quality indicators in different types of research (e.g., for group design studies; Gersten et al., 2005) as well as efforts by key organizations such as the Council for Exceptional Children and the Response to Intervention Center to assess the quality of instructional practices to determine if they are of high enough quality to be considered evidence based. In recent years, various researchers have used scoring rubrics to judge the quality of the evidence in support of specific practices (e.g., Chard, Ketterlin-Geller, Baker, Doabler, & Apichatabutra, 2009, who assessed the research in support of repeated reading, and Jitendra, Burgess, & Gajria, 2010, who evaluated the research on comprehension strategy instruction). Yet these efforts may not be focusing enough on external validity, particularly population validity, when evaluating the effectiveness and appropriateness of practices for English language learner (ELL) populations. The purpose of this examination is to compile and study the research on reading interventions to determine who the participants are and for whom interventions are being generalized. We consider whether or not the results of studies are being generalized as if they apply broadly to unique populations such as ELLs without having being tested with ELLs.
Population Validity
Establishing external validity by creating a research study that can be generalized from a smaller group of respondents within a certain set of conditions to a larger population is paramount for educational researchers (Bracht & Glass, 1968). External validity refers to the degree to which research findings developed in one contextual setting and time with one particular population or culture can be generalized to other contextual settings, times, populations, or cultures (Fyans, 1983; Parker, 1990). Addressing external validity ensures that the success of the reading intervention being studied can be adequately translated from the sample population of participants to different populations of students (see Note 1).
According to Bracht and Glass (1968), population validity grants researchers the ability to gain knowledge about a larger target population by studying a small section of that population. Kempthorne (1961) discussed the need for researchers to carefully examine the characteristics of both the target population and the experimentally accessible population. The target population, as defined by Kempthorne, is the larger group of persons with whom the researcher is interested and to whom the researcher intends to generalize findings. The experimentally accessible population is a smaller and more readily available version of the target population (Kempthorne, 1961). Ensuring population validity allows the researcher to make two leaps in generalizations (Bracht & Glass, 1968). First, the researcher can generalize findings from the sample to the experimentally accessible population from which the sample was drawn. Second, research findings can then be generalized from the experimentally accessible population to the larger target population (Bracht & Glass, 1968). These leaps in generalizations cannot be made until the researcher has gained a thorough knowledge of the demographic characteristics of both the experimentally accessible population and the target population (Bracht & Glass, 1968). Without this level of examination, the results of the experiment can pertain only to the sample population and cannot be generalized to the larger population (Bracht & Glass, 1968).
To ensure threats to population validity are adequately addressed, researchers must clearly describe the characteristics of both the sample and the target populations (Bracht & Glass, 1968; Lysynchuk, Pressley, d’Ailly, Smith, & Cake, 1989; Troia, 1999). This description must include all demographic information such as gender, race/ethnicity, language use, socioeconomic status, and the geographic locale from where the sample was drawn (Troia, 1999). Artiles, Rueda, Salazar, and Higareda (2005) also argued for researchers to describe and disaggregate results for subpopulations of ELLs because of the diversity within ELL populations (e.g., race/ethnicity, language spoken, language proficiency levels) and because the term ELL is often not clearly defined or specifically described in much educational research. Pedraza and Rivera (2005) similarly noted that much research on Latino/a populations emphasizes language proficiency with little regard given to other aspects of diversity (e.g., recent immigrant, long-term American citizen, Mexican, Puerto Rican, Cuban, Central American, South American, Caribbean). There is no doubt that demographic characteristics of target, accessible, and sample populations, the “variables” of population validity noted above, influence the effects of interventions. For example, Charity, Scarborough, and Griffen (2004) examined how racial difference, particularly in terms of familiarity with school English, had an impact on the reading achievement of African American children. E. C. Crowe, Connor, and Petscher (2009) conducted a comparative analysis of evidence-based core reading curricula for children from lower socioeconomic households, noting differences in reading growth based on the variable of socioeconomic status (SES). Rogoff, Paradise, Mejía Arauz, Correa-Chávez, and Angelillo (2003) discovered differences in how children learn through participation with adults across cultures (European American, indigenous communities around the world) with schools in the United States privileging European American participation patterns, Finally, Gutiérrez and colleagues (e.g., Gutiérrez, Asato, Santos, & Gotanda, 2002; Gutiérrez, Baquedano-López, & Tejeda, 1999) have consistently documented the impact of English-only pedagogy on ELLs while calling for more robust research in developing effective classrooms for both mono- and multilingual students. In educational research, it can be challenging to define or describe target and accessible populations. Students in general are the usual target population for educational research, yet demographic characteristics of students vary widely across schools, districts, states, and the nation. Equally problematic for educational researchers is the accessible population. Acquiring an accessible population that is representative of a specific target population may be difficult if the available accessible population has demographic characteristics that differ from those of the intended target population. However, these threats to population validity can be minimized if researchers specifically define for whom they intend the research to be applied and if they purposely search for an accessible population that matches the demographic characteristics of the population of students for whom they are interested.
Calls to focus more on population validity in special education research are not new. As long ago as 1978, Keogh, Major, Reid, Gandara, and Omori suggested that researchers use a system of marker variables to address concerns about the lack of uniformity in the reporting of educational research study samples. A marker variable, as defined by Bell and Hertz (1976, as cited in Keogh et al., 1978), is not the specific focal variable of the study but a background variable that is sufficiently relevant to what is being measured. Without the uniform and consistent use of marker variables in the field, Keogh et al. argued that individual researchers would define their samples differently, creating data that would be inconsistent and problematic to replicate: “Wide variability in children’s characteristics predictably leads to a variability of program effects” (p. 6). When researchers neglect to use descriptive marker variables, such as gender, age, race/ethnicity, SES, and language of the student, the data are not empirically anchored, and it is uncertain which background variables are influencing the focal variables being measured (Keogh et al., 1978). Furthermore, marker variables of importance for ELLs or Latino/a populations must consider diversity within ELL populations and describe how language proficiency is determined (Artiles et al., 2005).
In 1999, Troia noted that a certain degree of methodological rigor is necessary to be able to interpret and generalize the findings of experimental studies. In particular, Troia looked at threats to population validity in terms of participant selection and description, specifically defining the following categories: how participants are selected, age, grade, gender, race, SES, locale of participants, IQ, achievement status, history in special education, and disability criteria. Each category was weighted from 1 (an unlikely threat to external validity) to 3 (a serious flaw and major threat to validity). Findings indicated that researchers’ neglect to address threats to external validity created potentially fatal flaws in the research of instructional interventions, thus limiting what could be understood about what works for struggling students. Troia’s research, however, did not consider language use or proficiency.
More recently, researchers in the field of special education have called for the development of guidelines for conducting and evaluating the quality of research to ensure that instructional practices emerging from research are evidence-based (Odom et al., 2005). Gersten et al. (2005) developed a set of quality indicators for this purpose, proposing the use of weighted criteria for evaluating special education experimental and quasi-experimental research, but they stopped short of creating a rubric for use in evaluating studies. They deemed certain indicators within experimental and quasi-experimental studies as essential for the study to be determined evidence based, whereas they considered other indicators to be desirable but less crucial.
Building on this initial work with quality indicators, several studies designed quality indicator rubrics to apply to reviewing specific special education research interventions. Montague and Dietz (2009) evaluated cognitive strategy use in mathematical problem solving with a rubric noting whether studies conformed to the quality indicators developed by Gersten et al. (2005) or did not. Chard et al. (2009) conducted a review of research on repeated reading, creating a 4-point rubric with three categories under description of participants: information about diagnosis of disability or difficulty, comparability of samples across conditions, and information about comparability of the intervention across conditions. Using the rubric developed in the Chard et al. (2009) study, Baker, Chard, Ketterlin-Geller, Apichatabutra, and Doabler (2009) evaluated the quality of research conducted on self-regulated strategy development in writing. Jitendra et al. (2010) conducted a review of research on cognitive strategies by creating a 3-point rubric with three categories for participants and setting: description of participants (i.e., age, gender, IQ, disability, diagnosis), how participants were selected, and descriptions of the setting. Although the use of quality indicators as a tool for measuring the quality of special education research is relatively new, only the Jitendra et al. (2010) study specifically used ELL status as an important demographic marker for researchers to include.
Historically, researchers have not always described variables that reflect the diversity of students with disabilities. For example, Artiles, Trent, and Kuan (1997) conducted a review of 22 years of empirical research published in special education journals and noted a substantial disregard for diversity issues. They argued that the proportion of published articles that addressed issues of diversity was “alarmingly low,” despite the publication of several special issues on ethnic minority students. They also suggested that many educational researchers draw conclusions about the performance of culturally and linguistically diverse students without attending to external validity or without designing research that is sensitive to cultural differences. The Artiles et al. (1997) study was recently replicated with research conducted from 1995 to 2009, with findings suggesting there has been a slight increase in the number of studies specifically reporting findings on ethnic minority students, yet the majority of studies conducted during this time span continued to fail to disaggregate results for many subpopulations of learners (Vasquez et al., 2011). In analyzing research conducted with ELLs, Bos and Fletcher (1997) and Artiles et al. (2005) found a scarcity of research on within-group diversity among ELLs, noting the neglect of many studies to adequately describe the demographics and language proficiency of ELL populations. Yet relative language proficiency is an important variable that can and does affect treatment outcomes (Ortiz, 1997).
In this article, we focus specifically on the population validity of reading intervention research and how it is being applied to ELLs. We look to see if recent reading research articles include more information about the language proficiency of participants than did articles of the past. We also note whether authors overgeneralize their findings as if they apply to ELLs even when ELLs were not included in their samples. We emphasize that it is not enough to find out “what works” in a general sense; it is crucial to know what reading intervention practices work for whom (Artiles et al., 2005; Gandara & Bial, 2001; Klingner & Edwards, 2006). We argue that the results of reading intervention research studies that do not include ELLs in their sample populations should never be generalized to ELLs.
English Language Learners
ELLs are the fastest growing segment of the U.S. school-age population (National Clearinghouse for English Language Acquisition [NCELA], 2007). About two thirds of ELLs come from low-income families (NCELA, 2007). Although ELLs are often thought to be immigrants, most ELLs are born in the United States and are often second- or third-generation U.S. citizens. Many become long-term ELLs, students who are not reclassified as fluent in English even after 7 years (Menken & Kleyn, 2009). ELLs may have adequate conversational skills in English but lack the vocabulary and academic language needed for success in school. They generally score lower on academic achievement tests than their fluent English peers (Hemphill & Vanneman, 2011). The majority of ELLs speak Spanish as their first language (73%; NCELA, 2011). On the National Assessment of Educational Progress, reading scores have increased significantly for both White and Hispanic students, but the achievement gap between Hispanic and White students has not changed for fourth or eighth graders when comparing 1992 to 2009 data. The gap in 2009 for fourth graders was 25 points, and the gap for eighth graders was 24 points (Hemphill & Vanneman, 2011). Only a very small percentage of eighth-grade ELLs tested as proficient in reading (Aud et al., 2011). Research that is designed to examine reading interventions being used in response to intervention models in schools must be cognizant of these gaps in reading achievement. When interventions that have not been validated with ELLs are implemented broadly in schools, ELLs may be erroneously identified and placed into special education simply because the interventions are not meeting their language and learning needs (Orosco & Klingner, 2010).
Although identification and placement into special education may be beneficial and necessary for children with disabilities, it can be stigmatizing and lead to inequitable educational opportunities, particularly for students who are misidentified. The evidence suggests that culturally and linguistically diverse students are more likely to be overrepresented in certain special education categories compared with their White, English-dominant peers (Donovan & Cross, 2002). Older ELLs who are likely to be long-term ELLs may be those most at risk of overrepresentation in special education (Artiles et al., 2005). Furthermore, culturally and linguistically diverse students are more often placed in separate classrooms from their White, English-dominant special education peers (Fierros & Conroy, 2002; Parrish, 2002), thus missing essential exposure to general education curricula (Donovan & Cross, 2002). The broad application of “evidence-based” research, particularly research that has failed to address threats to population validity, places ELLs at risk of misidentification and overrepresentation into special education categories. Below we outline the response to intervention model, specifically considering the reading needs of second language learners.
Response to Intervention and Reading Interventions for English Language Learners
The response to intervention (RTI) model originated in part from the recommendations of the National Research Council report that questioned the validity of the existing special education classification system and expressed concerns about the disproportionate representation of culturally and linguistically diverse students in special education (Donovan & Cross, 2002; Individuals with Disabilities Education Improvement Act, 2004). Vaughn and Fuchs (2003) described RTI as a three-tier prevention model for identifying and assisting children with possible learning disabilities. The first tier involves evidence-based quality instruction in the general education classroom for all students. The second tier of instruction and assessment is introduced for a student if that student struggles to make gains in the general education curriculum. In the second tier, the struggling student receives more explicit instruction through the use of evidence-based interventions in addition to the quality instruction in the general education classroom. If this student fails to respond to this second layer of intervention, the student moves to the final tier of intervention and assessment. This third tier involves assessment for and placement in special education. When a student fails to respond to intervention at this level, it is a signal of a possibly disability, and he or she is often identified and placed into special education. Furthermore, reading intervention research that provides evidence of success has the potential to influence instructional practices as it is scaled up for use in broader settings, specifically in general education Tier 1 classrooms (Denton, Vaughn, & Fletcher, 2003). Consequently, it is imperative that reading intervention research attend to threats to population validity because of the implications of scaled-up, broadly applied intervention practices on linguistically diverse student populations.
Before identifying and placing an ELL in special education, the student must have received quality, evidence-based reading intervention instruction (Vaughn & Fuchs, 2003). This notion that instruction must be based on scientific evidence, or research, is a central one. Students who fail to respond to evidence-based general education literacy instruction should receive more intensive, explicit, evidence-based reading interventions prior to placement in special education (Torgesen et al., 2001). Placement into special education should come only after the student has received evidence-based reading interventions that have been validated on similar struggling students and it has been demonstrated that students have received an adequate opportunity to learn (Klingner & Edwards, 2006). However, this evidence-based reading intervention must be validated with students similar to those with whom it will be used. Drawing conclusions about the success of a reading intervention for ELLs without carefully validating the practice with similar students is problematic (Artiles et al., 1997). In the next section, we provide an overview of learning to read in English as a second language to highlight why validating reading interventions with ELLs is relevant.
English Language Learners and Learning to Read in English as a Second Language
Although elements of generic good teaching apply to ELLs and learning to read in English as one’s first language in many ways resembles learning to read in English as a second or additional language, there also are important differences that should be considered (August & Shanahan, 2006; Goldenberg, 2008). A common misconception applied to ELLs is that learning to read is simply learning to read, regardless of the language the student speaks predominantly. Reports on reading such as that by the 1998 National Research Council perpetuate this idea (National Institute of Child Health and Human Development [NICHD], 2000). The five major reading components emphasized in the report, phonological awareness, fluency, comprehension, vocabulary, and word study, have been at the center of education policies such as Reading First and No Child Left Behind, suggesting that reading instruction and interventions should rely heavily on these reading components to foster reading development in all children. Although the report notes that it “did not address issues relevant to second language learning” (NICHD, 2000, p. 3) and the panel omitted research with ELLs from their synthesis, recommendations are generalized as if they apply to all populations. More recently, the National Early Literacy Panel (NELP, 2008) report noted the limited number of studies focusing on specific subpopulations, such as ELLs, and recommended making some practices effective with monolingual English-speaking children (i.e., code-focused instruction in phonics) “available to all populations of young children at least until research more directly addresses this question” (p. 120). This is problematic because policy recommendations such as these may be erroneously taken as fact that “what works” in reading instruction and intervention for some works for all. Although the NRC and NELP reports acknowledge the limited knowledge base of effective reading instructional practices for ELLs and also advise that generalizing findings from existing studies to ELLs is recommended only until more definitive research emerges, we believe that there are other options available to practitioners that better take into account what is known about supporting young ELLs’ language and literacy development. Overemphasizing phonics and isolated word reading and underemphasizing other components of literacy, such as vocabulary, comprehension, and oral language, can contribute to a gap between ELLs’ word reading and these other skills (Crosson & Lesaux, 2009; Mancilla-Martinez & Lesaux, 2011a, 2011b).
The National Literacy Panel on Language-Minority Children and Youth (August & Shanahan, 2006) summarized the existing body of research on learning to read in a language other than English and recommended that ELLs receive instruction in the same components of reading as their monolingual peers, but with some key differences. They noted, “Instructional approaches found to be successful with native English speakers do not have as positive a learning impact on language-minority students” (August & Shanahan, 2006, p. 10). This may be particularly true for comprehension instruction. ELLs appear to need more support with oral language and vocabulary than their fluent English peers to benefit from comprehension strategies. Along with other scholars in the field, August and Shanahan (2006) suggest differences necessary for ELLs and recommend more research specifically targeting the needs of ELLs (Goldenberg, 2008; Klingner, Artiles, & Barletta, 2006).
Thus, it is the premise of this examination that before reading interventions can be adopted, scaled up, and broadly used with ELLs, they must clearly demonstrate their success with ELLs, and must also explicitly state to whom the findings can be applied. In particular, reading interventions either need to be found to apply broadly to diverse populations because they have been tested with ELLs and the data have been disaggregated for such populations or must have been specifically developed for ELLs, tested with ELLs, and found to be effective for ELLs. Threats to population validity can be reduced by attending to descriptive marker variables and directing attention to the demographic characteristics of the sample, including the accessible and target populations, and by openly discussing how the research findings can be generalized from the sample population to the target population.
Although many evidence-based reading interventions have been deemed successful generally, a question remains: Do they work as well for subpopulations of students, particularly for ELLs? The purpose of this article is threefold. First, we examine the research conducted on reading interventions since 2001 being used with struggling or at-risk elementary students with the intent of determining how well these research studies describe their sample and target populations to determine if they are adequately addressing threats to population validity. Second, we seek to determine the extent to which researchers are conducting reading interventions research with ELLs as a means toward building a better knowledge base of effective reading for ELLs. And finally, we look specifically to whom researchers are generalizing their findings and whether they appear to be overgeneralizing. This examination of reading intervention research addresses the following main question:
How generalizable are the research findings from reading intervention studies for subpopulations, especially ELLs? In particular, this examination looks at the following subquestions: (a) How are threats to population validity addressed in each study (i.e., target, accessible, and sample populations defined and described)? (b) How many reading intervention studies include ELLs in their samples or are specifically targeting ELLs, and how many disaggregate their findings for ELLs? and (c) To whom are the reported findings being generalized?
Method
Data Collection
For this examination of reading intervention research, we conducted a three-step comprehensive search of the literature from the years 2001 to 2010 (Cooper, 1998). The year 2001 was selected because RTI models emerged in the early part of the decade in response to the problems with using IQ discrepancy as an identification and placement method for special education (Vaughn & Fuchs, 2003). RTI reading interventions differ from earlier reading interventions in terms of being directed toward struggling general education students and not necessarily students already identified and placed into special education settings. This is relevant for our synthesis as we are seeking to understand how interventions are being studied and to whom the findings are being applied. We specifically considered reading interventions that target students who are not making adequate progress within the general education curriculum, seeking to understand to whom these interventions are being generalized. Consequently, we searched for research articles that targeted smaller groups of students as opposed to whole-group, general education instructional approaches.
Initially, the search was broad, with the purpose of locating all possible articles, and then moved to a more selective search of research articles meeting the criteria for inclusion in this examination of reading intervention research. The first step involved a systematic electronic search using the ERIC (Cambridge Scientific Abstracts), PsycINFO, SAGE Journals, MEDLINE, and Wilson Web online databases. The following descriptors and keywords were used to capture any article pertaining to students who are at risk of developing or who have been identified as having a reading disability: reading instruction, RTI, reading intervention, reading remediation, reading disability, struggling reader, and at-risk reader. The term English language learner was also used as a descriptor to seek additional articles specifically addressing the needs of ELLs and not initially found using the above descriptors. Along with relevant articles from the initial search, a list of the 11 most frequently cited journals containing articles matching the descriptors and keywords was compiled. This initial search revealed 2,875 articles matching the search terms. We cast our net wide in this search by not limiting our subject matter to education only.
The second step involved a hand search of the following journals: Annals of Dyslexia, Elementary School Journal, Exceptional Children, Journal of Educational Psychology, Journal of Learning Disabilities, Journal of Special Education, Learning Disability Quarterly, Learning Disabilities Research & Practice, Reading Research Quarterly, Remedial and Special Education, and Scientific Studies of Reading. The hand search consisted of first searching within each journal using EBSCOhost Academic Search Premier’s online tool and the descriptors and keywords and then following up by searching through the table of contents of each issue within the above journals from the years 2001 to 2010 to ensure that no article was overlooked. This second search revealed 580 articles, many of which were already located in the initial search.
The third step involved searching the reference lists and footnotes from relevant reading intervention studies to locate additional research articles that had not previously been located. We were then able to narrow down the relevant articles to 67 articles that met the following criteria (see Table 1).
Studies Synthesized Showing Target Populations, Sample Population Marker Variables Noted, and How Findings Were Reported.
Note: CLD = culturally and linguistically diverse; ELLs = English language learner; LD = learning disability; RD = reading disability; RTI = response to intervention.
Specifically noted the exclusion of ELLs from their sample.
1. Research studies of reading intervention
The articles included in this synthesis are research studies of a reading instructional intervention specifically targeting the reading needs of students not making growth. Intervention, therefore, pertains to additional reading support to supplement the regular reading curriculum. Studies described as “intervention” but that were applied to whole classrooms of students were excluded (e.g., Silverman, 2007a). Although we started our search with research from 2001, marking the beginning of RTI, we did not exclude intervention studies not referring to RTI. Articles describing the inception of the RTI model, ways to use the model, and merits of or problems with the model were not included in this synthesis (e.g., Hollenbeck, 2007; Kavale, Holdnack, & Mostert, 2006; Vaughn & Fuchs, 2003). Research studies describing predictor variables and/or the identification of children with possible learning disabilities were also not included (e.g., Chard et al., 2008; Fuchs, Fuchs, & Compton, 2004). Essays, literature reviews, meta-analyses, book reviews, and editorials were not included.
2. Studies conducted in the United States
Because our focus in this synthesis is on the population validity of reading intervention research used in RTI models of identification and placement into special education, we specifically selected studies conducted in the United States. Research studies examining reading difficulties in children located in other countries were eliminated (e.g., Iversen, Tunmer, & Chapman, 2005; Jiménez et al., 2003). Studies conducted with ELLs in other countries were also eliminated (e.g., D’Angiulli, Siegel, & Maggi, 2004; Lovett et al., 2008).
3. Participants
Included in this article were reading intervention studies using participants in elementary grades (kindergarten through 5th grade). These grade levels are relevant because students who struggle to develop successful reading skills and strategies in the elementary years have continuing academic problems when they move into secondary grades (NICHD, 2000). Studies involving research that focused on secondary level (6th grade through 12th grade) reading intervention strategies were not included. Because the focus of this synthesis was to determine if the needs of ELLs are being met, studies were not limited to monolingual, English-speaking students. Studies focusing on reading interventions targeting ELLs and/or limited English proficient populations were included in this synthesis.
4. Research design
Articles included experimental and quasi-experimental research studies using treatment/comparison designs or single-group designs. As a research methodology, experimental and quasi-experimental studies are often taken as the “gold standard” in research, and findings from such studies are most often broadly applied to other populations as “what works” (Klingner & Boardman, 2011). Consequently, we were interested in understanding how such studies applied their findings. Small case studies with fewer than 20 participants were not included in this article (e.g., L. K. Crowe, 2005).
Data Analysis
Coding procedures
To answer our research questions, we developed comprehensive coding procedures for organizing relevant information from each study. Our interest was in examining the population validity of reading intervention research particularly for ELLs. We did not synthesize the findings or effect sizes of any of the reading intervention studies examined in this article. Instead, as we read each study, we added information to a coding spreadsheet that included the following two major variables: population information and reported findings.
Within the population information variable, we coded studies based on the demographic data or descriptions given about target, accessible, and sample populations to determine if the studies addressed population validity (Bracht & Glass, 1968) or used marker variables (Keogh et al., 1978). The marker variables of interest to us in this study pertained specifically to population validity and how researchers addressed and described their ELL populations (see Artiles et al., 2005) and are described in more detail below (e.g., race/ethnicity, ELL status, how ELL status is determined by the researchers). In this variable, we coded studies in three steps. First, we determined and coded the target population as noted in the article title, abstract, and problem and purpose statements (e.g., at-risk first grade students, struggling kindergarteners). Next, we coded the accessible population in two ways: (a) if the study provided demographic information on race/ethnicity and (b) if the study provided demographic information on ELLs. Finally, we coded the sample population in the following manner: (a) numbers provided for students by race/ethnicity (i.e., White, Hispanic, Black, Asian, Other), (b) number of ELLs in the study, and (c) how ELLs were described (e.g., languages spoken, how language proficiency was determined).
In the reported findings variable, we looked specifically to whom the studies attempted to apply their findings. We did not attempt to synthesize findings in terms of effect sizes or particular intervention strategies to use for ELLs. For the purpose of our examination, we coded studies in three steps. First, we noted the level of disaggregation of results for their sample populations, which revealed four codes: (a) “no—no demographics” for studies that did not provide any demographic information about their sample populations, (b) “no—no ELLs in study” for studies that did not disaggregate their findings based on the demographic information provided about their sample population and did not note ELLs in their sample, (c) “no—have ELLs” for studies that did not disaggregate their findings based on the demographic information provided for their sample population but did have ELLs noted in their sample, and (d) “yes—for ____” for studies that did disaggregate their findings for a specific population (e.g., ELLs, Hispanic). Next, we extensively searched and coded the discussion sections of each article and recorded the specific wording used in the study to determine to whom (the target population) findings were applied. Finally, we compared the wording used in studies to their stated target populations to determine if the study was applicable to ELLs.
Analysis procedures
Our purpose in this examination of reading intervention research was not to determine quality of research in terms of the work being done by Gersten et al. (2005) and quality indicators nor to synthesize the results or suitability of the studies we found that did consider the needs of ELLs, but instead to determine how studies addressed threats to population validity and to determine how generalizable findings are for ELL populations. Thus, we did not weight variables as we coded. Instead, we tallied each category entered into our coding spreadsheet. For the population information variable we first collapsed the target population descriptions found in studies into two broad themes: general and specific. General target population descriptions consisted of “at risk” or “struggling” and often contained a grade level. Specific target population descriptions included “ELL,” “Hispanic,” and “Native American.” These were then tallied and noted. For the accessible population category, we tallied the number of studies in each category (e.g., included race/ethnicity demographics, included ELLs demographics) to determine how many studies included neither category, only one category, or both categories. Finally, we tallied the number of studies that provided race/ethnicity of their sample populations, the number that provided the number of ELL students in their study and the number of studies providing information about their ELL sample populations in terms of language proficiency and languages spoken.
For the reported findings variable, we tallied the number of studies in each sample disaggregation coding category (e.g., no—no demographics, no—no ELLs in study, no-have ELLs, yes—for ELLs). We then compared wording in the discussion sections (to whom they applied findings) to stated target populations (e.g., at risk, struggling, ELLs) and tallied the number of studies that could be generalized to ELL populations.
Reliability
The accuracy of coding was determined through interrater agreement. All coding was initially done by the first author with the second author randomly selecting 50% of the articles to cross-check to determine if the article met the search criteria and to verify the codes obtained and recorded in the spreadsheet. Interrater agreement of 100% was obtained on coding categories requiring tallying (e.g., number of studies noting geographic locale, providing information on number of ELL participants). Discrepancies that emerged in coding the findings section were the result of lack of clarity in the studies included in this synthesis, and such disagreements were discussed by the authors and resolved to meet interrater agreement of 100%.
Results
We located 67 intervention studies published since 2000 with students considered to be struggling or at-risk readers. We present our findings with our research questions as an organizational framework. First we describe how authors addressed population validity. Then we discuss whether they included ELLs in their studies. We finish by noting to whom authors generalized their findings.
Population Validity
To determine if the findings stated in reading intervention studies are appropriate for larger or different populations of struggling students, particularly ELLs, the studies must have addressed threats to external validity. The first subquestion in this synthesis was as follows: How are threats to population validity addressed in each study (i.e., target, accessible, and sample populations defined and described)? A study that has taken into consideration the target, accessible, and sample populations and described them in detail by stating the number of participants and their gender, SES, and race/ethnicity and described the demographics based on language use of the sample population can more readily be replicated and can generalize findings to other similar populations of students ensuring that other struggling readers can be expected to respond to the intervention in a similar manner as the sample of participants responded.
Target populations
Of the 67 reading intervention studies reviewed in this article, 32 referred to their target populations in general terms, failing to state specifically to whom they intended their findings to be applied. A total of 26 studies described their target populations as “at-risk,” “struggling,” or “low-achieving” elementary students. Six studies described their target populations as “below grade level,” students with “language difficulties,” or students who demonstrated poor spelling, fluency, or phonemic awareness.
The remaining 35 studies described their target populations in more specific terms. Eight described their target populations as students whose reading scores fell in the lowest quartile or students who had not previously responded to interventions. Two studies specifically targeted students who had previously been identified as having reading disabilities. Three studies noted their target populations as urban students. One study specifically targeted Spanish-speaking students (Linan-Thompson, Bryant, Dickson, & Kouzekanani, 2005). Three studies defined their target population as Hispanic students, with all studies drawing their samples from predominantly Hispanic classrooms and intentionally targeting similar populations of students (Calhoon, Otaiba, Greenberg, King, & Avalos, 2006; Gunn, Smolkowski, Biglan, & Black, 2002; Gunn, Smolkowski, Biglan, Black, & Blair, 2005). Finally, 18 studies targeted ELLs struggling after instruction in the elementary grades as their population of interest. Demographic characteristics of the ELL target populations were described for these studies mostly as students whose primary language is a language other than English or students who were learning English as a second language. Table 1 lists the studies synthesized in this article and denotes the populations indicated within each study.
Accessible populations
The experimentally accessible population, as defined by Kempthorne (1961), is a smaller representative version of the target population that is more readily available to the researcher. For each of the studies included in this article, the accessible population generally included at-risk elementary students attending elementary schools in which the studies were conducted. Of the reviewed intervention studies, only 15 provided adequate demographic information about race/ethnicity and ELLs in their accessible populations. These studies can more readily make leaps in generalization from their samples to their accessible populations. In all, 35 studies provided no demographic information about race/ethnicity and no information about ELLs in their accessible populations. Consequently, it is difficult to determine if their sample populations were similar enough to their accessible population to make a leap in generalization from sample to accessible and therefore to target. Of the remaining 17 studies, some provided only race/ethnicity demographics and some provided only numbers of ELLs in their accessible populations, making leaps in generalization from sample to accessible populations problematic. However, several exceptions should be noted (i.e., near the Mexico border: Calhoon, Otaiba, Cihak, King, & Avalos, 2007; predominantly Mexican American and Puerto Rican American: Carlo et al., 2004; primarily Spanish-speaking community: Leafstedt, Richards, & Gerber, 2004).
Sample populations
Of the 67 studies examined, 12 did not use marker variables to note the racial/ethnic demographic characteristics of their sample populations. And 10 studies partially described the racial/ethnic demographics of their populations. Of these studies, 6 did so in vague terms such as “other,” “minority,” or “non-White.” Of the studies reporting partial demographics, 4 specifically targeted Hispanic student populations (as noted above) and described their samples in terms of Hispanic students and “other, non-Hispanic students” (Calhoon et al., 2006; Gunn et al., 2002; Gunn et al., 2005; Wanzek & Vaughn, 2008). Of the 67 studies, 45 did provide demographic information about their sample populations.
Out of the 67 examined, 25 studies either did not denote whom they included or did not include ELLs in their sample populations. Five of these latter studies noted that they intentionally excluded ELLs from their sample populations.
Studies Specifically Attending to the Needs of ELLs
Of the 67 studies included in this examination of reading intervention research, 42 included and specifically noted the ELLs in their sample populations. This represents 63% of the studies we reviewed. Marker variables were used to indicate the language use of sample participants in 22 studies (e.g., Spanish, Hmong, Vietnamese). In all, 8 studies used home language surveys and 14 used oral language proficiency scores (from school or state testing) to determine the ELL status of their sample populations. Of the studies, 22 specifically targeted the reading needs of ELLs (see Table 1).
To Whom Reported Findings Are Being Generalized
The final subquestion for this examination was as follows: To whom are the reported findings applied? To answer this question, we looked at which studies disaggregated their data, to whom authors generalized their findings, and whether or not the findings could be applicable to ELL populations. Differences in the demographic characteristics of students are not necessarily seen by researchers to be as relevant as academic differences of students in term of the importance of conducting reading intervention research. Yet research has consistently indicated that variables of population validity (e.g., race/ethnicity, language use, language proficiency) have an influence on the reading achievement of students (e.g., Charity et al., 2004; E. C. Crowe et al., 2009; Gutiérrez et al., 1999; Gutiérrez et al., 2002; Rogoff et al., 2003). Although the target population of students may be generally considered “at risk,” variability in demographics of “at-risk” populations exists and can have an influence on the outcome of the study. Therefore, providing demographic characteristics of sample populations is necessary, and disaggregating findings informs the field about what works for ELLs.
Out of the 67 studies included in this examination, 4 (coded as “no—no demographics”) did not provide demographic information about the race/ethnicity or ELL status of their sample populations and were consequently not able to disaggregate their findings. These studies attempted to apply their findings in general to “at-risk” students, “struggling students,” or “students with poor decoding.” Yet they did not address threats to population validity, so their findings cannot be replicated and can apply only to the students who participated in their study.
In all, 21 studies in this article did provide adequate demographic information about the race/ethnicity of their sample populations but either did not note the number or presence of ELLs in their sample or simply did not include them. Of these studies, 5 specifically excluded ELLs from their sample. Of these studies, 11 attempted to apply their findings in general terms to “at-risk” students; 5 attempted to apply to “lowest performing” or “struggling” students. The remaining 5 studies attempted to apply their findings to more specific populations, using wording such as “non-responders” or “students with severe reading disabilities.” By providing demographic information about their samples, these studies could be replicated. However, their findings cannot be applied to ELLs because they were either excluded or simply not mentioned.
Of the 67 studies synthesized, 17 noted the inclusion of ELLs in their sample populations but did not disaggregate their findings for ELLs, thus making it problematic to assume that results will be relevant for such populations. Of these studies, 8 attempted to apply their findings to “at-risk” students; 7 tried to generalize their findings to “struggling,” “low-performing,” or “low-responding” students; and 1 noted that their findings could be applied to students with severe reading disabilities. However, 1 study is particularly noteworthy: Reis et al. (2007) specifically targeted urban, culturally and linguistically diverse students, described the race/ethnicity and the language status of their sample populations, and suggested their findings are applicable to other culturally and linguistically diverse, urban populations.
The remaining 25 studies in this article are noteworthy because they did address threats to population validity, used marker variables to describe their sample populations, disaggregated their findings, and applied them to diverse student populations. By doing so, these studies are more readily replicable and their findings can be generalized to other similar populations of students. Three studies noted earlier specifically targeted Hispanic student populations, drew their samples from such populations, and provided and applied findings specifically for those populations (Calhoon et al., 2006; Gunn et al., 2002; Gunn et al., 2005). The remaining 22 of the 67 studies also addressed threats to population validity by either specifically targeting ELLs or by purposely drawing their sample from such populations and providing results for such populations. These 25 studies and the 1 noted in the paragraph above specifically considered the needs of ELLs.
Discussion
In an era of high-stakes testing and accountability for the academic growth of all students, policy makers have raised the bar, calling for instructional and intervention practices that are rigorous, relevant, and “evidence based” (ESRA, 2002). Evidence-based practices that have been deemed successful are then scaled up and applied to broader populations of students through the use of RTI models of identification and placement of struggling students into special education. Researchers testing the reliability and validity of instructional reading interventions prior to their application for broader populations of students must ensure that threats to external validity, more specifically, population validity, have been addressed.
The purpose of this examination of reading intervention research was to take a closer look at reading intervention research studies targeting struggling readers with the intention of determining if the reading intervention strategies being recommended have considered the needs of ELL populations. The literacy needs of ELLs learning to read in English differ from the needs of their monolingual, English-speaking peers and research must take this into consideration. When studies fail to disaggregate their findings based on the subpopulations of their sample, it is difficult to determine how well the intervention works for diverse groups of students. If reading interventions are being scaled up and applied to student populations without having been tested and validated with similar sample populations of students, the needs of struggling students, specifically ELLs, may not be adequately addressed in schools.
It remains significant that 25 reading intervention studies out of 67 in the past 10 years or so attempted to apply their findings to nonspecific at-risk, struggling, low-performing elementary students either without clearly noting who these students are demographically or without disaggregating those results when they did note them. We advocate for the field to know what works for whom before broadly applying intervention approaches. These 25 studies did not provide enough information about their target, accessible, and sample populations to adequately address threats to population validity. Therefore, these studies do not meet the standards of evidence-based research, so their findings can be applied only to their sample populations of students.
Another issue we wish to call attention to is the exclusion of ELLs in a study’s sample population. Among the 67 studies examined, 5 specifically noted that they excluded ELLs from their samples. Of these studies, 4 did not provide a rationale for why ELLs were excluded. The other study noted that a variety of student groups were excluded. The authors stated that such exclusions “were necessary to control for extraneous effects on students’ reading skills once intervention took place” (Case et al., 2010, p. 405). Although it is not our position to argue against such exclusions in educational research, we merely wish to note that when ELLs are excluded from a sample population, the results cannot be applied to other ELLs in learning environments elsewhere.
Of the 67 studies, 17 included ELLs in their samples and could have provided some information on how well the ELLs in their study did in terms of the intervention being studied. Although these studies provided some evidence about reading interventions, this information cannot readily be assumed to work as well with ELLs. By not addressing threats to population validity, findings from these studies are really applicable only to the students on whom the intervention was studied.
Among the 67 studies reviewed in this article, 25 specifically defined and described their target and accessible populations and adequately defined and described their sample populations using marker variables to a point where their findings could be replicable and generalizable to a broader, at-risk population of ELLs. The majority of these studies also heeded the advice of Artiles et al. (2005) to attend to within-group diversity by noting the race/ethnicity, language spoken, and language proficiency of their ELL sample populations.
This article points to a more positive direction in special education research than have past examinations on population validity. As Bracht and Glass (1968), Kempthorne (1961), and Keogh et al. (1978) advised years ago, we seem to be doing a better job considering the merit of defining our target and accessible populations and adequately describing the demographic characteristics of our sample populations. More and more research is taking into consideration the diverse needs of other struggling student populations.
Implications for Practice and for Research
This examination of reading intervention research looked closely at how reading intervention research studies addressed issues of population validity with the purpose of determining how applicable the findings were for other student populations such as ELLs. Issues pertaining to ecological validity, which ensures that the experimental environment that the respondents experience is the environment expected and designed by the researchers (Washington & McLoyd, 1982), were not considered for this article because of space limitations and because our desire was to draw attention to the issue of population validity and how it potentially affects ELL student populations. Future research addressing reading intervention strategies should take into consideration threats to ecological validity as well as threats to population validity.
Although most reading intervention research has not necessarily set the specific goal of targeting ELL student populations, much of the research has searched for ways to help struggling, at-risk elementary readers. Although this is certainly relevant and valuable, it is also imperative that educational researchers recognize and specifically target other, diverse student populations (i.e., based on marker variables such as race, ethnicity, SES, gender) when conducting reading intervention research to help us build a quality, evidence-based foundation of knowledge about what works for different struggling reader populations. And research that does address ELL needs should consider and describe the demographic subpopulations within ELL populations (Artiles et al., 1997).
The introduction of quality indicators into special education research planning and evaluation shows some promise. However, very few of the quality indicators suggested by Gersten et al. (2005) evaluate potential threats to external validity in terms of population validity or applying findings to broader, more diverse populations of students. Their essential quality indictors attend to describing the disabilities of the sample participants and to ensuring that “relative characteristics of participants in the sample were comparable across conditions” (p. 152). The rubric designed by Chard et al. (2009) specifically uses the quality indicators suggested by Gersten et al. Although these are indeed essential to high-quality research, we argue that more specific demographic details about the sample population (i.e., race/ethnicity, language use, languages spoken, oral language proficiency) and more clearly defined and described target and accessible populations will increase the quality of research and propel forward our knowledge about what works for ELL student populations. The rubric designed by Jitendra et al. (2010) accounts more for potential threats to population validity by attending to how research studies describe their populations in terms of race, gender, SES, and ELL status. We suggest that researchers consider ways to address potential threats to population validity by more accurately describing target, accessible, and sample populations through the use of marker variables.
This examination revealed 25 studies that addressed population validity to a point where their findings could be replicable and generalizable to a broader, at-risk population of ELLs. Our purpose was not to synthesize the suitability of the findings in these studies but to draw attention to the ways research addresses threats to population validity, specifically for ELL populations. A future effort should look more closely at the body of research on reading interventions designed specifically for ELLs to determine what seems to be the most effective. Furthermore, a future synthesis of the findings of reading intervention research that does specifically target the needs of ELLs should take into consideration the work of the National Literacy Panel on Language-Minority Children and Youth (August & Shanahan, 2006) to determine the components of reading interventions that are necessary for ELLs.
Finally, consumers of reading intervention research must carefully evaluate how research has been conducted and to which student populations findings can be readily applied. Educators who are savvy consumers of research know to read descriptions of participants and determine which samples seem most like their own. They realize that although some interventions show great potential for the academic success for some struggling readers, differing student populations may not show such gains. When educators of ELLs use standard treatment protocol RTI models or hybrid models that include a menu of acceptable, evidence-based reading interventions from which to choose interventions for particular students or groups of students, they need to make sure that the list includes interventions found to be effective with ELLs like their students.
Conclusion
Threats to population validity can more readily be accounted for when researchers attend to how they describe and define their populations and to whom they apply their findings. Overall, from 2001 to the present, we found several studies that successfully addressed the needs of ELLs, helping build our knowledge base about what works for diverse student populations. We are encouraged that researchers seem to be describing diverse samples in more detail than in the past. However, we remain concerned that in several research studies, the authors did not give adequate consideration to student variation and overgeneralized their findings as applying to students not part of their research populations. It is not enough to ask, “What works?” We must consistently ask, “What works with whom?”
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
