Abstract
For bilingual children, the results of language and literacy screening tools are often hard to interpret. This leads to late referral for specialized assessment or inappropriate interventions. To facilitate the early identification of reading difficulties in English, we developed a method of screening that is theory-driven yet suitable for first-language (L1) and second-language learners of English. We administered five conventional tests (phonological awareness, vocabulary, Wide Range Achievement Test–4 [WRAT-4] spelling, letter identification, rapid naming of digits) to 127 five-year-olds (60 English-L1, 67 Mandarin-L1) about 6 months after they started kindergarten, and used the WRAT-4 word reading score 6 months later as the outcome measure. Consistent with previous research, and with children with reading disabilities defined as below the 25 percentile on the reading outcome, logistic regression revealed that the full set of screening measures predicted reading disability status. However, when each predictor was taken as a single measure, spelling scores provided the best fit in terms of the compromise between sensitivity (.75) and specificity (.73) for an optimal cutoff point. Based on this exploratory study, group-administered spelling tasks could provide an efficient solution to screening difficulties in large classes of bilingual children.
Reading skills are important for intellectual development and academic success. Many kindergarteners are at risk of reading difficulties as a result of general cognitive disabilities, specific language impairments, experiential disadvantages, or simply a lack of exposure to print and the target language (Geva, 2000; Vellutino, Fletcher, Snowling, & Scanlon, 2004; Vellutino, Scanlon, Zhang, & Schatschneider, 2008). Although some children may have long-term difficulties in reading, many first-language (L1) and second-language (L2) learners of English do catch up with their peers through early interventions implemented in preschool and first grade (e.g., Hindson et al., 2005; Jenkins, Hudson, & Johnson, 2007; Lesaux & Siegel, 2003; Vellutino et al., 2008). Vellutino, Scanlon, Small, and Fanuele (2006) have also shown that monitoring at-risk children’s response to intervention (RTI) can help determine if reading difficulties originate from experiential factors or cognitive problems.
These findings are encouraging and they highlight the importance of early screening and intervention for children at risk of reading difficulties regardless of language background. In this article we explore the utility of a new method of screening for kindergarten children that is theory-driven yet practical in bilingual settings. Before presenting the rationale for the study in more detail, we first summarize the elements of a good screening tool, and then describe the basic approaches to screening used in the United States and Europe.
What Makes a Good Screening Tool?
The purpose of mass screening tools is to help teachers and parents routinely monitor the progress of groups of children. In bilingual settings, the heterogeneity in language proficiency thwarts the precision afforded by normative data, but screening tools can still be judged according to technical, theoretical, and practical criteria (see Albers & Glover, 2007).
First, in terms of technical characteristics, screening tools should have high predictive and classification validity. This means, in the absence of intervention, scores attained on a screening tool for reading difficulties should predict the child’s future reading performance status. Predictions of this kind will facilitate judicious allocation of resources for intervention toward the children in need without omissions. However, this kind of classification validity depends on a predetermined cutoff point between normally achieving children and at-risk children. Cutoff points are not arbitrary; they depend on several factors including the proportion of children perceived to have reading difficulties in the population, as well as considerations of cost, time, and likely benefit. In terms of overall accuracy, classification validity represents a combination of two metrics, sensitivity and specificity. In terms of signal detection theory, sensitivity is the ratio of the true positives (hits) to the sum of true positives and false negatives, that is, the proportion of actual positives that are correctly identified as such; specificity is the ratio of the true negatives to the sum of true negatives and false positives, that is, the proportion of negatives which are correctly identified as such. To put these metrics in the context of identifying normally achieving children and at-risk children, the at-risk status of the child is the signal, positives and negatives are the presence and absence of the signal. Thus, true positives will be the correctly identified at-risk children, false negatives will be the at-risk children not identified, true negatives will be the correctly identified normally achieving children (i.e., identified as not at risk), and false positives will be the normally achieving children identified incorrectly as being at risk. Therefore, the sensitivity of the measure refers to the percentage of correctly identified at-risk children, and the specificity of the measure refers to the percentage of correctly identified normally achieving children.
Both sensitivity and specificity vary with the cutoff scores chosen for a particular screening test. A strict (high) cutoff point may identify fewer at-risk children and miss some children who are really at risk, thus reducing sensitivity; a lenient (low) cutoff point may identify more at-risk children, but also misidentify some normally achieving children as at risk, thus reducing the specificity. In view of the necessary compromise between sensitivity and specificity, an objective means of determining a cutoff point is required. The Youden index method, which gives equal weight to both sensitivity and specificity to optimize a test’s differentiating ability, has been widely used by researchers to determine the optimal cutoff point (see Schisterman, Perkins, Liu, & Bondell, 2005, for details).
Second, to have a strong theoretical foundation, a screening tool for reading difficulties should measure constructs that are known to underpin reading development. These constructs include a range of cognitive-linguistic abilities, notably phonological awareness (e.g., Cataldo & Ellis, 1998; Chiappe, Siegel, & Wade-Woolley, 2002; Dickinson, McCabe, Anastasopoulos, Peisner-Feinberg, & Poe, 2003; Kirby, Parrila, & Pfeiffer, 2003), fluency and processing (e.g., Georgiou, Parrila, Kirby, & Stephenson, 2008; Kirby et al., 2003), language proficiency (e.g., Dickinson et al., 2003), and memory (e.g., De Jong, 2006; Swanson, Saez, & Gerber, 2006). The multicomponent model, developed empirically by Vellutino, Tunmer, Jaccard, and Chen (2007), captures the relative potency of different constructs (for reviews, see also Rayner, Foorman, Perfetti, Pesetsky, & Seidenberg, 2001; Scarborough, 2001). Many of these constructs are already in use as screening tools but, as Jenkins et al. (2007) pointed out, the predictive relationships for reading development in a population need not apply to individual children. This is because the collection of items in a screening tool and the cutoff point must be sensitive to the heterogeneity among at-risk children and normally achieving children, especially those with abilities in the middle to lower range. In other words, predictive validity is theory-driven, whereas classification validity is data-driven, and the ideal screening tool should have both.
Finally, in addition to technical and theoretical concerns, a screening tool must be practical for a particular setting in terms of local resources and language diversity. With unlimited time and money, the administration of a multifaceted screening battery (e.g., A. G. Bishop & League, 2006; Catts, Fey, Zhang, & Tomblin, 2001; O’Connor & Jenkins, 1999) or published screening tests for reading abilities (e.g., Dyslexia Early Screening Test—Nicolson & Fawcett, 2004; Phonological Awareness Literacy Screening–Pre-Kindergarten; Invernizzi, Sullivan, & Meier, 2001) appear preferable to a simple screening tool. These batteries tap into a range of constructs that underlie reading development and provide converging evidence from different measures. However, multifaceted batteries normally require at least 1 hour of one-on-one individual testing, and they often need to be administered by speech and language therapists or educational psychologists. This may mean that only children whose reading difficulties have already been noticed by their teachers or parents are assessed. In large heterogeneous classes of bilingual children, an easy-to-administer screening tool is needed to ensure that referrals for detailed assessment can be made with confidence by teachers.
Approaches to Screening
In the United States and Europe, two main approaches have been used for screening: RTI (e.g., Vellutino et al., 2006; Vellutino et al., 2008; see review by Jenkins et al., 2007) and multifaceted assessment (e.g., in the United States: A. G. Bishop & League, 2006; Catts et al., 2001; O’Connor & Jenkins, 1999; in Europe: Gijsel, Bosman, & Verhoeven, 2006; Puolakanaho et al., 2007). These approaches differ in terms of the scale of screening and intervention programs, as well as methods of identifying children with reading difficulties. Within the framework of RTI, only one brief test is administered as a screening measure during the first year of the schooling; for example, a letter identification task was used by Vellutino et al. (2006, 2008) for Tier 1. In multifaceted assessment, at least three tests are administered in the first or second year of school, but the choice of tasks is not uniform. A. G. Bishop and League (2006) found that a combination of letter identification, phonological awareness, and rapid naming made the most parsimonious screening battery; Catts et al. (2001) showed that letter identification, sentence imitation, phonological awareness, rapid naming, and mother’s level of education provided the best set of predictors; and O’Connor and Jenkins (1999) found the best screening battery consisted of two phonological tasks (phoneme segmentation and syllable deletion) and rapid naming.
These differences suggest the choice of the single best screening task is not an easy decision, nor is the outcome measure during kindergarten years. For example, Vellutino et al. (2006, 2008) evaluated children at the end of Grade 1 (6 years old) using Word Identification and Word Attack subtests from the Woodcock Reading Mastery Tests–Revised (WRMT-R; Woodcock, 1987) and the Reading comprehension subtest from the Wechsler Individual Achievement Test (WIAT; Wechsler, 1992), whereas O’Connor and Jenkins (1999) used the full battery of WRMT-R. A. G. Bishop and League (2006) evaluated reading disability status longitudinally from Grade 1 through to Grade 4 using the Test of Word Reading Efficiency (TOWRE; Torgesen, Wagner, & Rashotte, 1999) and the Qualitative Reading Inventory–II (Leslie & Caldwell, 1995); Catts and colleagues (2001) evaluated children at Grade 2 (7 years) using WRMT-R, Gray Oral Reading Tests–3 (Wiederholt & Bryant, 1992), and Diagnostic Achievement Battery–2 (Newcomer, 1990). Those assessed using a multifaceted battery are not usually documented for specific follow-up, but under the framework of RTI a three-tier program is sometimes implemented with appropriate screening and outcome measures at each juncture (e.g., Vellutino et al., 2008).
The outcome for the RTI approach is impressive. For example, 84% of the children identified as at risk on entry to kindergarten (i.e., 84% of the lowest 30%) who went through Tier 2 and/or Tier 3 intervention later performed on par with their peers on reading tests by the end of Grade 1 (Vellutino et al., 2008). The authors suggest that the initial low performance of these children might be due to lack of exposure to spoken English, low socioeconomic status, or delayed development, whereas the reading difficulties of those who continued to have problems after intervention (16% of those identified as at risk) are likely to be a result of long-term cognitive difficulties. Children in this residual group would need more intensive long-term support.
Of importance, the improvement in reading of the majority of at-risk children following Tier 2 and 3 intervention attests to the value of mass screening at Tier 1 in kindergarten followed by early intervention. Large-scale screening programs for all kindergarteners are feasible if teachers can administer a single, brief, and literacy-related test in the RTI settings. Multifaceted batteries provide more detailed information about the nature of the child’s difficulties, but they can be administered to only a small number of children by trained speech therapists or psychologists and incur delay. They also have higher predictive validity than mass screening because the concomitant RTI interventions will confound long-term reading disability status. In what follows, we describe how the merits of the two approaches can be integrated for screening bilingual children in large classes.
Rationale for Integrated Screening Tool
In view of the benefits of early large-scale universal screening and intervention reported by Vellutino et al. (2006, 2008), we wanted to identify a screening tool that can be administered to bilingual kindergarten children. Lack of exposure to English is a common problem in multilingual settings, and reliable normative data are often hard to compile. Letter identification has been used as an RTI screening measure in the United States, whereas multifaceted batteries comprising tests of phonological awareness, rapid naming, and memory have proved reliable with unilingual English-speaking children. For our Asian bilingual population, it was not clear whether any of these tools would be reliable or valid for identifying children at risk of reading difficulties because parenting styles, language background, and methods of teaching literacy are different.
Parenting Styles
Regardless of home language, parents often teach their children the English alphabet before they enter kindergarten. This means that scores for letter identification are often close to the ceiling, and the lack of variance would make the task unreliable as an RTI screening tool for many children.
Language Backgrounds
Almost all children in kindergarten classes will have been exposed to one or more Asian languages at home, in addition to colloquial English (Gupta, 1998). The heterogeneity in language backgrounds would lead to differences in the salience of rapid naming (Yeong & Rickard Liow, 2011), phonological representations (Caravolas & Bruck, 1993), and the metalinguistic awareness (Rickard Liow & Lau, 2006) used for subsequent literacy development. So, although phoneme awareness is the best predictor of reading ability for young unilingual English-speaking children (e.g., Byrne, 1998), and is therefore a critical component of multifaceted batteries, the same principle might not apply to English as a second language (ESL) bilingual children.
Teaching Methods
Look-say methods of teaching early reading skills in English are still prominent in Asian classrooms because of the marked differences between the phonological structure of the home language (e.g., Mandarin or Cantonese) and English. Teachers of ESL bilingual children often recognize, implicitly or explicitly, that developing phonemic awareness for decoding is confusing when the mapping between speech and the writing system is so different for the child’s home language and the target language.
The Present Study
These three differences between our bilingual children and children screened elsewhere using either letter knowledge (RTI) or multifaceted batteries suggested that an empirical study was needed to establish a mass screening tool. Based on the model developed by Vellutino et al. (2007), performance on tests of phonological awareness, vocabulary (semantic knowledge), and spelling was selected as likely to be the most useful. Phonological awareness has often been cited as the most important predictor of literacy for English-speaking children (see Rayner et al., 2001), although differences in language background, phonological representations, and teaching methods are known to reduce its impact in young Mandarin-L1/English-L2 children (Yeong & Rickard Liow, 2012). Vocabulary was shown to correlate with real word reading (r = .41) in a meta-analytic review (Swanson, Trainin, Necoechea, & Hammill, 2003), and semantic knowledge might be especially important for reading skills in ESL bilinguals. Spelling has been used infrequently as a screening tool, perhaps because it is perceived as either very similar to reading (e.g., Ehri, 2000; Perfetti, 1997) or too advanced to predict reading. However, Vellutino et al. (2007) argued that spelling fosters “structural analysis of the type that facilitates precision in encoding the letters in printed words, the order in which they occur, and word-specific spellings” (p. 8). In other words, through spelling experience with letter sounds, children establish and strengthen orthographic knowledge, and this has a positive impact on early reading of isolated words.
To these three tasks (phonological awareness, receptive vocabulary, and spelling), we added letter identification because this measure has been used for RTI Tier 1 screening (e.g., Vellutino et al., 2008), and rapid automatic naming, the serial rapid naming of highly familiar visual stimuli such as letters or numbers. Rapid naming has been reported as a strong predictor of reading ability (see Wolf, Bowers, & Biddle, 2000, for a review) and also has proven utility as a screening tool (e.g., Berninger, Abbott, Thomson, & Raskind, 2001; McBride-Chang & Manis, 1996).
This battery of five tasks was administered to two different groups of bilingual children to identify which measure might serve as a reliable predictor of reading ability: children with English as first language and Mandarin as second language (English-L1/Mandarin-L2) and children with Mandarin as first language and English as second language (Mandarin-L1/English-L2). Children from these two language backgrounds attend the same classes and represent about 75% of the local kindergarten population in Singapore, whereas the remainder are from families where Malay or an Indian language is spoken at home. Vellutino et al.’s (2008) Tier 1 screening is used at entry to kindergarten, but we decided to test children 6 months after entry to ensure that they had been exposed to English and had received some formal instruction in literacy skills. Their reading disability status was then determined 1 year after they entered kindergarten (6 months after screening).
No formal intervention program akin to the Tier 2 and Tier 3 of RTI was offered after the screening battery was administered in this trial phase. With the reading disability status as the dependent variable, the predictive validity and classification validity of each of the five screening measures, and as a full set of screening battery, were determined using logistic regression (Catts et al., 2001), as well as the sensitivity and specificity measures. Logistic regression was chosen because children at risk for reading disabilities score in the outlier range on the variables and the distributions would violate the normality assumption of multiple regression.
To summarize, the purpose of the logistic regression analyses was to identify which single measure could be used as a reliable mass screening tool with large classes of young bilingual children. Provided the measure demonstrates adequate sensitivity and specificity, those considered at risk of reading difficulties can then be given early intervention, as in the RTI framework, during their first year in kindergarten, or be referred for further assessment using a multifaceted battery.
Method
Participants
A total of 138 ethnic Chinese children from five government-aided kindergartens were recruited for this 6-month longitudinal study with parental permission. From this pool, there were missing data for 11 children (8%), so the final analyses were based on 127 participants. The data from these 127 bilingual children (60 English-L1/Mandarin-L2 and 67 Mandarin-L1/English-L2) were collected as part of a longitudinal study of bilingual children looking at the development of spelling (Yeong & Rickard Liow, 2011) and phonological awareness (Yeong & Rickard Liow, 2012). Although these children were bilingual in Mandarin and English, they tended to be exposed to one of the languages more than the other, so balanced bilingualism was uncommon. Parents of the 60 English-L1 children (23 girls, 37 boys) reported English as the child’s first language and the language the child had been exposed to since birth. At home the average amount of time that caregivers and parents spoke to these children in English was reported as 73.61%, with Mandarin for 21.21% of the time and other Chinese languages or Malay for 5.18% of the time. Parents of the 67 Mandarin-L1 children (38 girls, 29 boys) reported Mandarin as the child’s first language and the language the child had been exposed to since birth. At home the average amount of time the caregivers and parents spoke to these children in Mandarin was reported as 72.48% of the time, with English 22.37% of the time and other Chinese languages or Malay 5.15% of the time. Both groups of parents also rated their children as having better proficiency in their first language than their second language. At Time 1 screening, the average ages of English-L1 children and Mandarin-L1 were 63.45 months (SD = 3.49 months) and 63.20 months (SD = 3.38 months), respectively. An independent t test confirmed no significant difference between the two groups in age, t(125) = 0.404, p = .687.
Materials
Phonological awareness
Syllable deletion and phoneme isolation were used to measure the children’s phonological awareness. In all, 10 bisyllabic and 15 trisyllabic words were used in the syllable deletion task. For the bisyllabic words, the children were asked to delete either the first syllable or the last syllable (5 items each); for the trisyllabic words, the children were asked to delete the first, middle, or last syllable (5 items each) and say what was left (e.g., say clean bedroom without saying clean). The internal reliability of the syllable deletion task was .81. A total of 10 monosyllabic words with a CVC structure were used for the phoneme isolation. Children were asked to identify the first sound in the word (e.g., what is the first sound in tall?). Cronbach’s alpha internal reliability of the phoneme isolation task was .96. The phonological awareness task was scored by a weighted average of the syllable deletion and phoneme isolation.
Receptive English vocabulary test
At Time 1, the children’s English vocabulary was measured by the Bilingual Language Assessment Battery (Rickard Liow & Sze, 2009), a locally developed test with culturally and linguistically appropriate stimuli. The test comprises 100 target words arranged in order of difficulty. The children were shown four pictures on the computer screen and listened to a single word played over headphones simultaneously. They were asked to choose the picture that best matches the word. The internal reliability of the test using Cronbach’s alpha was .77.
Letter and word spelling
On separate occasions, the blue and green parallel forms of the Wide Range Achievement Test–4 (WRAT-4) Spelling (Wilkinson & Robertson, 2006) were both administered at Time 1 to all children to increase reliability. The children were asked to write their own name, and two points were scored as long as the children could write down at least two letters. For each form, the children then wrote down the letters and the words they heard according to standard instructions. Raw scores from the name writing, letter writing, and word writing of both forms were summed and were used as a predictor in data analysis.
Letter identification
The letter reading part of the WRAT-4 Reading (Wilkinson & Robertson, 2006) subtest was administered as a predictive measure at Time 1.
Rapid automatic naming (RAN)
Four single-syllable digits (2, 5, 4, and 3) were used in the rapid naming. A total of 36 digits were randomly arranged in four rows of nine items. The children were asked to complete two different forms of rapid naming with different arrangements of the same digits. The children were asked to name the digits as fast as they could from left to right, beginning from the top row and moving on until the last row. For each form of the task, the time taken to complete the naming of all 36 items and the number of errors were recorded. Performance was converted to the average time in seconds to name an item correctly, that is, the time used to complete articulating the 36 items divided by the number of correct items named. It was then averaged for the two forms to increase reliability. Test–retest reliability between the two forms was .90.
Letter and word reading
For the outcome measure of reading ability, the children were asked individually to read aloud letters and words from both the blue and green forms of the WRAT-4 Reading (Wilkinson & Robertson, 2006). The raw scores for word reading in both forms were summed as the outcome measure for reading disability status at Time 2.
Results
Classification of Children With Reading Difficulties at Time 2
Following Vellutino et al.’s (2008) methods of identifying children with continued risk after intervention, we classified the status of having reading difficulties according to the children’s single word reading in WRAT-4 at Time 2 when they had just started Year 2 of the kindergarten. However, we used only single word reading as the reading outcome measure, instead of a composite of reading comprehension and fluency because a pilot study suggested there would be floor effects on the latter for our kindergarten children.
The descriptive statistics and t test comparisons for the two language background groups on the screening measures and outcome measure are shown in Table 1. The English-L1 and Mandarin-L1 groups differed significantly on the outcome measure so the reading disability status of the children was determined using their word reading of WRAT-4 at Time 2 with reference to their own language background group. More specifically, those who scored lower than the 25th percentile in their respective language background group on the outcome measure 6 months later were classified as having reading disability (n = 32); the rest were classified as normally achieving children (n = 95). The cutoff of the 30th percentile used by Vellutino et al. (2008) was not followed because the number of children under the 30th percentile is not an integer number in the present study. In terms of predictors and screening measures, the letter identification/naming, which was the screening measure in Vellutino and colleagues’ (2008) study, was at the ceiling for both groups of children in the present study. Up to 86.9% of the whole sample scored at least 13 out of 15 correct in the tasks. Therefore, 30% of the lowest scoring children on this screening measure could not be identified as they were by Vellutino and colleagues (2008).
Means (Standard Deviations) and t Test Comparison Between the Language Background Groups on Time 1 (T1) and Time 2 (T2) Measures.
Note. Phonological awareness = weighted score of syllable deletion (25 items) and phoneme isolation (10 items); rapid naming = average time in seconds to name an item correctly; Wide Range Achievement Test–4 (WRAT-4) Spelling = letter and word writing.
Logistic Regression Analyses
Multifaceted screening—overall model evaluation
Logistic regression with letter naming, phonological awareness, rapid automatic naming, vocabulary, and spelling test as predictors, and reading disability status 6 months later as the outcome was conducted for the whole sample (see Table 2 for the bivariate correlation between the Time 1 [T1] predictors and the Time 2 [T2] outcome reading scores). Separate models for the two language background groups would mean that results would not be useful for teachers unless they have detailed access to reliable information about home language. The model for 127 children was significant, χ2 = 54.856, p < .001, likelihood ratio = 88.524, indicating that this multifaceted battery of five predictors (see Table 3) reliably distinguished between children with reading disabilities and normally achieving children at T2. The logistic regression model also allows the calculation of the probability of a particular child being at risk,
Bivariate Correlations Between Time 1 Predictors and Time 2 Outcome Reading Scores.
Note. Phonological awareness = weighted score of syllable deletion (25 items) and phoneme isolation (10 items); rapid naming = average time in seconds to name an item correctly; it is negatively related to other variables because the shorter the time used, the better the performance in the test; Wide Range Achievement Test–4 (WRAT-4) Spelling = letter and word writing.
p < .05. **p < .001.
Logistic Regression Model With 5 Predictors of Reading Disability.
Note. Phonological awareness = weighted score of syllable deletion (25 items) and phoneme isolation (10 items); rapid naming = average time in seconds to name an item correctly; Wide Range Achievement Test–4 (WRAT-4) Spelling = letter and word writing.
Only phonological awareness, rapid naming, and spelling are significant.
where Vocab, LettRead, RNam, WRATSpell, and PA are the scores of vocabulary, letter reading, rapid naming of digits, WRAT-4 Spelling, and phonological awareness of the child, respectively. The higher the probability score a child attains, the more likely he or she is going to have a reading disability 6 months later (see Catts et al., 2001). Using these probability scores that take into account all five measures in the predictive model, instead of raw scores on any of the measures, a cutoff point could be determined. More details of this cutoff point are provided in the “Sensitivity and Specificity” section below.
Single screening task—statistical tests of individual predictors
Teachers of bilingual children in Asia are unlikely to have sufficient resources for multifaceted screening of all children in large kindergarten classes. For this reason the main aim of the study was to identify which single task among the five we administered might serve as the best predictor of reading disability status 6 months after screening. To make a comparison of effect of each predictor in the overall model, the Wald test, likelihood ratio test, and the odds ratio of each predictor in the model are shown in Table 3. Although the Wald test of the coefficient (B) has often been used to test for the effect of individual predictors in a model, the likelihood ratio test is a more powerful test according to Peng and So (2002). The likelihood ratio statistic is an indicator of “badness of fit”; the higher the likelihood ratio statistic, the weaker a model is. It is based on the differences of the likelihood ratios between the models with and without the predictor of interest. A significant and a larger deviance rise (i.e., ratio difference) when a predictor is dropped suggests that the predictor has a significant and larger effect in predicting the outcome status. Another indicator of effect of each predictor is the odds ratio. When all other predictors are kept constant, the odds ratio indicates the change in probability of having reading disability with one-unit increase in a predictor (Tabachnick & Fidell, 2007). Nonetheless, the size of the odds ratio is not to be interpreted as the effect size of individual predictors; it is a measure of change and is not standardized. In the full model with all five predictors (Table 3), the phonological awareness task, rapid naming, and spelling each contributed significantly to the model, with all ps < .018 for the Wald test and all ps < .012 for the Likelihood ratio test. Of interest, however, the contribution of the spelling test is the largest, with the rise in the likelihood ratio (13.787) the largest if it is dropped out of the overall model.
Given that the main aim of the present study is to look for a single brief screening test, each of the three predictors was subjected to separate logistic regression model as the sole predictor to test if each of the three predictors could still significantly predict the reading disability status without the other predictors in the model. The results for the three models when each task is the sole predictor, together with the relevant equations used to calculate the probability of a child being at risk of reading difficulty, are presented in Table 4. Although all three models were significant in predicting the reading disability status (ps < .001), they differ in the measure of badness of fit, that is, likelihood ratio. Of importance, the model with WRAT-4 Spelling as the sole predictor has the lowest badness of fit, with a likelihood ratio of 103.234.
Comparison of Different Screening Measures as the Only Predictor in a Logistic Regression Model.
Note. Likelihood ratio = badness of fit; phonological awareness = weighted score of syllable deletion (25 items) and phoneme isolation (10 items); rapid naming = average time in seconds to name an item correctly; Wide Range Achievement Test–4 (WRAT-4) Spelling = letter and word writing.
Sensitivity and Specificity
In addition to logistic regression, which facilitated the examination of the predictive validity of the full screening battery and individual tasks, measures of sensitivity and specificity also shed light on the classification validity of the tests. As mentioned earlier, these two metrics vary with the cutoff points; therefore in Table 5, the sensitivity and specificity associated with various cutoff points on the full screening battery (overall model), and the three individual tests that predicted reading disability in logistic regression analyses were listed. The hit, true negative, sensitivity, and specificity are calculated based on the number of at-risk children predicted by the screening and the outcome reading disability status at T2: 32 children with reading disability and 95 normally achieving children at T2. For the full screening battery, the probability score calculated from the logistic model must be used as the score of each child, as it is the composite score from all five tests, whereas for the single screening task as spelling test, rapid naming, or phonological awareness, either the probability score or the raw score of the test could be used as the score of each child because only a single test is assumed to be used to screen the children. The probability score and the raw score of phonological awareness and spelling are inverse; the higher the probability score of a child, the more likely he or she would have reading disability 6 months later, whereas the higher the raw score of a child on the two single screening tasks, the better performance he or she has on the tasks. The probability scores and the raw scores are in the same direction in the rapid naming because the lower the raw scores, the faster they name the item, the better the performance, and so the lower the probability scores of being at risk. This list of metrics would allow teachers or screeners to choose the cutoff point for the screening test that best fit their purpose and circumstances. Those who have sufficient intervention resources may want to have a lenient cutoff to have a higher sensitivity despite a lower specificity, and vice versa for those who have limited resources.
Sensitivity and Specificity Associated With Various Cutoffs on Full Screening Battery and Three Single Screening Tests.
Note. More cutoff points could be listed for the full battery, phonological awareness, and rapid naming because these three are average scores and have more variance than the spelling test. Raw scores of phonological awareness = weighted score of syllable deletion (25 items) and phoneme isolation (10 items); raw scores of rapid naming = average time in seconds to name an item correctly; raw scores of Wide Range Achievement Test–4 (WRAT-4) Spelling = letter and word writing.
Number of children predicted as at risk by the test according to the cutoff. bThe optimal cutoff point was determined by Youden index method.
The cutoff point that was calculated by the Youden index method is highlighted here. The method picks the cutoff point such that both sensitivity and specificity are maximized (for more detailed mathematics and concepts, see Schisterman et al., 2005), and an R package, OptimalCutpoints (Lopez-Raton & Rodriguez-Alvarez, 2013), was utilized to calculate it. This cutoff point could be considered the optimal cutoff when no other circumstances such as costs and benefits of screening and intervention come into play and would be useful for comparison of diagnostic validity of different tests. Under this cutoff, the sensitivity and specificity of the whole screening battery (.81 vs. .87), the spelling test (.75 vs. .73), and rapid naming (.75 vs. .61) seem to be a compromise, but not the phonological awareness (.53 vs. .92). Given that the spelling test has a higher specificity than the rapid naming, whereas both show the same sensitivity under this optimal cutoff, this would be the best single measure for predicting reading status. In other words, for our bilingual children, spelling is a better predictor than either the phonological awareness task or the rapid naming test.
To summarize, when the optimal cutoff point was used to compare the classification validity across the different single screening measures, both the sensitivity and specificity of the spelling test were greater than .70, whereas the sensitivity of phonological task and specificity of rapid naming were less than .70, indicating that spelling is superior to the other two tests, despite being inferior to the full set of screening tests.
Discussion
To find a single screening tool for bilingual children that can be administered by class teachers to identify those in need of early intervention and/or further assessments for reading disability, we compared the predictive utility of five tasks: phonological awareness, vocabulary, spelling, letter identification, and rapid naming. Tests of phonological awareness, vocabulary, and spelling were chosen with reference to Vellutino et al.’s (2007) multicomponent model of early reading and previous research (e.g., A. G. Bishop & League, 2006; Catts et al., 2001; O’Connor & Jenkins, 1999); tests of letter identification (Vellutino et al., 2008) and rapid naming have proven utility as predictors of reading disability (e.g., Berninger et al., 2001).
The battery of five tests was administered to English-L1 and Mandarin-L1 bilingual children 6 months after they had entered kindergarten (aged 5 years 5 months) and then 6 months later, a single word reading test (WRAT-4) was administered as an outcome measure for reading disability status. The reading disability status of each child was determined with reference to their language background because there was a significant difference in the reading scores between the English-L1 group and the Mandarin-L1 group. Those who scored under the 25th percentile in each group (n = 15/60 for the English-L1 group; n = 17/67 for the Mandarin-L1 group) were classified as having a reading disability after 1 year in kindergarten.
The results of the logistic regression analyses showed that together the five screening tests reliably predicted the reading disability status for the total sample of 127 children, with a sensitivity of .81 and specificity of .87 when the optimal cutoff point was used. More specifically, spelling, phonological awareness task and rapid naming reliably predicted reading disability status, either in a full battery with all five tests or as a single test, but neither letter identification nor vocabulary at T1 predicted reading disability status at T2. Past research shows that letter identification is a reliable single measure for RTI (Vellutino et al., 2008) and that phonological awareness and rapid naming are the best predictors when multifaceted batteries are administered (Catts et al., 2001; O’Connor & Jenkins, 1999). Performance on spelling tests has seldom been incorporated into early screening batteries designed for unilingual English-speaking children, even though it appears to predict reading performance (Swanson et al., 2003). However, in the full model, and as a single measure, the WRAT-4 Spelling had the highest predictive validity in the logistic regression model and showed an acceptable compromise between sensitivity (.75) and specificity (.73) when the optimal cutoff point was used. Whenever time and resources permit, a multifaceted battery comprising spelling, phonological awareness task, and rapid naming could be used as a screening tool for bilingual children. However, when a brief universal screening tool is needed for the first stage of RTI, a spelling test is likely to be more reliable than letter identification, and easier to administer than tests of rapid naming or phonological awareness.
Of importance, the results from the multifaceted screening battery employed in the present study with bilingual children differed from those found by A. G. Bishop and League (2006) and Catts et al. (2001) with unilingual children. For their populations, letter identification was one of the predictors of reading disability status whereas for our children, performance on this task was at ceiling. This is probably because parents place emphasis on learning the letters of the alphabet even when the home language is not English. More consistent with theoretical models and research elsewhere (e.g., A. G. Bishop & League, 2006; Catts et al., 2001; O’Connor & Jenkins, 1999), phonological awareness task and rapid naming were found to be reliable predictors in our version of a multifaceted battery. Phonological awareness and rapid naming have long been shown to be predictive of reading achievement in English unilinguals (e.g., Dickinson et al., 2003; Speece, Ritchey, Cooper, Roth, & Schatschneider, 2004) and bilingual children (Lesaux, Rupp, & Siegel, 2007; Lipka & Siegel, 2007). Thus, the main contribution of the present study is the demonstration that spelling ability has more predictive validity than these two tests for bilingual children, both in the full model and as a single measure.
Why Is Spelling a Reliable and Valid Screening Tool?
Here we describe how spelling performance meets the technical, theoretical, and practical requirements of a reliable and valid screening tool. The results of our exploratory study showed that in comparison to phonological awareness task and rapid naming, spelling at T1 is a better predictor of reading disability status at T2. Of importance, in terms of classification validity, when used as a single measure with the optimal cutoff point, spelling has the same sensitivity and a higher specificity than rapid naming, and has a higher sensitivity and lower specificity than phonological awareness. For technical reasons, therefore, spelling might prove to be a better screening tool if a compromise between the sensitivity and specificity is needed.
Theoretical support is growing for the idea that the assessment of early spelling abilities might be a useful metric for predicting subsequent reading difficulties. Prior to the development of Vellutino et al.’s (2007) model, Mehta, Foorman, Branum-Martin, and Taylor (2005) provided evidence that spelling development is closely related to reading development. There have also been other reports showing that spelling scores predict reading in young unilingual children (e.g., Abbott, Berninger, & Fayol, 2010; Caravolas, Hulme, & Snowling, 2001; Cataldo & Ellis, 1988; Ehri, 2000; Foorman & Petscher, 2010) even though they are not often used for screening purposes. If, as Shahar-Yames and Share (2008) suggested, spelling is a self-teaching mechanism for learning to read, any delay in spelling acquisition would have an adverse impact on reading acquisition (see also Frith, 1980). Of interest, recent work on individual differences in the quality of lexical representations in adults provides evidence that spelling proficiency is a useful index of precision (e.g., Andrews & Hersch, 2010). Likewise the outcome of the present study suggests that the concept of lexical quality (Perfetti, 1992) is best captured by early spelling ability.
In terms of practical considerations, spelling tasks can be administered in a group setting and are not susceptible to ceiling effects. Spelling tests also require less specialist knowledge to interpret and score for accuracy than tests of phonological awareness or rapid naming. Paradoxically, although spelling appears to be more cognitively demanding, or a more advanced precise skill than reading, it seems to require complexity (Jongejan, Verhoeven, & Siegel, 2007; Lervåg & Hulme, 2010) and precision (Vellutino et al., 2007), which leads to such good predictive power as a screening tool. Recent work on spelling errors (e.g., Treiman & Bourassa, 2000; Yeong & Rickard Liow, 2011) also suggests that spelling tests can provide valuable insights into underlying processes. Fine-grained analyses of spelling errors would be too time-consuming for routine screening, but once the child has been identified as at risk, spelling responses could be further examined and monitored during RTI. For example, ESL Mandarin-speaking children struggle more with sounds that are not represented in their home language (e.g., /v/ and /b/; see Yeong & Rickard Liow, 2010) as well as consonant clusters.
Remarkably few researchers have compared children’s reading and spelling using the same word pool, but the conventional view is that they are two sides of the same coin (e.g., Ehri, 2000; Perfetti, 1997). Vaessen and Blomert’s (2013) recent cross-sectional study (Grades 1 to 6) of Dutch-speaking children suggests that the relationship between the two skills changes during the course of development. The mappings between orthography and phonology are more transparent in Dutch than in English, but their findings show that early reading and spelling abilities depend on both phonological awareness and the use of letter-sound knowledge. Later in the course of development, the importance of these two underlying processes is maintained for spelling but declines for reading, presumably because the children have established reliable word-specific orthographic representations and can rely on recognition.
Future Research
Future work examining kindergarten children’s reading and spelling of the same pool of English words will shed light on the precise relationship between the two skills, and whether spelling does provide a self-teaching mechanism for learning to read (see Shahar-Yames & Share, 2008). Meanwhile, the paucity of research on spelling development, compared to reading development, is surprising given that it is often thought to be the more difficult task for children at risk of literacy difficulties. Spelling problems are more likely to persist into adolescence (D. V. M. Bishop & Clarkson, 2003) and go untreated (Berninger, Nielsen, Abbott, Wijsman, & Raskind, 2008). Again, the reason for the asymmetry appears to be that reading processes involve more lexical (word-specific) recognition abilities, whereas spelling involves more sublexical speech-based phonological knowledge (Jalil & Rickard Liow, 2008; Treiman, Goswami, Tincoff, & Leevers, 1997) and production abilities. It seems that young children’s relatively weak orthographic representations are often sufficient for successful reading using a whole-word strategy (and guessing), whereas correct spellings require holding the exact sound sequence in working memory while converting the phonemes to graphemes. Reading may involve sublexical processes, especially for less familiar words, but spelling requires segmentation and precision (Vellutino et al., 2007).
Conclusion
This exploratory study provides support for the use of a spelling test as a screening tool instead of letter identification performance in an RTI framework, or as a screening tool for more in-depth assessments using a multifaceted battery that requires more resources. It seems likely that spelling could become the preferred screening tool, over phonological awareness and rapid naming, in some bilingual settings because it is easily administered by a teacher on a one-to-many basis.
Footnotes
Acknowledgements
We thank the two anonymous reviewers for their helpful comments on previous versions of this article.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was funded by the National University of Singapore (NUS), Tier 1 Research Grant R-581-000-063-112, and the NUS’s Graduate Research Scholarship awarded to the first and the third authors.
