Abstract
The purpose of this study was to analyze published decoding tests, at the item level, to determine what decoding skills and discrete letter-patterns are assessed, and identify potential instructional implications of these measures. Twenty published word list decoding tests, either used in research or within a Multi-Tiered Systems of Supports (MTSS) framework, were included for analysis. Test items were coded at the syllable level to identify discrete letter-patterns that represent common decoding skills. Frequency of these skills, along with administration features, were analyzed to identify commonalities and differences between published decoding tests. Results show the published decoding tests analyzed vary in the number of decoding skills they measure. All tests showed potential for providing some diagnostic information to inform instructional practices. Implications for practice, directions for future research, and limitations are discussed.
Reading proficiency attainment continues to be an educational need. In 2015, only 36% of fourth graders scored at or above the proficient level in reading on the National Assessment for Educational Progress (NAEP). Of perhaps greater concern is that 31% of fourth graders scored below the basic level (National Center for Education Statistics, 2015). With a majority of students failing to meet reading proficiency standards, it is prudent to examine how educators are making instructional decisions to prevent reading failure. Examples of instructional decisions include selecting skills for instructional focus, intervention selection, placing students in instructional groups, determining intensity and frequency of interventions, and determining the need for additional assessments (Hamilton, Halverson, Jackson, Mandinach, & Supovitz, 2009).
Deficits in word reading skills are often the root cause for readers who struggle with fluency or comprehension (Carver, 1998; Murray, Munger, & Clonan, 2012). Up to 80% of students with a specific learning disability in reading struggle at the word level (Moats & Tolman, 2009). The ability to accurately and efficiently read words affects reading development, and is highly correlated with overall reading ability (Fuchs, Fuchs, Hosp, & Jenkins, 2001). While accurate word reading does not guarantee reading fluency and comprehension will occur, they are not possible without intact word reading skills. Readers who devote complete attention at the word, or even letter level do not read fluently and have few cognitive resources left for connecting or generating meaning to what they have read (Perfetti, 1986).
With the critical importance of word reading to overall ready ability, it is prudent that teachers make accurate and informed instructional decisions to ensure their students acquire effective word reading skills. Word reading skills, which provide readers with reliable strategies for identifying words in text, are often referred to as decoding skills. While there are many decoding assessments that can inform teachers instructional decisions for students with word reading problems, the type of specific information they provide has not been studied. The purpose of this study was to analyze published decoding tests for their diagnostic potential and identify possible instructional implications of these measures.
Decoding
Word reading skill deficits are most often linked to problems with decoding. Decoding refers to the ability to apply phoneme-grapheme connections to pronounce words. Phonemes are the smallest units of sound in language, and graphemes are the visual representations (i.e., letters) of those sounds in printed text. Decoding is a necessary, but not sufficient, component in the reading process as it assists with word identification (Gough & Tunmer, 1986). Decoding skills provide readers with a strategy to use when they encounter an unknown word that cannot be immediately retrieved from their lexicon (Pritchard, Coltheart, Palethorpe, & Castles, 2012). Decoding proficiency is also most often the source of difficulty for students who are struggling with reading (Christo & Davis, 2008; Fletcher, Lyon, Fuchs, & Barnes, 2007; Good, Simmons, & Kame’enui, 2001). Therefore, it is essential that decoding skills be intact to promote reading development and higher order reading skill acquisition.
In addition to the strong link between decoding and overall reading ability, there are also practical considerations that increase interest in decoding skills. Most reading curriculum include instruction on decoding patterns both in single and multisyllabic word patterns, and these same decoding patterns are also included in Grades K–5 in the English and Language Arts Common Core State Standards (CCSS) under the “Phonics and Word Recognition” category (National Governors Association Center for Best Practices & Council of Chief State School Officers, 2010). Organized hierarchically, the skills included in the upper grades rely on mastery of skills in previous grade levels. Students who experience difficulty with basic decoding skills, or fail to master the specific skills in their current grade level, are unlikely to master more complex skills in subsequent grades.
Similar to the organization of skills included in the CCSS, decoding typically follows a predictable developmental progression where disruptions or breakdowns at any point in the developmental continuum can prohibit progress (Ehri, 1997). There are roughly 44 sounds in spoken English (i.e., phonemes), and decoding requires a reader to learn to efficiently recognize the hundreds of printed letter-patterns that represent those sounds (Blevins, 2006). There are substantially more printed patterns than there are sounds. For example, the sound /k/ can be represented by many different spellings, such as “c” in “cat,” “k” in “kite,” “ch” in “chord,” or “que” in “opaque.” Decoding skills provide readers with word analysis strategies for interpreting and accurately applying the correct pronunciation to each letter-pattern.
The most basic decoding skill is the ability to connect a single letter with a single phoneme in a word. From that foundation, readers progress to associating more complex letter-patterns with pronunciations. Syllable rules, knowledge of root words, affixes, and specific digraph and blend pronunciations all assist a reader in identifying words. Proficient readers have mastered basic decoding skills (e.g., application of letter-sound associations in one syllable words) by the end of second grade (Chall, 1996). Advanced decoding skills (e.g., application of syllable rules and identification of words parts) continue to be refined throughout the middle grades, and are associated with gains in reading ability and improvements in spelling (Ehri & Wilce, 1987).
Testing to Inform Instructional Practice
Tests are used for both identifying which students are in need of additional support and what skills that support should focus on. The increase in prevalence of Multi-Tiered Systems of Supports (MTSS) and Response to Intervention (RtI) procedures in schools has resulted in an increase in the number and type of tests used in classrooms. Within the MTSS/RtI model, test data serve a variety of instructional purposes.
Purpose of testing
The purpose, or function, of a test can typically be categorized into one of four categories: screening, progress monitoring, diagnostic, and outcome. Screening tests are designed to identify students at-risk for learning failure. Progress-monitoring tests are designed to detect gradual changes in student performance over time. Diagnostic tests are designed to inform instructional practice by providing detailed results on specific skill proficiency. Outcome tests are summative, and are designed to measure if a student has acquired a skill or met a goal (Hosp, Hosp, & Howell, 2016).
Components of diagnostic tests
Of all the functions of a test, diagnostic tests provide the most specific information that links directly to instructional practice (Fuchs & Fuchs, 1996). Before a diagnostic test is given, typically students are screened to identify which students are at-risk. Diagnostic tests are then given to assist teachers in selecting appropriate interventions related to specific skills (Mathers, Sammons, & Schwartz, 2006; Van der Lely & Marshall, 2010). Although screening tests are useful in identifying students at-risk, they lack adequate sensitivity to inform which specific reading skills require attention (Deeney, 2010; Hosp & Fuchs, 2005; Wilson & Lonigan, 2010). Unlike screening tests, diagnostic tests are not intended to be given to all students. Instead they are most appropriate for those students who exhibit chronic skill deficits, and do not respond to high-quality intervention practices. In a MTSS framework, diagnostic tests are most appropriate for students in Tiers 2 and 3.
Previous studies have shown that diagnostic information on specific performance deficits when used to inform instruction has positively affected student outcomes in spelling and math (Fuchs & Fuchs, 1990; Fuchs, Fuchs, Hamlett, & Allinder, 1989, 1991). The complex process of reading can make diagnostic measures complicated. Focusing on a specific reading component, like decoding, can improve the diagnostic quality of reading tests. Previous findings on using diagnostic information show that teachers who use diagnostic data are able to tailor instruction to meet the individual needs of learners (Fuchs, Fuchs, & Hamlett, 2007). The use of high-quality diagnostic tests allows for low levels of inference regarding how results identify potentially beneficial instructional practices (Hosp & Ardoin, 2008). For decoding measures, one way to reduce the degree of inference for instructional decision-making is to score test items at the error level. Test items scored at the error level identify the discrete error the student made when reading. This allows teachers to identify specific error patterns and provide targeted instruction to remediate the specific skill deficit. Research has demonstrated that close matching of skill deficits to targeted instructional interventions positively affects student performance (Fuchs, Compton, Fuchs, Bouton, & Caffrey, 2011; Hosp & Ardoin, 2008; Olinghouse, Lambert, & Compton, 2006; Parker & Burns, 2014). In contrast, tests that are dichotomously scored at the item level as correct or incorrect do little to inform instructional practice, because it is impossible to determine what specific error the student made while reading the word.
In addition to scoring at the error level, diagnostic decoding tests that sample a broad range of skills allows for improved instructional insight (Zumeta, Compton, & Fuchs, 2012). Inclusion of a range of skills allows teachers to identify where the student is experiencing difficulty and may require additional instruction. For example, a student who has difficulty applying short vowel sounds in single syllable words is likely to experience difficulty applying short vowels in multisyllabic words. For this student, instruction on short vowel sounds should be provided to help address the skill deficits, with hopes that mastery of that skill will generalize to gains when reading more complex words. Vowels are the most common source of pronunciation errors in words, and students with poor decoding skills are likely to experience difficulty applying vowel patterns (DiBenedetto, Richardson, & Kochnower, 1983; Willson, Rupley, Rodriguez, & Mergen, 1999). However, assuming that all students who experience difficulty with decoding do so because of vowel errors is ill advised, as there are many decoding patterns (e.g., blends, digraphs and affixes) that students may need additional assistance with to address their specific skill deficits. This highlights the importance of scoring at the error level to ensure scores accurately represent skills that require instructional attention versus those that do not.
Test components
When determining the number of items on a diagnostic test, it is important each discrete skill be represented with enough frequency to determine if the skill has either been mastered or requires additional instruction. The number of items required to determine mastery varies across tests. Some curriculum-based measurements (CBMs) provide recommendations for criteria for mastery (Good, Simmons, Kame’enui, Kaminski, & Wallin, 2002). Other general recommendations for making mastery decisions include the criterion that students respond correctly 70% of the time on their first opportunity to respond. In addition, mastery determinations must be greater than chance (Engelmann, 2007). In a practical approach, this requires readers be given a minimum of three opportunities to demonstrate each skill. If a skill was represented only twice, and the student responded correctly on the first opportunity, and incorrectly on the second opportunity (i.e., 50% accuracy), there would be no way to accurately infer if that skill was still emerging, if there was a true skill deficit, or if the error was a fluke. The practical explanation of the three-response opportunity minimum is also mathematically sound. Examination of binomial probability distributions reveals that given three opportunities, the probability of randomly achieving a score of 2/3 on three dichotomous tasks is 37.5%, and random chance of a score of 3/3 is 12.5%. Thus, when a student is given three response opportunities, the confidence in identifying true skill deficits improves to better than chance odds. However, while three responses is an acceptable minimum, increasing the number of opportunities can improve confidence in the accuracy of the diagnostic information provided by the test. In addition to performance on a specific decoding test, educators can increase their confidence in student mastery of a skill by collecting additional samples of student skill demonstration (e.g., spelling samples).
An additional consideration in decoding tests is accounting for fluency of skills. Students who are able to quickly and accurately apply decoding rules to read words in isolation (e.g., reading word lists) are more likely to be proficient with applying skills in connected text. Diagnostic decoding tests should provide some indication of skill automaticity. However, under timed constraints (e.g., 1-min measures), readers will have limited opportunities to demonstrate skills. This is particularly problematic for readers who struggle, and take more time to read words. Timed fluency tests for decoding limit the number of words the student is able to read and decreases the number of patterns or skills that the teacher may observe during the testing session (Reutzel, Brandt, Fawson, & Jones, 2014; Ritchey, 2008). However, decoding fluency is still an important consideration because laborious application of grapheme-phoneme correspondences does not indicate proficiency (Joshi & Aaron, 2002). Thus, a need for balance between the number of skill demonstrations offered and administration time exists for decoding tests.
Test items
Test items on decoding tests are typically formatted as word lists. The type of words included in those word lists varies by test. Tests may use lists of real words, nonsense words, or a combination of both. Real word decoding tests provide test-takers with a range of words that can be pronounced using decoding skills, or recognized from memory. Beyond memorizing the word, real word decoding tests may also allow the test-taker to use vocabulary and comprehension skills to assist in reading the words. One of the trade-offs of using decoding tests with real words is that students may recognize words from memory or use other skills to assist in reading the word other than relying solely on their decoding skills. This may lead to inflated scores and can be problematic for teachers who are interested in identifying decoding skill deficits that may be affecting reading ability. A potential solution is to control for memory in word identification. This is accomplished by tests that use nonsense words as test items. Using nonsense words requires readers to apply their grapheme–phoneme correspondences to read the word, and eliminates the possibility that the reader is able to read the word based on memory or other reading skills (e.g., vocabulary, comprehension). Reading nonsense words can indicate the range of a reader’s knowledge of word patterns (Pierce, Katzir, Wolf, & Noam, 2010). Struggling readers display more impairment when reading nonsense words than proficient readers do (Gottardo, Chiappe, Siegel, & Stanovich, 1999). The ability to read nonsense words is predictive of the ability to read real words and of overall reading level (Carver, 2003). However, to strike a balance between real and nonsense words, a third option is for decoding tests to use a combination of real and nonsense words as test items.
Specific decoding skills
Previous examinations of decoding tests noted that considerable variability exists between decoding tests with regard to the number and type of skills they include (Martens, Steele, Massie, & Diskin, 1995). As previously mentioned, decoding is a category that includes a variety of skills. The specific decoding skills (e.g., short vowels, long vowel, digraphs, vowel teams) measured by decoding tests should be relevant to instruction. The specific decoding skills of interest are dependent on grade level and instructional expectations. The CCSS includes a variety of decoding skills across grade levels. For example, the standards in Grade 1 include consonant digraphs (e.g., th, ch), final e and vowel team representations of long vowels, and decoding one and two syllable words. For comparison, the standards in Grade 3 include knowledge of prefixes, suffixes, and decoding mulitsyllable words. The grade level of the test-taker will inform which specific decoding skills are of high interest for inclusion on a test.
In addition to indicating what specific decoding skills should be tested, the CCSS also provide guidance on the types of words in which those decoding skills should appear. Beginning in first grade, the standards state that students are expected to decode two syllable words. Then in second grade, the standards state that decoding skills should be applied to multisyllabic words. It is important for decoding tests to include test items that reflect word complexity expectations.
Decoding skills are critical in relation to general word analysis skills (i.e., breaking down long or complex words into smaller segments that decoding skills can be applied to). Students who are exposed to effective word structure and word analysis instruction are likely to develop strategies that will help them accurately decode most words in the English language (Henry, 2010). It follows then that decoding tests should include skills that can inform instructional practice. The more specific the decoding test results can be with identifying the decoding skill, or even letter-patterns that are sources of difficulty for struggling readers, the more tailored instruction can be to remedy skill deficits. For example, if a reader has difficulty identifying words with short vowels, it is instructionally beneficial to know if particular letters (e.g., a, e, i, o, u) are the source of the difficulty, or if the source of the difficulty relates to general decoding skills (e.g., recognizing short vowels are associated with vowel-consonant [VC] patterns).
Purpose of the Current Study
Research on decoding assessment has highlighted several aspects to consider regarding decoding test composition. As detailed above, the type of words (e.g., real and nonsense, single- and multisyllabic), error level scoring, fluency, and the number and frequency of discrete skills, are all important features to consider when selecting a decoding test. The purpose of this study was to analyze published decoding tests for their diagnostic potential and identify instructional implications of these measures. It is prudent for teachers and other practitioners who rely on assessment results to inform their instructional practice know what and with what frequency discrete decoding skills are included on available decoding tests. This study sought to answer the following questions:
Method
To answer our research questions, we compiled decoding test information from a variety of sources. To identify decoding tests for our study, the following procedures were used:
A review of current decoding research literature was conducted by completing a computer search of ERIC and PsychINFO databases. The following search terms were used: “phonics,” “decoding,” “word reading,” “word identification,” “word recognition,” “pseudoword reading,” “word reading fluency,” “assessment,” and “test.” Following the search return, the results were narrowed by only including studies published after 2000 (to identify those assessments currently used in research). In addition, only peer-reviewed articles were included.
Various CBM tools used in MTSS and RtI processes were identified by cross-referencing the results of the search in Step 1, with a web search for CBM publishers, as well as published CBM literature (Hosp et al., 2016).
The four most recent volumes of the Mental Measurements Yearbook were searched to identify published tests that report to measure decoding.
Compiled lists were then subject to the inclusion criteria.
Inclusion Criteria
Test protocols and word lists were gathered from identified assessments that met the following inclusion criteria:
The tests were used in recent decoding research or practice. As identified by the review of the decoding literature and web searches for tests used in MTSS or RtI procedures.
The test items were presented as word lists (i.e., no connected text).
The test reported psychometric properties for both reliability and evidence of validity. This was done to eliminate informal and practitioner-made decoding measures.
Coding Procedure for Determining Skills Measured on Included Tests
Only those test items that were words were coded (i.e., any letter identification items were omitted). All test items were coded at the syllable level. For all codes, consonants were denoted as “C” and vowels as “V.” Codes were developed based on the most common letter-patterns in English that are commonly included in curriculum scope and sequences, and included in the CCSS. We identified four broad skill categories: short vowel, long vowel, consonant groupings, and affixes. Several decoding skills were included in each broad skill category. For example, VC, CVC, and CVCC are decoding skills within the short vowel category. In addition, we identified several decoding skills that stand alone: vowel team, r-controlled vowel, consonant +le, and schwa vowels. Within the 14 decoding skills included for coding, two are represented by single discrete letter-patterns: consonant +le (Cle) and schwa vowels (ə), the other 12 have more than one letter-pattern associated with the decoding skill. For example, within the three short vowel decoding skills (VC, CVC, and CVCC), there are five letter-patterns (a, e, i, o, u) representing the five vowel sounds. Table 1 includes the 14 decoding skills and specific letter-patterns coded.
Decoding Skills and Letter-Patterns Analyzed.
Note. C = Consonant; V = Vowel; Cle = Consonant +le.
To code single syllable words, the number of letters in the word determined the coding procedure. For example, the nonsense word: “zilk” would be coded as a four-letter word with the CVCC letter-pattern, and a short “i” vowel. The number of multisyllabic words that appeared on each test was recorded. Multisyllabic words were first divided into syllables using the specified syllable breaks provided by the test publication. If no guide for breaking words apart was provided, then multisyllabic nonsense words were broken into individual syllables based on the most common syllable rules, and multisyllabic real words were broken based on dictionary notation. Once broken into individual syllables, multisyllabic words were coded using the same rules used for single syllable words (e.g., the nonsense word “punmag” was split into two syllables: pun/mag, the first syllable “pun” was coded as a CVC word with a short “u” vowel, and the second syllable “mag” was coded as a CVC word with a short “a” vowel). Additional skills coded included schwa vowels, consonant +le syllables, blends, digraphs, vowel teams, prefixes, and suffixes.
Categorizing sight words
For all real word test items, we cross-referenced each test item with complete Dolch (393 words) and Fry (1,000 words) high-frequency word lists. If the word appeared on either word list, it was counted as a sight word. Next, the identified sight words were coded on the syllable level to identify any decoding skills in the word. The words that were identified as both decodable and high-frequency (e.g., side) were categorized as decodable sight words in the database. Words that could not be coded (i.e., words with pronunciations that could not be associated with standard decoding rules, such as “all”) were categorized as irregular sight words.
Administration features
In addition to coding the test items, the scoring procedure for each test was noted. Tests were either scored on the item level (i.e., dichotomously if the whole word was correct or incorrect), or on the error level (i.e., specific letters within a word were marked incorrect/correct where partial credit was possible). Other administration features of the tests were also noted. The time requirements of tests with timing rules (e.g., 1-min) were noted as tests that provided an indication of fluency. In addition, discontinue or ceiling rules, which might limit the number of test items given, were also documented.
Data were entered into a database and frequency counts were used to determine the number of times letter-patterns and specific decoding skills appeared in tests. Frequency of skills was compared across included tests. Twenty percent of tests were randomly selected to check for inter-rater reliability. Point-by-point reliability was 97.8%.
Results
A total of 20 tests met the inclusion criteria for this study: eight tests that use nonsense words, 11 tests that use real words, and one test that uses a combination of real and nonsense words as test items. In the case where a test had multiple forms (e.g., Test of Word Reading Efficiency–2nd Edition [TOWRE-2]), Form A was included and coded. For all tests that had alternate forms for different grade levels (e.g., Aimsweb, FastBridge, easyCBM), Grade 1 forms were used. Fall screening forms were used for FastBridge, DIBELS 6th, easyCBM, and Word Identification Fluency (WIF) CBMs. For the Aimsweb Nonsense Word Fluency (NWF) CBM, the first progress-monitoring probe was used.
Research Question 1
Table 2 displays the number of decoding skills and administration features coded on published decoding tests included in this study. The range of total test items was from 26 (Woodcock Reading Mastery Test [WRMT]-3, Word Attack) to 129 (Consortium on Reading Excellence–Phonics Survey [CORE-PS]). As designed, only three of the tests (Names Test, Early Names Test, and CORE-PS) are given in their entirety per standard administration procedure. Two tests include the use of proper nouns (i.e., on Names Test and Early Names Test), the third uses a combination of nonsense and real words (i.e., CORE-PS).
Skills (Assuming All Test Items Administered) and Administration Features of Word List Decoding Tests.
Note. Tests with nonsense word items are in boldface. CORE-PS = Consortium on Reading Excellence–Phonics Survey; TOWRE-2 = Test of Word Reading Efficiency–2nd Edition; KTEA III = Kaufman Tests of Educational Achievement III; WIAT-III = Weschler Individual Achievement Test III; WJ-IV = Woodcock Johnson Tests of Achievement IV; WRMT-3 = Woodcock Reading Mastery Test–3rd Edition; DIBELS = Dynamic Indicators of Basic Early Literacy Skills; NWF = Nonsense Word Fluency; DWR = Decodable Word Reading; SWR = Sight Word Reading; WIF = Word Identification Fluency (Fuchs, Fuchs, & Powell, 2004).
Test includes a combination of real and nonsense word test items. bThe number of items completed at 30 s is used to compute fluency; the test continues until the discontinue rule is met.
The other 17 tests in this study have standard administration procedures that typically result in test-takers completing a reduced number of test items, either through time restrictions, the use of a discontinuation rule, starting students beyond the first item, or a combination of these. Eleven tests use time to provide an indicator of fluency. All the CBM measures (i.e., Aimsweb NWF, DIBELS 6th NWF, FastBridge NWR, DWR, and SWR, WIF, and easyCBM, Word Reading) are 1-min timed measures where the student is instructed to read as many of the test items as they can in 1 min. The TOWRE-2 Phonemic Decoding Efficiency and Sight Word Efficiency subtests are 45-s timed tests. The Word Reading and Pseudoword Decoding subtests of the Weschler Individual Achievement Test–III (WIAT-III) have timing components where the administrator indicates how many words the test-taker has read in 30 s, but testing continues until the discontinue rule has been met. Other tests have mandated discontinue or ceiling rules for administration that indicate how many errors in a row a student must commit before administration is ended. The range of discontinue rules on tests included in this study was from four to six errors.
The scoring procedures were also noted for each test. Only Aimsweb NWF and DIBELS 6th NWF are designed to be scored at the error level. All other included tests are designed to be scored dichotomously: correct/incorrect at the whole word level for all test items.
The range of decodable sight words was from 1 to 91 words represented on a total of 11 tests. The majority of these 11 tests also include irregular sight words, ranging from 1 to 16 on a total of nine tests. Every test that used real words included at least some sight words.
The range of multisyllabic words was from 0% (n = 4) to 72% (Kaufman Tests of Education Achievement [KTEA]-III, Letter and Word Recognition). Multisyllabic words were represented most often on decoding test using real words (10 out of 11) versus nonsense words (five out of eight) with the CORE-PS having both real and nonsense words. The majority of tests (n = 13) had fewer than 50% of their items as multisyllabic, seven of these are nonsense word tests. In addition, five tests (30% of total tests included) have no multisyllabic test items, four being CBMs and the fifth being the Early Names Test.
Research Question 2
The number and frequency of discrete letter-patterns found on each test is reported in Table 2. The number of discrete letter-patterns that appear at least 3 times on the test are reported. These data are reported following the assumption that all test items are administered.
Short vowels are represented by three decoding skills: CVC, CVCC, and VC. Of these three decoding skills, CVC is represented most often on 18 (90%) of the tests, CVCC is represented on 11 (55%) of the tests, and VC is represented on eight (40%) of the tests. CVC is the only short vowel category where some assessments (five out of 18) have diagnostic potential in all five short vowels. No test provided diagnostic information on all five short vowel letter-patterns in VC and CVCC decoding skills.
Long vowels, are represented by three decoding skills: V, CV, and CVCe. Of these three decoding skills, CV is represented most often on nine (45%) of the tests, CVCe is represented on six (30%) of the tests, and V is represented on three (15%) of the tests. The tests that measured the most discrete letter-patterns within long vowels are the TOWRE-2, Sight Word Efficiency (n = 4), and the KTEA-III, Letter & Word Recognition (n = 4). No test provided diagnostic potential on all possible long vowels. The other vowel decoding categories that were coded are r-controlled and vowel team. The r-controlled decoding skills are represented frequently with 10 (50%) of the tests including this skill while the vowel team is only represented on five (25%) of the tests. No test provided diagnostic information on all five r-controlled letter-patterns or all 21 vowel team letter-patterns coded.
Consonant groupings are represented by two decoding skills: digraphs and blends. Blends are represented on nine (45%) of the tests and digraphs are represented on six (30%) of the tests. The TOWRE-2, Sight Word Efficiency and Phonemic Decoding Efficiency tests, and easyCBM, Word Reading, all measure three blends. The other six tests that include blends provide diagnostic information on only one blend letter-pattern. No test provided diagnostic information on all 15 of the blend letter-patterns coded. The Names Test measured the most digraphs (n = 2). All other tests that measure digraphs provide diagnostic information on only one digraph letter-pattern. No test provided diagnostic information on all eight of the digraph letter-patterns coded.
Affixes are represented by two decoding skills: prefixes and suffixes. Suffixes are represented more often than prefixes on decoding tests. Four (20%) of the tests measured suffixes. No test provided diagnostic information on all 35 of the suffix letter-patterns coded. Only two (10%) of the decoding tests measured prefixes. No test provided diagnostic information on all 28 of the prefix letter-patterns coded.
The final two decoding skills represented on the decoding tests were consonant +le (Cle), and schwa vowels (ə). Three (15%) of the decoding tests measured consonant +le. Two of the three used real words and one test used nonsense words. Seven (35%) of the decoding tests measure schwa vowels. Schwa vowels were only measured on real word tests.
Research Question 3
The instructional implications based on diagnostic information provided by the included decoding tests is based on the number and frequency of decoding skills and letter-patterns included on each test. The number of discrete letter-patterns represented at least 3 times within a decoding test ranged from two to 24 skills. Ten percent (n = 2) of the tests measure more than 20 discrete letter-patterns. Both of these tests use real words, one is easyCBM, Word Reading, the other is the TOWRE-2, Sight Word Efficiency. Forty percent (n = 8) of the tests measured between 10 and 17 discrete letter-patterns. Of these eight, four use real words, three use nonsense words, and one uses a combination of real and nonsense words. For the 10 tests that measure 10 or more discrete letter-patterns, the majority of these tests also include a high percentage of multisyllabic words.
Table 2 shows the total number of discrete letter-patterns included on decoding tests is much larger than the number of discrete letter-patterns that are represented at least 3 times on each test. To illustrate the diagnostic potential of decoding tests, Table 3 displays the 14 decoding skills, and the discrete letter-patterns represented at least 3 times within each skill. By providing three opportunities, these letter-patterns show potential for providing some diagnostic information to inform instructional practices.
Discrete Skills Represented ≥ 3 Times on Word List Decoding Tests Assuming All Test Items Administered.
Note. Tests with nonsense word items are in boldface. C = consonant; V = vowel; VT = vowel team; RC = r-controlled vowel; DG = digraph; BL = blend; PF = prefix; SF = suffix; Cle = Consonant +le; ə = schwa vowel; CORE-PS = Consortium on Reading Excellence–Phonics Survey; TOWRE-2 = Test of Word Reading Efficiency–2nd Edition; KTEA-III = Kaufman Tests of Educational Achievement III; WIAT-III = Weschler Individual Achievement Test III; WJ-IV = Woodcock Johnson Tests of Achievement IV; WRMT-3 = Woodcock Reading Mastery Test–3rd Edition; NWF = Nonsense Word Fluency; DWR = Decodable Word Reading; SWR = Sight Word Reading; WIF = Word Identification Fluency.
All test items given. bTest includes a combination of real and nonsense word test items.
Only two tests—easyCBM, Word Reading, and TOWRE-2, Sight Word Efficiency—covered at least 75% of the 14 decoding skills coded in this study. Both easyCBM, Word Reading, and TOWRE-2, Sight Word Efficiency, have diagnostic potential in 11 of the 14 decoding skills. Six of the 20 tests have diagnostic potential in seven to 11 of the decoding skills coded. Of those six tests, two are nonsense word tests and four are real word tests. Five of the 20 tests have diagnostic potential in four to six of the decoding skills. Of those five tests, one is a nonsense word test, three are real word tests, and one has a combination of real and nonsense words. Seven of the 20 tests have diagnostic potential in less than four of the decoding skills. Five of those seven are nonsense word tests, and five are CBM tests.
Discussion
The purpose of this study was to analyze published decoding tests for their diagnostic potential and identify instructional implications of these measures. None of the 20 decoding tests included in this study cover all the skills students are expected to master. However, all the tests have the potential to provide at least some diagnostic information to inform instruction. The amount of diagnostic information available is dependent on the test.
Some of the most widely used tests, like CBMs, also have the most limited diagnostic information. Many of the CBMs focus only on short vowels within CVC and VC patterns. The one exception is easyCBM, Word Reading, that measured 23 discrete letter-patterns across 11 of the 14 decoding skills while also including 28% multisyllabic words. If the only decoding skill of interest is short vowels within single syllable words, there are many CBM decoding tests to choose from. However, given the complexity of decoding skills and the expectations even for first-grade students, educators will need to look beyond most CBM measures to obtain diagnostic information to inform instruction on discrete letter-patterns.
Fluency on a task implies a level of mastery that is also important to consider when making instructional decisions. While just over half the decoding tests provide some measure of fluency, with the majority being CBMs, it is done at the expense of potential diagnostic information. This is because the student typically only reads some, not all, of the words on the test. This is true for tests with a timing component as well as tests with discontinue or ceiling rules. While these tests may produce a fluency score and/or reduce the time required for test administration, they also decrease the ability to obtain diagnostic information for instructional decision-making. Given the competing requirements for measuring fluency and diagnostic information, it is not surprising we did not find any decoding tests that did both well.
Additional considerations for using a decoding test for diagnostic information is the use of real and nonsense words, as well as single and multisyllabic words. Varying the type of items on a decoding test provides an opportunity to observe word analysis skills within different word types. Results of this study show the decoding tests that show the greatest potential to measure decoding skills use real words. These decoding tests also include more multisyllabic words which provide natural opportunities to include skills like affixes (i.e., prefixes and suffixes) and consonant groupings (i.e., digraphs and blends) that may be harder to achieve in nonsense words. While no decoding test measured all the affixes coded, the suffixes included on the KTEA-III, tests are identified as some of the most commonly occurring (Fry & Kress, 2006). A similar trend was observed for vowel teams and the schwa vowel, which only occurred on tests using real words. Interestingly, only easyCBM Word Reading, provided potential for diagnostic information on multiple vowel teams (i.e., ea, ee, oo, ou, ow) that represent some of the most commonly occurring vowel teams (Fry & Kress, 2006).
Implications for Practice
The results of this study show that test selection should be a critical consideration for practitioners. The diagnostic information available for informing instructional decision-making is dependent on the test used. In practice, this study can assist teachers with aligning their instructional expectations (i.e., what they teach) to tests that will provide them with information on student attainment of those expectations. When selecting a test, teachers should first identify decoding skills of interest. It is not necessary to test every possible decoding skill with every student, as this would be cumbersome and impractical. After determining which decoding skills are of interest, teachers should review tests to determine the number and type of those decoding skills and letter-patterns measured on them. Next, word complexity (i.e., single syllable vs. multisyllabic) should be examined. Given that standards require application of decoding skills in multisyllabic words beginning in first grade, it is important for teachers to test skills in words that match instructional expectations. The CCSS also include expectations for sight-word and irregular-word reading, so consideration of the number of those words included on a decoding test should be given. Sight word data from tests may be instructionally beneficial to inform instruction on fluency and phrasing. Then, consideration for the word type (i.e., real and nonsense) should be considered. Using nonsense words on decoding tests not only controls for memory in word identification but also limits inclusion of other decoding skills (e.g., schwa vowels, prefixes, and vowel teams). In addition, results of this study show that decoding tests that use real words measure more skills than tests that use nonsense words. Additional points of consideration in decoding test selection include availability of a fluency score, total administration time requirements, test cost, and item scoring procedures.
Tests that are not scored at the error level have limitations in their ability to inform instructional practice. Dichotomous scores at the item level do not provide teachers with information about the source of error in the word. Thus, identifying error patterns across words to determine students’ instructional needs is not possible. This is because the absence of error patterns requires high levels of inference to determine the specific skills a student is struggling with. Absent such a pattern, it is possible an error was a fluke and not a true skill deficit in need of remediation. In addition, test items often contain multiple decoding skills, and if the test item is simply scored correct or incorrect, there is no information that a teacher can take from that item to inform the focus of their instruction. This may lead to inaccurate assumptions about the source of the student’s difficulty with decoding, and time spent on instruction on skills that do not address the students’ actual deficits.
Consider the following simple example to illustrate our point: two students reading the CVC nonsense word “bip.” Student 1 incorrectly pronounces the vowel, while Student 2 substitutes /d/ for /b/. Both students incorrectly pronounce the word. The teacher provides additional instruction on short vowel pronunciations (i.e., the most common source of error in reading CVC words). Following the instruction, Student 1 is able to correctly read CVC words with short /i/. However, Student 2 continues to mispronounce words. The additional instruction did nothing to remediate the deficit Student 2 was experiencing. Instead, Student 2 spent instructional time participating in an intervention targeting a skill they had mastered. This wasted instructional time puts Student 2 further behind as they continue to struggle with misidentification of “b” and “d.” The test results were not specific enough to inform instructional practice, resulting in a mismatch between instructional needs and instructional practice. Teachers, and other practitioners who use the tests included in this study, should be aware of the importance of making sure that data collected during testing accurately and adequately inform their instructional decisions.
Future Directions and Limitations
Researchers have shown that the use of diagnostic information improves instructional decision-making (Fuchs, Fuchs, Hosp, & Hamlett, 2003; Capizzi & Fuchs, 2005). In addition, decoding is a foundational reading skill, and is highly predictive of reading proficiency. Based on the importance of decoding and our analysis of items on published decoding tests, there is a necessity for improving diagnostic decoding tests. Researchers and publishers interested in assessing decoding skills could use this information to understand the gaps that exist. In addition, our methods for identifying the letter-patterns and decoding skills of test items could be applied to other tests, such as informal or unpublished measures, CBMs at alternate grade levels, tests of decoding in other languages, and spelling tests. Considerations for future decoding test development, based on the gaps we identified in our current study, would be the following: increasing the variety of decoding skills and letter-patterns included, providing multiple opportunities for skill demonstration, ensuring that skills are represented in single and multisyllabic words, and developing informative scoring procedures that allow for efficient connections between skill needs and instructional practice.
There are limitations to our study that should be addressed in future research. In identifying tests for inclusion, we were limited to the results of the search. It is possible that additional or alternate search terms may have returned additional studies for our review. In addition, the coding rules used to code the included tests may have affected results by forcing items that contain multiple skills to only be coded by a single skill (e.g., item “droy” was coded per the rules as blend-dr, but also contains the vowel team “oy” as well). Last, we selected Grade 1 CBMs (e.g., NWF) because the majority of decoding CBMs occur at this grade level. Limiting our inclusion criteria to Grade 1 CBMs may have led to an underrepresentation of multisyllabic words, as well as underrepresentation of specific decoding skills (e.g., prefixes and suffixes) that are typically not the focus of instruction during first grade. However, the majority of the assessments reviewed are used with students at all grade levels. Therefore, including additional CBMs beyond first grade would not significantly change our overall results.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
