Abstract
Measurement in early childhood is an increasingly large-scale endeavor addressing purposes of accountability, program improvement, child outcomes, and intervention decision making for individual children. The Early Communication Indicator (ECI) is a measure relevant to intervention decision making for infants and toddlers, including response to intervention approaches. The widespread use of the ECI is growing in multiple programs and states. Local program staff members collect ECI data and, with their program directors, manage their own system of ECI measurement. Program-level implementations represent independent ECI measurement replications, and the success of each potentially influences the quality of data produced and, ultimately, the validity of the inferences made thereof. The purpose of this research was to examine program-level influences on child-level ECI total communication growth and 36-month outcomes in a large sample of children, including those with individual family service plans served by multiple Early Head Start programs in two states. Results indicated variation in programs’ sociodemographic composition, ECI implementation quality, ECI total communication growth, and 36-month outcomes. Program-level sociodemographic composition was found not to be an influence on ECI growth or 36-month outcomes, whereas state location and implementation quality were. Implications are discussed.
Keywords
Measurement plays an increasingly important role in documenting the early educational outcomes of infants and toddlers with and without special needs. The need for measures to collect, use, and report meaningful child outcomes data in early childhood has increased due to federal and state policies supporting the use of child-level data for documenting program-level outcomes (Hebbeler, Barton, & Mallik, 2008). Examples are Head Start (Administration for Children and Families, 2006), Part C early intervention, and Part B preschool services (Individuals with Disabilities Education Improvement Act, 2004). Other examples include universal screening and progress monitoring needed to support the increasing interest in response-to-intervention approaches to decisions about service delivery (Berkeley, Bender, Peaster, & Saunders, 2009). As a result, measurement in early childhood has become a large-scale endeavor (Rous, LoBianco, Moffet, & Lund, 2005).
Measurement theory tells us that the validity of large-scale efforts such as these is not simply a property ascribed to the tests or specific measurement practices used with individual children but rather to the inferences that are made from the data produced through measurement (American Psychological Association, American Educational Research Association, & National Council on Measurement in Education, 1999): “Validity refers to the degree to which evidence and theory support the interpretation of test scores entailed by the proposed uses of the test. . . .The process of validation involves accumulating evidence to provide a sound scientific basis for the use of the tests” (p. 9). Validity in this context is evidence that decisions based on data collected and analyzed actually lead to attaining the system’s intended goals and purposes (i.e., program- and child-level intervention decisions leading to better child outcomes). The case for validity in a large-scale system of measurement is made through a set of logical and evidentiary statements that support the assertion that (a) the system is implemented as designed at high levels of quality and (b) the decisions based on data from the system lead to intended and not unintended results (Marion et al., 2002, p. 105).
Measurement in large-scale systems should also meet the professional standards and criteria outlined by key professional organizations guiding early intervention assessment practice (i.e., Division for Early Childhood of the Council for Exceptional Children, National Association for the Education of Young Children, National Head Start Association, and the American Speech–Language–Hearing Association). Included in recommendations for assessment in early childhood is that assessment needs be useful for informing individualized intervention and programmatic decision making (Bagnato, 2007; Carta et al., 2002; Division of Early Childhood, 2007; Sandall, McLean, Smith, & McLean, 2005; VanDerHeyden, 2005).
Individual growth and development indicators (IGDIs) are an increasingly used set of universal screening and progress-monitoring measures for children younger than kindergarten (VanDerHeyden & Snyder, 2006). IGDIs have been developed for preschoolers (McConnell & Missall, 2008; Missall, McConnell, & Cadigan, 2006) and for infants and toddlers (Carta, Greenwood, Walker, & Buzhardt, 2010). Both preschool and infant/toddler IGDIs are supported by websites that make them accessible and usable nationwide. 1 Conceptually and empirically, IGDIs were based on the general outcome measurement framework (Fuchs & Deno, 1991; McConnell, Priest, Davis, & McEvoy, 2002). As a general outcome measurement, IGDIs were designed to assess socially valid outcomes. They may be administered frequently for screening (e.g., quarterly) and even more frequently to assess the effects of an intervention (e.g., weekly, monthly); they produce scores that are reliable and valid; and they are intended for use by local service providers in intervention decision making (e.g., Buzhardt et al., 2010). Early interventionists can use IGDIs to monitor the short-term growth and development of young children (Priest et al., 2001; Snyder, Wixson, Talapatra, & Roach, 2008; VanDerHeyden & Snyder, 2006). Because IGDIs are sensitive to growth, they are well suited for planning and modifying interventions for individuals and entire programs (Walker, Carta, Greenwood, & Buzhardt, 2008). Thus, IGDIs are emerging in early childhood as resources for program accountability as well as for individual response-to-intervention purposes (Greenwood et al., 2008; McConnell & Missall, 2008).
The Early Communication Indicator (ECI) is one technically sound and scalable IGDI for measuring infants’ and toddlers’ growth in early communication proficiency (Greenwood et al., 2008; Greenwood, Carta, Walker, Hughes, & Weathers, 2006; Walker et al., 2008). Previous reports on the development and validity of the ECI have focused on evidence at the child level of analysis. For example, child-level psychometrics include score reliability and validity, as well as demonstrations of sensitive to growth over time. Also, normative estimates of children’s growth over time have been reported and used to create benchmarks for individual child intervention decision making. The influences of child-level variables have been reported (i.e., gender, individual family services plan [IFSP] status, and home language; Greenwood, Walker, & Buzhardt, 2010). Results indicated that individual growth trajectories and outcomes at 36 months of age for children in Early Head Start (EHS) were significantly influenced by whether they had IFSPs and were receiving Part C early intervention services but not by gender or the language (e.g., Spanish, English, etc.) heard at home (Greenwood et al., 2010). As a group, children with IFSPs used words in communication beginning months later, were slower accelerating over time, and were significantly lower in ECI total communication attained at 36 months of age.
At this juncture, comparatively more is known about child-level influences on psychometrics of the ECI (Greenwood et al., 2008) than about the influences of program-level factors. This may be an accurate statement of many existing general outcome measurements as well. For example, one may hypothesize that sociodemographics of a program, defined as an aggregate of its children’s characteristics, might influence the psychometrics of ECI scores. Differences in early language learning of boys and girls are well known (Fenson et al., 1994; Galsworthy, Dionne, Dale, & Plomin, 2000; Huttenlocher & Lyons, 1991). For example, the quantitative and qualitative early advantage reported for girls compared to boys on verbal measures (Galsworthy et al., 2000) and the association between differences in early learning capacity and level of adult input (Huttenlocker, Haight, Bryk, & Seltzer, 1991) may result in detectable programmatic differences on the ECI if some programs serve more boys than girls. Similarly, the number of children whose home language is other than English or who are bilingual may challenge a program’s resources in ways that adversely affect ECI administration. For example, assessors who speak the native language may or may not be readily available to be play partners during administration of the ECI or to score it. These factors may result in a compromised administration if the child is paired with a play partner who cannot converse in his or her native language, and scoring may not be accurate, because the person doing it is not proficient in the child’s language. The percentage of children with disabilities in a given program may also result in differences across programs. For instance, at the child level of analysis, we reported that children with an IFSP acquire communication skills at rates lower than those of children of the same age not receiving early intervention services (Greenwood et al., 2010). Thus, programs serving a larger percentage of children with disabilities may in turn reflect these differences at the program level. Given that program-level sociodemographic variables are known to vary by location (e.g., region and state), for example, we reported finding more non-English-speaking children in EHS programs in one state compared to another, 14% versus 2% (Greenwood et al., 2010). Therefore, we investigated the potential impact of these factors on program-level ECI outcomes.
Given implementation science (e.g., Fixsen, Naoom, Blase, Friedman, & Wallace, 2006), one may also hypothesize that program-level contextual factors surrounding the use of the ECI may influence ECI scores. For example, the fidelity and quality of its use by program staff may influence the psychometrics of ECI scores. Such factors may include the quality of ECI training and certification, the use of interobserver agreement checking, fidelity in the administration of the ECI, and the number of ECI scores falling outside statistically reasonable ranges of values (i.e., outliers). The relative effectiveness of a program’s intervention practices and their fidelity of use is another factor capable of influencing ECI scores. Thus, it is reasonable to suspect that the psychometrics of the ECI scores are a function of multiple program-level situational factors, and this has yet to be investigated.
The goal of this study was to examine program-level influences in the context of large-scale use of the ECI. We sought to advance knowledge about these potential effects and the implications that single program-level variables and combinations thereof may have on ECI score estimates and inferences made from the data. The following questions were addressed:
Research Question 1: What was the variation in programs’ sociodemographic characteristics by program and state location?
Research Question 2: What was the variation in implementation quality by program and state location?
Research Question 3: What was the variation in ECI total communication trajectories (i.e., 36-month outcomes, linear slope, and acceleration) across programs?
Research Question 4: To what extent were ECI total communication trajectories (i.e., 36-month outcomes, linear slope, and acceleration) influenced by program-level variables, including sociodemographic characteristics, state location, and implementation quality?
Method
Overview
This report is based on a secondary analysis of ECI data previously reported (Greenwood et al., 2010). The ECI is one of five IGDIs developed specifically for infants and toddlers. IGDIs are also available for preschoolers and are supported by a website (http://ggg.umn.edu/). Making this research possible was the fact that EHS programs in two states adopted the ECI as their universal measure of children’s outcomes in expressive communication and implemented data collection from 2002 through 2007. The ECI is supported by the Individual Growth and Development Indicators for Infants and Toddlers website, which provides access to ECI information and protocols as well as a data management system (http://www.igdi.ku.edu). The data management system is a hierarchical system containing the implementation tools associated with one’s role and level of access permission as decided by each program director (Buzhardt et al., 2010). Access is controlled with a login ID and password.
For this study, the two state-level directors in adjacent states each set up their own IGDI projects, populated by their local programs and coordinators. The state-level administrators were responsible for oversight of the assessment activities and had permission to view the data of all programs and children in reports they used for accountability purposes. At the program level, the program director used the website to authorize their ECI users, manage staff training certification, and monitor data collection activity. Authorized staff members entered children’s information, including sociodemographic information and all subsequent ECI data following each assessment administration. Program directors were able to generate group and individual child reports of progress, as well as authorized staff member information for purposes of management and reporting. Staff activity reports included information regarding training and staff certification, interobserver agreement checks, and the volume of data collected by individual staff. Program staff were able to access children’s data only for purposes of intervention decision making and reporting. The data management system was designed to provide support and evidence of success implementing the ECI data collection for use in making data-based decisions regarding the need for additional staff training as well as professional development in the use of intervention practices that might lead to improved services and better child outcomes.
Participants and Settings
The population of EHS programs (n = 27), staff members (n = 580), and enrolled children (n = 5,883) in two Midwestern states participated (see Table 1). Child participants consisted of children receiving Part C services (i.e., they had IFSPs; n = 471, 8% of the combined sample). EHS is a child development program serving low-income infants and toddlers and their families. Individual programs might vary, but family income at or below the federal poverty guidelines is a requirement for participation. Additionally, it is mandated that at least 10% of openings in EHS programs be available for children receiving Part C early intervention services under the Individuals with Disabilities Education Act regardless of family income. The objectives of EHS are to enhance children’s growth and development; strengthen families as primary caregivers; provide education, health, and nutritional supports; link child and family to services; and ensure well-managed programs that include parents in decision making.
Number of Children per Program and State
Child-level characteristics
The EHS programs often enrolled children as early as birth in the website, and depending on federal and state EHS policies, children were eligible to begin receiving EHS services as early as 6 months of age. Thus, more children were enrolled in programs than those who actually had ECI data. Children were served until the age of 36 months (State 2) or 42 months (State 1). The number of children in State 1 was 3,190 (ranging from 107 to 513 across programs); for State 2, 2,693 (ranging from 15 to 451 across programs). Table 1 shows the distribution of children by programs. The mean age of children at the first assessment was 18.4 months overall (SD = 9.3). Of these, 48% were male, and 90% heard English spoken at home; in both states, languages in the home included Spanish (9%) and other (1%; e.g., Arabic, Vietnamese, Somali, American Sign Language).
Program-level characteristics
Based on the volume of data collected, all but Program 34 of the 27 EHS programs reached programwide ECI implementation and maintained ECI data collection during the 2002–2007 period. As seen in Table 1, the mean was 218 (SD = 130), ranging from only 15 (Program 34) to 513 (Program 11). As described below, data for Program 34 were removed from further analyses of the research questions. The per-program distribution of staff registered in the website and implementing ECI measurement including the program director averaged 22 (SD = 15) per program over the period, ranging from 1 to 61. Program-level characteristics were calculated to represent the aggregate of the program’s child characteristics, thus representing the percentage of children with IFSPs, male gender, and non-English home language—the mean values were 8.4 (SD = 5.8), 50.5 (SD = 6.5), and 8.0 (SD = 13.7), respectively (see Table 2).
Descriptive Statistics for Level 3 Analyses
Note: No variables included at Level 2. ECI = Early Communication Indicator.
n = 16,677.
n = 26. In percentages (except state).
State 1.
State 2.
Measurement Procedures
In our study, the EHS programs used the ECI. The ECI is a 6-minute play-based measure of children’s growth in expressive communication (Carta et al., 2010). The ECI is an adaptive measure of individual progress designed to be administered quarterly for screening and as frequently as monthly to monitor the impact of a change in intervention. The original 5-year ECI development and validation effort for infants and toddlers (Greenwood, Carta, & Walker, 2005; Greenwood & Walker, 2010) involved (a) a national survey of parents of children with special needs and professionals in early childhood and early childhood special education that socially validated expressive communication as an important general outcome of early intervention for young children (Priest et al., 2001), (b) studies documenting the psychometric properties and feasibility of the ECI, including sensitivity to growth over time (Greenwood et al., 2006; Luze et al., 2001), 2 and (c) studies showing sensitivity to short-term early interventions (Greenwood, Dunn, Ward, & Luze, 2003; Harjusola-Webb, 2006; Kirk, 2006; Kosanic, 2000; Murray, 2002; Small, 2004).
The ECI total communication score reliability was reported to be r = .89 for mean level and r = .62 for linear slope (Luze et al., 2001). Alternate forms score reliability was reported to be r = .72. An estimate of the intercoder agreement on the scoring of ECI records was 90% for ECI total communication, ranging from a mean of 70% for single words to 81% for gestures. ECI total communication score validity is supported by evidence of sensitivity to age, to growth over 9 months, and to concurrently assessed criterion measures: r = .62 with the Preschool Language Scale–3 (Zimmerman, Steiner, & Pond, 1992) and r = .51 with a parent rating the child’s language skills (Luze et al., 2001).
ECI key skills elements
A child’s early communication skills (i.e., gestures, vocalizations, single-, and multiple-word utterances) are recorded by an observer in the course of interaction with an adult play partner during an ECI assessment (Greenwood et al., 2005; Walker & Carta, 2010). 3 Because these skills and ECI administration have been reported, we describe them briefly here. Gestures are defined as physical movements made by a child in an attempt to communicate with a play partner. Vocalizations are nonword verbal utterances voiced by the child to the play partner; they may occur alone or with gestures. Single-word utterances are single voiced or signed words by the child that are recognized and readily understood by the person hearing them. Multiple-word utterances are two or more different voiced or signed words by the child and readily understood by the coder. To count as a multiple-word utterance, the words/signs must fit together in a meaningful way to approximate a statement or sentence. At least two or more of the words/signs need to be understandable; however, they do not need to be grammatically correct (Walker et al., 2010). The frequency of occurrence of each key skill element is recorded on a paper-and-pencil data sheet over a 6-minute assessment. 4 The observed frequency is divided by 6 minutes to obtain the number of responses per minute: a rate score. Additionally, a total communication rate score is computed. These data are entered by program staff into the IGDI online data system.
Administration of the ECI involves a familiar adult who is taught to interact as a play partner with a child using either the Fisher-Price Barn (Form A) or Fisher-Price House (Form B). Familiarity is important in the administration of the ECI because of the potential for reactivity due to the stranger effect common in very young children (Greenwood et al., 2008). The play session takes place in a convenient setting with few distractions present in the home or early education program. The assessor times the session duration for 6 minutes using a digital timer capable of recording minutes and seconds. Accommodations made for children with sensory and/or physical impairments include moving toys closer to the child, supported positioning for the child in a manner that orients the child toward the toys and enables best access to them, and where needed, using toys that are larger and more identifiable and that make recognizable sounds. Procedures for administration with children whose primary home language is not English requires a bilingual assessor capable of discriminating utterances that are vocalizations, single words, and multiple words in the home language and in English (Walker & Buzhardt, 2010). 5
ECI assessor training
EHS staff members learned to use the ECI in a trainer-of-trainers model supported by ECI developers and website resources. This included attending a workshop conducted by the ECI developers for each program’s key staff. Key staff learned to administer the ECI, code children’s communicative events, and interpret results. These key staff returned to the local program and then submitted two videotaped ECI administrations as evidence of their use of the measure. These were reviewed by the developers for procedural fidelity and interobserver agreement. Passing these steps at a high standard (i.e., 85% or higher agreement on each tape) qualified the trainee as being ready for administering and coding child communication and training others to use the ECI. 6 Their certification was noted in the website by the developer.
Thereafter, the certified key staff trained additional local staff using materials given to them at the workshop or downloaded from the website. They also certified the assessors they trained. Directors and key staff registered all assessors expected to certify and collect data in the website and ensured that each could access the website and its online tools. By default, all assessors were registered as being uncertified and having yet to complete training. Local site training involved reading an administration manual and learning the key skill coding definitions, followed by observing and coding videotaped administrations. These materials included PowerPoint slides, manuals, and protocols. Following initial training at each local site and practice administering, assessors certified (as did their local trainers) on the basis of two videotaped administrations reviewed by the trainer for procedural fidelity as previously described. Certification results were reported to the assessor and the program director. The program director or the certifying staff member was ultimately responsible for recording the assessors’ change in certification status into the website data management system. At the time these data were collected, assessors registered in the website could administer and enter child ECI data whether they were documented in the website as being certified by the program director or not.
Interobserver agreement
After collecting data, assessors were encouraged to conduct annual interassessor coding agreement checks involving dual coding such that agreement between a reliability assessor and a primary assessor could be checked. Program directors were encouraged to monitor this reliability checking within their own programs and to provide input to individual assessors based on reports available at the website. Not all programs conducted reliability checks, however, as reported below.
In all, 246 paired interobserver assessments, 206 from State 1 versus 40 from State 2 were available for analysis (see Table 3). Pearson r was used to assess the order/interval of scores across pairs of primary and reliability assessors’ scores. Alpha was set at .01 because of the large sample size (e.g., Hartmann, 1977; Walker & Buzhardt, 2010). The overall Pearson r was strong at .96 (State 1: r = .96; State 2: r = .86), indicating a pattern of similarity between assessors’ scores. The equivalence of assessors’ ECI total communication scores was estimated by calculating confidence intervals around the mean difference between reliability versus primary assessor groups, overall and within each state (Seaman & Serlin, 1998). Perfect equivalence is indicated by a 0 at the low end of the interval. Overall and for State 1, the low end of the confidence interval was very near 0, at .03 and .01 respectively, indicating equivalence. The similar equivalence value for State 2 was −0.68, not nearly as strong (see Table 3).
Interobserver Agreement Summary by State and Overall
Program-level scores
In addition to ECI total communication rate, we calculated aggregate sociodemographic and implementation quality scores for each program. The percentages of children with an IFSP, male gender, and non-English home language were the aggregate of the child-level values for each program. Implementation quality scores at the program level were also aggregate percentages of specific activities recorded in the website. They were defined as follows: (a) the number of ECIs collected by certified assessors in each program, divided by the total number of assessments collected by certified plus uncertified assessors in the program, times 100; (b) the number of ECI total communication scores large enough to be considered outliers as determined in an earlier study (i.e., greater than 45 per minute; Greenwood et al., 2006), divided by total number of scores times 100; (c) the number of ECI assessments in a program with a paired agreement check involving a second assessor, divided by all total administrations, times 100; and (d) the number of ECI assessments lasting the exact 6-minute standard duration, divided by total administrations, times 100. It is worth noting that the number of ECI assessments administered by certified assessors may have been underreported because some program directors may have failed to change an assessor’s status from uncertified to certified.
With respect to these four ECI implementation variables, it was desirable from a fidelity perspective that all assessors collecting data were certified, that all administered the ECI for the standard amount of time, and that all conducted interobserver agreement checks. Doing so would in theory result in a minimum of outlier ECI scores in the program’s database. Last, the state in which a program was located was available for use as a geographical location indicator.
Statistical Approach
The data set reported by Greenwood et al. (2010) was used in a secondary analysis. In this data set, ECI assessments were nested under children nested under assessors (home visitors) who were nested under programs. However, because the observed relationship between assessors and the children was not strictly a person-to-person match—given that assessors occasionally assessed each other’s children or new assessors replaced original assessors who were on leave or who left the program over time—we chose to represent only program, child, and assessments in the design plan.
In this data set, the percentage of outlier ECI total communication scores was calculated first using all available ECI data. However, in subsequent analyses described below, we deleted outlier scores. Because ECI total communication scores of zero were considered legal scores for infants and toddlers first learning to communicate, outliers included only score values that exceeded three standard deviations above the ECI total communication mean score at each month of age inclusive of 6 to 42 months. These scores were removed (i.e., 176 scores, or 1%). The median number of ECI observations per child in the data set was 3, ranging from 1 to 12. The mean number of data points available per each month of age at test was 428 (SD = 198), ranging from 23 to 727 (Greenwood et al., 2010).
The first and second research questions regarding the variability in programs’ sociodemographic composition and implementation quality were addressed with simple descriptive statistics, chi-square, and graphical displays. Because the percentages of sociodemographic variables were based on child-level counts, we used graphical display and chi-square to examine differences between programs. Because the implementation quality data were composites of staff program-level activity, only graphical displays were used.
To address the remaining two research questions we used hierarchical linear modeling (HLM 6.08; Scientific Software International, Inc., Lincolnwood, IL) to conduct growth curve analyses (Bryk, Raudenbush, & Condon, 1996). The estimation of individual and group trajectories (linear and curvilinear) over time in multilevel data made growth curve analysis particularly appropriate. It also allows testing of fixed and random effects and use of covariate variables (Burchinal & Appelbaum, 1991). Because growth curve analysis is tolerant of missing data, individuals may have different numbers of observations, and it accommodates unequal time intervals separating occasions (Raudenbush & Bryk, 2002).
We group-mean centered the intercept at 36 months of age because this age is the federal EHS endpoint of eligibility for services; it is also the endpoint of early intervention services under the Individuals with Disabilities Education Act, Part C. Thus, we were able to test differences in programs’ mean outcome scores at this age. Given prior findings (Greenwood et al., 2010), we used a quadratic (curvilinear) unconditional growth model to fit ECI total communication in Level 1 analyses. The children included in the analyses were those with sufficient Level 1 data for a curvilinear analysis as internally determined by the HLM software, n = 5,180.
The third question regarding the variation in programs’ mean ECI total communication growth trajectories and outcomes at 36 months of age was addressed using a Level 2 growth curve analysis with the repeatedly measured ECIs at Level 1 nested under children at Level 2. The descriptive statistics for these data are shown in Table 4. Level 1 analysis provided an unconditional growth model, followed by Level 2 conditional analyses designed to compute each program’s unique ECI total communication trajectory and compare its deviation from Program 15. We selected Program 15 as a standard for comparison because its mean intercept, slope, and acceleration parameters (21.3, 0.95, and .011, respectively) most closely matched these same mean parameters at the program level (n = 26; 22.6, 0.93, and .008). To conduct independent comparisons to Program 15 and to uniquely identify each program, individual programs were dummy coded using zeroes and ones (Cornell Statistical Consulting Unit, 2008). For this purpose, we used 25 dummy codes, 1 for each program, with Program 15 assigned to 0, establishing it as the comparison program. Additionally, because of the 25 paired comparisons between programs, alpha of .05 was adjusted to .002 (alpha = .05/25) to account for the number of comparison-wise errors. The chi-square comparison test of improvement in fit was used to test the statistical significance of this Level 2 growth model (Hox, 2001).
The fourth question, regarding program-level influences, was addressed using a Level 3 growth curve analysis. Descriptive statistics for these data are shown in Table 2. The appropriateness of using a Level 3 growth model with these data was explored by examining the design effect and the intraclass correlation coefficients present in the unconditional model—that is, the variance in mean ECI total communication explained by each level in the model.
Descriptive Statistics for Level 2 Analyses
n = 16,677.
n = 5,180.
Comparison program.
Upper panel, individual family services plan; middle panel, male; and lower panel, non-English. The horizontal bars mark the mean of programs in percentages.
The grand mean ECI total communication rate across all repeated measures, all children, and all programs was 10.98 communications per minute. The estimated ECI variances in residual errors were 39.02, 51.29, and 1.27 for repeated measures (Level 1, n = 16,677), children (Level 2, n = 5,180), and program level (Level 3, n = 26), respectively. These values were used to compute the intraclass correlation for each level in the model. These intraclass correlation coefficients were .56, .43, and .01 for Levels 1, 2, and 3, respectively. These results indicated that most of the variability in the data occurred at the ECI repeated measures level (56%), followed by the child level (43%) and the program level (1%).
However, the need for accounting for levels of clustering is not determined solely by the intraclass correlation; it is also determined by average cluster size. The design effect in multilevel data takes this into account (Muthen & Satorra, 1995). According to our intraclass correlation coefficients and average sample sizes, our Level 3 design effect was 2.99. Design effect values greater than 2.00 indicate the need to include all levels in subsequent growth curve analyses. A design effect this large indicates that our nested sample variance is about 3 times larger than it would be if it were not a nested sample (Muthen & Satorra, 1995). Said differently, this design effect meant that children were growing in ECI skills across time (Level 1), that some children performed better than others (Level 2), and that children in some programs were performing better than children in other programs (Level 3).
Selecting variables used in the Level 3 growth model of program-level influences was informed by examining the collinearity among the variables using Pearson r. Collinearity was indicated for percentage of IFSP and male gender, r(25) = .52, and suggested for data by certified assessors and percentage of outliers, r(25) = –.24. There was a tendency for (a) more males to have IFSPs and (b) more outliers in programs with less data collected by certified assessors. This eliminated gender and data by certified assessors from further consideration. Selection was also informed using HLM’s Exploration tool, which evaluated the potential of variables not yet in the model in terms of t score. At this step, non-English and IFSP were eliminated in favor of outliers and state location.
Model building first examined the improvement in fit provided by each separate program-level variable (i.e., outliers and state) compared to the unconditional model using the chi-square model comparison test for nested models (Hox, 2001). This was followed by combining both outliers and state in the model including their interaction (Outliers × State) and assessing increase in fit at this step.
Results
What Was the Variation in Sociodemographic Characteristics by Programs and State Location?
Programs did vary in aggregate sociodemographic composition (see Table 2 and Figure 1). Variation in IFSP status was sufficiently large that programs differed, M = 8.4%, SD = 5.8, ranging from 0% to 20%, χ2(25) = 209.51, p = .0001. Gender showed the least variation, where the program-level mean was 50.5% male, SD = 6.5, ranging from 40% to 70%, with no significant differences between programs. As with IFSP, variation in the percentage of children who had languages other than English spoken at home was sufficiently large to be significant, M = 8.0%, SD = 13.7, ranging from 0% to 57.5%, χ2(25) = 1,085.28, p = .0001.

Composition of Programs’ Subpopulations
At the state level, both IFSP status and non-English were significantly different (see Figure 1, lower panel). The percentage of non-English children was lower in State 2 (e.g., Programs 15–24, 2.7%) compared to State 1 (e.g., Programs 2–14, 15.4%), χ2(1) = 209.51, p = .0001. The percentage of children with IFSPs (Figure 1, top left panel) also was lower in State 2 (5.7%) compared to State 1 (9.6%), χ2(1) = 31.47, p = .0001.
What Was the Variation in Implementation Quality by Programs and State Location?
With respect to variation in program-level implementation indicators, standard ECI administration time and the percentage of outliers were the least variable (see Figure 2). In the case of ECI administration time, it was high and stable across programs, M = 97.2%, SD = 4.1, ranging from 83% to 100%. In the case of outliers, it was very low and stable, with the exception of Program 33, M = 2.1%, SD = 6.6, ranging from 0% to 34%. These patterns for both variables were suggestive of uniformly high implementation quality. Variation in programs’ use of agreement checking was also relatively uniform, but with a mean of only 2.8% (SD = 4.4, ranging from 0% to 21%), agreement checking was simply not widely implemented by programs. The percentage of data by certified assessors was the most variable of all and only moderately implemented across programs (M = 59.6%, SD = 33.9, ranging from 0% to 100%). Six programs in all had no data collected by certified assessors (Programs 8, 9, 22, 23, 29, and 33) according to website records kept by the program director. No significant differences were found between states on these variables.

Programs’ Implementation Quality Indicators
What Was the Variation in ECI Total Communication Slope, Acceleration, and 36-Month Outcomes Across Programs?
The unconditional model for ECI total communication indicated a mean intercept of 22.29 communications per minute at 36 months of age, a slope of 0.98 communications per minute per month, and an acceleration of 0.010 communications per minute per month. The Level 2 analysis produced intercept, slope, and acceleration growth parameters for each individual program, producing 26 separate individual trajectories (see Figure 3). Based on the number of program trajectories, they were graphed by state in Figure 3 for visual clarity. The chi-square model comparison test indicated that, compared to the unconditional model, this Level 2 model that included the 26 programs significantly improved fit to the data, χ2(75) = 423.75, p = .0001.

Programs’ Early Communication Indicator Total Communication Growth Trajectories Grouped by State
Visual inspection of the individual programs’ growth patterns indicated relative similarity in slope, acceleration, and 36-month outcomes, as trajectories were relatively tightly clustered over increasing age (see Figure 3). As expected, review of each program’s deviations in growth parameters compared to those of Program 15 indicated variability (see Table 5). However, only 3 of the 25 comparisons to Program 15 (Programs 2, 5, and 33) were statistically significant at α = .002 for mean intercept at 36 months.
Children in Program 2 at 36 months of age underperformed Program 15 by −4.9 communications per minute, while Programs 5 and 33 outperformed Program 15 at 4.7 and 15.8 communications per minute more, respectively. Only Program 33 was close to significance for slope (p = .003), and none were significantly different for acceleration. Program 33’s mean intercept reflected actual values of 37.06 total communications per minute compared to 21.26 (Program 15) and 2.21 total communications per minute per month compared to 0.95 (Program 15) in slope.
Growth Parameters by Program Using Program 15 as the Comparison Program
Note: The coefficient values are deviations from the mean intercept, slope, and acceleration coefficients of the comparison program, P15. df = 5,154.
= Programs whose intercepts at 36 months were significantly different (α = .002) from the comparison program, P15.
To What Extent Were ECI Total Communication Slope, Acceleration, and 36-Month Outcomes Influenced by State Location, Sociodemographic Characteristics, and Implementation Quality?
Results of the Level 3 modeling indicated that variation in programs’ subpopulations of children with IFSPs and non-English home language were not significant influences on programs’ ECI total communication growth parameters (i.e., mean intercept, slope, or acceleration). Similarly, the percentage of data collected by certified assessors, 6-minute ECI administration durations, and use of agreement checking were not significant influences. However, the percentage of outliers was the strongest program-level implementation influence on 36-month outcomes, slope, and acceleration, followed by the state in which programs were located (see Table 6). The combination of Outliers + State proved to be the best fitting model overall. The Outliers × State interaction effect did not significantly improve model fit and was not included in the final model. The growth parameters for the unconditional Level 3 model and the best-fitting program influences model are seen in Table 7. Outliers significantly influenced intercept and acceleration (p = .001) but not slope, while state influenced slope (p = .001) and acceleration (p = .002) but not for the mean intercept at 36 months. This meant that every one-unit increase in outliers resulted in an increase in ECI total communication at 36 months of 0.48 communications per minute, in slope by 0.035 responses per minute per month, and in acceleration by 0.0001 communications per minute per month. This also meant that programs in State 2 compared to State 1 had greater slope and acceleration, producing a trajectory initially starting out slower and accelerating faster (see Table 7 and Figure 3).
Model Comparison Summary
Note: df = 3.
Not significant.
p = .001.
Best-Fitting Program-Level Growth Model
df = 1,431.
df = 23.
Discussion
The purpose of our investigation was to add to the evidentiary knowledge of children’s development of early communication skills as measured by the ECI in a secondary analysis of a large sample of ECI assessments, children, program staff, and EHS programs in two Midwestern states with multilevel ECI data collected by EHS staff with website support (Greenwood et al., 2010). The investigation centered on the question “Under what program-level conditions does the ECI adequately describe a child’s growth and outcome at 36 months of age?” Specifically, we investigated the variation in (a) program-level sociodemographic characteristics and (b) the fidelity of ECI implementation and (c) the relationships between these variables and ECI total communication growth trajectories (i.e., 36-month outcome, linear slope, and acceleration). We hypothesized that ECI scores would be sensitive to programs’ and states’ serving EHS subpopulations varying in gender, non-English home languages, and IFSP status composition. We hypothesized that program-level ECI scores would be influenced by the quality and fidelity of its use by local program staff. Both program-level and state-level factors are understudied areas. Results of the investigation informed our knowledge of the situational nature of ECI score validity in large-scale applications with early childhood staff implementing, supervising, and using the information for accountability and decision making.
Using a two-level growth model, we were able to examine 26 individual programs’ ECI total communication trajectories and to compare each to a normative, standard trajectory. Programs varied in ECI total communication trajectories within a reasonable range, as shown in Figure 3. Worth noting was each program’s ability to produce a trajectory with similar shape (i.e., slope and acceleration) and 36-month outcomes as predicted by prior ECI research and contemporary language development theory. Variation in shape was expected; however, a trajectory showing negative slope or acceleration, for example, would cause concern. No such trajectories were observed. Yet three programs’ mean intercepts at 36 months were significantly different from that of the comparison Program 15, suggesting concern with the internal measurement process as implemented and/or with subpopulation differences in children’s characteristics. Three programs (2, 5, and 33) showed significant deviations from Program 15’s standard, representing extreme cases. Program 33’s deviation from Program 15’s 36-month mean was most extreme at 15.8 communications per minute higher.
We also calculated four indicators of program-level ECI implementation quality based on programs’ website records and described their variation across programs. These indicators were the percentages of (a) data collected by certified assessors, (b) standard 6-minute ECI administration durations, (c) agreement checking between assessor pairs, and (d) outlier score values. Programs varied least in ECI administration duration and were uniformly high. Programs were also uniformly low in the number of outliers they produced, with the exception of Program 33. These two indicators showed that programs generally met high standards. Both agreement checking and data collected by certified assessors, however, were variable and often did not meet desired standards. Only Program 18 engaged in agreement checking at 20%, a frequency considered a high standard; the rest did relatively little or no checking at all, particularly in State 2. Data collected by certified assessors varied widely, ranging from 0% to 100% across programs, with a mean of 59% overall (see Figure 2). States did not differ on these variables; however, analyses of interobserver agreement indicated that State 2 assessors were lower in percentage agreement statistics (see Table 3).
Growth curve analyses indicated that even though programs varied in child subpopulation characteristics, they did not significantly influence ECI scores. However, the states in which programs were located did explain significant variation in the slope and acceleration of ECI trajectories. Similarly, the percentage of outliers explained significant variation in ECI scores. Programs producing greater numbers of outliers produced ECI trajectories with significantly higher 36-month outcomes and greater positive acceleration. The model containing both state and outliers proved to be the most parsimonious, best-fitting model.
Inspection of the convergence of the sociodemograhic and fidelity factors for these particular programs indicated that Program 33 also had an extreme number of outliers, that it did no agreement checking, and that only 10% of data were collected by certified assessors (see Table 5). Outliers were not a concern with the other two programs; however, they did no agreement checking, and only 40% and 50% of these data were collected by certified assessors. These same factors for the comparison Program 15 indicated generally better fidelity in comparison and no extreme values. Sixty-five percent of Program 15’s data were collected by certified assessors. It was also interesting to examine the convergence of these factors for programs whose mean intercept values were not significantly different from Program 15’s. Even for these programs, relatively few did agreement checking, and data by certified assessors were variable, sometimes very high and at other times low. Nearly all programs administered the ECI for the standard time and produced acceptably low numbers of outliers.
Using a three-level growth model, we were able to describe the variation in programs’ subpopulations of children with an IFSP, male gender, and non-English languages heard at home. Programs varied least in gender and most in IFSP status and non-English languages. Programs in State 2 served significantly fewer numbers of non-English-speaking children and fewer numbers of children with IFSPs than did State 1. By way of explanation, the finding that outliers inflated ECI 36-month intercepts and acceleration scores was somewhat expected, as by definition, outliers were inflated scores. Thus, efforts in future research and development need to focus more directly on the prevention of outliers in ECI data, and doing so should be effective in reducing this influence. Not nearly as clear-cut, however, was the relationship between the other ECI implementation quality indicators (i.e., certified assessors, administration duration, and agreement checking) and their influences on outliers and ECI scores. In this work, we looked simply at the direct relationship between the quality indicators and ECI total communication scores. An even more sophisticated approach would be to combine them in some ways to form a cumulative quality index, for example, in the same way as is often the case with research on the effects of cumulative risk (Liaw & Brooks-Gunn, 1994). Another approach would be a framework of ECI implementation fidelity wherein outliers play a moderator/mediator role in determining the relationship between indicators of staff implementation quality, the number of outliers, and inflationary effects on ECI scores. These remain topics for future research; however, the present findings argue strongly for the need for continual monitoring of program-level ECI implementation factors as well as other situational factors that may influence assessment over time. Just two of these situational factors are programs’ use of evidence-based practices and the fidelity of their implementation.
Results indicated that the state in which programs were located also influenced ECI scores, while program-level sociodemographic subpopulation differences did not. This was interesting because State 2 programs were on average serving fewer children who had IFSPs and were receiving early intervention and EHS services and fewer children hearing languages other than English at home, compared to State 1 programs. In this study, we were not able to untangle these relationships in association with different locations. Arguably, state is an imprecise indicator of location, but presumably, it and more regionally defined locations have meaning in association with multiple sociodemographic factors.
Limitations of the Study
While these findings advanced what is known about the implementation and validity of ECI measurement at the program level, a number of limitations must be considered. First, the website at the time of the study allowed assessors to collect, enter, and use ECI data once they were registered as assessors in the website but, in some cases, before being documented as certified (i.e., trained and meeting the standards of administration). This was partly a concern over the question of whether certification would be such a burden that it might delay or hinder large-scale data collection. As reported above, significant data were collected by assessors who were not documented in the website as being certified. The extent to which assessors were certified as described and their status simply not changed in the website, we do not know. Clearly, this could affect interpretation of findings.
Additionally, there was concern that frequency of agreement checking of these data was not rigorously updated or monitored by responsible local staff. The number of interobserver agreement checks conducted was small overall, and in some cases none were collected for some programs. Lack of this information introduces the potential for drift in accuracy in an ongoing data collection for many children over months and years. Because agreement checking is both a cost and an effort burden on programs, research is needed to generate methods of documenting and maintaining agreement between and among program staff. Other quality indicators (e.g., administration duration, outliers), however, seemed much more trustworthy because they were direct products of the data entry process necessary to add new children’s data into in the system.
Clearly, high-quality implementation is a determinant of ECI data quality and the inferences made from the data. A typical concern with most large-scale data collection efforts is that, because the data are collected and used completely by program staff working in local programs, the data collected by program staff introduce bias and inaccuracy, which is not the case when data are collected by research or evaluation staff. This outcome was supported by the performance of Program 33. Future research needs to focus on improving the accuracy and reliability of implementation quality indicators and, where possible, web solutions made to improve local monitoring, provide feedback on procedural errors, and concerns about data values and lagging certification issues, as well as reducing the burden to staff of using the system in a high-quality manner. Many improvements of this kind have been made to the website or are planned. One in particular was making the certification process possible online, where it is possible for the website to record certification records in the database as they occur.
Implications for Research and Practice
Based on the volume of data, the number of participants, and the long study duration, it was clear that the general approach to large-scale implementation and use of ECI data was successful in a sample of programs whose average location from the developers was 143 miles, ranging from 1 to 337. Programs were able to generate a rich source of useful data regarding children’s development of early communication skills for use in individual decision making, program improvement, and statewide accountability. This success was due to local administrative support for use of the ECI. Also contributing was the standard training, the ECI protocol, and the common website providing local program assessors with a uniform infrastructure for access and support such that large-scale implementation was achieved and continuity in the measurement process maintained over time.
A lesson learned in this work was the importance of both researchers and practitioners using the ECI and other large-scale accountability measures continuing to examine the validity of the data being collected. As illustrated here, this examination needs to be in terms of situational factors known to influence (or potentially influence) ECI scores. The analyses of ECI data did identify at least three programs whose values were significantly under and over expectation. In the fidelity data, this was helpful in pointing to quality indicators that were uniformly high, low, and variable. Such data reviewed by a program director, a state-level director, or a researcher has implications for monitoring implementation and quality control over the data collection. In our case, this examination happened only retrospectively, after the data were collected, and not used as a basis for discussion and improvement of identified programs’ implementation. The findings also support continued research, adding to the knowledge base of characteristics that affect assessment over time and understanding the ways that results may be situationally dependent.
Conclusion
The emergence of the ECI and its website infrastructure make it increasingly feasible for local programs to implement and manage the use of universal, programwide data on growth in expressive communication for use in individual intervention decision making and program accountability. This was a first look at program-level influences on growth in ECI total communication in a large-scale context. These findings strengthen and advance what is known about growth in EHS children’s early communication skills measured by the ECI at the program level. The subpopulation compositions of programs were not factors influencing measurement at the program level; however, ECI implementation factors varied widely and in some cases influenced program trajectories (e.g., percentage of outliers). These data collected by program staff clearly show room for improvement in implementation factors, and the tools exist for making this improvement. These findings at the program level of analysis provide new insights into the use and validity of the ECI as well as other measurement practices implemented large scale for accountability and universal screening and decision making.
Footnotes
All authors are affiliated with the Juniper Gardens Children’s Project, University of Kansas. This work was supported by the Office of Special Education Programs (H324C040095, H327A060051) and the Institute of Education Sciences, National Center for Special Education Research (R324A070085), U.S. Department of Education. Additional support was provided by the Kansas Intellectual and Developmental Disabilities Research Center, National Institutes of Health (HD002528), Schiefelbusch Institute for Life Span Studies, Kansas Social Rehabilitation Services, and the regional Early Head Start Association. We gratefully acknowledge the contributions of colleagues Judith Carta, Debra Montagna, Barbara Terry, Christine Muehe, Susan Higgins, Matt Garrett, and Chia-Fen Liu. We thank Todd Little of the Kansas Center for Research Methods and Data Analysis for assistance with the growth curve modeling aspects of the study. We acknowledge our Kansas and Missouri Early Head Start program research partners, participating programs, children and families, and Mary Weathers for her support in facilitating partnerships with Early Head Start.
1.
The infant and toddler website may be found online at http://www.igdi.ku.edu. The preschool website may be found online at
.
