Abstract
We explored whether and how cognitive measures of executive function (EF) can be used to help classify academic performance in Kindergarten and first grade using nonparametric cluster analysis. We found that EF measures were useful in classifying low-reading performance in both grades, but mathematics performance could be grouped into low, average, and high groups without the use of EF tasks. Membership in the high-performing groups was more stable through first grade than membership in the low or average groups, and certain Kindergarten EF tasks differentially predicted first-grade reading and mathematics cluster membership. Our results suggest a stronger link between EF deficits and low performance than between EF strengths and high performance. We highlight the importance of simultaneously using academic and cognitive skills to classify achievement, particularly since existing classification schemes have been largely based on arbitrary cutoffs using limited academic measures.
Keywords
An enduring goal in education research is distinguishing between students with low academic achievement who ultimately receive special education services and those who would be considered typically developing and thus ineligible for special education (e.g., Fuchs, Fuchs, Mathes, Lipsey, & Roberts, 2001; Hoskyn & Swanson, 2000; Ysseldyke, Algozzine, Shinn, & McGue, 1982). However, existing techniques for classifying or tracking academic performance largely rely on arbitrary cut points, sometimes using only one measure to group students (e.g., Kim et al., 2016; Moser, West, & Hughes, 2012). Serious questions remain regarding whether these techniques are the most effective means of classifying performance, despite a century’s worth of research on ability grouping (Steenbergen-Hu, Makel, & Olszewski-Kubilius, 2016). The current study explores how we might better classify low-performing students within the first 2 years of schooling using a wide variety of academic performance and executive function (EF) measures.
Low academic performers who are otherwise classified as typically developing are at risk for poorer educational attainment and school drop-out, which could lead to limited access to better-paying jobs, poorer health, and less social–political participation (Battin-Pearson et al., 2000). Low academic performance is often also accompanied by cognitive deficits in EF (Best, Miller, & Naglieri, 2011). Yet, whether and to what extent EF measures can be used to help classify academic performance among low-performing students without disabilities remains unclear. Although cluster analysis has been levied as a method that may help classify low-performing students to maximize accurate educational tracking, there has been limited research utilizing this methodology for these purposes beyond its application to select subsamples of children with disabilities (Loehlin, Wright, Hansell, & Martin, 2018). Therefore, we conducted an exploratory analysis evaluating the relations between EF tasks and low, average, or high academic performance to assess how using EF tasks might change the way these performance designations are made.
Classifying Academic Performance Beyond Identifying Disability
Researchers have long focused on how the co-development of cognitive and academic abilities helps identify, track, and group students with different skill levels. Some have argued that poor readers and students with learning disabilities share cognitive deficits irrespective of general cognitive abilities, which evinced the idea that low achievers are not clearly separable from students with disabilities (Fuchs et al., 2001; Hoskyn & Swanson, 2000; Siegel, 1992; Ysseldyke et al., 1982). While the bottom 10th percentile has been used to conservatively identify children who may have a disability (Morgan, Farkas, Hillemeier, & Maczuga, 2012), there is no “gold standard” for identifying children with learning disorders (Skibbe, Justice, Zucker, & McGinty, 2008). The lack of agreement on a definition for learning disabilities beyond “unexpected” or “specific” learning failure has led some to call for more objective assessments of performance during the identification process (Fuchs et al., 2001).
Today, many studies investigating performance classifications use a cutoff measure to differentiate low, average, and high performance among typically developing students. However, this cutoff is not standard across studies. Researchers have conceptualized “below average” or “low performing” students as those who perform below the median (e.g., Kim et al., 2016; Moser et al., 2012), below the 20th to 30th percentile (e.g., Geary, Hoard, Byrd-Craven, & DeSoto, 2004; Jordan, Hanich, & Kaplan, 2003; Navarro et al., 2012), or beyond one standard deviation from the sample mean (e.g., Gathercole et al., 2016; Skibbe et al., 2008).
The problem with the lack of standardization in grouping criteria is that low- or average-performing groups could be displaying average or high performance (respectively) in a study with a different classification scheme (e.g., tertiary sample splits of performance versus standard deviation cutoffs), which hinders replicability. This lack of standardization is further compounded by the fact that cutoffs may be made using only one measure at a time instead of examining performance more generally across several academic and nonacademic tests.
Educational Tracking
Classifying student performance has also long been used for the practice of educational tracking, or assigning students to groups based on their perceived academic ability under the guise that students will learn better when matched with students of similar skill levels (Ansalone, 2010). It is most commonly applied in the case of advanced, accelerated, or gifted students, but some also consider special education a unique form of tracking given that students receive these services in part based on academic ability (Lipsky & Gartner, 1989).
Tracking relies on teachers’ ability to place students into the tracks that most appropriately reflect their skill levels. Yet, whether and to what extent teachers are able to do this effectively and accurately is unclear. There has also been controversy and scholarly debate regarding the utility of educational tracking (e.g., see Oakes, 2005). For instance, whether tracking decisions are accurate is highly debated, but this is important to consider given that most students remain in their academic track once placed (Eccles, 2004; Oakes, 2005). Although grades and standardized test scores are considered the best predictors of tracking decisions (e.g., Hallinan, 1992; Southworth & Mickelson, 2007), teachers may be biased in their judgment of whether students should be tracked, especially when students also show inconsistent academic profiles (e.g., discrepancies between grades and standardized tests; Glock, Krolak-Schwerdt, Klapproth, & Böhmer, 2013). This means that there are practical questions regarding how likely it is that a low-performing student continues to display low performance, continually begetting a low academic track. Uncovering more objective and precise methods of classifying students, particularly those that do not rely on arbitrary cutoffs, may reduce bias, improve tracking decisions, and help resolve the debate regarding whether and to what extent we should rely on educational tracking in our schools.
Can EF Help Classify and Track Academic Performance?
Years of research have demonstrated substantial relations among young children’s EF skills and emerging and persistent mathematics and literacy achievement (e.g., Blair & Razza, 2007; Gathercole & Pickering, 2000; Jacob & Parkinson, 2015; Lan, Legare, Ponitz, Li, & Morrison, 2011; McClelland et al., 2007). EF refers to an individual’s ability to complete tasks and purposefully guide their mental thoughts and behaviors to achieve certain goals (Cartwright, 2012). EF skills include working memory (WM), attention control, and response inhibition (Blair, 2002). Strong EF skills are necessary in a learning environment where students are expected to pay attention, follow rules, and concentrate on both cognitive and behavioral tasks (Blair, 2002; Blair & Razza, 2007). Although stronger EF skills allow students to better adapt to the demands of early classrooms and schooling, deficits in EF have also been linked to reduced academic performance. This is particularly true for the relationship between WM and general achievement (Ahmed, Tang, Waters, & Davis-Kean, 2019; Best et al., 2011; Gathercole et al., 2016), WM and mathematics performance (Geary et al., 2004), attentional control and reading performance (Lam & Beale, 1991), and response inhibition and both reading and math (Blair & Razza, 2007; Cameron et al., 2012). Each of these EF components have also been linked to specific aspects of both literacy and numeracy, such as print knowledge (Purpura, Schmitt, & Ganly, 2017), vocabulary and phonological awareness (Allan & Lonigan, 2011), as well as counting, cardinality, subitizing, and set comparison (Purpura et al., 2017). In a recent meta-analysis, Allan, Hume, Allan, Farrington, and Lonigan (2014) found a moderate association (r = .34) between children’s inhibitory control and performance on standardized tests of math achievement. Furthermore, a study of attention control showed that kindergarteners with better attention scores outperformed students with poorer attention skills on standardized measures of math achievement (Howse, Lange, Farran, & Boyles, 2003). Blair and colleagues (2015) reported that attentional control predicted growth in scores on standardized tests of achievement from preschool to second grade, over and above demographic and early achievement covariates. Overall, there is ample evidence that EF may be closely connected to academic achievement, which provides even more impetus for considering these measures when classifying early performance.
The present study directly tested the relative contribution of EF to both low and high academic performance using cluster analysis, a discontinuous grouping methodology. Cluster analysis has been frequently used to establish groupings along many domains in educational research (e.g., Myers & Fouts, 1992; Ronning, 2004; Ullrich-French & Cox, 2009). Given the highly predictive relation between achievement and EF, and yet the lack of research using both cognitive and academic measures to group and track students without disabilities, we assessed the benefits of combining these measures to classify low performance.
Research Aims and Hypotheses
We explored classification strategies similar to those often used in both research and education (e.g., for determining special education eligibility). Although cutoff methods have typically been used in studies classifying academic and/or cognitive performance, the present study is novel in its application of nonparametric cluster analysis methodology. This allows us to differentiate students who display non-arbitrarily dissimilar performance from peers. We hope that this endeavor may help researchers and practitioners alike to classify performance with better precision and predictive validity.
Our first research aim was to identify the EF tasks that contribute to the differentiation of low, average, and high academic performance in early schooling. Our secondary aim was to assess the longitudinal stability of cluster membership. This latter aim is especially important given research indicating that students often remain in their educational track once placed and that tracking decisions are not free of bias. Our two hypotheses were as follows. First, because research has demonstrated that students with disabilities experience a persistent failure to learn, we expected to see a small group of students making little academic growth over time and maintaining stable low performance. Given the intensive intervention needed to remediate learning failure, we expected that membership in this low-performing cluster would more stable than membership in the average- or high-performing clusters. Second, in keeping with prior research, we expected that EF skills would both differentiate clusters and predict later cluster membership (i.e., that Kindergarten EF would predict first-grade clusters). We hypothesized this because attentional control and response inhibition have previously been found to influence both math and reading ability, and WM has been consistently demonstrated to more strongly influence math than reading (e.g., Blair & Razza, 2007; Nguyen & Duncan, 2018). Thus, we expected to see similar results.
Method
Sample
Data were drawn from two cohorts of children followed from Kindergarten into first grade (n = 120; Kindergarten Mage = 5.8 years; first grade Mage = 6.8 years) as part of a longitudinal study investigating the transition to school. The study included four socioeconomically and racially diverse schools in one Midwestern state. Socioeconomic disadvantage 1 in the student body ranged from 2.1% (School 1) to 87.4% (School 4), with a sample average of 56.1%. Nearly 58% of our sample was male. There was also considerable racial/ethnic diversity within and between sampled schools, with study participants attending schools that were on average 36.8% White, 44.9% African American, 3.7% Hispanic, and 12.1% Asian. Detailed descriptive sample statistics both within and across schools are presented in Appendix A.
Procedure
Approval for human subjects research was obtained by an Institutional Review Board prior to data collection, and parents of participants and teachers provided informed consent for supervised testing. Each child was assessed by trained research assistants once in the spring of Kindergarten and again in the spring of first grade using a battery of individually assessed EF tasks in quiet, unused, or multipurpose rooms in schools. Standardized tests of academic achievement and EF skills were administered to all participants. Teachers also provided information about whether study participants were observed with concerns regarding a potential special educational need and/or whether study participants received special education services during the years of data collection. On average, 26% of the study sample were observed during Kindergarten or first grade. Concerns for observation were low academic performance (15% of observations), behavior or emotional issues (11%), speech or language problems (3%), potential Autism (2%), or other concerns (2%). Of the low academic concerns, 53% were for low reading, 13% were for low mathematics, and 33% cited both reading and mathematics issues. Eleven percent of the sample received special education services in Kindergarten or first grade. Services were delivered for speech/language impairments (3%), learning disabilities (2%), emotional impairments or attention-deficit/hyperactivity disorder (ADHD) (2%), Autism (3%), or some other type of services (2%). Interestingly, though School 3 reported the most observational concerns about students (46%), they also reported the fewest students receiving services (8%).
Missing Data
Of 140 students sampled in Kindergarten, 120 students were tested again in first grade. The amount of data missing from each subtest ranged from 0% to 6.7% (see Table 1), with one exception: the number line task was missing 28.3% of data due to a protocol change implemented after 20 Kindergarten students had been tested, nullifying their scores. Twenty-two students were missing data on certain subtests due to school absences impeding data collection or misbehavior/inattention during testing that rendered data unusable, and eight of these students (36%) were reported to have been observed for or received special education. Unfortunately, data collectors experienced great difficulty obtaining individual demographic information from parents through online surveys, despite repeated efforts and incentivizing. Thus, 69% to 74% of variables measuring participants’ race/ethnicity, maternal education, income, and free/reduced lunch participation were missing data, and this was concentrated in Schools 2 to 4. For this reason, we provide school-level demographics in Appendix A, though we recognize this as a significant study limitation.
Descriptive Statistics for Analytical Measures.
Note. WJ-III Standard Scores are displayed.
Measures
A variety of academic and behavioral assessments were used to measure reading, math, and EF in both Kindergarten and first grade (see Table 1 for descriptive statistics and Appendix B for correlations). Select tests were administered from the Woodcock Johnson-III Tests of Achievement (WJ-III; Mather & Woodcock, 2001) and WJ-III Tests of Cognitive Ability (Woodcock & Mather, 2000). These types of standardized, normed tests are often used to identify gifted students and/or ability–achievement discrepancies for special education referral, so they are ideal for classifying performance clusters. Raw scores on the WJ-III tests were converted into standard scores (which are nationally norm-referenced to M = 100, SD = 15) based on when in the school year testing occurred. The same tests were given in both Kindergarten and first grade, and the testing battery was counterbalanced both within and across students.
Academic achievement
Reading performance was measured using two WJ-III tests: Letter-Word Identification, which assesses basic reading ability, and Passage Comprehension, which measures the ability to identify words in a sentence using contextual clues. Mathematics performance was assessed using two tasks: the WJ-III test of Applied Problems, which captures skill in analyzing and solving practical word problems; and the Number Line task (adapted from Siegler & Booth, 2004), which is often associated with general mathematics achievement. This test assesses students’ sense of number magnitude by asking students to mark the location of numbers on a line ranging from 0 to 20. Researchers then measured the distance away from the number’s true location in centimeters such that higher scores correspond to a poorer sense of number magnitude (e.g., a score of 3.50 would indicate that the student’s mark was three and a half centimeters away from the true location of the assessed number, and would be considered a poorer score than a score of 1.50). These scores were then averaged together across 16 number trials (one for each number 1–19, excluding 5, 10, and 15, which were used as trials). The sample range was 0.79 to 9.48 in Kindergarten, and 0.59 to 8.09 in first grade. Reliabilities were 0.87 in both Kindergarten and first grade.
EF
Attention Control
The Pair Cancellation task, which is a measure drawn from the Woodcock-Johnson III Tests of Cognitive Abilities (Woodcock & Mather, 2000), was used to test children’s attention control. In this task, children were presented with a testing sheet with small pictures of dogs, balls, and cups, and asked to circle all of the ball–dog pairs in which a dog is presented after a ball. After practicing to ensure that the child understood the task, they were given 3 min to complete the rest of the page, working as quickly as they could without making mistakes. There were 69 correct pairs, and the number of correct pairs identified within 3 min was recorded. In this study, we used W scores representing children’s ability level based on the Rasch measurement model, which provided comparative scores for each child, regardless of age. Test–retest reliability for this subtest is r = .78 (Mather & Woodcock, 2001).
Response inhibition
Response inhibition was measured using the Head-to-Toes, Knees-to-Shoulders task (HTKS; Ponitz et al., 2008), a game like Simon Says, in which students are instructed to touch the opposite body part than the one named by the researcher (e.g., when the researcher says to touch their head, they touch their toes). The task became increasingly challenging across 30 trials in three blocks (touching heads and toes, touching knees and shoulders, and a mix of the two). Incorrect responses were scored with a 0, 1 for a self-corrected response, and 2 for a correct response. The maximum score was 60. The internal consistency in the current study (Cronbach’s α) is 0.83 in Kindergarten and 0.73 in first grade. Reliability among overall scores obtained by different experimenters was 100%.
WM
The Counting Span subtest of the Weschler scales (1991) was administered to gauge short-term and WM skills, as the task requires the recall and manipulation of information stored in memory. This task was composed of two sections in which the instructor says a list of numbers and the participant was asked to recite the numbers forward and backwards. The list of numbers increases by one item for every correct response, and the largest set was six. If the participant answered incorrectly twice in a row, the researcher moved onto the next section. A score was assigned based upon the largest set at which the child successfully reported. This measure demonstrates acceptable test–retest reliability (r = .73; Lipsey et al., 2017). Short-term memory (forward digit span) ranged from 2 to 11 in Kindergarten and 4 to 11 in first grade, while WM (backward digit span) ranged from 0 to 7 in Kindergarten and 0 to 8 in first grade. Although this is a common WM test used in educational research, it is important to note that it only assessed verbal WM and not visuospatial WM.
Analytic Methods
Cluster analysis is a data-driven approach that allows students who are most alike to naturally cluster together into multidimensional groups instead of imposing group membership upon them. Heterogeneous samples are reorganized into smaller, more homogeneous clusters that are maximally similar within-group and maximally dissimilar across groups (Aldenderfer & Blashfield, 1984; Antonenko, Toy, & Niederhauser, 2012; Ronning, 2004). Hierarchical clustering algorithms sequentially combine each case with other clusters to construct a hierarchy of nested groups. They are useful when researchers do not have a preconceived idea about how many clusters to expect and want to explore how data cluster with few constraints, or if they have a relatively small sample size (Antonenko et al., 2012). Our goal was to determine whether three groups naturally emerge from data collected in early elementary samples. We were interested in this question because many previous studies had imposed artificial cutoffs to create three distinct groups for analysis of “low” versus “average” or “high” performers, but whether this approach is valid is unclear (i.e., whether these data cluster in this manner without constraint). Moreover, in the event that three distinct groups did not emerge from the data without constraints, we questioned whether it was possible to obtain a three-group solution by including measures of EF given their relation to academic achievement. This second aim was important because although EF has been suggested to be highly related to both strengths and deficits in academic performance, it is empirically unclear whether measures of cognitive skills can aid in distinguishing both low and high performance. If three groups did not emerge with or without the use of EF in these exploratory cluster analyses, then it is possible that there are not three distinct groups of low, average, and high performance during early schooling, which is an important consideration for both educators and researchers alike. We conducted an exploratory two-step cluster analysis in SPSS v. 24. This is a method that first groups and organizes the data in a quick pass and then applies a hierarchical clustering algorithm to arrive at a final solution. Cluster proximity was assessed using the log-likelihood linking algorithm given differences in subtest scales, and cases were arranged in a random order prior to clustering.
Results
Classifying Performance Using Cluster Analysis
Our main goal was to identify the cognitive and academic tasks that differentiate early schooling performance into low, average, and high groups without artificially imposing group membership upon students (e.g., using cutoffs). We entered different combinations of both cognitive and academic assessments into the clustering algorithm in an exploratory attempt to discover which combinations naturally produced three performance groups. First, the eight subtests were standardized to a sample mean of 0 and standard deviation of 1 to permit comparisons across task. Cluster analysis of all eight tasks resulted in only two clusters corresponding to above- and below-average performance. We suspect that these eight tests may oversaturate the clustering algorithm such that meaningful patterns in the data do not emerge. In other words, EF may function differently across tests of reading and mathematics, and these differences go unobserved when all eight tests are simultaneously included in a clustering algorithm.
We then separated the reading (Letter-Word Identification and Passage Comprehension) and mathematics tasks (Applied Problems and Number Line) and ran a series of cluster analyses on each in combination with the EF tests (Pair Cancellation, HTKS, DS-F, and DS-B) to ascertain what combination of measures produced groups of low, average, and high performance in reading and mathematics, respectively. Because most prior research imposes these three-group solutions onto their data, we were interested in understanding whether it was possible to obtain three groups either alone or in combination with these EF tasks. Cluster sizes are represented in Figure 1, and mean subtest scores within each cluster are represented in Figures 2 and 3. The results are as follows.

Cluster sample sizes by subject and by grade.

Mean z-scores of the tests used to create reading cluster membership at Kindergarten and first grade.

Mean z-scores of the tests used to create mathematics cluster membership at Kindergarten and first grade.
Reading
A three-cluster solution emerged for the Kindergarten reading data when combined with all four EF measures (N = 110; see Figure 1); cluster membership was driven primarily by performance on the Letter-Word Identification task, followed by DS-B, Passage Comprehension, HTKS, DS-F, and Pair Cancellation. Most students displayed relatively average performance (n = 53), while smaller groups displayed lower-than-average (n = 38) and higher-than-average (n = 19) performance. Figure 2 displays standardized mean task performance for each group within the Kindergarten reading cluster. Notably, the average-performing group had slightly below-average mean scores on the two reading subtests, but slightly higher-than-average performance on all four EF subtests. This may indicate that EF more strongly differentiates low from average readers. Indeed, when the EF tasks were removed from the cluster analysis, the algorithm was unable to differentiate average from low performance and produced only two clusters of above- and below-average students (Appendix C, see Figure C1).
This same pattern was evident in first grade, though a three-cluster solution was observed only for a combination of Letter-Word Identification, Passage Comprehension, DS-F, and DS-B (N = 119). Including HTKS and Pair Cancelation in the cluster analysis produced a two-cluster solution (see Appendix C, see Figure C2), suggesting that whereas all four EF measures helped to distinguish three different literacy-ability groups in Kindergarten, only the memory measures (DS-F and DS-B) continue to aid group differentiation into first grade. Unlike Kindergarten, Passage Comprehension primarily drove cluster membership, followed by Letter-Word Identification, DS-B, and DS-F. Most students again displayed average performance (n = 49), though there were relatively more high performers (n = 40) than low performers (n = 30) in first grade. The pattern of mean scores by subtest was also similar to Kindergarten (see Figure 2).
Several demographic features differed across reading clusters in both Kindergarten and first grade (e.g., male sex, school attended; see Table 2). Most relevantly, more than half of students who were observed for or received special education services were in the low-performing groups, though these statistics were only significant in first grade.
Demographic Differences by Cluster.
Note. Subscript denotes columns that significantly differ from each other at the p < .05 level.
Mathematics
A three-cluster solution emerged for math performance in both Kindergarten (n = 94) and first grade (n = 114) without including any EF subtests in the clustering algorithm. In both grades, Number Line contributed most strongly to cluster membership, followed by Applied Problems. There were no significant differences between the low (n = 28) and average (n = 49) Kindergarteners on the Applied Problems task, and no differences between the average and high (n = 17) Kindergarteners on the Number Line task (see Figure 3). This trend is similar in first grade, with no differences between the low (n = 32) and average (n = 60) performers on the Applied Problems task, and no differences between the average and high (n = 22) performers on the Number Line task. In other words, membership in the low groups at both grades seems to be characterized by poor number magnitude, while membership in the high groups is driven by very high Applied Problems scores.
Including all four EF tests in mathematics clusters produced a two-cluster solution of above- and below-average performance. Thus, we suspect that mathematics may be related to EF in a more domain-specific manner (i.e., related to distinct EF tasks like WM or response inhibition) than to global EF. To further explore this, each EF task was separately included with the two mathematics subtests. Only first grade HTKS (response inhibition) differentiated groups beyond an above- or below-average clustering solution, as it revealed more nuances in the low-performing mathematics group than were identified by previous analyses. Including only HTKS with the two mathematics measures revealed four distinct groups (see Appendix C, see Figure C3). Again, there appeared to be an average (n = 42) and high (n = 28) group, but now there appeared two distinct low-performing groups. Although both displayed poor Applied Problems performance, the first group also displayed a very poor sense of number magnitude (number line; n = 17), while the second group displayed very poor response inhibition (HTKS; n = 26).
At both Kindergarten and first grade, Mathematics cluster membership significantly differed according to the school students attended, but not by sex or age (see Table 2). However, and unlike reading, only a third of students who received special education services or who were observed with concerns were in the low-performing group in Kindergarten, and this number dropped to less than 20% in first grade (though no estimates were statistically significant).
Stability of cluster membership
Our second research aim was to assess the stability of clusters across the K-1 transition. A cross-tabulation is presented in Table 3, wherein the diagonal corresponds to stable group membership through first grade, or, students who were clustered into the same group during both study years. Contrary to our hypothesis, high performance was more stable during the first 2 years of schooling than was lower or average performance. One hundred percent of the high-performing Kindergarteners for reading remained high-performing in first grade, and nearly three quarters remained high-performing for math; none dropped into either low-performing group. In contrast, 63.5% and 55% of the average- and low-reading groups remained stable into first grade, respectively, compared with 58.7% and 32.1% of the average- and low-math groups. More students apparently remained stable or improved cluster membership than dropped in performance across the transition to first grade. This is especially true for those who were in the low-mathematics cluster in Kindergarten, as over two-thirds improved to the average- or high-performing groups.
Stability of Cluster Membership from Kindergarten to First Grade.
To understand which students changed cluster membership across the first 2 years of schooling, we computed difference scores between each first grade and Kindergarten subtest, which were then standardized for comparability. Briefly, students who dropped into a lower performing cluster (labeled “dropped”) exhibited significantly less growth than students who remained in the same cluster (“stable”) and vice versa for students who moved into a better-performing cluster (“improved”). Appendix D contains further description of which students exhibited stability and which changed cluster membership.
Analysis of Group Differences
Finally, we conducted two multivariate analyses of covariance (MANCOVA) to assess whether and how Kindergarten EF differentiated first-grade cluster membership while controlling for demographic characteristics (sex and school attended) and Kindergarten math and reading achievement (see Table 4). After correcting for multiple comparisons using the Holm–Bonferroni method, both Kindergarten digit span tasks predicted first-grade reading cluster membership (DS-F Effect Size [ES] = .15; DS-B ES = .07), while the HTKS and DS-B tasks predicted first-grade mathematics cluster membership (HTKS ES = .11; DS-B ES = .21). Examination of the ESs confirms our hypothesis that WM (DS-B) influenced both academic subjects but more strongly influenced mathematics. However, our hypothesis that attentional control (Pair Cancellation) and inhibitory control (HTKS) would influence both mathematical and reading ability was not supported (though inhibitory control did influence mathematics). Post hoc comparisons revealed that these effects were largely driven by the low-performing clusters, which were significantly different from both the average- and high-performing groups on all EF tasks. No Kindergarten EF task significantly differentiated first-grade average- and high- performing groups except for DS-B, which significantly differed across all first-grade mathematics clusters.
Using Kindergarten Executive Function (EF) Tasks to Predict First-Grade Achievement Cluster Membership.
Note. Marginal means displayed. Subscript denotes column proportions that differ significantly at the p < .05 level. Post hoc comparisons adjusted using the Sidak and Holm corrections. Analyses control for gender, school, and prior achievement. Reading: Wilks’ Lambda = .784, F(8, 196) = 3.177, p = .002, partial η2= .115. Math: Wilks’ Lambda = .719, F(8, 192) = 4.297, p = .000, partial η2 = .152. HTKS = Head-to-Toes, Knees-to-Shoulders task (Ponitz et al., 2008), DS-F = Digit-Span Forward, DS-B = Digit Span Backward (Weschler, 1991).
p < .05. **p < .01. ***p < .001.
Discussion
The present study examined performance classifications using cognitive and academic measures during the transition to early schooling. We sought to better understand how cognitive tasks can aid non-arbitrary classification of academic performance into low-, average-, and high-performance groups. We generally found that using cognitive measures helped to classify low-reading performance in both Kindergarten and first grade, but not mathematics performance. The high-performing groups exhibited more stability into first grade than the average- or low-performing groups. We also found that cognitive skills uniquely differentiated both concurrent and future academic classifications. Each of these findings is discussed further below.
Cluster Analysis of Both Cognitive and Academic Measures Sometimes Helps Differentiate Low Performance
Our primary research focus was to explore the efficacy of using cluster analysis to differentiate academic performance classifications in early schooling, specifically with respect to low performance and in combination with cognitive data. Cluster analysis presents a simple solution to the problem of arbitrary cutoff methodologies and may thus yield more meaningful groupings in academic data. Because many prior studies have investigated how EF was related to low, average, or high performance, we investigated whether including these measures would further differentiate the data into the three groups often arbitrarily assigned in other research. We conceptualized the inclusion of these EF measures as a form of validation, but found that EF measures only sometimes aided in distinguishing three distinct clusters. Many of our attempts at clustering these measures resulted in two-group solutions of “below” or “above” average performance. Thus, it may be that there are not three “natural” clusters of low/average/high during early elementary school, even though many researchers have constrained their data to behave in this way. This result has strong implications for educators tracking and observing students at these grade levels, particularly given the relative instability of these designations through first grade. Our results also imply that researchers who use cutoff methodologies may artificially impose group membership upon students, particularly for low-performing groups and among literacy measures that are not used in conjunction with cognitive data.
The Role of EF in Low, Average, and High Academic Classifications
Although continuous measurements typically reveal significant and positive relationships between achievement and EF (see correlations in Appendix B as an example), these methods do little to explain how these constructs are related. Most of the significant differences between our groups were driven by the low-performing clusters, aligning our results with previous literature linking EF deficits rather than strengths to achievement. However, in the three-group cluster solutions that we present here, we note that average-performing readers were (a) indistinguishable from low-performing readers without the addition of EF measures and (b) displayed EF skills more alike the high performers than the low performers. Moreover, including a measure of response inhibition (HTKS) with the first-grade mathematics subtests revealed further nuances among low performers, though not among the average- or high-performing groups. This differentiation illustrates the heterogeneity behind low achievement and may help explain why certain students perform poorly, thereby illuminating potential avenues for future intervention.
We expected that EF would play a large role in differentiating mathematics and reading cluster membership at each grade. This hypothesis was supported at both grades for the reading clusters, but not for the mathematics clusters. It is possible that EF (and, specifically, memory) contributes more globally to reading performance than to mathematics performance. Some prior research has only linked WM to advanced mathematics skills involving combinations of numbers and quantities (Purpura et al., 2017), a level that many of our sampled students might not yet have achieved. In contrast, there may be more domain-specific EF heterogeneity in low first-grade mathematics performance. This may be especially evident for response inhibition given the emergence of four clusters when including HTKS, two of which were specific to low performance. This is also consistent with prior research finding that response inhibition is closely linked to most components of mathematics skill (Purpura et al., 2017).
Kindergarten EF also significantly predicted first-grade cluster membership for both academic subjects. Consistent with our expectations, ESs for mathematics were larger than for reading. Although we also expected to see attention control and inhibitory control predict later performance in both subjects, this hypothesis was not supported. It may be that these relations were reduced to nonsignificance after accounting for WM, which is consistent with some recent investigations regarding the role of EF components on academic outcomes (e.g., Ahmed et al., 2019; Nguyen & Duncan, 2018). Alternatively, the tasks used in this study may differ from those used in prior research finding these significant relations, which call into question the replicability of these constructs.
Our findings also reinforce a growing body of literature underscoring the unique importance of WM for children’s academic skills. Specifically, we found that WM was significantly related to later reading and mathematics after controlling for early achievement and that it significantly differentiated low performance from average and high performance. This supports prior connections between WM deficits and struggling learners (e.g., Gathercole et al., 2016). An alternative explanation is that Kindergarten WM is predicting first-grade WM within the reading clusters, since digit span tasks were used in the first-grade reading cluster analysis. We find this less likely because Kindergarten WM only predicted differences between low and average reading in first grade, and not high reading. Moreover, EF tasks were not used in the creation of the mathematics clusters, and we see that WM significantly predicted all three groups. Our results invite future research to better examine the extent to which WM and other EF measures relate to non-arbitrarily-determined achievement classifications (ideally within larger, nationally representative samples). Because cluster analyses may allow a more nuanced and natural approach to group differences, researchers should especially consider this method to explore connections between deficits in EF and poor achievement.
High Performance is More Stable than Average or Low Performance
Relevant to research on and debate around educational tracking, we also investigated the stability of performance classifications across the K-1 transition. Given extensive documentation of low-performing, intervention-resistant students who later receive special education services, we expected the most stability among low performers. However, when exploring three-group solutions, results instead revealed that membership in the high-performance group was more stable than the average- or low-performing groups. Thus, it may be easier to objectively measure and classify young gifted students than young struggling students, particularly when cognitive and reading skills are both measured. Theoretically, it may be more common to measure “flukes” in low performance than in high performance without repeated testing (e.g., a student may just have a bad day during testing, and this could make them seem like a poorer performer but probably would not make them seem like a higher performer). Researchers and educators may want to revisit how low performers are differentiated from students with disabilities in early schooling given this instability.
Our results also suggest that teachers and special educators focus more on literacy problems than mathematics problems during early schooling, especially when observing students for potential special educational needs. It is therefore unsurprising that more low-reading students received special education services than low-mathematics students. In addition, and perhaps validating teachers’ focus on early literacy deficits, membership in the low-performing reading group was more stable through first grade than the low-performing mathematics group. It is possible that students received special education services because they had higher mathematics skills than reading skills, following the ability–achievement discrepancy model of learning disabilities. Future research investigating the consequences of this phenomenon is warranted, particularly among students eventually diagnosed with a mathematics disability.
Study Limitations and Conclusion
A significant limitation of this study was our inability to obtain individual demographic information, which we might have used to further externally validate these cluster solutions. Students attending schools with the least overall socioeconomic disadvantage (who were, presumably, the least disadvantaged themselves) displayed the best academic performance in both grades. However, it is unclear to what extent this trend is associated with a student’s school or their individual demographic background, and both are likely influential. We also caution against over-generalizing these results, given the relatively small and geographically limited nature of the present sample. We may not have had a large enough sample to capture a group of low-performing, treatment-resistant students, which would explain why high performance appeared more stable than low performance. A larger, nationally representative study would help establish whether our findings remain consistent for most schools and students.
Yet, we also offer counterpoints to these limitations that we believe validate our results. First, our sample is drawn from diverse racial and socioeconomic locations, which bolsters generalization to other diverse samples. Second, deficits in performance are most likely to be observed in high-risk, diverse schools, like those included in our sample. Third and finally, we argue that teachers may experience similar data issues, namely that there may also be a very small cohort of low-performing, intervention-resistant youth within their elementary schools. Future research methods should explore the accuracy and reliability of early classifications of these students despite their small sample size within grades and schools.
Despite these issues, few would question that existing classification and grouping schemes are arbitrary and that there is little evidence they are robust to alternative specifications. This study demonstrates the utility of combining cluster analysis with cognitive measures to examine early childhood performance classifications. Students without a formal Individualized Education Program (IEP) may still struggle under the typical curricula, so it is important to bolster the ability to better identify these students early in schooling. We hope that further research will continue to prioritize the use of available cognitive and academic performance data to better serve these learners.
Footnotes
Appendix A
Sample Demographic Information.
| Demographic | School 1 (n = 30) | School 2 (n = 40) | School 3 (n = 24) | School 4 (n = 25) | Total (N = 120) |
|---|---|---|---|---|---|
| Male (% of sample) | 63.3 | 55.0 | 54.2 | 48.0 | 57.7 |
| Percentage of school disadvantaged | 2.1 | 65.3 | 69.4 | 87.4 | 56.1 |
| White (%) | 45.7 | 40.5 | 46.9 | 14.1 | 36.8 |
| Black (%) | 4.2 | 47.9 | 44.5 | 82.9 | 44.9 |
| Hispanic (%) | 3.1 | 4.8 | 4.1 | 2.8 | 3.7 |
| Asian (%) | 45.3 | 2.3 | 0.7 | 0.0 | 12.1 |
| Age—Kindergarten: M (SD) | 5.71 (.31) | 5.82 (.36) | 5.74 (.44) | 5.94 (.40) | 5.80 (.38) |
| Age—First Grade: M (SD) | 6.80 (.30) | 6.81 (.30) | 6.83 (.37) | 6.95 (.38) | 6.84 (.34) |
| Observed for Special Ed (%) | 30.0 | 12.5 | 45.8 | 24.0 | 26.9 |
| Received Special Ed (%) | 16.7 | 5.0 | 8.3 | 12.0 | 10.8 |
Note. Data about schools were drawn from information made publicly available on statewide education reporting websites. Socioeconomic disadvantage was defined by this website as “those who have been determined to be eligible for free or reduced-price meals via locally gathered and approved family applications under the National School Lunch Program, are in households receiving food (SNAP) or cash (TANF) assistance, are homeless, are migrant, or are in foster care.”
Appendix B
Correlations Among Study Variables.
| Study Variables | (1) | (2) | (3) | (4) | (5) | (6) | (7) | (8) | (9) | (10) | (11) | (12) | (13) | (14) | (15) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| (1) K Passage Comp. | |||||||||||||||
| (2) First Passage Comp. | .613*** | ||||||||||||||
| (3) K Letter-Word ID | .756*** | .666*** | |||||||||||||
| (4) First Letter-Word ID | .618*** | .871*** | .759*** | ||||||||||||
| (5) K Applied Probs. | .517*** | .609*** | .630*** | .616*** | |||||||||||
| (6) First Applied Probs. | .516*** | .666*** | .583*** | .598*** | .668*** | ||||||||||
| (7) K Number Line | −.089 | −.096 | −.233* | −.104 | −.404*** | −.288** | |||||||||
| (8) First Number Line | −.203* | −.219* | −.349*** | −.222* | −.405*** | −.461*** | .294** | ||||||||
| (9) K Pair Cancel. | .240** | .187* | .298*** | .172 | .273** | .263** | −.279** | −.121 | |||||||
| (10) First Pair Cancel. | .282** | .313*** | .285** | .255** | .403*** | .374*** | −.017 | −.087 | .532*** | ||||||
| (11) K HTKS | .291** | .116 | .331*** | .111 | .440*** | .268** | −.151 | −.288** | .157 | .160 | |||||
| (12) First HTKS | .273** | .190* | .278** | .163 | .260** | .247** | −.094 | −.045 | .170 | .219* | .391*** | ||||
| (13) K Digit Span–F | .249** | .399*** | .348*** | .353*** | .438*** | .337*** | .021 | −.264** | .037 | .106 | .121 | .152 | |||
| (14) First Digit Span–F | .245** | .346*** | .317*** | .343*** | .444*** | .357*** | −.021 | −.237* | .058 | .098 | .129 | .148 | .681*** | ||
| (15) K Digit Span–B | .441*** | .406*** | .514*** | .406*** | .572*** | .508*** | −.454*** | −.480*** | .297*** | .194* | .342*** | .329*** | .333*** | .346*** | |
| (16) First Digit Span–B | .349*** | .424*** | .413*** | .366*** | .574*** | .467*** | −.155 | −.276** | .266** | .356*** | .238* | .265** | .385*** | .344*** | .487*** |
Note. K = Kindergarten; first = first grade; HTKS = Head-to-Toes, Knees-to-Shoulders task.
p < .05. **p < .01. ***p < .001.
Appendix C
Appendix D
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the National Science Foundation under grant number 1356118.
