Abstract
Previous research has shown that naïve views of math and science concepts coexist with more formal views. The current study extended this finding to the domain of mathematical equivalence and tested whether inhibitory control relates to using more formal views over naïve ones. In the current study, we report two experiments in which undergraduate students (n = 125 for Study 1 and n = 184 for Study 2) completed a priming task involving inhibitory control and math items, an inhibitory control flanker measure, and a comprehensive mathematical equivalence assessment. We found quantitative and qualitative evidence that adults hold both naïve operational views and formal relational views of equivalence across multiple measures and under timed and untimed conditions. In contrast to our hypotheses, we did not find evidence to support a strong association between individual differences in inhibitory control and mathematical equivalence knowledge. The results call into question the role of this domain-general cognitive skill in contributing to adults’ expression of naïve operational thinking.
When comparing children and adults, it can often appear that there are stark shifts in their thinking about the world. For example, there are many ideas (e.g., the world is flat, numbers only refer to discrete integers) that change over time and eventually resemble adult-like views that are more scientifically accurate (e.g., Carey, 2009; Wellman & Gelman, 1992). However, research on conceptual change in math and science suggests that these naïve, child-like conceptions can exist alongside the more formal, adult-like conceptions (e.g., Goldberg, Thompson-Schill, 2009; Kelemen & Rosset, 2009). Furthermore, specific cognitive skills, such as inhibitory control, are theorised to help deal with the interference of these child-like theories in adulthood (e.g., Zaitchik et al., 2016). Here, we extend these ideas to adults’ understanding of mathematical equivalence to determine whether child-like conceptions are still apparent in adulthood and whether inhibitory control is related to the propensity to think in these child-like ways.
Conceptual change
The literature on conceptual change is complex and (perhaps ironically) has updated in a variety of ways since its beginning in the 1960s (see Vosniadou et al., 2008). But at the heart of this work is a focus on how people update or revise their knowledge of the world, often in fairly radical ways and often in response to conflicting information (e.g., Carey & Spelke, 1994; Chi, 2008; Vosniadou & Brewer, 1992). Documenting these developmental changes has resulted in theoretical and educational advances, with one of the more surprising conclusions being that the naïve or child-like ways of thinking do not disappear. That is, these naïve views do not get replaced or overwritten with adult-like theories, but instead can exist alongside them (Shtulman & Valcarcel, 2012). For example, in the area of mathematics, misconceptions that appear in childhood can persist into adulthood (e.g., DeWolf & Vosniadou, 2015; Vamvakoussi et al., 2013), and these naïve childhood misconceptions can interfere with adults’ accurate knowledge of fractions, algebra, and geometry (e.g., Stricker et al., 2021). In the current study, we examine whether adults hold naïve views of a particular math topic—mathematical equivalence.
Knowledge of mathematical equivalence
Mathematical equivalence is the idea that both sides of the equal sign represent the same amount (Kieran, 1981). It is a foundational concept in arithmetic and algebra (e.g., Carpenter et al., 2003; Linchevski, 1995), and equivalence knowledge is related to math achievement (Fyfe et al., 2020; Hornburg et al., 2022; Matthews & Fuchs, 2020). Prior research has documented a trajectory of knowledge change (Rittle-Johnson et al., 2011, see Figure 1). One major shift that occurs is from operational to relational thinking. For example, a learner with a rigid operational view is only successful with problems in a traditional format (e.g., a + b = c) and defines the equal sign in a way that includes computations (i.e., “the total”). In contrast, a learner with a comparative relational view can solve problems in a variety of formats and recognizes relational definitions of the equal sign (e.g., “the same amount”) as optimal (Rittle-Johnson et al., 2011). Within this domain, the rigid operational view is considered a naïve child-like view, and the comparative relational view is considered a more sophisticated adult-like view.

Trajectory of mathematical equivalence knowledge.
Children in many countries struggle with mathematical equivalence and their errors often reflect rigid operational thinking (e.g., Simsek et al., 2021). For example, many children define the equal sign as “the total” or “answer” rather than a symbol relating two equivalent amounts (e.g., McNeil & Alibali, 2005a), and they often solve equivalence problems incorrectly by adding all the numbers (e.g., seeing 3 + 7 = 4 + __ and writing 14 in the blank; Fyfe & Rittle-Johnson, 2017; McNeil & Alibali, 2005b; Rittle-Johnson, 2006). These errors are not only prevalent, but children are confident that they are correct (Grenell et al., 2022) and they are often resistant to change (McNeil & Alibali, 2005b). The leading theory suggests strong adherence to these naïve operational views stems from children’s vast prior experiences with traditional arithmetic (McNeil & Alibali, 2005b). These experiences tend to be ubiquitous and extremely narrow, with students working almost exclusively with problems in a standard “operations = answer” format (a + b = c; Powell, 2012). Children tend to extract common patterns from these experiences (e.g., equal sign means add all the numbers) and then overgeneralize them. Because the strategies work so well in arithmetic, they become entrenched and resistant to change.
Although most research has been conducted with children, a few studies have examined adults’ understanding of mathematical equivalence to better understand the degree of entrenchment. Only three studies have assessed adults’ explicit definitions of the equal sign (Chesney et al., 2013; Fyfe et al., 2020; McNeil & Alibali, 2005a), and approximately 10%–18% of adults provided operational views (e.g., “the answer to the problem,” “the sum”). This rate was even higher (up to 45%) when asked if the equal sign had multiple meanings. Furthermore, these same studies indicate that adults sometimes solve mathematical equivalence problems incorrectly and that their errors are due to the persistence of operational thinking. For example, Chesney and colleagues (2013) found that 62% of the errors on timed mathematical equivalence problems reflected arithmetic-based strategies such as adding all the numbers. Thus, adults “did not merely make calculation errors but, rather, solved equations using the incorrect strategies typically used by children” (Chesney et al., 2013, p. 1084). Also, two studies using priming paradigms showed that adults’ performance declines when primed to think operationally (e.g., primed with operational words like “total”; McNeil & Alibali, 2005a; McNeil et al., 2010). These operational views may also have consequences; undergraduate students who produced an operational definition of the equal sign performed worse on algebra problems than students who provided a relational definition (62% vs. 79% accuracy; Fyfe et al., 2020).
From these few studies, it appears some adults still hold operational views of equivalence that reflect similar misconceptions as children. However, we still do not have a detailed understanding of individual differences in adults’ mathematical equivalence knowledge and what cognitive factors might predict variability in the expression of operational thinking.
Inhibitory control and mathematics achievement
Based on conceptual change theories (e.g., Vosniadou, 2014), one factor that may help individuals suppress this naïve operational thinking is inhibitory control, which refers to the ability to inhibit prepotent or automatic responses (e.g., Diamond, 2013). Inhibitory control may be important for conceptual change in mathematics more broadly to help learners inhibit old naïve knowledge to use more accurate knowledge (e.g., Zaitchik et al., 2016). Indeed, individual differences in inhibitory control sometimes relate to math achievement in children (see Allan et al., 2014) and in adults (e.g., Cragg & Gilmore, 2014). For example, a study with undergraduates found that more inhibitory control predicted better math performance after controlling for other skills, such as working memory and cognitive flexibility (Coulanges et al., 2021).
However, other studies have not found associations between inhibitory control and math achievement (e.g., K. Lee & Bull, 2016; K. Lee & Lee, 2019; Van Dooren & Inglis, 2015), suggesting that the math topic may matter. As Vosniadou et al. (2018) note, inhibitory control may “be particularly recruited in tasks in which the use of a scientific or mathematical concept contradicts an initial concept which must be rejected” (Vosniadou et al., 2018, p. 62). There is some evidence to support this idea as inhibitory control is particularly related to math topics with entrenched strategies (e.g., Laski & Dulaney, 2015). For example, in decimals and fractions, it helps learners inhibit their well-built up knowledge of whole numbers (e.g., knowing that 3/4 is bigger than 3/7 even though 4 is less than 7; Avgerinou & Tolmie, 2020; Gómez et al., 2015; Leib et al., 2023; Roell et al., 2017). Similarly, on certain arithmetic problems (e.g., a + b – c), it helps learners inhibit the standard “left-to-right” approach and use an order of operations that is most efficient (e.g., Eaves et al., 2022; Robinson & Dubé, 2013).
We speculate that inhibitory control may be similarly involved within the context of mathematical equivalence. Adults may simultaneously hold operational and relational views in mind, and expressing a relational view of the equal sign may require adults to inhibit their early-formed and well-practiced operational views of the equal sign. To the best of our knowledge, the association between inhibitory control and mathematical equivalence knowledge has only been examined in children (Devlin et al., 2023; Young & Shtulman, 2020), but not in adults. These studies reported modest correlations (rs = .25–.40) between math equivalence knowledge and inhibitory control in children ranging from age 5 to 12, but in both cases, the associations were no longer significant after controlling for other cognitive variables. It is an open question how the relation between math equivalence and inhibitory control operates in adults.
Testing the role of inhibitory control
A key question is how to test whether inhibitory control is related to adults’ understanding of math equivalence. That is, what evidence would suggest that inhibitory control helps adults suppress their operational thinking to express their relational thinking? One approach is to calculate the correlation between the two constructs to see if inhibitory control predicts differences in math equivalence knowledge. If inhibitory control is involved, we would expect a positive association, such that adults who have better inhibitory control are more likely to exhibit relational thinking and have higher scores on a math equivalence assessment.
A more direct experimental approach is to manipulate whether adults engage their inhibitory control and test if it influences their performance on equivalence tasks. It is well established that conflict tasks can engage inhibitory control. For example, in the classic Stroop task, adults are faster and more accurate to name the font color on non-conflict trials (e.g., RED written in red ink) than on conflict trials (e.g., BLUE written in red ink) because conflict trials require participants to inhibit reading the word (MacLeod, 1991). Critically, engaging inhibitory control on these conflict trials can influence performance on subsequent trials (referred to as the congruency sequence effect or the conflict adaptation effect; Braem et al., 2019; Egner, 2007). Participants often perform better on conflict trials within the Stroop task if they are preceded by a conflict trial (Duthoo et al., 2014). The idea is that engaging their inhibitory control to deal with the conflict on the initial trial allows them to deal with conflict more easily on the next trial.
Several studies suggest that this carry-over effect applies across tasks as well (e.g., Borst et al., 2012; Yu et al., 2020). For example, Linzarini and colleagues (2015) used a priming paradigm in which children completed a Stroop task (a conflict trial or a non-conflict trial) followed by a number conservation task (a conflict trial or a non-conflict trial). Consistent with the conflict adaptation effect, they found that children were faster on the conservation trials if they were preceded by a conflict Stroop trial, but only when it was a conflict conservation item. That is, engaging inhibitory control on a conflict Stroop item helped children perform better on the number conservation task when it required inhibitory control. A similar paradigm can be used in the context of math equivalence. If correctly solving a math equivalence item requires inhibitory control, then performance should be enhanced when the equivalence item is preceded by a conflict Stroop trial (i.e., when participants’ inhibitory control has been activated).
The current study
In the current study, we had two main aims which we tested in two experiments. The first aim was to characterize adults’ knowledge of math equivalence to examine the coexistence of child-like and adult-like views. We did so by assessing performance under speeded and non-speeded conditions, including both procedural (e.g., problem-solving) and conceptual (e.g., knowledge of equal sign) items, and comparing performance on equivalence problems and matched arithmetic problems that varied only in the location of the equal sign. We formed two hypotheses. First, under speeded conditions, we predicted that adults’ accuracy on equivalence problems would not be at mastery and would be poorer than their accuracy on matched arithmetic problems. Second, under speeded and non-speeded conditions, we predicted that adults’ errors on a broad set of math equivalence items would reflect operational thinking.
The second aim was to examine associations between inhibitory control and math equivalence knowledge, which we addressed in two ways. First, we used a correlational approach. We assessed inhibitory control using a stand-alone flanker task to see if individual differences on this task were associated with knowledge on a math equivalence measure. The flanker task requires participants to identify a central target while inhibiting distractors that are on each side. We used this task because it is a well-established measure that captures variability in inhibitory control across participants, and also because it provided a complementary measure relative to the Stroop task, which we employed in our experimental approach. We expected a positive association between inhibitory control scores and math equivalence performance.
Second, we used an experimental approach to test whether activating inhibitory control influenced adults’ responses on math equivalence items with an inter-task priming paradigm (e.g., Linzarini et al., 2015). Participants completed Stroop items as the prime followed by math items as the probe. The Stroop items were either conflict trials (e.g., BLUE written in red ink) or non-conflict trials (e.g., RED written in red ink), and the math items were either conflict trials (e.g., math equivalence items on which the default “add all the numbers” strategy produces an incorrect response) or non-conflict trials (e.g., traditional arithmetic items on which the default “add all the numbers” strategy produces the correct response). We reasoned that if math equivalence items require inhibitory control, then performance on equivalence items (but not arithmetic items) should be enhanced when they are preceded by a conflict Stroop trial.
Study 1 Method
Participants
Study 1 included 125 adult participants ranging in age from 17 to 32 years (M = 18.65 years, SD = 1.50) who completed a single online session lasting approximately 20–25 min. We aimed for a sample size of 100 based on previous studies using a similar paradigm. The stopping rule was based on the combination of this target sample size, the pool of available participants, and our timeline for being able to collect the data during one semester. Our final sample size of 125 aligned with this goal and exceeded the sample sizes in previous studies using a similar priming paradigm (e.g., 40 participants in Linzarini et al., 2015; 81 participants in Yu et al., 2020). A sensitivity analysis was carried out to determine the effect size that could be detected with a sample of 125 for the repeated measures analyses of variance (ANOVAs) for the priming task. This analysis revealed we had 80% power to detect a small effect size (f = 0.10, alpha = .05).
Participants were recruited from a university in a midwestern city, and they received course credit for participating in the study. Of those that reported, 73% self-identified as female, 26% as male, 0.8% as transgender, and 0.8% as genderqueer. Most participants were Caucasian (70%), freshman in college (68%), and Psychology majors (75%), but see the Supplemental File for detailed demographic information.
Materials and measures
All tasks and measures were programmed and presented using Psychopy v.2021.1.4 (Peirce et al., 2019). After agreeing to participate, students were able to complete the three tasks independently online by following instructions on the screen and typing their responses. The three tasks included a priming task, a flanker task, and an equivalence assessment.
Priming task
We used a priming paradigm that included two different tasks that both theoretically involved conflict and relied on inhibitory control. This paradigm tests for the presence of the conflict adaptation effect, which shows that engaging inhibitory control on an initial conflict trial can help participants deal with the conflict on subsequent trials (e.g., Braem et al., 2019; Egner, 2007), including conflict trials on a different task (e.g., Borst et al., 2012; Yu et al., 2020). For the priming task, participants were presented with either a congruent or incongruent Stroop item and then either an equivalence or an arithmetic math item. The incongruent Stroop items represented the conflict trials on the first task, and the equivalence items represented the conflict trials on the subsequent task. The Stroop items served as the primes and the math items served as the probes. Each prime-probe set is a single trial, and each participant completed 32 test trials (2 blocks of 16 trials each).
For the Stroop items, we created congruent and incongruent items, and each block of 16 trials contained 8 trials of each type. Each congruent item contained one of four color words (RED, BLUE, YELLOW, and GREEN) printed in its corresponding ink color (e.g., RED printed in red). Each incongruent item contained one of four color words printed in a different ink color (e.g., RED printed in blue). For each item, the word appeared in the middle of a black screen for 5,000 ms or until the participant responded (consistent with Linzarini et al., 2015). Participants had to indicate the color of the ink by pressing the r (red), b (blue), y (yellow), or g (green) keys on the keyboard.
For the math items, we created 16 equivalence items and 16 matched arithmetic items. Each block of 16 trials contained 8 trials of each type. Each item included four non-repeating addends between two and nine. To create matched items, we used the same four addends in the same order for an equivalence and arithmetic item but changed the location of the equal sign (see Table S4 in the Supplemental File). For each equivalence item, the equal sign was presented after three addends so that there were operations on both sides of the equal sign (8 + 9 + 7 = 5 + __). For each matched arithmetic item, the equal sign was presented after all four addends so that it was in the standard “operations = answer” format (8 + 9 + 7 + 5 = __). Each item was presented in white ink in the middle of a black screen for 1,500 ms. This duration time was modelled after McNeil et al. (2010). After the item went away, a box appeared, and participants had 10,000 ms to type their numerical response.
The task began with eight practice trials with corrective feedback (e.g., “Correct” or “Oops! That was wrong”) on both the Stroop and math items and then 32 test trials split into 2 blocks of 16 trials. In each block, there were four trials of each of the following types: Congruent Stroop then Arithmetic, Congruent Stroop then Equivalence, Incongruent Stroop then Arithmetic, and Incongruent Stroop then Equivalence. Within a block, the order of items was randomised with the contingency that participants did not get more than two of the same trial type in a row. Example trials are presented in Figure 2. Accuracy and response time data were collected for both the primes (congruent or incongruent Stroop items) and the probes (equivalence or arithmetic items). Trials that were more than three standard deviations from each participant’s mean reaction time for each task were excluded. In line with Chesney et al. (2013), we coded students’ errors on the equivalence items. Errors were categorised as arithmetic-based (i.e., adding all four addends and reporting the sum, or adding only the three addends before the equal sign and reporting the sum), other (e.g., idiosyncratic number), or blank. A second researcher independently coded 20% of the data and reliability was high (kappa = .98).

Priming task paradigm.
Flanker task
For the flanker task, participants were presented with either a congruent or incongruent trial. Each congruent trial contained five arrows pointing in the same direction (e.g., <<<<<). Each incongruent trial contained five arrows with the middle arrow pointing in a different direction (e.g., << > <<). For each item, the arrows appeared in the middle of a black screen for 5,000 ms or until the participant responded as was done in Zelazo et al. (2014). Participants had to indicate the direction of the middle arrow by pressing the right or left arrow keys on the keyboard. The task began with four practice trials with corrective feedback (e.g., “Correct” or “Oops! That was wrong”) and then 50 test trials split into 2 blocks of 25 trials each. In each block, there were 9 congruent and 16 incongruent trials presented in a randomised order.
Accuracy and response time data were collected for each item. We used the formulas reported in Zelazo et al. (2014) to create a total Flanker score that took into account both accuracy (across test trials) and reaction times (across correct incongruent trials). Prior to these scores being calculated, trials with reaction times that were less than 100 ms or greater than 3 standard deviations from each participant’s mean reaction time were excluded. This total score was out of a maximum of 10 points and was the sum of participants’ accuracy score (0–5 points possible) and reaction time score (0–5 points possible).
Mathematical equivalence assessment
Participants completed a math equivalence assessment using 11 items from a validated and comprehensive measure that has been used with adults (Fyfe et al., 2020; Matthews et al., 2012; Rittle-Johnson et al., 2011). The assessment included five conceptual items: two that assessed understanding of equation structures and three that assessed knowledge of the equal sign. It also included six procedural items that required open-ended problem-solving.
Each item was scored dichotomously (1 or 0 points). For procedural items, exact numerical solutions were scored as correct (allowing for rounding up or down when decimals were included). For conceptual items, responses were scored as correct if they were coded as relational, which meant the response relied on the relation between values on either side of the equal sign as opposed to computing answers (see Tables 1 and 2 for examples). The first and second authors independently scored all the data and inter-rater agreement was good (equation structure item 1 kappa = .85, equation structure item 2 kappa = .74, equal sign definition: relational or not kappa = .89, equal sign definition: operational or not kappa = .82, equal sign definition: vague or not kappa = .73). All discrepancies were discussed and resolved.
Items and performance on the mathematical equivalence assessment.
The order of the items in Table 1 (from top to bottom) is the order in which the participants solved the 11 math equivalence items.
Sample definitions of the equal sign provided by college students.
Procedure
Participants completed all tasks independently online. They completed the priming task first, followed by the flanker task, and then the math equivalence assessment. At the end of the session, participants were asked some optional questions about their identities (e.g., gender, ethnicity) and academic experiences (e.g., major, year in college).
Study 1 Results
We first report on the math equivalence assessment to characterise adults’ knowledge of this topic. Then, we report on the flanker task to assess the correlational association between inhibitory control and math equivalence knowledge. Finally, we report the priming task to provide experimental evidence on the association between inhibitory control and math equivalence knowledge. Table 3 includes the descriptive statistics for all study variables. The sample sizes varied somewhat across tasks because we used a pairwise deletion approach to handle missing data. Participants were included in each analysis if they had usable data for the tasks that were relevant for that analysis.
Descriptive statistics for study variables for Study 1 and Study 2.
RT = response time measured in seconds.
Mathematical equivalence assessment
On the equivalence assessment with no time constraints, most adults provided usable data, but some were excluded for providing disingenuous responses (e.g., typing “ggg” instead of a numerical response or typing “5” for their answer on every item) or for having missing data on more than half the items. For the six procedural items, 118 adults (out of 125) provided usable data. For the five conceptual items, 122 adults (out of 125) provided usable data.
On the procedural items, the average percent correct was 69% (SD = 28%). Performance varied as a function of item, with accuracies ranging from 45% to 96% (Table 1). On the conceptual items, the average percent of relational responses was 70% (SD = 23%). Item-level scores ranged from 24% to 88% (Table 1). The most difficult item required adults to spontaneously explain an equation in a relational way (e.g., “2 × 3 is 6 and if you were to multiply it by 4 it would be 6 × 4 so that equals 6 × 4 on the other side”), and only 24% did so spontaneously. To be clear, this does not mean 76% of the adults solved this problem incorrectly. Many adults provided valid explanations, but they relied on operational thinking (e.g., 2 × 3 × 4 = 24 and 6 × 4 = 24). When prompted to avoid computations on the second conceptual item, relational responses increased to 75%.
Performance on the untimed assessment was far from ceiling, and adults’ incorrect responses highlighted the presence of operational views. Here we provide qualitative descriptions of three items. First, consider the second procedural item: m + m + m = m + 12. Notice that the surface-level properties are similar to the first procedural item (n + n + n + 2 = 17), yet performance was markedly diminished on the second item because it is no longer in the traditional “operations = answer” format. Out of 116 attempts on this second procedural item, 67% provided the correct response, but of all the incorrect responses, a full 71% of them indicated an operational view either because the participant assumed the equal sign was at the end (e.g., 4 m = 12 so m = 3) or because the participant only focused on the left side of the equal sign (e.g., 3 m = 12 so m = 4).
Second, consider adults’ definitions of the equal sign on the third conceptual item. Table 2 provides examples and shows that some participants provided multiple views (e.g., “it means sum or that the numbers are equivalent”). Considering all the views, 90 out of 120 participants (75%) provided a relational definition, 19 out of 120 (16%) provided an operational definition, and 34 out of 120 (28%) provided vague responses. We also considered adults’ best interpretations—such that each adult received a single definition type with relational given priority, then vague, then operational. For these best interpretations, 75% were relational, 18% were vague, and 7% were operational. That means approximately 1 out of every 14 adults defined the equal sign exclusively in terms of operations or answers. Finally, consider adults’ ratings of equal sign definitions on the fourth conceptual item. Although 88% of participants endorsed the relational definition as good (“the same as”), 61% also endorsed the operational definition as good (“the answer to the problem”). These qualitative data support the notion that many adults hold operational perspectives that exist alongside a relational perspective.
Flanker task
We used the flanker task to capture individual differences in adults’ inhibitory control and to assess correlational associations with math equivalence knowledge. On the flanker task, 118 (out of 125) adults provided usable data. The remaining seven participants were excluded for providing disingenuous responses (n = 1) or for having missing data on more than half the test trials (n = 6). We first confirmed that the flanker task produced the classic congruency effect. Participants were in fact slower to respond on incongruent trials (M = 0.58 s, SD = 0.14 s) than on congruent trials (M = 0.51 s, SD = 0.14 s), t(117) = –5.89, p < .001, suggesting that the flanker task was appropriately engaging participants’ inhibitory control.
We then examined correlations, and we found that total flanker scores were significantly and positively correlated with their math equivalence conceptual scores, r(114) = .21, p = .02. A complementary Bayesian correlation was computed, and the Bayes factor (BF10 = 2.90) indicates only weak or “anecdotal” evidence in favor of a reliable positive correlation (see M. D. Lee & Wagenmakers, 2014 for interpretation suggestions). Total flanker scores were not significantly correlated with their math equivalence procedural scores, r(112) = .12, p = .20, and the Bayes factor (BF10 = 0.48) indicates “anecdotal” evidence in favor of a null correlation.
Priming task
Recall that the priming task included prime-probe pairs in which the prime contained either a congruent or incongruent Stroop trial and the probe contained either an arithmetic or equivalence item. For the priming task, 121 participants had usable data, which was defined as providing genuine responses and completing at least half of the prime-probe test trials. As shown in Table 3, the Stroop items functioned as expected: adults’ accuracy was near ceiling and they were faster to respond on congruent trials than on incongruent trials. Our primary focus was adults’ performance on the math items and whether it varied by trial type.
Accuracy
We first examined accuracy on the math items across all 3,872 trials (121 participants × 32 trials completed by each participant) using a 2 (Stroop Type: Congruent vs. Incongruent) × 2 (Math Type: Arithmetic vs. Equivalence) repeated measures ANOVA. There was no main effect of Stroop Type, F(1, 120) = 3.80, p = .05, ηp2 = .03, or a Stroop Type × Math Type interaction, F(1, 120) = 0.02, p = .89, ηp2 = .00. However, there was a large main effect of math type, F(1, 120) = 76.43, p < .001, ηp2 = .39. Under these speeded conditions, participants were less accurate on equivalence items (M = 49%, SE = 3%) than on matched arithmetic items (M = 77%, SE = 2%). We also explored their errors. Across all 1,936 equivalence trials (121 participants × 16 equivalence trials for each participant), adults used the operational add-all error 22% of the time. The remaining trials were solved correctly (47% of trials), left blank (6% of trials), or included a different error (6% were responses within one of the correct answers and 20% were idiosyncratic). If we only consider the trials on which adults provided a response (not blank), then the add-all error accounted for 46% of incorrect responses, which reflects the strategy typically used by children.
Response time
To answer our primary question about the association between inhibitory control and equivalence knowledge, we tested whether response times on the math items depended on whether they were preceded by an incongruent or congruent Stroop task. We reasoned that if math equivalence items require inhibitory control, then response times on equivalence items (but not arithmetic items) should be faster when they are preceded by an incongruent Stroop trial. We conducted a 2 (Stroop Type: Congruent vs. Incongruent) × 2 (Math Type: Arithmetic vs. Equivalence) repeated measures ANOVA with response time on the math items as the outcome measure. As is typical when examining response times, we focused only on the 2,307 trials that were solved correctly (i.e., participants answered the prime and probe correctly). 1 There was a main effect of Math Type, F(1, 85) = 64.86, p < .001, ηp2 = .43, as participants took longer to solve equivalence items (M = 5.08, SE = 0.14) than arithmetic items (M = 4.42, SE = 0.14). There was not a main effect of Stroop Type, F(1, 85) = 0.56, p = .46, ηp2 = .01. As expected, there was a significant Stroop Type × Math Type interaction, F(1, 85) = 5.37, p = .02, ηp2 = .06 (Figure 3).

Interaction between math type and congruency type for reaction time on math probes.
Follow-up comparisons revealed a main effect of Stroop Type for the arithmetic items, F(1,85) = 5.77, p = .02, ηp2 = .06. Participants took longer to solve an arithmetic item if preceded by an incongruent Stroop prime (M = 4.55, SE = .15) than a congruent Stroop prime (M = 4.29, SE = 0.14). However, there was not a main effect of Stroop Type for the equivalence items, F(1, 85) = 1.09, p = .30, ηp2 = .01. Although the difference was in the expected direction—participants took less time to solve an equivalence item if preceded by an incongruent Stroop prime (M = 5.02, SE = 0.15) than a congruent Stroop prime (M = 5.15, SE = 0.16)—it was not significant.
We conducted a complementary Bayesian analysis. The Bayes factor for the inclusion of the interaction effect (BF10 = 2.54) indicated anecdotal evidence in favor of including it in the model as a valuable predictor of participants’ response times.
Study 1 Discussion
In Study 1, we characterized adults’ math equivalence knowledge and examined associations with inhibitory control in a sample of U.S. college students. Our hypotheses about their math equivalence knowledge were supported. Under speeded conditions, adults’ accuracy solving equivalence problems was quite poor (49% correct), it was lower than their accuracy on matched arithmetic problems, and many of their errors were operational. Even under non-speeded conditions, performance on the math equivalence assessment was far below mastery and adults exhibited both relational and operational thinking. For example, 16% generated an operational definition of the equal sign and 61% endorsed an operational definition as good.
The results concerning associations with inhibitory control were much less robust. Individual differences in inhibitory control on the flanker task were positively correlated with their conceptual scores on the math equivalence assessment. However, the correlation was fairly weak. Furthermore, the correlation between inhibitory control and procedural scores on the math equivalence assessment was not significant. Also, we did not find evidence that activating inhibitory control on the Stroop task significantly improved performance on the equivalence items. Together, these results do not support a strong association between inhibitory control and equivalence performance. However, because these results were unexpected and because there were some descriptive hints in the expected direction (e.g., response times on equivalence items were shorter following an incongruent Stroop trial), we conducted a replication study.
Study 2 Method
Study 2 was nearly identical to Study 1. College students completed the same math equivalence assessment, a flanker task, and a priming task. However, in Study 2, we modified the priming task so that the math trials more closely resembled the Stroop trials (e.g., forced choice as opposed to free response), and so that the participants had to complete a larger number of trials. We made these changes to enhance the chances of observing the standard conflict adaptation effect, in which the effect of inhibitory control transfers across trials, as research suggests that the effect is larger to the extent the tasks are similar (Braem et al., 2019).
Participants
Study 2 included 184 adult participants ranging in age from 17 to 43 years (M = 19.18 years, SD = 2.07) who completed a single online session lasting approximately 60 min. Participants were recruited from the same university as Study 1, and they received course credit for participation. Of those that reported, 68.5% self-identified as female, 28% as male, 3% as non-binary, and 1% as genderqueer. Most participants were Caucasian (65%), freshman in college (54%), and Psychology majors (17%), but see the Supplemental File for detailed information.
Materials and measures
All the materials, measures, and procedures were the same as Study 1 with the exception of the priming task, which was modified to include more trials and to include a forced-choice paradigm on the math probe items. In Study 2, participants were asked to complete 160 prime-probe test trials. The congruent and incongruent Stroop items were identical to those used in Study 1. For the math items in Study 2, we created 80 equivalence items and 80 matched arithmetic items (see Table S4 in the Supplemental File). Each item was presented in white ink in the middle of a black screen for 5,000 ms or until the participant responded. Below each problem, there were two answer choices: one choice was the solution to the arithmetic version of the problem and the other choice was the solution to the equivalence version of the problem (counterbalanced for which choice appeared on the left versus right side of the screen). An example equivalence item is 4 + 8 + 6 = 3 + ___ with 15 and 21 as the 2 answer choices. This means for the equivalence items, one answer choice represented the correct solution that made both sides equivalent and the other answer choice represented the operational add-all error (i.e., adding up all the numbers in the problem). Participants had to indicate their answer choice (left-hand side or right-hand side) by pressing the left and right arrow keys on the keyboard.
As in Study 1, the task began with practice trials with corrective feedback on both the Stroop and math items, and then participants completed the 160 test trials. The test trials consisted of 40 trials of each of the 4 trial types (Congruent Stroop then Arithmetic, Congruent Stroop then Equivalence, Incongruent Stroop then Arithmetic, and Incongruent Stroop then Equivalence). The order in which items were presented was randomized with the contingency that participants did not get more than two of the same trial type in a row.
Study 2 Results
As in Study 1, we first report on the math equivalence assessment to characterize adults’ knowledge of this topic, then we report on the flanker task and the priming task. Table 3 includes the descriptive statistics for all study variables. The sample sizes varied somewhat across tasks because we used a pairwise deletion approach to handle missing data.
Mathematical equivalence assessment
On the math equivalence assessment, the results from Study 1 were replicated in Study 2. See Table 1 for the similarities in item-level performance on the procedural and conceptual items.
A qualitative examination of errors revealed similar patterns as Study 1. For example, on the second procedural item (m + m + m = m + 12), of all the incorrect responses, a full 63% of them indicated an operational view (compared with 71% in Study 1). Most of the operational errors occurred because the participant assumed the equal sign was at the end (e.g., 4 m = 12 so m = 3). Similarly, when defining the equal sign, 69% provided a relational definition, 21% provided an operational definition, and 24% provided vague responses. When considering their best interpretations, 69% were relational, 20% were vague, and 11% were operational (compared with 7% operational in Study 1). That means, in this sample, 1 out of every 9 adults defined the equal sign exclusively in terms of operations or answers. Finally, when rating equal sign definitions, 85% endorsed the relational definition as good, but 64% also endorsed the operational definition as good (compared with 61% in Study 1).
Flanker task
On the flanker task, 181 (out of 184) participants provided usable data. The remaining three participants were excluded for providing disingenuous responses or for having missing data on more than half the test trials. Like in Study 1, participants were in fact slower to respond on incongruent trials (M = 0.69 s, SD = 0.30 s) than on congruent trials (M = 0.62 s, SD = 0.29 s), t(180) = 5.37, p < .001. Similar to Study 1, total flanker scores were not significantly correlated with math equivalence procedural scores, r(165) = .06, p = .49, and the Bayes factor (BF10 = 0.18) indicates “moderate” evidence in favor of a null correlation. But unlike Study 1, flanker scores were not significantly correlated with math equivalence conceptual scores, r(175) = .12, p = .12, and the Bayes factor (BF10 = 5.73) indicates “anecdotal” evidence in favor of a null correlation.
Priming task
On the priming task, 181 participants provided usable data. As shown in Table 3, the Stroop items functioned as expected: adults’ accuracy was near ceiling, and they were faster to respond on congruent trials than on incongruent trials. As in Study 1, our primary focus was adults’ performance on the math items and whether it varied by trial type.
Accuracy
We first examined accuracy on the math items across all 28,960 trials (181 participants × 160 trials completed by each participant) using a 2 (Stroop Type: Congruent vs. Incongruent) × 2 (Math Type: Arithmetic vs. Equivalence) repeated measures ANOVA. There was no main effect of Stroop Type, F(1,180) = 0.48, p = .49, ηp2 = .003, or a Stroop Type × Math Type interaction, F(1,180) = 0.03, p = .85, ηp2 = .00. Consistent with Study 1, the only significant effect was a large main effect of Math Type, F(1, 180) = 36.70, p < .001, ηp2 = .17. Under speeded conditions and using a forced-choice paradigm, participants were less accurate at selecting the correct answer on equivalence items (M = 73%, SE = 2%) than on matched arithmetic items (M = 86%, SE = 1%).
Response Time
We then tested whether response times on the math items depended on whether they were preceded by an incongruent or congruent Stroop task. We conducted a 2 (Stroop Type: Congruent vs. Incongruent) × 2 (Math Type: Arithmetic vs. Equivalence) repeated measures ANOVA with response time on the math items as the outcome measure. As in Study 1, we focused only on the trials that were solved correctly (i.e., participants answered the prime and probe correctly) resulting in 22,501 trials. Consistent with Study 1, there was a main effect of Math Type, F(1,167) = 58.15, p < .001, ηp2 = .26, as participants took longer to select the correct answer on equivalence items (M = 2.16, SE = .05) than on arithmetic items (M = 1.98, SE = .04). There was not a main effect of Stroop Type, F(1, 167) = 1.38, p = .24, ηp2 = .01, and inconsistent with Study 1, there was not a significant Stroop Type x Math Type interaction, F(1,167) = 0.02, p = .89, ηp2 = .00 (Figure 4). We conducted a complementary Bayesian analysis. The Bayes factor for the inclusion of the interaction effect (BF10 = 0.08) indicated strong evidence in favour of a null interaction effect and for excluding it from the model as a valuable predictor of participants’ response times.

Interaction between math type and congruency type for reaction time on math probes.
Study 2 Discussion
Consistent with Study 1, we found that adults’ accuracy on equivalence items was lower than their accuracy on matched arithmetic items, even when we used a forced-choice paradigm instead of a production task. We also replicated the results on the untimed math equivalence assessment and showed that performance was not at mastery and that adults showed evidence of both relational and operational thinking.
With regard to inhibitory control, some of the specific analytic results differed across Study 1 and Study 2, but the general trend was the same and did not support our hypotheses: we did not find evidence of a strong association between inhibitory control and math equivalence knowledge. In Study 2, inhibitory control on the flanker task was not significantly correlated with either their procedural or conceptual scores on the math equivalence assessment. Moreover, on the priming task, we did not find evidence that activating inhibitory control on the Stroop trials influenced performance on either the arithmetic items or the math equivalence items.
General discussion
The current study examined the persistence of child-like views of mathematical equivalence into adulthood and whether inhibitory control relates to these views. In support of our hypotheses and in line with theories of conceptual change, in two independent studies, we found quantitative and qualitative evidence that adults hold both operational views and relational views of equivalence problems. The coexistence of these views was found across multiple measures and under speeded and non-speeded conditions. In contrast to our hypotheses, we found little evidence to suggest that inhibitory control relates to adults’ math equivalence knowledge. These results call into question the role of this domain-general cognitive mechanism in adults’ expression of naïve child-like views within the domain of math equivalence.
Contributions to understanding mathematical equivalence
Our findings add to the few studies on adults’ understanding of math equivalence by providing evidence that some adults hold both operational and relational views of the equal sign. This was evident in their accuracy rates and in the types of errors they made. Notably, the current studies made three unique contributions to this literature by (1) examining these child-like views in speeded and non-speeded conditions, (2) using items that tap procedural and conceptual knowledge, and (3) directly comparing equivalence items to matched arithmetic items.
Only one previous study included both speeded and non-speeded conditions within the same group of participants (Chesney et al., 2013). Consistent with that study, our results showed that accuracy was generally lower in the speeded conditions but still not at mastery in the non-speeded conditions. In addition, our error analysis indicated that operational views were still present in adults even when they had unlimited time to solve the equivalence problems. For example, only 67%–70% of college students solved an equation in a non-standard format correctly (m + m + m = m + 12) because many converted it to a standard “operations = answer” format. If we had found that operational thinking was only present under speeded conditions, it would be difficult to rule out the influence of the time pressure on the expression of these child-like views.
We also extended previous work by using a more comprehensive measure of equivalence that included both procedural and conceptual items. Procedural knowledge refers to “action sequences for solving problems” and conceptual knowledge refers to “understanding of principles that govern a domain” (Rittle-Johnson & Alibali, 1999, p. 175). Both types of knowledge are important for learning mathematics (Rittle-Johnson & Alibali, 1999), and we found evidence of operational thinking in both cases. For example, on the equivalence items in the Study 1 priming task, 46% of the errors reflected the operational “add all the numbers” that is often reported in children. Given that these types of procedural items involve calculations, it may be harder for adults to resist activating their operational views when solving these types of problems. However, operational views were also present on conceptual items. For example, 16%–21% of adults defined the equal sign operationally (e.g., “the sum,” “the answer to the problem”). Similarly, even when explicitly told not to perform calculations (e.g., “Without adding 67 + 86, can you tell if this is true: 67 + 86 = 68 + 85. How do you know?”), only 59%–75% of adults provided a relational response and the rest of the sample tended to rely on calculations. This tendency aligns with the “compulsion to calculate” phenomenon, which stems from “students’ attachment to the belief that problems are solved by direct calculation” due to their prior experiences with arithmetic (Stacey & McGregor, 1999, p. 151).
Finally, the current studies are the first to directly compare adults’ performance on matched arithmetic and equivalence problems. Furthermore, we can make causal claims about how item type influences performance given that (1) the priming task was a within-subjects task, (2) the only difference between the matched problems was the location of the equal sign, and (3) participants saw the problems in a randomized order. Across two independent samples, we found that adults had lower accuracy rates and were slower to respond on equivalence items than on arithmetic items—and this was true both when adults had to generate a numerical solution and when they had to select a numerical solution from two options. These results suggest the location of the equal sign causes differences in adults’ performance in ways that are consistent with the change-resistance theory (McNeil, 2014).
Contributions to understanding the role of inhibitory control
Although operational thinking was evident, relational thinking still tended to be the norm in these adults, and there was clear variation across individuals. What accounts for the tendency for some adults to express these child-like operational views on equivalence items? The current study examined whether individual differences in inhibitory control may be one factor. However, the results were somewhat mixed across our two studies and generally demonstrated weak to non-existent associations between mathematical equivalence and inhibitory control.
Based on our correlational approach using the flanker task, three of the four correlations were not significant and the fourth was relatively weak. Specifically, in Study 1, scores on the flanker task were significantly and positively correlated with adults’ conceptual equivalence scores (r = .21), but not with their procedural equivalence scores. And in Study 2, flanker scores were not correlated with either equivalence score. The lack of a strong correlational link is somewhat consistent with two published studies with children. In a cross-sectional study with 5- to 12-year-old children, inhibitory control (as measured by a flanker task) was modestly correlated with children’s accuracy across four mathematical equivalence problems (r = .40; Young & Shtulman, 2020). In a longitudinal study, inhibitory control (as measured by a Stroop task) at age 6 was weakly correlated with children’s math equivalence knowledge 2 years later (r = .25; Devlin et al., 2023). However, in both studies, the link between inhibitory control and math equivalence was no longer significant after accounting for other cognitive variables, raising questions as to the direct role of inhibitory control in this area.
We also employed a more rigorous experimental approach to examine the causal effects of inhibitory control by manipulating whether inhibitory control was activated prior to solving arithmetic or equivalence items. In Study 1, the math items included a production task (i.e., generating a numerical response), and in Study 2, the math items included a forced-choice task (i.e., selecting one of two numerical responses). In neither case did we find evidence for a causal connection between inhibitory control and equivalence performance. That is, performance on the equivalence items did not depend on whether the preceding Stroop trial was an incongruent task (i.e., relying on inhibitory control) or a congruent task. In Study 1, we found a significant interaction indicating that activating inhibitory control influences arithmetic and equivalence differently. However, it was the arithmetic performance that was influenced, not the equivalence items. For arithmetic items, there appeared to be a carry-over congruency effect from the primes where adults were slower on the arithmetic items after incongruent Stroop trials than congruent Stroop trials. However, we did not find significant evidence that activating inhibitory control on the primes facilitated performance on the subsequent equivalence probe items in either study, even in Study 2 when we used a much larger number of trials and used a forced-choice paradigm for both the Stroop primes and the math item probes.
One possible explanation of these findings is that inhibitory control plays a much smaller role in adults’ expression of equivalence knowledge than we expected. Given conceptual change theory, we predicted that expressing a relational view of the equal sign would require adults to engage their inhibitory control to suppress their well-formed operational views of the equal sign. Based on that premise, we expected performance on equivalence items would be enhanced if we had participants activate their inhibitory control beforehand by completing incongruent Stroop items. Similar theoretical ideas have been proposed in prior work; for example, Devlin et al. (2023) wrote, “We hypothesized that children who enter formal schooling with higher executive functioning skills will be less likely to become entrenched in operational patterns as they inhibit prepotent strategy responses and will, thus, develop a stronger formal understanding of mathematical equivalence” (p. 1429). However, when we tested these well-grounded theoretical ideas empirically, we did not find evidence of a link between inhibitory control and math equivalence in adults. This highlights the importance of testing theoretical ideas with relevant data because an intuitive link between constructs is not always supported by empirical evidence.
So one possibility is that inhibitory control is not related to math equivalence knowledge. A similar possibility is that this type of domain-general cognitive skill may play a role in adults’ ability to express relational views of equivalence over operational ones, but a much smaller and more indirect role relative to domain-specific skills. For example, it is possible that domain-specific skills—such as how much exposure adults have to traditional arithmetic or how they mentally organize their arithmetic facts—might be more important for their ability to express their math equivalence knowledge than inhibitory control. For example, adults with more exposure to traditional arithmetic may have a more entrenched operational view of equivalence and be more likely to express an operational view instead of a relational one (e.g.,McNeil, 2014). It is also possible that adults who tend to organize arithmetic facts into equivalent groups, such as putting 8 + 6 together and 10 + 4 together because they both add up to 14, may be more likely to express a relational view of the equal sign (Chesney et al., 2014). Future research that measures both domain-general and domain-specific skills would be able to better tease apart the role that these different skills play for explaining individual differences in math equivalence.
Of course, it can be difficult to interpret null effects and another possible explanation is that inhibitory control does influence math equivalence knowledge, but we were unable to detect it given the current study design. In fact, a review of the literature has revealed that the conflict adaptation effect can vary in magnitude depending on certain design choices, and in particular, more similarity between tasks is often associated with larger effects. Review articles have also suggested that cross-task interference can be hard to detect (see Braem et al., 2019). Although possible, we believe this explanation may be less likely because we employed two different priming task paradigms and found similar results (production task and forced-choice task), we had large enough sample sizes to be able to detect small effects, and the measures functioned appropriately and in other ways as expected.
Limitations and future directions
Several limitations of the current study can motivate future research to replicate and extend the findings. For instance, we only included one independent, stand-alone measure of inhibitory control (i.e., the flanker task). Future research should determine the role of different types of inhibitory control tasks that vary on the type of stimuli used (e.g., numerical versus non-numerical) and what participants are asked to inhibit (e.g., automatic associations, distracting information) to better understand the role that inhibitory control plays. For example, future research should include measures of both interference control (e.g., flanker task) and response inhibition (e.g., go/no-go task) to better understand how different types of inhibitory processes may relate to individuals’ ability to express their mathematical equivalence knowledge. The current study also did not include other executive function measures (e.g., working memory and cognitive flexibility) or control measures (e.g., IQ) that are related to inhibitory control and mathematics achievement (e.g., Cragg & Gilmore, 2014; Peng et al., 2016). Future research should include these additional measures to determine whether individual differences in mathematical equivalence knowledge are better explained by these other variables.
The current study also only examined the relation between inhibitory control and mathematical equivalence at a global level, but it is possible that the relation may be moderated by specific item characteristics or specific person characteristics. Also, similar to other studies on adults’ equivalence knowledge (e.g., Chesney et al., 2013; Fyfe et al., 2020; McNeil et al., 2010), the current study only included samples of college students. Studying a larger, more representative population will provide more precise insight into the persistence of child-like views of mathematical equivalence in adults with a broader range of educational and life experiences, and may also afford more fine-grained analyses focused on individual differences.
Additional limitations of the current study relate to the nature of the priming task. Although we modified the priming task in Study 2, future studies could include different variants of this task to better understand the mixed findings from the current study and to better understand how inhibitory control may or may not be recruited when solving different types of math problems. For example, a verification task used in previous studies could be used where participants have to indicate whether a solved arithmetic or equivalence problem is true or false (e.g., Goldberg & Thompson-Schill, 2009; Shtulman & Valcarcel, 2012; Stricker et al., 2021). One possibility with our forced-choice paradigm used in Study 2 is that some adults may have realized that they should pick the smaller solution when the equal sign was in the middle of the problem (i.e., an equivalence problem) and the bigger solution when the equal sign was at the end of the problem (i.e., an arithmetic problem). Therefore, using a verification task format where sometimes the equivalence problems are solved correctly and sometimes solved incorrectly using the add-all solution may allow us to remove this confound in a future study.
Conclusion
The current study adds to our understanding of how an operational, child-like view of equivalence still exists in adulthood. It also provides the first empirical test of a relation between inhibitory control and mathematical equivalence knowledge in adults. To our surprise, the results across two studies were largely consistent in demonstrating weak to non-existent links between inhibitory control and math equivalence. These findings are a testament to the need for empirical tests of well-grounded theoretical ideas, and they suggest a need for future research to focus on different potential mechanisms of conceptual change in this area. The broader results on adults’ equivalence knowledge also have practical implications for the education system and the need to revisit the concept of equivalence throughout students’ education.
Supplemental Material
sj-docx-1-qjp-10.1177_17470218241280941 – Supplemental material for Role of inhibitory control in adults’ mathematical equivalence knowledge
Supplemental material, sj-docx-1-qjp-10.1177_17470218241280941 for Role of inhibitory control in adults’ mathematical equivalence knowledge by Amanda Grenell and Emily R. Fyfe in Quarterly Journal of Experimental Psychology
Footnotes
Acknowledgements
We thank all the individuals who participated. We also thank undergraduate research assistants, Kendall Vance and Joules Emerson.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Grenell was supported by a training grant from the Eunice Kennedy Shriver National Institute of Child Health and Human Development of the National Institutes of Health under Award Number T32HD007475. The content is solely the responsibility of the authors and does not represent the official views of the National Institutes of Health.
Data availability statement
Data will be made available upon request.
Supplementary material
The Supplementary Material is available at: qjep.sagepub.com
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
