Abstract
In recent years, several researchers have proposed that skilled adults may solve single-digit addition problems (e.g., 3 + 1 = 4, 4 + 3 = 7) using a fast counting procedure. Practicing a procedure often leads to transfer of learning and faster performance of unpracticed items. Such transfer has been demonstrated using a counting-based alphabet arithmetic task (e.g., B + 4 = C D E F) that indicated robust generalization of practice (i.e., response time [RT] gains) when untrained transfer problems at test had been implicitly practiced (e.g., practice B + 3, test B + 2 or B + 1). Here, we constructed analogous simple addition problems (practice 4 + 3, test 4 + 2 or 4 + 1). In each of three experiments (total n = 108), participants received six practice blocks followed by two test blocks of new problems to examine generalization effects. Practice of addition identity rule problems (i.e., 0 + N = N) showed complete transfer of RT gains made during practice to unpracticed items at test. In contrast, the addition ties (2 + 2, 3 + 3, etc.) presented large RT costs for unpracticed problems at test, but sped up substantially in the second test block. This pattern is consistent with item-specific strengthening of associative memory. The critical items were small non-tie additions (sum ≤ 10) for which the test problems would be implicitly practiced if counting was employed during practice. In all three experiments (and collectively), there was no evidence of generalization for these items in the first test block, but there was robust speed up when the items were repeated in the second test block. Thus, there was no evidence of the generalization of practice that would be expected if counting procedures mediated our participants’ performance on small non-tie addition problems.
Whether repeated practice of a procedural process will eventually lead to memorization of specific facts or lead to automaticity of the procedure is a core question in human learning (Singley & Anderson, 1989). Development of basic arithmetic skill is a classic case in point. For decades, the general consensus among researchers has been that a fundamental element of basic numeracy—solving single-digit addition problems with sums less than or equal to 10 (e.g., 3 + 2 = ? or 4 + 3 = ?)—is based on direct retrieval from associative memory in most educated adults (see Ashcraft & Guillaume, 2009; Zbrodoff & Logan, 2005 for reviews). In recent years, however, several researchers have argued that the retrieval theory is wrong. Instead, it is proposed that the counting process commonly used by children first learning addition develops into an automatic or “compacted” counting procedure in skilled adults (Barrouillet & Thevenot, 2013; Thevenot, Barrouillet, Castel, & Uittenhove, 2016; Uittenhove, Thevenot, & Barrouillet, 2016; see also Fayol & Thevenot, 2012). If this new view is correct, it would not only overturn the widely accepted direct retrieval theory of adults’ simple addition but could also have implications for mathematics pedagogy, including changes to national standards for mathematics education according to its proponents (Fayol & Thevenot, 2012, p. 401; Thevenot et al., 2016, p. 55). Here, we present experiments to test a strong prediction of the counting theory; namely, the transfer of training (i.e., generalization of learning from practiced to unpracticed problems) that would be expected with counting-based procedures. To foreshadow our conclusion, we found no evidence to support the counting theory but good evidence for the direct retrieval theory.
Evidence of fast counting for simple addition by skilled adults
Barrouillet and Thevenot (2013) analyzed 92 adults’ response times (RTs) for addition problems involving the numbers 1 through 4. They found that RT to answer these very small addition problems increased linearly and monotonically with the sum of the operands. The RT slope for problems comprised different operands (i.e., the non-ties such as 1 + 4 and 3 + 2) was 20 ms, about twice the 11 ms slope observed for the tie problems with a repeated operand (i.e., 1 + 1, 2 + 2, 3 + 3, 4 + 4). Barrouillet and Thevenot proposed that the very small addition non-ties activated a fast, automatic counting procedure that produces a linear problem-size effect in RT because each counting increment takes a uniform amount of time. The ties were assumed to be solved by direct fact retrieval from memory and therefore less sensitive to problem size.
Uittenhove et al. (2016) pursued the findings of Barrouillet and Thevenot (2013) but tested the 81 addition problems from 1 + 1 and 9 + 9 and collected strategy self reports for each problem (e.g., recalled the answer and counted). The purpose of the strategy reports was to identify participants who reported use of conscious procedures, such as intentional counting or reconstructive strategies, because these could contribute to a problem-size effect for the very small problems and thereby contaminate measurement of a problem-size effect possibly owed to an automated counting procedure. Uittenhove et al. identified a subset of 51 participants among the 90 tested who they identified as “frequent retrievers” for small problems. For these individuals, the problem’s sum was used to predict participants’ mean problem RT for four types of problems. Very small non-tie problems (both operands ≤ 4) had a mean slope of 46 ms per increment in the sum, for tie problems the slope was 8 ms, n + 1 with n > 4 problems had a slope of 7 ms, but there was no problem-size effect (−5 ms) for “medium small” non-tie problems with sums from 7 to 10 and at least one operand >4. Thus, Uittenhove et al. replicated the sum-related problem-size effect for very small addition problems reported by Barrouillet and Thevenot and provided evidence that this subset of problems is a distinct category of simple addition problems apparently solved by an automatic counting algorithm.
Generalization of addition practice
The problem-size effect for very small addition problems is potentially consistent with a counting model, but it does not rule out a fact retrieval model. For example, the network interference model of memory for single-digit addition and multiplication facts (Campbell, 1995) predicts a strong relationship between retrieval time and the sum of operands for simple addition. According to this view, a problem-size effect on RT arises in both simple addition and multiplication fact retrieval because interference among competing arithmetic facts increases with problem size owing to the compression of the psychophysical magnitude scale as number size increases (e.g., Dehaene, 1989; Dehaene, Dupoux, & Mehler, 1990). 1 Figure 2 in Uittenhove et al. (2016, p. 294) showed a very strong linear relationship (we estimated r2 = 0.93, slope = 40 ms) between mean RT and the sum for non-tie problems across the full range of sums from 3 to 17 as predicted by the network interference retrieval model of addition. Thus, a retrieval model provides good prediction of the Uittenhove et al. self-reported retrieval RT data. Consequently, other kinds of converging evidence are needed to support the fast procedure theory.
To this end, Campbell and Beech (2014) examined generalization of practice. Practicing a procedural process usually results in its speed up (Singley & Anderson, 1989) and this should generalize to different, unpracticed problems that utilize that procedure. Campbell and Beech reasoned that if simple addition was based on procedures then practicing a subset of problems (e.g., 4 + 3) should facilitate subsequent performance of similar, unpracticed problems (e.g., 3 + 2). The results showed that the procedure-based 0 + N = N problems presented clear evidence of generalization (i.e., practicing a subset of 0 + N problems facilitated a different subset of 0 + N problems), but there was no generalization of practice for non-zero simple addition problems. Even n + 1 problems, for which a counting-based strategy seems most likely, have not yielded evidence of generalization in educated adults (Campbell & Beech, 2014; Campbell, Dufour, & Chen, 2015; Campbell & Therriault, 2013; Chen & Campbell, 2014, 2016). In contrast, Baroody, Purpura, Eiland, and Reid (2015) found that training of n + 1 problems transferred to novel n + 1 items in young children not initially fluent with the n + 1 = successor of n relation. This is evidence that counting-based procedural strategies used for n + 1 problems can produce generalization, but this was not observed in adults’ performance.
Campbell, Chen, Allen, and Beech (2016) sought direct evidence that practice of counting-based procedures can produce generalization that leads to the speed up of unpracticed problems. They examined generalization in an alphabet arithmetic task (e.g., B + 1 = C, D + 3 = G). Alphabet arithmetic problems are solved, at least initially, by a counting-based procedure to increment through successive letters the required number of steps (e.g., B + 5 = C D E F G; see Zbrodoff & Logan, 2005). In two experiments, participants received six blocks of training on a subset of problems followed by two test blocks on untrained problems. Mean RT for problems increased linearly with the numerical value of the addend, as would be expected given a counting strategy. Both experiments showed robust generalization at test (i.e., faster RTs relative to counterbalanced control problems) when a test problem’s letter augend and answer letter sequence overlapped with practiced problems (e.g., practice B + 5 = C D E F G, test B + 3 = C D E). In such cases, the solution steps for the transfer problems are executed each time counting is used to solve the practiced problem. This implicit practice of the transfer problems facilitates their solution when those problems are explicitly tested later. In Experiment 2, however, transfer items with an unpracticed letter but whose answer was in a practiced letter sequence (e.g., practice C + 3 = D E F, test D + 2 = E F) also displayed generalization, albeit smaller than that observed with both an augend and sequence overlap. These generalization effects were stronger in the initial test block than in the second test block. This occurred because the transfer problems showed little or no RT gains in Block 2 from having been practiced in Block 1 whereas control problems sped up substantially from Block 1 to Block 2. Speed up on transfer problems owing to generalization would partially or completely exhaust opportunities for further RT gains from a repeated test, whereas control problems are, in effect, at an earlier learning stage with more potential to benefit from repeated testing.
The present experiments
The Campbell et al. (2016) alphabet arithmetic experiments demonstrated strong generalization when counting was surely involved; therefore, if a counting process is involved in solving simple addition facts, there should be observable transfer. The alphabet arithmetic studies indicated that generalization effects of counting practice was greater when practice and transfer problems shared both a common letter augend and answer sequence (e.g., practice B + 3 = C D E, test B + 1 or B + 2 for generalization effects relative to control problems). The previous addition generalization experiments (Campbell & Beech, 2014; Chen & Campbell, 2014, 2016) examined generalization between items within several simple addition problem categories (0 + N, ties, small non-ties with sum ≤ 10 and large non-ties), but did not analyze for effects of augend-sequence matches between practiced and transfer problems. Campbell et al. presented a re-analysis of the collection of previously published addition generalization experiments (combined n = 172) but found no evidence of facilitation when problems were preceded by problems with a matching augend and counting sequence. Nonetheless, the implementation of these experiments did not design the practice and transfer problems to have the nested relation shown to promote strong generalization in alphabet arithmetic. The present experiments were designed specifically for this purpose.
To ensure high precision of RT measurement, we tested three groups of 36 participants in separate experiments (total n = 108). All three experiments had six practice blocks followed by two test blocks of new problems to examine generalization effects. There were four problem categories: 0 + N, ties, small non-tie transfer and small non-tie control problems. The single-digit tie addition problems (e.g., 2 + 2 and 5 + 5) are widely agreed to be solved by item-specific fact retrieval (e.g., Barrouillet & Thevenot, 2013; Campbell & Gunter, 2002). Consequently, speed up observed for the practiced subset of ties was not expected to generalize to unpracticed tie problems at test, but the new tie problems were expected to speed up in Test Block 2 owing to increased memory strength from successful retrieval in Test Block 1. In contrast, 0 + N problems, which are governed by the addition identity rule 0 + N = N, were expected to demonstrate generalization of practice producing speed up for unpracticed 0 + N problems, as was observed previously (e.g., Campbell & Beech, 2014). Here, however, we attempted to manipulate this effect with a between-experiment manipulation, which was the only procedural difference between the experiments. The purpose was to assess the efficiency of transfer for this procedural skill.
Specifically, in Experiment 1, the 0 + N practice problems were presented only in Practice Block 6, whereas in Experiments 2 and 2r (a replication of Experiment 2), the 0 + N problems were included in all six practice blocks. In Experiment 1, we expected the zero problems in Block 1 of the test phase to be faster than the 0 + N problems in practice Block 6 because of generalization of practice, replicating previous research (e.g., Campbell & Beech, 2014). In contrast, in Experiments 2 and 2r, we expected to measure generalization after sufficient practice that speed to apply the 0 + N = N rule was nearing a lower asymptote. If generalization of practice has 100% efficiency, then, we would observe no difference in mean RT between 0 + N problems in Practice Block 6 and the new 0 + N problems in Test Block 1. If the transfer efficiency is less than 100%, however, then 0 + N problems in Test Block 1 will be solved slower than the 0 + N problems in Practice Block 6. A replication of Experiment 2 was planned and conducted because this test of transfer efficiency has not been investigated previously whereas the predicted generalization effect for 0 + N problems in Experiment 1 has been repeatedly observed (Campbell & Beech, 2014; Chen & Campbell, 2014, 2016).
The critical items with respect to the counting hypothesis were the augend-matched transfer problems and controls analogous to the augend-sequence overlap alphabet arithmetic problems used by Campbell et al. (2016). These were small, non-tie additions with sums ≤10 (including n + 1 problems), which are more likely candidates for fast counting-based procedures than larger simple additions. Participants practiced a subset of the small non-tie problems (e.g., 5 + 4, 4 + 3) for six blocks and then were tested on yoked additions that had the same augend but a smaller addend (e.g., 5 + 2, 4 + 1). These cases are directly analogous to the alphabet arithmetic stimuli with an augend-sequence match (e.g., practice B + 4, test B + 2) that demonstrated robust generalization of practice in the Campbell et al. studies. 2 If our participants solved the small non-tie additions using a counting-based procedure during practice, then in the test phase, we would expect faster RTs for the yoked augend-matched transfer problems relative to control problems. As these items were treated identically in the three experiments, the effective sample size for these tests was n = 108.
We did not include large non-tie addition problems in these experiments. When adults do not use direct retrieval for these items (roughly 50% of the time on average for North American adults), they use a wide variety of problem-specific procedural strategies (see, for example, LeFevre, Sadesky, & Bisanz, 1996) rather than a common procedure shared across multiple problems. Consequently, it would be very difficult to measure generalization effects of procedure use for the larger simple additions without knowing in advance the specific strategies individuals used for specific problems so that the practice and test items could be selected appropriately. This is why we did not include the large non-tie problems in the present experiments.
The specific fast counting model proposed by Barrouillet and Thevenot (2013) and Uittenhove et al. (2016) applied to the “very small” non-tie problems with both operands ≤4 (i.e., 1 + 2, 1 + 3, 1 + 4, 2 + 3, 2 + 4, 3 + 4 and their commuted counterparts); consequently, the present experiments, which used small problems with sums up to 10 did not test this specific model. Generalization effects seem like a natural prediction of the model but the “very small” problem set does not afford construction of practice, transfer and control stimuli with counterbalanced manipulation of augend and sequence overlap in a repeated-measures design. Nonetheless, the very restricted range for fast counting is an assumption of their specific theory that could be wrong while its assumptions about a fast counting mechanism is correct, but it applies more widely than they assumed. Indeed, these researchers have proposed that the fast counting procedure might be directly linked to other addition phenomena that are observed for a wider range of addition problems (Barrouillet & Thevenot, 2013, p. 41; Mathieu, Gourjon, Couderc, Thevenot, & Prado, 2016, p. 237). Thus, it is worthwhile to seek evidence of generalization from counting for small addition problems more inclusively. Furthermore, while it is practically impossible to design a study of nested relation transfer for the very small additions in a repeated-measure design, the present experiments did afford a transfer analysis for the very small problems in a between-participant design. One group received the very small problems in the nested transfer condition (e.g., practice 4 + 3, test 4 + 2, 4 + 1), and the other received the very small problems in the control condition.
Method
Participants
Three groups of 36 participants recruited at the University of Saskatchewan received either course credit or CAD$5. Multilingual participants answered the arithmetic problems in English. Experiment 1 tested 25 women and 11 men, among whom 34 were right-handed. Ages ranged 18-47 years (M = 25.8 years, standard deviation [SD] = 7.3). In total, 25 participants reported English as their first language, seven reported Chinese and one each reported Persian, Bengali, Portuguese and Spanish.
Experiment 2 tested 26 women and 10 men, 31 right-handed, with ages of 17-45 years (M = 22.9 years, SD = 5.8). In total, 30 participants reported English as their first language. The first language reported by the others included one Chinese, two Persian, one Gujarati and two Vietnamese. Finally, Experiment 2r, which was a replication of Experiment 2, included 24 women and 12 men, 33 were right-handed, with ages of 18-33 years (M = 21.5 years, SD = 4.1). In total, 29 participants reported English as their first language, with the remaining language reports including two Korean, one French, one Arabic, two Urdu and one Turkish.
Stimuli
There were four sets of practice phase problems. These included two ties-and-zeroes addition practice sets that each contained four tie problems and four zero problems. One set included the problems 0 + 3, 0 + 4, 0 + 5, 0 + 8 and 2 + 2, 6 + 6, 7 + 7, 9 + 9 while the other included 0 + 2, 0 + 6, 0 + 7, 0 + 8 and 3 + 3, 4 + 4, 5 + 5, 8 + 8. The 0 + N and ties were included because generalization of practice to new problems was expected for 0 + N problems but not for ties (e.g., Campbell & Beech, 2014); thus, these problem types provided a demonstration that the current experiment could replicate both generalization and no generalization.
The other two practice sets were the critical stimuli designed to measure possible augend-sequence generalization of addition practice. These comprised the four small non-tie operand pairs 32, 43, 53, 54 and 82, 63, 73, 64. During the practice phase, one of these sets was practiced as addition problems (e.g., 3 + 2) and the other set was practiced as multiplication problems (e.g., 8 × 2). The two small non-tie practice sets each had five yoked test phase problems that overlapped with respect to augend and answer sequence given a counting strategy. For the 32, 43, 53, 54 practice pairs, the yoked transfer test pairs were 3 + 1, 4 + 1, 5 + 1, 4 + 2, 5 + 2, and for 82, 63, 73, 64, the yoked transfer test pairs were 6 + 1, 7 + 1, 8 + 1, 6 + 2, 7 + 2. If during practice participants solved the small addition non-ties by counting, then we would expect generalization and faster RT for the yoked addition test problems (the transfer set) relative to the addition test set yoked to the problems practiced as multiplications (the control set). The multiplication practice ensured that RT differences between the small non-tie transfer and control sets in the test phase, should they be observed, could not be attributed to an encoding advantage owing to exposure to only the transfer set’s operands during practice (see also Dehaene, Spelke, Pinal, Stanescu, & Tsivkin, 1999). Assignment of the two ties-zeroes sets to the practice or test phase was counterbalanced with assignment of the two small non-tie sets to addition or multiplication practice over groups of four participants.
Design
All three experiments had six practice blocks followed by two test blocks of new problems to examine generalization effects. Problem order was randomized in each block independently for each participant. In Experiment 1, the four 0 + N problems (e.g., 0 + 3) were presented only in practice Block 6, while the four addition ties, four addition augend-match problems and four multiplication controls were practiced in all six practice blocks. In Experiments 2 and 2r (a replication of Experiment 2), all four problems types were presented in all six practice blocks (16 trials per block). The test blocks included the four new (i.e., unpracticed) 0 + N and tie problems and the five yoked addition transfer (augend-sequence match) problems and five control additions (18 trials per block). In both the practice and test phases, successive blocks were separated by a 30-s interval during which the participant read newspaper stories aloud. The purpose of the interpolated reading task was to interfere with participants rehearsing or remembering a practiced problem and its answer from the preceding block (i.e., retrieval from episodic memory).
Apparatus and procedure
Participants were tested individually in a quiet room with an experimenter present. Testing required about 20 min to complete. The stimuli were presented on two cathode ray tube (CRT) monitors controlled by E-prime 2.0 (Schneider, Eschman, & Zuccolotto, 2012), one viewed by the participant and the other by the experimenter. Participants sat approximately 50 cm from the monitor and spoke into a hand-held microphone. This detected the onset of the participants’ verbal response and triggered the stop signal to a software clock accurate to ±1 ms. Prior to the arithmetic task, participants performed a letter naming warm up task. The eight letters from “a” to “h” were presented one at a time in a random order. Participants named each letter as quickly as possible.
Arithmetic problems were displayed in black, Courier New 14-point font on a white background. The displayed problem occupied five character spaces with the augend and addend separated by the plus sign with adjacent spaces (e.g., 3 + 2). Each trial started with a 1-s central fixation dot that flashed off for 500 ms and then flashed on and off twice for 250 ms before the problem appeared with the operator sign at fixation. Response timing began when the problem appeared and was stopped by the participant’s spoken response. Accuracy rather than speed was emphasized in the training phase, but in the test phase, participants were instructed to respond both quickly and accurately. After the verbal response was provided, the experimenter entered the participant’s answer and marked spoiled RTs when the stop signal was not activated by response onset. The fixation dot then appeared for the next trial. Participants did not receive feedback regarding their speed or accuracy.
Results
A total of 648 RTs (4.8%) were marked as spoiled and excluded from analysis. Another 95 RTs (0.7%) less than 200 ms or greater than 2000 ms were discarded from the analysis. We analyzed median RT owing to the relatively small number of observations per cell (Miller, 1988). Greenhouse-Geisser corrected statistics were reported when Mauchly’s Test indicated violation of the sphericity assumption. 3 Along with null hypothesis significance tests, we also reported a Bayes Factor (BF) for each test, calculated using MorePower 6.0 (Campbell & Thompson, 2012b). This program implements the Bayesian Information Criterion (BIC), which approximates the unit-information prior as a default, objective Bayes prior probability (Masson, 2011; Wagenmakers, 2007). The BIC formulation used by MorePower 6.0 favors H0 for small effect sizes making it conservative with respect to Type I errors (Nathoo & Masson, 2016). In this application, BF is the odds ratio of the null (H0) over alternative hypothesis (H1). For example, BF equal to 10 for a given analysis of variance (ANOVA) test indicates that the data favor H0 over H1 by 10 to 1, whereas a value of 0.1 indicates a 10 to 1 ratio in favor of H1. BF is a continuous scale but a conventional interpretation is that a BF greater than 10 or less than 0.1 provides relatively strong evidence for H0 or H1, respectively, whereas BF between 3 and 0.33 provides little evidence one way or the other for H0 or H1 (e.g., Wetzels, van Ravenzwaaij, & Wagenmakers, 2015).
Practice phase
The mean rate of incorrect answers during the six practice blocks was 1.7% for small non-tie addition, 0.8% for the tie additions, <0.1% for the 0 + N problems, and 3.5% for the multiplication filler trials. Figure 1 presents the mean median RT for correct trials for small non-tie addition, addition ties and multiplication fillers combining the three experiments given that these items were handled identically in the three experiments. Figure 1 also includes mean median RT for the 0 + N problems across practice blocks for Experiments 2 and 2r in which the 0 + N problems were practiced in all six practice blocks. Multiplication was slowest overall (mean RT = 801 ms, n = 106), 4 followed by the small non-tie additions (722 ms, n = 108), tie additions (695 ms, n = 108) and 0 + N problems (581 ms, n = 72). All problem types sped up across practice blocks, with the greatest gains from Block 1 to Block 2 and gains diminishing thereafter with little RT gain after Block 4. Separate one-way ANOVAs for each problem type confirmed robust linear, quadratic and cubic components in the speed up function across blocks for each type (all p ≤ 0.007). Thus, the practice phase was effective at producing faster performance for all problem types with each problem type approaching a local RT asymptote. The following analyses investigated if the learning expressed in RT speed up in the practice phase was generalized to new addition ties, 0 + N problems or small non-tie problems in the test phase.

Mean median RT and standard errors by problem type and block in the practice phase.
Generalization analysis for ties and 0 + N problems
Previous research found that the addition ties (i.e., 2 + 2 to 9 + 9) did not show evidence of generalization of practice whereas the rule-based 0 + N = N problems did (Campbell & Beech, 2014; Chen & Campbell, 2014, 2016). In Experiment 1, we expected the new 0 + N problems in Block 1 of the test phase to be faster than the 0 + N problems practiced only once in Block 6 because of generalization of practice. In contrast, in Experiments 2 and 2r, if transfer efficiency approached 100%, then we expected no RT difference between 0 + N problems in practice Block 6 compared to the new 0 + N problems in test Block 1 because speed up of executing the procedural rule for these problems was approaching an asymptote (see Figure 1). Accordingly, we conducted a Group (Experiment 1 vs Experiments 2 and 2r) × Block (Practice Block 6, Test Blocks 1 and 2) × Problem Type (ties vs 0 + N) ANOVA of median RT for correct trials (the mean error rate for the tie and 0 + N problems across these blocks was 0.7%). The corresponding RT means for the 2 × 3 × 2 cells of the design appear in Figure 2. Error bars are 95% confidence intervals based on the mean squared error (MSE) for the three-way interaction (Campbell & Thompson, 2012a). All the ANOVA tests were significant (p ≤ 0.05), with the exception of the Group main effect (p = 0.239) and the Group × Type interaction (p = 0.06). Most important, however, was the predicted three-way interaction depicted in Figure 2 (F(1.87, 198.46) = 4.097, p = 0.02, MSE = 3662.993,

Mean median RT for ties and 0 + N problems in Practice Block 6 and Test Blocks 1 and 2.
Whereas there was no effect of the group factor in the addition ties mean RT, the analysis of the 0 + N problems indicated a robust Group × Block interaction (F(1.867, 197.908) = 14.918, p < 0.001, MSE = 2231.066,
Generalization analyses for small non-tie addition
For the small non-tie problems, the transfer problems in the test phase were related to the practice problems by virtue of a common augend and the transfer problems having a smaller addend (e.g., practice 5 + 3, test 5 + 1). To assess generalization, RT for the transfer problems was compared to the small non-tie control problems.
5
The treatment of small non-tie problems was identical across the three experiments, which afforded combining the three sets of results into a single Problem Type (transfer vs control) × Test Block (1 vs 2) analysis. Additionally, we report the results separately for each experiment to assess the replicability of the pattern observed. Figure 3 presents mean median RT for correct responses by problem type and test block separately for each experiment. The overall error rate for the small non-tie problems in the test phase was 2.1%. The RT ANOVA including all 108 participants indicated faster mean RT in Block 2 (653 ms) than Block 1 (703 ms, F(1, 107) = 41.007, p < 0.001, MSE = 6800.634,

Mean median RT in Experiments 1, 2, and 2r for small non-tie transfer and control problems in Blocks 1 and 2 of the test phase.
In separate analyses of the three experiments, Experiment 2 indicated a significant RT advantage for transfer problems relative to controls (24 ms, F(1, 35) = 5.164, p = 0.029, MSE = 3968.937,
As Figure 3 shows, there was no evidence for generalization in Test Block 1 in Experiment 1 (9 ms, t(35) = 0.435, p = 0.666, SE = 20.175,
Generalization analysis for very small non-tie additions
The very small addition category, ignoring operand order, includes only the six non-tie additions problems with both operands between 1 and 4. As noted previously, the small size of this proposed problem category prohibits a nested relation study of generalization using a repeated-measures design. These studies, however, afforded a between-participant analysis of generalization for the very small additions. The experiment used two sets of 10 small non-ties with assignment of set to transfer and control condition counterbalanced across participants. One set included five of the six very small non-ties: two used at practice (4 + 3, 3 + 2) and three at test (3 + 1, 4 + 1, 4 + 2). The other set contained no very small non-ties (8 + 2, 6 + 3, 7 + 3, 6 + 4 for practice and 6 + 1, 7 + 1, 8 + 1, 6 + 2, 7 + 2 for test).
If we assume a counting process for the very small problems, the participants tested on the very small problems in the transfer condition (n = 54) should show generalization at test, with the other participants tested on those same very small problems in the control condition (n = 54), providing a no-generalization control group. This comparison is vulnerable to random group differences. Therefore, we selected 6 + 1, 7 + 1, 7 + 2 from the other small (but not “very small”) problem set as control problems for which no specific carry-over effect from practice would be expected for either group. RT differences between groups for these items would reflect random group differences. The critical test therefore was the Group × Set interaction; If analysis confirmed an RT advantage on the very small problems (i.e., 3 + 1, 4 + 1, 4 + 2) for the transfer relative to the control group, and this group difference was smaller for the control problems (i.e., 6 + 1, 7 + 1, 7 + 2), this would confirm generalization for the very small problem set.
The Group × Set ANOVA of median RT for correct answers indicated a main effect of set (F(1, 106) = 4.044, p = 0.047, MSE = 11,417.672,
Seeking confirming evidence that the transfer group in this analysis tended to be faster overall than the control group, we examined tie problems. The two tie problem sets (2 + 2, 6 + 6, 7 + 7, 9 + 9 and 3 + 3, 4 + 4, 5 + 5, 8 + 8) were counterbalanced across these groups in the practice and test phases. A t-test of median RT for correct answers in Test Block 1 indicated that, despite random assignment of participants to groups, the transfer group was significantly faster to correctly answer the tie problems (M = 692 ms) than the control group (M = 763 ms, t(106) = 2.056, p = 0.042, SE = 34.327,
The question remains, ‘how powerful was the test of the Group × Set interaction in the preceding ANOVA?’ We can gauge potential effect sizes for generalization from the transfer group’s speed up during practice of 3 + 2 and 4 + 2, the two very small-problem practice items whose speed up might generalize to their yoked, very small nested problems at test (i.e., 3 + 1, 4 + 1 and 4 + 2). For 4 + 3 and 3 + 2, the transfer group sped up 152 ms from Practice Block 1 (M = 838 ms) to Practice Block 6 (M = 686 ms, t(53) = 5.975, p < 0.00001, SE = 25.536,
In summary, given the combined analyses showing strong evidence of no generalization in the analysis of very small problems (i.e., no Group × Set interaction), and adequate power to detect a medium effect size for this test (7% of the variance), it seems that these experiments do not provide support for generalization of practice for the very small addition problems.
Discussion
Transfer of practice for 0 + N identity rule problems and addition ties
With respect to the 0 + N identity rule problems, these experiments investigated generalization to new problems after a small amount of practice (Experiment 1) and after a considerable amount practice when RT for these problems was approaching a lower asymptotic limit (Experiments 2 and 2r). This allowed us to determine if a switch to new 0 + N problems at this point incurred some RT cost or instead transferred to new problems with complete efficiency. Experiment 1 replicated previous demonstrations that practicing a subset of 0 + N problems once leads to RT gains for unpracticed 0 + N problems (Campbell & Beech, 2014; Chen & Campbell, 2014, 2016). Experiments 2 and 2r showed that after six blocks (24 trials) of practicing 0 + N problems, the learning acquired transferred completely to new, untrained problems with no RT cost (Figure 2). Thus, the generalization of procedural learning to execute the addition identity rule was extremely efficient. This is strong evidence that these problems are solved by a common procedure and that generalization is a signature of procedure use (Singley & Anderson, 1989).
In contrast, for the tie problems (e.g., 3 + 3, 7 + 7), after six blocks of practice, switching to new tie problems in the test phase incurred considerable RT costs. This was expected because previous studies showed no generalization of practice for the addition ties (Campbell & Beech, 2014), presumably because they are answered using an item-specific fact retrieval process that strengths memory for the retrieved fact only. Nonetheless, mean RT for ties in Test Block 1 (727 ms) was considerably faster than mean RT for tie problems in Practice Block 1 (797 ms). If ties are solved by direct fact retrieval from associative memory, as is widely assumed, the 70 ms RT advantage for new tie problems at the beginning of the test phase, compared to the new tie problems at the beginning of practice (t(107) = 4.56, p < 0.001, SE = 15.33) must reflect general rather item-specific transfer of learning (Pauli, Bourne, & Birbaumer, 1998; Pirolli & Anderson, 1985). This might include, for example, speed up owing to participants learning to calibrate their maximum speed while minimizing errors. If we assume there was no item-specific transfer for ties, then speed up of non-item-specific processes accounted for approximately half of the 142 ms speed up for the tie problems across the six practice blocks.
Transfer of practice for small non-tie addition
The critical results from these experiments are related to the small non-tie additions. Generalization for these items is relevant to recent claims that skilled adults’ simple addition may be based on fast counting procedures (e.g., Barrouillet & Thevenot, 2013; Mathieu et al., 2016; Uittenhove et al., 2016), rather than on direct fact retrieval. The specific model proposed by Uittenhove et al. and Barrouillet and Thevenot’s model were restricted to the very small additions with both operands less than or equal to 4, whereas these experiments sought evidence that fast counting for addition might occur more widely among small additions problems with sums ≤10. Based on the alphabet arithmetic results (Campbell et al., 2016), if generalization occurs in genuine addition then we would expect the RT advantage for transfer problems relative to controls to be maximum when they are initially tested but diminished in a subsequent retest. This is because the generalization effects for the transfer problems would reduce potential RT gains from practice that the control problems retain. In none of the three experiments reported here, however, there was any evidence in Test Block 1 of RT gains for transfer problems relative to controls. In contrast, all three experiments found generalization for the 0 + N problems, demonstrating the sensitivity of the design to detect such an effect. There was some evidence, mainly in connection with Experiment 2, that transfer problems were slightly (26 ms) faster than control problems in Test Block 2. A generalization effect in Block 2, but not Block 1, is opposite to the pattern observed in the alphabet arithmetic studies (Campbell et al., 2016) and contrary to the prediction that generalization effects (i.e., the RT advantage for transfer relative to control problems) ought to diminish with repeated testing. Indeed, the 0 + N problems in Experiment 1 confirmed diminished RT gains in Block 2 after the generalization effect seen in Block 1. As such, the results are inconsistent with effects owed to generalization of counting practice. Apart from the possibility of a Type I error, however, we can think of no straightforward explanation of a delayed benefit to the transfer problems (i.e., appearing only in the second test block) following training with their yoked practice problems. As the effect emerged inexplicably only in Test Block 2 and in only one of the three experiments, it may be an anomalous result.
If there was to be a generalization effect in the first block of the test phase, how large an effect might we expect? During the six practice blocks, the small non-ties sped up 153 ms (841-688 ms). This probably underestimates the potential speed up, however, because participants were performing under accuracy instructions during the practice phase. Nonetheless, if we assume the efficiency of transfer with genuine counting would be similar to that observed by Campbell et al. (2016; Experiment 2) for the counting process used in alphabet arithmetic (about 30%), then counting would be expected to produce a generalization effect of approximately 50 ms in these experiments. In this case, for the entire sample of 108 participants, the F-test of the difference between transfer and control conditions in Block 1 had power of 0.96 to detect an effect of this magnitude or larger. Thus, collectively the experiments were virtually certain to have detected a substantial generalization effect if it existed. The null generalization results in the first test block imply that our large sample of educated adults did not solve the small non-tie problems with counting-based procedures. When these problems were tested a second time in Test Block 2, however, they showed robust RT gains, (Figure 3) as would be expected if performance was based on an item-specific retrieval process.
Apart from the correlation between problem RT and the sum for very small problems (Barrouillet & Thevenot, 2013; Uittenhove et al., 2016), is there other evidence for automated procedures for simple addition? Fayol and Thevenot (2012) found that a 150 preview of the operator (+ or ×) facilitated non-tie simple addition but not multiplication and took this as evidence of automated procedures for addition. Chen and Campbell (2015), however, subsequently tested a large sample of adults (n = 144) with a 150 ms operator preview and found robust operator priming for both addition and multiplication, questioning the generalizability of Fayol and Thevenot’s results and conclusion. Mathieu et al. (2016) showed right and leftward shifts in visuospatial attention for simple addition and subtraction, respectively, and suggested possible links to the automatic counting theory (p. 270). Nonetheless, they concluded that their “data do not directly speak to the question of whether single-digit addition and subtraction problems are solved by means of procedural or retrieval strategies” (p. 237). Thus, there is no clear converging evidence for automated procedures for addition.
Conclusion
Campbell et al. (2016) used alphabet arithmetic to demonstrate that counting-based procedures can produce substantial generalization of learning to unpracticed items that would be implicitly practiced during training (e.g., solving B + 3 by counting implicitly practices B + 1 and B + 2). Here, we designed analogous stimuli using small non-tie addition problems (e.g., practice 4 + 3, test 4 + 1 and 4 + 2). With a large sample of 108 participants, there was no evidence of generalization of practice for the small non-tie problems in Block 1. In contrast, the experiments demonstrated robust generalization of practice for 0 + N problems, confirming the potential sensitivity of the experiments to detect generalization. We conclude that the participants in our sample did not employ fast counting procedures for the small non-tie addition problems. Thus, there remains no evidence of the generalization of practice that would be expected if counting procedures mediated adults’ performance of small simple addition problems.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by a grant from the Natural Sciences and Engineering Research Council of Canada to Jamie Campbell.
