Abstract
Being able to process multiword sequences is central for both language comprehension and production. Numerous studies support this claim, but less is known about the way multiword sequences are acquired, and more specifically how associations between their constituents are established over time. Here we adapted the Hebb naming task into a Hebb lexical decision task to study the dynamics of multiword sequence extraction. Participants had to read letter strings presented on a computer screen and were required to classify them as words or pseudowords. Unknown to the participants, a triplet of words or pseudowords systematically appeared in the same order and random words or pseudowords were inserted between two repetitions of the triplet. We found that response times (RTs) for the unpredictable first position in the triplet decreased over repetitions (i.e., indicating the presence of a repetition effect) but more slowly and with a different dynamic compared with items appearing at the predictable second and third positions in the repeated triplet (i.e., showing a slightly different predictability effect). Implicit and explicit learning also varied as a function of the nature of the triplet (i.e., unrelated words, pseudowords, semantically related words, or idioms). Overall, these results provide new empirical evidence about the dynamics of multiword sequence extraction, and more generally about the role of statistical learning in language acquisition.
Introduction
Humans are constantly exposed to and produce an unlimited number of novel utterances and this generative ability has long been considered as a hallmark of human language. For decades, generative linguists have argued that this phenomenon is explained by an innate system of abstract grammatical rules known as the “universal grammar hypothesis” (Chomsky, 1957). Distinct cognitive abilities supported by different neural systems may allow people to generate complex utterances (Ullman et al., 2005). For example, a mental lexicon including simple linguistic forms (e.g., individual words, morphemes) combined with a mental grammar including combinatorial rules would enable the formation of an infinite number of sentences (Pinker, 1991; Pinker & Ullman, 2002).
More recently, usage-based approaches to language have provided an alternative view to account for the mechanisms involved in language acquisition (Croft, 2001; Goldberg, 2006; Tomasello, 2003). According to this view, language gradually emerges through the interaction between general cognitive mechanisms and the repeated exposure to concrete items (Ibbotson, 2013). Learners are thought to store incoming utterances and to generate knowledge about the properties of these utterances (e.g., grammatical categories, semantics) by generalising over these stored multiword sequences (Abbot-Smith & Tomasello, 2006).
Over the last two decades, this approach has received multiple computational implementations to illustrate this learning and generalisation process. For instance, Solan et al. (2005) developed an algorithm (ADIOS for automatic distillation of structure) capable of generalising over different kinds of sentences from a given corpus using the statistical information present in the same data. In the same vein, Borensztajn et al. (2009) used an automatic data-oriented parsing procedure to identify the most likely multiword sequences used in child speech and model the evolution of their abstractness over time. Similarly, Meylan et al. (2017) developed a Bayesian statistical model to study the contribution of language productivity and abstractness to children’s linguistic knowledge by focusing on their early capacity to use the determiners “a” and “the” along with a noun.
Although these computational modelling studies have successfully captured multiword learning process, the emergence of grammatical knowledge and different developmental patterns more broadly, their reliance on mathematical algorithms and comprehensive corpus analysis undermines their psychological plausibility, as they lack realistic learning mechanisms and memory constraints inherent to the real-time nature of language processing (Christiansen & Chater, 2016). Chunk-based models, on the contrary, rely on a simple but a powerful mechanism (i.e., associative learning) that can account for both memory constraints and language processing, ranging from single word segmentation (Perruchet & Vinter, 1998) to multiword sequence acquisition (Jones & Rowland, 2017). For instance, McCauley and Christiansen (2019) developed a computational model of language perception and production that assumes language acquisition takes place in an incremental manner, through local shallow processes based on chunking and statistical learning mechanisms. Processing occurs on a word-by-word basis by assembling words into chunks (i.e., sequences of words), rather than via a full syntactic analysis as assumed by generativist theories. Given that language perception and production are thought to be interwoven processes in this model, both are assumed to rely on the same chunks and distributional statistics learnt during language acquisition. Thereby, this model relies on a chunk-by-chunk process instead of whole-sentence optimisation. Note that McCauley and Christiansen’s model is the first usage-based model having used a large number of natural language corpora (i.e., 79 single-child corpora for perception and 200 for production evaluation, representing a total of 29 languages).
In line with the model by McCauley and Christiansen (2019), numerous studies suggest that language users are sensitive to distributional properties at different levels of the linguistic input, and that statistical learning plays a key role in language acquisition (Aslin, 2018; Conway et al., 2010; Saffran et al., 1996). For instance, word frequency is known to affect word recognition (Grainger, 1990) and speech production (Jescheniak & Levelt, 1994). There is also evidence that linguistic processing is not only affected by word frequency but also by multiword frequency (Ambridge et al., 2015; Carrol & Conklin, 2020). In these studies, a multiword sequence is often defined as a number of consecutive words stored and retrieved from memory as a whole (Wray, 2002), acting as a single unit and resulting in a processing advantage (e.g., “How are you doing?”). It is worth noting, however, that it has also been suggested that this processing advantage could arise from either the simultaneous access to the component parts of a sequence, or from the priming of multiple combinations via the base components, rather than from storing the sequence as a whole (Wray, 2012, p. 234).
Many developmental studies have also tested this hypothesis. For instance, Bannard and Matthews (2008) used a sentence repetition task and found that 2- and 3-year-old children are more likely to repeat frequent sentences correctly (e.g., you want to play) compared with less frequent ones (e.g., you want to work). Arnon and Clark (2011), showed that 4-year-olds are better at producing irregular plurals when presented in a familiar context (e.g., On your feet). In the same vein, Janssen and Barber (2012) found multiword frequency effects in adults’ production latencies during a task where participants had to name drawings of noun and adjective pairs. Arnon and Snider (2010) also showed that comprehension is affected by multiword frequency. In a grammatical judgement task, adults processed frequent four-word phrases faster than less frequent ones, even when the frequency of the individual final words, bigrams and trigrams were controlled for. It is worth noting that sensitivity to statistical properties of multiword sequences seems to be present early on. Indeed, it has been shown that 11- and 12-month-olds can already discriminate frequent multiword sequences from infrequent ones (e.g., take it off vs shake it of, Skarabela et al., 2021). Moreover, it has been demonstrated that multiword sequences acquired early in childhood are processed faster in adulthood (Arnon et al., 2017).
Similarly, written language abounds with distributional cues (Arciuli & Simpson, 2012; Snell & Theeuwes, 2020; Treiman et al., 2014). Reading behaviour, for example, has also been shown to be influenced by the frequency and predictability of multiword phrases. For instance, frequent three-word binomial phrases (e.g., black and white) are read faster than their reversed forms (i.e., white and black; Siyanova-Chanturia, Conklin, & van Heuven, 2011) and idioms (e.g., at the end of the day—‘ultimately’) are read faster than non-idiomatic structurally equivalent counterparts (e.g., at the end of the war; Conklin & Schmitt, 2008; Siyanova-Chanturia, Conklin, & Schmitt, 2011).
In the past decades, research has mainly focused on isolated word learning (Pelucchi et al., 2009; Perruchet & Vinter, 1998; Saffran et al., 1996, 1997), leaving aside the question of how multiword sequences are acquired in real time. To date, only one study has addressed this issue in the context of first language acquisition. In an eye-tracking study, Conklin and Carrol (2020) presented participants with short stories containing existing English binomials in their canonical form (e.g., boys and girls), which were seen once, and novel binomials (e.g., goats and pigs), which were seen one to five times during the task. Participants were then presented with the existing and novel binomials in reverse (e.g., girls and boys, pigs and goats). They found that participants were sensitive to the co-occurrences of the novel binomials, which translated into faster reading times for the novel binomials as the number of co-occurrences increased. In addition, the results showed an advantage for forward novel binomials over their reverse forms after only four to five exposures, suggesting that participants very quickly detected and encoded the structure of the repeated pattern (see Sonbul et al., 2023, for a replication in second language acquisition).
Here, we propose to investigate how associations between multiword constituents other than binomials are established over time by using a visual lexical decision task. Based on the assumption that vocabulary acquisition and performance on the Hebb repetition learning paradigm (Hebb, 1961) are subserved by the same processes (Mosse & Jarrold, 2008; Norris et al., 2018; Page et al., 2013; Page & Norris, 2009; Smalle et al., 2016; Szmalec et al., 2009), we used an adaptation of the Hebb letter naming task by Rey et al. (2020) to study the learning dynamics of repeated words triplets.
In the original Hebb repetition task, participants had to recall sequences of digits where one particular sequence was repeated every third trial. Hebb (1961) found that participants’ performance gradually improved for the repeated sequences compared with the non-repeated ones. In the study by Rey et al. (2020), participants had to read aloud the names of single letters that were presented one at a time on a computer screen. Unknown to the participants, a triplet of letters (i.e., the Hebb sequence) was repeated with its constituent letters systematically presented in the same order. As in the standard Hebb learning paradigm, random letters (i.e., fillers) were inserted between two repetitions of the critical letter triplets. The extraction dynamics of the repeated triplet was tracked by looking at the evolution of response times (RTs) to the second and third letters of the triplet. RTs for these two letters decreased with repetition as they progressively became predictable when learning occurred. To study the extraction dynamics of multiword sequences in this experiment, we replaced the triplet of letters used in the study by Rey et al. (2020) by a triplet of words and instead of using a naming task, we used a lexical decision task hence simplifying online data collection and providing a better proxy for the silent reading that occupies the vast majority of skilled reading behaviour.
The reasons for using the Hebb paradigm to investigate multiword acquisition are twofold. First, as the Hebb paradigm is an implicit learning measure, it allowed us to study the extraction dynamics of multiword sequences in conditions where participants were not necessarily aware of the repetitions. Indeed, as participants are asked to read words without further instructions, knowledge of patterns of sequences can be attributed to implicit learning through regularity extraction. Second, it allowed us to study the online learning trajectory of multiword sequences rather than solely the “offline” end-product of what has been learned. Indeed, participants’ knowledge can be the same at the end of the task (offline knowledge), but their learning trajectories may differ (Siegelman et al., 2017). By using an online learning task, we sought to provide a comprehensive characterisation of the process of word-to-word associative learning.
Measuring the evolution of RTs for a repeated triplet of items also allowed us to study separately the repetition effect from the predictability effect. Indeed, because a random number of filler items occurred between two repetitions of the triplet, the first item in the triplet was not predictable and the evolution of RTs for this item can be considered as providing a good estimate of the repetition effect. In contrast, items occurring at Positions 2 and 3 of the triplet benefit from the immediately preceding item that systematically occurs before them and that should help participants anticipating and predicting the next item. Previous studies in sequence learning (Minier et al., 2016; Rey et al., 2019, 2020, 2022) even reported a stronger predictability effect on the third item of the triplet (i.e., a greater decrease in RTs) due to the richer contextual information provided by the two previous items. This experimental paradigm therefore allowed us to study the differential effect of repetition and predictability on the memory trace of each item belonging to a repeated triplet and on the processing gains generated by these effects.
Note that the predictability effect is closely linked to chunking mechanisms as it reflects the emergent association between several words that appear repeatedly in a sequence. As previously mentioned, chunking mechanisms are also considered central to several models of sequence learning and language acquisition (French et al., 2011; Jones & Rowland, 2017; McCauley & Christiansen, 2019; Perruchet & Vinter, 1998; Robinet et al., 2011). However, less is known about the precise dynamics related to the repeated presentation of a sequence of words and empirical evidence is needed to constrain models that assume a central role for chunking mechanisms in the development of language processing skills. This set of experiments has been designed to provide such empirical evidence about the dynamics of these fundamental associative learning mechanisms.
In this study, the learning and chunking dynamics of repeated triplets was studied in four Hebb lexical decision experiments. In Experiment 1, the repeated word triplet was composed of three unrelated words. In Experiment 2, the repeated triplet was composed of three pseudowords to test whether lexicality had an effect on the learning dynamics of the triplet. In Experiment 3, the repeated triplet was composed of three semantically related words to test whether semantic relatedness would facilitate the development of word associations. In Experiment 4, the repeated triplet corresponded to an existing idiomatic expression to test whether the learning trajectory of the repeated triplet would be facilitated by activating the pre-existing long-term memory representation of the triplet. These experiments were conducted remotely by using a platform for online experimentation that has been frequently used in experimental psychology to conduct experiments during the COVID-19 pandemic (Fournet et al., 2022; Isbilen et al., 2022; Ordonez Magro et al., 2022). It is worth noting that recent research has shown that JavaScript-based online experiment platforms, such as LabVanced and PsychoJS, allow researchers to collect reliable data that replicate the findings of in-lab studies (Angele et al., 2023; Mirault et al., 2018).
Experiment 1
Methods
Participants
Forty-two participants (20 females; Mage = 24 years, SD = 3) were paid for taking part in the experiment via Prolific (www.prolific.co). All participants reported to be native French speakers, having no history of neurological or language impairment.
Given that participants were recruited online, their proficiency in French was measured with the LexTALE language proficiency test (Brysbaert, 2013) before starting the main task. This test consists of a lexical decision task with no time pressure where participants are presented with 84 single-item trials (56 real French words, 28 French-looking pseudowords), and are instructed to decide whether each presented letter sequence is a real French word or not. Their average LexTALE vocabulary score was 86.53% (SD = 5.76). Any participant whose score was below 2.5 standard deviations from the average LexTALE vocabulary score was excluded from the analysis. No participant was excluded based on this criterion. The final data set consisted of 1,890 data points per condition, meeting the 1,600 measurements per condition recommendation from Brysbaert and Stevens (2018). A summary of the participants’ scores and standard deviations on the LexTALE task for each experiment is provided in the online Supplementary Material A.
Materials
We adapted the naming task by Rey et al. (2020) into a lexical decision task. The task was composed of three blocks of 120 trials, each trial corresponding to the presentation of a single word (or pseudoword) in the middle of the screen. A set of 66 words and 180 pseudowords were used as items in this experiment. All words were monosyllabic or disyllabic singular nouns. They were composed of four-to-six letters and were selected from the French database Lexique 3.83 (New & Pallier, 2020). Each word of the triplet had a freqfilms2 frequency ranging from 2 to 10 occurrences per million. We decided to use low-frequency words to maximise repetition effects and increase the chances of revealing any processing differences between positions within the triplet. Indeed, low-frequency words elicit larger repetition effects compared with high-frequency words in lexical decision tasks (Scarborough et al., 1977). Filler words had a frequency ranging from 10 to 100 occurrences per million. Pseudowords were drawn from the French Lexicon Project (Ferrand et al., 2010). They were monosyllabic or disyllabic and had a length from four to six letters.
A Latin-square design was used such that each word of the triplet appeared in every possible position within the triplet across participants, leading to six possible combinations of the same triplet of words (ABC, ACB, BAC, BCA, CAB, CBA). Seven triplets of words were used and were seen in one of the six possible combinations (for a total of 7×3 = 21 words). Each participant saw one triplet in a specific combination, leading to 7×6 = 42 participants (e.g., Participant 1 saw ABC while Participant 2 saw ACB instead throughout the task). Each triplet appeared 15 times per experimental block (resulting in a total of 45 repetitions across the three blocks) and was separated by three to six filler words or filler pseudowords (75 per block). Every block was composed of 60 words (the 15 repeated triplets, i.e., 45 words and 15 filler words) and 60 pseudowords. Therefore, there were an equal number of “yes” and “no” responses in the experiment (i.e., 180 for each type of response). Among the 66 selected words, 21 served to construct the 7 triplets and 45 served as filler words during the experiment. The set of word triplets and fillers are listed in the online Supplementary Material B.
To obtain more detailed information about participants’ explicit knowledge of the task, all participants responded to a short questionnaire after the experiment (similarly to Rey et al., 2020; Tosatto et al., 2022). The first question was: “Did you notice anything particular in this experiment?,” in case of a “Yes” response, the follow-up question was “Can you explain what you noticed?” If participants reported noticing the presentation of a repeated sequence of words, they were asked “Can you recall the words in their correct serial order?.” If the answer to the first question was “No,” the following questions were displayed “Did you notice that a sequence of words was systematically repeated?” and “Can you recall the words in their correct serial order?.”
Apparatus
The experiment was implemented in LabVanced, an online experiment builder (Finger et al., 2017) and participants were recruited via the Prolific platform (www.prolific.co). Participants participated via their personal computer and we made sure that the experiment would not work on smartphones or tablets to keep the testing conditions as similar as possible across participants. All words and pseudowords were presented in the centre of the computer screen using a 20-point Lato black font on a white background.
Procedure
Before the experiment, written instructions were displayed on the screen. Participants were instructed to decide as fast as possible whether the letter sequence displayed on the screen formed a French word or not. They were required to press “M” (for words) or “Q” (for pseudowords) on their keyboards (which are at extreme positions on the left and right of French AZERTY keyboards). RTs and accuracy were recorded for each word and pseudoword. Each target stayed on the screen until the participant’s response. Subsequently, the next target appeared immediately after the participant’s response. To encourage the participants, the number of remaining trials was displayed at the end of each block. The experiment lasted approximately 10 min. Figure 1 provides a schematic description of this experimental paradigm.

Experimental procedure for the Hebb lexical decision paradigm. Upper part: items are presented one at a time at the centre of the computer screen. Participants had to classify each string as a word or a pseudoword. A repeated triplet of three words (e.g., W1: “mule”—mule; W2: “proie”—prey; W3: “noeud”—knot) always appearing in the same order was intermixed with random filler words (WR) or random filler pseudowords (PWR). Words in blue belong to the repeated triplet. Lower part: one triplet of words (W1-W2-W3) is repeated several times and a variable number of random words or pseudowords (WR or PWR) are presented between two repetitions of the triplet.
Results
Only correct trials were analysed (97.06 % of the data), and we excluded RTs exceeding 1,500 ms (0.98 % of data) as well as RTs greater than 2.5 standard deviations above a participant’s mean per block and for each of the three possible positions within the triplet (2.47 %). The mean RTs and standard deviations computed over the entire sample and for each block are presented in Table 1. Data analysis was performed with the R software (version 4.2.1) using linear mixed-effects models (LMEs) fitted with the lmerTest (version 3.1-3; Kuznetsova et al., 2017) and the lme4 packages (version 1.1-29; Bates et al., 2015). The model included the maximum random structure that allowed convergence (Barr, 2013; Barr et al., 2013), that is, Position (1 to 3), Repetition (1 to 45) and their two-way interaction as fixed effects, participant and item sets were used as random effects. It is worth noting that Position was coded using repeated contrast coding (i.e., Position 1: –0.7 –0.3; Position 2: 0.3 –0.3; Position 3: 0.3 0.7) to perform pairwise comparisons (Schad et al., 2020), and Repetition was mean centred here and in the following analyses. Word length and log-transformed word frequency for each word in the triplet were included as covariates to control for any word-level differences. Given that the distribution of RTs was close to normal and provided good fit (established through visual inspection of QQ plots and histograms), no data transformation was performed prior to the analysis. The results of the model are shown in Table 2.
Mean response times (in milliseconds) and standard deviations (in parentheses) for each block and each position in Experiment 1.
Fixed effects of the mixed model for Experiment 1.
SE: standard error; CI: confidence interval.
We found a significant effect of Repetition with an overall decrease of RTs across the experiment. As predicted, RTs for Position 2 were significantly faster than those for Position 1, but they did not differ from Position 3. Moreover, there was a significant negative interaction coefficient for the difference between Position 2 and Position 1, and Repetition, and a significant positive interaction coefficient for the difference between Position 3 and Position 2, and Repetition, indicating that RT differences increased across repetitions. No significant effects were found for word length and word frequency. To investigate where the significant difference between Position 1 compared to Positions 2 and 3 emerges, we ran a series of paired sample t-tests on the RTs for Position 1 and the average RTs for Positions 2 and 3 on each repetition of the triplet. We found that a significant difference emerged on the fifth trial, t(38) = 5.26, Bonferroni-adjusted p < .001.
To get a clearer picture of the learning dynamics for each position in the triplet of words, Figure 2 represents the evolution of the mean RTs for each position in the triplet and for the successive 45 repetitions of the triplet. Given that linear regression only captures the overall change of position across repetitions, we conducted a broken-stick linear regression, using the segmented package (version 1.6-0; Muggeo, 2008), to account for the evolution of the learning pattern across the task. In broken-stick regression, multiple linear regressions are fitted and connected at certain estimated values referred as breakpoints. At the breakpoint, the relationship between the variables changes to model non-linear relationships between two variables. Thus, each position was regressed onto repetition separately. To estimate the number of breakpoints for each position, a broken-stick regression model was built incrementally (i.e., we added a breakpoint estimate to each successive model). For each model, an initial guess for the breakpoint was provided, and then the optimal breakpoints were calculated by the model using an iterative fitting procedure with the default package parametrisation (see Muggeo, 2008, for technical details). We compared each new model with the previous one (based on chi-square analysis) and selected the most parsimonious as the final model. For Position 1, the analysis revealed a breakpoint at repetition 18.46, 95% confidence interval (CI) = [14.26, 22.67], with RTs decreasing from repetitions 1 to 18.46, b = –4.22, 95% CI = [–5.70, –0.29], followed by a slow increase, b = 0.52, 95% CI = [–0.29, 1.33]. For Position 2, we estimated two breakpoints at repetitions 5.35, 95% CI = [3.54, 7.16] and 19.81, 95% CI = [13.63, 25.98], with RTs rapidly decreasing from repetitions 1 to 5.35, b = –32.40, 95% CI = [–46.24, –18.57], continuing to decrease, but at a slower rate, from repetitions 5.35 to 19.81, b = –7.92, 95% CI = [–10.78, –5.07], followed by a slower decrease until the end of the task, b = –3.20, 95% CI = [–4.34, –2.07]. For Position 3, we also estimated two breakpoints at repetitions 5.88, 95% CI = [3.52, 8.24] and 19.54, 95% CI = [15.36, 23.71], with a fast decrease in RTs from repetitions 1 to 5.88, b = –27.92, 95% CI = [–41.01, –14.85], continuing to decrease at a slower rate from repetitions 5.88 to 19.54, b = –8.18, 95% CI = [–10.86, –5.50], and with an even slower decrease from repetition 19.54 until the end of the experiment, b = –1.72, 95% CI = [–2.78, –0.65].

Upper panel: mean response times in Experiment 1 as a function of word position and number of repetitions of the triplet. The vertical dashed line indicates the first repetition at which there was a significant difference between Position 1 vs Positions 2 and 3. Error bars indicate 95% confidence intervals. Lower panel: results from the broken-stick regressions for each Position in the triplet. Vertical bars indicate the breakpoints.
Questionnaire
Forty-one of the 42 participants reported noticing a recurrent word sequence; 16 were able to recall the whole triplet, 12 correctly recalled one sub-sequence (Words 1 and 2 or Words 2 and 3), 4 could recall non-adjacent words (Words 1 and 3), 7 only recalled one word, and the 3 remaining participants did not recall any word.
Discussion
As expected, the results from Experiment 1 showed faster RTs for predictable words (i.e., Words 2 and 3) within the repeated triplet, and the difference between unpredictable (Word 1) and predictable items increased as the task progressed. Furthermore, this difference between unpredictable and predictable items emerges early on, around the fifth repetition of the triplet. The analysis of the mean RTs over the 45 repetitions of the triplet further indicated that learning occurred also for words appearing in Position 1 of the triplet. Although unpredictable, these words were repeated and their processing was facilitated by this repetition. The broken-stick regression analysis suggested that learning occurred during the first 18 repetitions and subsequently reached a plateau performance. While the mean RT for the first occurrence of these words was 682 ms, the mean RT was 561 ms after 18 repetitions, and 592 ms at the 45th repetition, indicating a processing speed-up of 90 ms between the first and last occurrence of the word. These data, therefore, provide an estimate of the dynamics of the repetition effect for words that are not predictable.
In contrast, RTs for predictable words (i.e., on Positions 2 and 3) followed a totally different dynamic. According to the broken-stick regression analysis, they indeed decreased very rapidly during the first five repetitions (640 ms at the first repetition, and 523 ms at the fifth repetition—RTs are averaged over Positions 2 and 3) and the decrease was slower between Repetitions 5 and 18 (419 ms at the 18th repetition). After the 18th repetition, RTs continued to decrease but at an even slower rate (347 ms at the 45th repetition). Clearly, compared with the results obtained for words at Position 1 of the repeated triplet, we found that the predictability effect was much larger than the repetition effect and followed different learning dynamics. For example, for the third position of the triplet, the mean RTs were 624 ms for the first occurrence of the word and 349 ms for the 45th repetition, resulting in a processing gain of 275 ms between the first and last occurrence of these words.
Interestingly, there was no evidence for an advantage of the third over the second word in the triplet, contrary to what was observed by previous studies. Indeed, prior findings indicated faster RTs for the final stimulus in a repeated triplet, as it benefits from the cumulative information provided by the two preceding stimuli (Minier et al., 2016; Rey et al., 2019, 2020, 2022). Regarding our study, although the words clearly benefitted from immediate contextual information (i.e., the preceding word in the triplet that systematically appeared before them), we did not observe any additional predictability effect regarding the final word of the triplet when the context was richer (i.e., words in Position 3 of the triplet benefit from the contextual information provided by words in Positions 1 and 2). This intriguing result likely reflects some limitations of associative and Hebbian learning mechanisms due to the specific time-scale of the present experimental paradigm. We will return to this issue in the general discussion.
Despite a clear decrease in RTs for the predictable positions in the triplet, indicating that learning of this repeated sequence occurred, most participants were unable to correctly recall the whole triplet, even though most of them noticed the presence of a repeated sequence. This result suggests that part of the triplet learning was explicit but that most of the learning was probably implicit. Participants did not have to explicitly encode the triplet repetition to anticipate the occurrence of words appearing on predictable positions.
In contrast to Experiment 1, which was conducted with triplets of unrelated words, Experiment 2 was conducted with triplets of pseudowords. We decided to use pseudowords because tasks consisting of the repetition and encoding of pseudoword sequences have been shown to mimic novel word learning (Norris et al., 2018; Schimke et al., 2021). Indeed, whereas words are likely to have long-term memory representations, pseudowords cannot benefit from such representations as they have not yet been encountered by participants. It is worth noting that the Hebb paradigm has also been described as a laboratory analogue of novel word learning (Szmalec et al., 2009, 2012). Therefore, studying triplets of pseudowords will allow us to compare the learning dynamics of completely novel multiword sequences with those obtained for already known words in Experiment 1.
Experiment 2
Methods
Participants
Forty-six participants (22 females; Mage = 25 years, SD = 3) were recruited from Prolific (www.prolific.co) for the experiment. All participants indicated that French was their native language and declared no neurological or language impairment. Four participants were excluded from the analyses due to chance-level performance on the main task.
As in Experiment 1, participants’ French proficiency was measured with the LexTALE test (Brysbaert, 2013). Participants’ average scores were 85.13% (SD = 7.08). No participant was excluded from the analysis. The final number of participants was 42, which corresponds to a data set of 1,890 data points per condition.
Materials
In contrast to the previous experiment, here the target triplets were composed of pseudowords whereas the words served only as fillers items. We selected 180 words from the French database Lexique 3.83 (New & Pallier, 2020). All words were monosyllabic or disyllabic singular nouns and had a length from four to six letters. Their freqfilms2 frequency was between 10 and 100 occurrences per million. A set of 66 pseudowords was selected from the French Lexicon Project (Ferrand et al., 2010). Twenty-one were drawn therefrom to construct triplets and the remaining 45 were used as filler pseudowords. All pseudowords were four-to-six letter long and monosyllabic.
Seven triplets were generated and counterbalanced across participants using a Latin-squared design. Every triplet repetition of pseudowords (15 per block) was always separated by three to six filler words or filler pseudowords (75 per block). As in Experiment 1, each block was composed of 60 words and 60 pseudowords. There were an equal number of “yes” and “no” responses in the experiment (i.e., 180 for each type of response). The sets of pseudoword triplets and fillers are listed in the online Supplementary Material C.
Apparatus and procedure
The apparatus and procedure were identical to the one used in Experiment 1.
Results
As the target triplets were made up of pseudowords, only correct “no” responses were analysed (95.87% of the data), and RTs exceeding 1,500 ms (1.43% of data), as well as RTs beyond 2.5 standard deviations from a participant’s mean per block and for each of the three possible positions within the triplet (2.01%) were excluded. Means and standard deviations per block are shown in Table 3. The linear mixed model we fitted included the maximum random effect structure allowing convergence (Barr, 2013; Barr et al., 2013). This model included position, repetition, and the interaction term as fixed effects. Item and participant were used as crossed random effects, with by-participant random slopes for position. The results of the mixed model are summarised in Table 4. Figure 3 provides the evolution of mean RTs for each position in the triplet and for the 45 repetitions of the triplet.
Mean response times (in milliseconds) and standard deviations (in parentheses) for each block in Experiment 2.
Fixed effects of the mixed model for Experiment 2.
SE: standard error; CI: confidence interval.

Upper panel: mean response times in Experiment 2 as a function of pseudoword position in the repeated triplet and number of repetitions. The vertical dashed line indicates the first repetition at which there was a significant difference between Position 1 vs Positions 2 and 3. Error bars indicate 95% confidence intervals. Lower panel: Results from the broken-stick regressions for each position in the triplet. Vertical bars indicate the breakpoints.
Results indicated a significant effect of repetition and faster RTs for pseudowords in Position 2 compared to those in Position 1, as well as for Position 3 compared to Position 2. Moreover, there was a significant negative interaction coefficient for the difference between Position 2 and Position 1, and repetition, and a significant positive interaction coefficient for the difference between Position 3 and Position 2, and repetition. Similarly to Experiment 1, paired sample t-tests comparisons showed a significant difference between Position 1 compared to Positions 2 and 3 on the sixth trial, t(36) = 3.12, Bonferroni-adjusted p = .021.
As for the first experiment, we conducted a broken-stick regression to study the evolution of the learning pattern of position across the task. The analysis revealed a breakpoint at repetition 5.18, 95% CI = [3.58, 6.78] for Position 1, with RTs decreasing from repetitions 1 to 5.18, b = –23.81, 95% CI = [–36.78, –10.86], followed by a slower decreasing rate, b = –1.82, 95% CI = [–2.38, –1.26]. For Position 2, two breakpoints were estimated at repetitions 7.44, 95% CI = [5.59, 9.29] and 22.00, 95% CI = [17.56, 26.44], with RTs rapidly decreasing from repetitions 1 to 7.44, b = –39.33, 95% CI = [–50.61, –28.05], continuing to decrease, but at a slower rate from repetitions 7.44 to 22.00, b = –10.16, 95% CI = [–14.00, –6.32], followed by a slower decrease until the last repetition, b = –1.16, 95% CI = [–2.87, 0.56]. Regarding Position 3, we estimated two breakpoints at repetitions 6.07, 95% CI = [4.40, 7.74] and 18.37, 95% CI = [14.74, 21.99], with RTs decreasing fast from repetitions 1 to 6.07, b = –41.73, 95% CI = [–54.71, –28.74], steadily decreasing at a slower rate from repetitions 6.07 to 18.37, b = –11.66, 95% CI = [–16.00, –7.33], followed by a slower decrease until the end of the task, b = –1.86, 95% CI = [–3.13, –0.59].
Given that usage-based theories postulate that novel items become lexicalised when they are encountered sufficiently often (Bybee, 2006; Zang et al., 2023), one might expect that after enough repetitions participants would begin to consider the target pseudowords to be almost as real words, resulting in more false “yes” judgements as the experiment progressed. Therefore, we conducted an additional analysis using a generalised (logistic) linear mixed model to compare the mean accuracy between positions across blocks (see Figure 4). The model was fitted with position and block, and the interaction term as fixed effects. The maximal random effects structure that converged was one that included by-participant and by-item random intercepts. To explore differences between positions within each block, we used the R package emmeans (Lenth, 2023). Helmert contrasts were used to compare Position 1 with both Positions 2 and 3, simultaneously, and to compare Position 2 with Position 3. The results of the contrasts are summarised in Table 5. The analysis showed that, systematically across the three blocks, participants made more false “yes” judgements for pseudowords in Position 1 than for those in Positions 2 and 3. In addition, in Block 3, participants made more false “yes” judgements for pseudowords in Position 2 compared to those in Position 3. Finally, false “yes” judgements for pseudowords in Position 1 increased across the blocks, in contrast to those in Positions 2 and 3.

Mean accuracy in Experiment 2 as a function of pseudoword position in the repeated triplet and block number. Error bars indicate 95% confidence intervals.
Summary of Helmert contrasts between positions across blocks for Experiment 2.
SE: standard error ; P: Position.
Questionnaire
Thirty-nine participants reported noticing a recurrent pseudoword sequence; 12 were able to recall the whole triplet, one could recall one subsequence (Words 2 and 3), eight correctly recalled non-adjacent pseudowords (Words 1 and 3), eight only recalled one pseudoword, and the 13 remaining could not recall any pseudoword.
Discussion
The results of Experiment 2 partly replicated those of Experiment 1. A first main difference between the two experiments concerns the overall slower RTs obtained for pseudowords compared with words: when averaging the RTs of all three positions, the mean RTs on their first occurrence was 654 ms for words and 810 ms for pseudowords; on their last occurrence (i.e., at the 45th repetition), the mean RTs was 429 ms for words and 470 ms for pseudowords. Apart from these longer RTs, the learning dynamics also produced noticeable differences compared with the one observed for words.
Regarding the repetition effect that is measured by the evolution of RTs for pseudowords occurring at Position 1 of the triplet, the dynamics was clearly different compared with words with a fast decrease of RTs during the first five repetitions (with a mean RT of 798 ms for the first occurrence and of 704 ms for the fifth repetition), followed by a smoother decrease until the last repetition (with a mean RT of 638 ms for the 45th repetition). While the beta coefficient of the first regression line was –4.22 for words, it was much larger for pseudowords (–23.81). The processing gain for pseudowords at Position 1 (i.e., the difference between mean RTs for the last repetition and the first occurrence) was 160 ms, which is much larger than the one obtained for words (90 ms). Pseudowords seem therefore to benefit to a larger extent from the repetition effect indicating that repetitions produced a fast change in the way these pseudowords were processed and in the way their trace developed in memory.
For predictable pseudowords (i.e., in Positions 2 and 3 of the triplet), the broken-stick regression analysis also identified two break points that were slightly different from those obtained with words (for pseudowords, 7.44 and 22 at Position 2, and 6.07 and 18.37 at Position 3; for words, 5.35 and 19.81 at Position 2, and 5.88 and 19.54 at Position 3). Apart from these differences, the learning dynamics were similar with a fast decrease in RTs during the initial repetitions followed by an intermediate decrease and a slower one during the last repetitions. Compared with the repetition effect, the predictability effect was again much larger and produced a much stronger processing gain (i.e., for the third position, when subtracting the mean RTs for the 45th repetition, 390 ms, from the mean RT for the first occurrence, 779 ms, the processing gain was 779 – 390 = 389 ms).
Contrary to Experiment 1, the data revealed a significant difference between Positions 2 and 3, with faster RTs on Position 3 of the triplet. This difference seems to emerge around the same time as in Experiment 1, namely on the sixth repetition of the triplet. Although this result is consistent with previous finding in sequence learning, here it might be an artefact due to the fact that participants were slower to classify the pseudowords in Position 2 at the beginning of the task, resulting in a higher estimation of the regression intercept compared to the one of Position 3. Due to this unexpected initial difference (that should have been cancelled by the Latin square design), this difference between Positions 2 and 3 is difficult to interpret.
In addition, we found that as the task progressed, it became more difficult for participants to classify the first item of the triplet as being a pseudoword. Indeed, they systematically made more false “yes” judgements for pseudowords in Position 1 than for those in Positions 2 and 3. Interestingly, false “yes” judgements for pseudowords in Position 1 increased over the course of the task, in contrast to those for pseudowords in Positions 2 and 3. This finding, consistent with usage-based theories, suggests that participants gradually became familiar with the first pseudoword of the repeated triplet, which presumably became lexicalised over time. As a result, participants were more likely to respond incorrectly to the first pseudoword in the triplet. Once they recognised the first pseudoword, they simply had to respond correctly to the rest of the triplet. It is worth noting that in Block 3, participants were also more likely to consider the second pseudoword in the triplet to be a word compared to the third, suggesting that the triplet was becoming progressively lexicalised as well.
As for Experiment 1, the number of participants who reported detecting a recurring sequence was high (93%) but the number of participants who were able to fully recall the triplet was much lower (29% in Experiment 2 compared to 38% in Experiment 1). Here again, the data suggest that learning occurred both implicitly and explicitly, and the rate of explicit learning (i.e., with a full recall of the triplet) was lower for pseudowords (29%) than for words (38%).
Overall, Experiment 1 and Experiment 2 yielded similar results regarding the learning dynamics of the repeated triplet, that is, a slower learning rate on the first unpredictable position due to a simple repetition effect, and a much larger learning rate for the predictable positions (i.e., the second and the third) due to the predictability effect. However, in both experiments and contrary to natural language, words and pseudowords were totally unrelated and apart from systematically occurring one after the other, there was no other reason to associate these items. In Experiment 3, we tested whether the use of a triplet composed of semantically related words (e.g., belonging to the same word category such as, for example, the fruit category: strawberry, banana, cherry) could have an effect on the learning dynamics of the triplet. We expected semantic relatedness to facilitate learning both at the implicit level (i.e., on RTs) and at the explicit level (i.e., on the recall of the triplet).
Experiment 3
Methods
Participants
Forty-two participants (22 females; Mage = 23 years, SD = 4) were paid and recruited via Prolific (www.prolific.co). All participants were native French speakers and reported having no neurological or language disorders. The average LexTALE vocabulary score (Brysbaert, 2013) was 85.08% (SD = 6.37), and no participant was excluded.
Materials
To construct seven semantically related triplets, we selected 21 low-frequency words from the database Lexique 3.83 (New & Pallier, 2020). All words were four-to-six letters monosyllabic or disyllabic singular nouns and had a freqfilms2 frequency ranged from 2 to 10 occurrence per million. Forty-five additional words and 180 pseudowords were selected and used as filler items between two repetitions of the target triplet. All filler words were monosyllabic or disyllabic singular nouns and were composed of four to six letters. Their freqfilms2 frequency ranged from 10 to 100 occurrences per million. Pseudowords were retrieved from the Lexicon Project (Ferrand et al., 2010), were monosyllabic or disyllabic, and were composed of four to six letters.
A Latin-square design was used, leading to the generation of seven triplets for the 42 participants (i.e., 6 participants per triplet). Every triplet repetition (15 per block) was separated by three to six filler words or filler pseudowords (75 per block). Sixty words and 60 pseudowords were presented in each block. There were an equal number of “yes” and “no” responses in the experiment (i.e., 180 for each type of response). Stimuli are listed in the online Supplementary Material D.
Apparatus and procedure
The apparatus and procedure were identical to that used in Experiments 1 and 2.
Results
Only correct responses were analysed (96.86% of the data). RTs exceeding 1,500 ms (1.32% of data) and RTs greater than 2.5 standard deviations from a participant’s mean per block and for each of the three possible positions within the triplet (2.26%) were removed. Means and standard deviations per block are shown in Table 6. We constructed a linear mixed-effects model with the maximum random effect structure allowing convergence (Barr, 2013; Barr et al., 2013). This model included position, repetition, and the interaction term as fixed effects; participant and item were used as random intercepts with by-participant random slopes for Position. We included word length and log-transformed word frequency for each word in the triplet as covariates. Given that word associations have been shown to influence processing times in multiword sequences (Carrol & Conklin, 2020), and that the order of presentation of the words in the triplets varied across participants (because of the Latin-squared design), potentially affecting processing times as some words were more strongly associated than others, we also included a measure of association strength between triplet words as a covariate. As existing free-association databases in French do not contain all the items we used, we decided to calculate the indirect association strength between the words using the JeuxDeMots database (Lafourcade & Joubert, 2008). This database is based on a collaborative online project where participants see a word and provide an association, which is only validated if other peers have suggested the same association. These associations are then weighted according to the number of associations given by the participants to obtain the association strength. To calculate the indirect association strength between two target words, we generated a list of the most frequently associated words with the target word, then selected the most frequent common word between two target words and averaged the association strengths to obtain the indirect association strength measure. For instance, both banana and strawberry were associated with fruit (i.e., 526 and 480, respectively). To obtain the indirect association strength, we then averaged the two values, resulting in an indirect association strength of 503. The results of the model are summarised in Table 7. Figure 5 provides the evolution of mean RTs for each position in the triplet and for the 45 repetitions of the triplet.
Mean response times (in milliseconds) and standard deviations (in parentheses) for each block in Experiment 3.
Fixed effects of the mixed model for Experiment 3.
SE: standard error; CI: confidence interval.

Upper panel: mean response times in Experiment 3 as a function of word position and number of repetitions. The vertical dashed line indicates the first repetition at which there was a significant difference between Position 1 vs Positions 2 and 3. Error bars indicate 95% confidence intervals. Lower panel: results from the broken-stick regressions for each position in the triplet. Vertical bars indicate the breakpoints.
The results showed a significant negative effect of repetition reflecting a decrease in RTs. We also found faster RTs for words in Position 2 compared to those in Position 1, but not to those in Position 3. Finally, there was a significant negative interaction coefficient for the difference between Position 2 and Position 1, and Repetition, and a significant positive interaction coefficient for the difference between Position 3 and Position 2, and Repetition. No significant effects were found for word length, word frequency and association strength for both bigrams. Paired sample t-tests comparisons showed that a significant difference between Position 1 compared to Positions 2 and 3 emerged on the third trial, t(39) = 3.39, Bonferroni-adjusted p = .005.
Following the same procedure as in Experiments 1 and 2, we performed a broken-stick regression on each position of the repeated triplet. For Position 1, the analysis revealed a breakpoint at repetition 16.72, 95% CI = [11.34, 22.10], with RTs decreasing from repetitions 1 to 16.72, b = –3.96, 95% CI = [–5.80, –2.11], followed by an almost flat slope, b = 0.01, 95% CI = [–0.76, 0.77]. Concerning Position 2, we estimated two breakpoints at repetitions 4.64, 95% CI = [3.08, 6.20] and 20, 95% CI = [16.20, 23.80], with RTs rapidly decreasing from repetitions 1 to 4.64, b = –43.01, 95% CI = [–63.16, –22.86], continuing to decrease, but at a slower rate from repetitions 4.64 to 20, b = –8.85, 95% CI = [–11.27, –6.42], followed by a slower decrease until the end of the task, b = –1.37, 95% CI = [–2.62, –0.11]. For Position 3, we also estimated two breakpoints at repetitions 5.57, 95% CI = [3.89, 7.24] and 18.85, 95% CI = [14.49, 23.22], with a fast decrease in RTs from repetitions 1 to 5.57, b = –35.23, 95% CI = [–49.00, –21.47], continuing to decrease at a slower rate from repetitions 5.57 to 18.85, b = –7.62, 95% CI = [–10.77, –4.46], followed by a slower decrease until the last repetition, b = –0.85, 95% CI = [–1.92, 0.21].
Questionnaire
Forty-one of the 42 participants reported noticing a recurrent word sequence; 29 were able to recall the whole triplet, one recalled one subsequence (Words 2 and 3), four could recall non-adjacent words (Words 1 and 3), five recalled all the words but in the wrong order, and the three remaining could not recall any word.
Discussion
Experiment 3 produced similar results as in Experiment 1. Concerning the repetition effect, we did not expect any advantage of the semantic relatedness because there is no reason to observe any effect of this variable on the first word of the triplet. And indeed, the dynamics of the repetition effect was very similar to the one obtained in Experiment 1.
For predictable items (in Positions 2 and 3 of the triplet), the beta coefficient of the first regression line (from the broken-stick regression analysis) was larger (–43.01 for related words compared with –32.4 for unrelated words) and the first breakpoint occurred earlier (4.64 compared with 5.35), suggesting that the initial learning phase was much steeper in the semantically related condition compared with the unrelated words from Experiment 1. The semantical relatedness between these words helped producing a larger predictability effect that certainly took advantage of the pre-existing semantic associations between these words. This was also confirmed by the fact that a difference between unpredictable and predictable items emerges earlier than in Experiment 1 (i.e., around the third rather than the fifth repetition of the triplet). Note that this advantage was only present at the early phase of learning because the processing gain for words in Experiment 1 is similar to the one obtained in Experiment 3. Indeed, the difference between the mean RTs on Position 3 for the first and last occurrence of these items was 624 – 349 = 275 ms in Experiment 1 and 624 – 360 = 264 ms in Experiment 3. Finally, as for Experiment 1, there was no additional advantage for items occurring in Position 3 of the triplet compared to those being in Position 2.
Like Experiment 1, the number of participants who reported detecting a recurring sequence was high (98%) but the number of participants who were able to fully recall the triplet was much larger (69% compared to 38% in Experiment 1). Clearly, the semantic relatedness may have helped participants encoding the triplet in an explicit way, which probably also explains the stronger predictability effect observed during the early phase of learning.
As expected, semantic relatedness had a facilitatory effect not only on the predictability effect but also on the ability of participants to explicitly memorise the repeated triplet and to recall it. However, this situation is rather artificial given that words belonging to the same semantic category rarely appear in a sequence when reading texts, apart from special cases such as binomials (e.g., salt and pepper, boys and girls, knife and fork), which are often composed of words belonging to the same semantic category. It has been shown that the association strength of the component words in binomials influences reading times in a natural reading task (Carrol & Conklin, 2020). We therefore tested whether the learning dynamics of a triplet would be improved by using words that often co-occur, such as idioms. A recent study has indeed shown that meaningful three-word sequences (e.g., idioms: on my mind; phrase: is really nice) are easier to process and lead to faster RTs compared with fragment sequences (e.g., because it lets) in a phrasal decision task (Jolsvai et al., 2020). Similarly, Northbrook et al. (2022) presented Japanese English speakers with a series of short stories containing repeated three-word lexical bundles, each seen three times, followed by a phrasal decision task. They found that repeated lexical bundles (e.g., set off home, tired and hungry) were processed faster than non-repeated bundles in the phrasal decision task, with faster RTs at each subsequent repetition. This advantage for repeated lexical bundles emerged from the first repetition and was still present a week later. In Experiment 4, we therefore used three-word idioms as repeated triplets to study whether the presence of frequently co-occurring words increases the predictability effect. We expected idioms to facilitate learning as they have already been encountered and encoded in memory as whole sequences by the participants.
Experiment 4
Methods
Participants
Forty-two participants (21 females; Mage = 24 years, SD = 4) were recruited and paid to take part in the study via Prolific (www.prolific.co). All participants were native French speakers and reported having no neurological or language impairments. Their average LexTALE vocabulary score (Brysbaert, 2013) was 86.18% (SD = 6.19), no participant was excluded from the analysis.
Materials
We constructed the triplets by selecting seven three-word idiomatic expressions from two databases of French idioms rated by native speakers (Bonin et al., 2013, 2018). Filler items that were inserted between two repetitions of the triplet were 45 words and 180 pseudowords. Words were monosyllabic or disyllabic singular nouns and were chosen from the database Lexique 3.83 (New & Pallier, 2020). All words were four to six letters long and had a freqfilms2 frequency between 10 and 100 occurrences per million. Pseudowords were selected from the French Lexicon Project (Ferrand et al., 2010). All pseudowords were monosyllabic or disyllabic and were composed of four to six letters.
In contrast to previous experiments in which we used a Latin-square design, here the triplets were not scrambled, and therefore participants saw the idioms in their canonical form. Indeed, reversing the word order of existing idiomatic expressions has been shown to result in a processing penalty (Conklin & Carrol, 2020). Each of the seven idiomatic expressions were presented to six participants (6×7 = 42). Every triplet repetition (15 per block) was separated by three to six filler words or pseudowords (75 per block). As in the previous experiments, every block was composed of 60 words and 60 pseudowords. Therefore, there were an equal number of “yes” and “no” responses in the experiment (i.e., 180 for each type of response). Stimuli are listed in the online Supplementary Material E.
Apparatus and procedure
The apparatus and procedure were identical to the one used in Experiments 1 to 3.
Results
Only correct responses were analysed (96.34% of the data). RTs exceeding 1,500 ms (1.54% of data), and RTs beyond than 2.5 standard deviations from a participant’s mean per block and for each of the three possible positions within the triplet (2.19%) were removed. Mean RTs and standard deviations per block and position are shown in Table 8.
Mean response times (in milliseconds) and standard deviations (in parentheses) for each block and position in Experiment 4.
We constructed a linear mixed-effects model with the maximum random effect structure allowing convergence (Barr, 2013; Barr et al., 2013). This model included position, repetition, and the interaction term as fixed effects. Item and participant were used as crossed random effects, with by-participant random slopes for position. In addition to word length and word frequency, we also included idiom frequency and bigram and trigram mutual information 1 (MI) scores as covariates in our analysis. Indeed, previous research on idioms has shown that these factors can influence the processing of multiword sequences (Carrol & Conklin, 2020). Idiom frequency and MI scores were calculated based on the French web corpus frTenTen20 (Jakubíček et al., 2013), which consists of 20.9 billion words. All frequencies were log-transformed prior to analysis. The results of the model are summarised in Table 9. Figure 6 provides the evolution of mean RTs for each position in the triplet and for the 45 repetitions of the triplet.
Fixed effects of the mixed model for Experiment 4.
SE: standard error; CI: confidence interval; MI: mutual information.

Upper panel: mean response times in Experiment 4 as a function of word position and number of repetitions. The vertical dashed line indicates the first repetition at which there was a significant difference between Position 1 vs Positions 2 and 3. Error bars indicate 95% confidence intervals. Lower panel: results from the broken-stick regressions for each Position in the triplet. Vertical bars indicate the breakpoints.
The results showed a significant negative effect of repetition reflecting a decrease in RTs with repetitions. We also found faster RTs for words in Position 2 compared to those in Position 1, but there was no difference between Positions 3 and 2. There were a significant interaction coefficient for the difference between Position 2 and Position 1, and Repetition, as well as for the difference between Position 3 and Position 2, and Repetition. Finally, there was a significant effect of Idiom frequency, with less frequent idioms eliciting faster responses, and of Bigram MI, with faster RTs for bigrams with stronger MI. This unusual pattern is most likely due to the fact that one of our less frequent idioms in the experiment (i.e., qui dort dîne) has a high bigram MI score (i.e., 4.72), which may have speeded up participants’ responses even though the idiom frequency was low. In fact, any collocation above an MI score of 3 is considered to be strong. When this idiom is excluded from the analysis, the effect of idiom frequency is no longer significant, b = 16.65, SE = 11.22, 95% CI = [–5.33, 38.64], p = .147. Similar to Experiment 3, paired sample t-tests comparisons showed a significant difference between Position 1 compared to Positions 2 and 3 on the fourth trial, t(40) = 2.70, Bonferroni-adjusted p = .04.
We then performed a broken-stick regression to better account for the evolution of the learning pattern throughout the task. A breakpoint was estimated at repetition 15, 95% CI = [6.51, 23.49] for Position 1, with RTs decreasing from repetitions 1 to 15, b = –3.38, 95% CI = [–5.81, –0.95], followed by a slower decrease in RTs, b = –0.51, 95% CI = [–1.25, 0.23]. Regarding Position 2, we estimated two breakpoints at repetitions 6.72, 95% CI = [4.88, 8.57] and 19, 95% CI = [15.37, 22.63], with a fast decrease in RTs from repetitions 1 to 6.72, b = –35.45, 95% CI = [–46.72, –24.17], continuing to decrease but at a slower rate from repetitions 6.72 to 19, b = –9.87, 95% CI = [–13.27, –6.47], followed by a slower decrease until the end of the task, b = –1.51, 95% CI = [–2.72, –0.31]. For Position 3, two breakpoints were estimated at repetitions 5.17, 95% CI = [3.92, 6.41] and 17, 95% CI = [13.50, 20.50], with a strong decrease in RTs from repetitions 1 to 5.17, b = –44.36, 95% CI = [–58.21, –30.51], continuing with a slower decrease from repetitions 5.17 to 17, b = –9.79, 95% CI = [–13.31, –6.28], followed by an even slower decrease until the end of the task, b = –1.74, 95% CI = [–2.73, –0.75].
Questionnaire
All participants reported noticing a recurrent word sequence; 37 were able to recall the whole triplet, two recalled one subsequence (Words 2 and 3), one could recall non-adjacent words (Words 1 and 3), one recalled all the words but in the wrong order, and the last one could not recall any word.
Additional analysis
To compare the predictability effects observed in Experiments 1, 3, and 4, we computed a predictability score for these experiments by calculating a difference between log-transformed RTs for unpredictable words (Position 1) versus the log-transformed mean RT for predictable words (Positions 2 and 3) for each repetition. Here, a positive score reflects a predictability effect. We decided to use log-transformed values to control for baseline differences in the participants’ responses (see Siegelman, Bogaerts, Kronenfeld, & Frost, 2018). For instance, let us consider two participants with a mean difference of 100 ms between predictable and unpredictable words, but with a different baseline RT: P1 unpredictable = 600 ms, predictable = 500 ms; P2 unpredictable = 400 ms, predictable = 300 ms. Without this transformation, these participants would have the same difference score, even if the relative acceleration of P2 to predictable words is much higher. After log-transformation, the difference between predictable and unpredictable words reflects better this acceleration: log difference of P1 = 0.18, P2 = 0.29.
We then ran a linear mixed-effects model on the predictability scores, using experiment, repetition, and the interaction term as fixed effects, and participant as random effect. Experiment was coded using repeated contrast coding (Experiment 1: –0.7 –0.3; Experiment 3: 0.3 –0.3; Experiment 4: 0.3 0.7). We observed higher predictability scores in Experiment 4 (idioms) compared with Experiment 3 (semantically related words), b = 0.07, SE = 0.01, p < .001, and higher scores in Experiment 3 compared with Experiment 1 (non-related words), b = 0.09, SE = 0.01, p < .001. In addition, there was a main effect of repetition, b = 0.01, SE = 0.00, p < .001, and a significant interaction between Experiment 4 – Experiment 3 and Repetition, b = 0.002, SE = 0.001, p = .008, indicating an increasing difference of predictability scores between both experiments (see Figure 7).

Predictability scores across repetitions for word triplets in Experiments 1, 3, and 4. Continuous lines represent loess fit for the predictability scores. Dashed lines represent the best linear fit and grey-shaded areas indicate 95% confidence intervals around linear regression lines.
Discussion
Experiment 4 produced results similar to Experiment 3. However, two notable differences suggest that idioms have benefitted to a larger extent from triplet repetition compared with semantically related words. First, the predictability score represented in Figure 7 indeed shows that when the repetition effect is subtracted from the predictability effect on each repetition trial, the remaining predictability score is stronger for idioms compared with semantically related words, which is also stronger than the score obtained for unrelated words from Experiment 1. Idioms, which are supposedly already coded in the brain as semantically coherent and frequent sequences of words, appear to derive a greater processing advantage from repetition. Second, while 69% of participants in Experiment 3 were able to recall the full triplet of semantically related words, 88% of participants in Experiment 4 managed to recall the full idiom. This improved performance for explicit correct recall of idioms is probably due to their pre-existing encoding as relevant linguistic sequences, or at least to a facilitated access to them in memory, and it suggests more generally that frequent multiword sequences (apart from idioms) do result in a different learning dynamic in this Hebb lexical decision task compared with less frequent multiword sequences.
General discussion
The goal of this set of experiments was to provide empirical evidence about the dynamics of multiword sequence extraction by studying the evolution of RTs for a repeated triplet of items in a task where participants were not informed about the presence of this regularity. Using a Hebb lexical decision task, where a word (Experiments 1, 3, and 4) or a pseudoword (Experiment 2) triplet was repeated throughout a noisy stream of random words and pseudowords, we found that RTs for the unpredictable first position in the triplet decreased over repetitions (i.e., the repetition effect) but more slowly and with a different dynamic compared with items appearing at the predictable second and third positions in the repeated triplet (i.e., the predictability effect). The learning dynamic also varied as a function of triplet type (i.e., unrelated words, pseudowords, semantically related words, or idioms) and there was no evidence of a difference between items appearing at Positions 2 and 3 of the triplets. Finally, these results, supported by implicit associative learning mechanisms, were accompanied by evidence of an explicit learning of the sequence that also varied as a function of the triplet’s type.
Repetition is a key mechanism for the development of memory traces for words and sequences of words. There is much recent evidence showing that we acquire not only memory traces for words but also for multiword sequences (Arnon & Snider, 2010; Bannard & Matthews, 2008; Conklin & Carrol, 2020; Conklin & Schmitt, 2008; Janssen & Barber, 2012; Siyanova-Chanturia, Conklin, & Schmitt, 2011; Siyanova-Chanturia, Conklin, & van Heuven, 2011). The development of these memory traces may facilitate their processing and this phenomenon is now considered by several models of language acquisition (Abbot-Smith & Tomasello, 2006; Ambridge, 2020; Bannard & Lieven, 2012; McCauley & Christiansen, 2019; Perruchet & Vinter, 1998) as being central for the processing of multiword sequences.
This set of experiments provides new empirical evidence allowing to better understand the effect of repetitions on the creation of memory traces in the processing of multiword sequences and notably, to differentiate the dynamics of the repetition effect and the predictability effect. The different dynamics of these effects were notably revealed by the broken-stick regression analyses that we conducted on mean RTs overall repetitions and for all positions in the repeated triplet. A summary of the main results from these analyses is provided in Table 10.
Broken-stick regressions results: breakpoints (BP) and beta coefficients (b) of the regression lines for each position in the repeated triplet and for each experiment.
Regarding the repetition effect (indexed by the evolution of RTs on the first unpredictable position of the triplet) for words in Experiments 1, 3, and 4, it was characterised by a late breakpoint (i.e., BP1) and a relatively slow decrease in RTs indexed by a small beta coefficient (i.e., b1). The dynamic was very different for pseudowords in Experiment 2 as it produced an earlier breakpoint (5.18) and a much larger beta coefficient (–23.81). Similarly, the processing gain (indexed by the difference in mean RTs between the 45th repetition and the first occurrence of the item) was smaller for words (i.e., 90, 67, and 90 ms, for Experiments 1, 3, and 4, respectively) than for pseudowords (160 ms). These results suggest that repetition will differentially affect the processing of items that are already encoded in memory (i.e., words) compared with novel items (i.e., pseudowords). Thus, in this study, we observe that the processing of novel items benefits very rapidly from repetition and certainly from the transitory development of a memory trace representing these items.
The dynamics of the predictability effect that is indexed by the evolution of RTs on the second and third predictable positions of the triplet, was characterised, for all items, by a fast decrease in RTs with an early breakpoint (around four to seven repetitions of the triplet) for the first regression line and a large beta coefficient. The processing gain, which can be computed by subtracting the mean RTs (averaged over Positions 2 and 3) for the last occurrence of the triplet (i.e., 45th repetition) from the mean RTs obtained for the first occurrence of the same items (e.g., 640 – 347 = 293 ms, for Experiment 1), indicates that the predictability effect was much larger than the repetition effect (i.e., it was 293, 417, 285, and 332 ms, for Experiments 1 to 4, respectively). The emergence of these early breakpoints for predictable items, as well as of the difference between unpredictable and predictable items (around three to five repetitions for words), is consistent with the findings of Conklin and Carrol (2020). Indeed, they found a rapid change in participants’ reading behaviour after only four to five repetitions of the repeated pattern. These results clearly illustrate that encoding multiword sequences in memory drastically accelerates the processing of these items and that the predictability effect goes far beyond the repetition effect.
We note that an alternative interpretation to the predictability effect described above can also be provided by the multiconstituent unit (MCU) hypothesis (Zang et al., 2023), which is very close to the assumptions made in the computational model by McCauley and Christiansen (2019). According to this hypothesis, frequently encountered linguistic units consisting of more than a single word can be lexically represented in memory and identified as single representations during reading. In the model by McCauley and Christiansen (2019), this lexicalisation process is driven by the central mechanism of chunking (see also Jessop et al., 2023; Perruchet & Vinter, 1998). Therefore, multiword and pseudoword sequences that co-occur repeatedly and frequently, as in our study, may gradually become lexicalised and represented as single units in the individual’s mental lexicon. Note that several studies by Liang et al. (2015, 2017, 2021, 2023) provide empirical data in favour of this hypothesis in the field of Chinese word reading.
In addition, it is worth noting that the different learning dynamics that we observed for pseudowords can be explained not only by the development of a new memory trace, but also by the lexical decision task itself and the cognitive processes underlying it. Indeed, while in Experiment 2, participants had to give a “no” response to the triplet consisting of pseudowords, it has been shown that producing a “yes” response involves different processes than producing a “no” response. For instance, based on the interactive activation model by McClelland and Rumelhart (1981), Grainger and Jacobs (1996) propose that the generation of a “yes” response occurs when a word is recognised as a result of surpassing a certain activation threshold. In contrast, a “no” response is generated on the basis of global lexical activation, which varies as a function of the likelihood that the stimulus is a word (see also Dufau et al., 2012). Experiment 2 is therefore not comparable to the other experiments in this regard. Nevertheless, like other studies of novel words and multiword sequences using pseudowords (Norris et al., 2018; Pellicer-Sánchez, 2017; Pellicer-Sánchez et al., 2022; Szmalec et al., 2012), it allows us to study the dynamics of the development of a trace in memory and its influence, in this case, on lexical decision processes. These data may also have direct consequences for computational models of language acquisition like, for example, the Parser model (Perruchet & Vinter, 1998). In this model, each time a unit is processed again (i.e., its processing is repeated), it receives a linear increase of its memory trace (indexed by a weight value). The present results suggest that this increase may not be linear but rather non-linear depending on the weight of the item’s memory trace. For new memory traces, the increase seems to be stronger and more rapid than for memory traces that are more strongly encoded in lexical memory (called “perceptual shaper” in this model).
Similarly, and beyond the repetition effect, the repeated temporal co-occurrence of items provides a strong and non-linear processing advantage for the predictable items. Following Hebbian learning principles (Brunel & Lavigne, 2009; Endress & Johnson, 2021; Tovar et al., 2018), the coactivation of populations of neurons coding for each item may result in the strengthening of the connection weights between these two populations, leading to the creation of a chunk. Another possibility is to assume that both populations of neurons are activating a third population of pair-coding neurons (Miyashita, 2004) that would code for the pairing of these items. Irrespective of these two possible implementations, the present data suggest that these learning dynamics are non-linear, with a fast development of the memory trace of the chunk followed by a slower regime of memory consolidation.
Although the broken-stick analyses did not permit differentiation of the processing dynamics of the three types of words used in Experiments 1, 3, and 4 (i.e., unrelated words, semantically related words, and idioms, respectively), the predictability scores reported in Figure 6 indicate that the processing of idioms benefitted more from the predictability effect than the processing of semantically related words, which also benefitted more from the predictability effect than the unrelated words of Experiment 1. This is in line with previous studies showing that prior linguistic knowledge influences and facilitates regularity extraction (Elazar et al., 2022; Siegelman, Bogaerts, Elazar, et al., 2018). Pre-existing associations between words would then support the predictability effect and notably for idioms, which are sequences that are supposedly already represented and supported by memory traces.
It is difficult, however, to determine whether the advantage for idioms was mainly supported by implicit associative learning or by the participants’ prior knowledge of idioms. Table 11 provides a summary of the participants’ responses to the final questionnaire, and it clearly suggests that participants’ explicit knowledge resulted in stronger learning for idioms compared with semantically related words, which only benefitted from implicit learning. Therefore, participants’ explicit knowledge of the sequence may have interacted with implicit associative learning mechanisms and the stronger predictability score obtained for idioms may be a product of both factors.
Participants’ responses to the questionnaire expressed in percentages for each experiment.
Finally, the present data did not reveal a processing advantage for the third position over the second, contrary to previous findings on regularity extraction in naming (Rey et al., 2020) and visuomotor tasks (Minier et al., 2016; Rey et al., 2019, 2022). This is likely due to the specific time-scale of the present experimental paradigm that does not allow chunking to occur beyond two items. Indeed, for Hebbian learning to occur between the first and third items in the repeated triplet, it certainly requires maintaining the activation of the neural population coding for the first item long enough to be coactivated with the neural population coding for the last item. However, contrary to previous experimental paradigms that have reported a learning advantage on the last position of a triplet sequence, lexical decision takes a longer processing time and requires greater attentional load. Both of these factors may lead to a fast deactivation of items that were processed two steps before, avoiding any possible association to occur between Items 1 and 3 of the repeated triplets. This is consistent with recent findings suggesting that long-distance associations are harder to establish and only occur under very specific conditions (Tosatto et al., 2022; Wilson et al., 2018).
The absence of effect on the third position of the triplets may also be related to a limitation of this study. Indeed, participants may have learnt two-item associations during the task because stimulus presentation was sequential. It has been argued that parallel presentation is essential for determining the creation of co-word dependencies (Snell et al., 2018). Therefore, sequential presentation might have influenced word extraction and hindered the formation of a three-word chunk.
In addition, a number of factors are likely to have influenced the learning dynamics during the task, and thus constitute limitations to our study. First, one-third of the words forming the triplets in Experiments 1 and 3 can be considered as being part of existing multiword sequences in French (e.g., collocations: “
Second, the fact that the triplets in Experiment 4 consisted of different parts of speech (i.e., noun, verb, and adjective) compared to those in Experiments 1 and 3 (i.e., nouns) may also have influenced their processing during the task. Indeed, reaction times have been shown to differ across parts of speech (Kauschke & Stenneken, 2008; Kostić & Katz, 1987; Monaghan et al., 2003; Sereno, 1999; Tyler et al., 2001). Similarly, because the triplets were not matched in terms of MI across the experiments, certain words in some triplets are much less predictive of the following words in the sequence. This is particularly the case in Experiment 4, where the verb faire (“to do” in English) is the first word in four triplets. Hence, it may be difficult to directly compare the learning dynamics observed in Experiment 4 with those of the other experiments. Future studies that control for these confounding factors are therefore needed.
Third, given the large number of triplet repetitions (i.e., 45), this task is far from mimicking a real reading situation in which multiword sequences are widely spaced from one another and occur much less frequently. Nevertheless, the use of a well-controlled environment allowed us to characterise the acquisition of multiword sequences in real time and to investigate in depth the process of word-to-word associative learning in different linguistic settings (i.e., unrelated words, novel words using pseudowords, semantically related words and idioms). To gain a fuller picture of how multiword sequences are acquired, studies employing more ecological presentation conditions, such as those of Conklin and Carrol (2020) and Sonbul et al. (2023), and using different types and larger multiword sequences are needed.
Conclusion
This study provides novel information about the learning dynamic of multiword sequences when presented in a noisy environment, as is the case in natural language. Our data suggest that multiword learning is carried out through chunking of local information and shows how repetition affects the development of memory traces and improves processing. To further explore and understand the dynamic of multiword sequences extraction, future research could manipulate different parameters from the present experimental Hebb lexical decision task like, for example, the spacing between two repetitions of the repeated sequence or the size of the sequence, to determine the limits of the conditions under which associative learning can occur between a sequence of words.
Supplemental Material
sj-docx-1-qjp-10.1177_17470218241228548 – Supplemental material for The dynamics of multiword sequence extraction
Supplemental material, sj-docx-1-qjp-10.1177_17470218241228548 for The dynamics of multiword sequence extraction by Leonardo Pinto Arata, Laura Ordonez Magro, Carlos Ramisch, Jonathan Grainger and Arnaud Rey in Quarterly Journal of Experimental Psychology
Footnotes
Acknowledgements
We would like to thank two anonymous reviewers for their constructive feedback on previous versions of this article.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: L.P.A. was supported by a doctoral fellowship of the French Ministry of Higher Education, Research, and Innovation. This research was supported by ERC grant 742141, the Convergence Institute ILCB (ANR-16-CONV-0002), the centre for research in education Ampiric, the CHUNKED ANR project (#ANR-17-CE28-0013-02), the COMPO ANR project (#ANR-23-CE23-0031) and the HEBBIAN ANR project (#ANR-23-CE28-0008). For the purpose of Open Access, a CC-BY 4.02 public copyright licence has been applied by the authors to this document and will be applied to all subsequent versions up to the Author Accepted Article arising from this submission.
Ethical approval and informed consent
Before starting the experiment, participants accepted an online informed consent form. Ethical approval was obtained from the “Comité de Protection des Personnes SUD-EST IV” (17/051).
Data accessibility statement
Supplemental Material
The supplementary material is available at qjep.sagepub.com.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
