Abstract
To segment words in unfamiliar speech, listeners are known to exploit both native prosodic cues and statistical cues available in the speech signal. However, how and when these cues are combined remains a matter of debate. Here, we studied how transitional probabilities (TPs) and prosodic phrasal boundaries are combined by French speakers to segment words. Since French does not have lexical stress, prosodic phrasal boundaries unambiguously signal word boundaries, providing a unique possibility to test whether prosodic cues can overcome statistical ones, and constrain further statistically based segmentation. We tested French adults in an artificial speech segmentation task, manipulating the consistency between prosodic and TP cues, signaling either the same or different word boundaries. Results showed that participants favored prosodic phrasal boundaries over TPs, regardless of exposure time to the speech stream (Experiment 1: 3.5 minutes; Experiment 2: 7 min), supporting a prosodically driven statistical segmentation of the speech stream.
Keywords
1 Introduction
Segmenting an unfamiliar language—dividing continuous speech into distinct units, such as words, syllables, or phonemes—is an arduous task. Adults learning a non-native language face the challenge of identifying words in fluent speech. Unlike most written languages, which use spaces to separate words, no single and obvious cue defines word boundaries in auditory speech, and therefore listeners must use all available speech cues (i.e., segmental and suprasegmental phonological cues, as well as distributional and statistical cues) to perform segmentation and learn a language (Morgan & Demuth, 1996). When speech is familiar, listeners can effortlessly extract new word forms based on previous lexical knowledge (e.g., Mattys et al., 2005; Spinelli et al., 2010). However, when speech is unfamiliar or artificial, segmentation must depend on speech cues such as transitional probabilities (TPs; the conditional probability of Y given X in the sequence XY), prosody or phonotactics. It is therefore necessary to investigate not only the role of each cue alone, but also how multiple cues interact (Mattys et al., 2005, 2007). In this study, we explore the interaction between TPs and prosodic phrasal boundaries in word segmentation by French-speaking adults.
TPs are probably the most studied cues in speech segmentation (for reviews see Frost et al., 2019; Saffran, 2020). Adults have been shown to use TPs to segment words from fluent speech (Aslin et al., 1998; Saffran et al., 1996). Saffran and colleagues (1997) presented adult listeners with an artificial language composed of six trisyllabic words, where syllables within each word had higher TPs than syllables across word boundaries. After 21 min of exposure, participants completed a forced-choice task in which they selected the more word-like option from pairs, consisting of a word (from the artificial language) and a non-word (using the same syllables that appeared in the artificial language but in a different order). Participants chose words significantly more often than non-words, providing evidence of speech segmentation based on TPs. This has been found in speakers of other languages including English (Saffran et al., 1997), Dutch (Johnson & Tyler, 2010), and French (Bonatti et al., 2005). Hence, statistical learning has been proposed as a core mechanism to explain how adults extract words from unfamiliar speech (for a review see Saffran & Kirkham, 2018).
However, concerns have been raised regarding its ecological validity (e.g., Endress & Hauser, 2010; Johnson & Tyler, 2010; Saksida et al., 2017; Wang et al., 2023; Yang, 2004). For instance, adults show better learning from artificial strings with uniform word length compared to irregular ones (Hoch et al., 2013), suggesting that the rhythmic properties of artificial speech—typically created by concatenating words of equal length and duration—can themselves serve as segmentation cues (Wang et al., 2023).
In fact, there is evidence that prosodic cues are also relevant for word segmentation (de Diego-Balaguer et al., 2015; Marimon et al., 2022; Mattys et al., 2005; Shukla et al., 2007). Prosodic cues are typically instantiated by variations in fundamental frequency (F0 or pitch), duration (lengthening), vowel quality, and/or intensity, relative to surrounding syllables (Lehiste, 1970). At the word level, these acoustic cues signal lexical stress, but they can also indicate phrase boundaries and pragmatic functions. As such, prosodic cues can occur both within and across prosodic phrases. 1
While the acoustic correlates of stress (duration, intensity, pitch, vowel quality) appear to be shared across languages (Gordon & Roettger, 2017; see Iambic-Trochaic Law, Hayes, 1985) their realization (i.e., which and how prosodic cues are used) varies across languages. Cross-linguistic perceptual studies show that the typical stress pattern of listeners’ native language shapes how they perceive and process prosodic cues (Ordin & Nespor, 2013; Sanders & Neville, 2000; Tyler & Cutler, 2009; Vroomen et al., 1998). For instance, in artificial speech segmentation tasks, English-speaking adults, whose native language typically features first-syllable stress, interpret a pitch rise as indicating a word onset. In contrast, French-speaking adults, whose language lacks lexical stress but features final-syllable lengthening at the end of prosodic phrases, tend to interpret increased syllabic duration as a final position marker (Tyler & Cutler, 2009).
However, the transfer of native-language prosodic patterns into segmentation strategies cannot be applied straightforwardly to all languages. In languages like English or Dutch, despite a dominant strong-weak lexical stress pattern, other patterns (e.g., weak-strong) are also legal, and thus cannot be totally excluded as word candidates.
Adding to the complexity, prominence cues indicating prosodic phrasal boundaries may or may not align with lexical stress cues (Beckman, 1992), potentially introducing alternative stress positions within words (e.g., at phrase-final syllables). This makes predicting how TPs and prosodic cues are perceived and weighted during speech segmentation complex. Indeed, how prosodic cues interact with statistical cues remains unclear (Ordin & Nespor, 2013; Sohail & Johnson, 2016), particularly because most experiments using artificial languages have tested these cues in isolation (e.g., Saffran et al., 1996; Thiessen et al., 2013).
To address this, several studies have investigated whether statistical cues show any dominance over other segmentation cues. Results remain mixed. Some studies report a dominance of statistical information when both prosodic and statistical cues are present. For example, Mattys et al. (2005) showed that English diphones with low phonotactic probability are interpreted as word boundaries regardless of stress pattern, while diphones with high within-word phonotactic probabilities suppress the perception of word onsets signaled by stress cues. In contrast, other studies have claimed that prosodic cues can override TPs in English, Finnish, German, and Italian speakers (Fernandes & Ventura, 2007; Gambell & Yang, 2005; Langus et al., 2012; Marimon et al., 2022; Shukla et al., 2007; Vroomen et al., 1998). For example, Vroomen et al. (1998) showed that segmentation was most successful when the phonological properties of an artificial language matched those of the listeners’ native one—Finnish speakers benefited from vowel harmony and word-initial stress, Dutch speakers from word-initial stress, and French speakers from neither.
Other studies have reported an interaction between prosody and statistics in word segmentation (Marimon et al., 2022; Shukla et al., 2007). Shukla et al. (2007) found that Italian adults recognized statistically coherent words only when they coincided with prosodic phrase boundaries, suggesting that prosodic cues can constrain TPs processing, acting as a filter on statistical learning.
French provides an optimal study-case for exploring the role of prosodic and statistical cues in speech segmentation because it does not have lexical stress at the word level (Cutler & Mehler, 1993; Féry, 2011; Van Der Hulst & Goedemans, 2009). While stress is typically a word-level feature, in French it is assigned at the phrase level (Grammont, 1965), exhibiting fixed phrasal stress on the final syllable of prosodic phrases (Dell et al., 1984; Tranel, 1987). Thus, successive syllables in a phrase are highly similar in duration, F0 and intensity, except for the last syllable. This final-syllable stress is acoustically realized by increased duration and F0 movement (pitch rise for sentence-internal phrases, pitch fall for the sentence-final phrases; (Delattre, 1963; Jun & Fougeron, 2002; Rolland & Loevenbruck, 2002; Spinelli et al., 2010). These acoustic cues in French are known to signal boundaries and serve grammatical functions (Peperkamp et al., 2010). For example, in the sentence: “[Le rat mar

Example of prosodic phrasal stress in French, showing pitch tracking (above) and syllable duration (below) for {[Le rat marron] accentual phrase (AP)1 [voulait manger] AP2 [le long mulot.] AP3. The brown rat wanted to eat the long field mouse. Reproduced from Rolland & Loevenbruck, 2002, with permission from the authors and editors.
Since the last syllable of a prosodic phrase is highly likely also the last syllable of a word, French speakers perceive prosodic cues (i.e., F0 movement and syllable lengthening) as word-final markers. They use these cues for lexical segmentation in fluent French speech (Spinelli et al., 2010) as well as in unfamiliar or statistical languages (Bagou et al., 2002; Bagou & Frauenfelder, 2006; Marimon et al., 2025; Michelas & d’Imperio, 2010; Tremblay et al., 2017; Tyler & Cutler, 2009). French speakers are also able to segment speech based on TPs alone (Bonatti et al., 2005; Mersad & Nazzi, 2012). In short, both cues (prosodic and statistical) are available to French listeners. However, the specific way speakers process, combine, and weight these two cues as a function of their language background to segment speech remains a matter of debate. Because prosodic cues in French mostly signal phrase-final boundaries, it provides an ideal test case for exploring how cross-linguistic differences affect the relative weight given to the prosodic and statistical cues in language processing.
In this study, we evaluated whether prosodic cues constrain the exploitation of TPs, whether TPs are computed independently of prosodic cues as well as the relative weight of TP cues and prosodic phrasal boundaries for speech segmentation by native French-speaking adults. We presented an artificial language learning task in three experimental conditions. In the flat condition, only TPs between syllables marked word boundaries. In the consistent condition, both TPs and lexical stress signaled the same word boundaries. In the inconsistent condition, TPs and lexical stress signaled different word boundaries. This condition allowed us to assess the relative weight of each cue for speech segmentation. We predicted that participants would show significant segmentation in both the flat and consistent conditions (i.e., choose words indicated by TPs), potentially showing higher segmentation scores in the latter. Crucially, performance in the inconsistent condition will give us insight on whether French speakers rely more strongly on TPs (i.e., choose words signaled by TPs) or on prosodic cues (choose words signaled by lexical stress) (see Figure 2). No clear preference in this condition would suggest that TPs and prosodic phrasal boundary cues have similar weight in the segmentation process, or that segmentation was impaired by the conflict of the two cues.

Cues present in the familiarization string across conditions. Capital letters indicate prosodic cues (lengthening and pitch change).
2 Experiment 1
2.1 Method
2.1.1 Participants
Ninety-four anonymous native speakers of French (they learned it before the age of 3) participated in Experiment 1 (mean age 20.5 years, range 18–28 years). We included participants that reported under the Prolific database being native monolingual speakers of French with minimal or no daily use of any other language. Participants were also required to finish the experiment within 10 min. Additional participants (17) were excluded because they answered incorrectly on at least one catch-trial (see below for details of the catch trials). Participants were randomly assigned to the flat (n = 30), consistent (n = 34), and inconsistent condition (n = 30). Participants were recruited and paid for their participation through the online platform Prolific. Following the General Data Protection Regulation (GDPR) guidelines, the experiment was undertaken with the understanding and consent of each participant and the ethical approval of the local ethical committee (number agreement: IRB00010290-2018-10-16-52).
2.1.2 Stimuli
Artificial Language: The artificial language consisted of 4 trisyllabic words, /gudola/, /logabi/, /batipu/, and /piduto/, concatenated pseudo randomly (all equally probable but never twice in a row) and without pauses for 3:25 min. To remove on/offset cues, we ramped up/down the initial/final 6 s of the stream. This stream was created with a text-to-speech synthesizer based on diphones from a native French female speaker as basic units and TD-PSOLA as a vocoder (Bailly et al., 1991). It did not contain any prosodic cue (flat condition; 235 Hz; consonant duration: 130 ms; vowel duration: 156 ms). We then created the two prosodic streams by modifying acoustic cues: increasing pitch (313 Hz) and duration (372 ms) of either the last syllable of each word (consistent) or the second syllable (inconsistent). The F0 contour on the prominent syllable was realized via a linear interpolation of three points per vowel (position at 20%, 50% and 80% of their durations; (Tournemire, 1994). The F0 is constant over 20%–80% of each vowel and accents are shaped as hat patterns (F0 baseline at 130 Hz and accent at 175 Hz). The two streams with modified acoustic cues lasted 3 min and 47 sec. Each stream (flat, consistent, and inconsistent) was created in four different versions to counterbalance the word order. Statistical words were defined by a TP of 1 between both syllabic transitions (e.g., between /gu/-/do/, and /do/-/la/ in /gudola/) and appeared 60 times in each speech stream. Coarticulation occurred between phonemes within a syllable but not between syllables, to ensure only statistical cues were provided. Pitch and duration values for all conditions were based on Rolland and Loevenbruck (2002).
Test pairs of the test phase. In all three conditions, test pairs consisted of isolated versions without prosodic information of the four statistical words /gudola/, /logabi/, /batipu/, and /piduto/, and four part-words, /lapidu/, /bigudo/, /puloga/, and /tobati/ constructed by concatenating the last syllable of a word with the first two syllables of another. Part-words were defined by a TP of 0.33 between the first and second syllable (between /bi/-/gu/ in /bigudo/), and a TP of 1.0 between the second and third (e.g., between /gu/-/do/ in /bigudo/), appearing 20 times per speech stream. We chose part-words that did not include phoneme repetitions. If participants rely more strongly on TPs, they should segment statistical words as defined as syllable pairs with a TP of 1.0 in both the flat and consistent condition. In contrast, if participants follow prosodic cues, they should segment the inconsistent condition into words that follow the French phrasal stress pattern (weak-weak-strong), which straddle the statistical word boundaries (Figure 2).
We created 16 test pairs, consisting of one of the four words followed by one of the four part-words or vice versa, separated by a 500 ms silence (4 appearances per word). The order was counterbalanced.
Image Stream. An image stream was presented during audio familiarization. The image stream was added to ensure participants’ attention on the task. We chose four distinguishable images (Moreno-Martínez & Montoro, 2012): two fruits (custard apple and pomegranate) and two animals (armadillo and tapir), with infrequent names in French, to prevent participants from using inner speech. These pictures were then rotated 30° clockwise and counterclockwise, resulting in eight pictures. A single pseudorandom order of presentation was generated with R, ensuring nearly equal repetition of each picture and orientation (range: 51–57 repetitions/image). Pictures were presented for 250 ms, with a white picture lasting 250 ms in-between (ISI 250 ms, SOA 500 ms). Image presentation rate was 4 Hz, misaligned with the syllabic and word/part-word boundaries, to ensure that it did not interfere with the segmentation process. We did not expect any interference of the image stream in our results because there is evidence that participants can learn the regularities in an auditory stream implicitly while watching a similar image stream (Toro et al., 2005). In addition, it was presented across all conditions.
Videos: The experiment was originally intended to be conducted in a laboratory setting, but due to the COVID-19 outbreak, we adapted it as an online questionnaire. To do so, we created two silent videos by running the laboratory-based experiment on OpenSesame and recording the video with OBS studio, for 3:26 s (flat condition) and for 3:48 s (consistent and inconsistent conditions). Then we added the audio streams, creating 12 videos (3 conditions with 4 orders each). The videos were presented online using Qualtrics.
2.1.3 Procedure
Participants first read a consent form and agreed to participate in the experiment. They were informed that they could leave the experiment at any time and that their data would be anonymous. They then indicated their age and whether French was their native language. In the learning phase, participants were randomly assigned to watch one of the 12 videos. They were informed that they were going to listen to a made-up language and watch an image stream but that there would only be questions about the language afterwards. They were asked to watch the video with minimal external disturbances. The test phase consisted of 20 forced-choice trials, presented in random order. Sixteen test trials consisted of presenting the word versus part-word pairs auditorily. Participants were then asked to click on a button labeled “sequence 1” or “sequence 2” depending on which auditory sequence they thought sounded the most like a word of the language they had heard. The 4 remaining test trials were catch trials designed to ensure participants performed the task attentively. Two of these were visually identical to the experimental trials (i.e., the same on-screen buttons and a sentence asking them to play the audio), but instead of audio sequences, the file contained a recorded voice instructing participants to select either the first (“sequence 1”) or second (“sequence 2”) option, and to disregard the usual task instructions (i.e., comparing the two sequences). These auditory catch trials used the same response buttons as the experimental trials and allowed us to assess whether participants were listening attentively to the audio files during the test phase. An incorrect response (e.g., clicking “sequence 2” when the voice said “sequence 1”) was interpreted as an indication of inattention.
The two other catch trials ensured that participants had paid attention during the familiarization phase and had not been distracted or engaged in unrelated activities—something not directly controllable on online platforms such as Qualtrics. Because participants’ responses to auditory stimuli during familiarization could reflect either a failure to segment or a failure to perform the task, we could not reliably assess auditory attention at that stage. We therefore chose to evaluate attention using visual content from the image stream shown during familiarization. These trials, also presented during the test phase, included very simple written control questions (e.g., whether any image had been repeated, or whether they saw an elephant at any point during the image stream).
Participants responded using dedicated “yes”/“no” on-screen buttons, which were only displayed during these specific catch trials. An incorrect response—such as clicking “no” when asked whether any image had been repeated (each image appeared over 50 times), or clicking “yes” when asked whether they had seen an elephant (only two fruits—custard apple and pomegranate—and two animals—armadillo and tapir—were shown)—was taken as an indication of inattention during the familiarization phase.
2.1.4 Data analysis
Participants’ answers of the test pairs (i.e., accuracy) were analyzed using generalized linear mixed models (GLMM), with the R software and the lme4 package (Bates et al., 2015). Post hoc forward model comparisons (i.e., likelihood ratio chi-square tests) were used to compare the base model fit to one including the Condition factor (flat, consistent, inconsistent), to assess whether it significantly improved the model. Next, each of the five hypotheses was investigated with contrasts, using the multcomp package (Hothorn et al., 2016). Mean and confidence intervals were estimated using bootstrapping. The data, stimuli, and code can be found at https://osf.io/426jb/?view_only=acab8f88bda64c34acfc741edbe39b7c. Experiment 1 was preregistered at (http://osf.io/7r4ya).
2.2 Results and discussion
Results from the post hoc likelihood model comparison confirmed that the model including condition fitted best, coded in R as glmer, accuracy ~ condition + (1|item) + (1|participant), family =“binomial.” The mixed-model predicted participant choices well (receiver operating characteristic (ROC) area under curve = 0.80). Results are presented in Figure 3(a) and in Tables 1 and 2. Surprisingly, performance (i.e., percentage of choices for statistical words over part-words) in the flat condition did not significantly differ from chance (chance at 0.5; β = 0.08, z = 0.33, p > .1). Instead, performance from the consistent and inconsistent condition did differ from chance, with participants choosing part-words over statistical words in the inconsistent condition (β = −0.90, z = −4.05, p < .01), and statistical words over part-words in the consistent condition (β = 0.63, z = 2.89, p < .01). In line with our hypothesis, choices of statistical words in the consistent condition were significantly higher than in the flat condition (β = 0.62, z = 2.89, p < .01), and choices of statistical words were significantly lower in the inconsistent condition than in the flat condition (β = −0.91, z = −4.05, p < .01).

Mean performance (proportion of choices for statistical words over part-words) in (a) Experiment 1 (exposure duration 3.5 min) and (b) Experiment 2 (exposure duration 7 min). Error bars represent 95% confidence intervals. Each point represents the mean performance of one participant. The dotted line represents 50% chance level.
Post Hoc Likelihood Ratio (Chi-Square Test) Forward Model Comparisons’ Statistics of Experiment 1.
Note. Signif. codes: 0 “***” 0.001 “**” 0.01 “*” 0.05 “.” 0.1 “ ” 1.
Fixed Effects’ Results for the Generalized Linear Mixed Model (GLMM) of Experiment 1.
Note. Signif. codes: 0 “***” 0.001 “**” 0.01 “*” 0.05 “.” 0.1 “ ” 1.
In Experiment 1, we present evidence that French speakers use prosodic cues for speech segmentation of an artificial language, regardless of their consistency with the statistical ones. As predicted, based on the rhythmic properties of the French language, participants processed the syllabic stress cues (duration and pitch) as an indication of a word boundary, and extracted the words according to the prosodic cues (typical final phrasal stress in French), regardless of whether statistical cues were signaling the same (consistent) or different (inconsistent) word boundaries. Interestingly, however, we found no evidence that participants at the group level used the statistical cues in the absence of prosody (i.e., flat condition).
3 Experiment 2
Experiment 1 showed that after 3:25 min of exposure to a new and artificial language, French speakers significantly segmented words when both statistical and prosodic cues were present, but not when only statistical cues were presented. One likely explanation for this lack of segmentation in the statistical-only condition is the duration of the familiarization. While phrasal prosody does not require learning during the experiment—due to previous learning through language-specific experience with French—statistical regularities require learning over time.
The choice of a short exposure duration in Experiment 1 was motivated by infant studies reporting successful statistical learning already after just 2–3 min (e.g., Johnson & Jusczyk, 2001; Johnson & Tyler, 2010; Saffran et al., 1996). We therefore aimed to test whether the same amount of information would suffice in adults’ speech segmentation. Given the lack of significant statistical-only segmentation, Experiment 2 explores whether a longer exposure duration (i.e., twice as long, 7 min) enables successful word segmentation based solely on statistical cues (flat condition), and, in turn, allows for a more robust evaluation of the interaction between statistical and prosodic cues in French speakers (consistent and inconsistent conditions).
3.1 Method
3.1.1 Participants
A total of 94 anonymous native speakers of French took part in Experiment 2 (mean age 20.2 years, range 18–27 years). As in Experiment 1, we included participants that reported under the Prolific database being native monolingual speakers of French and having little to no daily use of any other language. Participants were also required to finish the experiment in 15 min. Four additional participants were excluded from the statistical analyses because they answered at least one catch-trial incorrectly. Participants were randomly assigned to the flat (n = 30), consistent (n = 32) and inconsistent (n = 32) conditions. Participants were all native speakers of French. The experiment was conducted with ethical approval and with the informed consent of each participant, as in Experiment 1.
3.1.2 Stimuli, procedure, and data analysis
The stimuli, procedure, and data analysis were the same as in Experiment 1, except that all videos were twice as long. Video durations were 6 min and 50 s (instead of 3 min 25 s), and each word appeared exactly 120 times (instead of 60 times), and each part-word appeared 40 times (instead of 20).
3.2 Results and discussion
Results from the Post hoc likelihood model comparison confirmed that the model including Condition fitted best, coded in R as glmer, accuracy ~ condition + (1|item) + (1|participant), family =“binomial.” The mixed model predicted participant choices well (ROC area under curve = 0.86). Experiment 2 results are presented in Figure 3 (b) and in Tables 3 and 4. In contrast to Experiment 1, performance of the flat condition was above chance (β = 0.54, z = 2.16, p = .02), which reflects a successful use of statistical cues to segment words with a longer exposure time (7 min). Then, as found in Experiment 1, participants chose part-words significantly over statistical words in the inconsistent condition (β = −1.30, z = −5.20, p < .01) and statistical words in the consistent condition (β = 1.72, z = 6.62, p < .01). The model contrasts between conditions showed that choices of statistical words were higher in the consistent condition than in the flat condition (β = 1.18, z = 4.47, p < .01), and reasonably, lower in the inconsistent than in the flat condition (β = −1.84, z = .7.10, p < .01).
Post Hoc Likelihood Ratio (Chi-Square Test) Forward Model Comparisons’ Statistics of Experiment 2.
Note. Signif. codes: 0 “***” 0.001 “**” 0.01 “*” 0.05 “.” 0.1 “ ” 1.
Fixed-Effects Results for the Generalized Linear Mixed Model (GLMM) of Experiment 2.
Note. Signif. codes: 0 “***” 0.001 “**” 0.01 “*” 0.05 “.” 0.1 “ ” 1.
Last, to evaluate the effect of exposure time, we ran another GLMM model with data from Experiment 1 and 2 together, using condition (flat, consistent, inconsistent), experiment length (short, long) and their interaction, including random intercepts for Items, and Participants nested within experiment. Forward likelihood model comparisons confirmed that the full model, including the interaction between condition and experiment length, fitted the best (see Table 5) coded as: glmer, accuracy ~ condition * experiment + (1|item) + (1| participant:experiment), family =“binomial.” Table 6 summarizes the results of the model (ROC = 0.83). Because the interaction was significant (β = −0.94, z = −2.81, p < .01), we ran three additional mixed models, one per condition, coded as: glmer, accuracy ~ experiment + (1|item) + (1|participant), family =“binomial.” Results of the separate models (see Tables 7–9) confirmed that choices of statistical words over part-words (marked by prosodic cues) significantly increased with longer exposure time in the flat condition (β = 0.42, z = 2.33, p = .02) and the consistent condition (β = 1.09, z = 3.58, p < .01), but they significantly decreased in the inconsistent condition (β = −0.50, z = −2.15, p = .03), reflecting more frequent choices of part-words.
Post Hoc Likelihood Ratio (Chi-Square Test) Forward Model Comparisons’ Statistics of Experiment 1 and 2.
Note. Signif. codes: 0 “***” 0.001 “**” 0.01 “*” 0.05 “.” 0.1 “ ” 1.
Fixed-Effects Results for the Generalized Linear Mixed Model (GLMM) of Experiment 1 and 2.
Note. Signif. codes: 0 “***” 0.001 “**” 0.01 “*” 0.05 “.” 0.1 “ ” 1.
Fixed Effects Results for the Generalized Linear Mixed Model (GLMM) of the Flat Condition.
Note. Signif. codes: 0 “***” 0.001 “**” 0.01 “*” 0.05 “.” 0.1 “ ” 1.
Fixed Effects Results for the Generalized Linear Mixed Model (GLMM) of the Consistent Condition.
Note. Signif. codes: 0 “***” 0.001 “**” 0.01 “*” 0.05 “.” 0.1 “ ” 1.
Fixed Effects Results for the Generalized Linear Mixed Model (GLMM) of the Inconsistent Condition.
Note. Signif. codes: 0 “***” 0.001 “**” 0.01 “*” 0.05 “.” 0.1 “ ” 1.
Experiment 2 results suggest that a longer exposure time allowed participants to extract both TPs and prosodic cues from the speech stream, but that, as previously found in Experiment 1, when pitted against each other in the inconsistent condition, the prosodic cues had a stronger weight than the statistical ones, and participants continued to extract the prosodically marked words regardless of the statistical cues. Last, the statistical comparison of Experiment 1 and 2 showed that the effect sizes increased for the three conditions: That is, when exposure time was doubled, participants showed higher statistical segmentation scores for the flat and consistent conditions, and higher prosodic segmentation scores in the inconsistent condition.
3.3. General discussion
This study evaluates the role of TPs and prosodic cues in speech segmentation in French. French provides an optimal study-case for pitting TPs and prosodic cues against each other because it has no lexical stress but modulates syllable pitch and duration at the phrasal level, often indicating phrase-final (and thus word-final) boundaries. To test this question, we presented participants with an artificial language speech segmentation task in three different conditions: flat (TPs only), consistent (last-syllable stress, coinciding with French final phrasal boundary cues and aligned with TPs), and inconsistent (middle-syllable stress, where final phrasal boundary cues are misaligned with TPs). Results of Experiment 1 showed that, after 3.5 min of exposure, participants successfully segmented words according to the prosodic cues (consistent and inconsistent conditions), but failed to segment using TPs alone (flat condition). In Experiment 2, exposure time was doubled (7 min), and participants successfully segmented statistical words using TPs only (flat condition) and kept segmenting prosodically marked words in both consistent and inconsistent conditions.
Statistical cues have been proposed as a core mechanism for word segmentation, but their applicability in real-world language learning has been questioned (Johnson & Tyler, 2010; Smith et al., 2014; Wang et al., 2023). Notably, these studies have raised the concern that the input in typical experiments is highly simplified, inherently rhythmic and generally presents more word repetitions than natural speech (Erickson & Thiessen, 2015). Our results from the flat conditions offer new insights into this issue: evidence of statistical-only segmentation was not found in Experiment 1 with 60 word repetitions, and only emerged after 120 word repetitions (7 min) in Experiment 2, at a modest ratio of 61%. While recent studies using online measures have shown that adults can detect statistical patterns after just 2–4 exposures (e.g., Batterink, 2017; Hao Wang et al., 2024) our findings—based on a force-choice task—suggest that more extensive exposure (here 120 repetitions) may be required for such early sensitivity to translate into robust segmentation. Our results align with previous work using similar forced-choice paradigms, which also report modest or variable accuracy even with longer exposures—for example, Saffran et al. (1997) reported 71% accuracy after 21 min of exposure and 300 word repetitions, while Schön et al. (2008) found no evidence of segmentation with 84 repetitions, and generally conveying a great amount of individual variability (Siegelman et al., 2017). Studies with French-speaking adults show similar results, that is, significant statistical segmentation after 100–150 repetitions, and 7–19 min of exposure (Bagou et al., 2002; Bagou & Frauenfelder, 2006; Toro et al., 2009; Tyler & Cutler, 2009). Notably, even with consistent prosodic cues, performance rarely reaches ceiling (e.g., Bagou & Frauenfelder, 2006). This reinforces the view that statistical cues alone may not be sufficient for robust segmentation of unfamiliar speech and that prosodic cues play a dominant role in successful speech segmentation.
Indeed, our results from the consistent and inconsistent conditions align with previous speech segmentation studies in French population (Bagou et al., 2002; Bagou & Frauenfelder, 2006; Marimon et al., 2025; Tremblay et al., 2017; Tyler & Cutler, 2009) and confirm that French speakers use prosodic cues, specifically final syllable lengthening and pitch rise, to signal word boundaries and achieve successful segmentation. Specifically, the inconsistent conditions in Experiment 1 and 2 reveal a strong preference for prosodic cues even when they conflict with statistical ones. While prior studies had shown at-chance performance when the two cues predicted different word boundaries (e.g., Toro et al., 2009), our results in French-speaking adults demonstrate that participants can reach segmentation based on prosodic cues and in spite of statistically based predictions, as also shown by Marimon et al. (2025). These findings support the view that prosodic phrasal boundary cues constrain the exploitation of TPs. In addition, the effect size for part-words in the consistent and inconsistent conditions increased with exposure (Experiment 2). In the consistent condition, the greater segmentation scores with longer exposure could reflect either enhanced extraction of prosodic cues alone or the combined influence of prosodic and statistical cues. The segmentation performance in Experiment 2 (with longer exposure), suggests that such potential temporal statistical processing was swiftly overcome by the prosodic cues. Finally, the differing exposure durations required for TP and prosodic segmentation is also interesting considering other statistical segmentation processes. While TPs required 7 min or 120 repetitions to elicit segmentation, prosodic cues required half that time, even when inconsistent with TPs. This supports the claim that TPs need more time to be computed, to then be used for segmentation (Endress & Bonatti, 2007).
One possible explanation for these results is the acoustic salience of the stimuli used. While prior studies manipulated F0 change following the typical modulation of lexical stress—for example, 20–50 Hz difference in Toro-Soto et al., 2007, and Tyler and Cutler, 2009—the present experiment used typical phrasal-boundary stress—that is, 80 Hz Fo shift. Coupled with participants’ extensive experience with phrase-final stress cues, this difference could have increased the weight of prosodic cues for segmentation, and reduced reliance on TPs. It is also possible that the visual image stream in our paradigm reduced attentional resources for TP computation and or exploitation, thereby extending the necessary exposure time.
Nevertheless, this long learning phase for word segmentation seems to contrast with other statistical learning processes such as visual object categorization, which elicit significant learning after only three or four presentations (Perfors & Tenenbaum, 2009; Turk-Browne et al., 2005), as well as with computational accounts that successfully predict categorization after one or few examples (Kemp et al., 2007; Salakhutdinov et al., 2012). A potential interpretation of these differences is that statistical learning in the temporal domain—which is crucial for language learning (de Diego-Balaguer et al., 2016)—may be generally more arduous than in the visual domain (see also Ordin et al., 2021, for more discussion about this topic). However, at this stage, this interpretation remains entirely speculative. For instance, other disparities pertaining to the types of objects used and the task at hand (e.g., segmentation of auditory pseudo-words vs. categorization of artificial visual objects) could also explain these different results, rather than their modality of presentation.
Although we cannot entirely discard that in inconsistent conditions of Experiment 1 and 2, words compatible with TPs were temporarily stored before being discarded when shown incompatible with prosodic cues (Batterink & Paller, 2019), these results provide an interesting addition to the debate on the mechanisms of information processing and integration that lead to speech segmentation. While some models propose that prosodic and statistical cues are first computed in separate spaces, in parallel, and only thereafter merged to provide coherent predictions (e.g., Shukla, 2006), our results rather suggest that prosodic cues help to perform an initial chunking of the speech stream and therefore constraining the space domain where potential word candidates may be computed. The fact that when cues were inconsistent, prosodic cues led to segmentation supports the idea that statistical cues are not being computed independently of prosodic ones, but within the internal sequence or boundary points previously defined by the prosodic contour, in this case, impeding their use for segmentation.
Our results also align with the larger claim that statistical computations in adults are subsequent to an attention selection process (Palmer & Mattys, 2016; Smith et al., 2014, 2018; Toro et al., 2005). These studies show that segmentation is impeded under divided attention conditions—for example, performing a concurrent visual task while hearing the speech stream—and propose that attention filters the input, and statistics are only processed within the attended part. Even though other studies using online electroencephalography (EEG) recordings have shown that some levels of neural entrainment to statistical regularities can also occur outside the focus of attention, these seem to remain in the perceptual level and are not processed further for segmentation (Batterink & Paller, 2019; Benjamin et al., 2023). In our case, native-language prosody likely modulated attention allocation, gating the processing of the statistical cues.
In conclusion, the present results shed new light on adult speech segmentation by showing that regardless of exposure time, when TPs and prosodic cues are pitted against each other, French-speaking participants favor prosodic cues over statistical ones. These results provide further evidence supporting an early attentional chunking of the speech information based on prosodic cues. We highlight that other speech cues and mechanisms beyond statistical ones—such as prosody—need to be considered in speech segmentation and language acquisition.
Supplemental Material
sj-docx-1-las-10.1177_00238309251374295 – Supplemental material for French Speakers Prefer Prosody Over Statistics to Segment Speech
Supplemental material, sj-docx-1-las-10.1177_00238309251374295 for French Speakers Prefer Prosody Over Statistics to Segment Speech by Joan Birulés, Mireia Marimon, Alexandre Duroyal, Anne Vilain, Gérard Bailly and Mathilde Fort in Language and Speech
Footnotes
Acknowledgements
Our team is grateful to Marion Dohen and Hélène Lœvenbruck for taking the time to give us advice about the artificial language, to Silvain Gerber for checking our statistics, and to Elsa Spinelli for giving us feedback on the experiment. The authors thank the Screen-MSH-Aples platform for providing access to the Qualtrics software.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research has been supported by the NeuroCog, the Pole Grenoble Cognition, the SFR Santé Société, the Fondation de France (grant no.: 00112028), the Agence Nationale de la Recherche (grant no.: ANR-22-CE28-0004-01), and by Horizon Europe (grant no.:101108884).
Supplemental material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
