Abstract
This study examines how French listeners segment and learn new words of artificial languages varying in the presence of different combinations of sublexical segmentation cues. The first experiment investigated the contribution of three different types of sublexical cues (acoustic-phonetic, phonological and prosodic cues) to word learning. The second experiment explored how participants specifically exploited sublexical prosodic cues. Whereas complementary cues signaling word-initial and word-final boundaries had synergistic effects on word learning in the first experiment, the two manipulated prosodic cues redundantly signaling word-final boundaries in the second experiment were rank-ordered with final pitch variations being more weighted than final lengthening. These results are discussed in light of the notions of cue type, cue position and cue efficiency.
Keywords
1 Introduction
Lexical segmentation is the process of locating the boundaries separating words in continuous speech. It is a crucial mental operation for infants and adults in both word recognition and word learning that is based on multiple sublexical segmentation cues and on lexical information 1 . Lexically-based segmentation in word recognition has often been demonstrated (for a review, see Davis, 2000), and some studies have even revealed its primacy over sublexical segmentation cues (Mattys, White, & Melhorn, 2005). However, in word learning, listeners are confronted with a continuous speech stream with no or limited knowledge about words in the language they are hearing. Here segmentation consists in locating the potential boundaries of unknown lexical units that are to be stored in the mental lexicon. Since lexical knowledge is initially absent and only progressively constructed, segmentation at this early stage of learning cannot be driven lexically 2 . The characterisation of this segmentation process thus involves determining the role of sublexical cues and the progressive involvement of lexical information. At least four types of sublexical cues have been investigated in word learning experiments that include statistical information, phonotactic regularities (Finn & Hudson Kam, 2008), fine-grained acoustic-phonetic regularities (Fernandes, Ventura, & Kolinsky, 2007; Kim, Cho, & McQueen, 2012) and prosodic regularities, particularly rhythmic regularities (e.g., Saffran, Newport, & Aslin, 1996).
Firstly, learners have been shown to compute statistics between adjacent syllables. In the original study of Saffran et al. (1996), adult participants were presented with an artificial continuous stretch of speech in which “ words” were identifiable thanks to the high transitional probabilities (TPs) between their syllables but with no pauses or any other information about lexical boundaries. After a relatively short exposure to the new language, offline forced-choice tasks revealed that participants were able to distinguish words from non-word items. Adults were thus able to acquire artificial words based solely on the frequency of co-occurrence of adjacent syllables (Aslin, Saffran, & Newport, 1998; Saffran, Aslin, & Newport1996; Saffran, Newport, Aslin, Tunick, & Barrueco, 1997).
Secondly, many studies have shown that adults (McQueen, 1998) and infants (Jusczyk & Luce, 1994; Mattys, Jusczyk, Luce, & Morgan, 1999; Mattys & Jusczyk, 2001) make efficient use of familiar phonotactic cues to segment and extract words from fluent speech either in word recognition or learning. Phonotactics define restrictions on the ordering of phonemes within words and syllables. Since word boundaries generally coincide with syllable boundaries, a sequence of consonants that cannot occur within the same syllable in a particular language can be helpful in detecting word boundaries. When French adults encounter a sequence of consonants such as tf (e.g., porte fermée /
/, closed door) that do not occur word-initially or word-finally in French (i.e., consonant clusters with TPs of zero), they should be able to infer a word boundary separating these two consonants. Some artificial language (AL) learning studies have clearly shown that both adults (Finn & Hudson Kam, 2008; Weber & Cutler, 2006) and infants (Thiessen & Saffran, 2003) use their native phonotactic knowledge to extract and acquire words at initial exposure to a new language. For example, Finn and Hudson Kam (2008) showed that adult learners’ knowledge about phonotactic restrictions in their first language is strong enough to prevent them from segmenting words during the initial exposure to an AL that possesses diverging between-syllable statistics. However, unlike infants, adults are also able to adjust their own learning and rapidly acquire non-familiar phonotactic regularities when words are pre-segmented during exposure (Onishi, Chambers, & Fisher, 2002) or when the exposure to the new language is substantially increased (Finn & Hudson Kam, 2008).
Thirdly, learners also rely on a variety of subtle acoustic-phonetic word boundary cues to extract words from continuous speech. Studies have suggested that there is more coarticulation within than between words (Byrd & Saltzman, 1998) and more specifically, that word onsets carry fine acoustic-phonetic information that provides information about the likely locations of word-initial boundaries (Christiansen, Allen, & Seidenberg, 1998; Nakatani & Dukes, 1977; Quené, 1992; Shatzman & McQueen, 2006). The efficiency of these subtle acoustic cues has been largely attested in word recognition studies with adults (Shatzman & McQueen, 2006) and infants from the age of 10.5 months (Jusczyk, Hohne, & Bauman, 1999). Spinelli, McQueen, and Cutler (2003) showed that two homophonous sequences such as “dernier rognon” (last kidney) and “dernier oignon” (last onion) are discriminated by participants thanks to the duration of the critical consonant /
Finally, prosodic regularities also help listeners to segment words in continuous speech. High-level prosodic cues (phrasal prosody) have been shown to be relevant for segmenting and grouping speech into constituents at different levels of the prosodic hierarchy in many languages (e.g., Beckman & Pierrehumbert, 1986; Nespor & Vogel, 1986; Selkirk, 1984), from the phonological phrase (Millotte, René, Wales, & Christophe, 2008; Shukla, Nespor, & Mehler, 2007) up to the utterance (Langus, Marchetto, Bion, & Nespor, 2012). Many studies have also confirmed the Metrical Segmentation Strategy, initially developed for English by Cutler and Norris (1988), and applied it to other languages (for a review, see Cutler, Dahan, & van Donselaar, 1997; see also Cutler & Otake, 1994 in Japanese; Sebastián-Gallés & Costa, 1997 in Spanish; Vroomen, Tuomainen, & De Gelder, 1998 in Finnish and Dutch). These studies clearly established that prominent syllables guide lexical segmentation according to the stress pattern of the listeners’ native language. For example, English listeners generally segment correctly if they consider strong syllables as word onsets because most English lexical words begin with strong syllables (Cutler & Carter, 1987). AL studies investigating the integration of statistical and prosodic cues revealed that the use of strong syllables facilitates segmentation at least for languages where stress is aligned either with the beginning (e.g., English or Dutch) (Tyler & Cutler, 2009) or with the end of words (e.g., French) (Bagou, Fougeron, & Frauenfelder, 2002; Vroomen et al., 1998; Saffran et al., 1996).
As we have seen in this brief review of previous research, listeners and learners can exploit a wide range of sublexical cues. However, natural speech is complex: no segmentation cue type is available at every word boundary, and more than one cue is generally present at any given word boundary. Studying how each type of information impacts segmentation separately is not sufficient since cues are integrated with complex interactions between them. Unfortunately few studies have investigated how multiple language-specific segmentation cues interact in word recognition (but see Mattys et al., 2005), and even less in word learning.
Establishing the relative contribution of these different sublexical segmentation cues is a complicated enterprise with several methodological options. Most studies have involved pitting two cues against each other to test their relative primacy (Fernandes et al., 2007; Mattys et al., 2005). However, in attempting to develop a more complete view of lexical segmentation, we have developed experiments that include experimental conditions varying along the following four dimensions: (1) Cue convergence, having convergent rather than divergent cues; (2) Cue number, having not only two cues, but more than two cues; (3) Cue position, having cues that are both located word-initially or word-finally and; (4) Cue redundancy, having multiple cues that are either redundant in their positions (cues in the same location) or complementary (cues in two different positions).
Firstly, most studies of speech segmentation have provided listeners with opposing segmentation cues that point to different boundary locations. However, in natural speech, multiple sublexical cues rarely guide listeners towards different segmentations but generally converge on the same segmentation. Although pitting two cues pairwise against each other allows the researcher to establish whether a specific cue can overrule another cue under particular conditions, this approach has limited ecological validity and is not sufficient to establish the precise nature of the interaction between cues. Cues generally converge on a single segmentation in natural speech and can combine either in an additive or in a synergistic (or super-additive) manner. In the former case, the combined effect of multiple cues would be equal to the sum of the effect of each cue taken separately, whereas in the latter case the combined effect would be greater than the sum of their separate effects. Here, the availability of two sources of segmentation information modifies the individual contribution of each separate cue.
Secondly, most research has limited to two the number of cues being evaluated simultaneously in order to keep the experimental design manageable. However, word boundaries are usually conveyed by multiple cues, and the interactions between these cues are likely to be complex. Consider three convergent cues, say A, B and C, that occur together. Although A may well take precedence over B when these are the only two available segmentation cues, this relation may change when C is also present, particularly if B and C have synergistic effects. Indeed, even if A has a greater effect than B when only two cues are provided, the effect of B might be improved by the availability of C, and A may be relegated to a lower status. To our knowledge, the way that more than two convergent cues are integrated has never been addressed.
Thirdly, the precise nature of segmentation cues has not always been clearly specified, in part due to a terminological ambiguity relating to their position or function. It is necessary to distinguish between where the segmentation cue is located (word-initially or word-finally) and which boundary it signals (word onset or word offset). Cues located word initially can obviously signal the onset of that word, but also the offset of the preceding word. Further, word-final cues signal the offset of the word that contains them, but also the onset of the following word. In order to avoid this ambiguity, we will distinguish between word-initial and word-final cue for their position and word onset and word offset cues for what they signal, that is, their function.
Several studies have examined the relative efficiency of word-initial and word-final cues for segmenting the onset or the offset of the same word (Content, Dumay, & Frauenfelder, 2000). It was found that word-initial cues were more powerful than word-final cues, probably in large part because they came earlier. However, few, if any studies, have examined the informational value of segmentation cues (word-initial versus word-final of the preceding word) which serve to signal the same word-onset boundary.
Fourthly, it was of interest to compare conditions in which the multiple segmentation cues are located in the same or in different positions. To this end, we contrast segmentation in redundant conditions in which two cues in the same position (only word-initial or word final) function to signal the same word onset with complementary conditions in which the cues are in two different positions (word-initial and word-final). As previous results suggested (Mattys et al., 2005), listeners appear not to need all of the redundant cues to make a boundary decision. Indeed, when segmental (phonotactic and acoustic) and stress cues both signalled the same word-initial boundary in English, cues were found to be rank-ordered (Mattys et al., 2005) and the most informative (here, segmental) of the redundant cues was used by the participants. We would thus expect that the use of redundant cues signalling the same position should lead to a trading relation among cues (at least under good hearing conditions).
Unlike English, French allows the investigation of how both redundant and complementary cues interact and how their relative contribution varies according to their position and distribution. Indeed, primary accent is located on the last syllable of the word, whereas segmental cues are efficient for detecting word-initial boundaries (Dumay, Frauenfelder, & Content, 2002). In this case, word-initial (segmental) and word-final (prosodic) boundary cues are complementary and their combined effect could well be synergistic. Moreover, durational and pitch cues redundantly mark word-final boundaries since final lengthening, and final pitch rise are two acoustic cues that mainly make up prosodic information in French. Prosodic segmentation cues are thus redundant and should be rank-ordered.
In summary, the present study uses artificial language learning experiments (ALL) in which adult learners will be presented with languages containing convergent, multiple sublexical segmentation cues that vary in their nature, position and redundancy.
In the first experiment, we examine cue complementarity by investigating how various combinations of word-initial segmental cues (fine-grained acoustic-phonetic and phonotactic cues) and word-final prosodic cues influence word learning. In the second experiment, we examine cue redundancy by focusing on how multiple sublexical cues within one of these cue types, namely prosodic cues, contribute to the marking of word-final boundaries. We specifically investigate how multiple acoustic-prosodic cues (final lengthening and final pitch rise) simultaneously contribute to the identification of lexical boundaries.
In both experiments, French participants listened to continuous sequences of consonant-vowel-consonant (CVC) syllables, composed of recurring two- and three-syllable artificial words. Unlike our study, most AL studies have used artificial words of the same length (only trisyllabics). However, the experimental introduction of a prosodic prominence on every third syllable in our AL would have created an isochronous rhythm that could have dominated most other cues and exaggerated the influence of the prosodic cue under scrutiny. Therefore, the artificial words varied in length in the languages used in this study. Using the ALL paradigm allow us to: (1) eliminate semantic and lexical information and test adults acquiring a language about which they have no linguistic knowledge; and (2) control the participant’s exposure to the language and manipulate finely the linguistic and acoustic properties of the artificial words.
2 Experiment 1
2.1 Method
2.1.1 Participants
One hundred and forty-five native speakers of French (14 men and 131 women), ranging in age from 19 to 39 years old (mean (M) = 21.5; standard deviation (SD) = 2.5), voluntarily participated in the experiment for course credits. All were undergraduate or graduate students at the University of Geneva and none reported a history of speech or hearing difficulties. Participants were randomly assigned to one of the eight different experimental groups corresponding to eight different learning conditions.
2.1.2 Materials
Two different base artificial languages (ALs), one for each phonotactic condition (aligned and non-aligned), were created by concatenating 4 bisyllabic and 4 trisyllabic artificial words beginning with /
/. Each AL thus consisted of 8 artificial words varying in length. In the phonotactically aligned AL (Appendix A), each word ended with a fricative or a nasal that prevented the construction of Obstruent-Liquid (OBLI) clusters with the following word-onset /
/. Conversely, in the phonotactically non-aligned AL (Appendix B), each word ended with an obstruent to form OBLI clusters with the following word-onset. We used only OBLI clusters because these are the only clusters that are unanimously considered as indivisible (Dell, 1995). Hence, artificial words were constructed with 20 CVC syllables of which 11 were shared by both base AL whereas 9 syllables differed according to the alignment condition. TPs between syllables were matched across the different artificial languages. TPs were equal to 1 within words but varied from 0.1 to 0.4 between words. All syllables used were recorded in isolation in a sound-attenuated booth by a French female speaker and extracted at zero-crossing. All words were simply created by concatenating these recordings of single syllables.
Four versions of each base AL were constructed resulting in eight different learning conditions that varied in terms of the availability of TPs and of the three specific sublexical cues (see Table 1). Segmental cues, that is, phonotactic and acoustic cues, were manipulated at word onsets while supra-segmental cues or prosodic cues were manipulated on word-final boundaries.
Eight different versions of the artificial language (AL) according to the type and position of sublexical segmentation cues provided: prosodic cues (presence (+) or absence ( ) of a final primary accent at word-final boundaries); phonotactic cues (word-initial boundaries are aligned (+) or non-aligned (-) with a syllable boundary; acoustic-phonetic cues (syllables are coarticulated (coart.) or concatenated (concat.) at word boundaries).
The contribution of phonotactic cues was evaluated by comparing conditions in which word onsets were aligned with a phonotactically imposed syllable boundary (e.g.,
) with conditions in which word onsets were misaligned with the phonotactically imposed syllable boundary (e.g.,
). Indeed, a syllable boundary is required between /n/ and /
The contribution of acoustic-phonetic cues to segmentation was tested by comparing conditions in which the two adjacent consonants at the word boundary were either coarticulated or not. Stimuli in the non-coarticulated condition were created by simply concatenating the preceding word’s final syllable with the first syllable of the following word. Stimuli in the coarticulated non-aligned condition were created by cross-splicing part (i.e., CV) of the preceding word’s final syllable (CVC) onto the beginning of a produced consonant-consonant-vowel-consonant (CCVC) syllable whose initial C corresponded to the excised final C of the preceding word’s final syllable. For the coarticulated aligned condition, the preceding word-final CVC syllable was concatenated with the coarticulated CVC syllable excised from thee coarticulated CCVC sequence.
Finally, the contribution of word-final prosodic cues was evaluated by comparing conditions in which a final primary accent was assigned (Prosodic (Pro) versions) or not to the last syllable of each artificial word. Final primary accent was marked by increasing the syllable’s intrinsic duration by 30% and the F0 by 30 Hz (corresponding to 2.3 semitones).
Pitch manipulations were performed through PSOLA (Pitch Synchronous OverLap-Add), a technique implemented in Praat that removes naturally produced pitch variations so that voiced portions had a constant pitch of 210 Hz (the speaker’s mean pitch). A first pitch point was added at the onset of the first consonant and a second pitch point was added at the end of the vowel. The second pitch point was raised 30 Hz above the initial pitch.
Two categories of foils were also created for the forced-choice tests: eight non-words (NWs) and two sets of eight part-words each (PWs). NWs were created by concatenating syllables (from the collection of 20 CVC syllables) that never occur contiguously in the AL. Moreover, NWs never began or ended with the first or final syllable of an artificial word. Thus, TPs between syllables in the NWs were null. PWs consisted of sequences of syllables crossing an artificial word boundary in the language. They were constructed by concatenating the last or the two last syllable(s) of an artificial word respectively with the two first syllable(s) or the first syllable of another artificial word. For example, the PW /
/ was created with the final syllable of /
/ and the first syllable of /
/.
2.1.3 Design and procedure
Participants were tested individually in a sound attenuated booth. The experiment was divided into two phases: a learning and a test phase. In the learning phase, participants heard a 12 minutes sequence of continuous re-synthesized speech (around 3 syllables-per-second within the non-lengthened versions) consisting of concatenated artificial words randomly selected according to the TPs specified with no intervening pauses (e.g.,
). Each artificial word was repeated 96 times and all the repetitions of the same word were strictly identical. Participants were explicitly instructed to extract and remember artificial words from this continuous sequence of concatenated artificial words with no intervening pauses. No information about the structure, the length, or the number of the artificial words was given.
In the test phase, participants’ storage of the artificial words was assessed both by a classical non-speeded two-alternative forced-choice test (2AFC) and an auditory lexical decision task (LDT) for all participants. Both of these tasks have been used in the literature to assess word learning (2AFC: Saffran, 2002; LDT: Batterink & Neville, 2011). In the auditory lexical decision task, participants received 48 items presented in a random order: eight artificial words repeated three times; eight non-words; and sixteen part-words. They were required to decide, as fast as possible, whether each item corresponded or not to a word from the artificial language.
Two different versions of the two-alternative forced-choice test were constructed for each AL, respecting the same constraints but differing in the specific set of PWs used. Each test consisted in 32 Word-NW, 32 Word-PW and 32 NW-PW pairs. NW-PW pairs were included to encourage participants to evaluate the two choices carefully before responding. Participants were randomly assigned to one of two tests corresponding to the version of the AL they have learnt. Participants heard 96 pairs of bisyllabic or trisyllabic sequences, separated by an inter-stimulus interval of 500 milliseconds (ms). They had to indicate, as accurately as possible, which of the two sequences corresponded to (or sounded more like) a word from the artificial language. Response times were not measured. A four second interval was allowed for the answer before the beep indicating the beginning of the following trial. Four practice trials were given prior to the test in order to clarify the instruction and to acquaint participants with the structure of the test. These practice trials consisted of sequences of syllables which were part of the language’s syllable inventory but not exemplified either in the artificial language, or in the tests.
2.2 Results
2.2.1 Auditory lexical decision task
All 145 participants were included in the analyses. All statistical analyses were performed using the lmerTest package (version: 2.0-11) of the statistical software platform R (R Core Team, 2014). Analyses were run for accuracy and speed (reaction time (RT)) of responding to word stimuli by means of logit and linear mixed effects regression models (Baayen, Davidson, & Bates, 2008) respectively. The final model included the maximal random effects structure justified by the data, that is, random intercepts for participants and items, by-participant random slopes for phonotactic and prosodic cues predictors and their interaction and by-item random slopes for phonotactic and prosodic cues predictors. Only fixed predictors that were statistically significant or figured in statistically significant interactions were retained in the final model.
2.2.1.1 Accuracy
The number of data points on which further analyses were conducted totaled 3480 (145 participants listening to 3 repetitions of 8 artificial words). Both the availability of phonotactic cues and the availability of prosodic cues were entered as factorial predictors. Coarticulation cues and the trial number which represents the familiarity of participants across language exposure were excluded from the final model because these predictors were not statistically significant and did not figure in any statistically significant interactions. Table 2 summarizes the random-effects and fixed-effects structure of the final model, including β weights (estimate), standard errors, z-scores and p-values for main effects. Interactions were analyzed using the treatment contrast function, contr.treatment (). Simple effects were tested by using the relevel () function to change the reference level of the factors. Results are presented in Figure 1.
Summary of the random and fixed effects in the mixed logit model for accuracy on word items in the lexical decision task (Exp.1. N = 3480 observations. log. Likelihood = - 2058.6). The intercept represents artificial words produced with misaligned word onsets and without prosodic cue.

Accuracy (% correct) in the lexical decision task as a function of phonotactic alignment, and the availability of fine-grained acousticphonetic cues and prosodic cues (Exp.1).
Main effects of word-initial phonotactic cues (β = 0.35, z = 2.1, p < 0.05) and word-final prosodic cues (β = 0.13, z = 1.9, p < 0.05) were found suggesting that phonotactic alignment between words and syllables and the availability of a final primary accent on the last syllable of the artificial words have facilitated word learning. A significant two-way interaction between word-final prosodic and word-initial phonotactic cues (β = 0.17, z = 2.5, p = 0.01) suggested that these effects were modulated by the presence of other cues in the speech signal.
Simple effects were tested by using the relevel () function to better understand this two-way interaction. These additional analyses showed that word learning was improved when both cues were available (Effect of prosodic cues: β = – 0.6, z = – 3.08, p < 0.01: Effect of word-initial phonotactic cues: β = – 1.05, z = – 3.3, p < 0.001 with higher accuracy for aligned than non-aligned ALs) while no effect was found when either phonotactic or prosodic cues were not available (For prosodic cues: β = – 0.08, z = – 0.4, p = 0.7; for phonotactic cues: β = 0.34, z = 0.9, p = 0.4).
2.2.1.2 Response latencies (RTs)
Analyses were conducted on 2298 observations corresponding to RTs for correct responses to artificial word stimuli. The final model included the maximal random effects structure justified by the data, that is, random intercepts for participants and items and by-item random slopes for prosodic cues. Only fixed predictors that were statistically significant or figured in statistically significant interactions were retained in the final model. By being irrelevant for this model, coarticulation cues and the performance over the entire task (familiarization) were excluded from the analyses. The final model thus contained two fixed-effects predictors: the availability of word-initial phonotactic; and word-final prosodic cues. Table 3 summarizes the random-effects and fixed-effects structure of the final model, including β weights (estimate), standard errors, t-values and p-values. Results are presented in Figure 2.
Summary of the random and fixed effects in the generalized linear mixed model for response times on word items in the lexical decision task (Exp.1, N = 2298 observations). The intercept represents artificial words produced with misaligned word onsets and without prosodic cue.

Response times (RTs) on artificial words in the lexical decision task as a function of phonotactic alignment, and the availability of fine-grained acoustic-phonetic cues and prosodic cues (Exp.1).
A main effect of word-initial phonotactic cues (β = – 171.6, t = – 4.7, p < 0.0001) showed that participants were more rapid in identifying artificial words when syllable and word-initial boundaries were aligned during the learning phase. This effect was modulated by prosodic cues that significantly interacted with phonotactic cues (β = – 87.6, t = – 2.5, p = 0.01). Simple effects were tested to better understand this interaction. These additional analyses confirmed that participants were more rapid in identifying artificial words when word-final prosodic cues were simultaneously provided with phonotactic cues, that is when artificial word-onsets were aligned with a syllable boundary (Effect of prosodic cues: β = 191, t = 2.1, p = 0.03; effect of phonotactic cues: β = 517.8, t = 5.4, p < 0.0001), as compared to when either cue alone was available (for prosodic cues: β = 158.7, t = 1.6, p = .1; for phonotactic cues: β = – 168.04, t = – 1.6, p = 0.1).
2.2.2 Two-alternative forced-choice test
One participant was excluded from the analyses due to his failure to complete the task. Analyses were run for accuracy on the 64 Word-NW and Word-PW pairs. NW-PW pairs were not included in the analyses since we did not make any specific hypothesis about their effects. Responses were analyzed using logit mixed-effects regression models (Jaeger, 2008). In this and the following experiment, models were fit using the glmer function of the lmerTest package (version: 2.0-11) of the statistical software platform R (R Core Team, 2014). Linear mixed-effects models were fit by maximum likelihood and the bobyqa and Nelder_Mead optimizers. The final model included the maximal random effects structure justified by the data, that is, three random intercepts for participants, the first item (item 1) and the second item (item 2) of the pair, random slopes by-item 1 and by-item 2 for prosodic cues predictors. A richer random effect structure would have been possible (i.e., random slopes by-item 1 and by-item 2 for phonotactic and acoustic-phonetic cues predictors) but not computationally feasible. Only fixed predictors that were statistically significant or figured in statistically significant interactions were retained in the final model.
One-tailed t-tests were calculated for all 8 conditions (Appendix C) 3 . Importantly, they showed that participants performed above chance in the TP-coart condition, that is, corresponding to the TP-only condition, M = 59.5, t(1087) = 6.4, p < 00001. The availability of phonotactic, prosodic and coarticulation cues was entered as factorial predictors in the final model that contained three fixed-effects predictors. The number of data points on which further analyses were conducted totaled 9216 (144 participants listening to 64 pairs). Table 4 summarizes the random-effects and fixed-effects structure of the final model using the contrast function, contr.sum (), that gives orthogonal contrasts where every level is compared to the overall mean. Interactions were analyzed using the by default treatment contrast function, contr.treatment (), that determines a reference level and gives the contrast values for each level of a particular factor compared to this reference level. Simple effects were tested by using the relevel () function to change the reference level of the factors. Results are presented in Figure 3.
Summary of the random and fixed effects in the mixed logit model for accuracy in the forced-choice task (Exp. 1, N = 9216 observations, log. Likelihood = - 5518.3). The intercept represents artificial words consisting of concatenated syllables, produced with misaligned word onsets and without prosodic cue.

Accuracy (% correct) in the forced-choice task as a function of phonotactic alignment, the availability of fine-grained acoustic-phonetic cues and prosodic cues (Exp.1).
Main effects of prosodic (β = 0.2, z = 3.5, p < 0.0001) and phonotactic cues (β = 0.3, z = 3.2, p = 0.001) were found suggesting that phonotactic alignment at word-initial boundaries and prosodic cues at word-final boundaries have both facilitated word learning. Moreover, phonotactic cues interacted with both prosodic cues (β = 0.18, z = 2.9, p < 0.01) and coarticulation cues, β = 0.13, z = 2.5, p = 0.01. Finally, our data revealed a significant interaction between prosodic and coarticulation cues, β = – 0.1, z = – 2.07, p = 0.03. These interactions suggested that main effects were modulated by the presence of other sublexical cues in the speech signal. Simple effects were tested to better understand these interactions.
Additional analyses on the first interaction between phonotactic and prosodic cues showed that phonotactic alignment at word-initial boundaries (β = 0.74, z = 2.6; p = 0.008) and the presence of an accent at word-final boundaries (β = 1.03, z = 4.62, p < 0.0001) both improved word learning when the two cues were provided while neither phonotactic alignment (β = – 0.0004, z = 0.002, p = 0.9) nor prosodic cues (β = – 0.28, z = – 1.2, p = 0.2) showed an improvement on their own.
Additional analyses on the second interaction between acoustic-phonetic and phonotactic cues showed that phonotactic alignment at word-initial boundaries favored word learning when syllables were concatenated (β = 1.25, z = 4.2; p < 0.0001) whereas it did not have any effect when syllables were coarticulated, β = 0.0004, z = 0.002, p = 0.9.
Finally, simple effects analyses on the third interaction between word-final prosodic and word-initial coarticulation cues showed that the presence of prosodic cues at word-final boundaries improved word learning when syllables were coarticulated (β = 1.03, z = 4.6, p < 0.0001) whereas their effect was not significant when syllables were concatenated, β = 0.12, z = 0.55, p = 0.5. Moreover, coarticulation improved word learning when prosodic cues were present at word-final boundaries (β = – 0.6, z = 0.2, p = 0.005) whereas these cues did not have any effect on word learning when prosodic cues were absent, β = 0.4, z = 1.7, p = 0.08.
To sum up, word learning was clearly improved when both phonotactic (alignment between word and syllable boundary) and word-final prosodic cues were available. The presence of prosodic cues on the final syllable of the artificial words significantly improved word learning when word onsets and syllable boundaries were aligned, while no effect of prosodic cues was found when word and syllables boundaries were misaligned. Similarly, phonotactic alignment at word-initial boundaries improved word learning when word-final prosodic cues were provided while no effect of word-initial phonotactic alignment was found when prosodic cues were not available. Therefore, the combined effect of the two cues was greater than their individual effects. Moreover, phonotactic alignment at word-initial boundaries improved word learning when syllables were concatenated, whereas it was not relevant for word learning when syllables were coarticulated. Therefore, coarticulation at word-initial boundaries has prevented participants from efficiently using phonotactic alignment cues. Finally, the presence of word-final prosodic cues improved word learning when word-initial syllables were coarticulated, whereas these prosodic cues were not relevant when syllables were concatenated.
2.3 Discussion
This first experiment examined cue complementarity by investigating how various combinations of word-initial segmental cues and word-final prosodic cues influence word learning. To this end, we had participants perform an auditory lexical decision task and a two-alternative forced-choice task after a short exposure to a new mini-language. Accuracy analyses from these two tasks revealed a main effect of both word-initial phonotactic and word-final prosodic cues. Hence, word learning was improved by both the alignment of word-initial boundaries with a phonotactic boundary and the presence of prosodic cues at word-final boundaries. Importantly, accuracy and RTs from both tasks revealed a significant interaction between these two cues. The combined effect of these two sublexical segmentation cues was greater than their individual effects, providing support for synergistic effects of word-initial phonotactic and word-final prosodic cues on speech segmentation. Finally, the forced-choice task also revealed a weakly significant interaction between prosodic and coarticulation cues. Word-final prosodic cues improved word learning when syllables were coarticulated, whereas alone, neither prosodic nor coarticulation cues facilitated word learning. This result suggests that prosodic and coarticulation cues also have synergistic effects on speech segmentation since their combination leads to greater accuracy than when either cue is alone. However, since this interaction was not found in the auditory lexical decision task, it is premature to draw definitive conclusions from this weakly significant interaction.
These complex interactions observed between acoustic, prosodic and phonotactic cues confirmed that, to reach a complete view of speech segmentation, we must investigate the interactions between multiple cues rather than focusing on the separate contribution of each individual cue (Mattys et al., 2005). Importantly, our data revealed synergistic effects of word-final prosodic and word-initial phonotactic cues to lexical segmentation. In contrast, using word identification tasks in English, Mattys et al. (2005) showed that sublexical cues were rank-ordered and that phonotactic cues clearly outweighed prosodic cues in lexical segmentation. While these authors showed that phonotactic segmentation cues (low-frequency diphones) contribute equally to word boundary detection with or without congruent prosodic cues, our study rather showed that participants needed the combination of these two cues to efficiently extract artificial words from the new language. We cannot exclude that different experimental tasks (word identification versus word learning) could lead to differences in cue use for speech segmentation.This hypothesis that cue use is task-specific will be examined in the General Conclusion.
More importantly, there are at least two other possible theoretical explanations for these discrepant results. As suggested by Mattys et al. (2005), cross-linguistic differences could lead to differences in cue use for speech segmentation. The combined influence of multiple cues on lexical segmentation could reflect the familiarity of cues as markers of word-boundaries and could thus depend upon the specific phonetic and phonological properties of the language studied. Since acoustic-phonetic cues to word-onsets–such as the aspiration of voiceless stop consonants–are particularly regular in English (Keating, 1984), it is perhaps not surprising that word-initial segmental information in this language is highly weighted in the segmentation process (Mattys et al., 2005). Moreover, since stress placement varies in English, it may be less reliable than in French. Certain linguistic properties of English may thus explain why segmental cues are more weighted than prosodic cues in English as suggested by Mattys et al. (2005). However, the picture is quite different in French since its prosodic and segmental information differ substantially from those of English. Although some acoustic-phonetic cues to word boundaries might be available in French (e.g., Dumay, Content, & Frauenfelder, 1999; Fougeron, Bagou, Content, Stefanuto, & Frauenfelder, 2003), perceptual studies showed that fine-grained phonetic cues were not reliably used to differentiate between sequences varying in these cues (e.g., Bagou, Dufour, Fougeron, Content, & Frauenfelder, 2007). Moreover, accented syllables, when present, are typically word-final, and their placement is highly regular in French. Consequently, we expected accented syllables to be more weighted in French than in English, while acoustic-phonetic cues would be less informative in French than in English.
A second plausible alternative explanation of the observed differences between the study of Mattys et al. (2005) and ours, relates to the position of the segmentation cues in the speech signal. Indeed, the phonotactic, acoustic and stress cues in Mattys et al. (2005) all marked word-initial boundaries and were thus redundant. In our study, by contrast, phonotactic and acoustic-phonetic cues signaled word-initial boundaries while prosodic cues signaled word-final boundaries and therefore, these cues were complementary. Because the cues were redundant in Mattys et al. (2005), participants might have rank-ordered them by validity and thus used the most informative one to segment words, that is, the segmental cues in English. Because cues were in complementary distribution in our study, none of them, taken individually, led to better learning, whereas their combined effect has significantly enhanced word learning. This first experiment has revealed that cue complementarity leads to a synergistic integration of multiple cues. The second experiment was designed to study the effects of cue redundancy on new word learning by French listeners. Since word-final boundaries are redundantly marked by multiple prosodic cues in French, we specifically investigate how two acoustic–prosodic features, namely final lengthening and final pitch rise, simultaneously contribute to the learning of new artificial words.
3 Experiment 2
As established by many previous studies, the production and perception of prominent syllables in speech varies across languages (Dupoux & Peperkamp, 2002; Peperkamp, Dupoux, & Sebastián-Gallés, 1999). The acoustic-phonetic marking of prominent syllables is language-specific and varies depending on whether such syllables are associated with longer duration and/or changes in fundamental frequency and intensity. In French, syllables carrying a primary accent are fixed on the last full syllable of a word. These prominent syllables are typically longer, higher in intensity than preceding syllables, and bear pitch (F0) variations. These latter variations involve a fall in F0 if the phrase is utterance final or a rise if it is not (Delattre, 1940, 1941; Benguerel, 1971; Di Cristo, 1998; Jun & Fougeron, 2000; Post, 2000, 2002). Because accented syllables are redundantly marked by multiple acoustic-phonetic cues, it is crucial to examine how such redundant cues simultaneously contribute to the learning of new words in French. Our goal is to establish whether all redundant cues are relevant for speech segmentation and which cues are preferentially used by French listeners to segment words in fluent speech.
Regarding duration, many studies have concluded that final lengthening, typically observed on the last syllables of French words, contributes to segmentation (Tyler & Cutler, 2009). Indeed, previous studies (Banel & Bacri, 1994) have shown that French listeners exploit final lengthening to segment ambiguous one word/two words sequences (e.g., “bord#dur” versus “bordure”). Banel, Frauenfelder, and Perruchet (1998) also showed that French learners of a new language were more efficient in their acquisition when the language had word final lengthening rather than isochronous syllables. Regarding pitch, many studies have shown that the pitch variations typically observed at the end of phrases in French contribute to lexical segmentation across languages when word boundaries coincide with high-level constituent (e.g., syntactic) boundaries (Christophe, Peperkamp, Pallier, Bock, & Mehler, 2004). However, little is known about how pitch variations are exploited when word boundaries are aligned with lower-level constituents (e.g., accentual phrase), and the results are ambiguous. In their artificial language study, Tyler and Cutler (2009) showed that French learners efficiently used pitch movements in the right-edge position to acquire artificial French words, whereas results from Toro, Sebastian-Gallés, and Mattys (2009) revealed that French learners did not benefit from such final pitch variations. These discrepant results are quite surprising since both studies used similar experimental designs but might nevertheless be explained by three main methodological differences between the two studies.
First, the artificial words in Tyler and Cutler (2009) were longer than in Toro et al. (2009). Half of the artificial words from Tyler and Cutler (2009) were quadrisyllabic words, while all words were trisyllabic in Toro et al. (2009). Since French APs tend to contain an average of 3.5 to 3.9 syllables in natural speech (Jun & Fougeron, 2000), one could argue that French listeners segmented the artificial language into APs in Tyler and Cutler (2009) rather than into words.
Second, pitch was increased by 6 semitones on prominent syllables in Tyler and Cutler (2009), while Toro et al. (2009) used a much smaller F0 range corresponding to 1.7 semitones. Since greater pitch ranges are related to prosodic boundaries of higher levels, it is again likely that artificial words in Tyler and Cutler (2009) were considered as high-level prosodic constituents while those from Toro et al. (2009) rather had the syllabic and prosodic structure of words.
Finally, the artificial words in Tyler and Cutler (2009) varied in length (three or four syllables), while they all had the same syllabic length in Toro et al. (2009). It is likely that pitch movement was more relevant in Tyler and Cutler (2009), since participants had to deal with the variable syllabic length of the artificial words. On the contrary, the recurrent pitch pattern repeated every three syllables in Toro et al. (2009) may have been used during the initial phases of listening but then ignored because of its redundancy with statistical cues, as soon as listeners surmised that they were listening to a list of items of the same length.
On the basis of the comparison of these two studies, we can see that the question of the contribution of pitch cues to the marking of lexical boundaries in French is best addressed by: (1) using short rather than long artificial words; (2) applying a small rather a large rise in pitch on prominent syllables; and (3) varying the syllabic length of the artificial words.
The following study addresses this question by training French listeners to use different artificial languages varying in the acoustic-prosodic properties of artificial word boundaries. More precisely, we investigate the relative contribution to word learning of the two acoustic cues, final pitch rise and final lengthening, that constitute the main prosodic information to segmentation. Like for redundant word-initial cues (Mattys et al., 2005), we expect these two redundant word-final cues to be rank-ordered, and that participants use the most informative cue.
3.1 Method
3.1.1 Participants
Forty-eight native speakers of French (9 men and 39 women), ranging in age from 18 to 35 years old (M = 22.1; SD = 3.4), voluntarily participated in the experiment for course credits. All were undergraduate or graduate students at the University of Geneva and none reported a history of speech or hearing difficulties. Participants were randomly assigned to one of the four different experimental groups corresponding to four different learning conditions.
3.1.2 Materials (Appendix D)
A first version of the AL was created by combining 18 CVC syllables. Syllables were recorded in a sound-attenuated booth by a French male speaker and resynthesized by using the PSOLA resynthesis routine of PRAAT at 110 Hz, the mean fundamental frequency (F0) of the speaker. Four bisyllabic and four trisyllabic artificial words were constructed with these 18 syllables and then concatenated together without intervening pauses between syllables. We wanted to avoid the situation where the speaker produced perceptible word boundary cues. Hence, words were not recorded as units but spliced together from single syllables spoken in isolation. Each artificial word was repeated 100 times, resulting in ten minutes of continuous artificial speech (2000 syllables). Neither final lengthening nor final pitch variations could be used to determine artificial word boundaries, forcing participants to rely on statistical information (TP version). For bisyllabic words and for half of the trisyllabic words, within words TPs of syllable pairs was 1.0. For example, in
, the probability of hearing the syllable /daz/ immediately after the syllable /
/ was 100%. However, contrary to previous artificial language studies in which TPs within words were consistently 1.0 (as in Experiment 1), we introduce slight variations in these regularities to make the language more natural. Hence, for half of the trisyllabic words, TPs between the first and the second syllable were 1.0 while TPs between the second and the third syllable were 0.5. Finally, between words, TPs varied from 0.07 to 0.2.
Three other versions of this original version of the AL were constructed depending upon the availability of two prosodic cues on the last syllable of each artificial word (Table 5). The contribution of final lengthening was investigated by evaluating word learning in conditions in which the intrinsic duration of the last syllable of each artificial word was increased by 30% (-length versions). The role of final pitch rise was tested by evaluating word learning in conditions in which the F0 of the final syllable of each artificial word was increased by 20 Hz, corresponding to 2.8 semitones (-pitch versions).
Four versions of the artificial language (AL) according to the availability of acoustic-prosodic cues at word-final boundaries: final lengthening (present (+) or absent (-) on the last syllable of each artificial word; final pitch rise (present (+) or absent (-) on the last syllable of each artificial word).
Note: lengthened syllables are underlined; the presence of a final pitch rise is represented by superscripted characters.
“#” are artificial word boundaries.
“.” are syllable boundaries.
3.1.3 Design and procedure
Participants were tested in sub-groups of two to ten. They were equipped with high-quality individual audio systems. The experiment was divided into two phases: a learning and a test phase. In the learning phase, participants were instructed to extract and remember the words of the artificial language and were informed that their knowledge about the artificial words would be tested after the learning phase. In the test phase, participants’ storage of the artificial words was assessed by a non-speeded two-alternative forced-choice test. Participants had to indicate, as accurately as possible, which of the two bi- or trisyllabic sequences corresponded to (or sounded more like) a word from the AL. A four second interval was allowed for the answer before the beep indicating the beginning of the following trial. Thirty-two trials required discriminating between words and foils while 48 trials required discriminating between two different categories of foils. Trials were randomized across participants and stimuli were separated by an inter-stimulus interval of 500 ms. Four practice trials were given prior to the test in order to clarify the instruction and to acquaint participants with the structure of the test. These practice trials consisted of sequences of syllables which were part of the language’s syllable collection but were exemplified neither in the AL, nor in the tests.
Two categories of foils were created: Eight non-words (NWs) and 24 part-words (PWs). NWs were created by concatenating syllables (from the collection of 18 CVC syllables) that never occur contiguously in the AL. Moreover, NWs never began or ended with the first or final syllable of an artificial word. Thus, TPs between syllables in NWs were null. Three categories of part-words (PWs) were formed by concatenating: (1) the last syllable of an artificial word with the first syllable of another artificial word (PW2); (2) one or two syllables from the beginning of an artificial word with the first syllable of another artificial word (PW3); and (3) one or two syllables from the end of an artificial word with the last syllable of another artificial word (PW4). Hence, PW2 consisted of sequences of syllables crossing an artificial word boundary in the language; PW3 shared initial syllables with the artificial words while PW4 shared final syllables with the artificial words.
3.2 Results
The number of data on which analyses were conducted totaled 1536 (48 participants listening to 32 word-foil pairs). Analyses were run by means of logit mixed effects regression models (Baayen et al., 2008). The final model included the maximal random effects structure justified by the data, that is, random intercepts for participants and items, random slopes by-participant and by-item for intonational and durational cues predictors, and random sloped for their interaction. Only fixed predictors that were statistically significant or figured in statistically significant interactions were retained in the final model. The availability of final lengthening, final pitch rise and the foil-type and the trial number (item familiarity across language exposure) were included in the final model that contained four fixed-effects predictors. Table 6 summarizes the random-effects and fixed-effects structure of the final model, including β weights (estimate), standard errors, z-scores and p-values. Results are presented in Figure 4.
Summary of the random and fixed effects in the mixed logit model for accuracy in the forced-choice task (Exp. 2, N = 1536 observations, log. Likelihood = - 662.7). The intercept represents pairs of items produced with neither final lengthening nor final pitch rise in Word–NW1 pairs.

Accuracy for the forced-choice task as a function of the availability of final pitch rise and final lengthening (Exp. 2).
A one-tailed t-test showed that participants performed above chance in the TP-only condition that corresponds to the TP version of the AL in our study, M = 64.6, t(383) = 5.96, p < 0.0001. A main effect of the Foil type suggested that PW2 were more confounded with artificial words than other types of foils, β = 0.3, z = 2.4, p = 0.01. This result is not very surprising since TPs between syllables in PW2 were higher than in the other types of foils. This effect thus confirmed the use of TPs in word learning. More interestingly, a main effect of final pitch rise (β = – 0.5, z = – 3.7, p < 0.0001) was found suggesting that intonational cues at the end of the artificial words have facilitated word learning whereas durational cues did not, β = – 0.07, z = – 0.5, p = 0.6.
Moreover, analyses revealed a significant two-way interaction between final lengthening and final pitch rise (β = – 0.4, z = – 3.3, p = 0.001). Simple effects were tested to better understand this two-way interaction. These additional analyses showed that the presence of either final lengthening (β = 0.99, z = 3.4, p < 0.001) or final pitch rise (β = 1.9, z = 5.1, p < 0.0001) improved word learning, hence suggesting compensatory mechanisms between the two prosodic cues. The lack of an intonational or durational cue at the end of artificial words was indeed compensated for by a relevant durational or intonational contrast. Moreover, the presence of the two cues did not significantly improve word learning since neither final lengthening (β = 0.7, z = 1.5, p = 0.1) nor final pitch rise: β = – 0.3, z = – 0.6, p = 0.5) significantly improved word learning when the other cue was also provided by the speech signal. Hence, these new results revealed that durational and intonational cues do not have synergistic effects on word learning.
To explore more deeply the impact of different learning conditions on word learning, simple effects models were run on the same data set with the version of the AL as predictor. Comparing the TP-condition to the three prosodic conditions, these complementary analyses confirmed that the presence of prosodic cues on the last syllable of the artificial words clearly improved word learning (for final lengthening: β = 0.99, z = 3.4, p < 0.001), final pitch rise (β = 1.96, z = 5.1, p < 0.0001; both cues: β = 1.26, z = 3.02, p < 0.01). More interestingly, when the TP-length version was the intercept, analyses showed that participants were better in the TP-pitch condition than in the TP-length condition, β = 0.97, z = 2.4, p = 0.01. Finally, word learning was not better in the TP-length-pitch condition than either in the TP-length (β = 0.27, z = 0.7, p = 0.5) or in the TP-pitch conditions, β = – 0.7, z = – 1.5, p = 0.13. To sum up, results first showed that durational cues were less informative than intonational cues and second, that these two acoustic-prosodic cues do not have a synergistic effect on word learning.
3.3 Discussion
Our second experiment examined segmentation in redundant conditions, that is, when two cues in the same position signal the following word onset. More precisely, we investigated how two prosodic features of word-final boundaries, namely pitch rise and final lengthening, simultaneously contribute to the learning of new words by French learners. Our analyses revealed a main effect of the final pitch rise suggesting that the presence of intonational variations at the end of the artificial words facilitated learning. More importantly, a significant interaction between final pitch rise and final lengthening was found. While the sole presence of both cues contributed to word learning, their combined effect was not greater than their individual effects. Hence, final pitch rise and final lengthening did not produce synergistic effects on word learning. Rather they were rank-ordered with final pitch variations being more heavily weighted than final lengthening.
To date, the question of how multiple prosodic cues that simultaneously mark accented syllables contribute to lexical segmentation in French has remained unanswered with contradictory conclusions coming from acoustic and phonetic perceptual studies. Some perceptual studies, in which participants had to discriminate between accented and unaccented syllables, revealed that listeners relied on multiple cues simultaneously to decide whether a syllable was accented or not (Streeter, 1978), whereas others suggested that the availability of a single acoustic cue (either duration or pitch) was sufficient to perceive a syllable as accented (Mertens, 1991). Studies that tried to establish which acoustic cue was the best marker of accented syllables also reached contradictory conclusions. While acoustic studies have generally suggested that French accented syllables were principally characterized by their duration (Léon & Léon, 1979; Paradis, 1993), perceptual studies rather revealed that pitch variations occurred on 95% of the syllables that were perceived as accented (Guaïtella, Deshaie, & Paradis, 1997).
These discrepant results can be attributed to differences in the definition of accented syllables and in the methods used to identify them. Most phonetic studies compared the physical properties of accented syllables with those of their unaccented counterparts to establish the acoustic cues of accented syllables. In these studies, participants were generally exposed to spoken corpora of natural speech and had to indicate the accented syllables. However, the definition of the accented syllables lacked precision, taking into account neither their position within the linguistic unit (final/initial syllables) nor the prosodic unit (intonational, phonological or accentual phrases) from which these syllables were excised. Clearly these factors are crucial since the relevant acoustic cues depend upon them. Perceptual studies used metalinguistic tasks that required accent judgments despite the fact that French listeners generally exhibit stress ‘deafness’ (Dupoux, Pallier, Sebastian-Gallés, & Mehler, 1997). It is thus likely that participants from these perceptual studies were not able to distinguish clearly between accented and unaccented syllables. Despite their lack of success in tasks that require them to categorize accented and unaccented syllables, French listeners are undoubtedly able to use the demarcative function of accented syllables during speech processing. To establish which cues French listeners use, it is probably better to examine the effect of these cues on speech segmentation than on metalinguistic judgments. By exploring the effects of final pitch rise and final lengthening on word learning directly, our study thus avoided such biases and allows us to draw more valid conclusions.
Our findings are in line with previous results showing that French listeners can use final lengthening to distinguish between one-word and two-word sequences (Banel & Bacri, 1994; Rietveld, 1980). Our data thus confirmed that French listeners can apply a native iambic segmentation strategy based on lengthening to extract and learn new words from an AL (for similar conclusions, see Banel et al., 1998). However, results also showed that the listener’s use of final lengthening was restricted to conditions in which the final pitch cue was not available. This result is important since the role of final lengthening has only been tested previously in isolation despite the fact that prominent syllables are commonly marked by multiple cues concurrently (Banel et al., 1998; Tyler & Cutler, 2009).
Interestingly, our data suggest that final pitch rise provides French listeners with an efficient segmentation cue. This confirms the conclusions of Tyler and Cutler (2009) but contradicts those of Toro et al. (2009) who claimed that French listeners do not benefit from final F0 variations. It is important to mention that we used a substantially smaller F0 rise than in Tyler and Cutler (2009) but a greater F0 rise than in Toro et al. (2009). Final lengthening and final pitch rise were rank-ordered with final pitch variations being more heavily weighted than final lengthening. This conclusion is in line with previous prosodic studies that showed that word boundaries are commonly marked only by a pitch cue in French (Collier, Roelof de Pijper, & Sanderman, 1993) and that final lengthening at phrase-final boundaries is shorter in French than in other languages (Benguerel, 1971; Fant, Kruckenberg, & Nord, 1991; Fletcher, 1991). It is thus not surprising that French listeners relied more on pitch rise than on final lengthening in our study. However, this conclusion goes against Rietveld (1980) who claimed that French listeners mainly use final lengthening to distinguish between one-word (“le comtat/the country” in “le comtat saccagé/the devastated country”) and two-word sequences (“le conte a”/the count has” in “le conte a saccagé/the count has laid waste”). This difference may be accounted for by differences in task demands, a simple segmentation task (Banel & Bacri, 1994; Rietveld, 1980) versus our more complex word learning task. Learning words in continuous speech requires both segmenting but also remembering the segmented speech units. Spitzer, Liss, and Mattys (2007) recently showed that F0 flattening has a greater detrimental impact on lexical segmentation than the weakening of other prosodic cues when listening is effortful. Hence, among the different prosodic cues involved in segmentation, F0 variations might be the most resistant to greater task demands as suggested by Spitzer, Liss, Dorman, Spahr, and Lansford (2009) who showed that listeners with low-frequency hearing attend to F0 variations to guide segmentation.
Finally, in our word learning study, pitch and durational cues did not yield a synergistic effect. Since accent is often marked by multiple cues simultaneously, this result might be surprising but can, nevertheless, be better understood if one considers the relationship between F0 and syllable duration in the marking of accented syllables. Results from Alain (1993) revealed a negative correlation between F0 and syllable duration for accented syllables in French. In this phonetic study on French, 16 speakers were required to produce the same simple syllable /pa/ with different accentuation patterns. Data showed that F0 and duration both contributed to distinguishing between unaccented and accented syllables in French but were negatively correlated in accented syllables. This trade-off between pitch and duration reveals that speakers can use either pitch or duration to mark prosodic boundaries, at least at the accentual phrase level (Michelas & D’Imperio, 2010). Listeners should be able to adapt rapidly to the particular way a speaker marks prosodic breaks and focus their attention on the most informative cue. In order to test the relative contribution of these two cue types more directly, it would be necessary to put them in conflict in the same language learning experiment.
4 General conclusion
Two AL studies were designed to investigate how French learners extract and integrate multiple convergent sublexical segmentation cues during word learning. First, we examined how three types of sublexical cues, namely fine-grained acoustic-phonetic cues, phonotactic cues and prosodic cues, contribute concurrently to lexical segmentation in word learning. Second, we investigated how multiple sublexical cues within one of these cue types, namely prosodic cues, contribute to lexical segmentation. More precisely, we investigated how final lengthening and final pitch rise signal accentual phrase boundaries and guide word learning.
The first experiment revealed complex interactions between prosodic cues and both phonotactic cues and fine-grained acoustic-phonetic cues. Overall, this first experiment showed that word-initial segmental (both phonotactic and fine-grained acoustic) cues and word-final (prosodic) cues are used in a synergistic way in word learning. The first interaction–observed for both RT and accuracy in the two tasks–showed that word learning improved only when both prosodic and phonotactic cues were added to statistical information, whereas the sole presence of only one of these cues did not significantly improve word learning. This first interaction thus reveals that cue complementarity leads to a synergistic integration of multiple cues.
Similarly, the second interaction, found in the forced-choice task but not in the lexical decision task, revealed that prosodic cues also interacted synergistically with fine-grained acoustic-phonetic cues. Prosodic cues improved word learning when syllables were coarticulated but not when the syllables were concatenated. This latter result is somewhat unexpected. Indeed, the absence of coarticulation between adjacent segments, as in the case of the concatenated conditions, conveys information about the presence of a boundary in continuous speech (Fougeron & Keating, 1997; Keating, 2006). Accordingly, the concatenated conditions of the AL should have led to better performance on the two tasks than the coarticulated conditions. However, contrary to our expectations, we observed no benefit on segmentation in the concatenated conditions which were thought to favor segmentation. Moreover, poorer learning performance was observed when prosodic cues were available than when they were not. This poorer performance might be explained if one considers the nature of the concatenated conditions where all syllables, that is both between-word and within-word syllables were concatenated. Here, the acoustic-phonetic properties of these two types of syllables were in fact identical and presumably favored segmentation at every syllable. For the condition in which the prosodic cues were available, these cues and between-word acoustic-phonetic cues converge on the same segmentations. However, the inappropriate acoustic-phonetic segmentation cues for the within-word syllables were incompatible with the prosodic parsing. In this case, we assume that these conflicting acoustic-phonetic cues interfered with the use of prosodic cues. In contrast, in the coarticulated AL, the acoustic-phonetic properties of the within-word syllables were consistent with the prosodic parsing. Here, the coarticulation of the between-word syllables did not interfere with the use of prosodic cues. This might explain why prosodic cues facilitated lexical acquisition in the coarticulated conditions only.
To better understand how prosodic cues and coarticulation interact in word learning, further studies should compare the effects of such cues in ALs in which within-word syllables are coarticulated and between-word syllables are concatenated with ALs in which all syllables are concatenated. Particularly in the former case, the combination of prosodic boundary cues with coarticulated within-word syllables, should improve listeners’ learning considerably in a synergetic fashion. Moreover, to investigate how coarticulation influences segmentation further, we should compare ALs in which the degree of coarticulation between the between-word and within-word syllables is varied.
The second experiment showed that final pitch rise and final lengthening do not produce synergistic effects but rather are rank-ordered with final pitch variations being more heavily weighted than final lengthening. By establishing that final pitch cues to accentual phrase boundaries constrained word learning, we have extended previous findings that dealt with word recognition to word learning and to smaller prosodic units. Past studies (for French, see Christophe, 1993; Christophe et al., 2004; Millotte, 2005; Millotte, Wales, & Christophe, 2007; Millotte et al., 2008) have shown that major prosodic breaks–such as intonational and phonological phrases–were used to segment words in continuous speech and even influenced word recognition by changing lexical activation. Future work should investigate how these different types of prosodic boundaries, that is, intonational, phonological and accentual phrase boundaries are exploited by listeners during French word learning. To this end, an AL that mimics the prosodic structure of French sentences could be used to examine whether words at the edge of major prosodic boundaries are more accurately and rapidly acquired than words at the edges of minor prosodic boundaries.
Taken together, the results from these two experiments suggest that redundant word-final prosodic cues are hierarchically organized, whereas complementary word-initial segmental and word-final prosodic cues rather have synergistic effects on lexical segmentation in French. This pattern of results may help us understand previous findings in word recognition showing that multiple sublexical cues to word-initial boundaries in English were rank-ordered (Mattys et al., 2005). The redundancy of cues at word-initial boundaries probably reduced the listeners’ need to attend to multiple cues. Hence, listeners directed their attention more toward the information source of greater value, that is, segmental cues in English. Similarly, in our study, the redundancy of prosodic cues at word-final boundaries led participants to rely less on final lengthening. Since phonotactic, acoustic and stress cues all signalled word-initial boundaries in Mattys et al. (2005), determining the effect of complementary cues was not possible. By using multiple cues available in different positions, our study has provided some insight into how learners use both complementary and redundant cues. Since word-final cues signal the immediately upcoming word onset before the word-initial cues do, their predictive power may allow them to be weighted more heavily. Indeed, we can speculate that lexical segmentation is based mainly on word-final cues, while word-initial cues could confirm these segmentation decisions in a complementary and synergistic fashion. We cannot exclude, however, an alternative explanation according to which the interactions observed between cues relates not to their position, but rather to their type. In other words, it could be the type of the cue involved that determines the way in which they are combined. In our study, the cues for the same single position (final) were both prosodic cues, whereas the cues marking the different positions (initial and final) were two segmental cues namely, acoustic-phonetic and phonotactic cues. This cue distribution is imposed by the French language. To tease apart these alternative explanations, it would be necessary to test cue complementarity with languages which allow us to manipulate the position and nature of cues in order to explore alternative combinations of cues.
Our two experiments and those of Mattys et al. (2005) also diverge in the task with which segmentation was assessed: word learning versus word recognition. The weighting of segmentation cues for adults in learning a new language and for adults segmenting and identifying real words in their native language as in Mattys et al. (2005) is probably quite different. This conclusion is confirmed by recent word-spotting studies that investigated how French listeners recognized real French words inserted in nonsense sequences. Bagou and Frauenfelder (2008) showed that neither phonotactic cues nor acoustic-phonetic cues at word-initial boundaries influenced word identification when word-final prosodic cues were available in the speech input. Moreover, phonotactic cues only influenced word-spotting when neither prosodic nor acoustic-phonetic cues were provided. Unlike the synergistic effects of word-initial segmental and word-final prosodic cues in word learning results reported here, these word-spotting results rather suggest that sublexical segmentation cues are rank-ordered with word-final prosodic cues being more heavily weighted than word-initial segmental cues.
One likely explanation for these differences in cue use across these two different tasks is the differential availability of lexical cues for the participants. More precisely, in word recognition, since participants can rely heavily on word knowledge, a single or even no sublexical cue may be required for segmentation. In contrast, in the absence of lexical knowledge and in the presence of the unfamiliar phonology of the AL, participants must rely on the integration of several sublexical cues to learn the artificial words. Since sublexical regularities in the AL (i.e., prosodic, phonotactics, and acoustic-phonetic cues) are potentially novel, participants must first extract and learn them, for example, by computing TPs between syllables, in order to further use them efficiently in segmentation. We have seen that listeners tend to exploit all sublexical cues concurrently in a synergistic fashion to enhance this learning. Further systematic comparisons of word identification and word learning are required to confirm these claims and to investigate how the relative contribution of sublexical cues evolves across word learning while lexical knowledge progressively increases.
To conclude, this study has demonstrated the importance of evaluating lexical segmentation in learning conditions in which multiple cues are manipulated carefully and systematically. Second, our data confirmed that the relative strength of multiple sublexical cues is language-specific as previously proposed (Mattys et al., 2005). More importantly, we showed that cue use in word learning differs from that in word identification and is therefore task-specific. Third, we showed that, contrary to our expectations, final pitch and not final lengthening was the more important cue to the marking of word-final boundaries in French. Finally, our study revealed that listeners use cues that signal different boundaries (word-initial versus word-final boundaries) in a synergistic fashion. In contrast, they use the most efficient cue when two cues signal the same word-final boundaries. We conclude that examining cue contribution either through a cue-by-cue approach or by pitting conflicting segmentation cues against each other pairwise has clearly limited our understanding of how multiple sublexical cues contribute to lexical segmentation. Achieving a complete view of speech segmentation requires: (1) investigating the complex interactions between different linguistic cue types and within the same cue type; and (2) addressing the question of cue complementarity and cue redundancy by manipulating cues type and cue position at word boundaries orthogonally.
Footnotes
Appendices
Artificial words, part-words and non-words used in Experiment 2.
|
Acknowledgements
We thank Sophie Gallot for her assistance in running the experiments.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by grant 100014_152793 from the Swiss National Science Fundation.
