Abstract
As Chinese shows both satellite- and verb-framed properties (Slobin, 2004; Talmy, 2012, 2016), it provides a unique lens through which to observe the extent of first-language (L1) typological influence in second language (L2) acquisition of motion expressions. This study has dual purposes. First, it extends Wu’s (2016) investigation on motion expressions produced by 80 L1 satellite-framed English learners of L2 Chinese to include newly collected data produced by L1 verb-framed speakers, a sample comprised of 41 L1 Japanese learners of Chinese and 40 Japanese native speakers. Second, it synthesizes the data from both studies and comprehensively examines factors that have been proposed to affect development of L2 thinking-for-speaking (TFS) patterns. The results show that development of L2 TFS is best predicted by learners’ L1 type, but the effect is mitigated by L2 proficiency. While the L1 English learners outperform L1 Japanese learners in their development of target-like L2 Chinese TFS, learners with limited L2 proficiency in both groups tend to adopt verb-framed strategies to express only the core path information of a motion event and leave out the manner details. Analysis of L1 Japanese learners’ oral narratives in L1 Japanese and L2 Chinese also shows that reverse L2-to-L1 transfer is less likely to happen when learning a typologically closer L2 that requires minimal restructuring of their L1 TFS.
Keywords
I Introduction
Talmy’s (1985, 1991, 2000) seminal work on motion event typology and Slobin’s (1987, 1996, 2003) thinking-for-speaking (TFS) hypothesis demonstrate how motion events are encoded differently and how their speakers come to develop typologically distinct TFS patterns in their first language (L1). For instance, speakers of a satellite-framed language such as English describe the scene of an owl flying out from a tree hole as An owl flew out, with a focus on manner of motion encoded in the main verb flew. Speakers of a verb-framed language such as Spanish, by contrast, describe the same scene as Salió un buho, literally ‘An owl exited’, which underscores path in the main verb salió ‘exited’ and leaves manner unspecified (Slobin, 2004: 63). The crosslinguistic contrast in how manner or path is habitually encoded in the main verb has spurred considerable interest for second language (L2) researchers to study how the L1–L2 typological similarities and differences in conceptual representation of motion can facilitate or impede L2 learning. Jarvis and Pavlenko (2008) have referred to the influence of one’s L1 TFS in learning to express motion events in a typologically distinct L2 as a kind of ‘conceptual transfer’. It goes beyond learning new mappings of form and meaning and further involves restructuring one’s native extralinguistic mental representations: a process characterized as ‘rethinking for speaking’ by Robinson and N. Ellis (2008: 527).
While studies examining different pairings of L1 and L2 have convergently shown the influence of L1 TFS patterns on use of L2, two main issues lead to inconclusive findings in this line of research. The first issue has to do with the research design. To date, there are a considerable number of studies that include only two languages in their studies (e.g. how Spanish-speaking children learn L2 English motion expressions in Aveledo and Athanasopoulos, 2016; how German–Turkish bilinguals transfer conceptualization patterns of motion events from the dominant language in Daller et al., 2011). It is ideal to have a balanced three-language design that examines how learners of two distinct L1 types (i.e. satellite-framed language vs. verb-framed language) acquire a target language (see Cook, 2015). Only when comparing performance between learners with a typologically similar or different L1 can we pinpoint the extent of L1 typological influence on L2 learning. The challenge of recruiting many participants from different L1 backgrounds is likely to be the main reason there are limited studies that have adopted a three-language design (see also Jarvis, 2000). Moreover, it is difficult to synthesize the findings from different studies, because (1) various tasks have been developed to elicit use of motion language with a varying number of participants, and (2) the L2 proficiency of the participants, a critical modulating factor (e.g. Cadierno and Ruiz, 2006; Park, 2020), is often not independently assessed to evaluate its role in the learning process (e.g. Cadierno and Ruiz, 2006; Daller et al., 2011). The second issue has to do with the selection of languages. Research on motion events has concentrated on L1–L2 parings that fall in Talmy’s binary typology. The potential challenges associated with learning a third type of language, i.e. equipollently-framed languages such as Chinese and Thai, proposed by Slobin (2004, 2006), are largely overlooked. Examining how speakers of satellite-framed (S-framed) and verb-framed (V-framed) languages learn to express motion events in the so-called equipollently-framed (E-framed) Chinese, a language that straddles the characteristics of S- and V-framed languages, can provide a unique lens through which to explore the extent of L1 typological influence on L2 acquisition of motion expressions. That is, do L1 S-framed learners tend to retain their S-framed TFS in their L2 Chinese production, and L1 V-framed learners retain V-framed TFS?
We seek to address the aforementioned issues by conducting an extension study of Wu (2016) to explore how V-framed L1 Japanese speakers come to develop target-like L2 TFS in E-framed Chinese. Wu (2016) investigated how S-framed L1 English speakers (n = 80) learn to express motion events in L2 Chinese and adopted a Chinese elicited imitation task (Wu and Ortega, 2013) to assess the impact of L2 proficiency. Contrary to results reported in studies that focused on an S- or V-framed language as the target language, she found limited L1 typological influence and identified the importance of language socialization in facilitating target-like TFS. This extension study adopts the same instruments used in Wu (2016) to study V-framed L1 Japanese speakers’ acquisition of L2 Chinese motion expressions. With data collected from learners of both L1 typological backgrounds, the present study aims to provide a full picture of how speakers of S- and V-framed languages come to acquire an E-framed language and to comprehensively examine the impact of factors affecting L2 development of TFS. Notably, it also extends the scope of the previous study to survey bidirectional transfer in motion conceptualization by examining learners’ use of motion expressions in L1 Japanese and L2 Chinese. It explores not only whether the learners’ L1 Japanese influences their use of motion in L2 Chinese, but also whether their L2 Chinese learning experience influences their conceptualization of motion in L1 Japanese.
II Background
1 Typological differences in the expression of motion events
Although all languages offer ways to talk about where and how an object moves, Talmy’s (1985, 1991, 2000) binary typology of motion events demonstrates how languages in the world differ systematically in how they pack concepts of movement into linguistic forms. When describing the same motion scene, speakers of verb-framed languages (V-languages) such as Japanese, Spanish, and Greek tend to describe path of motion via the main verb (root), whereas speakers of satellite-framed languages (S-languages) such as English, German, and Russian prefer to reserve the main verb slot for manner of motion.
(1) Japanese as a verb-framed language Kare-wa kyōshitsu-ni haitta. he-CON classroom-CON entered ‘He entered the classroom.’ (2) English as a satellite-framed language He walked into the classroom.
Examples (1) and (2) show how speakers of the two types habitually pick different aspects of a motion event to fill in the main verb position; this is oftentimes an obligatory component of a sentence. V-framed Japanese speakers encode path in the sentence-final main verb haitta ‘entered’ and leave manner unspecified, while S-framed English speakers pack manner in the main verb walk and separate path using a satellite outside of the main verb (i.e. verb particle into). Talmy’s classification was then amplified by Slobin’s (1987, 1996, 2000, 2003) TFS hypothesis. Drawn from analysis of oral narratives produced by S- and V-language speakers from different age groups, Slobin identified distinct lexicalization patterns between S- and V-language speakers and demonstrated how acquiring a language shapes one’s attention to aspects of motion events that are readily codable in the language. He claimed TFS is a special form of thought for online communication developed in L1 acquisition and that language-specific TFS may affect acquisition of a new language (Slobin, 1996). It is worth noting that although TFS may exert general cognitive effects that are not confined to the use of language (e.g. Lai et al., 2014), the concept of TFS is concerned about verbal behavior such as online speech planning and language learning. Researchers (e.g. Bylund and Athanasopoulos, 2014; Cook, 2015; Lucy, 2016) have drawn clear distinctions between TFS and linguistic relativity. The former focuses on verbal evidence, and the latter on nonverbal evidence that shows how language influences thought.
2 Encoding motion events in Chinese
Talmy (2000) emphasizes that the binary motion event classification intends to describe the most characteristic modes of motion expressions in a given language that are colloquial in style, frequent in occurrence, and pervasive in their ability to explain semantic notions. However, intra- and inter-typological variations have been reported. The most significant revision to Talmy’s dichotomous framework was Slobin’s (2004, 2006) proposal that serial-verb languages such as Chinese and Thai exhibit different lexicalization patterns from S- and V-languages and should be classified into a third type: equipollently-framed languages (E-languages).
Talmy later (2012, 2016) argued that the serial verb constructions such as zǒu-jìn-lái that Slobin (2004, 2006) analysed as ‘walk-enter-come’ with both manner and path constituents receiving the same full verb status and thus regarded as E-framed are actually either S- or V-framed constructions. That is, zǒu-jìn-lái should be analysed as an S-framed construction ‘walk-into-hither’ with zǒu ‘walk’ as the only main verb encoding manner, jìn ‘into’ as the first directional complement (DC, a type of path satellite), and lái ‘hither’ as the second DC indicating the deictic path of whether the movement is toward or away from the speaker.
The authors of the present study agree with Talmy’s analysis of the Chinese motion construction. Most Chinese linguists and language textbooks treat the path component(s) that follow a manner verb as DCs, a small closed set
1
that have lost their full verb status and function as satellites attached to the manner verb to indicate path (e.g. Chao, 1968; Cheung et al., 1994; Liu and Yao, 2009; Liu et al., 1983). Motion constructions like zǒu-jìn-lái ‘walk-into-hither’ refer to a single movement of someone walking in towards the speaker, rather than a series of motions ‘walk-enter-come’. Compared to the head of the construction (i.e. the manner verb zǒu ‘walk’), the DCs jìn ‘into’ and lái ‘hither’ are syntactically more restricted in their abilities to take tense and aspect markers. For example, one can add the perfective aspect marker le or durative marker zhe to the manner verb zǒu ‘walk,’ but not the DC jìn ‘into’ (* zǒu-jìn-
However, different from English path satellites (e.g. up, down, into), the small set of Chinese post-verbal DCs, which were grammaticalized from full path verbs, can still function as full path verbs today (Beavers et al., 2010; Peyraube, 2006; Shi, 2011; Talmy, 2012). This renders Chinese either a case of intra-typological variation different from other S-languages or a case of an ‘equipollently-framed language’ because of the dual roles of the same set of path morphemes in S- and V-framing and the high accessibility of both framing types in the language. As shown in (3a) and (3b), one can describe the scenario of a person walking into a classroom through an S- or V-framed encoding, with the later V-framed option formed simply by omitting the manner verb. The same path morpheme jìn is used as a path satellite (i.e. DC) ‘into’ as in (3a), and as full path verb ‘enter’ as in (3b).
(3) a. S-framed encoding: jìn as a path satellite/directional complement 他 走 Tā zǒu he walk ‘He walked into the classroom.’ b. V-framed encoding: jìn as a full path verb 他 Tā he ‘He
The dual roles of the same set of Chinese morphemes as path satellites (an S-framed means) and full path verbs (a V-framed means) present a rare case where learners of Chinese can opt for either an S- or V-framed construction to describe a motion event in Chinese and be grammatically correct with either option. For L2 learners with limited L2 proficiency, producing V-framed constructions like (3b) is morpho-syntactically and conceptually easier than S-framed constructions like (3a). This is because S- and V-framed constructions share the same set of path morphemes. When one chooses to describe a motion event via an S-framed option, the more complex structure – i.e. incorporation of an additional manner verb such as zǒu ‘walk’ in (3a) – is likely to require more online processing effort than choosing a V-framed option. Learners would need to develop not only a sufficient L2 manner verb lexicon but also higher L2 competence to automatically process both manner and path components of a motion event during online speech.
Hypothetically speaking, if L1 habitual TFS plays a large role in L2 acquisition of motion expressions, learners with an S-language as the L1 would be more likely to transfer their L1 S-framed strategies to Chinese and be inclined to use more manner verbs and adopt motion constructions like (3a). Learners with a V-language as the L1, by contrast, would be inclined to use more path verbs and adopt V-framed constructions like (3b). Next, we will examine factors that have been proposed to play a role in the development of L2 TFS.
3 Crosslinguistic influence in the acquisition of motion expressions
To date, research on motion events has concentrated mostly on comparing L1s and L2s that fall in Talmy’s binary typology (1991, 2000). Most studies report distinct TFS patterns between V- and S-language speakers and observe L1-to-L2 transfer, even for very proficient L2 speakers (e.g. Cadierno, 2004, 2010; Choi and Lantolf, 2008; Hendriks and Hickmann, 2015; Stam, 2010; Yoshioka and Kellerman, 2006). For example, studying the speech and co-speech gesture patterns by L1 English learners of L2 Korean (V-language) and L1 Korean learners of L2 English, Choi and Lantolf (2008) reported that both groups of learners, despite their high proficiency in their respective L2, appeared to retain their L1 TFS in their L2 production. Likewise, analysing the narratives produced by L1 S-framed English learners of L2 V-framed French, Hendriks and Hickmann (2015) found that the learners used more English-like S-framed structures to express motion events even at an advanced level. Moreover, studies that include at least three languages to compare how learners of two different L1 types perform in learning the same L2 have shown that learners whose L1 is typologically closer to the L2 perform better than learners whose L1 is typologically more distant from the L2 (e.g. Cadierno, 2010; Jessen, 2014). In Hijazo-Gascón (2018), he identified that L1-to-L2 conceptual transfer can happen when there are intra-typological differences between genetically close L1 and L2.
On the other hand, reverse L2-to-L1 transfer has been reported in a few studies. Brown and Gullberg (2008) examined use of manner in speech and co-speech gesture by L1 Japanese learners of L2 English with intermediate proficiency. In terms of encoding of manner in speech and gesture, Japanese learners’ L1 Japanese and L2 English production was found to be less than monolingual English native speakers (NSs) but not different from monolingual Japanese NSs. And there were no significant differences between learners’ L1 and L2 production. The results suggest that learners transferred their L1 Japanese TFS to L2 English, and there was possible convergence between L1 and L2 systems. Most crucially, it was observed that learners differed from monolingual Japanese NSs by encoding manner in speech but often not in accompanying gesture in their L1 and L2 production, suggesting a shift toward English TFS and effects of L2 English on L1 Japanese. Brown and Gullberg concluded that bidirectional interaction between languages in bilingual minds can occur even with intermediate L2 proficiency. In a follow-up study by Brown (2015), she investigated encoding of manner in speech and co-speech gesture by two groups of L2 English learners with intermediate proficiency. She observed that L1 Chinese learners of L2 English produced more manner-highlighting gestures than monolingual Chinese NSs, and L1 Japanese learners of L2 English produced less manner-highlighting gestures than monolingual Japanese NSs. The observation, again, suggested effects of L2 English on learners’ L1.
In addition to typological effects, studies that included learners of different proficiency levels have identified that learners’ abilities to use target-like motion patterns are largely modulated by L2 proficiency (e.g. Cadierno and Ruiz, 2006; Park, 2020). The specific facets of L2 proficiency that have been proposed to connect to L2 development of target-like TFS include online processing competence and size of motion verb lexicon (e.g. Brown, 2015; Cadierno and Ruiz, 2006; Filipović and Vidaković, 2010; Lewandowski and Özçalışkan, 2021; Stam, 2010). It has been reported that learners at lower proficiency levels tend to omit encodings of manner and retain path information regardless of the L1–L2 typological similarities, as path is arguably the only mandatory core component to successfully describe a motion event (Talmy, 1991, 2000). For example, Brown (2015: 74) found both L1 V-framed Japanese learners and L1 E-framed Chinese learners produced significantly fewer manner encodings in L2 English speech than English NSs. Notably, the L1 Chinese learners encoded significantly less manner information in L2 English (M = 0.74), as compared to the manner encoding in Chinese by monolingual Chinese NSs (M = 0.95) and the manner encoding in English by monolingual English NSs (M = 0.98). The result suggested that the L1 Chinese learners did not transfer the manner elaboration from their L1 Chinese to L2 English and instead opted to leave out manner details and adopt more V-framed means in their L2 English. Brown (2015) considered omission of manner information a universal developmental constraint that was particularly prominent among learners at lower levels of L2 proficiency regardless of their L1 typological background, possibly due to insufficient lexical knowledge on manner verbs or limited online processing competence. In the same vein, Wu (2016) also identified that L1 English learners of L2 Chinese at the low-proficiency level deviated from their L1 S-framed English strategies and adopted more V-framed means in L2 Chinese. That is, they tended to use the set of path morphemes as full path verbs – e.g.
Finally, more language exposure or interaction with members of the target language community has been identified as a critical factor in facilitating the development of L2 TFS (e.g. Daller et al., 2011; Stam, 2010, 2015). Wu (2016) found that heritage language learners performed better in supplying target-like motion patterns than foreign language learners at the same proficiency level.
Research so far has shown that L2 acquisition of motion expressions is a dynamic learning process shaped by multiple interacting factors. It remains to be explored how these factors weigh against each other and how they matter in learning an E-framed language.
III The present study
Following up on Wu’s (2016) investigation into L1 S-framed English learners’ L2 acquisition of Chinese motion expressions, the present study collected comparable data from L1 V-framed Japanese learners so as to gain a full picture of L1 S- and V-language speakers’ development of TFS in L2 Chinese, a language that straddles the S- and V-language types.
To gain an overarching view of the crosslinguistic typological influence, the first two research questions focused on the synthesis and analysis of the data collected from both studies, including three baseline groups of NSs (English, Japanese, and Chinese) and two learner groups (L1 Japanese and L1 English).
Research question 1: When both S-framed and V-framed options are available in the target language, do learners with a V-language as their L1 (i.e. Japanese) choose to retain their L1 V-framed TFS in their L2 Chinese narratives? How does their performance compare to that of learners whose L1 is an S-language (i.e. English)?
Research question 2: Among the factors of L1 type, L2 proficiency, age of acquisition, language exposure, and length of in-country experience, what factor(s) make a significant contribution to L2 development of target-like Chinese TFS?
The next three questions closely examined L1 Japanese learners’ L1 Japanese and L2 Chinese narratives and explored whether there were traces of bidirectional crosslinguistic influence.
Research question 3: What types of Chinese motion constructions do L1 Japanese learners prefer? Do we see traces of L1-to-L2 crosslinguistic influence?
Research question 4: If L1-to-L2 crosslinguistic influence is observed, does the effect gradually diminish with increasing L2 proficiency?
Research question 5: Does the L1 Japanese learners’ conceptualization of motion events in L1 Japanese change due to their L2 Chinese learning experience? Specifically, does their use of Japanese motion expressions differ from that of Japanese NSs?
IV Method
1 Participants
This extension study recruited 40 Japanese NSs from a public university in Japan and 41 L1 Japanese learners of L2 Chinese from a public university in China. The newly collected data were then compared to data collected in Wu (2016), which included 80 L1 English learners of L2 Chinese, as well as two baseline groups of 40 English NSs and 40 Chinese NSs. Participants in all three baseline NS groups had lived in a foreign country for less than three years. To ensure the learner participants were able to complete the oral narrative task, all learner participants had studied Chinese for at least one year. Their L2 Chinese proficiency was independently assessed through a Chinese elicited imitation test (EIT; see details in Section IV.2) to ensure an accurate and consistent proficiency measure. Table 1 summarizes the two learner groups’ EIT scores and their biographical information.
Summary of biographical information.
Notes. a The highest possible EIT score was 120, based on 30 items polytomously scored from 0 to 4. b Exposure represents learners’ self-rating on their use of Chinese outside the classroom: ‘Not applicable’ and ‘never’ were coded as 0, ‘occasionally’ as 1, ‘sometimes’ as 2, ‘frequently’ as 3, and ‘almost always’ as 4.
As shown in Table 1, L1 Japanese learners had a mean EIT score of 72.98 (SD = 25.18), higher than the L1 English learners at 56.03 (SD = 24.16). An independent-samples t-test confirmed that the difference in EIT scores was statistically significant, t(119) = 3.601, p < 0.05, Cohen’s d = 0.947. Dividing the EIT scores into five proficiency levels (Novice = 0–24; Low = 25–48; Intermediate = 49–72; High = 73–96; Advanced = 97–120), while both groups of learners were at an Intermediate level, the L1 Japanese group’ proficiency in Chinese was close to High, and the L1 English group was at the lower end of Intermediate.
In terms of the average age when they started to study Chinese, the Japanese group started at 16.78 years old, slightly later than the English group at 14.91. Regarding their exposure to Chinese outside of the classroom, the Japanese group had a mean rating of 2.56, suggesting that they heard or used Chinese at a frequency of exposure between ‘sometimes’ and ‘frequently’, which was notably higher than their English counterparts’ mean rating at 1.47 that fell between ‘occasionally’ and ‘sometimes’. Finally, the index of in-country experience reports learners’ experience of staying in a Chinese-speaking country. The average length was 25.12 months for the Japanese group and 23.05 for the English group. Overall, the wide standard deviations for each index within the respective group suggest that both groups of learners exhibited a wide range of Chinese abilities and experiences.
2 Procedure and instruments
Following the same procedure and instruments used in Wu (2016), the Japanese NSs completed only the oral narrative task in their native language to provide baseline data. The learner participants first completed the oral narrative task, followed by the Chinese EIT and a background information questionnaire. As the extension study also intended to explore whether there is a bidirectional crosslinguistic influence in conceptualization of motion events, the 41 L1 Japanese learners of L2 Chinese completed the narrative task in both L1 Japanese and L2 Chinese with a minimum of one week apart. The language sequence was counterbalanced across the 41 L1 Japanese learners of L2 Chinese.
All instruments were downloaded from the IRIS database (http://www.iris-database.org) with the instructions translated into Japanese. The oral narrative task required participants to describe a picture story that depicted 12 different motion scenes of a boy looking for his missing dog. Participants were given a few minutes to prepare their story and were instructed to describe what happened in the pictures in as much detail as possible. The elicited oral narratives were recorded and transcribed. Following the coding conventions specified in motion event studies (see Berman and Slobin, 1994; Talmy, 1985, 2000; Özçaliskan and Slobin, 2003, 2004), motion events discussed in the study considered both voluntary motion (e.g. He walked to school) and caused motion (e.g. He kicked the ball). Clauses that contained information about an entity changing location from one place to another were identified as motion clauses, and motion verbs used in the motion clauses were classified into three types of verb, including manner verbs, path verbs, and neutral verbs. According to Slobin (2004, 2005), manner verbs involve a set of dimensions that modulate motion, including motor pattern, force dynamics, rate, rhythm, posture, affect, and evaluative factors, and path verbs comprise direction of the motion, deixis (i.e. deictic direction with regard to viewpoint of the speaker), and contour such as zigzag or curved. Neutral verbs such as stand or hold do not by their nature suggest movement from one point in space to another. The sense of movement is generated only when such verbs are followed by a path satellite. Following the coding procedures in Chen and Guo (2009) and Wu (2016), compound verbs such as luàn-chuàn (‘randomly-flee’) or pǎo-diào (‘run-gone’) in Chinese, and kake-oriru (‘run down’) or tobi-dasu (‘pop out’) in Japanese were excluded from the analysis of motion verbs. This is because phrasal verbs conflate both manner and path information and thus exhibit distinct syntactic and semantic properties and cannot be classified into the three types of motion verbs. Next, motion constructions in which the motion verbs appeared were categorized and tallied according to their syntactic structures (for the specific types of Chinese and Japanese motion constructions, see Table 3 and Table 4 in Section V). The Chinese data was coded and reviewed by two of the authors who are native speakers of Chinese. With consultation with her colleagues, the Japanese data was coded by one of the authors, who is a native speaker of Japanese. All percentage data were transformed using Arcsine transformation (Warton and Hui, 2011).
The Chinese EIT, also known as sentence repetition task, is one of the parallel EITs developed based on the English version by Ortega et al. (2002). This suite of parallel EITs are available in seven L2s and have been validated as an effective tool for measuring global L2 proficiency (e.g. Kostromitina and Plonsky, 2021; McManus and Liu, 2022; Yan et al., 2016) and have been employed as a short-cut L2 proficiency measure in different L2 studies (e.g. Huensch and Tracy-Ventura, 2017; Serafini and Sanz, 2016). As defined in Wu and Ortega (2013: 684), the Chinese EIT is consistent with Hulstijn’s (2015) Basic Language Cognition (BLC) in his binary proficiency model by drawing from ‘learners’ knowledge and automated ability for the use of core vocabulary and grammar delivered with reasonable pronunciation and fluency’. Notably, EITs are a proven and effective measure for L2 automatic processing competence 2 (e.g. Erlam, 2006; Van Moere, 2012) and therefore is arguably more closely tied to one’s L2 processing efficiency than the other proficiency measures such as cloze tests or C-tests (Van Moere, 2012).
The 10-minute Chinese EIT is comprised of 30 progressively longer sentences that feature a wide range of vocabulary and grammatical structures. Participants were instructed to listen to the stimulus sentences and repeat as much as they could within a given time for each repetition. Their repetition performance was transcribed and evaluated using a 5-point scoring rubric developed by Ortega et al. (2002).
The background information questionnaire solicited information about learners’ Chinese learning experience, including information about when they started to study the language, how long they had spent in a Chinese-speaking region, and their self-rating of language use outside of the classroom.
V Results
1 L1 Typological influence on development of L2 TFS
Research question 1 explored the typological influence of learners’ L1 on acquisition of L2 Chinese TFS. We first examined the use of manner verbs versus path verbs in the oral narratives produced by the three baseline groups of English NSs, Japanese NSs, and Chinese NSs to explore the typological differences among speakers of S-framed, V-framed, and E-framed languages. Next, we performed the same analysis on the L2 Chinese narratives produced by the L1 Japanese learners and L1 English learners and then compared the L2 results with those by the three NSs groups. Figure 1 shows the mean percentage distributions of manner verbs versus path verbs used in the narratives from each group. As the percentage use of neutral verbs did not reflect the typological choice between S-framed or V-framed options and accounted for only a small portion of the verbs supplied, following Özçalışkan and Slobin (2003), its distribution was not illustrated in the figure.

Mean percentage distributions of manner verbs vs. path verbs.
In agreement with the established literature on motion event typology, S-framed English NSs used manner verbs in the main verb slots at a much higher rate of 59% than V-framed Japanese NSs at 11%. S-framed English NSs not only encoded manner more frequently than Japanese NSs, a tally of types of manner verbs further demonstrated that English NSs used richer and more fine-grained manner verbs than did Japanese NSs (English NSs: 20 types vs. Japanese NSs: 5 types). While V-framed Japanese NSs produced 14 different path verbs – which is the same as English NSs – they appear to be much more accustomed to encoding path via main verbs than English NSs (Japanese NSs: 89% vs. English NSs: 31%). As for E-framed Chinese, in terms of frequency of use, Chinese NSs supplied manner verbs at an even higher rate than English NSs (Chinese NSs: 69% vs. English NSs: 59%) but employed fewer types of manner verbs than English NSs (Chinese NSs: 15 types vs. English NSs: 20 types) and fewer types of path verbs than Japanese NSs (Chinese NSs: 9 types vs. Japanese NSs: 14 types). The result of a one-way ANOVA test on the (arcsine transformed) percentage use of manner verbs confirmed that there was a statistically significant difference in use of manner verbs across the three NS groups, F(2, 117) = 187.85, p < 0.001, partial η2 = 0.763. Post-hoc comparisons using the Tukey HSD test indicated that the three NS groups differed significantly from one other. In sum, S-framed TFS was the preferred linguistic strategy for both Chinese and English NSs, with the Chinese NSs supplying manner verbs at a rate higher than the English NSs, and V-framed TFS was favored by Japanese NSs. Because the characteristic mode in Chinese is S-framed (i.e. encoding manner in the main verb), as shown in the Chinese NSs group as well as Lu (1984) and Shi (2011), we will then focus on the learner groups’ use of manner verbs and S-framed constructions to explore how well they can adjust to the Chinese TFS.
With the baseline data established, we then explored the extent of influence learners’ L1 typological preference has on their L2 use of Chinese motion expressions. Out of the 390 motion verbs produced by L1 Japanese learners, 33% were manner verbs, and 60% were path verbs. The results patterned with Japanese NSs’ V-framed TFS (manner verbs: 11% vs. path verbs: 89%). Although L1 Japanese learners used a higher rate of manner verbs in L2 Chinese than did Japanese NSs in Japanese, they still showed a notable difference from the Chinese NSs at 69%, suggesting they potentially transferred their L1 V-framed preference to L2 Chinese. Nevertheless, contrary to what one may predict based on the L1-to-L2 typological influence found among the L1 Japanese learners, L1 English learners did not prefer S-framed strategies in their L2 Chinese production. Out of the 936 motion verbs produced, L1 English learners used a lower percentage of manner verbs at 42% and a higher rate of path verbs at 46%. Compared with English NSs’ S-framed TFS patterns (manner verbs: 59% vs. path verbs: 31%), L1 English learners’ use of L2 Chinese leaned toward V-framed means. In sum, neither of the learner groups demonstrated S-framed TFS as did the Chinese NSs. It appears that with both S-framed and V-framed options available in the target language, the V-framed options were widely adopted across both learner groups. There was no substantial difference found in terms of their suppliance of types of manner and path verbs between the two groups of learners.
Although neither of the learner groups exhibited target-like S-framed TFS in their mean percentage use of manner verbs, the results can potentially be attributed to the different proficiency between the two groups. To control the confounding factor of L2 proficiency, we conducted a one-way between-group analysis of covariance (ANCOVA) analysis to compare the (arcsine transformed) percentage use of manner verbs between L1 Japanese learners (n = 41) and L1 English learners (n = 80). There was no interaction effect between L1 type and L2 proficiency (F(1, 117) =2.662, p = .105, partial η2 = 0.022), suggesting that the assumption of homogeneity of regression slopes was not violated. After the influence of the covariate (i.e. L2 proficiency) was controlled and all the other premises were checked and met, a significant difference between the two learner groups on their percentage use of manner verbs was found, (F(1, 118) =14.759, p < .001, partial η2 = 0.111). L1 Japanese learners used significantly fewer manner verbs than L1 English learners. The results of ANCOVA analysis confirmed the typological influence of learners’ L1 on L2 learning. L1 S-framed English learners were more likely to adopt target-like S-framed L2 Chinese TFS than L1 V-framed Japanese learners. 3
2 Factors affecting development of target-like TFS
Combining the biographical data collected from the L1 Japanese and L1 English learners, research question 2 further explored what factor(s) play a major role in L2 learners’ development of L2 Chinese TFS. Specifically, we conducted a standardized multiple regression analysis to explore how much variance in the (arcsine transformed) percentage use of manner verbs can be explained by five factors, including L1 type (i.e. Japanese vs. English), L2 proficiency (measured in EIT score), age of acquisition, self-rating of language exposure, and length of in-country experience. Preliminary analyses showed that all assumptions of normality, linearity, homoscedasticity, and multicollinearity were met. When entering all five factors simultaneously, the six variables contributed to 47.7% of the variance in the use of manner verbs, (R2 = 0.477, F(5, 115) = 20.948, p < .001).
As can be seen in Table 2, L1 type and EIT score were the only two factors that made a significant contribution to the use of manner verbs at a p < 0.05 level, with L1 type recording a higher beta value (β = −0.76) than EIT score (β = 0.251). The results of the multiple regression analysis show that L1 type was the best predictor in terms of how well learners can adopt target-like L2 Chinese TFS, and L2 proficiency also made a significant contribution to the prediction of percentage use of manner verbs.
Regression model of predictors of use of manner verbs.
Note. 95% confidence intervals are reported in parentheses.
We adopted a two-way ANOVA to further explore the impact of L1 type (English and Japanese) and L2 proficiency on learners’ (arcsine transformed) percentage use of manner verbs. All assumptions were checked and met. The interaction between L1 type and L2 proficiency was not significant, F(3, 112) = 0.699, p = .554, partial η2 = 0.018. There was a significant main effect for L1 type, F(1, 112) = 16.859, p < .001, partial η2 = 0.131. L1 English learners used manner verbs significantly more frequently than L1 Japanese learners. A significant main effect for L2 proficiency was also found, F(3, 112) = 7.074, p < .001, partial η2 = 0.202.
Figure 2 shows a plot of the percentage use of manner verbs for the two learner groups across the five proficiency levels. Both groups’ use of manner verbs steadily increased between the Novice and High level. However, at the Advanced level L1 English learners (M = 64.64, SD = 11.23) continued to become more aligned with the Chinese TFS (M = 69.41, SD = 13.74), whereas the L1 Japanese learners (M = 36.61, SD = 17.02) took a slight decrease, compared to their performance at the High level (M = 40.11, SD = 19.16). Note that the L1 English learners at the High and Advanced levels (High: M = 55.2, SD = 12.48; Advanced: M = 69.41, SD = 13.74) were the only cohort that showed target-like S-framed TFS, in which they showed a preference to encode manner in the main verb slot.

First language (L1) type by second language (L2) proficiency.
3 Examining L1-to-L2 transfer: Use of Chinese motion constructions by L1 Japanese learners
Research questions 3 and 4 shifted the focus to L1 Japanese learners’ use and acquisition of L2 Chinese motion expressions. Table 3 summarizes the distribution of the different types of Chinese motion constructions used by the Chinese NSs and L1 Japanese learners.
Distribution of Chinese motion constructions: Mean values (as percentages).
Notes. V = verb. DC = directional complement.
As shown in Table 3, Chinese NSs predominantly adopted the S-framed Type 1 constructions to encode manner via the main verb and path via DC(s) at the highest rate of 63% when describing motion events. Together with Type 2 constructions, they used S-framed means (i.e. Type 1 plus Type 2) 80% of the time and V-framed means (i.e. Type 3 plus Type 4) 20% of the ti. L1 Japanese learners, by contrast, employed Type 1 constructions at the lowest rate of 19% among the four types of motion constructions and favored V-framed Type 3 and Type 4 options at a rate of 27% and 34%, respectively. Japanese learners’ preference to use Type 3 ‘Path Verb Only’ and Type 4 ‘Path Verb + deictic DC’ constructions that structurally resemble the top two most widely used motion constructions by Japanese NSs (see Type 5 ‘Path Verb Only’ constructions at 50% and Type 4 ‘Path Verb-te + deictic Path Verb’ constructions at 16% in Table 4) suggests that they transferred their L1 V-framed TFS to encode path via the main verbs in L2 Chinese.
Distribution of Japanese motion constructions: Mean values.
Note. Percentages do not total 100 due to rounding.
Next, we explored if such L1-to-L2 typological transfer gradually diminishes as the learners improve their overall L2 Chinese proficiency. Analysis of the Pearson’s correlation coefficient between learners’ use of the most characteristic Type 1 constructions and their EIT scores showed a moderate positive correlation between the two variables, r = 0.548, p < 0.001. Overall, Japanese learners became more comfortable using the more complex Type 1 constructions where both manner and path components are encoded. Nevertheless, close scrutiny of the scatterplot (see Figure 3) that illustrates the relationship between the two variables suggests that there were quite a few learners with varying EIT scores who opted to avoid any use of S-framed Type 1 constructions (i.e. cases of 0% manner rate). Moreover, the highest percentage use of Type 1 constructions reached a ceiling at the rate of 50% even for the most proficient learners at the High and Advanced levels, which was still considerably lower than that of Chinese NSs at 63%. In short, analysis of L1 Japanese learners’ use of motion constructions showed strong L1-to-L2 typological influence. Different from Chinese NSs, they were inclined to use V-framed constructions that are structurally similar to their L1 Japanese encoding means.

Relationship between Type 1 constructions and elicited imitation test (EIT) scores by first language (L1) Japanese learners.
4 Examining L2-to-L1 transfer: Use of Japanese motion constructions by L1 Japanese learners
Research question 5 explored whether L1 Japanese learners’ conceptualization of motion events in their L1 Japanese changes due to their experience of learning a typologically different L2 Chinese. L1 Japanese learners of L2 Chinese are required to adjust their L1 V-framed TFS and learn to divert their attention from path to manner so as to become more aligned with the characteristic strategy in Chinese TFS. Hypothetically speaking, their L1 V-framed TFS can be affected, especially when they become more proficient and accustomed to attending to manner details in their L2 TFS. That is, their encodings of manner information may increase when they describe motion events in their L1 Japanese. This would suggest signs of L2-to-L1 conceptual transfer, or integration of L1 and L2 TFS to some extent. In response to this inquiry, we explored whether the encodings of manner in Japanese produced by L1 Japanese learners of L2 Chinese differ from those by baseline Japanese NSs who had no exposure to the Chinese language.
We first ran an independent-sample t-test to compare the (arcsine transformed) percentage use of manner verbs in Japanese produced by the two groups. The result showed no significant difference between Japanese NSs and L1 Japanese learners of L2 Chinese, t(79) = 0.882, p = .380, Cohen’s d = 0.198.
Next, we examined the distribution of five types of Japanese motion constructions produced by the two groups. As shown in Table 4, encoding path in the sentence-final main verbs (i.e. Types 2, 3, 4 and 5) was the preferred method and accounted for 88% of the motion constructions, showing strong V-framed tendencies. S-framed means with manner encoded in the main verb (i.e. Type 1) was possible but used at only 11%. The manner aspects of motion events can also be highlighted through constituents other than the sentence-final main verbs via the conjunctive particle te, as seen in Types 2 and 3. Therefore, Types 1, 2, and 3 constructions all arguably emphasize manner of a motion event. If we see L1 Japanese learners of L2 Chinese show heightened attention to manner, we would expect an increase in these three types of manner-focused constructions. Analysis of their (arcsine transformed) percentage use of the three types of manner-focused constructions (i.e. Types 1, 2 and 3), compared to those produced by Japanese NSs, did not show an increase. We summed up the three types of manner-focused constructions and compared the two groups through an independent-samples t-test. The result confirmed no significant difference between Japanese NSs and L1 Japanese learners of L2 Chinese, t(79) = 0.941, p = .350, Cohen’s d = 0.212.
We further investigated the possibility of L2-to-L1 transfer by exploring if there was a positive correlation between Japanese learners’ use of the three manner-focused constructions in L1 Japanese and their use of the most characteristic S-framed Type 1 constructions in L2 Chinese. The result revealed no significant correlation between the two variables, r = −0.179, p = .264. That is, an increase in use of Type 1 constructions in L2 Chinese did not result in more attention to the manner details in their L1 Japanese TFS. In sum, there was no detectable L2-to-L1 typological transfer found. The experience of learning a typologically different L2 Chinese did not appear to affect Japanese-speaking learners’ L1 TFS, and their encodings of manner overall did not differ from Japanese NSs.
VI Discussion
The present study extended Wu (2016) on L1 S-framed English learners’ use of L2 Chinese motion expressions to investigate how L1 V-framed Japanese learners come to express motion events in L2 Chinese. Combining the data collected from both studies, the results show that the three NS groups had distinct strategies to describe motion events in their oral narratives. English NSs preferred S-framed means and used manner verbs at a rate of 59%, whereas Japanese NSs preferred V-framed means and used manner verbs at a much lower rate at 11%. Chinese NSs opted to use manner verbs at 69%, a rate higher than English NSs. The results agree with Shi (2011) and Peyraube (2006) that the Chinese language is shifting toward S-framed while the grammaticalized path satellite DCs can still function as full path verbs. For both groups of L2 Chinese learners, it was found that neither of them showed S-framed TFS in their L2 Chinese narratives. On average, L1 Japanese learners used manner verbs in 33% of the events, and L1 English learners at 42%. L1 Japanese learners used significantly fewer manner verbs than L1 English learners. We analysed how well the factors of L1 type, L2 proficiency, age of acquisition, self-rating of language exposure, and length of in-country experience can predict use of manner verbs in L2 Chinese. It was found that only L1 type and L2 proficiency made a significant contribution to the regression model, with L1 type being the best predictor. The fact that both the Chinese and English languages’ characteristic means to encode motion is S-framed is likely to play a facilitatory role in L1 English learners’ acquisition of L2 Chinese TFS. The L1 English learners’ use of manner verbs became more aligned with the Chinese NSs at the High and Advanced levels, which were also the only cohort that showed target-like S-framed TFS. The L1 Japanese learners’ use of manner verbs, by contrast, capped at the High level and consistently stayed lower than their L1 English counterparts at the same proficiency level. Analysis of use of motion constructions revealed that L1 Japanese learners preferred to used V-framed Types 3 and 4 constructions in their L2 Chinese, which was likely due to their structural similarities with the two most frequently-used Japanese V-framed Types 4 and 5 constructions. As for L2-to-L1 reverse conceptual transfer, there was no sign of L2-to-L1 transfer identified. L1 Japanese learners of L2 Chinese and Japanese NSs behaved in very similar ways in terms of their use of Japanese motion constructions and how much attention was allocated to manner information.
The findings gleaned from the study offer new insights into the nature of L2 development of TFS. By examining L2 acquisition of Chinese produced by both L1 V- and S-framed speakers, the results show that L1 typological influence plays a significant role in determining how well learners can develop target-like TFS in the long run, but its impact may be mitigated by L2 proficiency. As shown in the analysis performed on both types of learners, L1 S-framed English learners did not transfer their L1 S-framed strategies to L2 Chinese at the Novice to Intermediate level. Both L1 V- and S-framed speakers preferred to adopt the simpler V-framed constructions that retain only the core path information and leave out manner details (e.g. jìn-jiàoshì ‘enter-classroom’), as compared to the morpho-syntactically and conceptually more complex S-framing that uses the same path morphemes with an additional manner verb (e.g.
By examining the uses of Chinese S- and V-framed options that present different levels of ease of processing and by employing an EIT that has been adopted to measure one’s L2 proficiency and processing competence, the present study pushes a step closer to reveal how ease of processing of the target motion constructions may interact with learners’ L2 proficiency to affect their L2 motion expressions. However, it awaits future research to investigate how complexity of the L2 motion constructions, L2 processing competence, and size of manner verb lexicon shape the development of L2 TFS. It is also likely that the V-framing preference is observed because the set of Chinese path morphemes are more salient as full path verbs than as path satellites (DCs) in the learners’ developing grammar, especially when they have relatively limited experience with the characteristic L2 motion mode at the lower proficiency levels. These factors, possibly all correlated with L2 proficiency, remain to be empirically examined.
Finally, the results contribute to the study of bidirectional conceptual transfer in conceptualization of motion events. The present study found L1-to-L2 transfer, but not reverse L2-to-L1 transfer. Learning typologically closer Chinese for Japanese speakers does not require drastic alternation to their L1 TFS because both S- and V-framed means are grammatically correct and available in the language. Despite their high exposure to the target language, the L1 Japanese learners transferred their L1 V-framed TFS strategies to L2 Chinese, and the V-framed encoding means remained favored by them even among the most proficient learners. The robust L1-to-L2 conceptual transfer effect helps explain the lack of reverse L2-to-L1 transfer. The learners overall retained their L1 V-framed TFS when describing motion events in L2 Chinese, and the V-framed TFS remained ingrained in their L1 and L2 production. Note that the results are overall in agreement with Brown and Gullberg (2008) and Brown (2015), in which traces of reverse L2-to-L1 transfer were observed in gesture but not in speech.
VII Limitations and conclusions
There are two factors that could potentially limit generalizations based on the current study. First, when recruiting the baseline NS participants, we used the criterion of living in a foreign country for less than three years to ensure the baseline participants were functionally monolingual. However, up to three years of residence in another country could potentially lead to changes in the L1. It is suggested that future studies set up a stricter control for international residence, preferably less than one month, for the monolingual baseline groups. Second, while the narrative data were carefully coded and reviewed in the study, the coding process can be improved by having two raters code the data independently and reporting inter-coder reliability.
Despite these limitations, through investigating L2 acquisition of Chinese which straddles the S- and V-language types, the study not only sheds new light on the dynamic relationship between the factors of L1 typological influence and L2 proficiency but also illuminates their role in shaping the development of L2 TFS. Regardless of their L1 typological background, learners at lower proficiency levels are prone to adopt simpler V-framed constructions available in the language that may not be aligned with their L1 TFS patterns. The effect of L1 typological influence becomes more crucial in determining development of target-like TFS when the constraint of limited L2 processing competence is alleviated, typically seen in learners with more advanced proficiency. The results also suggest that reverse conceptual transfer is less likely to happen when there is minimal restructuring of L1 TFS found in learners’ L2 motion expressions.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
