Abstract
The present work examines the impact of code-switching (CS) on novel word learning in adult second language (L2) learners of Spanish. Participants completed two sessions (1–3 days apart). In the first session, they were taught 32 nonwords corresponding to novel creatures. Training occurred across 4 conditions: (1) a sentence in English only, (2) a sentence in Spanish only (the L2), (3) a sentence that contained CS from Spanish-to-English, (4) a sentence that contained CS from English-to-Spanish. Immediately after training, participants were tested on their ability to identify the newly trained words using a looking-while-listening paradigm in which videos of participants’ looking patterns were collected remotely via Zoom. In the second session, re-testing of the trained words was completed. In the first session, training in the English-only condition led to better initial learning compared to the other conditions. In the second session, the English-only condition still had the highest accuracy, but performance in the two CS conditions was significantly better compared to the Spanish-only condition. These findings suggest that CS during vocabulary training may aid the retention of newly acquired word-object relations in the L2, compared to when training occurs entirely in the L2. This work has important implications for theories of L2 acquisition and can inform instruction practices in L2 classrooms.
I Introduction
For millions of individuals around the world, acquiring a second language (L2) is common practice. Reaching proficiency in the L2 requires learning the phonology, the vocabulary, the grammatical structures, and the social rules of the target language, and can be challenging for both child and adult learners. However, adults often demonstrate greater difficulty throughout the learning process, and typically reach lower levels of overall L2 proficiency compared to child learners (DeKeyser, 2013; DeKeyser et al., 2010; Han and Selinker, 2005; Monner et al., 2013). There are several factors that have been linked to the attainment of L2 skills later in life, including: (1) differences in cognitive abilities between children and adults (Elman, 1993; Newport, 1990; Smalle et al., 2021), and (2) differences in the methods of instruction (Norris and Ortega, 2000; Spada and Tomita, 2010). While prior research related to these elements has greatly contributed to theories of L2 acquisition, many questions remain regarding which factors might make it easier or harder for adult learners to acquire a second language. National guidelines for foreign language instruction in the United States encourage instructors to speak primarily in the target language, recommending its use at least 90% of the time in the classroom (American Council on the Teaching of Foreign Languages, n.d.), yet teachers and students alike frequently mix both the foreign and the native language during classroom interactions (Levine, 2003; Liebscher and Dailey-O’Cain, 2005; Sun et al., 2020; Thompson and Harrison, 2014), and it is unclear how the mixing of the two languages influences L2 learning. To address this question, the present work examines how language mixing (also known as code-switching - CS) contributes to novel vocabulary learning in novice adult L2 learners of Spanish.
Though understudied, the presence of CS during foreign language acquisition is not an unusual phenomenon. In fact, there is considerable data suggesting that CS is common among balanced-bilingual speakers (Auer, 1999) and L2 learners (Levine, 2003; Liebscher and Dailey-O’Cain, 2005; Thompson and Harrison, 2014). Broadly speaking, CS involves the use of two or more languages in a single sentence or conversation, and is used by speakers to achieve different communicative purposes (Auer, 1984, 1995, 1998, 1999). As discussed by Auer, CS can serve social or discourse-related functions, such as emphasizing a comment or indicating a change in the topic of discussion. Additionally, CS can be used to benefit the individual (i.e. to achieve a participant-related function). One example of this is to fall back into the language the speaker is most comfortable using (Liebscher and Dailey-O’Cain, 2005; Sarkis and Montag, 2021). When a speaker code-switches, there are two main manners in which CS is structurally used in utterances. These are called intra- and inter-sentential code-switches. When an individual uses intra-sentential CS, the code-switch is typically a single word insertion (e.g. ¿A ti te gustó la
Although CS is commonly encountered in dual-language environments, there is evidence suggesting that the presence of CS (in both written text – during reading, and in the spoken modality – while listening) might create added difficulties when completing certain language-related tasks. Several studies have relied on self-paced reading tasks to examine how CS is processed. In these studies participants saw sentences that either contained CS or were in a single language and read them silently. Overall, CS has been found to increase cognitive demand during silent reading compared to when sentences were presented in a single language (Moreno et al., 2002; Ng et al., 2014; Proverbio et al., 2004). Furthermore, factors such as the naturalness of the code-switch (Guzzardo Tamargo et al., 2016; Guzzardo Tamargo and Dussias, 2013; Valdés Kroff et al., 2018) as well as the direction of the code-switch (Bultena et al., 2015) impacted how quickly the participants were able to read code-switched sentences – as determined by eye-tracking data, or through a moving window paradigm where the sentence was presented to the participant one word at a time, and the participant controlled the rate of word presentation. The time taken to read code-switched sentences was taken to indicate ease of processing CS. Participants were faster when reading CS that were more likely to be used in natural conversation, than when reading CS that infrequently occurred in speech (Guzzardo Tamargo et al., 2016; Guzzardo Tamargo and Dussias, 2013; Valdés Kroff et al., 2018). Additionally, when the CS occurred from the participant’s dominant (i.e. the language learned first and most often used) into their non-dominant language (i.e. the language learned upon entry into the school system), reading times were slower than when the code-switch occurred in the opposite direction (Bultena et al., 2015).
In addition to self-paced reading studies, previous work has examined whether participants’ word recognition (Byers-Heinlein et al., 2017; Valdés Kroff et al., 2018) is affected when listeners hear sentences that contain CS. Byers-Heinlein and colleagues (2017) presented adults with a word recognition task that relied on eye-gaze to calculate the proportion of looking time to a target object on the screen, as a measure of accuracy in recognition. Looking times to the target object were significantly longer (i.e. greater accuracy) when sentences were heard in a single language compared to when CS was present – specifically when the code-switch occurred from the participant’s dominant to the non-dominant language. There were no differences in accuracy when the CS occurred in the opposite direction. Additionally, this study looked at pupil diameter during test trials, as changes in this measure have been linked to variation in cognitive load (Beatty, 1982; Mathôt, 2018). In the Byers-Heinlein study, increased pupil sizes were observed in trials that contained CS compared to when sentences were presented in a single language, which according to the authors suggested that processing code-switched sentences resulted in greater cognitive demands. This was independent of the direction of the CS. In another study, Valdés Kroff and colleagues (2018) examined bilingual adults’ word recognition via mouse click. Participants who were in environments where CS was more common were more accurate at clicking on the target item when CS was present, than participants who were not as frequently exposed to CS.
Findings from electrophysiological work align with the behavioral data discussed above. For example, Moreno et al. (2002), measured event-related potentials (ERPs) during sentence reading. When CS was present in the sentences, ERP components such as a Left Anterior Negativity (LAN) and a Late Positive Component (LPC) were present. These types of responses have been linked to the processing of a surprising or unanticipated event, and are thought to increase processing demands for the listener (Friederici, 2002; Tanner and Grey, 2017). Furthermore, according to data from Ng et al. (2014), the word class in which the CS takes place in the sentence is also important, with code-switches at the noun level leading to greater N400 effects compared to when the CS occurred in verbs. N400 effects are believed to be present when word meanings are being contextualized (Hagoort et al., 2004), suggesting that participants were experiencing higher processing costs when asked to integrate code-switched nouns into the rest of the sentence. Lastly, some studies have provided neural evidence supporting the directionality effect discussed earlier, with larger ERP effects identified when CS occur from the dominant to the non-dominant language (Litcofsky and Van Hell, 2017; Proverbio et al., 2004), compared to the opposite direction.
Taken together, the previous behavioral and ERP data suggest that during reading and auditory comprehension tasks, the presence of CS leads to greater processing costs. These costs may be associated with language activation, particularly in the studies that have found cost asymmetry depending on the direction of the CS. There is little work examining why CS leads to processing costs in the area of language comprehension and reading; the available evidence that might help explain this phenomenon comes primarily from studies that focused on speech production. Specifically, previous work suggests that CS requires monitoring language use and quick activation and inhibition of the two languages in conversation (Green, 1998; Meuter and Allport, 1999; Sarkis and Montag, 2021). As Byers-Heinlein et al. (2017) state in their article, this process likely increases cognitive load as speakers must control each of their languages and be prepared to rapidly change the activated language (Byers-Heinlein et al., 2017).
Less in known about the role that CS might play during other types of linguistic tasks – in particular, during vocabulary learning. This is an area that needs to be further studied given that the steps that listeners undergo during word learning and word recognition are different. For example, recognizing a familiar word requires in-the-moment processing of the speech stream or text, and then accessing the representation of a target word that has already been encoded in memory. On the other hand, learning a novel vocabulary item requires creating a new word-object representation and storing that representation so that it can be accessed later on. It is, therefore, possible that CS may have a different impact on how linguistic information is processed across the two types of tasks.
To date, the majority of work examining the role that CS plays on language learning comes from studies with young children, and offers insight as to how CS impacts word learning during childhood. Some studies conducted with adults have begun to examine how teacher’s CS patterns in classroom settings impact development in the L2, but there are limitations associated with this work. In order to best review the relation between CS and language acquisition, we summarize previous findings with both children and adults.
A common scenario in which language learners are regularly exposed to CS during early childhood is when they are growing up in bilingual environments. Several researchers have studied how the frequency of parental CS relates to vocabulary size in infants and toddlers. These studies have generated mixed results, sometimes reporting that CS has negative (Byers-Heinlein, 2013), neutral (Place and Hoff, 2011, 2016), and positive associations (Bail et al., 2015) with total vocabulary size. The variations in findings may be due to the different ways the studies quantified parental CS behaviors, as this was sometimes determined via parental report (Byers-Heinlein, 2013), language diaries and questionnaires (Place and Hoff, 2011, 2016), or through direct observation of parent-child interactions (Bail et al., 2015). But these findings are all with very young children, who are acquiring multiple languages simultaneously.
It is important to consider differences in vocabulary learning between bilinguals (who acquired both languages early in life) and L2 learners (who acquired a native language early on and are acquiring an L2 later). Theories such as the revised hierarchical model (Kroll et al., 2010; Kroll and Stewart, 1994), suggest that for L2 word learning, the link between the lexical form in the L2 and the underlying conceptual representation of that word are initially weak. Jiang (2000) poses that when L2 learners begin to acquire words in the L2, they rely heavily on the native language to infer word meaning. For example, if a native Spanish speaker was learning English as an L2, and they heard the word ‘cat’, they would have to translate the L2 word ‘cat’ into the native language (i.e. ‘gato’), in order to access the conceptual representation, or the meaning behind the word. As speakers become more proficient, the links between the words in the L2 and the conceptual meaning become stronger, and they then no longer need to rely as heavily on the lexical form in the native language to extract word meaning. For bilinguals (who have acquired both language systems simultaneously) lexical and conceptual links for words in each language are developed in tandem, and hence there is no need to rely on any one language in order to determine word meaning in the other (Kroll et al., 2010; Kroll and Stewart, 1994). Given the differences between theories surrounding bilingual and L2 learner’s lexicons, it is important to consider these groups separately when looking at word learning.
To our knowledge, a limited number of studies have examined the impact of CS on L2 learning in children (Read et al., 2020; Sun et al., 2020). Read et al. (2020) examined how toddlers and preschool aged children who were learning an L2 benefitted from encountering the names of novel animals in story books that were either entirely in the in the L2 or contained CS from the native language to the L2. The authors found that for the younger children, CS did not impact their learning or retention of the novel words. However, for the preschool-aged children, CS facilitated novel word retention. In another study, Sun and colleagues (2020) examined CS patterns in L2 classrooms in Singapore. Classroom video recordings were used to analyse teacher CS behaviors, and the students completed tasks that measured: cognitive flexibility, non-verbal attention, and receptive vocabulary skills. The authors found that most instances of teacher CS were habitual (i.e. teachers’ CS was unrelated to classroom instruction and often involved switching into their heritage language). Critically, there was no relation between the frequency of teacher’s CS and student language outcomes, but there was a positive relation between the frequency of intra-sentential CS and student performance on the cognitive flexibility task, which specifically measured executive functioning skills – such as the child’s ability to switch tasks.
There is some additional work with adolescents examining how CS impacts word learning during classroom instruction (Zhang and Graham, 2020a, 2020b). In both studies, high-schoolers were taught novel words either (1) entirely in the L2, with the word and definition presented in the L2,or (2) with CS from the L2 into the native language to provide a definition for the word. In some CS conditions participants were provided with additional context about word use and how it differed from use in the native language. The use of CS, particularly when greater context about L2 word use was presented, led to better word learning than when the L2 alone was presented. However, this benefit did not remain at delayed testing which occurred 3–4 weeks later. The adolescent participants included in this work were intermediate L2 learners with at least 7 years of previous English instruction.
Similar studies have been conducted with intermediate and advanced adult English learners in university classrooms. These studies examined whether CS from the L2 to the native language benefitted novel-word learning in comparison to instruction entirely in the L2 (Lee and Levine, 2020; Tian and Macaro, 2012; Zhao and Macaro, 2016). All of these experiments relied on a word learning task, where the target word was presented in the L2, and the word definition was provided either in the native language (CS condition) or entirely in the L2. Students who received training in the CS condition demonstrated better learning of the L2 words at testing than those in the L2-only condition. Additionally, studies that stratified participants based on L2 proficiency showed that learners with less L2 proficiency (in comparison to more proficient learners) acquired the novel words better in the CS condition than in the L2-only condition (Lee and Levine, 2020; Tian and Macaro, 2012). This suggests that less skilled L2 learners relied more heavily on the native language, and therefore showed a greater boost for word learning with CS. On the other hand, more proficient L2 learners appeared to rely less on the native language to infer L2 meaning, hence, showing less drastic differences in performance between the CS and L2 learning conditions. (Lee and Levine, 2020; Tian and Macaro, 2012). While these results demonstrate that CS is beneficial for L2 learning, there are some limitations. First, the participants had been learning English as an L2 for several years, and therefore were not considered to be novice. Second, this work has been conducted with L2 learners of English, and not those learning a language other than English as an L2. This is contrastive to a population of adult L2 learners in a country like the U.S., where their first exposure to an L2 and CS might occur in adulthood.
For children raised in bilingual environments, exposure to CS occurs as part of everyday conversation, and is normalized early on in life; but this is not the case for adults. Many adult learners may have grown up in monolingual settings, and their first exposure to CS may be in a foreign language classroom (Thompson and Harrison, 2014). How adult learners process CS while attempting to acquire an L2 may be different from what is observed in children, given that language learning during adulthood is a different process from learning languages as a child. First, advances in cognitive capacity take place throughout childhood (Gathercole, 1999), leading to children (but not adults) developing language and cognitive skills in parallel. This means that adult learners must face the process of acquiring an L2 with advanced cognitive capacities. It would appear that this should lend itself as an advantage to the adult learner, but their lower overall achievement in the L2 indicates that this is not the case (DeKeyser, 2013; DeKeyser et al., 2010; Han and Selinker, 2005). In fact, the ‘Less is More Hypothesis’ (Newport, 1990) suggests that approaching language with a limited cognitive system facilitates acquisition, and that greater attentional control may be negatively related to performance during implicit learning (Smalle et al., 2021).
In addition to developments of cognitive capacity, the type of instruction contributes to differences in language learning in children compared to adults. Particularly, adults are not as successful as children at acquiring a language via implicit learning (e.g. language immersion in more naturalistic interactions), as opposed to formal or explicit instruction, which involves definitive explanations of the language, such as the grammar or vocabulary (see Norris and Ortega, 2000 and Spada and Tomita, 2010 for meta-analyses on this topic). These differences have been discussed in theories such as the Noticing Hypothesis (Schmidt, 1990, 1993), which poses that adult L2 learners need to explicitly attend to linguistic information in the L2 in order to learn it, given that language learning in the L2 is less likely to be implicitly acquired. With this in mind, it is important for information to be salient/emphasized during instruction.
Given the differences in the language-learning process between children and adults, it is important to understand whether exposure to CS would affect adult L2 language learning in the same way described in the pediatric literature. Few studies have examined this topic. Those that have, have focused on measures of teacher beliefs about the use of CS, and how they describe their CS during instruction. In one study, researchers examined teachers’ CS behaviors in English as a foreign language classrooms at a university in Indonesia (Cahyani et al., 2018). Videos of classroom lessons were coded for instances where CS was used to (1) assist in conveying meaning or (2) facilitate communication. Analyses of these items, along with teacher interview responses appraising their own use of CS revealed that teacher CS was frequently used to support student learning and affirm student behavior by promoting clarity of communication. Similar findings were reported in a study looking at other English-learning classrooms in Indonesia (Suganda et al., 2018). Here, classroom lessons, teacher interviews and student responses to a questionnaire regarding the use of CS in the classroom were analysed. Students generally had a positive perception of CS in the classroom, and believed it supported the learning process. Teacher interviews indicated that teachers code-switched based on what they felt would be most engaging for their students, or would best facilitate learning. Additionally, data from classroom recordings in this same study indicated that CS was often used to clarify meaning, switch topics, attract student attention, or to represent cultural use of the L2. Lastly, a study conducted in Pakistan examined teacher’s CS patterns through observation of teacher lectures that were transcribed and coded for instances of CS (Bhatti et al., 2018). These data suggested that teachers often CS out of habit, to check in on student’s understanding of the topic, to explain difficult concepts, or to address student behavior. Taken together, these findings emphasize not only that CS occurs frequently in adult classrooms, but also that teachers self-report using CS as a teaching tool to promote L2 learning.
There are, however, limitations associated with these previous studies. First, this work has only relied on observational data and subjective analyses of the beliefs of students and instructors regarding the use of CS. Furthermore, prior work has focused on the acquisition of English as the L2 and has not been extended to other languages. To our knowledge, there are no empirical studies conducted with adults that directly measure how exposure to CS might affect different areas of L2 acquisition (e.g. vocabulary learning). With this in mind, the present work set out to investigate whether CS plays a role on vocabulary learning in the L2, and whether such an effect can be directly measured in adult learners of Spanish.
In the current study native speakers of English who were novice L2 learners of Spanish completed a word learning task. Participants were taught the names of novel animal-like creatures in sentences that were entirely in English (e.g. ‘Look at the table, on top of the table is the
Research question 1: Does CS influence vocabulary learning in novice adult L2 learners of Spanish?
Research question 2: Does the direction of the CS affect how well the novel words are learned?
Research question 3: Does the presence (or absence) of CS lead to differences in the retention of newly-learned words?
II Methods
1 Participants
Our sample included a total of 40 college students (M = 19.33 years, SD = 2.34. Male = 10) enrolled in 100-level Spanish classes at the University of Delaware. Data from an additional 29 participants were excluded due to (1) technical issues (n = 18) – which included problems with internet connectivity or inability to obtain codeable high-quality videos, (2) failure to complete both study sessions (n = 4), (3) failure to demonstrate learning in the first session (n = 5) – as measured by extremely low accuracy during testing at the first visit, or (4) insufficient knowledge of the familiar words (n = 2) – which was determined via a translation task. Of the participants included, 30 identified as White, 2 as Asian, 3 as African American, 1 as Hispanic, 3 reported being of mixed race and 1 participant declined to respond.
Students were eligible to participate in the study if they were native English speakers (based on self-report), who did not speak Spanish or any other romance language with proficiency, and who reported no history of language delays or disorders. All participants were enrolled in one of the following courses: Spanish 105, Spanish 106 or Spanish 107. These courses provide a basic knowledge of the Spanish language. Additionally, participants had minimal previous exposure to Spanish as a second language (M = 2.75 years of coursework, SD = 1.37). Lastly, in order to participate in the study, students needed to have a stable internet connection, a computer with a screen that was at least 12 inches wide, and a webcam. To be included in the analyses participants had to complete both testing sessions.
2 Stimuli
A total of 64 pictures of ‘novel creatures’ were created using the software SPORE™ Creature Creator and were assigned nonwords in English or in Spanish depending on the experimental condition (see additional information in the next paragraph regarding the characteristics of the nonwords). The creatures were designed to look animalistic without too closely resembling a real animal. Half of the creatures were assigned randomly to be target items, that is, creatures that were given a novel name that was taught to participants. The other 32 creatures were randomly assigned to be foils, and were never labeled during training. All novel creatures were presented alongside a familiar item (e.g. a table or house); the familiar items were stock images with white backgrounds. Familiar items were only used to provide a reference for the novel animal during training. Participants were expected to identify the reference item and use this knowledge to learn the name of the novel creature (i.e. the novel creature was on one of the familiar items). All stock images were presented in grayscale so that they did not overshadow the novel creatures (see Figure 1). Four creature/referent pairs were presented at once. Each pair was in a separate quadrant of the screen. Items were spaced out as much as possible (i.e. the pairs were closer to the perimeter of the screen rather than the center), so that the shift in the looking angle would be greater between objects, which facilitated offline coding.

Sample stimuli presented during training and test trials.
A Spanish–English bilingual female speaker recorded the auditory stimuli. Training items contained a sentence frame either in English or in Spanish, which instructed the participants to look towards one of the familiar items on the screen, and then told them the name of the novel creature accompanying that item (e.g. ‘Look at the ____, on top of the ____ is the ___’ for English, and ‘Mira la ____, sobre la ____ está la/el ___’). To ensure the same level of intelligibility, prosody, and speech rate of the carrier phrase across trials, one token was selected, and the familiar and novel words were cross spliced into the phrase. Given the structure of the sentence frame, the code-switch always occurred at the last word of the sentence. The speaker recorded the familiar/referent words in both English and Spanish, and the novel nonwords separately. Thirty-two nonwords were presented in the study, sixteen followed English phonotactic rules (e.g. tot͡ʃɪd, dʌget, titεks) and sixteen followed Spanish phonotactic rules (e.g. rit͡ʃa, laɲa, poɾa). Of the Spanish nonwords, eight had feminine grammatical gender (i.e. ending in -a) and eight had masculine grammatical gender (i.e. ending in -o). As we used nonwords (that did not exist in either language), the code-switch was largely based on the phonotactic differences between the sentence frame and the target word. Specifically, the Spanish nonwords contained phonemes that are not used in English (e.g. trilled r /r/), and English words contained phonemes not used in Spanish (e.g. retroflex r /ɹ/). These differences in phonotactic patterns are sufficient for indicating a code-switch as studies have shown speakers note language differences based on phonotactics (Weber and Cutler, 2006), and there are notable contrasts between these rules in English and Spanish (Prieto et al., 2012).
In addition to the nonwords, sixteen familiar items were selected from a vocabulary list included in a textbook that was used to teach Spanish 105 at the time the study took place (Goodall and Lear, 2015). These words were recorded in both Spanish and English, totaling 32 reference words. Additionally, Spanish familiar words were presented in blocks based on whether the words had feminine or masculine grammatical gender. This was done so that participants could not simply use gender cues to determine the familiar word, and instead needed to have knowledge of the item. Test trials, on the other hand, included the name of the target creature, which was produced three times without a sentence frame (e.g. tot͡ʃɪd! tot͡ʃɪd! tot͡ʃɪd’). Due to previous evidence suggesting that processing costs associated with the presence of CS are present during word recognition (Byers-Heinlein, 2017; Litcofsky and Van Hell, 2017; Proverbio et al., 2004), the sentence frame was omitted at test.
The duration of each trial was 7 seconds. This was true for both the training trials and the test trials. In training trials images appeared on the screen 500 ms prior to the onset of the auditory stimuli, and the onset of the novel word occurred 2.5 seconds after the onset of the phrase. During test trials, the onset of the first repetition of the target word occurred 1 second after the onset of the trial. Training and test trials were matched to have the same perceived amplitude. All recordings were created using a Shure MV51 microphone at a 44.1 kHz sampling rate, 16-bits precision, inside a sound-attenuated booth.
3 Procedure
Participants completed the study remotely from home, with an experimenter that led the appointments via a Zoom call. They were asked to complete the study alone in a quiet room to avoid distractions (e.g. roommates talking, or television in the background). All experimenters followed a written protocol during testing to ensure consistency across appointments. Additionally, experimenters received training on how to use Zoom and troubleshoot common computer issues.
During the Zoom call, before beginning the experiment, participants were administered a working memory measure, in this case a digit span task (Wilde et al., 2004). This was used to capture any potential differences in participants’ ability to remember words, which could contribute to performance during the word learning task. After completing the digit span, light, camera, and audio checks were completed. As part of this process, participants watched a 30 second video of a whale, which contained an alternating white and black background. The experimenter watched the participant to ensure that lighting changes could be seen on the participants’ face from the light to the dark background. These contrasts would be used to determine the start and end of trials during offline eye-gaze coding. If the lighting contrasts were not clearly visible to the experimenter, the participant was instructed to adjust the lighting in the room (e.g. by closing the curtains, or turning a light on/off), and the video was played a second time. Additionally, the sample video played music that was presented at the same intensity level as the study stimuli. Participants were instructed not to wear headphones and were asked to adjust their sound to a comfortable listening level. Once all checks were completed, participants were provided with the link to access the study running website, in this case Gorilla Experiment Builder (Anwyl-Irvine et al., 2020), which was used to design and host the study.
Prior to starting the study in Gorilla, participants were instructed to begin recording themselves using the webcam on their computer and a built-in software (e.g. Photobooth for Macs and Camera for Windows). The local recording eliminated the concern of video lag, which may have been present in a remote Zoom recording. Participants shared their computer screen so that experimenters could watch the study and to minimize technical issues. Experimenters then turned off their cameras and the participants proceeded to complete the word learning task. After completion of the study, experimenters gave participants instructions to upload their videos to a secure server, where recordings could later be accessed by members of the research team.
a Visit one
During the first testing session, participants completed a word translation task. In this task the names of the familiar/reference items were written in Spanish, and they were asked to translate them by typing the English equivalent of the Spanish word. This step was done using Gorilla. After translations were provided, participants progressed into the word learning portion of the study, which was divided into four blocks. Each block represented a language condition (English Only, Spanish Only, English to Spanish CS, Spanish to English CS), and began with eight training trials. In these eight trials, two identical training trials were used per novel object to ensure participants had enough exposure to the item and its name. Immediately after the eight training trials, participants were tested on their ability to identify trained objects, as determined via looking behavior. During test trials, four creatures appeared on the screen, this time without the familiar item (see Figure 1). Two target and two foil creatures were present per trial. This sequence repeated for four new creatures and that concluded a block. Eight novel creature names were taught per block and each block followed the same pattern but with a different language condition.
The following components were counterbalanced across participants: (1) the order of the language condition (i.e. which condition was presented 1st, 2nd, etc.), (2) the location of the target item, and (3) the gender of the familiar and novel words between the Spanish only and CS conditions (i.e. for some participants the items with female grammatical gender appeared in the CS conditions and in others it appeared in the Spanish-only condition). Additionally, the pairing of the creature and the familiar/reference item, as well as the pairing of creatures with novel words were also counterbalanced.
Lastly, a black screen with a fixation cross that lasted for 3 seconds was included in-between each trial, to center the participants attention and provide a lighting contrast to indicate the end of a trial during offline eye-gaze coding.
b Visit two
The second visit was shorter than the first and occurred 1–3 days after the initial session (M = 44 hours apart, SD = 14.77), once again via Zoom. In this later visit, participants were re-tested on the names of creatures learned in visit one, without any additional training. This visit consisted of presentation of the 32 test trials. Participants saw the test trials in the same order that they were presented during visit one (i.e. if tot͡ʃɪd was the first test trial in visit one, it was also the first test trial in visit two). Gorilla Experiment Builder was once again used for study presentation. We gave a window of 1–3 days to (1) lower attrition rates by allowing participants more flexibility in scheduling their second visit, and (2) 24 hours has been found to be enough time to demonstrate retention effects in word-learning task (e.g. Kurdziel and Spencer, 2016).
4 Data coding
Participant videos were downloaded from the secure server and coded offline on a frame-by-frame basis by two trained coders using Datavyu (Datavyu Team, 2014). Reliability between coders was checked across all coded files. A third coder recoded any trials where a discrepancy greater than .5 seconds was present between the initial two coders. A total of 14.3% of trials needed a third coder. To examine performance during testing, we analysed accuracy, which was defined as the proportion of time that participants looked at the target image averaged over a window of 5,860 ms, which started 240 ms after the onset of the first repetition of the target word, and went to the end of the trial. This window of analysis was used as it has been found that it often takes between 200–250 ms for adults to program a saccade (Deubel and Schneider, 1996; Walker et al., 2006). We chose to continue to examine looking behavior to the end of the trial, as time-course data of participant fixations indicated that the participants were maintaining focus to the target item until the trial ended (see Figures 2 and 3).

A graph showing time course of accuracy data for visit 1 for the analysis window.

A graph showing time course of accuracy data for visit 2 for the analysis window.
Additionally, the translation task was scored with one point awarded to a correct translation and 0 points to an incorrect translation. Translations were counted as ‘correct’ if the participants response would lead them to a correct item identification (e.g. desk instead of table, or child instead of boy). At maximum participants could receive 16 points and at minimum 0 points.
III Results
Before examining accuracy scores, we looked at participant performance on the digit span and the translation task. In the digit span, participants were able to recall approximately 6 numbers in the forward condition (M = 6.4, SD = 1.32) and about 5 numbers in the backwards condition (M = 4.73, SD = 1.13). These results are comparable to what has been reported in other adult studies that used the same task (e.g. Iverson and Tulsky, 2003). Furthermore, we did not find that performance on this measure was related to performance on the word learning task. We then examined performance on the translation task. Participants were able to translate an average of 68.9% of the Spanish words, indicating that they knew 11 of the 16 words (SD = 22%).
Initial evaluation of the accuracy data for visit 1 indicated that performance was highest for the English-only condition (M = .74, SD = .20), and similar across the English-to-Spanish (M = .60, SD = .17), Spanish-to-English (M = .59, SD = .19), and Spanish-only (M = .57, SD = .21) conditions. Two-tailed single sample t-tests indicated that participants’ accuracy was significantly above chance (in this case .25) when training occurred in English only (t(39) = 15.73, p = < .0001, Cohen’s d = 2.49), Spanish only (t(39) = 9.63, p =< .00001, Cohen’s d = 1.52), CS from English-to-Spanish (t(39) = 12.67, p =< .0001, Cohen’s d = 2.00), and CS from Spanish-to-English (t(39) = 11.21, p =< .0001, Cohen’s d = 1.77). Given that each of these differences had a large effect size, this indicates that there is a substantial difference between the score and chance. This suggests that regardless of the training condition, participants were able to reliably learn the names of the novel objects. Next, a one-way within-participants ANOVA revealed a significant main effect of language condition, with a large effect size (F(3, 117) = 13.34, p =< .001, ηp2 = .26). Pairwise comparisons indicated that significant differences in accuracy were only present between the English-only condition and the other three conditions: Spanish-only (t(39) = 5.78, p =< .0001, Cohen’s d = .91), English-to-Spanish (t(39) = 4.39, p =< .0001, Cohen’s d = .69), and Spanish-to-English (t(39) = 5.05, p =< .0001, Cohen’s d = .80) respectively (as evidenced from Cohen’s d, all effect sizes were medium to large). These results suggest that during the initial testing session, when immediate learning of the novel words was measured, hearing sentences in English alone during training led to the best performance at test compared to the other conditions. This is not surprising, given that English was the native language of the participants. Interestingly though, our data from this first visit suggest that there were no differences in immediate learning when training occurred only in Spanish, compared to the two conditions that contained CS.
As a next step, we examined whether the presence of CS would impact the retention of the newly-learned words, by examining accuracy scores from the second testing session. Accuracy was highest for the English-only condition (M = .58, SD = .19), followed by the two conditions that contained CS (English-to-Spanish: M = .43, SD = .18; Spanish-to-English: M = .43, SD = .19), and lowest for words that were trained in the Spanish-only condition (M = .34 SD = .18). Once again, two-tailed single sample t-tests indicated that participants’ performance was above chance for the English-only (t(39) = 11.31, p =< .0001, Cohen’s d = 1.79), Spanish-only (t(39) = 3.21, p =< .01, Cohen’s d = .51), English-to-Spanish (t(39) = 6.10, p =< .0001, Cohen’s d = .97), and Spanish-to-English (t(39) = 5.99, p =< .0001, Cohen’s d = .95) conditions. All of these effect sizes were medium to large, indicating that the scores were substantially different than chance. A one-way ANOVA revealed a significant main effect of language condition, with a large effect size (F(3, 117) = 15.81, p =< .001, ηp2 = .29). Post-hoc pairwise comparisons showed that accuracy for retention was significantly lower for the Spanish-only condition compared to the two CS conditions (English-to-Spanish: t(39) = 2.79, p =< .01, Cohen’s d = .44; Spanish-to-English: t(39) = 2.83, p =< .01, Cohen’s d = .45), with medium effect sizes, and no significant difference in performance between the two conditions that contained CS was present (t(39) = .75, p = >.05, Cohen’s d = .12). These analyses suggest that participants’ retention of the novel words was better when the training occurred in sentences that contained CS, compared to when words were taught exclusively in Spanish. Furthermore, the directions of the CS (Spanish-to-English vs. English-to-Spanish) did not seem to play a role. Lastly, the English-only condition continued to yield significantly higher accuracy compared to the other three conditions (Spanish-only: t(39) = 7.38, p =< .0001, Cohen’s d = 1.17; English-to-Spanish: t(39) = 4.24, p =< .001, Cohen’s d = .67; Spanish-to-English: t(39) = 4.34, p =< .0001, Cohen’s d = .69), and had medium effect sizes, which again was not surprising since participants had considerably more experience acquiring and retaining word-object relations in their native language.
IV Discussion
The present work examined the influence of CS on novel word learning in a group of college-aged Spanish L2 learners. CS is common in foreign language classrooms (Liebscher and Dailey-O’Cain, 2005; Thompson and Harrison, 2014), however, to our knowledge no other studies have directly examined the impact of CS on vocabulary learning in adult L2 learners. To explore how CS may impact word learning, participants completed a word-learning task, in which novel word-object pairs were trained in four experimental conditions: (1) entirely in the participant’s native language (English), (2) entirely in the L2 (Spanish), (3) code-switched from English to Spanish, and (4) code-switched from Spanish to English. Participants were tested on their ability to identify the word-object relations shortly after training (to examine initial learning), and re-testing occurred 1–3 days after training, serving as a measure of retention.
Accuracy during the first testing session suggested that participants were able to learn the novel words in all experimental conditions. Unsurprisingly, training words entirely in English (the native language) resulted in greater initial learning based on looking accuracy. However, the main goal of the study was to examine the role of CS during L2 learning. Training in the other three conditions (i.e. entirely in Spanish, and both CS conditions) resulted in comparable performance at visit 1. This suggests that initial word learning may not be impacted by CS and in fact, may be comparable to learning that takes place when the sentence is solely in the L2.
Another question that we examined was whether the presence of CS might influence retention of newly acquired words (compared to when no CS was present). During the second visit, participants also demonstrated retention of the novel word-object pairings. Across the board, accuracy scores went down for all conditions between testing session 1 and 2. This was expected given that it is more difficult to identify word-object relations that were trained 1–3 days earlier, compared to immediately prior to testing. The critical point however, was whether or not at the later time point the same pattern of results would be present (i.e. if retention would be affected equally across conditions). Once again, participants’ accuracy was highest when training occurred in the English-only condition – suggesting that there was better retention when all information was provided in the native language. Interestingly though, accuracy scores from this second visit were different for words that were trained in the code-switched conditions compared to when training occurred entirely in Spanish. Specifically, participants recalled novel words better when the training contained CS sentences, than when they were presented in the Spanish-only condition. No significant differences were present between the CS conditions. In other words, while CS did not facilitate nor hinder initial word learning compared to when training was entirely in the L2, it did facilitate the retention of the newly acquired word-object pairs, and this was true regardless of the direction of the CS.
One possibility is that the code-switch may increase the phonological salience of the target word. Specifically, if participants are learning an L2 that has a different phonological system, the sound variations between the sentence frame and the CS word in a sentence like ‘Look at the table, on top of the table is the roɲo.’ will likely create a noticeable contrast, facilitating parsing of the novel word in the sentence. Aiding in parsing may make it easier for the learner to detect and encode the novel vocabulary item. This explanation aligns with the ideas presented as part of the Noticing Hypothesis (Schmidt, 1990, 1993), which outlines the need to highlight/emphasize linguistic information in order for it to be successfully acquired in the L2. Under this view, the instances of CS help increase attention to the target word making that piece of information more noticeable, and facilitating learning. However, there may be other factors at play that lead to the faciliatory effect of CS, particularly the presence of the native language. For Spanish-to-English CS, we argue that the retention benefit likely lies in the phonotactics of the novel words. Studies have demonstrated that high probability phonotactics facilitate performance during nonword repetition tasks (Munson, 2001; Nora et al., 2015; Vitevitch and Luce, 2005), during novel word learning measures (Havas et al., 2018), and during nonword recall (Nora et al., 2015). Native-like nonwords activate the schema of familiar phonological forms and facilitate encoding. In other words, it is possible that the CS from Spanish to English resulted in increased accuracy (compared to the Spanish-only condition) at visit 2, due to the novel words being easier to recall, since they followed native phonotactic rules. This account helps explain why Spanish-to-English CS may boost retention of newly learned words in comparison to Spanish only phrases, but it does not offer an explanation for why we also identified improved retention when the CS took place from English to Spanish.
While novel words that contain phonotactic characteristics in the L2 likely do not facilitate retention, we argue that the sentence frame being in the native language likely boosts the comprehension of the instructions needed to identify the novel object during training. When CS occurred from English to Spanish, the native language was used to direct the participant to the familiar item that served as a referent for identifying target creature (e.g. Look at the table, on top of the table is the
Taken together, our data indicate that the use of CS and the use of only the L2 result in similar immediate learning. However, testing at the second visit suggests that CS facilitates long-term retention of the newly-acquired words better than training that occurred only in the L2. In other words, we find that CS does not appear to negatively impact the acquisition of new word-object relations, and in fact, can benefit retention of newly learned words. These findings are important and should be taken into account when thinking of L2 instruction practices. Specifically, our work suggests that instructors in L2 classrooms should incorporate (to some extent) the native language as way of supporting vocabulary acquisition in the L2 – particularly when working with novice learners. In other words, relying exclusively on the L2 during instruction may not be the best approach.
These results are some of the first to explore how word learning in the L2 is influenced by CS, and they provide important insight into the impact that CS has on vocabulary learning for adult L2 learners who are starting to acquire their second language. However, there are several limitations associated with this study that should be addressed in future work. First, the way in which students were taught the novel word-object pairings was not true to a real-world learning scenario. In the study, students received two identical instances of training (or exposure to the novel animal-word pairing), one immediately after the other. This type of word presentation likely does not occur in a real-world classroom. In foreign language classrooms new words may be taught through textbooks, classrooms activities or group projects at different points in time. When teachers do verbally provide new word meanings, students are likely not tested on their knowledge of the word after two back-to-back presentations. Therefore, while our study results point to a benefit of CS on novel word retention on the particular task presented to participants, future work should aim to examine the impact of CS on vocabulary learning in a more naturalistic or classroom environment.
Another limitation was that our study did not include a production measure. Due to the lack of this measure, our data only demonstrate that participants can recognize the word-object pairs, but not produce the names of the objects independently. In the real world, language learners need to be able to recall and produce newly-acquired words, not only identify them when asked to do so. Future work should expand on this topic, and examine how CS may impact the production of lexical items that have been recently added to the lexicon. It is possible that CS could affect production and comprehension of words differently, given that previous studies involving word learning tasks have reported asymmetrical performance during comprehension versus production measures (e.g. Gershkoff-Stowe and Hahn, 2013).
Future work should also examine retention over a longer interval of time. In the present work, retention was only tested 1–3 days after initial learning. Therefore, it is unclear whether a longer gap between training and testing would lead to the same pattern of results that we observed. One possibility is that extending the time between training and testing may lead to variations in retention abilities associated with changes in long term consolidation (Cotton and Ricker, 2021). The long-term impact of CS on word retention is important, given that when new words are learned, they need to be retained for an extended period of time, as in real life, we do not only need words for 1–3 days after we initially hear them.
Finally, our sample allows for limited generalization of the findings. The participant pool was mostly homogenous as all individuals were students enrolled in 100-level Spanish courses at the same university. Students from other colleges and universities were not tested. Additionally, students who participated in the study were all college aged, which means our findings may not be generalized to other age groups (e.g. children). Given the rise in popularity of language immersion programs for school-aged children, future work should explore the impact of CS on L2 learning in this population. Further, our findings cannot be extrapolated to classroom settings. Although the study was conducted remotely (i.e. not in the lab), the experimenters ensured that the participants completed the study in a quiet, isolated environment (e.g. a quiet room in their home or in the library), which is unlike ‘real world’ learning in a classroom. Specifically, there were no classmates or a teacher present during the study, taking away an element of a naturalistic learning environment. Hence, additional work is needed to inform how CS may impact learning in L2 classrooms.
V Conclusions
To conclude, findings from the present study suggest that CS facilitates the retention of newly-learned novel words better than when these words are taught in sentences that are entirely in the L2 in a controlled experiment. The benefit of CS on L2 word learning may be related to increased contrast between the novel word and the language of the carrier sentence. Additionally, the native language may serve as a support to boost learning in CS conditions in comparison to learning entirely in the L2. These findings suggest that CS could be used as an intentional tool to clarify word meaning and inform the learner, resulting in overall better learning outcomes and novel word retention. This work serves as a first step for understanding the impact of CS on L2 vocabulary learning and should be expanded to other groups of L2 learners as suggested in the discussion.
Supplemental Material
sj-docx-1-slr-10.1177_02676583221113334 – Supplemental material for From one language to the other: Examining the role of code-switching on vocabulary learning in adult second-language learners
Supplemental material, sj-docx-1-slr-10.1177_02676583221113334 for From one language to the other: Examining the role of code-switching on vocabulary learning in adult second-language learners by Mackensie Blair and Giovanna Morini in Second Language Research
Footnotes
Acknowledgements
We thank Silpa Annavarapu, Emily Arena, Amelia Ayala, Sarah Blum, Kathryn Catalino, Kate Chirinos-Cazar, Ben Cushman, Aashaka Desai, Erin Felter, Katherine Filliben, Talia Gillespie, Taylor Hallacy, Sydney Horne, Amanda Kalil, Samantha Kennedy, Nicole Khanutin, Claudia Kurtz, Jessica Michels, Ryan Moore, Shreeya Parekh, Brianna Postorino, Jessica Price, Chaithra Reddy, Aurora Reible-Gunter, Katherine Richard, Nicole Scacco, Alexandra Stone, and Jackson Xiao for assistance in scheduling and testing participants.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was funded by a graduate fellowship from the Unidel foundation awarded to the first author.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
