Abstract
This study presents a teaching intervention to maximize the learning of a set of target words (TW) in learners of English as a foreign language (EFL) in a secondary school by means of intentional vocabulary learning activities and additional captioned television (TV) viewing. In the course of one academic year, two groups of grade 10 EFL learners (N = 64) were introduced to new TWs each week through language-focused exercises. The experimental group (n = 33) was additionally exposed to a captioned TV series where these TWs appeared. To measure lexical growth, all students took pre- and post-tests evaluating both TW form and meaning recall. Vocabulary retention was measured with an eight-month delayed post-test. Results revealed that vocabulary was mainly learned intentionally, but that additional viewing of the captioned TV series significantly contributed to greater lexical gains at different testing times. Similar vocabulary retention rates were observed for both groups. Conclusions and implications for teaching are drawn on the role of extensive video viewing for vocabulary learning in instructional settings.
Keywords
I Introduction
The distinction between intentional and incidental vocabulary learning has been widely used in the field of lexical studies (e.g. Huckin & Coady, 1999; Hulstijn, 2001; Malone, 2018). Among the various definitions available, intentional learning is typically understood as the learning which takes place when the main intention is to learn new words or consolidate partially learned vocabulary, usually with a focus on the learning task (Schmitt, 2008). Incidental vocabulary learning is said to take place when the vocabulary is learned without this intention, as a by-product of another task, or when no specific attention is devoted to vocabulary (Ellis, 1994). Therefore, strictly speaking, only when there is no intention whatsoever to learn vocabulary is it possible to catalogue learning as incidental. However, it has also been suggested that this is unlikely to occur frequently because, when learners encounter new words, the chances are they notice them and try to decipher their meaning (Webb, 2020). Hence, purely incidental learning is thought to be virtually non-occurrent, since the two (incidental and intentional) need to be seen as complementary rather than mutually exclusive (Schmitt, 2008; Webb, 2020).
There is a great deal of research on incidental vocabulary learning from reading and listening (e.g. Godfroid et al., 2018; Vidal, 2011), and recent research (Montero Perez, 2022; Peters & Webb, 2018; Rodgers & Webb, 2020) has shown that incidental vocabulary acquisition can also take place from extensive viewing, i.e. the regular, silent, uninterrupted viewing of second language (L2) television (TV) inside and outside the classroom (Webb, 2015). However, not much research has been conducted on the combination of incidental and intentional learning from reading or listening (e.g. Barcroft, 2009), and even less from extensive viewing.
In light of the above, the present article describes a teaching intervention that combined intentional learning, promoted through language-focused activities, and incidental learning, through extensive video viewing, to explore how vocabulary learning could be maximized in English as a foreign language (EFL) classes.
1 Intentional vocabulary learning
Intentional vocabulary learning is said to occur when any activity is ‘geared at committing lexical information to memory’ (Hulstijn, 2001, p. 271), typically in a language learning setting. In practice, telling students before doing an activity that they will take a test afterwards has also been classified as intentional learning (Baddeley et al., 2009). Research has shown there are several advantages to learning vocabulary intentionally; for example, the number of words learned intentionally tends to be higher than the number of words learned incidentally (Agustín Llach, 2009; Barcroft, 2015; Hulstijn, 2003; Laufer & Rozovski-Roitblat, 2011; Lindstromberg, 2020), even if other authors have found that incidental learning is conducive to more robust learning (Ahmad, 2012; Sok & Han, 2020). Intentional learning, through explicit vocabulary teaching and teachers’ explanations (Lee & Lee, 2022), is said to lead to higher indices of vocabulary learning in comparison to less interventionist incidental approaches (Alemi & Tayebi, 2011; Barcroft, 2009; Bilgin & Bingol, 2022; Coyne et al., 2007; File & Adams, 2010), also as far as retention of the target vocabulary is concerned (Qian, 1996). Similarly, intentional learning has been claimed to facilitate reaching a threshold that enables learners to use these newly learned words in meaningful contexts more quickly (Schmitt, 2008). Furthermore, explicit vocabulary instruction allows for the possibility of familiarizing oneself with a relatively high number of words in a short period of time, although not many word features can be taught when time is limited (Schmitt, 2000; Sökmen, 1997); priority is then given to establishing at least a link between a word form and its most common meaning. Sökmen (1997) also argues that language-focused vocabulary tasks allow learners to integrate new words into the lexicon faster, promote a deeper level of processing, provide students with a relatively high number of encounters with the target vocabulary, and allow for the development of learning strategies more easily.
Word-focused instruction (e.g. vocabulary exercises) is defined as a ‘response that requires the understanding of the words on which the exercise focuses, with or without producing them’ (Laufer, 2020, p. 352). Research has shown that focused exercises lead to higher learning indices than unfocused ones (Laufer & Rozovski-Roitblat, 2011) and are less time-consuming. However, not all exercises are thought to be equally effective. According to the Involvement Load Hypothesis (Laufer, 2020; Laufer & Hulstijn, 2001), a vocabulary exercise will result in deeper learning if it requires the use of a particular word to comply with the demands of the task, if the learner has to look for it (i.e. the word is not provided), and if there is some degree of evaluation implied (e.g. comparing a given word with other items). As a result, a productive task (involving writing, for example) should be more conducive to learning than a simpler receptive task. Similarly, the Technique Feature Analysis (Nation & Webb, 2011) also categorizes different types of exercises as more or less effective depending on their demands.
Nevertheless, it has also been suggested that the learning potential of vocabulary activities is often overvalued. In a recent meta-analysis, Webb et al. (2020) found that gains on immediate form- and meaning-recall tests were limited to 58.5% and 60.1% respectively, and that such gains dropped to 25.1% and 39.4% respectively when assessed using delayed post-tests.
2 Incidental vocabulary learning
When students do not have the firm intention to learn vocabulary, because the main focus is on communicative meaning rather than on form, vocabulary is then learned incidentally (Huckin & Coady, 1999). Incidental learning is considered an effective way to enlarge one’s vocabulary, and it means that learners can also acquire new lexical items when reading (Mohamed, 2018; Pellicer-Sánchez, 2016; Waring & Nation, 2004), listening (Chang, 2012; van Zeeland & Schmitt, 2013; Vidal, 2011) or watching TV (Rodgers & Webb, 2020; Webb, 2015), with a recent study showing similar gains under the three conditions (Feng & Webb, 2020). Thanks to the massive exposure to target language input that learners may receive (Lindgren & Muñoz, 2013; Muñoz, 2020), they repeatedly encounter new words, which can lead them to gradually learn their written and spoken forms both receptively and productively, their various meanings, or how they combine with other words to form collocations (Puimège & Peters, 2020; Rott, 1999; Uchihara et al., 2019).
Even though incidental vocabulary learning is essential for L2 learners, as not all the necessary vocabulary can be taught in a language programme (Nation, 2013), it can be slow and laborious, as considerable exposure to L2 input is necessary (e.g. Horst et al., 1998). However, it is also possible that studies on incidental vocabulary learning tend to underestimate the vocabulary actually learned, since lexical aspects other than those being assessed may also have been learned. Therefore, as Webb (2020, p. 229) states, ‘there is value in both incidental and intentional learning, they should not be viewed as competitors with each other.’ Further, he adds, ‘it is important to understand that both forms of learning are useful and likely necessary.’ Nation (2011) argues as well that both should be part of a well-balanced vocabulary programme.
3 L2 vocabulary learning through video viewing
Audiovisual input (e.g. short videos, TV series, documentaries) is said to be a good tool to learn an L2 (Vanderplank, 2016): first language (L1) TV consumption is one of the most popular leisure activities (De Wilde et al., 2020). It is also popular among L2 learners, who can currently find a wide range of multimedia resources in different languages. TV viewing has been shown to favour various areas of L2 acquisition: listening comprehension (Wang, 2019), grammar (Lee & Révész, 2020; Muñoz et al., 2023), or pronunciation development (Wisniewska & Mora, 2020). In relation to vocabulary learning, TV programmes allow a high degree of repetition of lexical items (Webb & Rodgers, 2009) and provide learners with imagery (Peters, 2019; Rodgers, 2018). According to the Cognitive Theory of Multimedia Learning (Mayer, 2009), learning with the support of images is more effective than learning with words alone. In addition, as there is a good deal of repetition and linkage in TV series (Rodgers & Webb, 2011), more background knowledge about the storyline and characters is accumulated by viewers, facilitating the learning of new language through context cues (Nation, 2007).
Previous research on incidental vocabulary learning through video viewing has shown that in terms of lexical gains this learning is quite modest. Most studies so far have used short videos or are one-off studies, and they have mainly been carried out with undergraduate students (with few exceptions). For example, Montero Perez et al. (2014) found that, on average, intermediate university learners could recognise the meaning of 0.60 words (3.53%) and recall 0.16 (0.94%) after having seen three short clips with captions (i.e. L2 subtitles), but recognise 0.53 meanings (3.12%) and recall 0.13 (0.76%) when captions were not available, out of the 17 target words (TW) included in the study. Low indices of vocabulary learning have also been observed in Durbahn (2019), where 12.66% of the 26 TW forms were recalled when videos were presented with captions, as opposed to 3.84% when captions were not available. In Teng (2022), English majors saw a full-length documentary with and without captions and were tested on 60 TWs. Results showed that they recalled 12.7% of word forms and 35.75% of word meanings when captions were available (results without captions were even lower: 7.47% for form recall and 17.63% for meaning recall). Similarly, Fievez et al. (2020) showed that secondary school students could recall 14.02% of the 50 TW meanings they encountered after watching a set of short, captioned videos in French, as opposed to 9.07% when participants only took the meaning recall test. Finally, Peters and Webb (2018) found that watching an uncaptioned full-length documentary led to a gain of 8.31% of TW meanings (out of the 64 TWs included in the study), which was higher than the control group (CG) who only completed the tests. Higher scores have been reported with other types of tests (e.g. clip association), although these tend to assess the initial stages of word learning. In contrast, recall tests are said to be more challenging (Cabeza et al., 1997), and lower gains tend to be reported (Nation, 2013; Nation & Webb, 2011).
It should be noted, though, that not much research has been conducted on extensive viewing, and longitudinal studies with sustained exposure are scarce. Rodgers and Webb (2020) showed that intermediate university learners acquired an average of 26.32% of the 60 words assessed with meaning recognition tests after having been exposed to 13 episodes of a TV series. These participants outperformed those in the CG, who followed the regular curriculum and acquired 23.14% of the words tested. The university participants in Fievez et al. (2023) could recall 27.49% of forms and 36.47% of meanings out of the 78 TWs on which they were tested after viewing six episodes of a Netflix TV series with glossed captions. Similarly, Frumuselu et al. (2015) found that university participants could recognise and recall the meaning of a considerable number of colloquial target items (out of the 30 included in the test) after having been exposed to 13 episodes of a TV series for a period of seven weeks (results were better when captions were available: 48.93% of items correctly answered compared with 36.50% of correct responses with subtitles).
Furthermore, very few studies have compared incidental and intentional approaches with video viewing, and those carried out to date have obtained inconclusive results. Sinyashina (2020) compared the amount of vocabulary learned after watching a whole season of an American sitcom lasting around five hours to a one-hour session of explicit vocabulary teaching, and the results revealed higher benefits in the group that received explicit teaching. Likewise, Gesa (2019) also reported small vocabulary gains among primary school EFL learners exposed to a TV series for an academic year, together with language-focused instruction, and compared to explicit instruction only. It was found that the benefits of incidental vocabulary learning through TV viewing were only observed towards the end of the experiment, after learners had been exposed to a considerable amount of audiovisual input. Moreover, Pujadas and Muñoz (2019) showed that pre-teaching the target vocabulary prior to watching a sitcom led to more vocabulary learning than watching the TV series without this language-focused instruction (and this was so irrespective of whether the series had been watched with captions or subtitles).
Finally, Montero Perez et al. (2015, 2018) also assessed intentional and incidental vocabulary learning from video viewing in two one-off studies, although the authors adopted a different approach: in these two studies, telling or not telling participants about the upcoming vocabulary tests was operationalized as intentional and incidental learning (Hulstijn, 2001). In Montero Perez et al. (2015), test announcement proved to be a significant factor in word meaning recall, with intentional learning leading to higher scores when compared to incidental learning, but it did not yield significant differences in form recognition or clip association tests. It should be acknowledged that learners’ vocabulary gains were very low (an average of one or two TWs learned out of the 18 on which participants were tested). However, using the same instruments and a similar pool of participants as in the previous study, Montero Perez et al. (2018) did not find any differences between intentional and incidental learning. The authors tentatively attribute this result to the announcement of an upcoming comprehension test, which all participants were informed about before watching the clips. This may have led students to focus their attention resources on unknown vocabulary, as they wanted to perform well in the follow-up content test. Results of the unannounced vocabulary test may have thus been affected by the announcement of the comprehension test.
4 L2 vocabulary retention through video viewing
Most studies to date on vocabulary learning through video viewing have mainly assessed immediate learning, and very few have analysed the retention of the vocabulary learned: as Barclay and Pellicer-Sánchez (2021) point out, retention and decay are in need of further investigation. Those that have analysed retention and thus included a delayed post-test differ in the timespan that elapsed between immediate and delayed testing. For instance, Nagira (2011) and Baltova (1999) administered the delayed test one and two weeks after the immediate post-test respectively, and saw that there was a low level of attrition. The first study showed that participants remembered at least 97% of the words they answered correctly in the immediate post-test, while in Baltova (1999) participants were tested on 30 target items and remembered 81.5% of the words learned when captions were not available and 70.5% when they could resort to on-screen text. Feng and Webb (2020) and Montero Perez (2020) also administered the delayed test one week after the end of the experiment, and both saw that attrition rates were very low: participants remembered 95% of the 43 words in Feng and Webb (2020), and 91.9% in the meaning recognition test in Montero Perez (2020), with some learning reported in the form recognition test in the latter study, in which participants were tested on 15 pseudowords. This last result is in line with Peters and Webb (2018), who asked participants to complete a one-week delayed post-test and concluded that too much deliberate learning had taken place between the two testing times, indicating that the delayed test was an unreliable instrument for that study.
Of special relevance to the present study are Pujadas (2019) and Ahrabi Fakhr et al. (2021), who administered the delayed test some months after the end of the experiment. Pujadas (2019) tested middle-school students eight months after the immediate post-test. Her participants had been exposed to 24 episodes of a TV series including a total of 120 TWs (note, though, that the delayed test only included 40 of these TWs), and half of them had been pre-taught the target vocabulary. She saw that learners could retain 63.6% of word forms and 74.8% of word meanings when they were pre-taught the target vocabulary, and 69.6% and 62% respectively when there was no pre-teaching; she concluded that pre-teaching of the target items did not result in greater retention. Ahrabi Fakhr et al. (2021) exposed Iranian undergraduates to a one-hour documentary and administered form recognition and meaning recall tests tapping into 96 TWs. One group took the tests immediately after watching the documentary, with two other groups doing so one week and three months later respectively. At each testing time, there was a CG that did not watch the documentary but completed the tests. Regarding vocabulary retention, results showed that viewing still had an effect one week after watching the documentary, but not three months later. In fact, having watched a captioned documentary accounted for 54% of the variance in the immediate test, for 27% one week after exposure, and for 4% three months afterwards.
Overall, existing research has shown that learning gains from incidental video viewing tend to be low, especially if recall tests are used. This type of research has mostly been conducted with populations of adult university learners. Moreover, hardly any studies have combined intentional learning of the target vocabulary with extensive TV viewing over a long period of time (Gesa, 2019; Pujadas & Muñoz, 2019). As for vocabulary retention, there are only two studies the authors are aware of that have investigated long-term retention (Ahrabi Fakhr et al., 2021; Pujadas, 2019), as opposed to short-term, even though learning a foreign language (FL) takes time, and research should definitely investigate the effects of treatments in the long run.
II Research questions and methods
The present study was conducted in a secondary school classroom context with pre-intermediate EFL learners and aims to answer the following research questions:
Does intentional vocabulary learning, combined with sustained exposure to captioned TV series, enhance vocabulary learning more than intentional learning alone?
Does intentional vocabulary learning, combined with sustained exposure to captioned TV series, enhance long-term vocabulary retention more than intentional learning alone?
To answer these research questions, a longitudinal between-participants design was adopted, lasting one academic year (nine months), divided into three academic terms. The independent variable under investigation was the effect of additionally watching captioned TV series. There were two dependent variables: learning and retention of the target vocabulary. Both the experimental group (EG) and the CG were taught the target vocabulary, but only those in the EG watched the TV series in class, while the CG did not.
1 Participants
Participants for this study were Catalan/Spanish bilinguals who were learning English as a third language in a state-funded school near Barcelona. Two intact classes taught by the same qualified teacher were selected for the study, one being the EG and the other the CG. The initial sample consisted of 64 learners (33 in the EG and 31 in the CG). However, some students did not meet the study’s two inclusion criteria, namely having completed both the pre- and the post-tests in a given term, and having attended all the viewing sessions of a given term (n) or missed a maximum of one (n–1).
Hence, the number of students slightly varies depending on the term. In the first, 57 students (30 + 27) 1 met the conditions, 52 students (29 + 23) in Term 2 and 48 (28 + 20) in Term 3. Forty students (24 + 16) also completed the vocabulary delayed test.
Participants’ mean age was 15.03 years (SD = .37, min. = 14, max. = 17), and there was a balanced number of participants identifying as males (47.5%) and females (52.5%). All participants were attending their last year of compulsory education in the Catalan educational system (i.e. grade 10). At this point, participants had received a total of approximately 1,100 hours of formal instruction and, following the Catalan curriculum, they were expected to have attained a B1 level of the FL according to the Common European Framework of Reference for languages. X_Lex and Y_Lex tests (Meara, 2005; Meara & Miralpeix, 2006, respectively) showed that EG participants had a receptive vocabulary size of 3,376 words (SD = 970 words, min. = 1,400, max. = 5,600, 95% CI [3,032, 3,720]), and the CG one of 3,402 words (SD = 1,050 words, min. = 1,350, max. = 5,750, 95% CI [3,002, 3,801]). Further, an independent samples t-test with equal variances assumed (Levene’s test: p = .929) showed that there were no differences in terms of vocabulary size between the two groups (t(60) = −.101, p = .920, d = .026, 95% CI [−539.11, 487.18]). Regarding their viewing habits at the start of the academic year, participants in both groups were not frequent viewers of original version TV, either with or without on-screen text, as can be seen in Tables 6 and 7 in Appendix 1. During the intervention, there was a small increase in the amount of original version TV participants watched during their free time, although this was so irrespective of the experimental condition (again, see Tables 6 and 7 in Appendix 1).
2 Instruments
Authentic audiovisual materials were selected for the study (i.e. TV series) and other instruments (i.e. vocabulary tests and language-focused activities) were specifically devised for the teaching intervention; these latter instruments have been uploaded to the IRIS repository (Marsden et al., 2016).
a TV series and target vocabulary
Two TV series were used during the intervention. In Term 1, students watched seven episodes from season five of I Love Lucy (Oppenheimer & Arnaz, 1951). In the next two terms, Seinfeld (David et al., 1989) was used, and 14 episodes from seasons four and five were watched (seven each term). Thus, in total, participants watched 21 episodes lasting an average of 22 minutes each, so they were exposed to 7 hours 54 minutes of original version TV. These two TV series were chosen because they were last aired in Spain many years ago and they were not available on major streaming platforms at the time data were collected, so it was unlikely that participants would be familiar with them. The first had also been used in a previous study with learners of a similar profile (Cokely & Muñoz, 2019) and the second had been piloted in the same institution with a group of grade 11 learner: the results had shown that participants enjoyed it and were eager to watch more episodes, as it matched their interests (Cancino, 2021; Lee & Pulido, 2017). The 21 episodes were selected based on the amount of testable vocabulary and on their content and appropriateness for the age group under study. Priority was given to consecutive episodes.
In terms of vocabulary demands, the two series were also very similar, as shown by a corpus analysis of the different episodes selected, compiled with VocabProfile v.2 (Cobb, n.d.) and the British National Corpus/Corpus of Contemporary American English 1–25K frequency lists as baseline (Davies, 2009; Nation, 2012). Table 1 shows the coverage level in each of the terms for the first three frequency bands. Furthermore, comprehension tests 2 administered to EG participants showed that mean comprehension was high. Taken together, these results confirm that the two TV series were appropriate since EG participants had a mean vocabulary size of 3,376 words and could understand most of what was happening; therefore, coverage was adequate.
Cumulative coverage for the first three frequency bands, divided by term.
From each of the episodes, five TWs that were thought to be unknown to participants were selected, so learners were tested on a total of 105 TWs (35 each term). First, researchers attentively watched all the episodes and jotted down possible TWs, which were then analysed according to different criteria: frequency in the episode (the minimum number of encounters was set at two) and in larger English corpora (1K and 2K words were avoided whenever possible), concreteness (there was a higher number of concrete than abstract words) and cognancy (perfect cognates between English and Catalan/Spanish were not selected and near perfect cognates were avoided whenever possible). The five words that met most of these requirements were then chosen as TWs for that episode (for the complete list of TWs and their frequency in the episodes and in larger corpora, see Table 8 in Appendix 2).
b Vocabulary pre- and post-tests
Pre- and post-tests assessed all the 35 TWs for a given term, and they were administered in Spanish, one of the learners’ L1s. TWs were presented in a quasi-randomized order and four practice items were added at the beginning of the test to familiarise students with the format. The pre-tests were used to assess participants’ knowledge of the TWs at the beginning of each academic term.
Students had to listen to an audio file recorded by a native speaker in which each TW form was read aloud twice: first, the TW number was read out to facilitate students’ tracking of the task if they got lost; one second later, the first repetition of the TW was read aloud, followed by a five-second pause before it was read again. There was a ten-second interval between items. The participants’ task consisted in writing down the English form of the word and providing the Spanish/Catalan translation or definition if they knew it. Thus, participants had to access TW meanings in their L1 through the L2 forms to be able to complete the task correctly, applying a process of translation which Vandergrift et al. (2006) call ‘mental translation’ (given that learners could translate as they listened to the audio).
This format allowed for the calculation of two different scores: (1) knowledge of TW forms, and (2) knowledge of TW meanings. Concurring with previous literature on the research topic (Nation & Webb, 2011), the tests were designed to assess written form recall (prompted by the aural form of the words) and meaning recall. Hence, they tapped into different types of knowledge: receptive knowledge of the spoken form (although not directly assessed by the tests but necessary to complete them successfully), productive knowledge of the written form, and passive recall (Laufer & Goldstein, 2004; Webb, 2007) (for an example of pre- and post-tests, see Appendix 3). All pre- and post-tests had either acceptable or good internal consistency, as shown by the Cronbach’s alpha values displayed in Table 2 (Tavakol & Dennick, 2011).
Cronbach’s alpha values for the pre-, post-, and delayed tests, divided by term.
c Vocabulary delayed test
The delayed test assessed the 35 TWs which were presented in Term 3, and it followed exactly the same format as the pre- and post-tests presented above (see Appendix 3). In this way, it was possible to see how many TWs learners retained eight months after the end of the intervention.
d Intentional learning activities
In all viewing sessions, one vocabulary pre- and one post-task especially devised for each episode were carried out (for examples of both, see Appendix 4).
The vocabulary pre-task was designed to introduce the target vocabulary appearing in the episode (i.e. the five TWs) at the beginning of each session. It always consisted of one exercise following a focus-on-forms approach, as participants practised the vocabulary items in a non-authentic language situation (Laufer, 2006), and paid attention to the TWs because they were needed in the activity. Different types of tasks were used throughout the year: six fill-in-the-gaps, five cloze tests, four word-searches, three crosswords, two word–image matching exercises, and one exercise asking participants to match TWs with their definitions.
In the post-task at the end of each session, the TWs introduced using the pre-task were encountered again. Participants were asked to listen to an audio, in which TW forms were read aloud twice, and had to write down the forms in English and select the best Spanish translation out of five different possibilities plus an ‘I don’t know’ option (format adapted from Rodgers and Webb, 2020); this task was very similar to what learners were asked to do in the pre- and post-tests. The ‘I don’t know’ option was included to avoid guessing, since students were explicitly instructed to select it if they were not sure of the answer (Zhang, 2013). All the options were given in the L1 since ‘the use of the first language to convey and test word meaning is very efficient’ (Nation, 2001, p. 351). Thus, the task tapped into written form recall (prompted by the aural form of the words) and meaning recognition, and it analysed students’ receptive knowledge of the spoken forms of the TWs, their productive knowledge of the written forms, and their recognition of the meanings.
3 Procedure
The intervention lasted one academic year, from September to June, and it was divided into three terms, as mentioned above. In each of the terms, the same procedure was followed: at the beginning, the vocabulary pre-test was administered by the researcher. The audio was played in the overhead speakers of the classroom, and time was given before and after to read and revise the questions and answers. In the following weeks, the seven viewing sessions of the term were held, starting one week after the administration of the vocabulary pre-test.
In each of these sessions, and for the whole academic year, the EG started by doing the vocabulary pre-task individually or in small groups. It was corrected immediately afterwards, and students could ask questions or share doubts about the vocabulary presented, so all TWs were presumably attended to. After that, the pre-task was collected and kept by the researcher. Next, the EG watched the corresponding episode of the TV series: this was always presented in English (L2) and with captions, as participants already had a mean receptive vocabulary size of 3,376 words (following Webb and Rodgers, 2009). The episode was projected onto the classroom whiteboard using an overhead projector. After that, the vocabulary post-task was distributed, and the audio file was played. Students were also given time to complete their answers. The task was not corrected in class: the researcher was in charge of doing so, in order to check if TWs were learned immediately after being presented through language-focused instruction and encountered in the TV series. As expected, the scores on this task were always very high, and reached a ceiling effect; therefore, they will not be reported here.
The CG did the vocabulary pre- and post-tasks at similar timings within each class (start and end of the session respectively). However, learners in the CG did not watch the episode of the TV series and followed their regular curriculum instead: they worked with the textbook and did the activities included in it. Extreme care was taken to make sure that they had no further contact with the TWs. As we used real TWs, the CG was necessary to control for learning gains outside the intervention and to determine whether the gains experienced by the EG could be attributed to the videos rather than to the language-focused activities and testing effects (Nation & Webb, 2011).
At the end of each term, participants took the vocabulary post-test, following the same procedure adopted with the pre-test. Finally, data collection finished with the administration of the delayed post-test, which took place eight months after the completion of Term 3 post-test. All the tests were unannounced, so participants did not know that the vocabulary taught and seen in the TV series (by the EG) was being analysed for research. Given the nature of the study, at the end of Terms 2 and 3, participants might have expected a vocabulary exercise (i.e. the post-test), as they had taken one in Term 1. However, they knew that the scores on these tests did not affect their English course mark.
4 Test scoring and data analysis
a Scoring criteria
The vocabulary pre-, post- and delayed tests were scored in a consistent way. Form and meaning were assessed separately because they are two different aspects of lexical knowledge (Nation, 2013). Furthermore, most participants who knew the meaning of a TW also knew its form, so the scores for TWs whose form and meaning were correctly answered will not be reported here as they would be very similar to word meaning scores. For a TW form to be considered correct, it could not contain any spelling mistakes, following Webb (2007) and Pujadas and Muñoz (2019). More lenient criteria were adopted with TW meanings, since translations, synonyms, or definitions were accepted. In the case of polysemous words (just around 10% according to the Oxford Learner’s Dictionary), only the meaning shown in the pre-task, the post-task and the episode was accepted as a correct response. When administering the pre-tests, participants were reminded to include all the meanings they knew of the TWs and, in the post-test, they were told that, if words had more than one meaning, they should note down the one learned in the tasks. For instance, in the case of the TW ‘to break’, they should include ‘exchanging a large bill for bills or coins in smaller amounts’; other meanings such as ‘separating suddenly or violently into two or more pieces’ were not taken into account for the scoring.
Regarding the delayed test, only the TWs that were learned during the third term (i.e. unknown on the pre-test and known on the post-test) were considered. These were classified as ‘maintained’ or ‘lost’. Being classified as ‘maintained’ implied that the TW form or meaning was again answered correctly in the delayed test. In contrast, words which were not answered correctly were classified as ‘lost’. The calculation of this measure provided a clear picture of the number of words still remembered after eight months without further explicit instruction of the TWs in the classroom, since it discarded already known or ‘unlearned’ words.
b Absolute and relative gains
As the pre- and the post-tests were identical, lexical gains were computed (Nation & Webb, 2011) and students’ progress could be observed throughout the academic year. After data were screened and trimmed, absolute and relative gains were calculated. To determine absolute gains, the TW forms and meanings correctly answered on the post-test were divided into ‘known’ and ‘learned’ (see definitions in this same section), and the second variable was reported. Nevertheless, only relative gains were used in the statistical analyses since they are a more fine-grained measure of vocabulary knowledge: they control for learners’ previous knowledge of the items tested (Horst et al., 1998), and have been extensively used in vocabulary research (e.g. Peters & Webb, 2018; Rodgers & Webb, 2020; Shefelbine, 1990). The following formula was applied:
where: ‘TWs learned’ = N of items that were answered incorrectly on the pre-test and correctly on the post-test; and ‘TWs known’ = N of items which were answered correctly on both the pre- and the post-test. The ‘N of items tested’ was the total number of TWs on which participants were tested each term.
c Statistical analyses
In order to answer research question 1 (comparing EG and CG), a mixed between–within ANOVA, with time (academic terms) as within-participants factor and condition (EG vs. CG) as between-participants factor, was run to investigate the relative gains in each group throughout the intervention. The analyses were then conducted term-independently due to the size differences in the sample: relative gains were compared across experimental conditions in each term using independent-samples t-tests or Mann–Whitney U tests to determine whether differences were statistically significant.
For research question 2 (the extent to which L2 vocabulary was retained), the percentages of word forms and meanings ‘maintained’ on the delayed vocabulary test by the EG and the CG were compared using independent samples t-tests and Mann–Whitney U tests.
III Results
Results are presented according to the two research questions proposed: (1) on vocabulary learning and (2) on retention of the vocabulary learned. Regarding research question 1, the results of the mixed ANOVA are presented first, and then the results of each term are analysed in depth. In all the analyses, assumptions were checked, and decisions taken accordingly (e.g. non-parametric tests were selected when data failed to reach a normal distribution).
Research question 1: Vocabulary learning
The descriptive statistics of gains (for both word form and word meaning learning) and the three terms the intervention was divided into can be found in Table 3 (absolute gains) and Table 4 (relative gains) (for the descriptive statistics of the pre- and post-tests, and for the statistical tests confirming that the two groups were comparable at pre-test time, see Tables 9, 10 and 11 in Appendix 5).
Absolute gains for word form and word meaning learning.
Notes. Absolute gains are out of 35. EG = Experimental Group, CG = Control Group.
Relative gains for word form and word meaning learning.
Notes. Relative gains are shown in percentages. EG = Experimental Group, CG = Control Group.
a Academic year
Thirty-eight participants (24 in the EG and 14 in the CG) were followed throughout the academic year and their gains computed in all three terms. The data for word form were normally distributed, and homogeneity of variances was assumed following Levene’s tests (ps ranging from .053 to .914). Moreover, the sphericity assumption was not violated either, as shown by Mauchly’s test (χ2(2) = 3.335, p = .189). Results showed a non-significant main effect for time (F(2, 72) = 1.354, p = .265, partial eta squared = .036) and experimental condition (F(1, 36) = 1.548, p = .221, partial eta squared = .041). However, the time*condition interaction was statistically significant (F(2, 72) = 4.214, p = .019, partial eta squared = .105). Since there was a significant interaction, simple main effects for time and condition were inspected. It was seen that time had a significant effect in the EG (F(2, 46) = 7.168, p = .002, partial eta squared = .238), and pairwise comparisons with Bonferroni adjustments showed that scores for Term 3 were higher than those for Term 1 (p = .006, 95% CI [−14.65, −2.23]) and Term 2 (p = .014, 95% CI [1.28, 13.30]). Regarding the CG, time did not prove to be a significant main effect (F(2, 26) = .606, p = .553, partial eta squared = .045). As for the simple main effect played by condition, and in order not to duplicate results, these analyses are presented below, with a larger sample of participants in each of the terms.
The data for word meaning were not normally distributed, and so they were squared. After this transformation, normality was reassessed, and it was seen that data did not violate this assumption. Furthermore, equal variances could be assumed following the results of Levene’s tests (ps ranging from .128 to .857), and Mauchly’s test showed that the sphericity assumption was met (χ2(2) = 1.547, p = .461). Results showed a highly significant main effect for time (F(2, 72) = 10.492, p = .001, partial eta squared = .226) (Cohen, 1988). Bonferroni post-hoc corrected coefficients showed that the number of meanings learned in Term 3 was significantly higher than those learned in Term 1 (p = .001, 95% CI [.28, 1.32]) and in Term 2 (p = .001, 95% CI [.25, 1.14]). However, the results revealed that experimental condition was not a significant factor for the learning of TW meanings (F(1, 36) = 2.770, p = .105, partial eta squared = .071), and that there was no significant interaction between time and condition (F(2, 72) = 2.169, p = .122, partial eta squared = .057).
b Term analysis
In Term 1, data were normally distributed for word form, but not for word meaning, according to the results of Shapiro–Wilk tests (ps ranging from .013 to .428). The results of an independent samples t-test with equal variances assumed (Levene’s test: p = .998) showed that the difference between the number of TW forms learned by both experimental conditions did not reach statistical significance (t(55) = 1.210, p = .231, d = .321, 95% CI [−3.56, 14.42]). As regards TW meanings, a Mann–Whitney U test did not yield significant differences between the two conditions either (U = 282.5, z = −1.963, p = .050, d = .260).
In the second term, word meaning learning did not present a normal distribution, as seen by Shapiro–Wilk tests (ps = .002 and .005). That said, an independent samples t-test with equal variances assumed (Levene’s test: p = .377) showed that sustained exposure to captioned TV viewing was beneficial for word form learning, with the EG significantly outperforming the CG (t(50) = 2.465, p = .017, d = .688, 95% CI [1.95, 19.10]). However, the difference in the number of TW meanings learned by the two groups did not reach statistical significance (U = 259, z = −1.377, p = .169, d = .191), as indicated by a Mann–Whitney U test.
Finally, in the last term, Shapiro–Wilk tests showed that data were normally distributed (ps ranging from .176 to .858). The EG learned more TW forms than the CG, this difference being large enough to reach statistical significance, as shown by an independent samples t-test with equal variances assumed (Levene’s test: p = .150) (t(46) = 2.198, p = .033, d = .643, 95% CI [.95, 21.58]). The analysis also revealed that the difference in word meaning learning was statistically significant according to an independent samples t-test with equal variances not assumed (Levene’s test: p = .006) (t(44.773) = 2.821, p = .007, d = .762, 95% CI [3.03, 18.13]). A graphical summary of the results can be seen in Figure 1.

Relative gains for form and meaning, divided by term and condition.
Research question 2: Vocabulary retention
Forty participants in total completed the vocabulary delayed test eight months after the intervention (24 in the EG and 16 in the CG). The percentages of word forms and meanings maintained and lost based on Term 3 results were calculated, and the descriptive statistics can be found in Table 5.
Word forms and meanings maintained and lost on the delayed test.
Notes. Scores are shown in percentages. EG = Experimental Group, CG = Control Group.
The data for word form retention were normally distributed. That said, even though the EG remembered more word forms than the CG (+6.97%), an independent samples t-test with equal variances assumed (Levene’s test: p = .481) revealed no significant differences in the percentage of TW forms maintained (t(38) = .755, p = .455, d = .244, 95% CI [−11.72, 25.67]). The data for word meaning retention were not normally distributed, and a Mann–Whitney U test yielded no significant differences in the percentage of word meanings correctly remembered on the delayed test (U = 156.5, z = −.991, p = .331, d = .157), even though the EG was able to retain a higher number of word meanings (+11.60%).
IV Discussion
The results will be interpreted following the two research questions that the study aims to answer. First, we will address the effects of captioned TV viewing in addition to language-focused instruction and then discuss the findings related to the long-term retention of the vocabulary learned by both groups during the last term of the intervention.
1 Vocabulary learning through language-focused instruction and extensive TV viewing
The results show that learning of TWs was taking place in the language-focused exercises, although vocabulary learning also occurred through watching TV: learners in the EG significantly outperformed their peers on several occasions, albeit not at all testing times. Although all participants were presented with the TWs at the beginning and the end of each session, this was often not enough for pre-intermediate learners to learn the target vocabulary, in contrast to what may happen at more advanced levels (Gesa, 2019; Suárez & Gesa, 2019). Additional exposure to the captioned TV series (offering visual, audio, and textual support for learning) further helped EG students in learning the target vocabulary. Therefore, the double intentional–incidental approach was positive for lexical development, especially after several episodes, in line with the findings of recent studies on reading (Sok & Han, 2020), where the intentional–incidental learning condition proved to be the most beneficial.
Our findings suggest that most of the vocabulary learning that took place was a result of the language-focused activities, so intentional learning and thus explicit attention to the target vocabulary proved to be useful for the two experimental conditions. This supports previous research suggesting that intentional learning leads to faster and higher gains than incidental learning (e.g. Agustín Llach, 2009; Laufer & Rozovski-Roitblat, 2011), and it reinforces the value of teachers’ explanations of the target vocabulary (Lee & Lee, 2022). However, learning explicitly is challenging when TWs are new and, in instructional settings, repeated encounters are often a luxury, as vocabulary is not recycled to the extent it is in naturalistic settings. In this respect, the study shows as well that explicit vocabulary instruction without further vocabulary recycling was not enough for the CG to catch up with the EG, especially after some exposure to multimodal input had accumulated: repeated encounters and higher indices of vocabulary recycling are thus recommended in order to maximise gains (Schmitt & Schmitt, 1995; Webb, 2007; Webb & Nation, 2017).
In relation to the additional encounters with the TWs provided through extensive viewing, imagery may have also assisted learners in making form–meaning connections (Peters, 2019; Rodgers, 2018), and may offer an explanation as to why the EG significantly outperformed the CG in both word form and meaning recall in the third term. Furthermore, EG students were possibly helped by the presence of captions, which have been shown to facilitate the understanding of the video and to free attentional resources that could be devoted to other tasks (Caimi, 2006; Danan, 2004). Possibly, captions also helped to acquire the spelling of the TWs, thus facilitating form–meaning connections (Uchihara et al., 2022). Therefore, it is not surprising that the EG significantly outperformed the CG in form recall in the second and third terms. However, the design of the present study made it impossible to know which of the two (i.e. imagery or captions) made the largest contribution to vocabulary gains. To determine this, an additional group watching the TV series without captions would need to be included in the design. Moreover, a different pool of TWs would have been needed, selected as well according to the degree of imagery in the TV series. These two topics are worth scrutinising in future research.
On a different note, it is possible that the assessment method (i.e. listening to L2 forms and using L2 to L1 translation) benefited the less skilled learners who, research has shown, rely more on mental translation than more skilled learners, who rely on self-regulation and metacognition (Vandergrift et al., 2006).
Likewise, topic interest and, as a result, students’ motivation, may have also affected the results. The two TV series shown during the year were well received by the EG, who found them engaging, in line with previous studies using I Love Lucy (Cokely & Muñoz, 2019). In consequence, it is possible that the learners in the EG were exposed to a task they were very interested in, while those in the CG followed regular course instruction, a less enjoyable activity. Topic interest has been shown to affect vocabulary learning and retention (Cancino, 2021; Lee & Pulido, 2017), also when meaning recall was assessed; therefore, involving learners in extensive viewing may have boosted learning gains in the EG, but not in the CG.
It should also be emphasised that more significant differences between conditions were found towards the end of the intervention. This difference cannot be attributed to the nature and frequency of the TWs selected, because the same selection criteria were applied in all terms, or to the difficulty of the TV series, since coverage levels were similar (see Table 1). Consequently, the effects of repeated viewing are clearer in the long term. This could be explained by the task familiarity effect (Bygate & Samuda, 2005), that is, after some sessions and greater familiarity with the task, learners may have been better at processing the input they were exposed to, leading to higher vocabulary learning indices. In this respect, they may have needed fewer encounters with the TWs, as they already knew how to process audiovisual input. They were more proficient too, and possibly had a larger vocabulary size (after one academic year of formal instruction and extra exposure to TV viewing), factors that would have facilitated the learning of new TWs. Hence, participants in the CG were put at a disadvantage compared with those in the EG, since the latter probably needed less input to (partially) learn new vocabulary, which might be an instance of the Matthew effect (Stanovich, 1986), or the rich-get-richer principle, according to which those in the EG learned faster and more efficiently (i.e. they benefited from extra input and probably had a higher level of proficiency and a larger vocabulary).
The accumulation of meaningful input should also be taken into consideration. At the end of the intervention, EG participants had been exposed to 7 hours 54 minutes of audiovisual input, whereas the CG had not had the opportunity to learn from extensive viewing in the classroom. Hence, the gap between the EG and CG in terms of multimodal input was proportionally and increasingly larger than in the two previous terms, and the effects of the intervention were more noticeable. However, it should be noted that, even though the amount of multimodal input is much greater than in previous research studies, it is still small in comparison to what the population receives today via TV: over the course of the intervention, these L2 learners were exposed to the same amount of audiovisual input that, on average, Spanish citizens receive in their L1 in two or three days (European Commission, 2018). Bearing this difference in mind, it is reasonable to think that this practice can bring about more substantial benefits if regularly sustained over time, as learners would be able to encounter words repeatedly, thus fostering incidental vocabulary acquisition.
It is also interesting to note that providing students with additional exposure to TWs through TV viewing does not automatically lead to additional learning: no significant differences are found between the groups in the first term (and in the second term regarding word meaning). However, when significant differences are found, they are always in favour of the EG. This positive effect may also be related to proficiency level. When the same experiment is conducted with more proficient university learners, intentional learning through formal instruction is enough for the TWs to be learned: no differences are found between experimental conditions, as extra exposure to captioned TV series episodes did not show any significant effects on word learning (Gesa & Miralpeix, 2022; Suárez & Gesa, 2019). However, the results from the present study show that, at lower pre-intermediate levels, extra exposure to captioned input does promote vocabulary learning gains.
Nevertheless, vocabulary learning results, especially those from Terms 2 and 3, may also be influenced by uncontrolled factors inherent in longitudinal vocabulary research; for instance, a more intentional approach to the learning task. It is possible that the TWs in Terms 2 and 3 were more intentionally learned than those in the first term. As participants were already familiar with the mechanics of the intervention, they knew that the TWs tested at the beginning of the term and taught in each session would be the object of testing at the end of the term, so participants (in both the EG and the CG) could have made an extra effort to commit them to memory so as to perform well later on the post-test, leading to a more intentional approach to learning. However, this effect would have been the same for both groups and hence does not distort the results. In line with this, deliberate learning of TWs may also have occurred in both groups: although vocabulary pre- and post-tasks were collected immediately after their administration and kept by researchers, it is possible that some participants, when at home, looked them up in dictionaries to see whether their answers were correct or not, leading to a more in-depth processing of the target vocabulary and favouring learning (Peters, 2007; Zhang et al., 2021). Finally, extramural exposure to audiovisual input should be considered too, since participants had the opportunity to watch these and other TV series at home. However, as can be seen in Tables 6 and 7 in Appendix 1, they were not frequent viewers of original version TV, at least until the end of the intervention, and, although it is true that there was a small increase in the amount of exposure to audiovisual input, this was similar across experimental conditions and so cannot have affected the significant differences found as a result of the treatment.
2 Long-term retention of previously learned vocabulary
As pointed out by Sok and Han (2020), longitudinal studies examining the effects of long-term vocabulary interventions help to understand how vocabulary acquisition can be affected by intentional, incidental or combined approaches over time. The results of the present study indicate that having learned the TWs either solely intentionally, through classroom exercises, or through the double intentional and incidental condition did not actually make a difference in the vocabulary retention rate eight months later.
It should be acknowledged that the retention rates found in the present study are much lower than those reported in previous studies on incidental vocabulary learning through TV viewing (e.g. Baltova, 1999; Feng & Webb, 2020; Montero Perez, 2020; Peters & Webb, 2018). Nevertheless, these studies mostly investigated short-term rather than long-term retention (one or two weeks after the administration of the immediate post-test vs. eight months afterwards). When the results are compared to those of the very few studies with comparable timespans, the retention rates are similar (Ahrabi Fakhr et al., 2021), especially when some pre-teaching of the target vocabulary occurred (Pujadas, 2019). However, it is interesting to note that the retention rates found in this study are higher than those reported in studies with other modes of input, especially reading (Rott, 1999; Waring & Takaki, 2003), despite the differences between the studies, i.e. in reading studies, timespans were shorter (from 1 to 3 months), they were conducted with undergraduate participants, and they often involved the learning of pseudowords.
Besides the long timespan, there are other factors that may have influenced the results. In the third term, significant differences in both lexical aspects were found in favour of the EG; hence, there was more room for vocabulary decay in the EG than in the CG. The EG had to remember nine forms and eight meanings (those learned in Term 3), whereas the CG had to remember six forms and five meanings, which, in terms of task demands, was easier. In addition, the uncontrolled exposure to the TWs that learners may have had during these eight months, and the potential learning derived from it, can also contribute to explaining the results. We cannot rule out the possibility that learners were exposed to some of the TWs presented in Term 3 during the next academic year, when input was not controlled. Moreover, as a natural consequence of their learning development (i.e. participants were more proficient at the time of the delayed test than when the intervention ended), they may have learned some of the words which were partially learned in Term 3. In line with this, participants might have deliberately learned some TWs during the eight months between the Term 3 post-test and the delayed test, leading to more vocabulary recycling and higher retention rates.
Finally, we should note that principles of the Cognitive Theory of Multimedia Learning (Mayer, 2009), stating that instruction through multimedia materials facilitates learning, help to explain the results from research question 1. However, this learning is not preserved by pre-intermediate level learners in the long run when compared to form-focused instruction through classroom exercises.
V Conclusions, pedagogical implications and limitations
The present study is one of the few longitudinal classroom-based interventions on extensive viewing to explore how audiovisual materials in the FL classroom could help to enhance vocabulary learning; the urgent need for this type of research in the field was pointed out by Montero Perez (2022). The study also explores the value of the insufficiently researched intentional–incidental double approach, which seems to lead to more effective learning.
The results of this study have several implications for teaching that are worth considering. First, as the literature suggests (Webb, 2020), in order to maximise vocabulary learning it is particularly important that intentional and incidental approaches are combined. This is one of the first studies that explores focused instruction with extensive viewing over a long period of time and checks possible retention effects. Although active learning vocabulary tasks led to some learning (on average, 27.43% of forms and 11.19% of meanings per term), higher learning indices can be attained if the same vocabulary is encountered incidentally through video viewing (gains increased by 9.07% per term in relation to word forms and 7.89% to word meanings). Based on these figures, we can say that the gains attributable to viewing are modest in comparison to those achieved through intentional learning. This leads us to think that had intentional vocabulary exercises not been included, learning gains would have been even smaller, and incidental vocabulary learning virtually non-existent. However, to be able to make this claim conclusively, a third condition watching the TV series but not receiving explicit instruction would be needed in order to be able to distinguish between the effects of video viewing and explicit teaching. In other words, despite the many advantages of audiovisual input for language learning, it is not advisable to expose pre-intermediate learners to extensive viewing without further support. Indeed, as Montero Perez et al. (2018, p. 19) propose, ‘to stimulate the intentional learning of new words, form-focused activities before or after viewing videos could be a more effective approach’. Without this extra support, vocabulary acquisition may still take place, but at a lower rate, especially at the recall level (Hulstijn, 1992; Laufer, 2005). Therefore, extensive viewing alone is not recommended in an FL classroom setting where instruction time is very limited, and a great deal of curriculum content must be covered.
Second, we have seen that vocabulary learning through FL television, along with focused instruction, can enhance language learning, especially after repeated practice and large amounts of exposure. However, this is to be expected since the EG participants spent longer on the task. It has also been suggested that in-class viewing can have an impact on the amount of FL audiovisual input learners may watch extramurally (Shea, 2000; Webb, 2015). As Webb (2015) points out, students must be first aware of the benefits for language learning and need to develop strategies to support comprehension. Once developed, it is possible that these learners will engage in more extramural exposure to TV series in the long run, when more vocabulary could be learned incidentally (provided that their families support this option at home, and suitable materials are selected). In our study, though, results from a questionnaire before and after the treatment did not show a clear increment in the amount of FL television participants watched outside school (see Tables 6 and 7 in Appendix 1). However, the TV series that were chosen for the intervention matched participants’ interests, promoting engagement, and this has been suggested as a relevant factor for learning (Cancino, 2021; Lee & Pulido, 2017). It can also be a possible explanation for the significant differences found in favour of the EG. It is thus important that participants are asked beforehand about their preferences in order to foster a firm commitment to the viewing task and to increase learning opportunities.
It should be noted, though, that findings with pre-intermediate learners may not apply to other proficiency levels (Gesa, 2019; Gesa & Miralpeix, 2022; Vanderplank, 2016), as learners may need to be at a certain proficiency level to make sense of the input. Therefore, it is necessary that the multimodal materials they are exposed to are appropriate for their level. If input is far beyond their competence level, incidental vocabulary learning may not take place: that is why analysing coverage and the lexical demands the input presents and relating them to learners’ vocabulary knowledge is mandatory.
From the intervention conducted, we can also extrapolate that regular watching of TV in the FL can be preferable over a single session or sporadic experiences, although this point will need to be experimentally assessed by further research. If an educational programme needs to be designed from scratch, it is worth remembering that incidental vocabulary learning through video viewing was more notable towards the end of the academic year (also because participants were probably more proficient in June than in September). The vocabulary learning derived from sporadic viewing is likely to be insignificant, and learners may see it as a way to spend class time doing an activity unrelated to the curriculum rather than as a learning opportunity. Practitioners need to bear in mind that getting used to audiovisual input with textual support takes time; learners should be advised that the task will be challenging at first, but beneficial in the long term.
Finally, certain limitations of the study should be mentioned. First, the sample size was small, so a replication would be needed to see whether the double intentional–incidental approach is consistently beneficial in other larger samples of participants. In addition, we do not know whether this novel approach to learning impacted learners’ general vocabulary knowledge, and so controlling for vocabulary size at the end of the intervention is recommended to reflect on the potential effects of extensive viewing on overall lexical development. Likewise, although teachers and students both enjoyed the experience, we did not administer a satisfaction questionnaire at the end of the academic year that confirmed the feedback received informally. This would be necessary in future studies. Last, the TWs in the study also present some limitations: since in this type of research word selection is constrained by the input in each episode, 10% of TWs were polysemous, so the fact that some learners may have already known one meaning could have facilitated the learning of the new one (even if this is true for both the EG and the CG). However, a principled approach to the selection of representative TWs (105 in total) was followed. Future research could then scrutinise whether certain word-related factors may possibly mediate vocabulary learning from captioned video viewing.
Footnotes
Appendix 1. Experimental and control groups’ viewing habits of original version television
Control group’s viewing habits of original version television, divided by time.
| Frequency | September (n = 28) | June (n = 21) | ||||
|---|---|---|---|---|---|---|
| Subtitles | Captions | No subtitles | Subtitles | Captions | No subtitles | |
| Never | 25 | 60.7 | 64.3 | 33.3 | 28.6 | 33.3 |
| Monthly | 50 | 32.2 | 21.4 | 47.6 | 42.8 | 52.3 |
| Weekly | 25 | 7.2 | 14.2 | 19 | 23.8 | 14.3 |
| Daily | 0 | 0 | 0 | 0 | 4.8 | 0 |
Note. All figures are shown in percentages.
Appendix 2. List of target words
List of target words on which participants were tested, and their frequency of occurrence in the episode and in larger corpora.
| Term 1 | Term 2 | Term 3 | ||||||
|---|---|---|---|---|---|---|---|---|
| Target word | Frequency | Target word | Frequency | Target word | Frequency | |||
| Episode |
Band | Episode |
Band | Episode |
Band | |||
| beady eyes | 2 | 4K | anklet | 3 | 16K | accountant | 5 | 4K |
| bucket | 5 | 2K | awning | 4 | 10K | brassiere | 4 | 6K |
| cheapskate | 2 | 17K | blanket | 6 | 2K | to curse | 5 | 4K |
| choppy | 3 | 11K | blurb | 5 | 13K | deaf | 12 | 4K |
| conductor | 4 | 3K | to break | 2 | 1K | to flinch | 7 | 7K |
| crowbar | 3 | 12K | butler | 6 | 8K | flinty | 4 | 9K |
| crummy | 4 | 14K | cabin | 11 | 4K | frames | 7 | 2K |
| cue tip | 2 | 5K | to cherish | 2 | 6K | fungus | 8 | 5K |
| cuffs | 5 | 6K | to choke | 2 | 4K | godfather | 7 | 6K |
| curler | 4 | 2K | cigar | 16 | 5K | lineswoman | 8 | off-list |
| downtown | 4 | off-list | cleavage | 13 | 6K | lobster | 8 | 6K |
| dummy | 3 | 6K | coach | 4 | 2K | loop | 3 | 4K |
| to fool | 3 | 2K | conveyor belt | 2 | 3K | lure | 3 | 5K |
| forecourt | 6 | 13K | dill | 3 | 10K | mayor | 9 | 3K |
| to forge | 4 | 4K | filthy | 2 | 5K | medicine cabinet | 6 | 3K |
| furnace | 2 | 7K | gauge | 6 | 4K | to melt | 2 | 2K |
| gear | 2 | 2K | to leer | 4 | 9K | mohair | 4 | 13K |
| grapefruit | 2 | 9K | lieutenant | 5 | 4K | nametag | 9 | off-list |
| to hobnob | 2 | 16K | narc | 3 | 15K | napkin | 3 | 9K |
| hunk | 2 | 9K | nostrils | 3 | 6K | ply | 8 | 8K |
| newsstand | 2 | off-list | nut | 3 | 2K | podiatrist | 6 | 17K |
| to peek through | 5 | 7K | pickup | 4 | 5K | rabies | 4 | 11K |
| penthouse | 5 | 10K | to poke | 6 | 4K | rush | 4 | 2K |
| preview | 4 | 6K | prude | 2 | 12K | shot | 9 | 1K |
| razor | 3 | 6K | remote | 4 | 3K | to sniff | 14 | 5K |
| to roll | 6 | 1K | to rip | 3 | 2K | to spare | 13 | 2K |
| rubdown | 4 | off-list | script | 2 | 4K | spasm | 3 | 8K |
| rugged | 3 | 6K | shackles | 2 | 9K | to spit | 3 | 4K |
| seasick | 2 | off-list | to snub | 12 | 9K | to squint | 7 | 6K |
| sneaky | 2 | 5K | spot | 8 | 1K | to step off | 11 | 1K |
| to snore | 3 | 7K | stub | 11 | 7K | stitch | 3 | 4K |
| tenant | 4 | 4K | toll | 2 | 5K | to suck | 8 | 2K |
| to tuck in | 2 | 4K | treatment | 4 | 1K | to sweep | 8 | 2K |
| upper berth | 4 | 7K | weed | 2 | 2K | to wean | 6 | 8K |
| to vacuum | 2 | 5K | to wipe | 2 | 2K | to withdraw | 2 | 3K |
Note. K = thousand, i.e. 5K = 5,000, etc.
Appendix 3. Examples of pre- and post-tests
Appendix 4. Examples of vocabulary pre- and post-tasks
Appendix 5. Pre- and post-tests’ raw scores for word form and word meaning,and results of the independent samples t -tests and Mann–Whitney U tests between EG and CG at pre-test time
Results of the independent samples t-tests and Mann–Whitney U tests between EG and CG at pre-test time, divided by term and lexical aspect.
| Term | Form | Meaning |
|---|---|---|
| 1 | U = 339, z = –1.060, p = .289, d = .282 | U = 346, z = –1.059, p = .290, d = .252 |
| 2 | t(50) = .787, p = .435, d = .220, 95% CI [–1.58, 3.60] | U = 307.5, z = –.631, p = .528, d = .133 |
| 3 | t(46) = –.879, p = .384, d = –.257, 95% CI [–4.68, 1.84] | U = 231, z = –1.043, p = .297, d = .299 |
Notes. EG = Experimental Group, CG = Control Group.
Acknowledgements
The authors are grateful to the students and teachers of Institut Damià Campeny (Mataró) and to the GRAL Research Group at the Universitat de Barcelona. They are also very grateful to the anonymous reviewers for their valuable insights and comments.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Spanish Ministry of Economy, Industry and Competitivity (grant numbers BES-2014-068089, FFI2013-47616-P, and FFI2016-80564-R); and the Catalan Agency for Management of University and Research Grants (grant number 2017 SGR 560).
