Abstract
This study examines the effect of the listener’s mother tongue and competency level in second language (L2) English and L2 Spanish on the pronunciation ratings of Spanish-accented English. Pronunciation is operationalized by means of three constructs: intelligibility, comprehensibility and foreign-accentedness. Stimuli from 60 Spanish speakers of advanced English (levels B2–C2) were collected at a Spanish university. Subsequently, their speech samples were judged by 330 native and non-native speakers of English (Spanish, Polish and Other L1s) in an online test. Differences in intelligibility scores were associated with the listeners’ mother tongue, their English level, and their knowledge of Spanish with native English speakers outperforming all other groups. Foreign-accentedness was also affected by the listeners’ mother tongue: the Spanish listeners were the harshest accentedness judges although their understanding of the samples was not significantly hindered, which may suggest negative in-group attitudes towards Spanish-accented English among Spanish speakers. The native English listeners, conversely, were the most lenient raters. Moreover, comprehensibility differences were associated with the listeners’ mother tongue and their English level. Thus, evidence was found to support the existence of an interlanguage speech comprehensibility and intelligibility benefits, which stresses the role of the listener’s language background and their attitudes towards English as an international language.
Keywords
I Introduction
Over the past decades, the linguistic dominance of English as a global lingua franca (Jenkins, 2000; Seidlhofer, 2011) has led to a change in the way English itself and international communication by means of English have been approached by researchers. As a consequence of the acknowledgement that communication in English in today’s world occurs to a great extent among non-native speakers (NNSs) – or at least with non-native speakers as equal participants – applied linguists have shifted away from considering only native speaker (NS) accent models (Holliday, 2006; Levis, 2005) as a benchmark against which second language (L2) performance and successful acquisition should be measured. Thus, the role of nativeness as the learner’s (or the international user’s) ultimate goal or a measure of successful communication has gradually diminished – at least in research (Derwing, 2018; Jenkins, 2000). As a result, research into other aspects of speech, such as intelligibility (Jenkins, 2000; Levis, 2005) and comprehensibility (Derwing, 2018; Derwing & Munro, 1997, 2015) has flourished alongside accentedness studies.
The main purpose of this article is to investigate the relationship between the listener’s characteristics and the intelligibility, comprehensibility and foreign-accentedness ratings of Spanish-accented L2 English. The results presented herein are part of a research project dealing specifically with L2 Spanish pronunciation (Pietraszek, 2024a). Intelligibility (INT) was construed here as a measure of the extent to which speech samples were segmentally decoded (see Jenkins, 2000). This decision is based on two premises. First, reducing intelligibility to phonology by controlling contextual information follows Jenkins’ (2000, 2002) observation that even proficient L2 English (or English as a Lingua Franca (ELF)) speakers’ reliance on bottom-up (phonological) strategies when decoding messages is greater than among native speakers – a view derived from a paradigm earlier proposed by Smith and Nelson (1985) and confirmed in more recent studies (Deterding, 2013; Field, 2005; Meierkord, 2004). This interpretation of intelligibility is different from those studies where the linguistic context is not controlled and involves the decoding of meaning (e.g. Derwing & Munro, 1997, 2015). Second, a strictly phonological operationalization allows to focus the attention solely on pronunciation in what is essentially a pronunciation study and detect with more precision the segmental elements which may pose problems – although this will not be dealt with in this article. On the other hand, foreign-accentedness (FA) and comprehensibility (COM) are two subjective measures, usually obtained by means of semantic differential scales (Dörnyei, 2007; Poljak, 2019). The former gauges the perceived degree of a foreign accent in a speaker while the second assesses how easy a user is to understand (Derwing & Munro, 1997, 2015; Isaacs & Trofimovich, 2012; Munro & Derwing, 1995; Poljak, 2019; Trofimovich & Isaacs, 2012).
1 Literature review
Given the interactional nature of communication and understanding, numerous studies have focused on the figure of the listener and attempted to correlate various listener characteristics (demographic, linguistic and attitudinal) with intelligibility, comprehensibility and foreign-accentedness scores (see Sewell, 2010). As Trudgill (2008, p. 222) points out, a balance between the speaker’s and the listener’s needs is required for efficient communication to occur – a phenomenon he labelled the ‘speaker–listener equilibrium’. It seems only logical, for example, that the listener’s level of English competence will have a bearing on the pronunciation judgments issued by that listener. Indeed, a series of studies showed the association between the listeners’ level and pronunciation ratings (Van Wijngaarden et al., 2002; Xie & Fowler, 2013). It was recently revealed, for example, that high proficiency listeners were less sensitive to accent variation than intermediate listeners (Kang et al., 2020). Similarly, Eger and Reinisch (2019) discovered German listeners aligned with native listeners regarding the acoustic cues used when gauging the accentedness of German English. However, this was true only for proficient or experienced users, suggesting lower-level speakers perceived English through their first language (L1) lens to a greater extent. The rater’s familiarity with the speakers’ language has also been investigated with divergent results, showing both no effect of this variable on accentedness ratings (Bergeron & Trofimovich, 2017) and a familiarity bias in oral proficiency ratings (Winke & Gass, 2013). Similarly, the effect of previous exposure to the speakers’ accent remains unclear. Its association with comprehensibility (Kahng, 2023) or intelligibility (Kennedy & Trofimovich, 2008) has been reported alongside a lack of such effects on accentedness and comprehensibility (Kennedy & Trofimovich, 2008). Furthermore, there seems to be no consensus on the contribution of the rater’s native status to accentedness judgments. While a number of studies have revealed no clear differences in accentedness judgments between NSs and NNs (Crowther et al., 2016; Derwing & Munro, 2013; Eger & Reinisch, 2019), Western research on accentedness, oral performance, and irritation caused by non-native English suggests NNSs may be less tolerant of accented English than NSs (Fayer & Krasinski, 1987; Isaacs & Thomson, 2013; Kang, 2012). Moreover, variance in accentedness ratings has been linked to listener attitudes towards non-native speakers and their speech (Cammarata Philbrick & Ingvalson, 2021; Gluszek & Dovidio, 2010; Ingvalson et al., 2017, Simon et al., 2022). On the other hand, some studies conducted in Asian contexts suggest that NSs may actually be more lenient raters (Brown, 1995; Munro et al., 2006) and that Outer Circle judges might be less strict than Inner Circle raters (Saito & Shintani, 2016).
The listener’s language background has also been widely investigated in intelligibility research. In an analysis of English NSs, and Dutch and Chinese users of English, Wang (2007), found that listener-nationality is even more significant for mutual intelligibility scores than speaker nationality, a clear signal that the importance of the inner properties of the listener group should never be overlooked. It should be kept in mind from the onset that, as Pickering (2006) observed, the different findings on listener effects may be due to methodological diversity and it may be difficult to draw general conclusions – a statement likely to be true until today.
In her article under the self-explanatory title ‘The listener: No longer the silent partner in reduced intelligibility’, Zielinski (2008) analysed three NSs of Australian English listening to 47 NSs of Korean, Mandarin and Vietnamese proficient in English (TOEFL scores of 580 or above). The listeners had no previous exposure to English used by NSs of those languages. Zielinski’s (2008) exploratory study revealed a hierarchy of importance of segmentals depending on their position in the syllable and is an important contribution to the field as it gave voice to the listeners in the analysis of their own processing difficulties.
In an earlier study, Gass and Varonis (1984) analysed four types of familiarity effects on listeners depending on the topic and the language background of the speakers. They recorded ‘The North Wind and the Sun’ story and sentences (related and unrelated to the story) uttered by two Arabic and two Japanese speakers selected from a wider sample and of averagely good intelligibility to experienced ESL teachers. 142 NSs of English took part in the study. ‘Comprehensibility’ is used here according to the Smith and Nelson (1985) paradigm, whereas Derwing and Munro (2015) would call this ‘intelligibility’. Familiarity with the topic was demonstrated to be a factor significantly enhancing comprehensibility for the listener in the post-text transcription task (Gass & Varonis, 1984, p. 70) as was the familiarity with a particular speaker (p. 81). Experienced ESL teachers also scored higher on the tasks than naïve listeners (p. 79) – both the familiarity with non-native speech in general and with a particular accent proved to facilitate comprehension (p. 81).
The previously mentioned studies were concerned mostly with native English listener processing. Yet Rajadurai (2007, p. 94) argues it is a myth that ‘the native speaker is always the best judge (or representative) of what is intelligible’ (see Bamgbose, 1998, p. 11). In fact, in an ELF context, it seems an intuitive assumption that non-native speakers of English should be more accurate on intelligibility tests when gauging other speakers with whom they share the same language background, and thus, presumably, the same interlanguage. According to many researchers (see Zielinski, 2008, p. 70), NNSs rely on their L1 strategies when decoding messages in an L2. It would indeed be interesting to investigate what strategies listeners from different L1 backgrounds use – bearing in mind Jenkins’ (2000) suggestion that NSs use more contextual (top-down) cues than proficient NNSs, who rely on bottom-up processing to a greater degree.
Moreover, non-native speakers from a shared L1 background often claim to understand their fellow L1 speaker in English better than other NNs of English, a phenomenon dubbed ‘interlanguage speech intelligibility benefit’ (ISIB). Bent and Bradlow (2003) tested this claim under controlled conditions using sentences uttered by two native talkers of Korean and Chinese each and one NS of English, and playing those to 21 English, 21 Chinese, 10 Korean and a group of 12 mixed native listeners. The results showed that (1) NSs found it easier to understand NSs than they did NNSs; (2) NNSs understood their fellow L1 talkers and native talkers equally well, (3) NNSs understood people NNSs from other language backgrounds just as well or better than they did NS. This may be because of ‘certain tendencies in foreign accented English regardless of native language background’ (mismatched speech intelligibility benefit) (p. 1608). However, in a later study by Hayes-Harb, et al. (2008), it was found that proficient native Mandarin listeners did not find other proficient Mandarin speakers more intelligible than they did NSs. This was true only for low proficiency listeners and low proficiency speech suggesting an interaction between the level variable and ISIB. Yet Mandarin listeners did outperform English natives identifying Mandarin-accented English words in a task where final consonant voicing was targeted (p. 664). Xie and Fowler (2013) also revealed an ISIB for listeners of Mandarin residing in Beijing and showed how Mandarin speakers living in the US adopted a mix of native and non-native processing strategies thus diminishing the extent of ISIB (pp. 377–378). The hypothesis that exposure to an L2 accent has an impact on processing was tentatively confirmed by a study by Li and Mok (2015), which suggested that a shared phonological knowledge was not sufficient to explain ISIB and pointed to exposure to a specific L2 accent as another important contributing factor. In a response time (RT) study, Shu et al. (2016) also revealed that L2 English participants had faster RTs to same-L1 speakers’ stimuli than to signals coming from different-L1 talkers, which suggests their processing was easier and could be considered an objective measure of increased comprehensibility. Similar results were found by Ludwig and Mora (2017) in an experiment with Germans and Catalans where matching L1s were associated with quicker processing and better comprehensibility scores. Additionally, low level listeners benefited from listening to English input by speakers from a matching L1 background, while higher proficiency listeners outperformed native speakers under the same circumstances (p. 167). In this article, when comprehensibility, as opposed to intelligibility ratings are involved, this phenomenon will be dubbed the ‘interlanguage speech comprehensibility benefit’ (ISCB).
Turning to the Spanish university context, a preliminary study (Gómez Lacabex & Gallardo del Puerto, 2021) limited to a test of the comprehensibility and intelligibility of 10 technical terms as pronounced by 12 EMI students revealed interesting results. The participant groups’ intelligibility scores show that English, Spanish and Polish listeners’ intelligibility scores did not differ from each other. However, Chinese students had more difficulty and were at a disadvantage when understanding the terms. Although the findings of this research must be pondered with caution, they may be interpreted as evidence supporting ISIB (p. 137). Regarding comprehensibility, evidence for an advantage of Spanish listeners was also found (p. 138.) suggesting Spanish speakers perceived the sampled speech to be less difficult than all the other groups (p. 134).
However, evidence contradicting ISIB or ISCB can also be found. A quantitative study by Munro et al. (2006), extemporaneous speech stimuli from 48 native speakers of Cantonese, Japanese, Polish, and Spanish were assessed by 30 English listeners, 10 NSs of Cantonese, 10 NSs of Japanese and 10 NSs of Mandarin. The study extensively focused on listener characteristics and concluded that there is a high correlation between all but one group for intelligibility, comprehensibility and foreign-accentedness. Therefore, it was concluded that the ISIB was not systematically at work. For example, only the Japanese speakers exhibited familiarity benefits finding their fellow native speakers to be more comprehensible than did Mandarin and Cantonese listeners, they also considered their compatriots to be more intelligible than did the listeners from any other group and less accented than did the Cantonese listeners. One year earlier, Field (2005) also found scarce or non-existent impact of a shared L1 on intelligibility, whereas Wang (2007, p. 257) claims ISIB to be ‘pervasive’, although not absolute. All in all, no conclusive evidence has been proposed yet to fully support or explain the functioning of ISIB (or ISCB), although it undeniably preserves its intuitive appeal for further research and will hence be looked into in this article.
2 Research questions
The main hypothesis underlying the design of the article is that listener characteristics are consequential for the ratings issued (in the case of COM and FA) and the scores obtained (in the case of INT) by the different listener groups (henceforth referred to as ratings). The specific research questions to be answered on the following pages are the following:
Research question 1: What is the association between the listeners’ L1 and their COM-FA-INT ratings?
Research question 2: What is the association between the listeners’ level of L2 English and their COM-FA-INT ratings?
Research question 3: What is the association between the listeners’ knowledge of Spanish (i.e. the speakers’ L1) and their COM-FA-INT ratings?
Research question 4: Is there evidence that an Interlanguage Speech Intelligibility or Comprehensibility Benefit exists?
II Methodology
1 Instruments
Two recordings were used in the collection of the stimuli from the speakers in this research: (1) an elicitation paragraph (EP, Appendix A; Pietraszek, 2024b) consisting of a semantically meaningful and textually coherent sentences to be used in COM and FA ratings and (2) a series of 40 semantically unpredictable sentences (SUS, Appendix B; adapted from Wang, 2007) for phonological intelligibility testing. SUS (Benoît et al., 1996) are correctly formed sentences with little to no real world meaning, very much like Chomsky’s ‘colourless green ideas sleep furiously’, although the SUS meanings in this study are not intended to be contradictory. As such, they enable natural pronunciation and intonation, while impeding the use of contextual cues in their interpretation to a great extent (see Jenkins, 2000). As mentioned in the introduction, the construct of intelligibility was meant to eliminate as much contextual information as possible (see Jenkins, 2000). Thus, SUS were deemed an appropriate tool especially considering that previous research showed they are one of the most efficient methods to gauge segmental intelligibility (Kang et al., 2018). A total of 60 speakers were recorded whose description can be found in the following section. The recordings were conducted personally by the author in quiet university offices or classrooms using an H4N Pro Handy Recorder and they were stored in the .wav 16bit/44.1kHz format. The speakers were instructed to read the text and the sentences at a natural pace. They were given time to read the content before the start and were told they could repeat any input they were not satisfied with as only the final version would be used.
After processing the recorded input, a test was set up online. All the listeners (n = 330) were recruited around the world through social media and volunteered to participate. There was no contact between the researcher and the participants at this stage. They were asked to take the test in a quiet environment, preferably using headphones, although that was not controlled for (see Nagle & Rehman, 2021). For each listener, the test consisted of stimuli from a group of 5 speakers randomly assigned by the online platform (LimeSurvey GmbH, 2019) followed by a short questionnaire (Appendix C). The order of the items on the test was also randomized for every listener. On average, every speaker was assessed by 27.5 listeners.
There were a total of 30 questions in the online test: 5 FA ratings, 5 COM ratings (based on approximately 25–30-word excerpts from the EP), and 20 orthographic SUS transcriptions to test the speakers’ segmental intelligibility. None of the 20 SUS transcribed by each listener was repeated to avoid learning effects. The 20 transcriptions included 4 sentences by each of the 5 speakers assigned to the listener with 3 or 4 content words per sentence. This number of test items was chosen to minimize listener fatigue and increase completion rates on a test they were performing freely with no compensation. Before the test began, an illustration of the three item types was presented to them as training, the results of which were not included in the analyses. After taking the test, a brief sociodemographic questionnaire was completed inquiring about the listeners’ linguistic background and self-reported proficiency levels in English and Spanish. The average time required for the whole survey was 21 minutes. Although there was no time limit, the listeners had to complete the test in one browser session.
Both FA and COM ratings used semantic differential scales. For FA the listener was asked to directly ‘rate the speaker’s degree of foreign accent’ from (1) native or native-like to (6) very foreign. In the COM items, the listeners were required to complete the sentence ‘the speakers is ___ to understand’ on a scale from (1) very hard to (6) very easy. The same excerpt was used for each speaker in the rating of their COM and FA but a listener never heard the same fragment from different speakers once to prevent familiarity effects. Although 9-point scales have been extensively used in comprehensibility and accentedness research to date (see Poljak, 2019), methodological studies have shown that shorter scales may provide more reliable and valid results (Kermad & Bogorevich, 2022) and present fewer problems (Isaacs & Thomson, 2013). The use of an even or odd number of options has also been discussed (e.g. Chyung et al., 2017) as a middle option present in odd-numbered scales may be misused by the raters as not representative of a truly neutral judgment (Kulas & Stachowski, 2013). Overall, while the question of an ideal number of items remains unresolved, some research suggests that using more than six or seven responses in psychometric tests may be statistically unjustified and recommends the use of 6-point scales (Simms et al., 2019; Taherdoost, 2019). Scales with this number of options have previously been deployed in accentedness studies (Fuse et al., 2024; Kornder & Mennen, 2021; Mennen et al., 2023; Wrembel, 2008). The intra-class correlation coefficients across the listener language groups in this study showed excellent (Cicchetti, 1994) interrater reliability values (COM r = .828, p < .001; INT r = .896, p < .001; r = .916, p < .001) using the tools described in this section.
2 Participants: speaker and listener sample description
A group of 60 speakers of English enrolled in five different undergraduate degree programmes at a Spanish university were recorded for the study. They also completed a sociodemographic questionnaire (Appendix C) and signed a consent form. No assessment task was associated with the recording and the participation was voluntary. There were 26 females and 34 males in the sample. The average age in the speaking sample was 21.2 years. (Min = 19, Max = 26, SD = 1.44). 55 speakers reported that Spanish was the native language of both their parents. One speaker had a British father, and four others reported that one of their parents’ mother tongue was neither Spanish nor English. However, none of these five speakers had spent more than a total of 0.75 years in the non-Spanish-speaking parent’s country of origin. Since all of them had been born and raised in Spain, they were deemed valid representatives of English spoken in the country. All the talkers had a minimum B2 level of English and were not enrolled in degrees related to language or linguistics. Although no formal level-testing was conducted for this study, the institution’s admission criteria, placement tests and the self-reported language certificates were considered sufficient to qualify for recording. They all came from classes where a B2 level of English was an entry requirement or reported to be in possession of an external certificate. In each case, their instructor was also consulted regarding their level of English to confirm their eligibility for the recording as B2-or-above English speakers. Among the 50 speakers who reported to be language certificate holders, there were 27 B2, 15 C1 and 8 C2 certificate holders. The rest were assumed to have at least a B2 level of English according to the eligibility criteria mentioned above. 1
The total number of listeners who voluntarily participated in the online intelligibility, comprehensibility and foreign-accentedness tests amounted to 330. Their age ranged from 14 years 2 to 70 years (M = 29.67, SD = 11.14). Around a third of them (33%) had English as their mother tongue, around a quarter (25.2%) were L1 Spanish speakers and a fifth (20.3%) were Polish speakers. Those with L1s different from English, Spanish and Polish accounted for the remaining 21.5%. 3 Regarding gender, 206 females and 120 males completed the online questionnaire while 4 people preferred not to provide gender-related data. Polish listeners were included in the sample as a separate group not only because of the researcher’s background but also because of certain shared features of the Spanish and Polish phonological systems (e.g. the lack of stop aspiration and a similar vowel system), which were hypothesized to lead to a potential mismatched ISIB. See Table 1.
Listeners’ mother tongue and gender.
Regarding the listeners’ English knowledge, the 221 people whose mother tongue was not English had a median (and modal) C1 level in English, although the level distribution was different throughout groups (Table 2) as confirmed by group comparisons (H = 21.004, p < .001). The average period of exposure to English was 21.4 years (SD = 9.5) and the mean age at which the exposure started was 8.1 (SD = 5.8).
Listeners’ level of English (excluding native English listeners).
Notes. *Percentages in the total listener sample (n = 330); **Approximate within-group percentages.
As represented in Table 3, among non-Spanish L1 speakers, 103 declared no knowledge or previous exposure of L2 Spanish whereas 144 did know the language up to varying degrees of competency. The median level on a scale from 0 (NONE) to 6 (C2) was found to be A1, although it is worth noting almost 30% (or 73) of the 247 respondents whose mother tongue was not Spanish were situated at a B1 level or above. Among the 144 speakers who reported some knowledge of Spanish, the average time of exposure to Spanish was 12.1 years (SD = 9.9), while the mean age of onset of said exposure was 18.9 (SD = 8.6).
Listeners’ level of Spanish (excluding native Spanish listeners).
Note. *Percentages in the total sample (n = 330).
3 Scoring and analysis
In order to calculate COM and FA scores, the average values allotted by different (groups of) listeners were used, which resulted in scores with a maximum possible value of 6. The calculation of the intelligibility scores required the design of a detailed protocol (see Gass & Varonis, 1984; Huensch & Nagle, 2021; Munro & Derwing, 1995). A maximum of 2 points were assigned for a correctly transcribed word and 2 points for words with one segmental mistake. However, homophones were considered correct (e.g. meat and meet) as the context could not disambiguate the meaning and point the listener towards the correct spelling. Similarly, certain mistakes which were clearly typos (e.g. due to the letters proximity on the keyboard, such as *gor for got) were considered correct unless they resulted in an existing different word. Given that each speaker’s intelligibility was assessed using 4 sentences including 15 content words in total, the maximum possible per speaker was 30. This score was obtained by averaging the number of points obtained by all the listeners when transcribing that speaker’s sentences. For certain tests, for example, when comparing the differences between scores from different listener language groups, the means of the scores obtained by each speaker (n = 60) from listeners from a specific language background were used. These are analysed as listener variable section because what is indeed analysed here is whether the same speaker received different scores from different listener groups. When testing the impact of listener level, all scores issued by each listener were added up and treated separately as in the case of the listeners’ level of Spanish and English. To measure the relationship between the listeners’ language competency level and intelligibility, an overall aggregate of all intelligibility points obtained by each listener was used. That sum equals the total of points assigned by the researcher to the transcriptions of the SUS of all the speakers each listener assessed. The total was 150, as each listener assessed 5 speakers and for each speaker, a maximum score of 30 points could be assigned. This could be interpreted as ‘the listener’s intelligibility score’, i.e. a reflection of how many phonemes they properly recognized in the SUS they heard, regardless of any speaker-derived variability. While all comprehensibility and foreign-accentedness judgments were considered, certain intelligibility test transcriptions and, subsequently, listener scores were eliminated from the final tests as the listeners did not provide answers matching the task instructions.
The data was processed statistically using SPSS. Both parametric and non-parametric tests were deployed as indicated in the following section and report sizes were calculated. The ANOVA effect size η2 (eta squared) or ηp2 (partial eta squared), when multiplied by 100, ‘indicates the percentage of variance in the dependent variable explained by the independent variable’ (Tomczak & Tomczak, 2014, p. 22). The customary interpretation of the values of η2 and ηp2 is as follows: η2 = .01 small, η2 = .06 medium, η2 = .16 large (Lenhard & Lenhard, 2022). The epsilon squared determination coefficient was the measure of effect size for a Kruskal–Wallis test. According to Tomczak and Tomczak (2014, p. 24), its values are situated between 0 and 1. The former implies no relationship is present while the latter means a perfect relationship. The Mann–Whitney U effect size r, whose value is comprised between 0 and 1, is interpreted as a correlation (Herrera Soler et al., 2011).
III Results
1 Correlations between comprehensibility, intelligibility and foreign-accentedness
The speakers’ overall COM (M = 4.08, SD = .73) and FA (M = 4.13, SD = .87) scores were strongly negatively correlated, r = –.743, p < .001. On The other hand, there was a moderate positive correlation between INT (M = 21.68, SD = 2.9) and COM – r = .445, p < .001 – and a moderate negative correlation between INT and FA, r = –.465, p < .001. Thus, COM and FA were more strongly correlated with each other than either of them was with INT. All the correlations were highly statistically significant.
2 Native listeners and non-native listeners
A paired sample t-test did not show statistically significant differences in the COM scores issued by NSs (M = 4.02, SD = .95) and NNSs (M = 4.08, SD = .72), t(59) = –.879, p = .383, n = 60. For FA, however, a statistically significant difference was found between NS ratings (M = 3.92, SD = .91) and NNS ratings (M = 4.22, SD = .92), t(59) = –4.552, p < .001, d = –.558, n = 60. The NNSs rated the degree of foreign-accentedness higher than the NSs with a medium effect size of the difference between groups. Finally, the INT scores of the NS listeners were significantly higher (M = 22.32, SD = 3.56) than those of the NNS listeners (M = 21.29, SD = 2.8), t(59) = 4.995, p < .001, d = .322, n= 60. The effect size of the difference was small.
3 Listener ratings by mother tongue
A within-subjects one-way ANOVA showed there was a significant effect of the listeners’ mother tongue on COM scores, F(3,177) = 6.029, p < .001, ηp2 = .268, n = 60. The effect size proved to be large. Post-hoc analyses using the Bonferroni correction (Table 4) suggested that the Spanish listeners (M = 4.52, SD = .71) issued significantly higher COM scores than the native listeners (M = 4.02, SD = .95), the Polish listeners (M = 3.96, SD = .97) and those from other language backgrounds (M = 3.77, SD = .83). There were no differences between scores by the Polish and English listeners and the Polish and other language speakers. Less significant differences were also registered between the native group and the group of speakers of other languages.
Post-hoc pairwise comparisons for comprehensibility by listener language (with group means).
Notes. *p < 0.05; **p < 0.01.
Similarly, significant differences were found between FA scores by listeners with different mother tongues, F(3,177) = 8.302, p < .001, ηp2 = .123, n = 60. The effect size of the difference explained by the listeners’ mother tongue proved to be moderate. Pairwise t-tests with the Bonferroni correction (Table 5) suggested the English listeners (M = 3.92, SD = .91) assessed the speakers as significantly less foreign-accented than the Spanish listeners (M = 4.36, SD = 1.1) and those from other language groups (M = 4.14, SD = .91). No other significant between-group differences were found.
Post-hoc pairwise comparisons for foreign accentedness by listener language (with group means).
Notes. *p < 0.05; **p < 0.01.
Finally, differences were also detected between the scores of INT from different groups of listeners with a medium effect size, F(2.43,143.354) = 9.270, p < 001, ηp2 = .136, n = 60. Post-hoc pairwise comparisons (Table 6) showed speakers were rated higher on INT by the native listeners (M = 22.32, SD = 3.56) than by the Polish (M = 21.24, SD = 3.76) and other (M = 20.74, SD = 3.34) listeners. However, there was no statistically significant differences between the native and Spanish listeners (M = 21.88, SD = 2.36), the Spanish (M = 21.88, SD = 2.36) and Polish listeners (M = 21.24, SD = 3.76), or the Polish (M = 21.24, SD = 3.76) and other (M = 20.74, SD = 3.34) listeners.
Post-hoc pairwise comparisons for intelligibility by listener language (with group means).
Notes. *p < 0.05; **p < 0.01.
Analysis of covariance comparing means adjusted for L2 English level differences among the non-native sample groups revealed that the effect was still present even when the level variable was controlled for in the case of all three dependent variables. The effect sizes were moderate for INT (F(2, 201) = 7.9, p < .001, ηp2 = .073) and COM (F (2, 217) = 10.829, p < .001, ηp2 = .073) and small for FA (F(2, 217) = 3.824, p = .023).
4 Non-native listeners’ English level
Only levels B1 to C2 were taken into account as only one listener at an A2 level was present in the sample (0.5% of all non-native listeners). A Kruskal–Wallis test was conducted due to the non-normal data distribution. The results of the test provided evidence for a significant difference in the distribution of INT scores across the four levels, H(3) = 24.904, p < .001, n = 204, Eh R = 0.13, with the following median (mean) scores: B1 Mdn = 95 (M = 100.9), B2 Mdn = 103 (M = 101), C1 Mdn = 110 (M = 109) and C2 Mdn = 114 (M = 112.4). The effect size was moderate. Post-hoc group comparisons showed differences between B2–C1 (p = .005) and B2–C2 (p < .001), but no difference between B1–B2 (p = 1) or C1–C2 (p = .521), for example. See Figure 1

Intelligibility by English level (excluding native English listeners).
Regarding COM (Figure 2), the four groups also differed significantly in their ratings, H(3) = 11.612, p = .009, n = 220, Eh R = 0.05, with the following median (mean) scores: B1 Mdn = 20 (M = 19.62), B2 Mdn = 22.5 (M =21.95), C1 Mdn = 21.5 (M = 20.83) and C2 Mdn = 20 (M =19.01). The effect size was small. Post-hoc tests showed differences between B2–C2 (p = .007), but no difference between any other pair, for instance, B1–B2 (p = .556) or C1–C2 (p = .189). Lastly, no differences were found between the distributions of foreign-accented ratings amongst listeners of different levels, H(3) = 5.926, p = .115, n = 220 (Figure 3).

Comprehensibility by English level (excluding native English listeners).

Foreign-accentedness by English level (excluding native English listeners).
As in the case of group comparisons, in the calculation of the correlation between intelligibility comprehensibility, foreign-accentedness and the English level of non-native listeners, aggregate listener ratings were used. The level of English (B1–C2) of the NNS listeners was found to be positively correlated with their intelligibility scores, rs = .345, p < .001, n = 204, and their comprehensibility scores rs = –.178, p < .001, n = 220. Higher-level listeners’ intelligibility scores were better but the small – albeit significant – negative correlation between comprehensibility and level indicates listeners at higher levels were slightly stricter in their ratings. Finally, no correlation was present between foreign-accentedness ratings and the listeners’ English level, rs = .106, p= .116, n = 220.
5 The knowledge of L2 Spanish
A group comparison was carried out to confirm whether significant differences could be found between those who knew Spanish (YES group) and those who did not (NO group). Native Spanish listeners were excluded from this analysis. Due to the abnormal distribution of the data, non-parametric Mann–Whitney U tests were conducted. INT was the only variable whose measures differed between the two groups (U = 8316.5, p < .001, n = 228, r = .28). It was significantly higher in the YES group (Mdn = 113, M = 111, n = 136) than in the NO group (Mdn = 103, M = 104, n = 92) (Figure 4). The effect size was small reaching medium.

Intelligibility by non-Spanish speaking listeners’ knowledge of Spanish.
Neither COM (U = 8054.5, p = .247, n = 247) nor FA (U = 8000.5, p = .290, n = 247) differed significantly between both groups (Figure 5). The median COM was 20 in the YES group (n = 136) and 20 in the NO group (n = 92), while the median FA was 20.5 amongst those who knew Spanish (n = 236) and 21 amongst those who did not (n = 92).

Comprehensibility and foreign-accentedness by non-Spanish speaking listeners’ knowledge of Spanish.
In search of further evidence for the association between the level of Spanish and the ratings issued by the listeners, non-parametric Spearman correlation coefficients were also calculated. The level of Spanish of the non-Spanish listeners was found to be positively correlated with their INT scores, rs = .318, p < .001, n = 228. However, no positive correlations were found between the listeners’ level of Spanish and their aggregate COM (rs = .046, p = .474, n = 247) or FA (rs = .108, p = .089, n = 247) scores.
6 Summary of the findings
Table 7 displays the most relevant findings as far as listener-related variables are concerned. The gender variable was added here, although the tests showed that no significant differences between genders existed and it will not be further discussed. INT was the most affected variable as it was associated with four different variables in the group comparison tests.
Summary of listener group difference tests.
Notes. FA = foreign-accentedness; COM = comprehensibility; INT = intelligibility. **p < 0.01.
Table 8 recapitulates the correlation between listener level variables and the dependent variables. As can be seen, these mirror the group comparison results presented above as the level of language English competence was correlated COM and INT and the level of Spanish – with INT only. In both cases, the correlations are stronger for INT.
Summary of correlations between listener language levels and dependent variables.
Notes. FA = foreign-accentedness; COM = comprehensibility; INT = intelligibility. **p < 0.01.
IV Discussion
While all our speaker informants came from the same language background as native speakers of Spanish, the listener group was much more varied. Munro et al. (2006, p. 125) showed little to no listener effects, suggesting the listener’s impact on INT, COM and FA scores was minor (p. 129) and highlighting the need for further research. However, Wang’s (2007, pp. 256–257) assertion that listener background is crucial seems to be broadly supported by the present research, where the general question whether the listeners’ linguistic background had a bearing on their judgments was answered affirmatively.
The native judges differed from non-natives on the scores they assigned for foreign-accentedness and the ones they obtained on intelligibility tests. When split into two big native-non-native groups, no differences were found regarding comprehensibility, meaning that both the native and non-native group found the speakers in the sample equally easy or difficult to understand. Differences were present, however, in the listeners’ foreign-accentedness and intelligibility judgments. Surprisingly, the non-native judges were significantly stricter (M = 4.23/6) than, on average, the native listeners (M = 3.92/6) suggesting that native speakers may be more tolerant of foreign-accented speech than non-native speakers in line with previous research (Fayer & Krasinski, 1987; Isaacs & Thomson, 2013; Kang, 2012). Regarding intelligibility, the native listeners had a very slight advantage (1 point out of 30) when decoding SUS on the online tests, as confirmed by the small effect size (d = .322). Before suggesting any tentative explanations of this finding, the exact differences between groups should be considered. Comparing the average scores issued by the four language groups, significant differences were found for all three dependent variables: intelligibility, comprehensibility and foreign-accentedness. These results lend support to the hypothesis that the specific language background of the listener has a bearing on their intelligibility scores and comprehensibility and foreign-accentedness ratings (see Gass & Varonis, 1984; Wang, 2007). However, no significant differences were found between the Polish and Other groups, which is why both may be treated as one homogeneous listener group.
While the overall effect size of the differences of averaged comprehensibility ratings from the four language groups were large, it was the Spanish listeners (M = 4.52/6) who differed most significantly from each of the three other groups (English, Polish and Other listeners), where the average score was situated between 3.79 and 4.02. This might suggest that familiarity with the way English is spoken by other speakers of one’s language leads to an increase in the perceived ease of understanding (see Saito & Shintani, 2016). Yet, it should be stressed that the subjective impression of an individual may not always match their objective reality and real understanding/decoding of the message. What is evident is that familiarity may at the very least make speech ‘easy on the ear’. Thus, the advantage resulting from speakers sharing the same interlanguage when understanding a foreign language (Bent & Bradlow, 2003; see interlanguage speech intelligibility benefit referred to further on in this section) could be extended onto comprehensibility judgments. This parallel might even be taken further to posit the existence of a partially independent interlanguage speech comprehensibility benefit. It is also worth noting that the lowest comprehensibility marks were allocated by the group comprising speakers of languages other than English, Spanish and Polish (M = 3.79). Nevertheless, these last results are hard to interpret conclusively due to the linguistic heterogeneity of the Other group as well as the difference being less significant (p = .033) than the ones previously mentioned. As explained earlier, when comparing NSs and NNSs as two broad groups, no difference was found. This is easily accounted for by the group comparisons where the most significant and largest differences were present between the Spanish listeners and all the other groups (all p < .001). The Spanish listeners’ scores suggest they had the least reported difficulty understanding the speaker sample.
The mother tongue of the sampled listeners also had an impact on foreign-accentedness ratings, although the effect was moderate. The Spanish listeners, the same who found the speakers easiest to understand, were the ones to assess foreign-accentedness in the strictest fashion (M = 4.26/6) when considering the average rating per speaker. In other words, the speakers obtained the highest average marks for foreign-accentedness from their fellow Spanish-speaking listeners. As only Spanish natives were recorded and analysed in this study, it is difficult to predict whether the same would apply to other language configurations, e.g., Polish speakers assessing Polish speakers. The findings are, however, consistent with a popular belief held by Spanish people themselves about the poor quality of their English (see Cutillas Espinosa, 2017; El Confidencial, 2014; Galván, 2010; La Razón, 2019) alongside previous research on the impact of rater variables on accentedness (see Fayer & Krasinski, 1987; Isaacs & Thomson, 2013; Kang, 2012) including those studies that directly link attitudes towards non-native speech with variance accentedness judgments. Although a foreign accent is not inherently wrong or something to avoid or to be reduced at all costs – as long as it is intelligible (see Jenkins, 2000) – it is striking that the native speakers were the most lenient raters (M = 3.92), followed by the Polish and other listeners (both groups with M = 4.12), and Spaniards (M = 4.26), which implies that NSs are more tolerant of non-native speech than other NNSs. Finally, further evidence to support these findings is the fact that although overall foreign-accentedness and comprehensibility scores per speaker are highly negatively correlated (r = –.743, p < .001), the ratings obtained from Spanish listeners only moderately correlate with each other with a considerable drop in statistical significance (r = –.299, p = .02). What is clear is that our results show effect of a shared mother tongue as a factor leading to less favourable accent ratings in the Spanish context. In brief, Spanish speakers find their fellow Spaniards easy to understand but highly accented, which may be at least partially due to attitudinal variables (see Ingvalson et al., 2017, Simon et al., 2022).
As expected, the analyses showed slight yet statistically significant differences in the intelligibility scores by different language groups for each speaker. Considering the comparably similar means, those deviations could have been missed. However, the ANOVA test revealed that the exact differences were not only highly significant but also actually moderate in size and far from negligible. However, there were no differences between native listeners and Spanish listeners (p = 1). Neither were significant differences revealed between Polish and Spanish listeners’ scores. In the following ranking list from best scores to worst scores per speaker: (1) English listeners’ INT score (M = 22.32) (2) Spanish listeners’ INT score (M = 21.88) (3) Polish listeners’ INT score (M = 21.24) (4) Other listeners’ INT score (M = 20.74)
no differences were detected between adjacent groups. Thus, (1) was different from both (3) and (4) but not from (2), and so on. This might be due the small visible difference in means. Nonetheless, it should not be overlooked that those differences that did exist were highly significant, although the boundaries between the groups may not be clear-cut, just as it is in the case of speaker levels previously described. One possible reason why intelligibility scores were not better differentiated may have been the point assignment methodology. Had a more radical method been employed considering words as either fully correct or downright wrong, the results may have been different. The fact that English listeners were the most efficient at decoding speech in the sample could be explained by those listeners’ native proficiency. An alternative explanation might be the higher presence of top-down strategies in speech processing amongst native speakers (see Jenkins, 2000, p. 90; Jenkins, 2002, pp. 89–90). Although our intelligibility tests were devoid of semantic cues, the syntactic context was still present, which might possibly have helped native listeners to use more non-phonological cues than non-natives. Regarding Spanish speakers, the most appealing tentative explanation is that these results also fit with the working hypothesis that a matched interlanguage speech intelligibility benefit (ISIB) (Bent & Bradlow, 2003) does exist as the Spanish listeners in this study were better at transcribing Spanish-accented utterances that those of other nationalities, although none of them outperformed native listeners. This is different from comprehensibility, which was actually highest among the Spanish listeners, suggesting that the perceived comprehensibility is more affected by a shared L2 interlanguage than more objectively measured intelligibility, as confirmed not only by the position of the Spanish listener group in the ranking (first in comprehensibility and second in intelligibility), but also by the fact that the effect sizes which were larger in the case of between-group comprehensibility comparisons. All that said, it was beyond the scope if this study to make comparisons between speakers from different backgrounds so the evidence supporting the existence of ISIB is necessarily partial.
Leaving a restrictive interpretation of the statistical analyses aside, a simple practical implication should be put forward whereby the segmental intelligibility differences would probably be hardly noticeable with the naked eye in everyday student performance. Although tendencies exist as confirmed by tests, the difference between the highest scoring (English listeners: M = 22.32) and the lowest scoring (Other listeners: M = 20.74) group is 1.5 out of 30. Spanish listeners (M = 21.88) ranked second and Polish listeners third (M = 21.24). This might be tentatively and indirectly interpreted as a confirmation of Jenkins’ (2000, 2002) assumption – albeit not driven by quantitative tests and based solely on classroom interaction observation – that intelligibility amongst non-native speakers relies considerably more on phonological signal (bottom-up) than amongst native speakers, who rely more on contextual signal (top-down). The bottom line is that, when deprived of semantic context (accurate syntax was preserved) and forced to rely almost exclusively on phonological signals, native listeners were at a relatively small advantage over non-natives. Still, more quantitative research is needed to empirically corroborate Jenkins’ words given that this phonologically-oriented study is not concerned with performance differences between high-context/contextualized and low-context or decontextualized speech (see Bergeron & Trofimovich, 2017).
The second listener variable examined was the level of English competence of non-native speakers. It was expected that higher level listeners would do better decoding the provided input. Correlation tests provide clear evidence for a relationship between listener level and their intelligibility and comprehensibility judgments but not accentedness (see Eger & Reinisch, 2019; Van Wijngaarden et al., 2002). In the former case, the correlation coefficient indicated a moderate effect (rs = .345) while in the latter, a small one (rs = –.178). Yet, group comparisons reveal a much more intricate picture of this relationship. The intelligibility task performed by listeners for this study might be considered a comprehension check in a broad sense, although focused on phonological retrieval. Thus, it should come as no surprise that in group comparisons the level of English competence of those listeners whose mother tongue was not English was found to be clearly related to their performance with a moderate effect size. There was a clear difference between B2 and both of the higher levels (C1 and C2) but no other differences suggesting that there is a cut-off point whereby a significant increase in phonological intelligibility existed above B2, but was not detected across B1–B2 or C1–C2. In fact, the means were virtually the same at B1 (M = 100.9) and B2 (M = 101) and increased significantly at C1 (M = 109) and C2 (M = 112.4). The last two were not significantly different in our statistical analysis although the means themselves and the graphic representation of data do reflect a progressive pattern.
The listener level variable also affected one of the ratings measured by scales: comprehensibility. Foreign-accentedness scores were not sensitive to the issuers’ mother tongue suggesting NNSs at all levels have similar skills when gauging a person’s level of foreign accent. When it came to comprehensibility, only a small effect was registered and the only between-group comparison which rendered significant results was B2–C2. Surprisingly, it was the B2 level group which gave better comprehensibility scores. This might suggest that higher-level listeners were stricter when evaluating comprehensibility than lower-level listeners. However, analysing the means (and medians) of the grades, the following hierarchy from most lenient to strictest scores by level emerges: B2 > C1 > B1 > C2. The medians of B1 (Mdn = 20 out of 30) and C2 (Mdn = 20 out of 30) were identical. A possible explanation of the presence of B1 listeners amongst the low raters alongside the highest-level listeners (C2) were real comprehension problems. It could be tentatively ventured that while higher level students may have problems with the quality of the input and hence issue lower scores, in the case of B1 listeners, real comprehension problems may arise. This hypothesis would require further research – both quantitative and qualitative – into the motives behind the listeners’ choices.
Let us now turn to the hypothesis that the listeners' knowledge of Spanish could also be a relevant factor given the speaking sample’s Spanish background. Indeed, there was a significant (p < .001) and a small-reaching-moderate (r = .28) difference in intelligibility in favour of those who declared at least some knowledge of Spanish (Mdn = 113 vs. Mdn =103 out of 120). Foreign-accentedness or comprehensibility were not affected as shown by previous studies (see Bergeron & Trofimovich, 2017). This implies that knowledge of Spanish as a foreign-language is associated with a better understanding in the sense of phonological retrieval and recognition but not with a better or worse subjective assessment of Spanish-accented English by non-Spanish speakers. That variable, however, could be related to others, such as the time of residence in a Spanish speaking country and exposure to Spanish-accented English, which were not controlled for in this study. In spite of this limitation, what can be concluded from the data is that being familiar with the language – albeit imperfectly – was positively associated with the ability to decode Spanish-accented English speech. Ultimately, correlations between Spanish level and intelligibility (but not comprehensibility or foreign-accentedness) also proved to be significant and the coefficient was moderate, meaning that the higher the listeners’ level of Spanish, the more probable it was for them to score better on intelligibility tests (rs = .318, p < .001), which clearly matches the conclusions drawn from the group comparison tests.
Finally, neither the English language level nor the level of Spanish determined foreign-accentedness. This is aligned with previous research where speakers were found to be able to distinguish native from non-native speakers even in languages they were unfamiliar with (see Major, 2007), as there were no differences between accent rating amongst students with previous exposure to the language and the ones without. However, the study in question also suggests the listeners’ first and second languages (English native speakers, Brazilian Portuguese native speakers, i.e. L2 English speakers in this case) has little bearing on foreign-accentedness. As we have seen before, the mother tongues in our study did have a clear moderate effect on that variable (see Winke & Gass, 2013).
V Conclusions
In conclusion, the data support an association between listener-related factors and the dependent variables of intelligibility, comprehensibility and foreign-accentedness to varying degrees, which corroborates the importance of factoring in listener characteristics into pronunciation studies dealing with intelligibility, comprehensibility and foreign-accentedness (see Gass & Varonis, 1984; Pickering, 2006; Sewell, 2010; Trudgill, 2008; Zielinski, 2008)
Intelligibility – conceptualized as the decoding of auditory input for the purpose of this pronunciation-centred study – was quite naturally expected to be related to the listeners’ linguistic background. The listeners’ (1) native status, (2) mother tongue group, (3) English level and (4) knowledge of Spanish had small to medium significant effects on INT scores. COM was also affected by the listeners’ mother tongue and their English level, while FA – by the listeners’ mother tongue only. The fact that intelligibility is affected by the listener variables under scrutiny further stresses the role of the listener in effective communication and understanding and suggests that certain pressure can be taken off the speakers in their attempts to make themselves understood inside and outside the classroom. In Zielinsky’s (2008) terms, the listener is not a ‘silent partner’ or a passive receiver of the input and familiarity with the speaker’s pronunciation patterns enhances intelligibility while lowering the effort reported when decoding the speech signal. Broadly speaking, this research also makes a case for the existence of both an interlanguage speech intelligibility and comprehensibility benefit (see Bent & Bradlow, 2003, Ludwig & Mora, 2017; Shu et al., 2016; Wang, 2007; Xie & Fowler, 2013). Regarding the differences between listeners from different language backgrounds, one of the most interesting findings was that native speakers were less strict when assessing accent. Spanish speakers, in turn, were the strictest in the assessment of accent (see Fayer & Krasinski, 1987; Isaacs & Thomson, 2013; Kang, 2012), although their COM scores were the highest and their INT scores were similar to those of native speakers. The data, thus, confirm the existence of an interlanguage speech intelligibility benefit (Bent & Bradlow, 2003; Wang, 2007) – although natives are still at an advantage – and suggests that a shared mother tongue made the speakers subjectively easier to understand (a hypothesized matched interlanguage speech comprehensibility benefit). On the other hand, certain attitudinal variables may play a role in the issued ratings as suggested by the strictest accentedness scores from the Spanish listener sample. Differences in attitudes have been shown to affect accentedness scores in previous studies (see Ingvalson et al., 2017, Simon et al., 2022) and may tentatively explain some variance in the accentedness scores in this article. Surely, further research would be needed to gain insight into the purported attitudinal factors behind the listeners’ decisions.
Footnotes
Appendix A
Elicitation paragraph (originally designed for this research)
The sun was rising slowly and the birds were singing. The view from the top of the hill was amazing. Susan turned round and stepped heavily into the kitchen. She took a plastic cup and filled it with juice. Then, she sat down on a chair and started thinking what food she would cook for lunch. Suddenly, her dog, Zoe, jumped onto her, spilling her drink. She threw the empty cup away into a small waste bin. She hadn’t been feeling good for years. Her job as a university nurse didn’t give her pleasure and she hated starting in the early morning. At her age, it wasn’t easy to deal with people. As she couldn’t hear well, she was getting used to being on her own. She decided to skip work, stay in and enjoy the first day of spring. Then she remembered it was Saturday anyway.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
