Abstract
Scholars have argued for the inclusion of different spoken varieties of English in high-stakes listening tests to better represent the global use of English. However, doing so may introduce additional construct-irrelevant variance due to accent familiarity and the shared first language (L1) advantage, which could threaten test fairness. However, it is unclear to what extent accent familiarity and a shared L1 are related to or conflated with each other. The present study investigates the relationship between accent familiarity, a shared L1, and comprehensibility. Results from descriptive statistics and Mann–Whitney U test based on 302 second language (L2) English listeners’ responses to an online questionnaire suggested that a shared L1 meant high accent familiarity, but not vice versa. A path analysis revealed a complex relationship between accent familiarity, a shared L1, and comprehensibility. While a shared L1 had a direct effect on accent familiarity, and accent familiarity had a direct effect on comprehensibility, a shared L1 did not predict comprehensibility when accent familiarity was controlled for. These results disentangle accent familiarity from a shared L1. Researchers should consider both constructs when investigating fairness in relation to World Englishes for listening assessment.
Introduction
Language assessment scholars have argued that a more diverse set of spoken varieties of English should be used in listening proficiency tests when they are construct-relevant to represent the global use of English and thus to improve the ecological validity of listening tests (e.g., Harding, 2012). Nonetheless, testing agencies have been hesitant in making material changes in this regard. Specifically, many popular high-stakes language tests such as International English Language Testing System (IELTS), Test of English as a Foreign Language (TOEFL iBT), and Duolingo English Test (DET) only introduce inner-circle varieties (e.g., American, British English accents), and, to a lesser extent, outer- and expanding-circle varieties of English (see, for example, IELTS, n.d.-a, for publicly available sample listening tests). A noticeable exception is the Pearson Test of English (PTE, n.d.), which includes many different spoken varieties in its listening test. Some testing agencies have recently commissioned projects on fairness issues (e.g., Duolingo, 2020), that is, whether listening tests incorporating different spoken varieties would unduly disadvantage some test takers over others due to individual differences based on, for example, their first language (L1) background and familiarity with a particular accent variety. Nonetheless, there have been few changes made to the existing tests incorporating different spoken varieties thus far.
One major reason behind the lack of changes is perhaps the fear of introducing construct irrelevant variables such as test takers’ familiarity with a particular accent, which poses a threat to test validity and fairness (see Taylor & Geranpayeh, 2011). Accent familiarity has been defined as “a speech perception benefit developed through exposure and linguistic experience” (Browne & Fulcher, 2017, p. 39). However, the operationalizations of accent familiarity have been very different from study to study, including listeners’ ability to distinguish a particular accent variety from others (see Huang, 2013), their proficiency in speakers’ L1 (see Winke et al., 2013), and their self-reported contact and exposure to different varieties (see Ockey & French, 2016). This inconsistency reflects an incomplete understanding of accent familiarity from a theoretical and empirical perspective.
Related to accent familiarity is interlanguage speech intelligibility benefit or shared L1 advantage (see Bent & Bradlow, 2003), a concept claiming that listeners understand a second language (L2) speaker better if they share the same L1. This was empirically investigated in previous research, albeit with mixed findings. When test takers listened to speakers with the same L1, they sometimes (but not always) performed better than listeners with other L1s (see Kang et al., 2019; Major et al., 2002; Munro et al., 2006; Smith & Rafiqzad, 1979). In interpreting and explaining the findings, scholars often noted that with a shared L1, listeners were familiar with the accent, which explained their superior performance. Nonetheless, it was not clear to what extent accent familiarity and a shared L1 relate to or are conflated with each other. Specifically, Shin et al. (2021) found that across different L1 groups, the relationship between a shared L1 and accent familiarity was not straightforward. That is, people without shared L1 may also demonstrate high accent familiarity. This suggests that a shared L1 and accent familiarity may be two related but different constructs. Accordingly, the present study seeks to explore the relationship between accent familiarity and a shared L1. Moreover, it explores the relationship among accent familiarity, a shared L1, and comprehensibility (i.e., how easy a speaker is to understand; see Munro & Derwing, 2020) to reveal the differential effect of the two constructs on comprehensibility.
Clarifying the relationship between a shared L1 and accent familiarity in relation to comprehensibility is important for better describing issues related to fairness in listening assessment. Specifically, both a shared L1 and accent familiarity can potentially contribute to listening comprehension (see, for example, Bent & Bradlow, 2003; Ockey & French, 2016). However, it is unclear which construct (or a combination of both) could capture this more effectively. This study aims to provide a sounder theoretical foundation as to ways in which fairness can be maintained when incorporating different spoken varieties of English in high-stakes tests (e.g., in terms of how to select speakers for recordings for research as well as operational assessments).
Literature review
This section provides a conceptual and empirical review of relevant studies. First, it introduces the constructs of accent familiarity, a shared L1 advantage, and comprehensibility regarding their definitions and operationalizations in research contexts, encompassing studies focused on both speakers and listeners. Second, it reviews relevant empirical studies exploring the relationship among these constructs. This empirical review narrows its scope and solely discusses papers relevant to listeners, which is the focus of the present study.
Constructs, definitions, and operationalizations
Accent familiarity
The construct of accent familiarity has appeared in many studies investigating L2 speaking and listening (e.g., Huang, 2013; Ockey & French, 2016; Saito et al., 2019). However, most of these studies did not explicitly provide a definition of the construct. The lack of a clear definition might have resulted in the inconsistent operationalizations of accent familiarity in research as outlined above. The present study relies on a definition from Browne and Fulcher (2017), who synthesized different measurements of accent familiarity and defined it as “a speech perception benefit developed through exposure and linguistic experience” (p. 39). However, this definition has many limitations. First, it seems to describe the outcome of high accent familiarity (i.e., perception benefit), but not the internal structure of the construct itself (e.g., whether it is a multifaceted phenomenon). Second, it is not transparently operationalizable. In other words, it does not itself identify ways to quantify accent familiarity for research purposes. Third, the definition seems to make a claim that accent familiarity is associated with speech perception benefits. However, this claim is not always supported by empirical evidence, as demonstrated in the overviewing of the background literature below. These limitations reflect a developing understanding of the construct of accent familiarity. The present study includes this definition, acknowledging its many limitations, and treating it as a tentative hypothesis to be tested empirically. Below is a review of the many diverse and inconsistent measures of the construct in the field of applied linguistics (including language testing and speech perception research), which further highlight researchers’ incomplete understanding of the construct of accent familiarity itself.
For instance, Huang (2013) investigated whether raters’ familiarity with different English accents would influence their proficiency ratings of speakers in a language testing context. Huang asked listeners to guess the speakers’ L1 as one way of quantifying accent familiarity. Bearing a similar research question, Winke et al. (2013) quantified accent familiarity using listeners’ self-reported proficiency in the speakers’ L1. This inconsistency of how to operationalize accent familiarity extends from the language testing literature to general speech perception studies. Gass and Varonis (1984) investigated the relationship between accent familiarity, speaker familiarity, and comprehensibility. In their study, listeners were deemed familiar with a particular accent if they had listened to another speaker with the same L1 once during the experiment. Similarly, Matsuura et al. (1999) and Nejjari et al. (2012) categorized listeners with different accent familiarity based on their prior contact and exposure. Lastly, an often-used measure of accent familiarity is listeners’ self-reported familiarity with a particular accent variety, as measured using a Likert-type scale (see Saito & Shintani, 2016; Shin et al., 2021).
In sum, accent familiarity has attracted many scholars’ attention and is commonly researched within the fields of speech perception and language testing. However, there does not seem to be a consensus as to how to operationalize this construct. Without a validation study, taking into account many, if not all of these measures, it is difficult to evaluate what the best measures would be in a given context. This study adopts the Likert-type scale approach for measuring accent familiarity following many precursor studies (e.g., Saito & Shintani, 2016; Shin et al., 2021).
Shared L1 advantage
A relevant construct similar to accent familiarity is the shared L1 advantage, which was originated from the seminal work of Bent and Bradlow (2003). In the study, the researchers found evidence suggesting that a shared L1 was associated with an advantage in understanding L2 speech, but the result was inconclusive, as discussed below. Following Bent and Bradlow, many studies undertook similar investigations but yielded similarly mixed findings (e.g., Abeywickrama, 2013; Kang et al., 2019; Major et al., 2002; Smith & Rafiqzad, 1979). Overall, it was found that a shared L1 sometimes, but not always, contributes to better listener understanding. It is therefore possible that there are other potential confounding variables at play.
Regarding its measurement, a shared L1, by its name, is relatively unambiguously operationalized in research studies compared with accent familiarity. Scholars typically ask listeners to report their L1s and match them with the speakers’ L1 (e.g., Munro et al., 2006). However, one noticeable caveat of this measure is that it does not seem to consider the extent to which a listener is multilingual. To elaborate, it is possible that listeners’ (or speakers’) dialects and additional languages could have influenced the way they perceive and produce L2 speech, resulting in some “noise” in the data (e.g., L1 Mandarin listener who also speaks L1 Cantonese). However, the shared L1 has many limitations as a construct. First, it is a dichotomous variable that could limit the analyses that can be performed. Moreover, the shared L1 appears to be an overly broad construct. Listeners with or without a shared L1 could differ otherwise in terms of a wide range of individual differences. For this reason, it is possible that this construct alone could not adequately address fairness issues associated with World Englishes in listening assessment when listening tests incorporate different spoken English varieties. Nonetheless, this construct remains a relatively straightforward one to operationalize in research and testing contexts.
Understanding: Comprehensibility and intelligibility
In the L2 speech literature, the two currently dominant constructs probing understanding of spoken discourse are comprehensibility, focusing on listeners’ efforts (i.e., how difficult a speaker is to understand; see Munro, 2018) and intelligibility, focusing on utterance (generally word) recognition (i.e., actual understanding; see Munro, 2018). Comprehensibility has been conventionally measured via listeners’ self-reports on a Likert-type scale (see, for example, Crowther et al., 2018; Munro & Derwing, 1995). Saito (2021) in a meta-analysis found that different groups of raters (e.g., experts, laypeople, L1 and L2 English users) were all able to demonstrate good inter-rater reliability when they evaluated speakers’ comprehensibility using rating scales. However, reliability does not imply validity (Bachman & Palmer, 2010). Specifically, it is not clear to what extent one stand-alone Likert-type scale question could capture the construct of comprehensibility. Despite this noticeable methodological caveat, this measure remains popular in recent studies (e.g., Saito et al., 2019).
Compared with the relatively consistent operationalization of the comprehensibility construct, there are diverse approaches to measuring intelligibility. Of these, a most common one is speech transcription, where participants listen to many (usually short) speech samples and transcribe them verbatim (see, for example, Kang et al., 2019; Munro et al., 2006). The methodological limitations of a transcription task include the (a) assumption that mis-transcription of words equals misunderstanding (whereas in theory, understanding is constructed through not just linguistic resources, but also world knowledge; see Flowerdew & Miller, 2005) and (b) lack of generalizability to discourse-level speech because transcription tasks are time-consuming, so many researchers opt for shorter speech samples (e.g., Munro & Derwing, 2020).
Acknowledging these caveats, researchers have sought for alternative measures of intelligibility, including cloze dictation (Matsuura, 2007), and listening comprehension questions (Major et al., 2002). Kang et al. (2018) provide a good summary of more measures of intelligibility, which is beyond the scope of this article. In light of the diverse and inconsistent measures of intelligibility, the present study uses the construct of comprehensibility to probe listener understanding of L2 speech.
Empirical interrelationships among constructs
Accent familiarity and understanding
The relationship between accent familiarity, a shared L1, and understanding (i.e., comprehensibility and intelligibility) is complex and inconclusive in the current literature. While some studies have found that accent familiarity or a shared L1 positively contribute to understanding (e.g., Gass & Varonis, 1984; Ockey & French, 2016), other studies observed the opposite (e.g., Abeywickrama, 2013; Smith & Rafiqzad, 1979) or have produced mixed findings (e.g., Kennedy & Trofimovich, 2008; Matsuura et al., 1999). Below is a review of a sample of these studies.
Gass and Varonis (1984) investigated the relationship between accent familiarity and word recognition. They found that after L1 English listeners had listened to L2 speakers from one accent variety in the context of the experiment, they transcribed the words of subsequent speakers of the same variety more accurately than did listeners without such exposure. Similarly, Ockey and French’s (2016) large-scale study investigated the relationship between L2 English listeners’ accent familiarity and their performance in simulated tests of TOEFL iBT incorporating different regional varieties of English (i.e., American, British, and Australian English). Accent familiarity was measured using multiple Likert-type scales exploring test takers’ contact and exposure to these varieties in different contexts (although internal consistency across items was not reported). Based on this, listeners were categorized into two groups: familiar or unfamiliar. The results suggested that listeners’ accent familiarity was positively related to their test performance. This finding is especially relevant due to their similar measure of accent familiarity compared with the present study (i.e., via Likert-type scales).
Despite the positive findings observed above, other studies presented a mixed picture. For example, Kennedy and Trofimovich (2008) found that L1 English listeners’ accent familiarity (operationalized as experience in L2 English teaching) was related to their accuracy of word transcription (i.e., intelligibility) but not to perceptual judgments of comprehensibility. On the contrary, Matsuura et al. (1999) found that L2 English listeners’ accent familiarity (operationalized as classroom exposure) was related to comprehensibility, but not intelligibility. Overall, the relationship between listeners’ accent familiarity and their understanding of L2 speech is inconsistent and unclear. The different findings can be explained by (a) different operationalizations of comprehensibility or intelligibility) and (b) different measures of accent familiarity (e.g., reported familiarity, reported exposure). In other words, the different methodological choices in terms of construct operationalization may limit the comparability of results across studies.
Shared L1 and understanding
The previous section has revealed a relatively unclear relationship between accent familiarity and different measures of understanding. Also unclear is the relationship between a shared L1 and understanding. Bent and Bradlow’s (2003) article has often been cited as evidence to support the positive relationship between a shared L1 and understanding, but this study itself presents mixed findings. In the study, L1 Mandarin and L1 Korean listeners transcribed L1 English recordings by L1 Mandarin and Korean speakers of varying English proficiency. The shared L1 advantage was not observed in L1 Mandarin listeners. However, it was observed in L1 Korean listeners, and this effect was more pronounced when they listened to low-proficiency speakers (see also Stibbard & Lee, 2006).
Kang et al. (2019) investigated L2 listeners’ performance in TOEFL iBT listening tasks recorded in different accent varieties. Results provided mixed evidence for the shared L1 advantage. Specifically, only listeners from India and South Africa benefited from speakers with a shared L1, but not listeners from China and Mexico. A similarly inconclusive result was observed in Major et al. (2002), in which the researchers explored the shared L1 advantage in TOEFL listening tasks. The researchers argued that because accent familiarity might improve comprehension, listeners would understand speakers with a shared L1 better. The results suggested that their L1 Spanish listeners demonstrated the shared L1 advantage by performing better in comprehension questions than other listeners from a non-Spanish background. However, this shared L1 advantage did not generalize to other listener groups (i.e., L1 Mandarin and L1 Japanese). Similarly, many studies suggested a mixed picture where the shared L1 advantage was sometimes, but not always observed (see, for example, Harding, 2012; Stibbard & Lee, 2006). Yet other studies did not appear to find support of the shared L1 advantage (e.g., Abeywickrama, 2013; Butler, 2007). The discrepancy of the findings can be attributed to (a) random variation small listener sample size (e.g., 10 listeners per accent variety in Kang et al., 2019) and (b) lack of control for speakers’ dialect.
Accent familiarity and shared L1 advantage
The studies cited above have revealed that the relationship between accent familiarity and a shared L1 advantage is underexplored, and their relationships with understanding constructs (i.e., intelligibility, comprehensibility) need to be clarified. Accent familiarity and the shared L1 advantage appear to be similar constructs that are defined and operationalized differently in these studies. However, the extent to which a shared L1 advantage and accent familiarity are related to or conflated with each other remains underexamined.
It is possible that previous research intended to use a shared L1 as a proxy for accent familiarity. This is especially relevant because accent familiarity has not been widely understood as a construct. While studies have sought to conceptualize comprehensibility (e.g., Isaacs & Trofimovich, 2012) and intelligibility (e.g., Field, 2005), exploring what variables are associated with these constructs, little is known about accent familiarity. A shared L1, despite many notable methodological limitations, was still popularly used perhaps due to ease of operationalization in research contexts. Nonetheless, this argument is hypothetical, and construct-oriented studies regarding accent familiarity could provide a fuller understanding of accent familiarity and its many different measures.
On one hand, from the arguments made by Kang et al. (2019) and Major et al. (2002) among others, the shared L1 advantage and accent familiarity often seem to co-occur. That is, listeners having a shared L1 with a given speaker tend to be familiar with the speakers’ accent variety. This claim makes intuitive sense, as it is possible that listeners develop their familiarity with their accent varieties via extensive exposure in educational or naturalistic contexts. In short, a shared L1 is likely to be highly correlated with accent familiarity. Nonetheless, this claim warrants empirical support, as it appears that no study has, as yet, directly investigated the relationship between these two constructs.
On the other hand, it was argued that listeners without a shared L1 could also develop familiarity with a particular variety via contact and exposure (see, for example, Gass & Varonis, 1984; Matsuura, 2007). Therefore, there is reason to believe that accent familiarity is developed not just through a shared L1, but potentially through other channels as well (e.g., via social interactions). In other words, it is possible that the construct of accent familiarity encompasses much, if not all, of the shared L1 advantage. This claim again warrants empirical evidence, and further understanding of this claim would reveal the relationship between these constructs with implications for language assessment and speech perception.
An alternative way of looking at the relationship between a shared L1 and accent familiarity is that both constructs have the potential to explain listeners’ improved comprehension (if any) of a particular spoken variety. What is unclear, however, is whether one is more effective in capturing the improved comprehension than the other, which the current study sets out to explore.
Summary and a proposed model
The section above reviewed empirical evidence regarding the relationship between accent familiarity, the shared L1 advantage, and understanding. Two gaps in knowledge were evident from this review. First, it was unclear to what extent accent familiarity and a shared L1 related to or conflated with each other, despite scholars’ claims that a shared L1 could contribute to accent familiarity. Second, it was unclear to what extent accent familiarity or a shared L1 were related to comprehensibility. The present study seeks to address these two questions with a proposed model (see Figure 1), synthesizing existing knowledge from the literature, and then empirically tests this model. The model proposes the following hypotheses:
A shared L1 is likely associated with comprehensibility. However, the evidence presented in the literature is inconclusive.
A shared L1 is likely associated with accent familiarity. This claim is oftentimes assumed, but not empirically tested.
Accent familiarity is likely associated with comprehensibility. However, the evidence presented in the literature is inconclusive.
Other variables, such as speakers’ linguistic repertoire, are likely associated with comprehensibility.
The present study is guided by the following research questions:
What is the relationship between L2 English users’ accent familiarity and a shared L1?
What is the relationship among accent familiarity, a shared L1, and comprehensibility?

Hypothesized relationship among accent familiarity, a shared L1, and comprehensibility.
Participants
Listeners
Convenience sampling and snowballing sampling were used to recruit listening participants. The research team introduced the online survey to potential participants at two Masters-level programs in applied linguistics and language education at the University of Oxford and at other UK universities. Adult L1 and L2 English users with no self-reported hearing difficulties were invited to take the survey, although this manuscript only reports on results from the L2 participants.
To maximize implications of the study for operational language testing in higher education, the ideal population is L2 English users who take English language proficiency tests and who do not have a degree from an English-speaking country, although this criterion can be differently applied across institutions. However, the present sample only included L2 English users with at least high-intermediate proficiency. The main reason why the research team decided to exclude lower-proficiency participants, as attested through participants’ self-perceived proficiency and self-reported proficiency scores (if any; see below for details), was that some contacts with a lower proficiency reported comprehension difficulties with the baseline recording, which was used to control for listener severity with a between-subjects design.
After excluding 29 recruited participants from further analysis on the grounds of proficiency, a final sample of 302 L2 English users were retained for inclusion in the study (86 male, 214 female, 1 non-binary; Mage = 25.25, SD = 5.98). Of all the participants, 127 (42.05%) reported having language teaching experience, and 143 (47.35%) reported having taken language/linguistics-related classes at university, of which 70 (48.95%) were undertaking or had received a linguistics-related degree. In terms of L1 background, 103 (34.11%) were L1 Mandarin speakers; the remaining 199 (65.89%) participants reported diverse L1s as detailed in Supplementary Material 4. Overall, participants reported a wide range of familiarity with Mandarin-accented English (1 = very unfamiliar, 9 = very familiar; M = 5.61, SD = 2.70). All participants had at least intermediate proficiency in English. This was indicated achieving an above-B2 proficiency in the CEFR level based on their self-reported English proficiency scores (operationalized as an IELTS score of above 6.0, TOEFL iBT of above 60, TOEFL CBT of above 170, and TOEIC of above 570; see The Edge, n.d.; ETS, n.d., 2010) or self-reported proficiency of above 7.0 on a 9-point scale (see Miao, 2020, for further details).
All participants were randomly assigned to one of the six experimental groups. Within each group, listener participants listened to (a) the same baseline recording and (b) a different recording controlled for content but manipulated by accent (North American, moderately Mandarin, and heavy Mandarin accents) and lexicogrammar (grammatical and ungrammatical). The section below details how the research team manipulated the recordings. Table 1 provides more details regarding group assignment. For a more comprehensive description of the participants, see Miao (2020).
Group assignment.
Stimulus preparation: Recording listening tasks
Seven audio recordings were used in this study, including one baseline and six manipulated recordings. The speakers of the texts were two L1 English speakers from California (Speakers 1 and 2), one L1 speaker of Mandarin (Speaker 3), and one L1 speaker of Cantonese (Speaker 4). Communication with the research team suggested that the two L2 speakers demonstrated different degrees of accent in English. This was further supported by phonological coding and expert ratings as elaborated below.
To confirm objective differences between the moderate accent and heavy accent of the speakers, linguistic experts’ phonological coding and experienced teachers’ ratings of the two speakers’ grammatical recordings were triangulated. First, phonological analysis of the recordings was performed. This included researchers coding speakers’ segmental features, syllable structures, word stress features, vowel reductions, and pitch contour, following the coding scheme developed by Isaacs and Trofimovich (2012). Supplementary Material 5 provides the instructions for the coders along with examples. Two coders were used, one with L1 Mandarin (i.e., the author of the current manuscript), and the other, L1 English. They had both received a postgraduate degree in applied linguistics and TESOL and had experience teaching English phonetics at the time of the study. They first coded all recordings independently, and after this, a reconciliation meeting was held to address any coding discrepancies. The coding suggested that the heavily accented L1 Mandarin speaker demonstrated more phonological errors across coded categories than the moderately accented speaker, particularly for her syllable structure errors (see Table 2). Percentage exact agreement performed by two independent coders ranged from 91.18% to 100% across the coded categories.
Phonological coding of the grammatical recordings by Speakers 3 and 4.
An additional analysis involved nine experienced L1 Mandarin teachers of English (Mage = 33.56, SD = 5.92; eight female). They listened to the recordings and evaluated the three speakers who made the manipulated recordings regarding pronunciation on a 9-point scale (1 = very bad, 9 = very good; for similar methodological practice, see Gass & Varonis, 1984). Their ratings were rank ordered. The results suggested that all listeners unanimously gave higher ratings to Speaker 3 than to Speaker 4. Moreover, listeners unanimously gave higher ratings to Speaker 2 than to Speaker 3, except for one listener who provided equal ratings to both Speakers 2 and 3. Overall, this supported that Speakers 2, 3, and 4 had different levels of pronunciation.
The baseline recording was spontaneous speech (about 60 seconds) from one of the L1 English speakers of Northern American variety describing her morning routine (see Supplementary Material 1). The manipulated recordings were produced by having the other three speakers (with Northern American, moderate Mandarin, and heavy Mandarin accents) read aloud two scripts. The scripts were equally long (192 words) and described a party experience based on many prompt questions visually presented to the interviewees (e.g., what you did at the party, how did you feel about the party). This unpublished task was taken from a task in use at an IELTS teaching language institute in Shenzhen, China, designed to resemble the IELTS long-turn task (see IELTS, n.d.-b). Importantly, the two scripts were identical, except that one contained grammatical and lexical errors, and the other did not (see Supplementary Material 2).
To create the script with errors, the research team had six high-intermediate L1 Mandarin undergraduates (66.67% female, Mage = 21.33, SD = 0.52) complete the simulated IELTS task. Their utterances were transcribed, and based on it, the research team formulated a script incorporating the content and lexicogrammatical errors. To make the script as close to spoken discourse as possible, three steps were taken. First, most of the expressions, including errors, were taken directly from the interviewees. Second, if adding phrases for elaboration or coherence was necessary, colloquial expressions and less complex vocabulary was used (see Arndt & Woore, 2018). Third, the script was sent to two L1 English speakers for naturalness check. They indicated that the script appeared to be reasonably natural despite the presence of many errors. To create the error-free script, two L1 English speakers were asked to correct the ungrammatical script. They were instructed not to make the script “sound better” by replacing basic words with unnecessary, sophisticated vocabulary and solely corrected the errors. Moreover, they were asked to ensure that the ungrammatical and grammatical scripts were of equal length (measured in word tokens). The revised script was taken to another L1 English speaker for naturalness and grammar check. No further modifications were made.
Because speech sample duration has been shown to influence listener perceptual evaluations of L2 speech (Kormos & Dénes, 2004), the duration of reading the script was fixed at 75 seconds. This was done by having speakers adjust their speech rate when recording (see Kang et al., 2019).
Procedure
Listener participants were asked to listen to two recordings (depending on their assigned group) and evaluate the speaker’s comprehensibility on a 9-point Likert-type scale (i.e., how easy is the speaker to understand; 1 = very difficult, 9 = very easy). Listeners also evaluated the speakers’ grammar, vocabulary, discourse-organization, and fluency on four separate 9-point Likert-type scales via an online questionnaire (see Supplementary Material 3). Before the actual rating, they had a chance to listen to a test recording to adjust the volume of their device.
All listeners rated the same baseline recording before manipulated recordings. This is because the baseline recording is a planned speech from an L1 English speaker, which was predicted to receive relatively high ratings. This set a relatively high standard against which the manipulated recordings were rated, which minimized the likelihood of any ceiling effect on listeners’ ratings of the manipulated recordings. The research team acknowledged that the addition of some practice ratings and explanation of constructs would have been useful and benefited the validity of the obtained ratings.
After rating, listeners were asked to provide their background information such as gender, age, L1, and L2, as reported above. The online questionnaire took less than 10 minutes to complete, and all listeners had the opportunity to opt into a raffle for £100.
Data analysis
In the present study, the alpha level was set at .05; however, the research team acknowledges the arbitrariness of this threshold and provides descriptive statistics and measures of effect sizes, where possible, to provide additional insights into the findings. Data analysis was performed using SPSS and R studio (version 4.1.2; 1 November 2021), and the “lavann” package was used to compute path analysis (Rosseel, 2012).
The first research question, which explored the relationship between accent familiarity and a shared L1, was analyzed using descriptive statistics, including measures of central tendency and a box-and-whisker plot to visualize the data for ease of interpretation. Because a shared L1 was operationalized as a dichotomous variable and accent familiarity was operationalized as a continuous variable, a Mann–Whitney U test was performed for additional insights due to the non-normally distributed data.
The second research question investigated the relationship between accent familiarity, a shared L1, and comprehensibility solely of the L2 English speakers. Therefore, only participants in Groups 2 (Moderate Mandarin accent × ungrammatical), 3 (Heavy Mandarin accent × ungrammatical), 5 (Moderate Mandarin accent × grammatical), and 6 (Heavy Mandarin accent × grammatical) were analyzed (see Table 1).
The model proposed in Figure 1 hypothesized that multiple variables would predict comprehensibility. However, it also proposed that a shared L1 may predict accent familiarity. In light of the complex relationship between the predictor variables, a path analysis under the umbrella of structural equation modeling was deemed an appropriate analysis for this research question. While Figure 1 proposes a general model, Figure 2 provides a context-specific model relevant to the design of the present study. In this figure, a shared L1 is a dichotomous variable, which was hypothesized to be related to comprehensibility and accent familiarity. Accent familiarity, a continuous variable, was proposed to be related to comprehensibility. Other variables acting as covariates that needed to be controlled in the model included (a) listeners’ comprehensibility ratings of the baseline recording, (b) speaker pronunciation (i.e., moderate or heavy accents), and (c) speaker lexicogrammar (i.e., with or without errors). After the model was specified, preliminary analysis suggested that the model was not saturated, but overidentified, because not all the variables in the path diagram were linked with arrows (see Byrne, 2016; Kline, 2015).

Relationship between accent familiarity, a shared L1, and comprehensibility.
Results
The relationship between accent familiarity and a shared L1
This section reports descriptive and inferential statistics in response to the research questions. The first research question explored the relationship between listeners’ reported familiarity with Mandarin-accented English and a shared L1. For L1 Mandarin listeners (n = 103), their reported familiarity with Mandarin-accented English did not follow a normal distribution, with an observable negative skewness (−2.05, SE = 0.24) and a highly positive kurtosis (5.56, SE = 0.47). The data ranged from 1.85 to the maximum 9.00, with a mean of 7.95 (SD = 1.33) and a median of 8.18.
For listeners with other L1s (n = 199), their reported familiarity with Mandarin-accented English also did not show a reasonable normal distribution, with a marginally positive skewness (0.26, SE = 0.17) and an observable negative kurtosis (−1.13, SE = 0.34). The data ranged from the minimum 1.00 to the maximum 9.00, with a mean of 4.39 (SD = 2.41) and a median of 4.04. A box and whisker plot of the data is shown in Figure 3. Overall, the figure shows that L1 Mandarin listeners reported relatively high familiarity with Mandarin-accented English, which clustered at the high end of the scale. Comparatively, listeners with other L1s reported a considerably wider range of accent familiarity.

Listeners’ reported accent familiarity by their shared L1 status.
A Mann–Whitney U Test was performed to compare L2 English listeners’ reported accent familiarity with Mandarin-accented English for Mandarin L1 listeners compared with listeners from all other L1 backgrounds. The results suggested a significant difference between L2 English listeners’ reported accent familiarity in L1 Mandarin listeners (Median = 8.18) and listeners with other L1s (Median = 4.04), U = 18,280, z = −11.55, p < .001, r = .644, 95% confidence interval [CI] = [.571, .702]. Taken together, a substantial difference in reported accent familiarity was found in the two groups (i.e., with and without a shared L1).
The relationship between accent familiarity, a shared L1, and comprehensibility
This section reports findings from the path analysis. Several coefficients were computed to assess the model’s goodness of fit because the model was not saturated, but over-identified. All measures indicated that the model fit reasonably well, including the chi-square statistics χ2(9, n = 207) = 14.77, p = .097 > .05, comparative fit index = .977 > .95, Tucker–Lewis index = .962 > .95, root mean square error of approximation = .056 < .06, and standardized root mean square residual = .064 < .08; see Hooper et al., 2008; Hu & Bentler, 1999, for suggested benchmarks). A path diagram based on the findings is presented in Figure 4. The dotted arrows indicate nonsignificant relationships, whereas the solid-line arrows indicate significant relationships (i.e., when the 95% CIs around the path coefficient did not cross zero).

Relationship among accent familiarity, a shared L1, and comprehensibility: A path diagram.
Of the variables of interest in the left of the path diagram, the findings suggested a direct effect of a shared L1 on accent familiarity (β = 0.66, 95% CI = [0.56, 0.76]). Second, a direct effect of accent familiarity on comprehensibility was significant (β = 0.32, 95% CI = [0.18, 0.44]), meaning that a one-unit increase in listeners’ reported familiarity was associated with an increase in listeners’ reported comprehensibility by 0.32 standard deviations (SDs). Third, the indirect effect of a shared L1 on comprehensibility via accent familiarity was also significant (β = 0.21, 95% CI = [0.10, 0.33]), meaning that listeners’ reported comprehensibility would likely increase by 0.21 SDs with a shared L1, accounting for its associated change with accent familiarity. Finally, after controlling for accent familiarity, the direct effect of a shared L1 on comprehensibility was nonsignificant (β = 0.10, 95% CI = [−0.03, 0.23]).
Again, variables in the right of the path diagram including “baseline,” ‘pronunciation,’ and “lexicogrammar” were entered in the model to control for variance due to listener rating severity and speakers’ pronunciation and lexicogrammar. They will not be interpreted in detail in the present study. Listeners’ ratings of the baseline recording (β = 0.26, 95% CI = [0.15, 0.35]) and speakers’ pronunciation (β = 0.47, 95% CI = [0.36, 0.56]) exhibited a direct effect on comprehensibility, but not speakers’ lexicogrammar (β = 0.06, 95% CI = [−0.04, 0.16]). This suggests that listeners’ idiosyncrasies reflected on their rating of the baseline recording (e.g., their leniency), which made a significant contribution to comprehensibility, were controlled. Second, many speakers’ idiosyncrasies reflected on different experimental conditions (i.e., pronunciation and lexicogrammar features), which largely made a significant contribution to comprehensibility when other variables were controlled for. Table 3 includes information about listeners’ ratings of the baseline and the manipulated recordings across the four groups included in the analysis. Table 4 includes correlations among variables in the path analysis. Table 5 provides a comprehensive account of the findings pertaining to the path analysis.
Comprehensibility ratings of baseline and manipulated recordings across groups.
Note: Medians were included to describe central tendency in listeners’ ratings of the baseline recording due to non-normality of the distribution.
Moderate: the manipulated recording was recorded by the moderately-accented speaker; Heavy: the manipulated recording was recorded by the moderately accented speaker; ungrammatical: the manipulated recording contained grammatical errors; grammatical: the manipulated recording did not contain grammatical errors.
Correlation among variables in the path analysis.
Note: [] = 95% CI around the correlation coefficients based on 1000 bootstrap samples.
p < .01.
Path analysis results.
Note: L1: shared L1; Fam: accent familiarity; Comp: comprehensibility; Pron: pronunciation; Gram: lexicogrammar; SE: standard error; CI: confidence interval.
Discussion
The present study investigated the relationship among shared L1, accent familiarity, and comprehensibility. First, descriptive statistics suggested that given a shared L1 Mandarin, listeners’ reported familiarity with Mandarin-accented English was consistently high, meaning that their ratings were clustered at the higher end (see Figure 1). This makes intuitive sense because with a shared L1, listeners might have more exposure to such L1-influenced varieties in both educational settings (e.g., learning with instructors with a shared L1), and social settings (e.g., socializing in a community with a shared L1). Interestingly, it was found that about 25% of the data (i.e., the lower-bound whisker in Figure 3) did not cluster at the high end as did the rest. This could be explained by the fact that “Mandarin-accented English” encompasses many sub-varieties (i.e., influenced by speakers’ regional dialects; see Deterding, 2006), and speculatively, some listeners were not confident that they were familiar with all of them. However, this hypothesis needs empirical support, and if corroborated, perhaps the wording of the question used to measure accent familiarity could be modified, taking into account regional dialects, to gauge more contextualized familiarity ratings via self-report to enhance methodological and psychometric robustness.
Nonetheless, this finding provides empirical evidence to suggest that a shared L1 and accent familiarity indeed often co-occur, supporting the assumptions made by Kang et al. (2019) and Major et al. (2002), among others. By implication, this means that a shared L1, which is easy to operationalize in research terms, serves as a reasonably reliable indicator of listeners’ familiarity with their own accent varieties. Without further empirical evidence, it is possible to continue using a shared L1 as one way to index L1 listeners’ accent familiarity to explore fairness issues related to World Englishes.
Second, descriptive statistics suggested that the absence of a shared L1 did not necessarily mean lack of accent familiarity. Specifically, listeners without a shared L1 reported a wide range of accent familiarity, from highly familiar to not familiar at all. Third, Mann–Whitney U test suggested that L1 Mandarin listeners versus listeners with other L1s differed substantially in terms of their reported familiarity of Mandarin-accented English. This means that in the sample in the present study, listeners who reported a shared L1 also reported a substantially greater familiarity with Mandarin-accented English. However, this finding needs to be interpreted with caution in relation to the reported descriptive statistics. Importantly, the finding from inferential statistics did not suggest that listeners were not familiar with Mandarin-accented English if they did not speak L1 Mandarin. Rather, descriptive statistics showed a wide range of reported familiarity for non-L1 Mandarin listeners. In fact, some listeners reported being highly familiar with Mandarin-accented English, even without a shared L1 (see Figure 1). In other words, it is possible that there is more to accent familiarity than a shared L1. In speech perception studies, many scholars have assumed that accent familiarity could develop via short-term (e.g., Gass & Varonis, 1984) and long-term contact and exposure (e.g., Matsuura, 2007; Nejjari et al., 2012) or via listeners’ knowledge of speakers’ L1s (e.g., Winke et al., 2013), even when listeners did not share the same L1 with the speakers. That is, a shared L1 alone may not provide a full picture of accent familiarity. However, these beliefs and assumptions warrant empirical evidence of the psychological, cognitive, and socio-cultural underpinnings of accent familiarity development.
The argument that a shared L1 did not fully capture accent familiarity could be used to explain many inconsistent findings in the literature. Several scholars have used a shared L1 as the only focal listener variable to explore fairness issues related to World Englishes in the context of listening assessment (see, for example, Kang et al., 2019; Major et al., 2002). They demonstrated mixed findings in relation to the shared L1 advantage; that is, the shared L1 advantage was sometimes, but not always observed. If both a shared L1 and accent familiarity had been controlled for in these studies, the results might have shown a more consistent pattern.
Alternatively, while a shared L1 can index accent familiarity, and both can contribute to listening comprehension, a shared L1 may additionally tap into listeners’ attitudes, especially related to the status dynamics between speakers and listeners. Because listeners’ attitudes can impact their speech perception and evaluation (see, for example, Rubin, 1992; Schmidgall, 2013), the research team argues that both a shared L1 and accent familiarity should be included to explore fairness issues related to World Englishes in listening assessment.
Fourth, the path analysis suggested a direct effect of a shared L1 on accent familiarity. In other words, with a shared L1, listeners’ reported accent familiarity would likely increase by 0.66 SDs. This corroborates the above findings and perhaps disentangles accent familiarity and a shared L1, suggesting that they are indeed two different constructs. Fifth, a direct effect of accent familiarity on comprehensibility was observed. Specifically, a one-unit increase in listeners’ reported familiarity was associated with an increase in listeners’ reported comprehensibility by 0.31 SDs. This finding found support in precursor studies that reported accent familiarity to be associated with understanding of L2 speech (e.g., Gass & Varonis, 1984; Matsuura et al., 1999; Ockey & French, 2016); nonetheless, the different operationalizations of accent familiarity may confound this interpretation.
Finally, an indirect effect of a shared L1 on comprehensibility via accent familiarity was found; however, a direct effect of a shared L1 on comprehensibility, with accent familiarity controlled for, was not significant. Specifically, listeners’ reported comprehensibility would likely increase by 0.21 SDs with a shared L1, accounting for its associated change with accent familiarity. This seems to suggest that accent familiarity is a mediating variable in the relationship between a shared L1 and comprehensibility. In other words, it is possible that a shared L1 did not affect comprehensibility directly, but rather improved listeners’ accent familiarity, which, in turn, improved comprehensibility.
By way of implication, although previous research suggested that both a shared L1 and accent familiarity have the potential to improve understanding, it was unclear which construct could better capture the improved understanding. The findings from our study suggest two sources of improved comprehensibility: (a) from accent familiarity directly and (b) from the shared L1 via accent familiarity. This suggests that improved understanding should be understood in relation to both accent familiarity, and a shared L1. Many scholars have investigated the relationship between a shared L1 and test takers’ comprehension scores in high-stakes listening tests, seeking to evidence the shared L1 advantage (e.g., Harding, 2012; Kang et al., 2019; Major et al., 2002). This provides some evidence that incorporating different spoken varieties of English into listening tests could threaten test fairness. Building upon this claim, it is important to also acknowledge that besides a shared L1, test takers can be familiar with different spoken varieties via other channels. That is, it is possible that the construct of a shared L1 may not sufficiently capture fairness issues regarding test takers’ familiarity of a given spoken variety. This is perhaps why a mixed picture was observed in the literature. where test takers’ shared L1 did not always positively benefit their test scores. That is, the inclusion of both a shared L1 and accent familiarity might have improved the models used in previous studies and explained their inconsistent findings. Notably, the present study does not indicate whether accent familiarity or a shared L1 is more important than the other. Rather, both constructs have the potential to influence listeners’ understanding of L2 speech.
Limitations and future directions
Overall, the findings of this study should be interpreted with caution. First, accent familiarity was measured with one stand-alone Likert-type scale. However, it is possible that accent familiarity has multiple dimensions that cannot be captured with a single measure. As argued in the literature review, there have been many different measures of accent familiarity in the field, but scholars have yet to agree upon what the best measure should be in a given context (see, for example, Huang, 2013; Winke et al., 2013). The field would thus benefit from a validation study evaluating different accent familiarity measures, 1 which has direct implications for language assessment and speech perception in general. Specifically, the research team is undertaking a study to conceptualize accent familiarity, incorporating different measures of accent familiarity in the design. The research team found that among these measures were two main underlying factors—self-reported, impressionistic measures (e.g., via Likert-type scales; see Shin et al., 2021) and audio-prompted measures where listeners are exposed to different varieties before they assign ratings (see, for example, Bogorevich, 2018). Therefore, the research team acknowledges that the present study only captured a part of accent familiarity with its Likert-type scale measure.
Second, the present study used listeners’ perceptual judgment of comprehensibility, but not more “objective” measures of listening comprehension (e.g., comprehension questions). Thus, the generalizability of the findings applicable to language assessment is limited. Nonetheless, this study empirically disentangled accent familiarity from a shared L1, two sometimes conflated constructs. Thus, future studies need to replace the perceptual judgment of comprehensibility used in this study with more objective measures as used in authentic language assessment.
Third, the speech stimuli were elicited from only one task, which limits the generalizability of the findings. This is especially relevant because the shared L1 advantage may be more pronounced in transcription tasks than in listening comprehension tasks (see Dai & Roever, 2019). That is, speech elicitation tasks could potentially mediate the relationship between a shared L1, accent familiarity, and understanding. Future research could build on this finding and explore the differential effects that elicitation tasks have on the relationship between these three constructs.
Overall, a conceptual replication study incorporating all the above-mentioned ideas is warranted. Specifically, accent familiarity in the current path model (see Figure 4) could be operationalized using different measures. Understanding could be represented by listening comprehension questions (ideally ones that are authentic operational or “retired” test items), rather than comprehensibility as used in the present study. Moreover, the speaker population could be made more diverse, and “task” could also be included as a variable to see whether it mediates the relationship between understanding, a shared L1, and accent familiarity.
Conclusion
Several main findings were revealed in this study. First, accent familiarity and a shared L1 were related but independent constructs. While a shared L1 meant accent familiarity reasonably well, the absence of shared L2 did not. This suggested that a shared L1 alone only partially explained accent familiarity. Second, a complex relationship between a shared L1, accent familiarity, and comprehensibility was observed. Specifically, there seemed to be a direct effect of a shared L1 on accent familiarity, a direct effect of accent familiarity on comprehensibility, and an indirect effect of a shared L1 on comprehensibility via accent familiarity. However, a shared L1 did not contribute to comprehensibility with accent familiarity controlled.
In terms of implications, some assessment scholars have voiced the need to incorporate accented varieties of English into listening proficiency tests to reflect the real-life use of English where L2 accents were prevalent (e.g., Harding & McNamara, 2018; Jenkins & Leung, 2019). However, there are many complications that resulted in the lack of changes, including (a) a reconceptualization of the target language use domain to pinpoint what spoken varieties to be included, (b) determination of how many varieties to include, and (c) to ensure that the inclusion of different varieties do not threaten fairness. That is, both practical issues and test fairness issues come into play in the decision-making process.
Fairness has been investigated in previous research on whether test-takers’ shared L1 could positively or negatively benefit their listening comprehension scores in accented listening materials (see Kang et al., 2019; Major et al., 2002). The present study methodologically problematized previous studies by arguing that a shared L1 (a) may not be sufficient to capture accent familiarity, (b) did not contribute to comprehensibility when accent familiarity was controlled for, and (c) contributed to comprehensibility indirectly via accent familiarity.
Thus, future studies should consider both a shared L1 and accent familiarity in exploring fairness issues related to World Englishes in listening assessment. The consideration of both constructs could allow us to understand fairness issues more comprehensively. After this, researchers can proceed to investigate how to maintain fairness in listening assessments incorporating different spoken varieties (e.g., by applying screening criteria on speaker selection; see Kang et al., 2019). The research team believes that domain-relevant accent can be part of the target construct in listening assessment, and that it is possible for listeners without shared L1s to develop familiarity with different spoken varieties, thus reducing threats to fairness.
Following the inclusion of different spoken varieties of English in listening tests, the research team hopes that (a) the ecological validity of listening tests will be improved due to improved construct representativeness, meaning that test takers’ test performance will bear more resemblance to real-life contexts, which include different accent varieties; (b) this will lead to positive washback effects where different spoken varieties of English are included in teaching, learning, and test preparation; and (c) the field of language teaching can move away from the standard language ideology and empower multilingual users of English regardless of their L1(s).
Supplemental Material
sj-docx-1-ltj-10.1177_02655322231156105 – Supplemental material for The relationship among accent familiarity, shared L1, and comprehensibility: A path analysis perspective
Supplemental material, sj-docx-1-ltj-10.1177_02655322231156105 for The relationship among accent familiarity, shared L1, and comprehensibility: A path analysis perspective by Yongzhi Miao in Language Testing
Footnotes
Acknowledgements
I would like to thank journal Editors Luke Harding, Talia Isaacs, and four anonymous reviewers for their insights and support throughout the editorial process. I would also like to thank Okim Kang for recommending path analysis, Ekaterina Sudina and Jesse Egbert for very carefully reviewing earlier drafts of the manuscript, and Heath Rose for being an amazing supervisor of my MSc dissertation at the University of Oxford on which this manuscript is based. All remaining errors are mine.
Author’s note
The current manuscript is based on the data collected in my master’s dissertation at the University of Oxford. The research questions are new and were not published nor under review elsewhere.
Declaration of conflicting interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
Supplemental material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
