Abstract
This study investigated foreign-accentedness, comprehensibility, and intelligibility in second language (L2) Arabic speech. More specifically, it was designed to 1) measure how foreign-accented, comprehensible, and intelligible L2 Arabic speech was, 2) explain the relationships among those three global aspects, and 3) determine the extent to which foreign-accentedness and comprehensibility – individually or combined – explain or predict intelligibility. Thirty adult L2 Arabic speakers were audio-recorded while having an unofficial oral proficiency interview (OPI). A subset of three speech samples were extracted from the entire OPI for each speaker. Next, 10 adult native speakers of Arabic listened to and transcribed all speech samples in Arabic as a measure of intelligibility; in addition, they rated the samples for foreign-accentedness and comprehensibility on a 9-point scale. Results showed that L2 Arabic speech was rated more positively for comprehensibility than for foreign-accentedness, on average; in addition, most of it was less than highly intelligible. Also, results showed statistically significant relationships among the three global aspects under investigation, although they varied in strength: foreign-accentedness and comprehensibility showed the strongest relationship, followed with comprehensibility and intelligibility; foreign-accentedness and intelligibility was found to have the weakest relationship. Furthermore, foreign-accentedness and comprehensibility demonstrated a statistically significant explanatory or predictive power in intelligibility. However, both the unique contribution and added power of foreign-accentedness were not statistically significant, deeming a model with only comprehensibility being the best fit for the obtained assessment data of L2 Arabic speech. Consistent with existing evidence, the results support the partial independence of foreign-accentedness, comprehensibility, and intelligibility as global aspects of L2 speech while also indicating L2-specific patterns. After discussing the results and implications, directions for future research are provided.
Keywords
I Introduction
Foreign-accentedness is a salient feature of second and foreign language (L2) speech. Among others, this salience can be due to first language (L1) influence, especially among adult L2 learners (Flege et al., 1995). Therefore, comprehensible and/or intelligible, rather than native-like, speech for such learners seems to be a realistic criterion for assessment and, in turn, a practical goal for instruction (Isaacs, 2018; Levis, 2005, 2018). This conceptualization is supported by the existing research into L2 speech, especially the growing line of research into the influence foreign accent on communicative effectiveness of L2 speech (Munro & Derwing, 2015).
Adopting listener-based speech assessment framework, studies in this line of research have drawn upon foreign-accentedness (degree of foreign accent), comprehensibility (degree of easiness or difficulty of understanding), and intelligibility (degree of actual understanding) as central global aspects of L2 speech. Issues under investigation in those studies have developed over the past three decades, starting with the relationships those three global aspects have with one another (e.g. Derwing & Munro, 1997; Munro & Derwing, 1995a, 1995b). Subsequent studies have examined linguistic features associated with those aspects, mainly comprehensibility and foreign-accentedness (e.g. Isaacs & Trofimovich, 2012; Trofimovich & Isaacs, 2012), speaker-related and listener-related influence (e.g. Crowther, Trofimovich, Saito, & Isaacs, 2015; Saito et al., 2017), and research design (e.g. Crowther, 2020; Crowther, Trofimovich, Isaacs, & Saito, 2015; Kang et al., 2018; O’Brien, 2016).
As far as the current study is concerned, the existing evidence indicates that foreign-accentedness, comprehensibility, and intelligibility are partially independent aspects of L2 speech. In a groundbreaking study, Munro and Derwing (1995a) were the first to demonstrate evidence for the partial independence of those three aspects in L2 English speech. A key finding in that study was that strongly foreign-accented speech could be still understood. More recent studies have showed similar evidence. Overall, evidence related to the degree of association among foreign-accentednesss, comprehensibility, and intelligibility has been influential in the ongoing paradigm shift, from the nativeness principle toward the intelligibility principle (Isaacs, 2018; Levis, 2005), in L2 speech theory and research with important implications for both instruction and assessment.
Still, there is more research to be done to further advance our understanding of the nature of L2 speech. This includes expanding the scope of investigation to new L2s and/or new contexts, adopting more spontaneous tasks for speech elicitation, and using more advanced approaches for data analysis. To date, the investigation of all these three constructs has been limited in terms of the target L2 and context. To illustrate, a few studies have investigated foreign-accentednesss, comprehensibility, and intelligibility: Munro and Derwing, 1995a; Derwing and Munro, 1997; Jułkowska and Cebrian, 2015. What is more, all these studies were conducted on L2 English speech in a non-classroom context. One exception is that of Huensch and Nagle (2021) which examined the L2 Spanish speech of adult learners in a classroom context. Another limitation is that most of studies have adopted picture-narration as a speech elicitation task, which is not as spontaneous as other tasks, such as oral proficiency interview, undermining the ecological validity of research findings. Taken together, addressing these limitations have the potential of advancing the current understanding of listener-based assessment of L2 speech.
To that end, this study investigated foreign-accentedness, comprehensibility, and intelligibility in spontaneous L2 Arabic speech produced by L1 English speakers who were learners of Arabic as a foreign language in the United States. Oral proficiency interview was used for speech elicitation. Arabic is an underrepresented target language in L2 research, in general, and in this line of research, particular. However, the number of people who speak or learn Arabic as an L2 is ever increasing not only in the United States but also around the world. In addition, the L1-English/L2-Arabic pair investigated here represent an interesting case of study. This is due to the fact that these two languages are typologically different in many respects including sound system. Yet, to the best of our knowledge, no studies have explored listener-based assessment of global aspects in L2 Arabic speech.
II Literature review
1 L2 speech assessment framework
According to Saito and Plonsky (2019), an overall L2 speech assessment research framework can be conceptualized with respect to three sets of parameters: (1) the constructs under investigation (global vs. specific aspects of pronunciation), (2) the type of speech elicited (controlled vs. spontaneous), and (3) the scoring method (listeners vs. acoustic analyses). Drawing on this methodological framework, the specific research design adopted in the current study is delineated below.
a Key constructs
Given the complexities of L2 pronunciation proficiency as well as the oral proficiency more broadly, empirical research in this area has investigated a wide range of dimensions including global and specific aspects. While studies on global aspects tap into overall pronunciation phenomena such as how foreign-accented, comprehensible, or intelligible L2 speech is (e.g. Hayes-Harb & Watzinger-Tharp, 2012; Derwing & Munro, 1997; Munro & Derwing, 1995a, 1995b; Huensch & Nagle, 2021), those on specific aspects tap into more fine-grained linguistic properties of L2 speech such as segmental and suprasegmental features (e.g. González-Bueno, 1997; Munro & Derwing, 2006; Munro et al., 2015). From a proficiency-oriented perspective, global aspects appear to be more ecologically valid constructs to investigate as they are related to real-life communication and require no linguistic training on the part of the listeners to evaluate (Munro & Derwing, 2015; Saito & Plonsky, 2019). Therefore, the current study was designed to investigate three main global aspects of L2 speech: foreign-accentedness, comprehensibility, and intelligibility.
While widely investigated, there is a lack of consensus on how those three constructs are defined and operationalized (for a comprehensive review, see Thomson, 2017). Responding to calls in literature for greater consensus (Munro & Derwing, 2015), in this study, all three constructs under investigation are defined according to Munro and Derwing (2015, p.14) as follows:
foreign-accentedness is the perceived degree of differences in pronunciation as compared with a local variety;
comprehensibility is the perceived degree of ease or difficulty experienced by the listener in understanding speech; and
intelligibility is the extent to which listeners’ perceptions match speakers’ intentions (actual understanding).
b Task type
This parameter refers to the L2 speech elicitation tasks employed in empirical investigations. In general, the elicited speech depends on target content and means of elicitation (Munro & Derwing, 2019). Target content can be individual words, sentences, monologues or interactions. This target content can be elicited through reading, picture narration, speaking prompt, or unrehearsed interaction. As such, in a rough distinction, two major types can be identified: controlled and spontaneous elicitation tasks.
Controlled tasks include those in which speakers are asked to read aloud a list of words, sentences, or paragraphs. As the name suggests, the elicited speech content is controlled such that it contains specific pronunciation features of interest. Unlike controlled tasks, in spontaneous tasks speakers are asked to speak freely in response to a prompt. However, the degree of freedom can vary to some extent. For example, speakers can be given a picture to describe, a picture-story frame to narrate, asked to tell a story on their own, or interact with another speaker. Also, within this type of tasks there can be a distinction between monologic (e.g. picture description) and interactive (e.g. oral interview) spontaneous tasks.
Crucially, speech performance varies depending on task type with it being more native-like in the case of controlled tasks giving the less demanding nature of them compared to spontaneous tasks (Saito & Plonsky, 2019). While controlled tasks are more convenient, spontaneous tasks are more ecologically valid and thus make research findings applicable or generalizable to the target population or context (Munro & Derwing, 2019) as they are representative of real-life communication, unlike controlled tasks (Levis & Barriuso, 2011). In this regard, interactive spontaneous (as in oral interview) appear to be the most ecologically valid.
c Scoring methods
Two major methods of scoring L2 speech can be identified: human-based and computer-based. The former involves having individuals perform audio-based assessment tasks (e.g. rating and transcription). In addition, within the human-based approach, there is a distinction between listeners and raters. On one hand, listeners are those who provide holistic assessment of speech samples with zero to little linguistic knowledge. On the other hand, raters are experts who have had extensive linguistic training and/or certification training. With respect to global aspects of L2 speech, the focus of the current study, Munro and Derwing (2015) noted that they are to be measured based on listener-based approach as it provides insights into L2 use representative of real-life communication, unlike computer-based approach. In the computer-based approach, speech samples are submitted to acoustic analysis of quantifiable features, temporal and spectral, such as speech rate, voice onset time, format frequencies, pitch, and duration, using a speech analysis software such as Praat (Boersma & Weenink, 2019). This second approach best fits the measurement of specific aspects of L2 speech. For this study, the methodological framework adopted for assessing L2 speech has the three global aspects of foreign-accentedness, comprehensibility, and intelligibility as constructs, spontaneous task, oral proficiency interview (OPI), as speech elicitation task, and listener-based ratings and transcription of speech samples as scoring method.
2 Relationship among L2 speech global aspects
Munro and Derwing (1995a) was the first to investigate the communicative effectiveness of L2 as a function of the three global aspects for foreign-accentedness, comprehensibility, and intelligibility. Results showed a significant relationship between foreign-accentedness ratings and intelligibility scores for only five out of 18 listeners, with the strength of the correlations being quite weak, r = .37 – .48. For foreign-accentedness and comprehensibility, significant correlations for 15 out of 18 listeners were found, with the strength of correlations ranging from weak to strong, r = .41 – .82. The study concluded that comprehensibility was more related to intelligibility than foreign-accentedness. In addition, the study found that even some utterances were rated as strongly foreign-accented, they were still transcribed with perfect accuracy (i.e. fully intelligible). Overall, findings from this study established evidence for the partial independence of foreign-accentedness, comprehensibility, and intelligibility.
These results were replicated in Jułkowska and Cebrian (2015). In that study, comprehensibility and intelligibility were significantly related for 15 out of 18 listeners, r = .67 – .82, while intelligibility was related to foreign-accentedness significantly for 5 listeners only, r = .09 – .68. Most recently, Huensch and Nagle (2021) found that the relationship between comprehensibility and intelligibility was significant for all listeners while that between intelligibility and foreign-accentedness was not for any listener. In addition, though the relationship between foreign-accentedness and comprehensibility in that study was found to be significant, the strength of the relationship varied by listener.
Overall, findings from previous studies agreed on the partial independence of foreign-accentedness, comprehensibility, and intelligibility. However, these findings also showed L2-specific patterns. For example, while the most L2 Spanish was highly intelligible in Huensch and Nagle (2021) corresponding to L2 English in Munro and Derwing (1995a), the same trend was not found for comprehensibility. In Huensch and Nagle (2021) L2 Spanish was moderately but not highly comprehensible as in Munro and Derwing (1995a, 2020).
III The current study
To address those issues and gain a better understanding of L2 speech global aspects, this study is designed to expand the scope of investigation in L2 pronunciation research by examining foreign-accentedness, comprehensibility, and intelligibility in the spontaneous L2 Arabic speech of adult American learners of Arabic as a foreign language. The goal of the study is threefold:
first, as it is the first to examine L2 Arabic speech, the study is interested in finding out how much foreign-accented, comprehensible and intelligible L2 Arabic speech is based on native Arabic-speaking listeners’ perceptions;
second, the study is concerned with how foreign-accentedness, comprehensibility, and intelligibility are related to one another; and
third, it is interested in the extent to which foreign-accentedness and comprehensibility ratings, isolated and combined, could predict intelligibility.
The following research questions guide this study:
Research question 1: How foreign-accented, comprehensible, and intelligible spontaneous L2 Arabic speech is as rated and transcribed by L1 Arabic listeners?
Research question 2: To what extent are foreign-accentedness, comprehensibility and intelligibility related to one another?
Research question 3: To what extent can foreign-accentedness and comprehensibility, individually and combined, predict intelligibility?
IV Methods
1 Participants
Typical of studies in this line of research, this study included two groups of participants: speakers and listeners (Munro & Derwing, 2015). All participants were located in the United States at the time of data collection. Only listeners received compensation for participation in the study.
a Speakers
Thirty-five participants were recruited to provide speech samples for this study. Thirty of them were L1 English / L2 Arabic speakers (19 males, 11 females) who were graduate or undergraduate students from second- and third-year Arabic classes at a large public university located in the midwestern United States. These L2 Arabic speakers were between the ages of 19 and 28 years (
b Listeners
Ten L1 Arabic speakers (9 males, 1 female) were recruited as listeners. All of them were born and raised in an Arabic-speaking country, Egypt. Their mean age was 35 years (SD = 11.30, range: 19–57). The rationale for recruiting listeners from Egypt is that it represented the place of origin of the L2 Arabic learners’ instructor and, therefore, the Arabic dialect those learners were exposed to at the time of data collection. All listeners reported zero to little familiarity with L2 Arabic speech. While none of these listeners reported having teaching experiences, four of them reported having language-related training. All listeners reported having no hearing impairments.
2 Materials
a Language background questionnaire
Both the speakers and listeners completed a language background questionnaire, based on Huensch and Nagle (2021). The questionnaire elicited information relevant to describing the sample population including biographical information, language learning and use, and language-related training and language teaching experience. In addition, the listeners responded to questions related to familiarity with L2 Arabic speech.
b Speaking task
An oral proficiency interview was used to elicit naturalistic, spontaneous speech, following the American Council on the Teaching of Foreign Languages (ACTFL) Oral Proficiency Interview (OPI). All interviews were conducted individually by the author who is an ACTFL-certified OPI tester. The entire interview was in Arabic except for the introduction, which was in English, per the OPI protocol. Also, to simulate the official OPI, the interviews were administered online, on Zoom, which also allowed for high-quality recording using the recording function available on that platform.
c Oral proficiency interview
Based on the American Council on the Teaching of Foreign Languages (ACTFL, 2012) guidelines, the Oral Proficiency Interview (OPI; https://www.languagetesting.com/oral-proficiency-interview-opi) was developed as a standardized assessment to measure spoken language proficiency. In this study, it was additionally used for speech elicitation. Based on their interviews, while the L1 Arabic speakers were rated at the Superior Level, the L2 Arabic speakers were rated at the Intermediate Level of proficiency for speaking, according to ACTFL (2012) proficiency guidelines. It should be noted the proficiency level rating the L2 Arabic speakers in this study received is typical of learners of Arabic as a foreign language at that stage of learning, second- and third-year of study. As such, those L2 speakers s in this study is representative of the L2 Arabic learner population with the same characteristics.
3 Procedures
a Speech elicitation
The speaker participants were scheduled for individual one-hour long sessions on Zoom. All were instructed to be in a quiet room and have their headphones on during the entire session. First, they completed the language background questionnaire hosted on Qualtrics, which took between 10 and 15 minutes to complete. Next, they were introduced the oral proficiency in English and the interviewer answered any questions they had. Following OPI structure and protocol, each speaker was presented with a series of tasks. The topics of those tasks varied depending on the information gathered during a warm-up. Also, the tasks were adaptive such that speakers did not respond to the same exact speaking prompts. However, they were still asked questions that elicited similar types of language functions and content (i.e. present narration, description, role-play). The interviews ended with a wind-down and lasted between 20 and 35 minutes.
Three speech samples per speaker were extracted from each interview for the purpose of this study. These samples represented different language functions and content: present narration, past narration, and role-play, one sample each per speaker. An extraction protocol was followed such as that the speech samples contained full utterances, excluding any initial hesitations and/or dysfluencies. As such, the final set of speech samples included 105 speech samples with an average length of 13 seconds and an average word count of 15 words. Both length and word count of the samples were in line with previous L2 speech research using audio-based assessment (Derwing et al., 2008), while also being long enough to allow for reliable judgments (Munro & Derwing, 2019). In addition, all samples were extracted and normalized for peak amplitude, set to 0.99, on Audacity® software (Audacity Team, 2020).
b Speech assessment
There were three main speech assessment tasks completed by the listeners in this study: foreign-accentedness rating, comprehensibility rating, and transcription. Following Munro and Derwing (1995), among others, a 9-point Likert scale with endpoint descriptors was adopted for the two rating tasks: foreign-accentedness (1 = not foreign-accented at all, 9 = very strongly foreign-accented) and comprehensibility (1 = very easy to understand, 9 = very difficult to understand). Likewise, listeners were asked to transcribe speech samples in standard Arabic orthography as a measure of intelligibility, as commonly used in previous studies (e.g. Huensch & Nagle, 2021; Derwing and Munro, 1997; Munro & Derwing, 1995a). All assessment tasks were completed on Qualtrics.
A custom-designed interface was created on Qualtrics to present speech samples and to collect their ratings and transcriptions. The interface included (1) consent form, (2) overview of the three assessment tasks with instructions for completing each task, (3) three speech samples for practice, and (4) the main speech samples. All 105 speech samples (30 L2 Arabic speaker × 3 samples; 5 L1 Arabic speakers × 3 samples) were arranged in one block so that all listeners rate and transcribe all samples; however, each sample with three associated tasks was presented in a different page, in a randomized order. Also, speech samples were arranged to be played only once, with no stop and pause control enabled. In addition, listeners were not able to backtrack to revise their ratings and/or transcriptions of previous samples. Furthermore, listeners had between a maximum of 30 seconds to rate and a maximum of 60 second to transcribe each speech sample.
c Intelligibility scoring
In preparation for data analysis, each sample was also transcribed manually by the author in standard Arabic orthography. All transcriptions were checked by a research assistant for reliability. Next, those transcriptions were compared to those completed by the listeners to determine an intelligibility score. Following previous studies (e.g. Huensch & Nagle, 2021; Munro & Derwing, 1995a, Derwing & Munro, 1997), an intelligibility score for each sample was computed as the percentage of the number of words correctly transcribed, by the listener, divided by the total number of words of that sample, as transcribed and coded by the researcher. Foreign-accentedness and comprehensibility scores were more straightforward as they were directly retrieved from Qualtrics as given by each listener per each speech sample. A research assistant checked a subset of 45 speech sample transcriptions with their associated intelligibility scores for agreement which reached 98%. In case of disagreement, the author’s scores were adopted.
V Analysis
In this study, three research questions were addressed by analysing foreign-accentedness, comprehensibility, and intelligibility of spontaneous L2 Arabic speech as rated and transcribed by L1 Arabic listeners. This involved preliminary analysis and main analysis. To explain, we confirmed the reliability of foreign-accentedness and comprehensibility through intraclass correlation coefficients (ICC) analysis. More specifically, a two-way, consistency, average-measure ICC model was adopted. For foreign-accentedness, ICC was .98 (95% confidence interval (CI): [.97,.99]) and .96 (95% CI: [.93,.97]) for comprehensibility, both exceeding the .70 threshold in L2 research (Larson-Hall & Plonsky, 2015) and corresponding to those of previous studies, especially Huensch and Nagle (2021) for L2 Spanish (for a review, see Saito, 2021). As such, foreign-accentedness and comprehensibility ratings and intelligibility scores were averaged across all listeners to compute a single score for each speaker on each of these two measures. The L1 Arabic speakers’ data were excluded as a common practice (Munro & Derwing, 2015) from the following main analysis.
For the main analysis, parametric statistic procedures were used, specifically Pearson correlation and linear regression model (LRM). Pearson correlations were used to examine the bivariate relationships among foreign-accentedness (1–9, low to high), comprehensibility (1–9, high to low), and intelligibility (0–100, low to high) to answer research question 2. To answer research question 3, we used linear multiple regression analyses to explore how foreign-accentedness and comprehensibility, individually and combined, contributed to or predicted intelligibility. With the understanding that intelligibility is the goal of communication (Munro & Derwing, 2015; Levis, 2005, 2018, 2020), total of three models were constructed with intelligibility being the outcome variable. The first model included both foreign-accentedness and comprehensibility as predictor variables to examine the combined contribution and predictive power of foreign-accentedness and comprehensibility as related to intelligibility. In the second model, comprehensibility was entered first followed with foreign-accentedness, also as predictors of intelligibility, to investigate the contribution and predictive power of foreign-accentedness, after controlling for comprehensibility. Conversely, in the third model, foreign-accentedness was entered first and comprehensibility second to examine the unique contribution and predictive power of comprehensibility, after controlling for foreign-accentedness.
Prior to conduct correlation analyses and model-fitting, the aggregated scores on each of the three variables were verified to be normally distributed. To illustrate, normality tests, Kolmogorov–Smirnov and Shapiro–Wilk, and visual inspection, of histograms and boxplots, indicated normal distribution for all three variables of interest. Furthermore, skewness and kurtosis ratios showed the same. For these two ratios, values below −1.96 and above 1.96 are considered indicators of non-normal distribution (Field, 2013). For foreign-accentedness, Kolmogorov–Smirnov and Shapiro–Wilk tests of normality (p = .200 and p = .205, respectively) and skewness (
In addition to significance level, the results, reported in the following section, included 95% confidence intervals (CIs) and effect sizes, in line with L2 research reporting guidelines (Larson-Hall & Plonsky, 2015). Effect sizes were interpreted following Plonsky and Oswald’s (2014) benchmarks for small (.25), medium (.40) and large (.60) effect sizes. All analyses were performed using the Statistical Package for the Social Sciences (SPSS) version 27.0 (IBM Corp, Armonk, New York, USA) with alpha level set at less than 0.05.
VI Results
Regarding research question 1, as shown in Figure 1, the foreign-accentedness scores were moderately skewed to the right, the comprehensibility scores demonstrated a rather even distribution (around the moderately easy/difficult to understand point of the scale), and the intelligibility scores were heavily left-skewed. In other words, the L2 Arabic speakers were rated toward the higher, harsher, end of the scale (i.e. 9 = very strongly foreign-accented). On the other hand, they were rated almost evenly around the mid-point of the scale for comprehensibility. More importantly, they were still highly intelligible as their scores fell mostly in the 67–95 range of score, out of 100. Descriptive statistics, Table 1, supported these distributions. Therefore, overall, L2 Arabic speech was strongly foreign-accented, moderately comprehensible, and highly intelligible. These results are line with those in previous studies (e.g. Munro & Derwing, 1995a, 2020; Huensch & Nagle, 2021), supporting that even strongly foreign-accented speech can be still highly intelligible.
Descriptive statistics for all measures (N = 30).

Distribution of foreign-accentedness, comprehensibility, and intelligibility scores.
As to research question 2, the Pearson’s correlation analysis showed significant relationships among all three global aspects in L2 Arabic speech. As displayed in Table 2, results demonstrated that the relationship between foreign-accentedness (1–9, low to high) and comprehensibility (1–9, high to low) was positive and strong, r(28) = .75, r2 = .56, p < .001, indicating the more foreign-accented L2 Arabic speech was the more difficult to understand it was, as perceived by our listeners. The relationship between comprehensibility and intelligibility (0–100, low to high) was negative and strong, r(28) = −.66, r2 = .44, p < .001, indicating the more difficult to understand the speech was perceived by the listeners the less correctly it was transcribed by these same listeners (i.e. less intelligible). Finally, the relationship between intelligibility and foreign-accentedness negative and weak, r(28) = −.36, r2 = .13, p < .05, indicating that speech that was perceived to be more foreign-accented tended to less intelligible.
Pearson’s correlation (r [CI]) among the three measures.
Note. *p < .05. **p < .001.
Therefore, these initial correlation analyses showed foreign-accentedness, comprehensibility, and intelligibility to be related yet distinct, similar to findings in previous studies (e.g. Huensch & Nagle, 2021; Munro & Derwing, 1995a). The relationship between foreign-accentedness and comprehensibility was the strongest, followed with that between comprehensibility and intelligibility, while that between intelligibility and foreign-accentedness was the weakest. Following the Plonsky and Oswald’s (2014) benchmarks for r effect sizes, these correlations can be considered in the medium to large range.
Finally, to explore the combined and individual influence of foreign-accentedness and comprehensibility on intelligibility, a series of standard multiple regression models was conducted with intelligibility as a dependent variable and foreign-accentedness and comprehensibility as predictor variables. Additionally, we checked for multicollinearity using test statistics and no concerns were found with (VIF = 2.26, Tolerance = .44) or outliers (residual statistics between −3.0 and 3.0, Cook’s distance < 1.0).
First, the assessment data was fitted to a multiple regression model, model 1, with both foreign-accentedness and comprehensibility as predictors of intelligibility. The results showed that the model was significant with the foreign-accentedness and comprehensibility combined predicting 47% of the variance (F(2.27) = 12.19, R = .69, R2 = .47, p < .001), a medium effect size (Plonsky & Ghanbar, 2018). However, only comprehensibility (β = –.88, t = –4.20, p < .001) significantly contributed to the model, as shown in Table 3.
Multiple regression model 1 correlation coefficients.
Next, we examined the individual contribution of the two predictors variables, foreign-accentedness and comprehensibility, more closely by conducting two hierarchal regression models, which allowed for the estimation of the added predictive power, R2 change, of each of those two predictors while controlling for the other. As presented in Tables 4 and 5, foreign-accentedness only added 3.8% to predictive power of the overall model which is not statistically significant (R2 Change = .038, p = .174). On the hand, comprehensibility significantly added to the predictive power of the model explaining alone 34% of the variance in intelligibility (R2 Change = .34,
Hierarchical regression model summary (controlling for comprehensibility).
Hierarchical regression model summary (controlling for foreign-accentedness).
Considering the results from the hierarchal regression models above, a simple model with only comprehensibility scores (M = 4.89, SD = .76) as a predictor variable of intelligibility scores (M = 84.27, SD = 6.68) was deemed to the best fit for the L2 Arabic speech assessment data in this study, R = .66, R2 = .44, p < .001, explaining 44% of the variance in intelligibility with a medium effect size (for the model summary, see Table 6). Based on this model correlation coefficients, shown in Table 7, one point increase in comprehensibility ratings (1 = very easy to understand; 9 = very difficult to understand) is associated with 5.81 points decrease in intelligibility score, as evidenced by the observed unstandardized correlation coefficient, B. Alternatively, it can be said that for each one SD (.76) increase in comprehensibility score, intelligibility score is predicted to decrease by .66 SD, based on the standardized correlation coefficient, β.
Final regression model summary.
Multiple regression model 1 correlation coefficients.
VII Discussion
The purpose of this study was to examine the foreign-accentedness, comprehensibility, and intelligibility of L2 Arabic speech. More specifically, the study investigated the degree of association among those three global aspects of L2 speech. Adopting oral proficiency interview (OPI) for speech elicitation as well as proficiency level assessment, the aim was to cross-validate existing evidence of their partial independence, mainly with English as the target L2 in the English as a second language (ESL) context, to see if that would generalize to a new L2 and context (i.e. L2 Arabic in the Arabic as foreign language (AFL) context). In addition, another goal of this study was to explore the unique contribution of foreign-accentedness and comprehensibility, individually or combined, as predictors of intelligibility.
Building on existing literature (Huensch & Nagle, 2021; Munro & Derwing, 1995a, 2020), foreign-accentedness was defined as the degree of difference in pronunciation between the speaker and the native-speaker norm; it was rated on a 9-point scale with only endpoint descriptors (1 = not foreign-accented at all, 9 = very strongly foreign-accented); comprehensibility, defined as the degree to which the speaker was easy or difficult to understand, was also rated on a 9-point scale with only endpoint descriptors (1 = very easy to understand, 9 = very difficult to understand); and intelligibility was defined as the amount of L2 speech that was actually understood by listeners as intended; it was operationalized as the percentage of correctly transcribed words divided by the total number of words in the speech sample, yielding a score out of 100.
Apart from the individual measures descriptive findings from this study, in general, cross-measures comparisons appear to support the findings from previous studies regarding the partial independence of these three L2 speech dimensions. To illustrate, when considering the subset of speech samples that received a score of 90% or higher, it is evident that highly intelligible L2 Arabic speech does not necessarily mean the speech is not foreign-accented at all. In fact, the average foreign-accentedness rating of these speech samples was 5.7, more toward the very strongly foreign-accented end of scale. In addition, those same speech samples were rated 4.12, on average, on the comprehensibility scale, more toward the easy-to-understand end of the scale. Taken together, these findings suggest that while the L2 Arabic learners in this study were almost perfectly to perfectly intelligible, they were not perceived as much comprehensible, not to mention as much native-like in their pronunciation.
1 Foreign-accentedness, comprehensibility, and intelligibility in L2 Arabic speech
Regarding the strength of relationships among all three L2 speech dimension, simple LRM results indicated statistically significant linear relationships among foreign-accentedness, comprehensibility, and intelligibility. However, those relationships were different in both direction, strength and, in turn, effect size. Foreign-accentedness and comprehensibility showed the strongest relationship, R = 0.75, followed with the comprehensibility and intelligibility relationship, R = 0.66, and the least strong relationship was between foreign-accentedness and intelligibility, R = 0.36. Looking at the correlation coefficients, it was demonstrated that there was a strong positive relationship between foreign-accentedness (high to low) ratings and comprehensibility (high to low) ratings such that more foreign-accented speech is also more difficult to understand and vice versa. Based on the obtained value, (R = 0.75), 56% of the variance in comprehensibility is accounted for by foreign-accentedness, and the reverse is true, with a large effect size. Conversely, there was a strong but negative relationship between comprehensibility (high to low) ratings and intelligibility (low to high) scores such that speech that is more difficult to understand is also less intelligible, and vice versa. Again, based on the obtained point estimate this relationship, (R = 0.66), comprehensibility explains 44% of the variance in intelligibility, and vice versa, with a medium effect size. Likewise, there was negative yet weak relationship between foreign-accentedness ratings and intelligibility such that speech that is more foreign-accented is also less intelligible, and vice versa. Based on the obtained value of this relationship, (R = 0.36), 13% of the variance in intelligibility is explained by foreign-accentedness, and vice versa, which has a small effect size. Overall, the findings provide further evidence for the partial independence among foreign-accentedness, comprehensibility, and intelligibility.
Despite the overall convergence with existing findings, findings from this study showed unique patterns in terms of the distribution of as well as the degree of overlap among the three global aspects under investigation. In the current study, foreign-accentedness ratings were rather skewed toward the moderately to strongly foreign-accented, ranging between the 5 and 7 points of the rating scale, suggesting harsher judgements of L2 Arabic speech than those of L2 English in Munro and Derwing (1995a, 2020), for example, where foreign-accentedness ratings were evenly distributed, mainly around the 2–8 range of the scale. In contrast, comprehensibility ratings of the L2 Arabic were evenly distributed over the 2 and 7 points of the scale, with a moderate trend toward the more easy-to-understand (i.e. comprehensible) end of the scale. While this finding is not in line with findings from studies in L2 English context, including both Munro and Derwing (1995a) and Derwing and Munro (1997), whose results showed extreme skewness toward the more comprehensible end of the scale, it is in line with the most recent findings in L2 Spanish context, Huensch and Nagle (2021). Finally, the average L2 Arabic speech intelligibility scores were rather positively skewed, ranged between 67.75 and 95.73. However, most L2 Arabic speech, 87%, scored below 90. It should be noted the 90% cut-off has been established in the literature based on the observation that this is the minimum score an average native speaker gets (Munro & Derwing, 2020). As such, a score of less than 90 is then indicative of failing the intelligibility threshold and thus, of potential communication breakdown. Looking into individual speech samples that were transcribed with at least 90% accuracy, it was found that only about 30% of L2 Arabic speech met that threshold. Compared to findings from other L2 contexts, L2 Arabic speakers in this study were far less intelligible. In those studies, the majority intelligibility scores were above 90. For example, in Munro and Derwing (1995a, 2020) intelligibility scores were highly skewed toward the more intelligible range; notably, most L2 English speech in that study, 64%, was transcribed with over 90% accuracy; in fact, 53% of the speech was perfectly transcribed. As such, findings from this study suggest L2-Arabic-specific patterns of outcome measures distribution which could be related to L1–L2 pairing, learning context, and/or language proficiency level factors. These factors will be examined more closely in a follow-up study. For instance, regarding L2 Arabic speech comprehensibility, one explanation for the divergence from L2 English research findings and convergence with L2 Spanish research findings could be the commonalities between the L2 Spanish studies and this study in terms of proficiency level and/or learning context. To explain, as in this study, the level of proficiency of L2 speakers in both is intermediate and the learning context is that of a foreign language. In contrast, most speakers in the L2 English studies are highly proficient L2 English speakers in a second language context.
2 The relationships among foreign-accentedness, comprehensibility, and intelligibility
The L2-Arabic-specific patterns are also found when considering the interrelationships among those three constructs in the above comparison studies. Munro and Derwing (1995a, 2020) reported significant relationships among all three constructs with various degrees; the inter-listener correlations, Pearson r, between foreign-accentedness and comprehensibility (for 17 listeners) ranged from .41 to .82, those between comprehensibility and intelligibility (for 15 listeners) from .44 to .90, and finally those between foreign-accentedness and intelligibility (for 5 listeners) from .37 to .48. Later, these two researchers re-analysed the same data from the initial study, adopting a more advanced statistical analysis (i.e. mixed-effects modelling). The re-analysis showed that comprehensibility and foreign-accentedness, individually, were significant predictors of intelligibility was an odd ratio of 1.75 and 1.25, respectively. Given the non-normal distribution of intelligibility scores, they coded intelligibility as a categorical variable such that scores below 90 coded as unintelligible and those above or equal to 90 as intelligible. Huensch and Nagle’s (2021) results showed similar relationships, in general, as those found in the L2 English context (i.e. Munro & Derwing, 1995a); however, comprehensibility, but not foreign-accentedness, was found to be a significant predictor of intelligibility with an odd ratio of 2.07. The results from regression modelling in the current study align more with those from L2 Spanish context as comprehensibility, either individually or after controlling for foreign-accentedness, was found to be a significant predictor of intelligibility, with a medium predictive power, R2 = .44, for comprehensibility-only model and also a medium size effect added predictive power, R2 (Change) = .34, beyond that of foreign-accentedness which had a small predictive power of 13%, for foreign-accentedness-only, but a non-significant added predictive power of .038. Accordingly, in this study, a simple regression model with only comprehensibility appeared to be the best-fitting model. In this model, one unit increase on the comprehensibility scale (high to low) is associated with about 6 points decrease in intelligibility (low to high) scores.
Taken together, the findings from this study provide support for the partial independence of the three dimensions under investigation. That is, while these three dimensions are related, they are still distinct constructs such as L2 speech can be strongly foreign-accented (i.e. high in foreign-accentedness) but not necessarily very difficult to understand (i.e. low in comprehensibility) and/or not recognized as intended (i.e. low in intelligibility). Likewise, L2 speech can be strongly foreign-accented and rather difficult to understand yet it can be still correctly recognized eventually. This is not to claim that foreign-accentedness has no bearing on intelligibility at all, though. Based on results from this study, no conclusive statement can be made on the unique contribution of foreign-accentedness to intelligibility. On the other hand, and with understanding that intelligibility rather than nativeness is the goal for instruction and focus of assessment, it can be concluded that comprehensibility is a better predictor of intelligibility and, therefore, it should be prioritized when it comes to communicative effectiveness.
VIII Limitations and directions for future research
Caution should be taken when generalizing these findings to the entire population of L2 Arabic learners. First, the learners recruited for this study are native speakers of US English who represent a homogenous group of adult American learners of Arabic as a foreign language in university setting; obviously, this group does not represent the entire population of adult Arabic learners in the US. In addition, their proficiency was limited to intermediate level. While purposefully sampled to represent most learners, the findings from this study should not be generalized to other types of learners, that is, learners with lower or higher proficiency levels, or learners in second language contexts. Therefore, future studies should expand the investigation to L2 Arabic learners with novice and/or advanced proficiency, besides intermediate proficiency investigated in this study. This would be beneficial in extending and comparing findings from this study. Another feature of the study worth noting is the lack of control of listener-related factors. Though the study attempted reducing listener’s variability, a construct-irrelevant factor, by recruiting listeners who have the same L1 background, L2 speech familiarity, among others, there is still a potential of listener variance that might have influenced the outcome measures. That said, future studies on L2 Arabic speech should examine listener-based factors such as the Arabic language variety listeners speak, proficiency in the speaker’s L1, their familiarity with L2 Arabic speech. Finally, while the proficiency level of the L2 Arabic speakers in these studies was assessed, this was only done for reporting purposes. The influence of proficiency level on L2 speech assessment was not probed. Still, future research may also need to adopt more standardized proficiency measures like the one used in this study, oral proficiency interview
Specific to L2 Arabic future research, there is much to be revealed about L2 speech global dimensions, their interrelationships across different proficiency levels and different types of learners. For instance, it will be interesting to investigate how those three global constructs of L2 speech are assessed for different groups including native speakers, heritage speakers and non-native speakers. In addition, a consequent step for this study could be examining the linguistic correlates of each dimension; this would be particularly important given that it has been established that those correlates can differ significantly from one dimension to another. For example, regarding foreign-accentedness, which was found to be relatively high in L2 Arabic speech it is still an empirical question to determine what pronunciation aspects, segmental and suprasegmental, result in learners being perceived as strongly foreign-accented, especially when such perception is also aligned with comprehensibility, affecting communicative effectiveness.
IX Conclusions
To the best of our knowledge, this is the first study to investigate foreign-accentedness, comprehensibility, and intelligibility of L2 speech in the context of Arabic as a foreign/second language. In addition, the study extended on existing research methodology by adopting spontaneous speech and more advanced statistical analyses to investigate the relationships among those three L2 speech dimensions and the explanatory or predictive power of foreign-accentedness and comprehensibility for intelligibility in a completely new target language, Arabic, and less represented context, foreign language context. In doing so, the aim was to enhance the reliability and validity of assessment which is needed before any conclusions can be made. As such, the study added to knowledge and understanding of those three global constructs of L2 speech, rendering implications for language assessment, and by extension, instruction. The results confirmed existing evidence on the partial independence of those three L2 speech dimensions in English and non-English contexts while also demonstrating the uniqueness of the nature and shape of their relationships in the L2 Arabic context. It is hoped that this study spearhead empirical research on these global aspects of L2 speech in L2 Arabic and other non-Western languages, more broadly.
Footnotes
Acknowledgements
We would like to thank all L1 and L2 Arabic speakers for their participation. We would also like to thank the anonymous reviewers for their comments and suggestions during the peer-review process.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Data collection was sponsored by the University of Kansas Office of Graduate Studies Summer Research Award.
