Abstract
Stuttering is a multifactorial disorder that is characterized by disruptions in the forward flow of speech believed to be caused by differences in the motor and linguistic systems. Several psycholinguistic theories of stuttering suggest that delayed or disrupted phonological encoding contributes to stuttered speech. However, phonological encoding remains difficult to measure without controlling for the involvement of the speech-motor system. Eye-tracking is proposed to be a reliable approach for measuring phonological encoding duration while controlling for the influence of speech production. Eighteen adults who stutter and 18 adults who do not stutter read nonwords under silent and overt conditions. Eye-tracking was used to measure dwell time, number of fixations, and response time. Adults who stutter demonstrated significantly more fixations and longer dwell times during overt reading than adults who do not stutter. In the silent condition, the adults who stutter produced more fixations on the nonwords than adults who do not stutter, but dwell-time differences were not found. Overt production may have resulted in additional requirements at the phonological and phonetic levels of encoding for adults who stutter. Direct measurement of eye-gaze fixation and dwell time suggests that adults who stutter require additional processing that could potentially delay or interfere with phonological-to-motor encoding.
1 Introduction
Phonological encoding is the widely studied theoretical process by which individual phonemes are selected and incrementally assembled in preparation for speech. Interest in the phonological encoding skills of adults who stutter has been inspired by a number of psycholinguistic theories of stuttering (Howell & Au-Yeung, 2002; Karniol, 1995; Kolk & Postma, 1997; Perkins, Kent & Curlee, 1991; Postma & Kolk, 1993; Wingate, 1988). These theories differ in their details, but all suggest that a potential cause for stuttering may be a delay or difficulty at the level of phonological encoding. A number of studies have investigated whether adults who stutter exhibit a deficit or anomaly in phonological encoding that could contribute to stuttering. A range of tasks have been used to probe phonological encoding in adults who stutter, including silent phoneme monitoring (e.g., Sasisekaran & De Nil, 2006; Sasisekaran, De Nil, Smyth, & Johnson, 2006), rhyme monitoring (e.g., Bosshardt, 2002; Bosshardt & Fransen, 1996; Weber-Fox, Spencer, Spruill, & Smith, 2004; Jones, Fox, & Jacewicz, 2012), nonword repetition (e.g., Byrd, McGill, & Usler, 2015; Byrd, Vallely, Anderson, & Sussman, 2012; Pelczarski, 2011; Sasisekaran, 2013; Sasisekaran & Byrd, 2013), phonological priming (e.g., Hennessey, Nang, & Beilby, 2008), and phoneme elision (e.g., Byrd et al., 2012, 2015; Pelczarski, 2011). The findings have been mixed, with several studies pointing to aberrant phonological encoding whereas other studies suggest it does not differ in adults who do not stutter. In this study, an eye-tracking paradigm was used to test whether phonological encoding in adults who stutter differs from adults who do not stutter in way that contributes to our further understanding of the disorder.
Both general and specific challenges are encountered when attempting to measure phonological encoding in adults who stutter. The general challenge is that phonological encoding is embedded within the processes of language formulation, making direct measurement of phonological encoding difficult with behavioral tasks (Coles, Smid, Scheffers, & Otten, 1995; Meyer, 1992). The specific challenge is that the speech-motor system of adults who stutter appears to be unstable during perceptually fluent speech (e.g., Smith, Sadagopan, Walsh, & Weber-Fox, 2010). This makes it difficult to dissociate a potential aberration in phonological encoding from aberrant speech motor control when using tasks that require overt responses. Despite these limitations, several psycholinguistic accounts of stuttering posit delays or disruptions in phonological encoding as an explanation for stuttering (Howell & Au-Yeung, 2002; Karinol, 1995 Perkins et al., 1991; Postma & Kolk, 1993). These accounts have typically followed the theoretical framework of Levelt, Roelofs, and Meyer (1999) instantiated within their WEAVER++ language formulation model. Linguistic formulation in WEAVER++ is generally composed of three consecutive stages: (a) conceptual preparation (deciding the intended meaning of the utterance); (b) lexical retrieval (selecting the desired lemma); and (c) lexical encoding (morphological, phonological, and phonetic processing). Phonetic encoding begins after phonological encoding has been initiated and provides the speech motor plan that is transmitted to the articulators for production of the utterance (articulatory stage). The phonological encoding stage preceding phonetic encoding and articulation is hypothesized to be more error-prone, or delayed, in stuttering from both theoretical (e.g., Howell & Au-Yeung, 2002; Perkins et al., 1991; Postma & Kolk, 1993) and empirical (e.g., Byrd et al., 2015; Sasisekaran, 2013; Sasisekaran & Weisberg, 2014) reports.
Several experiments have explored phonological encoding in adults who stutter through silent behavioral tasks, nonword stimuli, or a combination of both in an attempt to isolate the phonological encoding stage of processing. Silent tasks do not require overt speech, which hypothetically removes the influence of the speech motor system, whereas nonword stimuli require the assembly of novel phonological segments that mitigates the influence of pre-existing lexical entities. Bosshardt (1990) compared silent and oral reading of a noun plus its definite article. In both the silent and overt reading conditions, adults who stutter took significantly longer to read the article + noun pairs than did adults who do not stutter. Sasisekaran and De Nil (2006) investigated phonological encoding in adults who stutter and adults who do not stutter via silent phoneme monitoring during perceptual and silent picture naming. For both the silent and overt tasks, participants were asked to press a button to indicate when a specified target phoneme was detected. Average response time and percent accuracy were used to evaluate group differences. Adults who stutter demonstrated significantly slower response times for phoneme monitoring during silent reading, but not during a similar perception task with aurally presented stimuli. Sasisekaran et al. (2006) further reported significantly slower phoneme monitoring by adults who stutter during silent picture naming. Brocklehurst and Corley (2011) asked adults who stutter and adults who do not stutter to read tongue twisters silently using silent and overt speech. Although both groups showed more errors in the overt condition than the silent condition, adults who stutter produced significantly more errors than adults who do not stutter in both conditions. Byrd et al. (2012) investigated phonological memory using nonword repetition and phoneme elision and found that adults who stutter were significantly less accurate in repeating seven-syllable nonwords than were adults who do not stutter, but no group differences were reported for the phoneme elision task. Byrd et al. (2015) then replicated and extended this study using nonword repetition and phoneme elision in both silent and overt conditions. Adults who stutter and adults who do not stutter did not differ in silent identification of the nonwords, but adults who stutter demonstrated significantly more errors than adults who do not stutter during overt repetition of seven-syllable nonwords. In the silent phoneme elision task, differences between adults who stutter and adults who do not stutter were evident at the seven-syllable length, whereas the overt phoneme elision task revealed differences in task performance with both four- and seven-syllable nonwords.
Contrary to the studies reviewed above, a smaller set of important studies have not found phonological encoding differences in adults who stutter, although these studies both used a rhyme judgment task, which may account for differences (Bosshardt & Fransen, 1996; Weber-Fox et al., 2004). Bosshardt and Fransen (1996) asked adults who stutter and adults who do not stutter to monitor for phonologically related (rhyming word), semantically related, or identical target words during a reading task. The authors reported that adults who stutter required more time to complete the semantic condition, but no between-group differences for either condition were present. Weber-Fox et al. (2004) conducted a behavioral rhyme judgment task and reported no differences in accuracy between adults who stutter and adults who do not stutter. Reaction time and event-related brain potentials were also not significantly different between groups, but adults who stutter were reported to demonstrate a right-hemisphere asymmetry that was not present in typically fluent adults. It is possible that electrophysiological measures are more sensitive to certain aspects of phonological encoding that may not manifest behaviorally. In summary, most of the phonological encoding studies that did not require overt responses suggest that adults who stutter display slower phonological encoding compared to adults who do not stutter (Sasisekaran & de Nil, 2006; Sasisekaran et al., 2006). Studies including an overt component point to group differences between adults who stutter and adults who do not stutter with larger effects than silent processing. Potentially, additional factors beyond phonological encoding or even the requirement for phonetic encoding is influencing group differences in overt conditions (Bosshardt, 1990; Brockelhurst & Corley, 2011; Byrd et al., 2015).
1.1 Eye-tracking and phonological encoding
Several eye-tracking measures can be used to infer the time course of phonological encoding during naming and reading (e.g., Duchowski, 2007; Griffin & Spieler, 2006; Meyer, Roelofs, & Levelt, 2003; Roelofs, 2008a, 2008b; Venker & Kover, 2015). Saccades are ballistic eye movements that shift the gaze from one location to another at a speed approaching 500 degrees per second. No information gathering or phonological processing is believed to occur during saccades due to the high speed of these transitions (Rayner, 2009). Fixations are typically defined as the period during which the eyes remain still within an area of interest (AOI) designated around the stimulus item. Fixations occur in a non-linear fashion and vary in duration. More frequent and longer fixations are considered to be directly related to the difficulty of the task (Just & Carpenter, 1980; Rayner, 2009). Dwell time is the duration that an individual’s gaze is focused within a given AOI that can include one or multiple fixations of varying lengths (including saccades). Thus, dwell time reflects the duration of fixations on a stimulus, including regressive saccades, as another distinct indication of overall processing difficulty, with longer dwell times being positively associated with task difficulty (Rayner, 1998). Fixation number is another important variable in measuring phonological encoding because it measures the number of times the eyes have made an attempt to collect information. Dwell time, number of fixations, and overall task duration of individual fixations provide valuable measurements of the time course of phonological encoding and are reflective of the synchronous relationship between eye gaze and phonological processing.
Eye-gaze tracking is frequently used in psycholinguistic studies as a measure of tracking visual attention that reflects processing of the stimulus properties. In the case of reading nonwords, it involves visual processing of the orthography and the associated phonological processing to encode the stimulus (e.g., Griffin, 2001; Meyer & van der Meulen, 2000; Roelofs, 2008a, 2008b). Precise temporal measures of eye-gaze fixation position and duration have been used to establish a well-documented time course for phonological encoding during reading and naming (Blanchard, 1985; Duchowski, 2007; Griffin, 2001; Griffin & Spieler, 2006; Korvorst, Roelofs, & Levelt, 2006; Meyer et al., 2003; Rayner, 2009; Reichle, Pollatsek, Fisher, & Rayner, 1998; Roelofs, 2004, 2008a, 2008b; Venker & Kover, 2015). Studies that employ eye tracking infer the length of time needed to encode a phonological plan from the duration of an individual’s fixation upon a word or object, prior to their gaze shifting, or saccade, to the next word. Phonological encoding is believed to be initiated as soon as an individual’s gaze falls upon a stimulus item and ends when gaze shifts to the next stimulus (e.g., Roelofs, 2008a, 2008b). This process occurs separately from the execution of the articulatory plan and gaze will often shift even before the onset of speech production (Korvorst et al., 2006; Levelt & Meyer, 2000), suggesting that phonological encoding can be completed even if the articulatory plan has not been fully executed or initiated.
Several eye-tracking studies have investigated the close association between gaze duration and phonological encoding through various methodological means. Griffin (2001) conducted a study that explored eye-gaze behavior and speech onset during production of high- and low-frequency pictures positioned on a screen. Frequency effects occur when high-frequency words are produced more quickly than low-frequency words reflecting the speed of phonological processing. Griffin reported that the frequency effect was demonstrated in both more rapid speech onset time as well as shorter dwell time for high-frequency words, providing evidence for the synchronicity of eye-gaze durations and phonological encoding.
Priming studies using eye-tracking also lend support to the notion that phonological encoding occurs in the time between gaze onset and offset. Meyer and van der Meulen (2000) conducted an experiment where speakers were asked to name aloud pairs of objects. Phonologically distracting, related, or neutral primes were played at trial onset to determine if the known priming effects would influence eye-gaze duration and speech onset of picture naming. The phonologically related primes revealed the expected effect of shorter speech onset latencies and shorter dwell times. The authors argued that speakers fixated on the picture until phonological encoding was complete, before initiating a saccade to the next picture to be named.
Word length, another phonological variable that can effect speech onset and processing time, has been studied in the context of eye-tracking as well. Zelinksy and Murphy (2000) reported that phonological effects influence gaze-shift latencies even when looking at pictures of words silently during a recognition task. Participants were asked to familiarize themselves with a display of four pictures (two with one-syllable names and two with three-syllable names) and then asked to look at a single picture to indicate if it was present in the previous display. Participants demonstrated significantly more fixations and dwell time on pictures with three syllables than one syllable. These findings demonstrate that gaze durations were affected by phonological variables that influence phonological encoding. Meyer et al. (2003) asked participants to name aloud two successive pictures of one- and two-syllable objects. The authors reported that there were no differences in recognition time between the one- and two-syllable pictures; however, eye-gaze durations (i.e., dwell times) were longer for the two-syllable words than the one-syllable words in initial position. The researchers attributed this effect of phonological length (i.e., increased dwell times) to participants requiring more time to phonological process the longer, two-syllable words. Meyer and colleagues also argued that gaze shifts to the next word are not initiated until the phonological encoding for the first word is completed. Roelofs (2008b) also reported phonological effects that influence the gaze fixations and durations during phonological encoding. Phonologically related photographs in a picture-naming task will facilitate speech onset whereas semantically related pictures will hinder naming speed. He reported that the distractor pictures affected gaze-shift latencies as well as speech onset times in similar ways, again reflecting the close relationship between gaze and phonological encoding.
Gaze durations in typical populations have been shown to be demonstrably longer when reading aloud than when reading silently (see Rayner, 1998; Vorstis, Radach, & Lonigan, 2014). An explanation of differences between silent and overt single word reading is not fully known, but likely involves increased time to develop a highly specified phonological code for phonetic encoding and/or speech-motor execution. Although the measurement of eye-movement variables as correlates for phonological encoding does not necessarily require spoken output, a highly specified phonological plan may be necessary to reveal subtle phonological processing challenges for adults who stutter. Nonword production can be particularly challenging for individuals who stutter, and differences in the phonological systems of adults who stutter are often more evident when nonword, as opposed to real-word, stimuli are used (e.g., Byrd et al., 2012; 2015). Consequently, nonword targets were selected to provide the hypothesized challenge to the phonological system of adults who stutter. Based on previous phonological studies (e.g., Byrd et al., 2015; Sasisekaran et al., 2006) and certain articulatory kinematic studies (e.g., Kleinow & Smith, 2000; Sasisekaran & Weisberg, 2014, Smith, Sadagopan, Walsh, & Weber-Fox, 2010), we hypothesize that adults who stutter will demonstrate delays in phonological encoding through significantly longer eye-gaze duration during silent and overt reading of nonwords compared to adults who do not stutter. The overt condition could be potentially challenging due to additional requirements at the phonological, phonetic, and articulation levels, each of which have to be ruled out systematically. Therefore, both silent and overt reading conditions were important for this project. This is the first eye-tracking study, to our knowledge, to measure phonological processing in stuttering during silent reading and overt articulation of nonwords.
2 Method
2.1 Participants
The participants included 18 adults who stutter (mean age: 29.4 y; SD: 9.2 y) and 18 adults who do not stutter (mean age: 27.1 y; SD: 12.1 y). All participants were monolingual, spoke Standard US English, and did not possess any speech, language, hearing, or neurological disorders other than stuttering. Thirteen out of 18 adults who stutter reported previous speech therapy for stuttering. A history of treatment is very common among adults who stutter but there is no evidence that previous treatment influences phonological encoding performance (e.g., Byrd et al., 2015; Jones et al., 2012; Logan, Byrd, Mazzocchi, & Gillam, 2011). Years of education ranged from 13 to 22 years (adults who stutter: 14–22; mean: 16.5; SD = 2.6; adults who do not stutter: 13–22; mean: 16.9, SD = 2.3).
Each participant’s fluency status was determined through analysis of a spontaneous speech sample of at least 400 syllables. Participants were determined to be an adult who stutters if they received a score of 18 (mild) or higher on the Stuttering Severity Instrument 4 (SSI-4; Riley, 2009); demonstrated at least three stutter-like disfluencies (part-word repetitions, sound prolongations, or blocks) per 100 words during a conversational speech sample; and self-identified as a person who stutters. Participants were deemed an adult who does not stutter if they received a score of 17 or below (i.e., less than mild) on the SSI-4; demonstrated fewer than three stutter-like disfluencies per 100 words of conversational speech; and did not self-identify as a person who stutters. The SSI-4 was used to assign a severity rating to the participants who stutter (mild: n = 9; mild-moderate: n = 6; moderate: n = 3).
2.2 Materials and design
Each condition included 15 nonwords (three one-syllable, four two-syllable, four three-syllable, four four-syllable) for a total of 30 stimuli across the overt and silent conditions. Due to a coding error, one of the one-syllable nonwords in each condition was incorrectly coded as a practice item and therefore were not included in data analysis. Nonwords were selected from a list generated by the English Lexicon Project (Balota et al., 2007). Nonwords were sorted by phoneme length and bigram frequency by position. Nonwords were excluded if they contained additional morphemes or irregular spelling, and the lists were narrowed down to 60 possible stimuli. Ten untrained participants were shown the potential list of stimuli and asked to read them aloud. Responses were recorded and transcribed to determine nonword stimuli that possessed the most straightforward pronunciations. Selected stimuli were then balanced across conditions (silent and overt) and syllable number (1–4 syllables). Nonword presentation was randomized across syllable lengths, and presentation order was counter-balanced, with half the participants assigned to the silent condition first and the remainder completing the overt condition first. Detailed verbal instructions were provided prior to the session and in written form immediately prior to word presentation.
2.3 Apparatus
All eye-tracking measures were recorded with an SR EyeLink 1000 system operating at a 1000-Hz sampling rate following the guidelines outlined by Stampe (1993). Eye gaze was calibrated for position and drift at the start of each experimental session using a 9-point calibration process and recalibrated as necessary during administration of the individual trials in case of head movement or excessive blinking. To minimize head movement, participants rested their head on a forehead bar and chin rest. Head motion was further minimized in the overt condition—when the chin rest could not be used—by using a head strap to hold the head closer to the forehead bar in a comfortable, yet secure way. In most cases, there was no need for additional calibration during the course of the study. Stimuli were presented on a ViewSonic VX2265wm monitor positioned 92 cm from the participants’ forehead. An Audio Technica AT813a condenser microphone was positioned 6 inches from each participant’s mouth to optimize recording. An M-Audio M-Track Plus sound card was used to record the spoken response of each participant during the overt condition.
2.4 Procedure
The time course of phonological encoding has been determined to begin when visual attention is focused on a stimulus and end when the gaze shifts away from the stimulus, regardless of spoken production requirements (e.g., Griffin, 2001; Meyer et al., 1998; Roelofs, 2008b). Therefore, an important component of the eye-tracking task is to obtain measurements from the time spent gazing solely at the nonword to provide an accurate duration. The eye-tracking experiment consisted of two conditions: a silent reading condition and an overt reading condition. Participants were required to first fixate on a white ampersand symbol (“&”) located on the left-hand, central side of a black computer screen (fixation screen). Once a steady gaze was achieved, the trial was initiated. The fixation point remained, joined by the nonword stimulus in the center of the screen, and secondary stimulus in the right-hand center of the screen that required a manual response (Figure 1). The positioning of the fixation point, stimulus, and manual response task required participants to shift their gaze onto and off the nonword providing a clear indicator of gaze duration on each stimulus. All stimuli were displayed in white 60-point font on a black background. Participants were asked to shift their gaze from the fixation point “&” to the nonword stimulus (e.g., “hekai”) in the center of the screen and then read the item silently or overtly. Participants then immediately looked to a secondary stimulus in the right side of the screen (“XX>XX” or “XX<XX”) and pressed an arrow button to indicate the direction of the embedded arrow (“XX>XX” or “XX<XX”). Oral responses were recorded in the overt condition to assess accuracy and measure duration of word production, while also monitoring for disfluencies. In the overt condition, participants were directed to carefully read the nonword aloud and then look to the manual button-press task positioned to the right of the nonword. In the silent condition, participants were instructed to carefully read the nonword silently to themselves and then look to the manual button-press task. Selection of the button corresponding with the direction the arrow was pointing completed a trial.

Example of stimulus presentation on the eye-tracking screen.
Phonological encoding duration was estimated to correspond to the duration of all fixations located within an AOI that surrounded the nonword stimulus, that is, dwell time. The start and end points were automatically calculated as follows: the start time corresponded to the end of the saccade from the fixation point on the left to the nonword in the center. The end point was determined as the start time of the saccade from the nonword to the secondary (decision) stimulus. The button press indicating a left or right arrow button after the final saccade to the right side ended each trial. The overall interval from gaze initiation until button press was labeled as “response time” (see below). After indicating the direction, the stimulus screen would disappear and the participant would again focus on the “&” of the fixation screen until the next trial began.
2.5 Dependent variables
Dwell time. The total time (in milliseconds) in which fixation(s) on the nonword stimulus occur within the rectangular AOI based on the start and end time described above.
Fixation number. The number of times the eyes remained still (focused) while gazing at the word and remaining within the rectangular AOI.
Response time. Response time is the overall time (in milliseconds) from onset of the ampersand signal until the button press, which represents the total time required for the task. As a follow-up analysis, Dwell time was subtracted from response time to also focus on the timing components of the gaze shift from the fixation point (left), to word (center) and to symbol string (right), apart from dwell time. This subtraction variable assessed whether dwell time was a primary determinant of response time.
Duration. The duration of word production (in milliseconds) during the overt condition measured from the acoustic recording. Two trained research assistants marked the onset and offset of the oscillogram in the software PRAAT (Ver. 6.0; Boersma & Weenink, 2016) and calculated the duration. The duration estimates were compared between the two raters. No statistical differences between the two raters were found. Additionally, the raters were instructed to flag any disfluencies during the eye-tracking task so that those trials could be excluded. No disfluencies were observed in any of the productions of either adults who stutter or adults who do not stutter.
Accuracy. The nonwords produced in the overt condition were transcribed separately by two trained research assistants skilled in phonetic transcription. Nonwords were then scored to determine overall accuracy (i.e., was nonword production accurate) and percent phonemes correct (PPC; i.e., the percentage of correct phonemes divided by the total number of phonemes). Self-correction of nonword production was included in the analysis; however, only the final production was included (whether it was correct or incorrect). Any differences between the transcriptions and PPC were discussed and resolved by consensus with the first author. Results revealed that there were no statistical differences between the raters.
2.6 Data analysis
Repeated analysis of variance (ANOVA) measures were used with response type (overt, silent) and number of syllables (1–4) as within-subject factors and group (adults who stutter, adults who do not stutter) as a between-subjects factor. A separate ANOVA was conducted for each dependent variable: dwell time, fixation number, response time, duration, and accuracy. The variables Duration and Accuracy could not be compared for the silent condition, so only the effect of number of syllables was compared for this variable across the two groups.
3 Results
There were several general findings that illustrate both groups were performing the nonword reading task as anticipated. First, in terms of dwell time (see Figure 2A), overt production elicited significantly longer overall dwell time than silent reading, in regard to response type: F(1, 34) = 71.6, p < 0.001, η2 = 0.68. Second, increasing syllable number resulted in longer dwell time, with the number of syllables: F(3, 102) = 103.2, p < 0.001, η2 = 0.78. Pairwise contrasts for dwell time indicated Syllable 1 < Syllable 2 < Syllable 3 < Syllable 4 (p < 0.05). Similar effects were found for fixation number (Figure 2B). The number of fixations increased significantly from the silent to the overt condition, F(1, 34) = 39.9, p < 0.001, η2 = 0.92; and with syllable number, F(3, 102) = 83.7, p < 0.001, η2 = 0.97, again increasing consecutively from Syllable 1 to Syllable 4 (p < 0.05). No significant correlations were found between stuttering severity, dwell time, and number of fixations for the adults who stutter. A significant correlation between dwell time and number of fixations was not found for the adults who do not stutter.

The primary eye-tracking variables from all participants are displayed for syllables 1–4 in the silent and overt conditions.
3.1 Dwell time
Kolmogorov–Smirnov testing for normality of distribution did not indicate departures from normality. adults who stutter had significantly longer dwell times than adults who do not stutter as shown by a significant main effect of the group, F(1, 34) = 4.70, p < 0.05, η2 = 0.12, but this group difference was mediated by a significant interaction between group and response type, F(1, 34) = 4.18, p < 0.05, η2 = 0.11 (Figure 3). The stuttering group showed significantly longer dwell time than the fluent group in the overt condition, whereas group differences were not detected in the silent condition. No other interactions with the group variable were found. Inspection of the Greenhouse Geisser and Hyunh–Feldt adjusted results did not change the general findings or group effects.

Dwell time in milliseconds for adults who stutter and adults who do not stutter in silent and overt reading conditions.
3.2 Fixation number
Adults who stutter made significantly more fixations on the nonwords than adults who do not stutter indicated by a main effect of the group, F(1, 34) = 6.2, p < 0.05, η2 = 0.15, and shown in Figure 4. Kolmogorov–Smirnov testing did not indicate departures from normality. None of the other factors interacted with the group, suggesting adults who stutter made more fixations across all experimental conditions. Inspection of the Greenhouse Geisser and Hyunh–Feldt adjusted results did not change the general findings or group effects. For both groups, fixation number was significantly correlated with dwell time in the overt condition: adults who stutter – r(18) = 0.74, p < 0.01 and adults who do not stutter – r(18) = 0.71, p < 0.01; and the silent condition for both groups: adults who stutter – r(18) = 0.68, p < 0.01 and adults who do not stutter – r(18) = 0.62, p < 0.01. The correlations between these variables and stuttering severity were not significant.

Mean number of fixations for adults who stutter and adults who do not stutter across both silent and overt reading conditions.
3.3 Response Time
Overall response time was significantly longer as syllable number increased, F(3, 102) = 92.6, p < 0.001, η2 = 0.55; and in the overt condition, F(1, 34) = 134.3, p < 0.001, η2 = 0.65, across both groups. The response type by group interaction, F(1, 34) = 7.1, p < 0.05, η2 = 0.15, indicated that the significant increase in response time for overt production was greater for adults who stutter (Figure 5a). This interaction was similar to dwell time results, so it was investigated further by subtracting dwell time from response time. Repeated measures of ANOVA of the subtraction variable indicated a significant increase in experiment time as syllable number increased, F(3, 102) = 4.8, p < 0.05, η2 = 0.11 (see Figure 5b) and in the overt condition, F(1, 34) = 16.6, p < 0.01, η2 = 0.12. In particular, in the silent or overt conditions, there were no group differences in the timing of the experimental task (Figure 5c).

The timing variables across phases of the trial.
3.4 Duration
Speech duration increased significantly with the number of syllables in the overt condition (Figure 6), F(3, 90) = 114.1, p < 0.001, η2 = 0.92). Group differences in speech duration were not detected.

Acoustic duration (ms) of the words spoken in the overt condition of the adults who stutter and adults who do not stutter.
3.5 Accuracy
PPC was compared between groups and across syllables. Group differences in PPC were not found but the effect of syllables was significant, F(3, 90) = 14.2, p < 0.001, with PPC decreasing as syllables increased. Post hoc testing of PPC indicated the following: Syllable 1 > Syllable 3, Syllable 1 > Syllable 4 and Syllable 2 > Syllable 4 (p < 0.05).
4 Discussion
The current study investigated silent and overt reading of nonwords in adults who stutter using eye-gaze dwell time and fixation number as measures of phonological encoding. The primary hypothesis, that adults who stutter take longer to encode the nonwords in both the silent and overt conditions than typical speakers, was partially supported. Group differences were found for both conditions, but were most evident in the overt condition. The adults who stutter had significantly longer dwell times during overt articulation but not during silent reading. Adults who stutter also produced more fixations on the nonwords than adults who do not stutter across both silent and overt conditions. Otherwise, both groups demonstrated general patterns of longer dwell times and more fixations in the overt condition and as syllable number increased. This pattern is a primary indicator of differences in processing demands when overt responses are required that aligns well with previous studies (e.g., Breen & Clifton, 2011; Rayner, 1998; Vorstius, Radach, & Lonigan, 2014).
One particular advantage eye-tracking research confers is the ability to temporally parse eye-gaze parameters on the linguistic task of nonword reading from the other components in the task (e.g., gaze-shifts, viewing the fixation point, and arrows). In contrast to behavioral studies of phonological encoding, eye-tracking is a physiological measure that enabled us to focus on the time period when participants were actively engaged in viewing the word. Given the remaining temporal components could be separated from the nonword dwell time measure, the processing demands occurring during the reading condition distinguished the groups, rather than task performance just being generally slower for adults who stutter.
4.1 Additional processes
Different cognitive processes are engaged when reading silently versus aloud. Silent reading requires visual interpretation of orthographic stimuli, orthographic coding, phonological encoding, and potentially subvocal rehearsal. Overt reading also requires visual interpretation of orthographic stimuli, orthographic coding, and phonological encoding, but additionally there will be syllabification of the newly constructed novel stimulus, phonetic encoding (establishment of a motor plan), and actual motor execution. The additional processing required for reading could increase dwell time for overt speech relative to silent speech as per our general finding that dwell time increased significantly in the overt condition (e.g., Rayner, 2009). However, deeper consideration is needed about whether it can explain the group differences. The first of two primary findings for stuttering from this study is that adults who stutter only had longer dwell times in the overt condition. Dwell time or reading time represents the total length of eye-gaze fixation in a pre-established area of interest (Duchowski, 2007). The second finding is that the number of fixations on the nonword was significantly higher for adults who stutter across both conditions. Attributing additional processes for overt production only contributes to the dwell-time finding, but not the fixation-number finding. Essentially, the overt condition reveals a potentially relevant temporal lag in processing time for adults who stutter that could reflect either slower phonological encoding or slower phonetic encoding, which lead to different conclusions about the cause of stuttering. The same adults who stutter, however, also made more eye movements in processing the nonword in both conditions and accordingly, more nuanced explanations may be required for these findings.
Attentional load may also be a contributing factor. The number of fixations during reading could reflect the location and degree of an individual’s overt visual attention. The larger number of fixations made by adults who stutter in both conditions may reflect a need to maintain a higher degree of task-related visual attention on nonword stimuli than adults who do not stutter. There is some evidence that an increase in attentional load conceivably contributes to slower overt dwell time in adults who stutter (Jones et al., 2012). Under this line of thought, attentional focus needs to be maintained in both the silent and overt conditions with the main difference between response conditions being the requirement to speak the word aloud. As the number of fixations increased for both groups in the overt condition, the phonetic plan or the speech motor production itself might require further attentional resources. There is still little evidence that speech motor encoding or speech production itself increases attentional load relative to silent or overt reading; therefore, a more parsimonious interpretation is that the nonword reading task itself required a greater degree of visual attention for the adults who stutter than for adults who do not stutter and, accordingly, was more challenging for the adults who stutter.
4.2 Phonetic versus phonological encoding
A possible difference in phonetic encoding is supported by studies that have reported anomalies in speech motor control in adults who stutter. An especially relevant finding showed that adults who stutter produced more variable lower lip movements when producing difficult nonwords than adults who do not stutter (Smith et al., 2010). This finding has also been reported for utterances composed of real words (Kleinow & Smith, 2000. Therefore, a delay in generating the phonetic code and perhaps subsequent production of the word could be operating, rather than a phonological issue.
Some potential counter evidence against a strict phonetic explanation for the current study is that group differences in spoken word duration were not found and there were no disfluencies. There are other reasons to consider a phonological encoding explanation for the group differences in dwell time even with overt production. The first is a relatively new idea that phonological representations are only fully engaged in association with production, whereas they are only engaged to a ”shallow degree” with silent production. Oppenheim and Dell (2010) investigated inner speech via production of tongue twisters using silent speech with articulatory movements (mouthing) and entirely silent reading. The articulated silent speech condition elicited more tongue twister errors than the silent (non-articulated) condition in adults who do not stutter. The authors interpreted this finding as evidence of both lexical and phonemic access during the production task, whereas in the silent condition (non-articulated silent speech) there was only evidence for lexical access. Thus, there appears to be some support that activation of a speech motor plan, even if not fully realized (silent mouthing), distinguishes phonological encoding for overt production from phonological encoding for silent reading. Therefore, if phonological encoding is not fully activated in silent conditions, or only activated to a shallow degree, then group differences due to phonological encoding might not arise accounting for the current findings as well as those of Bosshardt and Fransen (1996) and Weber-Fox et al. (2004).
Another possibility might be that only the overt condition was challenging enough to engage phonological encoding to the degree that group differences could be identified, although differences are still present for silent processing. Even though the types of tasks and difficulty levels have varied widely across phonological encoding studies (including varied silent and overt conditions), the majority reported delayed, unusual, or error-prone phonological conditions. Additionally, differences in silent conditions only emerged during the most challenging aspects of the silent tasks, whereas group differences consistent with aberrant phonological processing were reported frequently during overt tasks (e.g., Weber-Fox et al., 2004). The nonwords used in the present study only ranged from 1 to 4 syllables, and accordingly may not have been challenging enough in the silent condition. This is a different position from the reasons given above, but essentially posits that there is an identifiable phonological encoding problem even in silent reading as adults who stutter produced more fixations on the stimuli than adults who do not stutter during the silent condition. More investigation with a wider range of words and syllable lengths is needed to test if this condition was challenging to adults who stutter differentially. This accords with the behavioral studies with silent task rhyme judgements and phoneme monitoring that reported significant delays in phonological processing for adults who stutter (Bosshardt, 1990; Brockelhurst & Corley, 2011; Byrd et al., 2015; Sasisekaran & de Nil, 2006; Sasisekaran et al., 2006). It remains possible, however, that the less pronounced findings in the silent condition argue against differences in phonological encoding in people who stutter. Considering the diversity of findings to date and the emergence of combined physiological-behavioral paradigms to study phonological encoding, ongoing phonological encoding studies are needed to determine if there a phonological problem in stuttering. Future eye-tracking studies will need to increase the number of syllables to make the silent and overt tasks more difficult with continuing attention given to avoiding phonotactic violations.
4.3 Phonemic-phonetic interaction
A second reason, which deviates from the serial interpretation that phonological processes precedes phonetic production, is that overt production reveals interactions between these mechanisms. The Weaver ++ model may not fully capture the interactions between language formulation and motor speech processes, which may be a limitation of the model, as the intra-syllabic breakdowns of stuttering are not typical disfluencies. Howell and Au-Yeung (2002) proposed the EXPLAN model of stuttering to account for stuttering in terms of phonological complexity. The authors outlined two concurrently occurring stages: the planning stage (e.g., PLAN) and the execution stage (e.g., EX), and suggested that stuttering would occur if the interface of linguistic planning and motoric execution (i.e., the interaction between phonological and phonetic encoding) was mistimed. Howell and Au-Yeung argued that dis-synchrony can occur when a segment is challenging and requires additional processing time. The authors defined difficult segments as phonologically complex segments containing consonant clusters, late-developing consonants, or that possess increased word or segment length. Thus, according to the theory, increased phonological difficulty would result in slower retrieval of those segments during phonological encoding and a dis-synchrony between the speech and motor systems would occur. No discernable stuttering occurred during nonword reading in the current study, but our findings of significant differences in the overt condition would only be explained by the EXPLAN model because there was no opportunity for mistiming of the phonological and phonetic plans in the silent condition. The longer dwell time in the overt condition for the adults who stutter could have resulted in longer word duration under EXPLAN if there was a phonological/phonetic mismatch. However, it is not clear whether word duration differences are a characteristic of the disorder or a reaction to stuttering or a by-product of previous therapy. Longer nonword utterances may be necessary to determine if duration differences arise.
The alternative interpretation given by Smith et al. (2010) is relevant in that the higher variability of speech movements for nonwords could be a window into faulty interactions between phonology and motor processing in stuttering. The current results provide further evidence that an interaction between these mechanisms could differentiate individuals who stutter, but it is still challenging to identify and test the mechanism, because it involves transitions. In Smith et al. (2010), an interaction interpretation is strengthened because group differences were only found for the more complicated nonwords. We could not evaluate an interaction in the same way in this study even though dwell times increased with word length in both conditions, but to a greater degree for the overt condition. Inspection of Figure 3 also shows the difference in overt dwell time between the groups was greatest for the 3- and 4-syllable conditions. Further considerations of an interaction should be theoretically fruitful for understanding stuttering.
Other studies have reported longer overt reading times for adults who stutter that could arise due to faulty interactions between phonological and phonetic processes. Bosshardt and Nandyal (1988) reported that adults who stutter required longer durations for single word reading than adults who do not stutter in both silent and overt conditions. Postma, Kolk, and Povel (1990) measured speaking rates for silent, lipped, and overt reading of tongue twisters and matched control sentences in adults who stutter and adults who do not stutter. The authors reported that adults who stutter were slower than adults who do not stutter in all conditions, but particularly for the overt condition. Unfortunately, the methodology we used could not adequately identify the point of speech/vocalization onset relative to the eye-tracking task, which might have been related to eye-gaze events. In future studies, a sensitive and reliable threshold for speech onset needs to be calibrated in order to determine how it relates to fixation and dwell time on the stimulus. For example, if speech onset is reliably occurring towards the end of dwell time on the stimulus, then dwell time is potentially separable from the overt articulation.
4.4 Inner speech
There have been several other studies that did not use typical evaluations of phonological processing that are pertinent to this argument. The intriguing results from the tongue twister tasks of Brockelhurst and Corley (2011) suggest phonological encoding is engaged actively for silent and overt production, in contrast to Oppenheim and Dell (2010). The adults who stutter reported more phonological errors in the silent condition or inner speech (following the term of Brocklehurst & Corley, 2011) and had more actual errors in the overt condition. Other researchers into stuttering have found evidence that inner speech might differ in adults who stutter, even if errors or timing are not monitored (Brocklehurst & Corley, 2011; Hartsuiker & Kolk, 2001; Ingham, Fox, Ingham, & Zamarripa, 2000; Postma & Kolk, 1993). de Nil, Kroll, Kapur, and Houle (2000) reported cortical and subcortical activity from positron emission tomography scans during reading of single words in silent and overt reading conditions. Different patterns of activation during overt and silent reading (inner speech) were found between groups. During the inner speech condition, the adults who stutter demonstrated significantly more activation in the left anterior cingulate cortex than adults who do not stutter. A different pattern was revealed between groups in the oral reading condition in that adults who stutter demonstrated more right-hemisphere activation as compared to adults who do not stutter. The brain activity differences suggests that reading can be different for adults who stutter whether the task is overt or silent. Ingham et al. (2000) reported an unusual positron emission tomography study of imagined stuttering. The adults who stutter showed aberrations in brain activity during overt speech as well as during imagined stuttering, but the inner speech of adults who do not stutter did not show the same aberrations. Along with other physiological measurements taken during silent and overt reading, eye-tracking can be used to highlight subtle differences between the phonological process of adults who stutter and adults who do not stutter.
5 Conclusions
In this preliminary study, measurement of eye-gaze fixations and dwell times revealed a potential delay in phonological to motor encoding that affects adults who stutter. Further study of phonological encoding with real words and longer utterances are needed to isolate whether adults who stutter have aberrant or delayed phonological encoding. It is also not clear if the observed differences depend on speech motor control differences. We reasoned that longer dwell time and increased fixation number could follow from anomalies in phonological encoding or anomalous interactions between phonological and phonetic processes. It has not been determined how anomalous phonological processes theoretically contribute to stuttering disfluencies, but our interpretation is that some aspect of phonological or phonological-to-phonetic encoding is more difficult for adults who stutter, apart from whether there is a difference in speech motor control. Our findings show that eye-tracking is a valuable paradigm for investigating stuttering by making important progress in isolating components of the language formulation process. Dwell time and fixation number are important indicators of phonological processes, but a way of tracking speech onset and speech motor control relative to the ongoing eye-movements is needed.
