Abstract
Objective:
This study evaluated the individual and combined effects of voice (vs. manual) input and head-up (vs. head-down) display in a driving and device interaction task.
Background:
Advances in wearable technology offer new possibilities for in-vehicle interaction but also present new challenges for managing driver attention and regulating device usage in vehicles. This research investigated how driving performance is affected by interface characteristics of devices used for concurrent secondary tasks. A positive impact on driving performance was expected when devices included voice-to-text functionality (reducing demand for visual and manual resources) and a head-up display (HUD) (supporting greater visibility of the driving environment).
Method:
Driver behavior and performance was compared in a texting-while-driving task set during a driving simulation. The texting task was completed with and without voice-to-text using a smartphone and with voice-to-text using Google Glass’s HUD.
Results:
Driving task performance degraded with the addition of the secondary texting task. However, voice-to-text input supported relatively better performance in both driving and texting tasks compared to using manual entry. HUD functionality further improved driving performance compared to conditions using a smartphone and often was not significantly worse than performance without the texting task.
Conclusion:
This study suggests that despite the performance costs of texting-while-driving, voice input methods improve performance over manual entry, and head-up displays may further extend those performance benefits.
Application:
This study can inform designers and potential users of wearable technologies as well as policymakers tasked with regulating the use of these technologies while driving.
Keywords
Introduction
Crashes related to driver distraction are a large concern in the United States, with 3,328 people killed and over 421,000 injured due to driver distraction in 2012 (National Highway Traffic Safety Administration [NHTSA], 2014). Factors such as complexities in the driving environment and types of interaction methods used for in-vehicle tasks contribute to the level of driver distraction (Tsimhoni & Green, 2001). As drivers commonly perform secondary tasks while driving, often for navigation and communications purposes, characteristics of in-vehicle displays and interfaces can significantly impact driver distraction and safety. The current study investigated how voice input and head-up display functionalities affect driver distraction.
Human information processing theory offers insights into how distraction occurs and why it is so problematic for drivers (e.g., Hurts, Angell, & Perez, 2011). For example, “limited capacity” models suggest that the ability to timeshare driving and other tasks depends on the types and magnitudes of demands imposed by the task set on limited mental resources (Wickens, 2002). These resources can be broadly defined as a common but limited pool of attentional resources (Kahneman, 1973) or described as semi-separable perceptual, cognitive, and response resources (Wickens, 1980, 2002). As driving heavily loads the visual perceptual channel, requires spatial working memory processing, and demands manual interaction for controlling the vehicle, these theories suggest that secondary tasks will present the most difficulties when they require the same visual, spatial, and manual resources as driving (Hurts et al., 2011; Wickens & Horrey, 2008), potentially resulting in performance decrements due to structural or cognitive interference. Structurally, eyes can only support one field of view, and individual appendages can only perform one motor activity at a time. Cognitively, interferences may arise from secondary tasks that require similar working memory resources as the primary task (Drews, Yazdani, Godfrey, Cooper, & Strayer, 2009; Hurts et al., 2011; Sawyer, Finomore, Calvo, & Hancock, 2014). These cognitive interferences may result from confusion between tasks or when the combined mental workload imposed by the tasks exceeds the capacity for a cognitive resource (i.e., the “cognitive redline”; e.g., Grier et al., 2008; Rodriguez, Yang, Tippey, & Ferris, 2015).
This theoretical discussion provides insight into the NHTSA’s identification of secondary activities that pose the greatest risk for driver distraction: those that can produce structural interference with visual (e.g., eyes-off-road) or manual (e.g., hands-off-wheel) resources and those that can produce cognitive interference (e.g., mind-off-driving) with driving-related processing (NHTSA, 2013). This explains why texting on a mobile device, a task that requires visual perceptual, spatial cognitive, and manual resources, is especially problematic to attempt while driving (see Figure 1). The detrimental effects of texting and driving on multitask performance and driving safety have been demonstrated in more than several controlled experiments (e.g., Drews et al., 2009; Fitch, Hanowski, & Guo, 2015; Horrey & Wickens, 2007; Lyngsie, Pedersen, Stage, & Vestergaard, 2013; Tsimhoni & Green, 2001; Yager, 2013). Naturalistic studies provide further evidence of these detrimental effects, showing texting to be the secondary activity associated with the largest increase in crash risk (23-fold) compared to non-distracted driving (e.g., Fitch et al., 2013; Olson, Hanowski, Hickman, & Bocanegra, 2009).

Representation of how the tasks of texting and driving compete for similar visual perceptual, spatial cognitive, and manual resources. The dashed arrows indicate that manual and vehicle control activities require the engagement of spatial working memory and visual perception but that the automatic nature of these activities suggests relatively little demand is imposed on those resources.
Despite the documented risks and legal implications of texting while driving, this activity remains a common practice. Some professions even implicitly encourage and expect that the two tasks (driving and messaging) be performed concurrently as part of job duties (Yager, Dinakar, Sanagaram, & Ferris, 2015). While legal and social efforts continue to make strides toward reducing this problematic practice, researchers can take advantage of the familiarity most drivers have with texting to learn more about how to support multitasking while driving. Texting provides a convenient experimental “surrogate” for more objectively beneficial secondary tasks that also require visual, manual, and cognitive resources, such as consulting a GPS or other in-vehicle information systems. This study therefore sought to understand how interface characteristics (e.g., display and input modalities) for a secondary texting task impact driving safety and performance. This understanding can provide insight into ways to make secondary tasks that impose similar resource demands as texting less unsafe to perform while driving (e.g., Liu & Wen, 2004).
First, consider display characteristics in the texting task. Reading or manually responding to texts requires reorienting focal visual attention away from the road to the texting device. Both the magnitude of this reorientation and duration of eyes-off-road time impact safety (Horrey, Wickens, & Consalus, 2006; Hosking, Young, & Regan, 2009; Wittmann et al., 2006). Since smartphones and other in-vehicle technologies are often positioned outside the field of view used for driving, these devices require more disruptive reorientations than displays that at least partially share the same field of view as the roadway. This suggests that performance and safety may be improved with technologies that allow secondary tasks to be conducted within a field of view that partially overlaps with the roadway, thus making visual resources more “sharable” with the concurrent driving task.
Second, consider input characteristics. Many complex mobile devices have replaced physical keyboards with touchscreens. This input change makes texting while driving even more detrimental to driving performance (Lyngsie et al., 2013), potentially because verifying touchscreen button activation puts a higher demand on visual resources to compensate for diminished haptic feedback. An alternative to manual input is to use voice-to-text input, which substantially reduces structural interference by freeing the hands. Some evidence suggests voice-to-text input may be less detrimental to driving performance (He, Chaparro, et al., 2014), promoting more eyes-on-road time and reducing subjective mental workload when compared to manual input methods (Tsimhoni & Green, 2001). However, voice-to-text methods still require eyes-off-road time to read incoming messages and verify the correctness of speech-to-text translations (Yager, 2013). Additionally, while mental resource competition may be reduced with verbal compared to manual entry, the load imposed by verbal annunciation is similar to that required for phone conversations while driving and can result in similar performance decrements (Filtness, Lenné, & Mitsopoulos-Rubens, 2013). Moreover, as device input and output characteristics can affect the level of demand imposed on perceptual (e.g., vision, audition), cognitive (e.g., spatial and verbal working memory), and response (e.g., hands, voice) resources in complex ways, the effects of device characteristics on driver safety and performance must be considered as an emergent property (e.g., Leveson, 2011). Hence, the impact of these characteristics on driver safety depends on the individual and interacting demands that they impose on human information processing resources.
Recent years have seen a growth in the development of advanced technologies that have the potential to better support device interactions while driving. Of note, Bluetooth-connected smartphone extensions such as Google Glass, Sony’s SmartEyeglass, and Samsung’s Galaxy Glass are becoming more commonplace, and the usage of these devices while driving is an increasingly urgent issue that policymakers, transportation engineers, and hardware and software designers are working to address. These devices combine head-up display (HUD) functionality—which provides output close to the same field of view as the roadway, thus improving the driver’s ability to perceive events in the forward scene (e.g., Flannagan & Harrison, 1994; Kiefer, 1991; Kiefer & Gellatly, 1996; Okabayashi, Sakata, Furukawa, & Hatada, 1990; Sojourner & Antin, 1990)—with alternative input methods (e.g., voice input). The combination of these and other device characteristics may result in emergent performance and safety effects that are not fully observed with the characteristics in isolation (Leveson, 2011).
Recent research conducted with Google Glass—a wearable technology with a transparent prism screen mounted in front of the right eye, multi-axis accelerometers for head-based gesture controls, and voice input and output functionalities—illustrates the potential of advanced interfaces to support some in-vehicle tasks while driving. He, Ellis, Choi, and Wang (2015) compared reading performance on a smartphone versus Glass and found that while medium and long text messages both impaired driving performance, using Glass resulted in smaller driving performance decrements than using a smartphone. Building on this study, He, Choi, McCarley, Chapparo, and Wang (2015) sought to compare manual text entry using a smartphone, verbal text entry using a smartphone, and verbal text entry using Glass when engaged in a short answer texting task on a simulated three-lane freeway. They found that while all texting conditions negatively impacted driving performance, Glass did so the least. Similarly, Sawyer et al. (2014) evaluated the use of Glass’s voice-to-text versus smartphone manual entry while driving and performing a secondary arithmetic task. They found Glass users were equally impaired during brake events but exhibited superior performance in terms of subsequent recovery. Finally, Beckers et al. (2014) found similar results with a different in-vehicle secondary task: GPS destination entry. The study showed how using the voice input functionality with either Glass or a smartphone to enter an address resulted in significantly smaller driving performance decrements than did manual input.
The current study was designed to compare the effects of voice-to-text input (vs. manual input) and head-up display (vs. head-down display) on performance in a task set that included navigating a driving simulation while reading and responding to short, semi–open-ended text messages in a manner that reflected each participant’s natural response tendencies. Participants completed driving scenarios under four task conditions: (1) a baseline (no-texting) condition and conditions that included driving plus a secondary texting task using (2) a smartphone with manual input, (3) a smartphone with voice-to-text input, and (4) Google Glass with voice-to-text input. In addition to texting task measures, driving performance was assessed according to common metrics associated with driving safety, including the root mean square (RMS) of average absolute steering rate and standard deviation of lane position (SDLP) (Angell et al., 2006; Menhour, Lechner, & Charara, 2009; Regan, Lee, & Young, 2008). Additionally, the impact of texting on the orientation of visual attention was inferred via the mean following distance and brake reaction times to a lead “pace car” vehicle (Horrey et al., 2006; Hurts et al., 2011) and via video-based analysis of eyes-off-road glance durations.
As in previous studies, the addition of a secondary texting task was expected to negatively impact driving performance (e.g., Drews et al., 2009; Filtness et al., 2013; Lyngsie et al., 2013; Tsimhoni & Green, 2001; Yager, 2013). This study also used aspects of driving safety, such as the number of eyes-off-road glances exceeding a critical duration (1.6 seconds; Horrey & Wickens, 2007) and performance on the texting task, as measures of dual-task performance with the intent that the results can be generalized to more practical in-vehicle text entry tasks. Driving while texting manually on a smartphone was expected to produce the largest performance decrements as this imposes the largest amount of conflicting resource demands with driving (Figure 1). The manual texting condition was therefore expected to be associated with the highest RMS of average absolute steering rates, highest SDLP, slowest reaction times to pace car events, largest following distances, longest texting reaction times, and highest count of safety-critical eyes-off-road glances. Driving and texting performance were expected to improve in at least some of these measures with the addition of voice-to-text input on the smartphone as this reduces the manual and visual demands associated with entering and verifying input text. Finally, the addition of the Google Glass’s HUD functionality was expected to further improve driving and texting performance by reducing structural interference for visual resources because the device allows messages to be read close to the same field of view as the driving scene.
By comparing driving and texting performance with different texting devices, this study attempts to distinguish the benefits of voice-to-text input from those of combined voice-to-text input and HUD. This study adds to a growing body of knowledge regarding wearable technologies, mobile device usage in vehicles, driver performance and safety, human information processing theory, and multitasking theory. Additionally, these findings can be used to inform policymakers, app developers, and the general public about the implications of interacting with Glass and similar technologies while driving.
Method
Data collection and analysis activities were completed for 24 participants (15 men and 9 women) aged 20 to 32 years (men: M = 24.5, SD = 3.11; women: M = 23.8, SD = 1.92) from Texas A&M University. This research complied with the American Psychological Association Code of Ethics and was approved by the Institutional Review Board at Texas A&M University. All participants reported normal or corrected-to-normal vision and familiarity with smartphone texting and had a valid driver’s license.
Participants completed a primary driving task in all four experimental conditions and a secondary texting task in three of those conditions. The driving scenarios were constructed in STISIM Drive, a medium-fidelity, stationary desktop driving simulator displayed on a 30-inch screen. Drivers used a Logitech G27 force-feedback steering wheel and floor-mounted pedals to control the vehicle (see Figure 2).

Experimental setup in the Glass test condition. In conditions involving smartphone interaction, the device was placed on the table next to the steering wheel. The picture in the right corner is a still shot from the video recordings collected and used for coding the eye movement data.
After signing an informed consent form and completing a background questionnaire, participants received a short training session with the driving simulator. Simulator training involved completing a short scenario with a pace car that was repeated until it was both satisfactorily completed (i.e., without observing collisions or other unsafe behaviors), and each participant stated that they were comfortable driving in the simulation environment. All participants were able to demonstrate proficiency in the driving task. Participants then completed four device test conditions, the order of which was completely counterbalanced, with each participant performing a unique permutation of the test conditions (i.e., 24 total device permutations were used): (1) baseline (driving only) and driving plus (2) reading texts on a smartphone and responding via the smartphone’s touchscreen keyboard, (3) reading texts on a smartphone and responding via the smartphone’s voice-to-text input (no manual input was permitted), and (4) reading and listening to texts with Google Glass and responding via Glass’s voice-to-text input. Prior to each texting condition, participants were trained on how to use and tested for proficiency in use of the respective texting device. As nearly all participants had never used Glass, experimenters aided participants in physically adjusting the device so that the prism was properly positioned within a “sharable” field of view with the roadway. Then each participant completed a short tutorial, which included learning how to navigate the Glass interface and practicing texting. Participants then repeated the simulator training scenario while receiving and sending practice text messages with Glass; the scenario was repeated until participants were able to correctly send two successive text messages while driving. All participants were able to demonstrate proficiency in using all the required texting methods. Prior to every scenario, participants were instructed that their first priority was to drive safely and that their second priority was to answer the texts in a timely manner but understood that their reaction times for texting were being recorded. After driving all four scenarios, participants completed a post-experiment questionnaire, which included questions about their experiences using each texting device. The experiment lasted approximately one hour.
Information collected in the background questionnaire was used to categorize participants into “experience” levels in driving, texting, and multitasking contexts. The levels ranged from 1 to 5, with 1 being the least experienced and 5 being the most experienced, and each experience rating was used as a covariate.
Driving Task
The driving task involved driving a scenario that was 6.3 miles in length, took approximately 7 minutes to complete, and spanned both urban and rural driving environments. Participants were instructed to drive safely and near posted speed limits. Each of the four device scenarios (one for each test condition) included three stages that were presented in a randomized order: (1) a winding mountain road (mountain-road), (2) a city highway with interchanges (ramp-highway), and (3) a town square with pedestrians and a sharp left-hand turn (town-square). Each stage involved varying densities of vehicle and pedestrian traffic and standard traffic control devices, such as speed limit signs, stop signs, and traffic lights. A “pace car” led the driver’s vehicle throughout each scenario, and participants were instructed to follow this vehicle at a comfortable distance. The pace car periodically braked, with roughly half of brake events occurring at relatively unpredictable times (i.e., when no other roadway events would have suggested a braking response) and the other half occurring at contextually predictable locations (e.g., when approaching steep curves or a sharp turn). All of the brake event data were considered in a single data set.
Prior to each scenario, drivers were instructed to drive safely and obey traffic rules as their highest priority task. Data were sampled from the driving simulator for every 1 foot driven in the scenario, which resulted in an average sampling rate of approximately 75 Hz. Dependent measures are described in Table 1.
Dependent Measures of Driving Performance (as Generated by STISIM Software)
Texting Task
The secondary texting task required participants to read and respond to incoming text messages. The order of these device test conditions (baseline, touch_keyboard, voice-to-text, and Glass) was completely counterbalanced among participants, with each participant performing 1 of 24 unique permutations of the four conditions. Participants used their own smartphones for the touch_keyboard and voice-to-text conditions, except for four participants whose phones did not contain the voice-to-text feature. These participants instead used an experimenter-provided phone with the same operating system as their respective phones. In total, 18 participants used Android phones, and 6 participants used Apple phones.
Participants were trained with each device prior to starting the corresponding scenario and demonstrated proficiency in a baseline texting task. Modeled after Drews et al. (2009), this baseline texting task involved starting on the smartphone home screen, navigating to the texting interface, and entering and sending the message, “The quick brown fox jumps over the lazy dog.” Prior to the Glass condition, participants were trained on how to access received text messages (i.e., by tilting one’s head up or tapping the side of the glasses frame), how to read messages visually or with Glass’s read-aloud functionality, and how to compose and send responses using Glass’s voice-to-text input. Participants were allowed to use Glass’s read-aloud function to listen to incoming messages but were encouraged and tended to use Glass’s visual display either to quickly read or visually verify displayed content. This visual verification behavior was confirmed by reviewing video recordings of participant’s eye movements after receiving a text message. All but 4 of the 24 participants showed evidence of visually sampling incoming messages and entered text. The 4 exceptions presented difficulties in video-coding glances when using Glass due to eye characteristics or to Glass obstructing the eyes, and similar visual verification behavior was assumed because none were associated with outlier data for any dependent measure.
The texting task consisted of reading and responding to six text messages sent by the experimenters, two during each of the three stages of the scenario, using the condition’s assigned texting method. Messages were delivered at predetermined locations that were designed to impose higher workload (e.g., merging onto a highway, approaching an intersection, completing a turn) and at intervals that allowed for at least 45 seconds for participants to respond between messages. Across the three texting conditions, participants received 18 messages that were selected and randomly ordered from a set of 20 prewritten questions. Each question was designed to be of roughly equivalent difficulty for the participant population (refined through pilot testing), involved reading at least three lines of text (~50 characters), and required responses of several words. See examples of these text message questions in Table 2.
Example Text Messages Used in This Study and Common Expected Responses
Participants were told to respond as they naturally would in a texting conversation with a familiar party, maintain a consistent response style throughout the experiment, and address each message completely in their response (i.e., participants could not send multiple texts to address a single question). In order to determine if requiring clarity and accuracy in messages impacted voice input versus manual texting methods, half of the participants (N = 12) were instructed to correct typing and transcription errors in entered text until they felt the response was satisfactory while the other half were to send responses without correction, regardless of whether the text was entered as intended. Dependent measures of texting performance included texting reaction times, defined as the time from when the device announced the arrival of a message to when the participant submitted a response, and accuracy of response content. Since response content accuracy rates were all near 100%, ultimately this measure was not analyzed.
Video recordings of participant eye movement during all texting scenarios were collected, and glances away from the roadway toward the texting device were manually coded by counting frames in QuickTime. The frame count was then used to tally the number of glances away from the road of 1.6 seconds or greater during each stage. This is considered a critical eyes-off-road glance duration linked with impaired vehicle control and increased crash risk (Horrey & Wickens, 2007). As only one experimenter coded each participant’s eye movements, these data were not formally tested for interrater reliability; however, an alternative testing procedure was used to determine if “coder” impacted the model.
Results
All data were analyzed using SAS 9.3 with a significance level of α = 0.05. The model was a 2 (accuracy: adjust text, do not adjust text) × 4 (device: baseline, touch_keyboard, voice-to-text, Glass) × 3 (stages: mountain-road, ramp-highway, town-square) design. Accuracy of participants’ texting response was a between-subjects variable; the four device conditions and three stages were within-subjects variables. Each of the participants performed the device conditions in a different order, with 12 of the 24 total unique device sequences assigned to subjects in each of the two accuracy conditions. Driving performance, texting response, and eye movement metrics were calculated for each participant in each stage.
Repeated-measures ANOVAs for the driving performance measures were performed using the Proc Mixed function (reduced maximum likelihood [REML] estimates) with an unstructured covariance matrix and using a Kenward-Roger degrees of freedom approximation. The Proc Mixed function was chosen because it is capable of appropriately adjusting for repeated-measures designs and because it compensates for missing data values and violations of sphericity without imposing assumptions that make the test statistics overly conservative. Bonferroni post hoc tests were used to determine differences between means using an α = 0.05. Post hoc means and standard deviations are reported using Cousineau’s adjustment so as to better reflect the results of statistical tests (Cousineau, 2005). Cousineau’s adjustment removed the between-subjects variability from the within-subjects factors by subtracting out the subject average and adding the grand mean, thus normalizing those values and making it easier to compare treatment effects across subjects (Cousineau, 2005). Significant experience level covariates (i.e., driving, texting, and multitasking experience) coded from the background questionnaire were noted, but the data sets were too small and the subgroups formed by the covariates within them too unbalanced to further test their impact on the statistical model. As each participant completed the four device test conditions in a unique order, the model was unable to test for the impact of these sequences. Subjective rankings from the post-experiment survey were analyzed using a nonparametric Friedman test. A summary of significant values across all variables can be found in Table 3.
Normalized Means and Standard Deviations for Root Mean Square (RMS) of Average Absolute Steering Rate (deg/s), Standard Deviation in Lane Position (Feet), Mean Follow Distance (Feet), Driving Reaction Time (Seconds), Texting Response Time (Seconds), and the Number of Glances Greater Than 1.6 Seconds (Count)
Note. Only significant categories are reported with precedence given to device as this was the primary experimental variable in question.
As Proc Mixed uses REML estimates, this procedure does not produce sums of squares values, limiting the user’s ability to directly compute effect size (SAS Institute, 2005). Due to this limitation, an ad hoc method was used that estimated the sums of squares from the model by computing them backwards from the F value. These estimates were then used to compute eta-square (η2) values, which can then be used to compare the relative impact of each experimental variable on the models within this paper (SAS Institute, 2005).
Driving Performance
Driver control metrics.
RMS of average absolute steering rate, which indicates the speed of steering activity and is higher when drivers have more erratic corrections, was significantly affected by the interaction of accuracy and device, F(3, 20) = 3.66, p = .030, η2 = .014, and the interaction of device and stage, F(6, 17) = 8.32, p < .001, η2 = .063, with the accuracy F(1, 21.4) = 4.91, p < .038, η2 = .006; device, F(3, 20) = 30.98, p < .001, η2 = .118; and stage, F(3, 21) = 226.93, p < .001, η2 = .577, variables also being independently significant. The covariate multitasking experience, F(1, 21) = 5.61, p < .028, η2 = .007, was also significant. No other comparisons reached significance, with p values ranging from .142 to .279. Post hoc comparisons were only performed for the significant interaction effects so as to avoid confounding in the results (see Figure 3).

Normalized means of root mean square of average absolute steering rate (deg/s) for each device by stage. Error bars represent 95% confidence intervals.
When not required to correct the accuracy of the text, mean RMS of average absolute steering rate was significantly greater in both the touch_keyboard (M = 6.25 deg/s, SD = 2.88 deg/s) and voice-to-text (M = 5.36 deg/s, SD = 2.26 deg/s) conditions than in the baseline (M = 3.58 deg/s, SD = 1.98 deg/s) and Glass (M = 3.90 deg/s, SD = 1.98 deg/s) conditions, which did not significantly differ from each other. When correct text accuracy was required, a similar pattern is observed except that Glass conditions did not statistically differ from any other condition, with steering rates still significantly larger in the touch_keyboard (M = 5.36 deg/s, SD = 2.19 deg/s) and voice-to-text (M = 4.95 deg/s, SD = 1.93 deg/s) conditions than in baseline (M = 4.00 deg/s, SD = 1.65 deg/s) conditions.
The device and stage interaction displays a similar pattern for devices across stages, with varying levels of significance. The mountain-road stage, characterized by winding roads, showed higher RMS of average absolute steering rate in the touch_keyboard (M = 8.29 deg/s, SD = 3.08 deg/s) and voice-to-text (M = 7.10 deg/s, SD = 1.30 deg/s) conditions compared to the baseline condition (M = 5.18 deg/s, SD = 0.78 deg/s). Steering rates were also greater in touch_keyboard conditions compared to Glass conditions (M = 5.92 deg/s, SD = 1.16 deg/s). The ramp-highway stage, characterized by high speeds and lane changes, showed similar patterns in the results, with touch_keyboard (M = 2.99 deg/s, SD = 0.89 deg/s) and voice-to-text (M = 2.81 deg/s, SD = 1.02 deg/s) showing significantly higher steering rates than the baseline condition (M = 1.55 deg/s, SD = 0.97 deg/s). In the town-square stage, characterized by lower speeds but more visual distractions, the touch_keyboard condition again showed the highest steering rate (M = 6.14 deg/s, SD = 1.17 deg/s) but here only significantly differed from the baseline condition (M = 4.63 deg/s, SD = 0.82 deg/s).
Standard deviation of lane position was significantly affected by stage, F(2, 20.7) = 357.71, p < .001, η2 = .361. The covariate driving experience, F(1, 21) = 11.61, p < .003, η2 = .004, was also significant. No other factors reached significance, with p values ranging from .196 to .874. Post hoc comparisons show significantly higher SDLP measures in the town-square stage (M = 6.49 ft, SD = 3.23 ft) than in either the ramp-highway (M = 5.38 ft, SD = 2.85 ft) or mountain-road (M = 4.13 ft, SD = 3.36 ft) stages.
Pace car metrics.
Mean following distance behind the pace car, with longer distances associated with driving behavior under higher workload, showed a significant main effect due to stage, F(2, 21) = 21.52, p < .001, η2 = .137. The covariate driving experience, F(1, 21) = 5.16, p < .034, η2 = .016, was also significant. No other variables reached significance, with p values ranging from .068 to .646. Post hoc comparisons indicated that mean follow distance in the ramp-highway (M = 251.49 ft, SD = 82.75 ft) stage was significantly greater than in the mountain-road (M = 198.21 ft, SD = 68.71 ft) stage (see Figure 4, left).

Normalized means of mean follow distance (feet) for each stage (left) and of driving reaction time for each accuracy requirement by device (seconds) (right). Error bars represent 95% confidence intervals.
Reaction time to pace car events was calculated as the time between the onset of the pace car’s brake lights and when participants lifted their foot from the accelerator (with faster times indicating better roadway vigilance). Data were missing when (a) the participant did not brake because they were far enough behind the braking pace car that it was unnecessary and when (b) the participant was already actively applying the brake at the moment the pace car braked. The latter case primarily occurred when roadway elements, such as curves, led some participants to initiate an early brake response. Due to this missing data, a simplified version of the Proc Mixed repeated-measures ANOVA was used that employed a blocking factor instead of the standard repeated statement, thus resulting in differences in the degrees of freedom in the model.
Reaction time to pace car events was significantly impacted by the interaction of accuracy and device, F(3, 176) = 3.93, p = .010, η2 = .018. No other variables were significant, with p values ranging from .096 to .454. Post hoc comparisons indicated that when correcting responses for accuracy, participants reacted significantly slower to brake events in the touch_keyboard condition (M = 1.41 s, SD = 0.92 s) than in both the baseline (M = 0.88 s, SD = 0.37 s) and voice-to-text (M = 0.93 s, SD = 0.61 s) conditions (see Figure 4, right).
Texting Response Time
Texting response times, which represent time spent with partial attention devoted to the secondary task, with worse texting performance indicated by longer times, were significantly impacted by device, F(2, 21) = 19.02, p < .001, η2 = .148. No other variables were significant, with p values ranging from .160 to .998. Post hoc comparisons suggest that response times were significantly longer in both the touch_keyboard (M = 27.11 s, SD = 8.64 s) and Glass (M = 23.39 s, SD = 4.44 s) conditions than in the voice-to-text (M = 19.92 s, SD = 5.58 s) condition (see Figure 5, right). Response times were also significantly longer in the touch_keyboard condition than in the Glass condition.

Normalized means of texting response time (seconds) (left) and for the number of eyes-off-road glances greater than 1.6 seconds for each device (right). Error bars represent 95% confidence intervals.
Eyes-Off-Road Time
Glances away from the road that were greater than 1.6 seconds were tallied and compared. This threshold was defined by Horrey and Wickens (2007) as indicating increased crash risk (with more glances over 1.6 seconds indicating heightened risk). This threshold, however, differs from the criteria for long glances used in the NHTSA Phase 1 Voluntary Guidelines and is one of three glance metrics suggested by those guidelines (NHTSA, 2013). In lieu of formal methods to test interrater reliability, the impact of the experimenter who coded each participant’s glances was tested as a fixed factor and was found to be not significant.
The number of eyes-off-road glances was significantly impacted by the main effects of accuracy, F(1, 26.1) = 6.07, p < .021, η2 = .026; device, F(2, 21.3) = 15.98, p < .001, η2 = .097; and stage, F(2, 21.8) = 6.34, p < .007, η2 = .033. No other effects were significant, with p values ranging from .157 to .907. Post hoc comparisons suggest the number of eyes-off-road glances was greater for those who did not have the texting accuracy requirement (M = 1.76, SD = 1.55) than those who did have the texting accuracy requirement (M = 1.13, SD = 1.46). The number of eyes-off-road glances was also significantly greater in the touch_keyboard (M = 2.58, SD = 1.49) condition than both the voice-to-text (M = 1.01, SD = 0.74) and Glass (M = 0.79, SD = 0.94) conditions, which did not significantly differ from each other (see Figure 5, right). Additionally, the number of eyes-off-road glances was greater in the town-square (M = 1.67, SD = 1.47) stage than in the mountain-road (M = 1.19, SD = 1.14) stage.
Subjective Measures
All of the subjective measures differed significantly among texting methods (see Table 4). The touch_keyboard condition was rated as significantly more difficult and involving more dual-task interference than the voice-to-text and Glass conditions. The voice-to-text condition was also rated significantly worse than the Glass condition for both metrics. Glass’s overall rank was significantly more highly preferred than the touch_keyboard and voice-to-text devices, which did not significantly differ.
Summary of Analyses of Subjective Ratings and Rankings
Discussion
Over 3,000 people were killed and 431,000 people injured in distracted driving crashes in 2014, and interactions with in-vehicle technologies play a significant role in these crashes, with texting posing considerable concern as it can induce all three categories of driver distraction: visual, manual, and cognitive (NHTSA, 2013, 2014). This study used a texting task to determine how interface characteristics and differing demands for visual and manual resources impact multitasking in the vehicle.
With regard to input characteristics, this study’s results are in accordance with prior studies involving general in-vehicle tasks (Horrey & Wickens, 2007; Tsimhoni & Green, 2001) and studies involving interactions with Glass (Beckers et al., 2014; He, Ellis, et al., 2015; Sawyer et al., 2014) showing benefits to voice input (voice-to-text and Glass conditions) over manual input (touch_keyboard condition) methods in both driving and texting metrics. As expected, the baseline condition (no texting, driving only) showed the best driving performance, and the performance decrement associated with adding a secondary texting task depended on the texting device. The touch_keyboard condition was consistently associated with the worst performance in both driving metrics and texting metrics and was subjectively rated as inducing the highest workload. Conditions involving voice input (voice-to-text and Glass) generally showed smaller performance decrements than did the touch_keyboard condition. This relative improvement in driving safety afforded through use of voice input was best illustrated via the number of safety-critical eyes-off-the-road glances: The touch_keyboard was associated with nearly three times the number of glances as in the other texting conditions.
The performance benefits of voice input over manual input can be partially explained via multiple resource theory (Wickens, 1980, 2002), as using speech input for a secondary task reduces competition for the manual resources also used in driving, thus mitigating this potential source of structural interference. The time required to verbally enter messages is faster than manual entry and involves less focal visual reorientation than during text entry and verification, resulting in fewer safety-critical glances away from the road being observed.
The impact of display characteristics can be most directly compared between the Glass and voice-to-text conditions, which each involved voice input but differed in how text was displayed (e.g., head-down on a smartphone or HUD on Glass). In general, similar measures of performance and safety were found under the two conditions, making any advantages (or disadvantages) offered by Glass’s HUD functionality difficult to observe. RMS of average absolute steering rate was generally smaller (i.e., driving was less erratic) with Glass than with voice-to-text, but the difference only reached significance when accuracy in text message input was not required. This finding corresponds with findings from prior studies investigating the effects of Glass’s HUD functionality in the driving environment (e.g., Beckers et al., 2014; He, Ellis, et al., 2015; Sawyer et al., 2014) and suggests that while HUDs may improve performance on the driving task, the higher workload imposed by reviewing and revising errant text entries (as when accuracy was required) may limit the degree of improvement.
Despite these similarities in performance between the voice-to-text and Glass conditions, the noted statistical differences in driving performance and the significantly better subjective ratings and higher user preference rankings for Glass over voice-to-text may reflect an overall reduction in mental workload when using Glass. While not reaching significance when comparing the Glass and voice-to-text conditions, overall the number of safety-critical eyes-off-road glances when texting with Glass was fewer than with a smartphone, which may also reflect a reduction in the visual reorientation required to switch between the driving and texting tasks when the device is closer to the field of view of the roadway. By only requiring a change in eye gaze direction to view text on the device and not a more disruptive change in head or body posture, HUDs may better support in-vehicle tasks that use a visual display and thus potentially reduce structural interference by minimizing the size and number of overt focal reorientations between the texting device and the roadway. Prior studies concerned with HUD use, however, note that affording a more “shareable” visual field has the potential to result in attentional detriments, such as visual and cognitive capture (Tufano, 1997; Yantis & Egeth, 1999). Visual cues on a HUD that are salient or task-relevant, such as the onset of an icon indicating that a text message has arrived, may be more likely to lead to a reorientation of visual attention away from the roadway, thus cancelling out their potential benefits.
Of all the objective dependent measures, the highest effect sizes for driving performance were found with the variable stage, which reflects the diverse types and amounts of loading imposed by the three driving contexts. These effects may have weakened the impact of the main independent variable device but are nonetheless noteworthy as they indicate the sensitivity of different workload measures in different contexts. All factors considered, the mountain-road stage was assumed to be the most difficult, and the ramp-highway and town-square stages were assumed to be approximately equal but more sensitive to different driving measures. All driving measures supported these assumptions. The SDLP and RMS of average absolute steering rate results showed the mountain-road stage to be the most difficult. While SDLP has a complex relationship with driver workload, it often is smaller when high workloads cause drivers to more heavily rely on automatized processes for perceptual-motor control of highly practiced tasks, such as lane keeping, resulting in smaller errors (e.g., Engström, 2011; He, Chaparro, et al., 2014). RMS of average absolute steering rate further showed difference in the interaction between stage and device, demonstrating an increased sensitivity to device characteristics within higher workload driving contexts. Notably, texting with Glass impacted measures of RMS of average absolute steering rate the least across all of the stages, never differing from the baseline. In contrast, the touch_keyboard and, to a lesser extent, the voice-to-text conditions led to higher values (more erratic steering) when compared to the baseline condition, often reaching significance; this was particularly pronounced in the mountain-road stage.
For texting and eye movement measures, the highest effect sizes were found with the variable device, which may provide the strongest evidence for the theoretical inferences tested in this study. When using Glass, visual attention reorientation between the display and road were expected to be minimal. Moreover, the number of safety-critical eyes-off-road glances while using Glass for messaging was significantly reduced compared to when using the touch_keyboard. Using the smartphone voice-to-text method also showed a significant safety improvement over using the touch_keyboard but did not significantly differ from Glass, suggesting that the larger benefit Glass provides could be attributable to its voice input functionality.
Despite the fact that multitasking with Glass was less detrimental than multitasking with the other devices, it should be noted that performance in the Glass condition was not equivalent to performance in the baseline (no-texting) condition. This finding is consistent with other studies whose results indicate that the best overall driving performance occurs in driving-only, no-texting conditions (e.g., Drews et al., 2009; Filtness et al., 2013; Lyngsie et al., 2013; Tsimhoni & Green, 2001; Yager, 2013).
The subjective metrics indicate that participants perceived texting with Glass as easier and as interfering less with driving than texting with the other devices. Glass received the highest overall preference rating for supporting the two concurrent tasks among the devices tested. These positive rankings occurred despite participants’ lack of familiarity with the Glass; however, some of the positive reviews for Glass were likely due to the novelty and perceived “coolness” of the device. Glass’s positive assessments were also despite operational problems, such as the device’s tendency to overheat while in use, which slowed its processing time and sometimes required experimenters to stop and restart a scenario after allowing it to cool.
Qualitative observations noted while reviewing the video when coding the eye glance data suggested that participants exhibited several differences in behaviors due to their level of “trust” in the Glass technology. Some participants did not appear to trust the functional features of Glass and thus repeatedly visually sampled the display, re-reading incoming texts and verifying their location within the interface. Other participants exhibited higher trust in the system and visually verified input and output infrequently, most often glancing to confirm the content of an outgoing message. With increasing familiarity using advanced interfaces such as Glass, general user trust levels will likely increase (e.g., Riley, 1996). As display format and quality are two factors known to affect trust in technologies (Lee & See, 2004), future technological developments should consider how both task-related and trust-related factors may impact the frequency and duration of visual reorientations away from the roadway and hence driver safety.
The experimental nature of the Glass technology was a limitation in the current study. As Glass was not commercially available at the time of this study, nearly all participants were unfamiliar with the technology, and some design issues, such as overheating problems, will likely be resolved in future head-mounted wearable technologies. Participants were also primarily from a younger demographic, which affects both driving behavior and familiarity with texting tasks and devices; hence, future work should evaluate performance over broader age range and vary technological experience.
The fidelity of the driving simulator and design of the experimental scenarios were also study limitations. The driving simulator was presented on one (large) monitor and did not include side or over-the-shoulder views; the simulator was also unable to provide the vestibular feedback that results from real-world vehicle movement. The scenarios were designed to emphasize realism, which resulted in a reduction in experimental control of the pace car’s behavior; this led to some larger than expected variances and missing data, which was most problematic in computing the response times to pace car braking events. Additionally, texts were sent during particularly high-workload contexts during the scenarios, which is not necessarily representative of real-world texting, and participants may have responded to texts more quickly than in the real world because their safety was not truly at risk.
Future research on this topic might address the individual versus interacting effects of voice input and head-up display components by testing each functionality independently and then together. Additionally, while the results of this study indirectly suggest potential negative impacts of texting while driving on cognitive interference, the use of a basic texting task that was independent from the driving task limited the ability to draw further inferences about this relationship.
Conclusion
Both the voice-to-text and HUD characteristics evaluated in this study benefitted multitasking performance presumably by reducing the competition for resources in the driving environment. While the combined benefits afforded by HUD + voice input resulted in the smallest driving performance decrements, the addition of voice-to-text input alone resulted in the most significant improvements in both driving and texting metrics, including a reduction in the number of extended eyes-off-road glances.
The findings from this study can inform decisions for developers, users, and policymakers regarding how technology interface characteristics may support a wide range of interactions with in-vehicle information systems (e.g., GPS, comfort control systems) as well as their effects on driving safety. The results provide designers of in-vehicle devices with information about the extent to which their choice of interface characteristics will impact driver performance and suggest that designers should consider implementing voice commands and minimize the driver’s need for eyes-off-road time when interacting with such devices, with HUDs being a potential avenue for further reducing the amount of visual reorientation required to perform a secondary task. All parties should strongly encourage drivers to minimize their interaction with additional technologies they bring into vehicles while driving.
Future research should investigate the individual and combined impacts of voice input and HUD with broader participant age groups and technological familiarity and should include secondary tasks that are more objectively beneficial to the driver.
Key Points
While the risk of distracted driving crashes substantially increases when conducting secondary tasks with in-vehicle technologies, voice input and head-up display (HUD) features have the potential to individually and jointly produce fewer driving performance decrements due to these secondary tasks, with the greatest benefits seen through adding voice input to multitasking task sets.
Google Glass (HUD + voice input) supported the smallest decrement to primary driving task performance, but HUD use provided minimal additional benefits over solely using voice-to-text input in terms of either texting performance or the number of eyes-off-road glances.
This study’s findings can inform developers, users, and policymakers on decisions regarding the potentially beneficial characteristics of new technologies, particularly wearables, as well as their impact on driving safety.
Footnotes
Acknowledgements
This research was supported in part by the National Science Foundation Graduate Research Fellowship Program (NSF GRFP) (Grant No. DGE-1252521). The authors would also like to thank faculty and students at Texas A&M University for their participation in this research and anonymous reviewers for providing valuable feedback that helped improve this article.
Kathryn G. Tippey is a postdoctoral research fellow at the Center for Research and Innovation in Systems Safety under the Department of Anesthesiology at Vanderbilt University Medical Center. She earned her PhD in industrial and systems engineering from Texas A&M University in 2016.
Elayaraj Sivaraj is a manufacturing engineer at Tesla Motoes in Fremont, California. He completed his master’s in industrial and systems engineering at Texas A&M University, where he was a member of the Human Factors & Cognitive Systems Laboratory. He obtained an undergraduate degree in mechatronics engineering from SRM University, India.
Thomas K. Ferris is an assistant professor of industrial and systems engineering at Texas A&M University, where he is the director of the Human Factors & Cognitive Systems Laboratory. He earned his PhD in industrial and operations engineering from the University of Michigan in 2010.
