Abstract
Objective
This paper proposes an objective method to measure and identify trust-change directions during takeover transitions (TTs) in conditionally automated vehicles (AVs).
Background
Takeover requests (TORs) will be recurring events in conditionally automated driving that could undermine trust, and then lead to inappropriate reliance on conditionally AVs, such as misuse and disuse.
Method
34 drivers engaged in the non-driving-related task were involved in a sequence of takeover events in a driving simulator. The relationships and effects between drivers’ physiological responses, takeover-related factors, and trust-change directions during TTs were explored by the combination of an unsupervised learning algorithm and statistical analyses. Furthermore, different typical machine learning methods were applied to establish recognition models of trust-change directions during TTs based on takeover-related factors and physiological parameters.
Result
Combining the change values in the subjective trust rating and monitoring behavior before and after takeover can reliably measure trust-change directions during TTs. The statistical analysis results showed that physiological parameters (i.e., skin conductance and heart rate) during TTs are negatively linked with the trust-change directions. And drivers were more likely to increase trust during TTs when they were in longer TOR lead time, with more takeover frequencies, and dealing with the stationary vehicle scenario. More importantly, the F1-score of the random forest (RF) model is nearly 77.3%.
Conclusion
The features investigated and the RF model developed can identify trust-change directions during TTs accurately.
Application
Those findings can provide additional support for developing trust monitoring systems to mitigate both drivers’ overtrust and undertrust in conditionally AVs.
Keywords
Introduction
Automated driving technology has the potential to increase traffic efficiency (Teoh and Kidd, 2017), safety, and fuel efficiency (Fagnant and Kockelman, 2015). However, we still have a long way toward fully autonomous driving being widespread. Besides, vehicles with conditionally automated driving systems such as traffic jam pilots have been developed. According to the vehicle automation level by the Society of Automotive Engineers (SAE) International, the role of drivers may change from manual operators to disengaged supervisors in the SAE Level 3 automated driving (SAE, 2018). Several new challenges related to human factors have been raised, such as trust and acceptance of users, especially for situations in which drivers are asked to resume control of the vehicle due to system limitations or failures (Abe et al., 2018). Trust is one of the important challenges that impacts the public acceptance and use of AVs (Choi and Ji, 2015).
Trust in automation is defined as “the attitude that an agent will help achieve an individual’s goals in a situation characterized by uncertainty and vulnerability” (Lee and See, 2004). To systematically understand factors that influence trust in automation, Hoff and Bashir (2015) proposed a three-layered trust model, including dispositional, situational, and learned trust. Dispositional trust denotes an individual’s overall tendency to trust automation, independent of context or specific systems. Situational trust describes trust as a combination of the external environment and operator states. Learned trust is established by the operator’s evaluations of a system based on historical experience or the current interaction. The takeover transition, as the main component of the interaction of the driver with the SAE level 3 automated driving system, has a potential impact on situational trust and learned trust (Körber et al., 2018c; Jin et al., 2020). When trust in AVs decreases during the TTs, drivers may avoid using the system and return to manual control, thereby reducing the safety benefits and other goals of automated driving pursued (Seet et al., 2020). Thus, it is essential for exploring the effects and relationships between the trust-change direction and takeover-related factors or physiological parameters, and to establish a model identifying the changes in trust during TTs, which can provide additional support for the designing of the vehicle-driver interface and trust monitoring system to reduce the trust erosion due to TORs.
Factors Influencing Driver Trust
Previous studies have examined the factors that affect the trust of drivers, such as systems transparency, training, driving style, and perceived risk, (Hergeth et al., 2017; Ekman et al., 2019; Li et al., 2019; Zhang et al., 2019; Kraus et al., 2020; Azevedo-Sa et al., 2021) and provided some interesting findings. For instance, the “defensive” automated driving style was perceived as more trustworthy than the “aggressive” driving style (Ekman et al., 2019). Li et al. (2019) and Zhang et al. (2019) found a significant negative relationship between perceived safety risk and driver trust. Besides, a few researchers paid attention to the effects of takeover-related factors on trust in AVs, such as the TOR lead time and takeover frequencies (i.e., the number of takeovers).
The TOR lead time is defined as the time available between one TOR and the system limit (Gold et al., 2018b). It is considered to affect the system-perception reliability and drivers’ takeover safety and then impact trust. For example, untimely TOR lead time represents a risky takeover situation, which reflects the unreliability of the system and then can reduce trust (Lee and See, 2004). Jin et al. (2020) examined the effect of the TOR lead time (i.e., 7 s and 10 s) on driver trust, showing that when drivers experienced the shorter TOR lead time, they had a lower subjective trust rating. Besides, driver trust would be related to not only the TOR lead time but also the takeover frequencies. Hergeth et al. (2017) and Körber et al. (2018c) suggested that the mean of the pre-and post-takeover trust rating increased with takeover frequencies.
Trust Measurements in Automated Vehicles
To understand how trust impacts the acceptance of AVs, reliable metrics are needed to measure trust. Multidimensional questionnaires containing sub-dimensions (e.g., trust, perceived risks) were used to measure driver trust (Jian et al.,2000; Körber, 2018a; Holthausen, 2020). Furthermore, to reduce the frustration of participants facing repeated measurements, other researchers used a single-dimensional questionnaire during the experiment (Hergeth et al., 2016). However, self-reported trust ratings cannot provide continuous measurements, and therefore cannot estimate real-time changes in trust, and also are intrusive, which may be difficult to use in applied settings (Hergeth et al., 2016; Walker et al., 2019). To continuously estimate driver trust in real-time, the objective measurements of trust such as monitoring behaviors, skin conductance, and heart rate have been explored in many studies.
Previous empirical evidence demonstrated that driver trust is closely linked to monitoring behaviors (Hergeth et al., 2016; Körber et al., 2018b; Walker et al., 2019). For example, Hergeth et al. (2016) found a higher subjective trust rating associated with a lower monitoring behavior. Körber et al. (2018b) reported that the trust promotion group induced by the initial information paid less attention to roads or dashboards. Besides, many scholars have explored relatively novel methods to measure trust, such as skin conductance (SC) and heart rate (HR) (Waytz et al., 2014; Khawaji et al.,2015; Morris et al., 2017; Wintersberger et al., 2017; Gupta et al., 2019; Petersen et al., 2019; Perello-March et al., 2021). However, attempts at using these measures to capture trust in automation are still mixed results. For example, drivers showed lower self-reported trust ratings and higher SC in risky driving situations (Morris et al., 2017). Humans have more trust and weakened HR increases during a vehicular accident when they drive an anthropomorphism AV (Waytz et al., 2014). However, Gupta et al. (2019) and Perello-March et al. (2021) could not report any significant effects between trust and the SC. Petersen et al. (2019) also found a disconnection between trust and HR. Furthermore, Kohn et al. (2021) comprehensively reviewed trust measurements in existing empirical works, suggesting that the monitoring behavior is a more established measure than other physiological measures (e.g., the SC and HR), since it can provide a stable correlation of trust and is less disruptive to the original task.
Prediction Models for Trust in Automated Vehicles
To calibrate the levels of driver trust, it is also vital to develop a trust prediction model for AVs. There is some research on the predictive model of driver trust during the non-takeover period in AVs (Azevedo-Sa et al., 2020; Ayoub et al., 2021; Liu et al., 2021). For instance, Azevedo-Sa et al. (2020) established the trust model in AVs based on interactive experiences (e.g., system malfunction types and system usage time) and drivers’ behavior (e.g., non-driving-related tasks performance) by the Kalman filter approach, which can estimate driver trust during the non-takeover process. Liu et al. (2021) proposed a customized trust model by clustering humans based on their trust dynamics to improve trust prediction performance during the non-takeover.
Research Motivations and Main Contributions
Considering that self-reported data might always have limitations due to individual and situational biases, which may not accurately reflect driver trust (McBride et al., 2010; Rosenman et al., 2011; Petersen et al., 2019). Yu et al. (2021) proposed a combination of self-reported trust ratings and system usage frequencies to measure driver trust in advanced driver assistance systems. Drivers, as disengaged supervisors in AVs, will almost always use the system for a short period or longer. It may be difficult to obtain a vailed trust measurement through system usage frequencies and apply it to AVs. In SAE level 3 automation, drivers always have difficulty negotiating TTs safely because they are out of the control loop (Du et al., 2020a), which can undermine trust and in turn affect drivers’ reliance and use of AVs. Thus, we need to find reliable indicators to measure the change values of driver trust during TTs. However, the measurement of trust in AVs still relied on subjective evaluation as the ground truth. To measure trust in AVs accurately, a hybrid method of combining subjective ratings with the objective behavior indicator is introduced into our study. Lee and See (2004) concluded that users’ reliance on automation depends to a certain extent on trust. That is, when trust reduces, users may decrease their reliance on automation and increase the monitoring behavior. Thus, it is reasonable to believe that when drivers reduce trust during TTs, after completing takeover missions and reactivating the autopilot mode, they will pay more attention to the road or the dashboard. Moreover, as noted previously, the monitoring behavior can provide a more stable correlation of trust than the other physiological measures such as the SC and HR. Thus, we will combine the change values of the subjective ratings and the monitoring behavior before and after takeovers to measure the trust-change directions during TTs.
In addition, previous studies have explored the effects of potential factors on driver trust, such as systems transparency, training, TOR lead time, and takeover frequencies. However, the influence of takeover scenario types on driver trust is rarely mentioned. Vogelpohl et al. (2018) found that scenario types can affect the takeover criticality, while perceived risks can influence driver trust (Azevedo-Sa et al., 2021). Thus, we speculate that the takeover scenario may affect the changes in driver trust. Critical or unexpected TORs could be considered rather risky situations (Maule & Hockey, 1993). In these cases, drivers would feel safety risk, which might lead to a sharp drop in trust, inducing the increasement in stress, and then their physiological parameters would change accordingly (Mühl et al., 2020). If driver trust is not improved timely, misoperations might occur due to the distrust of the system, and then increase the risk of accidents.
Besides, existing studies have examined the relationships between physiological responses (e.g., the SC and HR) and trust to predict driver trust in AVs. Nevertheless, as mentioned above, these results are still mixed conclusions that need to be further explored. Also, as far as we knew, the connection between physiological parameters (i.e., SC and HR) and the trust-change directions during TTs is rarely considered because the relationships among the trust-change direction during TTs, psychological parameters, and takeover-related factors could be complex. Surely, exploring their associations can be beneficial to understand changes in driver trust due to altering the driving situation, providing guidance for the development of a trust monitoring system to calibrate trust, and ultimately, decreasing the risk of driving accidents.
Finally, to mitigate the trust miscalibration issues such as undertrust and overtrust, a few researchers established the trust model focused on estimating driver trust during non-takeover processes, which, however, cannot estimate the trust-change direction during TTs. The takeover missions will be recurring events that could undermine the trust and acceptance of drivers, and, therefore, the successful introduction of AVs. It is vital to establish a model identifying the trust-change direction during TTs to foster better driver-vehicle trust. To the best of our knowledge, there is no research on the model of the trust-change direction during TTs. As aforementioned, stress-mediated the SC and HR may also change when driver trust is altered. Besides, the trust-change direction during TTs is closely linked with takeover-related factors (i.e., the TOR lead time and takeover frequencies), and scenario types may affect the changes in driver trust during TTs.
Thus, we will comprehensively examine the relationships and effects among physiological parameters (i.e., the SC and HR), takeover-related factors (i.e., TOR lead time, takeover frequencies and scenario types), and the trust-change direction during TTs. Furthermore, we will establish an identifying model of the trust-change direction during the takeover period using the random forest, based on the SC, HR, and takeover-related factors. The main contributions of this study are summarized in brief as follows: 1) A hybrid method of combining the changes in the subjective rating and objective behavior indicator (i.e., the monitoring behavior) before and after takeovers is proposed to measure the trust-change direction during TTs. 2) The effects and relationships between takeover-related factors, drivers’ physiological responses, and the trust-change direction during TTs are explored and analyzed. 3) A model for identifying the trust-change direction during TTs is developed based on the driver’s multimodal physiological data and takeover-related factors.
Method
Participants
A total number of 34 graduate students participated in the experiment (mean age = 23.8; standard deviation [SD] = 3.9; 3 females and 31 males). Each participant received a cash bonus of 100 yuan. All of participants had a valid driver’s license and normal or corrected-normal vision. Participants had their driver’s licenses for at least 1 year. This study was approved by the Institutional Review Board in the college of mechanical and vehicle engineering at Hunan University.
Apparatus
The experiment was conducted in a static driving simulator. The virtual world was projected on three primary displays (size 27 inches). The simulation of driving scenarios was programmed in Prescan software 8.6.0. The different driving parameters (e.g., ego speed, lead speed, acceleration, braking) were displayed in the graphical user interface (GUI) (Figure 1a). The vehicle was programmed to simulate the SAE L3 automation (SAE, 2018), which can handle the lateral and longitudinal control. The automation could be toggled via the first red button on the right side of the steering wheel (Figure 1b) and was also shut off by steering or braking input. The TOR signal was emitted by the Arduino Uno (Figure 1c). All questionnaires were collected online with two mobile phones. Equipment: A fixed driving simulator, SMI Eye-Tracker, ECG, EAD sensor. (a) GUI. (b) Steering wheel. (c) Arduino Uno. (d) Eye‐Tracker. (e) ECG sensor. (f) EDA sensor.
As shown in Figure 1, the driver was equipped with psychophysiological sensors. Gaze behaviors were recorded using SMI glass with a tracking frequency of 60 Hz (Figure 1d). The SC and electrocardiogram (ECG) were measured with a system (MP150, Biopac Systems Inc, CA) with a sampling rate of 200 Hz (Figure 1e, Figure 1f).
The non-driving-related task (NDRT) in the study was a visual Surrogate Reference Task (SuRT) (ISO 14198, 2012), which was presented on a tablet (Figure 1). In this task, participants were instructed to select, by tapping on the larger circle (diameter 47 pixels) among 49 distractor circles (diameter 40 pixels). The placement of the display needed participants to reallocate their visual focus completely away from the road environment if they wanted to engage in the task. The reason for utilizing the SuRT with manual input was to simulate the hands-off wheel and the eyes-off the road condition.
Experimental Design
Information Material
We used a 2 Procedure and measures of the experiment. (Notes: AD = automated driving; NDRT = non-driving-related task; TOR = takeover request).
The System’s Limitation Information
Scenarios
Three takeover scenarios were designed based on the above system’s limitations. In the stationary vehicle scenario, one TOR was triggered by the broken-down vehicle directly in front of the same lane (Gold et al., 2018a, Figure 3a). For the roadwork scenario, one TOR was triggered by a temporary change in lane lines. An information board with words (WORK ZONE AHEAD), roadwork sign, and turn left ahead sign directly in front of the same lane was used as a static obstacle (Körber et al., 2018c, Figure 3b). In the dynamic truck scenario, the temporary sensor failure of the AV resulted in its longitudinal control failing and triggered one TOR (Gold et al., 2018a, Figure 3c). The slow leading truck (60 km/h) right in front of the same lane was regarded as a dynamic obstacle. Illustrations of three takeover situations: (a) stationary vehicle; (b) roadwork; and (c) dynamic truck.
A leading vehicle with 100 km/h obscured these situations to increase the criticality of the takeover situations. With nearly a TTC of 5 s or 9 s, the leading vehicle suddenly swerved to the left lane, allowing the subjects a view of the obstacles at 4 s or 8 s TTC (Figure 3). When an auditory TOR in form of buzzing (2700 Hz, 74 dB) was emitted at 4 s (44.44 m for the dynamic obstacle, 111.12 m for the static obstacle) or 8 s TTC (88.88 m for the dynamic obstacle, 222.24 m for the static obstacle), requesting drivers manually regain control, the automated driving system was shut down for drivers to conveniently take over control of the vehicle. Each participant experienced six takeover situations altogether, which were formed by the arrangement and combination of two different TOR lead time levels and three takeover scenarios. To minimize the ordering effect, the sequence of the takeover scenarios across participants was counterbalanced via Latin Square.
Non-Driving-Related Task Surrogate Reference Task
In each takeover event, the NDRT appeared every interval 120 seconds, and 2 or 3 times in total, each time lasting 60 seconds. One of the presentations lasting 30 seconds was interrupted by a TOR. In the next 45 seconds, drivers coped with takeover missions and observed the road environment without engaging in NDRT. Immediately after, the autopilot system was reactivated, and the NDRT was presented for another 30 seconds (Figure 2). Besides, to reduce the predictability of takeover situations, the driving time before a TOR (ranging from 2.5 to 6.5 min) was manipulated by performing at most 3 NDRT periods.
The experiment had a gamification aspect to it. Participants competed for monetary bonuses to increase the real-world consequences of the simulated experiment. Participants got two points for completing an NDRT correctly. Ten points were deducted if the driven vehicle left the lane boundary. Besides, a penalty of 50 points was given if a crash with another obstacle occurred in the lane. The scoring mechanism encouraged subjects to rely on automation, and they could focus on engaging the NDRT. After the experiment, the participant with the highest score was rewarded cash bonuses of 300 yuan.
Instructions and Experiment Track
The vehicle drove in the right lane at 100 km/h on a two-lane freeway. The headway was set to 1.5 s. Automated lane changing or overtaking was not implemented. Participants were told to drive in the right lane and were only allowed to change lanes in case of emergency. After passing the obstacle, participants were instructed to come back to the right lane. And then, drivers reactivated the autopilot mode by themselves or were prompted by the experimenter 45 s after the TOR was issued. Participants were asked to safely drive at all times. Furthermore, participants must perform the NDRT whenever the experimenter required the participant to engage in the NDRT.
Questionnaires
In this study, a single-dimensional questionnaire of trust was used to reduce the frustration of the subjects in face of repeated measurements during the trial. In reference to earlier studies, the subjects were asked to slide the slider to rate their trust levels on a scale from 0 to 100 (“On a scale from 0 to 100, how much do you trust the system?”) (Hergeth et al., 2016). A single-item trust rating was both filled at the beginning and end of each NDRT (Figure 2).
To assess the criticality of takeover situations, an 11-point scale proposed by Naujoks et al. (2015) was used. The scales ranged from 0 (imperceptible) to 10 (uncontrollable) with the anchor categories: imperceptible (0), harmless (1–3), unpleasant (4–6), dangerous (7–9), and uncontrollable (10). The criticality ratings of takeover situations were filled out at the end of the NDRT after a TOR.
Procedure
After participants have been welcomed and signed an informed consent form, participants filled out an online questionnaire on demographic data, watched the introductory video, and received introductory information in text form. The eye-tracker was worn and calibrated. Next, two SC electrodes were attached to participants’ distal phalanges to obtain the SC signal (Figure. 4a; Scerbo et al., 1992). Three electrodes were placed on the body in the lead II configuration (Figure. 4b) to record the ECG (Dupre et al., 2005). The participants received several instructions, and the introductory drive started. And then, the participants were invited to exercise the NDRT. If the participants said that they were comfortable with engaging in the NDRT, they were asked to sit quietly in the driver’s seat for 5 minutes to collect the physiological resting baseline. The experimental drive started. The experiment duration lasted nearly 1.5 hours. Recording sites for physiological signals. (a) SC electrodes and (b) ECG SC electrodes.
Feature Generation
In this subsection, we first preprocessed the physiological data to remove the noise and then standardized it. Finally, the feature extraction of the preprocessed data was implemented.
The SC signal was downsampled to 25 Hz and smoothed by a median smoothing filter of 1-second width for eliminating the movement artifact (Lascurain-Aguirrebeña et al., 2019). The phasic component of the SC was extracted using a 0.05 Hz high-pass filter. A bandpass filtered 0.5–35 Hz was used to clean the ECG signal noise (Howells et al., 2014). Heart rate was extracted by the auto-threshold detect method after removing the baseline. As shown in Figure 1, we defined two areas of interest (AOI) to assess participants’ fixations: (1) the road part which included the middle part of the driving simulator screen and two side rearview mirrors (size 500*550 mm); (2) the tablet (size 248*179 mm), on which the NDRT was presented. Outcome variables were monitoring frequency (i.e., percentage of fixations in the road AOI) and monitoring ratio (i.e., percentage of time spent viewing the road AOI).
Descriptions of Produced Features (HR = heart Rate; SC = Skin conductance; min = minimum; max = maximum; standard deviation = SD; Median = med)
Type a: con = continuous variable; cat = categorical variable.
Data Description
The takeover process stage started with a “TOR” and ended when drivers negotiated the takeover mission and reactivated the AV system (see Figure 5). We showed the physiological data during the takeover transition and the monitoring behaviors before and after takeover in this subsection. Takeover situation in Level 3 conditional automation (Gold et al., 2016).
Data from three participants were excluded because of the malfunctions of the physiological sensors. The physiological parameter distributions (i.e., SC and HR) of 31 subjects during TTs are shown in Figure 6a and b, respectively. In general, drivers’ SC and HR distribution are relatively scattered, which do not follow normality through viewing their QQ plots. Distribution of physiological data in TTs: (a) SC; (b) HR (31 subjects).
Two drivers’ fixation time distributions before and after takeovers are shown in Figure 7. Driver’s fixation time of road after experiencing TORs has changed. The changing trend may highly depend on the external driving environment. For example, the fixations on road for subject 5 were unchanged after experiencing the first TOR (Figure 7(a)). However, after experiencing the second TOR, this driver’s fixations on road significantly increased. Fixation time of average before and after six takeovers:(a) Subject 5; (b)Subject 21.
Ground Truth and Statistical Analyses
We assume that driver trust remains unchanged within 1 minute when the AVs work in a normal state. If a driver experiences one TOR or a system error, the trust level may change suddenly, as shown in Figure 8. To measure drivers’ trust responses to TORs accurately, we combined the changes in the subjective trust rating and monitoring behavior before and after takeovers to obtain the trust-change direction during TTs by a commonly unsupervised learning method: the K-means clustering (Guo and Fang, 2013). Trust-change trends in conditionally AVs:(a) Trust without a TOR ; (b) Trust with a TOR.
A GLMM was implemented by the SPSS version 25 to examine the effects and relationships between takeover-related factors, drivers’ physiological responses, and the trust-change direction during TTs (i.e., trust increasement or trust reduction). Besides, the bootstrapping method (Hayes, 2012) was used to test the mediating effect of the takeover criticality on the relationship between scenario types and the trust-change directions.
Model Development and Evaluation
Machine Learning Algorithm and Training Process
As mentioned above, each driver experienced six takeover situations, so a total of 186 samples are available (31 subjects). Considering the challenge of human behavior data collection, a 10-fold nested cross-validation method was used to train models and compare test results (Lee et al., 2013; Du et al., 2020b). The 9-fold training set was used to tune the hyper-parameters with the inner loop and generate classifiers. The model evaluation was based on the remaining 1-fold testing set. With 10-fold cross-validation, all the data samples in the dataset only appeared once in the test dataset. The average F1-score and accuracy of the 10-folds were used to evaluate the model performance.
Results
The result section mainly includes four parts. Section 3.1 demonstrated the clustering result of the trust-change direction during TTs. Section 3.2 showed the effects and relationships between takeover-related factors, physiological responses, and the trust-change direction during takeover missions. Furthermore, the mediating effect analysis of the takeover criticality on the trust-change direction was also shown in section 3.3. Finally, section 3.4 presented the result of the identification model for the trust-change direction.
K-Means Clustering
Consistent with previous studies (Hergeth et al., 2016; Körber et al., 2018b), the subjective trust rating is also significantly correlated with the monitoring behaviors (i.e., monitoring ratio and the monitoring frequency) in this study, with correlation coefficients of −0.35 and −0.39. To mitigate the potential bias of the self-reported trust rating, we used a hybrid measurement that combined the differences in the subjective rating and monitoring behavior before and after takeovers to better represent the changes in driver trust during TTs. The Spearman correlation analysis showed that the correlation coefficient between the monitoring ratio and the monitoring frequency is 0.90 (p < 0.001), indicating multicollinearity between the two features. Thus, only the monitoring frequency was selected as the objective measurement of driver trust since it is stronger correlated with trust.
The proposed trust measurement way and only the subjective rating way both used the K-means method to obtain the ground truth of the trust-change directions during TTs. Among all numbers of clusters varying from two to nine, two clusters had the largest average silhouette coefficient with 0.45 and 0.75 in the two measurement ways, and therefore the trust-change direction was classified into two groups (i.e., the trust-increasing and trust-decreasing groups). For the proposed measurement way, as shown in Figure 9a, the trust-increasing group showed a decreasing trend in monitoring frequency (Mean = −1.92%, SD = 3.11%) and nearly remained unchanged in subjective ratings (Mean = 0.34, SD = 6.88). The trust-decreasing group had an increasing trend in monitoring frequency (Mean = 3.76%, SD = 4.16%) and a reducing trend in self-reported trust ratings (Mean = −12.20, SD = 10.73). For the subjective measurement way, the changes in the self-reported trust rating were similar to the result of our proposed clustering way (Figure 9b). It was almost unchanged (Mean = 0.32, SD = 5.18) in the trust-increasing group, and a drop (Mean = −15.59, SD = 10.03) in the trust-decreasing group. Clustering results of the trust-change directions. (a) Subjective and objective. (b) Subjective.
Sample Distributions of the Trust-Change Directions During TTs in Two Clustering Ways
Correlation Analysis Results Between the Trust-Change Directions and Physiological Parameters, Takeover-Related Factors in Two Clustering Ways (N = 186)
Note: The feature abbreviation adopts the name _max/mean/median, etc.; DV1 = Dependent variable of the proposed way; DV2 = Dependent variable of the subjective clustering way; Scenario types = ST; TOR lead time = TORLT; Takeover frequencies = TF. The following symbol for all the figures and tables applicable were used: *p < 0.05; **p < 0.01; ***p < 0.001.
And then, the samples with alteration in the direction of trust-change were further analyzed by crosstab to examine the distribution difference of takeover-related factors (i.e., scenario types and TOR lead time) in the trust-increasing and trust-decreasing groups. As shown in Figure 10, the crosstab analysis result illustrated that the TOR lead time of 4 s emerged 18 times in the trust-decreasing group, accounting for 62.1% of the group, which was higher than the TOR lead time of 8 s (37.9%). It is consistent with the results of previous studies that the TOR lead time is positively associated with driver trust. Crosstab analysis results for the scenario types and TOR lead time at samples with alteration in the trust-change directions: (a) Scenairo types; (b) TOR lead time (N = 36).
In summary, these results may be provided with suggestive pieces of evidence, indicating that a combination of the subjective rating and objective measure (i.e., the monitoring frequency) can improve the classification results of the trust-change directions.
Besides, drivers’ physiological responses to different TOR lead time conditions (i.e., 4 s and 8 s) were plotted in Figure 11. Both the SC_med and HR_SD of 4 s TOR lead time were higher than those of 8 s TOR lead time (p < 0.1, p < 0.05). Drivers’ physiological responses to the TOR lead time.
Mixed Model Analysis
To increase the stability of the GLMM, physiological features were divided into three fragments using the quartile, representing the three levels of low (first quartile), medium (second–third quartile), and high (fourth quartile), respectively. And then, all variables in Table 5 are regarded as input variables for the GLMM, using the stepwise method to remove the insignificant variables and obtain the final model.
Estimation Result of the GLMM
Note: Standard error = SE; Confidence interval = CI.
aReference category.
Results of Takeover-Related Factors
The significant main effect of the scenario types on the trust-change directions during TTs was found through the GLMM analysis (F (2,177) = 11.19, p < 0.001). The estimated marginal mean (EMM) of the likelihood of trust elevation during TTs is 72.7%, 35.3%, and 33.0% in the stationary vehicle, the roadwork, and the dynamic truck scenarios (Figure 12a). What is more, compared to the reference category, the odds of trust increasing during TT are reduced by 80% and 82% in the roadwork and dynamic truck scenarios (Table 6). It was implied that drivers were more likely to increase trust when dealing with the takeover situation of the stationary vehicle. The estimated marginal mean of the likelihood of increased trust during TTs in different the scenario types and TOR lead time conditions. Error bars are 95% confidence interval. (a) Scenario types. (b) TOR lead time.
The GLMM revealed that the TOR lead time significantly affects the trust-change directions (F (1,181) = 36.05, p < 0.001). Figure 12b shows that the EMM of the likelihood of increased trust during the takeover transition is 24.5% and 71.1% in the TOR lead time of 4 s and 8 s, and their difference is 46.6% (p < 0.001). As demonstrated in Table 6, compared with 4 s TOR lead time, the odds of the trust elevation during the TTs are increased by 7.56 times in 8 s TOR lead time. It was indicated that the TOR lead time is a good predictor variable for trust-change directions during TTs.
Furthermore, as shown in Table 6, the regression coefficient of takeover frequencies is a positive value, indicating that the probability of increased trust during TTs is elevated with the increment of the takeover frequencies. Specifically, each increment in the number of takeovers increases the odds of trust elevation by 36% during the takeover transition. Any interactions were not found between the takeover-related factors.
Results of Physiological Parameters
For the skin conductance, only a significant relationship between the SC_median and trust-change directions was found through the GLMM analysis (F (2,177) = 4.29, p < 0.05). Figure 13a shows that the EMM of the likelihood of increased trust is 60.1% and 32.9% in the low and the medium SC_median. The difference in the EMM of the likelihood of increased trust in the low SC_median is 27.2%, compared with the medium SC_median ( p < 0.01). In addition, when drivers were in the medium SC_median, the odds of increased trust were reduced by 67% compared with drivers in the low SC_median. The interaction effects of the SC were not found (Table 6). The estimated marginal mean of the likelihood of increased trust during TTs in different SC_median and HR_SD levels. Error bars are 95% confidence interval. (a) SC_med. (b) HR_SD.
Besides, the GLMM analysis result showed that the trust-change direction is significantly related to the driver’s HR (F (2,177) = 5.00, p < 0.01). As demonstrated in Figure 13b, the EMM of the likelihood of increased trust is 68.7% in the low HR_SD and is significantly higher than 40.7% and category, the odds of trust elevation are reduced by 78% and 69% in the medium and high HR_SD (Table 6).
The Mediating Effect of the Takeover Criticality
To confirm whether the influence of takeover scenario types (i.e., stationary vehicle, dynamic truck, and roadwork) on the trust-change directions could be implemented by the mediating effect of the takeover criticality. As shown in Figure 14, after controlling for the TOR lead time and takeover frequencies, the bootstrapping method (Hayes, 2012) was used to test the mediating effect of takeover criticality, which assesses the significance through the range of the bootstrap confidence interval. If the range includes zero, the relationship is not significant. The indirect effect of the scenario types on the trust-change directions via the takeover criticality.
The Mediating Effect of Scenario Types on the Trust-Change Directions via the Takeover Criticality
Note: IV = independent Variable; M = mediating variable; DV = dependent variable; CF = control factors; roadwork = RW; dynamic truck = DTK; takeover criticality = TC; trust-change directions = TCD; TOR lead times = TORT; takeover frequencies = TF; confidence interval = CI; d = Significant mediation effect.
Recognition Models of the Trust-Change Direction During Takeover Transitions
Based on physiological features and takeover-related factors in Table 6, the identification models for the direction of trust-change during TTs were established by the RF method and four other machine learning approaches, which were implemented using the packages of the sklearn library. To improve the robustness of machine learning results, the 10-fold cross-validation was run 50 times (i.e., 50 different random seeds) for every machine learning method (Du et al., 2020b). Besides, Wilcoxon’s Signed-Rank test was used to compare the performance of five machine learning models and different feature subset models (Japkowicz and Shah, 2011). Since we compared the performance of the RF model to the other four machine learning models, and the full feature model to other subset feature models-resulting in four and two statistical tests, respectively, the Bonferroni correction was used to counteract the increased probability of Type Ⅰ, which adjusted the significance level to
Model Performance Comparisons
The Recognition Accuracy and F1-Score of Machine Learning Approaches and their Comparisons to the Random Forest Model (10-Fold)

Comparison of Receiver operating characteristic curve between random forest (RF) and the four other models (support vector machine [SVM], Naive Bayes [NB], k-nearest neighbors [KNN], decision tree [DT]).
The Confusion Matrix and Feature Importance
Figure 16a demonstrates the confusion matrix of the RF model. The recall is 80.6% and the precision is 74.1%, showing that the model has better completeness compared to the exactness. By randomly permuting one feature in the out-of-bag data (i.e., one-third of the samples are not in the bootstrap data), and measuring how much this arrangement reduces the accuracy of the model. We represent the feature importance through the proportion of accuracy decline. (a) Confusion matrix; (b) Feature importance using permutation method (TOR lead time = TOTLT; Takeover frequencies = TF; Scenario types = ST).
Figure 16b illustrates the estimated result of all variables’ feature importance using the out-of-bag estimate method. The most important is the TOR lead time followed by the takeover frequencies. Besides, the scenario types and the standard deviation of the HR are also important features for identifying the trust-change directions during TTs with the feature importance value of 0.06 and 0.05. Finally, we found that the median of the SC also contributed to the model performance.
Besides, the identification accuracy of the RF model at different takeover scenario types is also shown in Figure 17, which is 62.9%, 74.2%, and 83.9% in the roadwork, stationary vehicle, and dynamic truck scenarios. The recognition accuracy of the RF model in different takeover scenario types.
Effects of Features on the Random Forest Model Result
The accuracy and F1-score for the RF algorithm using the full feature set were higher than using other feature subsets (Figure 18 and Table 9). Specifically, if only the features of physiological parameters (i.e., SC_med and HR_SD) were used, the prediction accuracy and F1-score were 0.608 and 0.670. If only takeover-related factors (i.e., TOR lead time, scenario types and takeover frequencies) were used, the prediction accuracy and F1-score were 0.696 and 0.736. This suggested that combining the environmental aspects with physiological features reflecting the driver’s internal mental state is necessary for the establishment of a high-performance model. The mean accuracy and F1-score of the model using takeover-related factors and physiological parameters as input features increased to 0.737 and 0.773. Identification accuracy and F1-score of random forests with different feature subsets. Random Forest Recognition Accuracy and F1-Score With Different Feature Subsets and their Comparisons to the Full Feature Model
Discussion
This paper aimed to propose an objective method to measure and identify the trust-change direction during TTs in conditionally AV, mainly including three main aspects: (1) combined the change values in the monitoring behavior and subjective ratings before and after takeovers to estimate the trust-change direction during TTs; (2) explored the relations among takeover-related factors (i.e., the TOR lead time, takeover frequencies, and scenario types), drivers’ physiological responses (i.e., the SC and HR), and trust-change directions during TTs; and (3) developed an RF model to identify the trust-change direction during TTs, which is based on the multimodal physiological signals and takeover-related factors.
Discussion in Trust Measurements
Previous studies used mostly subjective ratings to measure trust in automation, but these self-reported trust ratings have some limitations: (1) they cannot provide continuous information for real-time estimation of driver trust (Walker et al., 2019) and (2) they always have biases from different drivers and driving situations (Rosenman et al., 2011). To alleviate the problem of relying only on self-reported ratings that may not accurately reflect the changes in driver trust during TTs, the differences in the monitoring behavior and subjective rating before and after the takeovers were combined to measure and classify the trust-change direction. The trust-change direction (i.e., trust increment and reduction) during TTs was obtained by the K-means clustering algorithm. Our results showed that integrating subjective ratings and monitoring behaviors can change the classification of the trust-change direction and achieve more objective and reliable trust-change measurements. This is consistent with the result of a study on trust measurement for advanced driver assistance systems (Yu et al., 2021), and they found that a combination of self-reported trust ratings and system usage frequencies can more accurately measure driver trust in this system.
Discussion in Statistical Analyses
Takeover-Related Factors
As expected, the analysis result of the GLMM showed that drivers were more likely to increase trust during TTs when they experienced a longer TOR lead time level. The result might be explained that the longer TOR lead time can reduce the difficulty and risk of takeover missions, and then affect situational trust (Hoff and Bashir, 2015). This result is consistent with the previous study (Jin et al., 2020). We also observed that familiarity with the takeover increased the odds of trust elevation during TTs. It may be associated with experiencing multiple takeover missions that can help drivers understand the purpose and process of the automated driving system, and then affect the learned trust of drivers (Hoff and Bashir, 2015). This matches with the results of research by Hergeth et al. (2017) and Körber et al. (2018c).
The GLMM analysis also revealed that drivers were more likely to decrease trust when dealing with roadwork and dynamic scenario than the stationary vehicle situation. Furthermore, the mediating effect analysis indicated that the takeover criticality rating increased when drivers experienced the roadwork and dynamic truck scenario, compared with the stationary vehicle situation, and then the odds of their trust elevation during TTs correspondingly were decreased. Consistent with previous studies, the roadwork and dynamic truck scenarios had a higher takeover criticality rating than the stationary vehicle situation (Vogelpohl et al., 2018; Yi et al., 2022), and then drivers were more likely to reduce trust ratings in critical situations (Morris et al., 2017; Li et al., 2019; Azevedo-Sa et al., 2021).
Physiological Parameters
For the SC, it was found that drivers were more likely to increase trust during TTs in the low SC_med, compared to the medium SC_med through the GLMM analysis. This is consistent with the research results of Khawaji et al. (2015) and Morris et al. (2017), which reported that trust increased with the reduction of the SC. However, we found that when drivers were in the low and high SC_med, there was no significant difference in their odds of trust increasement during TTs. One possible reason for this result is that the high SC level is only associated with emotional arousal and cannot distinguish emotional valences such as excitement or stress (Lang et al., 1993). Drivers may be excited or happy that they can safely handle the takeover mission, which may increase driver trust during TTs. This in turn may obscure the negative correlation between the SC and trust-change directions. Besides, the GLMM analysis result also suggested that the trust-change direction is negatively correlated with the standard deviation of the HR during TTs. This may be attributed to the increased stress on drivers resulting from trust reduction during TTs (Mühl et al., 2020), which then leads to the acceleration of stress-mediated HR (Healey and Picard, 2005; Rastgoo et al., 2018).
Discussion in the Trust-Change Direction Model During Takeover Transitions
Compared with the other four popular machine learning approaches (e.g., SVM, DT, and KNN), the performance of the RF method seems to outperform the other classification methods in identifying the trust-change direction during TTs. This is lined with prior studies on recognizing driver trust in ADAS and predicting the takeover performance (Du et al., 2020b; Yu et al., 2021). The RF approach also shows its supremacy in identifying the trust-change direction during TTs. This may be associated with the advantages of bagging and bootstrapping techniques, which reduce the effects of overfitting and improve generalization.
Our results showed that the model performance with the full set of features is improved compared to partial features (i.e., physiological data only and takeover-related data only). One possible reason for this result is that the data from different modalities can reflect the same change of trust and each modality carries trust-related information. Besides, there is some trust-related information across the modalities that can provide complementary information to improve the model performance of the trust-change directions. For example, full physiological data reflect the internal mental state of drivers and interactions with driving environments, such as stress, which may be related to the change in driver trust (Morris et al., 2017; Mühl et al., 2020). Thus, fusing multimodal data can improve the identified performance of the model. Using vehicle sensors and advanced wearable technology, the above features can be recorded in a minimally invasive way to identify the trust-change directions during TTs in conditionally AVs.
Limitations and Future Work
Several limitations of this study need consideration and should be addressed in future research. First, as only the influences of the TOR lead time and takeover scenario type on driver trust were examined, the study merely covered a small number of possible takeover situations when drivers were engaged with a visual Surrogate Reference Task in low traffic density. Further study could be extended to exploring the impacts of these two factors on driver trust when they are performing different types of NDRT (e.g., reading and proofreading text) or encountering different levels of traffic density (e.g., low and high traffic densities). Second, the study was conducted with a fixed-base driving simulator in a controlled laboratory, and the research samples comprised mainly young males. The physiological signals of drivers may be different when they are of diverse ages and drive in a real-world environment, as the physiological signals collected are affected by the combination of some factors, such as temperature, and age (Gavazzeni et al., 2008; Doberenz et al., 2011). Future studies can recruit diverse drivers in a natural driving environment to increase the generalization of models. Finally, this study used the statistics of the physiological signals of drivers during the takeover period as the model input and did not consider the sequence dependence among the time-series data. Future studies could try to use long-short-term memory (LSTM) or a combination of convolutional neural networks and LSTM to identify the trust-change directions during TTs.
Application Potential
Those findings of the takeover-related factors revealed two layers of variability in driver-vehicle trust (situational trust and learned trust), which may provide additional support for the concept design of the driver-vehicle interface in conditionally AV. And the model of the trust-change directions may provide implications for developing trust monitoring systems in SAE level 3 AVs to mitigate both overtrust and undertrust. For example, if the model detects that the driver loses trust after experiencing TTs, the SAE level 3 AVs can foster better driver-vehicle trust by proactively communicating (Collet & Musicant, 2019). Besides, the physiological parameters can be applied to not only identify the trust-change direction during TTs but also monitor the internal mental state of drivers by receiving continuous and real-time physiological signals. Based on the monitoring result, a mild onboard environment can be offered to drivers when they need to relieve the sub-optimal mental state caused by the takeover transition to improve the driver’s experience (Chung et al., 2019).
Conclusions
The study provides a unique perspective for objectively and reliably measuring and identifying the trust-change directions during TTs. A combination of the change in the monitoring behavior and subjective trust rating before and after takeovers can reliably estimate the changes of trust during TTs in conditionally AVs, which may help to reduce the biases of using only subjective evaluations.
Besides, the present study has examined the effects and relationships between takeover-related factors, drivers’ physiological responses, and trust-change directions during TTs. Drivers are more likely to increase trust during the takeover transition when they are in the longer TOR lead time condition, with more takeover frequencies and dealing with the stationary vehicle scenario. The change in driver trust during the takeover process has negative correlations with the standard deviation of the HR and the median of the SC.
More importantly, a recognition model of the trust-change direction during TTs is established by the RF algorithm, based on the driver’s multimodal physiological data (i.e., the SC and HR) and takeover-related factors (i.e., TOR lead time, takeover frequencies, and scenario types). The identification performance of the RF classifier has an F1-score of 77.3% and an accuracy of 73.7%.
KEY POINTS
Combining the changes in the subjective trust rating and monitoring frequency before and after takeovers can objectively and reliably measure the trust-change direction during takeover transitions. The influence of scenario types on the trust-change direction during takeover transitions can be implemented by the mediating effect of the takeover criticality. The present results indicate negative connections between drivers’ physiological parameters (i.e., the SC and HR) and the trust-change direction during takeover transitions. The study finds, that fusing driver’s physiological data and takeover-related factors (i.e., TOR lead time, takeover frequencies, and scenario types), can achieve a relatively accurate classification for the trust-change direction during TTs, with an F1-scores of 77.3%.
Footnotes
Acknowledgments
This work was supported by the National Natural Science Foundation of China under grant number 51975194, the Natural Science Foundation of Hunan Province under grant number 2021JJ30121, and the State Key Laboratory of Automotive Safety and Energy under Project No. KFZ2203. We sincerely thank constructive comments and valuable suggestions from anonymous reviewers.
Author Contributions
Binlin Yi: Methodology, Data curation, Formal analysis, Writing original draft. Haotian Cao: Writing—review & editing. Xiaolin Song: Conceptualization, Project administration, Funding acquisition. Jianqiang Wang: Writing—review & editing. Song Zhao: Writing—review & editing. Wenfeng Guo: Visualization, Investigation. Dongpu Cao: Software, Validation.
Ethical Approval
This study was approved by the Institutional Review Board in the college of mechanical and vehicle engineering at Hunan University.
