Abstract
While quantitative research on trust in HATs is well-established, there is a significant gap in understanding individuals’ subjective experiences during trust violations and repair. This study explores how people perceive and respond to these events, guided by three key questions: (1) How do participants describe emotional and behavioral responses to trust violations and repair? (2) How do personality traits influence reactions? (3) What qualitative themes emerge in the dissolution of trust? Using a mixed-methods approach, participants engaged in a 1-hour simulated search-and-rescue mission with four autonomous agents. During the task, one agent committed two trust violations, followed by repair attempts across five conditions (none, individual agents, or full team). Trust ratings were collected throughout, and participants provided written reflections at the end. Qualitative responses and behavioral data were examined using regression and ANOVA. Five themes emerged: emotional responses, performance changes, blame attribution, loss of trust, and expectations of the violation. Exploratory analyses showed strong correlations among themes, and regression models suggested that individual differences and thematic responses explained significant variance in trust dissolution.
Keywords
Introduction
Autonomous systems are becoming more commonplace among complex and interdependent teams across various domains. These systems encompass a range of technologies, including unmanned aerial vehicles and ground robots, that operate with varying levels of autonomy. In this study, we use the term agent to refer to autonomous systems that serve as teammates, whether physically or virtually, that interact with humans to achieve shared goals (O’Neill et al., 2020; Wildman et al., 2024). As the use of agents expands, it is crucial to understand the trust dynamics of human-agent teams (HATs), exploring how trust is developed, violated, and potentially repaired. We define HATs as at least one autonomous agent working interdependently with at least one human, totaling two or more members (McNeese et al., 2018; Wildman et al., 2024).
Trust in HATs, defined as “the attitude that an agent will help achieve an individual’s goals in a situation characterized by uncertainty and vulnerability” (Lee & See, 2004, p. 54), plays a central role in influencing behaviors like reliance and decision-making (Madhavan & Wiegmann, 2004). Prior research has extensively examined trust quantitatively (see O’Neill et al., 2020 for a review), such as tracking changes in numeric trust scores in response to violations and repairs. However, there is a notable gap in exploring the subjective, lived experiences of individuals navigating these dynamics in HAT contexts, which can influence our understanding of trust in HATs (Carmody et al., 2022). For example, research has found perspective differences in those that view technology as teammates versus tools. Lyons et al. (2019) found that those who saw technology as a teammate discussed benevolence and interdependence more often than those who saw technology as a tool.
This study seeks to fill that gap by focusing on the descriptive experiences individuals have when trust violations and repair attempts occur within HATs. Unlike quantitative studies that measure trust as an abstract score, this research aims to delve into how individuals perceive and interpret these events to uncover the nuanced ways that trust manifests in interactions with agents. We explore the following research questions:
How do participants describe their emotional and behavioral responses to trust violations and repair?
How do individual differences, such as personality traits, influence expectations and reactions to trust violations and repair events?
What, if any, emerging qualitative themes influence the dissolution of trust in the violating agent?
Understanding these subjective experiences can provide richer insights into the mechanisms underlying trust and inform the design of more effective trust repair strategies in HATs.
Methods
A simulation study was conducted in which participants performed a 1-hour search and rescue mission on a desktop simulator using our Multi-Agent Team Trust Emergence Research (MATTER) testbed. While attempting to work with four agents, two unmanned aerial vehicles and two ground robots, to achieve mission goals, participants experienced two trust violations from the same agent. Five conditions were presented in this study: three conditions with one agent repairing trust, one condition where all agents repaired trust, and one condition where no agents repaired trust. In all conditions, the same agent violated trust twice. A one-item measure rating trust in each agent from (1) not at all to (5) a great deal was administered every 10 min. Qualitative measures were collected at the end of the experiment. In the last survey, if participants indicated that any teammate made a mistake, they were asked to explain in detail how the teammate’s error made them feel. Similarly, if participants indicated that the agent attempted to correct the issue, they were prompted to provide an in-depth explanation of what the agent did to repair the mistake, and how the attempted correction made the participant feel. These measures provided a view of other factors that may be affecting participant’s trust. This research complied with the American Psychological Association Code of Ethics and was approved by the Institutional Review Board at the Florida Institute of Technology.
Participants
Our study included 161 participants (99 male, 61 female, 1 NA) with ages ranging from 18 to 44 (M = 21.59, SD = 3.93) who were recruited through various channels such as personal networks, distributing flyers across a university campus, emails, LinkedIn posts, online forums, and snowball sampling methods. Participants were either compensated with a $25.00 Amazon gift card or extra credit for approved university courses.
Qualitative Analysis Procedure
The data were analyzed using a thematic analysis approach (Braun & Clarke, 2023) and included four steps. First, the data were carefully cleaned by thoroughly reviewing the, organizing, and de-identifying participant responses. In cases where responses were missing, partial data were retained and used in the analysis. Second, using an inductive thematic analysis (Braun & Clarke, 2023), codes were generated from participants’ responses, independently identified by three researchers with regular consensus meetings. During the consensus meetings, a fresh spreadsheet was used to add the identified codes and they were tallied with frequency counts. Third, the codes were integrated into overarching themesq (see Table 1). For example, codes that identified emotions (e.g., annoyed, disappointed) were merged into the theme Emotional Response. To maintain traceability, participant IDs and code locations were added to the spreadsheet. During this process, the raw data were reviewed to ensure the context in which a response was captured was accounted for. Lastly, three researchers met and comprehensively reviewed the themes for consistency, and the themes were accepted. The process involved reviewing each theme, its description, and the context of the raw data one final time to ensure wuane were reflecting the participant responses accurately, such as reflecting both semantic meaning and latent meaning (Braun & Clarke, 2023). For example, a response such as, “I tried to look for the problem and fix it,” indicated a general behavioral change in approach Performance/Behavioral Change - General change) with an undertone of taking responsibility to “fix” the agent’s mistake (Blame/Responsibility – Self).
Overarching Themes from Qualitative Responses.
Results and Discussion
Qualitative Findings
The recurring themes identified are in Table 1. In response to violations, we found Emotional Response was the most prevalent theme overall, with 86 occurrences across all conditions. Of all instances of themes, emotional response had a 50% frequency of occurrence. Emotional responses included various emotions such as annoyed, disappointment, surprise, and upset. For example, one participant stated, “I felt a little frustrated especially when it made the same mistake twice.” Interestingly, conditions in which individual agents repaired had more emotional responses recorded than conditions in which all members or none repaired. Notably, the condition in which the violating agent repaired had fewer participants feeling “Surprised” than other conditions. The findings suggest that observable emotions such as frustration and disappointment underscore the importance of addressing emotional impacts in designing agent behaviors. The reduction of surprise when the violating agent performed the repair indicates a shared expectation among the agent and participant, which could have implications for enhancing trust by setting expectations in human-agent interactions.
In addition to their emotional responses, participants also reported their behavioral responses in which they indicated a change in their behavior in response to the mistake or the consequences of the mistake. We recorded 25 instances of performance/behavioral changes in total. Participants reported becoming more alert or cautious following the agent’s mistake, indicating an adjustment in their behavior to mitigate potential risks. For example, a participant stated that their reliance on the violating agent decreased, stating, “I only used it when I had to from then on. I felt it was bringing us all down as a team.” Such findings indicate a behavioral adaptation when the team or human participant experiences a violation, suggesting that participants perceive and response to violations as potential risks to performance. These findings highlight the notion that a noticeable change in the participants’ behavior with regards to agent tasking and communication, may indicate a change in trust among the human agent team, and may be a useful behavioral indicator to monitor.
Participants assigned blame or responsibility to entities for the mistake for 10% of recorded themes. Participants sometimes attributed the mistake to the agent (e.g., “I felt they should have been more careful and must try avoiding the same mistake twice”), to themselves (e.g., “Thought I should’ve motivated them to do [search and rescue] SAR more slowly.”), or to the human builders of the technology (e.g., “I also knew that it was not entirely its fault, for it was built by humans.”). Moreover, participants suggested they expected the violation in 5 instances across all conditions. For example, one participant wrote, “Also the fact that the same accident that happened was also in the tutorial, so it did not shock that much.” Moreover, we recorded only 10 instances of participants explicitly mentioning a loss of trust. These findings indicate that this is a nuanced attribution process, one in which participants recognize the broader network of working in a HAT, and it in turn may influence how they perceive accountability in human-agent teams and interactions. The results suggest that it may be beneficial to assess how best to mitigate or rectify a violation across all dimensions of the human agent team, whether that be from the agent itself directly, the manufacturer, or on a broader level.
In response to repairs, we found that participants’ perception of a repair occurring changed based on the agent who repaired. When the violating agent enacted a repair, roughly 1/3 of participants (8/28) did not perceive this as a repair. For participants in the condition in which the full team repairs, only one participant did not perceive this as a repair. Another interesting finding was that participants in which there was no repair occasionally answered yes to the previously mentioned qualitative question and cited themselves and behavioral changes they pushed onto the agents as an act of repair. The findings suggest that the recognition of a repair event in HATs depends heavily on the agent or team member enacting it. Showcasing, contrary to considerable research, the agent who violates may not be the most ideal candidate to perform the repair, and it is, in fact, more influential when all members of the team contribute to a repair.
Exploratory Quantitative Results
We conducted further exploratory analyses by combining survey measures with qualitative themes, aligning with recommendations for mixed methods integration (Creswell & Creswell, 2018). Specifically, we converted qualitative themes into dummy-coded quantitative variables to examine potential associations with individual differences, such as personality traits (Extraversion, Agreeableness, Conscientiousness, Neuroticism, and Openness; International Personality Item Pool, IPIP; Donnellan et al., 2006), domain experience, propensity to trust automation (Jessup et al., 2019), and perfect automation schema (Merritt, 2015), or potentially confounding factors such as caffeine amount consumed. The only significant correlation between individual differences and qualitative themes was between Agreeableness and the theme of expecting a violation (r = −.256, p = .002). Though weak, the negative correlation between Agreeableness and expectations suggests that more agreeable individuals may not expect violations. This concept is supported by previous literature, which states what when individuals are more agreeable, trust in automation is higher, (Hanna & Richards, 2015). The weak negative correlation between agreeableness and expectations also indicates that agreeable-leaning personalities may yield delayed, or less severe reactions when a trust violation occurs. In practicality, this concept may point to the fact that certain interventions for repairing trust may be customizable depending on the personality type of the user.
The themes showed strong positive correlations with each other (ranging from r = .865 to .982). The strong, positive correlations among the themes may suggest that participants experience a trust violation holistically, affecting multiple facets such as emotions, cognitive appraisals of the situation, and behaviors.
Moreover, we investigated the relationships between themes and the overall trust dissolution in the violating agent. Trust dissolution was calculated by subtracting the survey after the second violation from the survey before either violation and dividing by 10. There was a significant negative correlation between the theme of trust loss and trust dissolution in the violating agent (r = −.229, p = .006). The results indicate that as expressions of trust loss increase, trust in the violating agent deteriorates further, for example, as the qualitative data mentioned trust loss more, quantitative trust scores changes. This finding emphasizes the importance of addressing trust in reducing the erosion of trust over time. Additionally, we conducted a hierarchical regression consisting of controls (IPIP, caffeine consumption, domain experience, propensity to trust, perfect automation schema) and the qualitative themes. This model was not statistically significant (p = .059), but explained 67.5% of the variance in the dissolution of trust in the violating agent. While this should be interpreted cautiously due to the lack of statistical significance, the variance accounted for is large.
Lastly, an analysis of variance (ANOVA) was conducted to examine the factors influencing the dissolution of trust in the violating agent. The assumption of equal variances was met for trust dissolution across groups, as Levene’s Test of Equality of Error Variances was non-significant, F(4,137) = 0.989, p = .416. The overall model was significant, F(41, 100) = 1.671, p = .020, explaining approximately 40.7% of the variance in the dissolution of trust (R2 = .407). The adjusted R2 = .163 indicated a more modest fit when accounting for the number of predictors. Significant main effects were observed for the themes Performance/Behavioral Change and Trust Loss. Changes in performance behavior significantly predicted the dissolution of trust F(1,100) = 4.734, p = .032, which indicates that participants who discussed performance change experienced greater trust dissolution than those who did not. Trust Loss had a significant impact on the trust dissolution, F(1,100) = 7.182, p = .009. This suggests that those who discussed trust loss experienced a different pattern of trust dissolution than those who did not mention losing trust in the violating agent.
A significant three-way interaction was found among conditions, Emotional Response, and Performance/Behavioral Change, F(4,100) - 3.262, p = .015. This suggests that the effect of performance behavioral changes on trust dissolution depends on both the emotional response and the specific condition. Lastly, there was a significant interaction between Performance/Behavioral Change and Blame/Responsibility, F(1,100) = 3.754, p = .005. Interestingly, when examining the same factors’ influence as above on the overall slope of trust restoration in the violating agent, the overall model was not significant, F(1,139) = 1.129, p = .301.
Conclusion
As autonomous agents become more integrated into complex, interdependent teams, understanding how people experience trust development, violation, and restoration is vital for effective collaboration. This study highlights the nuanced ways individuals experience trust violations and repairs in HATs, demonstrating that these events impact emotional responses, cognitive appraisals, and behavioral adaptations. Participants’ experiences of emotions, behavioral changes, and differing perceptions of repair effectiveness suggest a holistic experience of trust in HATs. These findings suggest that trust repair strategies should address both cognitive and emotional dimensions to be effective. Additionally, tailoring interventions to individual differences, such as agreeableness, may increase effectiveness. Overall, these insights provide an initial foundation to understanding the lived experiences of humans in HATs, which can inform designing strategies that foster trust for more effective human-agent interactions. Future research should extend and repeat this analysis within more complex, such as multi-human multi-agent teams (MH-MATs), where coordination demands, role diversity, and subgroup dynamics may further influence trust-related challenges.
Footnotes
Acknowledgements
We would like to thank the Air Force Research Lab’s Gaming Research Integration for Learning Laboratory (GRILL) for developing the HAT M.A.T.T.E.R simulation testbed that was utilized to complete this research.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This material is based upon work supported by the Air Force Office of Scientific Research (AFOSR) under award number FA9550-21-1-0294.
