Abstract
With the growing adoption of autonomous technologies across various domains, an increasing number of studies have explored collaborations between humans and agents working together to achieve shared goals, forming human-agent teams (HATs). While much of the research has focused on dyadic relationships involving a single human and a single agent, the current study examines multi-human multi-agent teams where multiple humans and agents collaborate to achieve team goals. The shift from dyadic to more complex teams makes research topics studied in triadic or larger human-human teams relevant to HATs. Of particular interest in this study is the concept of “trust in team” and its relationships with trustor and trustee characteristics. The study used an adapted version of the Blocks World for Teams (BW4T) testbed, where a team of four (two humans and two agents) performs a collaborative block-moving task. This study first validated the use of the existing interpersonal trust scale and team trust scale for evaluating trust in human/agent teammates and in the team, respectively. The next step involved examining how trust in the team is formed in relation to trust in individual teammates. Additionally, the associations between team trust, communication, and performance were investigated.
Keywords
Introduction
With the increasing introduction of scenarios where humans and agents collaborate to achieve a common goal, numerous studies on human-agent teaming have explored ways to maximize human-agent team (HAT) performance (Chung et al., 2024; Lyons et al., 2021; O’Neill et al., 2022). Many have examined the effects of team composition, task design, automation behavior, and agent characteristics on performance (Bhat et al., 2024a, 2024b; Demir et al., 2023; Guo et al., 2024; Johnson et al., 2023; Song et al., 2022), often measured by time or accuracy metrics. Also, with the size of the HAT being investigated increasing due to advanced technology, it is becoming increasingly important to understand trust in different types and levels of referents to better capture team dynamics (Wildman et al., 2024). However, although trust in the AI and/or human teammate has been explored, limited research has investigated trust in the team in HAT and its association with other team metrics. This gap leaves the assessment and understanding of HAT dynamics largely unexplored.
In human organization and team research, the multi-level multi-referent theory of trust (Klein & Kozlowski, 2000), which explains that trust operates across different levels (individual, team, organization) and referents, is widely accepted. Studies have explored the antecedents and consequences of trust in a higher-level referent (Fulmer & Gelfand, 2012). Trustor and trustee characteristics are known to influence team trust, which in turn affect outcome metrics (e.g., team performance, satisfaction, and decision-making effectiveness) (Costa, 2003). Regarding trustor characteristics, an individual trustor’s identification with the team influences their trust in the team (Colquitt et al., 2011). For trustee (i.e., team) characteristics, factors such as team integrity and benevolence (Colquitt et al., 2011), group membership, and role structure (Kramer, 1999; Meyerson, 1996) play a role in the development of trust in the team.
Building on prior research, it can be hypothesized that the performance and characteristics of individual teammates contribute to an individual’s overall trust in the team. This may be particularly relevant in team tasks characterized by high unpredictability, where the integrity of individual team members has a greater impact (Colquitt et al., 2011).
However, this line of inquiry into the trust in the team has not yet been fully extended to the context of HATs. Relatively few studies have examined the differential influence of trust in individual teammates on trust in the overall team, primarily due to the limited studies on multi-human multi-agent collaboration (Chung et al., 2024). Specifically, it remains unclear whether trust in a high-performing and influential teammate carries more weight than trust in a low-performing teammate, or whether all teammates’ trust levels contribute equally. Additionally, little is known about the proper evaluation scales that can be used to evaluate trust in each teammate and the team in HATs.
This study builds on prior research on human teams by focusing on the “trust in team.” Our goal is to offer a comprehensive perspective, from validating interpersonal and team trust scales commonly used in human team research to understanding the development of trust in the team. Specifically, through a human-subjects experiment, this study aims to: (1) validate the appropriateness of using existing trust scales, (2) examine how trust in the team is formed in relation to trust in individual teammates, and (3) explore the associations between trust in the team and team outcome metrics.
Methods
Experimental Testbed and Protocol
Thirty teams, each consisting of two participants, participated in the experiment. Each team was paired with two agents, forming a group of four players tasked with completing a series of block-moving trials (Figure 1).

Overview of the team composition and experimental task.
The study utilized an adapted version of the Blocks World for Teams (BW4T) testbed (Johnson et al., 2009), which has been used in previous HAT studies (Figure 2). Humans searched rooms to locate target-colored blocks and issued commands to one of the agents to pick up and deliver them. Agents then acted based on these commands while providing transparent updates on their actions. Communication was limited to text chat, where humans designated the recipient and message type, drafted the message, and sent it. Task interdependence was manipulated through the following setup: only the player present in a room could see which blocks were inside, necessitating active engagement and collaboration. Additionally, human and agent roles were distinct, requiring both to perform their tasks to deliver a block, making human-agent collaboration essential.

BW4T testbed used in this study. (a) Screenshot of the participant’s display before the onset of an experimental trial. (b) Example of the environment and how participants engaged in the experiment.
Each trial required moving seven differently colored blocks. All teams completed three sets of ten trials each, totaling thirty trials. For each set, they worked with agents assigned to different reliability pairing conditions: under the perfect condition, both agents executed commands flawlessly. In the mixed condition, one agent had 100% reliability while the other had a 50% chance of deviating from commands. For half of the teams, agent 1 was imperfect; for the other half, agent 2 was imperfect. In the imperfect condition, both agents had 50% reliability.
Experimental Procedure
Participants provided consent, completed a demographic survey, and received task instructions, including the opportunity to earn a performance-based bonus. After being seated at separate laptops with screen partitions, they completed a practice session so that they could become familiar with the task and agent behaviors. The practice included four trials with agents exhibiting perfect and zero reliability, and participants were informed that actual agent reliability would differ during the experiment. They then completed three sets of trials, followed by a debriefing and compensation. The entire experiment lasted approximately 2 hours.
Measures
In the study, a variety of measures were collected, including personal characteristics, post-trial and post-session trust scores, communication frequency, and performance.
For the personal characteristics, seven dimensions that were previously identified as significantly associated with users’ trust dynamics (Chung & Yang, 2024a, 2024b) were collected before the experiment.
Individual trust ratings in four referents, each three teammates (human, agent 1, agent 2) and the team, were assessed both post-trial and post-session. The post-trial question was a one-item survey on a scale of 0 to 100 for each referent. For the post-session survey, well-established scales from the human team research were used. The current study focused on analyzing the post-session trust, specifically on validating the appropriateness of using these scales and investigating how the scores correlate with other key metrics.
Concerning the trust in human/agent teammate, McAllister (1995)’s interpersonal trust scale was used (Table 1). The 11-item scale consists of two dimensions, cognition- and affect-based trust in individual teammates.
Interpersonal Trust Scale (McAllister, 1995) which was Used to Evaluate Trust in Human/Agent Teammates.
Regarding the trust in the team, the 21-item team trust scale from Costa and Anderson (2011) was employed, which includes four dimensions (propensity to trust, perceived trustworthiness, cooperative behaviors, monitoring behaviors) (Table 2).
Team Trust Scale (Costa & Anderson, 2011) which was Used to Evaluate Trust in the Team.
Additionally, communication logs and task completion time were recorded for each trial. Communication frequency was determined by the total number of messages issued by humans. Performance was measured by the time taken to complete the delivery of all seven blocks.
Analyses
The appropriateness of using the existing trust scales was first assessed by evaluating both the reliability and structural validity of the measures. Specifically, three types of validation were conducted: (1) validation of the interpersonal trust scale in assessing trust in agent teammates in HATs, (2) validation of the same scale in assessing trust in a human teammate in HATs, and (3) validation of the team trust scale in assessing trust in a team in HATs.
For all three, internal consistency using Cronbach’s alpha was assessed to examine the coherence of responses across all survey items. Subsequently, confirmatory factor analyses (CFA) were conducted to evaluate the model fit of the entire scale and evaluate the construct validity. In cases where an item was flagged as a candidate for removal in the first phase (i.e., its deletion led to an increase in Cronbach’s alpha), the model fits of the full and reduced versions of the scale were compared to determine whether the removal improved overall model fit.
Following scale validation, to examine how individual trust in the team is formed in relation to trust in each of the three teammates, and to investigate whether the pairing condition influences the contributions of these trust ratings to overall trust in the team, a linear mixed-effects model was developed. Participant and team IDs were included as random effects to account for variability across individuals and teams, while the variables of interest were treated as fixed effects. Additionally, interaction terms between the pairing condition and trust in each referent were included to evaluate the moderating effect of the reliability pairing condition on the contributions of trust in each referent to trust in the team.
Results
Trust Scales Validation
Table 3 presents the results of the reliability analyses. While the responses to all items within each scale and dimension exceeded the acceptable reliability threshold (Cronbach’s alpha ≥ 0.7), the results suggest that deleting certain items may lead to higher internal consistency. For the use of the interpersonal trust scale in assessing trust in both the agent and the human teammate, it is recommended to remove one item: Cognition (6). For the team trust scale, the deletion of two items is recommended: Perceived trustworthiness (10) and Cooperative behaviors (16).
Reliability Analysis Results (The “N/A” Cell Indicates that None of the Item Deletions Led to an Increase in Reliability).
Subsequently, CFA was conducted to evaluate the model fit statistics for both the full and reduced versions of each scale across the three types of validation (Table 4). The results indicate that, for all scales, the reduced versions demonstrated better model fit, meeting commonly accepted thresholds. Accordingly, all subsequent analyses are based on the statistics from the reduced versions.
Confirmatory Factor Analysis Results.
Trust in Team
A linear mixed model analysis was conducted to model trust in the team as a function of trust scores in the three teammates. Figure 3 presents the distribution of the scores by the condition.

Average trust in the team and teammates (H: human, A1: agent 1, A2: agent 2) by the agent reliability condition.
During the analyses, none of the interaction terms involving the reliability pairing condition were statistically significant (p > .05) and thus were excluded from the final model. The final model revealed that trust in all three team members (human teammate, agent 1, and agent 2) were significant predictors of trust in the team (human: F(1, 176.99) = 335.51, p < .001; agent 1: F(1, 151.21) = 45.91, p < .001; agent 2: F(1, 141.73) = 32.43, p < .001) (Table 5).
Effects of Trust in Each Teammate on Trust in the Team.
Additionally, trust in the team was significantly associated with communication frequency (r = −0.44, p < .001) and performance (r = −0.43, p < .001). In other words, individuals with higher trust in the team communicated more efficiently with their teammates, which resulted in significantly shorter task completion times.
Discussion
This study adopts a multi-referent trust perspective to examine individual trust in teammates and the team, and their associations with communication and performance in multi-human multi-agent collaboration.
Well-established trust scales developed for human teams were first validated in the context of HATs. For assessing trust in individual teammates, regardless of whether they are humans or agents, the interpersonal trust scale developed by McAllister (1995) could be appropriate. Using the same scale across teammates with distinct roles and identities allows for direct comparisons of trust in different referents, which is increasingly important in the field of HAT. Additionally, while there has been limited effort to evaluate individual trust in the overall “team” within HAT contexts, the results suggest that the team trust scale developed by Costa and Anderson (2011) can serve this purpose effectively.
Using the validated survey scores, the results in modeling trust in the team in relation to trust in the teammates support the importance of “trust in the team.” This trust construct presents two important information about the team: it reflects aggregated trust in individual teammates and acts as an indicator of the team outcomes.
First, it captures how a trustor evaluates each teammate and their relative importance in the team’s functioning. In this study, trust in the human teammate contributed most significantly to trust in the team, compared to trust in the two agents. This is likely because, in the experimental task, human teammates played a central role in locating correctly colored blocks and issuing timely commands to the agents. In situations where agents were imperfect, close coordination and corrective actions from humans became even more critical for quickly recovering from errors caused by agents executing commands incorrectly. As a result, the team’s success heavily depended on the human teammate, which likely led to trust in the human carrying greater weight in shaping trust in the team. Along this line of inquiry on how trust in multiple distinct trustees is aggregated, PytlikZillig et al. (2024) proposed the “Perceived Influence Model of Trust.” It outlines three possible mechanisms: additive (trust combines additively), compulsory (trust in the least trusted trustee is weighted more heavily), and compensatory (trust in the most trusted trustee is weighted more heavily). The results of this study align with the compensatory scenario, in which trust in the most trusted trustee, which in this case was the human teammate who also contributed the most to the team task’s success, diminished the effect of trust in less trusted trustees when forming aggregate trust (i.e., trust in the team). However, as the study features a single collaboration scenario, future studies can build on this approach to evaluate team dynamics under various conditions.
Second, the significant associations observed between trust in the team, communication, and performance support prior research showing that trust within the team enhances team effectiveness by facilitating efficient collaboration and goal alignment (De Jong et al., 2016; Dirks, 1999). Regarding communication frequency, sending more prompts than necessary can be interpreted as inefficient collaboration, leading to longer task completion times. Thus, the findings suggest that trust in the team is a meaningful indicator of teamwork quality.
Overall, this study offers meaningful guidelines for evaluating trust in multi-human multi-agent collaboration. However, it has several limitations that can inform future research. The study employed agent reliability pairing as the primary independent variable, which is fairly straightforward and expected to have a linear relationship with trust. To gain a more comprehensive understanding of various trust constructs, future studies could explore other important team characteristics, such as leadership and role structures. Additionally, while the agents in this study only acted in response to human commands, future research could examine more intelligent and proactive agents. With such agents, the contribution of trust in each teammate to trust in the overall team may differ.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This material is based upon work supported by the National Science Foundation under Grant No. 2045009.
