Abstract
Humans can effectively search visual scenes by spatial location, visual feature, or whole object. Here, we showed that visual search can also benefit from fast appraisal of relations between individuals in human groups. Healthy adults searched for a facing (seemingly interacting) body dyad among nonfacing dyads or a nonfacing dyad among facing dyads. We varied the task parameters to emphasize processing of targets or distractors. Facing-dyad targets were more likely to recruit attention than nonfacing-dyad targets (Experiments 1, 2, and 4). Facing-dyad distractors were checked and rejected more efficiently than nonfacing-dyad distractors (Experiment 3). Moreover, search for an individual body was more difficult when it was embedded in a facing dyad than in a nonfacing dyad (Experiment 5). We propose that fast grouping of interacting bodies in one attentional unit is the mechanism that accounts for efficient processing of dyads within human groups and for the inefficient access to individual parts within a dyad.
Recognition of biological entities is vital to human survival. Ancestral motives, lifelong experience, and internal motivations are thought to account for the prioritization by human visual attention of properties and stimuli such as biological motion, eye gaze, faces, and bodies (Birmingham & Kingstone, 2009; Downing, Bray, Rogers, & Childs, 2004; New, Cosmides, & Tooby, 2007; Papeo, Wurm, Oosterhof, & Caramazza, 2017).
We propose that, for humans, recognition of ongoing social exchanges is no less important than recognition of biological entities. Observed interactions between two or more individuals require rapid discrimination to activate adaptive behaviors such as defense, assistance, or cooperation (Quadflieg & Koldewyn, 2017). Through third-party interactions, individuals can rapidly infer group affiliation, social conventions, and conformity (Powell & Spelke, 2018).
Specific perceptual adaptations might meet the requirement for efficient processing of third-party interactions, allowing, for example, efficient detection (Papeo, Stein, & Soto-Faraco, 2017). Such mechanisms could be particularly crucial for parsing crowded scenes (i.e., scenes with multiple preferred stimuli such as faces and bodies) and selecting and prioritizing the portion in which a social exchange is unfolding.
The mechanisms for parsing and selection in cluttered environments have been extensively studied with visual search. Typically, the subject searches for a target among a set of distractors. Search efficiency, the rate at which items are processed, can be affected by task demand, properties of the array (e.g., spatial frequency and predictability of target location), perceptual or semantic properties of target and distractors, and the relation between them (Duncan & Humphreys, 1989). Moreover, given two kinds of stimuli, searching for an A among Bs can be more efficient than searching for a B among As (Treisman & Souther, 1985). Such asymmetry may occur because target A is more salient than target B, or distractor Bs are easier to check and reject than distractor As. Moreover, depending on the search protocol, one effect or the other can express the asymmetry between two stimuli.
Here, we addressed whether two bodies within a crowd are detected and processed more efficiently if they appear to interact than if they appear unrelated. The asymmetry between searching for a facing dyad among nonfacing dyads and searching for a nonfacing dyad among facing dyads was tested in visual search tasks, which emphasized the effect of processing either the targets or the distractors. In Experiments 1, 2, and 4, different versions of visual search shared one feature: The target always appeared in a subset of central locations. This was done to favor target detection (Carrasco, McLean, Katz, & Frieder, 1998; Neider & Zelinsky, 2008) and, thus, probe whether recruitment of attention by the facing target was stronger than by the nonfacing target. In Experiment 3, we increased the target-location uncertainty to push search through the whole array. Under these circumstances, search times tend to reflect how quickly distractors—the majority of items in the array—can be checked out. In Experiments 1 to 4, we also varied the set size to measure whether search efficiency (i.e., time spent per item) changed when subjects searched through facing versus nonfacing distractors, as we specifically predicted would be the case in Experiment 3.
Finally, in Experiment 5, we sought to unravel the mechanism beyond the putative priority of facing dyads in search. We assessed whether such advantage is mediated by perceptual grouping, the processing of two items as one attentional unit. We reasoned that because the face-to-face positioning is a perceptual cue to human interactions, two facing bodies could be processed as a unit while a scene is parsed. A similar kind of grouping mechanism has been suggested for (seemingly interacting) objects (Riddoch, Humphreys, Edwards, Baker, & Willson, 2003). Grouping resolves in more efficient processing of the configuration but may involve a cost for individuating parts within it. For example, searching for a face is more efficient than searching for a component within it, such as the mouth (Suzuki & Cavanagh, 1995), suggesting that the search mechanism accesses the highest level of representation available (to the detriment of local features). Thus, we hypothesized that the advantage of accessing a facing dyad has a cost in terms of access to its local features (the individual bodies).
In sum, the present study provided the first test of whether visual search affords a fast appraisal of relations among people in a crowd. Experiments 1 to 4 provide twofold information on the role of visual relations among bodies in visual search and the perceptual advantage of facing over nonfacing dyads. In Experiment 5, we addressed grouping of multiple bodies in visual search by targeting the expected cost of accessing a local feature (an individual) of a hierarchically higher representation (a dyad). The local ethics committee (Le Comité de Protection des Personnes Sud-Est II) approved all five experiments.
Experiment 1
Here, we set conditions to test whether, in a crowd, two facing bodies were more likely to recruit attention than the same two bodies facing away from each other.
Method
Subjects
Eleven healthy adults (10 female; age: M = 23 years, SD = 2.48) with normal or corrected-to-normal vision participated as paid volunteers after giving informed consent. The sample size was decided on the basis of previous studies that provided the background knowledge for the current research (Kaiser, Stein, & Peelen, 2014; Suzuki & Cavanagh, 1995; Wolfe, 2001). A sensitivity power analysis (G*Power 3.1; Faul, Erdfelder, Buchner, & Lang, 2009) showed that a sample size of 11 (β = 0.80, α = .05) would give us power to detect a minimal effect size (η p 2) of .293 in the critical interaction among target, distractor orientation, and set size.
Stimuli and apparatus
Stimuli were search arrays composed of 8 or 12 body dyads (see examples in Fig. 1a). Thirty-three dyads were created using 2 of 10 bodies in different poses in left and right profile, for a total of 20 bodies. Bodies were gray-scale models created with Daz3D (Daz Productions, Salt Lake City, UT) and the MATLAB image-processing toolbox (The MathWorks, Natick, MA). All body poses were biomechanically possible and suitable to a social context. By varying the poses in targets and distractors across trials and across subjects, we could test the processing of spatial relations (facing dyads vs. nonfacing dyads) across different instances of those relations, minimizing the impact of low level features of individual stimuli.

Example search arrays and results from Experiments 1 and 2. The example search arrays (a) show a trial in which the target was a facing dyad (left half of left array) and a trial in which the target was a nonfacing dyad (right half of right array). Subjects had to indicate whether the target was on the left or the right of the array. The graphs in (b) show the mean proportion of correct responses (accuracy) and response time (RT) in Experiment 1 as a function of target type (facing target or nonfacing target), set size (8 or 12), and distractor orientation (upright or inverted). The graphs in (c) show the mean proportion of correct responses and RT in Experiment 2 as a function of target type and set size. In both (b) and (c), each dot corresponds to a group mean, and each gray bar corresponds to one subject; the order of subjects is kept constant across conditions. Asterisks indicate significant effects or interactions (p ≤ .05). Error bars represent ±1 SEM within subjects (Cousineau, 2005).
Each dyad included one body oriented leftward and one body oriented rightward, and dyad members could face toward or away from each other. Distances between two bodies in a dyad were matched across facing and nonfacing dyads. Distance was defined as both the distance between the centers of the two bodies (facing: M = 210 pixels, SD = 1.47; nonfacing: M = 210 pixels, SD = 2.37), t(64) = 0.49, p > .250, and the distance between the two closest extremities of the two bodies (facing: M = 62.57 pixels, SD = 13.27; nonfacing: M = 62 pixels, SD = 13.43), t(64) = 0.17, p > .250.
Each search array was divided into two halves with a central fixation cross between them. Each half was divided into 8 cells (two columns of 4 cells each) with slightly shifted onsets along the vertical and horizontal axes, for a total of 16 cells. Four or six dyads, all facing or all nonfacing, appeared on one side of the array. The other side was the mirror version of the first. The two halves differed by only 1 cell: On either half, this cell featured a facing dyad (the target) when all other cells featured nonfacing dyads (50% of trials) or a nonfacing dyad (the target) when all other cells displayed facing dyads (50% of trials). In each array, the distractors could appear in any of the 16 cells, whereas the target could appear in only one of the eight locations around central fixation, where the greater spatial resolution of vision could favor detection (Carrasco et al., 1998).
For each subject, we created a unique set of stimuli that contained 400 arrays with a facing target among nonfacing distractors (facing-target condition) and 400 arrays with a nonfacing target among facing distractors (nonfacing-target condition). Distractors could appear upright or inverted (rotated by 180°). Inversion disrupts body representation by disrupting the spatial relations between parts but leaves unaltered all the low-level visual features (Reed, Stone, Grubb, & McGoldrick, 2006). The condition with inverted distractors served to control whether search asymmetry depended on low-level properties of the distractors or a general bias for one dyad or the other rather than the hypothesized body-to-body relationship.
In each of the two conditions, half of the trials included 8 pairs (one target among 7 distractors) and the other half included 12 pairs (one target among 11 distractors). In summary, each subject saw 800 unique arrays including a facing target among 7 or 11 upright or inverted nonfacing distractors (100 trials per condition) and a nonfacing target among 7 or 11 upright or inverted facing distractors (100 trials per condition). Images were displayed on a 17-in. CRT monitor (1,024 pixel × 768 pixel resolution, 85-Hz refresh rate) positioned 60 cm from the subject’s eyes. Individual dyads subtended approximately 1.5° of visual angle (~0.6° for a single body) and were separated by approximately 3.5° of visual angle. The array did not exceed 18° of visual angle. Stimulus presentation and response collection were controlled through the Psychophysics Toolbox extension of MATLAB (Brainard, 1997).
Procedure
Subjects sat on a height-adjustable chair, 60 cm from the computer screen, with their eyes aligned to the center of the screen. In two separate blocks, subjects were instructed to search for the only facing dyad among nonfacing dyads or the only nonfacing dyad among facing dyads and report whether the target was on the left or on the right side of the screen. They were to respond by pressing one of two keys (the “Z” key located on the left side or the “1” key located on the right side of the computer keyboard) with their left or right index finger, respectively. The key assignment (responding to left-side targets with the left key and hand and to right-side targets with the right key and hand) was the same for all subjects to avoid stimulus-response incongruence (i.e., responding to left-side targets with the right key and hand and to right-side targets with the left key and hand). This choice had no effect on the results because responses to trials with left-side and right-side targets were averaged across all experimental conditions. The order of blocks (facing target first or nonfacing target first) was alternated across subjects. Each trial began with a central fixation cross (200 ms) followed by a blank screen (700 ms) and then a search array, shown for 800 ms. After the search array disappeared, a blank screen was shown until the subject responded. The next trial began after 1,400 ms. Subjects were invited to take a break every 40 trials and in the interval between the two blocks. The experiment began with a familiarization block including 16 stimuli, 2 stimuli for each of the eight experimental conditions. The entire experiment lasted approximately 75 min.
In sum, Experiment 1 included three important features (for a summary, see Table 1). First, the target always appeared in a subset of central locations, where spatial resolution is higher than at greater eccentricities (Carrasco et al., 1998), to favor detection. In this condition, subjects may implicitly learn to focus on a subset of locations, which can further increase search efficiency (Neider & Zelinsky, 2008; Wolfe, Alvarez, Rosenholtz, Kuzmova, & Sherman, 2011). Second, detection was further promoted by short stimulus presentation, which minimized eye movements and covert shifts of attention away from the task-relevant central area (Carrasco et al., 1998; Kwak, Dagenbach, & Egeth, 1991). Third, distractors were presented upright or inverted. Inversion disrupts body representation by disrupting the spatial relations between parts but leaves unaltered all the low-level visual features (Reed et al., 2006). The condition with inverted distractors ensured that any search asymmetry found in the condition with upright distractors depended on the relationship between target and distractors rather than on low-level properties of the distractors or a general bias for one dyad or the other.
Summary of Task Settings in Experiments 1 to 5
Note: Targets could be in 1 of 8 locations around central fixation or in any of 16 locations throughout the search array. Stimuli (search arrays) stayed on the screen for 800 ms or until the subject provided a response. Subjects had to indicate whether the target appeared in the left or right side of the array or whether the target was present or absent, depending on the experiment.
Results
We computed the mean proportion of correct responses (accuracy) and the mean response time (RT) for each condition for each subject. All subjects’ mean values were within 2.5 standard deviations from the group mean; therefore, they were all included in the following analyses. We conducted 2 (target type: facing dyad, nonfacing dyad) × 2 (distractor orientation: upright, inverted) × 2 (set size: 8, 12) repeated measures analyses of variance (ANOVAs) on accuracy values and RTs.
For the accuracy analysis, the results showed a search asymmetry only in the condition with upright distractors: Accuracy was higher in a search for the facing target among nonfacing distractors than in the opposite condition (see Fig. 1b). This pattern was confirmed by a significant interaction between target type and distractor orientation, F(1, 10) = 8.43, p = .016, η p 2 = .457. A significant interaction among the three factors showed that this interaction was stronger with a set size of 8 than with a set size of 12, F(1, 10) = 30.86, p < .001, η p 2 = .755. Pairwise comparisons showed that the search advantage for the facing target among nonfacing distractors was significant only when there was a set size of 8 and upright distractors, t(10) = 4.15, p = .002 (all other ps > .11). This effect was found in 9 of 11 subjects. The overall ANOVA also showed a trend for a main effect of target type, F(1, 10) = 4.31, p = .065, η p 2 = .301, and significant effects of distractor orientation, F(1, 10) = 38.18, p < .001, η p 2 = .792, and set size, F(1, 10) = 9.29, p = .012, η p 2 = .481. Also significant was the interaction between distractor orientation and set size, F(1, 10) = 5.42, p = .042, η p 2 = .351. The interaction between target and set size was not significant, F(1, 10) = 2.27, p = .163, η p 2 = .185. All significant effects and interactions were qualified by the above three-way interaction.
For the RT analysis, we considered values from trials in which subjects provided a correct response that was within 2 standard deviations from the individual’s mean (79% of total values). No asymmetry was found. The only significant results were the main effect of distractor orientation, F(1, 10) = 31.68, p < .001, η p 2 = .760; the main effect of set size, F(1, 10) = 14.26, p = .004, η p 2 = .588; and the interaction between the two, showing that the RT difference across different set-size conditions was larger with upright than inverted distractors, F(1, 10) = 10.36, p = .009, η p 2 = .509 (see Fig. 1b). All other effects or interactions were far from significant—effect of target, Target × Distractor Orientation, and Target × Set Size: Fs(1, 10) < 1, n.s.; Target × Distractor Orientation × Set Size, F(1, 10) = 1.77, p = .213, η p 2 = .150. Given the lack of interaction between target type and set size, the search efficiency reflected by the search slope (time spent per item) was comparable in the two target-type conditions.
Discussion
Experiment 1 revealed a search asymmetry in a task in which the relation between targets and distractors concerned the relative positioning of bodies (facing vs. nonfacing) in a dyad. The asymmetry was such that subjects were more accurate in detecting a facing dyad among nonfacing dyads than in the opposite condition. This effect was not found when distractors were inverted. With inverted distractors, both targets were quite prominent, as shown by high accuracy rates. This rules out the possibility that the asymmetry found with upright distractors was due to low-level properties of the stimuli (which were preserved in the condition with inverted distractors) or to a general bias for facing dyads, independently of the relation between targets and distractors. Moreover, the asymmetry was reliable only with a set size of 8 dyads (i.e., 16 bodies); the search through 12 dyads (i.e., 24 bodies) within 800 ms was probably too difficult to reveal an appreciable effect. Indeed, the search for the facing target was better overall with a set size of 8 than 12, t(10) = 3.94, p < .003.
Both RT and accuracy showed a main effect of set size (the more distractors, the slower or less accurate the performance). This could reflect a general effect of crowding (or stimulus density), given that the area of the array and the size of the bodies were constant across set sizes. With the increase of spatial density, effects of crowding and lateral interference increase, particularly when items have high visual similarity, such as in the current case (Carrasco et al., 1998). However, the effect of set size did not interact with the target type. The absence of this interaction, particularly in RTs, implies that the search was equally efficient in the two target-type conditions, but subjects were more likely to detect or correctly recognize the facing than the nonfacing target. A possibility is that, in the short stimulus-presentation time, the facing target around central fixation could recruit attention more strongly than the nonfacing target. This interpretation is supported by prior findings showing that recognition of facing-body dyads is more efficient than recognition of nonfacing-body dyads, using a different protocol with brief and spatially certain presentations of the stimuli (Papeo, Stein, & Soto-Faraco, 2017).
Experiment 2
Here, we focused on RTs and tested whether subjects detected facing dyads better than nonfacing dyads when they had unlimited time to explore the whole array. We expected high accuracy and the asymmetry effect to shift to response latencies.
Method
Subjects
Ten healthy adults (7 female; age: M = 22 years, SD = 3.63) with normal or corrected-to-normal vision participated as paid volunteers after giving informed consent. A sensitivity power analysis (using G*Power 3.1) showed that the current test of the interaction between target and set size (sample size = 10, β = 0.80, α = .05) could detect a minimal effect size (η p 2) of .318.
Stimuli and apparatus
Using the same dyads as in Experiment 1, we constructed search arrays in Experiment 2 with a set size of either four dyads (one target and three distractors, with two dyads at each side of central fixation) or eight dyads (one target and seven distractors, with four dyads at each side of central fixation). A total of 800 unique arrays were presented to each subject: 400 with a facing target among nonfacing distractors (200 with a set size of four and 200 with a set size of eight) and 400 with a nonfacing target among facing distractors (200 with a set size of four and 200 with a set size of eight). The target was on the left side of the screen in half of the arrays and on the right side in the remaining arrays. As in Experiment 1, the target always appeared at any of eight locations around central fixation.
Procedure
Subjects were tested in the same setting as in Experiment 1. In two separate blocks, they were instructed to search for the facing dyad among nonfacing dyads and the nonfacing dyad among facing dyads, respectively. The order of blocks was counterbalanced across subjects. Task instructions were identical to those in Experiment 1: report whether the target was on the left or right of central fixation by pressing a key with the left or right index finger, respectively. In each trial, a central fixation cross (200 ms) was followed by a blank screen (700 ms) and then by the search array, which stayed on the screen until the subject responded. The time to respond was unlimited. The next trial began 1,400 ms after the response. Subjects could take a break every 40 trials and in the interval between the two blocks. The experiment began with a familiarization block of 16 trials.
In summary, Experiment 2 differed from Experiment 1 by the following features. Arrays with a set size of 12 were replaced by arrays with a set size of 4 because results of Experiment 1 suggested that subjects found it too difficult to appreciate differences in the former across conditions. Conditions with inverted distractors were discarded because Experiment 1 showed that search asymmetry occurred only when both target and distractors were upright. Moreover, closely matched facing and nonfacing dyads (see the Stimuli and Apparatus section for Experiment 1) were used to create a unique set of 800 arrays for each subject. This made it very unlikely that systematic differences in visuospatial properties of the arrays—besides the relative positioning of bodies in targets and distractors (facing vs. nonfacing)—could bias the visual search performance. Finally, in Experiment 2, search arrays were presented for an unlimited time to induce an effect in RTs. As in Experiment 1, the target was always in one of eight central locations.
Results
All subjects performed within 2.5 standard deviations from the group mean in terms of accuracy (mean proportion of correct responses) and RTs; thus, all were included in the analyses. A 2 (target type: facing dyad, nonfacing dyad) × 2 (set size: four dyads, eight dyads) repeated measures ANOVA on accuracy values yielded a main effect of target type, F(1, 9) = 7.11, p = .026, η p 2 = .441 (see Fig. 1c). This effect, reflecting higher accuracy in detecting the facing-dyad target than the nonfacing-dyad target, should be taken with caution because accuracy was overall high and variability low. The set size and the interaction between the two factors yielded no significant effects, F(1, 9) < 1, n.s.
RTs within 2 standard deviations from the individual’s mean and associated with correct responses (95% of total values) were entered in a 2 × 2 ANOVA with factors of target type and set size. A main effect of target type showed an advantage of approximately 115 ms when subjects searched for the facing dyad among nonfacing distractors, relative to the reverse condition, F(1, 9) = 8.39, p = .018, η p 2 = .482. This effect was found in 9 of 10 subjects. There was a significant effect of set size, F(1, 9) = 186.69, p < .001, η p 2 = .954, but no interaction between the two, F(1, 9) < 1, n.s. (see Fig. 1c). As anticipated by the lack of a Target Type × Set Size interaction, the search slope (time spent per item) was comparable in the two target-type conditions.
Discussion
Generally, search was visibly less efficient (slower) in Experiment 2, relative to Experiment 1. In Experiment 2, subjects had unlimited time to search, and they used that time. As a result, the effect in accuracy seen in Experiment 1 transferred to RTs in an analogous condition of Experiment 2 (set size of eight) and generalized to a new condition (set size of four); in both experiments, the performance revealed an advantage in detecting facing-dyad over nonfacing-dyad targets. Comparable search slopes for the two target-type conditions imply that search throughout the array (i.e., through facing vs. nonfacing distractors) was equally efficient. Therefore, as in Experiment 1, the asymmetry is likely to reflect stronger recruitment of attention by the facing than the nonfacing target.
Experiment 3
An asymmetry can occur because one type of distractor is easier to process than another. When this happens, the asymmetry increases with the number of distractors to check and reject. Here, we set conditions to increase the processing of distractors and tested whether search through facing distractors was more efficient than search through nonfacing distractors.
Method
Subjects
Eleven healthy adults (8 female; age: M = 22 years, SD = 4.69) with normal or corrected-to-normal vision participated as paid volunteers after giving informed consent. One subject was excluded because he was an outlier in terms of RTs (M > 2.5 SD below the group mean). The minimal detectable effect size (η p 2) for the critical interaction in the current experiment was .203 (sample size = 10, β = 0.80, α = .05).
Stimuli and apparatus
The same dyads used in Experiment 1 (except the set with inverted distractors) were used in Experiment 3. Search arrays were structured as in Experiment 1. Each half of the array contained eight cells. In 50% of arrays, one half was the mirrored version of the other. Thus, these arrays included only nonfacing dyads (facing-target absent) or facing dyads (nonfacing-target absent). In the remaining 50% of trials, the two halves of the array differed by one cell, which displayed the target on one side and a distractor on the other side: When the target was a facing dyad, all other items in the array were nonfacing dyads (facing-target present); when the target was a nonfacing dyad, all other items were facing dyads (nonfacing-target present). When present, the target could appear at any (central or more peripheral) location in the array. Set sizes varied from four to eight dyads. We matched the level of crowding across set sizes by always presenting four dyads in each half of the array. Thus, with a set size of four (50% of the trials), one half of the array was blank (the left one in 50% of trials, the right one in 50% of trials). For each subject, a unique set of arrays was created: 100 with one facing target among three nonfacing distractors, 100 with one nonfacing target among three facing distractors, 100 with four nonfacing distractors, and 100 with four facing distractors. These conditions were replicated with a set size of eight dyads, including either one target and seven distractors or eight distractors. In total, each subject saw 800 unique arrays.
Procedure
Subjects were tested in the same setting as in Experiments 1 and 2. In Experiment 3, they were instructed to report whether the target was present or absent. They were to respond by pressing one of two keys on a computer keyboard with the left or right index finger, respectively (the mapping of a key—right or left arrow—to a response—present or absent—was counterbalanced across subjects). The task consisted of two blocks, one with the demand to search for the facing dyads and the other to search for the nonfacing dyad (the order of the blocks alternated across subjects). Each trial began with a central fixation cross (200 ms) followed by a blank screen (700 ms) and then by the search array, which stayed on the screen until the subject pressed a response key. The next trial began 1,400 ms after the response. RTs and accuracy were recorded. Subjects could take a break every 40 trials and in the interval between the two blocks. The experiment began with a familiarization block of 16 stimuli, 2 stimuli for each of the eight experimental conditions. The experiment lasted about 90 min.
In sum, to promote processing of distractors during search, we introduced three changes in the paradigm, relative to Experiment 1. First, we increased target-location uncertainty by having the target appear at any location, so that it was not always available around central fixation. Second, the target was present in only 50% of trials (target present vs. target absent task). In such a task, target absence can be reported only after checking and rejecting every item in the array. To further promote this strategy, we gave subjects an unlimited time to search.
Results
The mean proportion of correct responses and mean RTs for every subject but 1 were within 2.5 standard deviations from the group mean; therefore, the following analyses included a total of 10 subjects. A 2 (target type: facing dyad, nonfacing dyad) × 2 (set size: four dyads, eight dyads) × 2 (target presence: present, absent) repeated measures ANOVA was conducted on accuracy values. The analysis yielded a main effect of set size, F(1, 9) = 48.32, p < .0001, η p 2 = .843; a main effect of target presence, F(1, 9) = 78.89, p < .0001, η p 2 = .898; and an interaction between the two, F(1, 9) = 89.26, p < .0001, η p 2 = .908, reflecting a larger difference between the two set-size conditions in trials in which the target was present, relative to those in which it was absent (see Fig. 2b). No other effect or interaction was significant—effect of target: F(1, 9) < 1, n.s.; interaction between target type and set size: F(1, 9) = 1.58, p = .240; interaction between target type and target presence: F(1, 9) < 1, n.s.; interaction among target, set size, and target presence: F(1, 9) = 1.58, p = .240.

Example search arrays and results from Experiments 3 and 4. The example search arrays (a) show a trial with a set size of 4 in which the target was present (facing dyad in left array) and a trial with a set size of 8 in which the target was absent (right array). Subjects had to indicate whether the target was present or absent. The graphs in (b) show the mean proportion of correct responses (accuracy) and response time (RT) in Experiment 3 as a function of target type (facing target or nonfacing target), set size (four or eight), and target presence (target present or target absent). The graphs in (c) show the mean proportion of correct responses and RT in Experiment 4 as a function of target type, set size, and target presence. In both (b) and (c), each dot corresponds to a group mean, and each gray bar corresponds to one subject; the order of subjects is kept constant across conditions. Asterisks indicate significant Target Type × Set Size interactions (p ≤ .05). Error bars represent ±1 SEM within subjects (Cousineau, 2005).
RTs within 2 standard deviations from the individual’s mean and associated with a correct response (90% of total values) were entered in a 2 × 2 × 2 ANOVA with the factors of target type, set size, and target presence. First, a main effect of target type showed a general search advantage for the nonfacing target among facing distractors than for the facing target among nonfacing distractors, F(1, 9) = 10.86, p = .009, η p 2 = .547. Eight of 10 subjects showed this advantage. There was a significant effect of set size, F(1, 9) = 275.43, p < .001, η p 2 = .968, and, importantly, a significant interaction between target type and set size, F(1, 9) = 13.26, p = .005, η p 2 = .596 (see Fig. 2b). In particular, subjects were faster in finding the nonfacing target among facing distractors than in the opposite condition, and this difference was more pronounced in trials with eight dyads, t(9) = 3.90, p = .004, than with four dyads, t(9) = 2.17, p = .058. Put in another way, RTs increased from the smaller to the larger set size to a lesser extent when subjects searched through facing distractors than when they searched through nonfacing distractors. Therefore, the search slope was shallower (i.e., more efficient search) for the search through facing-dyad distractors than through nonfacing-dyad distractors. Also significant were the effect of target presence, F(1, 9) = 239.90, p < .0001, η p 2 = .964, and the interaction between set size and target presence, F(1, 9) = 106.09, p < .0001, η p 2 = .922. The last effect showed that the RT increase as a function of set size was more pronounced in the target-absent condition (1,406 ms on average) than in the target-present condition (760 ms on average). This is a well-known effect when using detection tasks (present, absent) in serial search (Treisman & Gelade, 1980). The remaining effects were not significant—Target Type × Target Presence: F(1, 9) = 3.87, p = .081; Target Type × Set Size × Target Presence: F(1, 9) = 2.32, p = .162.
Discussion
Subjects were faster at finding the nonfacing than the facing target. Different search slopes for the two target-type conditions demonstrate that the above asymmetry reflected more efficient search through facing distractors than through nonfacing distractors. Indeed, the search slope, estimating the time spent per item, primarily reflected the processing of distractors that constituted the majority of items in the array. Significantly different search slopes, also when the target was absent, confirmed that subjects’ performance reflected how efficiently distractors were processed, aside from target detection. Thus, the task setting in Experiment 3 succeeded in driving search through the distractors. In this process, subjects were faster at checking and rejecting distractors when they were facing dyads than when they were nonfacing dyads. As a result of more efficient search through the facing distractors, nonfacing targets, when present, were found faster than facing targets.
Experiment 4
Experiments 1 and 2 showed asymmetric detection of facing versus nonfacing targets; Experiment 3 showed asymmetric processing of facing versus nonfacing distractors. We argue that the two outcomes are expressions of the same fundamental difference between processing facing versus nonfacing dyads, which manifested in one way or the other depending on task settings. In our view, varying the target-location uncertainty was critical to put more or less demand on processing of distractors: Higher uncertainty promoted processing of distractors during search, yielding faster search for the nonfacing target (Experiment 3). In Experiment 4, we supported this finding by making everything identical to Experiment 3, except for the target location, constrained to the central area (as in Experiments 1 and 2). With this change, we expected to restore the asymmetry in target detection seen in Experiments 1 and 2 using the task of Experiment 3.
Method
Subjects
Eleven healthy adults (7 female; age: M = 21 years, SD = 2.77) with normal or corrected-to-normal vision participated as paid volunteers after giving informed consent. One subject was excluded because he was an outlier in terms of RTs. Therefore, the minimal detectable effect size (η p 2) for the Target Type × Set Size × Target Presence interaction was .203 (sample size = 10, β = 0.80, α = .05).
Stimuli, apparatus, and procedure
All details of the stimuli, apparatus, and procedure were identical to those in Experiment 3, except for the target, which could appear in only one of the eight locations around central fixation. By minimizing the effect of eccentricity, we made the target readily available for detection.
Results
The mean proportion of correct responses and mean RTs were within 2.5 standard deviations from the group mean for every subject but 1, who was excluded from the analysis. A 2 (target type: facing dyad, nonfacing dyad) × 2 (set size: four dyads, eight dyads) × 2 (target presence: present, absent) repeated measures ANOVA was conducted on accuracy. The analysis showed an effect of target type, F(1, 9) = 7.65, p = .022, η p 2 = .459, reflecting higher accuracy in search for the facing than for the nonfacing target. The analysis also showed a main effect of set size, F(1, 9) = 26.37, p < .001, η p 2 = .745, and target presence, F(1, 9) = 66.11, p < .0001, η p 2 = .880, and interactions between target type and set size, F(1, 9) = 5.73, p = .040, η p 2 = .389; target type and target presence, F(1, 9) = 5.29, p = .047, η p 2 = .370; and set size and target presence, F(1, 9) = 23.14, p = .001, η p 2 = .720. Finally, the three-way interaction was significant, F(1, 9) = 7.27, p = .025, η p 2 = .447 (see Fig. 2c). Pairwise comparisons showed that, analogously to Experiment 1, the advantage for the facing target among nonfacing distractors was significant in the condition with a set size of eight and with the target present only, t(9) = 2.85, p = .019 (all other ps > .250). Eight of 10 subjects showed this advantage.
The same analysis was performed over mean RTs, considering values within 2 standard deviations from the individual’s mean and associated with a correct response (90.45% of total values). Results show an effect of set size, F(1, 9) = 136, p < .0001, η p 2 = .938, and of target presence, F(1, 9) = 161.32, p < .0001, η p 2 = .947, and an interaction between the two, F(1, 9) = 62.37, p < .0001, η p 2 = .874. No other effect was significant—target type: F(1, 9) < 1, n.s.; Target Type × Set Size: F(1, 9) < 1, n.s.; Target Type × Target Presence: F(1, 9) = 2.58, p = .143; Target Type × Set Size × Target Presence: F(1, 9) = 3.12, p = .111. As implied by the lack of interaction between target type and set size, the slope was comparable for the two target-type conditions.
Discussion
The pattern of effects found in accuracy analysis showed that when present and in predictable (central) locations (i.e., readily available for detection), the facing target among nonfacing detractors was correctly recognized more often than the nonfacing target among facing distractors. Hence, as expected, the asymmetry was expressed through a difference in detecting facing versus nonfacing targets. The lack of difference in search slope (i.e., interaction between target type and set size in RT) contributes to ruling out an account of the performance based on different search efficiency among facing versus nonfacing distractors. In sum, Experiment 4 confirmed that reducing target-location uncertainty abolished performance differences that could be ascribed to processing of distractors.
Experiment 5
We tested whether grouping could mediate the search benefit for facing dyads by measuring a putative cost in accessing local features of grouped dyadic representations. If facing bodies are perceptually grouped, searching for one body should be more difficult when it is a member of a facing than of a nonfacing dyad.
Method
Subjects
Eleven healthy adults (9 female; age: M = 20 years, SD = 3.20) with normal or corrected-to-normal vision participated as paid volunteers after giving informed consent. One subject was excluded because of significantly slower performance with respect to the group average. The current test had power to detect a minimal effect size (η p 2) of .199 (sample size = 10, β = 0.80, α = .05) for the critical comparison between the two conditions (target in facing dyad vs. target in nonfacing dyad).
Stimuli and apparatus
The search arrays of Experiment 5 were created analogously to those in the previous experiments and consisted of eight dyads—four on each side of the array. This time, every array included four facing and four nonfacing dyads. The target was one of two individuals: “punching body” (i.e., an individual in a punching pose) or “begging body” (i.e., an individual in a begging pose). Within an array, the same individual could appear several times; the target appeared only once, either in a facing dyad or in a nonfacing dyad, on the left side (50% of trials) or the right side (50% of trials) of the array. The two halves of the array were separated by a central fixation cross. The dyad carrying the target could appear at one of the eight locations around central fixation. Each subject saw a total of 400 unique arrays with the target in either a facing dyad (200 arrays) or a nonfacing dyad (200 arrays).
Procedure
Subjects were tested in the same setting as in Experiments 1 to 4. They performed one block with the two experimental conditions randomly interleaved. In each trial, after a central fixation cross (200 ms) was displayed, one of two words appeared for 1,400 ms to cue the target (“begging” or “punching”). After a blank screen (700 ms), the search array appeared and stayed on the screen for 800 ms. Finally, a blank screen was shown until the subject provided a response. In each trial, subjects were instructed to search for the target cued by the word and to press the left or right key (with the left or right index finger, respectively) to indicate the positioning of the target (left or right, relative to the central cross). Subjects could take a break every 40 trials. Before starting the experiment, subjects were shown the images of the target individuals for as much time as they needed and completed a familiarization block of 16 trials.
Results
All subjects but 1 performed within 2.5 standard deviations from the group mean in terms of accuracy and RTs. Therefore, the following analyses included 10 subjects. We used a pairwise t test to contrast the mean accuracy and RTs for conditions with the target in a facing versus a nonfacing dyad. The target individual was correctly detected more often when in a nonfacing than in a facing dyad, t(9) = 3.41, p = .008, η p 2 = .563 (see Fig. 3). Nine of 10 subjects showed this effect. The RT analysis, including values within 2 standard deviations from the individual’s mean and associated with a correct response (71% of total values), yielded no significant difference between the two conditions, t(9) = 1.13, p > .250 (see Fig. 3).

Example search arrays and results from Experiment 5. The example search arrays (a) show one trial in which the target was the begging body (left side of the search array) and one trial in which the target was the punching body (right side of the search array). Subjects had to indicate whether the target was on the left or the right of the array. The graphs in (b) show the mean proportion of correct responses (accuracy) and response time (RT) as a function of whether the target was in a facing dyad or in a nonfacing dyad. Each dot corresponds to a group mean, and each gray bar corresponds to one subject; the order of subjects is kept constant across conditions. The asterisk indicates a significant effect of target type (p ≤ .05). Error bars represent ±1 SEM within subjects (Cousineau, 2005).
Discussion
Here, analogously to Experiment 1, the target was presented at predictable (central) locations and the array was shown briefly. Because of the predictable target location and because, in this experiment, arrays always displayed the same number of facing and nonfacing distractors, the contribution of distractor processing to the effect was equal across the two conditions. Thus, the asymmetry in accuracy is attributable to processing of targets. In particular, we propose that individual (target) bodies were easier to detect in the nonfacing dyad than in the facing dyad because in the latter case, pairs of bodies were perceptually grouped.
General Discussion
In cluttered environments, humans show priority and particularly efficient detection and recognition of socially relevant entities such as faces and bodies. But how do humans parse a crowd, in which socially relevant stimuli are everywhere?
Experiments 1 to 4 demonstrate that during visual search, there is rapid access to relations between multiple bodies, with particularly efficient processing of facing—seemingly interacting—body dyads. Experiment 5 indicates that perceptual grouping may be a likely candidate as the mechanism underlying efficient processing of facing dyads.
The occurrence of search asymmetry between facing and nonfacing dyads suggests that the relative positioning of bodies accounts for how individuals parse people in visual environments. A search asymmetry can occur because a stimulus type carries a property that is salient or easier to detect and is absent in another stimulus (Treisman & Souther, 1985). In Experiment 1, we asked whether, in a crowd, facing dyads are more likely to recruit attention than nonfacing dyads. In this test, the target could appear in only a subset of central locations, for a short time, to minimize eye movements and covert shifts of attention away from the central area and favor detection (Carrasco et al., 1998; Wolfe et al., 2011). We found that subjects correctly detected facing targets more often than nonfacing targets. Given comparable search slopes across conditions, we conclude that the asymmetry reflected how strongly the target could recruit attention, rather than how efficiently the distractors were processed. Experiment 2, in which arrays were presented for an unlimited time, replicated this pattern in both RTs and accuracy.
Experiment 3 was designed to exploit another source of asymmetry: the speed at which distractors are checked and rejected. Different search slopes depending on the type of distractors showed that the new task successfully promoted processing of distractors. As a result, the pattern in Experiment 3 was the opposite of that in Experiments 1 and 2, with faster search for nonfacing targets among facing distractors than for facing targets among nonfacing distractors. In Experiment 4, we replicated the conditions of Experiment 3 but with target appearance constrained to central locations, as in Experiments 1 and 2. This change was sufficient to restore, in a different task, the effect seen in Experiments 1 and 2: More predictable (and central) target location favored target detection, which emphasized a stronger recruitment of attention by facing than nonfacing targets.
Although we do not know which specific feature of facing dyads accounts for the asymmetry, our results encourage the thinking that facing dyads fall in the same biologically relevant category as faces or bodies, which are stimuli associated with high visual sensitivity, rapid discrimination, and spontaneous recruitment of attention (Birmingham, Bischof, & Kingstone, 2008; Downing et al., 2004; Gobbini, Gors, Halchenko, Hughes, & Cipolli, 2013; Langton, Law, Burton, & Schweinberger, 2008; New et al., 2007). In line with this view, interacting bodies are recognized faster and better than unrelated bodies in low-visibility conditions (Papeo, Stein, & Soto-Faraco, 2017; see also Vestner, Tipper, Hartley, Over, & Rueschemeyer, 2019). Moreover, perception of interacting dyads (such as faces or bodies) has been associated with a behavioral signature of high visual sensitivity (i.e., large cost of inversion; Papeo & Abassi, 2019) and with neural selectivity in posterior temporal regions (Isik, Koldewyn, Beeler, & Kanwisher, 2017; Walbrin, Downing, & Koldewyn, 2018).
High visual sensitivity to a stimulus indexed by behavioral and neural effects predicts an attentional-perceptual advantage. The current study substantiates that prediction: Whether they were targets to detect or distractors to check and reject, facing dyads were processed more efficiently than nonfacing dyads. A mechanism that could account for particularly efficient processing of multipart objects is perceptual grouping, the automatic combination of parts into a whole (Green & Hummel, 2006; McMains & Kastner, 2010). In Experiment 5, we exposed this mechanism through one of its diagnostic features, the cost of individuating components of the group, indicated by poorer access to single elements (i.e., bodies) within the dyad.
An analogous phenomenon has previously been reported for faces: Efficient search for a face is contrasted by inefficient search for one part of that whole (e.g., the mouth; Suzuki & Cavanagh, 1995). Those findings demonstrated that visual search gives priority to a more global level of representation over a more local one.
In conclusion, various lines of research have demonstrated the attentional and perceptual benefit of faces and bodies when they appear among other entities. The present study describes a mechanism to deal with real-world crowded environments, in which faces and bodies are everywhere. Human attention has rapid access to people as well as relations among people and can use those relations to parse crowded scenarios. In this process, high priority is given to seemingly interacting bodies over noninteracting bodies. The former trigger a stronger recruitment of attention and appear to be perceptually grouped as a unit, with concurrent benefits (efficient processing of the whole) and costs (poor access to local features).
Footnotes
Acknowledgements
We thank Leonor Castro and Lisa Boinon for help with data collection and Jean-Remy Hochmann for comments on an earlier version of the manuscript.
Action Editor
Philippe G. Schyns served as action editor for this article.
Author Contributions
L. Papeo and S. Soto-Faraco developed the study concept. L. Papeo and N. Goupil designed the study and collected and analyzed the data. L. Papeo drafted the manuscript, and S. Soto-Faraco provided critical revisions. All the authors approved the final manuscript for submission.
Declaration of Conflicting Interests
The author(s) declared that there were no conflicts of interest with respect to the authorship or the publication of this article.
Funding
This study received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program (Grant Agreement No. 758473 awarded to L. Papeo). S. Soto-Faraco received support from the Ministerio de Economía y Competitividad (PSI2016- 75558-P AEI/FEDER) and Agència de Gestió d’Ajuts Universitaris i de Recerca Generalitat de Catalunya (2017 SGR 1545).
Open Practices
Data and materials for these experiments have not been made publicly available but can be requested from the corresponding author. The design and analysis plans for the experiments were not preregistered.
