Abstract
Choosing between candidates for a position can be tricky, especially when the selection test is affected by irrelevant characteristics (e.g., reading speed). One can correct for this irrelevant attribute by penalizing individuals who have unjustifiably benefited from it. Statistical models do so by including the irrelevant attribute as a suppressor variable, but can people do the same without the help of a model? In three experiments (total N = 357), participants had to choose between two candidates, one of whom had higher levels of an irrelevant attribute and thus enjoyed an unfair advantage. Participants showed a substantial preference for the candidate with high levels of the irrelevant attribute, thus choosing the less suitable candidate. This bias was attenuated when the irrelevant attribute was a situational factor, probably by making the correction process more intuitive. Understanding the intuitive judgment of suppressor variables can help candidates from underprivileged groups boost their chances to succeed.
In May 2019, the College Board, which oversees the SAT, announced the creation of an adversity score—a comprehensive score that reflects students environmental, educational, and familial differences—that will accompany students’ SAT scores and allow a “cleaner” look at their raw potential (Pascus, 2019). David Coleman, the CEO of the College Board, said that the adversity score “shines a light on students who have demonstrated remarkable resourcefulness to overcome challenges and achieve more with less. It enables colleges to witness the strength of students who would otherwise be overlooked” (para. 7).
The adversity score is an example of a correction in a potentially biased selection process performed by using a statistical model. In the current study, we examined whether people can perform such a correction intuitively, without the help of a model in similar although less complex situations. We considered the way people incorporate potentially biasing information, such as the ethnicity of a candidate, in their prediction of future performance. Specifically, we were interested in situations that involve a characteristic that influences the information used for prediction but is irrelevant to the predicted performance. In statistical terms, such a variable may serve as a suppressor variable (Azen & Budescu, 2003; Darlington, 1968; Tzelgov & Henik, 1991). Such situations are common in everyday life—when selecting between candidates for a position or a promotion or evaluating students applying for a scholarship. In some cases, these decisions are made by an algorithm (e.g., a multiple regression equation that ranks all candidates on the basis of their predicted scores), but in many cases, they are made intuitively without the aid of a statistical tool. We examined how people make intuitive decisions, compared with statistics-based decisions, in the presence of such suppressors.
Such intuitive judgments are made every day—physicians prescribing medicine, psychiatrists diagnosing patients and recommending treatments, and so on—and have meaningful consequences. Thus, the question of how well intuitive (also known as clinical) judgments fare, compared with how well statistical (also known as model-based) judgments fare, has been a subject for debate among researchers for quite some time. There is robust support for the inferiority of intuitive predictions. For example, when future success of undergraduate students is predicted on the basis of information from various sources (e.g., GRE and grade point average), previous studies have shown that intuitive judgments are less accurate than statistically based ones (Alexakos, 1966; Dawes, 1971; Hess & Brown, 1977; Highhouse & Kostek, 2013).
This inferiority has also been found in other fields. Examples include predictions of performance among soldiers in the Israeli army (Ganzach, Kluger, & Klayman, 2000) and survival time of cancer patients (Dawes, Faust, & Meehl, 1989; Einhorn, 1972), as well as differentiating between psychosis and neurosis, predicting the progression of brain dysfunction on the basis of intellectual testing, and predicting chances of winning a contest (Hogarth, Mukherjee, & Soyer, 2013).
Why do intuitive judgments not fare as well as judgments based on a statistical model? One explanation is that people are good at identifying the variables that are relevant for prediction and are even able to code them properly (Dawes, 1979). When assessing a job candidate, for example, they are able to identify the relevant information (e.g., previous experience, past job performance) and assess it properly (Hogarth et al., 2013). However, people are not as good when it comes to integrating this information and assigning weights to the various predictors and make inaccurate judgments (Dawes, 1979; Dawes et al., 1989; Kahneman, 2003; Kahneman, Slovic, & Tversky, 1982).
We examined whether people are able to accurately weigh suppressor variables that are irrelevant for the performance being predicted but affect the other predictors. The classical suppressor variable is uncorrelated with the criterion (the variable being predicted), but when entered into a multiple regression model, it receives a nonzero weight and consequently increases the proportion of criterion variance being reproduced (Conger, 1974; Darlington, 1968; Tzelgov & Stern, 1978). For example, consider a written test as a predictor of knowledge in history as part of a selection process for history teachers. The test itself measures the knowledge, but its scores are also affected by reading speed, which is in itself irrelevant for the position. Now consider two examinees, both of whom achieved the same score on the test but one of whom is a faster reader than the other. A decision based solely on the test score would suggest that both are equally qualified. However, it can be argued that the slower reader is the more knowledgeable and should be preferred. This will be reflected in a regression model predicting knowledge in history (y) according to the test score and reading speed with a negative coefficient of reading speed: y = β1 × Test Score – β2 × Reading Speed + ε. Adjusting the test scores for reading speed will represent knowledge in history more accurately and will lead to a more optimal decision. (For a statistical description of the suppressor-variable effect, see Section 1 in the Supplemental Material available online.)
Many selection processes may also measure irrelevant aspects that have the potential to be suppressor variables. Examples include achievement tests that may benefit privileged groups or discriminate against minorities and interviews that are biased by verbal abilities. Unless adjusted, such biases will affect the validity of the tests. Effectively, one needs to “penalize” the candidate who is higher on the irrelevant attribute. Such adjustments may be difficult to perform intuitively without a statistical model because of their counterintuitive nature.
Human behavior in such situations has not been studied. Better understanding the way irrelevant information is considered in prediction tasks in general and specifically in selection processes is valuable. It can also help groups that are discriminated against because of the use of certain selection methods receive a fair chance not as a result of affirmative action but rather because they are higher on the relevant and more important attribute.
Current Research
The current experiments tested the intuitive understanding of attributes that affect selection processes but are irrelevant to the situation. In three experiments, participants saw different selection scenarios and had to choose between two candidates for a certain position. An attribute was described as affecting the selection test and was either relevant or irrelevant (i.e., a potential suppressor) to the position, and one of the candidates was higher on this attribute and enjoyed a biased advantage. In Experiment 1, we manipulated the score achieved by both candidates on the selection test. In Experiment 2, we manipulated the extent to which the selection test was affected by the irrelevant attribute (i.e., strength of suppression) and tested whether participants were sensitive to the phenomenon. In Experiment 3, the irrelevant attribute was not a characteristic of the candidate per se (e.g., reading speed) but rather a situational factor that affected performance (e.g., background noise).
We hypothesized that participants would fail to understand the complexity of a suppressor variable and would make “wrong” decisions (decisions different from those suggested by a statistical model) compared with situations that did not have the potential for a suppressor variable. We also tested whether thinking style and numeracy affected participants’ understanding. For full data and materials for all three experiments, see our Open Science Framework project (https://osf.io/wt4e5/).
Experiment 1
Method
Participants
One hundred twenty-one participants (77 men; mean age = 40 years, SD = 12.01) received $0.50 for answering the questionnaire on Amazon Mechanical Turk (MTurk). All participants were native English speakers. Sample size was determined by an a priori power analysis using G*Power software (Version 3.1; Faul, Erdfelder, Lang, & Buchner, 2007) with a .05 criterion of statistical significance, power of .80, and an effect size (f) of .25.
Apparatus and procedure
We told participants they would read several scenarios and would be asked to answer questions after reading each scenario. Each scenario described a selection process for a specific position: a managerial position, a truck driver, or an online desk operator. The selection process in each of these scenarios included a selection test and a suppressor variable that, although irrelevant for the success in the position, affected the selection test.
Participants had to choose between two candidates, one who was high on the suppressor (HOS candidate) and one who was low on the suppressor (LOS candidate). For example, the managerial scenario read as follows: You are looking for candidates for a managerial position in your company. As part of the screening process you use a computer simulation to measure managerial capabilities. The simulation favors those candidates who have previous experience working with computers. You have to choose between two candidates with identical scores in the simulation: John who has a lot of experience working with computers and Mark who has almost no experience working with computers. Assuming the ability to work with computers plays no role in being a successful manager, who would you prefer?
For details regarding the other two scenarios, see Section 2.1 in the Supplemental Material.
There were three versions. Each version consisted of a different combination of scenarios and conditions. Each participant was randomly assigned to one of the three versions. Completion of the scenarios and questionnaire took approximately 7 min.
Design
For each scenario, we manipulated the scores achieved by the two candidates on the selection test. In Condition A, the two candidates’ scores were equal. In Condition B, the HOS candidate’s score was greater than the LOS candidate’s score. In Condition C, the HOS candidate’s score was lower than the LOS candidate’s score. We also included a fourth condition (Condition D) in which both candidates achieved equal scores but the other attribute was relevant for the position (i.e., not a suppressor). Table 1 shows the design.
The Four Conditions Used in Experiment 1
Note: Candidates were either high on the suppressor (HOS candidate) or low on the suppressor (LOS candidate).
Our main interest was Condition A. Because an irrelevant attribute gave the HOS candidate an advantage on the selection test and given that both candidates scored the same on the selection test, the relevant ability of the LOS candidate for the position was actually higher. Consequently, the LOS candidate should be chosen over the HOS candidate. If all variables were included in a regression model, the LOS candidate would be predicted to have a higher score on the predicted successes in the position.
Conditions B and C served as controls for Condition A. In Condition B, the HOS candidate achieved a higher score in the test than the LOS candidate. Thus, it was difficult to decide between the candidates, although in Condition C, it should have been easy to see that the LOS candidate should be selected. This candidate did not have the (irrelevant) advantage that the HOS candidate had, and the candidate achieved a higher score on the test and overcame this disadvantage.
Condition D was an additional control to Condition A, in which both candidates scored identically but there was no suppression—the third variable was relevant to the position and therefore served as another predictor without contaminating the selection process.
The within-participants design included the four conditions for each participant. For Conditions A through C, we had three different scenarios. The scenario for Conditions A and D was always identical, as was the order of conditions in each version. We had three versions that were manipulated as a between-participants variable. Each version included all four conditions and all three scenarios.
Measures
After each scenario, participants had to choose which of the candidates they would hire (the HOS candidate, the LOS candidate, no preference) and rate each candidate’s chance of success in the position (1–100). They were also asked how important the irrelevant or relevant attribute presented in the scenario was to the position (rated from 1 to 7). The last question enabled us to examine whether participants noticed that the third variable was irrelevant and whether those who noticed took that into consideration while making their choices.
Results
To make sure that participants were aware of the fact that the irrelevant attribute was irrelevant for the position (and thus a suppressor), for each scenario, we analyzed participants’ answers to the question that asked them to rate the importance of the irrelevant attribute. For Conditions A through C, a rating of 1 or 2 was considered appropriate, whereas for Condition D, a rating of 6 or 7 was considered appropriate (in this condition, the irrelevant attribute was described as relevant for the position and, thus, not a suppressor). Ninety participants (74%) rated the importance of the irrelevant attribute as irrelevant, which indicated that they noticed this detail, which is essential for considering this variable as a suppressor. We performed all analyses twice, once on all participants and once including only the subset of participants who correctly reported the importance of the irrelevant variable. These analyses led to the same conclusions; thus, we present the analysis for all participants.
Candidate selection
A generalized linear mixed model (GLMM) multinomial logistic regression was performed to compare participants’ preferences within each condition separately 1 (for detailed analysis, see Section 3.1 in the Supplemental Material). Figure 1 presents participants’ choices in each condition across all scenarios (the pattern of results was the same when the scenarios were analyzed separately).

Percentage of participants in Experiment 1 who chose the candidate high on the suppressor (HOS), chose the candidate low on the suppressor (LOS), and had no preference between candidates (none) in each of the four conditions, across all scenarios. Error bars represent ±1 SE. Asterisks indicate significant differences between conditions (**p < .01).
In Condition A, participants preferred the HOS candidate over the LOS candidate, b = −1.22, p < .001, 95% confidence interval (CI) = [−1.8, −0.64], although a statistical model would have suggested the opposite. Furthermore, 46% were indifferent between the two candidates. This choice is also suboptimal—because both candidates achieved identical scores and the HOS candidate was higher on the suppressor, the LOS candidate should be preferred. These results may suggest that participants weighted the irrelevant attribute as a “bonus” or simply chose to ignore it and failed to correct for the bias it induced.
Results of Conditions B and C suggest that participants relied mostly on the score of the selection test. In Condition C, the preference for the LOS candidate, b = 0.97, p < .001, 95% CI = [0.49, 1.44], corresponded to the normative model because that candidate was lower on the irrelevant attribute and yet scored higher on the test. In Condition B, there was not enough information to determine who was the right choice, but the strong preference for the HOS candidate, b = −2.5, p < .001, 95% CI = [−3.29, −1.73], suggests that participants combined the test score and the irrelevant attribute with positive weights to make their decision.
In Condition D, both candidates scored the same on the selection test, and the other attribute (suppressor) was relevant to the position. Thus, if this attribute had a unique contribution over the selection test, there should be a clear preference for the HOS candidate (i.e., the candidate high on the other variable). Yet because there was not enough information in the scenario to determine whether this attribute indeed had a unique contribution, the normative prediction was equivocal (for a statistical derivation, see Section 3.1.2 in the Supplemental Material). Indeed, most participants showed a strong preference for the HOS candidate, b = −2.6, p < .001, 95% CI = [−3.38, −1.83], as would be expected if the other attribute had a unique contribution for success in the position.
To summarize, it appears that participants relied on the selection-test score to make their decision (as supported by results of Conditions B and C). When the test scores were identical, participants used the information regarding the other attribute as a bonus (Conditions A and D). Moreover, in Condition C, in which the test score and the irrelevant attribute conflict, participants relied on the test score.
It can be argued that when participants in Condition A rated the suppressor variable as irrelevant, it was due to a demand effect and thus did not reflect the participants’ real beliefs regarding its importance. If this were the case, we would expect no differences between Conditions A and D, for which the irrelevant attribute was, in fact, relevant. The difference found between preferences in those two conditions for the HOS candidate (42% and 79%) suggests that this does not explain the results.
Chances of success
Participants also assessed each candidate’s chances of success in the position. The purpose of this question was to test whether the choice made by the participants reflected their perception of who was a better fit or whether the choice was affected by other factors, such as affirmative action. The pattern of results was similar to the one reflected in participants’ choices and therefore was not analyzed further.
Experiment 2
Experiment 1’s findings support our main hypothesis that people lack the intuitive understanding of a suppressor variable and thus make incorrect decisions in selection processes. Experiment 2 was designed to replicate these findings and to test whether the strength of suppression affects people’s judgments (this factor was not manipulated in Experiment 1). In addition, we measured individual differences in cognitive-thinking style and numeracy to examine their relation to understanding of the suppressor variable.
Method
Participants
One hundred twenty-four participants (79 men; mean age = 34.2 years, SD = 10.1) received $0.80 for answering a questionnaire on MTurk. One hundred twenty-two were native English speakers. Sample size was determined by an a priori power analysis, as in Experiment 1.
Apparatus and procedure
We used the same three scenarios as in Experiment 1 to create three conditions but made two changes: (a) Both candidates achieved the same score in all three conditions, and (b) we manipulated the strength of suppression (no suppression, weak suppression, strong suppression) by providing information regarding the extent to which each selection test was affected by the irrelevant variable (not affected, slightly affected, strongly affected). In all three conditions, the suppressor variable was irrelevant to the position (unlike Condition D in Experiment 1). We had three versions that were manipulated as a between-participants factor. Each version included all three conditions and all three scenarios (for the complete scenarios, see Section 2.2 in the Supplemental Material).
Measures
Participants had to choose which of the candidates they would hire, rate each candidate’s chance of success in the position (1–100), and rate how important the suppressor variable was to the position (1–7).
Individual differences
To assess the tendency of people to rely on rational thinking and intuitive thinking, we used the 10-item Rational-Experiential Inventory (REI) short form (Epstein, Pacini, Denes-Raj, & Heier, 1996). The inventory consists of two unipolar and independent subscales—Need for Cognition and Faith in Intuition. Participants had to rate their agreement with each statement on a 5-point scale.
We also measured participants’ numeracy using an eight-questions test (Soll, Keeney, & Larrick, 2013). Each question presented a numerical problem, and participants had to choose from seven possible answers. We tested whether people who scored higher on numeracy were more sensitive to the effect of suppression.
Procedure
Each participant was randomly assigned to one of the three versions. After answering the four questions following each scenario, participants completed the REI questionnaire and the numeracy test and were asked about their gender and age. Completion of the scenarios and questionnaire took approximately 9 min.
Results
As in Experiment 1, we wanted to make sure that participants noticed that the irrelevant attribute was irrelevant for the position and thus a suppressor. We therefore examined participants’ ratings on the question by asking them for the importance of this attribute to the position, according to the scenarios. In all three conditions, an appropriate rating would be 1 or 2 because in all three of them, the irrelevant attribute was described as irrelevant. Eighty-nine participants (70%) rated the importance as irrelevant. We performed all analyses twice, once including only this group of participants and once including all participants. These analyses led to the same conclusions, and thus, we present the analyses based on all participants.
Candidate selection
A GLMM multinomial logistic regression was performed to compare participants’ preferences within each condition separately (for detailed analysis, see Section 3.2.1 in the Supplemental Material). Figure 2 presents participants’ choices for each condition across all scenarios (the pattern of results was the same when each scenario was analyzed separately).

Percentage of participants in Experiment 2 who chose the candidate high on the suppressor (HOS), chose the candidate low on the suppressor (LOS), and had no preference between candidates (none) in each of the three suppression conditions, across all scenarios. Error bars represent ±1 SE. Asterisks indicate significant differences between conditions (**p < .01).
In both the weak- and strong-suppression conditions, participants’ preference was different from that predicted by a normative model. In these conditions, the LOS candidate should be preferred because this candidate was lower on the irrelevant attribute and therefore higher on the relevant one. However, the LOS candidate was chosen by only 19% of the participants in the weak-suppression condition (b = −0.74, p = .004, 95% CI = [−1.24, −0.23]) and 16% of the participants in the strong-suppression condition (b = −0.85, p = .002, 95% CI = [−1.38, −0.33]). More participants, 39% in the weak-suppression condition and 38% in the strong-suppression condition, chose the HOS candidate. In the no-suppression condition, we expected there to be no preference for one of the candidates over the other because the other attribute did not affect the selection test and thus the test score was unbiased. Although about half of the participants (56%) chose accordingly, among the other participants, there was still a preference for the HOS candidate over the LOS candidate, b = −1.39, p < .001, 95% CI = [−2.05, −0.71]. There was no significant Suppression Strength × Candidate interaction, χ2(4, N = 124) = 6.8, p = .15, suggesting that the strength of suppression did not affect participants’ judgments. However, this finding was reevaluated when thinking style was taken into consideration.
Chances of success
As in Experiment 1, participants’ assessments of success chances were consistent with their choices, supporting the idea that their decision reflects their beliefs of who has better chances of succeeding in the position, and therefore were not analyzed further.
Individual differences
Faith in Intuition, Need for Cognition, and numeracy scores were uncorrelated (see Table S4 in the Supplemental Material). Because the optimal choice was different for the condition with no suppressor variable and the conditions with a suppressor variable, we analyzed them separately. In a multinomial logistic regression, we entered Faith in Intuition, Need for Cognition, and numeracy as predictors of whether participants would choose the LOS candidate over the HOS candidate. No effects were found in the no-suppression condition (b = −0.18, p = .62, 95% CI = [−0.88, 0.53]) or in the weak-suppression and strong-suppression conditions (b = −0.319, p = .075, 95% CI = [−0.82, 0.04]). However, the pattern of results suggests that the more participants relied on their intuitive thinking, the less likely they were to choose the LOS candidate over the HOS candidate. For demonstration purposes, we divided participants on the basis of whether their Faith in Intuition score was above or below the median. Figure 3 presents participants’ choices in the three conditions according to their Faith in Intuition score. The erroneous preference for the HOS candidate (over the LOS candidate) was higher among those high on intuitive thinking compared with those low on intuitive thinking.

Percentage of participants in Experiment 2 who chose the candidate high on the suppressor (HOS), chose the candidate low on the suppressor (LOS), and had no preference between candidates (none) in each of the three suppression conditions, separately for participants high and low on intuitive thinking (according to a median split). Intuitive thinking was indexed by Faith in Intuition scores. Error bars represent ±1 SE.
The difference between the LOS candidate and the HOS candidate under both suppression conditions was 12% and 13%, respectively, for participants low on intuitive thinking, whereas for those high on intuitive thinking, it was 28% and 32%, respectively. No effects were found for Need for Cognition and numeracy (for detailed analysis, see Section 3.2.2 in the Supplemental Material). This pattern of results may point to the role intuition plays in understanding the effect of suppression.
Experiment 3
Experiments 1 and 2 show that when an irrelevant attribute—that in itself stands as a positive attribute—affects a selection test and gives one of the candidates an irrelevant and potentially biased advantage, participants have difficulty correcting for this advantage and choose the less suitable candidate. We associate this lack of ability to perform the appropriate correction to the fact that it is counterintuitive to “punish” someone for having a positive attribute. Would participants be able to make the right decisions if the irrelevant attribute were something that is in line with their intuitions? In Experiment 3, we made the irrelevant attributes a feature of the situation and not a property of the candidate. We expected these types of attributes to be more easily perceived as something that needs to be controlled for.
Method
Participants
One hundred twelve participants (67 men; mean age = 33.2 years, SD = 9.9) received $0.80 for answering a questionnaire on MTurk. One hundred ten were native English speakers. Sample size was determined by an a priori power analysis, as in Experiment 1.
Apparatus and procedure
We used the same three scenarios as in Experiments 1 and 2, but the irrelevant attribute that affected the selection test was a situational factor rather than a characteristic of the candidate. For example, in the managerial-position scenario, participants were told that during one of the candidate’s simulation, there was a background noise from the next room, which affects performance in this kind of simulation tests, yet both achieved the same score. Such a factor can be a suppressor because it affects the selection test but is obviously irrelevant for the position, and unlike in Experiment 2, here it was much easier to understand the noise’s irrelevancy and the need for correction (for detailed description of the scenarios, see Sections 2.3 and 2.4 in the Supplemental Material). We manipulated the strength of suppression and combined the different conditions and scenarios into three between-participants versions. In this experiment, the LOS candidate was the candidate whose performance was affected by the irrelevant factor (e.g., suffered background noise during his test), and the HOS candidate was the candidate who was not affected by this factor. Therefore, if the LOS candidate’s performance was affected by this irrelevant factor and yet he or she achieved the same score, that candidate should be preferred for being able to overcome this (irrelevant) disadvantage.
Measures and procedure
Measures and procedure were identical to those used in Experiment 2 except for the fact that as a manipulation check we asked participants to what degree the irrelevant attribute affected the selection method, according to the scenario presented. It made little sense to ask them about the relevancy of the irrelevant attribute to the position, as in Experiments 1 and 2.
Results
In this experiment, the irrelevant attribute was circumstantial (e.g., background noise) and thus was intuitively irrelevant for future performance. We therefore did not assess participants’ awareness of its irrelevance as we did in Experiments 1 and 2. Here, we made sure that participants were aware of the different possible levels of suppression, which changed across conditions (none, weak, strong). Approximately 85 participants (76%, differed across conditions) rated the effect of the irrelevant attribute on the selection method appropriately (≤ 2 in Condition A, 3–5 in Condition B, and ≥ 6 in Condition C). We performed all analyses twice, once including only these participants and once including everyone. These analyses led to the same conclusions; thus, we present the analysis using all participants.
Candidate selection
A GLMM multinomial logistic regression was performed to compare participants’ preferences within each condition separately (for detailed analysis, see Section 3.3.1 in the Supplemental Material). Figure 4 presents participants’ choices for each condition across all scenarios (the pattern of results was the same when we analyzed the scenarios separately).

Percentage of participants in Experiment 3 who chose the candidate high on the suppressor (HOS), chose the candidate low on the suppressor (LOS), and had no preference between candidates (none) in each of the three suppression conditions, across all scenarios. Error bars represent ±1 SE. Asterisks indicate significant differences between conditions (*p < .05, **p < .01), and the dagger represents a marginally significant difference between conditions (p < .10).
Participants were more likely to make the normative decisions in all conditions in this experiment compared with Experiments 1 and 2. In the no-suppression condition, because both candidates achieved identical scores on the selection test, the right choice was no preference because the irrelevant factor did not affect participants’ performance. Indeed, the percentage of participants who were indifferent between the two candidates (55%) was significantly higher than the percentage of participants who chose one of the candidates (LOS candidate: b = 0.53, p = .007, 95% CI = [0.11, 0.94]; HOS candidate: b = 1.4, p < .001, 95% CI = [0.83, 1.97]). In the two suppression conditions, the right choice was the LOS candidate because this candidate was able to overcome an inferior starting point. In the weak- and strong-suppression conditions, the percentage of participants who chose the LOS candidate (44% in both conditions) was higher than the percentage of participants who chose the HOS candidate: 21% in the weak-suppression condition (b = 0.71, p = .005, 95% CI = [0.22, 1.21]) and 30% in the strong-suppression condition (b = 0.40, p = .08, 95% CI = [0.05, 0.84]). These results support our hypothesis that if the suppressor variable is not a characteristic of the candidate but a situational factor, the correction needed probably conforms with participants’ intuitions and they are better able to perform it. Nevertheless, only about half of the participants performed well under these conditions, which again demonstrates the counterintuitive nature of the correction needed.
Chances of success
Unlike in Experiment 2, participants’ assessments of success chances were a little different from their choices. No difference was found in the success assessments of the different candidates, nor was there an effect of suppression strength. A possible explanation for this inconsistency is that participants understood that the performance of the LOS candidate was negatively affected by the irrelevant factor, but because this factor was a situational one and did not characterize the candidate, they did not consider it as something relevant for future performance and thus assessed both candidates’ future chances of success as equal.
Individual differences
As in Experiment 2, we analyzed whether Faith in Intuition, Need for Cognition, and numeracy affect participants’ tendency to choose the LOS candidate over the HOS candidate. Here, too, no correlations were found among Faith in Intuition, Need for Cognition, and numeracy scores (see Table S7 in the Supplemental Material). Thus, we examined their ability to predict the choice of the LOS candidate over the HOS candidate in a multinomial logistic regression. In the no-suppression condition, no effect was found for either variable (as in Experiment 2). However, for both suppression conditions, we found an effect for Need for Cognition and numeracy (Need for Cognition: b = 0.88, p = .003, 95% CI = [0.30, 1.37]; numeracy: b = 0.20, p = .04, 95% CI = [0.01, 0.38]). The chances of choosing the LOS candidate over the HOS candidate increased as a function of the participants’ scores on the rational thinking or numeracy scales. Figures 5 and 6 demonstrate these effects by presenting participants’ choices for all three conditions according to whether their score on the rational thinking and on the numeracy scales, respectively, was above or below the median. As can be seen, under weak suppression and strong suppression, for participants high on Need for Cognition and numeracy, the difference between preference for the LOS candidate and the HOS candidate ranges between 29% and 39%, whereas for those low on both scales, the difference is 12% or less. (No effect was found for Faith in Intuition; for detailed analysis, see Section 3.3.2 in the Supplemental Material.)

Percentage of participants in Experiment 3 who chose the candidate not affected by an irrelevant factor (HOS), chose the candidate affected by an irrelevant factor (LOS), and had no preference between candidates (none) in each of the three suppression conditions, separately for participants high and low on rational thinking (according to a median split; groups have different ns because of the scale’s form of distribution). Rational thinking was indexed by Need for Cognition scores. Error bars represent ±1 SE.

Percentage of participants in Experiment 3 who chose the candidate not affected by an irrelevant factor (HOS), chose the candidate affected by an irrelevant factor (LOS), and had no preference between candidates (none) in each of the three suppression conditions, separately for participants with high and low numeracy scores (high: correct answers ≥ 6; low: correct answers ≤ 5). Error bars represent ±1 SE.
Discussion
The inferiority of intuitive decision making compared with statistical models in selection processes has been previously demonstrated (Ægisdóttir et al., 2006; Dawes, 1971; Dawes et al., 1989; Ganzach et al., 2000; Grove, Zald, Lebow, Snitz, & Nelson, 2000; Highhouse & Kostek, 2013). The present study extended the generality of this pattern to situations that involve an attribute that affects the selection process but is irrelevant to the position. Such situations have not yet been studied despite being common in everyday life. In three experiments, we presented participants with two candidates for a position, and one candidate had an attribute that offered an advantage on the selection test, but this attribute was irrelevant for the position, and therefore the advantage biased the process in the candidate’s favor. If all attributes were entered as variables to a regression model, the irrelevant attribute would be assigned a negative weight to correct for this biased advantage.
By asking participants to choose between the two candidates, we were able to test whether they understood that the suppressor variable, despite being irrelevant for the position, affected the selection test and thus needed to be corrected for. The results support our hypothesis that when the irrelevant attribute is perceived as a valuable attribute, it is highly counterintuitive to make the necessary correction. This is especially prominent when both candidates achieve the same score on the selection test. If one of the candidates has an attribute that gives an unjustified advantage, the other candidate should be preferred because that candidate is higher on the relevant aspect that the test measures.
These finding are in line with Dawes’s suggestion (Dawes, 1979; Dawes et al., 1989) that people can identify the relevant factors but fail to aggregate them appropriately. It appears that participants considered the score achieved on the selection test and the information regarding the irrelevant attribute as important, but when this information countered their intuitive understanding, they failed to aggregate it correctly.
One can think of three possible ways by which a participant can treat a suppressor variable. First, it could be treated as a bonus (i.e., as if it has a positive coefficient). Second, it could simply be ignored (i.e., a coefficient of 0). Third, it might be treated as having a negative—but insufficiently negative—impact. The fact that 46% of participants in Experiment 1 and 43% and 46%, respectively, in the weak-suppression and strong-suppression conditions in Experiment 2 were indifferent between the two candidates partially supports the possibility that participants simply ignored the suppressor variable. The results of Experiment 2 suggest that when the suppressor variable characterizes the candidate, people seem to treat the suppressor variable as a bonus and thus tend to prefer the candidate high on the suppressor variable regardless of its strength as a suppressor. The results of Experiment 3 suggest that when the irrelevant attribute is a characteristic of the situation and thus easier to correct, the suppressor has a negative impact (it was easier to understand that background noise was an attribute with a negative effect that needed to be controlled for, unlike the candidate’s previous experience with computers). However, even in this experiment, almost half of the participants did not choose optimally. Yet, on the basis of the current data, it is difficult to estimate the extent to which this negative weight is insufficient. Another limitation is that although a great majority of participants were aware of the fact that the irrelevant attribute was not important for performance, it is not clear whether they understood the statistical meaning of irrelevance.
Further research is needed to estimate the bias more precisely. The findings support the idea that the thinking process in such situations is based mostly on intuitions. When the correction process is counterintuitive, the more it is necessary, the more biased the decision-making process will be.
Moreover, findings regard thinking style and numeracy align with this explanation, at least to some degree. In Experiment 2, in which the effect of the suppressor variable was highly counterintuitive, intuitive thinking seemed to moderate participants’ performance, but in Experiment 3, in which the effect of the suppressor variable was more in line with participants’ intuitions, intuitive thinking did not play a role. The effect of rational thinking and numeracy may suggest that regardless of intuitive thinking (all three variables were uncorrelated), some degree of analytic thinking was needed to aggregate the information presented into a decision that would be in line with the normative model. Nevertheless, these effects were moderate (especially the effect of intuitive thinking), and more research is needed to understand the role of cognitive thinking, perhaps using better methods to evaluate participants’ thinking styles and analytic abilities.
Conclusion
The results of the present experiments support our hypothesis that when choosing between candidates using a selection test that is affected by an irrelevant attribute, people fail to grasp the complexity at hand and make wrong judgments, choosing less suitable candidates. They are, however, able to make the right decision more frequently when the attribute is not a characteristic of the candidate. Such findings are of great importance. In the current study, we examined examples for potential irrelevant attributes, but such attributes can be found in a variety of selection methods, such as tests biased by cultural background, interviews that benefit English speakers, or simulations that benefit men over women. Acknowledging the existence of such characteristics and learning how to overcome the complexities that they present to the decision maker can help underprivileged groups have a fair chance based on their true abilities rather than affirmative action.
Supplemental Material
Rabinovitch_OpenPracticesDisclosure_rev – Supplemental material for Achieving More With Less: Intuitive Correction in Selection
Supplemental material, Rabinovitch_OpenPracticesDisclosure_rev for Achieving More With Less: Intuitive Correction in Selection by Hagai Rabinovitch, Yoella Bereby-Meyer and David V. Budescu in Psychological Science
Supplemental Material
Rabinovitch_Supplemental_Material_rev – Supplemental material for Achieving More With Less: Intuitive Correction in Selection
Supplemental material, Rabinovitch_Supplemental_Material_rev for Achieving More With Less: Intuitive Correction in Selection by Hagai Rabinovitch, Yoella Bereby-Meyer and David V. Budescu in Psychological Science
Footnotes
Acknowledgements
This study was part of H. Rabinovitch’s doctoral dissertation, which was done under the supervision of Y. Bereby-Meyer.
Transparency
Action Editor: Timothy J. Pleskac
Editor: D. Stephen Lindsay
Author Contributions
H. Rabinovitch was the lead author of the manuscript. He ran the experiments, collected the data, and ran the analyses. All of the authors developed the research concept and contributed equally to the study design. All of the authors approved the final manuscript for submission.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
