Abstract
A classic problem in the literature on authority is that those with the power to enforce cooperation and proper norms of conduct can also abuse or misuse their power. The present research tested the argument that concerns about legitimacy can help regulate the use of power to punish by invoking a sense of what is morally right or socially proper for power-holders. We tested this idea in a laboratory experiment using public goods games in which one person in each group was selected to be a “designated punisher” who could give out material punishment that was either costly or costless to the punisher. Results show that costly punishment is perceived as more legitimate (proper) than costless punishment and that designated punishers engaged in more proper (“prosocial”) punishment and less abusive (“antisocial”) punishment when punishment was costly. These results highlight the importance of legitimacy in both motivating and regulating the enforcement of cooperation.
Imagine witnessing a neighbor who failsto pick up after his dog, a colleague who leaves the coffee pot empty, or teenagers littering in the park. Part of our civic duty—as members of organizations, communities, or societies—is to help enforce basic norms of conduct. Yet, whether people actually do so depends on a number of issues, including the cost of enforcement (Is it worth my time to confront them or alert the authorities?) and people’s role (Is it appropriate for me or someone else to intervene?). Indeed, the enforcement of cooperation has long puzzled social scientists, for no rational actor should voluntarily exercise punishment in the sole interest of producing public goods that would be available to punishers and non-punishers alike (Fehr and Gächter 2000; Heckathorn 1990; Yamagishi 1986). Yet, controlled experiments have amassed abundant evidence that people punish each other significantly more than is often theorized (Balliet, Mulder, and Van Lange 2011; Chaudhuri 2011), suggesting that peer sanctioning may play a critical role in the success and survival of groups, communities, and institutions (Bowles and Gintis 2011; Ostrom 1990).
Recently, scholars have begun to question the efficacy of peer sanctions. Some have noted that peer sanctioning can sustain cooperation and increase efficiency only under favorable conditions; often, peer sanctions result in retaliation or miscoordination, creating excessive punishment (Egas and Riedl 2008; Gächter, Renner, and Sefton 2008; Heckathorn 1990; Rand and Nowak 2011). Others have argued that personally punishing wrongdoers is less common outside of laboratory studies than presumed by theory (Baldassarri and Grossman 2011; Kiyonari and Barclay 2008). These concerns have led scholars to consider enforcement by designated punishers. By restricting the power to punish to a solitary member (e.g., leaders, authority figures, regulatory agents), designated punishment has the potential to curtail excessive punishment. Studies have shown that designated punishment is sufficient to sustain cooperation without reducing overall collective welfare (Balliet et al. 2011; O’Gorman, Henrich, and Van Vugt 2009).
Designated punishment presents at least two critical weaknesses, however. First, it is vulnerable to abuse of power. For instance, designated punishers may engage in antisocial punishment (sanctioning members who cooperate more than punishers; Herrmann, Thoni, and Gächter 2008) rather than prosocial punishment (punishing defectors). While peer punishers may engage in antisocial punishment out of retaliation, designated punishers may engage in antisocial punishment because targets cannot retaliate. Second, vesting one person with the sole responsibility of enforcement also carries the risk that the punisher may not engage in enough enforcement (Devlin-Foltz and Lim 2008).
The present research builds on the idea that concerns about legitimacy can help regulate these patterns because legitimacy derives in part from propriety, namely, personal beliefs about what is normatively desirable, appropriate, or correct (Dornbusch and Scott 1975; Hegtvedt and Johnson 2009), such as how to exercise power in ways that meet the social approval of peers and subordinates (Biggart and Hamilton 1984; Blau 1964; Zelditch and Walker 2000). Supporting this argument, our public goods experiment shows that designated punishers engage in more proper (prosocial) punishment and less improper (antisocial) punishment when punishment is costly rather than costless to punishers. This is because compared to costless punishment, costly punishment is seen as more proper, which compels punishers to use it in more fair and appropriate ways. This pattern stands in stark contrast to the standard economics of costly punishment, which suggests less punishment of any kind as the cost of punishment increases (Egas and Riedl 2008; Horne and Cutlip 2002). Overall, by integrating research on the sociology of power and legitimacy (Hegtvedt and Johnson 2009) and the behavioral economics of public goods games (Fehr and Gächter 2000), our research advances our understanding of how concerns about legitimacy can help both motivate and regulate the enforcement of cooperation.
Legitimacy
Legitimacy is a matter of critical importance for authority. Legitimacy refers to perceptions about what others see—apparently or presumably—as proper and valid in given situations or roles (Dornbusch and Scott 1975). Propriety concerns personal beliefs about what is normatively right (e.g., what others view as fair or appropriate; Johnson, Dowd, and Ridgeway 2006; Zelditch 2001), whereas validity refers to the degree to which a person feels obliged to obey or comply with the demands of the situation (e.g., norms, rules, roles) as “matters of objective fact” (Zelditch and Walker 1984:219). Together, legitimate entities and acts are widely accepted and “taken for granted” as normative, fair, and appropriate while illegitimate ones are met with scrutiny, disapproval, and resistance (Zelditch and Walker 1984). Without legitimacy, enforcement is less effective, potentially backfiring by provoking resentment and reducing compliance (Baldassarri and Grossman 2011; Fehr and Rockenbach 2003). Without legitimacy, people with power are also more likely to alienate others and lose their support by engaging in overly aggressive or antisocial behaviors that deviate from proper norms of conduct (Fast, Halevy, and Galinsky 2012; Zelditch 2001; Zelditch and Walker 2000). Despite its obvious relevance, legitimacy has received scant attention in studies of costly punishment until recently (e.g., Baldassarri and Grossman 2011).
While validity and propriety are both important elements of legitimacy, they can vary orthogonally; something can be reasonably valid but more or less proper (Zelditch and Walker 2000), such as a political regime (Kuran 1995) or unpopular norms (Willer, Kuwabara, and Macy 2009) that people ostensibly follow without personally endorsing them. Past research (summarized in Zelditch and Walker 2000) suggests that propriety—independent of validity—can help regulate the use of power by authority. In laboratory studies simulating bureaucratic decision making, participants assigned to proper (vs. less proper) positions were more likely to use their power in ways that they believed was proper. Building on this insight, our goal in the present research is to see how concerns about propriety in particular affect patterns of punishment in public goods games. In doing so, we highlight the relevance of legitimacy to questions of fundamental importance about human cooperation and enforcement.
Enforcers can derive legitimacy fromsocial processes like peer election (Baldassarri and Grossman 2011; Kosfeld, Okada, and Riedl 2009) or appointment to formal roles by authority (Walker, Rogers, and Zelditch 1988). Here, we examine legitimacy that derives from using costly versus costless punishment. We argue that punishment is perceived as more proper when it is costly rather than costless to oneself because enforcing cooperation at one’s own expense is viewed as selfless, fair, and sincere. In collective action settings, such acts of self-sacrifice for collective welfare are often rewarded with peer approval and compliance (Barclay 2006; Willer 2009). In this view, costly punishment is a form of costly signal that helps convey prosocial motives and affirms legitimacy in one’s group if used properly (Jordan et al. 2016). Conversely, punishers may risk antagonizing their peers by misusing their power, namely, using punishment improperly or using punishment that lacks apparent legitimacy (Willer et al. 2012; Xiao 2013; Xiao and Tan 2014). Costless punishment, in other words, may be perceived as improper—selfish, unfair, and antisocial—because it penalizes the target without any cost to the punisher.
Designated Punishment
We further argue that making punishment costless rather than costly has markedly different effects on peer versus designated punishment. The reason is that the cost of punishment helps offset concerns about the unequal distribution of power to punish under conditions of designated punishment. In peer punishment, all members are given the same role and the power to punish each other, which diffuses concerns about the fairness and appropriateness of punishment (Kurzban, DeScioli, and O’Brien 2007; Molenmaker, de Kwaadsteniet, and van Dijk 2016). Instead, under the “default” case of peer punishment, punishment is driven primarily by the basic economics of punishment, punishing less as the cost to oneself increases and the cost to targets decreases (Egas and Riedl 2008; Horne and Cutlip 2002). For instance, Egas and Riedl (2008) find that in public goods games with punishment, participants are less likely to use punishment when it costs the punisher 3 points rather than 1 point, and this was the case whether targets were relative cooperators or defectors. Based on these findings, our baseline hypothesis is that making punishment costly to punishers will reduce both prosocial punishment (punishing relative defectors) and antisocial punishment (punishing relative cooperators).
Hypothesis 1: Costly versus costless punishment under peer punishment. Peer punishers will engage less in prosocial and antisocial punishment when punishment is costly rather than costless.
Designated punishment may depart from such rational predictions. Compared to peer punishment, designated punishment heightens concerns about legitimacy, such as appearing proper, because punishment is centralized and delegated to one entity (e.g., leaders or outside authorities) who take on a unique role, commensurate with special status and power within each group, that elicits greater scrutiny (Devlin-Foltz and Lim 2008; Kosfeld et al. 2009). Under such conditions, and insofar as their power to punish is costly and viewed as proper, designated punishers will tend to use it in ways that are apparently proper and justifiable (prosocial) and avoid losing peer approval and support by abusing it because one defining effect of legitimacy is voluntary compliance with the demands of the situation (e.g., rules, norms, role expectations); people are more likely to comply with what they consider legitimate (Tyler 2006). In contrast, making punishment costless will reduce such normative constraints and let designated punishers gravitate toward antisocial punishment because it is no longer possible or meaningful to use such punishment to affirm or maintain propriety if the punishment itself lacks apparent propriety; it is unclear how to properly use something that is improper. In short, costly (vs. costless) punishment will prompt designated punishers to use punishment more prosocially.
Hypothesis 2: Costly versus costless punishment under designated punishment. Designated punishers will engage in more prosocial punishment (relative to antisocial punishment) when punishment is costly rather than costless.
We are not the first to examine the legitimacy of designated punishers (Baldassarri and Grossman 2011; Kosfeld et al. 2009). However, we are not aware of studies that examine the cost of punishment as the basis of legitimacy for designated punishers. We focused on the cost of punishment for both methodological and theoretical reasons. First, it is relatively easy to manipulate in both peer and designated punishment systems; alternative ways to legitimate enforcement, like electing punishers (e.g., Baldassarri and Grossman 2011), make little sense in peer punishment. Second, examining the cost of punishment is interesting because it highlights how the logic of legitimacy can diverge from the calculus of costly punishment. Although we are not the first to compare costly versus costless punishment or peer versus designated punishment, past research on the cost of punishment precluded costless punishment (Egas and Riedl 2008) or compared different forms of punishment (social vs. economic sanctions; Noussair and Tucker 2005), thus confounding the cost and form of punishment. To our best knowledge (see Balliet et al. 2011), no study has examined costly versus costless punishment by designated punishers.
Methods
To test these ideas, our experiment modified the public goods game with punishment (Fehr and Gächter 2000) to manipulate the cost of punishment to enforcers (holding constant the cost to targets) in groups with peer versus designated punishment. In natural settings, punishment occurs in various forms, for example, formal versus informal, public versus private, material versus symbolic. It may be argued that punishment is never perfectly costless to the punisher, given administrative, psychic, relational, and various other nonmaterial costs (Adams and Mullen 2012). Our research is not designed to speak to this issue but to test the possibility that simply varying the material cost of punishment has consequences for how punishment is used.
To see whether perceptions of propriety varied across different enforcement systems, we examined propriety in three ways. First, participants completed an exit survey after the public goods games. The survey included questions about how proper people felt in their assigned roles during the public goods games. Second, we ran a pilot study in which volunteers read a written description of the public goods game with a different type of punishment (costly vs. costless and designated vs. peer) and rated the propriety of the enforcement system. Finally, we examined compliance, namely, the effects of punishment on cooperation, as a behavioral measure of legitimacy in the public goods experiment. By considering propriety at multiple levels, we view legitimacy as a property of an enforcement system as a whole rather than particular individuals, their acts, or their relations (Dornbusch and Scott 1975).
Participants and Procedure
Two hundred and thirty-four students (20.15 ± 1.90 years old, 36 percent male) from a large university were recruited forcash based on overall performance (average = $12.50). The experiment was described as a study of organizational teamwork and took place in a laboratory with a no-deception policy. Participants were scheduled in groups of 6 to 12 but seated at isolated computer terminals. The entire experiment took place over the Internet through a custom website, thus preventing any face-to-face interactions or verbal communication and ensuring anonymity.
Prior to the experimental task, participants completed the consent form, detailed instructions, and a comprehension test. Next, they were randomly sorted into groups of three in different experimental conditions and assigned roles as punishers or non-punishers before completing “up to 10 rounds” of public goods games (the experiment ended after 6 periods). Finally, they were given an exit survey, received debriefing and payment, and dismissed.
Design and Materials
We created four experimental conditions: P1, P0, D1, and D0, where P and D designate peer versus designated punishment, and 1 and 0 designate the cost of punishment. Our baseline condition (P1) was the standard public goods game with costly punishment (Fehr and Gächter 2000). In the first (“contribution”) stage of each round, each member was given an endowment of 20 monetary units (MUs), of which they could contribute any amount to a team project or keep. Each MU contributed to the team project yielded a marginal per capita return of .5 MU for each member, and thus 1.5 MUs for the whole team, whereas keeping 1 MU yielded 1 MU for that member only. Thus, the earning π it for member i in the first stage of period t is
In the second (“punishment”) stage, each punisher was given an opportunity to punish teammates, which entailed assigning “deduction” points (0 MUs to 10 MUs) out of one’s own earnings to each other member. Each punishment point cost the punisher 1 MU and the target 3 MUs. Thus, the final payoff in each period t for player i in peer punishment is
In the costless peer punishment condition (P0), the cost of punishment was 0 MU for punishers but 3 MUs for targets. In the two designated punishment conditions with costly (D1) or costless (D0) punishment, one member in each group was randomly chosen to be the sole punisher (“Leader”) across all rounds. Punishers did not receive any additional endowment for punishment (O’Gorman et al. 2009).
To make the public goods game more engaging and meaningful to participants, it was described as a series of team projects in which members of an organization are asked to split their time between individual projects and team projects. Each “week,” participants had 20 hours of unsupervised time (equivalent to 20 MUs) and earned different points from each hour contributed to team projects (depending on how much other members were contributing to team projects as well) versus individual projects, based on Equation 1. At the end of each week, punishers were given an opportunity to provide “feedback” to each other by assigning deduction points. The experimental materials are provided in Appendix A.
In our experiment, contributions and punishment were public knowledge. After each round, each member learned how much each person contributed, who was punished, by whom, and at what cost (in MUs) to the punisher and the punished teammate. This design was necessary to ensure comparability across peer and designated punishment conditions such that all punishers, not just designated punishers, were identifiable. It also served to reinforce the costliness of punishment, our key manipulation.
Results
Table 1 presents summary statistics. Following the literature (Herrmann et al. 2008), we operationalized prosocial and antisocial punishment as punishing a relative defector versus cooperator, namely, a target who contributed fewer versus equal or more MUs than the punisher. Although punishers may engage in antisocial punishment for a number of different reasons, such as retaliation (Herrmann et al. 2008) or intergroup competition (Meier et al. 2012), we view antisocial punishment primarily as a display or abuse of power (Rand and Nowak 2011) since our experimental setup for the designated conditions rules out retaliation and intergroup competition.
Means and Standard Deviations of Contributions, Punishment, and Earnings
Note: Standard deviations in parentheses. P0 = costless peer punishment, P1 = costly peer punishment, D0= costless designated punishment, D1 = costly designated punishment, MUs = monetary units.
We begin by comparing peer (Hypothesis 1) versus designated (Hypothesis 2) punishment separately because our hypotheses concern how punishment cost affects the use of punishment under different punishment regimes. We then introduce econometric tests that compare peer and designated punishment more directly. After the hypothesis tests, we present evidence that varying the cost of punishment changed the perceived propriety of punishment. All tests are two-tailed.
Effects of Propriety on Punishment
As expected, peer punishers punished less when punishment was costly rather than costless, but this was not the case for designated punishers. Over six periods, peer punishers used less prosocial punishment in P1 (M = 3.13 MUs, SD = 4.45) than P0 (M = 8.08 MUs, SD = 13.16), t(109) = 2.56, p = .01, d = .49, and antisocial punishment in P1(M = 2.02 MUs, SD = 7.08) than P0 (M = 5.17 MUs, SD = 8.50), t(109) = 2.10, p= .04, d = .40. In contrast, costly punishment increased prosocial punishment by designated punishers from 3.57 MUs (SD = 4.64) to 7.70 (SD = 6.83), t(39) = 2.27, p = .03, d = .71 whereas it decreased antisocial punishment from 4.24 MUs (SD = 9.34) to 1.15 (SD = 1.95), t(39) = 1.45, p= .16 d = .45, although this effect did not reach significance at the 5 percent level.
While these patterns are consistent with our hypotheses, the descriptive results may be biased because they do not control for the effects of contributions on punishment or the repeated measures nested in individuals and teams. It is also difficult to compare peer versus designated punishment directly because of methodological differences (e.g., 1 vs. 3 punishers). We addressed these issues as follows by using errors clustered at the level of individual punishers and teams and controlling for contributions from the punisher, the target, and the team total in each round and the fixed effects of rounds:
The subscript m indicates punisher 1 . . . 3 (only 1 under designated punisher), idenotes rounds 1 . . . 6, and t is the current round. Thus, pimt is punishment by i to m, and cit is contribution by i, both in round t. In this model, B0 is the constant, B1 is the team’s total contribution, B2 is the punisher’s contribution, B3 is the difference in contribution between the punisher and a target, B4 is punishment given in t– 1, and B5 . . . are the types of punishment (costly vs. costless, peer vs. designated, and their interaction effect). Following the literature (e.g., Ashley, Ball, and Eckel 2010), we submitted this model to tobit regression because the dependent variable, punishment per target and round in MUs, is censored at 0 and 10 MUs.
The results (Table 2) converge with the descriptive patterns. In peer punishment groups, costly punishment shows negative effects on both prosocial punishment, B = −2.39, robust SE = .67, p < .001, and antisocial punishment, B = −3.46, robust SE = 1.60, p = .03, supporting Hypothesis 1. Under designated punishment, in comparison, costly punishment increased prosocial punishment, B = 1.14, robust SE = .57, p = .047, while it had no effect on antisocial punishment, B = .67, robust SE = 1.50, p = .65, consistent with Hypothesis 2. Finally, in the pooled data, the effect of costly × designated punishment on prosocial punishment was significant and positive, B = 3.40, robust SE = 1.01, p = .001. Looking at antisocial punishment, we find a negative main effect of costly punishment, B = −3.50, robust SE = 1.61, p = .03, but no interaction effect, B = 2.80, robust SE = 2.53, p = .27. Altogether, these results show that costly punishment increased prosocial punishment by designated punishers but not peer punishers and reduced antisocial punishment under both peer and antisocial punishment. 1
Predictors of Punishment Per Punisher Per Round
Note: Results from tobit regression with errors clustered at the individual and team levels. Robust standard errors in parentheses.
p < .05. **p < .01, two-tailed tests.
Evidence of Propriety
Overall, our results support our reasoning that imbuing punishment with propriety helps regulate the use of power by motivating prosocial punishment without increasing antisocial punishment by designated punishers. An alternative explanation, however, is that the cost of punishment changed its credibility, not propriety. That is, punishment may feel more credible—the punisher really means it—if it is costly because, according to signaling theory (Spence 1974), signals are taken more seriously if they are costly. Creditability and propriety are different because both proper and improper acts can be credible (e.g., a mobster threatening a person’s life). To address this issue, we provide three lines of evidence for propriety.
First, prior to the laboratory experiment, we ran a pilot study in which 240 volunteers from Amazon Mechanical Turk (all North Americans; 3 did not finish the study) were recruited in exchange for monetary compensation and asked to provide feedback on a new experiment “designed to examine teamwork.” Participants were randomly assigned to read a description of a public goods game with a different type of punishment (costly vs. costless and peer vs. designated), taken from the main experiment (see Appendix B). Next, participants answered two questions about propriety: “How fair/appropriate is this enforcement system?” (1 = very unfair/inappropriate, 5= very fair/appropriate; Spearman’s ρ> .79). 2 Designated punishment was perceived as more fair and appropriate when costly (M = 4.16, SD = 1.26) than costless (M = 3.48, SD = 1.34), t(120) = 2.84, p = .005, d = .46. However, peer punishment was perceived as equally fair and appropriate when costly (M = 4.23, SD = 1.06) versus costless (M = 4.34, SD = 1.37), t(115) = .49, p = .62. Thus, the cost of punishment changed perceptions of propriety at the system level but only under designated punishment.
Second, in the exit survey after the public goods games, participants were asked how proper (1 = very unfair/disrespected, 7 = very fair/respected, Spearman’s ρ = .48) they felt while playing their assigned roles. 3 Designated punishers reported feeling more legitimate in D1 (M = 4.88, SD = 1.73) than in D0 (M = 3.55, SD = 1.50), t(39) = 2.63, p = .01, d = .82. Non-punishers also felt more proper in D1 (M = 4.61, SD = 1.21) than D0 (M = 3.57, SD = 1.48), t(80) = 3.47, p= .008, d = .77. These results were robust to controlling for total punishment received and final earning. No such patterns were found between P0 (M = 3.78, SD = 1.42) and P1 (M = 4.14, SD = 1.52), t(109) = 1.30, p= .20. In addition, in the designated punishment conditions only, non-punishers were asked how fair the punisher in their team was. They evaluated their punisher as more fair in D1 (M = 4.63, SD = .26) than D0 (M = 3.69, SD = .25), t(80) = 2.60, p = .01, d = .57.
Third, an indirect but consequential measure of legitimacy is compliance, namely, how much members increase their contributions after receiving punishment (Baldassarri and Grossman 2011; Zelditch 2001). If costly punishment is more credible but not legitimate, we should find costly punishment to increase compliance after receiving either prosocial or antisocial punishment. To the contrary, we find that costly punishment increased the efficacy of prosocial punishment only and only for designated punishers.
Table 3 shows results from tobit regression predicting contributions in MUs in round t as a function of punishment received in t– 1. We found a significant positive effect of punishment received in t– 1 × costly punishment under designated punishment, B = .36, robust SE = .16, p = .03, but not peer punishment, B = .15, robust SE = .14, p = .28, indicating that contributions increased more after receiving costly (vs. costless) punishment from designated punishers but not from peer punishers.
Effects of Punishment on Compliance
Note: Results from tobit regression with errors clustered at the individual level. Dependent variable is contribution in monetary units (MUs) in round t. Robust standard errors in parentheses.
p < .05. **p < .01, two-tailed tests.
An alternative explanation for the difference in compliance under designated punishment is that punishers engaged in more antisocial punishment under costless punishment, which perhaps alienated group members and reduced their compliance. In other words, compliance dropped not because punishment was costless (and therefore less proper) but because punishers were abusing it. To consider this issue, we re-specified our econometric model by replacing the term punishment received in Table 3 with prosocial punishment received and antisocial punishment received and ran the new model for the costless and costly designated punishment conditions separately. This model should help reveal differences in compliance under costless versus costly punishment due specifically to antisocial punishment.
Table 4 shows the results. First, contrary to the idea that punishers drove down compliance by engaging in antisocial punishment, antisocial punishment had no effect on compliance in either condition. Second, prosocial punishment is positive and marginally significant under costly punishment, B = .93, SE = .54, p=.08, but not under costless punishment, B = –.23, SE = .51, p = .65. Thus, prosocial punishment lost its efficacy when it was made costless. While post hoc, these results provide support for the idea that making punishment costless affected compliance directly (by reducing propriety) rather than indirectly (by increasing antisocial punishment). 4
Effects of Punishment on Compliance under Designated Punishment
Note: Results from tobit regression with errors clustered at the level of individual targets of punishment. Robust standard errors in parentheses.
p < .1. **p < .01, two-tailed tests.
Cooperation and Efficiency
As supplemental analysis, we examined cooperation and efficiency across conditions (Figure 1). Censored regression of individual contribution per round on experimental conditions with random effects at the individual participant level and team level found a positive effect of costly punishment on average contribution under designated punishment, B = 1.48, robust SE = .66, p = .026. We found no difference between P0 and P1, B = .59, robust SE = 1.03, p = .57, although peer punishment produced more cooperation than designated punishment under both costly punishment, B = 1.59, robust SE = 79, p = .044, and costless punishment, B = 3.62, robust SE = .88, p < .001.

Individual contributions and earnings by condition.
Next, we regressed individual earning in MUs per round on conditions, with random effects at the individual and team levels. The results reveal that earnings were higher under costly punishment, B= 3.02, SE = 1.40, p = .031, and under designated punishment, B = 3.44, SE = 1.93, p = .009, but there was no effect of costly × designated punishment, B = −2.74, robust SE = 1.93, p = .16. Figure 1b suggests that these patterns are driven by the low earnings in groups with costless peer punishment, which showed the highest levels of punishment. Prior research has found that costly peer punishment increases cooperation, but overall efficiency gains are erased by punishment (Egas and Riedl 2008; Herrmann et al. 2008), at least in experiments with short time horizons (6–8 rounds; Gächter et al. 2008). Our research suggests that this may be the case in particular when peer punishment is costless. In contrast, costly designated punishment increased cooperation without reducing efficiency, relative to costless designated punishment. 5
Discussion
An enduring insight from sociology is that legitimacy exerts powerful constraints on actors to comply with prevailing norms ofconduct. For instance, Zelditch and Walker (2000) suggest that concerns about legitimacy can help regulate the use of power by invoking a sense of what is morally right or socially proper for power-holders (also Fast and Chen 2009; Kuwabara et al. 2016). The present research extends this argument on the basis of legitimacy induced by making the punishment costly rather than costless to the punisher. Consistent with our hypotheses, designated punishers engaged in significantly more prosocial punishment but not antisocial punishment when punishment was costly and thus perceived as proper. In contrast, peer punishers were less likely to use costly rather than costless punishment. We also found that costly punishment was more effective than costless punishment in sustaining cooperation, but only under designated punishment.
Our results for designated punishers are related to the effect of explicit cost to invoke norms of economic rather than social exchange. Shampanier, Mazar, and Ariely (2007) found that when offered candies at 1 cent each, students took four on average; when offered free candies at 0 cent each, more students took candies, but almost none took more than one. The argument is that even a small cost can invoke a market mindset that helps justify and motivate consumption, whereas zero cost invokes social norms against taking more than one’s share. Enforcement may be subject to a similar psychology (Tenbrunsel and Messick 1999), invoking different norms when it is costly rather than costless. In our experiment, simply changing the cost of punishment from 0 MU to 1 MU amounted to a noticeable shift from costless to costly punishment that was not prohibitively large in economic terms yet salient enough to alter its moral significance, imbuing acts of punishment with legitimacy and reducing antisocial punishment.
These patterns call into question what the cost of punishment really represents in experiments on costly punishment. Before the recent surge of interest in costly punishment, exchange theorists pursued a productive line of work on the use of coercive power without explicit cost (Lawler, Ford, and Blegen 1988; Molm 1997). The experiment by Fehr and Gächter (2000) made a departure by assuming costly punishment to provide a more stringent test of the idea that people are willing to use punishment even at their own cost. As our research suggests, however, punishment cost is more than a price; it is also a social signal that changes how people interpret the act of punishment. Indeed, many acts of punishment in real life do not come with an explicit price tag. An important direction for future research is to better understand conditions under which punishment is actually viewed as costly or costless.
Our findings have implications beyond the laboratory. For instance, recent years have seen a phenomenal growth of reputation systems designed to regulate online markets (e.g., eBay, Amazon, Yelp) byharnessing peer-to-peer enforcement that is virtually costless. Despite the success of these systems, however, a persistent challenge is how to ensure prosocial enforcement, namely, feedback that is viewed as proper and legitimate (Resnick et al. 2000). Our research suggests that making feedback a little more costly may help curtail antisocial punishment by raising not only the economic cost of punishment but also concerns about legitimacy.
Although legitimacy may derive from various sources, the idea that legitimacy may also inhere in the cost of punishment is empirically novel and may shed light on the problem of enforcement. If costly punishment is a source of legitimacy in itself, people might engage in enforcement precisely because it is a costly and therefore effective signal of their prosociality and social status (Barclay 2006; Jordan et al. 2016). This points to conditions under which punitive sentiments may have evolved: in hierarchical groups that recognize selfless punishers as legitimate leaders (Traulsen, Röhl, and Milinski 2012). Such groups have received relatively scant attention in the literature on costly punishment even though flat groups with no clear hierarchical differentiation are rare outside of the laboratory (Gruenfeld and Tiedens 2010). Hierarchies emerge quickly and spontaneously, creating differentiations in power and status that become the basis of social arrangements like designated punishment. Our understanding of how groups enforce cooperation in hierarchical groups is incomplete without greater efforts to account for the social psychology of power and legitimacy. In particular, more work is needed to better understand the conditions under which costly versus costless punishment evolved in different types of groups.
Footnotes
Appendix
1
As robustness checks, we examined punishment frequency (whether punishment occurred or not) and severity (points assigned, given actual punishment). For frequency, we obtained similar results as
. For severity, we did not obtain robust results since the analysis considers instances of actual punishment only, reducing the sample sizes. These results are available on request.
2
It is worth noting that propriety is more than fairness because, despite the high correlation here, fairness and appropriateness are conceptually distinct. For instance, something fair can be inappropriate (e.g., equal division of illicit money). This is crucial because fairness alone cannot explain why designated punishers engaged in more prosocial punishment since it is possible to be “fair” in other ways, for example, by not engaging in punishment at all.
3
Rather than simply replicating the vignette study, we changed the level of analysis from the system as a whole to individual roles within thesystem to see if manipulating legitimacy at the system level would affect individual behaviors and experiences at the role level. In doing so, we used the term respected rather than appropriate to better capture how people think they are seen by others.
4
It is curious that punisher contribution in t– 1 is significant under costless punishment only. Our interpretation is that when punishment was costless and deemed illegitimate, group members paid greater attention to punisher contributions.
5
O’Gorman, Henrich, and Van Vugt (2009) found that costly designated punishment sustains as much cooperation as costly peer punishment. One key difference in our designs is that participants stayed in the same group across all rounds instead of rotating after every round. Peer punishment (but not designated punishment) may be more effective in fixed groups.
