Abstract
This study investigates the effectiveness of increasing police presence in hot spots typically managed with problem-oriented policing (POP) compared to using POP alone. Our block-randomized controlled trial reveals that overall, “saturated POP” does not significantly reduce crime compared to a POP-only approach. However, a more complex pattern emerged across different blocks, namely, an inverse relationship between the level of police presence and the level of recorded crime, regardless of the random assignment. Crime rates decreased in the experimental arm within two statistical blocks, where the assigned dosage of saturated POP was greater than in the control arm. However, in one block, the assigned control hot spots received a higher dosage of police presence than the assigned treatment hot spots, resulting in significantly lower crime rates in the control hot spots. These findings underscore the necessity of reporting field experiment results based on both treatment-as-assigned and some form of treatment-as-delivered models.
Keywords
Introduction
Hot spot policing, a strategy that involves ongoing visits to micro-locations that experience a disproportionate level of crime, has received significant attention in criminology. With numerous rigorous evaluations already meta-analyzed (Braga et al., 2019; Dau et al., 2021, 2022), place-based interventions are integral to evidence-based law enforcement (Weisburd et al., 2019). While the degree of implementation of the core concept of “hot spot policing” may be a concern (Ariel, 2023), the strategy has gained traction in both research and practice.
Despite the prominence of hot spot policing, much of the existing research narrows its focus to the application of a single type of policing strategy in these hot spots—for example, saturated presence, problem-oriented policing (POP), or community policing—compared to no treatment or business-as-usual conditions (Ariel, Sherman et al., 2020). Rarely are controlled tests conducted to compare the relative effectiveness of various treatments or to assess the additive effect of supplementing one treatment with another, because they raise a methodological impasse: When officers deliver a specialized intervention such as POP, they are physically present at the hot spot, conflating the activity with proactive, saturated presence. Therefore, it is challenging to separate the effect of presence from POP activities on crime rates in hot spots. A possible solution would be to compare the additive effect of two or more interventions (enhanced treatment; Peck, 2020) versus one treatment only.
A closely related issue is a fundamental methodological consideration: how to analyze the data from such an experiment. Two primary options are available. The first is the intention-to-treat (ITT) model, which operates under the “treatment-as-assigned” assumption—that all units were managed and treated according to the experimental protocol. The second is the treatment-on-the-treated (TOT) model, which follows the “treatment-as-delivered” assumption, analyzing the effects of the intervention only among the units that actually received the treatment as intended (Bloom, 1984, 2008). For the most part, experimental criminologists analyze data under the intention to treat (ITT) model. In practice, however, the literature on field experiments in the social sciences is unambiguously clear that dropping out is common, misassignment and misallocation of treatment conditions occur, and treatment providers do not maintain sufficient treatment content validity (Ariel et al., 2022). Moreover, the most rigorous experiments in criminology on the effectiveness of hot spot fail to publish both outcomes and outcomes on treatment effects under tightly controlled implementation conditions (Braga et al., 2019), despite the existence of published guidance (Angrist, 2006). Therefore, how treatments in hot spots behave despite implementation issues, such as subjects’ non-compliance with the treatment protocol, is unknown.
In response to these concerns, this article reports results of a field trial where officers delivered additional patrols in POP hot spots. The effect of this combination was measured against the impact of POP-only hot spots. Next, as officers were equipped with GPS trackers, we could measure the treatment dosage in both arms of the experiment, at least in terms of police presence in the hot spots. We could then estimate the effect of saturated POP relative to ordinary POP dosage under both the ITT (treatment-as-assigned) and TOT (treatment-as-delivered) models of “saturated POP” at hot spots.
Literature Review
Hot Spots and Hot Spot Policing
One of the significant findings in contemporary criminology is that most crime events are concentrated in very small areas: street segments, unique addresses, and other micro geographical units, commonly referred to as “hot spots” (Pierce et al., 1988; Sherman et al., 1989). Accordingly, Weisburd (2015, p. 138) indicated that there is a “law of crime concentration” at certain specific locations and in relatively similar concentration bandwidths across different cities and over time (Lee et al., 2017). As a result, our understanding of crime calls for greater attention to the spatial, temporal, and ecological manifestations that can be found at micro-locations such as hot spots or risky facilities (Eck et al., 2007; Townsley et al., 2014; Weisburd et al., 2016).
One practical application for hot spots emerged in the 1990s, when policing expanded beyond a person-focused approach to include a micro-location–based approach. A paradigmatic shift to “hot spot policing” and a deluge of experimental evidence on the benefits of focusing on hot spots only began with Sherman and Weisburd's (1995) seminal field trial in Minneapolis. Their study, and most replications that followed (Braga et al., 2019), demonstrated that implementing focused interventions, with police resources specifically allocated to hot spots, can reduce crime compared to control conditions. Recent reviews have suggested an average treatment effect in the magnitude of a 16% reduction in crime due to hot spot policing (Braga & Weisburd, 2020; Weisburd & Majmundar, 2018).
Not All Hot Spot Policing Interventions Are Equal
Studies of hot spot policing have recently moved beyond whether it is effective and started asking which strategies are more effective, reflecting a broader trend in policy evaluation. For example, Braga et al.'s (2019) systematic review of 78 hot spot policing evaluations found that the available evidence suggests two primary methods for delivering hot spot policing: saturation or POP interventions.
Saturated policing at hot spots refers to placing additional visible resources at designated locations to deter wrongful behaviours in the public domain. While officers may be directed to conduct specific interventions while patrolling the hot spots (Dau et al., 2021), the primary focus is on high visibility and, by implication, deterrence theory. Sherman and Eck (2003) observed that saturation interventions might increase criminals’ perceived risk of detection and reduce their likelihood of committing a crime within those spatiotemporal confines (Blattman et al., 2021; Huey et al., 2021).
The question of optimal dosage—i.e., how much police patrol should be delivered in the hot spots—has been a central question in the hot spot policing literature for more than 30 years (Sherman, 1991). Based on observational data from the Minneapolis experiment (Sherman & Weisburd, 1995), Koper (1995) found that the optimal time spent in a hot spot is about 15 minutes less or more time leads to diminishing returns. Similarly, in Sacramento, California, officers were instructed to rotate between hot spots randomly and spend about 15 minutes in each hot spot; Telep et al. (2012) found that these “Koper minutes” significantly impacted crime. This trial suggests that officers need not be at hot spots constantly or even for an extended period to prevent crime. At the same time, there is little systematic experimental evidence comparing different dosage levels (Williams & Coupe, 2017). Nor is it clear if there is a dose—response relationship between frequency and crime—i.e., the more officers visit the hot spot, the less crime will emerge (Ariel et al., 2016).
On the other hand, POP attempts more systematic and long-lasting solutions to crime by identifying the core reasons that lead to offences and then implementing change accordingly. As Hinkle et al. (2020, p. 1) notes, “POP interventions commonly use the SARA (scanning, analysis, response, assessment) model to identify problems, carefully analyze the conditions contributing to the problem, develop a tailored response to target these underlying factors, and evaluate outcome effectiveness.” Applying POP has reduced the incidence of violence (Taylor et al., 2011), drug markets (Hope, 1994; Weisburd & Green, 1995), and prostitution (Weisburd et al., 2006) compared to control conditions. The latest systematic review (Hinkle et al., 2020)—though not only in hot spots—suggested a 33.8% reduction in crime problems relative to the controls (Braga et al., 1999; Braga & Bond, 2008; Groff et al., 2015; Weisburd & Green, 1995; for implementation of POP in hot spots).
The question that arises is what delivers a stronger treatment effect on crime in hot spots: saturated police presence, POP, or a combination of the two? Braga et al.'s (2019) review indicated that POP generates larger overall effect sizes relative to increased policing interventions (Braga & Bond, 2008; Eck, 1997, 2002), likely because the POP approach promises to deliver tailored responses to particular recurring problems at crime hot spots (Braga & Bond, 2008; Weisburd et al., 2006). However, as Hinkle et al. (2020, p. 4) commented: “distinguish[ing] between the effects of bringing focused attention to hot spots and that of such focused efforts being developed using a problem-oriented approach” is challenging. Stating this criticism differently: the mere presence of police officers at hot spots may reduce crime without necessitating an expensive intervention designed to uproot the core antecedents of a crime problem (Hasisi et al., 2022).
Some studies have attempted to observe POP versus saturated patrol or components thereof (Groff et al., 2015; Kochel et al., 2015). Taylor et al. (2011) compared POP with direct saturation patrol in one trial and found that POP generated statistically significant reductions in crime while saturation patrol did not yield any statistically significant decreases. However, the trial incorporated a limited number of hot spots, and statistical power considerations may drive the non-significance. Furthermore, Rosenfeld et al. (2014, p. 446) discovered that enhanced directed patrols, when not accompanied by other enforcement activity in the hot spots, did not produce significant reductions in firearm assaults. Notably, this study's supportive results were limited to one type of crime, but they suggest that a proactive approach in hot spots is necessary for successful hot spot policing initiatives.
There is a third option that has not been entertained in the literature: Will combining two interventions reduce crime rates more than a solo treatment? Consider “saturated POP”: two treatments running simultaneously within the designated hot spots, with more police presence alongside problem-solving, all relative to “business-as-usual” conditions.
Clinical trials often utilize various intervention strategies to improve the effectiveness of treatments. However, the application of more complex experimental designs in policing experiments is infrequent. In research on multi-treatment interventions for offenders and crime hot spots, several meta-analyses have demonstrated the efficacy of combining various strategies (Jolliffe & Farrington, 2007; Lipsey & Cullen, 2007). However, these are the exceptions rather than the norm; experimenters commonly concentrate on isolating the impacts of a single crime policy to estimate its effectiveness. As a result, the investigation of combined effects or comprehensive approaches involving multiple treatments is not fully utilized, even though criminal justice system agents frequently apply a multisystemic or holistic approach.
Treatment-as-Assigned Versus Treatment-as-Delivered in Hot Spot Policing Experiments
Implementation is a related topic often neglected in the study of policing (Sherman, 2013). We primarily rely on the “treatment-as-assigned” assumption of the ITT model, namely, that officers comply with the experimental protocol, and that the research procedure is perfectly delivered. However, this is likely a lenient assumption (Ariel, 2023). In part, this convention aligns with the ITT analytical model: analyzing experimental data based on the assignment of hot spots according to the experimental protocol, rather than the actual delivery of treatment, following TOT (“treatment-as-delivered”). Of course, the ITT model holds a crucial position in policy evaluations due to its substantial significance for several convincing rationales (Bloom, 2008). The ITT model views the treatment-as-assigned as synonymous with treatment-as-delivered, making the estimators can be used for policy recommendations (Pocock, 2013).
Formally, the ITT model adheres to the principles of randomization and preserves the integrity of the original study design. Participants are analyzed based on their allocated treatment groups, irrespective of their adherence to the treatment procedure. This conservative estimate depicts the policy's performance when applied on a larger scale because, in real-life settings, there will be dropouts, treatment fallouts, spillovers and misassignment of interventions (Ariel et al., 2022). In addition, one of the main criticisms of utilizing the ITT model is a perception that it might underestimate treatment effects by including an entire treatment group that includes both takers and no-shows (Armijo-Olivo et al., 2019). The ITT approach's validity in accurately representing an intervention's actual impact may be compromised, especially in cases with significant rates of noncompliance. For example, in hot spot policing, this could occur when officers fail to deliver the expected levels of intervention, such as the prescribed frequency or duration of patrols. This might dilute the observed effects, resulting in smaller effect sizes. Underestimating the treatment effect can have considerable consequences in policymaking: When policymakers need to accurately assess the full extent of the benefits associated with an intervention, underestimating the treatment effect can lead to ill-informed policy decisions, resulting in insufficient allocation of resources (Friedman et al., 2015).
Similarly, ignoring treatment-as-delivered has the potential to decrease accuracy. Including noncompliant people or dropouts in the analysis may inject additional unpredictability into the data, diminishing the study's statistical power. Consequently, identifying statistically significant disparities across treatment groups may become more arduous. Finally, non-clinical conditions, especially in policing experiments, do not have double-blinding (Ariel et al., 2022), so assuming balanced and interference-free experimental groups is another lenient assumption (Ariel et al., 2019). In reality, when control conditions in field trials adhere to “business-as-usual” circumstances, errors in the delivery of the studied treatment inherently interact with the random allocation. Assignment into the experimental group is the reason for issues with implementation; therefore, ignoring treatment-as-delivered may increase the likelihood of invalid treatment effect estimates.
Lack of Treatment Delivery Data in Hot Spot Policing Trials
Many studies—and hot spot policing experiments in particular—fail to report sufficiently detailed treatment fidelity measures (Braga et al., 2019). Given the comments of Rosenfeld et al. (2014, p. 435) that “[p]oor fidelity to experimental procedures is the downfall of many otherwise promising field experiments,” we should do more to provide sufficient details about the treatment-as-delivered. However, hot spot policing experiments provide insufficient details of the quantity, quality, and method of administration of the intervention (Sherman & Weisburd, 1995). The examples are rare, and when evidence is published, it portrays an unwelcomed discovery; as Bland et al. (2021, p. 113) concluded, “getting the desired level of patrol [is] impossible, even during the experiment.”
The Present Experiment
In this study, we sought to accomplish three primary objectives. The first was to test the effect of saturated POP at crime hot spots relative to business-as-usual levels of POP. Based on hot spot maps in South Birmingham, UK, we asked the West Midlands Police to direct community officers to engage in 15-minute patrols three times per day during the afternoon shifts when crime peaks. We then measured the effect of their interventions against control conditions. These community support officers—PCSOs—were tasked to conduct problem-solving during their visits in both treatment and control conditions but increase the frequency and duration of their presence in the treatment condition. The experiment was relatively long—12 consecutive months—during which PCSOs were tasked to apply the SARA model (see below). Equipped with intelligence on “local nominals” and access to local resources, they were informed of the core issues at each hot spot based on police intelligence and were then asked to solve these local crime problems. When they were unable to do so, they were provided with a straightforward escalation process to resolve the issues.
Second, we capitalized on the availability of person-based GPS trackers that “pinged” the PCSOs’ location to precisely and accurately measure the dosage of patrol they delivered. While we assigned three visits per day at 15 minutes per visit, the research evidence reviewed above suggests that we cannot assume it would be offered as such. Therefore, careful measures of treatment integrity were incorporated. Importantly, we could access the dosage delivered in both the treatment and comparison conditions, and we anticipated that more time spent at hot spots or more frequent visits would cause a crime reduction: a dose–response relationship between saturation and the level of crime in the hot spot. Note that although the entire community in the treatment hot spots was exposed to the intervention (i.e., there were no “no-shows” offenders), the actual exposure to the treatment may vary, as officers could deviate from the protocol (delivering the prescribed number and duration of visits).
Finally, we applied a block design, with the level of crime in hot spots as the blocking criterion (Weisburd & Gill, 2014). This means that each hot spot was categorized into one of three groups (“blocks”) based on its crime level (see below), with random allocation occurring separately within each group. While the PCSOs were unaware of this, the procedure was helpful because it allowed us to reduce the variance in the data (hence increasing the statistical power of the test) and differentiate between outcomes at different types of hotspots.
Methods
We used the CONSORT (Consolidated Standards of Reporting Trials) statement (Altman et al., 2001) to report this trial's methods and findings. The CONSORT statement provides guidelines and recommendations to improve the clarity and transparency of randomized controlled trial reports, helping readers assess their validity and applicability (Gates et al., 2024; Pandis et al., 2017)
Trial Design
This experiment was designed as a block-randomized controlled trial (Ariel & Farrington, 2012), following guidance from Weisburd and Gill (2014) on the design of place-based randomized controlled trials. The randomization of hot spots within separate groups (“blocking”) was necessary because there are variations in the frequency of crimes committed at “high,” “medium,” and “low” level crime hot spots, and the blocking helps reduce the heterogeneity within the blocks. To clarify, “low” level hot spots are still “hot” regarding crime. As shown in Table 1, the threshold for a hot spot was defined as 36 crimes per year, and the blocks were defined as 36–50, 51–75, and 76 or more crimes in 12 months. The thresholds for the blocks were not arbitrary, but based on an empirical analysis, which revealed natural breaks in the distribution where crime counts clustered. These breaks guided the selection of thresholds, ensuring that each block represented a distinct level of crime intensity. This approach enabled a nuanced understanding of the intervention's impact across varying crime severity levels. The chosen thresholds also ensured that each block contained a sufficient number of hotspots for meaningful comparisons while adhering to a natural Pareto distribution (Norton et al., 2018).
Crime Counts in the Hot Spots, by Random Assignment and Period.
As discussed below, 81 hot spots fit these criteria. However, most hot spots were defined as low-level hot spots, with a relatively small n of hot spots defined as high-level hot spots (Nhigh = 14, Nmedium = 25, and Nlow = 42). This distribution fits the Pareto curve, with a 1:1 allocation ratio between the treatment and control arms.
Hot Spots
As we reviewed earlier, hot spots of criminal activity are defined as particular geographical regions that exhibit a significantly elevated occurrence of unlawful behaviours compared to their local environs (Sherman & Weisburd, 1995). In this trial, the unit of analysis was defined as an Euclidean polygon of 150 m2, with at least 36 crimes in 12 months. 1 The size of 150 m2 was chosen for its optimal balance: it is large enough to encompass significant criminal activity yet small enough to allow for targeted intervention (Ariel et al., 2016). Setting the lower bound at 36 crimes per year ensures that the hot spots identified are areas with substantial crime issues that merit targeted interventions. It is important to note that the 81 identified hot spots exhibited consistent crime patterns in the years leading up to the experiment. Therefore, any observed reduction in crime is unlikely to result from offenders relocating to different areas. Also, note that the vast majority of reported crimes used to define these hot spots occurred during the intervention's time frame (2 p.m. to 10 p.m., see below).
Around the hot spots, we drew at least 200 m of buffers to ensure adequate spatial separation between distinct hot spots. Spatial autocorrelation can introduce bias into analytical findings (Anselin, 1995; Brantingham & Brantingham, 1995). It also avoids the treatment spillover between treatment and control conditions to achieve statistical independence of hot spots and strengthen the reliability of our subsequent analyses (Ariel et al., 2016). We opted for a 200-m distance as it mirrors the typical length of two city blocks in urban settings, providing a tangible measure for visualization. However, given the presence of elements such as buildings, trees, and motor vehicles within this 200-m line of sight, we assumed that the deterrent effect would be reduced at this distance.
Notably, the hot spots identified for this experiment were separate from routine police operations. The West Midlands Police did not create hot spot maps before this project, so the designated 150 m2 polygons would have been arbitrary land areas insofar as routine police activities or presence are concerned. While specific facilities and addresses would have been a focus of policing activities as crime generators/attractors (Brantingham & Brantingham, 1995), the hot spot maps were entirely new to the Birmingham South Local Police Authority. Note that these maps were generated based on K-clustering, a widely used method for identifying spatial crime clusters (Chainey & Ratcliffe, 2013).
Eligibility Criteria for Participants
The types of crime incidents included were those that took place in the public domain and could be deterred by the presence of a uniformed police officer, such as shoplifting, prostitution, or street drinking. Therefore, domestic abuse, fraud, neighbor disputes, and similar offences were excluded. Furthermore, we omitted incidents considered “police generated” (Sherman & Weisburd, 1995), such as proactive incidents like stop-and-frisk, and so on. Police-generated events distort the outcome variable in hot spot experiments because if officers are directed to attend the treatment but not the control hot spots, there will be an interaction between treatment and measurement.
Setting
The trial took place in the locality of Birmingham South, which is part of the West Midlands Police jurisdiction. As of 2022 (Birmingham City Council, 2022), there were 208,304 residents in the area, comprising five localities and two primary constituencies, ranging between 68.7% and 85.8% white. Overall, 32% of the residents live in the most deprived decile in England and Wales, with unemployment levels differing between constituencies (13.9% in Northfield and 7.2% in Edgbaston). Violent crime is about 63 per 100,000, higher than the average for England (42 per 100,000). Overall, in the West Midlands Police jurisdiction, there were 6,846 officers and 484 PCSOs as of 2020 (Home Office, 2020). A high-resolution map of Birmingham South and the designated hot spots, in PNG format, can be found in the Supplementary Materials.
Interventions
The intervention was delivered exclusively by PCSOs, who were led by problem-solving sergeants. PCSOs assume a crucial function: they are non-sworn personnel appointed to collaborate with constables to augment community security and tackle problems involving the community's overall well-being (Association of Chief Police Officers of England, Wales and Northern Ireland, 2008). Although PCSOs are assigned distinct duties and authorities, they do not have the same jurisdictional authority as sworn law enforcement officials. PCSOs possess distinct powers conferred upon them through legislative measures, notably the Police Reform Act 2002 in England and Wales, but they lack the authority to make arrests. Instead, they are primarily responsible for participating in community interactions, fostering trust, and assisting in the resolution of matters that do not necessitate the involvement of fully authorized police officers. They can conduct surveillance of residential areas, resolve instances of disruptive conduct, provide guidance on crime prevention to residents, solve local problems, and provide aid in minor incidents such as acts of vandalism, reports of excessive noise, and misplaced belongings (Center for Problem-Oriented Policing, 2022).
For the objectives of the present study, utilizing PCSOs combines the two methods of hot spot intervention: increased presence in micro-places and an attempt to solve problems (in this case, alongside interested community members and partners). Wright and Decker (1994) have suggested that offenders are aware of police presence when they select their targets; therefore, PCSO presence can change an offender's choices. At the same time, trained PCSOs can apply POP to solve specific crimes based on a proactive approach to scanning and analyzing hot spots and then assessing and identifying solutions. Therefore, directing PCSOs to be present at a given hot spot more than they routinely visit other hot spots would provide an optimal test of the relative effectiveness of saturated POP. Their presence serves to establish a connection between the police force and the general public, thereby facilitating effective communication and fostering a sense of trust and cooperation.
PCSOs were trained to solve problems in collaboration with local partners. They were trained on using the SARA model (Goldstein, 1990), a problem-solving approach that involves scanning (identifying crime problems), analysis (understanding their underlying causes), response (developing appropriate strategies), and assessment (evaluating the response effectiveness). They were also provided with a formal process for addressing issues and, when needed, a method for escalating problems to a dedicated sergeant. Intelligence on local concerns was “fed” to the PCSOs about crime problems, local “nominals,” and the types of crime problems they were most likely to confront during their tour of duty. A typical tasking sheet is shown in Figure 1.

Saturated POP experiment: treatment group tasking sheet.
A team of PCSOs were then tasked to conduct their visits on foot at the hot spots three times per shift for 15 minutes (Ariel et al., 2016; Koper, 1995; Williams & Coupe, 2017), whereas each hot spot (without variation) was assigned each day of the week to a pair of PCSOs. Unlike previous US experiments on hot spot policing, the officers did not use vehicles, as PCSOs do not patrol in police cars. The assigned PCSOs were not tasked to do anything else during their shifts. However, they were responsible for attending to the designated treatment hot spots at the exact dosage throughout the experimental year, no matter the crime levels at the hot spots. All blocks in the experiment were assigned the same “dosage” of 15-minute patrols three times per shift. All patrols were conducted between 2 p.m. and 10 p.m. because (1) PCSOs are not allowed to patrol past 10 p.m. due to personal safety reasons and (2) there are far fewer crime reports before 3 p.m. Note that the PCSOs assigned to provide the saturated POP were directed to do so only in the treatment areas.
PCSOs were trained and subsequently tasked to proactively engage with people in the public domain, including juveniles, shopkeepers, and people of any background, age, and demographic, which made them highly visible within the hot spots. They were committed to solving issues the problem-solving sergeants raised and identifying suitable solutions in partnership with local partners. Following the abovementioned Koper (1995) study, the expectation was that 15 minutes per visit, for three visits per day (approximately 200 hours per hot spot during the year of the trial), provided the optimal time in which additional police visibility would be effective.
Beyond the additional presence, the PCSOs were already tasked with solving problems, including mediating between juveniles and neighbors, reporting to the neighborhood police officers for further tasking, and collecting intelligence. If PCSOs were confronted with crime problems, they could proactively engage with the issues and contact local leads as needed. As a result, the PCSOs were instructed by the relevant problem-solving sergeants based on local knowledge, intelligence, and dynamic risk assessment. They were also encouraged to engage, visit shops (including the alcohol-related facilities, i.e., pubs, present in some hot spots), and talk to community members. PCSOs also engaged with local managers (Eck, 1994) to solve crime problems.
POP-Only Group
The control group in our trial comprises hot spots in which no additional police presence was incorporated. PCSOs in the control condition would still provide POP, engaging with members of the public and local leads in the hot spots, as these locations still experienced high crime levels. However, the control sites were not exposed to a systematic increase in police presence, and a “business-as-usual” level of POP was expected.
We emphasize that the PCSOs and their commanders were not notified of the location of the no-treatment hot spots to avoid treatment contamination and crossover. The hot spot map we created for this trial was separate from regular police patrol routines, which further reassured us that the business-as-usual sites were not exposed to saturated POP.
Outcomes
The trial's primary outcome was victim-generated reported crime counts during the experimental period (Ariel & Bland, 2019). We excluded indoor incidents from the outcome measure as we defined the hot spots based on outdoor crimes only.
Outputs
Saturation. We gained a unique perspective on dosage delivery in this trial. PCSOs were equipped with hand-held GPS trackers that provided crucial evidence of their whereabouts at any given time. The hot spots’ geofence was mapped (Wain et al., 2017; Wain & Ariel, 2014), and every time an officer entered the hot spot, a “ping” was recorded and communicated to the research team. More than 1 million pings were shared with us, which we could tabulate for analytical purposes.
POP. One limitation of our study is the lack of POP implementation data (outputs). We cannot ascertain the degree to which the assigned interventions were applied, the degree to which officers administered tactics that would deal with the antecedents of the crime problems in the hot spots, nor the “POP dosage” itself. PCSOs were not tasked to record precisely what they administered based on their tasking sheets (Figure 1), so there is no way to ascertain how much POP was applied.
This is a standard limitation in POP studies (Hinkle et al., 2020), as well as other studies where police are instructed to apply interventions where we have no quantifiable measures of outputs (Braga et al., 2015; Gill et al., 2014; Kane et al., 2018). Even in the context of police presence, it is generally unknown what precisely officers do, other than “being there” (Dau et al., 2022). Thus, police performance and the application of treatments in this and other field trials suffer from what scholars have referred to as the “black box” phenomenon (Famega et al., 2017; Harachi et al., 1999): We know what POP measures were assigned, and we can measure the outcomes, but our output measures are suboptimal.
Crime problems in the designated hot spots. Our data contained information regarding the core crime problem in each hot spot. The police conducted these analyses, applying the SARA model to identify the core crime issues at the hot spots. Figure 2 displays the distribution of core problems in the hot spots (which constitute the outcome of this trial). As can be seen, antisocial behaviour and shop-related crimes were the most common problems. Note that these data were collected prior to the intervention.

Distribution of core crime problems in hot spots.
Sample Size
We took every hot spot that fit the threshold of 36 crimes in 12 months and was not more extensive than 150 m2, which amounted to 81 hot spots in the entire jurisdiction. A t-test power analysis for the total sample with alpha = 0.10, statistical power of 80%, and a one-tailed assumption 2 , revealed that our sample is large enough to detect effects at Cohen's d = 0.48. We note, however, that recently, Braga and Weisburd (2020) suggested that by implementing a natural logarithmic relative incidence rate ratio (log RIRR), the anticipated effect size of hot spot policing interventions is Cohen's d = 0.24. This means that in the current study, the necessary sample size should have been 103 in order to obtain statistically significant differences between treatment and no-treatment conditions Nevertheless, we used the entire population of units available for this experiment, and there were no additional units left that fit these inclusion criteria. Also, we used a block randomization protocol, which should reduce the variance between the participating units and increase the statistical power (Ariel & Farrington, 2012; Weisburd & Gill, 2014).
Randomization
As noted, we applied a block randomization protocol, with three statistical blocks based on the number of crimes in the pretrial period, on a 1:1 basis using simple randomization. No hot spots were removed ex-ante or ex-post randomization. These blocks comprised 36–50 crimes per year (Block A), 51–75 crimes per year (Block B), and 76 or more crimes per year (Block C) (Table 1). Thus, we treat every block as its own “mini experiment,” conducting a separate analysis for each while all conditions remain equal. Neither the assigned treatment nor the measurement was varied among the blocks.
We acknowledge that conducting a separate analysis for each block reduces the statistical power of the test, as the number of observations within each block is smaller than the total sample (Weisburd & Gill, 2014). However, by treating each block as its own mini experiment, we are effectively implementing a form of stratified analysis, which is a well-established method in experimental design (Cox & Reid, 2000). This approach allows for the potential variation across blocks to be handled in a straightforward manner, ensuring that the treatment effects are estimated as accurately as possible within each stratum (Rubin, 1974). If the blocks are significantly different in terms of the average number of crimes per year, treating each block as a separate experiment ensures that the specific characteristics of each block are appropriately addressed, by improving the ability to detect potential context-specific effects of the intervention over the control group. Analyzing them together could obscure these differences, leading to incorrect inferences (Senn, 1994).
Blinding. As noted, the participating officers were blind to the blocking procedure. They were not tasked to administer the interventions on a conditional basis on each block and were responsible for applying the same interventions across all treatment units homogenously. Furthermore, the length of this experiment—12 months—did not allow us to conceal the fact that officers were participating in an academic study. Finally, we blinded the participating police officers and their commanders to the control conditions, meaning that officers could not contaminate the non-treatment sites, whether consciously or unconsciously.
Statistical Methods
We applied a series of measurements to observe the treatment effect. First, we used descriptive statistics on the baseline characteristics of the participating hot spots within each statistical block. Then, we observed the treatment effect without the blocking to provide a general estimate of the intervention. Here, we applied analysis of covariance, with the post-treatment count as the outcome, the random allocation as a predictor, and the baseline counts as a covariate (Vickers & Altman, 2001). Specifically, given the over-dispersed count nature of our data, we used negative binomial regression. We then applied the same statistical approach within the high, medium, and low crime blocks.
Next, we took advantage of our output measures—GPS tracking of the foot patrol conducted by the officers—to estimate the effect of the “dosage” (Ariel et al., 2016; Sherman, 1990). We looked at the differences in the average number of minutes spent at the treatment hot spots versus the control hot spots and then the number of visits to the hot spots. Note that the correlation between the two dosage measures is strong (r = .6, 95% CI [0.4, 0.7], p = 5.2 × 10−8). Given the time data distribution, we used t-tests for these measures in each statistical block. We used one-tailed tests for all statistical significance calculations where we examined differences between treatment and control groups, as our hypotheses predicted less crime and more dosage in the treatment hot spots (Sherman & Weisburd, 1995). In these cases, one-tailed tests were appropriate because we had a priori directional expectations. For analyses where we did not have directional hypotheses, we used two-tailed tests. Across all analyses, the continuous predictors were mean-centered for ease of interpretation.
Data Availability Statement
The raw data used in this study are confidential police records and cannot be shared due to privacy and confidentiality concerns.
Results
Descriptive Statistics
Table 1 presents crime counts at each hot spot, broken down into statistical blocks (low, medium, and high crime levels), random assignment (treatment vs. control), and period (pre-treatment vs. post-treatment).
Treatment Effect Without the Blocking Criterion
The results in Table 2 show no detectable effect of the intervention on post-treatment crime counts (IRR = 1.0, p = 0.6). Marginal effect displays of the predicted post-treatment crime counts show that he predicted number of incidents in both the treatment and control hot spots is about 51.
Negative Binomial Regression of the Number of Crimes After the Treatment on Random Assignment and Number of Crimes Before the Treatment (N = 81).
Note. IRR = incidence rate ratio; CI = confidence interval; LR = likelihood ratio.
Treatment Effect with the Blocking Criterion
As shown in Table 3 and illustrated in Figure 3, there is considerable variation in the treatment effect between the blocks. In the medium (IRR = 0.8, p = 0.0) and high crime blocks (IRR = 0.8, p = 0.0), there are statistically significant differences between the conditions, with fewer crimes in the treatment hot spots compared to control hotspots. However, for the low crime block (IRR = 1.2, p = 1.0), the difference is not statistically significant, with fewer crimes at control hot spots.

Levels and impacts on crime counts, by crime-level block.
Negative Binomial Regression of the Number of Crimes After the Treatment on the Number of Crimes Before the Treatment and Random Assignment, by the Statistical Block.
Note. IRR = incidence rate ratio; CI = confidence interval; LR = likelihood ratio.
The surprising finding raises the question of what explains it. One possibility is treatment fidelity—whether officers in the treatment condition correctly followed their instructions. If they adhered to the protocol, they should visit treatment hot spots more frequently and spend more time there than in control hot spots. To test this, we used GPS data to track the number and timing of officer visits.
Tables 4 and 5 present the mean number of visits per day and the mean minutes per visit to hot spots by crime-level block (low, medium, and high) and condition (POP vs. saturated POP), respectively. In medium crime-level hot spots, the treatment group averages 1.4 more visits per day (p = 0.1) and 3.4 more minutes per visit (p = 0.0) compared to the control group. Similarly, in the high crime-level hot spots, the treatment group averages 2.0 more visits per day (p = 0.0) and 6.0 more minutes per visit (p = 0.0) compared to the control group. However, in low-crime hot spots, no significant differences were observed, and the direction contradicted expectations, with the control group averaging 1.3 more visits per day and 2.1 more minutes per visit than the treatment group.
Mean Number of Visits Per Day by the Statistical Block and Random Assignment.
Note. Treatment = Saturated POP condition; Control = POP-only condition df = degrees of freedom; CI = confidence interval.
Mean Number of Minutes Per Visit by the Statistical Block and Random Assignment.
Note. Treatment = Saturated POP condition; Control = POP-only condition; df = degrees of freedom; CI = confidence interval.
Finally, we evaluated how the delivered dosage aligned with the target of 15 minutes per visit for three visits per day by comparing prescribed targets with confidence intervals of the marginal means. In high-crime hot spots, officers exceeded the target for visits (M: 4.9, 95% CI [3.6, 6.3]) and met the time expectation (M: 13.7, 95% CI [6.2, 21.2]). In medium-crime hot spots, visits were consistent with the protocol (M: 4.7, 95% CI [2.8, 6.5]), but time per visit fell short (M: 9.4, 95% CI [5.8, 12.9]). Finally, in low-crime hot spots, both visits (M: 2.5, 95% CI [2.2, 2.8]) and time per visit (M: 5.6, 95% CI [4.4, 6.9]) were below target. These results suggest PCSOs adhered more closely to targets in higher-crime hot spots.
Discussion
Are two interventions applied within the same hot spot more effective in reducing crime than one of the interventions alone? Policing scholarship and place-based criminology seldom deal with the controlled and systematic application of two policing strategies for the same places. Under ideal experimental conditions, a four-arm trial would be conducted: POP only, saturated police presence only, saturated POP, and no treatment conditions. However, due to the limited number of eligible hot spots within the participating police jurisdiction, as well as the endogeneity problem (the difficulty of separating police presence from POP), 3 the optimal settings allow for a test of multiple interventions (saturation and POP) versus a single crime prevention policy (POP). These conditions were possible in Birmingham South, where the PCSOs ordinarily deliver business-as-usual POP at the community level, and we were able to increase the dosage of police presence at the treatment hot spots. The most straightforward and measurable outputs manifested this approach: more time and more visits to the treatment hot spots relative to control conditions, during which the PCSOs were asked to problem-solve.
Without the blocking criterion, that is, the categorization of hot spots into three groups based on their crime levels and then randomizing separately within each group, the results showed no significant overall variation between the two treatment conditions. However, further analyses revealed a beneficial effect concentrated in two out of three statistical blocks, whereas a backfiring effect was discovered in the third block. Curiously, the direction of the effects “follows” the stimulus: in whichever arm the additional patrol was delivered, the crime rates were lower. At face value, this is a puzzling result because the blocking was not shared with the police officers; the same crime problems emerged in all three blocks, although at varying frequency and officers were tasked with delivering the exact dosage and similar POP interventions as for the high- and middle-level blocks. Below, we discuss these findings.
Crime Policy Implications
What might explain the non-negligible counterproductive effect? Our primary response is empirical and touches on the fundamental concern of implementation. We assigned officers to visit the treatment hot spots at a more significant dose than the control hot spots in all blocks. However, the opposite effect occurred at the low-level hot spots: PCSOs delivered significantly more dosage in the POP-only conditions. Therefore, the outcome followed the intervention: wherever saturated POP occurred, crime rates went down. This is associated not with what we assigned but rather with what was delivered by the officers.
The backfiring effect of the protocol, in combination with the beneficial outcome of the delivered action, supports the proposed intervention. If we had seen a ubiquitous effect across all statistical blocks, regardless of the dosage delivered, the results would have been difficult to interpret. The exposure to the stimulus causes the outcome rather than the exposure to policy. When the intended programme (i.e., assignment) is significantly reduced in its practical application (i.e., delivery), an outcome in the hypothesized direction would have been illogical. By not delivering the intervention according to the protocol, the Birmingham South PCSOs may have deviated from their instructions, but they provided a more robust validation of the study's hypothesis. This might also explain the crime reduction in the control hot spots, which suggests potential influences beyond the applied treatment. While the precise cause remains uncertain, such variations could potentially stem from the discrepancy between the treatment protocol and the implementation carried out by PCSOs.
Much like other hot spot experiments, while we could obtain information regarding the whereabouts of officers, we do not have sufficient output data on the POP activities that officers conducted at the hot spots (the “how”), and as such, we could not obtain insight into the “black box” that is policing (Famega et al., 2017). Officers possess considerable discretionary authority over whom, how, and what to police, as becomes more evident as we delve deeper into what they do (or do not do). Whereas hot spot RCTs provide greater clarity, research on how remains limited. Highlighting this for the reader may yield a benefit (Dau, 2023).
Beyond the issue of treatment fidelity, this study highlights the effectiveness of “soft policing,” an approach that emphasizes non-coercive strategies and community engagement in crime reduction efforts (Burke, 2004). PCSOs exemplify this model; while they possess certain powers, such as issuing penalty notices, these are far more limited than those of fully warranted officers, who have powers of arrest and criminal investigation. Nevertheless, our findings revealed a reduction in crime rates in areas where PCSOs conducted interventions, suggesting that soft policing can be an effective method for lowering crime (Ariel et al., 2016). Practically, this indicates that integrating PCSOs into future policing strategies could enhance crime reduction efforts. Future research could further explore the relative effects of soft versus hard policing strategies to better understand their distinct impacts on crime reduction.
Research Implications
Future hot spot experiments (and, in fact, all social science field trials) should consider the potential for resistance among treatment providers to deliver interventions according to the trial protocol (Fixsen et al., 2013). We can only speculate why PCSOs are under-delivered in low-crime block treatment conditions relative to control conditions. One possibility is that they did not trust the protocol: Hot spots with 36–50 crimes per year—often fewer than one crime per week—may have been viewed by the officers as mistakenly identified, so they purposely and systematically did not proactively visit their allocated locations. Supporting this is our finding that PCSOs adhered less strictly to the protocol in low crime hotspots, which suggests that officers may have deprioritized locations they perceived as having less need for intervention. On the other hand, they continued to travel to the control group hot spots, unaware that these locations were part of the hot spot experiment. We find this explanation logical, given the evidence we have on some UK officers’ lack of support for GPS-led interventions and dislike of the routinization of their shifts under hot spot policing (Wain et al., 2017). Further research is needed, preferably with officer-level surveys on their proclivity for hot spot policing (Telep, 2017).
These issues are exacerbated in hot spot policing trials, but they cut across all prospective field experiments in the social sciences. While it may not be a new discovery (Baldassarri & Abascal, 2017), only a few scholars have attempted to overcome these concerns. ITT is the primary, and usually only, reported causal estimation model. Related to our study, if some hot spots were de facto misassigned to the wrong trial arm (i.e., treatment hot spots became non-treatment hot spots or vice versa), the treatment dosage deviated from the protocol (i.e., under or over-delivery of patrols), or hot spots were abandoned before their total exposure to the prescribed stimulus (Sherman & Weisburd, 1995), then such deviations would be disregarded from the causal estimates. This approach may be sensible in clinical trials because double-blinding is possible: Neither the treatment provider nor the participant is aware of what treatment they have received, treatment officers were unaware of what control officers were instructed to do, so misassignments, attrition, and errors in administrating the intervention are equally likely to occur in the treatment group as in the control group, due to the blinded, random allocation of units into the study arms (Zelen, 1979). But as we noted, clinical and field trials differ (Armitage, 1992; Raudenbush, 2007). The assumption of balanced errors in the treatment fidelity between treatment and control groups is a loose supposition when experiments occur in natural settings (Mills et al., 2013, 2019). In field trials, we rarely see conditions where “compliance with test treatment is all or none; everyone who does not take the test treatment takes the control treatment; the test treatment is not available to control-assigned individuals” (Sheiner, 2002, p. 208).
Given these considerations, what seems clear from our paper is that future field experiments, at least in criminology, should report both effectiveness and efficacy outcomes, and consider both design and statistical corrections whenever possible. There are options to overcome problems of misassignment, crossover, treatment fidelity, and other concerns with the delivery of the intervention as assigned by the experimental protocol (Ariel et al., 2022). Despite these issues, the traditional ITT model and the per-protocol analysis are—and should—remain the primary diagnostic model for controlled tests based on the CONSORT statement (Altman et al., 2001). Nevertheless, there is a growing recognition across multiple disciplines that, alongside the ITT results, we must report efficacy outcomes whenever possible, particularly when we do not know how exposure to stimulus behaves under the most controlled settings (Bland et al., 2023; Jagannathan & Camasso, 2005; Polit & Gillespie, 2010; Sagarin et al., 2014). In the case of policing experiments, mainly when our understanding of the efficacy of interventions in natural settings is limited (Famega et al., 2017), we call for the supplementary reporting of treatment-as-delivered results as well. Our study illustrates just that: without the treatment-as-delivered results, we would have missed the backfiring effect of the intervention in the low crime level hot spots that are directly linked to less treatment delivered than in control settings. One possible avenue for change is incorporating this item in the CONSORT statement requirements.
Furthermore, we call for future experiments to consider improved data collection on POP, community policing, and other law enforcement strategies. Expanding outcome measurement beyond conventional crime statistics is of paramount importance. Although crime rates are frequently employed as a measure, they fail to encompass the effects of policing endeavors, particularly in POP. Subsequent investigations may consider integrating a more extensive array of outcome indicators, encompassing community satisfaction and perceptions of safety, to evaluate the more significant ramifications of POP techniques. Surveys and community feedback systems are considered excellent instruments for gathering data, as they offer vital insights into the efficacy of POP activities in enhancing community well-being.
In addition, implementing standardized frameworks and indicators tailored to POP can significantly improve the consistency and comparability of measurements in various research investigations. As the systematic reviews of POP strongly indicate (Hinkle et al., 2020; Weisburd et al., 2010), researchers must collaborate with law enforcement organizations to produce mutually agreed upon outcome indicators specifically adapted to the goals of POP. Potential indications that might be utilized include quantifying recognized and resolved issues, the degree of community engagement in the problem-solving process, and the durability of problem-solving endeavors during a specific time frame. Standardization enables the comparison of POP's effectiveness across different jurisdictions and promotes its evaluation on a larger scope.
Finally, sophisticated data analytics and technology can enhance the assessment of outcomes in policing research. Law enforcement agencies are progressively adopting data-driven methodologies, such as geographic information systems (GIS), to improve their strategic decision-making processes (Ariel & Bland, 2019). We have the opportunity to analyze these datasets to inform of the delivery of these interventions, as we attempted to do in this study. This enables the identification of trends and patterns that may not be readily apparent using conventional crime statistics, thus offering significant contributions in assessing the efficacy of POP interventions and facilitating the prompt enhancement of methods. More attention to these data sources is warranted.
The current study has several limitations. First, this study did not examine the possibility of crime displacement (Cornish & Clarke, 1989). This concept suggests that following a successful intervention in a particular geographic area (such as a hot spot), crime may relocate to another location (Weisburd & Telep, 2014). While numerous studies have shown that crime does not simply “move around the corner” (Telep et al., 2014; Weisburd et al., 2006) and that the diffusion of intervention benefits is more likely (Johnson et al., 2014), displacement remains a concern within the scope of this study, warranting further investigation in future inquiries. Second, we could not obtain data on the implementation of the intervention. Therefore, we could not determine the extent to which the assigned POP interventions were executed in the treatment hotspots. This limitation, common in POP research (Hinkle et al., 2020), underscores the necessity for future studies to establish mechanisms for documenting police activities to enhance understanding of intervention effectiveness. Finally, we acknowledge that implementing more advanced analyses for the dosage data, such as analysis of variance, could offer deeper insights, for example into differences in delivered dosage across treatment hotspots by crime level blocks. However, this was beyond the scope of our study, and we believe that the current analysis effectively addresses the study's objectives while maintaining conciseness. Nonetheless, future research could benefit from employing different analytical techniques to achieve a more nuanced understanding of the findings.
Conclusion
Through a block-randomized controlled trial design, we evaluated the impact of saturated POP—comprising a combination of problem-solving activities and intensified police presence compared to standard POP levels—at the hot spot level. Over a 12-month period, our findings indicate a statistically significant effect in two out of three blocks. In the third block, the treatment appeared to counteract expectations, as the control group received a higher intensity of police presence than the experimental group. The findings suggest a pronounced treatment effect but, as importantly, illustrate the limits of applying a treatment-as-assigned model in the study of field trials.
Supplemental Material
sj-png-1-aje-10.1177_10982140241313452 - Supplemental material for Addressing the Treatment-as-Assigned Assumption in Field Experiments: Lessons Learned From the Birmingham South Saturated Problem-Oriented Policing Hot Spots Experiment
Supplemental material, sj-png-1-aje-10.1177_10982140241313452 for Addressing the Treatment-as-Assigned Assumption in Field Experiments: Lessons Learned From the Birmingham South Saturated Problem-Oriented Policing Hot Spots Experiment by Eran Itskovich, Esther Buchnik, Barak Ariel, Neil Wain, and Cristóbal Weinborn in American Journal of Evaluation
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
